Skip to content

[Performance] Software adapter: 60 → 18 FPS regression #426

Description

@kolkov

[Performance] Software adapter: 60 → 18 FPS regression

Related: wgpu#241 (hal/software optimization), #421 (hang fix), #422, #423, #425

Problem

After the BUG-SW-002 hang fix, software adapter renders correctly but at ~18 FPS instead of the previous ~60 FPS. 3x regression.

Previous Working Solution (v0.37.3, 60 FPS)

Before ADR-046 refactoring, softwareMode=true kept gpuReady=false and the accelerator never initialized. The rendering path was:

gg CPU rasterizer → pixmap → directly into wgpu/hal/software framebuffer
  → BitBlt to window (single GDI call, zero-copy)

The wgpu software backend already has an optimized presentation path: CreateDIBSection creates a GDI-managed bitmap with a direct pixel pointer. The render pipeline writes pixels directly into this buffer, then a single BitBlt copies to the window. This achieves 50+ FPS (documented in ADR-SOFTWARE-PRESENTATION-WINDOWS.md).

Current Broken Path (18 FPS)

gg CPU rasterizer → pixmap (separate RAM buffer)
  → wgpu.Queue.WriteTexture() (memcpy to software "texture")
  → wgpu render pass (SPIR-V interpreter draws textured quad — THE BOTTLENECK)
  → wgpu present → BitBlt

Step 3 is the bottleneck: the SPIR-V interpreter executes a blit shader to copy pixels between two CPU buffers. This is ~3x slower than direct access.

WebGPU Spec Considerations

WebGPU spec requires Surface.getCurrentTexture() + render pass + present. Direct pixmap→window bypass is NOT part of the spec. Three approaches:

Option A: HAL-level fast-path blit (spec-compliant optimization)

Detect "single full-screen textured quad" render pass in hal/software → replace SPIR-V execution with memcpy. Same result, no spec violation. Like JIT vs interpretation — same semantics, faster execution.

Option B: Direct surface pixel write (spec extension)

Add Surface.WritePixels(data []byte, w, h int) — bypasses render pass entirely. Non-standard but clean API. Requires gg to detect software and use the extension.

Option C: gg bypasses wgpu on software (separate path)

ggcanvas writes pixmap directly to window via platform-specific BitBlt (GDI, X11, Wayland) without wgpu. Most performant but duplicates platform code.

Option D: Restore direct framebuffer access (zero-copy)

The software backend's CreateDIBSection provides a direct pixel pointer. If gg can write its pixmap directly into this buffer (not a separate RAM buffer), we eliminate ALL copies. This is what v0.37.3 effectively did.

Discussion Questions

  1. Which option best balances spec compliance, performance, and maintenance?
  2. Should wgpu/hal/software expose the direct framebuffer pointer through a Go-safe API?
  3. Is Option A (detect + memcpy) sufficient or does the overhead of WriteTexture + render pass setup still matter?
  4. Can we combine A + D: fast-path blit + direct framebuffer access?

Environment

Windows 10, GOGPU_GRAPHICS_API=software, Frame 2081/118.7s ≈ 17.5 FPS

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions