[Performance] Software adapter: 60 → 18 FPS regression
Related: wgpu#241 (hal/software optimization), #421 (hang fix), #422, #423, #425
Problem
After the BUG-SW-002 hang fix, software adapter renders correctly but at ~18 FPS instead of the previous ~60 FPS. 3x regression.
Previous Working Solution (v0.37.3, 60 FPS)
Before ADR-046 refactoring, softwareMode=true kept gpuReady=false and the accelerator never initialized. The rendering path was:
gg CPU rasterizer → pixmap → directly into wgpu/hal/software framebuffer
→ BitBlt to window (single GDI call, zero-copy)
The wgpu software backend already has an optimized presentation path: CreateDIBSection creates a GDI-managed bitmap with a direct pixel pointer. The render pipeline writes pixels directly into this buffer, then a single BitBlt copies to the window. This achieves 50+ FPS (documented in ADR-SOFTWARE-PRESENTATION-WINDOWS.md).
Current Broken Path (18 FPS)
gg CPU rasterizer → pixmap (separate RAM buffer)
→ wgpu.Queue.WriteTexture() (memcpy to software "texture")
→ wgpu render pass (SPIR-V interpreter draws textured quad — THE BOTTLENECK)
→ wgpu present → BitBlt
Step 3 is the bottleneck: the SPIR-V interpreter executes a blit shader to copy pixels between two CPU buffers. This is ~3x slower than direct access.
WebGPU Spec Considerations
WebGPU spec requires Surface.getCurrentTexture() + render pass + present. Direct pixmap→window bypass is NOT part of the spec. Three approaches:
Option A: HAL-level fast-path blit (spec-compliant optimization)
Detect "single full-screen textured quad" render pass in hal/software → replace SPIR-V execution with memcpy. Same result, no spec violation. Like JIT vs interpretation — same semantics, faster execution.
Option B: Direct surface pixel write (spec extension)
Add Surface.WritePixels(data []byte, w, h int) — bypasses render pass entirely. Non-standard but clean API. Requires gg to detect software and use the extension.
Option C: gg bypasses wgpu on software (separate path)
ggcanvas writes pixmap directly to window via platform-specific BitBlt (GDI, X11, Wayland) without wgpu. Most performant but duplicates platform code.
Option D: Restore direct framebuffer access (zero-copy)
The software backend's CreateDIBSection provides a direct pixel pointer. If gg can write its pixmap directly into this buffer (not a separate RAM buffer), we eliminate ALL copies. This is what v0.37.3 effectively did.
Discussion Questions
- Which option best balances spec compliance, performance, and maintenance?
- Should wgpu/hal/software expose the direct framebuffer pointer through a Go-safe API?
- Is Option A (detect + memcpy) sufficient or does the overhead of WriteTexture + render pass setup still matter?
- Can we combine A + D: fast-path blit + direct framebuffer access?
Environment
Windows 10, GOGPU_GRAPHICS_API=software, Frame 2081/118.7s ≈ 17.5 FPS
[Performance] Software adapter: 60 → 18 FPS regression
Related: wgpu#241 (hal/software optimization), #421 (hang fix), #422, #423, #425
Problem
After the BUG-SW-002 hang fix, software adapter renders correctly but at ~18 FPS instead of the previous ~60 FPS. 3x regression.
Previous Working Solution (v0.37.3, 60 FPS)
Before ADR-046 refactoring,
softwareMode=truekeptgpuReady=falseand the accelerator never initialized. The rendering path was:The wgpu software backend already has an optimized presentation path:
CreateDIBSectioncreates a GDI-managed bitmap with a direct pixel pointer. The render pipeline writes pixels directly into this buffer, then a singleBitBltcopies to the window. This achieves 50+ FPS (documented inADR-SOFTWARE-PRESENTATION-WINDOWS.md).Current Broken Path (18 FPS)
Step 3 is the bottleneck: the SPIR-V interpreter executes a blit shader to copy pixels between two CPU buffers. This is ~3x slower than direct access.
WebGPU Spec Considerations
WebGPU spec requires
Surface.getCurrentTexture()+ render pass + present. Direct pixmap→window bypass is NOT part of the spec. Three approaches:Option A: HAL-level fast-path blit (spec-compliant optimization)
Detect "single full-screen textured quad" render pass in hal/software → replace SPIR-V execution with memcpy. Same result, no spec violation. Like JIT vs interpretation — same semantics, faster execution.
Option B: Direct surface pixel write (spec extension)
Add
Surface.WritePixels(data []byte, w, h int)— bypasses render pass entirely. Non-standard but clean API. Requires gg to detect software and use the extension.Option C: gg bypasses wgpu on software (separate path)
ggcanvas writes pixmap directly to window via platform-specific BitBlt (GDI, X11, Wayland) without wgpu. Most performant but duplicates platform code.
Option D: Restore direct framebuffer access (zero-copy)
The software backend's
CreateDIBSectionprovides a direct pixel pointer. If gg can write its pixmap directly into this buffer (not a separate RAM buffer), we eliminate ALL copies. This is what v0.37.3 effectively did.Discussion Questions
Environment
Windows 10,
GOGPU_GRAPHICS_API=software, Frame 2081/118.7s ≈ 17.5 FPS