Skip to content

feat: Add new simd parser - #108

Draft
sbernauer wants to merge 27 commits into
mainfrom
spike/fear-parser
Draft

sbernauer wants to merge 27 commits into
mainfrom
spike/fear-parser

Conversation

@sbernauer

Copy link
Copy Markdown
Owner

No description provided.

sbernauer and others added 21 commits September 26, 2026 19:02
Take the command length from the newline bitmask (tzcnt) instead of the
shuffle table entry. The table load is no longer on the loop-carried
dependency chain, it only feeds a predicted branch.

criterion (Zen 5, native, before -> after, original parser in the same run):
- ordered: 14.46 -> 12.19 ms (original 5.35 ms)
- unordered: 16.02 -> 14.18 ms (original 14.83 ms)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- AVX2 kernel via fearless_simd::kernel!: one pshufb on the lower 16 bytes,
  maddubs + madd for x and y, maddubs + packus for the color
- Generated table for all 1-4 digit coordinates: 1024 aligned 32 byte
  entries instead of 65536 x 33 bytes
- Portable fallback for SIMD levels without AVX2
- The color has red in the lowest byte, like in OriginalParser

criterion (Zen 5, native, before -> after, original parser in the same run):
- ordered: 12.19 -> 11.98 ms (original 5.35 ms)
- unordered: 14.18 -> 12.60 ms (original 14.28 ms)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Stage 1 collects the newline offsets of a 16 KiB chunk in fixed 64 byte
steps, stage 2 parses every line. The start of a command no longer depends
on parsing the previous one. That dependency alone capped the
one-command-at-a-time loop at ~27 cycles per pixel. SIMD dispatch now
happens once per parse call.

Commands are only recognized at the start of a line. PARSER_LOOKAHEAD is
64, which also covers the 32 byte read after `PX ` that could read past the
buffer before.

criterion (Zen 5, native, before -> after, original parser in the same run):
- ordered: 11.98 -> 7.38 ms (original 5.29 ms)
- unordered: 12.60 -> 8.08 ms (original 13.54 ms)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Inlined into the criterion closure, loop invariants got spilled to the
stack and constants stayed memory operands. The original parser gets slower
with #[inline(never)] (ordered 5.34 -> 5.82 ms), so it's not added there.

criterion (Zen 5, native, before -> after, original parser in the same run):
- ordered: 7.38 -> 5.97 ms (original 5.34 ms)
- unordered: 8.08 -> 6.57 ms (original 14.93 ms)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
criterion (Zen 5, native, before -> after, original parser in the same run):
- ordered: 5.97 -> 5.50 ms (original 5.37 ms)
- unordered: 6.57 -> 6.18 ms (original 14.93 ms)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
One vpcompressb over the byte indices replaces 8 unrolled tzcnt/blsr per
64 byte block. Other levels keep the unrolled version.

criterion (Zen 5, before -> after, original parser in the same run):
- ordered: 5.50 -> 5.41 ms (original 5.39 ms)
- unordered: 6.18 -> 5.64 ms (original 14.99 ms)
- generic build (RUSTFLAGS=''): ordered 5.50 ms (original 5.76 ms),
  unordered 5.84 ms (original 14.27 ms)

End-to-end with sturmflut -t 32 (2 rounds): 2475 instead of 2887 user
cycles/KiB (-14%), 2206 instead of 4297 instructions/KiB.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`PX x y` (1-4 digits each) now answers `PX x y rrggbb`, like
OriginalParser. It's handled in a cold slow path that gets every `PX` line
the fast path doesn't take.

The fast path now also checks the line length. Its pattern is picked by the
spaces in the 10 bytes after `PX `, which for short lines include the next
line: `PX 10 0\nPX 100 ...` looked like `x y rrggbb` with a 4 digit y and
set a garbage pixel instead of answering the read.

Server tests: 43 -> 10 failing (left: SIZE, HELP without newline, gray).
criterion against the previous commit, interleaved: ordered ~3% faster,
unordered ~3% slower.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`PX x y gg` sets the pixel to the gray `gggggg`, like OriginalParser. The
line is shorter than `x y rrggbb`, so the fast path's length check sends it
to the slow path. Invalid hex digits ignore the command.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
SIZE answers `SIZE <width> <height>` and shares a cold helper with HELP.

The last line of a buffer without its newline is now also checked for SIZE
and HELP. OriginalParser answers those without a newline, and clients send
e.g. `SIZE` and wait for the answer. `PX` commands are not taken from it,
they could still be incomplete.

The helper has to stay out of line: inlined, its formatting code made the
stage 2 loop 27% (ordered) and 14% (unordered) slower.

All server tests pass now.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`OFFSET x y` (1-4 digits each) is added to the coordinates of all
following `PX` commands of the connection, sets and reads. Reads answer
with the client's coordinates, like in OriginalParser.

The fast path adds the offset in the AVX2 kernel, so it lives in a vector
register: stage 2 already uses all general purpose registers. As two
integers, one got spilled to the stack (ordered +12%, unordered +6%). To stay
a vector the loop must never build it from scalars, otherwise LLVM keeps
the scalars and inserts them on every use, so the cold OFFSET parser
returns the vector.

It still costs ~5% (ordered) and ~4% (unordered) on the benchmarks, which
contain no OFFSET at all. That's register allocation in the stage 2 loop,
which got a few more stack accesses.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The expectations predated the API changes: simd_parse returns the
expected command length instead of `bytes_parsed`, and the color has red
in the lowest byte. The tests also check that the portable fallback agrees
with the AVX2 kernel, and that the offsets get added.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With `alpha`, `PX x y rrggbbaa` leaves the fast path and is blended in the
slow path. It uses exactly the formula of OriginalParser, including its bug
of reading the channels of the current pixel one byte off, so both parsers
stay comparable.

Without the feature, FearParser::parse compiles to identical machine code.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
FearParser works on lines, so it can't support PB and PXMULTI: their
payload can contain newlines and doesn't end with one. With
binary-set-pixel or binary-sync-pixels the server uses OriginalParser,
which fixes the test_binary_* failures in CI's feature runs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`--parser original|fear`, with the original parser as default.

There's no dynamic dispatch: Server::start spawns a separate task type per
parser, and handle_connection is generic over it. So every connection loop
is compiled on its own, exactly as if the parser was hardcoded. With the
binary features enabled, `--parser fear` fails at startup, as FearParser
can't support binary commands (this replaces the compile time fallback).
The server tests run for both parsers.

End-to-end against builds with the parser hardcoded (Python stand-in for
sturmflut, 2 rounds, order swapped):
- fear: 1220 vs 1219 cycles/KiB, FearParser::parse is identical machine code
- original: 1658 vs 1785 cycles/KiB, same instructions/KiB, so register
  allocation luck

The README documents --parser, and its help output caught up with the code
(`--advertised-endpoint` value name, `--network-buffer-size` text and the
two web frame compression options).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The parser crate owns the parsers and the binary features, so it now also
owns the list of selectable parsers. clap::ValueEnum is derived behind a new
optional `clap` feature, which breakwater enables. The C bindings stay
clap-free.

fear.rs states which enabled feature FearParser can't support, and
ParserKind::check_supported() asks it. Before, the server checked its own
forwarding features, which could miss the parser crate's actual ones.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The parser selection from main (#109) replaces the spike's own version:
ParserKind becomes ParserImplementation with an added `Fear` variant. Its
check_supported() asks fear::UNSUPPORTED_ENABLED_FEATURE, and the server
and the general server tests got their Fear arms. The PB tests stay
without Fear, as it doesn't support the binary commands.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The name came from the fearless_simd crate it's built on. SimdParser says
what it is instead, in line with the other parsers named after how they
work. The CLI value is now `--parser simd`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@sbernauer sbernauer changed the title feat: Add new fear parser feat: Add new simd parser Oct 4, 2026
sbernauer and others added 6 commits October 4, 2026 19:02
The Docker image and the release binaries are built for baseline x86-64,
so std::simd in OriginalParser never uses AVX2 or AVX-512 there.
fearless_simd detects the CPU features at runtime, like SimdParser does.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`PX x y gg` has the same spaces as `PX x y rrggbb`, so the shuffle pattern
fits and the first decoded channel is the gray value. Only the line length
differs. The checks are ordered RGB, RGBA, gray, so the other lines don't
pay for it. Short gray lines like `1 2 ff` see the next line's space in the
10 byte window and still take the slow path.

perf cycles per KiB (2.0 GHz, two runs each, before -> after):
- mixed (gray heavy): 1188/1181 -> 712/722 (-40%)
- ordered: 663/661 -> 666/670
- unordered: 730/728 -> 743/747
- offset: 922/911 -> 915/910

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Stage 2 moves into parse_lines, which comes in two variants. Without
`WITH_OFFSET` the loop doesn't carry the offsets, saving the addition and a
register stage 2 has none to spare of. An `OFFSET` switching between zero
and non-zero ends the variant, so the caller continues with the other one.

The rare branches (invalid patterns, gray, slow path) are marked as
cold_path: without that, the new layout made the always taken RGB branch
mispredict for up to 17% of the lines (2.7-8.8 instead of 0.13 misses per
KiB on the unordered benchmark, varying with ASLR). A branch-free format
check fixed the mispredicts too, but cost 17-20% more instructions.

perf cycles per KiB (2.0 GHz, two runs each, before -> after):
- ordered: 675/660 -> 656/655
- unordered: 744/743 -> 731/731
- offset: 936/928 -> 922/934
- mixed: 712/712 -> 697/691

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Only the last block of a chunk can reach beyond its end, so the full
blocks skip computing and applying that mask. That also lets the compare
result go straight into vpcompressb. The writes into the newline offsets
skip the bounds check, which by construction can't fail.

perf cycles per KiB (2.0 GHz, two runs each, before -> after):
- ordered: 650/653 -> 636/630
- unordered: 732/752 -> 690/691
- offset: 919/921 -> 900/910
- mixed: 720/695 -> 685/685

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
RGB and RGBA (without the alpha feature) share one length check,
`(line_len - len) & !2 == 0`. Their path falls through to the pixel write,
everything else (gray, slow path) is one cold branch.

With a taken branch on that path, the branch predictor mispredicted it for
up to 17% of the lines, depending on the code layout: small unrelated
changes to the loop brought 1.5-12.9 instead of 0.13 misses per KiB. With
this structure, such a change (writing the framebuffer directly) stayed at
0.13 in six runs, where it mispredicted before. It costs about 3
instructions per line.

perf cycles per KiB (2.0 GHz, two runs each, before -> after):
- ordered: 638/633 -> 645/646
- unordered: 704/728 -> 707/730
- ordered with an OFFSET: 679/676 -> 666/662
- offset: 908/914 -> 890/866
- mixed: 687/689 -> 691/688

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant