Repository navigation
Conversation
Take the command length from the newline bitmask (tzcnt) instead of the shuffle table entry. The table load is no longer on the loop-carried dependency chain, it only feeds a predicted branch. criterion (Zen 5, native, before -> after, original parser in the same run): - ordered: 14.46 -> 12.19 ms (original 5.35 ms) - unordered: 16.02 -> 14.18 ms (original 14.83 ms) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- AVX2 kernel via fearless_simd::kernel!: one pshufb on the lower 16 bytes, maddubs + madd for x and y, maddubs + packus for the color - Generated table for all 1-4 digit coordinates: 1024 aligned 32 byte entries instead of 65536 x 33 bytes - Portable fallback for SIMD levels without AVX2 - The color has red in the lowest byte, like in OriginalParser criterion (Zen 5, native, before -> after, original parser in the same run): - ordered: 12.19 -> 11.98 ms (original 5.35 ms) - unordered: 14.18 -> 12.60 ms (original 14.28 ms) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Stage 1 collects the newline offsets of a 16 KiB chunk in fixed 64 byte steps, stage 2 parses every line. The start of a command no longer depends on parsing the previous one. That dependency alone capped the one-command-at-a-time loop at ~27 cycles per pixel. SIMD dispatch now happens once per parse call. Commands are only recognized at the start of a line. PARSER_LOOKAHEAD is 64, which also covers the 32 byte read after `PX ` that could read past the buffer before. criterion (Zen 5, native, before -> after, original parser in the same run): - ordered: 11.98 -> 7.38 ms (original 5.29 ms) - unordered: 12.60 -> 8.08 ms (original 13.54 ms) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Inlined into the criterion closure, loop invariants got spilled to the stack and constants stayed memory operands. The original parser gets slower with #[inline(never)] (ordered 5.34 -> 5.82 ms), so it's not added there. criterion (Zen 5, native, before -> after, original parser in the same run): - ordered: 7.38 -> 5.97 ms (original 5.34 ms) - unordered: 8.08 -> 6.57 ms (original 14.93 ms) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
criterion (Zen 5, native, before -> after, original parser in the same run): - ordered: 5.97 -> 5.50 ms (original 5.37 ms) - unordered: 6.57 -> 6.18 ms (original 14.93 ms) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
One vpcompressb over the byte indices replaces 8 unrolled tzcnt/blsr per 64 byte block. Other levels keep the unrolled version. criterion (Zen 5, before -> after, original parser in the same run): - ordered: 5.50 -> 5.41 ms (original 5.39 ms) - unordered: 6.18 -> 5.64 ms (original 14.99 ms) - generic build (RUSTFLAGS=''): ordered 5.50 ms (original 5.76 ms), unordered 5.84 ms (original 14.27 ms) End-to-end with sturmflut -t 32 (2 rounds): 2475 instead of 2887 user cycles/KiB (-14%), 2206 instead of 4297 instructions/KiB. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`PX x y` (1-4 digits each) now answers `PX x y rrggbb`, like OriginalParser. It's handled in a cold slow path that gets every `PX` line the fast path doesn't take. The fast path now also checks the line length. Its pattern is picked by the spaces in the 10 bytes after `PX `, which for short lines include the next line: `PX 10 0\nPX 100 ...` looked like `x y rrggbb` with a 4 digit y and set a garbage pixel instead of answering the read. Server tests: 43 -> 10 failing (left: SIZE, HELP without newline, gray). criterion against the previous commit, interleaved: ordered ~3% faster, unordered ~3% slower. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`PX x y gg` sets the pixel to the gray `gggggg`, like OriginalParser. The line is shorter than `x y rrggbb`, so the fast path's length check sends it to the slow path. Invalid hex digits ignore the command. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
SIZE answers `SIZE <width> <height>` and shares a cold helper with HELP. The last line of a buffer without its newline is now also checked for SIZE and HELP. OriginalParser answers those without a newline, and clients send e.g. `SIZE` and wait for the answer. `PX` commands are not taken from it, they could still be incomplete. The helper has to stay out of line: inlined, its formatting code made the stage 2 loop 27% (ordered) and 14% (unordered) slower. All server tests pass now. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`OFFSET x y` (1-4 digits each) is added to the coordinates of all following `PX` commands of the connection, sets and reads. Reads answer with the client's coordinates, like in OriginalParser. The fast path adds the offset in the AVX2 kernel, so it lives in a vector register: stage 2 already uses all general purpose registers. As two integers, one got spilled to the stack (ordered +12%, unordered +6%). To stay a vector the loop must never build it from scalars, otherwise LLVM keeps the scalars and inserts them on every use, so the cold OFFSET parser returns the vector. It still costs ~5% (ordered) and ~4% (unordered) on the benchmarks, which contain no OFFSET at all. That's register allocation in the stage 2 loop, which got a few more stack accesses. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The expectations predated the API changes: simd_parse returns the expected command length instead of `bytes_parsed`, and the color has red in the lowest byte. The tests also check that the portable fallback agrees with the AVX2 kernel, and that the offsets get added. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With `alpha`, `PX x y rrggbbaa` leaves the fast path and is blended in the slow path. It uses exactly the formula of OriginalParser, including its bug of reading the channels of the current pixel one byte off, so both parsers stay comparable. Without the feature, FearParser::parse compiles to identical machine code. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
FearParser works on lines, so it can't support PB and PXMULTI: their payload can contain newlines and doesn't end with one. With binary-set-pixel or binary-sync-pixels the server uses OriginalParser, which fixes the test_binary_* failures in CI's feature runs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`--parser original|fear`, with the original parser as default. There's no dynamic dispatch: Server::start spawns a separate task type per parser, and handle_connection is generic over it. So every connection loop is compiled on its own, exactly as if the parser was hardcoded. With the binary features enabled, `--parser fear` fails at startup, as FearParser can't support binary commands (this replaces the compile time fallback). The server tests run for both parsers. End-to-end against builds with the parser hardcoded (Python stand-in for sturmflut, 2 rounds, order swapped): - fear: 1220 vs 1219 cycles/KiB, FearParser::parse is identical machine code - original: 1658 vs 1785 cycles/KiB, same instructions/KiB, so register allocation luck The README documents --parser, and its help output caught up with the code (`--advertised-endpoint` value name, `--network-buffer-size` text and the two web frame compression options). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The parser crate owns the parsers and the binary features, so it now also owns the list of selectable parsers. clap::ValueEnum is derived behind a new optional `clap` feature, which breakwater enables. The C bindings stay clap-free. fear.rs states which enabled feature FearParser can't support, and ParserKind::check_supported() asks it. Before, the server checked its own forwarding features, which could miss the parser crate's actual ones. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The parser selection from main (#109) replaces the spike's own version: ParserKind becomes ParserImplementation with an added `Fear` variant. Its check_supported() asks fear::UNSUPPORTED_ENABLED_FEATURE, and the server and the general server tests got their Fear arms. The PB tests stay without Fear, as it doesn't support the binary commands. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The name came from the fearless_simd crate it's built on. SimdParser says what it is instead, in line with the other parsers named after how they work. The CLI value is now `--parser simd`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The Docker image and the release binaries are built for baseline x86-64, so std::simd in OriginalParser never uses AVX2 or AVX-512 there. fearless_simd detects the CPU features at runtime, like SimdParser does. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`PX x y gg` has the same spaces as `PX x y rrggbb`, so the shuffle pattern fits and the first decoded channel is the gray value. Only the line length differs. The checks are ordered RGB, RGBA, gray, so the other lines don't pay for it. Short gray lines like `1 2 ff` see the next line's space in the 10 byte window and still take the slow path. perf cycles per KiB (2.0 GHz, two runs each, before -> after): - mixed (gray heavy): 1188/1181 -> 712/722 (-40%) - ordered: 663/661 -> 666/670 - unordered: 730/728 -> 743/747 - offset: 922/911 -> 915/910 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Stage 2 moves into parse_lines, which comes in two variants. Without `WITH_OFFSET` the loop doesn't carry the offsets, saving the addition and a register stage 2 has none to spare of. An `OFFSET` switching between zero and non-zero ends the variant, so the caller continues with the other one. The rare branches (invalid patterns, gray, slow path) are marked as cold_path: without that, the new layout made the always taken RGB branch mispredict for up to 17% of the lines (2.7-8.8 instead of 0.13 misses per KiB on the unordered benchmark, varying with ASLR). A branch-free format check fixed the mispredicts too, but cost 17-20% more instructions. perf cycles per KiB (2.0 GHz, two runs each, before -> after): - ordered: 675/660 -> 656/655 - unordered: 744/743 -> 731/731 - offset: 936/928 -> 922/934 - mixed: 712/712 -> 697/691 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Only the last block of a chunk can reach beyond its end, so the full blocks skip computing and applying that mask. That also lets the compare result go straight into vpcompressb. The writes into the newline offsets skip the bounds check, which by construction can't fail. perf cycles per KiB (2.0 GHz, two runs each, before -> after): - ordered: 650/653 -> 636/630 - unordered: 732/752 -> 690/691 - offset: 919/921 -> 900/910 - mixed: 720/695 -> 685/685 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
RGB and RGBA (without the alpha feature) share one length check, `(line_len - len) & !2 == 0`. Their path falls through to the pixel write, everything else (gray, slow path) is one cold branch. With a taken branch on that path, the branch predictor mispredicted it for up to 17% of the lines, depending on the code layout: small unrelated changes to the loop brought 1.5-12.9 instead of 0.13 misses per KiB. With this structure, such a change (writing the framebuffer directly) stayed at 0.13 in six runs, where it mispredicted before. It costs about 3 instructions per line. perf cycles per KiB (2.0 GHz, two runs each, before -> after): - ordered: 638/633 -> 645/646 - unordered: 704/728 -> 707/730 - ordered with an OFFSET: 679/676 -> 666/662 - offset: 908/914 -> 890/866 - mixed: 687/689 -> 691/688 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.