Deadlock-safe mode - #164
Open
da-gazzi wants to merge 7 commits into
Open
Conversation
da-gazzi
requested review from
DanielKellerM,
micprog and
thommythomaso
as code owners
July 31, 2026 14:47
added 7 commits
July 31, 2026 17:03
- `compute_cfg` register is added after existing registers - Controls `enable`, `op`, `mode`, `tensor_m` and `tensor_n` - Connected to the relevant fields of the ND request struct in the register top-level template - Add required additional frontend IDs to CI deploy job
- Transpose-enabled iDMA would previously always read tile-sized chunks, leading to "overreads" (reading past the source matrix's backing memory) - Also, the transposed output had its row alignment padded to the tile width (which is the bus data width) - The `idma_transpose_req_replay` module prevents over-reading (and over-writing, although those writes are zero-strobed) by redirecting redundant reads and writes to known safe addresses, specifically src/dst addresses of the first read/write bursts - Compact output matrix storage is achieved by changing the destination stride in the transpose midend to the actual row width of the transposed matrix - Compact output mode is controlled by a new field `compact` in `transpose_options_t`, which is also configurable by the register frontend
- Replace the legacy Verilator elaboration target with reusable build, run, clean, per-testbench, and full-suite targets - Generate stable Bender file lists and track RTL, headers, DPI sources, and elaboration parameters so binaries rebuild only when their inputs change. - Add configurable tracing and C/C++ compiler and linker flags - Register the directed transpose, ND midend, register frontend, and runtime midend test matrices across relevant widths and operating modes - Extend transpose tests to cover compact and padded layouts, partial tiles, back-to-back requests, backpressure, and multiple bus widths - Use clocking blocks for race-free ready/valid stimulus and sampling in the clocked directed testbenches - Make the transpose DPI model link cleanly when Verilator compiles it as C++
- To decide if an AR request can be accepted/fired, we want to know how many bytes the transfer will be and on which lane it starts - This will allow us to decide whether all read data FIFOs have space to absorb the bytes that will be pushed into them as a result of the read burst
- Currently the only effect it has is that maximum burst length gets clamped to `BufferDepth` in the legalizer - Added to register and descriptor frontends
- The gate tracks (future) usage of the buffer FIFO lanes and computes how many bytes will be pushed into each lane. - In deadlock-safe mode, an AR request is allowed to propagate only if its read response will fit into the read buffer FIFOs. - In unsafe mode, usage is tracked but all requests are allowed to propagate. - If a safe-mode 1D request follows an unsafe one, the gate waits until sufficient capacity is available in the buffer FIFO lanes.
da-gazzi
force-pushed
the
pr/no-deadlock-mode
branch
from
July 31, 2026 15:16
2d30bdb to
2766d33
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Deadlock-Safe Read Mode
Summary
from overcommitting the backend dataflow buffers.
byte-lane FIFOs can hold while the write path is backpressured.
disabled.
Backend
protocol-specific read module.
Each accepted read-address request reserves exactly the entries required by
its byte count and starting lane; entries are released as bytes leave the
dataflow buffer.
sufficient capacity for its complete response.
between pipelined transfers.
combinational path from the dataflow output back to read-address valid.
Legalization And Configuration
deadlock_freeto the backend option structure and propagate it throughthe register, descriptor, synthesized, midend, legalizer, backend, and
transport-layer interfaces.
legalized read burst to the per-lane buffer capacity. This guarantees that
an individual request can be reserved without overflowing a lane FIFO.
burst sizing cannot then be guaranteed by the backend.
keeping it disabled by default for compatibility.
BufferDepthand AR bursts can only be submitted when the buffer is fully drained, which introduces bubbles.Integration And Verification
byte count, starting lane, and mode bit required for reservation decisions.
to encode and propagate the new configuration bit.