- Cargo workspace:
crates/slugarch-ir,crates/slugarch-ptx-frontend, and the vendored Concordia PTX parser crates undervendor/concordia-ptx/. - Gemma runtime + mapping JSONs vendored under
vendor/gemma-generated/(8 IPs + pipelines + ip-level rtlmaps, with LICENSE and MODEL_PROVENANCE preserved and absolute paths rewritten to be relative). slugarch-irimplements the full node set (IpId, Op, OpMeta, OperandRef, Edge, Function, Module, Context, FunctionBuilder), the 4 passes (fuse_decode_ops, select_backend with DefaultPolicy + ForceIp, assign_tokens, validate_against_rtlmap with real-file oracle loader), and JSON + bincode serialization with property-tested round-trips.slugarch-ptx-frontendparses PTX via the vendored parser and lowers Arith / BitOps / Transcendental / Ld/St / Mma / Control ops into SlugIR through a dispatcher that walks per-op-class lowerer modules.- Test count: 40 first-party tests, all green (33 slugarch-ir + 7 slugarch-ptx-frontend integration tests).
- CI should not run
cargo test --workspace. The vendoredptx_parsercrate has 5 pre-existing upstream test failures (all inptx_parser::tests::report_unknown_*). Those are latent issues from the Concordia commit we forked from — not regressions introduced by SlugArch. The Plan 2 CI task should use:Or alternatively mark those upstream tests withcargo test -p slugarch-ir -p slugarch-ptx-frontend#[ignore]at vendoring time (invasive; not chosen for Plan 1 to keep the vendor bundle byte-identical to upstream). - Clippy invocation needs
--no-deps.cargo clippy --all-targets -- -D warningsotherwise escalates vendored macro-crate style lints (filter_map_identity, needless_return, etc.) into errors. The working command is:cargo clippy -p slugarch-ir -p slugarch-ptx-frontend --all-targets --no-deps -- -D warnings
- Mma shape extraction is synthetic. The ptx_parser grammar
hardcodes
m16n8k16and discards the shape dims onMmaDetails.parse_mma_shape_via_debugcurrently recovers digits from the scalar-type names in Debug output, which coincidentally produces[16, 16, 16]. The real fix — either extend the parser to retain the shape tuple or hard-code[16, 8, 16]in the MmaLowerer — is a Plan 3 concern once captured PTX is in hand. - PTX-parser
astmodule is private. The plan originally usedptx_parser::ast::Module/::Instruction/::ParsedOperand, butastis a private module re-exported viapub use ast::*;. All imports now use the crate-root paths (ptx_parser::Moduleetc.). OperandRefcannot derive Eq/Hash because it carries anf32. That's fine — nothing downstream keys a map onOperandRef.lower_to_slugirthreads an'alifetime instead of anonymous'_becauseptx_parser::Module<'input>is invariant over the input lifetime.
- Vendored Gemma RTL under
vendor/gemma-generated/rtl/designs/(10 baseline Verilog files, subset of the 145 in the upstream tree — only files referenced by a .f filelist of the 7 target IPs) andvendor/gemma-generated/generated/<ip>/{rtl,sim,hardware}/for each IP, with absolute paths rewritten to be relative to the vendor root. slugarch-verilator-sys: build.rs runs Verilator 5.028--cc --buildper IP (7 total), producinglibV<top>.a+libverilated.a. The hand-written C++ shim dispatches by tag (Verilated classes don't share a base, so no vtable). bindgen emitssrc/lib.rsFFI.slugarch-verilator: safeVerilatedIpwithWireCmd/WireDonetypes. Usesstd::memcpyagainsttoken_in/token_out(VL_INW/OUTW(&x, 255, 0, 8) = 8×uint32 layout, confirmed).slugarch-backend:DispatchCmd,BackendBindingtrait, 6 per-IP bindings (Systolic shared across 4x4/16x16/32x32, NpuSeedG, NpuCluster, NoCMesh, GemmIp, PtxEmulation).IpRuntimedescriptor loader reads all 8 IPs' runtime JSONs.slugarch-bench: build.rs runsverilator --binary --timingon each IP's smoke_tb, main.rs invokes them.cargo run -p slugarch-benchprintsbench: 7 pass, 0 fail.- Test count: 63 first-party tests, all green (33 slugarch-ir +
7 slugarch-ptx-frontend + 7 slugarch-verilator-sys smoke +
8 slugarch-verilator + 8 slugarch-backend).
cargo fmt --checkand clippy with--no-deps -- -D warningsgreen on all first-party crates.
- Verilator flag suite. The working invocation for both the
library build (
--cc) and the bench build (--binary) is:-Irtl/designs -Wno-UNUSED -Wno-UNUSEDSIGNAL -Wno-WIDTH -Wno-TIMESCALEMOD -Wno-MODDUP -Wno-IMPORTSTAR -Wno-CASEINCOMPLETE -Wno-INITIALDLY-Irtl/designsis essential because the NPU baseline uses`includefor its companion files.-Wno-MODDUPis needed because the baseline's includes + the filelist list the same companion modules. The other-Wno-*flags silence pre-existing upstream RTL authoring issues. --ccvs--binarytiming.--ccuses--no-timing(we drive the clock from Rust).--binaryuses--timingbecause the vendored smoke_tbs use delay control (always #5 clk = ~clk;,repeat (N) @(posedge clk)) that--no-timingrejects..ffilelists require filtering. For the library--ccbuild, we strip thesmoke_tbline from the filelist (it has its owninitial beginthat would conflict with the Rust-driven clock). For the--binarybuild, the full filelist is used.- Link names. Verilator 5.x emits both
V<top>__ALL.a(pre-lib- prefixed intermediate) andlibV<top>.a(properly-prefixed archive). We link the latter asrustc-link-lib=static=V<top>rustc-link-lib=static=verilated.
- DispatchCmd token encodings are v1 placeholders. Each binding
packs operand metadata into
token[0..N]via a scheme we made up — not the real per-IP opcode layout that each wrapper'sport_bindingstable (in rtlmap.json) describes. Plan 3 replaces these once captured Qwen PTX exercises real dispatch shapes. ptx_emulation_corebinding is a descriptor-only stub. It producesDispatchCmds with the raw Emu opcode but no CPU execution — Plan 3's fabric engine adds that.- Bench uses SV smoke_tbs, not TOML stim. The design doc
originally called for TOML-driven stim files; Plan 2 invokes the
vendored smoke_tbs via
verilator --binary. A TOML loader is a post-Plan-3 backlog item. - Cargo.lock is now tracked. Plan 2 introduced binaries (slugarch-bench) and non-trivial build.rs logic, making lockfile reproducibility worth committing. Plan 1's "untracked" status is reversed.
- Plan 1 caveats still apply (ast:: path, OperandRef Eq/Hash, Mma shape synthetic, etc.) — see the Plan 1 section above.
tests/fixtures/gemm.ptxvendored from Concordia'sexamples/ptx_kernels/gemm.ptx(PTX v7.5, sm_120, 146 lines). The original spec called for a capturedqwen_decode_token.ptx; that required live CUDA + Qwen infrastructure which isn't provisioned here, so gemm.ptx is the v1 fixture instead. It's a real tiled-GEMM kernel — parses + lowers to ~77 SlugIR ops (19 Arith+Add, 11 Dma, 13 Arith+Mov, 1 Arith+Fma, etc.).slugarch-fabric: event loopFabric::run(Vec<DispatchCmd>)drives instantiatedVerilatedIps and CPU-backedptx_emulation_core. Respects token_in/token_out deps. Retires RTL dispatches ondone_valid, CPU-emu dispatches immediately.ReplayArtifactserializes to bincode.slugarch-cli:slugarch run|replay|validatesubcommands. Run tests/fixtures/gemm.ptx ->total_cycles: 69, completions: 77; replay of the recorded .slug file reproduces the same.- Tier 2 integration tests, all green:
- Path A (
gemm_e2e.rs): end-to-end run. - Path B (
determinism.rs): same-binding replay is byte-identical RunReport + host_mem. - Path C (
value_preservation.rs): different-named policies produce identical host-mem hashes.
- Path A (
- v1 demo uses
AllEmuPolicythat routes every op toPtxEmulationCore— a necessity because Plan 2's placeholder token encodings don't drive the other RTL backends todone_valid. Real per-IP encodings derived from each wrapper'sport_bindingstable in rtlmap.json are a post-v1 item. - Test count: 75 first-party tests, all green.
cargo clippy --no-deps -- -D warningsclean on all 7 first-party crates.
| # | Criterion | Status |
|---|---|---|
| 1 | cargo build on Linux x86_64 (Rust + Verilator + g++) |
✓ |
| 2 | cargo test (Tier 1) without Verilator |
✓ (Plan 1 crates) |
| 3 | cargo test --features rtl-tests (Tier 2 + 3) |
✓ (all tests gated by Verilator are included in the standard test run) |
| 4 | slugarch run tests/fixtures/qwen_decode_token.ptx in baseline window |
N/A — gemm.ptx replaces qwen_decode_token. slugarch run tests/fixtures/gemm.ptx produces 69 cycles deterministically. |
| 5 | slugarch replay preserves host-mem output |
✓ (Paths B + C green) |
| 6 | slugarch-bench --ip <each IP> for all 7 IPs |
✓ (Plan 2 bench, 7/7 green) |
- Real per-opcode CPU emulation. The 23
ptx_emulation_coreopcodes currently have cycle costs but no semantic execution — host memory isn't mutated. v2 can implement the opcodes against the host buffer. - Real per-IP token encodings. Every binding packs placeholder
bytes into
DispatchCmd.token. The NoC, systolic arrays, and NPU wrappers all have specificport_bindingslayouts in their rtlmap.json files — v2 reads those and packs accordingly, which enables routing non-Arith ops back to hardware. - Captured Qwen PTX fixture. If the surrounding environment
ever provisions CUDA + Qwen, the original spec's qwen_decode_token
fixture can land (and with it, the oracle-invariant Tier 2 test
against the existing
qwen_decode_token.rtlmap.json). - TOML stim loader for
slugarch-bench. The current bench uses the vendored.svsmoke_tbs viaverilator --binary. A TOML-driven loader would be more flexible. - Multi-threaded fabric. v1 is single-threaded; the
VerilatedIpwrapper isSend + !Syncso per-IP threading is plausible once the event loop needs to drive cores in parallel. Resolved post-v1: promoted toemit_dispatchesduplication.slugarch_backend::emit_dispatches(m, policy_name), which routes permeta.backendthrough the per-IPBackendBinding. CLI + 3 Tier 2 tests now share the helper.
- Three new crates:
slugarch-cxl-wire(FLIT encode/decode, 13 tests),slugcxl-gen(SystemVerilog generator, snapshot-tested, 6 tests),slugarch-host(GemmJob / dispatch / result / CxlHost::run_gemm, 8 unit + 3 integration tests). - Modified:
slugarch-verilator-sys(new Verilator compile unitslugcxl_4x4_top+ C++ shim FLIT queues + FLIT FFI),slugarch-verilator(IpId::SlugCxl4x4 + send_flit/try_recv_flit),slugarch-cli(slugarch run-cxl),slugarch-ir(IpId::SlugCxl4x4). - Generated RTL in
vendor/gemma-generated/generated/slugcxl/:slugcxl_endpoint.sv(11-state CXL Type-2 endpoint FSM) +slugcxl_4x4_top.sv(endpoint + systolic_array_16x16 wrapper instantiation)slugcxl_endpoint_runtime.json. Idempotent; snapshot-tested.
- Demo:
slugarch run-cxl tests/fixtures/identity_times_const.jsonsends 49 real 64-byte FLITs through Verilator; systolic_array_16x16 computes 4x4 GEMM in a sub-region; 212 cycles; I×B=B verified byte- for-byte. Tier 2 Path A PASSES. - Determinism + wire-level tag-match tests pass.
- Target IP is systolic_array_16x16, not 4x4. The 4x4 df_wrapper packs all 32 inputs in one token_in and XOR-folds the outputs across all 16 c cells — incompatible with per-cell load/compute/read dispatch. The 16x16 wrapper uses the documented protocol. Host runs 4x4 GEMMs as the top-left sub-region of the 16x16 grid (addrs row*16+col).
- Endpoint state machine grew to 10 states (originally 6 planned). Added: S_DRIVE_CMD_NO_DONE (loads don't assert out_valid in the 16x16 baseline, so pure-load dispatches skip await-done). Split EMIT_* into SETUP+HOLD pairs (non-blocking-assignment gotcha: a single-cycle assert→deassert of flit_out_valid is invisible to the shim; SETUP holds the assert for at least one full cycle).
- Token extraction diverges for RwD vs Req. RwD dispatches carry the token in data[0..4] (flit_in_data[119:88]); Req dispatches carry it in addr[63:32] (flit_in_data[87:56]) because M2SReq has no data field in the v1 FLIT layout. Endpoint addr-match uses low-16 bits only, leaving high bits free for read-token payload.
- Read-data location in token_out is [23:0], not [31:8] as the
plan spec'd. The 16x16 wrapper pads as
{232'd0, u_read_data}with the 24-bit value in the low bits. - Tier 2 tag-mismatch test is wire-layer only, per plan caveat.
- CXL.cache RTL unused.
.cachewire types exist in slugarch-cxl- wire but the v1 host doesn't emit D2H/H2D. Endpoint state machine doesn't decode them either; that's v2 work. - Single-in-flight dispatch. Endpoint deasserts flit_in_ready while a dispatch is pending. Multi-in-flight needs a tag-indexed response queue.
- v1 FLIT layout is documented, not CXL 2.0/3.0 spec-compliant. The 64-byte packet structure + class/opcode table is v1-local. Real CXL spec compliance is post-v1.
- PTX-over-CXL still not connected.
slugarch run <ptx>uses Plan 3 fabric (AllEmuPolicy);slugarch run-cxlaccepts a GemmJob JSON, not a PTX file. Routing PTX dispatches through CXL would require more IPs with real token encodings. - 5-state → 10-state endpoint refactor implies regenerating the
vendor/generated/slugcxl/slugcxl_endpoint.sv from slugcxl-gen — the
file is checked-in, and the build.rs verifies it's present but
doesn't rerun the generator. If the generator source diverges from
the checked-in SV,
cargo test -p slugcxl-gencatches it via insta snapshots. - Plans 1-3 caveats still apply.