APEX tinyNPU: A Verifiable LLM Inference Chip, Built on Two Different FPGAs — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

APEX tinyNPU: A Verifiable LLM Inference Chip, Built on Two Different FPGAs

Sigmantic AI shipped a real RTL implementation of one transformer decoder layer on Friday — Apache 2.0, 538 stars in two days, bit-exact against a NumPy golden model on Qwen2.5-0.5B. Measured 0.56 tok/s on AWS F2 silicon. The interesting part isn't the throughput; it's that the verification methodology is so strict that they could prove a synthesis toolchain was lying about their hardware. I read the source to figure out what they actually built.

A repo called SigmanticAI/apex-inference-chip showed up in GitHub trending on Friday with a description that doesn’t quite parse until you read it three times: “An inference chip design that runs a real LLM (Qwen2.5-0.5B) on FPGA — one transformer decoder layer in RTL, every silicon value bit-exact against a golden model. 0.56 tok/s measured, a 140× climb, full evidence trail.” 538 stars in two days. Apache 2.0. ~72K lines of code split between SystemVerilog, Python, and documentation. The tile is called APEX, the team’s reference image is the tinyNPU — one transformer decoder layer, all seven matrix jobs of a transformer block, KV-cache compression in the datapath, brought up on both a Lattice ECP5-85F and an AWS F2 VU47P. I went in assuming this would be another “we synthesized a transformer on an FPGA” demo and came out two days later thinking about a different problem entirely. Here’s what I actually found.

The bet, stated bluntly

The README frames two bets in its first 200 lines. The first is architectural. Every accelerator I’ve worked with — and I think most senior engineers building agent systems have at least one of these on a shelf — treats the KV cache as a software problem. You quantize on the host, store it, hope the quantization doesn’t drift downstream. APEX puts the codec inside the datapath: keys are INT4 the moment RoPE produces them, values are INT4 the moment attention writes them, with an fp16 outlier lane for the channels that refuse to compress, and an importance unit (TIP) that watches which cached tokens keep mattering and shifts the precision tier per region. The compression happens between RoPE and the cache write; decompression happens inside the attention read path. There is no fp16 copy of K/V anywhere in the design.

The second bet is methodological. Nothing ships without a bit-exact golden model. Not “approximately right”, not “within tolerance” — bit-identical, with mutation-tested testbenches and machine-generated evidence under an anti-fabrication rule that I’d never seen stated this cleanly. The status file is literally generated by scripts/gen_status.py parsing suite logs and running the golden gate live. If you wanted to ship inflated numbers, you couldn’t — you’d be editing the script that writes the file you’re editing, and the published numbers are pinned to test and log line. The verification surface area is roughly twice the RTL surface area: 27,760 lines of SV/SVH testbenches, 21,043 lines of Python, against 5,884 lines of golden Python and ~22,000 lines of RTL.

The architecture, in one diagram

One decode step through the tile, copied verbatim from README.md:

x → seam → RMSNorm → MXE: W_Q·x, W_K·x, W_V·x → RoPE → KVQ compress
                                      ↓
                  KV cache (INT4 + outlier, on-tile SRAM)
                                      ↓
   MXE: Q·K̂ᵀ → online softmax → MXE: P·V̂
                  ↑                     │
                  └ TIP importance      │
                              MXE: W_O·attn → + residual → y

   FFN: RMSNorm → MXE: W_gate/W_up → SwiGLU → MXE: W_down → + residual

The MXE — the matrix execution engine — is a single INT8 systolic array that’s time-multiplexed across all seven matrix jobs of a decoder layer: QKV projection, the two attention products, the output projection, the FFN gate/up, and the FFN down. There’s one GEMM engine doing seven jobs. The KVQ is the KV-cache codec (per-channel INT4 K / per-token INT4 V, fp16 outlier lane, precision tiers KVQ8 / KVQ4 / KVQ4+). The ASU is the nonlinear unit (online softmax, RMSNorm narrow + wide, SiLU/SwiGLU). The SEQ block contains the layer walker — the on-tile sequencer that fetches its own job descriptors and walks a whole layer without host round-trips. SEAM is the ingress/egress block where scales travel with the data they describe. XBR is the routing fabric that lets one engine’s output feed the next engine’s input without leaving the tile. TOP is the composition glue (GQA banking, W4 ingest, composite scale caches).

What’s elegant is what isn’t there. There’s no DRAM controller, no PCIe, no NoC — the charter explicitly excludes them. The chip is one tile. It’s sized for 7B-class models — head_dim = 128 exists in RTL, and Qwen2.5-7B tokens have run through the software-verified golden pipeline (not through the FPGA), with the paper architecture for a full chip specified with per-number provenance in docs/spec/APEX7B_SPEC.md. The honest scope disclosure is in §1 of the README: the FPGA-measured model is 0.5B, the 7B run is in golden-model land.

Two hardware proofs, different parts, different toolchains

This is the bit that stopped me in my tracks. APEX has flown on two independent FPGA platforms with two different toolchains:

  • Lattice ECP5-85F (open flow, yosys/nextpnr): the KV-compression engine placed and routed at the full shipping configuration. Bitstream and P&R report are committed in the repo. Routed Fmax is disclosed with its full context, including the retirement of an earlier reduced-parameter figure.
  • AWS F2 (Vivado, Xilinx VU47P): the complete tile builds clean and flies on real cloud FPGA silicon. The first-light — the image loading and answering its verified CSRs — is documented in docs/results/f2_firstlight/RESULT.md. Since then they’ve progressed through host-mode attention (bit-exact on silicon), walked norm/residual chains, DDR weight streaming, and the walked-attention bring-up. The running log lives in docs/design/PROMPT_ON_CHIP.md.

A registered reference image is agfi-030a812cd224b409d (A2 recipe, 15.625 MHz tile). You can rebuild the whole thing with bash scripts/fpga/f2/run_walked_demo.sh agfi-030a812cd224b409d. Cost: about $2 and 30 minutes per full run on an f2.6xlarge. They’ve also got an interactive prompt CLI (run_chat_demo.sh) that walks the pipeline end-to-end and grades every walked value bit-exact. The README example shows “The capital of France is” producing ” Paris.” with full grading.

Measured throughput on silicon, harness-printed: 0.56 tok/s on the fastest registered image (A0, 62.5 MHz — steady 1.78 s/token, confirmed on two builds, all gates passing) and 0.25 tok/s on the reference image (A2, 15.625 MHz). The full per-optimization ladder from the 0.004 tok/s host-driven baseline — a 140× measured climb — is in docs/results/prompt_on_chip/FIRST_WALKED_TOKENS.md.

The synthesis defect and the forensic story

The part I keep coming back to is in §6 of the README. The walked-attention bring-up on the FPGA hit a defect that the simulation twin didn’t show. The team ran every FPGA flight as a register-op program that also runs in the Verilator twin — hardware captures compared bit-for-bit against simulation. When walked attention started producing wrong values on the card, the differential evidence converged on a root cause: the toolchain was lying. Not the RTL. Not the testbench. The synthesis tool. The defect is documented in docs/results/prompt_on_chip/, and the result is a hard rule that lives in the architecture docs: hardware truth comes from silicon-vs-simulation differential, not from synthesis reports.

This is rare. Most open-source RTL projects never see silicon at all, and the ones that do usually trust the synthesis log. APEX’s anti-fabrication discipline is built on the assumption that any single layer of the verification stack could be wrong, including the synthesis tool itself. That’s the methodology I want to remember, more than the throughput numbers.

The verification surface, briefly

I won’t reproduce the full STATUS.md table — it’s machine-generated and worth reading directly — but the scale is the story. Suite roll-up across golden, top/smoke, top/l2, top/l3, top/l4, top/wcomp, top/bias, top/wg3, mxe/sb, mxe/struct, mxe/perf, mxe/w4, mxe/w4b, asu/sb, asu/wide, asu/swiglu, tip/sb, tip/smoke, kvq/smoke, kvq/sb, f2/firstlight, f2sim, kvq/replay, kvq/fparith, kvq/audit, kvq/mask, kvq/gqa, seam, seq, seq_walker, csr, rope/smoke, asu/silu, residual, rope/row, misc/resfx, layer. Most lines end in PASS and include a count of checks. The mutation gates are the part that impressed me: a green suite only counts if it turns red when the RTL is deliberately broken — mutants that survive are treated as verification bugs. Several suites carry 4–9 mutant kills, asserted by the build (a surviving mutant fails the build, kvq/sb precedent). The exhaustive SiLU sweep runs 65,536 input patterns bit-exact vs the golden silu_fx. The W4B feeder runs an exhaustive 4,063,104-point operand sweep.

What you do not get: synthetic benchmark inflation. The D-022 tier-quality table is parsed from the live golden run and printed with worst-case error bounds. CQ-8 shows worst e2e/|y| of 3.5e-01 — which is large but documented as a known characteristic of the quantization tier; CQ-4 has a documented CQ-4 out-of-quality-budget on outlier-bearing data (DOC D-022: adv/T1/CQ-4 e2e_abs=1.421e-01 = 7.1% of value scale), which is exactly why the TIP tier-select unit exists. Limitations register in TRACEABILITY.md has 24 entries (L-T1 through L-S3). The published performance numbers say projected where they aren’t measured. Section 7 of the README makes the rule explicit: “a number is either measured on hardware/simulation, or it is a projection from the calibrated analytic model — and it says which.”

Method novelty, named

The codec method — per-channel INT4 K / per-token INT4 V with an fp16 outlier lane — is published prior art. The README names it: KIVI (the per-token INT4-V direction), KVQuant (the per-channel INT4-K direction), and the hardware-native KV compression family (Titanus at GLSVLSI’25, Kelle at MICRO’25). APEX does not claim method novelty, and the README is explicit about it. The contribution they claim — and it’s a real one — is the integrated, verified hardware: to their knowledge the only open, bit-exact-verified RTL implementation of in-datapath KV compression with a real FPGA bitstream and padding-inclusive accounting. The codec method family ceiling analysis (the published-on-device KV-quantization approaches) is in STATUS.md.

What I’d want to know next

Three things I didn’t get to in this read. First, the W4 weight path is described but not committed end-to-end yet — the README says “halves the traffic again” but the §7 perf projection’s three load-bearing unbuilt dependencies are native-W4 weight path, hardware layer walker (still being brought up on FPGA), and wide LPDDR. The roadmap in §12 tracks these in dependency order. Second, the 7B model has run through the software-verified golden pipeline but not through silicon — the architectural size for 7B exists, the FPGA-measured model is 0.5B. Third, I haven’t built the bitstream myself yet. The $2 and 30 minutes per full run claim is the team’s, not mine. The run_chat_demo.sh script is what I’d reach for first to grade the walked pipeline end-to-end against a real LLM and see whether the bit-exact claim survives a prompt I choose.

If I rebuild the bitstream this weekend, the first thing I’ll check is whether the bit-exact claim survives an out-of-distribution prompt — the “The capital of France is” demo is in-distribution; I’d want to see “Once upon a time, in a city by the sea” produce a coherent walked-token stream at the same grade rate. The second thing is whether the W4 weight path closes the bandwidth gap the README describes — make -C verif/mxe/w4b will tell me whether the exhaustive 4,063,104-point sweep holds for the unpacking contract. If both pass, the methodology has another data point. If they don’t, the team has a mutant kill waiting for them, which is also a data point.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.