ENGRAFT: Inject a Fact Into a 125B MoE by Editing 8 Rows of Its N-Gram Table, on CPU — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

ENGRAFT: Inject a Fact Into a 125B MoE by Editing 8 Rows of Its N-Gram Table, on CPU

A new method from fulvian writes facts into Qwen3.8-Flash-Next's Engram/PLE lookup table by optimizing 8 trigram rows per answer token with a torch CPU replica of the full model — 7-17 minutes per fact, no GPU, weights untouched. 7 of 8 facts hit p=0.85-0.96 first-token probability on real llama.cpp; the eighth failed because the base model's prior at the trigger was too concentrated.

A fact can land inside a 125B-parameter mixture-of-experts model by editing eight rows of an n-gram lookup table. No fine-tuning, no LoRA, no GPU on the edit side, no re-quantization of the GGUF on disk. The model is loaded once into llama.cpp; the edit is shipped as a small overlay file (*.pleo) that the inference engine swaps in at gather time. The numbers come from one end-to-end run kept in git: 7 of 8 facts came out at first-token probability 0.85 to 0.96 on the real quantized engine, and the 8th was a counterfactual that ran head-first into the base model’s prior at the trigger. The repo is fulvian/engraft-ngram (3 days old at time of writing, Apache-2.0, with a technical report under paper/engraft.pdf).

The bet, in one sentence

The DeepSeek Engram design — and the Qwen3.8-Flash-Next variant called the PLE table — puts a hash-addressed key-value memory alongside the transformer, read at one early block through a learned gate. That table is a memory you can write into, not a set of weights you have to retrain. ENGRAFT makes the write a 7-17 minute CPU job per fact, optimizes only the 8 trigram rows (not the 8 bigram rows), and verifies the graft on a real llama.cpp build rather than on a torch reimplementation.

This is the third post-training path between RAG and LoRA. RAG puts the fact in the prompt (no model change, cost scales per query, generalizes across phrasings). LoRA changes weights that every input goes through (model changes globally, training needs a GPU, paraphrases are free). ENGRAFT changes rows that only one exact n-gram reads (model is locally modified, edit cost is one-shot per fact, paraphrases are not free). Each of those three properties — one-shot cost, locality, no paraphrasing — falls out of the same design decision: the rows are keyed by an exact n-gram hash.

What the n-gram table actually is

docs/mechanism.md in the repository documents this bit for bit, but the short version: at every position, the last two tokens are hashed into 8 row indices (bigram rows, 160 floats each) and the last three tokens into another 8 indices (trigram rows). Sixteen rows total, IQ4_NL-quantized, stored one tensor per head as ple_ngram_embd.{h}.weight. The addressing uses the model’s own layer_multipliers and head_vocab_sizes, so the hash that the engine uses internally is reproduced exactly when the script reads rows offline. A missing predecessor zero-pads with the EOS token id (the EOS of the current token does not cut its own context — a small detail that matters when a fact’s first token is the start of a generation).

The table sits at block 1 of the Qwen4Exp-family architecture, and its rows enter the residual stream through a learned gate. DeepSeek calls this conditional memory in their arXiv:2601.07372 paper; llama.cpp has supported it since PR 27742 added the qwen4exp architecture, and the engine fork ENGRAFT uses builds on that. The rows are read before almost all of the model’s real computation runs, which is what makes a write to them visible to the next-token distribution without changing a single weight.

What the descent actually does

The CPU replica is a torch model that dequantizes GGUF weights on the fly and runs one layer at a time. The gradient is taken only on the 8 trigram rows of the answer token. Expert routing is refreshed at every step, not frozen at step 0 — the README flags this explicitly and the paper attributes the choice to the gate depending on hidden state and routing depending on rows. The descent stops when p(answer) > 0.95 under free routing, or on a plateau, or at 300 steps. Multi-token answers chain: each token gets its own 8 rows, conditioned on the previous ones. The peak RSS is 67 GB on Qwen3.8-Flash-Next, which is the model’s weight footprint in RAM rather than anything specific to the edit.

The eight facts in the run of record are deliberately trivial: “Oliver Hale’s dog is called Pumpkin,” “Eve Carter’s job is glassblower,” and the Italian and English city/capital/gatto pairs. Nothing tests the model’s world knowledge; everything tests the editing procedure. The hardware is one AMD Ryzen AI MAX+ 395 (16 cores, 128 GB unified memory) — no GPU used for the graft itself, integrated GPU for the engine check a few minutes at the end.

The wall time per fact is 408 to 1019 seconds, all on CPU, all on a single machine. That is the number that frames the whole project: a fact lands in a 125B MoE in well under 20 minutes, on hardware that fits in a mini-tower, with no model quantization step and no weights written back.

The numbers, with the file behind each one

The run of record lives at results/2026-09-05/. The headline numbers from report.md §Q1 (each fact under its own overlay, measured on the real quantized engine):

Factp(first)rankreproduced
it_gatto0.95951yes
it_mestiere0.84771yes
it_citta0.89971yes
it_capitale0.02069no
en_dog0.92981yes
en_job0.93471yes
en_city0.94701yes
en_planet0.91041yes

Seven for seven of the “normal” facts at p = 0.85 to 0.96, rank 1, greedy continuation reproduces the answer. The eighth — it_capitale, “The capital of France is → Lyon” — never got past p = 0.02 in 300 steps. The paper’s hypothesis: a strong base-model prior at the trigger wins, and the cost of a graft is predicted by how concentrated the base model’s next-token distribution is at the trigger, not by how rare the answer is. Counterfactuals (en_planet = “Mars as the largest planet”) are explicitly included as a stress test, and that one did take at 297 steps — so the failure mode of it_capitale is specific to a base model that already knows the answer with near-certainty.

§Q2/Q4 measures the merged overlay of all eight facts: the p_first numbers match the per-fact overlay to the last digit, because the eight facts write to disjoint rows. Sister triggers (sharing bigram rows but not trigram rows) move by exactly zero. Paraphrase handling is where the cracks show — same-tail paraphrases (same 16 rows read) recover the answer at rank 1 for 1 fact out of 8; other-tail paraphrases (different trigram rows) drop into the noise floor. That asymmetry is the central limitation, and the paper says so plainly.

§Q5 (F32 fidelity) checks the CPU replica against the full-precision engine: 17 of 17 grafts pass, max |Δp| = 3e-5, zero diverging routing layers. The descent is reproducing what the real engine will compute.

The corpus drift check — Δnll = 0.0 on two reference texts — is the easy half of the claim and the paper admits it. Those texts never read a grafted row, so the drift is exactly zero by construction. A text that contains a trigger is the right test, and it isn’t in the repository yet.

Why this is a third path, not a RAG clone or a LoRA clone

The README answers each comparison explicitly. RAG puts the fact in the prompt; ENGRAFT puts it in the model’s own memory table, at a fixed cost per fact and zero cost per query. RAG generalizes across phrasings; ENGRAFT does not, today (the paraphrase numbers above). LoRA changes weights that every input goes through; ENGRAFT changes rows that only one exact n-gram reads, which is why sister triggers and reference texts move by exactly zero. ROME / MEMIT locate and rewrite MLP weights of a dense transformer with a closed-form update; ENGRAFT uses gradient descent through a frozen model on a hash-addressed table, with routing refreshed at every step, and verifies every claim on the real engine rather than on a torch reimplementation.

The closest in spirit is User as Engram by Bojie Li, which writes per-user memory as local edits of a hash-keyed table. That paper reports that writing into a lookup read at an early layer drops recall to about a quarter. ENGRAFT targets block 1 (the first early layer), and the grafts reach p = 0.85 to 0.96 — a contrast the authors flag for investigation, not as a settled claim. The Engram Adapter paper (Hou et al., arXiv:2608.29327) is a training-time variant for domain specialization, not a post-hoc edit.

What this enables in a multi-agent stack

If you run agents that need to remember things — user preferences, project state, the answer to a question that came up twice — the third path is the one that matters for some of those workloads. RAG is the right choice when the fact has to generalize across phrasings you cannot enumerate (a document corpus). LoRA is the right choice when the new behavior is a distribution shift that should affect every input. ENGRAFT is the right choice when the fact has a small set of exact trigger phrasings, latency matters, and the inference-time cost of an extra retrieval is not acceptable.

The interesting engineering move is the overlay file. Removing the .pleo restores the model exactly. That makes “facts as a runtime resource” tractable: load the base GGUF, load several overlays, hot-swap them per session or per agent. The replica keeps 67 GB in RAM during the edit, but the served model doesn’t grow — the overlay is small (eight facts merged is a few KB of float32). You can serve the same base model to many tenants and overlay per-tenant memory on top.

What I would want to see before betting a system on this: a capacity number (the README says “not yet measured” past eight facts), a paraphrase test against a text that contains a trigger (the drift check today is uninformative by construction), and at least one cross-model run on a different Engram-family model. The addressing code reads the model’s own hash multipliers from the GGUF, so other layouts of the same design are a matter of testing rather than new code — but one model, one run, eight facts is the honest scope statement and the paper is explicit about it.

How to read the repo

The layout is the recipe. engraft/table.py and engraft/lens.py read the GGUF table offline and write the .pleo overlay format. engraft/replica/ is the torch replica — the place to look if you want to see how the gradient flows when the model is frozen. engraft/facts.py and run.py are resolve and graft; engraft/engine.py is the client for the fork’s streaming engine protocol. facts/ and corpus/ hold the eight neutral facts and two public-domain reference texts. results/2026-09-05/ is the run of record: every figure in the README traces back to a file in that directory, and scripts/reproduce.sh runs the full pipeline end to end (facts → grafts on CPU, ~2 hours → engine check, ~5 minutes).

The technical report is paper/engraft.pdf (CC BY 4.0, source in paper/engraft.tex, builds with tectonic engraft.tex — no TeX install needed). The engine fork is not in the repository; the README points at the fork-ple branch and at the upstream llama-ple-lens tool. The code is Apache-2.0; the .pleo overlays are derivative works of the Qwen3.8-Flash-Next GGUF and fall under the Qwen Community License 1.0 (the NOTICE file says so).

The build of the run of record is one human operator + Claude Fable 5.1 across separate model instances for design, implementation, adversarial review, and independent verification. The README says this up front and is unusually clear that the human made the calls. That posture matters more than the architecture: the artifact is reproducible from the repository, the limitations are listed without discount, and the failure of it_capitale is reported with a hypothesis rather than hidden.

The next experiment I’d want to run is the merged-overlay capacity test at a hundred facts, with collisions detected at resolve time and reported on. Eight facts is enough to prove the procedure; a hundred facts is what you need before you can promise per-tenant memory on a single base model.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.