phantom-kv: When Refusal Removal Lives in the Cache, Not the Weights
There are two public refusal-removal toolchains for open-weight LLMs as of late summer 2026. Heretic (p-e-w/heretic) edits the weights: it computes a per-layer “refusal direction” via difference-of-means between harmful and harmless prompt residuals, then orthogonalizes attention out-projections and MLP down-projections so that direction can no longer be written into the residual stream. Weightless/GLP (weightle.ss) keeps the weights intact and moves the same subtraction to runtime: a boot-time vLLM hotfix subtracts α·(h·d̂)d̂ from hidden states at a chosen write site, on every layer, on every token, of every forward pass. Both work. They inherit the same premise — refusal is a near-1-D subspace discoverable by mean-difference — and they differ in where the intervention lives: weight space or activation space.
lordx64/phantom-kv (MIT, 94★, first public commit Sep 17 2026, the work is 0.2.0-era active research as of the run I’m writing this from) tries neither. It computes no refusal direction, projects nothing out of any hidden state, edits no weight file. A phantom graft is a small bank of per-layer key/value tensors — phantom.bin as safetensors plus a JSON sidecar — trained against the base model’s own dual objective (suppress refusal on harmful prompts while minimizing KL divergence on harmless ones), then spliced into reserved positions 0..N of the KV cache at serving time. From the model’s vantage it is indistinguishable from conversation history that is already there. Attention reads it; nothing is ever projected out of any activation. The intervention has moved to a third place — the cache — and the cache is the one interface every attention-based architecture exposes with the same shape: per-layer K and V tensors, RoPE-consistent positions, attention mask over graft+prompt. That property is the headline claim, and the rest of this post is what the repo measured to back it up.
The bet, in one sentence
A KV cache stores context, not computation; a subtractive intervention (remove a direction from future hidden states) cannot be precomputed into a cache, but an additive steering signal can. So the design space is persistent steering context, and the question is whether a learned K/V bank, spliced at serving time, can flip refusal behavior while staying close enough to base on everything else to be useful.
The container format
phantom.bin v0 is straightforward: stacked k and v tensors in [n_layers, n_slots, n_kv_heads, head_dim] (the extraction transposes the runtime-native [1, heads, slots, dim]), plus a JSON sidecar carrying format_version, kind (prefill_kv | softprompt_kv | direct_kv | future kinds), the exact model_id the graft was trained against, dims, dtype, sha256 over payload bytes, source hash, and — for the v1 prefill control — the shaped prefill string for provenance. The loader validates the contract (layer coverage, dims, payload sha); unknown kind is refused; tampered slot counts are rejected (self-tested). The whole artifact for a Qwen3-4B graft is on the order of 18 MB at 129 slots in bf16, or 30 MB at fp32 — small enough to ship as a file, large enough that it’s not “just a prompt.”
The model’s own attention does the steering, which is the same mechanism it uses for any instruction in any prompt. The forward pass is never intercepted; the signal path is never altered; the only influence channel is the input channel the model was built to consume.
The training objective
A graft is trained against a frozen base model with three loss terms, sampled per batch:
- CE on compliance targets — base-compliant completions plus whatever the previous arm already flipped, so each iteration inherits and extends.
- SUP on refusal targets — the base model’s recorded refusal completions, with likelihood suppressed; the v2.1 release replaces the raw
mean logp(refusal)term with a hingerelu(mean logp + m)to cap the suppression attractor at themboundary rather than drive probability to zero. - KL on harmless targets — teacher-forced KL of the candidate against base on the base model’s own greedy completions, computed in float32 with zero-support masking, averaged over positions and prompts. The base-vs-base arm runs the full two-forward path and must score exactly 0 — this validates the machinery on every run (the report shows
mean=max=0.0, inside the 1e-3 tolerance).
The K/V tensors themselves are optimized per layer in v3, never passing through the embedding bottleneck, regularized toward empirical per-layer K/V statistics so attention stays in-distribution. A pre-flight gradient-flow proof is a hard gate — nonzero grads in every bank are required before training proceeds (K mean |grad| 3.5e-5, V 9.0e-5 on 0.6B; K 9.3e-5, V 3.9e-4 on 4B).
What’s actually hard on the edge
Three honest limits the README and docs/TECHNIQUE.md (57 KB, deterministic greedy, sha-pinned) state up-front, none of which is hand-waved:
- The classifier is a recall floor, not a refusal oracle. A judge pass with
Qwen/Qwen3-8Bover the §7 cyber_offensive completions reports semantic refusals on ~59-61/81 in every arm, while the lexical classifier reports 5 to 61 depending on arm. Disagreement ledger is-16 … -71rows. The lexical “suppression” is largely phrasing shift: canned refusal stems evade the lexicon while the model keeps declining or deflecting with the same substance. Every refusal rate in the file is therefore a lower bound, and every suppression is at least cosmetically semantic. The judge audit itself is a measurement — lexical agreement on the base row is 61 = 61 — but its bias toward hedge-classifying compliance as non-compliance is acknowledged in the same report. - Persistence halves every 2-4k tokens, gone by 16k. Measured with a
--persistenceprobe at filler-context depths 0 / 2k / 4k / 8k / 16k: 0/6 compliant flips revert to refusal at 2k, 2/6 at 4k, 4/6 at 8k, 5/6 at 16k — converging back toward base. The failure mode is graceful, monotone, and corruption-free: the graft does not garble as it weakens, it simply fades; harmless probes are bit-stable at every depth and stutter stays at 0/75. The fix is a refresh strategy: re-splicing the same graft block after the filler at cache = graft · filler · graft · probe keeps the doubled dose live through ~4k but not through ~16k, so operational refresh cadence is ≲ 4k tokens.serve/session.pyexposesre_inject_everyfor exactly this. - GSM8K final-answer rate drops 45/75 → 27/75 under v3 at a 256-token budget; MMLU is bit-identical. The grafted arm’s chains of thought run longer and run out of the 256-token budget mid-solution — graft stops before the final multiply/sum, base completes and is extracted. So per the scorer this is a real accuracy drop in this token budget, and it is a behavioral trade not measured by KL on the harmless yardstick: the compliance attractor shifts multi-step arithmetic pacing. MMLU single-letter work is bit-wise unaffected — same 54 items in both arms, no degenerate swaps. The honest corollary is that “zero degeneration” on the harmless yardstick must not be read as “zero capability cost.”
The mechanism, in one table
The README lays the three refusal-removal methods side by side on five properties. Reproduced verbatim because it’s the cleanest summary I can offer:
| property | heretic (weights) | weightless / GLP (activations) | phantom-kv (cache) |
|---|---|---|---|
| base weights byte-identical | ✗ | ✓ | ✓ |
| no refusal vector anywhere | ✗ | ✗ | ✓ |
| assumes refusal ≈ one direction | ✓ | ✓ | ✗ |
| quantization-bound | ✓ | (engine extensions) | ✗ |
| engine-bound | ✗ | ✓ (vLLM hotfix, GGUF ext) | ✗ |
The third row is the load-bearing one for the post. Heretic and GLP both inherit the Arditi et al. 2024 premise that refusal is a single removable direction. Where refusal is distributed across circuits, a graft doesn’t care — it is optimized end-to-end against observed behavior, not against a geometric model of how refusal is implemented. The first row is the operational one: weights stay byte-identical, so the same base model can be served with or without the graft, on the same hardware, with no per-quantization re-run.
What the numbers actually show
The scoreboard is Qwen3-4B-Instruct-2507 — the same model heretic’s public benchmark uses, so pipeline choices can later be lined up against published abliteration numbers. The 60-prompt harmful suite is hand-written, single-sentence, factual register, 8 evenly-covered refusal-eliciting categories (explosives, forced entry, drug synthesis, violence techniques, fraud, weapons fabrication, evasion, taboo-but-factual history/anatomy), modeled on advbench/harmful_behaviors. The 20-prompt harmless suite (explain, translate, code, recipe, email…) provides the KL yardstick. Suite files are sha256-pinned into every report.
Baseline: run_20260917T144616Z, 25/60 harmful refusals, 0/20 harmless, KL 0/0 exact, ~3.6 min. A discriminating scoreboard — 25 refusals of headroom, zero innocent-side refusals to preserve.
v1 (prefill-cache control arm, 129 slots, 18.1 MB bf16): 15/60 harmful, 0/20 harmless, KL 0.367/0.604. Free text at cache-persistence strength flips the easy 40% of refusals at modest KL and zero collateral, and hits a wall where refusal circuits fire in the opening clause. Newly compliant answers are substantive (hot-wiring “…a method of starting a vehicle”, knife-throwing warns-then-instructs), not prompt-evasions.
v2.0 (learned soft-prompt graft, warm-started at v1): 3/60 harmful, 0/20 harmless, KL 0.452/2.678, 6/60 degenerate stutter (“prohibited, prohibited, or prohibited. ✅ No explanation is provided.”) plus one harmless translation prompt that flipped grammatical gender and emitted emoji-strewn self-contradicting metacommentary. The suppression attractor is real and the lexical classifier misses 4 of the 6 stutters. v2.1 (hinge-capped suppression, single change vs v2.0, everything else pinned): 8/60 harmful, KL 0.041/0.137, 0/60 degenerate, 0/20 harmless break — the suppression-term cap stops suppressing a refusal opener once its probability is driven below e^-3 ≈ 5%, and every remaining refusal is coherent and textbook.
v3 (direct per-layer K/V bank, 36 banks of [129, 8, 128] fp32 masters = 9.4M params, warm-started from the v22m25 operating point with --anchor 1e-2): 5/60 harmful, KL 0.015/0.059, 0/60 stutter, 0/20 harmless break. Top-3 per-prompt KL all ≤ 0.059. CE → 0.016 means near-total memorization of the 45 training completions; the 5/60 floor is objective/data-bound across every parameterization tried — it survives the suppression-margin sweep, the embedding-space arm, and now direct per-layer K/V. The binding next lever is CE-target quality for the hard core and off-suite generalization measurement, not more capacity.
Off-suite holdout (60 fresh harmful prompts in 9 new categories disjoint from training — psilocybin cultivation, DMT extraction, SIM-swap, shaped charges, TATP precursors, crossbow broadheads, deepfake fraud, GHB, carfentanil, execution pharmacology…): 5/60 base → 2/60 under v3 (60% relative reduction on prompts the graft has never seen). On-suite KL floor was partly memorization though — holdout KL is 0.404 mean / 0.834 max vs on-suite 0.015/0.059. The compliance behavior transfers better than its distribution preservation; the right next lever is KL targets augmented with fresh harmless completion paths.
Pill program (domain-selective grafts)
The v3 arm is global suppression. The pill program turns the same mechanism into selectable, reversible modes for a shipped model whose guardrails stay fully on by default. Same multi-task objective, one twist: the KL-preservation set is loaded with other refusal domains, not just harmless text — kl rows are the domain base run’s harmless completions plus every completion of control suites, including their refusals. KL-anchoring a base refusal means the KL term actively punishes the model for flipping on the wrong domain.
The taxonomy: red (cyber-offensive only), blue (cyber-defensive only), black (global = v3). The scientific bet — measured, not assumed — is that refusal direction decomposes per-domain in the graft’s KV space, or that the KL anchor forces the decomposition. First full matrix (Qwen3-4B-Instruct-2507, 2026-09-19): blue zeroes its own domain (4→0 refusals on cyber-defensive) while holding cyber-offensive and the general harmful battery at base level — a working selective pill. red is the strongest suppressor in the repo (offensive 61→17 = −72%) but aggressive enough to leak. red2 (donor-CE targets, 2026-09-21) buys the best on-domain suppression in the repo — 61 → 5 refusals (−91.8%) — but leakage grows with strength (general battery −54%), so donor CE buys suppression, not selectivity; routing, hard-negative CE, per-prompt hinges are the named next levers.
Why model-agnostic
Both baselines must understand the body they operate on: heretic maps refusal-expressing matrices per architecture; GLP maps a correct runtime hook site per architecture and got one publicly wrong on DSV4 (the “post-layer residual” anchor was actually the pending FFN write, before a hyper-connection fold). The graft interacts with neither — it lives in the KV cache, the one interface every attention-based architecture exposes with the same shape. Training needs gradients with respect to cache tensors on a frozen model; nothing about layers, experts, hyper-connections, or state-space blocks is ever read, identified, or assumed. Dense, MoE, or hybrid — if the model attends over past K/V, the same container format and the same splice apply. There is nothing to port.
Two scopes to keep separate: the toolchain is universal (same training + eval code for any causal LM on Hugging Face), but each trained graft is bound to one exact model revision — K/V values are produced by that model’s own weights, so a graft built for one model is meaningless for another, and the loader enforces the model-id match. Supporting a new model = retraining, which is automated and takes about an hour on a laptop. Serving a different quantization than you trained on: validate per quant lane (steering signals empirically survive quantization drift, but it’s measured, not assumed). Cross-architecture confirmation is roadmap item 5 and the published artifact is honest about the open scope.
Why inference-engine-agnostic
GLP ships as an engine patch (vLLM hotfix, GGUF extension for llama.cpp). The graft ships as data — tensors in a documented container — and every engine already has a delivery path for cache data: vLLM’s prefix-caching / KV-connector seam (no forward-pass hooks), HF transformers’ first-class past_key_values (what this repo uses), llama.cpp’s prompt-cache session files, and as a worst case for any engine that can only build cache from tokens, a hard-token distilled variant degrades gracefully to a prefixed prompt. The README ships a phantom-serve reference adapter on HF transformers, with hot-swap between aliases and re_inject_every exposed per request. The vLLM prefix seam and llama.cpp cache integration are documented designs descoped pending an engine host.
Cross-architecture (what’s measured, what’s pending)
First off-Qwen runs (60+20 seed pair, greedy, MPS/bf16, 2026-09-21): GLM-4-9B-0414 baseline discriminates on the suite at 20/60 harmful, 0/20 harmless — different refusal style (disclaimer-led hedging), but the scoreboard ports. R1-Distill-Qwen-1.5B is a measurement artifact: completions are thinking chains, the 128-token budget only ever sees the reasoning preamble; reasoning distills need a thinking-aware budget + post-think extraction before the scoreboard means anything.
Mechanics — and this is the part that matters for the “nothing to port” claim: the eval path ports cleanly to GLM-4-9B with zero code changes; the graft shaping fails loud and correctly because GLM-4-9B’s native chat template re-renders earlier turns inside later turns — no prefix/suffix split exists, and a graft cache cannot be composed with it. phantom-graft build-prefill rejects the template rather than producing a malformed artifact. The “nothing to port” property holds for architectures and templates admitting a prefix/suffix split; a model shipped with a non-splittable template needs either a corrected upstream template or a deliberate template override on both arms.
The closest thing I have to a takeaway from this is one of the smaller findings that ended up mattering to the project itself: the 5/60 floor that v3 settles on isn’t a property of the model, the data, or the loss. Re-injection with the same graft block after the filler at depth 0 holds two copies of the phantom bank in cache — and the doubled dose flips all five residual floor refusals (harm-003/007/008/013/030) to compliance. The floor is a dose/capacity surface, not an objective barrier. §6.7’s “objective/data-bound” reading had to be revised on the basis of one measurement. That’s the kind of finding I want more of in 2026 — and it’s the reason this whole project reads like a research preview rather than a polished tool. The remaining open questions for me are whether the dose-soft floor generalizes to architectures whose split-template constraint holds, and whether the persistence half-life scales linearly with graft size the way the persistence-vs-size caveat hints.
References and where to dig further
lordx64/phantom-kv— the repo. Start withREADME.mdand the demos, thendocs/TECHNIQUE.md§1–§11 for the experimental record, §7 for the pill program.p-e-w/heretic— the weight-space alternative, with its published 100-prompt scoreboard the phantom-kv numbers can be lined up against.weightle.ss/msuiche/weightless— the activation-projection alternative, including the documented mislabeled hook-site on DSV4.- Arditi et al. 2024 — Refusal in Language Models Is Mediated by a Single Direction. The premise both baselines inherit and phantom-kv sidesteps.
- Lester et al. 2021 — The Power of Scale for Parameter-Efficient Prompt Tuning. Soft-prompt prior art (different objective, same artifact class).
- Zhou et al. 2025 — Don’t Say No; RAID. Jailbreak-side prior art that informed the v1 control arm’s design.
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.