Nemotron 3 Nano Omni: A Multimodal MoE That Decides Its Own Quantization Recipe — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

Nemotron 3 Nano Omni: A Multimodal MoE That Decides Its Own Quantization Recipe

NVIDIA shipped a 30B-A3B multimodal MoE in BF16, FP8, and NVFP4 at the same time. The NVFP4 recipe is the interesting one — routed experts get FP4, Mamba and shared experts stay FP8, vision and audio stay BF16. The model card says this is inspired by Nemotron 3 Super, and the accuracy table shows why.

I had the Nemotron 3 Nano Omni model card open in one tab and the vLLM recipe for it in another, and the thing that kept pulling my eye was not the multimodal scoreboard. It was a sentence near the bottom of the model card’s quantization section: the NVFP4 variant uses a mixed-precision recipe inspired by Nemotron 3 Super, and then it enumerates, layer by layer, which parts of the model get NVFP4 and which parts stay at FP8. The vision and audio encoders don’t get quantized at all. The MoE router doesn’t get quantized. The lm_head doesn’t get quantized. The o_proj in attention stays FP8, even though the experts feeding it go down to 4 bits.

That sentence is the real news of the April 28 release, and it is what I want to write about. The headline numbers — 30B-A3B, 256k context, Mamba2-Transformer hybrid, video + audio + image + text in one model, Apache-friendly license — are easy to find anywhere. The quantization recipe is the part that took me a second read, and it’s the part that tells you how NVIDIA is actually thinking about deploying 30B-class multimodal models in 2026.

What the model is

Nemotron 3 Nano Omni is a 31B-A3B model: 3.1 × 10¹⁰ parameters total, 3B active per token (NVIDIA spells this 31B A3B on the model card; the Hugging Face repo uses 30B-A3B). The 30B lives in a Mixture-of-Experts language model; the 3B is what gets touched on each forward pass. The architecture is a Mamba2-Transformer hybrid, not a vanilla transformer. The vision encoder is CRADIO v4-H, and the speech encoder is Parakeet. Outputs are text only — there’s no native audio or video generation, just unified understanding of all four input modalities.

The release went out on April 28, 2026 on three channels at once: build.nvidia.com, Hugging Face, and NGC. Three checkpoints shipped in the same drop: BF16, FP8, and NVFP4. That’s unusual. Most open-weight releases ship BF16 first and let the community figure out quantization later; this one shipped all three precisions as first-class artifacts, with accuracy and throughput numbers to back each one. The model card credits prior Qwen and gpt-oss models as teacher models used for post-training improvement — Qwen3-VL-30B-A3B-Instruct, Qwen3.5-122B-A10B, Qwen3.5-397B-A17B, Qwen2.5-VL-72B-Instruct, and gpt-oss-120b. The improvement path is the distillation-style chain that has become standard for open-weight 2026 drops.

The license is NVIDIA’s commercial-use terms, English-only, with a 256k-token context window on both input and output.

The four precisions, ranked by footprint

The model card gives a clean comparison table, and it’s worth pasting in full because the numbers tell the story:

FootprintBF16FP8NVFP4
Size (GB)61.532.820.9
Effective bpw16.008.54.98

BF16 is the reference. FP8 brings the model down to 32.8 GB at 8.5 effective bits per weight. NVFP4 brings it to 20.9 GB at 4.98 effective bits per weight. Three checkpoints, three footprints, three different deployment profiles.

What FP8 actually does: per-tensor E4M3 on every linear layer in the language model, with the exception of the MoE router and the lm_head. Paired with an FP8 KV cache. The MoE router and lm_head are sensitive to quantization — they’re small but they sit on the critical path for token routing and final projection — and the model card explicitly calls them out as kept at higher precision. Everything else goes to FP8. Vision encoder, audio encoder, and their MLP projectors stay in BF16 even in the FP8 variant.

NVFP4 is more interesting. The card describes it as a mixed-precision recipe: the routed MoE experts are quantized to NVFP4 (FP4 E2M1 values with per-block FP8 E4M3 scales over groups of 16 elements and a per-tensor FP32 global scale). But the Mamba in_proj and out_proj, the shared experts, and the attention o_proj are quantized to FP8, not NVFP4. Vision and audio encoders stay BF16. The MoE router and lm_head stay higher precision than the experts.

The decision is not “quantize everything to FP4 and hope for the best.” It’s a layer-by-layer choice driven by what each layer is doing.

What the accuracy table actually shows

The model card reports FP8 and NVFP4 accuracy against BF16 across nine multimodal benchmarks:

BenchmarkBF16FP8NVFP4
MathVista_MINI71.9071.0571.30
Charxiv Reasoning49.1048.0547.95
MMlongBench Doc46.1045.8445.78
OCRBenchV2 (EN)65.8065.6365.77
CVBench2D84.2085.6285.27
Video MME70.8069.4069.60
Daily Omni74.5074.0674.23
World Sense55.2054.4054.60
MMAU74.6274.5674.34
Tedium Long (WER↓)3.113.123.04
HF-ASR (WER↓)5.955.975.95
Mean (9 non-ASR)65.8065.4065.43
Median (9 non-ASR)70.8069.4069.60
Δ vs BF16 (mean)---−0.40−0.38

Both quantized variants stay within 1 point of BF16 on average across nine multimodal benchmarks, and NVFP4 is better than FP8 on the mean (Δ −0.38 vs −0.40). On CVBench2D — the 2D computer-vision reasoning benchmark — NVFP4 actually scores 85.27, beating BF16’s 84.20 by over a point. The trend is consistent: NVFP4 holds up because the parts of the model that need precision (router, lm_head, attention output projection, shared experts, multimodal encoders) get it, and the parts that can absorb low precision (the bulk of the routed expert weights) drop to NVFP4.

The audio-side Word Error Rates (Tedium Long, HF-ASR) are essentially unchanged across precisions: 3.11 → 3.12 → 3.04 for Tedium Long, 5.95 → 5.97 → 5.95 for HF-ASR. That’s because the speech encoder (Parakeet) stays BF16 in both quantized variants. Whatever happens in the language model stack, ASR quality doesn’t move.

This is the part of the release that’s worth thinking about. The NVFP4 recipe is not “we made a smaller model.” It’s “we picked which layers to make smaller, based on what each layer is doing.”

What runs on a DGX Spark

The interesting deployment story is the DGX Spark (GB10) numbers, which are the only ones NVIDIA forum users have published in detail. The eugr forum post from May 1, 2026 reports llama-benchy results for the NVFP4 variant on a single DGX Spark using stock vllm/vllm-openai:v0.20.0-aarch64-cu128 and a single-node recipe.

The headline text-generation throughput is 56.96 tok/s on the NVFP4 checkpoint (text-only path; the model is multimodal but the tok/s figure is from text generation). The full benchy breakdown for NVFP4 includes:

Testt/speak t/sttfr (ms)est_ppt (ms)e2e_ttft (ms)
pp20486625.34 ± 93.11314.26309.38314.26314.26
tg3260.10 ± 0.5162.06---------
pp2048 @ d40964467.19 ± 1279.75515.74510.86515.74515.74
tg32 @ d6553567.68 ± 5.1469.90---------

Prompt processing at 2k tokens hits 6,625 tok/s. Text generation at 32-token output is 60.10 tok/s. The 65,535-token context test (tg32 @ d65535) still gets 67.68 tok/s — full 256k context, NVFP4, on a single GB10 board. The FP8 variant on the same hardware, from the same forum thread, reports tg32 at 61.99 ± 3.81 tok/s and pp2048 at 3,631.92 ± 596.12 tok/s. The 56.96 tok/s figure quoted in the thread title is the aggregate text-generation throughput from the published benchmark — a slightly different metric than the llama-benchy tg32 cell.

For comparison: the NVFP4 variant at 56.96 tok/s on a single 128 GB unified-memory Spark box is the same ballpark as the FP8 30B-A3B reasoning models running on dedicated inference GPUs, and it’s substantially faster than what most open-weight 30B-A3B checkpoints achieve on a single Spark. The thread title’s “achieved 56.96 tokens/sec” reads as a marketing number; the benchy breakdown is what tells you where the time goes.

The actual vLLM command

The DGX Spark recipe is published as a single-file YAML that the run-recipe.sh script executes. The relevant block:

recipe_version: "1"
name: Nemotron-3-Omni-FP8
description: vLLM serving Nemotron-3-Omni-FP8 on a single node
model: nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8
container: vllm-node
solo_only: true
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.8
  max_model_len: 262144
  max_num_batched_tokens: 49152
command: |
  vllm serve nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8 \
    --max-model-len {max_model_len} \
    --port {port} \
    --host {host} \
    --trust-remote-code \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --reasoning-parser nemotron_v3 \
    --kv-cache-dtype auto \
    --video-pruning-rate 0.5 \
    --allowed-local-media-path / \
    --max-num-batched-tokens {max_num_batched_tokens} \
    --gpu-memory-utilization {gpu_memory_utilization}

A few things to notice. --reasoning-parser nemotron_v3 is a vLLM-side parser that splits reasoning tokens from output tokens — the model emits <think>...</think> blocks and the parser surfaces them separately. --tool-call-parser qwen3_coder lets the model emit structured tool calls; the choice of qwen3_coder is interesting because it means NVIDIA trained the model’s tool-call grammar against Qwen’s parser format, even though the model is otherwise Nemotron. --video-pruning-rate 0.5 drops half the video frames before they hit the language model — a server-side knob for trading accuracy against prefill cost on long video inputs.

The gpu_memory_utilization: 0.8 is conservative for a 32.8 GB model on a single 128 GB unified-memory board; the headroom is for the multimodal encoder cache and the KV cache at long contexts. --max-model-len 262144 is the maximum supported context; the model card claims 256k but vLLM pads it to a power-of-two-friendly 262144.

The Spark-optimized invocation adds VLLM_SPARK_EXTRA_DOCKER_ARGS="-e TORCH_ALLOW_TF32_CUBLAS_OVERRIDE=1 -e NVIDIA_TF32_OVERRIDE=1" before launching the recipe, which forces TF32 matmul even though the model is FP8/NVFP4 — a Spark-specific knob that recovers a few percent on Blackwell-class fused matmul.

Why the Mamba hybrid matters for quantization

This is the part that took me a second read. The model is described as a Mamba2-Transformer hybrid, which means some layers are attention and some are state-space. State-space layers don’t have a KV cache in the attention sense — they have a fixed-size recurrent state, scaled by a constant regardless of sequence length. For a 256k-context model, that matters a lot: the KV cache for a pure-transformer 30B-A3B at 256k tokens is enormous, but the recurrent state of a Mamba2 layer is O(1) per token.

The quantization recipe reflects this. The FP8 variant quantizes the language-model linear layers to FP8 and pairs it with an FP8 KV cache. But the NVFP4 variant doesn’t quantize the Mamba in_proj and out_proj to NVFP4 — it leaves them at FP8. State-space model weights have different quantization sensitivity than attention weights (the in/out projections are essentially MLP layers around the SSM kernel, and they sit on the path that the recurrent state flows through), and the recipe author chose to preserve more precision there even when the routed experts go down to 4 bits.

This isn’t documented in any single line of the model card — it’s the structural implication of “Mamba in_proj / out_proj … quantized to FP8” while “routed MoE experts are quantized to NVFP4.” The two are doing different things in the forward pass, and they get different precisions.

What the model is actually for

The model card’s own framing is worth quoting: “NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (OCR), and long-context video reasoning capabilities.”

The OpenRouter listing calls it “a perception and context sub-agent in enterprise agent systems.” That’s a more honest framing. The 30B-A3B MoE isn’t a frontier chat model — Kimi K3’s 2.8T parameters and GLM-5.2’s 744B parameters are above it on the open-weight stack — but it’s a multimodal perception engine that can sit underneath a larger orchestration model. Real-time video + audio + image + text ingestion at 60+ tok/s on a single Spark, with 256k context, is a specific kind of deployment profile.

The 9-benchmark accuracy table covers the multimodal evaluation surface: video understanding (Video MME, MMlongBench Doc), 2D/3D visual reasoning (CVBench2D, MathVista_MINI), document OCR (OCRBenchV2), long-context audio transcription (Tedium Long, HF-ASR), and embodied agent tasks (MMAU, World Sense). These are the workloads you’d actually want a perception sub-agent to handle. They’re not the workloads you’d want a chat model to handle — for that, you’d reach for Kimi K3 or one of the Claude/GPT frontier models.

Trade-offs and what the recipe doesn’t fix

Three honest limits.

First, the NVFP4 recipe is calibrated against BF16 on nine multimodal benchmarks, not against frontier chat benchmarks. The mean delta of −0.38 across multimodal evals is reassuring, but no SWE-bench, no AIME, no terminal-bench number is in the table. The model is positioned as a perception sub-agent, not a frontier reasoning model, and the benchmark surface reflects that positioning. If you’re trying to decide whether Nemotron 3 Nano Omni can replace a frontier chat model on coding or math, the released numbers don’t tell you.

Second, English-only. The model card says “Language support: English only” three times. For enterprise multilingual workflows, this is a hard limit, not a soft one. Qwen3.5-122B-A10B and the other distillation teachers are multilingual, but the student model only emits English.

Third, the multimodal encoders stay BF16 even in the NVFP4 variant. This is the right call for accuracy, but it means a 20.9 GB checkpoint still requires BF16 storage and compute for the CRADIO vision encoder and the Parakeet speech encoder. The “NVFP4 is small” framing applies to the language-model stack only. On a 128 GB Spark, this is fine. On a smaller edge box where you wanted a 21 GB everything-included checkpoint, you don’t get one — and NVIDIA doesn’t publish a fully-quantized multimodal variant.

The thing the recipe does fix: it lets a single 128 GB unified-memory board run a 30B-A3B multimodal model with 256k context, on three precision footprints that map cleanly to different deployment profiles. BF16 is the reference. FP8 is the production default. NVFP4 is the size-constrained deployment. Each one ships with accuracy numbers against the same nine multimodal benchmarks, so the user can pick the precision and know what they’re trading.

The structural decision — “which layers get NVFP4, which stay FP8, which stay BF16” — is what makes the release newsworthy. It’s not a single quantization pass. It’s a deliberate choice, layer by layer, with the reasoning traceable in the model card.

The question I want to leave open: this recipe works for a Mamba2-Transformer hybrid MoE with CRADIO + Parakeet encoders. It is explicitly “inspired by Nemotron 3 Super.” Does it transfer to a pure-transformer MoE? Does it transfer to a non-hybrid MoE with a different multimodal encoder? The accuracy table doesn’t tell you — it tells you that this recipe works for this model on these benchmarks. The generalization question is open, and the next 30B-class multimodal drop will probably answer it.

References and where to dig further

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.