Atria Dawn Preview: A 744B MoE Agent Built on GLM-5.2, MIT, Text-Only by Design — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

Atria Dawn Preview: A 744B MoE Agent Built on GLM-5.2, MIT, Text-Only by Design

Shanghai AI Lab open-sourced a 744B-parameter MoE foundation model built on GLM-5.2, focused on long-horizon research and engineering agents. It leads discovery, tool use, and cybersecurity benchmarks while explicitly shipping as text-only under MIT — with deployment recipes for SGLang, vLLM, Codex, and Kimi Code that reveal what 'agentic' means in practice.

I had to read the model card three times before the headline number landed. Atria Dawn Preview, released by the Shanghai Artificial Intelligence Laboratory on September 11, 2026, scored 96.0 on DeepSearchQA and 92.5 on BrowseComp — both topping the table of eight models — while sitting 7 points behind Opus 5 on JobBench (50.3 vs 68.0) and 11.4 points behind on SWE-bench Pro. The same release is a 744-billion-parameter MoE built on the GLM-5.2 foundation, ships under MIT, and explicitly declares input_modalities: ["text"] in its Codex catalog. The numbers aren’t the story. The shape of the numbers is.

I want to walk through what made me stop, because the model card breaks a pattern I’d gotten used to expecting from open-weight Chinese lab releases. Most of those releases come with a single big benchmark headline and a quant ladder; they assume multimodal input because the field has converged on it. Atria Dawn drops both: no headline superlative, and a deliberate text-only restriction that requires active config wiring to handle correctly in agent harnesses. That second piece — the explicit “this is text-only by design” + the integration docs that come with it — is what made me realize this release is closer to a harness contract than a model announcement.

The architecture under the hood

The README is unusually explicit for an InternLM-team release. Atria Dawn Preview is “Built on the 744B-parameter MoE GLM-5.2 foundation model,” which makes it a post-trained / RL-refined derivative of GLM-5.2 rather than a from-scratch 744B training run. The HF repo lists two checkpoints: the BF16 Instruct model and Atria-Dawn-Preview-FP8, an FP8-quantized variant — both at 256K context. The quantization step is the same recipe Z.ai ships for the GLM-5.x line; it isn’t a novel contribution, but it puts the deployment-side inference budget in line with vLLM and SGLang’s defaults rather than forcing an exotic quant path.

The deployment table is concrete: SGLang v0.5.13.post1+ and vLLM v0.23.0+. Both are recent versions — SGLang’s GLM-5.2 cookbook and vLLM’s recipes/vllm.ai/zai-org/GLM-5.2 page have been kept current, which means the standard serving stack handles a 744B MoE without a fork. That’s worth flagging because a year ago a 744B open-weight MoE needed a custom CUDA graph path or a vendor fork; now the path is documented as vllm serve Atria-Dawn-Preview with whatever KV-cache budget you can afford.

The model’s API is exposed at api.atria-asi.ai/v1 (international) and via InternAI’s discovery platform inside China. Both endpoints take OpenAI-compatible Chat Completions and a Responses-shaped stream for Codex. The wire format is standard. The model is the news.

Where the model actually wins — and where it doesn’t

The benchmark table in the README compares Atria Dawn Preview against DeepSeek V4 Pro 0813, KIMI K3, Qwen 3.8 Max, GLM 5.3, GPT 5.6 sol, and Claude Opus 5. That’s the right opponent set — a Chinese-frontier cross-section plus the two Western flagship APIs people actually pick from. Across 18 benchmark slots, the picture is:

Discovery (Atria Dawn leads or ties on 3/4):

  • DeepSearchQA: 96.0 (Atria) vs KIMI 95.9 vs GLM 5.3 94.7
  • BrowseComp: 92.5 (Atria) vs V4 Pro 83.4 vs Opus 5 90.8
  • WideSearch: GPT 5.6 sol 83.3 leads; Atria 81.9 tied with Qwen
  • DeepResearch Bench II: Opus 5 leads at 54.1; Atria 51.1 sits behind KIMI (51.3) and GLM (52.7)

Tool use (Atria Dawn leads 2/4):

  • BFCL v4: 77.0 vs V4 Pro 71.4 vs GLM 5.3 74.1
  • AutomationBench: 53.8 vs V4 Pro 41.7 vs KIMI 45.9
  • SkillsBench: Qwen 3.8 Max 66.7 leads; Atria 66.4 second
  • τ³-Bench Banking: Qwen 55.2 leads; Atria 41.2

Creation (mixed):

  • SWE-bench Pro: Opus 5 74.7 leads; Qwen 65.1, KIMI 61.6, GPT 5.6 61.4 — Atria 59.6
  • Terminal-Bench 2.1: Opus 5 90.2 leads; Qwen 89.3 — Atria 78.3
  • MLE-bench Lite: GPT 5.6 88.9 leads — Atria 86.2

Cybersecurity: CyberGym 86.5 (Atria) vs GLM 5.3 84.5 vs V4 Pro 83.3

The honest read: Atria Dawn Preview is strongest on the long-horizon evidence-retrieval tasks — DeepSearchQA, BrowseComp, BFCL v4, AutomationBench, CyberGym — and clearly behind Opus 5 on the SWE-bench Pro and JobBench axes. The job-readiness gap (50.3 vs 68.0 on JobBench, 1583 vs 1768 on GDPval points) is the load-bearing number: this is a research-and-tool-use model, not an office-work delivery model. The CyberGym lead is the most interesting cell in the whole table; I’ll come back to that.

What “agentic” actually means in the model card

The README defines the agentic surface as four dimensions: Discovery, Creation, Delivery, Cybersecurity. Three of these are standard; Cybersecurity is unusual. The CyberGym 86.5 lead is the proof point, and the README’s framing of it (“analyzing security issues, validating vulnerabilities, applying fixes, and performing re-validation in authorized environments”) is doing legal work. The release is explicitly scoped to authorized red-team / blue-team use, which keeps the safety discussion inside a permitted perimeter. For a Chinese frontier-lab release, that firewall between “We trained on cyber data” and “Use this only where you’re allowed to” is rare in the open-weight space, and the wording is closer to what Mythos-class releases do than what KIMI K3 or DeepSeek V4.1 do.

What I think is genuinely novel here — and what I’d flag as the part to copy if you’re building your own agent — is the explicit text-only restriction in the deployment docs. Atria Dawn Preview accepts text only. Multimodal input is rejected with 400 Atria-Dawn-Preview is not a multimodal model. Codex defaults every model to multimodal, which means a vanilla ~/.codex/config.toml will silently send image attachments that the endpoint then refuses. The README spells out the wiring:

{
  "models": [{
    "slug": "Atria-Dawn-Preview",
    "input_modalities": ["text"],
    "context_window": 256000,
    "max_context_window": 256000,
    ...
  }]
}

with model_catalog_json pointed at the catalog file and [features] view_image = false in the config to disable the image-view tool. The same restriction shows up in the Kimi Code integration, where capabilities = ["tool_use", "thinking"] explicitly omits image capability — adding image_in to the capabilities list makes Kimi send image attachments that Atria then rejects.

This isn’t a bug; it’s a feature. A text-only agent has a defined media surface, which makes it testable in a way that multimodal agents aren’t. You know what inputs you can simulate, what inputs are out-of-scope, and what the failure mode is for out-of-scope (the endpoint rejects with a specific error). For evaluation work, that’s a better contract than the usual “we accept all media and you’ll find out in benchmark gaps.”

The deployment corner I’ve been waiting for

The Codex integration has a trap worth flagging: model_catalog_json replaces the model catalog rather than merging. If you switch models with -m without updating the catalog, Codex logs warning: Model metadata for <slug> not found and falls back to multimodal defaults. The README has this exact failure mode documented inline. That’s the kind of detail you only see when a model release treats agent harness integration as a primary surface.

The Kimi Code config is denser. It lists five separate capabilities/effort fields that all have to be set or the model silently loads with the wrong defaults:

  • max_context_size = 256000 — without this the entry silently fails to load
  • max_output_size = 65536 — without this Kimi sends max_context_size on the wire and Atria rejects (valid range is 1–65536)
  • capabilities = ["tool_use", "thinking"] — without tool_use, Kimi treats capabilities as unknown; without thinking, the effort is forced to off
  • support_efforts = ["low", "medium", "high", "xhigh", "max"] — without this, effort stays the opaque “on” string and no reasoning_effort is sent
  • off_effort = "none" — required to allow turning thinking off; omitting it errors with “reasons by default but declares no off effort”

Five fields, five silent failures if any one is missing. The kimi doctor/kimi provider list recipe in the README walks through validating the config (OK config.toml /home/<user>/.kimi-code/config.toml / All checked config files are valid) — a level of operational integration doc that almost no other open-weight release has bothered with.

Trade-offs and what this doesn’t fix

The license is MIT for both code and weights. That’s the cleanest release in the open-weight Chinese-frontier lane right now. There’s no commercial-use carve-out like HyperFrames had, no separate “research only” clause like some GLM-5.x sub-weights carried. If you’ve been waiting on a 744B-class model you can legally fine-tune and ship in a paid product, this is the one to evaluate first.

The benchmark numbers don’t translate to “drop-in for Opus 5.” Atria Dawn is weaker on the office-work and software-engineering benchmarks — JobBench 50.3, GDPval 1583, SWE-bench Pro 59.6, Terminal-Bench 2.1 78.3 — all behind Opus 5. The discovery and tool-use leads are real but they’re for a specific kind of agent: long-horizon research, evidence-gathering, multi-step tool orchestration. If your agent workflow is dominated by “open the IDE, edit a file, run tests, edit again,” Atria is not a replacement; it’s a third option.

CyberGym 86.5 is impressive, but the authorized-environments scope is tight. The “analyzing security issues, validating vulnerabilities, applying fixes, and performing re-validation in authorized environments” wording is doing real work. If you want to use this for offensive security research, you need to confirm your authorization paperwork covers model-assisted work. I don’t have a citation for whether AISI or any other safety org has evaluated the model specifically — the model card on DeepMind-style sites doesn’t appear to be published yet, only the InternLM HF organization hosts one.

The paper (arXiv:2609.15818) is on a non-standard arXiv ID. 2609.15818 puts September 2026 in the 26 series — arXiv’s been rolling 25.x through 26.x throughout 2026, so the prefix is current, but a quick search of the canonical arXiv listings turned up the paper via the citation block in the README, not via the arXiv listing. If you’re citing it, link through the GitHub README rather than relying on the arXiv ID resolving through their search.

The “Preview” tag is doing what “Preview” usually does. The README is explicit that this is a preview release. I’d expect a v1 checkpoint, additional language coverage, and possibly a multimodal track later. The text-only restriction today is a stability decision, not a permanent architectural limit.

Where I’m going to keep watching

The integration-docs-as-feature framing is what I’d like to see other open-weight releases pick up. Most open-weight releases today ship a vllm serve line and call it done. Atria Dawn ships that line plus a full Codex catalog spec, a full Kimi Code TOML spec, an image-blocking PreToolUse hook recipe for Claude Code (a block_pdf_image_read.py Python file you register in ~/claude_dir/settings.json), and the exact kimi doctor validation command. That’s an integration surface — not a model announcement.

The CyberGym 86.5 lead also has me curious about adjacent cyber evals. GPT-6 cyber threshold (the controversial 50% number from the Mythos cycle) and the Claude Mythos-class cyber ban are both still active in 2026 — a Chinese-frontier model leading CyberGym by 3 points over GLM 5.3 is a data point that complicates the “Western labs hold the cyber lead” framing. I’d want to see the cyber benchmark methodology breakdown before drawing conclusions, but it’s worth flagging.

Last thing: the deployment version floors. SGLang v0.5.13.post1+ and vLLM v0.23.0+. If you’re already running older stacks, an upgrade is in your near future. Both engines are mature at those versions, but Atria Dawn Preview is a fine reason to upgrade if you’d been putting it off.

The paper is at arXiv:2609.15818, the HF weights are at huggingface.co/internlm/Atria-Dawn-Preview (FP8 variant in a sibling repo), and the deployment docs and integration recipes are all in the GitHub atria-asi/Atria-Dawn-Preview README. The model card on build.nvidia.com and the ModelScope mirror on Shanghai_AI_Laboratory/Atria-Dawn-Preview round out the distribution points. Worth an afternoon of poking at the SGLang cookbook and the Kimi Code wiring even if you don’t end up deploying — the integration-doc shape is what I’d want other vendors copying.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.