The Hacker News launch thread for TypeSafe’s System One models hit 1,924 points and 504 comments in five days. The model behind it — Jev 1.13 — is served from one endpoint at https://api.typesafe.ai/v1/systemone, costs $42 per billion input tokens, and is free on the output side. It does not generate text. You send it a state and a map of typed questions; you get back probabilities for each answer label. That contract is what makes the deploys, the open-source alternatives, and the integration patterns look so different from anything on the LLM side of the fence.
The angle I want to walk through isn’t “is calibrated probability useful” — that question’s been settled by every ranker and reranker that’s shipped in the last three years. The angle is the inference shape: TypeSafe’s own API contract, the system-one-adapter shim that lets you swap LLM providers in for A/B comparisons, and the openjev-sglang deployment recipe that runs the API surface on top of Qwen3.6-35B-A3B and SGLang 0.5.19’s Rust frontend. Three pieces of code, three different jobs, and they all converge on the same observation: when the model is required to return a typed distribution instead of free-form text, every layer of the stack changes shape.
The bet, in one sentence
A System One model takes natural-language state, evaluates it against one or more typed questions, and returns probabilities — not generated text. The contract is closer to a database query than to a chat completion. Jev accepts three primitive question types: Choice (which option fits), Score (where on an ordered scale), and Noul (yes/no with the probability of yes). Each answer carries the full distribution across the answer labels and a confidence number, computed as 1 - H(probabilities) / log(number_of_options), clamped to [0, 1].
You can see what the API contract looks like in the TypeSafe HTTP reference. The request is state, model, and a questions map whose keys you choose; the keys are never sent to the model and exist only to thread answers back to your code. This is the contract that the rest of the post is going to keep coming back to: typed inputs in, typed distributions out, your keys preserved on the response. There’s no “rewrite the response in third person” middle step. There’s no streaming. There’s no message role negotiation.
What “prefill only” actually means
The openjev-sglang README has a section called “How inference works” that reads almost like a paper writeup. The whole thing is five steps and they all happen on the request side, before any model output:
- Validate the schema, answer count, body size, context length, total token budget.
- Render the chat template once, with thinking disabled. Split out a common prefix and tokenize each question suffix independently.
- Send the common prefix to
/generatewithmax_new_tokens=1, await completion, discard the sampled token. This warms SGLang’s radix cache. - Concurrently send
prefix + question suffix + assistant header + "Answer:\n"for each question. Every call again hasmax_new_tokens=1. Requesttoken_ids_logprobfor every answer label andlogprob_start_len=-1, so no recomputation of prompt logprobs. - Renormalize the requested label logprobs with stable softmax. Noul returns
P(yes), Choice returns the argmax and full distribution, Score returnssum(level_index * probability).
The pattern is prefill plus first-token readout. There is no generated chain of thought. There is no autoregressive continuation after the first token. The sampled token itself is ignored. There are N+1 one-token calls for N questions, including the cache-warming call. Speculative decoding is not enabled.
That’s the deployment shape. It’s not the LLM shape. For 64 questions against a 64K-context state, you’re sending ~65 short prefill calls and reading one logprob vector per call. Token usage is dominated by usage.input_tokens, which “sums SGLang’s full prompt counts across the warm-up and all branches, including cached tokens.” Output tokens count to N+1. That’s a workload where vLLM’s PagedAttention advantages largely vanish — there’s no long autoregressive tail to schedule, no KV-cache pressure from generated sequences, and the prefix-share across branches is exactly the case SGLang’s radix cache is optimized for.
How the adapter shim is wired
The interesting part of the second repo, typesafe-ai/system-one-adapter-python, is what it does not do. It is “a drop-in replacement for typesafe_sdk’s system_one evaluation API, backed by LLM APIs instead of TypeSafe.” It returns the same response object as the real client — response.answers and the typed views are unchanged — but the answers come from gpt-4o-mini or claude-haiku-4-5 instead of Jev. The README is explicit about why this exists: “Useful for comparing TypeSafe against an LLM on cost/speed/intelligence.”
The mechanism is provider-specific structured-output modes. OpenAI gets the Responses API with strict JSON Schema for structured output; Anthropic gets native tool-use with max_tokens=8192 for larger question sets. The adapter measures latency, retries, and n_retries_malformed_structure per call. A failed output raises typesafe_sdk.TypeSafeError with a message that says “increase max_tokens or request fewer questions” — it doesn’t burn retry budget on malformed-output recovery. This is the layer where the cost-savings argument lives. If you can pin 70% of your classification traffic to a $42/Btok decision model and the remaining 30% to a reasoning model on a longer horizon, you’ve bought yourself a different cost curve than routing everything to a frontier chat model.
The license is MIT. The provider SDKs are optional extras — 'system-one-adapter[openai]' or 'system-one-adapter[anthropic]'. The shim is small enough that it sits comfortably inside the typesafe-ai/skills agent skill, which TypeSafe ships as a Claude Code plugin (claude plugin install typesafe@typesafe-ai) and via npx skills add typesafe-ai/skills --skill typesafe-ai for any agent that consumes skills.sh. 997 stars on the skills repo as of this morning, up from the 530 it had a week ago.
How the third repo maps the inference shape onto SGLang
openjev-sglang (ekzhang, 213 stars, MIT, created 2026-09-17) is the most opinionated piece of the three. It runs Qwen3.6-35B-A3B on SGLang 0.5.19’s Rust frontend with breakable prefill CUDA graphs, exposes the Jev HTTP API on top, and ships a Modal deployment script that scales a single B200 server to zero after five idle minutes. The defaults tell you how the workload is meant to be run:
- 64 questions per request maximum
- 2–64 answers per Choice or Score
- 32,768 tokens per branch including its output
- 262,144 total submitted input tokens
- 16 simultaneous evaluations, 64 simultaneous backend calls
What you notice reading the code, before reading any of the prose, is that every default is about keeping the prefill cache hot. The smoke test runs a 64-answer question, plus basic semantic sanity checks and rejection of 65 answers; it reports startup wait, inference latency, and cache usage. Cache warmups explicitly request one unused token probability to dodge SGLang’s mixed-logprob batch crash without patching SGLang itself. The x-openjev-prefix-tokens header exposes the requested common prefix size, and x-openjev-cached-tokens reports branch cache hits — with the caveat that SGLang 0.5.19’s Rust frontend omits these counts, so the header is absent and smoke reports null rather than a misleading zero. The scheduler logs still show actual cache hits, and the README confirms the live B200 deployment verifies that.
A detail that surprised me: answer labels are A–Z, then verified single-token letter combinations (AA, AB, …), because Qwen tokenizes 10 and 64 as multiple tokens. All 64 labels are checked against the actual tokenizer at startup. This keeps 64-way classification an exact one-token readout instead of comparing only the first digit of a multi-token number. The system-one-adapter README mentions the same constraint in a different place: an Anthropic evaluation that runs out of tokens raises a typed error that says so explicitly rather than producing a malformed JSON answer that downstream code has to detect.
Why the ecosystem looks the way it does
The HN search shows a single point of gravity: TypeSafe, with 13+ third-party repos built on top in 5 days. I counted six repos with over 100 stars on Friday: typesafe-computer-use (580 stars, “Computer use for about $0.0002 a step: OCR the screen, classify the next action with TypeSafe”), openJev-verdict-2.0 (149 stars, 151M open-source RLCD model hitting 77.10% accuracy and 1.44% ECE on a typed-decisions benchmark), pg-jev (236 stars, a Postgres extension that runs questions against your tables), and a half-dozen skill adapters. The fourth-largest show-HN right now is “Open-Source Alternative to TypeSafe.ai” with the description: “Jev, Fly Me to the Moon” — a community-built fine-tune of Qwen-2.5-1B with RLCD-style training. None of these are “another LLM wrapper.” All of them want the typed-decision contract; they argue about which model serves it best.
What I find interesting about the open-source side is the cost numbers. The typesafe-computer-use README quotes “$0.0002 a step” for screen classification — about 4,800 classifications per dollar at list price. That’s the kind of cost-per-decision that makes an action-agent loop feasible in a way that frontier-chat pricing never has been, even when the prompt is short. (Compare: at $5/$30 input/output per million tokens, a typical 500-token classification prompt costs ~$0.0025 on input alone, plus whatever the model wants to generate for an answer. Jev’s whole cost is in the input tokens, and the input budget is exactly the state plus the questions — there’s no generated answer to pay for.)
The PG extension pattern is what I’d expect to spread fastest. pg-jev lets you write SELECT * FROM classify_tickets(...) USING MODEL jev-latest WITH QUESTIONS ('urgency:noul', 'department:choice[billing|technical]'). That kind of integration was previously impossible without first running a chat completion, parsing free-form text, and reconciling against the database schema. With the Jev contract, the parsing is the contract.
What I’m not sure about yet
Three things I haven’t gotten a clean answer on. First, the confidence semantics page notes that “Choice/Score confidence summarizes distribution concentration, not overall workflow correctness or permission to act.” That’s a sharp caveat and it directly affects how you wire the answer into downstream code — a high-confidence Choice answer still needs your deterministic checks. The TypeSafe docs are unusually honest about this, but I haven’t seen a third-party benchmark that pins down the actual calibration error of Jev on production-shaped workloads. The openJev-verdict-2.0 repo quotes ECE 1.44% on its own benchmark, which is impressive but obviously not Jev’s number.
Second, the rate limits are explicit dynamic numbers: “We are serving a very large volume of demand, and the limits above can change without notice while we do.” 250,000 tokens per second and 1,200 requests per minute per account is a real ceiling, and the warning that “Higher limits are available on custom and enterprise plans” is the kind of sentence that appears in vendor pages right before a pricing conversation. For someone running 100,000 classifications an hour, the system-one-adapter shim against a self-hosted openjev-sglang deployment isn’t optional — it’s the path.
Third, the founder cred. Diogo Almeida co-invented RLHF at OpenAI (InstructGPT), which is the same training path Jev explicitly argues against in its AI primer. The startup TypeSafe AI launched publicly on September 15, 2026, with the manifesto framing: “RLHF teaches a model to say things that people prefer … An output can be compelling to a person without being reliable enough for unattended automation. Human preference and machine trustworthiness are different optimization targets.” That argument is genuinely interesting and I want to read more of it before I take a position. The third training path they call out — RLCD, Reinforcement Learning for Calibrated Decisions — is the one that’s worth watching. If the training recipe generalizes the way the JevBench numbers suggest, every classification-shaped workload in production has a different option than it did three weeks ago.
The thing I keep coming back to is the deployment shape. The whole API is one endpoint, one model field, and a typed question map. The whole ecosystem is a few hundred lines of code on each side. That’s not a coincidence — it’s the contract doing the work. When the contract is “send state, get probabilities,” the integration surface is small, the failure modes are constrained, and the open-source alternatives have something specific to compete on. That’s a better starting position than any of the LLM wrappers I can think of.
I’ll keep an eye on the RLCD training recipe writeup when it lands, and on whether the typed-decisions benchmark ends up running on a public leaderboard. The current leaderboard is at benchmarkheaven.com/jev-models and is community-run; the moment there’s a peer-reviewed version of the same protocol, the conversation shifts from “what is this thing” to “which model handles it best.”
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.