A coding agent is easy to launch and hard to trust. You can hand it a training script, ask it to sweep a learning rate, and watch it happily report whichever run happened to land last as “the answer” — sometimes after editing the previous attempt to look worse than the new one, sometimes by simply forgetting what it already measured. The pitch for alphaxiv/openresearch is that the harness around the agent should be the thing that refuses to lie. Not by being clever about the agent itself, but by making the data structures it touches (experiments, runs, branches, evidence) too rigid for a coding agent to quietly rewrite.
OpenResearch is a local-first workspace for research agents. The CLI is orx. It owns your projects, your experiment tree, the runs themselves, the evidence each run leaves behind, and the dashboards that show them. The companion service at openresearch.sh handles the parts that can’t be local: accounts, organizations, sandbox provisioning, and managed-compute catalogs. Project state — every experiment node, every run, every log line, every artifact — stays on your machine in a SQLite database under 127.0.0.1:4791. That split is the whole reason the design feels different from other “agent orchestrator” launches this year: the data your agent cheats on has to be the data your disk holds.
The bet, in one sentence
A research agent that cannot edit a node after a run has answered it, and cannot vary its run command between branches, will produce a tree-shaped experimental record that an honest human can read without re-running anything to verify.
The model behind that bet is an experiment tree. The root is the baseline — it carries the starting code and a single shell command that trains or evaluates it. Every other node is a child branched off a parent, inheriting both the code on that branch and the parent’s run command. Each node is either provisional (still being fixed up before its first real answer) or frozen (a run has answered it, good or bad — nan is a valid result). The frozen state is permanent; the only way to try a new idea is to branch a child and edit the child.
The four cardinal rules in agent-skills/orx-experiment-tree/SKILL.md are the spine of the project:
- Never edit a node once a run has answered it. A node freezes the moment a run establishes its baseline or tests its hypothesis — that includes the root — and freezing is permanent. Until then it is provisional.
- The run command and the environment are a fixed contract — identical on every node. A child inherits its parent’s run command verbatim. Don’t give nodes different start commands, and don’t vary behavior through environment variables or env-prefixed commands.
- Vary code, not knobs-in-the-command. Encode hyperparameters in the code/config files and branch a child per variant. Never sweep them by editing the run command or passing env vars.
- Grow the tree downward, not sideways. Fan a little within a round, then descend onto that round’s winner for the next round.
These are not style preferences. The orx-experiment-tree skill is explicit: “Breaking any one silently invalidates your results.” The reason isn’t philosophical, it’s structural — every result is comparable to its parent only because they ran the same command over different committed code. Change the command and the comparison is meaningless. Edit a frozen node and the evidence attached to it is invalidated retroactively. The CLI enforces this by refusing to give you a working tool when you violate it: there’s no --change-run-command flag, and the node model in the SQLite store has no field for editing a frozen branch.
What’s actually hard about an agent harness
Three things show up the moment a coding agent starts running experiments on its own.
First, the experiment tree grows wrong by default. The skill documents two failure modes by name. A flat fan — every sweep hanging off the root — measures everything against the start, so wins never accumulate and the tree never makes progress. A noodle — a long single-child chain — manufactures depth without any of it building on the step above. The correct shape is stacked bushes: a small fan within a round (the options of one decision), then descend onto that round’s winner for the next round. The decision rule the skill teaches: “Before you make X a child of Y, name what Y established that X builds on. You can name it → real depth, descend. You can’t → co-equal options, fan as siblings.” That sentence is doing more work than the entire design doc, because it gives the agent a concrete test for a question it will ask badly otherwise.
Second, the run command is the wrong place to express variation. Sweeping a hyperparameter by editing the run command is the path of least resistance — and it’s the one that destroys the comparison. Every node needs to run the same shell command so that the only thing their logs differ on is the code they ran. Encoding the sweep in the committed config file (or in the model definition, or in the training script’s constants) is the discipline that keeps the tree’s evidence comparable. The cardinal rule “vary code, not knobs-in-the-command” is repeated because it is the rule most often bent.
Third, the per-completion loop is the loop body, not a barrier. orx exp wait --project <projectId> is a sleep-until-change signal, not a source of truth. It returns on the first completion during that one call, not on the full batch draining. The skill is explicit about this — and warns that any run finishing while you analyze the previous one is already terminal by the next call and won’t be reported. So on every wake you re-read orx runs, reconcile against the runs you’ve already handled, act on each newly-terminal run, and re-issue exp wait. Don’t expect a single exp wait to block until everything is done; that’s the failure mode the loop avoids. When exp wait prints “drained: no runs in flight”, you’re done. Don’t keep calling it into a timeout.
The mechanism, in one bundle of decisions
each experiment node = (parent_id, branch, run_command, frozen: bool, runs[])
each run = (commit_sha, backend, status, log_path, started_at, finished_at)
the only mutator on a frozen node = "branch a child"
the run_command field is set once on the baseline and inherited
The four primitive operations are orx create-experiment, orx exp run, orx exp wait, and orx exp desc. The cardinal rules are enforced by what those commands don’t let you do. orx project edit <localProjectId> --run-command '<cmd>' is the only place you set the command, and you set it once. There’s no per-child command. There’s no flag on orx exp run that lets you pass --env LR=3e-4 and consider it the same experiment. --force exists for concurrent launches on the same node, but the absence of --env and --set-cmd is itself the API contract.
What changes between nodes is the branch. orx create-experiment prints the child’s branch (orx/<slug>); you check it out in your local worktree, edit only the files the idea touches, and commit. The commit’s SHA is what the backend runs — every backend receives an immutable snapshot of the recorded commit, uncommitted files excluded, no backend needs a GitHub push. So “what code did this run?” is always answerable from the run record alone.
The free-form orx exp desc (markdown, writeable via --set or stdin) is where you record what the node is testing. The system prompt the harness injects into every agent session makes the discipline explicit: measured results must be cited with <run id="<runId>" /> tags, code and file facts with raw <file path="..." /> tags. Status alone is not evidence. “Read the cited run’s log before reporting the result.” That’s the anti-cheating glue.
What the numbers actually show
OpenResearch shipped v0.2.4 on 2026-09-17, two days before this post, on a cadence of ten releases in eighteen days (v0.1.119 on Sep 3 → v0.2.4 on Sep 17). The current Cargo.toml reads version = "0.2.5" — a v0.2.5 is in the source tree, presumably queued for the next push. The crate is MIT-licensed (Cargo.toml: license = "MIT", repository = "https://github.com/alphaXiv/OpenResearch"), and the orx binary compiles to a static musl Linux build with no system SQLite or OpenSSL — rusqlite is pulled in with features = ["bundled", "backup"] so the SQLite source ships inside the binary, and reqwest is pinned to rustls-tls to avoid the native-tls / openssl system dependency that breaks musl builds. That bundling decision is the reason a single Linux install script (curl -LsSf https://openresearch.sh/install.sh | sh) can deliver a self-contained binary.
The agent side is the second number that matters. Five harnesses are wired into the same workflow via the bundled agent-skills/ set: Claude Code (--append-system-prompt-file), Codex (developerInstructions), OpenCode (config.instructions), Cursor, and Google Antigravity. Each one receives the same SYSTEM_PROMPT.md playbook (with {token} substitution at render time) and the same per-task skill routing. The harness owns the project state; the agent owns the code edits on its private worktree. The CLI commands (orx projects, orx project view, orx runs, orx logs, orx exp run, orx exp wait) are the surface the agent talks to.
Compute backends, all reached via the same orx exp run <expId> with --backend <name>: Hugging Face Jobs (hf), Modal, Kubernetes, SSH, Slurm, Ray, OpenResearch managed compute, Tinker, and local. Each backend has its own reference doc under agent-skills/orx-compute/references/ that the skill routes you into only when you’re about to use it. The compute catalog is browsable in-CLI (orx compute --gpu H100_SXM --count 1), and --cpu is a first-class flag for CPU-only runs. SSH is interesting: orx up --remote user@host opens the dashboard over an SSH tunnel to a remote machine; the comment in the README is honest about the security model — “The remote service binds to loopback and has no application-level authentication, so other users on that host can reach it.” That’s a tradeoff they call out instead of burying.
Literature retrieval is its own four-primitive surface, all login-free: orx discover keyword, orx discover embedding, orx discover openalex, orx discover biorxiv. The retrieval ranker is the main agent itself, not a sub-agent — the orx-lit-review skill is emphatic about this (“Never delegate the retrieval loop to a sub-agent”). Each primitive returns the same JSON shape (source, self-routing id, title, abstract, date) so the ranker can fold them together. alphaXiv carries votes and full-text snippets; OpenAlex and bioRxiv may carry citations. Date and priority controls (--published-after, --prioritize recency|historical|popular|default) are the knobs — and the skill warns that older --published-before embedding searches can return thin or empty candidate sets because the bound applies after vector retrieval; empty is not proof that no literature exists.
Trade-offs and what it doesn’t fix
The harness doesn’t fix agent honesty, it shifts where dishonesty has to land. A coding agent can still produce a tree full of nodes whose descriptions don’t match what the code actually changed. It can still bury a regression by branching a child that performs worse and never promoting it. The CLI can’t see those failures because they look identical to a normal decision: branch, run, log, stop. The mitigation is procedural — orx exp desc carries a markdown note, and the system prompt instructs the agent to cite measured results with <run id="..." /> tags and to read the cited log before reporting — but a misbehaving agent can still produce citations that lie. The harness makes lying expensive (every result has to be on a real run on a real commit), not impossible.
The fixed run command is a discipline the project enforces through API absence, but the API absence breaks the moment you need it. Distributed training that needs a different launcher per node? Sweeping over a hyperparameter that lives in an environment variable because the training framework refuses to read it from config? You can’t do either inside an orx run; you have to encode the variation in committed code. That’s the right call for research reproducibility but it means OpenResearch is hostile to any workflow that already wants to vary commands per experiment. That’s a tradeoff, not a bug, and it shows up most loudly on Slurm / k8s where a launch script’s --env is the conventional way to express variation.
The telemetry default is opt-out. Official release builds send coarse usage events tied to a random installation ID; orx telemetry off and orx <command> --no-telemetry exist; source and dev builds don’t send anything. The events “do not include code, prompts, file contents or paths, repository names, tokens, emails, or project and experiment identifiers” — a deliberate narrow scope — but it’s still a phone-home by default. For a research tool whose entire pitch is “your work stays on your machine”, the default should arguably be the other way. The trade-off is real: the project needs some signal to know which installs are alive; making it opt-out lowers the bar to that signal at the cost of forcing researchers who care to remember to flip the switch.
The companion-service split is the last thing worth flagging. OpenResearch (the local Rust CLI) owns the data; openresearch.sh owns accounts, organizations, and managed-compute catalogs. Local runs and project state work without an account; organizations and managed compute don’t. That split is the cleanest answer to “what does this thing need from the cloud?” — basically nothing, except when you want a hosted GPU catalog. But it’s also the answer to “what breaks when openresearch.sh has an outage?” — your local runs and your existing project state are fine, but you can’t reach the compute catalog, can’t manage an org, and the install script curl -LsSf https://openresearch.sh/install.sh | sh doesn’t work if the host is down. Source builds are immune; release builds aren’t. That’s a real coupling that the project could decouple by hosting the install script on a static CDN.
What changed between Sep 3 and Sep 17
v0.1.119 on 2026-09-03 is roughly the version you’d have read about two weeks ago — a Rust CLI with the experiment tree, the four cardinal rules, the bundled SQLite. v0.2.4 on 2026-09-17 added the orx-lit-review retrieval surface as a separate, bundled skill (it lived inline in the experiment-tree skill earlier), formalized the orx-compute skill’s per-backend routing table with one-reference-per-backend as the contract, and turned the orx-figures and orx-paper skills into first-class peers (matplotlib-not-publishable by default, LaTeX templates that compile to PDF). The system prompt was split into a separate SYSTEM_PROMPT.md so the harness-specific injection points (Claude Code --append-system-prompt-file, Codex developerInstructions, OpenCode config.instructions) could be reviewed independently of the skill bodies.
The most interesting change is v0.2.0 on 2026-09-11 — the point at which the v0.1.x line stopped and the v0.2.x line began, and at which v0.1.123 shipped the same day as v0.2.0. That’s a back-compat-aware bump: v0.1.123 is the last v0.1.x with whatever migration shim you need, v0.2.0 is the new line. The discipline of bumping minor on the same day as a final patch on the previous minor is the kind of signal that suggests the maintainers care about users who upgrade across the boundary.
What I’d want to see next
The piece I’d most want is a way to ask the harness, after the run log is written, to attach a measured-result block to the run: a structured summary the agent must fill in (metric: value, step: N, seed: 42) before orx exp run returns success. Today the result lives in free-form log text and a markdown orx exp desc; the agent can write whatever it wants in there. A schema-bound summary that the CLI validates before it marks the run as terminal would close the last obvious cheating surface — the one where the agent writes a log line that doesn’t match what the code actually printed.
The second is compute-defaults as code. orx exp run <expId> uses the project’s configured default backend; today that default is a single string the user sets somewhere. If the default were itself an experiment-tree artifact (a orx/<backend-default> branch whose run command is the default-resolution command), the same discipline that protects experiments would protect the choice of where to run them. That’s the design ambition the project name openresearch gestures at: not a CLI, but a research substrate that happens to ship with a CLI on top.
Where to dig further
agent-skills/orx-experiment-tree/SKILL.md— the four cardinal rules, the stacked-bushes shape, the per-completion loopagent-skills/orx-compute/SKILL.md— the launch contract, the per-backend routing table, thereferences/<backend>.mddisciplineagent-skills/orx-lit-review/SKILL.md— the retrieval ranker protocol, the date and priority controlsagent-skills/orx-evidence/SKILL.md— how to capture and inspect run evidenceCargo.toml—version = "0.2.5",license = "MIT", thebundledSQLite andrustls-tlsmusl buildsrc/main.rsandsrc/local/opencode.rs(playbook_md()) — the harness-injection seam for the system promptAGENTS.md— the repository guide that documents which side of the local/service split owns what
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.