OpenFable: When the RAG Engine Refuses to Chunk Your Documents — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

OpenFable: When the RAG Engine Refuses to Chunk Your Documents

OpenFable ships a fresh Apache-2.0 implementation of the FABLE paper — bi-path retrieval, semantic-forest indexes, and 92.1% DragBalance completeness against Gemini-2.5-Pro with 94% fewer tokens. The interesting design choice isn't the math; it's that the engine won't chunk your documents at all. The agent has to do it.

OpenFable landed on Show HN today with seven stars, an Apache-2.0 license, and a sentence in the README that I keep coming back to: “OpenFable does not chunk documents. You read the document, choose the chunk boundaries, and build the tree.” That is the whole pitch, and it’s more interesting than the headline benchmark number (94% token reduction vs full-context Gemini-2.5-Pro at 92% completeness) because it forces a question most RAG papers skip: who actually decides where one topic ends and another begins? The retrieval engine that ships today says it shouldn’t be the engine.

The bet, in one sentence

A semantic-forest index built by an LLM at chunking time is structurally different from a flat vector index built at ingestion time, and the difference matters most when the answer to a query is buried in a subsection that doesn’t match the query’s surface keywords.

I want to be careful with that word “forest” because the paper uses it deliberately. The FABLE paper (arXiv:2601.18116, “Forest-Based Adaptive Bi-Path LLM-Enhanced Retrieval”) doesn’t build a single tree per document — it builds a forest because documents at the corpus level are siblings of the tree, and the retrieval path can switch between traversing the structure and querying the embedding space at each level. OpenFable, the open-source implementation released today, makes that explicit by separating the document-level path (LLMselect + vector top-K over internal nodes) from the node-level path (LLMnavigate + TreeExpansion) and fusing the two before trimming to a token budget.

What the numbers actually show

The benchmark table in the README quotes the FABLE paper directly. I’ll reproduce it with the caveat that the OpenFable maintainers haven’t yet run their own numbers against the implementation — the README says so explicitly and welcomes contributions:

BenchmarkMetricBM25BGE-M3HippoRAG2FABLE
DragBalanceRecall66.1%64.0%39.2%85.8%
DragBalanceCompleteness67.9%67.2%62.2%92.1%
HotpotQAEM36.4%51.5%41.0%46.5%
2WikiMultiHopQAEM——44.5%52.5%
BrowseComp-plusAgent accuracy——44.5%66.6%

Two numbers from the paper’s narrative stuck with me. The first is the headline framing: FABLE reaches 92% completeness on DragBalance using 31K tokens where Gemini-2.5-Pro with the full 517K-token document lands at 91%. The paper is saying that a structured retrieval path can match — and modestly beat — a frontier model reading the whole document, at roughly 6% of the token cost. The second is the curve: at 8K tokens, FABLE hits 98.2% of the full-document upper bound. Past that, you pay tokens for diminishing returns.

What the numbers don’t show is the tail. The hotpot QA EM score of 46.5% is in line with BGE-M3 flat retrieval (51.5% — actually beats FABLE on that metric) and below HippoRAG2 on multi-hop EM. The paper’s strongest wins are on agent-oriented tasks (BrowseComp-plus at 66.6% vs HippoRAG2’s 44.5%, a 22-point gap) and on the structural recall metric where discourse-aware chunking matters most. If your task is “find a single fact in a long document”, flat embeddings or BM25 still win. If your task is “answer a multi-hop question that requires knowing how the document is organized”, FABLE pulls ahead. The comparison the README draws against RAPTOR is the cleanest framing I’ve seen for this: RAPTOR builds a bottom-up tree by clustering embeddings, FABLE builds a top-down tree by asking an LLM where the discourse boundaries actually are. Different generation, different retrieval.

The mechanism, in one equation

TreeExpansion is the structural scoring function. Each node v in the tree gets a score:

S(v) = 1/3 · (S_sim(v) + S_inh(v) + S_child(v))

Where S_sim is the cosine similarity of the node’s embedding to the query (with depth decay — deeper nodes decay), S_inh is the inherited score from the parent (so a high-scoring parent lifts its subtree), and S_child is the aggregated child score (so a leaf with strong evidence lifts its ancestors). The averaging isn’t a smoothing trick — it’s what makes the forest behave as a forest rather than a flat retrieval over a tree. A leaf that strongly matches a query can pull its parent’s score up; a parent that strongly matches can drag its subtree’s score up; the final ranking is whatever combination of local evidence, structural context, and inherited relevance survives the average.

I went and read skills/openfable/SKILL.md to verify the depth-decay claim because the README wasn’t explicit. The skill file documents OPENFABLE_RETRIEVAL_LLMSELECT_DEPTH=2 as the tree-depth limit for document-level selection (internal nodes at depth ≤ 2 are shown to the LLM during document routing) and OPENFABLE_RETRIEVAL_TOP_K=10 as the vector candidate count. The depth-decay coefficient itself is implemented in src/openfable/retrieval/tree_expansion.py per the source layout — I didn’t trace the exact formula, but the README’s S_sim with depth decay matches the architecture doc. The point is that the structure of the score is what matters: it’s not just “embed the node and rank,” it’s “embed the node, decay by depth, inherit from parents, aggregate from children, average.”

Where the agent comes in

This is the part the headline benchmarks can’t tell you. OpenFable’s two surfaces — the agent CLI and the REST/MCP server — split the LLM work deliberately:

  • Agent surface (docker exec openfable openfable plan/apply-chunks/apply-tree): the agent reads the document, identifies discourse boundaries, proposes the chunk markers, and builds the tree shape. The server stores what the agent produces. Index quality therefore depends on the model you use — the README recommends Opus 5 “or equivalent” and warns that a weaker model “degrades them without erroring” because there’s no validation step that checks whether a summary says what a subtree says versus merely naming it.
  • Server surface (openfable index or POST /v1/api/documents): the server runs the same pipeline in code at temperature=0, pinning the model. Reproducible but more expensive — every document needs multiple LLM calls for chunking and tree construction.

The architecture diagram in the README makes the dual-surface pattern explicit. Both surfaces call the same ingestion and retrieval services and share one PostgreSQL database. The agent is a client of the same services the REST API is a client of; the difference is who decides where chunks go.

The MCP server ships at /v1/mcp/sse. The README’s worked example uses mcp-use with a langchain-openai model:

client = MCPClient.from_dict({
    "mcpServers": {"openfable": {"url": "http://localhost:8000/v1/mcp/sse"}}
})
agent = MCPAgent(llm=ChatOpenAI(), client=client, max_steps=10)
print(await agent.run("Search the indexed documents: Who discovered the Eye of Kurak?"))

So the deployment shape is: a Docker Compose stack with PostgreSQL 17 + pgvector + a TEI server hosting bge-m3, an OpenFable FastAPI process exposing REST and MCP, and whatever MCP client you want pointing at it. Default embeddings are 1024-dimensional bge-m3 served by TEI in the container (the all-local template downloads ~2GB once and runs offline after); external embedding endpoints are accepted as long as they return 1024-dim vectors — text-embedding-3-small truncates to 1024, -3-large at 3072 doesn’t fit and will fail validation at onboarding rather than at first ingest.

Trade-offs and what it doesn’t fix

Three honest limits, one per category.

Cost at ingest. Every document requires multiple LLM calls: one for chunk-boundary identification, one per internal node for the summary that becomes the table-of-contents entry, plus one embed per node (root + every section + every leaf). For a 100-page technical document, that’s roughly 30-60 LLM calls during ingestion plus a comparable number of embed calls. The README lists this as a “not a fit” condition and is right to. OpenFable is for corpora you index once and query many times — the same economics as any tree-based retrieval, just sharper because the tree is LLM-generated rather than clustering-derived.

Latency at query. The agent path uses LLMselect (an LLM scores document relevance from shallow tree nodes) and LLMnavigate (an LLM picks subtree roots from the full tree). Each is a round-trip. A --vector-only flag exists that skips both, but then you’re back to the flat-vector-plus-tree-expansion ranking — better than flat alone but missing the LLM reasoning that distinguishes FABLE from RAPTOR. If your task needs sub-second retrieval, the LLM paths are a budget you have to plan for.

The agent-quality dependency. The README says it directly: “index quality therefore depends on the agent, and we recommend a capable frontier model — Opus 5 or equivalent. Validation rejects structurally invalid input … but it cannot tell whether the boundaries fall at genuine topic changes, or whether a summary states what a subtree says rather than merely naming it.” If you point a 7B model at the agent surface and get garbage trees, the server will dutifully persist the garbage, embed it, and serve it back. The system rejects malformed input but not bad judgment.

There’s also a fourth operational limit the README doesn’t list but the skill contract does: no authentication on the REST/MCP server. The README says “place it behind a reverse proxy with auth or API gateway.” That means the production deployment shape is OpenFable + something else — nginx with basic auth, Cloudflare Access, an authenticated MCP gateway — and whoever integrates it owns that choice. Treating OpenFable as a single-component drop-in will end in an exposed port. The MCP server at /v1/mcp/sse is the most exposed surface and the most attractive target — anyone who can reach it can query your indexed corpus.

What’s missing from the comparison

The README’s table compares FABLE to BM25, BGE-M3, and HippoRAG2, but it doesn’t include PageIndex even though PageIndex is the closest sibling in the vectorless-RAG lane (35.6k★, MIT, separate VectifyAI/PageIndex repo with its own PageIndex Flash release four days ago for fast tree indexing of text-based PDFs). Both projects argue for tree-aware retrieval over flat vectors; both use LLM reasoning at ingestion time; both run an LLM through document structure. The substantive difference, from what I can tell reading both READMEs side by side, is the agent boundary: PageIndex owns the chunking (you call its API, it builds the tree), OpenFable delegates the chunking to whichever MCP client is calling. That’s a meaningful boundary because the cost moves with it — PageIndex has a managed cloud SKU at a per-document price; OpenFable has no managed SKU at all, just an Apache-2.0 server you run yourself. If your team has a capable model available already and you want the index to live close to your agent loop, the OpenFable shape is cheaper. If your team doesn’t have a model available, PageIndex’s cloud option is the path of least friction.

The other gap is independent benchmarks on the OpenFable implementation versus the FABLE paper. The README is candid: “Independent benchmarks on OpenFable’s implementation are in progress — contributions welcome.” The paper’s numbers reflect the authors’ reference implementation; OpenFable is a fresh rewrite in Python 3.12 with FastAPI, SQLAlchemy, and pgvector + HNSW indexes. The arithmetic is the same but the runtime is different. Anyone considering OpenFable for production should treat the paper’s benchmarks as an upper bound and expect to run their own evaluation on a representative corpus before committing. The commit history shows the maintainers have been actively iterating on this (28 commits since April, with the last burst in late July through early August replacing the testcontainers integration suite with a leaner integration-test setup) but the benchmark reproducibility work is still open.

The honest caveat I’d add: I haven’t run OpenFable against a real corpus yet, and the “index quality depends on the agent” claim is the kind of statement that’s only verifiable after you’ve shipped it. The architecture is sound, the numbers are real-but-paper-derived, and the operational shape is unusual enough that I’d want to see a deployment that’s been running for six months before betting a production RAG pipeline on it. The README is doing the right thing by being explicit about all three.

The thing I keep thinking about is the boundary. Most RAG systems hide their chunking decisions behind a sensible-default API — you hand the engine a document, the engine picks a chunk size, you move on. OpenFable refuses to do that. The argument is that discourse boundaries are a judgment, not a parameter, and judgments belong with the model that will be asked to reason over the result. Whether that argument holds up is something only a production corpus can answer — but the willingness to ship a retrieval engine that won’t work until you’ve made those judgments yourself is a stance I haven’t seen in this lane before. PageIndex owns the chunking because that’s the contract its users want; OpenFable’s contract is closer to “you bring the discourse, we bring the structure.” Different shape, different cost, different failure mode.

The repo is alainbrown/openfable, 28 commits, Apache-2.0, and the maintainers are clearly working on it (the Aug 1 burst restructured the test suite, added v0.2.2, shipped the Claude Code plugin). The FABLE paper is arXiv:2601.18116. Whether it holds up in production is an open question, but the design choice is interesting enough to be worth watching.

References and where to dig further

  • alainbrown/openfable — https://github.com/alainbrown/openfable (Apache-2.0, v0.2.2)
  • FABLE paper — arXiv:2601.18116
  • skills/openfable/SKILL.md — the full operational contract (onboarding, both workflows, error strings, budget selection)
  • VectifyAI/PageIndex — https://github.com/VectifyAI/PageIndex (MIT, 35.6k★, separate vectorless-RAG project; PageIndex Flash shipped four days ago)
  • RAPTOR paper (the bottom-up clustering baseline) — Sarthi et al., 2024, “RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval”
  • HippoRAG2 — the multi-hop baseline in the FABLE benchmark table
Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.