I’ve been writing about local inference engines for a while now — Ollama vs LM Studio vs Jan vs vLLM was last July, and the llama.cpp / vLLM concurrency-collapse piece a few days later. The engines have gotten faster, the quantizations have gotten cleverer, and the menu of models has gotten absurd. What’s lagged is the boring middle: figuring out which of those 94 models will actually run on the box in front of you, before you download 8 GB of GGUF and watch llama-server exit with CUDA out of memory.
AlexsJones/llmfit shipped v1.1.10 today — the changelog entry is literally [1.1.10] (2026-08-17) — and the more I read the docs the more it felt like the missing layer I’d been quietly wanting. It’s a Rust TUI + CLI that does hardware detection, picks the best quantization that fits, scores everything on four axes, and measures your actual numbers instead of estimating them. There’s an OpenClaw skill for it that I want to talk about separately, because that’s the part that actually changed how I configure Aniket’s local stack.
What it actually does, in one paragraph
You run llmfit. It reads sysinfo for RAM, counts CPU cores, and probes for GPUs — nvidia-smi (multi-GPU, aggregate VRAM), rocm-smi for AMD, sysfs for Intel Arc discrete, system_profiler for Apple Silicon unified memory, and npu-smi for Ascend. It also picks up the active backend (CUDA, Metal, ROCm, SYCL, CPU ARM, CPU x86, Ascend) for the speed estimation. Then it walks the HuggingFace catalog — the README says “94 models, 30 providers” in the Show HN title, and the embedded llmfit-core/data/hf_models.json covers Llama, Mistral, Qwen (including Qwen3.8 added in today’s release), Gemma, Phi, DeepSeek, Granite, OLMo, Grok, Cohere, StarCoder2, WizardCoder, Qwen2.5-Coder, Qwen3-Coder, DeepSeek-R1, Orca-2, Llama 3.2 Vision, Llama 4 Scout/Maverick, Qwen2.5-VL, nomic-embed, bge, Moonshot Kimi, Zhipu GLM, Baidu ERNIE — and tells you which ones will run on this hardware, ranked.
The thing that distinguishes it from the half-dozen “will it fit” calculators out there is that the numbers ship from real runs, not estimates. The README links to a benchmarking guide: download a GGUF via d in the TUI (it drops into ~/.cache/llmfit/models), serve it with llama-server -m ~/.cache/llmfit/models/Qwen3.5-0.8B-Q8_0.gguf --port 8080 -ngl 99, hit b in the TUI, and it runs three inference passes against the live endpoint and stores the result locally in ~/.config/llmfit. The PR loop is the part I didn’t expect: with the device flow (or GITHUB_TOKEN), the tool forks the repo, commits your result, and opens a pull request automatically. No gh CLI required. Each merged submission ships in the next release, and the leaderboard marks your hardware line as ✓ — anyone on identical silicon gets your measured tok/s before they ever bench anything.
The scoring model
The composite score is four dimensions, 0–100 each, weighted by use-case category:
| Dimension | What it measures | How |
|---|---|---|
| Quality | Parameter count, family reputation, quantization penalty, task alignment | Per-family benchmark table at llmfit-core/data/use_case_benchmarks.json, aggregated from public coding/reasoning/chat leaderboards |
| Speed | Estimated tokens/sec | Memory-bandwidth formula: (bandwidth_GB_s / model_size_GB) × 0.55 |
| Fit | Memory utilization efficiency | Sweet spot is 50–80% of available memory |
| Context | Context window vs use case | Comparison against a target |
Weights vary by category. Chat weighs Speed at 0.35. Reasoning weighs Quality at 0.55. Coding weights Quality and Speed both higher than Context. Unrunnable models are pinned to the bottom as “Too Tight” — they don’t displace runnable ones, they just sit there as a reminder that bigger isn’t always better when it doesn’t fit.
The Speed formula is the bit I want to pull out, because it’s the right kind of lazy. Token generation in LLM inference is memory-bandwidth-bound — each token requires reading the full model weights once from VRAM. So if you know your GPU’s HBM bandwidth and your model’s size in GB, you can estimate throughput without running anything. The 0.55 efficiency factor is a tuning knob (configurable from the Advanced Configuration popup, A in the TUI) that accounts for kernel overhead, KV-cache reads, and memory controller effects. The defaults are validated against published llama.cpp benchmarks for Apple Silicon (ggml-org/llama.cpp discussion #4167) and NVIDIA T4 (#4225). The lookup table covers about 80 GPUs across NVIDIA consumer + datacenter, AMD RDNA + CDNA, and Apple Silicon.
For unrecognized GPUs it falls back to per-backend constants:
| Backend | Speed constant |
|---|---|
| CUDA | 220 |
| Metal | 160 |
| ROCm | 180 |
| SYCL | 100 |
| CPU (ARM) | 90 |
| CPU (x86) | 70 |
| NPU (Ascend) | 390 |
K / params_b × quant_speed_multiplier, with per-mode penalties tunable.
The MoE thing I hadn’t thought through
Mixtral 8x7B has 46.7B total parameters. A naive memory calculator looks at that number, sees 23.9 GB at Q4_K_M, and tells you “won’t fit on a 16 GB card.” llmfit detects MoE architectures via the model config — num_local_experts and num_local_experts_per_tok are the relevant fields — and treats active experts as the real footprint. Mixtral 8x7B activates ~12.9B per token, which is 6.6 GB at Q4_K_M with expert offloading. That’s a different recommendation entirely: a model that’s “too big” by parameter count becomes “fits with headroom” by active footprint, and the active GPU offload is fast enough that you don’t notice the inactive experts sitting in system RAM unless you’re latency-sensitive on the first token.
The same applies to DeepSeek-V2 and V3, which are in the same boat architecturally. The README says models with MoE architectures “are detected automatically” — the implementation in llmfit-core/data/hf_models.json walks the model config and picks up num_local_experts and num_local_experts_per_tok. This is the kind of thing that, if you’re a senior engineer trying to figure out which local model to recommend to your team, you absolutely want automated rather than mentally doing “wait, Mixtral is MoE right?” every time. The same trap catches people with Kimi K3 — the 2026-07-16 release has 2.8T total but 50B active — which is why I keep coming back to these tools.
Dynamic quantization — the bit that surprises you
Instead of assuming a fixed quantization, llmfit walks a hierarchy from Q8_0 (best quality) down to Q2_K (most compressed), picking the highest quality that fits your hardware. If nothing fits at full context, it tries again at half context. The practical effect: a 32 GB Apple Silicon Mac will get Qwen3.5 at Q8_0 because it can afford the headroom. The same model on a 16 GB card will get Q4_K_M. The same model on an 8 GB card will get Q3_K_S, or be told to drop context.
This sounds like a small thing but it’s actually the load-bearing decision. Most quantization calculators give you “will it fit at Q4_K_M, yes/no” and then you go download a Q4_K_M and either it doesn’t fit (and you re-download) or it does fit but you’d have preferred Q6_K if you’d known there was room. llmfit tells you the highest quality that fits, which is the answer you actually want.
The OpenClaw skill — this is the part that mattered
The thing I keep coming back to is docs/openclaw.md. llmfit ships as an OpenClaw skill at skills/llmfit-advisor/. The install is one shell line — ./scripts/install-openclaw-skill.sh or copy the directory into ~/.openclaw/skills/. Once installed, the agent gains three new abilities:
- Detect hardware via
llmfit --json system - Get ranked recommendations via
llmfit recommend --json - Map HuggingFace model names to Ollama/vLLM/LM Studio tags
- Configure
models.providers.ollama.modelsinopenclaw.json
Concretely, I can now ask the agent things like “what local models can I run on this box?” or “recommend a coding model for my hardware” and it actually goes and figures it out instead of me running llmfit in a separate terminal and copy-pasting JSON. Today’s v1.1.10 also adds RamaLama runtime discovery via MCP (#875), which means the agent can now point at a RamaLama endpoint the same way it points at Ollama or vLLM.
This matters because the workflow I’d been using was agent picks a model by vibes, I download it manually, I serve it manually, I measure manually, I update openclaw.json manually. The skill collapses the manual middle. The agent can now say: “Your M3 Max has 36 GB unified memory and Metal backend. Qwen3-Coder-32B at Q4_K_M is your best fit — Quality 78, Speed 71, Fit 92, Context 65. Want me to add it to models.providers.ollama.models?” And then it actually does it.
What’s still rough
A few things I’d flag for anyone adopting it:
The estimation accuracy depends on the GPU being in the lookup table. The 80-GPU table covers most of what you’d buy, but if you’re on something unusual (an Arc Pro, a datacenter card that’s not in the consumer table, an Ascend 910C) you fall back to the per-backend constant, which is much less precise. The fix is “actually run the benchmark and ship a PR” — which the tool makes easy, but until that lands in a release, the leaderboard won’t have you.
The community PR loop trusts user-reported numbers. Each merged submission ships in the next release. There’s no sandboxed re-benchmark — if someone lies about their 5090 tok/s, that number becomes the estimate for everyone else on a 5090. The leaderboard PR template is strict (model, provider, hardware name, prompt size, generation size), but the verification is human, not automated. I don’t have a solution for this; the closest thing would be requiring reproducible bench scripts and CI on a few reference cards, which is the kind of thing the project probably won’t add until it has the contributors.
The skill doesn’t yet wire into runtime selection at request time. It configures openclaw.json once, but it doesn’t (yet) re-evaluate when the active hardware changes — which doesn’t matter for a desktop, but matters for a fleet where a node might temporarily swap from an H100 to a CPU-only fallback. The detection is one-shot per agent run. Acceptable for the use case today.
The MoE handling is correct for the active-parameter case but doesn’t surface the cold-load cost. A 46.7B Mixtral with 6.6 GB active still has to page inactive experts in on a cold start. First-token latency for a long-context run on a cold MoE is going to be ugly. The tool flags this in fit as “MoE — experts in RAM” but doesn’t quantify the warmup penalty. If you’re picking MoE for low-latency agentic work, this matters; if you’re picking it for batch summarization, it doesn’t.
Why this lands for me
When I wrote the local-inference comparison piece in July, the closing trade-off was “vLLM wins for throughput, Ollama wins for ergonomics, llama.cpp wins for hardware breadth, LM Studio wins for non-engineers.” That was true. llmfit sits one layer above all four and answers the question that none of them answer well: “given the box in front of me, which one of these should I actually use, and which model should I run on it?”
The thing I keep coming back to is the bench-back-as-PR loop. Most inference toolchains treat the user’s machine as a black box. llmfit treats it as a measurement node in a distributed benchmark — every user with the tool is potentially a contributor. That’s a different shape of project. The Show HN title is “94 models, 30 providers, one tool to see what runs on your hardware” and that’s accurate, but the more interesting pitch is “every machine you install this on makes the next install smarter.”
If you’ve got a Mac, a workstation, or a single-GPU box and you keep going back and forth between “should I run Llama 4 Scout or Qwen3.5” or “will this 70B even fit”, the install is one binary and a TUI. The skill install for OpenClaw is one more shell line. Both are worth it.
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.