TimesFM 3.0: Google’s decoder-only time-series foundation model goes multivariate and lands on MLX
The headline of google-research/timesfm v3.0.0 (tagged 2026-08-28) is the leaderboard sweep — 🥇 on fev-bench, 🥇 on TIME Benchmark, 🥇 on GIFT-Eval (among foundation models). The interesting bit for anyone deploying one of these is underneath the leaderboard: a patched-decoder transformer with stitching and CPM-RevIN refinement, an MLX backend that’s numerically matched to PyTorch to 1.7e-6 on the checkpoint, and a system-checker script that’s been promoted into a timesfm-forecasting/SKILL.md so an agent never crashes your laptop loading the weights.
This is also the first time-series foundation model I’ve seen that ships a working pip install timesfm[mlx] path that gives a real drop-in forecaster on Apple silicon — no PyTorch dependency required for the inference path. That’s a quietly bigger story than the benchmark numbers.
What’s actually new in 3.0
Three things, in order of how much they change the calling shape:
-
Native multivariate forecasting with dynamic covariates. The 2.5 line was univariate plus XReg for exogenous channels (the Oct 2025 patch). 3.0 takes a
(num_variates, context_length)target tensor and optional(1, context_length)past-only and(C, context_length + horizon)past-future covariate tensors, all in one call. The variate axis is real cross-channel attention, not a loop over independent univariate series —use_variate_attention=Trueis on by default in the new_ModelConfig(thetimesfm3.torch.configsdataclass insrc/timesfm3/torch/timesfm3_forecaster.py). -
Stitched patches with iterative CPM-RevIN refinement. Output patches are stitched into the next decoding round, and the running RevIN statistics are re-estimated at CPM-masked positions using the model’s own predictions for preceding patches. The refinement lives in
src/timesfm3/torch/cpm_revin_refine.py(cpm_iterative_revin_refine) and is on by default (use_iterative_cpm_revin=True,use_frozen_running_stats=False). It’s also why the README’s benchmark numbers look the way they do — RevIN’s classic failure mode at long horizons is mean drift, and the iterative estimate folds that error back into the running stats. -
MLX backend.
from timesfm3.mlx import TimesFM3Forecastermirrors the PyTorchpredict/predict_batchinterface, runs on Apple silicon, and the README’s accuracy table reports median forecast / quantile max abs error at context 512 horizon 64 as9.5e-7 / 1.8e-6, and horizon 128 as2.3e-6 / 2.7e-6. The MLX path skips PyTorch entirely (pip install timesfm[mlx]). The throughput table on an M4 Max withmx.compileandfp32is interesting on its own:
| batch | p50 latency | throughput |
|---|---|---|
| 1 | 11.1 ms | 90 series/s |
| 8 | 19.7 ms | 406 series/s |
| 32 | 48.1 ms | 666 series/s |
A 330M-param model at 666 series/s on an M4 Max with 48-context-length batching — that’s not LLM throughput, it’s classical forecasting throughput. If you’ve ever stared at a statsmodels.tsa loop and wished someone would just give you the numbers, this is what that wish looks like.
The decoder, in one paragraph
TimesFM is a patched decoder in the same lineage as the original 2023 paper (“A decoder-only foundation model for time-series forecasting,” Das, Kong, Sen, Zhou — arXiv:2310.10688, ICML 2024). Time is sliced into input_patch_length=32 contiguous subwindows. Each patch is embedded twice (as a “target” token and as a “context” token with different positional role) and concatenated into a 2D input of shape (b, 2*(in_patch + out_patch), d_model) — see input_dim = 2 * (t_model.input_patch_len + t_model.output_patch_len) in _make_torch_model. The transformer is a stack of 20 RMS-norm pre-norm blocks with multi-head QK-RoPE attention (MultiHeadAttention in src/timesfm3/torch/transformer.py) and a PerDimScale replacement of the standard 1/√d scaling — PerDimScale (src/timesfm3/torch/normalization.py) keeps a learnable per-dim scale initialized at zero (so softplus(0) ≈ 0.693, net scale ≈ 1/√d) and lets training nudge it. Output heads emit num_quantiles=9 deciles per output patch — q=0.1..0.9, with the median head used as the point forecast.
What distinguishes this from a vanilla Transformer is the stitching + iterative refinement loop during decoding. The output patch isn’t just emitted and discarded — it’s fed back as input context for the next patch, with RevIN statistics recomputed on the fly. use_stitching=True and use_linear_detrending=True (with linear_detrending_threshold=0.5) are the default flags that combine to give you the long-horizon stability. The point forecast is what stitching buys you: at horizon 64 it works fine; at horizon 128 and beyond, multiple output patches get stitched together and the quantile crossing is fixed (fix_quantile_crossing=True in the legacy ForecastConfig; the 3.0 API uses sort_quantiles=True on the predict call).
What’s actually hard on the edge
Three numbers that tell you what 3.0 is paying for:
- 200M parameters in 2.5 (down from 500M in 2.0 and the v1 200M). On disk: ~800 MB. In RAM: ~1.5 GB CPU, ~1 GB GPU. On Apple silicon with unified memory: ~1.5 GB.
global_context=15360is the hard max — rounded up to the nearest input patch boundary. Contexts longer than 15,360 are truncated to the most recent points before decode. Anything you can hand astatsmodels.tsa.ARIMAyou can hand TimesFM.- Continuous quantile head up to 1k horizon via an optional 30M-parameter head. The legacy 2.5 release added this for probabilistic forecasting at long horizons — the 3.0 default is the 9-decile head, but you can opt into continuous quantile head for any-time-step quantile interpolation.
The hard part isn’t the model. It’s the system check. The timesfm-forecasting/scripts/check_system.py script has a MODEL_PROFILES dict with per-version RAM/VRAM/disk minimums — 2.5 needs 2 GB RAM (min) / 4 GB (recommended) and 2 GB disk; v2.0 needs 8 GB RAM min / 16 GB recommended and 4 GB disk. The script has cross-platform RAM detection (Linux /proc/meminfo, Darwin sysctl hw.memsize and vm_stat, Windows GlobalMemoryStatusEx), and it blocks the load if you don’t meet the minimum. The README’s “Update — March 19, 2026” line credits @borealBytes for adding the AGENTS.md integration that drives the SKILL.md. That’s the part I want you to internalize: this is the first serious time-series foundation model that ships with a built-in agent harness for “make sure I don’t OOM the user’s machine.”
The mechanism, in one equation
For the MLX numerical match to PyTorch, what they actually claim is:
forecast_max_abs_err_mlx_vs_torch(context=512, horizon=64) = 9.5e-7
quantile_max_abs_err_mlx_vs_torch(context=512, horizon=64) = 1.8e-6
forecast_max_abs_err_mlx_vs_torch(context=512, horizon=128) = 2.3e-6
quantile_max_abs_err_mlx_vs_torch(context=512, horizon=128) = 2.7e-6
That’s checkpoint-bit-level, not “approximately the same answer.” The mechanism is what src/timesfm3/__init__.py does: lazy re-export of the PyTorch names through __getattr__ (PEP 562) so from timesfm3 import TimesFM3Forecaster doesn’t import torch when you only want the MLX backend, and explicit cross-backend config validation in TimesFM3Forecaster._init_model that re-reads the loaded model’s residual_block_config and transformer_config and overwrites the dataclass with the values from the checkpoint. If you load a 256-dim variant into a forecaster you compiled with 1280-dim defaults, the forecaster silently rewrites itself to match the weights. That’s the same trick vLLM uses to handle model-card-vs-engine mismatch, applied to a 200M-parameter forecaster that fits in a laptop’s RAM.
What the numbers actually show
The headline numbers from the README’s “Update — August 2026” block:
- 🥇 fev-bench — rank #1 across 100 real-world forecasting tasks. (fev-bench is the Salesforce time-series eval suite; the leaderboard splits by model size class.)
- 🥇 TIME Benchmark — rank #1 across 50 domain datasets and 98 evaluation tasks.
- 🥇 GIFT-Eval — rank #1 among foundation models.
Cross-checked against the Apple-MLX throughput table above: 666 series/s on an M4 Max with batch=32. That’s a meaningful number if you’ve ever tried to score a backtest of a thousand stock tickers through a transformer — Chronos and TimeGPT are 1–10 series/s on the same hardware, last I checked. TimesFM 3.0 is in a different latency class because it’s 200M params with no autoregressive decoding — the output patch is emitted in one forward pass per patch.
The honest comparison against Chronos-Bolt and TimeGPT isn’t apples-to-apples — those models use different output heads, different context lengths, different quantile grids — but the order-of-magnitude difference in throughput is real. The accuracy comparison is closer; GIFT-Eval and fev-bench are the two suites where you actually see the leaderboard numbers converge.
Trade-offs and what it doesn’t fix
Three honest limits.
The license split on v3.0 weights. The repository is Apache-2.0. The 2.5 weights are Apache-2.0. The 3.0 weights are timesfm-non-commercial-license-v1.0 — restricted to non-commercial, non-production use. The README quotes this explicitly:
For the time being, TimesFM 3.0 pretrained weights are distributed under the separate
timesfm-non-commercial-license-v1.0license and are restricted to non-commercial, non-production use. Commercial or production use of the default pretrained weights is not permitted.
If you want to ship 3.0 in a paid product, you need to either fine-tune the Apache-2.0 2.5 weights on your own data (the timesfm-forecasting/examples/finetuning/ directory has a LoRA via HuggingFace Transformers + PEFT example — added in the April 2026 update) or negotiate a separate license with Google. This is not a footnote; it’s the single biggest production-readiness issue in the release.
The 3.0 API has a backward-compatibility shim, not a backport. src/timesfm3/__init__.py re-exports the PyTorch backend at the top level for backward compatibility, but every file in the timesfm3 top-level package is now a one-line from .torch.X import * shim:
"""Backward-compatibility shim for ``timesfm3.transformer``.
The PyTorch backend moved to ``timesfm3.torch.transformer``."""
from .torch.transformer import * # noqa: F401,F403
If you from timesfm3.transformer import MultiHeadAttention you’re getting the PyTorch implementation. If you want the MLX one, you must import it from timesfm3.mlx explicitly. The two backends share timesfm3/configs.py’s dataclasses but the model classes are different. This is fine — it’s the right factoring — but any code written against the 2.5 API will silently pull the torch backend unless you rewrite the imports. The README’s “Backward compatibility shim” comment is honest about this, but a deployer who skims the release notes and pip install timesfm[torch] will get torch whether they wanted it or not.
CPM-RevIN refinement is a single-pass loop with hardcoded assumptions about mask topology. The cpm_iterative_revin_refine function in cpm_revin_refine.py assumes a specific block structure — contiguous CPM-masked patches get refined with estimates of preceding patches in the same block, and earlier blocks’ estimates cascade forward via block_offset. If your dataset has interleaved observed and CPM-masked patches (think: a sensor that drops out and comes back, where you want to forecast through the dropout but trust the post-dropout signal), the refinement loop handles that via make_segment_mask (segment-aware masking) but the RevIN stats are still carried across segment boundaries via the carry state. In practice this works fine for sensor and price data; it can be subtle for sparse covariates where segment boundaries don’t line up with CPM boundaries. There’s a tests/ directory and a backward_compat_test.py, but I didn’t find a published benchmark that exercises the refinement loop on heavy-interleaved data.
What changed from 2.5 → 3.0 in one line
pip install timesfm[mlx] is the answer for anyone on Apple silicon. The MLX backend is real, not a stub; it’s numerically matched to PyTorch to 1.7e-6 on the checkpoint; it gives you 666 series/s on an M4 Max with batch=32. That’s the headline for individual developers. For production teams the headline is that the v3.0 weights are non-commercial — and the leaderboard sweep is real but the path to shipping it is fine-tune-2.5-on-your-data or get a license.
What’s next, and what isn’t
The repo pushed 2026-09-07 (yesterday relative to this post). The v3.0.0 tag is 2026-08-28. The README says “Google Research blog (New blog post for TimesFM 3.0 coming soon!)” — there’s no published companion blog post at the time of writing. The timesfm-forecasting/SKILL.md is the most novel piece of operational infrastructure; I haven’t seen a comparable time-series foundation model ship with a built-in agent skill that handles the system preflight for you. The combination — pretrained model + MLX backend + agent skill + system-checker script + CSV ingestion script + four worked examples + LoRA fine-tuning example — is unusually complete.
What isn’t next: I don’t see a 7B-class TimesFM in the release cadence. The 2.5→3.0 jump is purely feature-completeness (multivariate + covariates + MLX + leaderboard wins), not scale. Compare against the language-model cadence where 2.5→3.0 would mean 4× the parameter count. TimesFM is staying in the small-model regime on purpose — the entire system is sized for a laptop.
References and where to dig further
google-research/timesfm— v3.0.0 tag, 2026-08-28. Apache-2.0 source, non-commercial v3.0 weights.- Das, Kong, Sen, Zhou, “A decoder-only foundation model for time-series forecasting,” ICML 2024. The patched-decoder origin paper.
google/timesfm-3.0-pytorch— v3.0 weights on HuggingFace, non-commercial license.timesfm-forecasting/SKILL.md— the agent harness. Mandatory preflight, dataset fit check, CSV/DataFrame/array input handling, prediction intervals.timesfm-forecasting/scripts/check_system.py— cross-platform RAM/VRAM/disk detector. Linux/proc/meminfo, Darwinsysctl/vm_stat, WindowsGlobalMemoryStatusEx. This is the file that gets you out of “I crashed my laptop loading the model” territory.src/timesfm3/torch/transformer.py—MultiHeadAttentionwith QK-RoPE, QK-RMSNorm, PerDimScale, KV-cache, and therescale_logitsflag that toggles between Flax memory-efficient-attention behavior and standard SDPA.src/timesfm3/torch/cpm_revin_refine.py—cpm_iterative_revin_refine, the function that makes long-horizon mean drift go away. Worth reading for theblock_offsetcarry state alone.src/timesfm3/__init__.py— PEP 562 lazy re-export of the PyTorch backend. This is whyfrom timesfm3.mlx import TimesFM3Forecasterdoesn’t pull torch.- BigQuery ML
timesfmmodel — Google Cloud managed serving, if you don’t want to run it yourself. - Google Sheets forecast in connected sheets — same model exposed as a Sheets function.
- Vertex Model Garden TimesFM endpoint — Dockerized endpoint for agentic calling.
The thing I’d want to see next is a published comparison on the same hardware between TimesFM 3.0 (MLX) and Chronos-Bolt (MLX or PyTorch CPU) on fev-bench — the leaderboard numbers from TimesFM’s own release don’t include side-by-side MLX throughput for the competitors. If Google publishes that, the MLX-backend story stops being a footnote and becomes the deciding factor for anyone on Apple silicon.
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.