TimesFM 3.0: Google's decoder-only time-series foundation model goes multivariate and lands on MLX — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

TimesFM 3.0: Google's decoder-only time-series foundation model goes multivariate and lands on MLX

A 200M-parameter decoder that ships native multivariate forecasting, past-only and past-future covariates, an MLX backend numerically matched to PyTorch to 1.7e-6, and a top rank on fev-bench, TIME Benchmark, and GIFT-Eval. The interesting bit is the patched-decoder-with-stitching design and what the CPM-RevIN refinement loop is actually doing.

TimesFM 3.0: Google’s decoder-only time-series foundation model goes multivariate and lands on MLX

The headline of google-research/timesfm v3.0.0 (tagged 2026-08-28) is the leaderboard sweep — 🥇 on fev-bench, 🥇 on TIME Benchmark, 🥇 on GIFT-Eval (among foundation models). The interesting bit for anyone deploying one of these is underneath the leaderboard: a patched-decoder transformer with stitching and CPM-RevIN refinement, an MLX backend that’s numerically matched to PyTorch to 1.7e-6 on the checkpoint, and a system-checker script that’s been promoted into a timesfm-forecasting/SKILL.md so an agent never crashes your laptop loading the weights.

This is also the first time-series foundation model I’ve seen that ships a working pip install timesfm[mlx] path that gives a real drop-in forecaster on Apple silicon — no PyTorch dependency required for the inference path. That’s a quietly bigger story than the benchmark numbers.

What’s actually new in 3.0

Three things, in order of how much they change the calling shape:

  1. Native multivariate forecasting with dynamic covariates. The 2.5 line was univariate plus XReg for exogenous channels (the Oct 2025 patch). 3.0 takes a (num_variates, context_length) target tensor and optional (1, context_length) past-only and (C, context_length + horizon) past-future covariate tensors, all in one call. The variate axis is real cross-channel attention, not a loop over independent univariate series — use_variate_attention=True is on by default in the new _ModelConfig (the timesfm3.torch.configs dataclass in src/timesfm3/torch/timesfm3_forecaster.py).

  2. Stitched patches with iterative CPM-RevIN refinement. Output patches are stitched into the next decoding round, and the running RevIN statistics are re-estimated at CPM-masked positions using the model’s own predictions for preceding patches. The refinement lives in src/timesfm3/torch/cpm_revin_refine.py (cpm_iterative_revin_refine) and is on by default (use_iterative_cpm_revin=True, use_frozen_running_stats=False). It’s also why the README’s benchmark numbers look the way they do — RevIN’s classic failure mode at long horizons is mean drift, and the iterative estimate folds that error back into the running stats.

  3. MLX backend. from timesfm3.mlx import TimesFM3Forecaster mirrors the PyTorch predict/predict_batch interface, runs on Apple silicon, and the README’s accuracy table reports median forecast / quantile max abs error at context 512 horizon 64 as 9.5e-7 / 1.8e-6, and horizon 128 as 2.3e-6 / 2.7e-6. The MLX path skips PyTorch entirely (pip install timesfm[mlx]). The throughput table on an M4 Max with mx.compile and fp32 is interesting on its own:

batchp50 latencythroughput
111.1 ms90 series/s
819.7 ms406 series/s
3248.1 ms666 series/s

A 330M-param model at 666 series/s on an M4 Max with 48-context-length batching — that’s not LLM throughput, it’s classical forecasting throughput. If you’ve ever stared at a statsmodels.tsa loop and wished someone would just give you the numbers, this is what that wish looks like.

The decoder, in one paragraph

TimesFM is a patched decoder in the same lineage as the original 2023 paper (“A decoder-only foundation model for time-series forecasting,” Das, Kong, Sen, Zhou — arXiv:2310.10688, ICML 2024). Time is sliced into input_patch_length=32 contiguous subwindows. Each patch is embedded twice (as a “target” token and as a “context” token with different positional role) and concatenated into a 2D input of shape (b, 2*(in_patch + out_patch), d_model) — see input_dim = 2 * (t_model.input_patch_len + t_model.output_patch_len) in _make_torch_model. The transformer is a stack of 20 RMS-norm pre-norm blocks with multi-head QK-RoPE attention (MultiHeadAttention in src/timesfm3/torch/transformer.py) and a PerDimScale replacement of the standard 1/√d scaling — PerDimScale (src/timesfm3/torch/normalization.py) keeps a learnable per-dim scale initialized at zero (so softplus(0) ≈ 0.693, net scale ≈ 1/√d) and lets training nudge it. Output heads emit num_quantiles=9 deciles per output patch — q=0.1..0.9, with the median head used as the point forecast.

What distinguishes this from a vanilla Transformer is the stitching + iterative refinement loop during decoding. The output patch isn’t just emitted and discarded — it’s fed back as input context for the next patch, with RevIN statistics recomputed on the fly. use_stitching=True and use_linear_detrending=True (with linear_detrending_threshold=0.5) are the default flags that combine to give you the long-horizon stability. The point forecast is what stitching buys you: at horizon 64 it works fine; at horizon 128 and beyond, multiple output patches get stitched together and the quantile crossing is fixed (fix_quantile_crossing=True in the legacy ForecastConfig; the 3.0 API uses sort_quantiles=True on the predict call).

What’s actually hard on the edge

Three numbers that tell you what 3.0 is paying for:

  • 200M parameters in 2.5 (down from 500M in 2.0 and the v1 200M). On disk: ~800 MB. In RAM: ~1.5 GB CPU, ~1 GB GPU. On Apple silicon with unified memory: ~1.5 GB.
  • global_context=15360 is the hard max — rounded up to the nearest input patch boundary. Contexts longer than 15,360 are truncated to the most recent points before decode. Anything you can hand a statsmodels.tsa.ARIMA you can hand TimesFM.
  • Continuous quantile head up to 1k horizon via an optional 30M-parameter head. The legacy 2.5 release added this for probabilistic forecasting at long horizons — the 3.0 default is the 9-decile head, but you can opt into continuous quantile head for any-time-step quantile interpolation.

The hard part isn’t the model. It’s the system check. The timesfm-forecasting/scripts/check_system.py script has a MODEL_PROFILES dict with per-version RAM/VRAM/disk minimums — 2.5 needs 2 GB RAM (min) / 4 GB (recommended) and 2 GB disk; v2.0 needs 8 GB RAM min / 16 GB recommended and 4 GB disk. The script has cross-platform RAM detection (Linux /proc/meminfo, Darwin sysctl hw.memsize and vm_stat, Windows GlobalMemoryStatusEx), and it blocks the load if you don’t meet the minimum. The README’s “Update — March 19, 2026” line credits @borealBytes for adding the AGENTS.md integration that drives the SKILL.md. That’s the part I want you to internalize: this is the first serious time-series foundation model that ships with a built-in agent harness for “make sure I don’t OOM the user’s machine.”

The mechanism, in one equation

For the MLX numerical match to PyTorch, what they actually claim is:

forecast_max_abs_err_mlx_vs_torch(context=512, horizon=64)  = 9.5e-7
quantile_max_abs_err_mlx_vs_torch(context=512, horizon=64)  = 1.8e-6
forecast_max_abs_err_mlx_vs_torch(context=512, horizon=128) = 2.3e-6
quantile_max_abs_err_mlx_vs_torch(context=512, horizon=128) = 2.7e-6

That’s checkpoint-bit-level, not “approximately the same answer.” The mechanism is what src/timesfm3/__init__.py does: lazy re-export of the PyTorch names through __getattr__ (PEP 562) so from timesfm3 import TimesFM3Forecaster doesn’t import torch when you only want the MLX backend, and explicit cross-backend config validation in TimesFM3Forecaster._init_model that re-reads the loaded model’s residual_block_config and transformer_config and overwrites the dataclass with the values from the checkpoint. If you load a 256-dim variant into a forecaster you compiled with 1280-dim defaults, the forecaster silently rewrites itself to match the weights. That’s the same trick vLLM uses to handle model-card-vs-engine mismatch, applied to a 200M-parameter forecaster that fits in a laptop’s RAM.

What the numbers actually show

The headline numbers from the README’s “Update — August 2026” block:

  • 🥇 fev-bench — rank #1 across 100 real-world forecasting tasks. (fev-bench is the Salesforce time-series eval suite; the leaderboard splits by model size class.)
  • 🥇 TIME Benchmark — rank #1 across 50 domain datasets and 98 evaluation tasks.
  • 🥇 GIFT-Eval — rank #1 among foundation models.

Cross-checked against the Apple-MLX throughput table above: 666 series/s on an M4 Max with batch=32. That’s a meaningful number if you’ve ever tried to score a backtest of a thousand stock tickers through a transformer — Chronos and TimeGPT are 1–10 series/s on the same hardware, last I checked. TimesFM 3.0 is in a different latency class because it’s 200M params with no autoregressive decoding — the output patch is emitted in one forward pass per patch.

The honest comparison against Chronos-Bolt and TimeGPT isn’t apples-to-apples — those models use different output heads, different context lengths, different quantile grids — but the order-of-magnitude difference in throughput is real. The accuracy comparison is closer; GIFT-Eval and fev-bench are the two suites where you actually see the leaderboard numbers converge.

Trade-offs and what it doesn’t fix

Three honest limits.

The license split on v3.0 weights. The repository is Apache-2.0. The 2.5 weights are Apache-2.0. The 3.0 weights are timesfm-non-commercial-license-v1.0 — restricted to non-commercial, non-production use. The README quotes this explicitly:

For the time being, TimesFM 3.0 pretrained weights are distributed under the separate timesfm-non-commercial-license-v1.0 license and are restricted to non-commercial, non-production use. Commercial or production use of the default pretrained weights is not permitted.

If you want to ship 3.0 in a paid product, you need to either fine-tune the Apache-2.0 2.5 weights on your own data (the timesfm-forecasting/examples/finetuning/ directory has a LoRA via HuggingFace Transformers + PEFT example — added in the April 2026 update) or negotiate a separate license with Google. This is not a footnote; it’s the single biggest production-readiness issue in the release.

The 3.0 API has a backward-compatibility shim, not a backport. src/timesfm3/__init__.py re-exports the PyTorch backend at the top level for backward compatibility, but every file in the timesfm3 top-level package is now a one-line from .torch.X import * shim:

"""Backward-compatibility shim for ``timesfm3.transformer``.
The PyTorch backend moved to ``timesfm3.torch.transformer``."""
from .torch.transformer import *  # noqa: F401,F403

If you from timesfm3.transformer import MultiHeadAttention you’re getting the PyTorch implementation. If you want the MLX one, you must import it from timesfm3.mlx explicitly. The two backends share timesfm3/configs.py’s dataclasses but the model classes are different. This is fine — it’s the right factoring — but any code written against the 2.5 API will silently pull the torch backend unless you rewrite the imports. The README’s “Backward compatibility shim” comment is honest about this, but a deployer who skims the release notes and pip install timesfm[torch] will get torch whether they wanted it or not.

CPM-RevIN refinement is a single-pass loop with hardcoded assumptions about mask topology. The cpm_iterative_revin_refine function in cpm_revin_refine.py assumes a specific block structure — contiguous CPM-masked patches get refined with estimates of preceding patches in the same block, and earlier blocks’ estimates cascade forward via block_offset. If your dataset has interleaved observed and CPM-masked patches (think: a sensor that drops out and comes back, where you want to forecast through the dropout but trust the post-dropout signal), the refinement loop handles that via make_segment_mask (segment-aware masking) but the RevIN stats are still carried across segment boundaries via the carry state. In practice this works fine for sensor and price data; it can be subtle for sparse covariates where segment boundaries don’t line up with CPM boundaries. There’s a tests/ directory and a backward_compat_test.py, but I didn’t find a published benchmark that exercises the refinement loop on heavy-interleaved data.

What changed from 2.5 → 3.0 in one line

pip install timesfm[mlx] is the answer for anyone on Apple silicon. The MLX backend is real, not a stub; it’s numerically matched to PyTorch to 1.7e-6 on the checkpoint; it gives you 666 series/s on an M4 Max with batch=32. That’s the headline for individual developers. For production teams the headline is that the v3.0 weights are non-commercial — and the leaderboard sweep is real but the path to shipping it is fine-tune-2.5-on-your-data or get a license.

What’s next, and what isn’t

The repo pushed 2026-09-07 (yesterday relative to this post). The v3.0.0 tag is 2026-08-28. The README says “Google Research blog (New blog post for TimesFM 3.0 coming soon!)” — there’s no published companion blog post at the time of writing. The timesfm-forecasting/SKILL.md is the most novel piece of operational infrastructure; I haven’t seen a comparable time-series foundation model ship with a built-in agent skill that handles the system preflight for you. The combination — pretrained model + MLX backend + agent skill + system-checker script + CSV ingestion script + four worked examples + LoRA fine-tuning example — is unusually complete.

What isn’t next: I don’t see a 7B-class TimesFM in the release cadence. The 2.5→3.0 jump is purely feature-completeness (multivariate + covariates + MLX + leaderboard wins), not scale. Compare against the language-model cadence where 2.5→3.0 would mean 4× the parameter count. TimesFM is staying in the small-model regime on purpose — the entire system is sized for a laptop.

References and where to dig further

  • google-research/timesfm — v3.0.0 tag, 2026-08-28. Apache-2.0 source, non-commercial v3.0 weights.
  • Das, Kong, Sen, Zhou, “A decoder-only foundation model for time-series forecasting,” ICML 2024. The patched-decoder origin paper.
  • google/timesfm-3.0-pytorch — v3.0 weights on HuggingFace, non-commercial license.
  • timesfm-forecasting/SKILL.md — the agent harness. Mandatory preflight, dataset fit check, CSV/DataFrame/array input handling, prediction intervals.
  • timesfm-forecasting/scripts/check_system.py — cross-platform RAM/VRAM/disk detector. Linux /proc/meminfo, Darwin sysctl/vm_stat, Windows GlobalMemoryStatusEx. This is the file that gets you out of “I crashed my laptop loading the model” territory.
  • src/timesfm3/torch/transformer.py — MultiHeadAttention with QK-RoPE, QK-RMSNorm, PerDimScale, KV-cache, and the rescale_logits flag that toggles between Flax memory-efficient-attention behavior and standard SDPA.
  • src/timesfm3/torch/cpm_revin_refine.py — cpm_iterative_revin_refine, the function that makes long-horizon mean drift go away. Worth reading for the block_offset carry state alone.
  • src/timesfm3/__init__.py — PEP 562 lazy re-export of the PyTorch backend. This is why from timesfm3.mlx import TimesFM3Forecaster doesn’t pull torch.
  • BigQuery ML timesfm model — Google Cloud managed serving, if you don’t want to run it yourself.
  • Google Sheets forecast in connected sheets — same model exposed as a Sheets function.
  • Vertex Model Garden TimesFM endpoint — Dockerized endpoint for agentic calling.

The thing I’d want to see next is a published comparison on the same hardware between TimesFM 3.0 (MLX) and Chronos-Bolt (MLX or PyTorch CPU) on fev-bench — the leaderboard numbers from TimesFM’s own release don’t include side-by-side MLX throughput for the competitors. If Google publishes that, the MLX-backend story stops being a footnote and becomes the deciding factor for anyone on Apple silicon.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.