mrkeyoor.com_
Wed 30 Sept 20:35 UTC
Dataevaluationupdated 26 Aug 2026

timesfm review

TimesFM is Google's pretrained model for predicting future values of a time series without first training a model on that particular series. The open Python package returns point and quantile forecasts, with optional covariates for known influences on the target.

+341stars / 7d
Verdict

Our TimesFM run passed 26 tests with 0 assertion failures, but 6 modules could not collect because pandas was missing and the audit found 5 known vulnerabilities. Put it in a forecasting bake-off when a zero-shot point and quantile baseline could save training work. Do not let it drive purchasing, staffing, or risk decisions until rolling backtests beat simple seasonal methods on the horizons that matter.

We ran it

Lab card: what happened when we ran timesfmScreenshot of timesfm (research.google/blog/a-decoder-only-foundation-model-for-time-series-forecasting)
Install✓ · 36s49 packages · 110 MB
Build✓ · 11s
Tests✗ · 11s26 passed · 0 failed · 6 errors of 32 (pytest)
Known vulns5(pip-audit)
Repo103 files~15,377 lines of source · 4 MB · 2 CI workflows · tests dir

Answers from our run

Does timesfm build from source?

Dependencies installed in 36 seconds (49 packages), and the build succeeded in 11 seconds. We cloned commit 3dae50b into a clean Debian container with 3 CPUs and no project-specific setup.

Do timesfm's tests pass?

Yes: 26 of 32 passed when we ran the project's own test command (pytest), with 6 collection errors. Some failures need services or credentials a bare container does not have.

Does timesfm have known vulnerabilities in its dependencies?

pip-audit flagged 5 known advisories in the dependency tree at the time of our run.

Who should not use timesfm?

Teams that need an officially supported Google product: the README says the open release is not one and points support-sensitive users toward Google products such as BigQuery ML.

What are the alternatives to timesfm?

Chronos, Moirai, StatsForecast. Our TimesFM run passed 26 tests with 0 assertion failures, but 6 modules could not collect because pandas was missing and the audit found 5 known vulnerabilities.

Setup3/5Quick core install, but backends and pandas need attention
Docs4/5Clear current example, with model and package version mismatch
Community4/528,269 stars and recent fix activity after July's push
Maturity3/5Research-backed model; 6 collection errors and 5 advisories

Discussed on

  1. hnGoogle's 200M-parameter time-series foundation model with 16k context327 points
  2. hnTimesFM: Time Series Foundation Model for time-series forecasting317 points
  3. hnTimesFM (Time Series Foundation Model)7 points

Who it’s for

Forecasting teams that want a zero-shot baseline across many series.
Data scientists comparing pretrained models with seasonal and statistical methods.
Python users who need point forecasts and uncertainty ranges from one package.
Researchers prepared to backtest every model and configuration on their own data.

Who it’s NOT for

Teams that need an officially supported Google product: the README says the open release is not one and points support-sensitive users toward Google products such as BigQuery ML.
Analysts who need causal coefficients or a model whose parameters explain each driver: TimesFM forecasts future values rather than estimating causal effects.
Pipelines that reuse a mutable input list without checking it: pull request 468 reproduces forecast() appending padding to the caller's list and returning an extra bogus forecast on a later call.
Users expecting one dependency extra to cover every backend: the README separates Torch, Flax, and XReg installs, plus hardware-specific PyTorch or JAX setup.
Security-sensitive environments unwilling to investigate dependency advisories: our installed environment contained 5 known vulnerabilities.

Setup reality

Our sandbox installed 49 Python packages in 36 seconds and used 110 MB on disk. The build succeeded in 11 seconds. Tests ended after 11 seconds with 26 passed, 0 failed, and 6 collection errors out of 32; pip-audit reported 5 known vulnerabilities.

The package uses separate extras for Torch, Flax, and XReg. First model use needs a TimesFM checkpoint from Hugging Face, while accelerator use needs a matching PyTorch or JAX backend. The open package needs no hosted inference key.

The checkout had 103 files, about 15,377 source lines, 2 CI workflows, a tests directory, and no Dockerfile. All 6 collection errors named a missing pandas import, including v1 data-loader and TimesFM tests. Model 2.5 and package v2.0.2 also use different version numbers.

TimesFM is a zero-shot baseline, not a final forecast

TimesFM takes one-dimensional histories and predicts future values without fitting a separate model to each series. Version 2.5 uses 200 million parameters, accepts up to 16,000 context points, and can forecast up to 1,000 steps with an optional 30 million-parameter quantile head. That makes it attractive for a catalog of products, sensors, or operational metrics where training and maintaining one model per series would be expensive.

The result should enter a comparison, not skip one. A pretrained model may transfer useful patterns, but it does not know a company's stockouts, sensor replacements, accounting changes, or promotion calendar by default. Run rolling historical windows and compare the exact horizons used for decisions against seasonal naive, moving averages, statistical models, and the incumbent system. Break errors down by series type because one average can hide expensive failures.

The API covers point, quantile, and covariate paths

The README example loads a Torch checkpoint, compiles maximum context and horizon values, then sends a list of NumPy arrays to forecast(). It returns point forecasts and quantile output. Configuration flags cover normalization, sign-flip invariance, positive-value inference, and quantile-crossing correction. Those switches express assumptions: forcing positive values may suit demand but would be wrong for a series such as temperature deviations or profit changes.

TimesFM 2.5 removed the old frequency indicator and brought back covariate support through XReg. Known future inputs can improve a forecast when prices, holidays, weather, or regional attributes genuinely exist at prediction time. They also create a direct route for leakage. Every backtest must reconstruct what was known on that date, not attach the final cleaned dataset to an old prediction point. Quantile ordering does not prove the resulting intervals are calibrated.

What happened when we ran it

Our sandbox installed 49 packages in 36 seconds, consuming 110 MB on disk, and completed the build in 11 seconds. The checkout was small for a model package: 103 files, about 15,377 source lines, and 4 MB. We tested commit 3dae50b in a fresh unprivileged Debian container with 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets.

Pytest stopped after 11 seconds with 26 passed, 0 failed, and 6 collection or setup errors out of 32. Every listed error came from a module that could not import pandas, including the v1 data-loader and TimesFM tests. The log therefore shows an incomplete test environment rather than six failed assertions. Pip-audit found 5 known vulnerabilities, so an adopter should identify the packages and fixed versions before approving the environment.

Torch, Flax, and XReg are separate installations

The README provides timesfm[torch], timesfm[flax], and timesfm[xreg] extras. Users must also select PyTorch or JAX builds that match CPU, GPU, TPU, or Apple Silicon hardware. That separation is sensible because accelerator packages are not interchangeable, but it makes the one-line install only the start of a reproducible service image. Pin the backend, model checkpoint, package, NumPy stack, and optional data dependencies together.

Version names add another trap. TimesFM 2.5 is the latest model, while the July 2, 2026 package and GitHub release are numbered 2.0.2. Older 1.0 and 2.0 model code sits in the v1 directory, and the README directs those users to package version 1.3.0. Record both the checkpoint identifier and Python package version in every experiment instead of writing “TimesFM 2” in a report.

Reused input lists can change the next forecast

Open pull request 468 reports that the 2.5 forecast() path mutates the list supplied by its caller when a batch needs padding. Its example starts with 3 inputs for a global batch size of 4. The first call appends a short zero array to the original list; the next call then treats that padding as a real fourth series and returns an extra forecast. The proposed fix copies the list before adding padding.

That bug is easy to defend against while it remains unmerged: pass a fresh list, assert its length and contents after every call, and verify output cardinality. Repeated-call tests matter more than a single successful notebook cell. The same principle applies to XReg batches and quantile outputs. Test one series alone and in different batch groupings, and reject any unexplained change caused only by its neighbors.

The open release has active code, not product support

GitHub showed 28,269 stars, 224 combined issues and pull requests, and a last push on July 14, 2026. Release v2.0.2 arrived July 2 with package-loading and compilation updates. August pull requests were still updating 2.5 examples and fixing the input mutation, so work continued after the last push visible on the default branch. The repository has 2 CI workflows, a tests directory, no Dockerfile, and Apache 2.0 licensing.

Google states that this open version is not an officially supported product. BigQuery ML, Google Sheets, and Vertex Model Garden are listed as product routes for users who want a managed context. Self-hosters own backend packaging, model loading, scaling, dependency advisories, monitoring, and forecast validation. TimesFM is worth trying because the setup is modest and the model offers a common zero-shot reference. Adoption should follow measured improvement on your series, never the model name alone.

Alternatives

ProjectWhat it isPick it when
ChronosAmazon's pretrained family for zero-shot probabilistic time-series forecasting.pick this instead when you want a second foundation-model family for a direct zero-shot comparison.
MoiraiSalesforce's training and inference code for universal time-series models.pick this instead when multivariate forecasting or training within a broader research stack matters.
StatsForecastA collection of classical statistical forecasting models for large groups of series.pick this instead when interpretable statistical baselines and low compute matter more than pretrained transfer.

What people are saying

  1. [github-trending] google-research/timesfm

Sources

  1. TimesFM repository and README
  2. TimesFM v2.0.2 release
  3. Pull request 468: input list mutation
  4. Issue 382: 2.5 example and covariate documentation

More data reviews

TradeGenuis-box · awesome-submitlist · ccf-deadlines · instagram-private-graph · OpenBB · polyledger · the whole board →