/00 — boot sequence

Hello.

Article

TimesFM: Google's Decoder-Only Foundation Model for Time-Series Forecasting

July 7, 2026•3 min read
time-series forecasting foundation model transformers zero-shot learning Google AI

Time-series forecasting is a critical task across industries—retail, finance, healthcare, and natural sciences. While deep learning models have improved accuracy, they often require lengthy training cycles for each new dataset. Enter TimesFM (Time Series Foundation Model), a decoder-only transformer pre-trained on 100 billion real-world time points, designed to deliver strong zero-shot forecasts out of the box.

Why a Foundation Model for Time-Series?

Traditional deep learning forecasters, like DeepAR or PatchTST, need task-specific training and validation. This limits rapid prototyping and deployment. A foundation model, pre-trained on diverse time-series data, can generalize to unseen datasets without additional training—similar to how large language models (LLMs) like GPT handle zero-shot text tasks.

TimesFM is compact: only 200 million parameters. Yet it competes with state-of-the-art supervised models on public benchmarks, demonstrating that a well-designed decoder-only architecture can capture temporal patterns effectively.

Architecture: Patches as Tokens

TimesFM adapts the decoder-only transformer paradigm to time-series. Key ideas:

  • Patching: Instead of treating each individual time point as a token, contiguous groups (patches) are used. This reduces sequence length and captures local patterns.
  • Input/Output asymmetry: The output patch length is larger than the input patch length (e.g., input 32, output 128). This reduces the number of autoregressive generation steps during inference, minimizing error accumulation.
  • Residual MLP projector: A multilayer perceptron with residual connections converts each patch into a token embedding, then positional encodings are added before the transformer layers.

The model is trained to predict the next output patch given a history of patches, simultaneously across all offsets within a training sequence.

Pretraining Data: Synthetic + Real-World

TimesFM was trained on a corpus of 100 billion time points, combining:

  • Synthetic data from statistical models (ARIMA, etc.) to teach basic temporal grammar.
  • Real-world data including Google Trends and Wikipedia pageviews. These reflect human interest and activity, providing rich, diverse patterns that help generalization.

This mix ensures the model learns both fundamental dynamics and realistic variability.

Zero-Shot Performance

TimesFM was evaluated without any fine-tuning on unseen datasets from the Monash Forecasting Archive (traffic, weather, demand) and long-horizon benchmarks like ETT. Results show:

  • On Monash: zero-shot TimesFM outperforms most supervised methods (including DeepAR, PatchTST) and beats GPT-3.5-based forecasting (llmtime) by a large margin, despite being orders of magnitude smaller.
  • On ETT long horizon: TimesFM matches or exceeds supervised PatchTST and surpasses llmtime. For 96- and 192-steps ahead, it achieves lower MAE.

Key Takeaways

  • Decoder-only works for time-series: Treating patches as tokens and using causal attention enables autoregressive generation with minimal changes.
  • Long output patches reduce error: By predicting larger chunks per step, the model avoids compounding errors from many generation steps.
  • Zero-shot forecasting is practical: Pre-training on diverse data (including search trends) yields strong out-of-the-box performance across domains.
  • Size matters, but not too much: 200M parameters are sufficient for competitive results, making deployment feasible.

Conclusion

TimesFM demonstrates that a decoder-only foundation model, pre-trained on a large and diverse time-series corpus, can achieve remarkable zero-shot forecasts. It bridges the gap between the convenience of LLMs and the specificity of time-series tasks. Google plans to release TimesFM via Vertex AI later this year, opening the door for easier, faster forecasting in production.


Source: A decoder-only foundation model for time-series forecasting

Automated Transmission

This entry was synthesized and populated dynamically using native API integrations.

Resources & Links