Skip to content

UV / Pip

Python >= 3.12. Install TServe from PyPI with uv or pip, then start tserve. Editable installs from a clone stay on From source.

Install

uv pip install "tserve[server,hub]"
pip install "tserve[server,hub]"

server is enough to serve naive (a test baseline). The hub extra above covers Chronos Bolt/T5, TTM, and TimesFM 2.x. Do not add client on a machine that only serves.

Dependencies

The extra name is the CPU image tag. GPU tags are {extra}-gpu. There is no :base-gpu.

kronos is built on base, so it cannot load Chronos Bolt, TTM, or TimesFM. hub can, and so can every extra that includes it: chronos, granite, moirai, tirex, tirex2, toto, mantis, timesfm3, t0, tafsut, and full.

full is chronos, kronos, granite, moirai, tirex, tirex2, toto, mantis, timesfm3, t0, and tafsut. client, http, dev, docs, and all-extras are not model families. Moirai pins gluonts, lightning, and hydra-core when python_version < '3.14'.

added counts checkpoints that extra contributes. full is the total, including naive.

extra CPU tag GPU tag families added example
server :base — Naive 1 naive
hub :hub :hub-gpu Chronos Bolt, Chronos T5, TTM, TimesFM 2.x 81 chronos_bolt
chronos :chronos :chronos-gpu Chronos-2 3 chronos_2
kronos :kronos :kronos-gpu Kronos, WindFM 5 kronos
granite :granite :granite-gpu FlowState 2 flowstate
moirai :moirai :moirai-gpu Moirai 2, Moirai 1.x, Lag-Llama 8 moirai_2
tirex :tirex :tirex-gpu TiRex 2 tirex
tirex2 :tirex2 :tirex2-gpu TiRex-2 4 tirex_2
toto :toto :toto-gpu Toto-2 5 toto_2_0_4m
mantis :mantis :mantis-gpu Mantis 3 mantis_8m
timesfm3 :timesfm3 :timesfm3-gpu TimesFM 3 1 timesfm_3
t0 :t0 :t0-gpu T0 1 t0
tafsut :tafsut :tafsut-gpu Tafsut 1 tafsut
full :full :full-gpu all of the above 117 chronos_2

Replace hub in the install above with another extra from the table. full is the union extra. all-extras is a pip convenience for client,server,full and is not a Docker tag. Each extra's command and models: catalog.

A CPU build of torch: CPU-only install.

CPU-only install

Family extras pull torch. Pick the CPU build for your OS on the PyTorch install page, install it, then install TServe as above. If a later install replaces that build, run the PyTorch command again.

Swap hub for any other family extra from Dependencies. A GPU host can also use a *-gpu image: GPU images.

Serve from the command line

uv run tserve chronos_bolt ttm_r3
tserve chronos_bolt ttm_r3

Startup prints the URLs it binds:

Starting TServe
  Dashboard   http://127.0.0.1:8000/
  Swagger UI  http://127.0.0.1:8000/docs
  ReDoc       http://127.0.0.1:8000/redoc
tserve --host 0.0.0.0 --port 9000 --log-level warning chronos_bolt ttm_r3

--host 127.0.0.1 accepts local connections only. 0.0.0.0 also accepts them from your network. --log-level is debug, info, warning, error, or critical. Ctrl+C stops the process and exits 0. Every flag: CLI. A craft spec is id=spec: Craft specs.

Serve from Python

Server takes the same arguments:

from tserve.server import Server

server = Server(
    model=["chronos_bolt", "ttm_r3"],
    host="127.0.0.1",
    port=8000,
)
print(server.url)  # http://127.0.0.1:8000
server.run()

Models load during construction, so an unknown model or a missing dependency raises before the port is bound. run() blocks until the process stops. Omitting model still loads naive.

server.app is the FastAPI app:

import uvicorn

uvicorn.run(server.app, host="127.0.0.1", port=8000, workers=1)

Use one worker per process. Each worker loads its own copy of every model.

Next