UV / Pip¶
Python >= 3.12. Install TServe from PyPI with uv or pip, then start tserve. Editable installs from a clone stay on From source.
Install¶
server is enough to serve naive (a test baseline). The hub extra above covers Chronos Bolt/T5, TTM, and TimesFM 2.x. Do not add client on a machine that only serves.
Dependencies¶
The extra name is the CPU image tag. GPU tags are {extra}-gpu. There is no :base-gpu.
kronos is built on base, so it cannot load Chronos Bolt, TTM, or TimesFM. hub can, and so can every extra that includes it: chronos, granite, moirai, tirex, tirex2, toto, mantis, timesfm3, t0, tafsut, and full.
full is chronos, kronos, granite, moirai, tirex, tirex2, toto, mantis, timesfm3, t0, and tafsut. client, http, dev, docs, and all-extras are not model families. Moirai pins gluonts, lightning, and hydra-core when python_version < '3.14'.
added counts checkpoints that extra contributes. full is the total, including naive.
| extra | CPU tag | GPU tag | families | added | example |
|---|---|---|---|---|---|
server |
:base |
— | Naive | 1 | naive |
hub |
:hub |
:hub-gpu |
Chronos Bolt, Chronos T5, TTM, TimesFM 2.x | 81 | chronos_bolt |
chronos |
:chronos |
:chronos-gpu |
Chronos-2 | 3 | chronos_2 |
kronos |
:kronos |
:kronos-gpu |
Kronos, WindFM | 5 | kronos |
granite |
:granite |
:granite-gpu |
FlowState | 2 | flowstate |
moirai |
:moirai |
:moirai-gpu |
Moirai 2, Moirai 1.x, Lag-Llama | 8 | moirai_2 |
tirex |
:tirex |
:tirex-gpu |
TiRex | 2 | tirex |
tirex2 |
:tirex2 |
:tirex2-gpu |
TiRex-2 | 4 | tirex_2 |
toto |
:toto |
:toto-gpu |
Toto-2 | 5 | toto_2_0_4m |
mantis |
:mantis |
:mantis-gpu |
Mantis | 3 | mantis_8m |
timesfm3 |
:timesfm3 |
:timesfm3-gpu |
TimesFM 3 | 1 | timesfm_3 |
t0 |
:t0 |
:t0-gpu |
T0 | 1 | t0 |
tafsut |
:tafsut |
:tafsut-gpu |
Tafsut | 1 | tafsut |
full |
:full |
:full-gpu |
all of the above | 117 | chronos_2 |
Replace hub in the install above with another extra from the table. full is the union extra. all-extras is a pip convenience for client,server,full and is not a Docker tag. Each extra's command and models: catalog.
A CPU build of torch: CPU-only install.
CPU-only install¶
Family extras pull torch. Pick the CPU build for your OS on the PyTorch install page, install it, then install TServe as above. If a later install replaces that build, run the PyTorch command again.
Swap hub for any other family extra from Dependencies. A GPU host can also use a *-gpu image: GPU images.
Serve from the command line¶
Startup prints the URLs it binds:
Starting TServe
Dashboard http://127.0.0.1:8000/
Swagger UI http://127.0.0.1:8000/docs
ReDoc http://127.0.0.1:8000/redoc
--host 127.0.0.1 accepts local connections only. 0.0.0.0 also accepts them from your network. --log-level is debug, info, warning, error, or critical. Ctrl+C stops the process and exits 0. Every flag: CLI. A craft spec is id=spec: Craft specs.
Serve from Python¶
Server takes the same arguments:
from tserve.server import Server
server = Server(
model=["chronos_bolt", "ttm_r3"],
host="127.0.0.1",
port=8000,
)
print(server.url) # http://127.0.0.1:8000
server.run()
Models load during construction, so an unknown model or a missing dependency raises before the port is bound. run() blocks until the process stops. Omitting model still loads naive.
server.app is the FastAPI app:
Use one worker per process. Each worker loads its own copy of every model.
Next¶
- From source — editable install from a clone
- Live objects — Python can also serve estimators you configured in the session
- Craft specs — load a sktime craft spec as
(id, spec)or CLIid=spec - Models from a directory — serve saved sktime
.zipfiles - Dashboard — the console at http://127.0.0.1:8000/