Skip to content

Quick start

Start a server with two models, confirm that it is ready, and send a forecast. If TServe is not installed yet, begin with Installation.

1. Start the server

Docker is faster and preferred. The image already carries the dependencies, and CPU vs GPU is a tag. The UV and Pip tabs assume you already installed tserve[server,hub].

docker run --rm -p 8000:8000 sktime/tserve:hub chronos_bolt ttm_r3
docker run --rm --gpus all -p 8000:8000 sktime/tserve:hub-gpu chronos_bolt ttm_r3
uv run tserve chronos_bolt ttm_r3
tserve chronos_bolt ttm_r3

The first start downloads model weights from Hugging Face. The terminal then prints the local URLs for the dashboard, Swagger UI, and ReDoc.

2. Check what loaded

GET /models reports models loaded by this process, not every model in the catalog.

curl -s http://127.0.0.1:8000/models
curl.exe -s http://127.0.0.1:8000/models

The response includes naive, which is always available as a test baseline, plus chronos_bolt and ttm_r3.

3. Send a prediction

This request sends five days of sales and asks chronos_bolt for the next three:

curl -s http://127.0.0.1:8000/predict -H "Content-Type: application/json" -d '{
  "past": {
    "timestamp": ["2024-01-01", "2024-01-02", "2024-01-03", "2024-01-04", "2024-01-05"],
    "sales": [120, 135, 128, 142, 138]
  },
  "time": "timestamp",
  "target": ["sales"],
  "fh": 3,
  "model": "chronos_bolt"
}'
curl.exe -s http://127.0.0.1:8000/predict -H "Content-Type: application/json" -d '{"past":{"timestamp":["2024-01-01","2024-01-02","2024-01-03","2024-01-04","2024-01-05"],"sales":[120,135,128,142,138]},"time":"timestamp","target":["sales"],"fh":3,"model":"chronos_bolt"}'

The response contains three predicted days and identifies the model that served them:

{
  "predictions": {
    "timestamp": ["2024-01-06T00:00:00", "2024-01-07T00:00:00", "2024-01-08T00:00:00"],
    "sales": [139.96, 138.93, 138.26]
  },
  "quantiles": null,
  "model": "chronos_bolt",
  "request_id": "…"
}

4. Try the Python client

Install the client in a separate environment if the calling application does not share the server environment:

uv pip install "tserve[client]"
pip install "tserve[client]"

Send the same request:

from tserve.client import Client

past = {
    "timestamp": ["2024-01-01", "2024-01-02", "2024-01-03", "2024-01-04", "2024-01-05"],
    "sales": [120, 135, 128, 142, 138],
}

with Client("http://127.0.0.1:8000") as client:
    result = client.predict(
        past=past,
        time="timestamp",
        target=["sales"],
        fh=3,
        model="chronos_bolt",
    )

print(result.predictions)

The input may also be a pandas, polars, or pyarrow table. See the Python client for type-preserving responses and the data specification for every request field and table format.

5. Open the dashboard

Open http://127.0.0.1:8000/. Pick a loaded model, a horizon, and a sample series or your own CSV. The page posts POST /predict and plots the result. What you can do

timesfm_3 forecasts retail sales with a 90% prediction interval.

TServe dashboard: timesfm_3 with a 90% prediction interval

Next