Quick start¶
Start a server with two models, confirm that it is ready, and send a forecast. If TServe is not installed yet, begin with Installation.
1. Start the server¶
Docker is faster and preferred. The image already carries the dependencies, and CPU vs GPU is a tag. The UV and Pip tabs assume you already installed tserve[server,hub].
The first start downloads model weights from Hugging Face. The terminal then prints the local URLs for the dashboard, Swagger UI, and ReDoc.
2. Check what loaded¶
GET /models reports models loaded by this process, not every model in the catalog.
The response includes naive, which is always available as a test baseline, plus chronos_bolt and ttm_r3.
3. Send a prediction¶
This request sends five days of sales and asks chronos_bolt for the next three:
The response contains three predicted days and identifies the model that served them:
{
"predictions": {
"timestamp": ["2024-01-06T00:00:00", "2024-01-07T00:00:00", "2024-01-08T00:00:00"],
"sales": [139.96, 138.93, 138.26]
},
"quantiles": null,
"model": "chronos_bolt",
"request_id": "…"
}
4. Try the Python client¶
Install the client in a separate environment if the calling application does not share the server environment:
Send the same request:
from tserve.client import Client
past = {
"timestamp": ["2024-01-01", "2024-01-02", "2024-01-03", "2024-01-04", "2024-01-05"],
"sales": [120, 135, 128, 142, 138],
}
with Client("http://127.0.0.1:8000") as client:
result = client.predict(
past=past,
time="timestamp",
target=["sales"],
fh=3,
model="chronos_bolt",
)
print(result.predictions)
The input may also be a pandas, polars, or pyarrow table. See the Python client for type-preserving responses and the data specification for every request field and table format.
5. Open the dashboard¶
Open http://127.0.0.1:8000/. Pick a loaded model, a horizon, and a sample series or your own CSV. The page posts POST /predict and plots the result. What you can do
timesfm_3 forecasts retail sales with a 90% prediction interval.
