HTTP¶
Send JSON to POST /predict from any language. The fields match Client.predict(...). The Python client posts Arrow to /predict/bytes instead of this route.
Endpoints¶
| method | path | what it gives you |
|---|---|---|
POST |
/predict |
JSON prediction |
POST |
/predict/bytes |
Arrow prediction, used by the Python client |
GET |
/health |
process liveness |
GET |
/models |
loaded models |
GET |
/stats |
uptime, memory, per-model metrics |
GET |
/ |
browser dashboard |
GET |
/docs, /redoc, /openapi.json |
live OpenAPI |
Every route returns JSON except /predict/bytes, which speaks Arrow, and /, which serves the dashboard. The HTTP API reference lists the same routes with schema links.
Start a server¶
Point forecasts use chronos_bolt. Quantiles use timesfm_2_5, because Chronos Bolt cannot return them:
Other ways to start: Server.
Send a prediction¶
past is a table, one row per timestamp:
POST /predict returns a column-oriented JSON table:
{
"predictions": {
"timestamp": [
"2024-01-06T00:00:00",
"2024-01-07T00:00:00",
"2024-01-08T00:00:00"
],
"sales": [139.96, 138.93, 138.26]
},
"quantiles": null,
"model": "chronos_bolt",
"request_id": "…"
}
See Data specification for every request and response field.
Use row-oriented JSON¶
Tables can also use columns and data. Here time and target are omitted, so TServe uses the first column as time and the other column as the target:
The response is still column-oriented JSON. HTTP does not preserve the row-oriented request shape.
Request quantiles¶
Quantiles are a second result table. The estimator must support quantile prediction, so this example uses the loaded timesfm_2_5 model:
curl -s http://127.0.0.1:8000/predict -H "Content-Type: application/json" -d '{
"past": {
"timestamp": ["2024-01-01", "2024-02-01", "2024-03-01", "2024-04-01", "2024-05-01"],
"sales": [120, 135, 128, 142, 150]
},
"time": "timestamp",
"target": ["sales"],
"fh": 3,
"model": "timesfm_2_5",
"quantiles": [0.1, 0.5, 0.9]
}'
Many estimators name those columns {target}_{level} (sales_0.1, sales_0.5, sales_0.9). timesfm_2_5 currently uses a positional prefix (0_0.1, 0_0.5, 0_0.9). See Quantiles for the response shape and model limitation.
Inspect the server¶
Use the JSON status routes to check the process and its loaded models:
GET /health checks process liveness, not whether models are warm. GET /models lists loaded models, not the registry catalog.
Point a browser at / for the dashboard or /docs to try the endpoints from Swagger.
Arrow endpoint¶
POST /predict/bytes accepts multipart metadata and Arrow IPC tables and returns a TServe envelope with media type application/vnd.tserve.predict+arrow. This is the route used by the Python client; you normally do not construct its body yourself.
Errors¶
Prediction is POST-only. GET /predict returns 405 Method Not Allowed. Invalid JSON request shapes return 422. Coercion and prediction failures, including an unloaded model, return 400 with an error message and request_id.
See Errors for response bodies and Python exceptions.