Skip to content

Server

The server is the process you start. It loads models, keeps them warm, and answers predict requests. The dashboard, OpenAPI, and /predict are that process.

A bare tserve loads naive only. Name models to load them too. GET /models lists what loaded, not the catalog. Which extra or image each model needs: Dependencies.

Start

docker run --rm -p 8000:8000 sktime/tserve:hub chronos_bolt ttm_r3

Token, cache, GPU, other tags: Docker.

uv pip install "tserve[server,hub]"
uv run tserve chronos_bolt ttm_r3
pip install "tserve[server,hub]"
tserve chronos_bolt ttm_r3

This installs CUDA torch (MPS on macOS). A CPU build: CPU-only install.

Another family is the same command with that row's tag and model. moirai_2:

docker run --rm -p 8000:8000 sktime/tserve:moirai moirai_2

Startup prints:

curl -s http://127.0.0.1:8000/models

GET /health is liveness. Then predict.

CLI flags and Server: UV / Pip. A clone: From source. Each family's command: catalog.

In this section

  • Docker


    Pull a tag, mount the Hub cache, pass a token, or build the image yourself.

    Docker

  • UV / Pip


    Install from PyPI, then tserve or Server.

    UV / Pip

  • From source


    Install from a clone, then tserve or Server.

    From source

  • Live objects


    Serve an estimator you configured in Python. CLI cannot do this.

    Live objects

  • Craft specs


    Load a sktime craft spec as (id, spec) or CLI id=spec.

    Craft specs

  • Models from a directory


    Load sktime .zip files by stem. Mix them with registry models.

    Models from a directory

  • Which models to load


    One page per extra and image tag, with the models each one can serve.

    Catalog

  • Dashboard


    Browser console at GET /. It talks to the JSON endpoints of this process.

    Dashboard