Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 25 additions & 9 deletions Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -17,24 +17,40 @@ ENV UV_LINK_MODE=copy \
RUN apt-get update && apt-get install -y --no-install-recommends libgomp1 \
&& rm -rf /var/lib/apt/lists/*

# `server` and `sktime` are always in, so the process can load naive. Heavier extras
# come from the build arg, one word each, e.g.
# `server` is always in, and it pulls `sktime`, so the process can load naive. Heavier
# extras come from the build arg, one word each, e.g.
# docker build --build-arg TSERVE_EXTRAS=hub .
# docker build --build-arg TSERVE_EXTRAS="chronos gpu" .
ARG TSERVE_EXTRAS=""
# Empty keeps PyPI torch (CUDA on Linux). Any value appends the CPU index, e.g.
# docker build --build-arg TSERVE_CPU=1 .
ARG TSERVE_CPU=""

# Optional `--extra` for every sync. Empty when TSERVE_EXTRAS is unset.
ENV SYNC_EXTRAS="${TSERVE_EXTRAS:+--extra ${TSERVE_EXTRAS}}"

# Resolve and install third-party deps from pyproject.toml only. Source changes
# then do not rebuild this layer. uv.lock is not tracked, so this is not
# `--frozen` / `--locked`. printf repeats `--extra` once per remaining word.
# then do not rebuild this layer.
COPY docker/pytorch-cpu.toml /pytorch-cpu.toml
COPY pyproject.toml ./
RUN --mount=type=cache,target=/root/.cache/uv \
uv sync --no-dev --no-install-project --no-editable \
$(printf -- '--extra %s ' server sktime $TSERVE_EXTRAS)
if [ -n "$TSERVE_CPU" ]; then cat /pytorch-cpu.toml >> pyproject.toml; fi \
&& uv sync \
--no-dev \
--no-install-project \
--no-editable \
--extra server \
$SYNC_EXTRAS

# Install tserve now that the source tree is present. Third-party deps stay in
# the layer above, so a source edit rebuilds only this step.
COPY . .
RUN --mount=type=cache,target=/root/.cache/uv \
uv sync --no-dev --no-editable \
$(printf -- '--extra %s ' server sktime $TSERVE_EXTRAS)
if [ -n "$TSERVE_CPU" ]; then cat /pytorch-cpu.toml >> pyproject.toml; fi \
&& uv sync \
--no-dev \
--no-editable \
--extra server \
$SYNC_EXTRAS

# Runtime image: no uv, no source tree. `--no-editable` baked tserve
# into the venv, so only `.venv` is copied. Python path must match the builder.
Expand Down
6 changes: 3 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -122,9 +122,9 @@ Multivariate series, covariates, and quantiles differ by family: [Capabilities](

Docker needs no local Python. uv and pip need Python 3.12 or newer. Install the extra, or pull the tag, for the family in the table above.

- **Docker.** [Pull an image](https://tserve.readthedocs.io/en/latest/server/docker/#pull-an-image), then [run the server](https://tserve.readthedocs.io/en/latest/server/docker/#run-the-server). CPU and GPU are separate tags.
- **uv or pip.** [UV / Pip](https://tserve.readthedocs.io/en/latest/installation/#uv-pip). A PyPI install takes CUDA torch (MPS on macOS). A CPU wheel: [CPU-only install](https://tserve.readthedocs.io/en/latest/server/pip/#cpu-only-install).
- **A clone.** [From source](https://tserve.readthedocs.io/en/latest/installation/#from-source). On a clone, uv selects the torch index with the [`gpu` extra](https://tserve.readthedocs.io/en/latest/server/source/#gpu).
- **Docker.** [Pull an image](https://tserve.readthedocs.io/en/latest/server/docker/#pull-an-image), then [run the server](https://tserve.readthedocs.io/en/latest/server/docker/#run-the-server). Each family has a CPU tag and a `-gpu` tag.
- **uv or pip.** [UV / Pip](https://tserve.readthedocs.io/en/latest/installation/#uv-pip). A CPU-only build for your OS on [CPU-only install](https://tserve.readthedocs.io/en/latest/server/pip/#cpu-only-install).
- **A clone.** [From source](https://tserve.readthedocs.io/en/latest/server/source/). An editable install for development and unreleased changes.

## Load a model

Expand Down
55 changes: 28 additions & 27 deletions docker-bake.hcl
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
# TServe images: same Dockerfile, different TSERVE_EXTRAS.
# Every tag is linux/amd64 + linux/arm64 (Linux, Mac, Windows Docker Desktop).
# `--extra gpu` is a torch index choice (PyPI / CUDA / MPS), not an arch pin.
# Empty TSERVE_CPU keeps PyPI torch (CUDA on Linux). CPU tags set TSERVE_CPU
# so torch comes from the CPU index. :base has no torch.
#
# First time on a new machine:
# docker run --privileged --rm tonistiigi/binfmt --install all
Expand Down Expand Up @@ -55,156 +56,156 @@ target "base" {

target "hub" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "hub" }
args = { TSERVE_EXTRAS = "hub", TSERVE_CPU = "1" }
tags = ["${TSERVE_IMAGE}:hub"]
}

target "chronos" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "chronos" }
args = { TSERVE_EXTRAS = "chronos", TSERVE_CPU = "1" }
tags = ["${TSERVE_IMAGE}:chronos"]
}

target "kronos" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "kronos" }
args = { TSERVE_EXTRAS = "kronos", TSERVE_CPU = "1" }
tags = ["${TSERVE_IMAGE}:kronos"]
}

target "granite" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "granite" }
args = { TSERVE_EXTRAS = "granite", TSERVE_CPU = "1" }
tags = ["${TSERVE_IMAGE}:granite"]
}

target "moirai" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "moirai" }
args = { TSERVE_EXTRAS = "moirai", TSERVE_CPU = "1" }
tags = ["${TSERVE_IMAGE}:moirai"]
}

target "tirex" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "tirex" }
args = { TSERVE_EXTRAS = "tirex", TSERVE_CPU = "1" }
tags = ["${TSERVE_IMAGE}:tirex"]
}

target "tirex2" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "tirex2" }
args = { TSERVE_EXTRAS = "tirex2", TSERVE_CPU = "1" }
tags = ["${TSERVE_IMAGE}:tirex2"]
}

target "toto" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "toto" }
args = { TSERVE_EXTRAS = "toto", TSERVE_CPU = "1" }
tags = ["${TSERVE_IMAGE}:toto"]
}

target "mantis" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "mantis" }
args = { TSERVE_EXTRAS = "mantis", TSERVE_CPU = "1" }
tags = ["${TSERVE_IMAGE}:mantis"]
}

target "timesfm3" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "timesfm3" }
args = { TSERVE_EXTRAS = "timesfm3", TSERVE_CPU = "1" }
tags = ["${TSERVE_IMAGE}:timesfm3"]
}

target "t0" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "t0" }
args = { TSERVE_EXTRAS = "t0", TSERVE_CPU = "1" }
tags = ["${TSERVE_IMAGE}:t0"]
}

target "tafsut" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "tafsut" }
args = { TSERVE_EXTRAS = "tafsut", TSERVE_CPU = "1" }
tags = ["${TSERVE_IMAGE}:tafsut"]
}

target "full" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "full" }
args = { TSERVE_EXTRAS = "full", TSERVE_CPU = "1" }
tags = ["${TSERVE_IMAGE}:full"]
}

target "hub-gpu" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "hub gpu" }
args = { TSERVE_EXTRAS = "hub" }
tags = ["${TSERVE_IMAGE}:hub-gpu"]
}

target "chronos-gpu" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "chronos gpu" }
args = { TSERVE_EXTRAS = "chronos" }
tags = ["${TSERVE_IMAGE}:chronos-gpu"]
}

target "kronos-gpu" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "kronos gpu" }
args = { TSERVE_EXTRAS = "kronos" }
tags = ["${TSERVE_IMAGE}:kronos-gpu"]
}

target "granite-gpu" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "granite gpu" }
args = { TSERVE_EXTRAS = "granite" }
tags = ["${TSERVE_IMAGE}:granite-gpu"]
}

target "moirai-gpu" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "moirai gpu" }
args = { TSERVE_EXTRAS = "moirai" }
tags = ["${TSERVE_IMAGE}:moirai-gpu"]
}

target "tirex-gpu" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "tirex gpu" }
args = { TSERVE_EXTRAS = "tirex" }
tags = ["${TSERVE_IMAGE}:tirex-gpu"]
}

target "tirex2-gpu" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "tirex2 gpu" }
args = { TSERVE_EXTRAS = "tirex2" }
tags = ["${TSERVE_IMAGE}:tirex2-gpu"]
}

target "toto-gpu" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "toto gpu" }
args = { TSERVE_EXTRAS = "toto" }
tags = ["${TSERVE_IMAGE}:toto-gpu"]
}

target "mantis-gpu" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "mantis gpu" }
args = { TSERVE_EXTRAS = "mantis" }
tags = ["${TSERVE_IMAGE}:mantis-gpu"]
}

target "timesfm3-gpu" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "timesfm3 gpu" }
args = { TSERVE_EXTRAS = "timesfm3" }
tags = ["${TSERVE_IMAGE}:timesfm3-gpu"]
}

target "t0-gpu" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "t0 gpu" }
args = { TSERVE_EXTRAS = "t0" }
tags = ["${TSERVE_IMAGE}:t0-gpu"]
}

target "tafsut-gpu" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "tafsut gpu" }
args = { TSERVE_EXTRAS = "tafsut" }
tags = ["${TSERVE_IMAGE}:tafsut-gpu"]
}

target "full-gpu" {
inherits = ["_common"]
args = { TSERVE_EXTRAS = "full gpu" }
args = { TSERVE_EXTRAS = "full" }
tags = ["${TSERVE_IMAGE}:full-gpu"]
}
10 changes: 10 additions & 0 deletions docker/pytorch-cpu.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
# Appended to pyproject.toml when the image is built with TSERVE_CPU set.
# `explicit` keeps every package except torch on PyPI.

[[tool.uv.index]]
name = "pytorch-cpu"
url = "https://download.pytorch.org/whl/cpu"
explicit = true

[tool.uv.sources]
torch = [{ index = "pytorch-cpu" }]
46 changes: 9 additions & 37 deletions docs/installation.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,7 @@ Other model families use different image tags. Choose the model first, then use

## UV / Pip

TServe requires Python 3.12 or newer. Install the `server` extra and the extra for the model family you need. The examples below install the `hub` family. A plain install pulls the CUDA build of torch (MPS on macOS); the CPU tabs skip that download on a machine without a GPU.
TServe requires Python 3.12 or newer. Install the `server` extra and the extra for the model family you need. The examples below install the `hub` family. A CPU build of torch: [CPU-only install](server/pip.md#cpu-only-install).

=== "uv"

Expand All @@ -36,23 +36,9 @@ TServe requires Python 3.12 or newer. Install the `server` extra and the extra f
uv venv
```

Then install TServe:

=== "GPU (default)"

```bash
uv pip install "tserve[server,hub]"
```

=== "CPU only"

```bash
uv pip install torch --index-url https://download.pytorch.org/whl/cpu
```

```bash
uv pip install "tserve[server,hub]"
```
```bash
uv pip install "tserve[server,hub]"
```

=== "pip"

Expand All @@ -72,29 +58,15 @@ TServe requires Python 3.12 or newer. Install the `server` extra and the extra f
.venv\Scripts\Activate.ps1
```

Then install TServe:

=== "GPU (default)"

```bash
python -m pip install "tserve[server,hub]"
```

=== "CPU only"

```bash
python -m pip install torch --index-url https://download.pytorch.org/whl/cpu
```

```bash
python -m pip install "tserve[server,hub]"
```
```bash
python -m pip install "tserve[server,hub]"
```

The `server` extra alone supports the `naive` test baseline. Replace `hub` with another [family extra](models/index.md#dependencies), or use `full` for every family. The same CPU-first order is documented as [CPU-only install](server/pip.md#cpu-only-install).
The `server` extra alone supports the `naive` test baseline. Replace `hub` with another [family extra](models/index.md#dependencies), or use `full` for every family.

## From source

Use a source install when developing TServe or testing unreleased changes. It is also the only path where the `gpu` extra selects the torch index, because that choice lives in the repository's uv lockfile. The [From source](server/source.md) guide covers cloning the repository, editable installs, dependency extras, and GPU setup.
Use a source install when developing TServe or testing unreleased changes. The [From source](server/source.md) guide covers cloning the repository, editable installs, and dependency extras.

## Next

Expand Down
2 changes: 1 addition & 1 deletion docs/models/hub.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ Four Hugging Face families, 81 of the catalog's 117 models. The usual starting p
| --- | --- | --- | --- | --- | --- |
| `hub` | [`:hub`](https://hub.docker.com/r/sktime/tserve/tags?name=hub) | [`:hub-gpu`](https://hub.docker.com/r/sktime/tserve/tags?name=hub-gpu) | Chronos Bolt, Chronos T5, TTM, TimesFM 2.x | 81 | `chronos_bolt` |

Builds on [`base`](base.md), so `naive` is available here too. Every extra that pulls `hf` builds on `hub`, so those pages can load these models as well.
Builds on [`base`](base.md), so `naive` is available here too. Every extra that includes `hub` can load these models as well.

## Start a server

Expand Down
2 changes: 1 addition & 1 deletion docs/reference/development.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,7 @@ make docs-serve

## Docker images

`Dockerfile` always installs `--extra server --extra sktime` and adds whatever `TSERVE_EXTRAS` names, which is how one file produces every tag. `docker-bake.hcl` holds the published matrix: one target per tag, plus `cpu` and `gpu` groups, with `TSERVE_IMAGE` defaulting to `sktime/tserve`.
`Dockerfile` always installs `--extra server` and adds whatever `TSERVE_EXTRAS` names. `TSERVE_CPU` selects the CPU torch index; leaving it empty keeps the PyPI wheel. That is how one file produces every tag. `docker-bake.hcl` holds the published matrix: one target per tag, plus `cpu` and `gpu` groups, with `TSERVE_IMAGE` defaulting to `sktime/tserve`.

```bash
TSERVE_IMAGE=sktime/tserve docker buildx bake --push hub
Expand Down
14 changes: 10 additions & 4 deletions docs/server/docker.md
Original file line number Diff line number Diff line change
Expand Up @@ -90,7 +90,7 @@ The `*-gpu` tags install torch from PyPI instead of the CPU wheel index. They ne
docker run --rm --gpus all -p 8000:8000 sktime/tserve:hub-gpu chronos_bolt ttm_r3
```

That covers Linux and Windows through WSL2. Docker on macOS has no GPU passthrough, so Apple silicon acceleration means a [local UV / Pip install](pip.md#install), which takes the MPS build of torch.
That covers Linux and Windows through WSL2. Docker on macOS has no GPU passthrough, so Apple silicon acceleration means a [local UV / Pip install](pip.md#install).

## Models from a directory

Expand All @@ -113,16 +113,22 @@ cd tserve

### One image with `docker build`

The [Dockerfile](https://github.com/sktime/tserve/blob/main/Dockerfile) always installs `--extra server --extra sktime`; `TSERVE_EXTRAS` adds the heavier ones:
The [Dockerfile](https://github.com/sktime/tserve/blob/main/Dockerfile) always installs `--extra server`. `TSERVE_EXTRAS` adds the heavier ones, one word each. Leave `TSERVE_CPU` empty for the PyPI torch wheel. Any value installs torch from the CPU index, which is what the published CPU tags do:

```bash
docker build --build-arg TSERVE_EXTRAS=hub -t tserve:hub .
docker build --build-arg TSERVE_EXTRAS=hub --build-arg TSERVE_CPU=1 -t tserve:hub .
```

A GPU image is the same build with `TSERVE_CPU` unset:

```bash
docker build --build-arg TSERVE_EXTRAS=chronos -t tserve:chronos-gpu .
```

Several extras go in one quoted argument:

```bash
docker build --build-arg TSERVE_EXTRAS="chronos gpu" -t tserve:chronos-gpu .
docker build --build-arg TSERVE_EXTRAS="chronos moirai" --build-arg TSERVE_CPU=1 -t tserve:custom .
```

This builds for the architecture of the machine you are on and leaves the image in the local store, which is all you need to run it locally:
Expand Down
2 changes: 1 addition & 1 deletion docs/server/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ A bare `tserve` loads `naive` only. Name models to load them too. `GET /models`
tserve chronos_bolt ttm_r3
```

This installs CUDA torch (MPS on macOS). A CPU wheel: [CPU-only install](pip.md#cpu-only-install).
This installs CUDA torch (MPS on macOS). A CPU build: [CPU-only install](pip.md#cpu-only-install).

Another family is the same command with that row's tag and model. `moirai_2`:

Expand Down
Loading
Loading