Files
gpu-turnstile/SPEC.md
T
mram 3aaa5d80a9
ci / test (push) Successful in 49s
ci / docker (push) Successful in 1m7s
CI: lowercase image repository name; compose example too
Docker registry names must be lowercase, so PUBLIC/gpu-turnstile was
rejected by buildx. The workflow now lowercases gitea.repository.
2026-09-20 19:25:37 +02:00

228 lines
10 KiB
Markdown

# gpu-turnstile — GPU arbitration proxy for Ollama + ComfyUI
## Problem
One consumer GPU (RTX 5080, 16 GB) is shared by an LLM server (Ollama) and an
image generator (ComfyUI). Both assume they own the card. When both hold models
at once, the NVIDIA Windows driver falls back to system memory and everything
becomes very slow; on Linux it would OOM instead.
## Goal
A single Go binary that sits in front of **both** services and guarantees that
at any moment the GPU is in exactly one of three states:
- `idle` — nothing in flight
- `llm` — N ≥ 1 Ollama inference requests in flight (concurrency allowed)
- `image` — exactly one ComfyUI job in flight, Ollama models unloaded
Clients (LiteLLM, Open WebUI, n8n) point at gpu-turnstile instead of at the
services. gpu-turnstile is transparent for everything that does not touch the
GPU.
## Non-goals
- Not a scheduler across multiple GPUs or hosts. One lock, one card.
- No auth, TLS, rate limiting. Runs on an internal network behind Traefik or a
Docker bridge.
- No request rewriting, caching, or protocol translation.
- No persistence. Restart = idle state.
## Architecture
```
LiteLLM / Open WebUI ──► :11434 ─┐ ┌─► Ollama :11435
├── gpu-turnstile (1 lock) ──┤
Open WebUI / n8n ────► :8188 ───┘ └─► ComfyUI :8189
```
gpu-turnstile listens on the ports the services normally use; the actual
services run one port higher. Two listeners, one process, one lock. Each
listener is an `httputil.ReverseProxy` to its upstream. Websocket upgrades
(ComfyUI `/ws`) and streaming bodies (Ollama NDJSON / SSE) must pass through
unbuffered (`FlushInterval = -1`).
### Lock semantics
Two-mode lock with image priority (writer-preferring RW lock, where "readers"
are LLM requests and the single "writer" is an image job):
- **LLM request** (see endpoint list): `AcquireLLM()` blocks while state is
`image` **or while an image job is waiting**. Then state := `llm`, n++.
On completion (response fully written, including streamed bodies, or client
disconnect) n--; if n == 0 state := `idle`.
- **Image job**: `AcquireImage()` marks "image pending" (so no new LLM
requests start), waits until n == 0, sets state := `image`. Released after
the ComfyUI job finished and models were freed.
- Concurrent image jobs queue FIFO behind each other.
- All waits are context-aware: a client that disconnects while waiting is
removed from the queue.
### Endpoint classification
Ollama listener (`:11434``OLLAMA_URL`):
| Path | Handling |
|---|---|
| `POST /api/generate`, `/api/chat`, `/api/embed`, `/api/embeddings` | LLM lock |
| `POST /v1/chat/completions`, `/v1/completions`, `/v1/embeddings` | LLM lock |
| everything else (`/api/tags`, `/api/ps`, `/api/show`, `/api/version`, `/v1/models`, `/api/pull`, …) | pass-through, no lock |
ComfyUI listener (`:8188``COMFY_URL`):
| Path | Handling |
|---|---|
| `POST /prompt` | image lock (see flow below) |
| everything else (`/ws`, `/history/*`, `/view`, `/system_stats`, `/queue`, `/free`, …) | pass-through, no lock |
### Image job flow (`POST /prompt`)
1. `AcquireImage()`.
2. Unload Ollama: `GET /api/ps`; for each model `POST /api/generate
{"model":M,"keep_alive":0}`; if that returns non-2xx (embedding-only
models), `POST /api/embed {"model":M,"input":"x","keep_alive":0}`. Poll
`/api/ps` every `UNLOAD_POLL_INTERVAL` (default 500 ms) until empty or
`UNLOAD_TIMEOUT`. On timeout: log and
continue (degrade, don't fail the user's request).
3. Forward the original request body to ComfyUI `/prompt`, return status,
headers and body to the caller unchanged, flush.
4. If the response is 200 and contains `prompt_id`: in a goroutine, poll
`GET /history/<prompt_id>` every `HISTORY_POLL_INTERVAL` (default 1 s)
until the entry has
`status.completed == true`, `status.status_str == "error"`, or
`JOB_TIMEOUT`. Then `POST /free {"unload_models":true,"free_memory":true}`.
Then release the image lock.
5. If the response is not 200 or has no `prompt_id`: release the lock
immediately.
Optional (config flag `WARM_MODEL`): after releasing the image lock, if the
state is `idle`, send `POST /api/generate {"model":WARM_MODEL,"keep_alive":-1}`
with empty prompt to reload the chat model so the next chat doesn't pay the
load time. Off by default.
## Configuration (env)
| Var | Default | Meaning |
|---|---|---|
| `LISTEN_OLLAMA` | `:11434` | Ollama-facing listener |
| `LISTEN_COMFY` | `:8188` | ComfyUI-facing listener |
| `OLLAMA_URL` | `http://127.0.0.1:11435` | upstream |
| `COMFY_URL` | `http://127.0.0.1:8189` | upstream |
| `UNLOAD_TIMEOUT` | `60s` | wait for Ollama to unload |
| `JOB_TIMEOUT` | `15m` | wait for ComfyUI job |
| `LLM_WAIT_TIMEOUT` | `10m` | max time an LLM request waits for the lock before 503 |
| `WARM_MODEL` | `` | optional model to reload after an image job |
| `LOG_LEVEL` | `info` | `debug` logs every lock transition |
| `UNLOAD_POLL_INTERVAL` | `500ms` | `/api/ps` poll interval while unloading |
| `HISTORY_POLL_INTERVAL` | `1s` | `/history/<id>` poll interval while a job runs |
| `PROBE_TIMEOUT` | `5s` | startup probe of both upstreams |
| `FREE_TIMEOUT` | `30s` | `POST /free` call after an image job |
| `WARM_TIMEOUT` | `2m` | warm-model reload after an image job |
| `SHUTDOWN_TIMEOUT` | `10s` | graceful shutdown on SIGINT/SIGTERM |
| `PROMPT_CAPTURE_LIMIT` | `65536` | bytes of the `/prompt` response buffered to find `prompt_id` (pass-through is unaffected) |
Startup fails fast on unparsable values. Both upstreams are probed once at
start (`/api/version`, `/system_stats`); failure is logged, not fatal.
## Observability
- `GET /healthz` on both listeners: 200 with JSON
`{"state":"idle|llm|image","llm_inflight":N,"image_pending":B}`.
- `GET /metrics` on the Ollama listener: Prometheus text format, no external
dependency needed:
`gpu_turnstile_state{state="…"} 1`, `gpu_turnstile_llm_inflight`,
`gpu_turnstile_image_jobs_total`, `gpu_turnstile_lock_wait_seconds`
(histogram, label `kind="llm|image"`), `gpu_turnstile_unload_seconds`.
- Structured logs (`log/slog`, JSON when `LOG_FORMAT=json`), one line per
state transition and per image job phase with `prompt_id`.
## Edge cases to handle
- Client disconnects while streaming an Ollama response: request context is
cancelled, proxy aborts upstream, in-flight counter still decrements.
- Client disconnects while waiting for the lock: removed from wait, no
counter change.
- ComfyUI job finishes but `/history` never shows it (e.g. ComfyUI restarted):
`JOB_TIMEOUT` releases the lock; log at warn.
- Ollama unreachable during unload: continue with the image job; the whole
point is not to block users on a misbehaving neighbour.
- `POST /prompt` with a body that ComfyUI rejects (400): lock released
immediately, body passed back.
- Websocket `/ws` connections are long-lived and never take the lock.
- The Ollama OpenAI-compatible endpoints stream SSE; the proxy must not buffer.
## Repository layout
```
gpu-turnstile/
cmd/gpu-turnstile/main.go # wiring, config, listeners
internal/lock/lock.go # two-mode lock + tests
internal/ollama/client.go # ps / unload / warm
internal/comfy/client.go # history poll / free
internal/proxy/ # handlers for both listeners
Dockerfile
.gitea/workflows/ci.yml
README.md
SPEC.md # this file
```
`main.go` from the first prototype (ComfyUI-only) is the starting point for
`internal/comfy` and the `/prompt` handler; the lock and the Ollama listener
are new.
## Testing
- `internal/lock`: table tests plus a race test (`go test -race`) with
goroutines: image waits for LLMs to drain; new LLMs block while image is
pending; FIFO for images; context cancellation removes waiters.
- `internal/proxy`: `httptest.Server` fakes for Ollama (`/api/ps`,
`/api/generate`) and ComfyUI (`/prompt`, `/history/:id`, `/free`); assert
the call sequence for one image job and that a concurrent `/api/chat` is
held until `/free` was called.
- Streaming test: fake Ollama emits chunks with delays; assert the client
receives the first chunk before the last is sent (no buffering).
## Build and CI
- Go 1.23+, stdlib only. `CGO_ENABLED=0`, `-ldflags="-s -w"`, version from
`git describe` injected via `-X main.version=`.
- Dockerfile: multi-stage, final image `gcr.io/distroless/static` (or
`scratch`), non-root user, `EXPOSE 8188 11434`,
`ENTRYPOINT ["/gpu-turnstile"]`.
- `.gitea/workflows/ci.yml` (Gitea Actions):
1. on every push: `go vet`, `go test -race ./...`, `golangci-lint` if
available in the runner image
2. on a version tag only (`vX.Y.Z`, enforced): build the image with buildx
and push it to the Gitea registry
`git.rambossek.at/<owner>/gpu-turnstile` tagged `:<tag>` and `:latest`
(the repository path is lowercased in the workflow; Docker registry
names must be lowercase).
Login uses the repo secret `REGISTRY_TOKEN` (an access token with
`write:package` scope) because the automatic `GITEA_TOKEN` cannot push
packages; the username is just `gitea.actor`.
- Release: a git tag `vX.Y.Z` produces the versioned image; the Open WebUI
compose pins that tag. No images are built from branches.
## Deployment (target)
```yaml
gpu-turnstile:
image: git.rambossek.at/<owner>/gpu-turnstile:v0.1.0 # owner lowercased, e.g. "public"
environment:
OLLAMA_URL: http://<workstation-ip>:11435
COMFY_URL: http://<workstation-ip>:8189
networks: [internal]
```
LiteLLM `api_base` → `http://gpu-turnstile:11434`; Open WebUI
`COMFYUI_BASE_URL` → `http://gpu-turnstile:8188`. Nothing else talks to the
workstation directly.
## Open questions
- Should embedding requests (`/api/embed`, `/v1/embeddings`) count as LLM
traffic for the lock? They do in this spec (they hold VRAM); reconsider if
RAG indexing starves image jobs for too long.
- Whether to add a `POST /gpu-turnstile/release` admin endpoint to force-reset
the lock without restarting. Cheap to add; decide once it's been stuck once.