Files
gpu-turnstile/README.md
T
mram 5a0e181293
ci / test (push) Successful in 47s
ci / docker (push) Failing after 23s
CI: build release images only from vX.Y.Z tags
Tests still run on every push. The docker job is gated on ref_type=tag,
validates the tag is a strict vX.Y.Z semver, and publishes :<tag> plus
:latest. No branch images (sha-* / main latest) anymore.
2026-09-20 18:12:00 +02:00

3.5 KiB

gpu-turnstile

GPU arbitration proxy for Ollama + ComfyUI. One consumer GPU is shared by an LLM server (Ollama) and an image generator (ComfyUI); gpu-turnstile sits in front of both and guarantees the GPU is always in exactly one of three states: idle, llm (N ≥ 1 Ollama requests in flight), or image (exactly one ComfyUI job, Ollama models unloaded). See SPEC.md for the full design.

gpu-turnstile listens on the ports the services normally use; the actual services run one port higher (Ollama on 11435, ComfyUI on 8189).

LiteLLM / Open WebUI ──► :11434 ─┐                           ┌─► Ollama  :11435
                                 ├── gpu-turnstile (1 lock) ──┤
Open WebUI / n8n ────► :8188  ───┘                           └─► ComfyUI :8189
  • LLM endpoints (/api/generate, /api/chat, /api/embed, /v1/*) take the LLM lock: concurrent requests allowed, but blocked while an image job is active or waiting (image priority).
  • POST /prompt on the ComfyUI listener takes the image lock: new LLM requests block, in-flight LLMs drain, Ollama models are unloaded, the prompt is forwarded, and the lock is held until the job finishes and ComfyUI frees its VRAM.
  • Everything else (including websockets and all streaming) passes through transparently and unbuffered.

Configuration

All configuration is via environment variables; invalid values fail at startup.

Var Default Meaning
LISTEN_OLLAMA :11434 Ollama-facing listener
LISTEN_COMFY :8188 ComfyUI-facing listener
OLLAMA_URL http://127.0.0.1:11435 Ollama upstream
COMFY_URL http://127.0.0.1:8189 ComfyUI upstream
UNLOAD_TIMEOUT 60s Wait for Ollama to unload before an image job
JOB_TIMEOUT 15m Wait for a ComfyUI job to finish
LLM_WAIT_TIMEOUT 10m Max lock wait for an LLM request before 503
WARM_MODEL (empty) Model to reload after an image job (off by default)
LOG_LEVEL info debug logs every lock transition
LOG_FORMAT text json for structured JSON logs

Observability

  • GET /healthz (both listeners): {"state":"idle|llm|image","llm_inflight":N,"image_pending":B}
  • GET /metrics (Ollama listener): Prometheus text format — gpu_turnstile_state, gpu_turnstile_llm_inflight, gpu_turnstile_image_pending, gpu_turnstile_image_jobs_total, gpu_turnstile_lock_wait_seconds (histogram, kind="llm|image"), gpu_turnstile_unload_seconds.

Build and run

go build ./cmd/gpu-turnstile
./gpu-turnstile
docker build -t gpu-turnstile .
docker run --rm -p 11434:11434 -p 8188:8188 \
  -e OLLAMA_URL=http://<workstation-ip>:11435 \
  -e COMFY_URL=http://<workstation-ip>:8189 \
  gpu-turnstile

Releases are built by Gitea Actions (.gitea/workflows/ci.yml): every push runs go vet and go test -race, and pushing a semantic-version tag vX.Y.Z builds and publishes git.rambossek.at/<owner>/gpu-turnstile:vX.Y.Z (and updates :latest). No images are built from branches.

Development

go vet ./...
go test -race ./...

Stdlib only, Go 1.23+. Layout:

cmd/gpu-turnstile/main.go  wiring, config, listeners
internal/lock/             two-mode lock (LLM readers / image writer, FIFO)
internal/ollama/           ps / unload / warm client
internal/comfy/            history poll / free client
internal/proxy/            handlers for both listeners
internal/metrics/          Prometheus exposition, no dependencies