gpu-turnstile

GPU arbitration proxy for Ollama + ComfyUI. One consumer GPU is shared by an LLM server (Ollama) and an image generator (ComfyUI); gpu-turnstile sits in front of both and guarantees the GPU is always in exactly one of three states: idle, llm (N ≥ 1 Ollama requests in flight), or image (exactly one ComfyUI job, Ollama models unloaded). See SPEC.md for the full design.

gpu-turnstile listens on the ports the services normally use; the actual services run one port higher (Ollama on 11435, ComfyUI on 8189).

LiteLLM / Open WebUI ──► :11434 ─┐                           ┌─► Ollama  :11435
                                 ├── gpu-turnstile (1 lock) ──┤
Open WebUI / n8n ────► :8188  ───┘                           └─► ComfyUI :8189
  • LLM endpoints (/api/generate, /api/chat, /api/embed, /v1/*) take the LLM lock: concurrent requests allowed, but blocked while an image job is active or waiting (image priority).
  • POST /prompt on the ComfyUI listener takes the image lock: new LLM requests block, in-flight LLMs drain, Ollama models are unloaded, the prompt is forwarded, and the lock is held until the job finishes and ComfyUI frees its VRAM.
  • Everything else (including websockets and all streaming) passes through transparently and unbuffered.

Configuration

All configuration is via environment variables; invalid values fail at startup.

Var Default Meaning
LISTEN_OLLAMA :11434 Ollama-facing listener
LISTEN_COMFY :8188 ComfyUI-facing listener
OLLAMA_URL http://127.0.0.1:11435 Ollama upstream
COMFY_URL http://127.0.0.1:8189 ComfyUI upstream
UNLOAD_TIMEOUT 60s Wait for Ollama to unload before an image job
JOB_TIMEOUT 15m Wait for a ComfyUI job to finish
LLM_WAIT_TIMEOUT 10m Max lock wait for an LLM request before 503
WARM_MODEL (empty) Model to reload after an image job (off by default)
LOGLEVEL warn info logs every request (colored arrows in text mode), debug adds lock transitions. LOG_LEVEL works as an alias
LOG_FORMAT text json for structured JSON logs
UNLOAD_POLL_INTERVAL 500ms /api/ps poll interval while unloading
HISTORY_POLL_INTERVAL 1s /history/<id> poll interval while a job runs
PROBE_TIMEOUT 5s Startup probe of both upstreams
FREE_TIMEOUT 30s POST /free call after an image job
WARM_TIMEOUT 2m Warm-model reload after an image job
SHUTDOWN_TIMEOUT 10s Graceful shutdown on SIGINT/SIGTERM
BACKOFF_INITIAL 1s First retry wait when an upstream refuses a connection
BACKOFF_MAX 60s Cap for the exponential retry backoff
PROMPT_CAPTURE_LIMIT 65536 Bytes of the /prompt response buffered to find prompt_id (pass-through is unaffected)

Observability

  • GET /healthz (both listeners): {"state":"idle|llm|image","llm_inflight":N,"image_pending":B}
  • GET /metrics (Ollama listener): Prometheus text format — gpu_turnstile_state, gpu_turnstile_llm_inflight, gpu_turnstile_image_pending, gpu_turnstile_image_jobs_total, gpu_turnstile_lock_wait_seconds (histogram, kind="llm|image"), gpu_turnstile_unload_seconds.
  • Logs: startup logs the version and every setting (visible even at the default warn level). With LOGLEVEL=info or debug, every request logs a --> incoming line and a <-- response line with status and duration — ANSI-colored (cyan incoming; green/yellow/red by status class) in text mode, which renders in docker compose logs on Windows Terminal. Set NO_COLOR to disable colors.

Build and run

go build ./cmd/gpu-turnstile
./gpu-turnstile
docker build -t gpu-turnstile .
docker run --rm -p 11434:11434 -p 8188:8188 \
  -e OLLAMA_URL=http://<workstation-ip>:11435 \
  -e COMFY_URL=http://<workstation-ip>:8189 \
  gpu-turnstile

Releases are built by Gitea Actions (.gitea/workflows/ci.yml): every push runs go vet and go test -race, and pushing a semantic-version tag vX.Y.Z builds and publishes git.rambossek.at/<owner>/gpu-turnstile:vX.Y.Z (and updates :latest). No images are built from branches.

The registry login needs one repository secret (Settings → Actions → Secrets): REGISTRY_TOKEN — an access token with write:package scope. The automatic GITEA_TOKEN cannot push packages.

Development

go vet ./...
go test -race ./...

Stdlib only, Go 1.23+. Layout:

cmd/gpu-turnstile/main.go  wiring, config, listeners
internal/lock/             two-mode lock (LLM readers / image writer, FIFO)
internal/ollama/           ps / unload / warm client
internal/comfy/            history poll / free client
internal/proxy/            handlers for both listeners
internal/metrics/          Prometheus exposition, no dependencies
S
Description
No description provided
Readme
664 KiB
v0.1.8
Latest
2026-09-21 17:19:16 +02:00
Languages
Go 99.8%
Dockerfile 0.2%