gpu-turnstile

GPU arbitration proxy for Ollama + ComfyUI. One consumer GPU is shared by an LLM server (Ollama) and an image generator (ComfyUI); gpu-turnstile sits in front of both and guarantees the GPU is always in exactly one of four states: idle, llm (N ≥ 1 Ollama requests in flight), image (exactly one ComfyUI job, Ollama models unloaded), or external (a foreign process such as a game holds the GPU). See SPEC.md for the full design.

gpu-turnstile listens on the ports the services normally use; the actual services run one port higher (Ollama on 11435, ComfyUI on 8189).

LiteLLM / Open WebUI ──► :11434 ─┐                           ┌─► Ollama  :11435
                                 ├── gpu-turnstile (1 lock) ──┤
Open WebUI / n8n ────► :8188  ───┘                           └─► ComfyUI :8189
  • LLM endpoints (/api/generate, /api/chat, /api/embed, /v1/*) take the LLM lock: concurrent requests allowed, but blocked while an image job is active or waiting (image priority). Blocked requests either hang until the lock is free (LLM_BUSY_MODE=wait, default) or fail immediately with 503 (or 429) + Retry-After (LLM_BUSY_MODE=reject) — the latter lets routers like LiteLLM cool down and retry instead of holding a hung connection.
  • POST /prompt on the ComfyUI listener takes the image lock: new LLM requests block, in-flight LLMs drain, Ollama models are unloaded, the prompt is forwarded, and the lock is held until the job finishes and ComfyUI frees its VRAM.
  • Everything else (including websockets and all streaming) passes through transparently and unbuffered.

Each consumer is enabled by setting its URL (OLLAMA_URL, COMFY_URL) and disabled by leaving it empty — at least one is required. With only Ollama the proxy is a pass-through (no image jobs can arrive); with only ComfyUI the Ollama unload/warm steps are skipped. A third, URL-less consumer — detection of foreign GPU holders such as games — is enabled by GAME_PROCS and/or GPU_FOREIGN_VRAM_MB (see below).

Configuration

Configuration comes from environment variables and/or an .env-style config file (KEY=VALUE lines, # comments). File lookup order: -config <path> flag, then GPU_TURNSTILE_CONFIG, then gpu-turnstile.env next to the executable. Process environment variables override file values. Invalid values fail at startup.

Var Default Meaning
LISTEN_OLLAMA :11434 Listener for Ollama-compatible clients
LISTEN_COMFY :8188 Listener for ComfyUI clients
OLLAMA_URL (empty = disabled) Ollama upstream; set to enable the Ollama consumer
COMFY_URL (empty = disabled) ComfyUI upstream; set to enable the ComfyUI consumer
UNLOAD_TIMEOUT 60s Wait for Ollama to unload before an image job
JOB_TIMEOUT 15m Wait for a ComfyUI job to finish
LLM_WAIT_TIMEOUT 10m Max lock wait for an LLM request before 503 (wait mode)
LLM_BUSY_MODE wait wait = hold blocked LLM requests; reject = fail them immediately
LLM_BUSY_STATUS 503 HTTP status for rejected LLM requests in reject mode (400599, e.g. 429)
BUSY_RETRY_AFTER 30 Seconds sent as Retry-After on busy responses (both modes)
WARM_MODEL (empty) Model to reload after an image job (off by default)
COMFY_CMD (derived from COMFY_DIR; both empty = unmanaged) Supervise ComfyUI: start on demand, stop when idle to free VRAM. Requires COMFY_URL
COMFY_DIR (empty) Standard venv install root: set alone to supervise ComfyUI with the derived command (.venv + main.py); also the working directory for COMFY_CMD
COMFY_IDLE_TIMEOUT 5m Stop the managed ComfyUI after this long idle
COMFY_START_TIMEOUT 2m Max wait for the managed ComfyUI to come up
GAME_PROCS (empty = disabled) Process names (comma-separated); while any runs, the GPU counts as held: requests wait, Ollama unloads, managed ComfyUI stops
GPU_FOREIGN_VRAM_MB 0 (disabled) Also treat the GPU as held when a non-ignored process uses more VRAM than this (needs nvidia-smi)
GPU_IGNORE_PROCS ollama,ollama app,ollama_llama_server,python,pythonw Process names never counted as foreign GPU users
GAME_POLL_INTERVAL 15s How often game/VRAM detection runs (don't go below ~10s — nvidia-smi polls keep the GPU awake)
LOGLEVEL warn info logs every request (colored arrows in text mode), debug adds lock transitions. LOG_LEVEL works as an alias
LOG_FORMAT text json for structured JSON logs
LOG_FILE (empty) Append logs to this file instead of stderr
UNLOAD_POLL_INTERVAL 500ms /api/ps poll interval while unloading
HISTORY_POLL_INTERVAL 1s /history/<id> poll interval while a job runs
PROBE_TIMEOUT 5s Startup probe of both upstreams (also per-probe health check timeout)
HEALTH_INTERVAL 30s Periodic upstream probe; down/recovered changes are logged
FREE_TIMEOUT 30s POST /free call after an image job
WARM_TIMEOUT 2m Warm-model reload after an image job
SHUTDOWN_TIMEOUT 10s Graceful shutdown on SIGINT/SIGTERM
BACKOFF_INITIAL 1s First retry wait when an upstream refuses a connection
BACKOFF_MAX 60s Cap for the exponential retry backoff
PROMPT_CAPTURE_LIMIT 65536 Bytes of the /prompt response buffered to find prompt_id (pass-through is unaffected)
AUTO_UPDATE true Poll the Gitea releases API for signed updates
UPDATE_INTERVAL 6h Auto-update check interval
UPDATE_REPO https://git.rambossek.at/PUBLIC/gpu-turnstile Repository checked for releases
UPDATE_ASSET gpu-turnstile.exe Release asset to download
APP_VER stable dev disables updates, stable tracks latest, or pin an exact vX.Y.Z

Observability

  • GET /healthz (both listeners): {"state":"idle|llm|image|external","llm_inflight":N,"image_pending":B}
  • GET /metrics (both listeners): Prometheus text format — gpu_turnstile_state, gpu_turnstile_llm_inflight, gpu_turnstile_image_pending, gpu_turnstile_image_jobs_total, gpu_turnstile_lock_wait_seconds (histogram, kind="llm|image"), gpu_turnstile_unload_seconds.
  • Logs: startup logs the version and every setting (visible even at the default warn level). With LOGLEVEL=info or debug, every request logs a --> incoming line and a <-- response line with status and duration — ANSI-colored (cyan incoming; green/yellow/red by status class) in text mode, which renders in docker compose logs on Windows Terminal. Set NO_COLOR to disable colors.

Managed ComfyUI (COMFY_CMD / COMFY_DIR)

Don't want ComfyUI running 24/7 (it holds VRAM even when idle — and the Desktop app kills its server when you close it)? gpu-turnstile can supervise it: the first request starts it, it stops again after COMFY_IDLE_TIMEOUT (default 5m) without work, freeing the GPU for games or the LLM.

The easy way — point COMFY_DIR at a standard venv install (a folder with .venv and main.py, or .venv and ComfyUI\main.py) and the launch command is derived from it, including --port from COMFY_URL:

COMFY_URL=http://127.0.0.1:8189
COMFY_DIR=C:\ComfyUI

For other layouts, spell the command out yourself (run it once manually to confirm it works):

COMFY_URL=http://127.0.0.1:8189
COMFY_CMD="C:\ComfyUI\.venv\Scripts\python.exe" main.py --port 8189
COMFY_DIR=C:\ComfyUI

POST /prompt waits for the server to answer before taking the GPU lock (LLM traffic keeps flowing while torch loads); crashes are logged and the next request respawns. Shutting gpu-turnstile down stops the child too.

Running the ComfyUI Desktop app alongside is safe: if something already answers on the port, gpu-turnstile just uses it instead of spawning (and never kills it — it only ever stops its own child). If the managed instance already holds the port when you open the desktop app, the desktop's server is the one that fails to bind.

Game detection

Want to game on the same GPU without Ollama/ComfyUI squatting on the VRAM? gpu-turnstile can watch for foreign GPU holders and, while one is active, make LLM/image requests wait (or 503, per LLM_BUSY_MODE), unload Ollama's models and stop the managed ComfyUI so the game gets the memory. Two detection paths, each optional, polled every GAME_POLL_INTERVAL (15s):

GAME_PROCS=cyberpunk2077.exe,bg3.exe     # the reliable way on Windows
GPU_FOREIGN_VRAM_MB=1024                 # catch-all via nvidia-smi

GAME_PROCS matches running process names (case-insensitive, .exe optional). GPU_FOREIGN_VRAM_MB asks nvidia-smi which processes hold GPU memory and treats anything not in GPU_IGNORE_PROCS above the threshold as foreign — handy as a catch-all, but note that under Windows' WDDM driver graphics-only games may not show up in nvidia-smi's per-process list, so name your games in GAME_PROCS there; on Linux both paths work. When the game exits, requests resume automatically.

Build and run

go build ./cmd/gpu-turnstile
OLLAMA_URL=http://127.0.0.1:11435 COMFY_URL=http://127.0.0.1:8189 ./gpu-turnstile

Running the binary with no arguments in a terminal prints the help screen (same as -h/--help); without a terminal (services, containers) a bare invocation starts the proxy.

Run natively on Windows (current primary deployment)

Download gpu-turnstile.exe from a release and install it as a Windows service — no admin shell needed, a UAC prompt appears automatically and the elevated child does the work (its window waits for Enter so you can read the result):

gpu-turnstile.exe --install-service        # installs into Program Files, auto-start
gpu-turnstile.exe --install-service --no-copy   # register in place instead
gpu-turnstile.exe --remove-service

Layout: C:\Program Files\gpu-turnstile\ holds the exe and gpu-turnstile.env, logs go to C:\ProgramData\gpu-turnstile\. If you install without a config, the installer writes a sample env file with every setting commented and explained — only LOG_FILE is active (a service has no console). Your own LOG_FILE setting is always kept. The service always runs as the virtual account NT SERVICE\gpu-turnstile (low-privilege, per-service, no password); the installer automatically grants it write access to the install and data directories — nothing else to do.

Re-running --install-service is safe: it stops a running service, replaces the installed binary only if it changed, fixes the registration only where it drifted, and restarts the service only if it was running.

Run natively on Linux (systemd)

The same binary works on Linux. Install it as a systemd service as root:

gpu-turnstile --install-service            # installs into /var/lib/gpu-turnstile, enables + starts
gpu-turnstile --install-service --no-copy  # register in place instead
gpu-turnstile --remove-service

The unit (/etc/systemd/system/gpu-turnstile.service) is Type=notify: systemctl start blocks until the listeners are actually bound, a 30 s watchdog restarts the process if it wedges, and logs land in the journal (journalctl -u gpu-turnstile -f) unless LOG_FILE is set. Install copies the binary to /var/lib/gpu-turnstile/ and the config to /etc/gpu-turnstile.env (edit that one after installing). The service runs sandboxed with DynamicUser=yes — a transient low-privilege UID, read-only filesystem except its install dir (so self-update keeps working), no capabilities, syscall-filtered: same least-privilege idea as the Windows virtual account. The notify integration is a no-op in containers and interactive shells.

Auto-update is on by default: the binary checks the repo's latest release on startup and every UPDATE_INTERVAL, verifies the Ed25519 signature of the download against the public key embedded at build time, and — once the GPU lock is idle — restarts the service onto the new version. APP_VER controls the target: dev disables updates, stable (the default) tracks the latest release, and an exact vX.Y.Z pins that release (even as a downgrade or to replace a dev build). Disable entirely with AUTO_UPDATE=false. Releases are signed by CI with OpenSSL; the matching public key lives in internal/update/pubkey.go (one-time setup: openssl genpkey -algorithm ed25519 -out private.pem, openssl pkey -in private.pem -pubout -out public.pem; private key goes to the RELEASE_SIGNING_KEY repo secret, public key is committed). gpu-turnstile --force-update checks immediately, stages the new binary and restarts the running service (elevating via UAC only if needed).

Docker

docker build -t gpu-turnstile .
docker run --rm -p 11434:11434 -p 8188:8188 \
  -e OLLAMA_URL=http://<workstation-ip>:11435 \
  -e COMFY_URL=http://<workstation-ip>:8189 \
  gpu-turnstile

Releases are built by Gitea Actions (.gitea/workflows/ci.yml): every push runs go vet and go test -race, and pushing a semantic-version tag vX.Y.Z publishes the container image (git.rambossek.at/<owner>/gpu-turnstile:vX.Y.Z plus :latest) and a signed Windows binary attached to a Gitea release. Nothing is built from branches.

The registry login needs one repository secret (Settings → Actions → Secrets): REGISTRY_TOKEN — an access token with write:package scope. The automatic GITEA_TOKEN cannot push packages.

Development

go vet ./...
go test -race ./...

Go 1.23+; the only external dependency is golang.org/x/sys (Windows service integration, unused in the Linux build). Layout:

cmd/gpu-turnstile/main.go  wiring, config, listeners, service + updater
internal/lock/             two-mode lock (LLM readers / image writer, FIFO)
internal/ollama/           ps / unload / warm client
internal/comfy/            history poll / free client
internal/proxy/            handlers for both listeners
internal/metrics/          Prometheus exposition, no dependencies
internal/config/           env + .env file configuration
internal/update/           signed auto-updater
internal/service/          Windows service integration
S
Description
No description provided
Readme
922 KiB
v0.2.3
Latest
2026-09-21 23:03:45 +02:00
Languages
Go 99.8%
Dockerfile 0.2%