gpu-turnstile
GPU arbitration proxy for Ollama + ComfyUI. One consumer GPU is shared by an
LLM server (Ollama) and an image generator (ComfyUI); gpu-turnstile sits in
front of both and guarantees the GPU is always in exactly one of three states:
idle, llm (N ≥ 1 Ollama requests in flight), or image (exactly one
ComfyUI job, Ollama models unloaded). See SPEC.md for the full
design.
gpu-turnstile listens on the ports the services normally use; the actual services run one port higher (Ollama on 11435, ComfyUI on 8189).
LiteLLM / Open WebUI ──► :11434 ─┐ ┌─► Ollama :11435
├── gpu-turnstile (1 lock) ──┤
Open WebUI / n8n ────► :8188 ───┘ └─► ComfyUI :8189
- LLM endpoints (
/api/generate,/api/chat,/api/embed,/v1/*) take the LLM lock: concurrent requests allowed, but blocked while an image job is active or waiting (image priority). Blocked requests either hang until the lock is free (LLM_BUSY_MODE=wait, default) or fail immediately with 503 (or 429) +Retry-After(LLM_BUSY_MODE=reject) — the latter lets routers like LiteLLM cool down and retry instead of holding a hung connection. POST /prompton the ComfyUI listener takes the image lock: new LLM requests block, in-flight LLMs drain, Ollama models are unloaded, the prompt is forwarded, and the lock is held until the job finishes and ComfyUI frees its VRAM.- Everything else (including websockets and all streaming) passes through transparently and unbuffered.
Each consumer is enabled by setting its URL (OLLAMA_URL, COMFY_URL) and
disabled by leaving it empty — at least one is required. With only Ollama
the proxy is a pass-through (no image jobs can arrive); with only ComfyUI
the Ollama unload/warm steps are skipped. Future consumers (e.g. local game
detection) plug into the same lock the same way.
Configuration
Configuration comes from environment variables and/or an .env-style
config file (KEY=VALUE lines, # comments). File lookup order:
-config <path> flag, then GPU_TURNSTILE_CONFIG, then
gpu-turnstile.env next to the executable. Process environment variables
override file values. Invalid values fail at startup.
| Var | Default | Meaning |
|---|---|---|
LISTEN_OLLAMA |
:11434 |
Ollama-facing listener |
LISTEN_COMFY |
:8188 |
ComfyUI-facing listener |
OLLAMA_URL |
(empty = disabled) | Ollama upstream; set to enable the Ollama consumer |
COMFY_URL |
(empty = disabled) | ComfyUI upstream; set to enable the ComfyUI consumer |
UNLOAD_TIMEOUT |
60s |
Wait for Ollama to unload before an image job |
JOB_TIMEOUT |
15m |
Wait for a ComfyUI job to finish |
LLM_WAIT_TIMEOUT |
10m |
Max lock wait for an LLM request before 503 (wait mode) |
LLM_BUSY_MODE |
wait |
wait = hold blocked LLM requests; reject = fail them immediately |
LLM_BUSY_STATUS |
503 |
HTTP status for rejected LLM requests in reject mode (400–599, e.g. 429) |
BUSY_RETRY_AFTER |
30 |
Seconds sent as Retry-After on busy responses (both modes) |
WARM_MODEL |
(empty) | Model to reload after an image job (off by default) |
LOGLEVEL |
warn |
info logs every request (colored arrows in text mode), debug adds lock transitions. LOG_LEVEL works as an alias |
LOG_FORMAT |
text |
json for structured JSON logs |
LOG_FILE |
(empty) | Append logs to this file instead of stderr |
UNLOAD_POLL_INTERVAL |
500ms |
/api/ps poll interval while unloading |
HISTORY_POLL_INTERVAL |
1s |
/history/<id> poll interval while a job runs |
PROBE_TIMEOUT |
5s |
Startup probe of both upstreams |
FREE_TIMEOUT |
30s |
POST /free call after an image job |
WARM_TIMEOUT |
2m |
Warm-model reload after an image job |
SHUTDOWN_TIMEOUT |
10s |
Graceful shutdown on SIGINT/SIGTERM |
BACKOFF_INITIAL |
1s |
First retry wait when an upstream refuses a connection |
BACKOFF_MAX |
60s |
Cap for the exponential retry backoff |
PROMPT_CAPTURE_LIMIT |
65536 |
Bytes of the /prompt response buffered to find prompt_id (pass-through is unaffected) |
AUTO_UPDATE |
true |
Poll the Gitea releases API for signed updates |
UPDATE_INTERVAL |
6h |
Auto-update check interval |
UPDATE_REPO |
https://git.rambossek.at/PUBLIC/gpu-turnstile |
Repository checked for releases |
UPDATE_ASSET |
gpu-turnstile.exe |
Release asset to download |
Observability
GET /healthz(both listeners):{"state":"idle|llm|image","llm_inflight":N,"image_pending":B}GET /metrics(both listeners): Prometheus text format —gpu_turnstile_state,gpu_turnstile_llm_inflight,gpu_turnstile_image_pending,gpu_turnstile_image_jobs_total,gpu_turnstile_lock_wait_seconds(histogram,kind="llm|image"),gpu_turnstile_unload_seconds.- Logs: startup logs the version and every setting (visible even at the
default
warnlevel). WithLOGLEVEL=infoordebug, every request logs a-->incoming line and a<--response line with status and duration — ANSI-colored (cyan incoming; green/yellow/red by status class) in text mode, which renders indocker compose logson Windows Terminal. SetNO_COLORto disable colors.
Build and run
go build ./cmd/gpu-turnstile
OLLAMA_URL=http://127.0.0.1:11435 COMFY_URL=http://127.0.0.1:8189 ./gpu-turnstile
Running the binary with no arguments in a terminal prints the help screen
(same as -h/--help); without a terminal (services, containers) a bare
invocation starts the proxy.
Run natively on Windows (current primary deployment)
Download gpu-turnstile.exe from a release and install it as a Windows
service — no admin shell needed, a UAC prompt appears automatically and
the elevated child does the work (its window waits for Enter so you can
read the result):
gpu-turnstile.exe --install-service # installs into Program Files, auto-start
gpu-turnstile.exe --install-service --no-copy # register in place instead
gpu-turnstile.exe --remove-service
Layout: C:\Program Files\gpu-turnstile\ holds the exe and
gpu-turnstile.env, logs go to C:\ProgramData\gpu-turnstile\ (set
LOG_FILE in the env file — there is no console). The service always runs
as the virtual account NT SERVICE\gpu-turnstile (low-privilege,
per-service, no password); the installer automatically grants it write
access to the install and data directories — nothing else to do.
Re-running --install-service is safe: it stops a running service,
replaces the installed binary only if it changed, fixes the registration
only where it drifted, and restarts the service only if it was running.
Run natively on Linux (systemd)
The same binary works on Linux. Install it as a systemd service as root:
gpu-turnstile --install-service # installs into /var/lib/gpu-turnstile, enables + starts
gpu-turnstile --install-service --no-copy # register in place instead
gpu-turnstile --remove-service
The unit (/etc/systemd/system/gpu-turnstile.service) is Type=notify:
systemctl start blocks until the listeners are actually bound, a 30 s
watchdog restarts the process if it wedges, and logs land in the journal
(journalctl -u gpu-turnstile -f) unless LOG_FILE is set. Install
copies the binary to /var/lib/gpu-turnstile/ and the config to
/etc/gpu-turnstile.env (edit that one after installing). The service
runs sandboxed with DynamicUser=yes — a transient low-privilege UID,
read-only filesystem except its install dir (so self-update keeps
working), no capabilities, syscall-filtered: same least-privilege idea as
the Windows virtual account. The notify integration is a no-op in
containers and interactive shells.
Auto-update is on by default: the binary checks the repo's latest
release on startup and every UPDATE_INTERVAL, verifies the Ed25519
signature of the download against the public key embedded at build time,
and — once the GPU lock is idle — restarts the service onto the new
version. Disable with AUTO_UPDATE=false. Releases are signed by CI with
OpenSSL; the matching public key lives in internal/update/pubkey.go
(one-time setup: openssl genpkey -algorithm ed25519 -out private.pem,
openssl pkey -in private.pem -pubout -out public.pem; private key goes
to the RELEASE_SIGNING_KEY repo secret, public key is committed).
gpu-turnstile --force-update checks immediately, stages the new binary
and restarts the running service (elevating via UAC only if needed).
Docker
docker build -t gpu-turnstile .
docker run --rm -p 11434:11434 -p 8188:8188 \
-e OLLAMA_URL=http://<workstation-ip>:11435 \
-e COMFY_URL=http://<workstation-ip>:8189 \
gpu-turnstile
Releases are built by Gitea Actions (.gitea/workflows/ci.yml): every push
runs go vet and go test -race, and pushing a semantic-version tag
vX.Y.Z publishes the container image
(git.rambossek.at/<owner>/gpu-turnstile:vX.Y.Z plus :latest) and a
signed Windows binary attached to a Gitea release. Nothing is built from
branches.
The registry login needs one repository secret (Settings → Actions →
Secrets): REGISTRY_TOKEN — an access token with write:package scope.
The automatic GITEA_TOKEN cannot push packages.
Development
go vet ./...
go test -race ./...
Go 1.23+; the only external dependency is golang.org/x/sys (Windows
service integration, unused in the Linux build). Layout:
cmd/gpu-turnstile/main.go wiring, config, listeners, service + updater
internal/lock/ two-mode lock (LLM readers / image writer, FIFO)
internal/ollama/ ps / unload / warm client
internal/comfy/ history poll / free client
internal/proxy/ handlers for both listeners
internal/metrics/ Prometheus exposition, no dependencies
internal/config/ env + .env file configuration
internal/update/ signed auto-updater
internal/service/ Windows service integration