# gpu-turnstile GPU arbitration proxy for Ollama + ComfyUI. One consumer GPU is shared by an LLM server (Ollama) and an image generator (ComfyUI); gpu-turnstile sits in front of both and guarantees the GPU is always in exactly one of three states: `idle`, `llm` (N ≥ 1 Ollama requests in flight), or `image` (exactly one ComfyUI job, Ollama models unloaded). See [SPEC.md](SPEC.md) for the full design. gpu-turnstile listens on the ports the services normally use; the actual services run one port higher (Ollama on 11435, ComfyUI on 8189). ``` LiteLLM / Open WebUI ──► :11434 ─┐ ┌─► Ollama :11435 ├── gpu-turnstile (1 lock) ──┤ Open WebUI / n8n ────► :8188 ───┘ └─► ComfyUI :8189 ``` - LLM endpoints (`/api/generate`, `/api/chat`, `/api/embed`, `/v1/*`) take the LLM lock: concurrent requests allowed, but blocked while an image job is active or waiting (image priority). Blocked requests either hang until the lock is free (`LLM_BUSY_MODE=wait`, default) or fail immediately with 503 (or 429) + `Retry-After` (`LLM_BUSY_MODE=reject`) — the latter lets routers like LiteLLM cool down and retry instead of holding a hung connection. - `POST /prompt` on the ComfyUI listener takes the image lock: new LLM requests block, in-flight LLMs drain, Ollama models are unloaded, the prompt is forwarded, and the lock is held until the job finishes and ComfyUI frees its VRAM. - Everything else (including websockets and all streaming) passes through transparently and unbuffered. Each consumer is enabled by setting its URL (`OLLAMA_URL`, `COMFY_URL`) and disabled by leaving it empty — at least one is required. With only Ollama the proxy is a pass-through (no image jobs can arrive); with only ComfyUI the Ollama unload/warm steps are skipped. Future consumers (e.g. local game detection) plug into the same lock the same way. ## Configuration Configuration comes from environment variables and/or an `.env`-style config file (`KEY=VALUE` lines, `#` comments). File lookup order: `-config ` flag, then `GPU_TURNSTILE_CONFIG`, then `gpu-turnstile.env` next to the executable. Process environment variables override file values. Invalid values fail at startup. | Var | Default | Meaning | |---|---|---| | `LISTEN_OLLAMA` | `:11434` | Listener for Ollama-compatible clients | | `LISTEN_COMFY` | `:8188` | Listener for ComfyUI clients | | `OLLAMA_URL` | _(empty = disabled)_ | Ollama upstream; set to enable the Ollama consumer | | `COMFY_URL` | _(empty = disabled)_ | ComfyUI upstream; set to enable the ComfyUI consumer | | `UNLOAD_TIMEOUT` | `60s` | Wait for Ollama to unload before an image job | | `JOB_TIMEOUT` | `15m` | Wait for a ComfyUI job to finish | | `LLM_WAIT_TIMEOUT` | `10m` | Max lock wait for an LLM request before 503 (wait mode) | | `LLM_BUSY_MODE` | `wait` | `wait` = hold blocked LLM requests; `reject` = fail them immediately | | `LLM_BUSY_STATUS` | `503` | HTTP status for rejected LLM requests in reject mode (400–599, e.g. 429) | | `BUSY_RETRY_AFTER` | `30` | Seconds sent as `Retry-After` on busy responses (both modes) | | `WARM_MODEL` | _(empty)_ | Model to reload after an image job (off by default) | | `COMFY_CMD` | _(empty = unmanaged)_ | Supervise ComfyUI: start on demand, stop when idle to free VRAM. Requires `COMFY_URL` | | `COMFY_DIR` | _(empty)_ | Working directory for `COMFY_CMD` | | `COMFY_IDLE_TIMEOUT` | `5m` | Stop the managed ComfyUI after this long idle | | `COMFY_START_TIMEOUT` | `2m` | Max wait for the managed ComfyUI to come up | | `LOGLEVEL` | `warn` | `info` logs every request (colored arrows in text mode), `debug` adds lock transitions. `LOG_LEVEL` works as an alias | | `LOG_FORMAT` | `text` | `json` for structured JSON logs | | `LOG_FILE` | _(empty)_ | Append logs to this file instead of stderr | | `UNLOAD_POLL_INTERVAL` | `500ms` | `/api/ps` poll interval while unloading | | `HISTORY_POLL_INTERVAL` | `1s` | `/history/` poll interval while a job runs | | `PROBE_TIMEOUT` | `5s` | Startup probe of both upstreams (also per-probe health check timeout) | | `HEALTH_INTERVAL` | `30s` | Periodic upstream probe; down/recovered changes are logged | | `FREE_TIMEOUT` | `30s` | `POST /free` call after an image job | | `WARM_TIMEOUT` | `2m` | Warm-model reload after an image job | | `SHUTDOWN_TIMEOUT` | `10s` | Graceful shutdown on SIGINT/SIGTERM | | `BACKOFF_INITIAL` | `1s` | First retry wait when an upstream refuses a connection | | `BACKOFF_MAX` | `60s` | Cap for the exponential retry backoff | | `PROMPT_CAPTURE_LIMIT` | `65536` | Bytes of the `/prompt` response buffered to find `prompt_id` (pass-through is unaffected) | | `AUTO_UPDATE` | `true` | Poll the Gitea releases API for signed updates | | `UPDATE_INTERVAL` | `6h` | Auto-update check interval | | `UPDATE_REPO` | `https://git.rambossek.at/PUBLIC/gpu-turnstile` | Repository checked for releases | | `UPDATE_ASSET` | `gpu-turnstile.exe` | Release asset to download | | `APP_VER` | `stable` | `dev` disables updates, `stable` tracks latest, or pin an exact `vX.Y.Z` | ## Observability - `GET /healthz` (both listeners): `{"state":"idle|llm|image","llm_inflight":N,"image_pending":B}` - `GET /metrics` (both listeners): Prometheus text format — `gpu_turnstile_state`, `gpu_turnstile_llm_inflight`, `gpu_turnstile_image_pending`, `gpu_turnstile_image_jobs_total`, `gpu_turnstile_lock_wait_seconds` (histogram, `kind="llm|image"`), `gpu_turnstile_unload_seconds`. - Logs: startup logs the version and every setting (visible even at the default `warn` level). With `LOGLEVEL=info` or `debug`, every request logs a `-->` incoming line and a `<--` response line with status and duration — ANSI-colored (cyan incoming; green/yellow/red by status class) in text mode, which renders in `docker compose logs` on Windows Terminal. Set `NO_COLOR` to disable colors. ## Managed ComfyUI (`COMFY_CMD`) Don't want ComfyUI running 24/7 (it holds VRAM even when idle — and the Desktop app kills its server when you close it)? Point `COMFY_CMD` at a standalone launch command and gpu-turnstile supervises it: the first request starts it, it stops again after `COMFY_IDLE_TIMEOUT` (default 5m) without work, freeing the GPU for games or the LLM. Example for a Desktop install (run it once manually to confirm it works): ``` COMFY_URL=http://127.0.0.1:8188 COMFY_CMD="C:\ComfyUI\.venv\Scripts\python.exe ComfyUI\main.py --port 8188" COMFY_DIR=C:\ComfyUI ``` `POST /prompt` waits for the server to answer before taking the GPU lock (LLM traffic keeps flowing while torch loads); crashes are logged and the next request respawns. Shutting gpu-turnstile down stops the child too. Running the ComfyUI Desktop app alongside is safe: if something already answers on the port, gpu-turnstile just uses it instead of spawning (and never kills it — it only ever stops its own child). If the managed instance already holds the port when you open the desktop app, the desktop's server is the one that fails to bind. ## Build and run ```sh go build ./cmd/gpu-turnstile OLLAMA_URL=http://127.0.0.1:11435 COMFY_URL=http://127.0.0.1:8189 ./gpu-turnstile ``` Running the binary with no arguments in a terminal prints the help screen (same as `-h`/`--help`); without a terminal (services, containers) a bare invocation starts the proxy. ### Run natively on Windows (current primary deployment) Download `gpu-turnstile.exe` from a release and install it as a Windows service — no admin shell needed, a UAC prompt appears automatically and the elevated child does the work (its window waits for Enter so you can read the result): ```sh gpu-turnstile.exe --install-service # installs into Program Files, auto-start gpu-turnstile.exe --install-service --no-copy # register in place instead gpu-turnstile.exe --remove-service ``` Layout: `C:\Program Files\gpu-turnstile\` holds the exe and `gpu-turnstile.env`, logs go to `C:\ProgramData\gpu-turnstile\`. If you install without a config, the installer writes a sample env file with every setting commented and explained — only `LOG_FILE` is active (a service has no console). Your own `LOG_FILE` setting is always kept. The service always runs as the virtual account `NT SERVICE\gpu-turnstile` (low-privilege, per-service, no password); the installer automatically grants it write access to the install and data directories — nothing else to do. Re-running `--install-service` is safe: it stops a running service, replaces the installed binary only if it changed, fixes the registration only where it drifted, and restarts the service only if it was running. ### Run natively on Linux (systemd) The same binary works on Linux. Install it as a systemd service as root: ```sh gpu-turnstile --install-service # installs into /var/lib/gpu-turnstile, enables + starts gpu-turnstile --install-service --no-copy # register in place instead gpu-turnstile --remove-service ``` The unit (`/etc/systemd/system/gpu-turnstile.service`) is `Type=notify`: `systemctl start` blocks until the listeners are actually bound, a 30 s watchdog restarts the process if it wedges, and logs land in the journal (`journalctl -u gpu-turnstile -f`) unless `LOG_FILE` is set. Install copies the binary to `/var/lib/gpu-turnstile/` and the config to `/etc/gpu-turnstile.env` (edit that one after installing). The service runs sandboxed with `DynamicUser=yes` — a transient low-privilege UID, read-only filesystem except its install dir (so self-update keeps working), no capabilities, syscall-filtered: same least-privilege idea as the Windows virtual account. The notify integration is a no-op in containers and interactive shells. **Auto-update is on by default**: the binary checks the repo's latest release on startup and every `UPDATE_INTERVAL`, verifies the Ed25519 signature of the download against the public key embedded at build time, and — once the GPU lock is idle — restarts the service onto the new version. `APP_VER` controls the target: `dev` disables updates, `stable` (the default) tracks the latest release, and an exact `vX.Y.Z` pins that release (even as a downgrade or to replace a dev build). Disable entirely with `AUTO_UPDATE=false`. Releases are signed by CI with OpenSSL; the matching public key lives in `internal/update/pubkey.go` (one-time setup: `openssl genpkey -algorithm ed25519 -out private.pem`, `openssl pkey -in private.pem -pubout -out public.pem`; private key goes to the `RELEASE_SIGNING_KEY` repo secret, public key is committed). `gpu-turnstile --force-update` checks immediately, stages the new binary and restarts the running service (elevating via UAC only if needed). ### Docker ```sh docker build -t gpu-turnstile . docker run --rm -p 11434:11434 -p 8188:8188 \ -e OLLAMA_URL=http://:11435 \ -e COMFY_URL=http://:8189 \ gpu-turnstile ``` Releases are built by Gitea Actions (`.gitea/workflows/ci.yml`): every push runs `go vet` and `go test -race`, and pushing a semantic-version tag `vX.Y.Z` publishes the container image (`git.rambossek.at//gpu-turnstile:vX.Y.Z` plus `:latest`) and a signed Windows binary attached to a Gitea release. Nothing is built from branches. The registry login needs one repository secret (Settings → Actions → Secrets): `REGISTRY_TOKEN` — an access token with `write:package` scope. The automatic `GITEA_TOKEN` cannot push packages. ## Development ```sh go vet ./... go test -race ./... ``` Go 1.23+; the only external dependency is `golang.org/x/sys` (Windows service integration, unused in the Linux build). Layout: ``` cmd/gpu-turnstile/main.go wiring, config, listeners, service + updater internal/lock/ two-mode lock (LLM readers / image writer, FIFO) internal/ollama/ ps / unload / warm client internal/comfy/ history poll / free client internal/proxy/ handlers for both listeners internal/metrics/ Prometheus exposition, no dependencies internal/config/ env + .env file configuration internal/update/ signed auto-updater internal/service/ Windows service integration ```