Files
gpu-turnstile/SPEC.md
T
mram 83812cebf3 Add Linux systemd support and --install-service/--remove-service flags
The service package now has a Linux implementation alongside the Windows
one: systemd unit install/remove (/etc/systemd/system), readiness
notification (READY=1 via go-systemd) sent only after the listeners are
bound, a 30s watchdog, and STOPPING on shutdown. Listeners are pre-bound
so port conflicts fail fast and the readiness signal is truthful. The
notify calls are no-ops without NOTIFY_SOCKET (containers, shells) and
on non-Linux builds. --install-service/--remove-service work on both
platforms; the 'service install|remove' subcommand remains as an alias.
2026-09-20 22:35:59 +02:00

18 KiB
Raw Permalink Blame History

gpu-turnstile — GPU arbitration proxy for Ollama + ComfyUI

Problem

One consumer GPU (RTX 5080, 16 GB) is shared by an LLM server (Ollama) and an image generator (ComfyUI). Both assume they own the card. When both hold models at once, the NVIDIA Windows driver falls back to system memory and everything becomes very slow; on Linux it would OOM instead.

Goal

A single Go binary that sits in front of both services and guarantees that at any moment the GPU is in exactly one of three states:

  • idle — nothing in flight
  • llm — N ≥ 1 Ollama inference requests in flight (concurrency allowed)
  • image — exactly one ComfyUI job in flight, Ollama models unloaded

Clients (LiteLLM, Open WebUI, n8n) point at gpu-turnstile instead of at the services. gpu-turnstile is transparent for everything that does not touch the GPU.

Non-goals

  • Not a scheduler across multiple GPUs or hosts. One lock, one card.
  • No auth, TLS, rate limiting. Runs on an internal network behind Traefik or a Docker bridge.
  • No request rewriting, caching, or protocol translation.
  • No persistence. Restart = idle state.

Architecture

LiteLLM / Open WebUI ──► :11434 ─┐                           ┌─► Ollama  :11435
                                 ├── gpu-turnstile (1 lock) ──┤
Open WebUI / n8n ────► :8188  ───┘                           └─► ComfyUI :8189

gpu-turnstile listens on the ports the services normally use; the actual services run one port higher. Two listeners, one process, one lock. Each listener is an httputil.ReverseProxy to its upstream. Websocket upgrades (ComfyUI /ws) and streaming bodies (Ollama NDJSON / SSE) must pass through unbuffered (FlushInterval = -1).

Modes of operation

Each GPU consumer is enabled by setting its URL and disabled by leaving it empty — no separate flags. At least one URL must be set; a disabled consumer gets no listener, no startup probe, and no lock participation:

  • Both set (default deployment): full arbitration as described below.
  • Only OLLAMA_URL: pure pass-through for Ollama; the LLM lock never blocks since no image jobs can arrive.
  • Only COMFY_URL: image jobs are tracked and ComfyUI's VRAM is freed afterwards, but the Ollama unload and warm-reload steps are skipped.
  • Future consumers (e.g. detecting a local game holding VRAM) plug into the same lock the same way: enabled by their config knob, excluded when absent.

Lock semantics

Two-mode lock with image priority (writer-preferring RW lock, where "readers" are LLM requests and the single "writer" is an image job):

  • LLM request (see endpoint list): AcquireLLM() blocks while state is image or while an image job is waiting. Then state := llm, n++. On completion (response fully written, including streamed bodies, or client disconnect) n--; if n == 0 state := idle. LLM_BUSY_MODE selects what a blocked LLM request sees: wait (default) hangs until the lock is free or LLM_WAIT_TIMEOUT expires (then 503 + Retry-After); reject answers immediately with LLM_BUSY_STATUS (default 503; 429 works too) + Retry-After: BUSY_RETRY_AFTER, which routers like LiteLLM honor for cooldowns/retries.
  • Image job: AcquireImage() marks "image pending" (so no new LLM requests start), waits until n == 0, sets state := image. Released after the ComfyUI job finished and models were freed.
  • Concurrent image jobs queue FIFO behind each other.
  • All waits are context-aware: a client that disconnects while waiting is removed from the queue.

Endpoint classification

Ollama listener (:11434OLLAMA_URL):

Path Handling
POST /api/generate, /api/chat, /api/embed, /api/embeddings LLM lock
POST /v1/chat/completions, /v1/completions, /v1/embeddings LLM lock
everything else (/api/tags, /api/ps, /api/show, /api/version, /v1/models, /api/pull, …) pass-through, no lock

ComfyUI listener (:8188COMFY_URL):

Path Handling
POST /prompt image lock (see flow below)
everything else (/ws, /history/*, /view, /system_stats, /queue, /free, …) pass-through, no lock

Image job flow (POST /prompt)

  1. AcquireImage().
  2. Unload Ollama (skipped when OLLAMA_URL is unset): GET /api/ps; for each model POST /api/generate {"model":M,"keep_alive":0}; if that returns non-2xx (embedding-only models), POST /api/embed {"model":M,"input":"x","keep_alive":0}. Poll /api/ps every UNLOAD_POLL_INTERVAL (default 500 ms) until empty or UNLOAD_TIMEOUT. On timeout: log and continue (degrade, don't fail the user's request).
  3. Forward the original request body to ComfyUI /prompt, return status, headers and body to the caller unchanged, flush.
  4. If the response is 200 and contains prompt_id: in a goroutine, poll GET /history/<prompt_id> every HISTORY_POLL_INTERVAL (default 1 s) until the entry has status.completed == true, status.status_str == "error", or JOB_TIMEOUT. Then POST /free {"unload_models":true,"free_memory":true}. Then release the image lock.
  5. If the response is not 200 or has no prompt_id: release the lock immediately.

Optional (config flag WARM_MODEL): after releasing the image lock, if the state is idle, send POST /api/generate {"model":WARM_MODEL,"keep_alive":-1} with empty prompt to reload the chat model so the next chat doesn't pay the load time. Off by default.

Configuration (env)

Configuration comes from environment variables and/or an .env-style config file (KEY=VALUE lines, # comments). File lookup order: -config <path> flag, then GPU_TURNSTILE_CONFIG, then gpu-turnstile.env next to the executable. Process environment variables override file values. A missing file is fine; a malformed one is fatal.

Var Default Meaning
LISTEN_OLLAMA :11434 Ollama-facing listener
LISTEN_COMFY :8188 ComfyUI-facing listener
OLLAMA_URL (empty = disabled) Ollama upstream; set to enable the Ollama consumer
COMFY_URL (empty = disabled) ComfyUI upstream; set to enable the ComfyUI consumer
UNLOAD_TIMEOUT 60s wait for Ollama to unload
JOB_TIMEOUT 15m wait for ComfyUI job
LLM_WAIT_TIMEOUT 10m max time an LLM request waits for the lock before 503 (wait mode)
LLM_BUSY_MODE wait wait = hold blocked LLM requests; reject = fail them immediately
LLM_BUSY_STATUS 503 HTTP status for rejected LLM requests in reject mode (400599, e.g. 429)
BUSY_RETRY_AFTER 30 seconds sent as Retry-After on busy responses (both modes)
WARM_MODEL `` optional model to reload after an image job
LOGLEVEL warn info logs every request (colored arrows in text mode), debug adds lock transitions. LOG_LEVEL is accepted as an alias
LOG_FORMAT text json for structured JSON logs
LOG_FILE `` append logs to this file instead of stderr (useful as a service)
UNLOAD_POLL_INTERVAL 500ms /api/ps poll interval while unloading
HISTORY_POLL_INTERVAL 1s /history/<id> poll interval while a job runs
PROBE_TIMEOUT 5s startup probe of both upstreams
FREE_TIMEOUT 30s POST /free call after an image job
WARM_TIMEOUT 2m warm-model reload after an image job
SHUTDOWN_TIMEOUT 10s graceful shutdown on SIGINT/SIGTERM
BACKOFF_INITIAL 1s first retry wait when an upstream refuses a connection
BACKOFF_MAX 60s cap for the exponential retry backoff
PROMPT_CAPTURE_LIMIT 65536 bytes of the /prompt response buffered to find prompt_id (pass-through is unaffected)
AUTO_UPDATE true poll the Gitea releases API for signed updates
UPDATE_INTERVAL 6h auto-update check interval
UPDATE_REPO https://git.rambossek.at/PUBLIC/gpu-turnstile repository to check for releases
UPDATE_ASSET gpu-turnstile.exe release asset to download

Startup fails fast on unparsable values and when neither consumer URL is set. Enabled upstreams are probed once at start (/api/version, /system_stats); failure is logged, not fatal.

Native deployment (Windows and Linux)

The binary runs natively on Windows (the current primary deployment) and on Linux with systemd (the future GPU server), as well as in Docker.

Service management is the same on both platforms: gpu-turnstile --install-service [-config path] registers and starts an auto-start service; --remove-service stops and unregisters it (both need an elevated/root shell). The legacy form gpu-turnstile service install|remove does the same thing.

Windows

  • --install-service registers a Windows service; recovery actions restart it 5 s after any failure.
  • Layout: install to C:\Program Files\gpu-turnstile\ (exe plus gpu-turnstile.env); logs belong in C:\ProgramData\gpu-turnstile\ via LOG_FILE. The service must be able to write its install directory for self-updates — Program Files is writable by LocalSystem and admins, which is why running as the default LocalSystem account is the simple choice.
  • Account: the default LocalSystem works out of the box. For least privilege, create the service with the virtual account NT SERVICE\gpu-turnstile and grant it write access to the install and log directories only (no network logon, no user profile).
  • Use a config file (above) for the service — Windows services have no convenient environment. Logs go to LOG_FILE since there is no console.

Linux (systemd)

  • --install-service writes /etc/systemd/system/gpu-turnstile.service with ExecStart pointing at the current executable and the -config file, then runs systemctl daemon-reload and enable --now. The unit runs as root (it must be able to overwrite its own binary for self-updates); harden with ProtectSystem=strict plus a writable ReadWritePaths if desired.
  • The unit is Type=notify: the binary sends READY=1 via github.com/coreos/go-systemd only after the listeners are bound, so systemctl start blocks until the proxy accepts connections. A 30 s watchdog (WatchdogSec=) is pinged as long as the process runs; three missed pings make systemd restart it. STOPPING=1 is sent on shutdown. All notify calls are no-ops when NOTIFY_SOCKET is unset (containers, interactive shells), and the whole integration is Linux-only — Windows builds carry no-op stubs.
  • Logs go to the journal (journalctl -u gpu-turnstile) or to LOG_FILE.
  • Auto-update works the same as on Windows: Restart=on-failure with RestartSec=5s brings up the staged binary after the updater exits with code 3.
  • Auto-update: on startup and every UPDATE_INTERVAL, the binary checks UPDATE_REPO's latest release; if its tag is a newer vX.Y.Z, it downloads UPDATE_ASSET plus its .sig (and .sha256 when present) and verifies an Ed25519 signature against the public key embedded in internal/update/pubkey.go. A verified binary is swapped in next to the running exe (rename-aside, allowed on Windows), and once the GPU lock is idle the process exits with code 3 so the service recovery restarts it on the new version. Interactive runs only log "restart to apply". dev builds and builds without an embedded public key never update.
  • Signing setup (one time): openssl genpkey -algorithm ed25519 -out private.pem; openssl pkey -in private.pem -pubout -out public.pem. Private key → repo secret RELEASE_SIGNING_KEY; public key → committed into internal/update/pubkey.go. CI signs release binaries with openssl pkeyutl -sign -rawin.

Observability

  • GET /healthz on both listeners: 200 with JSON {"state":"idle|llm|image","llm_inflight":N,"image_pending":B}.
  • GET /metrics on both listeners: Prometheus text format, no external dependency needed: gpu_turnstile_state{state="…"} 1, gpu_turnstile_llm_inflight, gpu_turnstile_image_jobs_total, gpu_turnstile_lock_wait_seconds (histogram, label kind="llm|image"), gpu_turnstile_unload_seconds.
  • Structured logs (log/slog, JSON when LOG_FORMAT=json), one line per state transition and per image job phase with prompt_id. Startup logs the version and every setting (visible even at the default warn level). With LOGLEVEL=info or debug, every request logs a --> incoming line and a <-- response line with status and duration — ANSI-colored (cyan incoming; green/yellow/red by status class) in text mode, which renders in docker compose logs on Windows Terminal. Set NO_COLOR to disable colors.

Edge cases to handle

  • Client disconnects while streaming an Ollama response: request context is cancelled, proxy aborts upstream, in-flight counter still decrements.
  • Client disconnects while waiting for the lock: removed from wait, no counter change.
  • ComfyUI job finishes but /history never shows it (e.g. ComfyUI restarted): JOB_TIMEOUT releases the lock; log at warn.
  • Ollama unreachable during unload: continue with the image job; the whole point is not to block users on a misbehaving neighbour.
  • Upstream unreachable while proxying (connection refused, dial timeout, DNS failure, TLS handshake error): retry with exponential backoff — BACKOFF_INITIAL, doubling per attempt, capped at BACKOFF_MAX — until the upstream answers or the client disconnects. These are safe to retry: the request never reached the upstream application. 5xx responses are retried the same way, but only when the request body can be replayed (GETs, or bodies with GetBody); streamed POSTs are never replayed to avoid duplicate work such as a double-enqueued ComfyUI prompt.
  • POST /prompt with a body that ComfyUI rejects (400): lock released immediately, body passed back.
  • Websocket /ws connections are long-lived and never take the lock.
  • The Ollama OpenAI-compatible endpoints stream SSE; the proxy must not buffer.

Repository layout

gpu-turnstile/
  cmd/gpu-turnstile/main.go  # wiring, config, listeners, service + updater
  internal/lock/lock.go      # two-mode lock + tests
  internal/ollama/client.go  # ps / unload / warm
  internal/comfy/client.go   # history poll / free
  internal/proxy/            # handlers for both listeners
  internal/metrics/          # Prometheus exposition
  internal/config/           # env + .env file configuration
  internal/update/           # signed auto-updater (public key in pubkey.go)
  internal/service/          # Windows SCM + Linux systemd (notify/watchdog) integration
  Dockerfile
  .gitea/workflows/ci.yml
  README.md
  SPEC.md                    # this file

main.go from the first prototype (ComfyUI-only) is the starting point for internal/comfy and the /prompt handler; the lock and the Ollama listener are new.

Testing

  • internal/lock: table tests plus a race test (go test -race) with goroutines: image waits for LLMs to drain; new LLMs block while image is pending; FIFO for images; context cancellation removes waiters.
  • internal/proxy: httptest.Server fakes for Ollama (/api/ps, /api/generate) and ComfyUI (/prompt, /history/:id, /free); assert the call sequence for one image job and that a concurrent /api/chat is held until /free was called.
  • Streaming test: fake Ollama emits chunks with delays; assert the client receives the first chunk before the last is sent (no buffering).
  • internal/config: env-file parsing, precedence, fail-fast values.
  • internal/update: fake Gitea releases API; staged update happy path, tampered signature rejected, older versions and dev builds skipped.

Build and CI

  • Go 1.23+, two external dependencies: golang.org/x/sys (Windows service integration) and github.com/coreos/go-systemd (systemd notify/watchdog, Linux build only). CGO_ENABLED=0, -ldflags="-s -w", version from git describe injected via -X main.version=.
  • Dockerfile: multi-stage, final image gcr.io/distroless/static (or scratch), non-root user, EXPOSE 8188 11434, ENTRYPOINT ["/gpu-turnstile"].
  • .gitea/workflows/ci.yml (Gitea Actions):
    1. on every push: go vet, go test -race ./..., golangci-lint if available in the runner image
    2. on a version tag only (vX.Y.Z, enforced): build the image with buildx and push it to the Gitea registry git.rambossek.at/<owner>/gpu-turnstile tagged :<tag> and :latest (the repository path is lowercased in the workflow; Docker registry names must be lowercase). Login uses the repo secret REGISTRY_TOKEN (an access token with write:package scope) because the automatic GITEA_TOKEN cannot push packages; the username is just gitea.actor.
    3. on a version tag: also build the Windows binary, sign it with OpenSSL (RELEASE_SIGNING_KEY secret), and attach gpu-turnstile.exe, .sig and .sha256 to a Gitea release for the auto-updater.
  • Release: a git tag vX.Y.Z produces the versioned image and the signed Windows binary; the Open WebUI compose pins that tag. No images or binaries are built from branches.

Deployment (target)

  gpu-turnstile:
    image: git.rambossek.at/<owner>/gpu-turnstile:v0.1.0  # owner lowercased, e.g. "public"
    environment:
      OLLAMA_URL: http://<workstation-ip>:11435
      COMFY_URL:  http://<workstation-ip>:8189
    networks: [internal]

LiteLLM api_basehttp://gpu-turnstile:11434; Open WebUI COMFYUI_BASE_URLhttp://gpu-turnstile:8188. Nothing else talks to the workstation directly.

Open questions

  • Should embedding requests (/api/embed, /v1/embeddings) count as LLM traffic for the lock? They do in this spec (they hold VRAM); reconsider if RAG indexing starves image jobs for too long.
  • Whether to add a POST /gpu-turnstile/release admin endpoint to force-reset the lock without restarting. Cheap to add; decide once it's been stuck once.