Files
gpu-turnstile/SPEC.md
T
mram 065c294a96 gpulock: GPU arbitration proxy for Ollama + ComfyUI
Implements SPEC.md: two listeners, one writer-preferring two-mode lock,
Ollama unload before image jobs, ComfyUI history polling + VRAM free,
optional model warm-up, healthz/metrics endpoints, streaming-safe
reverse proxies, Dockerfile and Gitea Actions CI.
2026-09-20 18:05:20 +02:00

8.8 KiB

gpulock — GPU arbitration proxy for Ollama + ComfyUI

Problem

One consumer GPU (RTX 5080, 16 GB) is shared by an LLM server (Ollama) and an image generator (ComfyUI). Both assume they own the card. When both hold models at once, the NVIDIA Windows driver falls back to system memory and everything becomes very slow; on Linux it would OOM instead.

Goal

A single Go binary that sits in front of both services and guarantees that at any moment the GPU is in exactly one of three states:

  • idle — nothing in flight
  • llm — N ≥ 1 Ollama inference requests in flight (concurrency allowed)
  • image — exactly one ComfyUI job in flight, Ollama models unloaded

Clients (LiteLLM, Open WebUI, n8n) point at gpulock instead of at the services. gpulock is transparent for everything that does not touch the GPU.

Non-goals

  • Not a scheduler across multiple GPUs or hosts. One lock, one card.
  • No auth, TLS, rate limiting. Runs on an internal network behind Traefik or a Docker bridge.
  • No request rewriting, caching, or protocol translation.
  • No persistence. Restart = idle state.

Architecture

LiteLLM / Open WebUI ──► :11435 ─┐                     ┌─► Ollama  :11434
                                 ├── gpulock (1 lock) ──┤
Open WebUI / n8n ────► :8189  ───┘                     └─► ComfyUI :8188

Two listeners, one process, one lock. Each listener is an httputil.ReverseProxy to its upstream. Websocket upgrades (ComfyUI /ws) and streaming bodies (Ollama NDJSON / SSE) must pass through unbuffered (FlushInterval = -1).

Lock semantics

Two-mode lock with image priority (writer-preferring RW lock, where "readers" are LLM requests and the single "writer" is an image job):

  • LLM request (see endpoint list): AcquireLLM() blocks while state is image or while an image job is waiting. Then state := llm, n++. On completion (response fully written, including streamed bodies, or client disconnect) n--; if n == 0 state := idle.
  • Image job: AcquireImage() marks "image pending" (so no new LLM requests start), waits until n == 0, sets state := image. Released after the ComfyUI job finished and models were freed.
  • Concurrent image jobs queue FIFO behind each other.
  • All waits are context-aware: a client that disconnects while waiting is removed from the queue.

Endpoint classification

Ollama listener (:11435OLLAMA_URL):

Path Handling
POST /api/generate, /api/chat, /api/embed, /api/embeddings LLM lock
POST /v1/chat/completions, /v1/completions, /v1/embeddings LLM lock
everything else (/api/tags, /api/ps, /api/show, /api/version, /v1/models, /api/pull, …) pass-through, no lock

ComfyUI listener (:8189COMFY_URL):

Path Handling
POST /prompt image lock (see flow below)
everything else (/ws, /history/*, /view, /system_stats, /queue, /free, …) pass-through, no lock

Image job flow (POST /prompt)

  1. AcquireImage().
  2. Unload Ollama: GET /api/ps; for each model POST /api/generate {"model":M,"keep_alive":0}; if that returns non-2xx (embedding-only models), POST /api/embed {"model":M,"input":"x","keep_alive":0}. Poll /api/ps every 500 ms until empty or UNLOAD_TIMEOUT. On timeout: log and continue (degrade, don't fail the user's request).
  3. Forward the original request body to ComfyUI /prompt, return status, headers and body to the caller unchanged, flush.
  4. If the response is 200 and contains prompt_id: in a goroutine, poll GET /history/<prompt_id> every 1 s until the entry has status.completed == true, status.status_str == "error", or JOB_TIMEOUT. Then POST /free {"unload_models":true,"free_memory":true}. Then release the image lock.
  5. If the response is not 200 or has no prompt_id: release the lock immediately.

Optional (config flag WARM_MODEL): after releasing the image lock, if the state is idle, send POST /api/generate {"model":WARM_MODEL,"keep_alive":-1} with empty prompt to reload the chat model so the next chat doesn't pay the load time. Off by default.

Configuration (env)

Var Default Meaning
LISTEN_OLLAMA :11435 Ollama-facing listener
LISTEN_COMFY :8189 ComfyUI-facing listener
OLLAMA_URL http://127.0.0.1:11434 upstream
COMFY_URL http://127.0.0.1:8188 upstream
UNLOAD_TIMEOUT 60s wait for Ollama to unload
JOB_TIMEOUT 15m wait for ComfyUI job
LLM_WAIT_TIMEOUT 10m max time an LLM request waits for the lock before 503
WARM_MODEL `` optional model to reload after an image job
LOG_LEVEL info debug logs every lock transition

Startup fails fast on unparsable values. Both upstreams are probed once at start (/api/version, /system_stats); failure is logged, not fatal.

Observability

  • GET /healthz on both listeners: 200 with JSON {"state":"idle|llm|image","llm_inflight":N,"image_pending":B}.
  • GET /metrics on the Ollama listener: Prometheus text format, no external dependency needed: gpulock_state{state="…"} 1, gpulock_llm_inflight, gpulock_image_jobs_total, gpulock_lock_wait_seconds (histogram, label kind="llm|image"), gpulock_unload_seconds.
  • Structured logs (log/slog, JSON when LOG_FORMAT=json), one line per state transition and per image job phase with prompt_id.

Edge cases to handle

  • Client disconnects while streaming an Ollama response: request context is cancelled, proxy aborts upstream, in-flight counter still decrements.
  • Client disconnects while waiting for the lock: removed from wait, no counter change.
  • ComfyUI job finishes but /history never shows it (e.g. ComfyUI restarted): JOB_TIMEOUT releases the lock; log at warn.
  • Ollama unreachable during unload: continue with the image job; the whole point is not to block users on a misbehaving neighbour.
  • POST /prompt with a body that ComfyUI rejects (400): lock released immediately, body passed back.
  • Websocket /ws connections are long-lived and never take the lock.
  • The Ollama OpenAI-compatible endpoints stream SSE; the proxy must not buffer.

Repository layout

gpulock/
  cmd/gpulock/main.go        # wiring, config, listeners
  internal/lock/lock.go      # two-mode lock + tests
  internal/ollama/client.go  # ps / unload / warm
  internal/comfy/client.go   # history poll / free
  internal/proxy/            # handlers for both listeners
  Dockerfile
  .gitea/workflows/ci.yml
  README.md
  SPEC.md                    # this file

main.go from the first prototype (ComfyUI-only) is the starting point for internal/comfy and the /prompt handler; the lock and the Ollama listener are new.

Testing

  • internal/lock: table tests plus a race test (go test -race) with goroutines: image waits for LLMs to drain; new LLMs block while image is pending; FIFO for images; context cancellation removes waiters.
  • internal/proxy: httptest.Server fakes for Ollama (/api/ps, /api/generate) and ComfyUI (/prompt, /history/:id, /free); assert the call sequence for one image job and that a concurrent /api/chat is held until /free was called.
  • Streaming test: fake Ollama emits chunks with delays; assert the client receives the first chunk before the last is sent (no buffering).

Build and CI

  • Go 1.23+, stdlib only. CGO_ENABLED=0, -ldflags="-s -w", version from git describe injected via -X main.version=.
  • Dockerfile: multi-stage, final image gcr.io/distroless/static (or scratch), non-root user, EXPOSE 8189 11435, ENTRYPOINT ["/gpulock"].
  • .gitea/workflows/ci.yml (Gitea Actions):
    1. on push and tag: go vet, go test -race ./..., golangci-lint if available in the runner image
    2. build image with buildx, tags :sha-<short> and :latest on main, :<tag> on tags
    3. push to the Gitea registry git.rambossek.at/<owner>/gpulock using the workflow token (${{ secrets.GITEA_TOKEN }} / gitea.actor)
  • Release: a git tag vX.Y.Z produces the versioned image; the Open WebUI compose pins that tag.

Deployment (target)

  gpulock:
    image: git.rambossek.at/<owner>/gpulock:v0.1.0
    environment:
      OLLAMA_URL: http://<workstation-ip>:11434
      COMFY_URL:  http://<workstation-ip>:8188
    networks: [internal]

LiteLLM api_basehttp://gpulock:11435; Open WebUI COMFYUI_BASE_URLhttp://gpulock:8189. Nothing else talks to the workstation directly.

Open questions

  • Should embedding requests (/api/embed, /v1/embeddings) count as LLM traffic for the lock? They do in this spec (they hold VRAM); reconsider if RAG indexing starves image jobs for too long.
  • Whether to add a POST /gpulock/release admin endpoint to force-reset the lock without restarting. Cheap to add; decide once it's been stuck once.