gpulock: GPU arbitration proxy for Ollama + ComfyUI
Implements SPEC.md: two listeners, one writer-preferring two-mode lock, Ollama unload before image jobs, ComfyUI history polling + VRAM free, optional model warm-up, healthz/metrics endpoints, streaming-safe reverse proxies, Dockerfile and Gitea Actions CI.
This commit is contained in:
@@ -0,0 +1,211 @@
|
||||
# gpulock — GPU arbitration proxy for Ollama + ComfyUI
|
||||
|
||||
## Problem
|
||||
|
||||
One consumer GPU (RTX 5080, 16 GB) is shared by an LLM server (Ollama) and an
|
||||
image generator (ComfyUI). Both assume they own the card. When both hold models
|
||||
at once, the NVIDIA Windows driver falls back to system memory and everything
|
||||
becomes very slow; on Linux it would OOM instead.
|
||||
|
||||
## Goal
|
||||
|
||||
A single Go binary that sits in front of **both** services and guarantees that
|
||||
at any moment the GPU is in exactly one of three states:
|
||||
|
||||
- `idle` — nothing in flight
|
||||
- `llm` — N ≥ 1 Ollama inference requests in flight (concurrency allowed)
|
||||
- `image` — exactly one ComfyUI job in flight, Ollama models unloaded
|
||||
|
||||
Clients (LiteLLM, Open WebUI, n8n) point at gpulock instead of at the services.
|
||||
gpulock is transparent for everything that does not touch the GPU.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- Not a scheduler across multiple GPUs or hosts. One lock, one card.
|
||||
- No auth, TLS, rate limiting. Runs on an internal network behind Traefik or a
|
||||
Docker bridge.
|
||||
- No request rewriting, caching, or protocol translation.
|
||||
- No persistence. Restart = idle state.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
LiteLLM / Open WebUI ──► :11435 ─┐ ┌─► Ollama :11434
|
||||
├── gpulock (1 lock) ──┤
|
||||
Open WebUI / n8n ────► :8189 ───┘ └─► ComfyUI :8188
|
||||
```
|
||||
|
||||
Two listeners, one process, one lock. Each listener is an
|
||||
`httputil.ReverseProxy` to its upstream. Websocket upgrades (ComfyUI `/ws`)
|
||||
and streaming bodies (Ollama NDJSON / SSE) must pass through unbuffered
|
||||
(`FlushInterval = -1`).
|
||||
|
||||
### Lock semantics
|
||||
|
||||
Two-mode lock with image priority (writer-preferring RW lock, where "readers"
|
||||
are LLM requests and the single "writer" is an image job):
|
||||
|
||||
- **LLM request** (see endpoint list): `AcquireLLM()` blocks while state is
|
||||
`image` **or while an image job is waiting**. Then state := `llm`, n++.
|
||||
On completion (response fully written, including streamed bodies, or client
|
||||
disconnect) n--; if n == 0 state := `idle`.
|
||||
- **Image job**: `AcquireImage()` marks "image pending" (so no new LLM
|
||||
requests start), waits until n == 0, sets state := `image`. Released after
|
||||
the ComfyUI job finished and models were freed.
|
||||
- Concurrent image jobs queue FIFO behind each other.
|
||||
- All waits are context-aware: a client that disconnects while waiting is
|
||||
removed from the queue.
|
||||
|
||||
### Endpoint classification
|
||||
|
||||
Ollama listener (`:11435` → `OLLAMA_URL`):
|
||||
|
||||
| Path | Handling |
|
||||
|---|---|
|
||||
| `POST /api/generate`, `/api/chat`, `/api/embed`, `/api/embeddings` | LLM lock |
|
||||
| `POST /v1/chat/completions`, `/v1/completions`, `/v1/embeddings` | LLM lock |
|
||||
| everything else (`/api/tags`, `/api/ps`, `/api/show`, `/api/version`, `/v1/models`, `/api/pull`, …) | pass-through, no lock |
|
||||
|
||||
ComfyUI listener (`:8189` → `COMFY_URL`):
|
||||
|
||||
| Path | Handling |
|
||||
|---|---|
|
||||
| `POST /prompt` | image lock (see flow below) |
|
||||
| everything else (`/ws`, `/history/*`, `/view`, `/system_stats`, `/queue`, `/free`, …) | pass-through, no lock |
|
||||
|
||||
### Image job flow (`POST /prompt`)
|
||||
|
||||
1. `AcquireImage()`.
|
||||
2. Unload Ollama: `GET /api/ps`; for each model `POST /api/generate
|
||||
{"model":M,"keep_alive":0}`; if that returns non-2xx (embedding-only
|
||||
models), `POST /api/embed {"model":M,"input":"x","keep_alive":0}`. Poll
|
||||
`/api/ps` every 500 ms until empty or `UNLOAD_TIMEOUT`. On timeout: log and
|
||||
continue (degrade, don't fail the user's request).
|
||||
3. Forward the original request body to ComfyUI `/prompt`, return status,
|
||||
headers and body to the caller unchanged, flush.
|
||||
4. If the response is 200 and contains `prompt_id`: in a goroutine, poll
|
||||
`GET /history/<prompt_id>` every 1 s until the entry has
|
||||
`status.completed == true`, `status.status_str == "error"`, or
|
||||
`JOB_TIMEOUT`. Then `POST /free {"unload_models":true,"free_memory":true}`.
|
||||
Then release the image lock.
|
||||
5. If the response is not 200 or has no `prompt_id`: release the lock
|
||||
immediately.
|
||||
|
||||
Optional (config flag `WARM_MODEL`): after releasing the image lock, if the
|
||||
state is `idle`, send `POST /api/generate {"model":WARM_MODEL,"keep_alive":-1}`
|
||||
with empty prompt to reload the chat model so the next chat doesn't pay the
|
||||
load time. Off by default.
|
||||
|
||||
## Configuration (env)
|
||||
|
||||
| Var | Default | Meaning |
|
||||
|---|---|---|
|
||||
| `LISTEN_OLLAMA` | `:11435` | Ollama-facing listener |
|
||||
| `LISTEN_COMFY` | `:8189` | ComfyUI-facing listener |
|
||||
| `OLLAMA_URL` | `http://127.0.0.1:11434` | upstream |
|
||||
| `COMFY_URL` | `http://127.0.0.1:8188` | upstream |
|
||||
| `UNLOAD_TIMEOUT` | `60s` | wait for Ollama to unload |
|
||||
| `JOB_TIMEOUT` | `15m` | wait for ComfyUI job |
|
||||
| `LLM_WAIT_TIMEOUT` | `10m` | max time an LLM request waits for the lock before 503 |
|
||||
| `WARM_MODEL` | `` | optional model to reload after an image job |
|
||||
| `LOG_LEVEL` | `info` | `debug` logs every lock transition |
|
||||
|
||||
Startup fails fast on unparsable values. Both upstreams are probed once at
|
||||
start (`/api/version`, `/system_stats`); failure is logged, not fatal.
|
||||
|
||||
## Observability
|
||||
|
||||
- `GET /healthz` on both listeners: 200 with JSON
|
||||
`{"state":"idle|llm|image","llm_inflight":N,"image_pending":B}`.
|
||||
- `GET /metrics` on the Ollama listener: Prometheus text format, no external
|
||||
dependency needed:
|
||||
`gpulock_state{state="…"} 1`, `gpulock_llm_inflight`,
|
||||
`gpulock_image_jobs_total`, `gpulock_lock_wait_seconds` (histogram, label
|
||||
`kind="llm|image"`), `gpulock_unload_seconds`.
|
||||
- Structured logs (`log/slog`, JSON when `LOG_FORMAT=json`), one line per
|
||||
state transition and per image job phase with `prompt_id`.
|
||||
|
||||
## Edge cases to handle
|
||||
|
||||
- Client disconnects while streaming an Ollama response: request context is
|
||||
cancelled, proxy aborts upstream, in-flight counter still decrements.
|
||||
- Client disconnects while waiting for the lock: removed from wait, no
|
||||
counter change.
|
||||
- ComfyUI job finishes but `/history` never shows it (e.g. ComfyUI restarted):
|
||||
`JOB_TIMEOUT` releases the lock; log at warn.
|
||||
- Ollama unreachable during unload: continue with the image job; the whole
|
||||
point is not to block users on a misbehaving neighbour.
|
||||
- `POST /prompt` with a body that ComfyUI rejects (400): lock released
|
||||
immediately, body passed back.
|
||||
- Websocket `/ws` connections are long-lived and never take the lock.
|
||||
- The Ollama OpenAI-compatible endpoints stream SSE; the proxy must not buffer.
|
||||
|
||||
## Repository layout
|
||||
|
||||
```
|
||||
gpulock/
|
||||
cmd/gpulock/main.go # wiring, config, listeners
|
||||
internal/lock/lock.go # two-mode lock + tests
|
||||
internal/ollama/client.go # ps / unload / warm
|
||||
internal/comfy/client.go # history poll / free
|
||||
internal/proxy/ # handlers for both listeners
|
||||
Dockerfile
|
||||
.gitea/workflows/ci.yml
|
||||
README.md
|
||||
SPEC.md # this file
|
||||
```
|
||||
|
||||
`main.go` from the first prototype (ComfyUI-only) is the starting point for
|
||||
`internal/comfy` and the `/prompt` handler; the lock and the Ollama listener
|
||||
are new.
|
||||
|
||||
## Testing
|
||||
|
||||
- `internal/lock`: table tests plus a race test (`go test -race`) with
|
||||
goroutines: image waits for LLMs to drain; new LLMs block while image is
|
||||
pending; FIFO for images; context cancellation removes waiters.
|
||||
- `internal/proxy`: `httptest.Server` fakes for Ollama (`/api/ps`,
|
||||
`/api/generate`) and ComfyUI (`/prompt`, `/history/:id`, `/free`); assert
|
||||
the call sequence for one image job and that a concurrent `/api/chat` is
|
||||
held until `/free` was called.
|
||||
- Streaming test: fake Ollama emits chunks with delays; assert the client
|
||||
receives the first chunk before the last is sent (no buffering).
|
||||
|
||||
## Build and CI
|
||||
|
||||
- Go 1.23+, stdlib only. `CGO_ENABLED=0`, `-ldflags="-s -w"`, version from
|
||||
`git describe` injected via `-X main.version=`.
|
||||
- Dockerfile: multi-stage, final image `gcr.io/distroless/static` (or
|
||||
`scratch`), non-root user, `EXPOSE 8189 11435`, `ENTRYPOINT ["/gpulock"]`.
|
||||
- `.gitea/workflows/ci.yml` (Gitea Actions):
|
||||
1. on push and tag: `go vet`, `go test -race ./...`, `golangci-lint` if
|
||||
available in the runner image
|
||||
2. build image with buildx, tags `:sha-<short>` and `:latest` on main,
|
||||
`:<tag>` on tags
|
||||
3. push to the Gitea registry `git.rambossek.at/<owner>/gpulock` using the
|
||||
workflow token (`${{ secrets.GITEA_TOKEN }}` / `gitea.actor`)
|
||||
- Release: a git tag `vX.Y.Z` produces the versioned image; the Open WebUI
|
||||
compose pins that tag.
|
||||
|
||||
## Deployment (target)
|
||||
|
||||
```yaml
|
||||
gpulock:
|
||||
image: git.rambossek.at/<owner>/gpulock:v0.1.0
|
||||
environment:
|
||||
OLLAMA_URL: http://<workstation-ip>:11434
|
||||
COMFY_URL: http://<workstation-ip>:8188
|
||||
networks: [internal]
|
||||
```
|
||||
|
||||
LiteLLM `api_base` → `http://gpulock:11435`; Open WebUI
|
||||
`COMFYUI_BASE_URL` → `http://gpulock:8189`. Nothing else talks to the
|
||||
workstation directly.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should embedding requests (`/api/embed`, `/v1/embeddings`) count as LLM
|
||||
traffic for the lock? They do in this spec (they hold VRAM); reconsider if
|
||||
RAG indexing starves image jobs for too long.
|
||||
- Whether to add a `POST /gpulock/release` admin endpoint to force-reset the
|
||||
lock without restarting. Cheap to add; decide once it's been stuck once.
|
||||
Reference in New Issue
Block a user