Add GPU_FOREIGN_UTIL_PCT: game detection via per-process GPU 3D-engine usage

Reads the same PDH counters as Task Manager (\GPU Engine(*)\Utilization
Percentage, locale-independent via PdhAddEnglishCounterW), which cover
graphics work under WDDM — games are caught without an exe list and the
offender is named. Windows-only; a missing or failing counter disables
the path with one log line. dwm (the desktop compositor) joins the
default ignore list.
This commit is contained in:
mram
2026-09-22 10:53:29 +02:00
parent 8fccd333aa
commit d259b2e96c
11 changed files with 387 additions and 72 deletions
+17 -11
View File
@@ -34,8 +34,8 @@ Each consumer is enabled by setting its URL (`OLLAMA_URL`, `COMFY_URL`) and
disabled by leaving it empty — at least one is required. With only Ollama disabled by leaving it empty — at least one is required. With only Ollama
the proxy is a pass-through (no image jobs can arrive); with only ComfyUI the proxy is a pass-through (no image jobs can arrive); with only ComfyUI
the Ollama unload/warm steps are skipped. A third, URL-less consumer — the Ollama unload/warm steps are skipped. A third, URL-less consumer —
detection of foreign GPU holders such as games — is enabled by `GAME_PROCS` detection of foreign GPU holders such as games — is enabled by `GAME_PROCS`,
and/or `GPU_FOREIGN_VRAM_MB` (see below). `GPU_FOREIGN_VRAM_MB` and/or `GPU_FOREIGN_UTIL_PCT` (see below).
## Configuration ## Configuration
@@ -64,7 +64,8 @@ override file values. Invalid values fail at startup.
| `COMFY_START_TIMEOUT` | `2m` | Max wait for the managed ComfyUI to come up | | `COMFY_START_TIMEOUT` | `2m` | Max wait for the managed ComfyUI to come up |
| `GAME_PROCS` | _(empty = disabled)_ | Process names (comma-separated); while any runs, the GPU counts as held: requests wait, Ollama unloads, managed ComfyUI stops | | `GAME_PROCS` | _(empty = disabled)_ | Process names (comma-separated); while any runs, the GPU counts as held: requests wait, Ollama unloads, managed ComfyUI stops |
| `GPU_FOREIGN_VRAM_MB` | `0` (disabled) | Also treat the GPU as held when a non-ignored process uses more VRAM than this (needs nvidia-smi) | | `GPU_FOREIGN_VRAM_MB` | `0` (disabled) | Also treat the GPU as held when a non-ignored process uses more VRAM than this (needs nvidia-smi) |
| `GPU_IGNORE_PROCS` | `ollama,ollama app,ollama_llama_server,python,pythonw` | Process names never counted as foreign GPU users | | `GPU_FOREIGN_UTIL_PCT` | `0` (disabled) | Also treat the GPU as held when a non-ignored process uses more than this percent of the GPU 3D engine (Windows PDH counters — what Task Manager shows; catches games without an exe list) |
| `GPU_IGNORE_PROCS` | `ollama,ollama app,ollama_llama_server,python,pythonw,dwm` | Process names never counted as foreign GPU users |
| `GAME_POLL_INTERVAL` | `15s` | How often game/VRAM detection runs (don't go below ~10s — nvidia-smi polls keep the GPU awake) | | `GAME_POLL_INTERVAL` | `15s` | How often game/VRAM detection runs (don't go below ~10s — nvidia-smi polls keep the GPU awake) |
| `LOGLEVEL` | `warn` | `info` logs every request (colored arrows in text mode), `debug` adds lock transitions. `LOG_LEVEL` works as an alias | | `LOGLEVEL` | `warn` | `info` logs every request (colored arrows in text mode), `debug` adds lock transitions. `LOG_LEVEL` works as an alias |
| `LOG_FORMAT` | `text` | `json` for structured JSON logs | | `LOG_FORMAT` | `text` | `json` for structured JSON logs |
@@ -148,20 +149,25 @@ adds a `BindPaths=` to the systemd unit on Linux. Re-run it after changing
Want to game on the same GPU without Ollama/ComfyUI squatting on the VRAM? Want to game on the same GPU without Ollama/ComfyUI squatting on the VRAM?
gpu-turnstile can watch for foreign GPU holders and, while one is active, gpu-turnstile can watch for foreign GPU holders and, while one is active,
make LLM/image requests wait (or 503, per `LLM_BUSY_MODE`), unload Ollama's make LLM/image requests wait (or 503, per `LLM_BUSY_MODE`), unload Ollama's
models and stop the managed ComfyUI so the game gets the memory. Two models and stop the managed ComfyUI so the game gets the memory. Three
detection paths, each optional, polled every `GAME_POLL_INTERVAL` (15s): detection paths, each optional, polled every `GAME_POLL_INTERVAL` (15s):
``` ```
GAME_PROCS=cyberpunk2077.exe,bg3.exe # the reliable way on Windows GPU_FOREIGN_UTIL_PCT=30 # zero-config on Windows: any non-ignored
# process using >30% of the GPU 3D engine
GAME_PROCS=cyberpunk2077.exe,bg3.exe # explicit exe watch list
GPU_FOREIGN_VRAM_MB=1024 # catch-all via nvidia-smi GPU_FOREIGN_VRAM_MB=1024 # catch-all via nvidia-smi
``` ```
`GAME_PROCS` matches running process names (case-insensitive, `.exe` `GPU_FOREIGN_UTIL_PCT` reads the same per-process GPU engine counters as
optional). `GPU_FOREIGN_VRAM_MB` asks nvidia-smi which processes hold GPU Task Manager (PDH), which cover graphics work under Windows' WDDM driver —
memory and treats anything not in `GPU_IGNORE_PROCS` above the threshold as so games show up without naming them, and the offender is named in the log
foreign — handy as a catch-all, but note that under Windows' WDDM driver and monitor. Windows-only. `GAME_PROCS` matches running process names
graphics-only games may not show up in nvidia-smi's per-process list, so (case-insensitive, `.exe` optional). `GPU_FOREIGN_VRAM_MB` asks nvidia-smi
name your games in `GAME_PROCS` there; on Linux both paths work. When the which processes hold GPU memory and treats anything not in
`GPU_IGNORE_PROCS` above the threshold as foreign — handy as a catch-all on
Linux, but under WDDM graphics-only games may not show up in nvidia-smi's
per-process list, so on Windows prefer `GPU_FOREIGN_UTIL_PCT`. When the
game exits, requests resume automatically. game exits, requests resume automatically.
## Build and run ## Build and run
+19 -12
View File
@@ -55,9 +55,9 @@ consumer gets no listener, no startup probe, and no lock participation:
- **Only `COMFY_URL`**: image jobs are tracked and ComfyUI's VRAM is freed - **Only `COMFY_URL`**: image jobs are tracked and ComfyUI's VRAM is freed
afterwards, but the Ollama unload and warm-reload steps are skipped. afterwards, but the Ollama unload and warm-reload steps are skipped.
- **Game detection** is a third, optional consumer without a URL: enabled by - **Game detection** is a third, optional consumer without a URL: enabled by
`GAME_PROCS` and/or `GPU_FOREIGN_VRAM_MB` it watches for foreign processes `GAME_PROCS`, `GPU_FOREIGN_VRAM_MB` and/or `GPU_FOREIGN_UTIL_PCT` it
holding the GPU (see below) and plugs into the same lock the same way — watches for foreign processes holding the GPU (see below) and plugs into
excluded when both knobs are unset. the same lock the same way — excluded when all knobs are unset.
### Lock semantics ### Lock semantics
@@ -171,20 +171,26 @@ files are flagged in the startup log.
## Game detection (foreign GPU holders) ## Game detection (foreign GPU holders)
Games and other foreign GPU users sit outside the URL-based consumer model — Games and other foreign GPU users sit outside the URL-based consumer model —
nothing proxies through gpu-turnstile for them. Two independent detection nothing proxies through gpu-turnstile for them. Three independent detection
paths, polled every `GAME_POLL_INTERVAL` (default 15 s); either one being paths, polled every `GAME_POLL_INTERVAL` (default 15 s); any one being
configured enables the feature: configured enables the feature:
- **Process watch list** (`GAME_PROCS`, comma-separated, case-insensitive, - **Process watch list** (`GAME_PROCS`, comma-separated, case-insensitive,
`.exe` optional): while any listed process runs, the GPU counts as held. `.exe` optional): while any listed process runs, the GPU counts as held.
This is the reliable path on Windows. - **Foreign 3D-engine utilization** (`GPU_FOREIGN_UTIL_PCT`, Windows only):
per-process GPU engine counters from PDH — the same data Task Manager
shows — cover graphics work under WDDM, so any process not in
`GPU_IGNORE_PROCS` using more than the threshold percent of the 3D engine
counts as a foreign holder, with no exe list needed. Values are summed per
process across engines; a missing/failing PDH counter disables the path
(logged once).
- **Foreign VRAM threshold** (`GPU_FOREIGN_VRAM_MB`): `nvidia-smi - **Foreign VRAM threshold** (`GPU_FOREIGN_VRAM_MB`): `nvidia-smi
--query-compute-apps` lists per-process GPU memory; any process not in --query-compute-apps` lists per-process GPU memory; any process not in
`GPU_IGNORE_PROCS` (default: Ollama and python — ComfyUI runs under python) `GPU_IGNORE_PROCS` (default: Ollama and python — ComfyUI runs under python
holding more than the threshold counts as a foreign holder. Needs — plus dwm, the desktop compositor) holding more than the threshold counts
nvidia-smi on the PATH (absent: logged once, path disabled) and works best as a foreign holder. Needs nvidia-smi on the PATH (absent: logged once,
on Linux — under Windows' WDDM driver, graphics-only games may not appear path disabled) and works best on Linux — under Windows' WDDM driver,
in the per-process list. graphics-only games may not appear in the per-process list.
While a holder is detected, gpu-turnstile: While a holder is detected, gpu-turnstile:
@@ -230,7 +236,8 @@ override file values. A missing file is fine; a malformed one is fatal.
| `COMFY_START_TIMEOUT` | `2m` | how long a request waits for the managed ComfyUI to come up | | `COMFY_START_TIMEOUT` | `2m` | how long a request waits for the managed ComfyUI to come up |
| `GAME_PROCS` | _(empty = disabled)_ | comma-separated process names (case-insensitive, `.exe` optional); while any runs, the GPU counts as held by it: requests wait, Ollama unloads, the managed ComfyUI stops | | `GAME_PROCS` | _(empty = disabled)_ | comma-separated process names (case-insensitive, `.exe` optional); while any runs, the GPU counts as held by it: requests wait, Ollama unloads, the managed ComfyUI stops |
| `GPU_FOREIGN_VRAM_MB` | `0` (disabled) | also treat the GPU as held when a process not in `GPU_IGNORE_PROCS` uses more VRAM than this; needs nvidia-smi | | `GPU_FOREIGN_VRAM_MB` | `0` (disabled) | also treat the GPU as held when a process not in `GPU_IGNORE_PROCS` uses more VRAM than this; needs nvidia-smi |
| `GPU_IGNORE_PROCS` | `ollama,ollama app,ollama_llama_server,python,pythonw` | process names never counted as foreign GPU users | | `GPU_FOREIGN_UTIL_PCT` | `0` (disabled) | also treat the GPU as held when a process not in `GPU_IGNORE_PROCS` uses more than this percent of the GPU 3D engine (Windows PDH counters, as shown by Task Manager; catches games without an exe list) |
| `GPU_IGNORE_PROCS` | `ollama,ollama app,ollama_llama_server,python,pythonw,dwm` | process names never counted as foreign GPU users |
| `GAME_POLL_INTERVAL` | `15s` | how often game/VRAM detection runs (nvidia-smi polls keep the GPU awake; don't go below ~10s) | | `GAME_POLL_INTERVAL` | `15s` | how often game/VRAM detection runs (nvidia-smi polls keep the GPU awake; don't go below ~10s) |
| `LOGLEVEL` | `warn` | `info` logs every request (colored arrows in text mode), `debug` adds lock transitions. `LOG_LEVEL` is accepted as an alias | | `LOGLEVEL` | `warn` | `info` logs every request (colored arrows in text mode), `debug` adds lock transitions. `LOG_LEVEL` is accepted as an alias |
| `LOG_FORMAT` | `text` | `json` for structured JSON logs | | `LOG_FORMAT` | `text` | `json` for structured JSON logs |
+7 -6
View File
@@ -592,6 +592,7 @@ func run(ctx context.Context, cfg config.Config, log *slog.Logger, logOut io.Wri
"comfy_start_timeout", cfg.ComfyStartTimeout, "comfy_start_timeout", cfg.ComfyStartTimeout,
"game_procs", cfg.GameProcs, "game_procs", cfg.GameProcs,
"gpu_foreign_vram_mb", cfg.GPUForeignVRAMMB, "gpu_foreign_vram_mb", cfg.GPUForeignVRAMMB,
"gpu_foreign_util_pct", cfg.GPUForeignUtilPct,
"gpu_ignore_procs", cfg.GPUIgnoreProcs, "gpu_ignore_procs", cfg.GPUIgnoreProcs,
"game_poll_interval", cfg.GamePollInterval, "game_poll_interval", cfg.GamePollInterval,
"auto_update", cfg.AutoUpdate, "auto_update", cfg.AutoUpdate,
@@ -726,13 +727,13 @@ func run(ctx context.Context, cfg config.Config, log *slog.Logger, logOut io.Wri
go healthLoop(ctx, cfg.HealthInterval, cfg.ProbeTimeout, log, probes, health) go healthLoop(ctx, cfg.HealthInterval, cfg.ProbeTimeout, log, probes, health)
} }
// Foreign GPU holders (games, other ML jobs) — enabled by GAME_PROCS // Foreign GPU holders (games, other ML jobs) — enabled by GAME_PROCS,
// and/or GPU_FOREIGN_VRAM_MB — hold the lock externally while they run. // GPU_FOREIGN_VRAM_MB and/or GPU_FOREIGN_UTIL_PCT — hold the lock
// gw collects the VRAM reading and the last check result for the status // externally while they run. gw collects the VRAM reading and the last
// channel. // check result for the status channel.
gw := &gpuWatch{enabled: len(cfg.GameProcs) > 0 || cfg.GPUForeignVRAMMB > 0} gw := &gpuWatch{enabled: len(cfg.GameProcs) > 0 || cfg.GPUForeignVRAMMB > 0 || cfg.GPUForeignUtilPct > 0}
if gw.enabled { if gw.enabled {
det := game.New(cfg.GameProcs, cfg.GPUForeignVRAMMB, cfg.GPUIgnoreProcs, log) det := game.New(cfg.GameProcs, cfg.GPUForeignVRAMMB, cfg.GPUForeignUtilPct, cfg.GPUIgnoreProcs, log)
go gameLoop(ctx, cfg, log, det, lk, ollamaClient, comfySup, gw) go gameLoop(ctx, cfg, log, det, lk, ollamaClient, comfySup, gw)
} }
+19 -8
View File
@@ -70,12 +70,15 @@ type Config struct {
// them runs, the GPU is treated as held by a foreign process. The // them runs, the GPU is treated as held by a foreign process. The
// nvidia-smi path (GPUForeignVRAMMB, GPU_FOREIGN_VRAM_MB) does the same // nvidia-smi path (GPUForeignVRAMMB, GPU_FOREIGN_VRAM_MB) does the same
// when a process not in GPUIgnoreProcs (GPU_IGNORE_PROCS) holds more than // when a process not in GPUIgnoreProcs (GPU_IGNORE_PROCS) holds more than
// that many MiB of VRAM. GamePollInterval (GAME_POLL_INTERVAL) is how // that many MiB of VRAM, and the PDH path (GPUForeignUtilPct,
// often both checks run. // GPU_FOREIGN_UTIL_PCT) when such a process uses more than that many
GameProcs []string `env:"GAME_PROCS"` // percent of the GPU 3D engine (Windows only). GamePollInterval
GPUForeignVRAMMB int `env:"GPU_FOREIGN_VRAM_MB"` // (GAME_POLL_INTERVAL) is how often all checks run.
GPUIgnoreProcs []string `env:"GPU_IGNORE_PROCS"` GameProcs []string `env:"GAME_PROCS"`
GamePollInterval time.Duration `env:"GAME_POLL_INTERVAL"` GPUForeignVRAMMB int `env:"GPU_FOREIGN_VRAM_MB"`
GPUForeignUtilPct int `env:"GPU_FOREIGN_UTIL_PCT"`
GPUIgnoreProcs []string `env:"GPU_IGNORE_PROCS"`
GamePollInterval time.Duration `env:"GAME_POLL_INTERVAL"`
WarmModel string `env:"WARM_MODEL"` WarmModel string `env:"WARM_MODEL"`
LogLevel slog.Level `env:"LOGLEVEL"` LogLevel slog.Level `env:"LOGLEVEL"`
@@ -119,8 +122,9 @@ func Defaults() Config {
ComfyStartTimeout: 2 * time.Minute, ComfyStartTimeout: 2 * time.Minute,
// ComfyUI runs under python; excluding it (and Ollama) by name keeps // ComfyUI runs under python; excluding it (and Ollama) by name keeps
// our own consumers from tripping the foreign-VRAM check. // our own consumers from tripping the foreign-VRAM check. dwm is the
GPUIgnoreProcs: []string{"ollama", "ollama app", "ollama_llama_server", "python", "pythonw"}, // desktop compositor — it always shows some 3D-engine usage.
GPUIgnoreProcs: []string{"ollama", "ollama app", "ollama_llama_server", "python", "pythonw", "dwm"},
GamePollInterval: 15 * time.Second, GamePollInterval: 15 * time.Second,
LogLevel: slog.LevelWarn, LogLevel: slog.LevelWarn,
@@ -243,6 +247,13 @@ func Load(getenv func(string) string) (Config, error) {
} }
cfg.GPUForeignVRAMMB = n cfg.GPUForeignVRAMMB = n
} }
if v := getenv("GPU_FOREIGN_UTIL_PCT"); v != "" {
n, err := strconv.Atoi(v)
if err != nil || n < 0 || n > 100 {
return cfg, fmt.Errorf("GPU_FOREIGN_UTIL_PCT: must be an integer in 0-100 (percent of the GPU 3D engine, 0 = disabled)")
}
cfg.GPUForeignUtilPct = n
}
if v := getenv("PROMPT_CAPTURE_LIMIT"); v != "" { if v := getenv("PROMPT_CAPTURE_LIMIT"); v != "" {
n, err := strconv.ParseInt(v, 10, 64) n, err := strconv.ParseInt(v, 10, 64)
if err != nil || n < 0 { if err != nil || n < 0 {
+2 -1
View File
@@ -34,7 +34,8 @@ func sampleEntries(logFile string) []sampleEntry {
{"COMFY_START_TIMEOUT", "2m", "How long a request waits for the managed ComfyUI to come up", false}, {"COMFY_START_TIMEOUT", "2m", "How long a request waits for the managed ComfyUI to come up", false},
{"GAME_PROCS", "cyberpunk2077.exe,hl2.exe", "While a listed process runs, the GPU counts as held by it: requests wait, Ollama unloads, managed ComfyUI stops (default: empty = disabled)", false}, {"GAME_PROCS", "cyberpunk2077.exe,hl2.exe", "While a listed process runs, the GPU counts as held by it: requests wait, Ollama unloads, managed ComfyUI stops (default: empty = disabled)", false},
{"GPU_FOREIGN_VRAM_MB", "1024", "Also treat the GPU as held when a process not in GPU_IGNORE_PROCS uses more VRAM than this (needs nvidia-smi; 0/empty = disabled)", false}, {"GPU_FOREIGN_VRAM_MB", "1024", "Also treat the GPU as held when a process not in GPU_IGNORE_PROCS uses more VRAM than this (needs nvidia-smi; 0/empty = disabled)", false},
{"GPU_IGNORE_PROCS", "ollama,ollama app,ollama_llama_server,python,pythonw", "Process names never counted as foreign GPU users (ComfyUI runs under python)", false}, {"GPU_FOREIGN_UTIL_PCT", "30", "Also treat the GPU as held when a process not in GPU_IGNORE_PROCS uses more than this percent of the GPU 3D engine (Windows Task-Manager counters; catches games without an exe list; 0/empty = disabled)", false},
{"GPU_IGNORE_PROCS", "ollama,ollama app,ollama_llama_server,python,pythonw,dwm", "Process names never counted as foreign GPU users (ComfyUI runs under python; dwm is the desktop compositor)", false},
{"GAME_POLL_INTERVAL", "15s", "How often game/VRAM detection runs (nvidia-smi polls keep the GPU awake; don't go below ~10s)", false}, {"GAME_POLL_INTERVAL", "15s", "How often game/VRAM detection runs (nvidia-smi polls keep the GPU awake; don't go below ~10s)", false},
{"UNLOAD_TIMEOUT", "60s", "How long to wait for Ollama to unload a model", false}, {"UNLOAD_TIMEOUT", "60s", "How long to wait for Ollama to unload a model", false},
{"JOB_TIMEOUT", "15m", "Maximum time to wait for a ComfyUI job", false}, {"JOB_TIMEOUT", "15m", "Maximum time to wait for a ComfyUI job", false},
+84 -30
View File
@@ -1,8 +1,10 @@
// Package game detects processes outside gpu-turnstile's control that hold // Package game detects processes outside gpu-turnstile's control that hold
// the GPU — typically a game — so the proxy can block new GPU work and free // the GPU — typically a game — so the proxy can block new GPU work and free
// VRAM while they run. Two detection paths: an explicit process watch list // VRAM while they run. Three detection paths: an explicit process watch
// (GAME_PROCS) and a foreign-VRAM threshold via nvidia-smi // list (GAME_PROCS), a foreign-VRAM threshold via nvidia-smi
// (GPU_FOREIGN_VRAM_MB) that catches anything not on the ignore list. // (GPU_FOREIGN_VRAM_MB), and a per-process GPU 3D-engine utilization
// threshold via Windows PDH counters (GPU_FOREIGN_UTIL_PCT). The latter two
// catch anything not on the ignore list without naming individual games.
package game package game
import ( import (
@@ -11,6 +13,7 @@ import (
"fmt" "fmt"
"log/slog" "log/slog"
"os/exec" "os/exec"
"slices"
"strconv" "strconv"
"strings" "strings"
) )
@@ -28,28 +31,33 @@ type computeApp struct {
} }
// Detector checks whether a foreign process holds the GPU. The zero value // Detector checks whether a foreign process holds the GPU. The zero value
// (no watch list, no threshold) never detects anything; main only starts the // (no watch list, no thresholds) never detects anything; main only starts
// poll loop when at least one path is configured. // the poll loop when at least one path is configured.
type Detector struct { type Detector struct {
procs map[string]bool // normalized names from GAME_PROCS procs map[string]bool // normalized names from GAME_PROCS
vramMB int // foreign VRAM threshold; 0 = disabled vramMB int // foreign VRAM threshold; 0 = disabled
utilPct int // foreign 3D-engine utilization threshold; 0 = disabled
ignore map[string]bool // normalized names never counted as foreign ignore map[string]bool // normalized names never counted as foreign
log *slog.Logger log *slog.Logger
noNvidia bool // nvidia-smi was not found; VRAM path disabled for good noNvidia bool // nvidia-smi was not found; VRAM path disabled for good
sampler *gpuEngineSampler // open PDH query, opened lazily on first Check
noPDH bool // engine counters unavailable; util path disabled for good
} }
// New builds a Detector from the configured watch list, VRAM threshold in // New builds a Detector from the configured watch list, VRAM threshold in
// MiB (0 disables the nvidia-smi path) and ignore list. Names are matched // MiB, 3D-engine utilization threshold in percent (both 0 = disabled) and
// case-insensitively, with or without a trailing ".exe". // ignore list. Names are matched case-insensitively, with or without a
func New(procs []string, vramMB int, ignore []string, log *slog.Logger) *Detector { // trailing ".exe".
func New(procs []string, vramMB, utilPct int, ignore []string, log *slog.Logger) *Detector {
if log == nil { if log == nil {
log = slog.Default() log = slog.Default()
} }
return &Detector{ return &Detector{
procs: nameSet(procs), procs: nameSet(procs),
vramMB: vramMB, vramMB: vramMB,
ignore: nameSet(ignore), utilPct: utilPct,
log: log, ignore: nameSet(ignore),
log: log,
} }
} }
@@ -73,34 +81,63 @@ func nameSet(names []string) map[string]bool {
// description of each (empty when the GPU is free for gpu-turnstile's // description of each (empty when the GPU is free for gpu-turnstile's
// consumers). A failing nvidia-smi call is returned as an error only when // consumers). A failing nvidia-smi call is returned as an error only when
// the process list found nothing; a missing nvidia-smi binary disables the // the process list found nothing; a missing nvidia-smi binary disables the
// VRAM path permanently (logged once). // VRAM path permanently (logged once), as do missing engine counters.
func (d *Detector) Check(ctx context.Context) ([]string, error) { func (d *Detector) Check(ctx context.Context) ([]string, error) {
ps, psErr := processes() ps, psErr := processes()
if d.vramMB <= 0 || d.noNvidia { utils := d.engineUtil()
return d.detect(ps, nil), psErr var apps []computeApp
if d.vramMB > 0 && !d.noNvidia {
var err error
apps, err = queryComputeApps(ctx)
if errors.Is(err, exec.ErrNotFound) {
d.noNvidia = true
d.log.Warn("GPU_FOREIGN_VRAM_MB is set but nvidia-smi was not found; VRAM detection disabled")
} else if err != nil {
return d.detect(ps, nil, utils), err
}
} }
apps, err := queryComputeApps(ctx) return d.detect(ps, apps, utils), psErr
if errors.Is(err, exec.ErrNotFound) {
d.noNvidia = true
d.log.Warn("GPU_FOREIGN_VRAM_MB is set but nvidia-smi was not found; VRAM detection disabled")
return d.detect(ps, nil), nil
}
if err != nil {
return d.detect(ps, nil), err
}
return d.detect(ps, apps), nil
} }
// detect is the pure core of Check: given the process table and (optionally) // engineUtil samples per-process 3D-engine utilization via PDH. The first
// the nvidia-smi compute-apps list, it returns the foreign holders. // call only primes the rate counters and returns nil. A failing open or
func (d *Detector) detect(ps []Process, apps []computeApp) []string { // sample disables the path permanently (logged once).
func (d *Detector) engineUtil() map[int]float64 {
if d.utilPct <= 0 || d.noPDH {
return nil
}
if d.sampler == nil {
s, err := openGPUEngineSampler()
if err != nil {
d.noPDH = true
d.log.Warn("GPU_FOREIGN_UTIL_PCT is set but per-process GPU counters are unavailable; engine detection disabled", "err", err)
return nil
}
d.sampler = s
}
utils, err := d.sampler.sample()
if err != nil {
if errors.Is(err, errNotPrimed) {
return nil
}
d.noPDH = true
d.log.Warn("per-process GPU counters failed; engine detection disabled", "err", err)
return nil
}
return utils
}
// detect is the pure core of Check: given the process table and
// (optionally) the nvidia-smi compute-apps list and the PDH engine
// utilization, it returns the foreign holders.
func (d *Detector) detect(ps []Process, apps []computeApp, utils map[int]float64) []string {
var holders []string var holders []string
for _, p := range ps { for _, p := range ps {
if d.procs[normName(p.Name)] { if d.procs[normName(p.Name)] {
holders = append(holders, fmt.Sprintf("%s (pid %d)", p.Name, p.PID)) holders = append(holders, fmt.Sprintf("%s (pid %d)", p.Name, p.PID))
} }
} }
if d.vramMB > 0 && apps != nil { if (d.vramMB > 0 && apps != nil) || (d.utilPct > 0 && utils != nil) {
names := make(map[int]string, len(ps)) names := make(map[int]string, len(ps))
for _, p := range ps { for _, p := range ps {
names[p.PID] = p.Name names[p.PID] = p.Name
@@ -115,6 +152,23 @@ func (d *Detector) detect(ps []Process, apps []computeApp) []string {
} }
holders = append(holders, fmt.Sprintf("%s (pid %d) using %d MiB VRAM", name, a.PID, a.UsedMB)) holders = append(holders, fmt.Sprintf("%s (pid %d) using %d MiB VRAM", name, a.PID, a.UsedMB))
} }
// Sorted for stable output (map iteration order is random).
pids := make([]int, 0, len(utils))
for pid := range utils {
pids = append(pids, pid)
}
slices.Sort(pids)
for _, pid := range pids {
util := utils[pid]
name := names[pid]
if d.ignore[normName(name)] || util < float64(d.utilPct) {
continue
}
if name == "" {
name = "unknown process"
}
holders = append(holders, fmt.Sprintf("%s (pid %d) using %.0f%% GPU", name, pid, util))
}
} }
return holders return holders
} }
+43 -4
View File
@@ -45,7 +45,7 @@ func TestParseComputeApps(t *testing.T) {
} }
func TestDetect(t *testing.T) { func TestDetect(t *testing.T) {
d := New([]string{"Cyberpunk2077.exe", "hl2"}, 1024, d := New([]string{"Cyberpunk2077.exe", "hl2"}, 1024, 0,
[]string{"ollama", "python", "pythonw"}, nil) []string{"ollama", "python", "pythonw"}, nil)
ps := []Process{ ps := []Process{
{PID: 10, Name: "ollama.exe"}, {PID: 10, Name: "ollama.exe"},
@@ -58,7 +58,7 @@ func TestDetect(t *testing.T) {
{PID: 40, UsedMB: 2048}, // foreign, above threshold {PID: 40, UsedMB: 2048}, // foreign, above threshold
{PID: 50, UsedMB: 100}, // foreign but below threshold {PID: 50, UsedMB: 100}, // foreign but below threshold
} }
holders := d.detect(ps, apps) holders := d.detect(ps, apps, nil)
if len(holders) != 2 { if len(holders) != 2 {
t.Fatalf("got %v, want 2 holders", holders) t.Fatalf("got %v, want 2 holders", holders)
} }
@@ -70,13 +70,52 @@ func TestDetect(t *testing.T) {
} }
} }
func TestDetectEngineUtil(t *testing.T) {
d := New(nil, 0, 30, []string{"dwm", "python"}, nil)
ps := []Process{
{PID: 10, Name: "dwm.exe"},
{PID: 20, Name: "game.exe"},
{PID: 40, Name: "browser.exe"},
}
utils := map[int]float64{
10: 45, // ignored: dwm
20: 61, // foreign, above threshold
30: 82, // foreign, unknown name
40: 5, // below threshold
}
holders := d.detect(ps, nil, utils)
want := []string{
"game.exe (pid 20) using 61% GPU",
"unknown process (pid 30) using 82% GPU",
}
if !slices.Equal(holders, want) {
t.Errorf("got %v, want %v", holders, want)
}
}
func TestDetectNothingConfigured(t *testing.T) { func TestDetectNothingConfigured(t *testing.T) {
d := New(nil, 0, nil, nil) d := New(nil, 0, 0, nil, nil)
if got := d.detect([]Process{{PID: 1, Name: "game.exe"}}, nil); len(got) != 0 { if got := d.detect([]Process{{PID: 1, Name: "game.exe"}}, nil, nil); len(got) != 0 {
t.Errorf("got %v, want none", got) t.Errorf("got %v, want none", got)
} }
} }
func TestParseGPUEngineInstance(t *testing.T) {
pid, eng, ok := parseGPUEngineInstance("pid_1234_luid_0x00000000_0x00011A2B_phys_0_eng_0_engtype_3D")
if !ok || pid != 1234 || eng != "3D" {
t.Errorf("got %d, %q, %v", pid, eng, ok)
}
pid, eng, ok = parseGPUEngineInstance("pid_42_luid_0x0_0x0_phys_0_eng_1_engtype_Copy")
if !ok || pid != 42 || eng != "Copy" {
t.Errorf("got %d, %q, %v", pid, eng, ok)
}
for _, bad := range []string{"", "something", "pid_", "pid_x_luid", "pid_-1_luid_0"} {
if _, _, ok := parseGPUEngineInstance(bad); ok {
t.Errorf("%q parsed, want failure", bad)
}
}
}
func TestProcessesLive(t *testing.T) { func TestProcessesLive(t *testing.T) {
if runtime.GOOS != "windows" && runtime.GOOS != "linux" { if runtime.GOOS != "windows" && runtime.GOOS != "linux" {
t.Skip("no process listing on this platform") t.Skip("no process listing on this platform")
+36
View File
@@ -0,0 +1,36 @@
package game
import (
"errors"
"strconv"
"strings"
)
// errNotPrimed marks the first PDH sample after opening a query: rate-based
// counters (like engine utilization) need two collections before they
// return meaningful values.
var errNotPrimed = errors.New("GPU engine counter needs a second sample")
// parseGPUEngineInstance splits a PDH "GPU Engine" instance name —
// "pid_1234_luid_0x00000000_0x00011A2B_phys_0_eng_0_engtype_3D" — into PID
// and engine type ("3D", "Copy", "VideoDecode", ...). engType is empty when
// the name carries no engtype marker.
func parseGPUEngineInstance(name string) (pid int, engType string, ok bool) {
rest, found := strings.CutPrefix(name, "pid_")
if !found {
return 0, "", false
}
digits, rest, found := strings.Cut(rest, "_")
if !found {
return 0, "", false
}
pid, err := strconv.Atoi(digits)
if err != nil || pid < 0 {
return 0, "", false
}
const marker = "engtype_"
if i := strings.LastIndex(rest, marker); i >= 0 {
engType = rest[i+len(marker):]
}
return pid, engType, true
}
+16
View File
@@ -0,0 +1,16 @@
//go:build !windows
package game
import "errors"
// errNoEngineCounters marks platforms without per-process GPU engine
// counters (the PDH path is Windows-only).
var errNoEngineCounters = errors.New("per-process GPU engine counters are only available on Windows")
// gpuEngineSampler is a stub on non-Windows platforms.
type gpuEngineSampler struct{}
func openGPUEngineSampler() (*gpuEngineSampler, error) { return nil, errNoEngineCounters }
func (s *gpuEngineSampler) sample() (map[int]float64, error) { return nil, errNoEngineCounters }
+111
View File
@@ -0,0 +1,111 @@
//go:build windows
package game
import (
"fmt"
"unsafe"
"golang.org/x/sys/windows"
)
// Per-process GPU engine utilization via PDH — the same counters Task
// Manager's "GPU engine" columns read. Unlike nvidia-smi's compute-apps
// this covers graphics work under WDDM, so games show up. The counter is
// added with PdhAddEnglishCounterW, which is independent of the Windows
// display language.
var (
pdhDLL = windows.NewLazySystemDLL("pdh.dll")
procPdhOpenQuery = pdhDLL.NewProc("PdhOpenQueryW")
procPdhAddEnglishCounter = pdhDLL.NewProc("PdhAddEnglishCounterW")
procPdhCollectQueryData = pdhDLL.NewProc("PdhCollectQueryData")
procPdhGetFormattedCounterArray = pdhDLL.NewProc("PdhGetFormattedCounterArrayW")
procPdhCloseQuery = pdhDLL.NewProc("PdhCloseQuery")
)
const (
pdhFmtDouble = 0x00000200 // PDH_FMT_DOUBLE
pdhMoreData = 0x800007D2 // PDH_MORE_DATA
)
// pdhCountervalueItem mirrors PDH_FMT_COUNTERVALUE_ITEM (64-bit, double).
type pdhCountervalueItem struct {
name *uint16
cStatus uint32
_ uint32 // alignment padding
value float64
}
// gpuEngineSampler holds an open PDH query on the wildcard GPU Engine
// utilization counter. Keeping the query open across polls is what makes
// the rate-based values meaningful; the sampler lives as long as the
// process (PdhCloseQuery would only matter on unload).
type gpuEngineSampler struct {
query uintptr // PDH_HQUERY
counter uintptr // PDH_HCOUNTER
primed bool
}
// openGPUEngineSampler opens a query on the per-process GPU engine
// utilization counter (all instances).
func openGPUEngineSampler() (*gpuEngineSampler, error) {
var q uintptr
if r, _, _ := procPdhOpenQuery.Call(0, 0, uintptr(unsafe.Pointer(&q))); r != 0 {
return nil, fmt.Errorf("PdhOpenQuery: status %#x", r)
}
path, err := windows.UTF16PtrFromString(`\GPU Engine(*)\Utilization Percentage`)
if err != nil {
procPdhCloseQuery.Call(q)
return nil, err
}
var c uintptr
if r, _, _ := procPdhAddEnglishCounter.Call(q, uintptr(unsafe.Pointer(path)), 0, uintptr(unsafe.Pointer(&c))); r != 0 {
procPdhCloseQuery.Call(q)
return nil, fmt.Errorf("PdhAddEnglishCounter: status %#x", r)
}
return &gpuEngineSampler{query: q, counter: c}, nil
}
// sample collects the counter once and returns per-PID 3D-engine
// utilization in percent. The first call after open only primes the rate
// calculation and returns errNotPrimed. Processes can drive several 3D
// engines; their values are summed.
func (s *gpuEngineSampler) sample() (map[int]float64, error) {
if r, _, _ := procPdhCollectQueryData.Call(s.query); r != 0 {
return nil, fmt.Errorf("PdhCollectQueryData: status %#x", r)
}
if !s.primed {
s.primed = true
return nil, errNotPrimed
}
var size, count uint32
r, _, _ := procPdhGetFormattedCounterArray.Call(s.counter, pdhFmtDouble,
uintptr(unsafe.Pointer(&size)), uintptr(unsafe.Pointer(&count)), 0)
if r == pdhMoreData && size == 0 {
return nil, nil // no GPU engine instances at all
}
if r != pdhMoreData {
return nil, fmt.Errorf("PdhGetFormattedCounterArray(size): status %#x", r)
}
buf := make([]byte, size)
r, _, _ = procPdhGetFormattedCounterArray.Call(s.counter, pdhFmtDouble,
uintptr(unsafe.Pointer(&size)), uintptr(unsafe.Pointer(&count)),
uintptr(unsafe.Pointer(&buf[0])))
if r != 0 {
return nil, fmt.Errorf("PdhGetFormattedCounterArray: status %#x", r)
}
items := unsafe.Slice((*pdhCountervalueItem)(unsafe.Pointer(&buf[0])), int(count))
out := make(map[int]float64)
for i := range items {
if items[i].cStatus != 0 || items[i].name == nil {
continue
}
pid, engType, ok := parseGPUEngineInstance(windows.UTF16PtrToString(items[i].name))
if !ok || engType != "3D" {
continue // only the 3D engine marks game-like work
}
out[pid] += items[i].value
}
return out, nil
}
+33
View File
@@ -0,0 +1,33 @@
//go:build windows
package game
import (
"errors"
"testing"
"time"
)
// TestGPUEngineSamplerLive opens the real PDH query and takes two samples;
// the first only primes the rate counters. Skipped (not failed) when the
// machine has no GPU counters.
func TestGPUEngineSamplerLive(t *testing.T) {
s, err := openGPUEngineSampler()
if err != nil {
t.Skipf("no GPU engine counters: %v", err)
}
if _, err := s.sample(); !errors.Is(err, errNotPrimed) {
t.Fatalf("first sample: err = %v, want errNotPrimed", err)
}
time.Sleep(200 * time.Millisecond)
utils, err := s.sample()
if err != nil {
t.Fatalf("second sample: %v", err)
}
for pid, util := range utils {
if pid < 0 || util < 0 {
t.Errorf("pid %d: util %.2f", pid, util)
}
}
t.Logf("%d processes with 3D-engine usage", len(utils))
}