Nodes
Queues
Recent requests
Tokens over time
Breakdown by API key
Nodes
Add or update a node
A node with a MAC and Always on unticked is assumed to sleep when idle:
once its last resident model is about to idle out it is about to sleep and
takes no client requests (the scheduler still warms models onto it under demand). Tick
Always on for a machine that never sleeps even though it has a MAC — it
then stays a request target while empty, and the scheduler prefers it for a lone replica.
Both fields hot-reload; address and VRAM changes need a restart.
API keys
Add or rotate a key
Secrets are written to the .env file; config.yaml only ever
stores a ${VAR} reference. Adding or rotating a key takes effect
immediately — no restart needed.
Model placement policy
Pin or restrict a model
A capability tag containing Vectorization or Embedding (or a model
name with embed) marks a model as an embedder, so the router warms and serves it
via /api/embeddings. The router also self-corrects from the node's own
/api/show classification, so a missing tag no longer breaks an embedder — but
set it here to clear the mislabeled badge and record the dimension.
This is the real placement control: pin_nodes restricts a model to those
nodes, exclude_nodes keeps it off them. Distinct from the
pinned badge on a node card, which only means Ollama is holding the model
resident because of the keep-alive below.
Context window is the window the router loads this model at, whatever a
request asks for (still capped by what each node's VRAM can hold); blank sizes each load to
the request, with a client's options.num_ctx as a ceiling. Ollama otherwise applies a conservative
built-in default, and for an embedding model longer input is rejected rather than
truncated — the input length exceeds the context length error. Set the model's
real capability here (nomic-embed-text supports 8192) and every client gets it
without changing any application.
Ollama budget is what the node's Ollama reserves for the model at
load time, when that differs from the VRAM it really uses. The router measures the real
footprint — but the node plans from its own estimate of the GGUF, and for a model that keeps
big tensors in system RAM (Gemma 4 / 3n "E" models) that estimate is 2–3× reality: the node
then evicts idle neighbours and spills the next model to the CPU while the card reads half
empty. A node card shows evicted by Ollama · model when that happens and
Ollama budgets N GB next to its free figure. Set this to the number the node
acts on (its Ollama log during a load: sched.go … predicted="10.6 GiB") and
fit, eviction, and spill-heal planning use it. Ollama resident is the
smaller figure the node books once the model is up (updated VRAM based on existing
loaded models: total − available), so a neighbour it evicted to load this model can
come back beside it; blank means the measured size (plus what the runner holds beyond
/api/ps), which is what Ollama books. Blank either to clear it.
Raise it together with KV accounting. A bigger context multiplies the
KV-cache VRAM every concurrent request needs. If the model has no
kv_cache_bytes_per_token (and no cluster-wide
scheduler.kv_cache_bytes_per_token), admission costs each request at zero and
only max_concurrency applies — which counts requests, not bytes. On an 8 GB
card that combination does not queue under load, it runs out of memory. Measure the real
cost against /api/ps size_vram and set it in
config.yaml, or lower concurrency. The router logs
router.context_without_kv_accounting at startup naming any model in this state.
Runtime settings
-1m (any negative duration) keeps every warmed model resident forever —
which is why models show as pinned. A real duration like 10m
lets idle models unload on their own and frees VRAM without intervention.
The last three tune anti-thrash on a cluster whose nodes sleep. Protect a loaded model for keeps a freshly-loaded model from being evicted for that long; wake-before-evict grace is how long the scheduler waits for a sleeping node that holds a needed model to wake before it evicts one from a busy node instead (0 disables); queue hysteresis is how long a queue must persist before it rebalances.
Config file
Every save validates the whole config first and leaves a timestamped
.bak beside it, so a bad edit is rejected rather than applied.
Add a model
Pulls to every node with enough VRAM. Downloads run in the background and resume if interrupted. Inspect first pulls to a single probe node and shows what Ollama reports — the model's type, vector dimension, size — so you can verify it's the model you meant before distributing it to the cluster.
Active downloads
Model placement & speed
Each cell shows placement and, once measured, the model's decode speed on that node (tok/s). Click any cell where the model is present to benchmark that one pair — even if the node is asleep-cold and the model isn't loaded. A full run also happens automatically each week. Embedding/vectorization models are skipped (they generate no decode tokens), and a model too large for a node's VRAM is skipped with a reason.