Cluster Router

Sign in with an API key from your router config.

▦

Cluster Router

connecting

Nodes

Queues

Recent requests

Export CSV

Tokens over time

Breakdown by API key

Nodes

Add or update a node

A node with a MAC and Always on unticked is assumed to sleep when idle: once its last resident model is about to idle out it is about to sleep and takes no client requests (the scheduler still warms models onto it under demand). Tick Always on for a machine that never sleeps even though it has a MAC — it then stays a request target while empty, and the scheduler prefers it for a lone replica. Both fields hot-reload; address and VRAM changes need a restart.

API keys

Add or rotate a key

Secrets are written to the .env file; config.yaml only ever stores a ${VAR} reference. Adding or rotating a key takes effect immediately — no restart needed.

Model placement policy

Pin or restrict a model

A capability tag containing Vectorization or Embedding (or a model name with embed) marks a model as an embedder, so the router warms and serves it via /api/embeddings. The router also self-corrects from the node's own /api/show classification, so a missing tag no longer breaks an embedder — but set it here to clear the mislabeled badge and record the dimension.

This is the real placement control: pin_nodes restricts a model to those nodes, exclude_nodes keeps it off them. Distinct from the pinned badge on a node card, which only means Ollama is holding the model resident because of the keep-alive below.

Context window is the window the router loads this model at, whatever a request asks for (still capped by what each node's VRAM can hold); blank sizes each load to the request, with a client's options.num_ctx as a ceiling. Ollama otherwise applies a conservative built-in default, and for an embedding model longer input is rejected rather than truncated — the input length exceeds the context length error. Set the model's real capability here (nomic-embed-text supports 8192) and every client gets it without changing any application.

Ollama budget is what the node's Ollama reserves for the model at load time, when that differs from the VRAM it really uses. The router measures the real footprint — but the node plans from its own estimate of the GGUF, and for a model that keeps big tensors in system RAM (Gemma 4 / 3n "E" models) that estimate is 2–3× reality: the node then evicts idle neighbours and spills the next model to the CPU while the card reads half empty. A node card shows evicted by Ollama · model when that happens and Ollama budgets N GB next to its free figure. Set this to the number the node acts on (its Ollama log during a load: sched.go … predicted="10.6 GiB") and fit, eviction, and spill-heal planning use it. Ollama resident is the smaller figure the node books once the model is up (updated VRAM based on existing loaded models: total − available), so a neighbour it evicted to load this model can come back beside it; blank means the measured size (plus what the runner holds beyond /api/ps), which is what Ollama books. Blank either to clear it.

Raise it together with KV accounting. A bigger context multiplies the KV-cache VRAM every concurrent request needs. If the model has no kv_cache_bytes_per_token (and no cluster-wide scheduler.kv_cache_bytes_per_token), admission costs each request at zero and only max_concurrency applies — which counts requests, not bytes. On an 8 GB card that combination does not queue under load, it runs out of memory. Measure the real cost against /api/ps size_vram and set it in config.yaml, or lower concurrency. The router logs router.context_without_kv_accounting at startup naming any model in this state.

Runtime settings

-1m (any negative duration) keeps every warmed model resident forever — which is why models show as pinned. A real duration like 10m lets idle models unload on their own and frees VRAM without intervention.

The last three tune anti-thrash on a cluster whose nodes sleep. Protect a loaded model for keeps a freshly-loaded model from being evicted for that long; wake-before-evict grace is how long the scheduler waits for a sleeping node that holds a needed model to wake before it evicts one from a busy node instead (0 disables); queue hysteresis is how long a queue must persist before it rebalances.

Saved. A restart is required for these changes to take effect.

Config file

Every save validates the whole config first and leaves a timestamped .bak beside it, so a bad edit is rejected rather than applied.

Add a model

Pulls to every node with enough VRAM. Downloads run in the background and resume if interrupted. Inspect first pulls to a single probe node and shows what Ollama reports — the model's type, vector dimension, size — so you can verify it's the model you meant before distributing it to the cluster.

Active downloads

Model placement & speed

Each cell shows placement and, once measured, the model's decode speed on that node (tok/s). Click any cell where the model is present to benchmark that one pair — even if the node is asleep-cold and the model isn't loaded. A full run also happens automatically each week. Embedding/vectorization models are skipped (they generate no decode tokens), and a model too large for a node's VRAM is skipped with a reason.