Seven decisions, six systems
A single request passes seven stations. Something is decided at each one — and every one of those decisions rests on configuration that somebody entered by hand beforehand.
| Station | What gets decided | Where it is configured | |
|---|---|---|---|
| 1 | Authentication | Who is this, and which team do they belong to? | Keycloak |
| 2 | Quota check | Budget spent? RPM or TPM exceeded? | LiteLLM |
| 3 | Content check | Does the prompt carry protected data? | gateway rules |
| 4 | Model selection | May this user use this model for this data class? | LiteLLM · allowlist |
| 5 | Serving | Which model revision, and is an instance running? | KServe |
| 6 | Execution | Does the request fit the batch? Is there KV cache left? | vLLM arguments |
| 7 | Accounting | How many tokens, at what cost, charged to which team? | LiteLLM · DCGM · cost centre |
Six systems, six formats, six places.
No system sees more than its own slice. Raise a model's context length and you have to work out for yourself whether the quotas of the teams assigned to it still add up. Today that sum is done on a calculator, if it is done at all.
An appendix of the Enterprise AI Platform Lab Guide lists what does not belong in the repository. One line reads:
Is built from the catalog, not maintained.
Lab Guide, appendix B.5
In the book that is an instruction to the reader: write yourself a script. nyrvex is the answer to what happens when that “built” becomes a product — with state of its own, drift detection, a way back, and an interface.
The catalog is the only truth. Everything else is generated from it.
That book is the Enterprise AI Platform Lab Guide — 386 pages on how this work is done by hand today. What is in it.
Not a toolbox. A loop.
Catalog
what should hold
Models, teams, environments. A model has a purpose, a data class, a context length, an approval carrying a date and a signature, and an evaluation set it has to pass. A team has a cost centre, a data class, the models it may use, a budget and limits. None of this is invented here — it is the structure of the reference repository from the book, no longer a convention but a data model.
Sizing
whether it fits
Before a change is rolled out, it is worked through: usable KV cache, concurrent sequences, requests in the system by Little's law, price per million tokens. nyrvex knows both the model and the quotas of the teams allowed to use it — which is why it can tell you the two do not add up.
Reconcile
who carries it out
The configuration of every system involved is generated from the catalog and kept in step with it: the gateway's model list and quotas, the serving layer's InferenceServices, groups and mappers in the identity provider, provider keys in the secret store. When a system drifts from the catalog, that is visible and reversible.
Accounting
what actually happened
Usage and cost arise at the gateway, utilisation at the GPU, readiness at the serving layer. Only joined do they answer a question worth asking: what does a thousand tokens cost in-house, and which team spent them?
↺ Accounting supplies the measurements that correct sizing's own assumptions. That is why it is a loop and not a chain.
Kubernetes has a scheduler. It is no use for GPUs.
Kubernetes is three things: a declarative resource model, a scheduler that maps intent onto available capacity, and control loops that pull actual state towards desired state.
The scheduler cannot be borrowed here. It treats a GPU as an integral, non-overcommittable resource: it knows a card is free. It knows nothing about video memory, context length, weights or KV cache — which are exactly the quantities that decide whether a model actually runs.
This arithmetic is well understood. It is simply not part of the system yet; it is the operator's job.
$ nyrvex apply -f catalog/models.yaml # draft ✓ corp-chat sized · rolled out ⏸ corp-rag sized · waiting for capacity $ nyrvex describe model corp-rag State Waiting for capacity — for 4 min Sizing context 16384 · 2 replicas · pool a100-40g weights 7.1 GB activations 1.8 GB usable 36.0 GB KV cache left 27.1 GB per sequence 1.05 GB → 25 concurrent sequences per replica, 50 in total The teams assigned to it require 118 concurrent. Suggestion context 8192 → 51 per replica 5 replicas → 125 in total sizing: override (recorded)
Draft of the command-line output. Arithmetic follows the formulas in appendix C.8.
The declaration is accepted; the rollout is not. The model sits in the catalog and waits — with the arithmetic as its reason, rather than an error twenty minutes into loading. You can override the calculation; the override is recorded.
What nyrvex is not
- Not a replacement for the tools you run. Gateway, serving layer, inference server, identity provider and secret store stay exactly what they are. nyrvex configures them — the way Kubernetes did not replace the container runtime, but conducts it.
- Not an inference server. nyrvex runs no models and sits outside the request path. If it fails, requests keep flowing; you simply cannot change anything until it is back.
- Not an ML platform. No training, no fine-tuning, no experiment tracking. What happens inside the model matters only as far as it decides something an operator has to decide.
- Not a replacement for Kubernetes. The analogy describes the relationship, not the ambition.
An architecture, not a product
There is nothing to download and no timeline. The commands shown are drafts. If you came looking for a waiting-list form: there isn't one of those either.
Settled so far: a resource model of its own as the single truth, built on the repository structure from the Lab Guide; the four parts; a command line and an interface; and that sizing holds back the rollout rather than rejecting the declaration.
Three decisions, and the reasoning
None of them was obvious. If you would have decided differently, that is the most useful thing you can send me.
The catalog separates the file from the promise
A Model states what is true about the artefact; an Endpoint states what we promise about the service. The same weights serve chat at 8k and RAG at 16k without being declared twice — and raising a revision becomes one operation that marks every endpoint whose approval has to be renewed.
Sizing holds back the rollout — it does not reject the declaration
An apply always succeeds, so a GitOps sync never breaks because someone changed a context length. What waits is the rollout, carrying the arithmetic as its reason. The calculation can be overridden, and the override is recorded.
The control plane lives where the gateway database lives
One operator in the cluster that already carries gateway, identity and secrets — not one per cluster. A team has one budget, so it needs one counter. Endpoints and GPU pools stay local to their cluster; everything a quota depends on stays central.
If you run an AI platform, you have probably made at least one of these calls differently. I read every message: contact@nyrvex.com
The reasoning behind all of this is written out in the Enterprise AI Platform Lab Guide (German): thomaszachmann.de/buch