nyrvex.

nyrvex.

Control plane for AI infrastructure

Between your applications and your GPUs, one layer is missing.

Six systems run an enterprise AI platform. Each one knows its own slice: the gateway knows the quotas, the serving layer knows the replicas, the identity provider knows the teams. A person holds the rest together — by hand, every day.

nyrvex is that layer.

Written against Kubernetes 1.31 · KServe 0.14 · vLLM 0.8.x · LiteLLM 1.6x September 2026 — an architecture, not a product
APPLICATIONS · DEVELOPERS · BUSINESS UNITS nyrvex Catalog · Sizing · Reconcile · Accounting One catalog of models, teams and environments — everything below is generated from it, and kept in step. this layer is missing today its work is done by a person, by hand Keycloakidentity OpenBaosecrets LiteLLMgateway KServeserving vLLMinference DCGMmetrics KUBERNETES · NVIDIA GPU · VRAM
The target architecture of an enterprise AI platform — with the layer that does not exist yet.
The problem

Seven decisions, six systems

A single request passes seven stations. Something is decided at each one — and every one of those decisions rests on configuration that somebody entered by hand beforehand.

StationWhat gets decidedWhere it is configured
1AuthenticationWho is this, and which team do they belong to?Keycloak
2Quota checkBudget spent? RPM or TPM exceeded?LiteLLM
3Content checkDoes the prompt carry protected data?gateway rules
4Model selectionMay this user use this model for this data class?LiteLLM · allowlist
5ServingWhich model revision, and is an instance running?KServe
6ExecutionDoes the request fit the batch? Is there KV cache left?vLLM arguments
7AccountingHow many tokens, at what cost, charged to which team?LiteLLM · DCGM · cost centre

Six systems, six formats, six places.

No system sees more than its own slice. Raise a model's context length and you have to work out for yourself whether the quotas of the teams assigned to it still add up. Today that sum is done on a calculator, if it is done at all.

The thesis

An appendix of the Enterprise AI Platform Lab Guide lists what does not belong in the repository. One line reads:

Generated configuration

Is built from the catalog, not maintained.

Lab Guide, appendix B.5

In the book that is an instruction to the reader: write yourself a script. nyrvex is the answer to what happens when that “built” becomes a product — with state of its own, drift detection, a way back, and an interface.

The catalog is the only truth. Everything else is generated from it.

That book is the Enterprise AI Platform Lab Guide — 386 pages on how this work is done by hand today. What is in it.

The four parts

Not a toolbox. A loop.

Catalog

what should hold

Models, teams, environments. A model has a purpose, a data class, a context length, an approval carrying a date and a signature, and an evaluation set it has to pass. A team has a cost centre, a data class, the models it may use, a budget and limits. None of this is invented here — it is the structure of the reference repository from the book, no longer a convention but a data model.

Sizing

whether it fits

Before a change is rolled out, it is worked through: usable KV cache, concurrent sequences, requests in the system by Little's law, price per million tokens. nyrvex knows both the model and the quotas of the teams allowed to use it — which is why it can tell you the two do not add up.

Reconcile

who carries it out

The configuration of every system involved is generated from the catalog and kept in step with it: the gateway's model list and quotas, the serving layer's InferenceServices, groups and mappers in the identity provider, provider keys in the secret store. When a system drifts from the catalog, that is visible and reversible.

Accounting

what actually happened

Usage and cost arise at the gateway, utilisation at the GPU, readiness at the serving layer. Only joined do they answer a question worth asking: what does a thousand tokens cost in-house, and which team spent them?

↺ Accounting supplies the measurements that correct sizing's own assumptions. That is why it is a loop and not a chain.

The core

Kubernetes has a scheduler. It is no use for GPUs.

Kubernetes is three things: a declarative resource model, a scheduler that maps intent onto available capacity, and control loops that pull actual state towards desired state.

The scheduler cannot be borrowed here. It treats a GPU as an integral, non-overcommittable resource: it knows a card is free. It knows nothing about video memory, context length, weights or KV cache — which are exactly the quantities that decide whether a model actually runs.

This arithmetic is well understood. It is simply not part of the system yet; it is the operator's job.

$ nyrvex apply -f catalog/models.yaml                      # draft

  ✓ corp-chat   sized · rolled out
  ⏸ corp-rag    sized · waiting for capacity

$ nyrvex describe model corp-rag

  State      Waiting for capacity — for 4 min
  Sizing     context 16384 · 2 replicas · pool a100-40g

    weights              7.1 GB
    activations          1.8 GB
    usable              36.0 GB
    KV cache left       27.1 GB
    per sequence         1.05 GB
    → 25 concurrent sequences per replica, 50 in total

    The teams assigned to it require 118 concurrent.

  Suggestion context 8192  → 51 per replica
             5 replicas   → 125 in total
             sizing: override   (recorded)

Draft of the command-line output. Arithmetic follows the formulas in appendix C.8.

The declaration is accepted; the rollout is not. The model sits in the catalog and waits — with the arithmetic as its reason, rather than an error twenty minutes into loading. You can override the calculation; the override is recorded.

Boundaries

What nyrvex is not

  • Not a replacement for the tools you run. Gateway, serving layer, inference server, identity provider and secret store stay exactly what they are. nyrvex configures them — the way Kubernetes did not replace the container runtime, but conducts it.
  • Not an inference server. nyrvex runs no models and sits outside the request path. If it fails, requests keep flowing; you simply cannot change anything until it is back.
  • Not an ML platform. No training, no fine-tuning, no experiment tracking. What happens inside the model matters only as far as it decides something an operator has to decide.
  • Not a replacement for Kubernetes. The analogy describes the relationship, not the ambition.
Status

An architecture, not a product

There is nothing to download and no timeline. The commands shown are drafts. If you came looking for a waiting-list form: there isn't one of those either.

Settled so far: a resource model of its own as the single truth, built on the repository structure from the Lab Guide; the four parts; a command line and an interface; and that sizing holds back the rollout rather than rejecting the declaration.

Decisions

Three decisions, and the reasoning

None of them was obvious. If you would have decided differently, that is the most useful thing you can send me.

The catalog separates the file from the promise

A Model states what is true about the artefact; an Endpoint states what we promise about the service. The same weights serve chat at 8k and RAG at 16k without being declared twice — and raising a revision becomes one operation that marks every endpoint whose approval has to be renewed.

Sizing holds back the rollout — it does not reject the declaration

An apply always succeeds, so a GitOps sync never breaks because someone changed a context length. What waits is the rollout, carrying the arithmetic as its reason. The calculation can be overridden, and the override is recorded.

The control plane lives where the gateway database lives

One operator in the cluster that already carries gateway, identity and secrets — not one per cluster. A team has one budget, so it needs one counter. Endpoints and GPU pools stay local to their cluster; everything a quota depends on stays central.

If you run an AI platform, you have probably made at least one of these calls differently. I read every message: contact@nyrvex.com

The reasoning behind all of this is written out in the Enterprise AI Platform Lab Guide (German): thomaszachmann.de/buch