Monitor the Fleet
This document applies to the Modelplane main branch and not to the latest release v0.4.
Modelplane runs an OpenTelemetry collector on every inference cluster. It collects from
every component Modelplane installs, which is more than your engines. It renames each
component’s series to a single modelplane_* vocabulary and pushes to a collector on your
control plane. That collector is your fleet’s
single egress point, and it sends to any backend that speaks OTLP.
Modelplane has no API for this: nothing to write, and nothing to keep in sync as your deployments change.
What you get
Every series carries cluster. A series about a deployment also carries deployment,
replica, namespace, engine, and role.
Each replica publishes its own series. Combine them in the query, the way the metric’s
acrossReplicas says: sum by (deployment) for anything counted, avg by (deployment)
for a ratio, max by (deployment) for a saturation figure an alert fires on. The
collector doesn’t add them up for you, because a scrape of one replica is one batch, and
adding readings taken at different moments is not the traffic that happened.
The replica is an index, not a pod. It’s bounded by the replica count and it survives a restart and a rolling update, so the series count doesn’t grow every time you deploy.
Some of what you can read:
| Metric | Means |
|---|---|
modelplane_frontend_ttft_seconds | Time to the first token, measured at the gateway |
modelplane_frontend_tpot_seconds | Time per output token, measured at the gateway |
modelplane_frontend_request_duration_seconds | What the caller waited, end to end |
modelplane_request_queue_seconds | How long a request waited before the engine started |
modelplane_requests_waiting | Queue depth per engine |
modelplane_kv_cache_utilization_ratio | KV-cache occupancy, averaged over replicas |
modelplane_request_input_tokens | Prompt size, as a histogram |
modelplane_request_output_tokens | Generated length, as a histogram |
modelplane_gpu_memory_used_bytes | Framebuffer memory in use, per GPU |
modelplane_energy_joules_total | Energy drawn since the driver last reloaded |
Latency appears twice on purpose. The frontend_ series are what your caller experienced,
measured at the gateway. The engine’s own series are what the engine spent. When the
frontend number is slow and the engine number isn’t, the problem is routing, queueing, or
the network rather than the model.
Saturation gauges come as a pair. The average is what you plan capacity against; the _max
is what you alert on, because three replicas at 0.3 and one at 0.99 average to something
comfortable while the fourth evicts and recomputes. A high _max beside
modelplane_requests_preempted_total climbing is one replica thrashing.
No series names a pod. Replicas are interchangeable, so they’re summed before the metrics leave the cluster; a rolling update would otherwise leave a dead series behind for every pod it replaced.
Sending it somewhere
Create a TelemetryDestination naming whatever you already run:
apiVersion: modelplane.ai/v1alpha1
kind: TelemetryDestination
metadata:
name: default
spec:
sinks:
- name: primary
type: otlphttp
endpoint: https://otel.example.internaltype names a collector exporter, by the name OpenTelemetry gives it.
Put the credential in a Secret, name it with the sink’s secretRef, and say which key holds
the token. Modelplane composes the authenticator and wires it up, and the token never
appears in kubectl get -o yaml:
spec:
sinks:
- name: primary
type: otlphttp
endpoint: https://otel.example.internal
secretRef:
name: telemetry-credentials
auth:
bearerTokenKey: tokenIt reads the token from a file rather than the environment, so rotating it doesn’t need the collector restarted.
If you run Prometheus, export to that instead and query the fleet there:
spec:
sinks:
- name: prometheus
type: prometheusremotewrite
endpoint: https://prom.example.internal/api/v1/writeName more than one sink and every one gets the whole stream. Each carries its own credential, so a vendor and your own Prometheus don’t have to share a Secret:
spec:
sinks:
- name: vendor
type: otlphttp
endpoint: https://otel.vendor.example
secretRef:
name: vendor-token
auth:
bearerTokenKey: token
- name: prometheus
type: prometheusremotewrite
endpoint: https://prom.example.internal/api/v1/writeThat is two copies of the fleet’s metrics, billed twice.
Anything else the exporter takes goes under config, passed through as you wrote it:
- name: vendor
type: otlphttp
endpoint: https://otel.vendor.example
config:
compression: gzip
sending_queue:
queue_size: 10000
tls:
ca_file: /etc/ssl/certs/internal.pemModelplane doesn’t model what an exporter is, so its TLS, retry and queue settings all work,
and a sink keeps working when the collector gains a setting Modelplane has never heard of.
An authentication scheme Modelplane doesn’t compose works the same way: define the extension
under spec.extensions and name it from the sink’s config, which is what the auth block
above does for you.
Until you create one, Modelplane composes no collectors: nothing here stores anything, so collecting with nowhere to send it would spend GPU-cluster memory on samples nobody reads. Creating a destination turns collection on everywhere at once, and there’s no per-deployment opt-out.
Your clusters reach the control plane, and only the control plane reaches your backend. A cluster with no route to your observability stack still reports, and the backend’s credential lives in one place instead of on every GPU cluster.
Computing rates, quantiles, and ratios
A collector transforms each measurement as it passes it on. It holds no history, so it produces no rates and no quantiles. Your backend does that. A fleet-wide p99:
histogram_quantile(0.99, sum by (le) (
rate(modelplane_frontend_ttft_seconds_bucket{model="Qwen/Qwen3-8B"}[5m])))Modelplane has no dashboards of its own. What it exports is counters and histogram buckets, and your backend derives the rates and quantiles at query time. To precompute them instead, export to Prometheus and write recording rules there.
Engines
Modelplane renames vLLM’s and SGLang’s own metrics for you, so neither needs a mapping. SGLang needs one flag to publish them at all, below.
Any other OpenAI-compatible engine reports its top-line numbers with no configuration. The
gateway measures those, not the engine, so modelplane_frontend_* works for an engine
Modelplane has never seen.
To normalize that engine’s own metrics as well, create a MetricMapping:
apiVersion: modelplane.ai/v1alpha1
kind: MetricMapping
metadata:
name: my-engine
spec:
metrics:
- from: my_engine_queued_requests
to: modelplane_requests_waiting
- from: my_engine_kv_transfer_ms
fromUnit: Milliseconds
to: modelplane_request_kv_transfer_secondsModelplane renders every mapping into every cluster’s collector, so you write one once.
from is the name your engine emits and to is what Modelplane calls it.
Say fromUnit whenever the engine measures in something other than the unit the name
claims, and Modelplane converts to the base one. Skipping it is the expensive mistake here:
a series named _seconds that holds milliseconds reads a thousand times fast, and nothing
downstream can tell.
Rename only where the measurements agree. Two engines’ histograms under one name are worth less than nothing if their buckets disagree, because a quantile over them is wrong rather than approximate.
SGLang publishes /metrics only when it runs with --enable-metrics, so add that to its
engine args. vLLM needs nothing.
Why engine latency and gateway latency differ
modelplane_request_ttft_seconds comes from the engine, and engines bucket their
histograms differently. vLLM resolves down to a millisecond. SGLang resolves to a hundred
of them. A quantile across both is wrong, not approximate. Use the engine series to compare
one engine against itself, and the frontend_ series for anything fleet-wide.
Some measurements don’t translate at all. SGLang’s inter-token latency isn’t vLLM’s time per output token, so neither is renamed onto a shared name. The gateway measures time per output token for both.
Migrating from a hand-written PodMonitor
Modelplane used to have you write a PodMonitor and reach an in-cluster Prometheus over a
port-forward. Both are gone. Three steps to move across, and two of them fail quietly if
you skip them.
Keep your Prometheus, and point a destination at it. Collection becomes a push, so your store stops scraping and starts receiving. Same Prometheus, same retention, same Grafana:
spec:
sinks:
- name: prometheus
type: prometheusremotewrite
config:
endpoint: http://prometheus.monitoring.svc:9090/api/v1/writeDelete the monitors you wrote. A PodMonitor or ScrapeConfig pointed at your engines
keeps working against your own Prometheus, so nothing appears to break and you collect
everything twice, under vllm:* and under modelplane_*, paying for both. One written
against Modelplane’s Prometheus stops being read by anything, because the operator goes with
the stack.
Rewrite your dashboard queries. Names change, and so do the labels: group by
deployment rather than model_name, there’s no pod label, and every series carries
cluster and replica. A panel that showed one engine now shows one replica, so wrap it
in sum by (deployment) or the aggregation that metric’s acrossReplicas names.
| Was | Is |
|---|---|
vllm:time_to_first_token_seconds | modelplane_request_ttft_seconds |
vllm:e2e_request_latency_seconds | modelplane_request_duration_seconds |
vllm:request_queue_time_seconds | modelplane_request_queue_seconds |
vllm:request_prefill_time_seconds | modelplane_request_prefill_seconds |
vllm:request_decode_time_seconds | modelplane_request_decode_seconds |
vllm:num_requests_running | modelplane_requests_running |
vllm:num_requests_waiting | modelplane_requests_waiting |
vllm:kv_cache_usage_perc | modelplane_kv_cache_utilization_ratio |
vllm:num_preemptions_total | modelplane_requests_preempted_total |
vllm:prefix_cache_hits_total | modelplane_prefix_cache_hits_total |
DCGM_FI_DEV_FB_USED | modelplane_gpu_memory_used_bytes |
DCGM_FI_DEV_GPU_TEMP | modelplane_gpu_temperature_celsius |
DCGM_FI_DEV_POWER_USAGE | modelplane_gpu_power_watts |
DCGM_FI_PROF_PIPE_TENSOR_ACTIVE | modelplane_gpu_tensor_active_ratio |
envoy_cluster_upstream_rq_time | modelplane_frontend_request_duration_seconds |
Some have no replacement. A rename carries one metric to one name, so the counters that
would fold several series under one label - tokens by direction, responses by reason,
requests by status - aren’t part of this surface yet. Keep reading those from your engine
and your gateway directly. For prompt and output size, modelplane_request_input_tokens
and modelplane_request_output_tokens carry the same measurement as histograms.
vllm:inter_token_latency_seconds isn’t renamed, because
SGLang publishes a metric of the same name measuring something else; use
modelplane_frontend_tpot_seconds, which the gateway measures the same way for every
engine. DCGM_FI_DEV_GPU_UTIL isn’t renamed either, because it only tells you the card
wasn’t idle; use modelplane_gpu_compute_active_ratio and
modelplane_gpu_tensor_active_ratio.
You can also defer the rewrite. A recording rule rebuilds an old name from a new one, so a dashboard keeps working untouched while you migrate it:
- record: vllm:time_to_first_token_seconds_bucket
expr: label_replace(modelplane_request_ttft_seconds_bucket,
"model_name", "$1", "deployment", "(.*)")Load that into the Prometheus you already run and nothing on the dashboard changes. It’s one rule evaluation per metric over series your backend already holds, so it costs far less than collecting everything twice. Write one per name in the table above, and delete them once the panels use the new names.
A series no statement renames doesn’t leave the cluster. If a panel needs an engine’s
own name, write a MetricMapping that renames it onto the modelplane_* surface: a
mapping for an engine Modelplane already knows adds to the built-in renames rather
than replacing them.