MetricMapping Custom Resource
This document applies to the Modelplane main branch and not to the latest release v0.4.
How one component’s metrics become part of the modelplane_* surface. Modelplane renders every MetricMapping into each inference cluster’s collector, so a mapping is written once on the control plane and reaches the whole fleet. A mapping naming a component Modelplane already provides renames for is additive: its renames run after the built-in ones, and a later rename of the same metric wins.
Concept guide: Monitor the Fleet →
#Metadata
#Spec
How one component’s metrics become part of the modelplane_* surface. Modelplane renders every MetricMapping into each inference cluster’s collector, so a mapping is written once on the control plane and reaches the whole fleet. A mapping naming a component Modelplane already provides renames for is additive: its renames run after the built-in ones, and a later rename of the same metric wins.
How this metric combines over a deployment’s replicas. Each replica publishes its own series, told apart by the replica label, and a query over a deployment combines them. This says which combination is the right one: Sum for anything counted - requests, tokens, joules, a queue’s depth. Mean for a ratio, where summing reads two replicas at half capacity as one at full. Max for a saturation figure an alert fires on, where a mean hides the replica in trouble. Modelplane does not combine them in the collector. A scrape of one replica is one batch, so a collector that added them up would be adding readings taken at different moments, and two readings of one cumulative counter sum to twice the traffic that happened. The backend holds every replica’s series and combines them at query time, where the arithmetic is right. Required, with no default, because the wrong combination is silent: a deployment reports a number that looks entirely plausible.
The metric’s name as the component emits it, matched exactly. Nothing here declares which engine a deployment runs: a name that no component emits simply matches nothing. Held to the characters a metric name can contain. The name is matched inside the collector’s own query language, so a quote here would end the comparison early and rename whatever the rest of the line matched.
What the component measures this in, when that isn’t the unit the name claims. Modelplane converts to the base unit: millijoules and milliseconds are divided by a thousand, nanoseconds by a billion, and mebibytes multiplied out to bytes. Say it whenever the source disagrees with the target, even where the factor looks obvious. A name ending in _bytes that holds mebibytes is the kind of thing nobody notices until a capacity review, and stating the source unit is what makes the conversion happen at all.
A label the component already emits, carried onto the new name and dropped from the series under its old one.
The label to set.
A fixed value, the same on every series this mapping produces. This is what tells two folded metrics apart.
What each of that label’s values becomes, for putting an engine’s own vocabulary into Modelplane’s. A value with no entry here is left as the component wrote it.
Only meaningful alongside from.
Take a part of a histogram as a counter of its own, rather than the histogram itself. Count is how many observations it holds, which is a request count where the histogram measures request duration. Sum is their total. The histogram carries on unchanged under its own name. This adds a series beside it.
What Modelplane calls it. Only modelplane_* leaves a cluster, so a metric with no name here is one nobody downstream can read.