Skip to content

Instantly share code, notes, and snippets.

@rominf
Created October 6, 2026 07:43
Show Gist options
  • Select an option

  • Save rominf/82f01b6d52feb98e90f7463f639b4ce1 to your computer and use it in GitHub Desktop.

Select an option

Save rominf/82f01b6d52feb98e90f7463f639b4ce1 to your computer and use it in GitHub Desktop.
AMD input on the xpumd HW metric semantics

AMD input on the xpumd HW metric semantics

Against xpumd/receiver/intelxpu/semantics.md at 9770382. Section headings below match that document. Examples are one MI300X: 8 XCD dies, 8 HBM3 stacks, 8 XCP partitions, 7 XGMI links to peer GPUs, PCIe Gen5 x16 to the host.

Where AMD agrees, it says so briefly. The substance is in the three open questions and the four gaps at the end.

Solution alternatives — agreed

Dropping the common / "semi-common" / hw.<component>.<metric> split and linking everything the same way is the right call, and the reasoning in "The problem" holds for AMD hardware too. Same for the observation that there is no "HW exists" info metric type to act as a parent — that gap is what makes the linking question hard to answer cleanly.

Hardware status alternatives — option 3

Of the three listed, AMD's preference is the third: a separate attribute saying whether a state is an issue. Proposed as hw.severity on hw.status:

hw.status{hw.id="gpu-0-xgmi-0", hw.type="interconnect", hw.state="link_up",         hw.severity="ok"}       1
hw.status{hw.id="gpu-0-hbm-3",  hw.type="memory",       hw.state="ecc_correctable", hw.severity="degraded"} 1
hw.status{hw.id="gpu-0",        hw.type="gpu",          hw.state="reset_needed",    hw.severity="failed"}   1

hw.status{hw.severity!="ok"} then finds every problem across a fleet without the query enumerating state names, and without the consumer knowing the vendor. Consumers that ignore the attribute see exactly what they see today.

Option 1 (report only issues) makes absence carry meaning, which breaks on any gap in collection. Option 2 (enumerate OK names in the spec) needs a spec change for every new state, and each vendor's list differs.

The aggregate GPU ok added in 9770382 is close to this, and AMD would keep it. The part that argues for moving the decision onto an attribute is the note attached to it — that the state is not emitted at all when the source states cannot be read. A consumer seeing no ok cannot distinguish healthy from unreadable. With a severity attribute the producer can emit hw.severity="unknown" and say so. It also generalizes past GPU: the same rollup is needed per component type, and computing it producer-side does not scale to every subcomponent.

Metric namespacing alternatives — left column, uniformly

AMD's position is the <metric>{<attrib>="<component>"} column of the table, applied without carve-outs — including utilization, which the current OTel spec keeps as hw.gpu.utilization.

The argument is that per-component naming is a back-compat liability the moment a second component grows the same notion. Utilization is the live example: CPUs, NPUs, accelerators, disks and NICs all have it, GPU was just named first. One dashboard query shape (hw.power{hw.type="gpu"}, hw.utilization{hw.type="gpu"}) then works across vendors and component types.

Worth noting the spec already does this in places and is inconsistent with itself inside a single file. In model/hardware/gpu-metrics.yaml, GPU hw.errors and hw.status are declared under metric_refinements as refinements of the shared metric names, each with hw.type noted "MUST be set to gpu" — while hw.gpu.io, hw.gpu.utilization and hw.gpu.memory.{limit,usage,utilization} are declared as plain per-component metric names a few lines above. So this is asking the WG to apply its own convention uniformly rather than adopt a new one.

Two boundaries on that:

  • .info items stay per-component (hw.gpu.info, hw.gpu.memory.info, hw.cpu.info). Their attribute schemas differ by component type and there is no cross-component aggregation use case, so the argument above does not apply to them.

  • State items split by what they mean, not by namespace. A health state — one an operator would alert on — goes on hw.status with hw.severity, because it has to aggregate. A configuration or mode state carries no health signal by itself and goes with the .info class, per-component. ECC enabled / disabled / pending-reboot is configuration; ECC errors accruing is health.

    That second rule is what hw.gpu.ecc.state in 9391a58 already does in practice, so AMD agrees with the outcome. Flagging it because the commit message reasons from namespace reservation rather than from the kind of state, and the rule is what decides the next case — whichever metric that turns out to be.

Also agreed: the note asking for a standard attribute that says a metric belongs to a GPU sub-component without a JOIN. That is worth having and AMD would use it.

Linking alternatives — Alternative 1

Alternative 1 (shared hw.id + hw.name) with hw.parent for subcomponents that have their own metric streams. It handles GPU → die → HBM stack without extra machinery, and hw.subdevice_id is worth formalizing in the spec for the lightweight case.

Alternative 3's separate hw.link items double the series count for something the other two carry in attributes. Alternative 2 works, but making every item's hw.id unique means a consumer must resolve parentage before it can group a device's own metrics.

On the TODO in Alternative 1 — sub-component info under hw.<component> items: yes, and AMD would resolve it by splitting info from measurements. Info goes per-component (hw.gpu.memory.info), measurements stay in per-metric form (hw.memory.size{...}), exactly as in the second code block. That is consistent with the .info boundary above.

Four things AMD needs that are not in the draft

These are the cases where an MI300X currently has nothing conformant to emit.

  1. hw.partition.info — GPUs partition (XCP, MIG), and so do CPUs (sub-NUMA). Not hw.gpu.partition.*, for the same reason as below. Info item only; partition measurements use per-metric form. Needs a partition member on hw.type.

  2. hw.interconnect.info — XGMI, NVLink, Xe Link, PCIe, Infinity Fabric, CXL. Not GPU-prefixed; interconnects are not GPU-exclusive. Info item only. Needs an interconnect member on hw.type. CXL's protocol layering is not modelled yet and AMD does not have a proposal for it — flagging it as open rather than pretending otherwise.

  3. Per-engine split — hw.gpu.task currently has encoder/decoder/general. AMD needs gfx, compute, dma, video, jpeg distinguished, either by extending that enum or by a separate engine-type attribute.

  4. Derivation guidance — say explicitly that hw.power MAY be derived from hw.energy, that derived metrics SHOULD avoid double-counting across hierarchy levels, and that consumers computing from the same base counters SHOULD get the same value. Related to the hw.host.power / hw.power discussion in semantic-conventions#1055.

One blocker for all of the above

The GPU hw.state enum is closed — ok, degraded, failed, predicted_failure, as a normative MUST — and hw.type has neither interconnect nor partition. So an XGMI link or an XCP partition has no standard component type, which also removes the "use a non-GPU hw.type" escape from the closed enum.

The xpumd receiver is in the same position from the other side: it reports GPU reset_needed and throttled, neither of which the enum permits, and the document says as much. Two vendors working around the same enum is a reasonable basis for asking the WG to reopen it, and AMD would rather raise that jointly than separately.

Note also that the permitted values are themselves severity-shaped — ok / degraded / failed is a severity scale sitting in the state field. That conflation is the same one hw.severity is meant to resolve, which is why the two asks are related.


Backing this up, if useful: every position above is encoded in an OpenTelemetry Weaver registry that passes weaver registry check clean, and there is a working Collector config that translates AMD's existing Prometheus GPU exporter output into this shape with no changes to the exporter or its consumers. Happy to share either.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment