Against xpumd/receiver/intelxpu/semantics.md at 9770382. Section headings below match
that document. Examples are one MI300X: 8 XCD dies, 8 HBM3 stacks, 8 XCP partitions,
7 XGMI links to peer GPUs, PCIe Gen5 x16 to the host.
Where AMD agrees, it says so briefly. The substance is in the three open questions and the four gaps at the end.
Dropping the common / "semi-common" / hw.<component>.<metric> split and linking
everything the same way is the right call, and the reasoning in "The problem" holds for
AMD hardware too. Same for the observation that there is no "HW exists" info metric type
to act as a parent — that gap is what makes the linking question hard to answer cleanly.
Of the three listed, AMD's preference is the third: a separate attribute saying whether a
state is an issue. Proposed as hw.severity on hw.status:
hw.status{hw.id="gpu-0-xgmi-0", hw.type="interconnect", hw.state="link_up", hw.severity="ok"} 1
hw.status{hw.id="gpu-0-hbm-3", hw.type="memory", hw.state="ecc_correctable", hw.severity="degraded"} 1
hw.status{hw.id="gpu-0", hw.type="gpu", hw.state="reset_needed", hw.severity="failed"} 1
hw.status{hw.severity!="ok"} then finds every problem across a fleet without the query
enumerating state names, and without the consumer knowing the vendor. Consumers that
ignore the attribute see exactly what they see today.
Option 1 (report only issues) makes absence carry meaning, which breaks on any gap in collection. Option 2 (enumerate OK names in the spec) needs a spec change for every new state, and each vendor's list differs.
The aggregate GPU ok added in 9770382 is close to this, and AMD would keep it. The
part that argues for moving the decision onto an attribute is the note attached to it —
that the state is not emitted at all when the source states cannot be read. A consumer
seeing no ok cannot distinguish healthy from unreadable. With a severity attribute the
producer can emit hw.severity="unknown" and say so. It also generalizes past GPU: the
same rollup is needed per component type, and computing it producer-side does not scale
to every subcomponent.
AMD's position is the <metric>{<attrib>="<component>"} column of the table, applied
without carve-outs — including utilization, which the current OTel spec keeps as
hw.gpu.utilization.
The argument is that per-component naming is a back-compat liability the moment a second
component grows the same notion. Utilization is the live example: CPUs, NPUs, accelerators,
disks and NICs all have it, GPU was just named first. One dashboard query shape
(hw.power{hw.type="gpu"}, hw.utilization{hw.type="gpu"}) then works across vendors and
component types.
Worth noting the spec already does this in places and is inconsistent with itself inside a
single file. In model/hardware/gpu-metrics.yaml, GPU hw.errors and hw.status are
declared under metric_refinements as refinements of the shared metric names, each with
hw.type noted "MUST be set to gpu" — while hw.gpu.io, hw.gpu.utilization and
hw.gpu.memory.{limit,usage,utilization} are declared as plain per-component metric names
a few lines above. So this is asking the WG to apply its own convention uniformly rather
than adopt a new one.
Two boundaries on that:
-
.infoitems stay per-component (hw.gpu.info,hw.gpu.memory.info,hw.cpu.info). Their attribute schemas differ by component type and there is no cross-component aggregation use case, so the argument above does not apply to them. -
State items split by what they mean, not by namespace. A health state — one an operator would alert on — goes on
hw.statuswithhw.severity, because it has to aggregate. A configuration or mode state carries no health signal by itself and goes with the.infoclass, per-component. ECC enabled / disabled / pending-reboot is configuration; ECC errors accruing is health.That second rule is what
hw.gpu.ecc.statein9391a58already does in practice, so AMD agrees with the outcome. Flagging it because the commit message reasons from namespace reservation rather than from the kind of state, and the rule is what decides the next case — whichever metric that turns out to be.
Also agreed: the note asking for a standard attribute that says a metric belongs to a GPU sub-component without a JOIN. That is worth having and AMD would use it.
Alternative 1 (shared hw.id + hw.name) with hw.parent for subcomponents that have
their own metric streams. It handles GPU → die → HBM stack without extra machinery, and
hw.subdevice_id is worth formalizing in the spec for the lightweight case.
Alternative 3's separate hw.link items double the series count for something the other
two carry in attributes. Alternative 2 works, but making every item's hw.id unique means
a consumer must resolve parentage before it can group a device's own metrics.
On the TODO in Alternative 1 — sub-component info under hw.<component> items: yes,
and AMD would resolve it by splitting info from measurements. Info goes per-component
(hw.gpu.memory.info), measurements stay in per-metric form (hw.memory.size{...}),
exactly as in the second code block. That is consistent with the .info boundary above.
These are the cases where an MI300X currently has nothing conformant to emit.
-
hw.partition.info— GPUs partition (XCP, MIG), and so do CPUs (sub-NUMA). Nothw.gpu.partition.*, for the same reason as below. Info item only; partition measurements use per-metric form. Needs apartitionmember onhw.type. -
hw.interconnect.info— XGMI, NVLink, Xe Link, PCIe, Infinity Fabric, CXL. Not GPU-prefixed; interconnects are not GPU-exclusive. Info item only. Needs aninterconnectmember onhw.type. CXL's protocol layering is not modelled yet and AMD does not have a proposal for it — flagging it as open rather than pretending otherwise. -
Per-engine split —
hw.gpu.taskcurrently hasencoder/decoder/general. AMD needsgfx,compute,dma,video,jpegdistinguished, either by extending that enum or by a separate engine-type attribute. -
Derivation guidance — say explicitly that
hw.powerMAY be derived fromhw.energy, that derived metrics SHOULD avoid double-counting across hierarchy levels, and that consumers computing from the same base counters SHOULD get the same value. Related to thehw.host.power/hw.powerdiscussion in semantic-conventions#1055.
The GPU hw.state enum is closed — ok, degraded, failed, predicted_failure, as a
normative MUST — and hw.type has neither interconnect nor partition. So an XGMI link
or an XCP partition has no standard component type, which also removes the "use a non-GPU
hw.type" escape from the closed enum.
The xpumd receiver is in the same position from the other side: it reports GPU
reset_needed and throttled, neither of which the enum permits, and the document says
as much. Two vendors working around the same enum is a reasonable basis for asking the WG
to reopen it, and AMD would rather raise that jointly than separately.
Note also that the permitted values are themselves severity-shaped — ok / degraded /
failed is a severity scale sitting in the state field. That conflation is the same one
hw.severity is meant to resolve, which is why the two asks are related.
Backing this up, if useful: every position above is encoded in an OpenTelemetry Weaver
registry that passes weaver registry check clean, and there is a working Collector
config that translates AMD's existing Prometheus GPU exporter output into this shape with
no changes to the exporter or its consumers. Happy to share either.