Skip to content

feat: support metricRelabelings on the ServiceMonitor - #2987

Merged
rahulait merged 1 commit into
NVIDIA:mainfrom
gseidlerhpe:dsx-ws0930/issue-2938-metric-relabelings
Oct 1, 2026
Merged

rahulait merged 1 commit into
NVIDIA:mainfrom
gseidlerhpe:dsx-ws0930/issue-2938-metric-relabelings

Conversation

@gseidlerhpe

@gseidlerhpe gseidlerhpe commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Description

ServiceMonitorConfig only exposed relabelings, which maps to Prometheus target relabeling (Endpoint.RelabelConfigs) and runs before the scrape. Rewriting labels on scraped samples — e.g. collapsing exported_namespace/exported_pod back onto namespace/pod, the use case in #2938 — requires Endpoint.MetricRelabelConfigs, which had no API surface. Since the Helm chart passes dcgmExporter.serviceMonitor through to the ClusterPolicy verbatim, a metricRelabelings key was rejected by CRD schema validation.

This adds MetricRelabelings to ServiceMonitorConfig and applies it on both rendering paths:

  • the ClusterPolicy controller (applyServiceMonitorCustomEdits)
  • the DRA dcgm-exporter manifest template (manifests/state-dcgm-exporter/0600_service_monitor.yaml)

Because DCGMExporterServiceMonitorConfig is an alias of ServiceMonitorConfig, the field is also available to the operator-metrics ServiceMonitor.

The duplicated pointer-slice deref loop is factored into a shared derefRelabelConfigs helper so the two branches cannot drift.

Example:

dcgmExporter:
  serviceMonitor:
    metricRelabelings:
      - action: replace
        regex: (.+)
        sourceLabels: [exported_namespace]
        targetLabel: namespace

Supersedes #2939, which was closed. This revision additionally adds the rendered-manifest test coverage that review asked for on that PR, updates config/samples/nvidia_v1alpha1_gpucluster.yaml for consistency, and deduplicates the relabel-config conversion.

Checklist

  • No secrets, sensitive information, or unrelated changes
  • Lint checks passing (make lint)
  • Generated assets in-sync (make validate-generated-assets)
  • Go mod artifacts in-sync (make validate-modules)
  • Test cases are added for new code paths

Testing

  • make unit-test — all 21 packages pass.

  • make validate-generated-assets, make validate-modules, make validate-helm-values — pass.

  • make lint — only the 3 pre-existing SA4023 findings in cmd/nvidia-validator/main.go remain; none in the changed files.

  • Extended the ClusterPolicy-path TestServiceMonitor dcgm-exporter case to assert MetricRelabelConfigs.

  • Added TestDCGMExporterServiceMonitorRelabelings in internal/state, which renders the DRA ServiceMonitor template and asserts both relabelings and metricRelabelings land on the endpoint.

  • Verified end-to-end with helm template, confirming metricRelabelings reaches the rendered ClusterPolicy.

  • Verified end-to-end with helm upgrade, added metricRelabelings section in values.yaml, and Prometheus metrics query for a test pod:

Test Pod:

apiVersion: v1
kind: Pod
metadata:
  name: gpu-burn
spec:
  restartPolicy: Never
  containers:
  - name: gpu-burn
    image: nvcr.io/nvidia/k8s/cuda-sample:vectoradd-cuda12.5.0
    command: ["sh", "-c", "while true; do /cuda-samples/vectorAdd; done"]
    resources:
      limits:
        nvidia.com/gpu: 1

Before the fix:
Metrics Query:

nvidia@dsx-oss-workshop-02-gpu01:~$ curl -sG http://localhost:9090/api/v1/query \
  --data-urlencode 'query=DCGM_FI_DEV_POWER_USAGE{exported_namespace!=""}' \
  | jq '.data.result'
Handling connection for 9090
[
  {
    "metric": {
      "DCGM_FI_DRIVER_VERSION": "595.71.05",
      "UUID": "GPU-479e7af3-0688-ffec-0011-18ecf59165a8",
      "__name__": "DCGM_FI_DEV_POWER_USAGE",
      "container": "nvidia-dcgm-exporter",
      "device": "nvidia1",
      "endpoint": "gpu-metrics",
      "exported_container": "gpu-burn",
      "exported_namespace": "default",
      "exported_pod": "gpu-burn",
      "gpu": "1",
      "hostname": "dsx-oss-workshop-02-gpu01",
      "instance": "192.168.34.62:9400",
      "job": "nvidia-dcgm-exporter",
      "modelName": "NVIDIA H100 NVL",
      "namespace": "nvidia-gpu-operator",
      "pci_bus_id": "00000000:63:00.0",
      "pod": "nvidia-dcgm-exporter-h4l2c",
      "service": "nvidia-dcgm-exporter"
    },
    "value": [
      1790831487.752,
      "99.795"
    ]
  }
]

After the fix:

Helm metricrelabel-values.yaml

# Test payload for NVIDIA/gpu-operator#2938 (metricRelabelings).
#
# Do NOT apply this against the stock v26.7.1 chart — that CRD has no
# metricRelabelings field and the API server will silently prune it.
# Apply only after the patched ClusterPolicy CRD and the operator image built
# from the dsx-ws0930/issue-2938-metric-relabelings branch are in place.
#
# Only has a visible effect while honorLabels is false (the default) — with
# honorLabels=true Prometheus never creates the exported_* prefixes, so there
# is nothing to rewrite.
dcgmExporter:
  serviceMonitor:
    enabled: true
    honorLabels: false
    metricRelabelings:
      - action: replace
        regex: (.+)
        sourceLabels: [exported_namespace]
        targetLabel: namespace
      - action: replace
        regex: (.+)
        sourceLabels: [exported_pod]
        targetLabel: pod
      - action: replace
        regex: (.+)
        sourceLabels: [exported_container]
        targetLabel: container

Verify ClusterPolicy CRD and CR after Helm upgrade:

nvidia@dsx-oss-workshop-02-gpu01:~$ kubectl get crd clusterpolicies.nvidia.com -o json | jq -e '
  .spec.versions[0].schema.openAPIV3Schema.properties.spec.properties
  .dcgmExporter.properties.serviceMonitor.properties.metricRelabelings != null' \
  && echo "schema OK"
true
schema OK

nvidia@dsx-oss-workshop-02-gpu01:~$ kubectl get clusterpolicy cluster-policy \
  -o jsonpath='{.spec.dcgmExporter.serviceMonitor.metricRelabelings}' | jq .
[
  {
    "action": "replace",
    "regex": "(.+)",
    "sourceLabels": [
      "exported_namespace"
    ],
    "targetLabel": "namespace"
  },
  {
    "action": "replace",
    "regex": "(.+)",
    "sourceLabels": [
      "exported_pod"
    ],
    "targetLabel": "pod"
  },
  {
    "action": "replace",
    "regex": "(.+)",
    "sourceLabels": [
      "exported_container"
    ],
    "targetLabel": "container"
  }
]

Verify Servicemonitor:

nvidia@dsx-oss-workshop-02-gpu01:~$ kubectl -n nvidia-gpu-operator get servicemonitor nvidia-dcgm-exporter \
  -o jsonpath='{.spec.endpoints[0].metricRelabelings}' | jq .
[
  {
    "action": "replace",
    "regex": "(.+)",
    "sourceLabels": [
      "exported_namespace"
    ],
    "targetLabel": "namespace"
  },
  {
    "action": "replace",
    "regex": "(.+)",
    "sourceLabels": [
      "exported_pod"
    ],
    "targetLabel": "pod"
  },
  {
    "action": "replace",
    "regex": "(.+)",
    "sourceLabels": [
      "exported_container"
    ],
    "targetLabel": "container"
  }
]

Metrics Query:

nvidia@dsx-oss-workshop-02-gpu01:~$ curl -sG http://localhost:9090/api/v1/query \
  --data-urlencode 'query=DCGM_FI_DEV_POWER_USAGE' | jq '.data.result[].metric'
Handling connection for 9090
{
  "DCGM_FI_DRIVER_VERSION": "595.71.05",
  "UUID": "GPU-479e7af3-0688-ffec-0011-18ecf59165a8",
  "__name__": "DCGM_FI_DEV_POWER_USAGE",
  "container": "nvidia-dcgm-exporter",
  "device": "nvidia1",
  "endpoint": "gpu-metrics",
  "gpu": "1",
  "hostname": "dsx-oss-workshop-02-gpu01",
  "instance": "192.168.34.62:9400",
  "job": "nvidia-dcgm-exporter",
  "modelName": "NVIDIA H100 NVL",
  "namespace": "nvidia-gpu-operator",
  "pci_bus_id": "00000000:63:00.0",
  "pod": "nvidia-dcgm-exporter-h4l2c",
  "service": "nvidia-dcgm-exporter"
}
{
  "DCGM_FI_DRIVER_VERSION": "595.71.05",
  "UUID": "GPU-b4b175d2-91b1-9082-4274-9aaa8822e0e3",
  "__name__": "DCGM_FI_DEV_POWER_USAGE",
  "container": "gpu-burn",
  "device": "nvidia0",
  "endpoint": "gpu-metrics",
  "exported_container": "gpu-burn",
  "exported_namespace": "default",
  "exported_pod": "gpu-burn",
  "gpu": "0",
  "hostname": "dsx-oss-workshop-02-gpu01",
  "instance": "192.168.34.62:9400",
  "job": "nvidia-dcgm-exporter",
  "modelName": "NVIDIA H100 NVL",
  "namespace": "default",
  "pci_bus_id": "00000000:17:00.0",
  "pod": "gpu-burn",
  "service": "nvidia-dcgm-exporter"
}

nvidia@dsx-oss-workshop-02-gpu01:~$ curl -sG http://localhost:9090/api/v1/query   --data-urlencode 'query=DCGM_FI_DEV_POWER_USAGE{pod="gpu-burn"}'   | jq '.data.result'
Handling connection for 9090
[
  {
    "metric": {
      "DCGM_FI_DRIVER_VERSION": "595.71.05",
      "UUID": "GPU-b4b175d2-91b1-9082-4274-9aaa8822e0e3",
      "__name__": "DCGM_FI_DEV_POWER_USAGE",
      "container": "gpu-burn",
      "device": "nvidia0",
      "endpoint": "gpu-metrics",
      "exported_container": "gpu-burn",
      "exported_namespace": "default",
      "exported_pod": "gpu-burn",
      "gpu": "0",
      "hostname": "dsx-oss-workshop-02-gpu01",
      "instance": "192.168.34.62:9400",
      "job": "nvidia-dcgm-exporter",
      "modelName": "NVIDIA H100 NVL",
      "namespace": "default",
      "pci_bus_id": "00000000:17:00.0",
      "pod": "gpu-burn",
      "service": "nvidia-dcgm-exporter"
    },
    "value": [
      1790833827.139,
      "99.504"
    ]
  }
]

nvidia@dsx-oss-workshop-02-gpu01:~$ curl -sG http://localhost:9090/api/v1/query \
  --data-urlencode 'query=DCGM_FI_DEV_POWER_USAGE{namespace!="nvidia-gpu-operator"}' \
  | jq '.data.result[].metric | {namespace, pod, container}'
Handling connection for 9090
{
  "namespace": "default",
  "pod": "gpu-burn",
  "container": "gpu-burn"
}

Fixes #2938

@gseidlerhpe
gseidlerhpe requested a review from a team as a code owner September 30, 2026 18:38
@copy-pr-bot

copy-pr-bot Bot commented Sep 30, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

📝 Walkthrough

Walkthrough

The ServiceMonitor configuration now supports optional metric relabeling rules. The controller converts configured rules into endpoint metric relabel configs, and the DCGM Exporter ServiceMonitor template conditionally renders them. Sample and Helm configuration include empty defaults and commented examples. Tests check metric relabeling in the resulting ServiceMonitor endpoint.

Priority: ➖ Normal

Merge Risk: 🔵 Low · up to 1f3ac

The new metric relabeling configuration is supported by the generated assets and both rendering paths. Mergeable with bounded follow-up to strengthen the rendering assertions against incomplete rules.

  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
internal/state/dcgm_exporter_test.go (1)

260-260: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick win

Assert the complete rendered metric relabeling rule.

The assertion passes if rendering drops sourceLabels, changes regex, or changes action while preserving targetLabel. Apply the same complete-rule assertion to the target relabeling rule on line 255. This test must detect any change that prevents the configured label rewrite.


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: NVIDIA/gpu-operator/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 40bb04ba-c932-4911-be38-e8acee5df20b

📥 Commits

Reviewing files that changed from the base of the PR and between 75210df and 1f3acb0.

⛔ Files ignored due to path filters (7)
  • api/nvidia/v1/zz_generated.deepcopy.go is excluded by !**/zz_generated.*.go
  • bundle/manifests/nvidia.com_clusterpolicies.yaml is excluded by !bundle/manifests/nvidia.com_*.yaml
  • bundle/manifests/nvidia.com_gpuclusters.yaml is excluded by !bundle/manifests/nvidia.com_*.yaml
  • config/crd/bases/nvidia.com_clusterpolicies.yaml is excluded by !config/crd/bases/**
  • config/crd/bases/nvidia.com_gpuclusters.yaml is excluded by !config/crd/bases/**
  • deployments/gpu-operator/crds/nvidia.com_clusterpolicies.yaml is excluded by !deployments/gpu-operator/crds/**
  • deployments/gpu-operator/crds/nvidia.com_gpuclusters.yaml is excluded by !deployments/gpu-operator/crds/**
📒 Files selected for processing (7)
  • api/nvidia/v1/clusterpolicy_types.go
  • config/samples/nvidia_v1alpha1_gpucluster.yaml
  • controllers/object_controls.go
  • controllers/object_controls_test.go
  • deployments/gpu-operator/values.yaml
  • internal/state/dcgm_exporter_test.go
  • manifests/state-dcgm-exporter/0600_service_monitor.yaml

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no bugs or issues to report.

Devin Review

ServiceMonitorConfig only exposed relabelings, which maps to Prometheus
target relabeling (Endpoint.RelabelConfigs) and runs before the scrape.
Rewriting labels on scraped samples, e.g. collapsing exported_namespace
back onto namespace, requires Endpoint.MetricRelabelConfigs and had no
API surface, so the Helm value was rejected by CRD validation.

Add MetricRelabelings to ServiceMonitorConfig and apply it on both
rendering paths: the ClusterPolicy controller and the DRA dcgm-exporter
manifest template. The DCGMExporterServiceMonitorConfig alias means the
field is available to the operator-metrics ServiceMonitor as well.

Fixes NVIDIA#2938

Signed-off-by: Gernot Seidler <gernot.seidler@hpe.com>
@gseidlerhpe
gseidlerhpe force-pushed the dsx-ws0930/issue-2938-metric-relabelings branch from 1f3acb0 to 1c179f1 Compare October 1, 2026 06:03
@rahulait

rahulait commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

/ok to test 1c179f1

@tariq1890

Copy link
Copy Markdown
Contributor

Thanks for your contribution Gernot @gseidlerhpe !

@rahulait
rahulait merged commit 3e1873a into NVIDIA:main Oct 1, 2026
20 of 21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature]: add metricRelabelings field to the dcgm exporter service monitor.

3 participants