Skip to content

Report which GPUCluster operands are blocking readiness - #2988

Merged
rahulait merged 1 commit into
NVIDIA:mainfrom
abonillabeeche:gpucluster-operand-not-ready-names
Oct 1, 2026
Merged

rahulait merged 1 commit into
NVIDIA:mainfrom
abonillabeeche:gpucluster-operand-not-ready-names

Conversation

@abonillabeeche

Copy link
Copy Markdown
Contributor

Fixes #2978

When nothing errors out but some operands just aren't ready yet, the GPUCluster OperandNotReady condition only said "Waiting for operand pods to be ready", so you had to go poke at pods to figure out what it was actually waiting on. This lists the not-ready states in the message, same format ClusterPolicy already uses (clusterPolicyNotReadyMessage). States that return ignore (e.g. DCGM when it's disabled) are left out.

Changes

  • New gpuClusterOperandNotReadyMessage() helper in controllers/gpucluster_controller.go, used when setting the OperandNotReady condition
  • Unit tests: a table test for the helper, plus a reconcile test that checks the condition reason and message end to end with the fake state manager / condition updater

Testing

go test ./controllers/ passes, and golangci-lint is clean.

Also tried it on a couple of real clusters with a locally built operator image.

k3s (Rancher Desktop, no GPU, node labeled so the DaemonSets schedule): DCGM is disabled, so it's correctly left out of the list:

Error  True  OperandNotReady  Waiting for operand pods to be ready; states not ready: [state-dra-driver state-dcgm-exporter state-dra-validation]

RKE2 v1.36, single node with a T400: I switched from ClusterPolicy to GPUCluster and watched the condition as things came up:

18:08:28  notReady | Error/PrerequisiteNotMet: A ClusterPolicy CR "cluster-policy" exists; a ClusterPolicy CR and GPUCluster CR may not exist at the same time
18:09:02  notReady | Error/OperandNotReady: Waiting for operand pods to be ready; states not ready: [state-dra-driver state-dcgm-exporter state-dra-validation]
18:11:10  notReady | Error/OperandNotReady: Waiting for operand pods to be ready; states not ready: [state-dcgm-exporter]

The driver, the DRA plugin and the validator all came up and dropped off the list. dcgm-exporter stayed stuck, and the message pointed right at it. Its logs show it trying to list v1beta1.ResourceSlice, which that API server doesn't serve:

"Failed to watch" err="failed to list *v1beta1.ResourceSlice: the server could not find the requested resource"

That's unrelated to this PR (happy to open a separate issue), but it's a nice example of why the extra detail helps.

@abonillabeeche
abonillabeeche requested a review from a team as a code owner September 30, 2026 19:16
@copy-pr-bot

copy-pr-bot Bot commented Sep 30, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@coderabbitai

coderabbitai Bot commented Sep 30, 2026

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: NVIDIA/gpu-operator/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 993b226d-ba9a-410b-928e-82068f8438aa

📥 Commits

Reviewing files that changed from the base of the PR and between 75210df and d2c5f9e.

📒 Files selected for processing (2)
  • controllers/gpucluster_controller.go
  • controllers/gpucluster_controller_test.go

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


📝 Walkthrough

Walkthrough

When operand states are not ready, the GPUCluster controller includes the names of states with NotReady or Error status in the OperandNotReady condition message. If no states match, the message remains unchanged. Tests cover message content and the condition set during reconciliation.

Priority: ⬇️ Low

Severity of issue fixed: Low

Merge Risk: ⚪ Minimal · up to d2c5f

The GPUCluster NotReady condition will now name the operands that are blocking readiness. This is a small, well-tested change and is low risk to merge.

  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no bugs or issues to report.

Devin Review

@cdesiniotis

Copy link
Copy Markdown
Contributor

/ok to test d2c5f9e

@cdesiniotis cdesiniotis left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@abonillabeeche thank you for the contribution!

@rahulait

rahulait commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

@abonillabeeche can you please make sure your commit is verified? You can add a GPG key to github and sign your commits using it.
https://docs.github.com/en/authentication/managing-commit-signature-verification

When no state returns an error but some operands aren't ready yet, the
GPUCluster OperandNotReady condition only said "Waiting for operand pods
to be ready", so you had to go digging through pods to find out what it
was waiting on. List the not-ready states in the message, same format
ClusterPolicy already uses:

  Waiting for operand pods to be ready; states not ready: [state-dcgm-exporter]

Ignored states (e.g. DCGM when it's disabled) are left out.

Fixes NVIDIA#2978

Signed-off-by: Alejandro Bonilla <abonilla@suse.com>
@rahulait
rahulait force-pushed the gpucluster-operand-not-ready-names branch from d2c5f9e to 087d5fb Compare October 1, 2026 18:24
@rahulait

rahulait commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Force pushed to have the commit signed and also rebased the branch.

@rahulait

rahulait commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

/ok to test 087d5fb

@rahulait
rahulait enabled auto-merge October 1, 2026 18:26
@rahulait
rahulait merged commit d5f4924 into NVIDIA:main Oct 1, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Report which GPUCluster operands are blocking readiness

4 participants