Report which GPUCluster operands are blocking readiness - #2988
Conversation
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: NVIDIA/gpu-operator/.coderabbit.yaml Review profile: QUIET Plan: Enterprise Run ID: 📒 Files selected for processing (2)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review. 📝 WalkthroughWalkthroughWhen operand states are not ready, the GPUCluster controller includes the names of states with Priority: ⬇️ Low Severity of issue fixed: Low Merge Risk: ⚪ Minimal · up to The GPUCluster NotReady condition will now name the operands that are blocking readiness. This is a small, well-tested change and is low risk to merge.
Comment |
|
/ok to test d2c5f9e |
cdesiniotis
left a comment
There was a problem hiding this comment.
@abonillabeeche thank you for the contribution!
|
@abonillabeeche can you please make sure your commit is verified? You can add a GPG key to github and sign your commits using it. |
When no state returns an error but some operands aren't ready yet, the GPUCluster OperandNotReady condition only said "Waiting for operand pods to be ready", so you had to go digging through pods to find out what it was waiting on. List the not-ready states in the message, same format ClusterPolicy already uses: Waiting for operand pods to be ready; states not ready: [state-dcgm-exporter] Ignored states (e.g. DCGM when it's disabled) are left out. Fixes NVIDIA#2978 Signed-off-by: Alejandro Bonilla <abonilla@suse.com>
d2c5f9e to
087d5fb
Compare
|
Force pushed to have the commit signed and also rebased the branch. |
|
/ok to test 087d5fb |
Fixes #2978
When nothing errors out but some operands just aren't ready yet, the GPUCluster
OperandNotReadycondition only said "Waiting for operand pods to be ready", so you had to go poke at pods to figure out what it was actually waiting on. This lists the not-ready states in the message, same format ClusterPolicy already uses (clusterPolicyNotReadyMessage). States that returnignore(e.g. DCGM when it's disabled) are left out.Changes
gpuClusterOperandNotReadyMessage()helper incontrollers/gpucluster_controller.go, used when setting theOperandNotReadyconditionTesting
go test ./controllers/passes, andgolangci-lintis clean.Also tried it on a couple of real clusters with a locally built operator image.
k3s (Rancher Desktop, no GPU, node labeled so the DaemonSets schedule): DCGM is disabled, so it's correctly left out of the list:
RKE2 v1.36, single node with a T400: I switched from ClusterPolicy to GPUCluster and watched the condition as things came up:
The driver, the DRA plugin and the validator all came up and dropped off the list. dcgm-exporter stayed stuck, and the message pointed right at it. Its logs show it trying to list
v1beta1.ResourceSlice, which that API server doesn't serve:That's unrelated to this PR (happy to open a separate issue), but it's a nice example of why the extra detail helps.