gather GPUCluster and DRA resources in must-gather.sh - #2985
Conversation
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: NVIDIA/gpu-operator/.coderabbit.yaml Review profile: QUIET Plan: Enterprise Run ID: 📒 Files selected for processing (1)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review. 📝 WalkthroughWalkthroughThe must-gather script now collects GPUCluster, DeviceClass, and ResourceSlice resources. It saves matching YAML to resource-specific files and reports when no matching resources are found. Priority: ⬇️ Low Merge Risk: ⚪ Minimal · up to The added resource collection has no established merge-blocking issue. Missing resource APIs do not terminate the script under its current startup settings.
Comment |
|
Thanks @raunakkumar for the contribution, taking a look. |
|
/ok to test 116feb1 |
|
I was able to test this on the launchpad instance Two cases captured
Edit: I was able to get the results with the last commit with |
|
Thanks @raunakkumar, LGTM. Can you squash your commits into one and I can run the CI and have it merged. |
2c68969 to
3c28afb
Compare
|
/ok to test 3c28afb |
|
Thank you @raunakkumar ! Changes LGTM. Once your commit shows up as |
Thanks. I was just checking the CI failure on |
|
/ok to test 3c28afb |
|
@raunakkumar Re run passed, could you sign off your commit with a GPG key |
must-gather did not collect any information about the GPUCluster CRD (introduced in 26.7.0 for the NVIDIA DRA Driver) or the DRA objects it relies on. Add collection of: - the GPUCluster CR (spec and status) - DeviceClass resources - ResourceSlice resources for the NVIDIA DRA drivers, gathered per driver (gpu.nvidia.com and compute-domain.nvidia.com, the latter backing IMEX / compute domains) into resourceslices-<driver>.yaml files Each section is guarded with --ignore-not-found and a "not found" message so gather still succeeds on clusters without DRA. Fixes NVIDIA#2976 Signed-off-by: raunakkumar <raunak.kumar@hpe.com>
3c28afb to
f6d32ba
Compare
|
/ok to test f6d32ba |
1 similar comment
|
/ok to test f6d32ba |
Fixes #2976
hack/must-gather.sh had no awareness of the DRA-based stack added in
26.7.0. This adds three collection blocks after the NVIDIADriver section, modeled on the existing ClusterPolicy/NVIDIADriver handling as mentioned in the issue:All three are guarded with
--ignore-not-foundso the script still completes on clusters without DRA it's backward compatible.Checklist
make lint)make validate-generated-assets) — N/A (no API changes)make validate-modules) — N/A (no dependency changes)Testing
Ran against a live GPU Operator 26.7.0 cluster:
Empty case(CRD present, DRA not deployed): all three sections run and report "not found" cleanly; script continues without error.Populated case: created sample DeviceClass and ResourceSlice (spec.driver: gpu.nvidia.com) objects; confirmed they're gathered into deviceclasses.yaml/resourceslices.yaml and the field-selector filters correctly.Populated objects were representative samples (resource.k8s.io/v1), not produced by a running DRA driver — the test cluster had the GPUCluster CRD but DRA not fully deployed.