Skip to content

gather GPUCluster and DRA resources in must-gather.sh - #2985

Merged
rahulait merged 1 commit into
NVIDIA:mainfrom
raunakkumar:issue-2976-must-gather-gpucluster-dra
Oct 1, 2026
Merged

rahulait merged 1 commit into
NVIDIA:mainfrom
raunakkumar:issue-2976-must-gather-gpucluster-dra

Conversation

@raunakkumar

@raunakkumar raunakkumar commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #2976

hack/must-gather.sh had no awareness of the DRA-based stack added in 26.7.0. This adds three collection blocks after the NVIDIADriver section, modeled on the existing ClusterPolicy/NVIDIADriver handling as mentioned in the issue:

  • GPUCluster
  • DeviceClass
  • ResourceSlice (driver gpu.nvidia.com)
    All three are guarded with --ignore-not-found so the script still completes on clusters without DRA it's backward compatible.

Checklist

  • No secrets, sensitive information, or unrelated changes
  • Lint checks passing (make lint)
  • Generated assets in-sync (make validate-generated-assets) — N/A (no API changes)
  • Go mod artifacts in-sync (make validate-modules) — N/A (no dependency changes)
  • Test cases are added for new code paths — N/A (no test harness; validated on a live cluster)

Testing

Ran against a live GPU Operator 26.7.0 cluster:

  • Empty case (CRD present, DRA not deployed): all three sections run and report "not found" cleanly; script continues without error.
  • Populated case: created sample DeviceClass and ResourceSlice (spec.driver: gpu.nvidia.com) objects; confirmed they're gathered into deviceclasses.yaml/resourceslices.yaml and the field-selector filters correctly.
    Populated objects were representative samples (resource.k8s.io/v1), not produced by a running DRA driver — the test cluster had the GPUCluster CRD but DRA not fully deployed.

@copy-pr-bot

copy-pr-bot Bot commented Sep 30, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@raunakkumar raunakkumar changed the title hack: gather GPUCluster and DRA resources in must-gather.sh gather GPUCluster and DRA resources in must-gather.sh Sep 30, 2026
@raunakkumar
raunakkumar marked this pull request as ready for review September 30, 2026 17:23
@raunakkumar
raunakkumar requested a review from a team as a code owner September 30, 2026 17:23
@coderabbitai

coderabbitai Bot commented Sep 30, 2026

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: NVIDIA/gpu-operator/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: bcb4713a-aec4-412a-a1fe-eaf8a8a2cef1

📥 Commits

Reviewing files that changed from the base of the PR and between 75210df and 116feb1.

📒 Files selected for processing (1)
  • hack/must-gather.sh

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


📝 Walkthrough

Walkthrough

The must-gather script now collects GPUCluster, DeviceClass, and ResourceSlice resources. It saves matching YAML to resource-specific files and reports when no matching resources are found.

Priority: ⬇️ Low

Merge Risk: ⚪ Minimal · up to 116fe

The added resource collection has no established merge-blocking issue. Missing resource APIs do not terminate the script under its current startup settings.

  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 1 potential issue.

Devin Review

Comment thread hack/must-gather.sh
@rahulait

Copy link
Copy Markdown
Contributor

Thanks @raunakkumar for the contribution, taking a look.

@tariq1890

Copy link
Copy Markdown
Contributor

/ok to test 116feb1

Comment thread hack/must-gather.sh
@raunakkumar

raunakkumar commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor Author

I was able to test this on the launchpad instance dsx-oss-workshop-01-gpu01 as well.

Two cases captured

  1. Backward-compatible (no DRA driver) baseline-console.log
    All 3 new sections report "not found" cleanly, no artifact files created, script reaches "All done!".

  2. DRA enabled (installed nvidia-dra-driver-gpu v25.8.1) console.log

  • deviceclasses.yaml— 4 real DeviceClasses (gpu.nvidia.com, mig.nvidia.com, 2× compute-domain)
  • resourceslices.yaml— real ResourceSlice with 2 H100 NVL GPUs, driver-produced attributes, --field-selector spec.driver=gpu.nvidia.com filter confirmed working
  • Script reaches "All done!".

Edit: I was able to get the results with the last commit with xompute-domain.nvidia.com
Uploading resourceslices-compute-domain.nvidia.com.yaml…

@rahulait

Copy link
Copy Markdown
Contributor

Thanks @raunakkumar, LGTM. Can you squash your commits into one and I can run the CI and have it merged.

@raunakkumar
raunakkumar force-pushed the issue-2976-must-gather-gpucluster-dra branch from 2c68969 to 3c28afb Compare September 30, 2026 18:27
@rahulait

Copy link
Copy Markdown
Contributor

/ok to test 3c28afb

@tariq1890

Copy link
Copy Markdown
Contributor

Thank you @raunakkumar ! Changes LGTM. Once your commit shows up as Verified, this should be good to merge.

@raunakkumar

Copy link
Copy Markdown
Contributor Author

Thank you @raunakkumar ! Changes LGTM. Once your commit shows up as Verified, this should be good to merge.

Thanks. I was just checking the CI failure on build-gpu-operator-amd64 seems to be unrelated to the changes i made in the PR. CUDA yum mirror returned 404 on repodata (developer.download.nvidia.com). How do i go about this. Do a re-run or a known issue?

@kvalliyurnatt

Copy link
Copy Markdown
Contributor

/ok to test 3c28afb

@kvalliyurnatt

Copy link
Copy Markdown
Contributor

@raunakkumar Re run passed, could you sign off your commit with a GPG key

must-gather did not collect any information about the GPUCluster CRD
(introduced in 26.7.0 for the NVIDIA DRA Driver) or the DRA objects it
relies on. Add collection of:

- the GPUCluster CR (spec and status)
- DeviceClass resources
- ResourceSlice resources for the NVIDIA DRA drivers, gathered per driver
  (gpu.nvidia.com and compute-domain.nvidia.com, the latter backing IMEX /
  compute domains) into resourceslices-<driver>.yaml files

Each section is guarded with --ignore-not-found and a "not found" message so
gather still succeeds on clusters without DRA.

Fixes NVIDIA#2976

Signed-off-by: raunakkumar <raunak.kumar@hpe.com>
@raunakkumar
raunakkumar force-pushed the issue-2976-must-gather-gpucluster-dra branch from 3c28afb to f6d32ba Compare September 30, 2026 23:16
@raunakkumar

raunakkumar commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor Author

/ok to test f6d32ba

1 similar comment
@kvalliyurnatt

Copy link
Copy Markdown
Contributor

/ok to test f6d32ba

@rahulait
rahulait merged commit ea22c37 into NVIDIA:main Oct 1, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Gather GPUCluster and DRA resources in must-gather.sh

4 participants