cuda.bindings: support multiple CTK release lines on main - #2737
Conversation
Moon migration belongs in PR NVIDIA#2659. Restore the pre-Moon selective-CI planner and workflows from 727ef59.
|
/ok to test b87d0a1 |
|
There was a problem hiding this comment.
I started to comment on some individual things, but then decided to stop because I think there is a more fundamental change that needs to be made across this whole PR (and then I'm happy to come back and review further).
/Today/ the "current" version is 13, and the backport version is 12. But at some point in the future that will switch to 14 and 13. This pervasively hardcodes those version numbers all over this codebase, especially in CI, but in a bunch of the release scripts as well, and even the cuda_bindings_12 directory name as indicators of current vs. backport.
Instead, we should use the config we already have in versions.yml and use that to drive the numbers everywhere. That way when it's time to move on, all that should be required is updating versions.yml, and copying/overwriting the existing cuda_bindings to cuda_bindings_backport (or whatever we want to call it), and move on. I'm sure there are many details I'm missing, but that should be the goal and design -- it would be preferable to reduce it to as close to that as possible. The problem with this as-is is that there are hundreds of context-sensitive places that would need to be updated to do that update -- we are creating a massive pile of technical debt to pay later. I'm sure an agent might get that X% correct, but I always think it's better to engineer for flexibility, especially for something we know will happen. If versions.yml (which requires using yq to parse etc.) makes this too difficult, we could explore a simple VARIABLE=value format which would parse as both bash variables and Python variables and probably be more convenient to use from the many places it is needed. There are really only two actual values in versions.yml today, so that should be fine.
I'm also a little concerned (without any testing-based evidence) that this will break when we tag the same commit with v13.x.y and v12.x.y, which will be the common case, in fact, IMHO, one of the real benefits of moving to this approach. We should get an agent to do a thorough investigation of that use case and make sure it is covered. Ideally, it would be nice for a single release run to do both releases simultaneously but it's not a deal breaker if it still requires kicking off two runs.
Also what is this (from the agent's PR description):
The later NVML memoryview fix is reproduced byte-for-byte from cybind commit
6def52ca508c9e14ef67f4ce26a0c677f3fbad72 with Doxygen 1.17.0:
If there is something like this that wasn't backported, let's deal with that separately so it's not an unrelated tag-along to this PR.
Also a note for future agent reviewers of this PR: The interesting part of this PR is the part outside of the cuda_bindings_backport or cuda_bindings_12 directory. Those are just direct copies from the 12.9.x branch, and any differences between that and the cuda_bindings directory are likely intentional. When reviewing, focus on the scaffolding / CI / overall structure.
|
Archiving options related to a lychee chicken-and-egg issue. I'll go with Option 1 below. This comment is to explain why. codex: We have three sensible options. For PR 2737, I recommend keeping the canonical links unchanged and treating these as documented pre-merge exceptions.
Run lychee once with only these three URLs excluded, record that every other link passes, and rerun without exclusions after merge. This is reasonable because authored-source lychee is explicitly skipped by the GitHub CI job at .github/workflows/ci.yml, so these are not merge-gating failures. It avoids landing temporary configuration or compromising the final URLs.
Add three exact anchored patterns so This is practical, but creates a mandatory cleanup PR and briefly leaves three blind spots on
Teach the hook to map: Then links to newly introduced files are validated against the checkout before they exist online. This is exactly the future-URL use case for lychee’s remapping feature. Lychee remapping documentation It is the principled reusable solution, but needs a portable wrapper to calculate the absolute worktree path. I would pursue it separately only if this problem starts recurring. I would avoid:
So my recommendation is option 1: preserve the three correct final URLs, validate everything else with exact one-off exclusions, and rerun lychee from fresh |
|
/ok to test f4ddc4e |
|
/ok to test c6a0cf1 |
|
/ok to test 83c1cf0 |
mdboom
left a comment
There was a problem hiding this comment.
This PR is really challenging to review. Even ignoring the files that are just copied from the 12.9.x branch (which don't need review) there is a lot here.
I found a bunch of sources of unnecessary complexity and stopped reading after that, so still haven't done a full human pass. Even with agents, multiple stages of transformations between data formats causes multiple places that bugs can creep in. It makes it harder for agents or humans to understand the fundamental logic of what branches are covered and how this all works. I think we need a step back analysis of how things should be represented based on how things are needed downstream of that and adjust accordingly.
I know partly what is driving this complexity is GHA's design that forces things into small snippets. @kkraus14 has suggested elsewhere that maybe moving to a more formal mono-repo management tool like moon may be better than building out more and more CI complexity. Maybe an agent could build a prototype quickly to at least see whether it meets our use cases and what sort of complexity it requires so we can compare.
Additionally, there are many new scripts here in ci/tools with largely overlapping functionality where logic is spread between them and bash scripts and they sort of go back-and-forth. Moving more logic into fewer Python scripts, that output directly to what is most often needed (POSIX environment variables) would probably be preferable to the current state. It should be possible to see in one file how the variables that control the rest of the execution are computed. I'm thinking particularly of bindings_config.py and resolve_release_bindings_line.py -- why are the separate? And there is probably some value in combining it with compute_ci_plan.py in some way. Not necessarily that they need to be in the same source file, but that they would interact in the same process. Sorry to not have /concrete/ suggestions for that, but I am trying to suggest ways that this could be simplified to be more easily reviewable and maintained going forward.
If there is any way to break this up into multiple steps that could be reviewed independently, that would help a lot. Even when I used an agent to review, it struggled to understand the "why" of many of these changes. I asked my agents to offer some suggestions about making this easier to review. I think it's mostly ok (except I wouldn't consider the cuda_bindings_12/ import step as big -- I'm comfortable rubber stamping a direct copy from another branch). It doesn't really have good suggestions for incremental development -- most of what it suggests are just bugfixes.
A few concrete ways to cut the size and the duplicated-logic risk this PR introduces:
-
Split into sequential PRs. The
cuda_bindings_12/import, the registry (ci/versions.yml+bindings_config.py), and the workflow rewiring are three logically separable changes. Landing the registry + validator first (small, reviewable, testable in isolation against the existing single-line setup) then rewiring workflows to consume it, then importingcuda_bindings_12/last, would let each step get real scrutiny instead of one 255K-line PR where reviewers rubber-stamp the bulk. -
Stop tracking the tag family in two places.
ci/versions.yml'stag_seriesandcuda_bindings/pyproject.toml'stag_regex/git_describe_commandencode the same CUDA-major fact independently (flagged in the review — they can drift and did in thev13.4.0scenario). Either generate the pyprojecttag_regexfrom the registry at build time, or drop the registry'stag_seriesfield and derive it from the pyproject regex instead. One source of truth removes a whole class of the findings above. -
Don't model generality you don't use yet.
roles.maintenanceis schema'd as a list, butbuild-wheel.ymlhard-errors unless it has exactly one entry, and nothing in this PR needs more than one maintenance line. Collapsing the schema to a singlemaintenanceline (not a list) until a second one is actually needed removes validation code, removes a whole "what if maintenance has 2 entries" test surface, and can be widened later when there's a real second line to design against. -
Trim the transitional compatibility code in
compute_ci_plan.py. Thevariantsdict with the "OR aggregation... consumers migrate tolines" comment is scaffolding for old CUDA-major-keyed workflow consumers. If those consumers are being rewritten in this same PR anyway, migrating them straight to line-keyed output and deleting the compatibility shim removes a chunk of logic (and a source of the empty-matrix / dual-bookkeeping risk flagged earlier).
| - name: Checkout docs control plane | ||
| uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0 | ||
| with: | ||
| ref: ${{ github.sha }} | ||
| path: .ci-control | ||
|
|
||
| - name: Install CI tool dependencies | ||
| run: python3 -m pip install -r .ci-control/ci/tools/requirements.txt | ||
|
|
There was a problem hiding this comment.
This is treating ci/tools as a sort of poor version of a package, all just so it can use PyYAML.
The modern way to handle dependencies of standalone scripts is to use PEP 723 metadata and then use a PEP 723-supporting tool like uv, pixi or hatch to run it. I think that would be way less cumbersome than this. Or we go all in and make it a proper package which might have other benefits given how big it's getting. But this approach is sort of worst-of-both-worlds, IMHO.
| set -euo pipefail | ||
| if [[ -n "$BINDINGS_LINE" ]]; then | ||
| bindings_line="$BINDINGS_LINE" | ||
| elif [[ "${IS_RELEASE}" == "true" && "$RELEASE_TAG" == v* ]]; then | ||
| bindings_line=$(python3 .ci-control/ci/tools/resolve_release_bindings_line.py \ | ||
| --release-tag "$RELEASE_TAG" \ | ||
| --release-source-root . \ | ||
| --control-config .ci-control/ci/versions.yml) | ||
| else | ||
| echo "error: cannot find ci/versions.yml or ci/versions.json" >&2 | ||
| exit 1 | ||
| bindings_line=$(python3 .ci-control/ci/tools/bindings_config.py get --role current) | ||
| fi | ||
| BUILD_CTK_VER=$(jq -er '.toolkit_version' <<< "$bindings_line") | ||
| BINDINGS_COMPONENT_DIR=$(jq -er '.release_source_dir // .source_dir' <<< "$bindings_line") | ||
| BINDINGS_REGISTRY_ORIGIN=$(jq -er '.release_registry_origin // "tag"' <<< "$bindings_line") | ||
| if [[ ! "${BUILD_CTK_VER}" =~ ^[0-9]+\.[0-9]+\.[0-9]+$ ]]; then | ||
| echo "error: derived CTK build version ${BUILD_CTK_VER} does not match MAJOR.MINOR.MICRO" >&2 | ||
| exit 1 | ||
| fi | ||
| if [[ ! -d "$BINDINGS_COMPONENT_DIR" ]]; then | ||
| echo "error: resolved bindings source directory does not exist: $BINDINGS_COMPONENT_DIR" >&2 | ||
| exit 1 | ||
| fi | ||
| echo "BUILD_CTK_VER=${BUILD_CTK_VER}" >> "$GITHUB_ENV" | ||
| echo "BINDINGS_COMPONENT_DIR=${BINDINGS_COMPONENT_DIR}" >> "$GITHUB_ENV" | ||
| echo "BINDINGS_REGISTRY_ORIGIN=${BINDINGS_REGISTRY_ORIGIN}" >> "$GITHUB_ENV" |
There was a problem hiding this comment.
AFAICT, the resolve_release_bindings_line.py or bindings_config.py scripts read in the versions.yaml metadata and output json, which is then parsed with jq to convert to environment variables. And there is additional validation of the values coming out of scripts that we control. Can't we skip that middle steps and have the scripts output what is needed (envvar pairs)? The long standing env-vars script does that, for example. I'll probably need to read further to discover where using JSON as an intermediary might be relevant, though. My concern isn't efficiency, it's that there are so many transformation steps and therefore places for bugs to slip in.
I think fixing this will require stepping back and understanding all of the places these JSON-emitting scripts are used and providing output in the most convenient way possible for those contexts.
| tag_regex = "^(?P<version>v13\\.\\d+\\.\\d+(?:[ab]\\d+)?(?:\\.post\\d+)?)" | ||
| git_describe_command = ["git", "describe", "--dirty", "--tags", "--long", "--match", "v13.*"] |
There was a problem hiding this comment.
Updating the metadata in versions.yml will not affect this. How do we ensure they stay in sync?
| _TOOLKIT_VERSION_PATTERN = re.compile(r"[1-9][0-9]*\.[0-9]+\.[0-9]+(?:[.-][A-Za-z0-9]+)*") | ||
| _TAG_SERIES_PATTERN = re.compile(r"v[1-9][0-9]*(?:\.[0-9]+)*\.") | ||
| _FINAL_TAG_SUFFIX_PATTERN = re.compile(r"[0-9]+(?:\.post[0-9]+)?") | ||
| _ALPHA_BETA_TAG_SUFFIX_PATTERN = re.compile(r"[0-9]+(?:[ab][0-9]+)?(?:\.post[0-9]+)?") |
There was a problem hiding this comment.
This should be PEP 440 compliant so it can support anything we might want to do on PyPI, and I don't think it is. It would be better to use a library for this than reinventing here.
| build_bindings_current=$(jq -r --arg id "$current_line_id" '.modules.bindings.lines[$id].needs_build | if type == "boolean" then . else error("invalid current bindings build gate") end' <<< "$WORKPLAN") | ||
| build_bindings_maintenance=$(jq -r --arg id "$maintenance_line_id" '.modules.bindings.lines[$id].needs_build | if type == "boolean" then . else error("invalid maintenance bindings build gate") end' <<< "$WORKPLAN") |
There was a problem hiding this comment.
This got me thinking about whether the current and maintenance lines are built and tested every time, even if only one or the other changed? Part of Keith's recent work was to limit the amount that is rebuilt and tested every time, and it would be nice not to step back from that.
I had my agent investigate this and it sees that is more-or-less the case.
| for relative in shared_paths: | ||
| candidates = [(root, repo_root / root / relative) for root in roots] | ||
| symlinks = [root for root, path in candidates if path.is_symlink()] | ||
| if symlinks: | ||
| violations.append(f"{relative}: symlink in {', '.join(symlinks)}") | ||
| continue | ||
| missing = [root for root, path in candidates if not path.is_file()] | ||
| if missing: | ||
| violations.append(f"{relative}: missing from {', '.join(missing)}") | ||
| continue |
There was a problem hiding this comment.
From my agent:
Only the leaf path and root are checked for is_symlink(); an intermediate directory symlink is followed transparently by is_file()/read_bytes() and reported as "identical," defeating the stated symlink guard for everything under it.
| echo "SETUP_SANITIZER=${SETUP_SANITIZER}" | ||
| echo "BINDINGS_SOURCE=${BINDINGS_SOURCE}" | ||
| echo "CUDA_BINDINGS_ROOT=${CUDA_BINDINGS_ROOT}" | ||
| echo "CUDA_PYTHON_ARTIFACT_NAME=cuda-python-wheel-cuda${BINDINGS_BUILD_CUDA_VER:-${CUDA_VER}}" |
There was a problem hiding this comment.
From my agent:
CUDA_PYTHON_ARTIFACT_NAME falls back to ${CUDA_VER} (the test runner's CTK) rather than the build-time toolkit version in published mode, since BINDINGS_BUILD_CUDA_VER is only set in the local branch. Any workflow consuming this name in that mode downloads a nonexistent artifact.
| ' <<< "$BINDINGS_CONFIG") | ||
| maintenance_line_id=$(jq -er ' | ||
| .roles.maintenance | ||
| | if length == 1 then .[0] else error("wheel builder currently requires one maintenance line") end |
There was a problem hiding this comment.
From my agent:
roles.maintenance is schema'd as a list, but the wheel builder hard-errors unless it has exactly one entry, with the job otherwise hardwired to fixed CURRENT_/MAINTENANCE_ env pairs. Adding a second maintenance line — the stated point of a registry — breaks every build job. Fails loudly, so it's a documented design limit rather than silent corruption, but worth flagging as inconsistent with the registry's stated generality.
|
|
||
| ### CI Pipeline Flow | ||
|
|
||
|  |
There was a problem hiding this comment.
My agent flagged that this diagram is now out-of-date.
| active_section = "" | ||
| for line in pyproject.read_text(encoding="utf-8").splitlines(): | ||
| if match := _SECTION_PATTERN.fullmatch(line): | ||
| active_section = match.group(1).strip() | ||
| continue | ||
| if active_section == section and (match := _KEY_PATTERN.match(line)) and match.group(1) == key: | ||
| return True | ||
| return False |
There was a problem hiding this comment.
Should use tomllib rather than regexes to read toml.
REMINDER
Before merging, remove the temporary
.lycheeignorebefore triggering final CI. It excludes only three canonicalmain/cuda_bindings_12URLs that cannot resolve until this PR is merged. The authored-sourcelycheehook is skipped by CI, so removing the file will not prevent final CI from passing.After merging, run
pre-commit run lychee --all-fileson freshmainto validate those links.Summary
Closes #1199.
This PR is the writable continuation of Keith Kraus's original PR #2675, "cuda.bindings: build 12.9 and 13.x selectively from main". Most of the CUDA 12 source import and the initial build, test, and release integration originated in Keith's PR.
PR #2675 was automatically closed when its temporary
pull-request/2467base was deleted after #2467 merged. Its head branch was not maintainer-writable, so this replacement preserves that work and commit history, retargets it to currentmain, and completes the redesign requested during review.The result is one active development branch for both released CUDA bindings lines:
main.ci/versions.ymlis the authoritative release-line registry for CI and release tooling.12.9.xbranch becomes a read-only release record, not an active backport or artifact-source branch.This builds on the dependency-aware selective CI merged in #2467.
Release-Line Registry
ci/versions.ymlrecords the two released bindings lines:released-12maintenancecuda_bindings_12/v12.9.*released-13currentcuda_bindings/v13.3.*Each line also declares its exact toolkit pin and channel and its prerelease-tag policy.
ci/tools/bindings_config.pyvalidates the registry and emits normalized records for downstream consumers. The registry has onecurrentline and an orderedmaintenancelist. Every configured line must be assigned exactly one role.CI and release logic select stable line IDs or roles instead of treating CUDA versions, source-directory names, and
current/backportas interchangeable concepts. Updating a role is therefore centralized and reviewable.The monolithic wheel builder retains one explicit transitional boundary: it currently requires exactly one
currentline and onemaintenanceline with different CUDA ABI majors. Unsupported registry shapes fail closed instead of producing incomplete artifacts.Source Layout and Maintenance Model
The filesystem names identify their contents; the registry identifies their orchestration role:
cuda_bindings/contains the released CUDA 13 line.cuda_bindings_12/contains the released CUDA 12.9 line.currentandmaintenanceexist only as registry roles, not as directory aliases.The two complete package roots are an intentional transitional design. Most of
cuda_bindings_12/is imported fromNVIDIA/cuda-python@238955935bd903ac72817c0dfdfe4f6a54ee6bb1:cuda_bindings.cuda_bindings_12/MAINTENANCE.mdrecords ownership and generation provenance, including the portion reproduced by cybind commit95d8bb525de46a9ff7ae40d759a98cbe50cf8391.ci/cuda-bindings-shared-files.jsonlists the small handwritten subset that must remain byte-identical across the two roots, and a pre-commit/CI checker enforces that invariant. Generated and cybind-owned support files are guarded by their existing content seals and recorded provenance rather than being required to match across different CTK targets.For subsequent bindings fixes, contributors must update every applicable root or document concretely why one line is unaffected.
Build, Test, and Release Behavior
cuda-pythonpackage; exercise the relevant CUDA Core ABI as neededv12.9.*release tagv13.3.*release tagRelease selection uses the registry stored in the tagged source tree. A compatibility path supports older tags whose source trees predate the registry and still use the generic
cuda_bindings/directory.One commit may carry one CUDA 12.9 release tag and one CUDA 13.3 release tag. Each tag independently selects the matching line, package root, metadata, and artifacts. Each tag still triggers its own release run; combining both releases into one run is not required here. Two different releases from the same tag family should not be placed on one commit.
CUDA 12 development versions advance normally after the latest stable CUDA 12 tag becomes reachable instead of remaining pinned indefinitely to the same development version.
Decisions Requested From Reviewers
Please explicitly accept or reject these policies:
mainis the sole active source of truth. The historical12.9.xbranch receives no further routine or emergency backports. Applicable CUDA 12 fixes are made incuda_bindings_12/onmain, alongside any corresponding current-line change.ci/versions.yml.The generated NVML memoryview change discussed during the first review is not part of this PR and should be handled independently if needed.
Review Map
The high-value review surface is outside the imported CUDA 12 tree:
ci/versions.yml,ci/tools/bindings_config.py, and their testsci/tools/compute_ci_plan.py, matrix validation, and their testsci/tools/ci/cuda-bindings-shared-files.json, its checker, andcuda_bindings_12/MAINTENANCE.mdMost files under
cuda_bindings_12/are the direct CUDA 12.9 import from Keith's PR #2675. Differences fromcuda_bindings/are generally target- or generation-specific and intentional.Validation
TestVenv/bin/python -m pytest ci/tools/tests: 161 passed on final local headc6a0cf1pre-commit run --all-files, including registry validation, shared-file checks, generated-file seals, Ruff, actionlint, YAML/TOML/RST checks, andlycheewith the temporary exception described in the REMINDERc6a0cf1: run 33572804123 passed with all 104 jobs successfulOut of Scope
Checklist