Skip to content

Avoid nvJitLink cache crashes during incremental LTO - #3040

Open
GodlyDonuts wants to merge 2 commits into
NVIDIA:mainfrom
GodlyDonuts:codex/mitigate-incremental-lto-cache-crash
Open

GodlyDonuts wants to merge 2 commits into
NVIDIA:mainfrom
GodlyDonuts:codex/mitigate-incremental-lto-cache-crash

Conversation

@GodlyDonuts

Copy link
Copy Markdown

Related to #3038.

Disable nvJitLink's intermediate cache when incremental linking and link-time optimization are both enabled. Repeated native round trips reproduced access violations on Windows with caching enabled, including through direct calls that bypass cuda-bindings.

Apply -no-cache at every incremental LTO stage, document the behavior, and add regression coverage for option generation. Other linking modes retain their existing caching behavior.

Validation:

  • Cached native runs crashed in 7 of 8 processes; cache-disabled runs completed 1,000 round trips in each of 8 processes.
  • Free-threaded Python 3.14 linker tests: 81 passed, 3 skipped with four workers, using the edited option-generation method injected into the existing wheel.
  • Cython C++ translation and git diff --check passed.

Draft pending rebuilt-extension validation and confirmation against the original MCDM CI environment. This mitigates the reproduced failure; it does not establish that every crash in #3038 has the same cause.

@copy-pr-bot

copy-pr-bot Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@github-actions github-actions Bot added the cuda.core Everything related to the cuda.core module label Oct 7, 2026
@GodlyDonuts
GodlyDonuts marked this pull request as ready for review October 7, 2026 08:55
@coderabbitai

coderabbitai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Important

Review skipped

Review was skipped as selected files did not have any reviewable changes.

⚙️ Run configuration
  • Configuration used: Repository: NVIDIA/cuda-python/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: 2f4b39bd-ca6c-4eb3-87dc-a6e3dbfaa3b9
📥 Commits

Reviewing files that changed from the base of the PR and between f687c49 and 0e4a3d9.

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA/cuda-python/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: a494d414-9adb-40eb-874c-d78ad1a00deb
📥 Commits

Reviewing files that changed from the base of the PR and between 5078fe6 and f687c49.

📒 Files selected for processing (3)
  • cuda_core/cuda/core/_linker.pyi
  • cuda_core/cuda/core/_linker.pyx
  • cuda_core/tests/test_linker.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


📝 Summary

Summary by CodeRabbit

  • Bug Fixes
    • Incremental links using link-time optimization now automatically disable intermediate caching, avoiding incompatible cache behavior.
  • Documentation
    • Clarified when intermediate caching is disabled automatically.

Walkthrough

The nvJitLink options builder now emits -no-cache when incremental linking and link-time optimization are both enabled. The option documentation describes this behavior, and a test checks when the flag is present or absent.

Changes

Incremental LTO cache handling

Layer / File(s) Summary
Document and apply cache option
cuda_core/cuda/core/_linker.pyi, cuda_core/cuda/core/_linker.pyx, cuda_core/tests/test_linker.py
The option documentation describes automatic cache disabling for incremental LTO links. The nvJitLink options builder emits -no-cache when both options are enabled. The test checks that the flag is absent when either option is disabled.

Suggested reviewers: andy-jost

Priority: ➖ Normal

Change: Bug fix

Merge Risk: ⚪ Minimal · up to f687c

No actionable merge-blocking risk was identified in the reviewed changes.

✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@leofang leofang left a comment •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@GodlyDonuts this comes out of blue. Could you please present your analysis on why you think the linker cache is to blame? For example, a stack trace when the segfault happens would be nice.

@leofang leofang added the awaiting-response Further information is requested label Oct 7, 2026
@GodlyDonuts

Copy link
Copy Markdown
Author

The original CI failure happens during the first incremental LTO link, but that traceback alone doesn’t prove caching is the cause. I reduced the path to nvJitLink directly and tested it through both cuda-bindings and ctypes on Windows with nvJitLink 13.4.92.

With caching enabled, 7/8 processes failed across the two paths. With -no-cache on both incremental LTO stages, all 8/8 completed 1,000 round trips each. I also tried disabling the cache only for the first stage and still got a crash at the second cached stage, which is why the change applies whenever incremental=True and LTO is enabled.

The crash dump lands inside nvJitLink_130_0.dll and includes the Windows freed-memory pattern, so a native lifetime issue looks plausible, but I haven’t established the exact internal defect or which cache object is responsible.

This supports the workaround but actual confirmation against the CI failures is still pending.

@GodlyDonuts

Copy link
Copy Markdown
Author

Here is the native stack excerpt from the local reproduction, captured with caching enabled and one worker thread. Personal paths and machine metadata have been removed.

Exception: C0000005.ACCESS_VIOLATION

rdx=feeefeeefeeefeee
rip=00007ffdaa2813a2

nvJitLink_130_0!_nvJitLinkGetLinkedPtx_13_1+0x1434f92:
00007ffd`aa2813a2 ff5210  call qword ptr [rdx+10h]

Call stack (top frames):
nvJitLink_130_0!_nvJitLinkGetLinkedPtx_13_1+0x1434f92
nvJitLink_130_0!_nvJitLinkAPI+0xf17
nvjitlink_cp310_win_amd64_7ffe2b9b0000+0x2ecf
cynvjitlink_cp310_win_amd64+0x1250
nvjitlink_cp310_win_amd64+0x36f9
nvjitlink_cp310_win_amd64+0x3939
python310!_PyEval_EvalFrameDefault+0x1a3b

The faulting instruction dereferences rdx, which contains a freed-memory pattern. This points toward a native lifetime issue, but doesn’t identify the responsible cache object. The large symbol offset also means the displayed GetLinkedPtx name should not be treated as identification of the actual internal function.
segfault-stack-sanitized.txt

@GodlyDonuts
GodlyDonuts requested a review from leofang October 7, 2026 19:47
@GodlyDonuts

Copy link
Copy Markdown
Author

@claude review this PR

@leofang

leofang commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

Thanks, @GodlyDonuts. I have shared your traceback with the compiler team and they are looking into it. I'll keep you posted, though at this point we cannot yet establish that the cache is to blame and accept this PR; the compiler team is against turning off the cache.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting-response Further information is requested cuda.core Everything related to the cuda.core module

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants