Repository navigation
cuda.core: accept stream-captured memcpy nodes in MemcpyNode.update - #3029
Conversation
The driver records both operands of a captured cuMemcpyAsync with the unified memory type. The descriptor check in MemcpyNode.update admitted only host and device operands, so every memcpy node produced by stream capture was rejected with a message about multidimensional copies. Accept the unified type for one-dimensional descriptors and read such operands from the device-pointer field. A replaced operand keeps the unified type, in the definition node and in the executable view, because the driver rejects a change of an operand's memory type in an executable update; Graph.update therefore applies the change as well. The node repr shows U for unified operands. Closes NVIDIA#2649 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
/ok to test fb71159 |
|
The update() docstrings on MemcpyNode and ExecutableMemcpyNode now say only what a caller needs: nodes recorded by stream capture are supported. The driver's memory-type bookkeeping stays in code comments. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
/ok to test 46ca8ac |
|
Reviewed at 46ca8ac; built it locally and ran 1. The release note overstates what The note says a replaced operand stays unified because the driver rejects a memory-type change, "so
This is not caused by this PR. A control with an ordinary 2. (Optional) Add a test that documents that boundary. No test replaces a captured node's operand with a host buffer. A short one would capture a device-to-device copy, call 3. Add a comment in It reads the recorded memory types from the definition node ( Minor nits, no action needed: the three new tests unpack |
…erand boundary (#3035) * cuda.core: narrow the memcpy update release note and test the host-operand boundary Follow-up to #3029. The release note said Graph.update() applies any replaced operand of a captured memcpy node; the driver accepts only a device-to-device replacement in an executable update, and a replacement with host memory takes effect in a new instantiation. A test pins that boundary, and ExecutableMemcpyNode.update() records why it may read the memory types from the definition node. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * cuda.core: allow either executable-update outcome for a host operand CI showed that Windows TCC drivers accept a host buffer as the replacement operand of a captured memcpy node in an executable update, while Linux drivers reject it. The test now accepts either outcome and checks the copy when the update is accepted, and the release note says the driver decides. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * cuda.core: say what decides whether a host operand update is accepted A diagnostic run across Linux, Windows MCDM, and Windows TCC rows showed that the executable-update outcome follows the allocation kind, not the platform: the driver compares the memory class of pool-backed and virtual-memory operands and rejects a replacement by host memory or by a cuMemAlloc buffer, while plain cuMemAlloc operands accept it. The TCC runners report no memory-pool support, so the default memory resource is not pool-backed there, which is why the update was accepted on them. The release note and the test docstring now say so. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(cuda.core): split the host-operand update test by allocation kind Pool-backed operands go through the driver's memory-class comparison, so that case now requires the PARAMETERS_CHANGED rejection. cuMemAlloc operands skip the comparison and the driver may accept the change, so that case keeps the lenient check and verifies the copy when accepted. A device without memory-pool support runs only the cuMemAlloc case. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(cuda.core): instantiate before changing the captured node The executable node update ran against a graph instantiated after the definition node already pointed at host memory, so it restated the operand and tested nothing. Both instantiations now precede the change. Pool-backed operands require CUDA_ERROR_INVALID_VALUE from the node update, the code cuGraphExecMemcpyNodeSetParams reports for the same memory-class rejection. cuMemAlloc operands check the code when the driver rejects. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Summary
MemcpyNode.update()raisedNotImplementedErrorfor every memcpy node produced by stream capture. The driver records both operands of a capturedcuMemcpyAsync(Buffer.copy_from,Buffer.copy_to) with the unified memory type, and the descriptor check admitted only host and device operands. The error message blamed multidimensional, pitched, or array-backed copies, which did not describe the rejected node.Changes
_is_supported_memcpy_descriptoracceptsCU_MEMORYTYPE_UNIFIEDfor one-dimensional, unpitched, unoffset descriptors. The pitch and offset checks are unchanged, so multidimensional and array-backed nodes are still rejected with the existing message.MemcpyNode.update()reads a unified operand from the device-pointer field. A replaced operand keeps the unified type and its new address goes into that field. The driver rejects a change of an operand's memory type in an executable update (cuGraphExecMemcpyNodeSetParamsandcuGraphExecUpdateboth compare the type of each operand), so keeping the type letsGraph.update()apply the change to an instantiated graph. Operands recorded as host or device memory are classified as before.ExecutableMemcpyNode.update()reads the recorded descriptor of the node and keeps a unified operand unified for the same reason. Without this, the executable view rejected every node from stream capture withCUDA_ERROR_INVALID_VALUE.MemcpyNode.__repr__showsUfor unified operands andAfor array operands instead ofDfor every non-host type.Graph.update()on an earlier one copy what the updated node describes; the executable view of a captured node accepts new operands and a new size. The existing rejection test for pitched descriptors is unchanged.Related Work
Closes #2649. Multidimensional, pitched, and array-backed descriptors remain out of scope here and are tracked in #2420.
🤖 Generated with Claude Code