Is this a duplicate?
Type of Bug
Runtime Error
Component
General cuda-python
Describe the bug
On a machine where a non-NVIDIA GPU drives the display (for example, a hybrid-graphics laptop with a non-NVIDIA iGPU plus an NVIDIA dGPU), the CUDA-GL interop tests fail instead of being skipped:
cuda_bindings/tests/test_graphics_apis.py::test_cuda_gl_register_image_smoketest (2 cases)
- 24 tests in
cuda_core/tests/test_graphics.py
With DISPLAY set, pyglet creates the GL context through GLX on the GPU that drives the screen, which here is the AMD iGPU (GL_VENDOR='AMD'). CUDA-GL interop needs the GL context to be on an NVIDIA GPU, so cuGraphicsGLRegisterImage / cuGraphicsGLRegisterBuffer fail with the generic CUDA_ERROR_UNKNOWN. cudaGLGetDevices returns cudaErrorUnknown for the same context as well.
The tests already skip when the driver refuses interop (CUDA_ERROR_OPERATING_SYSTEM, for example on WSL), but they don't recognize this case. CI doesn't catch it because every CI runner uses NVIDIA GPUs only (see #2077).
This is the same kind of problem as #2864/#2865, where the GL context ended up on a different GPU from the CUDA device. #2865 fixed the headless EGL path by choosing the matching EGL device. With a display, pyglet picks the GPU itself, so the tests need to check where the context actually is.
How to Reproduce
- Use a Linux machine with a non-NVIDIA GPU driving an X11 display and an NVIDIA GPU for CUDA.
- Run
pixi run test (or pytest cuda_bindings/tests/test_graphics_apis.py cuda_core/tests/test_graphics.py) with DISPLAY set.
- The tests fail:
FAILED tests/test_graphics_apis.py::test_cuda_gl_register_image_smoketest[<cudaGraphicsRegisterFlags.cudaGraphicsRegisterFlagsNone: 0>] - AssertionError: cudaGraphicsGLRegisterImage returned cudaErrorUnknown
FAILED tests/test_graphics_apis.py::test_cuda_gl_register_image_smoketest[<cudaGraphicsRegisterFlags.cudaGraphicsRegisterFlagsWriteDiscard: 2>] - AssertionError: cudaGraphicsGLRegisterImage returned cudaErrorUnknown
FAILED tests/test_graphics.py::test_register_image - cuda.core._utils.cuda_utils.CUDAError: CUDA_ERROR_UNKNOWN: This indicates that an unknown internal error has occurred.
... (24 cuda_core tests in total)
With DISPLAY unset, the tests use headless EGL, which selects the NVIDIA GPU, and all of them pass. That points to the GL context's GPU, not to the bindings or GraphicsResource.
Expected behavior
When the current GL context is not on an NVIDIA GPU, the graphics tests should skip with a clear reason (the GL vendor and renderer) instead of failing with CUDA_ERROR_UNKNOWN. On a machine where the context is on an NVIDIA GPU, they should behave exactly as before.
Operating System
Pop!_OS 24.04 LTS (kernel 7.0.11), X11 session
nvidia-smi output
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 595.84 Driver Version: 595.84 CUDA Version: 13.2 |
+-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA GeForce RTX 3050 ... Off | 00000000:01:00.0 Off | N/A |
| N/A 37C P3 15W / 30W | 15MiB / 4096MiB | 0% Default |
| | | N/A |
+-----------------------------------------+------------------------+----------------------+
+-----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| 0 N/A N/A 2159 G /usr/lib/xorg/Xorg 4MiB |
+-----------------------------------------------------------------------------------------+
Is this a duplicate?
Type of Bug
Runtime Error
Component
General cuda-python
Describe the bug
On a machine where a non-NVIDIA GPU drives the display (for example, a hybrid-graphics laptop with a non-NVIDIA iGPU plus an NVIDIA dGPU), the CUDA-GL interop tests fail instead of being skipped:
cuda_bindings/tests/test_graphics_apis.py::test_cuda_gl_register_image_smoketest(2 cases)cuda_core/tests/test_graphics.pyWith
DISPLAYset, pyglet creates the GL context through GLX on the GPU that drives the screen, which here is the AMD iGPU (GL_VENDOR='AMD'). CUDA-GL interop needs the GL context to be on an NVIDIA GPU, socuGraphicsGLRegisterImage/cuGraphicsGLRegisterBufferfail with the genericCUDA_ERROR_UNKNOWN.cudaGLGetDevicesreturnscudaErrorUnknownfor the same context as well.The tests already skip when the driver refuses interop (
CUDA_ERROR_OPERATING_SYSTEM, for example on WSL), but they don't recognize this case. CI doesn't catch it because every CI runner uses NVIDIA GPUs only (see #2077).This is the same kind of problem as #2864/#2865, where the GL context ended up on a different GPU from the CUDA device. #2865 fixed the headless EGL path by choosing the matching EGL device. With a display, pyglet picks the GPU itself, so the tests need to check where the context actually is.
How to Reproduce
pixi run test(orpytest cuda_bindings/tests/test_graphics_apis.py cuda_core/tests/test_graphics.py) withDISPLAYset.With
DISPLAYunset, the tests use headless EGL, which selects the NVIDIA GPU, and all of them pass. That points to the GL context's GPU, not to the bindings orGraphicsResource.Expected behavior
When the current GL context is not on an NVIDIA GPU, the graphics tests should skip with a clear reason (the GL vendor and renderer) instead of failing with
CUDA_ERROR_UNKNOWN. On a machine where the context is on an NVIDIA GPU, they should behave exactly as before.Operating System
Pop!_OS 24.04 LTS (kernel 7.0.11), X11 session
nvidia-smi output