-
Notifications
You must be signed in to change notification settings - Fork 1.2k
Pull requests: deepseek-ai/DeepGEMM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
fix: SM90 paged mqa
prefix_sum out of range (CUDBG_EXCEPTION_WARP_OUT_OF_RANGE_ADDRESS)
#399
opened Aug 5, 2026 by
cjackal
Loading…
Make the JIT cache key independent of the install include path
#398
opened Aug 5, 2026 by
matteso1
Loading…
Don't enable the assertion check for release build
#390
opened Jul 22, 2026 by
zhangqingshan373
Loading…
Optimize sm100 version with token-aware BM240 tiling, native UE8M0, and direct combine
#386
opened Jul 17, 2026 by
qinqinwo
Loading…
Add B200-Calibrated Sparse-Routing Adaptive Wave Sizing
#381
opened Jul 15, 2026 by
qinqinwo
Loading…
SM120: MoE grouped GEMM decode optimizations: skip padding I/O, BLOCK_M=32
#380
opened Jul 15, 2026 by
leavelet
Loading…
SM120: fix FP8 MQA logits swizzle mode for head_dim < 128
#379
opened Jul 15, 2026 by
leavelet
Loading…
Fix non-deterministic results in k_grouped_fp8_gemm_nt_contiguous
#375
opened Jul 9, 2026 by
Functionhx
Loading…
fix: index CUDA sources/headers when generating .pyi stubs
#361
opened Jun 14, 2026 by
Osamaali313
Loading…
Draft: Add green-context split-kernel MegaMoE features
#357
opened Jun 12, 2026 by
RayWang96
Collaborator
Loading…
fix: use importlib to load scripts/generate_pyi.py instead of package…
#355
opened Jun 5, 2026 by
Mikezhang001
Loading…
Add shape-aware recommended alignment for SM90 small-M grouped GEMM
#350
opened Jun 2, 2026 by
qescccczmr
Loading…
Previous Next
ProTip!
Filter pull requests by the default branch with base:main.