[ExecuTorch][llm] Fuse w1+w3 into single GEMM in quantized_moe_ffn#21124
[ExecuTorch][llm] Fuse w1+w3 into single GEMM in quantized_moe_ffn#21124digantdesai wants to merge 2 commits into
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21124
Note: Links to docs will display an error until the docs builds have been completed. ❗ 1 Active SEVsThere are 1 currently active SEVs. If your PR is affected, please view them below: ❌ 1 New Failure, 10 Pending, 2 Unrelated FailuresAs of commit a4989f1 with merge base 21554e5 ( NEW FAILURE - The following job has failed:
FLAKY - The following jobs failed but were likely due to flakiness present on trunk:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This PR needs a
|
Stack from ghstack (oldest at bottom):
Fuse the up-projection (w1) and gate-projection (w3) into a single [2F, D] GEMM per expert. This halves the number of torchao activation quantizations per expert (from 2 to 1) and reduces total GEMM calls from 3 to 2 per active expert.
At AOT time, w1 and w3 are concatenated before packing: pack_fn(cat([w1, w3], dim=0)). At runtime, a single expert_linear_dispatch produces [m_e, 2F], then a fused swiglu_and_compact pass reads the interleaved h1/h3 and writes [m_e, F] for the w2 down-projection.
Schema changes from (packed_w1, packed_w3, packed_w2) to (packed_w13, packed_w2) — one fewer tensor arg (14 -> 13).
Differential Revision: D102799854