llama.cpp/tests/test-backend-ops.cpp at 835b2b915c52bcabcd688d025eacff9a07b65f52

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-10-30 08:42:00 +00:00

Files

Aman Gupta 077c94d0ca CUDA: add a fused top-K MoE kernel (#16130 )

* CUDA: add a fused top-K MoE kernel

This kernel does the following:
1. softmax over the logits per token [n_experts, n_tokens]
2. argmax reduce over the top-k (n_experts_used) logits
3. write weights + ids to global memory

It is intended as fusion of softmax->top-k->get_rows pipeline for MoE models

* Refactor into ggml_cuda_should_use_topk_moe

* Review: Use better coalescing pattern, use WARP_SIZE, store logits into registers before

* Review: format + micro-optimizations

* Fix bug: fix tie breakers

* Add optional norm + clean-up code

* Use smem for final write

* Add bounds check

* Use better memory pattern for writeback

2025-09-25 16:35:05 +02:00

268 KiB

Raw Blame History

View Raw

268 KiB Raw Blame History

268 KiB

Raw Blame History