llama.cpp/tests/test-backend-ops.cpp at gg/remove-k-quants-per-iter

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-11-01 09:01:57 +00:00

Files

Johannes Gäßler e11bd856d5 CPU/CUDA: Gemma 2 FlashAttention support (#8542 )

* CPU/CUDA: Gemma 2 FlashAttention support

* apply logit_softcap to scale in kernel

* disable logit softcapping tests on Metal

* remove metal check

2024-08-24 21:34:59 +02:00

93 KiB

Raw Permalink Blame History

View Raw

93 KiB Raw Permalink Blame History

93 KiB

Raw Permalink Blame History