llama.cpp/ggml-cuda/fattn-common.cuh at 7c7836d9d4062d6858e3fb337b135c417ccee6ce

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-11-02 09:12:03 +00:00

Files

Johannes Gäßler 1f0dabda8d CUDA: use tensor cores for MMQ (#7676 )

* CUDA: int8 tensor cores for MMQ (legacy quants)

* fix out-of-bounds writes

* __builtin_assume -> GGML_CUDA_ASSUME

* fix writeback returning too early

2024-06-10 11:45:13 +02:00

25 KiB

Raw Blame History

View Raw

25 KiB Raw Blame History

25 KiB

Raw Blame History