llama.cpp/ggml-cuda/common.cuh at 864a99e7a01d9422d2f55618dbe62c8099a2175c

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-11-13 10:57:15 +00:00

Files

Johannes Gäßler 1f0dabda8d CUDA: use tensor cores for MMQ (#7676 )

* CUDA: int8 tensor cores for MMQ (legacy quants)

* fix out-of-bounds writes

* __builtin_assume -> GGML_CUDA_ASSUME

* fix writeback returning too early

2024-06-10 11:45:13 +02:00

28 KiB

Raw Blame History

View Raw

28 KiB Raw Blame History

28 KiB

Raw Blame History