llama.cpp/ggml-quants.c at 2f34b865b62b1d2b5eb8a27885e4de220deeacbd

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-11-02 09:12:03 +00:00

Files

Justine Tunney 7733f0c760 ggml : support AVX512VNNI (#6280 )

This change causes some quants (e.g. Q4_0, Q8_0) to go faster on some
architectures (e.g. AMD Zen 4).

2024-03-25 07:39:56 +02:00

View Raw