llama.cpp/ggml-quants.c at 3d032ece8e6973441273601ca2981130608e287d

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-11-01 09:01:57 +00:00

Files

Justine Tunney 7733f0c760 ggml : support AVX512VNNI (#6280 )

This change causes some quants (e.g. Q4_0, Q8_0) to go faster on some
architectures (e.g. AMD Zen 4).

2024-03-25 07:39:56 +02:00

View Raw