ggml : refactor online repacking (#10446)

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-11-08 10:07:01 +00:00

* rename ggml-cpu-aarch64.c to .cpp

* reformat extra cpu backend.

- clean Q4_0_N_M and IQ4_0_N_M
  - remove from "file" tensor type
  - allow only with dynamic repack

- extract cpu extra bufts and convert to C++
  - hbm
  - "aarch64"

- more generic use of extra buffer
  - generalise extra_supports_op
  - new API for "cpu-accel":
     - amx
     - aarch64

* clang-format

* Clean Q4_0_N_M ref

Enable restrict on C++

* add op GGML_OP_MUL_MAT_ID for Q4_0_N_M with runtime repack

* added/corrected control on tensor size for Q4 repacking.

* Update ggml/src/ggml-cpu/ggml-cpu-aarch64.cpp

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* Update ggml/src/ggml-cpu/ggml-cpu-aarch64.cpp

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* add debug logs on repacks.

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

This commit is contained in:

Djip007

2024-12-07 13:37:50 +01:00

committed by

GitHub

parent c2a16c0bdb

commit 19d8762ab6

33 changed files with 1136 additions and 1049 deletions

									
										8

ggml/src/ggml-cpu/ggml-cpu-hbm.h
									
										Normal file
									
												View File
												
				@@ -0,0 +1,8 @@

				#pragma once

				#include "ggml-backend.h"

				#include "ggml.h"

				// GGML CPU internal header

				ggml_backend_buffer_type_t ggml_backend_cpu_hbm_buffer_type(void);

ggml : refactor online repacking (#10446)

8 ggml/src/ggml-cpu/ggml-cpu-hbm.h Normal file Unescape Escape View File

8

ggml/src/ggml-cpu/ggml-cpu-hbm.h Normal file

View File