llama : refactor model loading code (#2620)

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-11-03 09:22:01 +00:00

* llama : style formatting + remove helper methods

* llama : fix quantization using gguf tool

* llama : simplify gguf_file_saver

* llama : fix method names

* llama : simplify write_header()

* llama : no need to pass full file loader to the file saver

just gguf_ctx

* llama : gguf_file_saver write I32

* llama : refactor tensor names (#2622)

* gguf: update tensor names searched in quantization

* gguf : define tensor names as constants

* gguf : initial write API (not tested yet)

* gguf : write to file API (not tested)

* gguf : initial write API ready + example

* gguf : fix header write

* gguf : fixes + simplify example + add ggml_nbytes_pad()

* gguf : minor

* llama : replace gguf_file_saver with new gguf write API

* gguf : streaming support when writing files

* gguf : remove oboslete write methods

* gguf : remove obosolete gguf_get_arr_xxx API

* llama : simplify gguf_file_loader

* llama : move hparams and vocab from gguf_file_loader to llama_model_loader

* llama : merge gguf-util.h in llama.cpp

* llama : reorder definitions in .cpp to match .h

* llama : minor simplifications

* llama : refactor llama_model_loader (WIP)

wip : remove ggml_ctx from llama_model_loader

wip : merge gguf_file_loader in llama_model_loader

* llama : fix shape prints

* llama : fix Windows build + fix norm_rms_eps key

* llama : throw error on missing KV paris in model meta data

* llama : improve printing + log meta data

* llama : switch print order of meta data

---------

Co-authored-by: M. Yusuf Sarıgöz <yusufsarigoz@gmail.com>

This commit is contained in:

Georgi Gerganov

2023-08-16 14:34:03 +03:00

committed by

GitHub

parent ea5615a03a

commit 758ff1bbb5

9 changed files with 1944 additions and 1889 deletions

									
										2

convert-llama-h5-to-gguf.py
									
												View File
												
				@@ -135,7 +135,7 @@ if Path(dir_model + "/tokenizer.model").is_file():

				        toktype = 1 # defualt to normal token type

				        if tokenizer.is_unknown(i): toktype = 2

				        if tokenizer.is_control(i): toktype = 3

				        # TODO: How to determinate if a token is user defined?

				        # ref: https://github.com/google/sentencepiece/blob/master/src/sentencepiece_model.proto

				        # if tokenizer.is_user_defined(i): toktype = 4

llama : refactor model loading code (#2620)

2 convert-llama-h5-to-gguf.py Unescape Escape View File

2

convert-llama-h5-to-gguf.py

View File