llama.cpp

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-10-28 08:31:25 +00:00

Files

Georgi Gerganov 644fd71b44 sampling : refactor + optimize penalties sampler (#10803 )

* sampling : refactor + optimize penalties sampler

ggml-ci

* common : apply ignore_eos as logit bias

ggml-ci

* batched : remove penalties sampler

* params : allow penalty_last_n == -1 to be equal to context size

ggml-ci

* common : by default, move the penalties at the end of the sampling chain

ggml-ci

* common : ignore all EOG tokens

Co-authored-by: Diego Devesa <slarengh@gmail.com>

* common : move back the penalties at the front of the sampling chain

ggml-ci

* readme : restore hint about --ignore-eos flag [no ci]

* llama : minor

ggml-ci

* webui : update

---------

Co-authored-by: Diego Devesa <slarengh@gmail.com>

2024-12-16 12:31:14 +02:00

llama-cpp.h

Introduce llama-run (#10291 )

2024-11-25 22:56:24 +01:00

llama.h

sampling : refactor + optimize penalties sampler (#10803 )

2024-12-16 12:31:14 +02:00