llama.cpp/examples/server/server.cpp at 5cab3e4aaa892df4620b720f20a503a1bbbe7a52

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-10-31 08:51:55 +00:00

Files

Xuan Son Nguyen 57bb2c40cd server : fix logprobs, make it OAI-compatible (#10783 )

* server : fix logprobs, make it openai-compatible

* update docs

* add std::log

* return pre-sampling p

* sort before apply softmax

* add comment

* fix test

* set p for sampled token

* update docs

* add --multi-token-probs

* update docs

* add `post_sampling_probs` option

* update docs [no ci]

* remove --multi-token-probs

* "top_probs" with "post_sampling_probs"

* resolve review comments

* rename struct token_prob to prob_info

* correct comment placement

* fix setting prob for sampled token

2024-12-19 15:40:08 +01:00

160 KiB

Raw Blame History

View Raw

160 KiB Raw Blame History

160 KiB

Raw Blame History