llama.cpp

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-06-29 12:35:16 +00:00

Files

Xuan Son Nguyen 57bb2c40cd server : fix logprobs, make it OAI-compatible (#10783 )

* server : fix logprobs, make it openai-compatible

* update docs

* add std::log

* return pre-sampling p

* sort before apply softmax

* add comment

* fix test

* set p for sampled token

* update docs

* add --multi-token-probs

* update docs

* add `post_sampling_probs` option

* update docs [no ci]

* remove --multi-token-probs

* "top_probs" with "post_sampling_probs"

* resolve review comments

* rename struct token_prob to prob_info

* correct comment placement

* fix setting prob for sampled token

2024-12-19 15:40:08 +01:00

test_basic.py

server : add flag to disable the web-ui (#10762 ) (#10751 )

2024-12-10 18:22:34 +01:00

test_chat_completion.py

server : fix logprobs, make it OAI-compatible (#10783 )

2024-12-19 15:40:08 +01:00

test_completion.py

server : fix logprobs, make it OAI-compatible (#10783 )

2024-12-19 15:40:08 +01:00

test_ctx_shift.py

server : replace behave with pytest (#10416 )