llama.cpp

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-06-28 12:25:03 +00:00

Files

Xuan Son Nguyen 0da5d86026 server : allow using LoRA adapters per-request (#10994 )

* slot.can_batch_with

* lora per request

* test: force disable cache prompt

* move can_batch_with check

* fix condition

* add slow test with llama 8b

* update docs

* move lora change task to queue

* Apply suggestions from code review

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* lora_base

* remove redundant check

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

2025-01-02 15:05:18 +01:00

test_basic.py

server : add flag to disable the web-ui (#10762 ) (#10751 )

2024-12-10 18:22:34 +01:00

test_chat_completion.py

server : clean up built-in template detection (#11026 )

2024-12-31 15:22:01 +01:00

test_completion.py

server : add OAI compat for /v1/completions (#10974 )

2024-12-31 12:34:13 +01:00

test_ctx_shift.py

server : replace behave with pytest (#10416 )