llama.cpp

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-08-16 05:02:58 -04:00

Files

Xuan Son Nguyen 1e6f6554aa server : add lora hotswap endpoint (WIP) (#8857 )

* server : add lora hotswap endpoint

* handle lora_no_apply

* fix build

* updae docs

* clean up struct def

* fix build

* add LoRA test

* fix style

2024-08-06 17:33:39 +02:00

__init__.py

convert-*.py: GGUF Naming Convention Refactor and Metadata Override Refactor (#7499 )

2024-07-18 20:40:15 +10:00

constants.py

Stop the generation when <|eom_id|> token is encountered - needed for Llama 3.1 tool call support (#8858 )

2024-08-05 09:38:01 +02:00

gguf_reader.py

py : type-check all Python scripts with Pyright (#8341 )

2024-07-07 15:04:39 -04:00

gguf_writer.py

Stop the generation when <|eom_id|> token is encountered - needed for Llama 3.1 tool call support (#8858 )

2024-08-05 09:38:01 +02:00

gguf.py

…

lazy.py

convert_hf : faster lazy safetensors (#8482 )

2024-07-15 23:13:10 -04:00

metadata.py

server : add lora hotswap endpoint (WIP) (#8857 )

2024-08-06 17:33:39 +02:00

py.typed

…

quants.py

Fix conversion of unnormalized BF16->BF16 weights (#7843 )

2024-08-02 15:11:39 -04:00

tensor_mapping.py

convert_hf : faster lazy safetensors (#8482 )

2024-07-15 23:13:10 -04:00

utility.py

gguf-py : fix some metadata name extraction edge cases (#8591 )

2024-07-20 21:58:49 -04:00

vocab.py

Move convert.py to examples/convert-legacy-llama.py (#7430 )

2024-05-30 21:40:00 +10:00