examples : add passkey test (#3856)

* examples : add passkey test * passkey : better prints * passkey : select pass key pos from CLI * passkey : simplify n_past logic * make : add passkey target * passkey : add "self-extend"-like context extension (#4810) * llama : "self-extend"-like context extension * passkey : add comment * passkey : add readme
2025-08-12 03:21:10 -04:00 · 2024-01-08 11:14:04 +02:00
parent b7e7982953
commit b0034d93ce
9 changed files with 361 additions and 1 deletions
--- a/examples/batched/batched.cpp
+++ b/examples/batched/batched.cpp
@@ -69,6 +69,7 @@ int main(int argc, char ** argv) {

    std::vector<llama_token> tokens_list;
    tokens_list = ::llama_tokenize(model, params.prompt, true);
+
    const int n_kv_req = tokens_list.size() + (n_len - tokens_list.size())*n_parallel;

    // initialize the context