Default Branch

ee4c505a4f · server: add dedup-cache-models preset option (#27346) · Updated 2026-08-19 17:04:26 +08:00

Branches

de7e0912b6 · convert : ignore tokens if their IDs are within [0, vocab_size) · Updated 2023-10-28 20:01:36 +08:00    tqcq

9060
1

bbfc62ac2f · sampling : temp == 0.0 -> no probs, temp < 0.0 -> probs · Updated 2023-10-28 19:04:57 +08:00    tqcq

9068
3

cd3e20fb50 · cuda : fix multi-gpu with tensor cores · Updated 2023-10-28 04:11:50 +08:00    tqcq

9067
3

49af767fad · build : add compile option to force use of MMQ kernels · Updated 2023-10-27 18:21:04 +08:00    tqcq

9069
7

d798a17c34 · cuda : add TODO for calling cublas from kernel + using mem pool · Updated 2023-10-24 21:33:24 +08:00    tqcq

9083
10

6966474928 · cuda : play with faster Q4_0 dequantization · Updated 2023-10-24 15:29:40 +08:00    tqcq

9083
8

b9bb4cbe86 · Separate bug and enhancement template + no default title · Updated 2023-10-23 23:59:11 +08:00    tqcq

9083
1

c0f4d54870 · server : add comment about changing slot_state to bool · Updated 2023-10-23 03:24:39 +08:00    tqcq

9089
72

cb79f8a2d8 · llama : add SKIP_KQ_KQV option · Updated 2023-10-22 14:58:29 +08:00    tqcq

9089
3

56ba00b923 · sampling : hide prev behind API and apply #3661 · Updated 2023-10-20 23:53:27 +08:00    tqcq

9092
6

ad2727d091 · Merge branch 'master' into speculative-tree · Updated 2023-10-18 15:50:58 +08:00    tqcq

9103
18

932589c0ef · Honor -ngl option for Cuda offloading in llava · Updated 2023-10-14 08:12:10 +08:00    tqcq

9117
1

5261aee8d8 · sampling : one sequence per sampling context · Updated 2023-10-13 01:36:44 +08:00    tqcq

9120
1

2fcdf869cd · batched-bench : add mmq CLI arg · Updated 2023-10-12 00:42:33 +08:00    tqcq

9132
7

ee7456926e · ggml-alloc : fix assert in debug builds · Updated 2023-10-09 20:33:12 +08:00    tqcq

9141
1

ee268b5446 · llama : no longer perform uninitialized access to the KV cache · Updated 2023-10-08 16:49:38 +08:00    tqcq

9148
5

acead654d2 · Merge branch 'master' into fix-refact · Updated 2023-10-08 16:25:16 +08:00    tqcq

9148
4

6b9554a740 · metal : print more GPU info + disable mul_mm for MTLGPUFamiliy < Apple7 · Updated 2023-10-08 14:55:13 +08:00    tqcq

9155
5

ba44776dc2 · bump version · Updated 2023-10-08 02:47:48 +08:00    tqcq

9154
6

5ab6c2132a · server-parallel : add "--reverse-prompt" + compiler warning fixes · Updated 2023-10-06 19:32:19 +08:00    tqcq

9167
4