Default Branch

ee4c505a4f · server: add dedup-cache-models preset option (#27346) · Updated 2026-08-19 17:04:26 +08:00

Branches

f64e4f04e7 · ggml : testing GPU FP precision via quantized CPY · Updated 2023-12-31 01:11:40 +08:00    tqcq

8768
1

f32f30bc57 · test · Updated 2023-12-26 23:52:42 +08:00    tqcq

8798
1

ab1b75166f · Merge branch 'master' into gg/ggml_scale · Updated 2023-12-22 04:35:11 +08:00    tqcq

8821
4

7c87353e61 · common : remove incorrect --model-draft default · Updated 2023-12-22 01:17:12 +08:00    tqcq

8829
1

a40f6110f0 · ggml : force F32 precision for ggml_mul_mat · Updated 2023-12-19 22:34:59 +08:00    tqcq

8836
1

3c734f4941 · plamo : testing · Updated 2023-12-18 23:06:05 +08:00    tqcq

8841
13

a462159c43 · cuda : ggml_cuda_op_mul_mat_cublas support F32 precision · Updated 2023-12-18 20:24:29 +08:00    tqcq

8841
16

1b05817112 · decode : fix logits_valid for old API · Updated 2023-12-18 07:49:21 +08:00    tqcq

8842
1

865066621b · llama.swiftui : improve bench · Updated 2023-12-18 01:37:22 +08:00    tqcq

8856
12

f86b9d152c · lookup : minor · Updated 2023-12-17 23:25:28 +08:00    tqcq

8854
9

d2f1e0dacc · Merge branch 'cuda-cublas-opts' into gg/phi-2 · Updated 2023-12-17 14:41:46 +08:00    tqcq

8852
17

b0547d2196 · gguf-py : fail fast on nonsensical special token IDs · Updated 2023-12-16 07:06:42 +08:00    tqcq

8854
1

c8554b80be · Merge branch 'master' of https://github.com/ggerganov/llama.cpp into ceb/fix-cuda-warning-flags · Updated 2023-12-14 01:06:01 +08:00    tqcq

8866
12

e1241d9b46 · metal : switch to execution barriers + fix one of the barriers · Updated 2023-12-13 19:56:45 +08:00    tqcq

8877
47

fc5f334689 · readme : add API change notice · Updated 2023-12-07 18:35:02 +08:00    tqcq

8879
15

af99c6fbfc · llama : remove memory_f16 and kv_f16 flags · Updated 2023-12-06 00:18:16 +08:00    tqcq

8891
26

3cb1c348b3 · metal : try to improve batched decoding · Updated 2023-12-02 04:01:58 +08:00    tqcq

8896
2

eb594c0f7d · alloc : fix build with debug · Updated 2023-12-01 16:46:05 +08:00    tqcq

8920
14

5b74310e6e · build : enable libstdc++ assertions for debug builds · Updated 2023-12-01 07:18:24 +08:00    tqcq

8905
1

bb39b87964 · ggml : restore abort() in GGML_ASSERT · Updated 2023-11-28 08:27:09 +08:00    tqcq

8924
1