Default Branch

ee4c505a4f · server: add dedup-cache-models preset option (#27346) · Updated 2026-08-19 17:04:26 +08:00

Branches

32a392fe68 · try a differerent fix · Updated 2024-01-20 06:10:23 +08:00    tqcq

8572
2

4a3bc1522e · py : linting with mypy and isort · Updated 2024-01-20 04:18:58 +08:00    tqcq

8573
3

1453215165 · kompute : fix ggml_add kernel · Updated 2024-01-19 06:09:16 +08:00    tqcq

8689
105

ccc78a200e · hellaswag: speed up even more by parallelizing log-prob evaluation · Updated 2024-01-19 00:25:29 +08:00    tqcq

8589
1

2917e6b528 · Merge branch 'master' into gg/imatrix-gpu-4931 · Updated 2024-01-18 00:43:45 +08:00    tqcq

8596
10

23742deb5b · py : fix padded dummy tokens (I hope) · Updated 2024-01-17 21:44:22 +08:00    tqcq

8615
4

9fd1e83f6d · Use Q4_K for attn_v for Q2_K_S when n_gqa >= 4 · Updated 2024-01-17 18:16:08 +08:00    tqcq

8601
1

49bafe0986 · tests : avoid creating RNGs for each tensor · Updated 2024-01-17 16:40:55 +08:00    tqcq

8604
6

bb9abb5cd8 · imatrix: guard Q4_0/Q5_0 against ffn_down craziness · Updated 2024-01-16 15:56:05 +08:00    tqcq

8618
2

9998ecd191 · llama : add phixtral support (wip) · Updated 2024-01-13 20:24:07 +08:00    tqcq

8648
1

1fb563ebdc · py : try to fix flake stuff · Updated 2024-01-13 19:42:35 +08:00    tqcq

8649
2

9bfcb16fd3 · Add llama enum for IQ2_XS · Updated 2024-01-12 00:24:12 +08:00    tqcq

8698
11

24096933b0 · server : try to fix infill when prompt is empty · Updated 2024-01-09 17:27:29 +08:00    tqcq

8700
1

7216af5c09 · ggml : fix 32-bit ARM compat (cont) · Updated 2024-01-09 16:33:16 +08:00    tqcq

8703
2

d57cb9c294 · passkey : add readme · Updated 2024-01-08 17:13:44 +08:00    tqcq

8713
7

7cfde78190 · llama : remove redundant GQA check · Updated 2024-01-06 22:04:20 +08:00    tqcq

8721
1

9f51f3e695 · metal : opt mul_mm_id · Updated 2024-01-03 02:50:18 +08:00    tqcq

8747
17

4cc78d3873 · ggml : force F32 precision for ggml_mul_mat · Updated 2024-01-02 23:54:56 +08:00    tqcq

8746
1

b5af7ad84f · llama : refactor quantization to avoid <mutex> header · Updated 2024-01-02 21:56:57 +08:00    tqcq

8749
1

120a1a5515 · llama : auto download HF models if URL provided · Updated 2024-01-02 19:29:06 +08:00    tqcq

8750
1