Default Branch

ee4c505a4f · server: add dedup-cache-models preset option (#27346) · Updated 2026-08-19 17:04:26 +08:00

Branches

1621aade61 · server: (router) do not evict busy models · Updated 2026-08-07 20:38:06 +08:00

186
1

89983d5c9c · ggml : update ggml_prec specification · Updated 2026-08-06 19:31:31 +08:00

204
1

1a0e158115 · fix it · Updated 2026-08-06 19:30:01 +08:00

207
2

6751477adf · cont : avoid copy_state · Updated 2026-08-06 15:19:52 +08:00

209
14

f8fa4df82a · test: add duplicated index case for ggml_mul_mat_id · Updated 2026-08-06 07:16:17 +08:00

207
1

f7655b451d · server: fix empty response for /cors-proxy · Updated 2026-08-06 07:09:24 +08:00

207
1

91f82eb1b4 · mtmd: fix granite 4v grid assembly · Updated 2026-08-06 06:42:06 +08:00

208
1

0377fc17ad · sampler : add llama_sampler_copy · Updated 2026-08-05 19:40:53 +08:00

221
2

e77293d4cc · mtmd: weave deepseek-ocr rows in one shot instead of per row (#26615) · Updated 2026-08-05 19:07:53 +08:00

357
2

8045779cff · cleanup · Updated 2026-08-05 16:22:33 +08:00

220
19

5522498343 · metal : port new kernels into the split sources · Updated 2026-08-05 16:22:32 +08:00

220
8

ac6c66699d · vendor : apply patches for subprocess.h · Updated 2026-08-05 07:34:32 +08:00

221
1

58137f0f47 · quant : allow quantization of ffn_gate_inp tensors · Updated 2026-08-04 23:11:42 +08:00

229
1

285941de20 · nits 2 · Updated 2026-08-04 22:49:43 +08:00

228
3

56fe6a8187 · models : fix dflash wo_a reshape on load · Updated 2026-08-04 21:41:42 +08:00

229
1

201e50cc20 · clean up comments · Updated 2026-08-04 21:18:19 +08:00

245
56

eb75c420f0 · Shared default of 64 for history-based samplers, remove context_size · Updated 2026-08-04 14:02:53 +08:00

239
2

3dfb7c7fad · server: simplify get_info result handling · Updated 2026-08-04 00:22:56 +08:00

250
4

7f5903206c · correct to 9931 · Updated 2026-08-03 18:18:50 +08:00

258
3

9d74f019e6 · ggml-meta: make sure to call init_tensor for all new tensors · Updated 2026-08-03 14:22:59 +08:00

261
2