llama.cpp

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-06-29 12:35:16 +00:00

Author	SHA1	Message	Date
Johannes Gäßler	467576b6cc	CMake: default to -arch=native for CUDA build (#10320 )	2024-11-17 09:06:34 +01:00
Diego Devesa	eda7e1d4f5	ggml : fix possible buffer use after free in sched reserve (#9930 )	2024-11-17 08:31:17 +02:00
Georgi Gerganov	24203e9dd7	ggml : inttypes.h -> cinttypes (#0 ) ggml-ci	2024-11-17 08:30:29 +02:00
Georgi Gerganov	5d9e59979c	ggml : adapt AMX to tensor->grad removal (#0 ) ggml-ci	2024-11-17 08:30:29 +02:00
Georgi Gerganov	68fcb4759c	ggml : fix compile warnings (#0 ) ggml-ci	2024-11-17 08:30:29 +02:00
Johannes Gäßler	8a43e940ab	ggml: new optimization interface (ggml/988)	2024-11-17 08:30:29 +02:00
Georgi Gerganov	db4cfd5dbc	llamafile : fix include path (#0 ) ggml-ci	2024-11-16 20:36:26 +02:00
Jeff Bolz	772703c8ff	vulkan: Optimize some mat-vec mul quant shaders (#10296 ) Compute two result elements per workgroup (for Q{4,5}_{0,1}). This reuses the B loads across the rows and also reuses some addressing calculations. This required manually partially unrolling the loop, since the compiler is less willing to unroll outer loops. Add bounds-checking on the last iteration of the loop. I think this was at least partly broken before. Optimize the Q4_K shader to vectorize most loads and reduce the number of bit twiddling instructions.	2024-11-16 07:26:57 +01:00
Dan Johansson	1e58ee1318	ggml : optimize Q4_0 into Q4_0_X_Y repack (#10324 )	2024-11-16 01:53:37 +01:00
Srihari-mcw	74d73dc85c	Make updates to fix issues with clang-cl builds while using AVX512 flags (#10314 )	2024-11-15 22:27:00 +01:00
slaren	883d206fbd	ggml : fix some build issues	2024-11-15 21:45:32 +02:00
Georgi Gerganov	09ecbcb596	cmake : fix ppc64 check (whisper/0) ggml-ci	2024-11-15 15:44:06 +02:00
thewh1teagle	3225008973	ggml : vulkan logs (whisper/2547)	2024-11-15 15:44:06 +02:00
Eve	18429220bd	AVX BF16 and single scale quant optimizations (#10212 ) * use 128 bit loads (i've tried 256->128 to death and its slower) * double accumulator * avx bf16 vec dot * +3% q4_0 inference * +7% tg +5% pp compared to master * slower f16c version, kep for reference * 256b version, also slow. i tried :) * revert f16 * faster with madd * split to functions * Q8_0 and IQ4_NL, 5-7% faster * fix potential overflow (performance reduced) * 16 bit add for q4_0 only * merge	2024-11-15 12:47:58 +01:00
Romain Biessy	5a54af4d4f	sycl: Use syclcompat::dp4a (#10267 ) * sycl: Use syclcompat::dp4a * Using the syclcompat version allow the compiler to optimize the operation with native function * Update news section * Update CI Windows oneAPI version to 2025.0 * Reword doc * Call syclcompat::dp4a inside dpct::dp4a This reverts commit `90cb61d692`.	2024-11-15 11:09:12 +08:00
Charles Xu	1607a5e5b0	backend cpu: add online flow for aarch64 Q4_0 GEMV/GEMM kernels (#9921 ) * backend-cpu: add online flow for aarch64 Q4_0 GEMV/GEMM kernels --------- Co-authored-by: Diego Devesa <slarengh@gmail.com>	2024-11-15 01:28:50 +01:00
Diego Devesa	ae8de6d50a	ggml : build backends as libraries (#10256 ) * ggml : build backends as libraries --------- Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> Co-authored-by: R0CKSTAR <xiaodong.ye@mthreads.com>	2024-11-14 18:04:35 +01:00
Johannes Gäßler	4a8ccb37ad	CUDA: no -sm row for very small matrices (#10185 )	2024-11-14 13:00:15 +01:00
Jeff Bolz	af148c9386	vulkan: Optimize binary ops (#10270 ) Reuse the index calculations across all of src0/src1/dst. Add a shader variant for when src0/src1 are the same dimensions and additional modulus for src1 aren't needed. Div/mod are slow, so add "fast" div/mod that have a fast path when the calculation isn't needed or can be done more cheaply.	2024-11-14 06:22:55 +01:00
Jeff Bolz	66798e42fb	vulkan: Use macros to make the mat mul pipeline creation more concise (#10259 ) Also add vk_matmul_pipeline2 to hold f16/f32 accumulator versions of a pipeline. This isn't really used yet.	2024-11-13 21:59:47 +01:00
Alberto Cabrera Pérez	2e82ffa4af	sycl : Fixes to broken builds and test-backend-ops (#10257 ) * Fixes broken build for the SYCL CUDA backend caused by non-explicit gemm call in outprod (merged in with RWKV6 in Optimize RWKV6 Operator Naming and Implement Multi-core CPU/ SYCL Acceleration #10133) * Marks permuted MUL_MAT as unsupported to be able to run test-backend-ops * Fixes asserts in norm to fix debug builds.	2024-11-13 09:40:57 +00:00
Jeff Bolz	80dd7ff22f	vulkan: Optimize contiguous copies (#10254 ) * tests: Fix memory bandwidth calculation for perf tests Add a flops calculation for flash attention. Add one GGML_OP_CPY perf test. * vulkan: Optimize contiguous copies Add a variant of the copy shader for when the tensors are contiguous. Avoid the complex addressing calculations, and do four elements per invocation to hide some other overhead. Apply similar changes to the scale shader, since scale is always contiguous. Add a "progress bar" for shader compiles.	2024-11-13 07:58:57 +01:00
Jeff Bolz	54ef9cfc72	vulkan: Throttle the number of shader compiles during the build step. (#10222 ) Fixes #9582 Spawning too many concurrent copies of glslc leads to "Failed to create pipes" errors on Linux. This change applies the same throttling we use for multithreaded pipeline creation.	2024-11-11 18:13:51 +01:00
Georgi Gerganov	b0cefea58a	metal : more precise Q*K in FA vec kernel (#10247 )	2024-11-11 08:39:13 +02:00
Jeff Bolz	160687b3ed	vulkan: Fix newly added tests for permuted mul_mat and 1D im2col (#10226 )	2024-11-10 12:37:56 +01:00
Georgi Gerganov	6423c65aa8	metal : reorder write loop in mul mat kernel + style (#10231 ) * metal : reorder write loop * metal : int -> short, style ggml-ci	2024-11-09 11:53:13 +02:00
Georgi Gerganov	39a334a9aa	metal : fix build and some more comments (#10229 )	2024-11-09 11:53:02 +02:00
Georgi Gerganov	bb38cdd8ba	metal : fix F32 accumulation in FA vec kernel (#10232 )	2024-11-09 11:52:45 +02:00
Georgi Gerganov	46323fa9ef	metal : hide debug messages from normal log	2024-11-09 11:21:49 +02:00
SXX	5b359bb1e3	ggml: fix zero division in ‘dne’ calculation in CUDA COUNT_EQUAL operator when ‘ne’ is small (#10213 )	2024-11-09 08:35:46 +01:00
amritahs-ibm	e89213492d	ggml : optimize llamafile cpu matrix multiplication for ppc64le (#10156 ) This change upstreams llamafile's cpu matrix multiplication kernels for ppc64le using MMA builtins for FP32 datatype. This change results in a consistent 90% improvement in input processing time, and 20% to 80% improvement in output processing time, across various batch sizes. The patch is tested with Meta-Lllama-3-8B, Mistral-7B, Llama-2-7B-chat-hf models on a IBM POWER10 machine. Signed-off-by: Amrita H S <amritahs@linux.vnet.ibm.com>	2024-11-09 09:17:50 +02:00
Georgi Gerganov	ec450d3bbf	metal : opt-in compile flag for BF16 (#10218 ) * metal : opt-in compile flag for BF16 ggml-ci * ci : use BF16 ggml-ci * swift : switch back to v12 * metal : has_float -> use_float ggml-ci * metal : fix BF16 check in MSL ggml-ci	2024-11-08 21:59:46 +02:00
Georgi Gerganov	695ad752b2	metal : improve clarity (minor) (#10171 )	2024-11-08 18:37:41 +02:00
Georgi Gerganov	841f27abdb	metal : optimize FA kernels (#10171 ) * ggml : add ggml_flash_attn_ext_get_prec * metal : use F16 precision in FA kernels ggml-ci * metal : minor clean-up * metal : compile-guard bf16 FA kernels ggml-ci * build : remove obsolete compile flag [no ci] * metal : prevent int overflows [no ci] * cuda : disable BF16 FA ggml-ci * metal : fix BF16 requirement for FA kernels ggml-ci * make : clean-up [no ci]	2024-11-08 13:47:22 +02:00
Diego Devesa	97404c4a03	ggml : add ggml-cpu.h to the public headers (#10204 )	2024-11-07 18:16:08 +01:00
snadampal	2319126a70	fix q4_0_8_8 format for corrupted tokens issue (#10198 ) Co-authored-by: EC2 Default User <ec2-user@ip-172-31-62-167.us-west-2.compute.internal>	2024-11-07 09:02:08 +01:00
Zhiyuan Li	3bcd40b3c5	Optimize RWKV6 Operator Naming and Implement Multi-core CPU/ SYCL Acceleration (#10133 ) * rwkv6: rename to wkv6 * rwkv6: support avx2 avx512 armv8 armv9 * rwkv6: update cuda file name * rwkv6: rename params * wkv on sycl * sycl: add some ops * sycl: Enhance OP support judgment * wkv6: drop armv9 and tranfer to GGML style ggml-ci * sync : ggml * update the function to use appropriate types * fix define error * Update ggml/src/ggml-cpu.c * add appropriate asserts * move element-wise functions outside * put the declaration outside the loop * rewrite to be more inline with the common pattern for distributing threads * use recommended way GGML_TENSOR_LOCALS --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> Co-authored-by: Diego Devesa <slarengh@gmail.com> Co-authored-by: Plamen Minev <pacominev@gmail.com> Co-authored-by: Yuri Khrustalev <ykhrustalev@users.noreply.github.com> Co-authored-by: Meng, Hengyu <airdldl@163.com>	2024-11-07 15:19:10 +08:00
Georgi Gerganov	5c333e0140	metal : add BF16 support (#8439 ) * ggml : add initial BF16 support ggml-ci * metal : add mul_mat_id BF16 support ggml-ci * metal : check for bfloat support on the Metal device ggml-ci * metal : better var names [no ci] * metal : do not build bfloat kernels when not supported ggml-ci * metal : try to fix BF16 support check ggml-ci * metal : this should correctly check bfloat support	2024-11-06 19:53:51 +02:00
Diego Devesa	94d8cb8be1	metal : fix from ptr buffer name (#10189 )	2024-11-06 12:10:07 +01:00
Georgi Gerganov	1dc04b2dee	ggml : adjust is_first_call init value (#10193 ) ggml-ci	2024-11-06 11:20:10 +02:00
Georgi Gerganov	a1eaf6a960	metal : add quantized FA support (#10149 ) * metal : add quantized FA (vec) support ggml-ci * metal : add quantized FA (non-vec) support * metal : fix support check ggml-ci * metal : clean-up * metal : clean-up (cont) * metal : fix shared memory calc + reduce smem + comments * metal : float-correctness * metal : minor [no ci]	2024-11-06 10:24:23 +02:00
Diego Devesa	a9e8a9a030	ggml : fix arch check in bf16_to_fp32 (#10164 )	2024-11-04 23:17:01 +01:00
Eve	3407364776	Q6_K AVX improvements (#10118 ) * q6_k instruction reordering attempt * better subtract method * should be theoretically faster small improvement with shuffle lut, likely because all loads are already done at that stage * optimize bit fiddling * handle -32 offset separately. bsums exists for a reason! * use shift * Update ggml-quants.c * have to update ci macos version to 13 as 12 doesnt work now. 13 is still x86	2024-11-04 23:06:31 +01:00
Diego Devesa	d5a409e57f	ggml : fix gelu tables initialization (#10172 )	2024-11-04 20:06:58 +01:00
Diego Devesa	401558b7ba	ggml : fix q4xx mat mul, increase ggml_aligned_malloc alignment (#10167 )	2024-11-04 17:34:08 +01:00
snadampal	6a066b9978	fix build break on arm64 linux (#10166 ) This fixes the build break from the recent changes to move the CPU backend to separate files https://github.com/ggerganov/llama.cpp/pull/10144	2024-11-04 16:08:33 +01:00
Diego Devesa	ea02c753eb	cuda : clear error after changing peer access (#10153 )	2024-11-04 13:10:23 +01:00
Georgi Gerganov	05697f670b	metal : simplify f16 and f32 dequant kernels (#0 )	2024-11-04 13:49:34 +02:00
Georgi Gerganov	f8e58135cf	metal : move dequantize templates to beginning of MSL source (#0 )	2024-11-04 13:44:06 +02:00
leo-pony	329ed914c9	CANN: adjust backend registry refactor. (#10158 ) remove buffer->iface.get_name that used in cann as it was removed in backend registry refactor PR.	2024-11-04 19:08:22 +08:00

... 7 8 9 10 11 ...

718 Commits