CUDA: refactor mmq, dmmv, mmvq (#7716)

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-07-28 13:20:27 -04:00

* CUDA: refactor mmq, dmmv, mmvq

* fix out-of-bounds write

* struct for qk, qr, qi

* fix cmake build

* mmq_type_traits

This commit is contained in:

Johannes Gäßler

2024-06-05 16:53:00 +02:00

committed by

GitHub

parent 2b3389677a

commit 7d1a378b8f

112 changed files with 1783 additions and 1767 deletions

File diff suppressed because it is too large Load Diff