CUDA: use mma PTX instructions for FlashAttention (#11583)

* CUDA: use mma PTX instructions for FlashAttention * __shfl_sync workaround for movmatrix * add __shfl_sync to HIP Co-authored-by: Diego Devesa <slarengh@gmail.com>
2025-06-26 19:55:04 +00:00 · 2025-02-02 19:31:09 +01:00
parent 84ec8a58f7
commit 864a0b67a6
29 changed files with 2058 additions and 998 deletions
--- a/ggml/include/ggml.h
+++ b/ggml/include/ggml.h
@ -1775,7 +1775,7 @@ extern "C" {
            struct ggml_tensor  * a,
            int                   k);

-#define GGML_KQ_MASK_PAD 32
+#define GGML_KQ_MASK_PAD 64

    // q:    [n_embd, n_batch,     n_head,    1]
    // k:    [n_embd, n_kv,        n_head_kv, 1]