llama : initial Mamba-2 support (#9126)

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-07-09 13:02:12 +00:00

* llama : initial Mamba-2 support

* ggml : SIMD ggml_ssm_scan for Mamba-2

* ggml : improve ggml_mul speed when masking recurrent states

* llama : support running Mamba-Codestral-7B-v0.1

* llama : fix Mamba-2 conv state saving

* ggml : make the ggml_mul fast broadcast path more consistently formatted

* llama : remove unused variable

* llama : add missing break

* convert_hf : prefer SentencePiece tokenizer for Mamba-2 when present

The tokenzier.json of Mamba-Codestral-7B-v0.1 otherwise requires
workarounds to work correctly.

* llama : avoid redundant state copy for Mamba 1 and 2

* metal : attempt to adapt SSM_SCAN for Mamba-2

* metal : fix SSM_SCAN pipeline scope

* metal : use log and exp instead of log1pf and expf in SSM_SCAN

* metal : remove unused arguments for SSM_SCAN

The max index is 31, so trimming the arguments is necessary.

* metal : add back n_seqs to SSM_SCAN args

Whoops, this is needed for the offset in the concatenated output.

* metal : fix SSM_SCAN state head offset

* metal : fix wrong number of tokens per sequence in SSM_SCAN

* ggml : remove unused fast broadcast path in GGML_MUL

This was initially added because states were masked with ggml_mul,
but this is no longer done and so this "optimisation" is no longer
necessary, or at least not worth the additional code complexity.

* ggml : avoid multiply by D in GGML_OP_SSM_SCAN

This makes the weight buft detection in src/llama.cpp simpler.

* convert : transpose Mamba-2 A, D and reshape SSM_NORM

This breaks existing conversions of Mamba-2 models
to avoid some reshapes.

Not sure if it's a good idea,
but it makes the graph slightly cleaner.

* llama : more appropriate SSM_SCAN and SSM_CONV buft support checks

* convert : fix flake8 lint

* metal : fix confusion between ; and ,

* metal : add missing args for nb references in ssm_scan_f32_group

* metal : single-user mamba2 inference works

* kv-cache : remove const_cast when setting inputs for s_copy

And also fix multi-user inference for recurrent models
by using cell_id instead of i as the kv cell index
when populating s_copy.

* convert : avoid AutoConfig for Mamba and Mamba2 hparams

* kv-cache : allow context shift for recurrent models

* graph : fix recurrent state copies when avoiding copies

Works, but using lambda functions might not be that clean.

* ggml : fix mamba2 ssm scan when compiled with SVE

* ggml-cpu : reorder SVE FMA for consistency with other SIMD arches

* cuda : implement ssm scan for Mamba2

There is still room for improvement, but it works!

* cuda : adapt Mamba1 ssm scan to shape changes from Mamba2

* mamba : fix mismatched new and delete size for llm_build_mamba

Subclasses of llm_graph_context cannot have extra fields,
because the called destructor is not the one from the subclass.
This otherwise would cause problems when runnning Mamba-(1|2) inference
when compiled -DGGML_SANITIZE_ADDRESS=ON

* cuda : graceful fallback for Mamba-1 models with weird embd size

This commit is contained in:

compilade

2025-07-02 13:10:24 -04:00

committed by

GitHub

parent e17991c466

commit 5d46babdc2

24 changed files with 1075 additions and 311 deletions

									
										6

gguf-py/gguf/tensor_mapping.py
									
												View File
												
				@ -477,7 +477,7 @@ class TensorNameMap:

				            "encoder.layers.{bid}.norm2",                   # nomic-bert

				            "transformer.decoder_layer.{bid}.rms_norm_3",   # Grok

				            "encoder.layer.{bid}.mlp.layernorm",            # jina-bert-v2

				            "encoder.layer.{bid}.layer_norm_2"              # jina-v2-code

				            "encoder.layer.{bid}.layer_norm_2",             # jina-v2-code

				        ),

				        MODEL_TENSOR.PER_LAYER_TOKEN_EMBD: (

				@ -574,6 +574,10 @@ class TensorNameMap:

				            "backbone.layers.{bid}.mixer.D",

				        ),

				        MODEL_TENSOR.SSM_NORM: (

				            "backbone.layers.{bid}.mixer.norm",  # mamba2

				        ),

				        MODEL_TENSOR.SSM_OUT: (

				            "model.layers.{bid}.out_proj",

				            "backbone.layers.{bid}.mixer.out_proj",

llama : initial Mamba-2 support (#9126)

6 gguf-py/gguf/tensor_mapping.py Unescape Escape View File

6

gguf-py/gguf/tensor_mapping.py

View File