Qwen2.5-7B-base-mla-topk8-rank512

This repository contains the attention-only incremental weights for an MLAfication conversion of Qwen/Qwen2.5-7B. It is not a standalone full-model checkpoint.

Variant

Field Value
Base model Qwen/Qwen2.5-7B
Base revision d149729398750b98c0af14eb82c78cfe92750796
MLA latent rank 512
RoPE dimensions per KV head 8
Training checkpoint step 180
Training variant stage1+2-distill-QKV
Delta tensors 336
Delta size 1.47 GiB
SHA-256 09e3c288e055711f3413457e32f0f3397ead889d472b15ddabb625c6f96dc684

The training run froze every parameter outside attention and also froze each attention output projection. The uploaded delta therefore contains exactly the parameters matching "attn" in name and "o_proj" not in name; embeddings, MLP, normalization outside attention, output projections, and LM head are omitted. Frozen tensors from the historical full checkpoint were verified exactly against the pinned base weights after casting the base tensors to the checkpoint dtype.

Loading

Use the MLAfication implementation to instantiate the MLAfication architecture from the pinned base model, then load model.safetensors with strict=False:

from safetensors.torch import load_file

delta = load_file("model.safetensors")
load_result = mla_model.load_state_dict(delta, strict=False)

The architecture must be patched before loading the delta. Loading this repository directly with vanilla AutoModelForCausalLM.from_pretrained(...) is not supported. config.json contains the layer-wise RoPE indices and sanitized MLAfication metadata.

Files

  • model.safetensors: attention-only delta weights
  • config.json: base architecture, RoPE indices, and MLAfication metadata
  • manifest.json: provenance, byte size, tensor count, and checksum

License

The weights follow the Apache 2.0 license of the Qwen2.5 base model.

Downloads last month
23
Safetensors
Model size
0.8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OpenMOSS-Team/Qwen2.5-7B-base-mla-topk8-rank512

Base model

Qwen/Qwen2.5-7B
Finetuned
(970)
this model