Kimi-K2.5-P48NVFP4-MoESQ

A W4A4 + paired-4:8 sparse compressed checkpoint of moonshotai/Kimi-K2.5, produced with MoE-SQ. The MoE expert weights are NVFP4 with paired-4:8 structured sparsity, and they are stored sparse: only the kept values plus a small mask are on disk, not a dense NVFP4 tensor with zeros in place. The target is NVIDIA Blackwell (SM100) sparse tensor cores.

  • Base model: moonshotai/Kimi-K2.5 (MoE, 384 routed experts, 61 layers)
  • Compression: NVFP4 W4A4 plus paired-4:8 sparsity on the routed MoE experts
  • Effective weight precision: about 2 bits/weight on routed-expert linears
  • Checkpoint size: 347.6 GiB, vs 524.8 GiB for the same weights stored as dense NVFP4 (0.66×)
  • Kernel: paired-4:8 sparse NVFP4 grouped GEMM (CUTLASS, SM100) through vLLM's paired48_nvfp4 MoE backend

Links

Usage

This checkpoint does not load in upstream vLLM. The MoE-SQ repository installs a patched vLLM v0.30.0 (the paired48_nvfp4 MoE backend and sparse-storage loader) together with the kernels. It needs NVIDIA Blackwell (SM100) GPUs and a CUDA toolkit >= 12.8.

git clone --recurse-submodules https://github.com/IST-DASLab/MoE-SQ.git && cd MoE-SQ
bash integrations/vllm/install.sh && source .venv-vllm/bin/activate

# 8x B200, tensor + expert parallel
vllm serve ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ --trust-remote-code \
  --tensor-parallel-size 8 --enable-expert-parallel --kv-cache-dtype bfloat16

# or data + expert parallel (requires DeepEP)
vllm serve ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ --trust-remote-code \
  --data-parallel-size 8 --enable-expert-parallel --all2all-backend deepep_low_latency \
  --kv-cache-dtype bfloat16 --max-num-batched-tokens 256

vLLM selects the backend automatically. See Serving with vLLM for other parallel layouts.

Evaluation

OpenLLM Leaderboard v1: 6-task average

The six tasks are ARC-Challenge (25-shot), HellaSwag (10-shot), MMLU (5-shot), TruthfulQA-MC2 (0-shot), Winogrande (5-shot) and GSM8K (5-shot, greedy). Scores use lm-evaluation-harness with the full test sets, as mean ± std over 3 few-shot seeds (1234, 0, 1). Recovery is computed per seed against dense, then averaged.

Model Avg Recovery vs dense
Kimi-K2.5 (dense) 82.30 ± 0.18 —
This model (MoE-SQ) 79.25 ± 0.07 96.30 ± 0.23 %
SparseGPT + GPTQ (same W4A4 paired-4:8 target) 73.88 ± 0.18 89.77 ± 0.34 %
OBR (same target) 74.80 ± 0.11 90.89 ± 0.24 %
JSQ (same target) 67.84 ± 0.08 82.43 ± 0.26 %

Reasoning

The metric is average pass@1 over N samples per problem (AIME25 N=10, GPQA-Diamond N=5, MATH-500 N=5). Sampling uses temperature 1.0 and top_p 0.95, with up to 65,536 new tokens. Both models were run with this same protocol. ± is the lighteval standard error across problems; it was not recorded for the dense run.

Benchmark Kimi-K2.5 (dense) This model (MoE-SQ) Recovery
AIME25 95.67 84.67 ± 5.27 88.5 %
GPQA-Diamond 89.49 76.46 ± 2.51 85.4 %
MATH-500 96.36 96.28 ± 0.74 99.9 %

Compression details

Field Value
Weights NVFP4 (E2M1), one FP8-E4M3 scale per 32 dense K elements, FP32 per-tensor global scale
Activations NVFP4, dynamic per-32 group scales, FP32 per-tensor global scale stored per expert linear
Sparsity Paired 4:8 along K: in every 8 consecutive elements (4 pairs), exactly 2 pairs are kept
Compressed layers Routed MoE experts (gate_proj, up_proj, down_proj) in layers 1–60
Left in BF16 lm_head, embeddings, attention, norms, router (mlp.gate), shared experts, layer-0 dense MLP, vision tower / projector
Format compressed-tensors, nvfp4-pack-quantized, with a paired48_sparse storage marker

The scale group is 32 because the SM100 sparse NVFP4 MMA requires one scale per 32 dense K elements, which is 16 surviving elements after the 4:8 prune.

Sparse storage layout (paired48_sparse, pair-bitmask v1)

Packed NVFP4 stores two adjacent K elements per byte, so one pair is one byte. A paired-4:8 weight therefore has exactly 2 nonzero bytes in every 4-byte chunk. Each routed-expert linear stores:

Tensor dtype / shape Contents
weight_sparse_packed uint8 [out, K/4] the 2 kept bytes of every 4, in K order
weight_sparse_mask uint8 [out, K/16] 4 bits per 4-byte chunk (exactly 2 set; bit i means byte i is kept); the low nibble is the lower-K chunk
weight_scale float8_e4m3 [out, K/32] block scales (linear layout)
weight_global_scale, input_global_scale float32 per-tensor global scales

quantization_config carries "paired48_sparse": {"layout": "pair-bitmask", "version": 1}. If a 4-byte chunk has fewer than 2 nonzero bytes, its lowest-index zero bytes are marked as kept. The encoding is lossless: converting to and from dense NVFP4 is exact, and every tensor in this upload was checked with a round trip. At load time vLLM rebuilds each layer's dense packed weight and compresses it into the kernel's own layout. Load-time memory is therefore the sparse checkpoint plus one dense layer, not the whole dense model.

Recipe (MoE-SQ, arm "gw2")

The full MoE-SQ config for this run is in moe_sq_config.yaml.

  • Init: GPTQ, 4-bit, 512 calibration samples, percdamp 0.1, block size 128. Masks come from paired-4:8 pruning.
  • Refinement: masks and weight values are learned jointly for 10 epochs on 8,192 mixed calibration sequences of up to 4,096 tokens, with activations fake-quantized to NVFP4. The objective is the block-output reconstruction error weighted by the router gate (gate_weight_exponent = 2).

Citation

@article{moesq,
  title   = {TODO},
  author  = {TODO},
  journal = {arXiv preprint arXiv:TODO},
  year    = {TODO}
}

Contact

For questions, open a discussion here or contact kwanhee.lee@postech.ac.kr.

Downloads last month
109
Safetensors
Model size
646B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ

Quantized
(40)
this model

Collection including ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ