Instructions to use ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ", trust_remote_code=True) model = AutoModel.from_pretrained("ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ", trust_remote_code=True, device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ
- SGLang
How to use ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ with Docker Model Runner:
docker model run hf.co/ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ
Kimi-K2.5-P48NVFP4-MoESQ
A W4A4 + paired-4:8 sparse compressed checkpoint of
moonshotai/Kimi-K2.5, produced with
MoE-SQ. The MoE expert weights are NVFP4 with paired-4:8 structured sparsity, and they are
stored sparse: only the kept values plus a small mask are on disk, not a dense NVFP4
tensor with zeros in place. The target is NVIDIA Blackwell (SM100) sparse tensor cores.
- Base model: moonshotai/Kimi-K2.5 (MoE, 384 routed experts, 61 layers)
- Compression: NVFP4 W4A4 plus paired-4:8 sparsity on the routed MoE experts
- Effective weight precision: about 2 bits/weight on routed-expert linears
- Checkpoint size: 347.6 GiB, vs 524.8 GiB for the same weights stored as dense NVFP4 (0.66×)
- Kernel: paired-4:8 sparse NVFP4 grouped GEMM (CUTLASS, SM100) through vLLM's
paired48_nvfp4MoE backend
Links
- Paper: coming soon
- Code: IST-DASLab/MoE-SQ. It contains the compression code, the vLLM integration and the kernels.
- Other MoE-SQ checkpoints: ISTA-DASLab/moesq collection
Usage
This checkpoint does not load in upstream vLLM. The MoE-SQ repository installs a patched
vLLM v0.30.0 (the paired48_nvfp4 MoE backend and sparse-storage loader) together with the
kernels. It needs NVIDIA Blackwell (SM100) GPUs and a CUDA toolkit >= 12.8.
git clone --recurse-submodules https://github.com/IST-DASLab/MoE-SQ.git && cd MoE-SQ
bash integrations/vllm/install.sh && source .venv-vllm/bin/activate
# 8x B200, tensor + expert parallel
vllm serve ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ --trust-remote-code \
--tensor-parallel-size 8 --enable-expert-parallel --kv-cache-dtype bfloat16
# or data + expert parallel (requires DeepEP)
vllm serve ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ --trust-remote-code \
--data-parallel-size 8 --enable-expert-parallel --all2all-backend deepep_low_latency \
--kv-cache-dtype bfloat16 --max-num-batched-tokens 256
vLLM selects the backend automatically. See Serving with vLLM for other parallel layouts.
Evaluation
OpenLLM Leaderboard v1: 6-task average
The six tasks are ARC-Challenge (25-shot), HellaSwag (10-shot), MMLU (5-shot), TruthfulQA-MC2 (0-shot), Winogrande (5-shot) and GSM8K (5-shot, greedy). Scores use lm-evaluation-harness with the full test sets, as mean ± std over 3 few-shot seeds (1234, 0, 1). Recovery is computed per seed against dense, then averaged.
| Model | Avg | Recovery vs dense |
|---|---|---|
| Kimi-K2.5 (dense) | 82.30 ± 0.18 | — |
| This model (MoE-SQ) | 79.25 ± 0.07 | 96.30 ± 0.23 % |
| SparseGPT + GPTQ (same W4A4 paired-4:8 target) | 73.88 ± 0.18 | 89.77 ± 0.34 % |
| OBR (same target) | 74.80 ± 0.11 | 90.89 ± 0.24 % |
| JSQ (same target) | 67.84 ± 0.08 | 82.43 ± 0.26 % |
Reasoning
The metric is average pass@1 over N samples per problem (AIME25 N=10, GPQA-Diamond N=5, MATH-500 N=5). Sampling uses temperature 1.0 and top_p 0.95, with up to 65,536 new tokens. Both models were run with this same protocol. ± is the lighteval standard error across problems; it was not recorded for the dense run.
| Benchmark | Kimi-K2.5 (dense) | This model (MoE-SQ) | Recovery |
|---|---|---|---|
| AIME25 | 95.67 | 84.67 ± 5.27 | 88.5 % |
| GPQA-Diamond | 89.49 | 76.46 ± 2.51 | 85.4 % |
| MATH-500 | 96.36 | 96.28 ± 0.74 | 99.9 % |
Compression details
| Field | Value |
|---|---|
| Weights | NVFP4 (E2M1), one FP8-E4M3 scale per 32 dense K elements, FP32 per-tensor global scale |
| Activations | NVFP4, dynamic per-32 group scales, FP32 per-tensor global scale stored per expert linear |
| Sparsity | Paired 4:8 along K: in every 8 consecutive elements (4 pairs), exactly 2 pairs are kept |
| Compressed layers | Routed MoE experts (gate_proj, up_proj, down_proj) in layers 1–60 |
| Left in BF16 | lm_head, embeddings, attention, norms, router (mlp.gate), shared experts, layer-0 dense MLP, vision tower / projector |
| Format | compressed-tensors, nvfp4-pack-quantized, with a paired48_sparse storage marker |
The scale group is 32 because the SM100 sparse NVFP4 MMA requires one scale per 32 dense K elements, which is 16 surviving elements after the 4:8 prune.
Sparse storage layout (paired48_sparse, pair-bitmask v1)
Packed NVFP4 stores two adjacent K elements per byte, so one pair is one byte. A paired-4:8 weight therefore has exactly 2 nonzero bytes in every 4-byte chunk. Each routed-expert linear stores:
| Tensor | dtype / shape | Contents |
|---|---|---|
weight_sparse_packed |
uint8 [out, K/4] |
the 2 kept bytes of every 4, in K order |
weight_sparse_mask |
uint8 [out, K/16] |
4 bits per 4-byte chunk (exactly 2 set; bit i means byte i is kept); the low nibble is the lower-K chunk |
weight_scale |
float8_e4m3 [out, K/32] |
block scales (linear layout) |
weight_global_scale, input_global_scale |
float32 | per-tensor global scales |
quantization_config carries "paired48_sparse": {"layout": "pair-bitmask", "version": 1}.
If a 4-byte chunk has fewer than 2 nonzero bytes, its lowest-index zero bytes are marked as
kept. The encoding is lossless: converting to and from dense NVFP4 is exact, and every tensor
in this upload was checked with a round trip. At load time vLLM rebuilds each layer's dense
packed weight and compresses it into the kernel's own layout. Load-time memory is therefore
the sparse checkpoint plus one dense layer, not the whole dense model.
Recipe (MoE-SQ, arm "gw2")
The full MoE-SQ config for this run is in moe_sq_config.yaml.
- Init: GPTQ, 4-bit, 512 calibration samples, percdamp 0.1, block size 128. Masks come from paired-4:8 pruning.
- Refinement: masks and weight values are learned jointly for 10 epochs on 8,192 mixed
calibration sequences of up to 4,096 tokens, with activations fake-quantized to NVFP4. The
objective is the block-output reconstruction error weighted by the router gate
(
gate_weight_exponent = 2).
Citation
@article{moesq,
title = {TODO},
author = {TODO},
journal = {arXiv preprint arXiv:TODO},
year = {TODO}
}
Contact
For questions, open a discussion here or contact kwanhee.lee@postech.ac.kr.
- Downloads last month
- 109
Model tree for ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ
Base model
moonshotai/Kimi-K2.5