Instructions to use ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ") model = AutoModelForMultimodalLM.from_pretrained("ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ
- SGLang
How to use ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ with Docker Model Runner:
docker model run hf.co/ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ
Qwen3.5-397B-A17B-P48NVFP4-MoESQ
A W4A4 + paired-4:8 sparse compressed checkpoint of
Qwen/Qwen3.5-397B-A17B, produced with
MoESQ. The MoE expert weights are NVFP4 with paired-4:8 structured sparsity. They are
stored sparse: only the kept values plus a small mask are on disk, not a dense NVFP4
tensor with zeros in place. The target is NVIDIA Blackwell (SM100) sparse tensor cores.
- Base model: Qwen/Qwen3.5-397B-A17B (MoE, 512 routed experts with 10 active plus a shared expert, 60 layers with hybrid linear/full attention)
- Compression: NVFP4 W4A4 plus paired-4:8 sparsity on the routed MoE experts
- Effective weight precision: about 2 bits/weight on routed-expert linears
- Checkpoint size: 155.3 GiB, vs 222.7 GiB for the same weights stored as dense NVFP4 (0.70×)
- Kernel: paired-4:8 sparse NVFP4 grouped GEMM (CUTLASS, SM100) through vLLM's
paired48_nvfp4MoE backend
Links
- Paper: coming soon
- Code: IST-DASLab/MoESQ. It contains the compression code, the vLLM integration and the kernels.
- Other MoESQ checkpoints: ISTA-DASLab/moesq collection
Usage
This checkpoint does not load in upstream vLLM. The MoESQ repository installs a patched
vLLM v0.30.0 (the paired48_nvfp4 MoE backend and sparse-storage loader) together with the
kernels. It needs NVIDIA Blackwell (SM100) GPUs and a CUDA toolkit >= 12.8.
git clone --recurse-submodules https://github.com/IST-DASLab/MoESQ.git && cd MoESQ
bash integrations/vllm/install.sh && source .venv-vllm/bin/activate
# 4x B200, tensor + expert parallel
vllm serve ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ \
--tensor-parallel-size 4 --enable-expert-parallel
vLLM selects the backend automatically.
Evaluation
OpenLLM Leaderboard v1: 6-task average
The six tasks are ARC-Challenge (25-shot), HellaSwag (10-shot), MMLU (5-shot), TruthfulQA-MC2 (0-shot), Winogrande (5-shot) and GSM8K (5-shot, greedy). Scores use lm-evaluation-harness with the full test sets, as mean ± std over 3 few-shot seeds (1234, 0, 1). Recovery is computed per seed against dense, then averaged.
| Model | Avg | Recovery vs dense |
|---|---|---|
| Qwen3.5-397B-A17B (dense) | 80.24 ± 0.14 | — |
| This model (MoESQ) | 75.07 ± 0.21 | 93.55 % |
| SparseGPT + GPTQ (same W4A4 paired-4:8 target) | 70.61 ± 0.25 | 88.00 % |
| OBR (same target) | 73.42 ± 0.36 | 91.50 % |
| JSQ (same target) | 62.96 ± 0.14 | 78.46 % |
Most of the gap to the baselines comes from GSM8K, which moves by about ±1.3 points across few-shot seeds. Reasoning benchmarks have not been run for this model yet.
Compression details
| Field | Value |
|---|---|
| Weights | NVFP4 (E2M1), one FP8-E4M3 scale per 32 dense K elements, FP32 per-tensor global scale |
| Activations | NVFP4, dynamic per-32 group scales, FP32 per-tensor global scale stored per expert linear |
| Sparsity | Paired 4:8 along K: in every 8 consecutive elements (4 pairs), exactly 2 pairs are kept |
| Compressed layers | Routed MoE experts (gate_proj, up_proj, down_proj) in all 60 layers |
| Left in BF16 | lm_head, embeddings, full and linear attention, norms, router (mlp.gate), shared experts, vision tower / merger, the MTP head |
| Format | compressed-tensors, nvfp4-pack-quantized, with a paired48_sparse storage marker |
The scale group is 32 because the SM100 sparse NVFP4 MMA requires one scale per 32 dense K elements, which is 16 surviving elements after the 4:8 prune.
Sparse storage layout (paired48_sparse, pair-bitmask v1)
Packed NVFP4 stores two adjacent K elements per byte, so one pair is one byte. A paired-4:8 weight therefore has exactly 2 nonzero bytes in every 4-byte chunk. Each routed-expert linear stores:
| Tensor | dtype / shape | Contents |
|---|---|---|
weight_sparse_packed |
uint8 [out, K/4] |
the 2 kept bytes of every 4, in K order |
weight_sparse_mask |
uint8 [out, K/16] |
4 bits per 4-byte chunk (exactly 2 set; bit i means byte i is kept); the low nibble is the lower-K chunk |
weight_scale |
float8_e4m3 [out, K/32] |
block scales (linear layout) |
weight_global_scale, input_global_scale |
float32 | per-tensor global scales |
quantization_config carries "paired48_sparse": {"layout": "pair-bitmask", "version": 1}.
The encoding is lossless: converting to and from dense NVFP4 is exact, and every tensor in
this upload was checked with a round trip. At load time vLLM rebuilds each layer's dense
packed weight and compresses it into the kernel's own layout.
Recipe (MoESQ, arm "gw2")
The full MoESQ config for this run is in moe_sq_config.yaml.
- Init: GPTQ, 4-bit, 512 calibration samples, percdamp 0.1, block size 128. Masks come from paired-4:8 pruning.
- Refinement: masks and weight values are learned jointly for 10 epochs on 8,192 mixed
calibration sequences of up to 4,096 tokens, with activations fake-quantized to NVFP4. The
objective is the block-output reconstruction error weighted by the router gate
(
gate_weight_exponent = 2). Learning rates are 2e-4 for masks and 2.5e-5 for weights.
Citation
@article{moesq,
title = {Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts},
author = {TODO},
journal = {arXiv preprint arXiv:TODO},
year = {TODO}
}
Contact
For questions, open a discussion here or contact kwanhee.lee@postech.ac.kr.
- Downloads last month
- 31
Model tree for ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ
Base model
Qwen/Qwen3.5-397B-A17B