Disaggregated Quantization: Specializing LLM Prefill and Decode Paper • 2609.26333 • Published 10 days ago • 91
One Simple Trick for Improving the Performance of Energy-Limited Local Inference and Training Paper • 2609.11936 • Published Jul 2
WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms Paper • 2609.38121 • Published 3 days ago • 14
WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms Paper • 2609.38121 • Published 3 days ago • 14
ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF Image-Text-to-Text • 117B • Updated 3 days ago • 261k • 187
ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF Image-Text-to-Text • 177B • Updated 3 days ago • 1.14M • 438
MoESQ Collection Paired-4:8 + NVFP4 W4A4 expert compressed MoE models for NVIDIA Blackwell Sparse Tensor Cores. • 3 items • Updated 3 days ago • 1
MoESQ Collection Paired-4:8 + NVFP4 W4A4 expert compressed MoE models for NVIDIA Blackwell Sparse Tensor Cores. • 3 items • Updated 3 days ago • 1
MoESQ Collection Paired-4:8 + NVFP4 W4A4 expert compressed MoE models for NVIDIA Blackwell Sparse Tensor Cores. • 3 items • Updated 3 days ago • 1
Disaggregated Quantization: Specializing LLM Prefill and Decode Paper • 2609.26333 • Published 10 days ago • 91