What is the reason to use this model instead of simple mlx 4 bit?

#1
by zaskara - opened

Please explain; I didn't get it. On the provided benchmark, it has 89 vs. 91 for the basic 4-bit version of this model. And it seems to have the same size. Then, what is the reason to use it instead of basic 4 bit?

MLX Community org

This was still under active development the calibration mix has been improved and made more diverse. New benchmarks were also added that show clear improvement over uniform 4 bit quants.

How does this benchmark against a Qwen3.6-35B-A3B-oQ6e model. I got similar speed and quality results for both.

MLX Community org

oQ (Jundot's oMLX) and OptiQ are the same family of idea: mixed-precision, different bit-widths per layer instead of uniform. The difference is the signal each uses to place the bits. oQ uses an importance matrix (imatrix); OptiQ uses per-layer KL-divergence sensitivity. The writeup on that is here: https://mlx-optiq.com/blog/not-all-layers-are-equal

Both allocate bits by importance rather than uniformly, so similar quality between the two makes sense. One note on the numbers: the OptiQ Capability Score is just the mean of six tasks, and it's a different suite than the 83.88 on the oQ card, so those two aren't directly comparable.

Where OptiQ tends to pay off is less the raw 4-bit quality and more what ships around the quant: a bundled MTP head for ~1.4x speculative decode, per-layer mixed-precision KV cache so long context doesn't blow up memory, and it drops into optiq serve (OpenAI + Anthropic-compatible) with sensitivity-aware LoRA fine-tuning on the same checkpoint.

May i suggest to review https://github.com/defai-digital/axquant
I think it would be helpful to promote local llm with mlx

Sign up or log in to comment