SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
Paper • 2211.10438 • Published • 6
INT8-quantized ONNX of jinaai/jina-embeddings-v5-text-small-retrieval, built with the SmoothQuant α=0.8 outlier-migration recipe (same approach used for the F2LLM and Octen Int8 exports).
1.06 GB (50 % memory of FP32).
SmoothQuant α=0.8 + per-channel dynamic INT8 keeps cos_min ≈ 0.99 vs PyTorch FP32 reference on the 6-text canonical probe (vanilla quantize_dynamic collapses on the same probe — same Qwen3-style activation outlier issue documented for Octen / F2LLM Int8).
| File | Size | Description |
|---|---|---|
model.int8.onnx |
~5 MB | ONNX header (external data) |
model.int8.onnx.data |
~1 GB | SmoothQuant α=0.8 + INT8 weights |
tokenizer.json, config.json, tokenizer_config.json |
small | tokenizer + model config |
scripts/smoothquant_onnx.py --alpha 0.8 — graph rewrite that migrates activation outliers into weights via Mul nodes (mathematically (X * 1/s) @ (W * s) = X @ W).scripts/quant_smoothed_int8.py — standard ORT dynamic INT8 on the smoothed FP32.See Xiao et al. 2023, SmoothQuant for the underlying technique.
let embedder = TextEmbedding::try_new(
InitOptions::new(EmbeddingModel::JinaEmbeddingsV5SmallInt8))?;
Pooling: last-token. Asymmetric retrieval prefixes "Query: " / "Document: " are recommended.
Apache 2.0, inherited from the base model.
jinaai.cc-by-nc-4.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not.Base model
Qwen/Qwen3-0.6B-Base