Instructions to use Qwen/Qwen3-ASR-0.6B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Inference
- Notebooks
- Google Colab
- Kaggle
Remove the inactive `temperature` from generation_config.json
generation_config.json sets both "do_sample": false and "temperature": 0.000001.
transformers accepts that pair on load (warning only) but refuses to write it back out:GenerationConfig.save_pretrained() runs validate(strict=True), which raises for a
sampling-only flag while do_sample is not True. Any load -> modify -> save round trip on
this repo therefore fails -- quantization export, a dtype change followed by save_pretrained,
format conversion. It fails after config.json has been written, so the output directory is
left holding a config and no weights.
Reproduced on transformers 5.15.1, on main @ 945dac9117cb54196888c0e6c08035792a98c485, and on
4.57.6 (the version stamped in this repo's own config.json):
ValueError: GenerationConfig is invalid:
- `temperature`: `do_sample` is not set to `True`. However, `temperature` is set to `1e-06` --
this flag is only used in sample-based generation modes. You should set `do_sample=True` or
unset `temperature`.
Qwen/Qwen3-ASR-0.6B-hf and Qwen/Qwen3-ASR-1.7B-hf already ship generation_config.json
with no temperature key. This PR realigns the legacy repo with them and changes nothing else.
Why this is a no-op for inference
I checked every documented path before proposing the deletion:
| Path | Where temperature comes from |
Effect of the key today |
|---|---|---|
qwen-asr, transformers backend |
the model's generation_config |
ignored -- do_sample: false, so the temperature warper is never constructed |
qwen-asr, vLLM backend |
SamplingParams(temperature=0.0, ...), hardcoded in qwen_asr/inference/qwen3_asr.py |
never read |
qwen-asr-serve / vllm serve -> /v1/audio/transcriptions |
TranscriptionRequest.temperature is a non-optional field defaulting to 0.0, so the generation_config fallback inside to_sampling_params() is unreachable |
never read |
raw LLM.generate() with no sampling_params, or /v1/completions |
generation_config |
read, then clamped up to 0.01 in SamplingParams.__post_init__ (_MAX_TEMP = 1e-2) before the _SAMPLING_EPS = 1e-5 greedy test -- so this is T=0.01 sampling, not greedy |
On the ASR paths the key is inert; on the remaining paths it does not produce greedy decoding
anyway. After this PR those last two fall back to vLLM's own temperature=1.0 default.
If pinning greedy decoding for direct vllm serve is the intent, "temperature": 0.0 is the
value that actually does it (0 < 0.0 is false, so the clamp is skipped and the greedy branch
fires) -- though that form does not fix the transformers round trip. Happy to switch this PR to
it if you prefer.
- Coordination: https://github.com/QwenLM/Qwen3-ASR/issues/210