Instructions to use SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S # Run inference directly in the terminal: llama cli -hf SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S # Run inference directly in the terminal: llama cli -hf SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S # Run inference directly in the terminal: ./llama-cli -hf SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S
Use Docker
docker model run hf.co/SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S
- LM Studio
- Jan
- vLLM
How to use SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S
- SGLang
How to use SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Ollama
How to use SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF with Ollama:
ollama run hf.co/SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S
- Unsloth Desktop
- Pi
How to use SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF with Docker Model Runner:
docker model run hf.co/SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S
- Lemonade
How to use SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S
Run and chat with the model
lemonade run user.Qwen3-VL-32B-Thinking-heretic-GGUF-Q4_K_S
List all available models
lemonade list
- Hermes Agent
How to use SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "SerialKicked/Qwen3-VL-32B-Thinking-heretic-GGUF:Q4_K_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quant discussions
Sadly... I got a llama_model_load: error loading model: error loading model architecture: unknown model architecture: 'qwen3vl'
It seems that GGUF files for Qwen3-VL that are created directly with/using llama.cpp don't work when Ollama attempts to load them...
Wow... I moved from Ollama to llama.cpp and it’s like stepping into a different universe. The model works great; it’s incredible and beyond words. Thanks!!!
My bad, I'm a bit late, and I'm glad you figured out yourself!
Your Ollama might need to be updated, Qwen3vl is a very recent architecture. But at this point, if you moved to llama.cpp, might as well keep using it, it tends to have all the cutting edge features and runs pretty fast.
Thanks! Your model got me to switch to llama.cpp, and there’s no going back. :)
One thing... Is there any chance you could apply the heretic procedure to gpt-oss-20b? Heretically ablated gpt-oss-20b models available in huggingface are broken beyond recognition. Your Qwen3-VL model performs extremely well, and I’m curious whether your approach could unlock gpt-oss-20b potential. Pretty please?
Sorry, I noticed that you only quantized coder3101's previous Qwen3-VL-32B-Thinking-heretic. Please, forget my last message. ;)
Yeah, sorry, I don't have the compute to do more than just quantization. No problem.
Recently, Coder3101 kindly created an heretic version of gpt-oss-20b that you can find at https://huggingface.co/coder3101/gpt-oss-20b-heretic ... you might want to quantize it... Pretty please? ;)
I probably can. I'll give it a go and leave it working today. I'll have to run some tests on it too.
But I have to warn you, quantizing GPT-OSS-20B is much more damaging than with other models. Beyond the practicality of having a GGUF running on llama.cpp-powered backends, I wouldn't recommend it.
Why? Well, The official GPT-OSS model is already a 4-bit model (MXFP4), it was directly trained in it. So doing a GGUF pass will be like quantizing it twice, and trice with heretic. With the base model in this MXFP4 format, it's already a lot less tolerant to format conversion. The guy who used heretic on it had to convert it back to BF16 (and despite being "larger", it's still an approximation of the actual MXFP4 values). Now, I'd have to turn the thing back into a GGUF. And then quantize it back to something useable, memory-wise, making it an approximation of an approximation (which is usually a big no-no in this field).
That's why most quant you'll find behave weirdly or are plain unusable, GPT-OSS needs special care. Modern llama.cpp tools can convert the model keeping its MXFP4 model intact, which I would have done here, but given that coder3101 already converted and damaged the weights (I can only hope it was because it wasn't possible to use Heretic on it otherwise), I can't undo that damage.
Given I'm already downloading the files, I'll attempt a few different quants, but don't expect miracles. Q4 will probably behave very badly, Q6 or Q8 might be useable, though.
Thanks for the insightful answer. Now I truly see that GGUF’s quantization blocks most probably rarely align with MXFP4's original 32-element scaling blocks. That's what's destroying coherence entirely...
Pretty much. I still made the Quants. Q8 probably works okay, anything below will be degraded in different ways.
https://huggingface.co/SerialKicked/GPT-OSS-20B-Heretic-GGUF/
Cheers.
You're awesome. I'll give them a try now.
During testing I found that Q8_0 performs well, but for the MXFP4, the successive quantizations somehow compromised its natural defenses... and as a result the MXFP4 quant actually performs better and is astonishingly fast.
I’ll run more tests, but so far the pairing Coder3101 (heretic) + SerialKicked (quant) is outstanding. In the end you produced a gpt-oss-20b that truly works and is blazingly fast. No small feat.
Thanks .
I didn't try them much (let alone fully test them), but so far the Q5 and the Q4NL both ended up doing weird generations ("And so, ........, .......", losing focus, fucking up the instruction format) a bit too often for my tastes. This is interesting because it's the ones using mixed quantization (Q8 for attention layers and Q5_1/IQ4NL for the other ones), so I suspect it's linked. Q8 and MXFP4 don't seem to have those issues and they both have a fixed a quantization scheme.
I might attempt to make a pure Q5 or Q6 to see if those models behave correctly or not just to test the theory.
Edit: Nope. Pure Q5 also lead to funky broken generations (more rarely than with the other ones, but still). So, it's not the mixed quant scheme being the issue. Welp I guess it's stuck to Q8 or MXFP4. In a way it makes sense that a base model that's already in 4bit natively wouldn't be a fan to be quantized to something that's not a multiple of 4.