Inference Providers
Active filters: modelopt
NVFP4/Qwen3-30B-A3B-Instruct-2507-FP4
Text Generation
• 16B • Updated • 2.21k
• 11
NVFP4/Qwen3-Coder-30B-A3B-Instruct-FP4
Text Generation
• 16B • Updated • 19.3k
• 36
gesong2077/Qwen3-32B-NVFP4
19B • Updated • 13
• 1
54B • Updated • 17
nvidia/Phi-4-multimodal-instruct-NVFP4
4B • Updated • 278k
• 13
nvidia/Phi-4-multimodal-instruct-FP8
6B • Updated • 338
• 7
nvidia/Phi-4-reasoning-plus-FP8
15B • Updated • 893
• 7
nvidia/Phi-4-reasoning-plus-NVFP4
8B • Updated • 251k
• 11
nvidia/Llama-3.1-8B-Instruct-NVFP4
5B • Updated • 201k
• 15
Text Generation
• 5B • Updated • 159k
• 21
Text Generation
• 8B • Updated • 10.2k
• 6
Text Generation
• 15B • Updated • 78.4k
• 6
Text Generation
• 17B • Updated • 48.9k
• 17
nvidia/Qwen2.5-VL-7B-Instruct-FP8
Text Generation
• 8B • Updated • 1.04k
• 8
nvidia/Qwen2.5-VL-7B-Instruct-NVFP4
Text Generation
• 5B • Updated • 31.4k
• 15
nuphoto-ian/Qwen3-8B-QAT-NVFP4
5B • Updated • 7
txn545/Qwen3-Coder-30B-A3B-Instruct-NVFP4
16B • Updated • 11
• 1
shanjiaz/gpt-oss-120b-nvfp4-modelopt
59B • Updated • 2.24k
• 5
shanjiaz/gpt-oss-20b-nvfp4-modelopt
11B • Updated • 120
• 2
nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1-FP4-QAD
Image-Text-to-Text
• 6B • Updated • 149
• 15
baseten-admin/glm-4.6-fp4
177B • Updated • 8
baseten-admin/glm-4.6-fp8
353B • Updated • 4
baseten-admin/glm-4.6-fp4-mlp
183B • Updated • 5
shinedays1993/Qwen3-30B-A3B-nvfp4
16B • Updated • 8
shinedays1993/Qwen3-32B-nvfp4
17B • Updated • 13
Beambutbetter/Deepseek-V2-Lite-16B-NVFP4
Text Generation
• 8B • Updated • 144
• 3
ramblingpolymath/Qwen3-4B-Instruct-2507
2B • Updated • 5
literid/Qwen3-Coder-480B-A35B-Instruct_nvfp4_kv_fp8
241B • Updated • 3
DevQuasar/DeepSeek-R1-Distill-Llama-8B_nvfp4
Text Generation
• 5B • Updated • 8
DevQuasar/Qwen.Qwen3-4B-Thinking-2507_nvfp4
Text Generation
• 2B • Updated • 7