MLX is not shipped for this model
The MLX build exported and ran correctly, but Core ML is far faster here, so no
.pte or config.json is published in this directory. Use the Core ML variant
on iOS, or XNNPACK.
Measured
iPhone 16, same 640x640 image, 20 timed iterations after 3 warmups, median:
| variant | forward | end to end | peak memory | file |
|---|---|---|---|---|
coreml/rfdetr_keypoint_preview_coreml_fp16.pte |
137.6 ms | 141.6 ms | 248 MB | 75 MB |
| MLX fp32 (withdrawn) | 378.7 ms | 382.2 ms | 1153 MB | 145 MB |
So MLX was 2.70x slower on 4.6x the memory.
Why
RF-DETR is a convolutional backbone feeding a transformer decoder. Core ML lowers that backbone to the ANE, which is built for exactly this shape of network. MLX is a Metal GPU path whose advantage shows up in LLM decode, which is memory-bandwidth bound, not in a compute-bound convolutional stack.
Quantizing the MLX build would not have rescued it. The MLX delegate can quantize matrix multiplies but has no quantized convolution, so an int4 or int8 export would shrink only the decoder's linear layers and leave the backbone, which dominates at this resolution, running fp32. An fp16 MLX build would halve the weights and reach the convolutions, but it would still have to close a 2.70x gap against a purpose-built accelerator.
The same comparison has come out the same way for every convolutional model in this project that was tried on both backends.