CLEF Calibrated INT8 ConvRot
CLEF turns text, JSON, images, and video into probabilities over the allowed answers to typed questions, in a single forward pass.
Benefits
- Lower VRAM: 43.3% lower measured process peak than BF16.
- Self-contained weights and processor: loads directly into INT8/BF16 storage, without downloading or loading the original BF16 checkpoint first.
- Calibrated INT8 text projections reduce drift from BF16 compared with naive INT8. Embeddings, vision tower, small projections, and decision head stay BF16.
- H100 CuTeDSL kernels fuse calibrated scales, biases, and MLP SwiGLU.
| GPU Memory | BF16 | BF16 + CuTe | INT8, Torch | INT8 + CuTe |
|---|---|---|---|---|
| Loaded-model residency, PyTorch allocated | 51.19 GiB | 51.19 GiB | 28.76 GiB | 28.65 GiB |
| Inference peak, PyTorch allocated | 51.66 GiB | 51.59 GiB | 29.42 GiB | 29.05 GiB |
| Load + batch-1 process peak, NVML | 52.90 GiB | 52.72 GiB | 31.04 GiB | 29.99 GiB |
Measured in fresh processes on H100 80 GB, batch size 1, using the same 23 text records (298-2,270 tokens) and one 224x224 image loading/parity probe. NVML process memory is sampled every 5 ms. The BF16 baseline uses the unmodified source runtime, without CuTe patches.
| Input Tokens | BF16 | BF16 + CuTe | INT8, Torch | INT8 + CuTe |
|---|---|---|---|---|
| 298 | 77.38 ms | 67.76 ms | 116.24 ms | 83.57 ms |
| 1,502 | 164.46 ms | 139.84 ms | 278.91 ms | 119.80 ms |
| 2,270 | 241.09 ms | 208.46 ms | 401.41 ms | 170.62 ms |
Warm median latency over 30 repeats, including the decision head. Both INT8 paths use the same calibrated weights. The Torch path uses ATen integer GEMMs and native pointwise operations, without CuTe kernels. Both use Comfy Kitchen's CUDA ConvRot activation quantizer. Measurements. Usage.
Batchable throughput on H100: 1,502 tokens per record, eight questions per record, four options per question. Each answered question counts as one decision. Warm medians over 10 repeats include the backbone and decision head. Both INT8 curves use the calibrated weights. Batch measurements. These are equal-length synthetic batches; tokenization, request queuing, and batch assembly are excluded. Batch latency, peak VRAM, and decision agreement with batch size 1 are included in the measurements.
Limitations
- Inference-only custom loader; not a drop-in
AutoModel.from_pretrainedcheckpoint. - CuTeDSL backend requires H100/SM90 and the dependencies listed in Usage.
- BF16 can be faster, particularly on short inputs; memory reduction is the primary benefit, not a universal latency improvement.
- Calibration and evaluation use small synthetic text sets. Agreement below is with BF16, not ground-truth accuracy; image/video quality is not evaluated.
- One BF16 decision disagreement remains. Validate representative inputs before deployment.
Evaluation
16 held-out records and seven regression records, with 80 decisions in total. BF16 is the reference. Calibration uses separate training and selection records.
| Metric | BF16 | Naive INT8 | Calibrated INT8 |
|---|---|---|---|
| Held-out decision agreement | 48/48 | 48/48 | 48/48 |
| Regression decision agreement | 32/32 | 30/32 | 31/32 |
| Total decision agreement | 80/80 | 78/80 | 79/80 |
| Mean held-out relative logit L2 | 0 | 0.02166 | 0.01837 |
| Mean held-out relative hidden-state L2 | 0 | 0.10379 | 0.09733 |
Calibration reduces mean held-out logit drift by 15.2% and hidden-state drift by 6.2%. The direct loader preserves calibrated hidden states and logits exactly on all 23 text records and the image parity probe. Evaluation details.
- Downloads last month
- 25
