Inverse-LLaVA-7B
Inverse-LLaVA maps language features into the visual feature space and fuses them inside the language decoder. This is the full-data 7B checkpoint trained jointly on visual instructions, without a separate alignment-pretraining stage.
Paper · Source and evaluation instructions
Load and answer
On Linux with a CUDA-compatible PyTorch installation, install the source package:
git clone https://github.com/xuhuizhan5/Inverse-LLaVA.git
cd Inverse-LLaVA
python -m pip install -e .
from invllava.release import load_pretrained
model = load_pretrained(
"xuhuizhan5/Inverse-LLaVA-7B",
cache_dir="/workspace/cache",
device="cuda",
dtype="bfloat16",
lora_execution="unmerged",
max_new_tokens=128,
)
print(model.answer("/path/to/image.jpg", "What is shown in this image?"))
Use a commit hash in revision= to pin a particular Hub snapshot. The weights
file has the stable identity below; the loader verifies its checksums before
constructing the model. It downloads the pinned Vicuna and CLIP weights
separately when they are not cached. The 695 MB delta is not the total download
or GPU-memory requirement. Allow space for the approximately 14 GB language
backbone and the vision encoder, plus normal framework overhead. The measured
batch-1 inference allocation was about 14.4 GiB; device and workload headroom
is required. This is a custom package loader, not a Transformers AutoModel
checkpoint. It does not execute code downloaded from the model repository.
For offline use, cache the bundle and both backbones first, then set
local_files_only=True. Benchmark reproduction requires the exact benchmark
prompts and scorers in the source repository rather than the generic question
above. See checkpoint instructions.
Architecture and training
| Setting | Value |
|---|---|
| Language backbone | Vicuna-7B-v1.5, 3321f76e3f527bd14065daf69dad9344000a201d |
| Frozen vision encoder | CLIP ViT-L/14-336, ce19dc912ca5cd21c8a653c79e251e808ccabcd1 |
| Visual features | Final layer, patch tokens, width 1024 |
| Fusion | Decoder layer 0, separate Q/K/V branches, concatenation |
| Adaptation | LoRA rank 128, alpha 256, dropout 0.05; fusion and LoRA trained jointly |
| Training | LLaVA-v1.5-mix665k, 665,298 distinct rows, one epoch |
| Global batch / updates | 128 / 5,198 |
| Precision | BF16 |
| Total / trained parameters | 7,389,496,323 / 347,573,251 |
| Delta size | 695,206,838 bytes |
The recorded 665,344 sampled rows include distributed-sampler padding; they are not a different dataset size. The configuration, exact tensor names and training-input identities are retained in the bundle. Backbone weights and optimizer states are not included.
model_delta.safetensors SHA-256:
9ed8914aedbbeb55d01607da0aab96af2d55f638e0b827f20ac232d0a4f6741f
Evaluation
These complete evaluations use this checkpoint and the official LLaVA-1.5 LoRA and full-fine-tuning (FFT) checkpoints under shared evaluation protocols. Their training data and adaptation procedures differ.
| Benchmark | Inverse-LLaVA | LLaVA-LoRA | LLaVA-FFT |
|---|---|---|---|
| VQAv2 test-dev | 78.45 | 79.13 | 78.55 |
| GQA | 62.28 | 62.63 | 61.89 |
| VizWiz | 50.96 | 48.56 | 50.64 |
| ScienceQA-IMG | 69.61 | 68.82 | 69.11 |
| TextVQA | 56.96 | 58.47 | 58.21 |
| MMBench EN | 62.63 | 67.10 | 65.12 |
| MMBench CN | 54.04 | 58.93 | 58.33 |
| MME perception | 1453.82 | 1484.58 | 1507.28 |
| MME cognition | 279.29 | 258.21 | 344.64 |
| MM-Vet | 28.67 | 31.24 | 29.68 |
Values are percentages except MME's official score. MME perception and cognition are parts of one benchmark. VQAv2 uses official server scoring; MM-Vet uses the documented hosted-judge protocol. See the source repository's evaluation guide for splits, prompt templates and scoring conditions. The performance gains are task-dependent. OCR and several perception scores remain below the reference models. This single trained model does not establish seed-averaged performance or a universal efficiency advantage.
Intended use and limitations
This model supports research in multimodal learning and reproducible model comparison. It can produce incorrect descriptions, hallucinated answers, harmful content and biases inherited from its backbones and training data. It is not validated for safety-critical decisions. Inspect outputs and follow the applicable model and dataset terms.
Terms and attribution
The model is distributed under the Llama 2 Community License, with its acceptable-use policy and notice. It builds on Vicuna/Llama 2 and CLIP; backbone weights are retrieved from their original repositories and retain their terms. The source implementation is separately licensed Apache-2.0. Dataset annotations and images retain their source terms.
Authors: Xuhui Zhan and Tyler Derr. Please cite Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping.
Model tree for xuhuizhan5/Inverse-LLaVA-7B
Base model
lmsys/vicuna-7b-v1.5