Inverse-LLaVA-7B

Inverse-LLaVA maps language features into the visual feature space and fuses them inside the language decoder. This is the full-data 7B checkpoint trained jointly on visual instructions, without a separate alignment-pretraining stage.

Paper · Source and evaluation instructions

Load and answer

On Linux with a CUDA-compatible PyTorch installation, install the source package:

git clone https://github.com/xuhuizhan5/Inverse-LLaVA.git
cd Inverse-LLaVA
python -m pip install -e .
from invllava.release import load_pretrained

model = load_pretrained(
    "xuhuizhan5/Inverse-LLaVA-7B",
    cache_dir="/workspace/cache",
    device="cuda",
    dtype="bfloat16",
    lora_execution="unmerged",
    max_new_tokens=128,
)
print(model.answer("/path/to/image.jpg", "What is shown in this image?"))

Use a commit hash in revision= to pin a particular Hub snapshot. The weights file has the stable identity below; the loader verifies its checksums before constructing the model. It downloads the pinned Vicuna and CLIP weights separately when they are not cached. The 695 MB delta is not the total download or GPU-memory requirement. Allow space for the approximately 14 GB language backbone and the vision encoder, plus normal framework overhead. The measured batch-1 inference allocation was about 14.4 GiB; device and workload headroom is required. This is a custom package loader, not a Transformers AutoModel checkpoint. It does not execute code downloaded from the model repository.

For offline use, cache the bundle and both backbones first, then set local_files_only=True. Benchmark reproduction requires the exact benchmark prompts and scorers in the source repository rather than the generic question above. See checkpoint instructions.

Architecture and training

Setting Value
Language backbone Vicuna-7B-v1.5, 3321f76e3f527bd14065daf69dad9344000a201d
Frozen vision encoder CLIP ViT-L/14-336, ce19dc912ca5cd21c8a653c79e251e808ccabcd1
Visual features Final layer, patch tokens, width 1024
Fusion Decoder layer 0, separate Q/K/V branches, concatenation
Adaptation LoRA rank 128, alpha 256, dropout 0.05; fusion and LoRA trained jointly
Training LLaVA-v1.5-mix665k, 665,298 distinct rows, one epoch
Global batch / updates 128 / 5,198
Precision BF16
Total / trained parameters 7,389,496,323 / 347,573,251
Delta size 695,206,838 bytes

The recorded 665,344 sampled rows include distributed-sampler padding; they are not a different dataset size. The configuration, exact tensor names and training-input identities are retained in the bundle. Backbone weights and optimizer states are not included.

model_delta.safetensors SHA-256:

9ed8914aedbbeb55d01607da0aab96af2d55f638e0b827f20ac232d0a4f6741f

Evaluation

These complete evaluations use this checkpoint and the official LLaVA-1.5 LoRA and full-fine-tuning (FFT) checkpoints under shared evaluation protocols. Their training data and adaptation procedures differ.

Benchmark Inverse-LLaVA LLaVA-LoRA LLaVA-FFT
VQAv2 test-dev 78.45 79.13 78.55
GQA 62.28 62.63 61.89
VizWiz 50.96 48.56 50.64
ScienceQA-IMG 69.61 68.82 69.11
TextVQA 56.96 58.47 58.21
MMBench EN 62.63 67.10 65.12
MMBench CN 54.04 58.93 58.33
MME perception 1453.82 1484.58 1507.28
MME cognition 279.29 258.21 344.64
MM-Vet 28.67 31.24 29.68

Values are percentages except MME's official score. MME perception and cognition are parts of one benchmark. VQAv2 uses official server scoring; MM-Vet uses the documented hosted-judge protocol. See the source repository's evaluation guide for splits, prompt templates and scoring conditions. The performance gains are task-dependent. OCR and several perception scores remain below the reference models. This single trained model does not establish seed-averaged performance or a universal efficiency advantage.

Intended use and limitations

This model supports research in multimodal learning and reproducible model comparison. It can produce incorrect descriptions, hallucinated answers, harmful content and biases inherited from its backbones and training data. It is not validated for safety-critical decisions. Inspect outputs and follow the applicable model and dataset terms.

Terms and attribution

The model is distributed under the Llama 2 Community License, with its acceptable-use policy and notice. It builds on Vicuna/Llama 2 and CLIP; backbone weights are retrieved from their original repositories and retain their terms. The source implementation is separately licensed Apache-2.0. Dataset annotations and images retain their source terms.

Authors: Xuhui Zhan and Tyler Derr. Please cite Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for xuhuizhan5/Inverse-LLaVA-7B

Adapter
(185)
this model

Paper for xuhuizhan5/Inverse-LLaVA-7B