Open weights and reproducible evaluations · Updated October 2026
Text-to-vision fusion. One instruction-tuning stage.
Inverse-LLaVA maps language features into visual feature space and
combines them with frozen CLIP features inside decoder attention. The
released 7B model learns multimodal interaction without a separate
alignment-pretraining stage.
1 stageJoint fusion and LoRA training
665K examplesOne instruction-tuning epoch
9 benchmarksOfficial LoRA and FFT references
Two directions of multimodal mapping. The image,
question, and answers come from the saved VizWiz evaluation:
VizWiz_test_00006540.jpg. This selected example
illustrates the interface, not overall accuracy. Reference: White,
with full consensus credit (1). Image and question:
VizWiz-VQA,
CC BY 4.0. Example provenance.
Architecture
LLaVA projects visual features into the language model's continuous
hidden space. Inverse-LLaVA instead introduces text-to-vision branches
in the decoder's query, key, and value projections. Each branch maps
language states to the visual feature width, concatenates them with
normalized CLIP features at the corresponding sequence positions, and
adds a learned update to the attention projection.
Released 7B configuration. Vicuna-7B-v1.5 and CLIP
ViT-L/14 at 336 pixels; final-layer CLIP patch features (1,024
channels); fusion at decoder layer 0. The backbones remain frozen
while fusion parameters and LoRA adapters outside the fusion layer
train jointly. All decoder layers are retained.
Architecture and tensor shapes.
Results from the released 7B checkpoint
Complete evaluations compare Inverse-LLaVA with the official
LLaVA-1.5-7B LoRA and full-fine-tuning (FFT) checkpoints under shared
task protocols. Their training data, adaptation procedures, and CLIP
feature layers differ. Separate single-stage controls address those
training differences.
Nine primary benchmarks, with MME perception and cognition shown
separately. Higher is better. Values are percentages except MME's
official scores.
Benchmark
Inverse-LLaVA
LLaVA-LoRA
LLaVA-FFT
VQAv2 test-dev
78.45
79.13
78.55
GQA balanced test-dev
62.28
62.63
61.89
VizWiz test
50.96
48.56
50.64
ScienceQA-IMG
69.61
68.82
69.11
TextVQA validation
56.96
58.47
58.21
MMBench EN dev
62.63
67.10
65.12
MMBench CN dev
54.04
58.93
58.33
MME perception
1453.82
1484.58
1507.28
MME cognition
279.29
258.21
344.64
MM-Vet v1
28.67
31.24
29.68
Bold and underline mark
the highest and second-highest point estimates, not statistical
significance. Shading identifies Inverse-LLaVA. VQAv2 uses official
server scoring; MM-Vet uses the hosted GPT-4.1 judge.
Splits, prompts, and scorers
·
Saved answers and scores.
Without the additional alignment stage, Inverse-LLaVA reaches 78.45%
on VQAv2 and 62.28% on GQA. VizWiz exceeds LoRA by 2.40 percentage
points (95% paired interval: 1.58 to 3.21). The smaller VizWiz
difference from FFT and the ScienceQA differences have intervals
spanning zero. TextVQA and MMBench remain lower; MME cognition is
above LoRA's point score and below FFT's. Intervals condition on these
trained checkpoints and do not measure training-seed variation.
Published reference scores and broader model context
These LLaVA-1.5 values are from the
official model zoo. They remain separate from the evaluations above because source
protocols and judges can differ.
Originally published LLaVA-1.5-7B results.
Benchmark
LoRA
FFT
VQAv2
79.1
78.5
GQA
63.0
62.0
VizWiz
47.8
50.0
ScienceQA-IMG
68.4
66.8
TextVQA
58.2
58.2
MMBench EN
66.1
64.3
MMBench CN
58.9
58.3
MME perception
1476.9
1510.7
MM-Vet
30.2
31.1
Related systems include InstructBLIP, InternVL-Chat, EVE, and the
Qwen-VL family. They use different connectors, visual encoders or resolutions,
and supervision. Their published scores are not training-matched
measurements of the inverse-mapping contribution.
A smaller multimodal training recipe
Task-training examples; pretrained backbone data are excluded.
Recipe
Alignment stage
Instruction stage
Combined
LLaVA-1.5
558K
665K
1,223K
Inverse-LLaVA
0
665K
665K
This removes 45.6% of the task-training examples in LLaVA-1.5's
two-stage recipe. The released model trains for one epoch on 665,298
mixture rows, including text-only examples. Alignment and instruction
examples have different costs, and fusion adds computation: the sample
reduction is not a measured 45.6% runtime or FLOP saving.
Training recipe
·
Profiling procedure.
What the additional studies examine
Fusion and supervision
Component, fusion-depth, feature-layer, and single-stage projector
controls isolate design choices. Continued-training studies test
paired data and instruction replay. Caption-only continuation does
not establish TextVQA recovery; effects depend on the task and
supervision.
Representations
Features from the same 100 distinct image/prompt pairs support
PCA, t-SNE, CKA, and pairwise geometry analysis. Matched image
interventions test sensitivity. Independently fitted projection
panels do not share coordinates.
Scaling studies
Inverse-LLaVA-HD concatenates final and penultimate CLIP features
into 2,048 channels. The 13B variant uses a larger language
backbone. Both released variants use a matched 5% instruction
subset with intermediate checkpoints; neither is a full-data model
or a fitted scaling law.
Install the
source package
in a CUDA-compatible environment, then load the verified 7B snapshot:
from invllava.release import load_pretrained
model = load_pretrained(
"xuhuizhan5/Inverse-LLaVA-7B",
revision="d2ec3d3b27b61ccc08ea1677ad63132f1400307c",
cache_dir="/workspace/cache",
)
print(model.answer("image.jpg", "What is shown in this image?"))
The 695 MB download contains learned fusion and LoRA weights. The
loader verifies checksums and retrieves pinned Vicuna and CLIP
backbones separately, so allow storage and GPU memory for the complete
model. Fresh-download inference was verified; all 33 local comparison
cells reproduce from exported answers. VQAv2 and MM-Vet retain
external scoring.
Offline use and checkpoint formats.
Code: Apache-2.0. Model weights: Llama 2 Community License and
applicable upstream terms. Dataset assets retain their source
licenses. Generated answers may contain errors or biases.
Citation
@article{zhan2025inverse,
title={Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping},
author={Zhan, Xuhui and Derr, Tyler},
journal={arXiv preprint arXiv:2508.12466},
year={2025}
}
Results on this page correspond to the released checkpoint and
evaluation records. The paper is available as an arXiv preprint.