Inverse-LLaVA

Rethinking Multimodal Alignment
via Text-to-Vision Mapping

Xuhui Zhan · Tyler Derr

Vanderbilt University

Open weights and reproducible evaluations · Updated October 2026

Text-to-vision fusion. One instruction-tuning stage.

Inverse-LLaVA maps language features into visual feature space and combines them with frozen CLIP features inside decoder attention. The released 7B model learns multimodal interaction without a separate alignment-pretraining stage.

1 stageJoint fusion and LoRA training

665K examplesOne instruction-tuning epoch

9 benchmarksOfficial LoRA and FFT references

A shared VizWiz photograph and question, What color is it?, follow the two feature paths. Inverse-LLaVA answers White; both official LLaVA references answer Blue.
Two directions of multimodal mapping. The image, question, and answers come from the saved VizWiz evaluation: VizWiz_test_00006540.jpg. This selected example illustrates the interface, not overall accuracy. Reference: White, with full consensus credit (1). Image and question: VizWiz-VQA, CC BY 4.0. Example provenance.

Architecture

LLaVA projects visual features into the language model's continuous hidden space. Inverse-LLaVA instead introduces text-to-vision branches in the decoder's query, key, and value projections. Each branch maps language states to the visual feature width, concatenates them with normalized CLIP features at the corresponding sequence positions, and adds a learned update to the attention projection.

LLaVA-1.5 alignment and instruction stages alongside Inverse-LLaVA's single-stage Q, K, and V fusion branches. Frozen backbones and trainable adapters are identified.
Released 7B configuration. Vicuna-7B-v1.5 and CLIP ViT-L/14 at 336 pixels; final-layer CLIP patch features (1,024 channels); fusion at decoder layer 0. The backbones remain frozen while fusion parameters and LoRA adapters outside the fusion layer train jointly. All decoder layers are retained. Architecture and tensor shapes.

Results from the released 7B checkpoint

Complete evaluations compare Inverse-LLaVA with the official LLaVA-1.5-7B LoRA and full-fine-tuning (FFT) checkpoints under shared task protocols. Their training data, adaptation procedures, and CLIP feature layers differ. Separate single-stage controls address those training differences.

Nine primary benchmarks, with MME perception and cognition shown separately. Higher is better. Values are percentages except MME's official scores.
Benchmark Inverse-LLaVA LLaVA-LoRA LLaVA-FFT
VQAv2 test-dev 78.45 79.13 78.55
GQA balanced test-dev 62.28 62.63 61.89
VizWiz test 50.96 48.56 50.64
ScienceQA-IMG 69.61 68.82 69.11
TextVQA validation 56.96 58.47 58.21
MMBench EN dev 62.63 67.10 65.12
MMBench CN dev 54.04 58.93 58.33
MME perception 1453.82 1484.58 1507.28
MME cognition 279.29 258.21 344.64
MM-Vet v1 28.67 31.24 29.68

Bold and underline mark the highest and second-highest point estimates, not statistical significance. Shading identifies Inverse-LLaVA. VQAv2 uses official server scoring; MM-Vet uses the hosted GPT-4.1 judge. Splits, prompts, and scorers · Saved answers and scores.

Without the additional alignment stage, Inverse-LLaVA reaches 78.45% on VQAv2 and 62.28% on GQA. VizWiz exceeds LoRA by 2.40 percentage points (95% paired interval: 1.58 to 3.21). The smaller VizWiz difference from FFT and the ScienceQA differences have intervals spanning zero. TextVQA and MMBench remain lower; MME cognition is above LoRA's point score and below FFT's. Intervals condition on these trained checkpoints and do not measure training-seed variation.

Published reference scores and broader model context

These LLaVA-1.5 values are from the official model zoo. They remain separate from the evaluations above because source protocols and judges can differ.

Originally published LLaVA-1.5-7B results.
Benchmark LoRA FFT
VQAv2 79.1 78.5
GQA 63.0 62.0
VizWiz 47.8 50.0
ScienceQA-IMG 68.4 66.8
TextVQA 58.2 58.2
MMBench EN 66.1 64.3
MMBench CN 58.9 58.3
MME perception 1476.9 1510.7
MM-Vet 30.2 31.1

Related systems include InstructBLIP, InternVL-Chat, EVE, and the Qwen-VL family. They use different connectors, visual encoders or resolutions, and supervision. Their published scores are not training-matched measurements of the inverse-mapping contribution.

A smaller multimodal training recipe

Task-training examples; pretrained backbone data are excluded.
Recipe Alignment stage Instruction stage Combined
LLaVA-1.5 558K 665K 1,223K
Inverse-LLaVA 0 665K 665K

This removes 45.6% of the task-training examples in LLaVA-1.5's two-stage recipe. The released model trains for one epoch on 665,298 mixture rows, including text-only examples. Alignment and instruction examples have different costs, and fusion adds computation: the sample reduction is not a measured 45.6% runtime or FLOP saving. Training recipe · Profiling procedure.

What the additional studies examine

Fusion and supervision

Component, fusion-depth, feature-layer, and single-stage projector controls isolate design choices. Continued-training studies test paired data and instruction replay. Caption-only continuation does not establish TextVQA recovery; effects depend on the task and supervision.

Representations

Features from the same 100 distinct image/prompt pairs support PCA, t-SNE, CKA, and pairwise geometry analysis. Matched image interventions test sensitivity. Independently fitted projection panels do not share coordinates.

Scaling studies

Inverse-LLaVA-HD concatenates final and penultimate CLIP features into 2,048 channels. The 13B variant uses a larger language backbone. Both released variants use a matched 5% instruction subset with intermediate checkpoints; neither is a full-data model or a fitted scaling law.

Checkpoints, training histories, and feature tensors · Analysis instructions.

Try the released model

Install the source package in a CUDA-compatible environment, then load the verified 7B snapshot:

from invllava.release import load_pretrained

model = load_pretrained(
    "xuhuizhan5/Inverse-LLaVA-7B",
    revision="d2ec3d3b27b61ccc08ea1677ad63132f1400307c",
    cache_dir="/workspace/cache",
)
print(model.answer("image.jpg", "What is shown in this image?"))

The 695 MB download contains learned fusion and LoRA weights. The loader verifies checksums and retrieves pinned Vicuna and CLIP backbones separately, so allow storage and GPU memory for the complete model. Fresh-download inference was verified; all 33 local comparison cells reproduce from exported answers. VQAv2 and MM-Vet retain external scoring. Offline use and checkpoint formats.

Code: Apache-2.0. Model weights: Llama 2 Community License and applicable upstream terms. Dataset assets retain their source licenses. Generated answers may contain errors or biases.

Citation

@article{zhan2025inverse,
  title={Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping},
  author={Zhan, Xuhui and Derr, Tyler},
  journal={arXiv preprint arXiv:2508.12466},
  year={2025}
}

Results on this page correspond to the released checkpoint and evaluation records. The paper is available as an arXiv preprint.