Skip to content

Wire CLIP vision-layer hook aliases for LLaVA adapters - #1823

Merged
jlarson4 merged 2 commits into
TransformerLensOrg:devfrom
koriyoshi2041:fix/llava-vision-hook-aliases
Sep 30, 2026
Merged

jlarson4 merged 2 commits into
TransformerLensOrg:devfrom
koriyoshi2041:fix/llava-vision-hook-aliases

Conversation

@koriyoshi2041

@koriyoshi2041 koriyoshi2041 commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Problem

The LLaVA-family adapters declare attention and MLP hook aliases for CLIP vision encoder layers, but CLIPVisionEncoderLayerBridge does not register the corresponding submodules. The four aliases therefore cannot resolve.

Fix

Register the CLIP layer norms, q/k/v/output projections, and fc1/fc2 MLP projections with the same generalized bridge structure used by the SigLIP tower. The attention bridge receives the vision tower dimensions rather than the language-model dimensions. The now-obsolete strict xfails are removed for LLaVA, LLaVA-Next, and LLaVA-OneVision.

Test

  • uv run pytest tests/unit/model_bridge/test_hook_alias_resolution.py tests/unit/model_bridge/supported_architectures/test_llava_adapter.py tests/unit/model_bridge/supported_architectures/test_llava_next_adapter.py tests/unit/model_bridge/supported_architectures/test_llava_onevision_adapter.py -q (192 passed, 1 skipped, 3 unrelated xfailed)
  • make format
  • uv run mypy transformer_lens/model_bridge/generalized_components/clip_vision_encoder.py
  • git diff --check

Risk

The change is limited to CLIP vision-layer component registration and alias resolution. It does not change the wrapped Hugging Face forward path or claim end-to-end multimodal parity.

@koriyoshi2041
koriyoshi2041 force-pushed the fix/llava-vision-hook-aliases branch from b8f7332 to 52b6a9c Compare September 26, 2026 02:42

@jlarson4 jlarson4 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi @koriyoshi2041! Thanks for closing the CLIP half of this gap. A couple small items to address below before we can merge.

Comment thread transformer_lens/model_bridge/generalized_components/clip_vision_encoder.py Outdated
Comment thread transformer_lens/model_bridge/generalized_components/clip_vision_encoder.py Outdated
@jlarson4

Copy link
Copy Markdown
Collaborator

Looks good! Approved

@jlarson4
jlarson4 merged commit 2d0ad51 into TransformerLensOrg:dev Sep 30, 2026
27 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants