Skip to content

Did you ever try with distilling the VGGT knowledge to the visual tokens just after ViT, instead of the LLM final visual hidden states? #24

Description

@TyroneLi

Hi, Did you ever try with distilling the VGGT knowledge to the visual tokens just after ViT saying visual encoder siglip, instead of the LLM final visual hidden states saying inside LLM layers?

thx

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions