Skip to content

OmAI Lab

Open Multimodal AGI Research

Website Hugging Face X (Twitter)

Building the foundational brains for the physical world.


🌌 About Us

At OmAI Lab, we believe the future of AI extends far beyond pure text. We are dedicated to building the "brains" for next-generation systems by focusing on the intersection of Visual Reasoning and Embodied Agents.

Our research spans across open-vocabulary perception, reinforced vision-language models, and real-time inference. We aim to bridge the critical gap between high-level logical reasoning and fine-grained visual action—building models that don't just "see" the world, but intuitively understand and interact with it.


🚢 Flagship VLX Model Series

  • 📹 VLX-Flow: A real-time VLM for streaming video understanding .
  • 🔍 VLX-Seek: Fine-grained visual perception and grounding for physical AI.
  • 🚗 VLX-Go: Efficient general-purpose embodied navigation in the wild.

🚀 Core Research Tracks

🧠 Reinforced & Advanced Visual Reasoning

Models that think, reason, and understand the visual world at a granular level.

  • 🌟 VLM-R1: Solving Visual Understanding with Reinforced VLMs. (Highly active)
  • 🔎 ZoomEye: Enhancing Multimodal LLMs with human-like zooming capabilities through tree-based image exploration.
  • 🌍 ImageRAG: Enhancing ultrahigh-resolution remote sensing imagery analysis.

👁️ Real-Time Perception & Open-World Visual Detection

Foundational spatial understanding optimized for edge and on-premise speeds.

  • OmDet: Real-time, highly accurate, open-vocabulary end-to-end object detection.
  • 🔍 VLM-FO1: Bridging the gap between high-level reasoning and fine-grained perception in Vision-Language Models.
  • 📐 GroundVLP: Harnessing Zero-shot Visual Grounding from Vision-Language Pre-training.

🤖 Multimodal Agents & Embodied AI

Action-oriented intelligence for physical and virtual environments.

  • 🛠️ OmAgent: A comprehensive framework to build multimodal language agents for fast prototyping and production.
  • 🎯 OmTrackVLA: Open and reproducible research for tracking Vision-Language-Action (VLA) models.

📊 Benchmarks & Evaluation

Rigorous standards for the open-source multimodal community.

  • 🌍 RS5M: A pioneer work in VLM benchmark for remote sensing.
  • 📏 OVDEval: A comprehensive evaluation benchmark for Open-Vocabulary Detection.
  • 📝 VL-CheckList: Evaluating vision & language pretraining models with objects, attributes, and relations.

Pinned Loading

  1. VLM-R1 VLM-R1 Public

    Solve Visual Understanding with Reinforced VLMs

    Python 6k 384

  2. OmDet OmDet Public

    Real-time and accurate open-vocabulary end-to-end object detection

    Python 1.4k 118

  3. OmAgent OmAgent Public

    [EMNLP-2024] Build multimodal language agents for fast prototype and production

    Python 2.7k 292

  4. VLX-Seek VLX-Seek Public

    VLX-Seek is a device-native vision-language model that enables machines to see, understand, and reason about the visual world with high precision.

    Python 580 75

  5. OmTrackVLA OmTrackVLA Public

    Open & Reproducible Research for Tracking VLAs

    Python 277 14

  6. ZoomEye ZoomEye Public

    [EMNLP-2025 Oral] ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

    Python 91 10

Repositories

Showing 10 of 27 repositories
  • vlx-dev-kit Public

    SDK for VLX developer modules

    om-ai-lab/vlx-dev-kit's past year of commit activity
    2 0 0 0 Updated Sep 9, 2026
  • om-ai-lab.github.io Public

    Official website for the org

    om-ai-lab/om-ai-lab.github.io's past year of commit activity
    HTML 1 1 0 0 Updated Sep 7, 2026
  • VLX-Seek Public

    VLX-Seek is a device-native vision-language model that enables machines to see, understand, and reason about the visual world with high precision.

    om-ai-lab/VLX-Seek's past year of commit activity
    Python 580 Apache-2.0 75 1 0 Updated Sep 2, 2026
  • VLX-Flow Public

    VLX-Flow: streaming VLM for real-time general vision intelligence

    om-ai-lab/VLX-Flow's past year of commit activity
    115 Apache-2.0 5 2 0 Updated Sep 2, 2026
  • VLX-Go Public

    VLX-Go is a lightweight vision-language waypoint planner that enables robots to perceive their surroundings, track targets, and navigate the physical world in real time.

    om-ai-lab/VLX-Go's past year of commit activity
    140 Apache-2.0 2 2 0 Updated Sep 2, 2026
  • .github Public
    om-ai-lab/.github's past year of commit activity
    0 0 0 0 Updated Sep 2, 2026
  • OmTrackVLA Public

    Open & Reproducible Research for Tracking VLAs

    om-ai-lab/OmTrackVLA's past year of commit activity
    Python 277 14 14 0 Updated Jul 21, 2026
  • VLM-R1 Public

    Solve Visual Understanding with Reinforced VLMs

    om-ai-lab/VLM-R1's past year of commit activity
    Python 6,018 Apache-2.0 384 164 2 Updated Jul 7, 2026
  • VLM-FO1 Public

    VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs

    om-ai-lab/VLM-FO1's past year of commit activity
    Python 332 15 7 0 Updated Jun 18, 2026
  • Probing-VLM-VGM Public

    Probing VLM vs VGM for spatial understanding.

    om-ai-lab/Probing-VLM-VGM's past year of commit activity
    Python 15 0 1 0 Updated Jun 1, 2026

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…