Skip to content

Latest commit

 

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Live-Text OWL-ViT / OWLv2 on NVIDIA DeepStream 9.0

License: MIT DeepStream OWL-ViT

Open-vocabulary detection with OWL-ViT / OWLv2 on NVIDIA DeepStream 9.0. You say what to look for in plain words, and you can change those words while the video is running - the new classes show up on the next frame, no restart.

Three models, picked with --model:

--model model input
owlvit (default) google/owlvit-base-patch32 (ViT-B/32) 768x768
nanoowl google/owlvit-base-patch16 (ViT-B/16) 768x768
owlv2 google/owlv2-base-patch16-ensemble 960x960

owlvit is the fastest, owlv2 the most accurate, nanoowl in between. nanoowl is the ViT-B/16 model from NVIDIA's NanoOWL project, served here through DeepStream.

owlv2 detecting an owl and its eye

owlv2 on the demo clip with the prompt owl . eye . - open-vocabulary detection of the whole owl and a part (its eye).

How it works

OWL-ViT looks at the image and at the words in your prompt separately, then matches regions to phrases. Because the two are independent, changing the prompt only re-runs the small text part - the image model keeps running untouched, so new classes appear on the very next frame.

Requirements

DeepStream 9.0 requirements (source)

  • Ubuntu 24.04
  • NVIDIA driver >= 590.48.01
  • CUDA 13.1
  • TensorRT 10.14.1.48
  • GStreamer 1.24.2

This project uses the Docker image (nvcr.io/nvidia/deepstream:9.0-samples-multiarch), which bundles all of the above - only the host driver needs to meet the minimum version.

Host requirements

  • NVIDIA GPU
  • NVIDIA driver >= 590.48.01
  • Docker + NVIDIA Container Toolkit (--gpus all must work):
    docker run --rm --gpus all nvcr.io/nvidia/deepstream:9.0-samples-multiarch nvidia-smi
  • Internet on first setup: 02_make_onnx.sh pulls the OWL-ViT / OWLv2 weights from Hugging Face (anonymous, no account needed); get_test_videos.sh fetches the demo clips.

Quick start

Steps 0-4 are one-time setup.

chmod +x scripts/*.sh
M=owlv2                        # or owlvit / nanoowl

./scripts/00_get_model.sh                                  # CPU export image
./scripts/01_build_libs.sh                                 # build/*.so
./scripts/02_make_onnx.sh    --model $M                    # export the model to ONNX
./scripts/03_build_engine.sh --model $M --precision fp16   # build the TensorRT engines
./scripts/04_build_app.sh                                  # build/owl-app
./scripts/get_test_videos.sh                               # data/{owl,dog_park,tomatoes}.mp4

# detect the owl and its eyes, save an annotated MP4
./scripts/run.sh --model owlv2 --video file:///workspace/data/owl.mp4 "owl . eye ."

The same flags work for any model:

./scripts/run.sh --model owlvit --video file:///workspace/data/dog_park.mp4 "dog . person ."
./scripts/run.sh --model owlv2  --video file:///workspace/data/owl.mp4 "owl . eye ."

Changing what it detects, live

While a run is going, write new phrases to the control file from another shell:

echo "owl . beak ." > /tmp/owl_prompt

Each phrase (separated by . or ,) is one thing to detect. --switch "<prompt>" does this automatically once the video starts:

./scripts/run.sh --model owlv2 --video file:///workspace/data/owl.mp4 "owl . eye ." --switch "owl . beak ."

Multistream

Repeat --video to run several streams at once; they're shown in a tiled grid and share one prompt. Build the engine with --max-batch at least the number of streams:

./scripts/03_build_engine.sh --model owlvit --precision fp16 --max-batch 4

./scripts/run.sh --model owlvit \
  --video file:///workspace/data/owl.mp4 \
  --video file:///workspace/data/dog_park.mp4 \
  --video file:///workspace/data/tomatoes.mp4 \
  --video file:///workspace/data/owl.mp4 \
  "owl . dog . person . tomato ."

Tuning detection

Flags to run.sh that change the boxes:

Flag Default Effect
--thr 0.10 confidence threshold - raise it for fewer, surer boxes
--nms-iou 0.50 overlap to suppress; <=0 disables
--max-area 0.92 drop boxes larger than this fraction of the frame; >=1 disables

Other flags: --out FILE, --sink ELEM, --live (window), --out-w / --out-h, --bitrate, --force-sw, --log.

License

MIT - see LICENSE. Third-party components have their own licenses, listed in NOTICE:

Component License Notes
OWL-ViT / OWLv2 (Google) Apache-2.0 weights pulled from Hugging Face at export; not redistributed
CLIP tokenizer (assets/clip_*) MIT from openai/CLIP via Hugging Face
NVIDIA DeepStream SDK 9.0 + TensorRT NVIDIA Proprietary runs inside the Docker container

References

About

No description or website provided.

Topics

Resources

Stars

10 stars

Watchers

0 watching

Forks

Releases

Sponsor this project

Packages

Contributors

Languages