Open-vocabulary detection with OWL-ViT / OWLv2 on NVIDIA DeepStream 9.0. You say what to look for in plain words, and you can change those words while the video is running - the new classes show up on the next frame, no restart.
Three models, picked with --model:
--model |
model | input |
|---|---|---|
owlvit (default) |
google/owlvit-base-patch32 (ViT-B/32) |
768x768 |
nanoowl |
google/owlvit-base-patch16 (ViT-B/16) |
768x768 |
owlv2 |
google/owlv2-base-patch16-ensemble |
960x960 |
owlvit is the fastest, owlv2 the most accurate, nanoowl in between. nanoowl is the ViT-B/16 model
from NVIDIA's NanoOWL project, served here through
DeepStream.
owlv2 on the demo clip with the prompt owl . eye . - open-vocabulary detection of the whole owl
and a part (its eye).
OWL-ViT looks at the image and at the words in your prompt separately, then matches regions to phrases. Because the two are independent, changing the prompt only re-runs the small text part - the image model keeps running untouched, so new classes appear on the very next frame.
DeepStream 9.0 requirements (source)
- Ubuntu 24.04
- NVIDIA driver >= 590.48.01
- CUDA 13.1
- TensorRT 10.14.1.48
- GStreamer 1.24.2
This project uses the Docker image (
nvcr.io/nvidia/deepstream:9.0-samples-multiarch), which bundles all of the above - only the host driver needs to meet the minimum version.
Host requirements
- NVIDIA GPU
- NVIDIA driver >= 590.48.01
- Docker + NVIDIA Container Toolkit (
--gpus allmust work):docker run --rm --gpus all nvcr.io/nvidia/deepstream:9.0-samples-multiarch nvidia-smi
- Internet on first setup:
02_make_onnx.shpulls the OWL-ViT / OWLv2 weights from Hugging Face (anonymous, no account needed);get_test_videos.shfetches the demo clips.
Steps 0-4 are one-time setup.
chmod +x scripts/*.sh
M=owlv2 # or owlvit / nanoowl
./scripts/00_get_model.sh # CPU export image
./scripts/01_build_libs.sh # build/*.so
./scripts/02_make_onnx.sh --model $M # export the model to ONNX
./scripts/03_build_engine.sh --model $M --precision fp16 # build the TensorRT engines
./scripts/04_build_app.sh # build/owl-app
./scripts/get_test_videos.sh # data/{owl,dog_park,tomatoes}.mp4
# detect the owl and its eyes, save an annotated MP4
./scripts/run.sh --model owlv2 --video file:///workspace/data/owl.mp4 "owl . eye ."The same flags work for any model:
./scripts/run.sh --model owlvit --video file:///workspace/data/dog_park.mp4 "dog . person ."
./scripts/run.sh --model owlv2 --video file:///workspace/data/owl.mp4 "owl . eye ."While a run is going, write new phrases to the control file from another shell:
echo "owl . beak ." > /tmp/owl_promptEach phrase (separated by . or ,) is one thing to detect. --switch "<prompt>" does this
automatically once the video starts:
./scripts/run.sh --model owlv2 --video file:///workspace/data/owl.mp4 "owl . eye ." --switch "owl . beak ."Repeat --video to run several streams at once; they're shown in a tiled grid and share one prompt.
Build the engine with --max-batch at least the number of streams:
./scripts/03_build_engine.sh --model owlvit --precision fp16 --max-batch 4
./scripts/run.sh --model owlvit \
--video file:///workspace/data/owl.mp4 \
--video file:///workspace/data/dog_park.mp4 \
--video file:///workspace/data/tomatoes.mp4 \
--video file:///workspace/data/owl.mp4 \
"owl . dog . person . tomato ."Flags to run.sh that change the boxes:
| Flag | Default | Effect |
|---|---|---|
--thr |
0.10 |
confidence threshold - raise it for fewer, surer boxes |
--nms-iou |
0.50 |
overlap to suppress; <=0 disables |
--max-area |
0.92 |
drop boxes larger than this fraction of the frame; >=1 disables |
Other flags: --out FILE, --sink ELEM, --live (window), --out-w / --out-h, --bitrate,
--force-sw, --log.
MIT - see LICENSE. Third-party components have their own licenses, listed in NOTICE:
| Component | License | Notes |
|---|---|---|
| OWL-ViT / OWLv2 (Google) | Apache-2.0 | weights pulled from Hugging Face at export; not redistributed |
CLIP tokenizer (assets/clip_*) |
MIT | from openai/CLIP via Hugging Face |
| NVIDIA DeepStream SDK 9.0 + TensorRT | NVIDIA Proprietary | runs inside the Docker container |
- OWL-ViT: https://arxiv.org/abs/2205.06230 - OWLv2: https://arxiv.org/abs/2306.09683
- Transformers OWL-ViT docs: https://huggingface.co/docs/transformers/model_doc/owlvit
- DeepStream SDK: https://docs.nvidia.com/metropolis/deepstream/dev-guide/
