Skip to content

Localize all video frames against the reconstruction - #42

Draft
sarlinpe wants to merge 1 commit into
location-priorfrom
frame-localization
Draft

sarlinpe wants to merge 1 commit into
location-priorfrom
frame-localization

Conversation

@sarlinpe

@sarlinpe sarlinpe commented Oct 9, 2026

Copy link
Copy Markdown
Member

Summary

Optional step after mapping that registers every video frame, or a subsampled rate, against the finished (sub-)reconstructions, without more triangulation or bundle adjustment.

python -m vidmap.localize_frames --run-dir <run> [--fps 0] [--batch-size 8]
  • Matching: each frame is matched with RoMaV2 to its nearest keyframe. If covisibility is low, it also tries the other bracketing keyframe, which handles frames next to cuts. Matches are lifted to the keyframe's 3D points, and the absolute pose (optionally with a refined focal length) is estimated with robust PnP.
  • Throughput: batched matching with the compiled RoMaV2 graphs on the GPU, PnP in a thread pool, and frame I/O workers. On an A100, about 25-35 ms per frame once the graphs are loaded (all ~1,800 frames of a 60 s clip in about a minute). Keyframes keep their bundle-adjusted poses.
  • Output: frame_poses_*.npz/.txt (cam_from_world, focal length, model index, inliers) and a JSON summary with timings. --leave-one-out re-localizes the keyframes to validate accuracy.
  • Viewer: optional dense trajectory of all localized frames; frusta and the slider stay on keyframes.
  • Tests for the frame selection helpers.

@sarlinpe
sarlinpe added this pull request to stack #43 October 9, 2026 15:45
@sarlinpe
sarlinpe marked this pull request as draft October 9, 2026 22:42
New optional step after mapping (python -m vidmap.localize_frames --run-dir <run>): register
every video frame (or a subsampled rate, --fps) against the finished (sub-)reconstructions
without further triangulation or bundle adjustment.

- Each frame is matched with RoMaV2 to its nearest keyframe (falling back to the other
  bracketing keyframe when covisibility is low, which handles frames next to cuts); matches
  are lifted to the keyframe's 3D points and the absolute pose (optionally with a refined
  focal length) is estimated by robust PnP.
- Batched matching with the compiled RoMaV2 graphs on the GPU, PnP in a thread pool, frames
  read by I/O workers; keyframes keep their bundle-adjusted poses.
- Output: frame_poses_<fps>fps*.npz/.txt (cam_from_world, focal, model index, inliers) and a
  JSON summary with timings; --leave-one-out re-localizes keyframes to validate accuracy.
- Viewer: optional dense trajectory of all localized frames (frusta and slider stay on
  keyframes).
- Tests for frame selection helpers.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant