Skip to content

Add evaluation tests and run them in CI - #27

Open
nushakrishnan wants to merge 9 commits into
cvg:mainfrom
nushakrishnan:tests/eval-coverage
Open

nushakrishnan wants to merge 9 commits into
cvg:mainfrom
nushakrishnan:tests/eval-coverage

Conversation

@nushakrishnan

Copy link
Copy Markdown
Collaborator

The tests use a small synthetic scene (a stereo rig moving past a set of control points, with the estimate in a different Sim(3) frame), so no dataset download is needed.

Covered:

  • Control point triangulation on full and partial trajectories.
  • Control point evaluation: alignment recovery and scoring, missing points on partial trajectories, clean failure when too few points remain.
  • Regression test for Bugfix: Control Point Triangulation #22: a displaced triangulation keeps its error and the scored points are not modified by the refinement (fails on the pre-fix code).
  • Saving results: .npy round trip (relies on Fix loading saved evaluation results and script exit codes #24) and the saved alignment consumed by evaluate_wrt_pgt.
  • Partial trajectory evaluation w.r.t. pGT and MPS.

CI: new tests.yml workflow runs python -m pytest tests on Python 3.10 and 3.12, installing only the evaluation dependencies. Also adds a testpaths entry to pyproject.toml and a short "Running the tests" section to the README.

Comment thread tests/test_control_point_evaluation.py Outdated
Comment thread tests/test_control_point_evaluation.py
Comment thread .github/workflows/tests.yml Outdated
@nushakrishnan

Copy link
Copy Markdown
Collaborator Author

Done, when fewer than three control points with height triangulate, the evaluation now saves a result with no alignment and reports CP score and recall of 0 with exit code 0. evaluate_wrt_pgt also reports pose recall 0 for such a result instead of failing.

@B1ueber2y

Copy link
Copy Markdown
Member

Done, when fewer than three control points with height triangulate, the evaluation now saves a result with no alignment and reports CP score and recall of 0 with exit code 0. evaluate_wrt_pgt also reports pose recall 0 for such a result instead of failing.

Thanks. For the evaluation wrt pGT, when the CP alignment fails, should we still try pGT alignment instead of directly assigning zero?

@nushakrishnan

Copy link
Copy Markdown
Collaborator Author

Done, when fewer than three control points with height triangulate, the evaluation now saves a result with no alignment and reports CP score and recall of 0 with exit code 0. evaluate_wrt_pgt also reports pose recall 0 for such a result instead of failing.

Thanks. For the evaluation wrt pGT, when the CP alignment fails, should we still try pGT alignment instead of directly assigning zero?

Our standard alignment comes from the control points, so if an alignment from that isn't found then I don't think it makes sense to do separate pGT alignment. The pGT alignment is only for sequences without CPs.

@B1ueber2y

B1ueber2y commented Oct 9, 2026 •

Copy link
Copy Markdown
Member

Assuming that one has an output that covers a span of only two control points (say out of ten), the pGT recall should probably not be zero but somewhere around 20%. The point of CP alignment for pGT evaluation is just to find a good transformation for it.

@nushakrishnan

nushakrishnan commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator Author

Assuming that one has an output that covers a span of only two control points (say out of ten), the pGT recall should probably not be zero but somewhere around 20%. The point of CP alignment for pGT evaluation is just to find a good transformation for it.

Partially evaluating a trajectory in that case does make sense even though the CP alignment fails. Let me update it accordingly. It goes against our current eval strategy but it makes sense to give a partial pose recall and 0 for CP score.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants