Multimodal vehicle trajectory prediction in roundabouts with Conditional VAEs — latent interpretability and cross-roundabout zero-shot transfer.
▶ Live demo — watch the model prune 20 futures down to one, in real time.
Predicting a vehicle's trajectory in a roundabout is an intrinsically multimodal problem: given the same past, the driver may keep circulating or take the next exit, and a deterministic regressor trained with MSE collapses those incompatible futures into an average that never happens. We model the conditional distribution p(y|c) instead, and sample K = 20 plausible futures per vehicle.
- ~126,000 windows (2 s of past → 4 s of future) built from the rounD drone recordings, expressed in an agent-centered reference frame.
- A 13-model ladder: physical baselines (constant velocity, curvature-aware) → deterministic regressors → GMM/MDN → autoencoders → the CVAE family, so every gain is measured against the step below.
- Chosen model: a CVAE with a temporal-convolutional (TCN) encoder, a learned prior, and a latent space of only 2 dimensions.
- Interpretable latent, no dimensionality reduction needed: the 2D latent organizes itself as a maneuver selector (NMI 0.50 with the exit arm), and its variance explains the spread of the prediction fan.
- Zero-shot transfer: evaluated on two roundabouts never seen in training, the CVAE family retains the lowest absolute error.
- Honest calibration analysis: the model is conservative — its fans are wider than necessary rather than overconfident.
| Metric | CVAE-TCN |
|---|---|
| minADE₂₀ | 0.517 ± 0.020 m |
| minFDE₂₀ | 1.337 ± 0.040 m |
| Exit-arm accuracy (intent) | 0.991 |
Full experimental detail, ablations and the zero-shot study are in the published report: doi.org/10.5281/zenodo.21365671 (also in docs/report.pdf, with the AI Fest poster in docs/poster.pdf).
This repository ships no rounD data. The rounD dataset (RWTH Aachen / leveLXdata) is distributed under a non-commercial license that requires each user to request access individually — once granted, point data/raw/ to your download as described in data/README.md. Everything derived (caches, checkpoints' training data) is regenerated locally by the pipeline.
Trained model weights are included in checkpoints/ so the notebooks' evaluation sections can run without retraining.
├── src/ # reusable library: geometry, windowing, models (13), ELBO losses,
│ # metrics (minADE/minFDE/intent/calibration), trainer, evaluation
├── notebooks/ # the work, phase by phase (00 orchestrator → 09 final model)
├── checkpoints/ # trained weights incl. the final CVAE-TCN (~260 KB total)
├── scripts/ # CLI utilities
├── tests/ # unit tests for geometry, windowing and losses
├── demo/ # interactive web demo (synthetic scenarios, real predictions)
└── docs/ # published report (EN), AI Fest poster, key figures
cvae-roundabout-demo.vercel.app — vehicles drive
through a roundabout while the trained checkpoint proposes 20 futures per step and evidence
prunes the fan until one trajectory locks. Because the rounD license forbids redistributing
data, the demo's traffic is synthetic (procedural paths, see demo/); the
prediction fans are real samples from the published model.
pip install -r requirements.txt
# 1) request rounD access, then link your download:
# see data/README.md
# 2) open the notebooks in order (narrative is in Spanish):
jupyter lab notebooks/00_orquestador.ipynbLuis Embon Strizzi · Alan Yeger — Universidad de San Andrés.
If you use this work, please cite the Zenodo report (DOI above).
Built as our Machine & Deep Learning capstone at UdeSA and presented at the university's AI Fest (July 2026). Code under MIT; the report is CC BY 4.0.
