Stage 1 — Visual geometry world model (geometry pretraining). A DVGT-2 geometry encoder maps the multiview image history into multi-level geometry tokens, which describe the spatial structure seen from each camera, and ego tokens, which carry ego-motion context. Together they form a history memory. A future geometry decoder starts from learned geometry queries that carry temporal, view, and 2D positional embeddings; causal temporal self-attention models how each spatial location evolves across the future steps, and cross-attention to the history memory retrieves the relevant observed context. A shared geometry head decodes the resulting features into a dense point map and a per-pixel confidence map for every future step. Supervision comes from cosine feature alignment against targets produced by the same encoder on the future frames (detached, and never given to the forecasting branch or used at inference), together with a dense point-map loss (targets generated by MoGe-2) combining Euclidean regression, confidence-aware regression, and multi-scale surface-normal consistency, applied to both the future steps and the current frame.
Stage 2 — Geometry-conditioned action head (planning finetuning). GeoWAM adopts an inverse-dynamics-like formulation: rather than regressing a trajectory directly, learned future ego queries cross-attend to both the history memory and the predicted future geometry to produce future ego tokens that encode the ego motion compatible with the forecast scene evolution. A stop-gradient on the predicted geometry keeps the trajectory loss from reshaping the pretrained forecasting capability. The action head refines the historical ego-token sequence with a causal temporal transformer, conditions a learnable trajectory query on the predicted future ego tokens, and regresses a single trajectory directly — with no trajectory anchors, mode classification, or iterative sampling. Planning finetuning keeps the future and current geometry objectives, adds an ℓ1 trajectory loss, and adds an auxiliary ℓ1 loss on relative historical poses.
We evaluate on the nuScenes validation set; none of the compared models is trained on nuScenes. Each predicted future point map is converted to ray depth, and we report absolute relative error (Abs Rel, lower is better) and threshold accuracy δ < 1.25 (higher is better) at horizons from one to four seconds. Video world models generate future RGB frames, so we reconstruct their geometry with DVGT to give every method a common geometric output. PAI denotes the NVIDIA PhysicalAI-Autonomous-Vehicles dataset used for geometry pretraining.
| Method / Horizon | Abs Rel ↓ | δ < 1.25 ↑ | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 1s | 2s | 3s | 4s | mean | 1s | 2s | 3s | 4s | mean | |
| Epona + DVGT | 0.229 | 0.263 | 0.292 | 0.310 | 0.274 | 0.732 | 0.677 | 0.620 | 0.589 | 0.655 |
| Cosmos 3 + DVGT | 0.300 | 0.376 | 0.405 | 0.422 | 0.376 | 0.588 | 0.513 | 0.464 | 0.447 | 0.503 |
| VGGT-World | 0.272 | 0.329 | 0.342 | 0.357 | 0.325 | 0.612 | 0.553 | 0.513 | 0.497 | 0.544 |
| GeoWAM (ours) w/o PAI pretraining | 0.171 | 0.190 | 0.227 | 0.245 | 0.208 | 0.849 | 0.813 | 0.777 | 0.716 | 0.789 |
| GeoWAM (ours) w/ PAI pretraining | 0.146 | 0.176 | 0.208 | 0.276 | 0.202 | 0.873 | 0.835 | 0.797 | 0.704 | 0.802 |
Future geometry prediction on nuScenes at different horizons. Bold and underlined values indicate the best and second-best results. Both GeoWAM variants beat every baseline at every horizon. PAI pretraining helps most at short horizons (Abs Rel 0.171 → 0.146 at 1 s and 0.227 → 0.208 at 3 s), while the variant without pretraining remains slightly better at the 4 s horizon.
navtest
We finetune GeoWAM on the NAVSIM navtrain split and evaluate with the Extended Predictive Driver Model Score (EPDMS), which aggregates no at-fault collision (NC), drivable-area compliance (DAC), driving-direction compliance (DDC), traffic-light compliance (TLC), ego progress (EP), time-to-collision (TTC), lane keeping (LK), history comfort (HC), and extended comfort (EC). All metrics use the official human-penalty protocol; higher is better.
| Method | NC ↑ | DAC ↑ | DDC ↑ | TLC ↑ | EP ↑ | TTC ↑ | LK ↑ | HC ↑ | EC ↑ | EPDMS ↑ |
|---|---|---|---|---|---|---|---|---|---|---|
| Transfuser | 96.9 | 89.9 | 97.8 | 99.7 | 87.1 | 95.4 | 92.7 | 98.3 | 87.2 | 84.0 |
| Hydra-MDP++ | 97.2 | 97.5 | 99.4 | 99.6 | 83.1 | 96.5 | 94.4 | 98.2 | 70.9 | 81.4 |
| DriveSuprim | 97.5 | 96.5 | 99.4 | 99.6 | 88.4 | 96.6 | 95.5 | 98.3 | 77.0 | 83.1 |
| ARTEMIS | 98.3 | 95.1 | 98.6 | 99.8 | 81.5 | 97.4 | 96.5 | 98.3 | – | 83.1 |
| DiffusionDrive | 98.2 | 96.2 | 99.5 | 99.8 | 87.4 | 97.3 | 96.9 | 98.4 | 87.7 | 88.2 |
| WoTE | 98.5 | 96.8 | 98.8 | 99.8 | 86.1 | 97.9 | 95.5 | 98.3 | 82.9 | 87.7 |
| DriveVLA-W0 | 98.4 | 95.2 | 99.4 | 99.9 | 86.6 | 97.9 | 97.8 | 98.3 | 82.7 | 86.9 |
| PWM | 98.8 | 95.9 | 99.4 | 99.9 | 86.4 | 98.4 | 97.6 | 98.3 | 85.3 | 88.2 |
| DriveLaW | 98.7 | 96.9 | 99.6 | 99.8 | 87.5 | 98.3 | 97.6 | 98.4 | 77.4 | 88.6 |
| DVGT-2 | 98.7 | 97.9 | 99.7 | 99.9 | 87.9 | 98.0 | 98.2 | 98.2 | 77.0 | 89.6 |
| EponaV2 | 98.5 | 97.4 | 99.5 | 99.9 | 87.9 | 98.1 | 97.7 | 98.2 | 77.4 | 88.9 |
| GeoWAM (ours) w/o PAI pretraining | 98.7 | 97.7 | 99.7 | 99.9 | 87.0 | 98.1 | 97.9 | 98.3 | 86.8 | 90.2 |
| GeoWAM (ours) w/ PAI pretraining | 99.0 | 97.9 | 99.7 | 99.9 | 86.7 | 98.5 | 97.9 | 98.3 | 87.8 | 90.7 |
Results on the NAVSIM v2 navtest split. Without geometry pretraining GeoWAM reaches an EPDMS of 90.2; geometry pretraining on PhysicalAI raises it to 90.7, the best overall score in the table.
navhard
The navhard benchmark approximates closed-loop evaluation with scenes reconstructed by 3D Gaussian Splatting: after the planner predicts a trajectory, a new observation is rendered from the resulting ego pose and fed back for the next planning step, so planning errors propagate into subsequent observations. Stage 1 (S1) evaluates the original scenes and Stage 2 (S2) the synthetic reactive scenes.
| Method | Stage | NC ↑ | DAC ↑ | DDC ↑ | TLC ↑ | EP ↑ | TTC ↑ | LK ↑ | HC ↑ | EC ↑ | EPDMS ↑ |
|---|---|---|---|---|---|---|---|---|---|---|---|
| DriveVLA-W0 | S1 | 96.8 | 83.3 | 99.0 | 99.6 | 84.6 | 95.3 | 96.4 | 97.6 | 78.2 | 24.4 |
| S2 | 76.8 | 64.3 | 79.9 | 98.3 | 89.2 | 75.0 | 46.8 | 95.8 | 53.1 | ||
| DriveLaW | S1 | 97.3 | 89.1 | 99.2 | 99.6 | 84.3 | 97.1 | 96.2 | 97.8 | 67.6 | 30.6 |
| S2 | 82.5 | 67.6 | 83.5 | 98.1 | 84.8 | 78.5 | 45.8 | 96.4 | 57.3 | ||
| DVGT-2 | S1 | 97.2 | 91.3 | 98.4 | 99.8 | 84.8 | 95.5 | 95.5 | 97.5 | 71.4 | 31.7 |
| S2 | 77.8 | 73.8 | 81.3 | 98.3 | 91.5 | 73.2 | 48.0 | 83.9 | 45.1 | ||
| LTFv6 | S1 | 96.5 | 86.6 | 99.2 | 99.5 | 84.4 | 95.1 | 94.4 | 97.7 | 76.4 | 31.9 |
| S2 | 79.8 | 75.5 | 86.2 | 97.8 | 89.5 | 76.0 | 50.0 | 95.2 | 66.7 | ||
| NavFormer† | S1 | 96.2 | 92.4 | 95.7 | 99.6 | 83.8 | 96.0 | 94.7 | 96.4 | 60.9 | 34.1 |
| S2 | 85.7 | 81.0 | 83.5 | 97.6 | 90.1 | 82.4 | 48.2 | 94.9 | 48.4 | ||
| EponaV2† | S1 | 97.3 | 90.7 | 99.4 | 100.0 | 83.3 | 97.3 | 97.3 | 97.6 | 60.9 | 36.1 |
| S2 | 83.6 | 78.0 | 88.0 | 98.9 | 86.0 | 80.3 | 50.1 | 96.1 | 52.0 | ||
| RAP-DINO† | S1 | 97.1 | 94.4 | 98.8 | 99.8 | 83.9 | 96.9 | 94.7 | 96.4 | 66.2 | 36.9 |
| S2 | 83.2 | 83.9 | 87.4 | 98.0 | 86.9 | 80.4 | 52.3 | 95.2 | 52.4 | ||
| GeoWAM (ours) w/o PAI pretraining | S1 | 97.7 | 91.5 | 99.1 | 99.8 | 83.8 | 95.8 | 96.0 | 97.8 | 79.0 | 36.6 |
| S2 | 80.4 | 76.3 | 87.3 | 98.7 | 88.9 | 76.2 | 49.9 | 94.0 | 56.0 | ||
| GeoWAM (ours) w/ PAI pretraining | S1 | 98.9 | 92.9 | 99.6 | 100.0 | 82.0 | 97.3 | 96.9 | 97.8 | 79.5 | 39.6 |
| S2 | 84.1 | 76.0 | 87.4 | 98.6 | 84.0 | 80.2 | 50.5 | 96.9 | 63.7 |
navhard leaderboard. Methods trained with reinforcement learning or PDMS-score supervision are shown in gray and marked with †; among the remaining methods, bold and underlined values indicate the best and second-best results. GeoWAM uses no PDMS-derived supervision (no RL or learned PDMS scorer). Without PAI pretraining it reaches an EPDMS of 36.6, outperforming all baselines trained without PDMS supervision as well as two that use it, and trailing only RAP (36.9). Geometry pretraining on PAI raises the score by 3.0 points to 39.6, surpassing all baselines including RAP.
To assess cross-dataset generalization, we evaluate GeoWAM on nuScenes without any nuScenes-specific trajectory supervision. Both planners are trained only on NAVSIM navtrain; the pretrained variant additionally uses image sequences from navtrain and PhysicalAI for geometry pretraining. We report L2 displacement error and collision rate at 1, 2, and 3 seconds and their averages.
| Method | nuScenes Finetune |
Auxiliary Supervision |
L2 (m) ↓ | Collision Rate (%) ↓ | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 1s | 2s | 3s | Avg. | 1s | 2s | 3s | Avg. | |||
| ST-P3 | ✓ | Map&Box&Depth | 1.33 | 2.11 | 2.90 | 2.11 | 0.23 | 0.62 | 1.27 | 0.71 |
| UniAD | ✓ | Map&Box&Motion | 0.48 | 0.96 | 1.65 | 1.03 | 0.05 | 0.17 | 0.71 | 0.31 |
| OccNet | ✓ | 3D-Occ&Map&Box | 1.29 | 2.13 | 2.99 | 2.14 | 0.21 | 0.59 | 1.37 | 0.72 |
| OccWorld | ✓ | 3D-Occ | 0.52 | 1.27 | 2.41 | 1.40 | 0.12 | 0.40 | 2.08 | 0.87 |
| VAD-Tiny | ✓ | Map&Box&Motion | 0.60 | 1.23 | 2.06 | 1.30 | 0.31 | 0.53 | 1.33 | 0.72 |
| VAD-Base | ✓ | Map&Box&Motion | 0.54 | 1.15 | 1.98 | 1.22 | 0.04 | 0.39 | 1.17 | 0.53 |
| GenAD | ✓ | Map&Box&Motion | 0.36 | 0.83 | 1.55 | 0.91 | 0.06 | 0.23 | 1.00 | 0.43 |
| Doe-1 | ✓ | QA | 0.50 | 1.18 | 2.11 | 1.26 | 0.04 | 0.37 | 1.19 | 0.53 |
| Epona | ✓ | Future RGB | 0.61 | 1.17 | 1.98 | 1.25 | 0.01 | 0.22 | 0.85 | 0.36 |
| PWM | ✗ | Future RGB | 2.06 | 3.91 | 6.00 | 3.99 | 0.12 | 0.15 | 0.86 | 0.36 |
| DriveVLA-W0 | ✗ | Future RGB | 0.43 | 1.26 | 2.60 | 1.43 | 0.22 | 0.66 | 1.42 | 0.77 |
| GeoWAM (ours) w/o PAI pretraining | ✗ | Future Geometry | 0.42 | 1.19 | 2.24 | 1.28 | 0.00 | 0.10 | 0.60 | 0.24 |
| GeoWAM (ours) w/ PAI pretraining | ✗ | Future Geometry | 0.29 | 0.79 | 1.58 | 0.89 | 0.02 | 0.12 | 0.23 | 0.12 |
Zero-shot planning on nuScenes; bold and underlined values indicate the best and second-best results. Even without PAI pretraining, GeoWAM (avg. L2 1.28 m, collision rate 0.24%) outperforms the zero-shot video-based world models PWM and DriveVLA-W0 and has a lower collision rate than every baseline, including planners finetuned on nuScenes. Geometry pretraining further reduces the averages to 0.89 m and 0.12%, surpassing all nuScenes-finetuned baselines on both average metrics without any additional trajectory supervision.
Geometry pretraining learns future scene geometry directly from image sequences, so it can use additional driving data without trajectory annotations. We vary the amount of PhysicalAI data from 0 to 832 hours (the full training set). Scaling the pretraining data improves navhard EPDMS from 36.6 to 39.6 and reduces the nuScenes zero-shot collision rate from 0.24% to 0.12%. The collision rate decreases consistently with more data, whereas nuScenes L2 is not strictly monotonic since it measures deviation from the recorded human trajectory and may penalize other valid behaviors. GeoWAM has only 1.9B parameters, substantially fewer than DriveVLA-W0 and EponaV2, yet keeps improving with more pretraining data, without extra model capacity or trajectory supervision.
| PAI Data (hours) |
navhard |
nuScenes Zero-shot | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| EPDMS ↑ | L2 (m) ↓ | Collision Rate (%) ↓ | |||||||||
| S1 | S2 | Comb. | 1s | 2s | 3s | Avg. | 1s | 2s | 3s | Avg. | |
| 0 | 80.7 | 43.5 | 36.6 | 0.42 | 1.19 | 2.24 | 1.28 | 0.00 | 0.10 | 0.60 | 0.24 |
| 100 | 81.3 | 45.5 | 37.8 | 0.31 | 0.81 | 1.64 | 0.92 | 0.02 | 0.12 | 0.44 | 0.19 |
| 496 | 81.2 | 45.8 | 38.2 | 0.29 | 0.76 | 1.55 | 0.86 | 0.04 | 0.12 | 0.33 | 0.16 |
| 832 | 83.1 | 46.7 | 39.6 | 0.29 | 0.79 | 1.58 | 0.89 | 0.02 | 0.12 | 0.23 | 0.12 |
Effect of geometry-pretraining data scale on downstream planning. Table: absolute performance at different pretraining-data scales. Plot: relative improvement over the model without geometry pretraining — up to +8.2% combined EPDMS on navhard and a 50% lower collision rate on nuScenes at 832 hours.
| Variant | Current Geometry | Future Geometry | PAI Pretraining |
navtestEPDMS ↑ | navhardEPDMS ↑ |
|---|---|---|---|---|---|
| w/o Future Geometry | ✓ | 89.2 | 31.7 | ||
| w/o PAI Pretraining | ✓ | ✓ | 90.2 | 36.6 | |
| Full | ✓ | ✓ | ✓ | 90.7 | 39.6 |
Future geometry and PAI pretraining
Future geometry and PAI pretraining. Removing future-geometry prediction while keeping current-frame geometry supervision lowers EPDMS from 90.2 to 89.2 on navtest and from 36.6 to 31.7 on navhard, isolating the benefit of forecasting future geometric evolution. PAI pretraining adds a further 0.5 points on navtest and 3.0 points on navhard; the larger navhard gain suggests that more diverse road layouts and traffic scenes strengthen the learned geometric dynamics, particularly when planning errors affect subsequent observations.
Ego-motion recovery from predicted future states. Each row shows current observations, PWM-generated future video, GeoWAM-predicted future geometry, and trajectories recovered from each future representation, compared with the logged ground truth.
To examine how directly different predicted states expose future ego motion, we recover trajectories from PWM's generated video (with classical epipolar geometry and with VGGT) and from GeoWAM's future geometry (matching static 3D keypoints across predicted point maps and solving for rigid transforms with Kabsch alignment). The geometry-derived trajectories follow the ground truth much more closely, whereas those recovered from future video deviate substantially — future geometry makes ego motion more directly accessible than pixel-space predictions.
If you find our work helpful, please consider cite us:
@article{lu2026geowam,
title={GeoWAM: Visual Geometry World Action Models for Autonomous Driving},
author={Lu, Yiren and Ye, Xin and Liu, Jiaming and Jacobson, Philip and Yao, Jin and Chen, Yi-chung and Merino, Liam and Kurra, Dhruva Dixith and Cai, Min and Lampo, Tom and others},
journal={arXiv preprint arXiv:2608.23486},
year={2026}
}