1 · Metric geometry
Frame-wise metric-scale and relative-egomotion objectives adapt the geometry backbone to rapid head motion and dynamic egocentric video.
Frame-wise metric-scale and relative-egomotion objectives adapt the geometry backbone to rapid head motion and dynamic egocentric video.
The recovered camera trajectory anchors global body motion, while scene-aware register tokens condition diffusion-based SMPL-X reconstruction.
A frozen motion prior propagates kinematic supervision through the trajectory condition to refine camera tracking in a closed loop.
Select a scene to inspect synchronized prediction and ground truth. The GT point cloud is sparse because Ego-Exo4D derives it from Project Aria MPS, which triangulates static scene points from consecutive frames or stereo SLAM views; regions with weak texture, occlusion, or limited sensor coverage remain unobserved.
State-of-the-art metric geometry and full-body motion reconstruction.
| Metric | RESELF | Second best |
|---|---|---|
| ATE ↓ [mm] | 217 | 263Pi3X |
| RPE Rot. ↓ [deg] | 0.341 | 0.557Pi3X |
| Abs-Rel ↓ | 0.141 | 0.147Pi3X |
| Metric | RESELF | Second best |
|---|---|---|
| MPJPE ↓ [mm] | 109.8 | 116.1UniEgoMotion |
| PA-MPJPE ↓ [mm] | 75.3 | 80.4UniEgoMotion |
| Hand-MPJPE ↓ [mm] | 174.6 | 184.2UniEgoMotion |


Complete 3D perception from egocentric video requires recovering the surrounding scene and the wearer's full-body motion in a shared metric frame. Existing methods typically address scene reconstruction and motion estimation separately: scene reconstruction methods ignore the wearer, whereas motion estimation methods lack explicit scene geometry and often depend on external trajectories. We propose RESELF (REconstructing the Scene and the sELF), a unified framework that couples deterministic metric geometry reconstruction with geometry-conditioned motion generation. RESELF adapts a geometry foundation model to egocentric video using frame-wise scale and relative-pose consistency objectives. Its camera trajectory and latent geometric features condition a diffusion model that recovers the wearer's motion, while closed-loop kinematic feedback further refines the camera head. We also curate EE4D-JSM with aligned scene, camera, and full-body motion supervision. RESELF outperforms state-of-the-art methods designed for the individual tasks across depth estimation, camera tracking, and full-body motion estimation.
@article{guan2026reself,
title = {Seeing the World and the Self from Egocentric Video},
author = {Guan, Kai and Jiang, Minchao and Wang, Liruichen and Zhu, Wentao and Zhang, Lei},
journal = {arXiv preprint},
year = {2026}
}