Seeing the World and the Self from Egocentric Video

RESELF: REconstructing the Scene and the sELF.

Kai Guan1,2,*, Minchao Jiang2,*, Ruichen WangLi2, Wentao Zhu2,†, Lei Zhang1,†

1The Hong Kong Polytechnic University    2Eastern Institute of Technology, Ningbo    *Equal contribution    Corresponding authors

TL;DR. Joint metric reconstruction of scene geometry, camera motion, and full-body motion from monocular egocentric video.

RESELF contrasts self-blind scene reconstruction, world-blind motion reconstruction, and joint scene-and-self reconstruction.
Overview. Scene reconstruction recovers the visible environment but omits the wearer, while motion reconstruction estimates the unseen body without metric scene context. RESELF jointly recovers scene geometry, camera motion, and full-body motion in one shared coordinate frame.

Method

Overview of the RESELF geometry backbone, motion head, and closed-loop feedback.
Pipeline. Given a monocular egocentric video, a Pi3-based geometry backbone predicts metric pointmaps, camera poses, and scene-aware register tokens. The recovered trajectory and pooled geometry features condition a diffusion motion head to generate SMPL-X motion, whose kinematic signal is fed back to refine camera estimation.

1 · Metric geometry

Frame-wise metric-scale and relative-egomotion objectives adapt the geometry backbone to rapid head motion and dynamic egocentric video.

2 · Conditioned motion

The recovered camera trajectory anchors global body motion, while scene-aware register tokens condition diffusion-based SMPL-X reconstruction.

3 · Kinematic feedback

A frozen motion prior propagates kinematic supervision through the trajectory condition to refine camera tracking in a closed loop.

Interactive demo

Select a scene to inspect synchronized prediction and ground truth. The GT point cloud is sparse because Ego-Exo4D derives it from Project Aria MPS, which triangulates static scene points from consecutive frames or stereo SLAM views; regions with weak texture, occlusion, or limited sensor coverage remain unobserved.

RESELF reconstruction

Loading prediction: 0%

Ground-truth reconstruction

Loading ground truth: 0%

Exocentric views (for reference only)

0:00 / 0:00

Results

State-of-the-art metric geometry and full-body motion reconstruction.

Metric geometry

MetricRESELFSecond best
ATE ↓ [mm]217263Pi3X
RPE Rot. ↓ [deg]0.3410.557Pi3X
Abs-Rel ↓0.1410.147Pi3X

Full-body motion

MetricRESELFSecond best
MPJPE ↓ [mm]109.8116.1UniEgoMotion
PA-MPJPE ↓ [mm]75.380.4UniEgoMotion
Hand-MPJPE ↓ [mm]174.6184.2UniEgoMotion
GT UniEgoMotion RESELF
Qualitative comparison of UniEgoMotion and RESELF body articulation.
Local articulation. GT, UniEgoMotion, and RESELF denote ground truth, UniEgoMotion, and RESELF, respectively. After local alignment, the bike sequence shows incorrect body articulation and foot-ground inconsistencies for UniEgoMotion, whereas RESELF recovers a more plausible local pose using geometry-aware scene context.
Top-view comparison of global body and camera trajectories.
Global alignment. The same color coding applies: GT is ground truth, UniEgoMotion is the baseline, and RESELF is our method. In the cooking sequence, UniEgoMotion captures much of the local articulation but exhibits substantial global root drift, while RESELF better preserves root position and orientation. Together, these examples highlight the complementary roles of geometry-conditioned motion synthesis and kinematic feedback.

Abstract

Complete 3D perception from egocentric video requires recovering the surrounding scene and the wearer's full-body motion in a shared metric frame. Existing methods typically address scene reconstruction and motion estimation separately: scene reconstruction methods ignore the wearer, whereas motion estimation methods lack explicit scene geometry and often depend on external trajectories. We propose RESELF (REconstructing the Scene and the sELF), a unified framework that couples deterministic metric geometry reconstruction with geometry-conditioned motion generation. RESELF adapts a geometry foundation model to egocentric video using frame-wise scale and relative-pose consistency objectives. Its camera trajectory and latent geometric features condition a diffusion model that recovers the wearer's motion, while closed-loop kinematic feedback further refines the camera head. We also curate EE4D-JSM with aligned scene, camera, and full-body motion supervision. RESELF outperforms state-of-the-art methods designed for the individual tasks across depth estimation, camera tracking, and full-body motion estimation.

BibTeX

@article{guan2026reself,
  title     = {Seeing the World and the Self from Egocentric Video},
  author    = {Guan, Kai and Jiang, Minchao and Wang, Liruichen and Zhu, Wentao and Zhang, Lei},
  journal   = {arXiv preprint},
  year      = {2026}
}