In submission · 2026

PACT

End-to-End Learning of Human Pose, Contacts, and Forces from Video

  • 1MBZUAI
  • 2ETH Zürich
  • 3EPFL

TL;DR

PACT turns a single video into world-space 3D motion, hand and foot contacts, and the forces at those contacts, in one end-to-end model. No motion capture, no force plates, no staged pipeline.


Abstract

Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual pose reconstruction from contact and force estimation. This separation limits joint reasoning and can propagate errors between stages. We introduce PACT, an end-to-end model that jointly learns to estimate human pose, contacts and contact forces from monocular video. Our approach augments a human reconstruction foundation model with learnable contact-force tokens and a temporal transformer that integrates visual features with world-space motion. Joint prediction heads refine human poses and estimate contacts and forces, while physics-based supervision encourages consistency between the reconstructed motion and interaction forces. To address the scarcity of force annotations, we develop a data annotation pipeline that combines contact labeling with physics-based motion and force optimization, producing training supervision from synthetic and real-world videos. We also introduce a real-world climbing benchmark ForceWall with climbing videos and corresponding ground-truth contact forces obtained from the force sensors. Experiments demonstrate state-of-the-art contact and force estimation, outperforming staged reconstruction approaches and generalizing to interactions beyond the training distribution. These results support end-to-end joint learning as an effective approach to recovering human motion and physical interactions from video.

Video

PACT in motion

A narrated walkthrough of the motivation, the labeling pipeline, the model, and results.

Training data

Physics-based labels from ordinary video

Force labels are scarce, so we built a physics-based labeling pipeline that turns ordinary videos into supervision for pose, contacts, and forces.

Annotation pipeline: foundation models recover the person, the scene, and the camera, then a physics-based optimization solves for consistent motion, contacts, and forces used as training labels.

Annotation pipeline

  1. Foundation models recover the person, the scene, and the camera.
  2. An optimization solves for physically consistent motion, contacts, and forces.
  3. We use it to label synthetic BEDLAM2 videos and real climbing videos from the web.
  4. These labels are imperfect, with noisy forces and forces on limbs that touch nothing. Learning from many of them, PACT becomes more robust than its own teacher.

The pipeline on a real video

Every stage on the same bouldering clip, played in sync.

SAM 3person tracking
Sapiens 22D keypoints
SAM 3D Bodybody estimate
VGGT-Omegawall and camera
Physics optimizationcontacts and forces

Physics-based labels

The labels on the input video, the reconstruction seen from the camera, and a side view of the same motion in world space.

Red arrows = forces solved by the pipeline

Model

From stages to end-to-end

Existing methods estimate pose, contacts, and forces in separate stages, so errors in one stage propagate to the next. PACT learns all three jointly, end to end, from a single video.

PACT architecture: a frozen SAM 3D Body backbone with learnable contact-force tokens, camera motion lifting the pose to world space, a temporal transformer, and heads that predict pose, gravity, contacts, and forces with a physics loss.

Architecture

  1. PACT builds on the SAM 3D Body backbone, kept frozen.
  2. Learnable contact-force tokens capture interaction cues in every frame.
  3. Camera motion lifts the pose into world space, and a temporal transformer reasons over time.
  4. Heads refine the pose and predict gravity, contacts, and forces, with a physics loss tying motion and forces together.

Results

From a single video to physics

Every clip below is predicted by PACT from a single monocular video: the input with the prediction drawn over it, the reconstruction seen from the camera, and a side view in world space.

Climbing in world space

Yellow arrows = predicted forces

Beyond climbing

Interactions outside the training distribution, shown the same way: input with PACT, camera view, and side view.

Balance beam
Yoga
Box step
Backflip
Squat walks
Aerial
Back handspring
Back walkover
Tic-tac

Benchmark

ForceWall

ForceWall is our instrumented climbing wall: its holds measure the force applied by every limb, giving real ground-truth forces alongside the video. On a climb recorded on the wall, the forces PACT predicts from the video alone follow the sensor readings.

Interactive 3D

Explore in 3D

The input video plays in sync with the 3D scene. Drag to orbit and right-drag to pan; click the scene, then scroll to zoom. On touch screens, tap the scene first, then drag to orbit, pinch to zoom, and use two fingers to pan.

Annotation pipeline

The labels solved by the physics-based optimization: the recovered wall, the camera path, contacts, and forces.

3D viewer loads when visible

Tap to explore

PACT predictions

World-space pose, contacts, and forces predicted by PACT from the video alone.

3D viewer loads when visible

Tap to explore

Citation

BibTeX

@article{akizhanov2026pact,
  title   = {PACT: End-to-End Learning of Human Pose, Contacts, and Forces from Video},
  author  = {Akizhanov, Rikhat and Zhang, Yangsong and Kaliazin, Nikolai and
             Wolf, Peter and Nakamura, Yoshihiko and Fua, Pascal and
             Pizzati, Fabio and Laptev, Ivan},
  journal = {arXiv preprint},
  year    = {2026}
}