Understanding how humans interact with objects is key to building robust human-centric artificial intelligence. However, this area remains relatively unexplored due to the lack of large-scale datasets. Recent datasets focusing on this issue mainly consist of activities captured entirely in controlled lab environments, and contact annotations are mostly estimated using threshold clips. We introduce Contact4D, a multi-view video dataset for human-object interaction that provides detailed body poses and accurate contact annotations. We use a flexible multi-view capture system to record individuals performing furniture assembly tasks and provide annotations for human detection, tracking, 2D/3D pose estimation, and ground-truth contact. Additionally, we propose a novel processing pipeline to extract accurate hand poses even when they are severely occluded. Contact4D consists of 2M images captured from 19 synchronized cameras across 350 video sequences, spanning diverse environments, varioius furniture types, and unique subjects. We evaluate existing methods for human pose estimation and human-centric contact estimation, demonstrating their inability to generalize to our dataset. Lastly, we fine-tune a pretrained MultiHMR model on Contact4D and observe an improved performance of 56.6% body MPJPE and 26.4% hand MPJPE in scenarios under severe self-occlusion and object occlusion.
Subjects assembling furniture across distinct rooms and object types.
Per-frame intrinsics and extrinsics for all 18 exo cameras + the Aria egocentric device.
2D and 3D body + hand keypoints, in world and camera space.
Fitted SMPL, SMPL-X, and MANO parameters for every annotated frame.
17 synchronized exo cameras ring each subject, plus one head-worn Aria egocentric device.
Raw images from Contact4D — spanning distinct subjects, rooms, and furniture-assembly tasks.
Raw image, 2D body+hand keypoints, SMPL, SMPL-X, MANO, and per-fingertip contact state.
session_03, exo cam05, sequence 007.
Four panels side by side: raw image, 2D body+hand pose, MANO mesh, and SMPL-X mesh.
session_06, exo cam15, sequence 007.
Four panels side by side: raw image, 2D body+hand pose, MANO mesh, and SMPL-X mesh.
session_24, exo cam09, sequence 006.
Four panels side by side: raw image, 2D body+hand pose, MANO mesh, and SMPL-X mesh.
session_02, Aria egocentric view, sequence 002.
Four panels side by side: raw image, 2D body+hand pose, MANO mesh, and SMPL-X mesh.
session_26, exo cam09, sequence 003.
Four panels side by side: raw image, 2D body+hand pose, MANO mesh, and SMPL-X mesh.
session_30, exo cam09, sequence 003.
Four panels side by side: raw image, 2D body+hand pose, MANO mesh, and SMPL-X mesh.
session_34, exo cam09, sequence 003.
Four panels side by side: raw image, 2D body+hand pose, MANO mesh, and SMPL-X mesh.
session_20, exo cam09, sequence 003.
Four panels side by side: raw image, 2D body+hand pose, MANO mesh, and SMPL-X mesh.
session_01, exo cam03.
session_07, exo cam01.
session_09, exo cam09.
session_13, exo cam09.
session_14, exo cam05.
session_18, exo cam09.
Per-fingertip contact state shown as a fixed HUD panel per hand: red = in contact, green = not.
@inproceedings{song2026contact4d,
title={Contact4d: A video dataset for whole-body human motion and finger contact in dexterous operations},
author={Song, Jyun-Ting and Kim, JungEun and Cao, Jinkun and Lei, Yu and Yagi, Takuma and Kitani, Kris},
booktitle={2026 International Conference on 3D Vision (3DV)},
pages={904--914},
year={2026},
organization={IEEE}
}