# Deform360 > A massive multi-view visuotactile dataset for deformable world models — 198 everyday deformable objects, 1,980 interactions, 41 surround-view cameras, bimanual tactile grippers. Accepted at ECCV 2026. Deform360 is built by researchers at Brown University, Columbia University, and MIT. The dataset provides 215.7 cumulative multi-view hours of synchronized video and tactile recordings, with markerless 3D particle annotations, enabling a controlled comparison of 2D video world models and 3D particle world models on real-world deformable dynamics. ## Core resources - [Project page](https://deform360.lhy.xyz/): Full paper summary, motivation, methodology, benchmarks. - [Dataset (HuggingFace)](https://huggingface.co/datasets/brownu/deform360): Full-resolution 41-view video plus tactile and particle annotations. - [Code & pipeline (GitHub)](https://github.com/lhy0807/deform360): Capture, reconstruction, perception, and world-model baselines. ## Key facts - 198 daily-life deformable objects across 17 categories, grouped into three material-response classes: 1D linear (ropes, cables), 2D thin-shell (cloth, paper), and 3D volumetric (plush, foam). - 1,980 robotic interactions, 215.7 cumulative multi-view hours, ~23.3M frames, 74,850 raw videos at 720p / 30 FPS. - 41 synchronized RGB cameras with full 360° coverage; bimanual tactile grippers based on the UMI platform. - Markerless visuotactile tracking pipeline: per-frame 3D Gaussian Splatting + lifted CoTracker3 2D tracks + tactile contact constraints + physics-informed optimization. - Substantially increases the scale and sensory richness of real-world deformable benchmarks with 198 objects, 41 calibrated surround views, tactile sensing, and dense markerless 3D annotations. ## Benchmark findings - Models compared: 2D video world model Cosmos versus 3D models ParticleFormer, PGND, and PhysTwin where applicable, using Chamfer distance, track error, PSNR, SSIM, and LPIPS. - With very limited per-episode data, PhysTwin's explicit physical priors outperform the learned 3D baselines; Cosmos cannot be post-trained stably from so little data. - On held-out episodes, Cosmos reconstructs appearance most faithfully, while ParticleFormer leads every reported future-prediction metric. - On entirely unseen objects, pretrained Cosmos leads PSNR and LPIPS, while ParticleFormer retains the best geometric errors. Cosmos can drift from commanded robot actions over long horizons. - Explicit particle states also support geometric planning rewards such as Chamfer distance; cross-environment appearance shift and reward design in video space remain challenges for direct video-model MPC. ## How to cite ``` @inproceedings{li2026deform360, title = {Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models}, author = {Li, Hongyu and Fu, Wanjia and Cong, Xiaoyan and Li, Zekun and Huang, Binghao and Jiang, Hanxiao and He, Xintong and Liang, Yiqing and Fu, Rao and Lu, Tao and Sridhar, Srinath and Smith, Kevin A. and Konidaris, George and Li, Yunzhu}, booktitle = {European Conference on Computer Vision (ECCV)}, year = {2026} } ``` ## Authors Hongyu Li, Wanjia Fu, Xiaoyan Cong, Zekun Li, Binghao Huang, Hanxiao Jiang, Xintong He, Yiqing Liang, Rao Fu, Tao Lu, Srinath Sridhar, Kevin A. Smith, George Konidaris, Yunzhu Li. ## License Dataset and code released under the MIT License.