CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
Advances the frontier of generalizable video simulation for robotics, enabling transfer across hardware platforms without retraining.
AI Summary
CLAP is a cross-embodiment video world model framework that learns physical dynamics from diverse human and robot videos and achieves zero-shot simulation across different robot morphologies.
Excerpt
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents. CLAP is grounded in the insight that universal physical l
