Crosshair
Leaderboard
World ModelsOpen weights

V-JEPA 2

Self-supervised video joint-embedding predictive architecture; learns world dynamics for motion understanding, action anticipation, and zero-shot robot planning.

videoactionembodiedOfficial site
Crosshair Index
#1 of 10 · World Models
Provider
Meta AI
Released
2025-06-11
Parameters
1.2B

Capability web

Strengths across world-modeling capabilities — each axis is one benchmark domain. Coverage is sparse while these evaluations mature; empty axes read “no data yet.”

UnderstandingAnticipationGen. PhysicsPhysical AI

Video Understanding

81
skill

Recognizing motion and answering questions about real-world video — what is happening and how things move, beyond static appearance.

Scorecard

BenchmarkScoreSourceStatus
Something-Something v2
Motion Understanding
77.3%bestMeta — V-JEPA 2 (arXiv:2506.09985)paperunverified
EPIC-Kitchens-100 Anticipation
Action Anticipation
39.7bestMeta — V-JEPA 2 (arXiv:2506.09985)paperunverified
Perception Test
Video Understanding
84%bestMeta — V-JEPA 2 (arXiv:2506.09985)paperunverified
Physics-IQ
Generative Physics
not evaluated
PAI-Bench-G
Physical-AI Generation
not evaluated