Crosshair
Benchmarks
Physical-AI Generationhigher is better

PAI-Bench-G

Physical-AI video-generation benchmark (CVPR 2026): an MLLM-judged 'Domain' physical-plausibility score blended with a 'Quality' visual-fidelity score. Real source videos cap the scale near 83.9; higher is better.

Benchmark source
Domain
Physical-AI Generation
Metric
pts
Orientation
Higher is better
Results
4

Ranking

#ModelScoreSourceStatus
1Wan 2.2
Alibaba Qwen
82.3PAI-Bench (arXiv:2512.01989, Table 3)paperunverified
2Veo 3
Google DeepMind
82.2PAI-Bench (arXiv:2512.01989, Table 3)paperunverified
3Cosmos Predict 2.5
NVIDIA
81.4PAI-Bench (arXiv:2512.01989, Table 3)paperunverified
4HunyuanVideo
Tencent
77.6PAI-Bench (arXiv:2512.01989, Table 3)paperunverified