Benchmarks
Physical-AI Generationhigher is better
PAI-Bench-G
Physical-AI video-generation benchmark (CVPR 2026): an MLLM-judged 'Domain' physical-plausibility score blended with a 'Quality' visual-fidelity score. Real source videos cap the scale near 83.9; higher is better.
Benchmark source- Domain
- Physical-AI Generation
- Metric
- pts
- Orientation
- Higher is better
- Results
- 4
Ranking
| # | Model | Score | Source | Status |
|---|---|---|---|---|
| 1 | Wan 2.2 Alibaba Qwen | 82.3 | PAI-Bench (arXiv:2512.01989, Table 3)paper | unverified |
| 2 | Veo 3 Google DeepMind | 82.2 | PAI-Bench (arXiv:2512.01989, Table 3)paper | unverified |
| 3 | Cosmos Predict 2.5 NVIDIA | 81.4 | PAI-Bench (arXiv:2512.01989, Table 3)paper | unverified |
| 4 | HunyuanVideo Tencent | 77.6 | PAI-Bench (arXiv:2512.01989, Table 3)paper | unverified |
