The paper audits 12 physical AI benchmarks and finds that some are redundant enough to change model rankings when you merge them.
When people compare physical AI models, they often average scores across many benchmarks as if each one adds a new piece of evidence. The problem is that those benchmarks are not always independent: two tests can be measuring almost the same ability, so counting both can overweight that skill.
The authors build a score matrix for 51 models across 12 benchmarks, pulling numbers from model cards, benchmark papers, and their own runs when needed. They then look at how much information the benchmarks share. The main point is not that one benchmark is “bad,” but that some benchmarks are so similar that treating them as separate can distort the final ranking.
The non-obvious part is that this is not just a correlation report. They also test what happens to model orderings if you collapse redundant benchmarks, and they search for a smaller benchmark set that keeps most of the useful signal. That makes the paper about measurement quality, not just benchmark counting.