Benchmark results are presented as objective measurements, and they are, of a very specific thing. Comparing them across reviews requires conditions that are rarely stated and often not met.
Thermal state dominates repeated runs
Processors run faster while cool and reduce speed as temperature rises. A benchmark run on a cold device therefore reports a figure the device cannot sustain.
Running the same test repeatedly produces a declining series, and where a reviewer stops determines the published number. A first run and a fifth run can differ substantially.
This is why sustained performance tests exist separately, and why peak scores describe capability under ideal conditions rather than behaviour during long tasks.
Ambient conditions are part of the measurement
Room temperature changes how quickly heat leaves a device, so identical hardware tested in different rooms produces different sustained results.
Surface matters too, since a device on a soft surface sheds heat more slowly than one on a hard desk, and laptops in particular are sensitive to this.
Reviewers who report ambient temperature are not being pedantic; without it the sustained figures cannot be compared to anyone else's.
Background activity contaminates results
A newly configured device performs indexing, backup, synchronisation and update tasks in the background for some time after setup, all of which compete for exactly the resources being measured.
Consistent methodology therefore requires a settled device in a known state, with those tasks complete and network activity quiet, which is difficult to arrange inside a short review window.
Preinstalled software differs between markets and retail channels as well, so two nominally identical machines can carry different background loads before anyone has installed anything.
Versions change what is being measured
Benchmark suites are revised, and scores from different major versions are not comparable even when the name is identical.
Compilers, drivers and firmware also change, so the same hardware and same test can produce different numbers months apart without anything being wrong.
Long comparison tables assembled from results gathered over a year therefore mix conditions in ways the presentation does not reveal.
Synthetic and real workloads diverge
Synthetic tests isolate a component under controlled load, which makes them repeatable and only loosely connected to how applications behave.
Application-based tests measure something closer to real use but depend on the version of that application and on how the task was configured.
The reasonable use of both is as a rough ordering rather than a precise ranking, because the difference between adjacent results is usually smaller than the variation between test conditions.