Apple M6 Pro achieves the highest single-core CPU score in Geekbench 7
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Apple M6 Pro achieves the highest single-core CPU score in Geekbench 7
Unofficial Hacker News client; not affiliated with Y Combinator.
GeekyBear · · focus · HN ↗
M5 Ultra CPU:
<a href="https://browser.geekbench.com/search?k=parkdale_cpu&q=mac17%2C15" rel="nofollow">https://browser.geekbench.com/search?k=parkdale_cpu&q=mac17%...
M5 Ultra GPU:
<a href="https://browser.geekbench.com/search?k=grand_gpu&q=mac17%2C15" rel="nofollow">https://browser.geekbench.com/search?k=grand_gpu&q=mac17%2C1...
Base M6 CPU:
<a href="https://browser.geekbench.com/search?k=parkdale_cpu&q=mac18%2C5" rel="nofollow">https://browser.geekbench.com/search?k=parkdale_cpu&q=mac18%...
Base M6 GPU:
<a href="https://browser.geekbench.com/search?k=grand_gpu&q=mac18%2C5" rel="nofollow">https://browser.geekbench.com/search?k=grand_gpu&q=mac18%2C5
The Base M6 CPU single core is averaging a bit over 4000 on Geekbench 7.
For comparison, the AMD Ryzen 9 9950X3D2 averages 3161 on the same test.
<a href="https://browser.geekbench.com/processors/amd-ryzen-9-9950x3d2" rel="nofollow">https://browser.geekbench.com/processors/amd-ryzen-9-9950x3d...
The Intel Core Ultra 7 270K Plus averages 2940 on the same test.
<a href="https://browser.geekbench.com/processors/intel-core-ultra-7-270k-plus" rel="nofollow">https://browser.geekbench.com/processors/intel-core-ultra-7-...
revolvingthrow · · focus · HN ↗
According to geekbench 7800x3d is 2400 single core, 15500 multicore. M4 pro is 3350 single core and 24750 multicore. Yet when I convert video using libsvtav1 with ffmpeg I’m getting noticeably faster performance on the desktop. And that’s with mbp, which doesn’t thermally throttle within 15 seconds.
Is it the 96mb cache? Avx-512? Are benchmarks bullshit when comparing different architectures?
kakacik · · focus · HN ↗
Bingo, each CPU is too unique with its own strengths and weaknesses to make broad statements, marketing picks up what they like and ignore rest
galad87 · · focus · HN ↗
snek_case · · focus · HN ↗
GeekyBear · · focus · HN ↗
Apple chips don't have AVX-512.
They do have media engines (that don't support AV1 encode), so if you switch to H.265, it will pull way ahead.
ajross · · focus · HN ↗
Zen 4 isn't 512 bits wide though, it's a split cycle 256 bit wide SIMD engine otherwise very similar (except in register size) to M6's.
The answer is more that Geekbench is at this point[1] heavily tuned to exactly the code Apple silicon does well: implicitly parallel wide-issue scalar code that you typically get out of modern compilers and JIT engines when throwing mostly-unoptimized "regular source code" at them. Apple has an enormous amount of instruction issue parallelism compared with x86.
The grandparent is looking at transcoding tasks where the limit isn't instruction issue but actual compute hardware on the core. And Apple doesn't actually win by much there.
It's just hard to know what to measure. Geekbench tends to be a metric for "feels fast doing boring interactive user stuff", which probably matches Apple's marketing imperatives well.
[1] Really they keep moving harder in that direction with every release. The "cooling pauses" in v6 likewise seemed very much like an attempt to boost the score on fanless Apple devices. If one were the type to allege a dark conspiracy, this is a tempting spot.
thejazzman · · focus · HN ↗
MacBook Air/neo are the only ones without fans. Even the Studio Display has a fan
ajross · · focus · HN ↗
Now... you can make a reasonable case that this matches real world interactive load better. But it also happens to have the practical effect of making Apple's numbers better, and no one else's. The new benchmarks are measuring subtly different things, and the new thing they measure happens to be what Apple wants to sell. At best, that's backwards.
krunkcoin · · focus · HN ↗
GB's author, John Poole, has stated the intent of Geekbench CPU is to measure the CPU and CPU alone, as in he doesn't want to measure limits caused by the form factor of the device the CPU under test is embedded in. Since GB runs on phones, that means he had to implement the cooldown pauses.
Intel's turbo boost means this GB feature almost certainly inflates lots of x86 scores too, even with fans involved. Not sure why you're being so conspiratorial about it.
[deleted] · · focus · HN ↗
[deleted]
luxuryballs · · focus · HN ↗
BoingBoomTschak · · focus · HN ↗
From what I've been able to gather, Apple only supports SME and the small "streaming" part of SVE2 required by SME since the M4. Which seems to be mostly useless for video encoding (I only see NEON/SVE in <a href="https://gitlab.com/AOMediaCodec/SVT-AV1/-/tree/master/Source/Lib" rel="nofollow">https://gitlab.com/AOMediaCodec/SVT-AV1/-/tree/master/Source... or <a href="https://github.com/Multicorewareinc/x265/tree/master/source/common" rel="nofollow">https://github.com/Multicorewareinc/x265/tree/master/source/...) it's basically a GEMM engine.
So basically, video encoding is a bad show for Apple who seem to be saying "use crappy hardware encoding and buy amd64 if you need more".
Very hard to find benchmark data. Found <a href="https://openbenchmarking.org/vs/Processor/Apple+M4+Pro,AMD+Ryzen+9+9900X+12-Core" rel="nofollow">https://openbenchmarking.org/vs/Processor/Apple+M4+Pro,AMD+R... that shows a good x265 performance at 1080p but not 4K. Guess the wider SIMD units matter more there.
GeekyBear · · focus · HN ↗
Apple has dedicated hardware for video decode and encode for several video formats.
That's why the comparisons between PCs and Macs running video editing software favor the Macs so heavily.
BoingBoomTschak · · focus · HN ↗
GeekyBear · · focus · HN ↗
Video editing is an area that Apple dominates because PCs become a laggy mess on complicated high resolution edits.
AMD has started to copy the same strategy with the recent Ryzen AI chips. They include hardware video encode/decode units as well as unified memory.
BoingBoomTschak · · focus · HN ↗
Uh yes, it's called reality. Advanced video coding techniques are just too complex, full of serial and branch heavy algos to work in decently sized and priced ASIC or even in GPGPU. Which is the reason none exists outside of very expensive special stuff for professional render/streaming farms.
Not even mentioning that few coding tools are shared between codecs, so while you only need 1 CPU to handle all of them, you'd need almost an ASIC per codec...
GeekyBear · · focus · HN ↗
I'll wait.
I would recommend you find a floor plan for Apple's M series chips and take a look at how big Apple's media engine is.
BoingBoomTschak · · focus · HN ↗
Anyway, here, I'm feeling nice. Some guy comparing the M1 Pro's HEVC encoder vs x265 (amongst others): <a href="https://colinmckellar.com/2024/01/11/video-encoder-comparison/" rel="nofollow">https://colinmckellar.com/2024/01/11/video-encoder-compariso...
Here's the only graph you need to look at if you don't want to bother (encoding time vs file size at fixed perceptual quality, log scale axes): <a href="https://colinmckellar.com/wp-content/uploads/2024/01/VMAF_90.png" rel="nofollow">https://colinmckellar.com/wp-content/uploads/2024/01/VMAF_90...
GeekyBear · · focus · HN ↗
We've got five generations of Apple's Media Engine shipping in M series chips.
If things are as dire as you claim, one of the many hardware reviews in reputable publications over the years would have mentioned this unacceptable quality at some point.
timschmidt · · focus · HN ↗
It seems like you are interpreting this as a slight against the quality of Apple's hardware encoders, which may legitimately be very good. As are Nvidia's, Intel's and AMD's. But all of them will produce larger file sizes and lower quality than equivalently optimized non-realtime software encoders, which simply have more information and more time, memory, and flexibility to compute over it.
We're talking about fundamental properties of compression and computational time/space trade-offs. Even Apple can't design around them.
That doesn't mean Apple's hardware encoder is in any way bad or unusable. All lossy compression will be imperfect, yet much of it is useful. And most modern codecs and encoders seem to be capable of high quality results. The implications of the differences under discussion are percentages of a bitrate or tiny nearly imperceptible artifacts or breadth of available resolutions, refresh rates, and color modes or codec choice. Software encoders are always at the bleeding edge of what's possible. Hardware encoders are necessarily a snapshot frozen in silicon with limitations imposed by the implementation. The middle ground is largely already occupied by SIMD and other transform-specific ISA extensions already present in most CPUs.
api · · focus · HN ↗
AVX still wins though.
M series still wins on performance per watt and now apparently leads on general purpose code.
All these leading edge chips are very good. We have an embarrassment of riches when it comes to blistering fast chips here.
Archit3ch · · focus · HN ↗
On what, microbenchmarks? I have realtime audio workloads where the hot loop is essentially linear algebra. Exactly the kind of work that suits AVX2/AVX512. Guess what, Apple Silicon still pulls ahead because real workloads are branchy, cache-hungry, full of dependencies and do not line up in 8 neat f64 operations per cycle.
tom_ · · focus · HN ↗
Geekbench 7 results:
* AMD 2990WX: 1384 (single), 13052 (multi) (<a href="https://browser.geekbench.com/v7/cpu/181239" rel="nofollow">https://browser.geekbench.com/v7/cpu/181239)
* Apple M4 Max: 3552 (single), 29863 (multi) (<a href="https://browser.geekbench.com/v7/cpu/390256" rel="nofollow">https://browser.geekbench.com/v7/cpu/390256)
For parallelisable stuff that can occupy all cores for an extended period, the 2990WX typically takes about ~1.2x as long to do the same work/does ~0.83x the work per unit time, assuming code compiled with clang or gcc. Which isn't really coming across in the numbers here.
CPUMark is a bit better:
* AMD 2990WX: 2282 (single), 32040 (multi) (<a href="https://www.cpubenchmark.net/cpu.php?cpu=AMD+Ryzen+Threadripper+2990WX&id=3309" rel="nofollow">https://www.cpubenchmark.net/cpu.php?cpu=AMD+Ryzen+Threadrip...)
* Apple M4 Max: 4590 (single), 43911 (multi) (<a href="https://www.cpubenchmark.net/cpu.php?cpu=Apple+M4+Max+16+Core&id=6348" rel="nofollow">https://www.cpubenchmark.net/cpu.php?cpu=Apple+M4+Max+16+Cor...)
I haven't spent much time timing single core stuff, except - regarding clang, which looks like it contributes to the Geekbench 7 score, I did some measurements a few months ago suggesting that clang compiles for x64 more slowly than for ARM, all else being as equal as I could be bothered to try to make it: <a href="https://news.ycombinator.com/item?id=46938682">https://news.ycombinator.com/item?id=46938682 - and the single threaded test runs I did of my code suggest that the Geekbench 7 single core might be about right?
(Whether the clang timing discrepancy is actually relevant to Geekbench, I've no idea, but I thought it interesting anyway.)
If you need a benchmark that makes the PC look massively faster than the Mac, I'm sure those are available too.
bob1029 · · focus · HN ↗
Go check out blender CPU scores if you need any reassurance that AMD still makes some kind of sense.
<a href="https://opendata.blender.org/benchmarks/query/?compute_type=CPU&blender_version=3.1.0&group_by=device_name" rel="nofollow">https://opendata.blender.org/benchmarks/query/?compute_type=...
jltsiren · · focus · HN ↗
wtallis · · focus · HN ↗
ksec · · focus · HN ↗
Another point is that the multicore part uses all core including E-Core. On AMD the multicore are all the same. Meaning for some benchmarks this will flavour Apple more.
Again there is nothing that stop people from optimising it for ARM Mac. The problem is the usage of it is so small it probably doesn't make sense to focus on it. SVT took a really long time for it to reach quality parity with AOM's AV1 encoder and later exceed it.
TiredOfLife · · focus · HN ↗