‹ BackHN Continuity

Thread

Btrfs/ZFS/bcachefs under workloads classic benchmarks skip

179 points · 190 comments · farlight

  1. fenio · · focus · HN ↗
    The author of the benchmark here. I went over some comments and I'll try to tackle them here. I'm pretty clear that GH runner based benchmark is far from perfect due to noisy neighbours etc. Thus every test first is running so called calibration... to reject completely unreliable VMs. I'm fully aware that this can't completely fix the issue. Can limit it but not fix. But as of now there are 593 runs recorded so average should still be quite meaningful.

    Having that said I'm desperately trying to get REAL hardware to run that benchmark. With some successes ;)

    Few months ago I got Hetzner machine from Kent Overstreet and I was able to finish 3 runs before machine died... Results: <a href="https:&#x2F;&#x2F;bartosz.fenski.pl&#x2F;modern-fs-benchmark&#x2F;real-hw&#x2F;" rel="nofollow">https:&#x2F;&#x2F;bartosz.fenski.pl&#x2F;modern-fs-benchmark&#x2F;real-hw&#x2F;

    Currently I&#x27;ve got even more interesting machine with tons of disks and I&#x27;m running new set of benchmarks but it&#x27;s really in its initial stage.

    <a href="https:&#x2F;&#x2F;bartosz.fenski.pl&#x2F;modern-fs-benchmark&#x2F;sas-hdd&#x2F;" rel="nofollow">https:&#x2F;&#x2F;bartosz.fenski.pl&#x2F;modern-fs-benchmark&#x2F;sas-hdd&#x2F; 2nd run in progress... one run on REAL hardware takes much more time than on GH runner so it&#x27;s slow.

    But this new hardware has also so many disks that the plan is to try also more complex, tiered cache topologies. I&#x27;m working on it.

    I&#x27;m happy to answer any other questions, sources of every piece of this benchmark are freely available and I&#x27;m not saying they are 100% correct. I&#x27;m open to improvements.

    1. Grayskull · · focus · HN ↗
      Love to hear that there will be more real hw tests. At the moment I am building NAS and used your benchmark for evaluating the filesystems. I am glad to see that your data roughly matches mine (apart from scrub which on 4x 6tb HDDs took 15 hours for md-raid10 while CoW systems took seconds). Personally I found that array of HDDs behaves very differently than GH runner (my feeling is that since it runs on same disk you are testing theoretical throughput rather than ability to utilize disks). My tests gave an idea for following topologies:

        * 4 HDDs (for example dm-raid has read balancing optimized specifically for HDDs)
        * 5 HDDs (classical raid should see no improvement but btrfs and bcachefs should balance the load)
        * 4 SSDs
        * 3 HDDs + 1 SSD no tiering
        * 2 HDDs + 2 SSD no tiering
        * 1 drive 10x larger than others (since how bcachefs and btrfs allocators work)
        * nocow
      Thanks for awesome work
      1. ciupicri · · focus · HN ↗
        How can a scrub take only seconds when it has to read all the data?
        1. Polizeiposaune · · focus · HN ↗
          zfs-style scrub only reads allocated blocks and skips unallocated blocks; if the pool is mostly empty it can complete very quickly.

          Layered storage systems with a RAID layer that makes N disks look like one big disk generally don&#x27;t have visibility into which blocks are free and which are allocated so they must &quot;scrub&quot; all the disks on initialization and repair even if only 1% is used.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.