‹ BackHN Continuity

Thread

Making Tailscale Faster

250 points · 112 comments · yarapavan

  1. iscoelho · · focus · HN ↗
    In my opinion, this is Tailscale's largest issue.

    It is slow. It cannot achieve speeds of greater than 1Gbps on clients systems (Windows & Mac), where you'd normally see it being used. On Linux, it struggles to achieve 10Gbps even when using a synthetic large packet benchmark [1]. With an IMIX benchmark, it would not be competitive whatsoever.

    This problem is fixable. WireGuard achieves higher performance (Kernel vs Userspace implementation) and IPsec implementations can achieve 100Gbps/400Gbps (DPDK/XDP). Zero-copy networking.

    From this blog post, I can say Tailscale still seems to not have the appetite for that, which is a shame.

    [1] <a href="https:&#x2F;&#x2F;tailscale.com&#x2F;blog&#x2F;more-throughput" rel="nofollow">https:&#x2F;&#x2F;tailscale.com&#x2F;blog&#x2F;more-throughput

    1. boomer_joe · · focus · HN ↗
      Yes. Just fucking stop doing userspace wireguard on linux <a href="https:&#x2F;&#x2F;github.com&#x2F;tailscale&#x2F;tailscale&#x2F;issues&#x2F;426" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;tailscale&#x2F;tailscale&#x2F;issues&#x2F;426 - issue has been open for 6 years (and is locked now, lol), btw.

      And if any tailscale employees are reading this - <a href="https:&#x2F;&#x2F;github.com&#x2F;tailscale&#x2F;tailscale&#x2F;issues&#x2F;15724" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;tailscale&#x2F;tailscale&#x2F;issues&#x2F;15724 please fix this too. Regular users not using some sort of enterprise saas DNS (whatever their thing is?) deserve DNS privacy too.

      1. lokar · · focus · HN ↗
        Kernel networking is not automatically faster then userspace.
        1. iscoelho · · focus · HN ↗
          You&#x27;re correct, kernel isn&#x27;t faster by default. With that said, the following is true:

          1) the WireGuard kernel implementation, despite not even being zero-copy, exceeds the performance of the userspace implementation

          2) implementations utilizing the userspace network stack have a maximum potential performance (context switch + memcpy is very slow, and that affects UDP disproportionately). It&#x27;s the wrong approach for meaningful improvement.

          1. vlovich123 · · focus · HN ↗
            Io_uring doesn’t have context switching and may not have memcpy. The trickier thing I suspect to get with wireguard is the encryption and GSO offload
            1. iscoelho · · focus · HN ↗
              That is not true unfortunately. io_uring only avoids the userspace copy&#x2F;switch. There are many other copies in the Linux userspace network stack.
              1. vlovich123 · · focus · HN ↗
                What do you mean by “userspace network stack” when we’re talking about io_uring? That’s a contradiction. Unless you mean extra memcpy’s within the kernel network stack, but when talking about the buffer you supply I don’t believe that’s true - your data generally gets directly DMA’ed into the device because the buffer you supplied is pinned and can’t be released until the second CQE is delivered. It would be helpful if you clarified.
          2. Veserv · · focus · HN ↗
            I do not understand why people continue parroting this performance nonsense.

            Memory copying is on the order of 100 gigabytes per second. You can do 80 full payload copys and still out-pace a dog-slow 10 gigabit per second connection. If your bottleneck is memory copying you are either doing something very wrong and doing way too many copys or congratulations you have implemented one of the fastest network stacks.

            Supervisor calls are also very fast, on the order of 100 ns up to maybe 1 us with all the Spectre mitigations. Even if you did something as stupid as one supervisor call per packet, you would still be getting on the order of 10 Gbps at the long end there and 100 Gbps at the short end. Which, again, means congratulations are in order because you have implemented one of the fastest network stacks. If you do batching and add just 10 us (us, not ms) of latency then that entire cost is so small as to be irrelevant. You are going to bottleneck on your memory copying first.

            Network stacks are so slow almost entirely due to poor protocol design and poor protocol implementation. Usually both.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.