(Tailscale cofounder) I see a few comments here that using kernel wireguard would make it faster; it’s not really that simple. In fact, for a while (and we wrote a blog post about it), our optimizations made wireguard-go faster than kernel wireguard because it was better optimized. They adopted some of those improvements and now we’re on to the next order of magnitude together.
For really high bandwidth cases, things like DPDK are the long term best choice and are primarily userspace, for good reasons. Kernel mode is not the pure benefit it once was (if it ever was).
Separately, wireguard itself has a problem that the crypto suite it uses is not supported by hardware accelerators. So if we want to get into the hundreds of gigabits range, we will possibly need to switch packet formats entirely. (But, wireguard also needs to update to support post-quantum so maybe they’ll fix both problems at the same time and we can join in.)
> wireguard itself has a problem that the crypto suite it uses is not supported by hardware accelerators. So if we want to get into the hundreds of gigabits range, we will possibly need to switch packet formats entirely
Has hardware-offload (for AES et al) got faster still, or that keeping CPU busy in the data path for 100gbps workloads is not ideal, or something else? The WireGuard website claims ChaPoly is at least as fast as hardware-accelerated AES. And that it can be further sped up with SIMD.
> now we’re on to the next order of magnitude together
Curious: Is this work currently in progress? If so, what's more that's still lined up? The previous GRO/GSO(/LRO, too?) improvements were incredibly impressive (even to u/majke, <a href="https://news.ycombinator.com/item?id=35567268">https://news.ycombinator.com/item?id=35567268).
apenwarr · · focus · HN ↗
For really high bandwidth cases, things like DPDK are the long term best choice and are primarily userspace, for good reasons. Kernel mode is not the pure benefit it once was (if it ever was).
Separately, wireguard itself has a problem that the crypto suite it uses is not supported by hardware accelerators. So if we want to get into the hundreds of gigabits range, we will possibly need to switch packet formats entirely. (But, wireguard also needs to update to support post-quantum so maybe they’ll fix both problems at the same time and we can join in.)
ignoramous · · focus · HN ↗
Has hardware-offload (for AES et al) got faster still, or that keeping CPU busy in the data path for 100gbps workloads is not ideal, or something else? The WireGuard website claims ChaPoly is at least as fast as hardware-accelerated AES. And that it can be further sped up with SIMD.
<a href="https://www.wireguard.com/known-limitations" rel="nofollow">https://www.wireguard.com/known-limitations
> now we’re on to the next order of magnitude together
Curious: Is this work currently in progress? If so, what's more that's still lined up? The previous GRO/GSO(/LRO, too?) improvements were incredibly impressive (even to u/majke, <a href="https://news.ycombinator.com/item?id=35567268">https://news.ycombinator.com/item?id=35567268).
Thanks.
mxey · · focus · HN ↗
I haven’t used WireGuard but I have easily doubled OpenSSH performance by switching back to AES.