With regards to (~6:01) "congestion control is the responsibility of the sender" and "somehow we have to get the sending nodes to stop sending so fast". Does that not exist in Ethernet/RoCE?
> Link Level Flow Control: InfiniBand uses a credit-based algorithm to guarantee lossless HCA-to-HCA communication. RoCE runs on top of Ethernet. Implementations may require lossless Ethernet network for reaching to performance characteristics similar to InfiniBand. Lossless Ethernet is typically configured via Ethernet flow control or priority flow control (PFC). Configuring a Data center bridging (DCB) Ethernet network can be more complex than configuring an InfiniBand network.[19]
> A sending station (computer or network switch) may be transmitting data faster than the other end of the link can accept it. Using flow control, the receiving station can signal the sender requesting suspension of transmissions until the receiver catches up. Flow control on Ethernet can be implemented at the data link layer.
Flow control is better than nothing but it can cause congestion spreading and bufferbloat. QCN, Falcon, and Ultra Ethernet provide much better congestion control for RoCE but they also require newer hardware compared to Homa.
Ethernet flow control doesn't really fix congestion except in very special cases.
Consider a very simple topology:
A C
\ /
S1===S2
/ \
B D
Say hosts A and B are both sending data to C, as fast as they can, via switches S1 and S2 (which are connected via a high-speed link). And say the sum of these two flows is more than the capacity of the link to C.
S2 is receiving packets destined for C faster than it can forward them, but sending an Ethernet pause frame from S2 to S1 is not a very productive way to alleviate the situation, because it also disrupts any traffic that would be bound for D. It just moves the bottleneck elsewhere and causes collateral damage.
Maybe, but doesn’t RoCE / RDMA handle this at a higher level as well? It’s why PFC and ECN are required for these deployments, which effectively ensure everyone behaves nicely.
I agree that Ethernet flow control is insufficient, but given that NVidia in an act of brilliant foresight acquired Mellanox a decade ago, I’m fairly certain that this is how all these AI clusters are actually deployed, not using TCP, and maybe not even using Ethernet but infiniband instead.
if you spend the zillion dollars on switches and racks and power and cables and redundant power in your datacenter for infiniband(the only vendor: Nvidia), you get IB.
throw0101c · · focus · HN ↗
> Link Level Flow Control: InfiniBand uses a credit-based algorithm to guarantee lossless HCA-to-HCA communication. RoCE runs on top of Ethernet. Implementations may require lossless Ethernet network for reaching to performance characteristics similar to InfiniBand. Lossless Ethernet is typically configured via Ethernet flow control or priority flow control (PFC). Configuring a Data center bridging (DCB) Ethernet network can be more complex than configuring an InfiniBand network.[19]
* <a href="https://en.wikipedia.org/wiki/RDMA_over_Converged_Ethernet" rel="nofollow">https://en.wikipedia.org/wiki/RDMA_over_Converged_Ethernet
> A sending station (computer or network switch) may be transmitting data faster than the other end of the link can accept it. Using flow control, the receiving station can signal the sender requesting suspension of transmissions until the receiver catches up. Flow control on Ethernet can be implemented at the data link layer.
* <a href="https://en.wikipedia.org/wiki/Ethernet_flow_control" rel="nofollow">https://en.wikipedia.org/wiki/Ethernet_flow_control
lokar · · focus · HN ↗
When a link becomes saturated working out how to manage that is a hard problem.
I have only used RoCE once at scale, it was really finicky. We would get big waves to pause frames that stalled everything.
wmf · · focus · HN ↗
teraflop · · focus · HN ↗
Consider a very simple topology:
Say hosts A and B are both sending data to C, as fast as they can, via switches S1 and S2 (which are connected via a high-speed link). And say the sum of these two flows is more than the capacity of the link to C.S2 is receiving packets destined for C faster than it can forward them, but sending an Ethernet pause frame from S2 to S1 is not a very productive way to alleviate the situation, because it also disrupts any traffic that would be bound for D. It just moves the bottleneck elsewhere and causes collateral damage.
486sx33 · · focus · HN ↗
[dead]
stingraycharles · · focus · HN ↗
I agree that Ethernet flow control is insufficient, but given that NVidia in an act of brilliant foresight acquired Mellanox a decade ago, I’m fairly certain that this is how all these AI clusters are actually deployed, not using TCP, and maybe not even using Ethernet but infiniband instead.
etc-hosts · · focus · HN ↗
if not, you get RoCE over ethernet.