Prerequisites
Feature Description
The network stack has delays and a small frame size. If you apply RDMA, you can achieve the speed of hundreds of backends, as if running on a single server
Motivation
It would be good to be able to synchronize the execution results in layers between backends via RDMA to reduce delays
Possible Implementation
If you do not have support or hardware with RDMA, you can use the RXE kernel module for emulation
https://enterprise-support.nvidia.com/s/article/howto-configure-soft-roce
Prerequisites
Feature Description
The network stack has delays and a small frame size. If you apply RDMA, you can achieve the speed of hundreds of backends, as if running on a single server
Motivation
It would be good to be able to synchronize the execution results in layers between backends via RDMA to reduce delays
Possible Implementation
If you do not have support or hardware with RDMA, you can use the RXE kernel module for emulation
https://enterprise-support.nvidia.com/s/article/howto-configure-soft-roce