Skip to content

EC2 Instance which Support EFA with RDMA Read/Write , NVSWITCH and GPUDirect

0

EC2 Instance which Support EFA with RDMA Read/Write , NVSWITCH and GPUDirect

asked a year ago627 views

2 Answers
1
Accepted Answer

Hi Alok,

The EC2 instance that supports EFA with RDMA Read/Write, NVSWITCH, and GPUDirect is the p4d.24xlarge instance.

  • The p4d.24xlarge instance is specifically designed for high-performance computing and machine learning workloads.
  • EFA, NVSwitch, and GPUDirect are crucial for minimizing latency and maximizing bandwidth in such workloads.
  • Understanding the purpose of each technology helps in identifying the appropriate EC2 instance.

Step 1: Identify the relevant technologies. EFA with RDMA, NVSwitch, and GPUDirect are key technologies for high-performance computing and machine learning workloads.

Step 2: Determine the EC2 instance type that supports these technologies. The p4d.24xlarge instance is designed to support all three technologies: EFA with RDMA, NVSwitch, and GPUDirect.

Step 3: Explain why the p4d.24xlarge instance is the correct choice. EFA with RDMA: The p4d.24xlarge instance supports EFA, enabling low-latency, high-bandwidth communication between instances using RDMA. NVSwitch: The p4d.24xlarge instance utilizes NVSwitch to provide fast communication between its eight NVIDIA A100 GPUs. GPUDirect: The p4d.24xlarge instance supports GPUDirect, enabling efficient data transfer between the GPUs and EFA.


If the Answer is helpful, please click Accept Answer and up-vote, so that it can help others in the community looking for help on similar topics.

answered a year ago

EXPERT

reviewed a year ago

1

Based on the available information, the Amazon EC2 P5en instances support EFA with RDMA capabilities, NVSwitch, and GPUDirect RDMA.

The P5en.48xlarge instances feature:

  • 8 NVIDIA H200 Tensor Core GPUs with 1128 GB of high bandwidth GPU memory
  • EFAv3 networking with up to 3,200 Gbps of aggregate network bandwidth
  • GPUDirect RDMA support, which enables lower latency and efficient scale-out performance by bypassing the CPU for internode communication
  • 900 GB/s NVSwitch for GPU peer-to-peer communication
  • 192 vCPUs and 2 TiB (2048 GiB) of system memory
  • 8 x 3.84 TB NVMe SSD instance storage
  • 100 Gbps EBS bandwidth

Other instances in the P5 family that support similar capabilities include P5e.48xlarge and P5.48xlarge, though they use EFAv2 networking rather than the newer EFAv3 found in P5en instances.

The P6-B200 and P6e-GB200 instances also support EFAv4 networking with GPUDirect RDMA capabilities, with the P6e-GB200 UltraServers delivering up to 28.8 terabits per second of total EFAv4 networking.

These instances are designed for high-performance computing (HPC) and machine learning workloads, particularly generative AI applications, and are part of Amazon EC2 UltraClusters.
Sources
New Amazon EC2 P5en instances with NVIDIA H200 Tensor Core GPUs and EFAv3 networking | AWS News Blog
Cloud Compute Instances – Amazon EC2 Instance Types – AWS
Amazon EC2 UltraClusters Instances – AWS

answered a year ago

EXPERT

reviewed a year ago

You are not logged in. Log in to post an answer.

A good answer clearly answers the question and provides constructive feedback and encourages professional growth in the question asker.