Skip to main content
This feature is in public preview.
InfiniBand is a high-bandwidth, low-latency network that lets GPUs on different instances exchange data during multi-node training. It uses RDMA (remote direct memory access). With RDMA, one machine reads and writes another machine’s memory directly, including GPU memory, without involving either operating system. An InfiniBand partition is an isolated network for RDMA traffic between your instances. Instances in the same partition can reach each other over InfiniBand. An instance created without a partition has no InfiniBand connectivity, even when its SKU supports InfiniBand. A fabric is a physical InfiniBand network in one availability zone, with a name such as europe-north1-a.fab2. Separate fabrics are separate networks and cannot reach each other. To connect instances over InfiniBand:
  1. Create a partition.
  2. Attach instances on the same fabric to that partition.
sf infiniband-partitions is also available as sf ib-partitions and sf ibp.

Enroll in the InfiniBand preview

Enroll your account in the preview before you create a partition. Until you do, the sf infiniband-partitions commands and the --infiniband-partition flag are hidden, and the API returns a 403 error.

Find a SKU that supports InfiniBand

A SKU supports InfiniBand when its RDMA type is infiniband. List the SKUs that support it, with their availability.

Create a partition

Create a partition on the fabric your SKU is on. The partition belongs to the current workspace.
You can also run sf infiniband-partitions create with no flags. It asks for a name, then lists each fabric with the names of the SKUs on it. Pick the fabric of the SKU you plan to create instances on. In a zone with only one fabric, you can pass --region and --zone instead of --fabric. A partition name must be unique within its workspace, across all fabrics.

List partitions

List the partitions in the current workspace. Add --all to list across every workspace you can read.
Filter by --fabric to see the partitions on one network.

Get details for one partition

Get a partition by its name, its ibpart_... ID, or its resource path. A name resolves within the current workspace.
A resource path joins the account, the workspace, and the partition name, as in acme, production, and training above. Anywhere a command or API field takes a partition, it accepts the name, the ID, or the resource path.

Attach an instance to a partition

Attach an instance by passing --infiniband-partition when you create it. The partition must be on the same fabric as the instance’s SKU. See Instances for the other creation flags and the startup script.
You can also run sf instances create without --infiniband-partition. On an InfiniBand SKU, it lists the partitions on the SKU’s fabric, so you can pick one, create one, or choose none. An instance stays in its partition for its whole life. To move it to another partition, terminate it and create a new instance.

List the instances in a partition

List the instances attached to a partition. The output has the same columns as sf instances list.
A terminated instance drops out of this list once its node releases the partition.

Delete a partition

Delete a partition by its name, ID, or resource path. Terminate every instance in the partition first, since deletion fails while any instance is still attached.
Two kinds of instance can block a deletion without showing up in sf infiniband-partitions instances:
  • An instance you don’t have permission to read
  • A terminated instance whose node hasn’t released the partition yet

Check InfiniBand from an instance

The current Ubuntu images with CUDA 13.1 or 13.2, such as ubuntu-24.04-cuda-13.2, ship the InfiniBand stack. See Images for the full list. The stack includes:
  • Drivers: NVIDIA DOCA-OFED, with IP over InfiniBand
  • Tools: ibverbs-utils and perftest, for checking a link
  • GPUDirect RDMA: nvidia-peermem, loaded at boot with the InfiniBand modules
An instance booted from one of these images needs no setup inside the guest. Other base images, and older builds of these images, don’t include the stack. Run ibv_devinfo on the instance to confirm it sees an InfiniBand device.
A device with state: PORT_ACTIVE is connected to the fabric.

API reference

See the InfiniBand Partitions API for programmatic access.

Next steps

To train across instances in a partition, follow Run a PyTorch job across multiple nodes.