> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sfcompute.com/llms.txt
> Use this file to discover all available pages before exploring further.

# InfiniBand partitions

> Connect instances over an isolated InfiniBand network for distributed training

<div className="preview-notice">
  <Info>
    This feature is in [public preview](/preview/roadmap#feature-states).
  </Info>
</div>

InfiniBand is a high-bandwidth, low-latency network that lets GPUs on different instances exchange
data during multi-node training. It uses RDMA (remote direct memory access). With RDMA, one machine
reads and writes another machine's memory directly, including GPU memory, without involving either
operating system.

An InfiniBand partition is an isolated network for RDMA traffic between your instances. Instances in
the same partition can reach each other over InfiniBand. An instance created without a partition has
no InfiniBand connectivity, even when its SKU supports InfiniBand.

A fabric is a physical InfiniBand network in one availability zone, with a name such as
`europe-north1-a.fab2`. Separate fabrics are separate networks and cannot reach each other.

To connect instances over InfiniBand:

1. Create a partition.
2. Attach instances on the same fabric to that partition.

`sf infiniband-partitions` is also available as `sf ib-partitions` and `sf ibp`.

## Enroll in the InfiniBand preview

Enroll your account in the preview before you create a partition. Until you do, the
`sf infiniband-partitions` commands and the `--infiniband-partition` flag are hidden, and the API
returns a `403` error.

```bash theme={null}
sf preview set infiniband-preview=true
```

## Find a SKU that supports InfiniBand

A [SKU](/preview/instance-skus) supports InfiniBand when its RDMA type is `infiniband`. List the
SKUs that support it, with their availability.

```bash theme={null}
sf availability --rdma-type infiniband
```

## Create a partition

Create a partition on the fabric your SKU is on. The partition belongs to the current
[workspace](/preview/workspaces).

```bash theme={null}
sf infiniband-partitions create --name training --fabric europe-north1-a.fab2
```

You can also run `sf infiniband-partitions create` with no flags. It asks for a name, then lists
each fabric with the names of the SKUs on it. Pick the fabric of the SKU you plan to create
instances on.

In a zone with only one fabric, you can pass `--region` and `--zone` instead of `--fabric`.

A partition name must be unique within its workspace, across all fabrics.

## List partitions

List the partitions in the current workspace. Add `--all` to list across every workspace you can
read.

```bash theme={null}
sf infiniband-partitions list
```

```text theme={null}
  NAME      FABRIC                ZONE             CREATED
  training  europe-north1-a.fab2  europe-north1-a  Sep 28, 2:02pm
```

Filter by `--fabric` to see the partitions on one network.

```bash theme={null}
sf infiniband-partitions list --fabric europe-north1-a.fab2
```

## Get details for one partition

Get a partition by its name, its `ibpart_...` ID, or its resource path. A name resolves within the
current workspace.

```bash theme={null}
sf infiniband-partitions get training
```

```bash theme={null}
sf infiniband-partitions get sfc:infiniband_partition:acme:production:training
```

A resource path joins the account, the workspace, and the partition name, as in `acme`,
`production`, and `training` above. Anywhere a command or API field takes a partition, it accepts
the name, the ID, or the resource path.

## Attach an instance to a partition

Attach an instance by passing `--infiniband-partition` when you create it. The partition must be on
the same fabric as the instance's SKU. See [Instances](/preview/instances) for the other creation
flags and the [startup script](/preview/instances#cloud-init).

```bash theme={null}
sf instances create --pool training --sku sku_4UpxzQw7A8N \
  --image sfc:image:sfcompute:public:ubuntu-24.04-cuda-13.2 --cloud-init ./startup.sh \
  --infiniband-partition training
```

You can also run `sf instances create` without `--infiniband-partition`. On an InfiniBand SKU, it
lists the partitions on the SKU's fabric, so you can pick one, create one, or choose none.

An instance stays in its partition for its whole life. To move it to another partition, terminate it
and create a new instance.

## List the instances in a partition

List the instances attached to a partition. The output has the same columns as
[`sf instances list`](/preview/instances).

```bash theme={null}
sf infiniband-partitions instances training
```

A terminated instance drops out of this list once its node releases the partition.

## Delete a partition

Delete a partition by its name, ID, or resource path. Terminate every instance in the partition
first, since deletion fails while any instance is still attached.

```bash theme={null}
sf infiniband-partitions delete training
```

Two kinds of instance can block a deletion without showing up in
`sf infiniband-partitions instances`:

* An instance you don't have permission to read
* A terminated instance whose node hasn't released the partition yet

## Check InfiniBand from an instance

The current Ubuntu images with CUDA 13.1 or 13.2, such as `ubuntu-24.04-cuda-13.2`, ship the
InfiniBand stack. See [Images](/preview/images) for the full list. The stack includes:

* **Drivers**: NVIDIA DOCA-OFED, with IP over InfiniBand
* **Tools**: `ibverbs-utils` and `perftest`, for checking a link
* **GPUDirect RDMA**: `nvidia-peermem`, loaded at boot with the InfiniBand modules

An instance booted from one of these images needs no setup inside the guest. Other base images, and
older builds of these images, don't include the stack.

Run `ibv_devinfo` on the instance to confirm it sees an InfiniBand device.

```bash theme={null}
ibv_devinfo
```

A device with `state: PORT_ACTIVE` is connected to the fabric.

## API reference

See the
[InfiniBand Partitions API](/preview/api-reference/infiniband-partitions/list-infiniband-partitions)
for programmatic access.

## Next steps

To train across instances in a partition, follow
[Run a PyTorch job across multiple nodes](/preview/guides/multi-node-training).
