> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sfcompute.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Run a PyTorch job across multiple nodes

> Create an InfiniBand partition, start instances in it, and launch torchrun on every node

This guide takes you from no instances to a `torchrun` job on 2 nodes with 8 GPUs each. The nodes exchange data over an isolated [InfiniBand partition](/preview/networking/infiniband-partitions). By the end, you'll have trained [nanoGPT](https://github.com/karpathy/nanoGPT) on a small corpus of Shakespeare and generated text from it.

## Before you start

Set up your account, a pool, and a partition before you create any instances. You need the `sf` CLI, logged in to your account. See the [Quickstart](/preview/quick-start#install-the-cli) to install it.

<Steps>
  <Step title="Enroll in the InfiniBand preview">
    Enroll your account so the `sf infiniband-partitions` commands and the `--infiniband-partition` flag become available.

    ```bash theme={null}
    sf preview set infiniband-preview=true
    ```
  </Step>

  <Step title="Find a SKU that supports InfiniBand">
    List the [SKUs](/preview/instance-skus) with InfiniBand and when each becomes available.

    ```bash theme={null}
    sf availability --rdma-type infiniband
    ```

    Pick a SKU and note its `NAME` and `ID`. The examples use the SKU with ID `sku_4UpxzQw7A8N`.
  </Step>

  <Step title="Buy an allocation of 2 nodes">
    Create a [pool](/preview/pools) named `training` and buy an allocation of 2 nodes of your SKU for 4 hours. The `sf buy` command quotes the price and asks you to confirm before it charges your organization.

    ```bash theme={null}
    sf pools create --name training
    sf buy --pool training --sku sku_4UpxzQw7A8N --count 2 --start now --duration 4h
    ```
  </Step>

  <Step title="Create a partition on your SKU's fabric">
    Create a partition named `training`. Without `--fabric`, the command lists each fabric followed by the names of the SKUs on it. Pick the fabric that lists your SKU's name. When only 1 fabric exists, the command uses it without asking.

    ```bash theme={null}
    sf infiniband-partitions create --name training
    ```

    If you already know the fabric, pass it with `--fabric`, for example `--fabric europe-north1-a.fab2`. Only instances in the same partition can reach each other over InfiniBand. See [Create a partition](/preview/networking/infiniband-partitions#create-a-partition) for the other options.
  </Step>

  <Step title="Write a startup script">
    SF Compute public images ship with no SSH key. This script adds your public keys to the `root` user when the instance boots. The `$(cat ...)` lines run on your machine as you write the file.

    ```bash theme={null}
    cat >startup.sh <<SCRIPT
    #!/bin/bash

    mkdir -p /root/.ssh
    cat >>/root/.ssh/authorized_keys <<"EOF"
    $(cat ~/.ssh/id_rsa.pub 2>/dev/null)
    $(cat ~/.ssh/id_ecdsa.pub 2>/dev/null)
    $(cat ~/.ssh/id_ed25519.pub 2>/dev/null)
    EOF
    SCRIPT
    ```

    See [Cloud-init](/preview/instances#cloud-init) for other ways to configure an instance.
  </Step>
</Steps>

## Launch the job

Each instance runs `torchrun`, and the instance with rank 0 coordinates the job. That instance's IP over InfiniBand (IPoIB) address is the only address the job needs.

<Steps>
  <Step title="Create one instance per node">
    Create 2 instances in the partition. This loop names them `training-0` and `training-1`.

    ```bash theme={null}
    for i in 0 1; do
      sf instances create --name "training-$i" --pool training \
        --sku sku_4UpxzQw7A8N \
        --image sfc:image:sfcompute:public:ubuntu-24.04-cuda-13.2 \
        --cloud-init ./startup.sh --infiniband-partition training --yes
    done
    ```

    Wait until both instances show `running` in this list.

    ```bash theme={null}
    sf infiniband-partitions instances training
    ```

    The image download and boot can take up to 10 minutes before SSH is available.
  </Step>

  <Step title="Find the InfiniBand address and interface">
    Open one terminal per instance and connect as `root`, the user the startup script added your key to. Start with `training-0`.

    ```bash theme={null}
    sf instances ssh training-0 --login root
    ```

    On each instance, print the IPoIB interface and its address.

    ```bash theme={null}
    ip -br -6 addr show scope global | tr -s ' '
    ```

    ```text theme={null}
    ibs14 UP fdf9:c316:93ee:0:5f4:15ed:ad34:ab08/64
    ```

    Record 2 values:

    * **Interface**: the first field, `ibs14` here, which can differ on each instance
    * **Rank 0 address**: the last field on `training-0`, without the `/64`, which every instance uses
  </Step>

  <Step title="Set up nanoGPT on every instance">
    This guide trains [nanoGPT](https://github.com/karpathy/nanoGPT), a small GPT model, on 1 MB of Shakespeare. Install [uv](https://docs.astral.sh/uv/), clone nanoGPT, install its packages, and prepare the dataset. Run these commands on each instance.

    ```bash theme={null}
    curl -LsSf https://astral.sh/uv/install.sh | sh
    source ~/.local/bin/env
    git clone https://github.com/karpathy/nanoGPT ~/nanoGPT
    uv venv ~/venv
    uv pip install --python ~/venv torch numpy tiktoken requests
    cd ~/nanoGPT && ~/venv/bin/python data/shakespeare_char/prepare.py
    ```

    `prepare.py` downloads the text and writes `train.bin` and `val.bin` to `data/shakespeare_char/`. Each instance reads its own copy.
  </Step>

  <Step title="Run torchrun on every instance">
    Start `torchrun` on each instance with its own node rank: `0` on `training-0` and `1` on `training-1`. Replace `ibs14` with the instance's interface and the address with your rank 0 address.

    ```bash theme={null}
    cd ~/nanoGPT
    NCCL_SOCKET_IFNAME=ibs14 NCCL_DEBUG=INFO ~/venv/bin/torchrun \
      --nnodes 2 --nproc-per-node 8 --node-rank 0 \
      --master-addr fdf9:c316:93ee:0:5f4:15ed:ad34:ab08 --master-port 29500 \
      train.py config/train_shakespeare_char.py \
      --gradient_accumulation_steps=16 --max_iters=500 --lr_decay_iters=500 \
      --warmup_iters=50 --eval_interval=100 --eval_iters=20 --compile=False \
      2>&1 | grep -E "NET/IB|step "
    ```

    `NCCL_SOCKET_IFNAME` sets the interface that the NVIDIA Collective Communications Library (NCCL) uses to connect the instances at startup. The IPoIB interface is the one other instances in the partition can reach.

    The flags after the config file change nanoGPT's single-GPU defaults for 16 GPUs:

    * **`--gradient_accumulation_steps=16`**: nanoGPT splits this value across the GPUs, so it must be a multiple of the GPU count
    * **`--max_iters`, `--lr_decay_iters`, `--warmup_iters`**: shorten the run, since each step now covers 16 GPUs' worth of data
    * **`--eval_interval`, `--eval_iters`**: evaluate every 100 steps on 20 batches, so evaluation fits the short run

    Each instance waits until all `--nnodes` instances have joined, then training starts. Each GPU prints the NCCL network line, and `training-0` prints the loss every 100 steps.

    ```text theme={null}
    NCCL INFO NET/IB : Using [0]mlx5_0:1/IB [1]mlx5_1:1/IB … [7]mlx5_7:1/IB
    …
    step 0: train loss 4.2885, val loss 4.2831
    step 100: train loss 2.3211, val loss 2.3593
    …
    step 500: train loss 1.2833, val loss 1.5373
    ```

    The `NET/IB` line lists 1 InfiniBand device per GPU, so NCCL sends the job's traffic over InfiniBand. On 2 nodes with 8 H100 GPUs each, the run takes about 45 seconds. If the job exits or hangs with no output, run the command again without the final `grep` to see the full log.
  </Step>

  <Step title="Generate text from the model">
    `training-0` saves the checkpoint to `out-shakespeare-char/`. Generate 300 characters from it on `training-0`.

    ```bash theme={null}
    cd ~/nanoGPT
    ~/venv/bin/python sample.py --out_dir=out-shakespeare-char \
      --num_samples=1 --max_new_tokens=300
    ```

    ```text theme={null}
    GLOUCESTER:
    And will it there by the execution of your swell:
    There is the mark of your grieves hew you love.
    ```
  </Step>
</Steps>

## Clean up

Terminate the instances when the job ends, then delete the partition. Run these commands on your machine, not on an instance.

```bash theme={null}
for i in 0 1; do sf instances terminate "training-$i" --yes; done
```

Delete the partition once `sf infiniband-partitions instances training` no longer lists the instances.

```bash theme={null}
sf infiniband-partitions delete training
```

Deletion fails while an instance is still attached. See [Delete a partition](/preview/networking/infiniband-partitions#delete-a-partition) for the cases that block it.

Terminating an instance doesn't free its allocation: the pool keeps the 2 nodes until the order's 4 hours end. To get credits back for unused time, place a [sell order](/preview/orders#sell-orders) on the pool.

## Next steps

Use these pages to manage the job's resources:

* [InfiniBand partitions](/preview/networking/infiniband-partitions): fabrics, partitions, and link checks
* [Instances](/preview/instances): instance states, priority, and termination
* [Pools](/preview/pools): the allocation for the nodes in your job
