Skip to main content
This guide takes you from no instances to a torchrun job on 2 nodes with 8 GPUs each. The nodes exchange data over an isolated InfiniBand partition. By the end, you’ll have trained nanoGPT on a small corpus of Shakespeare and generated text from it.

Before you start

Set up your account, a pool, and a partition before you create any instances. You need the sf CLI, logged in to your account. See the Quickstart to install it.
1

Enroll in the InfiniBand preview

Enroll your account so the sf infiniband-partitions commands and the --infiniband-partition flag become available.
2

Find a SKU that supports InfiniBand

List the SKUs with InfiniBand and when each becomes available.
Pick a SKU and note its NAME and ID. The examples use the SKU with ID sku_4UpxzQw7A8N.
3

Buy an allocation of 2 nodes

Create a pool named training and buy an allocation of 2 nodes of your SKU for 4 hours. The sf buy command quotes the price and asks you to confirm before it charges your organization.
4

Create a partition on your SKU's fabric

Create a partition named training. Without --fabric, the command lists each fabric followed by the names of the SKUs on it. Pick the fabric that lists your SKU’s name. When only 1 fabric exists, the command uses it without asking.
If you already know the fabric, pass it with --fabric, for example --fabric europe-north1-a.fab2. Only instances in the same partition can reach each other over InfiniBand. See Create a partition for the other options.
5

Write a startup script

SF Compute public images ship with no SSH key. This script adds your public keys to the root user when the instance boots. The $(cat ...) lines run on your machine as you write the file.
See Cloud-init for other ways to configure an instance.

Launch the job

Each instance runs torchrun, and the instance with rank 0 coordinates the job. That instance’s IP over InfiniBand (IPoIB) address is the only address the job needs.
1

Create one instance per node

Create 2 instances in the partition. This loop names them training-0 and training-1.
Wait until both instances show running in this list.
The image download and boot can take up to 10 minutes before SSH is available.
2

Find the InfiniBand address and interface

Open one terminal per instance and connect as root, the user the startup script added your key to. Start with training-0.
On each instance, print the IPoIB interface and its address.
Record 2 values:
  • Interface: the first field, ibs14 here, which can differ on each instance
  • Rank 0 address: the last field on training-0, without the /64, which every instance uses
3

Set up nanoGPT on every instance

This guide trains nanoGPT, a small GPT model, on 1 MB of Shakespeare. Install uv, clone nanoGPT, install its packages, and prepare the dataset. Run these commands on each instance.
prepare.py downloads the text and writes train.bin and val.bin to data/shakespeare_char/. Each instance reads its own copy.
4

Run torchrun on every instance

Start torchrun on each instance with its own node rank: 0 on training-0 and 1 on training-1. Replace ibs14 with the instance’s interface and the address with your rank 0 address.
NCCL_SOCKET_IFNAME sets the interface that the NVIDIA Collective Communications Library (NCCL) uses to connect the instances at startup. The IPoIB interface is the one other instances in the partition can reach.The flags after the config file change nanoGPT’s single-GPU defaults for 16 GPUs:
  • --gradient_accumulation_steps=16: nanoGPT splits this value across the GPUs, so it must be a multiple of the GPU count
  • --max_iters, --lr_decay_iters, --warmup_iters: shorten the run, since each step now covers 16 GPUs’ worth of data
  • --eval_interval, --eval_iters: evaluate every 100 steps on 20 batches, so evaluation fits the short run
Each instance waits until all --nnodes instances have joined, then training starts. Each GPU prints the NCCL network line, and training-0 prints the loss every 100 steps.
The NET/IB line lists 1 InfiniBand device per GPU, so NCCL sends the job’s traffic over InfiniBand. On 2 nodes with 8 H100 GPUs each, the run takes about 45 seconds. If the job exits or hangs with no output, run the command again without the final grep to see the full log.
5

Generate text from the model

training-0 saves the checkpoint to out-shakespeare-char/. Generate 300 characters from it on training-0.

Clean up

Terminate the instances when the job ends, then delete the partition. Run these commands on your machine, not on an instance.
Delete the partition once sf infiniband-partitions instances training no longer lists the instances.
Deletion fails while an instance is still attached. See Delete a partition for the cases that block it. Terminating an instance doesn’t free its allocation: the pool keeps the 2 nodes until the order’s 4 hours end. To get credits back for unused time, place a sell order on the pool.

Next steps

Use these pages to manage the job’s resources: