torchrun job on 2 nodes with 8 GPUs each. The nodes exchange data over an isolated InfiniBand partition. By the end, you’ll have trained nanoGPT on a small corpus of Shakespeare and generated text from it.
Before you start
Set up your account, a pool, and a partition before you create any instances. You need thesf CLI, logged in to your account. See the Quickstart to install it.
1
Enroll in the InfiniBand preview
Enroll your account so the
sf infiniband-partitions commands and the --infiniband-partition flag become available.2
Find a SKU that supports InfiniBand
List the SKUs with InfiniBand and when each becomes available.Pick a SKU and note its
NAME and ID. The examples use the SKU with ID sku_4UpxzQw7A8N.3
Buy an allocation of 2 nodes
Create a pool named
training and buy an allocation of 2 nodes of your SKU for 4 hours. The sf buy command quotes the price and asks you to confirm before it charges your organization.4
Create a partition on your SKU's fabric
Create a partition named If you already know the fabric, pass it with
training. Without --fabric, the command lists each fabric followed by the names of the SKUs on it. Pick the fabric that lists your SKU’s name. When only 1 fabric exists, the command uses it without asking.--fabric, for example --fabric europe-north1-a.fab2. Only instances in the same partition can reach each other over InfiniBand. See Create a partition for the other options.5
Write a startup script
SF Compute public images ship with no SSH key. This script adds your public keys to the See Cloud-init for other ways to configure an instance.
root user when the instance boots. The $(cat ...) lines run on your machine as you write the file.Launch the job
Each instance runstorchrun, and the instance with rank 0 coordinates the job. That instance’s IP over InfiniBand (IPoIB) address is the only address the job needs.
1
Create one instance per node
Create 2 instances in the partition. This loop names them Wait until both instances show The image download and boot can take up to 10 minutes before SSH is available.
training-0 and training-1.running in this list.2
Find the InfiniBand address and interface
Open one terminal per instance and connect as On each instance, print the IPoIB interface and its address.Record 2 values:
root, the user the startup script added your key to. Start with training-0.- Interface: the first field,
ibs14here, which can differ on each instance - Rank 0 address: the last field on
training-0, without the/64, which every instance uses
3
Set up nanoGPT on every instance
This guide trains nanoGPT, a small GPT model, on 1 MB of Shakespeare. Install uv, clone nanoGPT, install its packages, and prepare the dataset. Run these commands on each instance.
prepare.py downloads the text and writes train.bin and val.bin to data/shakespeare_char/. Each instance reads its own copy.4
Run torchrun on every instance
Start The
torchrun on each instance with its own node rank: 0 on training-0 and 1 on training-1. Replace ibs14 with the instance’s interface and the address with your rank 0 address.NCCL_SOCKET_IFNAME sets the interface that the NVIDIA Collective Communications Library (NCCL) uses to connect the instances at startup. The IPoIB interface is the one other instances in the partition can reach.The flags after the config file change nanoGPT’s single-GPU defaults for 16 GPUs:--gradient_accumulation_steps=16: nanoGPT splits this value across the GPUs, so it must be a multiple of the GPU count--max_iters,--lr_decay_iters,--warmup_iters: shorten the run, since each step now covers 16 GPUs’ worth of data--eval_interval,--eval_iters: evaluate every 100 steps on 20 batches, so evaluation fits the short run
--nnodes instances have joined, then training starts. Each GPU prints the NCCL network line, and training-0 prints the loss every 100 steps.NET/IB line lists 1 InfiniBand device per GPU, so NCCL sends the job’s traffic over InfiniBand. On 2 nodes with 8 H100 GPUs each, the run takes about 45 seconds. If the job exits or hangs with no output, run the command again without the final grep to see the full log.5
Generate text from the model
training-0 saves the checkpoint to out-shakespeare-char/. Generate 300 characters from it on training-0.Clean up
Terminate the instances when the job ends, then delete the partition. Run these commands on your machine, not on an instance.sf infiniband-partitions instances training no longer lists the instances.
Next steps
Use these pages to manage the job’s resources:- InfiniBand partitions: fabrics, partitions, and link checks
- Instances: instance states, priority, and termination
- Pools: the allocation for the nodes in your job