Instances
An instance is a GPU machine that runs a VM image. Like instances on other clouds, it has hardware, an operating system image, startup configuration, logs, SSH access, and lifecycle state. You create it, connect to it, run your workload, and terminate it when you are done. The only difference is that an SF Compute instance consumes compute time when it runs. It won’t start until it has compute time allocated to it. That is what pools are for.Pools
A pool is a named container that holds an allocation of compute time. It is represented as a schedule: how many nodes you own, on which hardware, and for which windows of time. For example, a pool might have this schedule:
The pool’s schedule changes as compute is bought, sold, transferred, or expires.
Instance SKUs
Each pool tracks a separate schedule for each instance SKU. An instance SKU represents a specific set of interchangeable hardware, including:- Accelerator type, such as H100 or H200
- Node specs like number of CPUs, RAM, and local storage
- Interconnect type
- Zone or location
- Other operator-defined hardware properties
Starting an instance
When you create an instance, you attach it to a pool. The instance starts inawaiting_allocation
if the pool does not currently have available compute for the hardware it will run on.
To add compute, place a buy order on the pool. When the order fills, it adds to
the pool’s allocation schedule for the order’s instance SKU. Once the schedule covers the current
time and there is enough matching allocation, the instance starts booting.
Instance lifecycle
A pool’s allocation determines whether attached instances can run.
Terminating an instance does not free its allocation. The allocation remains in the pool’s
schedule until it is sold, transferred, or expires.
When allocation grows, waiting instances can start. If you use a
deployment, the deployment can also create more instances to match its
target count, and those instances start when allocation is available.
When allocation shrinks below the number of running instances, SF Compute terminates instances
until the running count fits the pool schedule. Lower-priority instances are terminated first, and
among instances with the same priority, newer ones go first. If you need control over which work
survives a planned shrink, set instance priority before selling,
transferring, or letting allocation expire.
Through this lifecycle, compute allocation is decoupled from what you run. You can buy compute
before you launch instances, recreate instances without losing reserved time, sell unused time
without deleting the pool, and move allocation between workloads.
You can think about it as analogous to a power grid. Pools are the circuits. Orders add compute to
those circuits, transfers move compute between circuits, and instances draw compute from the pool
they are attached to. If a circuit has enough scheduled compute, the attached instances can run.
If it does not, they wait or terminate.
One pool or many
You can attach multiple instances to a single pool, or create a separate pool for each instance or workload. Use a separate pool when the instance should be managed independently.- A personal dev box should usually have its own pool.
- A one-off debugging machine should usually have its own pool.
- A customer-isolated workload should use its own pool if selling or expiration should not affect other customers.
- An inference fleet should usually use one pool and one deployment.
- Batch workers for the same queue can usually share one pool.
- A training job with many identical workers can usually share one pool.
Transfers
Compute time can be transferred between pools. This lets you manage allocation like any other resource: assign it to projects, teams, or workloads, then move it when it is not being used. For example, a platform team might buy a large block of H100 time into a shared pool, then transfer pieces of that allocation to separate pools fortraining, inference, and
experiments. If experiments does not need its allocation, the team can transfer it back or
move it to another project.
Transfers affect the same schedule that orders affect. If compute leaves a pool, attached
instances may terminate if the remaining allocation no longer covers them. If compute arrives,
waiting instances can start.