Skip to main content
Keep datasets and checkpoints in SF Compute object storage, close to your compute. Upload data from your laptop or another storage provider, then use it on SFC instances. Uploaded data stays after an instance terminates.

Create a bucket and access key

Creating buckets or access keys requires a billing profile, primary billing contact, and card set as the default payment method. Have an organization admin complete this setup under Billing in the console. The examples use europe-north1 and your active workspace. Replace acme-training-data throughout with your own bucket name.
Install sf and log in, then create a bucket and key in the same region.
Save the returned access key ID and secret in your secrets manager. The secret is returned only once.
You may reuse an existing access key if it’s in the same workspace and region as the bucket. Access scope: An access key can read and write every bucket in its workspace and region. Use separate workspaces when you need separate access boundaries. On each machine below, set these variables using your saved key and your bucket’s region. Find the endpoint in Access bucket or with sf storage buckets get acme-training-data. sf storage access-key create --format env and the console’s Access bucket snippets print AWS_* names instead. Rename them as below so they don’t override your existing AWS credentials.

Restrictions

Usage and billing

SF Compute measures the bytes stored in each bucket hourly. Storage costs $0.14 per GiB-month, prorated by the amount of data stored and how long you keep it. For example, 100 GiB stored for a full calendar month costs $14; for half that month, $7. Usage is invoiced after each UTC calendar month. Your default card is charged when the invoice is due, seven days after it is sent. Storage charges are separate from prepaid compute credits. Track usage and month-to-date cost under Usage → Storage in the console. View invoices under Billing → Invoices.

Bring your dataset to SFC

Run these commands on your laptop or the machine doing the transfer. Install rclone, then configure an SFC remote using the credentials loaded above.
Choose your source:
Copy your dataset into a versioned directory in the bucket.
Rerun copy to transfer new or changed files. Use a new directory such as datasets/v2 for a new dataset version. To mirror the source exactly, use rclone sync; it also deletes destination files absent from the source.

Use the dataset on an instance

Connect to your instance, set the variables above, and configure rclone as above. Copy the dataset onto its local disk before starting training.
Point your training data loader at ~/datasets/v1. Repeat on each instance that needs a local copy.

Save and resume training checkpoints

From an SFC instance, save checkpoints to object storage in the same region as your training job. This keeps them available when you move to another instance or start another run. With PyTorch already installed, run python -m pip install boto3 in your training environment. Set the same variables there. Add this to your training loop, where model, optimizer, and step are your current training state:
Use a different run directory for each job. Upload only completed checkpoints; for distributed training, coordinate the writers so they don’t overwrite each other. Wait for the upload to succeed before terminating the instance. On a fresh instance, recreate the model and optimizer, configure the S3 client as above, and restore a completed checkpoint:
Include any other state your job needs to resume, such as its learning-rate scheduler. See PyTorch checkpointing. To retrieve a run on your own machine, use the same rclone remote:

Manage storage

List buckets and keys with sf storage buckets list and sf storage access-key list, or open Storage in the console. Revoke a key with sf storage access-key delete training or Revoke Key. When you no longer need a bucket, run sf storage buckets delete acme-training-data or select Delete Bucket in the console. Confirm the name when prompted. Deletion removes the bucket and all its objects.
Deleting a bucket is permanent. Its name becomes available for anyone to claim.

API reference

The SF Compute API doesn’t manage buckets or access keys yet. Use sf storage to script them, or contact us if you need programmatic access.