> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sfcompute.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Object storage

> Bring datasets to SF Compute and save training checkpoints between runs

Keep datasets and checkpoints in SF Compute object storage, close to your compute. Upload data from
your laptop or another storage provider, then use it on SFC instances. Uploaded data stays after an
instance terminates.

## Create a bucket and access key

Creating buckets or access keys requires a billing profile, primary billing contact, and card set as
the default payment method. Have an organization admin complete this setup under
[Billing](https://console.sfcompute.com/dashboard/billing) in the console.

The examples use `europe-north1` and your active [workspace](/preview/workspaces).

Replace `acme-training-data` throughout with your own bucket name.

<Tabs>
  <Tab title="CLI">
    [Install `sf` and log in](/preview/quick-start#install-the-cli), then create a bucket and key in
    the same region.

    ```bash theme={null}
    sf storage buckets create --name acme-training-data --region europe-north1
    sf storage access-key create --name training --region europe-north1
    ```

    Save the returned access key ID and secret in your secrets manager. The secret is returned
    only once.
  </Tab>

  <Tab title="Console">
    Open **Storage** → **Buckets** → **Create Bucket**. Enter a name and region, and choose a
    workspace if offered. If **Also create an access key** appears, select it and click
    **Create Bucket & Key**. Otherwise, click **Create Bucket** and reuse your existing key.

    Save a new key's **Access key ID** and **Secret access key** in your secrets manager; the secret
    is shown only once. If you need a new key, create one under **Access keys**.
  </Tab>
</Tabs>

You may reuse an existing access key if it's in the same workspace and region as the bucket.

**Access scope:** An access key can read and write every bucket in its workspace and region. Use
separate workspaces when you need separate access boundaries.

On each machine below, set these variables using your saved key and your bucket's region. Find the
endpoint in **Access bucket** or with `sf storage buckets get acme-training-data`.
`sf storage access-key create --format env` and the console's **Access bucket** snippets print
`AWS_*` names instead. Rename them as below so they don't override your existing AWS credentials.

```bash theme={null}
export SFC_ACCESS_KEY_ID="your-access-key-id"
export SFC_SECRET_ACCESS_KEY="your-secret-access-key"
export SFC_ENDPOINT_URL="https://object.europe-north1.sfcompute.com"
export SFC_REGION="europe-north1"
```

## Restrictions

| Resource | Restriction |
| - | - |
| Bucket capacity | 100 TiB per bucket by default. [Contact us](https://sfcompute.com/contact) if you need a higher limit. |
| Access keys | Two per workspace per region. Key names must be unique within a workspace. |
| Bucket names | Globally unique; 3–63 lowercase letters, numbers, or hyphens, starting with a letter or number. |

## Usage and billing

SF Compute measures the bytes stored in each bucket hourly. Storage costs **\$0.14 per GiB-month**,
prorated by the amount of data stored and how long you keep it. For example, 100 GiB stored for a
full calendar month costs **\$14**; for half that month, **\$7**.

Usage is invoiced after each UTC calendar month. Your default card is charged when the invoice is
due, seven days after it is sent. Storage charges are separate from prepaid compute credits.

Track usage and month-to-date cost under **Usage** → **Storage** in the console. View invoices under
**Billing** → **Invoices**.

## Bring your dataset to SFC

Run these commands on your laptop or the machine doing the transfer. Install
[rclone](https://rclone.org/install/), then configure an SFC remote using the credentials loaded
above.

```bash theme={null}
rclone config create sfc s3 \
  provider=Other \
  access_key_id="$SFC_ACCESS_KEY_ID" secret_access_key="$SFC_SECRET_ACCESS_KEY" \
  endpoint="$SFC_ENDPOINT_URL" region="$SFC_REGION" \
  force_path_style=true
```

Choose your source:

<Tabs>
  <Tab title="From your laptop">
    Copy your dataset into a versioned directory in the bucket.

    ```bash theme={null}
    rclone copy ./dataset sfc:acme-training-data/datasets/v1 --progress
    ```
  </Tab>

  <Tab title="From another S3 provider">
    Use `rclone config` to add a remote named `source` with that provider's endpoint and its own
    credentials. Then copy between the two remotes.

    ```bash theme={null}
    rclone copy source:source-bucket/dataset sfc:acme-training-data/datasets/v1 --progress
    ```

    The transfer runs through the machine running rclone.
  </Tab>
</Tabs>

Rerun `copy` to transfer new or changed files. Use a new directory such as `datasets/v2` for a new
dataset version. To mirror the source exactly, use `rclone sync`; it also deletes destination files
absent from the source.

### Use the dataset on an instance

[Connect to your instance](/preview/instances#ssh-into-an-instance), set the variables above, and
configure rclone as above. Copy the dataset onto its local disk before starting training.

```bash theme={null}
mkdir -p ~/datasets/v1
rclone copy sfc:acme-training-data/datasets/v1 ~/datasets/v1 --progress
```

Point your training data loader at `~/datasets/v1`. Repeat on each instance that needs a local copy.

## Save and resume training checkpoints

From an SFC instance, save checkpoints to object storage in the same region as your training job.
This keeps them available when you move to another instance or start another run.

With PyTorch already installed, run `python -m pip install boto3` in your training environment. Set
the same variables there. Add this to your training loop, where `model`, `optimizer`, and `step` are
your current training state:

```python theme={null}
import os
from pathlib import Path

import boto3
import torch
from botocore.config import Config

s3 = boto3.client(
    "s3",
    aws_access_key_id=os.environ["SFC_ACCESS_KEY_ID"],
    aws_secret_access_key=os.environ["SFC_SECRET_ACCESS_KEY"],
    endpoint_url=os.environ["SFC_ENDPOINT_URL"],
    region_name=os.environ["SFC_REGION"],
    config=Config(s3={"addressing_style": "path"}),
)

checkpoint = Path("checkpoints") / f"step-{step}.pt"
checkpoint.parent.mkdir(parents=True, exist_ok=True)
torch.save({
    "step": step,
    "model": model.state_dict(),
    "optimizer": optimizer.state_dict(),
}, checkpoint)
s3.upload_file(str(checkpoint), "acme-training-data", f"runs/run-42/step-{step}.pt")
```

Use a different run directory for each job. Upload only completed checkpoints; for distributed
training, coordinate the writers so they don't overwrite each other. Wait for the upload to succeed
before terminating the instance.

On a fresh instance, recreate the model and optimizer, configure the S3 client as above, and restore
a completed checkpoint:

```python theme={null}
s3.download_file("acme-training-data", "runs/run-42/step-1000.pt", "checkpoint.pt")
state = torch.load("checkpoint.pt", map_location="cpu", weights_only=True)
model.load_state_dict(state["model"])
optimizer.load_state_dict(state["optimizer"])
start_step = state["step"] + 1
```

Include any other state your job needs to resume, such as its learning-rate scheduler. See
[PyTorch checkpointing](https://docs.pytorch.org/tutorials/beginner/saving_loading_models.html#saving-loading-a-general-checkpoint-for-inference-and-or-resuming-training).

To retrieve a run on your own machine, use the same rclone remote:

```bash theme={null}
rclone copy sfc:acme-training-data/runs/run-42 ./run-42 --progress
```

## Manage storage

List buckets and keys with `sf storage buckets list` and `sf storage access-key list`, or open
**Storage** in the console. Revoke a key with `sf storage access-key delete training` or **Revoke
Key**.

When you no longer need a bucket, run `sf storage buckets delete acme-training-data` or select
**Delete Bucket** in the console. Confirm the name when prompted. Deletion removes the bucket and
all its objects.

<Warning>
  Deleting a bucket is permanent. Its name becomes available for anyone to claim.
</Warning>

## API reference

The SF Compute API doesn't manage buckets or access keys yet. Use `sf storage` to script them, or
[contact us](https://sfcompute.com/contact) if you need programmatic access.
