Skip to content

A dataset on a drive, shared by runs#

You create a drive food101 that holds the Food-101 dataset (about 5 GB) from Hugging Face. You never copy it by hand: the first time a run that mounts the drive lands on a machine, Astraeus starts a fill run there that downloads the dataset into the machine's copy, and the run waits until the copy is whole. Every later run on that machine mounts the copy read-only and starts at once, and Astraeus prefers machines that already hold a copy.

What you need:

  • A workspace with access to a cluster, with machines that have a data location chosen (Machines → machine → Data location → Confirm).
  • Outbound HTTPS from the machines to huggingface.co and its CDN.
  • The editor role in the workspace, an API token, curl and jq (How the recipes are written).

How it fits together#

sequenceDiagram
  participant R as Run "inspect-food101"
  participant S as Astralyx control plane (SaaS)
  participant M as Machine gpu-01
  participant F as Fill run "fill-food101-gpu-01"
  R->>S: mounts drive food101
  S->>M: places the worker (room in the data location)
  M-->>R: Preparing: Waiting for drive food101 to be filled on gpu-01
  S->>F: one fill run, pinned to gpu-01
  F->>M: downloads into the copy, writes .astralyx-filled last
  M-->>R: copy complete: the worker starts, /data read-only

A copy is not usable until the fill writes the marker file .astralyx-filled as its last step, so no run ever reads a half-downloaded dataset. One fill run exists per drive and machine; its logs are where a failed download says why.

1. Create the drive#

drive.json
{
  "metadata": {"name": "food101", "labels": {"dataset": "food101"}},
  "spec": {
    "sources": [
      {"node_scope": "placed", "path": "", "mount_path": "/data", "mode": "ReadOnly"}
    ],
    "size_hint_bytes": 6000000000,
    "evictable": true,
    "fill": {
      "template": {
        "image": "python:3.12-slim",
        "command": "sh",
        "args": [
          "-c",
          "set -e; pip install --no-cache-dir -q 'huggingface_hub>=0.24'; huggingface-cli download ethz/food101 --repo-type dataset --local-dir /drive/food101; touch /drive/.astralyx-filled"
        ],
        "env": {"HF_HOME": "/tmp/hf"},
        "time_limit_seconds": 7200,
        "requested_resources": {"cpu_cores": 2, "memory_bytes": 4294967296}
      }
    }
  }
}
Field Why
node_scope: placed A drive kept on each machine: its bytes live in each machine's data location, under drives/<namespace>.food101. Nothing to grant, nothing to mount by hand.
mode: ReadOnly Runs mount it read-only unless they ask otherwise. The fill run always gets it read-write, at /drive.
size_hint_bytes A machine without a copy is chosen only if its data location has 6 GB free.
evictable Copies no worker is using may be removed when a machine's disk fills, least recently used first, and are filled again when needed. Right for a dataset you can download again; wrong for results.
fill.template The fill run's worker: an image, a command and its resources, like any run. It must write /drive/.astralyx-filled last. It may not mount other drives.
  1. Open Drives → New drive, choose the cluster, and press Edit as JSON.
  2. Paste drive.json and press Create drive.

The form alone covers everything but the fill: Kind On each machine's data location, Mounted in workers at, Access, Expected size and A cache.

$ curl -fsS -X POST "$API/drives" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d @drive.json | jq -r .metadata.name
food101

2. Use it from a run#

The first run only looks at the data. It mounts the drive at /data, read-only.

inspect.json
{
  "metadata": {"name": "inspect-food101"},
  "spec": {
    "task_template": {
      "image": "busybox:1.36",
      "command": "sh",
      "args": ["-c", "du -sh /data/food101 && ls /data/food101/data && touch /data/should-fail || echo 'read-only, as asked'"],
      "restart_policy": "Never",
      "requested_resources": {"cpu_cores": 1, "memory_bytes": 268435456, "node_selection": {"mode": "Any"}},
      "datavolume_refs": [{"name": "food101", "mount_path": "/data", "mode": "ReadOnly"}]
    }
  }
}

Runs → New run → Edit as JSON, paste inspect.json, Start run. In the form, the same mount is Drives: food101:/data.

$ curl -fsS -X POST "$API/runs" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d @inspect.json | jq -r .metadata.name
inspect-food101

3. Watch the fill#

The fill run appears among the workspace's runs, named after the drive and the machine:

$ astra astraeus runs
NAME                          CLUSTER       STATE       REASON
fill-food101-gpu-01           main          Running     1 of 1 workers running
inspect-food101               main          Pending     Waiting for workers to start
$ astra astraeus logs fill-food101-gpu-01-0 | tail -n 3
Fetching 7 files: 100%|██████████| 7/7 [01:42<00:00, 14.6s/it]
/drive/food101

Meanwhile the worker inspect-food101-0 is Preparing with Waiting for drive <namespace>.food101 to be filled on gpu-01 (the drive's full name, with your workspace's namespace). When the fill completes, the machine reports the copy whole and the worker starts:

$ astra astraeus logs inspect-food101-0
4.7G    /data/food101
train-00000-of-00008.parquet
train-00001-of-00008.parquet
…
validation-00000-of-00003.parquet
…
touch: /data/should-fail: Read-only file system
read-only, as asked

4. Check that placement follows the data#

  1. Open the drive's page. Copies lists gpu-01, its folder (<data location>/drives/<namespace>.food101), its size and when it was last used. Used by lists the workers that mount it.

    $ curl -fsS "$API/drives/food101" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
        | jq '.copies[] | {node, bytes, complete}'
    {
      "node": "gpu-01",
      "bytes": 5046112837,
      "complete": true
    }
    
  2. Submit the same run again under another name. It goes to gpu-01, which holds a copy, and starts without a fill:

    $ jq '.metadata.name = "inspect-food101-2"' inspect.json \
        | curl -fsS -X POST "$API/runs" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
            -H 'content-type: application/json' -d @- | jq -r .metadata.name
    inspect-food101-2
    $ astra astraeus runs --all
    NAME                          CLUSTER       STATE       REASON
    inspect-food101-2             main          Completed   All workers completed
    inspect-food101               main          Completed   All workers completed
    fill-food101-gpu-01           main          Completed   All workers completed
    

    The preference for a machine with a copy is a score, not a rule: when gpu-01 has no room, the run goes to another machine and that machine gets its own fill.

5. Mount it in your training runs#

Any run mounts the drive the same way. Many runs can mount it at once, on the same machine or on several:

fragment of a training run
"datavolume_refs": [
  {"name": "food101", "mount_path": "/data", "mode": "ReadOnly"},
  {"name": "cifar-runs", "mount_path": "/out", "mode": "ReadWrite"}
]

With sub_path a run mounts one folder of the drive: {"name": "food101", "sub_path": "food101/data", "mount_path": "/parquet"}.

6. Clean up#

Delete the runs from their pages, then delete the drive from its page.

$ astra astraeus delete inspect-food101
inspect-food101 deleted
$ astra astraeus delete inspect-food101-2
inspect-food101-2 deleted
$ curl -fsS -X DELETE "$API/runs/inspect-food101" -H "Authorization: Bearer $ASTRAEUS_TOKEN"
$ curl -fsS -X DELETE "$API/runs/inspect-food101-2" -H "Authorization: Bearer $ASTRAEUS_TOKEN"
$ curl -fsS -X DELETE "$API/drives/food101" -H "Authorization: Bearer $ASTRAEUS_TOKEN"

Deleting the drive deletes its copy on every machine and its fill runs. A drive cannot be deleted while a live worker uses it.

Variations#

One copy, reached over the network. If the dataset is too large to keep on every machine, keep it on one machine and let runs elsewhere read it over NFS, over RDMA when both machines have it:

drive-on-one-machine.json
{
  "metadata": {"name": "food101-gpu01"},
  "spec": {
    "sources": [
      {"node": "gpu-01", "path": "/datasets/food101", "mount_path": "/data", "mode": "ReadWrite"}
    ],
    "transport": "Auto"
  }
}

The path must be under a host path your workspace was granted (Settings → Clusters → Change → Host paths, by an organisation admin). Load it once with a run pinned to that machine ("node_selection": {"mode": "Exact", "names": ["gpu-01"]}) that mounts it read-write, then mount it read-only everywhere. "transport": "Local" instead makes every run that mounts it go to gpu-01. Astraeus never deletes the files of such a drive.

A shared filesystem. If the dataset already sits on Lustre, GPFS or NFS mounted at the same path on every machine, describe it with "node_scope": "shared": each worker reads it directly.

From object storage. A fill can read from S3 or any store your machines can reach. Give the fill's template a credential with secret_refs and use the store's own tool (aws s3 sync s3://… /drive/…, then touch /drive/.astralyx-filled). See Use cloud credentials without storing them.

Fill again. After a fill completes, if the copy is still not whole ten minutes later (it was evicted, or the fill wrote no marker), the fill runs again. A fill that failed stays failed so you can read why; delete the fill run to have it made again.

Troubleshooting#

Symptom Cause Fix
The run waits with Choose where to keep data on gpu-01 The machine has no data location. Confirm one on the machine's page.
The worker never leaves Waiting for drive … food101 to be filled on gpu-01, and the fill run completed The fill did not write /drive/.astralyx-filled. Make the marker the last command, after a set -e so a failed download stops before it.
The fill run failed The download failed: read its logs (astra astraeus logs fill-food101-gpu-01-0). Fix the template (a drive cannot be edited: delete it and create it again), or delete the fill run to retry.
No machine is chosen; the worker's page says drive food101 needs … at this machine's data location, … free size_hint_bytes is more than the machines' data locations have free (or than their limit allows). Free space, raise the machine's data location limit, or lower the hint.