A dataset on a drive, shared by runs#
You create a drive food101 that holds the Food-101 dataset (about 5 GB)
from Hugging Face. You never copy it by hand: the first time a run that
mounts the drive lands on a machine, Astraeus starts a fill run there
that downloads the dataset into the machine's copy, and the run waits until
the copy is whole. Every later run on that machine mounts the copy
read-only and starts at once, and Astraeus prefers machines that
already hold a copy.
What you need:
- A workspace with access to a cluster, with machines that have a data location chosen (Machines → machine → Data location → Confirm).
- Outbound HTTPS from the machines to
huggingface.coand its CDN. - The editor role in the workspace, an API token,
curlandjq(How the recipes are written).
How it fits together#
sequenceDiagram
participant R as Run "inspect-food101"
participant S as Astralyx control plane (SaaS)
participant M as Machine gpu-01
participant F as Fill run "fill-food101-gpu-01"
R->>S: mounts drive food101
S->>M: places the worker (room in the data location)
M-->>R: Preparing: Waiting for drive food101 to be filled on gpu-01
S->>F: one fill run, pinned to gpu-01
F->>M: downloads into the copy, writes .astralyx-filled last
M-->>R: copy complete: the worker starts, /data read-only
A copy is not usable until the fill writes the marker file
.astralyx-filled as its last step, so no run ever reads a half-downloaded
dataset. One fill run exists per drive and machine; its logs are where a
failed download says why.
1. Create the drive#
{
"metadata": {"name": "food101", "labels": {"dataset": "food101"}},
"spec": {
"sources": [
{"node_scope": "placed", "path": "", "mount_path": "/data", "mode": "ReadOnly"}
],
"size_hint_bytes": 6000000000,
"evictable": true,
"fill": {
"template": {
"image": "python:3.12-slim",
"command": "sh",
"args": [
"-c",
"set -e; pip install --no-cache-dir -q 'huggingface_hub>=0.24'; huggingface-cli download ethz/food101 --repo-type dataset --local-dir /drive/food101; touch /drive/.astralyx-filled"
],
"env": {"HF_HOME": "/tmp/hf"},
"time_limit_seconds": 7200,
"requested_resources": {"cpu_cores": 2, "memory_bytes": 4294967296}
}
}
}
}
| Field | Why |
|---|---|
node_scope: placed |
A drive kept on each machine: its bytes live in each machine's data location, under drives/<namespace>.food101. Nothing to grant, nothing to mount by hand. |
mode: ReadOnly |
Runs mount it read-only unless they ask otherwise. The fill run always gets it read-write, at /drive. |
size_hint_bytes |
A machine without a copy is chosen only if its data location has 6 GB free. |
evictable |
Copies no worker is using may be removed when a machine's disk fills, least recently used first, and are filled again when needed. Right for a dataset you can download again; wrong for results. |
fill.template |
The fill run's worker: an image, a command and its resources, like any run. It must write /drive/.astralyx-filled last. It may not mount other drives. |
- Open Drives → New drive, choose the cluster, and press Edit as JSON.
- Paste
drive.jsonand press Create drive.
The form alone covers everything but the fill: Kind On each machine's data location, Mounted in workers at, Access, Expected size and A cache.
2. Use it from a run#
The first run only looks at the data. It mounts the drive at /data,
read-only.
{
"metadata": {"name": "inspect-food101"},
"spec": {
"task_template": {
"image": "busybox:1.36",
"command": "sh",
"args": ["-c", "du -sh /data/food101 && ls /data/food101/data && touch /data/should-fail || echo 'read-only, as asked'"],
"restart_policy": "Never",
"requested_resources": {"cpu_cores": 1, "memory_bytes": 268435456, "node_selection": {"mode": "Any"}},
"datavolume_refs": [{"name": "food101", "mount_path": "/data", "mode": "ReadOnly"}]
}
}
}
3. Watch the fill#
The fill run appears among the workspace's runs, named after the drive and the machine:
$ astra astraeus runs
NAME CLUSTER STATE REASON
fill-food101-gpu-01 main Running 1 of 1 workers running
inspect-food101 main Pending Waiting for workers to start
$ astra astraeus logs fill-food101-gpu-01-0 | tail -n 3
Fetching 7 files: 100%|██████████| 7/7 [01:42<00:00, 14.6s/it]
/drive/food101
Meanwhile the worker inspect-food101-0 is Preparing with Waiting for
drive <namespace>.food101 to be filled on gpu-01 (the drive's full name,
with your workspace's namespace). When the fill completes, the machine
reports the copy whole and the worker starts:
$ astra astraeus logs inspect-food101-0
4.7G /data/food101
train-00000-of-00008.parquet
train-00001-of-00008.parquet
…
validation-00000-of-00003.parquet
…
touch: /data/should-fail: Read-only file system
read-only, as asked
4. Check that placement follows the data#
-
Open the drive's page. Copies lists
gpu-01, its folder (<data location>/drives/<namespace>.food101), its size and when it was last used. Used by lists the workers that mount it. -
Submit the same run again under another name. It goes to
gpu-01, which holds a copy, and starts without a fill:$ jq '.metadata.name = "inspect-food101-2"' inspect.json \ | curl -fsS -X POST "$API/runs" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \ -H 'content-type: application/json' -d @- | jq -r .metadata.name inspect-food101-2 $ astra astraeus runs --all NAME CLUSTER STATE REASON inspect-food101-2 main Completed All workers completed inspect-food101 main Completed All workers completed fill-food101-gpu-01 main Completed All workers completedThe preference for a machine with a copy is a score, not a rule: when
gpu-01has no room, the run goes to another machine and that machine gets its own fill.
5. Mount it in your training runs#
Any run mounts the drive the same way. Many runs can mount it at once, on the same machine or on several:
"datavolume_refs": [
{"name": "food101", "mount_path": "/data", "mode": "ReadOnly"},
{"name": "cifar-runs", "mount_path": "/out", "mode": "ReadWrite"}
]
With sub_path a run mounts one folder of the drive:
{"name": "food101", "sub_path": "food101/data", "mount_path": "/parquet"}.
6. Clean up#
Delete the runs from their pages, then delete the drive from its page.
Deleting the drive deletes its copy on every machine and its fill runs. A drive cannot be deleted while a live worker uses it.
Variations#
One copy, reached over the network. If the dataset is too large to keep on every machine, keep it on one machine and let runs elsewhere read it over NFS, over RDMA when both machines have it:
{
"metadata": {"name": "food101-gpu01"},
"spec": {
"sources": [
{"node": "gpu-01", "path": "/datasets/food101", "mount_path": "/data", "mode": "ReadWrite"}
],
"transport": "Auto"
}
}
The path must be under a host path your workspace was granted
(Settings → Clusters → Change → Host paths, by an organisation admin).
Load it once with a run pinned to that machine
("node_selection": {"mode": "Exact", "names": ["gpu-01"]}) that mounts it
read-write, then mount it read-only everywhere. "transport": "Local"
instead makes every run that mounts it go to gpu-01. Astraeus never
deletes the files of such a drive.
A shared filesystem. If the dataset already sits on Lustre, GPFS or NFS
mounted at the same path on every machine, describe it with
"node_scope": "shared": each worker reads it directly.
From object storage. A fill can read from S3 or any store your
machines can reach. Give the fill's template a credential with
secret_refs and use the store's own tool (aws s3 sync s3://… /drive/…,
then touch /drive/.astralyx-filled). See
Use cloud credentials without storing them.
Fill again. After a fill completes, if the copy is still not whole ten minutes later (it was evicted, or the fill wrote no marker), the fill runs again. A fill that failed stays failed so you can read why; delete the fill run to have it made again.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
| The run waits with Choose where to keep data on gpu-01 | The machine has no data location. | Confirm one on the machine's page. |
| The worker never leaves Waiting for drive … food101 to be filled on gpu-01, and the fill run completed | The fill did not write /drive/.astralyx-filled. |
Make the marker the last command, after a set -e so a failed download stops before it. |
| The fill run failed | The download failed: read its logs (astra astraeus logs fill-food101-gpu-01-0). |
Fix the template (a drive cannot be edited: delete it and create it again), or delete the fill run to retry. |
| No machine is chosen; the worker's page says drive food101 needs … at this machine's data location, … free | size_hint_bytes is more than the machines' data locations have free (or than their limit allows). |
Free space, raise the machine's data location limit, or lower the hint. |