Drives from a data source (datasets)#
A dataset in Astraeus is a drive backed by a data source: the files under a folder of a bucket (S3, GCS, Azure Blob, OCI, any S3-compatible store), of shared storage your machines mount (NFS, Lustre, Weka, VAST, a local path), or of a Hugging Face repository. Runs mount it by name, read-only, like any drive.
What you get:
- Revisions. The drive's revision is the source's listing. A run reads exactly the files it was made with, and the drive says what changed between two revisions, file by file.
- The data on the machine before the worker starts. A worker waits until its machine has the files, and its reason says from where and how far. It never starts on half a dataset.
- The source read once. One machine reads the data source. Every other machine copies the files from a machine that already has them (or is copying them), over your machines' own network, at a capped rate. Sixteen machines reading a 2 TB dataset read 2 TB from the bucket, not 32 TB.
- Three ways to read it, per drive and per run:
copy(the whole revision on each machine),shard(each worker gets only its own slice of the files) andstream(read through a cache as files are opened). - Placement that prefers machines holding the revision, without ever waiting for one.
- Lineage. The drive lists the runs that read it (and which revision) and wrote it; a run lists the drives it read and wrote; a model says which drive its weights are in.
The files stay between your data source and your machines. Credentials are credentials resolved on the machines.
Before you begin#
- You need the admin or editor role in the workspace to create drives and to list them again. Viewers can see drives, revisions and lineage.
- A data source in the workspace,
Ready. Its kind holds files:S3,MinIO,GCS,Azure,Oracle,NFS(withhost_path),LocalorHuggingFace. - The machines that will read the drive need a data location with room for a copy (or for their shards).
- The agent's data part,
astraeus-agent-data, on each machine (installed by default). Machines copy from each other on TCP port 7479: allow it between your machines (see Copies between machines). - For the API examples, set
ASTRALYX_APIandASTRALYX_TOKENas in Drives.
Create a drive from a data source#
- Open Drives and click New drive.
- Enter a Name (for example
corpus). - Under What it is, choose From a data source (a dataset).
- Choose the Data source and, optionally, a Folder in it (a prefix such as
corpus/v2; empty means all of it). - Choose How runs read it: Copy, Shard or Stream (see Copy, shard or stream).
- Optional: List it again every (for example
1d). Empty means only when you ask. - Click Create drive. The drive's page opens; its data source is listed within seconds (a large one takes longer).
With the data toolset, ask your assistant to create the drive; it calls create_drive.
Field of spec.data_source |
Default | Description |
|---|---|---|
name |
required | The data source, in the same workspace. |
path |
all of it | A folder (prefix) within it: relative, without ... |
mode |
copy |
How runs read it unless they say: copy, stream or shard. |
refresh_seconds |
0 |
List the source again this often, at least 60. 0: only when asked. |
Such a drive is read-only, has no sources of its own (its copies are kept in each machine's data location, at <data location>/drives/<drive>/<revision>), and is a cache: a machine's copy no worker uses can be removed when its disk fills, and is copied again when needed. Deleting the drive deletes its copies, never the data source.
Errors when creating:
400 INVALID_DATAVOLUME—spec.data_source.name: a data source of the drive's own workspace,spec.data_source.path is a folder within the data source: relative, without '..',spec.sources[0].mode: a drive backed by a data source is read-only (write results to another drive),spec.data_source.refresh_seconds: 0 (only when asked) or at least 60.403 NAMESPACE_FORBIDDEN— the data source is another workspace's.
Revisions#
A revision is the data source's listing under the drive's folder: every file's path, size and identity — an object's ETag, a Hugging Face file's hash, a file's modification time on a filesystem — named r- and 16 hex digits (for example r-3f9a1c0d2b7e4a56). The same files are the same revision wherever they are listed; any change is a new revision.
The machine serving the data source lists it:
- when the drive is created,
- every
refresh_seconds, if set, - when you ask: List it now on the drive's page,
astra astraeus drives refresh corpus, orPOST /v1/drives/corpus/refresh(it waits up to?wait=seconds, 20 by default and at most 120, then answers202while a large source is still being listed), - five minutes after a listing failed.
A new revision is an event on the drive: Revision r-3f9a1c0d2b7e4a56: 20,412 files, 2.0 TB (+412 −0 ~3 since r-0b6e2f9c11d4a873). A failed listing keeps the current revision and shows why, in the machine's words. The drive keeps its newest 50 revisions. A Hugging Face revision also records the commit its branch resolved to, and is always read at that commit. A drive lists at most 10,000,000 files.
The drive's page shows the Data source (source, how it is read, the revision, when it was listed) and a Revisions tab with every revision, its files and size, and What changed — the paths added, removed and changed since the one before.
$ astra astraeus drives revisions corpus
REVISION FILES SIZE LISTED CHANGES
r-3f9a1c0d2b7e4a56 (current) 20412 2.0 TB 2026-10-07T09:12:00Z +412 -0 ~3
r-0b6e2f9c11d4a873 20000 1.9 TB 2026-10-01T09:12:00Z first
$ astra astraeus drives diff corpus
r-0b6e2f9c11d4a873 → r-3f9a1c0d2b7e4a56: +412 −0 ~3
+ shard-20000.parquet
…
GET /v1/drives/corpus/revisions returns current and items. GET /v1/drives/corpus/diff?from=<r>&to=<r>&limit=1000 returns counts, added, removed and changed. The paths are answered by the machine that lists the data source; it keeps the listings of the drive's last 20 revisions (503 DRIVE_DIFF_FAILED for an older one). File names stay on your machines.
Read it in a run#
A run mounts it in drives like any drive. Two fields apply to drives backed by a data source:
| Field | Default | Description |
|---|---|---|
revision |
the drive's current revision | The revision the run reads. Left empty, the current revision is written into the run's specification when the run is created, so every worker and every restart reads the same files. A revision the drive does not have is refused (400 VALIDATION_ERROR: drive corpus has no revision r-… (see its revisions)). |
data_mode |
the drive's mode |
copy, stream or shard, for this run. |
The mount is always read-only: mode: ReadWrite is refused (drive corpus is backed by a data source and is read-only: write results to another drive). A drive created without a folder of its own appears at /data/<drive> unless the run sets mount_path.
{
"metadata": {"name": "count-files"},
"spec": {
"task_template": {
"image": "python:3.12-slim",
"command": "sh",
"args": ["-c", "echo revision $ASTRAEUS_DRIVE_CORPUS_REVISION; find /data/corpus -type f -not -path '*/.astralyx*' | wc -l"],
"restart_policy": "Never",
"requested_resources": {"cpu_cores": 1, "memory_bytes": 1073741824},
"drives": [{"name": "corpus", "mount_path": "/data/corpus", "revision": "r-3f9a1c0d2b7e4a56"}]
}
}
}
With the CLI, --drive <drive>[@<revision>][:<path>][:<mode>]:
$ astra astraeus run --image python:3.12-slim --drive corpus@r-3f9a1c0d2b7e4a56:/data/corpus -- ls /data/corpus
Every worker gets, for each drive backed by a data source (<NAME> is the drive's name in capitals, other characters as _):
| Variable | Value |
|---|---|
ASTRAEUS_DRIVE_<NAME>_REVISION |
The revision mounted. |
ASTRAEUS_DRIVE_<NAME>_PATH |
Where it is mounted. |
Copy, shard or stream#
| Mode | On the machine | Use it for |
|---|---|---|
copy |
The whole revision, copied before the worker starts. | Data every worker reads (most training, evaluation). |
shard |
Only the files of the shards of that machine's workers. | Data parallelism where each worker reads its own part. |
stream |
Nothing ahead: files are read through a cache as they are opened. | Data larger than the machines' disks, or read once. |
Shards#
In shard mode each worker of the run gets a distinct, deterministic slice of the revision's files, by its rank and the run's size:
- The files are sorted by path and laid end to end. Shard
iofnholds the files whose first byte falls in thei-th ofnequal parts of the total bytes. Every file is in exactly one shard; shards are as even as whole files allow; the same revision and the samenalways split the same way. With more workers than files some shards are empty. - Only the shards of the workers on a machine are copied to it. Each machine reads its own from the data source when no machine holds the whole revision.
Each worker gets:
| Variable | Value |
|---|---|
ASTRAEUS_DRIVE_<NAME>_SHARD_INDEX |
Its shard: its rank. |
ASTRAEUS_DRIVE_<NAME>_SHARD_COUNT |
How many shards: the run's workers. |
ASTRAEUS_DRIVE_<NAME>_SHARD_FILES |
A file listing its shard's files, one path per line, relative to the drive's mount. |
ASTRAEUS_SHARD_INDEX, ASTRAEUS_SHARD_COUNT, ASTRAEUS_SHARD_FILES, ASTRAEUS_SHARD_DRIVE |
The same, when the worker has exactly one drive in shard mode (ASTRAEUS_SHARD_DRIVE is its mount path). |
import os
root = os.environ["ASTRAEUS_SHARD_DRIVE"]
with open(os.environ["ASTRAEUS_SHARD_FILES"]) as f:
mine = [os.path.join(root, line.rstrip("\n")) for line in f if line.strip()]
Mount a drive in shard mode without sub_path: the list's paths are relative to the drive.
Stream#
- A data source the machines already mount (
NFSwithhost_path,Local) is read as it is, at that path: nothing is copied. - An object store (
S3,MinIO,GCS,Azure) is mounted read-only through a cache on the machine. This needs rclone and FUSE on the machine; without them the worker waits with stream mode needs rclone on this machine … or use copy mode. The credential is given to rclone in its environment only. OracleandHuggingFacedata sources are copied (copyorshard).- A stream reads the source as it is now. The run records the revision current when it starts.
Before the worker starts#
A worker whose drive is not yet on its machine waits in Preparing, its reason saying where the copy comes from and how far it is:
- Waiting for drive corpus: copying from machine gpu-03: 1.2 TB of 2.0 TB (60%)
- Waiting for drive corpus: copying from its data source: 310 GB of 2.0 TB (15%)
- Waiting for drive corpus: copying from its data source (other machines could not be reached from here on port 7479): 310 GB of 2.0 TB (15%) — see When machines cannot reach each other
- Waiting for drive corpus: its data source has not been listed yet
- Waiting for drive corpus: listing its data source failed: … (the source's own words)
- Waiting for drive corpus: copying it here failed: the data source no longer lists as revision r-… (it lists as r-… now): refresh the drive and read r-…, or keep r-… on machines that hold it
A run of several workers that start together waits for every member's copy; it is not restarted while copies are being made. A failed copy is tried again after 30 seconds, doubling to 10 minutes, from another machine when there is one; a machine that cannot copy from the others reads the data source itself (see below).
A copy is resumed where it stopped, and files unchanged since a revision the machine already holds are linked, not copied again. A machine keeps the revisions its runs read, and the newest whole one after that; older ones go.
Copies between machines#
One machine reads the data source. Every other machine copies the files from a machine of the same organisation that already has them, or is copying them — each file is passed on as soon as it is whole, so the copies follow one another closely. A machine serves at most two others at a time.
- Over the address the machine registered with (its management network), on TCP port 7479, never over the InfiniBand or RoCE fabric your runs use.
- Capped, so it never competes with the work: by default each machine sends and receives at most 2 Gbit/s to and from other machines. Reading the data source is not capped by default.
- Only to the machine the copy is for, at its address.
The machine's settings, in /etc/astraeus/agent.env (then systemctl restart astraeus-agent-data):
| Setting | Default | Description |
|---|---|---|
ASTRAEUS_DATA_P2P_LISTEN |
0.0.0.0:7479 |
Where it serves copies to other machines; off: not served (other machines read the source). |
ASTRAEUS_DATA_P2P_MAX_MBPS |
2000 |
The cap on copying between machines, each way, megabits a second; 0: none. |
ASTRAEUS_DATA_SOURCE_MAX_MBPS |
0 |
The cap on reading data sources, megabits a second; 0: none. |
ASTRAEUS_DATA_MAX_DRIVE_FILES |
10000000 |
The most files a drive may list. |
The drive's page shows, under Copying to machines, each copy being made: from where (its data source, or a machine), its state and progress, and in all how much came from the source and how much from machines. A copy is checked against the revision's listing (every file and its size); file contents are not checksummed.
Model weights Eos pulls from Hugging Face are made the same way: one machine downloads them, the others copy them from it.
When machines cannot reach each other#
Copies between machines need TCP port 7479 open from each machine to the others, at the addresses they registered with. Machines in different sites or homes, behind NAT, or with a firewall between them cannot copy from one another. A machine in that position does not wait for ever:
- A machine that a copy failed to come from twice is not tried again for that copy; another machine is tried instead.
- After three copies from other machines failed one after another — or as soon as every machine it could copy from has failed it twice — the machine reads the data source itself. Its workers' reason says so: other machines cannot be reached from here (port 7479): reading the source.
- From then on that copy keeps reading the data source, even if that read fails and is tried again.
Each failed copy waits before the next try (30 seconds, then doubling), so a machine that reaches no other starts reading the data source after a few minutes. Each such machine reads the source in full, in addition to the one that normally does: open port 7479 between your machines to keep the source read once.
Model weights copied between machines follow the same rules. A machine that cannot copy them from the others downloads them itself, with the same download that made the first copy; its workers wait with other machines cannot be reached from here (port 7479): downloading it here until the download is whole.
Copy a drive ahead of a run#
To have the data in place before you submit:
- Console: on the drive's page, Copy to a machine ahead of runs…, choose the machine, Copy.
- CLI:
astra astraeus drives stage corpus gpu-01 gpu-02(--revision r-…for another than the current). - API:
POST /v1/drives/corpus/stagewith{"machine": "gpu-01"}(and"revision"). It answers202with the copy;GET /v1/drives/corpus/stagesfollows them all.
Where runs go#
A run is placed, preferably, on machines that already hold its revision whole; a machine holding another revision of the drive comes next (its unchanged files are linked). This is a preference: when no such machine is free, the run goes to another one, and the copy is made there.
Lineage#
| Where | What it shows |
|---|---|
A drive's Lineage tab, GET /v1/drives/<drive>/lineage, astra astraeus drives lineage <drive> |
The runs that read it (with the revision each read), the runs that wrote it (mounted it read-write), and the models whose weights are in it. The newest 200 runs, kept after the runs themselves are cleaned up. |
A run's page (Data), GET /v1/runs/<run>/lineage, astra astraeus drives lineage --run <run> |
The drives it read (and which revision) and wrote. |
A model's row on Models, GET /v1/eos/models/<model>/lineage |
The drive its weights are in, and the runs that wrote that drive. |
Your AI assistant reads the same with get_lineage (data toolset).
A Hugging Face repository#
{
"metadata": {"name": "fineweb-edu"},
"spec": {"kind": "HuggingFace", "huggingface": {"repo": "HuggingFaceFW/fineweb-edu", "repo_type": "dataset", "revision": "main"}}
}
Then a drive of one of its folders:
{"metadata": {"name": "fineweb-sample"}, "spec": {"data_source": {"name": "fineweb-edu", "path": "sample/10BT", "mode": "shard"}}}
A gated or private repository needs a credential whose secret holds token, named in the data source's credential.