Skip to content

Drives from a data source (datasets)#

A dataset in Astraeus is a drive backed by a data source: the files under a folder of a bucket (S3, GCS, Azure Blob, OCI, any S3-compatible store), of shared storage your machines mount (NFS, Lustre, Weka, VAST, a local path), or of a Hugging Face repository. Runs mount it by name, read-only, like any drive.

What you get:

  • Revisions. The drive's revision is the source's listing. A run reads exactly the files it was made with, and the drive says what changed between two revisions, file by file.
  • The data on the machine before the worker starts. A worker waits until its machine has the files, and its reason says from where and how far. It never starts on half a dataset.
  • The source read once. One machine reads the data source. Every other machine copies the files from a machine that already has them (or is copying them), over your machines' own network, at a capped rate. Sixteen machines reading a 2 TB dataset read 2 TB from the bucket, not 32 TB.
  • Three ways to read it, per drive and per run: copy (the whole revision on each machine), shard (each worker gets only its own slice of the files) and stream (read through a cache as files are opened).
  • Placement that prefers machines holding the revision, without ever waiting for one.
  • Lineage. The drive lists the runs that read it (and which revision) and wrote it; a run lists the drives it read and wrote; a model says which drive its weights are in.

The files stay between your data source and your machines. Credentials are credentials resolved on the machines.

Before you begin#

  • You need the admin or editor role in the workspace to create drives and to list them again. Viewers can see drives, revisions and lineage.
  • A data source in the workspace, Ready. Its kind holds files: S3, MinIO, GCS, Azure, Oracle, NFS (with host_path), Local or HuggingFace.
  • The machines that will read the drive need a data location with room for a copy (or for their shards).
  • The agent's data part, astraeus-agent-data, on each machine (installed by default). Machines copy from each other on TCP port 7479: allow it between your machines (see Copies between machines).
  • For the API examples, set ASTRALYX_API and ASTRALYX_TOKEN as in Drives.

Create a drive from a data source#

  1. Open Drives and click New drive.
  2. Enter a Name (for example corpus).
  3. Under What it is, choose From a data source (a dataset).
  4. Choose the Data source and, optionally, a Folder in it (a prefix such as corpus/v2; empty means all of it).
  5. Choose How runs read it: Copy, Shard or Stream (see Copy, shard or stream).
  6. Optional: List it again every (for example 1d). Empty means only when you ask.
  7. Click Create drive. The drive's page opens; its data source is listed within seconds (a large one takes longer).
$ astra astraeus drives create corpus --source lake --path corpus/v2 --mode copy --refresh 1d
corpus created; its data source is being listed (astra astraeus drives revisions corpus)
$ curl -sS -X POST "$ASTRALYX_API/drives" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H "Content-Type: application/json" \
    -d '{
      "metadata": {"name": "corpus"},
      "spec": {"data_source": {"name": "lake", "path": "corpus/v2", "mode": "copy", "refresh_seconds": 86400}}
    }'

With the data toolset, ask your assistant to create the drive; it calls create_drive.

Field of spec.data_source Default Description
name required The data source, in the same workspace.
path all of it A folder (prefix) within it: relative, without ...
mode copy How runs read it unless they say: copy, stream or shard.
refresh_seconds 0 List the source again this often, at least 60. 0: only when asked.

Such a drive is read-only, has no sources of its own (its copies are kept in each machine's data location, at <data location>/drives/<drive>/<revision>), and is a cache: a machine's copy no worker uses can be removed when its disk fills, and is copied again when needed. Deleting the drive deletes its copies, never the data source.

Errors when creating:

  • 400 INVALID_DATAVOLUME — spec.data_source.name: a data source of the drive's own workspace, spec.data_source.path is a folder within the data source: relative, without '..', spec.sources[0].mode: a drive backed by a data source is read-only (write results to another drive), spec.data_source.refresh_seconds: 0 (only when asked) or at least 60.
  • 403 NAMESPACE_FORBIDDEN — the data source is another workspace's.

Revisions#

A revision is the data source's listing under the drive's folder: every file's path, size and identity — an object's ETag, a Hugging Face file's hash, a file's modification time on a filesystem — named r- and 16 hex digits (for example r-3f9a1c0d2b7e4a56). The same files are the same revision wherever they are listed; any change is a new revision.

The machine serving the data source lists it:

  • when the drive is created,
  • every refresh_seconds, if set,
  • when you ask: List it now on the drive's page, astra astraeus drives refresh corpus, or POST /v1/drives/corpus/refresh (it waits up to ?wait= seconds, 20 by default and at most 120, then answers 202 while a large source is still being listed),
  • five minutes after a listing failed.

A new revision is an event on the drive: Revision r-3f9a1c0d2b7e4a56: 20,412 files, 2.0 TB (+412 −0 ~3 since r-0b6e2f9c11d4a873). A failed listing keeps the current revision and shows why, in the machine's words. The drive keeps its newest 50 revisions. A Hugging Face revision also records the commit its branch resolved to, and is always read at that commit. A drive lists at most 10,000,000 files.

The drive's page shows the Data source (source, how it is read, the revision, when it was listed) and a Revisions tab with every revision, its files and size, and What changed — the paths added, removed and changed since the one before.

$ astra astraeus drives revisions corpus
REVISION                        FILES       SIZE        LISTED                  CHANGES
r-3f9a1c0d2b7e4a56 (current)    20412       2.0 TB      2026-10-07T09:12:00Z    +412 -0 ~3
r-0b6e2f9c11d4a873              20000       1.9 TB      2026-10-01T09:12:00Z    first
$ astra astraeus drives diff corpus
r-0b6e2f9c11d4a873 → r-3f9a1c0d2b7e4a56: +412 −0 ~3
  + shard-20000.parquet
  …

GET /v1/drives/corpus/revisions returns current and items. GET /v1/drives/corpus/diff?from=<r>&to=<r>&limit=1000 returns counts, added, removed and changed. The paths are answered by the machine that lists the data source; it keeps the listings of the drive's last 20 revisions (503 DRIVE_DIFF_FAILED for an older one). File names stay on your machines.

Read it in a run#

A run mounts it in drives like any drive. Two fields apply to drives backed by a data source:

Field Default Description
revision the drive's current revision The revision the run reads. Left empty, the current revision is written into the run's specification when the run is created, so every worker and every restart reads the same files. A revision the drive does not have is refused (400 VALIDATION_ERROR: drive corpus has no revision r-… (see its revisions)).
data_mode the drive's mode copy, stream or shard, for this run.

The mount is always read-only: mode: ReadWrite is refused (drive corpus is backed by a data source and is read-only: write results to another drive). A drive created without a folder of its own appears at /data/<drive> unless the run sets mount_path.

A run reading revision r-3f9a1c0d2b7e4a56 of corpus, whole
{
  "metadata": {"name": "count-files"},
  "spec": {
    "task_template": {
      "image": "python:3.12-slim",
      "command": "sh",
      "args": ["-c", "echo revision $ASTRAEUS_DRIVE_CORPUS_REVISION; find /data/corpus -type f -not -path '*/.astralyx*' | wc -l"],
      "restart_policy": "Never",
      "requested_resources": {"cpu_cores": 1, "memory_bytes": 1073741824},
      "drives": [{"name": "corpus", "mount_path": "/data/corpus", "revision": "r-3f9a1c0d2b7e4a56"}]
    }
  }
}

With the CLI, --drive <drive>[@<revision>][:<path>][:<mode>]:

$ astra astraeus run --image python:3.12-slim --drive corpus@r-3f9a1c0d2b7e4a56:/data/corpus -- ls /data/corpus

Every worker gets, for each drive backed by a data source (<NAME> is the drive's name in capitals, other characters as _):

Variable Value
ASTRAEUS_DRIVE_<NAME>_REVISION The revision mounted.
ASTRAEUS_DRIVE_<NAME>_PATH Where it is mounted.

Copy, shard or stream#

Mode On the machine Use it for
copy The whole revision, copied before the worker starts. Data every worker reads (most training, evaluation).
shard Only the files of the shards of that machine's workers. Data parallelism where each worker reads its own part.
stream Nothing ahead: files are read through a cache as they are opened. Data larger than the machines' disks, or read once.

Shards#

In shard mode each worker of the run gets a distinct, deterministic slice of the revision's files, by its rank and the run's size:

  • The files are sorted by path and laid end to end. Shard i of n holds the files whose first byte falls in the i-th of n equal parts of the total bytes. Every file is in exactly one shard; shards are as even as whole files allow; the same revision and the same n always split the same way. With more workers than files some shards are empty.
  • Only the shards of the workers on a machine are copied to it. Each machine reads its own from the data source when no machine holds the whole revision.

Each worker gets:

Variable Value
ASTRAEUS_DRIVE_<NAME>_SHARD_INDEX Its shard: its rank.
ASTRAEUS_DRIVE_<NAME>_SHARD_COUNT How many shards: the run's workers.
ASTRAEUS_DRIVE_<NAME>_SHARD_FILES A file listing its shard's files, one path per line, relative to the drive's mount.
ASTRAEUS_SHARD_INDEX, ASTRAEUS_SHARD_COUNT, ASTRAEUS_SHARD_FILES, ASTRAEUS_SHARD_DRIVE The same, when the worker has exactly one drive in shard mode (ASTRAEUS_SHARD_DRIVE is its mount path).
Reading your shard
import os
root = os.environ["ASTRAEUS_SHARD_DRIVE"]
with open(os.environ["ASTRAEUS_SHARD_FILES"]) as f:
    mine = [os.path.join(root, line.rstrip("\n")) for line in f if line.strip()]

Mount a drive in shard mode without sub_path: the list's paths are relative to the drive.

Stream#

  • A data source the machines already mount (NFS with host_path, Local) is read as it is, at that path: nothing is copied.
  • An object store (S3, MinIO, GCS, Azure) is mounted read-only through a cache on the machine. This needs rclone and FUSE on the machine; without them the worker waits with stream mode needs rclone on this machine … or use copy mode. The credential is given to rclone in its environment only.
  • Oracle and HuggingFace data sources are copied (copy or shard).
  • A stream reads the source as it is now. The run records the revision current when it starts.

Before the worker starts#

A worker whose drive is not yet on its machine waits in Preparing, its reason saying where the copy comes from and how far it is:

  • Waiting for drive corpus: copying from machine gpu-03: 1.2 TB of 2.0 TB (60%)
  • Waiting for drive corpus: copying from its data source: 310 GB of 2.0 TB (15%)
  • Waiting for drive corpus: copying from its data source (other machines could not be reached from here on port 7479): 310 GB of 2.0 TB (15%) — see When machines cannot reach each other
  • Waiting for drive corpus: its data source has not been listed yet
  • Waiting for drive corpus: listing its data source failed: … (the source's own words)
  • Waiting for drive corpus: copying it here failed: the data source no longer lists as revision r-… (it lists as r-… now): refresh the drive and read r-…, or keep r-… on machines that hold it

A run of several workers that start together waits for every member's copy; it is not restarted while copies are being made. A failed copy is tried again after 30 seconds, doubling to 10 minutes, from another machine when there is one; a machine that cannot copy from the others reads the data source itself (see below).

A copy is resumed where it stopped, and files unchanged since a revision the machine already holds are linked, not copied again. A machine keeps the revisions its runs read, and the newest whole one after that; older ones go.

Copies between machines#

One machine reads the data source. Every other machine copies the files from a machine of the same organisation that already has them, or is copying them — each file is passed on as soon as it is whole, so the copies follow one another closely. A machine serves at most two others at a time.

  • Over the address the machine registered with (its management network), on TCP port 7479, never over the InfiniBand or RoCE fabric your runs use.
  • Capped, so it never competes with the work: by default each machine sends and receives at most 2 Gbit/s to and from other machines. Reading the data source is not capped by default.
  • Only to the machine the copy is for, at its address.

The machine's settings, in /etc/astraeus/agent.env (then systemctl restart astraeus-agent-data):

Setting Default Description
ASTRAEUS_DATA_P2P_LISTEN 0.0.0.0:7479 Where it serves copies to other machines; off: not served (other machines read the source).
ASTRAEUS_DATA_P2P_MAX_MBPS 2000 The cap on copying between machines, each way, megabits a second; 0: none.
ASTRAEUS_DATA_SOURCE_MAX_MBPS 0 The cap on reading data sources, megabits a second; 0: none.
ASTRAEUS_DATA_MAX_DRIVE_FILES 10000000 The most files a drive may list.

The drive's page shows, under Copying to machines, each copy being made: from where (its data source, or a machine), its state and progress, and in all how much came from the source and how much from machines. A copy is checked against the revision's listing (every file and its size); file contents are not checksummed.

Model weights Eos pulls from Hugging Face are made the same way: one machine downloads them, the others copy them from it.

When machines cannot reach each other#

Copies between machines need TCP port 7479 open from each machine to the others, at the addresses they registered with. Machines in different sites or homes, behind NAT, or with a firewall between them cannot copy from one another. A machine in that position does not wait for ever:

  • A machine that a copy failed to come from twice is not tried again for that copy; another machine is tried instead.
  • After three copies from other machines failed one after another — or as soon as every machine it could copy from has failed it twice — the machine reads the data source itself. Its workers' reason says so: other machines cannot be reached from here (port 7479): reading the source.
  • From then on that copy keeps reading the data source, even if that read fails and is tried again.

Each failed copy waits before the next try (30 seconds, then doubling), so a machine that reaches no other starts reading the data source after a few minutes. Each such machine reads the source in full, in addition to the one that normally does: open port 7479 between your machines to keep the source read once.

Model weights copied between machines follow the same rules. A machine that cannot copy them from the others downloads them itself, with the same download that made the first copy; its workers wait with other machines cannot be reached from here (port 7479): downloading it here until the download is whole.

Copy a drive ahead of a run#

To have the data in place before you submit:

  • Console: on the drive's page, Copy to a machine ahead of runs…, choose the machine, Copy.
  • CLI: astra astraeus drives stage corpus gpu-01 gpu-02 (--revision r-… for another than the current).
  • API: POST /v1/drives/corpus/stage with {"machine": "gpu-01"} (and "revision"). It answers 202 with the copy; GET /v1/drives/corpus/stages follows them all.

Where runs go#

A run is placed, preferably, on machines that already hold its revision whole; a machine holding another revision of the drive comes next (its unchanged files are linked). This is a preference: when no such machine is free, the run goes to another one, and the copy is made there.

Lineage#

Where What it shows
A drive's Lineage tab, GET /v1/drives/<drive>/lineage, astra astraeus drives lineage <drive> The runs that read it (with the revision each read), the runs that wrote it (mounted it read-write), and the models whose weights are in it. The newest 200 runs, kept after the runs themselves are cleaned up.
A run's page (Data), GET /v1/runs/<run>/lineage, astra astraeus drives lineage --run <run> The drives it read (and which revision) and wrote.
A model's row on Models, GET /v1/eos/models/<model>/lineage The drive its weights are in, and the runs that wrote that drive.

Your AI assistant reads the same with get_lineage (data toolset).

A Hugging Face repository#

A data source for a public Hugging Face dataset
{
  "metadata": {"name": "fineweb-edu"},
  "spec": {"kind": "HuggingFace", "huggingface": {"repo": "HuggingFaceFW/fineweb-edu", "repo_type": "dataset", "revision": "main"}}
}

Then a drive of one of its folders:

{"metadata": {"name": "fineweb-sample"}, "spec": {"data_source": {"name": "fineweb-edu", "path": "sample/10BT", "mode": "shard"}}}

A gated or private repository needs a credential whose secret holds token, named in the data source's credential.