Skip to content

Shared storage#

Astraeus does not ship a storage system. It uses the one you have: storage every machine mounts at the same path, or an object store. For a team that has none, this page gives a supported recipe built from open-source parts over your own object storage.

Storage you already have#

Your machines report the filesystems they mount. A filesystem other machines see alike is reported shared: NFS (including AWS EFS, GCP Filestore, Azure NetApp Files, VAST), SMB, Lustre (including FSx and Azure Managed Lustre), GPFS (Storage Scale), BeeGFS, CephFS, Weka, JuiceFS, SeaweedFS, GlusterFS, DAOS, and a cloud's shared filesystem over virtiofs. An object store mounted through FUSE (s3fs, Mountpoint for Amazon S3, gcsfuse, BlobFuse, rclone…) is reported too, as an object store.

A machine's page lists them under its disks, with the kind and how full they are. There are two ways to use them:

You want Use What happens
Runs that read and write the same files on every machine (checkpoints, shared results) A drive of kind Shared filesystem (machine_scope: shared) at the mount path Workers bind-mount the path. A run goes only to machines that report that filesystem mounted: elsewhere the machine is refused with drive …: /lustre/x is not on a shared filesystem mounted here.
A read-only dataset with revisions, diffs and lineage A data source of kind NFS (host_path) or Local at the mount path, and a drive from it in stream mode Workers read the path as it is, nothing copied; the drive's revision is the listing.
A dataset in a bucket (S3, GCS, Azure, OCI, S3-compatible) A data source of that kind and a drive from it in copy or shard mode The bucket is read once; machines copy from each other.

The paths must be inside the workspace's granted Host paths, as for any path on the machines.

A data location is never on shared storage

A machine's data location (where drive copies go) must be on its own disk: a path on a shared filesystem is refused. Shared storage is used through shared drives, not as a place for per-machine copies.

No shared storage yet: JuiceFS or SeaweedFS over your object storage#

A team without shared storage can build one from its own object storage, and Astraeus uses it like any other: the machines report it shared, and a shared drive goes only where it is mounted. This is a recipe, not a service: you run it, on your machines or alongside them.

Choose one:

  • JuiceFS Community Edition keeps file data in your bucket (S3, GCS, Azure, OCI, MinIO-compatible) and metadata in a database you run (Redis or PostgreSQL). It mounts through FUSE with a local cache. Choose it when you already have object storage.
  • SeaweedFS stores the data itself, on disks of machines you choose, and mounts through FUSE. Choose it when you have disks but no object storage.

JuiceFS over your S3 bucket#

Below, <bucket-url> stands for your bucket (https://<bucket>.s3.<region>.amazonaws.com), <meta> for the metadata engine's URL (for example redis://10.0.0.5:6379/1 or postgres://jfs:<password>@10.0.0.5:5432/juicefs), and /mnt/shared for the mount point.

  1. Run the metadata engine somewhere every machine reaches (one Redis or PostgreSQL; make it durable: it holds the file tree).
  2. Install JuiceFS on every machine (curl -sSL https://d.juicefs.com/install | sh -), then create the filesystem once, from one machine:

    $ juicefs format --storage s3 --bucket <bucket-url> <meta> shared
    

    Give the bucket's keys with --access-key/--secret-key, or let the machine's cloud identity reach it.

  3. Mount it on every machine, at boot:

    /etc/fstab
    <meta>  /mnt/shared  juicefs  _netdev,cache-size=102400  0 0
    
    $ sudo mount /mnt/shared
    
  4. Check that every machine reports it: on each machine's page, its disks list /mnt/shared as shared (fuse.juicefs).

  5. Grant the workspace the path (Host paths: /mnt/shared), then make a drive of kind Shared filesystem at /mnt/shared/<folder>.

SeaweedFS on your machines' disks#

  1. Run SeaweedFS on the machines whose disks hold the data: weed server -dir=/data/seaweed -filer on one machine to start, more volume servers as it grows (see SeaweedFS's documentation for replication).
  2. Mount it on every machine: weed mount -filer=<filer-host>:8888 -dir=/mnt/shared, as a service at boot.
  3. Check that every machine reports /mnt/shared as shared (fuse.seaweedfs), grant the path and make a shared drive, as above.

Datasets in either

For a read-only dataset kept in JuiceFS or SeaweedFS, make a data source of kind Local at the mounted path and a drive from it in stream mode: runs read it as it is, with revisions and lineage.