Skip to content

Inside a worker#

Each worker is a container on one machine. This page describes what your program finds when it starts there: the environment variables Astraeus sets, where drives, credentials and files are mounted, which GPUs it sees, which user it runs as, its limits, how it reaches other workers, and what happens when it is stopped. Use it when you write a training script or an entrypoint for Astraeus.

A quick look#

Start a run that prints its environment, then read its log:

$ astra astraeus run --name whoami --image busybox --cpus 1 --mem 256M -- sh -c 'env | grep -E "^(ASTRAEUS|SLURM|RANK|WORLD|MASTER)" | sort; id; df -h /dev/shm'
whoami submitted to gpu-east: https://console.astralyx.cloud/o/acme/w/vision/jobs/gpu-east/whoami
$ astra astraeus logs whoami
ASTRAEUS_GROUP_NAME=default
ASTRAEUS_GROUP_REPLICA_INDEX=0
ASTRAEUS_JOB_NAME=ws-7f3c2a.whoami
ASTRAEUS_LEADER_ADDRESS=10.0.4.17
ASTRAEUS_LEADER_TASK_NAME=ws-7f3c2a.whoami-0
ASTRAEUS_NODE_NAME=gpu-09
ASTRAEUS_REPLICA_INDEX=0
ASTRAEUS_TASK_NAME=ws-7f3c2a.whoami-0
ASTRAEUS_WORLD_SIZE=1
MASTER_ADDR=10.0.4.17
MASTER_PORT=29500
RANK=0
SLURMD_NODENAME=gpu-09
SLURM_CLUSTER_NAME=astraeus
SLURM_CPUS_PER_TASK=1
SLURM_JOBID=whoami
SLURM_JOB_ID=whoami
SLURM_JOB_NAME=whoami
SLURM_JOB_NUM_NODES=1
SLURM_LOCALID=0
SLURM_MEM_PER_NODE=256
SLURM_NNODES=1
SLURM_NODEID=0
SLURM_NPROCS=1
SLURM_NTASKS=1
SLURM_NTASKS_PER_NODE=1
SLURM_PROCID=0
WORLD_SIZE=1
uid=0(root) gid=0(root) groups=0(root),10(wheel)
Filesystem                Size      Used Available Use% Mounted on
shm                      64.0M         0     64.0M   0% /dev/shm

In a workspace, full names carry the workspace's namespace (ws-7f3c2a. above): ASTRAEUS_TASK_NAME, ASTRAEUS_JOB_NAME and ASTRAEUS_LEADER_TASK_NAME are full names; the SLURM_JOB_* variables hold the run's own name.

Environment variables#

Who and where#

Variable Value Set
ASTRAEUS_TASK_NAME The worker's full name (<namespace>.<run>-<rank> in a workspace). Always.
ASTRAEUS_JOB_NAME The run's full name. Always for a run's worker.
ASTRAEUS_NODE_NAME The machine's name. Always.
ASTRAEUS_REPLICA_INDEX The worker's rank, from 0, across all groups. Always.
ASTRAEUS_GROUP_NAME Its worker group; default without task_groups. Always for a run's worker.
ASTRAEUS_GROUP_REPLICA_INDEX Its index within its group. With a group.
RANK Same as ASTRAEUS_REPLICA_INDEX. When the run's size is known (always for a run).
WORLD_SIZE, ASTRAEUS_WORLD_SIZE How many workers the run has. Same.

Rendezvous#

Variable Value Set
ASTRAEUS_LEADER_TASK_NAME The leader's full name (rank 0 of the leader group). Always for a run's worker.
ASTRAEUS_LEADER_ADDRESS The address of the machine the leader runs on. Once the leader is placed.
MASTER_ADDR Same as ASTRAEUS_LEADER_ADDRESS. Same.
MASTER_PORT 29500. Same.

MASTER_ADDR is a machine address. When workers have addresses of their own (the node network), point your rendezvous at the leader's name instead; see Multi-machine runs.

GPUs and the network#

Variable Value Set
ASTRAEUS_GPU_COUNT GPUs given to this worker. With GPUs.
NCCL_DEBUG WARN With GPUs.
NCCL_ASYNC_ERROR_HANDLING 1 With GPUs.
NCCL_SOCKET_IFNAME ^lo,docker,virbr,veth,cni,wg With GPUs.
NCCL_IB_HCA The RDMA ports that are up and fast enough: =mlx5_0:1,=mlx5_2:1. Placed on an InfiniBand fabric or RDMA, or asking for RDMA, on a machine with RDMA ports up.
NCCL_IB_DISABLE 1 Placed across machines over Ethernet.
NVIDIA_IMEX_CHANNELS 0 Placed in a multi-node NVLink domain, on a machine whose GPUs are handed over by the NVIDIA runtime hook.

Arrays#

Variable Value
ASTRAEUS_ARRAY_TASK_ID This copy's index, 0 to size − 1.
ASTRAEUS_ARRAY_SIZE array.size.

Slurm's names#

Set for every worker, so scripts and libraries written for Slurm read the same facts. One worker per machine is Slurm's shape here.

Variable Value
SLURM_JOB_ID, SLURM_JOBID, SLURM_JOB_NAME The run's name (without the namespace).
SLURM_PROCID, SLURM_NODEID The rank.
SLURM_LOCALID 0.
SLURM_NTASKS, SLURM_NPROCS, SLURM_NNODES, SLURM_JOB_NUM_NODES The run's number of workers.
SLURM_NTASKS_PER_NODE 1.
SLURMD_NODENAME The machine.
SLURM_CLUSTER_NAME astraeus.
SLURM_CPUS_PER_TASK The worker's cores (when it has a core count).
SLURM_MEM_PER_NODE The worker's memory, MiB (when it has a memory size).
SLURM_GPUS_ON_NODE, SLURM_GPUS_PER_TASK The worker's GPUs (with GPUs).

SLURM_ARRAY_TASK_ID, SLURM_NODELIST and SLURM_STEP_* are not set. See Slurm compatibility.

Features you turn on#

Variable Value Set when
ASTRAEUS_MODEL_PATH Where the model's weights are (/model by default). spec.model.
ASTRAEUS_MODEL_FILE The model's single GGUF file. spec.model, for a GGUF model.
OPENAI_BASE_URL, OPENAI_MODEL The engine serving the model beside the run, and the model's name there. spec.model.engine, in the main group.
ASTRAEUS_IDENTITY_DIR /var/run/astraeus/identity. identity: true.
ASTRAEUS_OUTPUTS /var/run/astraeus/outputs/outputs, the file to write key=value outputs to. outputs.
ASTRAEUS_TOOL_GATEWAY The machine's tool gateway, http://<address>:8801. tools, on the node network.
<port_env_name> The port allocated to an external access. external_accesses[].port_env_name.
HESPERUS_TOKEN (or interactive.token_env) The session token of an interactive server. interactive.
HOME The user's home from the image's /etc/passwd, or /root. When the image does not set it (containerd runtime).

Which value wins#

From lowest to highest: the variables above that Astraeus injects; external-access port variables; your env; credentials' env_vars; ASTRAEUS_IDENTITY_DIR and ASTRAEUS_OUTPUTS. The network variables (NCCL_IB_HCA, NCCL_IB_DISABLE) and ASTRAEUS_TOOL_GATEWAY are only set when you have not set them. So "env": {"MASTER_PORT": "1234"} replaces the default port, and "env": {"NCCL_DEBUG": "INFO"} turns on NCCL's logging.

Command and working directory#

The container runs command with args; command replaces the image's ENTRYPOINT and CMD. Without command, the image's ENTRYPOINT runs with args, or with its CMD when args is empty. There is no shell unless you run one. The working directory is the image's WORKDIR (/workspace with astra astraeus run --code). See Command, arguments and environment.

User#

The worker runs as task_template.user (name, uid or uid:gid), else the image's USER, else root (uid 0). Files the worker writes to a drive belong to that uid.

Filesystem#

Path What Writable
/ The image. Changes are lost when the worker stops or restarts. Yes
volumes[].mount_path (EmptyDir) An empty directory on the machine, removed with the worker's container. Yes
volumes[].mount_path (Bind) A path of the machine (in a workspace, only granted paths). With mode: ReadWrite
datavolume_refs[].mount_path A drive (the console's default is /data/<drive>). With mode: ReadWrite, if the drive allows it
secret_refs[].mount_path A credential's keys, one file per key. No
configs[].mounts[] A literal file from the spec. No
/model (or model.mount_path) A model's weights. No
/workspace Your code, with astra astraeus run --code (unpacked from /astra/code.tgz.b64). Yes
/astra/job.sbatch The batch script, with astra slurm sbatch. No
/var/run/astraeus/identity svid.pem, svid-key.pem, bundle.pem, token.jwt, renewed in place; with identity: true. No
/var/run/astraeus/outputs Where $ASTRAEUS_OUTPUTS is; with outputs. Yes
/dev/shm Shared memory; see below. Yes
/etc/resolv.conf The cluster DNS, on the node network. No

Keep everything you need after the run on a drive: the container's own filesystem and EmptyDir volumes do not survive a restart.

Shared memory#

Worker /dev/shm
Without GPUs 64 MiB.
With GPUs Half its memory limit, at least 1 GiB and at most 64 GiB (8 GiB if it has no memory limit).
security_context.ipc_mode: host The machine's /dev/shm.

PyTorch data-loader workers and NCCL's shared-memory transport need the larger size; a CPU worker that needs more can use ipc_mode: host where the workspace is allowed privileged work.

GPUs#

A worker sees only the GPUs it was given, numbered from 0 inside the container; ASTRAEUS_GPU_COUNT says how many. Their ids are on the worker (assigned_gpu_ids). Astraeus does not set CUDA_VISIBLE_DEVICES: there is nothing else to hide. NVIDIA GPUs are handed over by the NVIDIA container runtime or CDI; AMD GPUs by CDI. A worker without GPUs sees none.

Limits#

Resource Limit
Memory memory_bytes (or per_gpu.memory_bytes × GPUs) is a hard limit. Above it the kernel kills a process: OOMKilled: the container exceeded its memory limit.
CPU cpu_cores is a hard quota, and the worker's weight against other workers of the machine. With cpu_burst: true it is the weight only, and the worker may use idle cores.
node_fraction workers No CPU or memory limit.
Locked memory (memlock) Unlimited for GPU workers and workers given RDMA devices (CUDA host buffers, RDMA registration).
Open files (nofile) 1 048 576 on machines using the bundled containerd runtime; the Docker daemon's default on machines using Docker.
Anything else security_context.ulimits: [{"name": "stack", "soft": 67108864, "hard": 67108864}]; -1 is unlimited. Your values replace the defaults above.
Run time time_limit_seconds.

Network#

How a worker is networked depends on its machine:

Machine's network The worker
Host network (machines with RDMA; machines installed with --network host) Shares the machine's network: its address, its ports and its resolver. Two workers on one machine cannot listen on the same port.
Node network (the installer's default without RDMA) Has an address of its own on the machine's subnet; machines are joined by WireGuard. Uses the cluster DNS.

On the node network, workers and endpoints have names in the cluster DNS, resolved within the workspace:

Name Resolves to
<run>-<rank> (or <run>-<group>-<rank>) That worker, while it runs.
<run> The run's leader while it runs; otherwise every running worker.
<endpoint> The endpoint's ready backends.

The full form is <name>.<namespace>.astraeus.local. See Names and service discovery.

Outbound traffic is open unless the run sets network.egress; see the run specification.

Signals on stop#

A worker is stopped when its run is deleted, when it is preempted, when it reaches its time limit, when another worker's failure stops its gang, or when it is stopped through the API.

  1. Its state becomes Stopping.
  2. The machine sends SIGTERM to the container's main process — or the signal the image's STOPSIGNAL names.
  3. After the machine's stop grace (30 s by default, ASTRAEUS_STOP_GRACE on the machine), it sends SIGKILL.
  4. The worker is Cancelled.

The signal goes to the main process only (PID 1 of the container). If that is a shell, start your program with exec so it receives the signal itself; see Write a checkpoint-friendly run.