Inside a worker#
Each worker is a container on one machine. This page describes what your program finds when it starts there: the environment variables Astraeus sets, where drives, credentials and files are mounted, which GPUs it sees, which user it runs as, its limits, how it reaches other workers, and what happens when it is stopped. Use it when you write a training script or an entrypoint for Astraeus.
A quick look#
Start a run that prints its environment, then read its log:
$ astra astraeus run --name whoami --image busybox --cpus 1 --mem 256M -- sh -c 'env | grep -E "^(ASTRAEUS|SLURM|RANK|WORLD|MASTER)" | sort; id; df -h /dev/shm'
whoami submitted to gpu-east: https://console.astralyx.cloud/o/acme/w/vision/jobs/gpu-east/whoami
$ astra astraeus logs whoami
ASTRAEUS_GROUP_NAME=default
ASTRAEUS_GROUP_REPLICA_INDEX=0
ASTRAEUS_JOB_NAME=ws-7f3c2a.whoami
ASTRAEUS_LEADER_ADDRESS=10.0.4.17
ASTRAEUS_LEADER_TASK_NAME=ws-7f3c2a.whoami-0
ASTRAEUS_NODE_NAME=gpu-09
ASTRAEUS_REPLICA_INDEX=0
ASTRAEUS_TASK_NAME=ws-7f3c2a.whoami-0
ASTRAEUS_WORLD_SIZE=1
MASTER_ADDR=10.0.4.17
MASTER_PORT=29500
RANK=0
SLURMD_NODENAME=gpu-09
SLURM_CLUSTER_NAME=astraeus
SLURM_CPUS_PER_TASK=1
SLURM_JOBID=whoami
SLURM_JOB_ID=whoami
SLURM_JOB_NAME=whoami
SLURM_JOB_NUM_NODES=1
SLURM_LOCALID=0
SLURM_MEM_PER_NODE=256
SLURM_NNODES=1
SLURM_NODEID=0
SLURM_NPROCS=1
SLURM_NTASKS=1
SLURM_NTASKS_PER_NODE=1
SLURM_PROCID=0
WORLD_SIZE=1
uid=0(root) gid=0(root) groups=0(root),10(wheel)
Filesystem Size Used Available Use% Mounted on
shm 64.0M 0 64.0M 0% /dev/shm
In a workspace, full names carry the workspace's namespace (ws-7f3c2a. above): ASTRAEUS_TASK_NAME, ASTRAEUS_JOB_NAME and ASTRAEUS_LEADER_TASK_NAME are full names; the SLURM_JOB_* variables hold the run's own name.
Environment variables#
Who and where#
| Variable | Value | Set |
|---|---|---|
ASTRAEUS_TASK_NAME |
The worker's full name (<namespace>.<run>-<rank> in a workspace). |
Always. |
ASTRAEUS_JOB_NAME |
The run's full name. | Always for a run's worker. |
ASTRAEUS_NODE_NAME |
The machine's name. | Always. |
ASTRAEUS_REPLICA_INDEX |
The worker's rank, from 0, across all groups. | Always. |
ASTRAEUS_GROUP_NAME |
Its worker group; default without task_groups. |
Always for a run's worker. |
ASTRAEUS_GROUP_REPLICA_INDEX |
Its index within its group. | With a group. |
RANK |
Same as ASTRAEUS_REPLICA_INDEX. |
When the run's size is known (always for a run). |
WORLD_SIZE, ASTRAEUS_WORLD_SIZE |
How many workers the run has. | Same. |
Rendezvous#
| Variable | Value | Set |
|---|---|---|
ASTRAEUS_LEADER_TASK_NAME |
The leader's full name (rank 0 of the leader group). | Always for a run's worker. |
ASTRAEUS_LEADER_ADDRESS |
The address of the machine the leader runs on. | Once the leader is placed. |
MASTER_ADDR |
Same as ASTRAEUS_LEADER_ADDRESS. |
Same. |
MASTER_PORT |
29500. |
Same. |
MASTER_ADDR is a machine address. When workers have addresses of their own (the node network), point your rendezvous at the leader's name instead; see Multi-machine runs.
GPUs and the network#
| Variable | Value | Set |
|---|---|---|
ASTRAEUS_GPU_COUNT |
GPUs given to this worker. | With GPUs. |
NCCL_DEBUG |
WARN |
With GPUs. |
NCCL_ASYNC_ERROR_HANDLING |
1 |
With GPUs. |
NCCL_SOCKET_IFNAME |
^lo,docker,virbr,veth,cni,wg |
With GPUs. |
NCCL_IB_HCA |
The RDMA ports that are up and fast enough: =mlx5_0:1,=mlx5_2:1. |
Placed on an InfiniBand fabric or RDMA, or asking for RDMA, on a machine with RDMA ports up. |
NCCL_IB_DISABLE |
1 |
Placed across machines over Ethernet. |
NVIDIA_IMEX_CHANNELS |
0 |
Placed in a multi-node NVLink domain, on a machine whose GPUs are handed over by the NVIDIA runtime hook. |
Arrays#
| Variable | Value |
|---|---|
ASTRAEUS_ARRAY_TASK_ID |
This copy's index, 0 to size − 1. |
ASTRAEUS_ARRAY_SIZE |
array.size. |
Slurm's names#
Set for every worker, so scripts and libraries written for Slurm read the same facts. One worker per machine is Slurm's shape here.
| Variable | Value |
|---|---|
SLURM_JOB_ID, SLURM_JOBID, SLURM_JOB_NAME |
The run's name (without the namespace). |
SLURM_PROCID, SLURM_NODEID |
The rank. |
SLURM_LOCALID |
0. |
SLURM_NTASKS, SLURM_NPROCS, SLURM_NNODES, SLURM_JOB_NUM_NODES |
The run's number of workers. |
SLURM_NTASKS_PER_NODE |
1. |
SLURMD_NODENAME |
The machine. |
SLURM_CLUSTER_NAME |
astraeus. |
SLURM_CPUS_PER_TASK |
The worker's cores (when it has a core count). |
SLURM_MEM_PER_NODE |
The worker's memory, MiB (when it has a memory size). |
SLURM_GPUS_ON_NODE, SLURM_GPUS_PER_TASK |
The worker's GPUs (with GPUs). |
SLURM_ARRAY_TASK_ID, SLURM_NODELIST and SLURM_STEP_* are not set. See Slurm compatibility.
Features you turn on#
| Variable | Value | Set when |
|---|---|---|
ASTRAEUS_MODEL_PATH |
Where the model's weights are (/model by default). |
spec.model. |
ASTRAEUS_MODEL_FILE |
The model's single GGUF file. | spec.model, for a GGUF model. |
OPENAI_BASE_URL, OPENAI_MODEL |
The engine serving the model beside the run, and the model's name there. | spec.model.engine, in the main group. |
ASTRAEUS_IDENTITY_DIR |
/var/run/astraeus/identity. |
identity: true. |
ASTRAEUS_OUTPUTS |
/var/run/astraeus/outputs/outputs, the file to write key=value outputs to. |
outputs. |
ASTRAEUS_TOOL_GATEWAY |
The machine's tool gateway, http://<address>:8801. |
tools, on the node network. |
<port_env_name> |
The port allocated to an external access. | external_accesses[].port_env_name. |
HESPERUS_TOKEN (or interactive.token_env) |
The session token of an interactive server. | interactive. |
HOME |
The user's home from the image's /etc/passwd, or /root. |
When the image does not set it (containerd runtime). |
Which value wins#
From lowest to highest: the variables above that Astraeus injects; external-access port variables; your env; credentials' env_vars; ASTRAEUS_IDENTITY_DIR and ASTRAEUS_OUTPUTS. The network variables (NCCL_IB_HCA, NCCL_IB_DISABLE) and ASTRAEUS_TOOL_GATEWAY are only set when you have not set them. So "env": {"MASTER_PORT": "1234"} replaces the default port, and "env": {"NCCL_DEBUG": "INFO"} turns on NCCL's logging.
Command and working directory#
The container runs command with args; command replaces the image's ENTRYPOINT and CMD. Without command, the image's ENTRYPOINT runs with args, or with its CMD when args is empty. There is no shell unless you run one. The working directory is the image's WORKDIR (/workspace with astra astraeus run --code). See Command, arguments and environment.
User#
The worker runs as task_template.user (name, uid or uid:gid), else the image's USER, else root (uid 0). Files the worker writes to a drive belong to that uid.
Filesystem#
| Path | What | Writable |
|---|---|---|
/ |
The image. Changes are lost when the worker stops or restarts. | Yes |
volumes[].mount_path (EmptyDir) |
An empty directory on the machine, removed with the worker's container. | Yes |
volumes[].mount_path (Bind) |
A path of the machine (in a workspace, only granted paths). | With mode: ReadWrite |
datavolume_refs[].mount_path |
A drive (the console's default is /data/<drive>). |
With mode: ReadWrite, if the drive allows it |
secret_refs[].mount_path |
A credential's keys, one file per key. | No |
configs[].mounts[] |
A literal file from the spec. | No |
/model (or model.mount_path) |
A model's weights. | No |
/workspace |
Your code, with astra astraeus run --code (unpacked from /astra/code.tgz.b64). |
Yes |
/astra/job.sbatch |
The batch script, with astra slurm sbatch. |
No |
/var/run/astraeus/identity |
svid.pem, svid-key.pem, bundle.pem, token.jwt, renewed in place; with identity: true. |
No |
/var/run/astraeus/outputs |
Where $ASTRAEUS_OUTPUTS is; with outputs. |
Yes |
/dev/shm |
Shared memory; see below. | Yes |
/etc/resolv.conf |
The cluster DNS, on the node network. | No |
Keep everything you need after the run on a drive: the container's own filesystem and EmptyDir volumes do not survive a restart.
Shared memory#
| Worker | /dev/shm |
|---|---|
| Without GPUs | 64 MiB. |
| With GPUs | Half its memory limit, at least 1 GiB and at most 64 GiB (8 GiB if it has no memory limit). |
security_context.ipc_mode: host |
The machine's /dev/shm. |
PyTorch data-loader workers and NCCL's shared-memory transport need the larger size; a CPU worker that needs more can use ipc_mode: host where the workspace is allowed privileged work.
GPUs#
A worker sees only the GPUs it was given, numbered from 0 inside the container; ASTRAEUS_GPU_COUNT says how many. Their ids are on the worker (assigned_gpu_ids). Astraeus does not set CUDA_VISIBLE_DEVICES: there is nothing else to hide. NVIDIA GPUs are handed over by the NVIDIA container runtime or CDI; AMD GPUs by CDI. A worker without GPUs sees none.
Limits#
| Resource | Limit |
|---|---|
| Memory | memory_bytes (or per_gpu.memory_bytes × GPUs) is a hard limit. Above it the kernel kills a process: OOMKilled: the container exceeded its memory limit. |
| CPU | cpu_cores is a hard quota, and the worker's weight against other workers of the machine. With cpu_burst: true it is the weight only, and the worker may use idle cores. |
node_fraction workers |
No CPU or memory limit. |
Locked memory (memlock) |
Unlimited for GPU workers and workers given RDMA devices (CUDA host buffers, RDMA registration). |
Open files (nofile) |
1 048 576 on machines using the bundled containerd runtime; the Docker daemon's default on machines using Docker. |
| Anything else | security_context.ulimits: [{"name": "stack", "soft": 67108864, "hard": 67108864}]; -1 is unlimited. Your values replace the defaults above. |
| Run time | time_limit_seconds. |
Network#
How a worker is networked depends on its machine:
| Machine's network | The worker |
|---|---|
Host network (machines with RDMA; machines installed with --network host) |
Shares the machine's network: its address, its ports and its resolver. Two workers on one machine cannot listen on the same port. |
| Node network (the installer's default without RDMA) | Has an address of its own on the machine's subnet; machines are joined by WireGuard. Uses the cluster DNS. |
On the node network, workers and endpoints have names in the cluster DNS, resolved within the workspace:
| Name | Resolves to |
|---|---|
<run>-<rank> (or <run>-<group>-<rank>) |
That worker, while it runs. |
<run> |
The run's leader while it runs; otherwise every running worker. |
<endpoint> |
The endpoint's ready backends. |
The full form is <name>.<namespace>.astraeus.local. See Names and service discovery.
Outbound traffic is open unless the run sets network.egress; see the run specification.
Signals on stop#
A worker is stopped when its run is deleted, when it is preempted, when it reaches its time limit, when another worker's failure stops its gang, or when it is stopped through the API.
- Its state becomes
Stopping. - The machine sends
SIGTERMto the container's main process — or the signal the image'sSTOPSIGNALnames. - After the machine's stop grace (30 s by default,
ASTRAEUS_STOP_GRACEon the machine), it sendsSIGKILL. - The worker is
Cancelled.
The signal goes to the main process only (PID 1 of the container). If that is a shell, start your program with exec so it receives the signal itself; see Write a checkpoint-friendly run.