A team login node#
On an HPC cluster, people land on a login node: a small, always-on
machine where they edit scripts and submit jobs, never compute. This
recipe makes one with Hesperus: a CPU-only environment on a pool of CPU
machines that never stops when idle, shared with the team. Everyone
connects with astra ssh and their own certificate, and sends GPU work to
the cluster as runs with astra astraeus run.
Before you begin#
- The admin role in the workspace: only admins make an environment that never stops when idle, and only an admin (or its owner) shares it.
- A pool of machines without GPUs (or with spare cores) that the workspace
may use, with a Data location. Below,
<cpu-pool>stands for its name. - The team are members of the workspace with the editor or admin
role. Below,
[email protected]and[email protected]stand for them. astrasigned in on each person's computer (Install the CLI).
1. Make the login node#
- Open Hesperus → Environments and press New environment.
- Kind: Shell. Name:
login. Idle stop:0(admins only: 0: never (a shared login environment)). - Share with: add each person.
- Under Machine, keep Any machine and choose
<cpu-pool>in or only machines of pool. - Under Environment, choose Python + SciPy. Set GPUs to
0, CPU cores to4, Memory to16GB. - Press Create and start.
$ astra env create login --preset python --cpu 4 --memory 16G --pool <cpu-pool> --idle 0 \
--share [email protected] --share [email protected]
login is pending: `astra ssh login` connects once it is ready (and waits for it)
With ASTRAEUS_TOKEN, CONSOLE and API as in
API. Members are user:<id>: find
each person's user_id in GET $CONSOLE/members.
$ curl -fsS -X POST "$API/notebook-runtimes" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
-H 'content-type: application/json' \
-d '{"metadata": {"name": "login"}, "spec": {"kind": "shell", "image": "python",
"resources": {"cpu_cores": 4, "memory_bytes": 17179869184},
"placement": {"pool": "<cpu-pool>"}, "idle_timeout_minutes": 0,
"members": ["user:<ana-id>", "user:<ben-id>"]}}' | jq -r .status.state
Pending
Its Idle stop reads never (stopped by a person). It counts against your limit of environments only (its owner's), and holds no GPU.
One login, one home
Everyone logs in as the environment's one user (root here), so
everyone shares its home (/content/home/root) and can read every
file on its drive, including each other's astra sessions. Share it
only with people who may act as each other in the workspace, and keep
a directory each under /content. Its sessions still record who
connected: each certificate names its person.
2. Each person connects#
They land in the environment's home, /content/home/root. The commands
below run there.
astra ssh config writes hosts only for environments you own, so for
people it is shared with, astra ssh login is the way in (and in an
editor, see Connect an editor).
3. Install astra on the drive, once#
Anything installed outside /content is gone when the container restarts.
Install astra onto the drive:
$ curl -fsSL https://console.astralyx.cloud/install-cli.sh | ASTRA_INSTALL_DIR=/content/bin sh
$ echo 'export PATH=/content/bin:$PATH' >> ~/.bashrc && . ~/.bashrc
$ astra --version
curl is one of the basic tools a runtime installs when the image lacks
it; on the first start it may take a minute to appear.
4. Each person signs in, in a directory of their own#
astra keeps its session in $XDG_CONFIG_HOME/astra/config.json. Point
it at a directory of your own, so people do not share one session:
$ mkdir -p /content/ana && export XDG_CONFIG_HOME=/content/ana/.config
$ astra login
Open https://console.astralyx.cloud/device?code=ABCD-EFGH and approve code ABCD-EFGH
Signed in as [email protected], working in lab/vision
Run export XDG_CONFIG_HOME=/content/ana/.config at each login (or keep it
in a script of yours under /content/ana). astra logout ends the
session when you leave.
5. Send GPU work to the cluster#
From the login node, runs go to the workspace's GPU machines like from anywhere else:
$ astra astraeus run --name hello-gpu --image pytorch/pytorch:2.14.1-cuda12.6-cudnn9-runtime \
--gpus 1 --wait -- python -c "import torch; print(torch.cuda.get_device_name(0))"
hello-gpu submitted to lab-a: https://console.astralyx.cloud/o/lab/w/vision/jobs/lab-a/hello-gpu
NVIDIA H100 80GB HBM3
Slurm users can install the shims too (astra slurm install-shims
/content/bin) and use sbatch and squeue as they do on a cluster
(Slurm compatibility).
6. Run it#
- Who connected: the environment's page lists its SSH sessions: who, from what, when, for how long (Sessions and audit).
- Add or remove someone: its page → People, or
astra env share login <e-mail>/astra env unshare login <e-mail>. Someone taken off is refused at their next connection. - Stop it when the team does not need it: it never stops by itself.