Skip to content

A team login node#

On an HPC cluster, people land on a login node: a small, always-on machine where they edit scripts and submit jobs, never compute. This recipe makes one with Hesperus: a CPU-only environment on a pool of CPU machines that never stops when idle, shared with the team. Everyone connects with astra ssh and their own certificate, and sends GPU work to the cluster as runs with astra astraeus run.

Before you begin#

  • The admin role in the workspace: only admins make an environment that never stops when idle, and only an admin (or its owner) shares it.
  • A pool of machines without GPUs (or with spare cores) that the workspace may use, with a Data location. Below, <cpu-pool> stands for its name.
  • The team are members of the workspace with the editor or admin role. Below, [email protected] and [email protected] stand for them.
  • astra signed in on each person's computer (Install the CLI).

1. Make the login node#

  1. Open Hesperus → Environments and press New environment.
  2. Kind: Shell. Name: login. Idle stop: 0 (admins only: 0: never (a shared login environment)).
  3. Share with: add each person.
  4. Under Machine, keep Any machine and choose <cpu-pool> in or only machines of pool.
  5. Under Environment, choose Python + SciPy. Set GPUs to 0, CPU cores to 4, Memory to 16 GB.
  6. Press Create and start.

$ astra env create login --preset python --cpu 4 --memory 16G --pool <cpu-pool> --idle 0 \
    --share [email protected] --share [email protected]
login is pending: `astra ssh login` connects once it is ready (and waits for it)

With ASTRAEUS_TOKEN, CONSOLE and API as in API. Members are user:<id>: find each person's user_id in GET $CONSOLE/members.

$ curl -fsS -X POST "$API/notebook-runtimes" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' \
    -d '{"metadata": {"name": "login"}, "spec": {"kind": "shell", "image": "python",
         "resources": {"cpu_cores": 4, "memory_bytes": 17179869184},
         "placement": {"pool": "<cpu-pool>"}, "idle_timeout_minutes": 0,
         "members": ["user:<ana-id>", "user:<ben-id>"]}}' | jq -r .status.state
Pending

Its Idle stop reads never (stopped by a person). It counts against your limit of environments only (its owner's), and holds no GPU.

One login, one home

Everyone logs in as the environment's one user (root here), so everyone shares its home (/content/home/root) and can read every file on its drive, including each other's astra sessions. Share it only with people who may act as each other in the workspace, and keep a directory each under /content. Its sessions still record who connected: each certificate names its person.

2. Each person connects#

$ astra ssh login

They land in the environment's home, /content/home/root. The commands below run there.

astra ssh config writes hosts only for environments you own, so for people it is shared with, astra ssh login is the way in (and in an editor, see Connect an editor).

3. Install astra on the drive, once#

Anything installed outside /content is gone when the container restarts. Install astra onto the drive:

$ curl -fsSL https://console.astralyx.cloud/install-cli.sh | ASTRA_INSTALL_DIR=/content/bin sh
$ echo 'export PATH=/content/bin:$PATH' >> ~/.bashrc && . ~/.bashrc
$ astra --version

curl is one of the basic tools a runtime installs when the image lacks it; on the first start it may take a minute to appear.

4. Each person signs in, in a directory of their own#

astra keeps its session in $XDG_CONFIG_HOME/astra/config.json. Point it at a directory of your own, so people do not share one session:

$ mkdir -p /content/ana && export XDG_CONFIG_HOME=/content/ana/.config
$ astra login
Open https://console.astralyx.cloud/device?code=ABCD-EFGH and approve code ABCD-EFGH
Signed in as [email protected], working in lab/vision

Run export XDG_CONFIG_HOME=/content/ana/.config at each login (or keep it in a script of yours under /content/ana). astra logout ends the session when you leave.

5. Send GPU work to the cluster#

From the login node, runs go to the workspace's GPU machines like from anywhere else:

$ astra astraeus run --name hello-gpu --image pytorch/pytorch:2.14.1-cuda12.6-cudnn9-runtime \
    --gpus 1 --wait -- python -c "import torch; print(torch.cuda.get_device_name(0))"
hello-gpu submitted to lab-a: https://console.astralyx.cloud/o/lab/w/vision/jobs/lab-a/hello-gpu
NVIDIA H100 80GB HBM3

Slurm users can install the shims too (astra slurm install-shims /content/bin) and use sbatch and squeue as they do on a cluster (Slurm compatibility).

6. Run it#

  • Who connected: the environment's page lists its SSH sessions: who, from what, when, for how long (Sessions and audit).
  • Add or remove someone: its page → People, or astra env share login <e-mail> / astra env unshare login <e-mail>. Someone taken off is refused at their next connection.
  • Stop it when the team does not need it: it never stops by itself.