Skip to content

A notebook in 2 minutes#

In this page you create a notebook, run it on a GPU of one of your machines, install a Python package from a cell and stop the runtime so the GPU is free again. It takes two minutes of your time; the first start on a machine also downloads the image, which takes a few minutes more.

Before you begin#

  • A workspace with a cluster and at least one machine (see the Astraeus Quick start). For the GPU steps, a Linux machine with an NVIDIA or AMD GPU (GPUs). A machine without a GPU works too: skip the GPU cell.
  • A data location on that machine: the notebook's file is kept on the machine's disk. Open Machines → the machine and, under Data location, press Confirm (Drives).
  • The editor or admin role in the workspace. Viewers see notebooks but cannot open them: opening one runs code on your machines.
  • For the API tab, a personal API token (Account → API tokens). Below, <org>, <workspace> and <cluster> stand for your organisation, workspace and cluster (as in the console's address), and ast_pat_… for your token.

1. Create the notebook#

  1. Open Hesperus → Notebooks and press New notebook.
  2. Title: First steps. The Name (first-steps) and File on the drive (notebooks/first-steps.ipynb) follow from it. Drive is notebooks, made on first use.
  3. Under Image, keep PyTorch + SciPy (Jupyter) (on a machine with an AMD GPU, choose PyTorch: it has a ROCm build). The form asks for 4 CPU cores, 16 GB of memory and 1 GPU. In a workspace with no GPU machine it starts from Python + SciPy and no GPU instead.
  4. Under Machine, keep Any machine, or pick yours: each machine shows its free cores, memory and GPUs, and whether the runtime fits now.
  5. Press Create and open. The notebook's page opens.

astra has no notebook commands: create and open notebooks in the console, or with the API.

$ export ASTRAEUS_TOKEN=ast_pat_…
$ export API="https://console.astralyx.cloud/api/v1/orgs/<org>/workspaces/<workspace>/clusters/<cluster>/api"
$ curl -fsS -X POST "$API/notebooks" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' \
    -d '{"metadata": {"name": "first-steps"},
         "spec": {"title": "First steps",
                  "runtime_defaults": {"image": "pytorch-cuda",
                    "resources": {"cpu_cores": 4, "memory_bytes": 17179869184, "gpus": {"count": 1}}}}}' \
    | jq -r '.spec.drive + ":" + .spec.path'
notebooks:notebooks/first-steps.ipynb

The drive and the file default to the workspace's notebooks drive and notebooks/<name>.ipynb.

2. Wait for the runtime#

Opening the notebook finds your runtime for it or starts one. It goes Pending (waiting for a machine) → Starting (on a machine; the first time there, the image is pulled) → Ready.

The top of the editor shows Waiting for a machine…, then Starting on gpu-01…, then the machine and its GPU. If it waits, Why is it waiting? says for what: a free GPU, a machine with a data location, a machine of the pool. See Resources and machines.

The runtime is an Astraeus run like any other:

$ astra astraeus runs
NAME                          CLUSTER       STATE       REASON
first-steps-3fa91c-1          <cluster>     Running     All workers running

$ curl -fsS -X POST "$API/notebooks/first-steps/open" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    | jq '{runtime: .runtime.metadata.name, state: .runtime.status.state, reused}'
{
  "runtime": "first-steps-3fa91c",
  "state": "Pending",
  "reused": false
}
$ curl -fsS "$API/notebook-runtimes/first-steps-3fa91c" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    | jq '{state: .status.state, reason: .status.reason}'
{
  "state": "Ready",
  "reason": "Ready on gpu-01 (1 × NVIDIA GeForce RTX 4090)"
}

The runtime's name is the notebook's and a random suffix. Opening the notebook again while this runtime is active returns it ("reused": true).

3. Run a cell on the GPU#

In the editor, type in the first cell and press Shift+Enter:

import torch
print(torch.cuda.is_available(), torch.cuda.get_device_name(0))
x = torch.ones(4096, 4096, device="cuda")
print((x @ x).sum().item())

The output names your GPU (here an RTX 4090; on an AMD GPU, torch.cuda is ROCm's and works the same):

True NVIDIA GeForce RTX 4090
68719476736.0

!nvidia-smi in a cell shows the GPU and the kernel's process: a runtime with GPUs always has nvidia-smi on an NVIDIA machine.

4. Install a package from a cell#

%pip install --quiet tabulate
from tabulate import tabulate
print(tabulate([["gpu", torch.cuda.device_count()]], headers=["resource", "count"]))
resource      count
----------  -------
gpu               1

No kernel restart is needed. The package is installed on the notebook's drive, under /content/.hesperus/python, so the next runtime on this machine has it too. See Install packages.

5. Stop the runtime#

Your cells and outputs are saved to notebooks/first-steps.ipynb on the drive as you work.

Open the runtime's menu (the machine and GPU at the top of the editor) and choose Stop the runtime (frees the GPU). Left alone, it stops by itself after 30 minutes without activity.

astra env stop stops any runtime by its name:

$ astra env stop first-steps-3fa91c
first-steps-3fa91c is stopping

$ curl -fsS -X POST "$API/notebook-runtimes/first-steps-3fa91c/stop" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    | jq -r .status.state
Stopping

Opening the notebook again starts a runtime again, on the same drive: the file and the installed packages are there; variables in the kernel's memory are not — run the cells again.

Next#