Skip to content

Troubleshooting#

Find the symptom, read the cause, apply the fix. Most answers start from the runtime's state and its reason, shown beside it in Hesperus → Runtimes, on an environment's page, and by astra env get <name>:

$ astra env get dev
dev: Starting — Pulling pytorch/pytorch:2.14.1-cuda12.6-cudnn9-runtime@sha256:1262abc856a428933c0477b4be87079173a0b087b7691c47fcb396b7b5775c56 (gpu-01)
machine: gpu-01
image:   pytorch/pytorch:2.14.1-cuda12.6-cudnn9-runtime@sha256:1262abc856a428933c0477b4be87079173a0b087b7691c47fcb396b7b5775c56
drive:   dev-home (home /content/home/root)
idle:    stopped after 60 min with nobody connected
connect: astra ssh dev   (or `astra ssh config`, then VS Code / Cursor: Remote-SSH)
owner:   [email protected]

A runtime's run is <runtime>-<n> (n counts its starts) and its worker <runtime>-<n>-0: astra astraeus status dev-1 and astra astraeus logs dev-1-0 show them as for any run. In a notebook, the runtime's menu has Its worker (logs).

A runtime stays Pending#

Pending means no machine has taken it yet. Nothing you ask for is refused because no machine has it now: the runtime waits until one fits, or until a machine that has it joins the workspace.

Press Why? beside its state (on an environment's page, or Why is it waiting? in a notebook's editor). It opens the same diagnosis as for a run: what it waits for, what each machine lacks, and what to change.

What Why? says Cause Fix
No machine with that many free GPUs, of that model or memory The GPUs are held by other work, or no machine has them. Wait, ask for fewer or any model (GPU model empty, no GPU memory, at least), or choose Any machine. The form's machine list shows what each machine has free.
Choose where to keep data on gpu-01 A notebook's drive and an environment's home are kept on each machine's data location, and this machine has none. An organisation admin chooses one: Machines → the machine → Data location → Confirm (Drives).
Waits for the workspace's quota, or behind other work The workspace's share of the cluster is in use. Stop a runtime you do not need; see The queue.
The machine it is pinned to is down or cordoned Machine (--machine) names one machine only. Start that machine, or choose Any machine or a pool.

A runtime fails at once#

Reason Cause Fix
Its run could not be made: RAPIDS has no build for AMD GPUs, and the machines it may run on have AMD's: choose PyTorch or Hugging Face or vLLM, or an image of your own A preset built for NVIDIA only (CUDA dev, NGC PyTorch, JAX, TensorFlow, RAPIDS…) where the machines' GPUs are AMD's. Choose a preset with an AMD build, or an image of your own.
ErrImagePull in the reason An image of your own that does not exist, or is private. Check the reference. A runtime cannot name a registry credential: use an image that pulls anonymously, or one already on the machine (Presets and setup).
Jupyter Server exited, or Its run failed The server inside stopped. Read the worker's log (above). For an image of your own in a notebook: it needs Python with pip, or Jupyter Server and ipykernel.

A runtime stays Starting#

Reason Cause Fix
Pulling <image> (gpu-01) The first start on a machine downloads the image: from 381 MB (python-3.12) to 20.5 GB (pytorch on AMD). Wait; later starts on that machine reuse it. Presets lists each download.
… runs on gpu-01, but its agent is too old to say it is ready: update the agent (Machines → gpu-01 → Update) The machine's agent predates runtimes' readiness. Update the machine (Update, drain and remove).
An IDE environment, for a minute more The first IDE on a machine downloads it once (about 220 MB). Wait.

astra ssh and Open IDE wait up to 15 minutes, then say dev was not ready in 15 minutes; run them again.

nvidia-smi is missing#

Cause Fix
The runtime asked for no GPU (GPUs 0). A runtime has GPUs only if it asks for them. Ask for one: Change runtime… for a notebook; for an environment, delete it and create it again with the same name and --gpus 1 (Change an environment).
It runs on an AMD GPU. nvidia-smi is NVIDIA's; the AMD builds (ROCm) do not have it.

On NVIDIA machines, a runtime with GPUs gets nvidia-smi and NVML whatever its image says. A preset built for GPUs runs on the CPU without one; a preset not built for GPUs (Python, Spark) gets the GPU, but its libraries do not use it.

The IDE fails on an Alpine image#

The run's log says: this image is built on musl (Alpine), and code-server's Node.js needs glibc: use an image built on Debian, Ubuntu, RHEL or another glibc distribution (glibc 2.28 or newer, with libstdc++). SSH works on any image: use kind shell for this one.

Make the environment from a glibc image (any preset, ubuntu:24.04, python:3.12-slim), or keep the Alpine image as a Shell environment:

$ astra env create tiny --image alpine:3.20 --cpu 1 --memory 1G
tiny is pending: `astra ssh tiny` connects once it is ready (and waits for it)

Other IDE failures in the log:

  • this machine's agent mounted no IDE (/.astraeus/code-server/lib/node is missing): update the agent (Machines → the machine → Update).
  • code-server's Node.js cannot run in this image (it needs glibc 2.28 or newer and libstdc++): an older or minimal glibc image; use a newer one.

An extension is not found in the IDE#

The IDE's extensions come from Open VSX; the Microsoft Marketplace is not available. Microsoft's own extensions that are only there — Pylance, the Remote extensions, C# Dev Kit — cannot be installed. Use an Open VSX alternative (basedpyright for Python), or connect your desktop VS Code over SSH instead (Connect VS Code, Cursor or JetBrains Gateway), which uses the Marketplace from your computer.

The editor cannot install its server#

VS Code, Cursor and JetBrains Gateway install their server inside the container on first connect, downloading it from their vendor (update.code.visualstudio.com for VS Code). The container must reach the internet over HTTPS. If the machine's network blocks it, the connection hangs or fails while installing. Let the machine reach those hosts, or use the IDE in the browser, which the machine brings. In an environment, the server is kept in your home on the drive: installed once.

Refused: SSH_FORBIDDEN#

runtime dev is user:…'s: only its owner, the people it is shared with and the workspace's admins may connect to it. SSH, the IDE and apps are given to the runtime's owner, the people it is shared with and the workspace's admins — not to every editor. Ask the owner to share it. Viewers are refused before that, whatever the runtime.

SSH: no SSH, or the login fails#

Symptom Cause Fix
dev has no SSH (turn it on, or use an environment: astra env create); 409 NOTEBOOK_CONFLICT … has no SSH (turn it on: ssh.enabled) A notebook's runtime without SSH, or an IDE environment made with --no-ssh. Turn on SSH into the runtime too (SSH into a notebook's runtime); or use the IDE.
The runtime fails with user dev does not exist in this image: set the runtime's ssh.user to one of its users Log in as (--user) names a user the image does not have. Create it again with a user of the image, or none (root).
ssh: … (is OpenSSH's client installed?) astra ssh runs your system's ssh. Install OpenSSH's client.
WARNING: REMOTE HOST IDENTIFICATION HAS CHANGED The runtime's host key is kept on its drive; on another machine (another copy of the drive), or after it was made again, the key is new. Forget the old key: ssh-keygen -R <host> -f ~/.config/astra/ssh/known_hosts.
The connection drops when you are taken off Sharing is checked on each connection. Ask the owner to share it again.

Certificates last 8 hours and astra renews them itself (when less than 10 minutes are left). If one seems stuck, delete the runtime's folder under ~/.config/astra/ssh/ (named after the host) and connect again: a new key and certificate are made.

Refused: RUNTIME_NOT_RUNNING#

dev is stopped: start it first, then open it once it is ready. An app or the IDE is opened only while the runtime is on a machine. Start it (or Open IDE, which starts it), then open the app again.

An app does not open#

Symptom Cause Fix
502 the run does not answer on port 8501 Nothing listens on that port in the container (yet). Start the server; check the port it listens on, and its log.
Streamlit loads, but uploads fail with 403 Streamlit's XSRF check needs a cookie, and no cookie reaches an app. Start it with --server.enableXsrfProtection false.
400 INVALID_NOTEBOOK port 8822 is the IDE's own (Open IDE), port 8888 is Jupyter's own (open the notebook), port 2222 is the runtime's SSH server's These ports are the runtime's own. Use Open IDE, the notebook, or SSH; run your app on another port.
The address asks you to sign in again Its sign-in lasts 8 hours. Open it again from the console or astra env port.

Refused: HESPERUS_LIMIT#

this workspace allows 2 environments running at once per person, and its owner has 2 (dev, code): stop one first — or the same for GPUs. The workspace's admins set limits per person; the message names the owner's runtimes to stop. See Set limits per person.

Refused: ADMIN_ONLY#

  • idle_timeout_minutes 0 (never stopped when idle) is for the workspace's admins: a shared environment that stays up. Choose 5 to 1440 minutes.
  • only the workspace's admins make a runtime for someone else.

The runtime stopped while I was working#

It was idle for its idle timeout (30 minutes for a notebook's runtime, 60 for an environment, unless set): nobody connected, no traffic. A notebook open in the console counts while you run cells and outputs are saved; a long cell that prints nothing, or runs on with the tab closed, does not. Print progress, keep an SSH session open, or ask for a longer timeout (5 to 1440 minutes: Idle stop or --idle for an environment; idle_timeout_minutes when opening a notebook through the API). See Idle stop.

My files or packages are gone after a restart#

Cause Fix
They were written outside /content. A runtime is a new container each start; only its drives are kept. Keep work under /content.
Python packages installed without --user in an environment. pip install --user, or the setup's pip packages (Install packages).
It started on another machine. A notebook's drive and an environment's home are kept on each machine: one copy per machine. Pin it to one Machine; or put shared files on a drive every machine reaches.

The setup failed#

The setup's output is in the run's log and in /content/.hesperus/env/<runtime>/setup.log on the drive. A failure never stops the runtime: it starts, and the setup is tried again at the next start.

$ astra ssh dev -- cat /content/.hesperus/env/dev/setup.log
Message (hesperus-env: …) Cause Fix
the container does not run as root: … not installed System packages need root; a notebook's runtime with an image of your own runs as the image's user. Use a preset, or an image that runs as root, or install in the image.
no package manager (apt, dnf, apk) in this image: … not installed A minimal or distroless image. Use an image with one.
no Python in this image: … skipped, no pip for … in this image pip packages need Python and pip. Choose an image with Python.
Jupyter Server could not be installed (see …): the image needs Python and pip, or Jupyter Server itself A notebook's runtime on an image of your own without Jupyter. As it says.
this Python has no user site (a virtual environment): installed into it, not onto the drive (again at the next start) The image's Python is a virtual environment. Works, but is installed at each start.
setup: failed (see …); it is tried again at the next start A requirement or the script failed. Read the log; fix the setup (for an environment, create it again with the same name).
basic tools missing here ( … ) cannot be installed (not root, or no package manager) As it says. Turn Basic tools off (--no-tools), or install them in the image.