Skip to content

External access#

External access publishes a service outside the cluster, on a port in the range 30000–32767 that is the same on every machine. Clients outside connect to that port on one of your ingress machines, where Envoy balances across the ready workers, or directly on the machine that runs the worker.

Use external access to:

  • open a notebook, a dashboard or a TensorBoard running in a run from your laptop;
  • serve an HTTP or gRPC API from a run or a replica group to clients outside the cluster;
  • put your own load balancer, TLS terminator or API gateway in front of a service that runs on your machines.

Three things publish a port:

What How Selects
A run's external_accesses Declared on the run's worker template The first worker (replica 0) of the worker group that declares it
A balanced endpoint Created by you, mode: balanced Every ready worker matching its selector
A replica group's load balancer load_balancer.enabled: true Every ready member of the group

All three are balanced endpoints in the end: a run's external access becomes an endpoint named <run>-<access>, owned by the run and deleted with it.

How traffic flows#

flowchart LR
    C[Client] -->|"ingress machine :30080"| E["Envoy (ingress machine)"]
    E -->|"mesh, target port"| W1["Worker on gpu-01"]
    E -->|"mesh, target port"| W2["Worker on gpu-02"]
    C -.->|"nodeport: gpu-01 :30080"| W1
    E -.->|"no ready worker, can wake"| A["Activator (same machine)"]
    A -.->|"wakes the replica group, then splices"| W1
  • Each balanced endpoint gets one listen port, chosen once by Astraeus and kept for the endpoint's life. It is the same on every machine.
  • With exposure: ingress (the default), Envoy on every ingress machine of your organisation listens on that port and sends each connection or request to a ready worker, round robin, at the worker's target port.
  • With exposure: nodeport, the machine that runs each selected worker publishes the port and forwards it to the container. With both, both happen.
  • When an endpoint can wake a replica group scaled to zero and no worker is ready, Envoy hands connections to the activator on the same machine. See Replica groups and scale to zero.

Before you begin#

  • You need the admin or editor role in the workspace.
  • For exposure: ingress, you need at least one ingress machine: a machine of the cluster (it runs the worker and joins the mesh, so it reaches the workers' addresses) that also runs the agent's edge part and Envoy. See Set up an ingress machine.
  • Your firewall must let clients reach TCP 30000–32767 on the ingress machines (or, for nodeport, on the machines that run the workers). Open only what you publish.
  • For the API examples, set TOKEN and API as described in Drives.

No CLI commands for external access

The astra CLI has no commands for external access. Declare it in the run (console New run → Edit as JSON, or the API).

Set up an ingress machine#

An organisation admin does this once per ingress machine.

  1. Install the machine with the agent's edge part among its parts:

    $ curl -fsSL <installer URL> | sudo sh -s -- \
        --apiserver https://api.astralyx.cloud --token-file ./token \
        --agents drives,credentials,data,edge
    

    --apiserver is the Astraeus address shown in the console when you add a machine; the join token comes from the same place (Install a machine). The edge part (the service astraeus-agent-edge) connects out to Astraeus over HTTPS as the machine, like the agent's other parts. It writes Envoy's configuration to /etc/astraeus/envoy/xds/: a bootstrap envoy.json, and cds.json and lds.json, which it rewrites whenever endpoints or their workers change.

  2. Install Envoy on the machine. The installer does not install it.

  3. Run Envoy with the bootstrap the agent wrote, for example as a systemd service:

    $ envoy -c /etc/astraeus/envoy/xds/envoy.json
    

    Envoy reads its listeners and clusters from the files and reloads them when they change. Its admin interface listens on 127.0.0.1:9901, where the agent reads its traffic counters.

  4. Check that a balanced endpoint answers on the machine:

    $ curl -s http://<ingress machine address>:30080/healthz
    ok
    

The edge part's settings, in /etc/astraeus/agent.env:

Variable Default Description
ASTRAEUS_ENVOY_XDS_DIR /etc/astraeus/envoy/xds Where Envoy's configuration is written.
ASTRAEUS_ENVOY_ADMIN_PORT 9901 Envoy's admin port, on loopback, written into the bootstrap.
ASTRAEUS_INGRESS_LISTEN_ADDRESS 0.0.0.0 The address Envoy's listeners bind.
ASTRAEUS_ACTIVATOR_PORT 9902 The activator's loopback port. 0 turns scale-to-zero wake-ups off on this machine.
ASTRAEUS_ACTIVATOR_TIMEOUT 2m How long the activator holds a connection while a replica group wakes.

An ingress machine that also runs workers

If the machine also runs workers whose endpoints use nodeport or both, Envoy and the worker would both try to open the same ports. Give each its own address: ASTRAEUS_INGRESS_LISTEN_ADDRESS for Envoy, and ASTRAEUS_NODE_PORT_ADDRESS for the worker.

Publish a run's port#

Declare the access on the worker template. This publishes a Jupyter server on port 8888:

run.json
{
  "metadata": {"name": "notebook"},
  "spec": {
    "lifetime": "Service",
    "task_template": {
      "image": "quay.io/jupyter/pytorch-notebook:latest",
      "requested_resources": {"cpu_cores": 8, "memory_bytes": 34359738368, "gpu_requests": {"count": 1}},
      "external_accesses": [
        {"name": "jupyter", "target_port": 8888, "enabled": true, "mode": "ingress"}
      ]
    }
  }
}
  1. Open New run, fill in the form, and click Edit as JSON.
  2. Add external_accesses to spec.task_template as above, and submit the run.
  3. Open Resources → Endpoints: the endpoint notebook-jupyter appears. Click it to see its runtime.listen_port.
$ curl -sS -X POST "$API/runs" \
    -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
    -d @run.json

Then ask where to connect:

$ curl -sS "$API/runs/notebook/external-accesses" -H "Authorization: Bearer $TOKEN"
{
  "items": [
    {
      "name": "jupyter",
      "protocol": "tcp",
      "external_host": "",
      "external_port": 30000,
      "target_port": 8888,
      "leader_task": "notebook-0",
      "backend": "10.240.3.7:8888",
      "status": "Active",
      "mode": "ingress"
    }
  ]
}

Connect to http://<ingress machine address>:30000. status is Pending while no port is assigned yet (no port assigned yet), the worker is not running (the worker is not running) or not ready (the worker is not ready). For nodeport and both, the answer also carries node_name, node_ip and node_port: the machine to connect to.

Inside the container, the published port is in the environment variable ASTRAEUS_EXT_<NAME>_PORT (here ASTRAEUS_EXT_JUPYTER_PORT=30000), or the name you set in port_env_name. The container waits to start until the port is assigned.

External access fields#

Field Type Default Description
name string required Unique in the worker template. The endpoint is <run>-<name>.
target_port integer required The container's port, 1–65535.
enabled boolean false Only enabled accesses are published.
external_port integer assigned A specific port in 30000–32767. Without it, the lowest free port is assigned and kept.
mode ingress | nodeport | both ingress Where the port opens.
protocol tcp tcp Only tcp. Envoy proxies the connection as is.
port_env_name string ASTRAEUS_EXT_<NAME>_PORT The variable that holds the published port. <NAME> is the name in capitals, other characters as _.
proxy object none An HTTP allow-list for calls through the Astraeus API: $API/workers/<worker>/proxy/<path>. Same fields and limits as an endpoint's proxy.

Publish an endpoint or a replica group#

Create a balanced endpoint with the exposure you need:

An HTTP API on the ingress machines and on the workers' machines
{
  "metadata": {"name": "api"},
  "spec": {
    "selector": {"app": "api"},
    "mode": "balanced",
    "port": 30080,
    "target_port": 8080,
    "protocol": "http",
    "health_check_path": "/healthz",
    "exposure": "both"
  }
}

A replica group's load balancer is an endpoint the group creates with the group's name, on the ingress machines. See Replica groups and scale to zero.

What Envoy does, and does not do#

Protocol Envoy
tcp Proxies each connection to a ready worker.
http Proxies each request (HTTP/1.1 or HTTP/2 from clients), round robin. No timeout on a whole response by default (route_timeout_seconds), and at most one hour between bytes (stream_idle_timeout_seconds, default 3600).
grpc As http, with HTTP/2 to the workers.

For each endpoint, Envoy allows up to 16,384 connections, pending requests and active requests, ejects a worker for 10 s after 5 consecutive 5xx answers, and with health_check_path checks each worker every 3 s.

Published ports are open to whoever can reach them

The listeners on the ingress machines do not terminate TLS, match host names, authenticate clients or filter source addresses. Anyone who can reach the port reaches the service. Restrict access with your firewall or security groups, and put your own load balancer or reverse proxy in front for TLS, host names and authentication. The service itself must authenticate its users.

Limits#

Limit Value
Port range 30000–32767, shared by every balanced endpoint and external access of the cluster (2,768 ports)
Accesses per run Each access selects one worker: replica 0 of the group that declares it
external_accesses[].protocol tcp only
Activator hold 2 minutes by default, then 503 with retry-after: 5 for HTTP and gRPC, or the connection is closed for TCP
external_host Not filled in by the current release: connect to an ingress machine's address

Troubleshooting#

Symptom Cause Fix
Access Pending: no port assigned yet The endpoint has just been created, or every port is taken (PORTS_EXHAUSTED). Wait a few seconds; free ports by deleting unused endpoints.
Connection refused on the ingress machine Envoy is not running, or not with the agent's bootstrap. Start Envoy with -c /etc/astraeus/envoy/xds/envoy.json; check astraeus-agent-edge runs (systemctl status astraeus-agent-edge).
Connection times out A firewall blocks 30000–32767, or the ingress machine is not on the mesh. Open the port; check the machine is Up in Machines.
503 or a closed connection after a while No worker became ready (for a replica group at zero: within the activator's hold). Check the workers' state and health checks.
The port is published twice on one machine An ingress machine also publishes nodeport ports. Set ASTRAEUS_INGRESS_LISTEN_ADDRESS and ASTRAEUS_NODE_PORT_ADDRESS to different addresses.
An access is never exposed An endpoint named <run>-<access> already exists and is not the run's. Rename the access or delete the other endpoint.