External access#
External access publishes a service outside the cluster, on a port in the range 30000–32767 that is the same on every machine. Clients outside connect to that port on one of your ingress machines, where Envoy balances across the ready workers, or directly on the machine that runs the worker.
Use external access to:
- open a notebook, a dashboard or a TensorBoard running in a run from your laptop;
- serve an HTTP or gRPC API from a run or a replica group to clients outside the cluster;
- put your own load balancer, TLS terminator or API gateway in front of a service that runs on your machines.
Three things publish a port:
| What | How | Selects |
|---|---|---|
A run's external_accesses |
Declared on the run's worker template | The first worker (replica 0) of the worker group that declares it |
| A balanced endpoint | Created by you, mode: balanced |
Every ready worker matching its selector |
| A replica group's load balancer | load_balancer.enabled: true |
Every ready member of the group |
All three are balanced endpoints in the end: a run's external access becomes an endpoint named <run>-<access>, owned by the run and deleted with it.
How traffic flows#
flowchart LR
C[Client] -->|"ingress machine :30080"| E["Envoy (ingress machine)"]
E -->|"mesh, target port"| W1["Worker on gpu-01"]
E -->|"mesh, target port"| W2["Worker on gpu-02"]
C -.->|"nodeport: gpu-01 :30080"| W1
E -.->|"no ready worker, can wake"| A["Activator (same machine)"]
A -.->|"wakes the replica group, then splices"| W1
- Each balanced endpoint gets one listen port, chosen once by Astraeus and kept for the endpoint's life. It is the same on every machine.
- With
exposure: ingress(the default), Envoy on every ingress machine of your organisation listens on that port and sends each connection or request to a ready worker, round robin, at the worker's target port. - With
exposure: nodeport, the machine that runs each selected worker publishes the port and forwards it to the container. Withboth, both happen. - When an endpoint can wake a replica group scaled to zero and no worker is ready, Envoy hands connections to the activator on the same machine. See Replica groups and scale to zero.
Before you begin#
- You need the admin or editor role in the workspace.
- For
exposure: ingress, you need at least one ingress machine: a machine of the cluster (it runs the worker and joins the mesh, so it reaches the workers' addresses) that also runs the agent's edge part and Envoy. See Set up an ingress machine. - Your firewall must let clients reach TCP 30000–32767 on the ingress machines (or, for
nodeport, on the machines that run the workers). Open only what you publish. - For the API examples, set
TOKENandAPIas described in Drives.
No CLI commands for external access
The astra CLI has no commands for external access. Declare it in the run (console New run → Edit as JSON, or the API).
Set up an ingress machine#
An organisation admin does this once per ingress machine.
-
Install the machine with the agent's edge part among its parts:
$ curl -fsSL <installer URL> | sudo sh -s -- \ --apiserver https://api.astralyx.cloud --token-file ./token \ --agents drives,credentials,data,edge--apiserveris the Astraeus address shown in the console when you add a machine; the join token comes from the same place (Install a machine). The edge part (the serviceastraeus-agent-edge) connects out to Astraeus over HTTPS as the machine, like the agent's other parts. It writes Envoy's configuration to/etc/astraeus/envoy/xds/: a bootstrapenvoy.json, andcds.jsonandlds.json, which it rewrites whenever endpoints or their workers change. -
Install Envoy on the machine. The installer does not install it.
-
Run Envoy with the bootstrap the agent wrote, for example as a systemd service:
Envoy reads its listeners and clusters from the files and reloads them when they change. Its admin interface listens on
127.0.0.1:9901, where the agent reads its traffic counters. -
Check that a balanced endpoint answers on the machine:
The edge part's settings, in /etc/astraeus/agent.env:
| Variable | Default | Description |
|---|---|---|
ASTRAEUS_ENVOY_XDS_DIR |
/etc/astraeus/envoy/xds |
Where Envoy's configuration is written. |
ASTRAEUS_ENVOY_ADMIN_PORT |
9901 |
Envoy's admin port, on loopback, written into the bootstrap. |
ASTRAEUS_INGRESS_LISTEN_ADDRESS |
0.0.0.0 |
The address Envoy's listeners bind. |
ASTRAEUS_ACTIVATOR_PORT |
9902 |
The activator's loopback port. 0 turns scale-to-zero wake-ups off on this machine. |
ASTRAEUS_ACTIVATOR_TIMEOUT |
2m |
How long the activator holds a connection while a replica group wakes. |
An ingress machine that also runs workers
If the machine also runs workers whose endpoints use nodeport or both, Envoy and the worker would both try to open the same ports. Give each its own address: ASTRAEUS_INGRESS_LISTEN_ADDRESS for Envoy, and ASTRAEUS_NODE_PORT_ADDRESS for the worker.
Publish a run's port#
Declare the access on the worker template. This publishes a Jupyter server on port 8888:
{
"metadata": {"name": "notebook"},
"spec": {
"lifetime": "Service",
"task_template": {
"image": "quay.io/jupyter/pytorch-notebook:latest",
"requested_resources": {"cpu_cores": 8, "memory_bytes": 34359738368, "gpu_requests": {"count": 1}},
"external_accesses": [
{"name": "jupyter", "target_port": 8888, "enabled": true, "mode": "ingress"}
]
}
}
}
- Open New run, fill in the form, and click Edit as JSON.
- Add
external_accessestospec.task_templateas above, and submit the run. - Open Resources → Endpoints: the endpoint
notebook-jupyterappears. Click it to see itsruntime.listen_port.
Connect to http://<ingress machine address>:30000. status is Pending while no port is assigned yet (no port assigned yet), the worker is not running (the worker is not running) or not ready (the worker is not ready). For nodeport and both, the answer also carries node_name, node_ip and node_port: the machine to connect to.
Inside the container, the published port is in the environment variable ASTRAEUS_EXT_<NAME>_PORT (here ASTRAEUS_EXT_JUPYTER_PORT=30000), or the name you set in port_env_name. The container waits to start until the port is assigned.
External access fields#
| Field | Type | Default | Description |
|---|---|---|---|
name |
string | required | Unique in the worker template. The endpoint is <run>-<name>. |
target_port |
integer | required | The container's port, 1–65535. |
enabled |
boolean | false |
Only enabled accesses are published. |
external_port |
integer | assigned | A specific port in 30000–32767. Without it, the lowest free port is assigned and kept. |
mode |
ingress | nodeport | both |
ingress |
Where the port opens. |
protocol |
tcp |
tcp |
Only tcp. Envoy proxies the connection as is. |
port_env_name |
string | ASTRAEUS_EXT_<NAME>_PORT |
The variable that holds the published port. <NAME> is the name in capitals, other characters as _. |
proxy |
object | none | An HTTP allow-list for calls through the Astraeus API: $API/workers/<worker>/proxy/<path>. Same fields and limits as an endpoint's proxy. |
Publish an endpoint or a replica group#
Create a balanced endpoint with the exposure you need:
{
"metadata": {"name": "api"},
"spec": {
"selector": {"app": "api"},
"mode": "balanced",
"port": 30080,
"target_port": 8080,
"protocol": "http",
"health_check_path": "/healthz",
"exposure": "both"
}
}
A replica group's load balancer is an endpoint the group creates with the group's name, on the ingress machines. See Replica groups and scale to zero.
What Envoy does, and does not do#
| Protocol | Envoy |
|---|---|
tcp |
Proxies each connection to a ready worker. |
http |
Proxies each request (HTTP/1.1 or HTTP/2 from clients), round robin. No timeout on a whole response by default (route_timeout_seconds), and at most one hour between bytes (stream_idle_timeout_seconds, default 3600). |
grpc |
As http, with HTTP/2 to the workers. |
For each endpoint, Envoy allows up to 16,384 connections, pending requests and active requests, ejects a worker for 10 s after 5 consecutive 5xx answers, and with health_check_path checks each worker every 3 s.
Published ports are open to whoever can reach them
The listeners on the ingress machines do not terminate TLS, match host names, authenticate clients or filter source addresses. Anyone who can reach the port reaches the service. Restrict access with your firewall or security groups, and put your own load balancer or reverse proxy in front for TLS, host names and authentication. The service itself must authenticate its users.
Limits#
| Limit | Value |
|---|---|
| Port range | 30000–32767, shared by every balanced endpoint and external access of the cluster (2,768 ports) |
| Accesses per run | Each access selects one worker: replica 0 of the group that declares it |
external_accesses[].protocol |
tcp only |
| Activator hold | 2 minutes by default, then 503 with retry-after: 5 for HTTP and gRPC, or the connection is closed for TCP |
external_host |
Not filled in by the current release: connect to an ingress machine's address |
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
Access Pending: no port assigned yet |
The endpoint has just been created, or every port is taken (PORTS_EXHAUSTED). |
Wait a few seconds; free ports by deleting unused endpoints. |
| Connection refused on the ingress machine | Envoy is not running, or not with the agent's bootstrap. | Start Envoy with -c /etc/astraeus/envoy/xds/envoy.json; check astraeus-agent-edge runs (systemctl status astraeus-agent-edge). |
| Connection times out | A firewall blocks 30000–32767, or the ingress machine is not on the mesh. | Open the port; check the machine is Up in Machines. |
503 or a closed connection after a while |
No worker became ready (for a replica group at zero: within the activator's hold). | Check the workers' state and health checks. |
| The port is published twice on one machine | An ingress machine also publishes nodeport ports. |
Set ASTRAEUS_INGRESS_LISTEN_ADDRESS and ASTRAEUS_NODE_PORT_ADDRESS to different addresses. |
| An access is never exposed | An endpoint named <run>-<access> already exists and is not the run's. |
Rename the access or delete the other endpoint. |