Skip to content

Endpoints#

An endpoint gives a stable name to the workers whose labels match its selector. Clients use the endpoint's name instead of worker addresses, which change every time a worker restarts or moves. Only workers that are ready receive traffic.

Use an endpoint to:

  • let other runs reach a service by name: a Redis, a vector database, a model server (redis.<namespace>.astraeus.local);
  • put several replicas of a service behind one name;
  • publish a service outside the cluster on a fixed port of your ingress machines (see External access);
  • let the console or an API client call a service's HTTP API through the cluster, within an allow-list of methods and paths.

A replica group with a load balancer creates and keeps its own endpoint. A run's external_accesses also become endpoints. This page covers endpoints you create yourself.

Headless and balanced#

mode Inside the cluster Outside the cluster
headless (default) The name resolves to the addresses of the ready workers. The client picks one and connects to the worker's port. Not exposed. port is informational.
balanced The same: the name resolves to the ready workers. A listen port (30000–32767) is reserved for the endpoint, the same on every machine. Envoy on your ingress machines listens on it and balances across the ready workers, and/or each worker's machine publishes it (exposure).

Inside the cluster, clients always connect to workers directly at target_port. Envoy is only used for traffic from outside.

Before you begin#

  • You need the admin or editor role in the workspace.
  • The workers you want to select must carry labels. Labels you set on a run (metadata.labels) are copied to every worker of the run. Every worker also carries job.astraeus.io/group, job.astraeus.io/replica-index and job.astraeus.io/group-replica-index.
  • Workers resolve names through the cluster DNS only on the node network (the default network mode). See Names and service discovery.
  • For the API examples, set TOKEN and API as described in Drives.

No CLI commands for endpoints

The astra CLI has no endpoint commands. Use the console or the API.

Create an endpoint#

The example exposes a run of an HTTP API server listening on port 8080.

  1. Label the run. Create it with a label, for example "metadata": {"name": "api-server", "labels": {"app": "api"}}, and add a health check to its worker template so that only healthy workers receive traffic:

    run.json (excerpt)
    {
      "metadata": {"name": "api-server", "labels": {"app": "api"}},
      "spec": {
        "lifetime": "Service",
        "task_template": {
          "image": "ghcr.io/acme/api:3.2",
          "requested_resources": {"cpu_cores": 2, "memory_bytes": 4294967296},
          "health_check": {"type": "HTTP", "path": "/healthz", "port": 8080}
        }
      }
    }
    
  2. Create the endpoint.

    1. Open Resources in the workspace, choose the cluster, and select the Endpoints tab.
    2. Click New. A JSON editor opens with a starting point.
    3. Enter the endpoint and click Create:

      {"metadata": {"name": "api"}, "spec": {"selector": {"app": "api"}, "port": 80, "target_port": 8080, "protocol": "http"}}
      

    The Endpoints tab of Resources

    $ curl -sS -X POST "$API/endpoints" \
        -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
        -d '{
          "metadata": {"name": "api"},
          "spec": {"selector": {"app": "api"}, "port": 80, "target_port": 8080, "protocol": "http"}
        }'
    

    The cluster answers 201 Created.

  3. Check its backends. GET $API/endpoints/api returns the endpoint with its runtime:

    {
      "runtime": {
        "address": "api.ws-3f9c2a7d41be.astraeus.local",
        "ready_backends": 1,
        "total_backends": 1,
        "backends": [{"task_name": "api-server-0", "ip": "10.240.3.7", "ready": true, "state": "Running", "node_name": "gpu-02"}]
      }
    }
    

    The second part of the address is the workspace's namespace (ws- and 12 hexadecimal digits).

  4. From any worker of the same workspace, call it by name at the workers' port:

    $ curl -s http://api:8080/healthz
    ok
    

A balanced endpoint#

Set mode to balanced to reserve a listen port and serve it from your ingress machines:

A balanced HTTP endpoint
{
  "metadata": {"name": "api"},
  "spec": {
    "selector": {"app": "api"},
    "mode": "balanced",
    "target_port": 8080,
    "protocol": "http",
    "health_check_path": "/healthz"
  }
}

With port omitted (or 0), the cluster assigns the lowest free port from 30000–32767 and keeps it for the endpoint's life. runtime.listen_port shows it. See External access for how clients outside reach it.

Which workers receive traffic#

A worker is a backend of the endpoint when:

  • it is in the same workspace as the endpoint. A selector never reaches another workspace's workers;
  • it carries every label of the selector (an empty selector is refused).

A backend is ready, and receives traffic, when:

  • the worker is Running;
  • it has an address;
  • if its worker template has a health_check, the last check passed.

The cluster updates backends as workers change state. In the cluster DNS, the endpoint's name has an address only while at least one backend is ready.

Health checks#

There are two kinds of check, and a balanced endpoint can use both:

  • The worker's own health_check decides readiness for every endpoint and for the DNS. It runs on the worker's machine:

    Field Default Description
    type required HTTP (a GET that must succeed), TCP (a connection) or Exec (a command in the container).
    path HTTP only: the path.
    port HTTP and TCP: the port, 1–65535.
    command Exec only: the command and its arguments.
    interval_seconds 10 Seconds between checks.
    timeout_seconds 5 Seconds a check may take. Must be less than the interval.
    failure_threshold 3 Consecutive failures before the worker is unhealthy.
    initial_delay_seconds 0 Seconds before the first check.
  • health_check_path on a balanced endpoint with protocol http or grpc adds an active check by Envoy on the ingress machines: every 3 s, 2 s timeout, unhealthy after 2 failures, healthy again after 2 successes.

Envoy also ejects a backend for 10 s after 5 consecutive 5xx answers.

Call a service through the cluster#

An endpoint with a proxy can be called through the cluster's API, as the user making the request. Use it to reach a service from the console or from a machine that is not on the mesh, without publishing a port. Only the methods and paths you allow are relayed:

Allow GET and POST on /v1/ only
{
  "metadata": {"name": "api"},
  "spec": {
    "selector": {"app": "api"},
    "target_port": 8080,
    "protocol": "http",
    "proxy": {"enabled": true, "allowed_methods": ["GET", "POST"], "allowed_paths": ["/v1/*"], "timeout_ms": 30000}
  }
}
$ curl -sS "$API/endpoints/api/proxy/v1/status" -H "Authorization: Bearer $TOKEN"
Field Default Description
enabled false A disabled proxy allows nothing.
allowed_methods ["GET"] Any of GET, POST, PUT, PATCH, DELETE, HEAD, OPTIONS.
allowed_paths required when enabled Patterns starting with /. /v1/* matches /v1 and everything below it. A pattern ending in a bare * (/v1*) is a plain prefix and also matches /v1backdoor: prefer /v1/*.
max_body_bytes 1000000 Largest request body, at most 1400000.
timeout_ms 10000 How long the cluster waits for the answer, at most 60000.

How the relay behaves:

  • The path is percent-decoded and . and .. are resolved before it is matched. A path that climbs above / is refused.
  • Only the request headers Content-Type, Accept, Accept-Encoding and X-* are forwarded.
  • Answers are cut at 700 KiB.
  • The request goes to one ready backend. When none is ready, the answer is 503 NO_READY_BACKEND (the endpoint has no ready backend). For an endpoint that can wake a replica group, the cluster wakes it and holds the request up to 45 s (see Replica groups).
  • Refusals: 403 PROXY_DISABLED, 403 PROXY_METHOD_NOT_ALLOWED, 403 PROXY_PATH_NOT_ALLOWED, 400 PROXY_BODY_TOO_LARGE, 400 PROXY_BAD_PATH.

Change or delete an endpoint#

On Resources → Endpoints, click a row to see the endpoint as JSON, or Delete to delete it. There is no edit form: use the API to change one.

Replace the spec. The body carries the whole new spec:

$ curl -sS -X PUT "$API/endpoints/api" \
    -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
    -d '{"spec": {"selector": {"app": "api"}, "mode": "balanced", "target_port": 8080, "protocol": "http"}}'

Delete it:

$ curl -sS -X DELETE "$API/endpoints/api" -H "Authorization: Bearer $TOKEN"

The cluster answers 204 No Content.

An endpoint that a run's external access created is changed through the run: a direct change is refused with 409 MANAGED_BY_JOB (endpoint web-http is run web's external access; change the run instead). An endpoint that a replica group owns is rewritten by the group on its next pass.

Reference#

Field Type Default Description
metadata.name string required Letters, digits, - and _, at most 63 characters, not starting or ending with -. It becomes a DNS name.
spec.selector map required Labels a worker must all carry. At least one.
spec.mode headless | balanced headless See Headless and balanced.
spec.port integer 0 Balanced: the listen port, 30000–32767, or 0 to have one assigned. Headless: informational.
spec.target_port integer port The workers' port. Required for a balanced endpoint.
spec.protocol tcp | http | grpc tcp How Envoy proxies it: TCP, or HTTP (HTTP/1.1 and HTTP/2 from clients; HTTP/2 to grpc backends).
spec.health_check_path string none Balanced, http or grpc: Envoy's active check path. Starts with /.
spec.route_timeout_seconds integer none (unbounded) Balanced, http or grpc: the most a whole response may take.
spec.stream_idle_timeout_seconds integer 3600 Balanced, http or grpc: the longest gap between bytes.
spec.exposure ingress | nodeport | both ingress Balanced only: where the listen port opens. See External access.
spec.proxy object none The HTTP allow-list for calls through the cluster. See Call a service through the cluster.
spec.wake_targets list none Replica groups to wake when a request finds no ready backend. Set by replica groups for their own endpoints.

Runtime#

Field Description
address <name>.<namespace>.astraeus.local.
ready_backends, total_backends Ready and selected workers.
backends Each worker: task_name, ip, ready, state, node_name.
listen_port Balanced: the reserved port, the same on every machine.

Errors#

Code When
400 INVALID_ENDPOINT An invalid spec, such as spec.selector must name at least one label or a balanced endpoint needs spec.target_port (or spec.port).
400 NAME_NOT_DNS_SAFE The name cannot be published in DNS.
409 PORT_TAKEN port 30080 is reserved by endpoint web.
409 PORTS_EXHAUSTED All 2,768 ports of 30000–32767 are reserved.
409 ENDPOINT_ALREADY_EXISTS The name is taken.