Skip to content

Serve an internal service with an endpoint#

You run a small HTTP API, a label lookup service, as a replica group of two replicas, put an endpoint named labels in front of it, and call it by that name from a batch run. Replicas can move, restart or be replaced; the name stays. Then you open it to callers outside the cluster and let it scale to zero when nobody uses it.

What you need:

  • A workspace with any machine (the service needs no GPU). Machines whose workers are on the mesh network (the default for machines without RDMA): cluster names resolve there. Workers on a machine's host network do not get the cluster DNS.
  • The editor role in the workspace, an API token, curl and jq (How the recipes are written).

The pieces#

flowchart LR
  client["run label-client<br/>curl http://labels:8080"] -- "DNS: labels → ready replicas" --> ep(("endpoint<br/>labels"))
  ep --> r0["labels-api-0-0<br/>healthy"]
  ep --> r1["labels-api-1-0<br/>healthy"]
  rg["replica group labels-api<br/>min 2, max 2"] -. "keeps 2 runs" .-> r0
  rg -.-> r1
Object API name What it does
Run with lifetime: Service job Runs until stopped; restarted whenever it ends. Never Completed.
Replica group scaling group (/v1/replica-groups) Keeps N identical service runs (labels-api-0, labels-api-1, …), between min and max; replaces them when the template changes.
Endpoint endpoint (/v1/endpoints) A stable name over the workers its label selector matches, ready ones only.
Health check health_check What makes a worker ready: with one, only workers passing it receive traffic.

1. Write the service#

server.py
import json
import os
import signal
import sys
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer

LABELS = {0: "tench", 1: "goldfish", 207: "golden retriever", 281: "tabby cat",
          409: "analog clock", 948: "Granny Smith"}
WHO = os.environ.get("ASTRAEUS_TASK_NAME", "local").rsplit(".", 1)[-1]


class Handler(BaseHTTPRequestHandler):
    def send_json(self, code, body):
        data = json.dumps(body).encode()
        self.send_response(code)
        self.send_header("content-type", "application/json")
        self.send_header("content-length", str(len(data)))
        self.end_headers()
        self.wfile.write(data)

    def do_GET(self):
        if self.path == "/healthz":
            return self.send_json(200, {"ok": True})
        parts = self.path.strip("/").split("/")
        if len(parts) == 2 and parts[0] == "labels" and parts[1].isdigit():
            i = int(parts[1])
            if i in LABELS:
                return self.send_json(200, {"id": i, "label": LABELS[i], "served_by": WHO})
            return self.send_json(404, {"error": f"unknown id {i}", "served_by": WHO})
        self.send_json(404, {"error": "not found"})

    def log_message(self, fmt, *args):
        sys.stderr.write(f"{self.address_string()} {fmt % args}\n")


# Stop cleanly when the machine stops the worker.
signal.signal(signal.SIGTERM, lambda *_: sys.exit(0))
port = int(os.environ.get("PORT", "8080"))
print(f"{WHO} listening on :{port}", flush=True)
ThreadingHTTPServer(("0.0.0.0", port), Handler).serve_forever()

2. Create the replica group#

group.json
{
  "metadata": {"name": "labels-api", "labels": {"app": "labels-api"}},
  "spec": {
    "template": {
      "lifetime": "Service",
      "task_template": {
        "image": "python:3.12-slim",
        "command": "python",
        "args": ["/app/server.py"],
        "env": {"PORT": "8080"},
        "requested_resources": {"cpu_cores": 1, "memory_bytes": 268435456, "node_selection": {"mode": "Any"}},
        "health_check": {
          "type": "HTTP", "path": "/healthz", "port": 8080,
          "interval_seconds": 10, "timeout_seconds": 2, "failure_threshold": 3, "initial_delay_seconds": 3
        },
        "configs": []
      }
    },
    "scaling": {"min": 2, "max": 2}
  }
}
$ jq --rawfile src server.py \
    '.spec.template.task_template.configs = [{"mounts": ["/app/server.py"], "value": $src}]' \
    group.json > group.full.json
Field Why
metadata.labels.app The group's labels go to its runs and their workers: the endpoint selects on app: labels-api.
lifetime: Service The server runs until stopped; if it exits, it is started again.
health_check An HTTP GET /healthz on port 8080 every 10 s (2 s timeout); 3 failures in a row make the worker unhealthy and the endpoint drops it. Defaults: 10 s, 5 s, 3. Types: HTTP (needs path and port), TCP (port), Exec (command).
scaling.min, max Exactly 2 replicas. With a metric, the group scales between them (see Variations). At most 1000.
  1. In the sidebar, open Replica groups (the Resources page).
  2. Press New, replace the starting point with group.full.json, press Create.

The group's runs, labels-api-0 and labels-api-1, appear under Runs.

$ curl -fsS -X POST "$API/replica-groups" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d @group.full.json | jq -r .metadata.name
labels-api

3. Create the endpoint#

endpoint.json
{
  "metadata": {"name": "labels"},
  "spec": {
    "selector": {"app": "labels-api"},
    "mode": "headless",
    "port": 8080,
    "target_port": 8080,
    "protocol": "http"
  }
}

A headless endpoint publishes the ready workers' addresses under its name in the cluster DNS, and the client connects to one of them directly, on the workers' own port. The selector only ever matches workers of your workspace.

Endpoints in the sidebar → New, paste endpoint.json, Create.

$ curl -fsS -X POST "$API/endpoints" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d @endpoint.json | jq -r .metadata.name
labels

Check that both replicas are behind it:

$ curl -fsS "$API/endpoints/labels" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    | jq '.runtime | {address, ready_backends, total_backends, backends: [.backends[] | {task_name, ip, ready}]}'
{
  "address": "labels.ws-3f9c2a1b7d4e.astraeus.local",
  "ready_backends": 2,
  "total_backends": 2,
  "backends": [
    {"task_name": "labels-api-0-0", "ip": "10.42.1.7", "ready": true},
    {"task_name": "labels-api-1-0", "ip": "10.42.2.4", "ready": true}
  ]
}

ws-3f9c2a1b7d4e is your workspace's namespace (shown on Settings). From a worker of the same workspace, the short name labels resolves, as does the full name labels.<namespace>.astraeus.local.

4. Call it from another run#

client.json
{
  "metadata": {"name": "label-client"},
  "spec": {
    "task_template": {
      "image": "curlimages/curl:8.10.1",
      "command": "sh",
      "args": ["-c", "for i in 207 281 948 5; do curl -sS http://labels:8080/labels/$i; echo; done"],
      "restart_policy": "Never",
      "requested_resources": {"cpu_cores": 1, "memory_bytes": 134217728, "node_selection": {"mode": "Any"}}
    }
  }
}
$ curl -fsS -X POST "$API/runs" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d @client.json > /dev/null
$ astra astraeus logs label-client-0
{"id": 207, "label": "golden retriever", "served_by": "labels-api-1-0"}
{"id": 281, "label": "tabby cat", "served_by": "labels-api-0-0"}
{"id": 948, "label": "Granny Smith", "served_by": "labels-api-0-0"}
{"id": 5, "error": "unknown id 5", "served_by": "labels-api-1-0"}

served_by changes: the name answers with both replicas' addresses.

5. Check it survives a replica's loss#

Delete one of the group's runs. The endpoint drops its worker at once, the group makes a replacement, and the client keeps working:

$ astra astraeus delete labels-api-1
labels-api-1 deleted
$ curl -fsS "$API/replica-groups/labels-api" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    | jq '.runtime | {desired, current, ready, reason}'
{
  "desired": 2,
  "current": 2,
  "ready": 1,
  "reason": "steady"
}

A few seconds later ready is 2 again and the endpoint lists the new worker.

6. Clean up#

$ curl -fsS -X DELETE "$API/endpoints/labels" -H "Authorization: Bearer $ASTRAEUS_TOKEN"
$ curl -fsS -X DELETE "$API/replica-groups/labels-api" -H "Authorization: Bearer $ASTRAEUS_TOKEN"
$ astra astraeus delete label-client
label-client deleted

Deleting the group deletes its runs and the endpoint it owns, if any.

Open it outside the cluster#

Two ways, which you can combine:

Through the console, with an allow-list. Give the endpoint an HTTP proxy; Astraeus then relays requests that match it, authenticated like any API call, to a ready worker. Nothing is opened on the machines.

$ jq '{spec: (.spec + {proxy: {enabled: true, allowed_methods: ["GET"], allowed_paths: ["/labels/*", "/healthz"]}})}' endpoint.json \
    | curl -fsS -X PUT "$API/endpoints/labels" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
        -H 'content-type: application/json' -d @- > /dev/null
$ curl -fsS "$API/endpoints/labels/proxy/labels/207" -H "Authorization: Bearer $ASTRAEUS_TOKEN"
{"id": 207, "label": "golden retriever", "served_by": "labels-api-0-0"}

allowed_methods defaults to GET only; requests are limited to 1 MB of body and 10 s by default (max_body_bytes, timeout_ms). Paths are normalised before they are matched, so /labels/../admin is refused.

On the ingress machines, as a port. Give the group a load balancer. It owns a balanced endpoint named like the group (labels-api): an Envoy listener on every machine that runs the agent's edge part (installed with --agents drives,credentials,data,edge), on a port from 30000–32767 kept for the endpoint's lifetime, in front of the ready replicas.

$ curl -fsS "$API/replica-groups/labels-api" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    | jq '{spec: (.spec | .load_balancer = {enabled: true, protocol: "http", target_port: 8080, health_check_path: "/healthz", drain_seconds: 20})}' \
    | curl -fsS -X PUT "$API/replica-groups/labels-api" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
        -H 'content-type: application/json' -d @- > /dev/null
$ curl -fsS "$API/endpoints/labels-api" -H "Authorization: Bearer $ASTRAEUS_TOKEN" | jq '.runtime.listen_port'
30017
$ curl -fsS http://ingress-01.example.net:30017/labels/281
{"id": 281, "label": "tabby cat", "served_by": "labels-api-1-0"}

external_port asks for a specific port in that range. drain_seconds keeps a removed replica serving that long before it is deleted. Open the port range on the ingress machines' firewall only to the networks that should reach it (Network and firewalls).

Scale to zero#

A service used now and then need not hold a machine all day. With min: 0 and a load balancer, the group goes to zero replicas after it has been idle, and the next request wakes it:

fragment of group.json
"scaling": {
  "min": 0,
  "max": 3,
  "cooldown_seconds": 60,
  "scale_to_zero": {"idle_grace_seconds": 600, "activation_replicas": 1}
},
"load_balancer": {
  "enabled": true, "protocol": "http", "target_port": 8080, "health_check_path": "/healthz",
  "proxy": {"enabled": true, "allowed_methods": ["GET"], "allowed_paths": ["/labels/*"]}
}
  • Idle means no request through the ingress or the console's proxy for 30 s, and, if the group scales on a metric, that metric at 0. After idle_grace_seconds of idleness (default 300) the group drains to zero.
  • Wake. A request to the ingress port, or through the proxy ($API/endpoints/labels-api/proxy/labels/207), finds no replica: it is held while activation_replicas (default 1) start, up to 45 s through the proxy, then served. POST $API/replica-groups/labels-api/activate wakes it ahead of time.
  • Cold start is the image pull (once per machine) plus the server's start-up and its first passing health check.
  • Calls by cluster name (http://labels-api:8080) do not wake a group at zero: with no replica, the name has no address. Callers that must wake it go through the ingress port or the proxy.

Variations#

Scale on load. Workers can report a gauge the group scales on: with "target": {"metric_name": "queue_depth", "target": 10} the group keeps about 10 per ready replica (the Kubernetes HPA rule, ±10 % tolerance), between min and max; scale-down waits out cooldown_seconds (default 60). See Replica groups and scale to zero for how workers report it.

One run, no group. For a single long-lived server, a run with "lifetime": "Service" and labels is enough: the endpoint selects its worker the same way. The group adds replacement, replica counts and rolling template changes.

Expose a run's port directly. A run can declare external_accesses on its template; each one becomes an endpoint it owns, named <run>-<access> (External access).

A GPU service. Ask for GPUs in the template as for any run. To serve a language model, Eos deployments do all of this for you (engine, weights, keys, scale to zero): see Eos.

Troubleshooting#

Symptom Cause Fix
curl: (6) Could not resolve host: labels The client runs on a machine whose workers use the host network, or in another workspace. Run the client on a mesh machine, or call through the proxy.
The endpoint shows ready_backends: 0 with workers running The health check fails (wrong port or path), or the selector does not match the workers' labels. Open a worker's page: its health status and labels.
PROXY_DISABLED: the endpoint declares no HTTP proxy No proxy on the endpoint. Add one as above.
port 30017 is already reserved by endpoint … external_port taken by another endpoint. Choose another, or leave it out.