Skip to content

A nightly flow with retries#

You will build nightly-digest: every night it exports the previous day's incidents, has an agent summarise them (skipped on a quiet day), asks the on-call admin, and publishes the digest. Each step is retried on failure; a machine going away mid-run costs one attempt, never a step done twice.

flowchart LR
  C[collect<br/>run · retry 3] --> S[summarise<br/>agent · when count ≠ 0 · retry 2]
  S --> O[ok<br/>approval · when summarised · 8 h]
  O --> P[publish<br/>run · retry 3]

Before you begin#

  • An agent to summarise — here researcher (see the quick start).
  • A shared drive of the workspace, here digests, so steps on different machines see each other's files (see Drives).
  • Images for the export and the publication: here ghcr.io/acme/incident-export:2.1 and ghcr.io/acme/digest-publisher:1.0.
  • The editor role, astra signed in (astra login, astra use acme/research).

1. Write the flow#

nightly-digest.yaml
apiVersion: anemoi.astralyx/v1
kind: Flow
metadata:
  name: nightly-digest
spec:
  description: Export a day's incidents, summarise them, ask, then publish.
  drive: digests
  inputs:
    - name: day
      description: The day to report on (YYYY-MM-DD)
      required: true
  timeout_seconds: 43200                 # the whole execution: 12 h
  steps:
    - id: collect
      kind: run
      retry: {max: 3, backoff_seconds: 300}
      timeout_seconds: 1800
      template:
        image: ghcr.io/acme/incident-export:2.1
        command: sh
        args:
          - -c
          - |
            set -e
            mkdir -p /flow/collect
            export --day "${{ inputs.day }}" --out /flow/collect/incidents.json
            echo "count=$(jq length /flow/collect/incidents.json)" >> "$ASTRAEUS_OUTPUTS"
        requested_resources: {cpu_cores: 1, memory_bytes: 536870912}
    - id: summarise
      kind: agent
      needs: [collect]
      when: steps.collect.outputs.count != '0'
      agent: researcher
      input: >-
        Summarise the ${{ steps.collect.outputs.count }} incidents of ${{ inputs.day }}
        in /flow/collect/incidents.json: what happened, impact, follow-ups. Markdown.
      retry: {max: 2, backoff_seconds: 600}
      timeout_seconds: 3600
    - id: ok
      kind: approval
      needs: [summarise]
      when: steps.summarise.state == 'Succeeded'
      target: "Publish the incident digest for ${{ inputs.day }} (${{ steps.collect.outputs.count }} incidents)?"
      approvers: ["role:admin"]
      wait_seconds: 28800
    - id: publish
      kind: run
      needs: [ok]
      when: steps.summarise.state == 'Succeeded'
      retry: {max: 3, backoff_seconds: 60}
      template:
        image: ghcr.io/acme/digest-publisher:1.0
        command: publish
        args: ["--title", "Incidents ${{ inputs.day }}", "/flow/summarise/answer.md"]
        requested_resources: {cpu_cores: 1, memory_bytes: 268435456}
  outputs:
    incidents: "${{ steps.collect.outputs.count }}"
    published: "${{ steps.publish.state }}"

What to notice:

  • collect writes count=<n> to $ASTRAEUS_OUTPUTS; summarise reads it in when and in its input. On a quiet day summarise is Skipped. A step skipped for its condition counts as done for the steps that need it — a branch not taken — so ok and publish carry their own when, and are skipped too.
  • The agent's answer is written to /flow/summarise/answer.md on the drive; publish reads it from there.
  • retry makes new attempts after a failure, with a pause; a step whose machine goes away gets a new attempt elsewhere.

2. Check it and apply it#

Anemoi → Flows → New flow, name nightly-digest, open YAML, paste the file (from spec), check the graph on Steps, and press Create the flow.

$ astra anemoi flows validate -f nightly-digest.yaml --remote
nightly-digest.yaml: flow nightly-digest is valid
$ astra anemoi flows apply -f nightly-digest.yaml
flow nightly-digest: version 1 written, now current
$ yq -o=json '{"metadata": .metadata, "spec": .spec}' nightly-digest.yaml \
    | curl -sS -X POST "$WS/flows" -H "Authorization: Bearer $ASTRA_TOKEN" -H "Content-Type: application/json" -d @-

3. Run it once by hand#

On the flow's page, Run flow, day = 2026-09-30.

$ astra anemoi flows run nightly-digest --input day=2026-09-30
execution nightly-digest-cb4855 of flow nightly-digest started
$ astra anemoi flows get nightly-digest-cb4855
$ curl -sS -X POST "$WS/flows/nightly-digest/executions" -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H "Content-Type: application/json" -d '{"inputs": {"day": "2026-09-30"}}' | jq -r .metadata.name

The execution's page draws each step as it goes. When it reaches ok, the execution is Waiting and the admins see the question in Approvals.

An execution waiting on its approval step

4. Start it every night#

Flows have no schedule of their own. Start the execution from a scheduler you already run, with a personal API token (API tokens) in ASTRA_TOKEN:

.github/workflows/nightly-digest.yml
on:
  schedule:
    - cron: "15 2 * * *"           # 02:15 UTC
jobs:
  digest:
    runs-on: ubuntu-latest
    steps:
      - run: |
          curl -fsSL https://console.astralyx.cloud/install-cli.sh | sh
          astra anemoi flows run nightly-digest --input day=$(date -u -d yesterday +%F)
        env:
          ASTRA_TOKEN: ${{ secrets.ASTRA_TOKEN }}
          ASTRA_ORG: acme
          ASTRA_WORKSPACE: research
crontab
15 2 * * * ASTRA_TOKEN=ast_pat_… ASTRA_ORG=acme ASTRA_WORKSPACE=research astra anemoi flows run nightly-digest --input day=$(date -u -d yesterday +\%F)

5. When a step fails#

  • A step that fails all its attempts fails the execution (on_failure: stop) and cancels steps under way. Retry step on the execution's page (or astra anemoi flows retry <execution> <step>) gives it one more attempt; the execution resumes from there — finished steps are not run again.
  • A denied or lapsed approval fails ok. Retrying it asks again.
  • astra anemoi flows executions --flow nightly-digest lists the nights, with their state and how far each got.