A support agent on an Eos model#
You will build support-agent: given a customer's question or a ticket
number, it searches your Notion knowledge base, drafts an answer, and
posts it to the ticket as an internal note — the post waiting for a support
lead's approval. It thinks with an open model served by
Eos on your GPUs: the model, the prompts and the
answers stay on your machines, and its calls cost nothing per token.
Before you begin#
- An Eos deployment of the workspace that is
Ready, heresupport-chat— a chat model good at tool use, such as Qwen2.5 32B Instruct or Llama 3.3 70B Instruct. See Serve an open chat model on one GPU machine. - A machine that can run agents (it need not have a GPU), and the editor role.
- A Notion integration token with access to the knowledge-base pages, in a
credential
notion-token(keytoken). - Your ticketing system's API key in a credential
helpdesk-key(keykey). This recipe uses a generic HTTP API athttps://helpdesk.acme.example/apithat takes the key inX-Api-Key.
1. Create the agent#
// Notion: read the knowledge base.
@id("notion-reads")
permit (principal, action == Action::"http", resource in Server::"notion")
when { resource.method == "GET" || (resource.method == "POST" && (resource.path == "/v1/search" || resource.path like "/v1/databases/*/query")) };
// Helpdesk: read tickets freely.
@id("helpdesk-reads")
permit (principal, action == Action::"http", resource in Server::"helpdesk")
when { resource.method == "GET" && resource.path like "/tickets/*" };
// Helpdesk: internal notes wait for a support lead.
@id("helpdesk-notes")
@approval("role:admin")
@approval_wait("4h")
permit (principal, action == Action::"http", resource in Server::"helpdesk")
when { resource.method == "POST" && resource.path like "/tickets/*/notes" };
- Anemoi → Agents → New agent, name
support-agent, Blank, Agent: Assistant. - Model: An Eos deployment,
support-chat. - Connections: Add connection → Notion, credential
notion-token, Read only. - Switch to Advanced → Tools → Add a tool: name
helpdesk, HTTP, URLhttps://helpdesk.acme.example/api, credentialhelpdesk-key, keykey, headerX-Api-Key. - Replace the Tool policy with the text above.
- Instructions: You answer customers' support questions for Acme. Search the Notion knowledge base first and answer only from it, citing the page. When the input is a ticket number, read the ticket, draft the reply, and post it as an internal note on the ticket. If the knowledge base does not cover it, say so.
- Limits: 15 minutes; Tokens per run
300000. - Create the agent.
$ jq -n --rawfile p support-agent.cedar '{
metadata: {name: "support-agent"},
spec: {
kind: "astralyx",
model: {deployment: "support-chat"},
instructions: "You answer customers support questions for Acme. Search the Notion knowledge base first and answer only from it, citing the page. When the input is a ticket number, read the ticket, draft the reply, and post it as an internal note on the ticket. If the knowledge base does not cover it, say so.",
tools: [
{name: "notion", kind: "http", url: "https://api.notion.com", credential: "notion-token",
headers: {"Notion-Version": "2026-03-11"}},
{name: "helpdesk", kind: "http", url: "https://helpdesk.acme.example/api",
credential: "helpdesk-key", credential_key: "key", header: "X-Api-Key"}
],
policy: {tools: $p},
budget: {max_seconds: 900, max_tokens: 300000}
}}' \
| curl -sS -X POST "$WS/agents" -H "Authorization: Bearer $ASTRA_TOKEN" -H "Content-Type: application/json" -d @-
2. Ask it something#
Run agent, input: A customer asks whether exports keep their
folder structure. The output is the answer with the page it came
from; the trace shows the Notion searches and pages read, each
allowed by notion-reads, and the model calls to support-chat.
With a ticket number, the run reads the ticket, drafts the reply and asks to post the note: a support lead approves it in Anemoi → Approvals, after reading the note's text in The call.
What stays where#
| What | Where |
|---|---|
| The model and its weights | Your GPU machine (the Eos deployment). |
| Prompts, the knowledge base's pages, answers, the note | Your machines: the agent's sandbox, the gateway, the deployment. Notion and your helpdesk see only their own API calls. |
| Notion's and the helpdesk's keys | Your secret store, read by your machine; added by the gateway; never in the sandbox. |
| What Astralyx receives | States, token counts, decisions, digests (the receipt). |
Calls to an Eos deployment cost $0 per token in Anemoi's cost views (its GPU hours are in usage), so a token budget is the limit that matters here.
Make it faster or sturdier#
- Replicas: give
support-chatmore replicas when several runs think at once (Multiple replicas behind one endpoint). - Fallback: in Advanced → Model → Routing, add a fallback to a
second deployment for when the first is down or busy (
429,5xx). - Evaluate it: turn good runs into cases with Add to eval suite and gate new versions on them (Evaluate and promote).