Skip to content

Give a model tools#

A deployment can declare server-side tools: functions of its workspace (by alias), and a built-in code interpreter. A chat completion through either gateway — the hosted one or the one on your edge machines — then gets their definitions; when the model calls one, the gateway runs it, gives the result back, and asks the model again, until it answers. Your client sends one request and gets one answer, streamed or whole: it never sees the calls, only which tools ran.

The model must call tools: the deployment's Capabilities say whether it does (a model the catalog marks for tools; for vLLM, served with a tool parser).

Declare the tools#

Below, weather is the function of Write and publish a function, published, with prod pointing at a version, and chat a deployment of a model that calls tools.

Eos → Deployments → chat → Tools → Add a tool: a function and its alias, or the code interpreter. Review schema shows what the model is told, and lets you change it.

astra eos tools chat add weather@prod
astra eos tools chat add code_interpreter
astra eos tools chat              # the tools and their limits

spec.tools of the deployment (PUT /v1/eos/deployments/chat, the spec whole):

"tools": ["weather@prod", "code_interpreter"]

An entry is "<function>@<alias>" (latest without an alias), "code_interpreter", or an object (function, alias, name, description, parameters, approval, approvers, approval_wait_seconds), stored as objects.

from astralyx import inference
inference.Client.from_config().set_tools("chat", ["weather@prod", "code_interpreter"])

Each tool is called by the function's name (- as _) unless you name it. Its description and the schema of its arguments are those of the version the alias names — generated from the handler's signature and docstring when the version was published (What a model is told) — unless the tool carries its own (description, parameters). Moving the alias changes what the tool runs, and what the model is told, from the next request.

Call the deployment#

As always — any OpenAI client, a workspace API key:

curl -sS https://inference.astralyx.cloud/v1/chat/completions \
  -H "Authorization: Bearer $ASTRAEUS_API_KEY" -H 'content-type: application/json' \
  -d '{"model": "chat", "messages": [{"role": "user", "content": "How hot will it be in Lisbon tomorrow?"}]}'

The answer is the model's final one, with usage summed over the rounds and what ran:

"astraeus": {
  "rounds": 1,
  "limit": null,
  "tools": [{"name": "weather", "function": "weather", "alias": "prod", "version": 1,
             "call_id": "call_8f2a", "outcome": "ok", "duration_ms": 412, "receipt": "chat-17"}]
}

Streamed (stream: true), each round's text passes through as the model writes it, the calls of server-side tools are not shown, and astraeus comes with the last chunk. While a tool runs, the stream carries a comment line every 10 s, so nothing between you and the gateway closes it.

Try it in the Playground#

Eos → Playground, the deployment: the Playground runs the same loop and shows every call in the conversation while it runs — weather@prod, the arguments, the result or error, the duration and the version; a call held for approval links to the approval. Its receipts say playground. Only the Playground shows arguments and results: a gateway's client gets the final answer and astraeus.tools, never the calls' contents.

Your own tools#

Your own tools still work: a request may have tools of its own (named differently from the deployment's). A round in which the model calls one of yours is returned to you as it is, with only your calls; the server-side calls of that round are not run.

Limits#

Limit Default Range Past it
tool_limits.max_rounds 8 1–32 The model is asked once more with tool_choice: none, for an answer; astraeus.limit is rounds.
tool_limits.max_seconds 120 5–1800 The completion ends with 504 tool_loop_timeout (an error event when streamed).
tool_limits.interpreter_seconds 30 5–300 The code is interrupted; its session is kept.
A call's timeout the function's The result given to the model is the error.
astra eos tools chat limits --rounds 4 --seconds 60

On the hosted gateway, a call that is not streamed still has 90 s for its whole answer: stream a loop that may run longer.

Approval for a tool#

A tool can make each call wait for a person:

{"function": "send-mail", "alias": "prod", "approval": true, "approvers": ["role:admin"], "approval_wait_seconds": 600}

The call is held — its arguments stay with the gateway holding it — and an approval of kind function_tool appears in Anemoi → Approvals (and its approvers are mailed), with the call's arguments shown to them. Approved, the call is made once; denied, the model is told who denied it and why; not decided within approval_wait_seconds (default 300, within the loop's time), it is not made, and the model is told so. Approvers: role:editor (the default), role:admin, user:<id>. See Approvals.

Receipts#

Every tool call has a receipt, signed by the cluster: the deployment, the completion, the tool, the function, its alias and the version that ran, digests of the arguments and of the result given to the model — never the arguments or the result — how it ended (ok, error, timeout, denied, not_approved, unavailable, rejected), the approval it waited on, where it ran (hosted, playground, or edge:<machine>), when and how long. Each is a compact JWS (EdDSA) with the key published at /v1/identity/jwks, chained to the deployment's previous receipt. Kept 90 days.

astra eos tools chat receipts

or GET /v1/eos/deployments/chat/tool-receipts, or Tools → Receipts.

The code interpreter#

code_interpreter runs Python in a Jupyter kernel kept for the conversation: its variables, imports and files persist from one call to the next, for 15 minutes of idleness. It runs sandboxed (gVisor) on the workspace's machines, in the Hesperus image with NumPy, pandas, SciPy, scikit-learn and matplotlib. The model sends {"code": "…"} and gets what was printed, the value of the last expression, and any error (each cut at 64 KiB).

A conversation is the request's metadata.conversation_id, an x-conversation-id header, or else its first messages (which stay the same from one turn to the next).

Conversations are kept apart#

The workspace's interpreter is one instance, holding up to 16 conversations' kernels. Inside it, each conversation's kernel:

  • runs as a user of its own, with no other group, never as root;
  • works in a directory of its own (also its home and temporary directory), that no other conversation can list, read or write;
  • reaches nothing of another kernel: its connection key and sockets are in its own directory, and it cannot signal another's processes;
  • cannot write to /tmp or /var/tmp. /dev/shm stays usable (multiprocessing needs it), but what a kernel puts there is readable by itself only.

When a conversation's kernel ends — 15 minutes idle, a 17th conversation taking the place of the longest idle one, or the instance stopping — every process it started is stopped and every file it made is deleted, before its place is given to another conversation. A conversation that comes back after that starts with an empty directory and a new kernel.

Limit Value
Kernels at once, per workspace 16 (a 17th ends the longest idle; if all 16 are running code, the call is refused: try again)
Idle before a kernel ends 15 minutes
One call tool_limits.interpreter_seconds (30 s by default, at most 300)
Memory and CPU 4 GiB and 2 cores for the whole instance, shared by its kernels
Output 64 KiB each of output, errors and the result
Network the instance's: as any function's
Files lost when the kernel ends; nothing is kept between conversations