Give a model tools#
A deployment can declare server-side tools: functions of its workspace (by alias), and a built-in code interpreter. A chat completion through either gateway — the hosted one or the one on your edge machines — then gets their definitions; when the model calls one, the gateway runs it, gives the result back, and asks the model again, until it answers. Your client sends one request and gets one answer, streamed or whole: it never sees the calls, only which tools ran.
The model must call tools: the deployment's Capabilities say whether it does (a model the catalog marks for tools; for vLLM, served with a tool parser).
Declare the tools#
Below, weather is the function of Write and publish a function,
published, with prod pointing at a version, and chat a deployment of a
model that calls tools.
Eos → Deployments → chat → Tools → Add a tool: a function and its alias, or the code interpreter. Review schema shows what the model is told, and lets you change it.
spec.tools of the deployment (PUT /v1/eos/deployments/chat, the spec
whole):
An entry is "<function>@<alias>" (latest without an alias),
"code_interpreter", or an object (function, alias, name,
description, parameters, approval, approvers,
approval_wait_seconds), stored as objects.
Each tool is called by the function's name (- as _) unless you name it.
Its description and the schema of its arguments are those of the version the
alias names — generated from the handler's signature and docstring when the
version was published (What a model is told) —
unless the tool carries its own (description, parameters). Moving the
alias changes what the tool runs, and what the model is told, from the next
request.
Call the deployment#
As always — any OpenAI client, a workspace API key:
curl -sS https://inference.astralyx.cloud/v1/chat/completions \
-H "Authorization: Bearer $ASTRAEUS_API_KEY" -H 'content-type: application/json' \
-d '{"model": "chat", "messages": [{"role": "user", "content": "How hot will it be in Lisbon tomorrow?"}]}'
The answer is the model's final one, with usage summed over the rounds and
what ran:
"astraeus": {
"rounds": 1,
"limit": null,
"tools": [{"name": "weather", "function": "weather", "alias": "prod", "version": 1,
"call_id": "call_8f2a", "outcome": "ok", "duration_ms": 412, "receipt": "chat-17"}]
}
Streamed (stream: true), each round's text passes through as the model
writes it, the calls of server-side tools are not shown, and astraeus
comes with the last chunk. While a tool runs, the stream carries a comment
line every 10 s, so nothing between you and the gateway closes it.
Try it in the Playground#
Eos → Playground, the deployment: the Playground runs the same loop and
shows every call in the conversation while it runs — weather@prod, the
arguments, the result or error, the duration and the version; a call held
for approval links to the approval. Its receipts say playground. Only the
Playground shows arguments and results: a gateway's client gets the final
answer and astraeus.tools, never the calls' contents.
Your own tools#
Your own tools still work: a request may have tools of its own (named
differently from the deployment's). A round in which the model calls one of
yours is returned to you as it is, with only your calls; the server-side
calls of that round are not run.
Limits#
| Limit | Default | Range | Past it |
|---|---|---|---|
tool_limits.max_rounds |
8 | 1–32 | The model is asked once more with tool_choice: none, for an answer; astraeus.limit is rounds. |
tool_limits.max_seconds |
120 | 5–1800 | The completion ends with 504 tool_loop_timeout (an error event when streamed). |
tool_limits.interpreter_seconds |
30 | 5–300 | The code is interrupted; its session is kept. |
| A call's timeout | the function's | The result given to the model is the error. |
On the hosted gateway, a call that is not streamed still has 90 s for its whole answer: stream a loop that may run longer.
Approval for a tool#
A tool can make each call wait for a person:
{"function": "send-mail", "alias": "prod", "approval": true, "approvers": ["role:admin"], "approval_wait_seconds": 600}
The call is held — its arguments stay with the gateway holding it — and an
approval of kind function_tool appears in Anemoi → Approvals (and its
approvers are mailed), with the call's arguments shown to them. Approved, the
call is made once; denied, the model is told who denied it and why; not
decided within approval_wait_seconds (default 300, within the loop's
time), it is not made, and the model is told so. Approvers: role:editor
(the default), role:admin, user:<id>. See Approvals.
Receipts#
Every tool call has a receipt, signed by the cluster: the deployment, the
completion, the tool, the function, its alias and the version that ran,
digests of the arguments and of the result given to the model — never the
arguments or the result — how it ended (ok, error, timeout, denied,
not_approved, unavailable, rejected), the approval it waited on, where
it ran (hosted, playground, or edge:<machine>), when and how long. Each is a compact
JWS (EdDSA) with the key published at /v1/identity/jwks, chained to the
deployment's previous receipt. Kept 90 days.
or GET /v1/eos/deployments/chat/tool-receipts, or Tools → Receipts.
The code interpreter#
code_interpreter runs Python in a Jupyter kernel kept for the
conversation: its variables, imports and files persist from one call to the
next, for 15 minutes of idleness. It runs sandboxed (gVisor) on the
workspace's machines, in the Hesperus image with NumPy, pandas, SciPy,
scikit-learn and matplotlib. The model sends {"code": "…"} and gets what
was printed, the value of the last expression, and any error (each cut at
64 KiB).
A conversation is the request's metadata.conversation_id, an
x-conversation-id header, or else its first messages (which stay the same
from one turn to the next).
Conversations are kept apart#
The workspace's interpreter is one instance, holding up to 16 conversations' kernels. Inside it, each conversation's kernel:
- runs as a user of its own, with no other group, never as root;
- works in a directory of its own (also its home and temporary directory), that no other conversation can list, read or write;
- reaches nothing of another kernel: its connection key and sockets are in its own directory, and it cannot signal another's processes;
- cannot write to
/tmpor/var/tmp./dev/shmstays usable (multiprocessing needs it), but what a kernel puts there is readable by itself only.
When a conversation's kernel ends — 15 minutes idle, a 17th conversation taking the place of the longest idle one, or the instance stopping — every process it started is stopped and every file it made is deleted, before its place is given to another conversation. A conversation that comes back after that starts with an empty directory and a new kernel.
| Limit | Value |
|---|---|
| Kernels at once, per workspace | 16 (a 17th ends the longest idle; if all 16 are running code, the call is refused: try again) |
| Idle before a kernel ends | 15 minutes |
| One call | tool_limits.interpreter_seconds (30 s by default, at most 300) |
| Memory and CPU | 4 GiB and 2 cores for the whole instance, shared by its kernels |
| Output | 64 KiB each of output, errors and the result |
| Network | the instance's: as any function's |
| Files | lost when the kernel ends; nothing is kept between conversations |