The official templates#
Astralyx's official templates, in the gallery's categories:
| Category | Templates |
|---|---|
| Tools for models | Web search, Fetch page, Wikipedia, Weather, Calculator, Current time |
| Data and search | Search my sites, Search your documents |
| Training and fine-tuning | Fine-tune with LoRA |
| Batch inference | Batch inference |
| Evaluation | Evaluate a model |
| Apps and services | Open WebUI |
No key is ever in a template: a service's key is a Credential you name when you install. Every image they run is pinned by its digest. What a run or a function runs is shown whole on the template's page (View the template): once installed, it is the workspace's to read and change.
Tools for models#
These give an open model what it lacks: the web, facts, arithmetic, the
date. Each makes a function in Python (published as version 1, prod
on it) that a model calls as a tool, with typed arguments and a description
generated from its signature, and can give it to a deployment at once
(Give it to a model now, attach_to). Each has a test. The code is
small and uses Python's standard library.
| Template | Tool | Reaches | Needs |
|---|---|---|---|
| Web search | web_search(query, max_results=5) |
SearXNG on your machines, or Tavily, Brave or Google's API | nothing (SearXNG), or a key |
| Fetch page | fetch_page(url, max_chars=8000) |
the public web only | nothing |
| Wikipedia | wikipedia(query, language="en") |
Wikipedia's API | nothing |
| Weather | weather(location, days=3, units="metric") |
Open-Meteo's API | nothing |
| Calculator | calculator(expression) |
nothing | nothing |
| Current time | current_time(timezone="UTC") |
nothing | nothing |
A tool's name is the installation's, - as _ (installed as web-search,
the tool is web_search). Every function's timeout is the template's
(below); an error comes back to the model as {"error": "…"} that says what
to do — a missing key names the Credential and the variable, a refused key
says so, an empty answer suggests other words.
Web search#
Search the web: the top results, each {title, url, snippet} (snippets cut
at 500 characters), and with Tavily a short answer.
| Input | Type | Default | Description |
|---|---|---|---|
backend |
choice | searxng |
Where searches go (below). |
searxng_url |
string (URL) | With searxng: a SearXNG you already run (its JSON format enabled). Empty: one is run on your machines. |
|
tavily_key |
secret | key api_key |
With tavily. |
brave_key |
secret | key api_key |
With brave. |
google_key, google_engine |
secret | keys api_key, engine_id |
With google: the API key and the search engine's id (cx). |
attach_to |
deployment | Give the tool to this deployment. |
backend |
What it is | Makes |
|---|---|---|
searxng — On my machines (SearXNG, free) |
SearXNG runs on the workspace's machines, reachable only inside the workspace (an endpoint, no external access). No account, no key. It asks public search engines on your behalf, as a browser would: their terms may limit automated use, and they may slow it down or refuse it. Good to start; for production, use an API or Search my sites. | A replica group (<name>-searxng: SearXNG's official image, pinned, 1 CPU core, 512 MiB, JSON answers on, its rate limiter off since only the workspace reaches it, a secret key made inside the container), its endpoint, the function. |
tavily — Tavily |
A search API made for AI agents: snippets are the relevant text of each page, plus a short answer. Free plan: 1,000 credits a month, no card. | The function. |
brave — Brave Search API |
Brave's own index of the web. A monthly credit; a card is required. | The function. |
google — Google Programmable Search |
Google's Custom Search JSON API. Google limits searching the whole web for new search engines: check yours in the Programmable Search console. | The function. |
No search engine's web page is ever scraped: only these APIs, and SearXNG's JSON answers. Timeout 30 s (each search 15 s). With On my machines, the tool is given to the model once SearXNG answers: the install waits for its worker to be healthy.
Fetch page#
Fetch a page and return {url, status, title, content_type, text,
truncated, chars}: its readable text — the page's main content when it
marks one (<main>, <article>), without scripts, styles, navigation,
footers or forms; headings as #, list items as - — cut at max_chars
(500 to 20,000; 8,000 by default). Text pages (plain text, Markdown, CSV,
JSON, XML) are returned as they are; PDFs, images and other files are
refused.
It reads the public internet only. Before connecting — and again at
every redirect, at most 5 — the host is resolved and every address it has
is checked; the connection then goes to the address that was checked, so a
name cannot be pointed elsewhere in between. Refused: loopback, private
networks (10/8, 172.16/12, 192.168/16, unique local IPv6), link-local, the
clouds' metadata services (169.254.169.254, Azure's 168.63.129.16,
Alibaba's 100.100.100.200…), carrier-grade NAT, multicast, reserved and
documentation ranges, IPv6 forms that carry one of those IPv4 addresses, and
names such as localhost, *.local and *.internal. At most 2 MiB read,
http and https only, no user name or password in the address. Timeout
30 s.
Wikipedia#
Look a topic up: {title, description, summary, url, other_results} — the
summary of the best matching article and the titles of the others; a
disambiguation page says so. language is a Wikipedia's code (en, pt,
de, ja…). Wikipedia's public API, no key. Timeout 20 s.
Weather#
The weather now and the daily forecast for a place: location (a city,
optionally with its country or region — Porto, Portugal — or
latitude,longitude), days (1 to 16), units (metric: °C, km/h, mm;
imperial: °F, mph, inch). The answer has the place found, the current
conditions, temperature, feels-like, humidity, precipitation and wind, and
each day's conditions, highs and lows, precipitation and its probability.
Open-Meteo's public APIs, no key. Timeout 20 s.
Calculator#
Evaluate an expression exactly: {expression, result}. Numbers, + - * /
// %, powers (** or ^), parentheses, pi, e, tau, and sqrt,
cbrt, exp, log (with an optional base), ln, log10, log2,
sin, cos, tan (radians) and their inverses and hyperbolics, atan2,
hypot, degrees, radians, abs, round, floor, ceil, trunc,
min, max, factorial, gcd, lcm, comb, perm. The expression is
parsed, never run as code: anything else is refused, and sizes are bounded
(1,000 characters, results of at most 4,000 digits, factorial up to
1,000). Integers are exact (2 ** 64 is 18446744073709551616). No
network. Timeout 10 s.
Current time#
Today's date and the time now in an IANA time zone (Europe/Lisbon,
America/New_York, Asia/Tokyo; UTC by default):
{timezone, iso, date, time, weekday, utc_offset, abbreviation,
daylight_saving, unix}. A zone named almost right is answered with a
suggestion (Lisbon → Europe/Lisbon). No network. Timeout 10 s.
Data and search#
Search my sites#
Search your own sites — documentation, a blog, a knowledge base — crawled
and indexed on your machines: site_search(query) answers the best
matching pages {title, url, snippet}. Searches never leave your machines;
the crawler only reads the sites you name. Free.
| Input | Type | Default | Description |
|---|---|---|---|
sites |
text (URLs) | required | One address per line (https://docs.example.com/); only pages under each are read. At most 50. |
refresh_hours |
number | 24 |
Crawl again every (1 to 720). |
max_pages |
number | 500 |
The crawl of a site stops there (1 to 20,000). |
attach_to |
deployment | Give the tool to this deployment. |
It makes a drive (<name>-index, kept on the machine's data location:
choose one first, Machines → Data location), a replica group
running the crawler and the index (Python's standard library, SQLite full
text search; 1 CPU core, 512 MiB), its endpoint inside the workspace,
and the function.
How it crawls: each site's robots.txt is read first and obeyed (its
Crawl-delay too) — if it cannot be read (other than not found), nothing
of that site is crawled; pages marked noindex are not indexed and
nofollow links not followed; one request at a time per site, at least a
second apart; same host and path; HTML and text only, at most 2 MiB a page.
Public addresses only, checked as Fetch page checks them. A new crawl builds
a new index and replaces the old one when it is done: searches meanwhile
read the old one. Until the first crawl ends, a search answers that the
sites are still being indexed. Timeout 15 s. The tool is given to the model once the index answers.
Search your documents#
Retrieval for a model, all on your machines: a drive of documents indexed
into Qdrant with an Eos deployment of an embedding
model, and search_docs(query, max_results=5) answering the closest
passages, each {path, text, score}. Neither the documents nor the
questions leave the workspace.
| Input | Type | Default | Description |
|---|---|---|---|
docs_drive |
drive | required | The drive holding the documents. |
docs_path |
string | A folder in it (empty: the whole drive). | |
embeddings |
deployment | required | An Eos deployment of an embedding model (served for /v1/embeddings): deploy one first, e.g. Qwen3 Embedding or nomic-embed-text from the catalog. |
chunk_size, chunk_overlap |
number | 1200, 200 |
Characters per chunk, and how many each repeats of the one before. |
attach_to |
deployment | Give the tool to this deployment once the index is built. |
It makes, in order: a drive for Qdrant's data, a replica group running Qdrant (its official image, pinned; 1 CPU core, 2 GiB), waited for until it is healthy, its endpoint inside the workspace, waited for until it answers; then an indexing run — text, Markdown, HTML, reStructuredText, CSV, JSON and PDF read, cut into overlapping chunks at paragraph and sentence ends, embedded by your deployment in batches, stored with their file and place — waited for until it finished successfully; the function; and last the tool on the deployment.
Indexing again replaces what changed: a chunk's id follows from its file and place, and chunks of removed or shortened files are deleted at the end. A collection made with another embedding model is refused rather than mixed. Install it again under another name to index again. The function's timeout is 30 s; before the index is built it answers that nothing is indexed yet.
Training and fine-tuning#
Fine-tune with LoRA#
Teach a model your data: LoRA adapters trained on a GPU, checked on held-out examples, merged, converted to GGUF and registered as an Eos model of the workspace — ready to deploy, or deployed at once.
It makes a drive (named after the installation; or writes to the one you give) and five runs, each made once the one before it has finished successfully — the installation's page shows the step it is on:
| Run | What it does | Image | Asks |
|---|---|---|---|
<name>-prepare |
The dataset read and made into chat examples; a share held out, never trained on (data/train.jsonl, data/eval.jsonl, data/summary.json). |
Python 3.12 + datasets 4.0.0 |
2 CPU cores, 8 GiB |
<name>-train |
LoRA on every linear layer (rank r, alpha 2r, dropout 0.05), cosine schedule, bf16 where the GPU has it, gradient checkpointing; the loss on the answers only; an epoch's checkpoint kept, so a run started again goes on from it (adapter/). |
PyTorch 2.8 (CUDA 12.8) + Transformers 4.56.2, PEFT 0.17.1 | 1 GPU of at least gpu_memory_gb, 8 cores, 32 GiB |
<name>-evaluate |
The base and the fine-tuned model's loss and perplexity on the held-out examples (eval/report.md, eval/report.json). |
as train | as train |
<name>-merge |
The adapter merged into the base weights, saved as safetensors (merged/). |
as train | as train |
<name>-convert |
Converted to GGUF by llama.cpp and quantised (gguf/model-<quantisation>.gguf). |
llama.cpp (its full image) | 4 cores, 32 GiB |
Then the model <name> (GGUF, from the drive's gguf/), and — with
Deploy the result — a deployment of it.
| Input | Type | Default | Description |
|---|---|---|---|
base_source |
choice | huggingface |
huggingface: a repository, downloaded by the runs. eos: a model of the workspace fetched from Hugging Face in safetensors (its weights read from its drive). |
base_model |
string | Qwen/Qwen3-1.7B |
The repository (owner/name). |
base_revision |
string | A commit, tag or branch (main by default). |
|
eos_model |
model | With eos. |
|
dataset_source |
choice | huggingface |
huggingface or drive. |
dataset, dataset_split |
string | train |
A dataset on the hub (HuggingFaceH4/no_robots) and its split. |
dataset_drive, dataset_path |
drive, string | A file — JSONL, JSON, CSV or Parquet — or a folder of one kind, on a drive. | |
hf_token |
secret | key token |
For a gated or private model or dataset. |
epochs, learning_rate, lora_rank |
number | 2, 0.0002, 16 |
|
max_length |
number | 2048 |
Tokens; an example longer is cut (one whose answer is cut off entirely is left out). |
eval_percent |
number | 5 |
Held out (1 to 30 %; at least one example, at most half). |
max_examples |
number | 0 |
The most examples used (0: all). |
gpu_memory_gb |
number | 24 |
The training GPU's memory at least. |
quantization |
choice | Q4_K_M |
Q4_K_M, Q5_K_M, Q8_0 or F16. |
output_drive |
drive | Write here instead of a drive made for it. | |
deploy |
bool | false |
Deploy the result. |
The dataset's records may be: messages (chat turns, role and
content) or ShareGPT's conversations (from, value); prompt with
completion or response; instruction (with input) and output
(Alpaca); question and answer; or text, trained whole. A system
field is kept. In a conversation the model learns the last answer; what
comes before is context. Records of no known shape are skipped and counted
(data/summary.json); fewer than 10 examples is refused.
Sizes. LoRA on a model of 1–3 B parameters needs about 16 GB of GPU memory, 7–8 B about 24 GB at 2,048 tokens. Each step may take: prepare 4 h, train 72 h, evaluate, merge and convert 8 h each; past it the install fails.
If a run fails — out of memory, a gated model without a token, a
dataset of no known shape — the install fails: its runs and the drive it
made are removed, and the installation says which step failed and why (the
run's state and reason). Read a run's log while it runs: Runs →
Batch inference#
Batch inference#
Run a model over every item of a file and keep every answer: one run
reads JSONL or JSON (objects), CSV (a header) or text (one prompt a line)
from a drive, makes each item a prompt with your template, asks the model,
and writes results.jsonl to the output drive — {"id", "input",
"output"} per item, or {"id", "input", "error"} when that item failed —
in the input's order. The installation is ready once the run has
finished successfully.
| Input | Type | Default | Description |
|---|---|---|---|
model_source |
choice | deployment |
deployment: an Eos deployment, called at its address inside the workspace (http://deploy-<name>:8000/v1, no key), concurrency requests at once; its tools are not used. vllm: a Hugging Face model loaded by vLLM in the run itself (offline batching, the same vLLM image Eos serves with), on gpus GPUs. |
deployment |
deployment | With deployment. Keep it at one replica at least while the run goes. |
|
hf_model, hf_token, gpus |
string, secret, number | 1 GPU |
With vllm. |
input_drive, input_file |
drive, string | required | The file on the drive. |
prompt_template |
text | {text} |
Each item's fields in braces: Summarise in one sentence: {text}; a CSV's columns by name. A literal brace is doubled ({{). |
system_prompt |
text | A system message for every request. | |
id_field |
string | A field naming each item in the results (unique); its line number otherwise. | |
max_tokens, temperature |
number | 512, 0 |
|
output_drive, output_file |
drive, string | made; results.jsonl |
Where the results go. |
An item without a field the template names is written with its error, not asked. Answers already written are kept: a run started again (its machine lost) asks only what is left. A deployment that does not answer at once (a replica starting) is asked again, up to six times.
Evaluation#
Evaluate a model#
Measure a deployment on standard benchmarks with EleutherAI's
lm-evaluation-harness
(0.4.9): one run against the deployment's API inside the workspace,
then a report on the output drive under eval/: report.md (each task's
metrics and their standard error), report.json, and the harness's
results.json.
| Input | Type | Default | Description |
|---|---|---|---|
deployment |
deployment | required | What is evaluated. |
api |
choice | chat |
chat (/v1/chat/completions, the chat template applied): generation tasks — gsm8k, ifeval, bbh_cot_zeroshot… completions (/v1/completions): also multiple-choice tasks scored by log-likelihood — arc_easy, hellaswag, mmlu —; needs log-probabilities (vLLM answers them; llama.cpp's server does not) and tokenizer. |
tasks |
string | gsm8k |
Comma-separated task names. |
tokenizer |
string | With completions: the model's Hugging Face repository. |
|
limit |
number | 200 |
Examples per task (0: all — gsm8k has 1,319, MMLU 14,042). |
num_fewshot |
number | 0 |
Few-shot examples (0: the task's default). |
concurrency |
number | 4 |
Requests at once. |
output_drive, report_dir |
drive, string | made; eval |
Where the report goes. |
A limit makes it an estimate, not a benchmark's score. The deployment serves the requests like any others: keep it at one replica at least. The installation is ready once the run has finished successfully; install it again under another name to evaluate again.
Apps and services#
Open WebUI#
Open WebUI, the open-source chat app, on your machines, talking to Eos through its gateway with a workspace API key: every deployment the key may reach is a model in its list, with your tools, budgets and usage counted as for any client. Its accounts, chats and uploads are kept on a drive.
| Input | Type | Default | Description |
|---|---|---|---|
gateway_url |
string (URL) | https://inference.astralyx.cloud/v1 |
The Eos gateway: the hosted one, or yours (a deployment's page, Gateway → Base URL) to keep requests on your machines. |
api_key |
secret | key api_key |
A Credential holding a workspace API key (ak-…). |
exposure |
choice | workspace |
workspace: reachable only inside the workspace. edge: opened on the workspace's edge machines like any endpoint opened there — reachable wherever they are, possibly the internet. |
name_shown |
string | Open WebUI |
Its name in the app. |
It makes a drive (<name>-data), a replica group running Open
WebUI (its image, pinned; 2 CPU cores, 4 GiB; telemetry off, community
sharing off, Ollama off), waited for until it is healthy, and its
endpoint: inside the workspace at http://<name>:8080 — from its runs,
notebooks and environments, or from your computer through an environment
(astra ssh <environment> -L 8080:<name>:8080, then
http://localhost:8080) — or, with edge, at the endpoint's address on
the edge machines.
The first account made becomes its administrator: make yours at once. Accounts made afterwards wait for the administrator's approval.