Skip to content

The official templates#

Astralyx's official templates, in the gallery's categories:

Category Templates
Tools for models Web search, Fetch page, Wikipedia, Weather, Calculator, Current time
Data and search Search my sites, Search your documents
Training and fine-tuning Fine-tune with LoRA
Batch inference Batch inference
Evaluation Evaluate a model
Apps and services Open WebUI

No key is ever in a template: a service's key is a Credential you name when you install. Every image they run is pinned by its digest. What a run or a function runs is shown whole on the template's page (View the template): once installed, it is the workspace's to read and change.

Tools for models#

These give an open model what it lacks: the web, facts, arithmetic, the date. Each makes a function in Python (published as version 1, prod on it) that a model calls as a tool, with typed arguments and a description generated from its signature, and can give it to a deployment at once (Give it to a model now, attach_to). Each has a test. The code is small and uses Python's standard library.

Template Tool Reaches Needs
Web search web_search(query, max_results=5) SearXNG on your machines, or Tavily, Brave or Google's API nothing (SearXNG), or a key
Fetch page fetch_page(url, max_chars=8000) the public web only nothing
Wikipedia wikipedia(query, language="en") Wikipedia's API nothing
Weather weather(location, days=3, units="metric") Open-Meteo's API nothing
Calculator calculator(expression) nothing nothing
Current time current_time(timezone="UTC") nothing nothing

A tool's name is the installation's, - as _ (installed as web-search, the tool is web_search). Every function's timeout is the template's (below); an error comes back to the model as {"error": "…"} that says what to do — a missing key names the Credential and the variable, a refused key says so, an empty answer suggests other words.

Search the web: the top results, each {title, url, snippet} (snippets cut at 500 characters), and with Tavily a short answer.

Input Type Default Description
backend choice searxng Where searches go (below).
searxng_url string (URL) With searxng: a SearXNG you already run (its JSON format enabled). Empty: one is run on your machines.
tavily_key secret key api_key With tavily.
brave_key secret key api_key With brave.
google_key, google_engine secret keys api_key, engine_id With google: the API key and the search engine's id (cx).
attach_to deployment Give the tool to this deployment.
backend What it is Makes
searxng — On my machines (SearXNG, free) SearXNG runs on the workspace's machines, reachable only inside the workspace (an endpoint, no external access). No account, no key. It asks public search engines on your behalf, as a browser would: their terms may limit automated use, and they may slow it down or refuse it. Good to start; for production, use an API or Search my sites. A replica group (<name>-searxng: SearXNG's official image, pinned, 1 CPU core, 512 MiB, JSON answers on, its rate limiter off since only the workspace reaches it, a secret key made inside the container), its endpoint, the function.
tavily — Tavily A search API made for AI agents: snippets are the relevant text of each page, plus a short answer. Free plan: 1,000 credits a month, no card. The function.
brave — Brave Search API Brave's own index of the web. A monthly credit; a card is required. The function.
google — Google Programmable Search Google's Custom Search JSON API. Google limits searching the whole web for new search engines: check yours in the Programmable Search console. The function.

No search engine's web page is ever scraped: only these APIs, and SearXNG's JSON answers. Timeout 30 s (each search 15 s). With On my machines, the tool is given to the model once SearXNG answers: the install waits for its worker to be healthy.

Fetch page#

Fetch a page and return {url, status, title, content_type, text, truncated, chars}: its readable text — the page's main content when it marks one (<main>, <article>), without scripts, styles, navigation, footers or forms; headings as #, list items as - — cut at max_chars (500 to 20,000; 8,000 by default). Text pages (plain text, Markdown, CSV, JSON, XML) are returned as they are; PDFs, images and other files are refused.

It reads the public internet only. Before connecting — and again at every redirect, at most 5 — the host is resolved and every address it has is checked; the connection then goes to the address that was checked, so a name cannot be pointed elsewhere in between. Refused: loopback, private networks (10/8, 172.16/12, 192.168/16, unique local IPv6), link-local, the clouds' metadata services (169.254.169.254, Azure's 168.63.129.16, Alibaba's 100.100.100.200…), carrier-grade NAT, multicast, reserved and documentation ranges, IPv6 forms that carry one of those IPv4 addresses, and names such as localhost, *.local and *.internal. At most 2 MiB read, http and https only, no user name or password in the address. Timeout 30 s.

Wikipedia#

Look a topic up: {title, description, summary, url, other_results} — the summary of the best matching article and the titles of the others; a disambiguation page says so. language is a Wikipedia's code (en, pt, de, ja…). Wikipedia's public API, no key. Timeout 20 s.

Weather#

The weather now and the daily forecast for a place: location (a city, optionally with its country or region — Porto, Portugal — or latitude,longitude), days (1 to 16), units (metric: °C, km/h, mm; imperial: °F, mph, inch). The answer has the place found, the current conditions, temperature, feels-like, humidity, precipitation and wind, and each day's conditions, highs and lows, precipitation and its probability. Open-Meteo's public APIs, no key. Timeout 20 s.

Calculator#

Evaluate an expression exactly: {expression, result}. Numbers, + - * / // %, powers (** or ^), parentheses, pi, e, tau, and sqrt, cbrt, exp, log (with an optional base), ln, log10, log2, sin, cos, tan (radians) and their inverses and hyperbolics, atan2, hypot, degrees, radians, abs, round, floor, ceil, trunc, min, max, factorial, gcd, lcm, comb, perm. The expression is parsed, never run as code: anything else is refused, and sizes are bounded (1,000 characters, results of at most 4,000 digits, factorial up to 1,000). Integers are exact (2 ** 64 is 18446744073709551616). No network. Timeout 10 s.

Current time#

Today's date and the time now in an IANA time zone (Europe/Lisbon, America/New_York, Asia/Tokyo; UTC by default): {timezone, iso, date, time, weekday, utc_offset, abbreviation, daylight_saving, unix}. A zone named almost right is answered with a suggestion (Lisbon → Europe/Lisbon). No network. Timeout 10 s.

Search my sites#

Search your own sites — documentation, a blog, a knowledge base — crawled and indexed on your machines: site_search(query) answers the best matching pages {title, url, snippet}. Searches never leave your machines; the crawler only reads the sites you name. Free.

Input Type Default Description
sites text (URLs) required One address per line (https://docs.example.com/); only pages under each are read. At most 50.
refresh_hours number 24 Crawl again every (1 to 720).
max_pages number 500 The crawl of a site stops there (1 to 20,000).
attach_to deployment Give the tool to this deployment.

It makes a drive (<name>-index, kept on the machine's data location: choose one first, Machines → Data location), a replica group running the crawler and the index (Python's standard library, SQLite full text search; 1 CPU core, 512 MiB), its endpoint inside the workspace, and the function.

How it crawls: each site's robots.txt is read first and obeyed (its Crawl-delay too) — if it cannot be read (other than not found), nothing of that site is crawled; pages marked noindex are not indexed and nofollow links not followed; one request at a time per site, at least a second apart; same host and path; HTML and text only, at most 2 MiB a page. Public addresses only, checked as Fetch page checks them. A new crawl builds a new index and replaces the old one when it is done: searches meanwhile read the old one. Until the first crawl ends, a search answers that the sites are still being indexed. Timeout 15 s. The tool is given to the model once the index answers.

Search your documents#

Retrieval for a model, all on your machines: a drive of documents indexed into Qdrant with an Eos deployment of an embedding model, and search_docs(query, max_results=5) answering the closest passages, each {path, text, score}. Neither the documents nor the questions leave the workspace.

Input Type Default Description
docs_drive drive required The drive holding the documents.
docs_path string A folder in it (empty: the whole drive).
embeddings deployment required An Eos deployment of an embedding model (served for /v1/embeddings): deploy one first, e.g. Qwen3 Embedding or nomic-embed-text from the catalog.
chunk_size, chunk_overlap number 1200, 200 Characters per chunk, and how many each repeats of the one before.
attach_to deployment Give the tool to this deployment once the index is built.

It makes, in order: a drive for Qdrant's data, a replica group running Qdrant (its official image, pinned; 1 CPU core, 2 GiB), waited for until it is healthy, its endpoint inside the workspace, waited for until it answers; then an indexing run — text, Markdown, HTML, reStructuredText, CSV, JSON and PDF read, cut into overlapping chunks at paragraph and sentence ends, embedded by your deployment in batches, stored with their file and place — waited for until it finished successfully; the function; and last the tool on the deployment.

Indexing again replaces what changed: a chunk's id follows from its file and place, and chunks of removed or shortened files are deleted at the end. A collection made with another embedding model is refused rather than mixed. Install it again under another name to index again. The function's timeout is 30 s; before the index is built it answers that nothing is indexed yet.

Training and fine-tuning#

Fine-tune with LoRA#

Teach a model your data: LoRA adapters trained on a GPU, checked on held-out examples, merged, converted to GGUF and registered as an Eos model of the workspace — ready to deploy, or deployed at once.

It makes a drive (named after the installation; or writes to the one you give) and five runs, each made once the one before it has finished successfully — the installation's page shows the step it is on:

Run What it does Image Asks
<name>-prepare The dataset read and made into chat examples; a share held out, never trained on (data/train.jsonl, data/eval.jsonl, data/summary.json). Python 3.12 + datasets 4.0.0 2 CPU cores, 8 GiB
<name>-train LoRA on every linear layer (rank r, alpha 2r, dropout 0.05), cosine schedule, bf16 where the GPU has it, gradient checkpointing; the loss on the answers only; an epoch's checkpoint kept, so a run started again goes on from it (adapter/). PyTorch 2.8 (CUDA 12.8) + Transformers 4.56.2, PEFT 0.17.1 1 GPU of at least gpu_memory_gb, 8 cores, 32 GiB
<name>-evaluate The base and the fine-tuned model's loss and perplexity on the held-out examples (eval/report.md, eval/report.json). as train as train
<name>-merge The adapter merged into the base weights, saved as safetensors (merged/). as train as train
<name>-convert Converted to GGUF by llama.cpp and quantised (gguf/model-<quantisation>.gguf). llama.cpp (its full image) 4 cores, 32 GiB

Then the model <name> (GGUF, from the drive's gguf/), and — with Deploy the result — a deployment of it.

Input Type Default Description
base_source choice huggingface huggingface: a repository, downloaded by the runs. eos: a model of the workspace fetched from Hugging Face in safetensors (its weights read from its drive).
base_model string Qwen/Qwen3-1.7B The repository (owner/name).
base_revision string A commit, tag or branch (main by default).
eos_model model With eos.
dataset_source choice huggingface huggingface or drive.
dataset, dataset_split string train A dataset on the hub (HuggingFaceH4/no_robots) and its split.
dataset_drive, dataset_path drive, string A file — JSONL, JSON, CSV or Parquet — or a folder of one kind, on a drive.
hf_token secret key token For a gated or private model or dataset.
epochs, learning_rate, lora_rank number 2, 0.0002, 16
max_length number 2048 Tokens; an example longer is cut (one whose answer is cut off entirely is left out).
eval_percent number 5 Held out (1 to 30 %; at least one example, at most half).
max_examples number 0 The most examples used (0: all).
gpu_memory_gb number 24 The training GPU's memory at least.
quantization choice Q4_K_M Q4_K_M, Q5_K_M, Q8_0 or F16.
output_drive drive Write here instead of a drive made for it.
deploy bool false Deploy the result.

The dataset's records may be: messages (chat turns, role and content) or ShareGPT's conversations (from, value); prompt with completion or response; instruction (with input) and output (Alpaca); question and answer; or text, trained whole. A system field is kept. In a conversation the model learns the last answer; what comes before is context. Records of no known shape are skipped and counted (data/summary.json); fewer than 10 examples is refused.

Sizes. LoRA on a model of 1–3 B parameters needs about 16 GB of GPU memory, 7–8 B about 24 GB at 2,048 tokens. Each step may take: prepare 4 h, train 72 h, evaluate, merge and convert 8 h each; past it the install fails.

If a run fails — out of memory, a gated model without a token, a dataset of no known shape — the install fails: its runs and the drive it made are removed, and the installation says which step failed and why (the run's state and reason). Read a run's log while it runs: Runs → -train → Log.

Batch inference#

Batch inference#

Run a model over every item of a file and keep every answer: one run reads JSONL or JSON (objects), CSV (a header) or text (one prompt a line) from a drive, makes each item a prompt with your template, asks the model, and writes results.jsonl to the output drive — {"id", "input", "output"} per item, or {"id", "input", "error"} when that item failed — in the input's order. The installation is ready once the run has finished successfully.

Input Type Default Description
model_source choice deployment deployment: an Eos deployment, called at its address inside the workspace (http://deploy-<name>:8000/v1, no key), concurrency requests at once; its tools are not used. vllm: a Hugging Face model loaded by vLLM in the run itself (offline batching, the same vLLM image Eos serves with), on gpus GPUs.
deployment deployment With deployment. Keep it at one replica at least while the run goes.
hf_model, hf_token, gpus string, secret, number 1 GPU With vllm.
input_drive, input_file drive, string required The file on the drive.
prompt_template text {text} Each item's fields in braces: Summarise in one sentence: {text}; a CSV's columns by name. A literal brace is doubled ({{).
system_prompt text A system message for every request.
id_field string A field naming each item in the results (unique); its line number otherwise.
max_tokens, temperature number 512, 0
output_drive, output_file drive, string made; results.jsonl Where the results go.

An item without a field the template names is written with its error, not asked. Answers already written are kept: a run started again (its machine lost) asks only what is left. A deployment that does not answer at once (a replica starting) is asked again, up to six times.

Evaluation#

Evaluate a model#

Measure a deployment on standard benchmarks with EleutherAI's lm-evaluation-harness (0.4.9): one run against the deployment's API inside the workspace, then a report on the output drive under eval/: report.md (each task's metrics and their standard error), report.json, and the harness's results.json.

Input Type Default Description
deployment deployment required What is evaluated.
api choice chat chat (/v1/chat/completions, the chat template applied): generation tasks — gsm8k, ifeval, bbh_cot_zeroshot… completions (/v1/completions): also multiple-choice tasks scored by log-likelihood — arc_easy, hellaswag, mmlu —; needs log-probabilities (vLLM answers them; llama.cpp's server does not) and tokenizer.
tasks string gsm8k Comma-separated task names.
tokenizer string With completions: the model's Hugging Face repository.
limit number 200 Examples per task (0: all — gsm8k has 1,319, MMLU 14,042).
num_fewshot number 0 Few-shot examples (0: the task's default).
concurrency number 4 Requests at once.
output_drive, report_dir drive, string made; eval Where the report goes.

A limit makes it an estimate, not a benchmark's score. The deployment serves the requests like any others: keep it at one replica at least. The installation is ready once the run has finished successfully; install it again under another name to evaluate again.

Apps and services#

Open WebUI#

Open WebUI, the open-source chat app, on your machines, talking to Eos through its gateway with a workspace API key: every deployment the key may reach is a model in its list, with your tools, budgets and usage counted as for any client. Its accounts, chats and uploads are kept on a drive.

Input Type Default Description
gateway_url string (URL) https://inference.astralyx.cloud/v1 The Eos gateway: the hosted one, or yours (a deployment's page, Gateway → Base URL) to keep requests on your machines.
api_key secret key api_key A Credential holding a workspace API key (ak-…).
exposure choice workspace workspace: reachable only inside the workspace. edge: opened on the workspace's edge machines like any endpoint opened there — reachable wherever they are, possibly the internet.
name_shown string Open WebUI Its name in the app.

It makes a drive (<name>-data), a replica group running Open WebUI (its image, pinned; 2 CPU cores, 4 GiB; telemetry off, community sharing off, Ollama off), waited for until it is healthy, and its endpoint: inside the workspace at http://<name>:8080 — from its runs, notebooks and environments, or from your computer through an environment (astra ssh <environment> -L 8080:<name>:8080, then http://localhost:8080) — or, with edge, at the endpoint's address on the edge machines.

The first account made becomes its administrator: make yours at once. Accounts made afterwards wait for the administrator's approval.