CLI#
astra inference manages Eos from a terminal: browse the library, add and
pull models, deploy them and make API keys. It works in the organisation,
workspace and cluster you chose with astra use
(Install the CLI). The console and the
API do everything the CLI does, and more.
$ astra login
$ astra use acme/vision@main
$ astra inference library llama --fits
$ astra inference add llama3.2:3b
$ astra inference deploy llama3-2-3b-q4-k-m --name chat --min 0 --max 2
$ astra inference keys create my-app --deployment chat
Commands#
astra inference library [words…]#
The model library, judged against your machines.
| Flag | Default | Description |
|---|---|---|
--category |
Repeatable: chat, code, reasoning, vision, embedding, tools, small. |
|
--publisher |
Repeatable. | |
--size |
Repeatable: le3, 3-9, 9-20, 20-40, 40plus, or a size such as 7b. |
|
--engine |
Repeatable: llama.cpp, vllm. |
|
--license |
permissive or restricted. |
|
--hide-gated |
off | Leave out gated families. |
--machine |
every machine | Repeatable: judge against these machines. |
--fits |
off | Only what fits. |
--sort |
popular |
popular, newest, smallest, name. |
--cluster |
the configured one | |
--json |
off | The raw answer. |
Each size is shown with its best variant, its size and a fit word: 1 GPU,
N GPUs, CPU, CPU (slow), no fit, or ?.
astra inference show <family>#
Every variant of a family: VARIANT ENGINE SIZE MEMORY FIT MACHINES. The
default variant is marked *; one you added says (added as <name>).
Flags: --machine (repeatable), --cluster, --json.
astra inference add <reference>#
Add a library entry: family:size-tag (llama3.2:3b-q4_k_m) or
family:size (its default variant).
| Flag | Default | Description |
|---|---|---|
--name |
<family>-<size>-<tag> |
The model's name in the workspace. |
--credential |
A credential holding a Hugging Face token under token, for a gated model. |
|
--cluster |
the configured one |
astra inference models#
The library's entries (MODEL PARAMS FORMAT ENGINE SIZE GPU MEM), then the
workspace's models (MODEL STATE SIZE ON REASON). --mine shows only the
workspace's.
astra inference pull <model>#
Download a model's weights onto a machine. --machine (default: the one
with a copy, then a GPU that fits, then the most room).
astra inference deploy <model>#
| Flag | Default | Description |
|---|---|---|
--name |
the model's name | The deployment's name. |
--cpu |
off | CPU only (llama.cpp, GGUF). |
--min |
1 |
Fewest replicas; 0 scales to zero. |
--max |
1 |
Most replicas. |
--context |
the engine's default | Context length, tokens. |
--gpus |
the fewest that hold it | GPUs per replica: 1, 2, 4 or 8. |
The engine, requests at once, GPU vendor, idle minutes, engine arguments and priority are set in the console or the API.
astra inference deployments#
DEPLOYMENT MODEL KIND STATE REPLICAS GATEWAY. REPLICAS is
ready/min-max; GATEWAY the first gateway's base URL. A deployment that
is not Ready has its reason on the next line.
astra inference keys#
| Command | Description |
|---|---|
keys create <name> |
--deployment (repeatable; none: every deployment), --expires (RFC 3339). Prints the key once on standard output. |
keys list |
KEY PREFIX DEPLOYMENTS LAST USED EXPIRES. |
keys revoke <name> |
Revoke it. |
Runs with a model#
astra astraeus run reads a model too:
| Flag | Description |
|---|---|
--model <model> |
Mount the model's weights read-only (at /model). |
--model-path <path> |
Mount them there instead. Needs --model. |
--with-engine |
Run the model's engine beside the run; the workers get OPENAI_BASE_URL, OPENAI_MODEL and OPENAI_API_KEY. Needs --model. |