Skip to content

CLI#

astra inference manages Eos from a terminal: browse the library, add and pull models, deploy them and make API keys. It works in the organisation, workspace and cluster you chose with astra use (Install the CLI). The console and the API do everything the CLI does, and more.

$ astra login
$ astra use acme/vision@main
$ astra inference library llama --fits
$ astra inference add llama3.2:3b
$ astra inference deploy llama3-2-3b-q4-k-m --name chat --min 0 --max 2
$ astra inference keys create my-app --deployment chat

Commands#

astra inference library [words…]#

The model library, judged against your machines.

Flag Default Description
--category Repeatable: chat, code, reasoning, vision, embedding, tools, small.
--publisher Repeatable.
--size Repeatable: le3, 3-9, 9-20, 20-40, 40plus, or a size such as 7b.
--engine Repeatable: llama.cpp, vllm.
--license permissive or restricted.
--hide-gated off Leave out gated families.
--machine every machine Repeatable: judge against these machines.
--fits off Only what fits.
--sort popular popular, newest, smallest, name.
--cluster the configured one
--json off The raw answer.

Each size is shown with its best variant, its size and a fit word: 1 GPU, N GPUs, CPU, CPU (slow), no fit, or ?.

astra inference show <family>#

Every variant of a family: VARIANT ENGINE SIZE MEMORY FIT MACHINES. The default variant is marked *; one you added says (added as <name>). Flags: --machine (repeatable), --cluster, --json.

astra inference add <reference>#

Add a library entry: family:size-tag (llama3.2:3b-q4_k_m) or family:size (its default variant).

Flag Default Description
--name <family>-<size>-<tag> The model's name in the workspace.
--credential A credential holding a Hugging Face token under token, for a gated model.
--cluster the configured one

astra inference models#

The library's entries (MODEL PARAMS FORMAT ENGINE SIZE GPU MEM), then the workspace's models (MODEL STATE SIZE ON REASON). --mine shows only the workspace's.

astra inference pull <model>#

Download a model's weights onto a machine. --machine (default: the one with a copy, then a GPU that fits, then the most room).

astra inference deploy <model>#

Flag Default Description
--name the model's name The deployment's name.
--cpu off CPU only (llama.cpp, GGUF).
--min 1 Fewest replicas; 0 scales to zero.
--max 1 Most replicas.
--context the engine's default Context length, tokens.
--gpus the fewest that hold it GPUs per replica: 1, 2, 4 or 8.

The engine, requests at once, GPU vendor, idle minutes, engine arguments and priority are set in the console or the API.

astra inference deployments#

DEPLOYMENT MODEL KIND STATE REPLICAS GATEWAY. REPLICAS is ready/min-max; GATEWAY the first gateway's base URL. A deployment that is not Ready has its reason on the next line.

astra inference keys#

Command Description
keys create <name> --deployment (repeatable; none: every deployment), --expires (RFC 3339). Prints the key once on standard output.
keys list KEY PREFIX DEPLOYMENTS LAST USED EXPIRES.
keys revoke <name> Revoke it.

Runs with a model#

astra astraeus run reads a model too:

Flag Description
--model <model> Mount the model's weights read-only (at /model).
--model-path <path> Mount them there instead. Needs --model.
--with-engine Run the model's engine beside the run; the workers get OPENAI_BASE_URL, OPENAI_MODEL and OPENAI_API_KEY. Needs --model.