Skip to content

Recipes#

Each recipe serves one real thing from start to finish: the choices, the specifications, the commands, what you should see, how to check it works and how to clean up. Start with the one closest to what you are doing.

  • Serve an open chat model on one GPU machine


    Llama 3.1 8B on a 24 GB GPU: check the fit, deploy, try it, call it, and the same model on vLLM.

    Uses: the library, fit, a deployment, the Playground, API keys.

  • An OpenAI-compatible endpoint for your app


    The gateway on your machine behind TLS, a key per application, and clients in Python and TypeScript with streaming and retries.

    Uses: the edge gateway, ASTRAEUS_GATEWAY_URL, scoped keys, errors.

  • A private coding assistant for your team


    Qwen2.5 Coder with a long context, called from VS Code (Continue) and Aider, with one key for the team.

    Uses: context length, parallel requests, replicas, tools.

  • Embeddings for retrieval (RAG)


    BGE-M3 on CPUs for /v1/embeddings, a small retrieval index, and answers from a chat deployment with the retrieved context.

    Uses: an embedding model, CPU replicas, batch runs that read a model.

  • Serve on AMD GPUs


    A Radeon RX 7900 XTX or an Instinct MI300X serving with llama.cpp's ROCm build or vLLM's, and which GPUs take which engine.

    Uses: gpu_vendor, ROCm and Vulkan builds, the fit per machine.

  • Fit a big model with quantization


    Llama 3.3 70B on two 48 GB GPUs: the memory arithmetic, the variant to choose, GPUs per replica, context and a quantized KV cache.

    Uses: quantization, gpus, context_length, parallel, engine arguments.

  • Multiple replicas behind one endpoint


    A deployment that scales from 1 to 4 replicas with load, a load test to watch it, rolling changes, and scale to zero overnight.

    Uses: replicas, autoscaling, balancing, idle_minutes.

  • Serve a fine-tuned model


    Take the model from Fine-tune a model on one GPU, convert it to GGUF, register it from its drive, deploy it and call it.

    Uses: Astraeus runs and drives, a model from a drive, llama.cpp or vLLM.

How the recipes are written#

  • Every step shows each way to do it, in the order Console, CLI, API. The CLI is astra inference (CLI); some settings (engine arguments, GPU vendor, a model from a drive) are console- or API-only.
  • API calls go through the console with an API token, to your workspace on a cluster. Each recipe assumes:

    $ export ASTRAEUS_TOKEN=$(cat ~/.astraeus-token)
    $ export CONSOLE=https://console.astralyx.cloud/api/v1/orgs/acme/workspaces/vision
    $ export API=$CONSOLE/clusters/main/api
    

    Replace acme, vision and main with your organisation, workspace and cluster. Create the token under Account → API tokens.

  • Calls to deployments go to a gateway with an API key of the workspace (ak-…), not the API token:

    $ export EOS_URL=http://10.0.4.12:8800/v1      # Gateway → Base URL, on the deployment's page
    $ export ASTRAEUS_API_KEY=ak-…