Skip to content

Embeddings for retrieval (RAG)#

You serve an embedding model, BGE-M3 (multilingual, 1024 dimensions, inputs up to 8K tokens), on CPUs, and build a small retrieval-augmented answerer: embed your documents, find the passages closest to a question, and ask a chat deployment to answer from them. Everything — documents, vectors, questions and answers — stays on your machines.

What you need:

1. Deploy the embedding model on CPUs#

The library's embedding category has BGE-M3, Nomic Embed Text, mxbai-embed-large, Snowflake Arctic Embed, all-MiniLM, Granite Embedding and Qwen3 Embedding. An embedding deployment serves /v1/embeddings only.

  1. Eos → Models → Library, Category: embedding, open BGE-M3, click the Q4_K_M cell (438 MB) and press Deploy.
  2. Name: embed. Runs on: CPUs only (small GGUF model).
  3. Clear Scale to zero when idle; Replicas at least 1, at most 2.
  4. Press Deploy. Each replica asks for 2 CPU cores and uses the machine's idle cores beyond them.
$ astra inference add bge-m3:567m
$ astra inference deploy bge-m3-567m-q4-k-m --name embed --cpu --min 1 --max 2
$ curl -fsS -X POST "$CONSOLE/clusters/main/models" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d '{"catalog": "bge-m3:567m"}' | jq -r .metadata.name
bge-m3-567m-q4-k-m
$ curl -fsS -X POST "$API/deployments" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' \
    -d '{"metadata": {"name": "embed"}, "spec": {"model": "bge-m3-567m-q4-k-m", "accelerator": "cpu", "replicas": {"min": 1, "max": 2}}}' > /dev/null

The deployment's page says Serves embeddings (/v1/embeddings).

2. Try it#

Eos → Playground → Embeddings, pick embed, type a few texts one per line and press Embed. The table shows each vector's dimensions and norm, and a cosine similarity matrix: related texts score higher.

Embeddings in the Playground

3. Embed your documents#

index.py
import json, os
import numpy as np
from openai import OpenAI

client = OpenAI(base_url=os.environ["EOS_URL"], api_key=os.environ["ASTRAEUS_API_KEY"])

def chunks(text, size=300, overlap=50):   # about 400 tokens: within llama.cpp's batch
    words = text.split()
    step = size - overlap
    for i in range(0, max(len(words) - overlap, 1), step):
        yield " ".join(words[i:i + size])

passages = []
for name in os.listdir("docs"):
    with open(os.path.join("docs", name), encoding="utf-8") as f:
        passages += [{"source": name, "text": c} for c in chunks(f.read())]

vectors = []
for i in range(0, len(passages), 32):                       # 32 passages per request
    batch = [p["text"] for p in passages[i:i + 32]]
    out = client.embeddings.create(model="embed", input=batch)
    vectors += [d.embedding for d in out.data]

v = np.array(vectors, dtype=np.float32)
v /= np.linalg.norm(v, axis=1, keepdims=True)                # unit length: dot product = cosine
np.save("index.npy", v)
json.dump(passages, open("passages.json", "w"))
print(f"{len(passages)} passages, {v.shape[1]} dimensions")
$ python index.py
412 passages, 1024 dimensions

4. Retrieve and answer#

ask.py
import json, os, sys
import numpy as np
from openai import OpenAI

client = OpenAI(base_url=os.environ["EOS_URL"], api_key=os.environ["ASTRAEUS_API_KEY"])
v = np.load("index.npy")
passages = json.load(open("passages.json"))

question = sys.argv[1]
q = np.array(client.embeddings.create(model="embed", input=[question]).data[0].embedding, dtype=np.float32)
q /= np.linalg.norm(q)
top = np.argsort(v @ q)[::-1][:4]
context = "\n\n".join(f"[{passages[i]['source']}] {passages[i]['text']}" for i in top)

reply = client.chat.completions.create(
    model="chat",
    temperature=0.1,
    messages=[
        {"role": "system", "content": "Answer only from the context. Cite sources in brackets. If the context does not say, say so."},
        {"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"},
    ],
)
print(reply.choices[0].message.content)
$ python ask.py "How long are audit logs kept?"

Passages of 300 words stay under llama.cpp's default physical batch of 512 tokens, which bounds one embedding input. For longer inputs, deploy with engine_args: ["--batch-size", "8192", "--ubatch-size", "8192"]. Four such passages fit easily in the chat deployment's context.

5. Check that it worked#

  • curl -sS $EOS_URL/embeddings -H "Authorization: Bearer $ASTRAEUS_API_KEY" -H "Content-Type: application/json" -d '{"model": "embed", "input": "hello"}' | jq '.data[0].embedding | length' prints 1024.
  • A chat request to embed is refused by the engine: an embedding deployment serves embeddings only.
  • Organisation → Usage → Tokens by model shows bge-m3-567m-q4-k-m with prompt tokens only.

Variation: embed a large corpus in a batch run#

For millions of passages, skip the deployment: a run can mount the model's weights and embed with its own code, on as many workers as you like. With "engine": true, the run gets the model's own engine beside it and your code calls it as above, at OPENAI_BASE_URL, without a key. The engine runs as a deployment's replica would, on one GPU:

embed-corpus.json
{
  "metadata": {"name": "embed-corpus"},
  "spec": {
    "model": {"name": "bge-m3-567m-q4-k-m", "engine": true},
    "task_template": {
      "image": "python:3.12-slim",
      "command": "sh",
      "args": ["-c", "pip install --quiet openai numpy && python /data/index.py"],
      "requested_resources": {"cpu_cores": 2, "memory_bytes": 4294967296, "node_selection": {"mode": "Any"}},
      "datavolume_refs": [{"name": "docs", "mount_path": "/data", "mode": "ReadWrite"}]
    }
  }
}

In index.py, use OpenAI() (it reads OPENAI_BASE_URL and OPENAI_API_KEY) and model=os.environ["OPENAI_MODEL"]. The run completes when your workers do, and its engine stops with it. See Batch runs that read a model.

Clean up#

Delete the embed deployment, and the model to free its weights.