Skip to content

Use the Playground#

The Playground lets people try a deployment from the console, without an API key: chat with it, compare two, embed texts, or transcribe speech. Use it to check a model before you wire it into an application, to tune its parameters, and to copy the exact request as code.

Before you begin#

  • The editor or admin role in the workspace. Viewers cannot talk to deployments.
  • A deployment of the workspace, or one another workspace shares with yours for the Playground.

Open it#

  • Eos → Playground in the sidebar, or
  • Try in the Playground on a deployment's page, or
  • Playground next to a deployment under Shared with this workspace on the Deployments page.

Pick the Deployment on the left: the deployments of this workspace, and under Shared with this workspace those shared with it (… · shared by …). Under the picker: its model, its state (N of M replicas serving, or at zero: the first message starts it) and what it can do: tools, reasoning, vision, audio, video, speech to text.

The tabs Chat, Compare, Embeddings and Transcribe appear when a deployment can serve them.

Chat#

  1. Write a System prompt (default You are a helpful assistant.).
  2. Type a message and press Send (Enter sends, Shift+Enter is a new line). The answer is written as the model writes it; Stop ends it.
  3. Under each answer: tokens in and out (and thinking), time to the first token, tokens per second and the whole time.

The Playground: a conversation with a deployment

Answers are shown as markdown — headings, lists, tables, code with a Copy button. As text shows exactly what the model wrote. Text from the model is sanitised before it is shown, and an image it links from the web is shown as a link, not loaded.

A deployment at zero starts with the first message: the page says Starting… a replica is being started for this message, and waits up to 15 minutes.

New conversation starts over; Export and Import move a conversation as a JSON file. Conversations are kept in your browser only; Astralyx keeps none.

Parameters#

Under Parameters: temperature (0 to 2), top_p, top_k, min_p, max_tokens (1 to 131072), seed, presence_penalty, frequency_penalty, repetition_penalty, logprobs (0 to 20) and up to 4 Stop sequences. An empty field is not sent: the engine's default applies. Apply a preset… offers Precise, Balanced, Creative, Deterministic, Qwen3 thinking and Qwen3 non-thinking; Save as a preset keeps your own in the browser. With logprobs, Show token probabilities under an answer shows each token's alternatives.

Reasoning#

For a model that reasons (the reasoning tag):

  • Thinking: default, on or off, sent as chat_template_kwargs.enable_thinking (Qwen3 and models like it).
  • Effort: reasoning_effort (gpt-oss): minimal, low, medium or high.

The thinking is folded above the answer (Thought).

Structured output#

Under Structured output, choose text, JSON (JSON mode) or a JSON schema (response_format), optionally strict. Answers are checked against the schema here, in the browser, and marked ✓ valid JSON, matches the schema.

Tools#

For a model that calls tools (the tools tag): Add a tool (a name, a description and the arguments' JSON Schema) or Add an example…, and choose Tool choice (auto, none, required, or one tool). When the model calls a tool, its call is shown: type the result in its box and press Send results and continue. The page never runs anything.

A vLLM deployment parses tool calls only with a tool-call parser. When it was deployed without one and the library knows the right one, the page offers Redeploy with --enable-auto-tool-choice --tool-call-parser ….

Images, audio and video#

For a model that takes them, Attach files, drop or paste them, or Record from the microphone:

Kind Tag What is sent Limits
Images vision Resized in the browser to at most 1568 px on the longer side, sent as JPEG (PNG when transparent) Files up to 40 MiB
Audio audio WAV and MP3 as they are; anything else (a recording, M4A…) converted in the browser to 16 kHz mono WAV Files up to 100 MiB before conversion
Video video MP4 or WebM, as it is 6 MiB; vLLM only

Up to 8 files at a time, and 8 MiB of attachments per request in all. Everything is sent inline in the request: the engine never fetches a URL. A model that takes text only says This model takes text only.

Which models listen and watch is in the library: for example Qwen2.5-Omni listens (vLLM, or llama.cpp with its projector), and Qwen2.5-VL, Qwen3-VL and Qwen2.5-Omni watch videos on vLLM.

Compare#

Compare sends each message to two sides, Side A and Side B: two deployments, or the same deployment with different settings. Copy side A's settings to side B starts from the same parameters.

Embeddings#

For an embedding deployment (a model with the embedding capability):

  1. Type Texts, one per line — at most 64 at once.
  2. Press Embed.
  3. The table shows each text's dimensions, norm and first values, and a Cosine similarity matrix between them. Export JSON saves the vectors.

Embeddings in the Playground, with their cosine similarities

Transcribe#

For a speech-to-text deployment (Whisper, served by vLLM):

  1. Record, or Upload audio (or drop a file). WAV, MP3, FLAC and OGG go as they are; anything else — a recording, M4A, a video's sound — is converted in the browser to 16 kHz WAV. At most 8 MiB: about four minutes of recorded speech.
  2. Choose the Language, or Detect.
  3. Optionally, Context: names and terms it should spell right (up to 4000 characters). Tick Timestamps for [m:ss.s → m:ss.s] segments.
  4. Press Transcribe. Copy or Download .txt the text.

The audio goes to the machine serving the model and is not kept.

Get the request as code#

Under every answer, Raw request and response (Raw and code for embeddings) shows the JSON request and response, and the same request as curl, Python (OpenAI SDK) and TypeScript against your gateway, with $ASTRAEUS_API_KEY. The Python version passes the fields the SDK does not know (such as top_k) in extra_body.

What passes where#

Playground messages go from your browser through the Astralyx console to the machine serving the deployment, as you; the answer is read back the same way, a piece at a time. The console keeps none of it; an answer is kept on its machine for 10 minutes after it ends. Your applications use the gateway instead, which on your machines never passes through Astralyx.

Playground tokens count toward the workspace's usage like a gateway's (a stopped answer counts what was written), and show in the side panel as N prompt · M completion tokens in this conversation.