Skip to content

Errors and limits#

This page lists Eos's limits in one place. Errors are listed with the API that returns them:

Limits#

Models#

Item Limit
Model name Lowercase letters, digits and -, at most 40 characters
Hugging Face revision A commit, 40 hexadecimal characters
Files of a model At least one; each path at most 512 characters
Chat template 64 KiB
Download progress in the fill run's log Every 10 s
Library fit judged at 8192 tokens (or the model's, when shorter)

Deployments#

Item Limit
Deployment name Lowercase letters, digits and -, at most 40 characters
Replicas 1 to 64 (max); min 0 to max
Idle minutes before scaling to zero 1 to 1440; default 15
GPUs per replica 1, 2, 4 or 8, of one machine
CPU cores per replica 1 to 1024
Context length At least 256 tokens, at most the model's; default llama.cpp ≤ 8192, vLLM ≤ 32768
Requests at once per replica 1 to 1024; default llama.cpp 4 on a GPU, 2 on a CPU; vLLM 256
Engine arguments 64, each at most 4096 characters, from the allow-list
Scale-up and scale-down At most once every 120 s
A replica counted as failing After 3 restarts without serving
deployment_down alert A failure, or 5 minutes without a serving replica (not when scaled to zero)

API keys and shares#

Item Limit
Key name, share name Lowercase letters, digits and -, at most 63 characters
Deployments listed in a key 100, and 100 shared deployments
Share note 500 characters
A revoked share stops calls Within half a minute
Key last used Recorded to the minute

Gateways#

Item Gateway on your machines Hosted gateway
Request body 16 MiB 1 MB
Non-streamed answer 64 MiB —
Connect to a replica 10 s —
Silence before an answer is cut 10 minutes —
A call — 90 s
Wait for a deployment at zero 2 minutes (configurable) Within the call's 90 s
Streaming Yes No
Unknown keys per client address — 30 per minute

Playground#

Item Limit
Attachments per request 8 MiB in all, up to 8 files
Images Files up to 40 MiB, resized to 1568 px on the longer side
Audio Files up to 100 MiB before conversion
Video 6 MiB, MP4 or WebM, vLLM only
Transcription audio 8 MiB (about four minutes of recorded speech)
Transcription context 4000 characters
Embeddings 64 texts at once
Wait for a deployment at zero 15 minutes
An answer kept on its machine 10 minutes after it ends