Sharing a GPU#
By default a worker takes its GPUs whole: no other work is placed on them while it runs. Small workloads — a speech-to-text model, a voice, a small embedding model — use a fraction of a GPU's memory. A GPU request can instead share the GPU by memory: the worker says how much GPU memory it takes, and other shared work is placed on the same GPU while the sum stays within what the GPU offers.
"requested_resources": {
"cpu_cores": 2,
"memory_bytes": 4294967296,
"gpu_requests": { "count": 1, "shared": true, "memory_gb": 2.5 }
}
| Field | Meaning |
|---|---|
gpu_requests.shared |
Share each GPU with other shared work. |
gpu_requests.memory_gb |
The GPU memory the worker takes on each GPU, GB. Required with shared, and only with it. |
The other GPU fields (vendor, models, healthy_only, ids) work as
for whole GPUs. policy: All cannot be shared.
How shared work is placed#
- What a GPU offers is its memory less what something outside Astraeus holds (a desktop, a browser): an 8 GB RTX 4060 whose desktop holds 1 GB offers 7 GB. The machine's page shows both.
- Shared workers are placed on a GPU while the sum of their
memory_gb, with the new one, is at most what it offers. Among GPUs that fit, the fullest is chosen, so shared work packs onto as few GPUs as it fits and leaves the others whole for work that needs them. - A GPU is either whole or shared, never both: a worker asking for whole GPUs does not land on a GPU shared work holds, and shared work does not land on a GPU held whole.
- When a sharer ends, its memory is free for the next; the GPU is whole again once the last one ends.
- On a Windows PC (WSL2) every container sees all of the PC's GPUs: shared work there takes its memory on each of them.
A worker that does not fit waits, and its pending reason says why:
GPU 0 of majin: 6.5 of 7 GB taken by shared work; it needs 1.5
GPU 0 of majin offers 7 of its 8 GB (1 GB held outside Astraeus); it needs 7.5
What is not enforced#
Consumer and most data-center GPUs have no partitioning a container can
be held to: every sharer sees the whole GPU. memory_gb is a promise
the workload keeps, by its own sizing — an inference engine that
allocates what it needs and no more. A workload that takes more than it
said can leave its neighbours without memory, and they fail to allocate.
The worker receives its share as ASTRAEUS_GPU_MEMORY_GB, to size itself
by.
Compute is shared as well: sharers run at the same time and each is slower while the others are busy. Sharing suits workloads that are idle most of the time, or light — speech, small models, a notebook.
Eos models choose their share themselves: see Voice on one GPU.
Seeing who shares a GPU#
The machine's page lists each GPU's sharers and the memory each takes
(shared: listen 2.2 GB, voice 4.2 GB — 6.4 GB taken), and its workers
list says 1 · shared 2.2 GB in the GPUs column. Through the API,
GET /machines/{name}/workers gives each worker's gpus and, when shared,
gpu_shared_gb.
In the usage report#
In the machine usage report a shared GPU counts for each sharer by its share of the GPU's memory: a worker taking 2 GB of an 8 GB GPU for an hour is a quarter of a GPU-hour, with a quarter of the GPU's use and energy.