Skip to content

Home and alerts#

Home#

Home shows the workspace at a glance (tap its name to switch):

Tile Shows You see it with
Machines up Machines answering, of all the workspace's machines, and which are down machines:read
GPUs in use GPUs held by work, of all of them machines:read
Goodput today Productive GPU time over GPU time held since midnight, your time (Goodput) runs:read or machines:read-metrics
Open incidents Faults on the machines being dealt with (Self-healing) machine-health:read

Above them, N decisions wait for you leads to the inbox; below, your latest chats on the phone and New chat. A cluster that does not answer is said, not hidden.

Alerts#

Alerts lists, newest first, the open incidents on the workspace's machines (a red dot) and the thresholds firing now — a GPU too hot, memory errors, goodput dropping (an amber dot). It is shown when you may read machines' health (machine-health:read) or alerts (alerts:read).

Tap an incident for its timeline, oldest first: when it opened and why, the evidence that came after, each repair proposed, approved (by whom) or taken by itself, done or failed, the notes on it, and when it was resolved. A threshold shows what was measured, its value against its limit, and since when.

I'm on it#

On an incident or an alert, I'm on it tells the team you are dealing with it, with a line if you like ("Looking at the PSU"). Everyone of the organisation who sees machines' health sees it — on the list (Ana is on it) and on the item — for 30 days. I'm no longer on it takes it back. It changes nothing on the machines: it is a word to the team.

By API: GET, POST ({"subject": "incident:<cluster>:<id>", "note": "…"}) and DELETE ?subject= on /organizations/{org}/acknowledgements — see the API reference.