TensorBoard for a running training#
In this recipe you train a small model in a development environment, logging to TensorBoard on its drive, start TensorBoard beside it, and open it in your browser as an app of the environment: at an address of its own, through the machine's own connection, with no port opened on the machine. The same works in a notebook's runtime.
Before you begin#
- A machine with a GPU in your workspace, with a Data location.
- The editor or admin role.
astraon your computer, signed in (Install the CLI).
1. Make the environment, with TensorBoard installed#
The PyTorch preset has no TensorBoard: the environment's setup installs it once onto its drive.
- Open Hesperus → Environments → New environment.
- Kind: Shell. Name:
train. - Under Environment, choose PyTorch; GPUs
1. - Open Setup, and in pip packages enter
tensorboard==2.21.0. - Press Create and start.
$ astra env create train --preset pytorch --gpus 1 --pip tensorboard==2.21.0 --app TensorBoard=6006
train is pending: `astra ssh train` connects once it is ready (and waits for it)
--app names the app now, so the console lists it with an Open
button.
With ASTRAEUS_TOKEN and API as in API:
$ curl -fsS -X POST "$API/notebook-runtimes" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
-H 'content-type: application/json' \
-d '{"metadata": {"name": "train"}, "spec": {"kind": "shell", "image": "pytorch",
"resources": {"gpus": {"count": 1}}, "setup": {"pip": ["tensorboard==2.21.0"]},
"apps": [{"name": "TensorBoard", "port": 6006}]}}' | jq -r .status.state
Pending
The setup runs before the SSH server starts; its log is
/content/.hesperus/env/train/setup.log on the drive.
2. Start a training that logs to the drive#
Connect, and write the script into the drive:
In the environment:
$ mkdir -p /content/projects && cd /content/projects
$ cat > fit.py <<'EOF'
import time, torch
from torch import nn
from torch.utils.tensorboard import SummaryWriter
dev = "cuda" if torch.cuda.is_available() else "cpu"
torch.manual_seed(0)
X = torch.linspace(-3, 3, 4096, device=dev).unsqueeze(1)
y = torch.sin(X) + 0.1 * torch.randn_like(X)
model = nn.Sequential(nn.Linear(1, 64), nn.Tanh(), nn.Linear(64, 64), nn.Tanh(), nn.Linear(64, 1)).to(dev)
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
writer = SummaryWriter("/content/runs/sine")
for step in range(3000):
loss = nn.functional.mse_loss(model(X), y)
opt.zero_grad()
loss.backward()
opt.step()
if step % 10 == 0:
writer.add_scalar("loss/train", loss.item(), step)
writer.flush()
time.sleep(0.05)
print("done, final loss", loss.item())
EOF
$ nohup python fit.py > fit.log 2>&1 &
It runs about three minutes and writes its events under
/content/runs/sine.
3. Start TensorBoard#
In the same session (the setup's pip installs into your home on the
drive, so its command is in ~/.local/bin):
TensorBoard listens on the container's 127.0.0.1:6006: that is enough.
The app is reached inside the container; nothing listens on the machine.
4. Open it#
On the environment's page, under Apps, press Open on the TensorBoard row. If you did not name it at creation, press Open port…, choose the suggestion TensorBoard · 6006, keep Remember it on this environment, and press Open.
TensorBoard opens in a new tab at an address of its own. The loss/train curve grows as the training runs.
From your computer:
--print prints the address instead (it works once, within a minute).
The address signs that browser in for 8 hours, to this app only. Its owner, the people it is shared with and the workspace's admins may open it; anyone else is refused.
In a notebook instead#
In a notebook's runtime, install and start it from cells:
import subprocess
subprocess.Popen(["/content/.hesperus/python/bin/tensorboard", "--logdir", "/content/runs", "--port", "6006"])
Then, from your computer, astra env port <runtime> 6006, with the
runtime's name from the notebook's runtime menu (Runtime …). The
console's Apps section is on the page of runtimes with SSH; a
notebook's runtime without SSH has no such page, so use astra env port.