Split the unattended run into a GPU-free phase and a GPU phase

The target machine's card is busy with someone else's job, so a single
end-to-end script stalls on work that does not actually need a GPU.

Compiling the CUDA extensions needs nvcc, not a device, and downloading 18 GB
of data needs neither. Those are the slow parts (~50 min + ~30 min), so phase A
now runs entirely without the card:

  run_setup.sh   bootstrap, conda, extensions, patches, data      no GPU
  run_train.sh   voxel_max measurement, training, evaluation      GPU

run_setup reports the GPU but never fails on it, and verify_env.py gained
SKIP_CUDA_CHECK so import coverage still runs when no device is visible.
TORCH_CUDA_ARCH_LIST is stated rather than probed, since the card may be
unavailable at build time.

run_train waits for the GPU instead of failing when it is busy: it polls until
enough VRAM frees up (12h default), so it can be queued ahead of time. Past the
deadline it proceeds anyway and lets the measured voxel_max adapt to whatever
is actually free.

keepalive.sh now takes the phase to supervise. Replaces run_all.sh and RUN.md
with SETUP.md and TRAIN.md. Adds selfcheck.sh, which syntax-checks every script
and flags CRLF endings - a shell script with either fails at its first line,
which for an unattended weekend run means losing the weekend.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
nbright
2026-08-21 12:33:22 +09:00
co-authored by Claude Opus 5
parent 53279a6b6a
commit 3fdd7ab3f0
11 changed files with 989 additions and 758 deletions
+15 -4
View File
@@ -62,12 +62,23 @@ def main() -> int:
if root is None:
ok = False
# SKIP_CUDA_CHECK exists for the GPU-free setup phase: the extensions can be
# built and imported without a card present, and querying the device would
# fail on a machine whose GPU is absent or still occupied. Import coverage
# is unaffected -- only the device query is skipped.
skip_cuda = bool(os.environ.get("SKIP_CUDA_CHECK"))
try:
import torch
print(f"torch {torch.__version__} | cuda {torch.version.cuda} | "
f"available {torch.cuda.is_available()}")
if torch.cuda.is_available():
print(f"device: {torch.cuda.get_device_name(0)}")
line = f"torch {torch.__version__} | cuda {torch.version.cuda}"
if skip_cuda:
print(line + " | device check skipped (SKIP_CUDA_CHECK)")
else:
print(line + f" | available {torch.cuda.is_available()}")
if torch.cuda.is_available():
print(f"device: {torch.cuda.get_device_name(0)}")
else:
print("device: none visible -- fine for setup, required to train")
except Exception as e: # noqa: BLE001
print(f"torch import failed: {e}")
return 1