Split the unattended run into a GPU-free phase and a GPU phase
The target machine's card is busy with someone else's job, so a single end-to-end script stalls on work that does not actually need a GPU. Compiling the CUDA extensions needs nvcc, not a device, and downloading 18 GB of data needs neither. Those are the slow parts (~50 min + ~30 min), so phase A now runs entirely without the card: run_setup.sh bootstrap, conda, extensions, patches, data no GPU run_train.sh voxel_max measurement, training, evaluation GPU run_setup reports the GPU but never fails on it, and verify_env.py gained SKIP_CUDA_CHECK so import coverage still runs when no device is visible. TORCH_CUDA_ARCH_LIST is stated rather than probed, since the card may be unavailable at build time. run_train waits for the GPU instead of failing when it is busy: it polls until enough VRAM frees up (12h default), so it can be queued ahead of time. Past the deadline it proceeds anyway and lets the measured voxel_max adapt to whatever is actually free. keepalive.sh now takes the phase to supervise. Replaces run_all.sh and RUN.md with SETUP.md and TRAIN.md. Adds selfcheck.sh, which syntax-checks every script and flags CRLF endings - a shell script with either fails at its first line, which for an unattended weekend run means losing the weekend. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+15
-4
@@ -62,12 +62,23 @@ def main() -> int:
|
||||
if root is None:
|
||||
ok = False
|
||||
|
||||
# SKIP_CUDA_CHECK exists for the GPU-free setup phase: the extensions can be
|
||||
# built and imported without a card present, and querying the device would
|
||||
# fail on a machine whose GPU is absent or still occupied. Import coverage
|
||||
# is unaffected -- only the device query is skipped.
|
||||
skip_cuda = bool(os.environ.get("SKIP_CUDA_CHECK"))
|
||||
|
||||
try:
|
||||
import torch
|
||||
print(f"torch {torch.__version__} | cuda {torch.version.cuda} | "
|
||||
f"available {torch.cuda.is_available()}")
|
||||
if torch.cuda.is_available():
|
||||
print(f"device: {torch.cuda.get_device_name(0)}")
|
||||
line = f"torch {torch.__version__} | cuda {torch.version.cuda}"
|
||||
if skip_cuda:
|
||||
print(line + " | device check skipped (SKIP_CUDA_CHECK)")
|
||||
else:
|
||||
print(line + f" | available {torch.cuda.is_available()}")
|
||||
if torch.cuda.is_available():
|
||||
print(f"device: {torch.cuda.get_device_name(0)}")
|
||||
else:
|
||||
print("device: none visible -- fine for setup, required to train")
|
||||
except Exception as e: # noqa: BLE001
|
||||
print(f"torch import failed: {e}")
|
||||
return 1
|
||||
|
||||
Reference in New Issue
Block a user