Reproduces the SUM Parts (CVPR 2025) face-labeling benchmark on a single consumer GPU, then applies it to drone-photogrammetry road survey meshes. Verified on RTX 3060 12GB / WSL2 Ubuntu 22.04 / CUDA 11.8 / torch 2.0.1: - CUDA extensions build (pointnet2_batch, pointops, chamfer_dist, emd, subsampling) - PointNet 100 epochs reaches mIoU 17.19, matching the paper's reported 15.1 - OBJ -> PLY conversion round-trips through the model and yields per-point predictions Four upstream source patches, all idempotent, originals preserved: - numpy aliases removed in 1.24 (np.long etc.) and collections ABCs moved in python 3.10 - the blind test split ships label = -1, which crashed ConfusionMatrix - mode=val referenced `epoch` before assignment Documents the traps that cost the most time, including VRAM overflow silently falling back to host RAM on WSL2 (25-100x slowdown, no OOM) and the colour scale mismatch between r/g/b float32 and red/green/blue uint8. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
33 lines
831 B
Bash
33 lines
831 B
Bash
#!/usr/bin/env bash
|
|
# SUM Parts - stop the unattended training run
|
|
#
|
|
# Kills the watchdog first so it does not treat the dying trainer as a crash
|
|
# and immediately resume it.
|
|
set -uo pipefail
|
|
|
|
if ! pgrep -f train_watchdog.sh > /dev/null && ! pgrep -f "main.py" > /dev/null; then
|
|
echo "nothing running"
|
|
exit 0
|
|
fi
|
|
|
|
echo "stopping watchdog:"
|
|
pgrep -af train_watchdog.sh || true
|
|
pkill -f train_watchdog.sh || true
|
|
sleep 2
|
|
|
|
echo "stopping trainer:"
|
|
pgrep -af "main.py" || true
|
|
pkill -f "examples/segmentation/main.py" || true
|
|
sleep 3
|
|
|
|
if pgrep -f "main.py" > /dev/null; then
|
|
echo "still alive, sending SIGKILL"
|
|
pkill -9 -f "examples/segmentation/main.py" || true
|
|
fi
|
|
|
|
echo
|
|
echo "remaining:"
|
|
pgrep -af "train_watchdog.sh|main.py" || echo " clean"
|
|
echo
|
|
echo "checkpoints are kept -- relaunch resumes from the latest one."
|