Adds STATUS.md as the handoff document: benchmark numbers, the bare-earth
metrics that actually matter for this project, the Korean-data domain gap that
retraining will not fix, and what to do on the 24 GB machine.
The memory problem was in how a prediction's colours were turned back into
class indices. Every consumer built an (N, 13, 3) float64 temporary:
d = ((rgb[:, None, :] - COLOR_MAP[None, :, :]) ** 2).sum(axis=2)
That is ~250 MB of intermediates per 800k-point tile, several live at once, and
a full 4.7M-point block pushes it into gigabytes. main.py writes exact palette
entries, so an exact hash lookup resolves nearly every point with no large
temporary; only leftovers fall back to a chunked distance search. Peak RSS on a
470k-point tile drops to 61 MB. Extracted to sumparts_palette.py and shared by
coarse_eval.py and split_by_class.py.
Also from this round:
- patch_cm_mutation.sh: ConfusionMatrix.update() rewrote the caller's pred
tensor in place, folding every ignore_index point into class num_classes-1.
test() saves its visualization from that same tensor afterwards, so an
unlabelled tile came out 100% wall and the model looked degenerate when it
was not.
- patch_class_mask.sh: SUMPARTS_MASK_CLASSES drops known-absent classes from
the argmax. Measured on Seosan and it does not help - the runner-up for
"water" is "wall", not "terrain" - but the experiment is worth keeping.
- split_by_class.py now writes .ply alongside .obj. A vertex-only OBJ has zero
faces and most viewers render nothing, which is why the first export looked
broken.
- verify_outputs.sh reads exported files back with a parser, so "here are your
files" can be checked rather than asserted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
53 lines
1.8 KiB
Bash
53 lines
1.8 KiB
Bash
#!/usr/bin/env bash
|
|
# SUM Parts - find the newest checkpoint and the cfg that matches it
|
|
#
|
|
# The eval scripts used to hardcode pointnet.yaml. Point them at a PointVector
|
|
# checkpoint and the weights load into the wrong architecture:
|
|
# RuntimeError: Error(s) in loading state_dict for BaseSeg
|
|
#
|
|
# main.py bakes the model name into the run directory, so the cfg can be read
|
|
# back out of the checkpoint path instead of guessed.
|
|
#
|
|
# Source it, don't run it:
|
|
# source scripts/resolve_ckpt.sh # newest checkpoint of any model
|
|
# CKPT_MATCH=pointvector source scripts/resolve_ckpt.sh
|
|
#
|
|
# Sets: CKPT, CKPT_CFG, CKPT_NAME
|
|
|
|
LOGROOT="$HOME/sum-parts/semantic_segmentation/PointNeXt_bundle/examples/segmentation/log/sumv2_triangle"
|
|
CKPT_MATCH="${CKPT_MATCH:-}"
|
|
|
|
if [ -n "$CKPT_MATCH" ]; then
|
|
CKPT=$(find "$LOGROOT" -name "*${CKPT_MATCH}*_ckpt_best.pth" -printf '%T@ %p\n' 2>/dev/null \
|
|
| sort -rn | head -1 | cut -d' ' -f2-)
|
|
else
|
|
CKPT=$(find "$LOGROOT" -name '*_ckpt_best.pth' -printf '%T@ %p\n' 2>/dev/null \
|
|
| sort -rn | head -1 | cut -d' ' -f2-)
|
|
fi
|
|
|
|
if [ -z "$CKPT" ]; then
|
|
echo "resolve_ckpt: no checkpoint found under $LOGROOT" >&2
|
|
return 1 2>/dev/null || exit 1
|
|
fi
|
|
|
|
CKPT_NAME=$(basename "$CKPT")
|
|
|
|
# run names look like:
|
|
# sumv2_triangle-train-<model>-ngpus1-<stamp>-<uuid>_ckpt_best.pth
|
|
# longest names first so pointnet++msg wins over pointnet
|
|
CKPT_CFG=""
|
|
for m in pointvector-xl pointnext-xl "pointnet++msg" pointnet; do
|
|
case "$CKPT_NAME" in
|
|
*"-${m}-"*) CKPT_CFG="$m"; break ;;
|
|
esac
|
|
done
|
|
|
|
if [ -z "$CKPT_CFG" ]; then
|
|
echo "resolve_ckpt: cannot tell which model wrote $CKPT_NAME" >&2
|
|
return 1 2>/dev/null || exit 1
|
|
fi
|
|
|
|
export CKPT CKPT_CFG CKPT_NAME
|
|
echo "checkpoint: $CKPT_NAME"
|
|
echo "cfg : ${CKPT_CFG}.yaml"
|