Files
sum-parts-test/scripts/final_eval.sh
T
nbrightandClaude Opus 5 f39b093106 Record results so far and fix the memory blowup in the palette decode
Adds STATUS.md as the handoff document: benchmark numbers, the bare-earth
metrics that actually matter for this project, the Korean-data domain gap that
retraining will not fix, and what to do on the 24 GB machine.

The memory problem was in how a prediction's colours were turned back into
class indices. Every consumer built an (N, 13, 3) float64 temporary:

    d = ((rgb[:, None, :] - COLOR_MAP[None, :, :]) ** 2).sum(axis=2)

That is ~250 MB of intermediates per 800k-point tile, several live at once, and
a full 4.7M-point block pushes it into gigabytes. main.py writes exact palette
entries, so an exact hash lookup resolves nearly every point with no large
temporary; only leftovers fall back to a chunked distance search. Peak RSS on a
470k-point tile drops to 61 MB. Extracted to sumparts_palette.py and shared by
coarse_eval.py and split_by_class.py.

Also from this round:

- patch_cm_mutation.sh: ConfusionMatrix.update() rewrote the caller's pred
  tensor in place, folding every ignore_index point into class num_classes-1.
  test() saves its visualization from that same tensor afterwards, so an
  unlabelled tile came out 100% wall and the model looked degenerate when it
  was not.
- patch_class_mask.sh: SUMPARTS_MASK_CLASSES drops known-absent classes from
  the argmax. Measured on Seosan and it does not help - the runner-up for
  "water" is "wall", not "terrain" - but the experiment is worth keeping.
- split_by_class.py now writes .ply alongside .obj. A vertex-only OBJ has zero
  faces and most viewers render nothing, which is why the first export looked
  broken.
- verify_outputs.sh reads exported files back with a parser, so "here are your
  files" can be checked rather than asserted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 09:07:31 +09:00

85 lines
2.7 KiB
Bash

#!/usr/bin/env bash
# SUM Parts - final evaluation of the trained checkpoint
#
# Two passes, because the splits differ in what they can tell you:
#
# val : labeled (0..12) -> produces real numbers you can quote locally
# test : label = -1 everywhere (blind set) -> produces predictions only.
# The authors score it; see the README's "send predictions to our
# email for local assessment".
#
# Requires patch_unlabeled_test.sh and patch_val_mode.sh to have run, otherwise
# the test pass dies in ConfusionMatrix on the -1 placeholders and mode=val dies
# with UnboundLocalError on `epoch`.
#
# No voxel_max override here on purpose. The cfg already validates with
# voxel_max: null (whole tiles), which is what we want for the final number --
# training capped it only to keep the allocator inside VRAM. Passing
# `dataset.val.voxel_max=null` on the command line does NOT work: it arrives as
# the string "null" and crop_pc then compares int >= str.
set -uo pipefail
CONDA_ROOT="$HOME/miniconda3"
SEG="$HOME/sum-parts/semantic_segmentation/PointNeXt_bundle/examples/segmentation"
DATA="$HOME/sum-parts/data/face_labeling/texsp_pcl"
OUT="$HOME/sum-parts/runs/final_eval"
source "$CONDA_ROOT/etc/profile.d/conda.sh"
conda activate sumparts
export WANDB_MODE=disabled WANDB_SILENT=true CUDA_HOME="$CONDA_PREFIX"
export PYTORCH_CUDA_ALLOC_CONF="garbage_collection_threshold:0.7,max_split_size_mb:128"
mkdir -p "$OUT"
# The cfg has to match the architecture that wrote the checkpoint; hardcoding
# one here loads the weights into the wrong model and torch raises on the
# state_dict.
source "$SCRIPTS/resolve_ckpt.sh"
echo "data : $DATA"
echo
cd "$SEG"
run_mode() {
local mode="$1" log="$OUT/${1}.log"
echo "=== mode=$mode ==="
set +e
python -u main.py \
--cfg "../../cfgs/sumv2_triangle/${CKPT_CFG}.yaml" \
mode="$mode" \
--pretrained_path "$CKPT" \
dataset.common.data_root="$DATA" \
wandb.use_wandb=False \
val_batch_size=1 \
> "$log" 2>&1
local rc=$?
set -e
if [ $rc -eq 0 ]; then
echo " ok"
else
echo " FAILED rc=$rc"
tail -12 "$log"
fi
grep -aE 'val_oa|test_oa|iou per cls|Best ckpt' "$log" | tail -6
echo
return $rc
}
# val first: this is the number we can actually stand behind locally
run_mode val
val_rc=$?
# test: predictions only, no score possible
run_mode test
test_rc=$?
echo "=== prediction files ==="
find "$SEG/log/sumv2_triangle" -name '*_pred.ply' -newermt '-30 minutes' \
-printf '%p (%s bytes)\n' 2>/dev/null | tail -12
echo
echo "logs in $OUT"
[ $val_rc -eq 0 ] && echo "FINAL EVAL: val OK" || echo "FINAL EVAL: val FAILED"
[ $test_rc -eq 0 ] && echo "FINAL EVAL: test OK" || echo "FINAL EVAL: test FAILED"