Files
nbrightandClaude Opus 5 f39b093106 Record results so far and fix the memory blowup in the palette decode
Adds STATUS.md as the handoff document: benchmark numbers, the bare-earth
metrics that actually matter for this project, the Korean-data domain gap that
retraining will not fix, and what to do on the 24 GB machine.

The memory problem was in how a prediction's colours were turned back into
class indices. Every consumer built an (N, 13, 3) float64 temporary:

    d = ((rgb[:, None, :] - COLOR_MAP[None, :, :]) ** 2).sum(axis=2)

That is ~250 MB of intermediates per 800k-point tile, several live at once, and
a full 4.7M-point block pushes it into gigabytes. main.py writes exact palette
entries, so an exact hash lookup resolves nearly every point with no large
temporary; only leftovers fall back to a chunked distance search. Peak RSS on a
470k-point tile drops to 61 MB. Extracted to sumparts_palette.py and shared by
coarse_eval.py and split_by_class.py.

Also from this round:

- patch_cm_mutation.sh: ConfusionMatrix.update() rewrote the caller's pred
  tensor in place, folding every ignore_index point into class num_classes-1.
  test() saves its visualization from that same tensor afterwards, so an
  unlabelled tile came out 100% wall and the model looked degenerate when it
  was not.
- patch_class_mask.sh: SUMPARTS_MASK_CLASSES drops known-absent classes from
  the argmax. Measured on Seosan and it does not help - the runner-up for
  "water" is "wall", not "terrain" - but the experiment is worth keeping.
- split_by_class.py now writes .ply alongside .obj. A vertex-only OBJ has zero
  faces and most viewers render nothing, which is why the first export looked
  broken.
- verify_outputs.sh reads exported files back with a parser, so "here are your
  files" can be checked rather than asserted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 09:07:31 +09:00

81 lines
3.1 KiB
Bash

#!/usr/bin/env bash
# SUM Parts - stop ConfusionMatrix.update from overwriting its caller's predictions
#
# openpoints/utils/metrics.py:
#
# if (true == self.ignore_index).sum() > 0:
# pred[true == self.ignore_index] = self.virtual_num_classes - 1
# true[true == self.ignore_index] = self.virtual_num_classes - 1
#
# Folding ignored points into the last bucket is fine for scoring, but it is
# done in place on the tensors the caller passed in. main.py's test() calls
# cm.update(pred, label) and *then* writes the visualization from that same
# pred, so the file on disk is the mutated copy, not what the model predicted.
#
# On a tile labelled entirely with the ignore class - which is exactly what an
# unlabelled tile of our own data looks like, label=0 with ignore_index=0 -
# every single point gets rewritten to class num_classes-1 (wall). The
# prediction file then reads 100% wall no matter what the network actually
# said, and the model looks degenerate when it is not.
#
# Fix: clone before masking. Scoring is unchanged; the caller's tensors survive.
#
# Idempotent.
set -euo pipefail
source "$HOME/miniconda3/etc/profile.d/conda.sh"
conda activate sumparts
REPO="${1:-$HOME/sum-parts/semantic_segmentation/PointNeXt_bundle}"
METRICS="$REPO/openpoints/utils/metrics.py"
[ -f "$METRICS" ] || { echo "error: $METRICS not found" >&2; exit 1; }
if grep -q 'SUMPARTS-NO-MUTATE' "$METRICS"; then
echo "already patched"
exit 0
fi
cp -n "$METRICS" "$METRICS.orig" 2>/dev/null || true
python - "$METRICS" <<'PY'
import sys
from pathlib import Path
p = Path(sys.argv[1])
src = p.read_text(encoding="utf-8")
old = """ true = true.flatten()
pred = pred.flatten()
if self.ignore_index is not None:
if (true == self.ignore_index).sum() > 0:
pred[true == self.ignore_index] = self.virtual_num_classes -1
true[true == self.ignore_index] = self.virtual_num_classes -1"""
new = """ # SUMPARTS-NO-MUTATE: clone before masking. These used to be written
# in place, which silently rewrote the caller's prediction tensor --
# main.py's test() saves its visualization from the same `pred` right
# after calling this, so the file on disk showed the folded values
# rather than the model's output. On a tile labelled entirely with the
# ignore class every point came out as num_classes-1.
true = true.flatten().clone()
pred = pred.flatten().clone()
if self.ignore_index is not None:
if (true == self.ignore_index).sum() > 0:
pred[true == self.ignore_index] = self.virtual_num_classes -1
true[true == self.ignore_index] = self.virtual_num_classes -1"""
if old not in src:
print("PATTERN NOT FOUND -- metrics.py differs from what this patch expects",
file=sys.stderr)
raise SystemExit(1)
p.write_text(src.replace(old, new), encoding="utf-8")
print("patched:", p)
PY
python -c "import ast,sys; ast.parse(open(sys.argv[1], encoding='utf-8').read())" "$METRICS" \
&& echo "syntax OK"
echo "PATCH DONE (original kept at $METRICS.orig)"