Record results so far and fix the memory blowup in the palette decode
Adds STATUS.md as the handoff document: benchmark numbers, the bare-earth
metrics that actually matter for this project, the Korean-data domain gap that
retraining will not fix, and what to do on the 24 GB machine.
The memory problem was in how a prediction's colours were turned back into
class indices. Every consumer built an (N, 13, 3) float64 temporary:
d = ((rgb[:, None, :] - COLOR_MAP[None, :, :]) ** 2).sum(axis=2)
That is ~250 MB of intermediates per 800k-point tile, several live at once, and
a full 4.7M-point block pushes it into gigabytes. main.py writes exact palette
entries, so an exact hash lookup resolves nearly every point with no large
temporary; only leftovers fall back to a chunked distance search. Peak RSS on a
470k-point tile drops to 61 MB. Extracted to sumparts_palette.py and shared by
coarse_eval.py and split_by_class.py.
Also from this round:
- patch_cm_mutation.sh: ConfusionMatrix.update() rewrote the caller's pred
tensor in place, folding every ignore_index point into class num_classes-1.
test() saves its visualization from that same tensor afterwards, so an
unlabelled tile came out 100% wall and the model looked degenerate when it
was not.
- patch_class_mask.sh: SUMPARTS_MASK_CLASSES drops known-absent classes from
the argmax. Measured on Seosan and it does not help - the runner-up for
"water" is "wall", not "terrain" - but the experiment is worth keeping.
- split_by_class.py now writes .ply alongside .obj. A vertex-only OBJ has zero
faces and most viewers render nothing, which is why the first export looked
broken.
- verify_outputs.sh reads exported files back with a parser, so "here are your
files" can be checked rather than asserted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -31,11 +31,11 @@ export PYTORCH_CUDA_ALLOC_CONF="garbage_collection_threshold:0.7,max_split_size_
|
||||
|
||||
mkdir -p "$OUT"
|
||||
|
||||
CKPT=$(find "$SEG/log/sumv2_triangle" -name '*_ckpt_best.pth' -printf '%T@ %p\n' \
|
||||
| sort -rn | head -1 | cut -d' ' -f2-)
|
||||
[ -n "$CKPT" ] || { echo "error: no checkpoint found" >&2; exit 1; }
|
||||
# The cfg has to match the architecture that wrote the checkpoint; hardcoding
|
||||
# one here loads the weights into the wrong model and torch raises on the
|
||||
# state_dict.
|
||||
source "$SCRIPTS/resolve_ckpt.sh"
|
||||
|
||||
echo "checkpoint: $CKPT"
|
||||
echo "data : $DATA"
|
||||
echo
|
||||
|
||||
@@ -46,7 +46,7 @@ run_mode() {
|
||||
echo "=== mode=$mode ==="
|
||||
set +e
|
||||
python -u main.py \
|
||||
--cfg ../../cfgs/sumv2_triangle/pointnet.yaml \
|
||||
--cfg "../../cfgs/sumv2_triangle/${CKPT_CFG}.yaml" \
|
||||
mode="$mode" \
|
||||
--pretrained_path "$CKPT" \
|
||||
dataset.common.data_root="$DATA" \
|
||||
|
||||
Reference in New Issue
Block a user