| Model | Split | Precision | Recall | F1 | mAP@0.5 | mAP@0.5:0.95 | FPS | Inference (ms) | Device |
|---|---|---|---|---|---|---|---|---|---|
| YOLO26n (exp-12) -- NYU domain-transfer | test | 0.388 | 0.428 | 0.407 | 0.388 | 0.280 | 22.0 | 45.5 | cpu (onnxruntime) |
| YOLO26n (exp-4) -- official validation | val | 0.512 | 0.399 | 0.449 | 0.398 | 0.278 | — | 0.2 | NVIDIA RTX PRO 6000 Blackwell Server Edition (training-time benchmark, not representative of on-device/browser inference) |
YOLO26n (exp-12) -- NYU domain-transfer: exp-12: fine-tuned from a COCO-pretrained YOLO26n checkpoint (not NYU-only from-scratch, superseding exp-8, see script docstring). Evaluated here against the NYU test split for a domain-transfer measurement -- exp-12 was not trained on NYU Depth V2, so this is a generalization number, not a same-distribution one.
YOLO26n (exp-4) -- official validation: The real, official validation from exp-4's own training run (100 epochs, Ultralytics 8.4.164) -- an IN-DISTRIBUTION measurement against its own val split. exp-4 replaces exp-12 (docs/11_decision_log.md, 2026-09-27/28 entries): scope expanded from 14 classes to the full 59-class insighter-indoor-balanced dataset (door rebuilt from clean indoor photos after the original source turned out to be mislabeled street/locker photos; stairs kept; window added as a genuinely new class). The lower overall mAP50 vs. exp-12's 14-class number (0.398 vs 0.626) reflects scope breadth diluting the instance-weighted average with harder, lower-support long-tail classes (book, spoon, knife, backpack, etc.) -- not a regression on any class exp-12 already covered; door in particular improved dramatically (mAP50 0.703 -> 0.932) after the source-data fix. Per-class real-world confidence thresholds (frontend/src/inference/detector.ts, ios/InSighter/Detection/DetectorService.swift) come from a SEPARATE, more rigorous held-out test-split sweep (ml/scripts/eval_test_thresholds.py, 24,790 test images never used in training or checkpoint selection), not from this val table.