Bachelor Thesis Prototype

See what's ahead. Hear it instantly.

InSighter is a real-time assistive-navigation prototype. It detects 59 common indoor obstacles — from people and furniture to appliances — from a live camera feed, estimates how far away each one is, and speaks a clear, spoken cue — entirely on-device, in real time.

A person walking outdoors with a guide dog, representing independent, assisted mobility
0.40
mAP@0.5, official validation
59
obstacle classes detected live
100%
on-device detection, no video upload
~0.4m
depth-model error, near range
What it detects

59 classes, tuned for indoor obstacles

Fine-tuned from a COCO-pretrained checkpoint onto the objects that matter most for everyday indoor navigation — every indoor-plausible COCO class, plus door, stairs, and window sourced and verified outside COCO.

Person

Person

Detects people ahead and announces how far away and which direction they are.

Chair

Chair

Flags seating obstacles before you walk into them, with a proximity warning up close.

Table

Table

Picks out tables and desks — common waist-height hazards that canes can miss.

Door

Door

Locates doorways so a room transition is never a surprise.

Also detected
StairsBenchCatDogBackpackUmbrellaHandbagTieSuitcaseFrisbeeSports ballBaseball batBaseball gloveSkateboardTennis racketBottleWine glassCupForkKnifeSpoonBowlBananaAppleSandwichOrangeBroccoliCarrotHot dogPizzaDonutCakeCouchPotted plantBedToiletTVLaptopMouseRemoteKeyboardCell phoneMicrowaveOvenToasterSinkRefrigeratorBookClockVaseScissorsTeddy bearHair drierToothbrushWindow
How it works

One pipeline, three responsibilities

Detection, distance estimation, and speech are each handled by the component best suited for the job — nothing is reimplemented twice across platforms.

01

Detect, on-device

A YOLO26n model runs entirely in the browser (onnxruntime-web/WASM) or on-device on iOS (CoreML) — no video ever leaves the phone.

02

Estimate distance, in Python

A monocular depth model runs server-side and streams back real distance and direction over a live WebSocket, validated against ground-truth depth data.

03

Speak, instantly

The native speech engine (Web Speech API / AVSpeechSynthesizer) reads the closest obstacle aloud — free, offline-capable, no cloud TTS.

Ready to try it live?

Grant camera access in your browser and hear real-time obstacle announcements.

Launch live demo →