Semantic Navigation (Line 2) — one-page brief

2026-08-26 · DualMap benchmark · voice loop · real Go2W data · full report (appendix)

SYSTEMArchitecture — models, inputs, outputs

RGB · D455 640×480 @ 10 Hz Depth D455 stream / LiDAR-synth Pose · FAST-LIVO2 map → camera TF RGB PER-FRAME PERCEPTION · fp16 · GPU YOLO-World-L in: RGB + class list out: boxes + names boxes MobileSAM in: RGB + boxes out: refined named masks FastSAM-s in: RGB (no classes) out: proposal masks → unknowns named masks + unnamed masks, per frame MobileCLIP-S2 · image in: each mask crop out: 512-d open-vocab feature back-projection in: mask × Depth × Pose (from sensors) out: 3D point cloud in map frame DualMap core — cross-frame fusion in: {mask, name?, 512-d feature, 3D points} each frame out: objects {point cloud · feature · class · position} + layout cloud MobileCLIP-S2 · text in: "go to the dining table" out: 512-d query vector cosine matched object → 3D goal + standoff out: PoseStamped (map frame) Nav2 → mux → stack_safety → robot motion always behind the safety gate all models local · fp16 · 1.2 GB VRAM

PIPELINEReal robot data, step by step

1 · RAW RGB (D455) Raw RGB frame from the robot

39 s pit_turn bag, frame 200

2 · DEPTH FROM LIDAR Synthesized depth image

317k-pt cloud projected via calibration; 59% coverage

3 · ALIGNMENT CHECK Depth overlaid on RGB

stair treads land on pixels → chain verified

4 · OBJECT + 3D POSITION Desk detected with bbox and 3D position

DualMap object, bbox + (x,y,z)

5 · SEMANTIC MAP + QUERY Semantic top-down map of the pit

"the stairs" → staircase (0.56)

Robot's own recording → queryable map, offline, no new hardware session. 45 s mapping time.

SCENESimulation demo scene — top-down

Top-down of the MolmoSpaces house with registry objects, robot spawn and voice demo route

MolmoSpaces house val_5 (MuJoCo): 13 registry objects (blue), voice-demo targets (red), Go2W spawn (green star) and the spoken "go to the bed" route. Label positions come straight from the grounding registry — the same coordinates the robot navigates to.

R1Mapping quality — benchmark scenes

Replica room0 semantic map

Replica room0: 62 objects / 20 classes · layout matches reality.

DataObjectss/frame
Replica ×8 scenes27–600.44–0.61
TUM real camera180.37
Go2W pit bag645 s total
PrecisionQualitySpeed
fp16lossless+11%USE
int8equal−6%VRAM only
fp4−1%crashesNO

R2Ground-truth audit — labels imprecise, navigation-safe

GT instances vs DualMap predictions graded green orange red

Left: GT instances. Right: every prediction graded — 31 correct · 3 near · 11 wrong (pred→GT annotated) · gray structural. Reading the reds: because grading is position-anchored, "wrong" labels are co-located confusions (vase→indoor-plant, ceiling→lamp, pillows→sofa) — of 14 graded pairs only 1 would actually misdirect a navigation query (coffee maker→candle). Strict label accuracy 69% · navigation-equivalent accuracy ≈96% (0.6 cm positions).

R3Failure case, verified on real data

Mislabeled staircase object with bbox on clutter

Mislabeled: "staircase" points sit on clutter at the stair's foot — YOLO crop caught the stair edge, mask grabbed the neighbor.

Unnamed but correctly located object

Unnamed but located: features stay queryable — "the six wheeled robot" resolves here. Naming is fixable offline (38→62% with top-3), no map rebuild.

R4Planner check — same house, two backends

Nav2 MPPISCAN-Planner
Furnished flat floorreaches goals (voice-driven)refuses / crawls
Root cause2D slice + sampling + recoveriesz-down inflation → tabletops become floor no-go zones
VerdictNav2 owns flat homes · SCAN owns stairs/ramps · demo stays on Nav2

PLANTimeline

DONE (this week) benchmark · voice loop · real-bag map planner autopsy · laptop lite mode NEXT runner_ros online mapping (D455+LIVO2) workstation demo · record depth stream THEN live query on robot · explore mode caption-based naming upgrade LATER Orin deployment (TensorRT) stairs → SCAN re-tuning

SUMMARYBottom line