Semantic Navigation (Line 2) — one-page brief
2026-08-26 · DualMap benchmark · voice loop · real Go2W data · full report (appendix)
SYSTEM Architecture — models, inputs, outputs
RGB · D455
640×480 @ 10 Hz
Depth
D455 stream / LiDAR-synth
Pose · FAST-LIVO2
map → camera TF
RGB
PER-FRAME PERCEPTION · fp16 · GPU
YOLO-World-L
in: RGB + class list
out: boxes + names
boxes
MobileSAM
in: RGB + boxes
out: refined named masks
FastSAM-s
in: RGB (no classes)
out: proposal masks → unknowns
named masks + unnamed masks, per frame
MobileCLIP-S2 · image
in: each mask crop
out: 512-d open-vocab feature
back-projection
in: mask × Depth × Pose (from sensors)
out: 3D point cloud in map frame
DualMap core — cross-frame fusion
in: {mask, name?, 512-d feature, 3D points} each frame
out: objects {point cloud · feature · class · position} + layout cloud
MobileCLIP-S2 · text
in: "go to the dining table"
out: 512-d query vector
cosine
matched object → 3D goal + standoff
out: PoseStamped (map frame)
Nav2 → mux → stack_safety → robot
motion always behind the safety gate
all models local · fp16 · 1.2 GB VRAM
PIPELINE Real robot data, step by step
1 · RAW RGB (D455)
39 s pit_turn bag, frame 200
2 · DEPTH FROM LIDAR
317k-pt cloud projected via calibration; 59% coverage
3 · ALIGNMENT CHECK
stair treads land on pixels → chain verified
4 · OBJECT + 3D POSITION
DualMap object, bbox + (x,y,z)
5 · SEMANTIC MAP + QUERY
"the stairs" → staircase (0.56)
Robot's own recording → queryable map, offline, no new hardware session. 45 s mapping time.
SCENE Simulation demo scene — top-down
MolmoSpaces house val_5 (MuJoCo): 13 registry objects (blue), voice-demo targets (red),
Go2W spawn (green star) and the spoken "go to the bed" route. Label positions come straight
from the grounding registry — the same coordinates the robot navigates to.
R1 Mapping quality — benchmark scenes
Replica room0: 62 objects / 20 classes · layout matches reality.
Data Objects s/frame
Replica ×8 scenes 27–60 0.44–0.61
TUM real camera 18 0.37
Go2W pit bag 6 45 s total
Precision Quality Speed
fp16 lossless +11% USE
int8 equal −6% VRAM only
fp4 −1% crashes NO
R2 Ground-truth audit — labels imprecise, navigation-safe
Left: GT instances. Right: every prediction graded —
31 correct · 3 near ·
11 wrong (pred→GT annotated) · gray structural.
Reading the reds: because grading is position-anchored, "wrong" labels are co-located
confusions (vase→indoor-plant, ceiling→lamp, pillows→sofa) — of 14 graded pairs only
1 would actually misdirect a navigation query (coffee maker→candle).
Strict label accuracy 69% · navigation-equivalent accuracy ≈96% (0.6 cm positions).
R3 Failure case, verified on real data
Mislabeled: "staircase" points sit on clutter at the stair's foot —
YOLO crop caught the stair edge, mask grabbed the neighbor.
Unnamed but located: features stay queryable —
"the six wheeled robot" resolves here. Naming is fixable offline (38→62% with top-3), no map rebuild.
The gap: edge-tier open-vocab recognition ~70% vs 90%+ on workstation models — segmentation & localisation already good enough on our hardware.
R4 Planner check — same house, two backends
Nav2 MPPI SCAN-Planner
Furnished flat floor reaches goals (voice-driven) refuses / crawls
Root cause 2D slice + sampling + recoveries z-down inflation → tabletops become floor no-go zones
Verdict Nav2 owns flat homes · SCAN owns stairs/ramps · demo stays on Nav2
PLAN Timeline
DONE (this week)
benchmark · voice loop · real-bag map
planner autopsy · laptop lite mode
NEXT
runner_ros online mapping (D455+LIVO2)
workstation demo · record depth stream
THEN
live query on robot · explore mode
caption-based naming upgrade
LATER
Orin deployment (TensorRT)
stairs → SCAN re-tuning
SUMMARY Bottom line
DualMap validated end-to-end on our hardware — keep it ; fp16 is the only config.
Spoken "go to the dining table" drives the robot in sim, through the safety gate.
Robot's own bag → queryable semantic map — the oracle gap is demonstrably closable.
Weak link is naming, not geometry — fix is async captioning, not a rebuild.