← All reports · → 1-page meeting brief
26 August 2026 · RTX 4050 Laptop 6 GB · Replica ×8 + TUM RGB-D · precision sweep fp32→fp4 · ground-truth label audit
DualMap (Eku127/DualMap, RA-L 2025, Apache-2.0) is the open-vocabulary semantic mapper selected for the speech-driven navigation line. This report benchmarks it end-to-end on a 6 GB laptop GPU: 5 numeric precisions × 9 scenes × 2 datasets, geometric accuracy against the Replica GT mesh, a face-level ground-truth audit of its semantic labels, and a head-to-head against VLFM's BLIP-2 scorer.
| Precision | CLIP feature fidelity (cos vs fp32) | Full pipeline | Verdict |
|---|---|---|---|
| fp32 | 1.0000 | 0.69 s/frame (room0) | baseline |
| fp16 | 1.0000 | 0.61 s/frame · 9/9 scenes | production default |
| bf16 | 0.9991 | detectors unsupported (ultralytics) | no benefit over fp16 |
| int8 (visual tower) | 0.9989 | 9/9 scenes · +6% slower | VRAM escape hatch only |
| fp4 / nf4 | 0.9903 | crashes without extra fixes | rejected |
Per-component timing over all 9 scenes: CLIP 0.118→0.126 s (+7%), full frame 0.492→0.519 s (+6%). bitsandbytes pays a quant/dequant transform on every forward pass; at DualMap's real workload (per-object crops, batch 1–10) that overhead exceeds the int8 arithmetic gain, while fp16 rides the tensor cores at full speed. int8's only win is model residency (0.43→0.34 GB). On Jetson the correct speed path is TensorRT (compile-time INT8 calibration), not runtime bnb.
int8 map quality is genuinely equivalent though: 382 vs 376 total objects across 9 scenes, and the only class flip in room0 was one object oscillating kettle ↔ coffee maker — semantic neighbors at the same location.
| Scene | Objects | Classes | Unknown | s/frame (fp16) | Top classes found |
|---|---|---|---|---|---|
| room0 | 57 | 18 | 26% | 0.61 | pillows×9, sofa×4, carpet×4 |
| room1 | 44 | 20 | 20% | 0.47 | pillows×7, nightstand×2, lamp×2 |
| room2 | 38 | 14 | 29% | 0.48 | chair×8, vase×4, blinds×2 |
| office0–4 | 27–60 | 11–15 | 11–47% | 0.44–0.58 | chair×11/×12, desk, trash can |
| TUM fr1_desk (real camera) | 18 | 9 | 17% | 0.37 | keyboard×3, computer×3, tv×2, mouse |
The class distributions track scene type — bedrooms yield nightstands and pillows, offices yield a dozen chairs and desks, and the real TUM desk sequence yields exactly what sits on that desk: keyboards, monitors, a mouse. The real-camera run (sensor noise, hand-held trajectory) was also the fastest and had one of the lowest unknown rates — the D455 integration risk is low.
Every object point cloud checked against the Replica GT mesh (954k-vertex KDTree): sofa, table, vase, lamp, chair, pillows all land at median 0.6 cm / p90 0.8 cm from the true surface. For a navigation stack whose goal tolerance is 25 cm and whose robot is 70 cm long, the semantic map is the most accurate link in the chain — the error budget is dominated by SLAM drift and planner tolerance, not by DualMap.
Using the face-level Replica semantic mesh (92 GT instances, via the vMAP data package), each DualMap object was assigned its true class by nearest-face vote. Two findings:
| Label source | GT-correct | Notes |
|---|---|---|
| YOLO-World names (43 objects) | 72% top-1 | most misses are semantic neighbors: window→blinds, carpet→floor, vase→plant-stand, AC→vent |
| Offline CLIP relabel of unknowns (19) | 37% top-1 · 53% top-3 | object-level only (excl. 6 structural fragments): 38% / 62% |
An earlier self-consistency estimate (84% top-1) overestimated relabeling quality: named objects are the easy samples YOLO already recognizes, while unknowns are hard by construction (small, occluded, background-contaminated crops — e.g. a ceiling lamp crop matching "beam"). The corrected claim: zero-cost relabeling recovers ~4 in 10 unknowns at top-1, ~6 in 10 with top-3 aliases.
A follow-up grading of the 14 "wrong" pairs shows the errors are overwhelmingly co-located confusions (vase→indoor-plant, ceiling→lamp, pillows→sofa — position-anchored grading means mistakes can only occur between things at the same spot): only one pair (coffee maker→candle) would misdirect a navigation query, putting navigation-equivalent accuracy near 96% against 69% strict. Crucially, all 19 unknowns sit at 0.01 m from their GT instance — the errors are purely in naming, never in position. With top-3 aliases stored in the grounding registry, a query for the right concept still resolves to the right place in most cases. Raising strict label accuracy further requires per-object captioning (Florence-2-class, +1 GB VRAM) — queued as a workstation/Orin upgrade, and the reason ConceptGraphs-style pipelines cost the VRAM they do.
The full spoken-language loop was assembled and exercised in simulation (MolmoSpaces procedural house, MuJoCo, the product Nav2/safety chain):
mic (5 s windows) → faster-whisper STT (zh+en, 2.3 s CPU) → /semantic_nav/query
→ grounding node (registry lookup + standoff) → /goal_pose
→ Nav2 MPPI → mux → stack_safety → robot
| Link | Status | Evidence |
|---|---|---|
| Speech → text (English) | verified | spoken "go to the dining table" transcribed correctly; JFK test clip verbatim |
| Speech → text (Chinese) | verified | 「去南边绿色标记」grounded at 0.85 (typed + spoken paths share the topic) |
| Text → goal grounding | verified | "dining table" → diningtable object, 1 m standoff goal; nonsense speech correctly rejected as unmatched |
| Goal → navigation | verified | "go to the bed": Nav2 SUCCEEDED across the house; demo1 marker runs SUCCEEDED (zh + en) |
| All four concurrently on this laptop | not viable | see compute verdict below |
In this demo the semantic layer is an oracle: object positions come from a registry generated offline from scene ground truth — the robot never "saw" the dining table. The geometric layer is real: the occupancy map is built live from the simulated Mid-360, and Nav2 plans through unknown space and avoids obstacles it discovers en route. The demo therefore exercises stages 2–3 of the real pipeline:
Swapping the registry for live DualMap queries (runner_ros.py) makes stage 1 real; a target not yet mapped then returns unmatched — which is exactly where the VLFM-style frontier exploration mode (Result 6) takes over.
At full LiDAR fidelity the XPS 15 cannot host the concurrent demo: the MolmoSpaces run's occupancy map produced zero messages in 150 s while odometry was ready at 6 s — the 1,411-geom house × 20k-ray CPU raycast starved the mapping pipeline. Follow-up fix: reducing the simulated Mid-360 to 5,000 rays/frame (the Demo5 precedent) restored a steady 10 Hz map at load 6/20 cores — the laptop can run the demo in this "lite" mode (slower Fast-LIO convergence, readiness ~140 s). The official demo still targets the lab workstation at full fidelity; checklist: repo + data/benchmarks/molmospaces (7.5 GB) + composed world + registry + faster-whisper-small (460 MB, pre-downloaded) + a USB microphone; all sim-only, no robot required.
Control experiment: the same MolmoSpaces house, the same goals, both navigation backends. Nav2 MPPI reached the bed (SUCCEEDED); SCAN-Planner refused the dining-table goal outright (zero output, no log) and crawled 18 cm in 60 s toward an open-space goal. SCAN itself is healthy — in its home arena (demo4) the identical command moves the robot within 5 s. Sparse LiDAR is not the cause (18.5k healthy points/frame at 5k rays). The divergence is architectural:
| Layer | Nav2 (works) | SCAN (fails here) |
|---|---|---|
| Map | octomap 2D slice (z 0.2–1.0 m) — obstacles are flattened without growing: a table occupies exactly its own footprint | full-height 3D voxels; z_down inflation grows floating surfaces downward by a body height |
| Inflation | planar 0.30 m | cylinder r=0.25 + obstacles_inflation_z_down=0.4 |
| The killer math | 0.9 m doorway − 2×0.30 = 0.3 m corridor survives | tabletop underside ~0.70 m inflated 0.4 m down → 0.30 m, inside the A* travel layer; chair seats (~0.45 m) → 0.05 m, solid to the floor → furniture's floating surfaces all land in the 0.4–0.8 m kill band |
| Global planner | NavFn finds the 0.3 m gap | A* faces furniture minefields; the table-side goal itself sits in one → no solution → silent zero |
| Local control | MPPI samples thousands of rollouts and takes the least-bad — it moves even without a perfect solution (the observed squeeze-through behavior) | B-spline optimizer needs an initial path; none → zero; weak → 18 cm/min crawling |
| Recovery | BT recoveries (clear/spin/backup) | none (aerial-planner lineage) — stuck is forever |
Why demo4/demo5 never exposed this: those arenas contain no floating surfaces — walls, floors, ramps only — so z_down=0.4 was free safety margin there and is lethal in a furnished home. The fix direction is known (z_down 0.4→0.15, radius 0.25→0.20) but the real lesson is fit-for-terrain: flat furnished floors are Nav2's home turf; SCAN's 3D traversability earns its keep on stairs, ramps and multi-level scenes (Demo5). The voice demo therefore stays on the Nav2 backend.
The Go2W's own recordings (pit_turn bag, 2026-08-10: Mid-360 + D455 RGB, 39 s) were pushed through the full pipeline offline: FAST-LIVO2 replay produced 459 poses + registered world cloud; since no depth stream was recorded, depth was synthesized by projecting the accumulated LiDAR cloud (317k pts) into each camera frame through the koide3 extrinsics — the projection lands on RGB contours (stairs, railing, mats) confirming the whole calibration chain. 457 synthetic RGB-D frames → DualMap dataset mode: 92 keyframes in 45 s, 6 map objects.
Honest notes: only 6 objects (a 39 s turning clip sees little; 4 are unknowns — lab clutter outside the class list), and this is offline replay, not live. But stage-1 "oracle" status is now demonstrably closable with data the robot already collects; the remaining step is the live runner_ros wiring. Next recording session should enable the D455 depth stream (one launch flag) to skip synthesis entirely.