← All reports · → 1-page meeting brief

DualMap Benchmark — semantic mapping for Line 2

26 August 2026 · RTX 4050 Laptop 6 GB · Replica ×8 + TUM RGB-D · precision sweep fp32→fp4 · ground-truth label audit

SUMMARYWhat was measured

DualMap (Eku127/DualMap, RA-L 2025, Apache-2.0) is the open-vocabulary semantic mapper selected for the speech-driven navigation line. This report benchmarks it end-to-end on a 6 GB laptop GPU: 5 numeric precisions × 9 scenes × 2 datasets, geometric accuracy against the Replica GT mesh, a face-level ground-truth audit of its semantic labels, and a head-to-head against VLFM's BLIP-2 scorer.

Verdict. fp16 is the production configuration: zero feature loss, 11% faster than fp32, 9/9 scenes green. Geometry is sub-centimeter (median 0.6 cm). Semantic labels are the weak layer — 72% GT-correct on YOLO-named objects, and offline CLIP relabeling recovers only ~38% (top-1) of unknowns — but every object sits at the right location, so navigation queries resolve to correct goals far more often than strict label accuracy suggests.

SETUPConfiguration

RESULT 1The semantic map itself

DualMap semantic map of Replica room0, top-down and 3D views
Replica room0: 62 objects across 20 classes. Layout matches the real scene — sofa group with pillows, rugs, wall blinds, shelf unit and coffee maker on the west wall.
Open-vocabulary queries hitting the correct objects
Open-vocabulary queries. Note the last panel: "a place to put flowers" hits an object YOLO never named — its CLIP feature alone carries the semantics. That mechanism is also what makes offline relabeling possible.

RESULT 2Precision sweep — fp16 wins outright

PrecisionCLIP feature fidelity (cos vs fp32)Full pipelineVerdict
fp321.00000.69 s/frame (room0)baseline
fp161.00000.61 s/frame · 9/9 scenesproduction default
bf160.9991detectors unsupported (ultralytics)no benefit over fp16
int8 (visual tower)0.99899/9 scenes · +6% slowerVRAM escape hatch only
fp4 / nf40.9903crashes without extra fixesrejected

Why int8 is slower, not faster

Per-component timing over all 9 scenes: CLIP 0.118→0.126 s (+7%), full frame 0.492→0.519 s (+6%). bitsandbytes pays a quant/dequant transform on every forward pass; at DualMap's real workload (per-object crops, batch 1–10) that overhead exceeds the int8 arithmetic gain, while fp16 rides the tensor cores at full speed. int8's only win is model residency (0.43→0.34 GB). On Jetson the correct speed path is TensorRT (compile-time INT8 calibration), not runtime bnb.

int8 map quality is genuinely equivalent though: 382 vs 376 total objects across 9 scenes, and the only class flip in room0 was one object oscillating kettle ↔ coffee maker — semantic neighbors at the same location.

RESULT 3Generalization — 9/9 scenes, including a real camera

SceneObjectsClassesUnknowns/frame (fp16)Top classes found
room0571826%0.61pillows×9, sofa×4, carpet×4
room1442020%0.47pillows×7, nightstand×2, lamp×2
room2381429%0.48chair×8, vase×4, blinds×2
office0–427–6011–1511–47%0.44–0.58chair×11/×12, desk, trash can
TUM fr1_desk (real camera)18917%0.37keyboard×3, computer×3, tv×2, mouse

The class distributions track scene type — bedrooms yield nightstands and pillows, offices yield a dozen chairs and desks, and the real TUM desk sequence yields exactly what sits on that desk: keyboards, monitors, a mouse. The real-camera run (sensor noise, hand-held trajectory) was also the fastest and had one of the lowest unknown rates — the D455 integration risk is low.

RESULT 4Geometric accuracy — sub-centimeter

Every object point cloud checked against the Replica GT mesh (954k-vertex KDTree): sofa, table, vase, lamp, chair, pillows all land at median 0.6 cm / p90 0.8 cm from the true surface. For a navigation stack whose goal tolerance is 25 cm and whose robot is 70 cm long, the semantic map is the most accurate link in the chain — the error budget is dominated by SLAM drift and planner tolerance, not by DualMap.

RESULT 5Ground-truth label audit — the honest numbers

Ground truth semantic instances vs DualMap predicted labels, colored by correctness
Left: ground-truth semantic instances (47 objects, face-level Replica annotations). Right: every DualMap object graded against GT — green = label correct (31) · orange = top-3 / semantic neighbor (3) · red = wrong, annotated pred→GT (11) · gray = structural fragments (17). Object-level strict accuracy 31/45 = 69%; the red mistakes are readable at a glance (coffee maker→candle, trash can→basket, beam→lamp) and all sit at the correct position.

Using the face-level Replica semantic mesh (92 GT instances, via the vMAP data package), each DualMap object was assigned its true class by nearest-face vote. Two findings:

Label sourceGT-correctNotes
YOLO-World names (43 objects)72% top-1most misses are semantic neighbors: window→blinds, carpet→floor, vase→plant-stand, AC→vent
Offline CLIP relabel of unknowns (19)37% top-1 · 53% top-3object-level only (excl. 6 structural fragments): 38% / 62%

An earlier self-consistency estimate (84% top-1) overestimated relabeling quality: named objects are the easy samples YOLO already recognizes, while unknowns are hard by construction (small, occluded, background-contaminated crops — e.g. a ceiling lamp crop matching "beam"). The corrected claim: zero-cost relabeling recovers ~4 in 10 unknowns at top-1, ~6 in 10 with top-3 aliases.

A follow-up grading of the 14 "wrong" pairs shows the errors are overwhelmingly co-located confusions (vase→indoor-plant, ceiling→lamp, pillows→sofa — position-anchored grading means mistakes can only occur between things at the same spot): only one pair (coffee maker→candle) would misdirect a navigation query, putting navigation-equivalent accuracy near 96% against 69% strict. Crucially, all 19 unknowns sit at 0.01 m from their GT instance — the errors are purely in naming, never in position. With top-3 aliases stored in the grounding registry, a query for the right concept still resolves to the right place in most cases. Raising strict label accuracy further requires per-object captioning (Florence-2-class, +1 GB VRAM) — queued as a workstation/Orin upgrade, and the reason ConceptGraphs-style pipelines cost the VRAM they do.

RESULT 6Versus VLFM's scorer

BLIP-2 vs MobileCLIP frame-level value maps on room0
Same trajectory, same queries: BLIP-2 ITM (VLFM's brain, top) vs MobileCLIP-S2 (DualMap's brain, bottom). Green star = DualMap's object-level hit.

RESULT 7Voice-to-navigation demo — and what is real vs. oracle

The full spoken-language loop was assembled and exercised in simulation (MolmoSpaces procedural house, MuJoCo, the product Nav2/safety chain):

mic (5 s windows) → faster-whisper STT (zh+en, 2.3 s CPU) → /semantic_nav/query
  → grounding node (registry lookup + standoff) → /goal_pose
  → Nav2 MPPI → mux → stack_safety → robot

LinkStatusEvidence
Speech → text (English)verifiedspoken "go to the dining table" transcribed correctly; JFK test clip verbatim
Speech → text (Chinese)verified「去南边绿色标记」grounded at 0.85 (typed + spoken paths share the topic)
Text → goal groundingverified"dining table" → diningtable object, 1 m standoff goal; nonsense speech correctly rejected as unmatched
Goal → navigationverified"go to the bed": Nav2 SUCCEEDED across the house; demo1 marker runs SUCCEEDED (zh + en)
All four concurrently on this laptopnot viablesee compute verdict below

Honest architecture note: which layer is oracle

In this demo the semantic layer is an oracle: object positions come from a registry generated offline from scene ground truth — the robot never "saw" the dining table. The geometric layer is real: the occupancy map is built live from the simulated Mid-360, and Nav2 plans through unknown space and avoids obstacles it discovers en route. The demo therefore exercises stages 2–3 of the real pipeline:

Swapping the registry for live DualMap queries (runner_ros.py) makes stage 1 real; a target not yet mapped then returns unmatched — which is exactly where the VLFM-style frontier exploration mode (Result 6) takes over.

Compute verdict → offline demo required

At full LiDAR fidelity the XPS 15 cannot host the concurrent demo: the MolmoSpaces run's occupancy map produced zero messages in 150 s while odometry was ready at 6 s — the 1,411-geom house × 20k-ray CPU raycast starved the mapping pipeline. Follow-up fix: reducing the simulated Mid-360 to 5,000 rays/frame (the Demo5 precedent) restored a steady 10 Hz map at load 6/20 cores — the laptop can run the demo in this "lite" mode (slower Fast-LIO convergence, readiness ~140 s). The official demo still targets the lab workstation at full fidelity; checklist: repo + data/benchmarks/molmospaces (7.5 GB) + composed world + registry + faster-whisper-small (460 MB, pre-downloaded) + a USB microphone; all sim-only, no robot required.

RESULT 8SCAN-Planner vs Nav2 in the same house — a layer-by-layer autopsy

Control experiment: the same MolmoSpaces house, the same goals, both navigation backends. Nav2 MPPI reached the bed (SUCCEEDED); SCAN-Planner refused the dining-table goal outright (zero output, no log) and crawled 18 cm in 60 s toward an open-space goal. SCAN itself is healthy — in its home arena (demo4) the identical command moves the robot within 5 s. Sparse LiDAR is not the cause (18.5k healthy points/frame at 5k rays). The divergence is architectural:

LayerNav2 (works)SCAN (fails here)
Mapoctomap 2D slice (z 0.2–1.0 m) — obstacles are flattened without growing: a table occupies exactly its own footprintfull-height 3D voxels; z_down inflation grows floating surfaces downward by a body height
Inflationplanar 0.30 mcylinder r=0.25 + obstacles_inflation_z_down=0.4
The killer math0.9 m doorway − 2×0.30 = 0.3 m corridor survivestabletop underside ~0.70 m inflated 0.4 m down → 0.30 m, inside the A* travel layer; chair seats (~0.45 m) → 0.05 m, solid to the floor → furniture's floating surfaces all land in the 0.4–0.8 m kill band
Global plannerNavFn finds the 0.3 m gapA* faces furniture minefields; the table-side goal itself sits in one → no solution → silent zero
Local controlMPPI samples thousands of rollouts and takes the least-bad — it moves even without a perfect solution (the observed squeeze-through behavior)B-spline optimizer needs an initial path; none → zero; weak → 18 cm/min crawling
RecoveryBT recoveries (clear/spin/backup)none (aerial-planner lineage) — stuck is forever

Why demo4/demo5 never exposed this: those arenas contain no floating surfaces — walls, floors, ramps only — so z_down=0.4 was free safety margin there and is lethal in a furnished home. The fix direction is known (z_down 0.4→0.15, radius 0.25→0.20) but the real lesson is fit-for-terrain: flat furnished floors are Nav2's home turf; SCAN's 3D traversability earns its keep on stairs, ramps and multi-level scenes (Demo5). The voice demo therefore stays on the Nav2 backend.

RESULT 9DualMap on REAL robot data — the oracle gap closes

The Go2W's own recordings (pit_turn bag, 2026-08-10: Mid-360 + D455 RGB, 39 s) were pushed through the full pipeline offline: FAST-LIVO2 replay produced 459 poses + registered world cloud; since no depth stream was recorded, depth was synthesized by projecting the accumulated LiDAR cloud (317k pts) into each camera frame through the koide3 extrinsics — the projection lands on RGB contours (stairs, railing, mats) confirming the whole calibration chain. 457 synthetic RGB-D frames → DualMap dataset mode: 92 keyframes in 45 s, 6 map objects.

LiDAR-synthesized depth overlaid on real RGB
Real RGB (left), LiDAR-synthesized depth (right), overlays (bottom): stair treads and railings align — extrinsics × FAST-LIVO2 poses × projection all check out. 59% pixel coverage from a 2 cm-voxel accumulated cloud.
DualMap semantic map built from real robot data
The resulting semantic map of the pit arena: the pit ring, walls and stair corridor in the layout cloud; "the stairs" query hits the correctly-named staircase object (0.56); box/railing/robot queries resolve to unnamed-but-located objects.

Honest notes: only 6 objects (a 39 s turning clip sees little; 4 are unknowns — lab clutter outside the class list), and this is offline replay, not live. But stage-1 "oracle" status is now demonstrably closable with data the robot already collects; the remaining step is the live runner_ros wiring. Next recording session should enable the D455 depth stream (one launch flag) to skip synthesis entirely.

NEXTDeployment recipe & open items

Honest limits