Poster
Overview
Abstract
Long-horizon target navigation requires a robot to sustain task execution across evolving observations, decisions, and physical interactions. This requires three coupled capabilities: maintaining valid scene memory, revising target beliefs under partial observability, and selecting interaction-feasible navigation endpoints. However, the state underlying each capability is only conditionally valid: scene representations become stale when objects move or disappear, unsuccessful searches alter beliefs over target locations, and geometrically convenient endpoints may still be infeasible for manipulation. To address these challenges, we present OmniNav, which formulates long-horizon navigation as continual inference over a factorized task state posterior coupling scene validity, target belief, and interaction feasibility. For representation, OmniNav incrementally constructs an updatable 3D object scene memory, preventing stale scene evidence from propagating to subsequent decisions. For exploration, it introduces an evidence-aware Bayesian belief-revision mechanism that derives dependency-aware region priors from semantic context, incorporates unsuccessful searches as negative evidence, and updates them for posterior-guided frontier selection. For interaction, OmniNav incorporates manipulation reachability and collision constraints into navigation-endpoint selection and propagates execution feedback through hierarchical closed-loop recovery. Extensive experiments demonstrate that OmniNav achieves the highest success rates among the compared methods on semantic ObjectNav and fine-grained instance navigation benchmarks, remains robust to target relocation, and improves real-world pick-and-place success from 53.3% to 71.7% over an adapted open-loop baseline.
Method
Framework
Continual inference connects scene validity, target belief, and interaction feasibility across a long-horizon task.
Overall framework of OmniNav. The system consists of three main modules: (i) a dynamic object-centric scene memory that maintains instance-level semantics, geometry, relations, posterior persistence beliefs, and an auxiliary 2D exploration-progress map; (ii) a task-grounded exploration–exploitation navigation strategy that balances semantic target reasoning with exploration progress evidence; and (iii) a reachability-aware closed-loop mobile manipulation module that selects interaction-feasible base poses and recovers from execution failures or scene changes.
Simulation
Navigation experiments
Object navigation
Fine-grained instance matching
An overview figure presents the query, candidate observations, and selected target instance.

Real world
Long-horizon deployment
Long-sequence navigation
Targets: keyboard → cola bottle → white bowl (4×)
Challenge White bowl relocation
Targets: white bottle → trash can → banana (4×)
Challenge Trash can relocation
Targets: toy → yellow stool → bottle (4×)
Challenge Bottle relocation
Standard mobile manipulation
Navigation-to-interaction demonstrations with reachability-aware endpoint selection.
Pick the cola and place it on the red plate. (3×)
Pick the bottle and place it on the red plate. (4×)
Target relocation
The target is moved after it has been observed.
Pick the apple and place it in the white bowl. (4×)
Challenge White bowl relocation
Pick the apple and place it in the white bowl. (4×)
Challenges Apple displacement · white bowl relocation · similar-fruit distractors
Pick the apple and place it in the white bowl. (4×)
Challenge Apple relocation
Interference
Stress tests with target displacement and distractors.
Find and pick the white bottle.
Challenge Small target displacement
Pick the apple and place it in the bowl.
Challenges Apple displacement
Failure case
Pick the cola and place it on the red plate.
Failure cause Red plate not recognized