Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6171 storiesLiPS: Lightweight Panoptic Segmentation for Resource-Constrained Robotics
arXiv:2604.00634v3 Announce Type: replace Abstract: Panoptic segmentation is a key enabler for robotic perception, as it unifies semantic understanding with obj…
Habitat-GS: A High-Fidelity Navigation Simulator with Dynamic Gaussian Splatting
arXiv:2604.12626v2 Announce Type: replace Abstract: Training embodied AI agents depends critically on the visual fidelity of simulation environments and the abi…
Proprioception-Anchored Cross-Modal Pretraining for Zero-Shot Sim-to-Real Contact-Rich Assembly
arXiv:2609.07534v2 Announce Type: replace Abstract: Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable in…
Real-World Reinforcement Learning with MPC Scaffolding for Dexterous Manipulation
arXiv:2609.14878v1 Announce Type: new Abstract: Real-world reinforcement learning (RL) offers a promising route to dexterous manipulation policies that can adap…
Language-Guided Representation Learning for Robust Cross-Sensor Material Recognition
arXiv:2609.14783v1 Announce Type: new Abstract: Robots need touch to manipulate objects safely and reliably, as many properties, such as softness, texture, and …
Beyond Dead Reckoning: A Point of View on Camera--DAS--GNSS Continuity in Road Tunnels
arXiv:2609.14748v1 Announce Type: new Abstract: Road tunnels remove satellite visibility where connected and automated vehicles still require continuous, attrib…
Learning Multi-Agent Task Assignment and Navigation in the Factory: from Simulation to Real Robots
arXiv:2609.14567v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown considerable promise for robotic decision-making, yet deploying multi-agen…
Language-Grounded Semantic Target Navigation for Autonomous Surface Vehicles
arXiv:2609.14558v1 Announce Type: new Abstract: Autonomous Surface Vehicles (ASVs) are increasingly expected to operate in ports and harbour environments, where…
Learning-Based Dynamic Obstacle Avoidance for a UAV Using Only Three Range Sensors
arXiv:2609.14426v1 Announce Type: new Abstract: We present a learning-based approach to kinodynamic online motion planning for an Unmanned Aerial Vehicle (UAV) …
BIG-CBF: Behavior-Imagination-Guided Control Barrier Function with Shared Uncertainty for Mobile Robot Navigation
arXiv:2609.14343v1 Announce Type: new Abstract: Control barrier functions (CBFs) provide a mathematically grounded framework for enforcing local collision-avoid…
VLBiMan++: Expanding the Generalization Boundary of Vision-Language Anchored One-Shot Bimanual Manipulation
arXiv:2609.14310v1 Announce Type: new Abstract: Generalizable bimanual robotic manipulation requires a reusable task prior that can persist across increasingly …
Task-Specified Active Metrological Inspection with Measurement-Steered VLA Manipulation and Deterministic Evidence Gating
arXiv:2609.14219v1 Announce Type: new Abstract: High-mix low-volume (HMLV) manufacturing requires inspection systems to adapt to changing parts, specifications,…
Vision-Force Admittance Learning for Peg Insertion into a Movable Hole
arXiv:2609.14133v1 Announce Type: new Abstract: Precise manipulation in dynamic environments, whether induced by a mobile robot base or a target with unknown mo…
ReWeight: Leveraging Human Data for VLA Post-Training via Demonstration Retrieval and Sample Weighting
arXiv:2609.13851v1 Announce Type: new Abstract: Post-training vision-language-action (VLA) models for specific robots and tasks requires in-domain demonstration…
LePlanner: An Iterative Amortized Controller For World Models
arXiv:2609.13845v1 Announce Type: new Abstract: World models trained with joint-embedding predictive architectures learn compact, structured latent representati…
Force-Aware Reinforcement Learning with Hybrid Sensorless Force Estimation for Wheeled-Legged Loco-Manipulation
arXiv:2609.13779v1 Announce Type: new Abstract: Force-controlled loco-manipulation requires a whole-body policy to coordinate locomotion and arm motion while re…
When Do Learned Priors Help Visual Inertial Estimation? A Controlled Study of Prior Integration, Calibration, Initialization, and Backend Consistency
arXiv:2609.13777v1 Announce Type: new Abstract: Learned components are increasingly integrated into geometric visual--inertial estimators to provide motion, dep…
Learning In-Hand Object Reaching to General 6D Poses
arXiv:2609.13761v1 Announce Type: new Abstract: In-hand manipulation allows multi-fingered dexterous hands to reconfigure grasped objects without releasing and …
WLA$^3$: World Latent Action Modeling for Semantics, Dynamics, and Kinematics
arXiv:2609.15870v1 Announce Type: new Abstract: Scaling generalist policy models with heterogeneous data is limited by the lack of unified, low-noise action sup…
Goal-Oriented Communications for Physical AI: Design and Testbed
arXiv:2609.15895v1 Announce Type: new Abstract: Physical AI relies on frequently-updated, latency-sensitive video stream to perceive, reason, and interact with …
SlipSense: Multimodal Tactile Learning for Low-Latency and Generalized Slip Detection
arXiv:2609.15910v1 Announce Type: new Abstract: Slip detection is fundamental to dexterous manipulation, yet existing systems often lack precise characterizatio…
DynEoMT: Learning Object Dynamicity from Online Segmentation Queries
arXiv:2609.14466v1 Announce Type: cross Abstract: Video segmentation models recognize and track objects over time, but they do not indicate whether each segment…
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
arXiv:2609.14973v1 Announce Type: cross Abstract: We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and pre…
LG-VLN: A Zero-Shot Vision-and-Language Navigation Framework with LangGraph State Orchestration
arXiv:2609.15098v1 Announce Type: cross Abstract: Continuous-environment vision-and-language navigation (VLN-CE) requires interpreting natural-language instruct…
GRAVA: Grounded Reasoning-to-Action Representation and Learning for Autonomous Driving
arXiv:2609.15169v1 Announce Type: cross Abstract: Driving vision-language-action (VLA) models increasingly reason before acting, but their intermediate reasonin…
Learning Agile Flight Maneuvers: Deep SE(3) Motion Planning and Control for Quadrotors
arXiv:2209.11097v2 Announce Type: replace Abstract: Agile flights of autonomous quadrotors in cluttered environments require constrained motion planning and con…
Trust-Region Neural Moving Horizon Estimation for Robots
arXiv:2309.05955v5 Announce Type: replace Abstract: Accurate disturbance estimation is essential for safe robot operations. The recently proposed neural moving …
Zonal RL-RRT: Integrated RL-RRT Path Planning with Collision Probability and Zone Connectivity
arXiv:2410.24205v2 Announce Type: replace Abstract: Path planning in complex environments poses significant challenges, particularly in achieving time efficienc…
Is Semantic SLAM Ready for Embedded Systems ? A Comparative Survey
arXiv:2505.12384v2 Announce Type: replace Abstract: Semantic SLAM holds promise for robust robot navigation in complex environments, but its practicality on emb…
ROVER: Robust Loop Closure Verification with Trajectory Prior in Repetitive Environments
arXiv:2508.13488v2 Announce Type: replace Abstract: Loop closure detection is important for simultaneous localization and mapping (SLAM), which associates curre…