Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6171 storiesRoboCaf\'e in the Open: Interaction Continuity in Long-Term Public Human-Robot Interaction
arXiv:2609.27475v1 Announce Type: new Abstract: As robots remain in public spaces over extended periods, they must maintain interaction continuity by preserving…
Behavior-Aligned Action Tokenization for Robot Policy Learning
arXiv:2609.27513v1 Announce Type: new Abstract: Autoregressive robot policies learn continuous control by predicting discrete action tokens from observations. D…
NavProbe: Evidence-Grounded Reasoning with Active Memory Retrieval for Zero-Shot Navigation
arXiv:2609.27526v1 Announce Type: new Abstract: Long-horizon navigation requires an agent to revise its intermediate objectives as evidence accumulates. Full vi…
Behaviora - A Conceptual Architecture for External and Internal Behavior of Robots and Agents
arXiv:2609.27536v1 Announce Type: new Abstract: Behaviora is a preliminary conceptual architecture for representing agent and robot behavior, external and inter…
Gray-Box Model Predictive Control for Articulated Dump Trucks via Gaussian Process Learning of Sideslip
arXiv:2609.27597v1 Announce Type: new Abstract: The growing demand for automation in the mining industry, particularly for the autonomous operation of articulat…
RegenHarness: A Robot Agent Harness with Evidence-Gated Recursive Self-Improvement
arXiv:2609.27612v1 Announce Type: new Abstract: Long-horizon robot execution requires a clear distinction between a model's proposal, a controller's termination…
InternW0: A Foundational Physical World Model for Efficient Real-World Interactions
arXiv:2609.27656v1 Announce Type: new Abstract: Physical intelligence requires more than predicting how the world may evolve: predictions must remain actionable…
GLoTouch: Global-to-Local Haptic Perception Using a Parallel Gripper for Object Search, Recognition, and Grasping Without External Vision
arXiv:2609.27695v1 Announce Type: new Abstract: Perceiving objects in the environment is a fundamental capability of autonomous robots. In dark or low-light env…
DAVIO: Dense Monocular-Inertial SLAM with Feed-Forward Initialization and Pose-Conditioned Mapping
arXiv:2609.27702v1 Announce Type: new Abstract: A camera and an IMU are the minimal sensor setup for metric localization and dense mapping, yet classical visual…
Wave-Robust Passive AUV Localization Using FP-MUSIC
arXiv:2609.27712v1 Announce Type: new Abstract: Localizing an autonomous underwater vehicle without pre-deployed seabed transponders, or direct access to onboar…
Safe Multi-Robot Coordination via VLM-LLM Reasoning and Reachability Analysis
arXiv:2609.27816v1 Announce Type: new Abstract: Safe coordination in heterogeneous machine-to-machine (M2M) robotic systems is challenging when robots differ in…
Remote Surfaces at Your Fingertips: Electrovibration-Based Tactile Feedback for Robot Teleoperation via Touchscreen Interfaces
arXiv:2609.27938v1 Announce Type: new Abstract: Enabling operators to perceive and interact with remote environments naturally is a fundamental challenge in rob…
AeRSoM: An Aerial Rigid-Soft Integrated Manipulator for Contact-Rich Manipulation
arXiv:2609.28044v1 Announce Type: new Abstract: Contact-rich aerial manipulation remains fundamentally challenging because interaction forces are directly trans…
DEAL-Grasp: Decoupled Alignment Representation for Geometry-Aware Dexterous Grasp Generation
arXiv:2609.28131v1 Announce Type: new Abstract: Synthesizing realistic articulated hand-object interactions is a fundamental problem in virtual reality, embodie…
DAVIS: A Depth-Only End-to-End Active-Vision Framework for Humanoid Soccer Skills
arXiv:2609.28175v1 Announce Type: new Abstract: Humanoid soccer contact skills require more than producing high-impact foot-ball contacts: the robot must close …
Large-Scale Geometric Map-Based Localization of UAVs in GNSS-Denied Urban Environments
arXiv:2609.28225v1 Announce Type: new Abstract: Unmanned aerial vehicles (UAVs) operating in GNSS-denied urban environments require alternative methods for posi…
Generalizable Robotic Insertion with World Models
arXiv:2609.28258v1 Announce Type: new Abstract: Robotic assembly in high-mixture settings requires adaptable systems that can handle diverse parts, yet current …
Talk2Escape: Conversational Grounding for Vision-and-Language Navigation
arXiv:2609.28296v1 Announce Type: new Abstract: While Vision-and-Language Navigation (VLN) has demonstrated remarkable success, the prevailing single-turn parad…
VGM-VS: Rethinking Visual Geometry Model for High-Precision Visual Servoing
arXiv:2609.28312v1 Announce Type: new Abstract: We present VGM-VS, a visual servoing method built on a pretrained feed-forward visual geometry model. Given the …
Beyond Future Prediction: Denoising as Generative Adaptation for Robot Control
arXiv:2609.28339v1 Announce Type: new Abstract: Pretrained generative Diffusion Transformers (DiTs) capture rich pixel-level visual and language-conditioned str…
Amplify: A Lightweight Library for Reproducible Nonlinear Programming Problems in Robotics
arXiv:2609.28377v1 Announce Type: new Abstract: Optimization problems (OPs) are key to solving many challenging research problems in robotics. However, reproduc…
Tractable Reinforcement Learning for Full Class of Signal Temporal Logic Specifications Using Spatiotemporal Tube Reward
arXiv:2609.28396v1 Announce Type: new Abstract: This paper addresses the control problem for robotic systems, including non-holonomic and underactuated platform…
Where Should I Join? Robot Group Joining via Language-Guided Goal Prediction
arXiv:2609.28467v1 Announce Type: new Abstract: Social navigation typically assumes a specified goal and focuses on reaching it while respecting social conventi…
Pro-Bench: Prompt-Robust Open-Vocabulary Visual Grounding Across Real-World Heterogeneous Environments
arXiv:2609.27076v1 Announce Type: cross Abstract: Open-vocabulary visual grounding enables robots to localise task-relevant entities from natural-language queri…
Surgical Kinematics from Monocular Video with Learned Articulated Motion Constraints
arXiv:2609.27227v1 Announce Type: cross Abstract: Objective assessment of robotic surgery uses instrument kinematics, which must be reconstructed when only vide…
Geometry-Conditioned Visual Place Recognition in Natural Environments
arXiv:2609.27370v1 Announce Type: cross Abstract: Visual Place Recognition (VPR) in natural environments remains challenging due to repetitive vegetation, spars…
Frozen Flows Forget: Diagnosing and Restoring Lost Motion in a Latent-flow World Model
arXiv:2609.28414v1 Announce Type: cross Abstract: Latent world models that integrate a flow in a frozen self supervised latent space train stably and cheaply, y…
A Scalable Multi-Robot Framework for Decentralized and Asynchronous Perception-Action-Communication Loops
arXiv:2309.10164v3 Announce Type: replace Abstract: We develop a decentralized Perception-Action-Communication (PAC) system for multi-robot teams that enables t…
Robust Trajectory Tracking of Autonomous Surface Vehicle via Lie Algebraic Online MPC
arXiv:2511.18683v3 Announce Type: replace Abstract: Autonomous surface vehicles (ASVs) are influenced by environmental disturbances such as wind and waves, maki…
Vision-Based Safe Human-Robot Collaboration with Uncertainty Guarantees
arXiv:2604.15221v3 Announce Type: replace Abstract: Safe human-robot collaboration (HRC) requires accurate human pose estimation and motion prediction to preven…