Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6273 storiesSurgVIL: Scaling Surgical Robot Imitation Learning with Open-source Surgical Videos
arXiv:2608.16058v1 Announce Type: new Abstract: Learning-based surgical robot autonomy requires large-scale demonstrations with synchronized videos and robot ac…
Tabletop Pen Manipulation With a Vision-Guided 4-DoF Arm
arXiv:2608.15968v1 Announce Type: new Abstract: Low-cost four-degree-of-freedom (DoF) arms are among the most accessible robotic platforms. But they are, in the…
Rotate Disks to Reach Farther: Design and Modeling of a Novel Reconfigurable Tendon Driven Manipulator
arXiv:2608.15946v1 Announce Type: new Abstract: Rerouting the tendon path in tendon driven continuum manipulators (TDCMs) enables a broad range of deformation m…
Tactile Sim2Real without Tactile Simulation via Bottlenecked Latent Reconstruction
arXiv:2608.15897v1 Announce Type: new Abstract: Robot sensor designs, particularly tactile sensors, are highly diverse and evolve rapidly. Modeling each sensor …
Grouping Auction-Consensus Algorithm for Decentralized Task Allocation in Multi-Robot Systems
arXiv:2608.15884v1 Announce Type: new Abstract: Decentralized multi-robot task allocation (MRTA) is essential for scalable and resilient autonomous systems. The…
SCORE: Shape-Conforming Regions for Flight in Enclosed, Degraded Environments
arXiv:2608.15289v1 Announce Type: new Abstract: Autonomous UAVs enter enclosed environments such as caves and collapsed structures that confine the vehicle and …
Max-Q Selective Imitation for Human-in-the-Loop Online Robot Learning
arXiv:2608.15088v1 Announce Type: new Abstract: Human-in-the-loop (HIL) online reinforcement learning for real robots must absorb human interventions quickly wh…
MotionGS-SLAM: Event-Modulated Gaussian Splatting for Motion-Blur Robust SLAM
arXiv:2608.15024v1 Announce Type: new Abstract: Current Vision-based SLAM systems fail catastrophically when motion blur corrupts the visual input, as they atte…
HP2-SLAM: Adaptive Hybrid ICP for Robust and Efficient LiDAR SLAM
arXiv:2608.14996v1 Announce Type: new Abstract: Achieving robustness, accuracy, and efficiency simultaneously remains a central challenge in light detection and…
Evidence of Absence: Cross-Modal Abductive Risk Perception to Sustain World Models When Vision Fails
arXiv:2608.14952v1 Announce Type: new Abstract: A structured world-state (entities, relations, context, and predictive cues) is designed to preserve prediction-…
Geometry-Aware Online Mapping for 3D Gaussian Splatting SLAM
arXiv:2608.14902v1 Announce Type: new Abstract: Recent 3D Gaussian Splatting (3DGS) has enabled efficient photorealistic view synthesis and is rapidly being ado…
Modeling and Control of an Eel-Inspired Soft Robot for Design Optimization
arXiv:2608.14860v1 Announce Type: new Abstract: Anguilliform locomotion is a highly efficient swimming mode; the advent of new materials for soft robots enables…
MISTac: A Vision-Based Tactile Sensor for Minimally Invasive Surgery
arXiv:2608.14772v1 Announce Type: new Abstract: Minimally invasive and robot-assisted surgery offer many advantages over traditional open surgery, but deprive s…
SkillComposer: Learning Reusable Skills for Natural-Language Robot Programming
arXiv:2608.14944v1 Announce Type: new Abstract: Natural-language interfaces can lower the barrier to programming robots, but existing systems struggle when user…
ForceU-VLA: A Force-Aware Vision-Language-Action Model for Embodied Ultrasound Scanning
arXiv:2608.15009v1 Announce Type: new Abstract: Embodied intelligent ultrasound scanning enables the automation and standardization of the ultrasound examinatio…
PACE: Phase-Progress-Aware Credit for Long-Horizon Embodied Manipulation
arXiv:2608.15026v1 Announce Type: new Abstract: Post-training of vision-language-action (VLA) models typically relies on expert demonstrations and policy intera…
Contact Modes Are Strata: What Geometric Structure Buys in Discrete-Continuous Planning
arXiv:2608.15541v1 Announce Type: new Abstract: Contact-rich manipulation poses a discrete question and a continuous one at once, namely which contacts are acti…
Pre-training Visual Dexterity in Simulation
arXiv:2608.15917v1 Announce Type: new Abstract: Large-scale pre-training has made robot policy fine-tuning increasingly data-efficient, but this progress has la…
Learning Varying Physical Therapist-Patient Interactions for Robot-mediated Upper Limb Task-Specific Training
arXiv:2608.15995v1 Announce Type: new Abstract: Upper extremity motor function recovery is positively linked to Task-Specific Training (TST) and sufficient ther…
US-VLA: An Ultrasound Vision-Language-Action Model for Embodied Abdomina
arXiv:2608.16074v1 Announce Type: new Abstract: Artificial intelligence-assisted ultrasound scanning enhances diagnostic reliability and efficiency by providing…
Deep Probabilistic Indoor Gas Source Localization via Physical Dependency-Guided Sequential Inference
arXiv:2608.16221v1 Announce Type: new Abstract: Reliable gas source localization (GSL) is critical to safety in industrial and urban environments, yet remains c…
HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction
arXiv:2608.16222v1 Announce Type: new Abstract: Humanoid intelligence requires learning over an extremely diverse space of whole-body motions and physically gro…
Robot-Body-Aware Traversal Risk Graph Planning for Wheeled-Legged Robots in Complex Terrain
arXiv:2608.16433v1 Announce Type: new Abstract: Traversal Risk Graphs (TRGs) provide a compact, terrain-aware representation for global navigation, but native T…
Exposing the Long-tail in Embodied Urban Navigation via Scalable Learning from In-the-Wild Videos
arXiv:2608.16476v1 Announce Type: new Abstract: Learning embodied urban navigation policies from real-world data is constrained by the cost of task-specific dat…
Closing the Affective Loop: Multimodal Speaker-Listener Emotion-Dynamics-Aware Empathetic Social Robots
arXiv:2608.16686v1 Announce Type: cross Abstract: Empathetic social robots should respond not only to what users say, but also to how their emotions dynamically…
Language-Guided Generation for Personalized Inspection Planning
arXiv:2506.02917v2 Announce Type: replace Abstract: We propose a training-free, Vision-Language Model (VLM)-guided approach for efficiently generating trajector…
Flow Motion Policy: Manipulator Motion Planning with Flow Matching Models
arXiv:2604.07084v2 Announce Type: replace Abstract: Open-loop end-to-end neural motion planners have recently been proposed to improve motion planning for robot…
Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis
arXiv:2510.08759v3 Announce Type: replace-cross Abstract: Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is cruci…
Formalisms for Robotic Mission Specification and Execution: A Comparative Analysis
arXiv:2603.15427v2 Announce Type: replace-cross Abstract: Robots are increasingly deployed across diverse domains and designed for multi-purpose operation. As r…
Geometric Reconstruction of Extrinsic Contact Trajectories using Tactile Sensing and Proprioception for Tool Manipulation
arXiv:2606.22251v2 Announce Type: replace Abstract: Tactile sensing enables robots to perceive rich contact information at the grasp, supporting tasks such as o…