Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6171 storiesReinforcement Learning for Real-Time Vision-Language-Action Policies
arXiv:2609.18207v1 Announce Type: new Abstract: Reinforcement learning fine-tuning on top of large, pretrained Vision-Language-Action (VLA) models offers promis…
UMI-Bridge: Action-Anchored Latent Alignment across Human and Robot Manipulation Data
arXiv:2609.18232v1 Announce Type: new Abstract: Real-robot demonstrations are limited, motivating the use of human manipulation data collected without robots, i…
RAFAIL: Relationship-Aware Failure Detection for Robotic Manipulation
arXiv:2609.18324v1 Announce Type: new Abstract: Detecting failures during execution is essential for reliable robotic manipulation. Vision-language models (VLMs…
GraphPoint: Semantic Entity Graphs and Point Trajectories for Compositional Robot Manipulation
arXiv:2609.18358v1 Announce Type: new Abstract: Robot manipulation policies often struggle to generalize beyond their demonstrations, even when new instructions…
Decoupling Vision, Language, and Action for Efficient Multi-Task Robot Policies
arXiv:2609.18374v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models attach an action module to a Vision-Language Model (VLM) with billions of pa…
FIERCE: From Generalist Robot Policies to Fast Specialists via Progress-Failure Feedback
arXiv:2609.18651v1 Announce Type: new Abstract: Generalist robot policies offer useful initialization, but refining compact specialists through limited physical…
Toward 3D Printable Non-Planar Electroadhesive Structures for Active Anchoring
arXiv:2609.18700v1 Announce Type: new Abstract: This paper investigates multi-material 3D printing as a method to fabricate non-planar structures with 3D-printe…
PASSAGE: Scaling Scene-Aligned Motion Learning for Perceptive Humanoid Traversal in Cluttered Environments
arXiv:2609.18732v1 Announce Type: new Abstract: Humanoid robots can step over, squeeze past, and duck under obstacles, but learning to select and coordinate the…
QMSR: Query-Conditioned Mask-wise Expert Routing for Robust Open-Vocabulary Underwater Object Retrieval
arXiv:2609.18752v1 Announce Type: new Abstract: Open-vocabulary object retrieval remains challenging in complex underwater environments. Although underwater ima…
TRACER: Adaptive Multi-Robot Social Navigation via Joint Human-Response Prediction and Interaction-Aware Replanning
arXiv:2609.18776v1 Announce Type: new Abstract: Multi-robot navigation in human-shared spaces is inherently interactive: coordinated robot motions influence how…
SOL-SLAM: Inverse Compositional Gauss-Newton Direct Registration for Fast Sonar-Only Local SLAM
arXiv:2609.18893v1 Announce Type: new Abstract: Autonomous underwater navigation typically relies on complex and expensive multi-modal sensor suites designed to…
Examining the Difference in Human Behavior Between Virtual and Real-World Human-Robot Teaming
arXiv:2609.18900v1 Announce Type: new Abstract: Prototyping and evaluating human-robot teaming (HRT) scenarios in the real-world is costly. Virtual simulation o…
"What's going to happen after I'm gone?": Parent Perspectives on Technology in Supporting Independent Living for Adults with Intellectual Disabilities
arXiv:2609.18970v1 Announce Type: new Abstract: Adults with intellectual and developmental disabilities (IDD) are increasingly transitioning from family homes t…
Learning to Stack: Cube-Stacking Imitation Learning from Virtual Reality Demonstrations
arXiv:2609.19040v1 Announce Type: new Abstract: Imitation learning is attractive for robot manipulation, but collecting demonstrations remains a bottleneck for …
ElastiQP: An Always-Feasible QP Solver for Constrained Robot Control
arXiv:2609.19080v1 Announce Type: new Abstract: As robot capabilities increase, quadratic programming (QP)-based controllers must account for a similarly increa…
Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation
arXiv:2609.19137v1 Announce Type: new Abstract: Recent advances in video generation allow robots to learn manipulation trajectories from generated videos. Howev…
REVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff
arXiv:2609.17745v1 Announce Type: cross Abstract: A central goal of autonomous reinforcement learning is continuous policy training without external resets. How…
RoboVAD: A Large Cross-Domain Evaluation Benchmark for Anomaly Detection in Robotic Arm Manipulation Videos
arXiv:2609.17843v1 Announce Type: cross Abstract: Video anomaly detection (VAD) is an actively studied task, having wide applications in typical scenarios such …
Investigating Adversarial Robustness of Heterogeneous Cooperative Perception
arXiv:2609.17856v1 Announce Type: cross Abstract: Heterogeneous cooperative perception (CP) enables connected vehicles with diverse sensor setups to share spati…
PESTO: Formally Correct Registration of LiDAR Point Clouds with Limited Overlap
arXiv:2609.18082v1 Announce Type: cross Abstract: In this paper we tackle the problem of aligning LiDAR point clouds also known as the point cloud registration …
PRISM: Predictive Representation of Interaction Style and Motion for Social Robot Navigation
arXiv:2609.18125v1 Announce Type: cross Abstract: Humans often observe others before interacting and adjust their behavior accordingly. Robot navigation in crow…
Risk-Aware World Modeling with Flow-Guided Occupancy Evolution for Selective Trajectory Planning in Automated Driving
arXiv:2609.18442v1 Announce Type: cross Abstract: Safe motion planning in automated driving requires anticipating evolving traffic risks and deciding when to re…
Towards Interaction Regulation from Human Feedback via Free Energy Minimization
arXiv:2609.18853v1 Announce Type: cross Abstract: A central challenge across control and learning is the design of mechanisms regulating the interactions betwee…
PointZero: 3D Point Track Completion for Learning Transferable 3D Dynamics
arXiv:2609.19142v1 Announce Type: cross Abstract: World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They ar…
VLEM: Real-Time 3D Vision-Language Embedding Mapping
arXiv:2508.06291v2 Announce Type: replace Abstract: Semantic scene understanding in robotics requires representations that are both metric-accurate and queryabl…
Real-Time Maneuver Planning for Fixed-Wing UAVs in Unsteady Flows Using a GPU-Accelerated Vortex Particle Model
arXiv:2509.16079v2 Announce Type: replace Abstract: Unsteady aerodynamic effects can have a profound impact on aerial vehicle flight performance, especially dur…
End2Race: An End-to-End Learning Framework for Multi-Vehicle Autonomous Racing
arXiv:2509.16894v2 Announce Type: replace Abstract: Autonomous racing serves as a compelling testbed for advancing autonomous vehicle systems. The 1/10-scale F1…
PACT-WAM: Predicting Actions and Visual Foresight with Compact Temporal Encoding for Robot Manipulation
arXiv:2602.15882v2 Announce Type: replace Abstract: Robot manipulation uses temporal context to select actions and visual foresight to assess their consequences…
Multimodal Behavior Tree Generation: A Small Vision-Language Model for Robot Task Planning
arXiv:2603.06084v2 Announce Type: replace Abstract: Large language models have been widely used for robotic task planning, often taking advantage of representat…
Switchable-Polarity Electropermanent Magnet: Reconfigurable Magnetic Fields for Scalable Fluidic Control
arXiv:2603.24811v2 Announce Type: replace Abstract: Scalable control of pneumatic and fluidic networks remains fundamentally constrained by architectures that r…