Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6171 storiesCompVLA: A Variable Compliance Vision-Language-Action Model for Contact-rich Manipulation
arXiv:2609.23614v1 Announce Type: new Abstract: Contact-rich manipulation, requiring robots to regulate not only motion but also how they yield to external forc…
AR-WAM: A Visual-Conditioned Agent-Ready World Action Model for Robotic Manipulation
arXiv:2609.23578v1 Announce Type: new Abstract: As AI agents become increasingly capable, agent-driven robotic control is emerging as a compelling paradigm. How…
PRIMO: Prior-Informed Odometry from Human-Motion Tracking for Humanoid Robots
arXiv:2609.23610v1 Announce Type: new Abstract: Simulation-trained humanoid proprioceptive odometry faces two transfer challenges: training trajectories generat…
Grounded Action Model: 3D Grounding as a Foundation for Robotics
arXiv:2609.23863v1 Announce Type: new Abstract: Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that curre…
HapticWAM: Distilling Imagined Touch into a World-Action Model without Inference-Time Tactile Sensing
arXiv:2609.23888v1 Announce Type: new Abstract: Contact-rich manipulation requires estimating forces, slip and contact geometry that can remain ambiguous in sce…
Dexterous Robot Manipulation from Human Demonstrations via Contact-Anchored Retargeting and Residual Policy Learning
arXiv:2609.24093v1 Announce Type: new Abstract: Learning dexterous manipulation from demonstrations is bottlenecked by data: the contact forces that determine w…
KerColle: Unlocking Fine-Grained GPU Concurrency in Vision-Language-Action Models
arXiv:2609.22335v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as foundational models for next-generation robotics. High VLA…
Riemannian Density-Driven Optimal Control: Tangent-Space LQR for Second-Order Multi-Agent Systems on Curved Manifolds
arXiv:2609.22678v1 Announce Type: cross Abstract: Density-Driven Optimal Control (D2OC) provides an effective framework for steering multi-agent systems toward …
G6D: Geometric Learning-Free RGB-D 6D Pose Solver for Robotic Manipulation
arXiv:2609.23566v1 Announce Type: cross Abstract: 6D object pose estimation is fundamental to robotic manipulation and automation. Recent zero-shot methods have…
Hybrid Imitation Learning: Teleoperation Augmentation Primitives that Policies Learn to Trigger
arXiv:2512.04960v2 Announce Type: replace Abstract: What an operator can demonstrate bounds what imitation learning can learn. Teleoperation interfaces map the …
Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models
arXiv:2411.05005v2 Announce Type: replace-cross Abstract: Beyond high-fidelity image synthesis, diffusion models have recently exhibited promising results in de…
Energy-based Regularization for Learning Residual Dynamics in Neural MPC for Omnidirectional Aerial Robots
arXiv:2604.14678v2 Announce Type: replace-cross Abstract: Data-driven Model Predictive Control (MPC) has lately been a core research subject in the field of con…
3D-MoE: Towards Spatial Intelligence with Mixture-of-Experts for 3D Reasoning and Action Generation
arXiv:2501.16698v2 Announce Type: replace-cross Abstract: Spatial intelligence, encompassing 3D perception and reasoning, is the essential next frontier of AI. …
A Multimodal Label Forecasting Method for Aperiodic Visuo-Motor Time Series
arXiv:2609.07930v2 Announce Type: replace Abstract: Deep learning models have been increasingly applied to Time Series Forecasting (TSF) in recent years. Transf…
AdaReP:Adaptive Re-Planning under Model Mismatch for Neural World-Model Predictive Control
arXiv:2606.23079v2 Announce Type: replace Abstract: Neural world models coupled with model predictive control (MPC) replan at every environment step to bound ac…
Learning Energy-Efficient Air--Ground Actuation for Hybrid Robots on Stair-Like Terrain
arXiv:2603.26687v2 Announce Type: replace Abstract: Hybrid aerial--ground robots can use thrust to cross obstacles that impede wheel-driven motion, but deciding…
MoE-ACT: Scaling Multi-Task Bimanual Manipulation with Sparse Task-Conditioned Mixture-of-Experts Transformers
arXiv:2603.15265v2 Announce Type: replace Abstract: Developing a unified policy for multi-task robotic manipulation remains challenging due to policy degradatio…
Observing and Controlling Features in Vision-Language-Action Models
arXiv:2603.05487v2 Announce Type: replace Abstract: Vision-Language-Action models (VLAs) have shown remarkable progress towards embodied intelligence. While the…
Design and Biomechanical Evaluation of a Lightweight Low-Complexity Soft Bilateral Ankle Exoskeleton
arXiv:2602.18569v2 Announce Type: replace Abstract: Many people could benefit from exoskeleton assistance during gait, for either medical or nonmedical purposes…
Heterogeneous Robot Collaboration in Unstructured Environments with Grounded Generative Intelligence
arXiv:2510.26915v2 Announce Type: replace Abstract: While heterogeneous teams have typically been designed for well-specified missions with known semantics, gen…
Enhancing Autonomous Driving Safety through World Model-Based Predictive Navigation and Adaptive Learning Algorithms for 5G Wireless Applications
arXiv:2411.15042v3 Announce Type: replace Abstract: Addressing the challenge of ensuring safety in ever-changing and unpredictable environments, particularly in…
Do LiDAR Language Models Really Understand Spatio-temporal Relationships?
arXiv:2609.24452v1 Announce Type: cross Abstract: Recent 4D LiDAR language models aim to reason about objects and their evolving spatial relationships. Yet, in …
AnalogDepth: Multi-view Geometry from FPV drones under Analog Video Transmission
arXiv:2609.24312v1 Announce Type: cross Abstract: Analog video transmission (VTX) remains widespread in FPV drones due to low latency, weight and low cost. Howe…
MoSAT: Human Motion Generation from Spatial Audio and Textual Description
arXiv:2609.23797v1 Announce Type: cross Abstract: Human motion is shaped by both external acoustic events and behavioral intent: spatial audio conveys environme…
Which Terrain Is Better? Preference Learning with VLM Prototypes for Off-Road Traversability Ranking
arXiv:2609.23673v1 Announce Type: cross Abstract: In vision-based off-road navigation, a robot needs to know not only which obstacles to avoid but also which te…
Towards robust multimodal 3D object detection via visual foundation models
arXiv:2609.23541v1 Announce Type: cross Abstract: Multimodal 3D object detection is fundamental to robust perception in autonomous driving because it integrates…
Algebraic Consistency Alone Does Not Certify Temporal Structure in Latent Action Models
arXiv:2609.23478v1 Announce Type: cross Abstract: Latent action models infer a code for the transition between two frames of action-free video. Recent methods r…
Identity Continuity in Long-Term Embodied AI Relationships: From Agent-Specific Identity Representation to Identity-Continuity Appraisal
arXiv:2609.23356v1 Announce Type: cross Abstract: Long-term embodied AI will undergo learning, model updates, memory compression, hardware repair, and migration…
M3GA-Wild: A Large-Scale Dataset and Benchmark for Multi-Modal Multi-session Ground-to-Aerial Place Recognition in Forests
arXiv:2609.23003v1 Announce Type: cross Abstract: We present M3GA-Wild, the first benchmark for multi-modal, multi-session ground-to-aerial place recognition in…
Learning and Control Beyond Linearity: Towards a Non-asymptotic Theory for Bilinear Systems
arXiv:2609.22338v1 Announce Type: cross Abstract: This tutorial provides a unified view of the emerging area of bilinear learning and control. Using linear syst…