Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6171 storiesFeeling the Unexpected: ResTacVLA for Contact-Rich Manipulation via Residual Tactile Representation
arXiv:2607.03387v1 Announce Type: new Abstract: Tactile perception is indispensable for contact-rich manipulation, yet integrating it into Vision-Language-Actio…
StageCraft: Execution Aware Mitigation of Distractor and Obstruction Failures in VLA Models
arXiv:2603.20659v2 Announce Type: replace Abstract: Large scale pre-training on text and image data along with diverse robot demonstrations has helped Vision La…
Closing the Reality Gap: Zero-Shot Sim-to-Real Deployment for Dexterous Force-Based Grasping and Manipulation
arXiv:2607.04940v1 Announce Type: new Abstract: Human-like dexterous hands with multiple fingers offer human-level manipulation capabilities but remain difficul…
U-Joint CAAMS: Experimental Evaluation of a Universal-Joint Continuum Manipulator for Aerial Manipulation
arXiv:2607.03321v1 Announce Type: new Abstract: Continuum manipulators mounted on multi-rotor UAVs enable compliant aerial manipulation, but payloads and propel…
DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation
arXiv:2607.02604v1 Announce Type: cross Abstract: Although vision-language-action (VLA) models have received widespread attention, many challenges remain in man…
!Imperio, smolVLA: The Implications of Data Poisoning on Open Source Robotics
arXiv:2607.04146v1 Announce Type: new Abstract: This work establishes that trigger-word data poisoning of vision language action models is practical, while at t…
MOSAIC: Modular Scalable Autonomy for Intelligent Coordination of Heterogeneous Robotic Teams
arXiv:2601.23038v3 Announce Type: replace Abstract: Mobile robots have become indispensable for exploring hostile environments, such as in space or disaster rel…
MAD-PINN: A Decentralized Physics-Informed Machine Learning Framework for Safe and Optimal Multi-Agent Control
arXiv:2509.23960v2 Announce Type: replace Abstract: Co-optimizing safety and performance in large-scale multi-agent systems remains a fundamental challenge. Exi…
GelNeuro: A Sensing-Computing Integrated Neuromorphic Tactile System for Texture Recognition
arXiv:2607.05241v1 Announce Type: new Abstract: Neuromorphic visuo-tactile sensing offers a promising paradigm for low-latency and low-power robotic perception.…
Training Verifiably Robust Agents Using Set-Based Reinforcement Learning
arXiv:2408.09112v2 Announce Type: replace-cross Abstract: Reinforcement learning policies parametrized by deep neural networks have achieved strong performance …
Neural LiDAR Bundle Adjustment
arXiv:2607.04169v1 Announce Type: new Abstract: Recent research has achieved remarkable novel view rendering and scene reconstruction results with Neural Radian…
HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models
arXiv:2607.04265v1 Announce Type: new Abstract: World-action (WA) models can generate long-horizon action chunks for general-purpose robotic manipulation, but t…
Compressing the Validation Bottleneck: An Agentic Self-Driving Lab for Scientific Discovery
arXiv:2607.04508v1 Announce Type: cross Abstract: Agentic AI-for-Science can automate ideation, planning, and analysis, but final validation still depends on re…
$M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills
arXiv:2604.24182v2 Announce Type: replace Abstract: Current Vision-Language-Action (VLA) models predominantly rely on end-to-end fine-tuning. While effective, t…
Sample-Efficient Pareto Front Modeling for Energy-Aware Reinforcement Learning Using Bayesian Optimization
arXiv:2607.03140v1 Announce Type: cross Abstract: Industrial automation increasingly demands control strategies that balance operational performance with strict…
CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors
arXiv:2604.21241v2 Announce Type: replace Abstract: Vision--Language--Action (VLA) models often use intermediate representations to connect multimodal inputs wi…
Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies
arXiv:2607.05122v1 Announce Type: cross Abstract: Vision-language-action (VLA) models enable robot navigation from natural language and visual goals, but remain…
Toward Personalized Social Robots for Child Well-being: Data Requirement Principles from a Recommender-System Perspective
arXiv:2607.05110v1 Announce Type: cross Abstract: Social robots are increasingly deployed in clinical settings to support the well-being of children, where effe…
MOSAIC: Skill-Centric Manipulation Planning with Physics Simulation
arXiv:2504.16738v3 Announce Type: replace Abstract: Planning long-horizon manipulation motions using a set of predefined skills is a central challenge in roboti…
Finite Reliability Representations: Noise-Calibrated Belief-Space Covers for Reliable Decision-Making
arXiv:2607.04019v1 Announce Type: cross Abstract: Physical sensing and actuation noise floors should inform how much belief resolution a decision-making system …
Agent-driven Long-tail Simulation for Autonomous Driving
arXiv:2607.04331v1 Announce Type: new Abstract: Evaluating autonomous driving systems in closed-loop settings requires realistic and interactive simulation, yet…
Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation
arXiv:2607.05377v1 Announce Type: new Abstract: While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they stru…
Athena-WBC: Capability-Aligned Policy Experts for Long-Tail Humanoid Whole-Body Control
arXiv:2607.04837v1 Announce Type: new Abstract: Large-scale humanoid motion-tracking controllers are commonly improved by reallocating training effort: difficul…
HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control
arXiv:2607.03449v1 Announce Type: new Abstract: Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian …
AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning
arXiv:2607.03182v1 Announce Type: new Abstract: Autonomous driving planning requires translating navigation intent, traffic rules, dynamic interactions, and lan…
TreeLoc++: Robust 6-DoF LiDAR Localization in Forests with a Compact Digital Forest Inventory
arXiv:2603.03695v2 Announce Type: replace Abstract: Reliable localization is essential for sustainable forest management, as it allows robots to revisit and mon…
SurgAM: Surgical Affordance Map Prediction with Multimodal Feature Fusion for Robot Autonomy
arXiv:2607.04378v1 Announce Type: new Abstract: Surgical automation is being increasingly studied, yet bridging visual scene understanding with autonomous actio…
PreSIST: Vision-Language-Informed Object Persistence Prediction in Open-World Scenes
arXiv:2607.04057v1 Announce Type: cross Abstract: Robots deployed over long periods must reason about environments that change over time. Existing long-term per…
A User-driven Design Framework for Robotaxi
arXiv:2602.19107v3 Announce Type: replace Abstract: Robotaxis are emerging as a promising form of urban mobility, but removing human drivers fundamentally resha…
$\mathcal{P}^3$: Toward Versatile Embodied Agents
arXiv:2508.07033v2 Announce Type: replace Abstract: Embodied agents have shown promising generalization capabilities across diverse physical environments, makin…