Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6171 stories3D Cal: An Open-Source Software Library for Depth Reconstruction on Vision-Based Tactile Sensors
arXiv:2511.03078v3 Announce Type: replace Abstract: Tactile sensing plays a key role in enabling dexterous and reliable robotic manipulation, but realizing this…
TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training
arXiv:2607.02840v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown promising generalization in robotic manipulation, but they still …
MorphQuad: Morphable Quadrotor for Superhuman Maneuverability, Manipulation, and Resiliency
arXiv:2607.02764v1 Announce Type: new Abstract: Infrastructure maintenance, contact-based inspection, and emergency response can benefit from aerial vehicles th…
Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis
arXiv:2607.05348v1 Announce Type: cross Abstract: Open-vocabulary 3D scene understanding aims to segment 3D scenes beyond predefined categories by transferring …
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
arXiv:2607.02642v1 Announce Type: new Abstract: Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficien…
SLAM: Structured and Localized Analytic Manifold Adaptation for Lifelong VPR
arXiv:2607.04764v1 Announce Type: new Abstract: Visual Place Recognition (VPR) in lifelong deployment requires continuous adaptation to new environments without…
Learning 3D Affordances for Blade Insertion in Cluttered Stowing
arXiv:2607.02549v1 Announce Type: cross Abstract: Many manipulation tasks require reasoning about free-space affordances: discovering volumes where an extended …
WSA$_1$: a 3D-Centric World-Spatial-Action Model for Generalizable Robot Control
arXiv:2607.03941v1 Announce Type: new Abstract: Recent advances in embodied AI have established robot foundation models (RFMs) as the dominant approach for gene…
FLOAT Drone for Physical Interaction: Lateral Airflow Reduction, Wrench Modeling, and Adaptive Control
arXiv:2607.04260v1 Announce Type: new Abstract: Aerial physical interaction represents a promising direction for next-generation unmanned aerial vehicles (UAVs)…
Real-World Perturbation Testing of Autonomous Driving Systems
arXiv:2607.04953v1 Announce Type: cross Abstract: Autonomous Driving Systems (ADS) must operate reliably under diverse conditions, yet representative data for r…
Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue
arXiv:2509.15061v3 Announce Type: replace Abstract: Embodied agents are intelligent systems designed to perceive, reason, and act within the physical world. Whi…
KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation
arXiv:2607.04652v1 Announce Type: new Abstract: Learning manipulation from few demonstrations requires visual priors that capture not only where to interact, bu…
Learning to Visually Connect Actions and their Effects
arXiv:2401.10805v4 Announce Type: replace-cross Abstract: We introduce the novel concept of visually Connecting Actions and Their Effects (CATE) in video unders…
ObjRetarget: An Object-Aware Motion Retargeting Framework with Anthropomorphic Arm Constraints and Polyhedral Hand Modeling
arXiv:2607.03828v1 Announce Type: new Abstract: Learning robot dexterous manipulation from human manipulation videos requires reliably retargeting human intent …
Verifier-free Test-Time Sampling for Vision-Language-Action Models
arXiv:2510.05681v2 Announce Type: replace Abstract: Vision-Language-Action models (VLAs) have demonstrated remarkable performance in robot control. However, the…
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
arXiv:2511.17384v2 Announce Type: replace Abstract: While Visual Large Language Models (VLLMs) show great promise as embodied agents, they continue to face subs…
Humanoid says KinetIQ Ascend reinforcement learning approaches human-level dexterity
Humanoid says its KinetIQ Ascend approach can reach 99.9% manipulation reliability at human speed and beyond for industrial tasks. The post Humanoid says KinetI…
Context is king: How Avride uses cloud VLMs as a safety net for delivery robots
Avride uses vision-language models, or VLMs, to improve the environmental awareness of its delivery robots. The post Context is king: How Avride uses cloud VLMs…
Choreographing the Way of Water: A Computational Framework for Aquatic Robotic Art
arXiv:2607.02174v1 Announce Type: new Abstract: Robotic choreography in open water is governed by nonlinear fluid dynamics, which impose significant challenges …
Learning to Localize Reference Trajectories in Image-Space for Visual Navigation
arXiv:2602.18803v2 Announce Type: replace Abstract: We present LoTIS, a model for visual navigation that provides robot-agnostic image-space guidance by localiz…
BIEVR-LIO: Robust LiDAR-Inertial Odometry through Bump-Image-Enhanced Voxel Maps
arXiv:2604.14421v2 Announce Type: replace Abstract: Reliable odometry is essential for mobile robots as they increasingly enter more challenging environments, w…
Simulation Based Reward Function Validation for Multi-Agent On Orbit Inspection
arXiv:2607.01367v1 Announce Type: cross Abstract: A proposed method for the control of groups of inspection spacecraft is Multi-Agent Reinforcement Learning (MA…
Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots
arXiv:2607.02501v1 Announce Type: new Abstract: Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical de…
Cross-Platform Control for Autonomous Surface Vehicles via Adaptive Reinforcement Learning
arXiv:2607.02037v1 Announce Type: new Abstract: Autonomous surface vehicles vary widely in hydrodynamic and actuation characteristics, yet most controllers are …
A Convex Obstacle Avoidance Formulation
arXiv:2512.13836v2 Announce Type: replace-cross Abstract: Autonomous driving requires reliable collision avoidance in dynamic environments. Nonlinear Model Pred…
Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models
arXiv:2512.01715v2 Announce Type: replace Abstract: Vision-language-action (VLA) models that generate continuous action chunks via flow matching lack an interna…
PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
arXiv:2607.01938v1 Announce Type: new Abstract: Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodie…
Learning 3D-Gaussian Simulators from RGB Videos
arXiv:2503.24009v3 Announce Type: replace-cross Abstract: Realistic simulation is critical for applications ranging from robotics to animation. Learned simulato…
Robust Image Processing Techniques for Construction Environment Monitoring Using Underwater Robots
arXiv:2607.01915v1 Announce Type: cross Abstract: This paper proposes a robust image processing framework for underwater robot-based construction environment mo…
Adaptive Companionship for Group-Following Robots: Handling Dynamically Changing Group Formations
arXiv:2607.01287v1 Announce Type: new Abstract: Accompanying a group of humans is an essential aspect of developing human-like social cognition in robots. Howev…