Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
3542 storiesTerraTransfer: Learning End-to-End Driving Policies Without Expert Demonstrations
arXiv:2606.17386v1 Announce Type: cross Abstract: End-to-end autonomous driving has achieved state-of-the-art performance on benchmarks and real-world deploymen…
WeaveLA: Event Driven Cross-Subtask Latent Memory Weaving for Repetitive Robot Manipulation
arXiv:2606.17463v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies have achieved remarkable single-step manipulation, yet they remain britt…
GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning
arXiv:2606.17480v1 Announce Type: cross Abstract: Generalist vision-language-action systems need object-centric 3D evidence and reusable manipulation experience…
Learn to Quantify Social Interaction with Constraints for Pedestrian Walking
arXiv:2606.17897v1 Announce Type: cross Abstract: Long-term human path forecasting in crowds is critical for autonomous moving platforms (like autonomous drivin…
Memory as a Wasting Asset: Pricing Flash Endurance for Embodied Agents, and the Limits of Doing So
arXiv:2606.18144v1 Announce Type: cross Abstract: A robot's flash endurance is a non-renewable stock: every persisted write spends one of a few thousand program…
SSIL: Self-Supervised Imitation Learning for End-to-End Driving
arXiv:2308.14329v4 Announce Type: replace Abstract: In autonomous driving, the end-to-end (E2E) driving approach that predicts vehicle control signals directly …
MOCHI: Motion Enhancement of Collaborative Human-object Interactions
arXiv:2606.18243v1 Announce Type: cross Abstract: Collaborative human-object interaction shows dynamic and complex movements that require mutual anticipation an…
OpenTie: Open-vocabulary Sequential Rebar Tying System
arXiv:2509.00064v2 Announce Type: replace Abstract: Robotic practices on the construction site emerge as an attention-attracting manner owing to their capabilit…
OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
arXiv:2509.26633v3 Announce Type: replace Abstract: A dominant paradigm for teaching humanoid robots complex skills is to retarget human motions as kinematic re…
Can Vision Foundation Models Navigate? Zero-Shot Real-World Evaluation and Lessons Learned
arXiv:2603.25937v2 Announce Type: replace Abstract: Visual Navigation Models (VNMs) promise generalizable, robot navigation by learning from large-scale visual …
A 3D Isovist World Model -- Revealing a City's Unseen Geometry and Its Emergent Cross-City Signature
arXiv:2606.03609v3 Announce Type: replace Abstract: Embodied agents that navigate cities rely on world models that predict how their surroundings will change as…
Simulating Infant First-Person Sensorimotor Experience via Motion Retargeting from Babies to Humanoids
arXiv:2604.27583v2 Announce Type: replace-cross Abstract: Motion retargeting from humans to human-like artificial agents is becoming increasingly important as h…
SimTO: A two-stage, simulation-driven topology optimization framework for bespoke soft robotic grippers
arXiv:2601.19098v2 Announce Type: replace Abstract: Soft robotic grippers are essential for grasping delicate, geometrically complex objects in manufacturing, h…
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
arXiv:2601.01762v3 Announce Type: replace Abstract: Practical autonomous driving requires models that generalize by reasoning through spatial-temporal possibili…
Contactless Respiratory Monitoring on Heterogeneous Mobile Robots: A Multimodal Edge-Computing Framework
arXiv:2606.17376v1 Announce Type: new Abstract: Respiratory-rate (RR) monitoring is a critical component of remote triage and victim assessment in emergency res…
Agent Utilities over Generalized Voronoi Regions and their Gradients
arXiv:2606.17388v1 Announce Type: new Abstract: In this paper, we generalize the concept of Voronoi regions, define agent utility as the integral of a utility d…
DexLink Hand: A Compact, Affordable, 16-DOF Linkage-Driven Hand with Human-Like Dexterity
arXiv:2606.17418v1 Announce Type: new Abstract: Dexterous robotic hands face a longstanding trade-off among dexterity, compactness, and affordability. Particula…
Embodiment Shapes Rolling Behavior in a Multimodal Infant Model
arXiv:2606.17456v1 Announce Type: new Abstract: Rolling over is one of the earliest milestones in infant motor development, reflecting the emergence of coordina…
When Robots Sleep: Offline Skill Consolidation for Shared-Policy Robot Learning
arXiv:2606.17493v1 Announce Type: new Abstract: Robots that learn over long deployments must add new skills without losing the shared policy structure that make…
ERQA-Plus: A Diagnostic Benchmark for Reasoning in Embodied AI
arXiv:2606.17639v1 Announce Type: new Abstract: Generalist embodied agents require more than object recognition: they must reason about spatial relations, actio…
Accountability in Autonomous Drone-Based Firefighting: Insights From a Field Trial
arXiv:2606.17831v1 Announce Type: new Abstract: There is a growing research field exploring how autonomous drones can enhance emergency response effectiveness. …
From Ad Hoc Pilots to Repeatable Patterns: Structuring Drone Collaboration in Emergency Services with DroneLets
arXiv:2606.17839v1 Announce Type: new Abstract: Drones hold promise for supporting emergency services, but their integration into workflows remains ad hoc and c…
PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space
arXiv:2606.17924v1 Announce Type: new Abstract: Current Vision-Language-Action (VLA) models face a trade-off between efficient action generation and explicit de…
Uncertainty Quantification for Flow-Based Vision-Language-Action Models
arXiv:2606.18043v1 Announce Type: new Abstract: Vision-language-action models (VLAs) combine vision-language backbones with expressive generative action heads t…
A Hybrid Optimization Framework for Grasp Synthesis under Partial Observations
arXiv:2606.18053v1 Announce Type: new Abstract: We propose a hybrid grasp synthesis framework that combines a learning-based Energy-Based Model (EBM) with an an…
Beyond Failure Recovery: An Engagement-Aware Human-in-the-loop Framework for Robotic Systems
arXiv:2606.18189v1 Announce Type: new Abstract: Conventional human-in-the-loop approaches typically involve users only when a robot encounters failure or uncert…
EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies
arXiv:2606.18239v1 Announce Type: new Abstract: We present EBench, a simulation benchmark that diagnoses generalist mobile manipulation policies beyond a single…
Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement
arXiv:2606.18247v1 Announce Type: new Abstract: Robots deployed in the real world should learn from their experience and improve over time. This requires a mech…
Credibility-Weighted Pricing of Autonomous Vehicle Liability Under Operational Design Domain Shift
arXiv:2606.17451v1 Announce Type: cross Abstract: Automated Driving System deployments create a foundational ratemaking challenge: sparse experience, shifting o…
Adaptive Volumetric Mechanical Property Fields Invariant to Resolution
arXiv:2606.18231v1 Announce Type: cross Abstract: Accurate mechanical properties (or materials) Young's modulus ($E$), Poisson's ratio ($\nu$) and density ($\rh…