Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6171 storiesA Hierarchical Spatiotemporal Action Tokenizer for In-Context Imitation Learning in Robotics
arXiv:2604.15215v4 Announce Type: replace Abstract: We present a novel hierarchical spatiotemporal action tokenizer for in-context imitation learning. We first …
Enhancing Human-Likeness in Reinforcement Learning Agents via Hierarchical Macro Action Quantization
arXiv:2605.30928v2 Announce Type: replace Abstract: Human-like agents are a long-standing goal of artificial intelligence. Despite strong performance, most rein…
RoverDevKit: An open, physics-grounded tradespace toolkit for conceptual design of lunar micro-rovers
arXiv:2606.21755v2 Announce Type: replace Abstract: Lunar micro-rovers under 50 kg are a rapidly growing vehicle class, yet open, benchmarked tools for running …
SurgRAW: Multi-Agent Workflow with Chain of Thought Reasoning for Robotic Surgical Video Analysis
arXiv:2503.10265v3 Announce Type: replace-cross Abstract: Robotic-assisted surgery (RAS) is central to modern surgery, driving the need for intelligent systems …
Curvature-aware Expected Free Energy as an Acquisition Function for Bayesian Optimization
arXiv:2603.26339v2 Announce Type: replace-cross Abstract: We propose an Expected Free Energy-based acquisition function for Bayesian optimization to solve the j…
Veo-Act: Enhancing VLA Policies with Frontier Video Models
arXiv:2604.04502v2 Announce Type: replace Abstract: Video generation models can produce coherent vi- sual sequences depicting object motion and interactions. We…
NanoBench: A Multi-Task Benchmark Dataset for Nano-Quadrotor System Identification, Control, and State Estimation
arXiv:2603.09908v2 Announce Type: replace Abstract: Existing aerial-robotics benchmarks target vehicles from hundreds of grams to several kilograms and typicall…
EmboAlign: Aligning Video Generation with Compositional Constraints for Zero-Shot Manipulation
arXiv:2603.05757v2 Announce Type: replace Abstract: Video generative models (VGMs) pretrained on large-scale internet data can produce temporally coherent rollo…
OmniPlanner: Universal Exploration and Inspection Path Planning Across Robot Morphologies
arXiv:2603.04284v2 Announce Type: replace Abstract: Autonomous robotic systems are increasingly deployed for mapping, monitoring, and inspection in complex and …
Learning Contact Dynamics through Touching: Action-conditional Graph Neural Networks for Robotic Peg Insertion
arXiv:2509.12151v3 Announce Type: replace Abstract: We present a learnable physics-based model that predicts motion of the robot end effector and reaction force…
Visual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks
arXiv:2508.11584v3 Announce Type: replace Abstract: Deploying multiple machine learning models on resource-constrained robotic platforms for different perceptio…
Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation
arXiv:2502.14254v3 Announce Type: replace Abstract: Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made them powerfu…
2Fast-2Lamaa: Large-Scale Lidar-Inertial Localization and Mapping with Continuous Distance Fields
arXiv:2410.05433v3 Announce Type: replace Abstract: This paper introduces 2Fast-2Lamaa, a lidar-inertial state estimation framework for odometry, mapping, and l…
In-Context Robot Learning with VLM Agents
arXiv:2609.19138v1 Announce Type: cross Abstract: Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied A…
A Convergence Framework for Deep $V$-Learning: Error Propagation and Sharp Action-Gap Bounds
arXiv:2609.18782v1 Announce Type: cross Abstract: We establish convergence bounds for deep $V$-learning with horizon $H$. The algorithm fits a scalar value func…
FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory
arXiv:2609.18623v1 Announce Type: cross Abstract: State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessi…
WetRobo: A Reproducible Robot Kit for Coding Agents in Biological Laboratories
arXiv:2609.18435v1 Announce Type: cross Abstract: Automating biological research requires general-purpose, reproducible robot systems that allow individual wet-…
Rollback the World, Keep the Reflection: Rollback-Induced Reflection for Long-Horizon LLM Agents
arXiv:2609.18304v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly tackle long-horizon tasks through multi-step environment intera…
Indicators of resilience for autonomous control systems
arXiv:2609.18264v1 Announce Type: cross Abstract: As modern societies rely more on autonomous systems to facilitate daily life, assuring their safe operation is…
rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference
arXiv:2609.19104v1 Announce Type: new Abstract: Factory work is a promising early scenario for embodied AI: assigning repetitive manual jobs to robots has clear…
Body-Motion Control of a Simulated Aerial Swarm from a First-Person View
arXiv:2609.18881v1 Announce Type: new Abstract: First-person-view (FPV) teleoperation of aerial swarms requires an operator to coordinate collective translation…
KINO: A Keyframe Interface for VLM Planning and Whole-Body Control in Humanoid Loco-Manipulation
arXiv:2609.18869v1 Announce Type: new Abstract: Humanoid loco-manipulation requires robots to interpret task instructions and scene semantics while executing co…
DynoFluxBench: Benchmarking Kinodynamic Space-Time Planners in Dynamic Environments
arXiv:2609.18549v1 Announce Type: new Abstract: Robots that leave structured, static environments must plan motions that are kinodynamically feasible and safe a…
ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
arXiv:2609.18487v1 Announce Type: new Abstract: Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both th…
VLM-MPPI: Grounding Natural Language in Behaviorally Diverse Trajectories for Aerial Navigation
arXiv:2609.18451v1 Announce Type: new Abstract: We present a hierarchical UAV navigation framework that aligns natural-language intent with dynamically feasible…
Hardware-Free Robotics Laboratories in Mixed Reality
arXiv:2609.18434v1 Announce Type: new Abstract: Teaching robotics relies on screen-based simulation, showing robot motion in an abstract coordinate frame rather…
DistAL: Distance-based Advantage Learning for VLA Fine-Tuning
arXiv:2609.18392v1 Announce Type: new Abstract: Vision-language-action models (VLAs) have trans- formed the field of robotic manipulation in recent years by com…
RecMorph: Topology-Guided Spatial Recurrence for Generalized Morphology Control
arXiv:2609.18359v1 Announce Type: new Abstract: Generalized morphology control requires a single policy to transform information across limbs with different phy…
Function-Preserving Data Generation for Zero-Shot Real-to-Sim-to-Real Manipulation
arXiv:2609.18293v1 Announce Type: new Abstract: Robotic data generation is a promising paradigm for scaling robot learning without collecting large-scale real-w…
Acting in Meters: Learning Metric Interactions for Precise Robotic Manipulation
arXiv:2609.18243v1 Announce Type: new Abstract: Vision-Language-Action models and World-Action Models have advanced language-conditioned robotic manipulation, y…