Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
3542 storiesDeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
arXiv:2605.31286v2 Announce Type: replace Abstract: Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable…
SCC-Loc: A Unified Semantic Cascade Consensus Framework for UAV Thermal Geo-Localization
arXiv:2604.03120v2 Announce Type: replace-cross Abstract: Cross-modal Thermal Geo-localization (TG) provides a robust, all-weather solution for Unmanned Aerial …
ConTrack: Constrained Hand Motion Tracking with Adaptive Trade-off Control
arXiv:2606.03177v2 Announce Type: replace Abstract: Human demonstrations provide strong priors for robot manipulation, yet it is non-trivial to transfer them to…
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
arXiv:2603.22281v2 Announce Type: replace-cross Abstract: Recent progress in latent world models (e.g., V-JEPA2) has shown promising capability in forecasting f…
Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
arXiv:2603.03485v3 Announce Type: replace-cross Abstract: Recent video diffusion models have achieved impressive capabilities as large-scale generative world mo…
Critique of World Model: A Generative Latent Prediction Architecture for World Modeling
arXiv:2507.05169v4 Announce Type: replace-cross Abstract: World Model, the algorithmic simulator of the real-world environment which biological agents experienc…
Any2Any: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking
arXiv:2605.23733v2 Announce Type: replace Abstract: Whole-body tracking (WBT) models have become a key foundation for humanoid robots, enabling them to imitate …
When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning
arXiv:2605.05172v2 Announce Type: replace Abstract: Behavior Cloning (BC) has emerged as a highly effective paradigm for robot learning. However, BC lacks a sel…
LAGO Policy: Latency-Aware Asynchronous Diffusion Policies with Goal-Directed Collision-Free Planning for Smooth Manipulation
arXiv:2606.17982v1 Announce Type: new Abstract: Diffusion-based visuomotor policies deployed with asynchronous inference often exhibit inter-chunk discontinuiti…
Physical Imitation Learning: Distilling Control Policies into Passive Elasticity
arXiv:2604.00611v2 Announce Type: replace Abstract: Due to brain-body co-evolution, animals' intrinsic body dynamics play a crucial role in their energy-efficie…
RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models
arXiv:2506.17639v2 Announce Type: replace Abstract: Vision-Language-Action models (VLA) have demonstrated remarkable capabilities and strong potential in comple…
K-VARK: Kernelized Variance-Aware Residual Kalman Filter for Sensorless Force Estimation in Collaborative Robots
arXiv:2512.13009v2 Announce Type: replace Abstract: Reliable estimation of contact forces is crucial for ensuring safe and precise interaction of robots with un…
ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation
arXiv:2606.17937v1 Announce Type: new Abstract: Most Vision-Language-Action (VLA) models map observations directly to actions without explicit reasoning, limiti…
WAM-RL: World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFT
arXiv:2606.17906v1 Announce Type: new Abstract: Recent World-Action (WA) models demonstrate strong generalization ability and data efficiency, but they typicall…
ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
arXiv:2606.17200v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajec…
Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
arXiv:2606.17846v1 Announce Type: new Abstract: Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data und…
FLAP: FOV-Constrained Active Perception Planning for Prior-Map-Free 3D Navigation
arXiv:2606.17630v1 Announce Type: new Abstract: Safe and efficient trajectory planning in unknown, cluttered 3D environments constitutes a critical bottleneck f…
GASE: Gaussian Splatting-Based Automated System for Reconstructing Embodied-Simulation Environments
arXiv:2606.17520v1 Announce Type: new Abstract: Training embodied agents in the real world requires skilled operators and expensive hardware. Simulation environ…
ParkingTransformer: LLM-Enhanced End-to-End Trajectory Planning for Autonomous Parking
arXiv:2606.17082v1 Announce Type: new Abstract: End-to-end autonomous parking has emerged as a critical task within the realm of autonomous driving. However, ex…
MagicSim: A Unified Infrastructure for Executable Embodied Interaction
arXiv:2606.17511v1 Announce Type: new Abstract: Robot learning and embodied agents now require simulation to serve as a shared execution substrate linking contr…
MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation
arXiv:2606.17598v1 Announce Type: new Abstract: Humans naturally leverage diverse sensing modalities to interact with the physical world, while most Vision-Lang…
Transformer-Based Warm-Starting for Feasible and Optimal Terminal Approach to Tumbling Objects with Space Manipulators
arXiv:2606.17317v1 Announce Type: new Abstract: Real-time trajectory generation for on-orbit robotic servicing is challenging due to the nonlinear coupling betw…
Abstention-Aware Personalized Object Rearrangement via Uncertainty-Guided LLM Assistance
arXiv:2606.17309v1 Announce Type: new Abstract: Robotic assistance in household environments requires not only predicting where objects should be placed, but al…
VISTA: Scale-Aware Visual Navigation via Action History Conditioning
arXiv:2606.17294v1 Announce Type: new Abstract: Vision Navigation Foundation Models (VNMs) promise end-to-end learned navigation policies capable of zero-shot d…
VL-MemKnG: Hybrid Memory with a Spatio-Temporal Knowledge Graph for Question Answering over Long Egocentric Navigation Trajectories
arXiv:2606.17183v1 Announce Type: new Abstract: Answering navigation-relevant questions over long egocentric videos requires retrieving and organizing evidence …
Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System
arXiv:2606.18112v1 Announce Type: new Abstract: Agentic navigation systems require a base navigation model whose observation strategy can be externally reconfig…
DC-Motion: Decoupling Semantics and Details via Discrete-Continuous Tokens for Human Motion Generation
arXiv:2606.14721v1 Announce Type: cross Abstract: Text-to-motion generation requires synthesizing physically realistic dynamics that strictly follow complex and…
QPILOTS: Efficient Test-Time Q-Steering for Flow Policies
arXiv:2606.14801v1 Announce Type: cross Abstract: Flow-matching and diffusion policies are expressive action generators, but optimizing them with temporal-diffe…
Human Universal Grasping
arXiv:2606.17054v1 Announce Type: new Abstract: Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality. We a…
Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes
arXiv:2606.17043v1 Announce Type: new Abstract: When pretrained VLA policies are fine-tuned through online RL, each rollout episode produces only a single binar…