Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
3542 storiesWorld Action Models in Real Time: An Empirical Study of Smooth Execution via Asynchronous Deployment
arXiv:2608.01880v1 Announce Type: new Abstract: World Action Models generate fixed-horizon action chunks through iterative denoising, creating substantial infer…
GraRe: Grasp Candidate Re-Ranking for Frozen 6-DoF Grasp Detectors
arXiv:2608.00946v1 Announce Type: new Abstract: Existing 6-DoF grasp detectors typically rank grasp candidates by detector confidence. However, our analysis on …
Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models
arXiv:2608.02197v1 Announce Type: new Abstract: Visual representations of VLA models remain unreliable for spatially precise robotic manipulation. We uncover th…
Uncertainty Quantification for Visual Object Pose Estimation: S-Lemma Ellipsoidal Bounds
arXiv:2511.21666v2 Announce Type: replace Abstract: Quantifying the uncertainty of an object's pose estimate is essential for robust control and planning. Altho…
Local-Canonicalization Equivariant Graph Neural Networks for Sample-Efficient and Generalizable Swarm Robot Control
arXiv:2509.14431v2 Announce Type: replace Abstract: Multi-agent reinforcement learning (MARL) policies for swarm control often learn inefficiently and generaliz…
KING: Embodiment-Aware Kinematic Graph Neural Network for Unified Motion Representation of Legged and Wheeled Robots
arXiv:2608.01015v1 Announce Type: new Abstract: Kinematic models provide reliable motion constraints for odometry estimation in featureless environments, where …
Spline Policy: A Structured Representation for Robot Policies
arXiv:2606.07386v2 Announce Type: replace Abstract: Modern imitation-learning policies for robot manipulation often represent actions as fixed-resolution action…
Staged Multi-Agent Training (SMAT) for Hip Exoskeletons: Metabolic and Biomechanical Validation of a Simulation-Trained Co-Adaptive Controller
arXiv:2608.00715v1 Announce Type: new Abstract: Learning-based controllers can deliver exoskeleton assistance after training entirely in physics-based simulatio…
Latent-Centroid Steering: Single-Pass Classifier-Free Guidance for Command-Aligned Autonomous Driving
arXiv:2608.00237v1 Announce Type: cross Abstract: Vision-language models (VLMs) have recently emerged as a promising paradigm for end-to-end autonomous driving,…
Altitude-Adaptive Vision-Only Geo-Localization for UAVs in GPS-Denied Environments
arXiv:2602.23872v4 Announce Type: replace-cross Abstract: Matching downward-looking unmanned aerial vehicle (UAV) images to georeferenced satellite or aerial ma…
AffordTrajDP: Dynamic Affordance-Guided Visuomotor Policy Learning for Robotic Manipulation
arXiv:2608.01603v1 Announce Type: new Abstract: Affordance-guided imitation learning has shown impressive performance in robotic manipulation tasks by compressi…
PRISM: Privileged Probabilistic Latent Supervision for End-to-End Autonomous Driving Motion Planning
arXiv:2608.01201v1 Announce Type: new Abstract: End-to-end autonomous driving (E2E AD) systems integrate perception, prediction, and planning into a single diff…
From Failures to Supervision: DynamicEnvPlan for Robust Long-Horizon Embodied Planning
arXiv:2608.00613v1 Announce Type: new Abstract: Physical-world interaction is inherently dynamic, as environments can evolve during execution, requiring agents …
ORCESTRA: VLM-driven Visual Robot programming in Mixed Reality
arXiv:2608.00775v1 Announce Type: new Abstract: ORCESTRA is a mixed-reality system for programming robot digital twins through no-code waypoint teaching and lan…
HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing
arXiv:2603.15257v2 Announce Type: replace Abstract: Tactile sensing is a crucial capability for Vision-Language-Action (VLA) architectures, as it enables dexter…
Hybrid Attention Estimation Pipeline for Adaptive HRI Using an Expressive Robotic Head
arXiv:2608.00284v1 Announce Type: new Abstract: This paper presents an applied case study on hybrid visual attention estimation for human-robot interaction usin…
VLAGuard: A Framework for Evaluating and Mitigating Physical Attention Hijacking in Vision-Language-Action Robots within Wireless Sensor Networks
arXiv:2608.01028v1 Announce Type: new Abstract: Deploying Vision-Language-Action (VLA) robots as mobile edge nodes within wireless sensor networks (WSNs) requir…
DriveCode: Domain Specific Numerical Encoding for LLM-Based Autonomous Driving
arXiv:2603.00919v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown great promise for autonomous driving. However, discretizing nu…
DiffPhysCam: Differentiable Physics-Based Camera Simulation for Inverse Rendering and Embodied AI
arXiv:2508.08831v2 Announce Type: replace-cross Abstract: Generating synthetic images that closely mimic those from real cameras is instrumental in training vis…
Faster-WAM: Do World Action Models Need Deep Action Modules?
arXiv:2608.02365v1 Announce Type: cross Abstract: World Action Models (WAMs) couple robot action prediction with video world models. Existing WAMs with shared-b…
Rake-Compress Riccati Recursions for Parallel Scenario-Tree Model Predictive Control
arXiv:2608.01332v1 Announce Type: cross Abstract: Scenario-tree model predictive control (MPC) represents future information by a rooted tree and optimizes a no…
Open-DiffLoco: Open-Source Differentiable Learning for Deployable Blind Quadruped Locomotion
arXiv:2608.02069v1 Announce Type: new Abstract: Developing deployable locomotion policies through conventional reinforcement learning often requires complex rew…
CoWAM: Coordination Contracts for Selective Policy Intervention with WAMs
arXiv:2608.02578v1 Announce Type: new Abstract: World Action Models (WAMs) augment robot policies with action-conditioned predicted futures, but a plausible fut…
Teleopit: A Full-Embodiment Humanoid Teleoperation System
arXiv:2608.01834v1 Announce Type: new Abstract: Humanoid teleoperation for demonstration collection requires coordinated whole-body motion, continuous dexterous…
CAAT: Contact-Aware Attention Scaling and Tactile Masking for Data-Efficient Contact-Rich Manipulation
arXiv:2608.01102v1 Announce Type: new Abstract: In contact-rich manipulation, visual observations primarily guide motion in free space, whereas tactile observat…
SSTG-Nav: Metric-Grounded Spatial-Semantic Topological Graphs for Reusable Object Navigation
arXiv:2608.00527v1 Announce Type: new Abstract: Service robots operating for months in the same homes, offices, and facilities should become more reliable with …
Disentangled Control of Multi-Agent Systems
arXiv:2511.05900v4 Announce Type: replace-cross Abstract: This paper develops a general framework with convergence guarantees for multi-agent control synthesis,…
SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
arXiv:2608.01397v1 Announce Type: new Abstract: World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depend…
WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA
arXiv:2608.01035v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; howe…
Latency-Tolerant Cloud-Edge Collaborative Vision-Language-Action Models via Emergent Representational Specialization
arXiv:2608.00569v1 Announce Type: new Abstract: Deploying billion-parameter Vision-Language-Action (VLA) policies on mobile robots creates a systems conflict: s…