Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
3542 storiesTDMA Based Communications Control Co-Design for Cooperative Carrying: Delay Calibration and Sampling-Rate Optimization
arXiv:2608.09556v1 Announce Type: new Abstract: Multi robot teams performing cooperative transportation face a fundamental challenge: maintaining stable control…
FactorDrive: Adaptive Multi-Step Reasoning Driven by Planning-Critical Factors for End-to-End Autonomous Driving
arXiv:2608.09591v1 Announce Type: new Abstract: Vision-language models (VLMs) have advanced scene understanding and enabled explicit reasoning in end-to-end aut…
Nonlinear Model Predictive Control of a Robotic Soft Esophagus
arXiv:2608.09602v1 Announce Type: new Abstract: Strictures caused by esophageal cancer can narrow down the esophageal lumen, leading to dysphagia. Palliation of…
Predictive safety filter enhanced curriculum learning control for efficient vehicle dynamics controller
arXiv:2608.09653v1 Announce Type: new Abstract: Recent advances in learning-based control have enabled impressive achievements in solving complex control proble…
TAMS: Task-Aware Multi-View Adaptive Streaming for Wireless Telerobotic Manipulation
arXiv:2608.09731v1 Announce Type: new Abstract: Wireless telerobotic manipulation relies on timely multi-view video feedback, but the available uplink bandwidth…
Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition
arXiv:2608.09762v1 Announce Type: new Abstract: Real-world online reinforcement learning (RL) provides a promising approach for training robotic manipulation po…
SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation
arXiv:2608.09771v1 Announce Type: new Abstract: Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language condi…
RoboSeg: Online Part-Level Semantic Reconstruction for Robotic Manipulation via a Single Eye-in-Hand Camera
arXiv:2608.09778v1 Announce Type: new Abstract: Robotic manipulation requires perception systemsthat identify actionable parts such as handles, rims, triggers,a…
WRAP: Wasserstein-Robust Adaptive Plug-in for Robot Localization
arXiv:2608.09807v1 Announce Type: new Abstract: Robotic localization under changing sensing conditions can suffer from biased errors and miscalibrated covarianc…
Hierarchical Fast--Slow ReAct Agent for Zero-Shot Object-Goal Navigation
arXiv:2608.09816v1 Announce Type: new Abstract: Zero-shot object-goal navigation (ZSON) requires a robot to find a named object category in a building it has ne…
Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy
arXiv:2608.09857v1 Announce Type: new Abstract: Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development …
Energy-Structured Latent World Models with Neural Time Fields for Physically Constistent Open-World Motion Planning
arXiv:2608.09876v1 Announce Type: new Abstract: Physically consistent motion planning remains a fundamental challenge in embodied AI, as generated trajectories …
RoSE: A Robotic Soft Esophagus for Endoprosthetic Stent Testing
arXiv:2608.09891v1 Announce Type: new Abstract: Soft robotic systems are well suited for developing devices for biomedical applications. A bio-mimicking robotic…
XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment
arXiv:2608.09892v1 Announce Type: new Abstract: Robot policy evaluation and deployment remain fragmented by model-specific software dependencies, data represent…
Emotion in an active inference model of human driving
arXiv:2608.07480v1 Announce Type: cross Abstract: Active inference has emerged as a principled framework for modeling adaptive behavior by balancing goal-direct…
An AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadruped Navigation Research Loop
arXiv:2608.07542v1 Announce Type: cross Abstract: Autonomous research loops driven by large language models can run machine-learning experiments at scale but te…
Contraction Analysis of Holomorphic Dynamical Systems via the Intrinsic Kobayashi Metric
arXiv:2608.07551v1 Announce Type: cross Abstract: This paper studies incremental stability of holomorphic dynamical systems through the infinitesimal Kobayashi …
Exact Contraction Rates via the Berkson--Porta Representation: A Sharp Threshold and Its Herglotz-Kernel Obstruction
arXiv:2608.07552v1 Announce Type: cross Abstract: Semigroups of holomorphic self-maps of the unit disc with an interior fixed point are, by the classical Berkso…
Impact of Dataset Composition on Embedded Real-Time UAV Wildfire Detection Using Compact YOLO Models
arXiv:2608.07554v1 Announce Type: cross Abstract: The development of vision-based wildfire detection systems for unmanned aerial vehicles is constrained by the …
The Field Knows: Cross-Dimensional Geometry from Navigation to Black Holes
arXiv:2608.07566v1 Announce Type: cross Abstract: We introduce a continuous metric field framework trained by a single causal contrastive loss. The framework en…
Multimodal Skin Lesion Classification with Swin Transformer and Clinical Metadata Fusion
arXiv:2608.07574v1 Announce Type: cross Abstract: Skin lesion classification plays an important role in supporting the early diagnosis of skin cancer. However, …
Open-World Hierarchical Perception: Taxonomic Abstraction over Class-Agnostic Proposals for the Safe Handling of Out-of-Vocabulary Road Objects
arXiv:2608.07577v1 Announce Type: cross Abstract: A closed-set detector for autonomous driving must assign every object one of a fixed set of labels. On an obje…
CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models
arXiv:2608.07621v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous dr…
LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation
arXiv:2608.07746v1 Announce Type: cross Abstract: Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level…
V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control
arXiv:2608.07870v1 Announce Type: cross Abstract: Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world …
GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning
arXiv:2608.07905v1 Announce Type: cross Abstract: Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to l…
Parameter-Dependent LMI Synthesis for Semi-Global Differential ISS Trajectory Tracking of Nonholonomic Mobile Robots Under Multiplicative Wheel Slip
arXiv:2608.08049v1 Announce Type: cross Abstract: This paper presents a parameter-dependent linear matrix inequality (LMI) framework for trajectory tracking of …
Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?
arXiv:2608.08077v1 Announce Type: cross Abstract: Theory of Space framework (ToS) assesses the spatial understanding of curiosity-driven Vision-Language Models …
Exploring LLM Capabilities for Situational Understanding and COLREG compliance on real-world maritime navigation scenarios
arXiv:2608.08281v1 Announce Type: cross Abstract: Recently, Large Language Models (LLMs) have shown considerable capability for situational understanding, reaso…
Ego-OSCAR: Egocentric Open source Stereo CAptuRe System
arXiv:2608.08285v1 Announce Type: cross Abstract: We present Ego-OSCAR, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric d…