Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
3542 storiesKitchen Robotic Manipulation utilizing Foundation Models
arXiv:2608.04042v1 Announce Type: new Abstract: Deploying robots in everyday human environments requires perception systems that are both robust and adaptable t…
Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models
arXiv:2608.04765v1 Announce Type: new Abstract: Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language unders…
ResPlan: A Large-Scale Vector-Graph Dataset of 17,000 Residential Floor Plans
arXiv:2508.14006v2 Announce Type: replace-cross Abstract: We introduce ResPlan, a dataset of 17,000 residential floor plans with vector geometry, room-connectiv…
A GitOps-Driven Annotation Catalog for Fully Automatic Railway Operations
arXiv:2608.04724v1 Announce Type: new Abstract: Automatic train operation (ATO) at grade of automation 3 and above (GoA3-GoA4) requires robust AI-based percepti…
RORA: Realistic Object Reconstruction with Articulation
arXiv:2608.04842v1 Announce Type: new Abstract: Replicating real-world environments into simulation by realistic visual representation like NeRF and 3D Gaussian…
SSC: A Verifiable Structured Representation for Bimanual Manipulation Labelling
arXiv:2608.04425v1 Announce Type: new Abstract: Subtask labels decompose a long-horizon manipulation demonstration into shorter semantic segments for policy tra…
Differential 6-DOF Pose Estimation with Provable First-Order Immunity to Camera Calibration Errors
arXiv:2608.04673v1 Announce Type: cross Abstract: Accurate six-degree-of-freedom (6-DOF) motion estimation is essential for robotic manipulation, autonomous sys…
DM$^3$-Nav: Decentralized Multi-Agent Multimodal Multi-Object Semantic Navigation
arXiv:2604.22014v2 Announce Type: replace-cross Abstract: We present DM$^3$-Nav, a fully decentralized multi-agent semantic navigation system supporting multimo…
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
arXiv:2511.15669v3 Announce Type: replace-cross Abstract: Does Chain-of-Thought (CoT) reasoning genuinely improve Vision Language Action (VLA) models, or does i…
Geometry-Aware Sampling-Based Motion Planning on Riemannian Manifolds
arXiv:2602.00992v3 Announce Type: replace Abstract: In many robot motion planning problems, task objectives and physical constraints induce non-Euclidean geomet…
A Multi-Sensor Dataset for Monitoring the Operational Environment of Rail Vehicles
arXiv:2608.04704v1 Announce Type: cross Abstract: Reliable environment monitoring is essential for the safe and efficient operation of automated railway systems…
Toward Integrating Adaptive Experience Replay and Online Uncertainty Estimation in Safe Actor-Critic Optimal Control
arXiv:2608.04732v1 Announce Type: cross Abstract: Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as sep…
Interpretable Fuzzy Inference for UAV Target Tracking Using Bounding-Box Geometry
arXiv:2608.04121v1 Announce Type: new Abstract: Vision-based guidance of unmanned aerial vehicles (UAVs) toward unmanned ground vehicles (UGVs) supports coopera…
Retrieve in Time, Correct in Frequency
arXiv:2608.04527v1 Announce Type: new Abstract: Frozen vision-language-action (VLA) policies generate temporally extended action chunks, but long-horizon manipu…
SpikingNav: Robust Embodied Navigation with Spiking Neural Policies
arXiv:2608.05078v1 Announce Type: new Abstract: Embodied navigation requires an agent to make sequential decisions from egocentric observations in a physical en…
BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
arXiv:2608.05042v1 Announce Type: new Abstract: Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerge…
SACK : Safe Active Continual Koopman Learning for Uncertain Systems with Contractive Guarantees
arXiv:2605.09659v2 Announce Type: replace Abstract: Koopman operator theory provides a powerful framework for representing nonlinear dynamics through a linear o…
SCOPE: Field-of-View-Aware Path Planning in Unknown 3D Environments via Safety-Volume Certification
arXiv:2608.04420v1 Announce Type: new Abstract: Safe navigation with a body-mounted limited-field-of-view sensor requires the complete robot-inflated volume of …
Enabling Urgency-aware Robot Swarm Intralogistics using Smart IoT Tags
arXiv:2608.04721v1 Announce Type: new Abstract: Warehouse items differ in how urgently they must be moved: perishable goods, pharmaceutical shipments, and just-…
RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation
arXiv:2510.17640v4 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have shown strong manipulation capability when trained with large-scale …
Overcoming Statistical Bias in Action-Controllable World Models
arXiv:2608.04653v1 Announce Type: cross Abstract: Action-conditioned world models aim to predict how visual environments evolve under an agent's actions. Yet fu…
Exact Model-Free Policy Iteration for Co-safe LTL Planning
arXiv:2608.05047v1 Announce Type: cross Abstract: This work studies model-free reinforcement learning for co-safe linear temporal logic (sc-LTL) objectives in f…
From Transparent Labware Segmentation to Collision Avoidance: A Real-Time Edge-Aware Perception Pipeline
arXiv:2608.04769v1 Announce Type: new Abstract: This paper presents an edge-aware instance segmentation framework that enables real-time robotic collision avoid…
Static Timing Orchestration for Tree-Structured Robot Control Firmware
arXiv:2608.04600v1 Announce Type: new Abstract: As robotic systems become increasingly complex, generating control firmware from structural description files ha…
Feasibility of Embedded Photoplethysmography Sensing in Short-Duration Tactile Interactions With Pocket-Sized Robots Using IMU- and Confidence-Based Filtering
arXiv:2608.04242v1 Announce Type: new Abstract: Ubiquitous companion robots offer a promising avenue for immediate anxiety relief in children, yet their effecti…
GASP: GPU-Accelerated Safe Planner for Real-Time Collision-Aware Motion Generation with Latent Trajectory Sampling
arXiv:2608.04612v1 Announce Type: new Abstract: We present GASP, a GPU-Accelerated Safe Planner for real-time, collision-aware joint-space motion generation in …
Approximate Multi-Objective Search Under Rulebooks
arXiv:2608.04398v1 Announce Type: new Abstract: Robotic planning often involves multiple objectives with complex priority relationships, such as safety, efficie…
Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models
arXiv:2608.04633v1 Announce Type: new Abstract: Recent Vision-Language-Action (VLA) methods improve generalization by aligning their representations with 3D sce…
A Systematic Review and Taxonomy of Reinforcement Learning-Model Predictive Control Integration for Linear Systems
arXiv:2604.21030v2 Announce Type: replace-cross Abstract: The integration of Model Predictive Control (MPC) and Reinforcement Learning (RL) has emerged as a pro…
Tactus: Open-Vocabulary Object Recognition from Low-Cost Pressure Arrays
arXiv:2608.04043v1 Announce Type: cross Abstract: Resistive pressure arrays are the cheapest and most widely shipped tactile sensors, yet tactile representation…