Robotics Briefing /
Learning from human motion, physical contact and failure.
Human demonstrations can teach coordinated movement. Touch can reveal contact that cameras miss. Failed attempts can guide more useful practice in simulation. Today’s research puts numbers behind all three approaches, while a new industry collaboration asks how to keep robot actions within human-set permissions.
Your quick takeaways
- DexRoam transfers whole-body human demonstrations into robot training. Across five physical tasks, the strongest training recipe raised average success from 29% to 56% with one model and 32% to 57% with another.
- F4R reconstructs failed attempts in simulation, then retrains the robot. It reached 90% success on tested failure conditions, versus 71.25% for a comparison using newly collected corrective demonstrations.
- Uni-VLaT adds distributed touch sensing to a humanoid. It achieved 75% average success across five physical tasks, compared with 68% using touch without its predictive training objective.
Papers
DexRoam: preserving the coordination in human demonstrations
Rui Zhou and colleagues, HKUST, BAAI and collaborators • September 28 • Physical tests; project reports CoRL 2026 acceptance
Carrying a basket or moving an object to a bin requires the body, arms and fingers to work together. DexRoam records human movement with a consumer VR headset and a head-mounted stereo camera, then translates it into coordinated robot actions. It adjusts for differences in body structure, action representation and timing.
The team tested five tasks with twenty physical trials per task and training setting. Adding aligned human demonstrations improved the strongest training recipe from 29% to 56% average success on GR00T N1.7 and from 32% to 57% on π0.5. A separate data-budget experiment showed that human examples could reduce the number of robot demonstrations needed.
This offers a route to collecting useful examples without operating a robot for every recording. It still needs careful conversion and robot training data; ordinary unprocessed video is not equivalent to the captured demonstrations.
The robot hands lacked force or touch feedback. The authors observed grasps that contacted an object without securing it, showing why accurate motion alone is insufficient. Capture and alignment code is publicly available.
F4R: turning failed attempts into targeted simulated practice
Zhuoyuan Yu and colleagues, NTU, Dexmal and Xi’an Jiaotong University • September 28 • Preprint with physical tests
When a robot fails in an unfamiliar arrangement, collecting another human demonstration is one possible fix. F4R instead diagnoses the failure, reconstructs the relevant scene in simulation and refines the policy there before returning it to the physical robot.
Across four tabletop tasks, it averaged 90% success on the tested failure conditions, compared with 71.25% for targeted behavior cloning using corrective demonstrations. Each task and condition used twenty physical trials. The comparison matched one hour of data-preparation time, not identical numbers of trajectories or total computation.
Concentrating simulated practice around the diagnosed problem also beat a broader simulation-randomization comparison, which reached 77.5%. The implication is that the choice of practice conditions can matter as much as simply generating more practice.
The system still depends on a faithful reconstruction and correct diagnosis. It misidentified six of eighty test cases after two attempts, mainly involving lighting or viewpoint changes. Evaluation covered mostly rigid tabletop objects, and an operator remained available for emergency stops. Deformable objects, moving scenes and mobile manipulation remain unvalidated.
Uni-VLaT: teaching a humanoid to use contact across its body
Zihao Wang and colleagues, Tsinghua and collaborators • September 28 • Preprint with physical tests
A camera may miss pressure against a robot’s back or a changing load in its arms. Uni-VLaT adds textile pressure sensors across a Unitree G1’s torso, shoulders, back and arms. Its controller learns actions alongside predictions of future touch, body state and visual information.
Across five tasks with twenty trials per main-table configuration, average success was 75%. The same comparison without touch reached 32%; adding touch without predictive supervision reached 68%. The smaller seven-point comparison better isolates the added training method from the benefit of installing sensors.
Task results matter here. A back tap explicitly instructs the robot to walk, so that task strongly favors access to touch. Excluding it, the full method averaged 72.5%, versus 40% without touch and 63.8% with touch alone. The complete cleanup task remained at 55%.
These tests suggest that distributed sensing can help robots respond to loads and sustained contact. Sensor noise and mounting differences remain limitations, and a small set of demonstrations involving people does not establish general contact safety. Project code is marked coming soon.
News
Gecko and NVIDIA test enforceable boundaries for robot actions
Gecko Robotics • September 28 • Company announcement
Gecko Robotics says it is working with NVIDIA’s Open Agent Safety Platform and OpenShell software to enforce human-defined permissions around AI-controlled systems. Its announcement describes an independent enforcement layer between the AI agent and hardware on its Komodo robot.
The practical question is whether a robot can be prevented from taking an unauthorized action even when its AI proposes one. Separating action permissions from the model’s own decisions could provide an additional control for industrial autonomy.
The announcement describes the collaboration and intended architecture, but provides no independent robot-security benchmark, failure rate or certification evidence. It should be read as an engineering direction being tested, rather than proof that all unsafe physical behavior is prevented.
Highlights
ForVis: a harder look at drone localization in forests
Arman Kiani and colleagues, University of Maine • September 28 • Dataset and benchmark preprint
ForVis records twelve physical drone flights through meadow, above-canopy and under-canopy conditions, totaling roughly nine minutes and 1.1 kilometers. It compares seven systems that estimate motion from cameras and inertial sensors across 504 benchmark runs.
The study found substantial differences between two camera configurations. Reported tracking failures were about 10% with the RealSense D435i and 3% with the OAK-D Pro Wide. Different resolutions, fields of view and sampling rates prevent attributing the gap to one feature.
The harder limitation is reference accuracy: reliable satellite positioning was unavailable under the trees. Those sequences use return-to-start consistency as a proxy, which cannot reveal every error along the path. The paper is accessible, but its linked dataset folder was unavailable when checked, so a usable public download is not yet confirmed.
Read this first
Read Uni-VLaT, especially its results table. The comparison with touch alone, and the results excluding the back-tap task, make it easier to judge how much comes from new sensors versus the learning method.
Read the paper