Robotics Briefing /
Helping robots act at the right moment.
A robot reaching for a moving bottle needs to know where it is going. Today’s research tests how recent motion and predictions of the future improve robot actions. There is also a new humanoid announcement, with a useful distinction between demonstrated performance and the capabilities still under development.
Your quick takeaways
- TEMPO uses recent observations and movement history to improve dynamic manipulation. In physical bottle handovers, it succeeded in 37 of 50 trials.
- ModAR predicts motion and scene structure before choosing actions. Adding human videos improved its results across three physical tasks.
- Agility unveiled Digit 5 with new safety features and early access planned for 2027. Neverwhere shows why realistic simulation still needs checks against hardware.
Papers
TEMPO: a moving object needs more than a snapshot
Zhenyang Feng, Jimin Heo and colleagues, UC Irvine • Submitted September 15 • Authors report CoRL 2026 acceptance; physical robot tests
A single image can show a bottle’s position while hiding its direction and speed. TEMPO gives a robot model two extra inputs: a summary of recent video and a short history of its own movement commands. Together, they help it track motion and distinguish task stages that look similar.
In physical bottle handovers, TEMPO succeeded in 37 of 50 trials, compared with 22 for RTC and 19 for VLASH, two controllers designed for responsive execution. The tests varied the person’s walking speed, bottle height, and holding pose. TEMPO also reduced complete misses, where the arm reached toward the wrong place.
The study covers four tasks, including catching balls and pouring beads into a moving container. It suggests that reducing computation delay alone leaves useful motion information missing.
The scope remains one laboratory setup with four designed tasks. For two tasks, the main comparison removes a confusing release-and-retract phase to give the baselines a workable test. Public code and a separate motion-perception benchmark make the approach available for further evaluation.
ModAR: predict motion and geometry before choosing an action
Adam Hung and colleagues, Carnegie Mellon University • Submitted September 15 • Preprint; simulation and physical tests
World-action models learn to predict what a robot will see and do next. ModAR builds those predictions in stages: point movement, visual features, and depth, followed by robot actions. Each stage can use the predictions that came before it.
Across three physical two-arm tasks, ModAR achieved 83.3% success, compared with 66.7% for a model that generated the predictions together. Both used the same robot and human demonstration data. Each method ran 30 trials per task: stacking cups, folding a towel, and placing an object in a drawer before closing it.
Human videos helped even without robot action labels. Adding task-specific human demonstrations raised ModAR’s average success from 70% to 81.1%; adding broader human footage brought it to 83.3%. The model uses those videos to learn how scenes change.
The findings support structured predictions as a useful route to robot learning. The physical tests cover only three tabletop tasks, with varied starting poses. Reliability across unfamiliar objects and workspaces still needs broader testing.
Weave: coordinating a humanoid’s body and fingers
Liu Cao and colleagues, Tsinghua, Dalian University of Technology and CUHK • Submitted September 15 • Preprint; simulation-only evaluation
Carrying an object requires a humanoid to coordinate its grip, body posture, and steps. Weave converts captured human interactions into robot movement references while preserving intended hand contact. One controller commands 29 body joints and 12 finger joints on a simulated Unitree G1 with dexterous hands.
Across nine objects, the controller completed 92.5% of trained interactions and 65% of unseen sequences involving those same objects. That gap shows the remaining difficulty of combining familiar movements in unfamiliar sequences.
The contribution is a unified way to learn grasping and whole-body motion together. This matters because a slipping grip changes the forces the rest of the body must handle. The team has published implementation code.
All evaluation takes place in simulation, despite the abstract’s description of physically executed rollouts. The controller receives exact robot and object state and follows supplied movement references. Physical sensing, transfer to hardware, and deciding which interaction to perform remain open steps.
News
Digit 5 targets closer work with people
Agility Robotics • Announced September 15 • Product in development; early access expected in the first half of 2027
Agility unveiled Digit 5 with human detection and an independent safety controller designed to trigger avoidance, stopping, or a seated position. The goal is to reduce reliance on physical barriers in factories and warehouses.
The company specifies a 50-pound payload, swappable grippers, and a battery designed for 90 minutes of operation followed by nine minutes of charging. These changes target practical constraints: how much a robot can move, which tasks it can handle, and how often it leaves work to recharge.
Agility expects general availability by the end of 2027. Its cited 65,000 operating hours come from Digit 4 deployments; they do not establish Digit 5’s field performance.
The product page says some safety features remain in development and specifications may change. Throughput, interventions, and safe operation in shared spaces will be key measures as deployments begin.
Highlights
Neverwhere: testing robot vision in reconstructed places
Ziyu Chen and collaborators • Submitted September 14; listed September 16 • Authors report IROS 2026 acceptance; public code and scene data
Neverwhere offers more than 60 reconstructed indoor and outdoor environments for testing visual locomotion. It combines realistic camera views with collision geometry, letting a robot controller repeatedly attempt routes through digitally recreated places. The public release includes code, scene assets, and policy checkpoints.
The authors compared a controller in simulation and on hardware across four matching scenes, with ten trials per scene. Hurdle completion matched at 75%. On stairs, simulation reported 85% completion while physical trials reached 60%.
That difference is the useful lesson: realistic images can improve testing while imperfect physics still overstate performance. Neverwhere makes repeated comparisons easier, but it cannot replace physical validation. Creating additional environments also requires some manual alignment, scaling, and task labeling.
Read this first
Read TEMPO. A person walking past with a bottle makes its central problem easy to see: position alone does not reveal motion. The physical trials and failure breakdown show how a short history changes the robot’s timing.
Read the paper