Robotics Briefing /
Taking robot skills beyond the training setup.
A robot can succeed in training and struggle when the parts, surroundings, or task history change. Today’s research tests ways to close that gap, from simulated construction work to navigation lessons collected with a walker. Figure also reports household trials that put humanoid generalization to a concrete test.
Your quick takeaways
- RebarSim trained entirely in simulation and completed 137 of 150 physical insertions. The tests cover two familiar designs with real manufacturing variation.
- Workspace Models helped a robot remember hidden cube counts, succeeding in 17 of 20 physical trials. Only one task tested the approach on hardware.
- Figure reports 56% full-task success across three behaviors in 30 unfamiliar homes. The result shows progress, alongside substantial remaining failures.
Papers
RebarSim: simulated practice transfers to real steel
Tao Sun and colleagues, McGill, Princeton and University of Washington • Submitted September 17 • Preprint; physical robot tests
Rebar insertion requires aligning imperfectly manufactured steel with narrow slots. RebarSim trains a controller in simulation across varied bar shapes, then teaches a camera-based model to reproduce its actions. Randomized appearances and viewpoints help bridge the visual differences between simulation and hardware.
On a Franka arm, it completed 137 of 150 insertions using factory-produced bars, with no real-world training examples. The two tested designs both appeared in simulation training; the physical bars introduced manufacturing variation. A separate background test completed 25 of 30 attempts across five altered scenes.
The result suggests simulation can reduce repeated data collection for demanding assembly work. It also shows why training across part variation matters, beyond making simulated images look realistic.
The robot starts with the bar already grasped, and its camera crops target a fixed slot group. Some failures arise when contact friction makes the bar shift in its grasp. This validates an insertion stage, with full construction workflows still ahead.
Workspace Models: remembering the details a task needs
Nitish Dashora and colleagues, MIT and Carnegie Mellon • Submitted September 17 • Authors list CoRL 2026; simulation and hardware
Once a cube disappears into an opaque box, a robot needs memory to know how many it has placed. Workspace Models train a compact memory representation using a larger vision-language model’s judgments about which past events and image regions matter.
The expensive analysis happens during training. During operation, the robot uses the learned memory to choose actions without repeatedly asking the larger model to summarize its history.
In a physical task dividing cubes evenly between two boxes, the method succeeded in 17 of 20 trials. Three additional tasks tested counting, drawer recall, and adapting a grasp in simulation. With batching and caching, reported policy latency was 49.2 milliseconds, versus 199.5 milliseconds for the batched keyframe-selection baseline.
This could make useful memory easier to include in responsive robot control. Evidence remains narrow: one physical task, task-specific supervision, and a memory architecture whose computation grows with history length. The larger model can also misjudge which visual details deserve attention.
UNI: a walker collects navigation lessons for robots
Sarvesh Prajapati and colleagues, Northeastern University • Submitted September 17 • Preprint; data collection and powered-wheelchair tests
A person pushing a four-wheeled walker naturally chooses routes that accommodate wheels. The Universal Navigation Interface, or UNI, adds a smartphone to record those journeys. The team collected 37.2 kilometers of demonstrations, including pauses and waiting as well as forward motion.
Training existing navigation models on this data reduced trajectory prediction error by 17.4–24.8% on held-out UNI recordings. Results on other datasets were mixed, so the improvement does not establish general navigation gains everywhere.
In physical tests, a UNI-adapted controller stopped before three of four uncut curbs and all four staircases. The original controller stopped at none. Both crossed all four accessible curb cuts.
These are small tests at 12 locations, with a 0.30-meter-per-second speed limit and an operator holding an enable button. One curb failure remains. The study supports inexpensive data collection while leaving dependable public-space navigation unproven.
The project says its dataset, app source, processing code, and trained weights remain unreleased.
News
Figure tests Helix 2.5 in 30 unfamiliar homes
Figure • Announced September 17 • Company-reported physical evaluation
Figure reports that Helix 2.5 performed room tidying, towel folding, and bed making in 30 Bay Area homes excluded from training. The company adapted one shared foundation model into three task-specific behaviors.
In its comparison, pretraining on the Index human-behavior dataset raised full-task success from 9% to 56%, with downstream training and evaluation otherwise held fixed. Each task used a fixed model across all homes. Safety interventions counted as failures.
Here, “zero-shot” refers to unfamiliar homes and objects. The robots still received task training elsewhere. That distinction matters when judging how much preparation a new deployment could require.
The findings suggest broad human experience can help skills transfer between environments. They remain company-reported results across three behaviors, with failure in 44% of trials under the stated aggregate metric. Independent testing and a wider range of household work will determine how far that transfer extends.
Highlights
HIL-UMI: collecting demonstrations around a model’s weak spots
Zimu Han and colleagues, Peking University, Xi’an Jiaotong, PrimeBot and JD Technology • Submitted September 17 • Preprint; four physical tasks
HIL-UMI lets a person demonstrate tasks with a handheld gripper while a robot model predicts actions from the same observations. When the predictions differ from the person’s movements, the system collects those examples for further training. A second model helps identify demonstrations that advance the task.
On table cleanup, collection took 73.40 milliseconds per recorded frame, compared with 412.99 milliseconds for HG-DAgger, a method involving physical robot execution and human intervention. HIL-UMI also finished about five points higher on the paper’s task-progress score.
That score awards partial credit; it is not a percentage of fully completed tasks. Each checkpoint received ten trials per task on one arm. Collection was also slower per frame than ordinary demonstrations, although it produced better progress scores.
The opportunity is more targeted training data without tying every collection session to a robot. The project provides comparison demonstrations, while code is marked coming soon.
Read this first
Read RebarSim alongside its demonstrations. The recovery attempts and failures make the transfer problem tangible: simulated practice can teach a useful motion, while friction and grasp stability still decide whether it works.
Read the paper