Robotics Briefing /
Better robot training, and a harder look at the tests.
Generated video can broaden what a humanoid learns. Human corrections can improve a trained robot’s weakest skills. A single physical probe can help predict how a material will move. Today’s papers show progress on all three, alongside an audit explaining why impressive benchmark scores need closer inspection.
Your quick takeaways
- PRISM turned four real videos into varied simulated training. A physical humanoid completed 55 of 60 trials with familiar object categories and 32 of 40 with new categories, under human joystick commands.
- Microsoft’s Rho improved two physical manipulation tasks after fifteen additional supervised episodes. Those gains came on known difficult configurations, following 150 demonstrations per task.
- A benchmark audit found that changing unrealistic object masses could reverse robot-controller rankings. What a test rewards can materially change which method appears best.
Papers
PRISM: expanding humanoid practice with generated video
Zihan Wang, Zhen Wu and colleagues • September 29 • Physical tests; paper lists CoRL 2026
Teaching a humanoid to lift and carry many kinds of objects normally requires varied demonstrations. PRISM starts with four real box-carrying videos and generates 256 variations involving boxes, bins, barrels and balls. It reconstructs the movements, adjusts them to the robot’s body and preserves the contacts needed to hold each object.
Filtering leaves 137 feasible trajectories; simulation training produces 129 successful simulated demonstrations used to train a controller that sees depth images. The physical test uses a Unitree G1 without further training on the test objects or their scans.
Adding the paper’s table entries gives 55 successes in 60 trials across familiar categories and 32 in 40 across new categories, including a helmet, chair and kettle. Success means reaching, grasping and carrying an object at least three meters without dropping it or falling. It does not require placing the object down successfully.
The useful result is that generated visual variation can become physical training experience when reconstruction and contact handling make it usable. A plausible-looking video alone is insufficient.
These were small, controlled tests with human joystick commands and depth processing on an external computer. The approach assumes rigid objects and struggles with articulated or deformable ones. Camera-angle changes also caused problems. Demonstrations are available, but the project’s linked code repository was unavailable when checked.
Rho: adapting a robot model to its weak spots
Microsoft Research Rho Team • September 29 • Technical report with physical tests
Rho is a family of roughly five-billion-parameter vision-language-action models: systems that turn visual observations and instructions into robot actions. Its training separates broad preparation, adaptation to a particular robot and learning individual tasks. The report evaluates three dual-arm robot setups.
One revealing experiment starts with FR3 Duo controllers trained on 150 demonstrations per task. Fifteen additional human-supervised episodes then adjust a small adaptation component. Test-tube assembly improves from nine successes in thirty trials to twenty-one; plug insertion improves from twenty to twenty-eight.
The qualification matters: corrections were collected on the same ten difficult starting configurations used for evaluation, with three trials per configuration. This measures improvement in known weak areas, not performance on unseen situations. It also does not mean the tasks were learned from only fifteen examples.
The broader promise is a practical way to refine an already trained system without retraining its entire action model. More varied robots, environments and instructions are still needed to establish how widely the results transfer.
Microsoft’s public repository contains implementation code, and the model collection includes base, robot-specific and benchmark checkpoints. The base model’s downloadable weight files are present. Those releases predate this report; the new paper supplies the detailed methods and evaluation.
FORM: learning material behavior from a physical probe
Stepan Tretiakov, Ruihan Zhao and colleagues • September 29 • Preprint with physical tests
Soft solids and liquids respond differently to the same robot movement. FORM estimates material properties from observed motion and contact forces, then uses the resulting physical model to simulate and plan another action. Its mathematical formulation avoids repeatedly running a simulator during parameter identification.
In a physical pouring test, one exploratory pour estimates glycerol’s viscosity and calibrates a contact parameter. Simulation then selects pouring angles for six requested volumes between 60 and 160 milliliters. Each target is tested five times.
The largest difference between a target and its five-trial average was 3.8 milliliters. That is not a bound on every individual pour. Cup markings and camera parallax were not calibrated, which limits confidence in the absolute measurement.
The team also demonstrates shaping Play-Doh, butter slime and plasticine. Together, the tests suggest that identifying how a material behaves can support new actions, rather than requiring a separate demonstration for every desired outcome.
The method still assumes a suitable material model and usable observations. These controlled experiments do not establish reliable handling of arbitrary mixtures or containers. Implementation code is available; the repository says its datasets are not yet publicly hosted.
Highlights
When fixing the test changes the winner
Qiwei Chen, Kaijun Zhou and colleagues • September 29 • Benchmark-audit preprint
A new audit examines seven robot-learning benchmarks and reports twenty-two bugs and four design limitations. Problems include instructions that disagree with success checks, permissive scoring and unrealistic physical parameters. These can distort comparisons of methods designed to make robot models run faster.
In RoboTwin’s Place Fan task, changing object masses while keeping model checkpoints and scoring rules fixed moved the baseline from 40% to 38% success. The accelerated DP-Cache-Fast method fell from 61% to 33%, reversing their ranking.
This does not show that acceleration methods are generally worse. It shows that a claimed improvement can depend on the particular test implementation. The findings apply to the audited versions and settings, and some proposed scoring changes measure movement quality rather than the same binary success rate.
For anyone following robotics progress, this is a useful reminder to look for physical trials, realistic conditions and a precise definition of success alongside the headline score.
Read this first
Read the benchmark audit, especially Tables 2 and 3. Its before-and-after comparisons show how apparently clear rankings can depend on what the simulator and success checker allow.
Read the paper