Robotics Briefing /
Learning from failures, making room for other robots, and testing physical judgment.
A robot can look capable until it meets the awkward object placement, crowded passage or instruction it should refuse. Today’s research examines those difficult cases: directing training toward failures, coordinating movement without giving up a safe stop, and measuring how language models respond to physical hazards. In industry, RobCo sets a commercial launch date while Rockwell focuses on connecting robots to the rest of the factory.
Your quick takeaways
- Mulligan raised physical marker-insertion success from 42% to 76% by combining targeted practice with learned action selection. Human supervision remains essential.
- GAMBIT coordinated 1,024 simulated robots and demonstrated eight physical drones. Its collision guarantees depend on the modeled motion and stopping assumptions.
- Inspect Robots found that the tested language models rarely refused several hazardous instructions. Better manipulation scores do not establish sound physical judgment.
Papers
Mulligan: spend more learning time on the failures
Lars Ankile and colleagues, Stanford • October 5 submission; newly listed October 6 • Preprint with physical tests
Once a robot handles most object placements, collecting more examples at random can mean repeating tasks it already understands. Mulligan starts later training rounds from observed failures and previously unexplored placements. A person resets the scene, supervises the attempt and intervenes when needed.
The full system also learns a value function—a predictor of how useful a candidate action will be—from successful and failed experience. This helps choose among possible actions, while the movement policy learns from successful human demonstrations and corrections.
Across three physical tasks, the study used 2,550 held-out, blinded evaluation episodes. In the final marker-insertion comparison, the full method completed thirty-eight of fifty trials, versus twenty-one of fifty for the uniformly sampled interactive-imitation baseline: 76% versus 42%. Nut placement reached 72% versus 62%; completing both cable clips reached 34% versus 18%.
Those gains compare combined systems, not targeted sampling alone. The final-round nut and cable differences also did not meet the conventional 5% significance threshold in the paper’s paired tests. The especially difficult cable task still failed more often than it succeeded.
The practical lesson is that selecting informative experience can matter alongside improving the learning algorithm. The method assumes people can recreate chosen starting conditions, and each physical collection strategy had one campaign; these results do not establish autonomous improvement across an uncontrolled workplace.
GAMBIT: coordination that keeps a way to stop
Rishabh Jain and colleagues, Cambridge, TU Berlin and AIST • October 5 • Preprint with simulation and physical demonstrations
In a crowded space, the fastest route for one robot can block everyone else. GAMBIT learns to select short movements from examples, then improves that selection through reinforcement learning. Robots can temporarily move away from their own destinations to let others pass.
A separate safeguard checks candidate movements together with the stopping paths that follow them. If no acceptable new movement is available, the robot continues its previously reserved backup trajectory. That separation lets the learned component focus on cooperation while an explicit mechanism handles collision avoidance.
In simulation, the system coordinated as many as 1,024 robots. For 512 robots, it computed coordinated actions in under 100 milliseconds per step. Separately, eight physical drones exchanged destinations around obstacles in a motion-capture arena, with planning at ten updates per second.
The distinction matters: the thousand-robot result is computational evidence, while the hardware demonstration is much smaller. The mathematical guarantee assumes the specified motion model and mutually collision-free backup paths; it is not a blanket guarantee against sensing errors or unexpected people.
Some robots also became stuck repeating movements without reaching their goals. Safe stopping and successful traffic flow remain separate requirements.
Inspect Robots: test judgment as well as task completion
Christopher Leet and colleagues, Robocurve and collaborators • October 5 • Preprint; submitted to the CoRL physical-AI safety workshop
Inspect Robots provides reusable components for running physical evaluations: task definitions, model interfaces, action checks, trial records and analysis. Its new paper uses a two-arm YAM robot to show why capability and safety need separate measurements.
The safety study tested six language-model policies on four hazardous instructions, with twenty trials per model and task. Several models usually refused an instruction involving a human-shaped doll. Yet no model refused any of the other three task types more than 5% of the time. Those tasks involved foreseeable electrical, battery or fire hazards.
The result exposes a gap between recognizing a visibly human target and recognizing danger through object interactions. It measures behavior in these particular scenes, rather than proving that every accepted instruction was successfully executed or that harm occurred.
The sample is small, task wording and robot interfaces can affect behavior, and command limits constrained the hardware. The framework makes such comparisons easier to reproduce; it does not itself certify a model as safe. Its public code predates this paper, which supplies new evaluation evidence.
News
RobCo sets a launch date for Alfie
October 5 • Company announcement
RobCo announced a transaction valuing it above $1 billion, combining new investment with an employee share sale. The company says its Alfie industrial robot will commercially launch in Munich on March 4, 2027, targeting variable, less structured factory work.
RobCo also reports customer operations across more than a dozen US states and manufacturing and assembly in Austin. Those statements describe its existing business, not a measured deployment record for the forthcoming Alfie.
The useful milestone to watch is whether the new system can handle changing production tasks with acceptable reliability and integration cost. This announcement provides no independent performance comparison, uptime figures or Alfie customer results.
Rockwell puts robot logistics inside the production workflow
October 5 announcement • Demonstrations planned for October 18–21
Rockwell’s PACK EXPO preview brings OTTO mobile robots together with FactoryTalk Orchestration, production controls and Emulate3D digital twins. The emphasis is on coordinating material movement with manufacturing equipment and workflows.
That addresses a practical constraint: a robot arriving safely is useful only if the material, machine and next process are ready. The announced demonstrations could make those connections easier to inspect. They remain upcoming trade-show demonstrations, with no new measured throughput or deployment outcomes in the release.
Highlights
A closer look at Mulligan’s actual experience
Companion resource to the October 5 paper
The project page shows recorded training examples, and its Hugging Face collection includes demonstration and evaluation datasets. These resources help connect the reported scores to what the robot actually did. The project’s code button is still a disabled placeholder, despite the paper’s availability statement.
Read this first
Read Mulligan’s final-round comparison table alongside its demonstrations. The marker-insertion improvement is substantial, while the lower cable success and uncertainty on smaller gains keep the result in perspective.
Read the paper