Robotics Briefing /
Better feedback. Better recovery.
Robots need better ways to notice mistakes and correct them. Today’s papers offer concrete tests and methods for doing that.
Your quick takeaways
- Tactile feedback helped TacPAC correct actions mid-execution. Its authors report 64% average success across five physical-robot tasks, versus 48% for the strongest tested baseline.
- Two new benchmarks expose hidden weaknesses: robots struggle to recover after failure, and reward models can change their judgment when a task description gets reworded.
- Industrial announcements focus on inspections, assembly, and easier programming. Their deployment claims still need evidence beyond demonstrations.
Papers
TacPAC: Tactile Prediction and Real-Time Action Correction in World-Action Models for Contact-Rich Manipulation
Zipei Ma and colleagues • Submitted September 4 • Preprint; physical-robot experiments
TacPAC compares live touch readings with the contact its plan expected. A small correction module then revises actions the robot has yet to execute. This lets a robot respond during a movement, without regenerating its entire action sequence.
The authors report 64% average success across five tasks, compared with 22% for their vision-only version and 48% for T-Rex, the strongest tested baseline. One correction took 30.4 milliseconds on their setup.
Why it matters: insertion and fragile-object handling depend on contact details that cameras miss. Limit: the evaluation used one Flexiv arm with specific tactile sensors and 20 trials per task per method. Broader hardware transfer remains unproven.
LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models
Lin Liu and colleagues • Submitted September 4 • Preprint; simulation benchmark
The team collects failure states from robot policies running in LIBERO, then asks whether those policies can recover. Its four levels range from retrying an action to restoring disrupted objects and surroundings.
The study reports more than 50% performance drops for all evaluated models when they face these failures. Rankings also change: doing well from a clean starting state does not reliably predict recovery.
Why it matters: ask for recovery results before treating a high task-success score as deployment readiness. Limit: these experiments use simulated robot states. They do not establish failure rates on physical production robots.
Same Trajectory, Contradictory Rewards (RoboRMBench)
Wonje Jeung and colleagues, including Yonsei University and Carnegie Mellon University • Submitted September 4 • Preprint; recorded physical-robot trajectories
RoboRMBench tests whether a vision-language reward model scores the same robot behavior consistently when the instruction changes wording but keeps its meaning. It contains 2,390 trajectories and 21,673 checked paraphrases.
The authors find widespread scoring instability, including reversals between success and failure. Larger models and explicit reasoning do not reliably fix it; dedicated robot reward models prove more stable.
Why it matters: a reward model can feed conflicting signals into robot learning even when the task stays the same. Limit: much of the evidence comes from offline scoring and trajectory selection, rather than fresh closed-loop robot training.
Physics filtering favors the generalization of robot learning
Jindou Jia and colleagues • Published September 4 • Peer-reviewed, npj Robotics; physical-robot experiments
PhyFilter corrects learned outputs using known physical structure and live state feedback. The researchers test it across quadruped locomotion, drone flight, aerial manipulation, and acceleration estimation.
The paper reports that a quadruped policy trained on flat simulated ground handles varied physical terrain with the filter. An aerial manipulation experiment reports a maximum error of 2.5 centimeters despite wind and mass uncertainty.
Why it matters: physics-based correction offers a route to better transfer without relying entirely on more training data. Limit: applying it requires useful physical equations and suitable state feedback; it is not a universal add-on for every robot policy.
News
Pudu brings D7 to Europe at IFA
Company announcement: September 7 • Event: September 4–8
Pudu’s semi-humanoid D7 makes its European debut alongside its commercial robot lineup. The company describes live object-pickup and visitor-interaction demos; its D5-W demonstrates mobility on stairs and loose surfaces.
Why it matters: a service-robot vendor is testing how broader manipulation fits its product range. Treat these as company-described trade-show demonstrations. The release does not provide independent task-success or uptime results.
FANUC previews AI-assisted manufacturing
Catch-up: announced September 3 • IMTS runs September 14–19
FANUC plans demonstrations of handwritten instructions directing parts kitting, natural-language robot programming, and dual-arm assembly using vision and force data. It also describes simulation work connecting NVIDIA Isaac Sim with ROBOGUIDE.
Why it matters: these are concrete attempts to reduce programming and commissioning work. The announcement previews upcoming exhibits; it supplies no measured factory-wide productivity gain.
Caterpillar and FieldAI target inspections and jobsite autonomy
Catch-up: announced September 2
The collaboration combines Caterpillar’s operational knowledge with FieldAI’s robot foundation models. Early applications include autonomous inspections, digital representations of jobsites, and better awareness of potential hazards.
Why it matters: inspection creates a practical path for robot learning in changing industrial spaces. Caterpillar gives intended applications, but no fleet size, intervention rate, or independently measured return.
Highlights
RoboSPA: a benchmark to watch
Submitted September 4; authors report EMNLP 2026 acceptance
Zhenxuan Fan and colleagues describe 527,000 trajectories across 280 task variants that test spatial reasoning and multi-step planning. The abstract reports difficulties with spatial relations and memory-heavy tasks. Release check: the linked repository still says code and data are in preparation, despite the abstract saying they are available. This summary uses the abstract and repository.
Watch PhyFilter’s physical-robot demonstrations
Companion material to the September 4 journal paper
The project page offers a visual check on the paper’s transfer claims. Look at how the quadruped handles terrain changes and how aerial manipulation responds to disturbances. These are author-provided demonstrations, not an independent replication.
Read this first
Read TacPAC first. It offers a concrete mechanism, physical-robot tests, a strong baseline comparison, and a clear hardware limitation. That makes its claimed gain easier to assess.
Read the paper