Robotics Briefing /
Robot learning meets the physical world.
An excavator shapes a full-size embankment, while a robot hand uses touch to handle small objects. Another controller keeps working when its camera view changes. Today’s research shows where learned skills survive physical tests, and where narrow trials still leave open questions.
Your quick takeaways
- A learned excavator controller built a 42-meter embankment in 45 minutes without human intervention. The result comes from one large field run.
- STAR improves dexterous manipulation with touch, while LIT helps robot models handle unfamiliar lighting and camera views.
- FoldNet++ reports real T-shirt folding from simulated demonstrations. Richtech plans to show manipulation and transport robots working as one fleet.
Papers
Material-state learning: an excavator reshapes soil autonomously
Lennart Werner and colleagues, ETH Zurich and Hexagon • Submitted September 11; listed September 14 • Preprint; physical field tests
Digging depends on how soil moves around the bucket. This work trains controllers in a particle-based soil simulation, where rewards reflect the shape and compactness of the material. The learned movements use multiple bucket surfaces to push, lift, and reshape soil, including material outside the bucket.
The team deployed the approach on an 11.5-metric-ton hydraulic excavator. It built a 42-meter embankment in about 45 minutes, completing 201 strokes across 67 locations without human intervention. The measured height from trench bottom to embankment top was 2.1 meters; conventional controllers handled driving, chassis balance, and path planning.
The same learned weights also ran on a 500-gram tabletop machine through a calibrated control interface. That suggests a useful separation between learning how to move soil and controlling a particular machine. Transfer still requires calibration and compatible bucket geometry.
This was one large field run, with no discarded earlier attempts reported. A separate single-site comparison matched an expert’s forward progress while producing a higher embankment. Repeated trials across different sites and soils would be needed to establish broader reliability.
STAR: making sparse touch signals useful to robot hands
Xiangcheng Liu and colleagues, Shanghai Innovation Institute and AgiBot • Submitted September 11; listed September 14 • Preprint; physical manipulation tests
A camera can show where an object sits while missing how it presses against a finger. STAR combines images, touch, and language, filtering inactive touch inputs and learning to predict future contact. The team collected 200 hours of demonstrations covering 65 tasks.
Across four physical tasks, STAR reported about 61% success, compared with 44% for the same underlying model trained on the dexterous dataset without STAR’s training recipe. Tasks included flipping an earbud, retrieving a tightly stacked book, picking up a postcard, and collecting multiple objects. Each method received 100 task-specific demonstrations and ran 20 trials per task.
The result supports better use of contact information when finger coordination matters. It also shows why adding sensors alone does not solve dexterity: the hand’s sensors miss its finger edges, and touch brought uneven benefits across tasks. These tests used task-specific training with the evaluation objects; they do not establish broad performance on unfamiliar tasks.
LIT: keeping robot actions focused when the scene changes
Jianman Lin and colleagues, SCUT, NUS and NTU • Submitted September 11; listed September 14 • Preprint; simulation and physical tests
Robot models can learn accidental links between a scene’s appearance and the right movement. Latent Interface Training first teaches movements using spatial goals without images. It then introduces vision through a compact representation trained to preserve the goal’s position and orientation.
In physical tests using MolmoAct2, success under changed lighting rose from 53.3% to 70%. With a changed camera setup, it rose from 30% to 46.7%. Each comparison covered three tasks with ten trials per task under each changed condition; both models trained on the same 300 demonstrations.
The approach could reduce sensitivity to ordinary changes in a workspace. Across four model architectures, simulated robustness also improved. Physical evidence remains limited to one platform and three tasks, and the camera-change result still falls below half of attempts.
News
Richtech brings manipulation and transport into one planned workflow
Catch-up • Announced September 10 • IMTS runs September 14–19 • Company demonstration preview
Richtech plans to demonstrate its DEX humanoid alongside the Titan 440 autonomous mobile robot at IMTS in Chicago. DEX handles object interaction using vision and depth sensing, while Titan moves parts and materials. A shared platform lets employees assign work and monitor both robots.
The practical question is how reliably one robot hands work to another. Coordinating manipulation with transport could reduce manual transfers between stations and give operators one view of job progress.
The announcement previews a demonstration and supplies no measured throughput, intervention rate, or sustained customer-operation results. Those measures will determine whether the combined workflow improves an operation beyond what the individual robots can already do.
Highlights
FoldNet++: simulated laundry reaches physical robots
Yuxing Chen and colleagues, Peking University and Galbot • Submitted September 11; listed September 14 • Authors report CoRL 2026 acceptance; dataset release pending
FoldNet++ describes 120,000 simulated demonstrations of T-shirt unfolding and folding across six robot designs and 1,000 shirts. A rule-based system uses garment landmarks to generate movements, providing varied training examples without collecting each one on hardware.
The authors report physical transfer to Galbot and ARX robots without real-world fine-tuning for those transfer tests. In a separate ten-trial comparison, the simulation-only training condition completed nine attempts. Main physical experiments generally used 15 trials, with people judging whether the final fold was neat.
This is promising evidence that synthetic demonstrations can teach a long sequence of cloth movements. Its scope covers T-shirts and one folding style; long sleeves, buttons, and other garments remain open challenges. The project has demonstration videos, but labels its code and dataset as coming soon.
Read this first
Read the excavator paper. Its 45-minute field run connects simulation, sensing, and conventional machine control to a visible construction result. The distinction between 201 successful strokes and one full-scale trial makes it especially useful for judging robotics claims.
Read the paper