Robotics Briefing /
Turning human experience into robot skills.
Human videos can teach robots useful movements, but the details still matter: what to avoid, which object to move, and when to stop. Today’s papers test those questions with converted video, greenhouse experiments, and a memory-guided controller. Industry updates show how the same challenges reach factory floors and warehouse fleets.
Your quick takeaways
- HuRo converts human videos into robot training examples. It improves physical-task completion scores, while its public dataset remains pending.
- ObstaDiff reduces collisions around plants. 2AM improves progress through simulated tasks, but often fails to stop at the right moment.
- Vention targets easier factory deployment, Hai Robotics plans a large warehouse installation, and Unitree releases two embodied-reasoning models.
Papers
HuRo: human videos become robot training examples
Catch-up • Jinho Jeong and colleagues, RLWRLD and Yonsei University • Submitted September 9; authors report CoRL 2026 acceptance • Physical-robot tests
HuRo estimates hand movements in human videos, converts them into robot motion targets, and replaces the visible human with a rendered robot. The team created about 630,000 converted episodes from five video sources. This supplies training examples before a model learns from smaller collections of physical-robot demonstrations.
Across four ALLEX manipulation tasks, adding the full HuRo pretraining set raised average completion from 51.5% to 80.3%. Under changed object positions and visual conditions, the score rose from 34.9% to 72.2%. Three tasks award credit for partial progress, so these figures do not measure the share of fully successful attempts.
The work suggests a way to broaden robot experience without collecting every example on hardware. Converted videos still lack force and touch signals, and their motion targets can contain self-collisions. The conversion code is public, but the repository labels the dataset as coming soon and notes restrictions on commercial use from its dependencies.
ObstaDiff: reaching through plants with fewer collisions
Catch-up • Jiawen Wang, Kevin Yao and Khalid Jawed, UCLA • Submitted September 10; authors report CoRL 2026 acceptance • Physical Sawyer-arm experiments
ObstaDiff separates the target, nearby obstacles, and background in the robot’s camera input. A learned controller uses that structure to approach a pepper through surrounding plants. A recorded movement then handles the final interaction, which the experiment defines as touching the pepper.
Across 61 distinct greenhouse configurations, it completed 46 trials, compared with 30 for a standard image-based diffusion policy. Obstacle collisions fell from 16 trials to five. Tests varied target position, plant layout, and pepper appearance; each configuration ran once per method.
This is useful evidence for reaching into clutter, with an important boundary: the study does not demonstrate harvesting. It learns the approach stage and replays the final motion. Results come from one indoor testbed, and a person or prompt still specifies the target.
2AM: remembering the task and telling the controller where to act
Catch-up • Yutong Hu and colleagues, KU Leuven and Meituan • Submitted September 10 • Preprint; simulation only
2AM gives a high-level agent the task history. That agent sends a separate movement model short instructions plus image points indicating where to grasp, place, or move. The design tests whether clearer guidance can turn remembered steps into useful action.
Across ten simulated LIBERO-Mem tasks, average ordered-step completion reached 76.3%, compared with 70.8% for the authors’ stronger baseline reproduction. Under a relaxed measure that accepts reaching the goal before extra movement, success rose from 37.4% to 63.0%.
Strict success, which also requires stopping at the correct point, was 11.8% versus the baseline’s 12.3%. Better progress therefore did not improve every measure of success. The study highlights why a robot must both reach the intended state and recognize when the job is done; physical-robot validation remains future work.
News
Vention’s new lab targets the work between research and deployment
Catch-up • Announced September 9 • Company research agenda and planned capabilities
Vention opened a Physical AI Lab focused on industrial data collection and adapting robot models to manufacturing tasks. Research director Jimmy Li identifies a practical obstacle: demonstrations collected with one robot and gripper do not automatically transfer to another setup.
He says current cell deployments use separate perception and motion-planning components. The lab hopes to combine those systems with learning from demonstrations by year-end, then expand reinforcement learning in 2027. These are research plans; the interview provides no measured reduction in deployment time.
For factories, the potential benefit is less custom work when adding parts or changing tasks. Testing that benefit requires measuring the data collection, setup, and human assistance each new job still needs.
Hai Robotics plans more than 1,500 rack-climbing robots at one site
Catch-up • Original announcement September 8 • Planned deployment; company specifications
Hai Robotics says an unnamed European fashion retailer selected its HaiPick Climb system for a new fulfillment center. The plan includes more than 1,500 robots and 1.2 million storage locations across a 30,000-square-meter automated operation.
The company specifies throughput above 24,000 totes per hour. The announcement describes a planned installation and supplies no operating logs, commissioning date, or intervention rate. Treat that throughput as a project claim, with performance in sustained operation still to establish.
The scale makes fleet coordination and recovery from blocked equipment especially consequential. Dense storage only helps fulfillment when robots can keep goods moving to people through changing order demand.
Highlights
Unitree releases weights for two embodied-reasoning models
Catch-up • Released September 11 • Public model files; broader robot-control release incomplete
Unitree has released UnifoLM-ER-1 and UnifoLM-ER-Flow weights. ER-1 handles spatial reasoning, such as identifying locations in images; ER-Flow adds predictions of changing scene regions and representations of robot actions. Both model repositories contain downloadable weight files.
These releases give researchers concrete components to evaluate. The broader UnifoLM-WLA project reports a six-billion-parameter controller spanning 64 robot tasks, but its checklist still marks the WLA base model and post-training code as unreleased. The available reasoning weights alone do not provide that complete robot-control system.
Read this first
Read HuRo, then inspect the project’s before-and-after video examples. They make the conversion from human activity to robot training data easy to understand, while the paper explains what that conversion still misses.
Read the paper