Robotics Briefing /
A grasp reflex, a steadier stride, and learning from what comes next.
A robot hand can secure an object without seeing it. A humanoid can slow down before a staircase. A manipulation model can learn from future scenes without generating a video for every move. Today’s papers show how these ideas translate into physical demonstrations, with important limits on what the numbers prove.
Your quick takeaways
- A blind grasp reflex uses finger-joint feedback to secure objects. It grasped 29 different objects on hardware, but that demonstration is not a measured 100% success rate.
- Caltech’s humanoid controller combines depth cameras with learned motion planning to climb stairs and jump onto boxes. Its reported 25-point improvement comes from simulation.
- InternW0-Δ combines more than 23,000 hours of robot and human data. Small physical comparisons show a benefit from pretraining; code and model downloads remain forthcoming.
Papers
See to Reach, Feel to Grasp: a reusable reflex for robot hands
Alexander Alexiev and colleagues, MIT and collaborators • September 25 • Preprint; September 28 arXiv listing
Many robotic hands depend on a clear view of an object throughout a grasp. This team separates the job: an arm controller positions the hand, while a learned hand controller uses joint-position history and command feedback to close and stabilize the fingers. It needs no hand-camera images, object geometry or dedicated touch sensors.
The hand also estimates when the grasp is secure, signaling the arm to lift or begin another motion. The same frozen controller grasped 29 physical objects and worked with three independently designed arm behaviors, including pick-and-place and screwing in a lightbulb.
Those hardware demonstrations establish variety, not a repeat-trial reliability score. The reported 96% on YCB objects and 95% on GraspXL objects are simulation results. The authors explicitly did not run repeated trials for each physical object.
The potential value is a grasping skill that can work with different reaching systems and tolerate an obscured contact area. The arm still needs guidance toward the object; the whole robot is not operating without perception. The project’s code is marked coming soon.
Generate, Track, Improve: choosing a motion that fits the terrain
Zachary Olkin, William D. Compton and Aaron D. Ames, Caltech • September 25 • Preprint; September 28 arXiv listing
A humanoid needs to recognize a staircase early enough to change its stride. This system uses two depth cameras, a model that generates whole-body motion plans, and a separate controller that follows them. Reinforcement learning then improves the generator by giving more weight to plans that work.
In simulation, that refinement increased successful terrain crossings by up to 25 percentage points. On a physical Unitree G1, the team demonstrated walking, running, box jumps and six different staircases, including one with fifteen consecutive steps.
The robot can reduce its speed as terrain approaches, then return toward the commanded speed. That matters for movement through buildings, where repeatedly switching between flat ground and steps is more useful than performing one isolated maneuver.
The physical demonstrations do not establish the same 25-point gain on hardware. A joystick supplies velocity commands, and the system uses separate Jetson Thor and Orin computers. Depth also reveals shape without understanding meaning: the authors note that the controller cannot distinguish a box suitable for jumping from a similarly shaped object it should avoid.
InternW0-Δ: using predicted change to improve robot actions
Physical Intelligence Team, Shanghai AI Laboratory • September 25 • Preprint; September 28 arXiv listing
World action models connect predictions about a scene with a robot’s movements. InternW0-Δ combines a video model, an action model, language-and-image understanding, and geometry guidance. During training, future observations teach it which changes matter. During execution, it chooses actions without generating a future video.
The team reports 23,072 hours of processed data spanning robot recordings, handheld demonstrations and human viewpoints. The same pretrained model is then adapted separately to each robot and task.
In two physical comparisons with twenty trials per condition, pretraining raised bread-toasting success from four to nineteen successes and a laboratory liquid-handling task from zero to nineteen. Those are encouraging small comparisons against versions without pretraining, not evidence of general household or laboratory autonomy.
Simulation results also vary sharply by benchmark: 92.8% on LIBERO-Plus and 23.9% on the more demanding RoboDojo suite. The authors have not fully isolated the contribution of different human-data sources. Although the paper promises an open release, the project currently marks source code and model checkpoints as forthcoming.
News
Seyond brings its robotics sensing portfolio to IROS
Seyond • September 25 • Industry catch-up; company announcement
Seyond is showcasing laser-based 3D sensing for mobile robots, forklifts and humanoids at IROS. Its Hummingbird D1R is designed for robotics, with a stated 140° by 100° field of view and close-range detection.
Broad coverage in a small package can help robots perceive nearby obstacles when installation space is limited. Seyond also describes smaller, lower-power sensors under development.
This is a product-positioning announcement, not an independent navigation benchmark. The release provides no measured improvement in robot collision rates or uptime. Its claim of more than one million sensors delivered covers the broader business, not one million robotics deployments.
Highlights
Cybflight: fast flight with control on one microcontroller
Yifan Lin and colleagues, University of Toronto • September 25 • New preprint and available open-source code
Cybflight is a modular autopilot written in Rust. It runs state estimation and nonlinear flight control on one STM32H743 microcontroller, without a companion computer. The design makes individual control algorithms replaceable while keeping the embedded system compact.
The paper reports a physical outdoor flight reaching 31.4 meters per second, about 113 kilometers per hour, using satellite-based positioning. Indoor tests used motion capture and reached 12.38 meters per second.
The result shows that sophisticated control can fit on modest onboard hardware. These flights followed planned paths with external positioning; they do not demonstrate navigating unknown obstacles or relying on onboard vision alone. The public repository contains implementation and setup documentation under an Apache 2.0 license.
Read this first
Explore the See to Reach, Feel to Grasp project. Its demonstrations make the division between reaching and grasping easy to understand, while the paper clearly separates simulation success rates from qualitative hardware results.
Read the paper