Robotics Briefing /
Better handoffs, harder tool tests, and robots that revise their plans.
A robot can choose the right tool and still fail to use it. A person can spot a mistake and still struggle to take control smoothly. This edition examines those gaps, along with navigation agents that retain useful experience and a new research partnership connecting robotics with logistics. The common question is how promising capabilities hold up through a complete task.
Your quick takeaways
- DITTO-X gives operators force feedback and aligns their fingers with the robot’s before takeover, helping preserve a grasp during human intervention.
- HumanoidToolBench separates choosing a tool from using it successfully. Walking while working exposes large gaps that stationary tests can miss.
- ASENA improves repeated navigation tasks by revising code and notes. Its physical missions remain supervised, and its strongest learning gains come from revisiting the same tasks.
Papers
DITTO-X: helping people take over without losing the grasp
Zhanpeng He, Joaquin Palacios and colleagues, Stanford and Columbia • September 30 • Weekend catch-up • Preprint with physical tests
When a robot already holds an object, switching to human control can disturb the grasp. DITTO-X uses a powered hand exoskeleton to convey joint forces and fingertip contact. Before takeover, it can move the operator’s fingers into the robot hand’s current configuration, reducing the mismatch at the transition.
The interface supports three commercial robot hands without a separate mechanical redesign for each. It combines continuous force feedback with touch cues from sensors already on the robot, giving operators information that an obstructed camera view can miss.
In a six-person study, participants completed a tong-based data-collection task in 70% of trials with DITTO-X, compared with 46.7% using a Manus tracking glove. Each participant attempted ten trials per interface. Time per successful trial fell from 85.5 to 56.1 seconds.
Robot policies also benefited from demonstrations and subsequent human corrections. On the full tong task, success rose from 63.3% after initial DITTO-X training to 86.7% after two correction rounds, measured over thirty rollouts per run.
The implication is that better human interfaces can improve both immediate control and the data used to train autonomy. The evidence is still narrow: six study participants and a handful of manipulation tasks. Corrective policy training was evaluated on one hand; supporting other hands does not establish equal intervention performance across all three.
HumanoidToolBench: picking the right tool is only the beginning
Kyochul Jang and colleagues, Seoul National University, UMass Amherst and Google Research • October 1 • Weekend catch-up • Preprint
HumanoidToolBench evaluates the full sequence from choosing a suitable tool to using it while stationary or moving. Its eighteen tasks combine three scenarios—moving a ball, retrieving a ball and breaking ice—with different execution demands and competing tool choices.
The distinction exposes a substantial gap. In standard simulated ball-moving tasks, FastWAM achieved 76% success when stationary but 9% when mobile. These are separate task conditions, each tested for one hundred episodes, rather than a controlled estimate of the cost of walking alone.
Physical testing was more limited: three policies attempted stationary ball-moving and retrieval tasks on a Unitree G1, with ten trials per condition. After additional training on ninety-one physical demonstrations, GR00T N1.7 completed six of ten standard ball-moving trials, compared with one of ten for ACT.
The useful contribution is a test that records where execution breaks down: contacting a candidate, lifting the suitable tool, applying it and completing the goal. A correct-looking choice can conceal poor grip placement or alignment.
The benchmark covers only three scenarios, and physical evaluation excludes mobile tool use. Policies also use different pretrained models and action durations. Its results describe these evaluated systems and training recipes; they do not establish a universal ranking of robot intelligence.
ASENA: navigation that improves through reusable code and memory
An-Chieh Cheng and colleagues, NVIDIA and UC San Diego • September 30 • Weekend catch-up • Preprint
ASENA connects coding agents to robot sensors, computation and supervised actions. The agent can inspect what happened, repair a program and retain notes or executable skills. An optional learned navigation model handles movement, while the larger system decides what to inspect and how to respond.
In simulation, ten passes over recurring hundred-task subsets raised success from 72% to 98% on R2R and from 65% to 89% on RxR, two instruction-following navigation benchmarks. Model weights remained fixed.
Those gains came from revisiting the same tasks, scenes, instructions and starting poses with simulator-provided outcome feedback. They show that retained experience helps repeated work; they are not success rates on entirely new missions after training.
Three physical demonstrations on a Unitree G1 involved finding locations or objects, inspecting evidence and returning an answer. Tasks took ten to twenty minutes. Proposed movements passed checks and required operator approval; generated gestures were checked in simulation before execution.
The practical idea is that a robot can accumulate useful procedures outside its model weights. Wider deployment still needs faster execution and stronger safeguards in changing environments. The physical demonstrations establish supervised feasibility, not unattended reliability.
News
Amazon and Georgia Tech establish a joint science hub
October 2 announcement • Weekend catch-up • Research partnership
Amazon will fund joint research through a new Georgia Tech–Amazon Science Hub, with initial interests including autonomous systems and logistics infrastructure. Georgia Tech is also joining Fauna Robotics’ ecosystem and using the Sprout research platform to study mobility, perception, manipulation and interaction with people.
For robotics, the opportunity is to connect academic research with operational problems and a shared physical platform. The announcement describes research support and collaboration, without measured improvements to robot performance or a deployment timetable. The launch celebration took place the previous week; October 2 is the announcement date.
Highlights
ToolBook makes the humanoid tool-use tests inspectable
October 1 paper • Weekend catch-up • Benchmark resources
HumanoidToolBench’s companion ToolBook contains 3,003 simulated demonstrations and ninety-one physical demonstrations. The public repository includes implementation and evaluation documentation, and its dataset page exposes the underlying data structure.
That makes it possible to examine how tools, tasks and demonstrations are organized, rather than relying only on leaderboard scores. The physical data remains a small, task-specific collection; its availability should not be confused with broad coverage of everyday tool use.
Read this first
Read DITTO-X’s intervention examples. They make a concrete point: bringing a person into the loop is more useful when the control interface preserves what the robot is already doing.
Read the paper