Robot models are now tested in unseen homes and judged on safety, yet benchmarks still show wide gaps. These figures and dated updates show what progressed, what failed and how to read the claims.
What moved in robot AI this year.
89.4% success on RLBench's 18 simulated tasks, up from about 50% in 2022 (Stanford HAI, AI Index 2026, 2025)
3–10× slower than humans: typical speed of robots in Epoch AI's 2026 review (Epoch AI, 2026)
56% success of Figure's Helix 2.5 in 30 unseen homes (9% without Index) (TechRepublic, 2026)
Figure's Helix 2.5 is tested in 30 unseen homes and succeeds 56% of the time
· Milestone
Figure tested its Helix 2.5 robot policy in 30 Bay Area homes it had never seen, on tidying toys, folding towels and making beds. Across 420 attempts with no partial credit, it succeeded 56% of the time, versus 9% without Figure's Index pretraining. Tidying scored lowest at 40%.
Why it matters: A 44% miss rate shows why home-style pilots need measured success rates on your own tasks, not demo clips. Ask any vendor for trial counts, failure definitions and results from sites the robot has not seen.
Safe-Stop paper teaches a humanoid to judge when stopping will not make it fall
· Research
Researchers posted a method that estimates whether a humanoid can still stop safely from its current motion, then falls back to another policy if it cannot. On a Unitree G1 in simulation it stopped successfully in 96.4% of 179,650 test starts, with 150 trials on hardware.
Why it matters: Emergency stops on legged robots are not automatic. Ask humanoid vendors how their stop behavior was tested, from which motions, and how often real-world trials disagreed with simulation.
Skild AI says its S1 model learns a 10-minute task from one human video
· Launch
Skild AI introduced S1, a robot model prompted with a single human video instead of retraining. It reports 66% per-step success on unseen tasks lasting up to 10 minutes, against 9% for language-prompted policies. The results come from Skild's own benchmarks, and S1 has no public weights or API yet.
Why it matters: A model that learns from one video could shorten setup for site-specific jobs, but today's evidence is vendor-run and the model is not yet public. Ask for independent trials before planning around it.
LIBERO-Safety tests 10 robot AI models on collisions, people and risky instructions
· Research
A new benchmark scores ten vision-language-action models on five safety suites, from careful grasping to refusing hazardous instructions, using 19,664 collision-free demonstrations. Success fell sharply at the hardest level: OpenVLA managed 4.0% to 12.7%, and RoboBrain2.0's refusal rate dropped from 80% to 20%.
Why it matters: Safety is measurable, and today's models degrade when scenes get harder. When a vendor cites an AI model, ask which safety tests it passed, at what difficulty, and whether results are public.
Physical Intelligence unveils π0.7, one model that matches specialists across tasks
· Research
Physical Intelligence released π0.7, a general-purpose robot model that combines skills it learned separately to handle new tasks. The company says one model matches fine-tuned specialists on jobs like laundry folding and espresso making, and folded laundry on a two-arm UR5e robot it had no data for.
Why it matters: One model that transfers between robot bodies would cut the cost of switching hardware. Treat cross-robot claims as research, and check reach and payload for the exact robot on its Spec Card at The Humanoid Group.
AgiBot opens its World 2026 robot dataset, including failed attempts and recoveries
· Launch
AgiBot released the first phase of AGIBOT WORLD 2026, an open-source dataset of hundreds of hours of real-world robot data from homes and commercial spaces. It pairs camera, tactile, lidar and joint data with digital-twin simulations, and keeps annotated error-recovery trajectories from free-form teleoperation.
Why it matters: Open datasets that include failures let smaller teams train and test models without building their own data pipeline. Check the license, which robot body the data came from, and whether it matches your hardware.
People also search for robot AI news, robot foundation model news, vision-language-action model, VLA model news, humanoid robot AI and physical AI news.