Foundation models now let robots attempt unseen homes and ten-minute chores, while safety tests expose what still fails. Follow the progress, then match hardware with The Humanoid Group.
What moved in robot AI this year.
89.4% success on RLBench's 18 simulated tasks, up from about 50% in 2022 (Stanford HAI, AI Index 2026, 2025)
3–10× slower than humans: typical speed of robots in Epoch AI's 2026 review (Epoch AI, 2026)
56% success of Figure's Helix 2.5 in 30 unseen homes (9% without Index) (TechRepublic, 2026)
Figure's Helix 2.5 is tested in 30 unseen homes and succeeds 56% of the time
· Milestone
Figure tested its Helix 2.5 robot policy in 30 Bay Area homes it had never seen, on tidying toys, folding towels and making beds. Across 420 attempts with no partial credit, it succeeded 56% of the time, versus 9% without Figure's Index pretraining. Tidying scored lowest at 40%.
Why it matters: A 44% miss rate shows why home-style pilots need measured success rates on your own tasks, not demo clips. Ask any vendor for trial counts, failure definitions and results from sites the robot has not seen.
Safe-Stop paper teaches a humanoid to judge when stopping will not make it fall
· Research
Researchers posted a method that estimates whether a humanoid can still stop safely from its current motion, then falls back to another policy if it cannot. On a Unitree G1 in simulation it stopped successfully in 96.4% of 179,650 test starts, with 150 trials on hardware.
Why it matters: Emergency stops on legged robots are not automatic. Ask humanoid vendors how their stop behavior was tested, from which motions, and how often real-world trials disagreed with simulation.
Skild AI says its S1 model learns a 10-minute task from one human video
· Launch
Skild AI introduced S1, a robot model prompted with a single human video instead of retraining. It reports 66% per-step success on unseen tasks lasting up to 10 minutes, against 9% for language-prompted policies. The results come from Skild's own benchmarks, and S1 has no public weights or API yet.
Why it matters: A model that learns from one video could shorten setup for site-specific jobs, but today's evidence is vendor-run and the model is not yet public. Ask for independent trials before planning around it.
AI Risk Management Framework— National Institute of Standards and Technology (NIST). Voluntary US framework for mapping, measuring and managing AI risks, with a playbook and companion profiles.
The 2026 AI Index Report— Stanford Institute for Human-Centered Artificial Intelligence (HAI). Annual data on AI research, technical performance, economics and policy, with public datasets behind each chapter.
On the Opportunities and Risks of Foundation Models— Stanford Center for Research on Foundation Models (arXiv). The Stanford report that introduced the term foundation model, covering capabilities, training methods and societal risks.