Why Robotics Still Isn't Solved - But Could Be Soon | YC Paper Club
In a Nutshell
The core message is that robotics faces four fundamental scaling barriers—physics modeling gaps, deformable object dynamics, sensory-motor deficits, and embodiment drift—preventing the "next year" predictions that have persisted for a decade. The most important breakthroughs presented include memory-augmented policies that decompose tasks into high-level text memory and low-level visual execution, self-supervised bootstrapping of embodiment-specific reasoning that prunes useless annotations while improving out-of-distribution performance, and simulation-trained goal-conditioned policies that achieve zero-shot tool manipulation without teleoperation data. These advances collectively suggest that solving robotics requires architectural decomposition and selective reasoning rather than pure scaling of existing vision-action models.
These notes were generated by AI and may contain inaccuracies.
The discussion opens with a 10-year history of robotics being predicted as solved 'next year.' AlphaGo's release prompted claims that scaling the algorithm would solve robotics. MuJoCo enabled training robots to walk in 3,000 iterations, reinforcing the prediction. The ALOHA system was hailed as a breakthrough, with demonstrations of watering plants, fixing bikes, and operating a Keurig. The diffusion policy paper and VALAS were cited as evidence that 2026 would be the year of robotics. However, halfway through 2026, only pre-orders for Neo 1X are available, with no consumer access to Pi or Figure robots yet.
Four primary barriers to scaling robotics are identified. The first is physical real-world modeling. Video models trained on games like Doom fail to respect physics when deployed in real-world scenarios, such as driving a car into a grocery store that magically transforms into a highway. The sim-to-real gap remains unsolved.
The second barrier involves deformable objects and action-conditioned dynamics. Estimating the transition function from state to state-plus-one becomes significantly harder when conditioned on actions, requiring substantially more data. Feature pyramid networks developed in 2013 for robotic policy bot war simulations addressed representation learning, but action space representation remains unsolved for rapid learning.
The third barrier is the sensory-motor gap. Humans possess nerve endings detecting normal force, tangent force, moisture, temperature, vibration, and friction coefficient across the entire body. Current robots have at most one coin force-torque sensor per fingertip and a wrist camera. Neuroscientists note humans build world models without vision, such as identifying backpack contents by touch alone. The absence of an artificial epidermis prevents robots from achieving comparable tactile capabilities.
Sign in to read the full notes
Get access to AI-generated notes, topic timestamps, and more.