Back to Y Combinator

Why Robotics Still Isn't Solved - But Could Be Soon | YC Paper Club

Y CombinatorAugust 8, 20261h 24m
Topics56
Opening: The Persistent 'Next Year' Robotics Prediction0:05Four Scaling Walls in Robotics2:00Speaker Introductions7:32Marcel: Multiscale Embodied Memory (MAM)8:01Memory Integration Challenges10:31Short-Term Visual Memory Architecture12:00Long-Term Memory via Recurrent Text Predictions13:31In-Context Adaptation Through Memory15:01Q&A: Memory Annotation and Implementation16:32Milan Gennai: Self-Supervised Bootstrapped Embodied Reasoning20:30Research: Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning22:31Embodiment-Specific Chain of Thought Reasoning24:06Evaluation Across Embodiments26:32Self-Driving Application28:01Key Takeaways28:30Discussion: Non-Textual Reasoning and Latency29:31Sim Tool Reel: Simulation to Real Tool Use34:00Goal-Conditioned Policy36:33Zero-Shot Generalization38:31Baseline Comparisons39:30Play to Perfect41:30Discussion: Recovery Behaviors and Generalization42:32Simulation Evaluation Challenges47:33LSTM Architecture Decision48:01Transformer vs LSTM Comparison48:30Goal Generation Process49:31Pose Tracking Failure Analysis50:00Rerun Company Background51:31Robotics Application Companies52:32Business Model Strategy53:31Paper Plane Factory Example55:01Learning Infrastructure Setup57:02Data Collection Strategy58:30Physical Data Infrastructure Challenges1:00:00Scaling and Iteration Process1:02:01Successful Company Characteristics1:03:30Market Opportunity1:04:30Early Success Areas1:05:30Market Timing Explanation1:06:31Data Scale Estimation1:07:01General Instinct Company Overview1:08:30World Action Models vs VLAs1:09:00Performance and Cost Analysis1:10:00World Model Architecture: Training and Inference Pipelines1:10:47Alternative Approaches to Full Video Prediction1:12:00Generative vs Latent World Action Models1:13:00Maintaining World Representations1:13:30Infrastructure Optimizations1:14:31Modality Alternatives for World Representations1:15:32Flow Matching Inference Mechanics1:17:00Future Kinematics Learning1:18:00Architecture Efficiency and Cross-Attention1:20:00Flow Matching as Test-Time Planning1:21:00Business Model Rationale1:21:30Autoregressive Prediction and Chunk Size1:22:31Temporal Difference Encoding1:23:30
In a Nutshell

The core message is that robotics faces four fundamental scaling barriers—physics modeling gaps, deformable object dynamics, sensory-motor deficits, and embodiment drift—preventing the "next year" predictions that have persisted for a decade. The most important breakthroughs presented include memory-augmented policies that decompose tasks into high-level text memory and low-level visual execution, self-supervised bootstrapping of embodiment-specific reasoning that prunes useless annotations while improving out-of-distribution performance, and simulation-trained goal-conditioned policies that achieve zero-shot tool manipulation without teleoperation data. These advances collectively suggest that solving robotics requires architectural decomposition and selective reasoning rather than pure scaling of existing vision-action models.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

The discussion opens with a 10-year history of robotics being predicted as solved 'next year.' AlphaGo's release prompted claims that scaling the algorithm would solve robotics. MuJoCo enabled training robots to walk in 3,000 iterations, reinforcing the prediction. The ALOHA system was hailed as a breakthrough, with demonstrations of watering plants, fixing bikes, and operating a Keurig. The diffusion policy paper and VALAS were cited as evidence that 2026 would be the year of robotics. However, halfway through 2026, only pre-orders for Neo 1X are available, with no consumer access to Pi or Figure robots yet.

Four primary barriers to scaling robotics are identified. The first is physical real-world modeling. Video models trained on games like Doom fail to respect physics when deployed in real-world scenarios, such as driving a car into a grocery store that magically transforms into a highway. The sim-to-real gap remains unsolved.

The second barrier involves deformable objects and action-conditioned dynamics. Estimating the transition function from state to state-plus-one becomes significantly harder when conditioned on actions, requiring substantially more data. Feature pyramid networks developed in 2013 for robotic policy bot war simulations addressed representation learning, but action space representation remains unsolved for rapid learning.

The third barrier is the sensory-motor gap. Humans possess nerve endings detecting normal force, tangent force, moisture, temperature, vibration, and friction coefficient across the entire body. Current robots have at most one coin force-torque sensor per fingertip and a wrist camera. Neuroscientists note humans build world models without vision, such as identifying backpack contents by touch alone. The absence of an artificial epidermis prevents robots from achieving comparable tactile capabilities.

Sign in to read the full notes

Get access to AI-generated notes, topic timestamps, and more.