Back to Y Combinator

5 Papers That Show Where AI Research Is Heading Right Now

Y CombinatorJune 12, 20261h 16m
Topics49
Introduction and Overview0:08Call for Presentations0:31Views on Human-Generated Subspace1:00AlphaGo vs AlphaZero and the Solution Space1:30Intelligence Per Sample and Per Watt2:32Experiments with LoRA and Training Approaches3:01Human Learning vs Current AI Approaches4:01Intelligence Per Watt and Learning Procedures4:30Interests in Novel Breakthroughs and Biology5:01Club Improvement Ideas5:30BioAI Talk Introduction6:00High-Level Pitch6:31The Bitter Lesson and Biology7:01Scaling Laws in Protein Biology7:31Three Vignettes8:30Biology Background on Proteins9:02ESMC Model Training10:01Learning Protein Grammar10:30Prior Work and Table Overview11:00Three Questions12:01Question One: Do Scaling Laws Hold12:01Performance Against Training Compute13:00Previous Models and Data Scaling14:01Question Two: Bitter Lesson Evaluation15:30ESM Approach Without MSA16:30Looped Model Architecture17:00Headline Results18:00Performance on Protein Complexes and Antibodies18:32MSA Status and Test-Time Compute19:30Mechanistic Interpretability Analysis20:30Protein Model Latent Space Analysis20:35Protein Atlas Construction22:31Scaling the Bitter Lesson to Biology23:01Post-Training Reinforcement Learning for Large Language Models25:31Self-Play for Language Models28:00The Promise and Reality of Self-Play30:00Baseline Self-Play Algorithm and Its Failure Mode32:00Self-Guided Self-Play (SGS)35:00SGS Experimental Results36:02Streaming RAG for Voice AI37:32Optimizing RAG for Streaming Audio Queries44:03Lean and the Era of Verified Intelligence47:30Rethinking Software Engineering with Agentic Systems58:32Managing Multiple Agents with RTS-Inspired Workflows1:08:40APM Tracking for Agent Productivity1:12:01Token Efficiency and Parallel Execution1:13:30Building and Maintaining Knowledge Bases1:14:02Team Principles and Output Metrics1:15:00Closing Remarks1:16:06
In a Nutshell

The video argues that AI progress hinges on maximizing intelligence per sample and per watt, rejecting the idea that scaling test-time compute or self-improvement on human data alone will reach the full solution space. Key evidence comes from protein language models that follow clean neural scaling laws once data volume reaches billions of sequences, enabling near-AlphaFold performance without handcrafted MSAs, plus mechanistic interpretability showing the models spontaneously learn real biological features. Parallel advances in self-play RL, streaming voice RAG, Lean-based formal verification, and RTS-style agent orchestration show the same pattern: general scaling beats domain-specific engineering when data, search, and verification can be made cheap and automatic.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

Thank you guys so much for coming. This one will have a much more applied bent based on the feedback. We have a bunch of really cool people that will be introduced. The session covers AI for biology by Yas Beg. Luke out of Tatsu's lab will talk about self-play, Alpha Zero style self-play for LLMs. Arnob will present on stream rag, a super different application thinking about real-time voice agents. Robert George is working on lean for science. Luke Worthwine is the AI token maxer.

A call for presentations is made to inspire interest and encourage others to present. Memory has been the hot topic for at least the last year and a half. There have been many papers from mem zero to recursive language models, cartridges out of the lab, hnet, and dynamic chunking stuff. There are many different ideas in this area.

A Nome Brown podcast launched a couple weeks ago where the view is that the human-generated subspace H is still viable—if we train on that we can test-time compute our way out of it and recursively self-improve all the way to F minus H. There is a struggle with this view and it does not seem probable that we will sample all of that. It is not that it will not be possible, but it is just not probable that we will sample all of it.

The left side is AlphaGo, the right side is AlphaZero. AlphaZero, unbiased by humans meandering, is the way to get to much more intelligent systems, maybe even AGI. If the full solution space F is F, training on known human solutions will limit you to some typical set H despite any feasible amount of test-time compute or recursive self-improvement. You won't feasibly sample F minus H, especially all of it. If there is infinite recursive self-improvement and infinite test compute, maybe, but we do not have infinite.

Sign in to read the full notes

Get access to AI-generated notes, topic timestamps, and more.