Back to Y Combinator

Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club

Y CombinatorJuly 29, 20261h 16m
Topics55
Introduction to YC Kernel and Chip Club0:00Specialization Trends in AI Infrastructure3:31Training vs Inference Data Center Requirements4:32Parallel Kittens: Multi-GPU AI Kernels7:07Single GPU Efficiency vs GPU Networking Bottleneck7:30Hardware Advancements in GPU Networking9:00Multi-GPU Kernel Challenges10:01GPU Fundamentals11:30Single vs Multi-GPU Kernel Optimization13:31Three Key Trade-offs for Multi-GPU Kernel Design14:30All Gather GEM and TMA Optimization16:01Intra-SM vs Inter-SM Overlapping Trade-offs17:01Scheduling Strategy Selection18:02Parallel Kittens Framework19:01Parallel Kittens Performance and Adoption20:31Intelligence Per Watt: Measuring Intelligence Efficiency21:31AI Infrastructure Demand and Investment22:31Shift from Mainframe to Personal Computing Analogy23:31Study on Local Inference Redistribution24:30Intelligence per Watt Framework24:53Study Methodology25:30Key Findings on Intelligence per Watt26:30Intelligence per Joule Improvements27:01Inference Redistribution Implications27:31Limitations of Local Accelerators28:01Future Directions29:00Broader Impact Assessment30:00Kernel Writing Capabilities of AI31:02Programming Language Trade-offs32:30Kernel Programming Language Spectrum33:31Kernel Bot Leaderboard Results34:01Kernel Evaluation Framework36:31Vector Mean Kernel Example37:30Reward Hacking Patterns38:01Kernel Guard Anti-Cheat System39:31QR Decomposition Optimization41:00Kernel Synthesis Challenge43:00Kernel Compilation and Development Challenges45:13Heterogeneous Inference Infrastructure47:00Arithmetic Intensity and Roofline Analysis50:01Inference Workload Diversity53:03SRAM Machines and Memory Architecture54:30Heterogeneous System Disaggregation57:31Attention-MLP Disaggregation59:03Speculative Decoding on Heterogeneous Systems1:01:30End-to-End Heterogeneous Infrastructure Co-Design1:03:01GPU-Based Game Engine Simulation1:04:30Entity Component System Design for GPU1:06:32Task Graphs and GPU Implementation1:10:32GPU Utilization Visualization1:12:18Baseline Environment Development1:12:45Performance Comparison Results1:13:32End-to-End Training Extensions1:14:00GPU Programming Language Challenges1:14:31Session Conclusion1:15:45
In a Nutshell

Multi-GPU kernels require fine-grained compute-communication overlap using TMA for all-gather GEMM and register instructions for all-reduce, with Parallel Kittens achieving hand-optimized performance in 50-100 lines of code versus thousands. Intelligence per watt has improved 3x in two years for local accelerators, with 80-90% of queries potentially routable to local open-source models for 50-70% energy savings. Heterogeneous inference disaggregates prefill-decode and attention-MLP phases across specialized hardware, while GPU-accelerated batch simulators deliver 100x throughput gains for RL environments using entity-component-system architecture.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

The event is themed as the YC kernel and chip club due to the trend of specialization happening at the chip level. The TPU8 I and T version represent specialization into different ASICs (Zebrafish and Sunfish), driven by sufficient demand for such splits. Training data centers have fundamentally different requirements than inference data centers, with training centers not requiring any bandwidth in and out, while inference centers require proximity to users.

There is significant optimization potential remaining on the CUDA side, kernel side, algorithm side, software side, chip side, and data center side. The hosts include Stuart Soul from Stanford's Hazy Research lab, researcher at Cursor training Composer; John from the CS PhD lab at Stanford, co-advised by Aelia, focusing on LMS, MLIS, and hardware, pioneering intelligence per watt and intelligence per joule; Mark, former PyTorch maintainer, co-founder of GPU mode with Casey Elward, co-founding core automation with OpenAI VP research Jerry Torque; Misha with two decades of hardware-software co-design experience, previously running AI infra at NVIDIA and working on hardware-software co-design at Meta; and Brennan from Stanford working on RL for self-driving cars, focusing on GPU acceleration for simulators.

The proliferation of specialization at the chip level is driven by sufficient demand for tokens. The Zebraish Sunfish TPU V8 represents the first visible split, expected to become more pronounced. Pre-existing patterns include inference workloads using Nvidia for prefill and Cerebras for decode engines. A critical but under-discussed distinction exists between batch size one inference and throughput-optimized inference, particularly relevant for latency-sensitive applications like voice agents where 8-second delays are unacceptable. Batch size one prioritization leads to GPU exhaustion and high costs, representing a chip-level issue requiring specialized solutions.

Sign in to read the full notes

Get access to AI-generated notes, topic timestamps, and more.