Back to Invest Like The Best

Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper

Invest Like The BestAugust 25, 20261h 23m
Topics45
Token Cost as the Core Mission0:00What Sal Research Builds0:32The Theme of Abundance1:30Token Cost as the North Star2:01Why the Opportunity Exists2:31Long Horizon Tasks as the Future4:00Test Time Compute Scaling5:30Market Share Projections6:02Background Task Examples6:31Dream Use Cases9:00Non-Verifiable Tasks12:31The Master Plan: Software, Hardware, and Power13:00Tensor Cores Explained13:31Nvidia Culture and Retention17:00Throughput vs Latency Trade-off18:02Why the Trade-off is Unbreakable19:00GPU Parallelism and Interconnect Strategies20:47NVLink's Role in Latency Performance22:01Cerebras and Alternative Memory Architectures23:31KV Cache Limitations and Hybrid Architectures28:01Transformers Architecture and Scaling Properties32:31Data Evolution and Self-Improvement36:32Kernel Development and Software Optimization39:01Kernel Engineering and Efficiency Optimization42:10The Shift to Rack-Scale Computing43:30Chip Market Dynamics and Allocation44:30Alternative Chip Markets and AMD Opportunities46:00Heterogeneous Computing Strategy47:32Investor Concerns and Market Valuation49:00Training vs Inference Spend Patterns50:30Data Center Evolution and Power Distribution53:00Physical Scale and Distributed Architecture55:00Power Source Innovation and Intermittency Tolerance59:00Scavenger Strategy and Vertical Integration1:00:00Custom Hardware for Inference1:03:14Ranking Inefficiencies in Token Production1:03:31Fab Capacity and Process Optimization1:06:01Building a Performance-Focused Organization1:08:01Closed vs Open Source Dynamics1:10:30Vision for Abundant Intelligence1:13:01Divergent Technical Views1:14:30Consensus Challenges1:16:31Advice for Compute Startups1:19:01Nvidia's Strategic Positioning1:20:30Mentorship and Full-Stack Understanding1:21:30
In a Nutshell

Ex-NVIDIA engineer is building Sal Research to deliver the industry's cheapest AI tokens by optimizing for throughput over latency, targeting background agent workloads that run for hours or days rather than interactive chatbots. The core thesis is that making intelligence 1000x cheaper unlocks unbounded token consumption in proactive agents, using every available chip, power source, and distributed data center in the US. Key insight: the shift from human-in-the-loop to background processing removes attention constraints, making throughput-optimized stacks and scavenger strategies for underutilized chips/power economically dominant.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

The speaker's job is to make tokens as cheap as humanly possible through every layer of the stack. This includes using every chip, every source of power, and every piece of suitable land in the United States. The goal is to move away from treating agents as expensive consultants that you only consult for hard questions. Instead, the aim is to get intelligence into as many hands as possible.

Sal Research operates as a token factory with an API where anyone can send requests to use open source large language models for any task. The service delivers tokens at prices that are unbeatable in the market. The company also supports building agents on top of this infrastructure by hosting what they call sandboxes—long-running agent virtual machines in the cloud designed for agents that run for hours, days, or weeks.

The core positioning is being a peer company to others that serve different kinds of inference, with the goal of being the absolute cheapest provider and enabler of a certain kind of intelligence use.

The company's theme is abundance—delivering intelligence as a new commodity to as many people as possible at costs sustainable for almost every industry. The belief is that making something 10 times cheaper creates a new product category. The aspiration is to do this for tokens. The profound shift is that machines can now think, and the job is to make as many machines as possible work toward thinking.

Sign in to read the full notes

Get access to AI-generated notes, topic timestamps, and more.