Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
In a Nutshell
Ex-NVIDIA engineer is building Sal Research to deliver the industry's cheapest AI tokens by optimizing for throughput over latency, targeting background agent workloads that run for hours or days rather than interactive chatbots. The core thesis is that making intelligence 1000x cheaper unlocks unbounded token consumption in proactive agents, using every available chip, power source, and distributed data center in the US. Key insight: the shift from human-in-the-loop to background processing removes attention constraints, making throughput-optimized stacks and scavenger strategies for underutilized chips/power economically dominant.
These notes were generated by AI and may contain inaccuracies.
The speaker's job is to make tokens as cheap as humanly possible through every layer of the stack. This includes using every chip, every source of power, and every piece of suitable land in the United States. The goal is to move away from treating agents as expensive consultants that you only consult for hard questions. Instead, the aim is to get intelligence into as many hands as possible.
Sal Research operates as a token factory with an API where anyone can send requests to use open source large language models for any task. The service delivers tokens at prices that are unbeatable in the market. The company also supports building agents on top of this infrastructure by hosting what they call sandboxes—long-running agent virtual machines in the cloud designed for agents that run for hours, days, or weeks.
The core positioning is being a peer company to others that serve different kinds of inference, with the goal of being the absolute cheapest provider and enabler of a certain kind of intelligence use.
The company's theme is abundance—delivering intelligence as a new commodity to as many people as possible at costs sustainable for almost every industry. The belief is that making something 10 times cheaper creates a new product category. The aspiration is to do this for tokens. The profound shift is that machines can now think, and the job is to make as many machines as possible work toward thinking.
Sign in to read the full notes
Get access to AI-generated notes, topic timestamps, and more.