Naveen Rao: 4D Computing, AI's Energy Wall & Beating Biology
In a Nutshell
Naveen Rao's Unconventional AI is building the first physical dynamical computer that integrates memory and compute in single oscillator elements, eliminating the von Neumann bottleneck that moves trillions of bits per second and consumes 50% of AI's energy costs. Their chip generates images at 500 nanojoules each—thousands of times more efficient than GPUs—by using time-varying dynamical systems instead of matrix multiplication. The company aims to reach within 1-2 orders of biological efficiency in 3.5 years, enabling distributed AI infrastructure rather than gigawatt data centers.
These notes were generated by AI and may contain inaccuracies.
Naveen Rao is co-founder and CEO of Unconventional AI, an AI chip startup. He is best known for building and selling two deep tech companies, including Nirvana Systems, the first AI chip company founded in 2014. He sold the company to Intel and subsequently ran Intel's AI group. After departing Intel in 2020, he built a GPU platforming company that joined Databricks in 2023, contributing to a quarter of Databricks' total revenue.
Naveen positions himself as "the opposite of an AI doomer," viewing AI as the next evolution of humanity that requires hardware substrate innovation to achieve true intelligence.
Naveen became an electrical engineer due to his early interest in science fiction and building intelligent machines. He later pursued a PhD in neuroscience to explore how to make computers intelligent. He describes his current work as being "right where I wanted to be my whole life."
Google processes 3.2 quadrillion tokens per month. At 10 joules per token, this consumes 12 gigawatts of power. The US allocates 40 gigawatts to data centers, representing half of global data center capacity. Currently, data centers worldwide consume under 100 gigawatts.
Energy now represents approximately 50% of the cost of serving a token. The business case for Unconventional AI is to monetize energy 1000x more efficiently than existing hardware.
The human brain operates on approximately 20 watts. A monkey brain runs on 1 watt, equivalent to a cell phone's power consumption. A squirrel brain operates on 8 milliwatts while performing precise actions such as jumping between branches with 100% accuracy.
The human cortex moves only 16 billion bits per second despite containing 13-14 billion neurons. A GPU moves nearly 30 trillion bits per second in and out of memory, with internal chip movement estimated at 10-100x more. This massive information movement drives the energy demands of current computing systems.
Computers have maintained fundamentally the same architecture since the 1940s, featuring separate memory and compute units with constant data movement between them. The ENIAC computer built in 1945 for artillery trajectory calculations established this paradigm, prioritizing speed over energy efficiency.
Moore's Law has largely ended, with efficiency gains from transistor scaling no longer occurring. Single-thread performance and frequency scaling have also plateaued.
Digital computing represents an abstraction of physical reality where transistors with continuous states are forced to behave as binary 1s and 0s. Each abstraction layer loses information about underlying complexity, creating inefficiency. Neural networks built on these abstractions compound this problem.
Nature demonstrates computation through simple rules creating emergent behavior, such as bird flocking patterns and ant colony intelligence. The brain operates similarly, where neuron physics gives rise to intelligence without linear algebra or floating-point operations.
Multiple metronomes placed on a movable plank synchronize their phases through physical interaction alone. This demonstrates dynamical systems where simple physical behaviors produce emergent synchronization without explicit programming.
The UNO image generation model was built using oscillator-based dynamical systems and released as open source. Analysis of state space trajectories shows how the system evolves differently when conditioned on outputs like "airplane," "car," or "bird."
Sparsity research revealed that removing connections from fully-connected systems (n² scaling problem) can improve both efficiency and trainability. This applies to both simulated and physical systems.
The first physical dynamical computer was built in five months, with design taped out on June 1st and chips returned for testing. The chip generates images using only 500 nanojoules per image, compared to millijoules required by GPUs - representing multiple orders of magnitude improvement in efficiency.
The dynamical computer integrates compute and memory in single elements, eliminating memory interfaces. Each computing element serves as memory. This approach uses the time dimension in dynamics plus three physical dimensions through die stacking, creating "4D computing."
The company aims to reach within 1-2 orders of magnitude of thermodynamic efficiency limits within three and a half years, ultimately seeking to exceed biological efficiency. Current AI systems are approximately 10 billion times away from thermodynamic limits.
The vision includes shifting from gigawatt-scale data centers to distributed smaller facilities, enabling billions of robots for dynamic problem-solving. Jevons Paradox suggests that 1000x efficiency improvements will create markets larger than the trillion-dollar AI market projected for 2030.
A full product is expected within two years, taking the form of a complete rack system functioning as a VM. Tokens flow in and out through network connections while internal architecture differs completely from conventional computers.
Existing model families can be supported through model-layer porting rather than operation-level porting, though significant compute is required for transitions. The system implements matrix operations as time-varying behaviors rather than traditional matrix multiplication.
The team combines dynamical systems theorists with chip designers - disciplines that traditionally don't collaborate. Python libraries enable expression of time-varying elements with stochastic behavior, serving as the interface between theoretical concepts and physical implementation.
Naveen Rao discusses the concept of 4D computing as a fundamental shift beyond traditional three-dimensional architectures. The approach involves adding a temporal dimension to computation, allowing systems to process information across time as well as space. This represents a departure from conventional von Neumann architectures that separate memory and processing.
The energy wall in AI refers to the physical limitations of current computing systems as they scale. Moore's Law has provided exponential improvements in transistor density, but power consumption and heat dissipation are becoming insurmountable barriers. AI training runs are consuming megawatts of power, and the trajectory suggests that future models will require power plants dedicated solely to their operation.
Biological Computing Comparison
Rao compares current AI systems to biological systems, noting that the human brain operates on approximately 20 watts of power while performing tasks that would require megawatts in silicon-based systems. The brain achieves this efficiency through massive parallelism, event-driven computation, and integrated memory-processing architectures. Biological neurons operate on electrochemical gradients rather than binary switching, providing inherent efficiency advantages.
The fundamental difference lies in how biology handles the energy-information tradeoff. Biological systems optimize for survival and reproduction rather than maximum computational throughput, leading to architectures that are dramatically more power-efficient than current AI implementations.
Overcoming the Energy Wall
Several approaches are being explored to address AI's energy constraints. Neuromorphic computing attempts to replicate biological neural architectures in silicon, using spiking neurons and event-driven processing to reduce power consumption. Photonic computing uses light instead of electrons for information transfer, potentially achieving higher bandwidth with lower energy per bit.
Quantum computing is mentioned as another potential solution, though Rao notes that current quantum systems are far from practical deployment for general AI workloads. The error rates and decoherence times remain prohibitive for large-scale applications.
Future Implications
The energy wall suggests that current scaling trends in AI cannot continue indefinitely. This may lead to a bifurcation in the field, with some applications remaining in the cloud while others move to edge devices with severe power constraints. The economics of AI deployment will increasingly be dominated by energy costs rather than computational costs.
Rao emphasizes that beating biology's efficiency will require fundamental advances in materials science, architecture design, and algorithmic approaches. Incremental improvements to existing paradigms are unlikely to close the efficiency gap between silicon and biological systems.
Keep All-In Podcast in your library
Save the videos and channels worth coming back to, and find them again in one place.





