The Future of Frontier Model Architectures with Walter Goodwin, Founder & CEO of Fractile
In a Nutshell
Fractile is building a new class of inference chips that deliver 25x higher memory bandwidth than HBM-based GPUs by pairing high-speed DRAM with custom architecture, targeting the rapid scaling of model parameters and context lengths. The company’s core bet is that memory bandwidth—not compute—has become the primary bottleneck, and unlocking dramatically higher bandwidth will enable fundamentally faster reasoning and new deployment paradigms for frontier models. To execute this, Fractile maintains a vertically integrated design team that can iterate architectural bets on a 3–6 month cycle, giving frontier labs a structural speed advantage over slower, outsourced ASIC efforts.
These notes were generated by AI and may contain inaccuracies.
Currently, companies are trying to build a single chip, but Nvidia systems contain between six and nine custom chips working together to build something extremely powerful. There is already a gap here. The idea of being able to have a sophisticated front of stakes at any given moment, which will be deeply compatible and ready to activate its escalation, represents a tremendous advantage if you can build that apparatus and that engine. This is equivalent to the advanced front model in the field of chips. If you can find a way to gain a structural advantage for 3 to 6 months, you will win all those deployments.
Hello listeners. Welcome back to "No Priors." Today, Sarah is here with Walter Goodwin, founder and CEO of Fractile, an AI chip integrator. They discuss what it means to be a technically integrated company, the technical bets they are making, why they are so focused on memory bandwidth, their expectations for future advanced model architectures, the structure of the front-end chip market today and tomorrow between Nvidia and AMD, internal efforts, and this new class of accelerators.
Walter founded Fractile, a company specializing in chips. They build very fast inference chips for the world's largest models. This reliance on speed is something they adopted from the beginning. The company started in the summer of 2022, and at that time they started to see two things: the arrival of basic, online-trained models that are generalized across everything, and some wise people saying they needed to find a way to pump more computing power into these models during testing time. They needed to find a way to take what they had in AlphaGo, where you have a capable neural network but it becomes superhuman when scaled up, and apply that to language modeling.
Fractile's great endeavor over the past four years has been to find a way to build chips that allow them to simultaneously handle these massive models and run them much faster than current chips. They want to do so in a scalable way, scaling up to models that go beyond current limits, and scaling up to exceptionally long contexts.
The chip scene has become much more interesting in the four years since Fractile was founded. There is currently an exceptional variety of options available. If you look at the field of AI ASIC circuits, one of the striking things is that there are a lot of nearly identical chips. This is a structural feature of the industry. If you look particularly at the efforts of the big companies (hyperscalers), Google started this trend with TPU units more than 10 years ago. You will see this model where there are many proprietary internal chips, such as Google's TPU, Meta's MTIA, Microsoft's Maya, and now OpenAI's Halapenio, chips that are developed and ultimately delivered in partnership with a relatively small number of what might be called an "ASIC design role." These are the companies that handle the final deliveries. Broadcom is the largest, a company valued at $2 trillion. It helps other parties implement these designs.
When you examine this apparent abundance in the range of what exists today, you will find that there is actually a great deal of similarity between these platforms. All of these chips use HBM memory, which is a type of DRAM memory common to graphics processing units from Nvidia, AMD, and all those other circuits. You have the same bets on Tensor Cores for performing matrix multiplication, and the same advanced packaging technologies with TSMC. One of the things noticed in this landscape is the continued relative lack of efforts that encompass the entire silicon ecosystem and attempt to build fundamentally new capabilities. This is structural in nature. There are only a few teams that have decided to say: "We will build starting from what you might call the architectural layer, the front-end design layer, which mostly looks like writing code, all the way down to what we call physical design, processing technology, and factory interactions, which are the things that make you travel back and forth to Taiwan or Korea every week." This is something that is still actually limited to a relatively small number of companies.
If you, say Google, are designing a Tensor Processing Unit (TPU), you have a number of people on your team who have a very strong understanding of the workloads they are trying to speed up. Going back more than 10 years, this is a world that makes you envision things like the "Tensor Core," a custom circuit created by an architect, that is exceptionally adept at matrix multiplication, because you have a vision that matrix multiplication represents the vast majority of the number of calculations in these models. This is the area of expertise of architects, and they are people who really understand how chips work, and ideally have a very strong understanding of the type of workloads targeted. What you will then see within those institutions is a group of very intelligent people who are converting that into what is considered a circuit-level description of how this chip works. Much of this is what is called front design. Mechanically, it's still similar to writing code on a computer, which is then delivered to a player like Broadcom.
That kind of front-end design that describes the basic intent of the chip and the logic behind it is eventually converted into something that is sent to TSMC, which is literally a kind of "bitmap." There is a type of coil called GDSII, which literally specifies where the metal layers are placed, and where each transistor is placed individually. Therefore, it is a complete schematic design. It goes through the process of installing this RTL in a set of circuits that are ultimately planned. A lot of the complications there revolve around, in the case of Broadcom, ownership of the analog intellectual property used for communication between chips. This type of physical layout is specific to a particular processing node in TSMC. So, talking about 3 nanometers, 5 nanometers, and so on. Therefore, this part is still, in general, a kind of outsourced activity for all these projects.
Creating a new chip from start to finish is a very large project to undertake. From a workload perspective, Fractile is working with several leading companies to define that and ensure they are able to serve them. There are many industries that follow the waterfall method of getting things done, and then suddenly, a new way of thinking comes along that is more flexible. For Fractile as a full-service company, they have a team that has a very deep understanding of workloads. They are actually trying to move forward in many places and look at where they can change the structure of the model to be exceptionally compatible with their bets. They think of a vector of the law of measurement for the specific bets they make. They can then advise their clients and partners on this. It also helps them direct a range of bets towards a chip that will not necessarily be available in large quantities for a year or two. This becomes a very important part of the capabilities that they need to build.
When looking at this ability to create a flexible workspace, it means that Fractile has front-end designers within the company. They have their own physical design team. They have their own back-end execution team. They do their own thing in advanced packaging, for example. That's not a huge number of employees. Fractile today has about 150 people. They are somewhat limited in each of those sectors. But what this allows them to do is have a more flexible closed loop. This has become increasingly necessary given the pacing needs imposed by this industry. The relentless pursuit of workloads, where in the field of AI chips, more than in any other chip field before, you really have to bet correctly. Therefore, you must have a lot of skill and a lot of luck. You also need to move very quickly. This type of organizational structure, where they own the entire story internally, is completely different from the previous approach which involved a point of delivery. You reach a certain level and then hand it over to another partner. To some extent, you become at the mercy of how that other partner behaves.
One of the most important technical bets the company has made is orientation. Four and a half years ago, they were in the midst of the deduction story, but they felt that they were teaching the world the meaning of deduction for almost two years afterward. Some of today's inference chips, including one that was recently introduced, were training chips until about 18 months ago. There was a real avoidance of the idea of small marginal cost, and marginal cost has two meanings. It may seem like a small cost, but what it really means is the cost you pay every time you publish these templates. The primary bet was that they would slip into some kind of publishing error. The bet is on speed. That was something that really helped shape much of what they did later as an architectural response.
On the architectural side, they went through a journey. In the first two years of the company's existence, they were working like Grok or Cerebrus on a chip based on SRAM. They have come to the observation that SRAM is a very high-bandwidth memory. It is located on the same piece of silicon as your logic unit, and therefore you have this extremely high bandwidth between the computing unit where the calculations are being performed while these models are running, and the model weights or the key-value cache (KV cache) while you are subtracting the model. This is what can then push you to reach thousands of symbols per second in these language models.
One of the things they started to worry about somewhat in late 2023 and certainly in 2024 is the scalability of this approach. There are two things that are growing with artificial intelligence today. The first is the model's parameters. The other, and this is what really made them concerned about this architectural approach, is the ever-increasing length of context that has become more and more part of the story about how they view the presentation of these models. This is what has driven them over the past two years to what they see as a more exciting bet, which is working more closely with memory suppliers as well as their logic manufacturing partners to find ways to gain extremely high-bandwidth access to larger-capacity memory.
Over the past two years, they have been involved in somewhat secretive projects to move away from SRAM memory, and research how to obtain, for example, much higher bandwidth for DRAM memory. This was a very exciting bet for them because it has now allowed them to assemble a platform that will begin production in the second half of next year, combining the scalability of these higher-capacity, lower-cost DRAM modules found in GPUs and TPUs with all the speed benefits you get from a Groq or Cerebras chip. They see the importance of this particularly if you look at where speed makes a real difference today. There is a version of this that resembles a faster, more responsive chatbot. But this is similar to when Henry Ford posed the rhetorical question of what people would say they wanted, and the answer was faster horses instead of a car. The fastest responsive chatbot is like the "fastest horse" in the world of rapid reasoning.
The ability to take a model with billions of transactions and run it smoothly at a speed of thousands of tokens per second represents a fundamental paradigm shift in artificial intelligence capabilities, by making these long-term proxies radically faster. This represents a kind of controversial discrepancy today between the characteristics of the fast inference chips available, which have very high bandwidth memory but very low capacity. If you look at the technical details of how it is actually employed and deployed today, you will find that it does not manage attention processes for long contexts. You still have to go back to the graphics processing unit (GPU) to do that. So there is this perplexing discrepancy where the area in which they can almost accelerate these things enormously is the same area in which they lack that technical capability.
In a sense, as with many technical challenges, it comes down to a fairly simple technical observation: they need chips that have this unique property—a very high memory bandwidth—to be able to load weights and cases thousands of times per second, while also having economical memory. When you look at running inference at the data center level for thousands of users, the economic viability ultimately boils down to the cost per gigabyte of memory used. This has been another key focus for them: unlocking this new core component, namely finding a path to obtaining extremely high bandwidth from the world's least expensive type of memory, dynamic random access memory (DRAM).
Walter has said a few things that do not align with prevailing conventional views. First, regarding their own bet on the direction of workloads, and second, the idea that you can, as the chip traditionalists believe, deliver chips in one generation per year or slightly faster than that. Major architectural changes usually come at a slower pace than this rate. But there is an ambition within Fractile that the pace of change in architecture and the technical bets they are making could be much faster than that.
What you want at any given time for an AI chip is to have a chip that accurately targets the workload you care about, and that is available to you today in large quantities. But if you get that chip, you know you'll want something else after 6 months. These workloads are developing very rapidly. Today they see a new model being released approximately every two weeks. Common sense comes from looking at those models and asking: Well, what do they all have in common? Fortunately, there are many commonalities. New large language models almost always urgently require much larger memory bandwidth to operate more quickly. The reality today is that these models tend to be self-regressive, with a very small processing push when generating texts, with that basic trade-off again between productivity efficiency, cost, and the speed at which you can deliver these models. There are certain things that these models agree on and demand, and the memory frequency bandwidth is one of them.
There are developments, especially if you look at the limitations of open-source Chinese models, the precise nature of the attention mechanism changes every few weeks with respect to what is considered to be the latest. The level of dispersion in the expert mixture models (MOEs) that it possesses, and the dispersion in the attention mechanism itself. They may also be facing even bigger shocks. There may be more fundamental changes. There is a duality about the boundaries of the physical world and the financial world. Then there are the demands of workloads that change almost at the pace of software, where there is a clear desire to be able to ship new platforms faster and more powerfully. There are also some fundamental obstacles to this.
At Fractile, they are very excited to adopt an AI approach in how they design chips. They have seen that being able to have the full scope of a problem from beginning to end really allows for a radical rethinking of how they do things. If you think about the classical laws of computer science, you'll find Amdahl's law. Everything that can be balanced becomes very fast, and the part that cannot be balanced remains the same. It's similar when you're in a long chain of organizations, where there isn't necessarily a pressing need or a huge payoff to radically change your processes as a chip design startup, even with AI being aggressively adopted to shrink timelines from 12 months of initial design work to a fraction of that. If there are then traditional bottlenecks and a normal pace in the next stage, that kind of bottleneck will always exist to some extent. They all face the same manufacturing cycle times as their foundry partners. From the moment the chip is sent to them until it is retrieved, it takes three to five months, even in ultra-fast production scenarios.
When that chip is recovered, it is essential to consider how to make these chips financially viable, which is to ensure they have a cost payback period. You need to have a depreciation window of 3 to 5 years for a chip to be a sensible financial decision. There are these two contradictory things that are believed at the same time. One of them is the enormous value in being able to take the entire chip design cycle, compress it, and shrink it. There is also something they cannot afford to lose, which is to place well-thought-out and precise architectural bets, because the chip that is manufactured must have a useful lifespan of up to 3 years or more. Being able to compress the chip design cycle over time essentially means you have greater opportunities for success. You want to always be prepared to pick a particular leading platform and say, "Okay, that's it. That is already the leading platform, and that is the one we will work to increase its production." But after that, it takes an increase in production of between 12 and 18 months. And you expect her to enjoy a long life. You expect it to continue to provide real value to your customers.
The key point here is that there is no belief in the idea that we'll eventually reach a point where we're shipping a completely new chip every few weeks just because we've shortened that time. Because this is a physical world. There are certain limitations on the power of data centers. There are delays in installing these things. You have to finance the actual underlying silicon, so there needs to be a payback period for the costs. But what you can get to is a world where the more you can shorten this delay, that is, the gap between observation and collectively achieving this goal, there is tremendous value to be gained. It's also an area where Fractile is positioned because the decisions made about what should be scaled up are the right decisions. You want to be, and perhaps you'll get more opportunities available as well.
One of the things that is very interesting about the increasing capabilities of artificial intelligence is that as productivity in doing things increases, there are clearly two responses in the economy as a whole. The first one is "Oh no, we won't have much to do. We may lose some jobs." The other is that we will be able to do more things. At Fractile they are now trying to build a single chip. They know that they are facing competitors who, if you open an Nvidia system you will find between six and nine custom chips, all made by Nvidia to work together to build something very powerful. So, there is already a gap there. It would be great if they were productive enough to start building their responses in the same way. The idea of being able to have a rolling front of bets at any given moment, which hopefully are deeply compatible and ready to start scaling up, is a huge advantage if you can build that machine and that engine. The six-month gap that it might consistently give you against your competitors is the kind of wedge that allows you to get in. They see this at Frontier Labs. So it is the equivalent of the frontier model in the field of chips. If you can find a way to capture an advantage that is three to six months structurally, then you will win all those deployments.
There's a meme circulating here of an executive saying, "Wow, this AI is awesome," and we're structuring the organization in a way that can consume it. Instead of one program, we will have six people and everyone will be clapping. There is some truth to that. Walter was speaking with one of the top CEOs of the Big Three semiconductor companies, and he did not allow making this prediction on air. But when asked how long will it take until we can convert the intent from the chief engineer to a fully usable GDSII file, he replied, "Ten years." That won't happen now.
The reasoning that has been successful over the past two years is always to question your logical assumption and then divide it by four with respect to timescales. So, there's a world where ten years seems like a very logical timeframe for that, but dividing it by four and maybe subtracting a little from it. There will be room for a kind of end-to-end prototyping in the next few years. One of the interesting things about chip design is that there are still loops in the middle that represent traditional solutions to NP-hard problems. Thus, if you really want to produce a GDSII file, today you are still working through a set of design tools that perform positioning and routing using a traditional algorithmic approach that works for days on end. This is really interesting to see, and to bring it back a bit to the kind of workloads that they're also excited about. Many of the world's difficult problems look like this, where there is a lot of thinking that can now be automated. There's a lot of mental work involved. But then there's also this kind of delay inherent in something you're doing. They see this in artificial intelligence to guide the development of AI models today. And so, the RSI worked. As we know, we can now have superior intelligence to guide our experiences, but we are now limited by experiences, and limited by computing time. Chip design is a bit like this. Perhaps if we look at the issue of engineers' intent to get to the GDSII files, it is similar to the RSI issue in that the things that will prevent us from getting there very quickly are the fact that there are things in that process that are computationally expensive in the traditional sense. What will generally be seen for all these types of problems is whether they will be replaced by alternative models. There are definitely these kinds of simulation models. Pushing approximations is all the models we see today for finite element analysis, thermal estimation, and so on. There is a great return on this particular type of work because of Amdahl's law, which is that we accelerate intelligence to the extent that we accelerate the period between these experiments and tests. It becomes the bottleneck. It is very important to speed up those experiments as well. They are looking internally at what is inside some of those algorithms, and whether there is a rudimentary approximation of how to do some of that. The chip design will not change in the
Chip design will not change in the final adoption phase for a long time. Companies like Cadence and Synopsys have been working with TSMC and other chip manufacturers for decades to build what is called the final mark that confirms a design is clean in terms of DRC and LVS. This compliance with factory rules represents enormous value. However, obtaining some kind of rough formatting algorithm in the meantime could help iterate faster, and this is exactly what people should be working on because otherwise it will remain the bottleneck.
The more intelligent the thinking that drives an experiment, and as the experiment becomes the bottleneck, the more one will think. This is Fractile's bet on workload, which is why they are excited about difficult problems and see Fractal as the chip that will drive acceleration of solving complex problems in general. It becomes morally necessary to think more deeply before launching each experiment. Baron Milledge may have said recently on the Dwarkish show that you can probably think for a hundred years before starting an experiment, and then think for another hundred years, equivalent to human thought, using these models about the results of that experiment, because the experiment itself is costly and takes a certain amount of time.
There are fundamental real-time bottlenecks for traditional algorithms in areas such as chip design. A lot of inference symbols will be generated before moving on to implementing a particular circuit. Betting on workloads and collaborating closely with design partners on architectural transformations of models, hoping to be one step ahead so that planning can actually occur.
The biggest thing being chased is speed by maximizing memory bandwidth to run models faster. One of the contradictions a chip company needs to address is building something that can handle a wide variety of upcoming workloads, under the concept of a hardware lottery where it will be trained on HBM memory-based GPUs. It's great that this chip is exceptional in handling today's advanced transformer models and operating them more quickly and efficiently. What's exciting about building chips with fundamentally new capabilities is exploring whether there are attractions that can be used to steer the new prototyping landscape toward certain characteristics.
For Fractal, the idea of achieving 25 times more bandwidth per chip compared to HBM-based chips allows exploration of concepts such as the laws of bandwidth expansion internally. Traditionally, expansion laws are thought of as laws specific to the number of arithmetic operations (flops), where better performance comes from pumping more calculations into these models during training and during testing. The current landscape of ideas shows that bandwidth expansion laws are already beginning to be noticed.
Mixture of Experts models represent one area where ideally models would become more and more scattered and dispersed. For the same level of intelligence, a lot of flops would be saved if moving from a 1 in 16 scattering level in a model to 1 in 128 or 1 in 256. However, one of the challenges is that rendering these models efficiently becomes very expensive for current HBM-based GPUs, often resulting in bandwidth bottlenecks and very low levels of MFU (Model FLOPS Utilization). There are areas where, even while building for today's world, ways can be looked for to unlock greater capabilities.
With today's large language models, it's very similar with regard to the mechanism of attention. There are forms of attention that consume less bandwidth, but they tend to be more consuming of flops to reach a certain level of intelligence. The excitement lies in upgrading memory bandwidth, which has not been expanded much in chips recently. Flops have been increased by a million times in the last twenty years, while memory bandwidth increased by only about 40 times in the same time frame. When this scope is expanded, more other resources can be conserved, reducing the number of flops used in these models to reach a certain level of intelligence, which is considered a multiplier for the global productivity of these models.
Five years from now, it is not clear what a major AI player or cloud service provider will be buying or what chips they will be consuming among companies like Nvidia and AMD, and the new class of accelerators, especially since at least four of these players have their own in-house efforts. Today, everyone who deploys technologies on a large scale tries to use as many different platforms as possible. One thing that will remain true is the real need for a diverse supply chain, especially since computing has become existential for these players.
There is a joke circulating today that the primary purpose of corporate self-help efforts is to reduce the price people pay to Nvidia, and there may be some truth to that because those efforts are quite similar architecturally. It would be a bet that does not rely on enabling a core capability that other chips cannot provide. This is a kind of maneuver where chips can be built internally or purchased from AMD to get a better price from Nvidia. It also plays a role in overall capability and control. When the shift occurs toward chips that actually deliver new capabilities, that's where a real need emerges for everyone who wants to put AI at the forefront to have some solution to run these models many times faster.
For Vanguard Laboratories, there is a window of superior capability and superior intelligence which represents the reason for their entire existence, otherwise everyone would be using chemist models all the time. With pressure from behind from open sources, it becomes even more important for people who want to publish at the forefront to have every aspect of what that means. Not only are the models better weighted, but also the fastest deployment allows for the most inference in the shortest possible time. This type of premium category in the field of chips, where speed is optimized above all else, is becoming a very important part of this world.
One of the standout things about the dynamic between the chip vendor and the Vanguard laboratory is that Fractal tries to be as compatible as possible with the needs of the Vanguard, looking a bit like some of those teams with people who really think deeply about these workloads. The issue that sometimes needs to be raised is explaining why the world will keep some outside chip players at the forefront for multiple decades, rather than integrating everything within these laboratories. This comes down to the kind of asymmetrical game all these labs and front-end players have to play, where they take enormous risks if they bet everything on a single piece of gear.
If Lab Number One has invested everything in special silicon chips, and then Lab Number Two discovers a new computing breakthrough that achieves much better computing efficiency for the same level of intelligence but only works on the chip they decided to use, Lab Number One may fail during the nine months needed to deploy enough of that chip to achieve that fivefold improvement in computing efficiency. There is a real need for these companies to deploy and use the same software platforms. They are playing different roles and now trying to compete in the modeling class. Betting on entirely different options and fully committing to those bets in the chip layer is an irrational and very dangerous act for anyone working at the forefront of this field.
Keep No Priors: AI, Machine Learning, Tech, & Startups in your library
Save the videos and channels worth coming back to, and find them again in one place.





