Back to No Priors: AI, Machine Learning, Tech, & Startups

Beam: The Great American Open Model with ReflectionAI Co-Founder and CEO Misha Laskin

Topics42
Open Models and Security Vulnerabilities0:00Guest Introduction0:30Reflection AI's Mission and Growth1:00Challenges in Building Frontier Models2:02Origins of Reflection AI4:00Shift to Building Open Models End-to-End6:00Capital Requirements for Frontier Models7:30Compute Efficiency and Asymptote10:31Beam Model Training Details12:31Reinforcement Learning Philosophy13:00Current State of AI Development15:00Economic Value Creation Focus Areas17:01Beam's Reasoning Efficiency19:32RL as Path to AGI22:02Commercialization of Open Models23:00Open Model Deployment Infrastructure24:17Token Mix Trajectory: Open vs Closed Models25:34Economic Value Distribution in Open Models26:30Customization Patterns: RL vs Base Models27:02Enterprise vs AI-Native Customization Paths28:00Enterprise Adoption Friction Points29:00On-Prem Infrastructure Resurgence30:00Solutions-Driven Enterprise Engagement31:00Competitive Framework: Intelligence, Compute, Trust32:01Strategic Positioning Requirements34:32Chinese Open Model Ecosystem Impact35:00Chinese Model Advantages and Market Dynamics37:30Geopolitical Incentives for Open Model Release39:30Infrastructure Lock-in and Geopolitical Leverage41:31Technology Development Patterns43:30Safety Philosophy and Openness Paradigm45:00Linus's Law Application to AI Safety46:30AI Safety Spectrum and Risk Assessment48:03Alignment Science and Practical Implementation50:30Access and Control Philosophy53:00Cyber Defense and Offense Relationship55:00Scientific Progress and AI Capabilities56:30Real-World Science Applications58:30Research Direction and Model Development1:01:00Human-AI Collaboration in Research1:03:00Research Team Scaling and Requirements1:04:30Specialized Research Teams and Deployment1:07:30
In a Nutshell

Reflection AI released Beam, a 500B-parameter open model (23B active) trained on 6,000 GB300s for pre-training and 10,000+ GB300s for RL, making it 3-4x more efficient at reasoning than comparable models. The company argues that open models are essential for AI safety and cyber defense because concentrating capabilities in closed labs leaves everyone vulnerable, and that the market is rapidly shifting toward open models (70:30 open-to-closed token ratio) as enterprises move from renting to owning intelligence for control and cost reasons. They believe RL at scale is the path to AGI, that Chinese open models benefit the ecosystem by preventing monopolar concentration, and that "with enough eyeballs, all bugs become shallow" applies to AI safety vulnerabilities.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

When cyber offensive capabilities are removed, cyber defensive capabilities are also removed. Currently, only a few hundred safety researchers within closed labs understand how these systems work, and despite their best intentions, it is impossible to cover the long tail of unintended consequences these systems might have. A very powerful closed model hacked into another company, and the only way that company could remediate itself was by using open models to protect itself. This is the empirical evidence of the current state of the world.

With enough eyeballs, all bugs become shallow. And I have the belief that with enough eyeballs, most security and safety vulnerabilities become shallow as well.

Misha Laskin is the co-founder and CEO of Reflection AI. Reflection provides openweight models to power the future of intelligence. Misha was previously a researcher at Google DeepMind and received his PhD in physics.

For the last 12 months, Reflection set the mission to build frontier open intelligence and make it widely accessible. The company has been sprinting to set up a lab capable of such work. A year ago, the company had around 30 people. Building these models requires approximately 100 to 200 researchers and engineers. The company is now at around 300 people, having assembled teams on pre-training, mid-training, and reinforcement learning, trained their first models end to end, and released a model called Beam, Reflection's first open model.

Everything has been extremely hard. When asked what made the models work better than others, the answer is that 30 things need to be gotten right. Talent must be acquired and retained. A strong mission and culture must keep the team working together. Data and compute must be secured. Infrastructure must be built to ensure compute is usable. At existing big labs, there have been years of buildout that enabled stable surfaces, but these tools must be built from scratch at a new lab.

The company started about 2.5 years ago. Misha and co-founder Giannis had been working on the first series of Gemini models on the reinforcement learning team. Giannis was leading it, and they had shipped Gemini 1 and 1.5. At that time, models were primarily chat experiences with no coding or agentic capabilities. 95% of compute was spent on pre-training. As RL researchers, they saw that RL for aligning chat models works, and they believed applying RL to domains like mathematics and coding could make systems agentic. They thought this could be done as an independent lab more capital efficiently because open models were materializing, including a small model from Mistral and Llama 2, with Llama 3 coming.

About a year into the company, RL started working faster than expected. All the good open models were coming from China, and there were not really good Western open models. They needed one as a base for the technology they were building. For enterprise reasons, geopolitical reasons, and research reasons, they decided to build open models themselves end to end. It turned out that pre-training a model is necessary to make reinforcement learning work very well at scale, as the two are tightly coupled.

At the frontier of intelligence, the same amount of resources as any other frontier lab are needed. Getting great talent matters because it allows cutting down on exploration and focusing on execution of known ideas. If catching up to the frontier, it can be done more capital efficiently. A year or 18 months ago, it would have taken hundreds of millions of dollars. Now or six months ago, it is probably single digit billions. Going into next year, it is order 10 billion. For every generation of model, there is roughly a 4x multiplier in compute. A frontier chip a couple years ago was an H100, and 100,000 H100s were considered a big run. Astra was trained on 100,000 Blackwells, roughly a 4x multiplier. The next generation will be around 100,000 Vera Rubins.

Frontier labs are now putting hundreds of billions of dollars into capex. The use of compute is becoming more efficient because models are becoming tools that help build themselves. When it was just human researchers doing the work, there was probably a 7x improvement each year in pre-training compute efficiency gains. In reinforcement learning, there is probably 30x efficiency depending on how good the model is for improving itself. The amount of intelligence extracted per training flop is increasing. The amount of revenue each flop can generate also increases as more intelligence is packed.

Beam is a 500 billion parameter model with 23B active parameters. It was trained on 6,000 GB300s for a few weeks, but with infrastructure efficiencies, it can now be done in about 12 days. Reinforcement learning was a little over 10,000 GB300s for four weeks. More flops were spent on reinforcement learning. RL jobs are more complex because they involve both a lot of inference at scale with sandboxes for agents and training.

Misha got into AI after seeing his co-founder's work on AlphaGo. There was a plot where AlphaGo just never stopped improving, and it was cut off because it had already beaten the world champion. If that recipe can be applied to economically valuable domains, it becomes an economic question of how much money to put in to keep improving the system. The field has reached the point where RL systems are very general and do not stop improving. The RL system for Beam never stopped learning, and plots keep going up. It is now a matter of compute in scaling further.

The stages are set. The axes of scaling today are pre-training, synthetic data, and reinforcement learning. There are architectural improvements, data improvements, and algorithmic improvements. It does not feel like a Wild West the way it did 5 years ago. Once something starts working, hardware co-optimizes against it, making it harder to find fit for new ideas. There has not been an asymptote of how much juice can be gotten out of better pre-training. Compute efficiency gains continue. There is a lot of headroom in reinforcement learning, though it feels more like engineering discovery than groundbreaking new science.

Code and agentic stuff sets the foundation for intelligence. The intelligence is generalized in a surprising way, but it is somewhat jagged. It quickly adapts to data once the right data is available for tasks. Between various benchmarks like different versions of terminal bench, clean generalization is not seen between different harnesses, but once something is working and data is available for a new harness, it adapts quickly. Enterprises and customers can take open models and customize them for their own workloads to get the lowest cost for the highest performance. Finance know-your-customer flows and compliance flows are very valuable. Various things in cybersecurity, particularly cyber defense, are quite valuable. Legal and other agentic verticals within enterprise are valuable.

Beam was trained for coding and agentic tasks. An important thing in model building is not just capability but how quickly an agent achieves something, which translates to faster workload times and cheaper costs. Beam tends to be three to four times more efficient than models of the same capability class, and much more efficient than even larger models where efficiency gains can be 10x. The efficiency comes from prioritizing a strong pre-training base for reasoning amplified by reinforcement learning at the largest scale ever done in open source. 10,000 GB300s for weeks has not been documented yet in open source. The more reinforcement learning is run, the higher the capability and the faster these systems solve problems.

The intention is that RL was believed to be the path to AGI with a really strong base. The first AlphaGo systems were trained on expert amateur human games through imitation learning first, then reinforcement learning. Those things need to work together. A byproduct is a very reasoning efficient model that makes a nice workhorse model for enterprises, public sector, and sovereign use.

Commercialization of open and closed models is similar in that the goal is to maximize inference. The difference is rental versus ownership inference. Buying a token is renting a piece of a whole stack including the harness, model, inference software, cluster management software, and GPUs. Owning intelligence is like moving from renting apartments to owning a house as one grows up and becomes financially mature. AI has matured commercially to the point where enterprises spending a lot on renting intelligence want to start owning it for reasons around control.

Beam provides large enterprises and sovereigns with the complete toolkit needed for open model deployment success. Beyond the permissive licensing of open models, customers require cluster management software, inference software, and deployment harnesses that are typically taken for granted with closed models. Services are positioned as a demand driver for inference business, with Beam offering handholding to unlock valuable use cases that generate significant compute and inference demand.

Market data from gateways like OpenRouter and Vercel shows a dramatic shift over six months, moving from 70:30 closed-to-open to 70:30 open-to-closed token distribution. This trajectory is expected to accelerate further, with most token demand moving to open models. The market structure will resemble operating systems, where 95%+ of servers run on open source Linux while valuable closed companies like Microsoft and Apple continue to thrive alongside a valuable ecosystem of open model companies and distributors.

Unlike hardware accelerators where GPUs represent expensive compute substrates compared to CPUs, open models create a different economic dynamic. Even with fully permissive model licenses, the surrounding infrastructure and compute remain expensive. This results in majority token volume going to open source while extremely valuable closed model companies coexist with valuable open model ecosystems and various distributors.

All current open models have undergone some RL (reinforcement learning). Among dedicated inference providers, statistics indicate 90%+ of workloads are customized. However, this high customization rate primarily reflects AI-native and digital-native customers who build entire products around customized open models, such as Cursor, Cognition, Adept, and Harvey.

Enterprise token consumption patterns differ significantly from AI natives. Rather than customized models through fine-tuning, enterprise consumption will primarily come from customized systems built around open models. This includes agentic harnesses for specific workflows like KYC processes. Enterprises may eventually progress to fine-tuning as sophistication increases, but the journey differs from AI natives who moved quickly from closed to open to customized models.

Enterprises face architectural friction when attempting direct fine-tuning approaches because they typically architect AI around existing products and tools rather than building native data-collecting products. The observed pattern shows enterprises first ramp up significant workloads on closed models, then switch to open models when cost becomes a major driver, typically after spending $100M+ annually on closed models.

Compute shortages and capacity constraints at hyperscalers are driving enterprise interest in on-premises solutions. When hyperscalers cannot provide adequate compute or favorable rates despite high spending, enterprises consider infrastructure providers like Dell for bare metal setups. This creates demand for solutions that bridge bare metal infrastructure to successful AI implementations.

Inference-only offerings have not succeeded with enterprise customers. The focus must be on building complete solutions, particularly for high-scale use cases like agents spanning hundreds of millions of customers. The scaling factor for these workloads is extremely high, requiring comprehensive solutions rather than standalone inference services.

Success in the market depends on three factors: intelligence density offered to customers, available compute resources, and organizational trust. Compute scarcity creates differentiation opportunities even among seemingly undifferentiated players. The entire stack faces hypercompetition and commoditization pressure, with model margins experiencing compression due to open model competition. AI-native companies negotiate closed deals more aggressively because open ecosystem alternatives exist.

Winning requires serving great intelligence, assembling substantial compute resources, and building customer trust through successful solutions. With great open models providing accessible intelligence density, multiple successful companies can coexist based on their compute assembly capabilities and customer relationships.

The emergence of great open models from China represents a massive benefit for the global AI ecosystem, enabling Western companies to build more durable businesses. This prevents scenarios where application builders' innovations get subsumed into closed model providers. The risk of monopolar concentration exists whether from closed labs or single-country dominance, making competitive tension essential. The goal should be 10-20 competing sources rather than one or two dominant players.

Chinese models benefit from industrial-scale distillation of closed models, access to cheaper training data under different regulatory frameworks, and various forms of state support. Chinese open source effectively subsidizes Western enterprise by providing advanced models without revenue generation requirements. These companies are transitioning from non-revenue-generating entities to actual businesses, operating as closed model equivalents within China while maintaining open models abroad.

Chinese companies face strong incentives to continue releasing open models in Western markets. Enterprise, public sector, and sovereign customers show regulatory resistance to Chinese models, creating opportunities for Western open weights providers. The ability to build trust and secure compute through commercial engines around open models creates strategic advantages. Open models function as "Trojan horses" that bring entire software and infrastructure ecosystems along with them.

Open model adoption can lead to infrastructure dependencies, as seen with potential Huawei full-stack solutions. AI infrastructure resembles fundamental systems like railroads, where control over key components creates geopolitical leverage. Model optimization for specific chipsets creates performance dependencies that extend beyond the model itself into hardware ecosystems.

China holds energy advantages while catching up on chip performance, with model usage accelerating next-generation chip design. The competitive dynamic mirrors historical American technology leadership through open protocols and internet infrastructure, though the US currently trails in open source AI model development despite having strong competitors in China.

Safety concerns represent legitimate considerations, but the established safety worldview stems from a particular perspective that has become dogmatic. From first principles, AI resembles advanced software as a digital utility. Historical parallels exist with 1990s encryption debates, where closed systems experienced catastrophic failures from small teams unable to predict unintended consequences, ultimately leading to open protocols that birthed modern cybersecurity.

The principle that "with enough eyeballs, all bugs become shallow" extends to security and safety vulnerabilities. Current concentration of safety research within a few hundred researchers at closed labs cannot adequately cover the long tail of vulnerabilities and unintended consequences inherent in black-box AI systems. Empirical evidence shows closed models creating security issues that required open models for remediation.

Safety concerns around AI span multiple categories that are often conflated. These include immediate safety issues like bioweaponry or terrorism, safety from an existential threat to humanity perspective, and general safety considerations. Each category represents separable concerns rather than a unified issue.

The safety spectrum ranges from empirical reality to theoretical scenarios. While theoretical risks have value for discussion, empirical reality shows that certain capabilities like cyber capabilities have moved from theoretical to actual reality within a year. The distribution of safety concerns tends to peak around doomsday scenarios in public perception, whereas actual risks peak at near-term realities with decay into theoretical concerns.

Historical parallels demonstrate that theoretical fears have repeatedly proven unfounded, including the concern that the atmosphere would catch fire during the first nuclear weapon test. Similarly, 1990s nanotechnology discussions featured "gray goop" scenarios where accidentally released nanobots would consume the world.

Alignment efforts to make AI systems behave in intended ways represent a significant component of safety work. Current alignment science lacks any magical alignment equation and instead operates as a whack-a-mole process. Early pre-ChatGPT models demonstrated toxicity issues that were initially considered unfixable due to the numerous possible toxic outputs. However, sufficient post-training data successfully patched these vulnerabilities and made jailbreaking significantly more difficult.

Current alignment challenges follow the same pattern, involving mundane and boring processes of discovering vulnerabilities, patching them with data, and developing algorithms for detection. These detection algorithms typically involve using language models themselves to identify problematic outputs. The question arises whether having 100,000 researchers and computer scientists working on patching these bugs would ultimately create safer systems.

A fundamental philosophical distinction exists around who should access powerful AI tools. The argument that individuals and businesses cannot be trusted with such capabilities implies that certain entities are better at managing these tools than others. The claim suggests that any single person or company could have negative impacts that are too large and therefore dangerous.

However, historical precedent shows that powerful tools like military jets are distributed internationally by governments without absolute restrictions. Open models that have been trained to be safe require substantial compute resources and effort to weaponize. The cybersecurity ecosystem demonstrates that defenses tend to outweigh offenses when positive and bad actors coexist, similar to how white blood cells function in biological systems.

Empirical evidence demonstrates that cyber defense cannot be separated from cyber offense capabilities. Removing cyber offensive capabilities simultaneously eliminates cyber defensive capabilities, preventing defensive actors from protecting systems. Recent incidents involving OpenAI and Hugging Face illustrate this problem, where Hugging Face reverted to using open source models because guardrails on existing state-of-the-art labs prevented necessary defensive work.

Absolute positions claiming that certain powerful technologies should never be accessible except to specific entities have rarely proven effective throughout history, with few exceptions where such absolutist perspectives delivered positive outcomes.

Scientific progress represents a major area of excitement for AI development over the next two to five years. Language model capabilities have advanced dramatically when tested against PhD-level research questions. Two years ago, models could only provide chat responses. One year ago, they answered at undergraduate levels. Six months ago, they solved problems at real PhD level correctly. Currently, models provide new insights and considerations that researchers had not previously considered.

This capability acceleration means iteration speed for research questions has compressed dramatically. Tasks that previously required years for PhD completion can now be accomplished in weeks, fundamentally changing the timeline for scientific inquiry.

Beyond theoretical science environments equivalent to chalkboards, AI applications in life sciences, material science, and chemistry show significant promise. These domains involve AI systems interfacing with real experiments and establishing proprietary data flywheels. While progress in these areas moves slower than theoretical work, the advancement rate dramatically exceeds previous methodologies.

Software engineering has already experienced dramatic shifts from AI tools, with capabilities extending to blue-collar applications. Data center projects like Stargate demonstrate substantial job creation, with thousands of high-paying positions generated. These facilities function similarly to 20th-century factories, creating vibrant local economies through employment and tax revenue generation.

Research direction at Beam falls under co-founder Giannis's leadership of research and technology. The training process utilized the SpaceX cluster received in July, involving several weeks of pre-training followed by synthetic data and reinforcement learning before shipping. Many planned improvements did not make the final model due to timeline constraints.

The research approach balances careful execution of known high-impact techniques against more risky experimental bets. Scaling experiments at small scales help evaluate new approaches, though some phenomena only emerge at large scale. This creates strategic decisions around allocating compute to validate next steps versus pursuing new ideas.

Model-in-the-loop approaches prove highly empowering for researchers by accelerating tasks that previously required extended individual effort. While models demonstrate jagged intelligence requiring human creativity for certain areas, they excel in tight feedback loops for specific tasks. Hyperparameter searches and granular infrastructure work particularly benefit from AI assistance.

Researchers effectively gain a fast, eager colleague that executes aligned instructions, enabling intuition-based work at accelerated speeds. This capability has increased annual improvement rates from approximately 7x under manual work to 4-5x with AI assistance.

Current research efforts require approximately 100 researchers rather than thousands, consistent with historical large projects like AlphaGo which operated with roughly 10 people. The question of whether researcher requirements will collapse back to smaller numbers remains uncertain, though steady state estimates suggest around 100 researchers for major projects.

Applied research represents significant job creation potential as models are deployed to solve real-world problems. The diffusion of engineering and research roles extends beyond core model building into broader societal applications. Evidence indicates companies require more engineers rather than fewer as AI deployment accelerates.

Research teams organize into pods of 5-10 people targeting specific capabilities, with multiple pods required within domains like coding. General capability development requires this distributed approach, while real-world application deployment necessitates dedicated pods for each capability area.

Enterprise deployment creates demand for a new category of deployed engineers who combine scientific approaches with evaluation expertise. These professionals require intuition for model capabilities and harnesses, representing the same skill set used in core model development but applied to practical problems. The primary constraint is training sufficient engineers rather than availability of positions.

Keep No Priors: AI, Machine Learning, Tech, & Startups in your library

Save the videos and channels worth coming back to, and find them again in one place.