Beam: The Great American Open Model with ReflectionAI Co-Founder and CEO Misha Laskin
In a Nutshell
Reflection AI launched Beam, a 500B-parameter open-weight model optimized for programming and agent tasks that delivers 3-4x inference efficiency gains over comparable models through heavy reinforcement learning on 10,000+ GB300 units. Misha Laskin argues open models are essential for security because closed systems hide vulnerabilities while open ones enable collective defense, and Western open models must compete with China's growing ecosystem to prevent infrastructure lock-in. The path to AGI runs through continued reinforcement learning scaling, with the company betting that models will keep improving indefinitely as more compute is applied to agentic tasks.
These notes were generated by AI and may contain inaccuracies.
When offensive cyber capabilities are removed, defensive cyber capabilities are also removed. Currently, only a few hundred safety researchers inside closed laboratories understand how these systems work. Despite good intentions, covering up the long series of unintended consequences these systems may have is impossible. A very strong closed model had compromised another company, and the only way for that company to fix itself was to use open models to protect itself. This is empirical evidence of the world we live in.
With enough observation, all errors become superficial.
With enough monitoring, most security and safety gaps also become superficial.
Misha Laskin is co-founder and CEO of Reflection AI, which provides open-weight models for powering the future of intelligence. Misha was previously a researcher at Google DeepMind and holds a PhD in physics.
Over the past twelve months, Reflection AI set a mission to build sophisticated open intelligence and make it widely available. The company was assembling an airplane while flying it. About a year ago, Reflection had approximately 30 people. Building these models requires hundreds of people or more—specifically a few hundred researchers and engineers. The company has grown to about 300 people and has assembled all the teams involved in initial training, intermediate training, reinforcement learning, and expanded the training of their first models from start to finish.
Reflection just launched a model called Beam, which is the first open model from the company.
The most challenging aspects include computing resources, talent acquisition, and data size. Everything is extremely difficult. Building a model requires succeeding at 30 things simultaneously: finding talent, keeping it, maintaining a strong mission and culture, obtaining data, acquiring computing power, and building infrastructure to ensure computing power is usable. Many tools taken for granted in large established laboratories must be built from scratch.
Reflection AI started about two and a half years ago. Misha and co-founder Yannis were working on the first series of Gemini models at DeepMind, specifically on the reinforcement learning team. They had just launched the Gemini 1 and 1.5 models. At that time, models were primarily chat-focused, with 95% of computing power going to pre-training.
Their bet was that reinforcement learning could be applied to fields like mathematics and programming to make systems more agentic. They believed they could do this as an independent laboratory with much greater capital efficiency because open-source models were emerging, including small Mistral models and Llama 2, with Llama 3 coming.
About a year after the company was founded, all good open-source models were coming from China. There were no really good Western open models. Several reasons drove the decision to build open models entirely themselves, combining a bet on reinforcement learning with end-to-end model building. Institutional and geopolitical reasons existed, but there was also a research reason: pre-training the model is necessary for reinforcement learning to work very well at scale.
To build something of value, significant resources are needed. Once at the forefront of intelligence, the same amount of resources as any other leading laboratory is required. Research involves exploring new ideas versus implementing known ones. Attracting great talent reduces scouting and enables focusing on things that work, making catching up at the forefront much more capital efficient.
A year ago or 18 months ago, hundreds of millions of dollars would have been required. Now, or even 6 months ago, it was in the billions—a few billion dollars. Looking ahead to next year, each new generation of models requires approximately 4 times the computing power increase.
When talking about chips currently considered flagship, the H100 was two years ago. It was probably about 100,000 units of H100. The Astra model was trained on 100,000 units of Blackwell chips, representing approximately four times the power of the H100. The next generation will rely on approximately 100,000 Vera Rubin units. This gives an idea of scale, going from hundreds of millions to billions and then to tens of billions.
Leading laboratories are now pumping hundreds of billions of dollars into this field. Whether they will be pumping trillions of dollars in the next year or two is unlikely. However, the use of computing has become more efficient because models themselves have become tools that help build themselves.
When work was limited to human researchers only, there was a 7-fold improvement every year. Initial training systems are now more well established, but there is still a lot of room for improvement in reinforcement learning. Efficiency gains of 30 times may have been reached, depending on how well the model actually improves itself.
The speed is probably four times faster or more than if researchers had done it on their own, and that pace is likely to accelerate. The amount of intelligence that can be extracted for each training process (flop) is increasing.
The Beam model has 500 billion parameters and 23 billion active parameters. Training required 6,000 GB300 units. It was run for a few weeks, but with infrastructure efficiencies, it can now be done in about 12 days or possibly less. This is a combination of scientific expertise and infrastructure expertise.
Reinforcement learning consumed slightly more than 10,000 GB300 units over 4 weeks. More training efforts were spent on reinforcement learning than initial training. Reinforcement learning tasks are more complex because they involve doing a lot of inference on a large scale with all kinds of experimental agent and training environments.
The real reason for entering the field of artificial intelligence was seeing the work of the co-founder of AlphaGo. The moment in one of AlphaGo's graphs where it never stopped improving was compelling. They stopped it at a certain point because what was the point of improving it further when it had already defeated the world champion. If that recipe is applied to things of economic value, it becomes an economic issue concerning how much money to invest to continue improving the system.
The technical report will describe how to build these things, and this reinforcement learning system never stops learning. The charts show that it continues to rise. It is basically just a matter of computing power in terms of expanding its scope further. The field has reached a stage where very general reinforcement learning systems never stop improving.
The field has moved on—it is still a science, but it looks more like an engineering discipline. Rocket building is a science, but most people think of it as an engineering discipline that involves some scientific work. The phases are somewhat defined in the current axes of expansion: pretraining, artificial data, and reinforcement learning. There are improvements in the data and improvements in the algorithms, but it does not look like the Wild West it did five years ago.
There is no limit to the benefits that could be achieved through better initial training. There is no slowdown in computing efficiency gains. There is a lot of room for improvement in reinforcement learning, though it seems more like engineering discovery than revolutionary scientific understanding.
Programming and proxy operations lay the foundation for intelligence. Intelligence is generalized in a surprising way. The model reached a level of capability that is relatively new and modern, discovered by other laboratories not too long ago. There is a kind of generalization happening. At the same time, it is somewhat zigzagging, adapting quickly to data once the right data for the task is available.
This is evident even between different benchmarks, such as different versions of Terminal Bench. There is no clear generalization between the different measuring instruments, but once something is working and there is general proxy capability, with data for a new measurement tool, it starts to adapt to it very quickly.
The question becomes where the datasets with economic value are. One of the things that makes open models so powerful is that companies and customers can take them and customize them to suit their needs, getting something ideal for their workloads. Lower cost for higher performance. It is an empirical matter—going to customers, seeing what is valuable to them, and finding out if it is possible.
Typically, when something has economic value, data can be generated because training is not just on customer data. Evaluation processes are prepared. If a good assessment of the tasks can be prepared, then synthetic data that approximates that can be generated, thus getting good generalization.
In the field of finance, various Know Your Customer flows and compliance flows are of very high value. There are many different things in cybersecurity, and cyber defense in particular, that are of great value. In the legal field, there are all kinds of agent-oriented sectors within companies that are considered valuable and have a very similar pattern in terms of how the work is done.
The Beam model is specifically trained for programming and agent-oriented tasks. What is important in model building is not just the ability, but how quickly the agent accomplishes the task. This translates into faster completion times for workloads and cheaper costs for customers.
Beam tends to be three to four times more efficient than models of the same power class, and much more efficient even compared to larger models currently in existence. Efficiency gains sometimes reach approximately 10 times. The reason it is so effective is that a strong pre-training base for inference was prioritized, then reinforced with reinforcement learning on a scale believed to be the largest ever in open source.
When setting up reinforcement learning, the goal is to extract the maximum amount of capabilities in the shortest possible time. This is the same thing that happened in previous systems such as AlphaGo. The early AlphaGo agents were rather clumsy in their problem-solving methods, but by the time they reached Lee Sedol's level, they had become very clever in their approach to research. The more reinforcement learning is implemented, the greater the ability and speed with which these systems can solve problems.
The path to artificial general intelligence was believed to be through reinforcement learning with a very strong foundation. The first AlphaGo systems were trained on human matches between experts and amateurs. They learned imitation first and then reinforcement learning. These things need to be done together. The result is something of great economic value, highly efficient in reasoning, and a strong and useful model for institutions and the sovereign public sector.
The marketing of open and closed models is somewhat similar. The goal is to maximize the inference. The difference is that it is more like renting in exchange for having the ability to reason. When buying a token, a part of the integrated infrastructure is being rented—whether it's an agent or an operating system, the model, the inference software, the cluster management software, and the GPUs it runs on are all built into the cost of the token.
When wanting to own intelligence for many reasons, it is like starting out by renting apartments, but when financially mature, wanting to own a house for different reasons. Artificial intelligence has matured commercially in the market to the point that companies spend a lot of money to rent its intelligence, and then begin to want to own it for various control-related reasons.
An open model is available, but for the customer to benefit from it, all the other things surrounding it are needed. Just as things are taken for granted in a closed model, a group management program, an inference program, utilities, etc. are needed. Large corporations and sovereign entities are provided with all the tools needed to successfully deploy open models.
Services are viewed as a way to enter and unlock extremely valuable use cases that drive a lot of demand for computing—demand for inference work.
Six months ago, the majority of models were closed and the minority were open when using any gateway like OpenRouter. This has completely reversed from 70/30 closed/open to 70/30 open/closed now. This trend will only accelerate. The world will not look much different from operating systems, where more than 95% of the world's servers and computers run on an open-source operating system like Linux.
This does not mean that closed models are not of great value. Microsoft and Apple are companies of enormous value. The market for this type of thing is really big. The majority of demand for tokens will move towards open source. There is a big difference here: even if the model is fully available or perhaps subject to some kind of license, everything else needed to run it is still expensive. Computers are still expensive, and hardware accelerators like graphics processing units (GPUs) are a more expensive computing architecture than central processing units (CPUs). A different type of cloud needs to be built around it.
The expectation is that the majority of the code will go to open source. There will be closed model companies of extremely high value. There will also be a valuable ecosystem of open-source companies and various distributors around them.
Every open-source model currently available has undergone some degree of reinforcement learning. From statistics from some providers of custom inference for open models, the vast majority, over 90%, are custom models. The biggest customers for customized inference today are AI startups and digital companies that have a complete product that is a large customization of an open model. Companies like Cursor, Cognition, or Harvey have products built around something big and dedicated.
The majority of enterprise token consumption will not come from custom models, but from custom systems. An organization can take an open model and not fine-tune it yet, but simply customize a system around it as a business agent to make it work in know your customer processes. Once sufficiently sophisticated, they might begin to fine-tune.
In institutions, the journey that the original companies undertook in artificial intelligence was: starting with a closed model, then moving to an open model, and then customizing it. They got through that very quickly. Nations created original products that function as data collectors, while organizations mostly structure this around their existing products and tools.
Organizations begin to think about open models when they have lifted large workloads using closed models. There has not been a direct and rapid move towards the ownership market. Organizations have to go through the leasing phase to become a large enough cost engine, then switch after that when spending becomes significant, such as more than $100 million a year on closed models.
There is an interesting resurgence of local systems, partly due to a shortage of computing resources. Organizations consuming services from a preferred cloud provider often find there are not enough computing resources and it is difficult to get good prices. Organizations spending a lot of money may turn to an infrastructure provider like Dell to set up basic infrastructure and deliver services within the organization.
The role is enabling organizations to build successful solutions. What enabled AI-based companies to adopt open models was building their own solutions. However, institutions need more and more help in this area. It is about going in and assessing what are the most valuable things that can be done. Many of these organizations may have built a smart agent that serves hundreds of millions of customers using closed models and really want to scale up. The expansion factor is very large.
The closed-system model is a competitive system, and it increasingly appears that the open-source modeling system will be the same. Large markets tend to be competitive.
The equation for winning in this field is: how much revenue is generated, which depends on the intelligence density offered multiplied by the amount of computing power available, multiplied by the amount of trust gained with organizations. Computing is rare in this market. Previous competitive situations might suggest that a few players are not differentiated, but if you are really good and have computing abilities, that is sufficient.
The entire field is competitive across its entire structure. Everything across the entire structure is so highly competitive that it becomes a commodity, and the profit margins of the models will shrink. There is tension surrounding open models, and with closed models this leads to a shrinkage in margins. The margins of software applications are shrinking, as is the inference layer, and then the underlying infrastructure. Many native AI companies are negotiating their closed deals in a very aggressive way because the open ecosystem exists.
Any company is evaluated based on the intensity of intelligence offered to customers, the extent of their computing capabilities, and how much confidence their clients have in them due to their successful solutions. As long as there is such a high concentration of intelligence available to everyone, there should be many successful companies. If you are a brilliant AI developer, you can achieve what happened with closed companies where revenues increase dramatically, and you can amass greater computing power than anyone else.
Two years ago, most open-source or open-weight models were of Western origin, such as Mistral developed in Europe, with Meta being a mix between Europe and the United States. Over the past year or two, there has been the rise of Chinese models with open weights. In some cases, there is a view that they were distilling a lot of advanced models, which allowed them to make very rapid progress. Large closed laboratories are now trying to introduce tools to prevent these distillation processes from occurring as much as possible.
The fact that a wonderful open-source ecosystem has emerged from China is actually a tremendous benefit to the world. Many companies in the West have been able to build more sustainable businesses as a result. This is viewed positively as a gift, because a world without open models at all is imaginable. There is a kind of long poppy syndrome that occurs when an application developer builds something great, then it gets absorbed by a closed-model provider.
A unipolar world should be avoided, whether about concentrating all resources and computing in closed labs, or having only one or two labs. A competitive ecosystem with 10 or 20 closed labs is acceptable, but having only one or two labs is concerning. Similarly, the source of intelligence that everyone builds on should not come from only one country. Ideally, there would be 10 or 20, but the capital expenditures are too high. Having at least two would be great.
It is extremely important to have an ecosystem or power as an ecosystem that rivals China in the West. It is not really about being a conflict between the West and China. If there was another country besides China that produced all those amazing open-source models, some kind of competition would still be desired. Competition is simply a good thing.
Chinese models will remain great. They have advantages and disadvantages. The advantage is that they are definitely still developing closed models on an industrial scale. They have access to much cheaper or free data, because training on copyrighted PDF files in China is acceptable under a different arrangement. There is a perception that these companies have government support, either directly or indirectly, as the companies have not been generating any revenue for a long time.
Chinese open source is essentially Chinese government support for American or Western institutions. This is continuing, and now these companies are becoming businesses. Although the prototypes are open abroad, inside China they are equivalent to closed prototype labs. There are real businesses being built there, and they will continue to build great models.
The incentive for Chinese companies to keep models open in the West relates to the huge demand from corporations, the public sector, and sovereign entities for open Western models for many reasons. Most Fortune 500 companies currently have a very low footprint with regard to open models due to resistance or aversion to Chinese models for many reasons, some logical and some illogical. Building these things is important. If an economic engine can be maintained that builds trust, then organizations will want to work with the model developer. This allows securing computing resources.
China represents a kind of restricted market where open models can be built and still make a lot of money in that market. Interestingly, many great companies in the West do not build open models but offer their services to other companies. This dynamic does not occur in China to the same extent. Open-source providers are considered the intelligence providers in that country. Continuing to release great open models is a huge geopolitical advantage for China for many reasons. One reason is the desire as a country for other countries to build their projects on your technologies.
There have been previous escalations or tensions between the United States and China, such as those that occurred with 5G networks, fiber optics, and the Belt and Road Initiative. When offering something cheap to another country so that it becomes tied to your infrastructure, there are commercial and geopolitical advantages that are very similar in importance to rare earth minerals. If you have a component that others will use for various purposes, then you have geopolitical leverage over it.
Open models are a Trojan horse for the infrastructure they bring with them. You have an open model, but it alone is not very useful. The entire software and infrastructure system of a particular country must be adopted. It might be acceptable today if a U.S. ally were to run a Chinese model on American chips. But soon, Chinese companies like Huawei will come to offer a complete solution. You will become restricted to relying on that country's supplies.
If artificial intelligence or chips and data centers are looked at as railways, they are essentially basic infrastructure. Everyone depends on it, but no other country can actually build its own railways. They should all turn to you, or seek refuge in another country. Then you possess a great deal of geopolitical influence. It may happen that a particular model is deduced with higher efficiency on a specific set of chips supplied by Huawei or others. These systems are fully optimized end-to-end in real time, because the dominant model has already been defined.
While China is at a disadvantage today because its chips are not performing as well, it has an advantage in power and is very resourceful. They are catching up on the chip side as well. This will only accelerate with models, in terms of using models to design the next generation of chips. There is also the aspect related to geopolitical competition with the United States. The idea of American companies and startups relying on foreign technology, and the rest of the world building on your innovations, is the American approach for a long time with technologies such as the Internet, open protocols, and open source.
The competitors in China are really good, remarkably so. Thus, an ecosystem is just beginning to take shape in the United States, but there is still some catching up to be done.
The biggest criticism from closed laboratories of open models tends to be about safety and controllability. This is a legitimate concern, and the security outlook that has been established has been founded on a single perspective and a particular viewpoint that has become somewhat dogmatic. If artificial intelligence is looked at from first principles before thinking about whether it is closed or open, this technology is being built with some kind of foresight about its capabilities.
Artificial intelligence can be thought of as a more advanced version of software. There were discussions about the dangers of software, especially in the early 1990s, about strong encryption protocols. It was actually closed at the time, and there were discussions between the National Security Agency and civil liberties advocates about whether it should be closed or opened. Ultimately, the decision came after some disastrous failures on the closed side, where a few engineers designed certain systems with unintended consequences that they could not foresee and which were easily hacked. Strong encryption protocols became open, and this actually led to the birth of the entire cybersecurity field.
The default state of the world is in fact that openness is security. Models have many similarities to software, but they are compounded, as the long list of flaws is even longer. It is difficult to understand in software because it is a black box, and the list of vulnerabilities is long, and therefore the unintended consequences are even greater. Linus's Law, as established by the creator of Linux, is that with enough eyes, all mistakes become simple. With enough eyes on the ground, most security and safety gaps also become simple.
The state of the world today is that there are a few hundred safety researchers inside closed laboratories who understand how these things work. Despite their best intentions, it is impossible to cover up the long list of loopholes or unintended consequences that may result from these systems. When looking at the major cybersecurity issues that have emerged, a very strong closed model has had unintended consequences when it has compromised another company. The only way that company could heal itself was to use open models to protect itself.
When people talk about safety, they really confuse the concepts. Safety means three or four different things to different people. There is safety in terms of cyberattacks or the use of artificial intelligence for hacking. There is safety in relation to biological weapons or terrorism. There is a kind of safety from the perspective of the existential threat to humanity. People talk about these things as if they are one thing, when in fact each one of them is somewhat separate.
On the safety spectrum, there is essentially a spectrum of reality, from empirical reality to theoretical scenarios. There are aspects of artificial intelligence that seem like science fiction, where what was entirely theoretical last year becomes a reality, such as the cyber capabilities of these systems. There are real, empirical things happening that can be predicted with reasonable confidence over the next six months. Then there are what might be called extreme theoretical matters.
Earlier versions of theoretical fears would have been the idea that the atmosphere might ignite when the first nuclear weapon was detonated. There was much debate among a small group of researchers in the physics community about that. When people talked about nanotechnology, they talked about grey jelly and how you might accidentally launch a nanobot and it would suddenly eat the whole world. This was discussed in the 1990s and in laboratories while people were working on microfluidics.
The challenge with artificial intelligence is that it occupies minds, it is present in everyday culture and it is very easy to humanize it. From the perspective of general perception there is an irregular distribution around the meaning of security, where attention is focused on doomsday scenarios, while the reality is that the focus should be on real matters and then gradually move towards the theoretical. It does not help when company leaders say that there is a 10% chance that we will all theoretically die. This is not supported by anything.
As a scientist, distributional thinking centered around reality is preferred over made-up numbers. For the next six months, there are concerns about misalignments or unintended consequences as systems grow in power, but they should be viewed as engineering tools rather than having consciousness or feelings, which are difficult to define precisely.
Theoretical doomsday scenarios make good dinner party topics, but the real issue is compatibility—how to adapt systems to work in intended ways. The science of alignment has been extremely boring and unsatisfying with no magic formula; it's a game of hit-and-run. Early pre-ChatGPT models were launched pre-trained and turned out toxic, leading to the belief they couldn't be repaired, but this proved untrue.
With enough post-training data, toxic models can be corrected, and this generalizes in ways that make it difficult to break their protection. Alignment issues will most likely be routine, boring matters where gaps are discovered, data is created to correct them, and algorithms (usually language models) detect issues.
There's a distinction between the theoretical, philosophical aspect of safety alignment and the routine reality of fixing software bugs. The question arises whether one hundred thousand researchers and computer scientists should monitor and correct errors for safety.
There's an important philosophical distinction: even if tools are very intelligent and the ecosystem can manage unintended consequences, some argue individuals and companies shouldn't have such powerful tools. This comes from the view that "we have access to powerful tools but others do not, simply because we are better than them" at managing them.
To be fair to this claim, the negative impact one individual or company could have is very large and frightening, but this already exists. It's similar to weapons scale, like the US government selling fighter jets to other countries.
It is not easy to take a model, especially one that is open and trained to be safe, and allocate resources to make it too dangerous. Going back to cybersecurity, clever hackers can cause widespread chaos, but defenses tend to outweigh attacks. The internet is ultimately a system similar to white blood cells, and enabling more people to defend helps protect against offensive capabilities.
Empirical reality shows it is very difficult to separate cyber defense from attack. Removing offensive cyber capabilities also removes defensive capabilities, leaving parties who wish to help with defense unable to do so. This was seen in recent OpenAI and Hugging Face incidents where they reverted to using open source because they couldn't use the most advanced existing laboratories due to protection restrictions placed on the models.
Arguments reach a point where they become very absolute, saying these things are so powerful that no one should have access to them except for certain entities. The reality is that this never happens; it is difficult to find truly rare exceptions where such absolute perspectives have succeeded, and usually it's quite the opposite.
Born in the last year of Soviet Russia in Leningrad (now Saint Petersburg) before emigrating, there's a parallel to the Communist Manifesto and Lenin's manifesto phrase about building a socialist state that will benefit everyone but needing a temporary phase of dictatorship where everything is centralized. The claim is that once benefits are distributed equally, the centralized control will no longer be needed. This thought reminds of "we'll take care of everyone. There will be a period when everything needs to be centered around us, but this is safer and better. At some point, the benefits will be so widely distributed that it will no longer matter."
Very excited about scientific progress. Having developed an interest in physics after moving to the United States, there has been continuous pursuit of science. For the past few years, language models have been tested with PhD thesis requests, where getting to the right thesis question is a big part of the work, but carrying out routine work also takes very long time.
Two years ago, models could only chat and couldn't do anything with the thesis. A year ago, they started answering at undergraduate level. Six months ago, they solved it at a real PhD level and solved it correctly. Now, when tested, they even give interesting and new information that wasn't considered at the time. With the right question, models can do all calculations and give really interesting results.
Given repetition speed, there's no need for a few years to complete a PhD—you can get a PhD in a week. Recent OpenAI news about mathematical theories proven in the last two weeks demonstrates what's happening is crazy and absolutely unbelievable for science.
Excited about sciences of the real world: life sciences, materials science, and chemistry. There are many areas where AI can interact with real-world experiences, setting up special data loops that will be slower than theoretical work but significantly faster than before.
Still underestimating how radically this has transformed software engineering and how much can be built now. It's an amazing and very positive tool in many ways already. This also applies to manual and professional labor.
Visited Stargate a few weeks ago and was surprised by the number of cars in parking lots. There are a lot of people working there. Data center projects create tens of thousands of jobs, and they are high-paying jobs. Data centers now seem similar to what factories used to be in the twentieth century, when railways and factories made cities enjoy vibrant economies.
When visiting data center sites, there are concerns, but the reality is they create lots of jobs and provide tax revenue for local communities. Certain things need intelligent handling, like noise requiring soundproofing or placement away from residential areas, but for local economies they create jobs and generate tax revenues just like factories used to.
Scientific and experimental efforts at Beam use the same models to improve training efforts. Research is under supervision of co-founder Yannis who leads research and technology. Many things known to be effective couldn't be included due to schedule constraints. When training the final model, the SpaceX computing suite acquired in July was used, taking a few weeks of initial training, then a few weeks of synthetic data and reinforcement learning before release.
There's a trade-off between precise execution of known things or less risky bets versus moving to riskier bets. The goal is to exhaust all high-impact things known to work, then start expanding the risk profile through small-scale expansion experiments. Some things only become apparent at wide scale, requiring art and science in deciding whether to invest more computing to discover the next thing or move to the next idea.
Having the model in the workshop gives researchers great capabilities because things that used to take a long time individually can now be done very quickly. Models vary in intelligence levels, so their strengths need to be understood along with the need for human creativity. Tasks should be automated where possible, especially searching for supertransactions and finer aspects of infrastructure.
Models are very exciting tools because they enable scientists to exercise intuition. They're like having a fast and enthusiastic colleague who actually listens, and if a good match, researchers gain much greater ability to move much faster than before. The annual improvement rate was roughly 7 times higher previously with manual work, now several times that—maybe 4 to 5 times higher.
The question is when researchers will no longer be needed within the workshop. For the same effort, fewer people might be needed, but ambition constantly expands. There's a ratio between employees and computing power that must be maintained. Having 3,000 researchers wouldn't be very useful, but having a few hundred would be extremely useful.
Given computing power allocated to each person and expanding scope of experiments, there's a distribution in terms of contributions—few people contribute the most ideas in each field. This leads to giving extra computing power to the most productive group, resulting in fixed or decreasing numbers of researchers over time. Projects like AlphaGo had around 10 people working on them. For efforts as big as current work, around 100 people or hundreds are probably needed.
Large projects have never required exceptional numbers of people. The number of researchers required hasn't changed significantly—large projects have always required about 100 people. There will be a steady level of around 100 people. There will be many job opportunities in applied research, taking techniques and applying them to real problems.
Many other types of models can be built, and other closed-loop systems developed. Evidence suggests companies need more engineers, not fewer. There's proliferation of engineers of certain quality levels to disseminate technology in institutions that couldn't previously access this level of talent. The focus should be on GDP transformation through spreading people who transfer technology abroad.
The number of researchers required hasn't changed significantly. Large projects have always required about 100 people. Much of the need for hundreds of people comes from teams conducting assessments and collecting data. Models don't gain capabilities by chance—small groups of 5 to 10 people work to target specific abilities. In programming, there are between five and ten abilities, but this aims to provide general capabilities.
When taking models and applying them to real-world capabilities, a unit is needed for each capability, and every organization has plenty. Creating job opportunities for field engineers who tend more towards scientific and evaluative sides with intuition about models and testing systems is what evaluation researchers do. These are the same skills as researchers who build basic models, but employed to work on real problems. There's a training gap at this level—engineers need training to become proficient, and as many as possible are ready to be employed today.
Keep No Priors: AI, Machine Learning, Tech, & Startups in your library
Save the videos and channels worth coming back to, and find them again in one place.





