Rust at OpenAI: Performance, Safety and Scale
In a Nutshell
Predrag Grkovic shows how OpenAI uses Rust for performance-critical systems and AI-assisted engineering workflows. He demonstrates that AI tools enable shipping production code without line-by-line review while maintaining higher correctness standards through comprehensive testing and verification. The key insight is that Rust's safety guarantees become more valuable as AI models gain cybersecurity capabilities, making Rust the optimal choice for foundational systems where correctness and efficiency compound at scale.
These notes were generated by AI and may contain inaccuracies.
Orhun, Rust developer advocate at JetBrains and lead maintainer of Ratatouille, hosts a livestream series in collaboration with the Rust Foundation exploring the intersection of Rust and AI. The series examines how Rust helps build AI systems and how AI can help build better Rust software. The current episode focuses on Rust engineering at OpenAI.
Predrag Grkovic works on Frontier systems at OpenAI and builds Rust developer tools. He created CargoStandardChecks and TrustFall, focusing on software correctness, performance, and security. He is an open source maintainer.
Predrag entered Rust because of its emphasis on building high-quality software with minimal maintenance, efficiency, and respect for user and maintainer time. The Rust community aligned with his preference to minimize debugging time and maximize productive building.
"The Rust community was very welcoming. It has a very strong ethos aligned like this. I think the Rust code that exists out there was very high-quality, the ethos of just make the types do the work, and if it compiles, it runs."
Predrag started with Advent of Code in Rust. Initial challenges included the borrow checker, but learning Rust improved his skills in other languages. Rust's strictness forced deeper thinking about code correctness, making him a better engineer overall.
Serde and PyO3 libraries transformed his view of Rust. Serde provides every serialization format needed. PyO3 enables Python bindings with minimal code. These libraries allowed rewriting performance-critical Python components in Rust while maintaining seamless integration with existing Python systems.
CargoStandardChecks aims to catch breaking changes between Rust releases. The project originated from a conversation with Luca Palmieri, who suggested using Predrag's TrustFall query engine to solve the semantic versioning problem in the Rust community.
The tool addresses human error in semantic versioning. Even with best intentions, developers make accidental breaking changes. The goal is to catch these before release, preventing downstream users from discovering broken functionality after publication.
"The thing that the tool tries to do is to make sure that we don't make mistakes."
CargoStandardChecks became part of Rust compiler CI. Historically, Predrag identified about six accidental breaking changes over five years that could have been caught. The integration required handling special cases for stable versus unstable features in the Rust standard library.
The Rust standard library has unique requirements. Items may be stable individually but have unstable default implementations or const usage restrictions. The tool needed to distinguish between stable and unstable features to avoid false positives on nightly-only changes.
Predrag used AI to identify edge cases during integration. The approach involved iterative questioning to uncover potential false positives before implementation. This proactive edge case discovery prevented future debugging time for Rust maintainers.
The core TrustFall query engine required minimal changes. Work focused on the interface layer to handle Rust standard library special cases transparently, allowing the same lints to work for both standard library and regular Rust libraries.
Predrag's AI usage evolved over time. Initially used for test case generation and autocomplete, AI capabilities improved significantly. Around six months before the interview, he began shipping code without complete manual review when instruction following became reliable.
Predrag developed a key insight: not every line of code deserves equal attention. Low-stakes code like GitHub Actions scripts can be AI-generated and tested, freeing time for high-stakes code like unsafe blocks where correctness is critical.
Recent CargoStandardChecks development included 30 PRs focused on GitHub Actions security hardening. AI helped identify potential security issues that Predrag wouldn't have discovered independently due to time constraints and lack of specialized GitHub Actions expertise.
The first instance of shipping AI-generated code without complete manual review occurred approximately six months before the interview. This coincided with improved instruction following capabilities in frontier models.
The speaker describes shifting from line-by-line code review to higher-level concerns about API design and testing strategies. Rather than manually reading every line, they focus on ensuring code correctness through comprehensive testing approaches including code coverage, branch coverage, mutation testing, and deterministic simulation.
The speaker references a talk given two years ago at EuroRust on advanced testing techniques. These methods were previously impractical due to high effort requirements, but AI tools now make them feasible. They describe describing processes to AI such as testing state machine structures, ensuring state and transition coverage, running coverage tools on non-trivial logic, and using cargo mutants to introduce intentional bugs and verify test coverage.
For unsafe code, the approach involves considering all possible cases and testing them with Miri, ASAN, and other verification tools. The emphasis is on structure and coverage rather than variable naming or implementation details.
The speaker advocates spending up to 45 minutes crafting detailed prompts that describe desired outcomes and criteria for "good enough" rather than specifying exact code. These prompts can generate billions or tens of billions of tokens, with the focus on having confidence in the code structure rather than reading every generated line.
The speaker identifies as a tool builder at heart, having created cargo sandwich checks as open source tools. The goal is to solve problems once and eliminate them permanently rather than dealing with recurring issues.
An example is given of using voice recording to generate prompts for a terminal UI project, allowing the developer to avoid fixing issues one by one and instead provide comprehensive instructions to an AI agent.
The speaker answers whether they still write code by hand, stating they do for unsafe code where consequences of errors are dramatic, sometimes spending four hours on 100 lines. However, most code doesn't require this level of manual attention due to confidence in the AI-assisted process.
The approach involves giving sketches of desired approaches in prompts, similar to working with a skilled colleague who needs context. The speaker provides big-picture guidance about types, functions, and data structures, then relies on the AI to expand and implement while following directions.
In early July, the speaker spent a week exploring Rust compiler and standard library speedups. Rather than simply asking AI to optimize the compiler, they spent 45 minutes writing a prompt describing what constitutes an acceptable performance optimization and criteria for upstream contribution.
The strategy focuses on identifying code that hasn't been recently optimized rather than well-trodden paths. The approach involves examining commit chains for performance-related comments, GitHub discussions about benchmark results, and identifying hot code paths that haven't received recent attention.
The preferred optimizations are those already proven successful elsewhere in the Rust project, applied to new locations. Bonus points are given for one-line or ten-line changes that demonstrate clear improvements, such as 5% speedups from single-line modifications.
The prompt specifies running benchmarks in interleaved fashion (baseline, new, new, baseline) to avoid order dependence, with five iterations yielding ten runs of each variant. Statistical requirements include 1% or greater improvement in geometric mean to ensure results aren't phantom optimizations or environment-specific.
The outcome was a single-line deletion in the new trait solver that resulted in 7.5% performance improvement. The speaker notes that while AI discovered the optimization, the change was obvious once identified - recomputing a value that had been calculated 40 lines earlier.
The speaker references the concept that AI isn't merely a faster horse but enables work that would have taken five years manually, accomplished instead with billions of tokens and 45 minutes of prompt crafting time. The emphasis is on leveraging human time and expertise more effectively rather than replacing it.
While acknowledging token costs, the speaker notes that human time and expertise also carry costs. The goal is to automate tedious manual work so experts can focus on higher-impact activities, ultimately producing faster, more correct, and more secure software.
The speaker responds to concerns about AI-assisted development being dismissed as "vibe coding." They distinguish between casual prototyping and serious engineering with AI tools, noting that their approach expects higher quality standards than manual development alone could achieve.
For personal applications like home automation, lower standards are acceptable since failures have limited consequences. However, the speaker's professional work demands higher bars, and AI tools enable achieving greater correctness than would be possible without them.
The speaker cites their CargoSanverChecks tool finding sanitizer violations both in user code and within the tool itself, preventing releases and requiring version bumps. This demonstrates that even experienced developers benefit from additional verification layers.
Rather than blanket statements about not reading code, the speaker explains paying attention selectively based on risk levels. Not every line of code carries equal potential for problems, so attention is focused where errors would have significant consequences.
The correct approach involves understanding where AI tools excel and where they may fall short, then augmenting engineering practices to achieve results greater than the sum of individual parts. Poor usage of Rust (like adding unsafe to silence the borrow checker) is discouraged, just as poor usage of AI tools should be avoided.
People frequently ask about advanced techniques for AI-assisted programming, also called vibe coding, and specifically how to prompt LLMs to write reliable Rust code. One key insight is that integrating testing tools into the development workflow is essential. Another effective approach is asking models to make pragmatic decisions rather than being overly thorough.
If you ask an AI to test everything, it will generate tests for unrealistic scenarios like meteor strikes coinciding with power outages during panic handling. While this level of rigor makes sense for aircraft code, it is not effective or helpful for every piece of code or every tool. The critical skill is telling the model what assumptions are acceptable, giving it more context, and approaching the interaction with empathy - if the model produces a bad result, it will lead to a bad outcome.
Many struggle with writing prompts that are precise enough. An example from Blue Sky illustrates this: when asked whether X, Y, and Z constituted a breaking change, the answer was yes with six caveats, including whether the trait already existed, was implemented in an existing type, was not sealed, and other conditions. Mastering the art of being precise about what is wanted and which assumptions are acceptable versus unacceptable leads to dramatically better outcomes.
An optimization prompt for the Rust compiler took 45 minutes to outline what acceptable looks like - requiring a tiny change, understood technique, statistically sound results, benchmarks, and other criteria. This iterative process involves building something, evaluating whether it meets expectations, analyzing what led to unsatisfactory decisions, asking for clarification when needed, and then amending prompts, personalization files, strategies, or skills accordingly.
The recommended approach is treating AI-assisted development as an engineering problem where all issues are solvable. Once you can describe why something was undesirable because of specific reasons and incorporate that into your process, it becomes something that cannot cause problems again. The most value comes from ensuring designs are sound through significant effort and iteration, after which code naturally falls into place. The focus should be on thinking about the problem space more deeply rather than simply typing faster.
Proficiency with AI tools is not innate - it required many billions of tokens to develop. OpenAI's own data shows skyrocketing token usage among researchers and engineers, reflecting the process of discovering better ways to use the tools, which leads to more applications and faster overall progress.
A frontier systems engineer focuses on using GPUs for training models as efficiently as possible. The more efficiently GPUs are utilized, the cheaper it becomes to perform the same amount of work. The goal is to save data center seconds and avoid wasting data center months. A small code optimization multiplied by the size of a data center can have enormous financial impact, but a mistake that occurs one time in a billion can cause data center months of corruption and damage.
This is why AI is used to catch problems proactively. A real example involved data corruption occurring 0.00000001% of the time - which was effectively all the time - caused by approximately 20 lines in an unsafe block. Despite being human-written, human-vetted, and having passed code review, the code was still incorrect. Recent work has involved reviewing unsafe code in OpenAI's codebase and dependencies, including libraries like zero-copy, IDDQD, and Tokio, where AI models discovered unsound code and potential memory safety problems.
None of these issues represented major security vulnerabilities like Heartbleed, but they were still incorrect code. The unsafe review skill developed through the Rust community is open source and used internally at OpenAI. As models improve, more bugs can be found. The philosophy is that when skilled, passionate, and diligent people make mistakes that AI can catch, those tools should be used.
A concrete optimization example involves training model recovery. When training crashes, state must be restored from checkpoints to all participating GPUs. This workload has the property that recovery cannot complete until the last GPU is online. As more GPUs are added, the worst-case latency worsens because there are more draws from the latency distribution. The cost in GPU hours grows super-linearly.
The goal is shrinking recovery time. Hedging, described in a Jeff Dean paper, involves requesting data from multiple replicas and using whichever responds first. If one server or network path is slow, selecting a different source minimizes the chance of getting unlucky twice. However, this is easy to explain but difficult to implement because the data can be gigabytes in size, streamed, and may involve gather-scatter semantics rather than linear memory access.
The naive approach of allocating two different buffers works poorly for large or nonlinear buffers. The implemented solution used unsafe code to allow two futures to share a buffer, writing data concurrently but not in parallel. This requires precisely walking the tightrope of unsafe code - being slightly wrong results in undefined behavior, but getting it exactly right delivers major performance improvements.
The process began by reading Tokio documentation and forming a hypothesis that this could be done. Codex was asked to build a small prototype and run it through Miri to check for undefined behavior. After the prototype passed, a two-page document was written outlining the need, with Codex iterating on it and finding several mistakes. The document was reviewed by the team.
The feature was implemented entirely with Codex without looking at the code, then shipped to a shakeout run rather than production to verify behavior at scale. After no issues emerged, all the code was discarded. A new approach was requested where Codex would structure the work so every incremental piece of unsafe code could be reviewed, with discussion of every 20 lines until both were confident in correctness. Weeks were spent on this process.
The result was massive performance impact that worked the first time in production. The final code was approximately 1,500 lines with a couple hundred lines of unsafe code, but weeks were spent with AI in the loop. The prototypes that were never read served to improve how Codex would think about the problem so the final implementation would avoid blind alleys.
TrustFall represents the second iteration of a system, with the first written in Python and open sourced separately. The quality of TrustFall stems from six or seven years of the prior system's existence, allowing all failure modes to be observed and incorporated. This represents speed-running the iterative process: build, apply learnings, build again with greater fidelity, until the final stage where the implementation must meet a very high quality bar.
Security work evolved gradually from performance optimization work. Cargo Sanford Checks demonstrated that even deeply caring and skilled humans can make mistakes. Security became a natural extension because unsafe code involves many rules that are difficult to uphold. When security news emerged about Mozilla and Chrome finding CVEs, thinking turned to unsafe Rust.
The approach was collaborative - consulting experts at conferences like Rust Week and Rust Conf, including Josh LF, Jack Ren, and Manish. Zerocopy was identified as an exemplary library for careful unsafe code. The hypothesis was that writing a skill codifying review rules that the biggest experts follow would level up model capabilities. The skill was developed collaboratively with community input, including 18 suggestions from Astra for clarity and edge cases.
Scanning began with Zerocopy itself, finding approximately a dozen incomplete safety proofs and edge cases no human would have conceived, such as complex alignment interactions between repper packed types and nested alignment requirements. This was eye-opening because Zerocopy represents a gold standard for unsafe in the Rust ecosystem. Work expanded to Tokio maintainers, resulting in fixes in latest versions, and to Rain on IDDQD where several issues were found.
Not all issues have equal impact. Well-maintained codebases have a long tail of bugs that nobody would encounter, alongside a small number of more serious issues. The consistent philosophy is preferring to know what is happening, whether regarding breaking changes or unsafe issues, and thoroughly analyzing concerns even if the conclusion is that something is not a concern.
A security issue was discovered involving Cargo Miri and GitHub Actions making different assumptions about caching. This resulted in GitHub personal access tokens potentially being saved in actions cache and made available to anybody opening a PR on a repository. The issue was found and patched through systematic security review efforts where findings were evaluated for severity, leading to the identification of this critical vulnerability. The Rust security response team later announced a vulnerability in Miri involving saving environment variables in Rust's build directory that could be cached in GitHub Actions, exposing secrets.
The Rust community needs to take security findings seriously without being dismissive or panicking. The appropriate response requires urgency without being oblivious, avoiding both hiding problems and creating panic and havoc. Every new tool has revealed problems that were previously not findable but always existed. The response should be to move quickly to assess how bad problems are and fix them, recognizing this is neither the first nor last time such discoveries occur.
Astra and Codex models were used to discover the Cargo Miri vulnerability that was then reported. There are ongoing discoveries of crate security vulnerabilities, similar to how CVEs are regularly disclosed and fixed in Linux kernel releases, including a zero-day in KVM virtualization stack. More such vulnerabilities should be assumed to exist.
When phones, computers, or browsers prompt for updates, prioritize installing them. Use the best tools available to scan controlled libraries, update running software, and update dependencies to newer versions. OpenAI has worked with Rust security engineers and the Rust Foundation to provide access to Codex Daybreak models, which are cybersecurity-specific models capable of scanning for vulnerabilities.
OpenAI provides the Codex for Open Source program that gives free Codex access to open source maintainers with generous rate limits, enabling automated triage and fixing of security issues. The most important approach is taking security seriously and extending grace in both directions - recognizing that open source maintainers receiving vulnerability reports may feel overwhelmed, while security teams face high volumes of findings across multiple projects.
Astra reaching critical cybersecurity capability threshold means it can discover large volumes of findings across both categories: the long tail of contrived, non-realistic, non-exploitable issues, and the smaller set of genuinely scary vulnerabilities requiring attention. The long tail of findings has gotten longer with more capable tools, making it draining to determine which bugs are security-relevant. Logic errors can be security relevant beyond just memory safety issues.
Rust's value increases with cyber-capable models because the language's guarantees become more valuable when models can combine multiple small issues. In C or C++, little slip-ups during normal operation may not cause problems but can be exploited by adversaries. In Rust, it's harder to have these slip-ups in multiple places, making it more difficult for models to cause severe problems. Instead of finding 20 remote code execution vulnerabilities, findings tend to be smaller issues that can't easily be combined for exploitation.
The Cargo Miri vulnerability was simultaneously bad and not bad - the potential impact could have been severe, but in practice most actions and configurations people used were not affected. When choosing languages for new projects or rewrites, Rust is recommended because fewer holes in the Swiss cheese with holes farther apart reduces risk.
An OpenAI team rewrote a Python service to Rust using two engineers, Codex, and GPT-5.5, achieving 6x better CPU efficiency and 15x better memory efficiency. An awesome Rust migrations repository documents these types of rewrites across the ecosystem.
Cyber-capable models make writing C and C++ dangerous, especially for code with wide audiences, though C and C++ won't be eliminated soon. Languages with stronger semantics should be emphasized for writing code with more confidence in correctness and efficiency over long periods. The more foundational code is, the more a language that is both efficient and secure makes sense. Rust currently checks these boxes as one of the best available tools.
PyO3 makes embedding Rust into other environments easy, and OpenAI sponsors its development. Full ground-up rewrites where old and new systems are developed in parallel then cut over are never perfect and usually problematic. Progressive rewrites are more effective - identifying hot pieces of Python code and rewriting them into Rust to create reusable Rust-based components with Python veneers that can later have Python interfaces replaced with direct Rust interfaces.
Using AI for prototyping helps determine if something is possible. Python was originally chosen for ease of experimentation, but some pieces became load-bearing over time, especially at fast-growing companies like OpenAI. High-traffic code running for every pull request and CI run represents significant compute spend. AI-assisted rewrites with skilled engineers in the loop are now easier to justify with good tools and libraries like PyO3.
Direct "write it in Rust" approaches without engineering consideration often result in bug-for-bug rewrites that don't provide desired value. Reproducing Python bugs in Rust can lead to ugly Rust code. The point of rewriting to Rust is often to eliminate Python limitations like the GIL, so exactly equivalent rewrites may not be optimal.
Approaching migrations as engineering processes to optimize produces better results. The programming community can be trusted to zero in on what works, iterate quickly, learn from positive and negative experiences, and reach good outcomes.
Astral, makers of UV and Ruff, was acquired by OpenAI. OpenAI is both a direct Rust user and serves users who use Rust through their products, making Rust ecosystem success important. OpenAI has worked with Charlie and Astral folks to continue GitHub sponsors programs funding important maintainers for Codex app and Astral tooling ecosystem.
OpenAI is a platinum member of the Rust Foundation with a board seat. A budget has been set aside for direct Rust development funding, including recent funding for Nicolas Nethercote and team to continue making Rust faster. OpenAI provides Codex for open source licenses to developers across ecosystems, not just Rust.
Fast, efficient, high-quality Rust experiences benefit both OpenAI internally and their users. Investment in faster compilers, faster code, faster iteration loops, and faster linkers like Wild and Mold (recently rewritten in Rust) benefits everyone through better software. Google, Microsoft, and AWS are also investing in Rust, indicating it's a major part of computing's future.
Daybreak cyber-capable models have been proactively made available to Rust Foundation security response and supply chain teams. OpenAI invests both monetarily and through tool access and token provision to achieve positive outcomes for the ecosystem.
Codex threads can message other threads and check out work trees, enabling powerful workflows. When finding problems in code, Codex can be instructed to triage the codebase, find all occurrences, and spawn separate threads in separate work trees to fix issues. Multiple threads can be organized in UI sections, collapsed, with notifications for merge issues.
Codex was used for house renovation lighting planning to avoid issues with inconsistent lighting. Codex researched lighting standards, determined requirements for a new office floor plan, calculated light quantities, arrangements, and qualities, and performed geometry calculations to determine if back-row lighting would cause monitor glare based on monitor specifications.
Current models represent the worst performance level they will ever achieve. Future models will be more aligned and skilled at using provided tools. People will simultaneously improve at using available tools. The most exciting development is the dramatic rise in everyone's skill level combined with model capabilities for advancing science and technology, with tools becoming massively better, faster, and more efficient.
Predrag expresses enthusiasm for the intersection of science and technology research, noting personal fascination with exploration and discovery. He references visible models of Curiosity rover and Mars on his bookshelf as symbols of this interest. He emphasizes that significant engineering and scientific work remains ahead.
He states that the future is fundamentally positive and that excitement about rapid improvement is appropriate. While acknowledging that rapid change can be unsettling, he places trust in engineers and humanity to direct tools toward beneficial outcomes. He also expresses confidence in security professionals working to protect systems, noting they are receiving improved resources and meaningful budgets to maintain safety.
When asked how Rust engineering will evolve as AI becomes more powerful, Predrag draws an analogy to the introduction of computers for mathematics. He explains that the result was not less math occurring, but dramatically more math happening throughout the world, with people becoming beneficiaries rather than direct performers of calculations.
He predicts there will be more Rust code than ever before, but with less human attention devoted to writing it. Simultaneously, the most critical Rust components will receive greater attention from both humans and machines, resulting in higher performance, correctness, security, and thoughtfulness than previously achieved.
He acknowledges that AI-generated code will include experiments and incomplete work, but argues against gatekeeping. Instead, he advocates for welcoming newcomers to the community and establishing productive norms for tool usage.
"It's kind of on all of us as tool builders and as participants in the community to be part of the solution and not just point to other people and say you're the problem and I have nothing to do with it."
Predrag shares his recently pinned social media post stating: "Some people use AI to write more code. I use AI to make my code more correct. We're not the same."
He encourages using AI tools not merely to increase code volume, but to identify potential failure modes, improve code quality across multiple dimensions, and establish proof of correctness for audiences who may lack domain context.
He describes his experience optimizing the Rust compiler and standard library, where proof consisted of short code snippets demonstrating improved benchmark performance, passing test suites, and readable code explaining the optimization rationale.
He advises working backwards from the question of how to convince others that work is sound, then adapting workflows to maximize tool effectiveness. The key insight is that the person being convinced may ultimately be oneself, validating that the tools have successfully built what was intended.
Predrag emphasizes maintaining high standards for code correctness, noting that every piece of code he pushes carries his signature through the Git client. He states that using AI has actually enabled him to confidently maintain these high standards rather than compromising them.
He hopes others will experiment with AI in this manner to maximize its value, viewing it as a catalyst for more interesting work rather than simply generating large quantities of code.
The conversation concludes with thanks to Predrag for contributions across the Rust ecosystem, open source, AI, and security research. Predrag expresses anticipation for future developments and continued engagement with the community.
Keep JetBrains in your library
Save the videos and channels worth coming back to, and find them again in one place.





