Back to Sequoia Capital

RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor

Sequoia CapitalAugust 12, 202626m
In a Nutshell

Mercor has shifted from crowdsourcing basic training data to building complex RL environments that teach AI agents how to use real-world tools like Salesforce, Microsoft 365, and Google Workspace across 205 economic domains. These environments require expert humans to create realistic "worlds," high-fidelity app clones, and precise verifiers—since models can't reliably grade their own work—with tasks costing $50-$10,000 each. Post-training on these environments has shown dramatic gains, like boosting corporate law performance from 4.7% to 26.6%, and the future demands ultra-long horizon tasks and virtual co-worker interaction capabilities.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

Record grew from a 1 to 2 billion dollar revenue run rate in the last 4 months. The conversation focuses on RL environments as a new frontier topic. In 2020, the era of crowdsourcing data for behavior cloning began, primarily involving supervised fine-tuning data with inputs and outputs, and RLHF data where annotators would select preferred model responses from several options. This approach enabled progress in fine-tuning GPT-3 toward ChatGPT and GPT-4.

By 2024, a significant transition occurred away from low-skilled crowdsourcing of behavior cloning data toward the agentic era of data. The focus shifted to finding the highest-skilled experts worldwide who could work collaboratively in teams to build frontier evaluations and RL environments for next-generation models. This includes software engineers, lawyers, doctors, bankers, and other professionals who can measure the frontier of intelligence and improve model capabilities.

Mercor originated with its first major project being deep research, becoming the primary agentic data vendor to all leading frontier labs and application layer companies including Harvey, Cera, Cognition, and Ramp.

Over the last 12 months, RLVR within the agentic data paradigm evolved to include RL environments featuring rich apps and worlds that teach agents how to use everyday laptop tools. This technology, initially developed in frontier labs, is now disseminating to application layer companies building their own intelligence.

Sign in to read the full notes

Get access to AI-generated notes, topic timestamps, and more.