Back to Sequoia Capital

How Harvey Built a Research Lab on a Budget | Gabe Pereyra

Sequoia CapitalAugust 11, 202629m
Topics59
Introduction to Gabe Pereyra and Harvey0:00Building a Research Lab on a Budget0:31Using the Frontier Ecosystem1:01Building Benchmarks First2:01Legal Agent Bench2:31Contracting and Diligence Data Sets2:31Challenge of Training on Customer Data3:01Domain Experts for Synthetic Data Generation3:30Training Lawyers on Coding Models4:00Scaling with Mercor and Snorkel4:33Efficient RL Environments5:01Open Sourcing Data Sets5:31Open Source Models Getting Competitive6:01Models for Post Training6:31Working with Neo Labs7:01Partnerships with Multiple Providers7:30Why Work with Multiple Neo Labs8:00Building Composer One8:32Model Serving Infrastructure9:02Multiple Models and Fallbacks9:31Pre-Production Evaluation10:01Generic Evaluations10:30Product Surface Testing11:01Production Monitoring11:30Simple Open-Source Switches12:01Building the Muscle12:30The Post-Training Flywheel13:02The Playbook for Building a Research Lab13:30Moneyball Reference14:01Q&A: Synthetic Data Generation Process15:01The Chicken-and-Egg Problem15:31Julio's Approach16:01Checking Model Performance16:30Interest from Law Firms17:00Hiring Challenges17:32Talent Pool Opening Up18:02Infrastructure APIs Reducing Barriers18:31Benchmark Rubric Creation19:01Building Intuition Over Time19:30Debugging the Pipeline20:31Product-Market Fit as Foundation21:01Building the Muscle for New Models21:30Treating Post-Trained Models as New Models22:00Overfitting and Out of Distribution22:31Philosophy on Open Sourcing Benchmarks23:00Customer Model Preferences23:30Bandwidth for Experimentation24:01Strategic Value of Data24:31Remaining Open Questions25:00Gap Between Benchmarks and Production25:31Managing Large Context Windows in Legal AI25:51Post-Training and Continual Learning26:01Operationalizing Customization at Scale26:31Foundation Readiness and Infrastructure27:01From Individual to Organizational Productivity27:30Project-Level Coordination Challenges28:00Organization-Level Resource Management28:31Enterprise Complexity29:01Hyper-Vertical Domain Focus29:01
In a Nutshell

Harvey's Gabe Pereyra shows how application-layer companies can compete with frontier labs by using open-source models, Neo Labs for post-training, and domain experts to generate synthetic legal data that bypasses customer data restrictions. The playbook starts with building realistic benchmarks like Legal Agent Bench and large diligence datasets, then moves through efficient RL environments, open-sourcing data for validation, and serving post-trained models alongside closed-source ones with proper evaluation infrastructure. Key insight: frontier ecosystem tools now make it possible to reach frontier-level intelligence on specific domains like legal work without the resources of big labs.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

Gabe Pereyra is co-founder and president of Harvey. He was previously a research scientist at DeepMind and then at Meta before showing his college roommates what GPT-3 could do. The duo then became Harvey.

The alternative title to this talk is building a research lab on a budget. It is an unfair game competing with the frontier labs if you are an application layer company. There are rich teams, poor teams, and then application layer companies. The frontier labs have more money, more talent, more compute infrastructure, and data.

The way to compete is by using the frontier ecosystem. When Harvey started 4 years ago, most of these companies either did not exist or were just getting started. Companies either had to build everything themselves or focus on something different like building their GTM org and a great product. Today using the frontier ecosystem you can compete with the frontier labs and build frontier intelligence.

The high-level playbook includes building benchmarks and training data, working with the Neo Labs to do post training, and serving models in production. To start, you want to build a benchmark. If you do not have a good benchmark, you cannot train models, and if you cannot train models, you do not need to serve them in production.

This year Harvey released three data sets. They started by building Legal Agent Bench, which is a taxonomy of tasks that associates would do at a large law firm. These cover multiple practice areas and include complex tasks like drafting complex fund formation documents and doing case law research.

Sign in to read the full notes

Get access to AI-generated notes, topic timestamps, and more.