Self-Improving Harnesses, Local Personal AI And YC's Agent For Work | YC Paper Club
In a Nutshell
Harnesses now drive most AI progress through self-improving architectures that enable persistent memory, programmatic sub-agents, and test-time learning—outperforming static prompt engineering. Prime Agent, Open Jarvis, and YC's QM demonstrate this with concrete results: 95%+ on ARC-AGI, 800x cost reduction for local AI, and fleet-scale agent deployment across 50+ VMs. The key shift is from fixed tool loops to agentic operating systems where models dynamically manage context, spawn sub-agents, and accumulate capabilities through continual refinement.
These notes were generated by AI and may contain inaccuracies.
The event opened with a discussion of the new YC Paper Club visual design created by Ev, head of design at YC. The speaker explained that harnesses represent scaffolding and prompt engineering rather than traditional research, despite recent community pushback against considering prompt engineering as legitimate research at top-tier machine learning conferences.
Harnesses deliver an 18% performance improvement between different implementations. This difference determines whether ARC-AGI functions or fails entirely. The speaker referenced meter plots showing release dates versus agent runtime duration, demonstrating that much recent progress stems from harness improvements rather than model intelligence gains alone.
The speaker distinguished between the static harness era, where harnesses remain fixed, and the recent six-month period focused on self-improving harnesses. A plot from Trajectory's CEO illustrated how research emphasizes model perplexity and IQ metrics while neglecting test-time experience adaptation. The speaker described experiments increasing sample counts online and the challenge of learning from batch size one, noting that in-context learning saturates after 40-50 examples without further improvement on validation sets.
ARC-AGI emphasizes rapid adaptation to new problem distributions. Claude Opus achieved 30% on the private holdout set verified by Greg and Chalet. Harness improvements elevated performance to 95%, while AVO from Nvidia reached 100%. The speaker mentioned forking Karpathy's auto-researcher in March and accidentally building a harness while creating a user interface.
Sign in to read the full notes
Get access to AI-generated notes, topic timestamps, and more.