Bring type safety to your agent stack. Build Python agents with structured outputs and a familiar developer experience.
FrameworksBest Of Agent Harnesses
🏆 Curated, ranked list of AI agent harnesses (100+) — plus an MCP server, llms.txt & JSON so agents can recommend them too. Rescored weekly.
Building and shipping AI agents
3 releases
Official GitHub release feed
Follow the changes
The release trail
Original notes, ready to explore. Open a release to see what changed.
Best Of Agent HarnessesMid-September 2026: the harness-matters-more catalog, QM, Prime Agent, OpenJarvis
From the release notes
What's new Three harnesses from the YC Paper Club talk are in. Prime Agent (Prime Intellect, 95.5% on ARC-AGI-3 with Opus 5), QM (Y Combinator's multiplayer harness for work), and OpenJarvis (Stanford's local-first stack). 167 harnesses, stars captured 2026-09-13. Three new decision guides (16 total). Why the harness matters more than the model: the measurements (ARC-AGI-3 30% to 95.5% on the same weights, SWE-bench Pro 23% to 52%, Cursor 46% to 80%, Harness-Bench rank correlation -0.05), a table of twenty labs, papers, and practitioners with checked links, and what the claim does not mean, from the YC Paper Club session of September 7. Managed vs self-hosted always-on agents: Grok Bot, Claude Managed Agents, QM, OpenClaw, Hermes, OpenJarvis. Who owns the computer, who pays for idle time. The best AI agent harnesses in 2026, ranked by category: the top three per category from the live data, regenerated on every weekly refresh. README and machine-readable surfaces. The intro now carries the long-horizon numbers, the two eras of harnesses, and the pairing rule. Two new FAQ entries (the thesis, Grok Bot) flow to the README, llms.txt, harnesses.json, and the site's FAQPage. llms.txt states the thesis in its blurb; the site adds llms-full.txt (llms.txt plus every guide). Site. Every guide page carries TechArticle JSON-LD with real published and modified dates, breadcrumbs on all project, category, and guide pages, and a per-URL sitemap date. Guide links now resolve on the site (they previously pointed at .md paths that 404 under /compare/). MCP. compare prefers a guide that names the compared projects in its title. The 0.5.2 package on PyPI already serves the new projects and guides from main; the tie-break ships in the next package release.
Best Of Agent HarnessesSeptember 2026 update
From the release notes
What changed since August 12 MCP server works again. agent-harnesses-mcp 0.5.2 pins the MCP SDK below 2.0. The 2.0 release (July 28) renamed FastMCP, and 0.5.1 crashed on startup for anyone who installed it. Install: claude mcp add agent-harnesses -- uvx agent-harnesses-mcp. Cold-tested: initialize and tools/list complete, 10 tools. 164 harnesses, 12 categories. New in the tables: YYLO (coding agent products), AnythingLLM (personal agent runtimes), ClawBench (evaluation and benchmarking). Flowise stays listed with its archived flag instead of moving to the Graveyard. Roo Code is out of the turnkey coding agent pick and stays in the tables with its archived flag. On the radar. DeepSeek Harness (pinned, not ranked, pending a star-velocity check on 2026-10-07), Google Antigravity SDK, L∞pGate, and eight community submissions below the ranked-table bar. Community submissions now reach the curation queue. Open "Add project" issues feed the weekly discovery run, so the biweekly curation pass vets them with live repo data and closes the issue when an entry lands. Site. The nav shows a live star count that links back to the repo. Weekly rescores on August 16, 23, 30, September 6 and 9 refreshed every star count and the landscape charts. Thanks to contributors HelpMatey (#61), InsightFactoryAPP (#111), reacher-z (#60), rxdt (#69), and danawoodman (#78).
Best Of Agent Harnessesmcp-v0.5.2
From the release notes
agent-harnesses-mcp 0.5.2: pin the mcp SDK below 2 so the server star…
A collected snapshot, not the complete archive.
Keep connecting the dots
There’s more where that came from.
Small library, big possibilities. Let agents solve tasks by writing and executing code with Hugging Face’s toolkit.
FrameworksA code-first toolkit for building, evaluating, and deploying agents, from a single task to a multi-agent system.
Frameworks