Top GitHub Repos for Building AI Agents in 2026
The top GitHub repos for building AI agents in 2026: agent frameworks, MCP, browser automation, memory, sandboxes, evaluation and observability tools.

The top GitHub repos for building AI agents in 2026 fall into five groups: agent frameworks (LangGraph, OpenAI Agents SDK, Claude Agent SDK, Pydantic AI, Google ADK, Microsoft Agent Framework, CrewAI), the Model Context Protocol SDKs and servers, browser automation (Browser Use, Playwright MCP, Stagehand), infrastructure such as memory and sandboxes, and evaluation and observability tools. You rarely need more than one or two from each group.
Every repository below was checked as of September 2026: it exists, it is actively maintained, and we note the licence where GitHub reports a standard one. We deliberately leave out star counts because they change daily and say little about fit.
How we picked these repositories
- Active maintenance — recent commits and releases, not an abandoned demo.
- Production use — clear docs, versioned releases, a real user base.
- Clear licence — permissive licences are listed; others are flagged “check repo”.
- Durability — backed by a company or a large, healthy community.
Agent frameworks
A framework handles the loop: calling the model, running tools, managing state and handing off between agents. Pick based on your language and how much control you need. Our AI agent frameworks compared post goes deeper.
| Repository | Language | Licence | Best for |
|---|---|---|---|
| langchain-ai/langgraph | Python, JS | MIT | Explicit graph-based workflows with state, checkpoints and human-in-the-loop |
| openai/openai-agents-python | Python | MIT | Lightweight multi-agent workflows with handoffs and guardrails |
| anthropics/claude-agent-sdk-python | Python (TypeScript version also available) | MIT | Building agents on the same tool harness as Claude Code |
| pydantic/pydantic-ai | Python | MIT | Type-safe agents with validated, structured outputs; model-agnostic |
| google/adk-python | Python | Apache 2.0 | Code-first agents with built-in evaluation and deployment paths |
| microsoft/agent-framework | Python, .NET | MIT | Enterprise multi-agent orchestration; the successor to AutoGen |
| crewAIInc/crewAI | Python | MIT | Role-based teams of agents, quick to prototype |
| huggingface/smolagents | Python | Apache 2.0 | Minimal agents that act by writing code |
| mastra-ai/mastra | TypeScript | Check repo | Agents and workflows in a TypeScript/Node stack |
| vercel/ai | TypeScript | Check repo | Tool calling and streaming UIs in web apps |
A note on AutoGen: microsoft/autogen is now in maintenance mode, and its README directs new users to Microsoft Agent Framework. Don’t start new projects on it.
Frameworks for retrieval-heavy agents
If your agent mostly searches and reasons over documents, these are worth a look:
- run-llama/llama_index (MIT) — document parsing, indexing and agentic retrieval.
- deepset-ai/haystack (Apache 2.0) — pipeline-based orchestration for search and RAG.
- stanfordnlp/dspy (MIT) — “programming, not prompting”: optimises prompts and pipelines against a metric.
See our RAG best practices for the retrieval side.
Model Context Protocol (MCP)
MCP is the open standard for connecting agents to tools and data. Write a tool server once and it works across Claude, ChatGPT, IDEs and most frameworks above.
- modelcontextprotocol/python-sdk (MIT) — official Python SDK for servers and clients.
- modelcontextprotocol/typescript-sdk — official TypeScript SDK.
- modelcontextprotocol/servers — reference servers (filesystem, git, fetch and more) that are also good examples to learn from.
Treat third-party MCP servers like any dependency: read the code, pin versions and limit their permissions. Our MCP servers guide covers what to install and what to avoid.
Browser automation
For agents that need to use websites without an API:
| Repository | Licence | What it does |
|---|---|---|
| browser-use/browser-use | MIT | Python library that lets an LLM agent drive a real browser |
| microsoft/playwright-mcp | Apache 2.0 | MCP server exposing Playwright, using the accessibility tree rather than screenshots |
| browserbase/stagehand | MIT | SDK mixing natural-language actions with deterministic Playwright code |
Browser agents are powerful but fragile and expensive. In our projects we prefer an official API or a deterministic Playwright script wherever one exists, and use an LLM only for the steps that genuinely change. Remember that everything on a web page is untrusted input to your agent.
Memory, state and sandboxes
- letta-ai/letta (Apache 2.0) — a platform for stateful agents with long-term memory.
- mem0ai/mem0 (Apache 2.0) — a memory layer you can add to existing agents and apps.
- e2b-dev/E2B (Apache 2.0) — isolated cloud sandboxes for running agent-generated code.
Add persistent memory only when you have a clear use for it, and decide up front what the agent is allowed to remember and for how long, particularly under GDPR.
Evaluation and observability
This is the group most teams underinvest in. Without traces and evaluations, you cannot tell whether a prompt or model change made the agent better or worse.
| Repository | Licence | Use it for |
|---|---|---|
| promptfoo/promptfoo | MIT | Test suites for prompts, agents and RAG, plus red teaming, from the CLI or CI |
| confident-ai/deepeval | Apache 2.0 | Pytest-style LLM evaluation metrics |
| vibrantlabsai/ragas | Apache 2.0 | Retrieval and RAG-specific metrics |
| langfuse/langfuse | Check repo | Self-hostable tracing, evals and prompt management |
| Arize-ai/phoenix | Check repo | Tracing and evaluation built on OpenTelemetry |
Our LLM evaluation guide shows how to wire these into CI.
Coding agents you can learn from
Open-source coding agents are excellent reference implementations of tool use, planning and context management:
- openai/codex (Apache 2.0) — OpenAI’s terminal coding agent.
- cline/cline (Apache 2.0) — coding agent as an IDE extension, CLI or SDK.
- continuedev/continue (Apache 2.0) — open-source coding agent.
We compare commercial and open-source coding tools in AI coding assistants compared.
A sensible starter stack
For a typical business agent in Python, a stack we’d be comfortable putting in production:
- One framework — LangGraph or Pydantic AI for explicit control, or the OpenAI Agents SDK / Claude Agent SDK if you are committed to that vendor.
- Tools via MCP — using the official SDK, with each server’s permissions kept narrow.
- Tracing from day one — Langfuse or Phoenix.
- An evaluation suite in CI — promptfoo or DeepEval, built from real tasks.
- A sandbox for any code execution.
Choose the design pattern before the library. Our agentic AI design patterns guide explains which patterns need which capabilities.
Key takeaways
- Pick one framework that matches your language and control needs; avoid stacking several.
- MCP is the standard way to give agents tools; use the official SDKs.
- AutoGen is in maintenance mode; new projects should look at Microsoft Agent Framework or alternatives.
- Browser automation works, but prefer APIs and deterministic scripts where possible.
- Evaluation and tracing tools matter as much as the framework.
- Check licences and maintenance status before adopting, and pin versions.
From repositories to a working agent
Open-source building blocks make agents much faster to build, but choosing, securing and evaluating them still takes experience. Our AI solutions team helps companies pick a lean stack, build agents around real business processes and keep them measurable in production. If you’re planning an agent project, get in touch.