AI Agent Frameworks Compared (2026): When to Use One vs Plain Code
AI agent frameworks compared as of September 2026: LangGraph, CrewAI, OpenAI Agents SDK, Claude Agent SDK, Pydantic AI, Microsoft Agent Framework and more.

Choosing between AI agent frameworks matters less than most comparison posts suggest. The core agent loop — call a model, run the tools it asks for, feed the results back, repeat — is a few dozen lines of code. A framework earns its place when you need what surrounds that loop: durable state, human approval steps, multi-agent orchestration, tracing and evaluation. This guide compares the main options as of September 2026 and explains when we reach for one versus writing plain code.
What an agent framework actually gives you
Every framework below wraps the same basic idea. What differs is the extra machinery:
- Tool calling with schemas generated from functions or types.
- State and memory across steps and sessions.
- Durable execution — resuming a long task after a crash or a deploy.
- Human-in-the-loop — pausing for approval before risky actions.
- Orchestration — graphs, handoffs or teams of agents.
- Observability — traces of every model call and tool call.
- MCP support — connecting to tools exposed via the Model Context Protocol.
If you need few of these, plain code is often simpler.
The main AI agent frameworks compared
Framework features and versions change quickly; the descriptions below are as of September 2026 and link to official documentation.
LangGraph
LangGraph describes itself as a low-level orchestration framework and runtime for long-running, stateful agents. You model your agent as a graph of nodes and edges with explicit state. Its strengths are durable execution (resume where you left off), human-in-the-loop interrupts, and short- and long-term memory. It reached a stable 1.0 release in October 2025 and does not require the rest of LangChain. It suits teams that want explicit control over every step and production features like checkpoints.
CrewAI
CrewAI is a standalone Python framework built around agents with roles, crews of agents that collaborate, and flows for event-driven orchestration with state. It is quick to prototype with, especially for role-based multi-agent setups (“researcher”, “writer”, “reviewer”). For production we tend to lean on its flows for predictable control rather than letting crews decide everything.
OpenAI Agents SDK
The OpenAI Agents SDK (Python, with a TypeScript version) is deliberately small: agents, handoffs between agents, guardrails for validating inputs and outputs, sessions for memory, and built-in tracing. It is the natural choice if you are mostly on OpenAI models and want minimal abstraction. Other providers can be used, but it is designed around OpenAI’s APIs.
Claude Agent SDK
The Claude Agent SDK (Python and TypeScript) exposes the same agent loop, built-in tools, context management, subagents, hooks, permissions, sessions and MCP support that power Claude Code. It shines for agents that work with files, run commands and operate on codebases or document workspaces, and it gives you permission controls out of the box. It is tied to Claude models.
Pydantic AI
Pydantic AI brings a type-safe, “FastAPI-like” feel to agents in Python. It is model-agnostic, validates structured outputs with Pydantic models, supports typed dependency injection and MCP, integrates with several durable execution engines (such as Temporal and DBOS), and pairs with Pydantic Evals and OpenTelemetry-based observability. A strong pick when your team already lives in typed Python.
Microsoft Agent Framework (and AutoGen / AG2)
Microsoft Agent Framework is the successor to both AutoGen and Semantic Kernel, built by the same teams, for .NET and Python (with Go in preview). It combines agents with graph-based workflows and enterprise features like middleware and telemetry, and supports multiple model providers. AutoGen itself is in maintenance mode, so we would not start a new project on it. AG2 is a separate community-led fork of AutoGen maintained by some of its original creators.
LlamaIndex
LlamaIndex started as a data framework for connecting LLMs to your documents and now positions itself as a framework for agents over your data, with event-driven workflows. It is a natural fit when retrieval is the heart of the agent — document research, extraction, question answering over large corpora.
Google Agent Development Kit (ADK)
Google ADK is an open-source, code-first toolkit with LLM-driven agents plus deterministic workflow agents (sequential, parallel, loop). It is optimised for Gemini and Google Cloud but can be used more broadly.
Comparison table
| Framework | Languages | Style | Model support | Stand-out strengths |
|---|---|---|---|---|
| LangGraph | Python, JS/TS | Explicit graph + state | Any | Durable execution, interrupts, memory |
| CrewAI | Python | Role-based crews + flows | Any | Fast multi-agent prototyping |
| OpenAI Agents SDK | Python, TS | Minimal: agents, handoffs, guardrails | OpenAI-first | Simplicity, built-in tracing |
| Claude Agent SDK | Python, TS | Claude Code’s agent loop as a library | Claude | Built-in file/shell tools, permissions, hooks, subagents |
| Pydantic AI | Python | Typed agents | Any | Type safety, structured output, evals |
| Microsoft Agent Framework | .NET, Python (Go preview) | Agents + graph workflows | Multiple | Enterprise features, Azure integration |
| LlamaIndex | Python, TS | Data-centric agents + workflows | Any | Retrieval and document pipelines |
| Google ADK | Python and others | Agents + workflow agents | Gemini-first | Google Cloud integration |
When plain code beats a framework
Microsoft’s own documentation puts it well: if you can write a function to handle the task, do that instead of using an agent. We would add: if your agent is a single loop with a handful of tools, a provider SDK and 50–100 lines of your own code are often enough.
def run_agent(task: str, tools: dict, max_steps: int = 10) -> str:
messages = [{"role": "user", "content": task}]
for _ in range(max_steps):
reply = llm.chat(messages=messages, tools=tool_schemas(tools))
messages.append(reply.message)
if not reply.tool_calls:
return reply.text
for call in reply.tool_calls:
result = tools[call.name](**call.arguments)
messages.append(tool_result(call.id, result))
raise RuntimeError("Agent hit step limit")
Plain code wins when:
- The workflow is mostly deterministic, with the model making one or two decisions.
- You need tight control over prompts, retries, costs and logging.
- Your team does not want to track a fast-moving dependency.
- You run in a constrained environment (for example, a PHP or Node backend where the best-known frameworks are Python-first).
A framework wins when:
- Tasks run for minutes or hours and must survive restarts.
- Humans need to approve steps and the agent must pause and resume.
- You orchestrate several agents or complex branching.
- You want tracing, evaluation and replay without building them yourself.
Many of the most reliable “agents” in production are really workflows with a few LLM steps. Our article on agentic AI design patterns covers these shapes — routing, orchestrator-workers, evaluator-optimiser — independent of any framework.
How we choose
- Start from the workflow, not the framework. Sketch the steps, decisions, tools and approval points on one page.
- Match your stack. Typed Python team? Pydantic AI or LangGraph. .NET shop on Azure? Microsoft Agent Framework. Coding or file-heavy agents on Claude? Claude Agent SDK.
- Check model portability. If you may switch providers for cost or privacy, avoid tying orchestration to one vendor.
- Demand observability. Whatever you choose, you must be able to see every model call, tool call and token count.
- Prototype the hardest step in two candidates for a day each, then decide.
- Keep business logic outside the framework so you can swap it later.
Key takeaways
- The agent loop is simple; frameworks add durability, approvals, orchestration and tracing.
- As of September 2026, LangGraph, CrewAI, OpenAI Agents SDK, Claude Agent SDK, Pydantic AI, Microsoft Agent Framework, LlamaIndex and Google ADK are all actively maintained options.
- AutoGen is in maintenance mode; Microsoft Agent Framework is its official successor.
- For simple, mostly deterministic tasks, plain code with a step limit is often the better engineering choice.
- Always set step limits, budgets and permission checks, whatever you use.
Building your agent
If you are deciding how to build an agent — or untangling one that grew too complex — we can help you pick the lightest approach that meets your reliability needs. See our AI automation service or contact us with a description of the workflow you want to automate.