AI Solutions

AI Agent Frameworks Compared (2026): When to Use One vs Plain Code

AI agent frameworks compared as of September 2026: LangGraph, CrewAI, OpenAI Agents SDK, Claude Agent SDK, Pydantic AI, Microsoft Agent Framework and more.

GPTLabAI team 6 min read

Choosing between AI agent frameworks matters less than most comparison posts suggest. The core agent loop — call a model, run the tools it asks for, feed the results back, repeat — is a few dozen lines of code. A framework earns its place when you need what surrounds that loop: durable state, human approval steps, multi-agent orchestration, tracing and evaluation. This guide compares the main options as of September 2026 and explains when we reach for one versus writing plain code.

What an agent framework actually gives you

Every framework below wraps the same basic idea. What differs is the extra machinery:

  • Tool calling with schemas generated from functions or types.
  • State and memory across steps and sessions.
  • Durable execution — resuming a long task after a crash or a deploy.
  • Human-in-the-loop — pausing for approval before risky actions.
  • Orchestration — graphs, handoffs or teams of agents.
  • Observability — traces of every model call and tool call.
  • MCP support — connecting to tools exposed via the Model Context Protocol.

If you need few of these, plain code is often simpler.

The main AI agent frameworks compared

Framework features and versions change quickly; the descriptions below are as of September 2026 and link to official documentation.

LangGraph

LangGraph describes itself as a low-level orchestration framework and runtime for long-running, stateful agents. You model your agent as a graph of nodes and edges with explicit state. Its strengths are durable execution (resume where you left off), human-in-the-loop interrupts, and short- and long-term memory. It reached a stable 1.0 release in October 2025 and does not require the rest of LangChain. It suits teams that want explicit control over every step and production features like checkpoints.

CrewAI

CrewAI is a standalone Python framework built around agents with roles, crews of agents that collaborate, and flows for event-driven orchestration with state. It is quick to prototype with, especially for role-based multi-agent setups (“researcher”, “writer”, “reviewer”). For production we tend to lean on its flows for predictable control rather than letting crews decide everything.

OpenAI Agents SDK

The OpenAI Agents SDK (Python, with a TypeScript version) is deliberately small: agents, handoffs between agents, guardrails for validating inputs and outputs, sessions for memory, and built-in tracing. It is the natural choice if you are mostly on OpenAI models and want minimal abstraction. Other providers can be used, but it is designed around OpenAI’s APIs.

Claude Agent SDK

The Claude Agent SDK (Python and TypeScript) exposes the same agent loop, built-in tools, context management, subagents, hooks, permissions, sessions and MCP support that power Claude Code. It shines for agents that work with files, run commands and operate on codebases or document workspaces, and it gives you permission controls out of the box. It is tied to Claude models.

Pydantic AI

Pydantic AI brings a type-safe, “FastAPI-like” feel to agents in Python. It is model-agnostic, validates structured outputs with Pydantic models, supports typed dependency injection and MCP, integrates with several durable execution engines (such as Temporal and DBOS), and pairs with Pydantic Evals and OpenTelemetry-based observability. A strong pick when your team already lives in typed Python.

Microsoft Agent Framework (and AutoGen / AG2)

Microsoft Agent Framework is the successor to both AutoGen and Semantic Kernel, built by the same teams, for .NET and Python (with Go in preview). It combines agents with graph-based workflows and enterprise features like middleware and telemetry, and supports multiple model providers. AutoGen itself is in maintenance mode, so we would not start a new project on it. AG2 is a separate community-led fork of AutoGen maintained by some of its original creators.

LlamaIndex

LlamaIndex started as a data framework for connecting LLMs to your documents and now positions itself as a framework for agents over your data, with event-driven workflows. It is a natural fit when retrieval is the heart of the agent — document research, extraction, question answering over large corpora.

Google Agent Development Kit (ADK)

Google ADK is an open-source, code-first toolkit with LLM-driven agents plus deterministic workflow agents (sequential, parallel, loop). It is optimised for Gemini and Google Cloud but can be used more broadly.

Comparison table

Framework Languages Style Model support Stand-out strengths
LangGraph Python, JS/TS Explicit graph + state Any Durable execution, interrupts, memory
CrewAI Python Role-based crews + flows Any Fast multi-agent prototyping
OpenAI Agents SDK Python, TS Minimal: agents, handoffs, guardrails OpenAI-first Simplicity, built-in tracing
Claude Agent SDK Python, TS Claude Code’s agent loop as a library Claude Built-in file/shell tools, permissions, hooks, subagents
Pydantic AI Python Typed agents Any Type safety, structured output, evals
Microsoft Agent Framework .NET, Python (Go preview) Agents + graph workflows Multiple Enterprise features, Azure integration
LlamaIndex Python, TS Data-centric agents + workflows Any Retrieval and document pipelines
Google ADK Python and others Agents + workflow agents Gemini-first Google Cloud integration

When plain code beats a framework

Microsoft’s own documentation puts it well: if you can write a function to handle the task, do that instead of using an agent. We would add: if your agent is a single loop with a handful of tools, a provider SDK and 50–100 lines of your own code are often enough.

def run_agent(task: str, tools: dict, max_steps: int = 10) -> str:
    messages = [{"role": "user", "content": task}]
    for _ in range(max_steps):
        reply = llm.chat(messages=messages, tools=tool_schemas(tools))
        messages.append(reply.message)
        if not reply.tool_calls:
            return reply.text
        for call in reply.tool_calls:
            result = tools[call.name](**call.arguments)
            messages.append(tool_result(call.id, result))
    raise RuntimeError("Agent hit step limit")

Plain code wins when:

  • The workflow is mostly deterministic, with the model making one or two decisions.
  • You need tight control over prompts, retries, costs and logging.
  • Your team does not want to track a fast-moving dependency.
  • You run in a constrained environment (for example, a PHP or Node backend where the best-known frameworks are Python-first).

A framework wins when:

  • Tasks run for minutes or hours and must survive restarts.
  • Humans need to approve steps and the agent must pause and resume.
  • You orchestrate several agents or complex branching.
  • You want tracing, evaluation and replay without building them yourself.

Many of the most reliable “agents” in production are really workflows with a few LLM steps. Our article on agentic AI design patterns covers these shapes — routing, orchestrator-workers, evaluator-optimiser — independent of any framework.

How we choose

  1. Start from the workflow, not the framework. Sketch the steps, decisions, tools and approval points on one page.
  2. Match your stack. Typed Python team? Pydantic AI or LangGraph. .NET shop on Azure? Microsoft Agent Framework. Coding or file-heavy agents on Claude? Claude Agent SDK.
  3. Check model portability. If you may switch providers for cost or privacy, avoid tying orchestration to one vendor.
  4. Demand observability. Whatever you choose, you must be able to see every model call, tool call and token count.
  5. Prototype the hardest step in two candidates for a day each, then decide.
  6. Keep business logic outside the framework so you can swap it later.

Key takeaways

  • The agent loop is simple; frameworks add durability, approvals, orchestration and tracing.
  • As of September 2026, LangGraph, CrewAI, OpenAI Agents SDK, Claude Agent SDK, Pydantic AI, Microsoft Agent Framework, LlamaIndex and Google ADK are all actively maintained options.
  • AutoGen is in maintenance mode; Microsoft Agent Framework is its official successor.
  • For simple, mostly deterministic tasks, plain code with a step limit is often the better engineering choice.
  • Always set step limits, budgets and permission checks, whatever you use.

Building your agent

If you are deciding how to build an agent — or untangling one that grew too complex — we can help you pick the lightest approach that meets your reliability needs. See our AI automation service or contact us with a description of the workflow you want to automate.

7 min

How Much Does an AI Chatbot for Business Cost?

What drives AI chatbot cost for a business: build, model API usage, hosting and maintenance, build vs buy, plus a simple formula to estimate your token costs.

Read article

Have a project in mind? Let’s talk.

Whether you run a business or a research group, tell us what you need built, fixed or evaluated. You get a free consultation and a clear written estimate — no obligation.

  • Free consultation
  • Written scope and estimate
  • We reply within one working day
Contact us