Skip to content Skip to footer

How to Choose Agent Frameworks for Production: Prioritize Evaluation, Deployment and Recovery

What Happened

Recent releases emphasize the tooling around agents, not just agent orchestration. CrewAI 1.15.24 lets teams run evaluations with alternative models and role-level model overlays; its evaluation command now returns a failing exit code when a gate does not pass. It also adds background reply queuing and experimental turn, reply and job-lifecycle tooling. [2]

LangGraph CLI 0.4.33 can deploy an already-pushed image with –image-uri, list deployment listeners, and reject Git dependencies that contain credentials. [3] Claude Code 2.1.293 focuses heavily on operational fixes across context compaction, MCP memory, background tasks, permissions, subagents and plugins. Its release notes also identify a remaining failure mode: after a container restart, a session can lose a pending scheduled wakeup and remain asleep without notifying Claude. [1]

These releases do not establish new capabilities for LangChain, LlamaIndex, AutoGen, PydanticAI, DSPy, Semantic Kernel or the OpenAI Agents SDK. Evaluate those tools against the same production requirements rather than assuming feature parity.

Why It Matters to Businesses

An agent that completes a demo can still fail in production when a model change degrades results, a background task stalls, or a deployment includes unsafe credentials. The useful pattern here is gated evaluation, reproducible deployment and observable job state: each addresses a failure that prompt quality alone cannot fix. CrewAI’s evaluation exit code makes a quality threshold enforceable in CI; LangGraph’s image option supports promotion of the same artifact between environments. [2][3]

Kimbodo Engineering Perspective

We would choose among LangChain and LangGraph, LlamaIndex, AutoGen, CrewAI, PydanticAI, DSPy, Semantic Kernel and the OpenAI Agents SDK based on the application’s workflow and operating constraints—not the number of agent abstractions. Retrieval-heavy systems need strong data access and citation checks; long-running workflows need durable state, explicit transitions and recovery. Claude Code is useful for development work, but its session behavior should not be mistaken for a guaranteed production job scheduler. [1]

How We Would Implement It

  • Define the agent’s tools, data permissions, human-approval points and measurable success criteria before selecting a framework.
  • Build a versioned evaluation set covering successful tasks, tool failures, unsafe requests and model substitutions. Run it in CI and block releases on failed gates; CrewAI’s updated evaluation command supports this pattern. [2]
  • Package and scan an immutable container image, then promote that image through environments. Where LangGraph is the runtime, deploy the pre-pushed image and inspect listeners as part of release verification. [3]
  • Store job state outside an interactive session, use idempotent tool operations, and alert on missed wakeups, stalled jobs and repeated retries. Claude Code’s documented wakeup failure illustrates why this matters. [1]

Risks, Costs and Security

Long contexts and background agents increase token spend and the amount of sensitive material exposed to tools and logs. Budget per task, limit retained context, redact traces and use short-lived credentials. Keep secrets out of dependency URLs—LangGraph CLI now rejects Git dependencies containing credentials—and scan both dependencies and deployed images. [3] Treat experimental lifecycle features as candidates for testing, not as guarantees of durable execution. [2]

Where Kimbodo Comes In

Kimbodo builds and operates this in production for businesses — see our Enterprise AI Agent Development practice, or Scope an Enterprise AI Agent.

Sources

  1. [1] v2.1.293
  2. [2] 1.15.24
  3. [3] langgraph-cli==0.4.33

Leave a comment

0.0/5