What Happened
Claude Code v2.1.287 expanded plugin customization with Claude Mods, added a side agent that flags potentially missed issues, and improved reliability across MCP connectors, SDK streaming, cloud sessions, and file delivery. It also added prompt text to an OpenTelemetry event, enabled URL prompts from compatible MCP servers, and made a 1M-token context window the default for specified models and gateways, with an option to retain 200K. These are Claude Code changes, not releases across the other frameworks named here. [1]
Why It Matters to Businesses
Agent tooling is moving beyond “model calls plus tools” toward configurable workflows with plugins, external connectors, telemetry, and long-running sessions. Those capabilities can make applications more useful, but they also expand the operational boundary: prompts may enter traces, connectors may request sign-in, and larger context windows can increase spend and the amount of sensitive material exposed to a model. Claude Code’s changes make those trade-offs concrete. [1]
Framework selection should therefore start with the application’s control needs. LangGraph and similar workflow-oriented tools suit explicit state and branching; LlamaIndex is often relevant when retrieval is central; PydanticAI and DSPy emphasize different forms of structured outputs and programmatic optimization. Multi-agent tooling such as AutoGen or CrewAI is worth considering when distinct roles improve a measurable task—not simply because multiple agents are available. These are architectural patterns, not claims about new releases.
Kimbodo Engineering Perspective
We would choose the smallest orchestration layer that makes execution understandable and testable. For a bounded assistant, an SDK and typed tool contracts may suffice. For a process with approvals, retries, or resumable work, explicit workflow state is more valuable than an open-ended agent loop. We would add specialist agents only after evaluating whether they improve accuracy enough to justify extra latency, tokens, and failure paths.
Claude Code’s reliability fixes also illustrate a production lesson: connector behavior, attachment delivery, and review retries affect outcomes as much as prompt quality does. Compatibility changes need regression tests; observability changes need a privacy review. [1]
How We Would Implement It
- Define the workflow: Specify tasks, permitted tools, success criteria, human approval points, and a maximum budget for steps, time, and tokens.
- Separate responsibilities: Put authentication and authorization at the application boundary; expose narrowly scoped tools to the agent; keep workflow state and audit records outside model context.
- Choose orchestration by need: Use direct SDK calls for simple flows, a stateful graph for recoverable multi-step work, and retrieval components when answers must be grounded in business data.
- Test the complete path: Evaluate answer quality alongside tool selection, connector failures, retries, partial file delivery, and permission denials. Pin MCP compatibility and test sign-in behavior before upgrading connectors. [1]
- Instrument safely: Record latency, cost, tool outcomes, and trace identifiers while masking or dropping prompt text under the same policy wherever it appears in telemetry. [1]
Risks, Costs and Security
Plugins and MCP servers extend the agent’s effective supply chain. Review their permissions, restrict network and data access, and treat URL-based sign-in prompts as security-sensitive interactions. The new prompt-text telemetry field requires particular attention because an existing masking rule may not automatically cover a new attribute. [1]
A 1M-token default should not become a default application design. Limit retrieved context, measure quality against smaller windows, and set per-task spend ceilings. Treat the reported reliability improvements as reasons to upgrade deliberately—not as substitutes for application-level retries, idempotency, and monitoring. [1]
Where Kimbodo Comes In
Kimbodo builds and operates this in production for businesses — see our Enterprise AI Agent Development practice, or Scope an Enterprise AI Agent.
Sources
- [1] v2.1.287