What Happened
Recent infrastructure announcements point to a broader shift: inference, agent execution, retrieval and evaluation are moving into managed cloud services. That reduces infrastructure work, but leaves businesses responsible for workflow reliability, access control and spending.
- Agent execution is moving off the laptop. Anthropic’s redesigned Cowork runs both inference and a separate per-session sandbox in the cloud. Its desktop app handles requests for local files. The previous architecture used a local VM, consuming device resources and stopping work when the laptop closed. [1]
- Managed model access is expanding. Z.ai’s 753B-parameter mixture-of-experts GLM 5.3 is available to eligible enterprise customers through Amazon Bedrock, with multiple APIs, service tiers and cross-Region inference profiles. Its reported coding gains are vendor benchmark results, not evidence of equivalent gains on enterprise workloads. [2]
- Deployment tooling is becoming agent-assisted. AWS’s SageMaker inference skill can benchmark endpoints, rank deployment configurations and generate reviewable SDK code. AgentCore examples add cross-account resource promotion and automated evaluation to the platform lifecycle. [4][5][8]
The architectural question is no longer just which model to deploy. It is which execution, retrieval and governance responsibilities to delegate—and which controls must remain explicit.
Why It Matters to Businesses
Cloud execution removes device dependence, not operational responsibility. A workflow can continue after a user disconnects, but production systems still need cancellation, durable state, retry policies and safe handling of partial tool execution. Moving tools into the cloud also changes where business data is processed and how local files cross the trust boundary. [1]
Endpoint selection can determine which controls are available. The GovCloud guidance distinguishes bedrock-runtime, which supports Guardrails and invocation logging, from bedrock-mantle, which supports the native Anthropic Messages API and server-side tools but lacks those runtime controls. API convenience is therefore a governance decision, not merely a developer preference. Model-level authorization and suitability must be checked against the organization’s obligations. [3]
More capable workflows can multiply cost. Agentic retrieval plans follow-up searches rather than relying on one similarity query. That can improve coverage of multi-part questions, but adds latency and spending. In the documented example, increasing standard retrieval from five to ten chunks also covered all six question intents, although with redundancy. Agentic retrieval is not automatically the best answer. [6]
Kimbodo Engineering Perspective
Separate model selection from execution architecture
We would treat managed inference and dedicated SageMaker endpoints as workload-specific options, not competing platform ideologies. Managed access is a practical starting point when demand is uncertain or model flexibility matters. Dedicated endpoints warrant investigation when sustained traffic, deployment control or measured performance can justify their operating burden.
Benchmark the complete configuration. AWS’s example found Qwen3-8B delivered higher throughput and lower request latency than Qwen3-1.7B, but slower time-to-first-token. The deployments used four A10G GPUs versus one L4 GPU, so the result does not isolate model size as the cause. [4]
Use the simplest workflow that meets the acceptance criteria
Start with bounded tool calls and standard retrieval. Add planning loops or specialist agents only when evaluations show a meaningful improvement. Bedrock’s retrieval guidance specifically favors standard Retrieve when lower latency, typed relevance scores or guardrail MASK behavior matters. [6]
Evaluate business correctness separately from answer quality. A supply-chain assistant can explain an infeasible recommendation convincingly. AgentCore’s example therefore checks inventory bounds, budgets, warehouse capacity, route feasibility and SQL correctness alongside helpfulness and explainability. [8]
How We Would Implement It
- Define workload contracts. Establish task-success criteria, latency targets, permitted data regions, maximum workflow duration and a cost budget per completed task.
- Build a controlled inference gateway. Pin model versions, allowlist approved endpoints and regions, and select service tiers deliberately. For Claude Code, explicitly select the intended model: the documented default is Opus 5.5, which can undermine a Sonnet-based cost plan. [2][3]
- Separate orchestration from tool execution. Persist workflow state outside the sandbox. Run tools in isolated sessions with scoped, short-lived credentials, bounded network access and explicit approval for destructive actions. Make retried writes idempotent.
- Route retrieval by measured need. Use standard retrieval for narrow questions; escalate complex questions to agentic retrieval. Capture citations, planner traces, search counts and latency. Treat full-document access as a separate permission because expansion requires bedrock:GetDocumentContent. [6]
- Benchmark representative traffic. Measure throughput, time-to-first-token, tail latency and cost per successful task using realistic prompt lengths and concurrency. Review agent-generated deployment code before provisioning or load-testing production endpoints. [4]
- Gate releases with evaluations. Combine deterministic business-rule checks with quality evaluators. Run development evaluations before promotion, then sample production traces for regressions and send actionable results to monitoring. AgentCore supports both on-demand and online evaluation patterns. [8]
- Promote configuration through separate accounts. Preview changes, preserve versioned backups and verify permissions. Quick Resource Migrator relinks dependencies, but target connectors still require re-authentication and S3 knowledge-base documents are not copied. Treat promotion as incomplete until these dependencies pass validation. [5]
Risks, Costs and Security
Measure total workflow cost, not just token price. Include inference, retrieval, embeddings, storage, sandbox compute, evaluation and idle endpoints. Repeated agent steps can make a cheaper model more expensive per successful outcome. Apply caching only where eligible prefixes recur; GLM 5.3’s documented explicit cache breakpoints require at least 1,024 tokens per marked prefix. [2][6]
Monitoring is not a hard spending limit. The GovCloud guidance describes invocation logging, CloudWatch, Lambda and DynamoDB for per-developer usage tracking and alerts. Enforce strict budgets before dispatch where required, accounting for concurrent requests and delayed usage reporting. [3]
Managed inference does not secure the whole application. Developer machines, tool credentials, retrieved documents and execution environments remain separate attack surfaces. Enforce authorization outside model reasoning, redact sensitive traces and require ownership or explicit written permission before automated security testing. [2][3]
Cleanup and access reviews are production controls. Remove idle SageMaker endpoints and Studio spaces, unnecessary benchmark artifacts and temporary credentials. Review Quick roles and transfer asset ownership before deleting or downgrading users. These practices reduce both recurring charges and unnecessary privileges. [4][7]
The production lesson is straightforward: adopt managed services to reduce infrastructure work, but retain explicit control over workflow state, permissions, evaluation and cost per successful business outcome.
Where Kimbodo Comes In
Kimbodo builds and operates this in production for businesses — see our AI Infrastructure & MLOps practice, or Estimate My Infrastructure.
Sources
- [1] Quoting Felix Rieseberg
- [2] Introducing GLM 5.3 on Amazon Bedrock
- [3] Supercharge regulated workloads with Claude Code and Amazon Bedrock
- [4] New agent skill: Amazon SageMaker optimized generative AI inference for your coding agent
- [5] Making Amazon Quick enterprise-ready: Automated, auditable cross-account resource promotion
- [6] Agentic retrieval with LangChain and Amazon Bedrock Knowledge Bases
- [7] Downgrading user roles in Amazon Quick
- [8] Evaluating multi-agent systems for explainability and helpfulness with Amazon Bedrock AgentCore