Skip to content Skip to footer

How to Scale Enterprise AI Agents Without Losing Control of Cost, Data and Workflows

What Happened

Google introduced Gemini as a unified agent and API spanning knowledge work, media creation, coding and long-running tasks. The announced architecture combines shared memory, reusable skills, tools and coordination between agents across Google Workspace and Gemini Enterprise. Google also emphasized agent identities, least-privilege access, audit trails, sandboxing, multi-model orchestration, Smart Routing and real-time spend caps. [1]

The enterprise data layer is central to that proposition: BigQuery, Knowledge Catalog, Smart Storage and Borderless Lakehouse are positioned to ground agents in company data and trusted definitions across platforms. Financial-services and legal specializations were described as previews. [1]

Deployment examples suggest a shift from isolated assistants toward operational workloads. Ryanair is rolling out Gemini Enterprise and Google Workspace to approximately 35,000 employees; Smyths Toys reports that its Codie agent resolves more than 60% of web inquiries; and Virgin Media Ireland reports deploying workloads three times faster. These are reported customer outcomes, not benchmarks that predict another organization’s results. [2]

Why It Matters to Businesses

The buying decision is moving beyond model quality to operational control. An agent that retrieves documents has a different risk profile from one that changes customer records, runs code or delegates work. Shared memory and long-running execution make authorization, state management and traceability architectural requirements rather than optional integrations.

Google’s reported examples—including Bradesco reducing document reviews from an hour to five minutes and Airwallex resolving customer issues ten times faster—illustrate potential workflow gains. They do not establish implementation cost, error rates or how much human supervision remains necessary. Business cases should measure those factors alongside speed. [1]

Enterprise leaders should evaluate three outcomes:

  • Cost per successful workflow: Include inference, retrieval, tool execution, retries and human review—not just token prices.
  • Safe completion rate: Measure whether tasks finish correctly within access policies and approval requirements.
  • Operational ownership: Establish who maintains connectors, investigates failures and can suspend an agent.

Kimbodo Engineering Perspective

Our engineering judgment is to use a managed enterprise platform where its identity, connectors and governance match the workload, while keeping business rules and execution safeguards explicit. Buying orchestration does not remove responsibility for application correctness.

Start with one bounded workflow, not a general-purpose autonomous agent. A document-review assistant or support triage workflow is easier to evaluate than an agent allowed to act across every business system. Introduce write access only after demonstrating reliable retrieval, appropriate abstention and effective escalation.

Multi-agent coordination should be justified by measurable gains. Specialized agents can separate responsibilities, but also add model calls, handoff errors and latency. A deterministic workflow with a few model-assisted steps is often easier to test and operate.

Similarly, shared memory should be treated as governed application data. Define its scope, retention, provenance and deletion behavior. Do not let an unverified model inference become an authoritative customer fact merely because it was remembered.

How We Would Implement It

1. Establish identity and data boundaries

Put an authenticated application or API gateway in front of the agent. Propagate the user’s authorization context and assign narrowly scoped service identities to execution components. Separate read-only retrieval from write-capable tools. Enforce document permissions during retrieval, rather than asking the model to respect them.

2. Build a governed retrieval layer

Use cataloged data and approved business definitions for analytics; Google’s announced data services are candidates where they fit the existing estate. [1] Return source references and freshness metadata with retrieved evidence. For numerical answers, prefer validated queries and calculated results over model-generated arithmetic.

3. Make execution durable and bounded

Use a durable workflow engine for long-running tasks, with persisted state, timeouts and bounded retries. Place approval gates before consequential actions. Give write operations idempotency keys where supported and reconciliation checks where they are not. Define compensating actions when a workflow can partially complete.

4. Route models against measured requirements

Evaluate lower-cost models on representative tasks before routing routine work to them. Escalate difficult cases using calibrated signals or validation failures—not model confidence alone. Apply per-request and per-workflow budgets for tokens, elapsed time, tool calls and delegation depth. Platform routing and spend controls can support these limits, but their exact behavior must be verified. [1]

5. Operate through evaluation and controlled releases

Version prompts, models, tools, retrieval configuration and policies together. Maintain regression tests for accuracy, permission isolation, prompt injection and tool failures. Trace each workflow across retrieval, inference and execution. Release changes through a limited cohort, with rollback and an operator kill switch.

Risks, Costs and Security

Prompt injection becomes more consequential when agents can act. Treat retrieved documents and tool responses as untrusted inputs. Validate tool arguments outside the model, restrict destinations and credentials, and sandbox code execution.

Long-running tasks, repeated context and agent delegation can make costs unpredictable. Budget controls should specify whether exhaustion causes a safe stop, human handoff or cancellation. An alert after spending has occurred is not equivalent to enforcement.

Managed platforms reduce some infrastructure work but introduce dependency on connectors, model behavior and proprietary orchestration interfaces. Preserve exportable workflow state and audit records, and isolate vendor-specific APIs behind application interfaces where practical.

Finally, verify availability, regional support, retention settings and contractual data protections before production use. Google states that proprietary inputs and outputs remain customers’ property; ownership alone does not resolve residency, processing or compliance requirements. [1] The production goal is accountable automation: every action authorized, every failure recoverable and every completed workflow economically defensible.

Where Kimbodo Comes In

Kimbodo builds and operates this in production for businesses — see our AI Infrastructure & MLOps practice, or Estimate My Infrastructure.

Sources

  1. [1] Welcome to Gemini at Work 2026: Introducing the Gemini agent
  2. [2] Innovation in Ireland: How Irish brands scale with Gemini Enterprise

Leave a comment

0.0/5