What Happened
Meta and Microsoft are reducing employees’ use of Anthropic’s Claude while promoting their own AI tools. Meta’s Claude Code user count reportedly fell from about 60,000 to 30,000 [1][3]. Cohere launched North 2 with cross-session memory, revised agent orchestration, cloud and on-premises deployment, and administrator token-spending caps [24][28].
OpenAI made text watermarking available as an opt-in API setting and plans to add it for ChatGPT and Codex users in the EU [14]. Separate testing found watermark detection fell from as high as 95% to 17% when a quarter of the words were replaced [9]. Meanwhile, a survey found that only 11% of 396 businesses could forecast AI spending [32].
Why It Matters to Businesses
These stories point to three linked procurement questions: Can we switch providers, predict operating costs and govern what agents do? A survey of 300 executives found that agent projects reach production at an average rate of 34%; fragmented data was the most-cited obstacle to expanding agents’ knowledge access [18]. Provider consolidation may simplify purchasing, but tighter coupling to one model or platform can make later changes expensive [1][3].
Oversight also needs more than an approval button. Researchers warn that overloaded reviewers can end up rubber-stamping agent decisions [17]. That concern becomes more consequential as AI moves into regulated actions: a Utah pilot plans to let AI diagnose acne and prescribe medication without direct human oversight [15].
Kimbodo Engineering Perspective
We would treat model choice as a replaceable component and the agent’s data access, permissions and audit trail as the durable system. Cross-session memory can improve continuity, but it also creates retention and access-control obligations [24]. Token caps are useful guardrails, not complete budgets: cheaper-priced models cost more on 32% of over 6,800 tasks in a Microsoft analysis [32]. Measure cost per successfully completed task, including retries and review.
How We Would Implement It
- Build a controlled knowledge layer: expose approved records through permission-aware retrieval and APIs; test answers against source documents before granting broader access [18].
- Use a model gateway: route by task, log model and prompt versions, enforce structured outputs, and attribute tokens and retries to each workflow [32][37].
- Separate recommendation from action: give agents narrowly scoped tools, require explicit authorization for consequential operations, and present reviewers with evidence and alternatives—not just an “approve” button [17].
- Evaluate before and after release: combine deterministic checks with judged assessments and traced production cases to catch regressions in multistep workflows [36].
Risks, Costs and Security
Budget for retrieval infrastructure, evaluation, monitoring and human review as well as inference. Limit persistent memory by purpose and retention period; enforce tenant-level permissions at every retrieval and action. Do not treat text watermarking as reliable proof of origin after editing, given the reported drop in detection [9]. For regulated or safety-critical decisions, establish an accountable escalation path and validate the workflow against applicable requirements before allowing autonomous action [15][17].
Where Kimbodo Comes In
Kimbodo builds and operates this in production for businesses — see our AI Consulting & Strategy practice, or Request an AI Roadmap.
Sources
- [1] Meta and Microsoft pull back from Claude as Anthropic transforms from partner into competitor
- [3] Sources: Meta and Microsoft are working to cut their employees' use of Claude; Meta employees using Claude Code have dropped to ~30K from ~60K earlier this year (The Information)
- [9] OpenAI will watermark ChatGPT text in the EU but makes it optional for API users worldwide
- [14] To comply with the EU AI Act, OpenAI plans to add text watermarking for ChatGPT and Codex users in the EU and an opt-in setting for API customers globally (OpenAI)
- [15] In a first-of-its-kind pilot in the US, Nolla Health will use AI to diagnose and prescribe acne medications to Utah patients without direct human oversight (Annika Inampudi/Bloomberg)
- [17] Attempts to Keep Humans in the AI Loop May Actually Push Them Out
- [18] Connecting AI agents to enterprise knowledge
- [24] Cohere launches North 2, an update to its enterprise agent platform with cross-session memory and a redesigned harness, available across cloud and on-premises (Sean Michael Kerner/VentureBeat)
- [28] Cohere unveils North 2 AI agent platform with rebuilt orchestration and token spending caps
- [32] Survey: only 11% of 396 businesses could forecast AI spending; Microsoft finds lower-priced models cost more than higher-priced ones on 32% of 6,800+ tasks (Wall Street Journal)
- [36] Presentation: Building Reusable Evaluation Frameworks for Agentic AI Products
- [37] Article: The Platform Engineering Playbook for Production LLMs