What Happened
Several recent papers point to the same engineering lesson: improving an AI model is not the same as improving the decisions or actions of an application.
- Agent failures can become useful training data. The Agent Error Dataset contains 50,228 error–diagnosis pairs with execution traces. In matched replays, proposed corrections raised verifier pass rates from 18.4% to 51.1%. That measures replayed corrections, not production reliability [2].
- The agent harness matters. MILO searches over agent harnesses and reported higher benchmark scores while using fewer tokens than its starting harness [4]. But adding more procedure is not automatically helpful: profession-specific scientific prompts showed no clear accuracy gain and cost 2.2–4.5 times more per successful call [35].
- Decision-relevant errors matter more than aggregate error. Across 55 world models, total prediction error weakly tracked planning success, while error on decision-critical state dimensions tracked it strongly. Models with nearly identical total error achieved 97% and 37% planning success [7].
- Evidence needs an independent check. EviGraph separates an agent’s evidence-gathering from a deterministic recommendation checker [26]. In a separate study, corrupted rationales changed verifier support judgments even when the underlying evidence and candidate answer stayed fixed [31].
- Specialized tools can outperform unaided generation. CineMR’s tool-supported cardiac MRI assessment reached 35.9% pass@1 versus 1.5% for its underlying vision-language model. The result is promising for that benchmark, not a clinical deployment claim [38].
Why It Matters to Businesses
These findings shift procurement and implementation questions away from “Which model scores highest?” toward “Which workflow produces a correct, authorized, verifiable outcome at an acceptable cost?” A model can predict well overall yet miss a rare condition that changes the required action [7]. An agent can produce a persuasive explanation that biases its verifier [31]. A larger prompt or more elaborate harness can consume more budget without improving completion rates [30][35].
There are also practical gains in narrowing the task. For building-operations questions, retrieved examples raised text-to-SPARQL exact-match accuracy from 0.2–20% zero-shot to 56–65% across three evaluated models [24]. For irregular clinical records, selecting source-linked measurements improved prediction performance over evaluated text baselines while making the evidence easier to trace [22]. Neither result makes a general-purpose assistant a substitute for domain validation.
Kimbodo Engineering Perspective
Optimize the governed workflow, not the demo. We would treat the model, prompts, retrieval, tools, permissions and verification rules as one system. The right benchmark must measure the business consequence of an error: a wrong recommendation, unauthorized tool call, missed exception or unreviewable result. The world-model findings make the case for weighting critical errors rather than averaging them away [7].
We would also resist moving safety decisions into informal agent messages. In an agent-skill attack study, a false claim of user approval passed between skills helped induce attacker-selected actions in 74.2% of attempts [37]. Small models should not inherit approval duties merely because they are cheap: none of 16 tested Qwen3 configurations met the paper’s qualification thresholds across four agent-harness microtasks [32].
How We Would Implement It
- Define outcomes and authority. Specify which decisions the agent may recommend, which actions require human approval, and which systems it may access. Keep approval in a trusted service, not in retrieved text, skill files or model-generated rationales [26][37].
- Build an evidence trail. Store source identifiers, relevant excerpts or measurements, tool inputs and outputs, proposed actions, authorization decisions and observed effects. Use deterministic checks where requirements can be encoded; allow abstention when required evidence is missing [22][26][30].
- Test on representative failures. Replay failed tasks with their traces, classify root causes, and test corrections against recorded evidence before using them for training [2]. Include rare, high-impact cases and score decision-critical mistakes separately [7].
- Compare complete configurations. Evaluate the baseline against retrieval, specialist tools, revised prompts and alternative harnesses on the same held-out tasks. Track verified completion, harmful or unauthorized actions, abstention, latency and cost per successful task—not just answer accuracy [4][35][38].
- Roll out with controls. Begin with read-only or proposal-only operation, require approval for consequential writes, monitor failure categories and regressions, and expand permissions only after measured improvement.
Risks, Costs and Security
Most cited results are research evaluations, often on particular benchmarks; they do not establish production safety. Added verification can also have a cost: one agent-harness pilot showed no strict-trial improvement despite substantially higher token use [30]. Security reviews must cover the assembled system, not only the base model. Research on latent communication found that links between safety-aligned agents could increase harmful compliance, while a separate skill-chain attack exploited untrusted progress records [41][37].
The operating budget should therefore include trace storage, evaluation sets, tool isolation, access controls, human review and periodic retesting. Those costs are justified when they prevent consequential errors—or demonstrate that a simpler workflow performs just as well.
Where Kimbodo Comes In
Kimbodo builds and operates this in production for businesses — see our AI Consulting & Strategy practice, or Request an AI Roadmap.
Sources
- [2] Agent Error Dataset: Scaling 50,000 Error–Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training
- [4] MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
- [7] Not All Errors Matter: Decision-Relevant Prediction Error Predicts Planning Quality
- [22] Learning to Select Source-Traceable Evidence for Language-Model Prediction from Irregular Clinical Time Series
- [24] Build2SPARQL: A Large-Scale Text-to-SPARQL Benchmark Dataset for Building Knowledge Graph Querying
- [26] EviGraph: Proof-Carrying Selective Recommendation over Temporal Public-Service Knowledge Graphs
- [30] From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution
- [31] What Do Rationales Communicate? A Message-Intervention Study in Role-Specialized QA
- [32] Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?
- [35] Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks
- [37] Chaining Skills to Hijack LLM Agents
- [38] CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment
- [41] Safety of Latent Communication in Multi-Agent Systems