18 August 2026

AI evaluation tools shift focus from model to system performance

  • New evaluation plugins and platforms now track how AI agents perform on real tasks across millions of sessions, measuring routing decisions and cost per task.
  • The field is moving away from testing individual AI models in isolation toward measuring complete agent systems that break down problems and route them to different tools.
  • Tools like eval-skills plugins and Agent Arena enable engineers to find errors, group similar failures, and understand expenses across large real-world deployments.

How it was covered

Latent Spaceswyx & Alessio

New tools like Hamel Husain's eval-skills plugin and Agent Arena add harness-level workflows for error discovery, clustering, and cost-per-task tracking across 1.7M+ real-world sessions. The newsletter frames this as a field-wide shift from model-level evals toward measuring routing, decomposition, and completion cost.