17 August 2026
Evaluation tools shift focus from single models to full systems
- New tools like eval-skills and Agent Arena measure how AI systems actually perform in real workflows, not just how well individual models score on tests.
- These tools track practical concerns: whether systems route questions correctly, break problems into steps, remember context, and verify their own answers.
- The shift matters because a great model inside a poorly designed system produces worse results than a mediocre model in a well-built one.
How it was covered
Latent Spaceswyx & Alessio
Tools like Hamel Husain's eval-skills plugin and Agent Arena's filters are shifting evaluation from model-level evals to harness-level measurement covering routing, decomposition, memory, verifier loops, and total completion cost. The newsletter notes this represents movement toward measuring full system behavior rather than isolated model performance.