Insights
Evaluation harnesses, agents that survive contact with real users, and the engineering that separates the 5% of AI projects that ship from the 95% that don't. Written by practitioners, not marketers.
Four layers — dataset, scorers, harness, gates. Which metrics matter, how LLM-as-judge lies to you, and an honest map of the tool landscape with every vendor's incentive declared. Written by a team that sells none of the tools.
The 95% problemMIT's Project NANDA report is the most-cited number in enterprise AI. Here's what it actually says, what it gets right, where the "learning gap" diagnosis stops short — and the measurement discipline we see in every project that crosses the divide.
