Insights

Field notes from the production side of AI.

Evaluation harnesses, agents that survive contact with real users, and the engineering that separates the 5% of AI projects that ship from the 95% that don't. Written by practitioners, not marketers.

Evaluation

What is an LLM evaluation framework? A vendor-neutral guide

Four layers — dataset, scorers, harness, gates. Which metrics matter, how LLM-as-judge lies to you, and an honest map of the tool landscape with every vendor's incentive declared. Written by a team that sells none of the tools.

Atlas Research · July 2026 · 11 min read
The 95% problem

The GenAI Divide, annotated: why 95% of AI pilots die — and the pattern in the 5% that live

MIT's Project NANDA report is the most-cited number in enterprise AI. Here's what it actually says, what it gets right, where the "learning gap" diagnosis stops short — and the measurement discipline we see in every project that crosses the divide.

Atlas Research · July 2026 · 9 min read

Let’s put a number on it.

Book a free 15-min consult