The harness can change the cost of success

Melissa Pan compares seven models across Claude Code, Codex, and Pi harnesses. Her reported results suggest similar task success can hide meaningful cost differences. Evaluating a model in isolation misses the effects of the surrounding coding workflow.

September 28, 2026 · Amit Naik

An agent evaluation needs more than a prompt

Alex Lieberman’s notes on agent evaluation emphasize the combination of the agent, its system, tasks, and a verifier. Rubrics and reproducible environments matter because a plausible answer is not the same thing as a successfully completed task.

September 28, 2026 · Amit Naik

Evaluate documentation with reader agents

LangChain’s OpenWiki article introduces WikiBench, evaluating generated repository knowledge through questions answered by reader agents. The interesting shift is to test whether documentation helps a downstream user solve a problem, rather than judging the document only by its appearance.

September 28, 2026 · Amit Naik