Melissa Pan compares seven models across Claude Code, Codex, and Pi harnesses. Her reported results suggest similar task success can hide meaningful cost differences. Evaluating a model in isolation misses the effects of the surrounding coding workflow.

Source: @melissapan. This note summarizes the linked material; images belong to their respective creators.