Melissa Pan compares seven models across Claude Code, Codex, and Pi harnesses. Her reported results suggest similar task success can hide meaningful cost differences. Evaluating a model in isolation misses the effects of the surrounding coding workflow.

The harness can change the cost of success — image from the original post

Source: @melissapan. This note summarizes the linked material; images belong to their respective creators.

Original tweet