Alex Lieberman’s notes on agent evaluation emphasize the combination of the agent, its system, tasks, and a verifier. Rubrics and reproducible environments matter because a plausible answer is not the same thing as a successfully completed task.
Source: @businessbarista. This note summarizes the linked material; images belong to their respective creators.