<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Agent Evaluation on Amit Naik</title>
    <link>https://amitnaik.com/tags/agent-evaluation/</link>
    <description>Recent content in Agent Evaluation on Amit Naik</description>
    <generator>Hugo -- 0.144.2</generator>
    <language>en-us</language>
    <lastBuildDate>Mon, 28 Sep 2026 08:59:56 -0700</lastBuildDate>
    <atom:link href="https://amitnaik.com/tags/agent-evaluation/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>The harness can change the cost of success</title>
      <link>https://amitnaik.com/til/2026-09-28-the-harness-can-change-the-cost-of-success/</link>
      <pubDate>Mon, 28 Sep 2026 08:59:56 -0700</pubDate>
      <guid>https://amitnaik.com/til/2026-09-28-the-harness-can-change-the-cost-of-success/</guid>
      <description>Melissa Pan compares seven models across Claude Code, Codex, and Pi harnesses. Her reported results suggest similar task success can hide meaningful cost differences. Evaluating a model in isolation misses the effects of the surrounding coding workflow.</description>
    </item>
    <item>
      <title>An agent evaluation needs more than a prompt</title>
      <link>https://amitnaik.com/til/2026-09-28-an-agent-evaluation-needs-more-than-a-prompt/</link>
      <pubDate>Mon, 28 Sep 2026 08:59:41 -0700</pubDate>
      <guid>https://amitnaik.com/til/2026-09-28-an-agent-evaluation-needs-more-than-a-prompt/</guid>
      <description>Alex Lieberman&amp;#39;s notes on agent evaluation emphasize the combination of the agent, its system, tasks, and a verifier. Rubrics and reproducible environments matter because a plausible answer is not the same thing as a successfully completed task.</description>
    </item>
    <item>
      <title>Evaluate documentation with reader agents</title>
      <link>https://amitnaik.com/til/2026-09-28-evaluate-documentation-with-reader-agents/</link>
      <pubDate>Mon, 28 Sep 2026 08:59:24 -0700</pubDate>
      <guid>https://amitnaik.com/til/2026-09-28-evaluate-documentation-with-reader-agents/</guid>
      <description>LangChain&amp;#39;s OpenWiki article introduces WikiBench, evaluating generated repository knowledge through questions answered by reader agents. The interesting shift is to test whether documentation helps a downstream user solve a problem, rather than judging the document only by its appearance.</description>
    </item>
  </channel>
</rss>
