<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Inference Economics on Amit Naik</title>
    <link>https://amitnaik.com/tags/inference-economics/</link>
    <description>Recent content in Inference Economics on Amit Naik</description>
    <generator>Hugo -- 0.144.2</generator>
    <language>en-us</language>
    <lastBuildDate>Mon, 28 Sep 2026 08:59:40 -0700</lastBuildDate>
    <atom:link href="https://amitnaik.com/tags/inference-economics/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Fewer tokens per tool call can cost more overall</title>
      <link>https://amitnaik.com/til/2026-09-28-fewer-tokens-per-tool-call-can-cost-more-overall/</link>
      <pubDate>Mon, 28 Sep 2026 08:59:40 -0700</pubDate>
      <guid>https://amitnaik.com/til/2026-09-28-fewer-tokens-per-tool-call-can-cost-more-overall/</guid>
      <description>GitHub describes a counterintuitive optimization failure: shorter tool outputs can lead agents to reread information and make more calls. Measure cost over a completed task, alongside task quality, rather than optimizing each individual response in isolation.</description>
    </item>
    <item>
      <title>Serving agents means serving repeated context</title>
      <link>https://amitnaik.com/til/2026-09-28-serving-agents-means-serving-repeated-context/</link>
      <pubDate>Mon, 28 Sep 2026 08:59:39 -0700</pubDate>
      <guid>https://amitnaik.com/til/2026-09-28-serving-agents-means-serving-repeated-context/</guid>
      <description>The vLLM team&amp;#39;s AgentX article explores serving infrastructure for multi-round agent workloads. Context reuse, parallelism, and the split between prefill and decoding become important when a task repeatedly sends overlapping context back to the model.</description>
    </item>
  </channel>
</rss>
