Fewer tokens per tool call can cost more overall

GitHub describes a counterintuitive optimization failure: shorter tool outputs can lead agents to reread information and make more calls. Measure cost over a completed task, alongside task quality, rather than optimizing each individual response in isolation.

September 28, 2026 · Amit Naik

Serving agents means serving repeated context

The vLLM team’s AgentX article explores serving infrastructure for multi-round agent workloads. Context reuse, parallelism, and the split between prefill and decoding become important when a task repeatedly sends overlapping context back to the model.

September 28, 2026 · Amit Naik