Optimizing agent cost and token usage

Updated: 2026-09-28Reading time: 5 min

Identify where tokens go, right-size model selection for each step, enable prompt caching, and set budget limits to keep your agent spending under control.

Understanding where tokens go

Token costs in an agent accumulate from three sources: the system prompt (sent with every LLM call), the human message (varies by run), and the conversation history (grows over a multi-turn session). Use the Cost Dashboard (Metrics > Cost) to see a per-node and per-run breakdown. Often a single LLM node with a verbose system prompt accounts for the majority of costs. Start your optimization effort by identifying your most expensive nodes and measuring how many tokens each prompt component contributes. Even small reductions in the system prompt — removing redundant instructions or unnecessary examples — compound significantly at scale.

Right-sizing model selection

One of the highest-leverage cost optimizations is using the smallest model that reliably completes the task. Evaluate your agent's LLM nodes individually: a node that classifies incoming requests into five categories may perform identically on a smaller, cheaper model compared to a large frontier model. Run the same test suite against both model tiers and compare accuracy. If the cheaper model achieves acceptable quality, switch it. For agents with many runs per day, even a 10x price difference between model tiers translates to significant savings. Reserve the most capable (and expensive) models for steps that genuinely require deep reasoning.

Prompt caching and output reuse

If your system prompt is long and mostly static, take advantage of Cotonity's prompt caching feature. Enable Prompt Caching on the LLM node, and Cotonity will cache the tokenized representation of your system prompt server-side. Subsequent calls that use the same system prompt bypass the tokenization step for that portion, reducing both cost and latency. Prompt caching is especially effective for agents where the system prompt contains a large document (a knowledge base article, a product catalog) that rarely changes. Additionally, use the Memory Store to cache the results of expensive LLM calls and return the cached result for identical inputs within a configurable TTL.

Monitoring and budgeting

Set a daily or monthly token budget for each agent in Settings > Budget Limits. When the budget is exceeded, Cotonity can either pause the agent (safe for non-critical workflows) or send an alert while allowing it to continue (preferred for production-critical agents). Review the Cost Dashboard weekly to spot cost spikes early — a sudden jump in token usage often indicates a prompt that has been accidentally extended, a loop that is iterating more times than expected, or an increase in run volume driven by a noisy trigger. Combine budget alerts with the metric alerts described in the monitoring article for comprehensive cost governance.