Agent loops are full of tiny, bounded decisions with a fixed set of options: which action to take on a form field, whether a step is already done, which of N tools fits. Paying for a full generative LLM call on each is wasteful โ a small discriminative model can score the fixed options instead, emitting zero tokens.
Score the Options, Don't Generate the Action: Offload Bounded Agent Decisions to a Tiny Classifier
Unlock this tip โ and 129 more
This is one of 130 advanced, fact-checked tactics reserved for Pro. Get the full 152-tip library, a searchable archive, and a new tip every morning. Free for 7 days, then $9/mo.
Prefer to browse? The 22 Beginner tips are free forever.
More in Model Selection
Stop Paying Frontier Prices for Boilerplate Work
Most of your token spend is on tasks a small model handles perfectly. Match the model to the job instead of defaulting to your most expensive option for everything.
Buy More Reasoning on the Cheap Model Before You Upgrade the Tier
When a cheap model stumbles on a hard task, the reflex is to jump to the frontier tier. Often the cheaper move is to keep the small model and turn its reasoning effort up โ its per-token rate is so low it can brute-reason through the problem and still cost far less.
Cascade: Try the Cheap Model First, Escalate Only When It Fails
Send every request to a small model first, programmatically check the answer, and only escalate to a frontier model when the cheap one falls short.