The usual pattern waits until context nears the limit, then does one large blocking compaction that summarizes a huge span at once. Instead, shave a little already-consumed history every single turn so the running total stays flat, always preserving a guaranteed tail of recent user messages. The summarizer cost is amortized across turns rather than paid as one stall, and per-model thresholds tune how hard each backend gets trimmed.
Micro-Compact Every Turn Instead of One Big Blocking Pause
Unlock this tip โ and 118 more
This is one of 119 advanced, fact-checked tactics reserved for Pro. Get the full 141-tip library, a searchable archive, and a new tip every morning. Free for 7 days, then $9/mo.
Prefer to browse? The 22 Beginner tips are free forever.
More in Context Management
Drop the Screenshot Once the Model Has Read It
Images, PDFs, and attachments are charged as tokens and re-sent every turn in a multimodal thread. After the model has described or transcribed one, you usually don't need to keep sending the pixels.
Paste the Function, Not the Whole File
Most coding questions need 20-40 lines, not your 800-line file. Send the relevant slice plus a one-line note about the rest, and your input shrinks dramatically without hurting the answer.
Start a New Chat When the Topic Changes
Chat apps re-send your whole conversation with every message. When you switch tasks, the old turns become dead weight you keep paying to re-transmit โ even with caching discounts.