Token compression is not a universal discount. Its value depends on where tokens accumulate and whether reducing them preserves the information a model needs.
The clearest gains appeared in workflows that repeatedly sent long instructions, retrieved documents, or unchanged context to the model.
Where compression worked
Research, support, and document workflows benefited when stable context could be reduced before each request. The savings compounded across high-volume, repeated runs.
Where compression changed little
Short prompts, image-heavy requests, and workflows dominated by long model outputs saw less benefit. Routing to a cheaper sufficient model often mattered more in those cases.
Measure the request, not the average
The useful unit of analysis was the individual workflow. Request-level measurement showed where compression improved economics without weakening result quality.
