Lilith.
⌕
Editorial illustration: Asana cut browser agent cost 76x by fixing history, not shrinking the model
Lilith illustration · editorial remix

Asana tested a browser agent across four models and found that its most expensive problem was not reasoning itself. The agent kept editing earlier history, repeatedly invalidating its own cache.

Stable history brought a run down to $0.47

The original agent removed screenshots and trimmed older text as it worked. Each edit changed the shared prompt prefix, forcing the provider to charge full input price for much of the history again.

Asana cached the growing history, raised its budget from 120,000 to 480,000 characters and pruned screenshots in batches. In the best setup, GPT-6.1 Sol read 89% of its input from cache. A run averaged $0.47, costing 76 times less and finishing five times faster than the original production setup on the anonymized Model B.

Instrumentation mattered more than swapping models

The study tested six history policies, two budgets and three repetitions per condition. That produced 144 runs plus a 12-run follow-up, with every agent asked to collect 192 facts.

For engineering teams, the useful lesson is that agent cost also lives in the application layer. A stable prefix, a model-appropriate history budget and cache-read telemetry can change operating economics without another round of fine-tuning or a retreat to a weaker model.

The 76x headline combines optimization with a model change

The 76x comparison pits the original Model B production setup against the optimized GPT-6.1 Sol setup. Holding Sol constant, the new caching policy produced a fourfold reduction, from $1.97 to $0.47. The authors also caution that three or four runs per condition establish broad patterns, not small percentage differences.

A larger history is not unlimited insurance. If an agent drifts or cache reuse breaks, the bill rises again, so caps on steps, tokens and per-run cost remain necessary.

Longer and messier tasks will decide whether the result travels

Asana says the changes have shipped in StackAI browser navigation. The next useful evidence will come from tasks that exceed the tested 480,000-character budget, traverse changing pages or run much longer than four minutes.

Teams should track completed-task cost, cache-read share, step count and answer quality together. That set of measures will show whether the optimization survives beyond one carefully instrumented book-catalog task.

Lilith's verdict

Asana found a $36 bill where an agent kept rewriting its own notebook. Sometimes the most expensive model problem is an application pressing Delete.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗