Free. No signup. No wallet.
The Skim Audit
Send us your agent usage dashboard. We send back an executive-level estimate of how much of your LLM bill is deterministic work billed at model rates — and what fixing that is worth.
Subject line "Skim Audit". A screenshot is enough.
How it works
- 1
Send your dashboard
Email a screenshot of your agent usage or spend dashboard to hello@skim402.com with the subject "Skim Audit". Any platform, any format — if it shows token or dollar spend, it works. A usage export (CSV, JSON) is even better, but a screenshot is enough.
- 2
We run the numbers
We estimate how much of your token spend is deterministic work billed at model rates — raw web reading, repeated context, parsing, data crunching — and what routing each slice to flat-priced infrastructure would save. Every assumption is stated. Every range is honest.
- 3
You get the analysis
Within two business days you get a short, executive-level write-up: the savings estimate for clean reading alone, the estimate for a full deterministic layer, and which change pays back first. If your spend is already lean, we say that instead.
What we usually find
The numbers below come from real analyses. Yours will differ — that is the point of running the audit on your data.
~99%
of a typical agent fleet's spend is LLM tokens
From a real (anonymized) startup dashboard we analyzed: $803 of weekly spend, of which compute was under a penny and storage was zero. The entire cost story is tokens.
~4x
raw HTML is about four times larger than the clean content inside it
Our measured median across real pages. An agent reading the web raw pays token rates for navigation bars, cookie banners, and scripts — roughly three-quarters of every web-reading token is boilerplate.
~20%
typical cut to the total LLM bill from clean reading alone
Assuming a conservative 25-35% of input tokens are raw web content for assistant- and research-style agents. The reads that produce the saving cost $0.002 each — a savings-to-cost ratio around 100 to one.
45-60%
honest combined range for a five-layer deterministic stack
Clean reading plus prompt caching, a search API, a retrieval store, and a code sandbox. Non-competing layers, each attacking a different slice of token waste. The savings compound rather than simply add.
The fine print, in plain language
- It is genuinely free. No signup, no wallet, no sales call unless you ask for one. The analysis is the product demonstration — if the routing advice is good, you already know what working with us is like.
- Your numbers stay private. We never publish, share, or reference your data without your written permission. Anonymized aggregate patterns (like the ones above) are the most that ever appears publicly.
- It is an estimate, not an invoice. We state every assumption and give ranges, not false precision. A screenshot supports an executive estimate; a usage export supports a sharper one.
- We recommend things we do not sell. Skim is a clean reader. If your biggest saving is prompt caching or a retrieval store, the audit says so — those are not our products.
Point at the elephant
Read the story behind the audit — a real startup dashboard, a 55-minute agent session that cost $5.80, and the token spend nobody is looking at.