Docs
TokenShrinker is a local proxy. It runs on your machine, sits between your tools and your AI provider, and does nothing you can't verify by pointing a request at it yourself.
Getting started
- 1. Download and install TokenShrinker. It works immediately with local Ollama models; no account, no key.
- 2. Open the app and add your provider's API key in Settings, if you want to route to a cloud model. Keys are encrypted with your OS credential store and stay on your device.
- 3. Turn the proxy on from the Dashboard, then point your SDK's base URL at the local address below.
- 4. Send a request. Figures on the Dashboard fill in once the first one comes through.
The one-line integration
Point your SDK, agent framework, or editor at the local proxy. Your keys, models, and code stay as they are; anything the proxy doesn't handle is forwarded untouched.
# Any OpenAI-compatible client
OPENAI_BASE_URL="http://127.0.0.1:1337/v1"
# Anthropic SDKs append /v1 themselves
ANTHROPIC_BASE_URL="http://127.0.0.1:1337"
Works with Cursor, Claude Code, Aider, LangChain, CrewAI, and anything else speaking the OpenAI or Anthropic
API. Streaming passes through chunk by chunk with no buffering step added. To route a request at your local
Ollama models instead of a cloud provider, call /v1/ollama/...
or add the header x-tokenshrink-local: true.
To forward a single call exactly as you sent it, add
x-tokenshrink: off. Nothing is
altered on an opted-out request; it is still measured, so it still appears in your figures.
What each feature does
Agent loop breaker
Detects near-identical requests, fingerprinted locally, and blocks a retry loop at the fourth repeat within a rolling window. A blocked call is refused with an error your agent framework can catch; it never silently drops the request. Blocks expire on their own, so a false positive can't strand you.
Structural compression
Applies to JSON, logs, and stack traces, the context that changes every request and that provider-side prompt caching can't help with. JSON transformations are verified by re-parsing the result and comparing it against the source; if that comparison fails, the original is sent unchanged, never a best-effort guess.
Textual distillation
Removes conversational filler, repetition, and low-value phrasing from prose. Three levels (Safe, Balanced, Aggressive) trade more reduction for more risk of changing what the model produces. See Limitations below before you rely on it for anything where the exact wording matters.
Secret pre-flight
Scans prompts locally for API keys, private key blocks, database URLs, and card numbers before anything leaves your machine. You choose what happens next: redact, send anyway, or cancel. It warns; it never silently scrubs your prompt without telling you.
Cache alignment
Providers discount cached input, and a tool array rebuilt from a dictionary on every request can silently destroy that discount by reordering. Tool arrays are sorted deterministically and cache breakpoints are placed automatically, with no SDK changes required.
Shadow Mode
Measures what compression would have saved while forwarding your original prompt untouched, so you see real numbers on your own traffic with zero risk before you turn compression on for real.
Prompt Studio
A drafting workspace rather than a chat window. Character and token counts update as you type, next to a read-out of what is costing you: filler phrases, repeated sentences, removable tokens, and the reduction Optimize would achieve if you ran it. Optimize rewrites the draft in place and Copy takes the result. The grading is entirely local and nothing is sent to a provider from here — to compare a real answer before and after compression, use the A/B test, which makes two billed calls and says so.
Atlas workspace index
Point Atlas at a folder and it parses the code on your machine into symbols, imports, and a call graph, so an agent can be handed the call path a task actually needs instead of every file that touches it. Parsing, indexing, and search all run locally: no code, and no fragment of it, is uploaded. Searches built from your own symbol names are stripped of local identifiers before they reach anything external, and the app shows you exactly what it removed.
Monthly audit and PDF report
A month-by-month breakdown by model and by day, exportable as a PDF. Every cost figure is estimated from the token counts your provider returned and its published rates at the time of the request, and the report says so on the page. Shadow-mode and opted-out requests are recorded as potential rather than realised, and excluded from the estimate.
Device limits
Paid plans allow a set number of signed-in devices at a time (see pricing). This is enforced client-side, which raises the cost of casually sharing an account; it is not, and is never described as, DRM. Manage or sign out devices from your account page.
What reduction to expect
These are end-to-end figures through the proxy at the Balanced level, on traffic meant to look like real work rather than a best case. Run the same shapes through your own proxy and the Dashboard will show you your own numbers; that is the figure to plan against, not this table.
| Workload | Tokens removed |
|---|---|
| 40 JSON records, fields varying per row | 34% |
| 57-frame Node stack trace | 35% |
| 66 log lines with per-request ids and latencies | 24% |
| 12-turn agent history, prose only | 7% |
Uniform data compresses far harder than this: rows that repeat the same keys and values, or a log where the same line recurs hundreds of times, reach 70 to 90 percent because there is that much redundancy to collapse. Varied data does not, and most production data is varied.
One default is worth knowing about, because it decides whether any of this applies to you. Never touch the latest user message is on out of the box, so on a single-turn request, where the JSON or the log tail is the only user message, nothing is compressed and you will see 0%. That default is deliberate: the newest message is the one most likely to be worded exactly as the model needs it. Turn it off in Settings to compress single-turn payloads, and use Shadow Mode first to see what it would have changed.
Limitations, stated plainly
These aren't hidden in fine print because they're not a flaw we're hoping you won't notice; they're the shape of what pattern-based, local, offline tooling can and can't promise.
- Compression is lossy on prose. Textual distillation removes words. Removing words can change the response a model produces. We don't warrant that a compressed prompt yields an identical, equivalent, or comparable completion, only that Shadow Mode and the A/B tool let you check before you trust it on your own workload.
- Secret detection is pattern-based. It will not identify every secret, and it may occasionally flag content that isn't sensitive. A clean scan means no known pattern matched, not that the prompt is free of sensitive information.
- Loop detection is heuristic. It may occasionally block a request you intended to make, and it may fail to identify some genuine loops. It's a cost safeguard, not a guarantee against unexpected charges.
For the full licence terms, see the Terms of Service and Privacy Policy. Questions the docs don't answer: [email protected].