I am on the $20 Claude Pro plan, and for as long as I have used Claude Code I have kept a folder of scripts next to it. One trimmed noisy test output. One checked how much of my session was left. They worked, but they always ran outside the tool, after the tokens were already spent.
When Claude Code shipped mods, small TypeScript plugins that run inside it and can see every tool call, I started moving those scripts in. Slowly, and only the ones I had actually used and measured. claude-pro-kit is the first five.
Look before cutting
The first thing I wanted was not a saving. It was a clear view of where my usage went.
What I found surprised me:
- One turn sent 947,448 tokens in. My context was nowhere near that size. Every tool call resends the whole conversation, and that turn made five of them. 98.5% of it was cached, so it cost about 1% of my session, but I would never have known without seeing it.
- A tool I never use went out with every request. Its description alone was 19,870 characters.
- In a long chat, the conversation itself is 95%+ of the context. Honestly, this was the most useful thing I learned. Starting a fresh chat more often beats any mod.
So two of the mods only show things. pro-hud draws live meters above the prompt in the desktop app: the 5-hour session, the week, the context window, and a running receipt for the current turn. context-xray opens the exact breakdown /context computes in a pane with /xray.


Three small cuts
| Mod | What it does |
|---|---|
| tool-diet | Lists tools you have not used in your last five sessions by name only; Claude loads them through ToolSearch when it needs one. |
| output-diet | Cuts long shell output to the first 30 lines, the last 50 and any error or warning lines, and saves the full text to a file Claude can open. |
| reread-guard | Skips a repeat Read of the same range of a file whose size and modification time have not changed. |
tool-diet made the biggest single difference. With one prompt in a fresh session, as the API reported it, every request went from 43,859 prompt tokens to 27,905 (−36%). On my setup that moved 38 tools. It decides once per tool per session, so it never disturbs the prompt cache mid-conversation.
output-diet kicks in when a Bash or PowerShell result runs past 120 lines or 8,000 characters. One caveat I found while building it: Claude Code already cuts the middle out of very long shell output before any mod sees it. So the saved file holds what Claude would have read, not the raw process output. I wrote that down in the README instead of pretending otherwise.
reread-guard has the smallest job. The one detail that mattered was letting a deliberate retry through. Claude Code can clear old tool results from context, so if Claude asks for the exact same read again straight after a skip, it gets the file.
Rules I held to
- No mod calls the model. They cost nothing to run.
- No estimated numbers. Every figure on screen is one Claude Code already reports, or a time the mod measured itself.
- Commands answer with toasts.
/hud,/tool-dietand/xrayadd nothing to the conversation they are trying to keep small. - Each mod is one file and stands alone. Mods are not sandboxed, so anyone installing them should be able to read the code first.
What the benchmark shows
bench/run.sh gives Claude one task, three times without the mods and three times with them: run an 800-line test suite, find the one failure, fix it, re-run.
| Mean input tokens | Mean output tokens | Mean cost | |
|---|---|---|---|
| Without mods | 252,815 | 862 | $0.2575 |
| With mods | 222,863 | 1,179 | $0.1713 |
| Change | −12% | +37% | −33% |
Every run with the mods cost less than every run without them, and every run fixed the bug.
It is still one task and three runs per side. Token counts swing a lot between runs because Claude takes a different path each time. Output went up, not down. reread-guard had nothing to skip here, so the difference is output-diet's. And on a subscription, the dollar figure is only a proxy for how fast the session meter moves.
What comes next
The next mod I am planning is a turn budget. One long agentic turn can quietly eat a large part of a session between your messages; I once saw a single turn run for 45 minutes and send 10.5 million tokens in. The idea is to pause before the next request once a turn crosses a limit and ask whether to keep going.
Same rules as the first five: exact numbers only, and nothing spent while it waits.
Source: the five mods, reread-guard in full and the benchmark script.