The subject of this piece was handed to me by a task I no longer run myself. Every morning a scan walks the Anthropic changelog, the docs and a handful of news sources, checks whether anything in there changes how I work, and puts at most three suggestions in front of me. I read the result on my phone and pick one. That is how this article started.
What I never knew about that task is what it costs. It runs, it produces something, and beyond that it was a line in a config file. Since Claude Code 2.1.243 there is a Loops section in /usage that lists each recurring task with how often it fires, how many times it ran, how many tokens that took in total and per run, and when it last went off. So for the first time there is a number attached to the work I automated away.
I wrote about that automating away earlier, in the piece on loops and repeat work. That one asked which work repeats often enough that it should stop being yours. This one is about the bill that keeps running afterwards.
Three lines in the same release
Three items in 2.1.243 are really about one question, which is where your cost figure comes from.
| What | What it does | Requires |
|---|---|---|
Loops section in /usage | One row per recurring task, ordered by total tokens, with run count, total tokens, tokens per run and last run | Claude Code 2.1.242 or later |
promptCacheTtl and subagentPromptCacheTtl | Pick the prompt cache lifetime yourself, separately for the main conversation and for everything outside it | Claude Code 2.1.242 or later, value 5m or 1h |
modelPricing | Makes /cost and the status line use your organization’s contracted rates instead of list price | Managed setting only |
The first is a meter. The second is a dial the meter points you toward once you have read it. The third fixes something almost nobody thinks about, which is that the dollar figure in /usage is computed locally at standard list rates. That number knows nothing about your discounts and nothing about promotional pricing, and the documentation says outright that it may differ from your actual bill. Anyone on a discounted contract has been looking at a figure that was structurally too high. Anyone on a subscription has been looking at a figure that is not relevant for billing at all.
Why a daily task never has a warm cache
The cache setting is the most interesting of the three, precisely because for my scan it points the wrong way.
Prompt caching works on the beginning of your request. Claude Code resends the whole context every turn, and the API checks whether the start of it matches something it processed recently. That match is exact, so a change anywhere near the front recomputes everything behind it. Which is why the parts that rarely change sit first: the system prompt with the tool definitions, then the project context from CLAUDE.md, and only then the conversation itself.
A cache entry expires after a period of inactivity, and every request that hits it resets the clock. The API offers two lifetimes. Five minutes is the default for anyone on an API key or a cloud provider. One hour keeps the cache warm across a longer break, and it carries a price: a cache write at the five-minute lifetime costs 1.25 times the base input price, at the one-hour lifetime 2 times, and reading from cache costs 0.1 times. For Sonnet 5 that works out to $2 per million tokens for base input, $2.50 for the short write, $4 for the long one, and 20 cents to read.
Put those two facts next to each other and my scan falls outside of it. It runs once every twenty-four hours. The gap between two runs is a full day, so the cache is cold on the next run whether I set five minutes or an hour. What the longer lifetime gets me is nothing at all. What it costs me is the difference between $2.50 and $4 per million tokens, every morning, on a prefix nobody ever reads back. The documentation puts it in one sentence. The longer lifetime costs more on short bursts of work that never idle past five minutes, where the higher write rate applies and the longer lifetime goes unused. My scan is the mirror image of that. It idles so long that even the hour does not help.
The setting is aimed at something else, at a long working session with pauses in it, where you come back to the same conversation after half an hour away. For recurring tasks spaced far apart, 5m is the cheapest choice, and that happens to be the default.
There is a second split worth naming. The two settings cover two buckets. The main conversation is your interactive turns, your -p runs and the Agent SDK. Everything else is subagents, workflows, compaction and session titles. On a subscription within your plan usage the main conversation already gets an hour by default and the rest gets five minutes. Go over your limit so that Claude Code starts drawing on usage credits, and the main conversation drops back to five minutes, because from that point you are paying for it directly.
The number the meter does not give me
I do not have a figure for my own scan yet, and that is not laziness.
The Loops section computes from the local session history on this machine. Work from another device does not count, work through claude.ai does not count, and the docs call the figures approximate themselves. There are two windows, the last 24 hours and the last 7 days. A scan that only just started running on this version does not have a week of runs to average over.
What I do know from the same documentation page is the order of magnitude Anthropic uses for a developer: roughly $13 per active day on average and $150 to $250 per month, with 90 percent of users staying below $30 per active day. That is an average across enterprise deployments with no sample size given, so it says something about scale and very little about my scan. I include it because it is the only public reference I have, not because it stands in for my own number.
The caveat
The meter arrives after the habit. My scan has been running since 10 August, and a row about it only existed in /usage on 25 August. That is the order in which this kind of tooling shows up. First you can automate something, then you can see what it does. Anyone who forgot a loop during that gap noticed nothing and gets no correction after the fact. The one brake that does exist is that a recurring task inside a session expires after seven days and removes itself.
A number steers behavior, including in the wrong direction. The moment I know what a run costs, I will want to make it shorter. But the reason this scan runs daily is that a quiet day would otherwise let a subject slip past, and the meter cannot account for that spend. The editorial value of a run that finds nothing is exactly zero on the meter and not zero in practice. A meter that can measure one thing turns that one thing into the criterion.
The cache split rewards the wrong design. By default the main conversation gets the long lifetime and everything outside it gets the short one. So talking on in a single session works out cheaper than pushing work to subagents, which is the very thing that keeps context clean and lets a judgement come from somewhere other than the writer of the text. You can correct it with subagentPromptCacheTtl, but the default is what most people will keep.
What I am doing
I am leaving promptCacheTtl alone. This scan gains nothing from a longer lifetime, and the five-minute default is the cheapest setting for it.
I will read the Loops section once a week of runs has accumulated, using the week window rather than the day window, because a day window on a daily task shows exactly one run. The figure that comes out of that goes into this article, and not before.
And I am moving none of my editorial checks to a model. The six gates that block the build here are node scripts with no token usage at all: the spacing between publications, reused sentences, the caveat in AI pieces, the provenance of the cover image, the slop patterns, and the numbers checked against the source bundle. A check that can be deterministic should not need a meter.
Frequently asked questions
Is a one-hour cache any use if my task runs daily?
No. The cache entry expires long before the next run starts, so you never read it back. You do pay the higher write rate of 2 times base input instead of 1.25 times. For recurring tasks with a long gap between them, 5m is the right setting.
Is the amount /usage shows me correct?
Only if you pay list price. Claude Code computes the figure locally at standard rates and knows nothing about discounts, contracted pricing or a temporary promotional rate. The new managed modelPricing setting lets an organization supply its own rates. For actual billing, the usage page in the Console remains the source.
Why is my loop missing from the Loops section?
Three possible reasons. You are on a version older than 2.1.242. Your task is not among the heaviest and falls into the count of the rest. Or you ran it on a different machine, since the figures come from local session history and do not include other devices or claude.ai.
Sources
- Claude Code CHANGELOG, versions 2.1.242 through 2.1.245 (accessed 25 August 2026) — https://github.com/anthropics/claude-code/blob/main/CHANGELOG.md
- Claude Code docs, “Manage costs effectively” (accessed 25 August 2026) — https://code.claude.com/docs/en/costs
- Claude Code docs, “How Claude Code uses prompt caching” (accessed 25 August 2026) — https://code.claude.com/docs/en/prompt-caching
- Claude Code settings reference,
promptCacheTtlandsubagentPromptCacheTtl(accessed 25 August 2026) — https://code.claude.com/docs/en/settings-reference - Claude API docs, prompt caching and pricing (accessed 25 August 2026) — https://platform.claude.com/docs/en/build-with-claude/prompt-caching
Checked on 25 August 2026. I read the three changelog lines at the primary source, and the cache rates and default lifetimes come from Anthropic’s own documentation. Two things I could not verify. modelPricing appears in no documentation page as of this date, only in the changelog, so I could neither look up its behavior nor try it out without an organization on contracted pricing. And my own token figure is missing, because the Loops section computes from the local history of a single machine and there is not yet a full week of runs on it. As soon as that number exists, it goes here.
This piece also appeared in Dutch: Wat mijn dagelijkse loop kost, en waarom ik dat nooit wist.
