Stop Paying a Model to Make Decisions You Already Made
A skill that spells out a fixed procedure in prose makes Claude re-decide it every run. Here's how to measure that cost using data Claude Code already emits.
Join the DZone community and get the full member experience.
Join For FreeIf your team distributes AI development skills (as a Claude Code or Cursor plugin or a shared rules file), you own a catalog. That catalog covers things like:
- How to structure a service
- What must pass before a commit
- Which internal library to use instead of rolling your own
Skills load cheaply, thanks to progressive disclosure. Only the name and one-line description sit in context until the description matches the work. So when token spend climbs, the intuitive read is context bloat, and the obvious lever is tuning descriptions to fire less.
That work is worthwhile as it improves routing quality. But it addresses only one of the two ways a skill costs money. The other is procedural amplification: a skill that encodes a fixed procedure as prose, so the model reconstructs it, deliberating, calling tools, checking results, on every invocation. That doesn’t show up as a large context payload. It shows up as extra inference round trips, and an aggregate cost dashboard cannot see it.
Predictability does not prove a step should be automated. It identifies places where you should ask whether you’re paying a model to make a decision the system has effectively already made.
A few terms come up more than once below, defined here so they don’t slow you down later:
| Term | Meaning |
|---|---|
| Skill | A packaged set of instructions Claude Code loads only when the task matches it. |
| OTel | Short for OpenTelemetry, the open standard Claude Code uses to report what it did and how much it cost. |
| API request / round trip | One call to the model and its response. This article counts these to measure cost. |
| Hook | A small script Claude Code runs automatically at a specific moment, such as right after every tool call. |
| p50 / p90 / p95 / p99 | Percentile timing. p50 is the typical (median) call; p99 is close to the worst case you’ll see. |
| Async | A setting that lets a hook run in the background instead of making the agent wait for it to finish. |
| NDJSON | One JSON record per line in a file. Simple to append to and to stream. |
Use the Vendor’s Telemetry With Small Customization
I’ll correct a claim I believed myself and have seen repeated: OpenTelemetry only provides aggregate counters, so you have to build your own event pipeline. That’s false, and starting from scratch costs you a week.
Claude Code’s OTel export includes events, and they’re richer than I expected:
| Event | Carries |
|---|---|
claude_code.api_request |
skill.name, cost_usd, input_tokens, output_tokens, duration_ms,event.sequence, prompt.id |
claude_code.tool_result |
tool_use_id, tool_name, success, duration_ms |
skill.name on the API request is the important one: inference round trips are already attributed to the skill that was active.
cost_usd is documented as an estimate, not billing data.
The One Field to Add
Getting command detail out of the native exporter requires OTEL_LOG_TOOL_DETAILS=1, which attaches full_command, bash_command, and file_path to tool events. That’s exactly the payload you cannot centralize off a fleet of developer laptops: branch names, commit messages, customer identifiers, the occasional pasted token.
So a small local hook fills one gap: a privacy-preserving shape of the command, joined to native telemetry on tool_use_id, which the docs describe as matching the tool_use_id passed to hooks, allowing correlation between OTel events and hook-captured data.
native OTel ──┐
│ api_request: skill.name, cost, tokens, round trips
├── join on tool_use_id ──▶ analysis
│ tool_result: tool_use_id, tool_name, success
local hook ──┘
The hook doesn’t rebuild the trace. It emits the join key plus the one thing native telemetry can’t safely give you. Anything OTel already reports (success, duration_ms, skill.name, token counts) is deliberately not duplicated. Two sources of truth for one field will eventually disagree, and you’ll trust the wrong one.
The command-to-shape transform itself is a per-binary allowlist, not a secret detector, and it has a sharp edge worth knowing:
- The transform:
git commit -m "fix auth for acme"becomesgit commit -m <ARG>. - The edge case: a value-taking flag like
--token abc123leaks its argument as a bare token unless you explicitly track which flags consume the next word. Get the allowlist wrong, and the safety argument for the whole pipeline goes with it.
What the Hook Costs (Measured)
Tool hooks run on the critical path, so I benchmarked one: 1,000 synthetic PostToolUse payloads through the real script, mixed across Bash/Read/Edit/Write/Grep/Skill.
p50 p90 p95 p99 max mean
subprocess 24.28ms 26.91ms 27.91ms 32.31ms 52.42ms 24.93ms
in-process 0.08ms 0.15ms 0.16ms 0.20ms 0.35ms 0.10ms
macOS, arm64, 10 cores, Python 3.9.6, n=1000 after 25 discarded warmups.
The interesting part isn’t the total; it’s the split. About 24.2 ms of that 24.3 ms p50 is Python interpreter startup. The actual work — scrubbing, serializing, appending — costs 0.08 ms. That points the fix somewhere I would not have guessed:
- Don’t optimize the scrubber. It’s already three orders of magnitude below the launch cost.
- Do fix how the hook is launched. Register it
async: trueso nobody waits, or make it a thin client to a long-lived local collector so you pay startup once per session instead of per tool call.
Had I estimated instead of measured, I’d have guessed low single-digit milliseconds and gone tuning the scrubbing code. Wrong target entirely, which is the argument for measuring rather than estimating, in miniature.
Event records came in at 364 bytes each. One engineer at roughly 40 sessions a week and 60 tool calls a session generates about 2,400 records, under 1 MB of raw NDJSON weekly. It stays small-data at any plausible team size, because the enrichment record only carries what native OTel doesn’t.
Measuring Your Own Catalog
Skills have two token costs that behave differently, and knowing which one dominates for you tells you where to look first:
| Cost component | Paid when | Scales with |
|---|---|---|
| Description | Every request, every session, whether the skill fires or not | Catalog size |
| Body (procedural amplification) | On invocation, then re-sent on each subsequent request in the session | How often the skill fires, and how long sessions run |
Measuring the 25 skills in Claude Code’s official plugin marketplace, a catalog anyone can install and re-measure, gives:
body chars: min 989 median 11,395 max 32,625 (33x spread)
desc chars: min 105 median 357 max 906
always-resident total (all 25 descriptions): 9,703 chars
Read those two numbers against each other. Carrying every description in the catalog costs less than a third of one large body. The single biggest skill is roughly 3.4 times the entire always-resident footprint, and it’s charged again on every request for the rest of any session that invokes it.
That’s the shape that makes procedural amplification worth hunting. If your always-resident total instead dwarfs your median body, the lever really is catalog size and description tuning. You can stop there.
(These are characters, not tokens. The ratio shifts with how much code versus prose a body contains. Count with the real tokenizer, messages.count_tokens, which is free and rate-limited only by requests per minute; don’t reuse a chars/4 estimate or another vendor’s tokenizer, and don’t reuse an old Claude count either: the tokenizer used by Claude 4.7 and later yields roughly 30% more tokens for the same text than earlier models.)
Why This Is Worth Measuring at All
The subject isn’t token optimization. It’s finding misplaced probabilistic computation: places where the system already knows the next operation and is paying a model to rediscover it.
| If the next operation is… | It belongs in… |
|---|---|
| Genuinely uncertain | A skill, policy and judgment, model reasoning |
| Already determined | A script, deterministic execution |
A skill that spells out a fixed procedure in prose has put that boundary in the wrong place. The industry is converging on the same line from the runtime side: programmatic tool calling exists precisely to keep deterministic multi-tool sequences out of the inference loop.
Key Takeaways
Aggregate cost accounting answers the question you already knew to ask. Event-level behavioral traces let you find the one you didn’t, and in Claude Code, most of that stream already exists, attributed to the skill, waiting on one privacy-preserving field to become useful.
Procedural amplification leaves no trace in a token counter, raises no error, and turns no dashboard red. It’s visible in the order of the calls, and in how many round trips it takes to get through them.
- Measured, not estimated: Hook overhead is about 24 ms per call, almost entirely interpreter startup, not the redaction logic itself.
- Measured, not estimated: In one real catalog, a single skill body costs 3.4× what the entire always-resident description set costs.
- Next step: Measure by the actual stretch of work, not by an arbitrary number of calls. Group everything that happens while one skill is active, however long that turns out to be, rather than picking a fixed count upfront. When one model turn fires off several tool calls at once, count that as the single decision it was, not several. And don't trim the long, expensive stretches out of the data before analyzing it. Those are usually the exact cases worth finding.
Opinions expressed by DZone contributors are their own.
Comments