Kratoq
← All Bench Notes
POST A / 4·REASONING EFFORT·2026

More reasoning, same answer

Cranking the reasoning-effort dial to max cost us 3.8× the tokens and roughly 3× the wait. The output wasn't bigger. Whether it was any better took four runs to answer — and the answer was mostly no.

modelclaude-opus-4-8
taskCh.1 Foundation synthesis
knobreasoning effort
replicatesn = 4

The reasoning-effort setting is the easiest knob to talk yourself into. More thinking sounds like more quality, and it costs you almost nothing to bump it up. So we bumped it — all the way up the ladder, from the default to max, on the same synthesis step, holding everything else fixed. (How we hold things fixed is the flagship post.) Then we watched two numbers: what we spent, and what actually came back.

1 · The bill climbs. The output doesn't.

effort ladder · opus-4-8

Six runs, default up to max, nothing else touched. Tokens quadrupled. Cost nearly tripled. The wait stretched from a minute and a half to five. And the thing the reader actually gets — the brief on the card — barely moved off where it started.

Fig.A1opus-4-8 · indexed to none = 100
100200300400nonelowmedhighxhighmax22,783 tok · 4.0×1,805 chars · 1.1×both start here
output tokens (what you pay)final brief (what you get)
Same start, opposite fates. none→max: tokens 5,629→22,783 (4.0×), cost 2.8×, wall-clock 87s→297s (3.4×). The brief itself goes 1,602→1,805 chars. You pay four times over for 13% more words.

2 · "But the prompt capped the length"

so we removed the cap

Fair objection. Our prompt asks each finding for "~400–800 chars," so maybe the output couldn't grow even if the model wanted it to. So we pulled the length limit out entirely and ran default against max one more time. With nothing capping length, max still burned 3.8× the tokens — and came back shorter, by 13%. The effort is real; it just goes into reasoning you never see, not into a longer or fuller brief.

Fig.A2length unbounded · max ÷ default · n = 4
1.0× default3.81×tokens2.82×cost3.03×latency0.87×brief size
what max effort costswhat it delivers
The dial moves cost, not output. Length fully unbounded, mean of four runs. Three amber bars you pay for; one short indigo bar you get back.

3 · Shorter, but sharper?

the part that fooled us first

Shorter is fine if it's denser. And on the very first max run it did look sharper — it caught two methodological points the default had skipped. We nearly bought it. The flagship post tells the rest: we ran it three more times and those points never came back, the gap fell into ordinary run-to-run scatter, and every load-bearing part of the analysis was already sitting in the default output. On a good day max buys you tighter phrasing. Not something you can plan around, and not worth triple.

When the dial is worth it

match effort to the step

Turn it up when the step is genuinely reasoning-bound and nothing downstream is squeezing the output into a fixed shape — a fact-check, a verdict, a proof, where the thinking is the product. For synthesis into a review card, the default is the setting, full stop. And whichever way you lean, run the whole ladder once. It's a few dollars, and then you're choosing from a curve instead of a hunch.

Bench note · honest edges

One case, one synthesis step, one model, small n. Costs are CLI-measured and run high — the relative multiples are what hold; treat the absolutes as directional. Directional field notes, not a study.