More reasoning, same answer
Cranking the reasoning-effort dial to max cost us 3.8× the tokens and roughly 3× the wait. The output wasn't bigger. Whether it was any better took four runs to answer — and the answer was mostly no.
The reasoning-effort setting is the easiest knob to talk yourself into. More thinking sounds like more quality, and it costs you almost nothing to bump it up. So we bumped it — all the way up the ladder, from the default to max, on the same synthesis step, holding everything else fixed. (How we hold things fixed is the flagship post.) Then we watched two numbers: what we spent, and what actually came back.
1 · The bill climbs. The output doesn't.
effort ladder · opus-4-8
Six runs, default up to max, nothing else touched. Tokens quadrupled. Cost nearly tripled. The wait stretched from a minute and a half to five. And the thing the reader actually gets — the brief on the card — barely moved off where it started.
2 · "But the prompt capped the length"
so we removed the cap
Fair objection. Our prompt asks each finding for "~400–800 chars," so maybe the output couldn't grow even if the model wanted it to. So we pulled the length limit out entirely and ran default against max one more time. With nothing capping length, max still burned 3.8× the tokens — and came back shorter, by 13%. The effort is real; it just goes into reasoning you never see, not into a longer or fuller brief.
3 · Shorter, but sharper?
the part that fooled us first
Shorter is fine if it's denser. And on the very first max run it did look sharper — it caught two methodological points the default had skipped. We nearly bought it. The flagship post tells the rest: we ran it three more times and those points never came back, the gap fell into ordinary run-to-run scatter, and every load-bearing part of the analysis was already sitting in the default output. On a good day max buys you tighter phrasing. Not something you can plan around, and not worth triple.
When the dial is worth it
match effort to the step
Turn it up when the step is genuinely reasoning-bound and nothing downstream is squeezing the output into a fixed shape — a fact-check, a verdict, a proof, where the thinking is the product. For synthesis into a review card, the default is the setting, full stop. And whichever way you lean, run the whole ladder once. It's a few dollars, and then you're choosing from a curve instead of a hunch.
Bench note · honest edges
One case, one synthesis step, one model, small n. Costs are CLI-measured and run high — the relative multiples are what hold; treat the absolutes as directional. Directional field notes, not a study.