Kratoq
← All Bench Notes
POST B / 4·OUTPUT LENGTH·2026

You're not setting the length. Your wording is.

We changed nothing but the phrasing of one length instruction. The output swung from 340 to 644 characters — same evidence, nearly twice as long. And a hidden floor had been quietly padding thin findings the whole time.

modelclaude-opus-4-8
taskCh.1 Foundation synthesis
knoblength instruction
evidencefrozen · identical

We thought we were telling the model how long each finding should be. What we were really doing was nudging it, and the nudge did most of the work. Feed it the exact same eight sources every time, change only the sentence that talks about length, and the output moves further than the content ever does. Here is the same brief at four different phrasings.

1 · The wording sets the length

same sources · default effort

"Up to 800, be concise" gave us findings around 340 characters. "400–800" — our current prompt — pulled them up to about 400. "At least 400, as long as it needs" pushed to ~520. Take the instruction out entirely and the model settled around 640. Nothing about the evidence changed between these. The length was set by how we asked, not by what there was to say.

Fig.B1mean chars / finding · by length instruction
200400600340≤800, concise401← current400–800519≥400, as needed644no instruction
1.9× on wording alone. Identical sources, default effort. The phrase describing length moved the output from 340 to 644 characters — the model has no stable "natural" length it reveals; it writes to whatever you imply.

2 · A floor quietly pads

400–800 vs no floor

The "400–800" range looks harmless. It isn't, because it hands the model a floor and the model leans on it. With the range in place the findings cluster around 400 and rarely drop below ~370, even when a point is thin. Take the floor away and the same findings fall to a natural 340, some down to 296. The floor wasn't holding a standard. It was making the model write filler to reach a number.

Fig.B2chars / finding · range + mean · default effort
400 floor369435with the 400 floor296399floor removed250300350400450
with the floorfloor removed
The floor buys padding, not quality. Dot = mean, bar = min–max across a brief's findings. With the floor, findings sit around 400; without it the whole distribution slides down ~60 characters and dips below 400 where the content is genuinely thin.

3 · Let go, and length follows content

the upside of no bounds

Removing the bounds does something better than shorten things. It lets the findings stop being uniform. Under "400–800" they landed within about 65 characters of each other — every finding roughly the same size no matter how much it had to carry. Unbounded, that spread jumps past 200. The finding with a lot to cover runs long; the thin one stays short. That is what you actually want: length carrying information instead of filling a quota. It isn't purely content-driven — the model still drifts toward a house length — but the wording dominates, and the floor is the part doing active harm.

Fig.B3length spread within one brief · by instruction
10020066400–800103≤800111≥400215free
The tighter the rule, the more uniform the output. Spread = the character gap between the longest and shortest finding in one brief. It climbs as you loosen the instruction — 66 with the full 400–800 range, ~105 with a single bound, 215 with none — and the loosest setting is the one that lets each finding run as long as its content earns.

Three dials, not one

separate what you fused

It is tempting to treat "how long the answer is" as a single setting. It is really three: how much the model thinks, how much it says, and how much the surface can show. We had fused the last two into one prompt clamp. That clamp padded thin answers, would have truncated rich ones, and — as the effort post found — quietly muddied a separate experiment we were running. If you have to bound length, bound it with a ceiling tied to a real layout, never a floor. Let substance decide the rest.

Bench note · honest edges

One case, one synthesis step, one model. Length numbers are single-run means at default effort; the spread figures are per-brief. The model still has a house length even when free — "unbounded" widens the variation, it doesn't hand length purely to content. Directional, not a study.