Kratoq Research · Bench Notes
Bench Notes
Field notes from building an AI research tool — small, honest evals of the LLM step behind the product. Directional, not a study.
More reasoning, same answer
Turning reasoning effort to max cost 3.8× the tokens and 3× the wait — for an output that wasn't bigger or, once we replicated, any better.
Read →How to test an LLM change without fooling yourself
We changed a knob, ran it once, and it looked better. Then we replicated and the gain vanished. The cheap setup we use to tell a real gain from a lucky run.
Read →You're not setting the length — your wording is
Same evidence, only the phrasing of one length instruction changed. Output swung from 340 to 644 characters — and a hidden floor was padding thin findings the whole time.
Read →The checker that can't read a verdict
We planted six known errors in a report and swapped the cheapest model into the fact-check. It caught two — and never once caught a result described by the wrong verdict.
Read →