Evaluation · 2026
Jev AI 101
Judgment is cheap now. Planning still is not.
Playable · full benchmark and method online · repo going public

Seven problems from the standard AI syllabus. A 270M-parameter decision model against the textbook algorithm on every one of them, and against a frontier LLM where that comparison makes sense. About 22,000 decisions, under three dollars.
The method
Every rate carries a 95% Wilson interval. Every sample size is stated. Every call’s tokens and list-price cost are recorded. The write-up includes a row for where other people’s work is ahead of mine, and a corrections list for the three numbers that changed when the sample grew.
The finding
Where the task is judgment, choosing well among options that are already on the table, the small model is on par with the textbook and a fraction of the cost. Where the task is planning, building a path through a space, it is not, and what buys planning is reasoning rather than size.
Try it
The seven exhibits are playable. The benchmark page carries the full method and every number.