Nikhil Deshpande

Evaluation · 2026

Jev AI 101

Judgment is cheap now. Planning still is not.

Playable · full benchmark and method online · repo going public

A hand picks a navy tile from six shapes on the left, while on the right a path of tiles steps up to a small orange flag.

Seven problems from the standard AI syllabus. A 270M-parameter decision model against the textbook algorithm on every one of them, and against a frontier LLM where that comparison makes sense. About 22,000 decisions, under three dollars.

The method

Every rate carries a 95% Wilson interval. Every sample size is stated. Every call’s tokens and list-price cost are recorded. The write-up includes a row for where other people’s work is ahead of mine, and a corrections list for the three numbers that changed when the sample grew.

The finding

Where the task is judgment, choosing well among options that are already on the table, the small model is on par with the textbook and a fraction of the cost. Where the task is planning, building a path through a space, it is not, and what buys planning is reasoning rather than size.

Try it

The seven exhibits are playable. The benchmark page carries the full method and every number.