Nikhil Deshpande

AI research · 2026

Jev AI 101

22,000 decisions, under three dollars

A hand picks a navy tile from six shapes on the left, while on the right a path of tiles steps up to a small orange flag.

The problem

Everyone was shipping demos on a new decision model. Almost nobody was measuring it against anything.

The system

Seven problems from the standard AI syllabus. The model against the textbook algorithm against a frontier LLM, given byte-for-byte the same state and the same legal options. Sample sizes and 95% intervals on every number. Every figure computed from the recorded run, nothing typed by hand.

What happened

Judgment is cheap and planning is not, and what buys planning is reasoning rather than model size. Switch reasoning off on the frontier model and it fails the same puzzles as the small one, at four times the cost. About 22,000 decisions, under $3.

Not claiming

Seven toy domains, one model version, one week. Three of the seven are deliberately out of scope by the vendor’s own documentation, so failures there corroborate them rather than catch them out.