Research

What we have measured, written to be read.

Every writeup states its numbers, its baselines, and its limits. Nothing here is a claim without a measurement behind it.

July 2026

A small AI that solves problems it was never trained on, by thinking longer.

We trained a small reasoning model, then tested it on problems harder than anything it had ever seen. Held to the amount of thinking it used in training, it got 11 of every 100 right. Allowed to keep thinking, it got essentially all of them. The full report covers the head-to-head against a model 2.3 times its size, the best published method, a model fifty times larger, why thinking longer works at all, and what is not proven yet.

Read the full report →
ReasoningEfficiencyResult report