Beta v1.12.0|Methodology v2.3.0

SeanPropApp benchmarks

What long-form AI analysis actually costs, and what you give up to make it cheaper.

Two controlled studies on chained, multi-step LLM work: what each model tier is worth per dollar, and how far you can compress carried context before the thinking degrades. Every output, score and correction is published in full, details below.


One finding neither study set out to look for

Run the identical analysis three times, on the same model version, with the same inputs, and you do not get the same quality of thinking. The chart below is the spread of those repeat runs. The strongest model is also the steadiest, by nearly three times.

How much the same model varies between identical runs

Standard deviation of the Insight score (judged 0 to 10) on the first module, across 48 repeat runs per model. The first module is the fair test: it receives almost no carried context, so all four context strategies are effectively the same condition there, and what is left is the model's own run-to-run variation. Lower is better. A high value means the single analysis you actually commission may land well below what the model averages.

Claude Opus 4.80.60 SDClaude Sonnet 51.67 SDClaude Sonnet 4.61.38 SDClaude Haiku 4.50.94 SD

Ordered strongest model first. Claude Sonnet 4.6 to Sonnet 5 raised mean quality and also raised spread, 1.38 to 1.67: newer, better on average, less repeatable. One version step is not a law, which is why the fixed inputs are published for re-running against future models.

Open data

Every module output, judge verdict, token count and drift audit behind these studies is published: 8,065 files across four model versions, three replicate runs each.

github.com/seanomich/seanpropapp-benchmarks · corrections recorded in the open · the analysis script reproduces the published figures byte for byte

Beta Feedback