Run 2026-09-05 · 8 categories · deterministic grading

Open models matched the frontier on revenue-operations work — at a fraction of the price.

103 models ran 40 real GTM tasks — forecasts, CRM audits, renewal-risk calls, executive briefs — built from live, redacted company data. Every answer was graded by deterministic code against computed ground truth. No LLM judged any response. Every score is inspectable, answer by answer.

models tested
103
real GTM tasks
40
graded answers
4,576
llm judges used
0
personal data exposed
0
01

Read the exercise

Start with the narrative, then drill into any number. Everything links back here.

02

Results

Every ranking footnotes its basis: 40 tasks, deterministic rubric, no LLM judge.

03

Choosing a model

04

Verify the work

A ranking without evidence is not a BonuslyBench ranking. Every score below can be traced to a graded answer.