Model Archetypes — How the Models Actually Behaved
91 models with ≥20 tests, clustered on behavioral dimensions — not just scores: per-category strengths, cost, latency, fabrication events, and math-vs-prose tilt. This is an acknowledgement of observed behavior, not a recommendation. Every model links to its full answer-by-answer record in the Response Reviewer.
Use these profiles conversationally: "these models perform like this." A Complete Analyst behaves like a senior operator; a Fast Communicator writes well but stumbles on rollforward math; Capable-but-Loose models need an entity-validation guardrail. Clustering is k-means (k=6, 20 seeds) over standardized category profiles + log-cost + log-latency + fabrication rate.
The Complete Analysts
26 models
elite-accuracyfastall-around-analyst
median mean 0.949median cost $6.60median latency 49sstrongest communicationweakest data & crmfabrication events 0
Elite-accuracy, fast, all-around performers. High scores in every category with no weak flank; communication and reporting both strong. These models behave like a strong senior analyst: correct math, disciplined formatting, honest caveats.
median mean 0.923median cost $0.63median latency 83sstrongest ops & maintenanceweakest data & crmfabrication events 0
Strong, budget-priced workhorses. Slightly below the elite tier, best on ops-maintenance and routine execution, weakest on data-crm edge cases. Behavior profile: steady, cheap, occasionally thin on the hardest multi-table reconciliation.
median mean 0.929median cost $0.90median latency 89sstrongest ops & maintenanceweakest reporting & analyticsfabrication events 12
Strong scores at budget prices, but this cluster carries fabrication events — models that invented a deal or company alias at least once. Correct on most work, but their output needs entity-validation guardrails before it touches production.
median mean 0.728median cost $0.43median latency 52sstrongest communicationweakest reporting & analyticsfabrication events 0
Budget, fast, prose-oriented. Comfortable writing digests and emails, weak on the heavy reporting math (weighted forecasts, rollforwards). Behavior profile: good words, shaky arithmetic — fit for drafts, not for numbers.
median mean 0.742median cost $0.51median latency 18sstrongest ops & maintenanceweakest reporting & analyticsfabrication events 12
Ultra-fast budget models (median 18s) with mid-tier accuracy and fabrication risk in-cluster. Trade correctness for latency; usable where a quick pass beats a right pass, with validation.
median mean 0.124median cost $0.00median latency 29sstrongest communicationweakest customer & successfabrication events 0
Endpoints that returned instant garbage or refused to work ($0 cost, 3-6s responses). Not real scores — broken routes. Listed for completeness and exclusion.
Where price and performance actually meet. $/point = total suite cost ÷ mean score — the price of one unit of benchmark performance. The value frontier is the pareto set: no model is both cheaper and better than these.
THE VALUE FRONTIER — cheapest at every score level
VALUE KINGS — ≥0.93 mean at ≤$2 total (12 of 16 shown)
These models deliver top-decile scores for the price of a coffee. mimo-v2.5-pro, glm-5.3-flash, tencent/hy3, qwen3.8-flash all outscore every frontier model on this bench at under $0.75 each.
High spend that the scores didn't justify. gpt-5.5-pro is the outlier of the entire exercise: $123.62 for the suite (129× the value kings' median) with a score below 16 models costing under $2. The premium tier's honest pitch is speed and vendor polish — on this bench, not accuracy.
Reading note: this is observed price-performance on these 40 tests, not a vendor verdict — list prices shift, and latency/SLA/enterprise features aren't priced here. The complete dollar map is in the Excel workbook and the value-map chart.