Benchmark Sessions →

Benchmark Dashboard

Live rollups across every content-generation session — competition, SOCRATES, and CDE performance by LLM.

Total sessions →
Total questions generated
LLMs competing
CDE duplicates caught
Avg tokens / question
Avg time / question
Total cost (all competing models)

LLM Leaderboard

Every round's winner still counts as a win — "Regen Rate" is a separate signal for how often that came after a recompetition round rather than clean on the first attempt.

Provider Model Wins Losses Win Rate Avg FUSION Avg SOCRATES Avg Validator Avg Latency Avg Cost Regen Rate

CDE Leaderboard — duplicates caught per LLM

Lower is better: counts how often each LLM's winning content was flagged by the Claim Discrimination Engine as a duplicate and replaced.

Provider Model CDE Catches

Cost per Session

Date / Time Topics Items Total Cost Cost / Item Session