Open model evaluation
Run one problem.
Publish the evidence.
Choose a benchmark problem and a model. Your TrustedRouter delegated key pays for the run; OpenEvalz publishes the full trace, score, captured gateway cost, and the model that actually served it.
Single problemYour spend capPermanent public trace
Evaluation
BFCL
Berkeley Function Calling Leaderboard problems for evaluating tool-use behavior.
Evaluation
AIME 2024
Competition mathematics problems from the 2024 American Invitational Mathematics Examination.
Evaluation
GPQA Diamond
Graduate-level, expert-written questions in biology, chemistry, and physics.
Evaluation
GSM8K
Grade-school math word problems that test multi-step arithmetic reasoning.
Evaluation
MATH
Challenging competition mathematics across algebra, geometry, counting, and more.