OpenEvalz

Open model evaluation

Run one problem.
Publish the evidence.

Choose a benchmark problem and a model. Your TrustedRouter delegated key pays for the run; OpenEvalz publishes the full trace, score, captured gateway cost, and the model that actually served it.

Single problemYour spend capPermanent public trace

Evaluation

BFCL

Berkeley Function Calling Leaderboard problems for evaluating tool-use behavior.

Evaluation

AIME 2024

Competition mathematics problems from the 2024 American Invitational Mathematics Examination.

Evaluation

GPQA Diamond

Graduate-level, expert-written questions in biology, chemistry, and physics.

Evaluation

GSM8K

Grade-school math word problems that test multi-step arithmetic reasoning.

Evaluation

MATH

Challenging competition mathematics across algebra, geometry, counting, and more.