Grader Labs hi@graderlabs.com

Rollouts

Evaluation runs of our RL environments against frontier models, replayed exactly as they came out of the harness.

recorded run · not a simulation
prime eval run

What you're looking at. Each environment turns a real exporter's paperwork into tasks with a programmatic grader, so a model can be scored automatically with no human in the loop. Every prompt, completion and reward above came from an actual run against the named model. The per-channel scores show where a model fails, not just that it did.