Write evaluation units: a task, the allowed tools, a gold outcome, pass/fail rules, and how an agent typically fails. This is the METR / Scale evals / internal-red-team market, packaged as a licensable module instead of a one-off labeling shift.
This venue is a Wing in your Human architecture. Every exercise creates Rooms organized by Hall type. Data quality is tracked via Knowledge Graph triples.
Meet all milestones to unlock buyer engagement. Hybrid threshold — both content volume AND time consistency required.
Entries
Active Days
Min Days
Leaderboards lie. Production evals are private, domain-specific, and written by people who have done the job. That is what labs and agent vendors buy when they are about to put an agent in front of a customer.
6–9 hours initial, then 40 min/week · 4 steps · Each step builds your architecture
Eval suite (JSON: task, tools, gold, rubrics[], failure_modes[])
When readiness is achieved, an AAAK-compressed summary of your data is generated for buyer evaluation — privacy-preserving, compact, and instantly readable by any AI system.
Runs evals and human review for labs and enterprises deploying agents.
Expert evaluators for agent and model quality programs.
Evaluation and human-review platform used by teams shipping agents.
Needs SWE and ops evals that are not public GitHub issues.
Open the Data Builder to complete guided exercises. Each entry builds Rooms, Halls, and Knowledge Triples in your Human architecture — reaching buyer readiness in 30–60 days.