MediumPro job$500–$1,200/mo

Agent Evals

Write evaluation units: a task, the allowed tools, a gold outcome, pass/fail rules, and how an agent typically fails. This is the METR / Scale evals / internal-red-team market, packaged as a licensable module instead of a one-off labeling shift.

Palace Wing: Agent Evals

This venue is a Wing in your Human architecture. Every exercise creates Rooms organized by Hall type. Data quality is tracked via Knowledge Graph triples.

Halls populated: Facts, Events, Discoveries, Preferences, Advice
Tunnels form when you work across multiple venue Wings

Readiness Requirements

Meet all milestones to unlock buyer engagement. Hybrid threshold — both content volume AND time consistency required.

30

Entries

21

Active Days

30

Min Days

Why Robotics & AI Companies Pay Premium

Leaderboards lie. Production evals are private, domain-specific, and written by people who have done the job. That is what labs and agent vendors buy when they are about to put an agent in front of a customer.

How to Create This Content

6–9 hours initial, then 40 min/week · 4 steps · Each step builds your architecture

Output Specification

Final Deliverable Format

Eval suite (JSON: task, tools, gold, rubrics[], failure_modes[])

Quality Checklist
  • Minimum 30 eval units from domains you actually work in
  • Each unit has a gold outcome and at least two failure modes
  • Rubric is binary or 1–5 with written descriptors — not vibes
  • No leaked proprietary test sets from your employer
  • At least five units require tool use, not chat
AAAK Buyer Preview

When readiness is achieved, an AAAK-compressed summary of your data is generated for buyer evaluation — privacy-preserving, compact, and instantly readable by any AI system.

Active Buyers (4)

Scale AI

AI Company
High Demand

Runs evals and human review for labs and enterprises deploying agents.

Looking for: Private, domain-specific eval units
Accepts: JSON
Min. volume: 30+ units
Visit

Surge AI

AI Company
High Demand

Expert evaluators for agent and model quality programs.

Looking for: Gold tasks with written rubrics
Accepts: JSON
Min. volume: 25+ units
Visit

Labelbox

AI Company

Evaluation and human-review platform used by teams shipping agents.

Looking for: Structured eval specs
Accepts: JSON
Min. volume: 20+ units
Visit

Cognition

AI Company
High Demand

Needs SWE and ops evals that are not public GitHub issues.

Looking for: Engineering evals with gold patches and tests
Accepts: JSON
Min. volume: 20+ units
Visit

Ready to start building your Agent Evals Wing?

Open the Data Builder to complete guided exercises. Each entry builds Rooms, Halls, and Knowledge Triples in your Human architecture — reaching buyer readiness in 30–60 days.