01 — Res · 412 tasks
Account research
Given a company URL, a 10-K excerpt and four news items, produce the account brief a rep would open before a first call: what changed, who owns the problem, what the trigger is. Graded on cited specifics and on penalising invented facts.
Rubric graders: 3 · Hallucination penalty: −12
02 — Deep · 288 dossiers
Company & persona research
Twenty tool-using minutes on one company and one named person: what the company just changed, who sits on the buying committee, what that persona is measured on, which of their own public words to quote back. Graded on cited primary sources, on committee coverage, and on a hard penalty for inferred job scope.
Sources per dossier: 30+ · Uncited claim: −15
03 — Out · 468 tasks
Outbound copy
Write the first-touch email and the follow-up under real constraints: a named persona, a 90-word ceiling, one ask, no unearned familiarity. Graded on specificity to the account, on the strength of the ask, and against a spam-signal classifier.
Human preference pairs: 1,404
04 — Call · 386 tasks
Call analysis
Turn a 40-minute discovery transcript into qualification fields, risks and next steps. Graded field-by-field against an operator's own notes, with credit for flagging what the rep failed to ask.
Exact-match fields: 14 per call
05 — Obj · 340 turns
Objection handling
Six turns against a scripted buyer who is polite, budget-constrained and already talking to a competitor. Graded on whether the model holds price, surfaces the real blocker, and stops selling when the buyer says yes.
Buyer personas: 22 · Turns per run: 6
06 — CRM · 460 records
CRM extraction
Read a messy thread and emit a schema-strict object: accounts, contacts, amounts, stage, close date. Graded on schema validity, on deduplication against existing records, and on abstaining rather than guessing.
Schema violations: hard zero
07 — Seq · 346 runs
Sequence execution
Plan and execute fourteen days of multi-threaded follow-up with eight tools available: enrichment, calendar, email, CRM write, pricing lookup. Graded on plan coherence, tool-call validity and recovery after a failed call.
Tools: 8 · Max steps: 30
08 — Plan · 246 plans
End-to-end account plan
One named account, three personas, four channels, thirty days: where to enter, how to sequence the committee, what each persona hears, and what to do when the champion goes quiet. Graded head-to-head against the plan an operator wrote for the same account, and on whether the plan survives its own first failure.
Operator head-to-head · Personas per plan: 3
Contamination control
Every account, transcript and CRM snapshot in a run is drawn from material first published in the 30 days before that run, or generated with operators under NDA. Nothing in the July 2026 set was on the public web when any evaluated model finished training. Prior sets are released after the following run.
Grading
Rubric items are graded three times — twice by a model panel that excludes the model under test, once by a human operator on a 20% sample. Panel and human disagree on 4.1% of items; those items are re-graded by a second operator and the human label wins.
Settings
Temperature 0 where supported, no system-prompt tuning per model, no retries. Reasoning models run at their default effort. Every model sees the same tool schemas. Runs are single-shot: the score you see is the first answer, not the best of five.