Results — July 2026 run

Nico 2.5 is the first model to clear 88 on GTM Workbench.

GTMBench scores language models on what revenue teams actually do: research the account, build the dossier on the company and the buying committee, write the outbound, run the call, handle the objection, keep the CRM clean, and plan a named account end to end. Eight task families, 2,946 graded tasks, rebuilt every month from accounts and transcripts that did not exist when the models were trained.

01Nico 2.5GTM foundation model

MacroDeep’s GTM foundation model leads three of the eight families at roughly a third the blended cost of the closed frontier. It cedes account research, company & persona research, call analysis, crm extraction and sequence execution — the families where raw frontier scale still beats specialisation.

88.4
Top GTM score
2,946
Graded tasks
8
Task families
10.5
Closed → open gap, best of each
12/30
Open-weight entries

The leaderboard

Sort any column. Select a row for the full profile. Scores are 0–100; the GTM score is the unweighted mean of the eight families.

License
GTMBench leaderboard — GTM score and the eight family sub-scores for every evaluated model, with blended cost per million tokens and context window.
#License
Showing 30 of 30 modelsBlended cost is USD per million tokens, assuming a 3:1 input-to-output ratio at July 2026 list prices.

Open vs. closed

The gap is not uniform — and most of it belongs to one specialist.

On account research the best open-weight model is 7.1 points off the closed frontier. On outbound copy the same comparison costs 21.1. Strip the purpose-built GTM model out of the closed column and the widest gap roughly halves: what is left is a pricing question, not a capability one.

Best-in-class score per license for each of the eight task families.
FamilyClosedOpenGap
Account research87.980.8+7.1
Company & persona research90.680.2+10.4
Outbound copy92.971.8+21.1
Call analysis88.880.7+8.1
Objection handling91.275.2+16.0
CRM extraction86.578.5+8.0
Sequence execution90.179.9+10.2
End-to-end account plan91.580.2+11.3

Best-in-class per license, July 2026 run.

Score against price

Blended cost per million tokens, log scale. Filled marks are closed models; outlined marks are open weights.

Scatter plot of GTM score against blended cost per million tokens on a log scale. Score rises with price, but the leading GTM specialist sits well above the trend line at mid-range cost, and the strongest open-weight models cluster near 78 at under a dollar.90807060$0.10$1$10$100Nico 2.5Fable 5DeepSeek-V4-ProDeepSeek-V4-Flash

The eight families

Every task is scored 0–100 against a rubric written by working GTM operators, then blind-graded three times.

01 — Res · 412 tasks

Account research

Given a company URL, a 10-K excerpt and four news items, produce the account brief a rep would open before a first call: what changed, who owns the problem, what the trigger is. Graded on cited specifics and on penalising invented facts.

Rubric graders: 3 · Hallucination penalty: −12

02 — Deep · 288 dossiers

Company & persona research

Twenty tool-using minutes on one company and one named person: what the company just changed, who sits on the buying committee, what that persona is measured on, which of their own public words to quote back. Graded on cited primary sources, on committee coverage, and on a hard penalty for inferred job scope.

Sources per dossier: 30+ · Uncited claim: −15

03 — Out · 468 tasks

Outbound copy

Write the first-touch email and the follow-up under real constraints: a named persona, a 90-word ceiling, one ask, no unearned familiarity. Graded on specificity to the account, on the strength of the ask, and against a spam-signal classifier.

Human preference pairs: 1,404

04 — Call · 386 tasks

Call analysis

Turn a 40-minute discovery transcript into qualification fields, risks and next steps. Graded field-by-field against an operator's own notes, with credit for flagging what the rep failed to ask.

Exact-match fields: 14 per call

05 — Obj · 340 turns

Objection handling

Six turns against a scripted buyer who is polite, budget-constrained and already talking to a competitor. Graded on whether the model holds price, surfaces the real blocker, and stops selling when the buyer says yes.

Buyer personas: 22 · Turns per run: 6

06 — CRM · 460 records

CRM extraction

Read a messy thread and emit a schema-strict object: accounts, contacts, amounts, stage, close date. Graded on schema validity, on deduplication against existing records, and on abstaining rather than guessing.

Schema violations: hard zero

07 — Seq · 346 runs

Sequence execution

Plan and execute fourteen days of multi-threaded follow-up with eight tools available: enrichment, calendar, email, CRM write, pricing lookup. Graded on plan coherence, tool-call validity and recovery after a failed call.

Tools: 8 · Max steps: 30

08 — Plan · 246 plans

End-to-end account plan

One named account, three personas, four channels, thirty days: where to enter, how to sequence the committee, what each persona hears, and what to do when the champion goes quiet. Graded head-to-head against the plan an operator wrote for the same account, and on whether the plan survives its own first failure.

Operator head-to-head · Personas per plan: 3

Contamination control

Every account, transcript and CRM snapshot in a run is drawn from material first published in the 30 days before that run, or generated with operators under NDA. Nothing in the July 2026 set was on the public web when any evaluated model finished training. Prior sets are released after the following run.

Grading

Rubric items are graded three times — twice by a model panel that excludes the model under test, once by a human operator on a 20% sample. Panel and human disagree on 4.1% of items; those items are re-graded by a second operator and the human label wins.

Settings

Temperature 0 where supported, no system-prompt tuning per model, no retries. Reasoning models run at their default effort. Every model sees the same tool schemas. Runs are single-shot: the score you see is the first answer, not the best of five.

Submit a model

Labs and vendors can request an eval for the next monthly run (2026-08-14). Open-weight models are run at our cost; hosted models need an API key with a 40M-token allowance. Results publish whether they flatter you or not.

Cite this run

@misc{gtmbench2026,
  title  = {GTMBench: benchmarking language models on
            go-to-market work},
  author = {GTMBench Collective},
  year   = {2026},
  note   = {Run 2026.07, 2946 tasks, 30 models},
  url    = {https://gtmbench.org}
}