ZoomInfo launches GTM Bench to evaluate AI agents on go-to-market tasks
ZoomInfo released GTM Bench to evaluate AI agents on go-to-market tasks like building target lists and enriching records. Version 1 covers over 20 jobs and 4 systems. GTM.AI led with a score of 77, delivering 478 verifiable records per 1,000 at a cost of $0.79 per task.

*this image is generated using AI for illustrative purposes only.
ZoomInfo (NASDAQ: GTM) has released GTM Bench, a versioned benchmark designed to evaluate Large Language Models (LLMs) and AI agents on the specific work performed by go-to-market teams. The benchmark assesses systems on tasks such as building target lists, enriching records, scoring accounts, and reaching decision-makers. Version 1 covers more than 20 jobs, 4 systems, and 3 models, with published methodology and grading rubrics.
The benchmark addresses the specific constraints of go-to-market operations, where roughly 70% of B2B contact data decays annually. Unlike closed-world reasoning benchmarks, GTM Bench measures performance based on data availability and verifiability. Results are graded against a senior GTM operator's work product on two independent axes: Answer, which measures the share of the requested work product delivered, and Grounding, which measures the share of returned data traceable to a real, current source.
In the v1 run, ZoomInfo's GTM.AI led every pillar with a GTM Bench Index of 77. This compares to 47 for Apollo, 36 for Exa, and 31 for open-web search. GTM.AI finished 98% of the operator's work product and returned 478 verifiable records per 1,000, compared to 7 to 35 for the field. The system also operated at the lowest cost of $0.79 per task. On the same 1,000-contact suite, the non-ZoomInfo field returned 720 wrong phone numbers.
| System | GTM Bench Index | Verifiable Records (per 1,000) | Cost per Task |
|---|---|---|---|
| GTM.AI | 77 | 478 | $0.79 |
| Apollo | 47 | 7 - 35 | - |
| Exa | 36 | 7 - 35 | - |
| Open-web search | 31 | 7 - 35 | - |
ZoomInfo acknowledges that GTM Bench is a vendor-run benchmark and publishes it accordingly, including areas where its advantage is thin or absent, such as pure copywriting or owned CRM data. Grounding is graded against ZoomInfo's verified records. The benchmark is versioned and will be re-run on major model releases, with version 2 set to add agentic multi-step workflows, international coverage, and an owned-data axis.
GTM.AI serves as ZoomInfo's headless GTM context layer, exposing the GTM Context Graph and agentic orchestration through API and Model Context Protocol. The system powers integrations including Salesforce Agentforce, HubSpot Breeze, Microsoft Copilot, Claude, ChatGPT, Gong, LeanData, and Google Workspace. Other data providers and agent builders can submit their systems for inclusion in the leaderboard.
How will competitors like Apollo and Exa respond to these findings, and will they release their own benchmarks to challenge GTM Bench's methodology?
Will the inclusion of agentic multi-step workflows and international coverage in Version 2 significantly alter the current rankings on the leaderboard?
To what extent will third-party agent builders adopt the GTM Context Graph API given ZoomInfo's dominant performance in the benchmark?




























