Thunderdome B2B SaaS AI Perception Index

← LLM & Agent Evals ranking

Brand report · September 2026

DeepEval DeepEval: AI Visibility Report

deepeval.com ↗

⚔ DeepEval appears in the Perception Lab: forced-choice runs vs Braintrust →

When buyers ask ChatGPT, Claude, and Gemini about llm & agent evals, DeepEval ranks #9 of 15, a Visibility Score of 19 in LLM & Agent Evals .

These rankings are measured from what the three models know from training, not a live web search. The live-web grounded surface is rolling out across the index. Methodology.

How AI sees DeepEval

Mid-pack eval tool fighting for scraps in a crowded field.

Where it wins. DeepEval earns a 32% mention rate in LLM & Agent Evals, meaning models surface it in roughly one of every three relevant conversations. That foothold is real, even if the rank sits at 9 of 15.

Where it loses. It draws zero first-pick selections, so buyers hear the name but LLMs consistently prefer other tools when forced to commit. Enterprise and feature-driven prompts skip it entirely, cutting off the two buyer segments most likely to drive serious deal flow.

AI-generated analysis of the September 2026 measurements.

LLM & Agent Evals · #9 of 15

19 / 100 first snapshot (September 2026)
32%
mention rate (72 answers)
5
avg. position
0%
first pick
5%
share of voice
First snapshot: September 2026. The trend line appears with the second monthly measurement; no modeled history.
ChatGPT
16
Claude
6
Gemini
34

One of the index's bigger model splits: see where the models disagree →

How the models portray DeepEval

Being named is not the same as being recommended. Each mention is graded endorsed (a strong pick), listed (a neutral option), or caveated (named with a reservation).

48% 35% 17%
endorsed · 95% band 29–67% listed caveated graded across 23 mentions in LLM & Agent Evals answers
Prompt themes DeepEval would want to own

Share of each theme's answers that mention DeepEval. Hover a row for the exact prompt. Compare shapes on the head-to-head page.

Gaps Enterprise pick · Feature-led ask

Most co-mentioned competitors

Share of DeepEval's mentions where the AI models name this brand in the same answer. That is the real competitive set in the AI channel.

LangSmith
96%
Ragas
57%
Promptfoo
57%
Weights & Biases
43%
Langfuse
43%
Arize Phoenix
43%

Scores are measured from real AI answers, refreshed monthly. Methodology.

Get alerted on big movements

We rerun the index every month. Drop your email and pick what to watch in LLM & Agent Evals: the whole category or specific brands. Free, unsubscribe anytime.

Or watch specific brands:

Work at DeepEval? Put this on your site: a live "AI-recommended" badge, grounded in this data, that updates monthly and links back here. Free. Get the badge →