Braintrust is in the Perception Lab → aided, forced-choice, and grounded measurement of what Braintrust is to the models, beyond this ranking.
Platforms for evaluating and observing LLM apps and AI agents: eval suites, LLM-as-judge scoring, tracing, and production monitoring. Right now, LangSmith is the brand AI assistants recommend most, with a Visibility Score of 77 and the #1 spot in 46% of answers. Head-to-head: LangSmith vs Weights & Biases →
These rankings are measured from what the three models know from training, not a live web search. When the models DO search the web, see which sources they cite. Methodology.
LangSmith dominates; no one else is close.
LangSmith's 76.9 score sits 26 points above second-place Weights & Biases, and it shows up in 85% of answers while being picked first nearly half the time. Gemini is its loudest advocate at 93.8, while ChatGPT and Claude score it more modestly, yet even those modest scores beat most competitors. Behind LangSmith, Weights & Biases holds a stable second on the back of ChatGPT and Claude, but Braintrust, Langfuse, and Arize are essentially splitting thin air with scores bunched between 30 and 36.
AI-generated analysis of the September 2026 measurements.
| # | Brand | Visibility Score | Mention rate | Avg. position | #1 pick rate | Models |
|---|---|---|---|---|---|---|
| 1 | | ±5.8 | 85% | 1.9 | 46% | ChatGPT #2 Claude #1 Gemini #1 |
| 2 | | ±6.3 | 64% | 3.1 | 3% | ChatGPT #1 Claude #3 Gemini #5 |
| 3 | | ±6.1 | 44% | 2.9 | 15% | ChatGPT #5 Claude #2 Gemini #11 |
| 4 | | ±6.2 | 43% | 2.9 | 10% | ChatGPT #3 Claude #5 Gemini #8 |
| 5 | | ±6.1 | 36% | 2.7 | 13% | ChatGPT #4 Claude #7 Gemini #7 |
| 6 | | ±7.4 | 40% | 3.7 | 1% | ChatGPT #7 Claude #8 Gemini #2 |
| 7 | | ±6.2 | 36% | 3.3 | 11% | ChatGPT #11 Claude #4 Gemini #3 |
| 8 | | ±5.5 | 33% | 4.8 | 0% | ChatGPT #12 Claude #6 Gemini #6 |
| 9 | | ±5.3 | 32% | 5.2 | 0% | ChatGPT #8 Claude #12 Gemini #4 |
| 10 | | ±4.6 | 24% | 5.5 | 0% | ChatGPT #9 Claude #9 Gemini #9 |
| 11 | | ±4.3 | 14% | 4.4 | 0% | ChatGPT #6 Claude Gemini |
| 12 | | ±3.4 | 10% | 3.6 | 0% | ChatGPT #14 Claude #11 Gemini #10 |
| 13 | | ±4.6 | 10% | 5.0 | 1% | ChatGPT #10 Claude Gemini #13 |
| 14 | | ±3.3 | 10% | 6.1 | 0% | ChatGPT Claude #13 Gemini #12 |
| 15 | | ±2.6 | 7% | 4.6 | 0% | ChatGPT #13 Claude #10 Gemini |
The ± under each score is the reliability band: how much the number would move if we re-ran the identical measurement. It is tight because each prompt is sampled multiple times. It is not the same as how much the ranking depends on which prompts we ask — a separate, wider figure shown on hover.
Just below the cutoff: AgentOps (4), TruLens (3.9), Patronus AI (3.8), Portkey (3.5), Evidently AI (3.2). 42 brands were scored in this category; the leaderboard shows the top 15.
Visibility Score: position-weighted presence across all measured answers, 0 to 100. A score of 100 means the brand was the first recommendation in every answer. Every point in a trend line is a live monthly measurement; there is no modeled or backfilled history. Full methodology.
Openness 23/100: the share of this category's recommendation weight not held by the leader.
Compare brands in LLM & Agent Evals head-to-head → · Explore the prompt tree →
We rerun the index every month. Drop your email and pick what to watch in LLM & Agent Evals: the whole category or specific brands. Free, unsubscribe anytime.
Per-model Visibility Scores for the top 10. A long gray bar means the three assistants disagree about the brand. Hover a row for exact values.
When a brand is mentioned, where does it appear? Box shows the middle 50% of positions, the thick line is the median, whiskers are the extremes. Hover a row for the detail.
gpt-5.4
claude-sonnet-4-6
gemini-3-flash-preview
Each model answered every prompt 3 times in September 2026. Prompts are brand-neutral so no vendor gets seeded into the question.