⚔ Weights & Biases appears in the Perception Lab: forced-choice runs vs Braintrust →
When buyers ask ChatGPT, Claude, and Gemini about llm & agent evals, Weights & Biases ranks #2 of 15, a Visibility Score of 51 in LLM & Agent Evals .
These rankings are measured from what the three models know from training, not a live web search. The live-web grounded surface is rolling out across the index. Methodology.
Strong #2 in evals, but first-pick rate barely registers.
Where it wins. Weights & Biases holds the #2 rank in LLM & Agent Evals across 15 competitors, with a 64% mention rate that puts it near the top of how often AI models surface it to buyers in that category.
Where it loses. A 2.8% first-pick rate means models almost never open with Weights & Biases as the go-to evals tool, and the brand has no presence outside this single category to fall back on.
AI-generated analysis of the September 2026 measurements.
One of the index's bigger model splits: see where the models disagree →
Being named is not the same as being recommended. Each mention is graded endorsed (a strong pick), listed (a neutral option), or caveated (named with a reservation).
Share of each theme's answers that mention Weights & Biases. Hover a row for the exact prompt. Compare shapes on the head-to-head page.
Share of Weights & Biases's mentions where the AI models name this brand in the same answer. That is the real competitive set in the AI channel.
Head-to-head: Weights & Biases vs LangSmith
Scores are measured from real AI answers, refreshed monthly. Methodology.
We rerun the index every month. Drop your email and pick what to watch in LLM & Agent Evals: the whole category or specific brands. Free, unsubscribe anytime.