Skip to content

AI Recommendation Standings / archive

LLM observability and evals: August 2026

The final August 2026 standings for “What should I use for LLM observability and evals?”. These numbers are fixed and can be cited by permalink. See the current standings.

View as Markdown

AI share of voice · August 2026

610 sampled answers · 10 models

  1. 1LangSmithnew23.5%
    Claude Haiku 4.5 · Claude Opus 4.8 · Claude Sonnet 4.6 · GPT · Gem · DeepSeek V4 Flash 0731 · GLM 5.2
  2. 2Langfusenew22.8%
    Claude Haiku 4.5 · Claude Opus 4.8 · Claude Sonnet 4.6 · GPT · Gem · DeepSeek V4 Flash 0731 · GLM 5.2· quiet
  3. 3Heliconenew19.6%
    Claude Haiku 4.5 · Claude Opus 4.8 · Claude Sonnet 4.6 · GPT · Gem · DeepSeek V4 Flash 0731 · GLM 5.2
  4. 4Braintrustnew15.9%
    Claude Haiku 4.5 · Claude Opus 4.8 · Claude Sonnet 4.6 · GPT · Gem · GLM 5.2
  5. 5Arize Phoenixnew14.2%
    Claude Haiku 4.5 · Claude Opus 4.8 · Claude Sonnet 4.6 · GPT · Gem · DeepSeek V4 Flash 0731 · GLM 5.2
  6. 6DeepEvalnew9.7%
    Claude Haiku 4.5 · Claude Opus 4.8 · Claude Sonnet 4.6 · GPT · Gem · DeepSeek V4 Flash 0731 · GLM 5.2

Share of voice weights each compared model equally: it is the average, across the models, of how often each one named the tool in its own answers, so a model probed more than another can't skew it. Model chips list which AI models named it this month; the pulse is the sentiment of that month's developer mentions across Hacker News, Reddit, GitHub, Stack Overflow, Bluesky and more.

Quotable · copy freely

Averaged equally across 10 AI models, LangSmith is named in 23.5% of answers about llm observability and evals (7 of 10 name it).
Source: BackTalk AI Recommendation Standings, August 2026 · backtalk.sh/ai-recommends/llm-observability/2026-08
The 3 prompts behind these standings
  • What should I use for LLM observability? (recommendation)
  • What are the best LangSmith alternatives? (alternatives)
  • Which tool is best for evaluating LLM outputs? (recommendation)

Each prompt is sampled several times per model per probe because answers vary run to run; 610 answers across 8 probe runs fed this month.