LangSmith
LLM observabilityStep by step trace debugging, LLM as judge evals, and prompt management. Its Insights Agent now clusters production traces bottom-up into usage patterns and failure modes with no taxonomy defined up front, and multi-turn evals score semantic intent and task completion across a whole conversation. Reporting stops at the category: error rates and eval scores per cluster, no single score for the agent and no check that a shipped fix held.