Skip to content

Blog

Insights on AI agent analytics, product management for AI, and building better AI products.

Braintrust Pricing vs. the Alternatives: What Full Trace Coverage Actually Costs
evalsobservabilityai-product-management

Braintrust Pricing vs. the Alternatives: What Full Trace Coverage Actually Costs

Five observability platforms, five incompatible billing units. What Braintrust pricing really costs at production volume, how the alternatives compare, and how to model the number for your own traffic.

Itamar Kramer7 min read
Abstract illustration of labelled trays catching shapes that fall into categories defined in advance
engineeringproductagent-analyticsobservabilityunder-the-hood

Stop Asking the Model What the Categories Are

We rebuilt issue clustering by moving the taxonomy out of clustering entirely — into detection. What broke in the first version, what the rewrite bought, and what it cost.

Daniel Brodsky10 min read
The Multiplicative Failure Trap: Why Multi-Tool AI Agents Crash Midway
agent-reliabilityobservabilityengineering

The Multiplicative Failure Trap: Why Multi-Tool AI Agents Crash Midway

Why multi-tool AI agents scraping websites, inboxes, and SaaS tools keep failing midway, and how state checkpointing stabilizes multi-step workflows.

Dennis Zagiansky8 min read
Abstract illustration of a narrow beam passing through a stack of data columns beside the same stack fully lit
engineeringtoolsobservabilityunder-the-hood

ClickHouse Text Indexes, Direct Read, and the Compatibility Trap

We shipped a text index that did byte-for-byte nothing while every diagnostic said it was working. The culprit: one compatibility setting silently pinning a prerequisite off.

Daniel Brodsky8 min read
Abstract illustration of one structure at two zoom levels, a few large blocks resolving into many small ones
engineeringproductagent-analyticsai-product-managementunder-the-hood

The Granularity Problem, or: Why Your Embeddings Don't Know What a "Third-Party Integration" Is

Two VPs want the same dashboard at different zoom levels. Why dendrogram cuts can't give it to them — and what to build instead: leaf dedup, extracted facets, and a declared taxonomy.

Daniel Brodsky11 min read
Detect LLM Application Blind Spots Scorers Miss
evalsai-product-managementobservability

Detect LLM Application Blind Spots Scorers Miss

Automated LLM scorers create a false sense of security through high raw agreement rates. Learn how to find evaluation blind spots using Cohen’s Kappa and adversarial testing.

Dennis Zagiansky10 min read
Mastering Parallel Claude Code Sessions: A Guide to Multi-Agent Workflows
claudemcpagent-reliabilitytoolsopen-source

Mastering Parallel Claude Code Sessions: A Guide to Multi-Agent Workflows

Six terminals open and no idea which agent is waiting on you. How to run parallel Claude Code sessions with fleet.

Yuval Hayke14 min read
Agent Evaluations: Why Output-Only Checks Fail
engineeringproducttools

Agent Evaluations: Why Output-Only Checks Fail

Evaluating AI agents on final output hides silent failures and loops. Learn why trajectory evaluation and observability are essential for reliable agentic workflows.

Dennis Zagiansky9 min read