tools
10 articles tagged “tools”.

ClickHouse Text Indexes, Direct Read, and the Compatibility Trap
We shipped a text index that did byte-for-byte nothing while every diagnostic said it was working. The culprit: one compatibility setting silently pinning a prerequisite off.

Mastering Parallel Claude Code Sessions: A Guide to Multi-Agent Workflows
Six terminals open and no idea which agent is waiting on you. How to run parallel Claude Code sessions with fleet.

Agent Evaluations: Why Output-Only Checks Fail
Evaluating AI agents on final output hides silent failures and loops. Learn why trajectory evaluation and observability are essential for reliable agentic workflows.

AI Agent Error Handling: Why Strict Tool Standards Matter
Fix silent failures with a strict AI agent error handling contract. Learn how structured tool errors improve debugging and observability for LLM agents.

RAG Retrieval Optimization: The #1 Mistake AI Builders Make
Upgrading your LLM to fix hallucinations is a costly trap. Learn how RAG retrieval optimization and key evaluation metrics solve the real bottleneck.

AI Product Management: Fixing Silent Model Degradation
Stop AI model drift. Learn why AI product management requires continuous evaluation of probabilistic systems to prevent silent failure and user churn.

AI Agent Analytics: Why Traditional SaaS Metrics Fail
Stop slicing AI agent data by industry or ARR. Learn how intent-based AI agent analytics surface real product gaps and improve agentic workflow performance.

Continuous Evaluation for AI: Why Traditional QA Fails
Stop silent AI degradation. Learn why continuous evaluation for AI is replacing traditional QA to monitor semantic drift and maintain LLM product quality.

LLM as a Judge: Why Eval Scores Fail Your AI Roadmap
LLM as a judge is great for regression testing but terrible for discovery. Learn why high eval scores hide flat retention and how to surface real user intent.

Mission Control for Claude Code: How to Manage 10 Agents Without Losing Your Mind
Stop tab-switching and start orchestrating. fleet is a terminal mission control for managing parallel Claude Code sessions.