Skip to content

agent-reliability

6 articles tagged “agent-reliability”.

Mastering Parallel Claude Code Sessions: A Guide to Multi-Agent Workflows
claudemcpagent-reliabilitytoolsopen-source

Mastering Parallel Claude Code Sessions: A Guide to Multi-Agent Workflows

Six terminals open and no idea which agent is waiting on you. How to run parallel Claude Code sessions with fleet.

Yuval Hayke14 min read
Agentic AI Observability: Why Sampling Fails AI Agents
observabilityagent-reliabilityai-product-management

Agentic AI Observability: Why Sampling Fails AI Agents

Traditional APM sampling fails for probabilistic AI. Learn why agentic AI observability requires 100% visibility and wide events to fix silent failures.

Itamar Kramer10 min read
Why 100% Task Completion is the Most Dangerous Metric in Agent Evaluations
evalsobservabilityagent-reliability

Why 100% Task Completion is the Most Dangerous Metric in Agent Evaluations

Task completion is a dangerously misleading metric for AI agents. When agents are optimized to succeed, they will sometimes secretly alter the environment or rewrite test suites to guarantee a pass. Here is how to build a sabotage-proof evaluation framework.

Dennis Zagiansky8 min read
AI Agent Error Handling: Why Strict Tool Standards Matter
engineeringproducttoolsagent-reliability

AI Agent Error Handling: Why Strict Tool Standards Matter

Fix silent failures with a strict AI agent error handling contract. Learn how structured tool errors improve debugging and observability for LLM agents.

Dennis Zagiansky9 min read
RAG Retrieval Optimization: The #1 Mistake AI Builders Make
productengineeringtoolsragagent-reliabilityllm-cost

RAG Retrieval Optimization: The #1 Mistake AI Builders Make

Upgrading your LLM to fix hallucinations is a costly trap. Learn how RAG retrieval optimization and key evaluation metrics solve the real bottleneck.

Itamar Kramer8 min read
AI Product Management: Fixing Silent Model Degradation
productengineeringtoolsobservabilityagent-reliabilityai-product-management

AI Product Management: Fixing Silent Model Degradation

Stop AI model drift. Learn why AI product management requires continuous evaluation of probabilistic systems to prevent silent failure and user churn.

Daniel Shapira8 min read