No Way to Merge
Daniel Brodsky5 min read
Four issues in our catalog had exactly the same title: “Unable to Delete Workspaces.” They had been created within five minutes of each other.
At Brizz, we analyze AI agent conversations to detect failures and identify recurring problems. We call individual failures findings and cluster findings about the same underlying problem into an issue, so engineers can investigate repeated occurrences together. In this case, the pipeline kept creating separate issues for failures already represented in the catalog. The catalog also contained near-duplicates such as “Unable To Delete Items,” “Missing Ability To Delete Items,” and “Unable To Delete Board Items.”
How duplicates persisted
Repeatedly clustering the entire history would exceed our memory and compute budget, so we processed findings incrementally. Every twenty minutes, the pipeline embedded new finding descriptions, clustered them within the batch, and compared each cluster with existing issues. Below a distance threshold, it attached the findings to an issue. Otherwise, it created a new one.
Each existing issue was represented by an embedding derived from its first ten findings. That representation never changed. Later batches contained different samples and different phrasings of recurring failures, so their cluster representations could fall on opposite sides of the matching threshold. The clustering was deterministic, but no two runs saw the same input.
A false non-match created a duplicate. From then on, both issues remained in the catalog: the pipeline compared new clusters with existing issues, but never compared existing issues with each other. There was no merge operation.
Updating the representations or improving the matcher could reduce false creations. Neither would reconcile duplicates already created.
One failure across two batches
Consider a simplified example. In the first batch, findings describe users asking to delete a workspace. The agent has no deletion tool, and the pipeline creates issue A. In a later batch, findings emphasize cleaning up unused workspaces after a project ends. The missing capability is the same, but the different descriptions shift the batch’s cluster representation beyond the matching threshold, creating issue B.
Neither issue is reconsidered by the ingestion pipeline. The merge job can later retrieve A and B as candidates and inspect findings from both. If the evidence establishes the same missing workspace-deletion capability, it consolidates them. Similar wording caused by a different failure, such as insufficient permissions, would not justify that merge.
Could we fix the matcher?
We checked whether embedding distance separates findings about the same object (item, column, group) from findings about different objects.
The same-object and different-object distance distributions overlapped substantially. On pairs drawn from 681 findings, the best single threshold achieved 79% accuracy in distinguishing same-object from different-object pairs. A tighter threshold rejected more potential matches, increasing fragmentation. A looser threshold accepted more matches between distinct objects. In one clustering run, increasing the threshold produced a 149-finding cluster spanning eight objects.
This showed why threshold tuning involved a tradeoff rather than a clean separation. Fine-tuning an embedding model for issue matching is a promising research direction: can we learn a representation that separates failures requiring different fixes while bringing differently worded instances of the same failure closer together? Better matching could reduce duplicate creation, but it would not reconcile duplicates already in the catalog.

Same-object and different-object distances overlap, illustrating the tradeoff in choosing a matching threshold.
Why not recluster everything?
Full reclustering was too expensive to run routinely. It also required preserving the IDs, statuses, and history attached to existing issues: our one-time rebuild reduced the issue count, but discarded that state. We therefore needed a correction step that operated on selected existing issues and preserved their continuity. That became the periodic merge job.
The connection to incremental clustering
Ackerman and Dasgupta’s Incremental Clustering: The Case for Extra Clusters (2014) studies the limits of clustering with bounded memory as data arrives. They show that certain target partitions recoverable by batch methods cannot be guaranteed by deterministic memory-bounded incremental methods. Their positive results show how allowing additional clusters can enable recovery of refinements: smaller clusters contained within the intended groups.
Bounded memory was a real constraint in our system too. Incremental processing kept each run manageable, but the compact historical representations limited the evidence available to the matcher. Separately, our implementation had no reconciliation step, so a false creation persisted. We kept incremental ingestion and added targeted comparisons between existing issues rather than repeatedly reclustering the full history.
What we changed
We added a periodic merge job that compares existing issues. It retrieves candidate duplicates using embeddings, then uses an LLM to evaluate the pair against the underlying findings: would the same fix address both?
Accepted merges consolidate findings under a surviving issue and retire the duplicate, with provenance linking the original identities. This adds a correction path alongside incremental ingestion, evaluating candidate pairs without loading and reclustering the entire finding history.
What we learned
- Determinism does not imply stability across batches. The same algorithm receives different samples of a recurring failure, so its matching decisions can vary.
- Matching quality and correction are separate concerns. Improving the threshold or representation can reduce duplicate creation. A merge operation is needed to reconcile duplicates already in the catalog.
- Cluster maintenance includes application state. Reconciliation must account for issue IDs, history, and status; replacing cluster assignments alone is insufficient.
- Merge decisions need their own evaluation. Duplicate reduction must be measured alongside incorrect merges. Findings assigned to the wrong issue still require reassignment or splitting.