restartpten
zshdlq-doctor.md

status: complete

dlq-doctor

An agent that diagnoses and replays RabbitMQ dead letter queues — with human approval on every fix.

9/9 blocks178 testsMIT

The problem

Dead letter queues pile up with thousands of messages nobody investigates: the errors repeat, and replaying them by hand is risky — duplicated side effects, a poison message coming right back. dlq-doctor ingests the queue, clusters failures by embedding, uses an agent with read-only tools to propose a cause and a fix per cluster, and runs a safe replay: dry-run, human approval, canary, idempotency key.

Pipeline

ChaosGenerator→tickets.dlq→IngestionWorker→ClusteringWorker→DiagnosisAgent→ReplayExecutionService

Real numbers

N = 2,326 failing messages (partial from a 20k batch — declared as such in the README, not the plan's full volume).
ApproachClustersPurityARI
Group by exception_type372.4%0.506
Exact fingerprint hash8100.0%0.679
Embeddings + incremental threshold (real)472.4%0.505

Honest finding: at this volume, the real system ties the naive baseline of grouping by exception type — both merge two failure categories that throw the same exception, exactly as ADR-0003 predicted.

Technical decisions

ADR-0001

RabbitMQ's x-delivery-limit only counts a genuine connection loss, never an explicit nack(requeue: true) — verified against a real broker.

ADR-0003

Two failure types throw the same exception on purpose, so the clustering evaluation had a real chance to show a baseline failing.

ADR-0006

Embeddings via a local ONNX model (no key, no cost) and incremental-threshold clustering — verified empirically before trusting cosine similarity.

ADR-0010

Evaluation metrics computed per message, not per fingerprint, compared against two simpler baselines — not just the system itself.

Stack

.NET 10RabbitMQ 4PostgreSQL + pgvectorMicrosoft.Extensions.AIAnthropic SDKPollyTestcontainersxUnit