status: complete
dlq-doctor
An agent that diagnoses and replays RabbitMQ dead letter queues — with human approval on every fix.
The problem
Dead letter queues pile up with thousands of messages nobody investigates: the errors repeat, and replaying them by hand is risky — duplicated side effects, a poison message coming right back. dlq-doctor ingests the queue, clusters failures by embedding, uses an agent with read-only tools to propose a cause and a fix per cluster, and runs a safe replay: dry-run, human approval, canary, idempotency key.
Pipeline
Real numbers
| Approach | Clusters | Purity | ARI |
|---|---|---|---|
Group by exception_type | 3 | 72.4% | 0.506 |
| Exact fingerprint hash | 8 | 100.0% | 0.679 |
| Embeddings + incremental threshold (real) | 4 | 72.4% | 0.505 |
Honest finding: at this volume, the real system ties the naive baseline of grouping by exception type — both merge two failure categories that throw the same exception, exactly as ADR-0003 predicted.
Technical decisions
RabbitMQ's x-delivery-limit only counts a genuine connection loss, never an explicit nack(requeue: true) — verified against a real broker.
Two failure types throw the same exception on purpose, so the clustering evaluation had a real chance to show a baseline failing.
Embeddings via a local ONNX model (no key, no cost) and incremental-threshold clustering — verified empirically before trusting cosine similarity.
Evaluation metrics computed per message, not per fingerprint, compared against two simpler baselines — not just the system itself.