Blog · · 7 min read

Reading what Claude Code left behind

#The situation

You ask Claude Code why something is broken. It reads a dozen files, follows the call path, checks the config, and gives you a good answer: five findings, file and line numbers, a fix in the right order.

It gives it to you in the terminal, between the tool calls. You read half of it, switch to the editor to look at the line it mentioned, come back, and the answer is somewhere above two hundred lines of Read(app/orders.py). Tomorrow, when you actually sit down to fix it, the session is gone. Next month, when the same bug comes back in a different shape, nobody remembers what the agent found.

The investigation was the valuable part, and it was the part that nobody kept. The code change will be in git. The reasoning behind it lived in a scrollback buffer.

The fix is one paragraph in the prompt: tell the agent that the answer is a file.

#Ask the AI

Paste this into Claude Code, with your own question in the first two lines. It works the same in any agent that can write files in your repo.

Some customers are getting two receipt emails for one payment, and finance says
a few orders show the payment twice. Find out why.

Investigate only. Don't change any code, don't run migrations, don't commit.

When you're done, write what you found to a file:
notes/investigations/<short-topic>-<today's date YYYY-MM-DD>.md

The file:
- Starts with frontmatter: title, date, question (my question, one line),
  status: open, and files: (the files you read that matter).
- Then a # heading and a three-sentence answer: the cause, how sure you are,
  and the fix you'd make first.
- "## How it works today": the path the code takes, step by step, with
  file:line references so I can follow along.
- "## Findings": one ### per finding, most likely cause first. For each: what
  happens, the evidence (quote the lines), and how confident you are.
- "## What I'd change": a - [ ] checklist in the order you'd do it, smallest
  safe fix first.
- Put anything that could lose data or money in a > [!warning] callout.
- End with "## Couldn't verify": what you assumed, what you'd need to confirm it
  (logs, config, a test). Don't guess to fill a gap; list the gap here.

In the terminal, reply with the filename and one line on the likely cause.
Nothing else. The file is the answer.

Four lines in it do most of the work.

"Investigate only." An agent asked why will often start fixing, and then you are reviewing a diff before you have agreed on the cause. Separating the two gives you a moment to disagree with the diagnosis while it is still cheap.

The file:line references and quoted lines. A finding you can't check is an opinion. Quoting the evidence under each finding means you can verify it in ten seconds, and so can the colleague who reads the file after you.

"Couldn't verify". The agent read the code. It did not see production logs, the real config, or the email provider's latency. Giving those gaps a section of their own is what stops them from being quietly filled in with a plausible guess.

"The file is the answer." The terminal gets one line. That is deliberate: the moment the real answer is also in the terminal, that's where you read it, and the file becomes a copy nobody opens.

If you want this every time, put the file rules in your CLAUDE.md as a standing instruction. A prompt that makes your AI write files, not replies has the general version.

#What comes back

We ran this prompt on a small web shop's codebase with a real double-processing bug in it. In the terminal, Claude Code wrote exactly two lines:

notes/investigations/duplicate-receipts-2026-09-30.md

Likely cause: the webhook's duplicate check is an in-memory dict that each of the 4 gunicorn workers keeps separately, and it only records an event after the slow email call finishes, so provider retries get processed again.

The file is below, shortened for the page: the files: list, the fifth finding, three checklist items and half of the last section are cut, the evidence under findings 2 to 4 is trimmed, and a few lines are condensed.

---
title: Duplicate receipt emails and double payment rows
date: 2026-09-30
question: Why do some customers get two receipt emails for one payment, and why do some orders show the payment twice?
status: open
---

# Duplicate receipts and double payments

The cause is that the webhook handler's duplicate check is weak. It is an
in-memory dict that is private to each of the 4 gunicorn workers, and the event
is only remembered after the slow work (payment insert and email API call) has
finished, so provider retries of the same event id get processed again. I'm
fairly sure (about 85%) this is the cause, from reading the code only. The
first fix I'd make is a database-level guard: a UNIQUE constraint on the
provider event id, claimed atomically before any side effects.

## How it works today

1. The provider POSTs `payment.succeeded` to `/webhooks/payments`
   (`app/webhooks.py:16-17`). It waits 10 s for a 2xx and retries the same
   event id after 10 s, 1 min, 10 min and 1 h (`docs/payment-provider.md:4-5`).
2. `seen_recently(event_id)` looks in a module-level dict `_SEEN`
   (`app/dedupe.py:4`). If it finds the id, the handler returns 200.
3. `mark_paid` inserts a row into `payments` (`app/orders.py:6-12`).
4. `send_receipt` calls the email API with a 30 s timeout (`app/email.py:8-12`).
5. Only then is `remember(event_id)` called (`app/webhooks.py:34`).
6. gunicorn runs 4 workers (`gunicorn.conf.py:3`), each with its own `_SEEN`.

## Findings

### 1. The dedupe cache is per-process (most likely cause)

A retry is routed to a different worker than the original delivery. That
worker's `_SEEN` is empty, so it inserts a second payment and sends a second
receipt. With 4 workers, a retry has roughly a 3-in-4 chance of landing on a
worker that has not seen the event.

```
_SEEN: dict[str, float] = {}          # app/dedupe.py:4
workers = 4                           # gunicorn.conf.py:3
```

Confidence: high that the gap exists. Medium-high that it is the main source.

### 2. The event is remembered only after the slow work

If the email call takes longer than 10 s (its timeout is 30 s), the provider
counts the delivery as failed and retries. The original request is still
running and hasn't remembered the id, so both runs insert a payment and send a
receipt.

Confidence: high that the race exists. Whether real email latency exceeds
10 s is unverified.

### 3. A failed email means the payment is re-inserted on every retry

If `send_template` raises, `mark_paid` has already committed, and `remember()`
is skipped. Every retry inserts another payments row.

### 4. The database has nothing to stop duplicates

`payments` has no unique key on the order or the provider event. Every gap
above turns directly into a duplicate row.

## What I'd change

> [!warning] Money and customer data
> Existing `payments` rows may already contain duplicates. Adding a UNIQUE
> constraint will fail on those rows, and deleting or merging them changes
> financial records. Have finance review the duplicates first, take a backup,
> and do not auto-delete.

- [ ] Take a DB backup and list the duplicate `payments` rows for finance to
      review.
- [ ] Add a `processed_events` table with a UNIQUE event id, and claim the id
      with `INSERT ... ON CONFLICT DO NOTHING` at the top of the handler. Skip
      everything if no row was inserted.
- [ ] Record a `receipt_sent_at` on the order and only send the email if it is
      still empty, so a retry after an email failure can still send it.
- [ ] Move the email send out of the request, so the webhook returns 2xx well
      inside the provider's 10 s limit.
- [ ] Add tests: same event twice sequentially, twice concurrently, and with an
      email failure on the first attempt.

## Couldn't verify

- Whether production really runs `gunicorn.conf.py` with 4 workers. I'd need
  the deployed process list or config.
- How long the email API takes and how often it fails. I'd need application
  logs or the email provider's metrics.
- Whether the duplicate rows are about 10 s, 1 min or 10 min apart, which would
  confirm retries as the source. I'd need the provider's delivery logs and a
  query on `payments`.
- There are no tests in the repo, so I could not reproduce anything.

This is a good report. The diagnosis is right, the evidence is quoted, and the warning is the one a tired developer would skip: clean up the duplicate rows before adding the constraint, and let finance decide what counts as a duplicate.

It still deserves one careful read, and the file is what makes that possible. Two things stand out.

"About 85%" is invented precision. The agent can say it is fairly sure. It can't have measured 85 percent of anything. Read that number as fairly sure.

Two checklist items don't fit together yet. If the event id is claimed at the top and the email then fails, the retry is skipped, and the receipt is never sent. The receipt_sent_at item assumes the retry gets through. The fourth item (send the email outside the request) is what resolves it. That's the kind of thing you catch on a second read of a file, and never catch in a terminal you've already scrolled past.

#Save it as a file

The agent already did: it's in notes/investigations/. Connect the repo in MD Flow (what browsing a folder of Markdown looks like), filter the file list to Markdown, and the report sits in its folder next to the README instead of among the Python files.

#Why that matters next week

You read it as a document, not as terminal output. The frontmatter becomes a header, the warning is a coloured callout you can't skim past, the checklist is a checklist, and the quoted code sits in its own blocks. On a Mac, you can also select the file in Finder and press Space for a rendered preview, without opening anything.

It follows the agent. When you come back with "here are the delivery logs, update the report", Claude Code edits the same file, and MD Flow's preview follows each save without losing your place. The report grows as the investigation does. Nothing ends up in scrollback.

status: open means something. When the fix ships, change it to fixed and add the PR link. The folder becomes a record of what broke, why, and what was done about it. That's the context a new teammate (or a new agent session) needs before touching the webhook handler again.

The next investigation starts from this one. When duplicates show up somewhere else in six months, the file is there to find. Search by filename is free in MD Flow: duplicate finds it. With Pro, search also reaches inside the files, so gunicorn or idempotency finds every investigation that touched them, even if the filename doesn't say so. Hand that file to the agent as its starting point, and it doesn't have to rediscover what it already found once.

The terminal is where the agent works. The file is what it leaves behind, and that's the part worth reading.