At bemo we inherited a legacy reward-accrual path that had been quietly undercounting some users’ rewards. By the time it surfaced, the wrong numbers had been compounding for a while: stored balances disagreed with what users had actually earned, and nobody could say by how much without doing the math properly.
The tempting fix is a forward patch: correct the accrual formula, maybe add a one-off compensation, move on. We didn’t do that, because a forward patch preserves every historical error forever — you keep paying interest on a number you know is wrong.
Instead we treated the blockchain as the source of truth it is and recomputed everything from scratch:
- fetch the full on-chain transaction history for each affected user;
- replay it event by event — stakes, unstakes, reward accruals — applying the correct accrual rules at each block’s actual rate;
- diff the replayed balance against the stored one to get a per-user correction.
The replay engine was plain Python. The interesting work was in the guarantees around it:
- idempotency — the job could die and restart at any point without double counting; every correction was keyed by its derivation inputs;
- invariants — the sum of all corrections had to reconcile against pool totals before anything was written;
- a dry run first — corrections landed in a shadow table, we sampled users and verified them by hand against explorer data, and only then applied.
Users got their missing rewards, and the incident left behind something more useful than an apology: a replayable pipeline. The next time anyone doubts a balance, the answer is one replay away, not one spreadsheet-archaeology session away.