The problem nobody talks about

Every AML platform has fail-open logic. When the sanctions screening engine is unreachable, the transaction passes. This is deliberate. You can't block the entire payment rail because one service is down.

But here's what nobody puts in their sales deck: when fail-open fires, nothing happens. No record. No alert. No recovery. The transaction screens as if nothing was wrong, and if the engine was down for six hours, every transaction in that window passed with a silently degraded check.

The silent gap: for a bank facing a CBN examination, "our sanctions engine was down for six hours and we don't know which transactions were affected" is the nightmare scenario.

What we built

Three layers, each addressing a different aspect of the gap:

1. Evidence: record every fail-open

When the sanctions engine is unreachable or erroring during a transaction's screening, we now:

  • Write an EngineDegradationEvent row with the tenant, engine, reason, and the transaction ID
  • Stamp the verdict with a SANCTIONS_SCREENING_DEGRADED action, so investigators see it on the transaction itself
  • Log a warning with full context

This is the evidence trail. You can answer the examiner's question: "which transactions were screened with a degraded check?"

2. Recovery: automatic catch-up re-screening

A cron runs every 15 minutes. When the engine recovers, it:

  1. Finds every unresolved degradation event with a transaction link
  2. Re-runs the sanctions screening on the stored transaction data
  3. A confirmed match (≥95% confidence) creates a CRITICAL SANCTIONS_HIT case immediately, tagged retroactive
  4. A clear result marks the event resolved

Design rule: the original verdict is never rewritten. It is immutable evidence. The case is the outcome.

3. Visibility: alert the humans

Five or more fail-open events in 15 minutes triggers an email to the tenant's compliance officers. The email says what's happening, what it means, and what to do. The dashboard shows an amber banner when engines have degraded recently or a re-screen backlog exists.

The design constraints

Everything is fire-and-forget from the screening path's perspective. Integrity bookkeeping must never add latency or failure modes to the real-time verdict. A degradation event that fails to persist is logged at debug level; the transaction proceeds normally.

The re-screening cron is batched (100 per run) so it can't pin the event loop. Events older than 30 days without re-screening are expired, since the transaction may have been pruned by retention.

A real incident

During a production deploy, our sanctions engine OOM-killed repeatedly. In the old world, this would have been invisible: transactions pass, nobody knows. Instead:

  1. Every screened transaction was flagged with SANCTIONS_SCREENING_DEGRADED
  2. The dashboard showed the amber banner within 30 seconds
  3. The compliance team was alerted by email within 15 minutes
  4. When the engine recovered, the cron re-screened all affected transactions automatically
  5. The banner faded as events resolved

The remediation evidence, "we identified the gap, re-screened N transactions, opened M cases", was generated automatically.

Why this matters for procurement

When a bank evaluates AML platforms, they should ask: "what happens when your sanctions engine goes down?" Most platforms will say "we fail open", which is correct. The follow-up question is: "how do you know which transactions were affected, and what do you do about it?"

If the answer is "we don't" or "we check the logs," that's not good enough for a CBN examination. Our answer is a table, a cron, and an audit trail.

See a live trace on your own data

We'll run the tracer against a scenario from your institution in a 30-minute demo: blocked transaction, full fund trace, frozen destination, packaged case.