Real-time transaction monitoring means the system returns a verdict before the transaction completes. The payment enters your rails, the screening engine evaluates it inline, and a decision to approve, hold, or block comes back in a few hundred milliseconds, while the money is still inside your system. Batch screening works the other way: transactions settle first, and the monitoring runs afterwards, usually overnight. By the time a batch job flags a suspicious transfer, the funds have cleared, been split across mule accounts, and left the institution.

Vendors blur this line because "real-time" sells and "batch" does not. The difference is not cosmetic. It decides whether your compliance program blocks fraud or only reports it.

This post defines the term properly, walks through where batch screening breaks down in practice, shows what a real latency budget looks like, and ends with the questions to put to any vendor that claims real-time capability.

What real-time actually means

A monitoring system is real-time when it sits inline in the payment path. The core banking system or payment switch calls the screening API as part of processing the transaction, waits for the answer, and acts on it: approve, hold for review, or block. If the screening verdict cannot change the outcome of the transaction, the system is not real-time, whatever the sales deck says.

Inline operation comes with a hard constraint: the latency budget. Payment rails have their own timeouts. If screening takes longer than the rail will wait, either the payment times out or the screen gets skipped. A real-time system therefore commits to a ceiling, measured in milliseconds, and is engineered to stay under it at p95 and p99, not just on average.

In the market, "real-time" currently means three different things:

  • True inline verdict. Synchronous screening inside the payment path, verdict in milliseconds, transactions can be blocked before settlement. This is the only version that prevents loss.
  • Near-real-time alerting. Transactions post immediately; the monitoring system consumes them from a stream and raises alerts within seconds or minutes. Detection is fast, but the money has already moved. Useful for investigation, useless for interdiction.
  • Intraday batch, renamed. Files processed a few times a day instead of once overnight. The marketing improved; the architecture did not.

The one-line test: ask the vendor whether their system can stop a transaction before it settles. If the answer involves a queue, a webhook, or the phrase "within minutes," it is not real-time.

Why batch screening fails

Batch screening was designed for a slower era of payments. Against instant transfers and organized mule networks, it fails in four predictable ways.

The money is unrecoverable by the time you look

A fraudulent transfer that settles at 14:00 is layered through three accounts by 14:30 and withdrawn or moved off-platform by close of business. A batch job that runs at 02:00 is writing a report about money that left eleven hours earlier. Recall requests through the receiving institution rarely succeed once funds have moved through a mule chain. Detection after settlement is documentation, not prevention.

The 24-hour STR clock is consumed before anyone sees the alert

Under the Money Laundering (Prevention and Prohibition) Act 2022, a Suspicious Transaction Report must reach the NFIU within 24 hours of suspicion being formed. An overnight batch produces its alerts the next morning. The queue then waits for an analyst shift, a review, and an escalation decision. By the time compliance forms the suspicion, half the statutory window can be gone, and the filing itself still has to be assembled. The mechanics of that deadline are covered in our guide to filing STRs with the NFIU.

Mule accounts empty before the morning batch

Mule networks are built around the assumption that institutions check after the fact. Funds arrive, split, and exit within hours, and the account is often abandoned after a single pass. A batch screen catches the pattern reliably, and catches it reliably too late. The account it flags at 02:00 was drained at 19:00 the previous evening.

Examiners ask why a blockable transaction was only reported

CBN AML/CFT examinations increasingly look at what the institution's systems could have done, not only what staff did. When the transaction record shows a blatant pattern that settled unchallenged and surfaced only in a report days later, the examiner's question is straightforward: you had the data at the moment of the transaction, so why did the money move? "Our screening runs overnight" is a description of the problem, not an answer to it.

What real-time monitoring requires

Genuine inline screening is an engineering commitment, not a feature flag. Four things separate a real system from a renamed batch job.

A measured latency budget

The vendor should state a number and let you verify it. At Finhaq, six screening engines evaluate every transaction and return a verdict in 142ms, inside a 200ms latency budget. The full decomposition of that window is in Anatomy of a 142ms verdict. If a vendor cannot quote a p95 latency measured on production traffic, they are describing an aspiration.

Fail-open integrity, with automatic re-screening

Engines have dependencies, and dependencies fail. The dangerous design is one where an outage means transactions sail through unscreened, silently. The correct design fails open to keep payments moving, records every degradation, re-screens every affected transaction automatically when the engine recovers, and alerts compliance the moment an engine drops. An outage should never mean unscreened money, only money screened a few minutes late with a full paper trail. How that works in practice is covered in Screening integrity: never lose a transaction.

Explainable verdicts

A blocked customer deserves a reason, and so does the examiner who reviews the block six months later. Every verdict should decompose rule by rule: which engine fired, which threshold, which list entry. This is why Finhaq uses explainable rules rather than black-box ML, and it aligns with CBN's expectations on AI/ML governance. "The model said so" is not a defensible answer when a legitimate customer is standing in a branch asking why their transfer failed.

Behavioral context scored inline

Velocity rules alone cannot tell a salary payment from a mule pass-through of the same amount. The context that separates them (device fingerprint, IP, session behavior, and how this account normally transacts) has to be scored inside the same millisecond window, not appended afterwards. How behavioral signals work without opaque models is the subject of Behavioral intelligence without black-box ML.

The latency budget

The latency budget is where real-time claims become testable. Monitoring time must fit inside the payment rail's own timeout, with headroom for network travel and retries. Here is a sample 200ms budget, broken into the components that consume it:

ComponentBudgetWhat happens inside it
Data enrichment ~30ms Load account state, KYC tier, balance, and recent transaction history for sender and receiver.
Rules engine ~25ms Regulatory thresholds (CTR limits of ₦5m individual / ₦10m corporate), velocity rules, and the institution's custom rule set.
Watchlist screening ~40ms Sanctions and watchlist name matching against parties and counterparties.
Behavioral scoring ~30ms Device, IP, session, and account-pattern signals scored against the customer's own history.
Decision and logging ~17ms Combine engine outputs into a verdict, write the immutable audit record, return the response.
Total 142ms Leaves 58ms of headroom inside a 200ms budget for network overhead and rail timeouts.

Two things to notice. First, the budget is measured per component, so a regression in one engine is visible before it eats the headroom. Second, the figure that matters is p95, not the average. An average of 140ms with a p95 of 900ms means one transaction in twenty stalls the payment rail, and customers feel exactly that tail.

Shadow mode: the safe path to real-time

No institution should flip a screening system from nothing to blocking live traffic in one step. The safe path is shadow mode: the new system sits inline on live traffic and computes real verdicts, but those verdicts are logged rather than enforced. Payments flow exactly as before. The institution gets production-grade latency data from day one, without production-grade risk.

Over the following weeks, shadow mode gives you the two numbers that matter: catch rate and false positive rate, both measured against your existing batch baseline. Every transaction the shadow system would have blocked is compared against what the batch process actually caught. Every would-have-blocked verdict gets a human review. Rules are tuned until the false positive rate is one your operations team can live with.

Only then do you flip to blocking, and even then, gradually. Clear cases block automatically. Gray-zone verdicts route to humans with the full decomposition attached. In Finhaq's deployment model, automated interdiction is reserved for the strongest signals: accounts whose behavioral trust score falls below 20 freeze automatically, with humans in the loop on everything less certain.

Questions to ask any vendor that claims real-time

  • What is your measured p95 verdict latency on production traffic? Not the demo number, and not the average.
  • Is screening inline in the payment path, or asynchronous after posting? Async means the money settles first.
  • What happens when a screening engine goes down? If transactions pass unscreened, you carry the risk. Ask about re-screening and alerting.
  • Can I see a blocked verdict decomposed, rule by rule? If not, you cannot explain the block to the customer or the examiner.
  • Can we start in shadow mode and measure against our batch baseline before blocking? A vendor confident in their numbers will say yes.

These five are a starting point. The full evaluation list is in How to choose an AML solution: 14 questions that expose weak vendors.

Where Finhaq fits

Finhaq was built inline from the start. Six screening engines return an explainable verdict in 142ms inside a 200ms budget, engines fail open with automatic re-screening so an outage never means unscreened money, and every alert and decision lands in an immutable audit trail. At Buildbank MFB, that architecture blocked over ₦97 million in suspicious value with zero false negatives, money that a batch screen would only have reported after it left.

If you are evaluating monitoring systems and want to know where your current setup would fail this test, the free AML readiness assessment scores it in three minutes.


This article is general guidance for compliance and product professionals, not legal advice. Regulatory obligations referenced here, including STR deadlines and monitoring expectations, are set by the Money Laundering (Prevention and Prohibition) Act 2022, CBN AML/CFT regulations, and NFIU directives. Always check the current instruments and your regulator's circulars.

Watch a verdict come back in 142ms

A 30-minute demo on your scenarios: inline screening, shadow mode, and verdicts you can explain to an examiner.