Last Tuesday at 14:08 WAT, a salary account in Lagos sent ₦850,000 to a vendor it pays every month. Same device as always, same IP range, midday, an amount in line with six months of history. Our pipeline approved it in 142 milliseconds. Nobody noticed, which is exactly the point: the best verdict is the one a legitimate customer never feels.
What is less obvious is how much happened in those 142 milliseconds. That transfer was checked against regulatory thresholds, a custom rule set written by the institution's own compliance team, global sanctions lists, the account's KYC state, and a behavioral model of how that specific customer normally behaves. Six engines, each with its own data and its own failure modes, all inside a window shorter than a blink.
This post opens that window: the pipeline shape, the latency budget, the verdict format, and the failure design that keeps payments moving when a dependency does not.
The pipeline: normalize once, then fan out
The screening pipeline has four stages, and only one of them is expensive:
- Ingest. The transaction event arrives over the API, is authenticated, schema-validated, and assigned a monotonic event id. About 8ms.
- Normalization. We resolve the event into a canonical context object: sender and receiver accounts, KYC tier, device fingerprint, IP classification (residential, mobile, datacenter, VPN), geo, historical velocity counters. Every engine reads this same object, so nothing is computed twice. About 11ms.
- Parallel fan-out. The context is dispatched to five scoring engines concurrently: Regulatory Compliance, Custom Rules, Global Sanctions, KYC Verification, and Behavioral Analysis. The stage takes as long as its slowest engine, not the sum of them.
- Aggregation and verdict. The Decision Engine, the sixth engine, consumes the five results, applies the institution's policy weights, and emits a single verdict object. That verdict is persisted and returned. About 37ms combined.
The reason for parallel fan-out is arithmetic. Run sequentially, the five scoring engines cost roughly 230ms on their own, which already blows the budget before aggregation. Run concurrently, the stage costs the maximum of the five, not the sum. Our p50 lands at 142ms; p99 stays under 200ms. Here is where the p50 budget actually goes:
-- p50 latency budget, end to end: 142ms
ingest + auth 8ms
normalization 11ms
fan-out dispatch 4ms
-- parallel stage (wall clock = slowest engine)
regulatory_compliance 26ms # thresholds, reporting limits
custom_rules 21ms # institution-authored predicates
global_sanctions 58ms # fuzzy match over cached lists
kyc_verification 33ms # tier, document state, expiry
behavioral_analysis 82ms # slowest engine, sets the stage
decision_engine 27ms # policy weights, verdict assembly
persist + respond 10ms
total 142ms
Two things worth noting in that table. First, the parallel stage costs 82ms because Behavioral Analysis costs 82ms. If we make sanctions matching 20ms faster tomorrow, the verdict does not get faster. Second, Behavioral Analysis is the slowest engine by design: it is the only one that reads per-account history, and we chose to spend the budget there because behavior is where the signal lives.
The verdict is a structured object, not a score blob
A fraud verdict that says "score: 0.87" is useless to a compliance officer. It cannot be audited, it cannot be appealed, and it cannot be explained to a regulator. So every Finhaq verdict is a structured record of which engines ran, which rules fired, and what each one contributed to the outcome. A simplified approve verdict from the opening example:
{
"transaction_id": "txn_9f27c1a4",
"verdict": "APPROVE",
"trust_score": 88,
"latency_ms": 142,
"engines": [
{ "engine": "regulatory_compliance", "verdict": "PASS",
"rules_triggered": [], "score": 100 },
{ "engine": "custom_rules", "verdict": "PASS",
"rules_triggered": ["recurring_vendor_whitelist"], "score": 95 },
{ "engine": "global_sanctions", "verdict": "PASS",
"rules_triggered": [], "score": 100,
"list_snapshot": "2026-09-27T12:00Z" },
{ "engine": "kyc_verification", "verdict": "PASS",
"rules_triggered": [], "score": 100 },
{ "engine": "behavioral_analysis", "verdict": "PASS",
"rules_triggered": [], "score": 91 }
]
}
Contrast that with the verdict from the incident we described in part 1: a 03:12 WAT transfer of ₦4,250,000 from a new device on a datacenter IP. Same schema, but the behavioral engine came back with rules_triggered listing new_device_in_session, datacenter_ip, and night_time_request, the aggregate trust score collapsed to 18, below the auto-blacklist threshold of 20, and the verdict flipped to BLOCK in 155ms. The schema does not change with the outcome. Only the contents do, and every content is a fact a human can read.
Design rule: every rule that fires is recorded on the verdict with its contribution to the score. If we cannot point to the specific rule and the specific fact that triggered it, we do not act on it. Explainability is not a reporting feature layered on top; it is the format the decision is made in.
Fail-open, but never silently
Here is the uncomfortable question every screening vendor has to answer: what happens when a dependency is down? If your sanctions provider times out at 14:00 on a business day, do you stop approving payments?
Many platforms fail closed: no answer from the provider means no verdict, which means every payment queues or declines. That design converts a vendor outage into your outage. We refuse it. Finhaq fails open, with two hard constraints that make fail-open defensible to a compliance team.
First, there are no synchronous third-party calls in the decision path. Sanctions lists, IP intelligence, and device reputation data are all pre-materialized locally. A background refresher pulls updated lists every few minutes and publishes versioned snapshots to the engines. The hot path reads the local snapshot. A provider can be down for an hour and the only consequence is that we are screening against a snapshot that is sixty minutes old, which we know precisely, because the snapshot timestamp is stamped on every verdict, as in the JSON above.
Second, degradation is explicit, never silent. Every engine sits behind a circuit breaker with a 120ms trip budget, well inside the p99 envelope. If an engine breaks, say Behavioral Analysis hits a storage fault, the Decision Engine does not skip it and move on. It issues the verdict as APPROVE_WITH_REVIEW, records the degraded engine and the reason on the verdict object, and queues the transaction for a human case officer with the missing checks listed as outstanding work. The customer's payment completes. The compliance team sees exactly which check did not run, on exactly which transactions, and re-runs them when the engine recovers.
That distinction is the difference between fail-open as an engineering decision and fail-open as a compliance liability. Silent skipping means your audit trail claims a check happened when it did not. Approve-with-review-flag means your audit trail tells the truth: payment approved, sanctions snapshot from 12:00 UTC, behavioral engine degraded for 4 minutes, 37 transactions queued for re-screening, all re-screened clean by 14:11. An examiner can work with that. They cannot work with a gap in the log.
Custom rules: plain English in, predicates out
The Custom Rules engine exists because no vendor knows an institution's risk appetite better than its own compliance team. The problem is that compliance teams write requirements in prose, and engines execute predicates. We bridge that gap with a compiler, not a ticket queue.
An analyst writes: "Flag any transfer above ₦2,000,000 to an account less than 30 days old, unless the sender has paid that account before." The system drafts a compiled predicate against the normalized context object: amount, receiver account age, prior edge existence. The draft is shown to the analyst side by side with the English, with examples of historical transactions that would and would not have fired. A human activates it. From that point it executes in microseconds, versioned, with every activation and every firing recorded.
Two boundaries matter here. The AI drafts; it never activates. A rule that can block money requires a named human to turn it on. And the compiled predicate, not the prose, is what executes, so there is no model in the hot path interpreting policy at request time. The prose is documentation; the predicate is law.
What we deliberately do not do
The fastest way to blow a 142ms budget is to do something clever inside it. So the decision path has a short list of hard exclusions:
- No synchronous external calls. Not for sanctions, not for IP intelligence, not for device reputation. Everything the engines read is pre-materialized locally and refreshed asynchronously.
- No joins at request time. Normalization reads from pre-computed context stores, not from the transactional database. The pipeline never waits on a system of record.
- No unbounded work. Every engine has a trip budget and a circuit breaker. The pipeline degrades by policy, not by accident.
- No hidden scoring. If a contribution to the verdict cannot be named as a rule and a fact, it does not get a vote.
These exclusions are why p99 stays under 200ms even on bad days. Tail latency in screening systems almost never comes from the engines themselves; it comes from waiting on something outside the process. We removed the waiting.
What a verdict cannot tell you
A verdict answers one question well: should this transaction happen? It says nothing about the money that is already moving. The verdict that blocked the ₦4,250,000 transfer at 03:12 WAT was correct and complete in 155 milliseconds, and it still did not tell us that the sending account had been fed by a victim two hops up, or that ₦4.1 million was sitting at hop 3 waiting to cash out. That answer came from the tracing system, which is a different machine built for a different question. We wrote about how it works in part 1.
Next in the series: the signals the behavioral engine actually consumes, and how account takeover looks at the application layer before any money moves. Read part 3.
See a verdict rendered live
We'll run your scenarios through the six-engine pipeline in a 30-minute demo: clean transactions, hostile transactions, and a degraded engine, with the full verdict object for each.