The setup
Finhaq is a multi-tenant AML/CFT compliance platform for Nigerian banks. Every transaction that flows through a bank on our platform gets screened in real time against CBN-mandated rules, sanctions lists, behavioral signals, and custom bank rules. The screening pipeline produces a verdict: APPROVE, REVIEW, ESCALATE, or BLOCK. The bank's core banking system acts on it at authorization time.
The architecture was clean. Multiple engines (a rules engine, a sanctions screener, a behavioral engine, and an external decision engine) each evaluate the transaction independently. A verdict aggregator then combines them into one score using weighted averaging with highest-severity-wins logic.
The external decision engine was a third-party fraud/AML engine. We integrated it early because it promised scenario management, aggregation windows, and a proper decision framework.
The bug
The version we were running had a known issue: committing scenario iterations panicked (a runtime panic in CommitScenarioIterationVersion). Without committed scenarios, the /decisions endpoint always returned 400. So our connector worked around it:
- Ingestion: we POSTed transaction data to the engine for its audit trail (this worked)
- Evaluation: we evaluated the rules locally in TypeScript, interpreting the same AST formulas the engine would have computed
Here is the problem: our CustomRulesEngine already evaluated those same custom rules. The decision-engine connector re-evaluated them. Both verdicts entered the aggregation with separate weights:
ENGINE_WEIGHTS = {
CBN_RULES: 1.5,
CUSTOM_RULES: 1.2, // first evaluation
DECISION_ENGINE: 1.0, // same rules, second evaluation
}
Every triggered custom rule was counted twice in the weighted average.
The second bug (worse)
The duplicate used its own outcome thresholds: >=50 ESCALATE, >=70 BLOCK. Our platform's canonical mapping: >=60 ESCALATE, >=80 BLOCK.
Because aggregation is highest-severity-wins, a custom rule set summing to 50 to 59 got ESCALATE from the "decision engine" when the platform said REVIEW. At 70 to 79, it got BLOCK when the platform said ESCALATE.
The core problem: production verdicts were being escalated by a duplicate of our own engine wearing stricter thresholds.
How we found it
A manual code review during a broader platform assessment. The connector looked correct in isolation. It ingested data, evaluated rules, and returned results. The double-counting only became obvious when you traced a single custom rule through both code paths and saw it contribute to two separate engine verdicts.
The fix
We removed the decision-engine connector, bootstrap, and rule-sync entirely. Custom rules are evaluated exactly once by the CustomRulesEngine. We added a regression test that asserts ENGINE_WEIGHTS contains no two engines that derive from the same rule source, and that a CUSTOM_RULES score of 50 stays REVIEW.
We also white-labeled everything from the start. Bank users see "Regulatory Compliance" and "Behavioral Analysis", never engine vendor names. So the retirement had zero customer-facing impact.
Lessons
- Workarounds compound. The upstream bug forced a local-evaluation workaround. The workaround created a second evaluation path. The second path double-counted. Each step looked reasonable in isolation.
- Weighted averaging amplifies duplicates silently. If two engines evaluate the same signal and both have weights, the signal is over-represented. There is no error and no log line. Just a subtly wrong number.
- Outcome thresholds must live in one place. The platform's canonical mapping (60/80) and the connector's mapping (50/70) diverged, and highest-severity-wins let the stricter set flip final verdicts.
- If you are going to own the decision, own it completely. The marketing pitch "we run a real decision engine" was technically true, since we were running its ingestion pipeline. But the decisions were ours all along. Better to admit that and make it excellent than to maintain the facade.
See a live trace on your own data
We'll run the tracer against a scenario from your institution in a 30-minute demo: blocked transaction, full fund trace, frozen destination, packaged case.