Bug #1: The column that wasn't there
The code: a new behavioral_events table with a column for the event type.
The bug: the Prisma schema mapped eventType without @map("event_type"), so the ORM looked for a camelCase column while the hand-written migration created snake_case.
PrismaClientKnownRequestError P2022:
The column `eventType` does not exist in the current database.
Every POST /behavioral/events in production returned 500.
Why local tests passed: our test setups used prisma db push (which builds columns from the schema definition). Production used prisma migrate deploy (which runs the hand-written SQL). The two diverged on exactly this field.
Bug #2: Express on a Fastify app
The code: a DSAR export endpoint that set download headers.
The bug: the handler typed the reply as express.Response and called res.setHeader(), but the platform runs on Fastify, whose reply uses .header().
TypeError: res.setHeader is not a function
The export endpoint 500'd on every call.
Why local tests passed: our boot-smoke test checked that the app started and the health endpoint responded. It never made a real request to the new endpoint.
The pattern
Both bugs share the same DNA:
- The verification environment differed from production in one specific way (db push vs migrate deploy; boot-only vs real HTTP requests)
- The code compiled and all unit tests passed (TypeScript was happy; the schema types matched what the ORM client expected)
- The bug only manifested at runtime, on the first real use of the new endpoint
The fix: a CI pipeline that would have caught both
We rebuilt our CI boot-smoke job with two changes:
1. Build the test database with migrate deploy
- name: Push schema to disposable database
run: npx prisma migrate deploy --schema=packages/db/prisma/schema.prisma
Not db push. If the migration SQL and the Prisma schema disagree, migrate deploy creates the column the way production will, and the first query that expects a different name fails immediately.
2. Make one real HTTP request per new endpoint
- name: Boot the API (must reach listen)
run: |
node dist/main.js > boot.log 2>&1 &
APP_PID=$!
# ... health check, auth, then:
curl -sf -X POST .../behavioral/events -d '...' # must return 2xx
curl -sf .../dsar/$ID/export # must return 200 + PDF
If the handler crashes on the first real call (Express API on Fastify, undefined method, wrong column name), the CI job fails. The developer sees the error in the PR, not in production logs.
The meta-lesson
We'd been burned by this class of bug twice in one month. The first time, the fix was "add a @map annotation." The second time, the fix was "use the right framework's reply API." Both were one-line fixes that took hours to diagnose because the failure was in production, not in development.
The systemic fix wasn't "be more careful with @map annotations." It was to make the test environment match production in the specific ways that matter.
- Same migration path (
migrate deploy, notdb push) - Same HTTP framework (real requests, not just "the process started")
- Same runtime environment (Linux container, not a host Node process)
What our CI now covers
✅ npm ci (lockfile integrity, catches dropped transitives)
✅ Prisma generate + workspace library builds
✅ 394/394 unit tests (pure logic: algorithms, validation, patterns)
✅ nest build (TypeScript compilation)
✅ next build (dashboard compilation)
✅ Boot smoke (disposable Postgres/Redis, migrate deploy, real HTTP
requests to health + behavioral events, must reach `listen`)
✅ Docker Compose file validation (all 5 compose files parse)
The boot-smoke job costs ~5 minutes in CI. The two production incidents it would have prevented cost ~6 hours of diagnosis and debugging, plus the credibility cost of a 500 on a customer-facing endpoint.
See a live trace on your own data
We'll run the tracer against a scenario from your institution in a 30-minute demo: blocked transaction, full fund trace, frozen destination, packaged case.