Skip to content
Sagar Thakkar
← All writing
Architecture 12 min read
architecture lending human-in-the-loop audit model-risk ai-architecture

Approve and Deny Are the Easy Part: Loan Decisioning With a Human in the Loop

Approve and deny are the easy majority of loan volume. The architecture lives in the exceptions: classify, route, human review, re-entry, audit.

Sagar Thakkar
Sagar Thakkar
AI Systems Architect
Poster reading APPROVE. DENY. THEN THE HARD ONES: a thick flow splitting into two resolved branches, with a thin terracotta third branch that stops partway, unresolved
TL;DR

A loan decisioning model that outputs only approve or deny has nowhere to put the applications that arrive incomplete or contradictory. Those exceptions are where the architecture lives. Classify, route four ways, send exceptions to a reviewer, let corrected cases re-enter the pipeline, and append every transition to a record store.

For architects putting AI in front of loan applications, where the design problem is not the decision. It is everything the decision cannot resolve.


TL;DR

  • A model that outputs approve or deny handles the easy majority of the volume and lies about the rest. Real applications also arrive with conflicting data and missing fields.
  • The reframe: this is not a classifier, it is a stateful workflow. Classify, route four ways, hand exceptions to a reviewer, let corrected cases re-enter, append every transition to a record store.
  • The load-bearing constraint is lifecycle integrity, not auditability. An audit store bolts onto a binary classifier later. A four-outcome lifecycle does not.
  • What breaks first at ten times the volume is reviewer capacity, not compute. Workers scale by configuration. Reviewers scale by hiring.
  • For: architects and lending technology leads who must add speed and still explain a refusal.

Context

Picture a regulated non-bank lender. Underwriting is the bottleneck, and the queue is cleared by people reading applications one at a time. The ask is obvious: let a model decide the clear cases, so human attention lands only where needed. The awkward part surfaces in the first design session, when somebody asks what happens to an application whose stated income contradicts its bank statement.

A word on what is real here and what is not. The core of this pattern was built live, inside a time-boxed client engagement, for a regulated non-bank lender, with the client’s own underwriting team in the loop. That is the whole provenance claim, and it identifies nobody. Everywhere a detail would identify them it is withheld: the volumes, the outcome mix, the systems, the timing. No figure here is theirs.

What was built is the core: the four-way router, the reviewer queue, the re-entry path and the versioned append-only record. The hardening that follows, on erasure, abuse, fair lending and recovery, is what I would require before production volume. It is design, not a record of the build. Every number below is either arithmetic you redo with your own inputs, or a target, and each says which.


Four Outcomes, Not Two

Sit with a real queue and the two-outcome model falls apart. Alongside approved and denied are the two outcomes that consume underwriting time.

Mismatch. The data conflicts with itself or fails a policy rule. A declared employer does not match the salary credit, or a document says one thing and a field another. Not a decision. A correction waiting for somebody to make it.

Query. The application is short of something. A missing document, an unreadable upload, a field nobody filled in. Not a decision. A question to be answered first.

Neither is a failure state, and that is the point. Both must leave the pipeline, be touched by a human, and re-enter. Accept that, and the system stops being a model with an API in front. It becomes a state machine with a model inside.


Alternatives Considered

Five properties decide this, unequally. Explainability. An immutable record of what decided what. A real human path for ambiguity. Re-entry integrity. Speed on the clear majority.

PropertyA. Binary classifierB. Fully manual, model as advisory scoreC. Classify, route, review, re-enter
Explainable refusalPartial, a score is not a reasonYes, the reviewer writes itYes, reason codes plus the feature payload
Immutable decision recordRetrofittableYes, in whatever the ops tool logsYes, append only by construction
Human path for ambiguityNone, ambiguity is forced into a decisionEvery case, including the clear onesExceptions only
Re-entry integrityNot modelledManual, and invisible outside the reviewer’s memoryExplicit state, history carried forward
Speed on the clear casesYesNo, the bottleneck is untouchedYes

A is rejected because it has no home for mismatch or query. A missing payslip is not a marginal approve, and scoring it as one produces a wrong answer on the riskiest cases.

B is rejected because the bottleneck is untouched. Reviewers still read every application, and the model buys a suggestion nobody may act on.

C is chosen. It is more machinery, and the machinery is the deliverable: queue, worker, router, reviewer console, store, re-entry path.

A binary classifier is not a simpler version of the right architecture, it is the wrong one. Choosing it means discovering the exception path in production rather than in design.


Why the Exception Path Is the Whole Design

Anybody can build approve and deny. Labels and a threshold are enough. Everything expensive here exists because of mismatch and query.

They force a state machine, because an application awaiting information is in a state, not at a decision. They force a reviewer console, because somebody must see the conflict, the evidence and the history in one place. They force a re-entry path, so a corrected case re-flows without losing the first pass. They force a richer record, because the audit question is not what the model said. It is what the model said, what the reviewer then did, and why.

So the exception path is not the edge case around the design. It is the design.

This also names the load-bearing constraint. The obvious answer is auditability, and it is wrong. An append-only store bolts onto a binary classifier a year later, painfully but possibly. A four-outcome lifecycle does not. The states, the re-entry semantics and the reviewer’s authority over an outcome all have to exist before the first decision is written. Lifecycle integrity eliminates options. Explainability and immutability are obligations, and a correctly shaped system satisfies them.


Architecture

Loan decisioning pipeline: queued intake into a classifier worker, a four-way decision router, terminal approve and deny branches, an exception branch into the reviewer queue and ops console, a re-entry path back to the queue, and every transition appended to a decision record store

The pieces:

  • Intake queue. Intake never blocks on inference. That choice is why latency does not bind later.
  • Classifier worker. Emits one of four outcomes, reason codes, and a reference to the feature payload behind them. That reference is the explainability artifact.
  • Decision router. The four-way branch. Approve and deny terminate, mismatch and query divert.
  • Reviewer queue and ops console. The only path allowed to change an ambiguous outcome.
  • Re-entry path. A corrected application returns to the queue carrying its prior states, reason codes and the correction.
  • Decision record store. Append only. Every transition written once, stamped with model version, policy version and actor.

The lifecycle, as a state machine, is the artifact worth pinning on the wall:

RECEIVED -> QUEUED -> PROCESSING -> approved  -> APPROVED  -> DISBURSED
                                 -> denied    -> DENIED    -> CLOSED
                                 -> mismatch  -> UNDER_REVIEW  -> corrected -> QUEUED
                                 -> query     -> AWAITING_INFO -> corrected -> QUEUED

Two rules keep this honest. A state change and its record commit in one transaction, through an outbox, so the record is durable before anything downstream reacts. And only the reviewer console writes a terminal outcome onto an application sent to review.


Does It Hold at Scale, and What Breaks First

Start with what does not bind. Intake is queued and decisions are asynchronous, so no user waits on a first token. A classifier outage delays decisions and loses none, and the backlog drains as workers restore. Throughput is a configuration question: more workers, more partitions.

So do not state one blended latency budget. State two objectives, because different resources meet them. Clear-case time to decision, a target of two minutes at the ninety-fifth percentile, met by workers. Exception time to decision, a target of one working day, met by reviewers. Both are targets, not measurements.

The difficulty is people. Size the reviewer function first. It is the only part nobody provisions in an afternoon.

reviewers = peak-day applications × exception share × (1 + re-entry rate) × minutes per review ÷ productive reviewer minutes per day

Put round numbers through it so the arithmetic is checkable. None are measurements from any lender. They are inputs chosen to be easy to redo. Take two thousand applications on a peak day, an exception share of one in ten. Add a re-entry rate of one in five of those, twelve minutes of reviewer time per case, and three hundred productive reviewer minutes in a working day. That gives two hundred and forty reviews a day, two thousand eight hundred and eighty reviewer minutes, and roughly ten reviewers.

Now multiply the volume by ten. Workers scale by configuration. The reviewer function goes from ten people to about a hundred: a hiring plan, a training curriculum, a quality programme. That is what breaks first, and notice where. On the exception path. The thesis of this piece and its capacity model are one claim.

Two levers move that number. A lower exception share, a data-quality problem upstream of the model. And fewer minutes per review, a console design problem. Neither is a model problem.

The record store. Size a record before sizing the store. It holds reason codes, version stamps, actor, prior and next state, and a reference to the feature payload rather than the payload itself. Call it a few kilobytes. Count transitions from the state machine rather than guessing: the clean path writes four, an exception path more, so call it five. At the illustrative volume that is under four million records a year. Unremarkable. At ten times, partitioning and audit-path query performance become design decisions rather than defaults. Choose tiering on day one, because migrating an append-only store later means proving nothing changed.

Retention is an input, not a number I supply. It comes from the lender’s own record-keeping policy, and it runs longer than most teams expect. Architecturally, append-only storage and a right to erasure reconcile only through separation. Keep the decision record free of personal data. Keep the personal data in an encrypted store with a key per applicant. On an erasure request, destroy the key. The decision, its reason codes and its version stamps survive, holding nothing they have no right to hold.

RPO and RTO, stated rather than implied. RPO of zero for the decision record store, which is honest only because the state change and its record commit in one transaction. Best effort for everything else. RTO is a target, and a stated target of thirty minutes for the decision path is defensible because intake keeps queueing while the path is down. A failover exercised only on a diagram is not a failover, so it gets a quarterly drill.


The Security and Abuse Surface

Re-entry is a feature and an attack surface at once. A corrected application is a second attempt, and a second attempt is a probe. Submit, observe, adjust the figure, resubmit. Enough iterations and somebody is mapping the threshold.

Three controls, none exotic. Re-entry carries history, so the reviewer sees a field-level diff between attempts rather than a fresh-looking application. Velocity limits apply per applicant and per intermediary, because the interesting abuse concentrates there. And a re-entry count above a policy threshold becomes its own review reason, not a silent retry.

Reviewer override is privileged. Least privilege on who resolves which outcome, four-eyes review above a policy threshold, the reviewer identity on every record. An override without an actor is an anonymous decision.

One deliberate rejection, because it is the fashionable option. An autonomous agent that reads the file and decides is not a candidate. Its justification would be generated narrative rather than the features that drove the outcome, and narrative is not an adverse-action reason. The agentic pattern that earns a place is narrower. Drafting the information request for a query case, from the reason codes, for a reviewer to approve before it goes out. Generation in service of the human step, never in place of the decision.


Fair Lending and Model Risk

This is the sharp regulatory edge, sharper than the audit story. Credit scoring of individuals sits in the high-risk tier of the EU AI Act. Refusal regimes elsewhere demand a stated reason the applicant is able to act on.

What the architecture owes it:

  • Reason codes derived from the feature payload, never written after the fact. A reason composed to sound acceptable is a liability wearing a helpful tone.
  • Disparate outcome monitoring by segment, on approval, denial and exception rates. Watch proxy features: a postcode is a location field to an engineer and something else to a fair-lending reviewer.
  • Population stability and drift monitoring on the input distribution. A frozen model against a shifting applicant mix degrades quietly, and shows nothing on an uptime dashboard.
  • Shadow mode, then champion and challenger, on every model and policy change. This works only because every decision carries its model and policy version. The audit store paying for itself.

The human in the loop is a control and also a risk. Reviewers drift toward agreeing with the model, which is automation bias, and it turns a safeguard into a rubber stamp while every metric stays green. So the override rate by reviewer and by segment is monitored exactly as the model is. A reviewer who never disagrees is a finding. The oversight layer needs its own oversight.


When Not to Use This Pattern

  • Unregulated, low-stakes decisions. If nobody audits you and a wrong call is cheap, the record store and the reviewer path are overhead.
  • Volume a person clears comfortably. Below the point where a queue forms, the manual option is correct.
  • No reviewer capacity to fund. Without staffed reviewers, exceptions accumulate and the system becomes worse than the process it replaced.
  • Genuinely binary decisions. Rare in lending. If you have one, the four-outcome machinery is dead weight.

The Takeaway

The decision is the easy part. Build for what the decision cannot resolve, size the reviewer function before the compute, and make every transition write itself down once and never again. That system is faster on the clear cases and defensible on the hard ones. The pair is what the project was funded for, even when the brief named only the first.


About the Author

I architect production AI systems for regulated enterprises, with delivery experience in banking, lending and logistics. Focus on design that survives audit.

→ If you are standing one of these up: book a 30-minute architecture review. No pitch. Bring the constraint that is blocking you and we will work out whether this pattern fits or whether something else does.

→ More case studies: the blog

Newsletter

New essays, straight to your inbox.

Occasional, in-depth writing on distributed systems, AI agent architecture, and engineering leadership. No spam, unsubscribe anytime.