Hetu / Industries / Financial services
AI decision verification for lending, underwriting & collections
Hetu helps banks, NBFCs, and insurers put AI agents on real credit decisions — with deterministic checks, causal attribution, and confidence labels model risk committees can accept. For heads of AI, CROs, and model risk teams where the blocker on autonomy was never capability. It was accountability.
The model risk committee
The demo was excellent. The agent read the portfolio, found the delinquency spike, and wrote a paragraph explaining it that was clear, specific, and confident.
It was also wrong. And there was no way to tell that from the output — the wrong answer and the right answer are rendered in the same font, with the same certainty, at the same speed.
So the committee asked the only question that matters: when it is wrong, how will we know, and what will it have done by then? You did not have an answer. The pilot is still a pilot.
Why this one is different
The industry response has been to keep a human in the loop on everything, which means you have not built an agent — you have built an expensive suggestion box, and you are paying for both the model and the reviewer.
What Hetu does about it
Every conclusion carries a label: confirmed, high, medium — two candidates, or hypothesis — requires your input. The label is computed from explained variance and interval overlap, not from the model's self-report.
Below the sample-size floor, the conclusion is withheld and the thin metric is named. A regulatory or macro event overlapping the window suppresses attribution entirely rather than assigning it to the nearest plausible cause.
Figures come from a deterministic rule engine or a fitted causal model. The narration call is separate, receives only the verified object, and is rejected if its output contains a number or a causal claim not in the payload.
The autonomy envelope starts closed. Decision types move inside it as measured outcomes demonstrate accuracy, and fall back out when they don't. Irreversible, fat-tailed actions never auto-execute at any confidence level.
Every decision generated, approved, deferred, executed and measured — who, when, what, why, on what evidence, and what happened. Refusals are logged as first-class records, not as gaps.
The causal graph and the WHY traversals are configuration, hand-seeded by the people who already know your failure patterns. Adding one is an insert, not a deployment — and it stays auditable in a table your risk team can read.
Running today
Not a benchmark. A production deployment where the target variable is deliberately the hard one: not a disbursed loan, but a loan still performing at 90 DPD. Anything less lets a system claim success for originating bad credit quickly.
| Trigger | Outcome | Label |
|---|---|---|
| Performing rate down 14%, one zone | Sourcing quality at two field officers, 60 days prior | Confirmed |
| Throughput down while applications rise | Early delinquency dominant · 44%, CI 37–51 | High |
| Back-book DPD rising, policy change same window | Attribution suppressed. Two hypotheses, each with a falsification test. | Refused |
The third row is the one to take to your committee. A system that will say "I cannot attribute this, and here is what would settle it" is the only kind that can be trusted to act on the rows where it does not say that.
Deployment
The narration model is swappable and the reasoning does not depend on it, because it never did any reasoning. A model upgrade improves the sentences and changes none of the maths.
FAQ
Every decision carries a confidence label, a causal or deterministic rationale, and an immutable audit record — the artifacts a model risk committee asks for before an agent touches a real credit action.
Yes. Deployments support in-VPC and on-prem patterns. Your data is not used to train foundation models, and the reasoning layer does not depend on the narration model.
A pre-seeded causal graph and gate set for one regulated decision type — for example origination or collections — so you start from domain failure patterns instead of an empty prompt.
Most teams begin with one currently-manual decision via Consulting, then convert to the platform once the verification layer earns trust in shadow.
Explore
Design partners
We start with one: a decision your team already makes manually, where you can articulate the failure patterns and where being wrong is expensive. We encode the constraint framework, seed the graph with your experts, and show you the first refusal — which is, reliably, the moment the room understands what this is.