Hetu.

Platform / Verification guard

Nothing reaches the agent until it passes every check.

One shared service, called by all three tiers, so verification stays auditable in one place. Failing a single check refuses the conclusion rather than downgrading it.

The checks

What each one asserts, and what happens on failure.

Temporal ordering

The proposed cause precedes the observed effect by a plausible lag for this decision type.

Refuse. Reverse-causality candidates are logged.

Magnitude consistency

The size of the cause can account for the size of the effect within the model's stated bounds.

Refuse. Residual is reported.

Hard gates

No policy-forbidden conclusion, action or entity appears in the result.

Refuse. Gate id recorded.

Sample floor

The cohort behind the conclusion is above the per-decision-type minimum n.

Refuse and name the thin metric.

Confounder overlap

No declared policy change or macro event spans the anomaly window.

Suppress attribution, return hypotheses.

Label binding

The narration's strength of claim matches the computed confidence label.

Reject the narration and regenerate at the correct label.

Provenance

Every figure in the output traces to a named source in the verified payload.

Refuse. The unsourced figure is quoted in the log.

Refusal logging

A refusal is a record, not a gap.

Every refusal is written with the check that failed, the values that failed it, the tier that produced the candidate conclusion, and the falsification test that would settle the question. Refusals are queryable and countable per decision type.

That count is the most useful diagnostic in the system. A decision type refusing 40% of the time is telling you exactly where the graph is underspecified, and it is telling you before an agent acts on the gap rather than after.

One refusal, end to end

A back-book indicator deteriorating, with a policy change in the window.

The trigger arrives: days-past-due rising across the whole book over a 45-day window, with an intended action of tightening the sourcing filter for two channels.

Tier 1 reaches its second branch and the metric returns null — the feed is incomplete for this window. Logged as a data gap, escalated.

Tier 2 fits the graph, then the confounder check finds a declared underwriting policy change spanning the entire window. Attribution is suppressed rather than assigned to the nearest plausible cause. Escalated.

Tier 3 returns two ranked hypotheses, neither stated as a cause, each with a test: whether unaffected portfolios in the same segment moved too, and whether affected cases concentrate on specific reviewers.

The response is REFUSED. Auto-execution is blocked, the intended action is held, and the record names the failed check, the confounder, the two hypotheses and their tests. The agent tells its user what it does not know.

Next

Then decide what the agent is allowed to do unsupervised.

A verified conclusion and an authorised action are two different questions.