Skip to content

Equilens note

Five things to fix before you test a credit workflow for fair outcomes

Before choosing a metric: five preconditions for testing one credit workflow without confusing activity with evidence.

Choosing a fairness metric is not the first step.

Before a credit-risk, model-risk, Consumer Duty or AI-governance team starts an outcomes test, it needs to know whether the workflow can produce a result that people can understand, challenge and act on.

That is a practical problem, not a new regulatory framework. The FCA says its existing frameworks continue to apply to AI, including expectations around governance and controls. It is asking firms how they govern AI, test models, monitor outcomes, ensure fair treatment and explain AI-driven decisions. Its recent Consumer Duty work also warns against presenting large amounts of data without showing what they say about customer outcomes or what action follows.

The independent Financial Services AI Adoption Plan published by HM Treasury reports the same implementation gap: firms want clearer, practical application of existing expectations around Consumer Duty, model risk, explainability and accountability.

A useful readiness check asks whether five preconditions are in place.

1. One decision

“Our credit model” is too broad.

Choose one decision or treatment: approval or decline, affordability, pricing, a credit-limit change, a manual override, collections or forbearance treatment, or another clearly bounded workflow.

Record the population, time period, relevant policy or model version, and any exclusions. If the decision boundary is vague, the result will be vague too.

2. One evidence question

State the question before choosing the metric. For example:

  • Are relevant customer groups experiencing materially different outcomes?
  • Did a policy or model change alter those outcomes?
  • Are manual overrides concentrated in a way that needs investigation?
  • Are the available inputs good enough to support a defensible comparison?

Metrics should be selected because they help answer the question. Their uncertainty and limits should be stated. No single ratio or threshold creates a compliance conclusion by itself.

3. A safe data boundary

Credit evidence can involve borrower data, model information and, where lawful and appropriate, protected-characteristic data.

Decide in advance:

  • where the data and model information will remain;
  • who can access them;
  • what, if anything, may leave the controlled environment;
  • what will be retained and for how long; and
  • whether an external provider needs the raw data at all.

An evidence test should not create a new privacy, security or procurement problem while examining an existing one.

4. Named owners for each judgement

Software can calculate and preserve evidence. It should not silently become the decision-maker.

Name who will:

  • define the customer outcome being tested;
  • confirm the population and the meaning of the data;
  • review statistical and operational limitations;
  • decide whether a finding needs investigation; and
  • own any change to policy, model or customer treatment.

Without those owners, even a technically correct report can go nowhere.

5. A pre-agreed action

Agree the possible outcomes before seeing the analysis:

  • stop because the data or workflow is not ready;
  • fix a named gap and rerun;
  • investigate a defined finding;
  • proceed to a bounded test or pilot; or
  • close the question because no further work is justified.

Record the action criteria before the test: the minimum data quality needed to proceed, what magnitude and uncertainty would trigger investigation, and who can approve a bounded next step. These are internal decision thresholds, not regulatory safe harbours or automatic compliance verdicts.

A stop decision can be useful. It prevents weak evidence from being presented as certainty.

The output: a one-page test brief

The readiness check should end with a short brief recording:

  • the decision and evidence question;
  • the population, period and inputs;
  • the data boundary;
  • the named owners;
  • the action criteria; and
  • the known limitations.

Its conclusion should be simple: ready, ready after named gaps are fixed, or not ready.

This is not model validation, legal advice or compliance certification. It is a way to find out whether one fair-outcomes test can be run safely and produce evidence that leads to a real decision.

Equilens builds FL-BSA, a customer-hosted, simulation-only evidence appliance for regulated credit. FL-BSA does not make or override live lending decisions, provide legal advice or certify regulatory compliance.

Sources