Pre-registered · Version 1.0 · 8 September 2026

How the six risk states can be shown to be wrong

This protocol was written and published before any results were examined, so that what counts as satisfactory evidence was not decided afterwards by looking at what the data happened to support.

Why this document exists

The six-state model makes two claims that are easy to blur and must not be. The first is structural: the 2×3 grid is built from two independent dimensions with defined cut-offs, so every goal falls in exactly one state at any moment. That follows from the construction, and it is established.

The second is empirical: that the states meaningfully discriminate between later outcomes. That has not been demonstrated, and it cannot be settled by argument. It needs data that did not exist, because risk state was computed on read and never stored.

The distinction was raised in an external review of the methodology in June 2026 and remained the largest open point in September 2026. Rather than argue the question, we built the instrument that can answer it — and published the conditions first.

What is being claimed

De-Risk Matrix positions the six states as a decision-framing device with an empirically testable structure, not as a predictive classifier. A model can be predictively modest and still decision-relevant. These are three different propositions requiring three different kinds of evidence, and we report them separately — neither is used to argue the other.

A

Structural validity

Are the six states logically coherent, mutually exclusive, and capable of classifying the situations the methodology intends to classify?

Follows from the construction. Established.

B

Empirical validity

Do the states meaningfully discriminate between subsequent outcomes or transitions?

Requires outcome data over time. Not yet demonstrated.

C

Decision usefulness

Does distinguishing Optimistic from Dire lead decision-makers to materially better decisions than a simpler classification would?

Requires decision and intervention data. Not yet demonstrated.

The data

Every observation records the state and the inputs it was derived from: the forecast and its band, the target and threshold in force at the time, the evidence strength and how it resolved, the data point count, the assumption counts, priority and period end.

Storing the inputs rather than only the conclusion is what makes test 5 possible at all. A stored label can never be compared against a simpler model built on the same variables. Observations are written by every action that can change a state, and by a nightly job — history cannot depend on someone happening to log a value.

The five tests

A protocol that cannot fail is not a protocol. Each test states its own falsification condition.

1

Calibration and outcome discrimination

Of goals classified Dire at time T, what proportion ended below threshold at the end of their period? Of goals classified Optimistic?

For each goal-period with at least one observation, take the state at a fixed fraction of the period elapsed and the realised outcome at period end. Report the outcome distribution per state with confidence intervals.

If the states do not separate outcomes, the model is wrong on proposition B. That result gets published.

2

Transition structure

Which states flow into which, and at what rates?

Build the transition matrix from recorded history, excluding transitions caused by definition changes — moving a target moves the state without anything happening in the world.

If two states have statistically indistinguishable transition profiles and indistinguishable outcome distributions, they are not two states.

3

Classification reliability

How consistently do different people classify evidence?

Evidence is the one genuinely subjective axis — it can be set manually, derived from a 14-factor checklist, or left to the automatic rule. Compare manual overrides against what the checklist and the automatic rule would have produced on the same data.

Separates variance in the model from variance in the observer. A team that consistently overrides upward is telling you about its own optimism, not its data.

4

Decision impact

Did the classification actually change a decision, escalation, resource allocation or intervention — and compared with what would otherwise have happened?

Link state changes to decisions and actions created within a defined window afterwards. Report the rate of response per state, and whether responses match the state’s prescribed action.

Observational. It can show association between classification and response; it cannot by itself establish causation, nor that the response improved the outcome.

5

Incremental explanatory value

Does the six-state structure carry information beyond the variables it was derived from — and does it beat a simpler three- or two-state model built on those same variables?

Compare predictive performance of the six-state label, a three-state position-only model, a two-state model, and the raw underlying variables. Report the increment, not just the absolute performance of the six-state model.

Complexity should earn its place. If two states predict as well as six, the extra four must justify themselves on decision usefulness alone — and be described that way.

Horizon, sample size, and pooling

Horizon

No single evaluation period is imposed. Three months may be meaningful for an operational objective and close to meaningless for a three-year strategic one. The horizon is defined relative to each objective's own cycle, and results are reported per frequency band before any pooling.

Sample size

Deliberately not fixed in advance as a single number. The required n depends on observation frequency, the distribution across states, the number of organisations, the independence of observations, and which outcome is being tested — none of which are known yet. The stopping rule is structural instead:

  • Each test reports its own achieved precision — confidence intervals, not point estimates.
  • No test is reported as supporting or refuting the model while any state cell holds fewer than 30 goal-periods. Below that, the cell reads insufficient data.
  • Analysis runs at pre-declared intervals, not whenever results look interesting.

Organisation-level versus pooled

Kept separate and reported separately. A model that works well for one type of objective or organisation could otherwise conceal weak performance elsewhere behind a healthy pooled average. Pooled analysis also requires a lawful basis under GDPR, disclosure in the privacy policy, and coverage in the DPA including sub-processors — it is not run before that is in place.

What this protocol does not test

Stated explicitly, because the omission matters more than the inclusions.

The architecture records assumptions, dependencies, scenarios and reconsideration triggers — that is, it represents what management knows it does not know. There will always be uncertainty that was never represented as an assumption, scenario or dependency in the first place.

No result here can support the claim:

If all assumptions remain valid, the forecast is reliable.

The most any of it can support is:

None of the assumptions we explicitly identified has yet been invalidated.

This is why the fragility indicator in the product has a floor above zero and reports no identified fragility rather than no fragility. It is a constraint on the architecture, not a caveat on the wording.

Provenance

Claims in the methodology are labelled as one of three kinds, and this protocol tests only the third: alignment with ISO 31000; external research by others, which supports underlying concepts but does not validate our application of them; and De-Risk's own proposed practice — the six states, the state-to-leadership-behaviour mapping, the assumption impact scale and the fragility indicator, all of which are validation targets rather than established findings.

See the full methodology →

Changelog

Any change to this protocol after data collection began is recorded here, with a date and a reason.

DateChangeReason
2026-09-08Protocol v1.0 pre-registeredBefore any state-history data was examined