Reliable change, in one step
Is a change between two questionnaire scores real improvement, or within the measure's normal wobble? The NHS Talking Therapies thresholds for PHQ-9 and GAD-7, and the Jacobson–Truax reliable change index for any measure — with nothing about the client entered.
From the measure's manual or a published norm study. Use figures from a population like your client's.
Clinical significance (optional)
Jacobson–Truax criterion c: the point between the clinical and non-clinical population means, weighted by their spread.
—
No client information, by design. This page takes two numbers and nothing else — no names, dates or identifiers — and sends nothing anywhere. It supports clinical judgement; it doesn't replace it, and one score pair is never the whole picture.
The methods, and where the thresholds come from
PHQ-9 and GAD-7. NHS Talking Therapies (formerly IAPT) treats a change of 6 or more points on the PHQ-9, and 4 or more on the GAD-7, as reliable, and uses 10 (PHQ-9) and 8 (GAD-7) as the caseness thresholds. Recovery means moving from at or above the threshold to below it; reliable recovery means that and a reliable improvement together. The programme reports these on paired depression and anxiety measures together; this page looks at one measure at a time.
Any other measure. Jacobson & Truax (1991): the standard error of measurement is SD × √(1 − reliability); the standard error of the difference is √2 times that; the reliable change index is the change divided by it. An RCI beyond ±1.96 is unlikely (p < .05) to be measurement error alone. The smallest reliable change, in the measure's own points, is 1.96 × Sdiff.
Results depend heavily on the SD and reliability you use. Check that the direction matches your measure — on some, higher is better.
Why "reliable" matters
Every questionnaire has measurement error. A client can score 14 one week and 11 the next with nothing changing at all. Reliable change asks whether a difference is bigger than that noise, so that "she's improving" rests on more than a few points of drift — and so a real deterioration isn't missed.
It's also the backbone of measurement-based care: routine scores reviewed with the client, used to notice when therapy isn't helping early enough to change course.
Using it with clients
Clients can track their own PHQ-9 and GAD-7 in the private check-ins, which apply the same thresholds and keep everything on their own device. The guide to measuring progress in therapy explains reliable change and recovery in terms you can share.
Related
For the practice side, the private practice planner works out caseload, income and the real cost of no-shows.
Questions people actually ask
What counts as reliable change on the PHQ-9?
NHS Talking Therapies uses a change of 6 points or more. On the GAD-7 it uses 4 points or more.
What's the difference between reliable change and clinically significant change?
Reliable change says a difference is bigger than measurement error. Clinically significant change adds that the person has moved into the range of people without the problem — past a threshold or criterion c. Both together is the strongest result.
Does this page store anything?
No. It holds two numbers while the page is open and sends nothing anywhere. There is nowhere to type a name.