Practice management13 min read · 05 August 2026

SUDS Scale: How to Measure Anxiety During Exposure

By Equipo VRET

LinkedIn X / Twitter
TL;DR

The SUDS scale is the field instrument a clinician uses inside the session to record, in subjective units of distress from 0 to 100, the discomfort the patient reports. Its value lies not in the number but in the anchoring done beforehand and in the consistency of the record. This article sets out what the SUDS scale is good for as a process indicator, which other data turn that series into a decision criterion, and where a beginner's record falls apart.

Squared recording sheet on a consulting-room desk with a column of hand-written ratings and a line joining the values, beside a pen, in natural light

How is anxiety measured during exposure?

It is measured by asking, and that answer tends to disappoint anyone expecting a device. In exposure work the figure that governs the session does not come from a sensor: it comes from a short question the clinician repeats at regular intervals while the patient stays in contact with the stimulus. “Where are you right now, from 0 to 100?” The patient gives a number, and that number is written down with the time and with the task that was running.

That number is expressed in subjective units of distress, and the SUDS scale is the format in which they are collected. What the SUDS scale does is turn a private experience into a series of values that can be compared against itself over time. It does not measure physiological arousal, it does not measure the severity of the disorder and it does not stand in for any standardised instrument. It records what the patient reports feeling at that instant, which is precisely the datum an exposure task needs in order to govern itself.

The practical advantage is latency. A self-report questionnaire arrives too late: it is completed before or after the task, never during it. The SUDS scale is asked for during, and that is why it supports a decision on the spot about whether the task continues, steps up a level or is held a while longer. In the height exposure scenario, for instance, that reading arrives every minute or two without interrupting the scene or taking the headset off.

What follows sets out how the scale is anchored, what gets recorded from each session and, above all, what can and cannot be concluded from a series that falls. That last part is what turns the record into a clinical criterion rather than decoration in the notes.

Macro detail of a printed 0-to-100 numerical scale on a clinical recording sheet, with a pen mark indicating an intermediate value and very shallow depth of field

What is the SUDS scale?

The SUDS scale — subjective units of distress scale — is a 0-to-100 numerical scale on which the patient places the distress felt at that moment. Zero corresponds to complete calm and 100 to the worst distress that particular patient is able to imagine. The format descends from the behavioural tradition Wolpe formalised in 1958 with systematic desensitisation, where the anxiety hierarchy already needed a common unit in order to put its steps in order.

Three properties define how it is used, and all three have consequences in the consulting room. It is an ordinal scale: the distance between 30 and 40 is not equivalent to the distance between 80 and 90. It is self-referential: one patient's 70 is not another patient's 70, nor does it claim to be. And it is a repeated measure: the information lives in the whole series, not in the isolated value written down at a quarter past eleven.

It is worth being plain about what the SUDS scale is not. It is not validated in the psychometric sense and does not aspire to be: it carries no norms, no cut-off points and no published reliability, because what it captures is private by definition. Presenting it as an objective test is a framing error. Presenting it as a useful clinical datum, held up by a stable anchor and a constant cadence, is the use that fits it.

It should also be kept apart from records of peripheral arousal. Where a practice runs biofeedback recording during exposure, the clinician handles two series that do not always move together: a patient may report an 80 with flat skin conductance, or the reverse. That discordance between what is said and what is measured is clinical information of the first order, not a fault in the instrument.

Anchoring the scale: the step that decides whether the number is usable

Anchoring the scale is the procedure by which patient and clinician agree what the endpoints and one midpoint mean, using examples drawn from the patient and phrased in the patient's own words. Without that prior agreement the later series cannot be interpreted, because every session would be measuring with a different ruler and nobody would know which of the two had moved.

Anchoring is done in the consulting room, before the first exposure, and it is written into the notes. A usable anchor takes this form:

  • 0 — the specific situation in which the patient describes himself or herself as calm, of the order of “reading on the sofa on a Sunday afternoon”.
  • 50 — distress the patient recognises, tolerates, and that does not stop him or her carrying on with whatever was under way.
  • 100 — the worst episode the patient genuinely remembers, named by the patient and not proposed by the clinician.

Two precautions hold up everything else. The first: the 100 is anchored in a real memory and not in a catastrophic hypothesis, because if the ceiling is imaginary the scale compresses and every reading stays below 40. The second: the anchor is not retouched halfway through treatment. Where there are grounds to correct it, the date of the change is recorded and the series is read as two separate stretches.

The natural moment for this work is the framing session, alongside the explanation of the procedure and the check on headset tolerance. The first-session acclimatisation protocol places the anchoring of the SUDS scale before any phobic stimulus, which is exactly where it belongs.

The session curve: what gets recorded, and at what cadence

Within-session recording is not a run of loose numbers. It is a minimal five-column layout: minute, task under way, rating reported, behaviour observed, and clinician manoeuvre. With those five columns the session can be reconstructed six months later in front of a referrer or a case review. Without them what survives is an impression, and impressions cannot be set against one another.

A reasonable cadence is one reading of the SUDS scale every one or two minutes while the task is under way, plus one at the start and one at closure. Asking every fifteen seconds turns the question itself into the main stimulus and pushes the patient into monitoring his or her insides instead of attending to the scene. Asking twice a session leaves blind stretches exactly where the interesting things tend to happen.

Alongside the session curve, two items that are not the rating deserve a line each. Before the task: what the patient predicts will happen, and with what estimated probability he or she predicts it. After the task: what actually happened. That pairing of expectancy and outcome is what gives the series its clinical sense; the mechanism by which it does so belongs to another discussion and is not developed here.

The shape of the curve orients the next manoeuvre. A fast rise with a high plateau and no descent indicates that the task exceeds the challenge calibration the case can take today. A flat series at low values points either to a task that activates nothing or to safety signals the patient is still holding on to without having declared them. And a clean within-session descent is pleasant to look at but, on its own, says less than it appears to.

Settling the format in advance saves this discussion in every new case. The checklist for a virtual reality practice collects the minimum fields of the within-session record and the order in which they are filled in while the session runs.

Downloadable resource

VR exposure protocol for dog phobia

A step-by-step clinical protocol with the full hierarchy, the within-session recording sheet, and the closure criteria for each level.

Download the protocol
Consulting-room table with two different session recording sheets laid side by side for comparison, with a clip and a pencil, under soft side light

Process indicator versus outcome indicator

This is where the real usefulness of the instrument is decided. A process indicator describes what happens inside treatment; an outcome indicator describes whether the patient's problem has changed. The SUDS scale is the first and is not the second, and confusing the two planes is the costliest error in the whole record.

The update Craske and colleagues published in 2022, with the inhibitory learning and retrieval framework and the OptEx Nexus, is explicit on this point: the aim of exposure is not to maximise the drop in arousal inside the session but to make it likelier that the inhibitory association is retrieved when the patient needs it outside the consulting room. Hence a within-session fall in the SUDS scale is a poor predictor of medium-term outcome. Why that should be so belongs to the discussion of mechanism; what matters here is the operational consequence.

And the operational consequence is twofold. First: the task is not closed because the number falls, but because what was agreed has been completed and the patient's prediction has been put to the test. Second: the evaluation of treatment does not rest on the session curve, but on outcome indicators gathered between sessions and away from the headset.

Those indicators are of a different order: effective behavioural approach, avoidance recorded in the interval between sessions, reported functional interference, and whatever standardised measure of the disorder the clinician has chosen. The catalogue of outcome assessment instruments for the practice is treated separately; here it is enough to retain that none of them is the SUDS scale and none of them lets itself be swapped for it.

It is worth noting how the literature that underpins the field went about this. Powers and Emmelkamp, in their meta-analysis of virtual reality exposure, rest their conclusions on outcome measures and not on within-session series. Carl and colleagues arrive at equivalence with in vivo exposure by the same route. Neither of those reviews was built on what distress was doing inside the headset.

How do you know whether exposure therapy is working?

The short answer is that it is known by convergence across sessions, and never by what happens inside a single one. The clinician looks for three series moving in the same direction over four to six sessions, and only then speaks of progress with any foundation.

  • Starting rating. The value the patient enters the same task with comes down session after session. This is the useful reading of the SUDS scale: between sessions, not within them.
  • Level reached. The patient completes tasks previously left unfinished, at the same reported rating or even a higher one. Tolerating more at equal distress is progress, even when the curve does not fall.
  • Estimated probability. The probability the patient assigns to the catastrophic belief comes down, and comes down as well when the question is put cold, days after the session.

To those three series is added the criterion that really decides discharge: clinically significant change in the patient's life. Using the lift, the car or the team meeting again without mapping an escape route first. No curve stands in for that item and no repeated measure of distress anticipates it with enough reliability.

When the three series fail to move over six or eight sessions, the problem is almost never the scale. It usually sits in the calibration of the challenge, in safety signals that are still operating, or in an objective that was badly formulated from the outset. There, a case review — and supervision, where the clinician has access to it — pays off considerably better than adding sessions to the same plan.

Common recording failures, and the operational close

Five failures account for almost every recording problem that surfaces when other people's cases are reviewed:

  • An unanchored scale. Numbers start being asked for before the endpoints have been agreed with the patient, and the series is illegible from the first day.
  • Closing on the number. Ending the task when the rating comes down, rather than when what was agreed has been completed and the prediction has been tested.
  • Oversampling. Asking so often that the question itself becomes the dominant stimulus of the session.
  • A moving anchor. Redefining the 100 halfway through treatment and then comparing stretches that are no longer comparable with each other.
  • The curve as outcome. Presenting the within-session fall in the SUDS scale to the patient, or to the referrer, as evidence of clinical improvement.

There is a sixth failure, quieter than the rest: writing it down by hand and never looking at it again. The within-session record only pays off if it is reread before the following session and if the whole series can be taken in at a glance. Where the rating is captured in the system itself while the scene is running, that review costs a minute instead of an afternoon of transcription.

In VRET the SUDS scale reading is entered from the clinician panel during the exposure, together with markers for the behaviour observed and the manoeuvre applied. Book a demonstration to see how a session curve is collected and how it is set against the previous one without leaving the clinical record. Current plans are Starter at $119 a month, Clinic at $289 and Enterprise at $1,499, with a 30-day refund window.

This article is for informational purposes for psychology professionals. It is not clinical advice for any individual case and does not replace the judgment of the licensed psychologist in charge. VRET is professional clinical-support software, not a CE-marked medical device.

Frequently asked questions

How often is the rating asked for during an exposure session?

One reading every one or two minutes while the task is under way, plus one at the start and one at closure, covers most exposure protocols. The exact figure matters less than the consistency: the same cadence across every session with the same patient, because a series is only comparable when it is sampled the same way. Two deviations are worth avoiding. Asking every few seconds turns the question into the main stimulus and pushes the patient into monitoring his or her insides instead of attending to the scene in front of them. Spacing the readings too far apart leaves blind stretches precisely where the clinician would need to know what happened and in what order. Where the task is short, an entry reading, one intermediate reading and an exit reading are enough, always written down with the minute and the task they belong to.

What happens when the patient cannot produce a number?

It happens often in the first sessions and it is almost never resistance: usually the scale has been anchored badly, or the cognitive demand is excessive in the middle of high arousal. The usual way out is to lower the resolution. The 0-to-100 scale gives way to four bands, of the order of none, some, quite a lot and a great deal, which are then translated back into the reference values agreed at anchoring. Offering an interval also works, along the lines of asking whether it is closer to 40 or to 70 rather than requesting an exact value. With patients who have fewer verbal resources, or with minors, a visual scale of faces or a colour bar performs the same function without losing comparability. What is not acceptable is for the clinician to fill the gap with a personal estimate and record it as though it came from the patient.

Does one anchoring hold for the whole of treatment?

Yes, and much of the value of the instrument rests there. The anchor is the unit of measurement for the case: if it shifts, the series stops being a repeated measure and becomes a collection of readings with no relation to one another. There is one reasonable exception. When treatment advances a long way, the patient reinterprets his or her own ceiling and the original 100 comes to look exaggerated. In that case re-anchoring is defensible, but the date is recorded, the previous anchor is kept in the notes, and the course is read as two separate stretches, never as one continuous series. What is not defensible is re-anchoring tacitly, session by session, because then any apparent fall may be an artefact of the ruler rather than a change in the patient.

Are ratings from two different patients comparable?

They are not. The SUDS scale is self-referential and its anchoring is idiosyncratic, so a 70 may describe very different experiences in two people. This has three practical consequences. In the clinical notes, comparisons are always made against the patient's own history and against the same task, never against another case. In reports to the referrer, the number travels alongside the behaviour observed, which is comparable: whether the task was completed, how long the patient stayed with it, and what he or she ended up avoiding. And in any attempt to aggregate data from one's own caseload, ratings do not admit averaging across patients; what does admit averaging is the change within each case, expressed as a difference from that patient's own starting point.

What does recording the scale in the system add over a paper sheet?

Three concrete things. Traceability: the reading is stored with its time, task and manoeuvre attached, without depending on somebody transcribing it at the end of the day. Legibility: the series can be taken in at a glance, and comparing it with the previous session does not mean digging through folders. And continuity: the information survives sick leave, cover arrangements and changes of professional inside the practice. Paper works perfectly well in a single-clinician practice with few active cases, and stops working when several professionals share a patient or when forty case histories have to be reviewed before the quarter closes. Anyone who wants to see what that record looks like inside a virtual reality exposure session can ask for a demonstration of the clinician panel and set it against their own sheet format.

VRET is professional clinical-support software, not a CE-marked medical device. Clinical supervision remains with the licensed psychologist in charge.