Efficacy of Exposure Therapy: What the Evidence Shows
By Equipo VRET
The efficacy of exposure therapy is among the best documented in all of clinical psychology: meta-analyses report large effect sizes against waiting-list controls, somewhat smaller ones against psychological placebo, and maintenance at twelve-month follow-up. The limits are just as clear — heterogeneous protocols, probable publication bias, and dropout reported inconsistently. And the drop in distress recorded inside a session is not the marker that predicts where the patient will be a year later.

Does exposure therapy work?
The short answer is yes, and it deserves to be given plainly: exposure sits among the psychological interventions with the deepest accumulated empirical support. The efficacy of exposure therapy has been under review since Wolpe (1958) described systematic desensitization, and the body of data has not stopped growing. Along the way it has held up against control groups, against other forms of psychotherapy and against the passage of months. The longer answer, which is the one that matters in the consulting room, requires three things to be pinned down: what is being compared with what, how much the improvement obtained is worth, and how long it lasts.
When someone asks whether exposure therapy works, they are usually asking two things at once. One is statistical: does the treated group improve more than the group that receives nothing? The other is practical: is that improvement visible in the person's life? The first question is settled comfortably. The second turns on clinical significance, a criterion that a good share of studies reports only partially. Keeping the two planes apart guards against easy enthusiasm and against lazy skepticism alike.
Support is also not uniform across presentations. It is firmest in specific phobia, where the feared stimulus is well bounded and improvement can be measured with a behavioral approach test. Height fear is the classic case: graded exposure to heights, now also available in virtual environments such as the glass-elevator height exposure scenario, produces measurable change within a handful of sessions. In more diffuse presentations — complex post-traumatic stress, heavily generalized social anxiety — the evidence on the efficacy of exposure therapy exists, but it is thinner and more scattered.

Is exposure therapy effective outside the trial setting?
English keeps two words that the research literature is careful to separate. Efficacy describes what happens under ideal conditions: selected patients, trained therapists, a closed protocol and a control group. Effectiveness describes what happens in a real caseload, with comorbidity, irregular attendance and people who arrive referred for something else. The efficacy of exposure therapy is well established. Its effectiveness is studied far less often, and that is precisely where the licensed clinician has to apply independent judgment.
Three factors account for almost the whole distance between the two planes. The first is comorbidity: trials routinely exclude active substance use, suicide risk or psychotic features, and those profiles do turn up in private practice. The second is adherence, because an eight-session protocol is poorly delivered when the patient misses week three. The third is the clinician's own fidelity, which tends to soften the challenge as soon as the patient protests. None of the three invalidates the pooled data; all three require it to be read with caution.
Hence the protocol calls for systematic measurement case by case. Aggregate evidence says what to expect on average; the clinician's own record says what is happening with this patient. That second task is a separate conversation with its own material: the review of outcome instruments for the consulting room sets out which scales to administer and at what point. What concerns us here is the aggregate.
One framing point matters before going further. NICE CG113 does not treat exposure as a standalone product but as a component of a broader cognitive behavioral treatment. The same holds for its variants: exposure with response prevention in obsessive-compulsive disorder, reviewed by Abramowitz (1996), shares the mechanism but adds an ingredient of its own. When meta-analyses estimate the efficacy of exposure therapy, they are pooling a family of procedures rather than a single technique.
How is the efficacy of exposure therapy measured?
Every claim about the efficacy of exposure therapy rests on one unit of account: the effect size. The most common here is Hedges's g, a difference between means divided by their pooled standard deviation and corrected for small samples. It is read against a simple convention: around 0.20 the effect is small, near 0.50 moderate, and from 0.80 upward large. The advantage is that studies using different scales can be added together. The risk is mistaking the number for an individual promise, when what it describes is the average displacement of a group.
The same treatment yields very different figures depending on the comparator. This is the ladder worth keeping in mind when reading any meta-analysis:
- Waiting list: the most forgiving contrast. It captures the treatment effect plus the effect of time simply passing for someone who is expecting something.
- Psychological placebo: therapeutic attention without the active ingredient. It asks more of the treatment and brings the figure down.
- Active treatment: another intervention with support behind it. Here the differences narrow and frequently stop reaching significance.
- Non-inferiority: the aim is not to win but to rule out that the newer modality falls below a margin fixed in advance.
That last point is misread often. A non-significant result does not demonstrate equivalence: it may reflect a lack of statistical power. Non-inferiority designs exist for exactly that reason, because they fix the margin before data collection and force an adequate sample size. When a paper concludes that two modalities are comparable, it is worth checking whether the design was one of non-inferiority or whether no difference was simply found.
Then comes the least cited and most decisive question: what counts as the outcome. For decades the criterion was within-session habituation, in the emotional-processing tradition of Foa and Kozak (1986) — activate the fear structure and wait for the activation to come down. Craske and colleagues (2014, with the 2022 update) moved that criterion. Extinction does not erase the original fear association; it builds a competing association that has to be retrievable afterwards, outside the office. What predicts medium-term outcome is expectancy violation, the discrepancy between what the patient predicted and what actually happened. This is why the efficacy of exposure therapy is now judged at follow-up rather than from the slope of a single session.
What effect sizes support the efficacy of exposure therapy?
With those cautions on the table, the numbers. The first published controlled trial of exposure in virtual reality was Rothbaum and colleagues (1995), with patients presenting height fear; it opened the line of work that now carries much of the available arithmetic. Thirteen years later, the meta-analysis by Parsons and Rizzo (2008) pooled 21 studies and placed the effect size at around g = 0.95 in specific phobia — in the large band of the convention.
Powers and Emmelkamp (2008) grouped 13 controlled trials and found a large effect against the control conditions, with no relevant difference from exposure to real stimuli. Opriş and colleagues (2012) widened the pool of studies, confirmed the clear advantage over waiting list, and added the finding that weighs most in practice: the effect was still present at the 12-month follow-ups of the trials that got as far as measuring it.
The broadest count available is Carl and colleagues (2019): 30 randomized controlled trials and close to 1,057 participants, with specific phobia, social anxiety, panic, agoraphobia and post-traumatic stress in the same bag. Against waiting list they obtained g ≈ 0.90; against psychological placebo, g ≈ 0.78; against exposure to real stimuli, g ≈ −0.07, a difference that did not reach significance. That last number speaks of equivalence rather than advantage, and that is how it should be cited. The technical comparison between the two modalities has its own treatment in the review of the meta-analyses against in vivo exposure.
One finding tends to surprise: Öst (1989) showed that in some specific phobias a single long session with therapist modeling is enough to produce clinical change. It is not the rule, and how many sessions to plan is a separate discussion. It serves the point at hand: the efficacy of exposure therapy does not depend on one single way of delivering it. What repeats across all of this work is a large effect against control and a narrowing of the differences whenever the comparator is a good one.
Clinical VR software comparison
A side-by-side sheet on the clinical virtual reality platforms available today: scenarios included, the evidence each vendor states, equipment requirements and contract terms.
Download the comparisonDoes the effect hold at long-term follow-up?
A treatment that improves at discharge and fades within six months is worth little. Follow-up is therefore the hard test. Opriş and colleagues (2012) recorded maintenance at 12 months in the trials that reached that measurement point. Anderson and colleagues (2013) contrasted exposure in virtual reality with cognitive behavioral treatment in social anxiety and described comparable results both at the end of treatment and at the one-year follow-up. The general picture is one of reasonable persistence.
Reasonable does not mean assured. The extinction literature describes three phenomena any clinician recognizes at once: the return of fear with the passage of time, its return in a context different from the treatment setting, and relapse after an unexpected encounter with the stimulus. Craske and colleagues explain it well: if extinction creates a competing association instead of erasing the original one, then maintenance is a problem of retrieval. The new association has to be available at the moment and in the place where it is needed.
Concrete protocol decisions follow from that reading. The clinician varies contexts and stimuli instead of repeating the same scene, removes safety signals before closing, introduces retrieval cues the patient can evoke elsewhere, and schedules spaced booster sessions. In social anxiety, where the real context cannot be reproduced inside an office, that variation leans on graded audiences such as those in the social anxiety exposure scenario. The efficacy of exposure therapy measured at one year depends a good deal on these decisions.

Where is the available evidence weakest?
The efficacy of exposure therapy is always read through meta-analyses, and no meta-analysis improves on the quality of the trials it pools. The first problem is heterogeneity: the number and length of sessions, the progression criteria, the therapist's role and the quality of the stimulus vary widely between studies. When all of that is grouped under one label, the result describes a family of interventions rather than a concrete protocol that can be copied as it stands.
The second is blinding. In psychotherapy the patient cannot be kept unaware of which treatment is being received, nor the therapist of what is being delivered. Randomization controls selection, not expectation. Published effect sizes therefore carry the weight of expectancy and alliance, which are part of the real treatment but make it impossible to attribute all of the change to the procedure.
The third is publication bias. Null studies are published less often and later, which inflates the average of any young literature. Asymmetry analyses applied to this field point to a moderate bias, and not a negligible one [CITATION TO VERIFY]. The fourth is dropout, which is reported irregularly and, where the analysis is not intention-to-treat, favors the treatment. The discussion of tolerability and dropout in exposure develops that point in more detail.
Then there is representativeness. A trial participant agrees to be randomized, has schedule availability and a relatively clean diagnosis. The patient walking into a private practice in Manchester or in Denver rarely meets all three conditions. None of this brings down the central conclusion: the efficacy of exposure therapy remains among the best documented in the psychotherapy catalog. What these limits trim is the precision with which a pooled figure transfers to a single case.
What can the clinician state to a referrer?
The efficacy of exposure therapy can be summarized without inflating it and without playing it down. The technique has broad and sustained support in specific phobia, social anxiety and panic disorder; the effects are large against not treating and they hold at the follow-ups available; the average improvement is substantial, and there are patients who do not respond. That last clause does not weaken the message — it is what makes it credible to a general practitioner or a psychiatrist who has read the literature.
It is worth stating just as plainly what does not hold. Aggregate evidence does not predict the outcome of a single case, does not fix a timetable, and does not exempt anyone from the usual frame: informed consent, prior assessment, exclusion criteria and documentation in the clinical record. The indication also requires judging whether the feared stimulus can be approached with the means at hand. A large effect in a meta-analysis transfers nothing about this particular patient.
For anyone weighing whether to move part of that progression into a virtual environment, the useful conversation is about fit rather than figures. It is worth booking a demonstration with the clinical team to see how the levels of a scenario are configured, what gets recorded in each session and what information reaches the report. Plans start at $119 a month for the independent practitioner and $289 for a clinic, with a 30-day money-back window; the full breakdown of the plans sets out what each tier includes.
This article is for informational purposes for psychology professionals. It is not clinical advice for any individual case and does not replace the judgment of the licensed psychologist in charge. VRET is professional clinical-support software, not a CE-marked medical device.
Frequently asked questions
What effect size counts as clinically meaningful in exposure research?
The statistical convention places a small effect around g = 0.20, a moderate one near 0.50 and a large one from 0.80 upward. In the exposure literature the values against waiting list sit in the upper band and fall when the comparator is an active treatment. That convention is not a clinical criterion. A large effect can correspond to an improvement the patient does not experience as sufficient, and a moderate one can be enough for someone to step into an elevator again. Reading the efficacy of exposure therapy should therefore be accompanied by indices of clinical significance: the proportion of patients who no longer meet criteria, the minimal important change on the scale used, and behavioral tests passed. A good share of studies reports these incompletely.
Is exposure effective in obsessive-compulsive disorder?
Support in obsessive-compulsive disorder comes from the variant with response prevention, not from exposure on its own. Abramowitz (1996) reviewed the different ways of delivering it and concluded that preventing the compulsive behavior is part of the active ingredient rather than an optional addition. Transferring to this presentation the figures obtained in specific phobia is a frequent reading error: trial design, outcome measures and protocol length all differ. The indication here calls for specific training in response prevention and a plan built on this patient's obsessions rather than on a catalog of generic stimuli.
How much does publication bias weigh in the exposure literature?
It is a factor worth discounting, although it does not invalidate the whole. Null trials are published less often, later and in journals with smaller circulation, so any pooled average tends to sit above the true value. The usual methods for detecting it, funnel plots and asymmetry tests, point to a moderate bias in this field [CITATION TO VERIFY]. The magnitude of the effect against not treating is large enough to survive a correction of that order. Comparisons against active treatments, which already yield small differences, are far more vulnerable. That is where close reading pays off.
Can exposure be considered a first-line treatment?
Guidelines place it there, with a qualification worth keeping. NICE CG113 does not recommend exposure as an isolated intervention but as a component of a structured cognitive behavioral treatment. What the support backs, then, is the therapeutic package within which exposure does the specific work of disconfirming the patient's prediction. Presenting it as a loose technique blurs the frame that holds it up, including the prior assessment and the closure. With that qualification, the efficacy of exposure therapy justifies an early place in the sequence of decisions, always under the indication of a licensed psychologist.
How much should the efficacy of exposure therapy weigh when deciding whether to bring virtual reality into the practice?
The evidence does not make that purchase decision; it bounds it. The available meta-analyses do not describe the virtual modality as a different route but as another way of presenting the same stimulus, with equivalent results in the phobias where it has been studied. What the virtual environment adds is control of the stimulus, reproducibility between sessions and access to situations that do not fit inside an office. The reasonable decision is made on operational criteria: which presentations the practice sees, how many patients would benefit, what record the report requires, and what the equipment costs against the time it saves. A guided demonstration with the clinical team settles more doubts than another list of effect sizes.
Keep reading
Virtual Reality Therapy: The Complete 2026 Clinical Guide
What VRET is, the clinical evidence (Cochrane meta-analyses, UJI research), building a VR exposure hierarchy, contraindications, and ROI for private practice.
Software comparisonsThe Difference Between Exposure Therapy, CBT, ERP and VR
The difference between exposure therapy and systematic desensitisation, response prevention, CBT and virtual reality, sorted level by clinical level.
Practice managementHow Exposure Therapy Works: Inhibitory Learning Model
How exposure therapy works under the inhibitory learning model: what habituation explains, what extinction never erases, and why expectancy runs the show.
VRET is professional clinical-support software, not a CE-marked medical device. Clinical supervision remains with the licensed psychologist in charge.