How Exposure Therapy Works: Inhibitory Learning Model
By Equipo VRET
How exposure therapy works is no longer explained as fear wearing itself out. The current account is inhibitory learning: extinction does not erase the original association, it builds a competing one that the patient has to retrieve at the right moment. That is why the decisive factor is expectancy violation rather than how long the patient endures a trial or how much relief arrives before the session ends.

How does exposure therapy work?
The short answer takes two sentences. Exposure does not wear a learned fear down until there is nothing left of it: it installs a new learning that competes with the old one and that, at best, wins the race when the stimulus turns up again. What counts is not how much distress the patient tolerates, but how much information contrary to their own prediction reaches the alarm system.
The distinction sounds academic and is not. Very concrete protocol decisions hang on it. A clinician who assumes the mechanism is attrition keeps each trial running until distress subsides, and closes the session there. A clinician who assumes the mechanism is a competing learning works to a different criterion: what did the patient expect, what actually happened, and how far apart were the two.
That shift reorders the reading of trials that look like failures. A trial spent at high distress throughout may have produced excellent learning if it disproved a sharply stated prediction. A comfortable trial, one the patient walks out of unruffled, may have contributed nothing if they already took for granted that nothing would happen. In the dog phobia exposure scenario the contrast needs no explaining: the patient who expects the animal to lunge and watches it stay still has learned something; the one who assumed it would stay still has not.
How exposure therapy works is therefore described on two levels today: an observable within-session phenomenon, the fall in distress, and a learning process that runs underneath it and does not track it. The sections that follow take both levels in the order the research established them.

What does habituation mean in exposure therapy?
Habituation is the decline of a response to a stimulus that repeats without aversive consequences. Exposure work describes it on two planes: within-session habituation, the fall in distress while a trial is running, and between-session habituation, the fall in the peak reached from one encounter to the next.
Foa and Kozak gave those two planes the status of indicators when they set out emotional processing theory in 1986. Their proposal was that fear reduction requires two conditions at once: activation of the fear structure, meaning the patient has to be genuinely aroused by the stimulus, and the intake of corrective information incompatible with that structure. Initial activation, the decline inside the session and the decline across sessions were duly enshrined as the three signs that processing was under way.
The model had an enormous run, and for good reasons: it is operational, it condenses into a single number, and it can be taught in ten minutes. From it comes the classic rule still printed in most training manuals [CITATION TO VERIFY]: do not end the exposure until anxiety has come down appreciably, roughly to half the peak reached.
The trouble with the rule is empirical. Craske and colleagues reviewed the relationship between those indicators and long-term outcome in 2014, and the within-session decline came out a poor predictor. The between-session decline behaves somewhat better, but neither carries the weight the model assigned to it. Habituation remains a real, observable event; what it is not is the explanation of therapeutic change, nor a dependable criterion for ending a trial.
The practical consequence is uncomfortable for anyone who has spent years closing sessions on that criterion. One trial can end with distress on the floor and leave nothing usable behind; another can end with the patient still aroused and turn out to be the most productive of the series. A good deal of adherence is decided at this point, so it repays reading alongside the tolerance and dropout mechanisms of exposure work.
What does inhibitory learning actually mean?
Inhibitory learning names what happens during extinction according to the framework Craske and colleagues put forward in 2014. The original association between the conditioned stimulus and the aversive outcome, the CS-US association in the notation of classical conditioning, is not erased. Alongside it a competing association forms, CS-no US, which states that this stimulus no longer announces that outcome.
The two associations coexist and compete every time the stimulus appears. The new one is secondary and more fragile: it depends on the context in which it was acquired and it loses strength over time. The old one stays intact, available, waiting for conditions that favour it.
That it stays intact is not a theoretical conjecture; it follows from three well documented findings in the extinction literature. In spontaneous recovery, fear returns with the mere passage of time. In contextual renewal, it returns when the stimulus is met in a context other than the one where extinction took place. In reinstatement, it returns after an aversive experience unrelated to the treated stimulus. If extinction rewrote the original learning, none of the three could happen.
The update the same group published in 2022 moves the emphasis from learning to retrieval. What decides the outcome is less the forming of the inhibitory association than making sure the patient retrieves it at the moment it is needed, outside the consulting room and months later. Hence the name of the revised framework: inhibitory retrieval.
Hence too the central place of expectancy violation. If what gets learned is that a particular stimulus no longer announces what the patient fears, then the unit of learning is the discrepancy between what they predicted and what in fact occurred. With no discrepancy there is nothing to learn, however long the trial is drawn out.
How is expectancy violation engineered in the session?
Expectancy violation is not improvised, it is prepared. The first move is to obtain an explicit prediction before the trial begins. Anticipating a lot of anxiety is not enough. A usable prediction is specific and falsifiable: what the patient fears will happen, how likely they believe it is, and how they would know it had happened. A vague prediction cannot be disproved, and what cannot be disproved teaches nothing.
The second move is to design the trial so that experience, not argument, does the disproving. The clinician does not debate the belief; the clinician puts it to the test. If the patient holds that they will faint by the second minute, the trial has to reach the third. If they hold that they will be unable to let go of the handrail, letting go belongs in the trial.
The third move is the review afterwards, and it is where most is lost. The pertinent question is not how it went, but what they expected, what happened, and how they account for the gap. The aim is for the discrepancy to settle as the patient's own datum rather than the therapist's conclusion. Verbal persuasion is no substitute for surprise.
The habitual beginner errors are of this order rather than technical: reassuring before the trial, calling the outcome out loud in advance, tolerating a discreet safety signal that lets the good result be credited elsewhere. All three shield the patient from distress and withdraw the learning. A pass through the mistakes that keep surfacing in clinical supervision tidies this ground considerably.
VR exposure protocol for dog phobia
Step-by-step clinical protocol, with the prediction-and-outcome record for every trial and the graded sequence of levels.
Download the protocol
Which optimization strategies follow from the inhibitory framework?
The framework does not stop at explanation: it proposes a set of strategies aimed at maximising learning and, above all, its later retrieval. Craske and colleagues listed them in 2014 and keep them, with qualifications, in the 2022 update.
- Expectancy violation: build every trial around a specific prediction that the experience contradicts.
- Deepened extinction: combine within one trial two stimuli already extinguished separately.
- Removal of safety signals: withdraw the object, person or behaviour that lets the outcome be credited to something other than the patient's own coping.
- Variability: vary intensity, duration and trial order instead of walking a monotonous progression.
- Multiple contexts: repeat the learning in different settings, at different hours, in different company.
- Retrieval cues: tie the learning to a portable cue the patient can call up outside the consulting room.
- Affect labelling: name the emotion out loud during the trial, which dampens the response without recourse to cognitive reappraisal.
Two of them tend to surprise. Variability runs head-on into the intuition of always moving from less to more; the framework holds that an irregular route leaves a more retrievable learning, even if it is less tidy to administer. And removing safety signals forces a review of small details the patient does not declare: the phone in the hand, the door left ajar, the clinician standing one step away.
Multiple contexts is the strategy that sits worst with the conventional consulting room, because the consulting room is a single, heavily marked context. Here a controlled environment contributes something material: it allows the same trial to be repeated with the lighting, the background noise, the number of figures present or the distance altered, without moving the patient from the chair. The clinical protocol for dog phobia exposure develops a series of exactly this kind, and the practice checklist for the consulting room collects the equipment and framing checks that come first.
What does the clinician record when learning, not relief, is the target?
The change of model shows up on the record sheet before it shows up anywhere else. Three columns join the distress rating: the patient's prediction before the trial, what actually happened, and how much credibility they still grant that prediction once it is over. The third column is the one that reports whether learning occurred.
The within-session fall in distress does not leave the record, but it is demoted: it is logged as a phenomenon, not as a criterion of success. The scale that quantifies that distress, subjective units of distress, and the clinical reading of its curve are a separate matter with rules of their own.
Between sessions, the pertinent question is not whether the patient feels better but whether the learning survives a change of conditions. A trial repeated in another context, at another hour or with another person present is far more informative than an identical repetition in the same room.
It is also worth anticipating the return of fear rather than living it as a failure. Spontaneous recovery and contextual renewal are properties of the learning system, not signals that the intervention has come apart. The plan allows for them from the outset: spaced booster trials, agreed retrieval cues, and a framing that explains to the patient why a flare-up does not cancel what was learned.
What are the limits of this model, and what stays a clinical decision?
The inhibitory framework is the best supported account we have of how exposure therapy works, but its reach is worth measuring. Most of the mechanistic evidence comes from laboratory paradigms with experimental conditioning, and the jump to a clinical picture with years of accumulated avoidance is not automatic. The derived strategies have been tested with uneven intensity: some have trials of their own, others still rest largely on theoretical reasoning.
On the overall efficacy of the procedure the evidence is firmer than on its mechanism. The review by Maples-Keller and colleagues maps the clinical applications of the virtual medium, and the meta-analysis by Carl and colleagues finds no significant difference between exposure in virtual reality and in vivo exposure across the conditions where the two have been compared. Equivalence, not advantage: the medium changes the control the clinician holds over the stimulus, not the mechanism that does the work.
That leaves a decision no model settles on the clinician's behalf: which medium to choose for which patient and which step. A virtual environment supplies identical repetition, controlled variation and an immediate exit from the trial; in vivo exposure supplies the world as it is, noise and unpredictability included. The inhibitory framework does not rank the two media by quality, it only specifies what has to happen inside either of them.
To see how all of this translates into a working system, with prediction and outcome logged trial by trial, context varied without leaving the room and safety signals withdrawn in steps, you can book a guided demonstration and walk through the setup of a trial series on a worked example.
This article is for informational purposes for psychology professionals. It is not clinical advice for any individual case and does not replace the judgment of the licensed psychologist in charge. VRET is professional clinical-support software, not a CE-marked medical device.
Frequently asked questions
Does extinction erase the original fear memory?
No. The inhibitory learning framework holds that the original association between the stimulus and the aversive outcome stays in store, and that extinction adds a competing association beside it asserting the opposite. The two coexist and compete every time the stimulus reappears. The three classic return-of-fear findings, spontaneous recovery with the passage of time, renewal on a change of context and reinstatement after an unrelated aversive experience, only make sense if the original learning is still available. The clinical implication is direct: the work is not to delete a memory but to strengthen the new learning and to make it the one that fires first when the patient meets the stimulus outside the consulting room.
Why does fear come back weeks after an exposure that went well?
Because the inhibitory association is context-bound and more fragile than the original. The passage of time produces spontaneous recovery; a change of setting produces contextual renewal; a fresh aversive experience produces reinstatement. None of the three means the series of exposures was wasted. What the clinician weighs at that point is whether the series ran in a single context, whether safety signals stayed active throughout, and whether retrieval cues were agreed that the patient can call up. The intervention plan allows for the flare-up from the start, with spaced booster trials and a framing that explains it in advance. Anticipating it is what stops the patient reading it as relapse and walking away.
What is expectancy violation?
It is the discrepancy between what the patient predicts will happen and what happens during the trial. In Craske's framework it is the engine of inhibitory learning: the wider the discrepancy, the more corrective information the system takes in. It operates in three steps. Before the trial, the clinician obtains a specific, falsifiable prediction with an estimated probability and a test the patient would recognise. During the trial, the design works to have that prediction contradicted by experience. Afterwards, the review sets what was predicted against what was observed without turning it into a rational debate. A vague prediction is useless, because it admits no disproof and leaves the trial with no learning content.
Does anxiety have to come down within the session?
Not as a criterion of success. The 2014 review by Craske and colleagues established that the within-session fall in distress predicts long-term outcome poorly, and that the between-session fall behaves only somewhat better. A trial that ends with high arousal may have been the most productive of the series if it disproved a sharply stated prediction, and a comfortable trial may have contributed nothing. The decision to stop rests on whether the corrective information has been delivered and on the patient's tolerance at that moment, not on a number reaching a threshold. The decline is still logged, but as an observed phenomenon rather than the aim of the trial.
How does the inhibitory framework transfer to exposure in virtual reality?
The mechanism does not change with the medium: it remains the formation and the retrieval of a competing association. What changes is the control the clinician holds over the stimulus. A virtual environment allows the same trial to be repeated with one condition altered, the context to be changed without changing rooms, safety signals to be withdrawn in steps, and prediction and outcome to be logged trial by trial. The meta-analysis by Carl and colleagues finds no significant difference against in vivo exposure in the conditions compared, so the choice of medium is a matter of indication and logistics rather than of efficacy. Walking the full sequence in a guided demonstration is what lets a clinician judge that fit before committing.
Keep reading
Can Online Exposure Therapy Work? The Clinical Limits
Whether online exposure therapy can be done at all, what professional support actually contributes, and which credentials authorise a clinician to deliver it.
Practice managementContraindications to Exposure Therapy: Risks and Cautions
Contraindications to exposure therapy, whether it can make anxiety worse and who it does not suit: exclusion criteria and clinical cautions.
Practice managementSUDS Scale: How to Measure Anxiety During Exposure
What the SUDS scale is, how anxiety is measured during exposure, and which indicators actually show whether a course of treatment is moving.
VRET is professional clinical-support software, not a CE-marked medical device. Clinical supervision remains with the licensed psychologist in charge.