Solutions
1.
Mean = (54 + 61 + 58 + 63 + 59)/5 = 295/5 = 59
Deviations: −5, +2, −1, +4, 0. Squares: 25, 4, 1, 16, 0; sum = 46.
s² = 46/4 = 11.5
s = √11.5 = 3.391
SE(Δ) = s√2 = 3.391 × 1.4142 = 4.796
MDE₉₅ = 1.96 × 4.796 = 9.40
A Δ of 7 is smaller than the smallest movement this instrument can distinguish from its own noise. It is not a small effect; it is not an effect. Report it as unresolved, not as weak.
2.
SE(Δ) = 4.0 × √2 = 5.657
(a) Uncorrected threshold = 1.96 × 5.657 = 11.09. Clearing it: colour (12), motion (16), distance (18).
(b) α' = 0.05/8 = 0.00625; two-tailed, 0.003125 per tail, z ≈ 2.734.
Threshold = 2.734 × 5.657 = 15.47
Surviving: distance (18) and motion (16). Colour drops out. Tempo (10) and size (8) were never close.
(c) 8 × 0.05 = 0.4 expected false positives per sweep — meaning that roughly two sweeps in five will hand a practitioner a spurious "driver" even when every dial on the board is inert. This is the arithmetic behind the field's most durable folk beliefs about which submodalities matter.
3.
(a) E[X₂ | X₁ = 82] = 55 + 0.55 × (82 − 55) = 55 + 0.55 × 27 = 55 + 14.85 = 69.85. The expected drop with no intervention is 82 − 69.85 = 12.15 points.
(b) Observed Δ = 82 − 64 = 18. Attributable = 18 − 12.15 = 5.85.
(c) SD(X₂|X₁) = 14 × √(1 − 0.55²) = 14 × √(1 − 0.3025) = 14 × √0.6975 = 14 × 0.8352 = 11.69.
z = 5.85 / 11.69 = 0.50
(d) Nothing about efficacy. The defensible statement is: "Her reading fell eighteen points; about twelve of those were expected from her own week-to-week variation given how high she was at intake. The remainder is well inside the measurement's noise. If the work did something, this design cannot see it." Say this to the client. It costs you the moment of applause and buys you a client who believes you the next time you say something worked.
4.
I = Δ(a,b) − (Δa + Δb) = 18 − (16 + 14) = 18 − 30 = −12
R = −I / min(Δa, Δb) = 12 / 14 = 0.857
Floor check: S(a,b) = 70 − 18 = 52; the client's effective floor, from the neutral memory, is 6. The combined probe finished 46 points above the floor, so the scale had abundant room and the sub-additivity is not a ceiling artifact.
Verdict: a and b are ~86% redundant. Treat them as one driver, not two, and do not build an intervention that "stacks" them — you will pay twice for a single effect and attribute the shortfall to the client. Next probe: isolate — change b while explicitly holding a's parameter fixed. If b's isolated Δ collapses below threshold, b was a's shadow, exactly as location was distance's shadow in Worked Example 1.
5.
(a) With n occasions each side and the stable component cancelling in the difference:
Var(Δ) = 100/n + 100/n = 200/n
Required: 2.80 × SE(Δ) ≤ 15, so
SE(Δ) ≤ 15/2.80 = 5.357
Var(Δ) ≤ 5.357² = 28.70
200/n ≤ 28.70 → n ≥ 6.97 → n = 7
Seven occasions on each side — fourteen separate measurement days to certify a fifteen-point effect in one person. Report that number honestly to anyone who claims single-session verification.
(b) Spearman–Brown for the mean of n comparable indicators with average intercorrelation r:
ρ = n·r / (1 + (n−1)·r)
n = 2: ρ = 2(0.55)/(1 + 0.55) = 1.10/1.55 = 0.710
n = 3: ρ = 3(0.55)/(1 + 2(0.55)) = 1.65/2.10 = 0.786
(c) Improving the instrument (b) and repeating the occasion (a) attack different variance components, and (a) buys far more here. Adding a third indicator lifts reliability by 0.076, which shaves the ε² term — and ε² is the small term. The occasion variance ω² is untouched by better instruments and only ever falls as 1/n across genuinely separate days. The general rule: a better ruler cannot fix a fluctuating object. Measure on more days.
6.
Complete space: 2⁷ − 1 = 127
OAT probes: 7
Screen + top-3 pairs: 7 + C(3,2) = 7 + 3 = 10
Total pairs: C(7,2) = 21 → tested 3/21 = 14.3%
Total triples: C(7,3) = 35 → tested 0/35 = 0%
The sentence owed: "I tested each of the seven changes on its own and three of the twenty-one possible pairings. If two of these work only in combination — and that does happen — my design will report both of them as inert. I have looked at ten of a hundred and twenty-seven possibilities and I have chosen which ten."
7.
The probes. Two isolation probes, each holding one parameter while moving the other:
- P1 (angle changes, distance held): shrink the image without letting it recede — same felt proximity, smaller apparent size.
- P2 (distance changes, angle held): let the image recede while scaling it up so its apparent size is exactly unchanged.
Predictions. Under H1 (visual angle is the driver), P1 carries the full effect and P2 carries none, since P2 by construction leaves the angle fixed:
H1: Δ(P1) ≈ 20–22 , Δ(P2) ≈ 0
Under H2 (felt proximity is the driver), the reverse:
H2: Δ(P1) ≈ 0 , Δ(P2) ≈ 20–22
With ε = 4.5, SE(Δ) = 4.5√2 = 6.36 and the 95% threshold is 1.96 × 6.36 = 12.5.
Undecidable region. Any pair of results in which both probes land between 0 and 12.5 — say Δ(P1) = 9 and Δ(P2) = 11 — is consistent with either hypothesis plus noise, and also with a third possibility neither hypothesis contains: that the true driver is a joint quantity (a size-distance ratio) that both isolation probes partially disturb. Do not adjudicate by preference. Either raise n on both probes until the threshold falls below the gap, or accept the joint quantity as the driver and intervene on it directly by moving size and distance together in their natural coupling.
A note on the awkwardness of P2: many clients report that scaling an image up as it recedes is effortful or "won't hold." That effort is data, not an obstacle — it suggests the two parameters are yoked in that client's imagery, which is itself the answer to the question you asked.
8.
(a) The sweep is invalid, not merely hard, because it presupposes the thing it measures. Every probe asks the client to alter a property of an object that, for him, has no such properties. What he will do — what almost anyone will do, out of cooperativeness — is report changes: he will answer "dimmer" when asked to dim, and his S will move, because the asking itself directs attention, implies an expected direction, and takes effort. You will obtain a full table of plausible Δ values describing nothing. This is the demand-characteristic failure mode in its purest form, and note that it produces a richer-looking dataset than the valid procedure, not a poorer one.
Congenital aphantasia is a real and reasonably well-characterised variation — Zeman and colleagues named it in the mid-2010s, and self-report instruments for imagery vividness go back to Marks's VVIQ in the 1970s. Prevalence estimates sit in the low single-digit percent, which means that in ordinary practice you will meet these clients regularly, and by chance alone some of them will be the ones whose sweeps produced your most impressive-looking tables.
(b) Run the hunt in the channels he actually has. Concretely: (i) spatial — relative position, direction, near/far, above/below, which of two events stands closer; (ii) auditory — if present, the full set of location, volume, tempo, timbre, distance; (iii) kinaesthetic — location in the body, extent, temperature, weight, direction of movement, tempo; (iv) propositional — tense and person of the verbal encoding ("it is happening to me" versus "it happened to her"), which is a genuine parameter with a genuine literature. Establish ε as before, threshold as before, sham dial as before.
(c) The thing to take seriously rather than route around: spatial imagery is not degraded visual imagery. He is not a visual client with the lights off. Spatial and visual imagery dissociate — congenitally blind people navigate rich spatial representations with no visual content whatever — and his spatial layout is a first-class representational system with its own drivers, not a residue of one. Treat "which is nearer" as the primary datum in its own right. In practice this client is often easier to work with than a vivid visualiser, because the spatial parameters are few and unambiguous, and there is no decorative brightness or colour to waste probes on.
9.
In order:
- Do not make the change today. Answer (2) is load-bearing and answer (3) is empty. That combination is the stop condition, and it is a stop condition regardless of how large or how clean the driver is. D = 4.9 is a reason to be confident you could, not a reason to think you should.
- Say why, plainly. "I found the thing that would turn this down, and I'm not going to turn it down yet, because you just told me it's what makes you walk the building. I'd be taking a check off your work."
- Separate the signal from its intensity. The walk-the-building behaviour is the asset; the intensity is merely its current delivery mechanism, and a costly one. The question is not whether the caution should persist — it should — but whether it must be delivered at this amplitude.
- Build the replacement first. A procedural trigger that does not depend on dread: the walk becomes a fixed step in the pre-entry sequence, prompted by the sequence rather than by the feeling. Rehearse it, and — this is the part usually skipped — measure that it fires on several real occasions while the dread is still fully intact.
- Only then revisit the driver, and re-run the sweep, because the baseline will have moved.
Admissibility condition: the change becomes admissible when the protective behaviour has been observed to occur, in the field, on at least three separate occasions, driven by the replacement trigger and not by the intensity.
The symmetrical error. The stop rule can be applied too readily, and when it is, it becomes a permanent stay of execution: every charged memory can be narrated as protective by someone motivated to keep it, and the practitioner who accepts every such narration will leave clients suffering under intensities that protect nothing. The discriminator is question (1) answered behaviourally and specifically — "it makes me walk the building" is a protective function; "it keeps me on my toes," "it reminds me not to get complacent," "it keeps me humble" are not functions, they are the dread describing itself in flattering terms. Demand the observable behaviour. If none can be named, the memory is not protecting anything, and declining to work is not caution but abandonment.
10.
Setup. Let (X₁, X₂) be bivariate normal with equal means μ, equal SDs σ, correlation r, |r| < 1. Write z₁ = (x₁ − μ)/σ and z₂ = (x₂ − μ)/σ. The joint density is
f(x₁,x₂) = 1/(2πσ²√(1−r²)) · exp{ −(1/(2(1−r²))) · [ z₁² − 2r z₁ z₂ + z₂² ] }
The marginal of X₁ is
f(x₁) = 1/(σ√(2π)) · exp{ −z₁²/2 }
Form the conditional.
f(x₂|x₁) = f(x₁,x₂)/f(x₁)
The constants first:
[1/(2πσ²√(1−r²))] ÷ [1/(σ√(2π))] = √(2π)·σ / (2πσ²√(1−r²)) = 1/(σ√(2π)√(1−r²))
which is already the normalising constant of a normal with SD σ√(1−r²). Now the exponent:
E = −(1/(2(1−r²)))·[z₁² − 2r z₁ z₂ + z₂²] + z₁²/2
Put both terms over the common denominator 2(1−r²):
E = −(1/(2(1−r²)))·[ z₁² − 2r z₁ z₂ + z₂² − z₁²(1−r²) ]
Expand the last product: z₁²(1−r²) = z₁² − r²z₁². So
z₁² − 2r z₁ z₂ + z₂² − z₁² + r²z₁² = z₂² − 2r z₁ z₂ + r²z₁²
which is a perfect square:
= (z₂ − r z₁)²
Therefore
E = −(z₂ − r z₁)² / (2(1−r²))
and
f(x₂|x₁) = 1/(σ√(2π)√(1−r²)) · exp{ −(z₂ − r z₁)²/(2(1−r²)) }
Reading off the standard normal form, in z-units:
z₂ | z₁ ~ N( r z₁ , 1 − r² )
Un-standardising, z₂ = (X₂ − μ)/σ, so E[z₂|z₁] = r z₁ gives
E[X₂ | X₁ = x] = μ + σ · r · (x − μ)/σ = μ + r(x − μ)
Var[X₂ | X₁ = x] = σ²(1 − r²)
(a) r = 0: E[X₂|X₁] = μ. The second reading is entirely predicted by the population mean and nothing is carried over; every apparent change is complete regression.
(b) r = 1: E[X₂|X₁] = x. No regression at all; whatever you observe as change is real change (and the conditional variance is zero, so the model has no room for noise either — an idealisation no real measure meets).
(c) Because 0 < r < 1 for every real instrument, the coefficient on (x − μ) is strictly less than one, so any reading selected for being extreme is expected to move toward μ on its own. Practitioners see clients when clients are at their worst — that is what brings people through the door — so their intake readings are systematically selected on an extreme. The expected drift back is (1 − r)(x − μ), which grows with both the severity at intake and the unreliability of the measure. The less reliable your instrument and the worse your clients are when they arrive, the more spontaneous improvement you will observe and take credit for. A field with no control condition will therefore converge on high confidence, and its most confident practitioners will be the ones who take the most severe cases with the sloppiest measures.
11.
(a) R(a,b) = 0.91 means 91% of the smaller of the two effects is already contained in the larger — a and b are one thing seen twice. R(a,c) = 0.12 and R(b,c) = 0.08 are near enough to zero that c acts independently of both. So: two independent dimensions. Cluster 1 = {a, b}; cluster 2 = {c}.
(b) Under "fully redundant within cluster, additive across clusters," the cluster's contribution is its largest member, not its sum:
Δ(a,b,c) ≈ max(Δa, Δb) + Δc = 22 + 15 = 37
(c) Residual = 34 − 37 = −3.0. With ε = 3.9,
SE(Δ) = 3.9 × √2 = 5.515
z = −3.0 / 5.515 = −0.54
Well inside noise. The two-dimensional model survives this test — which is a weak statement and should be made weakly: a design with an SE of 5.5 cannot distinguish the two-dimensional model from any model predicting between about 26 and 45. It has failed to falsify, not confirmed.
Worth checking for internal consistency: R(a,b) = 0.91 implies I(a,b) = −0.91 × 20 = −18.2, so Δ(a,b) = 42 − 18.2 = 23.8, which is close to max(22, 20) = 22 — the pairwise data and the "take the max" rule agree, as they must if the cluster is nearly one-dimensional.
(d) If a is distance and b is apparent size, the latent dimension is plausibly egocentric proximity — a single "how near is this to me" quantity that both dials happen to address. The probe that tests the name rather than the maths: find a third, previously untested submodality that the name predicts should load on the same dimension and that has no obvious surface relation to either a or b — for instance, the image's occlusion of the visual field, or whether the client can see the scene's edges past it. If proximity is the real dimension, that probe should return a large Δ and be highly redundant with a. If it returns a large independent Δ, the name is wrong and you have found a third dimension. Naming a latent factor is a hypothesis with a prediction attached; a name that predicts nothing new is decoration.
12.
The claim under test. Altering a specified submodality parameter of a mental representation changes state intensity by an amount attributable to the parameter itself, over and above attention, expectancy, effort, and regression.
Sham condition. Not "do nothing" — that controls for nothing. The sham must be a change request identical in every respect except the parameter's plausibility: same operator, same wording, same duration, same demand to visualise and to report, same attentional load. Worked Example 1's "widen the frame by 2 mm" is the shape of it. Better still, use a yoked sham: for each participant, a sham drawn from another participant's decorative submodalities, so that shams are drawn from the same distribution of dial-names as targets and cannot be identified by content.
Blinding. The client cannot be blind to which change was made — they make it. They can, however, be blind to which change is predicted to matter, provided target and sham are indistinguishable in form and the operator does not signal. The operator can be fully blind: they read from a sealed, randomised script and do not know whether the current probe is target or sham. The rater of the state measure should also be blind, which is possible for the observable component of the Ch. 6 composite and impossible for the self-report component — a limitation to state, not to hide. Order must be counterbalanced or randomised within participant, because carryover and fatigue are both directional.
Pre-registered predictions, in advance and in numbers:
- Δ(target) − Δ(sham) > 0, with a stated minimum effect of interest — say 10 points, since anything smaller is clinically inert and, per Worked Example 2, unverifiable anyway.
- The rank order of parameters found in the individual sweep predicts the rank order in a held-out session of the same individual (a split-half within person). This is the discriminating prediction: expectancy predicts a general effect of being asked to change something, but has no reason to predict which specific dial replicates within a person across sessions.
- Dose-response: a graded distance change (1 m, 3 m, 10 m) produces a monotonic graded Δ. Expectancy predicts a step, not a gradient.
Prediction 2 is the one worth building the study around. It is the prediction that non-specific accounts cannot make.
Sample and analysis. A within-person crossover: each participant contributes multiple target probes and multiple shams across at least two sessions, analysed with a mixed model with random intercepts by participant, participant-by-session as a nested random effect, and probe order as a fixed covariate. Report the target-minus-sham contrast with its confidence interval, and report the within-person rank correlation (Spearman's ρ) between session 1 and session 2 driver orderings with a permutation null.
The result that obliges abandonment. State it before running:
- If the target-minus-sham contrast's 95% CI is contained within ±10 points, the specific-parameter claim is dead in its practically relevant form, whatever the p-value says.
- If the within-person driver ordering fails to replicate across sessions — ρ not distinguishable from zero — then "finding the driver" is finding nothing, and the entire clinical procedure of this chapter is a ritual, because a driver that does not persist to the next session cannot be intervened on.
- If the dose-response is flat or non-monotonic, the parameter is not acting as a parameter.
Any one of these is sufficient. Write them down and mean them.
The strongest objection, stated at full strength. It is this: the specific empirical record of NLP is bad, and it is bad in exactly this area. NLP's two most testable signature claims — that eye movements index representational systems, and that matching a client's preferred representational system improves rapport and outcome — were tested repeatedly through the 1980s, and the reviews (Sharpley's, in the mid-1980s; Heap's, at the end of the decade; Witkowski's survey of thirty-five years, in 2010) found the evidence unsupportive. Submodality claims are of the same family, generated by the same method — introspection by a small number of practitioners, propagated by demonstration — and have received far less testing than the claims that failed. Meanwhile, every element of the observed effect has a mundane explanation already in hand: demand characteristics, effort justification, attentional redirection, expectancy, and regression to the mean, the last of which we have just shown can manufacture an eight-point improvement from nothing at all. On the base rates, the prior on submodality specificity should be low.
The answer, which is partly a concession. Concede the historical point entirely: the provenance of these claims is poor, the field's own evidence for them is thin, and anyone teaching them as established is misrepresenting the record. What does not follow is that the effects are absent, because the general claim — that manipulating the parameters of a mental image changes its emotional force — has independent support from outside NLP entirely, and that support is not weak. Andrade, Kavanagh and Baddeley showed in 1997 that loading the visuospatial sketchpad with eye movements reduces both the vividness and the emotionality of an autobiographical image, and reduces them together, which is a parametric manipulation of imagery with an emotional consequence and a specified working-memory mechanism. Nigro and Neisser's 1983 distinction between field and observer perspective in personal memory is a submodality by another name, and Kross and Ayduk's programme on self-distanced versus self-immersed recall has found repeatedly that shifting perspective on a negative memory reduces emotional reactivity. Holmes and Mathews's work on the special relationship between imagery and emotion supplies the theoretical spine.
So the honest position, and the one this chapter takes: the general mechanism has real support from outside the tradition that named it; the specific catalogue of two dozen dials, the claim that each is a separable parameter, and the practice of reading a person's drivers from a single in-room sweep have essentially none. That asymmetry is precisely why the measurement discipline of Chapter 6 is not an accessory to this material but a condition of practising it at all. You are working with a mechanism that is probably real, through a technique whose specifics are unvalidated, on a single person, with an instrument you built this morning. The arithmetic in these three worked examples is not pedantry. It is the only thing standing between you and thirty-five years of confident practitioners.