2. Five Channels and the Shape of a Thought
A woman describes her marriage as having "gone quiet." Not cold, not distant, not broken — quiet. Pressed for more, she says the house "sounds different now," that she can "hear him not saying things," that when he does speak it "lands flat." Her husband, describing the same marriage in the same room an hour later, says he feels "boxed in," that there's "no room to move," that every conversation "has walls." Neither of them is being metaphorical on purpose. Neither has noticed that they are describing the same eighteen months in two languages that do not share a noun.
This is the observation that chapter one's method — read the model off the behaviour of the person who has one — first cashes out into something you can actually do. The model is private. The sentence is public. And the sentence carries, in a place most listeners never look, a report on the sensory format in which the model is currently being run.
The exposed surface
Look at the verbs.
Not the nouns, which carry the topic, and not the adjectives, which carry the evaluation. The verbs and their attached prepositions and the odd predicate adjective riding along with them — this is where the format leaks. I see what you mean. That doesn't sound right to me. I can't get a grip on it. It's all a bit hazy. Let's touch base. Something rings false. That's a bright idea. I'm not following you. He rubbed me the wrong way. It struck me as odd. I need to get clear on this.
Every one of those is a sensory claim about a non-sensory situation. Nobody literally sees a meaning or touches a base. The speaker has reached, unconsciously and in under a second, for a verb belonging to a specific sense channel, and the reach is not random. In NLP's vocabulary these are predicates, and the classes are conventionally abbreviated: V for visual, A for auditory, K for kinesthetic (which bundles tactile sensation, proprioception and emotion, an overloading we will have to unpack shortly), O for olfactory, G for gustatory. A sixth class matters enormously and is often mis-sorted: A<sub>d</sub>, auditory-digital, meaning internal language — words about words. I understand. That doesn't make sense. Let me think it through. The logic doesn't hold. Understanding, sense, logic, thinking-through: these are not pictures or sounds or feelings. They are propositions about propositions.
And then there is the large, uninformative remainder. I think we should proceed. I'm considering it. That's important. It matters to me. I'd like to work on this. These are unmarked predicates — verbs with no sensory allegiance at all — and the honest thing to say about them is that they tell you nothing. This matters more than it sounds. A practitioner freshly taught predicates will begin sorting every sentence into a bin, and roughly half of ordinary speech has no bin to go in. Forcing it into one is the first corruption of the data. When someone says "I think it matters," the correct entry in your log is a blank.
Here is the first thing that is simply, boringly true and requires no defence from any contested literature: people do this, it is measurable, and it varies within a single conversation. Count predicates in a transcript and you get a distribution. Change the topic and the distribution changes. That is not a theory. That is arithmetic performed on a recording.
Four working systems and a fifth that mostly sleeps
The visual, auditory, kinesthetic and auditory-digital systems carry nearly all the load in ordinary talk, and they are not interchangeable in what they can hold.
The visual system holds relations simultaneously. A whole diagram arrives at once; you can inspect its parts in any order; spatial arrangement — above, beside, behind — is native to it and costs nothing extra. This is why people reach for it when they want to see a situation whole, and why the phrase "I can't picture it" so often accompanies a genuine inability to hold several elements in relation.
The auditory system is inherently sequential and carries tempo, pitch, rhythm and — critically — tone, which is to say the entire register of relational meaning. Sarcasm has no visual form. The auditory channel is where other people's attitudes toward you are stored. When someone recalls a humiliation, the sentence that did it usually comes back with its intonation intact.
The kinesthetic system is doing at least three jobs under one letter, and NLP's failure to consistently distinguish them is a real defect in its notation. There is tactile sensation, external and locatable: the roughness of a fabric, warmth on the forearm. There is proprioceptive sensation: balance, weight, effort, the felt position of the body in space. And there is emotion, which is what practitioners almost always mean by K and which has a spatial and qualitative signature of its own — the tightness in the throat, the drop in the chest. These behave differently and respond differently to intervention. When you read K in a transcript, note which of the three it is; the technique chapters will need to know.
Auditory-digital deserves its own treatment and gets it below.
Which leaves the fifth: gustatory-olfactory, G/O, and here you deserve honesty rather than symmetry. Most textbook treatments list it alongside the others as though the five formed a tidy set. In practice, for the overwhelming majority of the conversations you will have as a coach, therapist, negotiator or teacher, smell and taste are not load-bearing. People do not solve a career decision by smelling it. The predicates exist — that leaves a bad taste, something smells off about this deal, a sweet arrangement — but they are largely frozen idiom, and idiom is precisely the thing predicate analysis cannot read.
There are two domains where this reverses hard, and they are not trivial exceptions. The first is trauma. Olfactory input reaches limbic and cortical structures by an unusually direct route — the olfactory bulb projects to piriform cortex and onward to amygdala and entorhinal cortex without the obligatory thalamic relay that vision and audition pass through — and odour-cued autobiographical memories are consistently reported as more emotionally vivid and more strongly felt than memories cued by word or image, even where they are not more accurate. Anyone who has worked with combat veterans or with survivors of assault has met the smell that reinstates the whole scene in one second. The second is appetite: eating, drinking, craving, disgust, and every disorder built on top of those. In both domains, G/O is not decorative. Everywhere else, treat it as a rare bird — record it when it appears, and don't go hunting.
That asymmetry is worth sitting with, because it is a small instance of something the field does at scale: a category is completed for the sake of the category's elegance rather than because the fifth member does comparable work. Five channels is a prettier system than four-and-a-half. Pretty is not a warrant.
The voice in the room where nobody is speaking
Internal dialogue is the channel most often waved at and least often examined. Practitioners call it A<sub>d</sub> and move on. It rewards three specific questions, and none of them is "what is it saying."
Tempo. Internal speech has a rate, and the rate is not constant. Rumination characteristically runs fast, above the speaker's normal speaking rate, with compressed pauses; the content loops rather than developing. Deliberation runs slower with real gaps. Ask someone how fast the voice is talking and you will frequently get an immediate, confident answer about a parameter they have never once considered — which is itself diagnostic, and which chapter seven will teach you to alter directly.
Grammatical person. Internal speech comes in first person (I can't do this), second person (you can't do this, or its harsher cousin you idiot), and occasionally third. This distinction is not cosmetic. There is a reasonably robust experimental literature on distanced self-talk — work associated with Ethan Kross and colleagues, using second-person and own-name self-address — finding that non-first-person self-talk during stressful tasks is associated with reduced distress and better performance under evaluation than first-person self-talk. Note the direction carefully, because it cuts against the intuition most practitioners carry: distancing is generally the helpful condition in that research. Meanwhile, clinically, second-person internal criticism is the standard grammar of a harsh internal critic. Both can be true — the same grammatical distance that cools a stress response also permits a voice to address you as an other, and whether that other is a coach or a prosecutor is a different variable. What you should take from this is that person is a real dial, that it is not a simple good/bad axis, and that you must ask rather than assume.
Whose voice is it. Ask a person to slow their internal speech down and attend to its timbre, and a striking proportion will report — often with visible surprise — that it is not their own. It is a parent's. A teacher's. A particular coach, a particular ex. This is not a mystical finding and needs no exotic mechanism: internal speech is assembled from heard speech, and heard speech carries a speaker. It is, however, enormously consequential for what you do next, and chapter twelve will build directly on it. For now, the discipline is this: the question "whose voice is that?" is a question, not a leading suggestion. If the person says "it's mine," it's mine. You do not get to install a parent because the technique would be more interesting with one.
When a picture manufactures a feeling
Watch what happens when the woman in the opening paragraph is asked what "quiet" is like. She pauses, and says the house is quiet the way it was when her mother stopped speaking to her father. Then her eyes fill.
Nothing kinesthetic was reported. A sound was reported, and then a scene, and the feeling arrived downstream of it. This is a synaesthesia pattern — in NLP's usage, not the neurological condition of the same name, but a habitual cross-channel link in which activity in one representational system reliably triggers activity in another. The notation is an arrow: A→K for a sound that produces a feeling, V→K for a picture that produces one, K→A<sub>d</sub> for a sensation that starts the internal commentary running.
V→K is the workhorse and worth learning first. Someone reports anxiety about a presentation. Ask what happens just before the anxiety and you frequently get an image — a specific one, often oddly detailed, often not obviously about presentations: the audience seen from an angle nobody would actually occupy, faces enlarged, the room dim. The image is not a report on the anxiety. It is upstream of it. And this is precisely why chapter seven's submodality work has any purchase at all: if the feeling is manufactured by a picture, and the picture has adjustable parameters, then the feeling has adjustable parameters at one remove, which is far easier to reach than the feeling itself.
Record the arrow, don't assume it. The direction is empirical. Some people's images follow their feelings rather than causing them, and running a V→K intervention on a K→V person produces a puzzled client and a practitioner who concludes the client is resistant. Chapter ten will need this notation formally, because a strategy is nothing but a chain of these arrows with tests attached; introduce it now so that the notation is old news by the time it has to carry weight.
Three questions that are not one question
Here is where practitioners routinely go wrong, and where a small amount of precision pays for the whole chapter.
There are three different things one might mean by "which system is she using," and they can have three different answers in the same thirty seconds.
The lead system is the channel that initiates retrieval — the one that goes first when the person reaches for information. Ask someone how many windows are in their house and something happens before the answer: most people construct an image and count. The image is the lead.
The representation system is the channel in which the information is consciously held and reported — what the person is aware of having, and what their predicates describe.
The reference system is the channel used to evaluate — the one consulted to decide whether the answer is right, whether the decision is good, whether the thing is true. It is the exit test, and chapter ten will name it as such.
These come apart routinely. Watch:
You: How was the flat you saw yesterday?
Him: Yeah — good, I think. Good light. High ceilings, big windows at the front. [V representation]
You: How are you working that out — what happened just then, before you answered?
Him: I sort of walked back in. Came through the door and down the hall. [The lead is K — proprioceptive, a route through space. The pictures came after the walk.]
You: And how will you know whether to take it?
Him: I'll know it's right when I stop hearing the objections. When I can run it past my sister in my head and there's nothing coming back. [Reference is A<sub>d</sub> — an internal dialogue with a specific interlocutor, tested for silence.]
One person, one topic, sixty seconds: leads K, represents V, references A<sub>d</sub>. Now consider what a practitioner does who has collapsed these three into "he's visual." They match visual predicates, offer visual homework, and put all their intervention into a channel that is neither where his information enters nor where his decisions are made. The elicitation above took three questions. The collapse takes none, which is exactly why it is popular.
Note also what the third question actually did. It did not ask what he wants. It asked how he will know — and got back a testable criterion in a specific channel. Hold onto that; chapter five is built out of it.
What the evidence actually says
Now the harder half of this chapter, and the reason it exists in this form.
From the mid-1970s onward, NLP made a further claim that goes considerably beyond anything above: that each person has a preferred representational system — a stable, characteristic channel that is theirs — that it can be read from their predicates and, more famously, from their eye movements, and that a practitioner who detects it and matches it will obtain better rapport, more influence and better outcomes.
That claim was tested. Repeatedly. It is one of the more thoroughly investigated propositions in the history of applied psychology, in the specific sense that it attracted a genuine, sustained programme of studies in the late 1970s and 1980s.
Christopher Sharpley reviewed the accumulated research twice, in 1984 and again in 1987 after Einspruch and Forman's methodological rebuttal. His conclusion across both reviews was that there was little support for the preferred-representational-system construct, and correspondingly little for the claim that matching predicates to it produced the promised gains in rapport or counselling effectiveness. The 1987 review engaged the methodological objections directly and did not find that they rescued the result.
Michael Heap's 1988 survey, which sat within the broader body of work he assembled reviewing NLP's empirical status, reached a similar verdict on the same core claims: the preferred-system hypothesis and the predicate-matching prescription were, as stated, unsupported by the research that had been conducted to test them.
The eye-movement claim — the famous one, the diagram in every seminar handout, up-left for remembered images, up-right for constructed, down-right for feelings, down-left for internal dialogue — was tested most directly and most publicly by Richard Wiseman and colleagues in a study published in 2012, which examined the further folk-elaboration that eye direction indicates lying. It found no supporting relationship between eye direction and deception, and a second study training observers on the pattern produced no improvement in lie detection. The eye-accessing-cue model, in the strong form in which it has been sold, does not stand.
Now be equally precise about what was not shown, because sloppiness here is how the field's defenders and its debunkers both get to be wrong.
What failed was a hypothesis with a specific shape: stable trait, single dominant channel, readable from a fixed universal code, exploitable by mechanical matching. Every one of those four elements is separable, and the studies falsified them as a bundle.
What was not tested, or was tested only glancingly: whether within-person, within-task channel shifts are detectable and meaningful; whether sensory-specific questioning recovers information that abstract questioning does not; whether the systematic within-session eye-movement patterns of an individual, established against that person's own baseline rather than against a universal chart, carry any information at all. These are answerable questions. They were largely skipped, because the field was busy defending the strong claim and its critics were busy — legitimately — knocking the strong claim down. A generalisation that would have been modest and possibly true was overwritten by one that was dramatic and false, and then the whole thing went down together.
The error, and its signature
Here is what actually went wrong, and it is worth more than the finding itself.
The preferred-system hypothesis did not fail because sensory representation is unreal. It is manifestly real; the mental-imagery literature has been accumulating evidence for decades that visual imagery engages overlapping neural machinery with visual perception, that imagery has measurable properties — Kosslyn's scanning studies, Shepard and Metzler's rotation work, in which response time scales linearly with the angular difference between two shapes as though something were literally being turned — and that individual differences in imagery vividness are real and correlate with performance on imagery-dependent tasks. People with aphantasia, who report no voluntary visual imagery at all, differ measurably from controls on such tasks. Nothing about the falsification of preferred systems touches any of this. Sensory-format differences exist.
The hypothesis failed because it took a variable that is state-dependent and task-dependent and reified it as a stable personality trait.
Ask someone to describe their kitchen and they go visual. Ask them to recall an argument and they go auditory. Ask them how the interview went and, if it went badly, they go kinesthetic. The channel is a function of what is being retrieved and what state the person is in while retrieving it. Measure it once and you have measured a moment. Label the person from that moment and you have converted a reading into an identity — and the reading will not replicate on Tuesday, which is precisely what the studies found.
Once you can see that error cleanly, you will find you cannot stop seeing it. It is the single most common failure in the entire business of human typology, and it has a signature you can check in about ninety seconds. Take any instrument that sorts people into kinds — learning styles as visual, auditory or kinesthetic learners; personality types with four-letter codes; colour-coded temperaments; love languages; whatever the current season is selling. Ask three questions of it. Does the instrument measure a behaviour that plausibly varies by context and mood? Does it then assign a durable label? And is its test-retest reliability across a meaningful interval either weak or quietly unreported? Where the answers are yes, yes, and yes, you are looking at the same error, and the fact that the underlying phenomenon is real will not save the typology built on top of it — learning-styles research is the largest instance, where genuine modality differences in stimulus processing failed utterly to produce the predicted matching effect on instruction.
This is the reason to teach the falsification in this book rather than around it. You did not come here to memorise which NLP claims survived. You came to acquire a discrimination, and this is the first one that transfers well beyond the field: state-dependent variable reified as trait. Half of what you will ever be sold, inside this discipline and well outside it, is that mistake wearing a chart.
What is still standing
Strip out the trait claim and something durable remains, and it is the thing you will actually use.
A private state is not directly observable. You cannot see someone's dread. What you can obtain is a description, and descriptions come at different resolutions. "I'm anxious about the meeting" is a low-resolution report: it names a category, offers no structure, and gives you nothing to change. "There's a picture of the room, dim, from up near the ceiling, everyone's face turned toward me, and it goes still — and then something drops in my chest and my throat tightens" is a high-resolution report of the same state. It has channels, order, spatial properties, a V→K arrow, and about nine separately alterable components.
Sensory-specific description is the best available handle on a private state precisely because it converts a category into a structure. That is not a claim about which channel someone prefers. It is a claim about resolution, and it survives every study cited above intact, because none of them tested it.
And predicates remain worth tracking for a reason unrelated to matching: they tell you what the person can currently access. When somebody's account of a problem is entirely A<sub>d</sub> — all reasoning, all should and logic and makes-sense — and contains no K at all, that absence is data. Something has been walled off. When a person who has described a project visually for ten minutes shifts abruptly to K on one sub-topic, the shift is the finding. Not the channel: the change. Which is the whole subject of chapter three, and the reason it comes next.
The failure mode
Predicate matching, performed as a technique, consumes the exact attentional resource that listening requires.
Watch a newly trained practitioner do it. There is a small delay before each of their replies. Behind the delay is a search: she said "see," I need a visual verb, what's a visual verb. The reply arrives grammatically matched and half a beat late, and it addresses the form of what was said while missing the content. Clients report this as a specific and unpleasant experience — being handled. They rarely have a name for it. They notice that the person opposite seems to be doing something rather than listening, and they close down by a degree that no amount of matched predicates will reopen.
There is a compounding cost. Attention is finite, and the practitioner running a predicate-search loop has spent theirs. They will miss the pause before the answer, the change in breathing, the moment the topic shifted and the channel shifted with it — all the material chapter three is about to insist is the actual signal. They have traded a large real skill for a small doubtful one. And note the irony precisely: the technique's justification was rapport, and its execution destroys rapport. That is what a failure mode looks like — not a technique that does nothing, but one that inverts past a threshold and does damage in the name of its own purpose.
The edge is sharp and easy to state. Track predicates receptively — as information about what channel this person's account is currently running in — and you are doing something cheap, honest and useful. Deploy them productively, choosing your words to match theirs while the search occupies your foreground, and you are performing a parlour trick with your client's hour. If matching ever becomes natural it will be because your own language has loosened enough to follow another person's without effort, and that is a consequence of long practice, not a technique you install on a Tuesday.
The practice
Carry a log through ten ordinary conversations — work, family, the phone, the queue, it does not matter, and the more ordinary the better. Record verbs only. Not topics, not judgments, not your interpretation: the verb and its class, V, A, K, A<sub>d</sub>, G/O, or a blank where the predicate is unmarked, and let the blanks stay blank. You will resent them at first; they are half your data and they are teaching you the limits of the instrument.
Then attend to one thing above all others. Mark the moment the channel shifts, and mark what the topic was doing when it shifted. Someone talks about their quarter in clean visual terms and then, arriving at one particular colleague, goes kinesthetic for two sentences and back. That is not noise. You do not yet know what it means, and chapter three will spend itself teaching you the discipline of not deciding. But you have found the seam, and you found it by attending to change rather than to state — which is, in miniature, the correction to the entire error this chapter has been dismantling.
When the log has ten conversations in it, turn the instrument on yourself. Take a problem of your own that is genuinely stuck — not a difficult one, a stuck one; the distinction matters, because a difficult problem yields to effort and a stuck one has structure holding it in place. Describe it fully three times, on paper or aloud, once in each of the three main channels. Visually: what is the picture, where is it, how far away, what is in the frame and what has been cropped out. Auditorily: what is being said and by whom, in what tone, at what tempo, and what the silences are doing. Kinesthetically: where in the body it sits, what it weighs, whether it moves, what direction it moves in.
Then record which rendering makes the problem feel workable — not which is most accurate, not which is most comfortable, but which one leaves you with something you could actually take hold of tomorrow morning.
Whatever that answer is, do not conclude that it is your channel. It is this problem's channel, today, and the next stuck problem may answer differently. Log it as a reading. That restraint, held consistently, is the difference between a practitioner and a convert, and you have just practised it on the only subject who will never let you get away with cheating.
Worked Examples — Chapter 2
The three examples below take the chapter's two facts and put numbers under them. The first shows what a predicate census can and cannot license. The second shows what a genuine modality effect looks like when you measure it. The third shows, in arithmetic, why the preferred-representational-system claim died — and, more precisely, what kind of death it was, because the usual summary ("the research didn't support it") hides the interesting part.
A note on the data. The transcript in 2.1 is a composite, written for this exercise; the coding procedure is identical to the one you would apply to a real recording, and that procedure is the content. The reaction-time figures in 2.2 are condition means patterned on the published mental-rotation literature (Shepard & Metzler, Science, 1971), rounded for hand computation. The reliability figure in 2.3 is stipulated at a value consistent with what the review literature reports about classification instability; it is not a quoted statistic from any single paper. Where a number is stipulated, it is labelled. Never let a teaching number migrate into a claim.
Worked Example 2.1 — The predicate census, and the null hypothesis nobody chose
The material. R. is a hospital unit manager, twelve minutes into a supervision session, describing why a rota redesign stalled. Excerpt:
"I can see where it went wrong, honestly. The picture we showed the board in March was clear — three shifts, colour-coded, everybody could look at it and get it. Then it went quiet. Nobody rang, nobody said anything, and I just had this heavy feeling in my stomach that we'd lost them. When I finally got hold of Priya they sounded fine, said it all looked good from their end, but something didn't sit right. I keep going back over the images from that meeting. Marcus leaning back. That's when it slipped. And now I'm holding the whole thing together with — I don't know — with tape. Every time I try to grasp what actually needs to happen next it goes blurry on me."
Step 1 — Define the coding unit. A predicate here means a verb, adjective, or adverb whose literal sense belongs to a sense modality, used to describe a mental or interpersonal process rather than a physical event. "The picture we showed the board" counts as visual: the object was a slide, but the process being described is one of display and apprehension. "Marcus leaning back" does not count: that is a report of a physical fact in the world, not a description of a process in R.'s mind or in the relationship. This boundary is where most inter-rater disagreement lives, so it is written down before coding, not after.
Step 2 — Assign each predicate to a class. Let the classes be V (visual), A (auditory), K (kinesthetic, including tactile and affective-somatic), and U (unclassifiable or cross-modal).
- V: see, picture, showed, clear, colour-coded, look at it, looked, images, blurry → 9
- A: quiet, rang, said, sounded, said → 5
- K: heavy feeling in my stomach, got hold of, didn't sit right, slipped, holding, grasp → 6
- U: "with tape" (a noun-phrase metaphor, not a predicate), "leaning back" (physical report) → 2
Excerpt totals: V = 9, A = 5, K = 6, U = 2.
Step 3 — Scale to the full segment. Coding the whole twelve-minute segment (about 1,800 words) gives:
- V = 21, A = 8, K = 13, U = 6.
- Classified total N = 21 + 8 + 13 = 42. U is excluded from the modality test, because the test asks how the classified predicates distribute; including a residual category as a fourth cell answers a different question and inflates degrees of freedom.
- Predicate density: 42 classified predicates / 1,800 words ≈ 2.3 per 100 words. Hold on to this figure; it does the real work at the end.
Step 4 — Test against the naive null. The naive null is that the three modalities are equiprobable: p_V = p_A = p_K = 1/3. Under H₀ the expected count in each cell is
E = N/3 = 42/3 = 14.
Pearson's goodness-of-fit statistic is
χ² = Σᵢ (Oᵢ − Eᵢ)² / Eᵢ
where Oᵢ is the observed count in cell i and Eᵢ the expected count. Term by term:
V: (21 − 14)² / 14 = 49 / 14 = 3.5000
A: (8 − 14)² / 14 = 36 / 14 = 2.5714
K: (13 − 14)² / 14 = 1 / 14 = 0.0714
χ² = 3.5000 + 2.5714 + 0.0714 = 6.1428
Degrees of freedom: df = (number of cells) − 1 − (parameters estimated from the data) = 3 − 1 − 0 = 2. For df = 2 the chi-square distribution is exponential, and the p-value has a closed form:
p = exp(−χ²/2) = exp(−3.0714) = 0.0463
That clears the conventional threshold. On this test, R. "is a visual."
Step 5 — Test against the null you should have chosen. The naive null assumes that English distributes sensory predicates evenly across modalities. It does not, and neither does any speaker discussing a project with slides in it. So build an empirical baseline: code the four other speakers in the same meeting, on the same topic, and pool their predicates. Suppose that pooling gives
p_V = 0.55, p_A = 0.17, p_K = 0.28
(illustrative values, derived within this exercise from the comparison speakers — not a published corpus statistic; when you do this for real you compute your own from your own comparison sample). Expected counts for N = 42:
E_V = 42 × 0.55 = 23.10
E_A = 42 × 0.17 = 7.14
E_K = 42 × 0.28 = 11.76
(check: 23.10 + 7.14 + 11.76 = 42.00 ✓)
Then:
V: (21 − 23.10)² / 23.10 = 4.4100 / 23.10 = 0.19091
A: (8 − 7.14)² / 7.14 = 0.7396 / 7.14 = 0.10358
K: (13 − 11.76)² / 11.76 = 1.5376 / 11.76 = 0.13075
χ² = 0.19091 + 0.10358 + 0.13075 = 0.42524
p = exp(−0.42524 / 2) = exp(−0.21262) = 0.808
Nothing. R.'s predicate distribution is indistinguishable from that of the people sitting around the same table. The apparent effect in Step 4 was not a fact about R.; it was a fact about English, and about what people say when the object under discussion is a colour-coded rota on a screen.
Step 6 — Put an interval on the estimate. Even setting the null aside, ask how precisely we have measured R.'s visual proportion. The point estimate is
p̂ = 21 / 42 = 0.500
The standard error of a sample proportion, assuming independent observations, is
SE = √( p̂(1 − p̂) / N ) = √( 0.25 / 42 ) = √0.0059524 = 0.07715
and the 95% Wald interval is
p̂ ± 1.96 × SE = 0.500 ± 1.96(0.07715) = 0.500 ± 0.1512 → [0.349, 0.651]
Thirty percentage points wide. The interval includes 0.35 — barely above the equiprobable 0.333 — and 0.65, which is emphatically visual. Twelve minutes of speech does not distinguish those two people.
Step 7 — Cost out the precision you would need. To get a margin of error m = 0.05 at 95% confidence, invert the formula. From m = 1.96 √(p(1−p)/N),
N = (1.96)² p(1 − p) / m²
At the worst case p = 0.5, which maximises p(1−p):
N = 3.8416 × 0.25 / 0.0025 = 0.9604 / 0.0025 = 384.16 → 385 predicates
At the observed density of 2.3 classified predicates per 100 words:
words = 385 / 0.023 ≈ 16,700 words
and at a conversational rate of roughly 150 words per minute:
minutes = 16,700 / 150 ≈ 111 minutes
Just under two hours of continuous speech from one person to establish that person's visual proportion to ±5 points — and only under the assumption that the proportion is constant across those two hours, which is the very thing in dispute.
Step 8 — Remove the assumption that inflated the precision. Predicates are not independent draws. They cluster: once a speaker is inside a visual description, the next three predicates are likelier to be visual. Let m̄ be the mean number of predicates per conversational turn and ρ the intraclass correlation among predicates within a turn. The variance inflation factor (design effect) is
D = 1 + (m̄ − 1)ρ
With m̄ = 6 and a modest ρ = 0.15:
D = 1 + 5(0.15) = 1.75
SE_eff = SE × √D = 0.07715 × 1.32288 = 0.10206
95% CI = 0.500 ± 1.96(0.10206) = 0.500 ± 0.2000 → [0.300, 0.700]
Forty points wide. And the required sample scales by D as well: 385 × 1.75 ≈ 674 predicates, roughly 3.3 hours of speech per person.
What this example establishes. The mixture is real and measurable — R.'s speech genuinely runs 50% visual, and you heard it before you counted it. What is not measurable at any tolerable cost is R. as a type. And the deeper point is the one hiding in Step 5: the predicate census, done against the wrong null, manufactures individual differences out of shared linguistic base rates. This is not a small methodological slip. It is the engine that made the preferred-representational-system idea look confirmed to everyone who tried it informally, because informal testing has no comparison speakers in it. You listen to one person, you hear a lot of visual language, and you have no way of knowing that everybody sounds like that.
Worked Example 2.2 — What a real modality effect looks like, in milliseconds
The claim that survived Chapter 2's demolition is stronger than the one that died, and it has a quantitative signature. Here is how you compute it.
The paradigm. A participant sees two line drawings of three-dimensional block figures. The figures are either identical or mirror images, and one is rotated relative to the other by an angle θ. The participant presses "same" or "different" as fast as they can. Let RT(θ) be the mean reaction time in milliseconds for correct "same" responses at angular disparity θ, measured in degrees.
The data (condition means over many trials; means are far less noisy than single trials, which is why the fit below is tighter than any single participant's data would be):
θ (deg): 0 40 80 120 160
RT (ms): 1000 1650 2280 2950 3560
Step 1 — Set up the least-squares problem. Fit the model
RT(θ) = a + bθ + ε
where a is the intercept in milliseconds (the time consumed by everything that is not rotation: encoding the figures, comparing them, choosing and executing a keypress), b is the slope in milliseconds per degree, and ε is residual error. Least squares chooses a and b to minimise
S(a, b) = Σᵢ (yᵢ − a − bxᵢ)²
where xᵢ = θᵢ and yᵢ = RT(θᵢ), over the n = 5 conditions.
Step 2 — Derive the estimators. Differentiate and set to zero. For a:
∂S/∂a = −2 Σᵢ (yᵢ − a − bxᵢ) = 0 ⟹ Σᵢ yᵢ = na + b Σᵢ xᵢ ⟹ ȳ = a + b x̄
For b:
∂S/∂b = −2 Σᵢ xᵢ(yᵢ − a − bxᵢ) = 0 ⟹ Σᵢ xᵢyᵢ = a Σᵢ xᵢ + b Σᵢ xᵢ²
Substituting a = ȳ − b x̄ into the second equation and rearranging gives the standard forms
b = Sxy / Sxx, where Sxy = Σᵢ (xᵢ − x̄)(yᵢ − ȳ) and Sxx = Σᵢ (xᵢ − x̄)²
a = ȳ − b x̄
Step 3 — Compute the means.
x̄ = (0 + 40 + 80 + 120 + 160)/5 = 400/5 = 80 degrees
ȳ = (1000 + 1650 + 2280 + 2950 + 3560)/5 = 11440/5 = 2288 ms
Step 4 — Compute the deviations.
x − x̄: −80, −40, 0, +40, +80
y − ȳ: −1288, −638, −8, +662, +1272
Step 5 — Compute the sums of products and squares.
Sxy = (−80)(−1288) + (−40)(−638) + (0)(−8) + (40)(662) + (80)(1272)
= 103,040 + 25,520 + 0 + 26,480 + 101,760
= 256,800
Sxx = (−80)² + (−40)² + 0² + 40² + 80²
= 6,400 + 1,600 + 0 + 1,600 + 6,400
= 16,000
Step 6 — Solve for slope and intercept.
b = 256,800 / 16,000 = 16.05 ms per degree
a = 2288 − (16.05)(80) = 2288 − 1284 = 1004 ms
Step 7 — Check the residuals. Fitted values ŷ = 1004 + 16.05θ:
θ = 0: ŷ = 1004 residual = 1000 − 1004 = −4
θ = 40: ŷ = 1646 residual = 1650 − 1646 = +4
θ = 80: ŷ = 2288 residual = 2280 − 2288 = −8
θ = 120: ŷ = 2930 residual = 2950 − 2930 = +20
θ = 160: ŷ = 3572 residual = 3560 − 3572 = −12
Σ residuals = −4 + 4 − 8 + 20 − 12 = 0 ✓ (a necessary consequence of the normal equations)
Sum of squared errors:
SSE = 16 + 16 + 64 + 400 + 144 = 640
Step 8 — Put an interval on the slope. The residual variance estimate is
s² = SSE / (n − 2) = 640 / 3 = 213.33, s = 14.61 ms
(n − 2 because two parameters were estimated from the data.) The standard error of the slope is
SE(b) = s / √Sxx = 14.61 / √16,000 = 14.61 / 126.49 = 0.1155 ms/deg
With df = n − 2 = 3, the two-tailed 5% critical value is t = 3.182, so
95% CI for b = 16.05 ± 3.182(0.1155) = 16.05 ± 0.368 → [15.68, 16.42] ms/deg
t = b / SE(b) = 16.05 / 0.1155 = 139.0
Step 9 — Convert the slope to a rate, which is where the meaning is. The slope has units of milliseconds per degree; its reciprocal is degrees per millisecond, and scaling by 1,000 gives degrees per second:
rate = 1000 / b = 1000 / 16.05 = 62.3 degrees per second
Applying the same conversion to the interval endpoints (and noting that the reciprocal reverses the order):
1000 / 16.42 = 60.9 deg/s and 1000 / 15.68 = 63.8 deg/s → [60.9, 63.8] deg/s
Step 10 — Read what this means. The intercept, 1004 ms, is the fixed overhead: encode, compare, decide, press. The slope says that something in the participant is turning the figure, and turning it at a rate you can quote to three significant figures — about one full revolution every six seconds. Nothing in the stimulus rotates; nothing in the retina rotates. The linearity is the evidence: if the comparison were made on an abstract, modality-free description of the figures, there would be no reason for time to scale with angle at all. A list of vertices and edges is as easy to compare at 160° as at 0°. The cost is paid only if the representation preserves spatial structure — which is what "the thought runs in the channel that built it" means, stated so that it could have come out false.
And this is the honest boundary. What we have measured is a property of the task. It says that this participant, doing this task, rotates at 62 deg/s. It says nothing whatever about which modality that participant prefers when they are not being asked a spatial question — and the between-participant spread in rotation rate, while real, is far smaller than the within-participant swing produced by changing the question. That asymmetry is the entire subject of the next example.
Worked Example 2.3 — Why the trait claim died, and what kind of death it was
Three separate arithmetics, each of which independently sinks the preferred-representational-system claim. They are worth doing separately because they fail it in different ways, and the third one contains the chapter's actual point.
Part A — The attenuation ceiling, derived
Step 1 — Set up classical test theory. Let X be an observed measure of a person's preferred representational system, and Y an observed outcome measure (say, rated rapport). Write each as a true score plus error:
X = T + E
Y = U + F
with the standard assumptions: E and F have mean zero, are uncorrelated with each other, and are uncorrelated with T and U.
Step 2 — Derive what the errors do to the covariance.
Cov(X, Y) = Cov(T + E, U + F)
= Cov(T, U) + Cov(T, F) + Cov(E, U) + Cov(E, F)
= Cov(T, U) + 0 + 0 + 0
= Cov(T, U)
Errors do not touch the covariance.
Step 3 — Derive what the errors do to the variances.
Var(X) = Var(T + E) = Var(T) + Var(E) + 2Cov(T, E) = Var(T) + Var(E)
Reliability is defined as the proportion of observed variance that is true variance:
r_XX = Var(T) / Var(X), hence σ_T = σ_X √r_XX
r_YY = Var(U) / Var(Y), hence σ_U = σ_Y √r_YY
Step 4 — Combine. The observed correlation is
ρ_XY = Cov(X, Y) / (σ_X σ_Y)
= Cov(T, U) / (σ_X σ_Y)
= [ρ_TU σ_T σ_U] / (σ_X σ_Y)
= ρ_TU (σ_X √r_XX)(σ_Y √r_YY) / (σ_X σ_Y)
= ρ_TU √(r_XX · r_YY)
This is the attenuation formula. Errors do not bias the correlation toward or away from anything; they shrink it, by exactly the geometric mean of the two reliabilities. Note the immediate corollary: since ρ_TU ≤ 1, the observed correlation can never exceed √(r_XX · r_YY), no matter how true the underlying theory is.
Step 5 — Put the numbers in. Suppose a preferred-representational-system classification into three categories shows 45% agreement when the same person is classified twice, two weeks apart. Chance agreement for three roughly equiprobable categories is 1/3. Cohen's kappa corrects for that:
κ = (p_o − p_e) / (1 − p_e) = (0.45 − 0.3333) / (1 − 0.3333) = 0.1167 / 0.6667 = 0.175
(These figures are stipulated for the exercise at a level consistent with what the review literature reports about classification instability — Sharpley, Journal of Counseling Psychology, 1984, and Heap's interim verdict of 1988 both turn on this kind of instability. Do not quote 0.175 as a published value.) Treating κ as a reliability index — a rough but standard move for categorical measures — and taking a well-constructed outcome measure at r_YY = 0.85, with a genuinely respectable true effect of ρ_TU = 0.30:
ρ_XY = 0.30 × √(0.175 × 0.85) = 0.30 × √0.14875 = 0.30 × 0.38568 = 0.1157
And the ceiling, if the theory were perfectly true (ρ_TU = 1):
ρ_max = √0.14875 = 0.386
Part B — The power arithmetic
Step 6 — Convert to Fisher's z. For hypothesis tests on a correlation, transform:
z_r = ½ ln[(1 + r)/(1 − r)] = artanh(r)
z_r = artanh(0.1157) = ½ ln(1.1157 / 0.8843) = ½ ln(1.26166) = ½(0.23244) = 0.11622
The sampling standard deviation of z_r is 1/√(n − 3).
Step 7 — Solve for the sample size at 80% power, α = .05 two-tailed. The requirement is
z_r √(n − 3) ≥ z_{α/2} + z_β = 1.9600 + 0.8416 = 2.8016
so
n − 3 ≥ (2.8016 / 0.11622)² = (24.106)² = 581.1
n ≥ 584.1 → n = 585 participants
Step 8 — Compute the power the actual studies had. Typical studies in this literature ran somewhere between twenty and sixty participants. At n = 45:
z_r √(n − 3) = 0.11622 × √42 = 0.11622 × 6.4807 = 0.7532
power = Φ(0.7532 − 1.9600) = Φ(−1.2068) ≈ 0.114
About eleven percent. A study with 11% power that reports a null result has told you almost nothing about the world; it has told you about its own sample size. If you stop here, the honest verdict is "untested," not "refuted" — and that was exactly Sharpley's own framing in his 1987 reply, where he raised the possibility that the theory was untestable rather than false.
Part C — The move that settles it
Step 9 — Notice what Step 5 already contains. The preferred-representational-system claim is not "modality shows up in speech." That claim is true and Example 2.1 measured it. The claim is that each person has one dominant preferred system — a stable individual characteristic. Read that as a measurement statement and it says: the person-level variance in modality preference is large relative to the occasion-level variance.
Now look at κ = 0.175 again. That figure is not a limitation on our ability to test the claim.
That figure is the test of the claim, and the claim failed it.
You do not need a single outcome study. You do not need rapport ratings, or therapy results, or matched-versus-mismatched predicate conditions. If a person classified as visual on Monday classifies as auditory two weeks later at a rate barely above coin-flipping, then whatever is being measured is not a stable property of the person, and the whole downstream apparatus — match their system, pace their channel — has no referent to attach to. The low reliability is not noise obscuring the signal. It is the signal, and it reads: there is no trait here of the size the theory requires.
This inverts how the null literature is usually narrated. The standard story is "dozens of studies looked for a matching effect and didn't find one, so the effect is probably absent, though small samples leave room for doubt." The better story is: the construct was disconfirmed at the measurement stage, before any outcome question was asked, and the outcome studies were downstream of a variable that had already failed to exist. Elich, Thompson and Miller's 1985 finding — that eye movements did not predict self-reported imagery modality — belongs to the same stage. It is not a failure to find an effect. It is a finding that the independent variable does not hold still.
Part D — The eye-movement chart, as a diagnostic test
Step 10 — Set up Bayes. The chart in Frogs into Princes (1979) claims that gaze up and to the speaker's left indicates remembered visual material. Treat this as a diagnostic test. Let V be the event that the speaker is at this moment retrieving a visual memory, and C the event that their gaze goes up-left. Grant the proponent generous operating characteristics:
sensitivity = P(C | V) = 0.70
specificity = P(not C | not V) = 0.70, hence P(C | not V) = 0.30
prior = P(V) = 0.25
Step 11 — Compute the marginal probability of the cue.
P(C) = P(C | V)P(V) + P(C | not V)P(not V)
= (0.70)(0.25) + (0.30)(0.75)
= 0.175 + 0.225 = 0.400
Step 12 — Apply Bayes' theorem.
P(V | C) = P(C | V)P(V) / P(C) = 0.175 / 0.400 = 0.4375
Equivalently, in odds form: prior odds = 0.25/0.75 = 1/3; likelihood ratio LR⁺ = 0.70/0.30 = 2.333; posterior odds = (1/3)(2.333) = 0.7778; posterior probability = 0.7778/1.7778 = 0.4375 ✓
Grant the practitioner every claim they make, and after seeing the cue they are still more likely wrong than right. 43.75% is not a diagnosis.
Step 13 — Quantify the information, which is the sharper way to see it. Shannon entropy of a binary variable with probability p, in bits:
H(p) = −p log₂ p − (1 − p) log₂(1 − p)
Prior uncertainty:
H(0.25) = −0.25 log₂(0.25) − 0.75 log₂(0.75)
= 0.25(2.00000) + 0.75(0.41504)
= 0.50000 + 0.31128 = 0.81128 bits
Posterior after seeing the cue:
H(0.4375) = 0.4375(1.19264) + 0.5625(0.83007)
= 0.52178 + 0.46691 = 0.98869 bits
Note that this is higher than the prior — seeing the cue moved you toward maximum uncertainty. That is why you must average over both possible observations. After not seeing the cue:
P(not C) = 1 − 0.400 = 0.600
P(V | not C) = P(not C | V)P(V) / P(not C) = (0.30)(0.25) / 0.600 = 0.075 / 0.600 = 0.125
H(0.125) = 0.125(3.00000) + 0.875(0.19265) = 0.37500 + 0.16857 = 0.54357 bits
Expected posterior entropy:
E[H] = P(C)·H(0.4375) + P(not C)·H(0.125)
= 0.400(0.98869) + 0.600(0.54357)
= 0.39548 + 0.32614 = 0.72162 bits
Mutual information — the expected reduction in uncertainty:
I = H(prior) − E[H] = 0.81128 − 0.72162 = 0.0897 bits
as a fraction of prior uncertainty: 0.0897 / 0.81128 = 0.111 → 11.1%
Eleven percent of one binary question's worth of uncertainty, at the proponent's own generous numbers. At the accuracy actually measured — Wiseman and colleagues tested the related eye-movement-and-deception claim in PLoS ONE in 2012 and found nothing — LR⁺ = 1, the posterior equals the prior, and I = 0 bits exactly. The chart transmits no information at all.
Step 14 — Explain why it nonetheless looked confirmed. The chart specifies six gaze positions crossed with three modality classes: eighteen cells. If a researcher tests each cell against chance at α = 0.05, and the complete null is true (the chart is entirely empty), the probability of at least one "significant" result is
P(at least one) = 1 − (1 − α)^k = 1 − (0.95)^18
Compute via logarithms:
ln(0.95) = −0.051293
18 × (−0.051293) = −0.923274
e^(−0.923274) = 0.39716
P = 1 − 0.39716 = 0.603
Sixty percent. A study of a chart with nothing in it will, more often than not, produce a publishable-looking confirmation of some cell. The Bonferroni-corrected threshold is α′ = 0.05/18 = 0.00278, which almost none of the early demonstrations would have survived.
The failure mode, named. The edge past which the surviving observation inverts and does harm is precisely here: a practitioner who has correctly learned that predicates carry modality, and who then reads that fact backward into the person, has recreated the dead claim inside the live one. They will hear a client say "I don't see it" and conclude something about the client rather than about the sentence. The discipline that prevents this is one question, asked every time: am I describing this utterance, or this person? The first is data. The second is a hypothesis with a κ of about 0.18 behind it.