The Worked Example — Taking a Baseline on Meta-Model Deletion
Here is the whole thing, done in front of you, with the reasoning left in.
The pattern under test. Simple deletion in the Meta Model: a sentence where a required argument of the verb has gone missing. "I'm disappointed." Disappointed by whom, about what? The word "disappointed" is a two-place predicate wearing a one-place coat. The recovery move is a question that restores the missing slot: "Disappointed about what specifically?"
That is the knowledge. You can read it once and hold it. Holding it is not the skill.
Step 1 — Build the stimulus set before you know what's in it.
I want twelve sentences, some containing simple deletions and some clean, in an order I cannot anticipate. If I write them and then respond to them, I am testing recall of my own list, not detection under load. So I take twelve index cards, write one sentence per card — six with deletions, six without — shuffle face-down, and don't look.
Why this matters more than it seems: anticipation is the single largest contaminant in self-testing. If you know a deletion is coming, your response time collapses by a factor that has nothing to do with competence and everything to do with priming. The shuffle is not fussiness. It is the difference between a measurement and a flattering story.
Step 2 — Define the response before the clock starts.
The output is a spoken question, out loud, at conversational volume, that recovers the missing material. Not "I would ask about the object." Not a written note. Spoken. The mouth is part of the circuit and it is slower than the mind by a margin that will surprise you.
I also decide in advance what counts as a hit: the question must restore a missing argument and must not add content. "Disappointed about what specifically?" is a hit. "Are you disappointed about work?" is a miss — I supplied the content, which is a different pattern entirely, and a worse one.
Deciding the criterion first is not bureaucracy. If you decide afterward, you will decide in favour of the answer you gave. Everyone does. The criterion set in advance is the only thing standing between you and a baseline that means nothing.
Step 3 — Instrument the timing.
Phone recording, audio only. I flip a card, read the sentence aloud, then respond aloud as fast as a clean response will come. The recording captures both. Afterward I open the audio in any editor with a waveform — the free ones are fine — and measure from the end of my reading to the start of my response.
I read the sentence aloud rather than silently for a reason worth stating. Reading aloud puts the stimulus into the auditory channel, which is where it lives in real conversation. Silent reading is a visual-linguistic task and recruits a partly different route. The measurement would be cleaner and would tell me less.
Step 4 — Run it. Here is the actual first pass.
Card 1: "The decision was made." (deletion — and a nominalisation, and a passive; I notice this and file it) — gap 2.1 s — response: "Made by whom?"
Card 2: "I drove to Manchester on Tuesday." (clean) — gap 1.4 s — response: none needed. But note the 1.4 s. I spent it deciding there was nothing to do.
Card 3: "She's better." — gap 3.3 s — response: "Better than what?" And a false start before it: an audible "um."
Card 4: "He never listens." (universal quantifier — not the target pattern) — gap 0.9 s — response: "Never?" Fast, because this one is over-practised from years of ordinary argument.
Card 5: "I'm frightened." — gap 2.7 s — "Frightened of what?"
Card 6: "We finished the report at four." (clean) — gap 1.1 s.
Card 7: "It's too expensive." — gap 4.0 s — "Too expensive compared to what?" The four seconds is the interesting one and I will come back to it.
Cards 8–12 run between 1.0 and 3.1 s, with two misses: on "They're not happy about it" I said "Who's not happy?" — which recovers the referential index, not the deletion I was hunting. Correct-adjacent. Still a miss by my own pre-set criterion.
Step 5 — Read the numbers honestly.
Deletion cards: 2.1, 3.3, 2.7, 4.0, 2.4, 3.1. Mean 2.9 s. Two misses out of six on criterion, so accuracy 67%, and the two hits I am proudest of were the slowest.
Conversation gives me roughly 400 ms. I am running at seven times the budget with a third of my responses wrong.
Step 6 — Interpret, which means finding the mechanism, not the excuse.
The easy read is "I need more reps." True but useless — it names the remedy without naming the disease.
Look at card 7, the four-second one. "It's too expensive." I know that pattern cold; I could teach it. What took four seconds was not retrieval of the pattern. It was deciding which pattern applied. "Too expensive" carries a comparative deletion, a lost performative, and an unspecified referential index in five syllables. Three doors, and I stood in the corridor.
That is the mechanism, and it is not what I expected to find. The latency is not in the knowing. It is in the selection. Declarative knowledge stored as a list — here are the twelve Meta Model violations — forces a serial scan at the moment of use, and a serial scan over twelve items cannot complete in 400 ms no matter how well you know each item. The list is the problem. The list is the thing you were taught, and it is stored in the system that cannot deliver.
Card 4 tells the same story from the other side. "He never listens" came back in 900 ms, my fastest, and it is not the pattern I was drilling. It was fast because I have said "never?" to people in real arguments a few hundred times. It never went through the list. It was never on the list. It was in my mouth.
So the honest conclusion from this baseline is not practise harder. It is: the reps must build direct stimulus-to-response bindings, one pattern at a time to fluency, not a taxonomy searched at the point of need. One door, opened a thousand times, until the door is not a decision.
Step 7 — Write the number down where it will embarrass you later.
Mean gap 2.9 s. Accuracy 67%. Date it. In eight weeks that piece of paper is either evidence or an indictment, and both are useful.
Notice what did not happen anywhere in those seven steps: I did not feel good about my knowledge of the Meta Model. I know it well. The measurement did not care.