The Second Wave: Measuring Outcomes
The Second Wave: Measuring Outcomes
The shift from outputs to outcomes represented a genuine developmental leap. Beginning in the late 1970s and accelerating through the 1990s, evaluators and funders began asking a more penetrating question: What actually changed as a result of our intervention?
This was a crucial evolution. Instead of merely counting the number of students who attended a tutoring program, outcome measurement asked: Did those students actually learn to read? Did their grades improve? Did they stay in school? Instead of tracking the number of counseling sessions delivered, it asked: Did participants report reduced symptoms of depression? Did their relationships improve? Did they experience a greater sense of agency in their own lives?
The outcome movement was fueled by several converging forces. The rise of program evaluation as a professional discipline brought methodological rigor to the assessment of social interventions. The work of pioneers like Michael Scriven, who distinguished between formative evaluation (aimed at improvement) and summative evaluation (aimed at judgment), gave the field a more nuanced vocabulary. Meanwhile, the growing influence of economists and policy analysts introduced cost-effectiveness analysis and cost-benefit analysis, demanding that social programs demonstrate not only that they produced outcomes, but that they produced outcomes efficiently.
Perhaps the most influential development of this era was the emergence of logic models — visual frameworks that map the causal chain from inputs (resources invested) through activities (what the program does) to outputs (what is produced), outcomes (short and medium-term changes), and ultimately impacts (long-term, systemic changes). The logic model became the lingua franca of program evaluation, and for good reason: it forced program designers to articulate their theory of how change happens.
Yet the outcome era introduced its own shadows. The emphasis on measurable outcomes created a powerful incentive to focus on changes that could be captured by standardized instruments — survey scores, test results, behavioral indicators. Outcomes that resisted quantification — shifts in meaning-making, deepening of relational capacity, the slow emergence of collective wisdom — were systematically marginalized. Not because anyone believed they were unimportant, but because the measurement tools of the era could not see them.
There was also a subtle but significant epistemological assumption embedded in outcome measurement: the belief that causation could be cleanly established. If students who attended our program improved their reading scores, the program caused the improvement. This assumption drove the field toward increasingly rigorous research designs — most notably, the randomized controlled trial (RCT), borrowed from medical research, which became the gold standard of evidence-based practice.
The RCT is a powerful tool. In contexts where a discrete intervention produces discrete, measurable effects in a relatively short timeframe, it can provide robust evidence of causal impact. But human development is not a pharmaceutical trial. The changes that matter most in a person's life — the gradual integration of a new way of making meaning, the slow repair of a damaged relationship, the quiet emergence of a felt sense of belonging — unfold over years, not weeks. They are shaped by countless interacting factors that no randomized design can isolate. And they are often invisible to the instruments that the outcome movement deployed.
The Third Wave: Evidence-Based Practice and Its Shadow
By the early 2000s, the outcome movement had crystallized into a broader paradigm: evidence-based practice (EBP). The proposition was compelling. We should fund and implement programs that have been shown, through rigorous research, to produce positive outcomes. We should stop wasting resources on programs that have not been validated. In an era of scarce social funding and growing demands for accountability, evidence-based practice seemed like an unassailable commitment to both effectiveness and fiscal responsibility.
And in many respects, it was. The evidence-based movement helped to expose programs that were ineffective or even harmful. It elevated the importance of evaluation in organizational culture. It created a shared language between funders, practitioners, and policymakers. These contributions are real and should be honored.
But the shadow of evidence-based practice is significant, and it is a shadow that the Luminous approach takes seriously.
The hierarchy of evidence problem. Evidence-based practice established an implicit (and sometimes explicit) hierarchy of evidence, with randomized controlled trials at the top and qualitative, narrative, and experiential evidence near the bottom. This hierarchy privileged a particular kind of knowing — detached, objectified, quantitative — and systematically devalued other ways of understanding impact: the story a participant tells about how a program changed their sense of self, the felt shift in a community's collective energy, the somatic indicators that a practitioner notices in a group.
The replication fallacy. The evidence-based paradigm assumed that if a program worked in one context, it would work in another — provided the program was implemented with fidelity. This assumption underestimates the profound influence of context. A parenting program that works beautifully in a middle-class suburb may fail utterly in a community shaped by generational poverty, institutional racism, and historical trauma. Not because the program is bad, but because the living system into which it is introduced is fundamentally different. Impact is not a property of a program alone; it is an emergent property of the relationship between a program and the system it enters.
The accountability trap. Evidence-based practice, in its most rigid forms, can create a culture of assessment-as-surveillance rather than assessment-as-learning. When organizations are evaluated primarily to determine whether they deserve continued funding, evaluation becomes a high-stakes performance rather than a genuine inquiry into what is happening and why. In such environments, organizations have powerful incentives to measure only what makes them look good, to select indicators they know they can hit, and to avoid asking the deeper questions that might reveal uncomfortable truths.
The mechanistic assumption. Perhaps most fundamentally, evidence-based practice inherited a mechanistic worldview from the medical model: the assumption that social interventions work like treatments — discrete inputs that produce predictable, linear effects. This assumption works reasonably well for simple problems (if children lack vaccinations, vaccinate them). But most of the challenges that matter most to human flourishing — poverty, loneliness, meaning-deficit, developmental stagnation, ecological destruction — are complex, adaptive, and irreducible to simple cause-and-effect chains.
The evidence-based movement, for all its contributions, did not ask the question that the Luminous approach places at the center: What kind of consciousness is doing the measuring, and how does that consciousness shape what can be seen?
The Fourth Wave: Developmental Evaluation and Complexity
In the early 2000s, a powerful counter-narrative began to emerge. Scholars and practitioners working in complex social systems — community development, systems change, organizational transformation — recognized that the existing evaluation paradigms were inadequate for the challenges they faced.
The most influential articulation of this counter-narrative came from Michael Quinn Patton, whose concept of developmental evaluation represented a paradigm shift. Patton argued that in complex, emergent situations — where the intervention itself is evolving, where the context is shifting, and where the outcomes cannot be predetermined — traditional evaluation is not merely insufficient but actively misleading. What is needed instead is an evaluative approach that:
- Serves learning rather than judgment. The primary purpose of developmental evaluation is not to render a verdict on a program's effectiveness, but to support ongoing adaptation and learning.
- Embraces emergence. Rather than measuring against predetermined outcomes, developmental evaluation tracks what actually emerges from the interaction between an intervention and its context — including outcomes that no one anticipated.
- Operates in real time. Rather than conducting evaluations after the fact, developmental evaluation embeds evaluative thinking within the ongoing process of innovation and adaptation.
- Honors complexity. Developmental evaluation recognizes that in complex adaptive systems, causation is non-linear, effects are distributed, and the relationship between action and outcome is often unpredictable.
Patton's work, along with contributions from scholars such as Patricia Rogers, Bob Williams, and Glenda Eoyang, opened the door to what we might call complexity-aware evaluation — approaches that take seriously the insights of complexity science, systems thinking, and developmental theory.
Around the same time, several complementary frameworks emerged. Outcome Harvesting, developed by Ricardo Wilson-Grau, offered a methodology for identifying and verifying outcomes in situations where they cannot be predicted in advance. Most Significant Change (MSC), developed by Rick Davies and Jess Dart, provided a participatory approach to evaluation that privileges the stories of those most affected by a program. Contribution Analysis, developed by John Mayne, offered a middle path between the impossibility of proving causation in complex settings and the abdication of any causal reasoning at all.
These approaches represent a genuine developmental advance. They honor the messiness of real-world change. They recognize that impact is not a fixed quantity to be measured but an ongoing, emergent phenomenon to be tracked, interpreted, and learned from. And they acknowledge what the earlier waves suppressed: that the observer is always part of what is observed, and that the act of measurement shapes the reality it claims to describe.
Yet even the complexity-aware wave has its limitations. Most of these approaches remain primarily cognitive — they analyze, interpret, and report, but they do not systematically engage the somatic, relational, and spiritual dimensions of impact. They expand the what of assessment (more kinds of outcomes, more kinds of evidence) without fundamentally questioning the who — the consciousness, the developmental stance, the embodied presence of the assessor.
What Gets Lost When We Only Measure What's Easy to Count
Let us pause here and name, with as much precision as we can, what has been systematically excluded from the impact assessment story as it has unfolded.
Developmental depth. The vast majority of impact assessment frameworks measure change at the behavioral or attitudinal level: Did participants change their behavior? Did their attitudes shift? These are surface-level indicators of what may or may not represent a deeper transformation. A person who completes an anger management program and reports fewer angry outbursts has changed at the behavioral level. But has their relationship to anger changed? Have they developed the capacity to hold anger as an object of awareness rather than being subject to it? Have they moved from a meaning-making structure in which anger is an uncontrollable force to one in which anger is a signal to be attended to with curiosity? This deeper shift — the shift in how someone makes meaning — is invisible to conventional assessment tools, yet it is precisely this shift that determines whether the behavioral change will endure.
Relational quality. Programs that serve communities, teams, and families inevitably affect the quality of relationships among participants. Yet relational quality — the degree of trust, mutuality, generative conflict capacity, and collective intelligence present in a relational field — is extraordinarily difficult to measure with conventional instruments. Surveys can capture perceptions of relational quality, but they cannot capture the living, felt reality of a relational field in motion. A team may report high satisfaction on a survey while harboring unspoken tensions that are eroding their collective capacity. Conversely, a team in the midst of a painful but transformative conflict may report low satisfaction while undergoing a relational deepening that will bear fruit for years.
Somatic knowing. The body carries information that the mind cannot access. A practitioner who enters a community meeting and feels a constriction in the chest is receiving data about the emotional field of that community — data that no survey can capture. Participants in a healing program may notice shifts in their somatic experience — the easing of chronic tension, the restoration of appetite, the return of dreamlife — long before they can articulate what has changed cognitively. The body is, in a very real sense, the first instrument of impact assessment. Yet it has been almost entirely absent from the field's methodology.
Ecological and systemic effects. Most impact assessment focuses on the immediate beneficiaries of a program. But programs exist within larger systems, and their effects ripple outward in ways that are often invisible to conventional frameworks. A leadership development program that transforms a single executive may, through that executive's changed behavior, affect the culture of an entire organization, the wellbeing of hundreds of employees, and the quality of service experienced by thousands of customers. These downstream effects are real, but they are almost never captured.
Spiritual depth and sacred participation. There is a dimension of human flourishing that transcends the psychological, the relational, and the systemic — a dimension that has to do with the sense of sacred participation in something larger than oneself. Call it meaning, call it purpose, call it the numinous — it is the quality that makes life feel not merely satisfactory but radiant. Programs that touch this dimension — contemplative practices, nature-based interventions, arts and creativity programs, grief rituals — produce effects that are real and profound but almost entirely unmeasurable by conventional means.
The assessor's own development. This is perhaps the most radical omission of all. The field of impact assessment has almost never turned its gaze upon the consciousness of the assessor. Yet the developmental stage, cultural assumptions, emotional state, and somatic awareness of the person doing the assessing profoundly shape what they can perceive. An evaluator operating from a conventional, achievement-oriented meaning-making structure will naturally privilege metrics that reflect achievement — efficiency, scale, cost-effectiveness. An evaluator operating from a more complex, integral meaning-making structure may perceive dimensions of impact that the first evaluator literally cannot see. The instrument of assessment is not the survey or the interview protocol — it is the human being who designed, administered, and interpreted it.
The Luminous Critique: Why a New Framework Is Needed
The Luminous approach to impact assessment does not reject the contributions of the waves that preceded it. Each wave — output counting, outcome measurement, evidence-based practice, developmental evaluation — represents a genuine advance, and each continues to offer tools and insights that have value. The Luminous approach practices what we call transcend-and-include: honoring what each previous stage contributed while recognizing what it could not yet see.
But the Luminous approach insists that the field of impact assessment is ripe for a further developmental leap — one that is demanded not by academic fashion but by the actual complexity of the challenges we face. Climate change, mental health crises, the erosion of social cohesion, the search for meaning in a disenchanted world — these are challenges that cannot be adequately addressed, let alone assessed, by frameworks that reduce human flourishing to behavioral outcomes measured by standardized instruments.
What is needed is an approach that:
Honors multiple forms of evidence. Quantitative data, qualitative narrative, somatic knowing, relational sensing, and contemplative insight are all legitimate forms of evidence. A progressive impact framework does not privilege one over the others but cultivates the capacity to integrate them — much as a skilled clinician integrates lab results, patient history, physical examination, and intuitive impression into a comprehensive diagnosis.
Includes developmental depth. Measuring what people do differently is important. Measuring how people make meaning differently is essential. A framework that cannot distinguish between surface-level behavioral compliance and deep structural transformation will consistently overestimate the impact of programs that produce the former and underestimate programs that cultivate the latter.
Assesses systemically. The unit of analysis for progressive impact assessment is not the individual alone. It is the nested system of individuals, relationships, organizations, communities, and ecosystems within which any intervention operates. Impact ripples through these nested systems in non-linear ways, and a progressive framework must be equipped to track those ripples.
Engages the body as an instrument of assessment. Somatic indicators — the felt sense of a group's energy, the physical markers of safety or threat in a room, the bodily signatures of transformation — are data. They are not sufficient by themselves, but they are essential. A framework that ignores the body's testimony is operating with a fraction of the available evidence.
Practices cultural humility. What counts as "impact" is not a neutral, universal category. It is deeply shaped by cultural values, worldviews, and power structures. A progressive impact framework asks: Whose definition of flourishing is being used? Whose voices are centered in determining what success looks like? Who benefits from the current metrics, and who is rendered invisible?
Turns the gaze upon the assessor. The Luminous approach takes seriously the recognition that the observer shapes what is observed. This means that progressive impact assessment includes practices for cultivating the assessor's own developmental capacity, somatic awareness, cultural humility, and ethical discernment. The most sophisticated assessment framework in the world, wielded by an assessor who lacks these capacities, will produce results that are technically precise and substantively hollow.
A Developmental Story, Not Just a Historical One
We have traced the evolution of impact assessment through four waves, and we have named both the contributions and the shadows of each. But this is not merely a story about the past. It is a story about us — about the developmental trajectory of a field that mirrors, in many ways, the developmental trajectory of human consciousness itself.
The output era corresponds to what developmental theorists might call the concrete operational stage — the capacity to track tangible, visible, countable things. The outcome era corresponds to the formal operational stage — the capacity to think about causation, to construct logical models, to reason about things that are not directly observable. The evidence-based era represents the systematic stage — the commitment to rigorous methodology, standardized procedures, and replicable results. And the developmental evaluation era begins the movement into post-conventional consciousness — the recognition that reality is more complex than any single framework can capture, that context matters as much as content, and that the observer is always embedded in what is observed.
The Luminous approach invites a further step: the movement into what we might call integral assessment consciousness — a way of engaging with impact that holds multiple perspectives simultaneously, honors the body as well as the mind, includes the sacred as well as the secular, and recognizes that the deepest forms of human flourishing cannot be fully captured by any instrument yet devised.
This does not mean we abandon measurement. It means we mature in our relationship to it. We learn to hold our metrics with the same reverence and humility with which a poet holds language — knowing that the words are never quite adequate to the experience, but using them as skillfully as we can, always pointing beyond themselves toward the living reality they attempt to describe.
Common Pitfalls and Ethical Cautions
Before we move forward, several cautions must be named:
The romance of the unmeasurable. There is a temptation, particularly in contemplative and spiritual communities, to dismiss all measurement as reductive and to claim that the things that matter most are inherently beyond assessment. While there is a kernel of truth in this — the deepest experiences of human life do resist full quantification — it can also become a convenient excuse for avoiding accountability. The Luminous approach holds that most dimensions of human flourishing can be assessed, even if they cannot be reduced to a single number. The challenge is to develop assessment practices that are adequate to the complexity of what we are trying to understand.
Measurement as control. Assessment always carries the potential to become an instrument of surveillance and control rather than learning and liberation. When assessment data is used primarily to reward or punish, to rank or sort, to justify or defund, it ceases to serve the purposes of human flourishing and becomes a tool of institutional power. The Luminous approach insists that assessment must always be in service to the people and communities being assessed, not merely to the funders or institutions that commissioned the assessment.
The certainty trap. It is tempting to present assessment findings as definitive — as proof that a program works or doesn't work. But in complex human systems, certainty is almost always an overstatement. The Luminous approach practices what we call epistemic humility: presenting findings as our best current understanding, acknowledging the limitations of our methods, and remaining open to being surprised.
Cultural imperialism in assessment. The dominant frameworks for impact assessment were developed primarily in Western, industrialized contexts. When exported to other cultural settings without adaptation, they can impose alien definitions of success, wellbeing, and flourishing. A progressive impact framework must always ask: Is this framework serving the people it claims to assess, or is it serving the cultural assumptions of those who designed it?
Mental health humility. Programs that address deep developmental, relational, or spiritual dimensions of human experience may surface difficult psychological material. Assessment practitioners must be aware of their own limitations and know when to refer participants to qualified mental health professionals. Assessment is not therapy, and the assessment process should never inadvertently cause harm.
✨ Luminous Invitations
As you read this chapter, you are invited to pause and notice:
- What is your own relationship to measurement? Do you tend toward the pole of wanting everything quantified and proven, or toward the pole of resisting measurement as reductive? Neither pole is wrong, but noticing where you stand is the beginning of a more integrated relationship to assessment.
- What has been your experience of being assessed? Think of a time when you were evaluated — at school, at work, in a program. Did the assessment capture what was most important about your experience? What did it miss?
- Where in your body do you feel the tension between rigor and mystery? This is not a rhetorical question. The body carries real information about our relationship to knowing and not-knowing. Notice what arises.
Reflection Questions
- Think about a program or intervention you have been involved in — as a designer, implementer, participant, or evaluator. What dimensions of impact were measured? What dimensions were ignored? What might have been different if the assessment had been more comprehensive?
- Consider the four waves of impact assessment described in this chapter. Which wave most closely reflects the assessment culture of your organization or community? What would it look like to take the next developmental step?
- The chapter argues that the consciousness of the assessor shapes what can be seen. How do you experience this in your own practice? What aspects of impact are you naturally attuned to? What aspects might you be blind to?
Practical Exercise: Your Assessment Autobiography
Take thirty minutes to write a brief "assessment autobiography." Begin with your earliest memory of being measured or evaluated (a school test, a performance review, a medical examination). Trace the thread of assessment through your life. Notice the moments when assessment felt like a gift — when it helped you see something true about yourself. Notice the moments when assessment felt like a violation — when it reduced you to a number or category that missed your essential humanity.
This exercise is not therapy. It is preparation. To practice progressive impact assessment, we must first understand our own history with measurement — the biases, wounds, and gifts that we bring to the work.
In the next chapter, we will move from history to principles, articulating the five foundational commitments of the Luminous Impact Framework and exploring what it means to build an assessment practice that truly honors the wholeness of human flourishing.