Haute Lumière · The reading

Everything measurable improved, and the thing that mattered did not

Four generations of impact measurement, and the one question each generation learned not to ask.

An hour into the work. No instrument in the room can tell this apart from the first minute of it.

An hour into the work. No instrument in the room can tell this apart from the first minute of it.

THE COUNTING IMPULSE

There is nothing suspicious about wanting to count. Long before the first evaluation framework was written down, people counted harvests, tallied trades, and looked hard at whether the effort of a season had been worth the season. The impulse to quantify is an expression of care — of wanting to understand, to improve, to steward what has been entrusted to you. Anyone who treats measurement itself as the enemy has misread the history badly. The trouble starts somewhere else entirely, and it starts late.

It starts at the moment counting becomes good enough to stand in for judgement. Somewhere on the road from the tally stick to the impact dashboard, the field got so skilled at counting that it stopped asking what was worth counting. The question changed shape without anyone deciding it should. Is this working quietly became can this be shown to be working, and those are not the same question. The second one is far easier to answer, which is exactly why it wins.

Consider how this looks from inside an organisation that is doing honest work. The report is accurate. The numbers were collected carefully, checked, and presented without spin. Nobody has lied about anything. And the person who ran the program reads that report and recognises almost nothing of what actually happened in the room — not because the figures are false, but because the figures are about a different thing than the one they were in. That gap is the subject of this book.

A measure that is easy to collect will always out-compete a measure that is true, unless somebody decides otherwise and keeps deciding it.

The history that follows is a developmental story rather than a chronology. Impact assessment moved through four broad waves — counting outputs, measuring outcomes, evidence-based practice, and complexity-aware developmental evaluation — and each wave was a genuine advance over the one before it. Each also carried a shadow: a set of things it could no longer see precisely because of what it had learned to see so well. This is how fields develop. It is also how they get stuck.

None of this is academic. The framework a field inherits does not merely describe impact; it shapes what impact is taken to be. Narrow the frame and the programs that survive funding decisions will be the ones that fit the frame — not the ones that change lives. Every subsequent argument in this book follows from that single mechanism, and understanding where the frame came from is the only way to see the edge of it.

COUNTING OUTPUTS

The earliest formal approaches were, of necessity, simple. Through the middle of the twentieth century, as governments and philanthropic institutions began investing heavily in public health campaigns, educational initiatives and poverty reduction, the operating question was blunt and reasonable: did we do what we said we would do. How many vaccinations administered. How many students enrolled. How many meals served. How many wells dug.

These questions are not trivial and the book does not sneer at them. Outputs matter. If a program promises clean water to a village and installs no wells, no amount of philosophical sophistication rescues it from that failure. A field that lost the ability to ask whether the thing was actually delivered would have traded one blindness for a worse one. Output counting is a floor, and floors are load-bearing.

The limitation is precise: output counting confuses activity with achievement. A literacy program that distributes ten thousand books has produced an impressive output. If those books sit unread — because the reading levels were wrong, or the cultural context was never considered, or the homes have no light to read by after dark — then the output number tells you almost nothing about what the program contributed to anyone's flourishing. The number is true. It is also nearly empty.

Roughly the 1950s through the 1970s, this was the operating paradigm, and it lived inside bureaucratic structures that prioritised accountability over learning. Managers were rewarded for hitting output targets. That incentive did something subtle and expensive: it made asking the deeper question professionally unrewarding. Nobody forbade the question. The structure simply ensured that asking it could only ever cost you.

An output target rewards the person who hits it and never once asks whether hitting it helped anybody.

The shadow of the first wave has not lifted. Organisations under intense funding pressure still default to output metrics, because outputs are cheap to collect, simple to report and easy for a board to understand in ninety seconds. The seduction of the countable is powerful and it is not stupidity — it is a rational response to what funders actually read. Which means the fix is not exhortation. It is changing what gets asked for.

MEASURING OUTCOMES

The shift from outputs to outcomes was a real developmental leap, and it deserves to be described as one. Beginning in the late 1970s and accelerating through the 1990s, evaluators and funders started asking a more penetrating question: what actually changed as a result of this. Not how many students attended the tutoring program, but whether those students learned to read, whether their grades moved, whether they stayed in school. Not how many counselling sessions were delivered, but whether participants reported fewer symptoms, better relationships, a greater sense of agency in their own lives.

Several forces converged to make this possible. Program evaluation matured into a professional discipline with its own methodological standards. Michael Scriven's distinction between formative evaluation, aimed at improvement, and summative evaluation, aimed at judgement, gave the field a vocabulary it had lacked. Economists and policy analysts arrived with cost-effectiveness and cost-benefit analysis, insisting that a program demonstrate not only that it produced outcomes but that it produced them at a defensible price.

The most durable artefact of this era was the logic model — the visual chain running from inputs, the resources invested, through activities, what the program does, to outputs, what is produced, to outcomes, the short and medium-term changes, and finally to impacts, the long-term systemic ones. The logic model became the common language of the field, and it earned that position honestly. It forced program designers to write down their theory of how change actually happens, which many of them had never done.

A logic model is worth its cost the first time somebody notices that the chain does not connect.

The shadow arrived with the same motion. Emphasis on measurable outcomes created powerful pressure toward changes that standardised instruments could capture — survey scores, test results, behavioural indicators counted before and after. Outcomes that resist quantification were pushed to the margin: shifts in how a person makes meaning, the deepening of relational capacity, the slow arrival of collective wisdom in a group that used to argue badly.

The crucial point is that nobody decided these were unimportant. There was no meeting at which the field agreed that meaning-making did not matter. They were marginalised because the instruments of the era could not see them, and what an instrument cannot see does not appear in the report, and what does not appear in the report does not get funded. Invisibility is not a judgement. It behaves exactly like one.

The reading is unfinished; the notes beside it are the second attempt at a definition.

The reading is unfinished; the notes beside it are the second attempt at a definition.

EVIDENCE AND SHADOW

Underneath outcome measurement sat an epistemological assumption that mostly went unexamined: that causation could be cleanly established. If students who attended the program improved their reading scores, the program caused the improvement. That assumption pushed the field toward ever more rigorous designs, and above all toward the randomised controlled trial, borrowed from medical research, which became the gold standard of evidence-based practice.

The RCT is a genuinely powerful instrument and the argument here is not against it. Where a discrete intervention produces discrete measurable effects over a relatively short window, it gives robust evidence of causal impact, and it has settled questions that argument alone never would. The difficulty is the transfer. Human development is not a pharmaceutical trial. The changes that matter most in a life — the gradual integration of a new way of making meaning, the slow repair of a damaged relationship, the quiet arrival of a felt sense of belonging — unfold across years, shaped by countless interacting factors that no randomisation can hold still.

By the early 2000s this had crystallised into a broader paradigm. Fund and implement what has been shown through rigorous research to work; stop spending on what has not been validated. In an era of scarce social funding, evidence-based practice looked like an unassailable commitment to both effectiveness and fiscal seriousness, and in important respects it was. It exposed programs that were ineffective and some that were harmful. It made evaluation central to organisational culture. It gave funders, practitioners and policymakers a shared language. Those contributions are real and the book honours them before it turns.

The hierarchy of evidence is where the shadow falls first. Evidence-based practice established, sometimes explicitly, a ranking with randomised trials at the top and qualitative, narrative and experiential evidence near the bottom. That hierarchy privileged one kind of knowing — detached, objectified, quantitative — and systematically devalued the others: the story a participant tells about how a program changed their sense of self, the shift a facilitator feels in a room's collective energy, the somatic indicators a practitioner reads in a group before anyone speaks.

Then the replication fallacy. The paradigm assumed that a program shown to work in one context would work in another provided it was implemented with fidelity, and that assumption badly underestimates context. A parenting program that flourishes in a middle-class suburb may fail entirely in a community shaped by generational poverty, institutional racism and historical trauma — not because the program is bad, but because the living system it enters is a different system.

Impact is not a property of a program. It is a property of the meeting between a program and the life it enters.

Two further shadows complete the picture. The accountability trap: when evaluation exists chiefly to decide whether an organisation deserves continued funding, it becomes a high-stakes performance rather than an inquiry, and organisations respond rationally by measuring what makes them look good and choosing indicators they already know they can hit. And the mechanistic assumption inherited from the medical model — that social interventions work like treatments, discrete inputs producing predictable linear effects. That holds well enough for simple problems. It holds poorly for poverty, loneliness, meaning-deficit, developmental stagnation and ecological destruction, which are the problems that actually matter.

COMPLEXITY ARRIVES

A counter-narrative emerged in the early 2000s, and it came from practitioners rather than from methodologists. People working in community development, systems change and organisational transformation kept running into the same wall: the existing paradigms were not merely insufficient for their work, they were actively misleading about it. Measuring a shifting intervention against predetermined outcomes in an unstable context produces a number that is precise and wrong.

Michael Quinn Patton's developmental evaluation gave that recognition a name and a method. Where the situation is complex and emergent — where the intervention itself is still evolving, where the context shifts underfoot, where outcomes cannot honestly be specified in advance — evaluation has to change its purpose. It serves learning rather than judgement, supporting ongoing adaptation instead of rendering a verdict. It tracks what actually emerges from the meeting of intervention and context, including what nobody anticipated. It operates in real time, embedded in the work rather than arriving after it. And it treats non-linear causation and distributed effects as normal conditions rather than as measurement failures.

Patton was not alone. Patricia Rogers, Bob Williams and Glenda Eoyang contributed to what became complexity-aware evaluation — a body of practice that takes complexity science, systems thinking and developmental theory seriously rather than decoratively. Alongside it, several complementary methods appeared and proved durable. Outcome Harvesting, developed by Ricardo Wilson-Grau, identifies and then verifies outcomes in situations where they could not have been predicted. Most Significant Change, developed by Rick Davies and Jess Dart, is participatory and privileges the accounts of those most affected. Contribution Analysis, developed by John Mayne, walks a middle path between the impossibility of proving causation in complex settings and the abdication of causal reasoning altogether.

This wave is a real advance and the book says so plainly. It honours the messiness of actual change. It treats impact as an emergent phenomenon to be tracked, interpreted and learned from rather than a fixed quantity waiting to be captured. And it admits what the earlier waves suppressed — that the observer is part of what is observed, and that the act of measuring shapes the reality it claims merely to describe.

And yet the complexity wave stops at a particular threshold. Its methods remain largely cognitive: they analyse, interpret, report, converse. They widen the what of assessment — more kinds of outcome, more kinds of evidence, more honesty about causation — without seriously interrogating the who. The developmental stance of the person assessing, their embodied presence in the room, the structure of consciousness through which they perceive what they then write down: these stay outside the frame.

Every wave widened the lens. Not one of them turned it around.

WHAT GOES MISSING

It is worth naming, as precisely as possible, what four waves of development have systematically left out. Not as a complaint — as an inventory, because each omission is addressable and several are addressable immediately.

Relational quality is the most commonly felt and the least often captured. Programs serving communities, teams and families inevitably alter the quality of relationship among the people in them: trust, mutuality, the capacity to have a productive fight, the collective intelligence available when the group thinks together. Surveys can capture perceptions of relational quality. They cannot capture a relational field in motion. A team may report high satisfaction while carrying unspoken tensions that are quietly eroding everything; a team in the middle of a painful and transformative conflict may report low satisfaction while undergoing a deepening that bears fruit for years. Both readings are accurate. Both are misleading.

Systemic and ecological effects are the next. Most assessment attends to a program's immediate beneficiaries, but programs sit inside larger systems and their effects travel. A leadership development program that changes one executive may, through that executive's changed behaviour, alter the culture of an organisation, the wellbeing of hundreds of employees, and the experience of thousands of customers who will never know the program existed. Those downstream effects are real. They are almost never counted, and the program is funded or defunded as though they were zero.

Then there is the dimension the field is least comfortable naming: the sense of sacred participation in something larger than oneself. Call it meaning, call it purpose, call it the numinous — it is what makes a life feel not merely satisfactory but radiant. Contemplative practice, nature-based work, arts and creativity programs, grief ritual: these touch that dimension, produce effects that participants describe as among the most significant of their lives, and remain almost entirely invisible to conventional instruments.

The unmeasured is not the unimportant. It is only the part of the work no instrument was pointed at.

Each of these omissions has the same structure. Something real happens; the available instrument cannot register it; the report is silent; the silence is read as absence; funding follows the report. No one in that chain acts in bad faith, and the outcome is that the field consistently starves the programs doing the deepest work. The correction is not to abandon measurement but to widen the definition of evidence — and the book's claim is that this can be done with rigour rather than by waving the requirement away.

Nothing here is being measured. Something here is being understood.

Nothing here is being measured. Something here is being understood.

THE DEPTH PROBLEM

Of all the omissions, one is structural rather than instrumental, and it is the hinge of the book's argument. Nearly every assessment framework measures change at the behavioural or attitudinal level: did participants behave differently, did their stated attitudes shift. Both are real signals. Neither distinguishes between two entirely different things that produce identical readings.

Take a person who completes an anger management program and reports fewer angry outbursts. At the behavioural level, change has occurred and the instrument correctly records it. But the instrument cannot tell you whether their relationship to anger has changed at all. Have they developed the capacity to hold anger as an object of awareness, rather than being subject to it? Have they moved from a structure in which anger is an uncontrollable force that arrives and takes over, to one in which anger is a signal to be attended to with curiosity? That shift is invisible to the questionnaire — and it is the shift that determines whether the behavioural change survives the first hard year after the program ends.

Two participants can produce the same numbers for opposite reasons. One has learned what the program wants and is producing it, competently and sincerely, while the underlying structure is untouched. The other has reorganised how they make sense of their own experience, which is slower, harder, and initially less visible. Six weeks out, the first looks like the better outcome. Three years out, the comparison reverses, and by then nobody is measuring.

Compliance reports on time. Structural change reports after the grant has closed.

The consequence for the field is severe and entirely predictable. A framework that cannot distinguish surface compliance from deep structural transformation will systematically overestimate the programs that produce the former and underestimate the ones that cultivate the latter. The bias is not random noise; it points in one direction, every cycle, and it compounds. Over decades it reshapes which programs exist.

Assessing developmental depth is not mysticism and it is not impossible. It requires asking a different kind of question — not only what happened, but how the person accounts for what happened, and what that account reveals about the structure making the meaning. It requires measurement windows long enough for slow change to surface. And it requires assessors trained to hear the difference between a person repeating the program's language and a person who has been reorganised by it.

THE BODY KNOWS

A practitioner walks into a community meeting and feels a constriction in the chest before a word is spoken. That is not a mood. It is information about the emotional field of that room — information no survey administered at the end of the session will capture, and information the practitioner will act on all evening whether or not the methodology permits it.

The same channel runs on the participant side, and it runs earlier than language. People in a healing program notice shifts in somatic experience long before they can say what has changed cognitively: chronic tension easing, appetite returning, dreamlife coming back after years of absence. Ask them at three weeks what has changed and they may have no articulate answer. Their sleep has an answer.

The body is, in a real and defensible sense, the first instrument of impact assessment, and it has been almost entirely absent from the field's methodology. Not debated and rejected — simply never admitted to the room. An evaluation tradition that prizes detachment had no category for the evaluator's own nervous system except as a source of bias to be suppressed.

A practitioner's body registers the state of a room minutes before the methodology permits anyone to say so.

The discipline this requires is the interesting part, because somatic data is not self-validating and the book does not pretend otherwise. A felt sense is a signal to be checked, not a verdict to be reported. It gets corroborated against narrative, against observation, against what participants themselves say — and where it is contradicted, it is set down. Treated that way it is evidence. Treated as unquestionable it is the same authoritarianism the field spent fifty years escaping, wearing better clothes.

What integration looks like in practice is closer to clinical judgement than to either pole of the old argument. A skilled clinician holds lab results, patient history, physical examination and an intuitive impression together, weights them against each other, and reaches a diagnosis that no single input could have produced. Quantitative data, qualitative narrative, somatic knowing, relational sensing and contemplative insight are all legitimate forms of evidence. The capacity worth building is not preference among them but the ability to hold them at once.

WHOSE FLOURISHING

What counts as impact is not a neutral category and never has been. It is shaped by cultural values, by worldview, and by who holds the power to define. So a framework worth using asks three questions before it asks anything else: whose definition of flourishing is in force here, whose voices determined what success looks like, and who benefits from the current metrics while somebody else is rendered invisible by them.

The dominant frameworks were developed primarily in Western, industrialised institutions, for Western, industrialised purposes. Exported into other cultural settings without adaptation, they impose alien definitions of wellbeing and then score people against them — and the score will be low, and the low score will be read as a deficit in the community rather than a mismatch in the instrument. Cultural humility is not a courtesy in assessment. It is a condition of the finding being true.

There is a second and more common failure, and it has nothing to do with culture. Assessment always carries the potential to become surveillance rather than learning. When data exists mainly to reward or punish, to rank or sort, to justify or defund, it stops serving the people it describes and becomes an instrument of institutional power aimed at them. The test is simple and worth applying honestly: does the assessment serve the communities being assessed, or only the funder who commissioned it.

An assessment that only the funder can read is not an assessment. It is an audit with better manners.

The certainty trap catches the careful as easily as the careless. It is tempting to present findings as definitive — proof that a program works or does not — because that is what the request asked for and hedging reads as weakness. But in complex human systems, certainty is almost always an overstatement. Epistemic humility means presenting findings as the best current understanding, naming the limits of the method used, and staying genuinely open to being surprised by what comes next.

And the opposite temptation needs naming with equal force, because it is common in contemplative and spiritual communities and it is corrosive. The romance of the unmeasurable holds that what matters most is inherently beyond assessment, and there is a kernel of truth in it — the deepest experiences do resist full quantification. It also makes a convenient shelter from accountability. Most dimensions of human flourishing can be assessed, even where they cannot be reduced to a single number, and the work is to build practices adequate to the complexity rather than to declare the complexity off-limits.

One further caution, and it is practical. Programs that work at depth surface difficult psychological material, and an assessment process that goes looking for depth will sometimes find more than it can hold. Assessment is not therapy. Practitioners need to know the edge of their competence and know when to refer a participant to a qualified mental health professional, and the process itself must never become the thing that caused harm.

THE ASSESSOR

The most radical omission is the last one, and it is the one that reorganises everything before it. The field of impact assessment has almost never turned its gaze on the consciousness of the assessor. It has scrutinised instruments, sampling, design, analysis and reporting with enormous care, and left unexamined the one component present in every step of the process.

Consider what the assessor brings into the room: a developmental stage, a set of cultural assumptions, an emotional state on the day, a degree of somatic awareness or its absence. These are not contaminants to be controlled for. They determine what is perceivable. An evaluator operating from a conventional, achievement-oriented structure of meaning will naturally privilege metrics that reflect achievement — efficiency, scale, cost per participant — and will do so sincerely, believing the choice to be neutral. An evaluator operating from a more complex structure may perceive dimensions of impact that the first one cannot see at all. Not overlooked. Not deprioritised. Literally not available.

The instrument of assessment is not the survey. It is the human being who designed it, ran it, and decided what it meant.

Read the four waves again with this in view and they stop being a chronology of techniques. Output counting corresponds to the capacity to track tangible, visible, countable things. Outcome measurement corresponds to the capacity to reason about causation and construct models of what is not directly observable. Evidence-based practice represents the commitment to rigorous methodology, standardised procedure and replicable result. Developmental evaluation begins the move into post-conventional ground, where reality is understood to exceed any single framework, context weighs as heavily as content, and the observer is admitted to be inside the picture. The field has been developing in the same shape its subjects do.

The step the book argues for next is what it calls integral assessment consciousness: a way of engaging with impact that holds multiple perspectives at once, admits the body alongside the mind, includes the sacred alongside the secular, and accepts that the deepest forms of human flourishing exceed every instrument so far devised. It practices transcend-and-include — honouring what each earlier stage contributed while seeing clearly what it could not yet see. Nothing in the earlier waves is discarded. The tools still work; they are simply no longer mistaken for the whole of what can be known.

This is not an argument against measurement. It is an argument for maturing in relation to it — holding metrics the way a poet holds language, knowing the words are never quite adequate to the experience, using them as skilfully as possible anyway, and letting them point beyond themselves toward the living thing they are trying to describe.

Which suggests where the work actually starts, and it is not with a framework. It starts with your own history of being measured. Write the assessment autobiography: begin at the earliest memory of being evaluated — a school test, a performance review, a medical examination — and follow the thread forward. Mark the moments when assessment was a gift, when it showed you something true about yourself that you could not have seen alone. Mark the moments when it was a violation, when it reduced you to a number or a category that missed the whole of you. That exercise is not therapy. It is preparation, because the biases and the wounds and the gifts you carry about measurement will be present in every assessment you ever design, whether or not you have looked at them.


Free to read

Free to read, and free to hear. Every chapter of every book in this house, and every narration of it, is open to anybody. No account, no card, nothing to cancel.

Progressive Impact Assessment™ — 1 chapter, 7,101 words, read aloud in full.

Buying a volume is now for keeping it — the EPUB, the PDF and the press file, yours on disk. The reading is free either way.

Read it free Keep the files — $44.44

What is in it


A field that measures only what it can already see will keep discovering that it was right.
What cannot be counted can still be assessed. The two words have never been synonyms.
Every framework decides in advance which kinds of change are allowed to exist.
Ask who is holding the instrument before you ask how precise the instrument is.

Keep looking

Every phrase on this page opens into the house search. The shelf holds The Developmental Canon and six other shelves, and the reading is free.