← Back to News
opinion

September 30, 2026

From evidence to impact

Dr Lígia Teixeira

‍This is an adapted version of a speech by our CEO Ligia Teixeira at the International Evaluation Conference in Darwin, Australia.

In September, I had the honour of speaking at the International Evaluation Conference in Darwin. It was a particular pleasure to be in a room with people who, whether they work directly as evaluators or in one of the many roles around evaluation, share a real commitment to learning and improvement. Because, at their best, evaluators do something profoundly important: they help people, communities and societies get better at getting better.

That is one of the reasons I was so pleased to be asked to speak. Much of my work has sat at the intersection of evidence, policy and implementation: building evidence, shaping how it informs decisions and investment, and strengthening the systems through which learning becomes action. The question I keep returning to is not simply whether we can produce rigorous evidence, but whether we can make it useful enough to change what happens — and whether those changes ultimately add up to better outcomes.

In that sense, evaluation sits inside a much bigger ambition. It is part of how we try to do good better: to understand whether what we are doing is helping, to notice when it is not, to learn faster and to change course when what we learn requires it. The conference theme, Making space, valuing place, felt particularly fitting. Making space means being willing to notice perspectives, questions and forms of knowledge that our institutions may not naturally privilege. Valuing place means recognising that context is not simply noise around an intervention, something inconvenient to be controlled away in pursuit of a cleaner answer. Place shapes what is possible. Institutions, relationships, histories, incentives and lived experience all affect what happens when an intervention meets the world.

Over time, I have become increasingly interested in the distance between knowing something and actually changing an outcome. We talk a great deal about the knowledge–action gap: good evidence exists, but too often does not find its way into policy, programmes or practice. That gap is real. But there is another one that troubles me just as much. People can know what the evidence says, act on it, commission the intervention and implement it competently, and the outcome we ultimately care about can still fail to improve. I think of this as the action–impact gap. Evidence is a means to an end; the end is impact.

Before getting into what it takes to close those gaps, it is worth remembering just how consequential our ability to evaluate, compare and learn has been. Some of the most important advances in medicine, public health and social policy did not begin with an entirely new idea. They began when somebody found a better way of asking whether what we were already doing was actually helping. Evaluation, in that sense, has been one of the engines of progress. It has allowed us to challenge practices that felt intuitively right but caused harm; to distinguish genuine effects from coincidence, conviction or selective memory; to discover that outcomes once treated as inevitable could in fact be changed; and, over time, to build knowledge cumulatively rather than beginning again with every new programme or generation.

That history matters because it reminds us what is at stake. The capacity to compare, learn and change course is not an academic luxury. It is one of the ways societies become capable of improving themselves. Medicine gives us some of the clearest examples because the consequences of getting the answer wrong were often stark.

For centuries, bloodletting was accepted medical practice. By the early nineteenth century, it was deeply embedded in medicine. Looking back, it is tempting to treat this simply as evidence of how unenlightened medicine once was. But that lets us off too easily. The physicians who practised bloodletting were not generally stupid, careless or indifferent to suffering. They were operating within an internally coherent understanding of illness, and what they saw in front of them could appear to validate it. A patient was bled and recovered: the treatment seemed to have helped. Another was bled and died: perhaps the disease had simply been too severe. Everyday experience, sincerely interpreted, had a remarkable ability to reinforce what practitioners already believed.

The French physician Pierre-Charles-Alexandre Louis introduced a different kind of challenge. Rather than relying predominantly on individual cases and accumulated clinical judgement, he systematically compared groups of patients. His methods would look rudimentary beside a modern clinical trial, but the shift was profound. His work on pneumonia challenged confidence in bloodletting, and his “numerical method” - replacing anecdote with counting and comparison of outcomes - became an important precursor to quantitative medicine and clinical epidemiology. The important move was not simply replacing one medical opinion with another. Louis created a comparison capable of contradicting what everybody thought they already knew. Instead of asking, “Does this make sense to us?” or “Have I seen patients improve?”, the question became: when we compare what actually happens, do people do better?

That is why one of the most important questions in evaluation is also one of the simplest: what am I missing? Progress depends on asking that question honestly enough to learn from the answer and, when necessary, change course. It requires a difficult combination: enough conviction to act, but enough humility to remain genuinely interested in the possibility that our explanation is incomplete.

The history of oral rehydration therapy offers another version of the same lesson. Severe cholera can strip the body of fluid with terrifying speed. For a long time, intravenous fluid was the mainstay of treatment, and there was deep scepticism that a gut so affected by diarrhoeal disease could absorb enough fluid orally to matter. Then physiological research established something crucial: sodium and glucose are transported together across the intestinal wall. Research in Dhaka and Calcutta subsequently showed that this mechanism remained intact in people with cholera and that oral glucose-salt solutions could maintain hydration while dramatically reducing the amount of intravenous fluid required. During the 1971 Bangladesh liberation war, oral rehydration was used under extraordinarily difficult field conditions among refugees affected by cholera, helping demonstrate that a treatment grounded in sophisticated physiology could also work where intravenous therapy was simply not practical.

What became oral rehydration therapy is almost disarmingly simple: water, salts and glucose in the right proportions. But behind that simplicity was a profound shift in understanding. The question was no longer simply how to replace fluid more efficiently. It was: what if the intestine can still absorb water under the right conditions? Sometimes progress comes from becoming better, faster or more consistent at what we already do. Sometimes it comes because somebody notices that an assumption sitting underneath the whole model is wrong. Evaluation, at its best, helps create the conditions for both kinds of learning.

The history of evaluation is also a history of expanding ambition. Once we became better at asking whether a treatment or programme worked, new questions followed almost inevitably. If it worked, for whom did it work? Why did it work? Under what conditions? Would it work somewhere else? What happened to people who were not reached by it? What perspectives were missing from the original account? And then an even bigger question emerged: suppose we identify an intervention that works and implement it well. Does that actually change the outcome at the level we care about — across a community, a population or a whole system?

So the questions have not simply become more sophisticated; they have become wider. We still absolutely need to ask: did it work? What changed? What caused the change? But increasingly we also ask: for whom? Why? Under what circumstances? Whose perspective are we missing? What happens when we take something somewhere else? And ultimately: does all of this add up to the outcome we actually care about?

Take obesity. A gastric band can work. For an individual receiving one, it can produce meaningful weight loss. That is important evidence. But it does not tell us how to reduce obesity across a population. The two questions sit at different scales. Once we step back, we see food systems, transport, planning, employment, income, commercial incentives, behaviour, social norms and many other forces operating simultaneously. The point is not that every evaluation needs to encompass the whole system. It is that the thing being evaluated sits somewhere inside one. If we forget that, the boundary of the evaluation can quietly become the boundary of our thinking.

Homelessness gives me the same discomfort. In the United Kingdom, we know more about homelessness than we ever have. We have better data, more research, a far richer evidence base, increasingly sophisticated services and substantial public investment. Individual programmes change lives every day. And yet homelessness has continued to rise. Those statements are not contradictory. An intervention can work, an organisation can improve and more people can receive better help while the population outcome still moves in the wrong direction.

I often come back to the image of a river. If people are entering the river upstream faster than we can pull them out downstream, becoming ever more accomplished at rescue will never be enough. We must, of course, become very good at rescue because people in crisis need excellent help now. But prevention forces us to turn around and look upstream. What is producing the flow? Where is it coming from? Which institutions shape it? What would have to change for fewer people to enter the river in the first place? Those questions inevitably take us beyond the organisations that encounter the problem once it has already become visible. They force us to think about housing, welfare, health, justice, employment, family support, commercial pressures and the way decisions made in one part of a system can create consequences elsewhere.

Finland’s experience of Housing First is instructive because its progress on homelessness is so often compressed into the name of a single intervention. Housing First has been enormously important, but that shorthand can obscure the deeper lesson. Finland’s approach combined housing provision, support services, sustained political commitment, implementation across different levels of government and a deliberate shift away from long-term shelter use. The interesting story is therefore not merely that one intervention worked. It is that many parts of a system moved in the same direction.

And that changes what we ask of evaluation. If impact depends not simply on finding an effective intervention but on how that intervention interacts with context, institutions, implementation and other parts of the system, then producing the answer is only part of the job. The practical question becomes: what do we have to do differently if we actually want evidence to travel all the way to impact?

In Darwin, I organised my answer around five principles. They are not a methodology and certainly not a formula. They are lessons I have learned from commissioning evaluation, trying to use evidence in live decisions and repeatedly asking some very practical questions: what would make this useful? What would make it believable? What would make it last?

The first is to start with the bigger picture. People are often asked to evaluate something and, quite reasonably, move quickly into the machinery of evaluation. What are the outcomes? What data do we need? Which method should we use? These questions matter enormously, but another set should come first. Where does this intervention sit in the wider system? What else is shaping the outcome? What assumptions are built into the programme? What sits outside the formal evaluation brief but could materially change how we understand what we find?

Sometimes the most valuable contribution at the beginning of an evaluation is not a more sophisticated method. It is a better frame.

That was very much in my mind when I created the Centre for Homelessness Impact. I did not want us simply to become very good at evaluating individual homelessness programmes. I wanted us to begin with a theory of change for the whole outcome. If our ambition is that homelessness should be prevented wherever possible and, when it does occur, should be rare, brief and non-recurrent, what would have to be true across the system for that to happen?

That question led us to develop the SHARE framework: smart policy, a functioning housing system, everyone playing their part, relationships and an ecosystem of services, underpinned by leadership and resources. Its purpose was to provide a bird’s-eye view of what it takes to prevent and ultimately end homelessness, rather than allowing the field of vision to collapse around whichever programme happened to be under scrutiny. It lets us ask a different set of questions. Where does this programme sit? What can it plausibly influence? What else has to be true for it to succeed? What might disappear from view if we look only inside the programme boundary?

Once we began looking at homelessness through that wider frame, another question followed: what do we actually know across it? Where is the evidence strong? Where is it sparse? Where are we repeatedly evaluating similar kinds of interventions, and where do we know surprisingly little? That is what our Evidence and Gap Maps have helped make visible. And they exposed something slightly uncomfortable: what we have the most evidence about is not necessarily what matters most. It may simply be what has been easiest to turn into a bounded programme, fund and evaluate. An intervention with identifiable participants, a clear start and end point and measurable outputs is usually easier to study than housing supply, welfare policy, institutional incentives or the interaction between several parts of government. But ease of evaluation is not a measure of importance.

This is why “what am I missing?” belongs at the beginning of an evaluation as much as at the end. It can change not just how we interpret an answer, but which question we decide is worth asking. There is, however, a consequence to widening the frame. The larger the system we are trying to understand, the less likely it is that every important question will come with neat data, a perfect counterfactual or a definitive answer. Once we stop pretending that the programme boundary is the whole world, uncertainty inevitably increases.

That leads to the second principle: don’t let perfect be the enemy of good. I realise that can sound dangerous at an evaluation conference, so the qualification matters. I do not mean lowering standards. I do not mean making stronger claims than the evidence allows. I mean that rigour has to help decisions become better. It cannot simply consist of explaining why certainty remains unavailable.

When the Centre began, we confronted an interesting paradox. Many people in the homelessness system had strong views about what worked, yet our early mapping revealed just how sparse reliable causal evidence was in many areas. That could easily have led us towards paralysis: if the evidence was weak, perhaps the responsible thing was to wait until better studies arrived. But people were making decisions every day. Budgets were being allocated. Services were being commissioned. Lives were being affected. Waiting was itself a decision.

So from the beginning we pursued two ambitions at once. We invested in generating stronger evidence for the future while also trying to make the best available knowledge easier to use today. The Intervention Tool, Evidence Finder, evidence summaries and rapid synthesis work all grew from that instinct. The task was not to pretend the evidence was stronger than it was. It was to make uncertainty usable.

Decision-makers generally know that the world is uncertain. What they need is help understanding the shape of that uncertainty. What do we know with reasonable confidence? What is suggestive but less secure? What remains genuinely unknown? What are the consequences of being wrong? What would we risk by acting now, and what might we risk by waiting? Good decision-making does not require certainty. It requires judgement disciplined by evidence.

Sure Start is a particularly rich example of why this matters. Launched in England in 1999, it brought together a wide range of early-years services for families with young children. Early national evaluation produced some troubling findings, including evidence of possible adverse effects among some disadvantaged groups. Those findings mattered. We should never hide evidence because we dislike what it says.

But the programme was also still becoming the thing that was being evaluated. Delivery varied between places. Families had often had relatively little exposure. Some of the outcomes that mattered most would take many years to emerge. Much later, researchers were able to use linked administrative data to follow children well beyond the point at which they had left Sure Start. The picture became considerably richer. Research led by the Institute for Fiscal Studies has found improvements in educational attainment lasting through GCSEs and substantial reductions in hospital admissions at older ages.

That is not a story in which the early evidence was simply “wrong” and the later evidence “right”. It is a story about time. Different questions became answerable as the programme matured, the cohorts aged and better data became available. Evaluation is not conducted outside time. An intervention that is still developing cannot always tell us what a mature programme will become, and an outcome that takes ten years to emerge cannot be summoned into existence because our evaluation contract ends in 18 months.

The lesson is emphatically not “don’t evaluate early”. Evaluate from the beginning. But match the question, the method and the time horizon to what the intervention is mature enough to tell us. Move too slowly and evidence arrives after the decision. Move too quickly and our confidence in the answer outruns the intervention. The difficult part is calibration: using what we know today without mistaking it for the final word.

But even evidence that is technically excellent, appropriately cautious and available at the right moment can still fail to move people. Decisions are made by human beings, inside institutions, with histories, beliefs, professional identities and stories of their own. That is why the third principle matters: win hearts and minds.

Evidence never arrives in an empty room. People bring values, assumptions, accumulated experience and emotion with them. Embedding evidence in decision-making is therefore not simply a technical challenge; it is a human one. One of the strongest lessons from our work has been that trust in the messenger, shared purpose and the social conditions around evidence matter alongside the quality of the evidence itself. That is why we have treated networks, champions, leadership and shared language as part of evidence use rather than as an optional communications layer added at the end. This is also consistent with work on research use and diffusion of innovation, from Nutley, Walter and Davies to Greenhalgh and colleagues, which has long emphasised the role of relationships, intermediaries and trusted internal champions.

We learned this very clearly through our work on how homelessness is represented. Images used in public discussion have traditionally relied heavily on rough sleeping: dehumanising pictures of somebody in a sleeping bag, a person in a doorway, perhaps a cardboard sign. Rough sleeping is a devastating form of homelessness, but it is not the whole phenomenon. Families in temporary accommodation, people moving between insecure places to stay and forms of homelessness largely hidden from public view can disappear almost entirely from the picture.

Repeated images create repeated narratives. If homelessness is pictured primarily as rough sleeping, that begins to define what people think homelessness is. And once the problem has been framed in that way, some explanations and solutions become much easier to imagine than others.

That is why we developed an evidence-based image bank with people who had experienced homelessness themselves, involving them in creating and selecting how homelessness should be represented. But there was another part of the work that mattered just as much: we made the images free, rights-cleared and easy to use. The point was not simply to persuade people that another way of representing homelessness was better. It was to remove the friction that stopped them doing it. That may sound like an operational detail, but it reflects a much bigger lesson. If we want behaviour to change, we should not merely make the intellectual case for change. We should make the better choice easier to take.

Stories matter enormously in this work. They make evidence memorable. They help numbers acquire human meaning. They can allow somebody to see a problem differently. But stories can also make bad ideas almost irresistible.

Scared Straight programmes are a classic warning. The idea was intuitively powerful: take young people considered at risk of offending into prisons, expose them to the reality of incarceration and frighten them away from a criminal future. It was dramatic. It provided moral clarity. It made compelling television. It produced vivid testimonies. It felt like the sort of intervention that ought to work.

The problem was that controlled evaluations told a very different story. A Campbell Collaboration systematic review, building on work by Petrosino and colleagues, found no deterrent benefit and evidence that participation could increase later offending. That case matters precisely because the intervention was so persuasive.

The lesson is not that stories, qualitative evidence or lived experience are somehow inferior. They tell us things an impact estimate cannot: how people experience an intervention, why they participate or disengage, what mechanisms may be operating, what practitioners are observing and what researchers may have failed to notice. But different evidence has different jobs. If our claim is that an intervention caused an outcome, we need evidence capable of supporting that causal claim.

Trust becomes most important when the evidence is inconvenient. The real test of an evidence culture is not what happens when the findings confirm what everyone hoped. It is what happens when evidence challenges a cherished programme, a professional judgement or an idea around which people have invested years of identity and effort. You cannot manufacture trust at the moment you arrive with an unwelcome finding. It is built in advance: through whether you listen, whether you understand the context, whether you are transparent about uncertainty, whether people believe you are interested in helping rather than judging, and whether you are prepared to challenge your own assumptions as readily as theirs.

But trust and persuasion, important as they are, still do not guarantee action. Somebody may believe the evidence and want to respond, yet lack the authority, capability, time, resources or organisational conditions to do so. That brings me to the fourth principle: make others successful.

If you want to embed evidence, do not start with what you know. Start with what others need and what the situation demands. That sounds simple, but it represents quite a profound shift away from dissemination. Our job is not to produce an evidence product, put it on a website and then wonder why people have not used it. We need to understand what they are trying to achieve, what decision they actually face, when it has to be made and what constraints they are operating under.

The goal is not to be the hero of the story. It is to equip other people to lead.

A Spending Review is a good example. A Spending Review is not a seminar about research. It is a live political and fiscal process involving competing priorities, finite money, deadlines, negotiation and judgement. The strongest research paper in the evidence base may be intellectually impressive and practically useless if it arrives at the wrong moment or answers a question nobody is currently deciding.

Ahead of the UK’s 2025 Spending Review, we worked with HM Treasury and departments across Whitehall through evidence briefs, rapid analysis and behind-the-scenes advice shaped around the questions officials were actually grappling with. The Spending Review subsequently provided £100 million for early interventions to prevent homelessness, including funding through the Transformation Fund. I would be wary of pretending that policy ever follows a neat line from “evidence” to “decision”. It rarely does. The more important point is that evidence became useful because it was connected to a live decision, at the moment that decision was being shaped, through relationships in which difficult questions could be explored.

That is very different from sending a report and calling it knowledge mobilisation.

We have learned the same thing through evidence coaching. Sometimes what is needed is not another publication but a conversation that helps somebody surface an assumption they did not realise they were making. Sometimes it is a technical note. Sometimes it is help constructing an investment case or evaluation plan. Sometimes it is sitting alongside a team while they decide what a meaningful outcome would actually look like.

Over time, we have learned that our job is not really to deliver products. It is to enable people: to help clarify goals, navigate uncertainty and build systems that keep learning. Helping others succeed is not a side benefit. It is the strategy. This is very much in line with the literature on implementation science and knowledge brokering, which has long shown that evidence use depends on relationships, capacity, ownership and support embedded in real work, not publication alone.

The World Health Organisation’s Surgical Safety Checklist offers a different but illuminating example of the same principle. Modern surgery did not suffer from a shortage of sophisticated medical knowledge. Surgeons, anaesthetists and nurses already possessed enormous expertise. But safe surgery depends on many people coordinating many actions and pieces of information at precisely the right moments. Some serious failures were therefore not failures of knowledge; they were failures of execution.

The WHO checklist created deliberate moments at which teams confirmed critical information and anticipated risks. In the original eight-site study, major complications fell from 11 per cent to 7 per cent and inpatient deaths from 1.5 per cent to 0.8 per cent after its introduction. Much of the knowledge already existed. What changed was the infrastructure around using it. Knowing what to do is not the same as being able to do it reliably.

That is why implementation cannot be treated as the slightly dull stage that comes after the “real” intellectual work of producing evidence. Implementation is part of the intellectual challenge. It asks whether the routines, relationships, authority, resources and feedback loops around people actually allow them to do what the evidence suggests.

But there is another challenge still. Sometimes the problem is not that people are failing to implement a good model. Sometimes they are implementing it brilliantly — and the model itself is no longer the best route to the outcome.

That leads to the fifth principle: keep getting better at getting better.

I mean something more demanding here than continuous marginal improvement. Sometimes the difficult discovery is not that something fails. Sometimes a thing works, people become exceptionally good at doing it, and it genuinely produces value — yet there may still be a better way.

The history of containerisation is a useful analogy. Before containers transformed freight, ports were not simply badly run. Generations of people had become highly skilled at moving cargo: loading and unloading crates, barrels, sacks and boxes; transferring them between trucks, warehouses, docks and ships; improving cranes, storage and schedules. You could keep getting better at every one of those activities.

The more disruptive question was: why are we handling the contents so many times at all?

That question reframed the task. The significance of containerisation was not simply the invention of a metal box. It was the reorganisation of transport around a standardised unit that could move between ship, road and rail without its contents being repeatedly unpacked and handled. Suddenly, enormous amounts of activity that people had spent decades becoming more efficient at were no longer necessary in the same way.

I think that is one of the hardest lessons in public services. Something can work. We can become very good at delivering it. It can genuinely help people. And there may still be a better route to the outcome. Getting better at getting better therefore requires curiosity not only about performance — how do we make this model more effective? — but about the model itself: is this still the best way to organise the work?

That is why I distinguish between evaluation capacity and improvement capacity. Evaluation capacity asks: can we generate reliable learning? Can we design the study, gather the data, estimate the effect and understand implementation? Improvement capacity asks a different question: can we act on what we learn, adapt and keep learning?

Sometimes evidence tells us to implement the existing model more reliably. Sometimes it tells us to adapt it. Sometimes it tells us to stop. And sometimes it tells us we are becoming very good at solving the wrong problem. An organisation can therefore become excellent at commissioning evaluation without becoming particularly good at improving.

Once you see that distinction, the gap between evaluation and impact starts to look less like a communications problem and more like an organisational and system-design problem. This is what I think of as the missing middle: shared outcomes, useful data, leadership and accountability, implementation capability and continuous learning.

Evidence only improves outcomes if the system receiving it can absorb it and do something with it. If a contract cannot change, a budget cannot move, nobody owns the outcome, frontline learning has nowhere to travel and evaluation happens once every five years, we should not be surprised when an excellent evaluation produces very little change. Learning has to become part of the operating system rather than an occasional event.

That is closely connected to something I have come to think about as the movement from knowledge to norm. Evidence use is not simply a technical upgrade. It is cultural. Reports alone do not change norms. Knowledge has to become embedded in relationships, habits, identity and everyday decision-making. This is why evidence use depends not only on analysts and evaluators, but on networks, champions, leaders and practitioners who carry it into ordinary decisions until a new way of working becomes less exceptional and more instinctive.

We have tried to build some of that infrastructure in England through Test & Learn. Different places are testing different approaches, including employment support, personalised budgets, health outreach and data-led prevention. The point is emphatically not that every place should do the same thing. Quite the opposite: context matters.

But variation becomes extraordinarily valuable when we organise ourselves to learn from it. If 30 places innovate privately, we may simply end up with 30 stories. Some will sound impressive, some disappointing, and remarkably little cumulative knowledge may travel. If instead we connect local experimentation to rigorous evaluation and implementation learning while the work is still developing, variation becomes informative. What appears to work? For whom? Under what conditions? What is failing because the underlying idea is weak, and what is failing because the surrounding system makes implementation almost impossible? What needs adapting locally, and what seems robust enough to travel?

That is how experimentation begins to become infrastructure rather than an occasional project.

And so I return to the theme of valuing place. Valuing place does not mean assuming every place is so unique that nothing can ever be learned across boundaries. Nor does it mean pretending context does not matter in pursuit of a universal answer. The opportunity is to understand what travels, what needs adapting and why.

Local places need enough freedom to test, learn and respond to their circumstances while contributing to knowledge that others can use. But local learning only takes us so far because places do not control everything shaping their outcomes. Decisions about housing, welfare, migration, health, justice and public spending meet in the same communities and, ultimately, in the same people’s lives. One part of government can make perfectly rational progress against its own objective while creating pressure somewhere else. A local homelessness service can become better and better while housing supply deteriorates, rents rise or changes elsewhere in the system increase the number of people arriving at its door.

Local freedom therefore needs national coherence. That is also why systems-wide evaluation matters. Not because the evaluator somehow owns the system, but because evaluation can help make visible where policies reinforce one another, where they collide, where pressure is simply being displaced and where local delivery is being asked to solve a problem whose causes sit elsewhere.

All of that brings me back to the question I began with: what am I missing?

For me, that is more than an evaluation question. It is a way of approaching improvement. It requires humility, but humility is not passivity. It means being open to learning, grounded in evidence and committed to purpose over ego. It means being able to say, “We don’t know, but we are willing to work it out.” It means holding ideas firmly enough to test them, but lightly enough to change them. And it means recognising that evidence may require us not only to refine somebody else’s programme, but to revise our own assumptions.

This matters institutionally as much as personally. A good evidence organisation cannot ask everybody else to be open to learning while behaving as though its own model is beyond challenge. Progress sometimes depends on knowing when to hold firm and when to adapt. Adaptive leadership, implementation science and the literature on organisational learning all point in a similar direction: improvement depends on curiosity, feedback, ownership and the ability to respond when reality does not behave as expected.

None of us controls the whole system, but that does not make any of us passive. Individually, we can know our craft while remaining curious beyond it. Relationally, we can work with people who see things we cannot and who bring different forms of knowledge. Organisationally, we can build the routines, relationships and capabilities that allow learning to change what happens next. And at system level, we can keep asking the uncomfortable question: do all the individual things we are doing actually add up to the outcome we say we want?

Ultimately, evidence becomes powerful not when it merely exists, but when it becomes embedded in the way people think, decide and work every day; when learning is expected rather than exceptional; when people have the confidence and capability to act on what they know; and when changing course in response to what we discover is understood as strength rather than failure.

Good evidence does not change the world on its own. It needs people to carry it, relationships through which it can travel, institutions capable of responding to it and enough humility to notice when it is telling us something we did not expect.

That was the argument I wanted to make. We are not trying to create a society that is excellent at evaluation for its own sake. We are trying to create a better society for all. Evidence helps us get there when it makes us more rigorous about what we know, more curious about what we may be missing, more willing to listen to knowledge beyond our own expertise and more willing to change course when what we learn requires it.

The ambition, in the end, is not simply better evidence. It is better evidence, better questions, and getting better at getting better.

  • Ligia Teixeira is Chief Executive of the Centre for Homelessness Impact

‍

← Back to News