Nonprofit Outcome Measurement: What to Track, How to Track It, and Which Tools Actually Help

A team reviewing charts and data on laptops around a table
📖 19 min readProgram Design
FG
For Good Consultants
Published 1 September 2026 · Updated 1 September 2026

How should a nonprofit measure and track outcomes?

Work backwards from the change you exist to create, not forwards from the data you happen to collect. Write down your theory of change, pick the smallest number of indicators that would genuinely tell you whether it is working, decide how each one gets collected before you promise it to anyone, and build the collection into the delivery of the program rather than bolting it on afterwards. Most organisations do not have a tools problem. They have too many indicators, collected inconsistently, for an audience they have not defined.

Key takeaways

  • Outputs are not outcomes. How many people you served is an output. Whether their situation changed is an outcome. Funders increasingly ask for the second and most reporting still delivers the first.
  • Five to seven indicators is usually the right number. Organisations that track thirty track none of them well, and the data is too thin to use for decisions.
  • If it is not collected during delivery, it will not be collected. Measurement that depends on someone remembering to do it at month end does not survive a busy quarter.
  • Your tool matters far less than your definitions. A spreadsheet with clear, consistent definitions beats a case management system where three staff interpret ‘engaged’ differently.
  • Separate funder reporting from internal learning. They need different things, on different timelines, and collapsing them into one exercise usually means neither is done well.
  • Some of the most important change is qualitative. Structured stories and consistent short interviews are legitimate evidence, not a fallback for organisations that cannot count.

Outputs, outcomes and impact: the distinction that decides everything

Almost every difficult conversation about nonprofit measurement comes back to three words being used interchangeably when they mean different things. Getting the vocabulary straight is not pedantry. It is what lets a board, a staff team and a funder agree on whether a program is working.

Outputs

What you did and how much of it. Sessions delivered, people served, meals distributed, calls answered. Outputs are easy to count, entirely within your control, and tell you nothing about whether anything changed.

Outcomes

What changed for the people you served. Skills gained, housing secured, isolation reduced, income increased, confidence improved. Outcomes are harder to measure, partly within your control, and are the thing your organisation actually exists to produce.

Impact

The longer-term, population-level change your outcomes contribute to. Reduced homelessness in a city, improved graduation rates in a region. Impact is rarely attributable to one organisation and most small nonprofits should not claim it.

Indicators

The specific, defined things you count or ask in order to know whether an outcome happened. ‘Participants report increased confidence’ is an outcome. ‘Score on a five-point confidence scale at intake and at twelve weeks’ is an indicator.

The practical consequence is that outputs alone cannot answer the question a funder or a board is actually asking. If you report that you delivered two hundred and forty workshop hours to ninety participants, you have described your activity. You have not said whether the ninety people are better off. A program can deliver a great many hours and change very little, and organisations that only track outputs generally cannot tell the difference until a funder asks a question they cannot answer.

The reframe that helps most Stop asking ‘what data do we have?’ and start asking ‘what would have to be true for us to believe this program worked?’ Then ask what the cheapest honest evidence for that would be. That sequence produces a short list of useful indicators. The reverse sequence produces a long list of things you happen to be able to count.

“A program can deliver a great many hours and change very little. Outputs cannot tell you which one you are running.”

Why nonprofit measurement usually fails

Organisations rarely fail at measurement because staff do not care or because the right software has not been purchased. The failures are structural and they repeat.

The indicator list grew by accretion. Every funder added two requirements, every board member suggested one more thing worth knowing, and nobody ever removed anything. The result is a reporting template with thirty fields, most of them completed approximately, none of them examined. Data that is collected but never used teaches staff that the exercise is compliance rather than learning, and the quality drops accordingly.

Collection was designed after the program. A survey that has to be administered separately from the service will be administered inconsistently. A follow-up call at six months requires somebody whose job includes making it. When measurement is designed as an addition to delivery rather than a part of it, it is the first thing dropped when the caseload rises.

Definitions live in people’s heads. Three frontline staff record ‘engaged’ three different ways, one counts a participant at first contact and another at first completed session, and nobody notices until the annual numbers look strange. This is the single most common reason nonprofit data cannot be trusted, and it has nothing to do with the tool.

Nobody owns it. Measurement gets assigned to whoever is most comfortable with spreadsheets, on top of a full role, without time protected for it. It then runs well for as long as that person has capacity and stops when they do not.

The data is never fed back. If frontline staff never see what their data showed, they experience it as paperwork for head office. Teams that get a quarterly read on what the numbers said about their own program collect noticeably better data, because it has become theirs.

The compliance trap Once measurement is experienced as something done for funders rather than for the work, quality degrades in a way that is very hard to reverse. Fields get completed at year end from memory, definitions drift, and the organisation ends up with a dataset that is technically complete and practically useless for making any decision.

Start from your theory of change

A theory of change is a plain statement of how your activity is supposed to produce the change you want. It does not need to be a diagram or a consulting deliverable. It needs to be written down, agreed by the people delivering the program, and specific enough that it could be wrong.

The simplest useful form is a chain of sentences: we do this, which produces this, which leads to this, which contributes to this. Each arrow is an assumption, and each assumption is a candidate for measurement.

A worked example We run a twelve-week employment readiness program (activity) so that participants build interview skills and a current resume (immediate outcome), so that they secure interviews (intermediate outcome), so that they move into sustained employment (primary outcome), contributing to reduced long-term unemployment in our community (impact). Each arrow is testable. The most useful indicators sit at the intermediate outcomes, where you have real influence and can still measure within a funding cycle.

Our guide to building a theory of change walks through this in more depth, and program evaluation covers what to do once the framework exists. Writing this out surfaces the assumptions that everyone believed but nobody had stated. It is common for a team to discover that two staff hold materially different views of what the program is for, or that the chain contains a step nobody is actually supporting. That conversation is worth having on its own, before any indicator is chosen.

It also tells you where to stop. If your chain ends at population-level impact, you are not going to measure the last link and you should not promise it. Measure the links you influence directly and describe the rest as contribution rather than attribution. Funders who understand measurement respect that distinction; the ones who do not are usually reassured by an organisation that can explain it clearly.

Choosing what to measure

Once the chain is written, the selection question becomes tractable. For each outcome you have named, ask what would count as evidence, then apply four filters.

  1. Would it change a decision? If an indicator moved sharply in either direction, would anyone do anything differently? If not, it is trivia. This single filter usually removes a third of a bloated indicator list.
  2. Can you collect it consistently, every time, with the staff you have? An indicator that requires a data collection capability you do not have is a plan to have inconsistent data. Choose the version you can actually sustain.
  3. Is it defined tightly enough that two different staff would record it the same way? Write the definition down, including the edge cases. When does someone count as a participant? What counts as a completion? What happens if they return after six months?
  4. Does it cover the change, not just the activity? At least half your indicators should sit at outcome level. If they are all outputs, you have built an activity report.

For most small and mid-sized organisations, five to seven indicators per program is the right target. That is enough to describe activity, immediate outcomes and at least one intermediate outcome, and few enough that each one can be defined properly, collected consistently and actually looked at.

A reach indicator

Who you served and how many, with the demographic breakdown you need for equity questions and funder reporting.

A dose indicator

How much of the program each participant actually received. Frequently the most diagnostically useful number you have, because outcomes usually track dose and it tells you where people drop out.

One or two immediate outcome indicators

The change you expect by the end of the program. Usually a before-and-after measure on something specific.

One intermediate outcome indicator

The change that matters beyond the program. Often requires follow-up contact, which needs to be designed in from the start.

A quality or experience indicator

Whether participants found it useful and would recommend it. Cheap to collect, and a leading signal that something has gone wrong.

One equity indicator

Whether outcomes differ across the groups you serve. Aggregate results routinely conceal that a program works well for one group and not another.

Resist the temptation to add ‘while we are at it’ Every additional field has a cost paid by frontline staff at the moment of delivery, and the cost is not the thirty seconds of typing. It is the attention taken away from the person in front of them. If you cannot say which decision an indicator informs, do not collect it.

Collecting the data without breaking the program

The design principle that matters more than any other: measurement should be a by-product of delivering the service, not a separate activity performed afterwards. Every step away from that principle costs you data quality.

What this looks like in practice

  • Intake captures the baseline. If you need a before-and-after measure, the ‘before’ is part of the intake conversation, not a survey emailed later. Anything not captured at intake is usually lost.
  • Attendance is recorded at the session, by the person running it, in the same place every time. Not reconstructed at month end.
  • The exit measure is built into the last session, as a structured activity rather than a form handed out at the door. Response rates change dramatically based on this one choice.
  • Follow-up is somebody’s named job, with a scheduled window and a defined number of attempts. Follow-up that belongs to everyone belongs to no one.
  • The same questions are asked the same way each time. Changing your scale between years means you have two datasets, not a trend.

Follow-up is where most measurement plans quietly fail. Six and twelve month follow-ups are the indicators funders most want and the ones organisations are least able to deliver, because they require reaching people who have finished with your service. Design for it at intake: ask permission to follow up, collect more than one contact method, explain why you will call, and if your budget allows it, offer a small honorarium. Then set a realistic expectation about response rate and report it honestly rather than presenting a partial follow-up sample as if it were the whole cohort.

Write the data dictionary before you write the survey One page per program: every indicator, its exact definition, who collects it, when, in what system, and what counts as an edge case. This document costs a few hours and prevents the most expensive measurement problem there is, which is discovering at year end that your numbers were recorded inconsistently and cannot be compared.

A team reviewing charts and data on laptops around a table
Data becomes useful at the point a team sits down with it. Collection without a scheduled review is just storage.

Tools: what to use at each stage

Software is the part of measurement organisations most want to talk about and the part that matters least. A well-defined spreadsheet outperforms a poorly configured case management system every time, because the constraint is almost never storage or reporting features. It is definitional clarity and collection discipline.

That said, tools do matter once volume and complexity rise. A rough guide to what fits where.

Spreadsheets

Right for a single program, a manageable number of participants, and a team that can agree on one file with locked column definitions. Cheap, flexible, universally understood. They break down with multiple programs, multiple people editing, or any requirement to track an individual across services over time.

Form tools feeding a sheet

A survey or form tool collecting into a spreadsheet or database is the most common good-enough setup for small organisations. It enforces field types, timestamps entries and removes transcription errors, while keeping analysis somewhere familiar.

Case management systems

Right when you track individuals across multiple services over time, when several staff need concurrent access to the same records, or when confidentiality requirements make a shared spreadsheet inappropriate. Real cost is configuration and staff training, not licensing.

Dashboards and BI tools

Useful once your data is already clean and consistent. A dashboard built on inconsistent definitions produces confident-looking charts of unreliable numbers, which is worse than no dashboard.

Qualitative tools

Structured note templates, consistent interview guides and a simple tagging scheme. Most organisations do not need specialist qualitative software; they need a consistent format and somewhere to keep it.

How to sequence a tool decision Define your indicators, write the data dictionary, and run the collection manually for one full program cycle. Then buy. Organisations that purchase first almost always configure the system around the data they were already collecting, which locks in the problems they were trying to fix.

Two practical cautions. First, whatever you choose, confirm you can export your own data in a usable format, and test that export before you depend on it. Second, check where the data is stored and whether that satisfies your privacy obligations and any funder requirements, particularly if you hold sensitive personal information. These are questions to ask before purchase, not after a migration.

Measuring what is hard to measure

A great deal of nonprofit work produces change that does not reduce comfortably to a number. Reduced isolation, restored dignity, cultural continuity, community capacity, a person feeling safe enough to ask for help. Organisations doing this work often feel caught between reporting something trivial and reporting nothing at all.

The answer is not to abandon measurement or to invent a proxy that misses the point. It is to treat qualitative evidence as evidence, and to collect it with the same discipline you would apply to a number.

What rigorous qualitative measurement looks like

  • A consistent question set, asked the same way to everyone, at the same points in the journey. Consistency is what turns anecdote into data.
  • Structured story collection with a simple template: the situation before, what changed, what contributed to it, what is different now. Same shape every time, so stories can be compared.
  • A tagging scheme applied to responses so you can say how many people described a particular kind of change, rather than only quoting the most articulate participant.
  • Someone other than the person who delivered the service asking the questions where that is feasible, because participants are reluctant to tell their support worker that the support did not help.
  • Reporting the range, not just the highlights. Including what did not work is what makes the positive findings credible.

Validated scales are worth knowing about for the domains where they exist, including wellbeing, social connection and self-efficacy. Using an established instrument means your numbers are comparable to other organisations and are harder for a sceptical funder to dismiss. Check licensing terms before adopting one, since some are free for nonprofit use and some are not.

The single most common qualitative mistake Collecting only success stories. If every story in your annual report is a triumph, an experienced funder reads it as marketing rather than evidence. Organisations that report honestly on what did not work, and what they changed as a result, build considerably more credibility than those with an unbroken record of success.

Funder reporting and internal learning are different jobs

These two purposes get collapsed into one exercise in most organisations, and it is why measurement feels like a burden. They have different audiences, different timelines and different tolerances for bad news.

Funder reporting

Answers questions somebody else chose, on their schedule, in their format, usually annually or semi-annually. Backward-looking, accountability-oriented, and correctly cautious. It has to be accurate and it has to align with what you committed to in the proposal.

Internal learning

Answers questions you chose, on a rhythm that lets you act, usually quarterly. Forward-looking and improvement-oriented. It needs to make room for findings that are uncomfortable, because those are the ones worth acting on.

The practical resolution is one collection system serving two outputs. Collect a single, well-defined dataset. Produce funder reports from it on their schedule and in their format; our guide to nonprofit impact reporting covers how to present them well. Separately, hold a short internal review each quarter that asks three questions: what do the numbers say, what surprised us, and what will we change before next quarter. Write the answer to the third question down and check it at the next review.

The review is the part that creates value Data that is collected and never discussed produces compliance. Data that a team sits with for ninety minutes every quarter produces decisions, and it also produces better data, because staff who see their numbers used start caring about their accuracy.

One further point on funders. Where a funder’s required indicators do not fit your program, say so during the proposal stage rather than agreeing and then reporting something approximate. Most funders would rather negotiate a sensible indicator than receive a number that nobody believes, and the conversation is far easier before the money moves than in the final report.

A realistic first ninety days

If you are starting from a position where measurement is inconsistent or absent, the temptation is to design a complete framework across every program at once. That approach reliably produces a well-designed document that nobody implements. Start with one program and get it genuinely working.

  1. Weeks 1 to 2 — Write the theory of change for one program With the people who deliver it, not for them. One page. Argue about it until the chain is one everyone recognises.
  2. Weeks 3 to 4 — Choose five to seven indicators Apply the four filters. Write the data dictionary, including definitions and edge cases. Get frontline staff to challenge each definition before you finalise it.
  3. Weeks 5 to 6 — Design collection into delivery Decide the exact moment each indicator is captured, by whom, in what. Build the intake baseline and the exit measure into the program’s own materials.
  4. Weeks 7 to 8 — Run a pilot cycle Use it live with a small group. Expect to find that two definitions were ambiguous and one collection point does not work in practice. This is the point of a pilot.
  5. Weeks 9 to 10 — Fix and document Revise the dictionary. Write the one-page instruction that a new staff member could follow.
  6. Weeks 11 to 12 — Hold the first review Look at what you have, even though it is thin. Ask the three questions. Decide one thing to change. Then schedule the next review before the meeting ends.
  7. After ninety days — Extend to the next program With the pattern proven, the second program takes a fraction of the effort, because the definitions, the rhythm and the review habit already exist.

Name an owner and protect the time Measurement needs a named person with hours protected for it, not an assignment added to a full role. In a small organisation this might be a few hours a week. What it cannot be is nobody’s job in particular, because it is the work that is always reasonable to postpone.

Mistakes we see most often

The ones that cost the most

  • Measuring everything. A thirty-indicator framework produces thin, inconsistent data across the board and no capacity to interrogate any of it.
  • Designing measurement after the program. If collection was not built into delivery, it will be the first thing to go when the caseload rises.
  • Definitions that live in people’s heads. The most common reason nonprofit data cannot be trusted, and entirely preventable with a one-page dictionary.
  • Buying the system first. Software configured around your existing bad habits preserves them at greater expense.
  • Promising follow-up you have not resourced. Six-month outcomes require somebody whose job includes making the calls. Committing to them in a proposal without that person is a reporting problem you have scheduled for yourself.
  • Reporting only successes. It reads as marketing and it costs you credibility with exactly the funders whose opinion matters most.
  • Collecting data nobody ever discusses. If there is no scheduled review, the whole exercise is storage, and staff will correctly treat it as paperwork.
  • Claiming impact you cannot attribute. Overstating your causal role invites scrutiny you will not survive, and understates the contribution you can actually evidence.

The through-line in all of these is that measurement problems present as technical problems and are almost always design problems. Organisations reach for a new system when what they need is a shorter list, a clearer definition and a standing meeting where somebody looks at the numbers and decides something. That is unglamorous work, and it is the work that turns data collection into an organisation that knows whether it is doing any good.

Not sure what you should be measuring?

We help nonprofits build outcome frameworks that staff will actually use: theory of change, indicator selection, data dictionaries, collection design and the review rhythm that turns numbers into decisions.

Book a conversation

Frequently asked questions

What is the difference between outputs and outcomes?

Outputs are what you delivered, such as sessions run or people served. Outcomes are what changed for those people, such as skills gained, housing secured or isolation reduced. Outputs are entirely within your control and easy to count. Outcomes are the reason the organisation exists. Most reporting problems come from an organisation counting the first and being asked about the second.

How many indicators should we track?

Five to seven per program is the right target for most small and mid-sized organisations. Enough to cover reach, dose, at least one immediate outcome, one intermediate outcome and a quality measure. Beyond that, definitional quality and collection consistency drop, and nobody has capacity to interrogate the results.

Do we need a case management system?

Only if you track individuals across multiple services over time, need several staff in the same records concurrently, or have confidentiality requirements a shared spreadsheet cannot meet. Otherwise a form tool feeding a well-defined spreadsheet is usually sufficient. Run your indicators manually for one full cycle before buying anything.

How do we measure things that are hard to quantify?

Collect qualitative evidence with the same discipline you would apply to numbers: a consistent question set asked the same way to everyone, a structured story template, a tagging scheme so you can count how many people described each type of change, and honest reporting of the range rather than only the highlights. Validated scales exist for wellbeing, social connection and self-efficacy if a number is required.

How do we get follow-up data after people leave the program?

Design for it at intake. Ask permission to follow up, collect more than one contact method, explain why you will be in touch, and make the follow-up a named person’s scheduled job with a defined number of attempts. Then report your response rate honestly rather than presenting a partial sample as the full cohort.

What if a funder asks for indicators that do not fit our program?

Raise it at the proposal stage, before the agreement is signed. Explain what your program actually produces and propose an indicator that captures it. Most funders prefer a negotiated measure they can trust over a required one that nobody believes, and the conversation is much harder once you are writing the final report.

Should we report results that were not good?

Yes, with the context and what you changed as a result. An organisation that only reports success reads as promotional to experienced funders. Reporting a program that underperformed, alongside a clear account of what you learned and adjusted, is one of the more effective ways to build funder confidence.

Can we claim impact?

Be careful with the word. Impact usually means population-level change, which very few individual organisations can attribute to their own work. Measure the outcomes you influence directly and describe your relationship to the wider change as contribution. Funders who understand measurement respect the distinction, and overstating attribution invites scrutiny you are unlikely to withstand.

Scroll to Top