Published 14 July 2026 · Updated 14 July 2026
Most nonprofit programs are designed backwards. Funding becomes available, a proposal is written to fit it, and the program is assembled afterwards from whatever the proposal promised. It works often enough to be self-reinforcing, and it produces a sector full of activities that nobody can quite explain the logic of.
Program design done properly starts from a defined problem, tests an assumption about how change happens, and builds delivery around that. It is not more expensive than the improvised version. It is mostly a matter of sequence, and the sequence is what this guide sets out.
What is nonprofit program design?
Nonprofit program design is the process of turning a defined community need into a structured set of activities with a clear logic connecting what you do to the change you expect. It specifies the problem, the participants, the activities, the assumed causal chain, the resources required and how you will know whether it worked, before delivery begins.
Key takeaways
- Start with the problem, not the activity. If the design begins with a workshop, the problem was never properly defined.
- Name your assumptions explicitly. Every program contains a theory about why change happens, and most are never written down.
- Design for the participant you will actually get, not the motivated one who shows up in the proposal.
- Build measurement in at design stage. Retrofitting evaluation onto a running program never produces good data.
- Pilot small and expect to be wrong. A design that survives first contact unchanged usually was not specific enough to test.
What this guide covers
- Start with the problem, properly defined
- Who the program is actually for
- Naming your theory of change
- Building the logic model
- Designing activities that fit the logic
- Designing around real barriers
- Building measurement in from the start
- Resourcing honestly
- Piloting and iterating
- Designing for funders without designing for funding
- Co-design without theatre
- Designing across partners
- Delivery capability is part of the design
- When to redesign and when to stop
- Design mistakes that recur
- Documenting the design
- Designing for equity
- Scaling a program that works
- The first ninety days of delivery
- Frequently asked questions
Start with the problem, properly defined
A well-defined problem statement names who is affected, what specifically is happening to them, where and when it occurs, and what the consequence is if nothing changes. Most nonprofit problem statements name a topic instead: youth unemployment, food insecurity, social isolation. A topic is not a problem you can design against.
The test is whether two people reading your problem statement would design similar programs. “Newcomer youth in our region face barriers to employment” would produce twenty different programs. “Newcomer youth aged 18 to 24 with foreign credentials are being screened out at the application stage because Canadian employers do not recognise their qualifications” produces roughly one.
Getting there requires evidence rather than assumption. Talk to the people affected before you design anything, and talk to enough of them that you are hearing patterns rather than anecdotes. Programs designed from a single compelling story reliably serve that one person well and everyone else poorly.
A discipline worth adopting: Write the problem statement, then ask three people who experience the problem whether it describes their situation. Their corrections are the most valuable design input you will get, and they cost nothing.
Who the program is actually for
Every program has an implicit ideal participant, and it is almost always someone with time, transport, stable housing and enough confidence to walk into an unfamiliar room. Design that assumes that participant will reach the people who need it least.
Define your participants concretely: age range, circumstances, what is already true about their week, what they would have to give up to attend. That last question is the one that predicts uptake, and it is the one most often skipped.
Who you say it is for
Usually a broad demographic drawn from the funding call. Useful for proposals, useless for design.
Who it is designed for
The implicit participant your delivery model assumes. Often much narrower than intended.
Who actually shows up
The people for whom your delivery model is realistically accessible. This is the honest answer.
Who you meant to reach
Frequently the group facing the most barriers, and therefore the least likely to arrive without deliberate design.

Naming your theory of change
Every program contains a theory: if we do this, then that will happen, because of some mechanism. Most organisations hold that theory implicitly, which means it is never examined and never tested. Writing it down takes an hour and frequently changes the design.
State it as a chain with the “because” included. Not “we run workshops so participants gain skills” but “we run workshops because participants lack a specific skill, and that skill is the binding constraint on their employment, and a workshop is the format that fits their availability”. Each clause is an assumption you can check.
The value is in finding the weak link. Often the skill is not the binding constraint at all, and the program is well-executed against the wrong bottleneck. That discovery is much cheaper at design stage than after two years of delivery. Our guide to theory of change for nonprofits covers the full method.
“A program that works is usually a correct theory competently delivered. A program that does not is usually a wrong theory delivered beautifully.”
Building the logic model
A logic model turns the theory into something you can plan and budget against. It maps inputs to activities to outputs to outcomes, with the causal links made explicit. It is a thinking tool first and a funder deliverable second, and treating it in the reverse order is why so many are written after the fact.
- Inputs: what you need. Staff hours, space, materials, partner commitments, participant time.
- Activities: what you do. Specific and countable, not “support” or “engagement”.
- Outputs: what those activities directly produce. Sessions delivered, people reached, materials distributed.
- Short-term outcomes: what changes for participants within the program. Knowledge, confidence, connections.
- Longer-term outcomes: what changes afterwards. Employment, stability, wellbeing.
- Assumptions: written alongside each arrow, not in a footnote. This is where the real thinking lives.
Keep it on one page. A logic model that requires a legend has stopped being a design tool and become a compliance artefact, and nobody on the delivery team will look at it twice.
Designing activities that fit the logic
Only now should you decide what the program actually does. Designing activities first is the most common failure mode in the sector, and it is understandable: activities are concrete, energising and easy to describe to a board.
Test each proposed activity against the logic. Which outcome does it drive, through which assumed mechanism? Activities that cannot answer that question are usually there because they were in the last program, or because a staff member enjoys running them. Both are real reasons and neither is a design rationale.
Be equally rigorous about dosage. How many sessions, how long, how far apart. Under-dosed programs are extremely common in the sector, because funding covers six weeks and the change being sought realistically takes six months. Naming that gap honestly is better than delivering something you know is too short and reporting it as a success.
The question to ask about every activity: “If we removed this, which outcome would weaken?” If nobody can answer, the activity is decoration and it is consuming budget.
Designing around real barriers
A program that is well-designed on paper and inaccessible in practice will report low uptake and conclude there was insufficient demand. Almost always the demand existed and the design excluded it.
Barriers to design against explicitly
- Timing: does the schedule assume people are not working, or are working standard hours?
- Transport: what does attendance cost in bus fare and travel time, and who absorbs it?
- Childcare: is it provided, or is attendance effectively limited to people without young children?
- Language: are materials and delivery accessible to the people you named in the problem statement?
- Documentation: does registration require paperwork some participants will not have?
- Trust: does the program require walking into an institution people may have reason to avoid?
Each of these is a design decision with a cost attached. Budget for them at design stage, because retrofitting childcare or bus tickets mid-program means finding money you did not raise.
Building measurement in from the start
Measurement designed at the same time as the program is cheap and useful. Measurement bolted on in year two is expensive, awkward and produces data nobody trusts. The difference is entirely about sequence.
Decide three things during design: what evidence would convince a sceptic that this worked, when you will collect it, and what you will do if it shows the program is not working. That last commitment is what separates evaluation from reporting.
Keep the instrument proportionate. A fifteen-question survey administered consistently beats a forty-question one administered when someone remembers. Our program evaluation guide covers instrument design and analysis in detail.
Resourcing honestly
Programs are routinely costed on the assumption that existing staff will absorb delivery alongside their current work. Sometimes that is true. Usually it means something else quietly stops, and nobody records what.
Cost the real thing: delivery hours, preparation, participant support, data collection, supervision, and the administrative overhead of reporting. Then compare that to what the funding actually covers. If there is a gap, decide deliberately how to close it rather than discovering it in month four.
This is also where diversified funding matters. A program dependent on a single grant is a program with a defined end date whether or not it works, which our guide to revenue diversification addresses directly.
Piloting and iterating
Pilot small, deliberately, and with the explicit expectation of being wrong about something. A pilot that confirms every assumption was probably not specific enough to be tested, or was evaluated too gently.
Run it with a small cohort, collect data from the start, and schedule a redesign point before scaling. Protect that redesign point in the funding agreement if you can, because funders generally prefer a program that improved to one that was defended.
Expect three kinds of finding: things that did not work, things that worked for a different reason than you assumed, and participants you did not expect to reach. All three should change the design. Only the first usually does.
Designing for funders without designing for funding
The tension is real. Funding calls have priorities, and an organisation that ignores them will not be funded. The failure mode is not responding to funders; it is letting the funding call define the problem.
A workable discipline is to design the program you believe in first, then assess which funders it fits. Where a call requires adjustment, be explicit about whether the adjustment improves the program, leaves it neutral, or weakens it. Organisations that never ask this question drift over a decade into delivering whatever was fundable.
Where the fit is genuinely poor, declining is a legitimate strategic choice, and one that stronger organisations make more often. If your fundraising strategy is not diversified enough to make that possible, that is the real problem to solve, and our fund development guide is the place to start.
Co-design, and how to do it without theatre
Co-design has become a sector expectation, which means it is frequently performed rather than practised. A single consultation session held after the program is designed is not co-design; it is validation seeking, and participants can tell the difference immediately.
Real co-design means participants influence decisions that are still open. That requires holding decisions open, which is uncomfortable when a proposal deadline is approaching. The practical compromise most organisations can manage is to co-design the delivery model while holding the problem statement and outcomes relatively fixed.
Be explicit about what is negotiable. Telling a group that the funding requires an employment outcome but that everything about how the program runs is genuinely open produces far better input than pretending everything is on the table. People engage more seriously with a real constraint than with a false invitation.
Pay participants for design time. Expecting people with the least resources to donate expertise is both an equity problem and a data quality problem, because unpaid consultation attracts the people with the most spare capacity, who are rarely the people the program is for.
A test for genuine co-design: Can you name one significant decision that changed because of participant input? If not, the process was consultation, and it is worth being honest about that in your reporting.
Designing a program across partners
Multi-partner programs fail in predictable ways, and almost all of them trace back to design rather than delivery. Partners agree on the goal, assume they agree on everything else, and discover in month three that they had different theories about how the change happens.
Do the logic model together, in one room, before anything is written into a proposal. The disagreements that surface are the valuable part. A partner who believes the binding constraint is skills and one who believes it is employer bias will design incompatible programs while using identical language.
Agree these before the proposal, not after the grant
- Who is accountable for which outcome, not just which activity
- How participants move between partners and who owns the relationship
- What data each partner collects, in what format, and who consolidates it
- How disagreements about delivery get resolved and by whom
- What happens to the program if one partner withdraws
- Who speaks to the funder, and whether partners may contact them independently
Write the answers down even where partners know each other well. Long relationships make organisations less likely to document and no less likely to disagree, and the memorandum is worth most precisely when the relationship is under strain.
Delivery capability is part of the design
A design that assumes capability the team does not have is not a design, it is a wish. This is the least discussed constraint in program design and one of the most common causes of underperformance.
Be concrete about what delivery requires. Facilitation of a group with complex needs is a specific skill, not a general one. Trauma-informed practice is training, not attitude. Data collection that participants will complete honestly depends on how the person asking is perceived. Each of these is a design decision with a hiring or training implication.
Where capability is missing, the choices are to hire, to train, to partner, or to simplify the design. All four are legitimate. What does not work is proceeding on the assumption that a committed staff member will figure it out, which places an unreasonable burden on the person least able to refuse it.
“Most underperforming programs are not badly designed. They are designed for a team that does not exist, delivered by the team that does.”
When to redesign and when to stop
Nonprofits are considerably better at starting programs than ending them. A program that is not achieving its outcomes consumes staff time, occupies strategic attention and quietly crowds out the thing that would work better. Deciding in advance what would trigger a stop is the discipline that prevents that.
Distinguish between three situations. The theory is right and delivery is weak, which calls for operational fixes. The theory is wrong but the need is real, which calls for redesign. The need has changed or is being met elsewhere, which calls for stopping.
Organisations reliably diagnose the second and third as the first, because operational fixes feel like progress and do not require admitting a strategic error. Ask an outsider to make the call if the internal answer keeps coming back the same.
Ending well matters. Tell participants early, help them transition to alternatives, tell funders directly rather than letting a grant simply lapse, and write down what was learned. A program stopped deliberately with a clear account of why strengthens an organisation’s credibility. One that fades out damages it.
Design mistakes that recur
Activity-first design
Starting from a workshop, an app or an event and reverse-engineering a rationale. The most common error in the sector.
Unnamed assumptions
A theory of change that lists steps without stating why each step causes the next one.
Optimistic dosage
Six weeks of contact aimed at a change that realistically takes six months, because that is what the grant covered.
Measurement as reporting
Collecting only what the funder asked for, which tells you whether you delivered but not whether it worked.
Designing for the willing
Building for the participant who is already motivated and available, then reporting low uptake as low demand.
No stop criteria
Beginning without agreeing what evidence would cause the organisation to redesign or end the program.
Documenting the design so it survives staff turnover
Programs outlive the people who designed them, usually by several years. When the designer leaves, the reasoning frequently leaves with them, and what remains is a set of activities that the next team delivers faithfully without knowing why any of it is shaped the way it is.
Keep a short design record alongside the logic model: the problem statement, the theory including assumptions, what was considered and rejected, what the pilot changed and why, and what the stop criteria are. Three pages is enough and it is worth more than any operations manual.
Review it annually against what is actually being delivered. Programs drift, and drift is not always bad, but undocumented drift means nobody can tell whether the current version still matches the theory it was built on. That review is also the natural moment to check whether the original problem still exists in the form you described.
This is the same discipline that underpins strong evaluation and strong governance, and it belongs in your broader capacity building work rather than sitting with one program manager.
Designing for equity rather than assuming it
Equity in program design is not a values statement appended to a proposal. It is a set of concrete decisions about who can realistically participate, whose knowledge shaped the design, and who bears the cost of taking part. Programs that treat it as a statement reliably serve the least marginalised people in their target group.
Start by asking who is missing from the room where the design is happening. If everyone involved shares a professional background, a language and a set of assumptions about how services work, the design will encode those assumptions invisibly. That is not bad intent; it is the default outcome of homogeneous design teams.
Then examine the cost of participation. Time away from paid work, transport, childcare, the emotional cost of retelling a difficult history to a stranger, and the risk of being identified as someone who uses a service. Each of these falls unevenly, and each is a design variable you can adjust.
Finally, check who your measurement excludes. Surveys requiring literacy in one language, digital forms requiring a device and data, and outcome measures defined by professionals rather than participants all quietly remove people from your evidence base. A program can appear to work well precisely because the people it failed did not complete the survey.
A practical check: Compare the demographic profile of the people who completed your program against those who started it. The gap between the two is usually your equity finding, and it is data you already have.
Scaling a program that works
Success creates its own design problem. A program that works well with thirty participants in one neighbourhood does not automatically work with three hundred across a region, and the assumption that it will is how good programs get diluted into ineffective ones.
Identify what is actually doing the work before you scale. Frequently it is a specific relationship, a particular facilitator, or the fact that the group is small enough for people to speak. If the active ingredient is intimacy, scaling by increasing group size removes the very thing that made it effective.
Consider replication over expansion. Running five separate small programs preserves the mechanism; running one large program often does not. Replication is administratively harder and it is frequently the honest answer, particularly for programs whose logic depends on trust.
Scale the measurement alongside the program. A model that worked in one site needs evidence that it works in the next one, because the original result may have depended on local conditions nobody documented. Treat each new site as a test rather than as a rollout.
The first ninety days of delivery
The gap between a design and what actually happens opens in the first three months, and it opens quietly. Facilitators adapt sessions in the moment, registration processes get simplified under pressure, and data collection slips to whenever there is time. Each adjustment is reasonable and the cumulative effect is a program that no longer matches its own logic model.
Schedule a design review at day thirty and again at day ninety, with the logic model on the table. The question is narrow: what are we doing differently from what we designed, and does that change strengthen or weaken the causal chain? Some drift is improvement and should be written into the design. Some is erosion and should be corrected.
Ask delivery staff rather than inferring from data. The people running sessions know exactly which parts are working and which they have quietly stopped doing, and they will usually say so if asked in a way that does not sound like an audit.
Fix the data collection first if it has slipped, because everything else you will want to know in month twelve depends on it. A program with strong delivery and no data cannot demonstrate anything, and cannot be improved except by intuition.
Designing a new program, or rethinking one that is not working?
Our free nonprofit assessment looks at program design, evaluation and capacity together, and tells you plainly where the logic breaks.
Frequently asked questions
What is nonprofit program design?
It is the process of turning a defined need into structured activities with an explicit logic linking what you do to the change you expect. It covers the problem, participants, activities, assumptions, resources and measurement, decided before delivery begins.
What is the difference between a logic model and a theory of change?
A theory of change explains why you believe change will happen, including the mechanism and assumptions. A logic model maps the practical chain from inputs through activities and outputs to outcomes. The theory is the reasoning; the logic model is the plan.
How specific should a problem statement be?
Specific enough that two people reading it would design similar programs. It should name who is affected, what is happening, where and when, and the consequence of inaction. A topic such as food insecurity is not a problem statement.
Should we design the program or find the funding first?
Design first where you can. Responding to funder priorities is legitimate, but letting a funding call define the problem is how organisations drift into delivering whatever happens to be fundable.
How long should a pilot run?
Long enough to observe the short-term outcomes in your logic model, with a redesign point scheduled before any scaling. What matters more than duration is collecting data from the first cohort rather than starting measurement later.
How do we design for people who face the most barriers?
Name the barriers explicitly at design stage, including timing, transport, childcare, language, documentation and trust, and budget for addressing them. Barriers left to be solved during delivery are usually not solved at all.
What if the pilot shows the program is not working?
That is a successful pilot. Decide during design what you will do with a negative result, because that commitment is what separates genuine evaluation from reporting. Most funders respond better to a program that improved than to one that was defended.
Do small organisations need a formal logic model?
A one-page version, yes. It does not need to be elaborate, and keeping it to a single page is what keeps it useful to the delivery team rather than turning it into a compliance document.



