Evaluation: what to measure in a mentoring scheme
Mentor evaluation in a UK scheme takes three usable forms, and a decent report contains all three: counts of what the scheme did, a distance travelled measure completed by the person supported at intervals, and one validated scale used consistently from the start of the match to the end of it.
A coordinator who has been asked to prove that a scheme works, usually by a funder, usually with three weeks’ notice, and usually without a budget line for evaluation. It is also for a volunteer being handed a form and wondering what it is for.
Honesty about a soft outcome. A scheme that measures well can defend a modest result; a scheme that measures badly has to choose between exaggeration and silence, and most choose exaggeration.
The scheme counts hours because hours are countable, reports them as impact, and then cannot explain to the next funder why a service delivering thousands of hours has nothing to show. Counting the input and calling it the outcome is the sector’s most common evaluation failure.
Everything else, including the mentor feedback form and the volunteer exit survey, is quality assurance rather than evaluation, useful for running the scheme and not evidence that it worked.
The honest starting position is that mentoring and befriending produce small effects and that the measurement is genuinely difficult. Meta-analysis of youth mentoring by DuBois and colleagues, published in 2011, found an overall effect of around 0.2 of a standard deviation, small but real. What moved that effect was programme practice: matching, training, supervision and match length. An evaluation that measures only the young person and never the practice will miss the variable that decides the result.
Outputs and outcomes are different things, and funders know
Outputs are what the scheme did. Outcomes are what changed for the person because of it. The distinction sounds pedantic until a report is read by somebody who assesses twenty of them a year, at which point it is the first thing they check.
| Output, what the scheme did | Outcome, what changed |
|---|---|
| Volunteers recruited, trained and checked | Person reports feeling less lonely on a repeated measure |
| Matches made and matches still running at 6 months | Young person’s school attendance or engagement changes |
| Sessions delivered, and hours of contact | Person does one thing independently they could not do before |
| Referrals received and referrals declined | Person’s contact with other people increases, or does not |
Two outputs are worth more than the rest and are usually missing: the proportion of matches still running at six and twelve months, and the proportion that ended as planned rather than abruptly. Both are cheap to collect, both describe the quality of the scheme rather than its volume, and match length is one of the practice variables the evidence says actually moves outcomes.
Distance travelled, for outcomes that have no natural unit
Distance travelled measures movement rather than arrival, which is why it suits this work. The person supported rates themselves against a small number of statements at the start of the match, again in the middle, and again at the end, and the evaluation reports the change rather than the final score. Five rules make it usable:
- Keep it to six or eight statements, written in the first person and in ordinary words.
- Ask the person to score themselves; a volunteer scoring on their behalf measures the volunteer’s optimism.
- Use the same wording every time, because a rewritten question destroys the comparison.
- Record the date of each round and the number of sessions between them.
- Report the people who got worse as well as the people who improved.
Distance travelled is not a validated instrument and should never be described as one. It is a structured, repeated self-report, and its value is that it is sensitive to the small changes a validated scale will miss over six months.
Validated scales, and choosing one rather than four
A validated scale is a set of questions that somebody else has tested for reliability, so a change in the score means something outside the scheme. Real options a UK scheme can use include the ONS4 personal wellbeing questions published by the Office for National Statistics, the Warwick-Edinburgh Mental Wellbeing Scale, the De Jong Gierveld loneliness scale and the UCLA Loneliness Scale. Check the publisher’s terms before using any of them, since some require registration and some restrict commercial use.
Pick one, use it at the start and the end, and resist adding a second. Three scales on one form produce a twelve-minute questionnaire, a lower completion rate, and a dataset with holes in it.
The measurement distinction that matters most in befriending is between loneliness and isolation. Loneliness is how a person feels about their contact with others. Isolation is how much contact they actually have. A scheme can reduce loneliness without changing isolation at all, and it is perfectly coherent to report exactly that. The evidence on befriending shows modest benefit on loneliness measures, and weaker, less consistent findings on depression and on use of health services. For context, around 6% of adults in Great Britain report feeling lonely often or always, according to the Office for National Statistics, which is a figure for the size of the problem and not a target a scheme can be measured against.
A wellbeing questionnaire comes back with a concerning answer
The volunteer, if they are the one collecting it, does not read it in silence and file it. They tell the person that they will pass it on, and they pass it to the named safeguarding lead the same day. The coordinator treats the response as a disclosure rather than as data, acts on it before it reaches a spreadsheet, and records the action taken. Any scheme using a scale that touches mood or self-harm decides in advance who reads the answers and how fast, because an evaluation instrument that surfaces risk and then sits in an inbox is a hazard the scheme created.
What a funder will actually ask for
A funder asks for six things, and a scheme that keeps them as it goes will never have to reconstruct a year in a fortnight.
- State the numbers: referrals, matches made, matches still running, sessions delivered, endings by type.
- State the outcome measure used, why it was chosen, and how many people completed it at both ends.
- Report the change, including the people for whom nothing changed.
- Give a unit cost, expressed as cost per active match per year.
- Give one or two case examples, fully anonymised and used with consent.
- Say what the scheme learned and what it changed as a result.
The unit cost is the one most schemes flinch at. It should not be flinched at: a scheme costs £700 to £2,500 per active match per year all in, and where a particular scheme sits in that band depends on how complex its referrals are, whether volunteers travel, and whether the coordinator is full-time. Naming the figure and explaining the width is far stronger than avoiding it.
This site puts no number on social return on investment, and neither should a scheme. Ratios of the “every £1 produces £X of value” kind exist throughout this sector’s literature, they are built on assumptions chosen by whoever commissioned them, and none of them survives being asked how the assumptions were set. A funder who wants one will accept a clear unit cost and an honest outcome measure instead. For the same reason, no reoffending, attainment or hospital admission reduction should be attributed to a small volunteer scheme in numbers.
The mentor or befriender
- Records sessions accurately, including the ones that did not happen.
- Lets the person supported answer for themselves rather than answering on their behalf.
- Reports a concerning answer immediately instead of returning it with the paperwork.
- Gives their own honest view of the match at review, including when it is not going well.
The person supported
- Is told what the questions are for and who will see the answers, before answering them.
- Can decline the questionnaire and keep the service.
- Is asked for separate, specific consent before any part of their story is used in a report.
- Is not identifiable in a case study, including by the combination of small details.
The scheme decides its measures before the first match, not before the first report, budgets the coordinator time to collect them, and publishes the result it got rather than the result it hoped for. It also keeps evaluation data under the same retention rules as everything else: a written period, decided and justified under UK GDPR, which sets no fixed period of its own.