The Headquarters Counseling Center ArchiveLawrence, Kansas · 1969–2020

The Headquarters Counseling Center Archive

Measuring whether a crisis line works

How one crisis centre measured its own effect: an eleven-variable caller-rated instrument administered three times a year, its design, its 2010 results, and what it can and cannot show.

Overhead view of hand-ruled paper tally sheets, index cards, a pencil and a coffee mug on a grey desk
Three two-week administrations a year, scored by hand and entered into a database.

The measurement problem

A crisis line cannot prove what it prevents. The person who does not die is indistinguishable from the person who was never going to. The service keeps no identifying information, so nobody can be followed up. Sample sizes for the outcome that matters most are, mercifully, tiny. And the intervention is a conversation, which resists standardisation.

This is not a small methodological inconvenience. It is why crisis lines spent decades being funded on faith and testimonial, and why they are structurally vulnerable whenever a funder asks what the money bought.

From 1999 the centre took a specific position on this: measure the thing you can actually observe, which is the change in the caller between the beginning of a call and its end, reported by the caller, before they hang up and vanish.

How the instrument was built

The centre's account is unusually explicit about provenance. Its evaluation process developed from 1999, after completing training in United Way's Measuring Program Outcomes: A Practical Approach and reviewing evaluations of counselling and crisis programmes conducted over the previous twenty-five years. It stated that its model was in line with two national studies on the impact of hotlines funded by the federal Substance Abuse and Mental Health Services Administration.

That is the right way to build a small-organisation instrument: take an established outcomes framework, check it against the existing literature, and align it with whatever national work exists so the results are comparable to something.

The eleven variables

Adult callers were asked to respond to eleven statements on a one-to-five scale, from strongly disagree to strongly agree, with a not-applicable option. Most statements opened with a stem of the form after talking with the counsellor, I feel…

  1. more calm
  2. less alone
  3. more hopeful
  4. gained useful knowledge about the concern
  5. gained information about available resources that they will use
  6. was more prepared to manage the concern
  7. more likely to take actions for safety
  8. perceived the counsellor as knowledgeable
  9. perceived the counsellor as understanding the concern
  10. perceived the counsellor as caring
  11. believed talking was helpful

The set divides neatly into three groups, and the division is deliberate. Items one to three are affective state. Items four to seven are capability and action — knowledge, resources, preparedness, and the explicit safety item. Items eight to eleven are the caller's perception of the counsellor, which is the closest a crisis line can get to a quality-of-delivery measure.

Item seven is the one that carries the weight. More likely to take actions for safety is the only variable that reaches toward the outcome anyone actually cares about, and it is still a self-report of intention at the end of a phone call. The centre did not claim otherwise.

How it was administered

The survey ran during three two-week periods each year rather than continuously. After a call reached a point of closure, the counsellor asked whether the person was willing to help improve the service by answering eleven questions. Responses were entered into a database, and the analysis produced mean ratings per item plus information about which factors influenced overall helpfulness. The centre set itself a benchmark of means above 4.0 on every item.

The reporting was refreshingly complete about attrition, and this is what makes the published tables worth reading. For each administration the centre reported total calls, how many were eligible, how many surveys were offered, the percentage offered, how many were completed, and the completed surveys as a percentage of eligible calls. Across the three 2010 administrations, surveys were offered on roughly 39 to 45 per cent of eligible calls, and completed on roughly 27 to 32 per cent of them.

What it found

Across the three 2010 administrations every published mean cleared the 4.0 benchmark, and the shape of the results is consistent enough to be interesting.

Mean caller ratings, three 2010 administrations (1–5 scale)
VariableMarchJulyOctober
More calm4.534.554.58
Less alone4.544.234.40
More hopeful4.464.034.46
Gained useful information4.514.244.49
Gained referrals they will use5.004.444.85
More prepared4.404.354.36
Increased safety4.574.424.77
Counsellor knowledgeable4.764.654.71
Counsellor understanding4.884.804.83
Counsellor caring4.894.904.91
Talking was helpful4.824.794.77

The perception items — knowledgeable, understanding, caring, helpful — sit consistently at the top, with caring the highest variable in all three administrations. The capability items sit lower, with more prepared the weakest item in every administration. Hope is the most volatile, dropping to 4.03 in July before recovering.

The honest reading is that callers reliably experienced the counsellor as caring and the conversation as helpful, and less reliably came away feeling equipped to manage the thing that was wrong. That is exactly the pattern you would predict for a service whose intervention is a single conversation, and it is a more useful finding than a uniformly glowing one would have been.

What the instrument cannot show

  • It measures the end of a call, not a life. Nothing here says anything about what happened afterwards, and the design makes it impossible to find out.
  • It is self-report, collected by the person delivering the service. A caller who has just been helped by a stranger is not a neutral rater, and the counsellor asking is not a neutral instrument.
  • The eligible-call filter is doing unexamined work. Fewer than half of eligible calls were offered a survey, and under a third were completed. Which calls a counsellor judged eligible, and when they judged the moment right to ask, is not in the published record — and there is an obvious mechanism by which the hardest calls are least likely to end in a survey.
  • The benchmark was set by the organisation being measured. A 4.0 threshold that is consistently surpassed is a floor the organisation chose.

None of that makes the exercise worthless. It makes it what it is: a small organisation, with no research budget, publishing its methodology and its attrition alongside its results, in a field where most of its peers published neither. Twenty-five years later that is still more transparency than a great many services offer, and it is the main reason this particular page exists.