top of page

Your Best Feedback Number is Probably Lying to You

  • Writer: Steve Barbour
    Steve Barbour
  • Jul 7
  • 6 min read


Almost every training session ends with a feedback form, and almost every one of them is wasted. Not because feedback has no value, but because of what we ask for. Did you enjoy it. Was the room comfortable. Rate the facilitator out of five. We collect a mood, file it, and learn nothing we could act on. In the trade these are called happy sheets, and the name is the whole criticism. A happy sheet can only ever return a compliment, which means it can never tell you what to change.


I run human factors sessions, and human factors has a low tolerance for measuring the wrong thing. The entire discipline exists because comfortable assumptions get people hurt, and the cure is honest measurement followed by honest change. So when I built the feedback survey for my own delivery, I tried to make it ask questions that could actually come back and tell me something I did not want to hear. Four sessions and twenty responses later, it did exactly that. It also handed me a striking-looking statistic that, on closer inspection, was lying to me a little. Both of those turned out to be useful. Here is the full version, including the parts I would rather skip, because the skipping is where most people reading their own feedback go wrong.


Measure the change, not the mood


The single most useful move on any feedback form is to ask the same question twice. Before the session: how would you rate your understanding of this subject, one to ten. After it: the same question, again. The difference between the two is the only figure on the page that even attempts to measure learning rather than enjoyment. Everything else tells you whether people had a nice morning. This tells you whether the morning did anything.


Line chart showing mean self-rated understanding rising from 7.30 before the session to 8.65 after, a gain of 1.35 points.

Across my year of sessions, average self-rated understanding moved from 7.3 to 8.65 out of ten. A mean gain of 1.35 points. On its own that is a modest, slightly forgettable number, the kind you put on a slide and move past. The interesting thing was not the average. It was the shape underneath it.



The trend that emerged


When I plotted each person's starting score against how much they then gained, a clear pattern appeared. The people who rated their understanding lowest at the start gained the most. The people who started high barely moved. One attendee went from 4 out of ten to 9. Several who arrived at 8 or 9 gained nothing, or a single point. The relationship was strongly inverse, and if you want the technical figure, the correlation was minus 0.89, which is about as strong as messy real-world data ever produces.


That pattern, if it is real, is the hard thing to achieve in a room. Every group you will ever teach is mixed. Some people know the subject cold. Some are meeting it properly for the first time. Most training quietly picks one of those groups to serve and abandons the other. Pitch high and the novices drown while the experts nod along. Pitch low and the novices are reassured while the experienced people check their phones. A session that lifts the people who most need lifting, without losing the people who already know the material, is doing the genuinely difficult thing, and the format is why. When the learning comes out of the room rather than off a slide, every person meets the material at their own level automatically, because they are working from their own experience and not from my pitch. The novice builds a concept from scratch. The expert reframes something they thought they already understood. Nobody is aimed at the wrong height, because nobody is being aimed at.


That is the headline I would love to stop on. Discussion beats presentation, and here is a big number to prove it. But stopping there would make me guilty of exactly what I accused the happy sheet of: presenting a comforting figure as if it settled the question.


The stats, told honestly


Start with the weaknesses, because they are real and they matter more than the headline.

First, the sample is small and self-selected. Twenty responses across four cohorts is not nothing, but it is not much, and the people who fill in a voluntary feedback form are not a random slice of the room. Anyone who disliked a session is less likely to bother, which quietly inflates every positive figure in the dataset. The scores you are reading are the views of people willing to respond, not the whole group.


Second, understanding was self-rated, not tested. Nobody sat an exam. They told me how well they felt they understood the subject. That measures confidence, not competence, and the two are not the same thing. It is entirely possible to leave a good session feeling clearer while being just as wrong as when you arrived, only now more sure of it. Self-rated learning is a soft measure, and I would be over claiming to call it anything firmer.


Scatter plot of each attendee's understanding before the session against points gained. A downward orange trend line (correlation minus 0.89) shows lower starters gaining more, with a dashed ceiling line showing the maximum possible gain at each starting score.

Third, and this is the important one, that eye-catching minus 0.89 is partly an artefact of how it is built. The gain is calculated from the starting score, so the starting score sits inside both numbers being compared. And the scale stops at ten. Someone who starts at 9 physically cannot gain more than 1 point. Someone who starts at 4 has room to gain 6. So high starters are forced toward small gains and low starters are free to record large ones, regardless of how good or bad the teaching was. Statisticians call this regression to the mean, or a ceiling effect, and it guarantees a chunk of negative correlation for free. If I had taught the sessions badly, or not at all, the figure would still have come out negative. So the honest statement is not "minus 0.89 proves the session closes the gap." A meaningful part of that number was always going to be there.


Square the correlation, by the way, and you get roughly 0.79, which a naive reading would translate as "starting level explains 79 per cent of the variation in how much people gained." Treat that with the same suspicion, because the same built-in coupling inflates it too.


So what survives all that. More than you might expect. The built-in effect explains why the correlation is negative, but it does not explain the size of the gains at the bottom end, and it does not explain the words. The written comments, which are not subject to the ceiling problem at all, say the same thing the numbers do, in the attendees' own language. They valued the session because it was interactive rather than slide-driven, discussion-based, an open forum, more engaging than a presentation, a chance to hear other people's experience. That qualitative chorus is independent corroboration. And the pattern holds across four separate groups and a year of delivery, not a single lucky cohort. When a soft number, a structural caveat, and twenty unprompted comments all point the same way, the trend is real even though the headline figure overstates it. The correct reading is calm rather than triumphant: the data is consistent with discussion-led delivery serving a mixed-ability room well, the effect is genuine, and it is smaller than minus 0.89 makes it look. That sentence is less exciting than the big number. It is also true, which is the only thing that should matter.


The number I did not want to see



Horizontal bar chart of mean satisfaction across seven aspects, all between 4.45 and 5.0. The facilitator scored 5.0; use of technology was lowest at 4.45, highlighted in orange.

The same survey handed me something less flattering. My weakest measured area was use of technology, at 4.45 out of five, with three people dropping it to neutral. The written comments named the cause without ceremony: a bit of IT faff at the start, a video that ran slightly long. Small stuff, easy to wave away as the room's fault or a one-off.


Except an independent assessor had, separately, flagged exactly the same thing. Two unrelated sources pointing at one cause is not noise. It is a pattern, and a pattern is an instruction. So it changed. There is now a fixed pre-session protocol: kit confirmed and compatible, video cued, room tested, before a single person sits down. Five minutes of preparation that removes the one recurring friction my own data kept naming. That is the entire value of the exercise in a single move. Not the praise, which is pleasant and changes nothing, but the one uncomfortable number that tells you precisely what to fix.


The form is a debrief tool


None of this is really a marketing exercise. It is the debrief discipline that aviation runs on, turned back on my own delivery. You fly the sortie, you measure honestly, you resist the urge to believe your own good numbers, you find the one thing that did not work, and you change it before the next one. The feedback form is not a happy sheet. It is a debrief, and most people simply never ask it the questions a debrief would.


So before your next session, look at the form you hand out. If every question on it could only ever produce a compliment, it is measuring nothing. Ask people what they understood, not whether they were comfortable, and ask it twice. Then, when the numbers come back, distrust the flattering ones long enough to find out which part of them is real. And when the awkward number turns up, which it will, treat it as the most useful line on the page.

Comments


bottom of page