Can you measure feels?

Part 2 of unpacking a recent project…

The first real design decision on this project wasn't a metric. It was a filter for deciding which metrics were really metrics.

Workforce data is enormous and mostly useless for decision-making. You can pull dozens of numbers about any given role and find hundreds of job titles, skills, requirements, etc. There is no consistency and every organization is using different words to say the same thing. One could easily drown in the quest for clarity. When trying to ‘respond’ to this demand data it produces a very special kind of paralysis.

So before building anything, I set a rule: a metric only earns a place as decision support if it's both measurable — meaning you can pull it consistently, in a way that's comparable — and controllable, meaning someone can actually act on it or it is attributable to a preceding action. Something measurable but not controllable is fun trivia or headlines. A metric that's controllable but not consistently measurable is just an opinion wearing a number's clothes. The overlap of the two is the only place real management decisions can live.

That filter did a lot of work. It knocked out plenty of data points that sounded important but didn't pass the test — activity counting, variance dependent on who was reporting, what no one could influence even if they wanted to, and most importantly, what was readily available at the start. This client needed to build on top of systems that already existed — the learning management system, the registration system, financial reports, a labor-market analytics platform they were already paying for — rather than asking anyone to invent new reporting infrastructure.

What survived the filter was a small, deliberately short list. Roughly six or seven measures, each tied to a plain-language definition and a named source, covering distinct territory: how often an offering was cancelled for low uptake; whether staffing and administrative effort was proportional to what a program returned; how reliably a program could actually be delivered; whether capacity was full, empty, or had unmet demand behind it; how what was offered compared to actual job postings and hiring demand in the region; whether there was potential for repeat engagement; and, plainly, whether there was a positive net return.

Each of those became a defined metric, calculated the same way every time, from a source everyone agreed on in advance — no debating methodology in the middle of a review. Each got normed onto a 1-to-5 scale, so a labor-market ratio and a cancellation percentage could sit side by side without pretending they were the same kind of thing.

Just as important as what made the list is what we explicitly ruled out, in writing, before any metric was named. This was never going to be a performance review of individual staff. It wasn't a financial audit. It wasn't a one-time compliance exercise to satisfy someone upstream. And it would never be the sole input into strategy — it informs judgment, it doesn't replace it. Naming those boundaries up front mattered more than I expected. It's what let people trust the numbers enough to look at them honestly.

There was a second, quieter benefit to building it this way. Because the labor-market comparison used a platform the college already had, running the same report on the same programs every year turns into a baseline you can actually watch move. Is demand for this skill area growing or shrinking compared to last year? Is a program's job-market fit improving? For the first time, that question had a repeatable answer instead of a gut feeling revisited every few years when someone happened to wonder.

That's the measurement half of the problem, and on its own it's still incomplete. Because the moment you hand a group of very different programs the same scorecard, someone reasonably asks: are we comparing things that shouldn't be compared? That question — and why the honest answer is yes, deliberately — is where the next post picks up.

Next
Next

Never Assume