ESLA Loops branded hero graphic in navy and coral with a concentric loop motif.

Why I’ve grown wary of rubrics — and what I do about it

A few years ago I helped design and operationalise the results and reporting system for a sustainability initiative that spanned a dozen countries and focused on the built environment. The funder’s chosen progress and results reporting tool was a rubric: criteria down one side, performance levels across the top — from “harmful” at one end to “thriving” at the other — agreed with partners in advance. Green, amber, red. Everyone would know where they stood.

It did exactly what it was designed to do. And that turned out to be the problem.

The results that mattered most that year were the ones the rubric had no cell for. One national team, in response to an emergent window of opportunity, had quietly rewired how a whole coalition made decisions. Another had abandoned its original plan — correctly — and found a better route nobody had foreseen. On the rubric, both looked like underperformance. The tool couldn’t see the very thing we most needed to learn from.

I still use rubrics — to help organisations ‘map’ their portfolios and reflect on elements of strategy. Used well, they’re genuinely valuable: they make implicit judgements explicit, they help translate the results and knowledge of a diverse group into deeper portfolio-level strategic sense-making, and they guard against the bias of whoever happens to have the loudest voice in the room. When evaluators like Jane Davidson argue that rubrics democratise evaluative reasoning — dragging it out of the expert’s head and onto the table where everyone can see and challenge it — they’re right. For questions where we already know what “good” looks like, a rubric is a fine instrument.

My worry is really about what happens across the life of a rubric — from the moment it’s handed down to the moment it’s used to make a call. The change these initiatives are reaching for is uncertain, complex and deeply human: contested systems that shift as you act on them, where progress rarely moves in a straight line and the outcomes that matter can’t be fully known in advance. And for work of that kind, I’ve come to believe rubrics are broadly unsuitable. By their very nature they don’t help us explore, understand and adapt to the challenge in front of us — they ask us to fix in advance what can only be learned by engaging with it. Once you take that seriously, rolling a rubric out across a multi-country initiative resolves into four distinct stages, each with its own pitfalls and its own fixes — and underneath all four sits a fifth that runs deeper than the rest.

Five stages in the life of a rubric — the pitfall at each stage and the fix — with a fifth theme, space for emergence, underpinning them all.
The five stages of a rubric’s life — the pitfall at each, and the fix.

Stage one — Onboarding: what is this thing, and what is it for?

A rubric usually arrives the way ours did: passed down from the funder to the grantees and partners expected to use it, with an introductory call and little else — no step-by-step guidance on how to adapt it to a particular context, and, more tellingly, no clear account of how the funder itself intends to use the results. Is this headline reporting? An informal assessment of each partner? Internal sense-making? An input to future funding? Often nobody can say. That ambiguity isn’t harmless: it stalls operationalisation for months and lets anxious rumours fill the gap. It also wastes partners’ time, creating yet another format for them to report against while generating very little value for them in return.

What helps is unglamorous but decisive — a proper operational guidance document, worked use cases (how a score gets used in a grantee conversation, how it’s read across a portfolio), and honest discussion of the headline questions up front: what is this for, how much effort is expected of partners, what do we get out of it, and who do we call when we have questions?

Stage two — Translation: making a generic grid fit a real context

A rubric written to travel across a whole portfolio arrives generic and multi-clause: each performance level bundles several sub-clauses together, some relevant to a given team, some not at all. Teams are then asked to rate — and be held accountable against — things that were never part of their design and are not relevant to their context. That erodes trust fast.

Why it happens: the criteria are drafted once, centrally, to fit everywhere, so they fit nowhere well. What helps: allow partners to tailor the rubric to their own context and real strategy, strip out what genuinely doesn’t apply, and assess at that finer resolution — rather than collapsing everything into one blurry rating that no longer means anything specific.

Stage three — Assessment and reporting: the tyranny of the single score

Then comes the pull towards one comparable number or rating — a single rating for the whole initiative, so the portfolio can be lined up side by side. Or, even worse, used to compare and contrast different funder initiatives in a simple grid format. But collapse a dozen distinct contexts, each with its own strategy and contextual story, into one grade and you get something tidy, comparable, and close to meaningless — and useless for the course correction the teams actually need.

Why it happens: the rubric is being asked to serve the centre’s oversight and upward reporting rather than any team’s learning — often for people who are too busy, and too important, to engage with detailed complexity. I lost count of the number of times I offered the response “it’s more complicated than that,” only to be told to translate the complexity into a dashboard or a RAG rating. What helps: treat each local context as the real unit of insight and build the tools and process to support consistent, honest reporting at that level; if you must roll up, be explicit about the method — and about the detail the single number hides.

Stage four — Value and use: incentives, and who the rubric is really for

This is the stage where I’ve changed my mind the most. Hand a comparative rubric to grantees and partners for upward reporting and you have, in effect, built an incentive to perform. Teams read it — rightly — as a competitive exercise: their scores will sit next to their neighbours’, and may feed funding decisions. So a rational pattern sets in. Each reporting round, everyone edges one cell to the right; nobody wants to be the partner reporting slow progress, a stubborn constraint, or a plan that isn’t working — even though those are precisely the things a funder most needs to hear. You get steadily rising scores and steadily thinner truth. The real challenge here is aligning the incentives.

The usual responses treat the symptom: ask for the rationale and evidence behind each score, bring in an independent critical friend to moderate, insist on honesty. All worth doing. But the deeper fix is to change who holds the rubric. Much of the time it doesn’t belong in the grantee’s hands at all. The value of rubrics reporting is captured at the organisational level — it’s the funder or host that wants the portfolio picture — so that is where the ‘aggregation’ should sit, aligned to the funder’s own theory of change and outcomes, rather than pushed down as a compliance task onto partners who carry the burden and see little of the benefit.

That doesn’t mean cutting partners out; it means changing what you ask of them. Keep their reporting light and human: conversation-based reflection on what’s actually happening in their context. The organisation then draws on those conversations to inform its own rubrics-level view — so the data that feeds the rubric still comes from the ground, but the comparative scoring, and the perverse incentives that ride along with it, never land on the grantee. What happened? So what? Now what? On every initiative I’ve worked on, that conversation — not the score — turned out to be by far the most valued part, precisely because it wasn’t a test.

Stage five — The one underneath: leaving space for emergence

Here’s the problem that persists even when the first four stages go well. A rubric, by design, fixes attention on its cells. Teams report against what’s scored; conversations orbit the pre-agreed criteria; and the unplanned adaptation — the course correction, the surprising second-order win — has nowhere to be recorded and quietly stops being looked for. In complex, contested change, that emergent outcome is very often the result. Applied rigidly, a rubric acts like a fixed rule where the work needed a loose guide: it constricts and restricts the very space where the most important change tends to appear. Dave Snowden’s Cynefin framework calls this the complex domain — “the domain of emergence” — precisely because you cannot specify the good things in advance; the system hasn’t produced them yet.

The fix here turns the rubric’s usual role on its head. Instead of a grid you score against, treat it as a shared frame you narrate against. The real gift of a rubric in complex work isn’t the rating — it’s a common language. Give every team the same underlying framing — the same dimensions, the same sense of the direction of travel — and then ask them not to slot themselves into a cell but to tell the story of how change actually emerged in their context, in those shared terms. Because the framing is held in common, those stories become comparable and connectable across an initiative; because the content belongs to each team, every one keeps the contextual dynamics a single score would have flattened. That’s consistency of framing with freedom of content — a way of reflecting on emergence that unifies a programme without pretending every context is the same. It’s the version of a rubric I think is genuinely worth keeping: not a gavel that ends the conversation, but a shared frame that makes a richer one possible across every context at once.

I’ve come to think the deepest risk of a rubric isn’t that it gives the wrong answer. It’s that it ends the conversation too early, narrows everyone’s field of view to the cells on the grid, and convinces the room that the map was the territory all along.

That’s why I’ve grown wary of them. Not because they’re wrong — but because we so often reach for them in exactly the situations where, applied poorly, they close down the very learning we needed most.


Where have you seen rubrics genuinely earn their place — and where have you watched one close down the learning you needed? And if you’ve used a rubric as a shared frame for surfacing emergent change across very different contexts, I’d love to compare notes.

Robbie Gregorowski
Robbie Gregorowski
Articles: 4

Leave a Reply

Your email address will not be published. Required fields are marked *