Most talent reviews produce a grid and nothing else. Names move between boxes, the deck gets filed, and twelve months later the same people are discussed again in roughly the same terms.
The stakes are not academic. In DDI's Global Leadership Forecast 2025, only 20 percent of HR leaders said they had leaders ready to fill their most critical roles. Most organizations are short of successors and running a process that is supposed to find them.
A 9 box talent review earns its place when it forces a real argument about a small number of people and ends in decisions someone owns. It becomes theatre when it turns into an annual sorting exercise, rating people on a dimension nobody in the room has defined, with no action attached to any square.
What the grid was actually built for
The nine box did not begin as a people tool. McKinsey built a nine cell matrix for General Electric in the 1970s to decide where GE should put money across its business units. One axis was the attractiveness of the industry. The other was the strength of GE's position in it.
It was a capital allocation device. The people version, plotting performance against potential, came later and spread through corporate HR from the 1990s onward.
That history is worth keeping in mind. A portfolio matrix answers one question: where does the next unit of investment go? Used that way, the people version is honest. It ranks nobody's worth. It decides where development money, stretch assignments and senior attention get spent, and those are finite.
Used as a verdict on a person, it collapses. A box is not a judgment about someone's character or their value to the company. It is a statement about what the organization plans to invest in them next, and it should be revisable the moment the evidence changes.
The two axes are not the same kind of measurement
Performance is a record. Potential is a forecast. Plotting them on the same picture, with the same nine squares and the same confident borders, implies both are known to the same standard. They are not.
Even the record is shakier than it looks. Research published in the Journal of Applied Psychology in 2000 examined 4,492 managers, each rated on performance dimensions by two bosses, two peers and two subordinates. As the study is reported, 62 percent of the variance in those ratings traced to the individual rater's peculiarities of perception. Actual performance accounted for 21 percent.
So the horizontal axis carries real noise before anyone argues about the vertical one. Potential is worse, because most organizations never write down what they mean by it.
Ask five executives in a calibration room to define potential without preparation and you get five answers. Ambition, visibility, availability, tenure, and a general sense that the person reminds the speaker of themselves at that age. That is the documented weakness of the tool, and it is a weakness of execution rather than of the shape. A strong track record says less about readiness for a larger job than the grid implies. That is the whole reason a high performer and a high potential are not automatically the same person.
Define potential before anyone opens a spreadsheet
Write the definition down and circulate it before the review, not during it. One paragraph, agreed by the executive team, used verbatim in every session.
A workable definition has three parts. First, the capacity to perform in situations the person has not faced before, which shows up as learning speed rather than accumulated knowledge. Second, the appetite for a bigger and more exposed role, which is a question you have to ask the person rather than infer. Third, the willingness to be moved, including geographically and across functions, because succession plans that ignore this produce slates of people who will decline.
Whatever definition you choose, the test is whether two different leaders applying it to the same person land in the same row. If they do not, the definition is decoration.
Evidence rules: what counts and what does not
Half the value of a talent review comes from restricting what people are allowed to say. Agree the admissible evidence in advance and hold the room to it.
- Behavior in a situation the person had not handled before, described concretely.
- Outcomes with their context attached, including what was inherited and what was outside the person's control.
- Patterns across sources in structured assessment data, where the instrument was built to predict workplace behavior rather than describe communication style.
- Specific feedback from people who work for the person, not only alongside them.
Rule out the rest explicitly. Tenure is not evidence. Visibility to the chief executive is not evidence. One strong quarter is not evidence. Neither is a senior leader's enthusiasm, however genuine, when it cannot be traced to anything the person actually did.
How the calibration session actually runs
A good session is uncomfortable and short. Two hours for twenty to forty names is realistic. The sequence matters more than the software.
- Every leader submits proposed placements in writing, with evidence, seventy two hours before the room meets. Nobody arrives with an open mind that has not been written down.
- Open by reading the performance and potential definitions aloud. It takes ninety seconds and it changes the vocabulary for the next two hours.
- Skip the placements everyone already agrees on. Spend the time on the disagreements, which is where the information is.
- Require every claim in one form: what the person did, in what situation, with what result. Adjectives get challenged.
- Give every attendee explicit permission to challenge a placement, including their own boss's, and require the challenge to name the evidence it rests on.
- Move people. If nothing moved during the session, it was a briefing rather than a calibration, and you should say so out loud.
- Close by attaching an owner and one first action to every name on the grid, including the names in the boxes nobody enjoys discussing.
What has to happen in the following two weeks
The grid is worthless until it changes what someone does on a Tuesday. Three outputs should exist within ten working days, and if they do not, the review did not happen.
The first is a development action per person that involves real work rather than a course. A stretch assignment, a board exposure, a piece of the business to run. The second is a successor slate for each critical role, with named people and honest readiness timeframes, including the roles where the honest answer is that there is nobody.
The third is a set of conversations. Most organizations do not tell people their box, and that is defensible: a label is not useful feedback and creates a caste system fast. What is not defensible is telling someone nothing. The person should hear what the organization is investing in them and why, in plain terms, even when the answer is that the investment is going elsewhere this year.
The hard part is not the grid, it is the meeting
Most articles about this tool argue about the squares. The squares are fine. The difficulty lives in the social dynamics of the room, which nobody designs for.
The most senior voice anchors everything
Whoever speaks first on a name sets the range the discussion happens inside. If that is the chief executive, the rest of the room calibrates to their opinion rather than to the evidence. Written pre-submission is the only reliable defense, because it captures independent judgments before they can be influenced.
Potential quietly becomes proximity
People who sit near power get rated as higher potential because their work is observed and narrated by people with authority. Leaders of remote teams, back office functions and overseas units lose systematically. Check the demographics and the locations of your top row before you sign the grid off.
The bottom left box has no owner
Every organization has people in the low performance, low potential square, and almost every review reaches that box at the end when everyone is tired. Decisions get deferred to a manager who has already avoided them for a year. Run that box first, at full attention, or accept that it will still be there next cycle.
Nobody revisits a placement
A box assigned in March is treated as fact in November, long after the person has had a different year. Without a mid cycle check, the grid stops being a live assessment and becomes a filing system for old opinions.
Given the choice, fix the definition before you fix the grid
If one thing could change before your next cycle, make it the written definition of potential rather than the tooling, the template or the number of boxes. The definition is what determines whether the vertical axis carries information, and everything downstream, including succession slates and development spend, inherits its quality.
It is also the cheapest change available. It costs one executive team conversation and a paragraph. Buying a new talent platform costs considerably more and will faithfully record the same undefined judgments.
If you want a sharper read on whether your potential ratings are measuring anything, our leadership assessment work is built to put evidence behind exactly that judgment. A short conversation will tell you whether your current process needs repair or replacement.
Frequently Asked Questions (FAQs)
How often should a 9 box talent review be run?
Once a year for the full population, with a lighter mid cycle check on the names where something material changed. Running it quarterly produces churn rather than insight, because potential does not move that fast. Running it less than annually means the grid is describing a workforce that no longer exists.
Should employees be told which box they are in?
Most organizations do not share the box itself, and that is a reasonable position. A label is not actionable feedback and tends to become a permanent identity. What should always be shared is the substance: what the organization sees, what it is investing in, and what would have to change for that to be different.
What are the main disadvantages of the 9 box grid?
Subjectivity on the potential axis is the largest, followed by the tendency to treat a placement as a permanent label rather than a current investment decision. The grid also flattens people into two numbers, which hides the specific strength or derailer that actually matters. None of these are fixed by changing the tool.
Who should be in the calibration meeting?
The leaders who directly manage the people being discussed, their common boss, and one facilitator with the standing to challenge anybody in the room. Keep it small enough that everyone speaks. If an attendee has no first hand knowledge of any name on the list, they are an audience, not a calibrator.
Should the 9 box grid drive compensation decisions?
No. Performance already drives compensation through the performance process, and adding potential to the pay equation means paying people for a forecast. It also gives every participant a financial incentive to inflate placements, which destroys the honesty the session depends on.
What are the alternatives to the 9 box grid?
A simple successor slate per critical role, with readiness timeframes and named development actions, does most of the useful work with less ceremony. Some organizations use a skills or capability inventory instead, which answers a different question about what the company can currently do. The grid remains useful mainly as a forcing device for a conversation that otherwise never happens.
