Coaching is one of the few senior investments routinely bought without an agreed measure of success. Harvard Business Review asked 140 leading coaches how they report progress. Seventy percent give a qualitative assessment. Fewer than a third supply quantitative data on behavior, and less than a fourth provide any quantitative data on business outcomes.
That was 2009, and the pattern has not moved much. The same report noted that companies will not get formal progress reviews unless they ask for them. Most sponsors never ask, then find themselves unable to answer a finance question that arrives eleven months later.
The practical problem is not that coaching resists measurement. It is that measurement gets designed at the end, when the baseline is gone and the only evidence left is a satisfied executive. That evidence is real. It just does not survive a budget review.
The numbers everyone quotes, and where they come from
Two figures dominate every article on this subject. Both were published in 2001 and neither has been reproduced at that scale since.
The first is a study by Manchester Inc, presented at the ICF annual conference in August 2001. It surveyed 100 executives at Fortune 1000 companies, half of them vice president or above, who had been coached for six months to a year. It reported a return of six times the cost of coaching, with self-reported improvements in productivity (53 percent) and executive retention (32 percent).
The second is the MetrixGlobal case study by Merrill Anderson, dated November 2001, source of the 529 percent figure that circulates without context. Read the briefing and the sample is 43 leadership development participants at a single company, of whom 30 returned a questionnaire. They were drawn mostly from middle management, not the C-suite.
Neither study is fraudulent. Both are self-reported questionnaires asking participants who chose coaching to estimate its financial value afterward. Quote them to a chief financial officer and the sample size, the date and the method will all come back at you.
What the rigorous evidence actually shows
There is better evidence, and it is quieter. A 2023 meta-analysis in Academy of Management Learning and Education restricted itself to 37 randomized controlled trials covering 2,528 participants between 1994 and 2021. Its best estimate of the effect was g = 0.59, described as well within the moderate range.
Three findings inside that paper matter more to a buyer than the headline. The authors reported indications of significant publication bias. Effects were larger when outcomes were self-reported than when they were observed. And coaching duration had minimal impact on outcomes.
Read together, that is a defensible position: coaching works, at a moderate and real effect size, and the measured size shrinks when someone other than the participant does the measuring. A business case built on that is stronger than one built on 529 percent, because it does not collapse when examined.
Decide whose return you are measuring
The organization’s return
The sponsor is buying a change in how a leader operates. The value lands in retention, in decision speed, in team performance, or in a transition that does not fail. These are the measures a finance function recognizes.
The individual’s return
The leader is buying capability, standing and options. Promotion, scope, compensation and the quality of their own working life are all legitimate returns. They are also invisible on a company scorecard.
Most measurement arguments are really a confusion between these two. Decide at contracting which one the engagement is accountable for, and say so out loud. An engagement paid for by the employer but measured on the leader’s career outcomes will satisfy nobody.
A four part measurement design
1. Name the change in observable terms
Not “improve executive presence”. Something a third party could witness: runs a decision meeting to a conclusion, gives direct corrective feedback within a week, stops taking back delegated work. If you cannot describe what an observer would see, you cannot measure it.
2. Take the baseline before the first session
This is the step that gets skipped and it is the one that cannot be recovered. A short structured interview with four or five stakeholders, run before coaching starts, takes a few hours and becomes the only before-picture you will ever have. A formal instrument works too, and the difference between 360 feedback and a performance review is worth understanding before you pick one.
3. Re-measure the same thing, the same way
Same stakeholders, same questions, same scale, at the midpoint and at the close. Consistency is what makes the comparison mean anything. Changing the instrument halfway converts your evidence into an anecdote.
4. Attach one business measure, and only one
Pick the measure the change should move if the theory is right. Voluntary attrition on that leader’s team. Cycle time on a decision their function owns. The survival of a specific transition. One measure, agreed in advance, is far more persuasive than nine measured afterward.
The attribution problem, and the honest way to handle it
A coached leader is rarely the only thing that changed. They were also promoted, or the reorganization landed, or the market turned. Any competent finance partner will raise this, and a business case that pretends otherwise loses credibility on the spot.
The workable answer is a contribution estimate rather than a causal claim. Ask each stakeholder what share of the observed change they attribute to the coaching, discount the result, and show your working. A figure presented with its own confidence interval is treated as analysis. A clean 529 percent is treated as marketing.
State the counterfactual too. If the leader had not been coached and had left, what would replacement have cost in fees, lost time and disruption? That comparison is often more decisive than the upside.
What to compare coaching against
Coaching is almost always evaluated in isolation, which flatters it. The real decision is against the alternatives for the same money: a training program, a role redesign, a stronger manager above the leader, or replacing the person.
Each of those has a different cost, a different lead time and a different failure mode. Coaching competes well when the capability is present and the behavior is the constraint. It competes badly when the problem is structural, and no measurement design rescues an engagement aimed at the wrong thing. Deciding when to hire an executive coach is the first half of the return.
Why an honest coaching number is hard to produce
- The baseline has to exist before anyone is sure the engagement will happen. Measurement work has to be paid for and scheduled at the least certain moment, which is exactly when it gets dropped.
- The person best placed to report change has the strongest reason to report it favorably. The controlled trials show the effect shrinking when observers rather than participants do the rating, and that gap is structural, not a flaw you can instruct away.
- Confidentiality limits what evidence can travel. The content of sessions is closed to the sponsor, so progress has to be evidenced through observed behavior rather than through the conversation that produced it.
- The useful outcomes appear after the invoice. A succession that holds or a team that stops losing people shows up over years, while the budget cycle asks in months.
- Nobody wants to publish the engagements that failed. The literature is built almost entirely on participants who chose coaching and completed it. The base rate of no change is therefore unknown, and every published average is flattered by it.
What changes in your next sponsor conversation
If you act on one thing here, make it the baseline. Before the next engagement starts, book four stakeholder conversations and write down what each person currently observes. That is a single afternoon, and it converts every later claim from opinion into comparison.
Your next sponsor meeting then sounds different. Instead of asking whether the leader is finding it useful, you are looking at the same five questions asked twice, and at one business measure agreed in advance. The conversation stops being about satisfaction and starts being about evidence, which is also the point at which a coach who is not working becomes visible.
We build the measure into the engagement rather than bolting it on afterward, because proving the work is part of the work. That is what the last step of our method covers in executive coaching.
Frequently Asked Questions (FAQs)
What ROI should we realistically expect from executive coaching?
Expect a moderate, real effect on leader behavior rather than a specific multiple. The controlled trial evidence supports a moderate effect size across leadership and personal outcomes. Any provider quoting you a precise percentage return in advance is quoting a 2001 questionnaire, not your organization.
How long before results are measurable?
Observable behavior change usually registers with stakeholders inside three to four months, which is why a midpoint re-measure is worth scheduling. Business measures attached to that behavior move later and less cleanly. Set the expectation that the first credible read is behavioral, not financial.
Can we measure coaching without breaching confidentiality?
Yes, and the distinction is simple. The sponsor sees observed behavior, agreed goals and progress against them. The sponsor does not see session content. Stakeholder ratings and business measures both sit entirely outside the confidential conversation.
Is group coaching a better return than one to one?
On unit cost, almost always. On depth for a single leader with a specific constraint, no. The right comparison is not price per head but whether the format can move the particular thing you need moved, and for an individual behavioral constraint it usually cannot.
What if the leader does not improve?
Decide in advance what that looks like and what happens next. A midpoint re-measure exists partly to surface a failing engagement while there is still time to change the coach, change the goal, or stop. An engagement with no stopping rule will run to its end date regardless of effect.
