Seven questions to ask before you bank the AI savings

For operating partners and boards

Most portfolio companies have sized a team against an assumed productivity gain. Very few have measured one. These are seven questions that find out which you have, and what the gain is being paid for.

None of this is an argument against adopting the tools. The companies reporting real, measured gains without a reliability cost appear to be the ones that kept mandatory human review in place. The ones that dropped review to go faster tend to get neither, because review becomes the bottleneck anyway. The questions below are aimed at finding out which of the two you are running.

Every question asks for a specific event or a specific fact. None of them can be answered from a dashboard, and none of them can be answered honestly in the abstract. That is the design: a question that can be met with a view will be met with a view.

Ask for an instance, not an assessment. "Are we managing this well?" produces a reassuring answer from a competent executive every time. "When did that last happen, and what did we do?" produces either a story or a silence, and both are information.

Question one · The gain

We sized this team against a productivity gain. Who measured it, against what baseline, and when?

This is the ROI question and it comes first because everything else rests on it. A headcount decision made against an assumed gain is an unpriced bet. The honest version of the answer is usually that the tools obviously help, the team felt faster, and nobody ran a before-and-after. That is not a failure of management — almost nobody has a baseline — but it does mean the savings on the page are an estimate wearing the clothes of a measurement.

What a managed answer sounds like

"Everyone's shipping more. You can see it in the velocity numbers." Velocity measures output, not the counterfactual. The question is what the same team would have produced without the tools, and against what quality bar.

Question two · The dependency

If these tools were unavailable for a week, which functions slow down and which ones stop?

This separates augmentation from dependency, and the two look identical on a P&L. Slowing down is what a tool that helps looks like when it is removed. Stopping is what a tool that has replaced a capability looks like. Ask for the list by function and watch which ones the CEO has to think about.

What a managed answer sounds like

"We'd manage. People worked without these tools two years ago." Two years ago the people who knew how were still doing the work. The question is about now.

Question three · The formation

Of the work we handed to tools, how much of it was how somebody learned this job?

The junior work was rarely valuable as output. It was the apprenticeship, and it was cheap precisely because it was also training. When it moves to a tool, the output is preserved and the training is not, and nothing in the reporting distinguishes the two. This is the question that matters most over a hold period and the one least likely to have been asked.

What a managed answer sounds like

"We've freed our junior people up for higher-value work." Sometimes true. Ask what the higher-value work is, who is teaching it, and how someone gets to be good at it without the years of the lower-value work underneath.

Question four · The check

When did someone here last catch one of these tools being wrong, and what happened next?

If nobody can name an instance, there are two possible explanations and only one of them is good. The follow-up matters as much as the answer: whether anything changed in how the work gets reviewed, or whether it was fixed quietly and filed as a one-off. Fluent output suppresses the instinct to verify, which means verification erodes without anyone deciding to stop.

What a managed answer sounds like

"We always have a human in the loop." A human in the loop who has never found anything is not a check. They are a signature.

Question five · The concentration

Which decisions now rest on output that nobody in this building could produce or fully evaluate themselves?

The risk is not that a tool is used. It is that it is used where in-house judgment can no longer assess the result. That gap is invisible while the output is right and total when it is wrong. Ask specifically about pricing, forecasting, engineering specifications, and anything going to a customer or a regulator.

What a managed answer sounds like

"Nothing critical." Then ask which decisions in the last quarter used an analysis nobody re-derived. The list is usually longer than the first answer implies.

Question six · The metric

What are we measuring about AI use, and what behavior is that measurement producing?

Many companies now measure adoption directly: token spend per employee, a target share of code that must be machine-written, internal leaderboards. Every one of those measures consumption rather than result, and people optimize what is measured. The predictable effect is that the most careful people look like the worst performers, because care shows up as lower usage. Ask what happens to someone whose numbers are low and whose work is good.

What a managed answer sounds like

"We track it to understand adoption." Then ask whether anyone has been spoken to about their numbers. Tracking that carries a consequence is a target, and a target on consumption buys consumption.

Question seven · The exit

In three years, what does a buyer's diligence find about our capability, and does our answer depend on these tools staying cheap?

Capability risk is currently underwritten by gut feel in almost every mid-market deal, which means it is also unpriced in almost every exit. A company that runs on tools it does not control, with a second layer that has never operated without them, is a different asset from one that does not, and the discount arrives at the worst possible moment. The pricing half of the question is doing real work: the current cost of these tools is a market position, not a law of physics.

What a managed answer sounds like

"Everyone will be in the same position." Possibly. That is an argument about the market, not about this company, and it is not the one a buyer's diligence will be running.


How to use these

Ask them in one sitting, with the CEO and whoever actually runs the work. Seven questions take about forty-five minutes if the answers are real and about ten if they are not, which is itself the finding.

Write the answers down. The value is not in the first pass, where everyone is thinking out loud. It is in asking the same seven ninety days later and seeing which answers changed because something was fixed, and which changed because the story improved.

These are diagnostic questions about systems. They are not a performance review, they are not about anyone's capability, and they stop working the moment anyone in the room thinks they are being assessed.

The printable version

One page, the seven questions and the managed answers, sized for a board folder. Leave your email and I will send it.

If the answers were worse than you expected

Two of these questions have instruments behind them. The Signal Loss Check measures whether accurate concerns reach a decision, and the Offloaded Judgment Check measures what has been handed over against what still gets checked. Both are free, both take under ten minutes, and neither sends anything anywhere.

Running all of it across a leadership team, changing the practices underneath, and measuring again at ninety days is the engagement I run.

Peter Allen Mann spent thirty-seven years in leadership positions across the Navy, two Fortune 100 companies, and two companies he founded. He writes about why organizations misread the people inside them, and works with leadership teams on the systems that produce it.