Almost every conversation about AI return on investment lands on the same slide. It says something like: "AI saved us 200 hours last quarter." It gets a nod in the meeting, a slot in the board deck, and then nothing. A quarter later, someone asks what the AI investment actually returned, and the room goes quiet.
We think that silence has a cause, and it is not the technology. It is the metric.
The metric that quietly kills AI budgets
Marketing AI Institute's preview of Sandy Carter's MAICON 2026 session, written by Cathy McPhillips, opens on exactly this point: saving time is not the goal of AI, spending it on something better is. Time saved, Carter argues in that piece, is "a currency that expires if you never spend it."
That framing matches what we see in real implementations. Hours saved is an input. It only becomes ROI when it turns into something you can point to.
Why "hours saved" fails the C-suite test
Carter's diagnosis of failed pilots is uncomfortable and correct. The first mistake, she says, is treating saved hours as a result: "Teams tell me they saved 200 hours, and my question is always the same: What did those 200 hours become? A campaign that shipped? Pipeline? Headcount you redeployed? If the answer is nothing you can point to, you did not create value."
That question is the whole game. A CFO does not fund hours. A CFO funds outcomes. When you walk into a budget review with "we saved time" and no downstream artifact, you are asking the finance team to take the value on faith. They will not. And they should not.
The second failure she names is letting every function invent its own scorecard. As AI ownership shifts to line-of-business leaders, the article describes CFOs receiving six different definitions of success from six different teams, with no way to rank them. The third is running too many pilots with no named owners. In Carter's words, "Twelve pilots and zero decisions is not a portfolio."
We have felt all three of these pressures in our own work. It is very easy to spin up a dozen clever automations and end up with no single number anyone agrees on.
Measure something before you automate anything
The most practical thing in Carter's argument is the sequencing. Her starting point for leaders getting their arms around generative AI is simple: measure something before you automate anything. "Half the companies already running AI in production cannot measure its ROI," she says. "They did the hard part, the deployment, and skipped the easy part, the before-and-after."
This is exactly backwards from how most AI adoption actually happens. Teams get excited, deploy, and only later go looking for a baseline that no longer exists. Carter's rule: "Your baseline only exists if you capture it now." Pick one workflow you run every week, write down what it costs today in hours or dollars, define "better" as a number, and name the person who signs off. If you cannot do those three things, she says, "you are not running a pilot. You are running a demo."
We would go one step further. The baseline is not just a number in a spreadsheet. It should live in the same system you use to measure everything else, so the before-and-after is not a special exercise you run once and forget.
What this looks like when the plumbing is real
The point was never the draft. The point was more experiments shipped, more topics tested, and more content actually reaching an audience where we could measure whether it did anything. The reclaimed drafting time only counted once it turned into published, instrumented output we could evaluate.
That is the distinction Carter is drawing. The workflow does not create value by being fast. It creates value because the speed lets us do more of the strategic, revenue-adjacent work that used to be crowded out: choosing better topics, running more variants, and reading the results instead of scrambling to produce the next thing.
Define "better" as a number the workflow can prove
Carter says to define "better" as a number and name who signs off. For a marketing team, that number almost never lives inside the AI tool itself. It lives in your analytics stack — which is the same argument we made about the ROI measurement gap: everyone knows what they spent on AI tools, and almost nobody knows what those tools earned.
This is why we are opinionated about instrumentation. If you cannot see the outcome of the work in GA4 and BigQuery, you cannot prove what the reclaimed time produced. "We drafted more articles" is an input. "More articles shipped, more form events fired, and here is the before-and-after in the warehouse" is an outcome. The second version survives a budget review. The first does not.
Concretely, that means every meaningful action needs to be a tracked event that lands somewhere queryable. A published post is an event. A form submission is an event. A campaign that ships is a thing with a date attached. When those events flow into BigQuery, the before-and-after Carter demands stops being a debate and becomes a query. You are no longer arguing in the quarterly review about whether anything improved. You are reading the answer.
The mindset shift for marketing leaders
If you lead a team, the reframe is this. Stop asking your people how many hours AI saved them. Start asking what those hours became. Make them name the campaign, the experiment, the piece of pipeline, or the redeployed headcount. If the honest answer is "nothing yet," that is not a failure of the tool. It is a signal that you saved a currency and never spent it.
Then do the unglamorous part first. Before you automate the next workflow, capture its baseline. Cost it in hours or dollars. Decide what "better" means as a number. Name one owner. Put the outcome metric somewhere your whole team already trusts, which for most marketing teams is GA4 and BigQuery, not a slide.
And resolve the second and third failures deliberately. Agree on one definition of success across functions, so your finance team is not handed six scorecards. Do not run twelve pilots with no owners — run a few with clear ones, and be willing to make a decision at the end. Which workflow to hand to a model at all is its own decision, and we have written before about matching the task to the model rather than automating whatever is easiest to automate.
None of this is about proving AI works in the abstract. It is about proving what your team did with the time it handed back. That is the number a CMO can defend and a CFO will fund.
Talk to us
If you are trying to prove the ROI of your AI workflows and the honest answer is that you cannot yet measure the before-and-after, that is usually a plumbing problem, not an AI problem. We build the GTM, GA4 and BigQuery instrumentation, and the MCP and n8n workflows, that turn reclaimed time into outcomes you can query. Talk to us about MCP and analytics implementation, and let us help you define "better" as a number you can actually prove.