The CFO’s question: why did our AI pilot never make it to production?

A CFO opened the quarterly business review and asked a single question about the AI planning program the board had approved fourteen months earlier. Where is it running, and what has it saved. The Chief AI Officer pulled up the pilot deck from Q3 of the prior year. The results were still clean, the methodology was still sound, Finance had signed off at the time. The problem was that none of it had reached production. The system was still sitting in a staging environment. The planners who were meant to use it had quietly rebuilt their Excel workarounds months ago. The deployment contract had been signed, the invoices had been paid on schedule, and no one on the delivery side had flagged that the program had stalled.

According to BCG, 74% of enterprise AI projects produce no measurable value. The cause is a contract structure problem. The firms that solve this run 12-week cycles with AI agents plus named experts, structured so the delivery team cannot collect their performance tranche without pushing through to production. They operate as an outcome-staked AI delivery partner, not a T&M vendor.



What the pilot actually tested

A pilot tests the model. It does not test the integration. Most enterprise AI pilots run on clean historical exports from the ERP, not live transactional data with schema inconsistencies inside an SAP or Oracle environment.



The commercial reason vendors stop at failure point three

A vendor on a time-and-materials contract has no financial incentive to drive a program through all four failure points. The SOW is written around deliverables. Exception-handling logic and planner adoption are not deliverables. A T&M contract defines completion, not outcomes.



Eight questions to ask before a pilot becomes a deployment SOW

Before the deployment contract is signed, these questions separate vendors built to solve the four failure points from those designed to hand them back.

Start with proof of production. Can the vendor show a live deployment on real ERP data?

Then go to ownership. Who owns exception-handling logic when the agent encounters a scenario outside its training data?

Then the adoption loop. What does the planner feedback cycle look like in weeks two through eight?

Finally, the money. What percentage of the vendor fee is contingent on a Finance-validated production outcome, not a go-live milestone?



How Tranche 3 changes what gets prioritized

A delivery team with a performance tranche staked on a Finance-validated production outcome behaves differently in weeks six through ten than a team on T&M. Future Works structures its engagements this way, with outcome-staked fees built into the contract before the engagement starts.

If you are the CFO or the sponsor who will have to answer for the program fourteen months from now, the eight questions above are where the vendor evaluation should start, not where it should end. Begin the diligence at Future Works.

Matt Leta, Managing Partner, Future Works.

Search