— And Why It’s the Difference Between a Pilot and a Payback
AI Insider Blog Article
The Voice of Dr. Cindy Gordon, SalesChoice · saleschoice.com
I get some version of the same question in nearly every boardroom or C-level discussion I have. The dialogue often is like this: “We have 12 AI use cases on our roadmap — how do we know which ones are actually worth building?” I even had one Chief AI Data Scientist advise that they had identified over 100. Although this is a good question, many organizations we work with are not sharpening their decision-making enough to double down. We have even seen AI consultants approach CEOs with an idea, bypassing the AI Data Science functions to build an AI use case without thinking of the impact downstream, as AI investments require production integration and process ownership E2E. Often AI use cases are prioritized by executive enthusiasm, vendor demo polish, or whoever in the AI governance or innovation committee meeting wanted to sponsor a project and beat their chests. These behaviors are not demonstrating that a governance AI use case evaluation framework is guiding the organization carefully forward. This reflects more as a popularity contest with a technology budget attached.
Why This Matters More Than Most Executives Realize
Let’s start with the uncomfortable numbers, because they explain why this topic deserves a worthy discussion, rather than a bullet point in a strategy deck.
RAND Corporation’s analysis of more than 2,400 enterprise AI initiatives found that 80.3% fail to deliver their intended business value — roughly a third are abandoned before ever reaching production, another third complete but underdeliver, and the remainder produce some value that can’t justify the cost.
MIT’s Project NANDA research is even starker: only about 5% of generative AI pilots achieve measurable profit-and-loss impact. Gartner projects that over 40% of agentic AI projects will be cancelled by the end of 2027, and S&P Global Market Intelligence found that 42% of companies abandoned most of their AI initiatives in 2025 alone, up sharply from 17% the year before.
None of this is because the underlying AI models are weak. Model quality is, in my experience, rarely the reason a use case dies. Use cases die from unclear success criteria, ungoverned data, missing production infrastructure, and change management that was treated as an afterthought rather than a design requirement.
A robust use case evaluation framework exists to catch those failure modes before capital and credibility are spent, not after.
Cindy’s Reflection: In nearly every failed AI project that we have been asked to review or analyze, the technology was never the villain. The villain was an evaluation process — or the absence of one — that let a use case advance on enthusiasm rather than evidence. Boards don’t lose confidence in AI because a model underperforms. They lose confidence because nobody could explain, in plain language, why a project was greenlit in the first place. A rigorous evaluation framework is not bureaucracy. It is a core foundation that protects an organization’s credibility to keep building with confidence.
The Core Risk Parameters Every AI Use Case Evaluation Must Score
At SalesChoice, we anchor every AI Data Science engagement in a structured evaluation before a single model gets built.
Below are the parameters we consider non-negotiable.
1. Type of Use Case: Employee-Facing vs. Customer-Facing vs. Regulator-Facing
This is the first fork in the road, and it should shape everything downstream — risk tolerance, testing rigor, and rollback planning. An internal, employee-facing productivity tool that drafts sales call notes carries a very different risk profile than a customer-facing AI agent making pricing decisions or a model whose outputs feed a regulatory filing. Employee-facing use cases generally tolerate more experimentation because a human is still in the loop reviewing output before it reaches the outside world. Customer-facing and regulator-facing use cases require a materially higher bar for accuracy validation, explainability, and human override before go-live.
2. Data Coherence, Quality, and Lineage
Gartner estimates that 60% of AI projects unsupported by AI-ready data will be abandoned through 2026. Before evaluating whether a use case is technically feasible, ask: Where does this data originate? Has it been validated for completeness, accuracy, and bias? Is there a documented lineage from source system to model input? Is the data governed under a clear ownership model, or does it live in someone’s personal spreadsheet? A use case built on ungoverned data is not a use case — it is a liability with a user interface.
3. Organizational and Technical Readiness
AI Readiness has two dimensions that are too often conflated. Technical readiness asks whether your infrastructure, APIs, and integration layer can actually support the use case at production scale. Organizational readiness asks whether the business unit sponsoring the use case has the process maturity, defined ownership, and executive sponsorship to sustain it past the pilot. A brilliant use case sponsored by a team without the authority or bandwidth to operationalize it will stall in exactly the way the RAND and MIT data describes — a well-built pilot that never crosses into production as foundations were never clearly established with accountabilities.
4. Technology Footprint and Architecture Fit
Every use case should be scored against your existing technology footprint, not evaluated in isolation. Does it require a new model provider, adding to vendor concentration risk? Does it duplicate capability you already have licensed elsewhere? Does it integrate cleanly with your system of record, or does it require a parallel data pipeline that your team will be maintaining manually within six months? The most expensive AI use cases are rarely the ones with the biggest model bill — they are the ones that quietly require a permanent shadow IT function to keep running.
5. Performance Metrics Defined Before Build, Not After
This is where we often see the most consistent failure. A 2026 industry study found that 61% of AI projects were approved on projected ROI that was never actually measured after launch. If you cannot articulate, in advance, the specific metric that will define success — reduced cycle time, error rate reduction, forecast accuracy improvement, adoption rate — you are not ready to build. Vague goals like “improve efficiency” or “enhance customer experience” are not performance metrics; they are aspirations, and aspirations cannot be audited.
6. Change Management Enablers
Technology adoption is a human behavior problem wearing a technology costume. Every use case evaluation should score the change management plan with the same rigor as the technical architecture: Who are the champions? What is the training and enablement plan? How will resistance be identified and addressed? Is there an incentive structure that rewards adoption, or does the new AI tool simply add a step to someone’s existing workflow with no visible benefit to them? Our own agentic AI rollout work has taught us that value realization requires deliberate recognition and celebration built into the plan — momentum has to be engineered; it does not happen by accident.
7. Production Parameters: Governance, Monitoring, and the Ops Layer
The model is the easy part. The system around the model — data pipeline, drift monitoring, human-in-the-loop checkpoints for high-stakes decisions, incident ownership, and a documented rollback plan — is what determines whether a use case survives contact with production traffic. Every evaluation should require a clear answer to: who owns this at 2 a.m. when it breaks? If there is no answer, the use case is not production-ready, regardless of how well the demo performed.
Cindy’s Reflection: I want to underline production parameters specifically, because it is the category executives most often skip. Infrastructure cost overruns are one of the leading causes of production AI failure — deployments routinely run three to five times the initial cost projection once real monitoring, governance, and support requirements are added in. If your business case doesn’t include the ops layer, it isn’t a business case. It’s a demo with a budget attached.
The Technology Variable: Why Risk Escalates as You Move From Predictive AI to Agentic AI
One other key dimension that is crucial to understand is that not all AI is the same kind of risk, and treating a predictive model and an autonomous agent as points on the same scale is one of the more consequential mistakes I see enterprises make right now.
Traditional predictive and machine learning use cases follow a familiar, contained pattern: a fraud model flags a transaction, a recommendation engine suggests a product, a forecasting model produces a number — and in every case, a human reviews the output and decides what happens next. The AI proposes; the human disposes. Agentic AI inverts that relationship entirely. Agents plan, decide, and execute multi-step tasks with minimal human input, chaining tool calls, data pulls, and downstream actions together with no natural pause point for review. That is the capability that makes agentic AI so compelling — and it is exactly what makes the risk profile fundamentally different, not just larger.
The industry data on this is now substantial enough to be taken seriously. Gartner predicts that more than 40% of agentic AI projects will fail by 2027 due to runaway costs, unclear business value, and agents that behave in ways that violate policy — and, separately, that 40% of enterprises will demote or decommission autonomous AI agents by 2027 because governance gaps were only discovered after a production incident occurred.
Gartner’s own diagnosis is blunt: organizations are treating agent governance as binary — locked down or fully trusted — when the controls actually need to flex with the level of autonomy and the consequence of the action being taken. An agent that summarizes a document and an agent that negotiates a vendor contract are not the same risk, even if they run on the same underlying model.
The compounding-error problem is what makes this more than a theoretical concern. When one agent in a chain acts on faulty data, agents further down the workflow tend to treat that output as authoritative and build on it, so a single bad instruction can cascade through a multi-step process before a human ever sees it.
A 2026 insider-risk study from the Ponemon Institute and DTEX Systems found that 92% of organizations say generative AI has already changed how employees access and share data, yet only 13% have a formal enterprise AI policy in place — and agentic systems widen that gap further, because standing permissions let an agent read, transform, and move data across systems in seconds, with no human checkpoint at all. The same research flags over-permissioned agents — granted broad database or admin-scoped access simply so they can “just work” — as one of the most common and preventable failure patterns.
It’s also worth being honest about the limits of the fix. The International AI Safety Report 2026 makes an important, under-discussed point: human-in-the-loop oversight is frequently impractical at the speed and scale agentic systems operate at, and even when a human is present, they tend to exhibit “automation bias” — placing more trust in the AI’s output than is actually warranted. In other words, simply inserting a human checkpoint into a workflow is not, by itself, a control. It only functions as one if the human genuinely has the context, time, and independence to catch a bad decision — which is a design requirement, not a checkbox.
Cindy’s Reflection: The way I frame this for boards and C-Level executives is that risk does not scale linearly with capability — it scales with autonomy multiplied by irreversibility. A highly capable predictive model that only ever produces a recommendation is a fundamentally different governance problem than a far less capable agent that can independently execute an irreversible action. This is precisely why our own agentic AI work leans on a phased approach — prioritize by value and risk, prove value through agile proof-of-concept before granting broader autonomy, and only expand an agent’s authority once there is a demonstrated, monitored track record. Every use case evaluation now needs an explicit question that older predictive-AI checklists never had to ask: exactly how much decision authority is this system being given, and what happens if it gets that decision wrong before a human ever finds out?
Lessons Learned: What Separates the Use Cases That Scale
A few patterns show up consistently in the organizations that beat the odds:
-
They score use cases on a value-versus-risk matrix before committing engineering time, not after a proof of concept is already half-built.
-
They run pilots at production-shaped traffic and realistic data volumes, not a curated fifty-prompt demo that flatters the model.
-
They instrument evaluation metrics from day one, so a model change can be judged as an improvement or a regression rather than a matter of opinion.
-
They assign a single accountable owner for the use case’s full lifecycle — evaluation, pilot, production, and sunset — rather than handing it off between teams at each stage.
-
They treat human-in-the-loop design as a permanent feature of high-stakes use cases, not a temporary training-wheels phase to be removed once the model looks confident.
Cindy’s Reflection: The organizations that consistently get this right are not the ones with the biggest AI budgets. They are the ones with the most disciplined evaluation habit. It is far less glamorous than announcing a new model deployment, but it is the single highest-leverage capability an enterprise can build right now — because it compounds. Every use case you evaluate well makes the next one faster and cheaper to ship, and every use case you wave through on enthusiasm alone makes the next one harder to fund.
For those of you following my newsletter, you know I like to close with a bit of whimsy and a different voice to help the lessons land. So take a read below — I promise you’ll smile. Please share this with others; the more we learn together, and the more we hear these stories told in different voices, the more we remember them.
Lady Whistledown’s Society Paper on Suitors, Suitability, and the Perils of a Hasty Engagement
Dearest Gentle Readers,
One does so enjoy the season’s parade of eligible new technologies, each arriving at the ball dressed in its finest demo, promising the enterprise everything a modern household could wish for. And yet, dear readers, this columnist must report a most curious pattern: an alarming number of these engagements end not in a happy union, but in a quiet, expensive annulment before the year is through.
Four in five, by the most reliable counts available to this author. Four in five suitors who make it as far as the introduction, only to be abandoned at the altar of production.
One must ask: what, precisely, went wrong at the courtship?
The answer, gentle readers, is rarely a defect in the suitor. The finest technologies in the land have been jilted just as often as the mediocre ones. No — the trouble lies almost entirely with the household that failed to ask the proper questions before the engagement was announced. Was the family’s data in good order, or was it, shall we say, of questionable lineage? Had anyone determined who would manage the household accounts once the honeymoon ended? Was there so much as a plan for what happens should the marriage require adjustment?
Too often, the answer to all three is a rather sheepish silence.
It has become fashionable in certain circles to announce an engagement simply because a rival household has done the same. This, dear readers, is not strategy. It is vanity wearing the costume of ambition, and vanity has never once balanced a ledger.
The households that thrive, this author observes, are not the ones who court the most suitors, nor even the ones who court the most fashionable ones. They are the households with a discerning aunt — call her governance, call her evaluation, call her whatever you like — who insists on meeting the intended thoroughly before the engagement is permitted to proceed. She asks the unglamorous questions. She wants to know who will mind the household at two o’clock in the morning should something go amiss. She is, on occasion, accused of being no fun at all.
She is also, without exception, the reason the marriages that survive her scrutiny tend to last.
So to every executive currently drafting an announcement about the twelve thrilling new AI courtships planned for next quarter, this columnist offers a gentle word of caution: an engagement is not an achievement. A marriage that endures the season is.
Until next time, dear readers, may your evaluations be thorough, your data be tidy, and your ownership of the household accounts be entirely unambiguous.
Yours most observantly,
Lady Whistledown
Lady Whistledown’s Closing Counsel to Board Directors
“A suitor’s charm at the ball reveals nothing of his character at home. The wise household does not ask whether a technology impresses in the drawing room, but whether it can be trusted with the accounts, the servants, and the family’s good name once the music has stopped. Evaluation is not the enemy of romance — it is the only thing that makes a lasting one possible.”
Bibliography
1. RAND Corporation. “The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed.” 2025 analysis of 2,400+ enterprise AI initiatives. https://www.rand.org/
2. MIT Project NANDA. “The GenAI Divide: State of AI in Business 2025.” https://nanda.media.mit.edu/
3. Gartner. Predictions on generative AI and agentic AI project abandonment, 2024–2027. https://www.gartner.com/
4. S&P Global Market Intelligence. “AI Proof-of-Concept and Production Trends, 2025–2026.” https://www.spglobal.com/marketintelligence/
5. Covasant. “AI Pilot Purgatory: How to Reach Production Faster.” https://www.covasant.com/blogs/ai-pilot-purgatory-production-timelines
6. Institute PM. “Why 88% of Enterprise AI Pilots Never Reach Production — and How to Be in the 12%.” https://www.institutepm.com/knowledge-hub/why-enterprise-ai-pilots-fail
7. Folio3 AI. “AI Project Failure Rate in 2026: What the Data Shows.” https://www.folio3.ai/blog/ai-project-failure-rate-stats
8. Webpuppies. “From AI Pilot to Production: Why 80% of Enterprise AI Stalls in 2026.” https://webpuppies.com.sg/ai-pilot-to-production-enterprise-2026/
9. Pertama Partners. “AI Project Failure Rate 2026: 80% Fail.” https://www.pertamapartners.com/insights/ai-project-failure-statistics-2026
10. SalesChoice. “Navigating the Promise of Agentic AI in Sales: How to Get Ready and Lessons Learned.” https://www.saleschoice.com/navigating-the-promise-of-agentic-ai-in-sales-how-to-get-ready-and-lessons-learned/
11. SalesChoice. “Case Study: How SalesChoice Delivers Explainable AI in an Era of Increasing Regulation.” https://www.saleschoice.com/case-study-how-saleschoice-delivers-explainable-ai-in-an-era-of-increasing-regulation/
12. SalesChoice. “Purolator — Data Science Case Study.” https://www.saleschoice.com/case-studies-sales-enablement/purolator-data-science/
13. Gartner. “Gartner Says Applying Uniform Governance Across AI Agents Will Lead to Enterprise AI Agent Failure.” May 26, 2026. https://www.gartner.com/en/newsroom/press-releases/2026-05-26-gartner-says-applying-uniform-governance-across-ai-agents-will-lead-to-enterprise-ai-agent-failure
14. Joget (citing Gartner, IDC). “AI Agent Adoption 2026: What the Data Shows.” https://joget.com/ai-agent-adoption-in-2026-what-the-analysts-data-shows/
15. IndyKite. “Gartner Predicts an AI Agent Governance Reckoning.” June 2026. https://www.indykite.ai/blogs/gartner-predicts-an-ai-agent-governance-reckoning
16. Raconteur. “Autonomous AI Agents 2026: The New Rules for Business Governance.” March 30, 2026. https://www.raconteur.net/technology/autonomous-ai-agents-2026-the-new-rules-for-business-governance
17. Insider Risk Index / Ponemon Institute & DTEX Systems. “Agentic AI as an Insider Threat in 2026: When Autonomous Agents Go Rogue.” https://www.insiderisk.io/research/agentic-ai-insider-risk-2026
18. International AI Safety Report 2026. “Humans in the Loop Allow for Direct Oversight in High-Stakes Settings.” https://arxiv.org/pdf/2602.21012
19. Strata.io. “Human-in-the-Loop: A 2026 Guide to AI Oversight.” May 11, 2026. https://www.strata.io/blog/agentic-identity/practicing-the-human-in-the-loop/
