Most large organisations now hold a long list of AI ideas. Some have come from technology teams exploring new capabilities, some from business units that have seen what peers are doing, and many from vendors demonstrating features that are now built into platforms the organisation already owns. A growing number have progressed to proof of concept (PoC), and a smaller number have produced results impressive enough to prompt a request for production funding.
The distance between those promising results and a dependable production service remains considerable. Gartner found that by the end of 2025, at least half of generative AI projects had been abandoned after PoC, citing poor data quality, inadequate risk controls, escalating costs or unclear business value. McKinsey's latest global survey reports that nearly two-thirds of organisations have not yet begun scaling AI across the enterprise, and just 39 percent report any effect on financial return.
In our work with organisations assessing AI opportunities, we find that the causes of stalled use cases are largely predictable. An outcome was never clearly defined; the workflow the AI was meant to support had not been mapped; the data was less ready than the PoC implied; nobody in the business owned the result; or integration and running costs emerged late, once commitments had been made. Each of these causes can be identified early, and a consistent filter applied before significant investment allows leaders to concentrate budget, talent and attention on the use cases with the most credible route to operation.
Why a successful PoC can mislead
A PoC is designed to show what a model can do and typically runs on a curated subset of data, in a controlled environment, with a small expert team guiding it and little dependence on the systems or people around it. Under those conditions, pilots can perform impressively, and the leap from to a production business case can feel short.
However, operation depends on a different set of conditions. The model must work with the organisation's real data, including its gaps and inconsistencies, fit into a workflow that people follow under time pressure, and manage exceptions that the PoC never encountered. It must integrate with systems of record, satisfy security and data protection requirements, and be monitored, supported and improved over time. It must also cost less to run than the value it creates, at production volumes.
McKinsey's research highlights where the difference lies. Among the small group of organisations it identifies as AI high performers, workflow redesign is a key factor in success, with most of these organisations fundamentally redesigning workflows as part of their AI programmes. The value of AI emerges from the combination of capability and changed ways of working, and a filter for use cases needs to assess both.
The eight criteria
We assess use cases against eight criteria, each rated very weak, weak, moderate, strong or very strong. A strong rating indicates a strong case for investment on that criterion. For the two risk criteria, a strong rating therefore means the risks are well understood and proportionate, so the scale always reads in the same direction.
1. Value
Value assesses the business outcome the use case is expected to deliver, expressed in terms the organisation already measures, such as handling time, error rates, cost per transaction, revenue or customer satisfaction. The assessment should consider value net of the cost to build and run the use case at production volume, since Gartner has cautioned that generative AI costs are less predictable than those of other technologies, varying with the use cases chosen and the deployment approach taken. It is worth beginning with the end in mind, as a successful proof of value shows what a use case could achieve, and the value itself arrives only when the use case is running in operation, embedded in the way people work. For example, use cases framed around capability, such as "use AI to summarise documents", tend to rate weakly here until they are reframed around the outcome the summaries are meant to improve.
2. Ownership
Ownership considers whether a named business owner feels the problem, will champion the change and will be accountable for the result. In our experience, it is one of the most decisive criteria. McKinsey found that AI high performers are three times more likely to strongly agree that senior leaders demonstrate ownership of and commitment to their AI initiatives. A use case owned solely by the technology function may prove that the technology works, and it will rarely change how the business operates.
3. Complexity
Complexity reflects the effort and uncertainty involved in delivering the use case. This includes the number of systems it must integrate with, the maturity of the techniques involved and the scale of change required around it. Integration deserves particular attention, since PoCs often run as standalone applications. Many AI capabilities deliver value only when their outputs flow into a system of record or appear in the tool where people already work.
4. Delivery risk
Delivery risk concerns the likelihood that the use case will fail to reach a working outcome. Short-term failure at this stage is acceptable and often valuable. A PoC that fails quickly and cheaply, and produces clear lessons, is a good use of investment. The aim of this criterion is to confirm that failure would be contained and informative, with clear points at which the organisation can decide whether to continue.
5. Organisational risk
Organisational risk addresses the wider consequences of the use case for the organisation, its customers and its obligations. This covers data protection, bias and fairness, security, and the reputational and regulatory effect of errors. For organisations processing personal data, the ICO's guidance is clear that a data protection impact assessment is legally required in the majority of cases where AI systems process personal data. Updated guidance also expects organisations to show within that assessment that less risky alternatives were considered, with reasoning for why they were not chosen. Organisations operating in the EU will also need to consider their obligations under the EU AI Act.
6. Data requirements
Data requirements considers what data the use case needs and whether those needs can be met. Some use cases depend heavily on the organisation's own data, while others, such as many drafting and summarisation tools, need relatively little. Where significant data is required, the assessment should consider its availability, quality and lawful use. Gartner predicts that through 2026, organisations will abandon 60% of AI projects that are not supported by AI-ready data. A use case with modest data needs, well met, should rate as strongly as one with extensive data needs that are equally well met.
7. Process readiness
Process readiness assesses how well the organisation understands the process the use case will support, and how ready that process is to change. This includes the steps people follow, the decisions they make, the exceptions they handle and the systems they use. Where a process is poorly understood or resistant to change, the use case carries hidden scope, and the effort to understand and redesign the process should be counted as part of its cost.
8. Measurability
Measurability asks whether the effect of the use case can be measured reliably, from baseline through PoC and into operation. KPIs set at the outset remain essential, and their role continues beyond the start of the work. AI approaches frequently evolve during a PoC or build, particularly with large language models and agents, where prompts, models and architectures may change several times. Each change can alter both quality and cost. Ongoing measures of return and output quality confirm that each updated approach still adds value, and they give the team an objective basis for choosing between alternatives.
Running the evaluation
Evaluating is designed to be a quick exercise. For most organisations, a focused working session is enough to assess a shortlist of use cases, provided the right people are involved, such as the business owner of each use case, representatives from technology, data, and risk or compliance. Involving stakeholders from the outset brings a balanced view of each area. It also builds the shared understanding and commitment that the use cases selected will need later.
The filter also gives boards and executive teams a clearer basis for their AI ambitions. A mandate to "do AI" gives teams little to work with. A statement of the business drivers behind the AI programme, and of what good outcomes look like, allows use cases to be assessed against value the organisation has defined for itself. The ratings then give leaders a concise, comparable view of where investment is likely to produce results.
Three illustrative use cases show how the filter distinguishes between opportunities that may appear equally promising at PoC stage.
Call summarisation in a contact centre:
The outcome is clear: less after-call work and more consistent case notes, both of which the organisation already measures. The contact centre manager owns the result. The process is well understood, call transcripts are available and the effect on handling time can be tracked directly. Integration with the case management system adds moderate complexity, and the presence of personal data in call recordings requires careful handling.
Invoice matching in finance:
The value is clear and measurable, and the finance operations lead is committed. The filter reveals weaknesses elsewhere, however. Supplier data is inconsistent across two legacy systems, the exception process for mismatched invoices has never been documented, and integration with both systems adds complexity.
A policy drafting assistant:
The PoC impressed senior stakeholders, and its data requirements are modest, since it draws largely on general language capability and a small set of templates. The filter shows that no business owner has been identified and that the intended outcome has not been defined beyond "faster drafting". It also shows that nobody has assessed the regulatory consequences of errors in published policy, and that no measures are in place to track quality.
| Criterion | Call summarisation | Invoice matching | Policy drafting assistant |
|---|---|---|---|
| Value | Strong | Strong | Weak |
| Ownership | Very strong | Strong | Very weak |
| Complexity | Strong | Weak | Strong |
| Delivery risk | Strong | Moderate | Strong |
| Organisational risk | Moderate | Strong | Weak |
| Data requirements | Strong | Weak | Strong |
| Process readiness | Very strong | Weak | Moderate |
| Measurability | Very strong | Strong | Very weak |
| Outcome | Proceed | Reshape | Pause |
The ratings are best read together as a profile. In our experience, weak ratings on value or ownership are the hardest to compensate for elsewhere, so they warrant particular attention when deciding the outcome.
Proceed, reshape or pause
The filter produces one of three outcomes for each use case, and each carries a clear next step.
A use case that proceeds has rated well across the criteria, with no significant weaknesses in value or ownership. It should move forward with a defined scope, the measures agreed during assessment, and a plan for integration, monitoring and support.
A use case that is reshaped has shown genuine promise and revealed weaknesses that would undermine it in production. Reshaping may mean narrowing the scope, addressing data or process readiness first, or simplifying the approach to reduce complexity. The invoice matching example belongs here. Its value is clear, and a period of data and process work would give it a credible route to operation.
A use case that is paused is held until the conditions for success are in place. Pausing is a deliberate decision, recorded with the ratings and the reasons behind it, so that the use case can be reassessed when those reasons change. A business owner may emerge, a risk assessment may be completed, or measures of quality may be defined. We recommend revisiting paused use cases periodically, since the conditions that held them back often shift within a year.
Concentrating investment where it counts
The organisations that gain most from AI are those that choose carefully which opportunities to pursue, and then commit the resources, ownership and workflow change needed to carry those opportunities into operation.
This article builds on our earlier discussion of AI as an operating model, and it shares a common principle with our latest articles - evidence should come before commitment. At Audacia, we apply this filter as part of the way we help organisations move from AI ideation to operation. We begin by understanding the organisation's goals and the opportunities already in view, define each use case against the eight criteria with the people who will own it, recommend where investment should be concentrated, and enable the organisation to deliver the use cases that proceed.


