Evaluate an AI investment by asking whether a defined operational problem has evidence of sufficient value, whether the proposed system can meet an agreed performance and risk threshold, and whether the organization can support the complete operating model. Treat benefits as hypotheses until they are measured against a documented baseline. Do not use industry averages, vendor percentages, or an impressive demonstration as a substitute for your own evidence.
This is an investment-governance question before it is a return calculation. The first useful outcome may be a decision not to build, to narrow the scope, or to solve the problem without AI.
Start with the decision, not the technology
Write the decision the organization needs to make. A useful statement includes:
- the workflow and accountable owner;
- the affected users or customers;
- the current failure, delay, capacity, or risk;
- the observable result that would justify further investment;
- constraints that cannot be traded away;
- the date or event that triggers a review of the decision.
“Adopt generative AI” is not an investment thesis. “Determine whether assisted document intake can meet this reviewed quality threshold without increasing unresolved exceptions” is testable.
The GAO AI accountability framework organizes oversight around governance, data, performance, and monitoring. That structure is useful for commercial projects even though the framework was developed for federal agencies and other entities.
Separate four kinds of evidence
Keep these categories distinct in a decision memo:
- Known baseline: observed volume, cycle time, corrections, handoffs, cost, incidents, and service constraints in the current process.
- Design assumption: expected adoption, automation rate, model usage, review effort, support burden, or demand growth that has not yet been measured.
- Pilot observation: results from a limited group, time period, data sample, and authority level.
- Operational result: repeated measurements after the system enters its intended environment, including exceptions and maintenance.
A pilot observation can support the next stage; it does not prove that the result will persist at broader scale. Record sample, exclusions, measurement method, and uncertainty so later readers know what the number means.
Identify value without double counting
Possible value can come from shorter queues, avoided re-entry, increased capacity, fewer corrections, faster access to information, improved consistency, or a capability the current process cannot support. Some value is financial; some is operational or risk-related.
Do not count the same effect twice. Time released from a task is not automatically a cash saving. It becomes realized value only if the organization can use that capacity, avoid a cost, improve service, or reduce a measured risk. Ask who receives the benefit and what will change in practice.
Avoid assigning a dollar value to trust, safety, or compliance merely to make a proposal appear complete. A threshold, control requirement, or qualitative decision may be more honest.
Include the complete obligation
The investment includes more than implementation and model usage. Consider:
- workflow discovery and source cleanup;
- integration, identity, permissions, and testing;
- review and exception handling;
- accessibility, privacy, security, and compliance work;
- vendor usage, hosting, storage, and monitoring;
- training, change management, and support;
- evaluation after model, prompt, source, or policy changes;
- failure recovery, reconciliation, and eventual replacement.
Some of these costs are estimates before a pilot. Mark them as ranges or assumptions supported by actual quotes and workload data; do not present them as measured fact.
Define risk and stop conditions
The NIST AI Risk Management Framework asks organizations to map context, measure performance and risk, and manage what remains. For an investment decision, define unacceptable outcomes and controls alongside value.
Examples of stop or redesign conditions include unsupported answers on consequential topics, inability to reproduce source evidence, excessive unresolved exceptions, uncontrolled data exposure, poor user adoption for a valid reason, or dependency on access a vendor does not reliably provide.
A lower-cost system is not a positive return if it transfers hidden work to customers, reviewers, or incident response.
Use stage gates
At each gate, decide whether to stop, narrow, redesign, or continue:
- Problem gate: Is the problem measured and material enough to investigate?
- Design gate: Is AI a better fit than a process change, configured product, integration, or deterministic software?
- Prototype gate: Does the approach work on representative examples under review?
- Pilot gate: Can real users operate it safely with manageable exceptions?
- Operations gate: Does the measured value justify recurring ownership and residual risk?
Do not pre-commit to later stages. A responsible discovery phase can produce a negative recommendation.
Require supportable claims
The Federal Trade Commission’s AI claims guidance warns businesses to support claims about what an AI product can do and whether it performs better than a non-AI alternative. Internally, apply the same discipline: retain the baseline, test design, exclusions, results, and approval behind any performance or value statement.
Our methodology explains how this site distinguishes observations, estimates, and examples. For implementation choices, see business automation and data intelligence. Bring an observed workflow and decision constraint to a scoping call; a target return invented before discovery is not a reliable starting point.