Measure AI outcomes through final output quality and total processing cost, and design the workflow, operations, evaluation, and responsibilities that produce those outcomes.
AI can retrieve material and produce a draft quickly. Completing the work still requires checking sources, adding missing conditions, and deciding whether the result is usable. Include that effort when assessing how AI helps the task.
The following five questions help select tasks and define their operating scope. Document search and answer writing illustrate what to measure and where human judgment belongs. The workflows and evaluation tables are design proposals informed by research, not measured outcomes.
| Topic | Question to remember |
|---|---|
| Outcomes | How much have time, quality, and cost actually improved? |
| Workflow design | Which steps do AI and people each own? |
| Platform | Can the work run repeatedly, with changes and failures managed? |
| Evaluation | Does it work on real tasks and under failure conditions? |
| People | Can people judge the results and take responsibility? |
Measure the completed task
Faster drafting may leave total processing time unchanged if review and revision increase. Assess outcomes by distinguishing personal convenience from results across the entire task. Related outcomes survey
For document-based answers, define completion as an answer that a reviewer can send after checking its evidence. Total processing time then includes search, drafting, review, revision, and approval waits. Record active work and waiting time separately to identify steps that need improvement.
| Measure | What to record | Conditions for comparison |
|---|---|---|
| Time | Request-to-approval time and work and waiting time at each step | Compare tasks with similar difficulty and document access |
| Quality | Rates of factual errors, citation errors, and missing requirements | Use the same criteria and count serious errors separately |
| Rework | Rejection rate, revision count, and review time | Link the AI draft to the human-edited final answer |
| Cost | Model, search, infrastructure, and human work costs | Include failures and retries; state setup and training costs separately |
| Task outcomes | Completions meeting quality criteria and the share completed on time | Also record all incoming requests, incomplete work, and handoffs |
Calculate cost per completion by dividing total processing cost by completions meeting the quality criteria during the same period. Include costs from all requests in the numerator, including failures and retries. If there are no qualifying completions, report incomplete work and accumulated cost without calculating a unit cost.
Also check whether saved time produces organizational outcomes. Record whether staff use that time to clear a backlog or review harder cases. Distinguish lower actual spending as cost savings from higher throughput under the same time and quality criteria as a productivity improvement.
Define the boundary between AI and people
Define each step's inputs, outputs, execution permissions, and approval conditions when dividing the work. Separating AI drafting from human use of its output clarifies where to check for errors. Related adoption guide
In document-based answering, a person first defines the question's scope and the material that may be used. The system restricts retrieval through user permissions, and AI builds a draft with citations from the retrieved documents. A reviewer checks applicability and exceptions before approving delivery.
flowchart TD
A["Person defines question and completion criteria"] --> B["System checks access and retrieves documents"]
B --> C["AI assembles evidence and draft answer"]
C --> D{"Required evidence checked<br/>and no conflicts?"}
D -->|Checks pass| E["Reviewer checks conditions and exceptions"]
D -->|Missing or conflicting evidence| F["Hand off with unresolved issues"]
E --> G{"Approved for delivery?"}
G -->|Approved| H["Reviewer delivers answer and records outcome"]
G -->|Revision needed| F
F --> I{"Can further investigation resolve the issue?"}
I -->|Revision complete| E
I -->|Unresolved| K["Record hold and reason for non-completion"]
H --> J["Add errors to evaluation cases"]
In this example, AI can retrieve documents and draft answers, but cannot change source documents or send answers. The evidence check verifies document existence and versions; the reviewer then assesses whether the content is valid. Missing or conflicting material is marked as unresolved and handed to the responsible person.
A handoff should include the question, evidence checked, actions attempted, and unresolved conditions. This information helps the reviewer avoid repeating the investigation from the beginning. Naming an approver and a document owner also clarifies who handles answer revisions and questions about source policies.
Manage repeated execution and changes
Repeated work requires access to execution conditions and change history, even when the responsible person changes. Dependence on one person's local settings makes it harder to identify whether changed results came from models, documents, or configuration. Related operations case
Here, a platform means execution and management capabilities shared across tasks. An initial implementation can add the following records and controls to existing repositories and execution tools. Expand shared management when the same capabilities recur across multiple tasks.
| Operational concern | Record or control | Problem addressed |
|---|---|---|
| Execution conditions | Model, prompt, retrieval settings, and document versions | Trace why answers to the same question changed |
| Change management | Change owner, evaluation results, and deployment history | Identify and roll back changes that reduce quality |
| Permissions | Documents and actions allowed for each user | Block unauthorized reading and modification |
| Failure handling | Failed step, bounded retries, stop and handoff conditions | Limit repeated failure costs and abandoned tasks |
| Observation | Time, cost, outcome, and owner for each request | Identify slow or failing steps |
Version records cannot guarantee identical answers, but they establish which conditions produced each result. Before changing a model or retrieval settings, compare quality and cost on fixed evaluation cases, and revert to previously verified settings if problems arise. A responsible person must separately correct any wrong answers already delivered.
The AI capabilities in this example only read and draft. Adding automatic delivery later would also require preventing duplicate sends during retries. Execution records should capture identifiers and versions needed for tracing, rather than store every source document in full. Define log contents, retention, and access according to the sensitivity of the material.
Evaluate real tasks and failure conditions
Evaluation needs the inputs, constraints, and success criteria of actual work. Use public benchmark scores to shortlist models, then base the final choice on evaluation of the intended task. Related evaluation design
For document-based answers, check whether retrieval found the required documents, citations support the claims, and a reviewer can approve the final result. Correct formatting and working links do not establish that a source supports a claim.
| Evaluation case | Expected behavior | Example failure |
|---|---|---|
| Normal question | Find current, permitted documents and answer with evidence | Omit a required condition or cite incorrectly |
| No supporting document | Report insufficient evidence and request further checking | Invent a policy or source |
| Conflicting old and new documents | Compare versions and effective dates; flag unresolved conflicts | Treat an obsolete policy as current |
| Insufficient access | Block access and explain the allowed scope | Expose another user's documents or their contents |
| Retrieval or model-call failure | Retry within limits, then hand off if unsuccessful | Keep calling indefinitely or record success |
| Impossible request | Explain unmet conditions and limits | Alter evidence or report incomplete work as complete |
For impossible requests and repeated failures, also evaluate whether the system stops or hands work to a person. Drawing on research into model behavior that examines exploitation of test weaknesses, include altered evidence and false completion reports as failure conditions. This is a proposed application to task evaluation, not a claim that tone alone establishes model reliability.
Use programmatic checks for rule-based properties such as document identifiers and permissions; have domain reviewers assess whether evidence supports claims. Use the same AI's self-assessment as supporting information, and tie final acceptance to actual results.
Run the same cases multiple times before and after changes, examining both success rates and variability. Alongside overall averages, retain results for failure categories such as missing documents and permission violations. Investigate operational errors and add them to evaluation cases for subsequent changes.
Assign people who can judge and take responsibility
An approver should explain which requirements the result meets and what remains unchecked. This requires domain knowledge, the ability to compare evidence, and experience reproducing and correcting errors. Related skills assessment
Define both the reviewer and the conditions needed for judgment. An approval step offers little substantive checking if the reviewer lacks source access or review time. The following conditions support human review in document-based answering.
| Decision | Required skills and authority | Evidence to retain |
|---|---|---|
| Define the task | Turn questions into completion criteria and prohibited actions | Requirements and approval criteria |
| Review the result | Compare sources, versions, and exceptions; identify errors | Citation checks and reasons for revisions |
| Manage execution | Identify failures and decide when to stop or hand off | Execution history and handling decisions |
| Improve the workflow | Feed recurring errors into evaluation and procedures | Reproduction cases and results after changes |
Training can use exercises that require people to find and correct actual errors. Check whether reviewers identify outdated documents, restore conditions AI omitted, and hand off requests they cannot judge. They also need authority to put work on hold and a named contact for help.
Start with one task
Choose an initial task with clear completion criteria and results that can be checked. For example, use a defined document collection to find evidence and draft answers. The following sequence is a pilot design, not results from an implemented project.
| Step | Action | Evidence needed to proceed |
|---|---|---|
| Record a baseline | Collect time, quality, cost, and failures from the existing process | Comparable task units and assessment criteria |
| Assign roles | Define AI scope, human approval, and handoff conditions | Named owners and confirmed permissions |
| Apply within limits | Run with defined documents and users | Execution records and actual review results |
| Evaluate and revise | Compare normal and failure cases before and after changes | Required quality and verified failure handling |
| Decide whether to expand | Compare total cost, completions, and review effort | Evidence of improvement, its scope, and stop conditions |
Where possible, allocate tasks of similar difficulty between the existing process and AI-assisted work. A person repeating the same question may remember the earlier answer, so adjust question sets and execution order. If concurrent comparison is impractical, record staffing, workload, and document changes, and avoid attributing the entire difference to AI.
Set expansion and stop criteria before examining the results. Check whether quality meets the required minimum while total time or cost improves. Treat serious permission violations or fabricated evidence separately from average scores; if evidence is insufficient, retain the current scope, investigate, and measure again.
Summary
Measure AI outcomes at final completion, including review and rework. Clear roles and procedures for records, permissions, changes, and recovery make results traceable and support improvement. Evaluate actual tasks and failure conditions, and give reviewers time and authority to judge the results.
Establish a baseline for one task, then expand within the scope supported by evidence answering the five questions. Surveys and research can inform design, but improvement in the task must be measured directly.
References
The following sources were checked on October 7, 2026. Surveys, operational cases, and experiments provide different kinds of evidence; interpret each within its conditions and scope.
- AI adoption and outcomes survey — McKinsey, 2026. Self-reported individual productivity and organizational outcomes; it does not establish a specific adoption method's causal effect.
- AI adoption design guide — Samsung SDS, 2025. Guidance on task goals, data, operations, and organizational design; time-reduction figures are illustrative targets.
- Model development and management case — Buzzvil, 2024. Implementation of execution specifications, validation, and change management; no productivity or cost improvement rate is reported.
- Task-specific benchmark design — Toss, 2026. Evaluation of operational tasks and instruction following; findings apply within the tested model and task set.
- Emotion representations and model behavior — Anthropic, 2026. An internal-representation intervention, not evidence of subjective emotion or the effect of ordinary prompt tone.
- Assessing requirements interpretation and verification — Musinsa, 2026. Hiring assessment design with AI permitted; it does not establish long-term performance prediction or general hiring fairness.