Skip to content
Curious Kaizer · guide

How to Scope an AI Automation Pilot

An AI pilot is useful when it tests a specific operational hypothesis. Start with a task, a baseline and a decision about whether to continue. A demonstration that produces a plausible answer is not enough to establish accuracy, value or permission to take actions.

Write the task and baseline

Describe the input, expected output, responsible person and current completion steps. Use representative examples that you have permission to process. Record where errors matter and how staff recover from them. Compare with simpler alternatives such as validation rules, templates or an API integration. Avoid assuming every repetitive task requires a language model.

Define allowed information and actions

List approved data sources and roles that may access them. Separate reading a record from drafting or sending a message. Decide which steps require human approval and how the workflow stops if context is missing. Treat retrieved documents and external pages as untrusted data, not instructions authorizing new actions.

Create the evaluation set

Include straightforward cases, ambiguous cases, unavailable information and deliberate attempts to bypass rules. Score factual correctness, appropriate abstention, task completion, latency and cost. Keep a record of the model and configuration used so comparisons remain meaningful. A small set of carefully reviewed examples is more useful than an unsupported accuracy percentage.

Calculate value cautiously

Estimate time saved only after observing actual completed tasks, including review and correction time. Subtract provider usage, maintenance, monitoring and support costs. Separate projected values from measured outcomes. Expand the pilot only when there is evidence that it is useful and its failure behavior is acceptable. Bring the baseline, approved examples and decision criteria to an automation review.