← Back to all field notes

Which Repetitive Work Should an AI Agent Take First? Use These Five Tests

Start from a task rather than a job title, then score business value, frequency, knowledge readiness, controllable risk, and process ownership.

Companies increasingly ask whether AI can enter a real workflow and remain useful, not whether someone can build another demo. The choice of pilot matters more than the choice of tool. A poorly selected pilot turns even a strong model into internal presentation material.

Start from a task, not an entire job. Look for a bounded group of repetitive actions where people currently spend time gathering information, applying rules, preparing a decision, or transferring context.

Three promising process types

1. Knowledge-intensive work with scattered rules

Examples include standard operating procedures, exception handling, customer answers, and cross-team handoff instructions. These processes quickly reveal whether company knowledge is actually readable and maintainable.

If naming, versions, and ownership are inconsistent, AI will amplify the disorder. The pilot may still be valuable because it forces the team to build a better knowledge structure.

2. High-frequency coordination

Status updates, reminders, material preparation, and repeated information transfer are often good starting points. They produce enough samples for a short pilot and their outputs are usually easy to inspect.

3. Decision preparation with a human final call

An agent can extract fields, retrieve rules, show conflicts, and prepare a recommendation while a person keeps the final authority. This is often safer and more informative than attempting full automation immediately.

Score the candidate on five dimensions

Use a 0–2 score for each dimension:

Dimension012
Business valueoutcome is unclearimproves a local experienceaffects an observable business result
Frequencyoccasionalweeklydaily or every business batch
Knowledge readinessmostly tacitdocuments exist but versions are messyrules and exceptions are traceable
Controllable riskfailure is hard to reversea person can reviewthe workflow can stop and roll back safely
Process ownerno one decidessomeone coordinatesa named owner controls rules and outcomes

The highest score is not automatically the best choice. A zero in controllable risk or ownership is a warning that should be repaired before development.

Define the smallest useful loop

A first pilot should fit into two to four weeks and answer one practical question. For example:

Can the system receive a freight inquiry, identify missing fields, retrieve the relevant rules, prepare a review package, and hand exceptions to the correct person?

This is much more testable than “automate sales.”

Define:

  • the input and sample;
  • the output and acceptance rule;
  • the human baseline;
  • the actions the agent may take;
  • the conditions that stop the workflow;
  • the owner who reviews exceptions;
  • the evidence collected after each run.

Avoid the easiest-looking process

The easiest demo is not always the best pilot. A task may be easy for a model but too infrequent to evaluate, too far from business value, or owned by no one.

Choose a process that is both bounded and consequential enough for the team to care about the result.

Use the free AI pilot priority scorecard to compare up to three candidates. Then continue through the pilot design and evaluation topic.

Continue reading
Pilot Design & Evaluation
How AI Agents Take Ownership of Real Work—and How People Reorganize Around Them Define a deliverable unit of work, expand agent responsibility through evidence-based authorization, and move human effort toward customers, products, judgment, and growth. Turning Freight Inquiries and Quotes into a First AI Workflow Start from the inquiry desk and separate field extraction, rule checks, exception handoff, and result write-back into a measurable workflow. How Should a Company Measure the Value of an AI Agent? Build an auditable evidence chain from work ownership and process speed to delivery quality, customer experience, and operating outcomes.