Back to Artificial intelligence

Artificial intelligence

AI Pilot Project: How to Define the Goal, the Sample, and the Final Decision

Test a use case before rolling it out across the whole company.

SqualiOnline editorial team · 2026-09-07

A pilot project exists to answer a question, not to prove that artificial intelligence works. The difference shows up at the end: if the criteria were written beforehand, the result is a decision; if they're written afterward, it's an interpretation, and whoever argues best in the meeting wins.

This guide covers how to organize the trial: what goes into it, what gets measured, who oversees it, and how the final decision gets made.

First thing: the question the pilot answers

"Let's see if AI can help us" isn't a question, because there's no answer that closes it. A useful question has an activity, a subject, and a verifiable condition.

  • If a system drafts the first replies to support requests, can the office close out the day's cases by end of day?
  • If vendor documents are read automatically, how many still end up in a person's hands anyway?
  • If repetitive quotes are drafted with assistance, does preparation time go down without an increase in errors on the amounts?

Each of these questions can get a negative answer, and that's the point. A pilot that can't have a negative outcome isn't a trial: it's a presentation.

The scope: what's in and what stays out

This is the part people tend to leave broad, and it's the mistake that makes results unreadable. The wider the scope, the less clear it is what to attribute the outcome to.

  • The activity: a single, well-defined one. Not "customer support," but "order-status requests that arrive by email."
  • The users: a small group, named individually, willing to flag problems rather than quietly work around them.
  • The data: which sources, with what permissions, and what must not be used.
  • The explicit exclusions: the cases the system shouldn't handle and that must go straight to a person, for example complaints, contractual matters, customers in dispute.

The time period should be chosen so it contains enough cases to say something meaningful, and at least one typical cycle of the company: if the monthly closing changes how people work, the pilot needs to run through it.

The starting situation gets recorded first

Without a starting point there's no comparison, and reconstructing it afterward always produces convenient numbers. Before switching anything on, measure how things work now: how much time the activity takes, how many cases get handled, how many errors or reworks occur, how long people wait.

If a measurement doesn't exist, it needs to be built by hand for a few weeks: two people logging times and cases on a sheet give a more useful baseline than an estimate pulled from memory. It's real work and needs to be in the plan, not sprung on people as a surprise.

The pilot sheet

It fits on one page. The example below is illustrative and concerns the automatic reading of incoming documents.

FieldHow it's filled in
QuestionCan vendors' delivery documents be recorded without manual data entry, while keeping the same level of accuracy?
ActivityRecording documents that arrive by email; paper documents stay out of scope.
UsersTwo people from the purchasing office, named individually.
DataDocuments from recurring vendors over the last few months. No data relating to staff or customers.
Out of scopeInvoices, credit notes, documents in a language other than Italian.
Starting situationRecording time and number of subsequent corrections, tracked by hand for three weeks before launch.
Acceptance criteriaRecording time goes down, corrections don't increase, every unrecognized document is passed to a person with a notification.
SupervisionEvery result stays checked by a person for the entire length of the trial.
Duration and close-outA defined period that includes at least one monthly closing; the decision meeting scheduled on the calendar right from the start.
OwnerThe purchasing manager, with a technical contact for anomalies.

The two rows that are almost always filled in worst are "out of scope" and "acceptance criteria": the first because it feels limiting, the second because it forces you to say in advance what counts as success. They're exactly the two that make the trial decidable.

Supervision and anomalies during the trial

During the pilot, someone has to watch, not just use it. You need a fixed check-in — weekly works — where you go through the cases that went wrong and decide whether they're fixable or whether they point to something deeper.

  • Every result stays checked by a person: the pilot isn't the moment to remove oversight.
  • Anomalies are collected in a single list, with the concrete case attached: without the real example, nothing gets fixed.
  • What would stop the trial needs to be defined in advance: an error that reaches a customer, data that gets out where it shouldn't have, a volume of corrections that exceeds the work saved.
  • Changes made during the trial should be logged with a date, otherwise at the end nobody will know which version the results refer to.

Deciding: extend, fix, or stop

The final meeting should be scheduled at the start, with the decision-makers already committed to attend. There are three possible outcomes, and none of the three is a failure.

  1. Extend: the criteria are met. Next comes defining who maintains the system, who's accountable when it's wrong, and what happens if a data source changes.
  2. Fix and repeat: the result is close but one specific case doesn't hold up. Run it again just once, with the same scope and one declared change.
  3. Stop: there's no benefit, or it costs more than doing it manually. Write down why, so that in a year nobody proposes the same trial again without knowing how it went.

The full cost needs to go into the count, not just the build: maintenance, people's checks, the time spent correcting. A system that saves half an hour and requires twenty minutes of checking has a thin margin, and that needs to be said before someone discovers it on their own.

What this guide doesn't cover

This guide covers organizing the trial. Choosing which activity to bring to the pilot — how the candidates are compared against each other and by what criteria they're ranked — is an earlier step and has a guide of its own. Testing a conversational assistant before making it available to customers also follows a dedicated method.

Frequently asked questions

How long should a pilot project last?

As long as it takes to gather enough cases and get through at least one typical cycle of the company. A duration decided in advance, without looking at the volume of work, produces results that end up impossible to read.

Who should take part in the trial?

A few people, named individually, who actually do that activity and are willing to flag problems. Too large a group makes it impossible to understand what happened; a group made up only of enthusiasts gives back nothing but good news.

What do you do if the pilot goes badly?

Write down why, with the concrete cases that didn't work, and choose between repeating it just once with a declared change or stopping. A negative trial that's closed out properly keeps the same idea from coming back in a year with no memory of what happened.

We define a measurable AI pilot project.

If you’d like to talk it through, the service that handles this is Artificial intelligence.

Related guides