Back to Artificial intelligence

Artificial intelligence

Off-the-Shelf AI Tool or Integrated Assistant: When You Actually Need to Customize

Decide whether an available product covers your need or requires specific integrations.

SqualiOnline editorial team · 2026-09-07

The question that leads you astray is "which AI is the best." The one that leads to a decision is more boring: does this task, done with the tool the company already pays for, produce a usable result? If the answer is yes, there's nothing to build. If it's no, the reason for the no tells you precisely what needs to be built.

Almost every evaluation runs aground because it compares products instead of tasks. A product is judged in the abstract, and the one with the best demo always wins; a task is judged on the result, and you're the one who knows how to recognize that result.

Define the task first, then look at the tools

A task worth evaluating is specific, recurring, and has a recognizable outcome. It needs to be written down, in four lines.

  • What goes in: a message, a document, a recording, a question.
  • What needs to come out and in what form: a text, a row in a system, a classification, an answer to a customer.
  • Who does it today, how many times a week, and how long it takes them.
  • What makes a result acceptable and what makes it unacceptable. This is the part almost nobody writes down, and it's the only one that makes a comparison possible.

"Use artificial intelligence for customer service" isn't a task, it's a category. "Classify incoming requests and draft a reply for the three most frequent categories" is one, and it can be tested in a week.

The test: same case, same criteria, same people

A comparison only makes sense if it's actually comparable. Here's a minimal protocol that requires no technical expertise.

  1. Collect twenty or thirty real past cases, including the difficult ones and the ones where the person got it wrong. Only easy cases produce a false conclusion.
  2. Decide in advance how you'll judge it: which errors are acceptable and which aren't. An error the recipient wouldn't notice and one that sends a wrong figure to a customer don't carry equal weight.
  3. Give the same material to every solution being tested, instructions included. If one gets better guidance than the others, you're measuring whoever wrote the instructions.
  4. Have whoever does that job today do the judging, not whoever proposed the tool.
  5. Log the review time too. A solution that produces slightly better results but needs to be checked line by line can end up costing more than the work it replaces.

The four checks that separate an off-the-shelf product from an integrated system

When testing with the generic tool isn't enough, the reason almost always falls under one of these four headings. Recognizing which one keeps you from building more than you need to.

CheckConcrete questionIf it's missing
SourcesCan it access your documents, price lists, and up-to-date history?The answers are plausible but not yours: a connection to your sources is needed
AccessWho sees what, and does confidential data stay confidential?The product can't be used on data that isn't allowed to circulate
ActionsDoes it just need to answer, or also write into one of your systems?Integrations and controls are needed: it's a project of a different nature
LimitsDoes it hold up under the volumes, formats, and language you actually use?It works in testing and buckles in daily use: it needs to be checked under real load

Many situations get resolved with a standard product plus a connection to the right sources, without building a system at all. It's the solution nobody proposes, because it's the least sellable, and it's often the most sensible.

What an off-the-shelf product does better than custom development

It's worth listing, because in the rush to customize, real advantages get thrown away.

  • It improves without you having to do anything, and it doesn't depend on a single person who knows how it's built.
  • It costs less to get started and lets you change your mind. A trial that ran for three weeks and was abandoned is a good outcome, not a waste.
  • It's already been used by many others: the obvious problems have surfaced elsewhere, not on you.
  • It doesn't generate maintenance for you. Everything built custom needs maintaining, and the maintenance never ends.

Custom development is justified when the task depends on data, rules, or systems that are yours alone, when the data can't leave, or when the same operation repeats often enough to make manual work unsustainable.

The cost you don't see: supervision and maintenance

The honest comparison isn't between a product's subscription fee and a development quote. It's the recurring items that decide, and they apply to both paths.

  • Who checks the results, how often, and for how long before the checking eases off.
  • Who updates the sources when price lists, procedures, or documents change. A system that answers based on outdated documents is worse than one that doesn't answer at all.
  • Who steps in when a connection breaks or an access credential expires, and how quickly.
  • What happens if the person handling it isn't around. This applies to the vendor just as much as to you.

Off-the-shelf and custom can coexist

The choice is rarely all-or-nothing. The most stable solution keeps standard everything that doesn't set you apart, and builds only the piece that depends on you: your sources, your rules, the connection to your systems. It's also worth writing down from the start what would stay yours if you switched tools — the documents prepared, the instructions, the test cases, the data collected — because that's what lets you change your mind without starting over from scratch.

What this guide doesn't cover

This guide covers the choice between what's available and what needs to be built, for your specific case. You won't find model rankings or general product comparisons here: they change over time and wouldn't be comparable to your work. The distinction between traditional automation and artificial intelligence, and the cost items for maintaining an assistant over time, are covered elsewhere.

Frequently asked questions

How long should a test run?

Long enough to run into the difficult cases, which are usually about one in ten. With twenty or thirty real cases, you reach an answer within one or two weeks; beyond a month, the test stops being a test and becomes undecided use, with the risk that nobody ends up owning the final call.

Can we use our own documents with a generic tool?

It depends on the service's terms and the type of data. Before uploading anything, two points need checking: what the vendor states it does with the material you send it, and whether those documents contain personal data or confidential information about your customers. For testing, you can almost always work with anonymized material.

Is it worth waiting for the tools to improve?

Waiting doesn't produce any of the things you need regardless: organized documents, test cases, acceptance criteria, people who know how to recognize a good result. That work stays valid with any tool, and whoever has already done it adopts the next tool in days instead of months.

We compare the tools available with custom development.

If you’d like to talk it through, the service that handles this is Artificial intelligence.

Related guides