How to test an AI agent in my company before investing in a full implementation

29 SEPT, 2026
cover

Let's say you have a fairly clear idea.

You want to use an AI agent for customer service, sales, support, administration, or an internal process.

The idea sounds good.

We may even know that, technically, it can be done.

But a fairly logical question comes up:

What if we make the full investment and then, in practice, it doesn't work as expected?

To me, that's a very good question.

Because one thing is seeing a demo where everything works perfectly.

And another very different thing is putting that agent in front of your real customers, employees, data, and processes.

That's why I wouldn't necessarily start with a full implementation.

I would start with something much smaller:

a specific problem, a first version, and clear metrics to decide whether it's worth continuing.

First: don't try to automate the entire company

This is probably one of the easiest mistakes to make.

We start by thinking:

“We want to incorporate AI.”

And we quickly end up talking about:

  • customer service;
  • sales;
  • collections;
  • support;
  • administration;
  • reports;
  • internal processes.

All at once.

And the project becomes huge before we've tested anything.

To validate an AI agent, the more specific the first problem is, the better.

Not:

“We want to use AI in sales.”

Yes:

“We want to know whether an agent can receive inquiries, gather initial information, and send salespeople only the leads that meet certain criteria.”

Now we have something we can test.

Choose a problem, not a tool

Before talking about models, agents, prompts, or integrations, I would start with a much more boring question:

What problem do we want to solve?

Because we can build an amazing agent that doesn't solve anything important.

And then we have a nice AI project that nobody needs.

A good first use case usually has several of these characteristics:

  • it happens frequently;
  • it consumes time;
  • it has a certain degree of repetition;
  • we can measure it;
  • the cost of making a mistake is manageable;
  • we have data or information to solve it.

For example:

  • classifying inquiries;
  • answering frequently asked questions;
  • summarizing sales opportunities;
  • reading documents;
  • preparing follow-ups;
  • generating reports;
  • searching internal information.

As with Product Discovery, the test shouldn't prove that Artificial Intelligence “is good.”

It should prove something much more specific:

that it works for that process inside your company.

Let's think about a sales example

Suppose a company receives many inquiries through WhatsApp.

The final goal could be to have an agent that:

  • answers inquiries;
  • qualifies leads;
  • checks the CRM;
  • schedules meetings;
  • creates opportunities;
  • follows up;
  • hands conversations off to salespeople.

Perfect.

But we don't need to build all of that to know whether the idea makes sense.

We can start with a first version that only does this:

  1. receives the inquiry;
  2. understands what the person needs;
  3. asks a few questions;
  4. classifies the lead;
  5. prepares a summary for the salesperson.

That's it.

Then we test it with real inquiries.

And that's when the interesting questions begin:

Did it correctly understand what the person wanted?

Did it ask the right questions?

Did the salesperson receive better information?

Did we save time?

Did users finish the conversation or leave halfway through?

If it works, we continue.

If it doesn't, it's better to find out here than after building everything else.

“I prefer to start with an agent that does one thing very well rather than build one that promises to do everything and then needs intervention at every step.”

Santiago Sola — Tuxdi

Another case: an internal agent

Suppose you want to create something like:

“An internal ChatGPT that answers any employee question.”

It sounds great.

It's also huge.

Because “any question” can involve Human Resources, administration, sales, operations, legal, and twenty other things.

To test it, we could start only with:

questions about sales processes.

We give it access to:

  • procedures;
  • documents;
  • policies;
  • frequently asked questions.

And we give it to five or ten people on the team.

Then we see what happens.

What do they ask?

What does it answer well?

Where does it make mistakes?

What information is missing?

Do they actually use it?

Did they stop asking the team certain questions?

In a few weeks, we'll probably know much more than after months of discussing what the ideal agent should look like.

Before testing it, define what “works” means

This is key.

Because if we don't define this beforehand, we end the pilot saying things like:

“I thought it was good.”

“It answered pretty well.”

“There are a few things to improve.”

That isn't very useful for deciding whether to invest more money.

Before starting, we should define what we want to measure.

For example:

“We want the agent to correctly resolve at least 60% of these inquiries.”

Or:

“We want to cut in half the time salespeople spend qualifying leads.”

Or:

“We want it to process documents in under one minute with an error rate lower than the current one.”

It doesn't have to be perfect.

It has to be measurable.

“A pilot shouldn't prove that AI can do something. It should prove that it makes sense to do it inside your company.”

Santiago Sola — Tuxdi

What should you measure?

It depends on the agent.

If it serves customers

I would look at:

  • how many inquiries it resolves;
  • how many it hands off;
  • how long it takes to respond;
  • how many customers abandon the conversation;
  • what errors appear.

If it automates an administrative process

I would look at:

  • how much time it saves;
  • how many cases it processes;
  • how many exceptions it needs;
  • what percentage requires human intervention;
  • how many errors it generates.

If it works in sales

I would look at:

  • how many leads it processes;
  • how many it manages to qualify;
  • how much time it saves the salesperson;
  • what information it obtains;
  • how many end up moving forward.

If it's an internal agent

I would look at:

  • what percentage of inquiries it answers;
  • how reliable the answers are;
  • how much it is used;
  • what questions it cannot answer;
  • how much work it takes off the team.

The right KPI depends on the problem.

Also measure where it makes mistakes

This is just as important.

In a demo, we normally show the cases where it works.

In a pilot, we need to actively look for the cases where it doesn't work.

We want to know:

  • when it makes up information;
  • which instructions it misinterprets;
  • when it hands off too often;
  • when it should have handed off but didn't;
  • which data it confuses;
  • where it needs more context.

Because that's exactly why we're testing.

An error during a controlled pilot is information.

An error after deploying it at scale can be a problem.

Test it with real people

This also changes things a lot.

When we test an agent, we know how to ask it questions.

The real user doesn't.

They'll write:

“hi this isn't working”

They'll send a voice message.

They'll ask two things at once.

They'll explain the problem poorly.

They'll change the subject.

They'll send “???” because it took five seconds.

And that's fine.

That's exactly what we want to test.

A demo proves that it technically works.

A pilot proves that it works in the real world.

Want to validate an AI agent before scaling it?

We can help you define a specific, measurable pilot.

You don't need to launch it across the whole company

Another mistake would be going from:

nobody uses the agent

to:

300 people use it tomorrow.

There's no need.

We can start with:

  • one salesperson;
  • one department;
  • ten employees;
  • a small group of customers;
  • 10% of inquiries.

If something goes wrong, we fix it.

If it works, we expand it.

For example:

First week: internal team only.

Second week: small group of users.

Third week: more volume.

There's no need to go from zero to one hundred.

At first, a person should supervise quite a lot

Especially during the first tests.

We can have the AI:

suggest → a person validates.

Then, once we see that certain cases work consistently:

the AI executes them automatically.

This lets us increase autonomy gradually.

For example, a sales agent can start by generating:

“I think this lead should be classified as an opportunity.”

A person confirms it.

After reviewing hundreds of cases, we may discover that this classification works very well.

Then we can automate it.

We don't have to trust it blindly from day one. It's also worth defining which tasks should not be automated with AI.

Don't fall in love with the fact that “it answers nicely”

This point is extremely important to me.

An agent can be amazing at conversation.

It can write perfectly.

It can seem super intelligent.

And still generate zero value for the company.

If after implementing it:

  • nobody uses it;
  • it doesn't save time;
  • it creates more work;
  • it needs constant corrections;
  • it doesn't improve any metric;

then it didn't work.

Even if the demo was impressive.

The question isn't:

“How intelligent does it seem?”

The question is:

“What changed because of the agent?”

A good pilot can also end with “no”

This can be hard to accept.

Let's say we test an agent for a month.

We measure it.

And we discover that:

  • it only automates 10% of cases;
  • it needs too much intervention;
  • the savings are low;
  • the cost isn't justified.

Perfect.

We don't scale it.

That doesn't necessarily mean the pilot failed.

Quite the opposite.

It prevented us from building a much more expensive solution only to discover exactly the same thing six months later.

A pilot can also tell us:

“It doesn't make sense to keep investing here.”

Define the criteria before you start

We can keep it fairly simple.

For example:

We're going to test this agent with 300 conversations.

And we define:

We continue if:

  • it correctly resolves at least X%;
  • it reduces response time;
  • it keeps critical errors below a certain level;
  • the team genuinely finds it useful.

We don't continue yet if:

  • it needs too much intervention;
  • it generates significant errors;
  • it doesn't save enough time;
  • users don't adopt it.

Now we have an experiment.

Not just a feeling.

“Scaling a solution before measuring it means taking unnecessary risk. First we need evidence that it works; then it makes sense to invest in making it bigger.”

Santiago Sola — Tuxdi

How much does it cost to test an AI agent?

Obviously, it depends greatly on the use case.

But one of the advantages of running a pilot is precisely that we don't need to build, from day one:

  • every integration;
  • every permission;
  • every scenario;
  • all the infrastructure;
  • every channel.

We can start with the core.

For example:

an inquiry comes in → AI interprets it → retrieves information → prepares a response → records the result.

If that works, we move forward.

Then we add:

  • more systems;
  • new actions;
  • more users;
  • WhatsApp;
  • CRM;
  • automations;
  • advanced permissions;
  • greater autonomy.

The investment grows as our certainty that it's worthwhile grows too.

So, how would I test an agent?

I would try to complete this sentence:

We want to test whether an agent can ________ to improve/reduce ______. We're going to test it with ______ and measure ________.

For example:

We want to test whether an agent can receive sales inquiries and gather initial information to reduce the time salespeople spend on poorly qualified leads. We're going to test it with 200 conversations and measure accuracy, time saved, and the number of leads correctly handed off.

That's already much more specific than:

“We want to implement AI in sales.”

What happens next?

There are three possibilities.

It works

Great.

We scale.

More users.

More integrations.

More cases.

More autonomy.

It partly works

That's probably the most normal outcome.

We discover:

  • what information is missing;
  • which instructions need improvement;
  • what a person should continue doing;
  • which cases are worth automating.

We adjust and measure again.

It doesn't work

We stop.

Or we look for another use case.

Maybe the problem we chose wasn't the right one.

That's also a valid answer.

“The biggest implementation isn't always the best decision. Often, the best decision is to find the smallest test that lets us know whether it's worth continuing.”

Santiago Sola — Tuxdi

Frequently asked questions

What's the difference between an MVP and an AI proof of concept?

A proof of concept mainly seeks to validate whether something is technically possible. An MVP seeks to put a first functional version in front of real users or processes and measure what happens.

Do I need to connect all my systems to test an agent?

No. Only the ones needed to validate the first use case.

Does the agent have to execute actions automatically?

No. It can start by only recommending or preparing actions for a person to validate.

How long should a pilot last?

It depends on volume. Rather than thinking only in days, it's better to define a sample: a certain number of conversations, documents, users, or operations.

What happens if the pilot fails?

You analyze why. The agent may need adjustment, the process may need to change, or you may simply decide that the use case doesn't justify the investment.

Have an AI use case but don't know where to start?

We define the scope, metrics, and a first test.

Before investing big, get evidence

To me, this is the most important idea in the entire article.

You don't need to start by thinking:

“We're going to implement an AI agent across the whole company.”

You can start by thinking:

“There's a problem here. Let's see whether an agent can actually solve it.”

Choose a use case.

Build something small.

Put it in front of real situations.

Measure.

Correct.

Then decide.

It's much less spectacular than announcing a “complete transformation with Artificial Intelligence.”

But it's probably a much smarter way to invest.

At Tuxdi, we work with this logic: start with a specific scope, validate, and scale only when we see real value.

If you have an idea for implementing an AI agent but still don't know whether it justifies a full project, tell us which process you want to improve. We can help you turn that idea into a concrete test, with a clear scope and metrics to decide whether it makes sense to move forward.

Want to test an AI agent with a real use case?

Contact us

let's work together

You are one step away from taking your project to success

Tuxdi LLC+54 (249) 469 8992[email protected]

2201 Menaul Blvd NE STE Albuquerque, NM 87107