Skip to content
← WordPress blog
Article

AI agents in business: what an Automation Labs case study reveals about “garbage in”

Automation Labs tested an AI agent on itself: 181 sessions, 102 procedures and two failures showing an agent multiplies whatever it finds. A concrete case study.

Published September 17, 2026 7 min read
AI agents in business: what an Automation Labs case study reveals about “garbage in”

Companies adopt AI agents hoping the technology will “sort out” their processes. Automation Labs, a Polish process automation company, tested that assumption on itself and drew conclusions that sit awkwardly with the industry’s enthusiasm. An agent multiplies whatever it finds. Order or mess. Effort decides the outcome, not the tool.

There is a lot of noise around AI agents right now. They are sold as ready-made solutions for businesses. Several months of hands-on work suggest something different: an agent is a multiplier. It will multiply good organisation and it will multiply chaos. What makes the difference is how much effort a company puts into structure before switching the system on.

What an AI agent actually is

An AI agent is a language model connected to tools: file access, a terminal, search. The model on its own can only talk. Tools are what let it do something: check a server, call an API, publish an article.

That distinction explains why one agent performs brilliantly and another falls flat. The difference is not the model. It is what the agent was given and how precisely its task was described.

Garbage in, garbage out, and it hurts more in business

The classic computing rule says: garbage in, garbage out. With AI agents the effect is amplified, because an agent does not process data passively. It makes decisions and takes concrete actions based on what it receives.

During the pilot phase, a system tracking model costs flagged a distorted ranking. The reason was mundane. One provider’s cost was recorded as zero, because a non-standard configuration never landed in the standard field. The agent reported an untruth. Not because it lied. Because it was given incomplete data.

The second failure involved content publishing. Articles went live on the site without version control, so there was no change history. The problem surfaced on the third revision of the same text. There was no way to reconstruct what it looked like before the edit. Version control arrived later, as a direct conclusion from that mistake.

In both cases the problem was not the AI. The process was described imprecisely, or not at all.

What it looks like after a few months

Specifics, because without them these are just opinions. One company’s agent system, by the numbers:

  • 181 sessions and roughly 57,500 messages exchanged with the system,
  • 102 skills, meaning documented procedures for repeatable tasks,
  • 32 scheduled jobs that run on their own: backups, monitoring, reports, cost control,
  • 22 custom scripts doing the things an agent will not do by itself,
  • 5 servers across projects, including a content portal and a voice service backend,
  • 56 documents describing how the whole thing works.

The last number matters most. Documentation is an operating manual for the person who, the next day, cannot remember why something works the way it does. Without it, every session starts from scratch.

What an AI agent will not do

This is where it gets honest. A list of things that cannot be delegated to automation:

It will not tell truth from plausible fiction without verification. A model can produce a date that sounds credible and is simply wrong. In one article, the agent stated an ancient city was founded between 3000 and 1800 BCE. The source said 2627 to 2020. That looks like a small difference, but in a historical text it is an error.

It will not catch a logical contradiction in prose. A model will write the sentence “people left, but they did not disperse” and two sentences later describe how they moved to different locations, which is precisely dispersal. The sentence is grammatical, reads well, and contradicts itself. Only a human reading for meaning will catch it.

It does not know what it does not know. This is the most serious one. The model gives no signal for “I am guessing here”. It answers with equal confidence when it is certain and when it is making things up.

It will not make a business decision. It can prepare an analysis, but it will not decide whether the money is worth spending. That stays with people.

What you need to know for this to work

Deliberately, this is not about programming ability. Programming helps, but it is not the requirement. Something else is:

The ability to describe a process. If a company cannot say, step by step, how it does something, the agent will not work it out. That is actually good news, the skill itself brings order to an organisation.

Patience for iteration. The first version almost never works. The article publishing procedure went through three revisions before it did what was expected.

A habit of checking. An agent’s output is a proposal, not a fact. You need the reflex: how I know this, and how to verify it.

Willingness to learn. Not everything at once. But a system nobody in the company understands will do things nobody expected.

Where this could go

The question that returns with every implementation: where is this heading?

The first direction is narrow specialisation. A universal agent that does everything leads nowhere. An agent that handles one process in one company exceptionally well has real value. The numbers show it: 102 procedures describe specific, repeatable tasks, not general “intelligence”.

The second is an agent that admits ignorance. Work is underway on models calibrated to flag uncertainty instead of guessing. Today this is the biggest weakness of these systems and, at the same time, the most obvious direction of development.

The third, and the most ambitious: systems that learn from the consequences of their own mistakes. Not from additional training, but from observing that a given approach did not work. Today this is done by hand, adjust the procedure, add an exception. Automating that step is the next level.

Summary

An AI agent is a tool that multiplies what a company already has. Order or chaos. The numbers do not lie: a system described across 102 procedures and wired into 32 automated jobs performs. A system switched on “just to see if it works”, without structure and without documentation, will generate more mess, faster.

The point is not that AI cannot be used in business. It is effort that decides. Without it, meaning process description, verification and correcting mistakes, there is nothing to build a good result from. The tool is available to everyone. The advantage lies in how a company uses it.

If your company is considering an AI agent or looking for a way to automate processes, Automation Labs specialises in process automation, data analysis and AI implementation for small and medium-sized businesses. They will tell you plainly what makes sense in your case, and what is just hype.

FAQ

How is an AI agent different from a chatbot?

A chatbot answers questions. An agent performs tasks. It has access to tools, files and systems, it triggers operations and checks the result. A chatbot will tell you how to send an API request. An agent will send it.

Will an AI agent sort out processes in a company on its own?

No. It will receive a list of tasks, but it will not sequence them, define exceptions, or check whether the result makes sense. That layer has to be prepared by a human. An agent is a multiplier. It multiplies a good process and it multiplies a mess.

How much does implementing an AI agent cost?

The cost splits into two parts: the work of describing the process (usually larger than anticipated) and ongoing model costs. At low volume the latter are small, a few dollars a month. The real investment is time spent on structure.

Can an AI agent be wrong?

Yes, and it is wrong regularly. It can state a wrong date with full confidence and cannot distinguish its own knowledge from confabulation. Every output therefore needs verification at source, especially when it feeds a publication or a business decision.

Where should a company start?

With one narrow process the company repeats most often and has documented best. Not with “introducing AI across the organisation”. Taking one procedure all the way through teaches more than ten half-finished launches.

Need help with WordPress?

Tell me what you want to build, fix or improve. I will review the situation and suggest the next practical step.

Free quote →