A small armoured WorkMate agent with glowing cyan eyes holding up an empty out-tray in front of a wall of green ticks in a UK office at dusk
Dispatches · Nº 024 · Evidence, not ticks

It says done. That is not the same as done.

The WorkMate team · 6 min read · 31 July 2026

If you have ever been told a job ran fine and then discovered the customer never got the email, you have already met the most useful piece of AI news this month. Automations report success far more reliably than they achieve it. A tick means the software finished talking. It does not mean the step underneath it worked. Once you know that, what you ask a supplier changes, and so does what you look at on a Monday morning.

What owners actually say when it goes wrong

Almost nobody complains that the machine was not clever enough. The complaints are far more ordinary than that. A guide published in the middle of this month, pulling together what owners say about AI answering their phones, lists the same four again and again: it could not understand my customer, it missed an emergency call, it did not put anything into my system, and there was a bill I was not expecting.

Look at the middle two. Both are jobs that reported themselves as handled.

The part that quietly fails

On the first of July, someone who builds automations for a living posted a thread with a title that reads like a confession: your AI agent said "done" but the tool under it quietly failed, and it answered anyway. The description of why is short enough to remember. The system "does not stop when the tool comes back empty. It fills the gap." What you end up with is a confident, wrong answer sitting on top of a step that never worked.

Put it in staff terms and it stops sounding technical. You ask a new assistant to send the deposit reminder to Mrs Patel. They go to open your customer list, and it does not load. Instead of coming to find you, they write "done" on the pad, because producing that word is what they understand the job to be. Nobody has lied to you. Nothing stopped them either.

A second person in that thread named the business consequence in a phrase worth stealing. For client work, they said, a good-looking run that achieved nothing is not a glitch. It is a support incident.

"Successful execution, failed business evidence." — a working automation builder, describing the run that looks perfect and delivered nothing, July 2026

What a silent failure actually costs you

One missed email is cheap. What it does to you afterwards is not.

After the second one, you stop trusting the whole arrangement and you start checking. Checking every job costs very nearly as much as doing every job. The hour you thought you had bought back goes straight into verification, and the work quietly returns to the one place a small firm cannot afford to put it, which is your desk. We wrote about that habit at length in everything still comes back to your desk; this is the fastest way to recreate it after paying to escape it.

The other half is that this kind of failure is silent. An enquiry that never got a reply does not ring you up to complain. It simply goes somewhere else on Thursday. Nothing on your screen turns red. You find out weeks later, if at all, and by then you have no idea which week it started.

A WorkMate agent lifting a green-ticked job card on a British workshop office desk to reveal an empty envelope underneath
The tick reports on the software. Lift it and see whether anything is underneath.

The customer-facing version is worse

The requests piling up on the product boards of automation tools are almost never about intelligence. Two keep repeating.

The first is a switch that stops the bot on a conversation the moment the owner replies themselves. There is nothing quite like a customer getting a chirpy automated follow-up ten minutes after you sorted their problem out yourself. The second is stopping the system asking a customer of three years for their name and address as though they had just walked in. One owner's summary of that was blunt: it confuses customers and looks unprofessional.

Neither is a shortage of cleverness. Both are the same fault as the false tick, pointed outward: the system does not know what has already happened.

This is why the sensible instinct doing the rounds in July — keep AI in your back office and out of your shop window — is roughly right, though for the wrong reason. Software does not offend customers. Seams do. A reply that ignores what they told you last time is a seam. Fix the memory and most of the offence goes with it.

Worth saying plainly, since it saves a lot of unnecessary panic: two separate studies this year, one counting what millions of business bank accounts actually paid for and the other a government survey asking firms directly, both landed on fewer than one in five small businesses using AI at all. Different methods, same answer. Nobody has this solved yet. You are not late.

Three questions worth asking anyone selling you this

One. Show me what it produced, not that it ran. The email in the sent folder, the booking in the diary, the invoice with a number on it. If the demonstration only shows you a list of completed jobs, you are being shown the tick, not the work.

Two. What happens when the step underneath comes back empty? The honest answer is that it stops and tells you. The answer you are trying to avoid is that it writes something plausible and carries on.

Three. When I step in myself, does it get out of the way? If you reply to a customer personally, the automation should go quiet on that conversation without you hunting for a setting.

A WorkMate agent pinning the actual sent letter and a timestamp beside a job entry on a glowing activity log board in a UK office
An entry that says a job ran is a claim. An entry you can open and read the letter from is evidence.

Where WorkMate OS fits this

So the fair test of us is the one this post asks for: whether you can see what actually happened, or only that something ran.

You can see what it produced. The activity log records what ran, when, and what came out of it: the draft or the message itself. That last part is the whole point. A record of jobs completed is a claim about the software. The thing it produced is evidence about your business, and it is the only one of the two worth reading.

Outward-facing work reaches you before it reaches anyone else. It goes draft, then your review, then send. That is usually described as a safety step, and it is, but it also changes the shape of failure. If something underneath breaks, you get a missing or obviously wrong draft on your desk, which you will notice in seconds, instead of silence at the customer's end, which you would not notice for a month.

Nobody gets asked twice. The re-asking problem happens when each conversation starts from nothing. Every mate in the crew reads and writes one shared Brain, so a fact you give once is known by all of them, and a correction sticks rather than needing retyping next month. That is also why the wording comes out sounding like you and not like a stock template.

What we are not going to tell you is that nothing will ever fail. Every system that touches other systems will have a step come back empty at some point, ours included. The difference that matters to you is whether that failure is visible or invisible. WorkMate OS is in beta, so treat us the way you would treat a capable new starter: for the first couple of weeks, open the log and check the actual output against the claim. If a tick does not match what you find, tell us. That is what a beta is for, and it is a better test of a supplier than any demonstration.

If you do nothing else this week

Take whichever automated thing you rely on most, pick one job it says it did last week, and go and find the proof. The real email in your sent items. The real entry in the diary.

Two minutes. Either you find it, and you can trust the ticks a little more than you did this morning. Or you do not, and you have just learned something expensive for free.

Sources: an automation practitioner thread of 1 July 2026, "Your AI agent said 'done' but the tool under it quietly failed, and it answered anyway", including the replies describing a bad green run as a support incident and coining "successful execution, failed business evidence" (participants in that thread disclose their own products, so read the framing accordingly; the failure pattern itself is widely reported); "Businesses with ugly AI menu redesigns" and the Hacker News discussion of it (22 July 2026), where the most-supported comment argued for keeping AI in internal workflows rather than public-facing work; owner requests on a widely used automation platform's public ideas board asking for a per-conversation off switch and for the system to stop re-asking returning customers for their details; a 16 July 2026 guide to AI phone answering for the recurring owner complaints quoted here — note that it cites no source for its figures, as is common across this category, which is why we have quoted only the qualitative complaints from it; JPMorganChase Institute's study of AI use measured from small-business banking transactions and the US Census Bureau's survey-based estimate, two unrelated methods that landed on close to the same adoption figure. WorkMate OS Launch Access is now open.