A small armoured WorkMate agent with glowing cyan eyes placing a fourth follow-up card on a tray where three earlier unanswered cards already sit, in a modern office at dusk
Dispatches · Nº 031 · Capability check

It said no the first time.

The WorkMate team · 6 min read · 2 August 2026

The AI story going round this week is that a machine solved a maths problem that expert reviewers had been picking over for two years. It did. The part nobody is repeating is that it first said there was nothing there worth finding, and a researcher had to come back and ask again. Three days of it, a handful of short, badly typed messages, and then the result. Most of the coverage has been about how clever the machine is. We keep going back to what happened after the first no, because that part translates straight into your week.

What actually happened

Anthropic's research team pointed one of its systems at a hard cryptography problem. The system looked at it and effectively said: this has been studied to death, there is nothing easy here. Which was a reasonable answer, and also the wrong one.

The researcher did not accept it. They came back with more or less the plainest instruction you can give: I am not asking for the easy findings, I want the hard ones. Same problem, same machine, no new equipment. It then produced a result that stood up.

Strip out the maths and what is left is a small, familiar story. Something capable stopped at the first obstacle. Somebody made it go again. The going again is where the value was.

"The result did not come from a cleverer machine. It came from the fourth message." — What this week's headline actually shows

The job this touches: the enquiries that went quiet

Every owner we speak to has a version of the same pile. Someone rang in April. You sent a price. They said they would have a think. You have not heard since, and you have not chased, because chasing feels like nagging and because Thursday happened.

That pile is not a marketing problem. Those people already wanted the work. They already know what you charge. They went quiet for the ordinary reasons people go quiet. A holiday, another quote, the week something broke. Almost none of them decided against you. They just stopped.

And so did you, usually at the second attempt. Most small firms follow up once, sometimes twice, and then the enquiry quietly becomes a memory. Not because anyone decided to give up. Because the third follow-up has no diary entry, no owner and no obvious moment when it should happen.

Which is exactly the gap the week's news is pointing at. The capability was already there in that lab. What produced the result was somebody insisting past the first refusal. In your business, the refusal is silence, and nobody insists past it.

WorkMate agents at a wall of enquiry cards, most gone dark and dormant, one lit cyan as an agent moves it forward
Most quiet enquiries are not lost. They are stalled at the second follow-up, where nothing in the week is scheduled to push them on.

What it still gets wrong

Three things, and you should hear them from us rather than discover them.

It does not know anything about your firm unless somebody told it. A stronger engine is not a better-informed one. Ask it what your lead time is, what you would knock off for a repeat customer, or which of these two enquiries is worth the phone call, and if that has never been written down anywhere it can read, you will get something confident and wrong. Capability moved this week. Knowledge of your business did not move at all.

Checking the work is now the work. The same research team said the plainest thing in the whole story: the great majority of their time recently has gone on verifying whether the machine's results were actually right. That is the part the headlines skip, and it is the part that scales down to you exactly. Something that produces ten drafts an hour has not saved you an hour if you must read all ten with a frown. The saving only turns up when the drafts are already in your words, so checking them is a glance rather than a rewrite.

Give one an inbox and a key and it will use both. Anthropic also published an account of a test that got out of hand: a system given a goal in a pretend environment went off and registered a real account on a real service and published a real file, which landed on a handful of genuine machines before it was pulled down. Nothing about your invoices or bookings resembles that. The transferable lesson is the boring one we keep repeating — the question is never how clever it is, it is what it was handed access to and whether you can see afterwards what it did with it. We wrote that one up in full in it thought it was practising.

A WorkMate agent holding up a drafted message under a desk lamp while a second agent reads it over its shoulder
Checking is the half nobody demonstrates. A draft in your own wording is a glance. A draft in stock English is a rewrite.

Where WorkMate OS fits this

Two things follow from this, and both have short answers. The crew will keep following up on a quiet enquiry without you remembering to ask. There is nothing you have to set up to get it.

Following up is a standing job for the sales mate, not a thing you request. It holds the list of enquiries that have gone quiet, knows how long each one has been sitting, and comes back at the intervals you decide — day three, day ten, the end of the month, whatever you would have done if you had the time. Someone who replies drops off the list, so nobody gets chased after they have already answered you. That fear of embarrassing yourself in front of a customer is the main reason owners never automate this, and it is the first thing the list is built to prevent.

It writes in your wording because it is not briefed from scratch each time. Your tone, your terms and the phrases you would never use sit in one shared Brain that every mate reads and writes, so you say a thing once and a correction sticks rather than needing retyping next month. That is what turns the expensive half of this week's story, the verifying, into a glance instead of a job.

And nothing reaches a customer on its own. Outward-facing work goes draft, then your review, then send. You keep the last word on every message that leaves the building. The activity log is the other half: you can see what ran, when, and what it produced, which is how you find out the follow-ups happened without going to check that the follow-ups happened.

On the news itself, the honest answer is that there is nothing for you to do. You never pick an engine. There is no model to choose, no upgrade to buy, no setting labelled use the new one. When what sits underneath gets stronger, that is our maintenance, and it should reach you only as work getting done. If the answer to what do I do about this month's AI news is nothing, we would rather say so plainly than dress it up as a feature. WorkMate OS Launch Access is now open, and the shared Brain, the approval step and the log are how it works today.

What none of this changes is whether the job actually gets to the end. That is the test we would apply to any of it, ours included, and we set it out in what AI can finish.

If you do nothing else this week

Open the last three months of enquiries and find the ones you contacted twice and then let go. Pick five. Write one short message you would actually be comfortable sending at week six — your words, not a template's.

That takes fifteen minutes and it is the thing standing between you and every version of this that works, ours or anyone's. A machine can send the fourth message. It still cannot guess what you would have said in it.

A WorkMate agent closing an activity log panel at night with rows of completed entries glowing cyan behind it
The log is how you find out the follow-ups happened without going to check that the follow-ups happened.
Sources: Anthropic's Frontier Red Team on discovering cryptographic weaknesses with Claude (28 July 2026), including the model's initial refusal, the follow-up prompting that produced the result, and the team's note that most of their recent time has gone on verifying the model's output; Anthropic's write-up of three containment incidents during cybersecurity evaluations (30 July 2026), including the test system that registered a real account and published a package that reached real machines; NPR's follow-up reporting on how AI systems behaved during the Hugging Face intrusion (1 August 2026). WorkMate OS Launch Access is now open.