How to test an AI team before you pay.
Search for an honest review of any AI employee product and most of what comes back is paid. Vendors here run partner programmes that pay recurring revenue share for life, and they recruit course creators and agencies to write about them. The result is a review layer where the "brutally honest" verdict arrives with tracking parameters on every button. One such review admits mid-article that it was written by the product's own AI blog writer. Several share section headings word for word, which points to a shared content kit rather than independent testing.
Bias is not the worst of it. That corpus goes stale and nobody updates it. Reviews still selling "unlimited tasks" for a flat monthly fee are describing plans the vendor retired and replaced with capped ones. Read ten of them and you can come away confidently wrong about the price.
So run your own test. It takes an afternoon, and it works on any product in this category, ours included.
1. Check who is paying for the review
Hover the buttons. A query string like ?via=, or a discount code in the comments, means revenue share. That does not make the review wrong, but it does mean the author never had a reason to publish the failure. Then compare the prices quoted against the vendor's live pricing page. If they disagree, the review is stale, and so is everything else in it.
2. Read the help centre, not the landing page
Marketing pages describe the promise. Help articles describe the mechanism. That is where you learn what a credit costs, what happens when the balance empties, which features were quietly discontinued, which integrations actually work. When the pricing page and the help centre disagree, believe the help centre.
3. Cost one real task, not a demo
Pick a job you would genuinely hand over. Chasing last month's unpaid invoices, say, or triaging a day of enquiries. Run it end to end and note what it consumed. Multiply that by how often you would really do it. That number is what the product costs you, not the headline price. If the answer comes out at somewhere between twelve and fifty runs a month, you have found the ceiling before it found you.
4. Test whether it finishes
Ask it to send the email, book the slot, file the receipt, publish the post. Plenty of products stop at the draft and hand the last step back to you. That is a legitimate design choice, and it is how we handle outward-facing work too, but you need to know before you buy. "Prepares" and "does" are different products at the same price.
5. Break its memory on purpose
Tell it a specific fact about your business. Correct something it got wrong. Close the session. Come back the next day and ask a related question. If the correction survived, the product gets better the longer you use it. If it did not, you are the memory, permanently.
6. Ask it something it cannot know
This is the safety test, and it matters most if anything you buy will speak to your customers. A model's instinct is to be helpful rather than to admit ignorance. The documented failure mode is that AI is often most confident when it is completely wrong, with a voice agent cheerfully confirming a policy the business does not actually offer. The customer relies on the recording. The mismatch surfaces months later as a complaint or a legal letter. Ask your candidate an unanswerable question about your own business and see whether it says "I don't know" or invents something plausible.
7. Look for the audit trail
You want to see what it did while you were not watching: every message, every action, timestamped. Automations that fail quietly are the expensive kind. Builders swap stories about agents that spent money in the background for weeks before anyone noticed. No log, no purchase.
8. Check the analytics behind the claim
If the product reports on your marketing, click into a number. A headline reach figure with no breakdown underneath cannot inform a decision, and a report you cannot act on is a screenshot with extra steps.
9. Test whether it still sounds like you
This is the fear the research keeps finding: in a poll of 1,000 US adults, 65.5% of business owners worried that adopting AI would make their business feel less personal to customers, and a quarter said they had already lost business to customers using AI instead of paying them. Paste in three real replies you have written, ask for a fourth, and read it aloud. If you would not send it, the tool has a voice problem, and a voice problem is a churn problem with a delay on it.
10. Try to leave
Before you commit, find the export, find the cancel button, and find out what happens to your data and your content when you go. A product confident in its value makes both easy. Check on day one of the trial rather than month six.
How WorkMate OS answers these
Having published the test, here are our own answers to it, plainly:
No affiliate review layer. We would rather you ran the checks than
read a paid verdict.
No meter on the work. The crew's working day is in the plan — £19, £49
or £99 a month at launch. Only premium media generation, such as video, uses a separate
top-up balance, and it is stated up front rather than discovered in month two.
Memory that survives. One shared Brain the whole crew reads and writes,
so a correction sticks and compounds.
Draft, review, send. Outward-facing work is prepared by the crew and
approved by you, which is also our answer on hallucination. Nothing reaches a customer
without a human yes.
An activity log. You can see what ran, when, and what it produced.
Your voice, stored. Brand kit, tone and banned words live in the Brain,
not in a prompt you retype.
We are in beta, and we would rather be judged on that list than on adjectives. If something on it is not yet true for your setup, tell us and we will say so. You can check, which is the whole point of the test.
