- In Carnegie Mellon's TheAgentCompany benchmark, the strongest agent finished 30% of the tasks set inside a simulated company, and just 8.33% of its finance tasks.
- An AI employee is an AI agent sold under a job title. In UK law it isn't an employee, and the business answers for everything it says and does.
- Hire one for named tasks, never a whole role: rewrite the job description as a task list, then let it prove itself in shadow mode before it acts alone.
AI employees are AI agents packaged and sold as if they were staff: a name, a job title such as sales development rep or bookkeeper, and access to your inbox, CRM or accounts. Underneath, each one is a language model with instructions, tools and permissions. The best independent test of whether agents can do office jobs, Carnegie Mellon and Duke's TheAgentCompany benchmark, found the strongest one completed 30% of the tasks on its own. So buy an AI employee for a short list of named tasks, not for a role.
We build agents for small firms, and the guides on this site come out of a pipeline of agent workers we run ourselves. We have every reason to like the idea. This is our view of the category anyway, including the parts the sales pages leave out. If the underlying technology is new to you, read what agentic AI is first.
What an AI employee actually is
The phrase is a sales category, not a technical one. A vendor gives an agent a persona and a job title, prices it like a seat and presents it as a hire. Strip the persona away and the parts are the same as any agent: a model, a written brief, connections to your systems, and rules about what it may do without asking.
That gap matters. A job title promises the whole job, including the awkward bits nobody ever wrote down. An agent does what it has instructions and tools for, and nothing else.
Three questions the label tends to hide:
- Who checks its work, and how often?
- What can it touch? Reading the CRM is one thing. Sending email, issuing refunds or posting to the ledger is another.
- What does it do when it isn't sure?
Gartner's warning about agent washing bites hardest here. It estimates only about 130 of the thousands of agentic AI vendors are real, and a chatbot with a first name and a headshot passes for an "employee" very easily. Our guide to choosing an agent builder lists the questions that expose it, and AI agents for business covers how a genuine agent differs from a chatbot or a fixed workflow.
What happened when agents were given office jobs
TheAgentCompany is the closest thing to a job trial for AI employees that we know of. Researchers at Carnegie Mellon University and Duke built a simulated software company, with its own intranet, shared files, team chat and simulated colleagues, then set 175 tasks across six kinds of work. The paper went through peer review for the NeurIPS 2025 datasets and benchmarks track.
The best model completed 30% of the tasks, or 39% once partial credit is counted. Each task took it almost 27 steps on average. Split by kind of work, the picture is stranger than the headline.
| Kind of work | Tasks the best agent completed | Read across to an AI employee |
|---|---|---|
| Project management | 39.29% | The best result in the test, and still under half |
| Software engineering | 37.68% | Work with clear finish lines and tests that prove it ran |
| Human resources | 34.48% | About a third, close to the overall rate |
| Data science | 14.29% | Analysis that needs the right data found first |
| Administration | 13.33% | The assistant and office manager jobs the label suggests |
| Finance | 8.33% | The worst of the six, and the job a "bookkeeper" persona claims |
The authors put the surprise plainly: "software engineering tasks, which may seem like much harder tasks for many humans, result in a higher success rate." The admin and finance samples are small, so read those two rows as a direction rather than a precise rate.
Two more honest caveats. This is one simulated company with tasks its authors wrote, not your business. And models improve quickly, so the rates describe the models tested in 2025, not the ones on sale next year. Even so, it's the only public test we know of that scores agents job by job, and the order of the rows matters more than the decimals.
Why the office jobs scored worst
What follows is our reading, not the paper's finding. Admin and finance work is mostly glue. Chase a colleague for a figure. Find the latest version of a spreadsheet. Decide which of two conflicting numbers is right, then tell someone. None of that comes with a test that says "done".
Code does. A task either runs or it doesn't, and the agent can check its own work before handing it over.
That's why we build agents around tasks with a checkable output and an approval gate on anything that leaves the building. Narrow finance tasks behave very differently from a finance role. Reading an invoice into fields has a right answer you can test against, which is why automated invoice processing works when an all-purpose "AI bookkeeper" struggles. The same goes for phones: answering, qualifying and booking is a bounded job, covered in our AI receptionist guide.
What an AI employee costs against a person
We don't publish prices, ours or anyone else's, because every build depends on the process behind it. The comparison most people are really making is with a salary, though, so here's that side.
The ONS puts median full-time pay at £19.67 an hour. Add employer National Insurance and the minimum pension contribution and the cost to the employer is about £22.75 an hour. One full-time receptionist on the National Living Wage costs about £28,300 a year before any Employment Allowance.
| Cost line | A person | An AI employee |
|---|---|---|
| Getting started | Recruitment and induction | Connecting systems, writing the brief, building a test set |
| Running cost | Pay, employer National Insurance and pension | Model usage and hosting, which rise with volume |
| Checking the work | Light once they're trained | Heavy at first, on every task it didn't finish |
| Cover | Holiday and sickness | None needed, but someone must own it when it breaks |
| Judgement and relationships | Included | Not included |
The third row sinks most business cases. If an agent finishes a third of what it starts, a person still handles the rest, plus the time spent checking. Compare the cost of each accepted task, before and after, the way our note on cost per task sets out. A seat price set against a salary tells you very little.
It isn't an employee, and that cuts both ways
UK employment law defines an employee as an individual working under a contract of employment (Employment Rights Act 1996, section 230). Software isn't an individual. So there's no PAYE, no notice period, no sick pay and no unfair dismissal claim. There's also nobody else to blame.
The clearest case so far is Canadian. In a February 2024 decision, British Columbia's Civil Resolution Tribunal dealt with Air Canada's argument that it couldn't be held liable for wrong fare advice its website chatbot gave a customer. The tribunal called that "a remarkable submission" and added: "It should be obvious to Air Canada that it is responsible for all the information on its website." UK courts aren't bound by it, but the reasoning is the one any customer will expect.
So treat an AI employee the way you'd treat a powerful piece of software, because that's what it is:
- Name one person who owns it and answers for its output.
- Split its permissions into read, draft, approve and execute, and grant execute last.
- Keep a log of what it did and why, so a mistake can be traced and put right.
How to hire an AI employee
Start with the job advert you were about to post, then rewrite it as tasks. The process mapping method in our journal goes step by step. Here's a typical office administrator advert put through it.
| Line in the job description | Task an agent can own | Who checks it |
|---|---|---|
| Manage the shared inbox | Sort and tag routine email and draft the replies | Staff approve drafts until the error rate is known |
| Book meetings and manage diaries | Offer slots and send invitations under agreed rules | Runs alone once proven; a person handles key clients |
| Raise and chase invoices | Draft invoices from job sheets and send scheduled reminders | Accounts sign off before anything is sent |
| Keep the CRM up to date | Log calls and email against the right records | Weekly spot checks |
| Handle supplier queries | Stays with a person at first: it needs judgement and relationships | Not applicable |
| Support the team with ad hoc tasks | Not a task, so not something to hire an agent for | Not applicable |
Then run a probation, in this order:
- Pick the two or three tasks with the most volume and a checkable output.
- Run them in shadow mode: the agent drafts while the team carries on, and you compare the results.
- Put an approval step on anything that sends money, makes a commitment or reaches a customer.
- Widen what it does alone only when the logs show it has earned it, one task at a time.
The last line of the table is the useful one. Every job has a slice of "anything else that comes up", and that slice is exactly where the benchmark's agents fell over.
Before it peaks: what we'd do now
The forecasts point both ways. Gartner expects at least 15% of day-to-day work decisions will be made autonomously by agentic AI by 2028, up from 0% in 2024. It also predicts over 40% of agentic AI projects will be cancelled by the end of 2027. Both can come true at once: a lot of real work handed over, and a lot of oversold projects abandoned.
In the UK, ONS survey data puts AI use at around 35% of UK businesses with 10 or more employees, up from around 12% in late 2023. Most firms are still early, and there's no prize for being first to hire a persona.
Our opinion, then. Skip any AI employee sold as a whole role with no task list behind it. Skip anything that can move money or make a commitment without a person approving it. Buy the narrow task first. Our own publishing pipeline is built that way: separate research, writing and review workers, each checked against written criteria, with a person signing off where the criteria demand it. That isn't one "AI writer". It's several tasks with checks between them, and that's the version that holds up. If you want the order of work for a first project, how to implement AI in a business sets it out, and an AI audit finds which tasks cost you most time.
Common questions
What is an AI employee?
An AI agent sold under a job title and a persona, such as an AI receptionist or AI bookkeeper. Underneath it is a language model with instructions, connections to your systems and permissions. The title describes what the vendor hopes it will do; the instructions and tools decide what it actually does.
Can an AI employee replace a member of staff?
Rarely the whole person. In Carnegie Mellon's TheAgentCompany benchmark the best agent finished well under half of the realistic office tasks it was set, and fared worst on admin and finance work. It can take over specific, repetitive tasks, which frees staff time rather than removing the role.
How much does an AI employee cost?
It depends on the tasks, the systems it connects to and how much checking it needs, which is why we scope before quoting. Compare it with the full employer cost of a person, including National Insurance and pension, and judge both on the cost of each accepted task.
Is an AI employee legally an employee in the UK?
No. The Employment Rights Act 1996 defines an employee as an individual under a contract of employment, and software is not an individual. It has no employment rights, and the business is responsible for what it says and does, as Air Canada found when a tribunal rejected its defence over its chatbot.
What is an AI workforce?
Several agents, each owning a narrow task, working alongside staff and handing work between them. It works best when each agent has its own brief, permissions and named human owner, and when a person approves anything that reaches a customer or moves money. It is a design choice, not a product.
How do you hire an AI employee?
Rewrite the job description as a list of tasks, keep those with high volume and a checkable output, and run them in shadow mode while staff carry on. Add an approval step for anything external, and widen what the agent does alone only once its logs show it has earned it.
Before you post that job advert, write down what the person would do in a normal week and tick the tasks whose output you could check in a minute. Bring that list to an informal scoping chat (pick Business automation on the form), tell us which inbox, CRM and accounts systems you run, and we'll send a price range for building the ticked tasks afterwards.
Sources
Every figure in this article links back to the source below it was checked against.
- Carnegie Mellon University and Duke University: TheAgentCompany benchmark (NeurIPS 2025)
- Gartner: over 40% of agentic AI projects will be cancelled by end of 2027 (June 2025)
- ONS: Employee earnings in the UK, 2025 (ASHE)
- Orwell Lab calculation from GOV.UK 2026 to 2027 pay, NI and pension rates (inputs: GOV.UK, national minimum wage rates)
- ONS: Artificial intelligence in UK businesses, 2023 to 2026
Written for Orwell Lab’s practical guide collection. Benchmark results, legal sources and statistics checked on 5 October 2026.
Discuss your own requirements ↗