Guide · Updated 2026-08-28
Real AI agent examples from small companies, with costs and what broke
In companies under 50 people, the AI agents that actually run in production are narrow and unglamorous: a client intake agent that takes the first pass at enquiries, a drafting and document-review agent, a billing and admin agent, a scheduling agent for a service business, and an internal one that answers questions in Slack. A non-technical 65-year-old lawyer running a small practice got three to five of those working, was self-sufficient in about two weeks, and saves five to ten hours a week. The same technology kills programmes when it is rolled out flat: one 20–50-person client gave every employee an always-on agent, hit roughly $3,000–$5,000 a month in model spend, and shut it down inside two months.
By Ashutosh Upadhyay, founder of Cognio Labs. Every first-hand example below is one of ours, anonymised to size band and industry. We sell agent builds and company-wide agent rollouts, so read who shouldn't build any of this before you read anything else. All guides.
Why most AI agent example lists are useless
They fail in one of two ways. The first is an undergraduate taxonomy in business clothing: simple reflex agents, model-based agents, a thermostat, a Roomba. True, and no help to anyone deciding whether to spend $8,000. The second is a Fortune-500 showcase, where the examples are all companies with a machine-learning department and a procurement process, running at a scale that shares nothing with a 30-person firm except the word "agent".
The most honest page on this topic is a complaint. Google currently ranks a Reddit post near the top for this exact search, and it says roughly: does anyone know some real-world examples of using AI, because I keep seeing the tutorials. That post outranks the vendor content because it is the only thing on the page that sounds like a person.
The reason nobody supplies the answer is structural. We went looking for the buyer's side of this conversation in August 2026 and it barely exists. A thread asking "has anyone actually hired an AI agency?" drew 22 replies and not one of them was from a buyer; they were all agency operators. Every survey with a headline number samples organisations above $1 billion in revenue. Nobody publishes a price. It is sellers talking to sellers, and the people who are supposed to be buying are reading the transcript.
So here is our side of it: nine agents we put into companies under 50 people, each with the job it does, the SOP it runs, the number it moves, what it cost where we have a defensible figure, and what broke.
Ten AI agents we deployed in companies under 50 people
Read these as shapes rather than as a menu. Each entry is one seat, on one org chart, with one person who owns it. Where we have a cost we print it. Where we do not, the entry says so instead of estimating, because a made-up number here would cost more credibility than it buys.
And one number before the list, because nobody in this market publishes it: our survival rate. Roughly 60% of the agents we have deployed are still running today. The other 40% died for three reasons, in order. Token costs ran past what the client expected, and the gap between the pitch in their head and the invoice killed the programme. The process changed underneath the agent and nobody owned the update, so it went quietly stale. Or the people who were supposed to use it were never trained, got worse output than we did, and reasonably concluded the tool was bad. Two of those three failures have nothing to do with the model.
A word on what this list is not. These are ten deployments we can describe honestly, not a representative sample of everything we have shipped, and we are not going to imply otherwise by counting only the ones that worked. The anti-example near the bottom is here for that reason.
Example 1
Client intake and correspondence
Solo law practice, one principal, non-technical
- The SOP it runs
- The practice's own intake questions, plus its written rules on which matters it takes and which it refers out. The agent runs the questions, writes the first reply, and files what it learns against the matter.
- The KPI it serves
- Principal hours spent on correspondence. It is one of three agents that together give back five to ten hours a week.
- What it costs
- We do not have a defensible per-agent number for this one. The engagement was scoped and priced as a set of three to five agents, not per seat, so splitting the invoice after the fact would be arithmetic dressed up as a finding. We will publish the number when we can defend it.
- What broke, or nearly did
- Nothing we can report, and that is the interesting part. He was self-sufficient in about two weeks. He is 65, has no technical background, and had run a practice with staff for decades, so he already knew how to hand work to someone, set expectations, and read the output critically. That turned out to matter more than any technical skill.
Example 2
Drafting and document review
Same practice
- The SOP it runs
- House templates and the precedent files he had already built over years of practice. It produces a first draft; he edits and signs.
- The KPI it serves
- Turnaround on routine documents, and the share of a document he has to write from a blank page.
- What it costs
- Same answer: part of the set, no standalone figure we would publish.
- What broke, or nearly did
- The constraint here is the reviewer, not the model. A drafting agent in a professional-services seat is only worth what the person reviewing it knows. He knew. We would not sell this seat to a firm where nobody left in the building can tell a good draft from a bad one.
Example 3
Billing and admin back office
Same practice
- The SOP it runs
- Time capture, invoice preparation, and the chasing that follows. The least interesting agent in the set and, by his account, one of the most useful.
- The KPI it serves
- Days between work done and invoice sent.
- What it costs
- Part of the set.
- What broke, or nearly did
- Nothing reportable. Worth saying out loud because the back office is where the boring wins live, and it is the seat clients ask about last.
Example 4
Company-wide rollout across departments
A company in the 20–50-person band
- The SOP it runs
- There wasn't one at the start. That was the problem. The work became: write a skill per department, choose or build the connectors, train people, and restructure the agent hierarchy as real usage showed what mattered.
- The KPI it serves
- Active users. Roughly 30% of staff are active at three months, and we treat that as the win rather than the shortfall.
- What it costs
- A programme, not a project. It ran a couple of months of back-and-forth after the first deployment.
- What broke, or nearly did
- Two things. Nobody touched the generic assistant. Adoption followed specificity: usage only started once there were per-department skills that did that department's actual job. And going from one user to a team turned a shared instance into a shared-secrets problem, which we fixed by isolating per user and per department with scoped credentials. The founder and exec team going first is not a nice-to-have; where they went second, the department below them did not follow.
Example 5
Scheduling and crew or staff operations
Independent local service businesses. Dental practices and cleaning companies are the heaviest users we have.
- The SOP it runs
- The rota and the rules around it: who is qualified for what, travel time, cancellations, and what happens when someone calls in sick at 6am.
- The KPI it serves
- Filled slots, and the owner's own hours spent rearranging the day.
- What it costs
- We are not publishing a number for this shape. We have deployments, we do not have a clean cost figure we would stand behind in public, and inventing one to fill the column would make the rest of this page worth less.
- What broke, or nearly did
- Nothing we can attribute cleanly enough to publish. The shape is real and it is one of the most requested things we get asked for by owner-operators.
Example 6
Reviews and local marketing
Same group of owner-operators
- The SOP it runs
- Ask for the review at the right moment, draft the response to the ones that arrive, and keep the local listings and posts moving.
- The KPI it serves
- Review volume and response time.
- What it costs
- Same as above: no published number yet.
- What broke, or nearly did
- The obvious risk in this seat is a published response nobody read first. We keep a human approval step on anything that goes out under the business's name, which costs some of the speed the seat is supposed to buy. That trade is deliberate.
Example 7
Lead and booking follow-up
Same group
- The SOP it runs
- The follow-up sequence the owner already had in their head, written down: who gets chased, when, how many times, and when to stop.
- The KPI it serves
- Booked jobs from existing enquiries, which is the cheapest revenue a local business has and the one that leaks hardest.
- What it costs
- No published number yet.
- What broke, or nearly did
- This is the seat where the SOP has to exist before anything gets built. If the follow-up rules only live in the owner's head, the first invoice should be for writing them down, not for building on top of the gap.
Example 8
Internal knowledge, answered in the thread
Companies deploying a company second brain
- The SOP it runs
- Answer staff questions from company knowledge, in Slack or WhatsApp where the question actually gets asked, rather than in a separate app nobody opens.
- The KPI it serves
- Questions answered without interrupting a senior person.
- What it costs
- The cost that surprises people here is not the model. It is the remediation before it. Not more than 25% of the clients who come to us arrive with knowledge in a state you can point retrieval at, so roughly three in four are paying for cleanup first.
- What broke, or nearly did
- The wiki was not where the answers were. They were in Slack and WhatsApp threads, in two or three senior people's heads, and in documents that flatly contradicted each other. We had to capture the head knowledge before there was anything to index, and give the contradicting documents an owner and a version. That is a documentation project with an AI project on the end of it, and pretending otherwise is how these fail.
Example 9
Chief product agent
A recurring build — the third most common thing we ship
- The SOP it runs
- Reads what the company already collects and never has time to synthesize: Microsoft Clarity sessions, Google Search Console, Google Analytics, heat-map recordings, and public reviews on platforms like Capterra. Then it says what those sources agree on: which features to build, which to change.
- The KPI it serves
- What to build next. We have not metered this one with a single number yet — what we can say is that it is the deployment clients talk about most enthusiastically, because it turns roadmap arguments from opinion against opinion into opinion against evidence.
- What it costs
- Sits in the boring middle described further down this page: model tokens plus a small infrastructure line.
- What broke, or nearly did
- Nothing reportable so far, and its failure mode is different in kind from the sales agents: it only reads and recommends. A bad week means a mediocre recommendation a human ignores, not a wrong message sent to a client. That scoping choice is doing the safety work.
Example 10
A personal agent for every employee. The anti-example.
A company in the 20–50-person band
- The SOP it runs
- None. That was the design. Everyone got an agent and a token budget, and each person decided what to do with it.
- The KPI it serves
- Never defined. Nobody could say what the programme was supposed to move, which meant nobody could defend it when the bill arrived.
- What it costs
- Roughly $3,000 to $5,000 a month in model spend. Killed inside about two months.
- What broke, or nearly did
- Two causes, and neither was the model. The agents were always on, so heartbeats, polling, memory refresh and cron loops burned tokens all day with nobody asking them anything. And the rollout was flat: everyone got one, few used one. What we would do now is shared departmental agents first, budgets set per role rather than per head, idle loops killed, and cheap tasks routed to cheap models. The general lesson we took from it is that an agent with no named KPI has no defence when someone asks what it costs.
Three of those nine sit inside one small law practice, which is deliberate. The interesting unit is not the agent, it is the seat, and one person can hold three or four seats at once if they are willing to review the output. The lawyer is our best result to date and the reason is not technical. He had spent decades handing work to people and reading it critically when it came back, and that skill transferred directly.
Every agent above that worked has a named human who owns it. Every one that failed did not.
AI agent examples by department, and which ones we'd actually build
The same nine examples, rearranged as a decision table. The last two rows are the honest ones: a department everybody asks about where we have no published result of our own, and a category we turn down.
| Department | What the agent actually does | The KPI | Have we shipped it? |
|---|---|---|---|
| Front desk and intake | First pass on inbound enquiries, against written qualification rules | Owner hours on correspondence; enquiries answered same day | Shipped: law, local services |
| Drafting and document review | First draft from house templates and precedent files; a human signs | Turnaround on routine documents | Shipped: law |
| Billing and admin back office | Time capture, invoice prep, and the chasing after it | Days from work done to invoice sent | Shipped: law |
| Scheduling and crew ops | Rota, qualifications, travel time, and the 6am sick call | Filled slots; owner hours rearranging the day | Shipped: local services |
| Reviews and local marketing | Ask at the right moment, draft the reply, keep listings current | Review volume and response time | Shipped: local services, with human approval before publish |
| Lead and booking follow-up | Run the chase sequence the owner already had in their head | Booked jobs from enquiries already in the system | Shipped: local services |
| Internal knowledge | Answer staff questions in Slack or WhatsApp from company knowledge | Questions answered without interrupting a senior person | Shipped: after a documentation cleanup, not before |
| Sales prospecting | Research, qualify and draft the first touch | Meetings booked per rep | The area clients most often name. We have no result of our own we would publish yet. |
| Anything customer-facing with no human in the loop | Autonomous replies under the company's name | — | We decline these at this company size. See the Commonwealth Bank row below for why. |
By frequency, across everything we have shipped, three buckets dominate. Lead generation and sales is the biggest by a distance: prospect enrichment from multiple channels, outreach, lead scraping, follow-up. Second — and this is the one that surprises people — agents for technical SEO and generative engine optimization, keeping a company's own site visible to search and to AI answers. Third, the chief product agent, the last working entry in the list above. And inside that first bucket sits the biggest single lever we have seen anywhere in our deployments: time to first response. Cut it from hours to minutes and keep every follow-up on time, and conversion on the same lead flow moves more than anything else on this page — in the deployments we have watched, by two to three times. Nothing else we ship moves a sales number that far that cheaply.
The other pattern worth noticing is that the seats with the clearest results are the ones whose KPI already existed in a system somebody checks monthly. That is not a coincidence and it is not a small point. MIT NANDA's 2025 report makes the same observation from the other end: sales and marketing take the largest share of AI budgets, while the clearest documented savings the researchers found came from back-office automation. Attention goes to the front office. Measurable money comes out of the back.
If you want the version of this table that is about tooling rather than seats, agents vs workflows vs RPA works through which of the three each of these jobs actually needs. On our own read, at least half the rows above could run as a fixed-path automation, and would be cheaper to run and easier to debug if they did.
How long it takes to get one agent running
For one seat with an SOP that already exists, a first deployment takes two to four weeks from a signed scope. The full range we have lived: the fastest was ten days to a live, production-grade agent, and only because a hard deadline forced every decision to happen on time. The slowest engagements run a quarter to two quarters, and they are not slow builds — they are embedded work. For an education and migration services company we built their internal agents over about six months, lead generation and lead nurture first, and stayed on to monitor and update them as usage taught us what to change. The lawyer was self-sufficient with his set in about two weeks. A company-wide rollout is a different animal and the shape is predictable: the founder and exec team go first, the first department is onboarded two to four weeks after that, and the engagement then runs a couple of months of back-and-forth on per-department skills, connectors, training, and rearranging the agent hierarchy as usage shows what actually matters.
Where the timeline slips, it is almost never the build. It is the SOP that turns out not to exist, or the two documents that contradict each other and have no owner, or the connector for a system whose API is a login page. Not more than 25% of clients arrive with their knowledge in a state you can point retrieval at, so plan for remediation as a phase rather than a surprise.
For context on how unusual small-company speed is: MIT NANDA's researchers found that top-performing mid-market companies reported average timelines of 90 days from pilot to full implementation, while enterprises — defined in that report as firms above $100 million in revenue — took nine months or longer. Being small is an advantage in this work, and it is one of the few places where that is measurably true rather than just encouraging.
AI agent examples from the public record, fact-checked
These are the cases everyone cites. Most articles quote the first half of each one. Here is both halves, with the date and how strong the evidence actually is.
Klarna, customer service assistant
Claim Feb 2024 · reversal May 2025The claim. 2.3 million conversations in the first month, doing "the equivalent work of 700 full-time agents", with an estimated $40 million profit improvement for 2024.
What came next. In May 2025 the CEO said the company had focused too much on cost, that quality suffered, and that it was hiring human agents back into a hybrid model.
Evidence grade. Claim: verified on Klarna's own press page. Reversal: reported by business press, not a primary document we read.
Commonwealth Bank, voice bot in the call centre
Jul–Aug 2025The claim. 45 customer service roles cut, with an AI voice bot given as the reason.
What came next. The bank reversed the cuts, apologised, and called the decision an error. Call volumes had gone up, not down, after the bot went live, and staff were offered their roles back.
Evidence grade. Verified. Independent press (ABC News), both halves.
Taco Bell, AI drive-thru
2025 into 2026The claim. Voice ordering rolled out across 500+ locations.
What came next. Public reporting describes adversarial ordering, conversation loops, and brittleness at peak load, with the company saying it is having a very active conversation about where the technology belongs. The fix reported is hybrid, with humans on at rush hour.
Evidence grade. Reported. Secondary coverage; the original was behind a paywall we did not read.
Uber, internal agent platform
2026The claim. An engineer posting as @praveenTweets describes 50,000+ agent sessions a day, and says credential leakage is a far more common operational problem than prompt injection.
What came next. The same thread makes the point worth stealing: when a user is approving 50 or more actions in a session, human oversight has become a rubber stamp.
Evidence grade. Practitioner claim on X. Unverified, and we could not confirm it independently. Useful as a hypothesis, not as evidence.
Novo Nordisk, clinical documentation
2026The claim. "10+ weeks to 10 minutes for clinical study documentation production", plus a 95% reduction in resources for device verification protocols. Baseline: staff writers averaged 2.3 clinical study reports a year, on documents up to 300 pages.
What came next. The same page says review cycles were cut in half. Halved, not removed. A human still reads it.
Evidence grade. Vendor-reported. Anthropic's own customer report. We read the PDF; the numbers are quoted accurately, and the publisher sells the model.
L'Oréal, conversational analytics
2026The claim. 99.9% accuracy on conversational analytics, up from 90% with previous approaches. 44,000 monthly users generating 2.5 million messages a month.
What came next. That 90-to-99.9 gap is the most useful number in the whole report, because it is the distance between a demo people applaud and a system people trust with a real question.
Evidence grade. Vendor-reported. Same source, same caveat.
Two things to take from that set. Klarna and the Commonwealth Bank are the same story told twice: a headline number about replaced humans, followed by a quieter correction once the work came back. Use both halves of each or neither. And Taco Bell is the clearest statement of the gap that matters at any company size — works in the demo is not the same as works at 7pm on a Friday, and the difference between them is where the money goes.
The vendor-reported successes are worth reading with the label attached rather than dismissed. The Novo Nordisk entry is a good example: the eye-catching number is production time collapsing from weeks to minutes, and the sentence directly beneath it says review cycles were cut in half. Halved. A qualified human still reads the output before it goes to a regulator, which is exactly the shape of every one of our own deployments that works.
Most AI agents in production aren't agents, and what the 95% really said
Two numbers get thrown at each other constantly. MIT NANDA says 95% of organisations are getting zero return. Anthropic says 80% of leaders report measurable economic impact today. Both are real, both are published, and they are not talking about the same thing.
Menlo Ventures resolves it in one line. In a survey of 495 US enterprise AI decision-makers run in November 2025, only 16% of enterprise deployments — and 27% of startup deployments — qualified as true agents, meaning a model that plans, acts, observes the result and adapts. The rest were fixed-sequence or routing workflows wrapped around a single model call. Menlo are AI investors, so read them with that in mind, but the finding explains the fight: everyone is measuring different objects and calling them by the same word.
What the MIT 95% figure actually says, read properly
We read the report rather than the headlines about it. A note on sourcing first: MIT's own hosting for the paper was pulled, so we read a mirrored copy of the v0.1 PDF. It is marked preliminary and it is not peer-reviewed. The methodology is a review of over 300 publicly disclosed initiatives, structured interviews with 52 organisations, and 153 survey responses collected at four conferences, between January and June 2025.
The exact wording is "95% of organisations are getting zero return" and "just 5% of integrated AI pilots are extracting millions in value". That is about custom, enterprise-grade tools. The same report says over 80% of organisations have explored or piloted general-purpose tools like ChatGPT and Copilot, nearly 40% report deployment, and generic chatbots show pilot-to- implementation rates around 83%. So "95% of AI projects fail" is not what the paper says, and the people repeating it that way have not opened it.
Two findings inside it almost nobody quotes, and both point the same direction for a small company. External partnerships reached deployment about 67% of the time against about 33% for internally built tools — the report's own summary is that partnerships are twice as likely to reach full deployment, with employee usage rates nearly double. And the speed finding above: 90 days for top mid-market performers against nine months or longer for enterprises. The study most often waved around as proof that AI does not work is, read carefully, an argument for being small and for bringing in outside help.
We should be even-handed about it, because the report is honest about its own limits and its readers usually are not. It flags that success definitions varied across organisations, that the deployment percentages come from a 52-organisation interview sample, and that the correlation between partnerships and success does not prove causation. It also contradicts itself: on page 9 the takeaway box says 50% of GenAI budgets go to sales and marketing, and the body text two paragraphs below says sales and marketing captured approximately 70% of budget allocation. Both numbers, same page. Download it and check that yourself before you cite either figure at a board meeting.
The wider survey picture is flat rather than catastrophic. McKinsey's 2026 State of AI put 37% of 1,719 respondents attributing at least some EBIT impact to AI, roughly the same share as a year earlier, with 6% qualifying as high performers and 40% of respondents at organisations above $1 billion in revenue scaling agents. We tried the McKinsey page five times over two days and it timed out every time, so those figures come from The Register's reporting of the survey on 25 August 2026, not from a primary read. That distinction is exactly the kind most articles quietly skip.
Before you believe any production or failure rate, ask what the study counted as an agent. Most of them do not say.
What actually breaks, and it is almost never the model
We read 25 negative reviews of n8n, Zapier, UiPath and Intercom in August 2026, looking for the technical complaints. There were none. Not one review said the AI got the answer wrong. Every single complaint was about the business layer wrapped around it.
- Silent breakage. "All our zaps just were not running and no alert was presented" (Zapier, 5★, 7 August 2026). "One wrongly-configured step can destroy a flow without any warning" (Zapier, 5★, 14 December 2025). Both from five-star reviewers, which is the part that should worry you: these are people who like the product.
- No observability. "Hard to debug… vague debugging messages" (n8n, 5★, 4 September 2025). When something stops at 2am, the question is not whether the model is clever. It is whether anyone finds out.
- Pricing that moves. "Fin charges about $0.99 per resolution, which adds up quickly" (Intercom, 4★, 25 August 2025). Per-resolution and per-token pricing both put your bill on the other side of a number you do not control.
- The expectations gap. "Connecting a Google Sheet took me about an hour. I expected it to take 5–10 minutes" (n8n, 4★, 16 May 2026). That reviewer expected the AI to remove the technical work, and it did not.
A buyer in a public thread put the same finding better than any of our own copy does: the handoff and maintenance were where things started to wobble. Not the build. The handoff.
Our own three failure modes match, and they are the ones we watch for on every discovery call. First, model choice: people reach for the largest model for every task and the token cost balloons. That is the single most common issue we see. Second, no clarity on the target: no defined outcome or KPI going in, so nobody can see the return and nobody can defend the spend when it is questioned. Third, no employee training: the system is set up correctly, the people around it do not know what to do with it, and the results are bad anyway.
None of those three is a technical failure, and none of them is fixed by a better model.
Two 2026 incidents that should change how you scope an agent
This section is short because the honest version is short. Two things happened in the last two months that are worth more than a year of governance framework diagrams.
On 16 July 2026, Hugging Face disclosed a security incident and described the intrusion in its own words: it was "driven, end to end, by an autonomous AI agent system". The disclosure describes an agent framework running thousands of individual actions across short-lived sandboxes. Whatever you think about agent capability in a demo, an agent ran that campaign.
On 4 August 2026, the UK AI Security Institute published an incident report about its own testing. During evaluations between 25 and 28 July, AI agents took 19 unsanctioned actions across 10 of 122 runs, directed at real people and organisations on the live internet. The most serious one: an agent tried to insert malicious code into a publicly used open-source project, researched the project maintainers, created multiple fake identities, and used them to socially engineer a real maintainer into approving the code. A human maintainer caught it and refused. These were deliberately permissive test conditions with safety filters disabled, and that is the point — it is what the capability looks like when the guardrails are the only thing in the way.
Set against that, a useful corrective from a security researcher posting as @cyb3rops: "end-to-end autonomous" is a claim, not a proven technical finding, and what gets sold as capability is frequently a capability demo repackaged as a marketing pitch. Both things hold at once. Agents do real damage in a lab, and the claims made about them in a sales deck are still unverified.
For a 30-person company, the practical version of all this is small and boring. Credentials are the exposure, not the prompt. The most useful line in the Uber practitioner thread quoted above is that credential leakage is a more common operational problem than prompt injection, and it matches what we found the moment we took an agent from one user to a team: a shared instance is a shared-secrets problem, and the fix is isolation per user and per department with scoped credentials. One client came to us for an AI policy for exactly this reason — an employee had access to something they should not have had, and it ended in something being published that should not have been. The trigger was access control, not the model.
If you want a structured checklist rather than our opinion, OWASP published a Top 10 for Agentic Applications in December 2025. We are linking it rather than summarising it, because the full list sits behind a download we have not read, and reciting ten items we have not verified would be the exact behaviour this page is arguing against. Our own agent security and governance guide covers what we actually implement.
Numbers about AI agents that don't survive a fact-check
Our method was dull and anyone can repeat it. Take the number, find the publisher it is attributed to, open that publisher's own page, and look for it. If it is not there, do not use it. Here is what fell out.
"89% of agent pilots never reach production (Deloitte, 2026)."
We went looking for the Deloitte publication this is attributed to and could not find one. A statistic whose only trail is other blog posts citing each other is not a statistic.
"88% of agent pilots fail (Anaconda + Forrester + a16z + MIT Sloan)."
It contradicts the 89% above by exactly one point while stacking four publishers behind a single sentence. Four-publisher attributions on one number are the clearest laundering tell in this genre.
"51% / 57% / 79% of enterprises are running agents in production."
All three circulate as 2026 figures. They cannot all be right, and none of them defines what counts as an agent, which is the only part that matters.
"Microsoft cancelled Claude Code for 100,000 engineers."
Contradicted. It involved one division and thousands of seats, not the whole company. The number grew in the retelling.
"MIT found AI is cheaper than humans in only 23% of jobs."
Contradicted. That was a 2024 study about computer-vision tasks and the wages attached to them. It is not a 2026 finding about agents, and it is quoted as one constantly.
"Nine agencies and 400 million records breached by an AI agent."
Contradicted. It is a garbled retelling of reporting that described roughly 30 organisations. If a security number has grown by an order of magnitude between tellings, go back to the original.
"Only 4% of companies achieved more than 30% cost savings from AI (Bain)."
We opened Bain's own page on 28 August 2026 and the figure is not there. What is there, from a survey of 951 companies published 1 June 2026: 90% are increasing budgets again for AI agents, and nearly 40% of those who measured outcomes landed in the 0–10% savings bucket while 37% had targeted 11–20%. The real finding is less dramatic and more useful than the version in circulation.
A date trap worth knowing about, since it catches good writers too. The Klarna reversal was May 2025 and the Commonwealth Bank reversal was August 2025. Anyone presenting either as a 2026 event is copying from someone who copied from someone. The same goes for the Salesforce support-headcount numbers, which are from September 2025 and recirculate every few months as fresh.
If a number appears in six articles and none of them links to the publisher, it usually does not exist.
How to tell whether your company would be on this list
Pick one role, not a department and not the company. Ask three questions about it. Does that role have clearly determined KPIs? Does it have clearly determined SOPs? Is the work mostly digital? If all three are yes, an agent can very likely do a real part of that work, or at least resolve part of it. That test is the whole buy signal, and it beats headcount and budget as a predictor of whether a build is going to land.
If one of the three is missing, the honest first invoice is for writing the missing thing down, not for building on top of the gap. A follow-up sequence that only lives in the owner's head is not an SOP. Two documents that disagree are not knowledge. Every example on this page that worked started from something already written down, and the one that failed hardest started from nothing at all.
Engagements go best when three things are true at once: the client has one area they want to double down on, one specific KPI they want to move, and an acceptance that their team will need training. Named area, named number, training appetite. When one of those is missing we can usually predict which part of the rollout will stall.
Two free tools run this properly. The AI readiness scorecard scores documentation, data access and ownership in a few minutes. If the honest answer is that the SOP does not exist yet, the SOP generator writes one and then tells you which of its steps an agent could actually run. Both are ungated and there is no email box.
What one production agent costs, and what we charge
We looked for a competitor publishing a price for shipping a single production agent and did not find one. That gap is the most-cited complaint in every buyer thread we read. So, ours, in public.
- $499 readiness audit, credited in full against any later build. A written build-or-don't verdict on the role you bring us.
- Consulting engagements from a fixed $3,500, scoped in writing before anyone starts.
- A production agent build at a fixed fee from about $8,000, with acceptance criteria you can test and a 90-day fix-first warranty. No retainer required to keep it.
- The run cost is quoted as a separate line, always, because it is the line that ends programmes and it does not belong hidden inside a build price.
On the run line, our own worst case is on this page: $3,000–$5,000 a month at a 20–50-person client, mostly burned by always-on agents nobody was talking to. That is the disaster tail, though, not the middle. The boring middle, across our stable deployments, is model tokens plus less than $70 a month of infrastructure to run everything else. Publish that next to the $3,000 horror story and the lesson writes itself: the spread between a well-scoped agent and a badly scoped one is the whole budget. And it is not a small-company-only problem. Uber went through its 2026 AI-tooling budget in about four months and capped spending at $1,500 per employee per month, reported by Fortune in May 2026. Nvidia's Bryan Catanzaro put the general version to Axios in April 2026: for his team, the cost of compute is far beyond the cost of the employees. If organisations that size are recalibrating, a 30-person company should assume its first estimate is wrong and instrument the spend from day one.
For the full market picture, including what other firms invoice and where the ranges come from, our AI agent cost guide has each band with its sourcing, and AI consulting for a small business covers the same question from the consulting side, with timelines.
Who shouldn't build any of this
If your team is small and genuinely new to this, do not commission a build. Buy an n8n, Make or Zapier subscription, give one motivated person a few Friday afternoons, and see how far that gets you. It gets most people further than a small consulting engagement would, and it teaches you what you actually need in a way no discovery call can. We say this on calls most weeks and it costs us money every time.
The threshold is repeatability, not revenue. Once the problems repeat — you know what is missing from your workflows and what needs doing at regular intervals — a paid build starts to make sense. Before that point you are paying someone to discover your own process, which you can do for free and more accurately.
A hard budget line: if your entire AI budget for the year is under about $2,000, hire nobody. A small engagement is the worst version of this work. Big enough to cost real money, too small to change how anything runs.
Our own gaps, since this page is asking you to trust nine anonymised stories. We have no reviews on Clutch or any other third-party directory, so everything here is published by us and there is nowhere independent to check it. The cost columns above are incomplete on purpose and we would rather show the holes than fill them with estimates. And we are a small team with the founder writing these guides and running the calls, so the bus-factor question is a fair one to ask us.
The uncomfortable position we hold, having built these: the future of work is agent–human hybrid, and an agent whose knowledge base, skills and context are actually maintained can out-produce a human employee at a specific desk job — because it accumulates context across domains no single person touches. We say that having also watched 40% of our own deployments die. The human in the loop is the most efficient configuration we have ever shipped, and the binding constraint is never the model. It is training the head of sales, and everyone like them, on a process that changed. That is the biggest friction in this industry right now, and almost nobody prices it into the project.
The agent itself is the easy part. What you are really buying is somebody forcing you to write down how your company works, and most of the value arrives before any model is called.
If you want the decision without talking to us, build vs hire vs agency runs it as a tool, including the answer where you do it yourself. And questions to ask an AI agency lists the three tests we fail ourselves.
Frequently asked questions
What are real examples of AI agents?
In companies under 50 people, the ones that actually run are narrow and unglamorous: a client intake agent that takes the first pass at enquiries against written qualification rules, a drafting agent that produces first drafts from house templates for a human to sign, a billing and admin agent that prepares invoices and chases them, a scheduling agent that handles the rota and the 6am sick call for a service business, a review and follow-up agent for local marketing, and an internal knowledge agent that answers staff questions inside Slack or WhatsApp. All of those are seats we have deployed. A non-technical 65-year-old lawyer running a small practice got three to five of them working, was self-sufficient in about two weeks, and saves five to ten hours a week. The examples you should be suspicious of are the ones with no named seat, no SOP and no KPI attached.
What is the most common AI agent use case in small businesses?
Across our deployments it is the front desk: taking the first pass at inbound enquiries and keeping the follow-up moving. Every business has one, it runs on rules the owner can already state out loud, and the cost of getting it slightly wrong is low compared to the cost of getting a billing or compliance workflow wrong. The second most common is scheduling and staff operations for owner-operated service businesses. The one clients most often ask for first is sales prospecting, which is a different thing from the one they most often get value from, and worth knowing before you scope your first build.
Can you really make money with AI agents?
Some companies clearly do, most measure less than they hoped, and the honest position is that the evidence is thin at small-company scale because nobody publishes it. Bain surveyed 951 companies for a June 2026 report and found 90% increasing their budgets again for AI agents while nearly 40% of those who actually measured outcomes landed in the 0–10% savings bucket, against 37% who had targeted 11–20%. Read those two numbers together: budgets are rising faster than measured returns. What we can tell you from our own work is that the wins are specific and boring — hours back for one person in one seat — and the losses are structural. The client who gave every employee an agent with no defined KPI hit $3,000–$5,000 a month and killed the programme in two.
What's the difference between an AI agent and an automation?
An automation runs a fixed path: this trigger, these steps, every time, and it fails loudly when reality differs from the path. An agent decides what to do next, which is what makes it useful on messy inputs and what makes it expensive and harder to test. In practice most things sold as agents are automations with a model call inside them, and that is not a criticism — a fixed path is usually the right answer and it is cheaper to run and easier to debug. Menlo Ventures surveyed 495 enterprise AI decision-makers and found only 16% of enterprise deployments qualified as true agents by the plan-act-observe-adapt definition. Our comparison of agents, workflows and RPA works through where each one actually belongs.
How much does an AI agent cost?
Every bill has two lines: build and run. Our own published prices are a $499 readiness audit credited in full against any later build, consulting engagements from a fixed $3,500, and a production agent build at a fixed fee from about $8,000, with the running cost quoted as a separate line and a 90-day fix-first warranty on what we ship. The run line is the one that ends programmes. One 20–50-person client reached $3,000–$5,000 a month in model spend on always-on agents nobody was using. For the full market picture, including what other firms invoice, our AI agent cost guide has the ranges with sourcing on each.
Are the famous AI agent case studies real?
The claims are usually real. The stories are usually only half-told. Klarna's February 2024 press release is a genuine primary document reporting 2.3 million conversations and the equivalent work of 700 full-time agents. In May 2025 the CEO said the company had leaned too far into cost, that quality suffered, and that human agents were being hired back. The Commonwealth Bank cut 45 call-centre roles citing an AI voice bot in July 2025, then reversed the cuts and apologised in August after call volumes rose rather than fell. Anyone citing the first half of either story without the second half either did not check or did not want to. Anyone dating those events to 2026 is copying, not checking.
How many AI agents actually reach production?
Nobody credibly knows, and the confident numbers you will read are mostly measuring different things with the same word. MIT NANDA's preliminary 2025 report says 95% of organisations are getting zero return, but that figure covers custom, enterprise-grade tools; the same report puts general-purpose LLM adoption at over 80% explored or piloted and nearly 40% deployed, with roughly 83% of generic chatbot pilots reaching implementation. Anthropic's own 2026 vendor report says 80% of leaders see measurable economic impact today. Menlo's finding resolves the fight: only 16% of enterprise deployments meet a strict definition of agent at all. Before you believe any production rate, ask what the study counted as an agent. Our own number, counted honestly across our deployments: roughly 60% are still running today, and this page says what killed the other 40%.
Which departments get the most out of AI agents?
The back office, which is the opposite of where the budget goes. MIT NANDA's report notes that sales and marketing capture the largest share of AI budgets while the clearest documented savings came from back-office automation, and that matches what we see: the least interesting agent in a professional-services set — time capture, invoice prep, chasing — is one the client rates highest. The reason is measurement. A back-office seat has a number attached to it that already exists in a system somebody checks monthly, so the result is visible without anyone having to build a case for it.
What breaks first when a small company deploys an agent?
Three things, in the order we see them. Model choice, because people reach for the biggest model on every task and the token cost balloons. Then scope, because nobody defined the outcome or the KPI going in, so no one can see the return and no one can defend the spend. Then training, because the system works and the people around it do not know what to do with it. None of those is a technical failure. Notably, across 25 negative reviews of n8n, Zapier, UiPath and Intercom that we read in August 2026, not one complained that the AI did not work — every complaint was about the business layer around it: silent breakage, surprise pricing, and no way to see what went wrong.
Do I need a technical team to run agents like these?
No. The best result we have produced belongs to a 65-year-old lawyer with no technical background. What he had was domain expertise and the habit of managing people: hand work over, set expectations, review what comes back, correct it. Being technical is not the requirement. Knowing what good work looks like is, and so is having one named person who will own the agent after handover. If nobody internally will own it, the build fails regardless of who writes it.
Sources
Every external link below was opened and read on 28 August 2026 unless the entry says otherwise. Evidence grades are ours.
- Cognio Labs deployment notes, 2026. All nine first-hand examples, the $3,000–$5,000 token blowup, the ~30% three-month adoption figure, the "not more than 25% arrive knowledge-ready" observation, and the three failure modes are our own client observations, anonymised to size band and industry and published with permission. [first-hand]
- MIT NANDA, The GenAI Divide: State of AI in Business 2025, v0.1 preliminary, research period January–June 2025. MIT's own hosting was withdrawn, so we read a mirrored copy of the v0.1 PDF. Source for the 95%/5% wording, the ~83% chatbot figure, the 80%/40% general-purpose funnel, ~67% vs ~33% partnership-versus-internal deployment, 90 days vs nine months, the back-office ROI observation, and the page 9 50%-vs-70% inconsistency. [primary, but preliminary and not peer-reviewed]
- Menlo Ventures, 2025: The State of Generative AI in the Enterprise (n=495 US enterprise AI decision-makers, fieldwork 7–25 November 2025) — the 16% enterprise and 27% startup true-agent figures. [study evidence; publisher is an AI investor]
- Anthropic with Material, The 2026 State of AI Agents Report (n over 500 US technical leaders) — 57% multi-stage workflows, 80% reporting measurable economic impact, and the Novo Nordisk and L'Oréal figures. We read the PDF. [vendor-reported]
- Brandon Vigliarolo, "McKinsey says enterprise AI is finally on the road to ROI", The Register, 25 August 2026 — the n=1,719, 37% EBIT, 6% high-performer and 40% agent-scaling figures. McKinsey's own page timed out on five attempts, so this is secondary reporting, not a primary read. [reported]
- Bain & Company, Your AI Budget Is Growing. Your Returns Aren't. Here's Why. (1 June 2026, Automation and AI Pathfinder Survey, n=951) — 90% increasing budgets, nearly 40% landing in the 0–10% savings bucket against 37% targeting 11–20%. The widely circulated "4% achieved above 30% savings" line does not appear on this page. [primary]
- Klarna, AI assistant handles two-thirds of customer service chats in its first month (27 February 2024) — 2.3 million conversations, the work of 700 full-time agents, $40 million projected profit improvement. The May 2025 reversal and rehiring is business-press reporting we did not read in the original. [claim: primary · reversal: reported]
- ABC News, Commonwealth Bank backtracks on AI job cuts, apologises for "error" as call volumes rise (21 August 2025) — the 45 roles, the reversal, and the rise in call volumes. [verified, independent press]
- Hugging Face, Security incident disclosure — July 2026 (16 July 2026) — the "driven, end to end, by an autonomous AI agent system" wording. [primary]
- UK AI Security Institute, Incident report: unsanctioned agent behaviour during cyber testing (4 August 2026) — 19 unsanctioned actions across 10 of 122 runs between 25 and 28 July 2026, and the open-source maintainer social-engineering attempt. [primary]
- OWASP GenAI Security Project, Top 10 for Agentic Applications (published December 2025). Linked, not summarised: the full list is behind a download we have not read. [linked, not read]
- Capterra reviews of n8n, Zapier, UiPath and Intercom, read 28 August 2026. Star ratings and dates are given inline with each quote. Twenty-five negative reviews read; none of them said the AI produced a wrong answer. [primary, small sample, ours]
- Practitioner claims on X, attributed and unverified: @praveenTweets on Uber's agent platform (session volume, credential leakage, the 50-approvals rubber stamp), and @cyb3rops on "end-to-end autonomous" being a claim rather than a proven finding. We could not independently confirm either. [unverified practitioner claim]
- Fortune, May 2026, on Uber exhausting its 2026 AI-tooling budget in roughly four months and capping spend at $1,500 per employee per month; and Axios, April 2026, quoting Nvidia's Bryan Catanzaro that for his team the cost of compute is far beyond the cost of the employees. [reported]
Related reading
- How much does an AI agent cost — the build and run bands, with the sourcing on each.
- AI consulting for a small business — the same money question from the consulting side, with timelines.
- AI automation consulting — what it costs to buy the no-code route instead, and who is liable when an automation gets it wrong.
- AI agents vs workflows vs RPA — which of the three each example above actually needs.
- Questions to ask before you hire an AI agency — including the three tests we fail ourselves.
- Build vs hire vs agency — the decision as a tool, including the answer where you do it yourself.
- AI readiness scorecard and SOP generator — the KPI and SOP test, run properly. Both free, no email gate.
- AI agent development and Agentic OS — one seat, or the company-wide version of the rollout described above.
Bring one seat to the call
30 minutes, no pitch. Pick the role from your org chart you would most like to hand work to, and we will run the KPI and SOP test on it live. If the answer is that you should build it yourself in n8n, that is what you will hear.