Free tools · 10 minutes · Updated 2026-08-30
The Agent-Ready Score: find the one role an AI agent can actually run
A role is agent-ready when three things are true: it has clearly determined KPIs, it has a written standard operating procedure, and the work happens mostly in software. This free test scores any role in your company duty by duty against those three questions and hands back a verdict per duty — agent-coverable, partial, or human — plus a score out of 100 for the role. Ten minutes, no email, nothing leaves your browser.
Fair warning before you start: about half the tasks owners bring us fail this test. We publish it anyway, because knowing which half you're in before you spend a dollar is the entire value.
By Ashutosh Upadhyay, founder of Cognio Labs. This is the same test we run before we agree to build anything — it disqualifies work for us as often as it qualifies it.
Score a role
Name the role, break it into the duties it actually consists of (five to eight is the useful range), and answer the three questions for each. Two minutes per duty. No email — the score and every verdict show right here.
Is there a number that says this duty was done well?
Could a new hire do this duty from what's written down today?
How much of this duty happens in software?
Answer all three questions for at least one named duty.
What are the three questions in the test?
Every duty gets the same three questions. Each answer scores 0, 1 or 2 points, so a duty scores 0–6.
1. Is there a number that says this duty was done well?
A KPI someone could read today — not a goal, a reading.
- 2 ptsYes — named, and somebody reads it regularly
- 1 ptsIt could be measured. It isn't yet
- 0 ptsNo — whether it went well is a feeling
2. Could a new hire do this duty from what's written down today?
A standard operating procedure they could follow on their first Monday without asking you anything.
- 2 ptsYes — a written procedure exists and is current
- 1 ptsPartly — something's written, but the real rules live in someone's head
- 0 ptsNothing is written down
3. How much of this duty happens in software?
Email, CRM, documents, spreadsheets, browsers — versus phone calls, site visits, hands.
- 2 ptsAlmost all of it
- 1 ptsAbout half
- 0 ptsMostly not — it's calls, visits, or physical work
A role with determined KPIs, a written SOP and mostly-digital work is a role an agent can very likely run better than it is being run today. That single sentence matters more than your headcount, your industry, or your budget. A missing yes is not a no — it tells you which document to write first.
One modifier sits outside the score: whether the duty touches client-confidential or regulated data. It doesn't change the arithmetic. It changes the build — permissions get designed before anything is connected, never after.
How does the scoring work?
Per duty: 0–6 points, three verdicts. One rule keeps the test honest — a zero on any axis caps that duty at “partial” no matter what the total says. One dead axis is a build-stopper on its own, and averaging over it is how demos get sold.
| Duty score | Verdict | What it means |
|---|---|---|
| 5–6, no zeros | Agent-coverable | All three axes hold. An agent can very likely do this duty better than it is being done now, or take most of it off someone's desk. |
| 3–4, or any zero | Partial — fix one axis first | The work is close, but one axis fails. Usually it's the SOP: the duty runs on rules nobody has written down. Write them first — that costs a Friday afternoon, not an invoice. |
| 0–2 | Human, for now | Two or more axes fail. Don't build here yet. This is not a defeat — knowing which duties to leave alone is most of what the test is for. |
The role score is the duty points earned divided by the points available, out of 100. Three bands:
| Score | Band | What it means |
|---|---|---|
| 75–100 | Agent-ready — build the top duty first | Most of this role's duties pass the test. Pick the single duty with the highest score and the clearest KPI, and build that one first. Not the whole role — one duty, one number, then the next. |
| 45–74 | Ready after documentation | The role is a real candidate, but duties are failing on the SOP axis. You can't automate what isn't documented. This is the most common result we see: not more than 25% of the clients who come to us arrive with their knowledge in a usable state, so roughly three in four start exactly where you are. |
| 0–44 | Don't build yet | Most duties fail on two or more axes. An agent built on this role today would be a demo: impressive in the meeting, useless by month two. About half the tasks owners bring us land here or in the band above, so you're in the majority, not behind. |
A worked example: inbound lead follow-up, scored honestly
The role we get asked about most — by funded startups calling it an SDR seat and by small companies calling it “whoever answers the enquiries”. Lead generation and sales is the biggest bucket across everything we have shipped. Here is the role broken into six duties and run through the test, with the reasoning we would give on a call.
| Duty | KPI | SOP | Digital | Verdict |
|---|---|---|---|---|
| First reply to a new inbound enquiry | 2 | 2 | 2 | Agent-coverable |
| Research and enrich the prospect | 1 | 2 | 2 | Agent-coverable |
| Chase leads that went quiet | 2 | 0 | 2 | Partial — write the SOP first |
| Qualify: decide who's worth a call | 1 | 0 | 2 | Partial — write the SOP first |
| Run the discovery call | 2 | 0 | 0 | Human, for now |
| Log everything in the CRM | 2 | 2 | 2 | Agent-coverable |
- First reply to a new inbound enquiry. Time to first response is a real number somebody already watches. Reply templates exist. This is the duty to build first — in the deployments we have watched, cutting first response from hours to minutes moved conversion on the same lead flow by two to three times.
- Research and enrich the prospect. The lookup steps are written down; nobody measures enrichment quality yet. Buildable now — name the number during the build, not after.
- Chase leads that went quiet. Follow-ups-on-time is measurable and the work is all in software. But who gets chased, when, how many times, and when to stop lives in the founder's head. If those rules aren't written down, the first invoice should be for writing them down — not for building on top of the gap.
- Qualify: decide who's worth a call. Pure judgment today. It usually turns out to be four rules and a tie-breaker once someone writes them out loud.
- Run the discovery call. A live conversation with a stranger. The agent preps the brief and books the slot; the call stays human. Scoring this honestly is the point of the test.
- Log everything in the CRM. The duty everyone hates and every agent does gladly. Perfect score, zero glamour.
Total: 26 of 36 points — a role score of 72, which lands in “Ready after documentation”, not in “Agent-ready”. That is the typical honest result for this role, and it is worth sitting with. The most-requested agent build in the market usually needs two SOPs written before it deserves a build. The first response duty alone is worth building — it carries the single biggest lever we have seen anywhere in our deployments, moving conversion on the same lead flow by two to three times when first response drops from hours to minutes. But the chasing and qualifying duties, built on rules that live in a founder's head, produce an agent that guesses. Guessing agents get switched off.
Why is “write the SOP first” the most common verdict?
Because the work is usually clear and the documentation usually isn't. Not more than 25% of the clients who come to us arrive with their knowledge in a state you can point an agent at — roughly three in four need the writing-down done first. The answers live in Slack and WhatsApp threads, in two senior people's heads, and in documents that contradict each other. None of that is retrievable by software.
The good news is the fix is free. A duty that fails on the SOP axis needs a document, not a vendor: the trigger, the steps, the exceptions, and what done looks like, written so a new hire could follow it on their first Monday. Our SOP generator walks you through it in eight sections and then tells you which steps an agent could run. Re-score after. An SOP fix typically moves a duty from partial to agent-coverable without anyone writing code.
What does it cost to skip the test?
Two exhibits from our own files, both anonymized, both cleared for publication.
The rollout that skipped it. A 20–50-person company gave every employee a personal agent, each with its own token budget, no per-role scoring first. Spend reached roughly $3,000–$5,000 a month and the program was abandoned inside two months. Two causes: always-on agents burning tokens on heartbeats, polling and memory refresh while nobody asked them anything, and a flat rollout — everyone got one, few used one. Duty-level scoring exists to prevent exactly this: it forces the build to start where the KPI and the SOP already are.
What a rollout that works looks like. Founder and exec team first, one department 2–4 weeks later, and roughly 30% of staff actively using agents at three months. If you expected 100%, that number reads as failure. It isn't. It's the win, and planning around the real number is cheaper than discovering it.
Who should not use this — and where DIY wins
If your role breaks down to one duty and it's a fixed, repeatable path — form comes in, record gets created, email goes out — you don't need this test and you don't need us. Build it in n8n, Make or Zapier with one motivated person's Friday afternoons. The test earns its ten minutes when a role has five or more duties, when judgment sits inside some of them, and when someone is about to put real money behind “an AI employee”.
And if every duty in your company scores 0 on the digital axis — crews, sites, hands — an agent isn't your next hire. Software can schedule the crew. It can't pour the concrete.
What is the paid version of this?
This page is the self-serve version of the first hour of our Agent Readiness Audit. The audit runs the same test across your whole operation with us in the room, checks whether your data and tools can actually support what scores well, and ends in a written build-or-don't-build verdict with a build spec for whatever passes. It is $1,500, fixed, takes a week, and the fee is credited in full to a build within 30 days.
If your top role scored “don't build yet”, the audit is where we tell you that to your face. You pay us to hear that too — it's cheaper than hearing it from a dead pilot. Details on the audit page, or and bring your score.
Frequently asked questions
What is the Agent-Ready Score?
A free scoring test that tells you which role in your company an AI agent can actually run. You break the role into five to eight duties and score each duty on three questions: is there a number that says the duty was done well (a KPI), could a new hire do it from what's written down today (a standard operating procedure), and does the work happen in software? Each duty gets a verdict — agent-coverable, partial, or human — and the role gets a score out of 100. It takes about ten minutes, runs in your browser, and needs no email.
What makes a role a good fit for an AI agent?
Three things, and they matter more than your headcount or your budget: clearly determined KPIs, a written standard operating procedure, and work that is mostly digital. Three yeses and an agent can very likely do that work much better than it is being done now, or take a large part of it off someone's desk. This is the test we run before we agree to build anything, and it disqualifies revenue for us regularly — about half the tasks owners bring us fail it on the first pass.
What do I do when a duty fails the SOP question?
Write the procedure. That is the whole fix, and it costs a Friday afternoon rather than an invoice. Write it the way you would for a new hire on their first Monday — trigger, steps, exceptions, and what done looks like. If a person could follow it without asking you a question, an agent has something to work from. Our free SOP generator produces one in eight sections, then tells you which steps an agent could run.
How is this different from the AI readiness scorecard on this site?
Different altitude. The readiness scorecard asks whether your company can start an AI project at all — ownership, cost control, training, governance. The Agent-Ready Score asks a narrower question about one role: which of its duties can an agent take over, starting now. Run the company one first if you haven't; run this one when you're deciding where the first build actually goes.
Do I have to give you my email to use it?
No. The test, every per-duty verdict, and the full score are free on this page with no signup. The email gate covers only the take-away files: the print-ready worksheet and the editable markdown version with the SOP-first checklist. We built it that way deliberately — a score you can only see after a sales call isn't a diagnostic, it's bait.
What share of roles actually pass?
About half the tasks owners bring us fail the test on the first pass, usually on the SOP axis. That number is the most useful thing on this page. It means a mediocre score puts you in the majority, and it means the fix is usually documentation — which is free — rather than technology, which isn't.
If a role scores high, will the agent definitely work?
No, and be suspicious of anyone who says otherwise. The score tells you the work is automatable; it says nothing about whether your team will adopt it or whether your knowledge is in a usable state. From our own rollouts: roughly 30% of staff actively using agents at three months is what success looks like, and not more than 25% of the clients who come to us arrive with their documentation in a state you can point an agent at. The score is the right first filter. It is not a guarantee, from us or anyone.
Related reading
- The AI Vendor Interrogation Kit — once you know which role to automate, this is how you vet whoever offers to build it. Twelve questions, our own answers included.
- AI readiness assessment — the company-level version: ownership, cost control, training, governance.
- Real AI agent examples — nine agents we deployed in companies under 50 people, each with its SOP, its KPI and what broke.
- The Agentic OS — what a whole company built duty-by-duty on this test looks like.