Guides · 18 min read · Updated 2026-08-30

12 Questions That Expose a Demo-Only AI Agency

You vet an AI agency by asking for four things in writing — a named workflow, a fixed price, a measurable result, and ownership of what gets built — and by asking twelve questions a demo shop can't answer: production history, what broke, the worst-case monthly run cost, who sees your data, and what's warranted after go-live. This page is the full kit: every question, what a real builder's answer sounds like, when to walk away, and — because we'd rather be tested than trusted — our own answers to all twelve.

Nothing here is gated. Score any vendor 0–2 per question: 20 or better, proceed; 14–19, fix the gaps in writing first; below 14, walk. Including if the vendor is us.

By Ashutosh Upadhyay, founder of Cognio Labs. The four commitments are the buyer community's own checklist, not ours — we just committed to them in writing.

The four commitments — demand all of them

Buyers in small-business communities circulate this checklist among themselves. No signature without all four, in writing:

  1. 1. A named workflow. Not “AI transformation” — the specific process, named in the contract.
  2. 2. A fixed price. Build quoted against written acceptance criteria, run cost capped.
  3. 3. A measurable result. One KPI, with its current reading, before anyone builds.
  4. 4. You own what gets built. Code, prompts, workflows, and the accounts they run in.

And one honest addition from us: any vendor should also name a thing agents can't do yet. Ours: an agent can't rescue an undocumented process — if the rules live in your head, the writing-down comes first, and no model fixes that.

I. Proof of production

A demo takes a weekend. Production takes discipline the demo never shows. These three separate the shops that have run agents for months from the shops that have recorded them for minutes.

1. “Show me one agent that has been running in production for 90+ days — and its numbers.

Anyone can show software working. Only a builder can show software that has kept working after the novelty wore off, the edge cases arrived, and the client stopped watching it.

A real builder answers like

Names a specific deployment, the KPI it serves, and roughly what that number reads now. Offers a walkthrough of the workflow, anonymized if it must be.

Walk away if

Every example is a screen recording. Every client is too confidential to describe even in outline. The portfolio is demos of what the software could do, not records of what it did.

How we answer it

We publish nine of them, with the SOP each runs, the KPI it serves, and what broke, in our AI agent examples guide. The one we cite most: a 65-year-old non-technical lawyer running three to five agents across intake, drafting and billing — self-sufficient in about two weeks, saving five to ten hours a week, for over a year now. Where we don't have a defensible number, that guide says so in the table rather than inventing one.

2. “What broke in production on a past build, and what did it cost?

Every production system breaks. A vendor with no failure story either hasn't shipped anything real or won't tell you the truth about it. Both answers matter.

A real builder answers like

A specific failure, a number attached, and what changed in how they build because of it.

Walk away if

“Nothing has ever really gone wrong.” That sentence should end the call.

How we answer it

A 20–50-person client rolled out a personal agent per employee. Token spend hit roughly $3,000–$5,000 a month — always-on agents burning tokens on heartbeats, polling and memory refresh while nobody asked them anything — and the program was abandoned inside two months. What changed for us: shared departmental agents first, per-role budgets, idle loops killed, cheap tasks routed to cheap models. We'd rather you hear that from us than find it in your own bill.

3. “What will you refuse to build?

A shop with no disqualifying rule sells to everyone, which means it sells demos. The refusal rule is the fastest proof that a vendor has watched builds fail.

A real builder answers like

Names the rule, and it costs them revenue. You should be able to imagine failing their test.

Walk away if

Yes to everything, company-wide rollout in month one, no questions about your process.

How we answer it

Any role that fails the three-question test: no determined KPI, no written SOP, or work that isn't mostly digital. About half the tasks owners bring us fail it, and we say so before quoting. We also won't put an agent on judgment calls where a wrong answer is irreversible — client-facing legal positions, final refund authority above a threshold, anything that can't be undone by a human the same day. You can run the free version of our test yourself before talking to anyone.

II. Cost honesty

The build price is the visible number. The run cost is the one that kills programs. A vendor's willingness to talk about the second number, worst case, in writing, tells you nearly everything.

4. “What will this cost to run each month, worst case — and will you put a cap in writing?

Agents consume tokens every time they act, and badly-scoped ones consume them while doing nothing. “No surprise bills” is only real when a ceiling and an alert exist on paper.

A real builder answers like

Quotes two lines — build and run — with the run line split into model spend, infrastructure and maintenance, plus a monthly cap with an alert that lands in a named inbox. Can describe what a bad month looked like on a past build.

Walk away if

One number on the proposal. “It depends on usage” with no ceiling attached. Any hesitation at the words “in writing.”

How we answer it

Two lines, always. From practitioner-reported invoices we've published: a simple single-purpose agent runs about $300–$800 a month on top of a $1,500–$5,000 build; production-grade builds start around $8,000. Every deployment gets a per-role token budget and a spend alert, because we've watched the alternative: the $3–5k-a-month bill in question 2 is what “depends on usage” costs.

5. “Is the build a fixed price — and what exactly triggers a change order?

Open-ended time and materials on a first agent build transfers all the risk to you, on the project where you can least judge progress.

A real builder answers like

Fixed price against written acceptance criteria, agreed before the build starts. Change orders only when the scope changes, in writing, priced before the work.

Walk away if

A price in the first call before anyone has read your SOPs — or no price at all until you've signed something.

How we answer it

Fixed price, and the acceptance criteria get written down before we build anything. If what we deliver misses them, you don't pay. We quote after reading how the work is done today, not before.

6. “Which model runs each task — and who decides when a cheaper one will do?

Routing everything to the biggest model is the single most common way agent running costs balloon. Model choice is an engineering decision your vendor should be making deliberately, task by task.

A real builder answers like

A cheap default model, a short written list of steps that justify an expensive one, and a review of the routing after the first month of real usage.

Walk away if

“We use the best model available” — for everything. That's not quality, that's your budget doing their engineering.

How we answer it

Cheap default, expensive by exception. Routing cheap tasks to cheap models is one of the four things we changed after the token blowup, and it's now in every build spec we write.

III. Security and data

For most owners this is the deal-blocker nobody addresses head-on: who sees what. Ask it plainly and watch whether the answer names people and systems or waves at the word “encrypted.”

7. “Who — at your company and in your toolchain — can see my data?

Your client data will pass through the vendor's people, their model providers, and every third-party tool in the chain. Each of those is an answer they should already have written down.

A real builder answers like

A named list: which of their staff, which model provider under which data-processing terms, which tools, what's stored where and for how long. Scoped credentials by default.

Walk away if

“Don't worry, everything's encrypted.” Encryption answers a different question. If they can't say who can see it, the answer is: more people than you'd like.

How we answer it

Named people per engagement, scoped credentials per system, and the model provider's data-processing terms in the spec. We don't train anything on your data, and access is per user and per department — never one shared pool.

8. “How are my existing permissions mirrored? Can the agent read what its user can't?

An agent connected with admin credentials can surface any document to any employee who asks nicely. Permissions have to be designed before anything is connected — after is too late.

A real builder answers like

The agent inherits the permissions of the person using it, enforced at the connector level. They can explain how, specifically, in your stack.

Walk away if

One super-admin API key “for simplicity.” Simplicity for whom?

How we answer it

We learned this one in production: a shared agent instance at a client became a shared-secrets problem — one pool of credentials, every user effectively seeing everything. The fix, now standard in our builds, is isolation per user and per department with scoped credentials. An agent must never be able to read what its user can't.

9. “What happens to credentials, data and accounts when we part ways?

Offboarding is where lock-in hides. If the workflows live in the vendor's tenancy and the keys live in their vault, leaving them means rebuilding.

A real builder answers like

Everything runs in your accounts from day one. Offboarding is a runbook: keys rotated, access revoked, and you tested that runbook before the final payment.

Walk away if

It all runs in their environment and data export is “a conversation we can have.” That conversation has a price, and you'll learn it at the worst moment.

How we answer it

Your repository, your keys, your prompts, your accounts — and a runbook you tested before the final payment. If we disappeared tomorrow, your agents wouldn't.

IV. Ownership and terms

The four commitments buyers now circulate as their own checklist: a named workflow, a fixed price, a measurable result, and you own what gets built. No agency should get your signature without all four in writing.

10. “Do I own what gets built — code, prompts, workflows, and the accounts they run in?

Custom work billed as a project should be yours. Some shops build your workflows inside “their platform” and convert a build fee into a hostage situation with a monthly ransom.

A real builder answers like

Yes, in the contract: code, prompts, workflow definitions, and the accounts everything runs in. Any genuinely proprietary platform component is named and separable.

Walk away if

You're “licensing their solution.” Ask what happens to your workflows if you stop paying, and time how long the answer takes.

How we answer it

Yes. It's one of the four commitments, verbatim: you own what gets built. Code in your repository, prompts and workflow definitions in your accounts.

11. “What is warranted after go-live, and for how long?

The gap between “we delivered” and “it keeps working” is where agent projects die. A vendor with no written warranty is telling you which side of that gap they plan to stand on.

A real builder answers like

A named warranty period, a definition of what counts as a defect, and a hard commitment on fixes — separate from any optional maintenance retainer.

Walk away if

The warranty question gets answered with a retainer quote. Fixing their own defects is not a subscription service.

How we answer it

Every agent we hand off carries a 90-day warranty: defects in our work get fixed at no charge, we fix first rather than argue, and the fix has a hard cap — if an accepted defect is still broken after 15 business days, you get that agent's build fee back. Maintenance beyond defects is optional and priced separately, only ever after a build.

12. “Can I talk to someone who went from pilot to production with you?

Most AI pilots never reach production. A reference who crossed that gap with this vendor is worth more than any case-study page — and how a vendor handles not having one is worth almost as much.

A real builder answers like

Yes — or an honest no, with what they can show instead and no theatre about it.

Walk away if

References are promised on every call and never materialize, or every “reference” turns out to be a testimonial paragraph they wrote themselves.

How we answer it

Sometimes, honestly, no — we're eight people, our production deployments are countable, and not every client takes calls. What we offer instead: the published examples with their failures included, these twelve answers in writing, and the cheapest possible way to test us — a $1,500 fixed-price audit, one week, ending in a build-or-don't-build verdict, credited in full to a build. And run this kit on us in the first call. If we fail your scoring sheet, you've spent nothing.

How to score a vendor on these questions

0 for a dodge, 1 for a partial answer, 2 for a specific one. Out of 24:

ScoreVerdict
20–24Proceed. This is a builder. Get the four commitments into the contract and start small.
14–19Proceed only with the gaps fixed in writing before signing — a missing cap or a vague warranty doesn't improve after the deposit clears.
0–13Walk away, whatever the demo looked like.

Three zeros end the conversation regardless of total: no written run-cost cap (question 4), you not owning what gets built (question 10), and a shop that refuses to build nothing (question 3).

Where we fail our own kit

Score us honestly and we drop points too. We have no reviews on Clutch or any other third-party directory, so everything you can read about our work is published by us — exactly the pattern question 1 warns about. Our named case studies are software product builds for technology companies, not internal agent rollouts at 20–50-person firms, so they prove we can build rather than proving we've done your exact job. And we're eight people, so the bus-factor question is fair; our only real answer is question 9's ownership rule — your repository, your keys, your prompts, a runbook you tested before the final payment.

If another shop answers all twelve better than we do, hire them. That outcome is this page working as intended.

Take the kit into your vendor calls

Everything above stays free on this page. If you want the kit as files — for a meeting, a forward, or your own edits — leave a first name and an email:

Free, in exchange for an email

You get the printable scoring sheet (all 12 questions, 0–2 boxes, the red-flag thresholds), the complete kit as an editable markdown file, and a paste-ready email that puts the twelve questions in front of any vendor before your first call.

Two follow-up emails at most, written by a person. No sequence, no daily newsletter.

The cheapest way to test any vendor — including us

Any agency should pass this kit before you sign a $20,000 build. The lowest-risk way to test one is a small fixed-price engagement with a written verdict at the end. Ours is the Agent Readiness Audit: $1,500, one week, the KPI/SOP test run across your whole operation, ending in a build-or-don't-build answer — done for you, across every role, and credited in full to a build within 30 days. Details on the audit page, or and open with question 1.

Frequently asked questions

Are AI automation agencies a scam?

Some are, most aren't, and the useful question is different: is this specific shop a builder or a demo shop? The scam isn't usually fraud — it's a prototype sold at production prices, built by someone who has never run an agent past the first month. The twelve questions on this page separate the two in one call, because a demo shop can't fake production history, a written run-cost cap, or a defect warranty. Score the answers 0–2 each: 20 or better, proceed; 14–19, proceed only with the gaps fixed in writing; below 14, walk.

How do I vet an AI agency before hiring one?

Send the twelve questions by email before the first call, then score the answers in the call: 0 for a dodge, 1 for a partial answer, 2 for a specific one. Three answers are instant red flags whatever the total: no written cap on monthly run cost, you not owning what gets built, and a shop that refuses to name anything it would refuse to build. A real builder will enjoy several of these questions. A demo shop will try to get past them to the screen share.

What is the single biggest red flag when hiring an AI agency?

One number on the proposal. An agent has two costs — the build and the monthly run — and a vendor who quotes only the first either doesn't know the second or doesn't want you to. We've watched the run line reach $3,000–$5,000 a month on a rollout nobody had capped; the program died inside two months. Ask for the worst-case monthly number, a written ceiling, and an alert that lands in a named inbox.

Why would an agency publish answers to its own interrogation kit?

Because we can and most can't — honest answers to these twelve cost a demo shop the deal. Publishing ours costs us too, which is what makes it credible: question 12's answer includes a plain no sometimes, and further up the page we admit we have no third-party reviews and only eight people. If our answers read like marketing anywhere, hold us to the same 0–2 scoring you'd use on anyone else.

What is a fair price for an AI agent build?

From practitioner-reported invoices we've compiled: a simple single-purpose agent runs $1,500–$5,000 to build plus $300–$800 a month; a production-grade build with written acceptance criteria runs $8,000–$15,000. Below those floors you're usually buying a demo; far above them, at 20–50 people, you're usually buying an enterprise firm's overhead. A paid diagnostic before a big build should be fixed-price and credited — ours is $1,500, one week, credited in full to a build within 30 days.

Do I have to give you my email to use the kit?

No. All twelve questions, the answer keys, and our own answers are on this page free. The email gate covers only the take-away files: the printable scoring sheet, the editable markdown version of the kit, and a paste-ready email that sends the questions to any vendor before a call.

Related reading