Agent washing is Gartner's term for vendors rebranding chatbots, robotic process automation, and AI assistants as "agentic AI" without real agentic capability, and it is not a fringe problem: Gartner estimates only about 130 of the thousands of vendors currently marketing agentic AI actually build genuinely agentic systems. A real agent has six traits working together: autonomy, planning, real tool use, persistent memory, a feedback loop, and documented human-in-loop intervention rates. Score a vendor 0 to 5 on each. Under 10 total means you are looking at a relabeled script. This matters beyond marketing accuracy: the SEC and FTC have already brought cases against companies for overstating exactly these claims.
Before you sign anything, see how we approach AI strategy and executive advisory, including vendor evaluation, so this decision gets made with a checklist instead of a sales deck. The full framework, plus the exact questions to ask in a demo, is below.
What is agent washing?
Gartner's own definition is precise: agent washing is "the rebranding of existing products, such as AI assistants, robotic process automation (RPA) and chatbots, without substantial agentic capabilities." The name deliberately echoes greenwashing, because the problem is the claim, not the underlying technology. A chatbot is a perfectly fine chatbot. It becomes a problem the moment it is sold to you as something it is not.
The scale is the part most buyers underestimate. Of the thousands of vendors currently marketing themselves as agentic AI providers, Gartner's June 2025 research puts the number that actually qualify at around 130. That means the overwhelming majority of the products carrying the word "agent" in their pitch deck do not meet the bar the word implies. If you are evaluating a vendor right now with no way to check their claim, the odds are not in your favor.
What actually makes an AI agent real?
Six capabilities, working together, separate a genuine agent from a relabeled chatbot or automation script. A vendor's marketing page will claim all six. A five-minute demo, run correctly, will only confirm the ones that are real.
| Trait | Washed (score 0) | Genuinely agentic (score 5) |
|---|---|---|
| Autonomy | Responds only when prompted; a human initiates every step | Owns a goal end to end, initiates steps itself, escalates only at defined checkpoints |
| Planning | Follows a fixed flow authored entirely by humans in advance | Decomposes a novel goal into steps and replans when a step fails |
| Tool use | No live tool calls; a human reads the output and executes it manually | Selects the right system or API per step, sets the parameters, and handles errors itself |
| Memory | Resets to a blank state every session | Retains durable state across sessions that shapes future decisions |
| Feedback loop | No self-check; all verification is done by a human afterward | Evaluates its own output against the original goal and corrects failures |
| Human-in-loop transparency | Intervention rate is undisclosed or unmeasured | Checkpoints are documented and intervention rates are measured and reported |
Score a vendor 0 to 5 on each of the six, for a total out of 30. A score of 0 to 10 is a rebranded chatbot or RPA tool wearing the word "agent." A score of 11 to 20 is a partial or assisted agent, genuinely useful, but not the fully autonomous system it may be marketed as. A score of 21 to 30 is genuinely agentic. Run this scorecard on any vendor claiming "agentic AI" before you sign, and you will have an evidence-based answer instead of a sales pitch's word for it.
Want to see what a 21+ score actually looks like? Hire AI agents built with real planning, tool use, and memory, not a demo script.
What are the most common disguises agent washing wears?
Agent washing tends to show up in three recognizable shapes, each borrowing the word "agent" for a system that was built to do something narrower.
- A chatbot relabeled as an agent. It generates advice, drafts, or answers in a conversation, but it never actually acts. It can tell you what email to send; it cannot send the email, check whether it landed, and follow up on its own. If the product's entire job ends at generating text, it is a chatbot, however fluent the text is.
- RPA rebranded as agentic. Robotic process automation executes a fixed, scripted sequence of steps reliably and fast, which is genuinely valuable for stable, high-volume, rule-based work. What it does not do is reason about a step it was not scripted for or adapt when the underlying process changes. Calling a script "agentic" does not give it judgment it was never designed to have.
- A workflow-automation tool presented as autonomous. These products connect several systems together and move data between them on a trigger, which is useful plumbing. But orchestrating a fixed sequence between tools is not the same as planning a novel path through a goal and replanning when a step fails. The tell is usually the same: ask what happens when a step encounters something the flow was not built for, and the answer is "it stops and someone gets an alert," not "it figures out the next step."
None of these three are bad products. A good chatbot, a good RPA script, and good workflow automation are all legitimate, useful tools that solve real problems at a lower cost than a full agent. The issue is only ever the label. Priced and sold honestly as what they are, they are fine. Priced and sold as autonomous agents, they set an expectation the architecture cannot meet, and that gap is exactly where the Presto and DoNotPay cases came from.
What five questions expose a washed product in a live demo?
The scorecard tells you what to look for. These five questions, brought into an actual vendor demo, tell you whether what you are being shown is real.
- "Show me the decision trace on your last ten production decisions." Ask for the input state, the logic applied, the output, and a timestamp. A genuinely agentic platform pulls this from production in real time. A washed product defers to "implementation" or shows you demo-only metrics instead.
- "How does the system resolve conflicting outputs between two agents?" A real platform demonstrates exception routing built into the platform itself. A washed product needs a human to manually configure a resolution for every conflict, because there is no real coordination layer underneath.
- "Show me 30 days of live production monitoring: decision volume, exceptions, baseline drift, alerts." A genuinely agentic platform displays this continuously, because it has to monitor itself to operate safely. A washed product offers you anonymized data later, if at all.
- "Can my business team change a policy without filing an engineering ticket?" Independent policy updates are a sign of real governance built into the product. If every change requires an engineer, what is being called "governance" is closer to hard-coded logic with a dashboard on top.
- "Retrieve a specific decision from six months ago, with full context." A real platform retrieves it instantly from its own records. A washed product has to reconstruct it from scattered logs, if it can be reconstructed at all.
None of these five questions require you to understand model architecture. They require the vendor to show you something real, and a washed product simply cannot, because the infrastructure the answer depends on was never built.
Has agent washing actually caused real harm?
Yes, and it is no longer just a reputational risk. Two U.S. regulators have already brought formal cases against companies for overstating exactly the kind of autonomy claims this checklist is designed to catch.
The SEC charged Presto Automation in January 2025 for materially false and misleading statements about Presto Voice, its AI drive-thru ordering product. The company marketed high "automated order completion" and "non-intervention" rates. What those rates actually measured was whether an order was completed without on-site restaurant staff getting involved, not without any human involvement. In reality, off-site human agents, largely based in the Philippines, were assisting in more than 70% of the customer interactions the company was marketing as autonomous AI. That is a direct failure of the human-in-loop transparency trait on the scorecard above: a real intervention rate existed, it was large, and it was not disclosed.
The FTC finalized an order against DoNotPay the same month, prohibiting its deceptive "robot lawyer" claims. DoNotPay had marketed itself as able to help consumers "sue for assault without a lawyer" and claimed it would "replace the $200-billion-dollar legal industry with artificial intelligence." The FTC's investigation found DoNotPay never tested whether its chatbot's output actually matched a human lawyer's competence, and had not hired or retained any attorneys to validate the product at all. DoNotPay was ordered to pay $193,000 in monetary relief and notify subscribers from 2021 to 2023 about the settlement. That is a direct failure of the planning and tool-use traits: the product claimed to handle a legal goal end to end, a claim it had never demonstrated, let alone tested.
Both cases map cleanly onto the scorecard above. That is not a coincidence. The six traits exist because they are exactly the claims vendors are most tempted to overstate, and exactly the claims regulators are now willing to test in court.
Why does agent washing matter beyond the marketing claim?
Because getting it wrong carries three distinct kinds of risk, and only one of them is embarrassment.
- Trust risk. A failed deployment built on an overstated product does not just hurt that vendor relationship. It erodes confidence in AI generally, inside your own organization, which makes the next genuinely good project harder to greenlight.
- Operational risk. Overestimating a system's real capability in customer service, finance, or security can cost real revenue or create real legal exposure, exactly as it did for the customers left holding legal documents from an untested "robot lawyer."
- Innovation risk. Every dollar and every quarter spent on a washed product is a dollar and a quarter not spent with one of the roughly 130 vendors actually building the real thing, and it makes it harder for those vendors to get a fair hearing the next time, because the buyer has been burned once already.
A product that gets some work done is not the same as a product you can safely hand a business process that has real consequences when it fails. The distance between those two things is exactly what the scorecard measures.
How do I actually run this checklist?
Before your next vendor call, print the six-trait table above and score it live during the demo, not from the pitch deck afterward. Ask all five demo questions in the room, and pay closer attention to what the vendor cannot show you than to what they can. A vendor who is confident in a real product will welcome the decision trace, the production monitoring dashboard, and the six-month-old audit retrieval, because those are the artifacts a real agentic system produces as a byproduct of operating safely. A vendor who deflects, reschedules, or offers you a "custom demo" instead of live production data is telling you the answer without saying it out loud.
If the total score lands under 10, walk away or renegotiate the pitch down to what the product actually is: a well-built chatbot or automation tool, which may still be useful, just not for the price or the promise of a genuine agent. If it lands between 11 and 20, ask specifically which of the six traits is weakest and whether the vendor's roadmap closes that gap on a timeline you can verify. Only a score of 21 or higher earns the word "agentic" without an asterisk.
How to get started
You do not need to become an AI architect to avoid becoming the next agent-washing headline. You need six questions, run consistently, in every vendor demo, before a contract is signed. Score autonomy, planning, tool use, memory, the feedback loop, and human-in-loop transparency, ask the five questions that force a vendor to show you production reality instead of a script, and treat "we'll show you that after signing" as a straight answer to your score.
If you would rather skip the vendor gauntlet entirely, that is the case for working with a partner who is happy to be scored by this exact checklist, because the answer holds up. We build agents with real autonomy, real planning, real tool use, real memory, and a documented, reported intervention rate, the same six traits this article asks you to demand from anyone else. Book a free consultation below and bring the scorecard. We will walk through it with you, live.
