How to evaluate an agentic AI platform
Six criteria, three grades of evidence and one bounded pilot
Evaluate an agentic AI platform on six criteria: shared context, human and agent collaboration, governance, bounded autonomy, a learning loop and depth in your domain. Grade every claim as documented, measured or asserted, run one bounded pilot with the controls your process needs, and judge the result on the business number the team owns rather than on the demo.
Overview
Every agentic AI platform demos well. The agent plans, the tools fire, the result appears, and the room nods. The evaluation starts after the demo, when a buyer asks what the agent may do alone on a Tuesday, who answers when it is wrong, and whether the number moved. This guide is the method the Nagent comparison hub applies to 64 platforms, written so a buyer can apply it to any vendor, Nagent included.
Start from the work, not the demo
Write down, before the first call, the one workflow the pilot will run and the one number it will be judged on. Meetings booked. Pages cited in AI answers. Campaigns shipped without an edit. Tickets resolved without a hand-off. A pilot without a number is a demo with a longer run time.
Then write the always-human list: the actions that must wait for a person whatever the agent earns. Spending more budget, sending to a customer, changing a live site are the usual three. Every vendor conversation is measured against that list.
The six criteria
Score each criterion on evidence, in this order, because it is the order a team feels the gaps.
- Shared context. Does every agent read one memory, and is there one ledger of the decisions people took? Ask to see the memory and to correct one entry.
- Human and agent collaboration. Can two people join an agent's task in progress, change its direction, hand it off and leave an instruction it keeps? Ask for a live session, not a recording.
- Governance. Roles, per-agent budget caps, a kill switch, tenant isolation, an append-only audit record, a choice of data residency. Ask to see the audit entry an approval leaves.
- Bounded autonomy. Graded levels of what an agent may do alone, earned on a record and taken back on a slip. Ask what this agent earned its level on, and what would take it away.
- A learning loop. Does the agent improve from feedback without retraining, is it scored against the number it was hired for, and is a change evaluated before release? Ask to see one documented downgrade.
- Depth in the work. Does the platform arrive knowing your function, its tools and its proof, or is it a canvas? Count the agents you would have to build before the first result.
Three grades of evidence
Every claim a vendor makes lands in one of three grades:
| Grade | Meaning | Where it comes from |
|---|---|---|
| Documented | The vendor's own pages or documentation state it, and you read them | Product pages, docs, pricing pages, security pages |
| Measured | A number with a method, a sample and a date | Case studies with baselines, independent reviews, your own pilot |
| Asserted | Someone said it | Decks, calls, testimonials without a number |
Most decks are asserted. Most dossiers in the comparison hub are documented, and they say so. Very little in this market is measured, which is why the pilot exists.
The pilot
- One workflow, one number, one always-human list, agreed before day one.
- Every agent at suggest only. Count the approval hours a person spends in week one; they are the baseline the ladder is measured against.
- Controls switched on as the vendor documents them: budget cap, kill switch, audit record. Trigger each once on purpose and keep the entry.
- Four weeks. Record what the agent carried to the result, what a person intervened in, what the controls queued or stopped, and what the number did.
- One promotion, if the record supports it, in week three. Watch what the agent may now do alone and whether the approval hours fell.
- A written result on day thirty: the number, the hours, the interventions, the breaches, and whether you would run the workflow on the platform for a quarter.
Pricing and security questions
On pricing, ask what the unit of work is, what is included, what the overage costs, whether the rate is published, and whether the pricing page matches the proposal. Compare on the monthly cost of the pilot workflow. Never compare on a seat price.
On security, ask where the data lives and whether you can choose, whether a private cloud or on-premise option exists, how tenants are isolated, what the audit record holds, and the exact status of each certification: held, in audit or planned, with a date. A badge without a status is a claim. Nagent's own answer is that SOC 2 Type II and ISO 27001 programmes are underway, with tenant isolation, role-based access, an append-only audit and private VPC deployment available today.
Using the comparison hub
The hub applies this method to 64 platforms. Each dossier grades its claims, each guide compares a platform with Nagent on the six criteria, the category pages list the players in each stack role, and the side-by-side page reads any two or three on the same rows. Start from the category your pilot workflow belongs to, shortlist two or three, and take the pilot above to each.
Frequently asked questions
What are the six criteria for evaluating an agentic AI platform?
Shared context, so every agent reads the same memory; collaboration, so people and agents work in one place; governance, with roles, budgets, a kill switch and an audit record; bounded autonomy that is earned; a learning loop without retraining; and depth in the work your team does. Score each on evidence, not on the demo.
How do I tell a documented capability from a claim?
Grade it. Documented means the vendor's own pages or documentation state it and you can read them. Measured means a number with a method, a sample and a date behind it. Asserted means someone said it in a meeting. Most decks are asserted; most dossiers in the comparison hub are documented; very little in this market is measured.
How long should a pilot run?
Four weeks on one workflow is enough to see whether the agent carries work to the result, how often a person had to intervene, what the controls did when something went wrong, and whether the number moved. Longer pilots without a fixed number drift into demos.
What should the pilot measure?
The business number the team owns, such as meetings booked, pages cited in answers or campaigns shipped, beside the approval hours a person spent, the actions the controls queued or stopped, and the agent's record over the window. A pilot that reports only output has measured the wrong thing.
What pricing questions matter?
What the unit of work is (a seat, a credit, a conversation, an outcome), what is included, what the overage costs, whether the vendor publishes the rate, and whether the figure on the pricing page matches the figure in the proposal. Compare on the monthly cost of the pilot workflow, never on a seat price.
What security and deployment questions matter?
Where the data lives and whether you can choose; whether a private cloud or on-premise deployment exists; tenant isolation; role-based access; an append-only audit record; and the exact status of each certification, stated as held, in audit or planned, with the date. Treat a badge without a status as a claim.
Sources
- Nagent comparisons research, the six criteria and how 64 platforms were reviewed https://nagent.ai/comparisons/research
- The Earned Autonomy Ladder https://nagent.ai/artefacts/earned-autonomy-ladder
- Nagent security and compliance https://nagent.ai/dev-technology/security-and-compliance
Cite this page
Plain:
Nagent AI. How to evaluate an agentic AI platform. Definitions, no. 2. 2026. https://nagent.ai/artefacts/evaluating-an-agentic-ai-platform
BibTeX:
@misc{nagent2026evaluatinganagenticaipla,
author = {Nagent AI},
title = {How to evaluate an agentic AI platform},
series = {Definitions},
year = {2026},
url = {https://nagent.ai/artefacts/evaluating-an-agentic-ai-platform},
note = {Published 2026-10-04, updated 2026-10-04}
}The direct answer at the top of this page is written to be quoted as one sentence with this URL as its source.
About Nagent
Nagent is Multiplayer AI for end to end growth: a team of AI coworkers and your own people, working together in one workspace across marketing, sales and customer experience. Three things make it different. You approve the AI coworkers' work until they earn the right to act on their own. They carry the work all the way to pipeline and customers, not just content. And where your plan includes it, a Nagent marketer joins your team and owns the number with you. Founded in Bengaluru in 2024, Nagent is an Anthropic partner, holds four filed patents on orchestration and memory, and deploys in the customer's private cloud.
Published 4 October 2026. All rights reserved. Quote with attribution to Nagent AI and a link to this page.
