For business For enterprise Solutions Apps Pricing Developers Blog Docs Launch a workspace
Blog / Agentic ERP

How to evaluate an AI ERP without falling for the marketing

One test survives any demo: ask the vendor to complete a real outcome, across modules, with nobody at the screen, using an agent you bring. Here is how to choose the outcome, run the day, score it and read what the score means.

6 min readUpdated September 4, 2026Sois engineering, the team that builds the platform

A boardroom table before a meeting: dark wood, four printed proposals face down in a row, a jug of water, a single pencil, tall windows with grey morning light
Short answer

Evaluate an AI ERP by asking each vendor to complete one real outcome, end to end, with nobody at the screen. Choose the outcome yourself, make it cross at least three modules, bring the agent your team already uses, and state the request once. Then open the records, sign in as a restricted user and repeat, and read the log. Score what you saw on a fixed card. A product that needs a person to click through part of the outcome is an assistant; a product that finishes it and shows its work in the log, within your permissions, is an agentic system.

The reason this works is that everything a slide can claim, the test can verify or falsify in an hour, in your environment, on your data. It also has the useful property of being the same test for every vendor, which is what makes the scores comparable.

The one test that survives a demo

A demo is a performance. The presenter chose the data, rehearsed the request, and knows which features to steer away from. None of that is dishonest, and none of it tells you what the product does when it is your data and your request. If you want to know how to evaluate an AI ERP, the short answer is to take the request away from the presenter.

Pick one outcome that matters in your business and that a competent junior could do in an afternoon: invoice a customer and chase them, book stock against a purchase order and match the supplier's invoice, or turn an accepted quote into a job, a schedule and a deposit request. Bring the agent your team already uses, connected from outside the vendor's product over the Model Context Protocol, the open standard the mainstream clients speak. State the request in one message and take your hands off the keyboard. Whatever happens next is the evaluation.

Three things can happen. The outcome completes and the records are right. The system does part of it and hands the rest to a person, usually at a module boundary or at a Save. Or the vendor's own agent completes it but nothing from outside can connect. Each of those is a different product, and each scores differently on the card below.

Choosing the outcome

The outcome does most of the work, so choose it before you speak to any vendor and use the same one for all of them. It should meet four conditions.

  • It crosses modules. At least three: a request that stays inside one screen tests the copilot, not the agent. Invoicing plus contacts plus email plus a scheduled task is a good shape. Stock receipt plus purchase order plus supplier invoice matching is another.
  • It has a write in it. Reading and summarising is the easy half. The evaluation is about whether the system will let an agent change records, and whether it checks who is asking before it does.
  • It has an obvious exception. Plant one: a customer with two near-identical records, or a delivery quantity that does not match the order. You want to see whether the agent asks, guesses, or silently corrects.
  • It has a follow-up in the future. A reminder, a chase, a scheduled check. That tests whether the agent can leave something for later and whether the system will run it when the date arrives.

Write the request down in plain language before the day, exactly as an operations lead would type it, and do not let the vendor edit it. A vendor who asks to rephrase the request is telling you where the edges are.

Running the day

The sequence below takes about an hour per vendor and needs nobody technical. Insist on doing it in a trial workspace you control, on data you loaded, with your own agent connected. If any of those three is refused, that refusal is a result and goes on the card.

  1. Connect your own agentAdd the vendor's MCP server to Claude, ChatGPT in developer mode, or whichever client your team uses, and sign in. Note whether it is an OAuth sign-in or a token you have to paste, and whether the vendor could do it at all.
  2. List the toolsAsk the agent what it is allowed to do. Read the list for length and coverage of the modules in your outcome. Then sign in as a restricted user and ask again; the list should shrink.
  3. State the outcome onceType the prepared request as the full-permission user and stop. Answer only genuine questions from the agent, such as which of two matching customers you meant. Count the hand-backs.
  4. Check the recordsOpen every record the request should have touched. Confirm the invoice, the email, the stock movement, the task. Then run the same request as the restricted user and watch where it is refused.
  5. Read the log and the meterFind the log entry: who asked, which agent, which tools, inputs, results, refusals. Find what the run cost and whether a cap could have stopped it.
  6. Score on the dayFill in the card before you leave the room. Memory is kind to good presenters.

The scorecard

Score each row 0, 1 or 2. Zero means you did not see it; one means you saw it with a caveat; two means you saw it clearly, in your environment, with the evidence in front of you. Ten rows, so the maximum is twenty. Do not weight the rows before you have run the test on at least two products; weighting in advance is how the marketing gets back in.

RowWhat a 2 looks likeCommon reason for a 1 or 0
Outcome completedAll records correct, no person at a screenHand-back at a module boundary or a Save
Your agent connectedYour own MCP client, OAuth sign-in, no token to pasteVendor's assistant only, or a pasted API key
Tool list is realLong, typed, covers your modules, machine-readableA dozen features, or a list only on a slide
Permissions per callRestricted user refused at the call; the rest completesRefusal only on screen load, or an approval that goes through
Exception handledThe planted ambiguity produced a question, not a guessSilent correction, or a wrong record chosen
Future action scheduledThe follow-up exists and will run on the dateA note in a summary, nothing in the system
Log is completeWho asked, agent, tools, inputs, results, refusalsRecords changed but no call sequence; a generic integration user
Spend is cappedA per-integration cap the agent cannot exceed; per-action costMonthly total only, or no cap
Own-agent costNothing charged when your agent does the reasoningAI charged regardless of who reasons
Extendable by othersA third-party app installs and appears in the tool listRoadmap or custom work only

Ten rows, two points each. Score every vendor on the same outcome, on the same day if you can, and compare totals only after every row has a number.

Reading the result

Totals matter less than the pattern in the first four rows, because those four decide what kind of product you are looking at. A product that scores zero on the outcome and zero on your agent connecting is an assistant inside a screen, and the rest of its score describes controls it does not need. It may still be the right purchase if what your team wants is a faster screen, but you should buy it as that.

A product that completes the outcome but scores zero on your agent connecting is agentic with a closed door. It works, on the vendor's terms, with the vendor's agent, at the vendor's price for reasoning. The question to ask yourself is what happens in two years when your team's agent is the one they want to use, and whether the vendor has said anything about opening the door.

A product that scores two on all four is agent-native, and the remaining six rows are where the real comparison happens: how well it handles the exception you planted, how complete the log is, whether spend can be capped, what it charges when your agent reasons, and whether outsiders can extend it. Two agent-native products can differ a lot on those six, and they are the rows that predict what living with the product will be like.

One more reading. If a vendor declines the test, or offers a recorded version, or asks to substitute their request for yours, score the rows you could not observe as zero and say why in the notes. Declining is information. A product that can do this will want to show you.

What we would show you

Sois is one product you could run this evaluation against, and since we build it we can say what you would see. A workspace is an MCP server; you add its address to Claude, ChatGPT or another client and sign in once over OAuth, with no token to paste. The tool list is filtered by your role before the agent sees it and checked again when each tool runs; access fails closed. Spend can be capped per integration, every action is logged, and when your own agent does the reasoning the platform performs no AI on your behalf and charges nothing for it. Apps from the marketplace appear in the tool list once installed. Here is the invoicing outcome from the flow above, as it runs.

Claudeconnected toapp.sois.aiover MCP
YouInvoice Acme for the September work, email Sarah Cole, and chase it if it isn't paid by the 18th.
Agent
  • Reading the project's billable work and the customer record
  • Invoice created from the billable lines
  • Sent to the customer contact by email
  • Chase scheduled for the due date
Done.
Sois records
INV-1064Invoice created, Acme Ltd
Sarah ColeEmail sent from the workspace inbox
TaskChase if unpaid on September 18

Four tools across accounting, contacts, inbox and tasks, with the log showing each call. Run it as a user who cannot raise invoices and the first write is refused, with the reason.

Then score it on the same card as everyone else. The point of a fixed test is that it does not care who built the product, and an evaluation you can defend is one that treated every vendor, including this one, the same way.

Questions people ask

How long does this evaluation take per vendor?

About an hour once the outcome is written and a trial workspace with your data exists. Loading realistic data and planting the exception is the preparation; the test itself is one request, two sign-ins and a read of the log.

What if the vendor cannot let my own agent connect?

Score that row zero and run the test with their agent so you still see the outcome, permissions, log and cost rows. Then decide whether a closed door is acceptable for your team, and ask the vendor whether and when it opens.

Should I weight the scorecard rows?

Not before you have run the test on at least two products. Score all ten rows first, then decide which rows matter most to your operation. Weighting in advance is how a strong presentation gets back into a decision the test was meant to keep it out of.

Sources
  1. Model Context Protocol specification: tools tool lists, authorisation-dependent listing, and the recommendation that a human stays in the loop with the ability to deny tool calls
  2. Anthropic: getting started with custom connectors using remote MCP adding a remote server, OAuth sign-in and per-tool approval in Claude
  3. OpenAI: ChatGPT developer mode full MCP client support in ChatGPT; write actions require confirmation by default
  4. Sois documentation: the workspace MCP server what the test observes on one implementation: OAuth, role-filtered tools, fail-closed execution, budget caps

This article is reviewed when the products it describes change. Next scheduled review: December 4, 2026.

Start

Explore the Sois platform.

How the platform works, what the permission layer does, and what it costs, in plain terms.

  • Free to start
  • Bring your own agent
  • No vendor lock-in