Most enterprise AI projects do not fail on the model. They fail on a decision made months earlier, in a procurement meeting, when nobody could explain how three shortlisted products actually differed.

The numbers back this up. MIT’s 2026 study found that 95% of enterprise generative AI pilots delivered no measurable return. IDC put the pilot failure rate at 88%. Gartner expects more than 40% of agentic AI projects to be cancelled by 2027, and Deloitte found only 21% of organisations have a mature governance model for autonomous agents. Fiddler AI’s production data is the most sobering: agents that succeed 60% of the time on a single run drop to 25% across eight consecutive runs under production load.

None of that is a model problem. It is a fit problem — the wrong platform for the data, the permissions, the deployment constraints, and the budget the organisation actually has.

So we did the work we would want a vendor to do for us: a 15-criteria framework, and a feature-by-feature matrix of all twelve platforms — ours included — built from public documentation and vendor announcements, with every claim linked to its source.

Method and caveat. Features below were compiled from vendor documentation, release notes, and public announcements current to August 2026, and every claim links to its source. This market ships monthly, and almost every enterprise contract is custom-quoted, so treat prices as order of magnitude and verify features against current documentation before you sign. “Partial” means the capability exists but is limited to one ecosystem, gated behind a paid tier, or narrower than a competitor’s equivalent.

First, the market is four markets

The word “platform” is doing a lot of hiding. Four distinct product shapes are sold under it, and they are not substitutes.

Suite-native assistants. AI added to a stack you already bought: Microsoft 365 Copilot with Copilot Studio, Google Gemini Enterprise (which absorbed Agentspace), Salesforce Agentforce 360, ServiceNow Otto — the unified layer that now folds in Now Assist and the Moveworks assistant ServiceNow acquired for $2.85B — and IBM watsonx Orchestrate. Cheapest to start, strongest inside their own boundary, weakest outside it.

Work-AI and search-first platforms. Vendor-neutral layers that index everything and answer across it: Glean (roughly $200M ARR at a $7.2B valuation, 100+ connectors), Dust, Writer, ChatGPT Enterprise, and Sana, now part of Workday. They win on cross-system retrieval and vary enormously on whether they can do anything with what they find.

Sovereign and private-deployment platforms. Built for organisations that cannot send data to a shared cloud: Cohere North, which runs on-premise, in a VPC, or air-gapped on as few as two GPUs, and our own platform. Smaller ecosystems, much harder constraints solved.

Build-your-own toolkits. Onyx, Dify, n8n, Flowise, LibreChat. Free or nearly free at the licence line, excellent if you have a platform team. The cost moves from the invoice to the payroll.

Knowing which of the four you are shopping for eliminates most of the confusion before you write a single requirement.

The 15 criteria

Four questions a buyer has to answer: what does it know, what can it do, who controls it, what does it really cost.

What it knows

1. Connector breadth and freshness. How many systems, and does the index re-sync automatically or drift stale? Ask for the re-sync interval, not the connector count.

2. Permission inheritance. Does the assistant respect the source system’s access rules per document, or flatten everything into one index? This is the most common cause of a stalled security review. Some platforms charge extra for it.

3. Citation and verifiability. Does every answer link back to a source a human can open? Without this you cannot use the output in a regulated process.

4. Answer format. Text only, or tables, charts, forms, slides, spreadsheets? A dataset rendered as a paragraph is a worse answer than a dataset rendered as a table.

What it can do

5. Write operations. Can it take action in your tools, or only read from them? A large share of “AI agent” products are read-only retrieval with an agentic name.

6. Approval model. An explicit pending → approved → executed gate before anything is written, or does the agent act and log afterwards?

7. Identity for actions. Does the agent sign in as the individual user, or as one shared service account? This determines whether an audit can attribute an action to a person. Shared service accounts are the quiet governance failure of 2026.

8. Autonomy modes. Chat only, or scheduled runs and event triggers? Real savings come from work that happens before anyone asks.

9. Extensibility. Custom connectors, MCP support, installable skill packages, sandboxed code execution. Ask what happens when your workflow needs the one integration nobody has built yet.

Who controls it

10. Deployment and data residency. SaaS-only, EU region, your VPC, on-premise, air-gapped. Note that if your cloud provider is US-incorporated, an EU region alone does not give you full sovereignty.

11. Model flexibility. One vendor’s models, or your choice per workload — and can you switch without rebuilding the knowledge base?

12. Governance and audit depth. Workspace isolation, RBAC, SSO and SCIM, replayable audit trail, configurable retention. From 2 August 2026 the EU AI Act’s transparency obligations apply, and deployers of high-risk systems must retain generated logs for at least six months. Log retention stopped being a nice-to-have.

13. Compliance posture. SOC 2 Type II, ISO 27001, ISO 42001, GDPR, HIPAA where relevant. Ask for the trust centre link, not the sales claim.

What it costs

14. Cost model and predictability. Per seat, per conversation, per credit, per vCPU-hour, or a mix. Consumption models are where budgets break, because the cost of an answer varies with how much retrieval and tool use it does.

15. Time to value and who does the work. Weeks or quarters, and does the vendor configure it or hand you a console? This is where the 88% pilot failure rate lives.

Feature comparison

Knowledge and retrieval

PlatformConnector reachPermission-aware retrievalCitationsWeb searchDeep research modeRich answer formats
Mitigate AIFiles, SharePoint, Drive, Notion, Confluence, Jira, web crawl; daily re-syncYesper-workspace isolationYesYesPartialYescharts, tables, forms, tabs streamed in chat (OpenUI)
Microsoft 365 CopilotDeep in Microsoft Graph, thinner outsideYeswithin MicrosoftYesYesYesResearcher and Analyst agentsYesOffice artifacts
Gemini EnterpriseValidated connectors plus BYO-MCPYeswithin GoogleYesYesYesYesWorkspace artifacts
ChatGPT EnterpriseApps and connectors: Salesforce, ServiceNow, enterprise RDBs, vector DBsYesYescompany knowledge cites sourcesYesYesYesdocs, spreadsheets, canvas
Glean100+ connectors, permissions-aware knowledge graphYesstrongest in categoryYesYesYesAgent Sandbox, programmatic tool callingYesslides, docs, Canvas co-authoring, image generation
Dust60–100+ connectors plus MCPYesRBAC with dual-layer agent permissionsYesYesPartialPartial
Cohere NorthBuilt-in connectors plus Compass search indexYesYesPartialPartialYestables, documents, presentations
Agentforce 360Deep in Salesforce; Data 360 for the restYeswithin SalesforceYesPartialPartialYesIntelligent Context over unstructured data
ServiceNow OttoDeep in ServiceNow plus Moveworks enterprise searchYesYesPartialPartialYesAI Data Explorer, plain-language querying
watsonx OrchestrateIntegration catalog plus custom toolsYesYesPartialPartialPartial
Onyx (open source)40–50+ connectors, MCPYesper-document mirroring, paid tierYesYesSerper, Brave, SearXNG, Firecrawl, ExaYesDeep Research modePartial
Dify / n8n / FlowiseBuild your ownBuild your ownBuild your ownYesBuild your ownBuild your own

Agents and automation

PlatformNo-code agent builderMulti-agent orchestrationScheduled runsEvent triggersApproval gate on writesCode execution sandboxComputer use
Mitigate AIYesper-workspace configurationPartialYesYesincluding email eventsYeson every writeYessecure sandboxesNo
Microsoft 365 CopilotYesCopilot StudioYesagent-to-agent via Work IQYesYesConfigurableYesYesGA, browser and desktop
Gemini EnterpriseYesAgent Designer, visual flowchartYesAgent Registry and GatewayYesYesSemantic governance policiesYesPartial
ChatGPT EnterprisePartialagent mode from a promptPartialPartialPartialPartialYesYesagent mode
GleanYesAgent Builder with debug modeYesAgentic Engine 2YesYesVaries by agentYesAgent SandboxNo
DustYesno-code, anyone on the teamYeschained multi-step agentsYesYesVaries by agentYesvia MCP code interpreterNo
Cohere NorthYesassistant builderYesNorth AutomationsYesYesYesYesPartial
Agentforce 360YesAgentforce Builder, Agent ScriptYesAtlas Reasoning Engine, hybrid reasoningYesYesYesPartialPartial
ServiceNow OttoYesYesAction Fabric, "agent of agents"YesYesYesapprovals across departmentsPartialPartial
watsonx OrchestrateYesdrag-and-drop and natural languageYesAgentic Control PlaneYesnative schedulingYesYesYesADK, PythonNo
OnyxYescustom agents with instructions and actionsPartialPartialPartialBuild your ownYesNo
Dify / n8n / FlowiseYesvisual canvasYesYesYesBuild your ownYesPartial

Surfaces and extensibility

PlatformExternal-facing widgetSlack / TeamsMCP supportCustom connectorsSkill or agent catalogAPI and SDK
Mitigate AIYesSSO, guest, or anonymous modesYesSlack and TeamsYesYesbuilt to orderYesinstallable skillsYes
Microsoft 365 CopilotVia published Copilot Studio agentsTeams nativeYesYesPower Platform connectorsYesagent storeYes
Gemini EnterprisePartialGoogle Chat nativeYesBYO-MCPYesYesAgents gallery, Skills RegistryYesAgent Development Kit
ChatGPT EnterpriseNoVia appsYesYesYesapps and connectorsSeparate OpenAI API
GleanPartialinternal-firstYesYesremote MCP serversYesYesLibrary, agent marketplaceYes
DustPartialSlack nativeYesadmin-added MCP serversYesREST, webhooks, OAuth2Yesshared agent workspaceYes
Cohere NorthPartialYesPartialYesflexible APIsPartialYes
Agentforce 360Yesstrong customer-facing story, voice with SIPSlack nativeYesYesMuleSoftYesAgentExchangeYes
ServiceNow OttoYesportal and voice agentsYesYesAction Fabric, runs headlessYesYesYes
watsonx OrchestrateYesYesYesYesYesAgent Catalog, 150+ prebuilt agents and toolsYesADK, Python
OnyxPartialSlack and Discord botsYesincluding its own MCP serverYesOpenAPI actionsPartialYes
Dify / n8n / FlowiseBuild your ownBuild your ownYesYesCommunity templatesYes

Governance and security

PlatformSSO / SCIMPer-user identity for actionsWorkspace isolationReplayable audit trailRetention and residency controlsPublished certifications
Mitigate AIEntra ID, Okta, Auth0, Google, KeycloakYesagent signs in as each person, never a shared accountYesend-to-end, knowledge and users separatedYesevery action, replayableYesconfigurable per workspaceNot published — residency solved by self-hosting
Microsoft 365 CopilotEntra nativeYeswithin MicrosoftYesYesPurviewYesExtensive
Gemini EnterpriseGoogle nativeYeswithin GoogleYesYesYesExtensive
ChatGPT EnterpriseYesPartialYesworkspacesYesdetailed audit logsYesdata residency, no training on your dataSOC 2
GleanYesPartialYesYesYesSOC 2, enterprise controls
DustSAML, OIDC, SCIMYesYesSpacesYes365-day retentionYesEU or US residencySOC 2 Type II, GDPR, HIPAA-ready
Cohere NorthYesYesYesfull data isolationYesYesincluding air-gappedSOC 2 Type II, ISO 27001, ISO 42001, GDPR, HIPAA, CCPA
Agentforce 360Salesforce nativeYesYesYesAgent Health Monitoring, error and escalation ratesYesExtensive
ServiceNow OttoServiceNow nativeYesYesYesAI Control TowerYesExtensive
watsonx OrchestrateYesYesYesYesAgentic Control PlaneYesExtensive
OnyxPaid tiers onlyDepends on setupRBAC in paid tiers; community edition is all-or-nothingPaid tiersYou control it entirelySOC 2 Type II in enterprise edition
Dify / n8n / FlowisePaid tiersBuild your ownFree tiers are all-or-nothingBuild your ownYou control it entirelyVaries

Deployment, models, and commercials

PlatformDeployment optionsModel choiceCost modelReported priceTime to production
Mitigate AIOur cloud, or self-host via Docker Compose, Kubernetes, AWS EKSOpenAI, Anthropic, Google, Mistral, DeepSeek, OpenRouter — per workspace, swap anytimeFlat subscription with per-workspace budget capsFrom €400/month up to 50 users; e-commerce workspace from €600/month; enterprise customUnder a week, vendor-run
Microsoft 365 CopilotMicrosoft cloud onlyMicrosoft-hosted, plus Anthropic Claude in Researcher and Copilot StudioPer seat plus consumption credits$18–30/user/month plus M365 base; Copilot Studio credits $200 per 25,000, or $0.01 eachDays to start, months to govern
Gemini EnterpriseGoogle Cloud onlyModel Garden, 200+ modelsPer seat plus agent compute, memory, storage$21–30/user/month typical, up to $60+; vCPU-hour, GiB-hour and GiB-month metering on topWeeks
ChatGPT EnterpriseOpenAI cloud, residency optionsOpenAI models onlyPer seat, negotiated~$60/user/month reported, ~150-seat minimum, ~$108K/year floor; Frontier customDays to weeks
GleanVendor cloud, private cloud optionsMultiple providersPer seat plus AI add-on plus FlexCredits~$45–50/user/month plus ~$15 AI add-on, ~100-seat minimum; median contract ~$97.5K/yearWeeks to months
DustVendor cloud, EU or US residencyClaude, GPT, MistralPer seat, credits metered$30/seat/month Pro, $150/seat/month Max; enterprise customDays to weeks
Cohere NorthOn-premise, VPC, hybrid, air-gapped, from 2 GPUsCohere modelsCustomCustomWeeks, with deployment engineering
Agentforce 360Salesforce cloud, plus AWSSalesforce-managedPer conversation, Flex Credits, or per seat$2/conversation, or $500 per 100K Flex Credits, or from $125/user/month; Data Cloud extraWeeks to months
ServiceNow OttoServiceNow cloudServiceNow-managedCustomNot publishedMonths
watsonx OrchestrateIBM cloud, hybridwatsonx plus third-party, any frameworkSubscription plus usageEssentials from $500/month; standard customWeeks to months
OnyxYour infrastructure, or Onyx CloudAny — including local Ollama, vLLM, LiteLLMOpen source plus paid tiersCommunity $0 (MIT); cloud ~$16–20/user/month; enterprise customDays to self-host, months to harden
Dify / n8n / FlowiseSelf-host or vendor cloudAnyOpen source plus paid cloudFree self-hosted; paid cloud tiersFast to prototype, slow to production

What the tables do not show

Three things matter more than any single row.

The suite trap is real, but not always a trap. Microsoft Copilot has the widest deployment in the world — around 150 million seats, with daily utilisation reported at 28–32%. That gap is not a Microsoft failure. It is what happens when an assistant can see Outlook, Teams and SharePoint but not Salesforce, Confluence, your ERP, or the custom database where the answer actually lives. If 90% of your knowledge genuinely is in one vendor’s stack, buy that vendor’s assistant and stop reading. If it is not, a suite-native assistant will be excellent at a fraction of the job.

Search is not action. The search-first generation is superb at finding things. Glean’s retrieval is the strongest in the category — and its reported 25% first-year churn tells you what happens when finding the answer is where the product stops. The question that separates a demo from a deployment: after the assistant finds the invoice discrepancy, who fixes it? Note how many cells in the approval-gate column read “varies by agent” — that is the difference between a governed action and a hopeful one.

Consumption pricing is a governance question, not a finance one. Copilot Studio bills credits per feature, so a single agent response can burn several depending on how much retrieval it does. Agentforce charges per conversation or per credit. Gemini’s agent platform bills compute by vCPU-hour, memory by GiB-hour, storage by GiB-month, with billing switches flipping through 2026. A mid-market Agentforce deployment can run $6,650–18,800 per month once Data Cloud is included. None of that is unreasonable — but you cannot forecast it, and an unforecastable line item is the one finance kills first.

Where we fit

We are in the sovereign bucket, deliberately.

Our platform. Model-agnostic by design — pick the model per workspace and swap it when something better ships, with no retraining and no re-ingestion. Every write passes a pending → approved → executed gate, and the agent signs into each tool as the individual user rather than a shared service account, so an audit can name a person instead of a robot. Workspaces are isolated end to end: the finance agent cannot see HR data. Deployment runs in our cloud or on yours via Docker Compose, Kubernetes or AWS EKS, with configurable retention and residency. Cost is tracked per workspace, per provider, per user, with budget caps — a flat subscription rather than a credit meter. Answers render as charts, tables and forms inside the conversation rather than walls of text. Pricing starts at €400/month rather than a six-figure floor, and we run the deployment ourselves — ingestion, connectors, onboarding — so production is a week away, not a quarter.

How to actually run the evaluation

Score each of the 15 criteria 0–3 for your context, weight the ones your compliance function cares about at double, and shortlist three. Then, before signing anything:

  • Pick one real workflow that crosses at least three systems. Not a demo. Something a named person does weekly and hates.
  • Run it end to end on each shortlisted platform with your own data. Retrieval quality on your documents is the only benchmark that predicts anything.
  • Force a write. Make the agent change something in a live system, then audit who did it. Many products fail here quietly.
  • Model the cost at 10x the pilot volume. Ask the vendor to do it in writing.
  • Ask what happens in month 13. Renewal price, data export, connector ownership.

The organisations that reach production are not the ones with the best model. They are the ones that picked a platform matching their permissions, their infrastructure and their budget — then deployed one workflow properly instead of ten badly.

If you want a second opinion on your shortlist, including an honest read on whether we belong in it, get in touch.


Sources

Market data and adoption

Microsoft

Google

OpenAI

Glean, Dust, Cohere

Salesforce, ServiceNow, IBM

Open source

Regulation and sovereignty