Playbooks
How to Evaluate an AI Customer Service Vendor
Turn a polished demo into a defensible buying decision by defining the workflow, normalizing claims, inspecting evidence, testing failures, and agreeing on exit terms.

An AI customer service vendor evaluation should answer one question: can this purchased configuration handle your defined work, boundaries, failures, and evidence requirements under conditions that resemble deployment? Start with scenarios and acceptance criteria, then evaluate claims, architecture, governance, integrations, operations, commercials, and exit. A demo is useful discovery, but it is not acceptance evidence.
Define the job before the product category
- Audience, channel, language, hours, geography, and accessibility needs.
- Questions to answer and sources allowed to determine each answer.
- Fields to collect and the minimum evidence required before an action.
- Actions permitted, approval limits, and destination systems.
- Safety, privacy, legal, regulated, emotional, and technical handoff boundaries.
- Expected volume shape, concurrency, seasonality, and recovery behavior.
- Customer disclosure, consent, recording, retention, and deletion requirements.
- Success, failure, and abstention outcomes for each scenario.
If the team has not decided whether it needs a scripted chatbot, a conversational agent, or a complete front-desk operating layer, use the category guides first. The vendor checklist assumes the buyer can describe the reader or caller job and the systems involved. chatbot versus AI agent guide · what an AI front desk includes
Convert claims into evidence requests
| Claim area | Request | Acceptance evidence | Common ambiguity |
|---|---|---|---|
| Availability and scale | Purchased scope, architecture, limits, incident history, recovery commitments | Buyer-run load and dependency tests plus contract terms | A platform maximum presented as every customer's entitlement |
| Knowledge accuracy | Source controls, precedence, freshness, permissions, citations, evaluation method | Representative test set with expected source and fallback | A curated demo mistaken for production quality |
| Integrations and actions | Implementation label, permissions, fields, retries, duplicate controls, audit trail | Sandbox and failure test verified in the destination system | A logo or API access mistaken for an implemented workflow |
| Human handoff | Triggers, destination, context, consent, availability, timeout and recovery | Live acceptance tests with unavailable owners and missing data | A transfer button presented as completed ownership |
| Security and privacy | Data flow, subprocessors, access, encryption, retention, deletion, incident process | Reviewed artifacts, configuration, contract, and technical validation | A certification used as a substitute for use-case review |
| Performance and outcomes | Metric definition, denominator, test population, baseline, exclusions, date | Reproducible method and buyer data where appropriate | Unscoped percentages or best-case examples |
The FTC advises businesses to substantiate claims about what AI products can do, whether they outperform alternatives, and the risks they create. Buyers should apply the same discipline in reverse: ask what evidence supports the claim and whether it covers the configuration, population, channel, and operating condition being purchased. FTC: Keep your AI claims in check
Inspect governance and human ownership
NIST's AI RMF organizes work around governing, mapping, measuring, and managing risk. Use those functions as diligence prompts: who is accountable, which impacts have been mapped, how behavior is measured, and how problems are managed after launch. Ask for named owners on both sides, change control, release notes, incident routes, evaluation cadence, and the ability to pause or constrain an action class. NIST AI Risk Management Framework
- Who approves knowledge, prompts, policies, actions, and escalation boundaries?
- Which configuration changes can the vendor make without buyer approval?
- Can the buyer see versions, audit events, exceptions, and failed actions?
- How are model or dependency changes evaluated before rollout?
- What can frontline staff override, and how is that decision recorded?
- Who receives a critical incident, and what happens if that owner is unavailable?
Test accessibility in the actual experience
A conformance statement or accessible website does not automatically prove that every conversational, voice, document, authentication, handoff, or embedded experience meets the buyer's needs. WCAG 2.2 provides web accessibility criteria, while applicable obligations and test scope depend on the implemented service. Include people using relevant assistive technologies and alternate input or communication methods in the evaluation. W3C Web Content Accessibility Guidelines 2.2
Run an acceptance pilot, not a showcase
- Baseline the current workflow using disclosed definitions and a representative sample.
- Lock the pilot configuration and scenario set before scoring.
- Include ordinary, ambiguous, adversarial, emotional, boundary, accessibility, and unsupported questions.
- Disable a destination system and verify truthful recovery without false confirmation.
- Test duplicate requests, retries, cancellation, permission changes, and unavailable human owners.
- Compare source conversation, generated summary, destination record, action state, and customer message.
- Review failures by class; do not hide them inside one average score.
- Define launch, remediation, and stop conditions before viewing results.
The knowledge-base audit provides test questions for retrieval, contradiction, scope, and permissions. The human-handoff guide helps specify destination, context, consent, timeout, and recovery evidence. knowledge base audit checklist · AI agent to human handoff guide
Normalize the commercial and exit model
- One-time discovery, configuration, integration, migration, security, training, and launch work.
- Recurring platform, seat, channel, number, model, storage, support, or environment fees.
- Usage units, included amounts, rounding, minimums, concurrency, overages, and third-party pass-throughs.
- Support hours, response commitments, named escalation, maintenance, and change notices.
- Data ownership, export format, deletion verification, logs, knowledge portability, and transition assistance.
- Renewal, price change, suspension, termination, post-termination access, and dependency replacement.
Compare a realistic usage range and a failure-heavy month, not one optimistic total. Keep unverified pricing, volume, integration, and performance claims as verification tasks until the vendor supplies scoped artifacts. Browse the fundamentals hub for related implementation and governance guides. customer service fundamentals guides
Bring your own calls, edge cases, destination systems, reviewers, and acceptance thresholds to the evaluation.
Plan a workflow evaluationQuick answers
Frequently asked
What should I ask an AI customer service vendor?
Ask what configuration and use cases are included, which evidence supports capability and outcome claims, how knowledge and actions are governed, how failures and handoffs recover, what data and accessibility controls apply, and how pricing and exit work.
How long should an AI customer service pilot run?
There is no universal duration. Run long enough to cover representative volume shapes, channels, reviewers, edge cases, dependency failures, and operational cycles, using launch and stop conditions defined before the results.
Does an integration logo prove the workflow works?
No. Determine whether it is a native adapter, API integration, webhook interoperability, configurable workflow, marketplace application, or planned integration, then test the exact permissions, fields, actions, retries, and failure behavior you need.
Evaluate the purchased workflow, not the demo category
Define scenarios, evidence, reviewers, acceptance tests, commercial terms, and exit before selecting a vendor.








