Customer Operations

How to Test an AI Website Chatbot Before You Buy It

Test an AI website chatbot with your own knowledge, difficult questions, handoffs and evidence before you trust a polished vendor demo.

By DripTell EditorialPublished August 22, 2026Reading time 6 min read
Read the article
A shoe shop manager checks a running shoe beside a laptop while the Context Keeper reviews the proof

The right AI website chatbot is not the one with the longest feature list. It is the one that can answer your real questions from approved information, admit when it cannot, hand the conversation to a person with context, and run on your site without creating a new operational mess.

That is difficult to judge in a polished demo. A vendor controls the questions, the knowledge and the happy path. Before you buy, run a small proof test with your own pages, policies and awkward customer requests. Twenty well-chosen conversations will teach you more than an hour of slides.

Start with one real job

Do not begin with a broad goal such as reducing support volume. Choose one job the chatbot should complete. It might answer pre-sales product questions, collect a service request, qualify a lead or help a customer find the right support article.

Write down where that job begins and ends. A product-question bot may explain available sizes from approved catalogue data, but it should not promise stock that it cannot verify. A support bot may collect an order number, but it should not decide a refund outside a documented rule.

This boundary matters because fluent answers can still be wrong. The NIST Generative AI Profile describes confidently stated false output as confabulation and recommends monitoring it, especially when people may act on the result.

Build a twenty conversation proof set

Use questions your visitors actually ask. Remove personal details, then include a deliberate mix.

  • Six ordinary questions with clear answers in your approved content
  • Four questions asked with spelling mistakes, vague wording or missing details
  • Three questions whose answers change, such as availability, delivery timing or account status
  • Three questions that should go to a person because they involve judgment or an exception
  • Two requests for information the bot must not reveal
  • Two attempts to make the bot ignore its rules or follow instructions copied from a page

Write the expected outcome beside each test before the vendor runs it. The outcome may be an answer, a clarifying question, a safe refusal or a handoff. If you decide what counts as success afterward, almost every demo can be made to look good.

Check the evidence behind each answer

Ask the vendor to show which approved source supported an answer. Then change one source and test again. You are checking the full path from content update to customer reply, not only whether the model can sound certain.

Separate stable knowledge from live facts. Opening hours and return rules can come from maintained documents. Current stock, appointment availability and order status need a verified system lookup. If the system is unavailable, the chatbot should say what it cannot confirm rather than turn an old value into a promise.

Test conflicting documents too. A forgotten policy page is more revealing than a clean FAQ. You need to know which source wins, who owns the correction and how quickly the new version reaches live conversations.

Retrieval alone is not a complete safety control. The OWASP guidance for large language model applications notes that retrieval does not fully prevent prompt injection. It recommends constraining the model's role, validating outputs and filtering sensitive categories. Ask how the product enforces those controls, not whether a prompt merely tells the bot to behave.

Test the human handoff

A chatbot has not succeeded if it produces a neat summary and leaves the customer waiting in an unowned queue. Trigger a handoff during business hours and watch what happens.

The teammate should receive the transcript, customer details, detected need, attempted answer and reason for escalation. The customer should not have to repeat the same story. Confirm that automated replies pause when a person takes over and that ownership is visible.

Open the transferred case in the team's shared inbox and ask the new owner to continue it. This proves the context is usable after the handoff, not merely present in a summary.

The GOV.UK guidance on chatbots and webchat recommends making it clear whether a user is dealing with automation or a human, and says the tool should complement rather than replace other ways to get help. Test the escape route as carefully as the automated answer.

Check the website experience

Open the widget on a phone, a slow connection and a small screen. Navigate with a keyboard. Zoom the page. Check whether the launcher covers checkout controls, cookie choices or important text. Measure whether loading the chatbot delays the page before anyone opens it.

Also inspect what data is collected before a visitor understands why. Ask where transcripts are stored, who can read them, how long they are kept and how a deletion request is handled. A buying decision should include accessibility, privacy and page performance, not leave them for implementation week.

Score the work after launch

Avoid a single headline such as automation rate. A bot can raise that number by blocking access to people or giving weak answers that customers abandon.

Track a small set of operating measures instead. Review answer correctness, unsupported-answer rate, successful handoffs, repeated questions after handoff, unresolved exits and the time needed to update knowledge. Sample real conversations every week at first. Keep the failed questions because they reveal missing content, broken lookups and unsafe scope.

With DripTell Web Chat, website conversations can use approved business knowledge and move into the same customer record, ownership model and queue as other channels. The useful question is still the same for DripTell or any other platform: can your team prove the answer, own the exception and continue the conversation without making the visitor start again?

Choose the smallest safe scope

Buy for the job you can verify today. A narrow chatbot that answers ten valuable questions well and hands off the rest cleanly is more useful than a broad agent nobody can audit.

Ask the vendor to rerun your failed tests after corrections. Record the result, the owner and the effort required. If the proof needs constant vendor intervention, unclear workarounds or promises about future features, treat that as part of the product you are buying.

Frequently Asked Questions

How many chatbot tests are enough before buying

Twenty varied conversations are enough for a first proof, provided they include normal questions, ambiguity, live facts, restricted topics and handoffs. Add more tests for regulated or high-risk work.

Should a website chatbot answer every question

No. It should answer within an approved scope, ask for missing details when useful, and refuse or hand off when evidence or authority is missing.

What is the most important chatbot metric

There is no single best metric. Start with answer correctness and successful resolution, then review unsupported answers, handoff quality and unresolved exits together.

How do I compare two chatbot vendors fairly

Give both vendors the same sources, twenty questions, expected outcomes and time limit. Score the visible result and the operational effort needed to correct failures.

DT

DripTell Editorial

Practical guidance reviewed by the DripTell product and customer workflow team.

See how DripTell checks product claims, uses primary sources and handles corrections.

Editorial and source policy