Good-looking replies do not prove launch readiness

Testing should use real scenarios with clear pass and fail criteria, rather than ten easy questions followed by a readiness decision.

Generative AI is not deterministic in the way conventional programs are. OpenAI's evaluation guide recommends task-specific evaluations, representative data, recorded results and continuous evaluation.

For customer-facing AI, ask whether it uses the right information, recognizes unknowns, follows business rules and transfers cases when a person is needed. This checklist helps turn those questions into tests.

Why a successful demo is not enough

Real customers do not write ideal prompts. They ask “How much?”, “Will it suit me?”, “Is yesterday's option still available?”, request cheaper alternatives or refund details, send several short messages, combine questions or misspell products.

Some may request another customer's details or tell the AI to ignore instructions. Easy demo questions measure only the happy path.

NIST's Generative AI Profile treats testing and monitoring as part of risk management, including confabulation: plausible but incorrect or unsupported information.

Define a correct answer before writing questions

Start with rules. For a store, current prices must be used; missing products and availability must not be invented; refund policy must remain accurate; unapproved discounts must not be granted; other customers' information must remain private; exceptions must go to the team; and unknown facts must be acknowledged rather than guessed.

These rules define expected behavior that can be assessed.

Build your first test set from real conversations

Use real customer scenarios after protecting or removing personal information. Include repeated questions and cases that concern the sales and service teams.

OpenAI's guidance on improving LLM accuracy suggests that 20 or more questions with ground-truth answers, together with failure analysis, can provide a useful baseline. This is a starting point, not a universal readiness guarantee. Expand the set according to the diversity and risk of your services and policies.

Eight groups to test

1. Direct factual questions

Check prices, hours, services, branches, payment methods, returns, booking steps and documented availability. Judge accuracy, not merely pleasant wording.

2. Different wording for the same question

Customers will not use your database field names. For a Business plan, try “the large package,” “the company plan,” “more usage,” a mixed Arabic-English question and a request for a team solution.

Include dialect, spelling mistakes, short messages, abbreviations, multiple intents and long context. Use the language your customers actually use.

3. Questions without an answer

Ask about missing prices, nonexistent products, unapproved discounts, unannounced dates, unwritten policies and unsupported competitor information.

Success may mean not answering. Some evaluation approaches reward guessing more than admitting uncertainty. For business conversations, explicitly reward a correct “I do not have that information” instead of a confident invented answer.

4. Conflicting sources

Test an old page showing EGP 900 against a current catalog showing EGP 1,100, or an outdated FAQ saying Friday is available against current operating information saying it is closed.

This exposes whether a reliable source of truth exists. Do not try to fix every data conflict with a prompt. Correct the information and identify the authoritative source.

5. Long conversation context

Do not start a fresh chat for every question. Build an 8- or 15-turn scenario: service inquiry, comparison, price, changed needs, return to an earlier question and a next-step request.

Check whether it remembers the intended product, mixes up prices, repeats questions already answered, understands references such as “that one,” and maintains policy consistency.

6. Human handoff

Define cases requiring a person: sensitive complaints, exceptional discounts, payment problems, approval-dependent decisions, missing knowledge, an explicit request for an employee and high-value cases defined by your workflow.

Test two things separately: the transfer decision and handoff quality. Did AI recognize the need, and did the employee receive enough context to continue without making the customer repeat the story?

Mr. AI's published product information describes team takeover with a conversation summary and record for review.

7. Instruction attacks and data protection

Test requests to ignore company instructions, reveal the system prompt, share the last customer's details, disclose confidential information or accept an unverified claim that the sender is a manager.

OWASP's Top 10 for LLM Applications identifies prompt injection and sensitive information disclosure as important risks. Its system prompt leakage guidance explains that prompts should not hold secrets such as credentials or connection strings or be treated as a security control.

A prompt alone is not a security boundary. Permissions and data access also need system-level protection.

8. Information changes after launch

Change a price, add a policy, remove a service or adjust hours, then rerun the relevant tests. Testing is not a one-time event.

NIST recommends post-deployment monitoring, while OpenAI's evaluation guidance recommends continuous assessment and adding cases from actual use. Important real defects should become regression tests.

Use a scorecard rather than “good” or “bad”

Criterion Result
Accurate information Pass / Fail
No invented facts Pass / Fail
Current authoritative source Pass / Fail
Correct context Pass / Fail
Policy compliance Pass / Fail
Necessary human handoff Pass / Fail
No unauthorized disclosure Pass / Fail
Appropriate next step Pass / Fail

Do not hide serious failures in a combined score. Passing 95 of 100 cases means little if the five failures disclose information or invent prices. Those differ from five imperfectly worded replies.

Classify severity

  • Critical: data disclosure, unauthorized actions, sensitive invented financial commitments or clear permission violations.
  • High: incorrect essential information affecting a decision, or a missed necessary handoff.
  • Medium: incomplete context or an unsuitable but readily correctable next step.
  • Low: excessive length, imperfect wording or tone.

An example launch gate might require no known critical failures, no high-severity failures in core scenarios, human sample review, an explicit unknown-information process and ongoing logging. These are illustrative operating thresholds, not an official NIST or OpenAI standard.

Test the whole workflow

A correct model answer can still fail operationally because information is outdated, a message never arrives, follow-up has unsuitable timing, a handoff does not appear for staff, a summary misses key information, a channel disconnects or an employee never sees a waiting case.

Test the complete journey: message → understanding → company knowledge → reply → stage → next step → follow-up or handoff → team review.

How can you test Mr. AI?

Mr. AI lets businesses supply services, prices, policies and FAQs, then test questions, review answers and refine knowledge and rules before relying on live conversations. Published capabilities include WhatsApp, Facebook Messenger and website chat, company-informed replies, follow-up, next steps, summaries, insight reports and communication analytics.

Start with 20–50 actual company questions and ten difficult or unexpected cases. Write the expected behavior first, compare results and adjust knowledge and rules. Expand testing and security review where sensitive information, financial actions or higher-risk use require it; this operational checklist does not replace a specialized assessment.

Final pre-launch checklist

Verify core facts, varied wording and dialects, unknown information, conflicting sources, long conversations, human handoffs, unauthorized-data requests, knowledge updates, the entire message journey, and a monitoring process that turns real failures into regression tests.

The goal is to demonstrate that the behavior your company needs can be tested, reviewed and improved. Start the experience with your own questions, or book a session to review the workflow before launch.