Do not measure AI by the number of messages it answered
Ask how many conversations were resolved correctly, when a person was needed and whether wasted work fell without damaging the customer experience. A useful dashboard connects coverage, resolution, escalation, response time, follow-up, business outcomes and quality.
An AI agent answering 90% of conversations is not necessarily successful if half end with inadequate answers or late handoffs. A single attractive number can hide a poor operation.
Why is this harder than measuring a traditional chatbot?
Traditional chatbots often follow a fixed flow. Modern AI handles varied wording, questions and contexts and may reply, follow up, summarize or transfer a case.
Salesforce State of Service 2025, based on a global survey of 6,500 service professionals and decision-makers, reported teams estimating AI handled approximately 30% of cases, with an expectation of 50% by 2027. Saudi Arabia and the UAE were included; this is not Egyptian research or a benchmark for Egypt.
As AI takes a larger role, management needs to measure the quality of the division of work between AI and people.
1. Conversation coverage: does AI handle suitable cases?
Of incoming conversations appropriate for AI, how many involved it?
AI coverage = conversations involving AI ÷ conversations eligible for AI.
The eligible denominator matters. Sensitive complaints, special negotiations or decisions needing human authority may intentionally be excluded. Higher coverage is not an independent objective; appropriate coverage is.
2. Resolution rate: was the request actually resolved?
Definitions matter. Intercom's August 2026 reporting changes distinguish AI involvement, opportunities to answer, resolved conversations and automation across total volume.
Its official example uses 1,000 conversations, 250 resolved by AI and 500 where AI had an opportunity to answer: 50% resolution rate, but 25% automation rate across all conversations.
Resolution rate = AI-resolved conversations ÷ conversations AI had an opportunity to resolve.
Automation rate = AI-resolved conversations ÷ all conversations in the measurement scope.
Use both. A high resolution rate may represent limited overall value if AI participates in very few cases. Do not import vendor benchmarks as mandatory targets; sector, question mix, information quality and permissions change the outcome.
3. Escalation rate: a human handoff is not automatically failure
Salesforce's analysis of AI-agent use in the first half of 2025 reported handoffs increasing from 22% in Q1 to 32% in Q2 in the data examined. It described this as part of improved collaboration rather than proof of worsening AI performance.
A useful system knows when to stop: missing or uncertain information, a customer requesting a person, complaints, negotiation, exceptions, financial authority or a high-value opportunity.
Escalation rate = conversations transferred to people ÷ AI conversations.
Also track correct escalation rate = appropriate handoffs ÷ reviewed handoffs. Review whether transfer was needed, timely and accompanied by enough context.
4. Response time: distinguish a reply from a useful answer
Separate first response time, time to the first useful answer and time to resolution or the required next step. A fast acknowledgment should not be rewarded as if it resolved a request.
Zendesk's 2026 CX statistics compilation reports 72% wanting immediate service. This is global, not Egyptian. Speed has value when it preserves accuracy.
5. Follow-up coverage: are stalled opportunities remembered?
Value can disappear after an adequate first reply because no next step occurs.
Follow-up coverage = completed follow-ups ÷ eligible follow-ups.
Follow-up outcome rate = follow-ups producing a useful next step ÷ completed follow-ups.
A useful step might be a customer reply, appointment, requested information, completed order or confirmation of no interest that closes a previously unresolved opportunity. It is not always a sale.
6. Human work saved: what actually left the team's queue?
Do not treat every AI message as a minute of saved employee work. Measure tasks people no longer had to perform or completed faster because of context and summaries.
Salesforce State of Service 2025 reported representatives using AI spent 20% less time on routine cases, estimated at around four hours weekly for more complex work. That is a global study result, not a promise to Mr. AI customers.
For internal measurement: human hours saved = routine cases resolved without an employee × previous average handling time.
Hypothetically, if a business has 300 monthly routine cases, previously requiring six minutes each, and AI resolves 120 without intervention, that frees 720 minutes, or 12 hours monthly. It measures capacity, not complete ROI.
7. Cost per useful outcome
Message or point cost alone is not a complete business metric. Consider operating cost divided by resolved conversations, qualified opportunities or resulting bookings.
One hundred messages may produce one important conversation; another conversation may need only three. The unit of value is not always a message.
Global AI-versus-human cost estimates vary widely. McKinsey's 2026 analysis of AI decision economics describes substantial differences between computation costs and human work, but computation alone omits platform, setup, data, integrations, quality and human review costs. It is not a ready-made ROI calculator for your business.
8. Quality score: was the answer accurate and helpful?
Review a regular sample with a consistent rubric: correctness, grounding in company information, relevance, completeness, tone, policy compliance, required handoff and an appropriate next step.
Each element could receive 0, 1 or 2, with weekly tracking. Preserve important failures as regression tests so a later knowledge or instruction change does not reintroduce them. See the pre-launch testing guide.
9. Business outcomes
For sales, track qualified conversations, bookings, orders, next-step conversion, recovered opportunities and stated reasons for loss. For support, track resolution, repeat contacts, resolution time, escalations, collected satisfaction scores and the human queue.
Do not automatically attribute improvement to AI. Campaigns, pricing, seasons, lead quality and staffing can change in the same period. Compare similar periods or cohorts, or use a more controlled experiment where practical.
A suggested weekly dashboard
| Layer | Metric | Question |
|---|---|---|
| Volume | Incoming conversations | How much demand is there? |
| Coverage | AI coverage | Is AI handling suitable cases? |
| Resolution | Resolution and automation | What was solved without an employee? |
| Safety | Escalation and correct escalation | Does AI know when a person is needed? |
| Speed | First useful answer | Does the customer receive value quickly? |
| Follow-up | Follow-up coverage | Are stalled opportunities remembered? |
| Quality | Reviewed quality score | Is the answer accurate and useful? |
| Capacity | Estimated human hours saved | Which routine work disappeared? |
| Business | Qualified, booked or resolved | What changed in the outcome? |
| Economics | Cost per useful outcome | What does a useful result cost? |
You do not need twenty charts to start. You need stable definitions and a consistent weekly review.
Where does Mr. AI fit?
Published product capabilities include WhatsApp, Facebook Messenger and website chat, company-informed answers, automated follow-up and next steps, chat summaries, AI insight reports and communication analytics. Use these to understand what happens after a customer arrives, rather than simply counting replies.
Current published Starter pricing is EGP 1,100 monthly for 1,300 points. One inquiry uses one point; one follow-up uses one point; a conversation summary uses two; a general insight report uses four. Verify current pricing and use actual usage and outcomes rather than assuming every point produces equal commercial value.
Read the ROI guide to separate financial return from usage measures.
A final operational rule
Rising automation with falling quality is a problem. Falling escalation with more complaints is a problem. Faster replies without better next steps mean speed was not the only gap. Freed team capacity does not become business value until it is used productively.
Good measurement helps you discover where AI succeeds, where it fails and what to improve next. Try the customer journey or book a session to review your current gaps.