Skip to main content
Waaru
WhatsApp business automation

How to measure WhatsApp automation beyond message delivery

Tanay PatelAI and full-stack engineer
The short answer

Measure WhatsApp automation by defining the task, counting eligible attempts, and checking how many reach a verified outcome. Track failures and human handoff alongside completion. A delivered message or clicked link shows activity; the system that owns an order, appointment or enquiry determines whether the business task actually finished.

Choose a few measures your team can verify and act on. This guide proposes a measurement worksheet, not a list of analytics features guaranteed to exist in your platform. You may need to combine workflow records, inbox timestamps and records from the business system.

Which outcome should you measure first?

Start with the reason for the workflow. For enquiry intake, success may mean the responsible person received the required details. For appointment scheduling, it may mean a confirmed appointment exists. Those are different outcomes and need different evidence.

Name the owning record before choosing the metric. A WhatsApp reply cannot prove a payment settled. Use the payment provider’s current record for that result. Likewise, use the reservation or appointment system to verify a confirmed booking.

How do you define a completion rate?

Use a consistent group of eligible task attempts and a fixed observation period. Define which attempts belong in the denominator before looking at results.

Task completion rate = verified completed tasks ÷ eligible task attempts. Report the underlying counts with the percentage. If an illustrative test has 20 eligible enquiries and 15 reach the defined completed-intake state, the rate is 75%. This is a calculation example, not a Waaru benchmark.

Decide how to treat a customer changing the request, returning later or abandoning the conversation. Do not count every message as a fresh task or remove failures from the denominator simply because they make the result look worse.

Which measures are useful alongside completion?

Use this worksheet to define each measure, its source and the action your team will take when it changes.

  • Completed intake: enquiries with the required information received by an owner, divided by eligible started enquiries.
  • Unowned work: unresolved requests with no responsible person or team at the time of review. Define the review time consistently.
  • Handoff acceptance time: elapsed time from requesting transfer to acceptance by a person. Record unaccepted requests separately rather than excluding them.
  • Verified action completion: actions confirmed by the owning business system, divided by eligible attempted actions.
  • Repeat contact: customers returning about the same unresolved task within your chosen period. Link by the relevant task record rather than message count.
  • Failure and opt-out counts: show these beside useful outcomes so increased activity does not hide a worse customer experience.

Why is handoff rate not automatically good or bad?

A handoff can be the intended successful outcome of an enquiry workflow. It can also reveal that the bot lacks a needed answer. Interpret the reason and the result before treating a lower rate as an improvement.

A useful review asks: was the transfer necessary, did someone accept it, and did the customer have to repeat information? The human-handoff guide helps you define those conditions. Measure the waiting time and unresolved queue as well as the count of transfers.

How do you avoid counting the same task twice?

Choose the record appropriate to the task: an enquiry reference, order reference or appointment request. Keep repeated messages and retries attached to that task when they are part of the same attempt. A new, distinct request may deserve a new record even if it comes from the same customer.

Use a written rule for reopened work. For example, if a customer returns because the original issue remains unresolved, review it as repeat contact rather than silently treating the first conversation as a complete resolution. Limit access to underlying customer records; publish aggregated results when individual details are not needed.

What should different industries verify?

For hotels, distinguish a collected stay enquiry, a quoted option and a confirmed reservation. The reservations record proves the final booking.

For ecommerce, an opened payment link differs from a provider-confirmed payment. A resolved delivery question also differs from a delivered marketing message.

For clinics, measure administrative appointment requests and confirmed scheduling. Do not use this worksheet to infer medical outcomes.

For property teams, distinguish a requested viewing from a confirmed viewing and from an eventual sale. The appropriate business record establishes each stage.

Can you attribute an improvement to automation?

A before-and-after change alone does not establish causation. Traffic source, staffing, seasonality and campaign mix may have changed too. Record those differences and compare similar groups where possible. Use your own baseline rather than an unsupported industry conversion promise.

Set the measurement definition before changing the workflow. Keep the observation window consistent and show unresolved work separately. If a result looks unexpectedly strong, inspect the underlying records for duplicates, missing failures or changed eligibility.

How do measurement and testing work together?

Before launch, the AI-agent test sheet checks whether specific scenarios produce the expected outcome. After launch, measurement helps identify which failures deserve another test case.

Anthropic’s evaluation guide treats final state as distinct from the transcript. That distinction also helps here: verify the task, not just the agent’s account of it.

If you cannot identify a clear outcome, return to the first-workflow worksheet. A narrower task is easier for customers to understand and for your team to measure.

FAQ

Frequently asked questions.

Is message delivery a conversion?

Delivery shows that a message reached its delivery state. It does not establish a completed purchase, appointment or enquiry. Define conversion using the record that owns the business outcome.

Should I try to eliminate human handoff?

No. Human handoff may be the correct outcome for exceptions or decisions requiring authority. Review whether it was needed, accepted and resolved with useful context.

Does Waaru provide every metric in this worksheet?

This is a measurement framework, not a promise of a particular report. Check the records available in your workflow, inbox and connected systems, and define a manual review where needed.

Research notes

Evidence checked for this article.

  • An agent’s reported result should be checked against the actual task outcome.

    Anthropic, Demystifying evals for AI agents · Accessed 6 September 2026

About the author

Tanay Patel

AI and full-stack engineer

Tanay builds production AI systems, agent workflows, and the full-stack software around them. His work focuses on reliable state, recovery, and human handoff.

Read Tanay's engineering notes

Put this into practice on your own number.

The platform is free. AI features are paid separately.