Skip to content
A sequence of rounded glass platforms broadens into a wider foundation in warm light, representing a measured pilot growing into a rollout.
Pilot guide

From the first call to a rollout decision

Run an AI calling pilot around one defined workflow, a fair comparison with today's process, and a decision the team agrees to make. The evidence should connect what happened in the conversation with the completed work, the customer's experience, and the effort needed to deliver it.

By ConnectX10 min read

Choose a workflow with an observable finish

A delivery company wants to improve the calls that arrive when customers need a different delivery time. The request is familiar to its operations team. A useful result is visible in the scheduling system. Some changes can be completed immediately; others need a supervisor's decision. That makes it a practical candidate for a pilot.

We will use this fictional operation throughout the guide. The question for its manager is whether ConnectX can handle that work consistently enough to take on more of it. Our handoff article follows the individual customer conversation. Here, we follow the decisions the business makes around those calls.

Choose a starting workflow whose outcome, available actions, and exceptions the team can explain. Delivery changes, service appointments, and requested customer callbacks can all provide a clear starting point when the relevant systems and operating rules are ready. The right choice depends on the business's actual workload and the access needed to finish the task.

Volume alone is insufficient. A busy queue with unclear ownership or no reliable way to verify completion will make the results difficult to interpret. Start where the team can connect a request to its outcome and identify who handles the work that remains.

Write down the decision before the first call

The pilot brief should state what the company wants to improve and what would justify a wider rollout. In our example, the objective might be to complete eligible delivery changes with less staff effort while preserving correct bookings and a good customer experience. The team needs to agree what an acceptable result means against its own baseline.

Define the initial customer group, call routing, supported hours, and workflow limits. Name the person who can approve an increase in scope. Agree on a review date, the evidence they will receive, and the conditions that would require an earlier intervention.

Use the worksheet below to prepare that discussion. Fill in the thresholds with the people responsible for the operation; a universal completion target cannot account for every company's call mix or consequences of error.

Pilot planning worksheet

Agree on the work and the decision

The first workflow
Which requests enter the pilot? Define the customer group, routing, hours, allowed actions, and completion event.
The comparison
Which equivalent requests form the baseline? Record the measurement period, exclusions, and follow-up window.
The owners
Who owns the operation, answers escalations, fixes issues, and approves the next scope?
The acceptance criteria
What outcome, customer experience, and staff effort would justify expansion? Which failures block it?
The review
When will the team decide? List the evidence required, the starting volume, and the route for unfinished requests.

The resulting brief should be short enough for the customer-service lead, supervisor, and implementation team to use during the same conversation. Keep it available as the reference for changes made during the pilot.

Make the comparison fair

Before live calls begin, establish how equivalent requests are handled today. For the delivery team, that means more than an average call length. It means the number of requests completed, whether customers called again about an unresolved change, how much work supervisors performed, and whether the booking was corrected later.

Where routing permits, compare the pilot with a concurrent group of similar requests handled through the existing process. Assign customers consistently so repeated contact about the same request stays in the same group. If a concurrent comparison is impractical, use a clearly described earlier period and record what differs: staffing, operating hours, route availability, demand, or customer mix. A before-and-after difference alone cannot establish that the agent caused the change.

Define the unit being measured. One delivery-change request may create two calls and a WhatsApp conversation. Count it once when measuring completed requests, and retain all three interactions when measuring contact volume and effort. Keep eligible failures and abandoned requests in the totals. Report out-of-scope contacts and routing exclusions separately, with their reasons.

For an outbound workflow, also separate attempted calls, answered calls, and completed requests. An unanswered attempt is part of outreach cost and coverage even though there was no conversation to evaluate.

Show the number of observations behind every rate. A small sample or an unrepresented customer group calls for more evidence. Agree on a follow-up window so repeat contacts and later corrections have time to appear before comparing results.

Prepare the people who will carry the work

An agent needs somewhere useful to turn when a request requires human judgment. In the delivery pilot, the operations lead defines which changes need approval; the supervisor knows when they are expected to answer; and the team agrees what happens when that person is unavailable. Test that fallback before customers depend on it.

ConnectX Soft Forwarding keeps the customer connected while the agent privately briefs a supervisor. The supervisor can guide the agent or take over the call. The pilot should measure the work in both paths, including the time needed to resolve the request after the transfer. A well-handled escalation can be the correct outcome.

Confirm that the pilot can use the relevant records and perform the agreed actions, with appropriate access and customer-data handling. Make the return to the existing service route practical: name the person who can pause routing and the team that receives unfinished requests. Discuss the organization's requirements through the Trust conversation.

If the workflow continues into WhatsApp, include that continuation in its scope. Unified Memory should carry relevant customer context, while the latest booking still needs verification. Treat the channels as parts of the same task when reviewing the result.

Keep a dated record of changes to the agent, business rules, and connected systems. ConnectX's Knowledge Loop scores learned solutions and makes high-scoring solutions usable automatically. During a pilot, review their use in comparable cases and in cases where an earlier exception should not apply. Record these changes so an improvement or regression can be investigated against the conditions that produced it.

Move into live calls deliberately

Begin with rehearsals that exercise the planned workflow, including corrections, interrupted calls, unavailable options, and failed system updates. People who understand the intended callers should listen to the audio and inspect the resulting action. Our Arabic voice evaluation guide explains how to examine language, conversation, and task performance together.

Sierra's enterprise voice guide similarly connects call readiness to observable business actions and realistic conversational conditions. Those principles inform preparation; the live pilot must then establish how the workflow performs in the customer's own operation.

Start with a defined portion of eligible calls and supervisor coverage. Review the first calls closely. As evidence accumulates, use representative samples across ordinary and difficult calls, languages, operating periods, and outcomes, alongside investigation of every identified serious failure. Automated scores can help locate calls for review; check them against human review and business records.

Schedule short operational reviews while the pilot runs. Separate configuration changes from the next measurement period and repeat the affected tests before increasing traffic. Keeping the scope stable long enough to observe its performance makes each decision easier to explain.

For planning, roughly four weeks of live measurement after setup and testing can be a useful starting point. This is an illustrative cadence, not a fixed ConnectX offer or a sufficient sample for every business. Call volume, variation, and the time needed to observe the outcome determine how much evidence is available at the review.

Read the outcome and the effort together

The delivery manager needs to see both the work completed and what it took to complete it. A request resolved with supervisor guidance belongs in the result, with the assistance visible. A conversation that ends pleasantly while the booking remains unchanged does not count as a completed delivery change.

Use a compact scorecard for the review. Compare the same definitions and observation windows across the pilot and the baseline.

Pilot decision scorecard

Five measures for the operating review

Verified completion
Verified completed requests divided by all eligible requests entering the workflow. Show completion with and without human assistance separately.
Customer effort
Repeat contact for the same unresolved request, repeated information, abandonment, and feedback where collected. Use a consistent observation window.
Escalation quality
Review whether help was requested when needed, the supervisor received useful context, and the request reached a recorded resolution.
Human work
Measure staff time across supervision, takeover, corrections, and follow-up. Keep pilot setup and training effort separate.
Cost per completion
Count the service, telephony, recurring review, and human work needed to deliver verified outcomes. Include the cost of failed attempts.

For cost per completed request, include the operating cost of all eligible pilot work, including failed attempts, human assistance, corrections, and follow-up. Divide by verified completed requests. If none completed, report the total spend and zero completions; a meaningful cost per completion cannot be calculated. Show one-time setup cost separately, and make recurring quality-review effort visible.

Keep the underlying cases available to explain the totals. An overall improvement may hide poorer results for a particular language, call type, or operating period. Review those groups before extending the service to them.

NIST's AI Risk Management Framework calls for evaluation in conditions relevant to deployment and continued measurement during operation. It also emphasizes documenting the methods and limits of the evidence. For a pilot team, that supports a review that can explain both the result and where confidence remains limited.

Decide what the evidence supports

Return to the decision agreed in the brief. The team should be able to show whether the workflow met its acceptance criteria, where human work remains, and which conditions have actually been tested. A reported average should never conceal an unresolved failure that the team agreed would block expansion.

The right next step depends on the evidence. Explore the four decisions below.

The rollout decisionIllustrative review paths

What does the evidence support?

Explore each decision and the work that follows

01Expand the scope

The agreed outcome and quality criteria are met across a representative sample. Human assistance, customer effort, and operating cost are understood. No blocking issue remains unresolved.

Next stepApprove one defined increase in scope. Keep the same measurement and review the new conditions.
02Improve the workflow

The evidence identifies a specific weakness. If the workflow remains acceptable at its current scope, agree a bounded correction and retain that scope while the change is evaluated.

Next stepName the issue and its owner, repeat the affected tests, and measure the revised workflow before expanding.
03Gather more evidence

The observed calls are too few, a relevant customer group is missing, or repeat-contact outcomes have not had time to appear. Record what remains uncertain.

Next stepAgree the missing evidence and a new review date. Keep the limits of the current findings visible.
04Pause and resolve

A blocking failure has occurred, such as changing the wrong record, claiming an action that did not complete, or failing a required handoff.

Next stepPause affected routing, hand unfinished requests to the agreed team, and correct and retest before resuming.
These are planning examples, not ConnectX pilot results. Apply the criteria and decision authority agreed for the actual operation.

When expansion is justified, change one part of the scope at a time. The delivery team might increase the volume of the same request type while keeping the existing hours and supervisor arrangement. A new workflow or customer group needs its own assumptions checked and its results tracked separately.

A completed pilot gives the business a working decision record: the task it evaluated, the conditions under which it worked, the effort involved, and the next scope it is ready to take on. Bring that starting workflow to a ConnectX voice demo, and use the conversation to define the evidence your team needs.

ConnectX

Product perspectives from the team building AI voice agents for business and clinical AI for care.

Examples are fictional; external results are attributed to their source.