Interaction mode

Human Handover

The measure of an assistant is not how rarely it fetches a person. It is whether the person arrives knowing what has already been said — and whether the customer had to fight to get them.

Arabic & EnglishGrounded answersEscalation to a person

The measure of an assistant is not how rarely it fetches a person. It is whether the person arrives knowing what has already been said — and whether the customer had to fight to get them.

What this actually means

Handover is the part of a deployment that gets least attention and determines most of the satisfaction. Every assistant will reach the edge of what it can answer. What happens in the next thirty seconds is what the customer remembers.

A good handover has three properties. It is easy to trigger — asking for a person is understood as an instruction, not answered as a question. It is fast — a queue that takes ten minutes is a rejection with extra steps. And it is complete — the agent who picks it up can see everything already said, so the customer never repeats themselves.

Most deployments fail on the third. The conversation is transferred, the agent gets a notification and an account number, and the first thing they say is "how can I help?" — at which point the customer explains their problem for the second time and concludes the assistant wasted their time.

What happens, step by step

  1. The trigger is recognisedEither the customer asks, or the assistant reaches a boundary it was configured not to cross — a complaint, a distressed tone, anything involving money moving, or simply a question outside its material.
  2. The customer is told plainlyNo pretending to be a person, no stalling. "I am connecting you to a colleague" is what is said, because the alternative is discovered and resented.
  3. The conversation joins a queueRouted by topic, language and agent skill. An Arabic conversation should reach someone who works in Arabic, which sounds obvious and is frequently not configured.
  4. An agent picks it upWith the full transcript visible — what was asked, what the assistant answered, what was retrieved, what the customer already provided.
  5. The agent continues, they do not restartThe first message from the person acknowledges what has already happened. This single detail separates a good handover from a bad one.
  6. The conversation is closed properlyResolved, recorded, and available for review, with the customer offered the chance to rate it.

Where it earns its place

When the customer asks

The simplest and most important trigger. Asking for a person should always work, immediately, without being talked out of it.

Complaints

A complaint is not a question to be answered. It is an escalation by definition, and attempting to resolve it automatically compounds it.

Distress

Where someone is upset, urgency outweighs efficiency. This should be tuned to escalate early rather than late.

Money and entitlement

Anything that moves money, changes an entitlement or creates a commitment belongs with a person, or at minimum passes through one.

Outside the material

Where your documents are silent, the honest response is a person rather than an improvised answer.

Repeated failure

If the assistant has failed to help twice, the third attempt will not go better. Escalating on repetition is one of the highest-value rules to configure.

Why two languages makes this harder

Language routing is where bilingual handover quietly fails. A conversation conducted in Arabic that reaches an English-speaking agent has achieved nothing — the customer now explains their problem again, in their second language, to someone who cannot read the transcript they are looking at. Queue routing has to carry language as a first-class attribute rather than as an afterthought.

The transcript itself has to be readable by the agent picking it up. If the conversation was in Arabic and the agent works in Arabic, that is straightforward. Where an organisation staffs mixed teams, the practical answer is routing by language rather than translating transcripts, because a translated transcript loses exactly the nuance that made the customer upset.

Working hours differ too. Where an organisation serves customers across the Gulf and staffs agents in one location, the honest configuration tells an out-of-hours customer when someone will be available rather than queueing them indefinitely into silence.

What good looks like

Time to a personMeasured from the request, not from when an agent happened to notice. Anything over a couple of minutes is experienced as being ignored.
Context completenessWhether the agent sees the full conversation. The test is simple: does the agent have to ask the customer to repeat anything?
Language routingWhether an Arabic conversation reaches an Arabic-speaking agent. Frequently unconfigured, and immediately obvious to the customer when it is not.
Escalation availabilityWhether asking for a person always works, or whether the assistant tries to deflect first. Deflection converts mild frustration into complaints.
Out-of-hours behaviourWhat happens when no agent is available. Silence is the worst answer; a clear statement of when someone will be there is the minimum.
Post-handover recordWhether the whole conversation — assistant and human — is retained as one record for review, or split across two systems.

The failure mode worth asking about

The failure that matters most is the silent queue. A customer asks for a person, the assistant says it is connecting them, and nothing happens — because presence was misconfigured, or every agent was marked away, or the queue had no fallback. The customer waits, then leaves, and is told by the system that nobody picked up. From their side it is indistinguishable from being ignored deliberately.

The second is the restart. The agent joins and asks what the problem is. Everything the assistant accomplished is discarded, the customer repeats themselves, and the reasonable conclusion is that the assistant was a delay rather than a help.

The third is the assistant pretending. Systems that avoid admitting they are automated, or that stall with filler while nothing happens behind the scenes, damage trust disproportionately when the truth becomes obvious — and it always becomes obvious.

What to ask any vendor about this

Ask to see the agent view

What the person picking up actually sees. If it is a notification and an account number rather than the conversation, handover is decorative.

Ask about language routing

Whether the queue understands language, and what happens when no agent speaks the customer's.

Ask what happens out of hours

Silent queueing is the most common and most damaging default.

Ask how escalation is triggered

Whether asking for a person always works immediately, or whether the assistant attempts to deflect first.

Ask about the record

Whether the assistant and human halves of a conversation stay as one reviewable record.

Ask about agent presence

How the system knows an agent is genuinely available, and what happens when that information is wrong.

How this fits industry practice

The direction of travel across the industry is consistent and worth understanding before you evaluate anyone. Assistants have moved from scripted decision trees, to retrieval over a knowledge base, to systems that can take actions against connected systems. Each step added capability and added a category of risk, and the risks did not replace each other — they accumulated.

Scripted systems were safe and useless: they could only answer what someone had anticipated, so they frustrated everyone and were quietly abandoned. Retrieval systems answer far more but introduce the possibility of answering wrongly with confidence. Action-taking systems answer and do, which raises the stakes again — an incorrect answer is embarrassing, an incorrect action is a liability.

The mature position, and the one worth insisting on, is that capability should be earned rather than assumed. Ground answers in material the organisation controls. Check the finished output rather than trusting an instruction given beforehand. Keep actions narrow, reversible and reviewable. Escalate rather than improvise. None of that is exotic; it is simply what separates a deployment that survives its first year from one that gets switched off after an incident.

The specific thing this market adds is language. Nearly every published benchmark, best practice and vendor claim was developed against English. A capability that performs well in English and adequately in Arabic is not bilingual — it is an English product with Arabic support, and the difference shows up precisely where it costs most: in the questions your customers actually ask, in the words they actually use.

Measuring whether it is actually working

Most assistant deployments are measured on the wrong number. Containment — the share of conversations resolved without a person — is easy to report and easy to improve for the wrong reasons. An assistant that makes escalation difficult will show excellent containment and a deteriorating relationship with customers, and the second effect takes months to surface while the first appears on a dashboard immediately.

The numbers worth watching are the ones that expose deferred problems. Repeat contact within forty-eight hours tells you whether a resolved conversation actually resolved anything. Abandonment mid-conversation tells you where people gave up. The share of escalations in which the customer had to repeat information already given tells you whether handover is working or merely happening. None of these flatter the system, which is precisely why they are useful.

Two more are worth instrumenting from the start. The first is grounding — what share of answers can be traced to a source document. In a regulated sector this is the number an auditor will eventually ask for, and retrofitting the ability to answer it is painful. The second is language accuracy: whether the reply came back in the language of the question. It sounds trivial and it is the single most common complaint about bilingual assistants in this market.

Set a baseline before launch. Without one, every improvement is an assertion. Take a month of existing contact, categorise it, count it, and time it. That baseline is what makes the business case defensible six months later when somebody senior asks what the assistant actually changed.

Misconceptions worth clearing up early

"It will replace the contact centre"

It will not, and deployments sold on that promise tend to fail. It removes the repetitive share of contact and leaves the judgement, the complaints and the complex cases — which is where trained people were always most valuable.

"More training data means better answers"

For a grounded assistant, answer quality depends on the quality and organisation of your own approved material, not on volume. A hundred well-maintained documents outperform a thousand contradictory ones.

"It needs to know everything on day one"

The opposite. Narrow deployments that answer a handful of questions extremely well build trust; broad ones that answer everything adequately lose it at the first confident mistake.

"Arabic support means Arabic works"

Nearly every vendor claims Arabic. The question is whether it handles the Arabic your customers actually type and speak — dialect, mixed Arabic-English sentences, and right-to-left interface behaviour — or only the formal written form.

"Accuracy is a single number"

Accuracy against what, asked by whom, in which language? A figure quoted without a described test set and a stated language is marketing, not measurement.

"We can add governance later"

Consent, retention and audit are far cheaper designed in than retrofitted. The most common reason a working pilot never reaches production is that it cannot pass legal review.

The numbers that matter

0
beyond which waiting is experienced as being ignored
0
of the conversation visible to the agent who picks it up
0
times a customer should have to repeat what they already said
0
record covering both the assistant and the human half of the conversation

Getting this right in practice

Configure the escalation rules before the answering rules. It is tempting to spend the deployment tuning what the assistant can answer and treat handover as plumbing. The opposite allocation produces better outcomes: an assistant that answers less but escalates cleanly is experienced as helpful, while one that answers more but escalates badly is experienced as an obstacle.

Rehearse the unhappy paths deliberately. No agent available. Every agent marked away. An Arabic conversation with only English agents online. A customer who asks for a person immediately without asking anything else. These are the situations that define perception, and they are the ones most pilots never test.

Watch the repeat-contact number rather than containment. A deployment where containment rises and repeat contact rises with it is not resolving anything — it is deferring conversations that will come back tomorrow, usually angrier.

Common questions

Yes, and it should be immediate. Asking for a person is treated as an instruction, not as a question to be answered. Systems that deflect first convert mild frustration into complaints.

Yes — the full transcript, what was retrieved, and anything the customer already provided. The test of a handover is whether the agent has to ask the customer to repeat themselves.

The customer is told plainly when someone will be available, rather than queued into silence. What the system does when nobody is there matters more than what it does when everyone is.

Where the queue is configured for it, yes. Language is a routing attribute rather than an afterthought — reaching an agent who cannot read the transcript achieves nothing.

No. It says it is connecting you to a colleague. Systems that obscure this damage trust disproportionately when it becomes obvious, and it always becomes obvious.

Yes, where you enable it. An agent watching a conversation that is going badly can step in before the customer has to ask.

See it on your own content

A working assistant on your own material, in both languages.