Interaction mode

Multiple Agents

One assistant that does everything is worse at every individual job than several that each do one. Separate agents let tone, knowledge and limits differ by function — which is how the organisation actually works.

Arabic & EnglishGrounded answersEscalation to a person

One assistant that does everything is worse at every individual job than several that each do one. Separate agents let tone, knowledge and limits differ by function — which is how the organisation actually works.

What this actually means

A multi-agent setup means you configure a distinct assistant per function: support, sales, HR, finance, technical. Each has its own name, its own voice, its own tone, its own knowledge base and its own limits on what it may say. Customers and staff reach the right one either by choosing, or by being routed automatically from what they asked.

The reason this matters is not organisational tidiness. It is that a single assistant forces you to average three things that should not be averaged: the material it answers from, the register it answers in, and the boundaries on what it may state. Averaging them produces something mediocre at each job and difficult to govern at all.

What happens, step by step

  1. You define the agentsOne per function, or per brand, or per language, or per business unit — whatever division matches how you actually operate.
  2. Each gets its own materialSupport answers from support content; finance from finance content. Retrieval quality falls as a corpus widens, so narrowing it is a direct quality improvement rather than an administrative convenience.
  3. Each gets its own characterName, avatar, voice, tone and formality. In Arabic this matters more than in English, because register carries social weight and getting it wrong reads as rudeness.
  4. Each gets its own limitsWhat it may state, what it must escalate, what it may never discuss. Set per agent, testable per agent, auditable per agent.
  5. Routing gets the person to the right oneEither by explicit choice or by understanding the question. A customer asking about a refund reaches the agent that knows about refunds.
  6. Handover works across agentsA conversation that starts with one and belongs with another moves across with its history, rather than starting again.

Where it earns its place

By department

The most common split. Support, sales, HR, finance, IT — each with different material and different rules about what may be said.

By brand

Groups running several brands need each to sound like itself. One assistant using one voice across brands undermines the reason the brands are separate.

By audience

Customer-facing and employee-facing assistants have almost nothing in common in tone, knowledge or permitted disclosure, and should not share a configuration.

By language

Sometimes the right split is Arabic and English personas with different registers, because expectations of formality genuinely differ between them.

By risk level

Separating agents that may only inform from agents that may act makes the boundary explicit and reviewable rather than buried in conditions.

By region or entity

Groups operating across jurisdictions need different disclosures and different limits, which is far cleaner as separate agents than as branching rules.

Why two languages makes this harder

Bilingual organisations often discover that the right unit of separation is not the department but the language. Arabic and English business registers are not equivalent. Arabic business communication carries higher baseline formality, more elaborate openings and closings, and stronger expectations about how seniority is acknowledged. An English persona translated directly into Arabic reads as curt, even when every word is correct.

The reverse is also true. An Arabic persona rendered literally into English reads as ornate and slightly archaic, which in a support context reads as insincere. Configuring the two separately — same facts, different register — is usually better than configuring one and translating it.

This is also where mixed-language conversation has to be handled deliberately. Saudi users routinely switch languages mid-sentence, and the agent needs a defined behaviour rather than an accidental one: follow the customer, or hold a language, but never flip unpredictably on a one-word reply.

What good looks like

Number of agentsWhether you can run as many as your structure needs without per-agent cost making the correct design unaffordable.
Separate knowledge per agentWhether each agent answers from its own material. Shared corpora measurably reduce answer quality for every agent using them.
Separate limits per agentWhether what an agent may say can be set and tested per agent rather than as conditions inside one large configuration.
Routing qualityWhether a customer who picks the wrong agent, or does not pick, still reaches the right one without starting over.
Cross-agent handoverWhether a conversation moves between agents with its history intact.
Per-agent analyticsWhether you can see how each is performing separately. Aggregate figures hide the one that is underperforming.

The failure mode worth asking about

The failure here is a routing loop: a customer asks something, is passed to another agent, which passes them back. It happens when routing is configured on keywords rather than intent, and it is experienced as being transferred around a switchboard — the exact frustration the assistant was meant to remove.

The second failure is inconsistency on shared facts. If support and sales answer differently about the same refund policy because they hold different documents, customers will notice and will quote one agent to the other. Facts that must be consistent should live in material shared across agents; only genuinely function-specific content should be separated.

The third is proliferation. Twelve agents, each maintained by someone different, drift apart within a year. The right number is the smallest that reflects real differences in tone, knowledge or risk — usually far fewer than an org chart would suggest.

What to ask any vendor about this

Ask how routing decides

Keyword matching produces loops. Understanding the question does not.

Ask about shared facts

How the same policy stays consistent across agents that hold different material.

Ask about per-agent limits

Whether what each may say is configurable and testable separately.

Ask about cross-agent handover

Whether history travels when a conversation moves.

Ask about per-agent reporting

Aggregate numbers hide the underperforming agent for months.

Ask what agents cost

If each additional agent carries meaningful cost, you will end up with the wrong architecture for commercial rather than design reasons.

How this fits industry practice

The direction of travel across the industry is consistent and worth understanding before you evaluate anyone. Assistants have moved from scripted decision trees, to retrieval over a knowledge base, to systems that can take actions against connected systems. Each step added capability and added a category of risk, and the risks did not replace each other — they accumulated.

Scripted systems were safe and useless: they could only answer what someone had anticipated, so they frustrated everyone and were quietly abandoned. Retrieval systems answer far more but introduce the possibility of answering wrongly with confidence. Action-taking systems answer and do, which raises the stakes again — an incorrect answer is embarrassing, an incorrect action is a liability.

The mature position, and the one worth insisting on, is that capability should be earned rather than assumed. Ground answers in material the organisation controls. Check the finished output rather than trusting an instruction given beforehand. Keep actions narrow, reversible and reviewable. Escalate rather than improvise. None of that is exotic; it is simply what separates a deployment that survives its first year from one that gets switched off after an incident.

The specific thing this market adds is language. Nearly every published benchmark, best practice and vendor claim was developed against English. A capability that performs well in English and adequately in Arabic is not bilingual — it is an English product with Arabic support, and the difference shows up precisely where it costs most: in the questions your customers actually ask, in the words they actually use.

Measuring whether it is actually working

Most assistant deployments are measured on the wrong number. Containment — the share of conversations resolved without a person — is easy to report and easy to improve for the wrong reasons. An assistant that makes escalation difficult will show excellent containment and a deteriorating relationship with customers, and the second effect takes months to surface while the first appears on a dashboard immediately.

The numbers worth watching are the ones that expose deferred problems. Repeat contact within forty-eight hours tells you whether a resolved conversation actually resolved anything. Abandonment mid-conversation tells you where people gave up. The share of escalations in which the customer had to repeat information already given tells you whether handover is working or merely happening. None of these flatter the system, which is precisely why they are useful.

Two more are worth instrumenting from the start. The first is grounding — what share of answers can be traced to a source document. In a regulated sector this is the number an auditor will eventually ask for, and retrofitting the ability to answer it is painful. The second is language accuracy: whether the reply came back in the language of the question. It sounds trivial and it is the single most common complaint about bilingual assistants in this market.

Set a baseline before launch. Without one, every improvement is an assertion. Take a month of existing contact, categorise it, count it, and time it. That baseline is what makes the business case defensible six months later when somebody senior asks what the assistant actually changed.

Misconceptions worth clearing up early

"It will replace the contact centre"

It will not, and deployments sold on that promise tend to fail. It removes the repetitive share of contact and leaves the judgement, the complaints and the complex cases — which is where trained people were always most valuable.

"More training data means better answers"

For a grounded assistant, answer quality depends on the quality and organisation of your own approved material, not on volume. A hundred well-maintained documents outperform a thousand contradictory ones.

"It needs to know everything on day one"

The opposite. Narrow deployments that answer a handful of questions extremely well build trust; broad ones that answer everything adequately lose it at the first confident mistake.

"Arabic support means Arabic works"

Nearly every vendor claims Arabic. The question is whether it handles the Arabic your customers actually type and speak — dialect, mixed Arabic-English sentences, and right-to-left interface behaviour — or only the formal written form.

"Accuracy is a single number"

Accuracy against what, asked by whom, in which language? A figure quoted without a described test set and a stated language is marketing, not measurement.

"We can add governance later"

Consent, retention and audit are far cheaper designed in than retrofitted. The most common reason a working pilot never reaches production is that it cannot pass legal review.

The numbers that matter

0
typical number of agents in a mid-sized deployment
0
languages per agent, each with its own register
0
shared limits — what each agent may say is set separately
0
conversation history, carried across every handover

Getting this right in practice

Start with two. Almost every organisation is tempted to model its entire structure on day one and almost every one regrets it. Two agents — typically a customer-facing one and an internal one — expose the real differences in tone and material quickly, and splitting further is easy once you have evidence about where the boundaries actually are.

Keep shared facts shared. The policies that must never differ between agents belong in common material, not duplicated per agent. Duplication is how two agents end up giving different answers to the same question six months apart, and it is almost always discovered by a customer rather than by you.

Give each agent an owner. Agents without a named owner drift: their material goes stale, their tone wanders and nobody notices until quality has visibly fallen. The owner does not need to be technical; they need to care whether the answers are still right.

Common questions

As many as your structure genuinely needs. The practical limit is maintenance rather than the platform — every agent needs an owner and current material, and agents without both go stale.

Yes, and usually they should. The voice appropriate for collections is not the voice appropriate for technical support, in either language.

They are routed to the right one and the conversation moves with them. Being bounced back and forth is a configuration failure, not an acceptable behaviour.

Only where you allow it. The default is separation, which is usually correct — an HR conversation has no business being visible to a sales agent.

Facts that must stay consistent should be shared deliberately. Everything function-specific stays separate, because retrieval quality falls as the corpus widens.

Yes. Sometimes the right split is by language rather than by department, because Arabic and English business registers differ enough that one configuration serves neither well.

See it on your own content

A working assistant on your own material, in both languages.