
Sovereign AI vs Cloud AI in Saudi Arabia: Two Gates, Seven Scores, and Why One Policy for the Whole Company Is Wrong
The same company can correctly run one workload on a cloud API and another on-premises. A seven-dimension scoring sheet, which form of hybrid survives staff turnover, and a worked example for a Saudi insurer.
“Sovereign or cloud” is usually presented as an identity — the kind of organisation you are — when it is a decision with inputs. Different workloads inside the same company will correctly land on different answers, and a policy that forces all of them one way is the actual mistake.
This article gives a decision procedure with a defined order, a scoring sheet you can take to a steering committee, and the two questions that decide the outcome before any of the scoring begins. It is deliberately opinionated about sequence, because getting the order wrong is how organisations end up optimising cost on a workload where cost was never the constraint.
Start by separating your workloads, not your company
The most common error is asking the question at organisation level. A single Saudi bank might run four AI workloads with genuinely different correct answers, and treating them as one procurement produces a policy that is wrong for three of them.
| Workload | Data involved | Where it lands | Why | |
|---|---|---|---|---|
| Public website FAQ | Published material only | Cloud API is fine | Nothing confidential is in the prompt | |
| Internal HR policy assistant | Employee records, salary bands | On-premises or a controlled region | Personal data of identifiable staff | |
| Customer account assistant | Account balances, transactions | On-premises | Regulated financial data, in every prompt | |
| Marketing copy drafting | Nothing sensitive | Cloud API is fine | The constraint is quality, not residency |
The same organisation, the same year, four correct and different answers. Any policy that produces one answer for all four is wrong three times.
The two questions that decide it before scoring
Most of the framework below is only worth running if you get past two gates. They are binary, and they eliminate options regardless of every other consideration.
- Gate 1Does regulated or personal data enter the prompt?If yes, cross-border transfer conditions apply to every single answer, not to an annual export. Options that cannot evidence residency are eliminated here, whatever they cost.
- Gate 2Must the answer path be verifiable by your own team?Some obligations are discharged by contract; some accountable teams need to check for themselves. If a regulator will ask you — not your vendor — then attestation is not evidence.
If both gates pass to cloud, run the scoring below and let cost and speed decide. If either gate closes, you are choosing between on-premises and a controlled in-Kingdom region, and the scoring is about which of those, not about whether.
The scoring sheet
Weights should be set by your own risk function before anyone sees the scores, not after. Setting weights after seeing scores is how a predetermined answer gets a process wrapped around it.
| Dimension | The question that scores it | Typical weight | |
|---|---|---|---|
| Residency evidence | Can you prove where the prompt went, without asking the vendor? | High when Gate 1 is live | |
| Time to production | Weeks to a working system with your own documents in it | High for a pilot, low for core systems | |
| Cost at your real volume | Not list price — your measured questions per day, on Arabic text | Medium; frequently a tie | |
| Model flexibility | Can you change model without changing vendor? | Higher than most buyers assume | |
| Continuity | If the link or the provider fails, does it keep answering? | High for anything customer-facing | |
| Operating burden | Who patches it, who watches it, at 2am | Often underweighted, and it is real | |
| Exit cost | Who holds the index, the tuning and the transcripts | Underweighted until the renewal |
Six of the seven are routinely scored. The last two are what organisations wish they had weighted properly, two years in.
Directional, drawn from how the conversation changes between a first purchase and a renewal. The pattern is consistent enough to plan around.
What each option is genuinely good at
An honest comparison has to include the cases where the other option wins, so here are both, stated as strongly as their advocates would put them.
| Cloud AI | Sovereign / on-premises | ||
|---|---|---|---|
| Best at | Reaching a working system fast, and access to the largest models | Verifiability, continuity, and flat cost at volume | |
| Genuinely wins when | You are piloting, volumes are low, or nothing sensitive is in the prompt | Regulated data is in every prompt, or you must evidence residency yourself | |
| Real weakness | Residency is contractual, and cost grows with your success | Slower to start, bounded by your hardware, and you operate it | |
| Failure mode | A dependency you cannot inspect and cannot fix | A box you outgrew and nobody planned the next one | |
| Underrated risk | Model deprecation forcing an unplanned migration | Skills concentration in one or two people |
If a vendor cannot articulate the row labelled “real weakness” about their own option, they are selling rather than advising.
The hybrid that actually works, and the one that does not
Hybrid is the answer most organisations reach, and it is right — but only in one of its two common forms. The distinction is worth being precise about, because both get called hybrid in the same meeting.
| Routing by data class — works | Splitting one answer — usually does not | ||
|---|---|---|---|
| How it works | Sensitive workloads on-premises; public workloads on cloud | One answer assembled from both a local and a remote model | |
| Where the sensitive data goes | Never leaves | Depends on the split, which is easy to get subtly wrong | |
| Can you explain it to a regulator | Yes, in one sentence | Only with a diagram, and diagrams drift from code | |
| Failure mode | A workload gets classified wrongly — visible and fixable | A prompt-construction change silently moves data across the line | |
| Our view | Recommended | Only with strong controls and continuous verification |
The first is a routing decision made once per workload. The second is a decision made implicitly on every request by whoever last edited the prompt template.
A worked scoring, end to end
To make it concrete: a mid-sized Saudi insurer deciding on a customer-facing claims assistant, weights set by their risk function in advance.
Illustrative of the method, not a recommendation for your organisation. Change the weights and the ranking changes — which is the point of setting them first.
Notice that the foreign API scores respectably on speed and model access and still loses decisively, because Gate 1 closed before scoring began. That is the framework working correctly: gates eliminate, scores rank what survives. Running the scores first and then noticing the gate is how organisations talk themselves into an architecture they cannot defend.
Weight reversibility explicitly. The cheapest decisions to get wrong are at the top; the ones at the bottom deserve the time the ones at the top usually get.
Questions to bring to the vendor meeting
- 1Does answering call a model you do not run?A one-word answer, then evidence. Hesitation here is the finding.
- 2Where does indexing run?Often a different answer from where answering runs, and a bulk transfer if outsourced.
- 3Can I change the underlying model?If no, you have two dependencies, not one.
- 4What happens when the link drops?“It queues” and “it stops” are different products.
- 5Show me the log for one answerReconstructing a single answer is the test of whether you can investigate a complaint.
- 6What do I keep if I leave?Index, embeddings, tuning, transcripts. Get it in the contract, not the conversation.
Honest limits
The weights, scores and regret figures above are illustrative of a method. They are not measurements, and they are not a recommendation for your organisation — a different risk appetite legitimately produces a different ranking from the same framework.
We have a commercial interest here, and it would be dishonest not to say so: we build on-premises systems. The framework above is written to be usable against us, which is why the cloud column includes the cases where cloud genuinely wins, and why Gate 1 passing to cloud is treated as an ordinary and correct outcome rather than a failure.
What we state as fact is our own architecture: Elbi answers from your material on hardware you control inside the Kingdom, with indexing on the same hardware and no external model call in the answer path. If your workload passes both gates to cloud, we are not the right answer for that workload, and we would rather say so than sell it.
Common questions
The question belongs at workload level, not organisation level. A single company can correctly run a public website FAQ on a cloud API while running a customer account assistant on-premises, because the deciding input is what personal or regulated data ends up inside the prompt. A policy that forces every workload one way will be wrong for most of them.
Two binary gates, applied before any cost comparison. First, does regulated or personal data enter the prompt — if yes, cross-border transfer conditions apply to every answer, not to an annual export, and options that cannot evidence residency are eliminated. Second, must the answer path be verifiable by your own team rather than by contractual attestation. If both gates pass to cloud, let cost and speed decide.
Hosting describes where the application runs. Sovereignty in the sense that matters describes whether producing an answer requires a call to a model you do not run. A system can be hosted in Riyadh and still send every customer record abroad inside a prompt. Ask whether answering calls an external model, and ask for a network diagram or firewall rule rather than an assurance.
One form does and one usually does not. Routing by data class — sensitive workloads on-premises, public workloads on cloud — works, because it is one decision per workload, explainable to a regulator in a sentence, and it survives staff turnover. Splitting a single answer between a local and a remote model is much harder to keep correct, because a prompt-construction change can silently move data across the line on every request.
Continuity, operating burden and exit cost. Time to production and price are heavily weighted at purchase, while what a system does when the link drops, who patches and watches it, and who holds the index, tuning and transcripts if you leave are rarely scored. Exit cost in particular is almost never weighted at purchase and is the most common source of regret at renewal.
Keep reading

HUMAIN, ALLaM and the sovereign AI stack
A Saudi model served over a foreign cloud is still a foreign network call. Here is what HUMAIN provides, where ALLaM actually runs, and the six questions that separate a real residency claim from a hosting address.
Read Article
PDPL and data residency for AI
The source records, the retrieval index, the prompt and the logs each leave the Kingdom by a different route. Why embeddings are usually personal data, why indexing is the largest transfer in the project, and twelve questions to ask.
Read Article
ChatGPT vs a custom AI chatbot
The honest answer is that they are not competing products. One is a brilliant general tool for your staff; the other is a narrow, accountable system for your customers. Choosing wrongly costs money in both directions.
Read Article