Cloud AI Setup on AWS, Azure and GCP
A governed environment on your chosen cloud where AI work can actually start — identity, networking, data residency, accelerator quota and cost guardrails settled before the first model rather than after the first invoice.
AI projects stall in cloud environments for reasons that have nothing to do with AI: a GPU quota request that takes three weeks, a network design that will not let a training job reach the data, an identity model where every data scientist is an administrator, and a bill nobody can attribute to a team.
Cloud AI setup is the configuration of a cloud environment for machine learning and language model workloads: account and project structure, identity and access, networking and private connectivity, data residency, accelerator capacity, managed AI service configuration, and cost attribution with guardrails.
Choosing between the clouds honestly
All three major providers can run essentially any AI workload competently, so the decision is rarely settled on capability. What actually decides it:
- Where your data already is. Moving large volumes between clouds is slow and expensive, and gravity usually wins the argument.
- Existing commitments and discounts. A committed spend agreement changes the effective price materially.
- Your team's existing skills. A capable platform your engineers cannot operate is worse than a slightly less capable one they can.
- Residency and sovereignty requirements. Which regions, which sovereign offerings, and which contractual terms your regulator will accept.
- Accelerator availability in your region. Which varies considerably and constrains plans more often than people expect.
We have no reseller relationship with any provider
The recommendation follows your data, your commitments, your team and your regulatory position. Where that means staying on the cloud you are on, or running a deliberate multi-cloud split, that is what the assessment says — and multi-cloud is more often a cost than a benefit, so it needs a reason.
What we set up
Account and project structure
Separation of environments and teams so that cost is attributable, blast radius is bounded and access can be granted at a sensible granularity. Retrofitting this onto a flat estate is disproportionately painful, which is why it comes first.
Identity and access with real least privilege
Roles for data scientists, engineers and services that permit the work without granting administrative rights to everyone. The common finding is an environment where every practitioner can delete production, which nobody intended and nobody noticed.
Networking and private connectivity
Private access to data stores and model endpoints, controlled egress, and connectivity to on-premises systems where training data lives there. Network design is the most frequent hidden blocker on cloud AI projects and the least visible from a project plan.
Data residency and sovereignty
Regions chosen and enforced by policy, with attention to where managed AI services actually process data — which is not always the region you deployed into. For EU obligations this needs to be verified against the provider's documented behaviour rather than assumed.
Accelerator capacity, requested early
GPU and accelerator quota varies by region and instance type and can take weeks to obtain. Requesting it in week one rather than week eight is one of the cheapest schedule protections available, and it is routinely forgotten.
Cost attribution and guardrails from the start
Tagging enforced by policy, budgets with alerts, and controls on the resource types that generate surprise bills. Attribution added later never covers the historical spend that prompted the request. See GPU cost optimisation.
Managed AI services configured deliberately
Each provider offers hosted model, training and endpoint services that can save substantial engineering effort, at the cost of some portability. We configure what earns its place and say plainly where the lock-in sits, so it is a decision rather than a discovery.
The landing zone, and what it prevents
| Without it | With it |
|---|---|
| GPU quota discovered as a blocker in month two | Capacity requested and confirmed in week one |
| Every practitioner effectively an administrator | Roles that permit the work and nothing else |
| Training jobs that cannot reach the data | Private connectivity designed before the first job |
| One infrastructure line item nobody can explain | Cost attributed by team, project and workload |
| Residency assumed from the deployment region | Processing locations verified against provider behaviour |
| Each team building its own environment differently | A repeatable, policy-enforced pattern |
The environment is delivered as infrastructure as code in your repositories, so it is reviewable, repeatable and yours to extend — not a configuration performed by hand that nobody can reproduce.
How the engagement runs
Accelerator quota is requested in the first week, because it is the most common schedule risk.
Assessment and quota
Requirements, residency and existing estate reviewed; accelerator quota requested immediately.
Design
Account structure, identity model, network topology and cost governance designed with your security and finance stakeholders.
Build
Landing zone implemented as infrastructure as code, with managed AI services configured where they earn their place.
Validation
Access model, private connectivity, residency behaviour and cost controls verified rather than assumed.
Handover
Documentation, provisioning patterns and working sessions with your platform team.
What you receive
A governed environment delivered as code, with quota confirmed and cost attributable from day one.
Landing zone as code
Account and project structure, policies and environments in your repositories.
Identity and access model
Least-privilege roles for practitioners, engineers and services, verified in practice.
Network design
Private connectivity to data and endpoints, controlled egress, on-premises links where needed.
Residency configuration
Region policy enforced, with processing locations verified against provider behaviour.
Cost governance
Enforced tagging, budgets, alerts and controls on the resources that cause surprises.
Provisioning patterns
Repeatable templates so the next team starts governed rather than improvising.
Is this the right engagement?
Worth being direct. Cloud AI Setup on AWS, Azure and GCP is the wrong spend in some situations, and those are listed rather than buried.
Good fit if
- AI work is starting on a cloud with no governed environment for it.
- Data scientists are blocked on access, quota or connectivity.
- Data residency or sovereignty obligations must be demonstrable.
- Cloud AI spend cannot be attributed to a team or project.
- Each team has built its own environment and none of them match.
Choose something else if
- A governed AI landing zone already exists and works.
- The requirement is reducing existing spend. See GPU cost optimisation.
- Workloads must run on premises or at the edge. See edge AI deployment.
- The cloud decision itself is unsettled at board level and no mandate exists.
Frequently asked questions
Marked up with FAQPage schema so these answers can surface directly in search results and inside AI assistant responses.
Which cloud is best for AI workloads?
All three run essentially any AI workload competently, so capability rarely decides it. What decides it is where your data already sits, your existing commitments and discounts, your team's skills, your residency obligations and accelerator availability in your region. We have no reseller relationship with any provider, so the answer can be 'stay where you are'.
Can we run AI workloads across more than one cloud?
You can, and it is usually more expensive than it looks: duplicated tooling, cross-cloud data transfer charges and a platform team maintaining two of everything. It is justified by genuine resilience or regulatory requirements, or by data that genuinely cannot be consolidated — not by a general wish to avoid lock-in.
How long does it take to get GPU capacity?
Anywhere from immediately to several weeks depending on the region, the instance type and your account history, which is exactly why we request it in the first week. Discovering a quota limit in month two is one of the most common and most avoidable delays in cloud AI projects.
How do we keep data in a specific region?
Through region policy enforced at the account or project level, plus verification of where managed AI services actually process data — which is not always the region you deployed into. For EU obligations this is checked against the provider's documented behaviour and contractual terms rather than inferred from the console.
Do you use managed AI services or build from components?
Both, chosen deliberately. Managed training, endpoint and model services save real engineering effort and increase portability cost; we configure what earns its place and state plainly where lock-in sits, so it is a decision your architects made rather than something discovered during a later migration.
Often paired with this
Most clients combine two or three engagements from the MLOps, LLMOps and AI Infrastructure pillar. These are the ones that most often run immediately before or after.
GPU Infrastructure and Cost Optimisation
Utilisation profiling, right-sizing, scheduling and purchasing strategy, measured before anything is changed.
Read more →MLOps Pipeline Implementation
Versioned data and models, automated retraining, tested deployment and a rollback anyone on the team can run.
Read more →Edge AI Deployment
Models sized to real hardware, offline operation, safe over-the-air updates and fleet-wide monitoring.
Read more →Is this the right engagement?
Tell us what you are trying to build. If a different service fits better, or if you do not need us at all, we will say so.