COE 558Lecture 02Part 08
Public cloud hyperscalers: AWS, GCP and Azure
Surveys the history and service portfolios of the three biggest hyperscalers.
- Concepts
- 5
- Slides
- 49-56
- Reading
- 30 min
Why this part matters
A dissertation prototype or an industrial edge-cloud system almost always ends up on one of three platforms: Amazon Web Services, Google Cloud or Microsoft Azure. Knowing them is not about memorising product logos. It is about knowing what is truly comparable between them, because that decides how hard a migration is, what a design costs, and how locked in you become.
This part first asks why only three companies dominate a market with countless providers. It then follows each one from its origin to its present portfolio, and ends with a single map that translates any service between the three. Expect the exam to hand you a service on one provider and ask for its equivalent on another, and to name the block it belongs to.
By the end you can
- Explain why only a few companies can operate at hyperscale, using economies of scale and the Berkeley motives for becoming a provider.
- Trace how AWS, Google Cloud and Azure began, with key dates, and link each origin to its strengths.
- Read any provider portfolio through the six building blocks.
- Translate a service between the three providers with an equivalence table.
- Spot outdated or misspelled service names on the slides.
Picture two data centers in 2006. The first is a medium-sized facility with about 1,000 servers. The second is a very large one with about 50,000. The first pays about $95 per Mbit/s of network bandwidth per month. The second pays about $13. Storage costs the first $2.20 per GB per month and the second $0.40. One administrator in the first looks after about 140 servers, while in the second one person runs more than 1,000. These are James Hamilton's estimates, reported in the Berkeley “Above the Clouds” report, the same table part 01 used to explain why the cloud stays centralized. Here it explains who can afford to be the cloud.
That gap is the whole story of the Hyperscaler. The slides note that there are a great many public cloud service providers, yet they focus on three. The reason is economic. A provider that buys bandwidth, disks, power and staff at one fifth to one seventh of a medium firm's price can rent capacity by the hour below what that firm pays to own it, and still make a profit. This is Utility computing sold at a wholesale cost base. It is also why the Public cloud model rewards size so strongly: every new customer spreads the fixed cost of software and operations over more machines.
Worked example
What scale buys
Network
A service needs 100 Mbit/s of sustained bandwidth. The medium data center pays 100 × $95 = $9,500 per month. The very large one pays 100 × $13 = $1,300.Storage
10 TB stored for a month costs about 10,000 × $2.20 = $22,000 in the medium facility and 10,000 × $0.40 = $4,000 in the large one.People
Running 5,000 servers takes about 36 administrators at 140 servers each, but at most 5 at 1,000 servers each.Result
The raw prices give ratios of 95 / 13 ≈ 7.3 and 2.20 / 0.40 = 5.5. The report prints 7.1, 5.7 and 7.1, so quote the advantage as “about 5.7 to 7.1 times, as reported”. Other estimates in the same report put it at 3 to 5 times.
Who could become a provider, and why
Big data centers are necessary but not sufficient. Armbrust and colleagues argue that a provider also needs large-scale software infrastructure (MapReduce, the Google File System, Bigtable, Dynamo) and the operational expertise to run and defend it. In the early 2000s only a handful of internet companies had both. Given those assets, the report lists six motives; the four that matter for the hyperscalers are:
- Make a lot of money from economies of scale, turning a cost advantage into margin.
- Leverage existing investment, adding a revenue stream on top of data centers already built for internal use.
- Defend a franchise, giving existing enterprise customers a cloud path before a rival does.
- Attack an incumbent, establishing a beachhead before a single “800 pound gorilla” emerges (the report's example is Google App Engine).
The other two are leveraging customer relationships (as with IBM) and becoming a platform (as with Facebook). Keep the four above in mind. The next concepts show that each hyperscaler's origin story fits one of them especially well. The market today reflects the same concentration: in the second quarter of 2026 the three largest providers took almost two thirds of a fast-growing market.
Cloud infrastructure services market, Q2 2026 (Synergy Research Group)
- Quarter
- Q2 2026
- Total cloud infrastructure services revenue
- $143.4B, +43% year on year
- Amazon (AWS)
- 28%
- Microsoft (Azure)
- 20%
- Google (Google Cloud)
- 15%
- Top three combined
- 63%
Recall
Which four Berkeley motives for becoming a cloud provider matter most for the hyperscalers?
By the early 2000s, Amazon's engineering leaders noticed something expensive. Teams building shopping features were spending about 70% of their time on the same basic plumbing: storage, compute, databases. Amazon later called this undifferentiated heavy lifting: work every team must do that does not set the product apart. Leaders began to think about building “a shared layer of infrastructure services that all these teams can rely on”. They soon saw that outside developers had exactly the same problem, and at a 2003 offsite at Jeff Bezos's house the top managers decided to sell those services to everyone else.
This is the Berkeley motive “leverage existing investment” in its purest form. Werner Vogels, Amazon's CTO, said that many AWS technologies were first developed for Amazon's internal operations. The slide adds the design pressures behind that platform: absorbing huge traffic spikes such as the holiday season, supporting a global store, and moving services onto commodity Linux hardware and open-source software. Building for those pressures forced Amazon to make infrastructure into reusable services with clean interfaces, which is exactly what a customer of a Cloud service provider (CSP) needs.
The public launches came in quick succession. A message queue, SQS, appeared in beta in November 2004. In March 2006 Amazon announced object storage as “storage for the Internet”. In August 2006 EC2 opened in limited beta, renting a Virtual machine (VM) at 10 cents per virtual CPU hour. EC2 made the Infrastructure as a Service (IaaS) model concrete: start a server with an API call in minutes, add more when load rises (Rapid elasticity), and pay only for the hours used (Measured service).
From internal platform to public cloud
- 2003
- Offsite at Jeff Bezos's house: leaders decide the planned shared infrastructure services can be an external business
- November 2004
- Simple Queue Service (SQS) public beta. Fortune still counts S3 as AWS's first official launch
- 13 March 2006
- Amazon S3 announced as “storage for the Internet”
- 25 August 2006
- Amazon EC2 limited beta at 10 cents per virtual CPU hour
- Today
- Over 200 services, used in about 190 countries (AWS overview whitepaper)
Recall
Why did Amazon build the shared platform that became AWS?
A portfolio slide with twenty product icons looks like a list to memorise. It is easier to read it as a floor plan. Take a small web shop and place each piece of it on the AWS portfolio from slide 51.
Worked example
Placing a web shop on the six blocks
Networking
Route 53 resolves the shop's name. CloudFront, a Content delivery network (CDN), caches product images close to shoppers. Elastic Load Balancing spreads requests over the web servers, which sit inside a private VPC. Direct Connect would add a dedicated line from the company's own office or data center.Compute
The web servers run on EC2 instances, each a Virtual machine (VM). A service packaged as a Container could run on ECS instead.Database
Orders go to a managed database: DynamoDB for key-value access at scale, or RDS for a relational schema. Shopping sessions live in ElastiCache, an in-memory cache.Application Services
SQS queues new orders so payment and shipping can process them at their own pace. SNS sends the “your order has shipped” notifications. CloudSearch would power product search.Storage
Product photos sit in S3 as objects. The servers' disks are EBS block volumes. Invoices older than a year move to cheap archive storage (Glacier on the slide).Deployment and Management
CloudFormation describes the whole setup as code so it can be rebuilt. CloudWatch monitors it, and IAM controls who may touch what. Elastic Beanstalk could instead deploy the app as a Platform as a Service (PaaS) with much of this wiring done for you.Result
Every part of the shop lands in one of six blocks: Deployment and Management, Networking, Application Services, Compute, Storage, Database.
The general rule is that every hyperscaler portfolio fits the same six blocks. The slides draw Google Cloud and Azure with exactly the same layout. Learn the categories once and the vendor names become vocabulary: what changes between providers is the name on the box, not the job it does. That is the point of slide 52, which skips a detailed tour of every CSP because the big ones offer similar or comparable services. Building on Infrastructure as a Service (IaaS) or Platform as a Service (PaaS) on any of the three means picking a box in each block.
Comparable is not identical, though. Microsoft's own guide for AWS professionals says the two clouds offer similar products built independently, and warns that “not every matched service has exact feature-for-feature parity”. Even the containers that hold resources differ: AWS organises them under accounts, Azure under subscriptions. These differences in APIs, limits and behaviour are exactly where Vendor lock-in comes from.
Recall
Name the six portfolio blocks and one service per provider for Compute, Storage and Database.
In April 2008 a developer could upload a Python web application to Google App Engine and never see a server. Google ran it on its own infrastructure, scaled it as traffic grew, balanced the load and stored the data in Bigtable and the Google File System. The free preview gave each application 500 MB of storage, 200 million megacycles of CPU per day and 10 GB of bandwidth per day, enough for about 5 million page views a month. Raw virtual machines came much later: Compute Engine (often abbreviated GCE) was announced at Google I/O on 28 June 2012 and became generally available on 2 December 2013.
So Google entered the cloud from the opposite end to Amazon. AWS started with hardware-like building blocks and later added managed services. Google started with a fully managed Platform as a Service (PaaS) and only later offered Infrastructure as a Service (IaaS). The Berkeley report turns this into a general way to tell providers apart: by the level of abstraction they present to the programmer, the spectrum part 07 used to place IaaS and PaaS. EC2 sits at the low end, where an instance looks like physical hardware and you control nearly the whole software stack from the kernel up. App Engine sits at the high end, where Google enforced a clean split between a stateless request-reply compute tier and a stateful storage tier, and in return gave you automatic scaling and high availability. Azure, as launched, sat in between.
In Berkeley terms, App Engine is the report's example of “attack an incumbent”: its appeal lay in automating the scalability and load balancing features that developers would otherwise build themselves.
| Platform | What you control | What the provider automates | Constraint |
|---|---|---|---|
| Amazon EC2 | Nearly the whole software stack, from the kernel upwards | Virtualized hardware through a thin API of a few dozen calls | None on application type; scaling and failover are your problem |
| Microsoft Azure (2009) | Your code and language choice, not the OS or runtime | Network configuration, some failover and scaling, once you declare properties | Code compiled to the .NET Common Language Runtime |
| Google App Engine | Only the request handlers of a web application | Servers, load balancing, automatic scaling, Bigtable-backed storage | Stateless request-reply compute tier, rationed CPU per request |
Azure automated some scaling and failover only after you declared your application's properties, so the figure leaves scaling with you.
The constraint is the price of the convenience. App Engine could scale your application automatically precisely because it knew its shape. A general-purpose program that keeps state in memory or needs long computations did not fit. EC2 would run anything, but scaling and failover were left to you because the provider cannot know how your application replicates state. Neither end is better; the right point depends on the task, which is why every hyperscaler now offers the whole range.
Google's portfolio on slide 54 shows that convergence. App Engine is still there in Deployment and Management, now beside Compute Engine for raw VMs, a managed Kubernetes service for each Container, Cloud Storage, Persistent Disk, Bigtable and Cloud SQL. The Application Services block also lists Pub/Sub for messaging and Dataflow and Dataproc for data processing, a reminder of Google's strength in large-scale data. The six blocks are the same as Amazon's; only the names differ.
Quick check
On the 2009 Berkeley abstraction spectrum, how were the three platforms ordered from lowest to highest abstraction?
Picture a company that runs Windows Server in its own data center, signs its staff in through Active Directory, and works in Office every day. Its developers write .NET. For this company, Azure is the shortest road to the cloud: the identities, licences and skills it already has carry straight over. That is the Berkeley motive “defend a franchise”, and the report names Azure as its example: it “provides an immediate path for migrating existing customers of Microsoft enterprise applications”.
Ray Ozzie unveiled Windows Azure at the Professional Developers Conference on 27 October 2008. It became generally available on 1 February 2010 in 21 countries, with full service level agreements. On 25 March 2014 Microsoft announced it would rename the platform Microsoft Azure, and the new name took effect on 3 April 2014. The rename signalled that Azure was no longer only about Windows: by then it ran Linux virtual machines and open-source stacks too. The enterprise focus on slide 55 remains its strongest card, with deep integration across Windows, Active Directory and Microsoft 365.
One map for all three providers
Slide 56 completes the set, and the portfolio once again falls into the same six blocks. Putting the three slides side by side gives the single most useful table in this part. Read each row as one need, then look up the local name on each provider. Where the slide's label is outdated or misplaced, the table gives the current service and notes the slide's version.
| Block | Need | AWS (slide 51) | Google Cloud (slide 54) | Azure (slide 56) |
|---|---|---|---|---|
| Deployment and Management | PaaS app hosting | Elastic Beanstalk | App Engine | App Service (slide: Web Apps) |
| Deployment and Management | Infrastructure as code | CloudFormation | Deployment Manager (retired, use Infrastructure Manager) | Resource Manager |
| Deployment and Management | Monitoring | CloudWatch | Cloud Monitoring | Azure Monitor |
| Deployment and Management | Identity and access | IAM | Cloud IAM | Microsoft Entra ID with Azure RBAC (slide shows AD Domain Services) |
| Networking | DNS | Route 53 | Cloud DNS | Azure DNS |
| Networking | Private network | VPC | VPC (slide: Virtual Network) | Virtual Network |
| Networking | CDN | CloudFront | Cloud CDN | Azure Front Door (slide: Content Delivery Network) |
| Networking | Load balancing | Elastic Load Balancing | Cloud Load Balancing | Load Balancer, Traffic Manager |
| Networking | Dedicated link | Direct Connect | Cloud Interconnect | ExpressRoute (not on slide) |
| Application Services | Messaging | SQS, SNS | Pub/Sub | Queue Storage (Service Bus per Google's table) |
| Application Services | Search | CloudSearch (closed to new customers) | Not on slide | Azure AI Search |
| Compute | Virtual machines | EC2 | Compute Engine | Virtual Machines |
| Compute | Containers | ECS (slide); EKS for Kubernetes | GKE | AKS (slide: Containers) |
| Storage | Object | S3 | Cloud Storage | Blob Storage |
| Storage | Block | EBS | Persistent Disk | Managed Disks |
| Storage | Archive or file | S3 Glacier classes | Cloud Storage Archive class | File Storage (the slides mix archive and file) |
| Database | NoSQL | DynamoDB | Bigtable | Cosmos DB |
| Database | Relational | RDS | Cloud SQL | SQL Database |
| Database | In-memory cache | ElastiCache | Memorystore | Azure Managed Redis (slide: Redis Cache; Azure Cache for Redis is being retired) |
The equivalences are close enough to plan a design and to answer exam questions, but each row hides differences. Azure's Cosmos DB, for example, is a fully managed NoSQL and vector database that offers document, key-value, graph and table models, and a 99.999% Availability SLA for multi-region setups. That is broader than DynamoDB or Bigtable, so a migration between them is a redesign, not a copy. This is the practical face of Vendor lock-in on any Hyperscaler. Slide 56 also shows Communication Services, Azure's APIs for chat, SMS, voice and video, which has no box on the other two slides.
Recall
Put the hyperscaler launches in order with dates.
Recall
Why is “AD Domain Services” not Azure's answer to AWS IAM?
Recall
Why is moving from DynamoDB to Cosmos DB a redesign rather than a copy, even though both sit in the same block?
Quick check
A team runs Amazon DynamoDB and must move to Azure. Which box on the Azure portfolio is the closest counterpart?
Quick check
Which Berkeley motive best explains Microsoft launching Azure?
Quick check
Which label on the slides names a service that has since been renamed?
Recap
If you remember nothing else
- Hyperscale is economics: very large data centers buy network, storage and staff at roughly 1/5 to 1/7 of medium-sized prices.
- The Berkeley report gives six motives for becoming a provider; the hyperscalers fit leverage existing investment (Amazon), attack an incumbent (Google App Engine) and defend a franchise (Microsoft).
- The top three hold 63% of cloud infrastructure revenue (Q2 2026: Amazon 28%, Microsoft 20%, Google 15%).
- AWS grew out of Amazon's internal shared platform: S3 and EC2 in 2006.
- Google started at the PaaS end with App Engine (Apr 2008) and added Compute Engine (GA Dec 2013).
- Azure was announced Oct 2008, reached GA Feb 2010, was renamed Microsoft Azure in 2014, and plays to Microsoft's enterprise base.
- Providers differ by abstraction level: EC2 gives control (you manage from the kernel up), App Engine gives convenience (automatic scaling, constrained app shape), and Azure (2009) sat between.
- Every portfolio uses the same six blocks. Services are comparable, not identical, and the differences drive lock-in.
- Slides go stale: Container Engine is now GKE, Deployment Manager gives way to Infrastructure Manager, “Cosmo DB” is Cosmos DB, and IAM maps to Entra ID plus RBAC, not AD Domain Services.
Sources
- Above the Clouds: A Berkeley View of Cloud Computing (Armbrust et al.)PaperUC Berkeley EECS, UCB/EECS-2009-28Table 2 economies of scale, the six motives for becoming a provider, Vogels on AWS origins, and the EC2, Azure, App Engine abstraction spectrum.(opens in a new tab)
- SP 800-145: The NIST Definition of Cloud ComputingDocsNISTDefinitions of public cloud and of SaaS, PaaS and IaaS.(opens in a new tab)
- Q2 cloud market passes $143 billion, highest growth rate in eight yearsArticleSynergy Research GroupQ2 2026 market size, growth and provider shares.(opens in a new tab)
- The history of Amazon Web ServicesArticleFortuneThe 70 percent figure, the 2003 offsite and the shared infrastructure layer.(opens in a new tab)
- The AWS Blog: The First Five YearsDocsAWS News BlogAmazon Simple Queue Service introduced in November 2004, before S3 and EC2.(opens in a new tab)
- Announcing Amazon S3, Simple Storage ServiceDocsAmazon Web Services13 March 2006 announcement.(opens in a new tab)
- Amazon EC2 betaDocsAWS News BlogLimited beta on 25 August 2006 at 10 cents per virtual CPU hour.(opens in a new tab)
- Overview of Amazon Web Services: introductionDocsAmazon Web ServicesAWS began offering IT infrastructure services in 2006; over 200 services today.(opens in a new tab)
- Transition from Amazon CloudSearch to Amazon OpenSearch ServiceDocsAWS Big Data BlogCloudSearch closed to new customers on 25 July 2024.(opens in a new tab)
- Amazon Glacier developer guide: introductionDocsAmazon Web ServicesThe vault-based service no longer accepts new customers; use the S3 Glacier storage classes.(opens in a new tab)
- Introducing Google App EngineDocsGoogle App Engine BlogApril 2008 launch, features and preview quotas.(opens in a new tab)
- Google Compute Engine is now generally availableDocsGoogle Cloud BlogGeneral availability on 2 December 2013.(opens in a new tab)
- Introducing Certified Kubernetes and Google Kubernetes EngineDocsGoogle Cloud BlogContainer Engine renamed Google Kubernetes Engine in November 2017.(opens in a new tab)
- Deployment Manager deprecationDocsGoogle CloudSupport ended 1 April 2026; replacement is Infrastructure Manager.(opens in a new tab)
- Compare AWS and Azure services to Google CloudDocsGoogle CloudCross-provider service mapping used for the equivalence table.(opens in a new tab)
- Hitting send on the next 15 years of GmailArticleGoogleGmail launched on 1 April 2004.(opens in a new tab)
- Microsoft unveils Windows Azure at Professional Developers ConferenceArticleMicrosoftAnnouncement on 27 October 2008.(opens in a new tab)
- Windows Azure general availabilityArticleMicrosoftGeneral availability on 1 February 2010 in 21 countries.(opens in a new tab)
- Upcoming name change for Windows AzureDocsMicrosoft Azure BlogRename to Microsoft Azure announced 25 March 2014, effective 3 April 2014.(opens in a new tab)
- Azure for AWS professionalsDocsMicrosoft LearnIndependent implementations without exact parity, accounts versus subscriptions, and IAM mapped to Entra ID and RBAC.(opens in a new tab)
- New name for Azure Active DirectoryDocsMicrosoft LearnAzure AD renamed Microsoft Entra ID; Azure AD Domain Services renamed Microsoft Entra Domain Services.(opens in a new tab)
- Azure Cosmos DB overviewDocsMicrosoft LearnFully managed NoSQL and vector database, multiple data models, 99.999 percent multi-region SLA.(opens in a new tab)
- Retirement of Azure Cache for Redis: FAQDocsMicrosoft LearnAzure Cache for Redis tiers are being retired in favour of Azure Managed Redis.(opens in a new tab)