Skip to main content

Choosing the Right Cloud Load Balancer Type

Learn to distinguish GCP's global external, regional external, and internal load balancer types based on traffic type and reach needed.

38 Terms

Overview and What You Will Learn

In this lab, you will create a Global External HTTP(S) Load Balancer with a health check and backend service, then compare its configuration against an internal load balancer meant only for traffic within the VPC - directly experiencing why "load balancer" in GCP is really a family of distinct products, not a single service.

Why This Matters in Production

A team builds an internal microservice meant to only ever be called by other services inside the same VPC, but configures it behind a Global External HTTP(S) Load Balancer - accidentally exposing an internal-only service to the public internet, discovered only during a security review months later. Choosing the correct load balancer type from the start is a security decision, not just a performance one.

Core Principles

GCP's load balancing products span from global, external, internet-facing options down to internal, VPC-only options - the right choice depends on traffic type, whether the backend needs to be reachable from outside the VPC at all, and whether traffic needs global or regional distribution.

◈ DIAGRAM
+------------------------------------------+
| Global External HTTP(S) Load Balancer |
| Public-facing, distributes globally |
| Single global IP address |
| Best for: public web applications |
+------------------------------------------+
| Regional External Load Balancer |
| Public-facing, but scoped to one region |
| Best for: regional compliance or latency |
| requirements |
+------------------------------------------+
| Internal Load Balancer |
| Reachable only from within the VPC |
| Best for: internal microservices, backend |
| tiers never meant to be public |
+------------------------------------------+

Detailed Step-by-Step Practical Lab

  1. Create a project and enable Compute Engine:
Bash
gcloud projects create gcp-lb-lab-2026 --name="Load Balancer Lab"
gcloud config set project gcp-lb-lab-2026
gcloud services enable compute.googleapis.com
  1. Reserve a static global external IP for a public-facing load balancer:
Bash
gcloud compute addresses create lb-ip-public --global
  1. Create a health check the load balancer will use to verify backend health:
Bash
gcloud compute health-checks create http health-check-web --port=80
  1. Create a backend service and a global HTTP(S) load balancer using it:
Bash
gcloud compute backend-services create backend-web-public \
--protocol=HTTP \
--port-name=http \
--health-checks=health-check-web \
--global
gcloud compute url-maps create lb-map-public \
--default-service=backend-web-public
gcloud compute target-http-proxies create lb-proxy-public \
--url-map=lb-map-public
gcloud compute forwarding-rules create lb-forwarding-public \
--address=lb-ip-public \
--global \
--target-http-proxy=lb-proxy-public \
--ports=80
  1. Now create an Internal Load Balancer for a service meant to stay reachable only within the VPC:
Bash
gcloud compute health-checks create tcp health-check-internal --port=8080
gcloud compute backend-services create backend-internal-service \
--protocol=TCP \
--health-checks=health-check-internal \
--load-balancing-scheme=INTERNAL \
--region=asia-south1
Note

--load-balancing-scheme=INTERNAL is what specifically makes this an internal-only load balancer - there is no global IP, no public-facing proxy, and the resulting backend is reachable only from within the same VPC (or a peered/connected network), never from the public internet.

  1. Confirm the difference in reachability by checking each load balancer's forwarding configuration:
Bash
gcloud compute forwarding-rules list \
--filter="name~'lb-forwarding' OR name~'backend-internal'"
  1. Clean up:
Bash
gcloud compute forwarding-rules delete lb-forwarding-public --global --quiet
gcloud compute target-http-proxies delete lb-proxy-public --quiet
gcloud compute url-maps delete lb-map-public --quiet
gcloud compute backend-services delete backend-web-public --global --quiet
gcloud compute backend-services delete backend-internal-service --region=asia-south1 --quiet
gcloud compute health-checks delete health-check-web health-check-internal --quiet
gcloud compute addresses delete lb-ip-public --global --quiet
gcloud projects delete gcp-lb-lab-2026 --quiet

Production Best Practices & Common Pitfalls

Common Mistake

Configuring a backend service meant only for internal traffic behind a Global External Load Balancer by default, accidentally exposing it to the public internet. Always confirm --load-balancing-scheme=INTERNAL (or the equivalent internal-only configuration) is used for any backend that should never be reachable from outside the VPC.

Tip

A Global External HTTP(S) Load Balancer is the right default for a public-facing web application needing traffic distributed across multiple regions - it uses a single global IP address and Google's own global network to route users to the nearest healthy backend, similar in concept to what Azure Front Door or AWS Global Accelerator provide, but natively built into GCP's core load balancing product.

  • A Regional External Load Balancer trades global reach for staying within a single region - appropriate when a compliance requirement or a very region-specific latency need makes global distribution unnecessary or undesirable.
  • Health checks differ by protocol and load balancer type. An HTTP health check (used for the public web load balancer) checks a specific path and expects an HTTP response; a TCP health check (used for the internal load balancer here) simply confirms a port accepts connections, a distinction worth matching to the actual backend protocol.

Quick Reference & Troubleshooting Commands

Command Description
gcloud compute addresses create --global Reserve a static global external IP
gcloud compute backend-services create --load-balancing-scheme=INTERNAL Create an internal-only backend service
gcloud compute health-checks create http Create an HTTP-based health check
gcloud compute forwarding-rules list List active load balancer forwarding rules

Resources

Artifact Registry

Google Cloud's current, unified package and container image storage service, the successor to the older Container Registry. It supports container images alongside language packages like npm, Maven, and Python packages in one service, and supports cleanup policies to automatically remove old, untagged image versions.

BigQuery

Google Cloud's serverless data warehouse for analytics, priced based on the volume of data a query actually scans rather than pre-provisioned compute capacity. Selecting only needed columns and querying a properly partitioned table's relevant date range are the two highest-impact levers for controlling BigQuery cost.

Bigtable (GCP)

A wide-column NoSQL database built for very high throughput at extremely low latency, suited to time-series data, IoT telemetry, and analytics workloads with massive write volume. It is a specialized tool distinct from Firestore's document model or Cloud SQL's relational model.

Billing Account (GCP)

A separate object from a Project that holds the actual payment method for Google Cloud usage. One Billing Account can be linked to many Projects, all of whose usage is billed to that same account, letting a company centralize payment while individual Projects retain their own separate IAM control over resources.

Cloud Armor

A security policy service attached to a Global External HTTP(S) Load Balancer, filtering malicious requests before they reach the backend. It supports rate-based rules blocking a single source IP exceeding a request threshold, and preconfigured WAF rules blocking common attack patterns like SQL injection, Google's equivalent of AWS WAF or Azure WAF.

Cloud Functions (GCP)

Google Cloud's event-driven, single-purpose compute service, the most granular and shortest-lived compute unit in GCP's compute spectrum. It fits a narrow niche compared to Cloud Run - a specific triggered action like resizing an uploaded image, rather than a full application with multiple routes and business logic.

Cloud Load Balancing

Google Cloud's family of load balancing products, spanning Global External HTTP(S) Load Balancers (public, distributed globally via a single IP), Regional External Load Balancers, and Internal Load Balancers reachable only from within the VPC. The right type depends on traffic type and whether the backend should ever be reachable from the public internet.

Cloud Logging

Google Cloud's service for collecting, storing, and querying log data from every GCP service and custom applications. Its query filter syntax lets logs be narrowed by resource, severity, and specific field values, and Log Sinks can route matching logs to BigQuery, Cloud Storage, or Pub/Sub for long-term retention beyond the default window.

Cloud Monitoring

Google Cloud's metrics, dashboards, and alerting service. An alert policy requires a Notification Channel explicitly attached to actually notify someone when it fires - without one, the policy still evaluates and logs its own firing history, but nobody is ever told, a common and easy-to-miss misconfiguration.

Cloud NAT

A service providing outbound-only internet connectivity for resources with no external IP address, GCP's equivalent of an AWS or Azure NAT Gateway. Cloud NAT requires a Cloud Router in the same region and provides zero inbound connectivity capability - a resource behind it remains completely unreachable from the internet.

Cloud Run

A fully managed, serverless platform for running containerized applications, scaling automatically to zero when idle and billing only for actual invocations. It combines the flexibility of running any container with zero OS management overhead, making it the right default for a standard stateless containerized web service.

Cloud SQL

Google Cloud's managed relational database service, supporting MySQL, PostgreSQL, and SQL Server. It is the right default for a traditional relational workload with moderate scale needs, handling patching, backups, and replication automatically without requiring OS-level access to the underlying database server.

Cloud Spanner

A globally distributed, strongly consistent relational database offering both SQL semantics and horizontal scalability beyond what a single Cloud SQL instance can provide. It is the right choice specifically when a workload needs both relational query capability and massive scale simultaneously, not simply because a workload feels large or important.

Cloud Storage Bucket

The top-level container for objects in Google Cloud Storage, GCP's object storage service equivalent to AWS S3 or Azure Blob Storage. Every object uploaded to Cloud Storage must belong to exactly one bucket, and bucket-level settings control default storage class, location, and access permissions for everything inside it.

Cloud VPN (GCP)

A service connecting an on-premises network to a GCP VPC over an encrypted tunnel, similar to VPN Gateway in Azure or a Site-to-Site VPN in AWS. It requires a Cloud Router as its foundation and is distinct from VPC Network Peering, which connects two GCP VPCs directly rather than an external network.

Compute Engine

Google Cloud's Infrastructure-as-a-Service offering, providing full virtual machines with complete OS-level control. It is the right choice specifically when a workload needs custom software, specific OS configurations, or hardware-level access that GKE and Cloud Run don't expose.

Error Reporting (GCP)

A Google Cloud operations service that automatically aggregates and groups application errors from logs and instrumented code, surfacing recurring error patterns and their frequency without requiring manual log searching for each individual occurrence.

Firestore

A serverless NoSQL document database built for mobile and web application data, flexible schemas, and real-time client synchronization. It is not a strong fit for workloads needing complex relational joins or multi-row transactions across unrelated entities, which point instead to Cloud SQL or Spanner.

GCP Audit Log

A specific category of Cloud Logging entry recording administrative actions and data access. Admin Activity logs (configuration changes) are always enabled and cannot be disabled; Data Access logs (who read or wrote actual data) are disabled by default for most services and must be explicitly enabled before that history exists.

GCP Firewall Rule

A rule defined at the VPC level controlling inbound or outbound traffic, applied using network tags or Service Accounts on instances rather than being tied to a specific subnet boundary. A firewall rule only affects VMs carrying the exact matching tag or identity - a VM without it is entirely unaffected, silently, with no error to flag the mismatch.

GCP Folder

An optional grouping node in the GCP resource hierarchy, sitting between the Organization and individual Projects. Folders are commonly used to group Projects by team, department, or environment, letting IAM permissions and Organization Policies be applied once at the Folder level rather than repeated across every Project inside it.

GCP IAM Role Types

Google Cloud IAM roles come in three types: Basic roles (Owner, Editor, Viewer) which are broad and legacy; Predefined roles curated by Google and scoped to a specific service's specific needs; and Custom roles built from individual permissions for access needs a predefined role doesn't precisely match.

GCP Organization

The top-level node in Google Cloud's resource hierarchy, tied to a Google Workspace or Cloud Identity domain. An Organization contains Folders and Projects, and policies or IAM permissions set at this level automatically apply to everything beneath it across the entire company.

GCP Project

The fundamental unit of organization in Google Cloud - every resource belongs to exactly one Project, which serves as the boundary for billing, IAM permissions, and API enablement. A Project has a permanent, globally unique Project ID chosen at creation and never changeable, distinct from its display name which can be updated freely.

GCP Quota

A default limit on how much of a specific resource (like VM CPUs or API calls) can be used per Project per region, designed to prevent runaway costs and protect shared infrastructure. Quotas can be checked and increased through the Console ahead of a known future need, such as a planned launch event.

GCP Storage Class

A pricing and access tier for Cloud Storage objects - Standard, Nearline, Coldline, or Archive - trading lower storage cost against higher retrieval cost and a longer minimum storage duration. Deleting or overwriting an object before its class's minimum duration elapses still bills as if it were stored the full minimum period.

GKE Autopilot

A mode of Google Kubernetes Engine where Google manages node provisioning, sizing, and patching entirely - you define workloads, and nodes appear automatically to run them, billed per pod resource request rather than per node. It is the lower-effort default for most containerized workloads, versus Standard mode's direct node pool control.

Managed Instance Group

A group of identical Compute Engine VMs created from an Instance Template and managed as a single unit, GCP's equivalent of an AWS Auto Scaling Group or Azure VM Scale Set. A Managed Instance Group can autoscale based on a metric like CPU utilization, adding or removing instances automatically within configured minimum and maximum bounds.

Network Tag (GCP)

A freeform string label attached to a Compute Engine VM at creation time, used to target firewall rules to specific VMs. Network tags are simple but can be forgotten or misspelled at scale, which is why Service Account-based targeting is often preferred for anything security-critical in a growing fleet.

OS Login

A Compute Engine feature that ties SSH access to a user's actual Google Identity and IAM permissions, rather than manually managing individual SSH public keys per VM. Enabling OS Login means revoking a departing employee's GCP access also immediately revokes their SSH access to every VM, with no separate key cleanup needed.

Organization Policy (GCP)

A constraint applied at the Organization, Folder, or Project level that restricts what resources are allowed to look like across everything beneath that level - for example, blocking external IP addresses on VMs or restricting allowed regions. Unlike IAM, which controls who can act, Organization Policy controls what configurations are permitted.

Persistent Disk (GCP)

The block storage attached to a Compute Engine VM, equivalent to an AWS EBS volume or Azure Managed Disk. Persistent Disks come in zonal (replicated within one zone, lower cost) and regional (synchronously replicated across two zones, survives a single zone failure) variants.

Security Command Center

Google Cloud's security posture management service, continuously scanning deployed resources against security best practices and surfacing specific, severity-ranked findings - like a publicly accessible storage bucket - rather than a single opaque score. Findings should be prioritized by actual severity, not treated as a count to minimize.

Service Account (GCP)

An identity used by applications, VMs, and automated processes to authenticate to Google Cloud APIs, rather than by a human user. A Service Account is itself also a resource with its own IAM permissions, and can be attached directly to a VM or impersonated temporarily, both safer alternatives to downloading its long-lived JSON key file.

Shared VPC

A configuration where one Project (the host) owns a VPC network, while other Projects (service projects) attach to it and deploy resources using its subnets. This centralizes network administration in the host Project while letting individual teams in service Projects still manage their own resources independently.

Spot VM (GCP)

A Compute Engine VM using Google's spare, unused capacity at a steep discount, with the trade-off that Google can reclaim the instance with short notice. Spot VMs are appropriate only for fault-tolerant, interruption-tolerant workloads like batch processing, never for customer-facing services needing reliable availability.

VPC (GCP)

A Virtual Private Cloud in Google Cloud is a global resource by default, unlike the regional VPC or VNet model in AWS or Azure - a single GCP VPC can span every region, with individual subnets defined per-region inside it. The default auto-mode VPC pre-creates a subnet in every region, while a custom-mode VPC starts empty until subnets are deliberately defined.

VPC Network Peering

A direct connection between two separate VPCs, potentially in different Projects or Organizations, with traffic never touching the public internet. Peering must be established from both sides to become active, and is explicitly non-transitive - if VPC A peers with B, and B peers with C, A cannot reach C without its own direct peering connection to C.

Explore More in GCP Networking and Connectivity

All 6 Topics

Frequently Asked Questions

Is Choosing the Right Cloud Load Balancer Type free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the Choosing the Right Cloud Load Balancer Type topic cover?

Learn to distinguish GCP's global external, regional external, and internal load balancer types based on traffic type and reach needed.