Skip to main content

Deploying Serverless Containers with Cloud Run

Learn to deploy a container to Cloud Run, configure traffic splitting for gradual rollouts, and set scaling parameters correctly.

38 Terms

Overview and What You Will Learn

In this lab, you will deploy a container to Cloud Run, deploy a second revision, split traffic between the two revisions for a gradual rollout, and configure scaling parameters to control cost and cold-start behavior.

Why This Matters in Production

A team deploys a new version of their API directly to 100% of production traffic, discovers a bug affecting live requests within minutes, and has to scramble to roll back while customers are actively affected. Cloud Run's traffic splitting lets that same team send just 10% of traffic to the new revision first, catching the issue with limited blast radius before it ever reaches most customers.

Core Principles

Cloud Run runs a container as a fully managed, serverless service - each deployment creates a new Revision, and traffic can be split between revisions by percentage, enabling gradual rollouts without any separate infrastructure to manage.

◈ DIAGRAM
+------------------------------------------+
| Cloud Run Service |
| |
| Revision 1 (previous version) - 90% traffic|
| Revision 2 (new version) - 10% traffic|
+------------------------------------------+

Unlike a traditional deployment where "the new version" simply replaces "the old version," Cloud Run keeps both revisions available simultaneously, letting traffic be shifted gradually and rolled back instantly by simply adjusting the percentage split back to the previous revision.

Detailed Step-by-Step Practical Lab

  1. Create a project and enable Cloud Run:
Bash
gcloud projects create gcp-cloudrun-lab-2026 --name="Cloud Run Lab"
gcloud config set project gcp-cloudrun-lab-2026
gcloud services enable run.googleapis.com
  1. Deploy an initial version of a service:
Bash
gcloud run deploy inventory-service \
--image=gcr.io/google-samples/hello-app:1.0 \
--region=asia-south1 \
--platform=managed \
--allow-unauthenticated
  1. Deploy a second revision without shifting any traffic to it yet:
Bash
gcloud run deploy inventory-service \
--image=gcr.io/google-samples/hello-app:2.0 \
--region=asia-south1 \
--no-traffic
  1. List the current revisions and confirm the new one is deployed but receiving zero traffic:
Bash
gcloud run revisions list \
--service=inventory-service \
--region=asia-south1
  1. Gradually shift 10% of traffic to the new revision:
Bash
gcloud run services update-traffic inventory-service \
--region=asia-south1 \
--to-revisions=inventory-service-00002-abc=10
  1. After confirming the new revision behaves correctly, shift all traffic to it:
Bash
gcloud run services update-traffic inventory-service \
--region=asia-south1 \
--to-latest
  1. Configure scaling parameters - a minimum instance count to avoid cold starts for a latency-sensitive service, and a maximum to cap cost:
Bash
gcloud run services update inventory-service \
--region=asia-south1 \
--min-instances=1 \
--max-instances=20
Note

--min-instances=1 keeps at least one instance warm at all times, eliminating cold-start latency for the first request after a quiet period, at the cost of paying for that one instance continuously rather than scaling fully to zero. This is a deliberate trade-off for latency-sensitive services, not the default for every workload.

  1. Clean up:
Bash
gcloud run services delete inventory-service --region=asia-south1 --quiet
gcloud projects delete gcp-cloudrun-lab-2026 --quiet

Production Best Practices & Common Pitfalls

Common Mistake

Deploying a new revision directly to 100% of traffic with no gradual rollout, discovering a critical bug only after customers are already affected. Cloud Run's traffic splitting costs nothing extra and gives a real, live testing path with limited blast radius before a full cutover.

Tip

Set --min-instances specifically for latency-sensitive, user-facing services where a cold start would be noticeable, and leave it at the default (scaling to zero) for internal or infrequently-used services where occasional cold-start latency is an acceptable trade-off for not paying for idle capacity.

  • Rolling back is as simple as shifting traffic back to the previous revision - since Cloud Run keeps prior revisions available (up to a retention limit), a bad deployment can be reversed in seconds without needing to rebuild or redeploy the old version from source.
  • --allow-unauthenticated makes a service publicly reachable with no authentication at all. For internal services, omit this flag and rely on IAM-based invoker permissions instead, restricting who can call the service to specific identities.

Quick Reference & Troubleshooting Commands

Command Description
gcloud run deploy Deploy a container as a new Cloud Run revision
gcloud run services update-traffic --to-revisions= Split traffic between revisions by percentage
gcloud run revisions list List all revisions for a service
gcloud run services update --min-instances= Set a minimum warm instance count

Resources

Artifact Registry

Google Cloud's current, unified package and container image storage service, the successor to the older Container Registry. It supports container images alongside language packages like npm, Maven, and Python packages in one service, and supports cleanup policies to automatically remove old, untagged image versions.

BigQuery

Google Cloud's serverless data warehouse for analytics, priced based on the volume of data a query actually scans rather than pre-provisioned compute capacity. Selecting only needed columns and querying a properly partitioned table's relevant date range are the two highest-impact levers for controlling BigQuery cost.

Bigtable (GCP)

A wide-column NoSQL database built for very high throughput at extremely low latency, suited to time-series data, IoT telemetry, and analytics workloads with massive write volume. It is a specialized tool distinct from Firestore's document model or Cloud SQL's relational model.

Billing Account (GCP)

A separate object from a Project that holds the actual payment method for Google Cloud usage. One Billing Account can be linked to many Projects, all of whose usage is billed to that same account, letting a company centralize payment while individual Projects retain their own separate IAM control over resources.

Cloud Armor

A security policy service attached to a Global External HTTP(S) Load Balancer, filtering malicious requests before they reach the backend. It supports rate-based rules blocking a single source IP exceeding a request threshold, and preconfigured WAF rules blocking common attack patterns like SQL injection, Google's equivalent of AWS WAF or Azure WAF.

Cloud Functions (GCP)

Google Cloud's event-driven, single-purpose compute service, the most granular and shortest-lived compute unit in GCP's compute spectrum. It fits a narrow niche compared to Cloud Run - a specific triggered action like resizing an uploaded image, rather than a full application with multiple routes and business logic.

Cloud Load Balancing

Google Cloud's family of load balancing products, spanning Global External HTTP(S) Load Balancers (public, distributed globally via a single IP), Regional External Load Balancers, and Internal Load Balancers reachable only from within the VPC. The right type depends on traffic type and whether the backend should ever be reachable from the public internet.

Cloud Logging

Google Cloud's service for collecting, storing, and querying log data from every GCP service and custom applications. Its query filter syntax lets logs be narrowed by resource, severity, and specific field values, and Log Sinks can route matching logs to BigQuery, Cloud Storage, or Pub/Sub for long-term retention beyond the default window.

Cloud Monitoring

Google Cloud's metrics, dashboards, and alerting service. An alert policy requires a Notification Channel explicitly attached to actually notify someone when it fires - without one, the policy still evaluates and logs its own firing history, but nobody is ever told, a common and easy-to-miss misconfiguration.

Cloud NAT

A service providing outbound-only internet connectivity for resources with no external IP address, GCP's equivalent of an AWS or Azure NAT Gateway. Cloud NAT requires a Cloud Router in the same region and provides zero inbound connectivity capability - a resource behind it remains completely unreachable from the internet.

Cloud Run

A fully managed, serverless platform for running containerized applications, scaling automatically to zero when idle and billing only for actual invocations. It combines the flexibility of running any container with zero OS management overhead, making it the right default for a standard stateless containerized web service.

Cloud SQL

Google Cloud's managed relational database service, supporting MySQL, PostgreSQL, and SQL Server. It is the right default for a traditional relational workload with moderate scale needs, handling patching, backups, and replication automatically without requiring OS-level access to the underlying database server.

Cloud Spanner

A globally distributed, strongly consistent relational database offering both SQL semantics and horizontal scalability beyond what a single Cloud SQL instance can provide. It is the right choice specifically when a workload needs both relational query capability and massive scale simultaneously, not simply because a workload feels large or important.

Cloud Storage Bucket

The top-level container for objects in Google Cloud Storage, GCP's object storage service equivalent to AWS S3 or Azure Blob Storage. Every object uploaded to Cloud Storage must belong to exactly one bucket, and bucket-level settings control default storage class, location, and access permissions for everything inside it.

Cloud VPN (GCP)

A service connecting an on-premises network to a GCP VPC over an encrypted tunnel, similar to VPN Gateway in Azure or a Site-to-Site VPN in AWS. It requires a Cloud Router as its foundation and is distinct from VPC Network Peering, which connects two GCP VPCs directly rather than an external network.

Compute Engine

Google Cloud's Infrastructure-as-a-Service offering, providing full virtual machines with complete OS-level control. It is the right choice specifically when a workload needs custom software, specific OS configurations, or hardware-level access that GKE and Cloud Run don't expose.

Error Reporting (GCP)

A Google Cloud operations service that automatically aggregates and groups application errors from logs and instrumented code, surfacing recurring error patterns and their frequency without requiring manual log searching for each individual occurrence.

Firestore

A serverless NoSQL document database built for mobile and web application data, flexible schemas, and real-time client synchronization. It is not a strong fit for workloads needing complex relational joins or multi-row transactions across unrelated entities, which point instead to Cloud SQL or Spanner.

GCP Audit Log

A specific category of Cloud Logging entry recording administrative actions and data access. Admin Activity logs (configuration changes) are always enabled and cannot be disabled; Data Access logs (who read or wrote actual data) are disabled by default for most services and must be explicitly enabled before that history exists.

GCP Firewall Rule

A rule defined at the VPC level controlling inbound or outbound traffic, applied using network tags or Service Accounts on instances rather than being tied to a specific subnet boundary. A firewall rule only affects VMs carrying the exact matching tag or identity - a VM without it is entirely unaffected, silently, with no error to flag the mismatch.

GCP Folder

An optional grouping node in the GCP resource hierarchy, sitting between the Organization and individual Projects. Folders are commonly used to group Projects by team, department, or environment, letting IAM permissions and Organization Policies be applied once at the Folder level rather than repeated across every Project inside it.

GCP IAM Role Types

Google Cloud IAM roles come in three types: Basic roles (Owner, Editor, Viewer) which are broad and legacy; Predefined roles curated by Google and scoped to a specific service's specific needs; and Custom roles built from individual permissions for access needs a predefined role doesn't precisely match.

GCP Organization

The top-level node in Google Cloud's resource hierarchy, tied to a Google Workspace or Cloud Identity domain. An Organization contains Folders and Projects, and policies or IAM permissions set at this level automatically apply to everything beneath it across the entire company.

GCP Project

The fundamental unit of organization in Google Cloud - every resource belongs to exactly one Project, which serves as the boundary for billing, IAM permissions, and API enablement. A Project has a permanent, globally unique Project ID chosen at creation and never changeable, distinct from its display name which can be updated freely.

GCP Quota

A default limit on how much of a specific resource (like VM CPUs or API calls) can be used per Project per region, designed to prevent runaway costs and protect shared infrastructure. Quotas can be checked and increased through the Console ahead of a known future need, such as a planned launch event.

GCP Storage Class

A pricing and access tier for Cloud Storage objects - Standard, Nearline, Coldline, or Archive - trading lower storage cost against higher retrieval cost and a longer minimum storage duration. Deleting or overwriting an object before its class's minimum duration elapses still bills as if it were stored the full minimum period.

GKE Autopilot

A mode of Google Kubernetes Engine where Google manages node provisioning, sizing, and patching entirely - you define workloads, and nodes appear automatically to run them, billed per pod resource request rather than per node. It is the lower-effort default for most containerized workloads, versus Standard mode's direct node pool control.

Managed Instance Group

A group of identical Compute Engine VMs created from an Instance Template and managed as a single unit, GCP's equivalent of an AWS Auto Scaling Group or Azure VM Scale Set. A Managed Instance Group can autoscale based on a metric like CPU utilization, adding or removing instances automatically within configured minimum and maximum bounds.

Network Tag (GCP)

A freeform string label attached to a Compute Engine VM at creation time, used to target firewall rules to specific VMs. Network tags are simple but can be forgotten or misspelled at scale, which is why Service Account-based targeting is often preferred for anything security-critical in a growing fleet.

OS Login

A Compute Engine feature that ties SSH access to a user's actual Google Identity and IAM permissions, rather than manually managing individual SSH public keys per VM. Enabling OS Login means revoking a departing employee's GCP access also immediately revokes their SSH access to every VM, with no separate key cleanup needed.

Organization Policy (GCP)

A constraint applied at the Organization, Folder, or Project level that restricts what resources are allowed to look like across everything beneath that level - for example, blocking external IP addresses on VMs or restricting allowed regions. Unlike IAM, which controls who can act, Organization Policy controls what configurations are permitted.

Persistent Disk (GCP)

The block storage attached to a Compute Engine VM, equivalent to an AWS EBS volume or Azure Managed Disk. Persistent Disks come in zonal (replicated within one zone, lower cost) and regional (synchronously replicated across two zones, survives a single zone failure) variants.

Security Command Center

Google Cloud's security posture management service, continuously scanning deployed resources against security best practices and surfacing specific, severity-ranked findings - like a publicly accessible storage bucket - rather than a single opaque score. Findings should be prioritized by actual severity, not treated as a count to minimize.

Service Account (GCP)

An identity used by applications, VMs, and automated processes to authenticate to Google Cloud APIs, rather than by a human user. A Service Account is itself also a resource with its own IAM permissions, and can be attached directly to a VM or impersonated temporarily, both safer alternatives to downloading its long-lived JSON key file.

Shared VPC

A configuration where one Project (the host) owns a VPC network, while other Projects (service projects) attach to it and deploy resources using its subnets. This centralizes network administration in the host Project while letting individual teams in service Projects still manage their own resources independently.

Spot VM (GCP)

A Compute Engine VM using Google's spare, unused capacity at a steep discount, with the trade-off that Google can reclaim the instance with short notice. Spot VMs are appropriate only for fault-tolerant, interruption-tolerant workloads like batch processing, never for customer-facing services needing reliable availability.

VPC (GCP)

A Virtual Private Cloud in Google Cloud is a global resource by default, unlike the regional VPC or VNet model in AWS or Azure - a single GCP VPC can span every region, with individual subnets defined per-region inside it. The default auto-mode VPC pre-creates a subnet in every region, while a custom-mode VPC starts empty until subnets are deliberately defined.

VPC Network Peering

A direct connection between two separate VPCs, potentially in different Projects or Organizations, with traffic never touching the public internet. Peering must be established from both sides to become active, and is explicitly non-transitive - if VPC A peers with B, and B peers with C, A cannot reach C without its own direct peering connection to C.

Explore More in GCP Compute and Container Services

All 6 Topics

Frequently Asked Questions

Is Deploying Serverless Containers with Cloud Run free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the Deploying Serverless Containers with Cloud Run topic cover?

Learn to deploy a container to Cloud Run, configure traffic splitting for gradual rollouts, and set scaling parameters correctly.