Skip to main content

Querying and Routing Logs with Cloud Logging

Learn Cloud Logging's query filter syntax to actually find an incident's root cause, and how to route logs to BigQuery for long-term analysis.

38 Terms

Overview and What You Will Learn

In this lab, you will generate log entries from a VM, write increasingly specific query filters to narrow down to exactly the relevant entries, and configure a log sink routing specific logs to BigQuery for long-term, queryable retention.

Why This Matters in Production

An engineer is told "checkout is slow sometimes," with no more specific information than that. Without the ability to actually filter logs - by resource, by severity, by a specific time window - finding the actual root cause becomes a matter of scrolling through a huge volume of raw text manually. Cloud Logging's filter syntax turns that into a precise, repeatable question asked directly against the data.

Core Principles

Cloud Logging collects log entries from every GCP service and your own applications into a searchable, filterable store - and its query filter syntax is the tool that turns "something is wrong" into "here is the exact log entry explaining why."

◈ DIAGRAM
+------------------------------------------+
| Resources emit log entries continuously |
+------------------------------------------+
|
v
+------------------------------------------+
| Cloud Logging stores and indexes them |
+------------------------------------------+
|
v
+------------------------------------------+
| Query filters narrow down to exactly the |
| relevant entries by resource, severity, |
| time range, or specific field values |
+------------------------------------------+
|
v
+------------------------------------------+
| Log Sinks can route matching entries to |
| BigQuery, Cloud Storage, or Pub/Sub for |
| long-term retention or further processing |
+------------------------------------------+

Detailed Step-by-Step Practical Lab

  1. Create a project and a VM that will generate some log activity:
Bash
gcloud projects create gcp-logging-lab-2026 --name="Cloud Logging Lab"
gcloud config set project gcp-logging-lab-2026
gcloud services enable compute.googleapis.com logging.googleapis.com
gcloud compute instances create vm-logging-test \
--zone=asia-south1-a \
--machine-type=e2-small \
--image-family=debian-12 \
--image-project=debian-cloud
  1. Start with the simplest possible query - view recent logs for this specific VM:
Bash
gcloud logging read \
'resource.type="gce_instance" AND resource.labels.instance_id="'$(gcloud compute instances describe vm-logging-test --zone=asia-south1-a --format="value(id)")'"' \
--limit=10
  1. Narrow the query to only error-severity entries from the last hour:
Bash
gcloud logging read \
'severity>=ERROR' \
--freshness=1h
  1. Combine multiple conditions - error-severity entries specifically for Compute Engine resources:
Bash
gcloud logging read \
'resource.type="gce_instance" AND severity>=ERROR' \
--freshness=24h \
--limit=20
  1. Create a BigQuery dataset to serve as a log sink destination for long-term retention:
Bash
bq mk --dataset --location=asia-south1 gcp-logging-lab-2026:log_archive
  1. Create a Log Sink routing error-severity logs to that BigQuery dataset:
Bash
gcloud logging sinks create error-logs-to-bigquery \
bigquery.googleapis.com/projects/gcp-logging-lab-2026/datasets/log_archive \
--log-filter='severity>=ERROR'
  1. Grant the sink's service account permission to write to the BigQuery dataset:
Bash
SINK_SA=$(gcloud logging sinks describe error-logs-to-bigquery --format="value(writerIdentity)")
bq add-iam-policy-binding \
--member="$SINK_SA" \
--role="roles/bigquery.dataEditor" \
gcp-logging-lab-2026:log_archive
Note

Every Log Sink has its own dedicated service account (the writerIdentity) used specifically to write matching logs to the destination - this identity must be explicitly granted write permission on the destination, since the sink itself doesn't automatically have access just by existing.

  1. Clean up:
Bash
gcloud logging sinks delete error-logs-to-bigquery --quiet
bq rm -r -f -d gcp-logging-lab-2026:log_archive
gcloud compute instances delete vm-logging-test --zone=asia-south1-a --quiet
gcloud projects delete gcp-logging-lab-2026 --quiet

Production Best Practices & Common Pitfalls

Common Mistake

Relying solely on Cloud Logging's default retention window for logs that genuinely need long-term retention for compliance or historical analysis. Cloud Logging's own storage has a default retention period - a Log Sink routing relevant logs to BigQuery or Cloud Storage is necessary for retention beyond that default window.

Tip

Build query filters incrementally - start broad, confirm results look reasonable, then add conditions one at a time (resource type, severity, specific field values) rather than trying to write the complete, precise filter in one attempt.

  • Log Sinks only route logs matching their filter going forward from creation, not retroactively. Historical logs that existed before the sink was created are not backfilled into the new destination automatically.
  • Audit Logs are a specific category within Cloud Logging worth filtering for separately during a security investigation - logName:"cloudaudit.googleapis.com" narrows a query specifically to administrative and data-access audit entries, distinct from regular application or system logs.

Quick Reference & Troubleshooting Commands

Command Description
gcloud logging read Query and filter log entries
gcloud logging sinks create Route matching logs to an external destination
gcloud logging sinks describe --format="value(writerIdentity)" Get a sink's service account for permission grants
gcloud logging read --freshness= Limit query results to a recent time window

Resources

Artifact Registry

Google Cloud's current, unified package and container image storage service, the successor to the older Container Registry. It supports container images alongside language packages like npm, Maven, and Python packages in one service, and supports cleanup policies to automatically remove old, untagged image versions.

BigQuery

Google Cloud's serverless data warehouse for analytics, priced based on the volume of data a query actually scans rather than pre-provisioned compute capacity. Selecting only needed columns and querying a properly partitioned table's relevant date range are the two highest-impact levers for controlling BigQuery cost.

Bigtable (GCP)

A wide-column NoSQL database built for very high throughput at extremely low latency, suited to time-series data, IoT telemetry, and analytics workloads with massive write volume. It is a specialized tool distinct from Firestore's document model or Cloud SQL's relational model.

Billing Account (GCP)

A separate object from a Project that holds the actual payment method for Google Cloud usage. One Billing Account can be linked to many Projects, all of whose usage is billed to that same account, letting a company centralize payment while individual Projects retain their own separate IAM control over resources.

Cloud Armor

A security policy service attached to a Global External HTTP(S) Load Balancer, filtering malicious requests before they reach the backend. It supports rate-based rules blocking a single source IP exceeding a request threshold, and preconfigured WAF rules blocking common attack patterns like SQL injection, Google's equivalent of AWS WAF or Azure WAF.

Cloud Functions (GCP)

Google Cloud's event-driven, single-purpose compute service, the most granular and shortest-lived compute unit in GCP's compute spectrum. It fits a narrow niche compared to Cloud Run - a specific triggered action like resizing an uploaded image, rather than a full application with multiple routes and business logic.

Cloud Load Balancing

Google Cloud's family of load balancing products, spanning Global External HTTP(S) Load Balancers (public, distributed globally via a single IP), Regional External Load Balancers, and Internal Load Balancers reachable only from within the VPC. The right type depends on traffic type and whether the backend should ever be reachable from the public internet.

Cloud Logging

Google Cloud's service for collecting, storing, and querying log data from every GCP service and custom applications. Its query filter syntax lets logs be narrowed by resource, severity, and specific field values, and Log Sinks can route matching logs to BigQuery, Cloud Storage, or Pub/Sub for long-term retention beyond the default window.

Cloud Monitoring

Google Cloud's metrics, dashboards, and alerting service. An alert policy requires a Notification Channel explicitly attached to actually notify someone when it fires - without one, the policy still evaluates and logs its own firing history, but nobody is ever told, a common and easy-to-miss misconfiguration.

Cloud NAT

A service providing outbound-only internet connectivity for resources with no external IP address, GCP's equivalent of an AWS or Azure NAT Gateway. Cloud NAT requires a Cloud Router in the same region and provides zero inbound connectivity capability - a resource behind it remains completely unreachable from the internet.

Cloud Run

A fully managed, serverless platform for running containerized applications, scaling automatically to zero when idle and billing only for actual invocations. It combines the flexibility of running any container with zero OS management overhead, making it the right default for a standard stateless containerized web service.

Cloud SQL

Google Cloud's managed relational database service, supporting MySQL, PostgreSQL, and SQL Server. It is the right default for a traditional relational workload with moderate scale needs, handling patching, backups, and replication automatically without requiring OS-level access to the underlying database server.

Cloud Spanner

A globally distributed, strongly consistent relational database offering both SQL semantics and horizontal scalability beyond what a single Cloud SQL instance can provide. It is the right choice specifically when a workload needs both relational query capability and massive scale simultaneously, not simply because a workload feels large or important.

Cloud Storage Bucket

The top-level container for objects in Google Cloud Storage, GCP's object storage service equivalent to AWS S3 or Azure Blob Storage. Every object uploaded to Cloud Storage must belong to exactly one bucket, and bucket-level settings control default storage class, location, and access permissions for everything inside it.

Cloud VPN (GCP)

A service connecting an on-premises network to a GCP VPC over an encrypted tunnel, similar to VPN Gateway in Azure or a Site-to-Site VPN in AWS. It requires a Cloud Router as its foundation and is distinct from VPC Network Peering, which connects two GCP VPCs directly rather than an external network.

Compute Engine

Google Cloud's Infrastructure-as-a-Service offering, providing full virtual machines with complete OS-level control. It is the right choice specifically when a workload needs custom software, specific OS configurations, or hardware-level access that GKE and Cloud Run don't expose.

Error Reporting (GCP)

A Google Cloud operations service that automatically aggregates and groups application errors from logs and instrumented code, surfacing recurring error patterns and their frequency without requiring manual log searching for each individual occurrence.

Firestore

A serverless NoSQL document database built for mobile and web application data, flexible schemas, and real-time client synchronization. It is not a strong fit for workloads needing complex relational joins or multi-row transactions across unrelated entities, which point instead to Cloud SQL or Spanner.

GCP Audit Log

A specific category of Cloud Logging entry recording administrative actions and data access. Admin Activity logs (configuration changes) are always enabled and cannot be disabled; Data Access logs (who read or wrote actual data) are disabled by default for most services and must be explicitly enabled before that history exists.

GCP Firewall Rule

A rule defined at the VPC level controlling inbound or outbound traffic, applied using network tags or Service Accounts on instances rather than being tied to a specific subnet boundary. A firewall rule only affects VMs carrying the exact matching tag or identity - a VM without it is entirely unaffected, silently, with no error to flag the mismatch.

GCP Folder

An optional grouping node in the GCP resource hierarchy, sitting between the Organization and individual Projects. Folders are commonly used to group Projects by team, department, or environment, letting IAM permissions and Organization Policies be applied once at the Folder level rather than repeated across every Project inside it.

GCP IAM Role Types

Google Cloud IAM roles come in three types: Basic roles (Owner, Editor, Viewer) which are broad and legacy; Predefined roles curated by Google and scoped to a specific service's specific needs; and Custom roles built from individual permissions for access needs a predefined role doesn't precisely match.

GCP Organization

The top-level node in Google Cloud's resource hierarchy, tied to a Google Workspace or Cloud Identity domain. An Organization contains Folders and Projects, and policies or IAM permissions set at this level automatically apply to everything beneath it across the entire company.

GCP Project

The fundamental unit of organization in Google Cloud - every resource belongs to exactly one Project, which serves as the boundary for billing, IAM permissions, and API enablement. A Project has a permanent, globally unique Project ID chosen at creation and never changeable, distinct from its display name which can be updated freely.

GCP Quota

A default limit on how much of a specific resource (like VM CPUs or API calls) can be used per Project per region, designed to prevent runaway costs and protect shared infrastructure. Quotas can be checked and increased through the Console ahead of a known future need, such as a planned launch event.

GCP Storage Class

A pricing and access tier for Cloud Storage objects - Standard, Nearline, Coldline, or Archive - trading lower storage cost against higher retrieval cost and a longer minimum storage duration. Deleting or overwriting an object before its class's minimum duration elapses still bills as if it were stored the full minimum period.

GKE Autopilot

A mode of Google Kubernetes Engine where Google manages node provisioning, sizing, and patching entirely - you define workloads, and nodes appear automatically to run them, billed per pod resource request rather than per node. It is the lower-effort default for most containerized workloads, versus Standard mode's direct node pool control.

Managed Instance Group

A group of identical Compute Engine VMs created from an Instance Template and managed as a single unit, GCP's equivalent of an AWS Auto Scaling Group or Azure VM Scale Set. A Managed Instance Group can autoscale based on a metric like CPU utilization, adding or removing instances automatically within configured minimum and maximum bounds.

Network Tag (GCP)

A freeform string label attached to a Compute Engine VM at creation time, used to target firewall rules to specific VMs. Network tags are simple but can be forgotten or misspelled at scale, which is why Service Account-based targeting is often preferred for anything security-critical in a growing fleet.

OS Login

A Compute Engine feature that ties SSH access to a user's actual Google Identity and IAM permissions, rather than manually managing individual SSH public keys per VM. Enabling OS Login means revoking a departing employee's GCP access also immediately revokes their SSH access to every VM, with no separate key cleanup needed.

Organization Policy (GCP)

A constraint applied at the Organization, Folder, or Project level that restricts what resources are allowed to look like across everything beneath that level - for example, blocking external IP addresses on VMs or restricting allowed regions. Unlike IAM, which controls who can act, Organization Policy controls what configurations are permitted.

Persistent Disk (GCP)

The block storage attached to a Compute Engine VM, equivalent to an AWS EBS volume or Azure Managed Disk. Persistent Disks come in zonal (replicated within one zone, lower cost) and regional (synchronously replicated across two zones, survives a single zone failure) variants.

Security Command Center

Google Cloud's security posture management service, continuously scanning deployed resources against security best practices and surfacing specific, severity-ranked findings - like a publicly accessible storage bucket - rather than a single opaque score. Findings should be prioritized by actual severity, not treated as a count to minimize.

Service Account (GCP)

An identity used by applications, VMs, and automated processes to authenticate to Google Cloud APIs, rather than by a human user. A Service Account is itself also a resource with its own IAM permissions, and can be attached directly to a VM or impersonated temporarily, both safer alternatives to downloading its long-lived JSON key file.

Shared VPC

A configuration where one Project (the host) owns a VPC network, while other Projects (service projects) attach to it and deploy resources using its subnets. This centralizes network administration in the host Project while letting individual teams in service Projects still manage their own resources independently.

Spot VM (GCP)

A Compute Engine VM using Google's spare, unused capacity at a steep discount, with the trade-off that Google can reclaim the instance with short notice. Spot VMs are appropriate only for fault-tolerant, interruption-tolerant workloads like batch processing, never for customer-facing services needing reliable availability.

VPC (GCP)

A Virtual Private Cloud in Google Cloud is a global resource by default, unlike the regional VPC or VNet model in AWS or Azure - a single GCP VPC can span every region, with individual subnets defined per-region inside it. The default auto-mode VPC pre-creates a subnet in every region, while a custom-mode VPC starts empty until subnets are deliberately defined.

VPC Network Peering

A direct connection between two separate VPCs, potentially in different Projects or Organizations, with traffic never touching the public internet. Peering must be established from both sides to become active, and is explicitly non-transitive - if VPC A peers with B, and B peers with C, A cannot reach C without its own direct peering connection to C.

Explore More in GCP Operations, Security, and Production Readiness

All 6 Topics

Frequently Asked Questions

Is Querying and Routing Logs with Cloud Logging free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the Querying and Routing Logs with Cloud Logging topic cover?

Learn Cloud Logging's query filter syntax to actually find an incident's root cause, and how to route logs to BigQuery for long-term analysis.