BigQuery
Google Cloud's serverless data warehouse for analytics, priced based on the volume of data a query actually scans rather than pre-provisioned compute capacity. Selecting only needed columns and querying a properly partitioned table's relevant date range are the two highest-impact levers for controlling BigQuery cost.
Frequently Asked Questions
Why does BigQuery charge per query instead of per provisioned server, and what does that change about how you write queries?
BigQuery separates storage and compute entirely — you're billed for bytes scanned by a query (on-demand pricing) rather than for keeping a cluster running, which is why an unindexed `SELECT *` across a multi-terabyte table can generate a surprising bill even though it 'just' ran once. This flips the usual database instinct: instead of optimizing for query latency on a fixed-cost cluster, you optimize for bytes scanned, since that's the actual cost driver.
What's the single most common expensive mistake in BigQuery usage?
Running `SELECT *` against a large table instead of naming only the columns needed — BigQuery is a columnar store, so it only reads referenced columns, and `SELECT *` forces a full-table scan even if you display three fields downstream. The second most common mistake is querying a partitioned table without a `WHERE` clause on the partition column, which scans every partition instead of the relevant date range and can turn a query that should cost cents into one that costs dollars.