Google Professional Data Engineer: Exam Objectives

The Professional Data Engineer blueprint is one data lifecycle. Architecture defines security, governance, reliability and migration. Ingestion and processing move and transform data. Storage preserves it according to access patterns. Analytics/AI preparation makes it useful. Automation and maintenance keep the system reliable after deployment. The current exam weights are approximately 22%, 25%, 20%, 15% and 18% across those stages.

Business and regulatory requirements sit above technical design

Data sovereignty, privacy, legal requirements, IAM and encryption can determine where data is stored and who can process it. These decisions belong before service selection.

A technically efficient architecture can still be wrong if it violates residency or access requirements.

Governance connects projects, datasets, tables and catalog

Project/dataset/table organization, Dataplex Catalog, discovery and profiling help teams know what data exists, who owns it and how it may be used.

Governance is therefore part of architecture and analytics, not a separate documentation exercise.

Sources, transformations and sinks form the pipeline core

A pipeline starts with source characteristics, defines batch/streaming transformation and ends at one or more sinks. Windowing, late data and event-time behavior matter in streaming systems.

A Dataflow pipeline and a Dataproc workload may process similar data but differ in framework, operational model and migration needs.

Storage should be chosen from access patterns

BigQuery fits analytical warehouse workloads, Cloud Storage object data, Cloud SQL relational transactional workloads, Spanner globally scalable relational systems and Bigtable high-throughput wide-column access patterns.

The Spanner and Bigtable distinction is easier when consistency/model/query requirements come first.

Warehouse and lake patterns can coexist

A lake stores diverse/raw data with discovery and governance; a warehouse supports curated analytical modeling and SQL access. BigLake and Dataplex can bridge governance or access across storage patterns.

The map should show data movement and governance rather than force every organization into one architecture.

Analytics preparation connects data engineering with consumers

BI Engine, materialized views, precalculated fields, query optimization, masking and IAM shape data prepared for analysts and dashboards.

The engineer needs to optimize both compute cost and user experience without weakening data protection.

ML and RAG are downstream data-consumer patterns

BigQuery ML needs curated features and data quality; embeddings/RAG need suitable unstructured documents, chunking/representation and governed retrieval inputs.

The data engineer’s role is preparing trustworthy, accessible data for those systems rather than becoming the model engineer.

Automation makes data systems repeatable

Cloud Composer DAGs, Workflows, CI/CD and job scheduling turn manual steps into predictable processes. Repeatability should include parameterization, environment separation and controlled deployment.

A successful pipeline is one that can be rerun or recovered without undocumented operator knowledge.

Observability and cost operate across every layer

Cloud Monitoring, Cloud Logging, BigQuery administrative views, quota/billing signals and job metrics show whether systems are healthy and economical.

Cost problems can originate in inefficient queries, idle clusters, excessive retention or poorly chosen storage, so optimization is cross-cutting.

Failure handling closes the data lifecycle

Data corruption, missing data, partial pipeline failure, regional outage or database failover should trigger recovery appropriate to the business requirement. Replication, restart behavior and multizone/multiregion execution are design choices rather than afterthoughts.

The map should also show data contracts between stages. A source schema, event format or business definition becomes an interface between producers and pipelines. If that contract changes silently, downstream jobs can fail or—worse—continue producing incorrect results. Schema management is therefore both reliability and governance.

Encryption and IAM should wrap every storage and pipeline component. Service accounts need only required access, keys need lifecycle management and sensitive datasets need appropriate boundaries. Security is not a separate edge firewall around the data platform.

Batch and streaming should be drawn as processing modes, not competing technologies. The same business domain can use streaming for current events and batch for backfills or reconciliation. An architecture can contain both when latency requirements differ.

Windowing belongs on the streaming branch because unbounded event streams still need finite analytical groupings. Fixed, sliding or session-style windows solve different questions, and late-arriving data complicates when results are considered complete.

Dataflow and Dataproc should be separated by execution model and ecosystem. Beam/Dataflow provides a unified managed data-processing model; Dataproc provides managed Spark/Hadoop-compatible clusters. Existing codebase and operational requirements can matter more than raw service feature count.

BigQuery and Bigtable sit at very different query layers. BigQuery is optimized for analytical SQL scans/aggregations; Bigtable serves low-latency key/range access at massive scale. Choosing one because both can hold large datasets ignores the access pattern.

Spanner and Cloud SQL form another useful contrast. Cloud SQL is managed relational database service for common engines and conventional scale; Spanner is designed for distributed relational workloads requiring strong consistency and horizontal/global scale. Operational complexity and cost should match the business need.

Memorystore should be drawn as an in-memory cache/store, not a durable system of record for every workload. Caching can reduce database load and latency but adds consistency/expiration considerations. The map should show where cached data is derived from authoritative storage.

Analytics consumers should be grouped by purpose: dashboards, ad-hoc analysts, ML training, RAG retrieval or downstream sharing. Each consumer can require different freshness, latency, masking and performance. One “gold table” is not necessarily optimal for every use case.

Data masking and Cloud DLP should connect privacy with analytics usability. Sensitive values can be detected or protected while allowing approved aggregate analysis. The control should be chosen from the actual privacy requirement rather than indiscriminately removing useful fields.

Composer and Workflows should sit on orchestration rather than transformation. They coordinate jobs and dependencies; the actual transformation may run in Dataflow, BigQuery, Dataproc or elsewhere. This separation helps diagnose whether a failure is orchestration or processing.

CI/CD should connect source control with pipeline definitions and environment promotion. A data pipeline is software and needs versioning, review, testing and controlled release. Manual console edits create drift that makes incidents difficult to reproduce.

Observability should include freshness as well as infrastructure health. A pipeline can be running but delivering yesterday’s data because an upstream source stalled. Business-level data freshness is often more important than CPU utilization in serverless managed pipelines.

Recovery should map to the failed artifact. A failed transform may be replayed, a corrupted table may be restored/rebuilt, a database outage may fail over and a missing source may require upstream remediation. “Restart the pipeline” is not a universal data-recovery strategy.

Use the completed map to debug from consumer backward: report wrong → curated table wrong → transform wrong → source or schema changed. This reverse lineage path is often faster than inspecting every service independently.

Migration should be drawn as a temporary but important lifecycle branch. Before steady-state pipelines exist, the engineer needs to move historical data, capture changes, validate counts and reconcile cutover. Migration tooling should feed the same governed storage model that production pipelines will use afterward.

Data sovereignty should appear near storage and processing location. A dataset might be encrypted and access-controlled yet still violate a residency obligation if replicated or processed in the wrong geography. Region selection can therefore be a compliance control.

Validation should sit between every major stage. Source checks, transformation checks, sink reconciliation and downstream quality metrics reduce the chance that bad data travels silently through the platform. A pipeline status of “success” is only one signal.

Lineage should connect analytics consumers back to transformations and sources. When a metric changes unexpectedly, lineage can show which upstream table or job contributed. This reduces troubleshooting time and supports audit or data-governance questions.

Cost controls should appear at both storage and compute. Lifecycle policies manage long-lived objects, partitioned analytical tables reduce scanned data, ephemeral clusters reduce idle compute and reservations can shape BigQuery capacity. Cost is an architectural dimension across the platform.

Data sharing should connect governance with consumer autonomy. Analytics Hub or governed shared datasets let producers expose authoritative data while consumers build their own analysis. This is often stronger than manually exporting copies that become stale and hard to revoke.

The map should also show operational ownership. Platform engineers may manage networking and IAM, data engineers manage pipelines and data models, analysts consume curated data and ML teams consume features/documents. Clear ownership helps incidents move to the right team quickly.

Use one end-to-end example: Pub/Sub events → Dataflow transformation → BigQuery warehouse → BI dashboard and BigQuery ML → Composer orchestration → Monitoring/freshness alerts. Add a private source, sensitive field and recovery requirement. That single path touches nearly every exam section.

Source systems should be shown with change rate and ownership. A frequently changing operational database may need CDC, while a static archive may need one-time bulk transfer. The same “ingestion” label hides very different freshness and migration semantics.

Staging belongs between source and transformation when raw data needs to be preserved for replay, validation or audit. Immutable raw storage can become a recovery point when downstream transformations fail or business logic changes.

Data quality should connect schema, completeness, uniqueness, timeliness and business-rule validation. Different consumers care about different dimensions. A model-training pipeline may tolerate some nulls that a financial report cannot.

Query generation with LLMs should be placed on the productivity layer, not treated as a replacement for schema knowledge or access control. Generated SQL still needs validation, cost awareness and authorization, especially when it can query sensitive datasets.

Federated governance should sit above distributed domains. Teams can own data products locally while shared policies, catalog standards and access principles maintain enterprise consistency. Central control of every transformation is not the only way to achieve governance.

The final map should show data moving and metadata moving together. Records flow through pipelines, while schema, lineage, ownership, quality and policy information should remain discoverable. A technically correct table with unknown provenance is a weak enterprise data product.

The modern data-engineer map is therefore requirements → governed data → pipeline → storage → analytics/AI → automation/monitoring → recovery and improvement.