Google Professional Data Engineer: What to Practice Hands-On

The Professional Data Engineer exam is not a coding interview, but hands-on experience is critical because many questions describe operational behavior. A small lab can demonstrate ingestion, transformation, storage, orchestration, governance, query performance and recovery. The current PDE guide should control which exercises deserve time.

Lab one: build a governed BigQuery dataset

Create a small dataset, load harmless data, apply IAM at an appropriate scope and document the owner/classification. Use separate dev and prod-style datasets or projects conceptually.

A BigQuery lab makes governance, analytical storage and query behavior tangible at the same time.

Lab two: create one batch pipeline

Use Dataflow/Beam, Data Fusion or another current managed path to ingest, clean and write data to a sink. Record source schema, transformation rules and output validation.

Break one record or schema assumption and observe failure handling.

Lab three: create one streaming pipeline

Use Pub/Sub plus Dataflow or a training sandbox to process events. Experiment with windows and late-arriving data using synthetic messages.

Document event time versus processing time and how the pipeline handles duplicate or delayed input.

Lab four: compare storage systems

Take the same hypothetical application and map parts of it to BigQuery, Cloud SQL, Spanner, Bigtable and Cloud Storage. Build one small example in two systems if budget allows.

A Cloud SQL example and Bigtable comparison make relational versus wide-column behavior concrete.

Lab five: create a data lake/governance view

Place several file types in Cloud Storage and inspect how Dataplex/catalog concepts can help discovery and governance. Define access groups and lifecycle rules.

The goal is not creating a huge lake; it is understanding how uncontrolled object storage becomes a governed data platform.

Lab six: optimize a BigQuery query

Run a query over a small table, then compare partitioning/clustering, materialized view or query rewrite concepts. Inspect bytes processed and execution behavior.

Performance work should always be evaluated against business need and cost.

Lab seven: prepare data for ML or RAG

Create a small feature dataset for BigQuery ML or prepare text documents for embeddings/retrieval. Identify PII and remove or protect it before downstream use.

Data engineers should validate quality and governance even when another team builds the model.

Lab eight: orchestrate with Composer or Workflows

Create a simple DAG or workflow that runs two dependent jobs, logs state and handles a retry. Keep the environment small or use a training lab to control cost.

The PDE hands-on mindset is to understand dependency, retry and observability rather than memorize UI steps.

Lab nine: simulate failure and recovery

Introduce a failed job, missing file, permission denial or bad schema. Use logs and monitoring to identify the first failure, then recover safely.

Record whether replay is idempotent and whether partial output needs cleanup.

Lab ten: audit cost, security and cleanup

Review IAM, unused resources, cluster uptime, query cost, storage lifecycle and logs. Delete disposable resources and preserve only non-sensitive notes or code.

Add a project-and-dataset IAM exercise. Give a service account access to only the dataset it needs, then test a denied action. This demonstrates least privilege and helps distinguish project-level from dataset/table-level governance decisions.

Add a customer-managed encryption-key tabletop. Identify which service account needs key permissions and what happens if the key becomes unavailable or is rotated incorrectly. Encryption design includes operational dependency, not only the algorithm.

Add a schema-evolution lab in the batch pipeline. Introduce a new nullable field, then an incompatible type change. Observe which change succeeds and where validation should stop unsafe data before the sink.

Add a dead-letter or error-output pattern for malformed records. Do not let one bad event crash the entire pipeline indefinitely. Capture enough context for remediation while protecting sensitive data in error logs.

Add a late-data streaming exercise with fixed windows. Send events after the expected window and observe or model allowed lateness/watermark behavior. Record how business totals change as late data arrives.

Add a Dataproc job using a training lab or small ephemeral cluster. Run a simple Spark transformation, inspect logs and delete the cluster. Compare startup/operational overhead with a serverless managed pipeline.

Add a BigQuery partitioning/clustering test with the same query before and after optimization. Record bytes processed and timing. The result makes storage/query design consequences visible without huge datasets.

Add a Spanner or Bigtable tabletop if full deployment is expensive. Model a global relational transaction for Spanner and a key-range time-series/device workload for Bigtable. The objective is choosing the right data model, not incurring unnecessary cloud spend.

Add a data-catalog exercise by tagging or documenting dataset metadata and ownership. Ask another person to find the trusted table without verbal guidance. If discovery depends on tribal knowledge, governance is incomplete.

Add masking or DLP to a synthetic sensitive field. Compare removing the field, masking it and granting restricted access. The correct privacy control depends on what analysts still need to accomplish.

Add an Analytics Hub/data-sharing tabletop. Define a producer, consumer, approved dataset and revocation requirement. Compare controlled sharing with exporting CSV copies to every team. The managed share should preserve a clearer source of truth.

Add a Composer DAG with a sensor or conditional dependency if the lab supports it. Make one upstream dataset unavailable and verify the workflow waits or fails visibly rather than generating partial downstream output.

Add CI/CD around one pipeline definition using a lightweight repository and validation step. Change transformation logic, review the diff and promote to a test environment before production. This is enough to practice repeatability without building a large delivery platform.

Add quota and billing troubleshooting. Run a harmless workload that reaches a small configured limit or examine a training example, then identify whether the error comes from quota, permission or budget. Similar symptoms can require different owners.

Add a data-freshness alert. Instead of monitoring only job failure, detect when the newest record is older than expected. This catches upstream silence that a technically successful pipeline can miss.

Add a replay exercise from raw immutable input. Re-run a processing interval and confirm the sink does not duplicate records. Idempotent replay is one of the most practical recovery skills for a data engineer.

Add a regional-failure tabletop. Decide which data is replicated, which jobs can run elsewhere, how DNS/connections change and what data-loss period is acceptable. Recovery design should follow business requirements rather than automatically duplicate everything globally.

Finish by documenting the lab as a data product: source, schema, owner, classification, pipeline, storage, consumers, SLOs, cost notes, runbook and recovery. Another engineer should be able to operate it without relying on your memory.

Add a Dataform or SQL-transformation exercise around a curated BigQuery layer. Version transformation logic, run assertions or validation and compare a development change with production promotion. This demonstrates that SQL-based data transformation still benefits from software-engineering discipline.

Add a Pub/Sub retention/replay tabletop. Assume a downstream processor is unavailable for a period and decide how messages remain recoverable. Then compare that with replaying from immutable Cloud Storage. Recovery design should identify the authoritative reprocessing source.

Add a BigQuery data-sharing exercise with an authorized view or shared dataset. Let one consumer query only the fields they need. Compare this controlled access with exporting a full copy and explain which is easier to revoke or keep fresh.

Add a query-quota scenario. A team exceeds a quota or reservation capacity during a business-critical load. Inspect the error or monitoring signal, identify whether to reschedule batch work, adjust capacity or optimize queries, and avoid treating every quota message as a code bug.

Add a data-lineage record to the project. For one dashboard metric, document source table, transform, curated table and report. Then change the source definition and identify downstream assets. This makes lineage useful rather than a catalog checkbox.

Add one fault-tolerance design for Composer or pipeline orchestration. Define retry count, backoff, idempotency and alert behavior. Excessive retries can waste cost or amplify bad data, while no retry can make transient failures operationally noisy.

Finish with a deliberate handoff: give another person the runbook and ask them to recover from a simulated failure. If they need undocumented help, turn that knowledge into the runbook or automation. Operability is part of professional data engineering.

Add a Database Migration Service or Datastream tabletop using a small relational source. Define initial load, ongoing change capture, cutover and validation. Compare this with a one-time export so continuous migration requirements become obvious.

Add a Data Transfer Service scenario for bringing supported SaaS or Google data into BigQuery. Identify schedule, destination dataset, credentials and freshness. The lab does not need a real commercial source; the purpose is understanding managed recurring transfer.

Add a Cloud Storage lifecycle policy for temporary staging files. Move or delete synthetic objects after a short period and document why. This makes storage lifecycle a cost/governance mechanism rather than an abstract feature.

Add a Composer versus Workflows comparison on paper. Use a data DAG with many scheduled tasks for Composer and a service-to-service orchestration example for Workflows. The goal is selecting orchestration style based on ecosystem and dependency complexity.

Add a BigQuery reservation/capacity tabletop. Split interactive analysts and scheduled batch workloads into separate capacity assumptions and decide how priority or reservation can protect business-critical work. This is more useful than memorizing edition names alone.

Add a missing-data incident. Remove one expected input file while leaving the pipeline technically runnable. Make the freshness/completeness check fail and investigate. This demonstrates why data SLOs must measure business data, not only task status.

Professional practice includes operating the system responsibly after the successful data pipeline demo.