Google Professional Cloud Architect: Hands-On Practice

Hands-on practice for the Professional Cloud Architect exam should teach architecture judgment rather than turn the candidate into a specialist operator for every Google Cloud service. The current guide spans infrastructure, AI, security, migration, CI/CD, observability, and business process design. A useful lab therefore reuses one environment and asks why each component exists, how it is governed, and what happens when it fails.

Use the current Professional Cloud Architect scope as the boundary. Small experiments are enough if they expose the decision, evidence, and trade-off behind the architecture.

Lab one: create a resource hierarchy with deliberate ownership

Model an organization with folders, projects, environments, and teams. Assign roles at different scopes and observe inheritance. Create one overly broad role assignment and replace it with a narrower design.

The IAM exercise should make one principle clear: role capability and assignment scope are separate dimensions of privilege.

Lab two: build a Shared VPC or equivalent network design on paper and in practice

Create or model host and service projects, subnets, firewall policy, routing, load balancing, and private service connectivity. Trace application, administration, and service-to-service traffic separately.

Use Google Cloud load balancing as one practical component, but keep the lab focused on why traffic enters, exits, or stays private.

Lab three: compare GKE and Cloud Run for the same stateless service

Deploy or model a simple containerized application with two operating models. Compare scaling, deployment complexity, network control, runtime assumptions, and team ownership.

The contrast between GKE and Cloud Run is valuable because it turns platform selection into a workload-and-operations decision instead of a preference.

Lab four: create a BigQuery-centered analytical workload

Load a small dataset, build a query, and document the source, transformation, access controls, cost behavior, and downstream consumers. If the environment includes data pipelines, note where orchestration and quality live.

A BigQuery exercise teaches that analytical architecture includes data ownership, query patterns, and cost—not merely SQL syntax.

Lab five: design an AI or agentic workload with permission boundaries

Choose a business use case and define the model, enterprise data, retrieval or tool access, human-approval points, and monitoring. If platform access allows, experiment with a managed AI API or agent builder; otherwise build the design in a clear architecture diagram.

Add a security review: which service account calls which tool, which data is available, and what happens when the agent attempts an action outside its approved scope?

Lab six: deploy one resource with Terraform

Represent part of the environment in infrastructure as code. Plan, apply, change, and remove the resource. Store the configuration in version control and review the change before applying it.

The Terraform exercise should demonstrate repeatability and drift control. The exam is architecture-focused, but implementation quality matters because manual divergence creates operational risk.

Lab seven: build a simple CI/CD and rollback path

Create a minimal application or infrastructure pipeline with automated validation and staged deployment. Deliberately fail a test or deployment and observe the stop condition.

The CI/CD concept becomes architectural when it controls change risk. A faster pipeline is not useful if it allows unreviewed, untested changes into production.

Lab eight: add observability before causing failure

Define the metrics, logs, alerts, and service-level indicators that should show healthy behavior. Then cause a controlled failure—bad configuration, unavailable dependency, excessive load, or permission problem—and verify that the evidence points toward the correct layer.

This teaches why observability is designed before incidents. Logging everything without a clear question can create noise rather than operational clarity.

Lab nine: run a migration tabletop

Take one legacy application and define dependencies, licensing, data volume, downtime tolerance, target architecture, migration method, rollback, testing, and cutover. Decide which components should be rehosted, modified, replaced, or retired.

The lab should show that migration is a program of dependency and risk management rather than a copy operation.

Lab ten: perform a Well-Architected review

Review the finished environment against operational excellence, security, reliability, performance, cost optimization, and sustainability. Identify one improvement per pillar, then rank those improvements by business value and implementation risk.

Add a policy-inheritance exercise to the resource-hierarchy lab. Apply a broad policy at folder or organization level, then create a project-level need that conflicts with it. Observe or model which controls can be overridden and which should remain centrally enforced. This turns resource hierarchy into a real governance mechanism rather than an organizational diagram.

Add a Workload Identity Federation tabletop for an external CI system or on-premises workload. Define the external identity, trust configuration, Google Cloud service account, role scope, and expected short-lived credential path. Compare it with storing a long-lived service-account key. The operational and security benefits become much clearer when the credential lifecycle is drawn explicitly.

In the network lab, include one failure caused by DNS or load-balancer health rather than by firewall rules. Architects need to avoid treating every connectivity complaint as a VPC route problem. Trace name resolution, endpoint selection, health checks, and backend reachability separately so the evidence points to the correct layer.

For the GKE versus Cloud Run comparison, add deployment and incident ownership. Record who patches nodes, who controls runtime versions, how scaling is configured, and which logs are available during failure. Operational responsibility is often the decisive difference between two technically capable platforms.

For the AI lab, add a red-team or misuse review. List prompt-injection, data-exposure, tool-abuse, hallucination, and unauthorized-action risks, then identify which control belongs in application design, model configuration, identity, network, or human workflow. This keeps secure AI inside normal architecture discipline.

In the Terraform lab, introduce drift deliberately through a manual console change. Run a plan and observe the difference between intended and actual state. Then decide whether the manual change should be codified or reverted. This is a concrete demonstration of why infrastructure as code improves architecture governance.

Finish the hands-on cycle with a cost review using actual or estimated resource consumption. Identify which components create fixed cost, usage-based cost, data-transfer cost, or operational labor. Then propose one cost reduction and state which reliability, performance, or support trade-off it introduces.

Document each lab as an architecture decision record: context, decision, alternatives, consequences, and verification. A Professional-level lab is most valuable when it teaches the reasoning behind the deployment, not only the deployment steps.

Add a secrets-and-key-management lab to the security layer. Store one application secret in an appropriate managed service, protect one dataset or storage object with a customer-managed key where practical, and document who can administer the key versus use it. This makes separation of duties visible instead of theoretical.

Add a Private Service Connect or private-service-access tabletop. Start with a workload that should consume a managed or producer service without broad public exposure. Draw the producer, consumer, DNS, IAM, and network boundaries. Compare that with a simpler public endpoint and state the security and operations trade-offs.

Include one organization-policy experiment if your environment permits it. Apply or model a policy that constrains resource location, external IP use, or another organization-level behavior, then observe how it affects individual projects. This demonstrates why hierarchy-wide governance can be stronger and more consistent than relying on every project owner to remember the same control.

For observability, define one technical SLO and one business KPI. Then create the monitoring path that would show whether each is healthy. The technical metric might be request latency or error rate; the business metric might be completed transactions. This prevents “green infrastructure” from being mistaken for business success.

Add a recovery test to the migration tabletop. Assume the cutover fails halfway. Which data changed, which system is authoritative, how does traffic move back, and what evidence proves the rollback is safe? Recovery planning exposes hidden dependencies that a forward-only migration diagram can miss.

Finish the environment with an ownership map. Label every major component with the team responsible for build, security, operations, cost, and incident response. If one resource has no clear owner or several teams assume another team owns it, the architecture has an operational defect.

Add a quota experiment or tabletop. Identify one regional or project-level quota that could limit an otherwise elastic workload, then decide how it would be monitored and increased. This teaches that autoscaling is only as useful as the surrounding limits allow it to be.

Add a decommissioning exercise. Remove a test project or service and verify that data, service accounts, DNS records, secrets, firewall rules, monitoring, and billing artifacts are cleaned up. Architecture ownership includes retirement as well as deployment.

Repeat one lab using a case-study narrative instead of a pure technical requirement. Start from business language, extract constraints, then build only the minimum environment needed to prove the design. This keeps hands-on practice aligned to the way the exam presents architecture decisions.

Add one security incident tabletop after the hands-on labs. Assume a service account is overprivileged or a secret is exposed. Identify containment, credential rotation, log review, affected resources, and the architecture change that reduces recurrence. This keeps identity, observability, and incident response connected.

Keep the final lab small enough to rebuild. If the environment becomes so complex that you cannot explain every dependency, reduce it. Architectural understanding is more valuable than the number of services deployed.

That rebuildability also makes failure testing safer because you can restore the intended state without depending on hidden manual steps.

Keep one final note for every exercise: what would have happened if the failure occurred in production and which team would have been responsible for the first response.

The Professional Cloud Architect perspective is useful because architecture maturity comes from trade-off review. The lab is successful when you can defend what you would change first and why.