DOP-C02 hands-on preparation should prove that you can automate delivery and operations, not merely identify AWS service names. One small reference application can cover most of the blueprint if you deploy it with CI/CD, define its infrastructure as code, monitor it, trigger event-driven remediation, protect it with security controls and test recovery.
Use the current DOP-C02 six-domain guide as the lab map and keep costs bounded with disposable resources.
Lab one: create a complete CI/CD path
Connect source, build, tests, artifact storage and deployment. Add one manual or automated quality gate and deliberately fail a test.
Use CodePipeline or an equivalent AWS-native flow and record exactly which stage blocks the release.
Lab two: deploy with two strategies
Compare rolling with blue/green or canary behavior. Generate a harmless faulty release and observe health checks, traffic shift and rollback.
A CodeDeploy exercise should focus on release safety rather than only successful deployment.
Lab three: build the infrastructure from code
Define network, compute, IAM and monitoring resources with CloudFormation/CDK. Commit the template, create a change, inspect the diff/change set and deploy.
Introduce one missing permission or invalid dependency and diagnose the stack event instead of fixing it manually in the console.
Lab four: operate at scale with Systems Manager
Use a safe command, inventory or automation document against tagged test instances. Scope the action narrowly and inspect success/failure output.
The Systems Manager lab should demonstrate that fleet automation needs identity, targeting and observability.
Lab five: test resilience and recovery
Use multiple Availability Zones, Auto Scaling or a managed resilient service where practical. Stop one test component and confirm traffic remains available or the system recovers.
Then restore from backup or run a tabletop for a larger regional failure so HA and DR remain separate concepts.
Lab six: create an observability baseline
Build CloudWatch metrics/logs/alarms and add CloudTrail. If the application is distributed, add tracing where practical.
A healthy CloudWatch baseline should make later latency, error-rate or resource anomalies immediately visible.
Lab seven: automate a low-risk remediation
Trigger EventBridge from an event or alarm and invoke Lambda or Systems Manager to correct a known test condition. Log every step.
Add a failure path so operators are notified if remediation cannot complete.
Lab eight: enforce security as code
Use IAM least privilege, encrypted storage, secret management and Config/policy controls. Deliberately create one noncompliant test resource and detect it.
Correct the issue through automation or a versioned infrastructure change rather than one undocumented console edit.
Lab nine: simulate a cross-account workflow
If multiple accounts are unavailable, diagram a pipeline that assumes a deployment role in another account and centralizes logs. Define trust, permissions and artifact access.
The professional exam expects organizational-scale patterns even when a personal lab is small.
Lab ten: run one blind production incident
Introduce a safe failure without recording which layer changed. Start from alarms and logs, inspect recent deployments/API changes, identify the failing component, remediate, validate users and update the pipeline/IaC if the incident exposed a repeatable weakness.
Add source-control protections to the pipeline lab. Require pull-request review or a protected branch before a production-affecting change can enter the pipeline. Tag releases and preserve the commit/artifact relationship so a deployed version can be traced back to exact source.
Add an artifact-promotion experiment. Build once, store the artifact under an immutable version, deploy the same artifact to staging and production, and confirm that only configuration changes between environments. This demonstrates reproducibility better than rebuilding at each stage.
Add a pipeline-permission failure. Remove one required KMS or S3 permission from a test role and observe exactly where the pipeline fails. Restore the minimum permission instead of attaching an administrator policy. This is one of the best ways to practice IAM under DevOps conditions.
Add CloudFormation drift. Modify one harmless resource manually, run drift detection, and decide whether code or resource state should be authoritative. Then restore through IaC. This lab reinforces why console changes undermine repeatability.
Add a StackSet or multi-account tabletop. Define administrator and execution roles, target organizational units/accounts, Regions, failure tolerance and rollback. The purpose is understanding organization-scale deployment even if a personal lab has only one account.
Add Systems Manager patch or inventory reporting in a safe test fleet. Tag instances, target only the intended group, run the operation and inspect results. Change one tag and verify the automation scope changes as expected.
Add an Auto Scaling failure-recovery lab. Terminate a test instance and observe replacement, target registration, health checks and alarms. Then compare with a database failure where instance replacement is not the whole recovery story.
Add a backup-and-restore drill for one stateful component. Restore to a new resource, update connection information safely, and validate application behavior. Record actual recovery time against the hypothetical RTO.
Add a synthetic canary or simple end-to-end health check. Resource metrics can be normal while the user flow is broken. A canary that performs the user action provides evidence closer to business availability.
Add distributed tracing to a two-service path. Introduce latency in one downstream call and verify the trace identifies the slow segment. Compare that evidence with aggregate CloudWatch latency metrics.
Add an EventBridge remediation failure. Let the event match successfully but make the target lack permission or fail internally. Verify logs, retry behavior, dead-letter or alert path. This demonstrates why event routing and target execution must be monitored separately.
Add a Config compliance lab. Create a harmless noncompliant resource, detect it, then remediate by IaC or automated action. Preserve the finding and fix evidence so the process is auditable.
Add a secrets rotation tabletop. Assume an application credential must rotate without redeploying source code. Map secret storage, workload permission, rotation event, application refresh behavior and rollback. This combines security with operational continuity.
Add one cost guardrail to the lab: budget alert, automated cleanup of ephemeral test resources, or log-retention policy. DevOps labs can quietly create spend, and the exam expects cost-aware operational judgment even without a dedicated cost domain.
Document every lab in the same format: objective, starting state, code/config change, expected behavior, evidence, failure, rollback and lesson. This creates a personal incident/runbook library and makes repeated experiments much more valuable than screenshots alone.
Finish by rebuilding the environment from source and IaC after deleting disposable resources. If the pipeline, infrastructure definition, secrets, configuration and observability can recreate a working service without hidden manual steps, the lab has reached true DevOps maturity.
Add a controlled IAM least-privilege review after the pipeline works. Record the API actions actually used by build and deployment roles, remove unnecessary permissions and rerun the pipeline. This turns “least privilege” from theory into evidence that the workflow still functions with smaller authority.
Add log centralization across at least two sources. Send application logs and CloudTrail-style audit events to a protected destination, then use one incident to query both. Different log types should retain their purpose while sharing governance and retention.
Add one synthetic transaction that crosses the entire application path. Trigger an alert when the transaction fails even if individual resources are healthy. This demonstrates how user-facing availability can differ from server-level health.
Add a configuration-drift incident caused by an emergency manual edit. Document why the edit happened, detect it, restore desired state through IaC and update the runbook so the emergency path is still auditable. Real systems sometimes require exceptions; mature DevOps makes them temporary and visible.
Add a security-finding remediation where the fix enters source control. Instead of correcting one resource directly, change the template or base image, deploy through the pipeline and verify the finding closes. This is the clearest example of DevSecOps feedback.
Run a cleanup validation at the end of the lab. Ensure ephemeral environments, old artifacts, test logs, unused roles and temporary secrets are removed or retained according to policy. Operational maturity includes decommissioning, not only creation.
Package the final lab evidence as if handing it to another engineer: repository, pipeline diagram, IaC, runbooks, alarms, dashboard, IAM roles, recovery procedure and one incident report. If another engineer can rebuild and operate the service, your practice has moved beyond a demo.
Add a deployment-metadata check. Tag or record the application version, source commit and pipeline execution on the running workload, then query that metadata during a simulated incident. This allows operators to correlate a regression with the exact release instead of guessing from deployment time.
Add a failure-budget exercise for alarms. Create one alarm that is too sensitive and another that misses real user impact, then tune thresholds or composite conditions. Monitoring quality is measured by useful action, not the number of alarms in the account.
Add a cross-account artifact-access tabletop. Keep the artifact encrypted and define which pipeline/deployment roles can read it in each account. The exercise ties together KMS, S3 policy, IAM trust and deployment flow without requiring a large organization lab.
Add one restore-after-region-loss tabletop that includes DNS, certificates, secrets, IAM roles, logs and external dependencies. A recovered database alone does not restore the service. This makes disaster recovery a complete business-system exercise.
Before final cleanup, intentionally rerun the full deployment from an empty test environment. Note any manual prerequisite you forgot to encode. Hidden prerequisites are exactly what Infrastructure as Code and automated pipelines are supposed to eliminate.
Keep a final list of every manual exception discovered during the labs and decide whether it belongs in code, configuration, a runbook, or documented approval. The fewer undocumented special steps remain, the more closely the environment reflects the automation principles DOP-C02 is designed to test.
Within the AWS certification portfolio, this closed-loop lab is the strongest DOP-C02 preparation because it combines delivery, operations, security and learning.