Hands-on MLA-C01 practice is still valuable even during the MLA-C02 transition because the traditional ML-engineering lifecycle remains foundational. The most useful lab takes one dataset from ingestion through feature engineering, model training, deployment, monitoring, and security. English candidates should map the same lab onto MLA-C02, while candidates still taking C01 in Japanese, Korean, or Simplified Chinese can use it directly against the C01 objectives.
Keep the MLA-C01 domain structure visible during the lab so each exercise has an exam purpose: 28% data preparation, 26% model development, 22% deployment/orchestration, and 24% monitoring/maintenance/security.
Lab one: ingest one dataset in two formats
Store a tabular dataset in CSV and Parquet, then compare schema handling, file size, and query or processing behavior. If practical, use S3 as the source and load it through Glue or SageMaker tooling.
The goal is to see why format and storage are architecture choices rather than file-extension trivia.
Lab two: create a repeatable preprocessing pipeline
Handle missing values, outliers, scaling, encoding, and feature creation in code or Data Wrangler. Save the transformation logic so training and later inference use compatible preprocessing.
A Glue or Spark-based path is useful when the workload has larger-scale transformation needs.
Lab three: measure class imbalance and bias
Create or identify an imbalanced classification dataset and compare label distribution before training. Try resampling or another mitigation, then evaluate whether the change improves the metric that matters.
Use SageMaker Clarify concepts to connect pretraining bias checks with production monitoring.
Lab four: train two model approaches and compare them
Choose a baseline and a more capable model, train both, and compare validation metrics, training time, interpretability, and inference requirements.
The SageMaker environment can support reproducible jobs and model artifacts instead of relying only on an interactive notebook session.
Lab five: tune one hyperparameter intentionally
Change a small set of hyperparameters and record the expected effect before running the experiment. Keep a model version, metric result, and configuration for each run.
The lesson is traceability. Tuning is engineering when the experiment can be reproduced and compared.
Lab six: deploy the model with the wrong pattern, then fix it
Model a real-time endpoint for a batch-only workload or a batch design for a latency-sensitive API. Compare cost, latency, and operations.
The deployment exercise should make endpoint choice follow service requirements rather than habit.
Lab seven: express infrastructure as code
Use CloudFormation or another AWS IaC pattern to represent a small part of the environment. Review the template, deploy it, change it, and understand how the desired state is recorded.
Production ML should not depend on manually remembering how training or endpoint resources were created.
Lab eight: orchestrate the workflow
Connect preprocessing, training, evaluation, registration, and deployment steps conceptually or with Step Functions and ML pipeline tooling. Add a failure path so one bad evaluation prevents deployment.
This creates a visible MLOps control flow rather than a chain of manual notebook actions.
Lab nine: add monitoring and alarms
Track endpoint latency, error rate, throughput, resource utilization, data quality, and model drift. Use CloudWatch and SageMaker monitoring concepts to define which conditions require investigation.
Monitoring should connect back to the training assumptions that justified the model.
Lab ten: audit access and encryption
Review IAM roles, S3 or model-artifact permissions, KMS encryption, network access, and sensitive-data handling. Remove one unnecessary permission and verify that the intended workload still functions.
Add a second ingestion path using a streaming source concept such as Kinesis or Kafka. Compare the assumptions with file ingestion: events can arrive continuously, ordering and buffering matter, and downstream transformations may need windowing or state. Even a conceptual comparison strengthens Domain 1 because the exam expects several ingestion mechanisms.
Add a Feature Store exercise conceptually or practically. Define two reusable features, note how they are computed, who owns them, and how training and inference consume consistent values. The key lesson is avoiding training-serving skew and duplicated feature logic across teams.
Add an evaluation experiment where overall accuracy looks good but recall is poor for the positive class. Change the decision threshold or model and observe the trade-off. This teaches why metric selection and thresholding should follow the business consequence rather than a single leaderboard score.
Add a model-registry step. Record model version, dataset reference, feature code, hyperparameters, metric result, and approval state. Then imagine an incident after deployment and ask whether you can identify exactly which artifact and configuration were serving traffic.
Add a canary or staged-deployment thought experiment. Route a small portion of traffic to a new model version, compare metrics, and define rollback criteria. This connects CI/CD, monitoring, and deployment risk in a way that a one-step endpoint update cannot.
Add an autoscaling scenario where request volume doubles. Decide which metric drives scale, what minimum and maximum capacity make sense, and how quickly the system should react. A model endpoint is only useful if infrastructure can keep service-level objectives during demand changes.
Add a drift scenario by changing the distribution of one important input feature. Predict how model performance might change and which monitoring signal would detect it. Then decide whether the response is retraining, investigation of source data, or a business-rule change.
Add an infrastructure-cost review. Identify training compute, notebook time, storage, endpoint uptime, batch jobs, logs, and data transfer. Then choose one optimization that does not violate latency or reliability requirements. Cost is part of production engineering, not an accounting exercise performed after launch.
Add a permissions audit from the point of view of three identities: data-preparation job, training job, and inference endpoint. Reduce access so each role can reach only required data, artifacts, keys, and logs. This turns IAM from theory into an explicit trust model.
Close the lab by documenting the entire pipeline in a one-page runbook: source, feature process, model version, deployment type, scaling, monitoring, rollback, IAM, and encryption. The lab is mature when another engineer could operate the system without knowing how you clicked through the console.
Add a batch-inference exercise alongside the online endpoint. Score a large offline dataset, record where inputs and outputs live, and compare operational cost with keeping an endpoint running. The difference makes it easier to recognize when batch transform is the natural deployment pattern.
Add an asynchronous-inference scenario for large payloads or longer processing. Compare request handling, user expectations, and scaling with a synchronous real-time endpoint. The deployment mode should reflect interaction pattern, not simply the model type.
Add a small infrastructure-failure scenario: endpoint capacity is exhausted or a pipeline step loses permission. Use logs and metrics to distinguish infrastructure failure from model-quality failure. ML engineers need to diagnose the serving platform as well as the model.
Add one sensitive-data field and trace it through source, transformed dataset, training artifact, logs, and endpoint payload. Decide where masking, encryption, or exclusion is needed. This exercise ties compliance directly to the lifecycle and makes data-protection objectives far more concrete.
Add one model-performance regression after release. Keep infrastructure healthy but change production input characteristics so predictions degrade. This demonstrates why endpoint health and model health are different monitoring dimensions and why both need production thresholds.
Finally, compare the finished lab with the MLA-C02 additions. Ask where Bedrock, RAG, foundation-model evaluation, agent workflows, or responsible-AI controls would fit. English candidates can extend the same lab toward C02 without rebuilding the foundational data, deployment, CI/CD, monitoring, and security work from scratch.
Add one data-quality gate before training. Fail the pipeline when a required column is missing or class balance moves beyond an agreed threshold. This makes quality a release condition rather than a dashboard reviewed after a poor model is already trained.
Add one model-explainability or bias review using the available SageMaker tooling. Compare what the model uses with the business expectations and decide whether a feature or subgroup deserves deeper investigation before promotion.
Add one infrastructure-as-code change that modifies endpoint capacity or networking, then review the diff before deployment. The exercise demonstrates that infrastructure changes should receive the same review discipline as model code.
Add one rollback drill after a failed canary. Restore the prior model version, confirm endpoint health, verify key business and technical metrics, and document why rollback was triggered. Recovery criteria should be defined before the release starts.
Add one access-denied incident intentionally by removing a required IAM permission in a test environment. Use the resulting logs and error message to identify the missing permission, then restore only the needed action instead of broadening the role excessively.
The strongest final lab report should include not just screenshots but decisions: why this data format, why this model, why this endpoint, why these metrics, why these permissions, and why this release gate. That reasoning survives the transition from C01 to C02 better than console-specific steps.
Add one final peer review. Give another engineer the runbook, model metrics, IaC, and monitoring plan and ask them to identify a hidden dependency or rollback gap. Production-readiness improves when the system can be understood and challenged by someone other than its author.
Record the peer-review findings as concrete actions—missing metric, excessive permission, unclear rollback, or undocumented dependency—and close them before calling the lab production-ready.
That closure turns experimentation into a repeatable engineering practice.
Keep the evidence with the runbook for later review.
Preserve that record.
The IAM and KMS layers complete the lab because an accurate model is not production-ready if access and data protection are weak. C01’s practical value remains strong even as English certification shifts to C02.