Hands-on PMLE preparation should focus on the decisions that turn an experiment into a production AI system. One reference problem can cover data preprocessing, BigQuery ML or AutoML, custom training, foundation-model evaluation, endpoints, pipelines, monitoring and responsible AI. The current Professional Machine Learning Engineer guide rewards this lifecycle approach.
Lab one: solve one problem with two model approaches
Choose a tabular classification or regression problem. Build a simple BigQuery ML or AutoML baseline, then compare it with a custom model.
Record quality, development effort, interpretability, serving requirements and cost. The point is model-selection judgment, not maximizing benchmark score.
Lab two: create a reproducible data-preparation path
Store raw data, clean it with SQL/Dataflow/Python, version the transformation logic and create train/evaluation splits. Check for leakage, missing values, class imbalance and sensitive fields.
Then confirm the same preprocessing can be applied during inference.
Lab three: track experiments and lineage
Run several parameter or feature experiments and log model version, dataset, metrics and notes. Use the current Google Cloud experiment/metadata concepts or a simple equivalent.
A good lab should let another engineer reproduce why one model was selected.
Lab four: compare training hardware
Run or estimate the same training job on CPU versus accelerator-capable infrastructure where appropriate. Record runtime, cost and utilization.
Do not use GPU/TPU simply because they are available; match hardware to workload characteristics.
Lab five: prototype a foundation-model solution
Use an approved Gemini or Model Garden model on non-sensitive data. Create a small task with prompts/context and evaluate several outputs against explicit criteria.
Then compare prompting, retrieval and fine-tuning as possible adaptation strategies.
Lab six: deploy batch and online inference
Serve the conventional model through an online endpoint and run batch inference separately. Compare latency, throughput, cost and operational complexity.
Package a custom model with an appropriate container only if the managed serving path does not meet the requirement.
Lab seven: test a safe model rollout
Register two versions and use a canary or A/B-style approach to route limited traffic to the candidate version. Define rollback criteria before switching traffic.
This makes model deployment look like production software delivery rather than notebook publication.
Lab eight: build an automated pipeline
Create a pipeline for data validation, training, evaluation and deployment approval. Add a retraining trigger based on schedule, new data or monitored drift.
A data-to-deployment workflow should be versioned and reproducible enough to rebuild the model from known inputs.
Lab nine: monitor model and data behavior
Establish a baseline and introduce synthetic drift or data-quality change. Observe model monitoring, serving metrics and evaluation signals.
Decide whether the right response is retraining, data correction, rollback or investigation rather than automatically retraining every anomaly.
Lab ten: add generative-AI safety and evaluation
Create prompts containing harmless adversarial or policy-sensitive cases and test safety filters or Model Armor-style protections. Evaluate quality with task rubrics and human review.
Add a BigQuery ML versus AutoML comparison using the same dataset. Keep the train/evaluation split and metric consistent so you can compare engineering effort as well as model quality. The best tool is the one that meets the objective with acceptable complexity, not automatically the one with the most control.
Add a privacy exercise before training. Identify PII or sensitive columns, decide whether they should be removed, transformed, access-controlled, or justified for the model, and document the decision. Privacy should be resolved before the data is copied into notebooks or prompts.
Add a feature consistency test. Compute one feature in the training pipeline and again in the serving path; deliberately make them differ, then observe the effect on predictions. This makes training-serving skew a real engineering problem rather than a vocabulary term.
Add an experiment lineage drill. Reproduce a prior result using only recorded metadata. If you cannot identify the dataset, code, parameters, framework/container, and metric, your experiment tracking is incomplete. Reproducibility is a practical MLOps competency.
Add one hyperparameter-tuning exercise where the search is bounded by budget. Compare the gain from tuning with additional compute cost. Professional ML engineering requires knowing when another search round is unlikely to deliver enough value.
Add a distributed-training tabletop. Take a model too large or slow for one accelerator and decide whether data parallelism or model parallelism is appropriate. Identify where communication overhead could reduce scaling efficiency. The exam tests the concept more than implementation syntax.
Add a RAG experiment for the foundation-model project. Use a small document corpus, retrieve relevant passages, and compare groundedness with and without retrieval. Include one document the user is not authorized to see and ensure the application cannot retrieve it.
Add a prompt-injection test using a harmless document that tells the model to ignore application instructions. Confirm that application-level authorization and tool restrictions remain enforced even if the model responds unpredictably. This is a security-boundary exercise, not a prompt-writing contest.
Add a model-selection experiment with two foundation models of different capability/cost. Measure task success, latency, and token/serving cost on the same test set. A smaller model may be the better production choice if quality remains sufficient.
Add preprocessing/postprocessing to the deployed endpoint. Validate inputs, transform features, invoke the model, convert the raw output into the application’s expected schema, and handle failures. This emphasizes that production inference is a system, not only a model file.
Add a private-endpoint or network-control design if you cannot implement it. Identify which clients need access, how identity and network boundaries interact, and what logs prove legitimate use. Public-versus-private serving is both an architecture and security decision.
Add an A/B evaluation where business behavior is measured in addition to model metric. A model with slightly better offline accuracy can still perform worse for users because latency or downstream behavior changes. Production validation should include the whole experience.
Add a pipeline failure in the validation stage. Feed malformed or shifted data and ensure the pipeline stops before training/deployment. Data validation is valuable because failing early is cheaper than diagnosing a bad production model later.
Add a controlled retraining trigger. Generate enough synthetic data drift to start a pipeline, retrain a candidate, evaluate it, and deliberately fail the promotion criterion. This demonstrates that retraining and deployment are separate decisions.
Add monitoring for both system and model signals. Track endpoint latency/error rate plus data drift or prediction-quality proxy. For the generative application, add groundedness/safety/task-success evaluation. One dashboard should not collapse these different health dimensions into a single number.
Add a rollback runbook. Assume the new model causes unacceptable errors or unsafe outputs. Identify how to route traffic back, preserve evidence, notify owners, and prevent the bad version from being promoted again until root cause is understood.
Add a cost review across training, storage, serving, feature retrieval, foundation-model calls, and monitoring. ML systems can spend money at several layers. Record which cost grows with data size, model size, request rate, context length, or retained telemetry.
Add an explainability/fairness exercise to the predictive model. Compare performance across relevant groups where appropriate, inspect feature attribution, and decide whether an observed disparity is acceptable, explainable, or requires redesign. Responsible AI needs evidence, not a checkbox.
Add a generative-evaluation dataset with normal, edge, adversarial, and policy-sensitive examples. Score outputs with a rubric and human review rather than evaluating only fluent answers. Convert any real failure into a permanent regression case.
Finish by packaging the project for another engineer: data description, preprocessing code, experiment results, model registry/version, endpoint, pipeline, monitoring, safety controls, cost notes, and rollback procedure. If the system cannot be handed over, it is not yet production-quality PMLE practice.
Add one notebook-to-pipeline migration exercise. Start with exploratory code that works interactively, then extract preprocessing and training into repeatable components with explicit inputs/outputs. This shows the engineering transition from prototype to production rather than treating the notebook itself as the deployed system.
Add one schema-evolution case. Change an input column type or add/remove a feature and observe which pipeline stage should detect the change. Production ML should fail clearly when the contract changes rather than silently producing incompatible features.
Add one endpoint autoscaling experiment or tabletop. Vary request volume, observe latency and utilization, and decide the minimum/maximum capacity strategy. Serving cost and user experience often depend on the scaling policy as much as the model.
Add one access-control review around model artifacts and endpoints. Decide who may train, register, deploy, invoke, or delete models. Separate human development permissions from runtime service-account permissions. Least privilege applies to the ML lifecycle too.
Add a handoff review with a security or data-governance perspective. Ask whether the project documents data sources, PII handling, model limitations, monitoring, approvals, and incident contacts. Production readiness is organizational as well as technical.
Add one final disaster-recovery tabletop for the AI service. Assume a regional endpoint or critical pipeline dependency is unavailable. Decide how traffic, models, artifacts, data, secrets, and monitoring recover. The exact architecture can vary, but the exercise reinforces that an ML system is still a production service with availability and continuity requirements.
Run the complete practice project once more from clean inputs to deployed output without relying on notebook state or undocumented manual steps. Any hidden step you discover should become code, configuration, documentation, or an explicit approval. That final reproducibility test is a strong measure of professional ML engineering readiness.
The current professional role expects traditional and generative AI systems to be secure, responsible and observable. Applied practice is complete when model quality, software reliability and governance are validated together.