Applied preparation for AI-200 should look like a small production system, not a collection of disconnected portal exercises. The exam’s four domains are designed around the work of an Azure AI cloud developer: package and host code, retrieve and manage data, connect services asynchronously, secure configuration, and collect enough telemetry to troubleshoot the complete solution.
A useful lab therefore has a business-shaped flow. Imagine an internal knowledge service that receives documents, creates or stores embeddings, retrieves relevant context for a user request, caches safe reusable results, and processes ingestion asynchronously. It is simple enough to build in stages but rich enough to touch nearly every objective.
The value of this kind of practice is not the finished demo. The value comes from the decisions and failures encountered along the way. Each exercise below should produce evidence that you can explain, not just a green status indicator.
Lab 1: create a versioned container delivery path
Build a small Python API and containerize it. Push versioned images to Azure Container Registry and use a naming convention that makes rollback possible. Practice an ACR Task so the image lifecycle is not limited to manual local builds.
Deploy the container to App Service and provide configuration through environment settings rather than editing the artifact. Introduce a missing variable, observe the startup behavior, and document where the failure appears. Then correct the configuration without rebuilding the image.
This exercise teaches separation of code, artifact, and runtime configuration—a production discipline that later makes troubleshooting much faster.
Lab 2: compare Container Apps and AKS with the same service
Deploy the API to Container Apps. Create a new revision and observe how traffic and configuration are associated with versions. Configure an event-driven scaling scenario with KEDA so you understand why the platform can increase or reduce replicas based on work rather than only CPU.
Next, deploy the same application to Azure Kubernetes Service using manifests. Inspect pod status, events, service configuration, and logs. Break the service selector or another connectivity element and diagnose the difference between an application failure and a Kubernetes routing problem.
Write down what extra control AKS provides and what extra responsibility it creates. The exam is easier when platform choice is understood as a trade-off rather than a hierarchy of “simple” and “advanced.”
Lab 3: build semantic retrieval in Cosmos DB
Create a small document collection in Cosmos DB and connect through the SDK. Store embeddings with the records and run vector similarity search. Track request-unit behavior as you change queries or indexing so semantic retrieval remains connected to cost and performance.
Add an ordinary field such as department, product, or tenant and compare an unfiltered vector query with one constrained by application context. This demonstrates a practical truth of AI retrieval: the nearest semantic match is not necessarily the correct authorized result.
Then enable a change feed processor for new or updated documents. Use it to trigger a small downstream action such as logging, enrichment, or an embedding refresh. This ties the data and event-processing ideas together.
Lab 4: implement RAG retrieval with PostgreSQL and pgvector
Create a relational dataset in Azure Database for PostgreSQL. Use appropriate types and indexes, then add vector storage. If relational fundamentals need refreshing, PostgreSQL essentials provides useful context before moving into pgvector-specific tuning.
Build a query that combines semantic similarity with metadata filters and use the result as context for a simple RAG flow. Measure query latency and inspect how index choices and workload size affect performance. The exam explicitly includes reducing pgvector compute overhead, so performance should be an observed property, not a theoretical note.
Finally, test connection behavior under concurrency. A production API that opens database connections carelessly can fail even when every query is individually correct.
Lab 5: make Redis freshness visible
Put a frequently requested value in Azure Managed Redis and set an expiration policy. Update the source record without invalidating the cache, then observe the stale result. Add explicit invalidation and repeat the test.
This is a better lesson than simply proving that cached reads are fast. AI applications often combine fast-changing business data with expensive model or retrieval calls. The team needs an intentional freshness policy so performance optimizations do not quietly undermine correctness.
If your environment allows, add a vector index in Redis and compare retrieval latency and architecture with your Cosmos DB or PostgreSQL version. Focus on reasoning about the workload rather than declaring a permanent winner.
Lab 6: design asynchronous ingestion with failure handling
Use Azure Service Bus to queue document-processing work. Build a consumer that intentionally rejects malformed messages and confirm that dead-letter handling preserves the failed work for investigation. Then use a topic with two subscriptions to model fan-out.
Add Event Grid to publish a state-change event and apply a filter so only relevant events reach a consumer. Configure retry behavior and compare what happens when the endpoint is temporarily unavailable.
The key skill is explaining why the two services exist in different roles. Durable work queues, pub/sub messaging, and event notification solve related but not identical problems.
Lab 7: use Functions for focused event-driven processing
Create an Azure Function with a message or HTTP trigger and deploy it as a function app. Use bindings where they reduce boilerplate, but understand the underlying service interaction so the convenience does not hide the architecture.
A good exercise is to process a queue message that describes an updated document, recalculate a value, write to the data store, and emit telemetry. Keep the unit of work small and idempotent so retries do not corrupt state.
Then compare the operational behavior with your containerized service. Functions can simplify bursty event work, while containers may be more appropriate for long-running APIs or components with different runtime needs.
Lab 8: prove that secrets and configuration can change independently
Move credentials and keys into Key Vault and implement retrieval through the application identity. Rotate a secret and confirm that the system continues to work using the intended retrieval pattern. Store ordinary settings in Azure App Configuration.
Now break access deliberately. Remove the permission or point the application at a missing configuration key. Capture the error and decide what telemetry would make the incident obvious in production.
This exercise builds security and troubleshooting together. A protected secret is useful only if the application can consume it correctly and operators can distinguish an authorization problem from a code defect.
Lab 9: trace a failure across the entire system
Instrument the API and event-driven components with OpenTelemetry. Generate requests that touch the container, cache, database, queue, function, and another downstream dependency. Use trace context to follow a transaction across those boundaries.
Then introduce one slow dependency and one failing dependency. Query logs and metrics with KQL to isolate which component changed. Avoid jumping straight to the resource that “looks most AI-related.” The bottleneck may be connection handling, message backlog, cache misses, or container scaling.
This end-to-end habit is what separates AI-200 from a fundamentals exam such as AI-901. The candidate is expected to build and operate the cloud components that make an AI solution dependable.
Lab 10: run a controlled production incident
After the individual labs work, create one end-to-end incident in the combined system. For example, deploy a revision that increases database calls, disable part of the cache path, or reduce consumer capacity so the Service Bus backlog grows. Generate a realistic amount of traffic and observe the user-facing symptom before inspecting individual resources.
Start with telemetry, not guesses. Use a trace to identify where the request is spending time, then correlate that observation with resource metrics, application logs, queue depth, or database behavior. Write at least one KQL query that narrows the time window and isolates the failing component. Record what evidence would have alerted an operator earlier and which metric would make a useful production threshold.
Then recover deliberately. Roll back the revision, restore the cache path, or increase consumer capacity. Confirm that latency and backlog return to normal and that no messages or state changes were lost. Finally, write a short incident note with cause, impact, evidence, recovery, and prevention. This may feel more operational than a traditional “AI lab,” but that is precisely the point of the current blueprint. AI-200 validates developers who can keep AI-enabled cloud services dependable after the interesting model work is already inside the application.
After completing the labs, repeat one of them without using the portal as your primary source of truth. Work from code, deployment artifacts, configuration, and telemetry, then use the portal to confirm rather than discover what happened. The goal is not to avoid graphical tools; it is to build a mental model strong enough that a changed interface does not erase your understanding. AI-200 is testing durable cloud-development concepts—deployment state, data access, asynchronous work, security boundaries, and observability—and those concepts should remain clear regardless of which management surface you use on exam day or at work.
Repeat the incident until diagnosis becomes systematic.