Microsoft AZ-305: Azure Architecture Troubleshooting

AZ-305 is a design exam, but troubleshooting skill matters because Microsoft expects Azure solutions architects to evaluate existing solutions and understand how decisions across networking, identity, data, business continuity, and governance affect the whole system. The strongest troubleshooting approach identifies the first architectural assumption that is false and changes the design rather than applying a one-off workaround.

The cases below stay within the current AZ-305 objectives. Use them to practice locating the owning layer before choosing a service or fix.

If a workload can reach the network but receives authorization errors

Separate network reachability from identity. Check managed identity or service principal, role assignment, scope, Key Vault access, resource-specific authorization, and any policy constraints.

Changing virtual-network rules is unlikely to fix an RBAC problem. The evidence should show whether authentication succeeded and where authorization failed.

If an application cannot reach a private service

Inspect virtual-network integration, DNS, private endpoint configuration, route tables, network security, and hybrid connectivity. A correct managed identity does not create a network route.

A virtual network diagram should make the intended path visible before configuration is changed.

If logs exist but the incident timeline is unclear

Review log routing, timestamps, correlation identifiers, workspace design, retention, and ownership. Different services may emit evidence to different destinations, making one incident difficult to reconstruct.

The Azure Monitor design should support correlation across the critical transaction path, not only individual resource dashboards.

If a relational database is healthy but the application is slow

Determine whether the bottleneck is database compute, query design, connection limits, network latency, application concurrency, or repeated reads. Read scaling, caching, or service-tier changes address different causes.

Do not move databases or increase tiers without evidence about which resource is limiting performance.

If a globally distributed Cosmos DB workload experiences hot partitions

Inspect partition-key choice, request distribution, throughput, consistency, and access pattern. A globally scalable service can still perform poorly when most traffic targets one logical partition.

The Cosmos DB architecture should be corrected at the data-model level when the partition strategy is the root cause.

If queue backlog grows continuously

Compare producer rate, consumer rate, downstream capacity, retries, poison messages, and dead-letter activity. A queue can absorb temporary spikes but cannot solve a system whose steady-state processing capacity is below demand.

The Service Bus design may need more consumers, better batching, faster downstream processing, or a revised workflow.

If API consumers fail after a backend change

Check API contract, versioning, policies, authentication, throttling, transformations, and deployment sequencing. API Management can protect consumers only if compatibility and version behavior are designed deliberately.

The API Management layer should make breaking changes visible before they reach every client at once.

If disaster-recovery testing misses the RTO

Break the recovery process into detection, decision, data recovery, compute recovery, network or DNS failover, credential availability, validation, and user restoration. The slowest dependency determines the real RTO.

A backup may be healthy while application recovery remains too slow. Business continuity is end-to-end.

If migration stalls despite sufficient transfer bandwidth

Look for identity, licensing, application coupling, database compatibility, network dependency, source downtime, data validation, and rollback constraints. More bandwidth cannot resolve a blocked dependency.

Architecture troubleshooting during migration should start from the dependency graph and the next irreversible step.

If the same incident returns after every deployment

Move the durable correction into architecture and automation: update policy, infrastructure definition, deployment pipeline, monitoring, runbook, or service boundary. A manual emergency action is incomplete if the next release recreates the same failure.

If a workload receives secrets successfully in development but not production, compare managed identity, Key Vault permissions, network access, private endpoint DNS, and environment-specific configuration. The same application code can behave differently because the control plane around it changes.

If monitoring alerts fire constantly without user impact, review thresholds, aggregation windows, dependency context, and service-level indicators. Alert fatigue is an architectural operations problem. A smaller set of actionable alerts is stronger than a large set of low-value notifications.

If a cache outage brings down the whole application, the architecture may have treated a performance layer as an irreplaceable dependency. Consider graceful fallback, cache warm-up, redundancy, and whether the application can temporarily read from the source of truth.

If a container platform repeatedly requires specialist intervention, evaluate whether the workload genuinely needs that level of orchestration control. A simpler managed runtime may reduce support burden if platform-specific features are not essential.

If network performance degrades after adding hybrid connectivity, inspect route selection, bandwidth, latency, gateway sizing, asymmetric paths, DNS, and where data crosses regions. The issue may be topology rather than compute or application code.

If a regional failover works technically but users cannot authenticate, the recovery design omitted an identity or key dependency. Business-continuity testing should include authentication, secrets, DNS, certificates, and application configuration, not only compute and data.

If a migration creates duplicate or inconsistent data, inspect synchronization direction, cutover timing, authoritative source, retry behavior, and write freeze. The transition architecture needs one clear source of truth at each stage.

If cost rises sharply after increasing resilience, identify which redundancy or data-transfer mechanism caused it. The added cost may be justified by the RTO, or the design may be overprotected. Cost troubleshooting should return to business continuity requirements rather than removing redundancy blindly.

If an application change repeatedly breaks consumers, strengthen API versioning, contract tests, staged deployment, and compatibility policy. The durable fix belongs in the release architecture, not in emergency client updates after every deployment.

After each architecture incident, record the assumption that failed. Was it “the identity always has access,” “the consumer processes faster than producers,” “the cache is optional,” or “regional failover restores everything”? Precise failed assumptions lead to precise design improvements.

If a governance policy blocks a deployment unexpectedly, inspect inheritance and the business reason for the policy before creating an exception. The policy may be enforcing a location, security, or compliance requirement that the workload design failed to consider. A technical workaround can become a governance violation.

If a managed identity works in one subscription but not another, compare role scope, resource hierarchy, tenant context, conditional policies, and target-resource permissions. Recreating the identity may not fix an authorization boundary.

If a virtual machine or container tier scales correctly but users still see timeouts, inspect downstream databases, APIs, queues, DNS, and external services. Scaling the front tier can increase pressure on the true bottleneck and make the incident worse.

If data integration jobs deliver duplicates, inspect source change detection, checkpoint or watermark state, retry semantics, idempotence, and sink behavior. A pipeline can be highly available yet still violate correctness.

If regional recovery works but cost is unsustainable, revisit which workloads actually require hot standby, which can use warm or cold recovery, and which data must be replicated continuously. Continuity architecture should reflect business tiers rather than one universal recovery pattern.

If an API change passes backend tests but breaks clients, the architecture lacked contract validation. Add versioning, consumer tests, staged rollout, and compatibility checks so the API surface is governed as a long-lived interface.

If alerts show rising latency but no resource is saturated, inspect network path, DNS, downstream calls, cache hit rate, and retry storms. User latency is a transaction property and can emerge from several healthy components interacting poorly.

If a migration leaves unused gateways, replicated databases, temporary identities, or duplicated storage after cutover, the project has a decommissioning gap. Transitional resources need an explicit retirement checklist so temporary architecture does not become permanent cost and risk.

If administrators repeatedly bypass the intended deployment process, investigate why. The process may be too slow, lack emergency paths, or fail to represent environment-specific configuration. Architecture governance should be usable enough that teams do not need undocumented workarounds.

End every troubleshooting exercise by retesting the original business transaction and updating the architecture record. A component-level fix is not enough if the end-to-end path still violates latency, security, recovery, or operational requirements.

If a private endpoint is created but traffic still uses the public path, inspect DNS resolution from the actual client network, linked private DNS zones, custom DNS forwarders, and routing. The endpoint resource alone does not guarantee that clients resolve or reach it correctly.

If backup restores succeed in tests but application recovery still misses the target, time the entire sequence including infrastructure, configuration, identity, DNS, validation, and user access. The recovery architecture should be measured end to end against the business RTO.

If an alert is assigned to a team that does not own the failing dependency, diagnosis will be slow even when telemetry is perfect. Update service ownership and escalation paths as part of the architecture correction.

If cost grows because logging volume increases, reduce noisy categories, adjust retention, archive where appropriate, or route different logs according to value. Do not remove the critical evidence needed for security and incident response simply to reduce storage charges.

If a migration design requires long-lived synchronization after cutover, question whether the transition ever truly ends. Persistent dual-write or replication can create permanent complexity. Define the event that retires the old path or consciously adopt the hybrid state as a supported architecture.

If one architecture team becomes the bottleneck for every operational change, improve templates, guardrails, self-service, and ownership. The architecture operating model should enable teams to make safe routine changes without bypassing governance or waiting for bespoke approval each time.

Use troubleshooting as a design feedback loop. Every incident reveals something about assumptions, observability, ownership, or failure boundaries. The best architect turns that evidence into a stronger standard pattern for the next workload rather than fixing only the affected resource.

The AZ-305 architecture mindset is improvement-oriented. Troubleshooting is complete when the original business path works again and the design is less likely to fail in the same way.