DP-750 is Microsoft’s intermediate Azure Databricks data-engineering exam, and its current scope is much broader than notebook-based transformation. The certification expects candidates to integrate and model data, select and configure compute, secure Unity Catalog objects, build batch and streaming ingestion, create production pipelines, apply data-quality controls, use Git and Databricks Asset Bundles, and troubleshoot performance and cost across live workloads.
As of October 3, 2026, the live English exam still uses the skills measured as of March 11, 2026. Microsoft has already announced an English-language update for October 19, 2026, but that date is still in the future. Candidates preparing now should therefore use the current DP-750 scope while also being aware that some objective wording will change later in October. The four current domains are environment setup at 15–20%, Unity Catalog security and governance at 15–20%, data preparation and processing at 30–35%, and pipeline/workload deployment and maintenance at 30–35%.
The certification is built around production data engineering, not isolated Spark knowledge
Microsoft’s audience profile expects experience with SQL, Python, Git, Microsoft Entra, Azure Data Factory, Azure Monitor, and Azure Databricks. That combination matters because the role is not just “write transformations.” The engineer has to choose compute, govern access, ingest from several kinds of sources, manage data quality, deploy repeatable workloads, and operate them after release.
The broader Microsoft certification catalog contains many data and cloud roles, but DP-750 is specifically tied to Microsoft Certified: Azure Databricks Data Engineer Associate. A separate Databricks certifications track can provide platform context, yet the Microsoft exam adds Azure-specific identity, monitoring, orchestration, and certification expectations.
Environment setup begins with choosing the right compute for the workload
The first domain includes job compute, serverless, SQL warehouses, classic compute, and shared compute, along with CPU, node count, autoscaling, termination, node type, pooling, Photon, runtime versions, libraries, and permissions. The exam is therefore likely to present workload requirements rather than simply ask which compute types exist.
Study compute as a trade-off among workload type, isolation, performance, cost, startup behavior, governance, and operational simplicity. A short scheduled job has different needs from an interactive development session or a persistent SQL analytics workload. The right answer follows the workload, not a universal preference for the newest option.
Unity Catalog object design sits beside compute in the environment domain
Microsoft includes catalogs, schemas, volumes, tables, views, materialized views, external catalogs, DDL operations on managed and external tables, naming conventions, and even AI/BI Genie instructions for data discovery. This shows that environment setup is also an information-architecture task.
Practice deciding how environments, teams, domains, and sharing requirements should be reflected in the catalog hierarchy. Object placement affects discoverability, security inheritance, lineage, and deployment. A table is not just a data structure; it lives inside a governed namespace.
Security and governance are separate because permission and policy solve different problems
Security objectives include privileges for users, groups, and service principals, row- and column-level controls, Azure Key Vault secrets, service principals, and managed identities. Governance objectives add descriptions, tags, ABAC policies, row filters, column masks, retention, lineage, audit logging, and Delta Sharing.
This domain should be studied as “who can do what?” plus “how do we prove, classify, retain, and share data correctly?” A user can technically have permission to query a table while the organization still needs masking, retention, lineage, or audit controls around that access. Governance adds context and policy to raw permissions.
Data modeling and ingestion make up the largest preparation area
The 30–35% data-preparation domain begins with ingestion design: extraction type, file type, ingestion tool, batch versus streaming, table format, partitioning, slowly changing dimensions, granularity, temporal history, clustering, and managed versus unmanaged tables. Those choices define how data will behave long after the first load.
The current blueprint names Lakeflow Connect, notebooks, Azure Data Factory, CTAS, CREATE OR REPLACE TABLE, COPY INTO, CDC feeds, Structured Streaming, Azure Event Hubs, Lakeflow Spark Declarative Pipelines, and Auto Loader. A review of Azure Data Factory can reinforce when external orchestration or ingestion makes sense, while Databricks data-engineering material can strengthen core lakehouse concepts.
Transformations are tested together with quality, not as a separate cleanup phase
Microsoft includes profiling, data types, duplicate and null handling, filtering, grouping, aggregation, joins, unions, intersects, excepts, denormalization, pivots, unpivots, and merge/insert/append operations. It then immediately connects those transformations to nullability, cardinality, range checks, type checks, schema enforcement, schema drift, and pipeline expectations.
This reflects production reality. A transformation is not complete when the code runs; it is complete when the output satisfies the contract expected by downstream consumers. Candidates should be able to decide whether bad data should fail a pipeline, be quarantined, be corrected, or be tolerated under a documented rule.
Pipeline design and Lakeflow Jobs form the second major 30–35% block
The deployment domain covers task ordering, notebook versus declarative pipeline choices, Lakeflow Jobs, error handling, precedence, triggers, schedules, alerts, and automatic restarts. The important theme is orchestration. A data engineer must turn individual transformations into a dependable workflow with explicit dependencies and failure behavior.
Study pipelines by drawing task graphs. Which step must finish first? What data does the next task depend on? What should happen after a partial failure? Can the failed task be repaired without rerunning everything? Those questions are more useful than memorizing every scheduling option.
Git, testing, and Asset Bundles make data engineering a software lifecycle discipline
Microsoft explicitly expects Git best practices, branching, pull requests, conflict resolution, unit/integration/end-to-end/UAT testing, Databricks Asset Bundles, CLI deployment, and REST API deployment. The exam therefore treats data workloads as versioned software rather than as scripts owned by one notebook author.
A review of Git fundamentals and CI/CD principles can be useful if those practices are weak. For DP-750, connect them directly to notebooks, jobs, pipelines, infrastructure definitions, tests, and environment-specific configuration. The goal is repeatable promotion, not merely source control for code snippets.
Operations and optimization close the lifecycle
The blueprint ends with cluster-consumption monitoring, Lakeflow job repair/restart/stop/run functions, Spark job and notebook troubleshooting, resource bottlenecks, cluster restart, caching, skew, spilling, shuffle, DAG and Spark UI analysis, query profiles, OPTIMIZE, VACUUM, Log Analytics, and Azure Monitor alerts.
This is where Azure Databricks operational knowledge becomes especially useful. A pipeline that is logically correct can still be too slow or expensive. Candidates should know how to distinguish poor partitioning, data skew, excessive shuffle, bad compute sizing, stale caches, or inefficient table layout from functional errors.
The current exam is a lifecycle: configure, govern, ingest, transform, deploy, observe
The most useful final mental model is one end-to-end dataset. Decide where it arrives, which compute processes it, how it is represented in Unity Catalog, who can access it, how quality is enforced, which transformations create downstream tables, how the pipeline is scheduled and deployed, and what evidence proves it remains healthy.
Microsoft’s current certification page also makes the operational expectation clear by assigning 120 minutes to the proctored exam and describing the credential as intermediate. The label “intermediate” should not be mistaken for narrow scope. The candidate is expected to move comfortably between data engineering, cloud identity, software lifecycle, governance, and workload operations. A person who knows Spark well but has never worked with service principals, Git-based deployment, or Azure Monitor will still have visible gaps.
The announced October 19 update deserves practical planning. Microsoft already publishes a change log showing no change to the audience profile and major skill areas, with minor changes to data modeling and lifecycle objectives. That means candidates sitting before the update should not abandon the March blueprint, while those testing after October 19 should re-open the English study guide and compare the detailed bullets. The underlying role is stable even if some feature wording changes.
Lakeflow terminology appears throughout the live objectives, so candidates using older Databricks material should normalize product names. Older resources may discuss Delta Live Tables or different job and ingestion labels. The concept can still be useful, but the exam answer should reflect the current Microsoft terminology and service boundaries. This is especially important around Lakeflow Connect, Lakeflow Spark Declarative Pipelines, and Lakeflow Jobs.
Data sharing is another cross-domain topic. Delta Sharing appears under governance, but its design depends on what data is stored, who owns it, which external consumer needs it, and how the organization protects sensitive information. Treat it as a governed delivery method rather than a simple export feature. A share that is easy to create but exposes excessive columns or stale data is not a sound solution.
AI/BI Genie instructions are a small current objective, but they illustrate a broader theme: data engineering increasingly supports discoverability and downstream AI-assisted analysis. The engineer’s responsibility is still to create clear, governed data objects with accurate descriptions and semantics. Good metadata improves both human and AI-assisted discovery, while poor metadata can make a technically correct platform harder to use.
The current role also assumes collaboration. Administrators, architects, data scientists, and analysts may own adjacent concerns, but the data engineer is responsible for handing off governed, reliable datasets and workloads. That makes documentation, naming, lineage, testing, and alerting part of engineering quality rather than optional polish.
That role breadth is why narrowly memorized feature lists are not enough.
Operational breadth is the point.
That breadth matters.
Because Microsoft has announced a future October 19 update, candidates close to that date should recheck the study guide before sitting the exam. As of October 3, however, the March 11 skills remain the live blueprint, and preparation should prioritize the two 30–35% engineering domains while still treating compute, governance, identity, and lifecycle controls as essential parts of the same production system.