{"id":17809,"date":"2026-09-21T11:35:55","date_gmt":"2026-09-21T11:35:55","guid":{"rendered":"https:\/\/www.examlabs.com\/certification\/?p=17809"},"modified":"2026-09-21T11:35:55","modified_gmt":"2026-09-21T11:35:55","slug":"databricks-certified-associate-developer-for-apache-spark-practice-test-questions-and-exam-dumps-part4-q61-80","status":"publish","type":"post","link":"https:\/\/www.examlabs.com\/certification\/databricks-certified-associate-developer-for-apache-spark-practice-test-questions-and-exam-dumps-part4-q61-80\/","title":{"rendered":"Databricks Certified Associate Developer for Apache Spark Practice Test Questions and Exam Dumps Part4 Q61-80"},"content":{"rendered":"<h2><b>View Full <\/b><a href=\"https:\/\/www.examlabs.com\/certified-associate-developer-for-apache-spark-exam-dumps\"><b>Databricks Certified Associate Developer for Apache Spark Exam Dumps <\/b><\/a><b>\u00a0and Practice Test Dumps<\/b><\/h2>\n<p>&nbsp;<\/p>\n<p><b>Question 61.<\/b><\/p>\n<p><b>Which Spark SQL function is used to return the first non-null value from a list of expressions?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> coalesce()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> collect_list()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> first()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> nvl2()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 1. coalesce()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">The coalesce() SQL function returns the first non-null expression from the arguments provided. It is commonly used to substitute fallback values or combine multiple nullable columns. This function is different from DataFrame.coalesce(), which changes the number of partitions. collect_list() aggregates values into an array, while first() returns the first value in an aggregation context. Using coalesce() is helpful when data contains multiple possible sources for the same logical field.<\/span><\/p>\n<p><b>Question 62.<\/b><\/p>\n<p><b>Which operation is most appropriate for changing the number of partitions upward from 4 to 20?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> coalesce(20)<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> repartition(20)<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> cache()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> persist()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 2. repartition(20)<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">repartition(20) is the appropriate operation when increasing the number of partitions. repartition() performs a shuffle to redistribute data across the requested number of partitions. coalesce() is generally intended for reducing partition counts while minimizing data movement and may not efficiently increase them. cache() and persist() control storage of computed results rather than partition count. Increasing partitions can improve parallelism when an existing dataset has too few partitions for the available cluster resources.<\/span><\/p>\n<p><b>Question 63.<\/b><\/p>\n<p><b>Which Spark function is commonly used to calculate the sum of a numeric column?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> avg()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> count()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> sum()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> max()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 3. sum()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">sum() calculates the total of a numeric column and is commonly used as an aggregation function. It can be applied across an entire DataFrame or within groups created by groupBy(). avg() calculates the arithmetic mean, count() counts rows or values, and max() returns the highest value. sum() is therefore the appropriate function for calculations such as total revenue, total units sold, or total processing time.<\/span><\/p>\n<p><b>Question 64.<\/b><\/p>\n<p><b>Which DataFrame method creates a temporary SQL view that can be queried using Spark SQL?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> cache()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> alias()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> checkpoint()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> createOrReplaceTempView()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 4. createOrReplaceTempView()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">createOrReplaceTempView() registers a DataFrame as a temporary view associated with the current SparkSession. Once registered, the view can be queried using spark.sql(). If a temporary view with the same name already exists, it is replaced. cache() stores data for reuse, alias() changes how a DataFrame or expression is referenced, and checkpoint() truncates lineage. Temporary views provide a convenient bridge between the DataFrame API and Spark SQL.<\/span><\/p>\n<p><b>Question 65.<\/b><\/p>\n<p><b>Which method is commonly used to execute a SQL query through a SparkSession?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> spark.sql()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> spark.query()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> spark.execute()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> spark.runSQL()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 1. spark.sql()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">spark.sql() executes a SQL statement using the active SparkSession and returns the result as a DataFrame when appropriate. It can query temporary views, tables, and other registered data sources available to the session. The other listed methods are not the standard SparkSession API for running SQL statements. spark.sql() is useful when developers prefer SQL syntax or need to combine SQL-based transformations with DataFrame operations in the same application.<\/span><\/p>\n<p><b>Question 66.<\/b><\/p>\n<p><b>Which window function assigns consecutive row numbers within each window partition?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> dense_rank()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> row_number()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> collect_list()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> first()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 2. row_number()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">row_number() assigns a unique sequential number to each row within a window partition according to the specified ordering. It is often used to identify the first or latest record per group by combining it with partitionBy() and orderBy(). dense_rank() assigns ranks while handling ties differently, collect_list() is an aggregation function, and first() returns a value rather than generating sequential row numbers. row_number() is widely used for deduplication and top-record selection.<\/span><\/p>\n<p><b>Question 67.<\/b><\/p>\n<p><b>Which function assigns the same rank to tied rows while leaving gaps after ties?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> row_number()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> dense_rank()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> rank()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> lag()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 3. rank()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">rank() assigns the same rank to rows with equal ordering values and leaves gaps in the ranking sequence after ties. For example, ranks may appear as 1, 2, 2, 4. dense_rank() also assigns equal ranks to ties but does not leave gaps, while row_number() assigns a unique sequence number to every row. lag() accesses a value from a previous row within the window. Understanding these differences is important when implementing analytical ranking logic.<\/span><\/p>\n<p><b>Question 68.<\/b><\/p>\n<p><b>Which window function returns a value from a preceding row in the same window?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> lead()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> rank()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> row_number()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> lag()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 4. lag()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">lag() returns the value of an expression from a preceding row within a defined window. It is useful for comparing a current record with a previous record, such as measuring changes in sales, timestamps, or balances. lead() accesses a following row, while rank() and row_number() generate ranking information. lag() depends on the ordering defined in the window specification, so a meaningful order should be provided.<\/span><\/p>\n<p><b>Question 69.<\/b><\/p>\n<p><b>Which window function is used to access a value from a following row?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> lead()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> lag()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> rank()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> dense_rank()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 1. lead()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">lead() returns a value from a subsequent row within the same window partition based on the defined ordering. It is useful when comparing the current record with the next event, transaction, timestamp, or state. lag() performs the opposite operation by accessing a preceding row. rank() and dense_rank() assign ranking values rather than retrieve neighboring records. lead() is therefore the appropriate function when forward-looking row comparison is required.<\/span><\/p>\n<p><b>Question 70.<\/b><\/p>\n<p><b>Which method defines the grouping columns for a Spark window specification?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> orderBy()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> partitionBy()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> groupBy()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> repartition()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 2. partitionBy()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">Within a Window specification, partitionBy() defines how rows are divided into independent groups for window calculations. Each partition is processed separately for functions such as row_number(), rank(), lag(), or running totals. orderBy() defines ordering within those partitions, while groupBy() performs aggregation and changes the result structure. DataFrame.repartition() changes physical data partitions and is conceptually different from logical window partitioning.<\/span><\/p>\n<p><b>Question 71.<\/b><\/p>\n<p><b>Which Spark operation produces the Cartesian product of two DataFrames?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> left join<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> union()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> crossJoin()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> left semi join<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 3. crossJoin()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">crossJoin() produces the Cartesian product of two DataFrames, meaning every row from the first DataFrame is paired with every row from the second. The resulting number of rows can become extremely large, so this operation should be used carefully. A left join matches rows according to a condition, union() appends rows vertically, and a left semi join returns matching rows from the left dataset. crossJoin() is appropriate only when all pairwise combinations are intentionally required.<\/span><\/p>\n<p><b>Question 72.<\/b><\/p>\n<p><b>Which function can create a map column from alternating key and value expressions?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> array()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> struct()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> explode()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> create_map()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 4. create_map()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">create_map() builds a map column from alternating key and value expressions. This is useful when related key-value data should be represented within a single nested column. array() creates an array, struct() creates a struct, and explode() expands array or map elements into separate rows. Map columns are useful for semi-structured data and can be accessed using key-based expressions during later transformations.<\/span><\/p>\n<p><b>Question 73.<\/b><\/p>\n<p><b>Which function returns the number of elements in an array or map column?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> size()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> length()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> count()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> cardinalityBy()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 1. size()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">size() returns the number of elements contained in an array or map expression. It is useful for filtering or validating nested collections, such as identifying rows with empty arrays or unusually large collections. length() is primarily used for strings or binary values, while count() is an aggregation function. size() is therefore the correct function when the number of elements inside an array or map needs to be measured.<\/span><\/p>\n<p><b>Question 74.<\/b><\/p>\n<p><b>Which function aggregates values from multiple rows into an array while retaining duplicates?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> collect_set()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> collect_list()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> array()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> explode()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 2. collect_list()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">collect_list() gathers values from multiple rows into an array and retains duplicate values. It is commonly used after groupBy() when all values belonging to each group should be collected. collect_set() also creates an array but removes duplicate values. array() combines expressions within an individual row, while explode() expands array elements into multiple rows. collect_list() is appropriate when duplicate values are meaningful and should be preserved.<\/span><\/p>\n<p><b>Question 75.<\/b><\/p>\n<p><b>Which function aggregates values into an array while removing duplicates?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> concat_ws()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> collect_list()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> collect_set()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> array_repeat()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 3. collect_set()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">collect_set() aggregates values into an array while eliminating duplicates. This is useful when only unique values associated with each group are needed. collect_list() retains duplicates, concat_ws() combines strings with a separator, and array_repeat() creates an array containing repeated instances of the same value. The ordering of values returned by collect_set() should generally not be relied upon unless additional sorting logic is applied.<\/span><\/p>\n<p><b>Question 76.<\/b><\/p>\n<p><b>Which DataFrame method is commonly used to rename multiple columns through SQL-style expressions in a single selection?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> cache()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> dropDuplicates()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> coalesce()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> selectExpr()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 4. selectExpr()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">selectExpr() accepts SQL-style expressions as strings and can conveniently rename multiple columns using aliases such as &#8220;old_name AS new_name&#8221;. It can also perform calculations and casts within the same selection. cache() stores a DataFrame for reuse, dropDuplicates() removes duplicate rows, and coalesce() commonly reduces partition count. selectExpr() is especially useful when developers want concise SQL-like syntax for projections and transformations.<\/span><\/p>\n<p><b>Question 77.<\/b><\/p>\n<p><b>Which expression is most appropriate for replacing values based on multiple conditional branches?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> when().otherwise()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> broadcast()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> explode()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> checkpoint()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 1. when().otherwise()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">when() combined with otherwise() provides conditional logic similar to SQL CASE expressions. Multiple when() calls can be chained to handle several branches before specifying a default value with otherwise(). broadcast() is used for join optimization, explode() expands nested collections, and checkpoint() truncates lineage. Conditional expressions are commonly used to classify records, create status columns, normalize categories, or apply business rules directly within DataFrame transformations.<\/span><\/p>\n<p><b>Question 78.<\/b><\/p>\n<p><b>Which persistence level concept determines where cached Spark data can be stored?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> Query hint<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> Storage level<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> Join type<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> Window frame<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 2. Storage level<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">A storage level defines how persisted data is stored, such as in memory, on disk, or using combinations and serialization options depending on the selected level. persist() can accept an explicit storage level, while cache() uses a default caching behavior. Query hints influence optimization, join types control relational combination semantics, and window frames define subsets of rows for window calculations. Choosing an appropriate storage level can help balance speed, memory consumption, and recomputation cost.<\/span><\/p>\n<p><b>Question 79.<\/b><\/p>\n<p><b>Which method removes a DataFrame from cache or persisted storage?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> clear()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> release()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> unpersist()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> drop()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 3. unpersist()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">unpersist() removes cached or persisted blocks associated with a Dataset or DataFrame, allowing those storage resources to be reclaimed. It is useful when a cached dataset is no longer needed and memory should be made available for other workloads. drop() removes columns rather than cached storage, and clear() or release() are not the standard DataFrame methods for this purpose. Properly unpersisting large reusable datasets can help manage executor memory efficiently.<\/span><\/p>\n<p><b>Question 80.<\/b><\/p>\n<p><b>Which operation is generally most appropriate when reducing a DataFrame from many partitions to a much smaller number while minimizing shuffle overhead?<\/b><\/p>\n<ol>\n<li><b><\/b><span style=\"font-weight: 400;\"> collect()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>2.<\/b><span style=\"font-weight: 400;\"> groupBy()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>3.<\/b><span style=\"font-weight: 400;\"> repartition()<\/span><span style=\"font-weight: 400;\"><br \/>\n<\/span><b>4.<\/b><span style=\"font-weight: 400;\"> coalesce()<\/span><\/li>\n<\/ol>\n<p><b>Correct Answer: 4. coalesce()<\/b><\/p>\n<p><b>Explanation:<\/b><\/p>\n<p><span style=\"font-weight: 400;\">coalesce() is generally appropriate when reducing the number of DataFrame partitions while attempting to minimize data movement. It can combine existing partitions without requiring the same full redistribution typically associated with repartition(). repartition() may produce better-balanced partitions but usually requires a shuffle. collect() sends data to the driver, while groupBy() performs grouping and typically introduces a shuffle for a different purpose. coalesce() is therefore useful when efficiently decreasing partition count.<\/span><\/p>\n<p>&nbsp;<\/p>\n","protected":false},"excerpt":{"rendered":"<p>View Full Databricks Certified Associate Developer for Apache Spark Exam Dumps \u00a0and Practice Test Dumps &nbsp; Question 61. Which Spark SQL function is used to return the first non-null value from a list of expressions? coalesce() 2. collect_list() 3. first() 4. nvl2() Correct Answer: 1. coalesce() Explanation: The coalesce() SQL function returns the first non-null [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":[],"categories":[1648,1647],"tags":[],"_links":{"self":[{"href":"https:\/\/www.examlabs.com\/certification\/wp-json\/wp\/v2\/posts\/17809"}],"collection":[{"href":"https:\/\/www.examlabs.com\/certification\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.examlabs.com\/certification\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.examlabs.com\/certification\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.examlabs.com\/certification\/wp-json\/wp\/v2\/comments?post=17809"}],"version-history":[{"count":1,"href":"https:\/\/www.examlabs.com\/certification\/wp-json\/wp\/v2\/posts\/17809\/revisions"}],"predecessor-version":[{"id":17810,"href":"https:\/\/www.examlabs.com\/certification\/wp-json\/wp\/v2\/posts\/17809\/revisions\/17810"}],"wp:attachment":[{"href":"https:\/\/www.examlabs.com\/certification\/wp-json\/wp\/v2\/media?parent=17809"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.examlabs.com\/certification\/wp-json\/wp\/v2\/categories?post=17809"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.examlabs.com\/certification\/wp-json\/wp\/v2\/tags?post=17809"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}