CERTIFIED-ASSOCIATE-DEVELOPER-FOR-APACHE-SPARK-3-5-QUESTIONS practice question 8 of 25
A data engineer is running a Spark job to process a dataset of 1 TB stored in distributed storage. The cluster has 10 nodes, each with 16 CPUs. Spark…
Choose your answer, then check it against the explanation.
A data engineer is running a Spark job to process a dataset of 1 TB stored in distributed storage. The
cluster has 10 nodes, each with 16 CPUs. Spark UI shows:
Low number of Active Tasks
Many tasks complete in milliseconds
Fewer tasks than available CPUs
Which approach should be used to adjust the partitioning for optimal resource allocation?
Question 8 of 25
Keep practicing
Take the free CERTIFIED-ASSOCIATE-DEVELOPER-FOR-APACHE-SPARK-3-5-QUESTIONS practice test
Ten exam-style questions with answers and explanations, plus the exam facts and study guides.
More questions
Other CERTIFIED-ASSOCIATE-DEVELOPER-FOR-APACHE-SPARK-3-5-QUESTIONS practice questions
- Question 1You have: DataFrame A: 128 GB of transactions DataFrame B: 1 GB user lookup table Which s…
- Question 2A Spark DataFrame df is cached using the MEMORY_AND_DISK storage level, but the DataFrame…
- Question 344 of 55. A data engineer is working on a real-time analytics pipeline using Spark Struct…
- Question 7A developer wants to test Spark Connect with an existing Spark application. What are the …
- Question 10A data engineer is building a Structured Streaming pipeline and wants the pipeline to rec…
- Question 11Given a DataFrame df that has 10 partitions, after running the code: result = df.coalesce…
