117 Spark Tuning questions from the Big Data bank, written for Indian campus drives and tech interviews. Every question has a verified answer and an AI-tutor explanation on placd.
Free to start: the 2-minute IT readiness check — six questions and a result.
A.it spills to disk, serializes data and transfers it over the network
B.the garbage collector recommended for large Spark executor heaps, dividing the heap into regions to keep pause times bounded better than the older parallel collector
C.adding a random prefix to a hot key to spread skew across partitions
D.an uneven key distribution that makes one task a straggler
Answer + AI explanation with Pro
2. Which term means: "it spills to disk, serializes data and transfers it over the network"?
Mid
A.why a shuffle is expensive
B.spark.memory.fraction
C.shuffle tracking
D.Garbage collection tuning
Answer + AI explanation with Pro
3. Which statement is correct?
Mid
A.why a shuffle is expensive — an AQE feature that merges many small shuffle partitions into fewer right-sized ones based on runtime statistics
B.why a shuffle is expensive — a task that runs much slower than its peers in a stage and delays the whole stage from completing
C.why a shuffle is expensive — it spills to disk, serializes data and transfers it over the network
D.why a shuffle is expensive — writing an RDD or DataFrame to reliable storage to truncate its lineage, preventing stack-overflow and recomputation in long iterative or streaming jobs
Answer + AI explanation with Pro
4. What is spark.sql.shuffle.partitions?
Junior
A.the mechanism (spark.dynamicAllocation.shuffleTracking.enabled) that lets dynamic allocation release executors safely without a separate external shuffle service by tracking which executors hold shuffle blocks
B.the act of writing shuffle or aggregation data from memory to disk when it exceeds the available execution memory, slowing the job
C.the config capping the number of bytes packed into a single partition when reading files, defaulting to 128 MB and controlling read-side parallelism
5. Which term means: "the setting controlling post-shuffle partition count (default 200)"?
Junior
A.spark.sql.shuffle.partitions
B.Garbage collection tuning
C.dynamic allocation
D.why a shuffle is expensive
Answer + AI explanation with Pro
6. Which statement is correct?
Junior
A.spark.sql.shuffle.partitions — the AQE feature that detects skewed join partitions at runtime and splits them into smaller subpartitions to balance task time
B.spark.sql.shuffle.partitions — the config that sets how many concurrent task slots each Spark executor gets, commonly tuned to 4-5 to balance HDFS throughput against memory pressure
C.spark.sql.shuffle.partitions — the config capping the number of bytes packed into a single partition when reading files, defaulting to 128 MB and controlling read-side parallelism
D.spark.sql.shuffle.partitions — the setting controlling post-shuffle partition count (default 200)
Answer + AI explanation with Pro
7. What is data skew?
Mid
A.an uneven key distribution that makes one task a straggler
B.the garbage collector recommended for large Spark executor heaps, dividing the heap into regions to keep pause times bounded better than the older parallel collector
C.a skew-mitigation technique that appends a random prefix to hot keys to spread them across more reducers before aggregating
D.writing an RDD or DataFrame to reliable storage to truncate its lineage, preventing stack-overflow and recomputation in long iterative or streaming jobs
Answer + AI explanation with Pro
8. Which term means: "an uneven key distribution that makes one task a straggler"?
Mid
A.spark.executor.memory
B.dynamic allocation
C.partition sizing
D.data skew
Answer + AI explanation with Pro
9. Which statement is correct?
Mid
A.data skew — an uneven key distribution that makes one task a straggler
B.data skew — it spills to disk, serializes data and transfers it over the network
C.data skew — the AQE setting that merges small post-shuffle partitions at runtime so the output partition count matches the actual data volume
D.data skew — adding a random prefix to a hot key to spread skew across partitions
Answer + AI explanation with Pro
10. What is salting?
Mid
A.an uneven distribution of keys across partitions that overloads a few tasks and creates straggler-driven slowness
B.the act of writing shuffle or aggregation data from memory to disk when it exceeds the available execution memory, slowing the job
C.adding a random prefix to a hot key to spread skew across partitions
D.adjusting JVM collector and heap settings to reduce long GC pauses that stall Spark executors
Answer + AI explanation with Pro
11. Which term means: "adding a random prefix to a hot key to spread skew across partitions"?
Mid
A.salting
B.checkpointing
C.why a shuffle is expensive
D.spark.executor.memoryOverhead
Answer + AI explanation with Pro
12. Which statement is correct?
Mid
A.salting — adding a random prefix to a hot key to spread skew across partitions
B.salting — the config that sets the JVM heap size per Spark executor, with task working memory drawn from it after reserving overhead and storage
C.salting — a task that runs much slower than its peers in a stage and delays the whole stage from completing
D.salting — an AQE feature that merges many small shuffle partitions into fewer right-sized ones based on runtime statistics
Answer + AI explanation with Pro
13. What is autoBroadcastJoinThreshold?
Junior
A.the size limit (default 10MB) under which Spark auto-broadcasts a table
B.the AQE feature that detects skewed join partitions at runtime and splits them into smaller subpartitions to balance task time
C.a performance issue where too many tiny output files create excessive task and metadata overhead on reads
D.the AQE setting that merges small post-shuffle partitions at runtime so the output partition count matches the actual data volume
Answer + AI explanation with Pro
14. Which term means: "the size limit (default 10MB) under which Spark auto-broadcasts a table"?
Junior
A.spark.sql.files.maxPartitionBytes
B.autoBroadcastJoinThreshold
C.Bucketed join
D.Salting
Answer + AI explanation with Pro
15. Which statement is correct?
Junior
A.autoBroadcastJoinThreshold — the garbage collector recommended for large Spark executor heaps, dividing the heap into regions to keep pause times bounded better than the older parallel collector
B.autoBroadcastJoinThreshold — the LRU process by which Spark drops the least-recently-used cached blocks from storage memory when space is needed, with MEMORY_AND_DISK spilling rather than recomputing
C.autoBroadcastJoinThreshold — the practice of keeping each Spark partition near 128-200 MB so tasks are neither too small (scheduler overhead) nor too large (spilling and skew)
D.autoBroadcastJoinThreshold — the size limit (default 10MB) under which Spark auto-broadcasts a table
Answer + AI explanation with Pro
16. What is Data skew?
Junior
A.memory managed outside the JVM heap by Spark (enabled with spark.memory.offHeap.enabled and sized by spark.memory.offHeap.size) to reduce garbage collection pressure on large workloads
B.the AQE feature that detects skewed join partitions at runtime and splits them into smaller subpartitions to balance task time
C.an uneven distribution of keys across partitions that overloads a few tasks and creates straggler-driven slowness
D.a task that runs much slower than its peers in a stage and delays the whole stage from completing
Answer + AI explanation with Pro
17. Which term means: "an uneven distribution of keys across partitions that overloads a few tasks and creates straggler-driven slowness"?
Junior
A.localCheckpoint
B.Data skew
C.checkpointing
D.data skew
Answer + AI explanation with Pro
18. Which statement is correct?
Junior
A.Data skew — an uneven distribution of keys across partitions that overloads a few tasks and creates straggler-driven slowness
B.Data skew — the config reserving off-heap memory per executor for JVM internals, shuffle buffers, and Python/native processes, defaulting to about 10 percent of executor memory
C.Data skew — a skew-mitigation technique that appends a random prefix to hot keys to spread them across more reducers before aggregating
D.Data skew — a faster Spark checkpoint variant that truncates lineage using executor local storage instead of reliable storage, trading fault tolerance for speed
Answer + AI explanation with Pro
19. What is Salting?
Mid
A.a skew-mitigation technique that appends a random prefix to hot keys to spread them across more reducers before aggregating
B.a shuffle-free join made possible when both tables are pre-bucketed on the join key with the same number of buckets
C.the fraction of usable heap shared by execution and storage in Spark unified memory management, defaulting to 0.6
D.an uneven key distribution that makes one task a straggler
Answer + AI explanation with Pro
20. Which term means: "a skew-mitigation technique that appends a random prefix to hot keys to spread them across more reducers before aggregating"?
Mid
A.Spill
B.shuffle tracking
C.Bucketed join
D.Salting
Answer + AI explanation with Pro
21. Which statement is correct?
Mid
A.Salting — a skew-mitigation technique that appends a random prefix to hot keys to spread them across more reducers before aggregating
B.Salting — the garbage collector recommended for large Spark executor heaps, dividing the heap into regions to keep pause times bounded better than the older parallel collector
C.Salting — an uneven key distribution that makes one task a straggler
D.Salting — writing an RDD or DataFrame to reliable storage to truncate its lineage, preventing stack-overflow and recomputation in long iterative or streaming jobs
Answer + AI explanation with Pro
22. What is Spill?
Mid
A.an uneven key distribution that makes one task a straggler
B.the act of writing shuffle or aggregation data from memory to disk when it exceeds the available execution memory, slowing the job
C.the LRU process by which Spark drops the least-recently-used cached blocks from storage memory when space is needed, with MEMORY_AND_DISK spilling rather than recomputing
D.the config that sets how many concurrent task slots each Spark executor gets, commonly tuned to 4-5 to balance HDFS throughput against memory pressure
Answer + AI explanation with Pro
23. Which term means: "the act of writing shuffle or aggregation data from memory to disk when it exceeds the available execution memory, slowing the job"?
Mid
A.salting
B.off-heap memory
C.Spill
D.data skew
Answer + AI explanation with Pro
24. Which statement is correct?
Mid
A.Spill — adding a random prefix to a hot key to spread skew across partitions
B.Spill — the Spark feature (spark.speculation) that relaunches copies of slow-running stragglers on other executors and keeps whichever finishes first to limit tail latency
C.Spill — a faster Spark checkpoint variant that truncates lineage using executor local storage instead of reliable storage, trading fault tolerance for speed
D.Spill — the act of writing shuffle or aggregation data from memory to disk when it exceeds the available execution memory, slowing the job
Answer + AI explanation with Pro
25. What is Coalescing post-shuffle partitions?
Mid
A.a task that runs much slower than its peers in a stage and delays the whole stage from completing
B.an AQE feature that merges many small shuffle partitions into fewer right-sized ones based on runtime statistics
C.the Spark feature (spark.dynamicAllocation.enabled) that scales executors up and down at runtime based on pending and idle tasks, requiring an external or shuffle-tracking service to preserve shuffle data
D.a faster Spark checkpoint variant that truncates lineage using executor local storage instead of reliable storage, trading fault tolerance for speed
Answer + AI explanation with Pro
26. Which term means: "an AQE feature that merges many small shuffle partitions into fewer right-sized ones based on runtime statistics"?
Mid
A.localCheckpoint
B.Salting
C.Coalescing post-shuffle partitions
D.speculative execution
Answer + AI explanation with Pro
27. Which statement is correct?
Mid
A.Coalescing post-shuffle partitions — an AQE feature that merges many small shuffle partitions into fewer right-sized ones based on runtime statistics
B.Coalescing post-shuffle partitions — the size limit (default 10MB) under which Spark auto-broadcasts a table
C.Coalescing post-shuffle partitions — the practice of keeping each Spark partition near 128-200 MB so tasks are neither too small (scheduler overhead) nor too large (spilling and skew)
D.Coalescing post-shuffle partitions — a faster Spark checkpoint variant that truncates lineage using executor local storage instead of reliable storage, trading fault tolerance for speed
Answer + AI explanation with Pro
28. What is Small files problem?
Senior
A.an uneven distribution of keys across partitions that overloads a few tasks and creates straggler-driven slowness
B.the config reserving off-heap memory per executor for JVM internals, shuffle buffers, and Python/native processes, defaulting to about 10 percent of executor memory
C.the config that sets how many concurrent task slots each Spark executor gets, commonly tuned to 4-5 to balance HDFS throughput against memory pressure
D.a performance issue where too many tiny output files create excessive task and metadata overhead on reads
Answer + AI explanation with Pro
29. Which term means: "a performance issue where too many tiny output files create excessive task and metadata overhead on reads"?
Senior
A.Salting
B.Bucketed join
C.Small files problem
D.dynamic allocation
Answer + AI explanation with Pro
30. Which statement is correct?
Senior
A.Small files problem — a performance issue where too many tiny output files create excessive task and metadata overhead on reads
B.Small files problem — the mechanism (spark.dynamicAllocation.shuffleTracking.enabled) that lets dynamic allocation release executors safely without a separate external shuffle service by tracking which executors hold shuffle blocks
C.Small files problem — adding a random prefix to a hot key to spread skew across partitions
D.Small files problem — it spills to disk, serializes data and transfers it over the network
Answer + AI explanation with Pro
Showing 30 of 117 Spark Tuning questions — the full set, with answers, explanations and an AI tutor on every question, is inside.
Free to start
Start with a free readiness check
Sign up free for the 2-minute IT readiness check and a scored result. Answers, explanations and the AI tutor on every Spark Tuning question come with Pro.