75 Spark Core questions from the Big Data bank, written for Indian campus drives and tech interviews. Every question has a verified answer and an AI-tutor explanation on placd.
Free to start: the 2-minute IT readiness check — six questions and a result.
A.an operation that triggers execution and returns a result (collect, count, save)
B.a shared write-only variable in Spark that workers add to and the driver reads, used for counters and metrics
C.a set of tasks that can run without a shuffle, bounded by shuffle (wide) dependencies in the DAG
D.a transformation needing no shuffle, where one input partition maps to one output partition
Answer + AI explanation with Pro
2. Which term means: "an operation that triggers execution and returns a result (collect, count, save)"?
Junior
A.Executor
B.Spark action
C.Checkpointing
D.Unified memory management
Answer + AI explanation with Pro
3. Which statement is correct?
Junior
A.Spark action — the Spark process that runs the main function, creates the SparkContext, builds the DAG, and coordinates tasks across executors
B.Spark action — an operation that triggers execution and returns a result (collect, count, save)
C.Spark action — a long-running auxiliary process that serves shuffle blocks so executors can be removed under dynamic allocation without losing shuffle data
D.Spark action — the basic unit of parallelism in Spark, a logical chunk of a dataset that one task processes
Answer + AI explanation with Pro
4. What is Spark transformation?
Junior
A.truncating an RDD's lineage by saving it to reliable storage so recovery does not recompute a long chain of transformations
B.the smallest unit of execution in Spark, operating on a single partition of data within a stage
C.reduces the number of partitions without a full shuffle
D.a lazy operation that returns a new RDD/DataFrame without executing (map, filter, join)
Answer + AI explanation with Pro
5. Which term means: "a lazy operation that returns a new RDD/DataFrame without executing (map, filter, join)"?
Junior
A.Accumulator
B.wide transformation
C.Executor
D.Spark transformation
Answer + AI explanation with Pro
6. Which statement is correct?
Junior
A.Spark transformation — the Spark process that runs the main function, creates the SparkContext, builds the DAG, and coordinates tasks across executors
B.Spark transformation — a lazy operation that returns a new RDD/DataFrame without executing (map, filter, join)
C.Spark transformation — memory Spark manages outside the JVM heap (spark.memory.offHeap) to reduce garbage-collection overhead for execution and storage
D.Spark transformation — a transformation requiring a shuffle, which creates a stage boundary
Answer + AI explanation with Pro
7. What is narrow transformation?
Junior
A.a transformation requiring a shuffle, which creates a stage boundary
B.the point created by a shuffle (wide dependency) that splits a job into stages
C.memory Spark manages outside the JVM heap (spark.memory.offHeap) to reduce garbage-collection overhead for execution and storage
D.a transformation needing no shuffle, where one input partition maps to one output partition
Answer + AI explanation with Pro
8. Which term means: "a transformation needing no shuffle, where one input partition maps to one output partition"?
Junior
A.Speculative execution
B.Off-heap memory
C.narrow transformation
D.coalesce()
Answer + AI explanation with Pro
9. Which statement is correct?
Junior
A.narrow transformation — redistributes data into N partitions with a full shuffle
B.narrow transformation — a transformation requiring a shuffle, which creates a stage boundary
C.narrow transformation — a transformation needing no shuffle, where one input partition maps to one output partition
D.narrow transformation — a lazy operation that returns a new RDD/DataFrame without executing (map, filter, join)
Answer + AI explanation with Pro
10. What is wide transformation?
Junior
A.the process that builds the DAG and schedules tasks
B.an operation that triggers execution and returns a result (collect, count, save)
C.a transformation requiring a shuffle, which creates a stage boundary
D.a shared write-only variable in Spark that workers add to and the driver reads, used for counters and metrics
Answer + AI explanation with Pro
11. Which term means: "a transformation requiring a shuffle, which creates a stage boundary"?
Junior
A.Unified memory management
B.Spark transformation
C.wide transformation
D.RDD lineage
Answer + AI explanation with Pro
12. Which statement is correct?
Junior
A.wide transformation — truncating an RDD's lineage by saving it to reliable storage so recovery does not recompute a long chain of transformations
B.wide transformation — a transformation requiring a shuffle, which creates a stage boundary
C.wide transformation — reduces the number of partitions without a full shuffle
D.wide transformation — a Spark mechanism that scales the number of executors up and down at runtime based on pending and running task load
Answer + AI explanation with Pro
13. What is RDD lineage?
Mid
A.a Spark feature that re-launches duplicate copies of slow-running tasks (stragglers) and uses whichever finishes first
B.the recorded chain of transformations used to recompute lost partitions
C.redistributes data into N partitions with a full shuffle
D.a long-running auxiliary process that serves shuffle blocks so executors can be removed under dynamic allocation without losing shuffle data
Answer + AI explanation with Pro
14. Which term means: "the recorded chain of transformations used to recompute lost partitions"?
Mid
A.Off-heap memory
B.Spark driver
C.External shuffle service
D.RDD lineage
Answer + AI explanation with Pro
15. Which statement is correct?
Mid
A.RDD lineage — a lazy operation that returns a new RDD/DataFrame without executing (map, filter, join)
B.RDD lineage — a transformation needing no shuffle, where one input partition maps to one output partition
C.RDD lineage — the recorded chain of transformations used to recompute lost partitions
D.RDD lineage — the model where execution and storage share one region and can borrow each other's space, with execution able to evict cached blocks
Answer + AI explanation with Pro
16. What is stage boundary?
Mid
A.the point created by a shuffle (wide dependency) that splits a job into stages
B.a JVM process launched on a worker node that runs tasks and holds cached data in memory or on disk for a single application
C.the recorded chain of transformations used to recompute lost partitions
D.a transformation needing no shuffle, where one input partition maps to one output partition
Answer + AI explanation with Pro
17. Which term means: "the point created by a shuffle (wide dependency) that splits a job into stages"?
Mid
A.Barrier execution mode
B.RDD lineage
C.coalesce()
D.stage boundary
Answer + AI explanation with Pro
18. Which statement is correct?
Mid
A.stage boundary — a lazy operation that returns a new RDD/DataFrame without executing (map, filter, join)
B.stage boundary — the point created by a shuffle (wide dependency) that splits a job into stages
C.stage boundary — a long-running auxiliary process that serves shuffle blocks so executors can be removed under dynamic allocation without losing shuffle data
D.stage boundary — a shared write-only variable in Spark that workers add to and the driver reads, used for counters and metrics
Answer + AI explanation with Pro
19. What is coalesce()?
Mid
A.the recorded chain of transformations used to recompute lost partitions
B.a transformation needing no shuffle, where one input partition maps to one output partition
C.the Spark process that runs the main function, creates the SparkContext, builds the DAG, and coordinates tasks across executors
D.reduces the number of partitions without a full shuffle
Answer + AI explanation with Pro
20. Which term means: "reduces the number of partitions without a full shuffle"?
Mid
A.coalesce()
B.Accumulator
C.stage boundary
D.Spark executor
Answer + AI explanation with Pro
21. Which statement is correct?
Mid
A.coalesce() — an operation that triggers execution and returns a result (collect, count, save)
B.coalesce() — memory Spark manages outside the JVM heap (spark.memory.offHeap) to reduce garbage-collection overhead for execution and storage
C.coalesce() — a long-running auxiliary process that serves shuffle blocks so executors can be removed under dynamic allocation without losing shuffle data
D.coalesce() — reduces the number of partitions without a full shuffle
Answer + AI explanation with Pro
22. What is repartition()?
Mid
A.the process that builds the DAG and schedules tasks
B.the process that runs tasks and holds cached data
C.a Spark mechanism that scales the number of executors up and down at runtime based on pending and running task load
D.redistributes data into N partitions with a full shuffle
Answer + AI explanation with Pro
23. Which term means: "redistributes data into N partitions with a full shuffle"?
Mid
A.Spark executor
B.Barrier execution mode
C.repartition()
D.stage boundary
Answer + AI explanation with Pro
24. Which statement is correct?
Mid
A.repartition() — a transformation needing no shuffle, where one input partition maps to one output partition
B.repartition() — redistributes data into N partitions with a full shuffle
C.repartition() — a transformation requiring a shuffle, which creates a stage boundary
D.repartition() — the basic unit of parallelism in Spark, a logical chunk of a dataset that one task processes
Answer + AI explanation with Pro
25. What is Spark driver?
Junior
A.the process that builds the DAG and schedules tasks
B.a JVM process launched on a worker node that runs tasks and holds cached data in memory or on disk for a single application
C.a scheduling mode that launches all tasks of a stage together (gang scheduling) for distributed ML workloads like deep learning
D.a long-running auxiliary process that serves shuffle blocks so executors can be removed under dynamic allocation without losing shuffle data
Answer + AI explanation with Pro
26. Which term means: "the process that builds the DAG and schedules tasks"?
Junior
A.External shuffle service
B.Spark action
C.Broadcast variable
D.Spark driver
Answer + AI explanation with Pro
27. Which statement is correct?
Junior
A.Spark driver — the smallest unit of execution in Spark, operating on a single partition of data within a stage
B.Spark driver — a long-running auxiliary process that serves shuffle blocks so executors can be removed under dynamic allocation without losing shuffle data
C.Spark driver — the model where execution and storage share one region and can borrow each other's space, with execution able to evict cached blocks
D.Spark driver — the process that builds the DAG and schedules tasks
Answer + AI explanation with Pro
28. What is Spark executor?
Junior
A.the process that runs tasks and holds cached data
B.a scheduling mode that launches all tasks of a stage together (gang scheduling) for distributed ML workloads like deep learning
C.the point created by a shuffle (wide dependency) that splits a job into stages
D.the recorded chain of transformations used to recompute lost partitions
Answer + AI explanation with Pro
29. Which term means: "the process that runs tasks and holds cached data"?
Junior
A.narrow transformation
B.Speculative execution
C.Spark executor
D.Spark action
Answer + AI explanation with Pro
30. Which statement is correct?
Junior
A.Spark executor — the basic unit of parallelism in Spark, a logical chunk of a dataset that one task processes
B.Spark executor — a read-only variable cached on each executor so a large lookup table is shipped once rather than with every task
C.Spark executor — a lazy operation that returns a new RDD/DataFrame without executing (map, filter, join)
D.Spark executor — the process that runs tasks and holds cached data
Answer + AI explanation with Pro
Showing 30 of 75 Spark Core questions — the full set, with answers, explanations and an AI tutor on every question, is inside.
Free to start
Start with a free readiness check
Sign up free for the 2-minute IT readiness check and a scored result. Answers, explanations and the AI tutor on every Spark Core question come with Pro.