Spark Core interview questions

75 Spark Core questions from the Big Data bank, written for Indian campus drives and tech interviews. Every question has a verified answer and an AI-tutor explanation on placd.

Free to start: the 2-minute IT readiness check — six questions and a result.

Take the free IT readiness check

or take a mock interview set up for this area

1. What is Spark action?

Junior
  1. A.an operation that triggers execution and returns a result (collect, count, save)
  2. B.a shared write-only variable in Spark that workers add to and the driver reads, used for counters and metrics
  3. C.a set of tasks that can run without a shuffle, bounded by shuffle (wide) dependencies in the DAG
  4. D.a transformation needing no shuffle, where one input partition maps to one output partition

Answer + AI explanation with Pro

2. Which term means: "an operation that triggers execution and returns a result (collect, count, save)"?

Junior
  1. A.Executor
  2. B.Spark action
  3. C.Checkpointing
  4. D.Unified memory management

Answer + AI explanation with Pro

3. Which statement is correct?

Junior
  1. A.Spark action — the Spark process that runs the main function, creates the SparkContext, builds the DAG, and coordinates tasks across executors
  2. B.Spark action — an operation that triggers execution and returns a result (collect, count, save)
  3. C.Spark action — a long-running auxiliary process that serves shuffle blocks so executors can be removed under dynamic allocation without losing shuffle data
  4. D.Spark action — the basic unit of parallelism in Spark, a logical chunk of a dataset that one task processes

Answer + AI explanation with Pro

4. What is Spark transformation?

Junior
  1. A.truncating an RDD's lineage by saving it to reliable storage so recovery does not recompute a long chain of transformations
  2. B.the smallest unit of execution in Spark, operating on a single partition of data within a stage
  3. C.reduces the number of partitions without a full shuffle
  4. D.a lazy operation that returns a new RDD/DataFrame without executing (map, filter, join)

Answer + AI explanation with Pro

5. Which term means: "a lazy operation that returns a new RDD/DataFrame without executing (map, filter, join)"?

Junior
  1. A.Accumulator
  2. B.wide transformation
  3. C.Executor
  4. D.Spark transformation

Answer + AI explanation with Pro

6. Which statement is correct?

Junior
  1. A.Spark transformation — the Spark process that runs the main function, creates the SparkContext, builds the DAG, and coordinates tasks across executors
  2. B.Spark transformation — a lazy operation that returns a new RDD/DataFrame without executing (map, filter, join)
  3. C.Spark transformation — memory Spark manages outside the JVM heap (spark.memory.offHeap) to reduce garbage-collection overhead for execution and storage
  4. D.Spark transformation — a transformation requiring a shuffle, which creates a stage boundary

Answer + AI explanation with Pro

7. What is narrow transformation?

Junior
  1. A.a transformation requiring a shuffle, which creates a stage boundary
  2. B.the point created by a shuffle (wide dependency) that splits a job into stages
  3. C.memory Spark manages outside the JVM heap (spark.memory.offHeap) to reduce garbage-collection overhead for execution and storage
  4. D.a transformation needing no shuffle, where one input partition maps to one output partition

Answer + AI explanation with Pro

8. Which term means: "a transformation needing no shuffle, where one input partition maps to one output partition"?

Junior
  1. A.Speculative execution
  2. B.Off-heap memory
  3. C.narrow transformation
  4. D.coalesce()

Answer + AI explanation with Pro

9. Which statement is correct?

Junior
  1. A.narrow transformation — redistributes data into N partitions with a full shuffle
  2. B.narrow transformation — a transformation requiring a shuffle, which creates a stage boundary
  3. C.narrow transformation — a transformation needing no shuffle, where one input partition maps to one output partition
  4. D.narrow transformation — a lazy operation that returns a new RDD/DataFrame without executing (map, filter, join)

Answer + AI explanation with Pro

10. What is wide transformation?

Junior
  1. A.the process that builds the DAG and schedules tasks
  2. B.an operation that triggers execution and returns a result (collect, count, save)
  3. C.a transformation requiring a shuffle, which creates a stage boundary
  4. D.a shared write-only variable in Spark that workers add to and the driver reads, used for counters and metrics

Answer + AI explanation with Pro

11. Which term means: "a transformation requiring a shuffle, which creates a stage boundary"?

Junior
  1. A.Unified memory management
  2. B.Spark transformation
  3. C.wide transformation
  4. D.RDD lineage

Answer + AI explanation with Pro

12. Which statement is correct?

Junior
  1. A.wide transformation — truncating an RDD's lineage by saving it to reliable storage so recovery does not recompute a long chain of transformations
  2. B.wide transformation — a transformation requiring a shuffle, which creates a stage boundary
  3. C.wide transformation — reduces the number of partitions without a full shuffle
  4. D.wide transformation — a Spark mechanism that scales the number of executors up and down at runtime based on pending and running task load

Answer + AI explanation with Pro

13. What is RDD lineage?

Mid
  1. A.a Spark feature that re-launches duplicate copies of slow-running tasks (stragglers) and uses whichever finishes first
  2. B.the recorded chain of transformations used to recompute lost partitions
  3. C.redistributes data into N partitions with a full shuffle
  4. D.a long-running auxiliary process that serves shuffle blocks so executors can be removed under dynamic allocation without losing shuffle data

Answer + AI explanation with Pro

14. Which term means: "the recorded chain of transformations used to recompute lost partitions"?

Mid
  1. A.Off-heap memory
  2. B.Spark driver
  3. C.External shuffle service
  4. D.RDD lineage

Answer + AI explanation with Pro

15. Which statement is correct?

Mid
  1. A.RDD lineage — a lazy operation that returns a new RDD/DataFrame without executing (map, filter, join)
  2. B.RDD lineage — a transformation needing no shuffle, where one input partition maps to one output partition
  3. C.RDD lineage — the recorded chain of transformations used to recompute lost partitions
  4. D.RDD lineage — the model where execution and storage share one region and can borrow each other's space, with execution able to evict cached blocks

Answer + AI explanation with Pro

16. What is stage boundary?

Mid
  1. A.the point created by a shuffle (wide dependency) that splits a job into stages
  2. B.a JVM process launched on a worker node that runs tasks and holds cached data in memory or on disk for a single application
  3. C.the recorded chain of transformations used to recompute lost partitions
  4. D.a transformation needing no shuffle, where one input partition maps to one output partition

Answer + AI explanation with Pro

17. Which term means: "the point created by a shuffle (wide dependency) that splits a job into stages"?

Mid
  1. A.Barrier execution mode
  2. B.RDD lineage
  3. C.coalesce()
  4. D.stage boundary

Answer + AI explanation with Pro

18. Which statement is correct?

Mid
  1. A.stage boundary — a lazy operation that returns a new RDD/DataFrame without executing (map, filter, join)
  2. B.stage boundary — the point created by a shuffle (wide dependency) that splits a job into stages
  3. C.stage boundary — a long-running auxiliary process that serves shuffle blocks so executors can be removed under dynamic allocation without losing shuffle data
  4. D.stage boundary — a shared write-only variable in Spark that workers add to and the driver reads, used for counters and metrics

Answer + AI explanation with Pro

19. What is coalesce()?

Mid
  1. A.the recorded chain of transformations used to recompute lost partitions
  2. B.a transformation needing no shuffle, where one input partition maps to one output partition
  3. C.the Spark process that runs the main function, creates the SparkContext, builds the DAG, and coordinates tasks across executors
  4. D.reduces the number of partitions without a full shuffle

Answer + AI explanation with Pro

20. Which term means: "reduces the number of partitions without a full shuffle"?

Mid
  1. A.coalesce()
  2. B.Accumulator
  3. C.stage boundary
  4. D.Spark executor

Answer + AI explanation with Pro

21. Which statement is correct?

Mid
  1. A.coalesce() — an operation that triggers execution and returns a result (collect, count, save)
  2. B.coalesce() — memory Spark manages outside the JVM heap (spark.memory.offHeap) to reduce garbage-collection overhead for execution and storage
  3. C.coalesce() — a long-running auxiliary process that serves shuffle blocks so executors can be removed under dynamic allocation without losing shuffle data
  4. D.coalesce() — reduces the number of partitions without a full shuffle

Answer + AI explanation with Pro

22. What is repartition()?

Mid
  1. A.the process that builds the DAG and schedules tasks
  2. B.the process that runs tasks and holds cached data
  3. C.a Spark mechanism that scales the number of executors up and down at runtime based on pending and running task load
  4. D.redistributes data into N partitions with a full shuffle

Answer + AI explanation with Pro

23. Which term means: "redistributes data into N partitions with a full shuffle"?

Mid
  1. A.Spark executor
  2. B.Barrier execution mode
  3. C.repartition()
  4. D.stage boundary

Answer + AI explanation with Pro

24. Which statement is correct?

Mid
  1. A.repartition() — a transformation needing no shuffle, where one input partition maps to one output partition
  2. B.repartition() — redistributes data into N partitions with a full shuffle
  3. C.repartition() — a transformation requiring a shuffle, which creates a stage boundary
  4. D.repartition() — the basic unit of parallelism in Spark, a logical chunk of a dataset that one task processes

Answer + AI explanation with Pro

25. What is Spark driver?

Junior
  1. A.the process that builds the DAG and schedules tasks
  2. B.a JVM process launched on a worker node that runs tasks and holds cached data in memory or on disk for a single application
  3. C.a scheduling mode that launches all tasks of a stage together (gang scheduling) for distributed ML workloads like deep learning
  4. D.a long-running auxiliary process that serves shuffle blocks so executors can be removed under dynamic allocation without losing shuffle data

Answer + AI explanation with Pro

26. Which term means: "the process that builds the DAG and schedules tasks"?

Junior
  1. A.External shuffle service
  2. B.Spark action
  3. C.Broadcast variable
  4. D.Spark driver

Answer + AI explanation with Pro

27. Which statement is correct?

Junior
  1. A.Spark driver — the smallest unit of execution in Spark, operating on a single partition of data within a stage
  2. B.Spark driver — a long-running auxiliary process that serves shuffle blocks so executors can be removed under dynamic allocation without losing shuffle data
  3. C.Spark driver — the model where execution and storage share one region and can borrow each other's space, with execution able to evict cached blocks
  4. D.Spark driver — the process that builds the DAG and schedules tasks

Answer + AI explanation with Pro

28. What is Spark executor?

Junior
  1. A.the process that runs tasks and holds cached data
  2. B.a scheduling mode that launches all tasks of a stage together (gang scheduling) for distributed ML workloads like deep learning
  3. C.the point created by a shuffle (wide dependency) that splits a job into stages
  4. D.the recorded chain of transformations used to recompute lost partitions

Answer + AI explanation with Pro

29. Which term means: "the process that runs tasks and holds cached data"?

Junior
  1. A.narrow transformation
  2. B.Speculative execution
  3. C.Spark executor
  4. D.Spark action

Answer + AI explanation with Pro

30. Which statement is correct?

Junior
  1. A.Spark executor — the basic unit of parallelism in Spark, a logical chunk of a dataset that one task processes
  2. B.Spark executor — a read-only variable cached on each executor so a large lookup table is shipped once rather than with every task
  3. C.Spark executor — a lazy operation that returns a new RDD/DataFrame without executing (map, filter, join)
  4. D.Spark executor — the process that runs tasks and holds cached data

Answer + AI explanation with Pro

Showing 30 of 75 Spark Core questions — the full set, with answers, explanations and an AI tutor on every question, is inside.

Free to start

Start with a free readiness check

Sign up free for the 2-minute IT readiness check and a scored result. Answers, explanations and the AI tutor on every Spark Core question come with Pro.

Take the free IT readiness check

or take a mock interview set up for this area

24,000+ questions & coding problemsSoftware & IT16,274 questionsGovernment jobs26 examsAptitudenew questions every timeAI practice interviewwith feedback65 topics to practiseMechanical1,149 questionsGATE ME9 papersEngineering Mathematics381 questions2-minute checkfreeDSA Problems1,422Civil1,005 questionsGATE CE9 papersCS Fundamentals1,209 questionsYour scores6 skillsSystem Design25Electrical / EEE1,047 questionsGATE EE9 papersRun your codeC++ · Java · PythonLow-Level Design144Electronics & Comm.975 questionsGATE EC9 papersAI help on every questionFull-Stack6,282Chemical1,005 questionsGATE CH9 papersAI whiteboardsystem designWork abroadEurope · remote · transfersESE ME1 paperGATE practice papers2019–2026ESE CE1 paperDate alertsbefore the last dateESE EE1 paperBehavioural courseHR round practiceESE ET1 paperResume optimizerProSSC JE ME1 paperApplication trackerSSC JE CE1 paperCompany-wise prepSSC JE EE1 paperRole roadmapsRRB JE1 subjectPriced in ₹UPI · cardsISRO SC1 paperGATE CS9 papersIBPS SO IT1 paperUGC NET CS1 paperSSC CGL26 papersIBPS PO26 papersRRB NTPC26 papersSSC CHSL26 papersIBPS Clerk26 papersSBI Clerk26 papersRRB Group D26 papersSSC CPO26 papersSSC GD26 papers