Warehousing interview questions

78 Warehousing questions from the Data Engineering bank, written for Indian campus drives and tech interviews. Every question has a verified answer and an AI-tutor explanation on placd.

Free to start: the 2-minute IT readiness check — six questions and a result.

Take the free IT readiness check

or take a mock interview set up for this area

1. What is Snowflake micro-partition?

Mid
  1. A.an automatic ~50-500MB columnar unit with min/max stats for pruning
  2. B.A cloud-warehouse design where data sits in shared storage and elastic compute scales independently, so you pay for each separately.
  3. C.Massively parallel processing, an architecture that distributes data and query work across many nodes that compute in parallel, the basis of cloud warehouses.
  4. D.Redshift Workload Management, which routes queries into queues with memory and concurrency settings; automatic WLM tunes this dynamically.

Answer + AI explanation with Pro

2. Which term means: "an automatic ~50-500MB columnar unit with min/max stats for pruning"?

Mid
  1. A.Snowflake micro-partition
  2. B.Redshift distribution style KEY
  3. C.MPP
  4. D.BigQuery clustering

Answer + AI explanation with Pro

3. Which statement is correct?

Mid
  1. A.Snowflake micro-partition — an independent compute cluster, separate from storage
  2. B.Snowflake micro-partition — A fact table that records the occurrence of an event or a coverage relationship with no numeric measures, used for counting events or capturing eligibility.
  3. C.Snowflake micro-partition — copying a small table to every node to avoid redistribution
  4. D.Snowflake micro-partition — an automatic ~50-500MB columnar unit with min/max stats for pruning

Answer + AI explanation with Pro

4. What is Snowflake virtual warehouse?

Junior
  1. A.an independent compute cluster, separate from storage
  2. B.splitting a table by date/integer range to cut bytes scanned
  3. C.sorting data within partitions by up to four columns
  4. D.A fact table that records the occurrence of an event or a coverage relationship with no numeric measures, used for counting events or capturing eligibility.

Answer + AI explanation with Pro

5. Which term means: "an independent compute cluster, separate from storage"?

Junior
  1. A.RA3 node
  2. B.Snowflake virtual warehouse
  3. C.cost-based optimizer
  4. D.MPP

Answer + AI explanation with Pro

6. Which statement is correct?

Junior
  1. A.Snowflake virtual warehouse — A fact table tracking a process with multiple milestones, where rows are updated as each step completes, used for pipelines like order fulfillment.
  2. B.Snowflake virtual warehouse — a virtual CPU unit BigQuery uses to execute SQL
  3. C.Snowflake virtual warehouse — an independent compute cluster, separate from storage
  4. D.Snowflake virtual warehouse — Auto-maintained min/max metadata per block (Redshift) or micro-partition that lets the engine skip blocks not matching a filter.

Answer + AI explanation with Pro

7. What is Snowflake Time Travel?

Junior
  1. A.querying or restoring historical table data within a retention window
  2. B.an automatic ~50-500MB columnar unit with min/max stats for pruning
  3. C.sorting data within partitions by up to four columns
  4. D.Two distributed join strategies: broadcast copies a small table to every node, shuffle redistributes both tables by join key across nodes.

Answer + AI explanation with Pro

8. Which term means: "querying or restoring historical table data within a retention window"?

Junior
  1. A.zone map
  2. B.accumulating snapshot fact
  3. C.Snowflake Time Travel
  4. D.Snowflake virtual warehouse

Answer + AI explanation with Pro

9. Which statement is correct?

Junior
  1. A.Snowflake Time Travel — A Redshift node type that separates compute from managed storage (RMS), letting you scale compute without overprovisioning storage.
  2. B.Snowflake Time Travel — querying or restoring historical table data within a retention window
  3. C.Snowflake Time Travel — A warehouse feature that spins up transient extra compute to absorb query spikes, keeping latency stable under high concurrency.
  4. D.Snowflake Time Travel — an independent compute cluster, separate from storage

Answer + AI explanation with Pro

10. What is zero-copy cloning?

Junior
  1. A.a virtual CPU unit BigQuery uses to execute SQL
  2. B.A small dimension that consolidates assorted low-cardinality flags and indicators into one table to keep fact tables narrow.
  3. C.splitting a table by date/integer range to cut bytes scanned
  4. D.an instant metadata-only copy of a table/schema/database

Answer + AI explanation with Pro

11. Which term means: "an instant metadata-only copy of a table/schema/database"?

Junior
  1. A.BigQuery slot
  2. B.Redshift Spectrum
  3. C.separation of storage and compute
  4. D.zero-copy cloning

Answer + AI explanation with Pro

12. Which statement is correct?

Junior
  1. A.zero-copy cloning — an instant metadata-only copy of a table/schema/database
  2. B.zero-copy cloning — A Redshift node type that separates compute from managed storage (RMS), letting you scale compute without overprovisioning storage.
  3. C.zero-copy cloning — A small dimension that consolidates assorted low-cardinality flags and indicators into one table to keep fact tables narrow.
  4. D.zero-copy cloning — A fact table that records the occurrence of an event or a coverage relationship with no numeric measures, used for counting events or capturing eligibility.

Answer + AI explanation with Pro

13. What is BigQuery partitioning?

Mid
  1. A.sorting data within partitions by up to four columns
  2. B.A cloud-warehouse design where data sits in shared storage and elastic compute scales independently, so you pay for each separately.
  3. C.an instant metadata-only copy of a table/schema/database
  4. D.splitting a table by date/integer range to cut bytes scanned

Answer + AI explanation with Pro

14. Which term means: "splitting a table by date/integer range to cut bytes scanned"?

Mid
  1. A.broadcast vs shuffle join
  2. B.Redshift Spectrum
  3. C.Snowflake virtual warehouse
  4. D.BigQuery partitioning

Answer + AI explanation with Pro

15. Which statement is correct?

Mid
  1. A.BigQuery partitioning — an instant metadata-only copy of a table/schema/database
  2. B.BigQuery partitioning — splitting a table by date/integer range to cut bytes scanned
  3. C.BigQuery partitioning — Auto-maintained min/max metadata per block (Redshift) or micro-partition that lets the engine skip blocks not matching a filter.
  4. D.BigQuery partitioning — A fact table tracking a process with multiple milestones, where rows are updated as each step completes, used for pipelines like order fulfillment.

Answer + AI explanation with Pro

16. What is BigQuery clustering?

Mid
  1. A.A Redshift feature that queries data directly in S3 without loading it, extending the warehouse over the data lake.
  2. B.Uneven distribution of rows across nodes or partitions that causes some workers to do far more work, a common cause of slow distributed queries.
  3. C.an independent compute cluster, separate from storage
  4. D.sorting data within partitions by up to four columns

Answer + AI explanation with Pro

17. Which term means: "sorting data within partitions by up to four columns"?

Mid
  1. A.BigQuery clustering
  2. B.Snowflake micro-partition
  3. C.Redshift distribution style KEY
  4. D.data skew

Answer + AI explanation with Pro

18. Which statement is correct?

Mid
  1. A.BigQuery clustering — an automatic ~50-500MB columnar unit with min/max stats for pruning
  2. B.BigQuery clustering — sorting data within partitions by up to four columns
  3. C.BigQuery clustering — co-locating rows with the same key on the same node for joins
  4. D.BigQuery clustering — Two distributed join strategies: broadcast copies a small table to every node, shuffle redistributes both tables by join key across nodes.

Answer + AI explanation with Pro

19. What is BigQuery slot?

Junior
  1. A.A cloud-warehouse design where data sits in shared storage and elastic compute scales independently, so you pay for each separately.
  2. B.A small dimension that consolidates assorted low-cardinality flags and indicators into one table to keep fact tables narrow.
  3. C.splitting a table by date/integer range to cut bytes scanned
  4. D.a virtual CPU unit BigQuery uses to execute SQL

Answer + AI explanation with Pro

20. Which term means: "a virtual CPU unit BigQuery uses to execute SQL"?

Junior
  1. A.factless fact table
  2. B.result cache
  3. C.accumulating snapshot fact
  4. D.BigQuery slot

Answer + AI explanation with Pro

21. Which statement is correct?

Junior
  1. A.BigQuery slot — An architecture that adds warehouse-grade management (ACID, governance, BI) directly on open data-lake storage via open table formats.
  2. B.BigQuery slot — copying a small table to every node to avoid redistribution
  3. C.BigQuery slot — A Redshift feature that queries data directly in S3 without loading it, extending the warehouse over the data lake.
  4. D.BigQuery slot — a virtual CPU unit BigQuery uses to execute SQL

Answer + AI explanation with Pro

22. What is Redshift distribution style KEY?

Mid
  1. A.A fact table that records the occurrence of an event or a coverage relationship with no numeric measures, used for counting events or capturing eligibility.
  2. B.co-locating rows with the same key on the same node for joins
  3. C.A warehouse feature that spins up transient extra compute to absorb query spikes, keeping latency stable under high concurrency.
  4. D.splitting a table by date/integer range to cut bytes scanned

Answer + AI explanation with Pro

23. Which term means: "co-locating rows with the same key on the same node for joins"?

Mid
  1. A.BigQuery partitioning
  2. B.broadcast vs shuffle join
  3. C.cost-based optimizer
  4. D.Redshift distribution style KEY

Answer + AI explanation with Pro

24. Which statement is correct?

Mid
  1. A.Redshift distribution style KEY — An architecture that adds warehouse-grade management (ACID, governance, BI) directly on open data-lake storage via open table formats.
  2. B.Redshift distribution style KEY — A planner that uses table statistics to estimate row counts and pick the cheapest join order and access path for a query.
  3. C.Redshift distribution style KEY — splitting a table by date/integer range to cut bytes scanned
  4. D.Redshift distribution style KEY — co-locating rows with the same key on the same node for joins

Answer + AI explanation with Pro

25. What is Redshift distribution style ALL?

Mid
  1. A.copying a small table to every node to avoid redistribution
  2. B.A small dimension that consolidates assorted low-cardinality flags and indicators into one table to keep fact tables narrow.
  3. C.A warehouse cache that returns a previously computed result for an identical query without rescanning, often free and instantaneous.
  4. D.Massively parallel processing, an architecture that distributes data and query work across many nodes that compute in parallel, the basis of cloud warehouses.

Answer + AI explanation with Pro

26. Which term means: "copying a small table to every node to avoid redistribution"?

Mid
  1. A.junk dimension
  2. B.RA3 node
  3. C.zero-copy cloning
  4. D.Redshift distribution style ALL

Answer + AI explanation with Pro

27. Which statement is correct?

Mid
  1. A.Redshift distribution style ALL — A managed service that continuously reorganizes table data by clustering keys in the background so users never run manual recluster jobs.
  2. B.Redshift distribution style ALL — copying a small table to every node to avoid redistribution
  3. C.Redshift distribution style ALL — A small dimension that consolidates assorted low-cardinality flags and indicators into one table to keep fact tables narrow.
  4. D.Redshift distribution style ALL — sorting data within partitions by up to four columns

Answer + AI explanation with Pro

28. What is MPP?

Mid
  1. A.A fact table that records the occurrence of an event or a coverage relationship with no numeric measures, used for counting events or capturing eligibility.
  2. B.A managed service that continuously reorganizes table data by clustering keys in the background so users never run manual recluster jobs.
  3. C.A Redshift table property defining physical row ordering so range-restricted scans skip blocks via zone maps; compound or interleaved variants exist.
  4. D.Massively parallel processing, an architecture that distributes data and query work across many nodes that compute in parallel, the basis of cloud warehouses.

Answer + AI explanation with Pro

29. Which term means: "Massively parallel processing, an architecture that distributes data and query work across many nodes that compute in parallel, the basis of cloud warehouses."?

Mid
  1. A.MPP
  2. B.auto-clustering
  3. C.BigQuery slot
  4. D.sort key

Answer + AI explanation with Pro

30. Which statement is correct?

Mid
  1. A.MPP — Massively parallel processing, an architecture that distributes data and query work across many nodes that compute in parallel, the basis of cloud warehouses.
  2. B.MPP — A Redshift feature that queries data directly in S3 without loading it, extending the warehouse over the data lake.
  3. C.MPP — A warehouse cache that returns a previously computed result for an identical query without rescanning, often free and instantaneous.
  4. D.MPP — an automatic ~50-500MB columnar unit with min/max stats for pruning

Answer + AI explanation with Pro

Showing 30 of 78 Warehousing questions — the full set, with answers, explanations and an AI tutor on every question, is inside.

Free to start

Start with a free readiness check

Sign up free for the 2-minute IT readiness check and a scored result. Answers, explanations and the AI tutor on every Warehousing question come with Pro.

Take the free IT readiness check

or take a mock interview set up for this area

24,000+ questions & coding problemsSoftware & IT16,274 questionsGovernment jobs26 examsAptitudenew questions every timeAI practice interviewwith feedback65 topics to practiseMechanical1,149 questionsGATE ME9 papersEngineering Mathematics381 questions2-minute checkfreeDSA Problems1,422Civil1,005 questionsGATE CE9 papersCS Fundamentals1,209 questionsYour scores6 skillsSystem Design25Electrical / EEE1,047 questionsGATE EE9 papersRun your codeC++ · Java · PythonLow-Level Design144Electronics & Comm.975 questionsGATE EC9 papersAI help on every questionFull-Stack6,282Chemical1,005 questionsGATE CH9 papersAI whiteboardsystem designWork abroadEurope · remote · transfersESE ME1 paperGATE practice papers2019–2026ESE CE1 paperDate alertsbefore the last dateESE EE1 paperBehavioural courseHR round practiceESE ET1 paperResume optimizerProSSC JE ME1 paperApplication trackerSSC JE CE1 paperCompany-wise prepSSC JE EE1 paperRole roadmapsRRB JE1 subjectPriced in ₹UPI · cardsISRO SC1 paperGATE CS9 papersIBPS SO IT1 paperUGC NET CS1 paperSSC CGL26 papersIBPS PO26 papersRRB NTPC26 papersSSC CHSL26 papersIBPS Clerk26 papersSBI Clerk26 papersRRB Group D26 papersSSC CPO26 papersSSC GD26 papers