78 Warehousing questions from the Data Engineering bank, written for Indian campus drives and tech interviews. Every question has a verified answer and an AI-tutor explanation on placd.
Free to start: the 2-minute IT readiness check — six questions and a result.
A.an automatic ~50-500MB columnar unit with min/max stats for pruning
B.A cloud-warehouse design where data sits in shared storage and elastic compute scales independently, so you pay for each separately.
C.Massively parallel processing, an architecture that distributes data and query work across many nodes that compute in parallel, the basis of cloud warehouses.
D.Redshift Workload Management, which routes queries into queues with memory and concurrency settings; automatic WLM tunes this dynamically.
Answer + AI explanation with Pro
2. Which term means: "an automatic ~50-500MB columnar unit with min/max stats for pruning"?
Mid
A.Snowflake micro-partition
B.Redshift distribution style KEY
C.MPP
D.BigQuery clustering
Answer + AI explanation with Pro
3. Which statement is correct?
Mid
A.Snowflake micro-partition — an independent compute cluster, separate from storage
B.Snowflake micro-partition — A fact table that records the occurrence of an event or a coverage relationship with no numeric measures, used for counting events or capturing eligibility.
C.Snowflake micro-partition — copying a small table to every node to avoid redistribution
D.Snowflake micro-partition — an automatic ~50-500MB columnar unit with min/max stats for pruning
Answer + AI explanation with Pro
4. What is Snowflake virtual warehouse?
Junior
A.an independent compute cluster, separate from storage
B.splitting a table by date/integer range to cut bytes scanned
C.sorting data within partitions by up to four columns
D.A fact table that records the occurrence of an event or a coverage relationship with no numeric measures, used for counting events or capturing eligibility.
Answer + AI explanation with Pro
5. Which term means: "an independent compute cluster, separate from storage"?
Junior
A.RA3 node
B.Snowflake virtual warehouse
C.cost-based optimizer
D.MPP
Answer + AI explanation with Pro
6. Which statement is correct?
Junior
A.Snowflake virtual warehouse — A fact table tracking a process with multiple milestones, where rows are updated as each step completes, used for pipelines like order fulfillment.
B.Snowflake virtual warehouse — a virtual CPU unit BigQuery uses to execute SQL
C.Snowflake virtual warehouse — an independent compute cluster, separate from storage
D.Snowflake virtual warehouse — Auto-maintained min/max metadata per block (Redshift) or micro-partition that lets the engine skip blocks not matching a filter.
Answer + AI explanation with Pro
7. What is Snowflake Time Travel?
Junior
A.querying or restoring historical table data within a retention window
B.an automatic ~50-500MB columnar unit with min/max stats for pruning
C.sorting data within partitions by up to four columns
D.Two distributed join strategies: broadcast copies a small table to every node, shuffle redistributes both tables by join key across nodes.
Answer + AI explanation with Pro
8. Which term means: "querying or restoring historical table data within a retention window"?
Junior
A.zone map
B.accumulating snapshot fact
C.Snowflake Time Travel
D.Snowflake virtual warehouse
Answer + AI explanation with Pro
9. Which statement is correct?
Junior
A.Snowflake Time Travel — A Redshift node type that separates compute from managed storage (RMS), letting you scale compute without overprovisioning storage.
B.Snowflake Time Travel — querying or restoring historical table data within a retention window
C.Snowflake Time Travel — A warehouse feature that spins up transient extra compute to absorb query spikes, keeping latency stable under high concurrency.
D.Snowflake Time Travel — an independent compute cluster, separate from storage
Answer + AI explanation with Pro
10. What is zero-copy cloning?
Junior
A.a virtual CPU unit BigQuery uses to execute SQL
B.A small dimension that consolidates assorted low-cardinality flags and indicators into one table to keep fact tables narrow.
C.splitting a table by date/integer range to cut bytes scanned
D.an instant metadata-only copy of a table/schema/database
Answer + AI explanation with Pro
11. Which term means: "an instant metadata-only copy of a table/schema/database"?
Junior
A.BigQuery slot
B.Redshift Spectrum
C.separation of storage and compute
D.zero-copy cloning
Answer + AI explanation with Pro
12. Which statement is correct?
Junior
A.zero-copy cloning — an instant metadata-only copy of a table/schema/database
B.zero-copy cloning — A Redshift node type that separates compute from managed storage (RMS), letting you scale compute without overprovisioning storage.
C.zero-copy cloning — A small dimension that consolidates assorted low-cardinality flags and indicators into one table to keep fact tables narrow.
D.zero-copy cloning — A fact table that records the occurrence of an event or a coverage relationship with no numeric measures, used for counting events or capturing eligibility.
Answer + AI explanation with Pro
13. What is BigQuery partitioning?
Mid
A.sorting data within partitions by up to four columns
B.A cloud-warehouse design where data sits in shared storage and elastic compute scales independently, so you pay for each separately.
C.an instant metadata-only copy of a table/schema/database
D.splitting a table by date/integer range to cut bytes scanned
Answer + AI explanation with Pro
14. Which term means: "splitting a table by date/integer range to cut bytes scanned"?
Mid
A.broadcast vs shuffle join
B.Redshift Spectrum
C.Snowflake virtual warehouse
D.BigQuery partitioning
Answer + AI explanation with Pro
15. Which statement is correct?
Mid
A.BigQuery partitioning — an instant metadata-only copy of a table/schema/database
B.BigQuery partitioning — splitting a table by date/integer range to cut bytes scanned
C.BigQuery partitioning — Auto-maintained min/max metadata per block (Redshift) or micro-partition that lets the engine skip blocks not matching a filter.
D.BigQuery partitioning — A fact table tracking a process with multiple milestones, where rows are updated as each step completes, used for pipelines like order fulfillment.
Answer + AI explanation with Pro
16. What is BigQuery clustering?
Mid
A.A Redshift feature that queries data directly in S3 without loading it, extending the warehouse over the data lake.
B.Uneven distribution of rows across nodes or partitions that causes some workers to do far more work, a common cause of slow distributed queries.
C.an independent compute cluster, separate from storage
D.sorting data within partitions by up to four columns
Answer + AI explanation with Pro
17. Which term means: "sorting data within partitions by up to four columns"?
Mid
A.BigQuery clustering
B.Snowflake micro-partition
C.Redshift distribution style KEY
D.data skew
Answer + AI explanation with Pro
18. Which statement is correct?
Mid
A.BigQuery clustering — an automatic ~50-500MB columnar unit with min/max stats for pruning
B.BigQuery clustering — sorting data within partitions by up to four columns
C.BigQuery clustering — co-locating rows with the same key on the same node for joins
D.BigQuery clustering — Two distributed join strategies: broadcast copies a small table to every node, shuffle redistributes both tables by join key across nodes.
Answer + AI explanation with Pro
19. What is BigQuery slot?
Junior
A.A cloud-warehouse design where data sits in shared storage and elastic compute scales independently, so you pay for each separately.
B.A small dimension that consolidates assorted low-cardinality flags and indicators into one table to keep fact tables narrow.
C.splitting a table by date/integer range to cut bytes scanned
D.a virtual CPU unit BigQuery uses to execute SQL
Answer + AI explanation with Pro
20. Which term means: "a virtual CPU unit BigQuery uses to execute SQL"?
Junior
A.factless fact table
B.result cache
C.accumulating snapshot fact
D.BigQuery slot
Answer + AI explanation with Pro
21. Which statement is correct?
Junior
A.BigQuery slot — An architecture that adds warehouse-grade management (ACID, governance, BI) directly on open data-lake storage via open table formats.
B.BigQuery slot — copying a small table to every node to avoid redistribution
C.BigQuery slot — A Redshift feature that queries data directly in S3 without loading it, extending the warehouse over the data lake.
D.BigQuery slot — a virtual CPU unit BigQuery uses to execute SQL
Answer + AI explanation with Pro
22. What is Redshift distribution style KEY?
Mid
A.A fact table that records the occurrence of an event or a coverage relationship with no numeric measures, used for counting events or capturing eligibility.
B.co-locating rows with the same key on the same node for joins
C.A warehouse feature that spins up transient extra compute to absorb query spikes, keeping latency stable under high concurrency.
D.splitting a table by date/integer range to cut bytes scanned
Answer + AI explanation with Pro
23. Which term means: "co-locating rows with the same key on the same node for joins"?
Mid
A.BigQuery partitioning
B.broadcast vs shuffle join
C.cost-based optimizer
D.Redshift distribution style KEY
Answer + AI explanation with Pro
24. Which statement is correct?
Mid
A.Redshift distribution style KEY — An architecture that adds warehouse-grade management (ACID, governance, BI) directly on open data-lake storage via open table formats.
B.Redshift distribution style KEY — A planner that uses table statistics to estimate row counts and pick the cheapest join order and access path for a query.
C.Redshift distribution style KEY — splitting a table by date/integer range to cut bytes scanned
D.Redshift distribution style KEY — co-locating rows with the same key on the same node for joins
Answer + AI explanation with Pro
25. What is Redshift distribution style ALL?
Mid
A.copying a small table to every node to avoid redistribution
B.A small dimension that consolidates assorted low-cardinality flags and indicators into one table to keep fact tables narrow.
C.A warehouse cache that returns a previously computed result for an identical query without rescanning, often free and instantaneous.
D.Massively parallel processing, an architecture that distributes data and query work across many nodes that compute in parallel, the basis of cloud warehouses.
Answer + AI explanation with Pro
26. Which term means: "copying a small table to every node to avoid redistribution"?
Mid
A.junk dimension
B.RA3 node
C.zero-copy cloning
D.Redshift distribution style ALL
Answer + AI explanation with Pro
27. Which statement is correct?
Mid
A.Redshift distribution style ALL — A managed service that continuously reorganizes table data by clustering keys in the background so users never run manual recluster jobs.
B.Redshift distribution style ALL — copying a small table to every node to avoid redistribution
C.Redshift distribution style ALL — A small dimension that consolidates assorted low-cardinality flags and indicators into one table to keep fact tables narrow.
D.Redshift distribution style ALL — sorting data within partitions by up to four columns
Answer + AI explanation with Pro
28. What is MPP?
Mid
A.A fact table that records the occurrence of an event or a coverage relationship with no numeric measures, used for counting events or capturing eligibility.
B.A managed service that continuously reorganizes table data by clustering keys in the background so users never run manual recluster jobs.
C.A Redshift table property defining physical row ordering so range-restricted scans skip blocks via zone maps; compound or interleaved variants exist.
D.Massively parallel processing, an architecture that distributes data and query work across many nodes that compute in parallel, the basis of cloud warehouses.
Answer + AI explanation with Pro
29. Which term means: "Massively parallel processing, an architecture that distributes data and query work across many nodes that compute in parallel, the basis of cloud warehouses."?
Mid
A.MPP
B.auto-clustering
C.BigQuery slot
D.sort key
Answer + AI explanation with Pro
30. Which statement is correct?
Mid
A.MPP — Massively parallel processing, an architecture that distributes data and query work across many nodes that compute in parallel, the basis of cloud warehouses.
B.MPP — A Redshift feature that queries data directly in S3 without loading it, extending the warehouse over the data lake.
C.MPP — A warehouse cache that returns a previously computed result for an identical query without rescanning, often free and instantaneous.
D.MPP — an automatic ~50-500MB columnar unit with min/max stats for pruning
Answer + AI explanation with Pro
Showing 30 of 78 Warehousing questions — the full set, with answers, explanations and an AI tutor on every question, is inside.
Free to start
Start with a free readiness check
Sign up free for the 2-minute IT readiness check and a scored result. Answers, explanations and the AI tutor on every Warehousing question come with Pro.