60 Cost & Performance questions from the Data Engineering bank, written for Indian campus drives and tech interviews. Every question has a verified answer and an AI-tutor explanation on placd.
Free to start: the 2-minute IT readiness check — six questions and a result.
A.Charges incurred when moving data out of a cloud region or provider, a factor in multi-cloud and cross-region pipeline design.
B.Skipping entire partitions that cannot match a query's filter, the primary way partitioning reduces scanned data and cost.
C.A cache that returns a previously computed query result instantly when the same query runs again over unchanged data, scanning zero bytes.
D.Matching warehouse or slot capacity to workload so jobs finish quickly without paying for idle or oversized compute.
Answer + AI explanation with Pro
2. Which term means: "Skipping entire partitions that cannot match a query's filter, the primary way partitioning reduces scanned data and cost."?
Mid
A.SELECT star avoidance
B.partition pruning
C.broadcast join
D.materialization tradeoff
Answer + AI explanation with Pro
3. Which statement is correct?
Mid
A.partition pruning — Skipping entire partitions that cannot match a query's filter, the primary way partitioning reduces scanned data and cost.
B.partition pruning — Physically co-locating rows by frequently filtered columns so the engine reads fewer micro-partitions or blocks for selective queries.
C.partition pruning — The primary cost driver for on-demand scan-priced warehouses; minimized via partition pruning, clustering, and selecting only needed columns.
D.partition pruning — When a query exceeds available memory and writes intermediate data to local or remote disk, a major performance and cost red flag.
Answer + AI explanation with Pro
4. What is clustering?
Senior
A.Physically co-locating rows by frequently filtered columns so the engine reads fewer micro-partitions or blocks for selective queries.
B.The primary cost driver in serverless query engines (BigQuery, Athena), reduced by partitioning, clustering, and selecting fewer columns.
C.Matching warehouse or slot capacity to workload so jobs finish quickly without paying for idle or oversized compute.
D.Tagging and tracking compute spend by team, query, or workload so chargeback and optimization can target the biggest cost drivers.
Answer + AI explanation with Pro
5. Which term means: "Physically co-locating rows by frequently filtered columns so the engine reads fewer micro-partitions or blocks for selective queries."?
Senior
A.result cache
B.right-sizing warehouse
C.clustering
D.data egress cost
Answer + AI explanation with Pro
6. Which statement is correct?
Senior
A.clustering — An uneven distribution of values across partitions or keys that overloads a few tasks, creating stragglers that dominate job runtime.
B.clustering — Physically co-locating rows by frequently filtered columns so the engine reads fewer micro-partitions or blocks for selective queries.
C.clustering — A join that ships a small table in full to every executor so each partition of the large table joins locally, avoiding a shuffle.
D.clustering — Refreshing only changed partitions of a derived table instead of rebuilding it fully, the main lever for cheap, frequent transformations.
Answer + AI explanation with Pro
7. What is auto-suspend?
Mid
A.A join that ships a small table in full to every executor so each partition of the large table joins locally, avoiding a shuffle.
B.The expensive redistribution of data across the cluster by key (for joins, group-bys, or sorts) that moves records over the network.
C.Refreshing only changed partitions of a derived table instead of rebuilding it fully, the main lever for cheap, frequent transformations.
D.Configuring compute to pause after idle time so you stop paying for unused warehouse runtime, a key cost control.
Answer + AI explanation with Pro
8. Which term means: "Configuring compute to pause after idle time so you stop paying for unused warehouse runtime, a key cost control."?
Mid
A.auto-suspend
B.result cache
C.partition pruning
D.clustering
Answer + AI explanation with Pro
9. Which statement is correct?
Mid
A.auto-suspend — Configuring compute to pause after idle time so you stop paying for unused warehouse runtime, a key cost control.
B.auto-suspend — The primary cost driver for on-demand scan-priced warehouses; minimized via partition pruning, clustering, and selecting only needed columns.
C.auto-suspend — Refreshing only changed partitions of a derived table instead of rebuilding it fully, the main lever for cheap, frequent transformations.
D.auto-suspend — Choosing reserved capacity (lower unit cost, fixed commitment) versus pay-per-use (flexible, higher unit cost) based on workload predictability.
Answer + AI explanation with Pro
10. What is spilling?
Senior
A.Refreshing only changed partitions of a derived table instead of rebuilding it fully, the main lever for cheap, frequent transformations.
B.An uneven distribution of values across partitions or keys that overloads a few tasks, creating stragglers that dominate job runtime.
C.Skipping entire partitions that cannot match a query's filter, the primary way partitioning reduces scanned data and cost.
D.When a query exceeds available memory and writes intermediate data to local or remote disk, a major performance and cost red flag.
Answer + AI explanation with Pro
11. Which term means: "When a query exceeds available memory and writes intermediate data to local or remote disk, a major performance and cost red flag."?
Senior
A.shuffle
B.spilling
C.partition pruning
D.data skew
Answer + AI explanation with Pro
12. Which statement is correct?
Senior
A.spilling — When a query exceeds available memory and writes intermediate data to local or remote disk, a major performance and cost red flag.
B.spilling — Refreshing only changed partitions of a derived table instead of rebuilding it fully, the main lever for cheap, frequent transformations.
C.spilling — Skipping entire partitions that cannot match a query's filter, the primary way partitioning reduces scanned data and cost.
D.spilling — Reading only needed columns instead of all columns, cutting bytes scanned and improving performance in columnar engines.
Answer + AI explanation with Pro
13. What is materialization tradeoff?
Mid
A.Refreshing only changed partitions of a derived table instead of rebuilding it fully, the main lever for cheap, frequent transformations.
B.Reading only needed columns instead of all columns, cutting bytes scanned and improving performance in columnar engines.
C.Choosing between recomputing a query each time (cheaper storage, costlier compute) and precomputing it as a table or MV (faster reads, refresh cost).
D.An uneven distribution of values across partitions or keys that overloads a few tasks, creating stragglers that dominate job runtime.
Answer + AI explanation with Pro
14. Which term means: "Choosing between recomputing a query each time (cheaper storage, costlier compute) and precomputing it as a table or MV (faster reads, refresh cost)."?
Mid
A.auto-suspend
B.materialization tradeoff
C.right-sizing warehouse
D.partition pruning
Answer + AI explanation with Pro
15. Which statement is correct?
Mid
A.materialization tradeoff — Reading only needed columns instead of all columns, cutting bytes scanned and improving performance in columnar engines.
B.materialization tradeoff — Choosing between recomputing a query each time (cheaper storage, costlier compute) and precomputing it as a table or MV (faster reads, refresh cost).
C.materialization tradeoff — Skipping entire partitions that cannot match a query's filter, the primary way partitioning reduces scanned data and cost.
D.materialization tradeoff — A performance issue where many tiny files force excessive metadata and task overhead; compacting into larger files restores efficient scans.
Answer + AI explanation with Pro
16. What is query cost attribution?
Senior
A.Charges incurred when moving data out of a cloud region or provider, a factor in multi-cloud and cross-region pipeline design.
B.Tagging and tracking compute spend by team, query, or workload so chargeback and optimization can target the biggest cost drivers.
C.A cache that returns a previously computed query result instantly when the same query runs again over unchanged data, scanning zero bytes.
D.An uneven distribution of values across partitions or keys that overloads a few tasks, creating stragglers that dominate job runtime.
Answer + AI explanation with Pro
17. Which term means: "Tagging and tracking compute spend by team, query, or workload so chargeback and optimization can target the biggest cost drivers."?
Senior
A.query cost attribution
B.spilling
C.broadcast join
D.shuffle
Answer + AI explanation with Pro
18. Which statement is correct?
Senior
A.query cost attribution — A cache that returns a previously computed query result instantly when the same query runs again over unchanged data, scanning zero bytes.
B.query cost attribution — Tagging and tracking compute spend by team, query, or workload so chargeback and optimization can target the biggest cost drivers.
C.query cost attribution — A warehouse setting that idles compute after a period of inactivity so you stop paying for an unused cluster.
D.query cost attribution — Refreshing only changed partitions of a derived table instead of rebuilding it fully, the main lever for cheap, frequent transformations.
Answer + AI explanation with Pro
19. What is right-sizing warehouse?
Mid
A.Tagging and tracking compute spend by team, query, or workload so chargeback and optimization can target the biggest cost drivers.
B.Skipping entire partitions that cannot match a query's filter, the primary way partitioning reduces scanned data and cost.
C.Matching warehouse or slot capacity to workload so jobs finish quickly without paying for idle or oversized compute.
D.A join that ships a small table in full to every executor so each partition of the large table joins locally, avoiding a shuffle.
Answer + AI explanation with Pro
20. Which term means: "Matching warehouse or slot capacity to workload so jobs finish quickly without paying for idle or oversized compute."?
Mid
A.clustering
B.right-sizing warehouse
C.spilling
D.materialization tradeoff
Answer + AI explanation with Pro
21. Which statement is correct?
Mid
A.right-sizing warehouse — Matching warehouse or slot capacity to workload so jobs finish quickly without paying for idle or oversized compute.
B.right-sizing warehouse — Physically co-locating rows by frequently filtered columns so the engine reads fewer micro-partitions or blocks for selective queries.
C.right-sizing warehouse — A join that ships a small table in full to every executor so each partition of the large table joins locally, avoiding a shuffle.
D.right-sizing warehouse — The primary cost driver for on-demand scan-priced warehouses; minimized via partition pruning, clustering, and selecting only needed columns.
Answer + AI explanation with Pro
22. What is data egress cost?
Senior
A.The expensive redistribution of data across the cluster by key (for joins, group-bys, or sorts) that moves records over the network.
B.Configuring compute to pause after idle time so you stop paying for unused warehouse runtime, a key cost control.
C.Charges incurred when moving data out of a cloud region or provider, a factor in multi-cloud and cross-region pipeline design.
D.A cache that returns a previously computed query result instantly when the same query runs again over unchanged data, scanning zero bytes.
Answer + AI explanation with Pro
23. Which term means: "Charges incurred when moving data out of a cloud region or provider, a factor in multi-cloud and cross-region pipeline design."?
Senior
A.data egress cost
B.small-files problem
C.clustering
D.spill
Answer + AI explanation with Pro
24. Which statement is correct?
Senior
A.data egress cost — Matching warehouse or slot capacity to workload so jobs finish quickly without paying for idle or oversized compute.
B.data egress cost — Charges incurred when moving data out of a cloud region or provider, a factor in multi-cloud and cross-region pipeline design.
C.data egress cost — The primary cost driver for on-demand scan-priced warehouses; minimized via partition pruning, clustering, and selecting only needed columns.
D.data egress cost — When an operation's working set exceeds memory and is written to disk, slowing execution; reduced by tuning partitions or adding memory.
Answer + AI explanation with Pro
25. What is bytes scanned?
Mid
A.A performance issue where many tiny files force excessive metadata and task overhead; compacting into larger files restores efficient scans.
B.A join that ships a small table in full to every executor so each partition of the large table joins locally, avoiding a shuffle.
C.Tagging and tracking compute spend by team, query, or workload so chargeback and optimization can target the biggest cost drivers.
D.The primary cost driver in serverless query engines (BigQuery, Athena), reduced by partitioning, clustering, and selecting fewer columns.
Answer + AI explanation with Pro
26. Which term means: "The primary cost driver in serverless query engines (BigQuery, Athena), reduced by partitioning, clustering, and selecting fewer columns."?
Mid
A.bytes scanned
B.SELECT star avoidance
C.data skew
D.commitment vs on-demand
Answer + AI explanation with Pro
27. Which statement is correct?
Mid
A.bytes scanned — When a query exceeds available memory and writes intermediate data to local or remote disk, a major performance and cost red flag.
B.bytes scanned — Charges incurred when moving data out of a cloud region or provider, a factor in multi-cloud and cross-region pipeline design.
C.bytes scanned — A cache that returns a previously computed query result instantly when the same query runs again over unchanged data, scanning zero bytes.
D.bytes scanned — The primary cost driver in serverless query engines (BigQuery, Athena), reduced by partitioning, clustering, and selecting fewer columns.
Answer + AI explanation with Pro
28. What is SELECT star avoidance?
Mid
A.Physically co-locating rows by frequently filtered columns so the engine reads fewer micro-partitions or blocks for selective queries.
B.A warehouse setting that idles compute after a period of inactivity so you stop paying for an unused cluster.
C.Reading only needed columns instead of all columns, cutting bytes scanned and improving performance in columnar engines.
D.Charges incurred when moving data out of a cloud region or provider, a factor in multi-cloud and cross-region pipeline design.
Answer + AI explanation with Pro
29. Which term means: "Reading only needed columns instead of all columns, cutting bytes scanned and improving performance in columnar engines."?
Mid
A.broadcast join
B.query cost attribution
C.commitment vs on-demand
D.SELECT star avoidance
Answer + AI explanation with Pro
30. Which statement is correct?
Mid
A.SELECT star avoidance — A performance issue where many tiny files force excessive metadata and task overhead; compacting into larger files restores efficient scans.
B.SELECT star avoidance — Reading only needed columns instead of all columns, cutting bytes scanned and improving performance in columnar engines.
C.SELECT star avoidance — Tagging and tracking compute spend by team, query, or workload so chargeback and optimization can target the biggest cost drivers.
D.SELECT star avoidance — Choosing between recomputing a query each time (cheaper storage, costlier compute) and precomputing it as a table or MV (faster reads, refresh cost).
Answer + AI explanation with Pro
Showing 30 of 60 Cost & Performance questions — the full set, with answers, explanations and an AI tutor on every question, is inside.
Free to start
Start with a free readiness check
Sign up free for the 2-minute IT readiness check and a scored result. Answers, explanations and the AI tutor on every Cost & Performance question come with Pro.