72 Data Quality questions from the Data Engineering bank, written for Indian campus drives and tech interviews. Every question has a verified answer and an AI-tutor explanation on placd.
Free to start: the 2-minute IT readiness check — six questions and a result.
D.A check confirming a column or key has no duplicate values, protecting primary-key and grain assumptions.
Answer + AI explanation with Pro
2. Which term means: "completeness, accuracy, consistency, timeliness, uniqueness, validity"?
Junior
A.anomaly detection test
B.Great Expectations checkpoint
C.reconciliation
D.data quality dimensions
Answer + AI explanation with Pro
3. Which statement is correct?
Junior
A.data quality dimensions — A dbt test that feeds a model fixed mock rows and asserts the transformed output, validating logic independently of warehouse data.
C.data quality dimensions — A check verifying that foreign-key values in one table all exist in the referenced table, catching orphaned records.
D.data quality dimensions — A data-quality check that flags metrics deviating from their historical baseline rather than failing a fixed threshold, catching unexpected drifts.
Answer + AI explanation with Pro
4. What is schema drift?
Mid
A.Monitoring that alerts when upstream columns are added, removed, or change type so downstream transforms do not silently break.
B.an unexpected change in source schema that can break pipelines
C.A data-quality check that flags metrics deviating from their historical baseline rather than failing a fixed threshold, catching unexpected drifts.
D.A check that a table's row count falls within an expected range or relative change, catching partial loads and runaway duplications.
Answer + AI explanation with Pro
5. Which term means: "an unexpected change in source schema that can break pipelines"?
Mid
A.schema drift
B.anomaly detection test
C.freshness check
D.freshness SLA
Answer + AI explanation with Pro
6. Which statement is correct?
Mid
A.schema drift — Automated monitoring that flags metrics (row counts, nulls, distributions) deviating from learned historical patterns, a pillar of data observability.
B.schema drift — An open-source Python framework for declaring data expectations (assertions) and validating datasets, producing data docs and validation results.
C.schema drift — an unexpected change in source schema that can break pipelines
D.schema drift — A data-quality check that flags metrics deviating from their historical baseline rather than failing a fixed threshold, catching unexpected drifts.
Answer + AI explanation with Pro
7. What is data contract?
Mid
A.The practice of monitoring data health across freshness, volume, schema, distribution, and lineage to detect and resolve issues quickly.
B.an agreed schema/SLA between data producers and consumers
C.A documented maximum acceptable staleness for a dataset, breached when its last update falls outside the agreed window.
D.Asserting that incoming data matches an expected structure (columns, types, nullability) before it is accepted into a pipeline.
Answer + AI explanation with Pro
8. Which term means: "an agreed schema/SLA between data producers and consumers"?
Mid
A.anomaly detection test
B.data observability
C.data contract
D.Soda
Answer + AI explanation with Pro
9. Which statement is correct?
Mid
A.data contract — Comparing record counts or aggregates between source and target after a load to confirm no data was lost or duplicated in transit.
B.data contract — Monitoring that alerts when upstream columns are added, removed, or change type so downstream transforms do not silently break.
C.data contract — an agreed schema/SLA between data producers and consumers
D.data contract — A pipeline safeguard that halts downstream loads when quality checks fail, preventing bad data from propagating to consumers.
Answer + AI explanation with Pro
10. What is Great Expectations?
Junior
A.An open-source Python framework for declaring data expectations (assertions) and validating datasets, producing data docs and validation results.
B.Asserting that incoming data matches an expected structure (columns, types, nullability) before it is accepted into a pipeline.
C.an unexpected change in source schema that can break pipelines
D.The practice of monitoring data health across freshness, volume, schema, distribution, and lineage to detect and resolve issues quickly.
Answer + AI explanation with Pro
11. Which term means: "An open-source Python framework for declaring data expectations (assertions) and validating datasets, producing data docs and validation results."?
Junior
A.freshness SLA
B.Great Expectations
C.referential integrity test
D.anomaly detection test
Answer + AI explanation with Pro
12. Which statement is correct?
Junior
A.Great Expectations — an agreed schema/SLA between data producers and consumers
B.Great Expectations — An open-source Python framework for declaring data expectations (assertions) and validating datasets, producing data docs and validation results.
C.Great Expectations — The practice of monitoring data health across freshness, volume, schema, distribution, and lineage to detect and resolve issues quickly.
D.Great Expectations — A documented maximum acceptable staleness for a dataset, breached when its last update falls outside the agreed window.
Answer + AI explanation with Pro
13. What is Soda?
Mid
A.A reusable bundle pairing data batches with an expectation suite and actions, run to validate data and emit results in a pipeline.
B.A pipeline safeguard that halts downstream loads when quality checks fail, preventing bad data from propagating to consumers.
C.A declarative data-quality assertion written in SodaCL (such as row count, missing percent, or freshness) that runs against a dataset.
D.A data-quality tool using a declarative checks language (SodaCL) to scan datasets for anomalies, freshness, and rule violations in pipelines.
Answer + AI explanation with Pro
14. Which term means: "A data-quality tool using a declarative checks language (SodaCL) to scan datasets for anomalies, freshness, and rule violations in pipelines."?
Mid
A.reconciliation
B.Soda
C.schema validation
D.data observability
Answer + AI explanation with Pro
15. Which statement is correct?
Mid
A.Soda — A data-quality tool using a declarative checks language (SodaCL) to scan datasets for anomalies, freshness, and rule violations in pipelines.
B.Soda — A dbt test that feeds a model fixed mock rows and asserts the transformed output, validating logic independently of warehouse data.
C.Soda — A pipeline safeguard that halts downstream loads when quality checks fail, preventing bad data from propagating to consumers.
D.Soda — A reusable bundle pairing data batches with an expectation suite and actions, run to validate data and emit results in a pipeline.
Answer + AI explanation with Pro
16. What is data contract?
Mid
A.A check confirming a column or key has no duplicate values, protecting primary-key and grain assumptions.
B.A dbt test that feeds a model fixed mock rows and asserts the transformed output, validating logic independently of warehouse data.
C.A formal, versioned agreement between data producers and consumers specifying schema, semantics, and quality guarantees, enforced in CI or at ingestion.
D.an agreed schema/SLA between data producers and consumers
Answer + AI explanation with Pro
17. Which term means: "A formal, versioned agreement between data producers and consumers specifying schema, semantics, and quality guarantees, enforced in CI or at ingestion."?
Mid
A.dbt unit test
B.data contract
C.schema validation
D.Soda
Answer + AI explanation with Pro
18. Which statement is correct?
Mid
A.data contract — Asserting that incoming data matches an expected structure (columns, types, nullability) before it is accepted into a pipeline.
B.data contract — Comparing aggregate totals between source and target (counts, sums) to confirm a pipeline neither dropped nor duplicated data.
C.data contract — A dbt test that feeds a model fixed mock rows and asserts the transformed output, validating logic independently of warehouse data.
D.data contract — A formal, versioned agreement between data producers and consumers specifying schema, semantics, and quality guarantees, enforced in CI or at ingestion.
Answer + AI explanation with Pro
19. What is freshness check?
Junior
A.A test asserting that the newest data in a table is within an acceptable age threshold, catching stalled or delayed pipelines.
B.An open-source Python framework for declaring data expectations (assertions) and validating datasets, producing data docs and validation results.
C.Monitoring that alerts when upstream columns are added, removed, or change type so downstream transforms do not silently break.
D.Comparing record counts or aggregates between source and target after a load to confirm no data was lost or duplicated in transit.
Answer + AI explanation with Pro
20. Which term means: "A test asserting that the newest data in a table is within an acceptable age threshold, catching stalled or delayed pipelines."?
Junior
A.schema validation
B.reconciliation
C.freshness check
D.anomaly detection test
Answer + AI explanation with Pro
21. Which statement is correct?
Junior
A.freshness check — A pipeline safeguard that halts downstream loads when quality checks fail, preventing bad data from propagating to consumers.
B.freshness check — A test asserting that the newest data in a table is within an acceptable age threshold, catching stalled or delayed pipelines.
C.freshness check — A check that a table's row count falls within an expected range or relative change, catching partial loads and runaway duplications.
D.freshness check — A declarative data-quality assertion written in SodaCL (such as row count, missing percent, or freshness) that runs against a dataset.
Answer + AI explanation with Pro
22. What is anomaly detection?
Mid
A.A reusable bundle pairing data batches with an expectation suite and actions, run to validate data and emit results in a pipeline.
B.an unexpected change in source schema that can break pipelines
C.Automated monitoring that flags metrics (row counts, nulls, distributions) deviating from learned historical patterns, a pillar of data observability.
D.A dbt test that feeds a model fixed mock rows and asserts the transformed output, validating logic independently of warehouse data.
Answer + AI explanation with Pro
23. Which term means: "Automated monitoring that flags metrics (row counts, nulls, distributions) deviating from learned historical patterns, a pillar of data observability."?
B.anomaly detection — A named collection of Great Expectations assertions about a dataset that together define its quality contract.
C.anomaly detection — Automated monitoring that flags metrics (row counts, nulls, distributions) deviating from learned historical patterns, a pillar of data observability.
D.anomaly detection — A documented maximum acceptable staleness for a dataset, breached when its last update falls outside the agreed window.
Answer + AI explanation with Pro
25. What is referential integrity test?
Junior
A.an unexpected change in source schema that can break pipelines
B.A check that a table's row count falls within an expected range or relative change, catching partial loads and runaway duplications.
C.Automated monitoring that flags metrics (row counts, nulls, distributions) deviating from learned historical patterns, a pillar of data observability.
D.A check verifying that foreign-key values in one table all exist in the referenced table, catching orphaned records.
Answer + AI explanation with Pro
26. Which term means: "A check verifying that foreign-key values in one table all exist in the referenced table, catching orphaned records."?
Junior
A.schema drift detection
B.referential integrity test
C.Soda
D.volume test
Answer + AI explanation with Pro
27. Which statement is correct?
Junior
A.referential integrity test — A check that a table's row count falls within an expected range or relative change, catching partial loads and runaway duplications.
B.referential integrity test — A check verifying that foreign-key values in one table all exist in the referenced table, catching orphaned records.
C.referential integrity test — A named collection of Great Expectations assertions about a dataset that together define its quality contract.
D.referential integrity test — Automated monitoring that flags metrics (row counts, nulls, distributions) deviating from learned historical patterns, a pillar of data observability.
Answer + AI explanation with Pro
28. What is data observability?
Mid
A.an unexpected change in source schema that can break pipelines
B.A data-quality tool using a declarative checks language (SodaCL) to scan datasets for anomalies, freshness, and rule violations in pipelines.
C.The practice of monitoring data health across freshness, volume, schema, distribution, and lineage to detect and resolve issues quickly.
D.Asserting that incoming data matches an expected structure (columns, types, nullability) before it is accepted into a pipeline.
Answer + AI explanation with Pro
29. Which term means: "The practice of monitoring data health across freshness, volume, schema, distribution, and lineage to detect and resolve issues quickly."?
Mid
A.freshness check
B.schema drift
C.data observability
D.Soda check
Answer + AI explanation with Pro
30. Which statement is correct?
Mid
A.data observability — A formal, versioned agreement between data producers and consumers specifying schema, semantics, and quality guarantees, enforced in CI or at ingestion.
B.data observability — A dbt test that feeds a model fixed mock rows and asserts the transformed output, validating logic independently of warehouse data.
C.data observability — A named collection of Great Expectations assertions about a dataset that together define its quality contract.
D.data observability — The practice of monitoring data health across freshness, volume, schema, distribution, and lineage to detect and resolve issues quickly.
Answer + AI explanation with Pro
Showing 30 of 72 Data Quality questions — the full set, with answers, explanations and an AI tutor on every question, is inside.
Free to start
Start with a free readiness check
Sign up free for the 2-minute IT readiness check and a scored result. Answers, explanations and the AI tutor on every Data Quality question come with Pro.