36 File Formats questions from the Data Engineering bank, written for Indian campus drives and tech interviews. Every question has a verified answer and an AI-tutor explanation on placd.
Free to start: the 2-minute IT readiness check — six questions and a result.
A.Pushing filter conditions down to the storage layer so only matching row groups or files are read, using column statistics to skip data.
B.Storing values of each column together so analytic scans read only needed columns and compress better, the basis of Parquet and ORC.
C.An optimized row-columnar file format from the Hadoop ecosystem with strong compression, lightweight indexes, and built-in statistics.
D.Hierarchical structures (structs, arrays, maps) that columnar formats like Parquet encode with repetition and definition levels.
Answer + AI explanation with Pro
2. Which term means: "Storing values of each column together so analytic scans read only needed columns and compress better, the basis of Parquet and ORC."?
Mid
A.Avro
B.columnar storage
C.run-length encoding
D.compaction
Answer + AI explanation with Pro
3. Which statement is correct?
Mid
A.columnar storage — Pushing filter conditions down to the storage layer so only matching row groups or files are read, using column statistics to skip data.
B.columnar storage — A row-oriented binary format with a self-describing schema, well suited to streaming and write-heavy workloads and schema evolution.
C.columnar storage — An optimized row-columnar file format from the Hadoop ecosystem with strong compression, lightweight indexes, and built-in statistics.
D.columnar storage — Storing values of each column together so analytic scans read only needed columns and compress better, the basis of Parquet and ORC.
Answer + AI explanation with Pro
4. What is ORC?
Mid
A.Performance degradation from many tiny files that bloat metadata and scheduling overhead, mitigated by compaction into larger files.
B.Hierarchical structures (structs, arrays, maps) that columnar formats like Parquet encode with repetition and definition levels.
C.A compression technique replacing repeated values with small integer codes referencing a dictionary, very effective for low-cardinality columns.
D.An optimized row-columnar file format from the Hadoop ecosystem with strong compression, lightweight indexes, and built-in statistics.
Answer + AI explanation with Pro
5. Which term means: "An optimized row-columnar file format from the Hadoop ecosystem with strong compression, lightweight indexes, and built-in statistics."?
Mid
A.ORC
B.predicate pushdown
C.page
D.Avro
Answer + AI explanation with Pro
6. Which statement is correct?
Mid
A.ORC — A horizontal chunk of a Parquet file holding column chunks plus min/max statistics, the unit at which readers prune and parallelize.
B.ORC — A compression technique replacing repeated values with small integer codes referencing a dictionary, very effective for low-cardinality columns.
C.ORC — Performance degradation from many tiny files that bloat metadata and scheduling overhead, mitigated by compaction into larger files.
D.ORC — An optimized row-columnar file format from the Hadoop ecosystem with strong compression, lightweight indexes, and built-in statistics.
Answer + AI explanation with Pro
7. What is predicate pushdown?
Senior
A.A compression method storing repeated consecutive values as a value plus a count, paired with dictionary encoding in Parquet.
B.Applying structure to data at query time rather than at write time, the flexibility that lets lakes store raw semi-structured files.
C.Pushing filter conditions down to the storage layer so only matching row groups or files are read, using column statistics to skip data.
D.The smallest unit of storage and compression within a Parquet column chunk, holding encoded values, definition, and repetition levels.
Answer + AI explanation with Pro
8. Which term means: "Pushing filter conditions down to the storage layer so only matching row groups or files are read, using column statistics to skip data."?
Senior
A.predicate pushdown
B.schema-on-read
C.small files problem
D.columnar storage
Answer + AI explanation with Pro
9. Which statement is correct?
Senior
A.predicate pushdown — A maintenance operation that rewrites many small data files into fewer large ones to improve scan efficiency in lakes and open table formats.
B.predicate pushdown — Pushing filter conditions down to the storage layer so only matching row groups or files are read, using column statistics to skip data.
C.predicate pushdown — A horizontal chunk of a Parquet file holding column chunks plus min/max statistics, the unit at which readers prune and parallelize.
D.predicate pushdown — A row-oriented binary format with a self-describing schema, well suited to streaming and write-heavy workloads and schema evolution.
Answer + AI explanation with Pro
10. What is row group?
Senior
A.Hierarchical structures (structs, arrays, maps) that columnar formats like Parquet encode with repetition and definition levels.
B.A horizontal chunk of a Parquet file holding column chunks plus min/max statistics, the unit at which readers prune and parallelize.
C.A row-oriented binary format with a self-describing schema, well suited to streaming and write-heavy workloads and schema evolution.
D.Pushing filter conditions down to the storage layer so only matching row groups or files are read, using column statistics to skip data.
Answer + AI explanation with Pro
11. Which term means: "A horizontal chunk of a Parquet file holding column chunks plus min/max statistics, the unit at which readers prune and parallelize."?
Senior
A.columnar storage
B.compaction
C.row group
D.small files problem
Answer + AI explanation with Pro
12. Which statement is correct?
Senior
A.row group — The smallest unit of storage and compression within a Parquet column chunk, holding encoded values, definition, and repetition levels.
B.row group — A horizontal chunk of a Parquet file holding column chunks plus min/max statistics, the unit at which readers prune and parallelize.
C.row group — A row-oriented binary format with a self-describing schema, well suited to streaming and write-heavy workloads and schema evolution.
D.row group — An optimized row-columnar file format from the Hadoop ecosystem with strong compression, lightweight indexes, and built-in statistics.
Answer + AI explanation with Pro
13. What is dictionary encoding?
Mid
A.A maintenance operation that rewrites many small data files into fewer large ones to improve scan efficiency in lakes and open table formats.
B.A compression technique replacing repeated values with small integer codes referencing a dictionary, very effective for low-cardinality columns.
C.Hierarchical structures (structs, arrays, maps) that columnar formats like Parquet encode with repetition and definition levels.
D.An optimized row-columnar file format from the Hadoop ecosystem with strong compression, lightweight indexes, and built-in statistics.
Answer + AI explanation with Pro
14. Which term means: "A compression technique replacing repeated values with small integer codes referencing a dictionary, very effective for low-cardinality columns."?
Mid
A.compaction
B.dictionary encoding
C.columnar storage
D.run-length encoding
Answer + AI explanation with Pro
15. Which statement is correct?
Mid
A.dictionary encoding — The smallest unit of storage and compression within a Parquet column chunk, holding encoded values, definition, and repetition levels.
B.dictionary encoding — A compression technique replacing repeated values with small integer codes referencing a dictionary, very effective for low-cardinality columns.
C.dictionary encoding — A compression method storing repeated consecutive values as a value plus a count, paired with dictionary encoding in Parquet.
D.dictionary encoding — A maintenance operation that rewrites many small data files into fewer large ones to improve scan efficiency in lakes and open table formats.
Answer + AI explanation with Pro
16. What is small files problem?
Mid
A.The smallest unit of storage and compression within a Parquet column chunk, holding encoded values, definition, and repetition levels.
B.A compression method storing repeated consecutive values as a value plus a count, paired with dictionary encoding in Parquet.
C.A maintenance operation that rewrites many small data files into fewer large ones to improve scan efficiency in lakes and open table formats.
D.Performance degradation from many tiny files that bloat metadata and scheduling overhead, mitigated by compaction into larger files.
Answer + AI explanation with Pro
17. Which term means: "Performance degradation from many tiny files that bloat metadata and scheduling overhead, mitigated by compaction into larger files."?
Mid
A.nested data
B.schema-on-read
C.dictionary encoding
D.small files problem
Answer + AI explanation with Pro
18. Which statement is correct?
Mid
A.small files problem — Pushing filter conditions down to the storage layer so only matching row groups or files are read, using column statistics to skip data.
B.small files problem — Storing values of each column together so analytic scans read only needed columns and compress better, the basis of Parquet and ORC.
C.small files problem — Performance degradation from many tiny files that bloat metadata and scheduling overhead, mitigated by compaction into larger files.
D.small files problem — Hierarchical structures (structs, arrays, maps) that columnar formats like Parquet encode with repetition and definition levels.
Answer + AI explanation with Pro
19. What is compaction?
Senior
A.Hierarchical structures (structs, arrays, maps) that columnar formats like Parquet encode with repetition and definition levels.
B.Performance degradation from many tiny files that bloat metadata and scheduling overhead, mitigated by compaction into larger files.
C.Pushing filter conditions down to the storage layer so only matching row groups or files are read, using column statistics to skip data.
D.A maintenance operation that rewrites many small data files into fewer large ones to improve scan efficiency in lakes and open table formats.
Answer + AI explanation with Pro
20. Which term means: "A maintenance operation that rewrites many small data files into fewer large ones to improve scan efficiency in lakes and open table formats."?
Senior
A.page
B.ORC
C.compaction
D.dictionary encoding
Answer + AI explanation with Pro
21. Which statement is correct?
Senior
A.compaction — A compression technique replacing repeated values with small integer codes referencing a dictionary, very effective for low-cardinality columns.
B.compaction — Pushing filter conditions down to the storage layer so only matching row groups or files are read, using column statistics to skip data.
C.compaction — A maintenance operation that rewrites many small data files into fewer large ones to improve scan efficiency in lakes and open table formats.
D.compaction — A row-oriented binary format with a self-describing schema, well suited to streaming and write-heavy workloads and schema evolution.
Answer + AI explanation with Pro
22. What is schema-on-read?
Mid
A.Applying structure to data at query time rather than at write time, the flexibility that lets lakes store raw semi-structured files.
B.A compression method storing repeated consecutive values as a value plus a count, paired with dictionary encoding in Parquet.
C.A maintenance operation that rewrites many small data files into fewer large ones to improve scan efficiency in lakes and open table formats.
D.A row-oriented binary format with a self-describing schema, well suited to streaming and write-heavy workloads and schema evolution.
Answer + AI explanation with Pro
23. Which term means: "Applying structure to data at query time rather than at write time, the flexibility that lets lakes store raw semi-structured files."?
Mid
A.run-length encoding
B.Avro
C.schema-on-read
D.ORC
Answer + AI explanation with Pro
24. Which statement is correct?
Mid
A.schema-on-read — A horizontal chunk of a Parquet file holding column chunks plus min/max statistics, the unit at which readers prune and parallelize.
B.schema-on-read — Storing values of each column together so analytic scans read only needed columns and compress better, the basis of Parquet and ORC.
C.schema-on-read — The smallest unit of storage and compression within a Parquet column chunk, holding encoded values, definition, and repetition levels.
D.schema-on-read — Applying structure to data at query time rather than at write time, the flexibility that lets lakes store raw semi-structured files.
Answer + AI explanation with Pro
25. What is Avro?
Junior
A.An optimized row-columnar file format from the Hadoop ecosystem with strong compression, lightweight indexes, and built-in statistics.
B.Applying structure to data at query time rather than at write time, the flexibility that lets lakes store raw semi-structured files.
C.A row-oriented binary format with a self-describing schema, well suited to streaming and write-heavy workloads and schema evolution.
D.Pushing filter conditions down to the storage layer so only matching row groups or files are read, using column statistics to skip data.
Answer + AI explanation with Pro
26. Which term means: "A row-oriented binary format with a self-describing schema, well suited to streaming and write-heavy workloads and schema evolution."?
Junior
A.Avro
B.dictionary encoding
C.page
D.ORC
Answer + AI explanation with Pro
27. Which statement is correct?
Junior
A.Avro — A maintenance operation that rewrites many small data files into fewer large ones to improve scan efficiency in lakes and open table formats.
B.Avro — A horizontal chunk of a Parquet file holding column chunks plus min/max statistics, the unit at which readers prune and parallelize.
C.Avro — Applying structure to data at query time rather than at write time, the flexibility that lets lakes store raw semi-structured files.
D.Avro — A row-oriented binary format with a self-describing schema, well suited to streaming and write-heavy workloads and schema evolution.
Answer + AI explanation with Pro
28. What is page?
Mid
A.Applying structure to data at query time rather than at write time, the flexibility that lets lakes store raw semi-structured files.
B.A compression method storing repeated consecutive values as a value plus a count, paired with dictionary encoding in Parquet.
C.The smallest unit of storage and compression within a Parquet column chunk, holding encoded values, definition, and repetition levels.
D.A compression technique replacing repeated values with small integer codes referencing a dictionary, very effective for low-cardinality columns.
Answer + AI explanation with Pro
29. Which term means: "The smallest unit of storage and compression within a Parquet column chunk, holding encoded values, definition, and repetition levels."?
Mid
A.Avro
B.predicate pushdown
C.page
D.schema-on-read
Answer + AI explanation with Pro
30. Which statement is correct?
Mid
A.page — Hierarchical structures (structs, arrays, maps) that columnar formats like Parquet encode with repetition and definition levels.
B.page — A compression method storing repeated consecutive values as a value plus a count, paired with dictionary encoding in Parquet.
C.page — The smallest unit of storage and compression within a Parquet column chunk, holding encoded values, definition, and repetition levels.
D.page — Performance degradation from many tiny files that bloat metadata and scheduling overhead, mitigated by compaction into larger files.
Answer + AI explanation with Pro
Showing 30 of 36 File Formats questions — the full set, with answers, explanations and an AI tutor on every question, is inside.
Free to start
Start with a free readiness check
Sign up free for the 2-minute IT readiness check and a scored result. Answers, explanations and the AI tutor on every File Formats question come with Pro.