File Formats interview questions

63 File Formats questions from the Big Data bank, written for Indian campus drives and tech interviews. Every question has a verified answer and an AI-tutor explanation on placd.

Free to start: the 2-minute IT readiness check — six questions and a result.

Take the free IT readiness check

or take a mock interview set up for this area

1. What is Row group?

Junior
  1. A.a Parquet column encoding that replaces repeated values with small dictionary indexes, shrinking storage and speeding equality filters
  2. B.the contiguous storage of one column's data for a single Parquet row group, the unit of column pruning
  3. C.the Avro mechanism that lets readers and writers use different schema versions through writer and reader schema resolution
  4. D.the horizontal partition of rows in a Parquet file within which each column is stored together as a chunk

Answer + AI explanation with Pro

2. Which term means: "the horizontal partition of rows in a Parquet file within which each column is stored together as a chunk"?

Junior
  1. A.Row group
  2. B.Parquet page index
  3. C.Predicate pushdown via min-max
  4. D.ORC predicate pushdown

Answer + AI explanation with Pro

3. Which statement is correct?

Junior
  1. A.Row group — the Avro mechanism that lets readers and writers use different schema versions through writer and reader schema resolution
  2. B.Row group — a probabilistic index in Parquet and ORC that quickly rules out row groups not containing a sought value with no false negatives
  3. C.Row group — the horizontal partition of rows in a Parquet file within which each column is stored together as a chunk
  4. D.Row group — a columnar encoding that replaces repeated values with small integer codes referencing a per-chunk dictionary

Answer + AI explanation with Pro

4. What is Stripe?

Junior
  1. A.the equivalent of a row group in ORC, a self-contained set of rows holding column data plus its indexes
  2. B.a Parquet column encoding that replaces repeated values with small dictionary indexes, shrinking storage and speeding equality filters
  3. C.an Avro annotation layering semantic meaning (such as decimal, date, or timestamp-millis) on top of a primitive physical type
  4. D.a compression technique that stores a repeated value once together with the count of its consecutive occurrences

Answer + AI explanation with Pro

5. Which term means: "the equivalent of a row group in ORC, a self-contained set of rows holding column data plus its indexes"?

Junior
  1. A.Protobuf
  2. B.Stripe
  3. C.Footer metadata
  4. D.row group

Answer + AI explanation with Pro

6. Which statement is correct?

Junior
  1. A.Stripe — the equivalent of a row group in ORC, a self-contained set of rows holding column data plus its indexes
  2. B.Stripe — a file format that can be divided at block boundaries so multiple tasks read one file in parallel, as Parquet and ORC allow
  3. C.Stripe — the contiguous storage of one column's data for a single Parquet row group, the unit of column pruning
  4. D.Stripe — an Avro annotation layering semantic meaning (such as decimal, date, or timestamp-millis) on top of a primitive physical type

Answer + AI explanation with Pro

7. What is Column chunk?

Junior
  1. A.Protocol Buffers, a compact binary serialization format using numbered fields and a schema, common for Kafka messages and service APIs
  2. B.a probabilistic index in Parquet and ORC that quickly rules out row groups not containing a sought value with no false negatives
  3. C.the contiguous storage of one column's data for a single Parquet row group, the unit of column pruning
  4. D.the smallest unit of encoding and compression within a Parquet column chunk, also carrying its own statistics

Answer + AI explanation with Pro

8. Which term means: "the contiguous storage of one column's data for a single Parquet row group, the unit of column pruning"?

Junior
  1. A.Arrow Flight
  2. B.Stripe
  3. C.Column chunk
  4. D.row group

Answer + AI explanation with Pro

9. Which statement is correct?

Junior
  1. A.Column chunk — the contiguous storage of one column's data for a single Parquet row group, the unit of column pruning
  2. B.Column chunk — the in-memory columnar format enabling zero-copy data exchange between engines and languages without serialization overhead
  3. C.Column chunk — a file format that can be divided at block boundaries so multiple tasks read one file in parallel, as Parquet and ORC allow
  4. D.Column chunk — the equivalent of a row group in ORC, a self-contained set of rows holding column data plus its indexes

Answer + AI explanation with Pro

10. What is Dictionary encoding?

Mid
  1. A.the ORC reader capability of skipping stripes and row groups using built-in min/max indexes and bloom filters when a filter is supplied
  2. B.the contiguous storage of one column's data for a single Parquet row group, the unit of column pruning
  3. C.a columnar encoding that replaces repeated values with small integer codes referencing a per-chunk dictionary
  4. D.a modern compression codec offering high ratios at fast decompression speeds, increasingly the default over Snappy and Gzip in columnar formats

Answer + AI explanation with Pro

11. Which term means: "a columnar encoding that replaces repeated values with small integer codes referencing a per-chunk dictionary"?

Mid
  1. A.Schema evolution in Avro
  2. B.Page
  3. C.Dictionary encoding
  4. D.Run-length encoding

Answer + AI explanation with Pro

12. Which statement is correct?

Mid
  1. A.Dictionary encoding — Protocol Buffers, a compact binary serialization format using numbered fields and a schema, common for Kafka messages and service APIs
  2. B.Dictionary encoding — the ORC reader capability of skipping stripes and row groups using built-in min/max indexes and bloom filters when a filter is supplied
  3. C.Dictionary encoding — the contiguous storage of one column's data for a single Parquet row group, the unit of column pruning
  4. D.Dictionary encoding — a columnar encoding that replaces repeated values with small integer codes referencing a per-chunk dictionary

Answer + AI explanation with Pro

13. What is Run-length encoding?

Mid
  1. A.Protocol Buffers, a compact binary serialization format using numbered fields and a schema, common for Kafka messages and service APIs
  2. B.the in-memory columnar format enabling zero-copy data exchange between engines and languages without serialization overhead
  3. C.a Parquet column encoding that replaces repeated values with small dictionary indexes, shrinking storage and speeding equality filters
  4. D.a compression technique that stores a repeated value once together with the count of its consecutive occurrences

Answer + AI explanation with Pro

14. Which term means: "a compression technique that stores a repeated value once together with the count of its consecutive occurrences"?

Mid
  1. A.Protobuf
  2. B.Run-length encoding
  3. C.ZSTD
  4. D.Columnar vs row-oriented storage

Answer + AI explanation with Pro

15. Which statement is correct?

Mid
  1. A.Run-length encoding — the ORC reader capability of skipping stripes and row groups using built-in min/max indexes and bloom filters when a filter is supplied
  2. B.Run-length encoding — a compression technique that stores a repeated value once together with the count of its consecutive occurrences
  3. C.Run-length encoding — Protocol Buffers, a compact binary serialization format using numbered fields and a schema, common for Kafka messages and service APIs
  4. D.Run-length encoding — the tradeoff where columnar formats like Parquet excel at analytical scans while row formats like Avro suit record-at-a-time writes

Answer + AI explanation with Pro

16. What is Bloom filter?

Mid
  1. A.the smallest unit of encoding and compression within a Parquet column chunk, also carrying its own statistics
  2. B.a columnar encoding that replaces repeated values with small integer codes referencing a per-chunk dictionary
  3. C.a probabilistic index in Parquet and ORC that quickly rules out row groups not containing a sought value with no false negatives
  4. D.a modern compression codec offering high ratios at fast decompression speeds, increasingly the default over Snappy and Gzip in columnar formats

Answer + AI explanation with Pro

17. Which term means: "a probabilistic index in Parquet and ORC that quickly rules out row groups not containing a sought value with no false negatives"?

Mid
  1. A.ZSTD
  2. B.Row group
  3. C.Bloom filter
  4. D.Parquet dictionary encoding

Answer + AI explanation with Pro

18. Which statement is correct?

Mid
  1. A.Bloom filter — the horizontal partition of rows in a Parquet file within which each column is stored together as a chunk
  2. B.Bloom filter — using the per-chunk minimum and maximum statistics in Parquet or ORC footers to skip chunks that cannot match a filter
  3. C.Bloom filter — the ORC reader capability of skipping stripes and row groups using built-in min/max indexes and bloom filters when a filter is supplied
  4. D.Bloom filter — a probabilistic index in Parquet and ORC that quickly rules out row groups not containing a sought value with no false negatives

Answer + AI explanation with Pro

19. What is Predicate pushdown via min-max?

Mid
  1. A.the horizontal Parquet partition holding column chunks for a block of rows, the unit at which statistics and parallel reads operate
  2. B.using the per-chunk minimum and maximum statistics in Parquet or ORC footers to skip chunks that cannot match a filter
  3. C.a columnar encoding that replaces repeated values with small integer codes referencing a per-chunk dictionary
  4. D.the Avro mechanism that lets readers and writers use different schema versions through writer and reader schema resolution

Answer + AI explanation with Pro

20. Which term means: "using the per-chunk minimum and maximum statistics in Parquet or ORC footers to skip chunks that cannot match a filter"?

Mid
  1. A.ORC predicate pushdown
  2. B.Protobuf
  3. C.ZSTD
  4. D.Predicate pushdown via min-max

Answer + AI explanation with Pro

21. Which statement is correct?

Mid
  1. A.Predicate pushdown via min-max — the horizontal Parquet partition holding column chunks for a block of rows, the unit at which statistics and parallel reads operate
  2. B.Predicate pushdown via min-max — a file format that can be divided at block boundaries so multiple tasks read one file in parallel, as Parquet and ORC allow
  3. C.Predicate pushdown via min-max — using the per-chunk minimum and maximum statistics in Parquet or ORC footers to skip chunks that cannot match a filter
  4. D.Predicate pushdown via min-max — a columnar encoding that replaces repeated values with small integer codes referencing a per-chunk dictionary

Answer + AI explanation with Pro

22. What is Page?

Mid
  1. A.the horizontal partition of rows in a Parquet file within which each column is stored together as a chunk
  2. B.the Arrow RPC framework that streams columnar record batches over the network at high throughput without row-by-row serialization
  3. C.the smallest unit of encoding and compression within a Parquet column chunk, also carrying its own statistics
  4. D.an Avro annotation layering semantic meaning (such as decimal, date, or timestamp-millis) on top of a primitive physical type

Answer + AI explanation with Pro

23. Which term means: "the smallest unit of encoding and compression within a Parquet column chunk, also carrying its own statistics"?

Mid
  1. A.Schema evolution in Avro
  2. B.Page
  3. C.ORC predicate pushdown
  4. D.Predicate pushdown via min-max

Answer + AI explanation with Pro

24. Which statement is correct?

Mid
  1. A.Page — the horizontal Parquet partition holding column chunks for a block of rows, the unit at which statistics and parallel reads operate
  2. B.Page — a columnar encoding that replaces repeated values with small integer codes referencing a per-chunk dictionary
  3. C.Page — the in-memory columnar format enabling zero-copy data exchange between engines and languages without serialization overhead
  4. D.Page — the smallest unit of encoding and compression within a Parquet column chunk, also carrying its own statistics

Answer + AI explanation with Pro

25. What is Footer metadata?

Senior
  1. A.a file format that can be divided at block boundaries so multiple tasks read one file in parallel, as Parquet and ORC allow
  2. B.the smallest unit of encoding and compression within a Parquet column chunk, also carrying its own statistics
  3. C.an Avro annotation layering semantic meaning (such as decimal, date, or timestamp-millis) on top of a primitive physical type
  4. D.the trailing section of a Parquet or ORC file holding the schema, row-group offsets, and statistics needed to read it

Answer + AI explanation with Pro

26. Which term means: "the trailing section of a Parquet or ORC file holding the schema, row-group offsets, and statistics needed to read it"?

Senior
  1. A.Avro logical type
  2. B.Stripe
  3. C.Parquet dictionary encoding
  4. D.Footer metadata

Answer + AI explanation with Pro

27. Which statement is correct?

Senior
  1. A.Footer metadata — Protocol Buffers, a compact binary serialization format using numbered fields and a schema, common for Kafka messages and service APIs
  2. B.Footer metadata — the ORC reader capability of skipping stripes and row groups using built-in min/max indexes and bloom filters when a filter is supplied
  3. C.Footer metadata — the trailing section of a Parquet or ORC file holding the schema, row-group offsets, and statistics needed to read it
  4. D.Footer metadata — the in-memory columnar format enabling zero-copy data exchange between engines and languages without serialization overhead

Answer + AI explanation with Pro

28. What is Schema evolution in Avro?

Senior
  1. A.the Avro mechanism that lets readers and writers use different schema versions through writer and reader schema resolution
  2. B.the trailing section of a Parquet or ORC file holding the schema, row-group offsets, and statistics needed to read it
  3. C.a compression technique that stores a repeated value once together with the count of its consecutive occurrences
  4. D.a Parquet column encoding that replaces repeated values with small dictionary indexes, shrinking storage and speeding equality filters

Answer + AI explanation with Pro

29. Which term means: "the Avro mechanism that lets readers and writers use different schema versions through writer and reader schema resolution"?

Senior
  1. A.Parquet dictionary encoding
  2. B.Bloom filter
  3. C.row group
  4. D.Schema evolution in Avro

Answer + AI explanation with Pro

30. Which statement is correct?

Senior
  1. A.Schema evolution in Avro — the smallest unit of encoding and compression within a Parquet column chunk, also carrying its own statistics
  2. B.Schema evolution in Avro — the Avro mechanism that lets readers and writers use different schema versions through writer and reader schema resolution
  3. C.Schema evolution in Avro — the trailing section of a Parquet or ORC file holding the schema, row-group offsets, and statistics needed to read it
  4. D.Schema evolution in Avro — the in-memory columnar format enabling zero-copy data exchange between engines and languages without serialization overhead

Answer + AI explanation with Pro

Showing 30 of 63 File Formats questions — the full set, with answers, explanations and an AI tutor on every question, is inside.

Free to start

Start with a free readiness check

Sign up free for the 2-minute IT readiness check and a scored result. Answers, explanations and the AI tutor on every File Formats question come with Pro.

Take the free IT readiness check

or take a mock interview set up for this area

24,000+ questions & coding problemsSoftware & IT16,274 questionsGovernment jobs26 examsAptitudenew questions every timeAI practice interviewwith feedback65 topics to practiseMechanical1,149 questionsGATE ME9 papersEngineering Mathematics381 questions2-minute checkfreeDSA Problems1,422Civil1,005 questionsGATE CE9 papersCS Fundamentals1,209 questionsYour scores6 skillsSystem Design25Electrical / EEE1,047 questionsGATE EE9 papersRun your codeC++ · Java · PythonLow-Level Design144Electronics & Comm.975 questionsGATE EC9 papersAI help on every questionFull-Stack6,282Chemical1,005 questionsGATE CH9 papersAI whiteboardsystem designWork abroadEurope · remote · transfersESE ME1 paperGATE practice papers2019–2026ESE CE1 paperDate alertsbefore the last dateESE EE1 paperBehavioural courseHR round practiceESE ET1 paperResume optimizerProSSC JE ME1 paperApplication trackerSSC JE CE1 paperCompany-wise prepSSC JE EE1 paperRole roadmapsRRB JE1 subjectPriced in ₹UPI · cardsISRO SC1 paperGATE CS9 papersIBPS SO IT1 paperUGC NET CS1 paperSSC CGL26 papersIBPS PO26 papersRRB NTPC26 papersSSC CHSL26 papersIBPS Clerk26 papersSBI Clerk26 papersRRB Group D26 papersSSC CPO26 papersSSC GD26 papers