87 Cloud Platforms questions from the Big Data bank, written for Indian campus drives and tech interviews. Every question has a verified answer and an AI-tutor explanation on placd.
Free to start: the 2-minute IT readiness check — six questions and a result.
A.the AWS managed service for running Spark, Hive, and other big-data frameworks on provisioned or serverless clusters
B.the Azure unified SaaS analytics platform that brings data engineering, warehousing, and BI together on OneLake
C.the BigQuery capability to train and run machine-learning models using SQL directly on warehouse data
D.the single, tenant-wide data lake underlying Microsoft Fabric where all workloads store data in open Delta format
Answer + AI explanation with Pro
2. Which term means: "the AWS managed service for running Spark, Hive, and other big-data frameworks on provisioned or serverless clusters"?
Junior
A.Amazon EMR
B.Dataproc Serverless
C.BigLake
D.Amazon Athena
Answer + AI explanation with Pro
3. Which statement is correct?
Junior
A.Amazon EMR — the AWS S3 bucket type providing managed Apache Iceberg storage with built-in maintenance like compaction and snapshot expiration
B.Amazon EMR — a unit of BigQuery compute capacity that the engine allocates to execute the stages of a query
C.Amazon EMR — the AWS managed service for running Spark, Hive, and other big-data frameworks on provisioned or serverless clusters
D.Amazon EMR — the AWS option that runs Spark and Hive jobs on automatically provisioned capacity without managing clusters
Answer + AI explanation with Pro
4. What is AWS Glue?
Junior
A.the Google Cloud storage engine unifying governance and fine-grained access over open table formats across BigQuery and external engines
B.the AWS serverless ETL service with a managed Spark runtime and a central data catalog for schema discovery
C.the in-memory acceleration layer for BigQuery that caches data to deliver sub-second responses for dashboards
D.the AWS serverless query service that runs Trino-based SQL directly over data in S3 with no clusters to manage
Answer + AI explanation with Pro
5. Which term means: "the AWS serverless ETL service with a managed Spark runtime and a central data catalog for schema discovery"?
Junior
A.Azure Synapse Analytics
B.Athena
C.AWS Glue
D.S3 Tables
Answer + AI explanation with Pro
6. Which statement is correct?
Junior
A.AWS Glue — the AWS serverless ETL service with a managed Spark runtime and a central data catalog for schema discovery
B.AWS Glue — the GCP fully managed service for running Apache Beam batch and streaming pipelines with autoscaling
C.AWS Glue — the GCP serverless data warehouse that separates storage from compute and queries petabytes with standard SQL
D.AWS Glue — the Azure unified SaaS analytics platform that brings data engineering, warehousing, and BI together on OneLake
Answer + AI explanation with Pro
7. What is Amazon Athena?
Junior
A.the AWS Glue component that scans data stores, infers schemas, and populates table definitions in the Glue Data Catalog
B.the GCP managed service for running Spark and Hadoop clusters, including ephemeral and serverless options
C.the AWS serverless query service that runs SQL directly over data in S3 using Presto and Trino engines
D.the Google Cloud service that runs Spark batch and interactive workloads without provisioning or managing clusters
Answer + AI explanation with Pro
8. Which term means: "the AWS serverless query service that runs SQL directly over data in S3 using Presto and Trino engines"?
Junior
A.Trino
B.OneLake shortcut
C.Google BigQuery
D.Amazon Athena
Answer + AI explanation with Pro
9. Which statement is correct?
Junior
A.Amazon Athena — the GCP managed service for running Spark and Hadoop clusters, including ephemeral and serverless options
B.Amazon Athena — the AWS serverless query service that runs SQL directly over data in S3 using Presto and Trino engines
C.Amazon Athena — the AWS managed service for running Spark, Hive, and other big-data frameworks on provisioned or serverless clusters
D.Amazon Athena — the AWS option that runs Spark and Hive jobs on automatically provisioned capacity without managing clusters
Answer + AI explanation with Pro
10. What is Google BigQuery?
Junior
A.the GCP serverless data warehouse that separates storage from compute and queries petabytes with standard SQL
B.the GCP fully managed service for running Apache Beam batch and streaming pipelines with autoscaling
C.the Google Cloud service that runs Spark batch and interactive workloads without provisioning or managing clusters
D.a unit of BigQuery compute capacity that the engine allocates to execute the stages of a query
Answer + AI explanation with Pro
11. Which term means: "the GCP serverless data warehouse that separates storage from compute and queries petabytes with standard SQL"?
Junior
A.Google BigQuery
B.Fabric shortcut
C.BigQuery BI Engine
D.EMR Serverless
Answer + AI explanation with Pro
12. Which statement is correct?
Junior
A.Google BigQuery — the GCP serverless data warehouse that separates storage from compute and queries petabytes with standard SQL
B.Google BigQuery — the Google Cloud storage engine unifying governance and fine-grained access over open table formats across BigQuery and external engines
C.Google BigQuery — the GCP fully managed service for running Apache Beam batch and streaming pipelines with autoscaling
D.Google BigQuery — the AWS S3 bucket type providing managed Apache Iceberg storage with built-in maintenance like compaction and snapshot expiration
Answer + AI explanation with Pro
13. What is Google Dataproc?
Mid
A.the Google Cloud service that runs Spark batch and interactive workloads without provisioning or managing clusters
B.the in-memory acceleration layer for BigQuery that caches data to deliver sub-second responses for dashboards
C.the AWS metadata store, compatible with the Hive metastore, that many AWS analytics services share for schema information
D.the GCP managed service for running Spark and Hadoop clusters, including ephemeral and serverless options
Answer + AI explanation with Pro
14. Which term means: "the GCP managed service for running Spark and Hadoop clusters, including ephemeral and serverless options"?
Mid
A.Glue Data Catalog
B.BigQuery BI Engine
C.Google Dataflow
D.Google Dataproc
Answer + AI explanation with Pro
15. Which statement is correct?
Mid
A.Google Dataproc — the AWS option that runs Spark and Hive jobs on automatically provisioned capacity without managing clusters
B.Google Dataproc — the Google Cloud service that runs Spark batch and interactive workloads without provisioning or managing clusters
C.Google Dataproc — the AWS serverless query service that runs SQL directly over data in S3 using Presto and Trino engines
D.Google Dataproc — the GCP managed service for running Spark and Hadoop clusters, including ephemeral and serverless options
Answer + AI explanation with Pro
16. What is Google Dataflow?
Mid
A.the GCP fully managed service for running Apache Beam batch and streaming pipelines with autoscaling
B.the BigQuery capability to train and run machine-learning models using SQL directly on warehouse data
C.the distributed SQL query engine that federates queries across data lakes and databases without moving the data
D.the AWS serverless ETL service with a managed Spark runtime and a central data catalog for schema discovery
Answer + AI explanation with Pro
17. Which term means: "the GCP fully managed service for running Apache Beam batch and streaming pipelines with autoscaling"?
Mid
A.Google Dataflow
B.Glue crawler
C.Google BigQuery
D.AWS Glue
Answer + AI explanation with Pro
18. Which statement is correct?
Mid
A.Google Dataflow — the Amazon Redshift feature that queries external data in S3 directly, joining it with tables stored in the warehouse
B.Google Dataflow — a unit of BigQuery compute capacity that the engine allocates to execute the stages of a query
C.Google Dataflow — the GCP fully managed service for running Apache Beam batch and streaming pipelines with autoscaling
D.Google Dataflow — the single unified, tenant-wide data lake in Microsoft Fabric where all workloads store data in open Delta and Parquet formats
Answer + AI explanation with Pro
19. What is Azure Synapse Analytics?
Mid
A.a unit of BigQuery compute capacity that the engine allocates to execute the stages of a query
B.the Azure platform unifying data warehousing and big-data analytics over dedicated and serverless SQL and Spark pools
C.the Google Cloud service that runs Spark batch and interactive workloads without provisioning or managing clusters
D.a Microsoft Fabric reference that virtualizes external data into OneLake without copying it
Answer + AI explanation with Pro
20. Which term means: "the Azure platform unifying data warehousing and big-data analytics over dedicated and serverless SQL and Spark pools"?
Mid
A.Amazon Athena
B.BigQuery ML
C.AWS Glue
D.Azure Synapse Analytics
Answer + AI explanation with Pro
21. Which statement is correct?
Mid
A.Azure Synapse Analytics — the Azure platform unifying data warehousing and big-data analytics over dedicated and serverless SQL and Spark pools
B.Azure Synapse Analytics — a unit of BigQuery compute capacity that the engine allocates to execute the stages of a query
C.Azure Synapse Analytics — the AWS metadata store, compatible with the Hive metastore, that many AWS analytics services share for schema information
D.Azure Synapse Analytics — the distributed SQL query engine that federates queries across data lakes and databases without moving the data
Answer + AI explanation with Pro
22. What is Microsoft Fabric?
Mid
A.the single, tenant-wide data lake underlying Microsoft Fabric where all workloads store data in open Delta format
B.the AWS option that runs Spark and Hive jobs on automatically provisioned capacity without managing clusters
C.the AWS serverless ETL service with a managed Spark runtime and a central data catalog for schema discovery
D.the Azure unified SaaS analytics platform that brings data engineering, warehousing, and BI together on OneLake
Answer + AI explanation with Pro
23. Which term means: "the Azure unified SaaS analytics platform that brings data engineering, warehousing, and BI together on OneLake"?
Mid
A.EMR Serverless
B.Glue Data Catalog
C.Trino
D.Microsoft Fabric
Answer + AI explanation with Pro
24. Which statement is correct?
Mid
A.Microsoft Fabric — the GCP fully managed service for running Apache Beam batch and streaming pipelines with autoscaling
B.Microsoft Fabric — the Azure unified SaaS analytics platform that brings data engineering, warehousing, and BI together on OneLake
C.Microsoft Fabric — the distributed SQL query engine that federates queries across data lakes and databases without moving the data
D.Microsoft Fabric — the Azure platform unifying data warehousing and big-data analytics over dedicated and serverless SQL and Spark pools
Answer + AI explanation with Pro
25. What is OneLake?
Mid
A.the Google Cloud service that runs Spark batch and interactive workloads without provisioning or managing clusters
B.the AWS serverless ETL service with a managed Spark runtime and a central data catalog for schema discovery
C.the single, tenant-wide data lake underlying Microsoft Fabric where all workloads store data in open Delta format
D.the Google Cloud storage engine unifying governance and fine-grained access over open table formats across BigQuery and external engines
Answer + AI explanation with Pro
26. Which term means: "the single, tenant-wide data lake underlying Microsoft Fabric where all workloads store data in open Delta format"?
Mid
A.EMR Serverless
B.BigQuery BI Engine
C.OneLake
D.S3 Tables
Answer + AI explanation with Pro
27. Which statement is correct?
Mid
A.OneLake — the AWS serverless query service that runs SQL directly over data in S3 using Presto and Trino engines
B.OneLake — the AWS serverless ETL service with a managed Spark runtime and a central data catalog for schema discovery
C.OneLake — the single, tenant-wide data lake underlying Microsoft Fabric where all workloads store data in open Delta format
D.OneLake — the in-memory acceleration layer for BigQuery that caches data to deliver sub-second responses for dashboards
Answer + AI explanation with Pro
28. What is BigLake?
Senior
A.the GCP storage engine that lets BigQuery and open engines query data lake tables, including Iceberg, with unified governance
B.the AWS serverless query service that runs SQL directly over data in S3 using Presto and Trino engines
C.a Microsoft Fabric reference that virtualizes external data into OneLake without copying it
D.the in-memory acceleration layer for BigQuery that caches data to deliver sub-second responses for dashboards
Answer + AI explanation with Pro
29. Which term means: "the GCP storage engine that lets BigQuery and open engines query data lake tables, including Iceberg, with unified governance"?
Senior
A.BigLake
B.Glue crawler
C.AWS Glue
D.Google Dataproc
Answer + AI explanation with Pro
30. Which statement is correct?
Senior
A.BigLake — the Google Cloud storage engine unifying governance and fine-grained access over open table formats across BigQuery and external engines
B.BigLake — the GCP storage engine that lets BigQuery and open engines query data lake tables, including Iceberg, with unified governance
C.BigLake — the AWS option that runs Spark and Hive jobs on automatically provisioned capacity without managing clusters
D.BigLake — the AWS serverless query service that runs SQL directly over data in S3 using Presto and Trino engines
Answer + AI explanation with Pro
Showing 30 of 87 Cloud Platforms questions — the full set, with answers, explanations and an AI tutor on every question, is inside.
Free to start
Start with a free readiness check
Sign up free for the 2-minute IT readiness check and a scored result. Answers, explanations and the AI tutor on every Cloud Platforms question come with Pro.