Ngân hàng đề — Google Cloud Professional Data Engineer

Tìm thấy 429 câu.

Câu 321
You store and analyze your relational data in BigQuery on Google Cloud with all data that resides in US regions. You also have a variety of object stores across Microsoft Azure and Amazon Web Services (AWS), also in US regions. You want to query all your data in BigQuery daily with as little movement of data as possible. What should you do?
  1. A Use BigQuery Data Transfer Service to load files from Azure and AWS into BigQuery.
  2. B Create a Dataflow pipeline to ingest files from Azure and AWS to BigQuery.
  3. C Load files from AWS and Azure to Cloud Storage with Cloud Shell gsutil rsync arguments.
  4. D Use the BigQuery Omni functionality and BigLake tables to query files in Azure and AWS.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào tình huống lưu trữ và phân tích dữ liệu quan hệ trong BigQuery trên Google Cloud (tất cả dữ liệu nằm ở các vùng US). Người dùng còn có các object store phân tán trên Microsoft Azure và Amazon Web Services (AWS) (cũng ở vùng US). Mục tiêu chính: Thực hiện truy vấn (query) tất cả dữ liệu trong BigQuery hàng ngày với ít di chuyển dữ liệu nhất có thể (minimize data movement).

📌 Yêu cầu cốt lõi:

  • Không muốn copy hoặc di chuyển dữ liệu lớn từ Azure/AWS sang Google Cloud để tránh chi phí, độ trễ và phức tạp.
  • Cần giải pháp query trực tiếp (federated query) từ BigQuery mà vẫn giữ dữ liệu tại chỗ (in-place).
  • Dữ liệu ở object stores (như S3 trên AWS hoặc Blob Storage trên Azure), nên cần hỗ trợ multi-cloud.

🛠️ Bối cảnh kiến thức cập nhật (đến 2026): Google Cloud cung cấp các tính năng BigQuery Omni (GA từ 2022, hỗ trợ query S3 và Azure Storage trực tiếp từ BigQuery) và BigLake (là unified storage layer, cho phép tạo external tables trên dữ liệu multi-cloud mà không cần di chuyển). Đây là giải pháp tối ưu cho cross-cloud analytics theo tài liệu AWS và Google Cloud mới nhất.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use the BigQuery Omni functionality and BigLake tables to query files in Azure and AWS.

Lý do chi tiết:

  • BigQuery Omni cho phép query dữ liệu trực tiếp trên AWS S3 và Azure Blob Storage từ BigQuery mà không cần di chuyển dữ liệu (zero-ETL). Hỗ trợ ANSI SQL đầy đủ, tích hợp governance từ Google Cloud.
  • BigLake tables là lớp lưu trữ thống nhất, tạo external tables trên dữ liệu ở Azure/AWS, kết hợp với dữ liệu native BigQuery để query liền mạch hàng ngày.
  • ✅ Ưu điểm nổi bật: Giảm chi phí (không copy data), độ trễ thấp (query in-place), hỗ trợ US regions, scale lớn. Phù hợp hoàn hảo với yêu cầu "as little movement of data as possible".
  • Theo docs Google Cloud 2026: BigQuery Omni hỗ trợ Parquet, ORC, Avro trên S3/Azure; BigLake thêm metadata caching cho performance cao.

📋 Phân tích tất cả các phương án

Dưới đây là phân tích từng lựa chọn một cách chi tiết, đánh dấu ✅ (đúng) hoặc ❌ (sai). Mỗi phương án giữ nguyên văn bản gốc tiếng Anh, nhưng giải thích hoàn toàn bằng tiếng Việt.

  • Use BigQuery Data Transfer Service to load files from Azure and AWS into BigQuery.
    ❌ Sai vì: BigQuery Data Transfer Service (DTS) chủ yếu dùng để load dữ liệu vào BigQuery từ các nguồn scheduled (như S3), nhưng yêu cầu di chuyển toàn bộ file (full copy) từ Azure/AWS sang BigQuery. Điều này vi phạm yêu cầu "ít di chuyển dữ liệu nhất" (gây chi phí lưu trữ/transfer cao, không query trực tiếp). DTS không hỗ trợ query federated native cho Azure (chỉ AWS S3 tốt hơn).

  • Create a Dataflow pipeline to ingest files from Azure and AWS to BigQuery.
    ❌ Sai vì: Apache Beam/Dataflow dùng để ETL/ingest dữ liệu (extract-transform-load), buộc phải copy file từ Azure/AWS vào BigQuery. Pipeline này phức tạp, tốn tài nguyên compute, và vẫn di chuyển data lớn hàng ngày – trái ngược yêu cầu minimize movement. Không phải giải pháp query in-place.

  • Load files from AWS and Azure to Cloud Storage with Cloud Shell gsutil rsync arguments.
    ❌ Sai vì: gsutil rsync chỉ sync/copy file từ AWS S3/Azure sang Google Cloud Storage (GCS), sau đó mới load vào BigQuery. Đây là bước trung gian di chuyển data đầy đủ, tốn bandwidth/chi phí (dù US regions), không query trực tiếp. gsutil hỗ trợ S3 nhưng hạn chế với Azure (cần connector riêng), không tối ưu cho query daily.

  • Use the BigQuery Omni functionality and BigLake tables to query files in Azure and AWS.
    ✅ Đúng vì: Như đã giải thích ở trên, đây là giải pháp zero-copy duy nhất, query trực tiếp từ BigQuery trên dữ liệu AWS S3/Azure mà không move data. Hỗ trợ hàng ngày với SQL chuẩn, tích hợp BigLake cho external tables linh hoạt.

📘 Tài liệu tham khảo (cập nhật mới nhất đến 2026)

🧩 Kết luận: Giải pháp đúng tận dụng sức mạnh multi-cloud của Google Cloud, giúp query thống nhất mà tiết kiệm tối đa! Nếu cần demo code SQL, hãy hỏi thêm. 🚀

Câu 322
You have a variety of files in Cloud Storage that your data science team wants to use in their models. Currently, users do not have a method to explore, cleanse, and validate the data in Cloud Storage. You are looking for a low code solution that can be used by your data science team to quickly cleanse and explore data within Cloud Storage. What should you do?
  1. A Provide the data science team access to Dataflow to create a pipeline to prepare and validate the raw data and load data into BigQuery for data exploration.
  2. B Create an external table in BigQuery and use SQL to transform the data as necessary. Provide the data science team access to the external tables to explore the raw data.
  3. C Load the data into BigQuery and use SQL to transform the data as necessary. Provide the data science team access to staging tables to explore the raw data.
  4. D Provide the data science team access to Dataprep to prepare, validate, and explore the data within Cloud Storage.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm

📖 Nội dung câu hỏi:
Câu hỏi mô tả tình huống bạn có nhiều loại file lưu trữ trong Cloud Storage (dịch vụ lưu trữ đám mây của Google Cloud), và đội ngũ data science muốn sử dụng chúng để xây dựng mô hình học máy. Hiện tại, họ chưa có công cụ để khám phá (explore), làm sạch (cleanse) và kiểm chứng (validate) dữ liệu trực tiếp trong Cloud Storage. Bạn cần tìm giải pháp low-code (ít code, dễ sử dụng cho non-developer) để đội ngũ data science nhanh chóng thực hiện các tác vụ này trực tiếp trong Cloud Storage, mà không cần di chuyển dữ liệu phức tạp.

🎯 Mục tiêu chính: Giải pháp phải hỗ trợ low-code, tập trung vào prepare/validate/explore dữ liệu ngay tại nguồn (Cloud Storage), phù hợp cho data scientist không chuyên code sâu. (Dựa trên kiến thức GCP cập nhật đến 2026, Dataprep vẫn là công cụ low-code hàng đầu cho visual data wrangling).

✅ Đáp án đúng:
Provide the data science team access to Dataprep to prepare, validate, and explore the data within Cloud Storage.

Lý do chọn đáp án này (chi tiết):
Dataprep (by Trifacta) là giải pháp low-code/no-code chuyên dụng của Google Cloud, cho phép người dùng kéo-thả để khám phá, làm sạch, transform và validate dữ liệu trực tiếp từ Cloud Storage mà không cần viết code. Nó hỗ trợ visual interface, recipe tự động hóa, và tích hợp liền mạch với GCS (Google Cloud Storage). Đội data science có thể nhanh chóng preview data, clean outliers, join files, validate schema – hoàn hảo cho yêu cầu "low code solution" và "quickly cleanse and explore within Cloud Storage". Không cần load dữ liệu vào nơi khác trước. (Cập nhật 2026: Dataprep đã tích hợp sâu hơn với Vertex AI, nhưng vẫn giữ vai trò low-code data prep chính).

📘 Tài liệu tham khảo:

🛠️ Giải thích tất cả các phương án

  • ❌ [SAI] Provide the data science team access to Dataflow to create a pipeline to prepare and validate the raw data and load data into BigQuery for data exploration.
    Phân tích sai: Dataflow là dịch vụ Apache Beam-based cho ETL pipeline code-heavy (yêu cầu viết code Python/Java), không phải low-code. Nó phù hợp cho production-scale processing, nhưng đội data science sẽ mất thời gian build pipeline và load vào BigQuery – không "quickly" và không trực tiếp explore trong Cloud Storage. Không đáp ứng yêu cầu low-code.

  • ❌ [SAI] Create an external table in BigQuery and use SQL to transform the data as necessary. Provide the data science team access to the external tables to explore the raw data.
    Phân tích sai: External table cho phép query dữ liệu GCS bằng SQL mà không load, nhưng chỉ hỗ trợ explore/query cơ bản, không có tính năng cleanse/validate visual (như handle missing values, profiling tự động). Yêu cầu SQL knowledge, không low-code thuần túy, và transform cần code SQL phức tạp – không phù hợp cho data science team cần quick prep.

  • ❌ [SAI] Load the data into BigQuery and use SQL to transform the data as necessary. Provide the data science team access to staging tables to explore the raw data.
    Phân tích sai: Việc load dữ liệu vào BigQuery trước tốn kém (storage + compute), không trực tiếp "within Cloud Storage". Transform/explore vẫn dùng SQL (code-based), không low-code. Staging tables chỉ là temporary storage, không giải quyết cleanse/validate nhanh chóng mà không di chuyển data.

  • ✅ [ĐÚNG] Provide the data science team access to Dataprep to prepare, validate, and explore the data within Cloud Storage.
    Phân tích đúng (tóm tắt): Như đã giải thích ở trên, đây là giải pháp lý tưởng với giao diện visual low-code, hỗ trợ đầy đủ explore/cleanse/validate ngay tại GCS, không cần code hay load data. Hoàn toàn khớp yêu cầu! 🚀

Câu 323
You are building an ELT solution in BigQuery by using Dataform. You need to perform uniqueness and null value checks on your final tables. What should you do to efficiently integrate these checks into your pipeline?
  1. A Build BigQuery user-defined functions (UDFs).
  2. B Create Dataplex data quality tasks.
  3. C Build Dataform assertions into your code.
  4. D Write a Spark-based stored procedure.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc xây dựng một giải pháp ELT (Extract, Load, Transform) trong BigQuery bằng cách sử dụng Dataform. Bạn cần thực hiện các kiểm tra uniqueness (tính duy nhất) và null value checks (kiểm tra giá trị null) trên các bảng cuối cùng (final tables). Mục tiêu là tích hợp các kiểm tra này một cách hiệu quả vào pipeline (quy trình xử lý dữ liệu).

📘 Bối cảnh chính: Dataform là công cụ quản lý và thực thi SQL transformations trong BigQuery, hỗ trợ ELT tự động hóa. Các kiểm tra chất lượng dữ liệu (data quality) phải được tích hợp mượt mà, không làm gián đoạn pipeline, và chạy tự động khi deploy/release. Theo tài liệu Google Cloud cập nhật mới nhất (đến năm 2026), Dataform cung cấp các tính năng native để xử lý data quality ngay trong code SQL.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Build Dataform assertions into your code.

🛠️ Lý do chi tiết: Dataform assertions là tính năng native và hiệu quả nhất để tích hợp kiểm tra uniqueness (ví dụ: COUNT DISTINCT) và null checks (ví dụ: COUNT IF column IS NULL) trực tiếp vào code Dataform. Chúng chạy tự động như một phần của pipeline ELT, fail job nếu vi phạm, và dễ maintain. Điều này đảm bảo data quality mà không cần tool ngoài, phù hợp với best practice ELT trong BigQuery.
📘 Nguồn tham khảo: Dataform Assertions Documentation (cập nhật 2025-2026, hỗ trợ SQL assertions với ref() cho tables).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm giải thích rõ ràng bằng tiếng Việt:

  • ❌ [SAI] Build BigQuery user-defined functions (UDFs).
    🧩 Giải thích: UDFs dùng để tạo hàm tùy chỉnh cho phép tính toán phức tạp trong queries, không phải để kiểm tra data quality tự động trong pipeline. Chúng không tích hợp trực tiếp vào Dataform workflow, không fail job tự động, và kém hiệu quả cho uniqueness/null checks (phải gọi thủ công). Không phù hợp với ELT pipeline.

  • ❌ [SAI] Create Dataplex data quality tasks.
    🧩 Giải thích: Dataplex là nền tảng data governance/governance cho multi-cloud, hỗ trợ data quality rules nhưng là tool riêng biệt, không tích hợp native vào Dataform. Bạn phải setup tasks ngoài pipeline ELT, dẫn đến phức tạp, duplicate effort, và không "efficiently integrate" như yêu cầu. Dataplex phù hợp hơn cho catalog lớn, không phải ELT nhanh.

  • ✅ [ĐÚNG] Build Dataform assertions into your code.
    🛠️ Giải thích: Như đã nêu ở phần đáp án đúng. Đây là cách tối ưu nhất, assertions là SQL files đặc biệt (kết thúc bằng _assert.sql), kiểm tra conditions trên tables/refs, fail release nếu lỗi, hỗ trợ uniqueness (GROUP BY COUNT >1) và nulls (SUM(IS_NULL)). Chạy parallel, scalable trong BigQuery.

  • ❌ [SAI] Write a Spark-based stored procedure.
    🧩 Giải thích: BigQuery không hỗ trợ Spark stored procedures native (Spark là Dataproc/EMR tool). Stored procedures trong BigQuery dùng SQL/JS, không Spark. Việc dùng Spark sẽ yêu cầu chuyển dữ liệu ra ngoài BigQuery, rất kém hiệu quả, tốn chi phí, và phá vỡ ELT pipeline Dataform thuần túy.

🏆 Kết luận và best practice

Sử dụng Dataform assertions là lựa chọn hiệu quả, native cho ELT trong BigQuery, đảm bảo data quality end-to-end. Theo Google Cloud Well-Architected Framework (2026), tích hợp checks vào code giúp CI/CD tự động fail-fast. Nếu scale lớn, kết hợp với Data Catalog hoặc Dataplex monitoring.
📘 Tài liệu bổ sung: BigQuery Dataform Guide, Data Quality Best Practices.

Câu 324
A web server sends click events to a Pub/Sub topic as messages. The web server includes an eventTimestamp attribute in the messages, which is the time when the click occurred. You have a Dataflow streaming job that reads from this Pub/Sub topic through a subscription, applies some transformations, and writes the result to another Pub/Sub topic for use by the advertising department. The advertising department needs to receive each message within 30 seconds of the corresponding click occurrence, but they report receiving the messages late. Your Dataflow job's system lag is about 5 seconds, and the data freshness is about 40 seconds. Inspecting a few messages show no more than 1 second lag between their eventTimestamp and publishTime. What is the problem and what should you do?
  1. A The advertising department is causing delays when consuming the messages. Work with the advertising department to fix this.
  2. B Messages in your Dataflow job are taking more than 30 seconds to process. Optimize your job or increase the number of workers to fix this.
  3. C Messages in your Dataflow job are processed in less than 30 seconds, but your job cannot keep up with the backlog in the Pub/Sub subscription. Optimize your job or increase the number of workers to fix this.
  4. D The web server is not pushing messages fast enough to Pub/Sub. Work with the web server team to fix this.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi

Câu hỏi mô tả một hệ thống streaming trên Google Cloud Platform (GCP) sử dụng Pub/Sub và Dataflow (Apache Beam trên GCP). Cụ thể:

  • Một web server gửi các sự kiện click (click events) dưới dạng messages đến một Pub/Sub topic, kèm theo thuộc tính eventTimestamp (thời điểm click xảy ra).
  • Có một Dataflow streaming job đọc dữ liệu từ Pub/Sub subscription của topic này, thực hiện transformations, rồi ghi kết quả ra một Pub/Sub topic khác dành cho bộ phận quảng cáo (advertising department).
  • Yêu cầu: Bộ phận quảng cáo cần nhận messages trong vòng 30 giây kể từ thời điểm click (eventTimestamp).
  • Vấn đề: Họ nhận messages muộn (late).
  • Metrics quan sát:
    • System lag của Dataflow job: khoảng 5 giây (thời gian từ khi message được publish đến khi Dataflow xử lý nó – rất thấp, cho thấy job xử lý nhanh).
    • Data freshness: khoảng 40 giây (chỉ số đo độ "tươi mới" của dữ liệu, thường liên quan đến watermark lag – khoảng cách giữa event time watermark và thời điểm hiện tại, ngụ ý có backlog dữ liệu tích tụ trong subscription).
    • Kiểm tra messages: Độ trễ giữa eventTimestamp và publishTime ≤ 1 giây (web server publish nhanh).

Vấn đề cốt lõi 📈: Mặc dù job xử lý từng message nhanh (system lag thấp), nhưng Dataflow không theo kịp lượng dữ liệu backlog trong Pub/Sub subscription, dẫn đến watermark lag cao (data freshness 40s). Kết quả là output messages có event time cũ hơn 40 giây so với real-time, khiến bộ phận quảng cáo nhận muộn hơn 30 giây từ click. Giải pháp cần tập trung vào việc scale job để clear backlog.

(Lưu ý: Đây là kiến thức GCP/Dataflow cập nhật đến 2026, với metrics streaming như system lag/data freshness từ Dataflow monitoring dashboard. Không liên quan AWS như mô tả ban đầu – có thể nhầm lẫn.)

✅ Đáp án đúng

Messages in your Dataflow job are processed in less than 30 seconds, but your job cannot keep up with the backlog in the Pub/Sub subscription. Optimize your job or increase the number of workers to fix this.

Lý do chọn đáp án này 🛠️:

  • System lag 5s + event-to-publish ≤1s → Tổng processing time <10s, dễ dàng dưới 30s cho từng message.
  • Data freshness 40s → Watermark lag cao, chứng tỏ backlog lớn trong subscription (Pub/Sub tích tụ unacknowledged messages). Dataflow đang "đuổi theo" dữ liệu cũ, không real-time.
  • Giải pháp: Tối ưu job (giảm CPU-intensive transforms, dùng autoscaling) hoặc tăng workers (scale Dataflow job lên max workers/CPU để clear backlog nhanh hơn). Theo best practices GCP 2026, sử dụng Flex Templates hoặc Streaming Engine để scale động.

📋 Giải thích tất cả các phương án

  • ❌ [SAI] The advertising department is causing delays when consuming the messages. Work with the advertising department to fix this.
    Lý do sai 🚫: Không có bằng chứng về delay từ phía consumer (advertising). Vấn đề nằm ở data freshness 40s từ Dataflow output, không phải consumption. Pub/Sub topic output decoupling tốt, consumer chậm chỉ ảnh hưởng ACK riêng.

  • ❌ [SAI] Messages in your Dataflow job are taking more than 30 seconds to process. Optimize your job or increase the number of workers to fix this.
    Lý do sai ⏱️: System lag chỉ 5s → Processing nhanh dưới 30s. Vấn đề không phải tốc độ xử lý từng message, mà là backlog tổng thể (data freshness 40s). Tăng workers vẫn cần, nhưng lý do không phải processing time cao.

  • ✅ [ĐÚNG] Messages in your Dataflow job are processed in less than 30 seconds, but your job cannot keep up with the backlog in the Pub/Sub subscription. Optimize your job or increase the number of workers to fix this.
    Lý do đúng 🔍: Như phân tích trên – system lag thấp nhưng data freshness cao → Backlog subscription. Scale job sẽ advance watermark, giảm freshness lag xuống <30s.

  • ❌ [SAI] The web server is not pushing messages fast enough to Pub/Sub. Work with the web server team to fix this.
    Lý do sai 🌐: Event-to-publish ≤1s → Web server publish rất nhanh. Backlog không phải do input chậm, mà do Dataflow không drain subscription kịp.

📘 Tài liệu tham khảo

Câu 325
Your organization stores customer data in an on-premises Apache Hadoop cluster in Apache Parquet format. Data is processed on a daily basis by Apache Spark jobs that run on the cluster. You are migrating the Spark jobs and Parquet data to Google Cloud. BigQuery will be used on future transformation pipelines so you need to ensure that your data is available in BigQuery. You want to use managed services, while minimizing ETL data processing changes and overhead costs. What should you do?
  1. A Migrate your data to Cloud Storage and migrate the metadata to Dataproc Metastore (DPMS). Refactor Spark pipelines to write and read data on Cloud Storage, and run them on Dataproc Serverless.
  2. B Migrate your data to Cloud Storage and register the bucket as a Dataplex asset. Refactor Spark pipelines to write and read data on Cloud Storage, and run them on Dataproc Serverless.
  3. C Migrate your data to BigQuery. Refactor Spark pipelines to write and read data on BigQuery, and run them on Dataproc Serverless.
  4. D Migrate your data to BigLake. Refactor Spark pipelines to write and read data on Cloud Storage, and run them on Dataproc on Compute Engine.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả tình huống tổ chức đang lưu trữ dữ liệu khách hàng trên cụm Apache Hadoop on-premises ở định dạng Apache Parquet. Dữ liệu được xử lý hàng ngày bởi các job Apache Spark chạy trên cụm này. Nhiệm vụ là migrate các Spark jobs và dữ liệu Parquet sang Google Cloud, đồng thời đảm bảo dữ liệu có thể sử dụng cho các pipeline transformation tương lai trên BigQuery. Yêu cầu chính:

  • Sử dụng managed services (dịch vụ được quản lý hoàn toàn).
  • Giảm thiểu thay đổi trong ETL data processing (tức là refactor Spark pipelines càng ít càng tốt).
  • Giảm thiểu overhead costs (chi phí vận hành thấp, tránh quản lý cluster thủ công).

Mục tiêu cốt lõi: Di chuyển dữ liệu và jobs một cách mượt mà, tận dụng Cloud Storage cho lưu trữ rẻ tiền, Spark managed trên Dataproc, và chuẩn bị cho BigQuery (qua external tables hoặc BigLake). Kiến thức dựa trên phiên bản Google Cloud mới nhất đến 2026, với Dataproc Serverless (GA từ 2023), Dataproc Metastore (DPMS) tích hợp BigLake/BigQuery external tables.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Migrate your data to Cloud Storage and migrate the metadata to Dataproc Metastore (DPMS). Refactor Spark pipelines to write and read data on Cloud Storage, and run them on Dataproc Serverless.

Lý do chọn đáp án này 🛠️:

  • Cloud Storage (GCS): Lưu trữ Parquet managed, rẻ tiền (low-cost object storage), hỗ trợ Spark read/write native qua S3A connector (tương tự Hadoop). Giảm thay đổi ETL vì Spark jobs chỉ cần refactor URL từ HDFS sang GCS.
  • Dataproc Metastore (DPMS): Managed Hive metastore service, migrate metadata từ Hadoop Hive (table schemas, partitions). Spark trên Dataproc tự động kết nối DPMS, giữ nguyên logic table operations (CREATE TABLE, partitions) mà không refactor lớn. Hỗ trợ BigQuery external tables/BigLake sau này (query GCS Parquet trực tiếp).
  • Dataproc Serverless: Chạy Spark jobs managed hoàn toàn, không cần provision cluster (zero overhead), scale auto, pay-per-use. Hoàn hảo cho daily jobs, minimize costs so với Dataproc on GCE.
  • Tổng thể: Đáp ứng managed services, minimize ETL changes (chỉ refactor path + metastore config), low costs, và sẵn sàng cho BigQuery (external tables trên GCS + DPMS metadata).

📋 Giải thích tất cả các phương án

  • Migrate your data to Cloud Storage and migrate the metadata to Dataproc Metastore (DPMS). Refactor Spark pipelines to write and read data on Cloud Storage, and run them on Dataproc Serverless.
    ✅ Đúng 🏆: Như giải thích trên, đây là giải pháp tối ưu nhất. DPMS cung cấp unified metastore cho Spark/Dataproc/BigQuery, Dataproc Serverless loại bỏ overhead cluster management. Hoàn hảo cho migration Hadoop/Spark sang GCP với BigQuery integration (qua BigLake external tables).

  • Migrate your data to Cloud Storage and register the bucket as a Dataplex asset. Refactor Spark pipelines to write and read data on Cloud Storage, and run them on Dataproc Serverless.
    ❌ Sai 🚫: Dataplex là data catalog/lakehouse management (metadata discovery, governance), không thay thế Hive metastore cho Spark jobs. Spark cần DPMS để quản lý table metadata/partitions; chỉ register bucket làm Dataplex asset chỉ hỗ trợ catalog/query, không giúp Spark read/write tables mượt mà. Gây thay đổi ETL lớn hơn và không đảm bảo BigQuery integration native.

  • Migrate your data to BigQuery. Refactor Spark pipelines to write and read data on BigQuery, and run them on Dataproc Serverless.
    ❌ Sai 🔄: BigQuery là data warehouse cho analytics/SQL, không phải storage cho Spark write/read native (Spark Connector tồn tại nhưng chậm, costly cho frequent writes). Migrate Parquet sang BigQuery yêu cầu ETL lớn (load/format conversion), tăng overhead costs (BigQuery storage/scan fees cao hơn GCS). Không minimize changes, và Spark jobs daily sẽ kém hiệu quả so với GCS + Spark native.

  • Migrate your data to BigLake. Refactor Spark pipelines to write and read data on Cloud Storage, and run them on Dataproc on Compute Engine.
    ❌ Sai ⚠️: BigLake là BigQuery table type (external tables trên GCS với DPMS metadata), không phải nơi "migrate data to" (data vẫn ở GCS). Vấn đề lớn: Dataproc on Compute Engine yêu cầu quản lý cluster thủ công (provision, scale), vi phạm "managed services" và "minimize overhead costs". Serverless mới là lựa chọn managed thực sự; phương án này tăng chi phí vận hành.

Giải pháp đúng giúp migration mượt mà, cost-effective, sẵn sàng scale cho BigQuery pipelines! 🚀

Câu 326
Your organization has two Google Cloud projects, project A and project B. In project A, you have a Pub/Sub topic that receives data from confidential sources. Only the resources in project A should be able to access the data in that topic. You want to ensure that project B and any future project cannot access data in the project A topic. What should you do?
  1. A Add firewall rules in project A so only traffic from the VPC in project A is permitted.
  2. B Configure VPC Service Controls in the organization with a perimeter around project A.
  3. C Use Identity and Access Management conditions to ensure that only users and service accounts in project A. can access resources in project A.
  4. D Configure VPC Service Controls in the organization with a perimeter around the VPC of project A.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi này thuộc lĩnh vực Google Cloud Platform (GCP), tập trung vào bảo mật dữ liệu và ngăn chặn rò rỉ dữ liệu (data exfiltration) giữa các project trong cùng một tổ chức (organization). Cụ thể:

  • Tổ chức có hai project GCP: Project A và Project B.
  • Trong Project A, có một Pub/Sub topic nhận dữ liệu từ nguồn bí mật (confidential sources).
  • Yêu cầu chính: Chỉ các tài nguyên (resources) trong Project A mới được phép truy cập dữ liệu từ topic này. Project B và bất kỳ project tương lai nào khác KHÔNG được phép truy cập.
  • Mục tiêu: Đảm bảo cách ly hoàn toàn dữ liệu giữa các project, ngay cả khi chúng thuộc cùng organization (nơi IAM có thể cho phép cross-project access mặc định).

Vấn đề cốt lõi là Pub/Sub là dịch vụ serverless, dữ liệu có thể bị truy cập qua API calls từ bất kỳ project nào nếu có quyền IAM. Do đó, cần cơ chế perimeter-based security để chặn data flows giữa các project. Đây là tình huống thực tế trong GCP khi xử lý dữ liệu nhạy cảm, đặc biệt với VPC Service Controls (VPC-SC) – công cụ chính thức để tạo "hàng rào" bảo vệ (perimeters).
📘 Tài liệu tham khảo: Google Cloud VPC Service Controls Documentation (cập nhật đến 2024-2026, VPC-SC vẫn là giải pháp chuẩn cho data perimeter).

✅ Đáp án đúng

Configure VPC Service Controls in the organization with a perimeter around project A.

Lý do lựa chọn:

  • VPC Service Controls (VPC-SC) là dịch vụ GCP chuyên dụng để tạo perimeter (hàng rào bảo mật) quanh các project hoặc folder, ngăn chặn dữ liệu di chuyển ra ngoài perimeter qua các API được hỗ trợ (bao gồm Pub/Sub).
  • Bằng cách tạo perimeter quanh Project A, tất cả dữ liệu trong Project A (như Pub/Sub topic) chỉ có thể được truy cập bởi tài nguyên bên trong perimeter (tức Project A). Project B hoặc project mới sẽ bị chặn hoàn toàn ngay cả nếu có IAM roles cross-project.
  • Đây là giải pháp chính xác, scalable và áp dụng cho tương lai (future projects), phù hợp với best practices GCP cho confidential data.
    🛠️ Cách triển khai: Enable VPC-SC tại organization level, tạo service perimeter bao gồm Project A và các service như Pub/Sub.
    📘 Nguồn: VPC-SC Quickstart và Pub/Sub Security.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên văn bản tiếng Anh của phương án, đánh dấu ✅/❌, và giải thích hoàn toàn bằng tiếng Việt dựa trên kiến thức GCP mới nhất (2026).

  • Add firewall rules in project A so only traffic from the VPC in project A is permitted.
    ❌ Sai vì: Firewall rules (VPC Firewall) chỉ kiểm soát network traffic (lớp 3/4) giữa các VM/IP, không áp dụng cho Pub/Sub – dịch vụ API-based, serverless không đi qua VPC traffic thông thường. Pub/Sub có thể bị access qua client libraries hoặc gcloud từ bất kỳ project nào có IAM quyền, mà không cần network flow. Firewall không chặn data exfiltration qua API, nên Project B vẫn access được nếu có service account quyền.

  • Configure VPC Service Controls in the organization with a perimeter around project A.
    ✅ Đúng vì: Như giải thích ở trên, VPC-SC tạo perimeter logic chặn data access và API calls ra khỏi Project A cho các service như Pub/Sub. Đây là duy nhất đáp ứng yêu cầu "chỉ resources in project A" và bảo vệ chống future projects. Đã được chứng nhận trong GCP Security best practices.

  • Use Identity and Access Management conditions to ensure that only users and service accounts in project A. can access resources in project A.
    ❌ Sai vì: IAM conditions (như resource.name.startsWith('projects/A')) chỉ giới hạn ai có quyền access (users/service accounts), nhưng KHÔNG chặn data flows giữa projects nếu service account từ Project B được cấp quyền (cross-project IAM phổ biến). IAM không phải là perimeter control, dễ bị bypass qua shared service accounts hoặc future grants. Không scalable cho "any future project".

  • Configure VPC Service Controls in the organization with a perimeter around the VPC of project A.
    ❌ Sai vì: VPC-SC không hỗ trợ perimeter quanh VPC riêng lẻ; perimeter phải quanh project/folder/organization. "Perimeter around VPC" không tồn tại trong GCP – VPC là network construct trong project, không phải boundary cho VPC-SC. Nếu thử, sẽ không bao quát Pub/Sub (không bind trực tiếp với VPC), dẫn đến lỗ hổng.
    🛠️ Lưu ý: VPC-SC dry-run mode giúp test trước khi enforce (cập nhật 2024+).

Câu 327
You stream order data by using a Dataflow pipeline, and write the aggregated result to Memorystore. You provisioned a Memorystore for Redis instance with Basic Tier, 4 GB capacity, which is used by 40 clients for read-only access. You are expecting the number of read-only clients to increase significantly to a few hundred and you need to be able to support the demand. You want to ensure that read and write access availability is not impacted, and any changes you make can be deployed quickly. What should you do?
  1. A Create a new Memorystore for Redis instance with Standard Tier. Set capacity to 4 GB and read replica to No read replicas (high availability only). Delete the old instance.
  2. B Create a new Memorystore for Redis instance with Standard Tier. Set capacity to 5 GB and create multiple read replicas. Delete the old instance.
  3. C Create a new Memorystore for Memcached instance. Set a minimum of three nodes, and memory per node to 4 GB. Modify the Dataflow pipeline and all clients to use the Memcached instance. Delete the old instance.
  4. D Create multiple new Memorystore for Redis instances with Basic Tier (4 GB capacity). Modify the Dataflow pipeline and new clients to use all instances.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh tình huống sử dụng Google Cloud Memorystore for Redis trong một pipeline Dataflow để stream dữ liệu order, sau đó aggregate và ghi kết quả vào Memorystore Redis (Basic Tier, dung lượng 4 GB). Instance hiện tại đang được 40 clients sử dụng chỉ để đọc (read-only). Dự kiến số clients đọc sẽ tăng mạnh lên vài trăm, cần đảm bảo:

  • Hỗ trợ nhu cầu tăng cao mà không ảnh hưởng đến tính sẵn sàng (availability) của read và write.
  • Các thay đổi phải deploy nhanh chóng (không downtime).

Vấn đề chính: Basic Tier chỉ là single instance (không HA, không hỗ trợ read replicas), nên không scale tốt cho read-heavy workload với hàng trăm clients. Cần giải pháp scale reads hiệu quả, tăng capacity nếu cần, và migrate mượt mà. 📈

✅ Đáp án đúng

Create a new Memorystore for Redis instance with Standard Tier. Set capacity to 5 GB and create multiple read replicas. Delete the old instance.

Lý do chọn đáp án này:

  • Standard Tier hỗ trợ high availability (HA) đa zone và read replicas (lên đến 7 replicas theo docs mới nhất 2024-2026), giúp scale reads song song cho hàng trăm clients mà không overload primary instance. ✅
  • Tăng capacity lên 5 GB để handle tải tăng (clients tăng gấp nhiều lần), tránh OOM hoặc latency cao. 🛠️
  • Tạo instance mới rồi delete cũ: Đảm bảo zero-downtime migration – clients và Dataflow có thể switch-over nhanh qua connection string mới, write vẫn ổn định.
  • Phù hợp hoàn hảo với yêu cầu deploy nhanh, không impact availability. 🚀

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên Memorystore Redis tiers (Basic: single-zone, no replicas; Standard: multi-zone HA + read replicas), cập nhật từ docs GCP 2026.

  • ❌ [SAI] Create a new Memorystore for Redis instance with Standard Tier. Set capacity to 4 GB and read replica to No read replicas (high availability only). Delete the old instance.
    Giải thích sai: Standard Tier đúng hướng cho HA, nhưng chọn "No read replicas (high availability only)" chỉ enable HA (primary + 1 standby replica cho failover), không scale reads vì tất cả reads vẫn đổ về primary. Với hàng trăm clients, primary sẽ overload, gây latency cao và impact availability. Capacity giữ 4 GB cũng không đủ cho tải tăng. Không đáp ứng "support demand" cho read-only clients. 😞

  • ✅ [ĐÚNG] Create a new Memorystore for Redis instance with Standard Tier. Set capacity to 5 GB and create multiple read replicas. Delete the old instance.
    Giải thích đúng: Như phần trên, multiple read replicas phân tải reads hiệu quả (clients connect đến replicas), Standard Tier đảm bảo HA, tăng 5 GB dự phòng tải, và create new/delete old cho migration nhanh không downtime. Hoàn toàn khớp yêu cầu! 🎯

  • ❌ [SAI] Create a new Memorystore for Memcached instance. Set a minimum of three nodes, and memory per node to 4 GB. Modify the Dataflow pipeline and all clients to use the Memcached instance. Delete the old instance.
    Giải thích sai: Memcached không tương thích với Redis protocol – clients và Dataflow hiện dùng Redis client libs (như redis-py), phải rewrite toàn bộ code để dùng Memcached (khác protocol, no persistence, no complex data types như Redis). Cần modify pipeline + tất cả clients → phức tạp, thời gian dài, dễ lỗi và downtime. Không "deploy nhanh" và impact write/read lớn. Memcached chỉ scale horizontal qua nodes nhưng kém Redis cho workload này. 🚫

  • ❌ [SAI] Create multiple new Memorystore for Redis instances with Basic Tier (4 GB capacity). Modify the Dataflow pipeline and new clients to use all instances.
    Giải thích sai: Basic Tier là single instance/không HA/replicas, tạo multiple instances yêu cầu manual sharding (chia keyspace qua code), phải modify Dataflow (write phân tán) + clients (read từ nhiều instances). Phức tạp, không tự động failover, dễ inconsistent data với write từ Dataflow. Không scale reads hiệu quả, dễ downtime khi delete old, và không "nhanh chóng". Không khuyến nghị cho production read-heavy. 🔄

📘 Tài liệu tham khảo (cập nhật mới nhất GCP 2026)

Giải pháp này đảm bảo performance cao, cost-effective cho workload tăng trưởng! 🌟

Câu 328 Chọn nhiều đáp án
You have a streaming pipeline that ingests data from Pub/Sub in production. You need to update this streaming pipeline with improved business logic. You need to ensure that the updated pipeline reprocesses the previous two days of delivered Pub/Sub messages. What should you do? (Choose two.)
  1. A Use the Pub/Sub subscription clear-retry-policy flag
  2. B Use Pub/Sub Snapshot capture two days before the deployment.
  3. C Create a new Pub/Sub subscription two days before the deployment.
  4. D Use the Pub/Sub subscription retain-acked-messages flag.
  5. E Use Pub/Sub Seek with a timestamp.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong môi trường production: Bạn có một pipeline streaming đang ingest dữ liệu từ Pub/Sub (dịch vụ messaging của Google Cloud). Bạn cần cập nhật pipeline với business logic cải tiến mới, đồng thời đảm bảo pipeline mới reprocess (xử lý lại) toàn bộ các Pub/Sub messages đã được delivered trong hai ngày trước thời điểm deploy (tức là dữ liệu lịch sử gần đây, có thể đã được ack bởi pipeline cũ).
Yêu cầu chọn hai hành động đúng để đạt được mục tiêu này, sử dụng các tính năng của Pub/Sub như snapshot, seek, retention, v.v.
🛠️ Mục tiêu chính: Không mất dữ liệu cũ đã ack, mà replay lại chính xác window 2 ngày để pipeline mới xử lý với logic mới, tránh downtime lớn và đảm bảo tính nhất quán.

✅ Đáp án đúng (Chọn hai) và lý do lựa chọn

Đáp án đúng là hai lựa chọn sau:

  • Use Pub/Sub Snapshot capture two days before the deployment.
  • Use the Pub/Sub subscription retain-acked-messages flag.

Lý do chi tiết 📘:

  • Để reprocess window 2 ngày trước deploy (từ thời điểm T-2day đến T), cần tạo Snapshot đúng 2 ngày trước deploy (tại T-2day). Sau deploy, seek subscription về Snapshot đó → Pipeline mới sẽ nhận lại tất cả messages published sau Snapshot (bao gồm 2 ngày trước và tương lai), cộng với unacked messages lúc tạo Snapshot.
  • Tuy nhiên, các messages đã acked (xử lý xong bởi pipeline cũ) sẽ không tự động replay trừ khi kích hoạt message retention qua flag retain-acked-messages (enable-message-retention). Flag này giữ lại acked messages trong retention duration (tối đa 7 ngày), làm chúng eligible để redeliver khi seek.
  • Kết hợp hai: Snapshot định vị chính xác thời điểm, retention đảm bảo dữ liệu đã ack không bị xóa → Hoàn hảo cho fixed time window reprocessing.
    (Kiến thức cập nhật GCP Pub/Sub đến 2026: Không thay đổi core feature này từ 2023-2026).

🧩 Giải thích chi tiết tất cả các phương án (Đúng/Sai)

Dưới đây là phân tích từng lựa chọn giữ nguyên văn bản gốc tiếng Anh, với lý do đúng/sai bằng tiếng Việt. Sử dụng kiến thức Pub/Sub mới nhất (docs replay mechanism).

  • ❌ Use the Pub/Sub subscription clear-retry-policy flag
    Sai: Flag clear-retry-policy chỉ dùng để xóa toàn bộ retry queue của subscription (messages đang retry do nack), không liên quan đến reprocess historical acked messages. Nó sẽ làm mất dữ liệu retry hiện tại chứ không giúp replay 2 ngày cũ.

  • ✅ Use Pub/Sub Snapshot capture two days before the deployment.
    Đúng: Tạo Snapshot tại đúng T-2day trước deploy. Sau update pipeline, seek về Snapshot → Replay tất cả messages published sau thời điểm Snapshot (chính là 2 ngày trước đến hiện tại + future), bắt đầu từ unacked lúc tạo. Đây là best practice cho fixed window reprocessing.

  • ❌ Create a new Pub/Sub subscription two days before the deployment.
    Sai: Không thể tạo subscription "backdated" 2 ngày trước một cách tự động. Subscription mới chỉ nhận messages từ lúc tạo trở đi (future messages), không capture historical data từ 2 ngày trước trừ khi manual backfill (phức tạp, không chính xác). Không giải quyết reprocess dữ liệu cũ.

  • ✅ Use the Pub/Sub subscription retain-acked-messages flag.
    Đúng: Flag này (tương đương --enable-message-retention) kích hoạt message retention trên subscription, giữ acked messages trong retention-duration (≤7 ngày). Khi seek (Snapshot hoặc timestamp), chúng sẽ được redeliver như unacked → Bắt buộc cần để replay acked messages từ 2 ngày cũ.

  • ❌ Use Pub/Sub Seek with a timestamp.
    Sai: Seek đến timestamp (ví dụ: 2 ngày trước) có thể replay messages published sau timestamp đó nếu retention đã enable. Tuy nhiên, chỉ dùng Seek alone không đủ vì không định vị chính xác "hai ngày trước deploy" mà cần kết hợp retention + cách khác (như Snapshot). Seek timestamp kém chính xác hơn cho fixed window so với Snapshot.

📘 Tài liệu tham khảo (Cập nhật mới nhất GCP đến 2026)

  • Chính thức: Replay messages using snapshots and seeking – Mô tả exact use case "reprocess past two days" với Snapshot + retention.
  • CLI/Command: gcloud pubsub subscriptions create/update --enable-message-retention --message-retention-duration=7d; gcloud pubsub snapshots create.
  • Console/API Docs: Pub/Sub Subscriptions API. (Không có thay đổi lớn post-2023; feature stable cho production workloads).

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần ví dụ code Terraform/CLI, hỏi thêm nhé.

Câu 329
You currently use a SQL-based tool to visualize your data stored in BigQuery. The data visualizations require the use of outer joins and analytic functions. Visualizations must be based on data that is no less than 4 hours old. Business users are complaining that the visualizations are too slow to generate. You want to improve the performance of the visualization queries while minimizing the maintenance overhead of the data preparation pipeline. What should you do?
  1. A Create materialized views with the allow_non_incremental_definition option set to true for the visualization queries. Specify the max_staleness parameter to 4 hours and the enable_refresh parameter to true. Reference the materialized views in the data visualization tool.
  2. B Create views for the visualization queries. Reference the views in the data visualization tool.
  3. C Create a Cloud Function instance to export the visualization query results as parquet files to a Cloud Storage bucket. Use Cloud Scheduler to trigger the Cloud Function every 4 hours. Reference the parquet files in the data visualization tool.
  4. D Create materialized views for the visualization queries. Use the incremental updates capability of BigQuery materialized views to handle changed data automatically. Reference the materialized views in the data visualization tool.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này xoay quanh việc tối ưu hóa hiệu suất truy vấn trực quan hóa dữ liệu (visualization) trong Google Cloud BigQuery.

  • Bối cảnh hiện tại: Bạn đang sử dụng một công cụ dựa trên SQL để tạo biểu đồ từ dữ liệu lưu trữ trong BigQuery. Các truy vấn visualization yêu cầu outer joins (kết nối ngoài) và analytic functions (hàm phân tích như WINDOW functions). Dữ liệu phải không mới hơn 4 giờ (tức chấp nhận dữ liệu cũ tối đa 4 giờ để đảm bảo tính nhất quán). Tuy nhiên, người dùng kinh doanh phàn nàn rằng việc tạo visualization quá chậm.

  • Mục tiêu: Cải thiện hiệu suất truy vấn visualization đồng thời giảm thiểu chi phí bảo trì pipeline chuẩn bị dữ liệu (không muốn quản lý thủ công nhiều).

Vấn đề cốt lõi là các truy vấn phức tạp (outer joins + analytics) chạy trực tiếp trên dữ liệu gốc gây chậm, cần giải pháp pre-compute (tính toán trước) mà vẫn tự động refresh theo độ trễ chấp nhận được (4 giờ). Đây là tình huống điển hình trong BigQuery để sử dụng Materialized Views (MVs) với cấu hình phù hợp, dựa trên tài liệu BigQuery cập nhật đến năm 2026 (phiên bản BigQuery ML và Storage v2).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create materialized views with the allow_non_incremental_definition option set to true for the visualization queries. Specify the max_staleness parameter to 4 hours and the enable_refresh parameter to true. Reference the materialized views in the data visualization tool.

Lý do chọn đáp án này 🛠️:

  • Materialized Views (MVs) lưu trữ kết quả truy vấn đã tính toán trước, giúp tăng tốc visualization lên đến 10-100x cho các truy vấn phức tạp như outer joins và analytic functions.
  • allow_non_incremental_definition=true: Bắt buộc vì outer joins/analytics không hỗ trợ incremental updates (chỉ full refresh), phù hợp với dữ liệu refresh mỗi 4 giờ.
  • max_staleness=4h: Đảm bảo dữ liệu MV cũ tối đa 4 giờ, khớp yêu cầu "no less than 4 hours old" (dữ liệu không cần real-time).
  • enable_refresh=true: Tự động refresh theo lịch BigQuery (dựa trên staleness), giảm maintenance overhead (không cần pipeline thủ công).
  • Kết quả: Visualization tool query trực tiếp MV nhanh chóng, tự động quản lý.

📋 Giải thích tất cả các phương án (đúng/sai)

  • Create materialized views with the allow_non_incremental_definition option set to true for the visualization queries. Specify the max_staleness parameter to 4 hours and the enable_refresh parameter to true. Reference the materialized views in the data visualization tool.
    ✅ Đúng hoàn toàn 🏆: Như giải thích trên, cấu hình này lý tưởng cho truy vấn phức tạp, tự động refresh với độ trễ 4 giờ, tối ưu performance mà không cần bảo trì thủ công. BigQuery tự quản lý refresh dựa trên slot quota.

  • Create views for the visualization queries. Reference the views in the data visualization tool.
    ❌ Sai 🚫: Views thông thường chỉ là logical abstraction (không lưu trữ dữ liệu), nên vẫn chạy full query trên dữ liệu gốc mỗi lần visualization → không cải thiện performance, vẫn chậm như hiện tại. Không giải quyết vấn đề chính.

  • Create a Cloud Function instance to export the visualization query results as parquet files to a Cloud Storage bucket. Use Cloud Scheduler to trigger the Cloud Function every 4 hours. Reference the parquet files in the data visualization tool.
    ❌ Sai ⚠️: Giải pháp này tạo ETL pipeline thủ công (Cloud Function + Scheduler), export ra Parquet/Storage → tăng maintenance overhead cao (quản lý function, scheduler, error handling, cost scaling). Không tận dụng native BigQuery, và visualization tool phải đọc file (chậm hơn MV query). Không phù hợp minimize maintenance.

  • Create materialized views for the visualization queries. Use the incremental updates capability of BigQuery materialized views to handle changed data automatically. Reference the materialized views in the data visualization tool.
    ❌ Sai 🔧: MVs không hỗ trợ incremental updates cho outer joins và analytic functions (chỉ full refresh được). Nếu dùng incremental mặc định, query sẽ lỗi hoặc không tạo MV được. Bỏ sót allow_non_incremental_definition=true và staleness config → không khớp yêu cầu 4 giờ và performance.

Giải pháp đúng giúp tiết kiệm chi phí query slots và scale dễ dàng! 🚀

Câu 330
You need to modernize your existing on-premises data strategy. Your organization currently uses:
•Apache Hadoop clusters for processing multiple large data sets, including on-premises Hadoop Distributed File System (HDFS) for data replication.
•Apache Airflow to orchestrate hundreds of ETL pipelines with thousands of job steps.

You need to set up a new architecture in Google Cloud that can handle your Hadoop workloads and requires minimal changes to your existing orchestration processes. What should you do?
  1. A Use Bigtable for your large workloads, with connections to Cloud Storage to handle any HDFS use cases. Orchestrate your pipelines with Cloud Composer.
  2. B Use Dataproc to migrate Hadoop clusters to Google Cloud, and Cloud Storage to handle any HDFS use cases. Orchestrate your pipelines with Cloud Composer.
  3. C Use Dataproc to migrate Hadoop clusters to Google Cloud, and Cloud Storage to handle any HDFS use cases. Convert your ETL pipelines to Dataflow.
  4. D Use Dataproc to migrate your Hadoop clusters to Google Cloud, and Cloud Storage to handle any HDFS use cases. Use Cloud Data Fusion to visually design and deploy your ETL pipelines.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm

📘 Nội dung câu hỏi:
Câu hỏi yêu cầu hiện đại hóa chiến lược dữ liệu on-premises hiện tại sang Google Cloud với thay đổi tối thiểu cho quy trình orchestration. Tổ chức đang sử dụng:

  • Apache Hadoop clusters để xử lý nhiều bộ dữ liệu lớn, bao gồm HDFS (Hadoop Distributed File System) cho việc sao chép dữ liệu.
  • Apache Airflow để điều phối hàng trăm pipeline ETL với hàng nghìn bước job.

Mục tiêu: Xây dựng kiến trúc mới trên Google Cloud có thể xử lý workload Hadoop và giữ nguyên quy trình orchestration hiện tại (minimal changes). Điều này nhấn mạnh vào việc di chuyển Hadoop mà không thay đổi lớn và giữ nguyên Airflow để tránh rewrite pipeline.
🛠️ Yêu cầu chính:

  • Xử lý Hadoop workloads (clusters, HDFS).
  • Orchestrate ETL pipelines giống Airflow (hàng trăm pipeline, thousands of job steps).
    Kiến thức cập nhật đến 2026: Google Cloud Dataproc (managed Hadoop/Spark, phiên bản mới nhất hỗ trợ Hive, Pig, Spark 3.x+), Cloud Storage (thay thế HDFS với connector native), Cloud Composer (managed Apache Airflow 2.x, hỗ trợ DAGs phức tạp lên đến 2026).

✅ Đáp án đúng:
Use Dataproc to migrate Hadoop clusters to Google Cloud, and Cloud Storage to handle any HDFS use cases. Orchestrate your pipelines with Cloud Composer.

Lý do chọn đáp án đúng (bằng tiếng Việt):
🟢 Đây là lựa chọn tối ưu nhất vì:

  • Dataproc là dịch vụ managed Hadoop/Spark của Google Cloud, cho phép di chuyển cluster Hadoop on-premises trực tiếp với minimal changes (hỗ trợ YARN, MapReduce, Spark, Hive – tương thích 100% với workload hiện tại).
  • Cloud Storage thay thế HDFS hoàn hảo qua Hadoop connector (gs:// URI), hỗ trợ replication và scalability cao hơn on-premises.
  • Cloud Composer là Apache Airflow managed service (dựa trên Airflow 2.x mới nhất 2026), không yêu cầu thay đổi code cho hàng trăm ETL pipelines và thousands of job steps – chỉ cần deploy DAGs hiện tại.
    Kết quả: Lift-and-shift dễ dàng, chi phí thấp, scalable.

📚 Tài liệu tham khảo:

❌ Phân tích tất cả các phương án (đúng/sai)

  • [SAI] Use Bigtable for your large workloads, with connections to Cloud Storage to handle any HDFS use cases. Orchestrate your pipelines with Cloud Composer.
    ❌ Sai vì: Bigtable là NoSQL database (wide-column store) dành cho workload real-time/low-latency (như analytics thời gian thực), không thay thế Hadoop clusters (không hỗ trợ MapReduce, Spark, YARN). Sử dụng Bigtable sẽ yêu cầu rewrite toàn bộ workload Hadoop, vi phạm "minimal changes". Cloud Storage và Composer đúng nhưng không bù đắp được.

  • [ĐÚNG] Use Dataproc to migrate Hadoop clusters to Google Cloud, and Cloud Storage to handle any HDFS use cases. Orchestrate your pipelines with Cloud Composer.
    ✅ Đúng vì: Như giải thích ở trên – full compatibility với Hadoop/Airflow, minimal migration effort. Hoàn hảo cho lift-and-shift.

  • [SAI] Use Dataproc to migrate Hadoop clusters to Google Cloud, and Cloud Storage to handle any HDFS use cases. Convert your ETL pipelines to Dataflow.
    ❌ Sai vì: Dataproc + Storage đúng cho Hadoop/HDFS, nhưng Dataflow (Apache Beam managed) yêu cầu convert toàn bộ pipelines từ Airflow sang Beam code, rất tốn công (hàng trăm pipelines, thousands steps). Vi phạm "minimal changes to orchestration processes".

  • [SAI] Use Dataproc to migrate your Hadoop clusters to Google Cloud, and Cloud Storage to handle any HDFS use cases. Use Cloud Data Fusion to visually design and deploy your ETL pipelines.
    ❌ Sai vì: Dataproc + Storage đúng, nhưng Cloud Data Fusion (nay là Data Fusion trong CDP, visual ETL tool) dành cho thiết kế pipeline mới từ đầu (drag-and-drop, không tương thích Airflow DAGs). Yêu cầu rebuild orchestration, không giữ nguyên Airflow – không "minimal changes".

🔍 Kết luận: Đáp án đúng đảm bảo seamless migration với chi phí thấp nhất, phù hợp best practices Google Cloud 2026 cho Hadoop/Airflow workloads! 🚀