Ngân hàng đề — Google Cloud Professional Data Engineer

Tìm thấy 429 câu.

Câu 201
You need to copy millions of sensitive patient records from a relational database to BigQuery. The total size of the database is 10 TB. You need to design a solution that is secure and time-efficient. What should you do?
  1. A Export the records from the database as an Avro file. Upload the file to GCS using gsutil, and then load the Avro file into BigQuery using the BigQuery web UI in the GCP Console.
  2. B Export the records from the database as an Avro file. Copy the file onto a Transfer Appliance and send it to Google, and then load the Avro file into BigQuery using the BigQuery web UI in the GCP Console.
  3. C Export the records from the database into a CSV file. Create a public URL for the CSV file, and then use Storage Transfer Service to move the file to Cloud Storage. Load the CSV file into BigQuery using the BigQuery web UI in the GCP Console.
  4. D Export the records from the database as an Avro file. Create a public URL for the Avro file, and then use Storage Transfer Service to move the file to Cloud Storage. Load the Avro file into BigQuery using the BigQuery web UI in the GCP Console.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi yêu cầu thiết kế giải pháp an toàn (secure) và hiệu quả về thời gian (time-efficient) để sao chép hàng triệu bản ghi bệnh nhân nhạy cảm từ một cơ sở dữ liệu quan hệ (relational database) sang BigQuery. Tổng kích thước dữ liệu là 10 TB – một lượng dữ liệu rất lớn.

🔑 Yêu cầu chính:

  • Bảo mật cao: Dữ liệu nhạy cảm (patient records) nên tránh các phương pháp công khai hoặc truyền qua internet không an toàn.
  • Hiệu quả thời gian: Với 10 TB, việc tải lên qua internet thông thường sẽ mất rất nhiều thời gian (có thể hàng tuần), nên cần phương pháp tối ưu cho dữ liệu lớn.
  • Quy trình: Xuất dữ liệu từ DB → Lưu trữ trung gian (GCS) → Tải vào BigQuery.

Giải pháp phải tận dụng các công cụ GCP như Transfer Appliance (thiết bị chuyển dữ liệu vật lý cho dữ liệu lớn), gsutil, Storage Transfer Service (STS), và giao diện BigQuery UI. Định dạng Avro được ưu tiên vì hỗ schema phức tạp, nén tốt, và hiệu suất cao khi load vào BigQuery (theo docs GCP 2024-2026).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Export the records from the database as an Avro file. Copy the file onto a Transfer Appliance and send it to Google, and then load the Avro file into BigQuery using the BigQuery web UI in the GCP Console.

Lý do 🛠️:

  • Hiệu quả thời gian: Transfer Appliance là thiết bị vật lý (rackmount server) cho phép copy 10 TB dữ liệu trực tiếp qua USB/SAS, gửi đến Google qua đường bưu điện. Thời gian chỉ vài ngày thay vì hàng tuần upload internet. Sau khi Google load dữ liệu vào GCS, bạn chỉ việc load vào BigQuery qua UI – siêu nhanh!
  • Bảo mật tuyệt đối: Không truyền qua internet, dữ liệu được mã hóa AES-256 tại chỗ, Google xử lý an toàn trước khi lưu vào GCS riêng tư.
  • Tối ưu cho BigQuery: Avro hỗ schema evolution, nén tốt, load song song nhanh chóng (hàng TB chỉ trong giờ).

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Export the records from the database as an Avro file. Upload the file to GCS using gsutil, and then load the Avro file into BigQuery using the BigQuery web UI in the GCP Console.
    Phân tích: Phương án này sử dụng gsutil để upload trực tiếp qua internet, nhưng với 10 TB, thời gian sẽ rất lâu (có thể 1-2 tuần tùy băng thông), không time-efficient. Bảo mật ổn (GCS private), nhưng không phải lựa chọn tối ưu cho dữ liệu lớn. Avro tốt, nhưng bottleneck ở upload.

  • ✅ [ĐÚNG] Export the records from the database as an Avro file. Copy the file onto a Transfer Appliance and send it to Google, and then load the Avro file into BigQuery using the BigQuery web UI in the GCP Console.
    Phân tích: Hoàn hảo cho 10 TB nhạy cảm! Transfer Appliance giải quyết cả secure (vật lý, mã hóa) và time-efficient (offline ship). Google tự động load vào GCS customer-managed encryption, rồi load Avro vào BigQuery chỉ trong phút. Khuyến nghị chính thức của GCP cho >1 PB dữ liệu (cập nhật 2026).

  • ❌ [SAI] Export the records from the database into a CSV file. Create a public URL for the CSV file, and then use Storage Transfer Service to move the file to Cloud Storage. Load the CSV file into BigQuery using the BigQuery web UI in the GCP Console.
    Phân tích: Public URL cực kỳ không an toàn cho dữ liệu nhạy cảm (patient records) – ai cũng truy cập được! CSV kém hiệu quả (không schema, khó parse large data, load chậm hơn Avro). STS chỉ phù hợp public sources như web, không phải sensitive data.

  • ❌ [SAI] Export the records from the database as an Avro file. Create a public URL for the Avro file, and then use Storage Transfer Service to move the file to Cloud Storage. Load the Avro file into BigQuery using the BigQuery web UI in the GCP Console.
    Phân tích: Avro tốt, STS ổn cho transfer, nhưng public URL lại vi phạm secure (dữ liệu nhạy cảm lộ ra internet). Không time-efficient bằng Transfer Appliance cho 10 TB, vì STS vẫn phụ thuộc băng thông upload ban đầu để tạo public URL.

💡 Kết luận: Chọn Transfer Appliance là best practice GCP cho dữ liệu lớn + nhạy cảm. Nếu dữ liệu nhỏ hơn (<100 GB), gsutil có thể dùng, nhưng ở đây 10 TB thì không! 🚀

Câu 202
You need to create a near real-time inventory dashboard that reads the main inventory tables in your BigQuery data warehouse. Historical inventory data is stored as inventory balances by item and location. You have several thousand updates to inventory every hour. You want to maximize performance of the dashboard and ensure that the data is accurate. What should you do?
  1. A Leverage BigQuery UPDATE statements to update the inventory balances as they are changing.
  2. B Partition the inventory balance table by item to reduce the amount of data scanned with each inventory update.
  3. C Use the BigQuery streaming the stream changes into a daily inventory movement table. Calculate balances in a view that joins it to the historical inventory balance table. Update the inventory balance table nightly.
  4. D Use the BigQuery bulk loader to batch load inventory changes into a daily inventory movement table. Calculate balances in a view that joins it to the historical inventory balance table. Update the inventory balance table nightly.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi yêu cầu xây dựng một dashboard inventory gần real-time (near real-time) đọc dữ liệu từ các bảng inventory chính trong BigQuery data warehouse.

  • Dữ liệu lịch sử: Lưu trữ dưới dạng inventory balances (số dư tồn kho) theo item (mặt hàng) và location (vị trí).
  • Tần suất cập nhật: Hàng nghìn updates mỗi giờ (several thousand updates every hour).
  • Mục tiêu chính:
    • Tối ưu hiệu suất dashboard (maximize performance): Dashboard cần load nhanh, tránh scan dữ liệu lớn.
    • Đảm bảo dữ liệu chính xác (ensure data accurate): Phản ánh số dư gần real-time mà không bị lỗi hoặc delay lớn.

Vấn đề cốt lõi 🛠️: BigQuery là hệ thống analytical warehouse tối ưu cho batch processing và query lớn, không phải OLTP (transactional updates real-time). Cần giải pháp xử lý high-velocity updates mà không làm chậm dashboard, sử dụng các tính năng như streaming, views, partitioning một cách thông minh.

(Kiến thức cập nhật đến 2026: BigQuery hỗ trợ streaming inserts lên đến 1 triệu rows/giây với latency ~few seconds, và materialized views cho performance cao hơn từ phiên bản 2023+ – theo Google Cloud docs).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use the BigQuery streaming the stream changes into a daily inventory movement table. Calculate balances in a view that joins it to the historical inventory balance table. Update the inventory balance table nightly.

Lý do chọn 📘:

  • Streaming inserts của BigQuery cho phép insert dữ liệu gần real-time (latency <10 giây), lý tưởng cho thousands updates/giờ mà không cần full table scan.
  • Daily inventory movement table: Lưu changes (biến động) hàng ngày, join với historical balance table qua view để tính current balance = historical + sum(movements). View này query nhanh, tự động update khi có stream mới.
  • Nightly update balance table: Batch job hàng đêm để consolidate balances, giữ table chính gọn nhẹ, tránh bloat từ streaming.
  • Lợi ích: Dashboard query view → performance cao (chỉ scan movements mới), accurate (gần real-time), scale tốt cho high-volume.

Nguồn: BigQuery Streaming Inserts & BigQuery Views (Google Cloud Docs, cập nhật 2025).

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Leverage BigQuery UPDATE statements to update the inventory balances as they are changing.
    Phân tích sai: UPDATE trong BigQuery yêu cầu full table rewrite (scan toàn bộ table), rất tốn kém (chi phí cao theo bytes processed) và chậm (latency cao cho thousands updates/giờ). Không phù hợp near real-time, dễ timeout dashboard. BigQuery không thiết kế cho frequent DML updates (tốt hơn dùng streaming inserts).

  • ❌ [SAI] Partition the inventory balance table by item to reduce the amount of data scanned with each inventory update.
    Phân tích sai: Partition by item (thousands items) tạo nhiều partitions nhỏ, nhưng updates scattered (phân tán) vẫn scan nhiều partitions → không giảm scan hiệu quả. BigQuery partitioning tối ưu cho date/time, không phải high-cardinality như item. Vẫn chậm và costly cho updates liên tục.

  • ✅ [ĐÚNG] Use the BigQuery streaming the stream changes into a daily inventory movement table. Calculate balances in a view that joins it to the historical inventory balance table. Update the inventory balance table nightly.
    Phân tích đúng: Như giải thích ở trên – streaming cho near real-time, view join tính balance động (query nhanh), nightly batch consolidate. Hoàn hảo cân bằng performance + accuracy, scale đến petabytes.

  • ❌ [SAI] Use the BigQuery bulk loader to batch load inventory changes into a daily inventory movement table. Calculate balances in a view that joins it to the historical inventory balance table. Update the inventory balance table nightly.
    Phân tích sai: Bulk loader (như bq load) là batch process (delay phút/giờ), không near real-time (dashboard sẽ stale data). Phù hợp low-velocity, nhưng fail yêu cầu "several thousand updates/giờ" cần immediate visibility.

Kết luận tổng quát 🎯: Giải pháp đúng tận dụng immutable append-only nature của BigQuery (streaming + views), tránh mutable updates tốn kém. Đây là best practice cho inventory dashboards!

Câu 203
You have a data stored in BigQuery. The data in the BigQuery dataset must be highly available. You need to define a storage, backup, and recovery strategy of this data that minimizes cost. How should you configure the BigQuery table that have a recovery point objective (RPO) of 30 days?
  1. A Set the BigQuery dataset to be regional. In the event of an emergency, use a point-in-time snapshot to recover the data.
  2. B Set the BigQuery dataset to be regional. Create a scheduled query to make copies of the data to tables suffixed with the time of the backup. In the event of an emergency, use the backup copy of the table.
  3. C Set the BigQuery dataset to be multi-regional. In the event of an emergency, use a point-in-time snapshot to recover the data.
  4. D Set the BigQuery dataset to be multi-regional. Create a scheduled query to make copies of the data to tables suffixed with the time of the backup. In the event of an emergency, use the backup copy of the table.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc thiết kế chiến lược lưu trữ, sao lưu và khôi phục dữ liệu trong BigQuery (Google Cloud) sao cho dữ liệu trong dataset có tính sẵn sàng cao (highly available), đồng thời tối ưu hóa chi phí thấp nhất. Yêu cầu cụ thể là cấu hình bảng BigQuery với Recovery Point Objective (RPO) = 30 ngày, nghĩa là có thể khôi phục dữ liệu tối đa lùi lại 30 ngày mà không mất dữ liệu quá mức chấp nhận được.

🔑 Các yếu tố chính cần xem xét:

  • High availability (HA): Dữ liệu phải khả dụng ngay cả khi có sự cố vùng (region outage). BigQuery hỗ trợ location regional (giới hạn trong 1 region, HA nội bộ qua 3 zone) hoặc multi-regional (phân tán qua nhiều region, HA cao hơn với tự động failover).
  • Backup & Recovery: Sử dụng point-in-time snapshot (PIT snapshot) dựa trên Time Travel (lưu lịch sử thay đổi metadata, không tốn storage thêm, mặc định 7 ngày nhưng có thể kết hợp snapshot table cho retention dài hơn theo docs mới nhất). Hoặc scheduled query để copy dữ liệu (tốn storage gấp đôi).
  • Minimize cost: Tránh copy dữ liệu (scheduled query) vì tăng chi phí storage (~$0.02/GB/tháng active). PIT snapshot rẻ hơn vì chỉ dùng metadata. Storage price uniform giữa regional/multi-regional, nhưng multi-regional đảm bảo HA tốt hơn mà không tăng cost đáng kể.
  • RPO 30 ngày: Time Travel chuẩn 7 ngày, nhưng với table snapshots (tính năng mới từ 2023-2024, hỗ trợ retention tùy chỉnh qua copy-on-write, lên đến years nếu config), kết hợp multi-regional để HA toàn diện.

📘 Dẫn nguồn:

✅ Đáp án đúng

Set the BigQuery dataset to be multi-regional. In the event of an emergency, use a point-in-time snapshot to recover the data.

Lý do lựa chọn 🛠️:

  • Đảm bảo HA cao: Multi-regional tự động replicate dữ liệu qua nhiều region (ví dụ: US multi-region), chịu được outage toàn region, SLA 99.99%.
  • Tối ưu chi phí: PIT snapshot (Time Travel + table snapshot) chỉ dùng metadata/history, không duplicate storage, rẻ hơn scheduled query (tránh chi phí copy ~2x).
  • Đáp ứng RPO 30 ngày: Kết hợp Time Travel (7 ngày) + snapshot tables (retention tùy chỉnh đến 30+ ngày), phù hợp yêu cầu mà không tốn kém.
  • So với regional: Không đủ HA cross-region. So với copy: Đắt hơn không cần thiết.

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên HA, cost, và RPO.

  • [SAI] Set the BigQuery dataset to be regional. In the event of an emergency, use a point-in-time snapshot to recover the data.
    ❌ Sai vì: Regional chỉ HA nội bộ 1 region (3 zone), không chịu được outage region (ví dụ: us-central1 down → dữ liệu unavailable). PIT snapshot tốt cho recovery/RPO nhưng không giải quyết HA. Không minimize cost so với multi-regional vì storage price tương đương, nhưng fail HA requirement.

  • [SAI] Set the BigQuery dataset to be regional. Create a scheduled query to make copies of the data to tables suffixed with the time of the backup. In the event of an emergency, use the backup copy of the table.
    ❌ Sai vì: Regional thiếu HA cross-region. Scheduled query copy dữ liệu (tables với suffix timestamp) tốn storage gấp đôi (~$0.02/GB/tháng x2), không minimize cost. Có thể đạt RPO 30 ngày qua lịch copy, nhưng đắt đỏ và phức tạp quản lý (dùng Cloud Scheduler + query).

  • [ĐÚNG] Set the BigQuery dataset to be multi-regional. In the event of an emergency, use a point-in-time snapshot to recover the data.
    ✅ Đúng hoàn toàn: Multi-regional đảm bảo HA cao nhất (replicate multi-zone/region). PIT snapshot rẻ, hiệu quả cho RPO (Time Travel + snapshots retention dài), không extra storage. Tối ưu cost + HA + recovery.

  • [SAI] Set the BigQuery dataset to be multi-regional. Create a scheduled query to make copies of the data to tables suffixed with the time of the backup. In the event of an emergency, use the backup copy of the table.
    ❌ Sai vì: Multi-regional tốt cho HA, nhưng scheduled query tăng cost không cần thiết (duplicate storage). PIT snapshot đã đủ cho recovery/RPO, không cần copy thủ công phức tạp và đắt.

Kết luận 🎯: Lựa chọn đúng cân bằng hoàn hảo giữa HA (multi-regional), recovery (PIT), và cost thấp. Khuyến nghị config dataset location khi tạo: bq mk --location=US. Test qua bq show --format=prettyjson dataset.

Câu 204
You used Dataprep to create a recipe on a sample of data in a BigQuery table. You want to reuse this recipe on a daily upload of data with the same schema, after the load job with variable execution time completes. What should you do?
  1. A Create a cron schedule in Dataprep.
  2. B Create an App Engine cron job to schedule the execution of the Dataprep job.
  3. C Export the recipe as a Dataprep template, and create a job in Cloud Scheduler.
  4. D Export the Dataprep job as a Dataflow template, and incorporate it into a Composer job.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi mô tả tình huống: Bạn đã sử dụng Dataprep (công cụ của Google Cloud để chuẩn bị dữ liệu) để tạo một recipe (công thức xử lý dữ liệu) trên một mẫu dữ liệu từ bảng BigQuery. Bây giờ, bạn muốn tái sử dụng recipe này hàng ngày trên dữ liệu mới được tải lên (daily upload) với cùng schema, sau khi công việc tải dữ liệu (load job) hoàn thành. Vấn đề chính là thời gian thực thi của load job là biến đổi (variable execution time), nên cần một cơ chế lập lịch linh hoạt, phụ thuộc vào việc load job hoàn tất trước khi chạy recipe.
📌 Mục tiêu: Tìm cách tự động hóa quy trình này một cách đáng tin cậy, xử lý dependency động mà không bị miss dữ liệu hoặc chạy sớm.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Export the Dataprep job as a Dataflow template, and incorporate it into a Composer job.

Lý do:

  • Dataprep hỗ trợ xuất recipe thành Dataflow template (Apache Beam template), cho phép chạy scalable trên Cloud Dataflow.
  • Cloud Composer (dựa trên Apache Airflow) là công cụ orchestration mạnh mẽ, cho phép tạo DAG (Directed Acyclic Graph) với sensors (như BigQuerySensor) để chờ load job hoàn thành (kiểm tra trạng thái bảng BigQuery hoặc job ID) trước khi trigger Dataflow template.
  • Điều này xử lý hoàn hảo thời gian biến đổi của load job, đảm bảo recipe chạy đúng sau khi dữ liệu sẵn sàng, và dễ scale hàng ngày.
    🛠️ Ưu điểm: Linh hoạt, idempotent, monitoring tốt qua Airflow UI. Phù hợp best practice GCP cho ETL pipelines phức tạp (cập nhật đến 2026, Cloud Composer v3 vẫn hỗ trợ tích hợp Dataflow templates mượt mà).

📋 Giải thích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên tài liệu GCP mới nhất (2026):

  • [SAI] Create a cron schedule in Dataprep.
    ❌ Sai vì: Dataprep không hỗ trợ cron schedule trực tiếp cho việc chạy job hàng ngày. Nó chỉ cho phép chạy thủ công hoặc xuất recipe, không xử lý dependency với load job biến đổi (có thể chạy sớm dẫn đến lỗi dữ liệu chưa sẵn sàng). Dataprep tập trung vào visual prep, không phải scheduler.
    📘 Tham khảo: Dataprep documentation - Không có tính năng cron native.

  • [SAI] Create an App Engine cron job to schedule the execution of the Dataprep job.
    ❌ Sai vì: App Engine cron chỉ hỗ trợ lịch fixed (cron expression), không thể chờ dynamic dependency như load job hoàn thành (variable time). Phải hardcode thời gian, dễ fail nếu load job chậm. App Engine không tích hợp sâu với Dataprep job execution.
    📘 Tham khảo: App Engine cron docs - Giới hạn ở periodic fixed schedules.

  • [SAI] Export the recipe as a Dataprep template, and create a job in Cloud Scheduler.
    ❌ Sai vì: Dataprep template chỉ dùng cho reuse thủ công hoặc API call, nhưng Cloud Scheduler là fixed-time trigger (HTTP/PubSub), không hỗ trợ wait cho load job. Không có sensor để kiểm tra trạng thái BigQuery, dẫn đến race condition.
    📘 Tham khảo: Cloud Scheduler docs & Dataprep templates - Không hỗ trợ conditional execution.

  • [ĐÚNG] Export the Dataprep job as a Dataflow template, and incorporate it into a Composer job.
    ✅ Đúng vì: Như giải thích ở trên, kết hợp Dataflow template từ Dataprep với Composer DAG cho phép sensor chờ load job (ví dụ: BigQueryTableSensor), đảm bảo thứ tự chính xác. Hỗ trợ retry, alerting tự động. Đây là pattern chuẩn cho production ETL trên GCP.
    📘 Tham khảo:

🧩 Kết luận: Phương án đúng tận dụng orchestration mạnh mẽ của Composer để xử lý dependency động, phù hợp với best practices GCP cho data pipelines scaleable! 🚀

Câu 205
You want to automate execution of a multi-step data pipeline running on Google Cloud. The pipeline includes Dataproc and Dataflow jobs that have multiple dependencies on each other. You want to use managed services where possible, and the pipeline will run every day. Which tool should you use?
  1. A cron
  2. B Cloud Composer
  3. C Cloud Scheduler
  4. D Workflow Templates on Dataproc
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc tự động hóa thực thi một pipeline dữ liệu đa bước (multi-step data pipeline) chạy trên Google Cloud. Các yêu cầu chính bao gồm:

  • Pipeline sử dụng Dataproc (dịch vụ quản lý cluster Hadoop/Spark) và Dataflow (dịch vụ xử lý dữ liệu stream/batch theo mô hình Apache Beam).
  • Có nhiều dependencies lẫn nhau giữa các job (ví dụ: job Dataflow phải chờ Dataproc hoàn thành trước khi chạy tiếp).
  • Ưu tiên managed services (dịch vụ được quản lý bởi Google Cloud để giảm vận hành thủ công).
  • Pipeline chạy hàng ngày (daily schedule). Mục tiêu là chọn công cụ orchestration (điều phối) phù hợp nhất để quản lý luồng công việc phức tạp này một cách tự động và đáng tin cậy. ✅ Đây là tình huống điển hình trong Google Cloud Data Engineering, nơi cần tool hỗ trợ DAG (Directed Acyclic Graph) để xử lý dependencies.

✅ Đáp án đúng: Cloud Composer

Lý do lựa chọn:
Cloud Composer là dịch vụ managed Apache Airflow trên Google Cloud, được thiết kế chuyên biệt để orchestrate các pipeline dữ liệu phức tạp với nhiều steps có dependencies. 🛠️ Nó hỗ trợ native integration với Dataproc và Dataflow (qua operators như DataprocSubmitJobOperator và DataflowTemplatedJobStartOperator), cho phép định nghĩa DAG để tự động chạy theo thứ tự, retry failures, và monitor. Pipeline chạy daily chỉ cần schedule qua DAG trigger. Là managed service, không cần quản lý server, phù hợp hoàn hảo với yêu cầu "use managed services where possible". Theo tài liệu Google Cloud cập nhật 2024-2026, Cloud Composer 3 (dựa Airflow 2.7+) tối ưu cho các workload lớn với autoscaling environments.

📘 Tài liệu tham khảo:

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên khả năng đáp ứng multi-step dependencies, managed service, và hỗ trợ Dataproc + Dataflow.

  • cron ❌ Sai:
    Cron là công cụ schedule cơ bản (dựa Linux/Unix), chỉ trigger job theo thời gian cố định (như hàng ngày) mà không hỗ trợ dependencies phức tạp giữa các steps. Không orchestrate được pipeline multi-step với Dataproc/Dataflow (phải viết script thủ công, dễ lỗi). Không phải managed service Google Cloud thuần túy, thiếu monitoring/retry. Phù hợp job đơn lẻ, không cho workload này.

  • Cloud Composer ✅ Đúng:
    Như đã giải thích ở trên. 🏆 Hoàn hảo cho DAG-based orchestration, managed fully bởi Google, tích hợp sâu Dataproc/Dataflow. Cập nhật 2026: Hỗ trợ Airflow 2.9+ với GKE-based environments cho scalability cao.

  • Cloud Scheduler ❌ Sai:
    Cloud Scheduler là dịch vụ schedule HTTP/Cloud Pub/Sub/App Engine jobs đơn giản, tương tự cron trên cloud. Chỉ trigger một job cụ thể hàng ngày, không orchestrate multi-step hay handle dependencies giữa Dataproc/Dataflow. Thiếu DAG, monitoring chi tiết cho pipeline phức tạp. Không phù hợp dù là managed service.

  • Workflow Templates on Dataproc ❌ Sai:
    Workflow Templates là tính năng của Dataproc để định nghĩa workflows chỉ trong ecosystem Dataproc (như Spark/Hadoop jobs với dependencies). Tuy managed và hỗ trợ multi-step, nhưng không native hỗ trợ Dataflow jobs (pipeline có cả Dataflow), dẫn đến phải dùng workaround phức tạp. Không linh hoạt cho mixed workloads như yêu cầu. Cập nhật 2024+: Vẫn giới hạn trong Dataproc, không thay thế Composer cho broad orchestration.

Tóm tắt khuyến nghị 🚀: Sử dụng Cloud Composer để xây dựng DAG sample như sau (pseudocode):

Task1: DataprocSubmitJobOperator >> Task2: DataflowTemplatedJobStartOperator >> Task3: Dataproc...

Điều này đảm bảo pipeline chạy daily, scalable, và zero-ops! Nếu cần code ví dụ chi tiết, hãy cung cấp thêm yêu cầu.

Câu 206
You are managing a Cloud Dataproc cluster. You need to make a job run faster while minimizing costs, without losing work in progress on your clusters. What should you do?
  1. A Increase the cluster size with more non-preemptible workers.
  2. B Increase the cluster size with preemptible worker nodes, and configure them to forcefully decommission.
  3. C Increase the cluster size with preemptible worker nodes, and use Cloud Observability to trigger a script to preserve work.
  4. D Increase the cluster size with preemptible worker nodes, and configure them to use graceful decommissioning.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này thuộc chủ đề quản lý Cloud Dataproc cluster trên Google Cloud Platform (GCP). Bạn đang quản lý một cụm Dataproc và cần tăng tốc độ chạy job (make a job run faster) đồng thời giảm thiểu chi phí (minimizing costs), nhưng không được làm mất công việc đang tiến hành (without losing work in progress) trên các cụm máy.

🛠️ Yêu cầu chính:

  • Tăng tốc độ: Cần scale up cluster (tăng kích thước) để thêm tài nguyên tính toán.
  • Giảm chi phí: Ưu tiên sử dụng preemptible worker nodes (rẻ hơn ~80% so với non-preemptible, theo giá GCP mới nhất 2024-2026).
  • Không mất dữ liệu: Phải xử lý tình huống preemptible nodes bị thu hồi (preempted) bởi GCP mà không làm gián đoạn job đang chạy.

📘 Bối cảnh kỹ thuật: Cloud Dataproc hỗ trợ preemptible VMs để tiết kiệm chi phí, nhưng chúng có thể bị GCP preempt sau tối đa 24 giờ. Để tránh mất work, cần cơ chế graceful decommissioning (phân tích chi tiết dưới đây). Kiến thức dựa trên tài liệu GCP cập nhật đến 2026 (Dataproc 2.x series).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Increase the cluster size with preemptible worker nodes, and configure them to use graceful decommissioning.

Lý do 🏆:

  • Tăng tốc độ: Scale up với thêm preemptible workers tăng số lượng node xử lý song song → job chạy nhanh hơn.
  • Giảm chi phí: Preemptible nodes rẻ hơn đáng kể (80% tiết kiệm), phù hợp minimizing costs.
  • Không mất work: Graceful decommissioning (tính năng chính thức của Dataproc từ 2020, cập nhật 2024) cho phép node tự động drain tasks đang chạy sang node khác trong cluster trước khi bị preempt (quá trình graceful ~10-30 phút). YARN/Hadoop tự động reschedule tasks mà không mất tiến độ.
  • Đây là best practice chính thức của GCP để cân bằng performance/cost/reliability.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Increase the cluster size with more non-preemptible workers.
    Phân tích: Scale up với non-preemptible workers (ổn định, không bị preempt) sẽ tăng tốc độ job nhờ thêm tài nguyên. Tuy nhiên, chi phí cao (đắt gấp 5 lần preemptible), không minimizing costs. Không vi phạm "không mất work" nhưng thất bại ở tiêu chí chi phí.

  • ❌ [SAI] Increase the cluster size with preemptible worker nodes, and configure them to forcefully decommission.
    Phân tích: Sử dụng preemptible để rẻ và nhanh, nhưng forceful decommission (buộc tắt ngay lập tức) sẽ mất work in progress vì tasks không được drain/graceful → job fail hoặc restart từ đầu. Vi phạm yêu cầu "without losing work".

  • ❌ [SAI] Increase the cluster size with preemptible worker nodes, and use Cloud Observability to trigger a script to preserve work.
    Phân tích: Preemptible tốt cho tốc độ/chi phí, nhưng Cloud Observability (nay là Cloud Monitoring/Logging) chỉ monitor, không tự động preserve work. Script tùy chỉnh phức tạp, không reliable (preempt có thể xảy ra đột ngột <1 phút), dễ mất work và không phải best practice. GCP không khuyến nghị.

  • ✅ [ĐÚNG] Increase the cluster size with preemptible worker nodes, and configure them to use graceful decommissioning.
    Phân tích: Hoàn hảo cân bằng: Tăng node preemptible → nhanh + rẻ; graceful decommissioning tự động (qua Dataproc/YARN config: --properties=spark:spark.yarn.appMasterEnv.GCLOUD_DATAPROC_GRACEFUL_DECOMMISSION=true) drain tasks an toàn → không mất work. Hỗ trợ lên đến 70% workers preemptible/cluster.

🔗 Tài liệu tham khảo (cập nhật GCP 2024-2026)

Hy vọng phân tích giúp bạn ôn thi hiệu quả! 🚀

Câu 207
You work for a shipping company that uses handheld scanners to read shipping labels. Your company has strict data privacy standards that require scanners to only transmit tracking numbers when events are sent to Kafka topics. A recent software update caused the scanners to accidentally transmit recipients' personally identifiable information (PII) to analytics systems, which violates user privacy rules. You want to quickly build a scalable solution using cloud-native managed services to prevent exposure of PII to the analytics systems. What should you do?
  1. A Create an authorized view in BigQuery to restrict access to tables with sensitive data.
  2. B Install a third-party data validation tool on Compute Engine virtual machines to check the incoming data for sensitive information.
  3. C Use Cloud Logging to analyze the data passed through the total pipeline to identify transactions that may contain sensitive information.
  4. D Build a Cloud Function that reads the topics and makes a call to the Cloud Data Loss Prevention (Cloud DLP) API. Use the tagging and confidence levels to either pass or quarantine the data in a bucket for review.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế tại công ty vận chuyển sử dụng handheld scanners để quét nhãn vận chuyển và chỉ gửi tracking numbers đến Kafka topics khi có sự kiện xảy ra. Công ty có tiêu chuẩn bảo mật dữ liệu nghiêm ngặt, yêu cầu không được truyền PII (Personally Identifiable Information - thông tin cá nhân nhận dạng). Tuy nhiên, sau cập nhật phần mềm gần đây, scanners vô tình gửi thêm PII đến hệ thống analytics, vi phạm quy định bảo mật.

📌 Yêu cầu giải pháp: Xây dựng nhanh chóng một giải pháp scalable sử dụng cloud-native managed services (dịch vụ quản lý gốc đám mây) của Google Cloud để ngăn chặn PII tiếp cận hệ thống analytics. Giải pháp phải:

  • Xử lý dữ liệu thời gian thực từ Kafka topics.
  • Phát hiện và lọc PII tự động.
  • Quản lý dữ liệu nhạy cảm an toàn (pass hợp lệ hoặc quarantine để review).
  • Đảm bảo tính scalable (mở rộng dễ dàng) và nhanh chóng triển khai.

🛠️ Bối cảnh kỹ thuật: Kafka topics là nguồn dữ liệu streaming. Giải pháp cần tích hợp với các dịch vụ GCP như Cloud Functions (serverless), Cloud DLP (phát hiện mất mát dữ liệu), và Storage buckets để xử lý.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Build a Cloud Function that reads the topics and makes a call to the Cloud Data Loss Prevention (Cloud DLP) API. Use the tagging and confidence levels to either pass or quarantine the data in a bucket for review.

Lý do chọn đáp án này 🏆:

  • Cloud Functions là dịch vụ serverless, scalable tự động, trigger trực tiếp từ Pub/Sub topics (tích hợp dễ với Kafka qua connectors như Kafka Connect to Pub/Sub), đọc dữ liệu real-time mà không cần quản lý server.
  • Cloud DLP API là dịch vụ managed chuyên dụng để quét, de-identify và tag PII với confidence levels (mức độ tin cậy cao/thấp), hỗ trợ tagging để quyết định pass (gửi tiếp đến analytics) hoặc quarantine vào Cloud Storage bucket để review thủ công.
  • Giải pháp nhanh chóng triển khai (code vài hàm Lambda-like), cloud-native, prevent exposure ngay từ nguồn, phù hợp yêu cầu strict data privacy.
  • Cập nhật 2026: Cloud DLP v2 hỗ trợ real-time scanning với ML models cải tiến, tích hợp Pub/Sub/Functions tốt hơn (theo GCP docs 2025+).

📘 Tài liệu tham khảo:

❌ Giải thích tất cả các phương án (đúng/sai)

  • [SAI] Create an authorized view in BigQuery to restrict access to tables with sensitive data.
    ❌ Sai vì: Authorized views chỉ kiểm soát access quyền đến dữ liệu đã lưu trong BigQuery, không prevent PII truyền vào pipeline từ Kafka. Data đã "rò rỉ" đến analytics trước khi vào BigQuery, vi phạm yêu cầu ngăn chặn từ nguồn. Không scalable real-time cho streaming data. (BigQuery phù hợp batch analytics, không phải filtering live).

  • [SAI] Install a third-party data validation tool on Compute Engine virtual machines to check the incoming data for sensitive information.
    ❌ Sai vì: Không phải cloud-native managed service (phải tự quản lý VM, install tool bên thứ 3), không scalable nhanh (cần provision VM, scaling thủ công). Vi phạm yêu cầu quickly build và managed services. Tốn chi phí O&M cao hơn serverless.

  • [SAI] Use Cloud Logging to analyze the data passed through the total pipeline to identify transactions that may contain sensitive information.
    ❌ Sai vì: Cloud Logging chỉ analyze sau sự kiện (log đã qua pipeline), không prevent PII đến analytics (data đã expose). Là công cụ monitoring, không phải filtering real-time. Không đáp ứng scalable solution to prevent exposure.

  • [ĐÚNG] Build a Cloud Function that reads the topics and makes a call to the Cloud Data Loss Prevention (Cloud DLP) API. Use the tagging and confidence levels to either pass or quarantine the data in a bucket for review.
    ✅ Đúng vì: Xem giải thích ở phần đáp án đúng trên. Hoàn hảo khớp tất cả yêu cầu: real-time, managed, scalable, PII detection chính xác với DLP.

🧩 Tóm tắt kiến trúc đề xuất: Kafka → Pub/Sub trigger → Cloud Function → DLP scan → Pass to Analytics / Quarantine to Bucket. Hoàn toàn serverless! 🚀

Câu 208
You have developed three data processing jobs. One executes a Cloud Dataflow pipeline that transforms data uploaded to Cloud Storage and writes results to
BigQuery. The second ingests data from on-premises servers and uploads it to Cloud Storage. The third is a Cloud Dataflow pipeline that gets information from third-party data providers and uploads the information to Cloud Storage. You need to be able to schedule and monitor the execution of these three workflows and manually execute them when needed. What should you do?
  1. A Create a Direct Acyclic Graph in Cloud Composer to schedule and monitor the jobs.
  2. B Use Observability Monitoring and set up an alert with a Webhook notification to trigger the jobs.
  3. C Develop an App Engine application to schedule and request the status of the jobs using GCP API calls.
  4. D Set up cron jobs in a Compute Engine instance to schedule and monitor the pipelines using GCP API calls.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống bạn đã phát triển ba công việc xử lý dữ liệu trên Google Cloud Platform (GCP):

  • Công việc 1: Một pipeline Cloud Dataflow biến đổi dữ liệu được tải lên Cloud Storage và ghi kết quả vào BigQuery. 📊➡️BigQuery
  • Công việc 2: Công việc thu thập dữ liệu từ các máy chủ on-premises và tải lên Cloud Storage. 🖥️➡️Cloud Storage
  • Công việc 3: Một pipeline Cloud Dataflow lấy thông tin từ các nhà cung cấp dữ liệu bên thứ ba và tải lên Cloud Storage. 🌐➡️Cloud Storage

Yêu cầu chính: Cần lập lịch (schedule), giám sát (monitor) các workflow này, và thực thi thủ công khi cần. 🕒👀✨
Đây là bài toán orchestration workflow điển hình trên GCP, nơi các job phụ thuộc lẫn nhau hoặc cần chạy định kỳ, với khả năng trigger thủ công. Cloud Composer (dịch vụ managed Apache Airflow) là giải pháp lý tưởng vì hỗ trợ DAG (Directed Acyclic Graph) để định nghĩa, schedule, monitor và trigger jobs qua GCP operators. (Cập nhật đến 2026: Cloud Composer v3 vẫn là best practice cho data orchestration, tích hợp sâu với Dataflow, Storage, BigQuery).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create a Direct Acyclic Graph in Cloud Composer to schedule and monitor the jobs.
Lý do:

  • Cloud Composer cho phép tạo DAG để orchestrate các job Dataflow và GCS operations một cách managed, scalable. 🛠️
  • Hỗ trợ schedule (cron-like), monitor (UI dashboard, logs, metrics), và manual trigger qua Airflow UI hoặc CLI.
  • Tích hợp native với GCP services: DataflowRunJavaOperator, GCSToBigQueryOperator, GCSHook cho upload từ on-prem/third-party.
  • Best practice theo GCP Well-Architected Framework cho data pipelines (không cần quản lý infrastructure).
    📘 Nguồn tham khảo: Cloud Composer Documentation & Dataflow Orchestration with Composer (cập nhật 2025-2026).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng/best) hoặc ❌ (sai/không phù hợp), kèm lý do chi tiết bằng tiếng Việt:

  • Create a Direct Acyclic Graph in Cloud Composer to schedule and monitor the jobs.
    ✅ Đúng và là lựa chọn tốt nhất. Như đã giải thích ở trên, DAG trong Cloud Composer (Apache Airflow managed) xử lý hoàn hảo việc schedule, monitor real-time (qua Airflow UI, Stackdriver/Cloud Monitoring), và manual execution. Hỗ trợ dependencies giữa các job (ví dụ: job 2 chạy trước job 1). Scalable, serverless, không lo maintenance. 🏆

  • Use Observability Monitoring and set up an alert with a Webhook notification to trigger the jobs.
    ❌ Sai. Cloud Monitoring (Observability) chỉ dùng để giám sát metrics/logs/alerts, không phải để schedule hoặc orchestrate workflows. Webhook chỉ trigger đơn giản (như notification), không hỗ trợ DAG logic, dependencies, hay manual retry. Không phù hợp cho data jobs phức tạp. 🚫

  • Develop an App Engine application to schedule and request the status of the jobs using GCP API calls.
    ❌ Sai. App Engine có thể dùng cron jobs và API calls (Dataflow API, Storage API) để schedule/monitor, nhưng đây là custom solution, thiếu native orchestration UI, DAG support, và scalability cho complex workflows. Phải tự code error handling, retries – không phải best practice. Tốn công phát triển/maintain. 💻❌

  • Set up cron jobs in a Compute Engine instance to schedule and monitor the pipelines using GCP API calls.
    ❌ Sai. Cron trên VM Compute Engine dùng script + API calls có thể schedule, nhưng không managed: phải tự quản lý VM (uptime, scaling, security), monitor thủ công qua logs. Không có UI dashboard, dependencies handling, hay auto-retries như Composer. Rủi ro downtime cao, vi phạm nguyên tắc serverless. 🖥️🚫

🛠️ Kết luận & Lời khuyên

Sử dụng Cloud Composer là cách hiệu quả nhất để quản lý các data workflows trên GCP, giảm operational overhead. Nếu scale lớn, kết hợp với Cloud Scheduler cho simple triggers. Test DAG sample tại Composer Quickstart. 🎯

Câu 209 Chọn nhiều đáp án
You have Cloud Functions written in Node.js that pull messages from Cloud Pub/Sub and send the data to BigQuery. You observe that the message processing rate on the Pub/Sub topic is orders of magnitude higher than anticipated, but there is no error logged in Cloud Logging. What are the two most likely causes of this problem? (Choose two.)
  1. A Publisher throughput quota is too small.
  2. B Total outstanding messages exceed the 10-MB maximum.
  3. C Error handling in the subscriber code is not handling run-time errors properly.
  4. D The subscriber code cannot keep up with the messages.
  5. E The subscriber code does not acknowledge the messages that it pulls.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi mô tả một tình huống trong Google Cloud Platform (GCP): Bạn có các Cloud Functions viết bằng Node.js, chúng pull messages từ Cloud Pub/Sub topic và gửi dữ liệu vào BigQuery. Vấn đề quan sát được là tỷ lệ xử lý messages trên Pub/Sub topic (message processing rate) cao hơn dự kiến rất nhiều (orders of magnitude higher), nhưng không có lỗi nào được ghi log trong Cloud Logging.
📌 Ý nghĩa vấn đề: "Processing rate cao hơn anticipated" nghĩa là hệ thống dường như đang xử lý một lượng messages khổng lồ, nhưng không có lỗi log → có thể do messages bị lặp lại xử lý (redelivery) mà không được xác nhận (acknowledge), dẫn đến subscriber liên tục pull cùng messages. Câu hỏi yêu cầu chọn hai nguyên nhân likely nhất (choose two).

✅ Đáp án đúng và lý do lựa chọn

Hai đáp án đúng là:

  • Error handling in the subscriber code is not handling run-time errors properly.
  • The subscriber code does not acknowledge the messages that it pulls.

🛠️ Lý do chính:
Trong Pub/Sub, khi subscriber pull messages, nó phải ack (xác nhận) để xóa messages khỏi queue. Nếu không ack, messages sẽ được redelivered sau ack deadline (mặc định 10 giây), gây loop xử lý lặp vô tận → processing rate "bùng nổ" nhưng không có tiến triển thực tế. Đồng thời, nếu error handling kém (ví dụ: không catch lỗi runtime đúng cách), code có thể crash ngầm hoặc không ack mà không log lỗi vào Cloud Logging (vì Cloud Functions chỉ log những gì được explicit ghi). Điều này khớp hoàn hảo với triệu chứng: rate cao bất thường, không error log. Kiến thức cập nhật GCP 2024-2026: Pub/Sub vẫn giữ cơ chế ack/modack như vậy, Cloud Functions v2 hỗ trợ structured logging nhưng không tự log tất cả runtime errors nếu code không handle.

📋 Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Tôi đánh dấu ✅ cho đúng, ❌ cho sai, và giải thích rõ lý do bằng tiếng Việt:

  • ❌ Publisher throughput quota is too small.
    🧨 Sai vì: Publisher quota nhỏ sẽ giới hạn tốc độ publish messages từ producer, dẫn đến ít messages hơn trên topic chứ không phải processing rate cao bất thường. Vấn đề ở đây là subscriber-side, không phải publisher. Nếu quota nhỏ, sẽ có lỗi throttle ở publisher logs, không khớp triệu chứng.

  • ❌ Total outstanding messages exceed the 10-MB maximum.
    🔒 Sai vì: Pub/Sub có giới hạn 10MB per pull response (undelivered messages), nhưng vượt quá sẽ gây lỗi explicit ở subscriber (như DEADLINE_EXCEEDED) và log vào Cloud Logging. Triệu chứng không có error log, nên không phải nguyên nhân. (Cập nhật 2026: Giới hạn vẫn 10MB/pull, không thay đổi).

  • ✅ Error handling in the subscriber code is not handling run-time errors properly.
    🐛 Đúng vì: Trong Node.js Cloud Functions pull từ Pub/Sub, nếu không dùng try-catch đúng cách (ví dụ: không ack khi error xảy ra), runtime errors (như BigQuery insert fail) sẽ làm function silent-fail mà không log (Cloud Functions chỉ log console.error explicit). Kết hợp không ack → redelivery loop, rate cao. Khớp triệu chứng hoàn hảo.

  • ❌ The subscriber code cannot keep up with the messages.
    ⏳ Sai vì: Nếu subscriber chậm (cannot keep up), sẽ tạo backlog lớn trên topic (unack messages tăng), nhưng processing rate sẽ thấp hơn anticipated chứ không cao. Pub/Sub metrics sẽ show high unacked, và có thể có retry errors logged – không khớp "no error logged".

  • ✅ The subscriber code does not acknowledge the messages that it pulls.
    🔄 Đúng vì: Không gọi ack() sau pull thành công → messages redelivered liên tục sau ack deadline (10s default). Subscriber loop pull cùng messages → processing rate "vô hạn" cao hơn anticipated, nhưng không error vì code chạy bình thường. Classic Pub/Sub anti-pattern!

📘 Tài liệu tham khảo

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm ví dụ code Node.js fix issue, cứ hỏi nhé!

Câu 210
You are creating a new pipeline in Google Cloud to stream IoT data from Cloud Pub/Sub through Cloud Dataflow to BigQuery. While previewing the data, you notice that roughly 2% of the data appears to be corrupt. You need to modify the Cloud Dataflow pipeline to filter out this corrupt data. What should you do?
  1. A Add a SideInput that returns a Boolean if the element is corrupt.
  2. B Add a ParDo transform in Cloud Dataflow to discard corrupt elements.
  3. C Add a Partition transform in Cloud Dataflow to separate valid data from corrupt data.
  4. D Add a GroupByKey transform in Cloud Dataflow to group all of the valid data together and discard the rest.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả tình huống bạn đang xây dựng một pipeline mới trên Google Cloud để stream dữ liệu IoT từ Cloud Pub/Sub qua Cloud Dataflow đến BigQuery. Khi preview dữ liệu, bạn phát hiện khoảng 2% dữ liệu bị corrupt (hỏng). Nhiệm vụ là sửa đổi pipeline Cloud Dataflow để lọc bỏ (filter out) dữ liệu corrupt này, đảm bảo chỉ dữ liệu hợp lệ được xử lý tiếp.

Mục tiêu chính: Sử dụng transform phù hợp trong Cloud Dataflow (dựa trên Apache Beam) để loại bỏ dữ liệu hỏng một cách hiệu quả, scalable cho streaming pipeline. Không cần tách riêng hay group dữ liệu, chỉ đơn giản discard corrupt elements.

📘 Kiến thức cập nhật: Theo tài liệu Google Cloud Dataflow mới nhất (2024-2026), Cloud Dataflow sử dụng Apache Beam 2.58+ (phiên bản stable mới nhất), nơi ParDo là transform cốt lõi để xử lý từng element một cách tùy chỉnh, hỗ trợ filter/discard dễ dàng qua DoFn.

✅ Đáp án đúng và lý do chọn

Đáp án đúng: Add a ParDo transform in Cloud Dataflow to discard corrupt elements.

Lý do 🛠️:

  • ParDo là transform cơ bản nhất trong Beam/Dataflow, cho phép áp dụng hàm DoFn (ProcessElement) lên từng element riêng lẻ.
  • Trong DoFn, bạn kiểm tra điều kiện corrupt (ví dụ: parse JSON thất bại, giá trị null, format sai), rồi discard bằng cách không output element đó (return empty Iterable hoặc điều kiện if).
  • Hoàn hảo cho streaming (Pub/Sub → Dataflow → BigQuery), hiệu suất cao, không cần key/grouping, tránh overhead không cần thiết.
  • Ví dụ code:
    public class FilterCorruptFn extends DoFn<String, String> {
      @ProcessElement
      public void processElement(ProcessContext c) {
        String elem = c.element();
        if (isValid(elem)) {  // Kiểm tra corrupt
          c.output(elem);
        }  // Discard nếu corrupt
      }
    }
    
  • Đây là best practice cho filtering trong docs chính thức.

Nguồn tham khảo 📚:

❌ Phân tích tất cả các phương án

Dưới đây là giải thích chi tiết từng lựa chọn, với giữ nguyên văn bản gốc tiếng Anh:

  • [SAI] Add a SideInput that returns a Boolean if the element is corrupt.
    ❌ Sai vì: SideInput dùng để enrich data từ nguồn phụ (như lookup table), không phải filter/discard element chính. Nó trả Boolean nhưng không loại bỏ element corrupt trực tiếp – bạn vẫn cần ParDo để xử lý. Overhead cao cho streaming, không scalable cho 2% corrupt đơn giản. SideInput phù hợp validate phức tạp với external data, không phải discard inline.

  • [ĐÚNG] Add a ParDo transform in Cloud Dataflow to discard corrupt elements.
    ✅ Đúng vì: Như giải thích trên, ParDo linh hoạt nhất để xử lý từng element, kiểm tra và discard corrupt mà không ảnh hưởng pipeline. Hiệu quả, zero-overhead cho valid data (98%), tích hợp streaming hoàn hảo.

  • [SAI] Add a Partition transform in Cloud Dataflow to separate valid data from corrupt data.
    ❌ Sai vì: Partition dùng để chia collection thành N partitions dựa hàm trả index (0 đến numPartitions-1). Nó tách valid/corrupt nhưng tạo 2 output PCollections, yêu cầu downstream xử lý riêng (ví dụ: sink valid to BigQuery, corrupt to Dead Letter). Phức tạp hơn cần thiết, overhead shuffling, không discard trực tiếp – chỉ separate.

  • [SAI] Add a GroupByKey transform in Cloud Dataflow to group all of the valid data together and discard the rest.
    ❌ Sai vì: GroupByKey group elements theo key trước khi xử lý (gây shuffle lớn, stateful cho streaming). Không liên quan filter corrupt (không có khái niệm "valid data together"), dễ gây watermark issues/out-of-order ở streaming. Discard "the rest" không khả thi vì GroupByKey output grouped values, overhead khổng lồ cho IoT high-volume data.

Tóm tắt khuyến nghị 🚀: Sử dụng ParDo ngay sau Pub/Sub read để filter sớm, tối ưu chi phí BigQuery insert và pipeline performance! Nếu corrupt pattern phức tạp, kết hợp với TryParse/FlatMap.