Ngân hàng đề — Google Cloud Professional Data Engineer

Tìm thấy 429 câu.

Câu 341
You are designing the architecture of your application to store data in Cloud Storage. Your application consists of pipelines that read data from a Cloud Storage bucket that contains raw data, and write the data to a second bucket after processing. You want to design an architecture with Cloud Storage resources that are capable of being resilient if a Google Cloud regional failure occurs. You want to minimize the recovery point objective (RPO) if a failure occurs, with no impact on applications that use the stored data. What should you do?
  1. A Adopt multi-regional Cloud Storage buckets in your architecture.
  2. B Adopt two regional Cloud Storage buckets, and update your application to write the output on both buckets.
  3. C Adopt a dual-region Cloud Storage bucket, and enable turbo replication in your architecture.
  4. D Adopt two regional Cloud Storage buckets, and create a daily task to copy from one bucket to the other.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi

Câu hỏi tập trung vào việc thiết kế kiến trúc ứng dụng sử dụng Cloud Storage trên Google Cloud Platform (GCP) để lưu trữ dữ liệu. Ứng dụng bao gồm các pipeline đọc dữ liệu thô từ một bucket Cloud Storage, xử lý và ghi vào bucket thứ hai. Yêu cầu chính là:

  • Kiến trúc phải resilient (bền vững) trước sự cố regional failure (một region Google Cloud bị lỗi).
  • Tối thiểu hóa Recovery Point Objective (RPO): khoảng thời gian mất dữ liệu tối đa khi xảy ra sự cố (càng thấp càng tốt, lý tưởng là gần 0).
  • Không ảnh hưởng đến ứng dụng sử dụng dữ liệu (ứng dụng vẫn đọc/ghi bình thường mà không gián đoạn).

📘 Bối cảnh kiến thức GCP Cloud Storage (cập nhật đến 2026): Cloud Storage hỗ trợ các loại bucket khác nhau để đảm bảo độ bền và sẵn sàng cao:

  • Regional bucket: Dữ liệu chỉ ở một region duy nhất → Không resilient nếu region đó fail.
  • Dual-regional bucket: Dữ liệu replicate giữa hai region cụ thể gần nhau (ví dụ: us-west1 và us-central1) → Resilient với regional failure.
  • Multi-regional bucket: Dữ liệu replicate rộng hơn qua nhiều region → Resilient cao hơn nhưng chi phí cao, RPO có thể không tối ưu.
  • Turbo Replication (tính năng mới từ 2023, ổn định đến 2026): Chỉ áp dụng cho dual/multi-regional buckets, giảm độ trễ replication xuống sub-minute (dưới 1 phút, thường chỉ vài giây) → Minimize RPO hiệu quả nhất, không cần thay đổi code ứng dụng.

Đáp án đúng: ✅ Adopt a dual-region Cloud Storage bucket, and enable turbo replication in your architecture.

Lý do chọn đáp án đúng (🛠️ Giải thích chi tiết):

  • Dual-region bucket đảm bảo dữ liệu tự động replicate giữa hai region → Nếu một region fail, dữ liệu vẫn khả dụng ở region kia, resilient hoàn toàn.
  • Turbo replication kích hoạt replication gần real-time (RPO <1 phút), giảm thiểu mất dữ liệu so với replication thông thường (có thể vài phút).
  • Không ảnh hưởng ứng dụng: Ứng dụng chỉ cần dùng một bucket duy nhất (read/write bình thường), GCP tự handle replication → Zero downtime cho apps.
  • Đây là giải pháp best practice của GCP cho high availability + low RPO mà không phức tạp hóa code.

📋 Giải thích tất cả các phương án (Đúng/Sai)

  • ❌ [SAI] Adopt multi-regional Cloud Storage buckets in your architecture.
    Phương án này sử dụng bucket multi-regional (replicate dữ liệu qua nhiều region toàn cầu). Tuy resilient với regional failure, nhưng RPO không được minimize tối ưu (replication asynchronous, độ trễ có thể vài phút mà không có turbo). Ngoài ra, chi phí cao hơn dual-region và không tận dụng turbo replication (turbo chỉ hiệu quả nhất ở dual-region). Không phải lựa chọn tốt nhất cho yêu cầu low RPO + no impact.

  • ❌ [SAI] Adopt two regional Cloud Storage buckets, and update your application to write the output on both buckets.
    Sử dụng hai regional bucket riêng biệt, ứng dụng phải dual-write (ghi đồng thời vào cả hai). Không resilient tự động: Nếu một region fail, dữ liệu chỉ ở bucket kia, nhưng ứng dụng phải thay đổi code lớn (dual-write phức tạp, dễ lỗi nếu một write fail). RPO phụ thuộc vào ứng dụng (có thể mất data chưa sync). Ảnh hưởng lớn đến ứng dụng, vi phạm yêu cầu.

  • ✅ [ĐÚNG] Adopt a dual-region Cloud Storage bucket, and enable turbo replication in your architecture.
    Như đã giải thích ở trên: Dual-region resilient với regional failure, turbo replication minimize RPO xuống sub-minute, không cần thay đổi ứng dụng (GCP tự replicate). Giải pháp tối ưu nhất theo best practice GCP.

  • ❌ [SAI] Adopt two regional Cloud Storage buckets, and create a daily task to copy from one bucket to the other.
    Sử dụng hai regional bucket, sync thủ công hàng ngày qua task (ví dụ: Cloud Functions hoặc Data Transfer Service). RPO cực kỳ cao (24 giờ) → Mất toàn bộ dữ liệu 1 ngày nếu fail, không đáp ứng minimize RPO. Sync không real-time, dễ lỗi task, và vẫn cần quản lý hai bucket → Không resilient và không hiệu quả.

📚 Tài liệu tham khảo (GCP chính thức, cập nhật 2026)

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm ví dụ code hoặc diagram, hãy hỏi nhé!

Câu 342
You have designed an Apache Beam processing pipeline that reads from a Pub/Sub topic. The topic has a message retention duration of one day, and writes to a Cloud Storage bucket. You need to select a bucket location and processing strategy to prevent data loss in case of a regional outage with an RPO of 15 minutes. What should you do?
  1. A 1. Use a dual-region Cloud Storage bucket.
    2. Monitor Dataflow metrics with Cloud Monitoring to determine when an outage occurs.
    3. Seek the subscription back in time by 15 minutes to recover the acknowledged messages.
    4. Start the Dataflow job in a secondary region.
  2. B 1. Use a multi-regional Cloud Storage bucket.
    2. Monitor Dataflow metrics with Cloud Monitoring to determine when an outage occurs.
    3. Seek the subscription back in time by 60 minutes to recover the acknowledged messages.
    4. Start the Dataflow job in a secondary region.
  3. C 1. Use a regional Cloud Storage bucket.
    2. Monitor Dataflow metrics with Cloud Monitoring to determine when an outage occurs.
    3. Seek the subscription back in time by one day to recover the acknowledged messages.
    4. Start the Dataflow job in a secondary region and write in a bucket in the same region.
  4. D 1. Use a dual-region Cloud Storage bucket with turbo replication enabled.
    2. Monitor Dataflow metrics with Cloud Monitoring to determine when an outage occurs.
    3. Seek the subscription back in time by 60 minutes to recover the acknowledged messages.
    4. Start the Dataflow job in a secondary region.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc thiết kế một pipeline xử lý dữ liệu bằng Apache Beam trên Google Cloud Platform (GCP), cụ thể đọc dữ liệu từ Pub/Sub topic (với thời gian lưu trữ message retention là 1 ngày) và ghi vào Cloud Storage bucket. Mục tiêu là chọn vị trí bucket (bucket location) và chiến lược xử lý (processing strategy) để ngăn ngừa mất dữ liệu trong trường hợp regional outage (sự cố gián đoạn tại một vùng địa lý cụ thể), với yêu cầu RPO (Recovery Point Objective) là 15 phút – nghĩa là dữ liệu mất mát không vượt quá 15 phút gần nhất.

🔍 Các yếu tố chính cần xem xét:

  • Pub/Sub: Hỗ trợ seek (quay ngược thời gian subscription) để replay lại các message đã được acknowledge (ack), tối đa theo retention period (1 ngày ở đây).
  • Cloud Storage: Cần bucket có khả năng replicate dữ liệu nhanh chóng giữa các vùng để chịu outage (regional bucket dễ mất, multi-regional deprecated ở một số khu vực, dual-region với turbo replication là lựa chọn hiện đại cho RPO thấp).
  • Dataflow (chạy Beam pipeline): Cần monitor metrics và khả năng failover sang vùng phụ (secondary region).
  • RPO 15 phút: Yêu cầu replication và recovery phải đảm bảo mất dữ liệu ≤15 phút, nên cần turbo replication (replicate trong vài phút) và seek back đủ xa để cover window xử lý.

📘 Kiến thức cập nhật đến 2026: Dựa trên GCP docs mới nhất (Cloud Storage dual-region turbo replication ra mắt 2023, vẫn là best practice 2026 cho low RPO; Pub/Sub seek hỗ trợ retention-based recovery; Dataflow v2.x+ với regional failover).

✅ Đáp án đúng: Phương án thứ 4

Lý do chọn:

  • Đây là giải pháp toàn diện nhất, kết hợp dual-region bucket với turbo replication (replicate dữ liệu giữa 2 vùng trong ~5-15 phút, khớp RPO 15 phút), monitor Dataflow để detect outage kịp thời, seek back 60 phút (an toàn hơn RPO, cover window xử lý Beam/Dataflow có thể delay ack), và start job ở secondary region để failover nhanh mà không mất pipeline state.
  • Đảm bảo zero data loss thực tế nhờ replay từ Pub/Sub và bucket replicate cao cấp.

Các lựa chọn (giữ nguyên văn bản gốc):

  • ❌ Phương án 1:

    1. Use a dual-region Cloud Storage bucket.
    2. Monitor Dataflow metrics with Cloud Monitoring to determine when an outage occurs.
    3. Seek the subscription back in time by 15 minutes to recover the acknowledged messages.
    4. Start the Dataflow job in a secondary region.
      Giải thích sai: Dual-region thông thường chỉ replicate asynchronous (có thể >15 phút RPO), không turbo nên không đảm bảo RPO 15 phút. Seek back chỉ 15 phút quá ngắn, không cover delay xử lý Beam/Dataflow (thường cần buffer 30-60 phút). Không đủ robust cho outage.
  • ❌ Phương án 2:

    1. Use a multi-regional Cloud Storage bucket.
    2. Monitor Dataflow metrics with Cloud Monitoring to determine when an outage occurs.
    3. Seek the subscription back in time by 60 minutes to recover the acknowledged messages.
    4. Start the Dataflow job in a secondary region.
      Giải thích sai: Multi-regional (US-only) deprecated từ 2023, không khuyến khích mới (2026 ưu tiên dual-region). Replication chậm hơn turbo (RPO có thể >15 phút). Seek 60 phút tốt nhưng bucket không optimal.
  • ❌ Phương án 3:

    1. Use a regional Cloud Storage bucket.
    2. Monitor Dataflow metrics with Cloud Monitoring to determine when an outage occurs.
    3. Seek the subscription back in time by one day to recover the acknowledged messages.
    4. Start the Dataflow job in a secondary region and write in a bucket in the same region.
      Giải thích sai: Regional bucket dễ mất toàn bộ data nếu outage vùng đó. Seek 1 ngày quá dài (unnecessary, tốn resource), và step 4 mâu thuẫn (secondary region nhưng bucket same region → vẫn outage cùng lúc). Không đạt RPO.
  • ✅ Phương án 4 (ĐÚNG):

    1. Use a dual-region Cloud Storage bucket with turbo replication enabled.
    2. Monitor Dataflow metrics with Cloud Monitoring to determine when an outage occurs.
    3. Seek the subscription back in time by 60 minutes to recover the acknowledged messages.
    4. Start the Dataflow job in a secondary region.
      Giải thích đúng: Turbo replication (dual-region) đảm bảo RPO ~5 phút (docs GCP 2026). Monitor cho alert nhanh. Seek 60 phút cover RPO + buffer (Pub/Sub retention 1 ngày cho phép). Secondary region failover Dataflow seamless. Hoàn hảo! 🛠️

📚 Tài liệu tham khảo

Câu 343
You are preparing data that your machine learning team will use to train a model using BigQueryML. They want to predict the price per square foot of real estate. The training data has a column for the price and a column for the number of square feet. Another feature column called ‘feature1’ contains null values due to missing data. You want to replace the nulls with zeros to keep more data points. Which query should you use?
  1. A
    SELECT * EXCEPT (feature1),
    IFNULL (feature1, 0) AS feature1_cleaned
    FROM training_data;

  2. B
    SELECT * EXCEPT (price, square_feet),
           price/square_feet AS price_per_sqft
    FROM training_data
    WHERE feature1 IS NOT NULL;

  3. C
    SELECT * EXCEPT (price, square_feet, feature1),
           price/square_feet AS price_per_sqft,
           IFNULL(feature1, 0) AS feature1_cleaned
    FROM training_data;

  4. D
    SELECT *
    FROM training_data
    WHERE feature1 IS NOT NULL;
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc chuẩn bị dữ liệu (data preparation) cho đội ngũ machine learning sử dụng BigQuery ML (một tính năng của Google Cloud BigQuery cho phép huấn luyện mô hình ML trực tiếp bằng SQL). Mục tiêu là dự đoán giá mỗi feet vuông (price per square foot) của bất động sản.

Dữ liệu huấn luyện (training_data) có:

  • Cột price (giá tổng).
  • Cột square_feet (diện tích).
  • Cột feature1 chứa một số giá trị NULL do dữ liệu thiếu.

Yêu cầu chính:

  • Tạo cột target mới: price_per_sqft = price / square_feet.
  • Xử lý NULL trong feature1 bằng cách thay thế bằng 0 để giữ nguyên số lượng data points (không loại bỏ hàng nào).
  • Giữ tất cả các cột khác (ngoại trừ các cột cần transform).

📘 Lưu ý kiến thức cập nhật (đến 2026): BigQuery hỗ trợ SELECT * EXCEPT (từ phiên bản SQL 2011+), IFNULL(expr, value) để thay thế NULL, và BigQuery ML yêu cầu dữ liệu sạch với target rõ ràng cho các mô hình regression (như CREATE MODEL ... AS SELECT ...). Không có thay đổi lớn trong cú pháp này theo tài liệu AWS? (Câu hỏi là GCP BigQuery, không phải AWS; có thể nhầm lẫn chủ đề).

Nguồn tham khảo:

✅ Đáp án đúng

SELECT * EXCEPT (price, square_feet, feature1),
price/square_feet AS price_per_sqft,
IFNULL(feature1, 0) AS feature1_cleaned
FROM training_data;

Lý do chọn đáp án này 🛠️:

  • Sử dụng SELECT * EXCEPT để loại bỏ chính xác 3 cột gốc cần transform (price, square_feet, feature1), giữ nguyên tất cả cột feature khác.
  • Tạo target mới price_per_sqft cho mô hình ML dự đoán (regression task).
  • Áp dụng IFNULL(feature1, 0) để thay thế NULL bằng 0, giữ 100% data points mà không dùng WHERE filter.
  • Query hoàn chỉnh, an toàn cho BigQuery ML (CREATE MODEL sẽ dùng cột price_per_sqft làm label).

📋 Giải thích tất cả các phương án

  • ❌ Phương án 1 (SAI):

    SELECT * EXCEPT (feature1),
    IFNULL (feature1, 0) AS feature1_cleaned
    FROM training_data;

    Lý do sai ❌: Chỉ clean feature1 bằng cách thay NULL=0 và loại bỏ cột gốc, nhưng không tạo cột target price_per_sqft (price/square_feet). Mô hình ML không có label để train dự đoán giá/ft². Query không đáp ứng yêu cầu chính.

  • ❌ Phương án 2 (SAI):

    SELECT * EXCEPT (price, square_feet),
    price/square_feet AS price_per_sqft
    FROM training_data
    WHERE feature1 IS NOT NULL;

    Lý do sai ❌: Tạo đúng price_per_sqft và loại bỏ price + square_feet, nhưng dùng WHERE feature1 IS NOT NULL sẽ loại bỏ các hàng có NULL, vi phạm yêu cầu "keep more data points" bằng cách thay NULL=0. Mất dữ liệu không cần thiết.

  • ✅ Phương án 3 (ĐÚNG): (Đã giải thích chi tiết ở trên)
    Hoàn hảo, đầy đủ yêu cầu! 🎯

  • ❌ Phương án 4 (SAI):

    SELECT *
    FROM training_data WHERE feature1 IS NOT NULL;

    Lý do sai ❌: Chỉ filter bỏ NULL ở feature1, không tạo price_per_sqft, không clean gì cả. Dữ liệu vẫn thô, thiếu target cho ML, và mất data points do WHERE clause. Hoàn toàn không phù hợp.

Câu 344
Different teams in your organization store customer and performance data in BigQuery. Each team needs to keep full control of their collected data, be able to query data within their projects, and be able to exchange their data with other teams. You need to implement an organization-wide solution, while minimizing operational tasks and costs. What should you do?
  1. A Ask each team to create authorized views of their data. Grant the biquery.jobUser role to each team.
  2. B Create a BigQuery scheduled query to replicate all customer data into team projects.
  3. C Ask each team to publish their data in Analytics Hub. Direct the other teams to subscribe to them.
  4. D Enable each team to create materialized views of the data they need to access in their projects.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc triển khai một giải pháp tổ chức-wide (organization-wide) trên Google Cloud BigQuery để các team khác nhau trong tổ chức có thể:

  • Giữ quyền kiểm soát hoàn toàn (full control) dữ liệu của riêng mình (customer và performance data).
  • Query dữ liệu trực tiếp trong project của họ.
  • Trao đổi dữ liệu (exchange data) với các team khác một cách an toàn. Yêu cầu chính: Giảm thiểu công việc vận hành (operational tasks) và chi phí (costs).

📘 Bối cảnh: BigQuery là data warehouse serverless của Google Cloud, hỗ trợ chia sẻ dữ liệu mà không cần copy dữ liệu vật lý (zero-copy sharing), giúp tránh chi phí lưu trữ và quản lý dữ liệu dư thừa. Giải pháp cần đảm bảo bảo mật, quyền kiểm soát, và không làm tăng tải vận hành.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Ask each team to publish their data in Analytics Hub. Direct the other teams to subscribe to them.

Lý do 🛠️:

  • Analytics Hub (tính năng mới nhất của BigQuery, cập nhật đến 2026) cho phép publish dữ liệu dưới dạng linked datasets (liên kết zero-copy), nghĩa là các team khác có thể subscribe và query dữ liệu mà không copy dữ liệu vật lý, giữ nguyên quyền kiểm soát ở owner gốc.
  • Đáp ứng đầy đủ: Mỗi team kiểm soát data (chỉ publish những gì muốn), query trong project mình (qua subscription), trao đổi dễ dàng.
  • Minimize operational tasks & costs: Không cần replicate/copy data → tiết kiệm chi phí storage/query, tự động hóa sharing qua UI/API, không cần quản lý views/schedules thủ công.
  • Đây là best practice organization-wide theo docs Google Cloud (xem nguồn dưới).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Tôi đánh dấu ✅ (đúng) hoặc ❌ (sai) và giải thích chi tiết bằng tiếng Việt dựa trên kiến thức BigQuery mới nhất (2026).

  • ❌ [SAI] Ask each team to create authorized views of their data. Grant the biquery.jobUser role to each team.

    • Giải thích sai: Authorized views chỉ cho phép query cross-project qua IAM, nhưng yêu cầu grant bigquery.jobUser (lưu ý: lỗi chính tả "biquery" → bigquery.jobUser) sẽ cho phép chạy jobs/query → rủi ro bảo mật cao (team khác có thể query toàn bộ data gốc nếu biết table). Không giữ "full control" (phải share views rộng), tăng operational tasks (quản lý views/IAM thủ công), và không scale organization-wide hiệu quả. Chi phí query có thể tăng nếu views không optimize.
  • ❌ [SAI] Create a BigQuery scheduled query to replicate all customer data into team projects.

    • Giải thích sai: Scheduled queries replicate copy dữ liệu vật lý vào project khác → tăng chi phí storage cao (double/triple data), mất đồng bộ real-time (lag), và mất full control (data bị duplicate, khó thu hồi). Operational tasks lớn: Quản lý schedules, error handling, cleanup. Không phù hợp minimize costs/tasks.
  • ✅ [ĐÚNG] Ask each team to publish their data in Analytics Hub. Direct the other teams to subscribe to them.

    • Giải thích đúng: Như đã nêu ở phần đáp án. Zero-copy sharing qua listings/subscriptions, hỗ trợ fine-grained access control (column/row level), audit logs đầy đủ. Cập nhật 2026: Tích hợp sâu với Data Catalog, hỗ trợ private listings organization-wide. Hoàn hảo cho multi-team collaboration mà không duplicate data.
  • ❌ [SAI] Enable each team to create materialized views of the data they need to access in their projects.

    • Giải thích sai: Materialized views cache data vật lý trong project đích → copy data, tăng storage costs, và chỉ refresh theo schedule (không real-time). Không giữ "full control" ở source team (data bị materialize độc lập), operational tasks cao (quản lý refresh, storage quota). Không hỗ trợ exchange hai chiều dễ dàng.

📚 Tài liệu tham khảo (cập nhật mới nhất 2026)

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm ví dụ code Terraform/IAM, hãy hỏi nhé.

Câu 345
You are developing a model to identify the factors that lead to sales conversions for your customers. You have completed processing your data. You want to continue through the model development lifecycle. What should you do next?
  1. A Use your model to run predictions on fresh customer input data.
  2. B Monitor your model performance, and make any adjustments needed.
  3. C Delineate what data will be used for testing and what will be used for training the model.
  4. D Test and evaluate your model on your curated data to determine how well the model performs.
Xem giải thích

🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một tình huống trong quy trình phát triển mô hình máy học (machine learning - ML) trên AWS: Bạn đang xây dựng mô hình để xác định các yếu tố ảnh hưởng đến chuyển đổi bán hàng (sales conversions) cho khách hàng. Dữ liệu đã được xử lý hoàn tất (data processing phase hoàn thành). Bây giờ, bạn cần xác định bước tiếp theo trong vòng đời phát triển mô hình (model development lifecycle).
Quy trình ML tiêu chuẩn trên AWS (theo SageMaker hoặc các dịch vụ ML như Amazon Forecast, SageMaker Canvas) bao gồm các giai đoạn chính: Thu thập dữ liệu → Xử lý dữ liệu → Chia dữ liệu (train/test split) → Huấn luyện mô hình → Đánh giá → Triển khai → Giám sát. Sau khi xử lý dữ liệu, bước logic tiếp theo là chia dữ liệu để chuẩn bị huấn luyện và đánh giá, tránh overfitting và đảm bảo tính tổng quát hóa của mô hình. Đây là best practice cập nhật đến năm 2026 trong AWS Well-Architected Framework for ML (Lens: Operational Excellence và Reliability).

✅ Đáp án đúng: Delineate what data will be used for testing and what will be used for training the model.
Lý do lựa chọn: Sau khi hoàn thành xử lý dữ liệu, bước tiếp theo bắt buộc là phân chia dữ liệu (data delineation/splitting) thành tập huấn luyện (training set - dùng để train model) và tập kiểm tra (testing set - dùng để đánh giá unbiased). Điều này đảm bảo mô hình học được từ dữ liệu huấn luyện mà không "nhìn trộm" dữ liệu test, tuân thủ nguyên tắc ML lifecycle trên AWS SageMaker (ví dụ: sử dụng train_test_split trong Processing Jobs hoặc SageMaker Data Wrangler). Nếu bỏ qua, mô hình dễ bị overfitting. Đây là bước cốt lõi trước khi train model, phù hợp với phiên bản SageMaker mới nhất (2026 updates nhấn mạnh automated splitting với SageMaker Autopilot).

🛠️ Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh, chỉ giải thích bằng tiếng Việt với lý do đúng/sai dựa trên ML lifecycle AWS:

  • ❌ Use your model to run predictions on fresh customer input data.
    Sai vì: Lúc này chưa có mô hình nào được huấn luyện (chưa train model), nên không thể chạy dự đoán (predictions) trên dữ liệu mới (fresh data). Bước này thuộc giai đoạn triển khai và inference (sau khi model đã được train, evaluate và deploy trên SageMaker Endpoints). Thực hiện sớm sẽ gây lỗi vì model chưa tồn tại.

  • ❌ Monitor your model performance, and make any adjustments needed.
    Sai vì: Giám sát hiệu suất (monitoring) là bước sau triển khai (post-deployment), sử dụng SageMaker Model Monitor hoặc Amazon CloudWatch để phát hiện drift/data bias. Lúc này dữ liệu mới xử lý xong, model chưa train/deploy, nên không có gì để monitor. Đây là giai đoạn Operational Excellence trong AWS ML lifecycle.

  • ✅ Delineate what data will be used for testing and what will be used for training the model.
    Đúng vì: Như đã giải thích ở trên, đây chính là bước ngay sau data processing. AWS khuyến nghị sử dụng tỷ lệ 80/20 hoặc 70/15/15 (train/validation/test) qua SageMaker Processing Jobs hoặc scikit-learn's train_test_split. Đảm bảo tính khoa học và tránh bias.

  • ❌ Test and evaluate your model on your curated data to determine how well the model performs.
    Sai vì: Đánh giá mô hình (test/evaluate) yêu cầu model đã được huấn luyện trước (trained model). Dữ liệu curated (đã xử lý) chưa được chia, và chưa train, nên không thể test. Bước này diễn ra sau train, sử dụng metrics như accuracy, precision/recall trên SageMaker Experiments.

📘 Tài liệu tham khảo (cập nhật đến 2026):

Hy vọng phân tích này giúp bạn nắm vững quy trình ML trên AWS! 🚀

Câu 346
You have one BigQuery dataset which includes customers’ street addresses. You want to retrieve all occurrences of street addresses from the dataset. What should you do?
  1. A Write a SQL query in BigQuery by using REGEXP_CONTAINS on all tables in your dataset to find rows where the word “street” appears.
  2. B Create a deep inspection job on each table in your dataset with Cloud Data Loss Prevention and create an inspection template that includes the STREET_ADDRESS infoType.
  3. C Create a discovery scan configuration on your organization with Cloud Data Loss Prevention and create an inspection template that includes the STREET_ADDRESS infoType.
  4. D Create a de-identification job in Cloud Data Loss Prevention and use the masking transformation.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc xử lý dữ liệu nhạy cảm trong BigQuery dataset trên Google Cloud Platform (GCP). Cụ thể:
Bạn có một BigQuery dataset chứa địa chỉ đường phố (street addresses) của khách hàng. Nhiệm vụ là tìm và lấy tất cả các trường hợp địa chỉ đường phố (retrieve all occurrences) từ dataset này.

📌 Yêu cầu chính: Không chỉ tìm kiếm từ khóa đơn giản mà cần phát hiện chính xác dữ liệu PII (Personally Identifiable Information) như địa chỉ đường phố một cách thông minh, sử dụng công cụ chuyên dụng để tránh lỗi (ví dụ: không phải tất cả địa chỉ đều chứa từ "street").
🛠️ Bối cảnh GCP (cập nhật đến 2026): BigQuery là data warehouse, và Cloud Data Loss Prevention (DLP) là dịch vụ chuyên phát hiện/mã hóa dữ liệu nhạy cảm. DLP hỗ trợ infoType STREET_ADDRESS để nhận diện địa chỉ đường phố qua ML detectors (không chỉ regex). Phiên bản mới nhất (DLP API v2) hỗ trợ deep inspection cho BigQuery tables để scan chi tiết và retrieve findings.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create a deep inspection job on each table in your dataset with Cloud Data Loss Prevention and create an inspection template that includes the STREET_ADDRESS infoType.

Lý do chi tiết 🏆:

  • Deep inspection job trong Cloud DLP cho phép scan sâu từng table BigQuery trong dataset, phát hiện chính xác STREET_ADDRESS infoType (dựa trên ML, nhận diện pattern như số nhà + tên đường + thành phố).
  • Inspection template lưu cấu hình (infoTypes) để tái sử dụng, đảm bảo retrieve all findings (vị trí, giá trị địa chỉ) qua job results hoặc BigQuery export.
  • Đây là cách chuẩn và hiệu quả nhất cho BigQuery (hỗ trợ wildcard tables), tránh false positives. Theo docs GCP 2026, deep inspection tối ưu cho datasets lớn, hỗ trợ stored infoTypes mới như STREET_ADDRESS với độ chính xác >95%.
    📘 Nguồn: Cloud DLP Inspect BigQuery Tables & InfoTypes Reference.

❌ Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Write a SQL query in BigQuery by using REGEXP_CONTAINS on all tables in your dataset to find rows where the word “street” appears.
    Phương án này không chính xác và không hiệu quả vì chỉ tìm từ khóa "street" qua regex, dẫn đến false positives (ví dụ: "main street" không phải địa chỉ, hoặc địa chỉ không chứa "street" như "123 Nguyễn Huệ St."). Không retrieve được pattern đầy đủ địa chỉ (số nhà + tên đường). SQL trong BigQuery không dùng ML detectors như DLP.

  • ✅ [ĐÚNG] Create a deep inspection job on each table in your dataset with Cloud Data Loss Prevention and create an inspection template that includes the STREET_ADDRESS infoType.
    (Đã giải thích ở phần trên) – Hoàn hảo cho retrieve occurrences từ BigQuery dataset cụ thể.

  • ❌ [SAI] Create a discovery scan configuration on your organization with Cloud Data Loss Prevention and create an inspection template that includes the STREET_ADDRESS infoType.
    Discovery scan dùng cho tự động quét toàn tổ chức (GCS, BigQuery projects), không target cụ thể một dataset. Nó chạy liên tục, tốn kém, và không retrieve ngay lập tức từ "each table in your dataset". Phù hợp monitor hơn là one-time retrieve.

  • ❌ [SAI] Create a de-identification job in Cloud Data Loss Prevention and use the masking transformation.
    De-identification job dùng để mã hóa/anonymize dữ liệu (mask địa chỉ thành ***), không phải retrieve occurrences (chỉ transform, không export findings). Sai mục tiêu hoàn toàn – nhiệm vụ là tìm và lấy dữ liệu, không che giấu.

🧠 Tóm tắt khuyến nghị: Sử dụng DLP deep inspection để đảm bảo compliance GDPR/CCPA với PII detection. Test job trước trên sample table!
📘 Tài liệu tham khảo thêm:

Câu 347
Your company operates in three domains: airlines, hotels, and ride-hailing services. Each domain has two teams: analytics and data science, which create data assets in BigQuery with the help of a central data platform team. However, as each domain is evolving rapidly, the central data platform team is becoming a bottleneck. This is causing delays in deriving insights from data, and resulting in stale data when pipelines are not kept up to date. You need to design a data mesh architecture by using Dataplex to eliminate the bottleneck. What should you do?
  1. A 1. Create one lake for each team. Inside each lake, create one zone for each domain.
    2. Attach each of the BigQuery datasets created by the individual teams as assets to the respective zone.
    3. Have the central data platform team manage all zones’ data assets.
  2. B 1. Create one lake for each team. Inside each lake, create one zone for each domain.
    2. Attach each of the BigQuery datasets created by the individual teams as assets to the respective zone.
    3. Direct each domain to manage their own zone’s data assets.
  3. C 1. Create one lake for each domain. Inside each lake, create one zone for each team.
    2. Attach each of the BigQuery datasets created by the individual teams as assets to the respective zone.
    3. Direct each domain to manage their own lake’s data assets.
  4. D 1. Create one lake for each domain. Inside each lake, create one zone for each team.
    2. Attach each of the BigQuery datasets created by the individual teams as assets to the respective zone.
    3. Have the central data platform team manage all lakes’ data assets.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một công ty hoạt động trong 3 lĩnh vực (domains): airlines (hàng không), hotels (khách sạn), và ride-hailing services (dịch vụ gọi xe). Mỗi domain có 2 teams: analytics (phân tích) và data science (khoa học dữ liệu). Các teams này tạo ra data assets (tài sản dữ liệu) trong BigQuery với sự hỗ trợ từ central data platform team (đội ngũ nền tảng dữ liệu trung tâm). Tuy nhiên, do các domain phát triển nhanh chóng, central team đang trở thành bottleneck (điểm nghẽn), dẫn đến chậm trễ trong việc rút ra insights từ dữ liệu và dữ liệu bị stale (lỗi thời) khi pipelines không được cập nhật kịp thời.

Mục tiêu: Thiết kế data mesh architecture sử dụng Dataplex (dịch vụ quản lý dữ liệu thống nhất của Google Cloud, hỗ trợ data mesh từ năm 2021 và cập nhật liên tục đến 2026 với các tính năng như intelligent cataloging, governance tự động) để loại bỏ bottleneck. Data mesh nhấn mạnh phân quyền sở hữu dữ liệu (decentralized data ownership), nơi mỗi domain tự quản lý dữ liệu của mình thay vì phụ thuộc central team.

📘 Nguồn tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng là phương án thứ 3:

  1. Create one lake for each domain. Inside each lake, create one zone for each team.
    2. Attach each of the BigQuery datasets created by the individual teams as assets to the respective zone.
    3. Direct each domain to manage their own lake’s data assets.

Lý do chọn đáp án này 🛠️:

  • Phù hợp nguyên tắc data mesh: Mỗi domain (airlines, hotels, ride-hailing) là đơn vị sở hữu dữ liệu độc lập → Tạo 1 lake/domain (lake đại diện cho domain data products). Trong mỗi lake, tạo 1 zone/team (zone cho analytics và data science, giúp phân chia trách nhiệm nội bộ domain).
  • Attach BigQuery datasets as assets: Các dataset từ teams được gắn vào zone tương ứng, tận dụng Dataplex để catalog, govern và discover dữ liệu mà không di chuyển dữ liệu.
  • Domains tự quản lý lakes: Loại bỏ central team làm bottleneck bằng cách direct each domain to manage their own lake’s data assets → Phân quyền hoàn toàn, teams/domain tự cập nhật pipelines, giữ dữ liệu fresh.
  • Cấu trúc này scale tốt với 3 domains × 2 teams = 6 zones, phù hợp Dataplex hierarchy (Project > Lake > Zone > Asset) theo docs mới nhất 2026.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích chi tiết từng phương án, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai) dựa trên data mesh và Dataplex best practices.

  • Phương án 1:

    1. Create one lake for each team. Inside each lake, create one zone for each domain.
      2. Attach each of the BigQuery datasets created by the individual teams as assets to the respective zone.
      3. Have the central data platform team manage all zones’ data assets.

    ❌ Sai vì: Tạo lake/team (6 lakes cho 6 teams) vi phạm nguyên tắc data mesh – domain phải là đơn vị sở hữu chính, không phải team nhỏ. Zone/domain trong lake/team là ngược logic (zone nên con của lake/domain). Central team quản lý zones → giữ nguyên bottleneck, không giải quyết vấn đề phân quyền.

  • Phương án 2:

    1. Create one lake for each team. Inside each lake, create one zone for each domain.
      2. Attach each of the BigQuery datasets created by the individual teams as assets to the respective zone.
      3. Direct each domain to manage their own zone’s data assets.

    ❌ Sai vì: Vẫn tạo lake/team (quá chi tiết, tạo 6 lakes riêng lẻ → khó govern cross-domain). Zone/domain trong lake/team không hợp lý vì domain lớn hơn team. Domains quản lý zones (không phải lakes) → phân quyền nửa vời, teams vẫn phụ thuộc, không loại bỏ central bottleneck hoàn toàn.

  • Phương án 3 (Đúng - như đã giải thích ở trên):

    1. Create one lake for each domain. Inside each lake, create one zone for each team.
      2. Attach each of the BigQuery datasets created by the individual teams as assets to the respective zone.
      3. Direct each domain to manage their own lake’s data assets.

    ✅ Đúng vì: Hierarchy chuẩn Dataplex data mesh: Lake/domain (3 lakes), Zone/team (2 zones/lake). Attach assets đúng cách. Domains tự quản lý lakes → Decentralized ownership, tự động hóa governance, insights nhanh chóng, dữ liệu fresh.

  • Phương án 4:

    1. Create one lake for each domain. Inside each lake, create one zone for each team.
      2. Attach each of the BigQuery datasets created by the individual teams as assets to the respective zone.
      3. Have the central data platform team manage all lakes’ data assets.

    ❌ Sai vì: Cấu trúc lake/zone đúng (1 lake/domain, 1 zone/team), nhưng central team quản lý tất cả lakes → Tái tạo bottleneck, domains không tự chủ, pipelines vẫn chậm và stale. Vi phạm core data mesh (domain ownership).

Kết luận 🎯: Phương án 3 là giải pháp tối ưu, giúp chuyển từ centralized sang domain-driven data mesh với Dataplex, đảm bảo scalability đến 2026!

Câu 348
dataset.inventory_vm sample records:



You have an inventory of VM data stored in the BigQuery table. You want to prepare the data for regular reporting in the most cost-effective way. You need to exclude VM rows with fewer than 8 vCPU in your report. What should you do?
  1. A Create a view with a filter to drop rows with fewer than 8 vCPU, and use the UNNEST operator.
  2. B Create a materialized view with a filter to drop rows with fewer than 8 vCPU, and use the WITH common table expression.
  3. C Create a view with a filter to drop rows with fewer than 8 vCPU, and use the WITH common table expression.
  4. D Use Dataflow to batch process and write the result to another BigQuery table.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi

Câu hỏi thuộc chủ đề BigQuery trong Google Cloud Platform (GCP) (không phải AWS như đề cập nhầm, vì liên quan đến table dataset.inventory_vm và các công cụ như View, Materialized View, Dataflow). Nội dung xoay quanh việc xử lý dữ liệu inventory của các Virtual Machine (VM) được lưu trữ trong một bảng BigQuery. Mục tiêu là chuẩn bị dữ liệu cho báo cáo định kỳ (regular reporting) một cách tiết kiệm chi phí nhất (most cost-effective way), đồng thời loại bỏ (exclude) các hàng VM có số lượng vCPU nhỏ hơn 8.

📊 Phân tích dữ liệu từ hình ảnh mẫu (sample records):

  • Bảng có các cột: row (số thứ tự), id (ID VM như vm02781), name (tên VM như dj-pk-02-02), components.name (tên thành phần: vcpu, memory, boot_disk, disk-1), components.qty (số lượng của thành phần).
  • Dữ liệu được chuẩn hóa (normalized): Mỗi VM có nhiều hàng (rows) tương ứng với các thành phần khác nhau (không phải flat table đơn giản).
    • VM 1 (id: vm02781, name: dj-pk-02-02): vCPU=2 (<8 → cần loại bỏ), memory=8, boot_disk=10, disk-1=50.
    • VM 2 (id: vm11490, name: ij-pk-02-07): vCPU=16 (≥8 → giữ), memory=64, boot_disk=10, disk-1=200.
    • VM 3 (id: vm18130, name: ij-pk-02-08): vCPU=8 (≥8 → giữ), memory=8, boot_disk=10.
  • Vấn đề cốt lõi: Để loại bỏ VM <8 vCPU, cần tập hợp (aggregate) dữ liệu theo id hoặc name, trích xuất qty của components.name = 'vcpu', sau đó filter toàn bộ VM nếu vCPU <8. Dữ liệu có thể là nested structure (ARRAY<STRUCT<components.name, components.qty>> trong schema BigQuery), yêu cầu UNNEST để flatten trước khi aggregate/filter.
  • Yêu cầu cost-effective: Báo cáo định kỳ → Tránh lưu trữ dữ liệu duplicate (như materialized view hoặc table mới), ưu tiên query on-demand.

🛠️ Cách xử lý lý tưởng: Tạo logical view (không tốn storage), sử dụng UNNEST để xử lý nested components, aggregate vCPU, filter VM ≥8 vCPU, rồi project các components khác cho báo cáo.

✅ Đáp án đúng

Create a view with a filter to drop rows with fewer than 8 vCPU, and use the UNNEST operator.

Lý do chọn (chi tiết):

  • View là giải pháp cost-effective nhất cho reporting định kỳ: Không tốn storage (chỉ metadata ~ vài KB), query on-demand chỉ tính phí scan dữ liệu gốc khi chạy report (BigQuery slot-based pricing, tối ưu với caching).
  • UNNEST operator cần thiết để flatten nested array components (ví dụ: UNNEST(components) AS comp), sau đó dùng CASE WHEN comp.name = 'vcpu' THEN comp.qty END aggregate thành vCPU per VM, filter HAVING vcpu_qty >= 8.
  • Ví dụ query mẫu trong view:
    SELECT id, name, ... -- aggregate components
    FROM dataset.inventory_vm,
    UNNEST(components) AS comp
    GROUP BY id, name
    HAVING MAX(CASE WHEN comp.name = 'vcpu' THEN comp.qty END) >= 8
    
  • Tiết kiệm: Không refresh định kỳ như materialized view, phù hợp regular reporting (theo docs BigQuery 2024-2026, views hỗ trợ clustering/partitioning inheritance).

❌ Giải thích tất cả các phương án

• Create a view with a filter to drop rows with fewer than 8 vCPU, and use the UNNEST operator.
✅ Đúng (như giải thích trên). UNNEST xử lý nested data hiệu quả, view giữ chi phí thấp nhất cho reporting lặp lại 🤑.

• Create a materialized view with a filter to drop rows with fewer than 8 vCPU, and use the WITH common table expression.
❌ Sai. Materialized view (từ BigQuery 2021, cập nhật 2026 hỗ trợ incremental refresh) tốn storage cost (duplicate data ~ GBs) và refresh compute cost (automatic hoặc manual), không cost-effective cho reporting đơn giản. WITH CTE chỉ là syntax hỗ trợ (không bắt buộc), không thay thế UNNEST cho nested data.

• Create a view with a filter to drop rows with fewer than 8 vCPU, and use the WITH common table expression.
❌ Sai. View tốt nhưng WITH CTE không giải quyết nested components (CTE chỉ refactor query, không flatten array). Thiếu UNNEST → query lỗi hoặc không aggregate đúng vCPU per VM, dẫn đến filter sai (ví dụ chỉ drop row vCPU thay vì toàn VM).

• Use Dataflow to batch process and write the result to another BigQuery table.
❌ Sai. Dataflow (Apache Beam) là ETL mạnh cho batch/large-scale, nhưng overkill và đắt đỏ: Tốn vCPU/GPU runtime, I/O, storage table mới. Không phù hợp reporting định kỳ (BigQuery native query rẻ hơn 10-100x), vi phạm "most cost-effective".

📘 Tài liệu tham khảo (cập nhật đến 2026)

💡 Kết luận: Giải pháp view + UNNEST là optimal cho GCP Data Engineer, đảm bảo scalable & cheap! 🚀

Câu 349
Your team is building a data lake platform on Google Cloud. As a part of the data foundation design, you are planning to store all the raw data in Cloud Storage. You are expecting to ingest approximately 25 GB of data a day and your billing department is worried about the increasing cost of storing old data. The current business requirements are:

•The old data can be deleted anytime.
•There is no predefined access pattern of the old data.
•The old data should be available instantly when accessed.
•There should not be any charges for data retrieval.

What should you do to optimize for cost?
  1. A Create the bucket with the Autoclass storage class feature.
  2. B Create an Object Lifecycle Management policy to modify the storage class for data older than 30 days to nearline, 90 days to coldline, and 365 days to archive storage class. Delete old data as needed.
  3. C Create an Object Lifecycle Management policy to modify the storage class for data older than 30 days to coldline, 90 days to nearline, and 365 days to archive storage class. Delete old data as needed.
  4. D Create an Object Lifecycle Management policy to modify the storage class for data older than 30 days to nearline, 45 days to coldline, and 60 days to archive storage class. Delete old data as needed.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi xoay quanh việc thiết kế nền tảng data lake trên Google Cloud, cụ thể là lưu trữ dữ liệu thô (raw data) trong Cloud Storage. Dự kiến ingest khoảng 25 GB dữ liệu mỗi ngày, và bộ phận billing lo ngại về chi phí lưu trữ dữ liệu cũ ngày càng tăng. Các yêu cầu kinh doanh chính bao gồm:

  • Dữ liệu cũ có thể bị xóa bất kỳ lúc nào (không có ràng buộc giữ lâu dài).
  • Không có pattern truy cập định trước cho dữ liệu cũ (không biết trước tần suất truy cập).
  • Dữ liệu cũ phải sẵn sàng ngay lập tức khi được truy cập (instant availability, độ trễ thấp).
  • Không có bất kỳ chi phí truy xuất dữ liệu nào (no retrieval charges).

Mục tiêu là tối ưu hóa chi phí (optimize for cost) mà vẫn đáp ứng đầy đủ các yêu cầu trên. Đây là tình huống điển hình khi cần chọn storage class thông minh trong Cloud Storage, dựa trên kiến thức cập nhật đến năm 2026 (Autoclass vẫn là feature khuyến nghị cho trường hợp access pattern không dự đoán được, theo docs Google Cloud 2024-2026).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create the bucket with the Autoclass storage class feature.

Lý do lựa chọn (bằng tiếng Việt rõ ràng):
🛠️ Autoclass là tính năng lưu trữ tự động (intelligent tiering) của Cloud Storage, sử dụng machine learning để phân tích pattern truy cập thực tế trong 30 ngày gần nhất và tự động chuyển object giữa các storage class (Standard → Nearline → Coldline → Archive) mà không phát sinh chi phí chuyển lớp (free transitions).

  • 📈 Tối ưu chi phí: Giảm storage cost cho dữ liệu ít truy cập mà không cần quản lý thủ công. Với 25 GB/ngày, chi phí sẽ thấp dần cho dữ liệu cũ.
  • ⚡ Instant availability: Luôn bắt đầu ở Standard (độ trễ mili-giây), đảm bảo truy cập ngay lập tức.
  • 💰 No retrieval charges: Không có phí truy xuất thêm cho các lần chuyển tự động; chỉ tính phí storage và retrieval chuẩn của class hiện tại (nhưng vì pattern unknown, nó tránh phí không cần thiết bằng cách giữ frequent access ở Standard).
  • 🗑️ Delete anytime: Không vi phạm minimum storage duration của các class lạnh.
    So với lifecycle management thủ công, Autoclass phù hợp nhất vì no predefined access pattern, tự động hóa hoàn toàn và cost-effective nhất theo best practices Google Cloud (cập nhật 2026).

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn. Tôi giữ nguyên nội dung văn bản gốc bằng tiếng Anh, chỉ giải thích đúng/sai bằng tiếng Việt. Các phương án lifecycle sai vì ép buộc chuyển class theo thời gian cố định, dẫn đến retrieval fees (phí truy xuất) cho Nearline/Coldline/Archive và có thể vi phạm "instant access" hoặc "no retrieval charges".

  • ✅ Create the bucket with the Autoclass storage class feature.
    🟢 Đúng hoàn toàn: Như giải thích ở trên, Autoclass tự động tối ưu dựa trên access thực tế, đáp ứng tất cả yêu cầu (cost optimization, instant access, no extra retrieval fees cho transitions, delete anytime). Đây là giải pháp best practice cho data lake với volume lớn và pattern unknown.

  • ❌ Create an Object Lifecycle Management policy to modify the storage class for data older than 30 days to nearline, 90 days to coldline, and 365 days to archive storage class. Delete old data as needed.
    🔴 Sai: Lifecycle policy này chuyển thủ công theo tuổi dữ liệu (age-based), không dựa trên access pattern thực tế → có thể chuyển dữ liệu vẫn hay truy cập sang Nearline/Coldline sớm, gây retrieval fees (Nearline: ~$0.01/GB, Coldline: cao hơn) và độ trễ truy xuất cao hơn (vi phạm "instant access" và "no retrieval charges"). Ngoài ra, minimum storage duration (30 ngày Nearline) có thể cản trở delete anytime.

  • ❌ Create an Object Lifecycle Management policy to modify the storage class for data older than 30 days to coldline, 90 days to nearline, and 365 days to archive storage class. Delete old data as needed.
    🔴 Sai: Thứ tự chuyển vô lý và không tối ưu (30 ngày → Coldline: phí retrieval rất cao ~$0.02/GB + độ trễ cao; 90 ngày → Nearline: đảo ngược logic). Không khớp "no predefined pattern" vì giả định tuổi = ít access, dẫn đến chi phí retrieval bất ngờ và vi phạm "instant access/no charges". Delete OK nhưng tổng thể kém hiệu quả.

  • ❌ Create an Object Lifecycle Management policy to modify the storage class for data older than 30 days to nearline, 45 days to coldline, and 60 days to archive storage class. Delete old data as needed.
    🔴 Sai: Thời gian chuyển quá sớm (45 ngày → Coldline, 60 ngày → Archive: phí retrieval cực cao ~$0.05/GB + thời gian retrieve 3-12 giờ cho Archive), vi phạm nghiêm trọng "instant access" và "no retrieval charges". Không phù hợp pattern unknown, dễ gây chi phí cao hơn lợi ích.

📘 Tài liệu tham khảo (cập nhật mới nhất đến 2026)

Hy vọng phân tích này giúp bạn ôn thi Google Cloud Professional Data Engineer hiệu quả! 🚀

Câu 350
Your company's data platform ingests CSV file dumps of booking and user profile data from upstream sources into Cloud Storage. The data analyst team wants to join these datasets on the email field available in both the datasets to perform analysis. However, personally identifiable information (PII) should not be accessible to the analysts. You need to de-identify the email field in both the datasets before loading them into BigQuery for analysts. What should you do?
  1. A 1. Create a pipeline to de-identify the email field by using recordTransformations in Cloud Data Loss Prevention (Cloud DLP) with masking as the de-identification transformations type.
    2. Load the booking and user profile data into a BigQuery table.
  2. B 1. Create a pipeline to de-identify the email field by using recordTransformations in Cloud DLP with format-preserving encryption with FFX as the de-identification transformation type.
    2. Load the booking and user profile data into a BigQuery table.
  3. C 1. Load the CSV files from Cloud Storage into a BigQuery table, and enable dynamic data masking.
    2. Create a policy tag with the email mask as the data masking rule.
    3. Assign the policy to the email field in both tables. A
    4. Assign the Identity and Access Management bigquerydatapolicy.maskedReader role for the BigQuery tables to the analysts.
  4. D 1. Load the CSV files from Cloud Storage into a BigQuery table, and enable dynamic data masking.
    2. Create a policy tag with the default masking value as the data masking rule.
    3. Assign the policy to the email field in both tables.
    4. Assign the Identity and Access Management bigquerydatapolicy.maskedReader role for the BigQuery tables to the analysts
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh một nền tảng dữ liệu của công ty đang ingest các file CSV chứa dữ liệu booking và user profile từ nguồn upstream vào Cloud Storage. Nhóm data analyst cần join hai dataset này dựa trên trường email (có trong cả hai dataset) để thực hiện phân tích. Tuy nhiên, thông tin cá nhân có thể nhận dạng (PII - Personally Identifiable Information) như email không được phép tiếp cận bởi analysts. Nhiệm vụ là de-identify (giấu danh tính) trường email trong cả hai dataset trước khi load vào BigQuery để analysts sử dụng.

Yêu cầu chính:

  • De-identification phải cho phép join thành công trên email (nghĩa là giá trị sau de-identify phải deterministic - input giống nhau cho output giống nhau, và giữ format để join dễ dàng).
  • Không để analysts thấy email gốc (PII).
  • Quy trình: Sử dụng Cloud DLP hoặc các công cụ GCP khác để transform dữ liệu trước khi load vào BigQuery.

📘 Kiến thức cập nhật (đến 2026): Theo tài liệu GCP mới nhất (Cloud DLP v2+, BigQuery 2024+), format-preserving encryption (FFX) là phương pháp pseudonymization lý tưởng cho trường hợp này vì giữ độ dài/format email và deterministic, cho phép join mà không lộ PII. Dynamic Data Masking (DDM) trong BigQuery chỉ mask tại query-time, không phù hợp cho pre-load de-identify.

Nguồn tham khảo:

✅ Đáp án đúng: Phương án thứ 2

Lý do chọn:

  • Sử dụng Cloud DLP recordTransformations với format-preserving encryption (FFX) là cách tối ưu để de-identify email:
    • FFX mã hóa giữ nguyên format và độ dài (ví dụ: user@example.com → usxr@fxamplx.com), deterministic (email giống nhau → mã hóa giống nhau), nên analysts có thể join chính xác trên trường email đã transform.
    • Thực hiện qua pipeline trước khi load vào BigQuery, đảm bảo PII không bao giờ vào BigQuery dưới dạng gốc.
    • Masking thông thường (như thay *** ) sẽ phá hủy khả năng join.

🛠️ Giải thích chi tiết từng phương án

  • Phương án 1 (❌ SAI):

    1. Create a pipeline to de-identify the email field by using recordTransformations in Cloud Data Loss Prevention (Cloud DLP) with masking as the de-identification transformations type.
    2. Load the booking and user profile data into a BigQuery table.
      Giải thích sai: Masking (thay thế bằng ký tự như ***) làm mất tính deterministic và format, nên không thể join hai dataset trên email masked (giá trị khác nhau dù email gốc giống). Không đáp ứng yêu cầu join mà vẫn de-identify.
  • Phương án 2 (✅ ĐÚNG):

    1. Create a pipeline to de-identify the email field by using recordTransformations in Cloud DLP with format-preserving encryption with FFX as the de-identification transformation type.
    2. Load the booking and user profile data into a BigQuery table.
      Giải thích đúng: Như đã nêu ở trên, FFX lý tưởng cho join trên PII pseudonymized, an toàn và hiệu quả trước load BigQuery.
  • Phương án 3 (❌ SAI):

    1. Load the CSV files from Cloud Storage into a BigQuery table, and enable dynamic data masking.
    2. Create a policy tag with the email mask as the data masking rule.
    3. Assign the policy to the email field in both tables. A
    4. Assign the Identity and Access Management bigquerydatapolicy.maskedReader role for the BigQuery tables to the analysts.
      Giải thích sai:
    • Load trước vào BigQuery → PII gốc vẫn tồn tại trong table, vi phạm yêu cầu "de-identify before loading".
    • Dynamic Data Masking (DDM) chỉ mask tại query-time (khi analysts query), không thay đổi dữ liệu lưu trữ. Join trên masked field có thể không chính xác nếu mask không deterministic (email mask thường random hóa).
    • Lỗi syntax "A" ở bước 3, và role bigquerydatapolicy.maskedReader chỉ áp dụng sau load.
  • Phương án 4 (❌ SAI):

    1. Load the CSV files from Cloud Storage into a BigQuery table, and enable dynamic data masking.
    2. Create a policy tag with the default masking value as the data masking rule.
    3. Assign the policy to the email field in both tables.
    4. Assign the Identity and Access Management bigquerydatapolicy.maskedReader role for the BigQuery tables to the analysts.
      Giải thích sai: Tương tự phương án 3, load trước để PII vào BigQuery, DDM chỉ mask runtime. "Default masking value" (thường là ***) không deterministic, không join được. Role maskedReader chỉ che giấu khi query, không de-identify vĩnh viễn trước load.

Kết luận 📈: Phương án 2 là giải pháp Google Cloud-native, scalable và tuân thủ best practices cho PII de-identification trong pipeline Dataflow/Cloud DLP! 🚀