Ngân hàng đề — Google Cloud Professional Data Engineer

Tìm thấy 429 câu.

Câu 261
You need to choose a database to store time series CPU and memory usage for millions of computers. You need to store this data in one-second interval samples. Analysts will be performing real-time, ad hoc analytics against the database. You want to avoid being charged for every query executed and ensure that the schema design will allow for future growth of the dataset. Which database and data model should you choose?
  1. A Create a table in BigQuery, and append the new samples for CPU and memory to the table
  2. B Create a wide table in BigQuery, create a column for the sample value at each second, and update the row with the interval for each second
  3. C Create a narrow table in Bigtable with a row key that combines the Computer Engine computer identifier with the sample time at each second
  4. D Create a wide table in Bigtable with a row key that combines the computer identifier with the sample time at each minute, and combine the values for each second as column data.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi yêu cầu chọn một cơ sở dữ liệu (database) phù hợp để lưu trữ dữ liệu time series về mức sử dụng CPU và memory của hàng triệu máy tính (millions of computers). Dữ liệu được lấy mẫu mỗi giây một lần (one-second interval samples). Các nhà phân tích sẽ thực hiện phân tích ad hoc thời gian thực (real-time, ad hoc analytics) trực tiếp trên database. Các yêu cầu chính bao gồm:

  • Tránh bị tính phí cho mỗi truy vấn (avoid being charged for every query executed) – nghĩa là không muốn chi phí phát sinh theo số lượng hoặc dung lượng truy vấn.
  • Thiết kế schema linh hoạt cho sự phát triển tương lai (schema design allows for future growth) – dữ liệu sẽ tăng lớn theo thời gian, cần hỗ trợ scale dễ dàng.

Đây là tình huống điển hình cho dữ liệu high-velocity time series với throughput cao (hàng triệu mẫu/giây), cần low-latency reads/writes cho analytics thời gian thực, và chi phí tối ưu (không scan toàn bộ dữ liệu mỗi query).

📘 Dẫn nguồn tham khảo:

✅ Đáp án đúng

Create a narrow table in Bigtable with a row key that combines the Compute Engine computer identifier with the sample time at each second

Lý do lựa chọn:

  • 🛠️ Bigtable là NoSQL wide-column store lý tưởng cho time series data với high write throughput (hàng tỷ rows/ngày) và low-latency queries (milliseconds), không tính phí theo query mà chỉ theo storage và operations (ops/second).
  • Narrow table: Mỗi mẫu dữ liệu (CPU/memory tại 1 giây) là một row riêng biệt → Tránh hotspot (nóng cục bộ) khi write/read, dễ scale horizontally.
  • Row key = computer_id + timestamp_second (ví dụ: "vm123#2024-01-01T00:00:01"): Đảm bảo unique, ordered theo thời gian và device, hỗ trợ range scans nhanh cho analytics ad hoc (lấy dữ liệu theo device/time range).
  • Phù hợp future growth: Bigtable tự động scale đến petabytes, schema flexible (dynamic columns cho CPU/memory).
  • Không bị charge per query như BigQuery.

❌ Giải thích tất cả các phương án

  • [SAI] Create a table in BigQuery, and append the new samples for CPU and memory to the table
    ❌ Sai vì: BigQuery là columnar OLAP tốt cho batch analytics lớn, nhưng tính phí theo TB scanned mỗi query (on-demand model, ~$6/TB đến 2026) → Không tránh được charge cho ad hoc real-time queries (có thể tốn kém với millions rows). Append ok cho ingestion, nhưng không optimal cho high-frequency writes (1s intervals) và real-time analytics. Schema khó scale nếu dữ liệu tăng vọt.

  • [SAI] Create a wide table in BigQuery, create a column for the sample value at each second, and update the row with the interval for each second
    ❌ Sai vì: BigQuery không hỗ trợ updates hiệu quả (immutable, dùng MERGE/UPDATE tốn kém và chậm). Wide table (column per second) vi phạm best practice columnar storage → Scan toàn bộ columns không cần thiết, phí cao hơn. Không phù hợp real-time writes (update hàng giây), schema kém linh hoạt cho growth.

  • [ĐÚNG] Create a narrow table in Bigtable with a row key that combines the Compute Engine computer identifier with the sample time at each second
    ✅ Đúng vì: Như giải thích ở phần đáp án trên. Đây là best practice của Google cho time series (xem schema design guide), đảm bảo performance cao, chi phí thấp, và scale vô hạn.

  • [SAI] Create a wide table in Bigtable with a row key that combines the computer identifier with the sample time at each minute, and combine the values for each second as column data.
    ❌ Sai vì: Wide table (60 columns/phút cho 60s) tạo hotspots khi write (tất cả seconds cùng lúc vào 1 row) → Throttle writes với millions devices. Range scans kém hiệu quả (phải đọc full row cho 1s data). Narrow table tốt hơn cho granular 1s queries và future growth (dễ thêm metrics mới mà không restructure).

🧩 Tóm tắt khuyến nghị: Bigtable narrow table là lựa chọn tối ưu cho workload này trên Google Cloud! 🚀

Câu 262
You want to archive data in Cloud Storage. Because some data is very sensitive, you want to use the `Trust No One` (TNO) approach to encrypt your data to prevent the cloud provider staff from decrypting your data. What should you do?
  1. A Use gcloud kms keys create to create a symmetric key. Then use gcloud kms encrypt to encrypt each archival file with the key and unique additional authenticated data (AAD). Use gsutil cp to upload each encrypted file to the Cloud Storage bucket, and keep the AAD outside of Google Cloud.
  2. B Use gcloud kms keys create to create a symmetric key. Then use gcloud kms encrypt to encrypt each archival file with the key. Use gsutil cp to upload each encrypted file to the Cloud Storage bucket. Manually destroy the key previously used for encryption, and rotate the key once.
  3. C Specify customer-supplied encryption key (CSEK) in the .boto configuration file. Use gsutil cp to upload each archival file to the Cloud Storage bucket. Save the CSEK in Cloud Memorystore as permanent storage of the secret.
  4. D Specify customer-supplied encryption key (CSEK) in the .boto configuration file. Use gsutil cp to upload each archival file to the Cloud Storage bucket. Save the CSEK in a different project that only the security team can access.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc lưu trữ dữ liệu nhạy cảm trong Google Cloud Storage (GCS) với cách tiếp cận "Trust No One" (TNO). TNO nghĩa là không tin tưởng bất kỳ ai, kể cả nhà cung cấp đám mây (Google), nên dữ liệu phải được mã hóa sao cho nhân viên Google không thể giải mã ngay cả khi họ có quyền truy cập vào hệ thống.

  • Yêu cầu chính: Sử dụng mã hóa client-side hoặc CSEK (Customer-Supplied Encryption Key) để khách hàng tự cung cấp khóa mã hóa thô (raw key). Google chỉ sử dụng khóa này để mã hóa/giải mã dữ liệu tạm thời trong quá trình upload/download, nhưng không lưu trữ khóa và không thể truy cập dữ liệu nếu không có khóa từ khách hàng.
  • Công cụ liên quan: gsutil để upload, .boto config cho CSEK, Cloud KMS cho CMEK (nhưng không phù hợp TNO), và cách lưu trữ khóa an toàn.
  • Bối cảnh cập nhật 2026: Theo tài liệu GCP mới nhất (GCS Encryption docs, cập nhật 2024-2026), CSEK vẫn là lựa chọn chuẩn cho TNO vì Google xóa khóa ngay sau khi xử lý object (không cache hoặc lưu). CMEK (KMS) không đạt TNO vì Google quản lý metadata key và có thể hỗ trợ export/import theo chính sách (dù customer managed).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Specify customer-supplied encryption key (CSEK) in the .boto configuration file. Use gsutil cp to upload each archival file to the Cloud Storage bucket. Save the CSEK in a different project that only the security team can access.

Lý do 🛠️:

  • CSEK đạt chuẩn TNO: Khách hàng cung cấp khóa thô (base64-encoded AES-256) qua .boto, Google chỉ dùng tạm thời để mã hóa object, không lưu khóa → nhân viên Google không thể giải mã.
  • Lưu trữ khóa an toàn: Giữ CSEK ở project khác, chỉ security team truy cập (sử dụng IAM roles như roles/secretmanager.secretAccessor), tránh rủi ro cùng project.
  • Quy trình đơn giản: gsutil cp tự động áp dụng CSEK cho từng file.
  • Hoàn hảo cho dữ liệu archive nhạy cảm, tuân thủ nguyên tắc zero-trust.

📋 Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên kiến thức GCP 2026.

  • ❌ Phương án SAI: Use gcloud kms keys create to create a symmetric key. Then use gcloud kms encrypt to encrypt each archival file with the key and unique additional authenticated data (AAD). Use gsutil cp to upload each encrypted file to the Cloud Storage bucket, and keep the AAD outside of Google Cloud.
    Lý do sai ❌: Cloud KMS tạo CMEK (Customer-Managed Encryption Key), Google vẫn quản lý key lifecycle và metadata → không đạt TNO vì Google có quyền truy cập key (dù customer tạo). Encrypt client-side rồi upload không cần thiết, và giữ AAD ngoài GCP phức tạp, không giải quyết gốc rễ (Google vẫn decrypt được nếu có key). Không khuyến nghị cho TNO strict.

  • ❌ Phương án SAI: Use gcloud kms keys create to create a symmetric key. Then use gcloud kms encrypt to encrypt each archival file with the key. Use gsutil cp to upload each encrypted file to the Cloud Storage bucket. Manually destroy the key previously used for encryption, and rotate the key once.
    Lý do sai ❌: Tương tự trên, KMS CMEK không TNO. Destroy key chỉ làm dữ liệu unrecoverable (mất vĩnh viễn), không ngăn Google decrypt trước khi destroy. Rotate key không liên quan vì dữ liệu đã encrypt với key cũ. Quy trình thủ công, dễ lỗi, không phải best practice cho archive.

  • ❌ Phương án SAI: Specify customer-supplied encryption key (CSEK) in the .boto configuration file. Use gsutil cp to upload each archival file to the Cloud Storage bucket. Save the CSEK in Cloud Memorystore as permanent storage of the secret.
    Lý do sai ❌: CSEK đúng cho TNO, nhưng lưu khóa ở Cloud Memorystore sai lầm. Memorystore (Redis) là in-memory, không permanent (dữ liệu mất khi restart), không dành cho secrets (thiếu encryption at-rest mạnh, dễ expose qua network). Nên dùng Secret Manager hoặc external vault thay vì Memorystore.

  • ✅ Phương án ĐÚNG: Specify customer-supplied encryption key (CSEK) in the .boto configuration file. Use gsutil cp to upload each archival file to the Cloud Storage bucket. Save the CSEK in a different project that only the security team can access.
    Lý do đúng ✅: Như đã giải thích ở phần đáp án. Kết hợp CSEK TNO + lưu trữ khóa isolated (project riêng, IAM strict) là best practice. Hỗ trợ scale cho nhiều file archive mà không lộ key.

Câu 263
You have data pipelines running on BigQuery, Dataflow, and Dataproc. You need to perform health checks and monitor their behavior, and then notify the team managing the pipelines if they fail. You also need to be able to work across multiple projects. Your preference is to use managed products or features of the platform. What should you do?
  1. A Export the information to Cloud Monitoring, and set up an Alerting policy
  2. B Run a Virtual Machine in Compute Engine with Airflow, and export the information to Cloud Monitoring
  3. C Export the logs to BigQuery, and set up App Engine to read that information and send emails if you find a failure in the logs
  4. D Develop an App Engine application to consume logs using GCP API calls, and send emails if you find a failure in the logs
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này thuộc lĩnh vực Google Cloud Platform (GCP), tập trung vào việc giám sát và cảnh báo cho các data pipelines chạy trên các dịch vụ BigQuery, Dataflow và Dataproc.

  • Yêu cầu chính:
    • Thực hiện health checks (kiểm tra sức khỏe) và monitor behavior (giám sát hành vi) của các pipelines.
    • Notify team (thông báo cho đội ngũ quản lý) nếu pipelines fail (thất bại).
    • Hỗ trợ làm việc across multiple projects (qua nhiều dự án GCP).
    • Ưu tiên sử dụng managed products hoặc features của nền tảng (sản phẩm được quản lý tự động, không cần tự vận hành).

📘 Bối cảnh: Các dịch vụ BigQuery, Dataflow, Dataproc đều tích hợp sẵn metrics (chỉ số đo lường) và logs vào Cloud Monitoring (trước đây là Stackdriver Monitoring), cho phép giám sát thống nhất mà không cần công cụ bên thứ ba. Đây là giải pháp managed, hỗ trợ multi-project qua Cloud Monitoring workspaces. Kiến thức cập nhật đến năm 2026 vẫn giữ nguyên tính năng cốt lõi này (theo GCP docs phiên bản mới nhất).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Export the information to Cloud Monitoring, and set up an Alerting policy.

Lý do 🛠️:

  • Cloud Monitoring là dịch vụ managed hoàn toàn của GCP, tự động thu thập metrics/logs từ BigQuery, Dataflow, Dataproc (ví dụ: job status, failure rates, latency).
  • Export information (xuất metrics/logs) vào Cloud Monitoring qua managed export (như Cloud Monitoring Metrics API hoặc Logging sinks).
  • Alerting policy cho phép thiết lập quy tắc cảnh báo dựa trên thresholds (ngưỡng), tự động notify qua email/SMS/Slack/Pub/Sub cross-project (qua shared workspaces).
  • Hoàn hảo khớp yêu cầu: Managed, multi-project, health checks (qua uptime checks/SLA metrics), notify tự động. Không cần code custom.

📋 Giải thích tất cả các phương án (đúng & sai)

  • Export the information to Cloud Monitoring, and set up an Alerting policy
    ✅ Đúng 🟢: Như phân tích trên, đây là giải pháp managed native của GCP. Metrics từ Dataflow/Dataproc/BigQuery được export tự động vào Cloud Monitoring. Alerting policy hỗ trợ conditions phức tạp (MQL queries), notify multi-channel, và workspaces multi-project. Tiết kiệm chi phí, scalable đến 2026.

  • Run a Virtual Machine in Compute Engine with Airflow, and export the information to Cloud Monitoring
    ❌ Sai 🔴: Không ưu tiên managed – phải tự chạy VM Compute Engine với Airflow (self-managed orchestration), tốn công quản lý patching/security/scaling. Airflow không native GCP cho pipelines này (GCP recommend Dataflow/Cloud Composer). Export sau đó là thừa, vì Cloud Monitoring đã managed sẵn.

  • Export the logs to BigQuery, and set up App Engine to read that information and send emails if you find a failure in the logs
    ❌ Sai 🔴: Phức tạp và không managed cho monitoring. Export logs to BigQuery (qua Logging sinks) ok, nhưng phải tự build App Engine để query logs (SQL scans tốn kém), detect failure, gửi email – không scalable, không real-time, khó cross-project. Cloud Monitoring làm tốt hơn mà không code.

  • Develop an App Engine application to consume logs using GCP API calls, and send emails if you find a failure in the logs
    ❌ Sai 🔴: Hoàn toàn custom development trên App Engine, dùng Cloud Logging API để pull logs – tốn developer time, không reliable (polling API có rate limits), khó maintain cross-project. Không tận dụng managed alerting/metrics, vi phạm ưu tiên "managed products".

📚 Tài liệu tham khảo (cập nhật 2026)

Hy vọng phân tích này giúp bạn ôn thi Google Cloud Professional Data Engineer hiệu quả! 🚀 Nếu cần thêm chi tiết, hỏi nhé!

Câu 264
You are working on a linear regression model on BigQuery ML to predict a customer's likelihood of purchasing your company's products. Your model uses a city name variable as a key predictive component. In order to train and serve the model, your data must be organized in columns. You want to prepare your data using the least amount of coding while maintaining the predictable variables. What should you do?
  1. A Create a new view with BigQuery that does not include a column with city information.
  2. B Use SQL in BigQuery to transform the state column using a one-hot encoding method, and make each city a column with binary values.
  3. C Use TensorFlow to create a categorical variable with a vocabulary list. Create the vocabulary file and upload that as part of your model to BigQuery ML.
  4. D Use Cloud Data Fusion to assign each city to a region that is labeled as 1, 2, 3, 4, or 5, and then use that number to represent the city in the model.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh việc chuẩn bị dữ liệu cho mô hình hồi quy tuyến tính (linear regression) trên BigQuery ML để dự đoán khả năng khách hàng mua sản phẩm, với biến city name (tên thành phố) là thành phần dự đoán chính (key predictive component). 📊

  • Bối cảnh chính: Dữ liệu cần được tổ chức dưới dạng cột (columns) để huấn luyện (train) và phục vụ (serve) mô hình. Mục tiêu là chuẩn bị dữ liệu với ít code nhất (least amount of coding), đồng thời giữ nguyên các biến dự đoán (predictable variables) – đặc biệt là thông tin thành phố.
  • Thách thức: Biến city name là categorical (phân loại), không thể dùng trực tiếp dưới dạng chuỗi trong linear regression vì mô hình yêu cầu dữ liệu số hoặc binary. Cần chuyển đổi (encode) nó một cách hiệu quả trên BigQuery (Google Cloud's data warehouse).
  • Yêu cầu cốt lõi: Sử dụng công cụ GCP để xử lý nhanh chóng, không phức tạp, tận dụng SQL native của BigQuery. ✅ (Dựa trên BigQuery ML phiên bản mới nhất 2024-2026, hỗ trợ tự động hóa encoding categorical nhưng manual one-hot qua SQL vẫn là cách kiểm soát tốt cho linear models).

✅ Đáp án đúng

Use SQL in BigQuery to transform the state column using a one-hot encoding method, and make each city a column with binary values.

Lý do lựa chọn:

  • 🛠️ Đây là cách ít code nhất vì chỉ dùng SQL thuần túy trong BigQuery (không cần tool ngoài), tạo ra các cột binary (0/1) cho từng thành phố – phù hợp hoàn hảo cho linear regression trên BigQuery ML.
  • 📈 One-hot encoding giữ nguyên thông tin dự đoán của city (không mất dữ liệu), tránh multicollinearity, và dữ liệu đã sẵn sàng dưới dạng columns cho CREATE MODEL.
  • 🎯 BigQuery hỗ trợ pivot/group by để one-hot nhanh chóng (ví dụ: CASE WHEN hoặc ML.ONE_HOT_ENCODE), tối ưu cho scale lớn mà không cần Python/TensorFlow.
  • Cập nhật 2026: BigQuery ML vẫn khuyến nghị SQL preprocessing cho categorical cao cardinality như city names để control features (theo docs GCP).

📋 Phân tích tất cả các phương án

  • ❌ [SAI] Create a new view with BigQuery that does not include a column with city information.
    ❌ Sai vì: Loại bỏ hoàn toàn cột city sẽ mất key predictive component, làm mô hình kém chính xác. View chỉ reorganize data nhưng không giải quyết encoding, vi phạm yêu cầu "maintaining the predictable variables". Không cần thiết và counterproductive. 🗑️

  • ✅ [ĐÚNG] Use SQL in BigQuery to transform the state column using a one-hot encoding method, and make each city a column with binary values.
    ✅ Đúng vì: Như giải thích ở trên – least coding (SQL native), tổ chức dữ liệu thành columns binary chuẩn cho BigQuery ML linear regression. Giữ nguyên sức mạnh dự đoán của city, dễ train (CREATE OR REPLACE MODEL ...). Hoàn hảo cho GCP workflow. 🚀

  • ❌ [SAI] Use TensorFlow to create a categorical variable with a vocabulary list. Create the vocabulary file and upload that as part of your model to BigQuery ML.
    ❌ Sai vì: TensorFlow yêu cầu code phức tạp (Python/ notebooks), tạo file vocab rồi upload – không least coding. BigQuery ML không cần import model TF cho linear reg đơn giản; nó hỗ trợ native categorical. Lãng phí effort và không scale tốt trên BigQuery. 🤖

  • ❌ [SAI] Use Cloud Data Fusion to assign each city to a region that is labeled as 1, 2, 3, 4, or 5, and then use that number to represent the city in the model.
    ❌ Sai vì: Cloud Data Fusion là ETL tool phức tạp (pipeline building, nhiều config), không least coding. Label arbitrary (1-5) gây loss of information (nhiều city map chung 1 số), dẫn đến bias trong linear regression. Không tận dụng BigQuery SQL trực tiếp. 🔄

📘 Tài liệu tham khảo

Câu 265
You work for a large bank that operates in locations throughout North America. You are setting up a data storage system that will handle bank account transactions. You require ACID compliance and the ability to access data with SQL. Which solution is appropriate?
  1. A Store transaction data in Cloud Spanner. Enable stale reads to reduce latency.
  2. B Store transaction in Cloud Spanner. Use locking read-write transactions.
  3. C Store transaction data in BigQuery. Disabled the query cache to ensure consistency.
  4. D Store transaction data in Cloud SQL. Use a federated query BigQuery for analysis.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm

📖 Giải thích nội dung câu hỏi:
Câu hỏi mô tả tình huống bạn làm việc cho một ngân hàng lớn hoạt động khắp Bắc Mỹ, cần thiết lập hệ thống lưu trữ dữ liệu xử lý giao dịch tài khoản ngân hàng. Yêu cầu chính là ACID compliance (Atomicity - Nguyên tử, Consistency - Nhất quán, Isolation - Cách ly, Durability - Bền vững) để đảm bảo tính toàn vẹn dữ liệu giao dịch tài chính, đồng thời hỗ trợ truy vấn bằng SQL. Đây là nhu cầu điển hình cho cơ sở dữ liệu giao dịch (OLTP - Online Transaction Processing) ở quy mô lớn, phân tán địa lý (North America), đòi hỏi tính nhất quán mạnh mẽ (strong consistency) và khả năng mở rộng toàn cầu mà không mất dữ liệu. 🏦💳

✅ Đáp án đúng:
Store transaction in Cloud Spanner. Use locking read-write transactions.
Lý do lựa chọn: Cloud Spanner là cơ sở dữ liệu quan hệ phân tán toàn cầu của Google Cloud, hỗ trợ đầy đủ ACID transactions với tính nhất quán mạnh mẽ (external consistency). Sử dụng locking read-write transactions đảm bảo các giao dịch đọc-ghi được khóa để tránh xung đột, phù hợp hoàn hảo cho giao dịch ngân hàng yêu cầu độ tin cậy cao và truy vấn SQL chuẩn. Đây là lựa chọn tối ưu theo tài liệu GCP mới nhất (2024-2026), nơi Spanner được khuyến nghị cho workload tài chính global. 🛡️

🛠️ Giải thích tất cả các phương án (đúng và sai)

  • Store transaction data in Cloud Spanner. Enable stale reads to reduce latency.
    ❌ Sai: Mặc dù Cloud Spanner hỗ trợ ACID và SQL, nhưng stale reads (đọc dữ liệu cũ) chỉ cung cấp tính nhất quán cuối cùng (eventual consistency), không đảm bảo strong consistency cần thiết cho giao dịch ngân hàng (có thể đọc dữ liệu chưa cập nhật, dẫn đến lỗi tài chính). Điều này vi phạm yêu cầu ACID đầy đủ. 📉

  • Store transaction in Cloud Spanner. Use locking read-write transactions.
    ✅ Đúng: Như đã giải thích ở trên, phương án này tận dụng đúng tính năng locking transactions của Spanner để đảm bảo ACID với isolation và consistency mạnh mẽ, hỗ trợ SQL, và mở rộng toàn cầu mà không cần sharding thủ công. Hoàn hảo cho workload ngân hàng. 🌟

  • Store transaction data in BigQuery. Disabled the query cache to ensure consistency.
    ❌ Sai: BigQuery là kho dữ liệu phân tích (data warehouse) cho OLAP (Online Analytical Processing), không hỗ trợ ACID transactions thực sự (chỉ append-only, không update/delete atomic). Tắt query cache chỉ cải thiện freshness cho query, nhưng không giải quyết vấn đề thiếu transactional support và SQL OLTP. Không phù hợp cho giao dịch thời gian thực. 🚫

  • Store transaction data in Cloud SQL. Use a federated query BigQuery for analysis.
    ❌ Sai: Cloud SQL (MySQL/PostgreSQL) hỗ trợ ACID và SQL, nhưng chỉ phù hợp cho single-region hoặc regional replication, không phải global distribution như Spanner (dễ gặp latency cao và downtime khi scale North America). Federated query với BigQuery chỉ dùng cho phân tích, không giải quyết core storage transactional global. 🗺️

📘 Tài liệu tham khảo (cập nhật mới nhất GCP đến 2026):

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀

Câu 266
A shipping company has live package-tracking data that is sent to an Apache Kafka stream in real time. This is then loaded into BigQuery. Analysts in your company want to query the tracking data in BigQuery to analyze geospatial trends in the lifecycle of a package. The table was originally created with ingest-date partitioning. Over time, the query processing time has increased. You need to implement a change that would improve query performance in BigQuery. What should you do?
  1. A Implement clustering in BigQuery on the ingest date column.
  2. B Implement clustering in BigQuery on the package-tracking ID column.
  3. C Tier older data onto Cloud Storage files and create a BigQuery table using Cloud Storage as an external data source.
  4. D Re-create the table using data partitioning on the package delivery date.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một công ty vận chuyển có dữ liệu theo dõi gói hàng thời gian thực (live package-tracking data) được gửi vào Apache Kafka stream, sau đó tải vào BigQuery. Các nhà phân tích muốn truy vấn dữ liệu theo dõi trong BigQuery để phân tích xu hướng địa lý không gian (geospatial trends) trong vòng đời của gói hàng (lifecycle of a package).

Bảng dữ liệu ban đầu được tạo với partitioning theo ingest-date (ngày ingest dữ liệu). Theo thời gian, thời gian xử lý truy vấn tăng lên. Nhiệm vụ là thực hiện thay đổi để cải thiện hiệu suất truy vấn BigQuery.

🔍 Vấn đề cốt lõi:

  • Dữ liệu real-time từ Kafka → BigQuery, partitioning theo ingest-date giúp tối ưu query theo thời gian ingest.
  • Nhưng truy vấn tập trung vào xu hướng địa lý theo gói hàng cụ thể (geospatial trends per package lifecycle) → thường filter/group theo package-tracking ID, không chỉ theo date.
  • Partitioning theo date đã có, nhưng query scan nhiều partition không cần thiết → cần clustering để tối ưu sort/filter theo cột thường dùng như package ID.

🛠️ Giải pháp lý tưởng: Sử dụng clustering trên cột liên quan đến truy vấn chính (package ID) để giảm dữ liệu scan, cải thiện performance mà không cần recreate table.

📘 Kiến thức cập nhật (BigQuery phiên bản mới nhất 2024-2026): BigQuery hỗ trợ partitioned + clustered tables (dual optimization). Clustering tự động reorganize dữ liệu trong partition theo cột cluster key, giảm I/O lên đến 90% cho query filter/sort theo key đó. Không liên quan AWS (câu hỏi thuần GCP/BigQuery).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Implement clustering in BigQuery on the package-tracking ID column.

Lý do 🏆:

  • Truy vấn phân tích lifecycle của package (geospatial trends) thường filter/group theo package-tracking ID (ví dụ: WHERE package_id = 'XYZ' hoặc GROUP BY package_id).
  • Partitioning ingest-date đã có, nhưng thêm clustering trên package ID sẽ tối ưu sort/filter trong mỗi partition, giảm dữ liệu scan đáng kể (BigQuery tự động cluster data khi insert/query).
  • Không cần recreate table → nhanh, tiết kiệm chi phí. Hiệu suất cải thiện ngay lập tức cho query theo ID.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ Phương án SAI: Implement clustering in BigQuery on the ingest date column.
    Giải thích: Clustering trên ingest-date vô ích vì bảng đã partition theo ingest-date rồi (clustering trùng lặp không mang lợi ích). Query geospatial theo package lifecycle không chủ yếu filter theo ingest-date, nên không giảm scan data hiệu quả.

  • ✅ Phương án ĐÚNG: Implement clustering in BigQuery on the package-tracking ID column.
    Giải thích: Như trên, clustering trên package-tracking ID (cột chính cho truy vấn lifecycle/package-specific) tối ưu hóa chính xác filter/group theo ID trong mỗi partition date. BigQuery cluster tự động, cải thiện performance query lên đến 10x cho pattern này.

  • ❌ Phương án SAI: Tier older data onto Cloud Storage files and create a BigQuery table using Cloud Storage as an external data source.
    Giải thích: Chuyển data cũ sang external table (Cloud Storage) làm query chậm hơn (external tables scan toàn bộ file, không partition/cluster tự động như native BigQuery). Không phù hợp real-time tracking, tăng latency và chi phí query.

  • ❌ Phương án SAI: Re-create the table using data partitioning on the package delivery date.
    Giải thích: Recreate table tốn kém (downtime, ETL lại toàn bộ data). Partition theo delivery date (ngày giao hàng cuối) không phù hợp vì data là tracking real-time (nhiều record per package theo thời gian, ingest-date gần sát tracking time hơn). Query geospatial lifecycle vẫn scan nhiều partition nếu không filter date chính xác.

📚 Tài liệu tham khảo

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần ví dụ SQL cụ thể, hỏi thêm nhé!

Câu 267
Your company currently runs a large on-premises cluster using Spark, Hive, and HDFS in a colocation facility. The cluster is designed to accommodate peak usage on the system; however, many jobs are batch in nature, and usage of the cluster fluctuates quite dramatically. Your company is eager to move to the cloud to reduce the overhead associated with on-premises infrastructure and maintenance and to benefit from the cost savings. They are also hoping to modernize their existing infrastructure to use more serverless offerings in order to take advantage of the cloud. Because of the timing of their contract renewal with the colocation facility, they have only 2 months for their initial migration. How would you recommend they approach their upcoming migration strategy so they can maximize their cost savings in the cloud while still executing the migration in time?
  1. A Migrate the workloads to Dataproc plus HDFS; modernize later.
  2. B Migrate the workloads to Dataproc plus Cloud Storage; modernize later.
  3. C Migrate the Spark workload to Dataproc plus HDFS, and modernize the Hive workload for BigQuery.
  4. D Modernize the Spark workload for Dataflow and the Hive workload for BigQuery.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống một công ty đang chạy cluster lớn on-premises sử dụng Spark (xử lý dữ liệu lớn thời gian thực/batch), Hive (query engine cho dữ liệu lớn), và HDFS (hệ thống lưu trữ phân tán) tại cơ sở colocation. Cluster được thiết kế cho peak usage, nhưng hầu hết job là batch nên usage biến động mạnh. Công ty muốn di chuyển lên cloud để:

  • Giảm chi phí vận hành hạ tầng on-premises.
  • Tiết kiệm chi phí nhờ cloud (pay-per-use).
  • Hiện đại hóa sang serverless để tận dụng lợi ích cloud (scale tự động, không quản lý server).

Hạn chế thời gian: Chỉ 2 tháng do hợp đồng colocation sắp hết hạn. Mục tiêu: Chiến lược migration tối ưu cost savings (tiết kiệm chi phí cao nhất) + hoàn thành đúng hạn.

🛠️ Yêu cầu chiến lược lý tưởng:

  • Nhanh chóng (lift-and-shift để migrate kịp 2 tháng).
  • Tiết kiệm chi phí: Sử dụng serverless/managed services, tránh over-provisioning như on-prem.
  • Tương thích: Spark/Hive chạy tốt trên managed Hadoop ecosystem của cloud.
  • Serverless hướng: Thay HDFS bằng object storage rẻ tiền, scale vô hạn.

📘 Bối cảnh kiến thức GCP (cập nhật 2026): Google Cloud khuyến nghị Dataproc (managed Spark/Hive/Hadoop) kết hợp Cloud Storage (thay HDFS) cho migration nhanh từ on-prem Hadoop. Điều này cho phép pay-per-use, auto-scale, và dễ modernize sau (ví dụ: Dataflow cho Spark streaming/batch, BigQuery cho Hive queries).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Migrate the workloads to Dataproc plus Cloud Storage; modernize later.

Lý do 🏆:

  • Migrate nhanh (lift-and-shift): Dataproc hỗ trợ Spark + Hive trực tiếp, tương thích 100% code on-prem (chỉ thay HDFS config bằng Cloud Storage connector). Hoàn thành trong 2 tháng dễ dàng.
  • Tiết kiệm chi phí tối đa 💰: Cloud Storage rẻ hơn HDFS (không cần cluster storage riêng, object storage pay-per-use, durable cao). Dataproc serverless mode (2023+) auto-scale theo job, tránh idle cost – phù hợp usage biến động.
  • Serverless hóa dần: Giữ nguyên workload, modernize sau (Dataflow/BigQuery) mà không rush.
  • Phù hợp thời gian gấp: Không refactor code lớn.

🧪 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Migrate the workloads to Dataproc plus HDFS; modernize later.
    Phân tích: Phương án này giữ HDFS (on-prem storage) trên Dataproc, không tận dụng cloud-native. HDFS yêu cầu quản lý cluster storage, tốn kém (over-provision), không serverless, và không tiết kiệm chi phí so với on-prem. Dataproc khuyến nghị Cloud Storage thay thế để scale rẻ. Không tối ưu cost savings.

  • ✅ [ĐÚNG] Migrate the workloads to Dataproc plus Cloud Storage; modernize later.
    Phân tích: Như giải thích trên – lift-and-shift hoàn hảo. Cloud Storage (gs://) thay HDFS seamless (qua Hadoop connector), rẻ hơn 70-80% storage cost, auto-scale, multi-region. Dataproc workflows chạy Spark/Hive/HDFS commands y chang on-prem. Dễ modernize sau (Dataproc Serverless 2024+ hỗ trợ batch jobs auto-terminate).

  • ❌ [SAI] Migrate the Spark workload to Dataproc plus HDFS, and modernize the Hive workload for BigQuery.
    Phân tích: Phân tách workload phức tạp: Spark giữ HDFS (vẫn tốn kém, không serverless), Hive refactor sang BigQuery (cần ETL schema change, query rewrite – quá lâu >2 tháng). Không đồng bộ, tăng rủi ro migration, không maximize cost savings vì HDFS vẫn đắt.

  • ❌ [SAI] Modernize the Spark workload for Dataflow and the Hive workload for BigQuery.
    Phân tích: Modernize ngay lập tức (refactor Spark sang Dataflow Apache Beam, Hive sang BigQuery SQL) yêu cầu thay đổi code lớn, test ETL pipelines – không kịp 2 tháng (thường mất 6-12 tháng). Vi phạm deadline colocation, rủi ro downtime cao. Không phải "initial migration" nhanh.

📚 Tài liệu tham khảo (cập nhật GCP 2026)

🛡️ Khuyến nghị bổ sung: Sử dụng Dataproc Workflows + Cloud Composer orchestrate migration, và Transfer Service copy data từ HDFS sang Cloud Storage trong <1 tuần!

Câu 268
You work for a financial institution that lets customers register online. As new customers register, their user data is sent to Pub/Sub before being ingested into
BigQuery. For security reasons, you decide to redact your customers' Government issued Identification Number while allowing customer service representatives to view the original values when necessary. What should you do?
  1. A Use BigQuery's built-in AEAD encryption to encrypt the SSN column. Save the keys to a new table that is only viewable by permissioned users.
  2. B Use BigQuery column-level security. Set the table permissions so that only members of the Customer Service user group can see the SSN column.
  3. C Before loading the data into BigQuery, use Cloud Data Loss Prevention (DLP) to replace input values with a cryptographic hash.
  4. D Before loading the data into BigQuery, use Cloud Data Loss Prevention (DLP) to replace input values with a cryptographic format-preserving encryption token.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong hệ thống tài chính sử dụng Google Cloud Platform (GCP):

  • Khách hàng đăng ký trực tuyến, dữ liệu người dùng (bao gồm Government issued Identification Number, ví dụ như SSN - Social Security Number) được gửi đến Pub/Sub trước khi nạp vào BigQuery.
  • Yêu cầu bảo mật: Redact (che giấu/mã hóa) số ID này để bảo vệ dữ liệu, nhưng vẫn cho phép nhân viên dịch vụ khách hàng (customer service representatives) xem giá trị gốc khi cần thiết.
  • Mục tiêu: Tìm giải pháp tốt nhất để xử lý trước khi dữ liệu vào BigQuery, đảm bảo tính bảo mật cao, khả năng khôi phục có kiểm soát, và tuân thủ các quy định bảo mật dữ liệu (như GDPR hoặc PCI-DSS).

🛠️ Bối cảnh kỹ thuật (cập nhật đến 2026):

  • Pub/Sub là dịch vụ messaging để stream dữ liệu thời gian thực.
  • BigQuery là data warehouse serverless, hỗ trợ bảo mật tại mức hàng/cột qua Authorized Views và Column-level security (mới hơn từ 2023+), nhưng không lý tưởng cho mã hóa động.
  • Cloud DLP (Data Loss Prevention) là công cụ mạnh mẽ để de-identify dữ liệu, hỗ trợ nhiều phương pháp như hashing, masking, và format-preserving encryption (FPE) – phương pháp giữ nguyên định dạng (ví dụ: số ID vẫn là chuỗi số) và có thể decrypt với key phù hợp.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Before loading the data into BigQuery, use Cloud Data Loss Prevention (DLP) to replace input values with a cryptographic format-preserving encryption token.

Lý do chi tiết:

  • Phương pháp này sử dụng Cloud DLP để áp dụng format-preserving encryption (FPE) trước khi nạp dữ liệu vào BigQuery (từ Pub/Sub stream). FPE mã hóa giá trị gốc thành token giữ nguyên định dạng (ví dụ: "123-45-6789" thành token tương tự), không thể đoán ngược mà không có key.
  • Nhân viên dịch vụ khách hàng có thể decrypt token về giá trị gốc khi cần qua DLP API với quyền truy cập key (quản lý qua Cloud KMS).
  • ✅ Ưu điểm: Bảo mật cao (tuân thủ zero-trust), hiệu suất tốt cho BigQuery query (giữ format queryable), và xử lý trước ingestion tránh lưu trữ plaintext. Đây là best practice theo GCP 2025+ cho PII (Personally Identifiable Information).

❌ Giải thích tất cả các phương án (đúng/sai)

  • [SAI] Use BigQuery's built-in AEAD encryption to encrypt the SSN column. Save the keys to a new table that is only viewable by permissioned users.
    ❌ Lý do sai: BigQuery không có built-in AEAD encryption per column (AEAD là Authenticated Encryption with Associated Data, thường dùng trong client-side). BigQuery hỗ trợ CMEK (Customer-Managed Encryption Keys) hoặc CSEK cho toàn table/dataset, không phải column-level encryption tự động. Lưu key vào table mới là rủi ro cao (dễ bị lộ nếu table bị hack), vi phạm nguyên tắc key management (key phải ở KMS, không lưu plaintext). Không phù hợp với yêu cầu redact động từ Pub/Sub.

  • [SAI] Use BigQuery column-level security. Set the table permissions so that only members of the Customer Service user group can see the SSN column.
    ❌ Lý do sai: BigQuery có column-level lineage và masking (từ 2023+ qua Dynamic Data Masking), nhưng không hỗ trợ column-level permissions thực sự (chỉ row-level filters qua Authorized Views hoặc IAM policies). Không thể "ẩn cột" hoàn toàn cho nhóm user mà vẫn query table; dễ bị bypass qua export hoặc views. Không redact dữ liệu gốc, chỉ kiểm soát view – không đáp ứng "redact while allowing view original when necessary".

  • [SAI] Before loading the data into BigQuery, use Cloud Data Loss Prevention (DLP) to replace input values with a cryptographic hash.
    ❌ Lý do sai: Hashing (như SHA-256) là one-way function – không thể khôi phục giá trị gốc (irreversible). Dù DLP hỗ trợ hashing tốt cho de-identification, nó không cho phép customer service xem original values. Hash thay đổi format (chuỗi hex dài), làm khó query trong BigQuery. Không khớp yêu cầu "view original when necessary".

  • [ĐÚNG] Before loading the data into BigQuery, use Cloud Data Loss Prevention (DLP) to replace input values with a cryptographic format-preserving encryption token.
    ✅ Lý do đúng (như phần trên): Xử lý trước BigQuery qua DLP API trên Pub/Sub dataflow, FPE giữ format + reversible với key, lý tưởng cho bảo mật có kiểm soát. Hỗ trợ scale lớn, tích hợp KMS cho key rotation (2025+ features).

🛠️ Khuyến nghị triển khai: Sử dụng Dataflow để process Pub/Sub → DLP FPE → BigQuery, với IAM roles giới hạn decrypt cho customer service group. Test với DLP primitives như FFX-FP cho số ID.

Câu 269
You are migrating a table to BigQuery and are deciding on the data model. Your table stores information related to purchases made across several store locations and includes information like the time of the transaction, items purchased, the store ID, and the city and state in which the store is located. You frequently query this table to see how many of each item were sold over the past 30 days and to look at purchasing trends by state, city, and individual store. How would you model this table for the best query performance?
  1. A Partition by transaction time; cluster by state first, then city, then store ID.
  2. B Partition by transaction time; cluster by store ID first, then city, then state.
  3. C Top-level cluster by state first, then city, then store ID.
  4. D Top-level cluster by store ID first, then city, then state.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc mô hình hóa dữ liệu (data modeling) cho một bảng trong BigQuery (Google Cloud) khi di chuyển dữ liệu từ nguồn khác. Bảng lưu trữ thông tin giao dịch mua sắm tại nhiều cửa hàng, bao gồm: thời gian giao dịch (transaction time), mặt hàng mua, ID cửa hàng (store ID), thành phố (city) và bang (state) của cửa hàng.

Các truy vấn phổ biến:

  • Đếm số lượng từng mặt hàng bán ra trong 30 ngày qua (liên quan đến thời gian).
  • Phân tích xu hướng mua sắm theo bang (state), thành phố (city), và từng cửa hàng riêng lẻ (store ID).

Mục tiêu: Tối ưu hiệu suất truy vấn (query performance) bằng cách chọn partitioning (phân vùng) và clustering (nhóm dữ liệu) phù hợp.
📘 Kiến thức cốt lõi: Trong BigQuery (cập nhật đến 2024-2026), partitioning theo thời gian giúp loại bỏ nhanh các phân vùng không liên quan (partition pruning). Clustering sắp xếp dữ liệu trong phân vùng theo thứ tự selectivity giảm dần (từ rộng đến hẹp), hỗ trợ cluster pruning cho filter hiệu quả. Không partition chỉ dùng clustering sẽ kém hơn cho query thời gian lớn.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Partition by transaction time; cluster by state first, then city, then store ID.

🛠️ Lý do chi tiết:

  • Partition by transaction time: Hoàn hảo cho query "past 30 days" vì BigQuery tự động prune các partition cũ, giảm dữ liệu scan đáng kể (hàng triệu rows chỉ scan vài partition).
  • Cluster by state → city → store ID: Thứ tự lý tưởng vì query thường filter state (rộng nhất, ít giá trị unique) trước, rồi city, cuối cùng store ID (hẹp nhất, nhiều unique). Điều này tận dụng multi-level clustering (BigQuery hỗ trợ lên đến 4 cấp từ 2023+), pruning cluster hiệu quả, giảm chi phí và thời gian query.
    Kết quả: Query performance tốt nhất cho workload mô tả.

📘 Nguồn tham khảo:

❌ Phân tích tất cả các phương án

Dưới đây là phân tích từng lựa chọn giữ nguyên văn bản gốc tiếng Anh, đánh dấu đúng/sai và giải thích chi tiết bằng tiếng Việt:

  • ✅ [ĐÚNG] Partition by transaction time; cluster by state first, then city, then store ID.
    🟢 Đúng vì: Kết hợp partitioning thời gian (prune 30 days) + clustering theo thứ tự selectivity giảm dần (state rộng → store hẹp). Tối ưu nhất cho query mô tả, giảm scan dữ liệu lên đến 90%+ theo docs BigQuery.

  • ❌ [SAI] Partition by transaction time; cluster by store ID first, then city, then state.
    🔴 Sai vì: Partition thời gian tốt, nhưng clustering ngược thứ tự (store ID high-cardinality đầu tiên) làm pruning kém hiệu quả. Query filter state/city sẽ scan nhiều cluster hơn, tăng latency và chi phí so với sắp xếp state trước.

  • ❌ [SAI] Top-level cluster by state first, then city, then store ID.
    🔴 Sai vì: Chỉ clustering "top-level" (không partition) bỏ lỡ partition pruning cho transaction time – yếu tố chính của query 30 days. Với bảng lớn, scan toàn bộ dữ liệu chậm hơn nhiều, không tận dụng ingestion-time partitioning (miễn phí từ 2022+).

  • ❌ [SAI] Top-level cluster by store ID first, then city, then state.
    🔴 Sai vì: Tệ nhất – không partition + clustering ngược (store ID trước) làm pruning kém cho cả time, state/city. Query sẽ scan gần như toàn bảng, vi phạm best practices BigQuery về hybrid partitioning/clustering.

🧩 Kết luận: Lựa chọn đúng tận dụng partitioning + clustering đa cấp theo nguyên tắc "time first, geography hierarchical" – chuẩn mực cho analytics workload như vậy! 🚀

Câu 270
You are updating the code for a subscriber to a Pub/Sub feed. You are concerned that upon deployment the subscriber may erroneously acknowledge messages, leading to message loss. Your subscriber is not set up to retain acknowledged messages. What should you do to ensure that you can recover from errors after deployment?
  1. A Set up the Pub/Sub emulator on your local machine. Validate the behavior of your new subscriber logic before deploying it to production.
  2. B Create a Pub/Sub snapshot before deploying new subscriber code. Use a Seek operation to re-deliver messages that became available after the snapshot was created.
  3. C Use Cloud Build for your deployment. If an error occurs after deployment, use a Seek operation to locate a timestamp logged by Cloud Build at the start of the deployment.
  4. D Enable dead-lettering on the Pub/Sub topic to capture messages that aren't successfully acknowledged. If an error occurs after deployment, re-deliver any messages captured by the dead-letter queue.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh tình huống cập nhật code cho một subscriber (người nhận) của Pub/Sub feed trong Google Cloud Pub/Sub. Người dùng lo ngại rằng sau khi deploy code mới, subscriber có thể acknowledge (xác nhận đã nhận) nhầm các message, dẫn đến mất message vĩnh viễn vì subscriber không được thiết lập để retain (giữ lại) các message đã được acknowledge.

📌 Mục tiêu chính: Tìm cách đảm bảo có thể recover (khôi phục) từ lỗi sau khi deploy, nghĩa là có thể replay (phát lại) các message bị mất do lỗi acknowledge. Pub/Sub là dịch vụ messaging managed của Google Cloud, nơi message có thể được retain trong một khoảng thời gian (max 7 ngày) trên subscription, và các tính năng như snapshot và Seek cho phép replay message từ một điểm thời gian cụ thể.

🛠️ Bối cảnh kỹ thuật (cập nhật đến 2026): Theo tài liệu GCP Pub/Sub mới nhất (phiên bản 2024-2026), message retention mặc định là 7 ngày trên subscription. Snapshot là bản sao trạng thái subscription tại một thời điểm, cho phép Seek để replay message từ snapshot đó trở đi, ngay cả khi chúng đã được acknowledge hoặc expire.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create a Pub/Sub snapshot before deploying new subscriber code. Use a Seek operation to re-deliver messages that became available after the snapshot was created.

Lý do chọn đáp án này 🏆:

  • Trước khi deploy code mới, tạo snapshot của subscription để "chụp" trạng thái message tại thời điểm đó (message chưa được xử lý hoặc available).
  • Sau deploy, nếu có lỗi acknowledge dẫn đến mất message, sử dụng Seek operation để replay tất cả message từ sau snapshot (bao gồm cả message mới publish sau snapshot).
  • Điều này đảm bảo recover hoàn hảo vì snapshot không phụ thuộc vào retention policy của subscriber, và Seek có thể nhắm chính xác đến message "became available after snapshot". Tính năng này được thiết kế dành riêng cho replay trong Pub/Sub (hỗ trợ đến 2026, không giới hạn số lượng snapshot).

📘 Tài liệu tham khảo:

📋 Giải thích tất cả các phương án (đúng và sai)

  • Phương án 1: Set up the Pub/Sub emulator on your local machine. Validate the behavior of your new subscriber logic before deploying it to production.
    ❌ Sai vì: Emulator chỉ dùng để test local (không kết nối production data), không giúp recover message thực tế sau deploy. Nó chỉ validate logic trước deploy, nhưng không giải quyết vấn đề mất message sau khi đã deploy lên production (không có dữ liệu thực để replay). Không liên quan đến recover post-deployment.

  • Phương án 2 (Đúng): Create a Pub/Sub snapshot before deploying new subscriber code. Use a Seek operation to re-deliver messages that became available after the snapshot was created.
    ✅ Đúng vì: Như giải thích ở trên, snapshot + Seek là cách chuẩn và an toàn nhất để replay message từ điểm cụ thể, đảm bảo không mất dữ liệu dù subscriber không retain acknowledged messages. Hoàn hảo cho scenario này.

  • Phương án 3: Use Cloud Build for your deployment. If an error occurs after deployment, use a Seek operation to locate a timestamp logged by Cloud Build at the start of the deployment.
    ❌ Sai vì: Cloud Build là CI/CD tool, có log timestamp nhưng Seek trong Pub/Sub chỉ hỗ trợ seek đến snapshot hoặc timestamp có message publish time, không phải log deployment. Không đảm bảo chính xác message "available after deployment", dễ dẫn đến replay thừa/thiếu. Không phải best practice cho recover message loss.

  • Phương án 4: Enable dead-lettering on the Pub/Sub topic to capture messages that aren't successfully acknowledged. If an error occurs after deployment, re-deliver any messages captured by the dead-letter queue.
    ❌ Sai vì: Dead-lettering (DLQ) được enable trên subscription, không phải topic, và chỉ capture message không acknowledge sau max delivery attempts (mặc định 5-10 lần). Nó không capture message acknowledge nhầm ngay lần đầu (leading to loss), nên không recover được message đã mất do ack sai. DLQ dùng cho poison messages, không phải toàn bộ recover post-deploy.

🧠 Kết luận: Phương án đúng tận dụng snapshot + Seek – tính năng core của Pub/Sub để tránh downtime và data loss trong production (best practice theo GCP 2026). Các phương án sai chỉ là workaround gián tiếp hoặc không khớp scenario! 🚀