Ngân hàng đề — Google Cloud Professional Data Engineer

Tìm thấy 429 câu.

Câu 361
You want to migrate your existing Teradata data warehouse to BigQuery. You want to move the historical data to BigQuery by using the most efficient method that requires the least amount of programming, but local storage space on your existing data warehouse is limited. What should you do?
  1. A Use BigQuery Data Transfer Service by using the Java Database Connectivity (JDBC) driver with FastExport connection.
  2. B Create a Teradata Parallel Transporter (TPT) export script to export the historical data, and import to BigQuery by using the bq command-line tool.
  3. C Use BigQuery Data Transfer Service with the Teradata Parallel Transporter (TPT) tbuild utility.
  4. D Create a script to export the historical data, and upload in batches to Cloud Storage. Set up a BigQuery Data Transfer Service instance from Cloud Storage to BigQuery.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này xoay quanh việc di chuyển (migrate) kho dữ liệu Teradata hiện có sang BigQuery trên Google Cloud. Các yêu cầu chính bao gồm:

  • Chuyển dữ liệu lịch sử (historical data) một cách hiệu quả nhất (most efficient).
  • Yêu cầu ít lập trình nhất (least amount of programming).
  • Dung lượng lưu trữ cục bộ (local storage) trên kho dữ liệu Teradata hiện tại bị hạn chế.

📌 Bối cảnh: Teradata là một hệ thống data warehouse truyền thống, và BigQuery là dịch vụ data warehouse serverless của Google Cloud. Phương pháp lý tưởng phải tận dụng các công cụ tự động của BigQuery để tránh export dữ liệu lớn ra đĩa cục bộ (do hạn chế storage), đồng thời giảm thiểu code tùy chỉnh. Kiến thức cập nhật đến 2026: BigQuery Data Transfer Service (DTS) hỗ trợ kết nối trực tiếp với Teradata qua JDBC driver kết hợp FastExport, cho phép stream dữ liệu trực tiếp mà không cần lưu trữ trung gian lớn (theo docs GCP mới nhất).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use BigQuery Data Transfer Service by using the Java Database Connectivity (JDBC) driver with FastExport connection.

🛠️ Lý do chi tiết:

  • BigQuery Data Transfer Service (DTS) hỗ trợ Teradata connector chính thức, sử dụng JDBC driver kết hợp FastExport (tính năng export nhanh của Teradata).
  • Phương pháp này stream dữ liệu trực tiếp từ Teradata sang BigQuery, không yêu cầu export full data ra local storage (giải quyết hạn chế storage).
  • Ít lập trình nhất: Chỉ cần cấu hình qua console/UI hoặc gcloud CLI, không cần script phức tạp.
  • Hiệu quả cao: FastExport tối ưu hóa tốc độ export parallel, phù hợp dữ liệu lớn (petabyte-scale).
  • Đây là phương pháp khuyến nghị chính thức từ Google cho migration Teradata → BigQuery (cập nhật 2024-2026).

📘 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên docs GCP/BigQuery DTS mới nhất:

  • Use BigQuery Data Transfer Service by using the Java Database Connectivity (JDBC) driver with FastExport connection.
    ✅ Đúng (như đã giải thích ở trên). Phương pháp này tận dụng Teradata connector built-in của DTS, stream dữ liệu nhanh chóng, không cần storage local lớn và zero coding tùy chỉnh. Hoàn hảo khớp yêu cầu! 🏆

  • Create a Teradata Parallel Transporter (TPT) export script to export the historical data, and import to BigQuery by using the bq command-line tool.
    ❌ Sai. Yêu cầu viết script TPT tùy chỉnh (nhiều programming), export dữ liệu ra file → lưu local storage (vi phạm hạn chế storage). Sau đó dùng bq load import thủ công, không efficient bằng DTS tự động. TPT là tool Teradata nhưng không tích hợp trực tiếp với BigQuery.

  • Use BigQuery Data Transfer Service with the Teradata Parallel Transporter (TPT) tbuild utility.
    ❌ Sai. BigQuery DTS không hỗ trợ trực tiếp TPT tbuild utility (dùng để build job TPT). DTS chỉ hỗ trợ JDBC + FastExport cho Teradata, không có connector cho TPT. Sử dụng TPT vẫn cần export trung gian, tốn storage và setup phức tạp hơn.

  • Create a script to export the historical data, and upload in batches to Cloud Storage. Set up a BigQuery Data Transfer Service instance from Cloud Storage to BigQuery.
    ❌ Sai. Vẫn cần script export tùy chỉnh (programming cao), chia batches → tốn local storage tạm thời (không giải quyết hạn chế). Phần load từ GCS sang BigQuery qua DTS thì OK, nhưng toàn bộ quy trình không efficient và không ít code nhất so với DTS trực tiếp.

📚 Tài liệu tham khảo

Hy vọng phân tích này giúp bạn ôn thi chứng chỉ hiệu quả! 🚀 Nếu cần thêm ví dụ code/config, hãy hỏi nhé!

Câu 362
You are on the data governance team and are implementing security requirements. You need to encrypt all your data in BigQuery by using an encryption key managed by your team. You must implement a mechanism to generate and store encryption material only on your on-premises hardware security module (HSM). You want to rely on Google managed solutions. What should you do?
  1. A Create the encryption key in the on-premises HSM, and import it into a Cloud Key Management Service (Cloud KMS) key. Associate the created Cloud KMS key while creating the BigQuery resources.
  2. B Create the encryption key in the on-premises HSM and link it to a Cloud External Key Manager (Cloud EKM) key. Associate the created Cloud KMS key while creating the BigQuery resources.
  3. C Create the encryption key in the on-premises HSM, and import it into Cloud Key Management Service (Cloud HSM) key. Associate the created Cloud HSM key while creating the BigQuery resources.
  4. D Create the encryption key in the on-premises HSM. Create BigQuery resources and encrypt data while ingesting them into BigQuery.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này thuộc lĩnh vực Data Governance và Security trên Google Cloud Platform (GCP), cụ thể liên quan đến việc mã hóa dữ liệu trong BigQuery. Bạn là thành viên đội ngũ quản trị dữ liệu, cần triển khai yêu cầu bảo mật:

  • Mã hóa toàn bộ dữ liệu trong BigQuery bằng khóa mã hóa (encryption key) do đội ngũ của bạn quản lý.
  • Khóa mã hóa phải được tạo và lưu trữ CHỈ trên phần cứng HSM (Hardware Security Module) tại chỗ (on-premises) của tổ chức.
  • Phải sử dụng các giải pháp được Google quản lý (Google-managed solutions) để tích hợp mượt mà, không tự xây dựng từ đầu.

Mục tiêu chính là sử dụng Customer-Managed Encryption Keys (CMEK) cho BigQuery, nhưng với yêu cầu đặc biệt: khóa vật lý (key material) KHÔNG được lưu trữ trên cloud của Google, mà chỉ trên HSM on-premises. Điều này đảm bảo quyền kiểm soát tuyệt đối (BYOK - Bring Your Own Key) mà vẫn tận dụng dịch vụ Google như BigQuery và KMS.

📘 Kiến thức cập nhật (đến 2026): GCP hỗ trợ Cloud External Key Manager (Cloud EKM) từ năm 2022, cho phép liên kết khóa từ HSM ngoài (như Thales, Fortanix) với KMS mà không import key material vào Google. BigQuery hỗ trợ CMEK qua EKM từ phiên bản mới nhất (xem docs GCP 2026).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create the encryption key in the on-premises HSM and link it to a Cloud External Key Manager (Cloud EKM) key. Associate the created Cloud KMS key while creating the BigQuery resources.

🛠️ Lý do chi tiết:

  • Tạo khóa trên on-premises HSM (ví dụ: Thales Luna HSM) và liên kết (link) với Cloud EKM key trong Cloud KMS. Key material KHÔNG bao giờ rời HSM, Google chỉ gọi API đến HSM qua kết nối bảo mật (IP Allowlist hoặc VPC).
  • Sau đó, associate EKM key với BigQuery dataset/table khi tạo (qua IAM policy hoặc CMEK config). BigQuery sẽ dùng EKM để mã hóa/giải mã dữ liệu.
  • Hoàn toàn Google-managed: EKM tự động rotate, audit logs, tích hợp BigQuery mà đội ngũ chỉ quản lý HSM. Đáp ứng 100% yêu cầu "generate/store only on-premises" và "rely on Google solutions".

❌ Phân tích tất cả các phương án (đúng/sai)

  • [SAI] Create the encryption key in the on-premises HSM, and import it into a Cloud Key Management Service (Cloud KMS) key. Associate the created Cloud KMS key while creating the BigQuery resources.
    ❌ Sai vì: Việc import key vào Cloud KMS sẽ sao chép key material vào hệ thống Google (dù là CMEK), vi phạm yêu cầu "generate and store ONLY on on-premises HSM". Key material không còn độc quyền trên HSM nữa, Google sẽ quản lý bản sao. Không phù hợp BYOK thuần túy.

  • [ĐÚNG] Create the encryption key in the on-premises HSM and link it to a Cloud External Key Manager (Cloud EKM) key. Associate the created Cloud KMS key while creating the BigQuery resources.
    ✅ Đúng vì: Như giải thích ở trên. EKM chỉ link/reference (không import), key material luôn ở HSM. BigQuery gọi EKM qua KMS proxy. Hoàn hảo cho hybrid security (xem phần ✅).

  • [SAI] Create the encryption key in the on-premises HSM, and import it into Cloud Key Management Service (Cloud HSM) key. Associate the created Cloud HSM key while creating the BigQuery resources.
    ❌ Sai vì: Cloud HSM không tồn tại trên GCP (đây là dịch vụ của AWS - CloudHSM). GCP không có "Cloud HSM key". Import vào KMS cũng sai tương tự phương án 1. BigQuery không hỗ trợ dịch vụ giả định này.

  • [SAI] Create the encryption key in the on-premises HSM. Create BigQuery resources and encrypt data while ingesting them into BigQuery.
    ❌ Sai vì: Không chỉ rõ mechanism tích hợp với Google-managed solutions. Chỉ mã hóa thủ công khi ingest (ví dụ: client-side encryption) sẽ không áp dụng CMEK tự động cho toàn bộ dữ liệu BigQuery (at-rest và query). Không scalable, không audit được qua KMS, và không rely on Google services.

📘 Tài liệu tham khảo (Google Cloud Docs - cập nhật 2026)

Hy vọng phân tích này giúp bạn ôn thi chứng chỉ GCP Professional Data Engineer! 🚀 Nếu cần thêm ví dụ code Terraform/IAM, hãy hỏi nhé!

Câu 363
You maintain ETL pipelines. You notice that a streaming pipeline running on Dataflow is taking a long time to process incoming data, which causes output delays. You also noticed that the pipeline graph was automatically optimized by Dataflow and merged into one step. You want to identify where the potential bottleneck is occurring. What should you do?
  1. A Insert a Reshuffle operation after each processing step, and monitor the execution details in the Dataflow console.
  2. B Insert output sinks after each key processing step, and observe the writing throughput of each block.
  3. C Log debug information in each ParDo function, and analyze the logs at execution time.
  4. D Verify that the Dataflow service accounts have appropriate permissions to write the processed data to the output sinks.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc xử lý sự cố hiệu suất (performance troubleshooting) trong một ETL pipeline streaming chạy trên Google Cloud Dataflow. Cụ thể:

  • Vấn đề chính: Pipeline xử lý dữ liệu đầu vào (incoming data) chậm, dẫn đến trì hoãn output.
  • Nguyên nhân liên quan: Dataflow đã tự động tối ưu hóa pipeline graph bằng cách fusion (hợp nhất) các bước xử lý thành một bước duy nhất. Điều này làm khó xác định bottleneck (điểm nghẽn) vì metrics và execution details bị gộp chung, không thể phân tích riêng lẻ từng bước.
  • Mục tiêu: Tìm cách xác định vị trí bottleneck một cách hiệu quả, tận dụng các tính năng của Dataflow (không phải AWS, vì Dataflow là dịch vụ GCP thuần túy).

Đây là tình huống phổ biến trong Apache Beam trên Dataflow, nơi automatic fusion cải thiện hiệu suất nhưng làm phức tạp debugging. Giải pháp cần phá vỡ fusion để quan sát chi tiết mà không ảnh hưởng lớn đến production. (Kiến thức cập nhật đến 2026: Dataflow vẫn hỗ trợ fusion optimizations trong Beam SDK 2.50+ và Dataflow Runner v2).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Insert a Reshuffle operation after each processing step, and monitor the execution details in the Dataflow console.

Lý do 🛠️:

  • Reshuffle là một transform đặc biệt trong Apache Beam/Dataflow, nó phá vỡ fusion bằng cách shuffle lại dữ liệu (re-partition và re-key), buộc Dataflow tách riêng từng bước trong pipeline graph.
  • Sau khi insert Reshuffle, bạn có thể monitor execution details trong Dataflow console (qua tab "Jobs" > "Graph" hoặc "Metrics") để xem metrics riêng lẻ như CPU, memory, throughput, watermark lag từng step – giúp pinpoint bottleneck chính xác (ví dụ: ParDo chậm hoặc GroupByKey nghẽn).
  • Đây là best practice chính thức của Google cho streaming pipelines bị fusion, không overhead cao ở streaming mode, và tự động scale. Phiên bản mới nhất (2026) vẫn khuyến nghị điều này trước khi dùng custom metrics.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ✅ Insert a Reshuffle operation after each processing step, and monitor the execution details in the Dataflow console.
    Đúng vì lý do như trên: Reshuffle ngăn fusion, expose graph chi tiết để monitor ✅. Hiệu quả nhất cho streaming, low-impact.

  • ❌ Insert output sinks after each key processing step, and observe the writing throughput of each block.
    Sai vì: Thêm output sinks (như WriteToBigQuery) sau mỗi step sẽ tạo I/O overhead lớn, làm pipeline chậm hơn và tốn kém (nhiều writes không cần thiết). Không giải quyết fusion, chỉ quan sát throughput writes – không detect bottleneck ở compute/processing 🛑.

  • ❌ Log debug information in each ParDo function, and analyze the logs at execution time.
    Sai vì: Logging debug trong ParDo (qua Beam logging) chỉ cung cấp logs text-based, khó aggregate/analyze real-time cho streaming lớn (hàng triệu elements). Không expose metrics graph-level, và fusion vẫn che khuất bottleneck. Phù hợp debug nhỏ, không scale cho production ❌.

  • ❌ Verify that the Dataflow service accounts have appropriate permissions to write the processed data to the output sinks.
    Sai vì: Kiểm tra permissions (IAM roles như Dataflow Service Agent) chỉ fix lỗi authorization failures (ví dụ: 403 errors), không liên quan đến processing chậm do fusion/bottleneck nội bộ. Vấn đề là performance, không phải access denied 🚫.

Câu 364
You are running your BigQuery project in the on-demand billing model and are executing a change data capture (CDC) process that ingests data. The CDC process loads 1 GB of data every 10 minutes into a temporary table, and then performs a merge into a 10 TB target table. This process is very scan intensive and you want to explore options to enable a predictable cost model. You need to create a BigQuery reservation based on utilization information gathered from BigQuery Monitoring and apply the reservation to the CDC process. What should you do?
  1. A Create a BigQuery reservation for the dataset.
  2. B Create a BigQuery reservation for the job.
  3. C Create a BigQuery reservation for the service account running the job.
  4. D Create a BigQuery reservation for the project.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào BigQuery (dịch vụ kho dữ liệu của Google Cloud Platform - GCP), không phải AWS như mô tả ban đầu (có thể là nhầm lẫn). Tình huống:

  • Dự án đang sử dụng mô hình thanh toán on-demand (thanh toán theo lượng dữ liệu scan thực tế).
  • Quy trình Change Data Capture (CDC) liên tục ingest dữ liệu: Mỗi 10 phút load 1 GB vào bảng tạm thời (temporary table), sau đó merge vào bảng đích 10 TB.
  • Quy trình này rất intensive về scan (quét dữ liệu nhiều), dẫn đến chi phí biến động cao.
  • Mục tiêu: Chuyển sang mô hình chi phí predictable bằng BigQuery reservation (đặt chỗ slot để sử dụng flat-rate pricing).
  • Dựa trên dữ liệu utilization từ BigQuery Monitoring để tạo reservation và apply cho quy trình CDC.

📈 Vấn đề cốt lõi: Cần chọn mức độ (scope) phù hợp để tạo reservation sao cho nó được áp dụng tự động cho các job CDC, đảm bảo chi phí ổn định mà không ảnh hưởng đến các workload khác.

✅ Đáp án đúng và lý do chọn

Đáp án đúng: Create a BigQuery reservation for the project.

🛠️ Lý do chi tiết:

  • Trong BigQuery, reservation (đặt chỗ slot) được tạo và quản lý ở mức project (dự án). Khi tạo reservation cho project chứa các job CDC, nó sẽ tự động áp dụng cho tất cả các query/job chạy trong project đó (bao gồm load/merge CDC).
  • Dựa trên BigQuery Monitoring, bạn ước lượng slot cần (ví dụ: từ metrics như query/scanned_bytes hoặc slot_ms), tạo reservation với số slot phù hợp, rồi assign qua slot pool hoặc trực tiếp cho project.
  • Điều này chuyển project từ on-demand sang flat-rate (chi phí cố định theo slot/giờ), giúp predictable cost cho workload scan-intensive như CDC (merge 10TB table quét nhiều dữ liệu).
  • Theo tài liệu GCP mới nhất (2024-2026), không có thay đổi cơ bản: Reservations vẫn scope ở project/folder/organization, ưu tiên project cho trường hợp này.

📘 Nguồn tham khảo:

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Create a BigQuery reservation for the dataset.
    Phương án này sai vì reservation không hỗ trợ scope ở mức dataset. BigQuery chỉ quản lý reservation ở mức project/folder/organization, không phải dataset. Nếu tạo ở dataset, job CDC (load/merge) vẫn fallback về on-demand, không achieve predictable cost. Dataset chỉ dùng cho partitioning/clustering, không bind reservation.

  • ❌ [SAI] Create a BigQuery reservation for the job.
    Phương án sai vì reservation không tạo trực tiếp cho từng job riêng lẻ. Job chỉ sử dụng reservation từ project/slot pool đã assign. Tạo per-job sẽ không khả thi cho CDC (chạy liên tục mỗi 10 phút), dẫn đến quản lý phức tạp và không tự động apply từ Monitoring data.

  • ❌ [SAI] Create a BigQuery reservation for the service account running the job.
    Sai vì reservation không bind với service account (tài khoản dịch vụ). Service account chỉ dùng cho authentication/authorization job, không ảnh hưởng đến billing model. Job vẫn dùng reservation của project mà nó chạy, bất kể account nào.

  • ✅ [ĐÚNG] Create a BigQuery reservation for the project.
    (Đã giải thích chi tiết ở phần trên). Đây là cách chuẩn và hiệu quả nhất, phù hợp với best practice GCP cho workload predictable như CDC.

🔍 Lời khuyên thực tế: Sau khi tạo, dùng BigQuery Admin Console hoặc bq CLI để monitor utilization (--dry_run cho query planning) và scale slot nếu cần. Nếu multi-project, dùng slot pool để share reservation!

Câu 365
You are designing a fault-tolerant architecture to store data in a regional BigQuery dataset. You need to ensure that your application is able to recover from a corruption event in your tables that occurred within the past seven days. You want to adopt managed services with the lowest RPO and most cost-effective solution. What should you do?
  1. A Access historical data by using time travel in BigQuery.
  2. B Export the data from BigQuery into a new table that excludes the corrupted data
  3. C Create a BigQuery table snapshot on a daily basis.
  4. D Migrate your data to multi-region BigQuery buckets.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm về BigQuery (Google Cloud Platform)

📘 Giới thiệu ngắn gọn:
Là một Google Cloud Professional Data Engineer, tôi sẽ phân tích câu hỏi này dựa trên kiến thức cập nhật mới nhất về BigQuery (phiên bản tính năng đến năm 2026, theo tài liệu chính thức Google Cloud). Câu hỏi tập trung vào thiết kế kiến trúc chịu lỗi (fault-tolerant) cho dataset BigQuery regional, nhằm khôi phục dữ liệu từ sự cố hỏng (corruption) trong 7 ngày qua. Yêu cầu chính: sử dụng dịch vụ managed (quản lý tự động), RPO thấp nhất (Recovery Point Objective - thời gian mất dữ liệu ngắn nhất), và tiết kiệm chi phí nhất.

✅ Nội dung câu hỏi được giải thích chi tiết:

  • Bối cảnh: Bạn đang thiết kế hệ thống lưu trữ dữ liệu trong một dataset BigQuery regional (dữ liệu chỉ ở một vùng địa lý cụ thể, không phải multi-region).
  • Vấn đề cần giải quyết: Ứng dụng phải khôi phục (recover) từ sự cố hỏng bảng dữ liệu (table corruption) xảy ra trong vòng 7 ngày qua. Corruption có thể do lỗi DELETE/UPDATE nhầm, hoặc lỗi hệ thống.
  • Yêu cầu then chốt:
    • Sử dụng managed services (dịch vụ Google quản lý tự động, không cần tự vận hành).
    • Lowest RPO: RPO là khoảng thời gian dữ liệu có thể bị mất khi recover (ví dụ: RPO=0 phút nghĩa là recover chính xác đến giây phút gần nhất).
    • Most cost-effective: Giải pháp rẻ nhất, không tốn lưu trữ/snapshot thừa.
  • Mục tiêu: Fault-tolerant architecture giúp ứng dụng tiếp tục hoạt động mà không mất dữ liệu lịch sử gần (7 ngày). BigQuery hỗ trợ time-based recovery tự nhiên qua các tính năng built-in.

🏆 Đáp án đúng và lý do lựa chọn:
Đáp án đúng: Access historical data by using time travel in BigQuery.
✅ Lý do chi tiết:

  • Time Travel là tính năng managed của BigQuery (không cần cấu hình thủ công), cho phép query dữ liệu lịch sử lên đến 7 ngày cho standard tables (90 ngày cho partitioned/clustered tables - cập nhật 2023-2026).
  • Lowest RPO: Recover granular (đến mức phút), gần như RPO=0, vì dữ liệu được lưu metadata tự động mà không tốn thêm chi phí lưu trữ.
  • Cost-effective nhất: Không mất phí snapshot/export, chỉ tính phí query khi sử dụng (pay-per-use). Bạn có thể CLONE table từ timestamp cụ thể: CREATE TABLE new_table CLONE old_table FOR TIMESTAMP AS '2024-01-01 12:00:00'.
  • Hoàn hảo cho regional dataset, fault-tolerant cao.
    📚 Nguồn tham khảo: BigQuery Time Travel Docs (cập nhật 2024-2026).

🛠️ Giải thích tất cả các phương án (đúng/sai):
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh, chỉ giải thích bằng tiếng Việt với emoji đánh dấu.

  • ✅ Access historical data by using time travel in BigQuery.
    Đúng! Như đã giải thích ở trên: Managed, RPO thấp nhất (granular đến phút), chi phí thấp nhất (không lưu trữ thừa), hỗ trợ recover corruption trong 7 ngày qua một cách tự động và hiệu quả. Lý tưởng cho yêu cầu fault-tolerant.

  • ❌ Export the data from BigQuery into a new table that excludes the corrupted data.
    Sai! Đây là cách thủ công (không managed), phải query/export thủ công qua BigQuery Jobs hoặc Data Transfer Service. RPO không thấp (phụ thuộc tần suất export, có thể mất dữ liệu >1 ngày), tốn chi phí compute/export/storage cao hơn, không tự động recover historical data trong 7 ngày. Không phù hợp lowest RPO/cost-effective.

  • ❌ Create a BigQuery table snapshot on a daily basis.
    Sai! BigQuery hỗ trợ table snapshots (từ 2022, cập nhật 2026), nhưng tạo daily chỉ cho RPO=1 ngày (mất dữ liệu trong ngày corruption). Không lowest RPO (không granular), phải tự động hóa qua Scheduled Queries (vẫn semi-managed), tốn chi phí lưu trữ snapshot lâu dài (snapshot expire sau 90 ngày nhưng vẫn charge). Đắt hơn time travel.

  • ❌ Migrate your data to multi-region BigQuery buckets.
    Sai! BigQuery datasets là regional/multi-regional, không dùng "buckets" (buckets là Cloud Storage). Multi-region cải thiện high availability và disaster recovery zonal/regional failure, nhưng không recover corruption logic (như DELETE nhầm) trong 7 ngày qua - dữ liệu vẫn sync corruption. Tốn chi phí cao hơn (egress/storage multi-region), không giải quyết RPO thấp cho table-level corruption. Không managed cho historical recovery.

💡 Kết luận & Lời khuyên:
Giải pháp Time Travel là lựa chọn tối ưu nhất theo best practices GCP (Zero-ETL, serverless). Để triển khai fault-tolerant đầy đủ, kết hợp với BigQuery reservations cho slot và column-level lineage. Nếu cần demo code, tham khảo BigQuery Samples GitHub. 🚀

Câu 366
You are building a streaming Dataflow pipeline that ingests noise level data from hundreds of sensors placed near construction sites across a city. The sensors measure noise level every ten seconds, and send that data to the pipeline when levels reach above 70 dBA. You need to detect the average noise level from a sensor when data is received for a duration of more than 30 minutes, but the window ends when no data has been received for 15 minutes. What should you do?
  1. A Use session windows with a 15-minute gap duration.
  2. B Use session windows with a 30-minute gap duration.
  3. C Use hopping windows with a 15-minute window, and a thirty-minute period.
  4. D Use tumbling windows with a 15-minute window and a fifteen-minute .withAllowedLateness operator.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc xây dựng một pipeline streaming trên Google Cloud Dataflow (sử dụng Apache Beam) để xử lý dữ liệu mức độ ồn từ hàng trăm cảm biến gần các công trường xây dựng trong thành phố.

  • Dữ liệu đầu vào: Các cảm biến đo mức ồn mỗi 10 giây, chỉ gửi dữ liệu vào pipeline khi mức ồn > 70 dBA.
  • Yêu cầu xử lý:
    • Phát hiện trung bình mức ồn từ một cảm biến khi dữ liệu được nhận liên tục trong hơn 30 phút (tức là session dữ liệu kéo dài >30 phút).
    • Window kết thúc khi không nhận dữ liệu trong 15 phút (gap im lặng 15 phút).
      📘 Mục tiêu chính: Sử dụng loại windowing phù hợp trong Dataflow để nhóm dữ liệu theo session động, đảm bảo window chỉ đóng khi có khoảng im lặng 15 phút và chỉ xử lý (tính average) cho các session dài hơn 30 phút (có thể kết hợp filter sau window).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use session windows with a 15-minute gap duration.
Lý do:
🛠️ Session windows trong Apache Beam/Dataflow là loại window động, nhóm dữ liệu từ cùng một key (ví dụ: sensor ID) thành các session riêng biệt dựa trên gap duration (khoảng thời gian im lặng giữa các event).

  • Với gap duration = 15 phút: Window bắt đầu khi nhận event đầu tiên, tiếp tục mở rộng miễn là các event sau đến trong vòng <15 phút so với event trước. Window kết thúc ngay khi gap >15 phút (khớp yêu cầu "window ends when no data has been received for 15 minutes").
  • Để detect average chỉ khi duration >30 phút: Sau khi window đóng, áp dụng filter trên metadata của window (như PaneInfo hoặc kích thước window) để chỉ tính average cho session dài >30 phút.
    ✅ Điều này hoàn hảo cho dữ liệu burst-y (gửi không đều, chỉ khi >70 dBA), và phù hợp với streaming pipeline trên Dataflow (cập nhật Beam 2.54+ đến 2026).

📋 Giải thích tất cả các phương án

  • ✅ Use session windows with a 15-minute gap duration.
    🟢 Đúng vì khớp chính xác gap 15 phút để kết thúc window, và session tự động nhóm dữ liệu liên tục (dễ filter duration >30 phút sau). Hoạt động tốt trong streaming với watermark và late data handling mặc định của Dataflow.

  • ❌ Use session windows with a 30-minute gap duration.
    🔴 Sai vì gap 30 phút sẽ giữ window mở lâu hơn (chỉ kết thúc sau 30 phút im lặng), không khớp yêu cầu "no data for 15 minutes". Dẫn đến window kéo dài không cần thiết, tính average sớm hoặc muộn.

  • ❌ Use hopping windows with a 15-minute window, and a thirty-minute period.
    🔴 Sai vì hopping windows (sliding windows) có kích thước cố định 15 phút và period 30 phút (overlap 50%), không linh hoạt theo dữ liệu thực tế. Không tự kết thúc dựa trên gap 15 phút, và không detect session duration >30 phút động (chỉ fixed overlap).

  • ❌ Use tumbling windows with a 15-minute window and a fifteen-minute .withAllowedLateness operator.
    🔴 Sai vì tumbling windows là fixed-size non-overlapping (15 phút mỗi window), không nhóm theo session liên tục. .withAllowedLateness(15 phút) chỉ cho phép late data đến sau watermark 15 phút, nhưng không thay đổi logic kết thúc window dựa trên gap dữ liệu thực tế (vẫn emit đúng giờ fixed).

📚 Tài liệu tham khảo (cập nhật mới nhất đến 2026)

Câu 367
You are creating a data model in BigQuery that will hold retail transaction data. Your two largest tables, sales_transaction_header and sales_transaction_line, have a tightly coupled immutable relationship. These tables are rarely modified after load and are frequently joined when queried. You need to model the sales_transaction_header and sales_transaction_line tables to improve the performance of data analytics queries. What should you do?
  1. A Create a sales_transaction table that holds the sales_transaction_header information as rows and the sales_transaction_line rows as nested and repeated fields.
  2. B Create a sales_transaction table that holds the sales_transaction_header and sales_transaction_line information as rows, duplicating the sales_transaction_header data for each line.
  3. C Create a sales_transaction table that stores the sales_transaction_header and sales_transaction_line data as a JSON data type.
  4. D Create separate sales_transaction_header and sales_transaction_line tables and, when querying, specify the sales_transaction_line first in the WHERE clause.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc mô hình hóa dữ liệu (data modeling) trong BigQuery (dịch vụ kho dữ liệu của Google Cloud) cho dữ liệu giao dịch bán lẻ (retail transaction data). Cụ thể:

  • Có hai bảng lớn nhất: sales_transaction_header (chứa thông tin header của giao dịch, như ID giao dịch, khách hàng, tổng tiền...) và sales_transaction_line (chứa thông tin chi tiết dòng sản phẩm trong giao dịch, như sản phẩm, số lượng, giá...).
  • Mối quan hệ giữa chúng là tightly coupled (liên kết chặt chẽ), immutable (không thay đổi sau khi load), rarely modified (hiếm khi sửa), và frequently joined (thường được JOIN khi query).
  • Mục tiêu: Cải thiện hiệu suất truy vấn phân tích dữ liệu (data analytics queries).

Vấn đề chính là tránh JOIN thường xuyên giữa hai bảng lớn, vì JOIN có thể tốn kém về chi phí và thời gian trong BigQuery (slot usage cao). Giải pháp tối ưu là denormalize dữ liệu bằng cách sử dụng nested và repeated fields để lưu trữ dữ liệu liên quan trong một bảng duy nhất, tận dụng schema linh hoạt của BigQuery. Kiến thức này dựa trên best practices BigQuery đến năm 2026 (BigQuery schema evolution và nested structs hỗ trợ query hiệu quả hơn với DML/partitioning).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create a sales_transaction table that holds the sales_transaction_header information as rows and the sales_transaction_line rows as nested and repeated fields.

Lý do 🛠️:

  • Phương án này denormalize dữ liệu bằng cách lưu header làm row chính, và line items làm nested STRUCT (cho header) + REPEATED ARRAY (cho nhiều dòng line).
  • Lợi ích: Truy vấn chỉ scan một bảng duy nhất, không cần JOIN → giảm I/O, slot usage, và thời gian query lên đến 10x cho analytics workload.
  • Phù hợp hoàn hảo với đặc tính: tightly coupled, immutable → nested fields ổn định, query nhanh với UNNEST.
  • Đây là best practice của BigQuery cho parent-child relationship (header-lines), hỗ trợ clustering/partitioning hiệu quả.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ✅ Create a sales_transaction table that holds the sales_transaction_header information as rows and the sales_transaction_line rows as nested and repeated fields.
    Giải thích đúng 🏆: Như trên, đây là cách tối ưu nhất. Nested/repeated giảm scan volume, hỗ trợ query như UNNEST(line_items) siêu nhanh. Không duplicate dữ liệu thừa, tiết kiệm storage.

  • ❌ Create a sales_transaction table that holds the sales_transaction_header and sales_transaction_line information as rows, duplicating the sales_transaction_header data for each line.
    Giải thích sai 🚫: Đây là full denormalization (duplicate header cho mỗi line), dẫn đến storage waste lớn (header lặp lại nhiều lần nếu giao dịch có nhiều line). Query nhanh nhưng không hiệu quả bằng nested (storage cao hơn 5-10x), vi phạm nguyên tắc "store efficiently" trong BigQuery.

  • ❌ Create a sales_transaction table that stores the sales_transaction_header and sales_transaction_line data as a JSON data type.
    Giải thích sai ⚠️: JSON trong BigQuery (STRING hoặc JSON type từ 2023) không được tối ưu cho analytics. Query JSON cần JSON_EXTRACT/JSON_VALUE, chậm hơn nested 2-5x, không hỗ trợ indexing/clustering tốt. Phù hợp raw data hơn là structured analytics.

  • ❌ Create separate sales_transaction_header and sales_transaction_line tables and, when querying, specify the sales_transaction_line first in the WHERE clause.
    Giải thích sai 🔍: Giữ separate tables vẫn yêu cầu JOIN, và reorder WHERE clause chỉ tối ưu filter pushdown nhẹ (giảm scan line trước), nhưng với bảng lớn + frequent joins → vẫn tốn kém (cross-join potential). Không giải quyết gốc rễ, hiệu suất kém hơn denormalized model.

Kết luận 🎯: Sử dụng nested/repeated là chuẩn mực cho BigQuery analytics đến 2026, giúp scale petabyte queries hiệu quả! Nếu implement, thêm partitioning trên transaction_date và clustering trên customer_id để boost nữa nhé.

Câu 368
You created a new version of a Dataflow streaming data ingestion pipeline that reads from Pub/Sub and writes to BigQuery. The previous version of the pipeline that runs in production uses a 5-minute window for processing. You need to deploy the new version of the pipeline without losing any data, creating inconsistencies, or increasing the processing latency by more than 10 minutes. What should you do?
  1. A Update the old pipeline with the new pipeline code.
  2. B Snapshot the old pipeline, stop the old pipeline, and then start the new pipeline from the snapshot.
  3. C Drain the old pipeline, then start the new pipeline.
  4. D Cancel the old pipeline, then start the new pipeline.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc cập nhật một pipeline streaming Dataflow trên Google Cloud Platform (GCP). Cụ thể:

  • Pipeline hiện tại (phiên bản cũ) đang chạy production, đọc dữ liệu từ Pub/Sub (dịch vụ messaging real-time) và ghi vào BigQuery (data warehouse), sử dụng window 5 phút để xử lý dữ liệu theo batch nhỏ trong streaming.
  • Bạn tạo phiên bản mới của pipeline và cần deploy nó mà không mất dữ liệu (no data loss), không tạo inconsistency (dữ liệu không bị trùng lặp hoặc lệch lạc), và tăng latency tối đa không quá 10 phút.
  • Mục tiêu: Chuyển tiếp mượt mà từ old pipeline sang new pipeline trong môi trường streaming, nơi dữ liệu liên tục chảy vào Pub/Sub.

📘 Bối cảnh kiến thức GCP Dataflow (cập nhật đến 2026): Dataflow là dịch vụ managed Apache Beam cho batch/streaming processing. Với streaming, deploy new version cần kỹ thuật đặc biệt để xử lý dữ liệu đang "in-flight" (đang trong window/buffer). Window 5 phút nghĩa là dữ liệu tích lũy 5 phút mới trigger xử lý, nên bất kỳ gián đoạn nào cũng có thể tăng latency hoặc mất data nếu không xử lý đúng.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Drain the old pipeline, then start the new pipeline.

🛠️ Lý do chi tiết:

  • Drain là lệnh chính thức của Dataflow (gCloud CLI/API) dành riêng cho streaming pipelines. Nó cho phép pipeline cũ hoàn thành tất cả dữ liệu đang trong window/buffer (ở đây ~5 phút), sau đó tự động dừng mà không drop data.
  • New pipeline start ngay sau, đọc từ cùng Pub/Sub topic → zero data loss, no inconsistency (vì old hoàn thành trước khi new bắt đầu), và latency tăng chỉ ~5 phút (thời gian drain window), dưới ngưỡng 10 phút.
  • Đây là best practice được AWS... à GCP recommend cho migration streaming pipelines (không phải AWS, câu hỏi là GCP thuần túy).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một cách đầy đủ, dựa trên docs GCP Dataflow mới nhất (Apache Beam 2.54+ và Dataflow runner 2026):

  • ✅ [ĐÚNG] Drain the old pipeline, then start the new pipeline.
    🟢 Đúng vì: Như giải thích trên, drain đảm bảo unbounded data in-flight được process hết (watermark tiến tới EOF), tránh loss/inconsistency. Latency chỉ tăng bằng window size (5 phút).
    📘 Nguồn: Dataflow Streaming Pipelines: Drain & gcloud dataflow jobs drain.

  • ❌ [SAI] Update the old pipeline with the new pipeline code.
    ❌ Sai vì: Update trực tiếp (qua template/update job) có thể gây code change mid-flight, dẫn đến inconsistency (dữ liệu cũ/new mix trong cùng window) hoặc restart bundles, tăng latency >10 phút và rủi ro duplicate data nếu schema thay đổi.

  • ❌ [SAI] Snapshot the old pipeline, stop the old pipeline, and then start the new pipeline from the snapshot.
    ❌ Sai vì: Snapshot chỉ hỗ trợ batch pipelines hoặc experimental streaming (từ Beam 2.28+), không guarantee zero-loss cho unbounded streaming với Pub/Sub. Stop + resume từ snapshot có thể mất data sau watermark hoặc tăng latency lớn (snapshot overhead >10 phút), không phù hợp production migration.

  • ❌ [SAI] Cancel the old pipeline, then start the new pipeline.
    ❌ Sai vì: Cancel (hoặc stop) drop ngay tất cả dữ liệu in-flight/buffer (bao gồm window 5 phút đang process), gây data loss lớn và inconsistency (Pub/Sub message sau cancel mới được new pipeline đọc). Latency có thể tăng >10 phút nếu backlog tích tụ.

🏆 Kết luận & Best Practices bổ sung

  • Drain là lựa chọn tối ưu cho scenario này, đặc biệt với windowed streaming.
  • Lưu ý 2026: Dataflow hỗ trợ flex templates và dynamic work pools để deploy nhanh hơn, nhưng drain vẫn là core cho migration.
  • 📘 Tài liệu tham khảo chính:

Nếu cần demo code hoặc lab, hãy cho tôi biết nhé! 🚀

Câu 369
Your organization's data assets are stored in BigQuery, Pub/Sub, and a PostgreSQL instance running on Compute Engine. Because there are multiple domains and diverse teams using the data, teams in your organization are unable to discover existing data assets. You need to design a solution to improve data discoverability while keeping development and configuration efforts to a minimum. What should you do?
  1. A Use Data Catalog to automatically catalog BigQuery datasets. Use Data Catalog APIs to manually catalog Pub/Sub topics and PostgreSQL tables.
  2. B Use Data Catalog to automatically catalog BigQuery datasets and Pub/Sub topics. Use Data Catalog APIs to manually catalog PostgreSQL tables.
  3. C Use Data Catalog to automatically catalog BigQuery datasets and Pub/Sub topics. Use custom connectors to manually catalog PostgreSQL tables.
  4. D Use customer connectors to manually catalog BigQuery datasets, Pub/Sub topics, and PostgreSQL tables.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi

📖 Nội dung câu hỏi:
Câu hỏi mô tả tình huống tổ chức của bạn lưu trữ dữ liệu ở BigQuery (kho dữ liệu phân tích), Pub/Sub (dịch vụ messaging thời gian thực), và PostgreSQL chạy trên Compute Engine (máy ảo). Vấn đề là các đội ngũ từ nhiều domain khác nhau khó khám phá (discover) dữ liệu hiện có do thiếu catalog tập trung. Yêu cầu thiết kế giải pháp cải thiện khả năng discoverability với nỗ lực phát triển và cấu hình tối thiểu.
✅ Mục tiêu chính: Sử dụng Data Catalog (dịch vụ metadata management của Google Cloud) để tự động hóa catalog hóa càng nhiều càng tốt, giảm manual work. Data Catalog giúp tìm kiếm, tag metadata, và tích hợp liền mạch với các dịch vụ GCP native.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Data Catalog to automatically catalog BigQuery datasets and Pub/Sub topics. Use Data Catalog APIs to manually catalog PostgreSQL tables.

🛠️ Lý do chi tiết:

  • Data Catalog tự động catalog BigQuery datasets và Pub/Sub topics (từ phiên bản cập nhật 2023-2026, hỗ trợ native entry cho Pub/Sub topics qua Cloud Pub/Sub integration).
  • Với PostgreSQL trên Compute Engine (không phải dịch vụ managed như Cloud SQL), không có auto-catalog, nên dùng Data Catalog APIs để manual catalog tables một cách đơn giản, không cần custom code phức tạp.
  • Giải pháp này tối ưu nỗ lực: Auto cho 2/3 nguồn, manual nhẹ cho PostgreSQL, phù hợp best practice GCP đến 2026.

📘 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm lý do bằng tiếng Việt dựa trên tài liệu GCP mới nhất (2026):

  • ❌ [SAI] Use Data Catalog to automatically catalog BigQuery datasets. Use Data Catalog APIs to manually catalog Pub/Sub topics and PostgreSQL tables.
    Sai vì Pub/Sub topics được tự động catalog bởi Data Catalog (không cần manual APIs). Manual Pub/Sub sẽ tăng nỗ lực không cần thiết, vi phạm yêu cầu "minimum efforts".

  • ✅ [ĐÚNG] Use Data Catalog to automatically catalog BigQuery datasets and Pub/Sub topics. Use Data Catalog APIs to manually catalog PostgreSQL tables.
    Đúng hoàn toàn như giải thích ở trên: Tận dụng auto-integration cho BigQuery & Pub/Sub, chỉ manual nhẹ APIs cho PostgreSQL (hỗ trợ qua Tag API hoặc Entry API cho custom metadata).

  • ❌ [SAI] Use Data Catalog to automatically catalog BigQuery datasets and Pub/Sub topics. Use custom connectors to manually catalog PostgreSQL tables.
    Sai vì custom connectors yêu cầu phát triển code phức tạp (dùng Connector Development Kit), trong khi Data Catalog APIs đơn giản hơn nhiều (REST/gRPC calls). Không tối ưu "minimum configuration".

  • ❌ [SAI] Use customer connectors to manually catalog BigQuery datasets, Pub/Sub topics, and PostgreSQL tables.
    Sai toàn bộ vì bỏ qua auto-catalog native cho BigQuery & Pub/Sub. "Customer connectors" (có lẽ lỗi đánh máy của "custom connectors") sẽ làm tăng nỗ lực phát triển cao, không hiệu quả.

📚 Tài liệu tham khảo (cập nhật đến 2026)

Giải pháp này giúp tổ chức scale data discovery an toàn, tuân thủ IAM & metadata lineage! 🚀

Câu 370
You need to create a SQL pipeline. The pipeline runs an aggregate SQL transformation on a BigQuery table every two hours and appends the result to another existing BigQuery table. You need to configure the pipeline to retry if errors occur. You want the pipeline to send an email notification after three consecutive failures. What should you do?
  1. A Use the BigQueryUpsertTableOperator in Cloud Composer, set the retry parameter to three, and set the email_on_failure parameter to true.
  2. B Use the BigQueryInsertJobOperator in Cloud Composer, set the retry parameter to three, and set the email_on_failure parameter to true.
  3. C Create a BigQuery scheduled query to run the SQL transformation with schedule options that repeats every two hours, and enable email notifications.
  4. D Create a BigQuery scheduled query to run the SQL transformation with schedule options that repeats every two hours, and enable notification to Pub/Sub topic. Use Pub/Sub and Cloud Functions to send an email after three failed executions.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi yêu cầu xây dựng một SQL pipeline trên Google Cloud Platform (GCP), cụ thể là:

  • Chạy aggregate SQL transformation (chuyển đổi tổng hợp SQL) trên một bảng BigQuery mỗi 2 giờ.
  • Append (thêm dữ liệu vào cuối) kết quả vào một bảng BigQuery hiện có khác.
  • Pipeline phải retry (thử lại) nếu xảy ra lỗi.
  • Gửi email notification sau 3 lần thất bại liên tiếp.

Mục tiêu là chọn giải pháp tự động hóa, đáng tin cậy sử dụng các dịch vụ GCP như Cloud Composer (dựa trên Apache Airflow) hoặc BigQuery scheduled queries, với cơ chế retry và thông báo chính xác. 📊🕒

✅ Đáp án đúng

Use the BigQueryInsertJobOperator in Cloud Composer, set the retry parameter to three, and set the email_on_failure parameter to true.

Lý do lựa chọn:

  • BigQueryInsertJobOperator là operator chuẩn trong Cloud Composer (Apache Airflow trên GCP) để chạy SQL query tùy chỉnh trên BigQuery, hỗ trợ INSERT (append) dữ liệu vào bảng đích. Nó phù hợp hoàn hảo cho aggregate transformation và append mỗi 2 giờ qua DAG scheduler (schedule_interval='0 */2 * * *').
  • retry=3: Tham số của BaseOperator, tự động retry 3 lần nếu job thất bại (do lỗi tạm thời như quota, network).
  • email_on_failure=true: Gửi email ngay sau mỗi failure cuối cùng (sau 3 retries thất bại), khớp yêu cầu "sau three consecutive failures". Cloud Composer hỗ trợ email qua Airflow config (SMTP hoặc SendGrid). 🛠️
  • Giải pháp đơn giản, scalable, tích hợp retry/notification built-in mà không cần code thêm. Phù hợp kiến thức GCP mới nhất (Airflow 2.7+ trong Cloud Composer 3.x đến 2026).

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn một cách chi tiết:

  • Use the BigQueryUpsertTableOperator in Cloud Composer, set the retry parameter to three, and set the email_on_failure parameter to true.
    ❌ Sai. Operator BigQueryUpsertTableOperator không phải operator chuẩn trong Apache Airflow providers-google (chỉ có BigQueryCreateEmptyTableOperator hoặc BigQueryInsertJobOperator cho SQL jobs). Nó không hỗ trợ chạy aggregate SQL transformation tùy chỉnh và append; chủ yếu dùng cho upsert schema/table, không phù hợp pipeline SQL. Retry/email đúng nhưng operator sai dẫn đến không thực hiện được task. 🛑

  • Use the BigQueryInsertJobOperator in Cloud Composer, set the retry parameter to three, and set the email_on_failure parameter to true.
    ✅ Đúng. Như giải thích ở trên: Operator lý tưởng cho chạy SQL aggregate và INSERT/append, kết hợp DAG schedule mỗi 2 giờ, retry 3 lần, email sau failures. Hoàn hảo khớp yêu cầu! 🎯

  • Create a BigQuery scheduled query to run the SQL transformation with schedule options that repeats every two hours, and enable email notifications.
    ❌ Sai. BigQuery scheduled queries hỗ trợ chạy SQL định kỳ (mỗi 2 giờ) và append qua write_disposition='WRITE_APPEND', nhưng không có cơ chế retry tự động (chỉ rerun thủ công hoặc qua UI). Notification chỉ gửi mỗi failure ngay lập tức, không đếm "three consecutive failures" (BigQuery alerting chỉ basic, không track liên tiếp). Không linh hoạt cho pipeline phức tạp. ⏰

  • Create a BigQuery scheduled query to run the SQL transformation with schedule options that repeats every two hours, and enable notification to Pub/Sub topic. Use Pub/Sub and Cloud Functions to send an email after three failed executions.
    ❌ Sai. Scheduled query + Pub/Sub notification đúng cho alert, nhưng không có retry built-in (phải tự implement retry qua Dataflow/Composer). Phần Cloud Functions đếm 3 failures yêu cầu code custom phức tạp (state management với Firestore/Redis), không đơn giản và dễ lỗi. Quá rườm rà so với Composer native. 🔄

📘 Tài liệu tham khảo (cập nhật đến 2026)

Giải pháp này đảm bảo pipeline robust, observable trên GCP! 🚀