Ngân hàng đề — Google Cloud Professional Data Engineer
Tìm thấy 429 câu.
-
A
1. Check for duplicate rows in the BigQuery tables that have the daily partition data size doubled.
2. Schedule daily SQL jobs to deduplicate the affected tables.
3. Share the deduplication script with the other operational teams to reuse if this occurs to other tables. -
B
1. Check for code errors in the deployed pipelines.
2. Check for multiple writing to pipeline BigQuery sink.
3. Check for errors in Cloud Logging during the day of the release of the new pipelines.
4. If no errors, restore the BigQuery tables to their content before the last release by using time travel. -
C
1. Check for duplicate rows in the BigQuery tables that have the daily partition data size doubled.
2. Check the BigQuery Audit logs to find job IDs.
3. Use Cloud Monitoring to determine when the identified Dataflow jobs started and the pipeline code version.
4. When more than one pipeline ingests data into a table, stop all versions except the latest one. -
D
1. Roll back the last deployment.
2. Restore the BigQuery tables to their content before the last release by using time travel.
3. Restart the Dataflow jobs and replay the messages by seeking the subscription to the timestamp of the release.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả tình huống giám sát data lake được lưu trữ trên BigQuery trong tổ chức. Các ingestion pipelines đọc dữ liệu từ Pub/Sub và ghi dữ liệu vào các bảng trên BigQuery. Sau khi triển khai phiên bản mới của pipelines, lượng dữ liệu lưu trữ hàng ngày tăng 50%, mặc dù lượng dữ liệu trong Pub/Sub không thay đổi và chỉ một số bảng có kích thước partition dữ liệu hàng ngày bị nhân đôi. Nhiệm vụ là điều tra và khắc phục nguyên nhân gây tăng dữ liệu.
🔍 Vấn đề cốt lõi:
- Dữ liệu đầu vào (Pub/Sub) không đổi, nhưng dữ liệu lưu trữ tăng → Có khả năng duplicate data (dữ liệu trùng lặp) ở một số bảng cụ thể.
- Nguyên nhân có thể từ multiple pipelines (nhiều pipeline chạy song song) ghi vào cùng bảng, ví dụ: pipeline cũ chưa được dừng sau khi deploy phiên bản mới (thường gặp với Dataflow jobs xử lý từ Pub/Sub sang BigQuery).
- Cần xác định nguyên nhân gốc rễ (root cause) và fix triệt để, không chỉ dọn dẹp triệu chứng.
📘 Kiến thức liên quan (cập nhật đến 2026):
- BigQuery hỗ trợ daily partitions cho dữ liệu thời gian, dễ kiểm tra kích thước partition qua INFORMATION_SCHEMA.
- BigQuery Audit Logs ghi lại job IDs của các write operations.
- Cloud Monitoring theo dõi metrics của Dataflow như job start time, version code (qua labels/metadata).
- Dataflow pipelines từ Pub/Sub to BigQuery có thể chạy multiple instances nếu không quản lý deployment đúng (ví dụ: autoscaling hoặc fail-over).
- Tham khảo: BigQuery Audit Logs, Dataflow Monitoring, Cloud Monitoring for Dataflow.
✅ Đáp án đúng: Lựa chọn 3
1. Check for duplicate rows in the BigQuery tables that have the daily partition data size doubled.
2. Check the BigQuery Audit logs to find job IDs.
3. Use Cloud Monitoring to determine when the identified Dataflow jobs started and the pipeline code version.
4. When more than one pipeline ingests data into a table, stop all versions except the latest one.
Lý do chọn đáp án này 🛠️:
- Đây là quy trình điều tra toàn diện và fix root cause theo best practices GCP.
- Bước 1: Xác nhận duplicate rows ở tables affected (sử dụng SQL query trên partitions).
- Bước 2: Audit logs giúp trace job IDs từ các write operations vào BigQuery (lọc theo thời gian deploy mới).
- Bước 3: Cloud Monitoring xác định Dataflow jobs nào đang chạy (start time, code version qua job labels), chứng minh multiple pipelines parallel (old + new).
- Bước 4: Stop các pipeline cũ, giữ latest → Fix vĩnh viễn, tránh duplicate tiếp tục.
- Phù hợp với nguyên nhân phổ biến: Deploy mới mà quên dừng old Dataflow templates/jobs (Dataflow streaming jobs persist until cancelled).
❌ Giải thích tất cả các phương án
-
Lựa chọn 1 (SAI):
1. Check for duplicate rows in the BigQuery tables that have the daily partition data size doubled. 2. Schedule daily SQL jobs to deduplicate the affected tables. 3. Share the deduplication script with the other operational teams to reuse if this occurs to other tables.Tại sao sai 🚫: Chỉ tập trung dọn dẹp triệu chứng (deduplicate) mà không điều tra/fix nguyên nhân gốc (như multiple pipelines). Schedule daily jobs tốn kém (BigQuery slot usage cao), không scalable, và có nguy cơ mất dữ liệu nếu dedup sai. Không giải quyết vấn đề lâu dài.
-
Lựa chọn 2 (SAI):
1. Check for code errors in the deployed pipelines. 2. Check for multiple writing to pipeline BigQuery sink. 3. Check for errors in Cloud Logging during the day of the release of the new pipelines. 4. If no errors, restore the BigQuery tables to their content before the last release by using time travel.Tại sao sai 🚫:
- Bước 1-3 chỉ kiểm tra lỗi code/logging, nhưng vấn đề không nhất thiết là lỗi (có thể là config/deployment miss).
- Bước 4: Time travel chỉ hỗ trợ 7 ngày (không đủ nếu incident kéo dài), và restore không fix cause → duplicate sẽ tái diễn ngay sau restore. Không trace Dataflow jobs cụ thể.
-
Lựa chọn 3 (ĐÚNG): ✅ (Đã giải thích ở trên).
-
Lựa chọn 4 (SAI):
1. Roll back the last deployment. 2. Restore the BigQuery tables to their content before the last release by using time travel. 3. Restart the Dataflow jobs and replay the messages by seeking the subscription to the timestamp of the release.Tại sao sai 🚫:
- Rollback/restore không xác định cause, giả sử lỗi ở deployment mới (có thể không phải).
- Seeking subscription trên Pub/Sub chỉ rewind cho acknowledged messages (mất nếu đã ack), không phù hợp replay chính xác, dễ gây thêm duplicate hoặc data loss. Không dùng Cloud Monitoring/Audit logs để trace jobs.
🔄 Khuyến nghị bổ sung: Sau fix, implement CI/CD với auto-stop old jobs qua Cloud Build/Deploy, và set alerts trên Cloud Monitoring cho Dataflow job counts. Tham khảo: Dataflow Best Practices.
- A Create the “gdpr” tag template with private visibility. Assign the bigquery.dataViewer role to the HR group on the tables that contain sensitive data.
- B Create the “gdpr” tag template with private visibility. Assign the datacatalog.tagTemplateViewer role on this tag to the all employees group, and assign the bigquery.dataViewer role to the HR group on the tables that contain sensitive data.
- C Create the “gdpr” tag template with public visibility. Assign the bigquery.dataViewer role to the HR group on the tables that contain sensitive data.
- D Create the “gdpr” tag template with public visibility. Assign the datacatalog.tagTemplateViewer role on this tag to the all employees group, and assign the bigquery.dataViewer role to the HR group on the tables that contain sensitive data.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào việc quản lý quyền truy cập dữ liệu nhạy cảm trong Google Cloud BigQuery sử dụng Data Catalog tag templates. Cụ thể:
- Có dataset BigQuery tên "customers", tất cả bảng sẽ được gắn tag từ template "gdpr" với trường bắt buộc "has_sensitive_data" (giá trị boolean: true/false).
- Yêu cầu chính:
- Tất cả nhân viên (all employees group) phải có thể tìm kiếm đơn giản (search) và xác định các bảng có giá trị true hoặc false ở trường này → Họ cần xem metadata của tags mà không cần truy cập dữ liệu thực tế.
- Chỉ nhóm HR mới được xem dữ liệu bên trong (data inside) các bảng có "has_sensitive_data" = true.
- Đã cấp quyền cơ bản: Nhóm all employees có bigquery.metadataViewer (xem metadata) và bigquery.connectionUser (kết nối) trên dataset.
- Mục tiêu: Giảm thiểu overhead cấu hình (minimize configuration overhead) → Tránh cấp quyền phức tạp, lặp lại trên nhiều bảng.
- Bối cảnh cập nhật 2026: Dựa trên tài liệu GCP mới nhất (BigQuery & Data Catalog v2024-2026), tag templates có public/private visibility. bigquery.metadataViewer cho phép xem metadata/tags nếu template public; quyền dữ liệu tách biệt với metadata. 📘 Nguồn: GCP Data Catalog Docs, BigQuery IAM.
✅ Đáp án đúng
Create the “gdpr” tag template with public visibility. Assign the bigquery.dataViewer role to the HR group on the tables that contain sensitive data.
Lý do chọn:
- Public visibility cho tag template: Với quyền bigquery.metadataViewer đã cấp, tất cả nhân viên có thể search metadata/tags (bao gồm field "has_sensitive_data" true/false) trên các bảng mà không cần quyền thêm → Đơn giản, giảm overhead. 🛡️
- bigquery.dataViewer chỉ cấp cho HR group trên các bảng sensitive (true): Đảm bảo chỉ HR xem dữ liệu thực, tách biệt hoàn toàn metadata (ai cũng xem) và data (chỉ HR). Không ảnh hưởng bảng false. ⚡
- Minimize overhead: Chỉ 2 bước (tạo public template + cấp dataViewer cho HR trên bảng cần), không cần config per-tag hay per-user phức tạp. Hoàn hảo cho scale lớn! 🚀
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên IAM & Data Catalog GCP (cập nhật 2026).
-
❌ [SAI] Create the “gdpr” tag template with private visibility. Assign the bigquery.dataViewer role to the HR group on the tables that contain sensitive data.
Lý do sai: Visibility private làm tag không hiển thị với tất cả nhân viên (dù có metadataViewer), họ không search được field "has_sensitive_data" → Vi phạm yêu cầu search cho mọi người. Phải config thêm quyền xem tag riêng → Tăng overhead lớn. Không minimize! 😞 Nguồn: Data Catalog visibility rules. -
❌ [SAI] Create the “gdpr” tag template with private visibility. Assign the datacatalog.tagTemplateViewer role on this tag to the all employees group, and assign the bigquery.dataViewer role to the HR group on the tables that contain sensitive data.
Lý do sai: Private vẫn che giấu tag entries (tags trên bảng) với nhân viên; datacatalog.tagTemplateViewer chỉ xem template definition (cấu trúc), không xem giá trị field trên bảng (cần entry-level perms). Phải cấp thêm datacatalog.entriesViewer hoặc tương tự per-tag → Overhead cao, phức tạp scale. Không hiệu quả! 🔒 Nguồn: Tag permissions docs. -
✅ [ĐÚNG] Create the “gdpr” tag template with public visibility. Assign the bigquery.dataViewer role to the HR group on the tables that contain sensitive data.
Lý do đúng: Như phần đáp án trên – Public + metadataViewer = search dễ dàng cho tất cả; dataViewer chỉ HR trên bảng true → Tách biệt metadata/data, zero overhead thêm cho search. Lý tưởng! 🌟 Nguồn: BigQuery metadata access. -
❌ [SAI] Create the “gdpr” tag template with public visibility. Assign the datacatalog.tagTemplateViewer role on this tag to the all employees group, and assign the bigquery.dataViewer role to the HR group on the tables that contain sensitive data.
Lý do sai: Public visibility đã đủ cho search với metadataViewer (không cần tagTemplateViewer thừa thãi) → Cấp thêm role này là dư thừa, tăng overhead không cần thiết. Không minimize config! 🤦♂️ Nguồn: IAM least-privilege principle.
Kết luận: Phương án đúng tận dụng public visibility để leverage quyền metadata sẵn có, chỉ bổ sung data access targeted. Hoàn hảo cho governance GDPR-like trong GCP! 🏆 Nếu cần lab thực hành, dùng GCP console tạo tag template public nhé! 🔧
-
A
1. Use Cloud Build to copy the code of the DAG to the Cloud Storage bucket of the development instance for DAG testing.
2. If the tests pass, use Cloud Build to copy the code to the bucket of the production instance. -
B
1. Use Cloud Build to build a container with the code of the DAG and the KubernetesPodOperator to deploy the code to the Google Kubernetes Engine (GKE) cluster of the development instance for testing.
2. If the tests pass, use the KubernetesPodOperator to deploy the container to the GKE cluster of the production instance. -
C
1. Use Cloud Build to build a container and the KubernetesPodOperator to deploy the code of the DAG to the Google Kubernetes Engine (GKE) cluster of the development instance for testing.
2. If the tests pass, copy the code to the Cloud Storage bucket of the production instance. -
D
1. Use Cloud Build to copy the code of the DAG to the Cloud Storage bucket of the development instance for DAG testing.
2. If the tests pass, use Cloud Build to build a container with the code of the DAG and the KubernetesPodOperator to deploy the container to the Google Kubernetes Engine (GKE) cluster of the production instance.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào việc thiết lập quy trình CI/CD (Continuous Integration/Continuous Deployment) cho mã nguồn của các DAGs (Directed Acyclic Graphs) chạy trên Cloud Composer (dịch vụ quản lý Apache Airflow trên Google Cloud). Đội ngũ có hai môi trường Cloud Composer riêng biệt: một cho development (dev) để kiểm tra và một cho production (prod) để triển khai thực tế. Mã nguồn DAGs được lưu trữ và phát triển trên Git repository. Yêu cầu chính là tự động triển khai DAGs vào Cloud Composer khi một tag cụ thể được push lên Git repo.
🛠️ Các khái niệm chính cần nắm:
- Cloud Composer hoạt động bằng cách đồng bộ hóa các file DAG từ thư mục
dags/trong Cloud Storage bucket gắn liền với mỗi environment (không phải deploy trực tiếp qua Kubernetes). - Cloud Build là dịch vụ CI/CD của Google Cloud, có thể trigger tự động từ Git events (như push tag), và hỗ trợ các bước như copy file vào GCS.
- Quy trình mong muốn: Test trên dev trước → Nếu pass → Deploy sang prod.
- Lưu ý cập nhật 2026: Theo tài liệu mới nhất của Google Cloud Composer 2.x (phiên bản hiện hành đến 2026), DAGs vẫn được deploy bằng cách upload file vào GCS bucket của environment, không qua container hoặc Pod trực tiếp (📘 Tài liệu chính thức: Deploy DAGs in Cloud Composer).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng là phương án đầu tiên (đã đánh dấu [ĐÚNG]):
- Use Cloud Build to copy the code of the DAG to the Cloud Storage bucket of the development instance for DAG testing.
- If the tests pass, use Cloud Build to copy the code to the bucket of the production instance.
Lý do chọn ✅:
- Đây là cách chuẩn và đơn giản nhất để deploy DAGs vào Cloud Composer: Cloud Build trigger từ Git tag push, copy code trực tiếp vào GCS bucket của environment dev để test (Composer tự sync DAGs từ bucket).
- Nếu test pass (ví dụ: qua các bước validation trong Cloud Build như linting, unit test), tiếp tục copy sang bucket của prod. Quy trình này tự động, an toàn, tách biệt môi trường, phù hợp với best practices CI/CD cho Composer (không cần container phức tạp).
- 📘 Nguồn: Cloud Composer DAG deployment guide và Cloud Build triggers (cập nhật 2026).
📋 Phân tích chi tiết tất cả các phương án
Dưới đây là phân tích từng phương án một cách rõ ràng, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai kèm lý do cụ thể dựa trên kiến thức Cloud Composer mới nhất.
-
✅ Phương án ĐÚNG (Phương án 1):
- Use Cloud Build to copy the code of the DAG to the Cloud Storage bucket of the development instance for DAG testing.
- If the tests pass, use Cloud Build to copy the code to the bucket of the production instance.
Giải thích: Hoàn toàn chính xác vì deploy DAGs chỉ cần copy file vào GCS bucket của từng environment. Cloud Build xử lý trigger từ tag push và các bước test/copy một cách hiệu quả, đảm bảo tách biệt dev/prod. Đây là phương pháp được Google khuyến nghị (không rườm rà).
-
❌ Phương án SAI (Phương án 2):
- Use Cloud Build to build a container with the code of the DAG and the KubernetesPodOperator to deploy the code to the Google Kubernetes Engine (GKE) cluster of the development instance for testing.
- If the tests pass, use the KubernetesPodOperator to deploy the container to the GKE cluster of the production instance.
Giải thích: Sai hoàn toàn vì DAGs không deploy bằng container + KubernetesPodOperator trực tiếp vào GKE cluster của Composer. Composer sử dụng GKE nội bộ để chạy Airflow scheduler/workers, nhưng DAG code phải ở GCS bucket, không phải Pod/container. Cách này phức tạp, không chuẩn và có thể gây lỗi sync DAGs.
-
❌ Phương án SAI (Phương án 3):
- Use Cloud Build to build a container and the KubernetesPodOperator to deploy the code of the DAG to the Google Kubernetes Engine (GKE) cluster of the development instance for testing.
- If the tests pass, copy the code to the Cloud Storage bucket of the production instance.
Giải thích: Sai ở bước 1 vì lại dùng container + KubernetesPodOperator cho dev (không đúng cách deploy DAGs vào Composer). Bước 2 đúng (copy sang prod bucket), nhưng tổng thể không nhất quán và không khả thi cho dev testing. Cloud Composer không hỗ trợ deploy DAG qua Pod như vậy.
-
❌ Phương án SAI (Phương án 4):
- Use Cloud Build to copy the code of the DAG to the Cloud Storage bucket of the development instance for DAG testing.
- If the tests pass, use Cloud Build to build a container with the code of the DAG and the KubernetesPodOperator to deploy the container to the Google Kubernetes Engine (GKE) cluster of the production instance.
Giải thích: Sai ở bước 2 vì bước 1 đúng (copy dev bucket), nhưng prod lại dùng container + Pod (không chuẩn). DAGs prod vẫn cần GCS bucket, không phải deploy container vào GKE. Cách này làm phức tạp hóa prod deployment không cần thiết.
🧩 Tóm tắt: Phương án đúng tận dụng GCS bucket sync – core mechanism của Cloud Composer – kết hợp Cloud Build cho CI/CD tự động. Các phương án sai nhầm lẫn với deploy Kubernetes trực tiếp, không phù hợp! (📘 Tham khảo thêm: Cloud Composer architecture).
- A Use Cloud KMS encryption key with Dataflow to ingest the existing Pub/Sub subscription to the existing BigQuery table.
- B Create a new BigQuery table by using customer-managed encryption keys (CMEK), and migrate the data from the old BigQuery table.
- C Create a new Pub/Sub topic with CMEK and use the existing BigQuery table by using Google-managed encryption key.
- D Create a new BigQuery table and Pub/Sub topic by using customer-managed encryption keys (CMEK), and migrate the data from the old BigQuery table.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm
📘 Nội dung câu hỏi:
Câu hỏi mô tả một tình huống thực tế trong Google Cloud: Bạn có một bảng BigQuery đang nhận dữ liệu trực tiếp từ một subscription Pub/Sub (quá trình ingestion streaming). Dữ liệu hiện tại được mã hóa tại chỗ (at-rest) bằng khóa mã hóa do Google quản lý (Google-managed encryption key). Bây giờ, tổ chức áp dụng chính sách mới yêu cầu sử dụng khóa từ một dự án Cloud Key Management Service (Cloud KMS) tập trung để mã hóa dữ liệu tại chỗ.
Mục tiêu chính là thay đổi cơ chế mã hóa cho dữ liệu tại chỗ trong BigQuery table, vì chính sách tập trung vào "encrypt data at rest" (dữ liệu lưu trữ tĩnh). Lưu ý:
- BigQuery hỗ trợ Customer-Managed Encryption Keys (CMEK) từ Cloud KMS, nhưng không thể thay đổi khóa mã hóa cho bảng đã tồn tại sau khi tạo.
- Pub/Sub chỉ là nguồn dữ liệu streaming, mã hóa at-rest của nó không phải trọng tâm chính sách (chính sách nhắm đến BigQuery).
- Centralized KMS project nghĩa là khóa từ dự án KMS riêng biệt, cần cấp quyền IAM phù hợp (như
cloudkms.cryptoKeyEncrypterDecrypter).
(Kiến thức cập nhật đến 2026: BigQuery vẫn yêu cầu tạo table mới cho CMEK, theo docs Google Cloud BigQuery encryption - phiên bản mới nhất hỗ trợ CMEK cho routine tables và materialized views).
✅ Đáp án đúng:
Create a new BigQuery table by using customer-managed encryption keys (CMEK), and migrate the data from the old BigQuery table.
🛠️ Lý do chọn đáp án đúng:
- Đây là cách duy nhất và tối ưu để áp dụng CMEK từ dự án KMS tập trung cho dữ liệu at-rest trong BigQuery.
- Quy trình: Tạo bảng BigQuery mới với tham số
kms_key_namechỉ định khóa KMS từ dự án tập trung → Sử dụngbq cphoặc Data Transfer Service để migrate dữ liệu từ bảng cũ → Cập nhật subscription Pub/Sub để push dữ liệu vào bảng mới. - Không ảnh hưởng đến Pub/Sub hiện tại, tiết kiệm chi phí và thời gian.
- Tuân thủ chính sách: Dữ liệu mới và cũ đều được mã hóa bằng CMEK.
(Nguồn: BigQuery Customer-managed encryption keys (CMEK) & Migrate BigQuery tables).
🔍 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] Use Cloud KMS encryption key with Dataflow to ingest the existing Pub/Sub subscription to the existing BigQuery table.
Phương án này không giải quyết vấn đề mã hóa at-rest của bảng BigQuery hiện tại. Dataflow chỉ mã hóa dữ liệu trong pipeline (transient data), không thay đổi khóa CMEK cho bảng đích (vẫn dùng Google-managed key). Existing table không hỗ trợ thay đổi khóa sau khi tạo, dẫn đến vi phạm chính sách. -
✅ [ĐÚNG] Create a new BigQuery table by using customer-managed encryption keys (CMEK), and migrate the data from the old BigQuery table.
Như đã giải thích ở trên: Hoàn hảo, tạo bảng mới với CMEK từ KMS tập trung, migrate dữ liệu, và redirect ingestion từ Pub/Sub. Đơn giản, hiệu quả, không thừa thãi. -
❌ [SAI] Create a new Pub/Sub topic with CMEK and use the existing BigQuery table by using Google-managed encryption key.
Phương án này vô ích vì vẫn dùng existing BigQuery table với Google-managed key, không đáp ứng chính sách mã hóa at-rest cho dữ liệu trong BigQuery. Pub/Sub CMEK chỉ ảnh hưởng dữ liệu tạm thời trong topic/subscription, không liên quan đến data at-rest chính. -
❌ [SAI] Create a new BigQuery table and Pub/Sub topic by using customer-managed encryption keys (CMEK), and migrate the data from the old BigQuery table.
Thừa thãi và phức tạp không cần thiết. Tạo new Pub/Sub topic với CMEK là ok nhưng không bắt buộc (Pub/Sub không phải trọng tâm chính sách at-rest BigQuery). Việc migrate chỉ cần cho BigQuery table, không cần recreate toàn bộ Pub/Sub, dẫn đến downtime và chi phí cao hơn.
📚 Tài liệu tham khảo chính (cập nhật 2026):
- BigQuery encryption at rest ✅
- Pub/Sub to BigQuery streaming
- Cloud KMS for CMEK
- Best practices: Sử dụng
gcloud bigquery tables createvới--kms-keycho table mới.
Hy vọng phân tích này giúp bạn ôn thi Google Cloud Professional Data Engineer hiệu quả! 🚀
- A Import the ORC files to Bigtable tables for the data scientist team.
- B Import the ORC files to BigQuery tables for the data scientist team.
- C Copy the ORC files on Cloud Storage, then deploy a Dataproc cluster for the data scientist team.
- D Copy the ORC files on Cloud Storage, then create external BigQuery tables for the data scientist team.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc xây dựng một môi trường phân tích dữ liệu (analytics environment) trên Google Cloud để đội ngũ data scientist có thể khám phá dữ liệu mà không ảnh hưởng đến hệ thống Apache Hadoop on-premises. Dữ liệu gốc nằm trong HDFS cluster dưới định dạng ORC (Optimized Row Columnar) với nhiều cột phân vùng Hive (Hive partitioning). Yêu cầu chính là cho phép đội ngũ explore dữ liệu tương tự như trên HDFS on-prem, sử dụng SQL trên Hive query engine. Giải pháp cần cost-effective nhất cho storage và processing, nghĩa là ưu tiên chi phí thấp, không cần import dữ liệu (để tránh tốn kém copy/transform), và hỗ trợ query SQL nhanh chóng với partitioning giống Hive.
📌 Yêu cầu cốt lõi:
- Giữ nguyên định dạng ORC và Hive partitioning.
- Query SQL trực tiếp như Hive.
- Cost-effective: Tránh cluster luôn chạy (như Dataproc) hoặc import dữ liệu (tốn storage/compute).
Dựa trên kiến thức Google Cloud cập nhật đến 2026 (BigQuery hỗ trợ external tables cho ORC từ lâu, với schema auto-detection và Hive partitioning đầy đủ qua HiveQL-like syntax).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Copy the ORC files on Cloud Storage, then create external BigQuery tables for the data scientist team.
Lý do chi tiết 🛠️:
- Copy ORC files lên Cloud Storage (GCS): GCS là storage rẻ nhất, scalable, hỗ trợ định dạng ORC native. Không cần transform dữ liệu, giữ nguyên Hive partitioning (như
PARTITIONED BY (date, region)). - Tạo external BigQuery tables: BigQuery cho phép query external tables trực tiếp trên GCS mà không import dữ liệu (zero-copy), hỗ trợ ORC format đầy đủ (columnar, compression). BigQuery tự detect schema, partitioning (Hive-style), và partition pruning tự động – giống hệt Hive query engine.
- Cost-effective nhất 💰:
- Storage: Chỉ tốn GCS (~$0.023/GB/tháng).
- Query: On-demand pricing (~$5/TB scanned), serverless (pay-per-query), không cần cluster luôn chạy.
- So với on-prem Hive: Query nhanh hơn nhờ BigQuery engine (Dremel), hỗ trợ SQL chuẩn ANSI + HiveQL extensions.
- Phù hợp data scientist: Explore ad-hoc SQL queries, không impact HDFS on-prem.
❌ Phân tích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên tính phù hợp, cost, và khả năng hỗ trợ ORC + Hive partitioning + SQL-like Hive.
-
[SAI] Import the ORC files to Bigtable tables for the data scientist team.
❌ Lý do sai: Bigtable là NoSQL wide-column store, không hỗ trợ ORC format (chỉ native Bigtable format). Không có SQL query engine (dùng HBase-like API), không hỗ trợ Hive partitioning. Import sẽ yêu cầu transform toàn bộ dữ liệu (tốn kém compute/storage), và không phù hợp analytics SQL exploration. Bigtable dành cho high-throughput OLTP, không phải batch analytics như Hive. -
[SAI] Import the ORC files to BigQuery tables for the data scientist team.
❌ Lý do sai: Mặc dù BigQuery hỗ trợ ORC (import được), nhưng import vào native tables yêu cầu copy dữ liệu vào BigQuery storage (tốn ~$0.02/GB/tháng + load job chi phí). Không cost-effective bằng external tables (vì duplicate storage). External tables mới là optimal cho trường hợp giữ nguyên files trên GCS mà vẫn query như Hive. -
[SAI] Copy the ORC files on Cloud Storage, then deploy a Dataproc cluster for the data scientist team.
❌ Lý do sai: Copy lên GCS OK, nhưng deploy Dataproc cluster (managed Hadoop/Spark) tốn kém vì cluster luôn chạy (master/worker nodes ~$0.1-1/GiB/giờ). Data scientist phải manage cluster (scale up/down), và dùng Hive-on-Dataproc để query – giống on-prem nhưng không cost-effective (pay idle time). BigQuery external tables serverless, nhanh hơn (query seconds vs minutes). -
[ĐÚNG] Copy the ORC files on Cloud Storage, then create external BigQuery tables for the data scientist team.
✅ Lý do đúng (như phần trên): Optimal cost, native hỗ trợ ORC/Hive partitioning, SQL query giống Hive, serverless.
📘 Tài liệu tham khảo (cập nhật 2026)
- BigQuery External Tables: Cloud BigQuery: External data sources – Hỗ trợ ORC, Hive partitioning từ 2018+, với auto-schema detection.
- ORC in BigQuery: Query ORC files – Partition pruning, columnar read.
- Cost comparison: BigQuery Pricing vs Dataproc Pricing.
- Best practices: Google Cloud Architecture Framework – "Decouple storage from compute with external tables".
Giải pháp này đảm bảo explore dữ liệu linh hoạt, rẻ tiền, không downtime on-prem! 🚀
- A Submit duplicate pipelines in two different zones by using the --zone flag.
- B Set the pipeline staging location as a regional Cloud Storage bucket.
- C Specify a worker region by using the --region flag.
- D Create an Eventarc trigger to resubmit the job in case of zonal failure when submitting the job.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc thiết kế một Dataflow pipeline (dịch vụ xử lý dữ liệu batch trên Google Cloud) cho công việc xử lý batch (batch processing job). Mục tiêu là giảm thiểu (mitigate) rủi ro từ nhiều sự cố zonal failures (lỗi xảy ra ở một hoặc nhiều zone cụ thể trong region) tại thời điểm submit job (job submission time).
- Zonal failure: Là sự cố chỉ ảnh hưởng đến một zone cụ thể (ví dụ: zone us-central1-a), có thể do vấn đề phần cứng, mạng hoặc bảo trì. Nếu pipeline chỉ chạy ở một zone, nó dễ bị gián đoạn.
- Batch processing: Xử lý dữ liệu theo lô lớn, không phải streaming real-time, nên cần độ tin cậy cao từ đầu.
- Job submission time: Tập trung vào cấu hình ngay khi submit job, không phải runtime monitoring sau đó.
- Yêu cầu chính: Chọn cách tối ưu để Dataflow tự động phân bổ workers (máy ảo xử lý) đa zone trong một region, tránh single point of failure từ zonal issues.
Kiến thức cập nhật đến 2026: Theo tài liệu Google Cloud Dataflow mới nhất (phiên bản 2025-2026), Dataflow hỗ trợ multi-zonal worker pools khi chỉ định --region, giúp tự động mitigate zonal failures mà không cần cấu hình thủ công phức tạp. 📘 Nguồn tham khảo:
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Specify a worker region by using the --region flag.
Lý do:
- Khi submit job Dataflow bằng lệnh
gcloud dataflow jobs submithoặc qua console/CLI, flag--region(ví dụ:--region=us-central1) chỉ định region cụ thể cho workers. Dataflow sẽ tự động phân bổ workers đa zone trong region đó (multi-zonal deployment), giúp chịu được nhiều zonal failures cùng lúc mà không làm job fail hoàn toàn. - Điều này được thực hiện ngay tại submission time, phù hợp chính xác với yêu cầu. Không cần duplicate jobs hay monitoring sau.
- Ưu điểm: Tiết kiệm chi phí, tự động scale, và tuân thủ best practices cho batch jobs lớn. 🛠️
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] Submit duplicate pipelines in two different zones by using the --zone flag.
Phương án này sai vì: Flag--zone(deprecated từ 2023, không khuyến khích dùng) chỉ định zone duy nhất, dẫn đến single zone failure dễ dàng làm job fail. Submit duplicate pipelines tốn kém (double chi phí compute/storage), phức tạp quản lý, và không mitigate multiple zonal failures hiệu quả (nếu cả hai zone cùng fail). Không phải best practice cho submission time. -
❌ [SAI] Set the pipeline staging location as a regional Cloud Storage bucket.
Phương án này sai vì: Staging location (nơi lưu temporary files, JARs, pipelines) dùng regional bucket (như gs://bucket/us-central1/) chỉ tăng availability cho storage (99.99% durability), nhưng không ảnh hưởng đến workers. Zonal failures vẫn làm workers ở zone cụ thể fail, job vẫn crash. Staging chỉ hỗ trợ recovery một phần, không mitigate tại submission time. -
✅ [ĐÚNG] Specify a worker region by using the --region flag.
(Đã giải thích chi tiết ở phần trên). Đây là cách chuẩn và hiệu quả nhất, Dataflow tự handle multi-zonal workers ngay từ đầu. -
❌ [SAI] Create an Eventarc trigger to resubmit the job in case of zonal failure when submitting the job.
Phương án này sai vì: Eventarc (dịch vụ event-driven trên GCP) dùng để trigger actions dựa trên events (như Pub/Sub, Audit Logs), nhưng không hoạt động tại submission time mà chỉ sau khi failure xảy ra (reactive, không preventive). Resubmit thủ công phức tạp, tốn thời gian, và không mitigate multiple failures ngay lập tức. Không phù hợp cho batch jobs cần reliability từ đầu.
Tóm tắt khuyến nghị: Luôn dùng --region cho Dataflow jobs để tối ưu resilience! 🚀 Nếu cần ví dụ code: gcloud dataflow jobs submit ... --region=us-central1 --staging-location=gs://my-bucket/staging/.
- A Group the data by using a tumbling window in a Dataflow pipeline, and write the aggregated data to Memorystore.
- B Group the data by using a hopping window in a Dataflow pipeline, and write the aggregated data to Memorystore.
- C Group the data by using a session window in a Dataflow pipeline, and write the aggregated data to BigQuery.
- D Group the data by using a hopping window in a Dataflow pipeline, and write the aggregated data to BigQuery.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả việc thiết kế một hệ thống thời gian thực (real-time) cho ứng dụng gọi xe (ride hailing app), nhằm xác định các khu vực có nhu cầu cao để điều hướng tài xế kịp thời. Hệ thống hoạt động như sau:
- Nguồn dữ liệu: Dữ liệu từ nhiều nguồn được ingest vào Pub/Sub (dịch vụ messaging thời gian thực của Google Cloud), bao gồm:
- Vị trí tài xế cập nhật mỗi 5 giây.
- Sự kiện đặt xe từ ứng dụng người dùng (riders).
- Xử lý dữ liệu: Thực hiện tổng hợp (aggregation) dữ liệu cung-cầu (supply-demand) cho khoảng thời gian 30 giây gần nhất, lặp lại mỗi 2 giây.
- Lưu trữ: Kết quả cần lưu vào hệ thống low-latency (độ trễ thấp) để hiển thị trên dashboard thời gian thực (visualization và analysis). Yêu cầu chính là chọn phương pháp windowing phù hợp trong Dataflow (dịch vụ stream/batch processing dựa trên Apache Beam) và nơi lưu trữ tối ưu cho truy vấn nhanh. Kiến thức dựa trên Google Cloud Dataflow phiên bản mới nhất (Apache Beam 2.58+ đến 2026), nhấn mạnh windowing cho streaming data với overlap để đảm bảo tính liên tục real-time.
📘 Tài liệu tham khảo:
- Dataflow Windowing Documentation (Apache Beam/Dataflow).
- Memorystore for Redis Overview (low-latency in-memory store).
- BigQuery Streaming Inserts (không phù hợp real-time dashboard).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Group the data by using a hopping window in a Dataflow pipeline, and write the aggregated data to Memorystore.
Lý do:
- Hoppping window (còn gọi là sliding window trong Beam) cho phép window kích thước 30 giây, hop/slide mỗi 2 giây, tạo overlap để tổng hợp dữ liệu "last 30s" liên tục mà không bỏ lỡ dữ liệu (phù hợp với cập nhật 5s của tài xế).
- Memorystore (Redis/Memcached) là in-memory store low-latency (<1ms), lý tưởng cho dashboard real-time visualization (query nhanh, scale cao).
- Phương án này đảm bảo real-time aggregation mượt mà, tránh độ trễ của các lựa chọn khác.
🛠️ Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), với lý do cụ thể dựa trên yêu cầu low-latency real-time và windowing phù hợp.
-
❌ [SAI] Group the data by using a tumbling window in a Dataflow pipeline, and write the aggregated data to Memorystore.
Lý do sai: Tumbling window là fixed-size, non-overlapping (ví dụ: window 30s không overlap). Không thể tổng hợp "mỗi 2 giây" cho last 30s vì các window không slide/overlap, dẫn đến dữ liệu bị "cố định" và bỏ lỡ cập nhật liên tục (như vị trí tài xế 5s). Memorystore đúng nhưng window sai, gây gián đoạn real-time. -
✅ [ĐÚNG] Group the data by using a hopping window in a Dataflow pipeline, and write the aggregated data to Memorystore.
Lý do đúng: Hopping window overlap (size 30s, hop 2s), lý tưởng cho aggregation rolling "last 30s every 2s". Dataflow xử lý streaming hiệu quả với watermarking. Memorystore đảm bảo low-latency read cho dashboard (Redis cluster hỗ trợ pub/sub real-time đến 2026). -
❌ [SAI] Group the data by using a session window in a Dataflow pipeline, and write the aggregated data to BigQuery.
Lý do sai: Session window dựa trên gap hoạt động (không fixed size), không phù hợp cho aggregation fixed "30s every 2s" – có thể tạo window động, không đều đặn. BigQuery là data warehouse với latency cao (giây đến phút cho streaming), không dành cho real-time dashboard (phù hợp batch/analysis hơn). -
❌ [SAI] Group the data by using a hopping window in a Dataflow pipeline, and write the aggregated data to BigQuery.
Lý do sai: Hopping window đúng (overlap cho real-time), nhưng BigQuery không low-latency: streaming insert có delay 1-90 phút cho query, không phù hợp visualization tức thì. Memorystore mới là lựa chọn tối ưu cho cache real-time.
🧠 Lưu ý bổ sung: Trong Dataflow (Beam SDK 2.58+), sử dụng Window.into(SlidingWindows.of(Duration.standardSeconds(30)).every(Duration.standardSeconds(2))) để implement hopping window. Kết hợp với Pub/Sub trigger cho late data handling!
- A Enable retaining of acknowledged messages in your Pub/Sub pull subscription. Use Cloud Monitoring to monitor the subscription/num_retained_acked_messages metric on this subscription.
- B Use an exception handling block in your Dataflow’s DoFn code to push the messages that failed to be transformed through a side output and to a new Pub/Sub topic. Use Cloud Monitoring to monitor the topic/num_unacked_messages_by_region metric on this new topic.
- C Enable dead lettering in your Pub/Sub pull subscription, and specify a new Pub/Sub topic as the dead letter topic. Use Cloud Monitoring to monitor the subscription/dead_letter_message_count metric on your pull subscription.
- D Create a snapshot of your Pub/Sub pull subscription. Use Cloud Monitoring to monitor the snapshot/num_messages metric on this snapshot.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một kịch bản trong Google Cloud Platform (GCP) liên quan đến xử lý dữ liệu streaming thời gian thực từ nhà máy sản xuất ô tô:
- Máy móc đẩy các phép đo (measurements) dưới dạng messages vào một Pub/Sub topic trong dự án GCP.
- Một Dataflow streaming job (viết bằng Apache Beam SDK) đọc các messages này, gửi acknowledgment (ack) cho Pub/Sub (xác nhận đã nhận và xử lý), áp dụng logic kinh doanh tùy chỉnh trong một DoFn instance (hàm xử lý ParDo), rồi ghi kết quả vào BigQuery.
- Yêu cầu chính: Đảm bảo nếu logic kinh doanh thất bại trên một message cụ thể, message đó sẽ được chuyển đến một Pub/Sub topic mới để giám sát và alerting (sử dụng Cloud Monitoring).
📌 Mục tiêu: Xử lý partial failures (thất bại cục bộ) trong DoFn mà không làm gián đoạn toàn bộ pipeline, đồng thời route message lỗi đến topic riêng để theo dõi. Đây là tình huống phổ biến trong streaming pipelines trên Dataflow, nơi cần resilience (khả năng phục hồi) cao. Kiến thức dựa trên Apache Beam 2.58+ và Dataflow runner cập nhật 2024-2026, hỗ trợ side outputs linh hoạt cho error handling.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng:
Use an exception handling block in your Dataflow’s DoFn code to push the messages that failed to be transformed through a side output and to a new Pub/Sub topic. Use Cloud Monitoring to monitor the topic/num_unacked_messages_by_region metric on this new topic.
🛠️ Lý do chi tiết:
- Trong DoFn của Apache Beam/Dataflow, bạn sử dụng try-catch block (xử lý exception) để bắt lỗi trong logic kinh doanh. Khi fail, emit message gốc qua side output (output phụ, tag riêng như "failed").
- Side output cho phép route selective (chọn lọc) messages lỗi mà không ack message gốc ngay, tránh mất dữ liệu. Sau đó, pipe side output đến Pub/SubIO.write để đẩy vào topic mới.
- Giám sát bằng metric topic/num_unacked_messages_by_region trên topic mới (theo vùng, unacked nghĩa là chưa xử lý, phù hợp alerting).
- Đây là best practice chính thức của Google cho custom error handling trong streaming, đảm bảo exactly-once semantics và scalability. Không ảnh hưởng đến ack của main path.
📋 Giải thích tất cả các phương án (đúng/sai)
-
SAI ❌ Enable retaining of acknowledged messages in your Pub/Sub pull subscription. Use Cloud Monitoring to monitor the subscription/num_retained_acked_messages metric on this subscription.
🧠 Phân tích: Tính năng "retaining acknowledged messages" chỉ giữ lại messages đã ack để replay nếu cần, nhưng ở đây Dataflow đã ack ngay sau đọc (trước DoFn), nên message lỗi không còn trong subscription gốc. Metric num_retained_acked_messages chỉ đếm messages đã ack bị giữ lại, không route đến topic mới. Không giải quyết được việc đẩy message lỗi đến topic riêng cho alerting. -
ĐÚNG ✅ Use an exception handling block in your Dataflow’s DoFn code to push the messages that failed to be transformed through a side output and to a new Pub/Sub topic. Use Cloud Monitoring to monitor the topic/num_unacked_messages_by_region metric on this new topic.
🛠️ Phân tích: Như đã giải thích ở trên. Đây là cách tùy chỉnh và chính xác nhất trong Beam/Dataflow, hỗ trợ fanout (phân nhánh) dữ liệu lỗi mà không cần cấu hình subscription. Metric theo dõi unacked messages trên topic mới lý tưởng cho alerting realtime. -
SAI ❌ Enable dead lettering in your Pub/Sub pull subscription, and specify a new Pub/Sub topic as the dead letter topic. Use Cloud Monitoring to monitor the subscription/dead_letter_message_count metric on your pull subscription.
🧠 Phân tích: Dead letter queues (DLQ) hoạt động ở mức subscription (pull mode), chỉ kích hoạt khi message fail sau nhiều retry (maxDeliveryAttempts, mặc định 5) ở tầng subscriber, không phải fail trong DoFn logic tùy chỉnh. Dataflow dùng push subscription nội bộ, và ack xảy ra trước DoFn, nên DLQ không bắt được lỗi business logic. Metric dead_letter_message_count chỉ đếm messages đã dead-lettered, không phù hợp cho alerting chi tiết. -
SAI ❌ Create a snapshot of your Pub/Sub pull subscription. Use Cloud Monitoring to monitor the snapshot/num_messages metric on this snapshot.
🧠 Phân tích: Snapshots lưu trạng thái subscription tại một thời điểm (unacked messages), dùng để restore sau downtime. Không tự động route message lỗi đến topic mới, và không handle failures realtime trong DoFn. Metric snapshot/num_messages chỉ đếm messages trong snapshot, không hỗ trợ alerting động cho business failures.
📘 Tài liệu tham khảo (cập nhật mới nhất 2026)
- Apache Beam Docs: Side Outputs in DoFn & Error Handling.
- Dataflow Docs: Streaming Pipelines Error Handling (hỗ trợ side outputs từ Beam 2.50+).
- Pub/Sub Metrics: Cloud Monitoring Metrics List (num_unacked_messages_by_region từ 2023).
- Official Cert Guide: Google Cloud Professional Data Engineer Study Guide (2024 edition, Wiley).
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần code sample DoFn, hãy hỏi thêm nhé!
- A Give analysts the BigQuery Data Viewer role at the project level. Create one other dataset, and give the analysts the BigQuery Data Editor role on that dataset.
- B Give analysts the BigQuery Data Viewer role at the project level. Create a dataset for each analyst, and give each analyst the BigQuery Data Editor role at the project level.
- C Give analysts the BigQuery Data Viewer role on the shared dataset. Create a dataset for each analyst, and give each analyst the BigQuery Data Editor role at the dataset level for their assigned dataset.
- D Give analysts the BigQuery Data Viewer role on the shared dataset. Create one other dataset and give the analysts the BigQuery Data Editor role on that dataset.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc quản lý quyền truy cập dữ liệu trong Google BigQuery (dịch vụ dữ liệu lớn của Google Cloud). Yêu cầu chính là:
- 📊 Lưu trữ các bảng chia sẻ (shared tables) trong một dataset duy nhất để các analyst dễ dàng truy cập.
- 🔒 Dữ liệu chia sẻ chỉ đọc được (readable) nhưng không chỉnh sửa được (unmodifiable) bởi analysts.
- 🏠 Cung cấp không gian làm việc cá nhân (individual workspaces) trong cùng một project, nơi mỗi analyst có thể tạo và lưu bảng riêng (create and store tables for their own use).
- 🚫 Các bảng cá nhân không bị truy cập bởi analyst khác (không accessible by other analysts). Mục tiêu là sử dụng IAM roles của BigQuery để kiểm soát quyền một cách chính xác, tận dụng dataset-level permissions thay vì project-level để tránh rò rỉ quyền. Đây là kịch bản thực tế trong enterprise data governance, đảm bảo least privilege principle (nguyên tắc quyền hạn tối thiểu).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng:
Give analysts the BigQuery Data Viewer role on the shared dataset. Create a dataset for each analyst, and give each analyst the BigQuery Data Editor role at the dataset level for their assigned dataset.
Lý do chọn đáp án này 🛠️:
- ✅ BigQuery Data Viewer trên shared dataset chỉ cho phép đọc dữ liệu (query tables/views), không chỉnh sửa (không insert/update/delete data hoặc schema) → Đáp ứng readable/unmodifiable cho dữ liệu chia sẻ.
- ✅ Tạo dataset riêng cho từng analyst, cấp BigQuery Data Editor tại mức dataset → Analyst chỉ create/modify tables trong dataset của mình (read/write data, manage tables), không ảnh hưởng dataset khác nhờ dataset-scoped IAM (quyền giới hạn dataset).
- ✅ Tất cả trong cùng project, dễ quản lý, tuân thủ BigQuery IAM best practices (cập nhật đến 2026, không thay đổi core roles). Không rò rỉ quyền giữa analysts.
📘 Nguồn tham khảo: BigQuery Access Control & IAM Roles for BigQuery (Google Cloud Docs, version 2024-2026).
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm lý do cụ thể dựa trên quyền IAM BigQuery (Data Viewer: read-only data; Data Editor: read/write/modify data/tables trong scope).
-
Phương án 1 ❌:
Give analysts the BigQuery Data Viewer role at the project level. Create one other dataset, and give the analysts the BigQuery Data Editor role on that dataset.
Tại sao sai? 🧨 Quyền Data Viewer tại project level cho phép analysts đọc tất cả datasets trong project (bao gồm dataset cá nhân của nhau) → Vi phạm yêu cầu "không accessible by other analysts". Dataset chung với Data Editor cũng khiến tất cả analysts chỉnh sửa lẫn nhau, không phải workspace riêng. -
Phương án 2 ❌:
Give analysts the BigQuery Data Viewer role at the project level. Create a dataset for each analyst, and give each analyst the BigQuery Data Editor role at the project level.
Tại sao sai? 🚫 Data Viewer project level lại cho phép đọc cross-dataset (analysts đọc dataset cá nhân của nhau). Data Editor project level còn tệ hơn: analyst có thể create/modify/delete tables ở mọi dataset trong project → Hoàn toàn mất kiểm soát, không isolate workspaces. -
Phương án 3 ✅:
Give analysts the BigQuery Data Viewer role on the shared dataset. Create a dataset for each analyst, and give each analyst the BigQuery Data Editor role at the dataset level for their assigned dataset.
Tại sao đúng? 🎯 Hoàn hảo khớp yêu cầu: Data Viewer dataset-scoped chỉ đọc shared dataset (không chỉnh sửa). Data Editor dataset-scoped cho phép analyst tự do trong dataset riêng, không leak quyền sang dataset khác (BigQuery enforce per-dataset IAM từ 2019, ổn định đến 2026). -
Phương án 4 ❌:
Give analysts the BigQuery Data Viewer role on the shared dataset. Create one other dataset and give the analysts the BigQuery Data Editor role on that dataset.
Tại sao sai? 🔒 Data Viewer shared dataset ok (read-only shared), nhưng một dataset chung với Data Editor khiến tất cả analysts chia sẻ và chỉnh sửa lẫn nhau → Không tạo "individual workspaces", vi phạm isolation.
Kết luận 🌟: Phương án đúng tận dụng fine-grained dataset IAM để balance accessibility và security. Khuyến nghị thêm BigQuery Admin cho owner quản lý roles! 📘 Nguồn bổ sung: BigQuery Dataset-level Permissions (Google Cloud Best Practices 2026).
- A Use watermarks to define the expected data arrival window. Allow late data as it arrives.
- B Change your windowing function to tumbling windows to avoid overlapping window periods.
- C Change your windowing function to session windows to define your windows based on certain activity.
- D Expand your hopping window so that the late data has more time to arrive within the grouping.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào vấn đề xử lý dữ liệu muộn (late data) trong một pipeline streaming sử dụng Google Cloud Dataflow (dựa trên Apache Beam). Bạn đang dùng hopping windows để nhóm dữ liệu khi chúng đến, nhưng một số dữ liệu muộn không được đánh dấu là late data, dẫn đến tổng hợp (aggregations) downstream bị sai lệch. Mục tiêu là tìm giải pháp bắt và xử lý late data vào đúng window mà không làm gián đoạn pipeline.
Chi tiết vấn đề:
- Streaming pipeline: Dữ liệu đến liên tục, không batch.
- Hopping windows: Các cửa sổ trượt (overlap), ví dụ mỗi 5 phút với slide 1 phút.
- Late data: Dữ liệu đến sau thời điểm mong đợi (sau watermark của window).
- Hậu quả: Aggregations sai vì late data bị bỏ qua hoặc không vào đúng window.
- Yêu cầu giải pháp: Phải capture late data vào window phù hợp, theo docs Dataflow cập nhật đến 2024 (vẫn áp dụng 2026, Beam 2.55+).
📘 Tài liệu tham khảo:
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use watermarks to define the expected data arrival window. Allow late data as it arrives.
Lý do 🛠️:
- Watermarks là cơ chế cốt lõi trong Dataflow/Beam để theo dõi tiến độ dữ liệu (event time). Nó định nghĩa "thời điểm mong đợi" dữ liệu phải đến (ví dụ: watermark hiện tại = max observed timestamp - allowed lateness).
- Allow late data: Sử dụng
allowedLateness(trong AfterWatermark trigger) để giữ window mở thêm thời gian, cho phép late data được xử lý và emit vào pane riêng (on-time pane + late pane). Late data sẽ được đánh dấu và tổng hợp đúng window. - Giải pháp này chính xác vì trực tiếp giải quyết vấn đề: capture late data vào window gốc mà không thay đổi windowing strategy. Áp dụng ngay với hopping windows, không cần refactor lớn.
- Cập nhật 2026: Beam/Dataflow vẫn dùng watermark + triggers làm standard cho late data handling.
❌ Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc:
-
Use watermarks to define the expected data arrival window. Allow late data as it arrives.
✅ Đúng (như giải thích trên). Đây là best practice chuẩn của Dataflow để xử lý late data một cách chính xác, linh hoạt với mọi loại window (hopping, sliding). -
Change your windowing function to tumbling windows to avoid overlapping window periods.
❌ Sai. Tumbling windows là non-overlapping (không trượt, ví dụ 5 phút cố định), giúp tránh overlap nhưng không giải quyết late data. Late data vẫn bị drop nếu đến sau watermark của tumbling window, dẫn đến mất dữ liệu và aggregations vẫn sai. -
Change your windowing function to session windows to define your windows based on certain activity.
❌ Sai. Session windows nhóm dựa trên khoảng lặng (gap) hoạt động (ví dụ: gap 10 phút), phù hợp cho user sessions nhưng không xử lý late data trực tiếp. Nó thay đổi logic windowing hoàn toàn, không capture late data vào window cũ mà tạo session mới, làm aggregations downstream lệch lạc hơn. -
Expand your hopping window so that the late data has more time to arrive within the grouping.
❌ Sai. Mở rộng hopping window (tăng size/period) chỉ trì hoãn vấn đề tạm thời (late data có thể vẫn muộn hơn), nhưng không đánh dấu hay capture late data đúng cách. Watermark vẫn quyết định lateness; dữ liệu sau watermark sẽ bị drop, không vào "appropriate window" như yêu cầu. Không scalable cho streaming real-time.