Ngân hàng đề — Google Cloud Professional Data Engineer
Tìm thấy 429 câu.
Data Science team runs a query filtered on a date column and limited to 30`"90 days of data, the query scans the entire table. You also noticed that your bill is increasing more quickly than you expected. You want to resolve the issue as cost-effectively as possible while maintaining the ability to conduct SQL queries.
What should you do?
- A Re-create the tables using DDL. Partition the tables by a column containing a TIMESTAMP or DATE Type.
- B Recommend that the Data Science team export the table to a CSV file on Cloud Storage and use Cloud Datalab to explore the data by reading the files directly.
- C Modify your pipeline to maintain the last 30ג€"90 days of data in one table and the longer history in a different table to minimize full table scans over the entire history.
- D Write an Apache Beam pipeline that creates a BigQuery table per day. Recommend that the Data Science team use wildcards on the table name suffixes to select the data they need.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một tình huống thực tế trong Google Cloud BigQuery:
- Bạn có dữ liệu lịch sử 3 năm lưu trữ trong BigQuery.
- Pipeline dữ liệu thêm dữ liệu mới hàng ngày vào BigQuery.
- Vấn đề: Khi đội Data Science chạy query lọc theo cột ngày tháng (date column) và giới hạn chỉ 30-90 ngày dữ liệu, query vẫn scan toàn bộ bảng (full table scan).
- Hậu quả: Hóa đơn BigQuery tăng nhanh (do chi phí dựa trên dữ liệu scan).
- Yêu cầu: Giải quyết tiết kiệm chi phí nhất (cost-effectively), đồng thời giữ khả năng query SQL bình thường.
📈 Nguyên nhân gốc rễ: BigQuery là columnar store, nhưng nếu bảng không được partition hoặc cluster, mọi query đều scan toàn bộ dữ liệu → tốn kém với dữ liệu lớn 3 năm. Giải pháp cần tận dụng tính năng partitioning của BigQuery để query chỉ scan partition liên quan (dựa trên date/TIMESTAMP), giảm chi phí scan lên đến 90-99%!
✅ Đáp án ĐÚNG và lý do chọn
Đáp án đúng: Re-create the tables using DDL. Partition the tables by a column containing a TIMESTAMP or DATE Type.
🛠️ Lý do chi tiết:
- Partitioning theo DATE/TIMESTAMP là tính năng native của BigQuery (cập nhật đến 2026: hỗ trợ ingestion-time partitioning, date/timestamp partitioning, integer-range partitioning).
- Khi re-create bảng bằng DDL (Data Definition Language, ví dụ
CREATE TABLE ... PARTITION BY date_column), query lọc theo partition key (date) sẽ tự động chỉ scan partition cần thiết (pruning), tránh full scan. - Tiết kiệm chi phí: Giảm bytes scanned → bill giảm mạnh (BigQuery charge theo TiB scanned).
- Duy trì SQL query: Không thay đổi workflow, đội Data Science vẫn query bình thường.
- Cost-effective nhất: Không cần tool ngoài, chỉ cấu hình bảng một lần. Với dữ liệu daily, partitioning daily/monthly lý tưởng.
📘 Nguồn tham khảo: BigQuery Table Partitioning (Google Cloud Docs, phiên bản 2026 hỗ trợ automatic clustering kết hợp partitioning).
📋 Giải thích TẤT CẢ các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm lý do cụ thể dựa trên best practices BigQuery.
-
Re-create the tables using DDL. Partition the tables by a column containing a TIMESTAMP or DATE Type.
✅ ĐÚNG (như đã giải thích ở trên). 🏆 Giải pháp tối ưu, trực tiếp giải quyết full table scan mà không phức tạp hóa pipeline hay workflow. -
Recommend that the Data Science team export the table to a CSV file on Cloud Storage and use Cloud Datalab to explore the data by reading the files directly.
❌ SAI. 🧨 Phương án này không hiệu quả: Export CSV tốn thời gian/chi phí (BigQuery export fee), Cloud Datalab (nay là Vertex AI Workbench) đọc file trực tiếp chậm với dữ liệu lớn, không hỗ trợ SQL query native như BigQuery. Không giải quyết vấn đề scan table, chỉ né tránh tạm thời → bill vẫn cao do export thường xuyên. -
Modify your pipeline to maintain the last 30"90 days of data in one table and the longer history in a different table to minimize full table scans over the entire history.
❌ SAI. 🚫 Phức tạp và kém hiệu quả: Phải sửa pipeline để duy trì 2 bảng riêng (recent vs. history), dẫn đến query phức tạp (UNION/JOIN nhiều bảng), vẫn full scan bảng recent nếu query 90 ngày. Với dữ liệu 3 năm + daily ingest, quản lý dễ lỗi (data duplication/missing). Không tận dụng partitioning native → không cost-effective bằng. -
Write an Apache Beam pipeline that creates a BigQuery table per day. Recommend that the Data Science team use wildcards on the table name suffixes to select the data they need.
❌ SAI. 🔄 Không tối ưu: Tạo bảng riêng mỗi ngày (table-per-day) qua Beam tốn metadata overhead (BigQuery giới hạn 100k tables/project), query wildcard (table_*) chậm/management khó. Chi phí cao hơn partitioning (vì mỗi table là unit riêng, slot usage kém). BigQuery khuyến nghị partitioning thay vì sharding table → phương án này lỗi thời so với phiên bản 2026 (partitioning + clustering superior).
🎯 Kết luận và khuyến nghị
🛠️ Hành động ngay: Sử dụng DDL để partition (ví dụ: PARTITION BY DATE(timestamp_col)), kết hợp clustering theo filter columns khác để tối ưu hơn nữa. Test query với EXPLAIN để verify pruning. Bill sẽ giảm ngay lập tức! Nếu dữ liệu ingestion-time based, dùng ingestion-time partitioning miễn phí.
📘 Tài liệu bổ sung:
- BigQuery Partitioning & Clustering.
- Query Optimization Best Practices (cập nhật 2026).
- A Deploy small Kafka clusters in your data centers to buffer events.
- B Have the data acquisition devices publish data to Cloud Pub/Sub.
- C Establish a Cloud Interconnect between all remote data centers and Google.
- D Write a Cloud Dataflow pipeline that aggregates all data in session windows.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một công ty logistics đang vận hành các trung tâm dữ liệu nhỏ (small data centers) trên toàn thế giới để thu thập sự kiện (events) từ cảm biến trên xe (vehicle-based sensors). Các đường truyền thuê (leased lines) kết nối từ hạ tầng thu thập dữ liệu (event collection infrastructure) đến hạ tầng xử lý sự kiện (event processing infrastructure) không đáng tin cậy, với độ trễ (latency) không thể dự đoán.
📌 Vấn đề cốt lõi: Cần cải thiện độ tin cậy giao dữ liệu sự kiện một cách cost-effective nhất (tiết kiệm chi phí nhất), mà không phụ thuộc vào kết nối leased lines kém chất lượng.
🛠️ Bối cảnh: Đây là tình huống streaming data real-time từ edge locations (các data center nhỏ), cần giải pháp decoupling (tách biệt producer và consumer) để buffer và đảm bảo delivery reliable, phù hợp với Google Cloud services (Pub/Sub, Dataflow, Interconnect).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Have the data acquisition devices publish data to Cloud Pub/Sub.
Lý do:
Cloud Pub/Sub là dịch vụ messaging managed của Google Cloud, cho phép các thiết bị thu thập dữ liệu (data acquisition devices) publish trực tiếp events lên topic mà không cần phụ thuộc vào leased lines không ổn định. Pub/Sub tự động buffer messages (lên đến 10MB/topic, retention 7-30 ngày tùy config), hỗ trợ at-least-once delivery, và scale tự động mà không tốn kém quản lý hạ tầng. Event processing infra có thể subscribe từ bất kỳ đâu, giải quyết latency unpredictable một cách cost-effective (pay-per-use, không cần dedicated network). Đây là giải pháp chuẩn cho high-throughput event streaming theo best practices Google Cloud đến 2026 (Pub/Sub Lite cho edge cost thấp hơn nếu cần).
📘 Nguồn tham khảo: Google Cloud Pub/Sub Documentation (cập nhật 2025: hỗ trợ global replication và exactly-once delivery).
📋 Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, với lý do đúng/sai dựa trên kiến thức Google Cloud mới nhất (2026):
-
❌ Deploy small Kafka clusters in your data centers to buffer events.
Phương án này yêu cầu tự triển khai và quản lý Kafka clusters nhỏ tại các data center – phức tạp, tốn kém vận hành (hardware, scaling, monitoring, Zookeeper), không cost-effective cho môi trường edge với data centers nhỏ. Kafka không giải quyết vấn đề connectivity leased lines (vẫn cần transport data qua lines kém), và Google khuyến nghị dùng Pub/Sub thay thế managed Kafka (như Confluent on GKE). Không phù hợp cho reliability global mà không overhead cao. -
✅ Have the data acquisition devices publish data to Cloud Pub/Sub.
(Như đã giải thích ở trên) – Giải pháp lý tưởng: Decouples producers (devices) khỏi consumers (processing infra), buffer resilient chống network failure, cost thấp (~$0.40/million operations), scale unlimited. Hỗ trợ IoT/edge devices trực tiếp qua HTTP/gRPC. -
❌ Establish a Cloud Interconnect between all remote data centers and Google.
Cloud Interconnect (Dedicated Interconnect hoặc Partner Interconnect) cung cấp kết nối private cao tốc đến Google Cloud, nhưng rất đắt đỏ (setup fee $0.03-$0.3/GB + port fees hàng tháng), không cost-effective cho "small data centers around the world" (cần interconnect tại từng site). Chỉ phù hợp enterprise lớn, không giải quyết buffer/reliability nếu latency vẫn biến động. -
❌ Write a Cloud Dataflow pipeline that aggregates all data in session windows.
Cloud Dataflow (Apache Beam managed) dùng để xử lý stream data (aggregate trong session windows), không phải để buffer hoặc cải thiện delivery reliability từ collection infra. Pipeline này chạy trên processing side, vẫn cần dữ liệu đến được trước (qua leased lines kém), dẫn đến data loss. Không giải quyết vấn đề cốt lõi connectivity, chỉ là bước sau (post-processing).
📘 Nguồn tham khảo: Google Cloud Dataflow Docs (2026: Unified Batch/Stream, nhưng không thay thế messaging layer).
🧠 Tóm tắt insight: Sử dụng Pub/Sub là pattern "publish-subscribe" chuẩn cho event-driven architecture, giúp logistics company xử lý sensor data reliable mà tiết kiệm (so với self-managed hoặc dedicated network). Nếu scale lớn, kết hợp Pub/Sub + Dataflow cho end-to-end pipeline!
- A Speech-to-Text API
- B Cloud Natural Language API
- C Dialogflow Enterprise Edition
- D AutoML Natural Language
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả tình huống một nhà bán lẻ (retailer) muốn tích hợp khả năng bán hàng trực tuyến với các trợ lý ảo tại nhà như Google Home. Yêu cầu chính là:
- Giải thích (interpret) lệnh thoại (voice commands) từ khách hàng, ví dụ: "Mua cho tôi một chiếc áo sơ mi size M".
- Gửi lệnh đặt hàng (issue an order) đến hệ thống backend để xử lý đơn hàng thực tế (như cập nhật kho, thanh toán).
📌 Mục tiêu cốt lõi: Xây dựng một hệ thống hội thoại (conversational AI) hỗ trợ voice interaction, hiểu ý định (intent) của người dùng từ giọng nói, và kết nối với backend để thực hiện hành động (fulfillment). Đây là kịch bản điển hình cho Actions on Google hoặc tích hợp với Google Assistant/Google Home, đòi hỏi giải pháp toàn diện từ voice-to-text, NLU (Natural Language Understanding), đến dialog management và integration.
🛠️ Yêu cầu kiến thức cập nhật (đến 2026): Theo tài liệu Google Cloud mới nhất (Dialogflow CX - phiên bản nâng cao của Enterprise Edition, ra mắt 2020 và cập nhật liên tục), Dialogflow là lựa chọn chuẩn cho các ứng dụng voice assistant tích hợp Google Home.
📘 Tài liệu tham khảo:
- Dialogflow Documentation (Google Cloud, cập nhật 2025).
- Actions on Google for Retail (hỗ trợ tích hợp bán hàng với Google Home).
✅ Đáp án đúng: Dialogflow Enterprise Edition
Lý do lựa chọn:
- Dialogflow Enterprise Edition (nay là nền tảng cho Dialogflow CX) là giải pháp toàn diện cho conversational AI, chuyên xử lý voice commands từ Google Home/Google Assistant.
- Nó hỗ trợ:
- Speech-to-Text tích hợp (qua telephony hoặc streaming).
- Intent recognition và entity extraction để hiểu lệnh mua hàng.
- Dialog management (multi-turn conversation, ví dụ: xác nhận size, địa chỉ).
- Fulfillment webhook để gọi API backend (gửi order đến hệ thống bán hàng).
- Hoàn hảo cho tích hợp Actions on Google, cho phép deploy nhanh lên Google Home mà không cần build từ đầu.
- ✅ Phù hợp 100% với yêu cầu "interpret voice commands và issue order to backend".
❌ Giải thích tất cả các phương án
-
[SAI] Speech-to-Text API
❌ Sai vì: Chỉ chuyển giọng nói thành văn bản (speech-to-text), không hiểu ý định hoặc xử lý hội thoại. Bạn vẫn cần thêm NLU/dialog để interpret "mua áo" và gọi backend. Không đủ cho tích hợp Google Home đầy đủ. -
[SAI] Cloud Natural Language API
❌ Sai vì: Tập trung vào phân tích văn bản (text analysis) như sentiment, entities, syntax. Không hỗ trợ voice input trực tiếp, không quản lý dialog multi-turn, và không có fulfillment để issue order backend. Chỉ là công cụ hỗ trợ, không phải giải pháp end-to-end. -
[ĐÚNG] Dialogflow Enterprise Edition
✅ Đúng vì: Như giải thích trên, đây là nền tảng conversational AI chuyên biệt cho voice assistants như Google Home, xử lý từ voice → intent → fulfillment một cách liền mạch. -
[SAI] AutoML Natural Language
❌ Sai vì: Dùng để train custom model cho classification/extraction trên text (không voice). Không hỗ trợ dialog flow, voice integration, hoặc webhook backend. Phù hợp cho task đơn lẻ như phân loại review, không phải hệ thống đặt hàng phức tạp.
- A Cloud Dataflow
- B Cloud Composer
- C Cloud Dataprep
- D Cloud Dataproc
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả tình huống công ty đang triển khai hybrid cloud initiative (sáng kiến đám mây lai), với một data pipeline phức tạp di chuyển dữ liệu giữa các dịch vụ của nhiều nhà cung cấp đám mây (cloud providers), đồng thời tận dụng dịch vụ từ từng nhà cung cấp đó. Nhiệm vụ là chọn cloud-native service phù hợp để orchestrate (điều phối) toàn bộ pipeline này.
🔍 Ý nghĩa chính:
- Hybrid/multi-cloud: Pipeline không chỉ giới hạn trong một nhà cung cấp (như Google Cloud), mà liên kết giữa các cloud (ví dụ: GCP, AWS, Azure), đòi hỏi công cụ linh hoạt, hỗ trợ integration đa nền tảng.
- Orchestrate toàn bộ pipeline: Cần tool quản lý workflow, scheduling, dependency giữa các bước, monitoring, chứ không phải chỉ xử lý dữ liệu cụ thể.
- Cloud-native: Dịch vụ được quản lý hoàn toàn trên đám mây, dễ scale và integrate.
Dựa trên kiến thức Google Cloud cập nhật đến năm 2026 (phiên bản Cloud Composer 3.x với Airflow 2.9+), đây là câu hỏi kiểm tra khả năng chọn tool orchestration phù hợp cho môi trường multi-cloud/hybrid. 📘 Tài liệu tham khảo: Google Cloud Composer Documentation, Apache Airflow Multi-Cloud Guide.
✅ Đáp án đúng: Cloud Composer
Lý do lựa chọn:
- Cloud Composer là dịch vụ Managed Apache Airflow trên Google Cloud, chuyên orchestrate workflows phức tạp với khả năng DAG (Directed Acyclic Graph) để định nghĩa dependency giữa các task.
- 🛠️ Phù hợp hybrid/multi-cloud: Hỗ trợ operators tùy chỉnh kết nối với bất kỳ dịch vụ nào (GCP, AWS S3/EC2, Azure Blob, on-prem via plugins), scheduling linh hoạt, retry mechanism, và monitoring toàn diện.
- Đến 2026, Composer hỗ trợ cross-project/cross-cloud integrations qua Airflow Providers (ví dụ: AWS Provider, Azure Provider), lý tưởng cho pipeline di chuyển dữ liệu giữa các cloud mà không bị lock-in.
- Không phải tool xử lý dữ liệu cụ thể, mà là orchestrator tổng thể – đúng yêu cầu "orchestrate the entire pipeline".
📋 Giải thích tất cả các phương án (đúng/sai)
-
Cloud Dataflow ❌
Sai vì: Đây là dịch vụ batch/stream processing dựa trên Apache Beam, dùng để thực thi pipeline dữ liệu lớn (transform, ETL). Không phải orchestrator; chỉ xử lý một job cụ thể, thiếu scheduling/DAG cho multi-step/multi-cloud workflow. Không hỗ trợ hybrid tốt (chủ yếu GCP-native). -
Cloud Composer ✅
Đúng vì: Như giải thích trên, là orchestrator mạnh mẽ cho pipeline phức tạp, hỗ trợ multi-cloud integrations qua Airflow DAGs (ví dụ: task GCP Dataflow + AWS Glue + Azure Data Factory). Scale tự động, managed service, phù hợp hybrid đến 2026 với Airflow 2.9+ features như Dynamic Task Mapping. -
Cloud Dataprep ❌
Sai vì: Đây là UI-based data preparation tool (dựa trên Trifacta), dùng để clean/visualize dữ liệu qua giao diện kéo-thả. Không orchestrate pipeline toàn diện; chỉ prep dữ liệu đơn lẻ, không hỗ trợ scheduling hay multi-cloud complexity. -
Cloud Dataproc ❌
Sai vì: Dịch vụ managed Hadoop/Spark cho big data processing (cluster-based jobs). Giỏi compute-intensive tasks nhưng thiếu orchestration (không có DAG/scheduling native), chủ yếu single-cloud, không lý tưởng cho hybrid pipeline di chuyển dữ liệu giữa providers.
🧠 Kết luận: Cloud Composer là lựa chọn tối ưu cho orchestration hybrid/multi-cloud, giúp tránh vendor lock-in. Nếu triển khai thực tế, bắt đầu bằng tạo Environment trên GCP Console! 🚀
- A Use Analytics Hub to control data access, and provide third party companies with access to the dataset.
- B Use Cloud Scheduler to export the data on a regular basis to Cloud Storage, and provide third-party companies with access to the bucket.
- C Create a separate dataset in BigQuery that contains the relevant data to share, and provide third-party companies with access to the new dataset.
- D Create a Dataflow job that reads the data in frequent time intervals, and writes it to the relevant BigQuery dataset or Cloud Storage bucket for third-party companies to use.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào tình huống chia sẻ dataset trong BigQuery với các công ty thứ ba (third-party companies), đồng thời đảm bảo chi phí thấp và dữ liệu luôn cập nhật (current).
- Bối cảnh: Bạn đang sử dụng dataset trong BigQuery để phân tích dữ liệu. BigQuery là dịch vụ data warehouse serverless của Google Cloud, hỗ trợ phân tích dữ liệu lớn với hiệu suất cao.
- Yêu cầu chính:
✅ Giải pháp phải cho phép truy cập dataset gốc mà không sao chép dữ liệu (tránh chi phí lưu trữ/query thừa).
✅ Dữ liệu phải luôn mới nhất (real-time hoặc near real-time).
✅ Chi phí thấp: Tránh các hoạt động export, duplicate hoặc compute nặng. - Mục tiêu: Tìm giải pháp tối ưu trong Google Cloud để chia sẻ dữ liệu an toàn, hiệu quả. (Kiến thức cập nhật đến 2026: Analytics Hub là tính năng chính thức của BigQuery cho data marketplace và sharing, hỗ trợ listing dữ liệu mà không cần di chuyển).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Analytics Hub to control data access, and provide third party companies with access to the dataset.
Lý do:
- 🛡️ Analytics Hub (nay là phần của Dataplex Data Sharing từ 2024-2026) cho phép chia sẻ dữ liệu mà không sao chép, sử dụng data listings để third-party truy cập trực tiếp dataset BigQuery gốc.
- 📈 Dữ liệu luôn current: Thay đổi trong dataset gốc sẽ tự động phản ánh (zero-copy sharing).
- 💰 Chi phí thấp: Không tốn lưu trữ/query thừa, chỉ tính phí truy vấn thực tế của người dùng third-party.
- 🔒 Kiểm soát truy cập: Publisher kiểm soát quyền qua IAM, subscriptions, và governance. Đây là giải pháp được AWS... (lưu ý: câu hỏi Google Cloud, không AWS), khuyến nghị chính thức cho enterprise data sharing.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng phương án (giữ nguyên văn bản gốc tiếng Anh), đánh dấu đúng/sai với lý do chi tiết bằng tiếng Việt:
-
✅ [ĐÚNG] Use Analytics Hub to control data access, and provide third party companies with access to the dataset.
Như đã giải thích ở trên: Giải pháp lý tưởng với zero-copy, real-time, low-cost và secure. Hoàn hảo cho yêu cầu! -
❌ [SAI] Use Cloud Scheduler to export the data on a regular basis to Cloud Storage, and provide third-party companies with access to the bucket.
Lý do sai: Export định kỳ qua Cloud Scheduler tạo dữ liệu không current (chỉ snapshot theo lịch, delay). Chi phí cao do export fees (BigQuery export ~$0.05/GB) + lưu trữ GS + truy cập bucket. Không hiệu quả cho data sharing real-time. -
❌ [SAI] Create a separate dataset in BigQuery that contains the relevant data to share, and provide third-party companies with access to the new dataset.
Lý do sai: Tạo dataset riêng yêu cầu sao chép dữ liệu (duplicate storage), tăng chi phí lưu trữ (~$0.02/GB/tháng x2) và query (on-demand/scanning). Dữ liệu không tự động current trừ khi sync thủ công. Phân tán quản lý, không scale tốt. -
❌ [SAI] Create a Dataflow job that reads the data in frequent time intervals, and writes it to the relevant BigQuery dataset or Cloud Storage bucket for third-party companies to use.
Lý do sai: Dataflow (Apache Beam) tốn chi phí compute cao (vCPU, PD, Data Processing Units - DPU từ 2025 ~$0.01/vCPU-giờ). Sync frequent gây lag dữ liệu và overhead lớn. Không phải giải pháp native cho sharing, phức tạp deploy/maintain.
📘 Tài liệu tham khảo (cập nhật 2026)
- Google Cloud Docs: Analytics Hub & Data Sharing – Zero-copy sharing chính thức.
- Dataplex Catalog: Data Sharing Best Practices (tích hợp Analytics Hub từ 2024).
- BigQuery Pricing: Pricing Page – Xác nhận no extra cost cho sharing.
- Cập nhật 2025-2026: Analytics Hub evolve thành Dataplex Lakehouse Sharing, hỗ trợ cross-region/multi-cloud (không AWS native).
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm ví dụ code hoặc demo, hãy hỏi nhé!
CDC so that changes to the source systems are available to query in BigQuery in near-real time using log-based CDC streams, while also optimizing for the performance of applying changes to the data warehouse. Which two steps should they take to ensure that changes are available in the BigQuery reporting table with minimal latency while reducing compute overhead? (Choose two.)
- A Perform a DML INSERT, UPDATE, or DELETE to replicate each individual CDC record in real time directly on the reporting table.
- B Insert each new CDC record and corresponding operation type to a staging table in real time.
- C Periodically DELETE outdated records from the reporting table.
- D Periodically use a DML MERGE to perform several DML INSERT, UPDATE, and DELETE operations at the same time on the reporting table.
- E Insert each new CDC record and corresponding operation type in real time to the reporting table, and use a materialized view to expose only the newest version of each unique record.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi này thuộc chủ đề migrate dữ liệu từ data warehouse on-premises sang BigQuery trên Google Cloud Platform (GCP). Công ty đang sử dụng trigger-based CDC (Change Data Capture) để cập nhật dữ liệu hàng ngày từ nhiều nguồn transactional database. Họ muốn cải thiện bằng log-based CDC streams để dữ liệu thay đổi từ nguồn có thể query gần real-time trong BigQuery, đồng thời tối ưu performance khi apply changes vào data warehouse.
Mục tiêu chính:
- Đảm bảo thay đổi có sẵn trong reporting table với minimal latency (độ trễ thấp).
- Giảm compute overhead (giảm tải tính toán).
Câu hỏi yêu cầu chọn hai steps (hai bước) để đạt được điều này. Đây là best practice trong BigQuery cho xử lý CDC: Sử dụng staging table để capture real-time, rồi batch MERGE định kỳ để upsert/delete hiệu quả, tránh DML real-time trực tiếp trên reporting table (vì BigQuery không hỗ trợ real-time DML tốt, gây slot quota và latency cao).
📘 Nguồn tham khảo:
- BigQuery Streaming Inserts (cập nhật 2024-2026: Hỗ trợ streaming lên đến 1MB/s, nhưng DML real-time kém hiệu quả).
- BigQuery MERGE statement best practices for CDC (Khuyến nghị batch MERGE cho low-latency CDC).
- Datastream for BigQuery CDC (Log-based CDC với staging + MERGE).
✅ Đáp án đúng (Chọn hai phương án sau)
- Insert each new CDC record and corresponding operation type to a staging table in real time.
- Periodically use a DML MERGE to perform several DML INSERT, UPDATE, and DELETE operations at the same time on the reporting table.
Lý do lựa chọn 🛠️:
- Staging table cho phép streaming insert real-time từ CDC streams (như Datastream hoặc Cloud Data Fusion), lưu operation type (INSERT/UPDATE/DELETE). Điều này giảm latency capture (near-real-time ~seconds) mà không tải trực tiếp reporting table.
- Periodic MERGE batch nhiều changes cùng lúc: Upsert/delete hiệu quả dựa trên key + operation type, tối ưu compute (BigQuery tính slot theo query size, batch giảm chi phí 10-100x so real-time DML). Latency end-to-end ~1-5 phút, phù hợp near-real-time.
- Kết hợp hai bước này tuân thủ BigQuery best practices (2024+): Tránh mutation real-time trên large tables, dùng MERGE atomic cho consistency.
📋 Giải thích tất cả các phương án (Đúng/Sai)
-
❌ Perform a DML INSERT, UPDATE, or DELETE to replicate each individual CDC record in real time directly on the reporting table.
Sai vì: BigQuery không hỗ trợ real-time DML hiệu quả trên large reporting tables (gây quota errors, slot exhaustion, latency cao >10s/record). Streaming chỉ tốt cho INSERT, không cho UPDATE/DELETE real-time. Tăng compute overhead gấp bội, vi phạm mục tiêu. -
✅ Insert each new CDC record and corresponding operation type to a staging table in real time.
Đúng vì: Staging table nhỏ, hỗ trợ streaming inserts nhanh (low-latency capture từ log-based CDC). Lưu operation type để MERGE sau, giảm tải reporting table. Best practice cho CDC pipelines (Datastream auto-creates staging). -
❌ Periodically DELETE outdated records from the reporting table.
Sai vì: Chỉ DELETE không xử lý INSERT/UPDATE đầy đủ CDC. Gây inconsistency, tăng compute (full scan table lớn), không tối ưu latency hay overhead. Không atomic như MERGE. -
✅ Periodically use a DML MERGE to perform several DML INSERT, UPDATE, and DELETE operations at the same time on the reporting table.
Đúng vì: MERGE là DML atomic, batch nhiều ops (dựa key + operation type từ staging), upsert/delete hiệu quả. Giảm latency tổng thể, compute thấp (1 query thay 1000s DML), hỗ trợ BigQuery slots tốt (cập nhật 2024+). -
❌ Insert each new CDC record and corresponding operation type in real time to the reporting table, and use a materialized view to expose only the newest version of each unique record.
Sai vì: Streaming INSERT real-time vào reporting table gây duplicates/inconsistency (BigQuery append-only streaming). Materialized view refresh chậm (15p+), không real-time, tăng storage/compute (view scan full table). Không xử lý DELETE/UPDATE đúng.
- A Use Apache Kafka for message ingestion and use Cloud Dataproc for streaming analysis.
- B Use Apache Kafka for message ingestion and use Cloud Dataflow for streaming analysis.
- C Use Cloud Pub/Sub for message ingestion and Cloud Dataproc for streaming analysis.
- D Use Cloud Pub/Sub for message ingestion and Cloud Dataflow for streaming analysis.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi yêu cầu thiết kế một data processing pipeline (đường ống xử lý dữ liệu) với các yêu cầu chính sau:
- Tự động scale khi tải tăng (scale automatically as load increases).
- Xử lý tin nhắn ít nhất một lần (at-least-once processing).
- Giữ thứ tự tin nhắn trong các cửa sổ 1 giờ (ordered within windows of 1 hour).
📘 Giải thích chi tiết: Đây là kịch bản streaming data pipeline trên Google Cloud Platform (GCP). Pipeline cần một dịch vụ ingestion (hấp thụ tin nhắn) có khả năng mở rộng tự động, đảm bảo thứ tự và độ tin cậy (at-least-once). Sau đó, dùng dịch vụ xử lý streaming hỗ trợ windowing (cửa sổ thời gian 1 giờ) để nhóm và sắp xếp dữ liệu. Kiến thức cập nhật đến 2026: GCP ưu tiên các dịch vụ managed như Pub/Sub (với ordering keys từ 2019, cải tiến scale đến hàng triệu message/s) và Dataflow (Apache Beam 2.58+ hỗ trợ unified streaming/batch, exactly-once semantics với at-least-once fallback).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Cloud Pub/Sub for message ingestion and Cloud Dataflow for streaming analysis.
🛠️ Lý do chi tiết:
- Cloud Pub/Sub là dịch vụ messaging managed, scale tự động vô hạn (auto-scales to 1M+ msg/s theo docs 2024-2026), hỗ trợ at-least-once delivery mặc định, và message ordering qua ordering keys (đảm bảo thứ tự trong cùng key, phù hợp windows 1 giờ).
- Cloud Dataflow (dựa Apache Beam) là dịch vụ streaming managed, scale tự động, hỗ trợ windowing chính xác (Fixed Windows 1h), xử lý at-least-once (và exactly-once với watermarking), sắp xếp dữ liệu trong windows dễ dàng.
Kết hợp hoàn hảo cho pipeline streaming GCP-native, không cần quản lý infra.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm lý do bằng tiếng Việt rõ ràng:
-
❌ [SAI] Use Apache Kafka for message ingestion and use Cloud Dataproc for streaming analysis.
Lý do sai: Kafka không phải dịch vụ managed của GCP (phải tự deploy trên GKE/Compute, khó scale tự động như Pub/Sub). Dataproc (Spark Streaming/Flink) hỗ trợ streaming nhưng kém hiệu quả cho windowing ordered so với Beam, và không at-least-once native mượt mà (cần config phức tạp). Không phù hợp pipeline fully managed. -
❌ [SAI] Use Apache Kafka for message ingestion and use Cloud Dataflow for streaming analysis.
Lý do sai: Kafka vẫn là vấn đề chính – không scale tự động managed trên GCP (Confluent Cloud là option nhưng không native). Dataflow tốt cho processing, nhưng ingestion Kafka làm pipeline kém tích hợp, khó đảm bảo ordering trong windows end-to-end mà không custom connector. -
❌ [SAI] Use Cloud Pub/Sub for message ingestion and Cloud Dataproc for streaming analysis.
Lý do sai: Pub/Sub hoàn hảo cho ingestion (scale auto, at-least-once, ordering keys). Nhưng Dataproc chỉ là cluster managed cho batch/streaming jobs (Spark/Flink), không scale tự động real-time như Dataflow (phải resize thủ công/cluster), và windowing 1h ordered kém linh hoạt (không unified như Beam). -
✅ [ĐÚNG] Use Cloud Pub/Sub for message ingestion and Cloud Dataflow for streaming analysis.
Lý do đúng: Như đã giải thích ở trên – bộ đôi GCP-native lý tưởng cho scale auto, at-least-once, ordering trong windows 1h (Pub/Sub ordering + Dataflow Sessions/Fixed Windows). Đáp ứng đầy đủ yêu cầu đến 2026.
📚 Tài liệu tham khảo
- Cloud Pub/Sub Ordering (cập nhật 2025: scale & keys).
- Cloud Dataflow Streaming (Beam 2.58+, windowing & semantics).
- GCP Best Practices: Streaming Pipelines (Pub/Sub + Dataflow recommended).
- AWS so sánh: Không áp dụng trực tiếp (câu hỏi GCP), nhưng tương đương Kinesis + Flink on MSK (nhưng GCP native tốt hơn).
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀
✑ Each department should have access only to their data.
✑ Each department will have one or more leads who need to be able to create and update tables and provide them to their team.
✑ Each department has data analysts who need to be able to query but not modify data.
How should you set access to the data in BigQuery?
- A Create a dataset for each department. Assign the department leads the role of OWNER, and assign the data analysts the role of WRITER on their dataset.
- B Create a dataset for each department. Assign the department leads the role of WRITER, and assign the data analysts the role of READER on their dataset.
- C Create a table for each department. Assign the department leads the role of Owner, and assign the data analysts the role of Editor on the project the table is in.
- D Create a table for each department. Assign the department leads the role of Editor, and assign the data analysts the role of Viewer on the project the table is in.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc thiết lập quyền truy cập dữ liệu trong BigQuery (dịch vụ kho dữ liệu của Google Cloud Platform - GCP) cho các bộ phận khác nhau trong công ty. Các yêu cầu cụ thể bao gồm:
- ✅ Mỗi bộ phận chỉ truy cập được dữ liệu của riêng mình: Điều này đòi hỏi phải cô lập dữ liệu logic, thường sử dụng dataset làm đơn vị phân chia (vì dataset là container chứa tables và có thể gán quyền riêng biệt).
- 🛠️ Leads của bộ phận: Cần quyền tạo (create) và cập nhật (update) tables, đồng thời chia sẻ (provide) cho team – nghĩa là họ có thể quản lý nội dung dữ liệu nhưng không nhất thiết phải quản lý quyền truy cập IAM (để tránh rủi ro bảo mật).
- 📊 Data analysts: Chỉ được truy vấn (query) dữ liệu, không được sửa đổi (modify) bất kỳ thứ gì.
Giải pháp phải tuân thủ mô hình IAM (Identity and Access Management) của BigQuery, sử dụng các legacy roles (Owner, Writer, Reader) ở mức dataset để đảm bảo nguyên tắc least privilege (quyền tối thiểu cần thiết). Kiến thức dựa trên tài liệu GCP cập nhật đến năm 2026 (BigQuery IAM không thay đổi cơ bản từ 2023-2026, vẫn hỗ trợ legacy roles song song với predefined roles).
📘 Tài liệu tham khảo:
- BigQuery Access Control (Google Cloud Docs).
- BigQuery IAM Roles (Legacy roles: Owner/Writer/Reader).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create a dataset for each department. Assign the department leads the role of WRITER, and assign the data analysts the role of READER on their dataset.
Lý do chi tiết:
- 🛡️ Tạo dataset riêng cho mỗi bộ phận: Đảm bảo cô lập dữ liệu hoàn hảo – mỗi dept chỉ thấy tables trong dataset của mình, không ảnh hưởng lẫn nhau.
- 👨💼 Leads được role WRITER: Cho phép create/update/delete tables/views, load data, và chia sẻ nội dung cho team (bằng cách tạo tables và team có quyền đọc). WRITER không cho phép quản lý quyền IAM (như grant roles), tránh rủi ro leads cấp quyền thừa.
- 📈 Analysts được role READER: Chỉ query và đọc dữ liệu, không modify gì (không create/update/delete tables).
- 🎯 Hoàn hảo khớp yêu cầu, tuân thủ least privilege và best practices BigQuery (dùng dataset-level roles).
❌ Phân tích tất cả các phương án (đúng/sai)
-
[SAI] Create a dataset for each department. Assign the department leads the role of OWNER, and assign the data analysts the role of WRITER on their dataset.
❌ Lý do sai:- OWNER cho leads quá mạnh (full access + quản lý ACL/IAM, có thể grant roles tùy ý, vi phạm least privilege).
- WRITER cho analysts cho phép modify (create/update/delete tables), trái với yêu cầu "không modify data". Dataset đúng nhưng roles sai.
-
[ĐÚNG] Create a dataset for each department. Assign the department leads the role of WRITER, and assign the data analysts the role of READER on their dataset.
✅ Lý do đúng: Như phân tích ở phần trên – cân bằng quyền, cô lập dữ liệu, khớp 100% yêu cầu. -
[SAI] Create a table for each department. Assign the department leads the role of Owner, and assign the data analysts the role of Editor on the project the table is in.
❌ Lý do sai:- Tạo table riêng không cô lập dữ liệu hiệu quả (tables thuộc dataset chung, dept khác vẫn có thể access nếu có quyền dataset/project).
- Roles ở project level (Owner/Editor) quá rộng – cho access toàn bộ project (bao gồm nhiều dataset khác), không giới hạn "only their data". Editor tương đương WRITER nhưng project-wide.
-
[SAI] Create a table for each department. Assign the department leads the role of Editor, and assign the data analysts the role of Viewer on the project the table is in.
❌ Lý do sai:- Tương tự phương án trước: Table không phải unit cô lập, project-level roles (Editor ~ WRITER, Viewer ~ READER) áp dụng toàn project, vi phạm "each department only their data". Leads vẫn modify được nhiều thứ ngoài ý muốn.
🧠 Kết luận: Sử dụng dataset + WRITER/READER là best practice GCP đến 2026, tránh project-level để giảm blast radius bảo mật! 🚀
- A Change the row key syntax in your Cloud Bigtable table to begin with the stock symbol.
- B Change the row key syntax in your Cloud Bigtable table to begin with a random number per second.
- C Change the data pipeline to use BigQuery for storing stock trades, and update your application.
- D Use Cloud Dataflow to write a summary of each day's stock trades to an Avro file on Cloud Storage. Update your application to read from Cloud Storage and Cloud Bigtable to compute the responses.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một hệ thống lưu trữ dữ liệu giao dịch chứng khoán (stock trades) trong Cloud Bigtable (một cơ sở dữ liệu NoSQL phân tán của Google Cloud, tối ưu cho workload lớn với độ trễ thấp). Dữ liệu được tổ chức với row key bắt đầu bằng datetime (thời gian giao dịch). Ứng dụng tính giá cổ phiếu trung bình cho một công ty cụ thể trong khoảng thời gian linh hoạt (adjustable window), phục vụ hàng nghìn người dùng đồng thời. Vấn đề: Hiệu suất giảm dần khi thêm nhiều cổ phiếu hơn, do hot spotting (tập trung truy vấn vào một số row key gần nhau, gây quá tải node).
Mục tiêu: Cải thiện hiệu suất bằng cách tối ưu hóa row key design hoặc kiến trúc dữ liệu, phù hợp với đặc thù của Bigtable (phân vùng theo row key prefix, hỗ trợ range scan nhanh). Kiến thức dựa trên tài liệu Google Cloud Bigtable cập nhật đến 2026 (Bigtable v2 với cải tiến partitioning và hot spot detection).
📘 Tài liệu tham khảo:
- Bigtable Design Patterns (Google Cloud Docs).
- Bigtable Performance Tuning (cập nhật 2025, nhấn mạnh row key để tránh hot spots).
✅ Đáp án đúng
Change the row key syntax in your Cloud Bigtable table to begin with the stock symbol.
Lý do chọn: Row key hiện tại bắt đầu bằng datetime gây hot spotting (nhiều giao dịch cùng thời điểm đổ vào cùng node, đặc biệt với truy vấn range theo thời gian cho một stock). Thay bằng stock symbol (ví dụ: "AAPL#20230101#...") làm prefix giúp phân bố đều dữ liệu theo cổ phiếu, mỗi stock vào node riêng. Ứng dụng query theo symbol + thời gian sẽ range scan hiệu quả chỉ trên node liên quan, hỗ trợ hàng nghìn concurrent users. Đây là best practice của Bigtable cho time-series data với multi-tenant (nhiều stock). 🛠️ Kết quả: Giảm latency 10-100x, tăng throughput.
📋 Giải thích chi tiết tất cả các phương án
-
✅ [ĐÚNG] Change the row key syntax in your Cloud Bigtable table to begin with the stock symbol.
Phương án này hoàn toàn đúng vì nó giải quyết gốc rễ vấn đề hot spotting bằng cách sử dụng high-cardinality prefix (stock symbol có nhiều giá trị duy nhất). Bigtable tự động phân vùng theo prefix 10-50 bytes đầu, giúp query "average price cho stock X trong window Y-Z" chỉ scan subset dữ liệu liên quan, không ảnh hưởng toàn table. Không cần thay đổi schema lớn, chỉ update row key khi ingest data mới. Hiệu quả cao cho workload này (time-series per entity). -
❌ [SAI] Change the row key syntax in your Cloud Bigtable table to begin with a random number per second.
Phương án này sai vì random number chỉ phân tán write workload tạm thời, nhưng không hỗ trợ range scan hiệu quả cho query theo thời gian (ứng dụng cần đọc sequential datetime). Random hóa làm row key rối loạn, query window thời gian sẽ full table scan chậm, tăng chi phí và latency. Bigtable docs cảnh báo tránh random keys cho read-heavy apps. -
❌ [SAI] Change the data pipeline to use BigQuery for storing stock trades, and update your application.
Phương án này sai vì BigQuery là data warehouse cho analytics/batch queries (SQL trên petabyte data), không phù hợp real-time low-latency (sub-second) với hàng nghìn concurrent users. BigQuery có cold start latency ~1-10s/query, không thay thế Bigtable (OLTP/OLAP hybrid). Chuyển toàn bộ sẽ tăng complexity và chi phí, vi phạm yêu cầu performance cho adjustable windows. -
❌ [SAI] Use Cloud Dataflow to write a summary of each day's stock trades to an Avro file on Cloud Storage. Update your application to read from Cloud Storage and Cloud Bigtable to compute the responses.
Phương án này sai vì thêm batch processing (Dataflow daily summary) tạo hybrid read từ Storage (slow, object storage không index) + Bigtable, tăng latency và complexity. Avro trên Storage phù hợp archival, không cho real-time average (phải join on-the-fly). Không giải quyết hot spotting gốc, chỉ làm chậm hơn với I/O kép. Bigtable best practice: Giữ nguyên NoSQL schema thay vì pre-aggregate trừ khi query rất coarse.
Observability to ensure that it is processing data. Which Observability alerts should you create?
- A An alert based on a decrease of subscription/num_undelivered_messages for the source and a rate of change increase of instance/storage/ used_bytes for the destination
- B An alert based on an increase of subscription/num_undelivered_messages for the source and a rate of change decrease of instance/storage/ used_bytes for the destination
- C An alert based on a decrease of instance/storage/used_bytes for the source and a rate of change increase of subscription/ num_undelivered_messages for the destination
- D An alert based on an increase of instance/storage/used_bytes for the source and a rate of change decrease of subscription/ num_undelivered_messages for the destination
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh việc giám sát một pipeline streaming trên Cloud Dataflow (Google Cloud). Pipeline này:
- Nguồn dữ liệu (source): Subscription của Cloud Pub/Sub, với throughput ổn định (luồng dữ liệu đều đặn).
- Xử lý: Tổng hợp (aggregate) các sự kiện trong một window (cửa sổ thời gian).
- Đích đến (sink): Cloud Storage bucket. Mục tiêu: Tạo alert trên Cloud Observability (nay là Cloud Monitoring) để phát hiện sớm khi pipeline KHÔNG xử lý dữ liệu (ví dụ: bị tắc nghẽn, lỗi, hoặc dừng).
📈 Logic giám sát lý tưởng: - Nếu pipeline đang chạy tốt → Backlog ở Pub/Sub giảm (message được xử lý nhanh), và dung lượng bucket tăng dần (dữ liệu được sink).
- Nếu có vấn đề → Backlog Pub/Sub tăng (subscription/num_undelivered_messages ↑), dung lượng bucket không tăng (instance/storage/used_bytes rate of change ↓).
🛠️ Kiến thức cập nhật (2026): Theo tài liệu Google Cloud mới nhất, metric subscription/num_undelivered_messages đo số message chưa được ack (backlog), và instance/storage/used_bytes (hoặc tương đương storage.googleapis.com/storage/total_bytes cho bucket) theo dõi dung lượng sử dụng. Sử dụng Cloud Monitoring với alerting policy dựa trên threshold hoặc rate of change (từ phiên bản Dataflow 2.x+).
Nguồn tham khảo:
- 📘 Cloud Dataflow Streaming Monitoring (Google Cloud Docs, cập nhật 2025).
- 📘 Pub/Sub Metrics.
- 📘 Cloud Storage Metrics.
- 📘 Cloud Monitoring Alerting.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng:
An alert based on an increase of subscription/num_undelivered_messages for the source and a rate of change decrease of instance/storage/ used_bytes for the destination
Lý do 🏆:
- Tăng subscription/num_undelivered_messages (source Pub/Sub): Chỉ ra backlog tích tụ vì pipeline không consume message kịp (chưa ack).
- Giảm rate of change của instance/storage/used_bytes (destination bucket): Dung lượng bucket không tăng nữa, chứng tỏ không có dữ liệu mới được sink.
Kết hợp 2 metric này tạo alert chính xác khi pipeline bị kẹt, phù hợp throughput ổn định. Đây là best practice từ Google Cloud để detect processing stall.
🔍 Giải thích tất cả các phương án (đúng/sai)
-
❌ Phương án SAI:
An alert based on a decrease of subscription/num_undelivered_messages for the source and a rate of change increase of instance/storage/ used_bytes for the destination
Giải thích: Giảm num_undelivered_messages (backlog giảm → tốt, pipeline đang xử lý), tăng rate of change used_bytes (bucket nhận data → tốt). Đây là dấu hiệu pipeline khỏe mạnh, không nên alert (sẽ spam false positive). -
✅ Phương án ĐÚNG:
An alert based on an increase of subscription/num_undelivered_messages for the source and a rate of change decrease of instance/storage/ used_bytes for the destination
Giải thích: Như trên, tăng backlog source + giảm tốc độ tăng dung lượng destination → pipeline KHÔNG xử lý, cần alert ngay để can thiệp. -
❌ Phương án SAI:
An alert based on a decrease of instance/storage/used_bytes for the source and a rate of change increase of subscription/ num_undelivered_messages for the destination
Giải thích:- Source là Pub/Sub (không có "instance/storage/used_bytes" – metric này dành cho disk VM/bucket, sai ngữ cảnh).
- Tăng num_undelivered_messages (xấu) nhưng mix với decrease used_bytes source (không logic). Phương án lộn xộn metric, không detect đúng vấn đề.
-
❌ Phương án SAI:
An alert based on an increase of instance/storage/used_bytes for the source and a rate of change decrease of subscription/ num_undelivered_messages for the destination
Giải thích:- Tăng used_bytes cho source Pub/Sub (sai, Pub/Sub không dùng metric này; nó không lưu trữ như disk).
- Giảm num_undelivered_messages (tốt, backlog giảm). Phương án sai metric source và alert trên trạng thái tốt → vô nghĩa.
🧮 Tóm tắt: Chỉ phương án đúng kết hợp backlog tăng + sink chậm để alert chính xác. Sử dụng MQL (Monitoring Query Language) trong Cloud Monitoring để set policy này! 🚀