Ngân hàng đề — Google Cloud Professional Data Engineer

Tìm thấy 429 câu.

Câu 191
You work for a bank. You have a labelled dataset that contains information on already granted loan application and whether these applications have been defaulted. You have been asked to train a model to predict default rates for credit applicants.
What should you do?
  1. A Increase the size of the dataset by collecting additional data.
  2. B Train a linear regression to predict a credit default risk score.
  3. C Remove the bias from the data and collect applications that have been declined loans.
  4. D Match loan applicants with their social profiles to enable feature engineering.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong lĩnh vực ngân hàng và rủi ro tín dụng (credit risk). Bạn có một dataset đã được gắn nhãn (labelled dataset) chứa thông tin về các đơn vay đã được phê duyệt (already granted loan applications) và tình trạng có default (mặc định trả nợ) hay không. Nhiệm vụ là train một mô hình để dự đoán tỷ lệ default (default rates) cho các người nộp đơn vay mới (credit applicants).

🔍 Điểm quan trọng cần lưu ý:

  • Dataset chỉ bao gồm các đơn vay đã được grant, nghĩa là không có dữ liệu từ các đơn bị từ chối (declined loans). Điều này tạo ra selection bias tự nhiên vì chỉ observe được default trên những đơn được approve (những đơn declined không được vay nên không thể default).
  • Mục tiêu là predict default risk score cho applicants mới trước khi quyết định grant, dựa trên đặc trưng (features) như thu nhập, lịch sử tín dụng, v.v.
  • Trong thực tế ML (áp dụng kiến thức AWS SageMaker mới nhất đến 2026, như SageMaker Canvas và Clarify cho bias detection), đây là bài toán classification hoặc regression cho risk score, thường train trên dữ liệu approved-only để tránh unobserved outcomes.

📘 Tài liệu tham khảo:

  • AWS SageMaker Documentation: "Built-in algorithms for credit risk modeling" (cập nhật 2025-2026, hỗ trợ Linear Learner cho regression scores).
  • AWS ML Best Practices: "Handling selection bias in lending models" (SageMaker JumpStart models for fraud/risk, nhấn mạnh train trên observed defaults).

✅ Đáp án đúng: Train a linear regression to predict a credit default risk score.

Lý do lựa chọn 🛠️:

  • Dataset hiện tại đã đủ để train model vì default chỉ có thể observe trên granted loans. Train trực tiếp trên dữ liệu này sẽ tạo ra risk score (điểm rủi ro tín dụng, thường là giá trị liên tục từ 0-1 đại diện probability default) cho applicants mới.
  • Linear Regression (Linear Learner trong AWS SageMaker) phù hợp để predict continuous risk score, dễ interpret và scale. AWS khuyến nghị cho credit scoring vì đơn giản, nhanh, và tránh overfitting trên dataset labelled (theo SageMaker built-in algorithms 2026).
  • Không cần thu thập thêm data hoặc chỉnh sửa bias phức tạp, vì bias ở đây là intentional (chỉ predict trên eligible applicants). Model sẽ giúp bank quyết định approve/decline dựa trên score threshold.

❌ Phân tích tất cả các phương án

Dưới đây là giải thích từng phương án một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi cái được đánh giá đúng/sai dựa trên best practices ML cho credit risk (AWS SageMaker và Fairlearn integration đến 2026).

  • Increase the size of the dataset by collecting additional data.
    ❌ Sai: Việc tăng kích thước dataset bằng cách thu thập thêm data không giải quyết vấn đề cốt lõi (selection bias). Thêm data tương tự (chỉ granted loans) chỉ làm dataset lớn hơn nhưng không cải thiện khả năng predict cho declined cases. AWS khuyên chỉ collect thêm nếu underfitting rõ rệt, không phải trường hợp này (SageMaker Data Wrangler docs 2026).

  • Train a linear regression to predict a credit default risk score.
    ✅ Đúng (như đã giải thích ở trên): Phương án tối ưu và trực tiếp, tận dụng dataset có sẵn để build risk model nhanh chóng, phù hợp với quy trình production ML trên AWS (SageMaker endpoints for real-time scoring).

  • Remove the bias from the data and collect applications that have been declined loans.
    ❌ Sai: Không thể "remove bias" đơn giản vì declined loans không có nhãn default (unobserved outcome – họ không vay nên không default). Thu thập declined data chỉ thêm noise, dẫn đến label bias và model kém chính xác. AWS Clarify (2026) detect bias nhưng không recommend mix observed/unobserved data cho lending (vi phạm causal inference principles).

  • Match loan applicants with their social profiles to enable feature engineering.
    ❌ Sai: Việc khớp dữ liệu với social profiles vi phạm quy định bảo mật (GDPR/CCPA) và fair lending laws (như ECOA ở Mỹ). Có thể giới thiệu privacy risks và bias mới (social data thường noisy/unreliable). AWS khuyến cáo tránh external data như vậy trong SageMaker Feature Store để tuân thủ compliance (updated 2026 guidelines).

🧠 Kết luận: Chọn đáp án đúng giúp deploy model nhanh trên AWS SageMaker, đảm bảo scalable và ethical cho production! 🚀

Câu 192
You need to migrate a 2TB relational database to Google Cloud Platform. You do not have the resources to significantly refactor the application that uses this database and cost to operate is of primary concern.
Which service do you select for storing and serving your data?
  1. A Cloud Spanner
  2. B Cloud Bigtable
  3. C Cloud Firestore
  4. D Cloud SQL
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi yêu cầu chọn dịch vụ phù hợp để di chuyển (migrate) một cơ sở dữ liệu quan hệ (relational database) dung lượng 2TB lên Google Cloud Platform (GCP). Các ràng buộc chính bao gồm:

  • Không có nguồn lực để refactor đáng kể ứng dụng đang sử dụng cơ sở dữ liệu này (nghĩa là ứng dụng hiện tại cần tương thích ngay mà không thay đổi lớn).
  • Chi phí vận hành là ưu tiên hàng đầu (cost to operate is of primary concern).

📘 Bối cảnh: Đây là tình huống điển hình khi migrate RDBMS truyền thống (như MySQL, PostgreSQL) lên cloud, cần dịch vụ managed relational database hỗ trợ quy mô lớn (2TB), dễ dàng migrate mà không thay đổi schema hoặc query, và chi phí thấp so với các dịch vụ NoSQL hoặc globally distributed cao cấp. Kiến thức dựa trên tài liệu GCP cập nhật đến 2026 (phiên bản Cloud SQL Enterprise Plus hỗ trợ instance lên đến 128TB, với tính năng autoscaling và high availability giá rẻ).

✅ Đáp án đúng: Cloud SQL

Lý do lựa chọn:
Cloud SQL là dịch vụ managed relational database của GCP, hỗ trợ các engine phổ biến như MySQL, PostgreSQL, SQL Server. Nó lý tưởng cho migrate 2TB RDBMS vì:

  • Tương thích hoàn hảo với ứng dụng hiện tại (không cần refactor code, schema, hoặc query SQL chuẩn).
  • Chi phí thấp nhất cho workload relational thông thường: Giá dựa trên vCPU, RAM, storage (khoảng 0.17 USD/giờ cho instance nhỏ, storage 0.17 USD/GB/tháng – rẻ hơn Spanner 5-10 lần). Hỗ trợ 2TB dễ dàng với High Availability (HA) và backup tự động.
  • Dễ migrate: Sử dụng Database Migration Service (DMS) để chuyển dữ liệu trực tiếp từ on-prem hoặc AWS RDS mà không downtime lớn.
    🛠️ Nguồn tham khảo: Cloud SQL Pricing & Migrate to Cloud SQL (cập nhật 2025-2026 với hỗ trợ AI-optimized instances).

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên yêu cầu câu hỏi (relational DB, no refactor, low cost).

  • Cloud Spanner
    ❌ Sai: Cloud Spanner là dịch vụ globally distributed relational database với strong consistency toàn cầu, phù hợp cho workload lớn cần horizontal scale (hàng PB). Tuy nhiên, nó đắt đỏ (giá node từ 0.9 USD/giờ, tổng chi phí cao gấp 5-10 lần Cloud SQL cho 2TB), yêu cầu refactor schema để dùng interleaved tables và Spanner SQL dialect (không chuẩn 100%). Không ưu tiên cho migrate đơn giản + low cost.
    🛠️ Nguồn: Cloud Spanner Pricing.

  • Cloud Bigtable
    ❌ Sai: Cloud Bigtable là NoSQL wide-column store (tương tự Cassandra/HBase), dành cho big data analytics với throughput cực cao (hàng triệu QPS). Nó không hỗ trợ SQL quan hệ (chỉ NoSQL API), buộc phải refactor hoàn toàn ứng dụng để dùng key-value model – vi phạm ràng buộc no refactor. Chi phí cũng cao cho relational workload (từ 0.65 USD/node/giờ). Không phù hợp migrate RDBMS 2TB.
    🛠️ Nguồn: Bigtable Overview.

  • Cloud Firestore
    ❌ Sai: Cloud Firestore là document NoSQL database (real-time, mobile-first), hỗ trợ queries linh hoạt nhưng không phải relational (không có JOIN, foreign keys chuẩn). Migrate 2TB RDBMS sẽ yêu cầu refactor lớn schema thành documents, và chi phí đọc/ghi cao (0.06 USD/100K reads) cho workload transactional. Không tối ưu low cost + no refactor cho relational app.
    🛠️ Nguồn: Firestore Pricing.

Tóm lại, Cloud SQL là lựa chọn tối ưu nhất ✅ nhờ tính managed, tương thích SQL chuẩn, và chi phí thấp cho migrate relational DB quy mô 2TB! 🚀

Câu 193
You're using Bigtable for a real-time application, and you have a heavy load that is a mix of read and writes. You've recently identified an additional use case and need to perform hourly an analytical job to calculate certain statistics across the whole database. You need to ensure both the reliability of your production application as well as the analytical workload.
What should you do?
  1. A Export Bigtable dump to GCS and run your analytical job on top of the exported files.
  2. B Add a second cluster to an existing instance with a multi-cluster routing, use live-traffic app profile for your regular workload and batch-analytics profile for the analytics workload.
  3. C Add a second cluster to an existing instance with a single-cluster routing, use live-traffic app profile for your regular workload and batch-analytics profile for the analytics workload.
  4. D Increase the size of your existing cluster twice and execute your analytics workload on your new resized cluster.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh việc sử dụng Google Cloud Bigtable (một cơ sở dữ liệu NoSQL columnar của GCP) cho ứng dụng real-time với tải trọng nặng kết hợp read và write. Gần đây, có thêm use case phân tích hàng giờ (analytical job) để tính toán thống kê trên toàn bộ database. Yêu cầu chính là đảm bảo độ tin cậy (reliability) cho cả ứng dụng sản xuất (production app) lẫn workload phân tích, tránh ảnh hưởng lẫn nhau.

🔍 Chi tiết vấn đề:

  • Production app cần latency thấp, consistent cho read/write real-time.
  • Analytical job là batch hourly, quét lớn dữ liệu toàn DB → có thể gây hotspot hoặc overload nếu chạy chung cluster.
  • Giải pháp cần scale riêng biệt, tận dụng tính năng multi-cluster và app profiles của Bigtable để tách traffic mà không làm gián đoạn production.

📘 Kiến thức cập nhật (Bigtable phiên bản mới nhất 2024-2026): Bigtable hỗ trợ instances với nhiều clusters (replicas), app profiles định nghĩa routing (single/multi-cluster). Single-cluster routing chỉ định traffic đến cluster cụ thể, lý tưởng cho workloads riêng biệt. Multi-cluster routing dùng cho replication nhưng có thể gây interference nếu scan lớn.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Add a second cluster to an existing instance with a single-cluster routing, use live-traffic app profile for your regular workload and batch-analytics profile for the analytics workload.

🛠️ Lý do chi tiết:

  • Thêm second cluster (replica) vào instance hiện tại → replicate data tự động, đảm bảo high availability và durability mà không downtime.
  • Sử dụng single-cluster routing trong app profiles:
    • live-traffic profile: Route đến cluster chính → production real-time giữ latency thấp, tránh scan lớn ảnh hưởng.
    • batch-analytics profile: Route đến cluster phụ → analytical job quét toàn DB thoải mái, không overload production.
  • Điều này tuân thủ best practice GCP: Tách workloads bằng app profiles + single-cluster để isolate traffic, hỗ trợ analytical mà không ảnh hưởng real-time (theo docs Bigtable 2024+).

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn (giữ nguyên văn bản gốc tiếng Anh). Mỗi phương án được đánh giá ✅ đúng hoặc ❌ sai, kèm lý do bằng tiếng Việt:

  • ❌ [SAI] Export Bigtable dump to GCS and run your analytical job on top of the exported files.
    🧨 Lý do sai: Export dump (qua Bigtable export tool) chỉ snapshot tại thời điểm, không real-time → analytical job hàng giờ sẽ lạc dữ liệu mới. Quá trình export tốn thời gian/thông lượng lớn, ảnh hưởng production. Không scale tốt cho DB lớn, phải dùng công cụ ngoài như Dataflow/BigQuery để process files → phức tạp, không efficient.

  • ❌ [SAI] Add a second cluster to an existing instance with a multi-cluster routing, use live-traffic app profile for your regular workload and batch-analytics profile for the analytics workload.
    🧨 Lý do sai: Multi-cluster routing cho phép read từ nearest replica, write sync leader → analytical scan lớn trên cluster phụ vẫn có thể fan-out hoặc compete resources với production (qua replication lag/hotspot). Không isolate hoàn toàn workloads, vi phạm yêu cầu reliability cho real-time app. Best practice khuyên dùng single-cluster cho separation rõ ràng.

  • ✅ [ĐÚNG] Add a second cluster to an existing instance with a single-cluster routing, use live-traffic app profile for your regular workload and batch-analytics profile for the analytics workload.
    🛠️ Lý do đúng: Như đã giải thích ở phần trên. Single-cluster routing pin traffic chính xác đến cluster riêng → production dùng cluster 1 (low-latency), analytics dùng cluster 2 (batch scan thoải mái). Data replicate async, chi phí thấp, scale độc lập. Hoàn hảo cho mixed workloads.

  • ❌ [SAI] Increase the size of your existing cluster twice and execute your analytics workload on your new resized cluster.
    🧨 Lý do sai: Scale up cluster (nodes CPU/SSD) chỉ tăng capacity chung, analytical batch vẫn compete với production read/write → gây throttling, latency spike cho real-time app. Không tách biệt workloads, vi phạm reliability. Scale up tốn kém hơn replicate cluster.

📚 Tài liệu tham khảo

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần ví dụ code hoặc demo, hãy hỏi thêm.

Câu 194
You are designing an Apache Beam pipeline to enrich data from Cloud Pub/Sub with static reference data from BigQuery. The reference data is small enough to fit in memory on a single worker. The pipeline should write enriched results to BigQuery for analysis. Which job type and transforms should this pipeline use?
  1. A Batch job, PubSubIO, side-inputs
  2. B Streaming job, PubSubIO, JdbcIO, side-outputs
  3. C Streaming job, PubSubIO, BigQueryIO, side-inputs
  4. D Streaming job, PubSubIO, BigQueryIO, side-outputs
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi yêu cầu thiết kế một Apache Beam pipeline để làm giàu (enrich) dữ liệu streaming từ Cloud Pub/Sub bằng dữ liệu tham chiếu tĩnh (static reference data) từ BigQuery. Dữ liệu tham chiếu này nhỏ đủ để fit vào memory của một worker duy nhất. Pipeline sau đó sẽ ghi kết quả đã enrich vào BigQuery để phân tích.

🔍 Chi tiết chính cần lưu ý:

  • Nguồn dữ liệu chính: Cloud Pub/Sub → Đây là nguồn streaming (dữ liệu liên tục, thời gian thực), nên pipeline phải hỗ trợ streaming job (không phải batch).
  • Dữ liệu enrich: Static từ BigQuery, nhỏ → Sử dụng side-inputs (cách hiệu quả để broadcast dữ liệu nhỏ đến tất cả elements trong pipeline mà không cần lookup liên tục).
  • Đích đến: BigQuery → Sử dụng BigQueryIO để write dữ liệu đã enrich.
  • Yêu cầu transform: PubSubIO (đọc Pub/Sub), BigQueryIO (write), và side-inputs (enrich).
  • Kiến thức cập nhật (Beam 2.55+ đến 2026): Apache Beam hỗ trợ streaming với Pub/Sub và BigQuery một cách native trên Google Cloud Dataflow, side-inputs lý tưởng cho static small data (theo docs Beam 2024-2026).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Streaming job, PubSubIO, BigQueryIO, side-inputs

Lý do chi tiết 🛠️:

  • Streaming job: Pub/Sub là nguồn streaming unbounded, batch job không xử lý được dữ liệu liên tục (sẽ fail hoặc không hiệu quả).
  • PubSubIO: Transform chuẩn để đọc từ Cloud Pub/Sub.
  • BigQueryIO: Transform native để write trực tiếp vào BigQuery (hỗ trợ streaming inserts với exactly-once semantics từ Beam 2.20+).
  • Side-inputs: Hoàn hảo cho static reference data nhỏ (fit in memory single worker), Beam sẽ broadcast toàn bộ table như một MultiMapSideInput, cho phép enrich O(1) lookup trên mỗi event mà không cần query BigQuery mỗi lần (tiết kiệm chi phí và latency thấp).

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn (giữ nguyên văn bản gốc bằng tiếng Anh), đánh dấu ✅ đúng hoặc ❌ sai, kèm lý do cụ thể bằng tiếng Việt:

  • ❌ [SAI] Batch job, PubSubIO, side-inputs
    ❌ Sai vì: Batch job chỉ phù hợp dữ liệu bounded (kết thúc), Pub/Sub là unbounded streaming → Pipeline sẽ không xử lý real-time data, có thể stuck hoặc fail khi checkpoint. Side-inputs đúng nhưng job type sai hoàn toàn.

  • ❌ [SAI] Streaming job, PubSubIO, JdbcIO, side-outputs
    ❌ Sai vì: JdbcIO dùng cho relational DB như MySQL/PostgreSQL, không phải BigQuery (BigQuery dùng BigQueryIO). Side-outputs dùng để split stream thành multiple outputs (ví dụ tag-based), không phải enrich reference data. Streaming và PubSubIO đúng nhưng các phần còn lại không khớp.

  • ✅ [ĐÚNG] Streaming job, PubSubIO, BigQueryIO, side-inputs
    ✅ Đúng hoàn toàn như giải thích ở phần trên: Kết hợp lý tưởng cho streaming enrich với static small data và write BigQuery. Hiệu suất cao trên Dataflow (runner mặc định GCP).

  • ❌ [SAI] Streaming job, PubSubIO, BigQueryIO, side-outputs
    ❌ Sai vì: Side-outputs dùng cho phân nhánh output (side output streams), không phải input để enrich. Dùng side-outputs ở đây sẽ không broadcast reference data, dẫn đến enrich thất bại hoặc không hiệu quả (phải dùng lookup riêng, kém hơn side-inputs).

🧠 Tóm tắt nhanh: Chọn đúng để pipeline chạy streaming hiệu quả, chi phí thấp trên GCP Dataflow! 🚀

Câu 195 Chọn nhiều đáp án
You have a data pipeline that writes data to Cloud Bigtable using well-designed row keys. You want to monitor your pipeline to determine when to increase the size of your Cloud Bigtable cluster. Which two actions can you take to accomplish this? (Choose two.)
  1. A Review Key Visualizer metrics. Increase the size of the Cloud Bigtable cluster when the Read pressure index is above 100.
  2. B Review Key Visualizer metrics. Increase the size of the Cloud Bigtable cluster when the Write pressure index is above 100.
  3. C Monitor the latency of write operations. Increase the size of the Cloud Bigtable cluster when there is a sustained increase in write latency.
  4. D Monitor storage utilization. Increase the size of the Cloud Bigtable cluster when utilization increases above 70% of max capacity.
  5. E Monitor latency of read operations. Increase the size of the Cloud Bigtable cluster of read operations take longer than 100 ms.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi tập trung vào việc giám sát pipeline dữ liệu ghi dữ liệu vào Cloud Bigtable (dịch vụ NoSQL database của Google Cloud, được thiết kế cho workload lớn với throughput cao). Pipeline sử dụng row keys được thiết kế tốt (well-designed row keys), giúp tránh hotspots. Mục tiêu là xác định thời điểm cần tăng kích thước cluster Bigtable (thêm nodes để scale horizontally).
Câu hỏi yêu cầu chọn hai hành động (actions) phù hợp để giám sát và quyết định scale cluster.
Bối cảnh quan trọng: Bigtable tự động scale theo CPU, nhưng người dùng cần chủ động monitor các metrics như latency, CPU, storage để tránh bottleneck khi workload tăng (dữ liệu ghi liên tục từ pipeline). Không nên scale chỉ dựa trên hotspots (Key Visualizer), vì row keys đã tốt.
📘 Kiến thức cập nhật (GCP Bigtable đến 2026): Theo tài liệu chính thức Google Cloud (phiên bản mới nhất 2024-2026), scaling cluster dựa trên sustained high latency, CPU >60%, storage utilization cao (per node ~70-80%). Key Visualizer chỉ để detect/design row keys, không dùng để scale trực tiếp.

✅ Đáp án đúng (Chọn hai phương án sau)

  • Monitor the latency of write operations. Increase the size of the Cloud Bigtable cluster when there is a sustained increase in write latency.
  • Monitor storage utilization. Increase the size of the Cloud Bigtable cluster when utilization increases above 70% of max capacity.

Lý do lựa chọn:
🛠️ Hai metrics này trực tiếp chỉ ra nhu cầu scale cluster khi write-heavy workload (pipeline ghi dữ liệu) gây bottleneck. Write latency tăng sustained (> vài phút/giờ) cho thấy nodes quá tải → thêm nodes để phân tải. Storage >70% max capacity (per tserver node) dẫn đến throttling → scale ngay để tránh mất dữ liệu. Đây là best practice từ GCP docs, phù hợp với pipeline writes lớn.

📋 Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể bằng tiếng Việt, dựa trên tài liệu GCP Bigtable Monitoring & Key Visualizer (cập nhật 2026).

  • ❌ Review Key Visualizer metrics. Increase the size of the Cloud Bigtable cluster when the Read pressure index is above 100.
    Sai: Key Visualizer dùng để phát hiện hotspots (read/write pressure index >100 chỉ ra row keys kém, gây imbalance). Vì row keys đã "well-designed", pressure cao không phải do scale thiếu mà do cần redesign keys. Scale cluster ở đây chỉ lãng phí, không giải quyết gốc rễ. (Không áp dụng cho read-heavy ở pipeline writes).

  • ❌ Review Key Visualizer metrics. Increase the size of the Cloud Bigtable cluster when the Write pressure index is above 100.
    Sai: Tương tự trên, Write pressure >100 từ Key Visualizer báo hiệu hotspots trong writes, yêu cầu tối ưu row keys (như salt keys). Không scale cluster vì Bigtable đã tự balance nếu keys tốt; scale chỉ làm tăng chi phí vô ích.

  • ✅ Monitor the latency of write operations. Increase the size of the Cloud Bigtable cluster when there is a sustained increase in write latency.
    Đúng: Write latency tăng sustained (qua Cloud Monitoring metrics như bigtable.googleapis.com/table/write_latency) là dấu hiệu nodes quá tải CPU/QPS. Pipeline writes lớn → scale cluster để giảm latency xuống <10ms (target Bigtable). Best practice cho write-heavy workloads.

  • ✅ Monitor storage utilization. Increase the size of the Cloud Bigtable cluster when utilization increases above 70% of max capacity.
    Đúng: Storage utilization >70% (metrics bigtable.googleapis.com/server/disk_utilization per node) gây write throttling (SSD/HDD limits). Bigtable quota ~10TB/node (SSD); vượt 70% → thêm nodes để phân tán storage, tránh mất dữ liệu. Threshold 70% là khuyến nghị GCP cho proactive scaling.

  • ❌ Monitor latency of read operations. Increase the size of the Cloud Bigtable cluster of read operations take longer than 100 ms.
    Sai: Read latency (thường <10ms ở Bigtable) có thể cao do scan lớn hoặc cache miss, không phải lý do chính scale cluster (pipeline focus writes). Threshold 100ms không chuẩn (GCP không định nghĩa fixed ms cho scale); ưu tiên CPU/latency tổng trước. Có thể fix bằng tăng replicas thay vì nodes.

📘 Tài liệu tham khảo

  • Bigtable Monitoring (GCP Docs 2026): Metrics cho latency/storage/CPU.
  • Key Visualizer : Chỉ để fix hotspots, không scale.
  • Scaling Clusters : CPU>60%, latency sustained, storage high → increase nodes.
  • Cloud Monitoring Dashboard samples cho Bigtable (tích hợp Prometheus-style metrics).

Hy vọng phân tích giúp bạn ôn thi certification GCP Data Engineer! 🚀 Nếu cần thêm ví dụ thực tế, hỏi nhé!

Câu 196
You want to analyze hundreds of thousands of social media posts daily at the lowest cost and with the fewest steps.
You have the following requirements:
✑ You will batch-load the posts once per day and run them through the Cloud Natural Language API.
✑ You will extract topics and sentiment from the posts.
✑ You must store the raw posts for archiving and reprocessing.
✑ You will create dashboards to be shared with people both inside and outside your organization.
You need to store both the data extracted from the API to perform analysis as well as the raw social media posts for historical archiving. What should you do?
  1. A Store the social media posts and the data extracted from the API in BigQuery.
  2. B Store the social media posts and the data extracted from the API in Cloud SQL.
  3. C Store the raw social media posts in Cloud Storage, and write the data extracted from the API into BigQuery.
  4. D Feed to social media posts into the API directly from the source, and write the extracted data from the API into BigQuery.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc thiết kế một hệ thống xử lý dữ liệu hàng trăm nghìn bài đăng mạng xã hội mỗi ngày với chi phí thấp nhất và ít bước nhất trên Google Cloud Platform (GCP). Các yêu cầu cụ thể bao gồm:

  • Batch-load dữ liệu một lần mỗi ngày và chạy qua Cloud Natural Language API để trích xuất topics (chủ đề) và sentiment (cảm xúc).
  • Lưu trữ raw posts (dữ liệu thô) để archiving (lưu trữ lịch sử) và reprocessing (xử lý lại nếu cần).
  • Tạo dashboards chia sẻ nội/ngoại bộ tổ chức dựa trên dữ liệu đã trích xuất.
  • Cần lưu trữ cả dữ liệu thô (raw social media posts) và dữ liệu đã trích xuất (extracted data từ API) để phân tích.

Mục tiêu chính: Tối ưu chi phí lưu trữ (raw data lớn, unstructured), dễ phân tích (extracted data có cấu trúc), và hỗ trợ dashboard (như Looker Studio hoặc Connected Sheets). Đây là bài toán data lake + data warehouse điển hình trên GCP, với kiến thức cập nhật đến 2026 (BigQuery hỗ trợ AI/ML tích hợp tốt hơn, Cloud Storage với lifecycle policies tự động).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Store the raw social media posts in Cloud Storage, and write the data extracted from the API into BigQuery.

Lý do 🛠️:

  • Cloud Storage lý tưởng cho raw posts (unstructured text lớn): Chi phí thấp (Nearline ~$0.001/GB/tháng), scalable vô hạn, hỗ trợ versioning/archiving, dễ batch-load hàng ngày và reprocessing (chỉ cần đọc file JSON/CSV).
  • BigQuery hoàn hảo cho extracted data (topics, sentiment - semi-structured): Serverless analytics, SQL queries nhanh cho hàng triệu rows, tích hợp trực tiếp Looker Studio/Data Studio để tạo dashboards chia sẻ (public sharing qua links).
  • Lowest cost & fewest steps: Batch-load raw → API → extract → load BigQuery (1 pipeline đơn giản với Dataflow hoặc Cloud Functions). Tổng chi phí thấp vì raw data ở storage rẻ, chỉ query extracted data thường xuyên.
  • Phù hợp batch daily và reprocessing (raw ở Storage dễ truy xuất lại).

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Store the social media posts and the data extracted from the API in BigQuery.
    Phân tích: BigQuery không phù hợp lưu raw posts lớn (text unstructured, hàng trăm nghìn posts/ngày → storage đắt ~$0.02/GB/tháng, on-demand scanning tốn query cost). Raw data chỉ cần archiving, không cần query thường xuyên → lãng phí chi phí cao, vi phạm "lowest cost". Extracted data thì OK, nhưng tổng thể không tối ưu.

  • ❌ [SAI] Store the social media posts and the data extracted from the API in Cloud SQL.
    Phân tích: Cloud SQL (MySQL/PostgreSQL) là RDBMS relational, kém scalable cho big data unstructured (hàng trăm nghìn rows/ngày → cần sharding phức tạp, chi phí cao ~$0.1+/instance/tháng). Không hỗ trợ archiving rẻ, khó batch-load lớn, và không lý tưởng cho analytics/dashboard (query chậm so BigQuery). Vi phạm "fewest steps" và "lowest cost".

  • ✅ [ĐÚNG] Store the raw social media posts in Cloud Storage, and write the data extracted from the API into BigQuery.
    Phân tích: Như đã giải thích ở trên – kết hợp hoàn hảo data lake (Storage cho raw) + data warehouse (BigQuery cho analysis). Hỗ trợ đầy đủ yêu cầu: archiving/reprocessing (Storage lifecycle auto-tier to Archive), dashboards (BigQuery + Looker), batch API processing. Tối ưu nhất theo best practices GCP 2026.

  • ❌ [SAI] Feed to social media posts into the API directly from the source, and write the extracted data from the API into BigQuery.
    Phân tích: Bỏ qua lưu raw posts → không đáp ứng "store raw for archiving and reprocessing" (không thể xử lý lại nếu API thay đổi hoặc lỗi). "Direct feed" không phù hợp batch-load daily, có thể tốn kém real-time processing, và thiếu historical raw data cho audits/compliance. Chỉ lưu extracted → mất tính linh hoạt.

Kết luận 🎯: Giải pháp đúng tận dụng hybrid storage của GCP để cân bằng chi phí, hiệu suất và tính năng, phù hợp certification Professional Data Engineer!

Câu 197
You store historic data in Cloud Storage. You need to perform analytics on the historic data. You want to use a solution to detect invalid data entries and perform data transformations that will not require programming or knowledge of SQL.
What should you do?
  1. A Use Cloud Dataflow with Beam to detect errors and perform transformations.
  2. B Use Cloud Dataprep with recipes to detect errors and perform transformations.
  3. C Use Cloud Dataproc with a Hadoop job to detect errors and perform transformations.
  4. D Use federated tables in BigQuery with queries to detect errors and perform transformations.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc xử lý dữ liệu lịch sử (historic data) được lưu trữ trong Cloud Storage trên Google Cloud Platform (GCP). Yêu cầu chính là thực hiện phân tích dữ liệu (analytics), cụ thể là phát hiện các mục dữ liệu không hợp lệ (invalid data entries) và thực hiện chuyển đổi dữ liệu (data transformations). Giải pháp phải không yêu cầu lập trình (no programming) hoặc kiến thức SQL.
✅ Mục tiêu chính: Tìm công cụ no-code/low-code, dễ sử dụng, hỗ trợ phát hiện lỗi và biến đổi dữ liệu trực quan từ dữ liệu trong Cloud Storage.
🛠️ Bối cảnh: Dữ liệu lớn, cần xử lý ETL (Extract, Transform, Load) tự động mà không cần code phức tạp. (Kiến thức cập nhật GCP 2024-2026: Cloud Dataprep đã tích hợp sâu với Trifacta, hỗ trợ AI-driven profiling và visual recipes).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Cloud Dataprep with recipes to detect errors and perform transformations.

Lý do chi tiết:

  • Cloud Dataprep (nay là phần của Google Cloud Dataflow Prep) là công cụ no-code visual data preparation dựa trên Trifacta, chuyên phát hiện lỗi dữ liệu (như anomalies, invalid entries qua data profiling và suggestions tự động) và thực hiện transformations qua recipes (các luồng visual drag-and-drop).
  • Không cần SQL hay code: Chỉ cần import dữ liệu từ Cloud Storage, sử dụng giao diện đồ họa để clean, transform, và export ra BigQuery/Dataflow.
  • Hoàn hảo cho yêu cầu: Hỗ trợ analytics trên historic data lớn, tích hợp AI/ML để detect issues tự động (cập nhật 2025: Tích hợp Gemini AI cho suggestions).
    📘 Nguồn: Cloud Dataprep Documentation, GCP Data Engineer Exam Guide 2024.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên khả năng đáp ứng no-programming/no-SQL:

  • ❌ [SAI] Use Cloud Dataflow with Beam to detect errors and perform transformations.
    Lý do sai: Cloud Dataflow sử dụng Apache Beam, yêu cầu viết code pipeline (Python/Java/Go) để detect errors và transform. Không phải no-code, phù hợp cho developer hơn là user không code. Không đáp ứng yêu cầu "no programming".

  • ✅ [ĐÚNG] Use Cloud Dataprep with recipes to detect errors and perform transformations.
    Lý do đúng: Như đã giải thích ở trên, đây là công cụ visual recipes no-code, tự động detect invalid data qua profiling (duplicate, null, outliers) và transform dễ dàng. Tích hợp trực tiếp Cloud Storage → BigQuery.

  • ❌ [SAI] Use Cloud Dataproc with a Hadoop job to detect errors and perform transformations.
    Lý do sai: Cloud Dataproc là managed Hadoop/Spark cluster, yêu cầu viết job code (Spark SQL/Python) hoặc Hive để xử lý. Không no-code, phức tạp cho analytics historic data mà không cần SQL/programming.

  • ❌ [SAI] Use federated tables in BigQuery with queries to detect errors and perform transformations.
    Lý do sai: Federated tables trong BigQuery cho phép query dữ liệu ngoài (như Cloud Storage) bằng SQL, nhưng detect errors/transform cần viết query phức tạp (CTEs, UDF). Vi phạm yêu cầu "no knowledge of SQL".

🧠 Kết luận nổi bật: Cloud Dataprep là lựa chọn tối ưu cho data engineers/citizen data scientists cần xử lý nhanh, không code. Các phương án khác đều yêu cầu kỹ năng lập trình/SQL!
📘 Tài liệu tham khảo bổ sung:

Câu 198
Your company needs to upload their historic data to Cloud Storage. The security rules don't allow access from external IPs to their on-premises resources. After an initial upload, they will add new data from existing on-premises applications every day. What should they do?
  1. A Execute gsutil rsync from the on-premises servers.
  2. B Use Dataflow and write the data to Cloud Storage.
  3. C Write a job template in Dataproc to perform the data transfer.
  4. D Install an FTP server on a Compute Engine VM to receive the files and move them to Cloud Storage.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống: Công ty cần upload dữ liệu lịch sử (historic data) lên Cloud Storage (dịch vụ lưu trữ của Google Cloud). Quy tắc bảo mật nghiêm ngặt không cho phép truy cập từ IP bên ngoài vào tài nguyên on-premises (tức là không mở port inbound từ internet vào server nội bộ). Sau lần upload ban đầu, họ sẽ thêm dữ liệu mới hàng ngày từ các ứng dụng on-premises hiện có.

📌 Yêu cầu giải quyết: Tìm phương án an toàn, hiệu quả, hỗ trợ upload ban đầu lớn + đồng bộ tăng dần (incremental) hàng ngày, mà không vi phạm quy tắc bảo mật (không cần mở inbound connection từ GCP vào on-prem). Phương án phải tận dụng kết nối outbound từ on-prem ra GCP để tránh rủi ro.

🛠️ Bối cảnh kỹ thuật (cập nhật đến 2026): Google Cloud khuyến nghị sử dụng các công cụ như gsutil cho transfer dữ liệu on-prem-to-GCS với chế độ rsync hỗ trợ delta sync (chỉ copy file mới/thay đổi). Điều này phù hợp với Google Cloud Storage Transfer best practices và gsutil version 5.x+ (hỗ trợ multi-threaded rsync nhanh hơn, checksum verification).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Execute gsutil rsync from the on-premises servers.

Lý do 🏆:

  • gsutil rsync chạy trực tiếp từ server on-premises, khởi tạo kết nối outbound đến Cloud Storage (GCS), không cần mở inbound port trên on-prem → Tuân thủ quy tắc bảo mật.
  • Hỗ trợ upload ban đầu toàn bộ dữ liệu lịch sử và đồng bộ tăng dần hàng ngày (rsync chỉ copy file mới/thay đổi dựa trên checksum/size/modtime, tiết kiệm băng thông).
  • Hiệu quả cao: Multi-threaded (-m flag), hỗ trợ resume nếu gián đoạn, authentication qua service account key hoặc ADC. Có thể schedule qua cron job hàng ngày.
  • Best practice GCP 2026: Được khuyến nghị trong docs cho hybrid transfer (on-prem to GCS) thay vì các tool phức tạp hơn như Storage Transfer Service (yêu cầu agent nếu cần pull).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên tính khả thi, bảo mật, hiệu quả và phù hợp với yêu cầu (upload initial + daily incremental, no inbound access).

  • ✅ Execute gsutil rsync from the on-premises servers.
    Đúng vì: Như giải thích trên, đây là giải pháp tối ưu nhất – đơn giản, an toàn outbound-only, hỗ trợ rsync incremental hoàn hảo cho daily adds. Không cần infra thêm.

  • ❌ Use Dataflow and write the data to Cloud Storage.
    Sai vì: Dataflow (Apache Beam trên GCP) là ETL processing service chạy trên GCP, cần pull data từ on-prem (qua JDBC/HTTP source), đòi hỏi inbound access vào on-prem → Vi phạm security rules. Không phù hợp cho simple upload, overhead cao cho daily sync, và không incremental tự nhiên mà không custom pipeline phức tạp.

  • ❌ Write a job template in Dataproc to perform the data transfer.
    Sai vì: Dataproc (managed Hadoop/Spark) chạy cluster trên GCP, job cần access nguồn dữ liệu on-prem (qua HDFS connector hoặc network), yêu cầu inbound/open firewall từ GCP → Không an toàn. Phù hợp big data processing chứ không phải simple file transfer daily; setup template phức tạp, tốn chi phí cluster.

  • ❌ Install an FTP server on a Compute Engine VM to receive the files and move them to Cloud Storage.
    Sai vì: Dù on-prem có thể connect outbound đến FTP trên GCE VM (sau copy thủ công sang GCS), nhưng FTP không an toàn (plaintext, dễ tấn công), không hỗ trợ incremental rsync tự động, cần script/move job thêm (dùng gsutil hoặc gcsfuse). Vi phạm best practice GCP (khuyến nghị HTTPS/SCP thay FTP), tốn VM chi phí, và phức tạp quản lý daily.

📘 Tài liệu tham khảo (cập nhật 2026)

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần ví dụ lệnh gsutil cụ thể, hỏi thêm nhé!

Câu 199
You have a query that filters a BigQuery table using a WHERE clause on timestamp and ID columns. By using bq query `"-dry_run you learn that the query triggers a full scan of the table, even though the filter on timestamp and ID select a tiny fraction of the overall data. You want to reduce the amount of data scanned by BigQuery with minimal changes to existing SQL queries. What should you do?
  1. A Create a separate table for each ID.
  2. B Use the LIMIT keyword to reduce the number of rows returned.
  3. C Recreate the table with a partitioning column and clustering column.
  4. D Use the bq query --maximum_bytes_billed flag to restrict the number of bytes billed.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống thực tế trong BigQuery (dịch vụ kho dữ liệu của Google Cloud):
Bạn đang chạy một truy vấn (query) lọc dữ liệu từ một bảng BigQuery bằng mệnh đề WHERE trên hai cột timestamp (thời gian) và ID. Khi sử dụng lệnh bq query --dry_run, BigQuery báo rằng truy vấn sẽ quét toàn bộ bảng (full scan), dù bộ lọc chỉ chọn ra một phần dữ liệu rất nhỏ (tiny fraction).
🎯 Mục tiêu: Giảm lượng dữ liệu bị quét (data scanned) bởi BigQuery, với thay đổi tối thiểu cho các truy vấn SQL hiện tại (minimal changes to existing SQL queries).
🛠️ Vấn đề cốt lõi: BigQuery mặc định quét toàn bộ bảng nếu không có cơ chế tối ưu hóa như partitioning hoặc clustering, dẫn đến chi phí cao và thời gian chậm, ngay cả khi filter selective.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Recreate the table with a partitioning column and clustering column.

Lý do chi tiết:

  • Partitioning (phân vùng bảng): Chọn cột timestamp làm partitioning column (thường là time-based partitioning như DATE hoặc TIMESTAMP). BigQuery sẽ tự động chia bảng thành các partition nhỏ theo thời gian, và bộ lọc WHERE trên timestamp sẽ prune (loại bỏ) các partition không liên quan, giảm đáng kể data scanned.
  • Clustering (tổ chức dữ liệu trong partition): Chọn cột ID làm clustering column. Trong mỗi partition, dữ liệu được sắp xếp và nhóm theo ID, giúp BigQuery prune clusters dựa trên filter ID, tối ưu hóa thêm mà không cần thay đổi SQL.
  • Minimal changes: Chỉ cần recreate bảng (sử dụng CREATE TABLE ... PARTITION BY timestamp CLUSTER BY ID), copy dữ liệu từ bảng cũ, và truy vấn SQL giữ nguyên – BigQuery tự động áp dụng tối ưu hóa.
  • Hiệu quả cao: Theo tài liệu BigQuery 2026, partitioning + clustering có thể giảm scanned data lên đến 99% cho filter selective như thế này.

📋 Giải thích tất cả các phương án (đúng và sai)

Dưới đây là phân tích từng phương án một cách rõ ràng. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích hoàn toàn bằng tiếng Việt với lý do đúng/sai dựa trên tính năng BigQuery mới nhất (2026).

  • ❌ Phương án SAI: Create a separate table for each ID.
    🧨 Lý do sai: Tạo bảng riêng cho từng ID không khả thi vì số lượng ID có thể rất lớn (hàng triệu), dẫn đến hàng nghìn bảng khó quản lý, không scale được. BigQuery giới hạn số bảng per project/dataset (hàng trăm nghìn), và không hỗ trợ tự động prune cross-table. Thay đổi lớn, không minimal.

  • ❌ Phương án SAI: Use the LIMIT keyword to reduce the number of rows returned.
    🚫 Lý do sai: LIMIT chỉ giới hạn số rows trả về (output), không ảnh hưởng đến lượng dữ liệu BigQuery quét (scanned) từ storage. --dry_run vẫn báo full scan vì BigQuery phải đọc toàn bộ bảng trước khi áp dụng LIMIT. Không giải quyết vấn đề gốc.

  • ✅ Phương án ĐÚNG: Recreate the table with a partitioning column and clustering column.
    🎉 Lý do đúng: Như giải thích ở phần đáp án trên. Kết hợp partitioning (prune partition theo timestamp) + clustering (prune cluster theo ID) là giải pháp chuẩn của BigQuery cho full scan trên filter multi-column. --dry_run sau khi áp dụng sẽ chỉ scan fraction nhỏ. Hỗ trợ đầy đủ trong BigQuery v2.0+ (2026), không cần thay đổi SQL.

  • ❌ Phương án SAI: Use the bq query --maximum_bytes_billed flag to restrict the number of bytes billed.
    💸 Lý do sai: Flag này chỉ giới hạn chi phí billing (nếu vượt sẽ fail query), không giảm lượng dữ liệu thực tế bị scanned. BigQuery vẫn full scan, chỉ là bạn có thể bị chặn billing chứ không tối ưu performance hoặc cost thực (vì slot usage vẫn cao).

📘 Tài liệu tham khảo (cập nhật đến 2026)

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần ví dụ code cụ thể, hãy hỏi thêm.

Câu 200
You have a requirement to insert minute-resolution data from 50,000 sensors into a BigQuery table. You expect significant growth in data volume and need the data to be available within 1 minute of ingestion for real-time analysis of aggregated trends. What should you do?
  1. A Use bq load to load a batch of sensor data every 60 seconds.
  2. B Use a Cloud Dataflow pipeline to stream data into the BigQuery table.
  3. C Use the INSERT statement to insert a batch of data every 60 seconds.
  4. D Use the MERGE statement to apply updates in batch every 60 seconds.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này thuộc lĩnh vực Google Cloud Platform (GCP), cụ thể là xử lý dữ liệu streaming thời gian thực với BigQuery – một kho dữ liệu serverless mạnh mẽ cho phân tích lớn.

Yêu cầu chính:

  • Nhập dữ liệu từ 50.000 cảm biến với độ phân giải mỗi phút (minute-resolution data), nghĩa là lượng dữ liệu lớn và tăng trưởng nhanh (significant growth).
  • Dữ liệu phải có sẵn trong vòng 1 phút sau khi ingest (ingestion) để hỗ trợ phân tích thời gian thực các xu hướng tổng hợp (real-time analysis of aggregated trends).
  • Thách thức: Cần giải pháp streaming (luồng dữ liệu liên tục) với độ trễ thấp (low latency), khả năng scale cao cho volume lớn, và đảm bảo dữ liệu sẵn sàng query ngay lập tức mà không cần batch processing chậm chạp.

Bối cảnh GCP mới nhất (cập nhật đến 2026): BigQuery hỗ trợ streaming inserts qua API với độ trễ ~1-2 giây, nhưng cho high-throughput như 50k records/phút từ nhiều nguồn, Cloud Dataflow (dựa trên Apache Beam) là lựa chọn tối ưu vì xử lý streaming scalable, exactly-once delivery, và tích hợp native với BigQuery. Không liên quan AWS như đề cập (có thể nhầm lẫn), đây thuần GCP. 📘 Nguồn: BigQuery Streaming Inserts, Dataflow BigQuery IO (phiên bản 2025+ hỗ trợ improved autoscaling).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use a Cloud Dataflow pipeline to stream data into the BigQuery table.

Lý do 🛠️:

  • Cloud Dataflow là dịch vụ stream processing managed, sử dụng Apache Beam để xử lý dữ liệu liên tục từ 50k sensors với throughput cao (hàng triệu records/giây), độ trễ <1 phút (thường vài giây).
  • Tích hợp native sink vào BigQuery với exactly-once semantics (tránh duplicate), tự động scale theo growth.
  • Dữ liệu sẵn sàng query ngay lập tức cho real-time queries/aggregations (ví dụ: materialized views hoặc BI Engine).
  • Phù hợp nhất cho real-time analysis mà không cần batch. Hoàn hảo cho scenario này! 🚀

❌ Phân tích tất cả các phương án (đúng/sai)

  • [SAI] Use bq load to load a batch of sensor data every 60 seconds.
    ❌ Sai vì: bq load là công cụ batch loading từ file (CSV/JSON), mất thời gian scan/load (vài phút cho large files), không đáp ứng <1 phút availability. Không scale tốt cho 50k sensors liên tục, dễ backlog khi growth. Phù hợp offline ETL, không real-time. 🕒

  • [ĐÚNG] Use a Cloud Dataflow pipeline to stream data into the BigQuery table.
    ✅ Đúng vì: Như giải thích trên, Dataflow xử lý streaming end-to-end với low latency, autoscaling, và BigQuery streaming sink. Dữ liệu append ngay lập tức, hỗ trợ aggregations real-time. Best practice GCP 2026! 🌟

  • [SAI] Use the INSERT statement to insert a batch of data every 60 seconds.
    ❌ Sai vì: BigQuery không hỗ trợ SQL INSERT trực tiếp cho streaming (chỉ DML cho batch queries, giới hạn 10k rows/request). Batch every 60s sẽ gây high latency và quota exceed (max 1M rows/ngày streaming inserts thủ công). Không scalable cho 50k/min. 🔒

  • [SAI] Use the MERGE statement to apply updates in batch every 60 seconds.
    ❌ Sai vì: MERGE là DML cho upsert (update/insert), chạy batch mode (mất 1-5 phút/operation lớn), không dành cho append-only sensor data. Quota nghiêm ngặt (1.5GB/query), không real-time, dễ throttle với high volume. Phù hợp CDC, không streaming ingest. ⚠️

Kết luận 💡: Chọn Dataflow để đảm bảo performance, scalability và low-latency – tiêu chuẩn cho IoT/sensor streaming vào BigQuery! Nếu implement, dùng template Dataflow "Streaming insert into BigQuery". 📘