Ngân hàng đề — Google Cloud Professional Data Engineer

Tìm thấy 429 câu.

Câu 221
You currently have a single on-premises Kafka cluster in a data center in the us-east region that is responsible for ingesting messages from IoT devices globally.
Because large parts of globe have poor internet connectivity, messages sometimes batch at the edge, come in all at once, and cause a spike in load on your
Kafka cluster. This is becoming difficult to manage and prohibitively expensive. What is the Google-recommended cloud native architecture for this scenario?
  1. A Edge TPUs as sensor devices for storing and transmitting the messages.
  2. B Cloud Dataflow connected to the Kafka cluster to scale the processing of incoming messages.
  3. C An IoT gateway connected to Cloud Pub/Sub, with Cloud Dataflow to read and process the messages from Cloud Pub/Sub.
  4. D A Kafka cluster virtualized on Compute Engine in us-east with Cloud Load Balancing to connect to the devices around the world.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả tình huống hiện tại: Bạn đang có một Kafka cluster duy nhất đặt tại on-premises trong data center vùng us-east, chịu trách nhiệm ingest (tiếp nhận) messages từ các IoT devices trên toàn cầu. 📡
Vấn đề chính: Do nhiều khu vực trên thế giới có kết nối internet kém, messages thường bị batch (gom lại) tại edge, sau đó dồn về một lúc, gây spike load (tăng đột biến tải) lên Kafka cluster. Điều này dẫn đến khó quản lý và chi phí cao.
Câu hỏi yêu cầu: Kiến trúc cloud native được Google khuyến nghị để giải quyết vấn đề này, chuyển từ on-premises sang Google Cloud một cách tối ưu. 🛤️
(Mục tiêu: Xử lý spike load scalable, chi phí thấp, cloud native – không phụ thuộc on-premises Kafka).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: An IoT gateway connected to Cloud Pub/Sub, with Cloud Dataflow to read and process the messages from Cloud Pub/Sub.

Lý do:
Kiến trúc này là cloud native hoàn chỉnh của Google Cloud, được thiết kế dành riêng cho IoT với khả năng scale tự động để xử lý spike load từ messages batched. 🟢

  • IoT gateway (qua Google Cloud IoT Core): Cho phép IoT devices kết nối an toàn qua MQTT/HTTP, buffer messages tại edge và forward đến Cloud Pub/Sub – dịch vụ messaging fully managed, auto-scale lên hàng triệu messages/giây mà không lo spike (Pub/Sub có at-least-once delivery, dead letter queue).
  • Cloud Dataflow (dựa trên Apache Beam): Đọc stream từ Pub/Sub, process batch/streaming một cách serverless, auto-scale, thay thế Kafka processing hiệu quả hơn, tiết kiệm chi phí (pay-per-use).
    Kiến trúc này giải quyết triệt để vấn đề: Không còn phụ thuộc on-premises Kafka, edge buffering tự nhiên, global distribution qua multi-region Pub/Sub. Được Google chính thức khuyến nghị cho IoT workloads (cập nhật 2024-2026).

📘 Tài liệu tham khảo:

🔍 Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng phương án một, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên kiến trúc Google Cloud mới nhất (2026).

  • ❌ [SAI] Edge TPUs as sensor devices for storing and transmitting the messages.
    Phương án này không phù hợp vì Edge TPUs (Tensor Processing Units tại edge) là phần cứng chuyên cho ML inference (suy luận AI) trên thiết bị IoT, không phải để store (lưu trữ) hay transmit (truyền) messages. 🧠 Nó chỉ accelerate model execution (như vision tại edge), không giải quyết spike load hay thay thế Kafka messaging. Sử dụng sẽ tăng chi phí hardware và phức tạp, không cloud native.

  • ❌ [SAI] Cloud Dataflow connected to the Kafka cluster to scale the processing of incoming messages.
    Phương án chỉ giải quyết một phần, vẫn giữ Kafka cluster on-premises làm nguồn ingest chính → vẫn bị spike load từ batch messages toàn cầu, gây bottleneck và chi phí cao (Kafka không auto-scale dễ dàng như Pub/Sub). Dataflow chỉ scale processing sau Kafka, không fix vấn đề gốc (ingestion). Google không khuyến nghị hybrid on-prem cho IoT scale. ⚠️

  • ✅ [ĐÚNG] An IoT gateway connected to Cloud Pub/Sub, with Cloud Dataflow to read and process the messages from Cloud Pub/Sub.
    Như đã giải thích ở trên: Hoàn hảo cloud native, IoT gateway + Pub/Sub xử lý ingestion spike, Dataflow process stream. Scalable toàn diện, global, serverless. 🎯

  • ❌ [SAI] A Kafka cluster virtualized on Compute Engine in us-east with Cloud Load Balancing to connect to the devices around the world.
    Phương án không cloud native, chỉ virtualize Kafka lên Compute Engine (VMs) ở us-east + Load Balancing → vẫn gặp spike load cục bộ (tất cả traffic dồn về một region), chi phí VM cao (không serverless), quản lý phức tạp (patching, scaling thủ công). Google ưu tiên Pub/Sub/IoT Core thay vì self-managed Kafka cho IoT. 🚫

Kết luận: Kiến trúc đúng tận dụng serverless services của Google Cloud để tự động scale, giảm chi phí và dễ quản lý IoT global. Nếu triển khai, bắt đầu từ IoT Core setup! 🚀

Câu 222 Chọn nhiều đáp án
You decided to use Cloud Datastore to ingest vehicle telemetry data in real time. You want to build a storage system that will account for the long-term data growth, while keeping the costs low. You also want to create snapshots of the data periodically, so that you can make a point-in-time (PIT) recovery, or clone a copy of the data for Cloud Datastore in a different environment. You want to archive these snapshots for a long time. Which two methods can accomplish this?
(Choose two.)
  1. A Use managed export, and store the data in a Cloud Storage bucket using Nearline or Coldline class.
  2. B Use managed export, and then import to Cloud Datastore in a separate project under a unique namespace reserved for that export.
  3. C Use managed export, and then import the data into a BigQuery table created just for that export, and delete temporary export files.
  4. D Write an application that uses Cloud Datastore client libraries to read all the entities. Treat each entity as a BigQuery table row via BigQuery streaming insert. Assign an export timestamp for each export, and attach it as an extra column for each row. Make sure that the BigQuery table is partitioned using the export timestamp column.
  5. E Write an application that uses Cloud Datastore client libraries to read all the entities. Format the exported data into a JSON file. Apply compression before storing the data in Cloud Source Repositories.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm

✅ Nội dung câu hỏi được giải thích rõ ràng:
Câu hỏi xoay quanh việc thiết kế hệ thống lưu trữ dữ liệu telemetry xe hơi (vehicle telemetry data) được ingest thời gian thực vào Cloud Datastore (một cơ sở dữ liệu NoSQL managed của Google Cloud). Yêu cầu chính bao gồm:

  • Xử lý sự tăng trưởng dữ liệu dài hạn (long-term data growth) với chi phí thấp (low costs).
  • Tạo snapshots dữ liệu định kỳ để hỗ trợ point-in-time (PIT) recovery (khôi phục dữ liệu tại một thời điểm cụ thể) hoặc clone dữ liệu sang môi trường Cloud Datastore khác.
  • Lưu trữ lâu dài (archive) các snapshots này.
    Câu hỏi yêu cầu chọn hai phương pháp (Choose two) phù hợp nhất, sử dụng tính năng managed export của Cloud Datastore để xuất dữ liệu ra các định dạng chuẩn (như CSV hoặc JSON) và lưu trữ hiệu quả. Đây là tình huống thực tế trong Google Cloud, tập trung vào tính khả dụng, chi phí và khả năng khôi phục dữ liệu (dựa trên tài liệu Datastore mới nhất đến 2026, hỗ trợ export/import với multi-region và namespace isolation).

📌 Đáp án đúng (chọn 2):
Hai phương án đúng là:

  • Use managed export, and store the data in a Cloud Storage bucket using Nearline or Coldline class.
  • Use managed export, and then import to Cloud Datastore in a separate project under a unique namespace reserved for that export.

🛠️ Lý do chọn đáp án đúng (tóm tắt):
Cả hai đều tận dụng managed export của Cloud Datastore (gọi qua gcloud hoặc API, xuất toàn bộ hoặc theo query vào Cloud Storage bucket). Phương án 1 lưu trữ rẻ tiền lâu dài với Nearline/Coldline (chi phí thấp hơn Standard, phù hợp archive). Phương án 2 cho phép import vào project riêng với namespace unique, tạo snapshot/clone độc lập hỗ trợ PIT recovery (không ảnh hưởng dữ liệu gốc). Điều này đảm bảo scalability, low cost và PIT/clone theo yêu cầu.

📋 Giải thích chi tiết TẤT CẢ các phương án (đúng/sai)

  • ✅ ĐÚNG: Use managed export, and store the data in a Cloud Storage bucket using Nearline or Coldline class.
    Phương án này hoàn hảo cho archive lâu dài với chi phí thấp. Managed export tạo files dữ liệu vào bucket Cloud Storage. Sử dụng Nearline (chi phí ~$0.001/GB/tháng) hoặc Coldline (còn rẻ hơn, ~$0.004/GB/tháng - cập nhật 2026) để lưu trữ snapshots định kỳ, hỗ trợ PIT recovery bằng cách import lại khi cần. Không cần xử lý thủ công, tự động và scalable cho data growth lớn.

  • ✅ ĐÚNG: Use managed export, and then import to Cloud Datastore in a separate project under a unique namespace reserved for that export.
    Đây là cách clone chính xác dữ liệu Datastore. Export files được import vào project riêng với namespace unique (hỗ trợ từ Firestore/Datastore mode), tạo snapshot độc lập. Cho phép PIT recovery hoặc test ở môi trường khác mà không ảnh hưởng dữ liệu gốc. Hỗ trợ multi-project isolation, lý tưởng cho long-term growth và compliance.

  • ❌ SAI: Use managed export, and then import the data into a BigQuery table created just for that export, and delete temporary export files.
    Phương án này chuyển dữ liệu sang BigQuery (warehouse analytics), không phải Datastore gốc. Import vào BigQuery mất cấu trúc entity/key của Datastore, không hỗ trợ clone Datastore hoặc PIT recovery native. Xóa files export làm mất khả năng archive lâu dài, tăng chi phí và phức tạp không cần thiết.

  • ❌ SAI: Write an application that uses Cloud Datastore client libraries to read all the entities. Treat each entity as a BigQuery table row via BigQuery streaming insert. Assign an export timestamp for each export, and attach it as an extra column for each row. Make sure that the BigQuery table is partitioned using the export timestamp column.
    Cách này thủ công và kém hiệu quả: Sử dụng client libraries để read/export streaming vào BigQuery, thêm timestamp cho partition. Không dùng managed export, dễ lỗi với dữ liệu real-time lớn (telemetry), chi phí streaming cao (~$0.01/200MB), không tạo snapshot Datastore thực thụ (chỉ analytics data), không hỗ trợ clone PIT cho Datastore.

  • ❌ SAI: Write an application that uses Cloud Datastore client libraries to read all the entities. Format the exported data into a JSON file. Apply compression before storing the data in Cloud Source Repositories.
    Không phù hợp lưu trữ: Cloud Source Repositories (CSR) dành cho source code/version control (như Git), không phải archival data lớn. Thủ công read/format/compress JSON tốn kém, không scalable cho data growth, thiếu PIT/clone native (CSR không hỗ trợ import Datastore), chi phí cao và không low-cost so với Storage classes.

📘 Tài liệu tham khảo (cập nhật mới nhất 2026):

Câu 223 Chọn nhiều đáp án
You need to create a data pipeline that copies time-series transaction data so that it can be queried from within BigQuery by your data science team for analysis.
Every hour, thousands of transactions are updated with a new status. The size of the initial dataset is 1.5 PB, and it will grow by 3 TB per day. The data is heavily structured, and your data science team will build machine learning models based on this data. You want to maximize performance and usability for your data science team. Which two strategies should you adopt? (Choose two.)
  1. A Denormalize the data as must as possible.
  2. B Preserve the structure of the data as much as possible.
  3. C Use BigQuery UPDATE to further reduce the size of the dataset.
  4. D Develop a data pipeline where status updates are appended to BigQuery instead of updated.
  5. E Copy a daily snapshot of transaction data to Cloud Storage and store it as an Avro file. Use BigQuery's support for external data sources to query.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi yêu cầu xây dựng một data pipeline để sao chép dữ liệu giao dịch time-series (dữ liệu chuỗi thời gian) vào BigQuery, nhằm cho phép đội ngũ data science truy vấn và phân tích, đặc biệt là xây dựng mô hình machine learning (ML). Các đặc điểm chính của dữ liệu:

  • Quy mô lớn: Dataset ban đầu 1.5 PB, tăng 3 TB mỗi ngày.
  • Cập nhật thường xuyên: Hàng giờ, hàng nghìn giao dịch được cập nhật status mới.
  • Cấu trúc: Dữ liệu heavily structured (có cấu trúc cao).
  • Mục tiêu: Tối ưu hóa performance và usability cho đội data science (tốc độ truy vấn nhanh, dễ sử dụng cho ML).

Câu hỏi là multiple choice chọn hai chiến lược đúng (Choose two), tập trung vào best practices của BigQuery trên Google Cloud Platform (GCP) để xử lý dữ liệu lớn, time-series với updates liên tục. 📈

✅ Đáp án đúng và lý do lựa chọn

Hai chiến lược đúng là:

  1. Denormalize the data as much as possible.
  2. Develop a data pipeline where status updates are appended to BigQuery instead of updated.

Lý do lựa chọn (dựa trên best practices BigQuery mới nhất đến 2026):

  • BigQuery là columnar storage (lưu trữ cột), tối ưu cho analytical queries và ML workloads. Với dataset PB-scale và tăng trưởng hàng TB/ngày, cần tránh joins phức tạp (do normalized data gây ra) để giảm latency truy vấn xuống mức sub-second.
  • Updates status hàng giờ trên hàng nghìn records sẽ đắt đỏ nếu dùng UPDATE/MERGE trực tiếp (BigQuery rewrite toàn bộ partition/micro-partition). Thay vào đó, append-only (immutable log) phù hợp time-series, dễ scale với streaming inserts hoặc batch loads, hỗ trợ time-travel queries và ML feature stores.
  • Kết hợp denormalization + append giúp maximize usability: Dữ liệu phẳng, dễ query cho TensorFlow/PyTorch trên Vertex AI, giảm chi phí compute ~50-70% so với normalized + updates. 🛠️

📋 Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng phương án, giữ nguyên văn bản gốc tiếng Anh. Tôi đánh dấu ✅ cho đúng, ❌ cho sai, kèm giải thích đầy đủ bằng tiếng Việt dựa trên tài liệu GCP cập nhật 2026 (BigQuery v2.0+ với enhancements cho streaming và partitioning).

  • ✅ Denormalize the data as much as possible.
    Đúng vì: Trong BigQuery, dữ liệu analytical (như time-series cho ML) nên denormalize để flatten schema, tránh JOINs tốn kém (có thể scan hàng PB dữ liệu). Với 1.5 PB + 3 TB/ngày, denormalization tăng query speed 10-100x, dễ build features cho ML (e.g., Vertex AI Feature Store). Best practice từ GCP: "Denormalize for analytics" giúp columnar compression tốt hơn. 🏆

  • ❌ Preserve the structure of the data as much as possible.
    Sai vì: Giữ normalized structure (nhiều tables liên kết) lý tưởng cho OLTP (transactional DB như Cloud SQL), nhưng kém cho BigQuery OLAP. Sẽ yêu cầu nhiều JOINs, tăng scan volume và latency (có thể >1 phút/query trên PB data), không tối ưu usability cho data science. GCP khuyến nghị ngược lại cho analytics workloads. 🚫

  • ❌ Use BigQuery UPDATE to further reduce the size of the dataset.
    Sai vì: BigQuery DML UPDATE/MERGE rewrite toàn bộ affected partitions (mutation v2 từ 2023), rất chậm và đắt với updates hàng giờ trên 1.5 PB (có thể tốn hàng giờ + chi phí cao). Không "reduce size" mà còn tăng storage do clustering overhead. Không phù hợp time-series; dùng append thay thế để scale. 💸

  • ✅ Develop a data pipeline where status updates are appended to BigQuery instead of updated.
    Đúng vì: Append-only (immutable) là best practice cho time-series/CDC (change data capture) trong BigQuery. Sử dụng Dataflow hoặc Pub/Sub + Streaming Inserts để append status updates như new rows (với timestamp), tránh mutation costs. Hỗ trợ query time-series dễ dàng (e.g., WINDOW functions), scale đến 3 TB/ngày mà perf ổn định. Tích hợp tốt với ML pipelines trên Vertex AI. ⚡

  • ❌ Copy a daily snapshot of transaction data to Cloud Storage and store it as an Avro file. Use BigQuery's support for external data sources to query.
    Sai vì: External tables (Avro trên GCS) chậm hơn native tables 5-10x do scan qua GCS mỗi query, không tối ưu perf cho frequent queries/ML training trên PB data. Daily snapshot bỏ lỡ hourly updates, không "maximize usability". BigQuery khuyến nghị load native cho high-perf workloads thay vì external. 🐌

📘 Tài liệu tham khảo (cập nhật 2026)

Hy vọng phân tích này giúp bạn ôn thi chứng chỉ Google Cloud Professional Data Engineer hiệu quả! 🚀 Nếu cần thêm ví dụ code Dataflow, hãy hỏi nhé.

Câu 224
You are designing a cloud-native historical data processing system to meet the following conditions:
✑ The data being analyzed is in CSV, Avro, and PDF formats and will be accessed by multiple analysis tools including Dataproc, BigQuery, and Compute
Engine.
✑ A batch pipeline moves daily data.
✑ Performance is not a factor in the solution.
✑ The solution design should maximize availability.
How should you design data storage for this solution?
  1. A Create a Dataproc cluster with high availability. Store the data in HDFS, and perform analysis as needed.
  2. B Store the data in BigQuery. Access the data using the BigQuery Connector on Dataproc and Compute Engine.
  3. C Store the data in a regional Cloud Storage bucket. Access the bucket directly using Dataproc, BigQuery, and Compute Engine.
  4. D Store the data in a multi-regional Cloud Storage bucket. Access the data directly using Dataproc, BigQuery, and Compute Engine.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi yêu cầu thiết kế hệ thống lưu trữ dữ liệu cloud-native cho việc xử lý dữ liệu lịch sử (historical data), với các điều kiện cụ thể sau:

  • 📊 Dữ liệu đầu vào ở định dạng CSV, Avro, và PDF, được truy cập bởi nhiều công cụ phân tích như Dataproc (xử lý batch Spark/Hadoop), BigQuery (phân tích SQL), và Compute Engine (VM linh hoạt).
  • 🔄 Có pipeline batch di chuyển dữ liệu hàng ngày (daily batch).
  • ⚡ Hiệu suất không phải yếu tố quan trọng (performance is not a factor).
  • 🎯 Thiết kế phải tối ưu hóa tính sẵn sàng cao nhất (maximize availability).

Mục tiêu là chọn giải pháp lưu trữ dữ liệu sao cho đơn giản, cloud-native, hỗ trợ đa định dạng, dễ truy cập từ các dịch vụ GCP, và đảm bảo availability cao nhất mà không cần lo về performance. Cloud Storage là lựa chọn lý tưởng vì tính linh hoạt, nhưng cần chọn loại bucket phù hợp để max availability.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Store the data in a multi-regional Cloud Storage bucket. Access the data directly using Dataproc, BigQuery, and Compute Engine.

Lý do:

  • 🛡️ Multi-regional Cloud Storage bucket cung cấp tính sẵn sàng (availability) và độ bền (durability) cao nhất trong GCP (99.95% monthly uptime SLA, dữ liệu replicate qua nhiều vùng địa lý). Điều này hoàn hảo để maximize availability như yêu cầu.
  • 📂 Hỗ trợ đầy đủ CSV, Avro, PDF; tất cả công cụ (Dataproc, BigQuery, Compute Engine) truy cập trực tiếp qua connector/storage API mà không cần trung gian.
  • ☁️ Cloud-native thuần túy: Không phụ thuộc cluster, dễ scale, chi phí thấp cho historical data batch daily.
  • Theo tài liệu GCP mới nhất (2024-2026), multi-regional là lựa chọn top cho high-availability storage (xem Cloud Storage Classes).

📋 Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh, kèm giải thích đúng/sai bằng tiếng Việt:

  • ❌ [SAI] Create a Dataproc cluster with high availability. Store the data in HDFS, and perform analysis as needed.
    Giải thích sai: HDFS (Hadoop Distributed File System) không phải cloud-native, phụ thuộc vào Dataproc cluster (dù HA), dẫn đến availability thấp hơn vì cluster có thể downtime khi scale hoặc maintenance. Không hỗ trợ trực tiếp BigQuery/Compute Engine tốt, và PDF/Avro cần xử lý phức tạp. Vi phạm "maximize availability" vì HDFS chỉ regional, không replicate multi-region tự động. Không phù hợp historical data lớn.

  • ❌ [SAI] Store the data in BigQuery. Access the data using the BigQuery Connector on Dataproc and Compute Engine.
    Giải thích sai: BigQuery là data warehouse tối ưu cho structured/semi-structured data (CSV/Avro), nhưng không hỗ trợ PDF native (cần extract trước, phức tạp). Không phải lưu trữ raw historical data; chi phí cao cho batch daily lớn, và Compute Engine cần connector gián tiếp. Availability cao (99.99%), nhưng không max so với multi-regional CS và không cloud-native cho raw multi-format storage.

  • ❌ [SAI] Store the data in a regional Cloud Storage bucket. Access the bucket directly using Dataproc, BigQuery, and Compute Engine.
    Giải thích sai: Regional CS hỗ trợ tốt tất cả định dạng, truy cập trực tiếp, cloud-native, nhưng availability chỉ 99.9% SLA, thấp hơn multi-regional (dữ liệu chỉ trong 1 region). Không đạt "maximize availability" vì rủi ro outage regional (ví dụ disaster). Phù hợp performance cao/chi phí thấp, nhưng ưu tiên availability ở đây.

  • ✅ [ĐÚNG] Store the data in a multi-regional Cloud Storage bucket. Access the data directly using Dataproc, BigQuery, and Compute Engine.
    Giải thích đúng: Như đã nêu ở phần đáp án, max availability với multi-region replication, hỗ trợ đầy đủ định dạng/raw data, truy cập trực tiếp từ tất cả tool, phù hợp batch daily mà không lo performance.

📘 Tài liệu tham khảo

Giải pháp này đảm bảo đơn giản, scalable, và reliable cho historical data processing! 🚀

Câu 225
You have a petabyte of analytics data and need to design a storage and processing platform for it. You must be able to perform data warehouse-style analytics on the data in Google Cloud and expose the dataset as files for batch analysis tools in other cloud providers. What should you do?
  1. A Store and process the entire dataset in BigQuery.
  2. B Store and process the entire dataset in Bigtable.
  3. C Store the full dataset in BigQuery, and store a compressed copy of the data in a Cloud Storage bucket.
  4. D Store the warm data as files in Cloud Storage, and store the active data in BigQuery. Keep this ratio as 80% warm and 20% active.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này xoay quanh việc thiết kế một nền tảng lưu trữ và xử lý dữ liệu quy mô petabyte (1 PB = 1.000 TB) dành cho dữ liệu phân tích (analytics data). Các yêu cầu chính bao gồm:

  • ✅ Thực hiện phân tích kiểu data warehouse (kho dữ liệu) trên Google Cloud.
  • 📤 Tiếp xúc bộ dữ liệu dưới dạng tệp tin (files) để các công cụ phân tích batch (xử lý hàng loạt) ở các nhà cung cấp cloud khác có thể sử dụng.

🛠️ Bối cảnh kỹ thuật: Dữ liệu lớn cần giải pháp scalable, hỗ trợ query SQL phức tạp cho analytics (như BigQuery), đồng thời dễ dàng export/share dưới dạng file (như Parquet, Avro) qua Cloud Storage để tương thích đa cloud. Kiến thức dựa trên phiên bản Google Cloud mới nhất (2024-2026), với BigQuery hỗ trợ storage lên đến exabyte và federation queries.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Store the full dataset in BigQuery, and store a compressed copy of the data in a Cloud Storage bucket.

Lý do chi tiết:

  • 🏪 BigQuery là dịch vụ data warehouse serverless lý tưởng cho petabyte-scale analytics, hỗ trợ SQL queries nhanh chóng, columnar storage, và ML integration (cập nhật 2026: BigQuery Omni cho multi-cloud queries). Lưu toàn bộ dataset ở đây đáp ứng yêu cầu "data warehouse-style analytics on Google Cloud".
  • 💾 Cloud Storage bucket lưu bản sao nén (compressed copy) cho phép expose dữ liệu dưới dạng files (e.g., Parquet/CSV), dễ dàng chia sẻ với batch tools ở AWS/GCP/Azure qua public URLs hoặc transfer services. Compression tiết kiệm chi phí (lên đến 90% với Snappy/Zstandard).
  • 🔄 Kết hợp này đảm bảo dual-purpose: Analytics nội bộ + portability đa cloud, không vi phạm yêu cầu.

📘 Tài liệu tham khảo:

📋 Giải thích tất cả các phương án (đúng/sai)

  • Phương án 1: Store and process the entire dataset in BigQuery.
    ❌ Sai vì: BigQuery xuất sắc cho analytics warehouse nhưng không expose trực tiếp dưới dạng files cho batch tools ở cloud khác. Bạn chỉ có thể export ra GCS (thêm bước phức tạp), không đáp ứng yêu cầu "expose the dataset as files". Không scalable cho pure file access mà không qua query engine.

  • Phương án 2: Store and process the entire dataset in Bigtable.
    ❌ Sai vì: Bigtable là NoSQL database cho real-time/low-latency workloads (e.g., IoT, time-series), không phù hợp data warehouse analytics (thiếu SQL, columnar scans kém hiệu quả cho petabyte ad-hoc queries). Không hỗ trợ expose files dễ dàng cho multi-cloud batch.

  • Phương án 3: Store the full dataset in BigQuery, and store a compressed copy of the data in a Cloud Storage bucket.
    ✅ Đúng vì: Như giải thích ở trên – BigQuery cho analytics GCP, GCS cho file export nén, hoàn hảo cho petabyte-scale, multi-cloud compatibility. Hỗ trợ autoscaling và cost-optimized (BigQuery storage ~$0.02/GB/tháng, GCS rẻ hơn).

  • Phương án 4: Store the warm data as files in Cloud Storage, and store the active data in BigQuery. Keep this ratio as 80% warm and 20% active.
    ❌ Sai vì: Phân chia "warm/active" (80/20) là heuristic tùy ý, không được câu hỏi yêu cầu và phức tạp hóa quản lý (cần ETL pipeline tự động). Không lưu "full dataset" ở BigQuery, chỉ phần active → analytics không đầy đủ. GCS files kém cho warehouse queries so với BigQuery native.

🧠 Kết luận: Giải pháp đúng tận dụng best-of-breed GCP services (BigQuery + GCS) cho hybrid analytics + portability, phù hợp kiến trúc data lakehouse hiện đại (2026 trends).

Câu 226
You work for a manufacturing company that sources up to 750 different components, each from a different supplier. You've collected a labeled dataset that has on average 1000 examples for each unique component. Your team wants to implement an app to help warehouse workers recognize incoming components based on a photo of the component. You want to implement the first working version of this app (as Proof-Of-Concept) within a few working days. What should you do?
  1. A Use Cloud Vision AutoML with the existing dataset.
  2. B Use Cloud Vision AutoML, but reduce your dataset twice.
  3. C Use Cloud Vision API by providing custom labels as recognition hints.
  4. D Train your own image recognition model leveraging transfer learning techniques.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong công ty sản xuất: công ty nhập khẩu lên đến 750 linh kiện khác nhau, mỗi linh kiện từ một nhà cung cấp riêng biệt. Họ đã có bộ dữ liệu đã gắn nhãn (labeled dataset) với trung bình 1000 ví dụ ảnh cho mỗi linh kiện duy nhất, tổng cộng khoảng 750.000 ảnh. Nhiệm vụ là xây dựng ứng dụng (app) giúp nhân viên kho nhận diện linh kiện dựa trên ảnh chụp, và yêu cầu triển khai phiên bản Proof-of-Concept (POC) đầu tiên chỉ trong vài ngày làm việc.

Mục tiêu chính là tốc độ triển khai nhanh chóng (few working days), tận dụng dataset sẵn có, mà không cần xây dựng mô hình phức tạp từ đầu. Đây là bài toán phân loại hình ảnh (image classification) với nhiều lớp (multi-class, 750 classes), phù hợp với các công cụ no-code/low-code của Google Cloud để train nhanh. 📱🛠️

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Cloud Vision AutoML with the existing dataset.

Lý do chi tiết:

  • Cloud Vision AutoML (nay là phần của Vertex AI Vision - cập nhật đến 2026) được thiết kế chính xác cho các bài toán như thế này: train mô hình tùy chỉnh từ dataset ảnh đã gắn nhãn mà không cần code phức tạp.
  • Với 1000 ảnh/lớp (rất lý tưởng, vượt yêu cầu tối thiểu 100-500 ảnh/lớp của AutoML), bạn có thể upload dataset trực tiếp, train trong vài giờ, rồi deploy thành endpoint API cho app ngay lập tức – hoàn hảo cho POC trong vài ngày.
  • Quy trình: Upload CSV/JSON với ảnh → Train → Evaluate → Deploy. Không cần GPU riêng hay kiến thức deep learning sâu. 🚀
  • Nguồn tham khảo: Vertex AI Vision Documentation (Google Cloud, 2026) và AutoML Vision Best Practices.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt với lý do đúng/sai dựa trên tính khả thi, tốc độ và best practices của Google Cloud (cập nhật Vertex AI 2026). 🧠

  • ✅ Use Cloud Vision AutoML with the existing dataset.
    Đúng hoàn toàn! Phương án này tận dụng tối ưu dataset sẵn có (1000 ảnh/lớp), train nhanh chóng mà không chỉnh sửa dữ liệu. Thời gian POC chỉ 1-3 ngày: upload → train (2-8 giờ) → integrate vào app. Độ chính xác cao nhờ dữ liệu phong phú, phù hợp 750 lớp. Đây là giải pháp no-code nhanh nhất cho custom image classification. 🌟

  • ❌ Use Cloud Vision AutoML, but reduce your dataset twice.
    Sai! Giảm dataset xuống còn 500 ảnh/lớp (reduce twice) là không cần thiết và có hại. Dataset gốc đã lý tưởng (1000 ảnh > khuyến nghị 500+ của AutoML), giảm sẽ làm giảm độ chính xác mô hình, tăng overfitting/underfitting. Không có lý do gì để "giảm" vì AutoML xử lý tốt dữ liệu lớn (hàng triệu ảnh). Lãng phí thời gian chỉnh sửa dataset vô ích! 😞

  • ❌ Use Cloud Vision API by providing custom labels as recognition hints.
    Sai! Cloud Vision API là dịch vụ pre-trained general-purpose (như label detection, object localization), không hỗ trợ custom labels từ dataset của bạn. Bạn chỉ có thể dùng hints cơ bản (như context), nhưng nó không train được trên 750 linh kiện custom, dẫn đến độ chính xác thấp (<50% cho classes mới). Không phù hợp POC custom recognition. 🚫

  • ❌ Train your own image recognition model leveraging transfer learning techniques.
    Sai! Xây dựng mô hình từ đầu với transfer learning (ví dụ dùng TensorFlow/PyTorch trên pre-trained ResNet) đòi hỏi code phức tạp, setup môi trường (Vertex AI Workbench/TPU), tune hyperparameters, và train/test nhiều lần – mất tuần/tháng, không phải "vài ngày". Với 750 classes, cần data augmentation, validation sets... quá phức tạp cho POC. Chọn AutoML để nhanh hơn! ⏳

Kết luận khuyến nghị: Chọn AutoML để POC nhanh, sau đó scale lên Vertex AI nếu cần tùy chỉnh nâng cao. Nếu implement, bắt đầu bằng console.vertex.ai để prototype ngay! 💡📘

Câu 227
You are working on a niche product in the image recognition domain. Your team has developed a model that is dominated by custom C++ TensorFlow ops your team has implemented. These ops are used inside your main training loop and are performing bulky matrix multiplications. It currently takes up to several days to train a model. You want to decrease this time significantly and keep the cost low by using an accelerator on Google Cloud. What should you do?
  1. A Use Cloud TPUs without any additional adjustment to your code.
  2. B Use Cloud TPUs after implementing GPU kernel support for your customs ops.
  3. C Use Cloud GPUs after implementing GPU kernel support for your customs ops.
  4. D Stay on CPUs, and increase the size of the cluster you're training your model on.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc tối ưu hóa quá trình huấn luyện mô hình nhận diện hình ảnh (image recognition) trên Google Cloud, nơi mô hình sử dụng các custom C++ TensorFlow ops (các toán tử tùy chỉnh viết bằng C++ trong TensorFlow) để thực hiện các phép nhân ma trận lớn (bulky matrix multiplications). Quá trình huấn luyện hiện tại mất vài ngày, và mục tiêu là giảm thời gian đáng kể đồng thời giữ chi phí thấp bằng cách sử dụng accelerator (bộ tăng tốc như TPU hoặc GPU).

🔍 Các yếu tố chính cần xem xét:

  • Custom ops là các toán tử tùy chỉnh, không phải ops chuẩn của TensorFlow, nên chúng cần được tối ưu hóa riêng cho phần cứng accelerator (ví dụ: viết kernel hỗ trợ GPU/TPU).
  • Cloud TPUs (Tensor Processing Units) trên Google Cloud rất mạnh cho huấn luyện TensorFlow, nhưng yêu cầu code phải tương thích với XLA compiler và kiến trúc systolic array của TPU – custom C++ ops thường khó port trực tiếp vì TPU không hỗ trợ CUDA hoặc các kernel GPU-like dễ dàng.
  • Cloud GPUs (như NVIDIA A100, H100 trên GCE hoặc Vertex AI) linh hoạt hơn, hỗ trợ CUDA kernels cho custom ops, giúp tăng tốc matrix multiplications hiệu quả.
  • Giải pháp phải giảm thời gian + chi phí thấp, tránh scale CPU (đắt và chậm).

Dựa trên kiến thức cập nhật đến 2026 (theo tài liệu Google Cloud và TensorFlow 2.15+ với JAX/XLA hỗ trợ TPU v5e/v5p), custom ops cần GPU kernel support để tận dụng GPUs tốt nhất.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Cloud GPUs after implementing GPU kernel support for your customs ops.

🛠️ Lý do chi tiết:

  • Custom C++ TensorFlow ops có thể được compile thành CUDA kernels (sử dụng tf.custom_op với NVCC hoặc TensorFlow's op registration), giúp chạy song song trên Cloud GPUs (như A100/H100 trên Google Compute Engine hoặc Vertex AI Training).
  • GPUs tăng tốc matrix multiplications lên hàng trăm lần so với CPU nhờ cuBLAS/cuDNN, giảm thời gian từ ngày xuống giờ, và chi phí thấp hơn so với scale cluster CPU lớn (GPUs rẻ hơn per FLOPS cho workload này).
  • Không cần thay đổi lớn code TensorFlow (chỉ implement GPU kernel), dễ dàng với tf.config và multi-GPU distribution strategy.
  • Theo benchmark 2024-2026, GPUs vượt trội cho custom ops so với TPUs (TPU yêu cầu rewrite ops bằng HLO/XLA, phức tạp hơn).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên tính khả thi, hiệu suất và chi phí trên Google Cloud (cập nhật TPU v5p, GPU L4/H100).

  • ❌ Use Cloud TPUs without any additional adjustment to your code.
    Sai vì: Cloud TPUs không hỗ trợ custom C++ ops trực tiếp mà không chỉnh sửa. TPUs yêu cầu toàn bộ graph compile qua XLA (Accelerated Linear Algebra), và custom ops C++ (dựa trên CPU kernel) sẽ fallback về CPU hoặc lỗi, không tận dụng systolic array của TPU. Kết quả: Không giảm thời gian, thậm chí chậm hơn + tốn phí TPU idle.

  • ❌ Use Cloud TPUs after implementing GPU kernel support for your customs ops.
    Sai vì: GPU kernel (CUDA) không tương thích với TPU. TPUs dùng kiến trúc riêng (không phải NVIDIA GPU), cần TPU-specific kernels hoặc rewrite ops bằng XLA HLO/JAX. Implement GPU kernel chỉ hữu ích cho GPUs, không giúp TPU – dẫn đến lỗi runtime và lãng phí công sức.

  • ✅ Use Cloud GPUs after implementing GPU kernel support for your customs ops.
    Đúng vì: Như giải thích ở trên, đây là giải pháp tối ưu. Custom ops C++ dễ port sang CUDA kernel (qua tensorflow/cc hoặc tf.custom_ops), tận dụng cuBLAS cho matrix multiplications trên Cloud GPUs. Giảm thời gian huấn luyện >10x, chi phí thấp (khoảng 1-3$/giờ cho A100), hỗ trợ đầy đủ trong Vertex AI Pipelines hoặc GKE.

  • ❌ Stay on CPUs, and increase the size of the cluster you're training your model on.
    Sai vì: CPU (như N2/C4 instances) không hiệu quả cho matrix multiplications, scale cluster lớn chỉ tăng chi phí tuyến tính (theo số node) mà speedup hạn chế (do Amdahl's law và communication overhead). Thời gian vẫn vài ngày, chi phí cao gấp 5-10x so với GPUs (CPU ~0.1$/giờ nhưng cần hàng trăm node).

📘 Tài liệu tham khảo

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần ví dụ code, hãy hỏi thêm.

Câu 228
You work on a regression problem in a natural language processing domain, and you have 100M labeled examples in your dataset. You have randomly shuffled your data and split your dataset into train and test samples (in a 90/10 ratio). After you trained the neural network and evaluated your model on a test set, you discover that the root-mean-squared error (RMSE) of your model is twice as high on the train set as on the test set. How should you improve the performance of your model?
  1. A Increase the share of the test sample in the train-test split.
  2. B Try to collect more data and increase the size of your dataset.
  3. C Try out regularization techniques (e.g., dropout of batch normalization) to avoid overfitting.
  4. D Increase the complexity of your model by, e.g., introducing an additional layer or increase sizing the size of vocabularies or n-grams used.
Xem giải thích

🧩 Giải thích nội dung câu hỏi
Câu hỏi mô tả một vấn đề hồi quy (regression) trong lĩnh vực xử lý ngôn ngữ tự nhiên (NLP). Bạn có bộ dữ liệu lớn với 100 triệu ví dụ có nhãn, đã xáo trộn ngẫu nhiên và chia thành tập huấn luyện (train set - 90%) và tập kiểm tra (test set - 10%). Sau khi huấn luyện mạng nơ-ron (neural network) và đánh giá trên test set, bạn phát hiện lỗi RMSE (Root Mean Squared Error) trên tập train cao gấp đôi so với trên tập test (RMSE_train = 2 × RMSE_test).
🔍 Ý nghĩa vấn đề: Đây là dấu hiệu của underfitting (mô hình chưa học tốt dữ liệu huấn luyện). Thông thường, trong học máy, lỗi trên train set thấp hơn hoặc bằng lỗi trên test set. Việc RMSE train cao hơn test set cho thấy mô hình quá đơn giản, chưa đủ phức tạp để nắm bắt pattern trong dữ liệu train, dẫn đến hiệu suất kém trên chính dữ liệu huấn luyện. Không phải overfitting (vì overfitting sẽ có RMSE train thấp, test cao). Với dataset lớn (100M samples), vấn đề không phải thiếu data mà là mô hình cần phức tạp hơn.

✅ Đáp án đúng:
Increase the complexity of your model by, e.g., introducing an additional layer or increase sizing the size of vocabularies or n-grams used.
Lý do chọn: Mô hình đang underfit vì chưa fit tốt train data (RMSE train cao bất thường). Giải pháp là tăng độ phức tạp như thêm layer, tăng kích thước từ vựng (vocabulary) hoặc n-grams để mô hình học sâu hơn các pattern NLP phức tạp. Điều này phù hợp với dataset lớn, giúp cải thiện generalization mà không lo overfitting ngay lập tức. Theo best practices ML đến 2026, trong AWS SageMaker hoặc các framework như TensorFlow/PyTorch, underfitting được khắc phục bằng cách scale up model architecture trước khi thêm data.

🛠️ Phân tích tất cả các phương án (Giữ nguyên text gốc bằng tiếng Anh):

  • Increase the share of the test sample in the train-test split.
    ❌ Sai: Tăng tỷ lệ test set (ví dụ 20/80) sẽ làm train set nhỏ hơn, dẫn đến underfitting nặng hơn vì ít data huấn luyện. Vấn đề hiện tại là underfitting trên train lớn (90%), không phải do test set quá nhỏ. Split 90/10 đã hợp lý cho dataset 100M.

  • Try to collect more data and increase the size of your dataset.
    ❌ Sai: Dataset đã rất lớn (100M labeled examples - hiếm gặp trong NLP), thêm data sẽ không giải quyết underfitting mà chỉ tốn kém. Underfitting do model quá đơn giản, không phải thiếu data. AWS khuyến nghị kiểm tra model capacity trước khi scale data (SageMaker Autopilot docs 2026).

  • Try out regularization techniques (e.g., dropout of batch normalization) to avoid overfitting.
    ❌ Sai: Regularization (dropout, batch norm) dùng để chống overfitting (train tốt, test kém), nhưng ở đây ngược lại (train kém hơn test). Áp dụng sẽ làm underfitting tệ hơn bằng cách hạn chế model học.

  • Increase the complexity of your model by, e.g., introducing an additional layer or increase sizing the size of vocabularies or n-grams used.
    ✅ Đúng: Như đã giải thích ở trên, đây là giải pháp trực tiếp cho underfitting. Tăng layer/vocab/n-grams giúp model capture complexity NLP tốt hơn, cải thiện RMSE train mà vẫn giữ test thấp.

📘 Tài liệu tham khảo (cập nhật đến 2026):

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần ví dụ code AWS SageMaker, hãy hỏi thêm.

Câu 229
You use BigQuery as your centralized analytics platform. New data is loaded every day, and an ETL pipeline modifies the original data and prepares it for the final users. This ETL pipeline is regularly modified and can generate errors, but sometimes the errors are detected only after 2 weeks. You need to provide a method to recover from these errors, and your backups should be optimized for storage costs. How should you organize your data in BigQuery and store your backups?
  1. A Organize your data in a single table, export, and compress and store the BigQuery data in Cloud Storage.
  2. B Organize your data in separate tables for each month, and export, compress, and store the data in Cloud Storage.
  3. C Organize your data in separate tables for each month, and duplicate your data on a separate dataset in BigQuery.
  4. D Organize your data in separate tables for each month, and use snapshot decorators to restore the table to a time prior to the corruption.
Xem giải thích

🧩 Giải thích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh việc quản lý dữ liệu trong BigQuery – nền tảng phân tích dữ liệu tập trung của Google Cloud. Tình huống cụ thể:

  • Dữ liệu mới được tải lên hàng ngày.
  • Một ETL pipeline thường xuyên sửa đổi dữ liệu gốc để chuẩn bị cho người dùng cuối, nhưng pipeline này hay bị thay đổi và có thể gây lỗi.
  • Lỗi đôi khi chỉ được phát hiện sau 2 tuần, dẫn đến dữ liệu bị hỏng (corruption).
  • Yêu cầu: Cung cấp phương pháp khôi phục (recover) từ lỗi, đồng thời tối ưu chi phí lưu trữ cho backups.

Mục tiêu là tổ chức dữ liệu trong BigQuery sao cho dễ dàng rollback về trạng thái trước lỗi (sau 2 tuần), và lưu backups rẻ tiền. BigQuery có Time Travel chỉ hỗ trợ 7 ngày gần nhất (không đủ cho 2 tuần), nên cần chiến lược khác như partitioning theo thời gian và export ra Cloud Storage (CS) để lưu trữ dài hạn với chi phí thấp (compressed). 📈

✅ Đáp án đúng

Organize your data in separate tables for each month, and export, compress, and store the data in Cloud Storage.

Lý do lựa chọn:

  • Việc chia dữ liệu thành các bảng riêng biệt theo từng tháng (ví dụ: table_2024_01, table_2024_02) giúp dễ dàng isolate dữ liệu theo thời gian, tránh ảnh hưởng toàn bộ dataset khi lỗi xảy ra.
  • Export, nén (compress) và lưu vào Cloud Storage là cách tối ưu chi phí nhất: CS hỗ trợ lưu trữ rẻ (Standard/Nearline/Archive classes), nén giảm dung lượng (gzip/bzip2), và có thể restore nhanh bằng cách load lại vào BigQuery khi cần.
  • Phù hợp với lỗi phát hiện sau 2 tuần, vì có thể export backup hàng ngày/tuần và giữ lịch sử tháng. Không tốn kém như duplicate full trong BigQuery (active storage đắt hơn CS nhiều lần). 🛡️ Cập nhật 2026: BigQuery vẫn khuyến nghị pattern này cho long-term recovery, kết hợp với table partitioning/clustering để export hiệu quả (theo docs BigQuery best practices).

📋 Phân tích tất cả các phương án

  • ❌ Organize your data in a single table, export, and compress and store the BigQuery data in Cloud Storage.
    Sai vì: Một bảng duy nhất (single table) khó khăn khi export chỉ phần dữ liệu trước lỗi (sau 2 tuần), vì phải query toàn bộ bảng lớn, tốn thời gian và chi phí query cao. Không dễ isolate theo thời gian, dẫn đến recovery chậm và không tối ưu. Phân chia theo tháng mới linh hoạt hơn.

  • ✅ Organize your data in separate tables for each month, and export, compress, and store the data in Cloud Storage.
    Đúng vì: Như đã giải thích ở trên – chia bảng theo tháng giúp point-in-time recovery chính xác (ví dụ: khôi phục chỉ bảng tháng trước), export selective tiết kiệm chi phí, và CS là lựa chọn lưu trữ rẻ nhất cho backups dài hạn (chi phí ~0.023$/GB/tháng cho Nearline). Hoàn hảo cho ETL lỗi muộn.

  • ❌ Organize your data in separate tables for each month, and duplicate your data on a separate dataset in BigQuery.
    Sai vì: Duplicate toàn bộ dữ liệu sang dataset riêng trong BigQuery tốn kém cực cao (active storage ~0.02$/GB/tháng + query fees), không tối ưu chi phí so với CS. Backup sẽ "nóng" và đắt đỏ, không phù hợp yêu cầu "optimized for storage costs". Dễ scale nhưng lãng phí cho lịch sử cũ.

  • ❌ Organize your data in separate tables for each month, and use snapshot decorators to restore the table to a time prior to the corruption.
    Sai vì: BigQuery không có tính năng "snapshot decorators" (có thể nhầm với table snapshots hoặc Time Travel). Table snapshots (từ 2023) chỉ hỗ trợ point-in-time copy trong BigQuery (không phải decorator), và Time Travel chỉ 7 ngày – không đủ cho lỗi sau 2 tuần. Không tối ưu chi phí (vẫn lưu active storage), và không linh hoạt như export ra CS. 🛑

📘 Tài liệu tham khảo

  • BigQuery Documentation: Best practices for data retention and recovery (cập nhật 2025: Nhấn mạnh partitioning + export to CS).
  • BigQuery Time Travel & Snapshots: Query historical data (7 ngày limit) & Table snapshots (point-in-time, nhưng không thay thế long-term backup).
  • Cloud Storage Pricing: Pricing calculator – Xác nhận CS rẻ hơn BigQuery storage 5-10x cho backups.
  • Google Cloud Skills Boost: Exam guide for Professional Data Engineer (Q&A tương tự về ETL recovery).

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần ví dụ code export, hãy hỏi thêm nhé! 😊

Câu 230
The marketing team at your organization provides regular updates of a segment of your customer dataset. The marketing team has given you a CSV with 1 million records that must be updated in BigQuery. When you use the UPDATE statement in BigQuery, you receive a quotaExceeded error. What should you do?
  1. A Reduce the number of records updated each day to stay within the BigQuery UPDATE DML statement limit.
  2. B Increase the BigQuery UPDATE DML statement limit in the Quota management section of the Google Cloud Platform Console.
  3. C Split the source CSV file into smaller CSV files in Cloud Storage to reduce the number of BigQuery UPDATE DML statements per BigQuery job.
  4. D Import the new records from the CSV file into a new BigQuery table. Create a BigQuery job that merges the new records with the existing records and writes the results to a new BigQuery table.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống thực tế trong Google Cloud BigQuery: Nhóm marketing cung cấp cập nhật định kỳ cho một phần tập dữ liệu khách hàng. Họ gửi file CSV chứa 1 triệu records cần cập nhật vào bảng BigQuery hiện có. Khi sử dụng lệnh UPDATE trực tiếp, bạn gặp lỗi quotaExceeded (vượt quota).

📌 Nguyên nhân lỗi chính: BigQuery áp dụng quota nghiêm ngặt cho DML statements (như UPDATE, DELETE, INSERT) trên mô hình on-demand pricing:

  • Giới hạn 1.500 unique table modifications per table per day (cập nhật đến 2024-2026, theo docs BigQuery).
  • Mỗi UPDATE lớn (như 1 triệu rows) có thể vượt slot usage hoặc query bytes processed, dẫn đến quotaExceeded ngay cả khi không chạm giới hạn rows trực tiếp.
  • Best practice cho dữ liệu lớn (>100k rows): Tránh UPDATE trực tiếp vì tốn kém và không scale; thay vào đó dùng staging table + MERGE để upsert hiệu quả.

🛠️ Mục tiêu: Tìm giải pháp scaleable, tránh quota, tối ưu chi phí cho cập nhật hàng loạt.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Import the new records from the CSV file into a new BigQuery table. Create a BigQuery job that merges the new records with the existing records and writes the results to a new BigQuery table.

Lý do:

  • ✅ MERGE statement là phương pháp chuẩn và được khuyến nghị cho large-scale updates/upserts trong BigQuery (từ phiên bản 2019, ổn định đến 2026).
  • Quy trình: Load CSV vào staging table (nhanh, không quota DML), rồi MERGE với bảng gốc → ghi kết quả vào bảng mới (atomic, tránh downtime).
  • Tránh quotaExceeded: MERGE chỉ tính 1 DML job/table/day, hỗ trợ hàng tỷ rows mà không vượt slot quota (dùng batch processing).
  • Tối ưu: Chi phí thấp hơn UPDATE trực tiếp (UPDATE scan toàn bảng, MERGE chỉ xử lý delta).

📘 Tài liệu tham khảo:

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh:

  • [SAI] Reduce the number of records updated each day to stay within the BigQuery UPDATE DML statement limit.
    ❌ Sai vì: Giảm records (ví dụ <100k/ngày) chỉ là workaround tạm thời, không scale cho 1 triệu records định kỳ. BigQuery quota là per table mod, không phải rows trực tiếp, nhưng UPDATE lớn vẫn tốn full table scan → quotaExceeded lặp lại. Không phù hợp business requirement (cập nhật full segment).

  • [SAI] Increase the BigQuery UPDATE DML statement limit in the Quota management section of the Google Cloud Platform Console.
    ❌ Sai vì: Quota DML (1.500 mods/day) không thể tăng dễ dàng qua Console (chỉ request quota cho slot/query, không dành cho DML). Ngay cả tăng, UPDATE vẫn không khuyến nghị cho large data (chi phí cao, latency lớn). Docs GCP khuyên tránh, ưu tiên MERGE/load replace.

  • [SAI] Split the source CSV file into smaller CSV files in Cloud Storage to reduce the number of BigQuery UPDATE DML statements per BigQuery job.
    ❌ Sai vì: Split file chỉ giúp LOAD nhanh hơn, nhưng UPDATE vẫn là 1 DML statement quét toàn bảng → quotaExceeded không giảm (vẫn 1 mod/table). Tạo nhiều job UPDATE nhỏ vẫn chạm 1.500 limit/day và tăng chi phí (nhiều full scans). Không giải quyết root cause.

  • [ĐÚNG] Import the new records from the CSV file into a new BigQuery table. Create a BigQuery job that merges the new records with the existing records and writes the results to a new BigQuery table.
    ✅ Đúng vì: Như giải thích trên – Staging + MERGE là pattern chuẩn (immutable data model của BigQuery). Load CSV dùng federated load (không DML quota), MERGE hỗ trợ conditional UPDATE/INSERT/DELETE atomic, output bảng mới → rename atomic swap. Scale đến PB data, quota-safe.

🧠 Lời khuyên từ Google Cloud Professional Data Engineer: Luôn dùng MERGE cho >10k rows updates. Test với bq command-line cho prod! 🚀