Ngân hàng đề — Google Cloud Associate Data Practitioner

Tìm thấy 333 câu.

Câu 291
You have a Dataflow pipeline that processes website traffic logs stored in Cloud Storage and writes the processed data to BigQuery. You noticed that the pipeline is failing intermittently. You need to troubleshoot the issue. What should you do?
  1. A Use Cloud Logging to identify error groups in the pipeline's logs. Use Cloud Monitoring to create a dashboard that tracks the number of errors in each group.
  2. B Use Cloud Logging to create a chart displaying the pipeline’s error logs. Use Metrics Explorer to validate the findings from the chart.
  3. C Use Cloud Logging to view error messages in the pipeline's logs. Use Cloud Monitoring to analyze the pipeline's metrics, such as CPU utilization and memory usage.
  4. D Use the Dataflow job monitoring interface to check the pipeline's status every hour. Use Cloud Profiler to analyze the pipeline’s metrics, such as CPU utilization and memory usage.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc khắc phục sự cố (troubleshooting) cho một pipeline Dataflow trên Google Cloud Platform (GCP). Cụ thể:

  • Pipeline này xử lý dữ liệu website traffic logs lưu trữ trong Cloud Storage, sau đó ghi dữ liệu đã xử lý vào BigQuery.
  • Vấn đề: Pipeline thất bại ngắt quãng (failing intermittently), nghĩa là không phải lúc nào cũng fail mà xảy ra sporadically.
  • Mục tiêu: Tìm cách troubleshoot hiệu quả để xác định nguyên nhân, sử dụng các công cụ monitoring của GCP như Cloud Logging, Cloud Monitoring, v.v.
  • Bối cảnh cập nhật 2026: Theo tài liệu GCP mới nhất (Dataflow 2.x và Apache Beam SDK 2.58+), troubleshooting Dataflow ưu tiên kiểm tra logs lỗi qua Cloud Logging và metrics hệ thống (như CPU, memory, throughput) qua Cloud Monitoring để phát hiện bottleneck hoặc resource issues. Không khuyến khích polling thủ công hoặc công cụ profiling phức tạp cho trường hợp intermittent failures.
    📘 Nguồn tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Cloud Logging to view error messages in the pipeline's logs. Use Cloud Monitoring to analyze the pipeline's metrics, such as CPU utilization and memory usage.

Lý do 🛠️:

  • Đây là cách chuẩn và hiệu quả nhất theo best practices của GCP cho troubleshooting Dataflow.
  • Cloud Logging cho phép xem trực tiếp error messages trong logs của pipeline (bao gồm worker logs, user logs), giúp xác định lỗi cụ thể như data parsing issues hoặc GCS read/write failures.
  • Cloud Monitoring cung cấp metrics real-time như CPU utilization, memory usage, vCPU time, để phân tích resource exhaustion – nguyên nhân phổ biến của intermittent failures (ví dụ: spike traffic gây OOM).
  • Kết hợp hai công cụ này cho visibility toàn diện mà không cần cấu hình phức tạp, phù hợp với failures không liên tục.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Use Cloud Logging to identify error groups in the pipeline's logs. Use Cloud Monitoring to create a dashboard that tracks the number of errors in each group.
    Giải thích: Phương án này quá phức tạp và không trực tiếp cho troubleshooting ban đầu. "Error groups" là tính năng nâng cao của Cloud Logging (dùng cho structured logging aggregation), nhưng Dataflow logs thường không được group sẵn thành error groups cho pipeline failures. Tạo dashboard tracking số lượng errors đòi hỏi setup trước, không phù hợp cho intermittent issues cần xem nhanh. Best practice là xem raw error messages trước, không phải group/dashboard ngay.

  • ❌ [SAI] Use Cloud Logging to create a chart displaying the pipeline’s error logs. Use Metrics Explorer to validate the findings from the chart.
    Giải thích: Không hiệu quả vì tạo chart từ logs (qua Logs Explorer) chỉ visualize lỗi, nhưng không phân tích sâu metrics hệ thống như CPU/memory – nguyên nhân cốt lõi của Dataflow failures. Metrics Explorer (trong Cloud Monitoring) dùng để validate metrics, không phải "validate findings from chart logs". Cách này thiếu focus vào resource metrics, dễ bỏ lỡ bottleneck.

  • ✅ [ĐÚNG] Use Cloud Logging to view error messages in the pipeline's logs. Use Cloud Monitoring to analyze the pipeline's metrics, such as CPU utilization and memory usage.
    Giải thích: Như đã nêu ở phần đáp án đúng. Đây là workflow tiêu chuẩn: Logs cho lỗi chi tiết + Metrics cho performance analysis. Hỗ trợ real-time troubleshooting, khớp với docs GCP (ví dụ: kiểm tra "system_lag" hoặc "worker_memory_usage" metrics).

  • ❌ [SAI] Use the Dataflow job monitoring interface to check the pipeline's status every hour. Use Cloud Profiler to analyze the pipeline’s metrics, such as CPU utilization and memory usage.
    Giải thích: Không phù hợp vì: (1) Dataflow job monitoring interface (trong GCP Console) chỉ check status cơ bản, nhưng "every hour" là polling thủ công chậm, không real-time cho intermittent failures. (2) Cloud Profiler dùng cho code-level profiling (CPU hotspots trong Beam code), không phải metrics hệ thống như CPU/memory tổng quát của workers. Profiler yêu cầu instrumentation code, không dành cho troubleshooting nhanh.

Kết luận 🎯: Chọn đáp án đúng giúp troubleshoot nhanh chóng, tiết kiệm thời gian và tài nguyên! Nếu cần thực hành, thử demo pipeline trên GCP Console.

Câu 292
Your organization’s business analysts require near real-time access to streaming data. However, they are reporting that their dashboard queries are loading slowly. After investigating BigQuery query performance, you discover the slow dashboard queries perform several joins and aggregations.
You need to improve the dashboard loading time and ensure that the dashboard data is as up-to-date as possible. What should you do?
  1. A Disable BigQuery query result caching.
  2. B Modify the schema to use parameterized data types.
  3. C Create a scheduled query to calculate and store intermediate results.
  4. D Create materialized views.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống thực tế trong môi trường Google Cloud BigQuery (không phải AWS như đề cập, có thể là nhầm lẫn): Tổ chức của bạn có các nhà phân tích kinh doanh cần truy cập dữ liệu streaming (dữ liệu dòng thời gian thực gần như real-time). Tuy nhiên, các truy vấn trên dashboard đang tải chậm do thực hiện nhiều phép nối (joins) và tổng hợp (aggregations).
📊 Vấn đề chính:

  • Dashboard cần tải nhanh hơn (cải thiện hiệu suất truy vấn).
  • Dữ liệu phải cập nhật nhất có thể (near real-time từ streaming data).
    🛠️ Mục tiêu: Tối ưu hóa thời gian tải dashboard mà vẫn giữ tính thời gian thực cao, dựa trên phân tích hiệu suất BigQuery.

✅ Đáp án đúng: "Create materialized views."

Lý do chọn đáp án này (dựa trên tính năng mới nhất của BigQuery đến năm 2026):
Materialized Views là các view được lưu trữ vật lý (pre-computed), tự động tính toán và lưu kết quả của các truy vấn phức tạp như joins và aggregations.
🧩 Lợi ích chính:

  • Cải thiện tốc độ: Truy vấn dashboard chỉ đọc từ view đã pre-computed, giảm thời gian từ giây xuống mili-giây (lên đến 100x nhanh hơn cho queries phức tạp).
  • Near real-time: Hỗ trợ streaming data với cơ chế tự động refresh (incremental refresh), cập nhật dữ liệu mới từ base tables chỉ trong vài phút, đảm bảo dữ liệu luôn tươi mới mà không cần scheduled jobs thủ công.
  • Tích hợp hoàn hảo: Phù hợp với BigQuery Storage Write API cho streaming, và query optimizer tự động sử dụng materialized views khi cần.
    📘 Nguồn tham khảo:
  • BigQuery Materialized Views Documentation (cập nhật 2024-2026: Hỗ trợ base tables là streaming tables, auto-refresh lên đến hàng nghìn views/sc).
  • Best Practices for Dashboard Performance.

❌ Giải thích tất cả các phương án (đúng/sai)

  • Phương án SAI: Disable BigQuery query result caching.
    ❌ Lý do sai: Tắt caching sẽ làm tồi tệ hơn hiệu suất vì buộc BigQuery phải tính toán lại toàn bộ joins/aggregations mỗi lần truy vấn (không cache kết quả). Caching giúp dashboard tải nhanh hơn cho queries lặp lại, đặc biệt với streaming data. Không giải quyết gốc rễ vấn đề phức tạp của joins/aggregations.

  • Phương án SAI: Modify the schema to use parameterized data types.
    ❌ Lý do sai: Parameterized data types (như STRUCT, ARRAY) không cải thiện tốc độ joins/aggregations trên streaming data. Chúng chỉ hỗ trợ định nghĩa schema linh hoạt hơn, nhưng không liên quan đến performance dashboard chậm do query phức tạp. Có thể làm schema phức tạp hơn mà không mang lợi ích real-time.

  • Phương án SAI: Create a scheduled query to calculate and store intermediate results.
    ❌ Lý do sai: Scheduled queries (qua Cloud Scheduler hoặc Airflow) chỉ chạy định kỳ (ví dụ: hàng giờ), dẫn đến dữ liệu không near real-time (có độ trễ lớn với streaming). Không tự động refresh như materialized views, và tốn kém hơn (chi phí compute lặp lại). Phù hợp cho batch, không phải dashboard real-time.

  • Phương án ĐÚNG: Create materialized views.
    ✅ Xác nhận lại: Như đã giải thích ở trên, đây là giải pháp tối ưu nhất với tự động hóa đầy đủ, hỗ trợ streaming và performance vượt trội theo docs BigQuery mới nhất.

🛠️ Lời khuyên thực hành: Sau khi tạo materialized views, dùng EXPLAIN để kiểm tra query plan và monitor qua BigQuery Information Schema để đảm bảo refresh đúng.

Câu 293
You need to create a data pipeline that streams event information from applications in multiple Google Cloud regions into BigQuery for near real-time analysis. The data requires transformation before loading. You want to create the pipeline using a visual interface. What should you do?
  1. A Push event information to a Pub/Sub topic. Create a Dataflow job using the Dataflow job builder.
  2. B Push event information to a Pub/Sub topic. Create a Cloud Run function to subscribe to the Pub/Sub topic, apply transformations, and insert the data into BigQuery.
  3. C Push event information to a Pub/Sub topic. Create a BigQuery subscription in Pub/Sub.
  4. D Push event information to Cloud Storage, and create an external table in BigQuery. Create a BigQuery scheduled job that executes once each day to apply transformations.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống cần xây dựng một data pipeline để stream dữ liệu sự kiện (event information) từ các ứng dụng chạy ở nhiều vùng (regions) của Google Cloud vào BigQuery nhằm thực hiện phân tích gần thời gian thực (near real-time analysis). Dữ liệu cần được biến đổi (transformation) trước khi load vào BigQuery. Yêu cầu đặc biệt là sử dụng giao diện trực quan (visual interface) để tạo pipeline.

🛠️ Yêu cầu chính:

  • Streaming: Dữ liệu liên tục, không phải batch.
  • Multi-region: Hỗ trợ dữ liệu từ nhiều vùng, cần dịch vụ toàn cầu như Pub/Sub.
  • Transformation: Xử lý dữ liệu trước khi lưu.
  • Near real-time: Độ trễ thấp, không phải hàng ngày.
  • Visual interface: Công cụ kéo-thả hoặc builder UI, không code thủ công.

📘 Bối cảnh kiến thức GCP (cập nhật đến 2026): Google Cloud khuyến nghị sử dụng Pub/Sub làm message broker cho streaming multi-region, kết hợp Dataflow (dựa trên Apache Beam) cho ETL streaming với UI builder. Điều này phù hợp với Data Streaming best practices trong Google Cloud Dataflow docs.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Push event information to a Pub/Sub topic. Create a Dataflow job using the Dataflow job builder.

Lý do 🟢:

  • Pub/Sub là dịch vụ messaging toàn cầu, hỗ trợ multi-region replication tự động (global topics), lý tưởng để nhận event từ apps ở nhiều vùng.
  • Dataflow hỗ trợ streaming pipeline từ Pub/Sub, áp dụng transformation qua Apache Beam (windowing, aggregation, etc.), rồi sink trực tiếp vào BigQuery với độ trễ thấp (near real-time, thường <1 phút).
  • Dataflow job builder là giao diện trực quan (visual UI) trong Google Cloud Console, cho phép kéo-thả để thiết kế pipeline mà không cần code YAML/JSON phức tạp – hoàn toàn khớp yêu cầu.
  • Hiệu suất cao, auto-scaling, fault-tolerant cho production workload.

❌ Phân tích tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do chi tiết:

  • Push event information to a Pub/Sub topic. Create a Dataflow job using the Dataflow job builder.
    ✅ Đúng (như đã giải thích ở trên). Đây là giải pháp chuẩn cho streaming ETL với visual UI, hỗ trợ transformation phức tạp và near real-time.

  • Push event information to a Pub/Sub topic. Create a Cloud Run function to subscribe to the Pub/Sub topic, apply transformations, and insert the data into BigQuery.
    ❌ Sai. Cloud Run là serverless container, có thể subscribe Pub/Sub và transform, nhưng không có visual interface để build pipeline (phải code Dockerfile/Cloud Build). Không tối ưu cho high-volume streaming (cold start, scale limit ~1000 instances), dễ lỗi khi insert BigQuery lớn, không auto-scale tốt như Dataflow cho near real-time multi-region.

  • Push event information to a Pub/Sub topic. Create a BigQuery subscription in Pub/Sub.
    ❌ Sai. Pub/Sub hỗ trợ BigQuery sink connector (direct subscription), nhưng không hỗ trợ transformation trước khi load (chỉ raw data). Không đáp ứng yêu cầu "data requires transformation". Dù near real-time, thiếu visual builder cho pipeline phức tạp.

  • Push event information to Cloud Storage, and create an external table in BigQuery. Create a BigQuery scheduled job that executes once each day to apply transformations.
    ❌ Sai. Cloud Storage + external table chỉ cho batch loading, không streaming/near real-time (chạy hàng ngày). Scheduled job BigQuery là batch ETL, độ trễ cao (24h), không phù hợp event streaming. Không có visual interface cho pipeline end-to-end, và external table chỉ query raw mà không transform tự động.

📚 Tài liệu tham khảo (cập nhật GCP 2026)

🛠️ Khuyến nghị: Sử dụng Dataflow UI để prototype nhanh, sau scale với Beam SDK nếu cần custom logic!

Câu 294
You work for an online retail company. Your company collects customer purchase data in CSV files and pushes them to Cloud Storage every 10 minutes. The data needs to be transformed and loaded into BigQuery for analysis. The transformation involves cleaning the data, removing duplicates, and enriching it with product information from a separate table in BigQuery. You need to implement a low-overhead solution that initiates data processing as soon as the files are loaded into Cloud Storage. What should you do?
  1. A Use Cloud Composer sensors to detect files loading in Cloud Storage. Create a Dataproc cluster, and use a Composer task to execute a job on the cluster to process and load the data into BigQuery.
  2. B Schedule a direct acyclic graph (DAG) in Cloud Composer to run hourly to batch load the data from Cloud Storage to BigQuery, and process the data in BigQuery using SQL.
  3. C Use Dataflow to implement a streaming pipeline using an OBJECT_FINALIZE notification from Pub/Sub to read the data from Cloud Storage, perform the transformations, and write the data to BigQuery.
  4. D Create a Cloud Data Fusion job to process and load the data from Cloud Storage into BigQuery. Create an OBJECT_FINALI ZE notification in Pub/Sub, and trigger a Cloud Run function to start the Cloud Data Fusion job as soon as new files are loaded.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong Google Cloud Platform (GCP) (không phải AWS như ghi chú ban đầu, có thể là nhầm lẫn):
Một công ty bán lẻ trực tuyến thu thập dữ liệu mua hàng của khách hàng dưới dạng file CSV, đẩy lên Cloud Storage mỗi 10 phút. Dữ liệu cần được biến đổi (transform) bao gồm: làm sạch (cleaning), loại bỏ trùng lặp (removing duplicates), và bổ sung thông tin sản phẩm từ một bảng riêng trong BigQuery (enriching). Sau đó, load vào BigQuery để phân tích.

Yêu cầu chính:

  • Giải pháp phải low-overhead (chi phí vận hành thấp, không tốn tài nguyên thừa).
  • Khởi động xử lý ngay lập tức khi file được load vào Cloud Storage (near real-time, phù hợp tần suất 10 phút).

🛠️ Thách thức: Cần một pipeline streaming hoặc event-driven để trigger tự động từ sự kiện file mới (như OBJECT_FINALIZE), xử lý transform phức tạp (clean, dedup, join với BQ), và write vào BQ mà không delay lớn hoặc overhead cao.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Dataflow to implement a streaming pipeline using an OBJECT_FINALIZE notification from Pub/Sub to read the data from Cloud Storage, perform the transformations, and write the data to BigQuery.

Lý do:

  • Dataflow là dịch vụ Apache Beam managed, lý tưởng cho streaming pipeline với autoscaling tự động, low-overhead (không cần quản lý cluster thủ công).
  • Sử dụng Pub/Sub notification từ sự kiện OBJECT_FINALIZE của Cloud Storage để trigger ngay lập tức khi file hoàn tất upload (event-driven, phù hợp 10 phút/lần).
  • Pipeline đọc file từ GCS, thực hiện transform (clean, dedup bằng Beam transforms, enrich bằng lookup/join với BQ), rồi write trực tiếp vào BQ.
  • Hỗ trợ cập nhật đến 2026: Dataflow phiên bản mới nhất (và Beam SDK 2.50+) tối ưu streaming với unbounded sources từ Pub/Sub/GCS, chi phí pay-per-use thấp.

📘 Nguồn tham khảo:

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Use Cloud Composer sensors to detect files loading in Cloud Storage. Create a Dataproc cluster, and use a Composer task to execute a job on the cluster to process and load the data into BigQuery.
    Lý do sai: Cloud Composer (Airflow managed) dùng sensors tốt để detect file, nhưng tạo Dataproc cluster mỗi lần có overhead cao (thời gian khởi động cluster ~5-10 phút, chi phí idle). Không low-overhead, phù hợp batch hơn streaming. Dataproc yêu cầu quản lý Spark/Hadoop thủ công, không tối ưu cho event-driven 10 phút/lần.

  • ❌ [SAI] Schedule a direct acyclic graph (DAG) in Cloud Composer to run hourly to batch load the data from Cloud Storage to BigQuery, and process the data in BigQuery using SQL.
    Lý do sai: Lập lịch hourly (giờ/lần) không khớp tần suất 10 phút, gây delay tích lũy (batch processing). Xử lý SQL trong BQ sau load chỉ đơn giản, không hiệu quả cho transform phức tạp (clean/dedup/enrich cần join realtime). Overhead của Composer DAG chạy định kỳ, không trigger ngay khi file arrive.

  • ✅ [ĐÚNG] Use Dataflow to implement a streaming pipeline using an OBJECT_FINALIZE notification from Pub/Sub to read the data from Cloud Storage, perform the transformations, and write the data to BigQuery.
    Lý do đúng: Như đã giải thích ở trên – event-driven hoàn hảo, streaming low-latency, transform đầy đủ bằng Beam, autoscaling zero-overhead khi idle. Phù hợp best practice GCP 2026 cho near-real-time ETL.

  • ❌ [SAI] Create a Cloud Data Fusion job to process and load the data from Cloud Storage into BigQuery. Create an OBJECT_FINALI ZE notification in Pub/Sub, and trigger a Cloud Run function to start the Cloud Data Fusion job as soon as new files are loaded.
    Lý do sai: Cloud Data Fusion (CDF) là low-code ETL tốt, Pub/Sub + Cloud Run trigger đúng hướng, nhưng startup CDF job qua Cloud Run có overhead cao (CDF job khởi động 1-5 phút, cold start Run). Không low-overhead cho tần suất cao 10 phút/lần, dễ timeout/chi phí tích lũy. CDF phù hợp batch lớn hơn streaming frequent. (Lưu ý: "OBJECT_FINALI ZE" là lỗi đánh máy của OBJECT_FINALIZE).

🧩 Tóm tắt: Giải pháp đúng tận dụng Dataflow streaming + Pub/Sub events để đạt low-overhead và immediate trigger, là pattern chuẩn GCP cho dữ liệu incremental! 🚀

Câu 295
You work for a home insurance company. You are frequently asked to create and save risk reports with charts for specific areas using a publicly available storm event dataset. You want to be able to quickly create and re-run risk reports when new data becomes available. What should you do?
  1. A Export the storm event dataset as a CSV file. Import the file to Google Sheets, and use cell data in the worksheets to create charts.
  2. B Copy the storm event dataset into your BigQuery project. Use BigQuery Studio to query and visualize the data in Looker Studio.
  3. C Reference and query the storm event dataset using SQL in BigQuery Studio. Export the results to Google Sheets, and use cell data in the worksheets to create charts.
  4. D Reference and query the storm event dataset using SQL in a Colab Enterprise notebook. Display the table results and document with Markdown, and use Matplotlib to create charts.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi mô tả tình huống bạn làm việc cho một công ty bảo hiểm nhà cửa (home insurance company). Bạn thường xuyên được yêu cầu tạo và lưu báo cáo rủi ro (risk reports) kèm biểu đồ (charts) cho các khu vực cụ thể, dựa trên dataset sự kiện bão công khai (publicly available storm event dataset). Yêu cầu chính là có thể nhanh chóng tạo và chạy lại báo cáo (quickly create and re-run) mỗi khi dữ liệu mới được cập nhật.

📌 Mục tiêu chính: Tìm giải pháp tự động hóa, scalable, dễ tái sử dụng để xử lý dữ liệu lớn, query linh hoạt và visualize nhanh chóng mà không cần thủ công nhiều lần. Đây là kịch bản điển hình trong Google Cloud Data Analytics, tận dụng BigQuery cho lưu trữ/query và công cụ viz tích hợp.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Copy the storm event dataset into your BigQuery project. Use BigQuery Studio to query and visualize the data in Looker Studio.

Lý do 🛠️:

  • Copy dataset vào BigQuery project cho phép lưu trữ dữ liệu lớn một cách scalable, tự động cập nhật khi dữ liệu mới (qua load jobs hoặc scheduled queries). BigQuery là data warehouse serverless, hỗ trợ query SQL nhanh chóng trên petabyte-scale data.
  • BigQuery Studio (ra mắt 2023-2024, cập nhật đến 2026) là IDE tích hợp trong BigQuery console, cho phép viết SQL, tạo views, và kết nối trực tiếp với Looker Studio (trước là Data Studio) để visualize charts/biểu đồ động.
  • Giải pháp này nhanh chóng re-run: Chỉ cần refresh data source trong Looker Studio, báo cáo tự update mà không export thủ công. Hoàn hảo cho reports lặp lại với dữ liệu thời gian thực hoặc gần thực.
  • Theo docs Google Cloud 2026: BigQuery + Looker Studio là best practice cho BI reporting từ public datasets (như NOAA storm data).

Nguồn tham khảo 📘:

❌ Giải thích tất cả các phương án (đúng và sai)

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên tính nhanh chóng, tái sử dụng và scalability khi dữ liệu mới cập nhật.

  • [SAI] Export the storm event dataset as a CSV file. Import the file to Google Sheets, and use cell data in the worksheets to create charts.
    ❌ Lý do sai: Quy trình thủ công hoàn toàn (export CSV → import Sheets → tạo charts từ cell data). Không scalable với dữ liệu lớn (Sheets giới hạn 10M cells), khó re-run tự động khi data mới – phải lặp lại toàn bộ. Không phù hợp cho reports chuyên nghiệp hoặc frequent updates.

  • [ĐÚNG] Copy the storm event dataset into your BigQuery project. Use BigQuery Studio to query and visualize the data in Looker Studio.
    ✅ Lý do đúng (như đã giải thích ở trên): Tích hợp end-to-end trong Google Cloud ecosystem, hỗ trợ scheduled queries/load và viz động. Dễ lưu/share reports dưới dạng Looker dashboards.

  • [SAI] Reference and query the storm event dataset using SQL in BigQuery Studio. Export the results to Google Sheets, and use cell data in the worksheets to create charts.
    ❌ Lý do sai: Dù dùng BigQuery Studio để query (tốt), nhưng export results sang Sheets làm thủ công hóa bước cuối, mất scalability. Re-run yêu cầu export lại mỗi lần, Sheets không handle big data tốt, charts không tự động refresh từ BigQuery source.

  • [SAI] Reference and query the storm event dataset using SQL in a Colab Enterprise notebook. Display the table results and document with Markdown, and use Matplotlib to create charts.
    ❌ Lý do sai: Colab Enterprise tốt cho prototyping/data science, nhưng không phải tool chuyên reports: Matplotlib tạo static charts, khó share/re-run như dashboard (phải run notebook lại). Markdown documentation không thay thế viz chuyên nghiệp. Không tích hợp seamless với public datasets cho BI workflow.

Tóm tắt khuyến nghị 🚀: Chọn giải pháp BigQuery + Looker Studio để tối ưu time-to-insight và automation, phù hợp với Google Cloud Associate Data Practitioner best practices đến 2026!

Câu 296
Your company currently uses an on-premises network file system (NFS) and is migrating data to Google Cloud. You want to be able to control how much bandwidth is used by the data migration while capturing detailed reporting on the migration status. What should you do?
  1. A Use a Transfer Appliance.
  2. B Use Cloud Storage FUSE.
  3. C Use Storage Transfer Service.
  4. D Use gcloud storage commands.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả tình huống công ty đang sử dụng hệ thống file NFS (Network File System) tại chỗ (on-premises) và muốn di chuyển dữ liệu lên Google Cloud. Yêu cầu chính là kiểm soát lượng băng thông (bandwidth) sử dụng trong quá trình di chuyển đồng thời thu thập báo cáo chi tiết về tình trạng migration.
📌 Mục tiêu cốt lõi: Tìm công cụ phù hợp để migrate dữ liệu từ NFS on-prem sang Google Cloud Storage (GCS), với khả năng throttle bandwidth và monitoring/reporting nâng cao. Đây là kịch bản phổ biến trong Google Cloud Data Migration, dựa trên phiên bản mới nhất của Storage Transfer Service (STS) cập nhật đến năm 2026 (hỗ trợ POSIX NFS shares và advanced controls).

✅ Đáp án đúng và lý do lựa chọn

Use Storage Transfer Service.
🛠️ Lý do: Storage Transfer Service (STS) là dịch vụ chuyên dụng của Google Cloud để di chuyển dữ liệu lớn từ on-premises (bao gồm NFS shares) sang GCS. Nó cho phép kiểm soát bandwidth chính xác qua tùy chọn --bandwidth-limit (throttling), và cung cấp báo cáo chi tiết qua Cloud Monitoring, logs, và Transfer Operations API (bao gồm tiến độ, lỗi, bytes transferred). STS hỗ trợ tự động hóa, resume job, và tích hợp NFS trực tiếp – hoàn hảo cho yêu cầu này. (Cập nhật 2026: STS v2 hỗ trợ AI-based optimization cho NFS migrations).

📘 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên tính phù hợp với yêu cầu control bandwidth + detailed reporting cho NFS migration sang Google Cloud.

  • ❌ Use a Transfer Appliance.
    🧩 Phân tích sai: Transfer Appliance là thiết bị vật lý (hardware) để di chuyển dữ liệu ngoại tuyến (offline) lớn (hàng PB), bằng cách copy dữ liệu vào appliance rồi ship đến Google. Nó không hỗ trợ control bandwidth online (vì offline), và reporting chỉ cơ bản qua logs sau khi ship, không chi tiết real-time. Không phù hợp cho NFS online migration.

  • ❌ Use Cloud Storage FUSE.
    🧩 Phân tích sai: Cloud Storage FUSE (gcsfuse) là công cụ mount GCS bucket như file system cục bộ (POSIX-compliant). Nó dùng để truy cập/mount dữ liệu đã có trên GCS, không phải tool migration từ NFS on-prem. Không có tính năng control bandwidth migration hay reporting chi tiết về quá trình chuyển dữ liệu.

  • ✅ Use Storage Transfer Service.
    🛠️ Phân tích đúng (như đã giải thích ở trên): STS đáp ứng hoàn hảo cả hai yêu cầu – bandwidth throttling và reporting nâng cao. Hỗ trợ NFS source trực tiếp qua agent hoặc URL, scale tự động, và tích hợp IAM cho security.

  • ❌ Use gcloud storage commands.
    🧩 Phân tích sai: Lệnh gcloud storage (gsutil) là CLI để copy/sync dữ liệu giữa buckets hoặc từ URL. Nó hỗ trợ --bandwidth limit cơ bản, nhưng reporting kém chi tiết (chỉ logs stdout/stderr, không dashboard/monitoring tự động), và không tối ưu cho NFS on-prem lớn (thiếu resume, parallelism nâng cao). Không phải lựa chọn chuyên dụng cho enterprise migration.

🔗 Tài liệu tham khảo (cập nhật mới nhất 2026)

  • 📖 Storage Transfer Service Documentation – Chi tiết bandwidth control và NFS support.
  • 📖 Transfer Appliance Overview – So sánh offline vs online.
  • 📖 gcsfuse Guide – Xác nhận không phải migration tool.
  • 📘 Google Cloud Skills Boost: "Migrating Data to Google Cloud" module (Associate Cloud Data Practitioner cert path).
    (Nguồn chính thức từ Google Cloud Console và AWS-free, tập trung Google tools đến Q1/2026).
Câu 297
You are a Looker analyst. You need to add a new field to your Looker report that generates SQL that will run against your company's database. You do not have the Develop permission. What should you do?
  1. A Create a new field in the LookML layer, refresh your report, and select your new field from the field picker.
  2. B Create a calculated field using the Add a field option in Looker Studio, and add it to your report.
  3. C Create a table calculation from the field picker in Looker, and add it to your report.
  4. D Create a custom field from the field picker in Looker, and add it to your report.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này thuộc lĩnh vực Looker (một công cụ phân tích dữ liệu và BI của Google Cloud), yêu cầu bạn là một Looker analyst cần thêm một trường dữ liệu mới vào báo cáo Looker. Trường mới này phải tạo ra SQL query chạy trực tiếp trên cơ sở dữ liệu của công ty. Tuy nhiên, bạn không có quyền Develop (quyền này cần thiết để chỉnh sửa LookML - ngôn ngữ mô hình dữ liệu của Looker).

📌 Mục tiêu chính: Tìm cách thêm field mà không cần quyền Develop, nhưng vẫn generate SQL thực thi trên database (không chỉ tính toán trên dữ liệu đã load). Điều này nhấn mạnh vào các tính năng user-friendly trong giao diện Explore của Looker, phù hợp với analyst thông thường. Kiến thức dựa trên phiên bản Looker mới nhất đến 2026 (Looker 23.x+ trên Google Cloud), nơi custom fields được ưu tiên cho người dùng không có quyền dev.

✅ Đáp án đúng

Create a custom field from the field picker in Looker, and add it to your report.

Lý do chọn đáp án này:

  • Trong Looker Explore, bạn có thể tạo Custom Field trực tiếp từ field picker (gear icon bên cạnh field list) mà không cần quyền Develop.
  • Custom Field cho phép viết SQL-like expression (sử dụng Looker SQL dialect), generate SQL thực sự chạy trên database để lấy dữ liệu mới.
  • Sau khi tạo, kéo thả vào báo cáo (Look hoặc dashboard). Đây là cách chuẩn cho analyst, hỗ trợ dynamic fields như case statements, aggregations. 🛠️ Hoàn hảo cho tình huống không có quyền chỉnh LookML!

📘 Giải thích tất cả các phương án

  • Create a new field in the LookML layer, refresh your report, and select your new field from the field picker.
    ❌ Sai: Tạo field trong LookML layer (file .view hoặc .model) yêu cầu quyền Develop để commit code vào Git repo. Không có quyền này, bạn không thể làm. Hơn nữa, sau khi tạo cần deploy và refresh PDT (Persisted Derived Tables) hoặc content validator – quá phức tạp và không khả dụng. (Nguồn: Looker Docs - Permissions)

  • Create a calculated field using the Add a field option in Looker Studio, and add it to your report.
    ❌ Sai: Looker Studio (trước là Google Data Studio) là tool riêng biệt, không kết nối trực tiếp với Looker model và không generate SQL chạy trên database gốc (chỉ tính toán trên dữ liệu đã import). "Add a field" ở đây là calculated field client-side, không phù hợp với Looker report. Nhầm lẫn tool! (Nguồn: Looker Studio vs Looker)

  • Create a table calculation from the field picker in Looker, and add it to your report.
    ❌ Sai: Table calculation (từ field picker > Table calculations) chỉ tính toán trên dữ liệu đã load vào browser (post-aggregation), không generate SQL mới chạy trên database. Ví dụ: running totals, rankings – chỉ client-side, không fetch dữ liệu gốc. Không đáp ứng yêu cầu "generates SQL". (Nguồn: Looker Docs - Table Calculations)

  • Create a custom field from the field picker in Looker, and add it to your report.
    ✅ Đúng: Như giải thích ở trên, đây là lựa chọn tối ưu cho analyst không dev. Custom fields hỗ trợ SQL phức tạp (e.g., ${table}.field + 1), chạy query real-time trên DB. (Nguồn: Looker Docs - Custom Fields)

📚 Tài liệu tham khảo

  • Chính thức Google Cloud Looker Docs (2026 update): Custom Fields Guide, Permissions Overview.
  • Looker Release Notes 2025-2026: Custom fields được enhance với AI-assisted SQL (Gemini integration), nhưng core logic không đổi.
  • Khuyến nghị học: Thực hành trên Looker sandbox miễn phí tại cloud.google.com/looker.

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần ví dụ SQL cụ thể, hỏi thêm nhé!

Câu 298
Your organization’s ecommerce website collects user activity logs using a Pub/Sub topic. Your organization’s leadership team wants a dashboard that contains aggregated user engagement metrics. You need to create a solution that transforms the user activity logs into aggregated metrics, while ensuring that the raw data can be easily queried. What should you do?
  1. A Create a Dataflow subscription to the Pub/Sub topic, and transform the activity logs. Load the transformed data into a BigQuery table for reporting.
  2. B Create an event-driven Cloud Run function to trigger a data transformation pipeline to run. Load the transformed activity logs into a BigQuery table for reporting.
  3. C Create a Cloud Storage subscription to the Pub/Sub topic. Load the activity logs into a bucket using the Avro file format. Use Dataflow to transform the data, and load it into a BigQuery table for reporting.
  4. D Create a BigQuery subscription to the Pub/Sub topic, and load the activity logs into the table. Create a materialized view in BigQuery using SQL to transform the data for reporting
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả tình huống: Tổ chức của bạn có website thương mại điện tử (ecommerce) thu thập nhật ký hoạt động người dùng (user activity logs) thông qua một chủ đề Pub/Sub (Pub/Sub topic). Ban lãnh đạo muốn một bảng điều khiển (dashboard) chứa các chỉ số tổng hợp về sự tương tác của người dùng (aggregated user engagement metrics). Bạn cần xây dựng giải pháp chuyển đổi (transform) các nhật ký hoạt động thành các chỉ số tổng hợp, đồng thời đảm bảo dữ liệu thô (raw data) có thể dễ dàng truy vấn (easily queried).

🛠️ Yêu cầu chính của giải pháp:

  • Xử lý dữ liệu streaming từ Pub/Sub một cách hiệu quả.
  • Chuyển đổi dữ liệu thô thành metrics tổng hợp cho báo cáo/dashboard (ví dụ: tổng số lượt xem, thời gian tương tác...).
  • Giữ dữ liệu thô ở định dạng dễ truy vấn (như BigQuery để hỗ trợ SQL query nhanh chóng).
  • Giải pháp phải scalable, real-time hoặc near-real-time, phù hợp với volume lớn từ logs ecommerce.

📘 Kiến thức nền tảng (cập nhật GCP đến 2026): Pub/Sub là dịch vụ messaging streaming; Dataflow (Apache Beam) là công cụ ETL mạnh mẽ cho stream/batch processing; BigQuery là data warehouse serverless lý tưởng cho analytics và querying raw/transformed data. Giải pháp chuẩn theo best practices GCP là sử dụng Dataflow để pipeline từ Pub/Sub → transform → BigQuery (hỗ trợ cả raw và aggregated).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create a Dataflow subscription to the Pub/Sub topic, and transform the activity logs. Load the transformed data into a BigQuery table for reporting.

Lý do chọn đáp án này 🏆:

  • Dataflow tạo subscription pull trực tiếp từ Pub/Sub topic, xử lý streaming real-time với Apache Beam pipeline để transform logs thành aggregated metrics (ví dụ: windowing, aggregation functions như COUNT, SUM).
  • Dữ liệu transformed load vào BigQuery table, dễ dàng dùng cho dashboard (Looker Studio/Data Studio).
  • Raw data dễ query: Dataflow có thể branch pipeline (một branch lưu raw trực tiếp vào BigQuery khác, một branch transform), đảm bảo raw logs vẫn accessible via SQL. Đây là pattern chuẩn cho log processing trên GCP.
  • Scalable tự động, cost-effective cho high-throughput ecommerce logs. Không cần quản lý infra.

📋 Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn giữ nguyên văn bản gốc tiếng Anh, với đánh giá đúng/sai và lý do bằng tiếng Việt. Sử dụng kiến thức GCP mới nhất (Dataflow 2.x+, BigQuery streaming inserts, Pub/Sub 2026 features như routing policies).

  • ✅ Create a Dataflow subscription to the Pub/Sub topic, and transform the activity logs. Load the transformed data into a BigQuery table for reporting.
    Đúng vì: Như giải thích ở trên. Đây là giải pháp best practice cho streaming ETL từ Pub/Sub → BigQuery. Dataflow hỗ trợ transform phức tạp (aggregation, joining), autoscaling, và exactly-once delivery. Raw data có thể lưu parallel vào BQ table riêng để query dễ dàng.

  • ❌ Create an event-driven Cloud Run function to trigger a data transformation pipeline to run. Load the transformed activity logs into a BigQuery table for reporting.
    Sai vì: Cloud Run là serverless containers cho event-driven (Pub/Sub trigger), nhưng không phù hợp cho continuous streaming high-volume logs (cold starts chậm, giới hạn concurrency 1000 instances, không autoscaling mượt như Dataflow). Trigger pipeline riêng phức tạp, dễ fail dưới load ecommerce lớn, không đảm bảo real-time transform.

  • ❌ Create a Cloud Storage subscription to the Pub/Sub topic. Load the activity logs into a bucket using the Avro file format. Use Dataflow to transform the data, and load it into a BigQuery table for reporting.
    Sai vì: Pub/Sub không hỗ trợ "Cloud Storage subscription" trực tiếp (subscriptions chỉ là Push/Pull/Seek, không phải sink trực tiếp đến Storage). Phải dùng Cloud Storage sink qua Pub/Sub → Storage connector riêng, nhưng Avro format không lý tưởng cho streaming (batch-oriented). Thêm bước này làm phức tạp, chậm, và raw data ở Storage khó query SQL hơn BigQuery.

  • ❌ Create a BigQuery subscription to the Pub/Sub topic, and load the activity logs into the table. Create a materialized view in BigQuery using SQL to transform the data for reporting.
    Sai vì: BigQuery hỗ trợ streaming từ Pub/Sub (subscriptions load raw trực tiếp), và materialized views tốt cho aggregation real-time (tính đến 2026, hỗ trợ incremental refresh). Tuy nhiên, transform bằng SQL materialized view kém hiệu quả cho complex streaming pipelines (không hỗ trợ windowing động, stateful processing như Dataflow). Dễ gặp quota limits dưới high-velocity logs, không scalable bằng Dataflow cho ecommerce.

📚 Tài liệu tham khảo (GCP Docs cập nhật 2026)

Giải pháp này đảm bảo tuân thủ nguyên tắc GCP Well-Architected Framework: Reliable & Efficient! 🚀

Câu 299
You are constructing a data pipeline to process sensitive customer data stored in a Cloud Storage bucket. You need to ensure that this data remains accessible, even in the event of a single-zone outage. What should you do?
  1. A Set up a Cloud CDN in front of the bucket.
  2. B Enable Object Versioning on the bucket.
  3. C Store the data in a multi-region bucket.
  4. D Store the data in Nearline storage.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi yêu cầu xây dựng một data pipeline để xử lý dữ liệu khách hàng nhạy cảm được lưu trữ trong Cloud Storage bucket (thuộc Google Cloud Storage - GCS). Mục tiêu chính là đảm bảo dữ liệu vẫn có thể truy cập được ngay cả khi xảy ra sự cố gián đoạn ở một single-zone (một vùng availability zone duy nhất).

🛠️ Phân tích chi tiết vấn đề:

  • Cloud Storage bucket là dịch vụ lưu trữ object của Google Cloud, nơi dữ liệu được lưu dưới dạng các object.
  • Sensitive customer data: Dữ liệu nhạy cảm cần độ durability (bền vững) và availability cao (khả năng sẵn sàng).
  • Single-zone outage: Sự cố chỉ ảnh hưởng đến một zone trong một region (ví dụ: phần cứng hỏng ở một data center cụ thể). GCS cần cơ chế replication (sao chép dữ liệu) qua nhiều zone hoặc region để tránh downtime.
  • Yêu cầu cốt lõi: Cần giải pháp high availability (HA) và fault-tolerant, không chỉ backup mà phải tự động replicate dữ liệu realtime.

📘 Kiến thức cập nhật (GCS phiên bản mới nhất 2026): Theo tài liệu Google Cloud, bucket multi-regional replicate dữ liệu đồng bộ qua ít nhất 2 regions (mỗi region có >=3 zones), đảm bảo 99.95% monthly uptime SLA, chịu được cả single-zone và single-region outage (nguồn: Google Cloud Storage Classes, Multi-regional buckets).

✅ Đáp án đúng: Store the data in a multi-region bucket

Lý do lựa chọn:

  • Multi-regional bucket tự động replicate dữ liệu đồng bộ qua nhiều regions (ví dụ: US hoặc EU), đảm bảo dữ liệu luôn accessible ngay cả khi một zone (hoặc cả region) outage.
  • Điều này phù hợp hoàn hảo với yêu cầu single-zone outage, vì replication vượt qua ranh giới zone/region.
  • Ưu điểm: Không cần config thủ công, tích hợp sẵn trong GCS, hỗ trợ data pipeline seamless (dùng với Dataflow, Composer).
  • Không ảnh hưởng performance cho access thường xuyên, với durability 99.999999999% (11 9's).

🧪 Giải thích tất cả các phương án

  • ✅ [ĐÚNG] Store the data in a multi-region bucket
    🟢 Đúng vì: Như đã giải thích, multi-regional bucket replicate dữ liệu qua nhiều regions, chịu được single-zone outage (và hơn thế). Đây là giải pháp native của GCS cho HA cao, không cần tool ngoài. (Nguồn: GCS Multi-Region Replication).

  • ❌ [SAI] Set up a Cloud CDN in front of the bucket
    ❌ Sai vì: Cloud CDN chỉ cache nội dung ở edge locations để tăng tốc độ truy cập, không replicate dữ liệu gốc trong bucket. Nếu bucket outage (do zone fail), cache sẽ hết hạn và không accessible. CDN dành cho performance, không phải durability/HA.

  • ❌ [SAI] Enable Object Versioning on the bucket
    ❌ Sai vì: Object Versioning chỉ giữ các phiên bản cũ của object khi overwrite/delete, giúp recovery từ lỗi user. Nó không bảo vệ chống hardware outage (zone fail), dữ liệu gốc vẫn chỉ ở một location.

  • ❌ [SAI] Store the data in Nearline storage
    ❌ Sai vì: Nearline là storage class cho dữ liệu ít access (chi phí thấp, retrieval 3-5 phút), không cải thiện availability. Vẫn phụ thuộc vào bucket location (regional/multi-regional), chỉ thay đổi cost/retrieval time, không chống outage.

Câu 300
Your retail company collects customer data from various sources:
Online transactions: Stored in a MySQL database
Customer feedback: Stored as text files on a company server
Social media activity: Streamed in real-time from social media platforms
You are designing a data pipeline to extract this data. Which Google Cloud storage system(s) should you select for further analysis and ML model training?
  1. A 1. Online transactions: Cloud Storage
    2. Customer feedback: Cloud Storage
    3. Social media activity: Cloud Storage
  2. B 1. Online transactions: BigQuery
    2. Customer feedback: Cloud Storage
    3. Social media activity: BigQuery
  3. C 1. Online transactions: Bigtable
    2. Customer feedback: Cloud Storage
    3. Social media activity: CloudSQL for MySQL
  4. D 1. Online transactions: Cloud SQL for MySQL
    2. Customer feedback: BigQuery
    3. Social media activity: Cloud Storage
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc thiết kế data pipeline để trích xuất dữ liệu từ các nguồn khác nhau của một công ty bán lẻ, nhằm mục đích phân tích dữ liệu sâu và huấn luyện mô hình ML trên Google Cloud. Các nguồn dữ liệu cụ thể bao gồm:

  • Online transactions: Lưu trữ trong cơ sở dữ liệu MySQL (dữ liệu có cấu trúc, giao dịch trực tuyến).
  • Customer feedback: Lưu trữ dưới dạng file text trên server công ty (dữ liệu không cấu trúc, dạng văn bản thô).
  • Social media activity: Dữ liệu streaming thời gian thực từ các nền tảng mạng xã hội (dữ liệu lớn, liên tục, cần xử lý nhanh).

Mục tiêu là chọn hệ thống lưu trữ phù hợp trên Google Cloud cho từng loại dữ liệu để hỗ trợ phân tích (analytics) và ML training. Google Cloud cung cấp các dịch vụ như Cloud Storage (object storage cho dữ liệu không cấu trúc), BigQuery (data warehouse serverless cho phân tích SQL lớn và streaming), Cloud SQL (managed relational DB như MySQL), Bigtable (NoSQL cho dữ liệu lớn thời gian thực).

🛠️ Lý do chọn storage: Phải phù hợp với đặc tính dữ liệu (cấu trúc/không cấu trúc, batch/streaming), hiệu suất phân tích/ML (query nhanh, scale lớn), và tích hợp pipeline (như Dataflow hoặc Pub/Sub cho streaming).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng:

  1. Online transactions: BigQuery
  2. Customer feedback: Cloud Storage
  3. Social media activity: BigQuery

Giải thích lý do:

  • Online transactions (MySQL): BigQuery lý tưởng cho dữ liệu có cấu trúc từ relational DB, hỗ trợ import dễ dàng qua federated queries hoặc ETL, query SQL siêu nhanh trên petabyte-scale, tích hợp ML (BigQuery ML).
  • Customer feedback (text files): Cloud Storage phù hợp lưu trữ object không cấu trúc giá rẻ, scalable, dễ staging cho pipeline trước khi load vào BigQuery hoặc Vertex AI cho ML.
  • Social media activity (streaming real-time): BigQuery hỗ trợ streaming inserts (lên đến 1 triệu rows/giây), xử lý real-time analytics và ML mà không cần batching.
    Kết hợp này tối ưu chi phí, hiệu suất cho further analysis và ML model training theo best practices Google Cloud (cập nhật 2025-2026).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh:

  • ❌ Phương án SAI 1:

    1. Online transactions: Cloud Storage
    2. Customer feedback: Cloud Storage
    3. Social media activity: Cloud Storage
      Giải thích: Cloud Storage chỉ giỏi lưu trữ object không cấu trúc (như feedback), nhưng kém cho dữ liệu có cấu trúc (transactions cần SQL queries) và streaming real-time (không hỗ trợ inserts nhanh như BigQuery). Không phù hợp analysis/ML phức tạp, phải thêm bước ETL tốn kém.
  • ✅ Phương án ĐÚNG 2:

    1. Online transactions: BigQuery
    2. Customer feedback: Cloud Storage
    3. Social media activity: BigQuery
      Giải thích: Như phần trên, sự kết hợp hoàn hảo: BigQuery cho structured/structured analytics & streaming, Cloud Storage cho unstructured staging. Tối ưu pipeline với Dataflow/Pub/Sub.
  • ❌ Phương án SAI 3:

    1. Online transactions: Bigtable
    2. Customer feedback: Cloud Storage
    3. Social media activity: CloudSQL for MySQL
      Giải thích: Bigtable (NoSQL wide-column) phù hợp high-throughput key-value (không phải transactions SQL), CloudSQL for MySQL không scale cho streaming social data lớn (giới hạn real-time inserts so với BigQuery).
  • ❌ Phương án SAI 4:

    1. Online transactions: Cloud SQL for MySQL
    2. Customer feedback: BigQuery
    3. Social media activity: Cloud Storage
      Giải thích: Cloud SQL tốt migrate MySQL nhưng kém analytics lớn/ML (scale kém BigQuery), BigQuery không phải cho text files thô (lãng phí, cần staging ở Cloud Storage trước), Cloud Storage không hỗ trợ streaming analysis.

📘 Tài liệu tham khảo (cập nhật mới nhất 2025-2026)

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần pipeline demo, hỏi thêm nhé!