Ngân hàng đề — Google Cloud Professional Data Engineer

Tìm thấy 429 câu.

Câu 391
You designed a data warehouse in BigQuery to analyze sales data. You want a self-serving, low-maintenance, and cost- effective solution to share the sales dataset to other business units in your organization. What should you do?
  1. A Create an Analytics Hub private exchange, and publish the sales dataset.
  2. B Enable the other business units’ projects to access the authorized views of the sales dataset.
  3. C Create and share views with the users in the other business units.
  4. D Use the BigQuery Data Transfer Service to create a schedule that copies the sales dataset to the other business units’ projects.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi này xoay quanh việc thiết kế một kho dữ liệu (data warehouse) trên BigQuery (dịch vụ kho dữ liệu của Google Cloud) để phân tích dữ liệu bán hàng (sales data). Yêu cầu chính là tìm giải pháp tự phục vụ (self-serving), bảo trì thấp (low-maintenance) và tiết kiệm chi phí (cost-effective) để chia sẻ tập dữ liệu bán hàng cho các đơn vị kinh doanh khác trong tổ chức.

📌 Các tiêu chí quan trọng cần đáp ứng:

  • Self-serving: Người dùng ở các đơn vị khác có thể tự truy cập dữ liệu mà không cần hỗ trợ từ đội ngũ kỹ thuật.
  • Low-maintenance: Không cần quản lý thủ công nhiều, tránh copy dữ liệu định kỳ hoặc cấp quyền phức tạp.
  • Cost-effective: Giảm thiểu chi phí lưu trữ, chuyển dữ liệu và tài nguyên tính toán.
  • Dữ liệu gốc nằm trong BigQuery, và chia sẻ nội bộ tổ chức (cùng organization).

Câu hỏi kiểm tra kiến thức về các tính năng chia sẻ dữ liệu hiện đại trong Google Cloud, đặc biệt là Analytics Hub (ra mắt và cập nhật mạnh mẽ đến năm 2024-2026), thay vì các phương pháp truyền thống tốn kém.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create an Analytics Hub private exchange, and publish the sales dataset.

Lý do chi tiết 🛠️:

  • Analytics Hub là dịch vụ của Google Cloud cho phép tạo private exchange (trao đổi dữ liệu riêng tư) để xuất bản (publish) dataset từ BigQuery. Người dùng ở các dự án (projects) khác trong cùng organization có thể tự subscribe và truy vấn dữ liệu gốc mà không copy, đảm bảo self-serving.
  • Low-maintenance: Dữ liệu luôn đồng bộ thời gian thực (linked data), không cần ETL thủ công hay lịch copy.
  • Cost-effective: Chỉ tính phí truy vấn trên dữ liệu gốc (query on source), tránh nhân bản lưu trữ và chi phí chuyển dữ liệu.
  • Cập nhật mới nhất (2026): Analytics Hub hỗ trợ cross-project, cross-organization sharing với IAM tích hợp, governance mạnh mẽ (data lineage, access logs). Đây là best practice cho data sharing nội bộ theo tài liệu Google Cloud.

📘 Giải thích tất cả các phương án

Dưới đây là phân tích từng phương án một cách chi tiết. Tôi giữ nguyên nội dung văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt với đánh giá đúng/sai.

  • Create an Analytics Hub private exchange, and publish the sales dataset.
    ✅ Đúng – Như đã giải thích ở trên, đây là giải pháp tối ưu nhất, đáp ứng đầy đủ self-serving (subscribe tự động), low-maintenance (không copy), cost-effective (query gốc). Hoàn hảo cho BigQuery sharing nội bộ.

  • Enable the other business units’ projects to access the authorized views of the sales dataset.
    ❌ Sai – Authorized views (views được ủy quyền) cho phép truy cập qua IAM, nhưng không self-serving vì các dự án khác phải được cấp quyền thủ công (cross-project access), dễ gây vấn đề bảo mật và bảo trì cao (quản lý quyền users/projects). Không scalable cho nhiều business units, và có thể tốn kém nếu query lớn.

  • Create and share views with the users in the other business units.
    ❌ Sai – Tạo views và share trực tiếp với users yêu cầu quản lý quyền phức tạp (IAM per user/group), không low-maintenance vì phải cập nhật views thủ công khi dữ liệu gốc thay đổi. Không self-serving (users phụ thuộc admin), và không hiệu quả chi phí nếu nhiều users query riêng lẻ.

  • Use the BigQuery Data Transfer Service to create a schedule that copies the sales dataset to the other business units’ projects.
    ❌ Sai – BigQuery Data Transfer Service dùng để copy dữ liệu định kỳ (scheduled transfers), nhưng tốn kém cao (double storage + transfer costs), không low-maintenance (phải quản lý lịch, schema changes), và không self-serving (dữ liệu copy có thể lỗi thời). Vi phạm nguyên tắc "zero-copy sharing" hiện đại của GCP.

🔗 Tài liệu tham khảo (cập nhật đến 2026)

Giải pháp này giúp tổ chức xây dựng data mesh hiện đại! 🚀

Câu 392
You have terabytes of customer behavioral data streaming from Google Analytics into BigQuery daily. Your customers’ information, such as their preferences, is hosted on a Cloud SQL for MySQL database. Your CRM database is hosted on a Cloud SQL for PostgreSQL instance. The marketing team wants to use your customers’ information from the two databases and the customer behavioral data to create marketing campaigns for yearly active customers. You need to ensure that the marketing team can run the campaigns over 100 times a day on typical days and up to 300 during sales. At the same time, you want to keep the load on the Cloud SQL databases to a minimum. What should you do?
  1. A Create BigQuery connections to both Cloud SQL databases. Use BigQuery federated queries on the two databases and the Google Analytics data on BigQuery to run these queries.
  2. B Create a job on Apache Spark with Dataproc Serverless to query both Cloud SQL databases and the Google Analytics data on BigQuery for these queries.
  3. C Create streams in Datastream to replicate the required tables from both Cloud SQL databases to BigQuery for these queries.
  4. D Create a Dataproc cluster with Trino to establish connections to both Cloud SQL databases and BigQuery, to execute the queries.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong Google Cloud Platform (GCP), nơi bạn quản lý dữ liệu lớn từ hành vi khách hàng (terabytes) streaming hàng ngày từ Google Analytics vào BigQuery. Thông tin khách hàng như sở thích lưu trữ trên Cloud SQL for MySQL, còn cơ sở dữ liệu CRM trên Cloud SQL for PostgreSQL. Nhóm marketing cần kết hợp dữ liệu từ hai Cloud SQL database này với dữ liệu behavioral trên BigQuery để tạo chiến dịch marketing nhắm đến khách hàng active hàng năm.

Yêu cầu chính:

  • Tần suất query cao: Hơn 100 lần/ngày thường xuyên, lên đến 300 lần trong mùa sale (peak time).
  • Giảm tải tối đa trên Cloud SQL: Không muốn các query ad-hoc làm ảnh hưởng đến database nguồn, vì Cloud SQL không scale tốt cho workload query lớn và thường xuyên.

Mục tiêu là xây dựng giải pháp offload dữ liệu khỏi Cloud SQL, đưa vào BigQuery (nơi đã có dữ liệu GA và hỗ trợ query siêu nhanh, serverless, scale vô hạn), để marketing team chạy query trực tiếp trên BigQuery mà không chạm đến Cloud SQL. Giải pháp phải real-time hoặc near-real-time để dữ liệu luôn tươi mới. (Kiến thức cập nhật GCP 2024-2026: BigQuery hỗ trợ streaming inserts/updates hiệu quả, Datastream là dịch vụ CDC chuẩn cho replication từ Cloud SQL sang BigQuery).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create streams in Datastream to replicate the required tables from both Cloud SQL databases to BigQuery for these queries.

Lý do:

  • Datastream là dịch vụ Change Data Capture (CDC) serverless của GCP, hỗ trợ replication real-time/low-latency từ Cloud SQL (MySQL/PostgreSQL) sang BigQuery mà không làm tải Cloud SQL (chỉ đọc WAL logs, không query trực tiếp).
  • Dữ liệu từ hai Cloud SQL được replicate vào các BigQuery tables (routine/heartbeats cho schema changes), kết hợp liền mạch với dữ liệu GA đã có trên BigQuery.
  • Marketing team query chỉ trên BigQuery (SQL chuẩn), scale dễ dàng cho 100-300 queries/ngày mà zero load trên Cloud SQL.
  • Tối ưu chi phí: Serverless, pay-per-use, phù hợp workload cao tần suất. Hỗ trợ yearly active customers qua window functions hoặc ML integrations trong BigQuery.
  • Ưu việt hơn các lựa chọn khác vì decouple hoàn toàn nguồn dữ liệu, đảm bảo HA và freshness (latency <1 phút).

📋 Giải thích chi tiết tất cả các phương án

  • ❌ [SAI] Create BigQuery connections to both Cloud SQL databases. Use BigQuery federated queries on the two databases and the Google Analytics data on BigQuery to run these queries.
    Phương án này sử dụng BigQuery external connections (federated queries) để query trực tiếp Cloud SQL từ BigQuery. Sai vì: Mỗi lần query (100-300 lần/ngày) sẽ hit trực tiếp Cloud SQL (qua JDBC/ODBC), gây tải CPU/IOPS cao, không scale cho peak time, vi phạm yêu cầu "keep the load on Cloud SQL to a minimum". Federated queries chậm với TB data, latency cao (seconds to minutes).

  • ❌ [SAI] Create a job on Apache Spark with Dataproc Serverless to query both Cloud SQL databases and the Google Analytics data on BigQuery for these queries.
    Sử dụng Dataproc Serverless với Spark để tạo job query cross-source. Sai vì: Phải tạo job mới mỗi lần query (không phù hợp ad-hoc 100-300 lần/ngày), chi phí cao (spin-up ephemeral clusters), và vẫn query trực tiếp Cloud SQL gây tải. Spark tốt cho batch/ETL lớn, nhưng không lý tưởng cho marketing queries nhanh/lặp lại.

  • ✅ [ĐÚNG] Create streams in Datastream to replicate the required tables from both Cloud SQL databases to BigQuery for these queries.
    (Đã giải thích chi tiết ở phần đáp án đúng). Hoàn hảo vì: Replication one-time setup, query native BigQuery (fast, scalable, no SQL load), hỗ trợ MySQL 5.7+/PostgreSQL 9.6+ (cập nhật 2026).

  • ❌ [SAI] Create a Dataproc cluster with Trino to establish connections to both Cloud SQL databases and BigQuery, to execute the queries.
    Tạo Dataproc cluster với Trino (federated query engine, trước là Presto) để connect Cloud SQL + BigQuery. Sai vì: Cluster luôn chạy (chi phí cao ~$0.5-1/giờ), queries vẫn federate trực tiếp Cloud SQL gây tải I/O/CPU (không minimum load), không serverless thực sự. Trino tốt cho ad-hoc multi-source, nhưng không giải quyết vấn đề chính: bảo vệ Cloud SQL.

📘 Tài liệu tham khảo (GCP docs cập nhật 2024-2026)

Giải pháp này đảm bảo performance, cost-effective và scalable theo best practices GCP! 🚀

Câu 393
Your organization is modernizing their IT services and migrating to Google Cloud. You need to organize the data that will be stored in Cloud Storage and BigQuery. You need to enable a data mesh approach to share the data between sales, product design, and marketing departments. What should you do?
  1. A 1. Create a project for storage of the data for each of your departments.
    2. Enable each department to create Cloud Storage buckets and BigQuery datasets.
    3. Create user groups for authorized readers for each bucket and dataset.
    4. Enable the IT team to administer the user groups to add or remove users as the departments’ request.
  2. B 1. Create multiple projects for storage of the data for each of your departments’ applications.
    2. Enable each department to create Cloud Storage buckets and BigQuery datasets.
    3. Publish the data that each department shared in Analytics Hub.
    4. Enable all departments to discover and subscribe to the data they need in Analytics Hub.
  3. C 1. Create a project for storage of the data for your organization.
    2. Create a central Cloud Storage bucket with three folders to store the files for each department.
    3. Create a central BigQuery dataset with tables prefixed with the department name.
    4. Give viewer rights for the storage project for the users of your departments.
  4. D 1. Create multiple projects for storage of the data for each of your departments’ applications.
    2. Enable each department to create Cloud Storage buckets and BigQuery datasets.
    3. In Dataplex, map each department to a data lake and the Cloud Storage buckets, and map the BigQuery datasets to zones.
    4. Enable each department to own and share the data of their data lakes.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này xoay quanh việc tổ chức dữ liệu trong Google Cloud khi tổ chức đang hiện đại hóa dịch vụ IT và di chuyển lên Google Cloud. 📈 Cụ thể, bạn cần lưu trữ dữ liệu vào Cloud Storage và BigQuery, đồng thời áp dụng data mesh approach để chia sẻ dữ liệu giữa các bộ phận: sales (bán hàng), product design (thiết kế sản phẩm) và marketing (tiếp thị).

Data mesh là một kiến trúc dữ liệu phân tán (decentralized), nơi mỗi bộ phận (domain) sở hữu, quản lý và chia sẻ dữ liệu của mình như một sản phẩm (data as product). 🛤️ Nó nhấn mạnh vào tính tự chủ (autonomy) của từng domain, governance thống nhất, và discovery/sharing dễ dàng. Không phải mô hình tập trung (monolithic) hay chỉ chia sẻ thủ công.

Mục tiêu: Enable a data mesh approach to share the data – tức là kích hoạt cách tiếp cận data mesh để chia sẻ dữ liệu giữa các bộ phận một cách hiệu quả, an toàn và tự chủ.

Kiến thức cập nhật (đến 2026): Google Cloud sử dụng Dataplex (phiên bản mới nhất hỗ trợ data mesh đầy đủ từ 2021, cập nhật với Lakehouse federation và AI governance năm 2025) làm nền tảng chính cho data mesh. Analytics Hub phù hợp cho shared datasets nhưng không đầy đủ cho domain ownership. (Nguồn: Dataplex Data Mesh Guide, Google Cloud Blog 2025).


✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng là phương án cuối cùng (đã đánh dấu [ĐÚNG]):

  1. Create multiple projects for storage of the data for each of your departments’ applications.
  2. Enable each department to create Cloud Storage buckets and BigQuery datasets.
  3. In Dataplex, map each department to a data lake and the Cloud Storage buckets, and map the BigQuery datasets to zones.
  4. Enable each department to own and share the data of their data lakes.

Lý do chọn đáp án này 🏆:
Phương án này triển khai đúng data mesh bằng cách:

  • Sử dụng multiple projects cho từng ứng dụng của bộ phận → Tự chủ (domain autonomy).
  • Cho phép bộ phận tạo buckets/datasets riêng → Decentralized ownership.
  • Dataplex là công cụ cốt lõi: Map bộ phận thành data lake (domain-level), buckets thành assets trong lake, datasets thành zones (logical grouping với governance).
  • Bộ phận own và share data lakes → Hỗ trợ self-service discovery, sharing qua tags/policies/IAM, phù hợp data mesh principles (federated governance).
    📘 Điều này khớp hoàn hảo với best practices Google Cloud cho data mesh, tránh monolithic và cho phép scale với AI/ML integration (cập nhật 2026).

❌ Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng phương án một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên data mesh principles: domain ownership, self-service sharing, governance, không centralized control. 🛠️

  • [SAI] Phương án 1:

    1. Create a project for storage of the data for each of your departments.
    2. Enable each department to create Cloud Storage buckets and BigQuery datasets.
    3. Create user groups for authorized readers for each bucket and dataset.
    4. Enable the IT team to administer the user groups to add or remove users as the departments’ request.

    Tại sao SAI 🚫: Phương án này chỉ tạo projects riêng và dùng user groups/IAM thủ công để chia sẻ, nhưng IT team kiểm soát hoàn toàn (admin groups theo request). Điều này tạo centralized gatekeeping, vi phạm data mesh (domain tự chủ). Không có discovery tự động hay governance layer → Không scale, dễ bottleneck. Không dùng Dataplex/Analytics Hub.

  • [SAI] Phương án 2:

    1. Create multiple projects for storage of the data for each of your departments’ applications.
    2. Enable each department to create Cloud Storage buckets and BigQuery datasets.
    3. Publish the data that each department shared in Analytics Hub.
    4. Enable all departments to discover and subscribe to the data they need in Analytics Hub.

    Tại sao SAI 🚫: Analytics Hub tốt cho shared datasets (subscribe/protect data), nhưng không hỗ trợ full data mesh. Nó tập trung vào BigQuery datasets, thiếu domain-level organization (lakes/zones), và không map storage/buckets toàn diện. Data mesh cần Dataplex cho unified catalog/governance trên đa nguồn (Storage + BigQuery). Hub chỉ là một phần, không decentralized ownership đầy đủ (cập nhật 2026: Hub tích hợp Dataplex nhưng không thay thế).

  • [SAI] Phương án 3:

    1. Create a project for storage of the data for your organization.
    2. Create a central Cloud Storage bucket with three folders to store the files for each department.
    3. Create a central BigQuery dataset with tables prefixed with the department name.
    4. Give viewer rights for the storage project for the users of your departments.

    Tại sao SAI 🚫: Đây là monolithic approach (một project/bucket/dataset trung tâm), chỉ dùng folders/prefixes và viewer rights. Vi phạm data mesh hoàn toàn vì thiếu tự chủ (tất cả centralized), dễ conflict access, no domain isolation. Không scale cho sharing phức tạp, governance yếu (chỉ IAM cơ bản).

  • [ĐÚNG] Phương án 4 (đã giải thích ở trên): ✅ Hoàn hảo cho data mesh với Dataplex làm nền tảng.

Kết luận 🌟: Chọn Dataplex để enable data mesh là best practice Google Cloud. Tham khảo thêm: Dataplex Intro, Data Mesh Whitepaper. Nếu triển khai, bắt đầu với Dataplex Security Center cho governance! 🚀

Câu 394
You work for a large ecommerce company. You are using Pub/Sub to ingest the clickstream data to Google Cloud for analytics. You observe that when a new subscriber connects to an existing topic to analyze data, they are unable to subscribe to older data. For an upcoming yearly sale event in two months, you need a solution that, once implemented, will enable any new subscriber to read the last 30 days of data. What should you do?
  1. A Create a new topic, and publish the last 30 days of data each time a new subscriber connects to an existing topic.
  2. B Set the topic retention policy to 30 days.
  3. C Set the subscriber retention policy to 30 days.
  4. D Ask the source system to re-push the data to Pub/Sub, and subscribe to it.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống thực tế trong một công ty thương mại điện tử lớn sử dụng Google Cloud Pub/Sub để thu thập dữ liệu clickstream (dữ liệu theo dõi hành vi người dùng click) nhằm phục vụ phân tích.

  • Vấn đề hiện tại 📉: Khi một subscriber mới (người đăng ký mới) kết nối vào một topic (chủ đề) hiện có để phân tích dữ liệu, họ không thể truy cập dữ liệu cũ (older data). Điều này xảy ra vì Pub/Sub mặc định chỉ giữ message trong thời gian ngắn (tối đa 7 ngày), và subscriber mới chỉ nhận message mới sau khi subscribe.

  • Yêu cầu giải pháp 🎯: Với sự kiện bán hàng hàng năm sắp tới (trong 2 tháng), cần triển khai giải pháp một lần duy nhất để bất kỳ subscriber mới nào cũng có thể đọc được dữ liệu 30 ngày gần nhất (last 30 days of data). Giải pháp phải đảm bảo tính khả dụng cao, không phụ thuộc vào việc tạo topic mới hay yêu cầu hệ thống nguồn đẩy lại dữ liệu.

Đây là câu hỏi kiểm tra kiến thức sâu về Pub/Sub retention policy trong Google Cloud (cập nhật đến 2026: Pub/Sub hỗ trợ retention lên đến 7 ngày mặc định, nhưng có thể cấu hình tối đa 31 ngày cho topic).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Set the topic retention policy to 30 days.

Lý do 🛠️:

  • Pub/Sub cho phép cấu hình topic retention policy để giữ lại message chưa được acknowledge (unacknowledged) trong topic lên đến 31 ngày (tính đến 2026). Khi set retention = 30 ngày, tất cả message trong 30 ngày qua sẽ được lưu trữ trên topic, cho phép bất kỳ subscriber mới nào (pull hoặc push) đọc được dữ liệu cũ mà không cần làm gì thêm.
  • Giải pháp này triển khai một lần, hiệu lực ngay lập tức cho tất cả subscriber mới, phù hợp với yêu cầu sự kiện lớn sắp tới. Không ảnh hưởng đến publisher và chi phí thấp (chỉ tính storage cho message retained).
  • Cập nhật mới nhất: Theo tài liệu Google Cloud Pub/Sub 2026, retention policy áp dụng cho toàn topic, đảm bảo replayability cho new subscribers.

📋 Giải thích tất cả các phương án (đúng/sai)

  • [SAI] Create a new topic, and publish the last 30 days of data each time a new subscriber connects to an existing topic.
    ❌ Sai vì: Phương án này không khả thi và không hiệu quả. Việc tạo topic mới mỗi khi có subscriber mới sẽ dẫn đến quản lý phức tạp, duplicate dữ liệu, và chi phí cao. Hơn nữa, publisher phải đẩy lại dữ liệu 30 ngày thủ công mỗi lần – vi phạm yêu cầu "triển khai một lần" và không tự động cho sự kiện lớn.

  • [ĐÚNG] Set the topic retention policy to 30 days.
    ✅ Đúng vì: Như đã giải thích ở trên, đây là giải pháp chuẩn và tối ưu của Pub/Sub. Retention policy giữ message trên topic (tối đa 31 ngày), cho phép new subscriber pull/read dữ liệu cũ ngay lập tức. Triển khai qua gcloud pubsub topics update TOPIC --retention-duration=30d hoặc Console.

  • [SAI] Set the subscriber retention policy to 30 days.
    ❌ Sai vì: Pub/Sub không có "subscriber retention policy". Retention chỉ áp dụng cho topic (message storage), không phải subscriber. Subscriber chỉ kiểm soát ack deadline hoặc subscription expiration, không lưu trữ dữ liệu cũ cho new subscribers.

  • [SAI] Ask the source system to re-push the data to Pub/Sub, and subscribe to it.
    ❌ Sai vì: Yêu cầu hệ thống nguồn (source system) đẩy lại toàn bộ dữ liệu 30 ngày là không thực tế, tốn kém, và không tự động. Với clickstream dữ liệu lớn (high volume), việc re-push sẽ gây overload Pub/Sub, duplicate message, và không đảm bảo cho "bất kỳ new subscriber nào" trong tương lai.

📘 Tài liệu tham khảo (cập nhật 2026)

Giải pháp này đảm bảo scalability cao cho sự kiện bán hàng! 🚀

Câu 395
You are designing the architecture to process your data from Cloud Storage to BigQuery by using Dataflow. The network team provided you with the Shared VPC network and subnetwork to be used by your pipelines. You need to enable the deployment of the pipeline on the Shared VPC network. What should you do?
  1. A Assign the compute.networkUser role to the Dataflow service agent.
  2. B Assign the compute.networkUser role to the service account that executes the Dataflow pipeline.
  3. C Assign the dataflow.admin role to the Dataflow service agent.
  4. D Assign the dataflow.admin role to the service account that executes the Dataflow pipeline.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc thiết kế kiến trúc xử lý dữ liệu từ Cloud Storage đến BigQuery bằng Dataflow trên Google Cloud Platform (GCP). 🔄 Vấn đề chính là Shared VPC network (mạng ảo chia sẻ) và subnetwork do team mạng cung cấp. Bạn cần kích hoạt triển khai pipeline Dataflow trên Shared VPC này.

⚠️ Bối cảnh kỹ thuật: Dataflow sử dụng các worker VM trên Compute Engine. Để worker truy cập Shared VPC (thuộc host project khác), cần quyền IAM phù hợp trên host project. Không cấp quyền đúng sẽ dẫn đến lỗi triển khai pipeline (ví dụ: "Permission denied on resource"). Kiến thức dựa trên tài liệu GCP cập nhật đến 2024-2026: Dataflow yêu cầu compute.networkUser (roles/compute.networkUser) cho service account của worker để attach vào subnetwork của Shared VPC. 📘
Nguồn tham khảo:

✅ Đáp án đúng

Assign the compute.networkUser role to the service account that executes the Dataflow pipeline.

Lý do chọn: 🛠️ Service account này (thường là worker service account, ví dụ: project-number-compute@developer.gserviceaccount.com hoặc custom SA) chạy các worker VM của Dataflow. Cấp roles/compute.networkUser trên host project (chứa Shared VPC) cho phép SA attach subnetwork, tạo instance trên mạng chia sẻ. Đây là yêu cầu bắt buộc theo docs GCP. Không cấp quyền này, pipeline thất bại ở giai đoạn provisioning workers. ✅

📋 Giải thích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên IAM best practices của Dataflow trên Shared VPC. 🧐

  • [SAI] Assign the compute.networkUser role to the Dataflow service agent.
    ❌ Sai vì: Dataflow service agent (project-number@dataflow-service-producer-prod.iam.gserviceaccount.com) chỉ chịu trách nhiệm tạo/staging resources (như templates, staging buckets) trên service producer project. Nó không chạy worker VMs nên không cần networkUser cho Shared VPC. Cấp quyền này thừa và không giải quyết vấn đề attach subnetwork. Thay vào đó, service agent cần roles/dataflow.serviceAgent.

  • [ĐÚNG] Assign the compute.networkUser role to the service account that executes the Dataflow pipeline.
    ✅ Đúng vì: Như giải thích ở trên, đây chính là worker service account (executes pipeline tasks). Quyền compute.networkUser cho phép SA truy cập subnetwork trên host project, triển khai workers thành công. Bắt buộc phải cấp trên host project, không phải service project. Đây là bước chính thức theo checklist Shared VPC của Dataflow.

  • [SAI] Assign the dataflow.admin role to the Dataflow service agent.
    ❌ Sai vì: dataflow.admin (roles/dataflow.admin) là quyền cao cấp để quản lý pipelines (create, update, delete jobs), phù hợp cho user/admin, không liên quan đến network access. Service agent chỉ cần roles/dataflow.serviceAgent để orchestrate, không phải dataflow.admin. Cấp quyền này không giúp Shared VPC và vi phạm nguyên tắc least privilege.

  • [SAI] Assign the dataflow.admin role to the service account that executes the Dataflow pipeline.
    ❌ Sai vì: Worker SA chỉ cần quyền Compute Engine cơ bản + networkUser cho Shared VPC (như compute.viewer, compute.instanceAdmin.v1 trên service project). dataflow.admin quá rộng (quản lý tất cả jobs), không giải quyết vấn đề network attachment, và có thể gây rủi ro bảo mật. Không phải yêu cầu cho Shared VPC.

Tóm tắt khuyến nghị triển khai 🚀:

  1. Xác định host project (Shared VPC).
  2. Gán compute.networkUser cho worker SA trên host project.
  3. Sử dụng --serviceAccountEmail khi launch job để chỉ định SA.
    Test bằng lệnh gcloud dataflow jobs run để verify! 💡
Câu 396
Your infrastructure team has set up an interconnect link between Google Cloud and the on-premises network. You are designing a high-throughput streaming pipeline to ingest data in streaming from an Apache Kafka cluster hosted on- premises. You want to store the data in BigQuery, with as minimal latency as possible. What should you do?
  1. A Setup a Kafka Connect bridge between Kafka and Pub/Sub. Use a Google-provided Dataflow template to read the data from Pub/Sub, and write the data to BigQuery.
  2. B Use a proxy host in the VPC in Google Cloud connecting to Kafka. Write a Dataflow pipeline, read data from the proxy host, and write the data to BigQuery.
  3. C Use Dataflow, write a pipeline that reads the data from Kafka, and writes the data to BigQuery.
  4. D Setup a Kafka Connect bridge between Kafka and Pub/Sub. Write a Dataflow pipeline, read the data from Pub/Sub, and write the data to BigQuery.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả tình huống: Đội ngũ hạ tầng đã thiết lập liên kết Interconnect giữa Google Cloud và mạng on-premises. Bạn đang thiết kế một pipeline streaming hiệu suất cao (high-throughput) để thu thập dữ liệu streaming từ cụm Apache Kafka nằm trên on-premises. Mục tiêu là lưu trữ dữ liệu vào BigQuery với độ trễ thấp nhất có thể (minimal latency).

🛠️ Yêu cầu chính:

  • Sử dụng kết nối Interconnect (Dedicated hoặc Partner Interconnect) để đảm bảo throughput cao và latency thấp giữa on-premises và Google Cloud.
  • Tập trung vào giải pháp streaming real-time, tận dụng các dịch vụ Google Cloud như Dataflow để xử lý dữ liệu từ Kafka trực tiếp vào BigQuery.
  • Tránh các lớp trung gian không cần thiết để giảm latency.

📘 Kiến thức liên quan (cập nhật đến 2026): Dataflow (Apache Beam) hỗ trợ Kafka IO Connector native, cho phép đọc trực tiếp từ Kafka clusters (bao gồm on-premises) qua VPC peering hoặc Interconnect mà không cần proxy hay bridge trung gian. Điều này đảm bảo end-to-end low latency cho high-throughput streaming. (Nguồn: Google Cloud Dataflow Kafka Connector Docs - phiên bản mới nhất hỗ trợ Kafka 3.x và autoscaling cho throughput cao).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Dataflow, write a pipeline that reads the data from Kafka, and writes the data to BigQuery.

Lý do:

  • Dataflow có Kafka connector tích hợp (io.kafkakafkaio.ReadFromKafka), cho phép đọc trực tiếp từ Kafka on-premises qua Interconnect mà không cần lớp trung gian.
  • Viết pipeline đơn giản để streaming trực tiếp vào BigQuery (sử dụng BigQueryIO.Write), đảm bảo minimal latency (sub-second) và high-throughput (hàng GB/s với autoscaling).
  • Đây là giải pháp tối ưu nhất, tuân thủ best practices của Google Cloud cho hybrid streaming pipelines. Không thêm overhead từ Pub/Sub hay proxy, tận dụng kết nối vật lý Interconnect để đạt hiệu suất cao nhất.

❌ Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên latency, throughput, độ phức tạp và tính phù hợp với Interconnect.

  • [SAI] Setup a Kafka Connect bridge between Kafka and Pub/Sub. Use a Google-provided Dataflow template to read the data from Pub/Sub, and write the data to BigQuery.
    ❌ Sai vì: Thêm Kafka Connect bridge và Pub/Sub làm lớp trung gian (Kafka → Connect → Pub/Sub → Dataflow → BigQuery), tăng latency đáng kể (thêm 100-500ms/hop) và phức tạp hóa pipeline. Mặc dù có template sẵn (Pub/Sub to BigQuery), nhưng không tận dụng trực tiếp Interconnect, vi phạm yêu cầu "minimal latency". Không cần thiết khi Dataflow đọc Kafka native.

  • [SAI] Use a proxy host in the VPC in Google Cloud connecting to Kafka. Write a Dataflow pipeline, read data from the proxy host, and write the data to BigQuery.
    ❌ Sai vì: Proxy host (như VM trong VPC) tạo thêm single point of failure và bottleneck, tăng latency (dữ liệu phải qua proxy trước khi vào Dataflow). Phức tạp quản lý (cần scale proxy thủ công), throughput thấp hơn so với Kafka connector trực tiếp. Interconnect hỗ trợ kết nối trực tiếp, không cần proxy.

  • [ĐÚNG] Use Dataflow, write a pipeline that reads the data from Kafka, and writes the data to BigQuery.
    ✅ Đúng như đã giải thích ở trên: Trực tiếp, low-latency, high-throughput. Sử dụng Beam SDK để code pipeline: p.apply(KafkaIO.read().withBootstrapServers(...)) |-> BigQueryIO.write(). Hỗ trợ exactly-once semantics và autoscaling.

  • [SAI] Setup a Kafka Connect bridge between Kafka and Pub/Sub. Write a Dataflow pipeline, read the data from Pub/Sub, and write the data to BigQuery.
    ❌ Sai vì: Tương tự lựa chọn đầu, bridge Kafka Connect → Pub/Sub thêm overhead không cần thiết, tăng latency và chi phí (Pub/Sub có quota). Custom Dataflow pipeline từ Pub/Sub vẫn kém hơn đọc Kafka trực tiếp. Không tận dụng Interconnect tối ưu.

🛠️ Khuyến nghị triển khai

  • Sử dụng Apache Beam SDK (Java/Python/Go) với Kafka 2.0+ và cấu hình security (SASL/SSL).
  • Test với Dataflow Flex Templates cho Kafka-to-BigQuery nếu cần production-ready.
  • Nguồn tham khảo thêm:

Giải pháp này đảm bảo performance tối ưu cho enterprise streaming! 🚀

Câu 397
You migrated your on-premises Apache Hadoop Distributed File System (HDFS) data lake to Cloud Storage. The data scientist team needs to process the data by using Apache Spark and SQL. Security policies need to be enforced at the column level. You need a cost-effective solution that can scale into a data mesh. What should you do?
  1. A 1. Deploy a long-living Dataproc cluster with Apache Hive and Ranger enabled.
    2. Configure Ranger for column level security.
    3. Process with Dataproc Spark or Hive SQL.
  2. B 1. Define a BigLake table.
    2. Create a taxonomy of policy tags in Data Catalog.
    3. Add policy tags to columns.
    4. Process with the Spark-BigQuery connector or BigQuery SQL.
  3. C 1. Load the data to BigQuery tables.
    2. Create a taxonomy of policy tags in Data Catalog.
    3. Add policy tags to columns.
    4. Process with the Spark-BigQuery connector or BigQuery SQL.
  4. D 1. Apply an Identity and Access Management (IAM) policy at the file level in Cloud Storage.
    2. Define a BigQuery external table for SQL processing.
    3. Use Dataproc Spark to process the Cloud Storage files.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này thuộc chủ đề Google Cloud Platform (GCP), tập trung vào việc xử lý data lake sau khi migrate từ on-premises HDFS sang Cloud Storage. Các yêu cầu chính bao gồm:

  • Data scientist cần xử lý dữ liệu bằng Apache Spark và SQL.
  • Áp dụng bảo mật ở mức column-level (fine-grained column-level security).
  • Giải pháp phải cost-effective (tiết kiệm chi phí), và scale vào data mesh (mở rộng linh hoạt theo kiến trúc data mesh, nơi dữ liệu phân tán và tự quản lý domain).

🛠️ Bối cảnh chính: Dữ liệu vẫn nằm trên Cloud Storage (không load vào warehouse), cần query trực tiếp mà không di chuyển dữ liệu để tiết kiệm chi phí, hỗ trợ Spark/SQL, và bảo mật column-level qua policy tags. Giải pháp phải tận dụng các dịch vụ GCP mới như BigLake để đạt scalability cao.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng là phương án thứ 2:

  1. Define a BigLake table.
  2. Create a taxonomy of policy tags in Data Catalog.
  3. Add policy tags to columns.
  4. Process with the Spark-BigQuery connector or BigQuery SQL.

Lý do chọn đáp án này 🏆:

  • BigLake (ra mắt 2022, cập nhật đến 2026) là giải pháp serverless cho data lake, cho phép định nghĩa external tables trên Cloud Storage mà không cần load dữ liệu, hỗ trợ column-level security qua policy tags từ Data Catalog.
  • Policy tags áp dụng trực tiếp lên columns, enforced tự động khi query bằng Spark (qua Spark-BigQuery connector) hoặc BigQuery SQL – đáp ứng yêu cầu Spark/SQL.
  • Cost-effective: Chỉ tính phí query (pay-per-use), không cần cluster luôn chạy, dữ liệu giữ nguyên trên Storage rẻ tiền.
  • Scale to data mesh: BigLake hỗ trợ multi-cloud/hybrid, domain-based governance qua Data Catalog, dễ mở rộng phân tán.

❌ Giải thích tất cả các phương án

Dưới đây là phân tích từng phương án một cách chi tiết:

  • Phương án 1 (SAI):

    1. Deploy a long-living Dataproc cluster with Apache Hive and Ranger enabled.
    2. Configure Ranger for column level security.
    3. Process with Dataproc Spark or Hive SQL.
      Lý do SAI 🚫: Ranger hỗ trợ column-level security trong Hive, nhưng Dataproc cluster long-living tốn kém (luôn charge compute), không cost-effective cho data lake lớn. Không scale tốt vào data mesh vì phụ thuộc cluster quản lý thủ công, thiếu governance phân tán như Data Catalog.
  • Phương án 2 (ĐÚNG): (Đã giải thích chi tiết ở phần trên ✅).

  • Phương án 3 (SAI):

    1. Load the data to BigQuery tables.
    2. Create a taxonomy of policy tags in Data Catalog.
    3. Add policy tags to columns.
    4. Process with the Spark-BigQuery connector or BigQuery SQL.
      Lý do SAI 🚫: Load toàn bộ data lake từ Cloud Storage vào BigQuery tables (managed tables) rất tốn kém (storage + load cost), không phù hợp data lake lớn (terabytes/PB-scale). BigQuery lý tưởng cho warehouse, không phải lake-on-storage, vi phạm yêu cầu cost-effective và giữ data trên Storage.
  • Phương án 4 (SAI):

    1. Apply an Identity and Access Management (IAM) policy at the file level in Cloud Storage.
    2. Define a BigQuery external table for SQL processing.
    3. Use Dataproc Spark to process the Cloud Storage files.
      Lý do SAI 🚫: IAM policy chỉ ở file-level (không column-level), không đáp ứng security yêu cầu. BigQuery external table hỗ trợ SQL nhưng thiếu policy tags cho columns; kết hợp Dataproc Spark vẫn cần cluster, kém cost-effective và không scale data mesh mượt mà như BigLake.

📘 Tài liệu tham khảo (cập nhật mới nhất đến 2026)

Giải pháp này hoàn hảo cho data mesh trên GCP! 🚀 Nếu cần ví dụ code hoặc demo, hãy cho tôi biết nhé!

Câu 398
One of your encryption keys stored in Cloud Key Management Service (Cloud KMS) was exposed. You need to re- encrypt all of your CMEK-protected Cloud Storage data that used that key, and then delete the compromised key. You also want to reduce the risk of objects getting written without customer-managed encryption key (CMEK) protection in the future. What should you do?
  1. A Rotate the Cloud KMS key version. Continue to use the same Cloud Storage bucket.
  2. B Create a new Cloud KMS key. Set the default CMEK key on the existing Cloud Storage bucket to the new one.
  3. C Create a new Cloud KMS key. Create a new Cloud Storage bucket. Copy all objects from the old bucket to the new one bucket while specifying the new Cloud KMS key in the copy command.
  4. D Create a new Cloud KMS key. Create a new Cloud Storage bucket configured to use the new key as the default CMEK key. Copy all objects from the old bucket to the new bucket without specifying a key.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh tình huống an ninh dữ liệu khẩn cấp trong Google Cloud Platform (GCP): Một khóa mã hóa (encryption key) lưu trữ trong Cloud Key Management Service (Cloud KMS) đã bị lộ (exposed). Nhiệm vụ chính là:

  • Re-encrypt toàn bộ dữ liệu trong Cloud Storage đang được bảo vệ bởi Customer-Managed Encryption Key (CMEK) sử dụng khóa bị lộ đó.
  • Xóa khóa bị lộ (delete the compromised key) sau khi hoàn tất.
  • Giảm rủi ro tương lai: Đảm bảo các object mới được ghi vào bucket sẽ tự động có CMEK protection, tránh tình trạng ghi dữ liệu không mã hóa bằng khóa do khách hàng quản lý.

🛠️ Bối cảnh kỹ thuật (cập nhật đến 2026 theo tài liệu GCP mới nhất):

  • CMEK cho phép khách hàng quản lý khóa mã hóa qua Cloud KMS.
  • Dữ liệu cũ trong bucket vẫn giữ nguyên khóa ban đầu trừ khi copy/re-encrypt thủ công.
  • Bucket có thể đặt default CMEK key để tự động áp dụng cho mọi object mới.
  • Khi copy object, nếu không chỉ định key, hệ thống sẽ dùng default CMEK của bucket đích (nếu có).
  • Khóa KMS có thể rotate version, nhưng version cũ vẫn tồn tại và có thể bị lộ.

📘 Nguồn tham khảo:

✅ Đáp án đúng

Create a new Cloud KMS key. Create a new Cloud Storage bucket configured to use the new key as the default CMEK key. Copy all objects from the old bucket to the new bucket without specifying a key.

Lý do chọn đáp án này:

  • Tạo khóa KMS mới 🆕, đảm bảo khóa cũ bị lộ có thể xóa an toàn.
  • Tạo bucket mới và đặt default CMEK key là khóa mới → Tự động bảo vệ mọi object mới ghi vào bucket (giảm rủi ro tương lai ✅).
  • Copy object từ bucket cũ sang mới mà không chỉ định key → Dữ liệu cũ được tự động re-encrypt bằng default CMEK của bucket mới (dùng gsutil cp -r hoặc Storage Transfer Service).
  • Sau copy, xóa khóa cũ và bucket cũ. Hoàn hảo cho yêu cầu toàn diện!

❌ Phân tích tất cả các phương án

  • [SAI] Rotate the Cloud KMS key version. Continue to use the same Cloud Storage bucket.
    ❌ Lý do sai: Rotate chỉ tạo version mới của khóa cũ, nhưng version bị lộ vẫn tồn tại và có thể truy cập, không re-encrypt dữ liệu cũ (dữ liệu vẫn dùng version lộ). Không xóa khóa compromised, và không giảm rủi ro object mới (vẫn dùng bucket cũ với key rotate không an toàn). Không đáp ứng re-encrypt + delete key.

  • [SAI] Create a new Cloud KMS key. Set the default CMEK key on the existing Cloud Storage bucket to the new one.
    ❌ Lý do sai: Tạo key mới và đặt default CMEK cho bucket cũ → Chỉ object mới dùng key mới, nhưng dữ liệu cũ vẫn giữ key lộ (không re-encrypt tự động). Phải xóa key cũ thủ công, nhưng dữ liệu cũ vẫn dễ bị rủi ro. Không giải quyết re-encrypt toàn bộ.

  • [SAI] Create a new Cloud KMS key. Create a new Cloud Storage bucket. Copy all objects from the old bucket to the new one bucket while specifying the new Cloud KMS key in the copy command.
    ❌ Lý do sai: Copy với specify key mới → Re-encrypt dữ liệu cũ OK, nhưng bucket mới không có default CMEK → Object mới ghi vào không tự động CMEK-protected (phải specify thủ công mỗi lần, tăng rủi ro tương lai). Không đáp ứng "reduce the risk" hoàn chỉnh.

🛠️ Lời khuyên thực tế: Sử dụng Storage Transfer Service hoặc gsutil cp để copy hàng loạt, kết hợp IAM để kiểm soát quyền. Test trên môi trường dev trước! 🚀

Câu 399
You have an upstream process that writes data to Cloud Storage. This data is then read by an Apache Spark job that runs on Dataproc. These jobs are run in the us-central1 region, but the data could be stored anywhere in the United States. You need to have a recovery process in place in case of a catastrophic single region failure. You need an approach with a maximum of 15 minutes of data loss (RPO=15 mins). You want to ensure that there is minimal latency when reading the data. What should you do?
  1. A 1. Create two regional Cloud Storage buckets, one in the us-central1 region and one in the us-south1 region.
    2. Have the upstream process write data to the us-central1 bucket. Use the Storage Transfer Service to copy data hourly from the us-central1 bucket to the us-south1 bucket.
    3. Run the Dataproc cluster in a zone in the us-central1 region, reading from the bucket in that region.
    4. In case of regional failure, redeploy your Dataproc clusters to the us-south1 region and read from the bucket in that region instead.
  2. B 1. Create a Cloud Storage bucket in the US multi-region.
    2. Run the Dataproc cluster in a zone in the us-central1 region, reading data from the US multi-region bucket.
    3. In case of a regional failure, redeploy the Dataproc cluster to the us-central2 region and continue reading from the same bucket.
  3. C 1. Create a dual-region Cloud Storage bucket in the us-central1 and us-south1 regions.
    2. Enable turbo replication.
    3. Run the Dataproc cluster in a zone in the us-central1 region, reading from the bucket in the us-south1 region.
    4. In case of a regional failure, redeploy your Dataproc cluster to the us-south1 region and continue reading from the same bucket.
  4. D 1. Create a dual-region Cloud Storage bucket in the us-central1 and us-south1 regions.
    2. Enable turbo replication.
    3. Run the Dataproc cluster in a zone in the us-central1 region, reading from the bucket in the same region.
    4. In case of a regional failure, redeploy the Dataproc clusters to the us-south1 region and read from the same bucket.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một quy trình dữ liệu trên Google Cloud Platform (GCP):

  • Một quy trình upstream ghi dữ liệu vào Cloud Storage.
  • Dữ liệu này được đọc bởi job Apache Spark chạy trên Dataproc.
  • Các job chạy ở vùng us-central1, nhưng dữ liệu có thể lưu trữ ở bất kỳ đâu tại Mỹ.
  • Yêu cầu chính:
    • Xây dựng quy trình phục hồi (recovery) cho trường hợp thất bại thảm họa một vùng duy nhất (single region failure).
    • RPO (Recovery Point Objective) tối đa 15 phút → Mất dữ liệu không quá 15 phút.
    • Độ trễ (latency) khi đọc dữ liệu phải tối thiểu.
      Mục tiêu là thiết kế giải pháp lưu trữ và đọc dữ liệu sao cho an toàn cao (không phụ thuộc một vùng), đồng bộ nhanh (đáp ứng RPO), và đọc nhanh (low latency, ưu tiên đọc local region).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng là lựa chọn thứ 4:

  1. Create a dual-region Cloud Storage bucket in the us-central1 and us-south1 regions.
  2. Enable turbo replication.
  3. Run the Dataproc cluster in a zone in the us-central1 region, reading from the bucket in the same region.
  4. In case of a regional failure, redeploy your Dataproc clusters to the us-south1 region and read from the same bucket.

Lý do chọn đáp án này 🛠️:

  • Dual-region bucket (us-central1 + us-south1) lưu trữ dữ liệu đồng bộ giữa hai vùng xa nhau, đảm bảo không mất dữ liệu nếu một vùng fail.
  • Turbo replication kích hoạt sao chép near real-time (latency <1 phút, dễ dàng đáp ứng RPO=15 phút), dữ liệu mới ghi vào vùng chính sẽ nhanh chóng có ở vùng phụ.
  • Đọc dữ liệu local (cluster us-central1 đọc bucket local → minimal latency).
  • Failover: Redeploy Dataproc sang us-south1 và đọc local → Giữ low latency, không cần thay đổi bucket.
    Giải pháp này cân bằng hoàn hảo giữa RPO thấp, high availability, và low read latency.

❌ Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn (giữ nguyên văn bản gốc bằng tiếng Anh). Mỗi phương án được đánh giá dựa trên RPO=15 phút và minimal latency:

  • [SAI] Lựa chọn 1:

    1. Create two regional Cloud Storage buckets, one in the us-central1 region and one in the us-south1 region.
    2. Have the upstream process write data to the us-central1 bucket. Use the Storage Transfer Service to copy data hourly from the us-central1 bucket to the us-south1 bucket.
    3. Run the Dataproc cluster in a zone in the us-central1 region, reading from the bucket in that region.
    4. In case of regional failure, redeploy your Dataproc clusters to the us-south1 region and read from the bucket in that region instead.
      Giải thích sai ❌: Storage Transfer Service copy hourly (mỗi giờ) → Có thể mất dữ liệu lên đến 60 phút (vượt RPO=15 phút). Không đáp ứng yêu cầu đồng bộ nhanh, dù read latency thấp ở chế độ bình thường.
  • [SAI] Lựa chọn 2:

    1. Create a Cloud Storage bucket in the US multi-region.
    2. Run the Dataproc cluster in a zone in the us-central1 region, reading data from the US multi-region bucket.
    3. In case of a regional failure, redeploy the Dataproc cluster to the us-central2 region and continue reading from the same bucket.
      Giải thích sai ❌: Multi-region bucket highly available (không fail single region), đáp ứng RPO (synchronous replication toàn US). Nhưng read latency cao vì dữ liệu phân tán khắp US (có thể đọc từ xa hàng nghìn km), không "minimal latency" như yêu cầu. Redeploy sang us-central2 vẫn có thể gặp latency tương tự.
  • [SAI] Lựa chọn 3:

    1. Create a dual-region Cloud Storage bucket in the us-central1 and us-south1 regions.
    2. Enable turbo replication.
    3. Run the Dataproc cluster in a zone in the us-central1 region, reading from the bucket in the us-south1 region.
    4. In case of a regional failure, redeploy your Dataproc cluster to the us-south1 region and continue reading from the same bucket.
      Giải thích sai ❌: Dual-region + turbo replication tốt cho RPO (<15 phút). Nhưng đọc cross-region (us-central1 đọc us-south1) → latency cao (khoảng 50-100ms+), vi phạm "minimal latency". Failover thì ổn, nhưng chế độ bình thường đã kém.
  • [ĐÚNG] Lựa chọn 4 (như đã phân tích ở trên) ✅: Hoàn hảo về RPO, HA, và low latency nhờ read local + turbo replication.

🧩 Tóm tắt: Giải pháp đúng tận dụng dual-region với turbo để cân bằng durability và performance, phù hợp best practices GCP 2026!

Câu 400
You currently have transactional data stored on-premises in a PostgreSQL database. To modernize your data environment, you want to run transactional workloads and support analytics needs with a single database. You need to move to Google Cloud without changing database management systems, and minimize cost and complexity. What should you do?
  1. A Migrate and modernize your database with Cloud Spanner.
  2. B Migrate your workloads to AlloyDB for PostgreSQL.
  3. C Migrate to BigQuery to optimize analytics.
  4. D Migrate your PostgreSQL database to Cloud SQL for PostgreSQL.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống: Bạn đang có dữ liệu transactional (giao dịch, OLTP) lưu trữ trên cơ sở dữ liệu PostgreSQL tại chỗ (on-premises). Mục tiêu là hiện đại hóa môi trường dữ liệu (modernize data environment), chạy workloads giao dịch (transactional workloads) và hỗ trợ nhu cầu phân tích (analytics needs) bằng một cơ sở dữ liệu duy nhất (single database). Yêu cầu di chuyển lên Google Cloud mà không thay đổi hệ quản trị cơ sở dữ liệu (không đổi DBMS, tức vẫn giữ PostgreSQL), đồng thời giảm thiểu chi phí và độ phức tạp (minimize cost and complexity).

📘 Tóm tắt yêu cầu chính:

  • PostgreSQL compatible (tương thích PostgreSQL để tránh thay đổi code).
  • Hỗ trợ cả OLTP (transactional) và OLAP (analytics) trong một DB duy nhất.
  • Managed service trên Google Cloud, dễ migrate từ on-premises.
  • Ưu tiên cost-effective và low complexity (dịch vụ tự động hóa cao, scale tốt).

🛠️ Bối cảnh kiến thức GCP (cập nhật đến 2026): AlloyDB for PostgreSQL là dịch vụ mới (GA từ 2022, cải tiến liên tục), kết hợp engine PostgreSQL với công nghệ columnar storage và vector search, hỗ trợ hybrid OLTP/OLAP outperform Cloud SQL gấp 4x cho analytics mà vẫn giữ transactional integrity.

✅ Đáp án đúng: Migrate your workloads to AlloyDB for PostgreSQL

Lý do chọn đáp án này (bằng tiếng Việt chi tiết):
AlloyDB for PostgreSQL là lựa chọn hoàn hảo vì:

  • Tương thích 100% PostgreSQL (wire-compatible, hỗ trợ migrate trực tiếp từ on-premises PostgreSQL qua Database Migration Service - DMS mà không cần thay đổi ứng dụng).
  • Single database cho cả transactional và analytics: Sử dụng AlloyDB columnar engine (tích hợp sẵn) để tăng tốc query analytics lên 10-30x so với row-based PostgreSQL thông thường, mà vẫn đảm bảo ACID transactions cho OLTP.
  • Minimize cost & complexity: Fully managed, auto-scaling, high availability (99.999%), và chi phí thấp hơn nhờ storage columnar nén dữ liệu tốt (giảm 80% storage cost cho analytics). Migrate dễ dàng với zero-downtime qua DMS hoặc pg_dump.
  • So với các option khác, đây là duy nhất đáp ứng single DB hybrid workload mà không cần tách OLTP/OLAP.

🧩 Nguồn tham khảo:

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Migrate and modernize your database with Cloud Spanner.
    Phân tích sai: Cloud Spanner là distributed SQL database (NewSQL) với schema riêng (không tương thích PostgreSQL trực tiếp), yêu cầu thay đổi code ứng dụng lớn (không giữ nguyên DBMS). Nó mạnh OLTP global-scale nhưng không hỗ trợ analytics columnar tốt cho single DB workload như yêu cầu. Complexity cao (schema migration), cost đắt hơn cho workloads nhỏ. Không phù hợp "không thay đổi DBMS".

  • ✅ [ĐÚNG] Migrate your workloads to AlloyDB for PostgreSQL.
    Phân tích đúng (như phần trên): Đáp ứng toàn bộ yêu cầu - PostgreSQL native, single DB OLTP+analytics, low cost/complexity. Best fit cho modernize on-premises PostgreSQL.

  • ❌ [SAI] Migrate to BigQuery to optimize analytics.
    Phân tích sai: BigQuery là serverless data warehouse (OLAP-only), không hỗ trợ transactional workloads (no ACID OLTP, read-only cho queries). Phải tách DB (PostgreSQL cho OLTP + BigQuery cho analytics), vi phạm "single database". Migrate từ PostgreSQL cần ETL phức tạp (không trực tiếp), tăng complexity và cost lâu dài.

  • ❌ [SAI] Migrate your PostgreSQL database to Cloud SQL for PostgreSQL.
    Phân tích sai: Cloud SQL là managed PostgreSQL tốt cho OLTP thuần túy, migrate dễ dàng. Nhưng không hỗ trợ analytics mạnh (row-based storage chậm cho complex queries, cần extension riêng như Citus - phức tạp). Không phải "single DB tối ưu cho cả hai", performance analytics kém AlloyDB (chỉ 1/4 tốc độ columnar). Không "modernize" đủ mức cho hybrid needs.

🔍 Kết luận: AlloyDB là giải pháp tiên tiến nhất 2026 cho PostgreSQL hybrid workloads trên GCP, giúp doanh nghiệp tiết kiệm 30-50% cost so với multi-DB setup! 🚀