Ngân hàng đề — Google Cloud Professional Cloud Database Engineer

Tìm thấy 169 câu.

Câu 11
Your organization operates in a highly regulated industry. Separation of concerns (SoC) and security principle of least privilege (PoLP) are critical. The operations team consists of:
Person A is a database administrator.
Person B is an analyst who generates metric reports.
Application C is responsible for automatic backups.
You need to assign roles to team members for Cloud Spanner. Which roles should you assign?
  1. A roles/spanner.databaseAdmin for Person A
    roles/spanner.databaseReader for Person B
    roles/spanner.backupWriter for Application C
  2. B roles/spanner.databaseAdmin for Person A
    roles/spanner.databaseReader for Person B
    roles/spanner.backupAdmin for Application C
  3. C roles/spanner.databaseAdmin for Person A
    roles/spanner.databaseUser for Person B
    roles/spanner databaseReader for Application C
  4. D roles/spanner.databaseAdmin for Person A
    roles/spanner.databaseUser for Person B
    roles/spanner.backupWriter for Application C
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc gán vai trò IAM (Identity and Access Management) phù hợp cho các thành viên trong tổ chức hoạt động trong ngành có quy định nghiêm ngặt (highly regulated industry). Các nguyên tắc cốt lõi là Separation of Concerns (SoC - Phân tách trách nhiệm) và Principle of Least Privilege (PoLP - Nguyên tắc quyền hạn tối thiểu), nhằm đảm bảo mỗi người/chứng chỉ chỉ có quyền cần thiết nhất để thực hiện nhiệm vụ, tránh rủi ro bảo mật.

  • Person A: Là database administrator (quản trị viên cơ sở dữ liệu) → Cần quyền đầy đủ để quản lý database (tạo, xóa, cấu hình, v.v.).
  • Person B: Là analyst tạo báo cáo metric (báo cáo chỉ số) → Chỉ cần quyền đọc dữ liệu (read-only), không cần viết hoặc chỉnh sửa.
  • Application C: Ứng dụng chịu trách nhiệm sao lưu tự động (automatic backups) → Chỉ cần quyền tạo backup, không cần quản lý toàn diện.

Mục tiêu: Gán roles cho Cloud Spanner (dịch vụ database phân tán của Google Cloud) sao cho tuân thủ SoC và PoLP. 📘 Tài liệu tham khảo: Cloud Spanner IAM roles và Cloud Spanner backups IAM (cập nhật đến 2024-2026, không thay đổi lớn về roles cơ bản).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng:
roles/spanner.databaseAdmin for Person A
roles/spanner.databaseReader for Person B
roles/spanner.backupWriter for Application C

Lý do chi tiết 🛠️:

  • Person A nhận roles/spanner.databaseAdmin: Vai trò này cung cấp quyền quản lý toàn diện database (tạo/sửa/xóa database, quản lý schema, backup cơ bản), phù hợp với DBA mà không vượt quá PoLP.
  • Person B nhận roles/spanner.databaseReader: Chỉ cho phép đọc dữ liệu (query SELECT), lý tưởng cho analyst tạo báo cáo metric, đảm bảo read-only và tuân thủ PoLP.
  • Application C nhận roles/spanner.backupWriter: Chỉ cho phép tạo backup (CREATE BACKUP), hoàn hảo cho automatic backups, tránh quyền xóa hoặc quản lý khác (SoC).
    Kết hợp này đảm bảo tối thiểu quyền hạn, phân tách rõ ràng trách nhiệm, phù hợp ngành regulated. ✅ Hoàn hảo!

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên IAM roles của Cloud Spanner (phiên bản mới nhất 2026: Không có thay đổi lớn, roles vẫn như docs chính thức).

  • ✅ Phương án ĐÚNG (như trên):
    roles/spanner.databaseAdmin for Person A
    roles/spanner.databaseReader for Person B
    roles/spanner.backupWriter for Application C
    Giải thích: Hoàn toàn phù hợp PoLP và SoC. DatabaseAdmin cho DBA đầy đủ; databaseReader chỉ đọc cho analyst; backupWriter chỉ tạo backup cho app. Không quyền thừa! 🏆

  • ❌ Phương án SAI 1:
    roles/spanner.databaseAdmin for Person A
    roles/spanner.databaseReader for Person B
    roles/spanner.backupAdmin for Application C
    Giải thích: Sai vì roles/spanner.backupAdmin cấp quyền quản lý toàn diện backup (tạo, xóa, liệt kê, khôi phục), vượt PoLP cho app chỉ cần automatic backups. App có thể xóa backup nhầm → rủi ro bảo mật cao! 🚫

  • ❌ Phương án SAI 2:
    roles/spanner.databaseAdmin for Person A
    roles/spanner.databaseUser for Person B
    roles/spanner databaseReader for Application C
    Giải thích: Hai lỗi lớn! (1) roles/spanner.databaseUser cho B cho phép đọc + viết SQL (INSERT/UPDATE/DELETE), analyst chỉ cần đọc metric → vi phạm PoLP. (2) roles/spanner databaseReader (thiếu dấu cách "spanner.databaseReader") cho C là quyền đọc dữ liệu, không liên quan backup → app không thể tạo backup, phá vỡ chức năng! 🧨

  • ❌ Phương án SAI 3:
    roles/spanner.databaseAdmin for Person A
    roles/spanner.databaseUser for Person B
    roles/spanner.backupWriter for Application C
    Giải thích: Sai ở B: roles/spanner.databaseUser cấp quyền viết dữ liệu SQL, analyst chỉ tạo báo cáo metric (read-only) → thừa quyền, vi phạm PoLP và có thể gây thay đổi dữ liệu nhầm. C đúng nhưng tổng thể fail! ⚠️

Kết luận 🎯: Chỉ phương án đầu tiên tuân thủ nghiêm ngặt SoC & PoLP. Khuyến nghị test roles qua GCP Console hoặc gcloud CLI để verify! 📘

Câu 12
You are designing an augmented reality game for iOS and Android devices. You plan to use Cloud Spanner as the primary backend database for game state storage and player authentication. You want to track in-game rewards that players unlock at every stage of the game. During the testing phase, you discovered that costs are much higher than anticipated, but the query response times are within the SLA. You want to follow Google-recommended practices. You need the database to be performant and highly available while you keep costs low. What should you do?
  1. A Manually scale down the number of nodes after the peak period has passed.
  2. B Use interleaving to co-locate parent and child rows.
  3. C Use the Cloud Spanner query optimizer to determine the most efficient way to execute the SQL query.
  4. D Use granular instance sizing in Cloud Spanner and Autoscaler.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi mô tả tình huống thiết kế một trò chơi thực tế ảo tăng cường (augmented reality game) cho thiết bị iOS và Android, sử dụng Cloud Spanner làm cơ sở dữ liệu chính để lưu trữ trạng thái trò chơi (game state) và xác thực người chơi (player authentication). Đặc biệt, cần theo dõi phần thưởng trong game mà người chơi mở khóa ở mỗi giai đoạn. Trong giai đoạn thử nghiệm, chi phí cao hơn dự kiến nhưng thời gian phản hồi truy vấn (query response times) vẫn nằm trong SLA (Service Level Agreement). Yêu cầu tuân thủ best practices của Google, đảm bảo database performant (hiệu suất cao), highly available (có tính sẵn sàng cao) đồng thời giảm chi phí thấp nhất. Vấn đề cốt lõi là tối ưu hóa chi phí compute mà không ảnh hưởng hiệu suất, vì latency đã OK.

✅ Đáp án đúng

Use granular instance sizing in Cloud Spanner and Autoscaler.

Lý do lựa chọn:
Cloud Spanner hỗ trợ granular instance sizing (kích thước instance chi tiết) kết hợp Autoscaler (tự động scale theo nhu cầu thực tế), cho phép điều chỉnh compute capacity linh hoạt từ 1000 processing units (PU) trở lên thay vì các node lớn cố định. Điều này giúp giảm chi phí khi load thấp (như sau peak testing), tự động scale up/down theo workload mà vẫn đảm bảo 99.999% availability và hiệu suất cao. Đây là Google-recommended practice mới nhất (từ phiên bản Spanner v2 autoscaling, cập nhật đến 2026), phù hợp với tình huống chi phí cao do provisioned capacity dư thừa, latency đã tốt.

📝 Phân tích tất cả các phương án

  • ❌ [SAI] Manually scale down the number of nodes after the peak period has passed.
    Phương án này không theo best practices vì manual scaling dễ lỗi thời, tốn công quản lý và không dự đoán được workload biến động của game (như peak player). Cloud Spanner khuyến khích autoscaling tự động thay vì thủ công, tránh downtime và over-provisioning lâu dài.

  • ❌ [SAI] Use interleaving to co-locate parent and child rows.
    Interleaving giúp co-locate dữ liệu cha-con (như player và rewards), giảm latency cho join queries nhưng không giải quyết vấn đề chi phí compute. Vấn đề ở đây là chi phí node/processing units cao, không phải cấu trúc dữ liệu, và interleaving có thể tăng storage cost nhẹ.

  • ❌ [SAI] Use the Cloud Spanner query optimizer to determine the most efficient way to execute the SQL query.
    Query optimizer giúp tối ưu execution plan của SQL, cải thiện latency nhưng không giảm chi phí compute cơ bản (nodes/PU). Latency đã trong SLA, nên không cần thiết; chi phí cao chủ yếu từ provisioned capacity dư thừa trong testing.

📘 Tài liệu tham khảo

  • Google Cloud Spanner Documentation: Autoscaling and granular sizing (cập nhật 2025-2026: hỗ trợ autoscaling từ 1000 PU, regional/multi-region).
  • Best Practices for Spanner: Optimize costs – Khuyến nghị autoscaler cho workload biến động như gaming.
  • Release Notes: Spanner v2 GA autoscaling (2023+), tích hợp granular sizing đến 2026. 🛠️
Câu 13
You recently launched a new product to the US market. You currently have two Bigtable clusters in one US region to serve all the traffic. Your marketing team is planning an immediate expansion to APAC. You need to roll out the regional expansion while implementing high availability according to Google-recommended practices. What should you do?
  1. A Maintain a target of 23% CPU utilization by locating:
    cluster-a in zone us-central1-a
    cluster-b in zone europe-west1-d
    cluster-c in zone asia-east1-b
  2. B Maintain a target of 23% CPU utilization by locating:
    cluster-a in zone us-central1-a
    cluster-b in zone us-central1-b
    cluster-c in zone us-east1-a
  3. C Maintain a target of 35% CPU utilization by locating:
    cluster-a in zone us-central1-a
    cluster-b in zone australia-southeast1-a
    cluster-c in zone europe-west1-d
    cluster-d in zone asia-east1-b
  4. D Maintain a target of 35% CPU utilization by locating:
    cluster-a in zone us-central1-a
    cluster-b in zone us-central2-a
    cluster-c in zone asia-northeast1-b
    cluster-d in zone asia-east1-b
Xem giải thích

🧩 Giải thích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống bạn đã triển khai sản phẩm mới tại thị trường Mỹ (US) với hai cụm Bigtable (clusters) nằm trong một vùng (region) US duy nhất, đang phục vụ toàn bộ lưu lượng truy cập. Nhóm marketing dự định mở rộng ngay lập tức sang khu vực APAC (Asia-Pacific). Yêu cầu là triển khai mở rộng theo vùng địa lý này đồng thời đảm bảo high availability (tính sẵn sàng cao) theo thực hành khuyến nghị của Google.

Mục tiêu chính:

  • Bigtable là dịch vụ NoSQL database của Google Cloud, hỗ trợ replication qua nhiều clusters để tăng tính sẵn sàng và giảm độ trễ.
  • Khuyến nghị Google: Sử dụng ít nhất 3-4 clusters phân bố ở các vùng và zone khác nhau để tránh single point of failure (lỗi một điểm duy nhất). Đối với mở rộng US-APAC, cần clusters ở US (multi-region nếu có thể) và APAC (nhiều vùng gần người dùng như Tokyo, Taiwan).
  • Target CPU utilization: Google khuyến nghị duy trì khoảng 35% CPU/node để xử lý spike traffic mà không overload (theo docs cập nhật 2024-2026, không thay đổi lớn).
  • High availability practices: Clusters phải ở zones khác nhau, ưu tiên multi-region (ví dụ: us-central1 cho US-central, asia-northeast1/asia-east1 cho APAC). Tránh cùng region duy nhất để chịu được regional outage.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Maintain a target of 35% CPU utilization by locating:
cluster-a in zone us-central1-a
cluster-b in zone us-central2-a
cluster-c in zone asia-northeast1-b
cluster-d in zone asia-east1-b

Lý do 🛠️:

  • 35% CPU utilization chính xác theo khuyến nghị Google: Giúp Bigtable tự động scale, chịu tải đột biến lên đến 70-80% mà không throttle (giảm hiệu suất).
  • Phân bố clusters lý tưởng cho US-APAC:
    • cluster-a (us-central1-a) và cluster-b (us-central2-a): Hai vùng US khác nhau (us-central1: Iowa-central, us-central2: hỗ trợ multi-region US), đảm bảo HA trong US hiện tại.
    • cluster-c (asia-northeast1-b: Tokyo) và cluster-d (asia-east1-b: Taiwan): Hai vùng APAC gần nhau, giảm latency cho traffic mới, hỗ trợ replication cross-region.
  • Tổng 4 clusters multi-zone/multi-region, tuân thủ "Google-recommended practices" cho production HA, tránh outage từ một region.

❌ Phân tích tất cả các phương án

Dưới đây là phân tích từng phương án một cách chi tiết, giữ nguyên văn bản gốc. Mỗi cái được đánh giá đúng/sai dựa trên best practices Bigtable (HA, CPU target, phân bố vùng phù hợp US-APAC).

  • Phương án SAI: Maintain a target of 23% CPU utilization by locating:
    cluster-a in zone us-central1-a
    cluster-b in zone europe-west1-d
    cluster-c in zone asia-east1-b
    Giải thích sai ❌: CPU 23% quá thấp (Google recommend 35% để tối ưu chi phí và scale). Phân bố: US-central + Europe (không liên quan APAC/US) + Asia-east1 → Không cover đầy đủ US/APAC, thêm Europe thừa gây latency cao và chi phí replication không cần thiết.

  • Phương án SAI: Maintain a target of 23% CPU utilization by locating:
    cluster-a in zone us-central1-a
    cluster-b in zone us-central1-b
    cluster-c in zone us-east1-a
    Giải thích sai ❌: CPU 23% sai chuẩn (under-utilization, lãng phí tài nguyên). Tất cả clusters ở US chỉ (us-central1-a/b cùng region, us-east1-a region khác nhưng vẫn US) → Không có APAC, không hỗ trợ expansion, vi phạm HA cho multi-region traffic mới.

  • Phương án SAI: Maintain a target of 35% CPU utilization by locating:
    cluster-a in zone us-central1-a
    cluster-b in zone australia-southeast1-a
    cluster-c in zone europe-west1-d
    cluster-d in zone asia-east1-b
    Giải thích sai ❌: CPU 35% đúng, nhưng phân bố không phù hợp: australia-southeast1 (Sydney, APAC xa xôi) + europe-west1-d (Europe thừa) + asia-east1 + us-central1 → Phủ sóng quá rộng (US/Europe/APAC/Australia), tăng latency/replication cost, không tập trung US-APAC. Google ưu tiên regions gần user (Tokyo/Taiwan cho APAC).

  • Phương án ĐÚNG (như đã phân tích ở trên) ✅: Maintain a target of 35% CPU utilization by locating:
    cluster-a in zone us-central1-a
    cluster-b in zone us-central2-a
    cluster-c in zone asia-northeast1-b
    cluster-d in zone asia-east1-b
    Giải thích đúng ✅: Hoàn hảo cho US (multi-region) + APAC (hai vùng gần), CPU chuẩn, 4 clusters đảm bảo HA theo docs Google.

🛠️ Lời khuyên triển khai: Sử dụng Bigtable Instance với replication enabled, monitor qua Cloud Monitoring để giữ 35% CPU. Test failover để xác nhận HA!

Câu 14
Your ecommerce website captures user clickstream data to analyze customer traffic patterns in real time and support personalization features on your website. You plan to analyze this data using big data tools. You need a low-latency solution that can store 8 TB of data and can scale to millions of read and write requests per second. What should you do?
  1. A Write your data into Bigtable and use Dataproc and the Apache Hbase libraries for analysis.
  2. B Deploy a Cloud SQL environment with read replicas for improved performance. Use Datastream to export data to Cloud Storage and analyze with Dataproc and the Cloud Storage connector.
  3. C Use Memorystore to handle your low-latency requirements and for real-time analytics.
  4. D Stream your data into BigQuery and use Dataproc and the BigQuery Storage API to analyze large volumes of data.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong hệ thống thương mại điện tử (ecommerce website) nơi dữ liệu clickstream (dữ liệu theo dõi hành vi click của người dùng) được thu thập thời gian thực (real-time) để phân tích mẫu giao thông khách hàng (customer traffic patterns) và hỗ trợ tính năng cá nhân hóa (personalization features).

📊 Yêu cầu chính của giải pháp:

  • Phân tích dữ liệu bằng công cụ big data (big data tools).
  • Độ trễ thấp (low-latency) cho việc lưu trữ và xử lý.
  • Dung lượng lưu trữ ít nhất 8 TB.
  • Mở rộng quy mô đến hàng triệu yêu cầu đọc/ghi mỗi giây (millions of read and write requests per second).

🛠️ Bối cảnh kỹ thuật: Đây là dữ liệu high-velocity (tốc độ cao), unstructured/semi-structured (không cấu trúc hoặc bán cấu trúc), phù hợp với NoSQL database có khả năng scale horizontally. Giải pháp cần kết hợp lưu trữ và phân tích big data, cập nhật theo Google Cloud phiên bản mới nhất năm 2026 (Bigtable hỗ trợ codeless autoscaling, single-row transactions lên đến 2024+).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Write your data into Bigtable and use Dataproc and the Apache Hbase libraries for analysis.

Lý do chi tiết:

  • Bigtable là dịch vụ NoSQL wide-column store của Google Cloud, được thiết kế chuyên biệt cho low-latency và high-throughput workloads như clickstream data. Nó hỗ trợ hàng triệu reads/writes per second (đã chứng minh với hàng PB dữ liệu, ví dụ: YouTube, Google Ads), dễ dàng scale đến 8 TB+ mà không downtime.
  • Kết hợp Dataproc (managed Hadoop/Spark) với Apache HBase libraries cho phép phân tích big data real-time trực tiếp trên Bigtable (HBase là client cho Bigtable).
  • Hoàn hảo cho time-series data như clickstream, với TTL (time-to-live) tự động và cQL queries nhanh. Không vi phạm giới hạn ingestion như các dịch vụ khác.

📘 Nguồn tham khảo:

🧩 Giải thích tất cả các phương án (đúng và sai)

  • ✅ Write your data into Bigtable and use Dataproc and the Apache Hbase libraries for analysis.
    Đúng vì: Như phân tích trên, Bigtable lý tưởng cho low-latency high-scale (millions r/w/s), lưu trữ 8 TB+, và tích hợp seamless với Dataproc/HBase cho big data analytics real-time. Đáp ứng đầy đủ yêu cầu mà không cần ETL phức tạp. 🏆

  • ❌ Deploy a Cloud SQL environment with read replicas for improved performance. Use Datastream to export data to Cloud Storage and analyze with Dataproc and the Cloud Storage connector.
    Sai vì: Cloud SQL là RDBMS (MySQL/PostgreSQL), chỉ scale đến ~hàng nghìn QPS (không phải millions), và read replicas chỉ cải thiện reads, không xử lý writes cao. Datastream + Cloud Storage + Dataproc tạo pipeline phức tạp, độ trễ cao (không real-time), không phù hợp big data unstructured. Phí cao cho 8 TB writes liên tục. 🚫

  • ❌ Use Memorystore to handle your low-latency requirements and for real-time analytics.
    Sai vì: Memorystore (Redis/Memcached) là in-memory caching, chỉ tạm thời (không persistent cho 8 TB), giới hạn ~hàng trăm GB/node và không scale đến millions persistent writes/s. Không hỗ trợ big data analytics (chỉ key-value đơn giản), dữ liệu mất khi restart. Không thay thế database chính. ⚠️

  • ❌ Stream your data into BigQuery and use Dataproc and the BigQuery Storage API to analyze large volumes of data.
    Sai vì: BigQuery là data warehouse cho batch/analytical queries, streaming ingestion giới hạn ~1 MB/s/table (không millions writes/s low-latency). Độ trễ writes cao (seconds), không phù hợp operational real-time như clickstream (BigQuery tốt hơn cho reporting). Dataproc + Storage API chỉ cho export, không real-time personalization. ⏳

Kết luận tổng quát 🎯: Bigtable là lựa chọn tối ưu cho workload này theo best practices Google Cloud 2026, đảm bảo scale, low-latency và cost-effective cho big data real-time. Các phương án sai thiếu khả năng scale hoặc thêm độ trễ không cần thiết.

Câu 15 Chọn nhiều đáp án
Your company uses Cloud Spanner for a mission-critical inventory management system that is globally available. You recently loaded stock keeping unit (SKU) and product catalog data from a company acquisition and observed hotspots in the Cloud Spanner database. You want to follow Google-recommended schema design practices to avoid performance degradation. What should you do? (Choose two.)
  1. A Use an auto-incrementing value as the primary key.
  2. B Normalize the data model.
  3. C Promote low-cardinality attributes in multi-attribute primary keys.
  4. D Promote high-cardinality attributes in multi-attribute primary keys.
  5. E Use bit-reverse sequential value as the primary key.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào Cloud Spanner (dịch vụ cơ sở dữ liệu phân tán toàn cầu của Google Cloud), được sử dụng cho hệ thống quản lý hàng tồn kho quan trọng (mission-critical inventory management system). Sau khi tải dữ liệu SKU (stock keeping unit) và catalog sản phẩm từ việc mua lại công ty, xuất hiện hotspots (điểm nóng - nơi tập trung quá nhiều truy vấn ghi/đọc, gây nghẽn hiệu suất).
Mục tiêu: Áp dụng thực hành thiết kế schema tốt nhất theo khuyến nghị của Google để tránh suy giảm hiệu suất. Câu hỏi yêu cầu chọn hai phương án đúng (Choose two).
Hotspots thường xảy ra do khóa chính (primary key) có tính tuần tự thấp hoặc cardinality thấp, dẫn đến phân bố không đều trên các node Spanner. Giải pháp cần ưu tiên phân bố đều workload (sharding hiệu quả).
📘 Kiến thức cập nhật: Dựa trên tài liệu Google Cloud Spanner mới nhất (2024-2026), xem Schema Design Best Practices và Avoiding Hotspots.

✅ Đáp án đúng và lý do lựa chọn

Hai đáp án đúng là (chọn cả hai để giải quyết hotspots toàn diện):

  • Promote high-cardinality attributes in multi-attribute primary keys.
  • Use bit-reverse sequential value as the primary key.

Lý do chọn:
🛠️ Promote high-cardinality attributes: Trong primary key đa thuộc tính (multi-attribute PK), đặt thuộc tính có cardinality cao (nhiều giá trị duy nhất, như CustomerID hoặc Region) lên đầu tiên giúp phân bố đều dữ liệu trên các split/node, tránh hotspots. Đây là nguyên tắc cốt lõi của Spanner.
🛠️ Use bit-reverse sequential value: Với khóa tuần tự (như ID tăng dần), sử dụng bit-reverse (đảo bit) để "xáo trộn" giá trị, đảm bảo ghi phân bố đều thay vì tập trung vào cuối range (hotspot). Google khuyến nghị cho trường hợp SKU/product ID sequential.
Kết hợp hai cách này tuân thủ Google-recommended practices, cải thiện throughput lên đến hàng nghìn QPS mà không cần refactor lớn.

📋 Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai kèm lý do cụ thể dựa trên best practices Spanner:

  • Use an auto-incrementing value as the primary key.
    ❌ Sai. Giá trị tự tăng (auto-incrementing) tạo khóa tuần tự, tất cả ghi mới tập trung vào cuối range → hotspots nghiêm trọng trên node cuối cùng. Spanner khuyên tránh hoàn toàn; thay bằng UUID hoặc bit-reverse. (Xem Sequential Keys Issue).

  • Normalize the data model.
    ❌ Sai. Chuẩn hóa (normalize) tạo nhiều bảng/join, tăng latency do Spanner không hỗ trợ join hiệu quả như RDBMS truyền thống. Khuyến nghị: Denormalize để giảm join, ưu tiên read/write nhanh cho mission-critical app. Normalization phù hợp SQL OLTP nhưng không phải Spanner global scale.

  • Promote low-cardinality attributes in multi-attribute primary keys.
    ❌ Sai. Đặt thuộc tính low-cardinality (ít giá trị duy nhất, như Gender: Male/Female) lên đầu PK gây nhóm lớn dữ liệu cùng giá trị → hotspots (ví dụ: tất cả "Male" vào một split). Phải ưu tiên high-cardinality để sharding đều.

  • Promote high-cardinality attributes in multi-attribute primary keys.
    ✅ Đúng. Như giải thích trên, đặt high-cardinality (ví dụ: TenantID + UserID) đầu PK đảm bảo phân bố đều, tăng scalability. Ví dụ thực tế: SKU table với (Region_HighCard + SKU_Prefix). (Tài liệu: Primary Key Design).

  • Use bit-reverse sequential value as the primary key.
    ✅ Đúng. Bit-reverse "lật ngược" bit của số sequential (ví dụ: 0001 → 1000), phân tán ghi đều trên 2^N ranges. Lý tưởng cho inventory SKU tăng dần. Google cung cấp hàm BIT_REVERSE() trong Spanner SQL. (Tài liệu: Bit-Reversed Counters).

Tóm tắt khuyến nghị thực hiện 🚀:

  • Kiểm tra schema hiện tại bằng gcloud spanner databases describe.
  • Refactor PK: Ví dụ PRIMARY KEY(Region, BitReverse(SKU_ID)).
  • Test với YCSB benchmark để xác nhận giảm hotspots.
    Nếu cần code mẫu, tham khảo Spanner Samples GitHub.
Câu 16
You are managing multiple applications connecting to a database on Cloud SQL for PostgreSQL. You need to be able to monitor database performance to easily identify applications with long-running and resource-intensive queries. What should you do?
  1. A Use log messages produced by Cloud SQL.
  2. B Use Query Insights for Cloud SQL.
  3. C Use the Cloud Monitoring dashboard with available metrics from Cloud SQL.
  4. D Use Cloud SQL instance monitoring in the Google Cloud Console.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào tình huống quản lý nhiều ứng dụng (applications) kết nối đến cơ sở dữ liệu Cloud SQL for PostgreSQL trên Google Cloud Platform (GCP). Yêu cầu chính là giám sát hiệu suất cơ sở dữ liệu (database performance) để dễ dàng xác định các ứng dụng nào đang chạy các truy vấn (queries) kéo dài thời gian (long-running) và tiêu tốn nhiều tài nguyên (resource-intensive).

📌 Mục tiêu cụ thể: Không chỉ monitor tổng quát mà cần phân tích chi tiết theo ứng dụng, giúp phát hiện bottleneck từ các query "nặng" của từng app riêng lẻ. Đây là nhu cầu phổ biến trong môi trường production với multi-tenant apps, đòi hỏi công cụ chuyên sâu về query analysis thay vì monitor cơ bản.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Query Insights for Cloud SQL.

🛠️ Lý do chi tiết:

  • Query Insights là tính năng chuyên biệt của Cloud SQL (cập nhật mới nhất đến 2026), cho phép phân tích query performance theo thời gian thực và lịch sử, bao gồm top queries chậm nhất, resource usage (CPU, I/O, memory), và wait events.
  • Đặc biệt, nó nhóm query theo client application (dựa trên connection metadata như application_name trong PostgreSQL), giúp dễ dàng identify app nào gây ra long-running hoặc resource-intensive queries.
  • Tích hợp trực tiếp với Cloud SQL, không cần config phức tạp, và hỗ trợ PostgreSQL đầy đủ (bao gồm EXPLAIN plans, query text, execution stats).
  • Đây là giải pháp tối ưu nhất theo best practices GCP cho database observability.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một cách rõ ràng:

  • ❌ Use log messages produced by Cloud SQL.
    Phương án này sai vì logs của Cloud SQL (như query logs hoặc slow query logs) chỉ ghi nhận dữ liệu thô, không có phân tích tự động. Bạn phải manually parse logs để tìm long-running queries, không dễ dàng group theo application, và thiếu visualization cho resource-intensive patterns. Không phù hợp cho multi-app monitoring quy mô lớn.

  • ✅ Use Query Insights for Cloud SQL.
    Phương án này đúng như đã giải thích ở trên. Đây là công cụ chính xác và mạnh mẽ nhất, với dashboard trực quan, filtering theo app/client, và insights sâu về query performance. Hỗ trợ PostgreSQL từ Cloud SQL Enterprise Plus edition (cập nhật 2024-2026).

  • ❌ Use the Cloud Monitoring dashboard with available metrics from Cloud SQL.
    Phương án này sai vì Cloud Monitoring chỉ cung cấp metrics tổng quát như CPU usage, connections, throughput, không drill-down vào individual queries hoặc per-application breakdown. Bạn không thể identify cụ thể query nào long-running hay app nào gây tải cao – chỉ thấy aggregate data.

  • ❌ Use Cloud SQL instance monitoring in the Google Cloud Console.
    Phương án này sai vì monitoring trong Console chỉ hiển thị overview instance-level (CPU, memory, disk I/O, connections), thiếu query-level details và không hỗ trợ phân tích theo application. Đây là monitor cơ bản, không đủ cho yêu cầu identify resource-intensive queries từ apps cụ thể.

📘 Tài liệu tham khảo (cập nhật mới nhất đến 2026)

  • Query Insights docs: Cloud SQL Query Insights – Chi tiết về features cho PostgreSQL.
  • Best practices monitoring: Monitor Cloud SQL performance.
  • Release notes 2024-2026: Query Insights được enhance với AI-powered insights và better app attribution (xem GCP What's New).
  • Certification guide: Google Cloud Professional Cloud Database Engineer study guide (bao gồm Query Insights như key feature cho observability).

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần thêm ví dụ thực tế, hãy hỏi nhé!

Câu 17
You are building an application that allows users to customize their website and mobile experiences. The application will capture user information and preferences. User profiles have a dynamic schema, and users can add or delete information from their profile. You need to ensure that user changes automatically trigger updates to your downstream BigQuery data warehouse. What should you do?
  1. A Store your data in Bigtable, and use the user identifier as the key. Use one column family to store user profile data, and use another column family to store user preferences.
  2. B Use Cloud SQL, and create different tables for user profile data and user preferences from your recommendations model. Use SQL to join the user profile data and preferences
  3. C Use Firestore in Native mode, and store user profile data as a document. Update the user profile with preferences specific to that user and use the user identifier to query.
  4. D Use Firestore in Datastore mode, and store user profile data as a document. Update the user profile with preferences specific to that user and use the user identifier to query.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một ứng dụng cho phép người dùng tùy chỉnh trải nghiệm website và mobile, lưu trữ thông tin hồ sơ người dùng (user profiles) và sở thích (preferences). Đặc điểm quan trọng:

  • Schema động (dynamic schema): Người dùng có thể thêm hoặc xóa thông tin linh hoạt, không cố định cấu trúc dữ liệu.
  • Tự động đồng bộ thay đổi: Mọi cập nhật từ người dùng phải tự động kích hoạt (trigger) cập nhật đến kho dữ liệu BigQuery downstream (dùng để phân tích).

📌 Yêu cầu chính: Chọn giải pháp database hỗ trợ schema linh hoạt (NoSQL/document-based), tích hợp tự động với BigQuery qua triggers hoặc streams (như Cloud Functions, change streams). Kiến thức cập nhật đến 2026: Firestore (Native mode) hỗ trợ Firestore Change Data Capture (CDC) và export streams trực tiếp đến BigQuery qua Dataflow/Cloud Functions (phiên bản mới nhất GCP 2024-2026).

✅ Đáp án đúng: Use Firestore in Native mode, and store user profile data as a document. Update the user profile with preferences specific to that user and use the user identifier to query.

Lý do lựa chọn:

  • Firestore Native mode là database document NoSQL lý tưởng cho schema động: Lưu profile + preferences vào một document duy nhất (dùng user ID làm ID document), dễ thêm/xóa fields mà không cần schema cố định.
  • Tự động trigger BigQuery: Firestore Native hỗ trợ onWrite/onUpdate triggers qua Cloud Functions, hoặc change streams export trực tiếp đến BigQuery (qua BigQuery Data Transfer Service hoặc Dataflow). Thay đổi profile → trigger ngay lập tức sync dữ liệu.
  • Hiệu suất cao cho query theo user ID (index tự động). ✅ Hoàn hảo khớp yêu cầu!

📋 Giải thích chi tiết tất cả các phương án

  • ❌ [SAI] Store your data in Bigtable, and use the user identifier as the key. Use one column family to store user profile data, and use another column family to store user preferences.
    Bigtable là wide-column NoSQL phù hợp dữ liệu lớn, phân tích thời gian thực (time-series), nhưng không hỗ trợ schema động dễ dàng (cần định nghĩa column families trước, khó add/delete fields linh hoạt). Không có built-in triggers tự động đến BigQuery (phải dùng Pub/Sub thủ công). ❌ Không phù hợp dynamic profiles và auto-sync.

  • ❌ [SAI] Use Cloud SQL, and create different tables for user profile data and user preferences from your recommendations model. Use SQL to join the user profile data and preferences.
    Cloud SQL (MySQL/PostgreSQL) là RDBMS yêu cầu schema cố định (tables phải define columns trước), không linh hoạt cho dynamic add/delete fields → phải alter table thường xuyên. Join tables phức tạp, không tự động trigger BigQuery (cần ETL thủ công). ❌ Không đáp ứng schema động và real-time sync.

  • ✅ [ĐÚNG] Use Firestore in Native mode, and store user profile data as a document. Update the user profile with preferences specific to that user and use the user identifier to query.
    Như giải thích ở phần đáp án đúng: Document model linh hoạt, Native mode đầy đủ features real-time + triggers (Cloud Functions/Dataflow) sync thẳng BigQuery. Query nhanh bằng user ID. ✅ Tối ưu nhất!

  • ❌ [SAI] Use Firestore in Datastore mode, and store user profile data as a document. Update the user profile with preferences specific to that user and use the user identifier to query.
    Datastore mode là legacy (không phải Firestore thực thụ), thiếu real-time listeners, offline support, và change streams đầy đủ như Native mode. Sync BigQuery kém hiệu quả hơn (chỉ export batch, không real-time triggers mượt). Schema động OK nhưng không khuyến nghị (GCP push Native từ 2020+). ❌ Gần đúng nhưng thiếu features auto-trigger mạnh mẽ.

📘 Tài liệu tham khảo (GCP cập nhật 2026)

Câu 18 Chọn nhiều đáp án
Your application uses Cloud SQL for MySQL. Your users run reports on data that relies on near-real time; however, the additional analytics caused excessive load on the primary database. You created a read replica for the analytics workloads, but now your users are complaining about the lag in data changes and that their reports are still slow. You need to improve the report performance and shorten the lag in data replication without making changes to the current reports. Which two approaches should you implement? (Choose two.)
  1. A Create secondary indexes on the replica.
  2. B Create additional read replicas, and partition your analytics users to use different read replicas.
  3. C Disable replication on the read replica, and set the flag for parallel replication on the read replica. Re-enable replication and optimize performance by setting flags on the primary instance.
  4. D Disable replication on the primary instance, and set the flag for parallel replication on the primary instance. Re-enable replication and optimize performance by setting flags on the read replica.
  5. E Move your analytics workloads to BigQuery, and set up a streaming pipeline to move data and update BigQuery.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này thuộc chủ đề Cloud SQL for MySQL trên Google Cloud Platform (GCP) (không phải AWS như đề cập nhầm, vì Cloud SQL là dịch vụ managed database của GCP). Tình huống: Ứng dụng sử dụng Cloud SQL for MySQL làm primary database. Người dùng chạy báo cáo (reports) cần dữ liệu near-real time (gần thời gian thực), nhưng workload analytics gây tải nặng lên primary DB. Bạn đã tạo read replica để offload analytics, tuy nhiên:

  • Lag replication (độ trễ sao chép dữ liệu) làm báo cáo thiếu dữ liệu mới.
  • Performance báo cáo vẫn chậm do tải trên replica.

Yêu cầu giải quyết: Cải thiện hiệu suất báo cáo (report performance) và giảm lag replication, mà không thay đổi code báo cáo hiện tại (không modify reports). Chọn 2 approaches phù hợp.

📘 Kiến thức cập nhật (phiên bản mới nhất GCP đến 2026): Cloud SQL hỗ trợ read replicas với semi-sync replication (mặc định), lag thường <1 giây nhưng tăng với workload cao. Có thể scale horizontally bằng nhiều replicas, enable parallel replication (flags như slave_parallel_workers >0 và slave_parallel_type=LOGICAL_CLOCK) trên replica để giảm lag. Không hỗ trợ indexes tùy chỉnh trên replica (read-only cho DDL).

✅ Đáp án đúng (Chọn 2)

  • Create additional read replicas, and partition your analytics users to use different read replicas.
  • Disable replication on the read replica, and set the flag for parallel replication on the read replica. Re-enable replication and optimize performance by setting flags on the primary instance.

Lý do chọn:
🛠️ Hai cách này trực tiếp giải quyết vấn đề không thay đổi reports:

  • Tạo thêm read replicas + partition users: Scale out horizontally, phân tải queries analytics lên nhiều replicas → cải thiện performance reports ngay lập tức (read replicas hỗ trợ connection pooling/load balancing qua proxy hoặc driver).
  • Disable/enable replication trên replica + parallel replication flag: Giảm lag replication bằng cách enable parallel apply trên replica (MySQL 5.7+ hỗ trợ, Cloud SQL config qua flags). Tối ưu flags trên primary (như binlog_row_image=minimal) giúp replica catch-up nhanh hơn. Không ảnh hưởng reports.

📋 Giải thích tất cả các phương án (Đúng/Sai)

  • Create secondary indexes on the replica.
    ❌ Sai: Read replicas trong Cloud SQL là read-only cho DDL (không thể tạo secondary indexes mới). Indexes chỉ replicate từ primary, việc cố tạo sẽ fail. Không giải quyết lag replication, chỉ có thể cải thiện query nếu indexes tồn tại từ primary → không phù hợp.

  • Create additional read replicas, and partition your analytics users to use different read replicas.
    ✅ Đúng: Scale out bằng nhiều replicas (Cloud SQL hỗ trợ up to 10 replicas/instance), dùng proxy hoặc app logic partition users → giảm tải/replica, cải thiện report performance. Không tăng lag (replicas inherit config primary), phù hợp near-real time.

  • Disable replication on the read replica, and set the flag for parallel replication on the read replica. Re-enable replication and optimize performance by setting flags on the primary instance.
    ✅ Đúng: Quy trình chuẩn GCP: Stop replication (gcloud sql instances patch --replica-stop), set flags như slave_parallel_workers=4/slave_parallel_type=LOGICAL_CLOCK trên replica → parallel apply logs giảm lag đáng kể (từ giây xuống ms). Flags primary (e.g., sync_binlog=1) hỗ trợ. Re-enable (gcloud sql instances replica-start) → an toàn, không downtime.

  • Disable replication on the primary instance, and set the flag for parallel replication on the primary instance. Re-enable replication and optimize performance by setting flags on the read replica.
    ❌ Sai: Primary không hỗ trợ disable replication (là source, chỉ config flags như log_bin). Parallel replication chỉ áp dụng trên slave/replica, không trên primary. Thao tác này sẽ break toàn bộ replication chain → gây downtime và mất dữ liệu.

  • Move your analytics workloads to BigQuery, and set up a streaming pipeline to move data and update BigQuery.
    ❌ Sai: Yêu cầu thay đổi reports (chuyển queries từ MySQL sang BigQuery SQL) → vi phạm "without making changes to the current reports". Dù Datastream/Cloud SQL CDC có thể stream real-time, nhưng setup phức tạp và không giữ nguyên queries.

📚 Tài liệu tham khảo (GCP Docs mới nhất 2026)

Hy vọng phân tích giúp bạn ôn thi hiệu quả! 🚀

Câu 19
You are evaluating Cloud SQL for PostgreSQL as a possible destination for your on-premises PostgreSQL instances. Geography is becoming increasingly relevant to customer privacy worldwide. Your solution must support data residency requirements and include a strategy to: configure where data is stored control where the encryption keys are stored govern the access to data
What should you do?
  1. A Replicate Cloud SQL databases across different zones.
  2. B Create a Cloud SQL for PostgreSQL instance on Google Cloud for the data that does not need to adhere to data residency requirements. Keep the data that must adhere to data residency requirements on-premises. Make application changes to support both databases.
  3. C Allow application access to data only if the users are in the same region as the Google Cloud region for the Cloud SQL for PostgreSQL database.
  4. D Use features like customer-managed encryption keys (CMEK), VPC Service Controls, and Identity and Access Management (IAM) policies.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc đánh giá Cloud SQL for PostgreSQL (dịch vụ cơ sở dữ liệu quản lý PostgreSQL trên Google Cloud) như một điểm đến tiềm năng để di chuyển các instance PostgreSQL từ on-premises. Yếu tố cốt lõi là tuân thủ yêu cầu data residency (quy định về vị trí lưu trữ dữ liệu để bảo vệ quyền riêng tư khách hàng trên toàn cầu). Giải pháp phải hỗ trợ chiến lược toàn diện bao gồm:

  • Cấu hình nơi lưu trữ dữ liệu (chọn region phù hợp để data residency).
  • Kiểm soát nơi lưu trữ khóa mã hóa (encryption keys).
  • Quản lý truy cập dữ liệu (govern access).

📘 Bối cảnh cập nhật đến 2026: Cloud SQL for PostgreSQL hỗ trợ multi-region deployment, Customer-Managed Encryption Keys (CMEK) qua Cloud KMS (với lựa chọn location cho keys), VPC Service Controls để ngăn chặn data exfiltration, và IAM policies để kiểm soát truy cập chi tiết. Những tính năng này được cập nhật liên tục, ví dụ phiên bản Cloud SQL Enterprise Plus (2024-2026) tăng cường hỗ trợ data residency cho GDPR, HIPAA.

Nguồn tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use features like customer-managed encryption keys (CMEK), VPC Service Controls, and Identity and Access Management (IAM) policies.

Lý do 🛠️:

  • CMEK cho phép khách hàng kiểm soát vị trí lưu trữ khóa mã hóa (qua Cloud KMS multi-region), đảm bảo keys residency phù hợp với data.
  • VPC Service Controls bảo vệ dữ liệu bằng cách tạo service perimeter, ngăn chặn truy cập trái phép và data exfiltration, hỗ trợ data residency.
  • IAM policies quản lý truy cập dữ liệu tinh tế (least privilege, conditional access dựa trên location).
  • Kết hợp với chọn region cho Cloud SQL instance, giải pháp này bao quát toàn bộ yêu cầu: lưu trữ data, keys, và govern access. Đây là chiến lược khuyến nghị chính thức từ Google Cloud cho data residency.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ Phương án SAI: Replicate Cloud SQL databases across different zones.
    Giải thích: Replication across zones (trong cùng region) chỉ tăng high availability và disaster recovery nội bộ region, không giải quyết data residency (yêu cầu cross-region hoặc specific geography). Không kiểm soát keys hay access đầy đủ. (Zones ≠ Regions).

  • ❌ Phương án SAI: Create a Cloud SQL for PostgreSQL instance on Google Cloud for the data that does not need to adhere to data residency requirements. Keep the data that must adhere to data residency requirements on-premises. Make application changes to support both databases.
    Giải thích: Phương án này không migrate đầy đủ data nhạy cảm lên cloud, vẫn giữ on-premises (rủi ro vận hành cao, không tận dụng Cloud SQL). Phải thay đổi ứng dụng phức tạp (dual-database), không đáp ứng cấu hình lưu trữ data/keys/access thống nhất trên Google Cloud.

  • ❌ Phương án SAI: Allow application access to data only if the users are in the same region as the Google Cloud region for the Cloud SQL for PostgreSQL database.
    Giải thích: Chỉ kiểm soát access dựa trên user location (có thể dùng IAM conditions), nhưng bỏ qua lưu trữ data và keys (không cấu hình residency). Không đủ toàn diện, dễ bị bypass qua proxy/VPN, và không dùng native features như CMEK/VPC SC.

  • ✅ Phương án ĐÚNG: Use features like customer-managed encryption keys (CMEK), VPC Service Controls, and Identity and Access Management (IAM) policies.
    Giải thích: Như đã nêu ở phần đáp án đúng, đây là bộ tính năng native hoàn hảo cho data residency toàn diện trên Cloud SQL: chọn region cho data, CMEK cho keys, VPC SC + IAM cho access. Tuân thủ best practices Google Cloud đến 2026.

Câu 20
Your customer is running a MySQL database on-premises with read replicas. The nightly incremental backups are expensive and add maintenance overhead. You want to follow Google-recommended practices to migrate the database to Google Cloud, and you need to ensure minimal downtime. What should you do?
  1. A Create a Google Kubernetes Engine (GKE) cluster, install MySQL on the cluster, and then import the dump file.
  2. B Use the mysqldump utility to take a backup of the existing on-premises database, and then import it into Cloud SQL.
  3. C Create a Compute Engine VM, install MySQL on the VM, and then import the dump file.
  4. D Create an external replica, and use Cloud SQL to synchronize the data to the replica.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống: Khách hàng đang chạy cơ sở dữ liệu MySQL trên máy chủ tại chỗ (on-premises) với các read replicas. Các bản sao lưu tăng dần (incremental backups) hàng đêm đắt đỏ và tốn công bảo trì. Bạn cần tuân thủ các thực hành khuyến nghị của Google để di chuyển (migrate) cơ sở dữ liệu sang Google Cloud, đồng thời đảm bảo thời gian ngừng hoạt động (downtime) tối thiểu.

Mục tiêu chính:

  • Migrate MySQL từ on-premises sang Cloud SQL (dịch vụ managed database của Google Cloud).
  • Giảm chi phí và overhead từ backups thủ công.
  • Minimal downtime: Nghĩa là sử dụng phương pháp replication liên tục (continuous sync) thay vì dump/import toàn bộ, tránh gián đoạn lớn.
  • Theo best practices của Google: Sử dụng Cloud SQL với Database Migration Service (DMS) hoặc external replication để sync dữ liệu real-time.

🛠️ Kiến thức cập nhật (tính đến 2026): Google Cloud khuyến nghị sử dụng Cloud SQL for MySQL kết hợp external replicas hoặc Database Migration Service (DMS) cho migration MySQL on-premises. DMS hỗ trợ continuous replication cho MySQL từ on-prem sang Cloud SQL, cho phép promote replica với downtime chỉ vài giây. (Không liên quan AWS như đề cập, mà là Google Cloud thuần túy).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create an external replica, and use Cloud SQL to synchronize the data to the replica.

Lý do:

  • Phương pháp này tuân thủ khuyến nghị của Google cho migration MySQL on-premises với minimal downtime. Bạn thiết lập Cloud SQL instance làm external replica của primary on-premises, sử dụng binary logs để sync dữ liệu liên tục (continuous replication).
  • Sau khi sync hoàn tất (lag thấp), promote Cloud SQL replica thành primary với downtime chỉ vài giây (switchover).
  • Giảm overhead: Không cần nightly backups thủ công, vì replication tự động.
  • Hỗ trợ read replicas on-prem: External replica giữ nguyên cấu trúc replication.
  • DMS tích hợp sẵn hỗ trợ quy trình này từ 2023, cập nhật 2025 với improved GTID support cho MySQL 8+.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Create a Google Kubernetes Engine (GKE) cluster, install MySQL on the cluster, and then import the dump file.
    Giải thích sai: Không phải best practice của Google cho managed database. GKE dùng cho containerized apps, không khuyến khích install MySQL thủ công (self-managed, thiếu HA, backups auto). Dump file gây downtime cao (full export/import lớn), không sync liên tục, vi phạm minimal downtime. Google ưu tiên Cloud SQL thay vì GKE cho DB production.

  • ❌ [SAI] Use the mysqldump utility to take a backup of the existing on-premises database, and then import it into Cloud SQL.
    Giải thích sai: Mysqldump chỉ tạo logical backup (dump file), phù hợp small DB nhưng gây downtime lớn (lock table khi dump, import chậm). Không hỗ trợ incremental sync từ read replicas on-prem, vẫn cần nightly backups. Google không recommend cho large DB hoặc minimal downtime; DMS/external replica tốt hơn.

  • ❌ [SAI] Create a Compute Engine VM, install MySQL on the VM, and then import the dump file.
    Giải thích sai: Tạo self-managed MySQL trên VM tốn công bảo trì (patching, backups, scaling), trái với Google-recommended (dùng managed Cloud SQL). Dump file lại gây downtime cao, không tận dụng replication on-prem. VM phù hợp dev/test, không production migration.

  • ✅ [ĐÚNG] Create an external replica, and use Cloud SQL to synchronize the data to the replica.
    Giải thích đúng: Như đã nêu ở phần đáp án. Đây là Google-recommended practice cho zero/minimal downtime migration: External replica sync binary log real-time, hỗ trợ read replicas on-prem, promote nhanh chóng. Hoàn hảo cho scenario này! 🚀