Ngân hàng đề — Google Cloud Professional Data Engineer
Tìm thấy 429 câu.
- A Schedule a daily export of the table to a Cloud Storage dual or multi-region bucket.
- B Schedule a daily copy of the dataset to a backup region.
- C Schedule a daily BigQuery snapshot of the table.
- D Modify ETL job to load the data into both the current and another backup region.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi này xoay quanh việc bảo vệ dữ liệu trong BigQuery (Google Cloud) đối với một dataset multi-region chứa bảng daily sales volumes (khối lượng bán hàng hàng ngày). Bảng này được cập nhật nhiều lần mỗi ngày. Yêu cầu chính là:
- Bảo vệ chống lại sự cố khu vực (regional failures), chẳng hạn như mất dữ liệu ở một vùng địa lý.
- Recovery Point Objective (RPO) < 24 giờ: Nghĩa là thời gian mất dữ liệu tối đa phải dưới 24 giờ (tức là dữ liệu mới nhất có thể khôi phục phải trong vòng dưới 1 ngày).
- Giảm thiểu chi phí ở mức tối thiểu.
Bối cảnh kỹ thuật (dựa trên kiến thức GCP cập nhật đến 2026):
- BigQuery multi-region dataset đã có tính năng automatic replication giữa các vùng trong cùng multi-region location (ví dụ: US hoặc EU), giúp chịu lỗi ở mức độ cao. Tuy nhiên, để bảo vệ toàn diện chống regional failures (mất toàn bộ multi-region), cần cơ chế backup riêng biệt.
- Bảng cập nhật thường xuyên → Không thể sao chép realtime (đắt đỏ), mà cần giải pháp periodic backup với tần suất hàng ngày để đạt RPO <24h.
- Ưu tiên chi phí thấp: Tránh duplicate data realtime hoặc sao chép full dataset.
Mục tiêu: Chọn giải pháp backup bền vững, rẻ tiền, đạt RPO bằng cách export dữ liệu ra lưu trữ ngoài an toàn cao.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Schedule a daily export of the table to a Cloud Storage dual or multi-region bucket.
Lý do chi tiết 🛠️:
- Export hàng ngày bảng ra Cloud Storage dual-region hoặc multi-region bucket (có độ bền 99.999999999% - 11 9's, replicate tự động qua nhiều vùng) đảm bảo dữ liệu an toàn chống regional failures.
- RPO <24h: Export 1 lần/ngày → Mất tối đa dữ liệu chưa export (dưới 24h).
- Chi phí tối thiểu 💰: Export BigQuery chỉ tính phí scan dữ liệu (rẻ), lưu trữ CS multi-region rẻ hơn so với duplicate BigQuery storage/processing. Có thể khôi phục bằng cách load lại từ CS vào BigQuery mới.
- Phù hợp với bảng cập nhật thường xuyên: Không cần realtime sync.
- Đây là best practice cho BigQuery backup theo tài liệu GCP (xem nguồn bên dưới).
📋 Giải thích tất cả các phương án (đúng và sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm lý do bằng tiếng Việt rõ ràng:
-
✅ Schedule a daily export of the table to a Cloud Storage dual or multi-region bucket.
Lý do đúng 🏆: Như đã giải thích ở trên, đây là giải pháp tối ưu nhất. Export định kỳ (cron job qua Cloud Scheduler + Data Transfer Service hoặc BigQuery scheduled queries) rẻ, linh hoạt, và CS multi-region bucket chịu lỗi khu vực tốt (replicate qua ít nhất 2 vùng). RPO đạt <24h mà không duplicate storage BigQuery. -
❌ Schedule a daily copy of the dataset to a backup region.
Lý do sai 🚫: BigQuery không hỗ trợ "copy dataset" trực tiếp đến "backup region" một cách rẻ tiền. Copy dataset (quabq cphoặc API) sẽ duplicate toàn bộ storage và metadata, dẫn đến chi phí gấp đôi (storage + query fees). Multi-region dataset đã replicate nội bộ, copy thêm không hiệu quả cho backup và không giảm chi phí. Không đạt yêu cầu "minimum costs". -
❌ Schedule a daily BigQuery snapshot of the table.
Lý do sai 📸: BigQuery snapshots (time-travel queries hoặc table snapshots từ 2023) chỉ là point-in-time copy trong cùng dataset, hỗ trợ query lịch sử (lên đến 7 ngày miễn phí). Chúng không bảo vệ chống regional failures vì vẫn nằm trong cùng multi-region location gốc. Nếu region fail toàn bộ, snapshot cũng mất. Chi phí snapshot thấp nhưng không giải quyết vấn đề cross-region durability. -
❌ Modify ETL job to load the data into both the current and another backup region.
Lý do sai 🔄: Dual-write ETL (sửa job load dữ liệu vào 2 regions) đạt RPO thấp (gần realtime), nhưng tăng chi phí gấp đôi (storage, ingestion, query ở 2 nơi). Với bảng cập nhật nhiều lần/ngày, overhead ETL lớn, không "minimum costs". BigQuery multi-region đã replicate tự động, dual-region thêm là thừa và đắt (vi phạm yêu cầu).
📘 Tài liệu tham khảo (cập nhật mới nhất GCP đến 2026)
- BigQuery Backup Best Practices: Exporting data from BigQuery & Multi-region datasets.
- Cloud Storage Durability: Dual/multi-region buckets – Độ bền 11 9's.
- RPO cho BigQuery: Disaster Recovery khuyến nghị export to CS cho low RPO/low cost.
- Scheduled Exports: Sử dụng Cloud Scheduler + BigQuery EXPORT DATA SQL.
Giải pháp này đảm bảo tuân thủ nguyên tắc GCP Professional Data Engineer! 🚀 Nếu cần ví dụ code/script export, hãy hỏi thêm nhé!
- A Determine whether your Dataflow pipeline has a custom network tag set.
- B Determine whether there is a firewall rule set to allow traffic on TCP ports 12345 and 12346 for the Dataflow network tag.
- C Determine whether there is a firewall rule set to allow traffic on TCP ports 12345 and 12346 on the subnet used by Dataflow workers.
- D Determine whether your Dataflow pipeline is deployed with the external IP address option enabled.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc khắc phục sự cố (troubleshooting) cho pipeline Dataflow trên Google Cloud, xử lý dữ liệu từ Cloud Storage sang BigQuery. Vấn đề cụ thể là các worker nodes của Dataflow không thể giao tiếp (communicate) với nhau. Đội ngũ networking sử dụng Google Cloud network tags để định nghĩa firewall rules. Yêu cầu là xác định nguyên nhân theo các thực hành bảo mật mạng được Google khuyến nghị (Google-recommended networking security practices).
🔍 Bối cảnh kỹ thuật: Dataflow sử dụng các worker VM trên Compute Engine, cần giao tiếp nội bộ qua các port cụ thể (như shuffle service cho dữ liệu lớn). Firewall rules phải cho phép traffic giữa các worker, ưu tiên sử dụng network tags thay vì quy tắc dựa trên subnet hoặc IP để đảm bảo bảo mật và scalability (theo docs Google Cloud mới nhất 2024-2026).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Determine whether there is a firewall rule set to allow traffic on TCP ports 12345 and 12346 for the Dataflow network tag.
Lý do 🛠️:
- Các port TCP 12345 và 12346 là ports chuẩn mà Dataflow worker nodes sử dụng để giao tiếp nội bộ (worker-to-worker communication), đặc biệt cho shuffle service trong các job streaming/batch lớn (xác nhận từ tài liệu Dataflow networking 2024+).
- Google khuyến nghị mạnh mẽ sử dụng network tag mặc định "dataflow-worker" (hoặc custom tag nếu set) cho firewall rules, thay vì quy tắc dựa trên subnet/IP để tránh mở rộng không cần thiết và tuân thủ least-privilege principle.
- Kiểm tra rule này sẽ xác định chính xác nếu thiếu allow traffic nội bộ → gây lỗi communication.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá với lý do cụ thể dựa trên best practices Dataflow networking (cập nhật đến 2026):
-
❌ [SAI] Determine whether your Dataflow pipeline has a custom network tag set.
Giải thích: Phương án này chỉ kiểm tra custom network tag (tag tùy chỉnh), nhưng Dataflow mặc định sử dụng tag "dataflow-worker" mà không cần custom. Vấn đề communication thường do thiếu firewall rule cho tag đó, chứ không phải việc có custom tag hay không. Kiểm tra custom tag không giải quyết root cause và không theo recommended practice (Google ưu tiên kiểm tra rule trước). -
✅ [ĐÚNG] Determine whether there is a firewall rule set to allow traffic on TCP ports 12345 and 12346 for the Dataflow network tag.
Giải thích: Như đã nêu ở phần đáp án đúng. Đây là bước troubleshooting chuẩn từ Google: Tạo VPC firewall rule với source/destination tag "dataflow-worker", allow TCP 12345-12346 (và các port khác như 443, 8080 nếu cần). Thiếu rule này gây lỗi worker communication ngay lập tức. -
❌ [SAI] Determine whether there is a firewall rule set to allow traffic on TCP ports 12345 and 12346 on the subnet used by Dataflow workers.
Giải thích: Sử dụng subnet-based rule (quy tắc dựa trên subnet) không được khuyến nghị vì mở rộng traffic cho toàn subnet → rủi ro bảo mật cao (vi phạm least-privilege). Google yêu cầu dùng network tags cho Dataflow để isolate traffic worker. Subnet rule chỉ dùng cho legacy setup, không phải best practice 2024+. -
❌ [SAI] Determine whether your Dataflow pipeline is deployed with the external IP address option enabled.
Giải thích: External IP liên quan đến traffic ra ngoài (egress đến GCS/BigQuery), không ảnh hưởng đến internal worker-to-worker communication (luôn dùng internal IPs). Dataflow mặc định dùng no external IP cho security, và vấn đề ở đây là internal networking → không liên quan.
📘 Tài liệu tham khảo (cập nhật mới nhất 2024-2026)
- Google Cloud Dataflow Networking Requirements: Dataflow worker ports and firewall rules – Chi tiết ports 12345/12346 cho shuffle.
- Troubleshooting Dataflow: Dataflow troubleshooting networking.
- Firewall best practices: VPC Firewall Rules for Dataflow – Nhấn mạnh network tags.
- Dataflow Security: Recommended network config.
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm ví dụ code Terraform cho firewall rule, hãy hỏi nhé!
SELECT date, order, status FROM customer_order
WHERE country = '' AND username = ''
How should you redesign the BigQuery table to support faster access?
- A Cluster the table by country and username fields.
- B Cluster the table by country field, and partition by username field.
- C Partition the table by country and username fields.
- D Partition the table by _PARTITIONTIME.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào vấn đề tối ưu hóa hiệu suất truy vấn trên bảng customer_order trong BigQuery (Google Cloud), với kích thước khổng lồ 10 PB và 10 triệu khách hàng. Bảng lưu lịch sử đơn hàng, và dashboard hỗ trợ cần hiển thị nhanh lịch sử đơn hàng khi áp dụng hai bộ lọc string: country_name và username.
Query hiện tại chậm vì scan toàn bộ bảng (full table scan) khi filter:
SELECT date, order, status FROM customer_order
WHERE country = '<country_name>' AND username = '<username>'
Mục tiêu: Redesign bảng để truy vấn nhanh hơn bằng cách sử dụng các tính năng partitioning hoặc clustering của BigQuery, giúp prune dữ liệu không cần thiết (data skipping), giảm lượng dữ liệu scan từ 10 PB xuống chỉ phần liên quan.
BigQuery (phiên bản mới nhất đến 2026) hỗ trợ:
- Partitioning: Chia bảng theo cột thời gian (DATE/TIMESTAMP/DATETIME/INTEGER), tối ưu cho filter thời gian.
- Clustering: Sắp xếp dữ liệu trong partitions theo tối đa 4 cột (string/int/...), lý tưởng cho filter trên non-time columns như country/username. Clustering tự động reorganize khi insert/update.
📘 Tài liệu tham khảo:
- BigQuery Partitioned Tables (cập nhật 2024-2026).
- BigQuery Clustered Tables (hỗ trợ multi-column clustering lên đến 4 fields).
✅ Đáp án đúng: Cluster the table by country and username fields.
Lý do chọn đáp án đúng 🛠️:
- Hai filter
countryvàusernameđều là string, không phải kiểu thời gian → không thể partition trực tiếp (partition chỉ hỗ trợ DATE/TIMESTAMP/...). - Clustering theo country và username (multi-level clustering) sẽ tổ chức dữ liệu physically theo thứ tự này, cho phép automatic data skipping khi query filter cả hai → giảm scan từ 10 PB xuống chỉ % nhỏ dữ liệu liên quan.
- Hiệu suất dashboard cải thiện gấp 10-100x cho filter equality trên clustered columns (theo benchmark BigQuery 2024+).
- Phù hợp bảng lớn không có time column rõ ràng trong query.
📋 Giải thích tất cả các phương án
-
✅ Cluster the table by country and username fields.
🟢 Đúng vì clustering lý tưởng cho string filters như country/username. BigQuery cluster theo hierarchical order (country trước, username sau), prune blocks không match → query siêu nhanh. Không cần partition trước, clustering độc lập hiệu quả cho bảng 10 PB. -
❌ Cluster the table by country field, and partition by username field.
🔴 Sai vì không thể partition bằng username (string) – BigQuery chỉ partition theo DATE/TIMESTAMP/INTEGER (pseudo-column như _PARTITIONTIME). Username string vi phạm rule, tạo lỗi khi setup. -
❌ Partition the table by country and username fields.
🔴 Sai vì partitioning không hỗ trợ multi-column string. Chỉ 1 partition column kiểu thời gian được phép; country/username string → lỗi syntax. Partition multi-level chỉ cho ingestion time/time-unit, không phải string. -
❌ Partition the table by _PARTITIONTIME.
🔴 Sai vì query không filter theo thời gian (chỉ country/username). _PARTITIONTIME (daily ingestion) chỉ prune nếu WHERE có date filter → full scan vẫn xảy ra, dashboard vẫn chậm. Không giải quyết gốc rễ filter string.
Kết luận 🚀: Clustering là giải pháp tối ưu nhất cho case này, kết hợp với query best practices (như materialized views nếu cần). Test trên BigQuery console để verify speedup!
- A Create a Standard Tier Memorystore for Redis instance in the development environment. Initiate a manual failover by using the limited-data-loss data protection mode.
- B Create a Standard Tier Memorystore for Redis instance in a development environment. Initiate a manual failover by using the force-data-loss data protection mode.
- C Increase one replica to Redis instance in production environment. Initiate a manual failover by using the force-data-loss data protection mode.
- D Initiate a manual failover by using the limited-data-loss data protection mode to the Memorystore for Redis instance in the production environment.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm
📖 Nội dung câu hỏi:
Câu hỏi tập trung vào việc mô phỏng (simulate) tình huống failover cho một instance Memorystore for Redis Standard Tier đang chạy trong môi trường production (sản xuất). Mục tiêu là tạo ra kịch bản disaster recovery (khôi phục thảm họa) chính xác nhất, đồng thời đảm bảo không ảnh hưởng đến dữ liệu production.
Memorystore for Redis là dịch vụ managed Redis của Google Cloud, với Standard Tier hỗ trợ high availability (HA) qua các replica tự động. Failover có thể được kích hoạt thủ công với hai chế độ bảo vệ dữ liệu:
- Limited-data-loss: Ưu tiên giảm thiểu mất dữ liệu bằng cách chờ replicas sync trước khi failover (an toàn hơn nhưng chậm).
- Force-data-loss: Buộc failover ngay lập tức, chấp nhận rủi ro mất dữ liệu chưa sync (mô phỏng tình huống khẩn cấp xấu nhất).
Vấn đề chính: Không được tác động production data, nên cần môi trường riêng biệt để test. Đây là kịch bản thực tế trong Google Cloud Professional Data Engineer để đảm bảo DR plan không làm gián đoạn dịch vụ live. (Kiến thức cập nhật đến 2026: Không thay đổi lớn về failover modes trong Memorystore docs).
✅ Đáp án đúng:
Create a Standard Tier Memorystore for Redis instance in a development environment. Initiate a manual failover by using the force-data-loss data protection mode.
Lý do lựa chọn:
🛠️ Phương án này chính xác nhất vì:
- Tạo instance mới trong dev environment → Không ảnh hưởng production data (an toàn 100%).
- Sử dụng force-data-loss mode → Mô phỏng failover DR chính xác nhất, tái hiện tình huống khẩn cấp thực tế (mất dữ liệu có thể xảy ra, replicas chưa sync), giúp test toàn diện khả năng khôi phục mà không cần chờ đợi như limited-data-loss.
- Standard Tier đảm bảo HA giống production, tạo môi trường test giống hệt (realistic simulation).
📘 Tài liệu tham khảo:
- Google Cloud Memorystore for Redis: Manual failover (cập nhật 2024-2026, xác nhận hai modes).
- Best practices for DR testing → Khuyến nghị test ở non-prod env.
🔍 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] Create a Standard Tier Memorystore for Redis instance in the development environment. Initiate a manual failover by using the limited-data-loss data protection mode.
Phương án này gần đúng nhưng không chính xác nhất vì dùng limited-data-loss mode chỉ mô phỏng failover "an toàn" (chờ sync, mất ít data), không tái hiện DR tình huống xấu nhất (như primary crash đột ngột). Không đủ "accurate" cho simulation toàn diện. -
✅ [ĐÚNG] Create a Standard Tier Memorystore for Redis instance in a development environment. Initiate a manual failover by using the force-data-loss data protection mode.
(Như đã giải thích ở trên: ✅ Môi trường dev an toàn + force-data-loss mô phỏng DR thực tế nhất). -
❌ [SAI] Increase one replica to Redis instance in production environment. Initiate a manual failover by using the force-data-loss data protection mode.
Rủi ro cao: Thay đổi replica trực tiếp trên production (tăng replica rồi failover) có thể gây gián đoạn hoặc mất data thật, vi phạm yêu cầu "no impact on production data". Không nên chỉnh sửa env live để simulate. -
❌ [SAI] Initiate a manual failover by using the limited-data-loss data protection mode to the Memorystore for Redis instance in the production environment.
Nguy hiểm nhất: Failover trực tiếp trên production với bất kỳ mode nào cũng ảnh hưởng data live (dù limited-data-loss an toàn hơn nhưng vẫn có downtime). Không simulate được mà còn phá hủy env sản xuất!
🎯 Kết luận: Phương án đúng ưu tiên test isolated (dev env) + worst-case scenario (force-data-loss) để DR plan robust. Luôn test non-prod trước! 🚀
- A Provide the partner organization a copy of your CMEKs to decrypt the data.
- B Export the tables to parquet files to a Cloud Storage bucket and grant the storageinsights.viewer role on the bucket to the partner organization.
- C Copy the tables you need to share to a dataset without CMEKs. Create an Analytics Hub listing for this dataset.
- D Create an authorized view that contains the CMEK to decrypt the data when accessed.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc quản trị một dataset BigQuery sử dụng customer-managed encryption key (CMEK) – một loại khóa mã hóa do khách hàng tự quản lý thông qua Cloud KMS. Vấn đề chính là cần chia sẻ dataset này với một tổ chức đối tác không có quyền truy cập vào CMEK của bạn.
✅ Mục tiêu: Đảm bảo đối tác có thể truy cập dữ liệu mà không cần chia sẻ khóa mã hóa (để tránh rủi ro bảo mật), đồng thời tuân thủ các tính năng chia sẻ an toàn của Google Cloud. BigQuery với CMEK yêu cầu người truy cập phải có quyền sử dụng chính xác cùng một CMEK để giải mã dữ liệu; nếu không, họ sẽ không đọc được dữ liệu. Giải pháp phải xử lý việc "thoát khỏi" CMEK để chia sẻ mà vẫn giữ nguyên tính toàn vẹn dữ liệu.
🛠️ Bối cảnh kỹ thuật (cập nhật đến 2026): Theo tài liệu BigQuery mới nhất (phiên bản 2024-2025), CMEK được quản lý qua Cloud KMS, và chia sẻ dataset CMEK không hỗ trợ trực tiếp với bên ngoài. Analytics Hub (ra mắt từ 2021, cập nhật 2025 với bảo mật nâng cao) là công cụ lý tưởng để liệt kê và chia sẻ dataset mà không cần grant quyền trực tiếp.
📘 Tài liệu tham khảo:
- BigQuery Customer-managed encryption keys
- Sharing BigQuery data using Analytics Hub
- BigQuery authorized views limitations with CMEK
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Copy the tables you need to share to a dataset without CMEKs. Create an Analytics Hub listing for this dataset.
Lý do chi tiết:
- Khi copy table từ dataset CMEK sang dataset mới không sử dụng CMEK (mặc định dùng Google-managed encryption keys – CMEK của Google), dữ liệu sẽ được giải mã tạm thời trong quá trình copy và lưu trữ lại với khóa mã hóa mặc định. Dataset mới này có thể được chia sẻ an toàn qua Analytics Hub – một dịch vụ chuyên chia sẻ dữ liệu BigQuery với bên ngoài mà không cần grant quyền IAM trực tiếp, hỗ trợ subscription model và kiểm soát truy cập chi tiết (cập nhật 2025 với linked listings).
- Phương án này ✅ an toàn, tuân thủ zero-trust, không chia sẻ khóa, và hiệu quả cho dữ liệu lớn mà không cần export thủ công.
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ Provide the partner organization a copy of your CMEKs to decrypt the data.
Sai vì: Việc chia sẻ CMEK (khóa do bạn quản lý) vi phạm nguyên tắc bảo mật cốt lõi của Cloud KMS – khóa không được copy hoặc chia sẻ ra ngoài. Đối tác cần quyền IAM chính xác trên cùng KMS key để decrypt, dẫn đến rủi ro lộ khóa và không được AWS/GCP khuyến nghị (tương tự IAM best practices). Thay vào đó, phải dùng cơ chế chia sẻ không lộ khóa. -
❌ Export the tables to parquet files to a Cloud Storage bucket and grant the storageinsights.viewer role on the bucket to the partner organization.
Sai vì:- Export table CMEK sang GCS yêu cầu quyền decrypt trước (bạn phải export thủ công), nhưng parquet files vẫn cần xử lý mã hóa riêng nếu apply CMEK trên bucket.
- Role storageinsights.viewer không tồn tại cho GCS (đó là role cho Storage Insights trong Operations Suite); role đúng cho đọc bucket là storage.objectViewer. Phương án này phức tạp, tốn kém (chi phí export/storage), và không tận dụng native sharing của BigQuery.
-
✅ Copy the tables you need to share to a dataset without CMEKs. Create an Analytics Hub listing for this dataset.
Đúng vì: Như giải thích ở trên. Copy table tự động re-encrypt với khóa mặc định (GOOGLE-managed), sau đó Analytics Hub cho phép tạo listing để đối tác subscribe và query dữ liệu trực tiếp mà không cần truy cập dataset gốc. Hỗ trợ selective sharing (chỉ table cần thiết), kiểm soát thời hạn, và audit logs đầy đủ (cập nhật 2025). -
❌ Create an authorized view that contains the CMEK to decrypt the data when accessed.
Sai vì: Authorized view (AV) kế thừa encryption settings từ base table, nên AV vẫn yêu cầu quyền CMEK để truy cập. Không có cơ chế "chứa CMEK" trong AV để tự decrypt – AV chỉ là logical view, không thay đổi mã hóa vật lý. Tài liệu BigQuery xác nhận hạn chế này với CMEK.
- A Set up VPC Network Peering between Project A and Project B. Add a firewall rule to allow the peered subnet range to access all instances on the network.
- B Turn off the external IP addresses on the Dataflow worker. Enable Cloud NAT in Project A.
- C Add the external IP addresses of the Dataflow worker as authorized networks in the Cloud SQL instance.
- D Set up VPC Network Peering between Project A and Project B. Create a Compute Engine instance without external IP address in Project B on the peered subnet to serve as a proxy server to the Cloud SQL database.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi
Câu hỏi mô tả tình huống phát triển một pipeline Apache Beam sử dụng JdbcIO để trích xuất dữ liệu từ instance Cloud SQL (không có public IP address). Pipeline được triển khai và chạy trên Dataflow thuộc Project A, trong khi instance Cloud SQL nằm ở Project B. Sau khi deploy, pipeline thất bại do lỗi kết nối. Không sử dụng VPC Service Controls hay Shared VPC. Yêu cầu giải quyết lỗi mà đảm bảo dữ liệu không đi qua public internet (tức là kết nối private network).
Vấn đề cốt lõi 📌:
- Dataflow workers (trong Project A) cần truy cập private IP của Cloud SQL (Project B) qua mạng nội bộ Google Cloud.
- Cross-project private access yêu cầu cấu hình mạng đặc biệt, vì các project mặc định không giao tiếp private trực tiếp.
- Giải pháp phải an toàn, không expose ra internet, và phù hợp với Dataflow (hỗ trợ private IP ranges).
✅ Đáp án đúng
Set up VPC Network Peering between Project A and Project B. Add a firewall rule to allow the peered subnet range to access all instances on the network.
Lý do lựa chọn 🏆:
- VPC Network Peering cho phép hai VPC (Project A và B) kết nối private trực tiếp mà không qua internet, hỗ trợ traffic cross-project. Dataflow workers (sử dụng IP ranges từ subnet của Project A) có thể truy cập private IP của Cloud SQL qua peering.
- Thêm firewall rule ở Project B để cho phép subnet peered (từ Project A) truy cập port 3306 (MySQL) hoặc tương ứng của Cloud SQL. Đây là cách đơn giản, hiệu quả nhất theo best practices Google Cloud (không cần proxy hay NAT).
- Đảm bảo dữ liệu private end-to-end, phù hợp với JdbcIO trên Beam/Dataflow.
🛠️ Giải thích tất cả các phương án
-
Set up VPC Network Peering between Project A and Project B. Add a firewall rule to allow the peered subnet range to access all instances on the network.
✅ Đúng: Như phân tích trên, peering + firewall rule giải quyết chính xác private cross-project access cho Dataflow đến Cloud SQL private IP. Không phức tạp, chi phí thấp, và được khuyến nghị chính thức. -
Turn off the external IP addresses on the Dataflow worker. Enable Cloud NAT in Project A.
❌ Sai: Tắt external IP trên Dataflow workers (qua--disable-public-ips) và dùng Cloud NAT chỉ hỗ trợ outbound traffic từ private subnet ra internet. Cloud SQL ở Project B có private IP (không public), nên NAT không giúp kết nối inbound private cross-project. Dataflow vẫn không reach được mà không có peering. -
Add the external IP addresses of the Dataflow worker as authorized networks in the Cloud SQL instance.
❌ Sai: Cloud SQL không có public IP, nên "authorized networks" chỉ áp dụng cho public IP access (TCP 3306 qua internet). Dataflow workers có dynamic external IPs (không fixed), nên không khả thi và vi phạm yêu cầu "không qua public internet". -
Set up VPC Network Peering between Project A and Project B. Create a Compute Engine instance without external IP address in Project B on the peered subnet to serve as a proxy server to the Cloud SQL database.
❌ Sai: Phức tạp hóa không cần thiết! VPC Peering đã cho phép Dataflow truy cập trực tiếp Cloud SQL private IP (với firewall rule). Tạo proxy Compute Engine thêm overhead (chi phí, quản lý, latency), không phải giải pháp tối ưu theo docs Google Cloud.
📘 Tài liệu tham khảo (cập nhật đến 2026)
- VPC Network Peering docs (Using VPC Network Peering - Cross-project).
- Cloud SQL Private IP access via peering.
- Dataflow connect to Cloud SQL privately (Configure private IP connectivity).
- Best practices: Google Cloud Architecture Center - Private connectivity for Dataflow (2025 updates hỗ trợ Beam 2.58+ với JdbcIO private).
Giải pháp này đảm bảo tuân thủ nguyên tắc least privilege và zero-trust networking mới nhất! 🚀
- A Create two separate authorized datasets; one for the data analytics team and another for the consumer support team.
- B Ensure that the data analytics team members do not have the Data Catalog Fine-Grained Reader role for the policy tags.
- C Replace the authorized dataset with an authorized view. Use row-level security and apply filter_expression to limit data access.
- D Remove the bigquery.dataViewer role from the data analytics team on the authorized datasets.
- E Enforce access control in the policy tag taxonomy.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi này thuộc lĩnh vực Google Cloud BigQuery, tập trung vào việc bảo mật dữ liệu nhạy cảm trong bảng BigQuery chứa thông tin khách hàng (tên, địa chỉ). Bạn cần chia sẻ dữ liệu an toàn với hai nhóm:
- Nhóm Data Analytics: Truy cập dữ liệu tất cả khách hàng, nhưng KHÔNG được truy cập cột nhạy cảm (sensitive columns).
- Nhóm Consumer Support: Truy cập tất cả cột dữ liệu, nhưng CHỈ khách hàng còn hợp đồng hoạt động (active contracts).
Giải pháp đã áp dụng: Sử dụng authorized dataset (để chia sẻ dữ liệu cross-project an toàn) và policy tags (để kiểm soát truy cập cột-level). Tuy nhiên, nhóm Data Analytics vẫn truy cập được cột nhạy cảm. Nhiệm vụ: Chọn 2 hành động để khắc phục, đảm bảo nhóm này không truy cập dữ liệu bị hạn chế.
Vấn đề cốt lõi 🛡️:
- Authorized dataset cho phép truy cập dữ liệu gốc mà không lộ quyền trực tiếp.
- Policy tags gắn vào cột nhạy cảm để enforce column-level security, nhưng quyền Data Catalog Fine-Grained Reader trên policy tags có thể cho phép bypass hạn chế.
- Cần cấu hình IAM chính xác trên policy tag taxonomy (cấu trúc phân cấp policy tags) để kiểm soát ai đọc được tag nào.
Kiến thức cập nhật (BigQuery phiên bản mới nhất 2026): Policy tags hỗ trợ fine-grained access control qua IAM trên Data Catalog, tích hợp chặt chẽ với authorized datasets/views. Không có thay đổi lớn từ 2024-2026 về cơ chế này (xem docs GCP).
✅ Đáp án đúng (Chọn 2)
Các đáp án đúng là:
- Ensure that the data analytics team members do not have the Data Catalog Fine-Grained Reader role for the policy tags.
- Enforce access control in the policy tag taxonomy.
Lý do lựa chọn 📝:
- Hai hành động này trực tiếp khắc phục vấn đề column-level access với policy tags. Nhóm Data Analytics có thể có quyền Fine-Grained Reader (cho phép đọc policy tags hạn chế), dẫn đến truy cập cột nhạy cảm. Việc loại bỏ quyền này và enforce IAM trên taxonomy (deny quyền đọc tag cho nhóm này) sẽ block truy cập hoàn toàn, mà vẫn cho phép đọc dữ liệu không nhạy cảm. Điều này phù hợp với yêu cầu: Họ đọc tất cả rows, chỉ block columns.
🔍 Giải thích chi tiết tất cả các phương án
-
❌ Create two separate authorized datasets; one for the data analytics team and another for the consumer support team.
Sai vì: Tạo authorized dataset riêng không giải quyết vấn đề column-level security. Authorized datasets chỉ proxy quyền đọc dataset gốc, nhưng policy tags vẫn apply trên tất cả datasets. Nhóm analytics vẫn truy cập cột nhạy cảm nếu họ có quyền đọc policy tags. Phương án này chỉ hữu ích cho row-filtering (như support team), không fix issue gốc. -
✅ Ensure that the data analytics team members do not have the Data Catalog Fine-Grained Reader role for the policy tags.
Đúng vì: Quyền Data Catalog Fine-Grained Reader trên policy tags cho phép user đọc cột có tag hạn chế. Nếu nhóm analytics có quyền này, họ bypass policy enforcement dù authorized dataset đã set. Loại bỏ quyền sẽ block truy cập cột nhạy cảm ngay lập tức, đúng yêu cầu "không access sensitive data". -
❌ Replace the authorized dataset with an authorized view. Use row-level security and apply filter_expression to limit data access.
Sai vì: Authorized view với row-level security (RLS via filter_expression) chỉ lọc rows (dòng dữ liệu), không block columns (cột nhạy cảm). Yêu cầu cho analytics là block columns cho tất cả rows, không phải row-filter. Policy tags đã phù hợp hơn cho column-level, thay thế sẽ không fix và phức tạp hóa (RLS dùng cho support team thôi). -
❌ Remove the bigquery.dataViewer role from the data analytics team on the authorized datasets.
Sai vì: Loại bỏ bigquery.dataViewer sẽ block toàn bộ truy cập dataset (bao gồm dữ liệu không nhạy cảm), vi phạm yêu cầu "access data of all customers". Authorized dataset cần quyền này để đọc cơ bản; vấn đề nằm ở policy tags, không phải dataset role. -
✅ Enforce access control in the policy tag taxonomy.
Đúng vì: Policy tag taxonomy (cấu trúc phân cấp tags trong Data Catalog) cần IAM policies cụ thể (allow/deny đọc tag cho groups). Hiện tại, thiếu enforce dẫn đến analytics đọc được tag sensitive. Áp dụng IAM trên taxonomy sẽ restrict chính xác, tích hợp với authorized dataset mà không ảnh hưởng quyền đọc rows không nhạy cảm.
📘 Tài liệu tham khảo
- BigQuery Policy Tags Documentation (cập nhật 2026: Fine-grained IAM trên taxonomy).
- Authorized Datasets/Views (tích hợp policy tags).
- Data Catalog IAM Roles (Fine-Grained Reader role chi tiết).
- Sample:
gcloud datacatalog tags policies set-iam-policyđể enforce taxonomy.
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần ví dụ code, hỏi thêm nhé!
- A Enable zonal high availability on the primary instance. Create a new read replica in a new region.
- B Create a cascading read replica from the existing read replica in Region3.
- C Create two new read replicas from the new primary instance, one in Region3 and one in a new region.
- D Create a new read replica in Region1, promote the new read replica to be the primary instance, and enable zonal high availability.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm
📖 Nội dung câu hỏi:
Câu hỏi mô tả một tình huống khôi phục thảm họa (Disaster Recovery - DR) trong Cloud SQL for PostgreSQL (dịch vụ cơ sở dữ liệu quản lý của Google Cloud). Cụ thể:
- Bạn có một primary instance ở Region1 (khu vực 1).
- Có hai read replica (bản sao chỉ đọc): một ở Region2 và một ở Region3.
- Xảy ra sự cố bất ngờ ở Region1, buộc phải promote (nâng cấp) read replica ở Region2 thành primary instance mới.
- Yêu cầu chính: Trước khi chuyển kết nối ứng dụng (switch over connections), phải đảm bảo ứng dụng có cùng dung lượng cơ sở dữ liệu (database capacity) như trước.
- Trước sự cố: Tổng cộng 3 instances (1 primary + 2 replicas), cung cấp khả năng đọc/ghi và mở rộng đọc.
- Mục tiêu: Sau DR, vẫn duy trì 3 instances để tránh giảm hiệu suất ứng dụng.
Câu hỏi kiểm tra kiến thức về quản lý replication trong Cloud SQL PostgreSQL, đặc biệt là hành vi sau khi promote replica và cách tái tạo topology replication để khôi phục capacity (dựa trên tài liệu GCP cập nhật đến 2026, replication logic không thay đổi cơ bản từ phiên bản 2023+).
✅ Đáp án đúng:
Create two new read replicas from the new primary instance, one in Region3 and one in a new region.
🛠️ Lý do chọn đáp án đúng:
- Sau khi promote read replica ở Region2 thành primary mới, read replica cũ ở Region3 sẽ ngừng replicate (vì nó chỉ kết nối với primary cũ ở Region1 đã fail). Do đó, cần tạo mới 2 read replicas trực tiếp từ primary mới ở Region2 để khôi phục topology:
- Một replica ở Region3 (thay thế replica cũ).
- Một replica ở region mới (thay thế primary cũ ở Region1, tăng tính sẵn sàng đa vùng).
- Điều này đảm bảo tổng 3 instances (1 primary + 2 replicas), giữ nguyên capacity đọc/ghi cho ứng dụng trước khi switch connections.
- Hợp lý cho DR cross-region, tránh single point of failure.
📘 Giải thích tất cả các phương án (A, B, C, D)
-
Enable zonal high availability on the primary instance. Create a new read replica in a new region.
❌ Sai. Phương án này không khả thi vì primary instance ở Region1 đã gặp sự cố bất ngờ (fail), không thể enable zonal HA (HA trong cùng zone) trên instance đã hỏng. Hơn nữa, chỉ tạo 1 replica mới ở region mới sẽ không khôi phục đủ 2 replicas (chỉ có primary Region2 + 1 replica = 2 instances, giảm capacity). Không giải quyết vấn đề promote replica Region2. -
Create a cascading read replica from the existing read replica in Region3.
❌ Sai. Cascading replica (replica chuỗi) từ replica Region3 chỉ tạo thêm 1 replica phụ thuộc vào replica Region3, nhưng replica Region3 đang stale (dữ liệu cũ) sau khi primary Region1 fail. Không đảm bảo sync dữ liệu mới từ primary Region2, và tổng instances vẫn thiếu (không đạt 3 instances độc lập). Không khôi phục capacity đúng cách cho DR. -
Create two new read replicas from the new primary instance, one in Region3 and one in a new region.
✅ Đúng. Như giải thích ở trên: Tạo 2 replicas mới trực tiếp từ primary Region2, thay thế replica Region3 cũ và thêm replica ở region mới (ví dụ Region4). Đảm bảo full sync dữ liệu, topology cân bằng đa vùng, và giữ nguyên 3 instances trước switch connections. Hoàn hảo cho DR. -
Create a new read replica in Region1, promote the new read replica to be the primary instance, and enable zonal high availability.
❌ Sai. Region1 đang có sự cố, tạo replica mới ở đây rủi ro cao (có thể fail luôn). Promote replica mới thành primary sẽ mất thời gian dài (crash recovery), không nhanh như promote replica Region2 sẵn có. Enable zonal HA chỉ trong zone, không giải quyết cross-region DR, và không khôi phục đủ 2 replicas khác.
📚 Tài liệu tham khảo (cập nhật GCP đến 2026)
- Cloud SQL PostgreSQL Replication Docs: cloud.google.com/sql/docs/postgres/replica-promotion – Chi tiết về promote replica và recreate replicas sau failover.
- High Availability & DR Guide: cloud.google.com/sql/docs/postgres/high-availability – Topology cross-region và capacity management.
- Release Notes 2024-2026: Không thay đổi core logic promote/replica creation (xác nhận qua GCP console & API v1beta4+).
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀
- A Assign a function with notification logic to the on_retry_callback parameter for the operator responsible for the task at risk.
- B Configure a Cloud Monitoring alert on the sla_missed metric associated with the task at risk to trigger a notification.
- C Assign a function with notification logic to the on_failure_callback parameter tor the operator responsible for the task at risk.
- D Assign a function with notification logic to the sla_miss_callback parameter for the operator responsible for the task at risk.
Xem giải thích
🧩 Giải thích chi tiết nội dung câu hỏi
Câu hỏi này thuộc chủ đề Google Cloud Composer (dịch vụ quản lý Apache Airflow trên Google Cloud Platform - GCP), không phải AWS như mô tả ban đầu (có thể là nhầm lẫn). Nội dung xoay quanh việc lập lịch và điều phối ETL pipelines bằng Cloud Composer. Cụ thể:
- Bạn đang sử dụng DAG (Directed Acyclic Graph) trong Apache Airflow để định nghĩa quy trình ETL.
- Một task trong DAG phụ thuộc vào dịch vụ bên thứ ba (third-party service), có nguy cơ thất bại cao.
- Yêu cầu chính: Thông báo (notify) ngay khi task không thành công (does not succeed), tức là khi task thất bại (fail).
Mục tiêu là chọn cách xử lý callback phù hợp trong Airflow operator để gửi thông báo kịp thời. Đây là kiến thức cốt lõi của Airflow phiên bản mới nhất (hỗ trợ đến Airflow 2.10+ trong Cloud Composer 3.x năm 2026), tập trung vào task lifecycle callbacks để tùy chỉnh hành vi khi task fail, retry hoặc miss SLA.
📘 Tài liệu tham khảo:
- Apache Airflow Docs - Task Callbacks (cập nhật 2024-2026).
- Cloud Composer Docs - Handling Task Failures (phiên bản mới nhất 2026).
✅ Đáp án đúng
Assign a function with notification logic to the on_failure_callback parameter for the operator responsible for the task at risk.
Lý do chọn đáp án này 🛠️:
- Trong Apache Airflow, on_failure_callback là hàm callback chạy ngay lập tức khi task thất bại hoàn toàn (sau khi hết số lần retry nếu có). Đây là cách chuẩn và trực tiếp nhất để gửi thông báo (notification) khi task "does not succeed".
- Bạn chỉ cần gán một Python function chứa logic gửi email/Slack/Pub/Sub vào tham số này của operator (ví dụ: BashOperator, PythonOperator).
- Phù hợp hoàn hảo với yêu cầu: task phụ thuộc third-party dễ fail → notify ngay khi fail.
❌ Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên hành vi callback của Airflow (không thay đổi đến 2026).
-
[SAI] Assign a function with notification logic to the on_retry_callback parameter for the operator responsible for the task at risk.
❌ Sai vì:on_retry_callbackchỉ chạy khi task bắt đầu retry (nếu retries > 0), không phải khi task fail hoàn toàn. Nếu task fail mà không retry (retries=0), callback này không chạy. Không đảm bảo notify khi "does not succeed" cuối cùng. -
[SAI] Configure a Cloud Monitoring alert on the sla_missed metric associated with the task at risk to trigger a notification.
❌ Sai vì:sla_missedmetric (trong Cloud Monitoring) chỉ kích hoạt khi task vượt quá thời gian SLA (Service Level Agreement) đã định nghĩa cho DAG/task. Đây là notify về trễ hạn, không phải thất bại (failure). Task có thể fail mà không miss SLA (nếu fail sớm), nên không phù hợp. -
[ĐÚNG] Assign a function with notification logic to the on_failure_callback parameter for the operator responsible for the task at risk.
✅ Đúng như đã giải thích ở trên: Callback chính xác cho failure, chạy sau khi task fail vĩnh viễn, hỗ trợ notify qua email, Slack, hoặc tích hợp GCP services như Cloud Functions/Pub/Sub. -
[SAI] Assign a function with notification logic to the sla_miss_callback parameter for the operator responsible for the task at risk.
❌ Sai vì:sla_miss_callbackchỉ chạy khi task miss SLA (trễ hơn thời gian dự kiến). Tương tựsla_missedmetric, nó không liên quan đến failure, mà chỉ về độ trễ. Task fail nhanh vẫn không trigger callback này.
🛠️ Lời khuyên thực hành
- Ví dụ code Airflow:
def notify_failure(context): print("Task failed! Sending alert...") # Logic gửi Slack/Email task = BashOperator( task_id='risky_task', bash_command='call_third_party', on_failure_callback=notify_failure # ✅ Gán ở đây ) - Kết hợp với Cloud Composer 3.x (2026): Hỗ trợ triggerer và dynamic task mapping để scale notify cho nhiều task.
- Test bằng
airflow tasks testđể verify callback! 🚀
- A Update your existing on-premises ETL tool to write to BigQuery by using the BigQuery Open Database Connectivity (ODBC) driver. Set up the proxy parameter in the simba.googlebigqueryodbc.ini file to point to your data center’s NAT gateway.
- B Use Datastream to replicate data from your on-premises MySQL database to BigQuery. Set up Cloud Interconnect between your on-premises data center and Google Cloud. Use Private connectivity as the connectivity method and allocate an IP address range within your VPC network to the Datastream connectivity configuration. Use Server-only as the encryption type when setting up the connection profile in Datastream.
- C Use Datastream to replicate data from your on-premises MySQL database to BigQuery. Use Forward-SSH tunnel as the connectivity method to establish a secure tunnel between Datastream and your on-premises MySQL database through a tunnel server in your on-premises data center. Use None as the encryption type when setting up the connection profile in Datastream.
- D Use Datastream to replicate data from your on-premises MySQL database to BigQuery. Gather Datastream public IP addresses of the Google Cloud region that will be used to set up the stream. Add those IP addresses to the firewall allowlist of your on-premises data center. Use IP Allowlisting as the connectivity method and Server-only as the encryption type when setting up the connection profile in Datastream.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc di chuyển data warehouse từ on-premises sang BigQuery (dịch vụ kho dữ liệu của Google Cloud). Một nguồn dữ liệu upstream là cơ sở dữ liệu MySQL chạy on-premises trong data center, không có public IP addresses (không thể truy cập trực tiếp từ internet công cộng). Yêu cầu chính là ingest dữ liệu vào BigQuery một cách an toàn, không đi qua public internet để tránh rủi ro bảo mật như lộ dữ liệu hoặc tấn công.
📌 Bối cảnh kỹ thuật:
- On-premises MySQL cần kết nối private (riêng tư) với BigQuery.
- Không dùng public internet nghĩa là phải sử dụng các phương thức kết nối dedicated như Cloud Interconnect, VPN, hoặc tunnel an toàn.
- Công cụ liên quan: Datastream (dịch vụ CDC - Change Data Capture của Google Cloud, hỗ trợ replicate real-time từ MySQL sang BigQuery), ETL tools, ODBC driver.
- Kiến thức cập nhật đến 2026: Theo tài liệu Google Cloud mới nhất (Datastream v1.5+ và Cloud Interconnect Premium/Partner), ưu tiên private connectivity để tuân thủ zero-trust security model.
Mục tiêu: Chọn giải pháp đảm bảo private pathway, mã hóa đúng cách, và tích hợp seamless với BigQuery.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Datastream to replicate data from your on-premises MySQL database to BigQuery. Set up Cloud Interconnect between your on-premises data center and Google Cloud. Use Private connectivity as the connectivity method and allocate an IP address range within your VPC network to the Datastream connectivity configuration. Use Server-only as the encryption type when setting up the connection profile in Datastream.
Lý do chọn 🛠️:
- Datastream lý tưởng cho replicate real-time từ MySQL on-premises sang BigQuery mà không cần ETL thủ công.
- Cloud Interconnect tạo kết nối dedicated private (Layer 3) giữa on-premises và Google Cloud VPC, hoàn toàn tránh public internet (dedicated fiber hoặc partner interconnect, throughput cao đến 100Gbps).
- Private connectivity: Gán IP range từ VPC cho Datastream profile, cho phép traffic private peering.
- Server-only encryption: Mã hóa dữ liệu ở server-side (MySQL), phù hợp cho kết nối private, giảm overhead mà vẫn secure (tuân thủ GCP best practices).
- Giải pháp này scaleable, low-latency, zero-downtime cho migration data warehouse.
📘 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm lý do chi tiết bằng tiếng Việt dựa trên docs GCP mới nhất.
-
❌ [SAI] Update your existing on-premises ETL tool to write to BigQuery by using the BigQuery Open Database Connectivity (ODBC) driver. Set up the proxy parameter in the simba.googlebigqueryodbc.ini file to point to your data center’s NAT gateway.
- Lý do sai: ODBC driver kết nối BigQuery qua TCP/IP public endpoints (port 443), ngay cả khi dùng proxy đến NAT gateway. NAT gateway vẫn route traffic qua public internet (không private), vi phạm yêu cầu "không đi qua public internet". Không hỗ trợ CDC real-time, chỉ batch ETL thủ công, dễ lỗi và không scale cho data warehouse lớn.
-
✅ [ĐÚNG] Use Datastream to replicate data from your on-premises MySQL database to BigQuery. Set up Cloud Interconnect between your on-premises data center and Google Cloud. Use Private connectivity as the connectivity method and allocate an IP address range within your VPC network to the Datastream connectivity configuration. Use Server-only as the encryption type when setting up the connection profile in Datastream.
- Lý do đúng: Như đã giải thích ở trên. Đây là best practice cho on-premises MySQL không public IP: Cloud Interconnect đảm bảo private IP-only traffic, Datastream tự động CDC với low latency (<1s), và Server-only encryption tối ưu cho private link.
-
❌ [SAI] Use Datastream to replicate data from your on-premises MySQL database to BigQuery. Use Forward-SSH tunnel as the connectivity method to establish a secure tunnel between Datastream and your on-premises MySQL database through a tunnel server in your on-premises data center. Use None as the encryption type when setting up the connection profile in Datastream.
- Lý do sai: Forward-SSH tunnel tạo secure channel qua public internet (Datastream outbound SSH đến tunnel server), không hoàn toàn private như yêu cầu. Đặc biệt, None encryption loại bỏ mã hóa SSL/TLS ở connection profile, dẫn đến dữ liệu plain-text – rủi ro bảo mật cao, vi phạm zero-trust. Chỉ dùng cho testing, không production.
-
❌ [SAI] Use Datastream to replicate data from your on-premises MySQL database to BigQuery. Gather Datastream public IP addresses of the Google Cloud region that will be used to set up the stream. Add those IP addresses to the firewall allowlist of your on-premises data center. Use IP Allowlisting as the connectivity method and Server-only as the encryption type when setting up the connection profile in Datastream.
- Lý do sai: IP Allowlisting yêu cầu whitelist public IP ranges của Datastream (ví dụ: 34.0.0.0/8 cho us-central1), nghĩa là traffic qua public internet (inbound từ GCP public IPs). Dù Server-only encryption tốt, nhưng vẫn expose firewall rules public, không đáp ứng "không đi qua public internet" và dễ bị DDoS/throttle.
🔗 Tài liệu tham khảo (cập nhật 2026)
- Datastream Docs: Datastream Connectivity Methods – Chi tiết Private connectivity & Cloud Interconnect.
- Cloud Interconnect: Dedicated Interconnect Overview – Private peering cho on-premises.
- BigQuery Migration Guide: Migrate On-Premises to BigQuery – Khuyến nghị Datastream cho MySQL CDC.
- Datastream Best Practices: Security & Encryption – Server-only cho private links.
Giải pháp này giúp migration an toàn, hiệu quả! 🚀 Nếu cần demo code hoặc setup chi tiết, hãy hỏi thêm.