Ngân hàng đề — Google Cloud Professional Data Engineer

Tìm thấy 429 câu.

Câu 301
You need to migrate a Redis database from an on-premises data center to a Memorystore for Redis instance. You want to follow Google-recommended practices and perform the migration for minimal cost, time and effort. What should you do?
  1. A Make an RDB backup of the Redis database, use the gsutil utility to copy the RDB file into a Cloud Storage bucket, and then import the RDB file into the Memorystore for Redis instance.
  2. B Make a secondary instance of the Redis database on a Compute Engine instance and then perform a live cutover.
  3. C Create a Dataflow job to read the Redis database from the on-premises data center and write the data to a Memorystore for Redis instance.
  4. D Write a shell script to migrate the Redis data and create a new Memorystore for Redis instance.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc di chuyển (migrate) một cơ sở dữ liệu Redis từ trung tâm dữ liệu on-premises (tại chỗ) sang instance Memorystore for Redis trên Google Cloud. Memorystore for Redis là dịch vụ managed Redis của Google Cloud, giúp dễ dàng triển khai và quản lý Redis mà không cần lo lắng về hạ tầng.

Yêu cầu chính:

  • Tuân thủ thực hành được Google khuyến nghị (Google-recommended practices).
  • Tối ưu hóa chi phí thấp nhất (minimal cost), thời gian ngắn nhất (minimal time) và công sức ít nhất (minimal effort).

Đây là kịch bản migration phổ biến trong Google Cloud, sử dụng các công cụ native như RDB backup (snapshot của Redis), gsutil (CLI cho Cloud Storage) và tính năng import sẵn có của Memorystore. Phương pháp này không downtime lớn, dễ thực hiện và an toàn, phù hợp với best practices cập nhật đến năm 2026 (theo tài liệu Memorystore Redis v7+).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Make an RDB backup of the Redis database, use the gsutil utility to copy the RDB file into a Cloud Storage bucket, and then import the RDB file into the Memorystore for Redis instance.

Lý do 🛠️:

  • Đây là phương pháp được Google khuyến nghị chính thức cho migration Redis on-premises sang Memorystore, với downtime tối thiểu (chỉ cần dừng Redis source ngắn để tạo backup fresh nếu cần).
  • Chi phí thấp: Sử dụng RDB (Redis Database file - snapshot binary) miễn phí tạo từ Redis CLI (redis-cli --rdb), gsutil copy rẻ (Cloud Storage giá thấp ~$0.02/GB/tháng), import qua gcloud CLI tự động.
  • Thời gian & effort ít: Quy trình chỉ 3 bước đơn giản, hỗ trợ automation qua script, hoàn thành trong vài phút đến giờ tùy kích thước dữ liệu (hỗ trợ instance lớn đến 300GB+ ở phiên bản 2026).
  • An toàn: RDB đảm bảo consistency, hỗ trợ persistence, và Memorystore tự động scale sau import.

📋 Giải thích chi tiết tất cả các phương án

  • Phương án 1 (Đúng ✅): Make an RDB backup of the Redis database, use the gsutil utility to copy the RDB file into a Cloud Storage bucket, and then import the RDB file into the Memorystore for Redis instance.
    Giải thích: Như đã nêu ở phần đáp án đúng. Đây là best practice native, nhanh, rẻ, không cần code custom hay dịch vụ trung gian. Hỗ trợ AOF/RDB đầy đủ ở Memorystore Redis 7.2+ (2025 update).

  • Phương án 2 (Sai ❌): Make a secondary instance of the Redis database on a Compute Engine instance and then perform a live cutover.
    Giải thích: Phương pháp này phức tạp và tốn kém hơn, yêu cầu setup Compute Engine VM (chi phí VM ~$0.1/giờ+), replicate data qua replication protocol (cần config master-slave), rồi live cutover (DNS switch). Không phải recommended vì effort cao (quản lý VM, network peering on-prem-GCP), thời gian dài (sync liên tục), và chi phí VM chạy lâu. Google ưu tiên RDB import thay vì replication thủ công.

  • Phương án 3 (Sai ❌): Create a Dataflow job to read the Redis database from the on-premises data center and write the data to a Memorystore for Redis instance.
    Giải thích: Không phù hợp vì Dataflow (Apache Beam) dành cho batch/streaming big data ETL (như từ Kafka/JDBC), không optimize cho Redis key-value store. Đọc Redis on-prem qua Dataflow cần connector custom (RedisIO chưa mature), write lại từng key tốn kém (Dataflow vCPU/GPU giá cao ~$0.1/vCPU/giờ), thời gian dài với dữ liệu lớn, và rủi ro data loss/inconsistency do streaming không atomic. Google không recommend cho Redis migration nhỏ.

  • Phương án 4 (Sai ❌): Write a shell script to migrate the Redis data and create a new Memorystore for Redis instance.
    Giải thích: Quá thủ công và không scalable. Shell script (dùng redis-cli SCAN/GET/SET loop) có thể migrate nhỏ lẻ nhưng chậm khủng khiếp với dataset lớn (hàng triệu keys), tốn bandwidth on-prem (export/import từng key), dễ lỗi (timeout, memory overflow), và không handle persistence/RDB/AOF chuẩn. Tạo instance OK nhưng migration script không minimal effort/cost so với RDB native tool. Google khuyên tránh custom script cho production migration.

🧠 Tóm tắt lợi ích phương pháp đúng: Tiết kiệm 90% effort so với các cách khác, phù hợp zero-downtime gần như (backup offline + import online). Nếu dataset >1TB, kết hợp Persistent Disk replication cho hybrid approach (theo docs 2026).

Câu 302
Your platform on your on-premises environment generates 100 GB of data daily, composed of millions of structured JSON text files. Your on-premises environment cannot be accessed from the public internet. You want to use Google Cloud products to query and explore the platform data. What should you do?
  1. A Use Cloud Scheduler to copy data daily from your on-premises environment to Cloud Storage. Use the BigQuery Data Transfer Service to import data into BigQuery.
  2. B Use a Transfer Appliance to copy data from your on-premises environment to Cloud Storage. Use the BigQuery Data Transfer Service to import data into BigQuery.
  3. C Use Transfer Service for on-premises data to copy data from your on-premises environment to Cloud Storage. Use the BigQuery Data Transfer Service to import data into BigQuery.
  4. D Use the BigQuery Data Transfer Service dataset copy to transfer all data into BigQuery.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm

📖 Nội dung câu hỏi:
Câu hỏi mô tả một nền tảng (platform) chạy trên môi trường on-premises (tại chỗ) tạo ra 100 GB dữ liệu mỗi ngày, bao gồm hàng triệu file JSON có cấu trúc. Môi trường on-premises không thể truy cập từ internet công khai (cannot be accessed from the public internet), nghĩa là nó được bảo vệ nghiêm ngặt, không cho phép kết nối inbound từ bên ngoài. Yêu cầu là sử dụng các sản phẩm Google Cloud để truy vấn (query) và khám phá (explore) dữ liệu này.

✅ Mục tiêu chính: Chuyển dữ liệu từ on-premises lên Google Cloud (cụ thể là Cloud Storage làm staging, sau đó vào BigQuery để query/explore), mà không cần môi trường on-premises expose ra internet. Với lượng dữ liệu liên tục hàng ngày (daily) và nhiều file nhỏ, cần giải pháp ongoing transfer (chuyển liên tục), an toàn, không yêu cầu inbound connection từ cloud vào on-prem.

🛠️ Bối cảnh kỹ thuật (cập nhật GCP đến 2026): Theo tài liệu GCP mới nhất (Google Cloud Storage Transfer options, BigQuery loading docs - phiên bản 2024-2026), ưu tiên Storage Transfer Service (STS) với agent on-prem cho trường hợp này vì agent chỉ cần outbound HTTPS từ on-prem đến Google (không cần inbound). Không dùng Transfer Appliance cho daily transfer vì quá chậm (ship vật lý).

✅ Đáp án đúng và lý do chọn

Đáp án đúng: Use Transfer Service for on-premises data to copy data from your on-premises environment to Cloud Storage. Use the BigQuery Data Transfer Service to import data into BigQuery.

Lý do chọn (chi tiết):
🟢 Giải pháp này sử dụng Transfer Service for on-premises data (tức Storage Transfer Service - STS với agent cài trên server on-prem) để chuyển dữ liệu an toàn, liên tục hàng ngày lên Cloud Storage. Agent STS chỉ khởi tạo outbound connection (từ on-prem ra Google), phù hợp với môi trường "không accessible từ public internet". Sau đó, BigQuery Data Transfer Service tự động load từ Cloud Storage vào BigQuery (hỗ trợ JSON structured, external tables hoặc federated queries).
📈 Hoàn hảo cho 100GB/ngày, millions files: STS hỗ trợ parallel transfer, resume, filtering. BigQuery lý tưởng cho query/explore JSON (schema auto-detect, partitioning).

📘 Tài liệu tham khảo:

❌ Phân tích tất cả các phương án (đúng/sai)

  • [SAI] Use Cloud Scheduler to copy data daily from your on-premises environment to Cloud Storage. Use the BigQuery Data Transfer Service to import data into BigQuery.
    ❌ Lý do sai: Cloud Scheduler chỉ là scheduler chạy job (như trigger Cloud Functions/HTTP), không copy data trực tiếp từ on-prem. Không có cơ chế agent hoặc connector an toàn cho on-prem private. Không khả thi cho daily 100GB/millions files (thiếu transfer engine, dễ fail với firewall). Phần BigQuery Transfer OK nhưng upstream sai.

  • [SAI] Use a Transfer Appliance to copy data from your on-premises environment to Cloud Storage. Use the BigQuery Data Transfer Service to import data into BigQuery.
    ❌ Lý do sai: Transfer Appliance (ship ổ cứng vật lý) phù hợp bulk one-time/large offline transfer (TB/PB scale), KHÔNG cho daily recurring (quá trình ship/upload mất 1-2 tuần/lần). Với 100GB/ngày, sẽ tích tụ backlog khổng lồ, không thực tế. Phù hợp air-gapped hoàn toàn, nhưng câu hỏi ngụ ý có outbound khả dụng.

  • [ĐÚNG] Use Transfer Service for on-premises data to copy data from your on-premises environment to Cloud Storage. Use the BigQuery Data Transfer Service to import data into BigQuery.
    ✅ Lý do đúng: Như đã giải thích ở trên. Transfer Service for on-premises = STS agent-based, hỗ trợ POSIX file systems, incremental sync, bandwidth throttle. BigQuery Transfer tự động hóa load JSON vào tables (URI wildcards cho millions files). Scalable đến 2026 với VPC Service Controls.

  • [SAI] Use the BigQuery Data Transfer Service dataset copy to transfer all data into BigQuery.
    ❌ Lý do sai: Không tồn tại "BigQuery Data Transfer Service dataset copy" cho on-prem. BQ Transfer chỉ hỗ trợ sources như Cloud Storage/S3/partner SaaS (không direct on-prem). Không có "dataset copy" feature; cần staging ở CS trước. Bỏ qua bước upload, dẫn đến fail hoàn toàn với môi trường private.

🎯 Kết luận: Giải pháp đúng tận dụng hybrid transfer pattern GCP (on-prem → CS → BQ), đảm bảo security, performance và cost-effective cho workload analytics! Nếu cần lab thực hành, dùng Qwiklabs GCP STS. 🚀

Câu 303
A TensorFlow machine learning model on Compute Engine virtual machines (n2-standard-32) takes two days to complete training. The model has custom TensorFlow operations that must run partially on a CPU. You want to reduce the training time in a cost-effective manner. What should you do?
  1. A Change the VM type to n2-highmem-32.
  2. B Change the VM type to e2-standard-32.
  3. C Train the model using a VM with a GPU hardware accelerator.
  4. D Train the model using a VM with a TPU hardware accelerator.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một mô hình học máy TensorFlow đang được huấn luyện trên máy ảo Compute Engine loại n2-standard-32 (cấu hình: 32 vCPU, 128 GB RAM), mất 2 ngày để hoàn thành. Mô hình có các operations tùy chỉnh (custom TensorFlow operations) phải chạy một phần trên CPU. Mục tiêu là giảm thời gian huấn luyện một cách tiết kiệm chi phí (cost-effective).

🛠️ Vấn đề cốt lõi:

  • Huấn luyện trên CPU thuần túy chậm (2 ngày), cần accelerator phần cứng để tăng tốc.
  • Custom ops bắt buộc một phần chạy trên CPU, nên không thể offload toàn bộ sang accelerator không hỗ trợ CPU (như TPU).
  • Giải pháp phải tối ưu tốc độ + chi phí, ưu tiên hardware accelerator phù hợp với TensorFlow và workload mixed (CPU + accelerator).

📘 Kiến thức cập nhật (GCP 2026): Theo tài liệu Google Cloud Compute Engine và AI Accelerators (cập nhật Q1/2026), GPU (như NVIDIA A100/H100) hỗ trợ linh hoạt cho TensorFlow với custom ops (chạy phần tương thích trên GPU, phần CPU giữ nguyên). TPU v5p (mới nhất) yêu cầu model compile đặc biệt và không hỗ trợ custom ops chạy partial CPU hiệu quả. Nguồn: Google Cloud Compute Engine GPU docs, Cloud TPU docs.

✅ Đáp án đúng

Train the model using a VM with a GPU hardware accelerator.
Lý do chọn:

  • GPU (như NVIDIA Tesla/A100/H100 trên Compute Engine) tăng tốc mạnh mẽ các ops TensorFlow chuẩn (matrix ops, convolutions), trong khi custom ops vẫn chạy mượt trên CPU mạnh của VM.
  • Giảm thời gian từ 2 ngày xuống đáng kể (thường 5-10x tùy workload), cost-effective hơn so với scale CPU thuần vì GPU có giá/GFlops thấp hơn (khoảng 0.5-1$/giờ tùy loại, so với CPU đắt đỏ khi scale ngang).
  • Hỗ trợ TensorFlow native qua CUDA/cuDNN, không cần refactor lớn. Phù hợp workload mixed CPU-GPU.

❌ Phân tích tất cả các phương án

  • Change the VM type to n2-highmem-32.
    Sai vì: n2-highmem-32 chỉ tăng RAM (256 GB so với 128 GB của n2-standard-32), nhưng không tăng CPU cores hoặc thêm accelerator. Thời gian huấn luyện vẫn ~2 ngày vì bottleneck là compute (ops tính toán), không phải memory. Không cost-effective (chi phí cao hơn ~20-30% mà lợi ích thấp).

  • Change the VM type to e2-standard-32.
    Sai vì: e2-standard-32 là dòng Ti4 gen tiết kiệm chi phí (32 vCPU, 128 GB RAM), nhưng hiệu suất CPU thấp hơn n2 (dùng CPU Tau T2A/Arm-based, optimized cho web/general, không phải ML heavy). Có thể chậm hơn do clock speed thấp, tăng thời gian huấn luyện. Chỉ rẻ hơn ~30% nhưng không giải quyết bottleneck compute.

  • Train the model using a VM with a GPU hardware accelerator.
    Đúng (như đã giải thích ở trên). ✅ Tăng tốc tối ưu cho TensorFlow mixed ops, cost-effective với spot/preemptible instances.

  • Train the model using a VM with a TPU hardware accelerator.
    Sai vì: Cloud TPU (v4/v5e/v5p mới nhất 2026) không hỗ trợ custom TensorFlow ops chạy partial trên CPU hiệu quả – TPU yêu cầu toàn bộ graph compile qua XLA (TensorFlow compiler) và chạy vectorized trên TPU systolic array, không có CPU fallback mạnh cho custom ops. Phải refactor code lớn (dùng tf.experimental.tensorrt hoặc TPU-specific), tốn kém thời gian/dev. Không cost-effective cho workload này (TPU đắt hơn GPU cho non-pure TPU models). Nguồn: TPU Compatibility docs.

🧩 Kết luận: Chọn GPU là giải pháp tối ưu nhất cho huấn luyện TensorFlow có custom CPU ops trên GCP, giảm thời gian mà không cần thay đổi code lớn. Nếu scale, kết hợp với Vertex AI Training cho managed GPU! 🚀

Câu 304
You want to create a machine learning model using BigQuery ML and create an endpoint for hosting the model using Vertex AI. This will enable the processing of continuous streaming data in near-real time from multiple vendors. The data may contain invalid values. What should you do?
  1. A Create a new BigQuery dataset and use streaming inserts to land the data from multiple vendors. Configure your BigQuery ML model to use the "ingestion" dataset as the framing data.
  2. B Use BigQuery streaming inserts to land the data from multiple vendors where your BigQuery dataset ML model is deployed.
  3. C Create a Pub/Sub topic and send all vendor data to it. Connect a Cloud Function to the topic to process the data and store it in BigQuery.
  4. D Create a Pub/Sub topic and send all vendor data to it. Use Dataflow to process and sanitize the Pub/Sub data and stream it to BigQuery.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi yêu cầu xây dựng một mô hình máy học (machine learning model) sử dụng BigQuery ML (công cụ xây dựng ML trực tiếp trong BigQuery) và triển khai endpoint để host mô hình này trên Vertex AI (nền tảng ML end-to-end của Google Cloud). Mục tiêu là xử lý dữ liệu streaming liên tục (continuous streaming data) từ nhiều nhà cung cấp (multiple vendors) với tốc độ near-real time (gần thời gian thực). Dữ liệu có thể chứa giá trị không hợp lệ (invalid values), nên cần xử lý (sanitize) trước khi đưa vào mô hình.

🛠️ Thách thức chính:

  • Dữ liệu streaming cao tải từ nhiều nguồn → Cần hệ thống queue/buffer như Pub/Sub.
  • Xử lý invalid values → Cần pipeline xử lý mạnh mẽ, scalable như Dataflow.
  • Đưa dữ liệu sạch vào BigQuery → Để train/inference BQML model.
  • Deploy model lên Vertex AI endpoint → Cho phép inference near-real time.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create a Pub/Sub topic and send all vendor data to it. Use Dataflow to process and sanitize the Pub/Sub data and stream it to BigQuery.

Lý do 🏆:

  • Pub/Sub làm trung tâm nhận dữ liệu streaming từ multiple vendors, hỗ trợ high-throughput, at-least-once delivery, near-real time (latency <1s).
  • Dataflow (dựa Apache Beam) lý tưởng cho streaming pipeline: đọc từ Pub/Sub, sanitize invalid values (filter/transform bằng Beam transforms), rồi stream trực tiếp vào BigQuery (sử dụng BigQueryIO sink). Dataflow auto-scale, fault-tolerant, xử lý volume lớn liên tục.
  • Dữ liệu sạch vào BigQuery → Train BQML model → Export/deploy endpoint trên Vertex AI cho inference near-real time.
  • Phù hợp best practice GCP cho streaming ML pipeline (cập nhật 2026: Dataflow hỗ trợ Flex Templates cho BQ streaming).

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm giải thích rõ ràng.

  • [SAI] Create a new BigQuery dataset and use streaming inserts to land the data from multiple vendors. Configure your BigQuery ML model to use the "ingestion" dataset as the framing data.
    ❌ Sai vì: Streaming inserts trực tiếp vào BigQuery từ multiple vendors có quota giới hạn (500,000 rows/sec/table, nhưng dễ throttle với high-volume/multi-source). Không có cơ chế sanitize invalid values (BQ chỉ buffer 1h, invalid data sẽ làm model train lỗi). "Framing data" không chuẩn cho BQML streaming; BQML cần dữ liệu sạch, ổn định để train/deploy Vertex AI. Không scalable cho continuous streaming.

  • [SAI] Use BigQuery streaming inserts to land the data from multiple vendors where your BigQuery dataset ML model is deployed.
    ❌ Sai vì: Tương tự phương án 1, streaming inserts trực tiếp thiếu xử lý invalid values (dữ liệu bẩn sẽ corrupt table/model). Deploy BQML trực tiếp trên dataset đang stream → Race condition (train/inference conflict), không near-real time ổn định. Không dùng Pub/Sub/Dataflow làm buffer, dễ vượt quota BQ (cập nhật 2026: vẫn 1MB/insert limit).

  • [SAI] Create a Pub/Sub topic and send all vendor data to it. Connect a Cloud Function to the topic to process the data and store it in BigQuery.
    ❌ Sai vì: Pub/Sub tốt cho ingestion, nhưng Cloud Function không phù hợp streaming continuous/high-volume: Cold start (latency >100ms), timeout 60s (9 phút với gen2), không auto-scale tốt như Dataflow cho unbounded data. Xử lý sanitize yếu (memory-limited 8GB), dễ fail với invalid values lớn. Không phải best practice cho ML streaming pipeline (GCP recommend Dataflow cho Beam-based processing).

  • [ĐÚNG] Create a Pub/Sub topic and send all vendor data to it. Use Dataflow to process and sanitize the Pub/Sub data and stream it to BigQuery.
    ✅ Đúng vì: Như giải thích ở phần đáp án đúng. Toàn diện: Decouple ingestion (Pub/Sub) + Transform/sanitize (Dataflow) + Sink (BQ). Hỗ trợ Vertex AI endpoint inference near-real time trên BQML model. Scalable đến petabyte-scale (Dataflow Unified 2026).

🧠 Kết luận: Phương án đúng tận dụng Streaming Data Pipeline pattern của GCP, đảm bảo dữ liệu sạch, low-latency cho ML workflow!

Câu 305
You have a data processing application that runs on Google Kubernetes Engine (GKE). Containers need to be launched with their latest available configurations from a container registry. Your GKE nodes need to have GPUs, local SSDs, and 8 Gbps bandwidth. You want to efficiently provision the data processing infrastructure and manage the deployment process. What should you do?
  1. A Use Compute Engine startup scripts to pull container images, and use gcloud commands to provision the infrastructure.
  2. B Use Cloud Build to schedule a job using Terraform build to provision the infrastructure and launch with the most current container images.
  3. C Use GKE to autoscale containers, and use gcloud commands to provision the infrastructure.
  4. D Use Dataflow to provision the data pipeline, and use Cloud Scheduler to run the job.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc provision (cung cấp) và quản lý hạ tầng cho một ứng dụng xử lý dữ liệu chạy trên Google Kubernetes Engine (GKE). Các yêu cầu cụ thể bao gồm:

  • Containers phải được khởi chạy với cấu hình mới nhất từ container registry (như Artifact Registry hoặc Container Registry), nghĩa là cần cơ chế tự động pull image mới nhất mà không can thiệp thủ công.
  • GKE nodes phải hỗ trợ GPU (cho tính toán song song), local SSD (lưu trữ nhanh, tạm thời), và 8 Gbps bandwidth (để đạt network performance cao, thường dùng machine types như n1-standard với accelerators).
  • Mục tiêu: Hiệu quả provision hạ tầng (tự động hóa, scalable) và quản lý deployment process (CI/CD pipeline để deploy containers mới nhất). Đây là kịch bản điển hình trong Google Cloud, nhấn mạnh IaC (Infrastructure as Code) và CI/CD để tránh manual operations, phù hợp với best practices của GKE Autopilot hoặc Standard clusters với node pools tùy chỉnh (cập nhật đến 2026: GKE hỗ trợ GPU như NVIDIA A100/H100, local SSD lên đến 8TB/node, và network tier Premium cho 8Gbps+).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Cloud Build to schedule a job using Terraform build to provision the infrastructure and launch with the most current container images.

Lý do chi tiết:

  • Cloud Build là dịch vụ CI/CD serverless của Google Cloud, hỗ trợ schedule jobs qua Cloud Scheduler hoặc triggers (cập nhật 2026: tích hợp sâu với GKE và Terraform).
  • Terraform build trong Cloud Build cho phép định nghĩa IaC để provision GKE cluster/node pools với GPU (qua google_container_node_pool + accelerator), local SSD (local_ssd_count), và high-bandwidth (machine_type như n2-highmem-16 với network bandwidth >8Gbps).
  • Tự động pull latest container images qua build steps (ví dụ: gcr.io/cloud-builders/docker pull hoặc ko.apply cho Kubernetes manifests), đảm bảo deployment luôn dùng config mới nhất.
  • Hiệu quả cao: Tích hợp end-to-end (provision infra + deploy), scalable, không cần quản lý servers, phù hợp data processing workloads trên GKE. 📘 Nguồn tham khảo: Cloud Build Terraform integration, GKE with GPU/SSD, Terraform GCP Provider v5.x (2026).

🛠️ Giải thích tất cả các phương án

Dưới đây là phân tích từng phương án, giữ nguyên nội dung gốc bằng tiếng Anh. Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm giải thích chi tiết bằng tiếng Việt:

  • Use Compute Engine startup scripts to pull container images, and use gcloud commands to provision the infrastructure.
    ❌ Sai: Phương án này dùng startup scripts trên Compute Engine VM để pull images và gcloud CLI thủ công để provision. Không liên quan đến GKE (chỉ VM thuần), không hỗ trợ tự động latest images hiệu quả, thiếu scalability cho GPU/local SSD/bandwidth cao. Quản lý thủ công, vi phạm nguyên tắc "efficient provision" và không phải best practice cho container workloads (manual, error-prone).

  • Use Cloud Build to schedule a job using Terraform build to provision the infrastructure and launch with the most current container images.
    ✅ Đúng: Như đã giải thích ở trên. Hoàn hảo khớp yêu cầu: IaC với Terraform provision GKE infra tùy chỉnh (GPU/SSD/8Gbps), Cloud Build schedule + pull latest images tự động. Đây là workflow CI/CD tiêu chuẩn cho GKE data processing (2026: hỗ trợ GKE Enterprise với Confidential GPUs).

  • Use GKE to autoscale containers, and use gcloud commands to provision the infrastructure.
    ❌ Sai: GKE autoscale (HPA/Cluster Autoscaler) chỉ scale pods/nodes sau khi cluster đã có, nhưng gcloud commands là thủ công để provision (không IaC). Không đảm bảo pull latest images tự động, thiếu scheduling, và không hiệu quả cho initial provisioning với hardware đặc biệt (GPU/SSD). Gcloud phù hợp dev/test, không production-scale.

  • Use Dataflow to provision the data pipeline, and use Cloud Scheduler to run the job.
    ❌ Sai: Dataflow là managed service cho Apache Beam (batch/stream processing), không chạy custom containers trên GKE với GPU/local SSD/bandwidth cụ thể. Chỉ schedule jobs qua Cloud Scheduler, nhưng không provision GKE infra hay pull images từ registry. Không khớp với "GKE-based application", Dataflow templates không hỗ trợ hardware tùy chỉnh như vậy (2026: Dataflow vẫn tập trung unified batch/stream, không thay thế GKE).

Kết luận 🚀: Phương án đúng tận dụng Cloud Build + Terraform để tự động hóa toàn bộ lifecycle (provision + deploy), phù hợp Professional Data Engineer best practices trên Google Cloud. Nếu triển khai thực tế, khuyến nghị dùng GKE Autopilot cho ít quản lý hơn!

Câu 306
You need ads data to serve AI models and historical data for analytics. Longtail and outlier data points need to be identified. You want to cleanse the data in near-real time before running it through AI models. What should you do?
  1. A Use Cloud Storage as a data warehouse, shell scripts for processing, and BigQuery to create views for desired datasets.
  2. B Use Dataflow to identify longtail and outlier data points programmatically, with BigQuery as a sink.
  3. C Use BigQuery to ingest, prepare, and then analyze the data, and then run queries to create views.
  4. D Use Cloud Composer to identify longtail and outlier data points, and then output a usable dataset to BigQuery.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc xử lý dữ liệu quảng cáo (ads data) trong môi trường Google Cloud Platform (GCP). Cụ thể:

  • Yêu cầu chính: Cần dữ liệu quảng cáo để phục vụ các mô hình AI (serve AI models) và dữ liệu lịch sử cho phân tích (historical data for analytics).
  • Thách thức: Phải xác định các điểm dữ liệu "longtail" (dữ liệu đuôi dài, thường là các giá trị hiếm gặp, tần suất thấp) và "outlier" (dữ liệu ngoại lai, bất thường).
  • Yêu cầu xử lý: Làm sạch dữ liệu (cleanse the data) ở chế độ near-real time (gần thời gian thực) trước khi đưa vào mô hình AI.

Mục tiêu là chọn giải pháp tối ưu, scalable cho xử lý dữ liệu streaming/batch, phát hiện bất thường programmatically, và lưu trữ kết quả cho analytics/AI. Đây là tình huống điển hình trong Data Engineering trên GCP, nhấn mạnh vào stream processing để đáp ứng near-real time. (Kiến thức cập nhật đến 2026: Dataflow hỗ trợ Beam 2.58+ với cải tiến Auto-scaling và Unified Streaming).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Dataflow to identify longtail and outlier data points programmatically, with BigQuery as a sink.

Lý do:
🛠️ Dataflow (dựa trên Apache Beam) là dịch vụ lý tưởng cho near-real time processing trên GCP, hỗ trợ xử lý streaming dữ liệu lớn, phát hiện longtail/outlier bằng code tùy chỉnh (programmatically, ví dụ dùng PTransforms, ML transforms như TensorFlow). Nó cleanse dữ liệu trước khi đưa vào AI models.
📊 BigQuery as a sink lưu trữ dữ liệu sạch cho historical analytics và serving AI (hỗ trợ ML inference qua BigQuery ML). Giải pháp này scalable, serverless, tự động scale cho ads data volume cao. Không có tool nào khác trên GCP xử lý near-real time tốt hơn Dataflow cho trường hợp này.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc:

  • ❌ Use Cloud Storage as a data warehouse, shell scripts for processing, and BigQuery to create views for desired datasets.
    Sai vì: Cloud Storage chỉ là object storage, không phải data warehouse (thiếu query engine, schema management). Shell scripts không scalable cho near-real time, dễ lỗi, không xử lý streaming/outlier programmatically. BigQuery views chỉ là reporting layer, không cleanse dữ liệu trước AI. Giải pháp này batch-oriented, thủ công, không đáp ứng yêu cầu thời gian thực.

  • ✅ Use Dataflow to identify longtail and outlier data points programmatically, with BigQuery as a sink.
    Đúng vì: Như đã giải thích ở trên. Dataflow xử lý near-real time cleansing, phát hiện longtail/outlier bằng code (ví dụ: custom DoFn với percentile stats hoặc anomaly detection libs). BigQuery sink đảm bảo dữ liệu sẵn sàng cho AI serving và analytics. Tối ưu nhất theo best practices GCP.

  • ❌ Use BigQuery to ingest, prepare, and then analyze the data, and then run queries to create views.
    Sai vì: BigQuery xuất sắc cho batch analytics/SQL queries, nhưng không hỗ trợ near-real time processing (streaming inserts chậm cho high-velocity ads data, prepare/outlier detection chỉ qua SQL – không programmatic linh hoạt). Views chỉ là virtual layer, không cleanse trước AI. Không phù hợp cho serving AI real-time.

  • ❌ Use Cloud Composer to identify longtail and outlier data points, and then output a usable dataset to BigQuery.
    Sai vì: Cloud Composer (Managed Apache Airflow) là orchestration tool cho workflow/DAGs, không phải data processing engine. Nó không xử lý streaming/near-real time trực tiếp, khó identify outlier programmatically (chỉ schedule tasks). Phụ thuộc operator bên ngoài, kém hiệu quả hơn Dataflow cho cleansing lớn.

📘 Tài liệu tham khảo (cập nhật 2026)

Giải pháp này đảm bảo end-to-end pipeline hiệu quả trên GCP! 🚀

Câu 307
You are collecting IoT sensor data from millions of devices across the world and storing the data in BigQuery. Your access pattern is based on recent data, filtered by location_id and device_version with the following query:

SELECT 
    MAX(temperature)
FROM 
    acme_iot_data.sensors
WHERE 
    create_date > DATE_SUB(CURRENT_DATE(), INTERVAL 7 day)
    AND location_id = "SW1W9TQ"
    AND device_version = "202007r3"


You want to optimize your queries for cost and performance. How should you structure your data?
  1. A Partition table data by create_date, location_id, and device_version.
  2. B Partition table data by create_date, cluster table data by location_id, and device_version.
  3. C Cluster table data by create_date, location_id, and device_version.
  4. D Cluster table data by create_date, partition by location_id, and device_version.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc tối ưu hóa truy vấn (query) trên BigQuery (dịch vụ kho dữ liệu của Google Cloud) cho dữ liệu IoT từ hàng triệu thiết bị. Dữ liệu được lưu trữ trong bảng acme_iot_data.sensors, với truy vấn điển hình lọc:

  • Dữ liệu gần đây (create_date > 7 ngày trước).
  • Theo location_id cụ thể (ví dụ: "SW1W9TQ").
  • Theo device_version cụ thể (ví dụ: "202007r3").
  • Tính toán MAX(temperature).

📈 Mục tiêu: Giảm chi phí (scan ít dữ liệu hơn) và tăng hiệu suất bằng cách cấu trúc dữ liệu sử dụng Partitioning (phân vùng theo thời gian để prune partition) và Clustering (nhóm dữ liệu theo cột filter phổ biến để tối ưu sắp xếp và lọc). Đây là best practice cho BigQuery ingestion-time partitioned tables (cập nhật đến 2026, BigQuery hỗ trợ hybrid partitioning + clustering với tối đa 4 cluster keys).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Partition table data by create_date, cluster table data by location_id, and device_version.

🛠️ Lý do:

  • Partition by create_date: Truy vấn luôn filter theo khoảng thời gian gần (7 ngày), nên partitioning theo ngày giúp BigQuery prune (loại bỏ) các partition cũ, chỉ scan dữ liệu 7 ngày gần nhất → Tiết kiệm chi phí lớn (giảm bytes scanned lên đến 90%+).
  • Cluster by location_id và device_version: Hai cột này là filter equality (bằng giá trị cụ thể), clustering sắp xếp dữ liệu vật lý theo chúng → Tăng hiệu quả lọc trong partition, giảm I/O.
  • Kết hợp partition + cluster là cách tối ưu nhất cho pattern này (theo docs BigQuery 2026: hỗ trợ multi-level clustering lên đến 4 keys).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên nguyên tắc BigQuery (partition ưu tiên thời gian, cluster cho filter non-time).

  • ✅ [ĐÚNG] Partition table data by create_date, cluster table data by location_id, and device_version.
    🟢 Đúng vì: Kết hợp hoàn hảo partition thời gian (prune date range) + cluster filter columns (location_id, device_version) → Query chỉ scan dữ liệu liên quan, tối ưu cost/performance. BigQuery tự động cluster khi insert/query.

  • ❌ [SAI] Partition table data by create_date, location_id, and device_version.
    🔴 Sai vì: BigQuery không hỗ trợ multi-column partitioning cho ingestion-time tables (chỉ partition theo single column như date/INT64). Partition multi-key chỉ khả dụng ở declarative partitioning (beta 2026), nhưng không phù hợp pattern thời gian + equality filters này → Không prune hiệu quả, tăng overhead.

  • ❌ [SAI] Cluster table data by create_date, location_id, and device_version.
    🔴 Sai vì: Clustering không prune partition như partitioning thực thụ. Filter create_date trên cluster chỉ sắp xếp dữ liệu, vẫn scan toàn bộ bảng → Không tối ưu cho recent data (7 days), chi phí cao với dữ liệu lớn từ millions devices.

  • ❌ [SAI] Cluster table data by create_date, partition by location_id, and device_version.
    🔴 Sai vì: Partition không thể theo location_id/device_version (string/non-time columns) một cách hiệu quả. Partition lý tưởng cho high-cardinality time/INT, không phải strings như location_id → Không prune date range chính, query chậm và đắt.

📘 Tài liệu tham khảo (cập nhật mới nhất đến 2026)

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần ví dụ code DDL table, hãy hỏi thêm.

Câu 308
A live TV show asks viewers to cast votes using their mobile phones. The event generates a large volume of data during a 3-minute period. You are in charge of the "Voting infrastructure" and must ensure that the platform can handle the load and that all votes are processed. You must display partial results while voting is open. After voting closes, you need to count the votes exactly once while optimizing cost. What should you do?

  1. A Create a Memorystore instance with a high availability (HA) configuration.
  2. B Create a Cloud SQL for PostgreSQL database with high availability (HA) configuration and multiple read replicas.
  3. C Write votes to a Pub/Sub topic and have Cloud Functions subscribe to it and write votes to BigQuery.
  4. D Write votes to a Pub/Sub topic and load into both Bigtable and BigQuery via a Dataflow pipeline. Query Bigtable for real-time results and BigQuery for later analysis. Shut down the Bigtable instance when voting concludes.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong Google Cloud Platform (GCP): Một chương trình truyền hình trực tiếp yêu cầu khán giả bình chọn qua điện thoại di động, tạo ra lượng dữ liệu khổng lồ chỉ trong 3 phút (burst traffic cao, cần xử lý tải lớn với độ trễ thấp). Bạn chịu trách nhiệm xây dựng Voting Infrastructure (phần ? trong hình ảnh) để:

  • Xử lý toàn bộ phiếu bầu mà không mất mát, đảm bảo exactly-once processing (xử lý chính xác một lần).
  • Hiển thị kết quả tạm thời (partial results) trong lúc bình chọn còn mở → yêu cầu real-time querying (truy vấn thời gian thực, low-latency).
  • Sau khi đóng bình chọn: Đếm phiếu chính xác một lần và tối ưu chi phí (không lãng phí tài nguyên dài hạn).

📸 Phân tích hình ảnh:

  • Bên trái: Serving Platform là Kubernetes Engine (GKE) cluster với multiple instances (tự động scale để phục vụ viewers).
  • Mũi tên hai chiều giữa Voters & Viewers (từ mobile) → GKE → Voting Infrastructure (?):
    • GKE nhận votes từ người dùng và gửi đến Voting Infrastructure để lưu trữ/xử lý.
    • Voting Infrastructure cần nhận dữ liệu từ GKE, xử lý real-time, và có thể gửi kết quả ngược lại cho GKE hiển thị.
  • Yêu cầu chính: Voting Infrastructure phải scale horizontally cho burst 3 phút, hỗ trợ real-time dashboard, exactly-once semantics, và tắt tài nguyên sau event để tiết kiệm chi phí (theo mô hình pay-as-you-go của GCP).

Kiến thức cập nhật đến 2026 (GCP phiên bản mới nhất): Sử dụng Pub/Sub cho decoupling, Dataflow (Apache Beam) cho streaming ETL với exactly-once (nhờ runner v2), Bigtable cho NoSQL real-time (millions QPS), BigQuery cho analytics batch/cost-effective.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Write votes to a Pub/Sub topic and load into both Bigtable and BigQuery via a Dataflow pipeline. Query Bigtable for real-time results and BigQuery for later analysis. Shut down the Bigtable instance when voting concludes.

Lý do 🛠️:

  • Pub/Sub: Decouples GKE từ storage, auto-scale ingestion millions messages/sec mà không mất dữ liệu.
  • Dataflow pipeline: Xử lý streaming votes với exactly-once semantics (windowing + dedup), branch data vào Bigtable (real-time aggregation/queries cho partial results, latency <10ms) và BigQuery (batch analysis chính xác sau event).
  • Bigtable: Hoàn hảo cho real-time dashboard (high QPS, strong consistency), nhưng tắt instance sau 3 phút → tối ưu chi phí (chỉ trả cho thời gian sử dụng).
  • Phù hợp burst ngắn: Scale up/down nhanh, tích hợp GKE qua service account.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ SAI: Create a Memorystore instance with a high availability (HA) configuration.
    Lý do sai: Memorystore (Redis) tốt cho caching/counting real-time (in-memory, sub-ms latency), nhưng không hỗ trợ exactly-once durable storage (dữ liệu mất nếu crash), không phù hợp đếm chính xác sau event. Chi phí cao cho HA liên tục, không optimize bằng shutdown. Không scale write tốt cho burst millions votes.

  • ❌ SAI: Create a Cloud SQL for PostgreSQL database with high availability (HA) configuration and multiple read replicas.
    Lý do sai: Cloud SQL (relational) có HA/read replicas cho queries, nhưng không scale horizontally cho burst cao (connection limits ~65k, write bottleneck). Latency cao hơn NoSQL cho real-time, chi phí fixed cao (provisioned IOPS), khó exactly-once ở scale lớn. Không tối ưu cho short-lived event.

  • ❌ SAI: Write votes to a Pub/Sub topic and have Cloud Functions subscribe to it and write votes to BigQuery.
    Lý do sai: Pub/Sub tốt cho ingestion, nhưng Cloud Functions có cold starts (latency spike ở burst), execution limits (540s), không guaranteed exactly-once (at-least-once + dedup manual). BigQuery streaming ok cho insert, nhưng real-time queries kém (latency 1-2s, không partial results mượt), thiếu low-latency dashboard. Không tối ưu chi phí cho real-time.

  • ✅ ĐÚNG: Write votes to a Pub/Sub topic and load into both Bigtable and BigQuery via a Dataflow pipeline. Query Bigtable for real-time results and BigQuery for later analysis. Shut down the Bigtable instance when voting concludes.
    Lý do đúng (tóm tắt): Toàn diện giải quyết burst → real-time → exactly-once → analytics → cost-optimize, tích hợp hoàn hảo với GKE qua Pub/Sub. Dataflow autoscaling, Bigtable cho live dashboard (query từ GKE), BigQuery cho final count rẻ tiền. Shutdown Bigtable via API sau 3 phút → zero waste.

Câu 309
A shipping company has live package-tracking data that is sent to an Apache Kafka stream in real time. This is then loaded into BigQuery. Analysts in your company want to query the tracking data in BigQuery to analyze geospatial trends in the lifecycle of a package. The table was originally created with ingest-date partitioning. Over time, the query processing time has increased. You need to copy all the data to a new clustered table. What should you do?
  1. A Re-create the table using data partitioning on the package delivery date.
  2. B Implement clustering in BigQuery on the package-tracking ID column.
  3. C Implement clustering in BigQuery on the ingest date column.
  4. D Tier older data onto Cloud Storage files and create a BigQuery table using Cloud Storage as an external data source.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một công ty vận chuyển có dữ liệu theo dõi gói hàng thời gian thực được gửi vào Apache Kafka stream, sau đó load vào BigQuery. Các nhà phân tích muốn query dữ liệu theo dõi trong BigQuery để phân tích xu hướng địa không gian (geospatial trends) trong vòng đời của gói hàng.

  • Bảng gốc được tạo với partitioning theo ingest-date (ngày dữ liệu được ingest vào BigQuery).
  • Theo thời gian, thời gian xử lý query tăng dần do lượng dữ liệu tích lũy lớn.
  • Nhiệm vụ: Copy toàn bộ dữ liệu sang một bảng clustered mới để tối ưu hiệu suất query.

🛠️ Vấn đề cốt lõi: Partitioning theo ingest-date giúp quản lý dữ liệu theo thời gian ingest, nhưng không tối ưu cho query theo package lifecycle (vòng đời gói hàng, liên quan đến package-tracking ID và geospatial data). Cần clustering để sắp xếp dữ liệu bên trong các partition, giảm scan dữ liệu không cần thiết và tăng tốc query.
📈 Mục tiêu: Sử dụng clustered table trong BigQuery (tính năng từ 2019, cập nhật mới nhất 2024-2026 hỗ trợ clustering lên đến 4 cột, tự động reorganize khi insert/update).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Implement clustering in BigQuery on the package-tracking ID column.

Lý do:

  • Clustering trên package-tracking ID sẽ nhóm dữ liệu liên quan đến cùng một gói hàng lại với nhau trong các partition (dựa trên ingest-date).
  • Điều này tối ưu hóa query geospatial trends theo lifecycle vì dữ liệu của một package thường được ingest ở nhiều thời điểm khác nhau → query chỉ scan dữ liệu liên quan, giảm thời gian xử lý đáng kể (có thể lên đến 90% theo docs BigQuery).
  • Khi copy dữ liệu sang bảng mới: Sử dụng CREATE TABLE ... CLUSTER BY package_tracking_id và INSERT INTO new_table SELECT * FROM old_table → BigQuery tự động cluster data.
  • Phù hợp nhất với yêu cầu "copy all the data to a new clustered table" mà không thay đổi partitioning gốc.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Re-create the table using data partitioning on the package delivery date.
    Phương án này thay đổi partitioning từ ingest-date sang package delivery date, nhưng:

    • Dữ liệu real-time từ Kafka có thể ingest trước delivery date → partitioning mới không khớp, gây lỗi khi load stream mới.
    • Không giải quyết query chậm vì vẫn thiếu clustering cho geospatial queries (delivery date chỉ là một điểm cuối, không cover full lifecycle).
    • Tốn kém: Re-partition toàn bộ data lớn → chi phí cao hơn clustering (chỉ cần trên bảng mới).
  • ✅ [ĐÚNG] Implement clustering in BigQuery on the package-tracking ID column.
    Như giải thích ở trên: Clustering trên ID nhóm data theo package → lý tưởng cho query trends theo lifecycle, giảm scan data hiệu quả. BigQuery hỗ trợ clustering kết hợp partitioning (hybrid model), giữ nguyên ingest-date partition và thêm cluster layer.

  • ❌ [SAI] Implement clustering in BigQuery on the ingest date column.
    Clustering trên ingest date vô ích vì bảng đã partition theo cùng cột này → BigQuery không gain lợi ích (cluster chỉ hiệu quả khi khác partition key).
    Query geospatial vẫn scan full partition không liên quan → thời gian xử lý không cải thiện.

  • ❌ [SAI] Tier older data onto Cloud Storage files and create a BigQuery table using Cloud Storage as an external data source.
    Phương án này tier data cũ ra Cloud Storage (external table), nhưng:

    • Không copy all data vào BigQuery clustered table như yêu cầu (external table chậm hơn native BigQuery table ~10x).
    • Query real-time + historical sẽ chậm với external data, không phù hợp geospatial trends full lifecycle.
    • Phức tạp: Cần ETL riêng cho tiering, tăng latency.

📘 Tài liệu tham khảo (cập nhật mới nhất 2024-2026)

Câu 310
You are designing a data mesh on Google Cloud with multiple distinct data engineering teams building data products. The typical data curation design pattern consists of landing files in Cloud Storage, transforming raw data in Cloud Storage and BigQuery datasets, and storing the final curated data product in BigQuery datasets. You need to configure Dataplex to ensure that each team can access only the assets needed to build their data products. You also need to ensure that teams can easily share the curated data product. What should you do?
  1. A 1. Create a single Dataplex virtual lake and create a single zone to contain landing, raw, and curated data.
    2. Provide each data engineering team access to the virtual lake.
  2. B 1. Create a single Dataplex virtual lake and create a single zone to contain landing, raw, and curated data.
    2. Build separate assets for each data product within the zone.
    3. Assign permissions to the data engineering teams at the zone level.
  3. C 1. Create a Dataplex virtual lake for each data product, and create a single zone to contain landing, raw, and curated data.
    2. Provide the data engineering teams with full access to the virtual lake assigned to their data product.
  4. D 1. Create a Dataplex virtual lake for each data product, and create multiple zones for landing, raw, and curated data.
    2. Provide the data engineering teams with full access to the virtual lake assigned to their data product.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi

Câu hỏi tập trung vào việc thiết kế data mesh trên Google Cloud sử dụng Dataplex – một dịch vụ quản lý data lakehouse hiện đại (cập nhật đến phiên bản mới nhất năm 2026). Data mesh là kiến trúc phân tán, nơi các domain team (nhóm data engineering riêng biệt) sở hữu và xây dựng data products độc lập, nhấn mạnh vào decentralized data ownership và self-service.

Mô tả tình huống:

  • Mỗi team xây dựng data products theo pattern tiêu chuẩn:
    • Landing: Files thô lưu vào Cloud Storage.
    • Transform: Xử lý raw data trong Cloud Storage và BigQuery datasets.
    • Curated: Data sản phẩm cuối lưu vào BigQuery datasets.
  • Yêu cầu chính: ✅ Isolation: Mỗi team chỉ truy cập assets cần thiết (GCS buckets, BigQuery datasets) cho data product của họ → tránh cross-team access không mong muốn. ✅ Sharing: Dễ dàng chia sẻ curated data products với các team khác hoặc consumers.
  • Dataplex concepts (cập nhật 2026):
    • Lake: Nhóm logic dữ liệu cross-projects/regions, đại diện cho một domain/data product.
    • Zone: Phân chia lake theo stages (landing/raw/curated), hỗ trợ governance, metadata, security riêng.
    • Asset: Tài nguyên cụ thể (GCS bucket/BigQuery dataset) trong zone.

📘 Nguồn tham khảo:

✅ Đáp án đúng

Phương án ĐÚNG:

  1. Create a Dataplex virtual lake for each data product, and create multiple zones for landing, raw, and curated data.
  2. Provide the data engineering teams with full access to the virtual lake assigned to their data product.

Lý do chọn 🛠️:

  • Tạo lake riêng cho mỗi data product → domain isolation hoàn hảo, mỗi team chỉ thấy lake của mình, tránh lẫn lộn assets giữa các team (tuân thủ data mesh principles).
  • Multiple zones (landing/raw/curated) → Phân tách rõ ràng stages dữ liệu trong lake, hỗ trợ governance (tags, policies) và transform pipeline độc lập.
  • Full access to their lake → Team tự quản lý toàn bộ lifecycle data product (landing → curated), nhưng sharing dễ dàng bằng cách grant fine-grained permissions cho curated zone/assets với consumers bên ngoài (qua IAM roles như roles/dataplex.assetViewer).
  • Hoàn hảo cho scalability và security trong multi-team setup.

❌ Phân tích tất cả các phương án

  • Phương án 1 (SAI):

    1. Create a single Dataplex virtual lake and create a single zone to contain landing, raw, and curated data.
    2. Provide each data engineering team access to the virtual lake.
      Giải thích sai ❌: Single lake + single zone → không isolate teams, tất cả assets lẫn lộn, teams có thể access data của nhau → vi phạm yêu cầu "each team can access only the assets needed". Không hỗ trợ data mesh decentralized.
  • Phương án 2 (SAI):

    1. Create a single Dataplex virtual lake and create a single zone to contain landing, raw, and curated data.
    2. Build separate assets for each data product within the zone.
    3. Assign permissions to the data engineering teams at the zone level.
      Giải thích sai ❌: Vẫn single lake/zone → assets riêng nhưng permissions ở zone level quá coarse-grained, teams dễ access assets ngoài data product của mình. Không tách stages dữ liệu, khó governance và sharing curated data một cách an toàn.
  • Phương án 3 (SAI):

    1. Create a Dataplex virtual lake for each data product, and create a single zone to contain landing, raw, and curated data.
    2. Provide the data engineering teams with full access to the virtual lake assigned to their data product.
      Giải thích sai ❌: Lake riêng tốt cho isolation, nhưng single zone chứa tất cả stages → không phân tách landing/raw/curated, vi phạm best practice Dataplex (zones dành riêng cho stages để apply policies riêng). Khó transform và share curated mà không expose raw data.

Tóm tắt 🎯: Chỉ phương án đúng cân bằng isolation per team/domain (lake riêng) + separation of concerns (multiple zones) + easy sharing (IAM trên zones/assets), phù hợp data mesh trên Dataplex 2026!