Ngân hàng đề — Google Cloud Associate Data Practitioner

Tìm thấy 333 câu.

Câu 321
You are working on a project that requires analyzing dally social media data. You have 100 GB of JSON formatted data stored in Cloud Storage that keeps growing. You need to transform and load this data into BigQuery for analysis. You want to follow the Google-recommended approach. What should you do?
  1. A Use Cloud Data Fusion to transfer the data into BigOuery raw tables, and use SQL to transform it.
  2. B Use Dataflow to transform the data and write the transformed data to BigQuery.
  3. C Manually download the data from Cloud Storage. Use a Python script to transform and upload the data into BigQuery.
  4. D Use Cloud Run functions to transform and load the data into BigOuery.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm

📘 Nội dung câu hỏi:
Câu hỏi mô tả một dự án phân tích dữ liệu mạng xã hội hàng ngày (daily social media data). Bạn có 100 GB dữ liệu định dạng JSON lưu trữ trong Cloud Storage, và dữ liệu này liên tục tăng trưởng (keeps growing). Nhiệm vụ là transform (chuyển đổi) và load (tải) dữ liệu này vào BigQuery để phân tích. Yêu cầu phải tuân thủ cách tiếp cận được Google khuyến nghị (Google-recommended approach).

🛠️ Bối cảnh kỹ thuật:

  • Dữ liệu lớn (100 GB+), định dạng JSON phức tạp, cần xử lý batch hoặc streaming.
  • Cloud Storage là nơi lưu trữ gốc, BigQuery là data warehouse cho phân tích SQL.
  • Google Cloud khuyến nghị sử dụng pipeline ETL/ELT tự động, scalable cho dữ liệu lớn và tăng trưởng, theo best practices từ Data Engineering trên Google Cloud (cập nhật đến 2026, với Dataflow hỗ trợ Apache Beam 2.58+ và tích hợp BigQuery IO native).

✅ Đáp án đúng:
Use Dataflow to transform the data and write the transformed data to BigQuery.

Lý do chọn đáp án đúng (chi tiết):
Dataflow là dịch vụ Apache Beam-managed được Google khuyến nghị chính thức cho các pipeline transform/load dữ liệu lớn từ Cloud Storage vào BigQuery. Nó hỗ trợ:

  • Xử lý batch/streaming JSON lớn (100 GB+ scalable tự động).
  • Transform linh hoạt (parse JSON, clean, enrich) bằng Beam SDK (Python/Java).
  • Write trực tiếp vào BigQuery với schema auto-detect/infer.
  • Tích hợp Pub/Sub cho dữ liệu real-time nếu cần.
    Theo tài liệu Google Cloud 2026, đây là ELT pipeline tiêu chuẩn cho BigQuery ingestion (xem Batch Dataflow for JSON to BigQuery guides).

📚 Tài liệu tham khảo:

🔍 Giải thích tất cả các phương án (đúng/sai)

  • Use Cloud Data Fusion to transfer the data into BigQuery raw tables, and use SQL to transform it.
    ❌ Sai. Cloud Data Fusion là công cụ no-code/low-code integration (dựa Dremel/Wrangler), phù hợp cho data integration phức tạp đa nguồn, nhưng không phải recommended cho simple JSON transform/load từ Cloud Storage vào BigQuery. Nó overhead cao (managed Hadoop), không scalable tốt cho 100 GB+ growing data, và transform chủ yếu qua SQL post-load kém hiệu quả. Google ưu tiên Dataflow cho ETL native.

  • Use Dataflow to transform the data and write the transformed data to BigQuery.
    ✅ Đúng. Như giải thích trên, đây là Google-recommended approach cho scalable, serverless ETL. Dataflow tự động scale, fault-tolerant, và tối ưu chi phí cho dữ liệu lớn JSON (hỗ trợ TextIO/AvroIO read từ GCS, BigQueryIO write). Hoàn hảo cho daily batch jobs.

  • Manually download the data from Cloud Storage. Use a Python script to transform and upload the data into BigQuery.
    ❌ Sai. Cách thủ công này không scalable cho 100 GB+ growing data: download/upload tốn thời gian/băng thông, dễ lỗi (memory overflow với JSON lớn), không tự động hóa, và vi phạm best practices (no orchestration). Google không khuyến nghị cho production; dùng Compute Engine/Cloud Functions chỉ cho POC nhỏ.

  • Use Cloud Run functions to transform and load the data into BigQuery.
    ❌ Sai. Cloud Run là serverless containers cho HTTP/microservices, không phù hợp cho batch transform dữ liệu lớn (giới hạn 32 GB RAM/container, timeout 60 phút, không native streaming). Nó kém hiệu quả cho 100 GB JSON (phải chunk thủ công), và Google recommend Dataflow thay thế cho data pipelines.

🎯 Kết luận: Chọn Dataflow để đảm bảo scalable, cost-effective, và tuân thủ Google best practices cho Data Lakehouse architecture (GCS + BigQuery). Nếu dữ liệu real-time, kết hợp Pub/Sub + Dataflow streaming! 🚀

Câu 322
You have a Cloud SQL for PostgreSQL database that stores sensitive historical financial data. You need to ensure that the data is uncorrupted and recoverable in the event that the primary region is destroyed. The data is valuable, so you need to prioritize recovery point objective (RPO) over recovery time objective (RTO). You want to recommend a solution that minimizes latency for primary read and write operations. What should you do?
  1. A Configure the Cloud SQL for PostgreSQL instance for multi-region backup locations.
  2. B Configure the Cloud SQL for PostgreSQL instance for regional availability (HA) with synchronous replication to a secondary instance in a different zone.
  3. C Configure the Cloud SQL for PostgreSQL instance for regional availability (HA) with asynchronous replication to a secondary instance in a different region.
  4. D Configure the Cloud SQL for PostgreSQL instance for regional availability (HA). Back up the Cloud SQL for PostgreSQL database hourly to a Cloud Storage bucket in a different region.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc bảo vệ dữ liệu nhạy cảm lịch sử tài chính trong cơ sở dữ liệu Cloud SQL for PostgreSQL (dịch vụ của Google Cloud). Yêu cầu chính là:

  • Đảm bảo dữ liệu không bị hỏng (uncorrupted) và có thể khôi phục (recoverable) nếu vùng chính (primary region) bị phá hủy hoàn toàn.
  • Ưu tiên RPO (Recovery Point Objective) hơn RTO (Recovery Time Objective): Nghĩa là chấp nhận thời gian khôi phục lâu hơn (RTO cao), nhưng phải giảm thiểu mất mát dữ liệu (RPO thấp, ví dụ dưới 1 giờ).
  • Giảm thiểu độ trễ (latency) cho các hoạt động đọc/ghi chính (primary read/write operations): Giải pháp không được ảnh hưởng đến hiệu suất của instance chính. Giải pháp cần cross-region để chống lại sự cố toàn vùng, nhưng vẫn giữ primary operations nhanh chóng. Theo tài liệu Google Cloud mới nhất (cập nhật 2025-2026), Cloud SQL sử dụng High Availability (HA) cho failover trong vùng và backup cho disaster recovery cross-region.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Configure the Cloud SQL for PostgreSQL instance for regional availability (HA). Back up the Cloud SQL for PostgreSQL database hourly to a Cloud Storage bucket in a different region.

🛠️ Lý do đúng:

  • Regional HA: Sử dụng synchronous replication sang instance phụ ở zone khác cùng region, đảm bảo zero data loss cho sự cố zone-level, không ảnh hưởng latency primary (failover <60 giây).
  • Backup hourly to GCS khác region: Tạo bản sao cross-region với RPO ≈ 1 giờ (ưu tiên RPO thấp), dữ liệu uncorrupted nhờ automated backup. Khôi phục bằng cách tạo instance mới từ backup GCS (RTO cao hơn nhưng chấp nhận được). Không ảnh hưởng primary latency vì backup là asynchronous.
  • Phù hợp hoàn hảo: Bảo vệ region failure, ưu tiên RPO, giữ hiệu suất primary.

❌ Phân tích tất cả các phương án

  • [SAI] Configure the Cloud SQL for PostgreSQL instance for multi-region backup locations.
    ❌ Sai vì: Cloud SQL không hỗ trợ trực tiếp "multi-region backup locations" cho automated backups (chúng lưu trong cùng region). Phải export thủ công sang GCS khác region, không tự động hourly và không đảm bảo RPO thấp. Không giải quyết region destruction đầy đủ.

  • [SAI] Configure the Cloud SQL for PostgreSQL instance for regional availability (HA) with synchronous replication to a secondary instance in a different zone.
    ❌ Sai vì: Regional HA chỉ sync replication trong cùng region, khác zone, bảo vệ zone failure chứ không chống region destruction. Không có cross-region sync, dẫn đến mất dữ liệu nếu primary region hỏng.

  • [SAI] Configure the Cloud SQL for PostgreSQL instance for regional availability (HA) with asynchronous replication to a secondary instance in a different region.
    ❌ Sai vì: Cloud SQL không hỗ trợ HA với async replication cross-region. HA chỉ sync trong region. Cross-region chỉ dùng read replicas (async), nhưng không phải HA config và gây RPO cao (lag giây/phút), tăng latency primary do network cross-region.

Giải pháp đúng cân bằng hoàn hảo giữa bảo mật, RPO và hiệu suất! 🚀

Câu 323
Your company has an on-premises file server with 5 TB of data that needs to be migrated to Google Cloud. The network operations team has mandated that you can only use up to 250 Mbps of the total available bandwidth for the migration. You need to perform an online migration to Cloud Storage. What should you do?
  1. A Use the gcloud storage cp command to copy all files from on-premises to Cloud Storage using the --daisy-chain option.
  2. B Use Storage Transfer Service to configure an agent-based transfer. Set the appropriate bandwidth limit for the agent pool.
  3. C Request a Transfer Appliance, copy the data to the appliance, and ship it back to Google Cloud.
  4. D Use the gcloud storage cp command to copy all files from on-premises to Cloud Storage using the --no-clobber option.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống công ty bạn có một máy chủ file on-premises chứa 5 TB dữ liệu cần di chuyển (migrate) lên Google Cloud Storage. Đội ngũ mạng (network operations team) yêu cầu giới hạn băng thông tối đa chỉ 250 Mbps trong quá trình di chuyển. Bạn phải thực hiện online migration (di chuyển trực tuyến qua mạng, không phải ngoại tuyến).

📌 Yêu cầu chính:

  • Phải là online migration → Sử dụng kết nối mạng từ on-premises đến Cloud Storage.
  • Tuân thủ giới hạn băng thông 250 Mbps → Cần công cụ hỗ trợ cấu hình limit bandwidth.
  • Dữ liệu lớn (5 TB) → Cần giải pháp hiệu quả, đáng tin cậy cho transfer lớn.

🛠️ Bối cảnh kiến thức Google Cloud (cập nhật đến 2026): Google Cloud cung cấp các công cụ như Storage Transfer Service (STS) cho online transfer từ on-premises, hỗ trợ agent-based để kiểm soát bandwidth. Transfer Appliance là offline (ship hardware). gsutil/gcloud storage cp là CLI cơ bản nhưng thiếu tùy chọn limit bandwidth tinh chỉnh.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Storage Transfer Service to configure an agent-based transfer. Set the appropriate bandwidth limit for the agent pool.

Lý do:

  • Storage Transfer Service (STS) là dịch vụ chuyên biệt cho online migration lớn từ on-premises đến Cloud Storage, hỗ trợ agent-based transfer (cài agent trên máy on-premises để push/pull dữ liệu an toàn).
  • Có thể cấu hình bandwidth limit chính xác cho agent pool (từ 1 Mbps đến hàng Gbps, dễ set 250 Mbps qua console/CLI/API).
  • Phù hợp 5 TB dữ liệu, resumeable, monitoring tốt, và tuân thủ chính sách mạng. Đây là best practice theo docs Google Cloud 2026.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên tính phù hợp với online migration và giới hạn 250 Mbps.

  • Use the gcloud storage cp command to copy all files from on-premises to Cloud Storage using the --daisy-chain option.
    ❌ Sai: Lệnh gcloud storage cp (hay gsutil cp) là CLI đơn giản cho copy file, không có tùy chọn --daisy-chain (tùy chọn này không tồn tại trong gsutil/STS đến 2026). Không hỗ trợ giới hạn bandwidth tinh chỉnh (chỉ có --parallel-process-count cơ bản), dễ vượt 250 Mbps và thiếu resume cho 5 TB lớn. Không phải giải pháp enterprise cho migration kiểm soát.

  • Use Storage Transfer Service to configure an agent-based transfer. Set the appropriate bandwidth limit for the agent pool.
    ✅ Đúng: Như giải thích trên, STS agent-based cho phép set bandwidth limit trực tiếp trên agent pool (qua console: Networking > Bandwidth throttle), lý tưởng cho online migration với constraint mạng. Hỗ trợ 5 TB hiệu quả, scalable, và an toàn (HTTPS, checksum).

  • Request a Transfer Appliance, copy the data to the appliance, and ship it back to Google Cloud.
    ❌ Sai: Transfer Appliance là giải pháp offline (ship thiết bị cứng chứa dữ liệu về Google), không phải online migration qua mạng. Không liên quan đến bandwidth limit (vì không dùng mạng), vi phạm yêu cầu "online" và đội ngũ mạng.

  • Use the gcloud storage cp command to copy all files from on-premises to Cloud Storage using the --no-clobber option.
    ❌ Sai: --no-clobber chỉ tránh ghi đè file (tùy chọn gsutil cơ bản), không kiểm soát bandwidth (dễ vượt 250 Mbps). Không scalable cho 5 TB (thiếu parallelism mạnh, monitoring), và không phải công cụ chuyên migration online lớn.

📘 Tài liệu tham khảo (Google Cloud docs cập nhật 2026)

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần thêm ví dụ CLI, hãy hỏi nhé!

Câu 324
Following a recent company acquisition, you inherited an on-premises data infrastructure that needs to move to Google Cloud. The acquired system has 250 Apache Airflow directed acyclic graphs (DAGs) orchestrating data pipelines. You need to migrate the pipelines to a Google Cloud managed service with minimal effort. What should you do?
  1. A Create a Google Kubernetes Engine (GKE) standard cluster and deploy Airflow as a workload. Migrate all DAGs to the new Airflow environment.
  2. B Create a Cloud Data Fusion instance. For each DAG, create a Cloud Data Fusion pipeline.
  3. C Create a new Cloud Composer environment and copy DAGs to the Cloud Composer dags/ folder.
  4. D Convert each DAG to a Cloud Workflow and automate the execution with Cloud Scheduler.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi mô tả tình huống sau khi công ty thực hiện mua lại một công ty khác, bạn kế thừa hệ thống dữ liệu on-premises (chạy tại chỗ, không phải cloud) sử dụng Apache Airflow với 250 Directed Acyclic Graphs (DAGs) để điều phối các pipeline dữ liệu. Nhiệm vụ là migrate (di chuyển) toàn bộ các pipeline này sang một dịch vụ được quản lý (managed service) của Google Cloud với nỗ lực tối thiểu (minimal effort).

🛠️ Yêu cầu chính:

  • Phải là dịch vụ managed (Google Cloud tự quản lý infrastructure, scaling, update, v.v.).
  • Minimal effort: Không cần rebuild, convert lớn, chỉ copy hoặc deploy đơn giản.
  • Airflow DAGs là các workflow định sẵn, cần giữ nguyên logic để chạy seamless.

Đây là câu hỏi kiểm tra kiến thức về Cloud Composer – dịch vụ managed Apache Airflow trên Google Cloud (cập nhật đến 2026: Cloud Composer 3 sử dụng Airflow 2.9+, hỗ trợ auto-scaling, integration với BigQuery, Pub/Sub...).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create a new Cloud Composer environment and copy DAGs to the Cloud Composer dags/ folder.

Lý do (🧩 Phân tích chi tiết):

  • Cloud Composer là dịch vụ fully managed Apache Airflow của Google Cloud, tự động xử lý Kubernetes, scaling, monitoring, security (IAM, VPC-SC).
  • Với minimal effort: Chỉ cần tạo environment mới (qua Console/CLI/gcloud), sau đó copy trực tiếp các DAG files vào thư mục dags/ (qua Cloud Storage bucket gắn với Composer). DAGs sẽ tự sync và chạy mà không cần chỉnh sửa code (hỗ trợ Python deps, plugins).
  • Hỗ trợ 250 DAGs lớn: Auto-scaling workers, scheduler HA. Thời gian migrate: Giờ thay vì tuần/tháng.
  • Cập nhật 2026: Composer 3+ hỗ trợ Airflow 2.10+, GKE Autopilot, giảm chi phí 40% so với self-managed.

❌ Phân tích tất cả các phương án

Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên tiêu chí managed service và minimal effort:

  • [SAI] Create a Google Kubernetes Engine (GKE) standard cluster and deploy Airflow as a workload. Migrate all DAGs to the new Airflow environment.
    ❌ Sai vì: Đây không phải managed service cho Airflow (bạn phải tự deploy Helm chart, quản lý GKE cluster, scaling, patching Airflow). Với 250 DAGs, effort cao: Cài đặt Postgres/Fernet, config secrets, monitoring – trái ngược "minimal effort". GKE chỉ là infra, không optimize cho Airflow như Composer.

  • [SAI] Create a Cloud Data Fusion instance. For each DAG, create a Cloud Data Fusion pipeline.
    ❌ Sai vì: Cloud Data Fusion (dựa trên CDAP) là cho ETL pipelines (drag-drop, no-code/low-code), không tương thích trực tiếp với Airflow DAGs (Python-based workflows). Phải rebuild từng DAG thành pipeline mới (250 cái = effort khổng lồ, mất tuần/tháng). Không phải Airflow managed, chỉ phù hợp data integration đơn giản.

  • [ĐÚNG] Create a new Cloud Composer environment and copy DAGs to the Cloud Composer dags/ folder.
    ✅ Đúng vì: Như giải thích trên – managed Airflow native, copy DAGs trực tiếp, zero-downtime migrate. Hỗ trợ plugins, variables, XComs đầy đủ.

  • [SAI] Convert each DAG to a Cloud Workflow and automate the execution with Cloud Scheduler.
    ❌ Sai vì: Cloud Workflows là YAML-based orchestration (cho serverless steps như API calls), không hỗ trợ Airflow DAGs (phải convert thủ công từng DAG thành YAML – effort cực cao cho 250 cái). Cloud Scheduler chỉ trigger, không thay thế Airflow scheduler. Không managed cho complex pipelines với operators tùy chỉnh.

🛡️ Lưu ý cuối: Chọn Composer để tối ưu chi phí (pay-per-use), bảo mật (private IP, CMEK), và tích hợp ecosystem (Dataflow, Dataproc). Nếu DAGs legacy, dùng Migration Toolkit for Composer!

Câu 325
You created a curated dataset of market trends in BigQuery that you want to share with multiple external partners. You want to control the rows and columns that each partner has access to. You want to follow Google-recommended practices. What should you do?
  1. A Publish the dataset in Analytics Hub. Grant dataset-level access to each partner by using subscriptions.
  2. B Grant each partner read access to the BigQuery dataset by using IAM roles.
  3. C Create a separate Cloud Storage bucket for each partner. Export the dataset to each bucket and assign each partner to their respective bucket. Grant bucketlevel access by using IAM roles.
  4. D Create a separate project for each partner and copy the dataset into each project. Publish each dataset in Analytics Hub. Grant dataset-level access to each partner by using subscriptions.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi tập trung vào tình huống thực tế trong Google Cloud BigQuery: Bạn đã tạo một curated dataset (dataset được chọn lọc và chuẩn hóa) về xu hướng thị trường (market trends). Mục tiêu là chia sẻ dataset này với nhiều đối tác bên ngoài (external partners), nhưng cần kiểm soát chính xác các hàng (rows) và cột (columns) mà mỗi đối tác có thể truy cập. Đồng thời, phải tuân thủ best practices được Google khuyến nghị (Google-recommended practices).

Vấn đề cốt lõi là chia sẻ dữ liệu an toàn, linh hoạt với granular access control (kiểm soát ở mức chi tiết), tránh rò rỉ dữ liệu nhạy cảm, và tối ưu chi phí/quản lý. Đây là kịch bản phổ biến trong data sharing trên Google Cloud, đặc biệt với Analytics Hub – dịch vụ chuyên biệt cho việc xuất bản và chia sẻ dữ liệu lớn.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Publish the dataset in Analytics Hub. Grant dataset-level access to each partner by using subscriptions.

Lý do chi tiết 🛠️:

  • Analytics Hub là dịch vụ chính thức của Google Cloud để publish (xuất bản) dataset từ BigQuery, cho phép chia sẻ với external partners (ngay cả ngoài tổ chức của bạn) qua subscriptions (đăng ký truy cập).
  • Mỗi subscription có thể được cấp dataset-level access riêng biệt, hỗ trợ row-level và column-level security thông qua authorized views, row access policies, hoặc column-level security (tính năng mới nhất từ 2023-2026).
  • Đây là Google-recommended practice vì: (1) Không cần copy dữ liệu (tiết kiệm storage), (2) Kiểm soát granular (mỗi partner chỉ thấy rows/columns được phép), (3) Tích hợp IAM an toàn, audit logs đầy đủ, và hỗ trợ linked datasets để truy cập live data.
  • Phù hợp với kiến thức cập nhật đến 2026: Analytics Hub v2+ hỗ trợ private listings và multi-region sharing với zero-copy semantics.

📋 Giải thích tất cả các phương án (đúng và sai)

Dưới đây là phân tích từng phương án một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi giải thích tập trung vào lý do đúng/sai dựa trên best practices Google Cloud (BigQuery & Analytics Hub, phiên bản mới nhất 2026).

  • ✅ [ĐÚNG] Publish the dataset in Analytics Hub. Grant dataset-level access to each partner by using subscriptions.
    Giải thích: Phương án này hoàn hảo vì Analytics Hub được thiết kế dành riêng cho data sharing với external partners, hỗ trợ granular control trên rows/columns qua subscriptions. Không tốn kém copy data, tuân thủ zero-ETL, và là recommended practice trong docs Google (ví dụ: Publisher có thể định nghĩa views với RLS/CLS cho từng subscriber). ✅ Hoàn toàn khớp yêu cầu!

  • ❌ [SAI] Grant each partner read access to the BigQuery dataset by using IAM roles.
    Giải thích: Sai vì IAM roles chỉ cấp dataset-level hoặc table-level read access chung chung, không hỗ trợ kiểm soát rows/columns chi tiết cho từng partner. External partners cần Google account hoặc federated identity, dễ dẫn đến over-privileging (quyền thừa) và rủi ro bảo mật. Không phải best practice cho multi-partner sharing. ❌ Thiếu granular control!

  • ❌ [SAI] Create a separate Cloud Storage bucket for each partner. Export the dataset to each bucket and assign each partner to their respective bucket. Grant bucketlevel access by using IAM roles.
    Giải thích: Sai vì phải export dataset từ BigQuery sang GCS (tốn thời gian, chi phí, và mất tính live data). Mỗi bucket chỉ cấp bucket-level access, không kiểm soát rows/columns (GCS là object storage, không phải query engine như BigQuery). Phức tạp quản lý nhiều bucket, không scalable cho data lớn, và vi phạm best practice (Google ưu tiên in-place sharing). ❌ Tốn kém và thiếu precision!

  • ❌ [SAI] Create a separate project for each partner and copy the dataset into each project. Publish each dataset in Analytics Hub. Grant dataset-level access to each partner by using subscriptions.
    Giải thích: Sai vì yêu cầu tạo project riêng và copy dataset (tốn storage gấp nhiều lần, chi phí cao, và duplicate data khó sync). Dù dùng Analytics Hub, việc copy làm phức tạp hóa quản lý (multi-project overhead). Best practice là single publish từ project gốc với subscriptions đa dạng, không cần copy. ❌ Không hiệu quả và không recommended!

📘 Tài liệu tham khảo (cập nhật đến 2026)

  • Google Cloud Docs - Analytics Hub: cloud.google.com/bigquery/docs/analytics-hub-publisher – Hướng dẫn publish & subscriptions với row/column security.
  • BigQuery Security Best Practices: cloud.google.com/bigquery/docs/best-practices-security – Nhấn mạnh Analytics Hub cho external sharing.
  • What's New in BigQuery (2025-2026): Column-level security native support và Analytics Hub enhancements (xem BigQuery release notes).
  • Certification Guide (Associate Cloud Data Engineer): Khuyến nghị Analytics Hub cho data marketplaces.

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần thêm ví dụ code hoặc demo, hãy hỏi nhé!

Câu 326
Your company wants to implement a data transformation (ETL) pipeline for their BigQuery data warehouse. You need to identify a managed transformation solution that allows users to develop with SQL and JavaScript, has version control, allows for modular code, and has data quality checks. What should you do?
  1. A Use Dataform to define the transformations in SQLX.
  2. B Use Dataproc to create an Apache Spark cluster and implement the transformations by using PySpark SQL.
  3. C Create a Cloud Composer environment, and orchestrate the transformations by using the BigQueryInsertJob operator.
  4. D Create BigQuery scheduled queries to define the transformations in SQL.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này tập trung vào việc triển khai một pipeline chuyển đổi dữ liệu (ETL) cho kho dữ liệu BigQuery trên Google Cloud Platform (GCP). Công ty cần một giải pháp managed transformation (quản lý hoàn toàn bởi GCP) với các yêu cầu cụ thể sau:

  • Cho phép phát triển bằng SQL và JavaScript.
  • Hỗ trợ version control (kiểm soát phiên bản mã nguồn).
  • Cho phép modular code (mã nguồn mô-đun hóa, tái sử dụng).
  • Có data quality checks (kiểm tra chất lượng dữ liệu).

📘 Mục tiêu chính: Tìm giải pháp phù hợp nhất để định nghĩa các chuyển đổi dữ liệu (transformations) một cách hiệu quả, an toàn và scalable trên BigQuery. Đây là câu hỏi kiểm tra kiến thức về các công cụ ETL managed trên GCP (cập nhật đến năm 2026, theo tài liệu chính thức GCP Dataform v2.x và BigQuery ecosystem).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Dataform to define the transformations in SQLX.

🛠️ Lý do chi tiết:
Dataform là dịch vụ managed transformation chuyên biệt cho BigQuery, sử dụng ngôn ngữ SQLX (kết hợp SQL + JavaScript) để phát triển pipeline ETL. Nó hỗ trợ đầy đủ:

  • Version control: Tích hợp Git repositories (GitHub, GitLab, Cloud Source Repositories).
  • Modular code: Sử dụng SQLX files, operations, và dependencies để xây dựng mã mô-đun (như JavaScript functions cho logic phức tạp).
  • Data quality checks: Assertions và tests tự động (ví dụ: kiểm tra NULLs, duplicates, schema validation).
    Dataform compile thành SQL native cho BigQuery, chạy serverless, và là lựa chọn tối ưu theo best practices GCP 2026.

Nguồn tham khảo:

📋 Giải thích tất cả các phương án (đúng/sai)

  • Use Dataform to define the transformations in SQLX.
    ✅ Đúng – Như đã giải thích ở trên, Dataform khớp 100% yêu cầu: SQL/JS (SQLX), version control qua Git, modular code (operations/declarations), và data quality checks (assertions). Đây là giải pháp managed lý tưởng cho BigQuery ETL.

  • Use Dataproc to create an Apache Spark cluster and implement the transformations by using PySpark SQL.
    ❌ Sai – Dataproc là dịch vụ managed cho Spark/Hadoop, hỗ trợ PySpark SQL cho ETL lớn, nhưng không phải managed transformation thuần túy cho BigQuery. Nó yêu cầu cluster management (không serverless hoàn toàn), không hỗ trợ SQL/JavaScript native (chủ yếu Python/Scala), thiếu version control tích hợp và data quality checks chuyên biệt. Không phù hợp cho phát triển SQL/JS modular.

  • Create a Cloud Composer environment, and orchestrate the transformations by using the BigQueryInsertJob operator.
    ❌ Sai – Cloud Composer (Managed Apache Airflow) giỏi orchestration (lập lịch workflow), sử dụng BigQueryInsertJob để chèn dữ liệu, nhưng không phải transformation tool. Nó thiếu SQL/JavaScript development, version control cho code transform (chỉ DAGs), modular code hạn chế, và không có data quality checks built-in. Phù hợp orchestrate hơn là define transformations.

  • Create BigQuery scheduled queries to define the transformations in SQL.
    ❌ Sai – BigQuery Scheduled Queries chỉ hỗ trợ SQL đơn giản để lập lịch chạy queries, rất cơ bản cho ETL nhẹ. Thiếu hoàn toàn: JavaScript support, version control (không Git), modular code (không tái sử dụng modules), và data quality checks (phải code thủ công). Không scalable cho pipeline phức tạp.

🧩 Kết luận: Dataform là lựa chọn tối ưu và managed nhất theo roadmap GCP 2026, giúp giảm chi phí vận hành so với các option tự quản lý như Dataproc hay Composer. Nếu triển khai thực tế, bắt đầu bằng tạo Dataform repository liên kết BigQuery! 🚀

Câu 327
Your team uses Google Sheets to track budget data that is updated daily. The team wants to compare budget data against actual cost data, which is stored in a BigQuery table. You need to create a solution that calculates the difference between each day's budget and actual costs. You want to ensure that your team has access to daily-updated results in Google Sheets. What should you do?
  1. A Download the budget data as a CSV file and upload the CSV file to a Cloud Storage bucket. Create a new BigQuery table from Cloud Storage, and join the actual cost table with it. Open the joined BigOuery table by using Connected Sheets.
  2. B Create a BigQuery external table by using the Drive URI of the Google sheet, and join the actual cost table with it. Save the joined table, and open it by using Connected Sheets.
  3. C Download the budget data as a CSV file, and upload the CSV file to create a new BigQuery table. Join the actual cost table with the new BigQuery table, and save the results as a CSV file. Open the CSV file in Google Sheets.
  4. D Create a BigQuery external table by using the Drive URI of the Google sheet, and join the actual cost table with it. Save the joined table as a CSV file and open the file in Google Sheets.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả tình huống: Nhóm của bạn sử dụng Google Sheets để theo dõi dữ liệu ngân sách (budget data), được cập nhật hàng ngày. Dữ liệu chi phí thực tế (actual cost data) được lưu trữ trong một bảng BigQuery. Nhiệm vụ là tạo giải pháp tính toán sự chênh lệch giữa ngân sách và chi phí thực tế mỗi ngày, đồng thời đảm bảo nhóm có thể truy cập kết quả cập nhật hàng ngày trực tiếp trong Google Sheets mà không cần thao tác thủ công lặp lại.

📌 Yêu cầu chính:

  • Tích hợp dữ liệu từ Google Sheets (cập nhật tự động hàng ngày) với BigQuery.
  • Tính toán chênh lệch qua phép JOIN giữa hai bộ dữ liệu.
  • Kết quả phải hiển thị tự động cập nhật trong Google Sheets (sử dụng tính năng Connected Sheets của BigQuery để liên kết trực tiếp, không cần export/import thủ công).

🛠️ Giải pháp lý tưởng: Sử dụng BigQuery external table từ Google Sheets (qua Drive URI như gsheets://), JOIN với bảng actual cost, lưu kết quả dưới dạng view hoặc table, rồi mở qua Connected Sheets để Sheets tự động refresh dữ liệu khi Sheets gốc thay đổi.

✅ Đáp án đúng

Create a BigQuery external table by using the Drive URI of the Google sheet, and join the actual cost table with it. Save the joined table, and open it by using Connected Sheets.

Lý do lựa chọn:

  • ✅ External table từ Drive URI (ví dụ: gsheets://drive.google.com/...) cho phép BigQuery query trực tiếp dữ liệu từ Google Sheets mà không cần import, nên tự động phản ánh thay đổi hàng ngày trên Sheets gốc.
  • ✅ JOIN với bảng actual cost (bảng native trong BigQuery) để tính chênh lệch, sau đó lưu joined table (có thể là view hoặc table) đảm bảo query có thể tái sử dụng.
  • ✅ Connected Sheets cho phép mở bảng BigQuery trực tiếp trong Google Sheets với tự động refresh (real-time hoặc scheduled), đáp ứng yêu cầu "daily-updated results in Google Sheets" mà không cần tải/xuất file thủ công.
  • 🛠️ Đây là cách tối ưu, tự động hóa cao, phù hợp với best practice của Google Cloud (phiên bản mới nhất 2024-2026 hỗ trợ Connected Sheets với external tables từ Drive).

📘 Giải thích chi tiết từng phương án

Dưới đây là phân tích tất cả 4 phương án, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai kèm lý do cụ thể bằng tiếng Việt:

  • Download the CSV file and upload the CSV file to a Cloud Storage bucket. Create a new BigQuery table from Cloud Storage, and join the actual cost table with it. Open the joined BigOuery table by using Connected Sheets.
    ❌ Sai: Yêu cầu tải CSV thủ công từ Sheets và upload lên Cloud Storage, tạo bảng mới – quá trình này không tự động, phải lặp lại hàng ngày. Mặc dù dùng Connected Sheets ở cuối, nhưng dữ liệu budget đã "đông cứng" sau khi import, không reflect cập nhật Sheets gốc. Không đáp ứng "daily-updated".

  • Create a BigQuery external table by using the Drive URI of the Google sheet, and join the actual cost actual cost table with it. Save the joined table, and open it by using Connected Sheets.
    ✅ Đúng: Như giải thích ở trên. External table từ Drive URI đảm bảo tự động sync với Sheets, JOIN và lưu kết quả, Connected Sheets hiển thị real-time trong Sheets. Hoàn hảo cho yêu cầu tự động hóa.

  • Download the budget data as a CSV file, and upload the CSV file to create a new BigQuery table. Join the actual cost table with the new BigQuery table, and save the results as a CSV file. Open the CSV file in Google Sheets.
    ❌ Sai: Toàn bộ quy trình thủ công kép (tải CSV budget → import BigQuery → JOIN → xuất CSV kết quả → mở Sheets). Không tự động cập nhật hàng ngày, phải làm lại mỗi ngày. Không dùng Connected Sheets, vi phạm yêu cầu "daily-updated results in Google Sheets".

  • Create a BigQuery external table by using the Drive URI of the Google sheet, and join the actual cost table with it. Save the joined table as a CSV file and open the file in Google Sheets.
    ❌ Sai: Phần external table và JOIN tốt (tự động sync budget), nhưng xuất CSV thủ công thay vì Connected Sheets làm mất tính tự động. Kết quả CSV không refresh theo Sheets gốc, phải export lại hàng ngày – không hiệu quả.

📚 Tài liệu tham khảo (cập nhật mới nhất đến 2026)

🧩 Kết luận: Giải pháp đúng tận dụng tích hợp native GCP để zero-ETL, tự động hoàn toàn! Nếu cần demo code SQL, hãy cho biết thêm chi tiết. 🚀

Câu 328
Your organization's website uses an on-premises MySQL as a backend database. You need to migrate the on-premises MySQL database to Google Cloud while maintaining MySQL features. You want to minimize administrative overhead and downtime. What should you do?
  1. A Use a Google-provided Dataflow template to replicate the MySQL database in BigOuery.
  2. B Install MySQL on a Compute Engine virtual machine. Export the database files using the mysqldump command. Upload the files to Cloud Storage, and import them into the MySQL instance on Compute Engine.
  3. C Use Database Migration Service to transfer the data to Cloud SQL for MySQL, and configure the on-premises MySQL database as the source.
  4. D Export the database tables to CSV files, and upload the files to Cloud Storage. Convert the MySQL schema to a Spanner schema, create a JSON manifest file, and run a Google-provided Dataflow template to load the data into Spanner.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc di chuyển (migrate) cơ sở dữ liệu MySQL đang chạy on-premises (tại chỗ) sang Google Cloud, với các yêu cầu chính:

  • Giữ nguyên các tính năng của MySQL (như engine, syntax, stored procedures, v.v.).
  • Giảm thiểu overhead quản trị (không muốn tự quản lý server, patching, scaling).
  • Giảm thiểu thời gian downtime (migrate với ít gián đoạn nhất có thể).

Tình huống: Website của tổ chức sử dụng MySQL on-premises làm backend. Giải pháp cần managed service để dễ quản lý, hỗ trợ replication liên tục (continuous migration) nhằm tránh downtime lớn. Đây là kịch bản phổ biến trong Google Cloud Data Migration, sử dụng các dịch vụ như Cloud SQL (managed MySQL) và Database Migration Service (DMS).
📘 Kiến thức cập nhật: Theo tài liệu Google Cloud năm 2026 (phiên bản DMS 2.0+), DMS hỗ trợ MySQL on-premises → Cloud SQL for MySQL với change data capture (CDC), cho phép migrate live data mà không downtime.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Database Migration Service to transfer the data to Cloud SQL for MySQL, and configure the on-premises MySQL database as the source.

Lý do:

  • Database Migration Service (DMS) là dịch vụ fully managed của Google Cloud, chuyên migrate databases với minimal downtime qua one-time migration hoặc continuous replication (CDC).
  • Cloud SQL for MySQL là dịch vụ managed MySQL, giữ nguyên 100% tính năng MySQL (InnoDB, replication, GTID, v.v.), không cần quản lý OS/server.
  • Phù hợp hoàn hảo: Giảm admin overhead (Google lo HA, backup, scaling) và downtime gần zero (promote Cloud SQL làm primary sau sync).
  • 🛠️ Quy trình: Cài DMS agent trên on-premises, config source MySQL → target Cloud SQL, run migration job.
    📘 Nguồn: Google Cloud DMS Documentation & Cloud SQL for MySQL.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ✅ [ĐÚNG] Use Database Migration Service to transfer the data to Cloud SQL for MySQL, and configure the on-premises MySQL database as the source.
    Như đã giải thích ở trên: DMS + Cloud SQL là giải pháp optimized cho migrate MySQL managed, giữ nguyên features, low overhead & downtime. Hoàn hảo match yêu cầu!

  • ❌ [SAI] Use a Google-provided Dataflow template to replicate the MySQL database in BigQuery.
    Lý do sai: BigQuery là data warehouse columnar (không phải transactional RDBMS như MySQL), mất hết tính năng MySQL (index, transactions ACID, queries OLTP). Dataflow template chỉ dùng cho batch ETL (export/import), gây downtime lớn và high overhead (cần transform schema). Không phù hợp migrate live database giữ features.

  • ❌ [SAI] Install MySQL on a Compute Engine virtual machine. Export the database files using the mysqldump command. Upload the files to Cloud Storage, and import them into the MySQL instance on Compute Engine.
    Lý do sai: Đây là cách manual/self-managed trên Compute Engine VM → high admin overhead (tự install, patch, monitor, scale HA). Mysqldump gây downtime lớn (export/import full DB), không hỗ trợ continuous sync. Không dùng managed service như Cloud SQL, vi phạm yêu cầu minimize overhead/downtime.

  • ❌ [SAI] Export the database tables to CSV files, and upload the files to Cloud Storage. Convert the MySQL schema to a Spanner schema, create a JSON manifest file, and run a Google-provided Dataflow template to load the data into Spanner.
    Lý do sai: Cloud Spanner là distributed NewSQL (không phải MySQL engine), yêu cầu convert schema phức tạp (Spanner dùng SQL ANSI, khác syntax MySQL). CSV export/Dataflow chỉ batch load, gây downtime cao, mất features MySQL-specific. Overhead lớn (custom schema + manifest), không giữ nguyên MySQL features.

🧩 Tóm tắt: DMS + Cloud SQL là lựa chọn best practice cho migrate MySQL on-premises sang Google Cloud managed! Nếu cần lab thực hành, thử free tier trên Google Cloud Console. 📘

Câu 329
Your retail company wants to analyze customer reviews to understand sentiment and identify areas for improvement. Your company has a large dataset of customer feedback text stored in BigQuery that includes diverse language patterns, emojis, and slang. You want to build a solution to classify customer sentiment from the feedback text. What should you do?
  1. A Preprocess the text data in BigQuery using SQL functions. Export the processed data to AutoML Natural Language for model training and deployment.
  2. B Develop a custom sentiment analysis model using TensorFlow. Deploy it on a Compute Engine instance.
  3. C Use Dataproc to create a Spark cluster, perform text preprocessing using Spark NLP, and build a sentiment analysis model with Spark MLlib.
  4. D Export the raw data from BigQuery. Use AutoML Natural Language to train a custom sentiment analysis model.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một công ty bán lẻ muốn phân tích cảm xúc (sentiment analysis) từ đánh giá khách hàng được lưu trữ trong BigQuery. Dữ liệu lớn, bao gồm đa ngôn ngữ, emoji và slang, cần xây dựng giải pháp phân loại cảm xúc từ văn bản phản hồi.
Mục tiêu chính: Xây dựng mô hình phân loại sentiment một cách hiệu quả, tận dụng các dịch vụ Google Cloud (GCP) để xử lý dữ liệu thô và huấn luyện mô hình.
🛠️ Thách thức: Dữ liệu "dơ" (noisy) với emoji, slang cần tiền xử lý (preprocessing) trước khi huấn luyện mô hình ML. Giải pháp phải đơn giản, scalable và managed (không tự code phức tạp).
📘 Kiến thức cập nhật (GCP 2026): BigQuery hỗ trợ SQL functions mạnh mẽ cho text processing (như REGEXP_REPLACE, NORMALIZE). AutoML Natural Language (nay tích hợp Vertex AI) lý tưởng cho custom sentiment models trên text đa ngôn ngữ, nhưng yêu cầu data sạch (không emoji/sl ang raw).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Preprocess the text data in BigQuery using SQL functions. Export the processed data to AutoML Natural Language for model training and deployment.

Lý do chi tiết 🏆:

  • Preprocess ngay trong BigQuery bằng SQL (ví dụ: REGEXP_REPLACE loại bỏ emoji, LOWER() chuẩn hóa chữ thường, TRANSLATE() xử lý slang) giúp tiết kiệm chi phí, nhanh chóng vì BigQuery serverless và scalable cho big data.
  • Sau đó export sang AutoML Natural Language (Vertex AI Sentiment Analysis) để train/deploy custom model: Dịch vụ managed, hỗ trợ đa ngôn ngữ, tự động handle patterns phức tạp sau khi clean. Không cần code thủ công!
  • Phù hợp nhất với yêu cầu: Xử lý noisy data trước → Train/deploy dễ dàng, chi phí thấp (pay-per-use).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên tính khả thi, hiệu quả và best practice GCP (cập nhật Vertex AI 2026).

  • ✅ Preprocess the text data in BigQuery using SQL functions. Export the processed data to AutoML Natural Language for model training and deployment.
    Đúng vì: Như đã giải thích trên – kết hợp BigQuery SQL cho preprocessing (rẻ, nhanh) + AutoML NL managed cho ML. Best practice cho text analytics trên GCP, tránh over-engineering.

  • ❌ Develop a custom sentiment analysis model using TensorFlow. Deploy it on a Compute Engine instance.
    Sai vì: Quá phức tạp và tốn kém – tự code TensorFlow từ đầu (cần expertise NLP), deploy trên Compute Engine VM (quản lý thủ công scaling, patching). Không tận dụng managed services như Vertex AI, dễ lỗi với dữ liệu noisy đa ngôn ngữ.

  • ❌ Use Dataproc to create a Spark cluster, perform text preprocessing using Spark NLP, và build a sentiment analysis model with Spark MLlib.
    Sai vì: Overkill cho task này – Dataproc/Spark phù hợp big data batch phức tạp, nhưng tốn kém (cluster spin-up), thời gian setup lâu. Spark NLP/MLlib mạnh nhưng không managed như AutoML, và preprocessing Spark không hiệu quả bằng BigQuery SQL cho text đơn giản.

  • ❌ Export the raw data from BigQuery. Use AutoML Natural Language to train a custom sentiment analysis model.
    Sai vì: Bỏ qua preprocessing – Dữ liệu raw (emoji, slang, đa ngôn ngữ) sẽ làm AutoML NL train kém hiệu quả (accuracy thấp, noise ảnh hưởng model). AutoML yêu cầu data sạch (docs khuyến nghị clean trước); export raw lãng phí và fail best practice.

📚 Tài liệu tham khảo (GCP chính thức, cập nhật 2026)

Giải pháp này giúp công ty bạn triển khai nhanh, scalable! 🚀 Nếu cần code mẫu SQL, hỏi thêm nhé!

Câu 330
You need to transfer approximately 300 TB of data from your company's on-premises data center to Cloud Storage. You have 100 Mbps internet bandwidth, and the transfer needs to be completed as quickly as possible. What should you do?
  1. A Use Cloud Client Libraries to transfer the data over the internet.
  2. B Compress the data, upload it to multiple cloud storage providers, and then transfer the data to Cloud Storage.
  3. C Request a Transfer Appliance, copy the data to the appliance, and ship it back to Google.
  4. D Use the gcloud storage command to transfer the data over the internet.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi yêu cầu tìm giải pháp tối ưu nhất để chuyển khoảng 300 TB dữ liệu từ data center on-premises (tại chỗ của công ty) sang Cloud Storage (dịch vụ lưu trữ của Google Cloud). Các ràng buộc chính:

  • Băng thông internet chỉ 100 Mbps (tương đương khoảng 12.5 MB/s thực tế sau overhead).
  • Yêu cầu hoàn thành nhanh nhất có thể (as quickly as possible).

Tính toán thời gian ước tính nếu dùng internet:
300 TB = 300 × 1024 GB ≈ 307.200 GB.
Với 100 Mbps, tốc độ thực tế ~10-12 MB/s → Thời gian ≈ 27-30 ngày (quá chậm cho nhu cầu khẩn cấp).

Vì vậy, cần giải pháp offline/physical transfer thay vì online để vượt qua hạn chế băng thông. Kiến thức dựa trên Google Cloud Transfer Appliance (cập nhật đến 2026: vẫn là lựa chọn hàng đầu cho petabyte-scale data với bandwidth thấp, theo docs Google Cloud 2024+).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Request a Transfer Appliance, copy the data to the appliance, and ship it back to Google.

Lý do:
🛠️ Transfer Appliance là thiết bị phần cứng chuyên dụng của Google (dung lượng lên đến 100-480 PB tùy model mới nhất 2026), cho phép copy dữ liệu on-premises qua mạng nội bộ tốc độ cao (fiber/NAS), sau đó ship về Google để upload trực tiếp vào Cloud Storage.

  • Thời gian: Copy on-prem ~vài ngày (tùy tốc độ nội bộ), ship ~3-7 ngày, upload tại Google ~1-2 ngày → Tổng <2 tuần, nhanh hơn hàng chục lần so với internet.
  • Phù hợp hoàn hảo với 300 TB + bandwidth thấp. Đây là best practice chính thức của Google cho dữ liệu lớn (>100 TB).

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Use Cloud Client Libraries to transfer the data over the internet.
    Phương án này dùng thư viện client (như gsutil hoặc API) để upload qua internet. Sai vì: Với 100 Mbps, thời gian >27 ngày, không đáp ứng "nhanh nhất". Overhead protocol (TCP/IP) làm chậm thêm 20-30%, không scale tốt cho 300 TB.

  • ❌ [SAI] Compress the data, upload it to multiple cloud storage providers, and then transfer the data to Cloud Storage.
    Nén dữ liệu rồi upload đa nhà cung cấp (như AWS S3), sau chuyển sang Google Cloud Storage. Sai vì:

    • Nén chỉ giảm 20-50% kích thước (vẫn ~150-240 TB), thời gian upload vẫn >10-20 ngày.
    • Phức tạp, tốn phí cross-provider (Cloud Interconnect hoặc Storage Transfer Service), và không nhanh hơn đáng kể do bottleneck bandwidth.
  • ✅ [ĐÚNG] Request a Transfer Appliance, copy the data to the appliance, and ship it back to Google.
    Như đã giải thích ở trên: Đúng vì vượt qua hoàn toàn hạn chế internet, thời gian ngắn nhất, an toàn dữ liệu (encryption at-rest/in-transit), hỗ trợ trực tiếp từ Google.

  • ❌ [SAI] Use the gcloud storage command to transfer the data over the internet.
    Sử dụng lệnh gcloud storage (gsutil tương đương) để upload qua internet. Sai vì: Tương tự phương án 1, chỉ là công cụ CLI, vẫn phụ thuộc 100 Mbps → thời gian dài (>27 ngày), multi-thread giúp chút nhưng không đủ cho 300 TB khẩn cấp.