Ngân hàng đề — Google Cloud Professional Data Engineer

Tìm thấy 429 câu.

Câu 411
Your chemical company needs to manually check documentation for customer order. You use a pull subscription in Pub/Sub so that sales agents get details from the order. You must ensure that you do not process orders twice with different sales agents and that you do not add more complexity to this workflow. What should you do?
  1. A Use a Deduplicate PTransform in Dataflow before sending the messages to the sales agents.
  2. B Create a transactional database that monitors the pending messages.
  3. C Use Pub/Sub exactly-once delivery in your pull subscription.
  4. D Create a new Pub/Sub push subscription to monitor the orders processed in the agent's system.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong hệ thống Google Cloud Platform (GCP), cụ thể là dịch vụ Pub/Sub (Publish/Subscribe messaging service). Công ty hóa chất cần kiểm tra thủ công tài liệu đơn hàng khách hàng, và sử dụng pull subscription trong Pub/Sub để các sales agents (nhân viên bán hàng) tự kéo (pull) thông tin chi tiết đơn hàng từ topic Pub/Sub.

Vấn đề chính cần giải quyết:

  • Tránh xử lý trùng lặp (duplicate processing): Không để một đơn hàng bị nhiều sales agents xử lý hai lần, dẫn đến lỗi hoặc lãng phí.
  • Không thêm phức tạp vào workflow: Giải pháp phải đơn giản, tận dụng tính năng sẵn có của Pub/Sub, tránh xây dựng thêm hệ thống mới làm phức tạp quy trình hiện tại (manual check + pull subscription).

Mục tiêu là đảm bảo exactly-once delivery (giao hàng chính xác một lần) cho pull subscription mà không cần code thêm hoặc dịch vụ phụ trợ. Đây là tính năng được GCP cập nhật mạnh mẽ từ năm 2021 và vẫn là phiên bản mới nhất đến 2026 (Pub/Sub v2 với exactly-once semantics). 📘

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Pub/Sub exactly-once delivery in your pull subscription.

Lý do chi tiết:

  • Pub/Sub hỗ trợ exactly-once delivery cho pull subscriptions (từ năm 2021, áp dụng đầy đủ đến 2026), đảm bảo mỗi message chỉ được giao chính xác một lần cho subscriber duy nhất, ngay cả khi có nhiều sales agents pull cùng lúc.
  • Tính năng này sử dụng message ID deduplication và subscriber-side acknowledgment tự động, ngăn chặn duplicate mà không cần thay đổi workflow hiện tại (vẫn giữ pull subscription đơn giản).
  • Hoàn hảo cho yêu cầu: Giảm rủi ro xử lý trùng lặp giữa các agents, không thêm complexity (chỉ enable flag khi tạo subscription). 🛠️
  • Theo docs GCP 2026: Exactly-once chỉ áp dụng cho pull (không phải push), phù hợp 100% với scenario.

❌ Phân tích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên nội dung văn bản gốc bằng tiếng Anh, chỉ giải thích lý do đúng/sai bằng tiếng Việt với emoji nổi bật:

  • [SAI] Use a Deduplicate PTransform in Dataflow before sending the messages to the sales agents.
    ❌ Sai vì: Dataflow (Apache Beam) với Deduplicate PTransform yêu cầu xây dựng pipeline mới để xử lý deduplication trước khi forward message sang sales agents. Điều này thêm complexity lớn (cần code Beam job, quản lý Dataflow resources), vi phạm yêu cầu "không thêm phức tạp workflow". Pub/Sub pull không cần Dataflow; giải pháp này overkill và không tận dụng native Pub/Sub. 🧩

  • [SAI] Create a transactional database that monitors the pending messages.
    ❌ Sai vì: Tạo transactional DB (như Cloud SQL/Spanner) để track pending messages đòi hỏi code thêm logic (insert/check DB trước khi process), dẫn đến latency cao, điểm nghẽn và complexity lớn (xử lý race condition giữa agents). Không phù hợp với workflow pull đơn giản; Pub/Sub native tốt hơn mà không cần DB external. 🚫

  • [ĐÚNG] Use Pub/Sub exactly-once delivery in your pull subscription.
    ✅ Đúng vì: Như đã giải thích ở trên. Đây là tính năng built-in của Pub/Sub (enable enableExactlyOnceDelivery: true), đảm bảo deduplication tự động cho pull subscription. Không thay đổi workflow, xử lý duplicate ở mức message ID, lý tưởng cho multi-agents. Hiệu suất cao, scale tốt đến 2026. 🎯

  • [SAI] Create a new Pub/Sub push subscription to monitor the orders processed in the agent's system.
    ❌ Sai vì: Tạo push subscription mới để monitor orders đã process sẽ thêm một layer phức tạp (cần endpoint nhận push, handle agent system integration). Không giải quyết duplicate ở pull gốc (vẫn có thể pull/process trùng), và push không hỗ trợ exactly-once tốt như pull. Workflow trở nên rối với 2 subscriptions. 🔄

📘 Tài liệu tham khảo (cập nhật mới nhất GCP đến 2026)

Giải pháp này tối ưu, tuân thủ best practices GCP! 🚀

Câu 412
You are migrating your on-premises data warehouse to BigQuery. As part of the migration, you want to facilitate cross-team collaboration to get the most value out of the organization’s data. You need to design an architecture that would allow teams within the organization to securely publish, discover, and subscribe to read-only data in a self-service manner. You need to minimize costs while also maximizing data freshness. What should you do?
  1. A Use Analytics Hub to facilitate data sharing.
  2. B Create authorized datasets to publish shared data in the subscribing team's project.
  3. C Create a new dataset for sharing in each individual team’s project. Grant the subscribing team the bigquery.dataViewer role on the dataset.
  4. D Use BigQuery Data Transfer Service to copy datasets to a centralized BigQuery project for sharing.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm

📖 Nội dung câu hỏi được giải thích rõ ràng:
Câu hỏi mô tả tình huống bạn đang di chuyển data warehouse từ on-premises sang BigQuery (dịch vụ kho dữ liệu serverless của Google Cloud). Mục tiêu chính là thiết kế một kiến trúc chia sẻ dữ liệu an toàn giữa các team trong tổ chức, cho phép:

  • Publish (xuất bản) dữ liệu từ team cung cấp.
  • Discover (khám phá) dữ liệu có sẵn.
  • Subscribe (đăng ký) để truy cập dữ liệu read-only (chỉ đọc) một cách self-service (tự phục vụ).
    Yêu cầu bổ sung: Giảm thiểu chi phí (minimize costs) và tối đa hóa độ tươi mới của dữ liệu (maximizing data freshness, nghĩa là dữ liệu luôn cập nhật realtime mà không cần copy).
    🛠️ Đây là kịch bản thực tế trong Google Cloud, tập trung vào data governance và collaboration trong môi trường multi-project/multi-team, sử dụng các tính năng native của BigQuery để tránh sao chép dữ liệu (tiết kiệm storage và compute costs).

✅ Đáp án đúng: Use Analytics Hub to facilitate data sharing.
Lý do lựa chọn (chi tiết):
Analytics Hub là dịch vụ chuyên biệt của Google Cloud (ra mắt từ 2021 và cập nhật liên tục đến 2026) để chia sẻ dữ liệu cross-project/cross-organization một cách an toàn, self-service. Nó hỗ trợ:

  • Publisher tạo data listings (danh sách dữ liệu) để teams khác discover qua catalog.
  • Subscriber subscribe để tạo linked dataset (liên kết chỉ đọc, không copy dữ liệu gốc → giữ freshness cao nhất, realtime sync).
  • Zero-copy sharing: Không tốn storage/copy costs, chỉ tính phí query trên dữ liệu gốc.
  • IAM-based security: Quyền granular (bigquery.dataViewer-like cho linked data).
    Điều này hoàn hảo khớp yêu cầu: self-service, read-only, low-cost, high-freshness. Theo docs cập nhật 2026, Analytics Hub hỗ trợ cả private listings (internal org) và public marketplace.
    📘 Nguồn tham khảo:
  • BigQuery Analytics Hub overview (Google Cloud Docs, cập nhật Q1 2026).
  • Best practices for data sharing.

🔍 Giải thích tất cả các phương án (đúng/sai)

  • ✅ Use Analytics Hub to facilitate data sharing.
    Phân tích đúng: Như đã giải thích trên, đây là giải pháp tối ưu nhất cho self-service publish/discover/subscribe với linked datasets (zero-copy, realtime freshness). Không cần copy dữ liệu → tiết kiệm chi phí storage/transfer lên đến 100% so với các phương án copy-based. Hoàn toàn phù hợp kiến trúc BigQuery hiện đại (2026), hỗ trợ data mesh patterns cho collaboration lớn.

  • ❌ Create authorized datasets to publish shared data in the subscribing team's project.
    Phân tích sai: Không tồn tại khái niệm "authorized datasets" chuẩn trong BigQuery (có thể nhầm lẫn với Authorized Views – views được ủy quyền chia sẻ cross-project). Authorized Views chỉ hỗ trợ chia sẻ views (không phải toàn bộ datasets), phải tạo thủ công ở project subscriber → không self-service, không discover dễ dàng, và vẫn cần query gốc (nhưng phức tạp hơn Analytics Hub). Không tối ưu freshness/costs cho datasets lớn, dễ gặp vấn đề governance.

  • ❌ Create a new dataset for sharing in each individual team’s project. Grant the subscribing team the bigquery.dataViewer role on the dataset.
    Phân tích sai: Phương án này yêu cầu tạo dataset riêng ở từng project subscriber và grant IAM role thủ công → không scalable/self-service (phải quản lý hàng trăm dataset cho multi-team). Tốn kém (duplicate storage), freshness kém nếu phải sync thủ công, và rủi ro security (chia sẻ cross-project cần VPC-SC hoặc authorized views bổ sung). Không khớp yêu cầu "publish once, subscribe many".

  • ❌ Use BigQuery Data Transfer Service to copy datasets to a centralized BigQuery project for sharing.
    Phân tích sai: BigQuery Data Transfer Service (nay là Data Transfers) dùng để copy/schedule transfer dữ liệu từ nguồn ngoài (như S3, on-prem) vào BigQuery → tạo duplicate datasets ở project trung tâm. Sai vì: tốn costs cao (storage + transfer fees), freshness thấp (chỉ sync định kỳ, không realtime), không self-service (phải quản lý transfer jobs), và không read-only native (cần thêm IAM). Không phù hợp migration/collaboration scenario.

🧠 Kết luận: Analytics Hub là best practice cho data sharing trong BigQuery (theo Google Cloud Well-Architected Framework 2026). Nếu triển khai, bắt đầu bằng publisher tạo listings trong Console/CLI! 🚀

Câu 413
You want to migrate an Apache Spark 3 batch job from on-premises to Google Cloud. You need to minimally change the job so that the job reads from Cloud Storage and writes the result to BigQuery. Your job is optimized for Spark, where each executor has 8 vCPU and 16 GB memory, and you want to be able to choose similar settings. You want to minimize installation and management effort to run your job. What should you do?
  1. A Execute the job as part of a deployment in a new Google Kubernetes Engine cluster.
  2. B Execute the job from a new Compute Engine VM.
  3. C Execute the job in a new Dataproc cluster.
  4. D Execute as a Dataproc Serverless job.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi yêu cầu di chuyển (migrate) một công việc batch Apache Spark 3 từ on-premises sang Google Cloud một cách tối thiểu thay đổi mã nguồn (minimally change the job). Công việc cần đọc dữ liệu từ Cloud Storage (GCS) và ghi kết quả vào BigQuery. Công việc đã được tối ưu hóa cho Spark với mỗi executor có 8 vCPU và 16 GB memory, và người dùng muốn chọn cấu hình tương tự. Mục tiêu chính là giảm thiểu nỗ lực cài đặt và quản lý (minimize installation and management effort) khi chạy job.

🛠️ Yêu cầu chính cần đáp ứng:

  • Hỗ trợ Spark 3 (phiên bản mới nhất đến 2026 vẫn tương thích qua Dataproc).
  • Tích hợp dễ dàng với GCS (đọc) và BigQuery (ghi) mà không cần thay đổi lớn.
  • Cho phép custom config executor (vCPU/RAM).
  • Serverless hoặc managed để tránh quản lý hạ tầng thủ công.

📘 Kiến thức cập nhật (Google Cloud 2026): Dataproc Serverless (ra mắt 2022, cập nhật liên tục) là lựa chọn lý tưởng cho Spark batch jobs, hỗ trợ Spark 3.x, auto-scale, không cần quản lý cluster, tích hợp native với GCS/BigQuery qua connectors. (Nguồn: Google Cloud Dataproc Serverless docs).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Execute as a Dataproc Serverless job.

Lý do:

  • 🟢 Serverless hoàn toàn: Không cần tạo cluster, cài đặt Spark thủ công → giảm thiểu nỗ lực quản lý tối đa (auto-provision, scale, cleanup).
  • 🟢 Hỗ trợ Spark 3 native, giữ nguyên config executor (8 vCPU/16GB) qua tham số --properties spark.executor.cores=8,spark.executor.memory=16g.
  • 🟢 Tích hợp liền mạch: Đọc GCS (gs://), ghi BigQuery qua Spark connector (không cần thay đổi job lớn).
  • 🟢 Tối ưu chi phí/effort: Chỉ submit job qua gcloud/CLI/API, phù hợp batch job. (Nguồn: Dataproc Serverless Spark guide).

🧪 Giải thích tất cả các phương án

  • ❌ [SAI] Execute the job as part of a deployment in a new Google Kubernetes Engine cluster.
    Phân tích sai: GKE yêu cầu deploy Spark on Kubernetes thủ công (sử dụng Spark Operator), phải config Pod/executor chi tiết, quản lý namespace/scale → effort cao, không "minimally change" và không giảm thiểu installation. Không serverless, phù hợp streaming hơn batch đơn giản.

  • ❌ [SAI] Execute the job from a new Compute Engine VM.
    Phân tích sai: Phải cài đặt Spark 3 thủ công trên VM (download, config YARN/K8s mode), quản lý VM (scale, update, networking cho GCS/BigQuery) → effort cực cao, không managed, dễ lỗi config executor. Không phù hợp migrate nhanh.

  • ❌ [SAI] Execute the job in a new Dataproc cluster.
    Phân tích sai: Dataproc cluster yêu cầu tạo và quản lý cluster (init script, resize, delete sau job) → vẫn có effort quản lý (dù managed hơn VM/GKE). Không serverless, tốn thời gian setup so với yêu cầu "minimize effort". (Tuy hỗ trợ Spark 3/GCS/BigQuery tốt, nhưng không tối ưu nhất).

  • ✅ [ĐÚNG] Execute as a Dataproc Serverless job.
    Phân tích đúng: Như giải thích trên, serverless 100%, custom config dễ dàng, tích hợp native GCS/BigQuery, giữ nguyên job Spark 3 → hoàn hảo khớp yêu cầu. (Nguồn: Dataproc Serverless concepts).

💡 Lời khuyên: Sử dụng lệnh gcloud dataproc batches submit spark để submit job nhanh chóng!

Câu 414
You are configuring networking for a Dataflow job. The data pipeline uses custom container images with the libraries that are required for the transformation logic preinstalled. The data pipeline reads the data from Cloud Storage and writes the data to BigQuery. You need to ensure cost-effective and secure communication between the pipeline and Google APIs and services. What should you do?
  1. A Disable external IP addresses from worker VMs and enable Private Google Access.
  2. B Leave external IP addresses assigned to worker VMs while enforcing firewall rules.
  3. C Disable external IP addresses and establish a Private Service Connect endpoint IP address.
  4. D Enable Cloud NAT to provide outbound internet connectivity while enforcing firewall rules.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc cấu hình mạng (networking) cho một Dataflow job trên Google Cloud. Dataflow là dịch vụ quản lý dữ liệu streaming và batch processing, sử dụng các worker VMs (máy ảo làm việc) để xử lý dữ liệu.

  • Bối cảnh cụ thể:
    • Pipeline sử dụng custom container images (hình ảnh container tùy chỉnh) đã cài sẵn thư viện cho logic biến đổi dữ liệu.
    • Đọc dữ liệu từ Cloud Storage (GCS) và ghi vào BigQuery.
    • Yêu cầu chính: Đảm bảo giao tiếp chi phí thấp (cost-effective) và bảo mật (secure) giữa pipeline (Dataflow workers) với các Google APIs và services (như GCS, BigQuery APIs).

🔑 Vấn đề cốt lõi: Dataflow workers cần truy cập các dịch vụ Google mà không cần IP công khai (public IP) để tránh rủi ro bảo mật (như tấn công từ internet) và giảm chi phí (không tốn phí egress traffic ra internet). Giải pháp phải tận dụng mạng riêng tư nội bộ của Google Cloud.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Disable external IP addresses from worker VMs and enable Private Google Access.

🛠️ Lý do chi tiết:

  • Disable external IP: Loại bỏ IP công khai trên worker VMs của Dataflow, ngăn chặn truy cập từ internet, tăng bảo mật (không expose ra public).
  • Enable Private Google Access (PGA): Cho phép workers truy cập các Google APIs/services (GCS, BigQuery) qua private IP ranges (như 199.36.153.4/30 cho APIs) mà không cần NAT hoặc internet. Điều này cost-effective vì:
    • Không tốn phí egress internet.
    • Không cần Cloud NAT (tiết kiệm chi phí NAT gateway).
  • Phù hợp hoàn hảo với Dataflow custom containers, vì workers vẫn pull images từ Container Registry qua PGA.
  • Đây là best practice khuyến nghị chính thức cho Dataflow để secure và optimize cost (không ảnh hưởng performance).

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm lý do bằng tiếng Việt:

  • Disable external IP addresses from worker VMs and enable Private Google Access.
    ✅ Đúng. Như đã giải thích ở trên, đây là giải pháp tối ưu: bảo mật cao (no public IP), cost-effective (private access miễn phí cho Google services), và hỗ trợ đầy đủ GCS/BigQuery mà không cần thêm dịch vụ trung gian. Dataflow docs khuyến nghị cấu hình này cho production jobs.

  • Leave external IP addresses assigned to worker VMs while enforcing firewall rules.
    ❌ Sai. Giữ external IP expose workers ra internet, dù có firewall rules vẫn rủi ro bảo mật cao (firewall chỉ kiểm soát traffic, không loại bỏ hoàn toàn threat). Ngoài ra, tốn kém hơn do egress fees khi traffic đi qua internet đến Google APIs (thay vì private routing). Không cost-effective và không secure bằng PGA.

  • Disable external IP addresses and establish a Private Service Connect endpoint IP address.
    ❌ Sai. Private Service Connect (PSC) dùng cho kết nối private đến specific services (như publisher services ngoài Google), nhưng không cần thiết và phức tạp cho Google APIs như GCS/BigQuery (chúng đã hỗ trợ qua PGA). PSC tốn thêm chi phí (endpoint management) và config overhead, không phải best practice cho Dataflow internal traffic.

  • Enable Cloud NAT to provide outbound internet connectivity while enforcing firewall rules.
    ❌ Sai. Cloud NAT dùng cho internet outbound (public internet), buộc traffic đến Google APIs đi qua NAT gateway → tốn NAT fees + egress fees, kém cost-effective. Không secure bằng PGA (vẫn phụ thuộc internet routing gián tiếp). Dataflow ưu tiên PGA cho Google services, NAT chỉ dùng khi cần third-party internet access.

🧠 Kết luận: Lựa chọn đúng tận dụng Private Google Access – giải pháp native, đơn giản nhất cho Dataflow trên VPC mà không cần IP public. Áp dụng ngay để đạt security + cost savings tối đa! 🚀

Câu 415
You are using Workflows to call an API that returns a 1KB JSON response, apply some complex business logic on this response, wait for the logic to complete, and then perform a load from a Cloud Storage file to BigQuery. The Workflows standard library does not have sufficient capabilities to perform your complex logic, and you want to use Python's standard library instead. You want to optimize your workflow for simplicity and speed of execution. What should you do?
  1. A Create a Cloud Composer environment and run the logic in Cloud Composer.
  2. B Create a Dataproc cluster, and use PySpark to apply the logic on your JSON file.
  3. C Invoke a Cloud Function instance that uses Python to apply the logic on your JSON file.
  4. D Invoke a subworkflow in Workflows to apply the logic on your JSON file.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc tối ưu hóa workflow trong Google Cloud Workflows để xử lý một quy trình cụ thể:

  • Gọi một API để nhận phản hồi JSON nhỏ (1KB).
  • Áp dụng logic kinh doanh phức tạp trên JSON này (thư viện chuẩn của Workflows không đủ khả năng).
  • Chờ logic hoàn thành.
  • Sau đó, thực hiện load dữ liệu từ file Cloud Storage vào BigQuery.

Mục tiêu chính là đơn giản hóa (simplicity) và tăng tốc độ thực thi (speed of execution), đồng thời sử dụng Python standard library thay vì giới hạn của Workflows (dựa trên YAML và thư viện chuẩn hạn chế).
🛠️ Vấn đề cốt lõi: Workflows không hỗ trợ trực tiếp code Python phức tạp, nên cần tích hợp dịch vụ bên ngoài để chạy logic này một cách serverless, nhanh chóng và dễ quản lý.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Invoke a Cloud Function instance that uses Python to apply the logic on your JSON file.

Lý do chi tiết:

  • Cloud Functions là dịch vụ serverless, hỗ trợ Python native (bao gồm standard library), cho phép chạy logic phức tạp chỉ với vài dòng code.
  • Workflows có thể invoke trực tiếp Cloud Functions qua bước cloud_functions.invoke hoặc HTTP, rất đơn giản và nhanh (thời gian cold start thấp ~ vài giây cho payload nhỏ 1KB).
  • Tối ưu simplicity: Không cần quản lý infrastructure, deploy function một lần và gọi từ workflow.
  • Tối ưu speed: Serverless scaling, execution nhanh cho workload nhỏ, sau đó tiếp tục load CS → BigQuery seamless.
  • Phù hợp phiên bản mới nhất (2024-2026): Cloud Functions Gen2 hỗ trợ VPC, event-driven, tích hợp sâu với Workflows/BigQuery (xem docs GCP 2024).

📋 Giải thích tất cả các phương án (đúng và sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh, kèm giải thích đúng/sai bằng tiếng Việt:

  • ❌ [SAI] Create a Cloud Composer environment and run the logic in Cloud Composer.
    Lý do sai: Cloud Composer (dựa trên Apache Airflow) quá phức tạp cho task nhỏ (1KB JSON), yêu cầu setup environment đầy đủ (DAGs, operators), thời gian khởi tạo lâu, chi phí cao. Không tối ưu simplicity/speed, phù hợp hơn cho orchestration lớn chứ không phải logic Python đơn lẻ trong Workflows.

  • ❌ [SAI] Create a Dataproc cluster, and use PySpark to apply the logic on your JSON file.
    Lý do sai: Dataproc dành cho big data processing (Spark/Hadoop), quá nặng nề/overkill cho 1KB JSON. Phải tạo cluster (thời gian provision ~5-10 phút), chi phí cao, không serverless. PySpark không cần thiết và chậm hơn Python thuần, vi phạm nguyên tắc simplicity/speed.

  • ✅ [ĐÚNG] Invoke a Cloud Function instance that uses Python to apply the logic on your JSON file.
    (Đã giải thích chi tiết ở phần trên – lựa chọn tối ưu nhất!)

  • ❌ [SAI] Invoke a subworkflow in Workflows to apply the logic on your JSON file.
    Lý do sai: Subworkflow vẫn dùng YAML syntax của Workflows, chỉ kế thừa thư viện chuẩn (không hỗ trợ Python standard library phức tạp). Không giải quyết vấn đề "thư viện chuẩn không đủ", dẫn đến lặp lại hạn chế gốc, không đơn giản hóa mà còn phức tạp hóa workflow.

📘 Tài liệu tham khảo (phiên bản mới nhất GCP đến 2026)

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần ví dụ code workflow, hãy hỏi thêm nhé!

Câu 416
You are administering a BigQuery on-demand environment. Your business intelligence tool is submitting hundreds of queries each day that aggregate a large (50 TB) sales history fact table at the day and month levels. These queries have a slow response time and are exceeding cost expectations. You need to decrease response time, lower query costs, and minimize maintenance. What should you do?
  1. A Build authorized views on top of the sales table to aggregate data at the day and month level.
  2. B Enable BI Engine and add your sales table as a preferred table.
  3. C Build materialized views on top of the sales table to aggregate data at the day and month level.
  4. D Create a scheduled query to build sales day and sales month aggregate tables on an hourly basis.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc quản trị môi trường BigQuery on-demand (mô hình tính phí theo lượng dữ liệu quét), nơi công cụ BI đang gửi hàng trăm truy vấn mỗi ngày để tổng hợp (aggregate) một bảng fact lớn 50 TB về lịch sử bán hàng theo mức ngày (day) và tháng (month). Các vấn đề chính:

  • Thời gian phản hồi chậm (slow response time).
  • Chi phí vượt dự kiến (exceeding cost expectations) do quét toàn bộ bảng lớn mỗi lần.
  • Yêu cầu: Giảm thời gian phản hồi, giảm chi phí truy vấn, và giảm thiểu bảo trì (minimize maintenance).

Mục tiêu là tối ưu hóa cho các truy vấn aggregate lặp lại trên dữ liệu lớn, tận dụng tính năng của BigQuery để pre-compute dữ liệu mà không cần quản lý thủ công nhiều. 📈

✅ Đáp án đúng và lý do lựa chọn

Build materialized views on top of the sales table to aggregate data at the day and month level.

Lý do:
Materialized Views trong BigQuery là lựa chọn tối ưu nhất vì chúng pre-compute và lưu trữ kết quả aggregate (như SUM, COUNT theo day/month) trên bảng gốc, tự động refresh nền (background refresh) dựa trên metadata mà không cần can thiệp thủ công. Điều này:

  • Giảm thời gian phản hồi: Truy vấn chỉ đọc view đã aggregate (dữ liệu nhỏ hơn nhiều so với 50 TB). ⚡
  • Giảm chi phí: Giảm lượng dữ liệu quét (on-demand pricing chỉ tính bytes scanned từ view). 💰
  • Giảm bảo trì: BigQuery tự quản lý refresh (hỗ trợ incremental refresh từ phiên bản mới nhất 2023-2026), không cần scheduler thủ công. Hoàn hảo cho workload BI hàng trăm query/ngày. 🛠️

(Dẫn nguồn: BigQuery Materialized Views Documentation - Cập nhật 2024-2026, hỗ trợ base table lên đến petabyte-scale với auto-refresh).

📋 Giải thích tất cả các phương án (Đúng/Sai)

  • ❌ [SAI] Build authorized views on top of the sales table to aggregate data at the day and month level.
    Authorized Views chỉ dùng để kiểm soát quyền truy cập (row/column-level security) trên dữ liệu gốc, không pre-compute hay lưu trữ aggregate. Mỗi query vẫn phải quét toàn bộ 50 TB bảng sales → thời gian chậm, chi phí cao, không giải quyết vấn đề. Không giảm maintenance vì chỉ là view logic.

  • ❌ [SAI] Enable BI Engine và add your sales table as a preferred table.
    BI Engine là in-memory acceleration cho BI tools (như Looker/Tableau), tăng tốc query trên dữ liệu đã partition/clustered. Tuy nhiên, nó không aggregate dữ liệu sẵn → vẫn quét 50 TB mỗi lần, chỉ nhanh hơn chút (phụ thuộc RAM reservations). Chi phí BI Engine thêm (slot-based), không giảm query costs on-demand, và cần cấu hình preferred tables thủ công → không minimize maintenance.

  • ✅ [ĐÚNG] Build materialized views on top of the sales table to aggregate data at the day and month level.
    Như đã giải thích ở trên: Pre-compute aggregate day/month, auto-refresh, lý tưởng cho bảng lớn và query lặp lại. Giảm slots scanned xuống mức tối thiểu, phù hợp on-demand. (Nguồn: BigQuery Pricing - Views chỉ tính 10% chi phí so với base table).

  • ❌ [SAI] Create a scheduled query to build sales day and sales month aggregate tables on an hourly basis.
    Tạo bảng aggregate thủ công qua scheduled queries (Cloud Scheduler + BigQuery) → cần bảo trì cao (xử lý failure, điều chỉnh schedule, duplicate data). Refresh hourly không real-time cho query BI hàng trăm lần/ngày, vẫn tốn chi phí build bảng định kỳ + query chậm nếu lag. Không tự động như materialized views. ⏰

Câu 417
You have several different unstructured data sources, within your on-premises data center as well as in the cloud. The data is in various formats, such as Apache Parquet and CSV. You want to centralize this data in Cloud Storage. You need to set up an object sink for your data that allows you to use your own encryption keys. You want to use a GUI-based solution. What should you do?
  1. A Use BigQuery Data Transfer Service to move files into BigQuery.
  2. B Use Storage Transfer Service to move files into Cloud Storage
  3. C Use Dataflow to move files into Cloud Storage
  4. D Use Cloud Data Fusion to move files into Cloud Storage.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc tập trung hóa dữ liệu không cấu trúc (unstructured data) từ nhiều nguồn khác nhau, bao gồm on-premises data center và trong cloud, với các định dạng như Apache Parquet và CSV, vào Cloud Storage (GCS).
✅ Yêu cầu chính:

  • Thiết lập một object sink (điểm đến lưu trữ object trong GCS).
  • Hỗ trợ sử dụng khóa mã hóa riêng (own encryption keys) – tức là Customer-Managed Encryption Keys (CMEK).
  • Sử dụng giải pháp dựa trên GUI (giao diện đồ họa) để dễ dàng thiết kế và quản lý pipeline mà không cần code nhiều.

Mục tiêu là di chuyển dữ liệu một cách an toàn, linh hoạt từ các nguồn đa dạng vào GCS, tận dụng giao diện trực quan. Đây là tình huống điển hình trong Google Cloud Data Engineering, phù hợp với các công cụ ETL/ELT GUI-based.
(Kiến thức cập nhật: Dựa trên Google Cloud docs phiên bản 2024-2026, Cloud Data Fusion v2.x hỗ trợ đầy đủ CMEK cho GCS sinks – xem Cloud Data Fusion Documentation).

✅ Đáp án đúng: Use Cloud Data Fusion to move files into Cloud Storage

Lý do chọn đáp án này:
🛠️ Cloud Data Fusion là nền tảng no-code/low-code ETL dựa hoàn toàn trên GUI (dùng giao diện drag-and-drop để thiết kế pipeline). Nó hỗ trợ:

  • Kết nối đa nguồn: On-premises (qua agents), cloud sources, Parquet/CSV.
  • Object sink vào GCS: Tạo pipeline với GCS sink, hỗ trợ CMEK trực tiếp (chọn encryption key khi config sink).
  • GUI-based 100%: Sử dụng Cloud Data Fusion UI hoặc integrated với Cloud Console để build, test, schedule pipeline mà không cần code.
    Hoàn hảo cho unstructured data, scalable, và tích hợp DataProc backend.
    📘 Nguồn: Cloud Data Fusion - Connectors & Sinks, CMEK in Data Fusion.

❌ Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh:

  • [SAI] Use BigQuery Data Transfer Service to move files into BigQuery.
    ❌ Sai vì: Dịch vụ này chỉ di chuyển dữ liệu vào BigQuery (warehouse), không phải Cloud Storage (object storage). Không hỗ trợ "object sink" vào GCS, và chủ yếu dành cho structured/semi-structured data từ sources như GCS/S3, không linh hoạt cho on-premises unstructured files. Không nhấn mạnh GUI cho custom pipelines phức tạp.
    📘 Nguồn: BigQuery Data Transfer Service Limits.

  • [SAI] Use Storage Transfer Service to move files into Cloud Storage.
    ❌ Sai vì: Storage Transfer Service (STS) hỗ trợ di chuyển files vào GCS từ on-prem/cloud (qua agent), nhưng không phải GUI-based đầy đủ – chủ yếu dùng CLI/API hoặc console cơ bản, thiếu pipeline designer drag-and-drop cho multi-source phức tạp. Hỗ trợ CMEK cho destination, nhưng không có "object sink" concept như ETL tool, và kém linh hoạt với formats như Parquet (chỉ copy raw).
    📘 Nguồn: Storage Transfer Service Overview, không đề cập GUI ETL.

  • [SAI] Use Dataflow to move files into Cloud Storage.
    ❌ Sai vì: Dataflow là serverless stream/batch processing dựa trên code (Apache Beam), không có GUI native (dù có template console, nhưng yêu cầu code/programming). Phù hợp transform data, nhưng không phải giải pháp GUI đơn giản cho "set up object sink", và config CMEK phức tạp hơn. Không lý tưởng cho non-engineers.
    📘 Nguồn: Dataflow Templates, nhấn mạnh code-based.

  • [ĐÚNG] Use Cloud Data Fusion to move files into Cloud Storage.
    ✅ Đúng vì (như phần trên): GUI hoàn chỉnh, hỗ trợ full yêu cầu, là lựa chọn tối ưu cho data integration scenarios.

Kết luận 🎯: Cloud Data Fusion là best practice cho hybrid/multi-cloud data pipelines GUI-based trên GCP (2026). Nếu cần thực hành, thử free trial trên Cloud Console!

Câu 418
You are using BigQuery with a regional dataset that includes a table with the daily sales volumes. This table is updated multiple times per day. You need to protect your sales table in case of regional failures with a recovery point objective (RPO) of less than 24 hours, while keeping costs to a minimum. What should you do?
  1. A Schedule a daily export of the table to a Cloud Storage dual or multi-region bucket.
  2. B Schedule a daily copy of the dataset to a backup region.
  3. C Schedule a daily BigQuery snapshot of the table.
  4. D Modify ETL job to load the data into both the current and another backup region.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc bảo vệ dữ liệu trong BigQuery (dịch vụ kho dữ liệu của Google Cloud) đối với một regional dataset chứa bảng dữ liệu về khối lượng bán hàng hàng ngày. Bảng này được cập nhật nhiều lần mỗi ngày, và yêu cầu chính là:

  • Bảo vệ chống lại sự cố khu vực (regional failures): Nghĩa là nếu vùng (region) chứa dataset bị lỗi, dữ liệu vẫn có thể khôi phục từ nơi khác.
  • Recovery Point Objective (RPO) < 24 giờ: Mức độ mất dữ liệu tối đa phải dưới 24 giờ (tức là dữ liệu mới nhất có thể mất không quá 23 giờ 59 phút).
  • Tối ưu chi phí thấp nhất: Không dùng giải pháp đắt đỏ như lưu trữ kép hoặc xử lý kép.

Tình huống thực tế: Regional dataset chỉ tồn tại ở một region duy nhất (ví dụ: us-central1), dễ bị ảnh hưởng bởi sự cố khu vực. Cần giải pháp sao lưu định kỳ, bền vững cao, chi phí thấp, phù hợp với cập nhật dữ liệu thường xuyên. Dựa trên kiến thức Google Cloud mới nhất đến năm 2026 (BigQuery version 2.x với các tính năng backup/DR cải tiến), giải pháp phải tận dụng Cloud Storage multi-region để đạt độ bền cao (11 nines durability) mà không tốn kém.

📘 Tài liệu tham khảo:

✅ Đáp án đúng

Schedule a daily export of the table to a Cloud Storage dual or multi-region bucket.

Lý do chọn đáp án này:

  • 🛡️ Đáp ứng RPO <24 giờ: Xuất dữ liệu hàng ngày đảm bảo mất dữ liệu tối đa dưới 24 giờ, phù hợp với bảng cập nhật nhiều lần/ngày.
  • 🌍 Bảo vệ regional failure: Cloud Storage dual/multi-region bucket tự động replicate dữ liệu qua nhiều region (ví dụ: US multi-region), chịu được sự cố khu vực.
  • 💰 Chi phí tối thiểu: Export từ BigQuery sang Cloud Storage chỉ tính phí xuất dữ liệu rẻ (khoảng $0.01/GB), lưu trữ multi-region Standard class chỉ ~$0.026/GB/tháng, rẻ hơn copy dataset hoặc dual-load. Không cần query hoặc compute thêm.
  • 🔄 Dễ khôi phục: Có thể load lại từ CS vào BigQuery region mới chỉ trong vài phút.
  • Đây là best practice chính thức cho low-cost DR trong BigQuery regional datasets.

📋 Giải thích chi tiết tất cả các phương án

  • ✅ Schedule a daily export of the table to a Cloud Storage dual or multi-region bucket.
    (Như đã giải thích ở trên) – Giải pháp lý tưởng, cân bằng RPO, độ bền và chi phí.

  • ❌ Schedule a daily copy of the dataset to a backup region.
    Phương án này sai vì: Copy dataset BigQuery cross-region (sử dụng bq cp hoặc scheduled query) tạo bản sao đầy đủ, dẫn đến chi phí lưu trữ gấp đôi (active storage ở cả hai region, ~$0.023/GB/tháng/region). Với dữ liệu cập nhật thường xuyên, copy daily sẽ tốn kém cao hơn export sang CS. Ngoài ra, copy không tự động replicate như multi-region bucket, và RPO vẫn chỉ ~24 giờ nhưng kém hiệu quả hơn.

  • ❌ Schedule a daily BigQuery snapshot of the table.
    Phương án này sai vì: BigQuery snapshots (logical table snapshots, ra mắt 2023-2024) chỉ là point-in-time copy trong cùng region/dataset, không bảo vệ chống regional failure (vẫn bị mất nếu region lỗi). Snapshots dùng cho time-travel/query lịch sử, tính phí lưu trữ riêng (~$0.01/GB/tháng snapshot storage), không đạt yêu cầu DR cross-region. Không phù hợp RPO với dữ liệu động.

  • ❌ Modify ETL job to load the data into both the current and another backup region.
    Phương án này sai vì: Dual-write ETL (sử dụng Dataflow hoặc scheduled queries load kép) tăng gấp đôi chi phí compute/load (query slots, Dataflow vCPU, ~2x bills), vi phạm "keeping costs to a minimum". Dù đạt RPO thấp hơn (near real-time), nhưng phức tạp triển khai, dễ lỗi đồng bộ, và tốn kém hơn export hàng ngày rất nhiều.

Kết luận 🏆: Chọn export sang Cloud Storage multi-region là cách thông minh, tiết kiệm nhất theo hướng dẫn DR chính thức của Google Cloud! 🚀

Câu 419
You are preparing an organization-wide dataset. You need to preprocess customer data stored in a restricted bucket in Cloud Storage. The data will be used to create consumer analyses. You need to follow data privacy requirements, including protecting certain sensitive data elements, while also retaining all of the data for potential future use cases. What should you do?
  1. A Use the Cloud Data Loss Prevention API and Dataflow to detect and remove sensitive fields from the data in Cloud Storage. Write the filtered data in BigQuery.
  2. B Use customer-managed encryption keys (CMEK) to directly encrypt the data in Cloud Storage. Use federated queries from BigQuery. Share the encryption key by following the principle of least privilege.
  3. C Use Dataflow and the Cloud Data Loss Prevention API to mask sensitive data. Write the processed data in BigQuery.
  4. D Use Dataflow and Cloud KMS to encrypt sensitive fields and write the encrypted data in BigQuery. Share the encryption key by following the principle of least privilege.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi này thuộc lĩnh vực Google Cloud Platform (GCP) (không phải AWS như mô tả ban đầu, vì toàn bộ dịch vụ đề cập là GCS, Dataflow, BigQuery, DLP API, KMS – các dịch vụ GCP tiêu chuẩn). Nội dung yêu cầu chuẩn bị một dataset toàn tổ chức từ dữ liệu khách hàng lưu trữ trong restricted bucket trên Cloud Storage (GCS). Dữ liệu này dùng để tạo phân tích người tiêu dùng (consumer analyses), nhưng phải tuân thủ yêu cầu bảo mật dữ liệu (data privacy):

  • Bảo vệ các yếu tố dữ liệu nhạy cảm (sensitive data elements) như PII (Personally Identifiable Information).
  • Giữ nguyên TẤT CẢ dữ liệu cho các trường hợp sử dụng tương lai (retain all data for potential future use cases). 📌 Thách thức chính: Preprocess dữ liệu để bảo vệ mà không xóa hoặc mất mát bất kỳ phần nào, đồng thời lưu trữ kết quả vào nơi dễ phân tích như BigQuery. Giải pháp cần sử dụng công cụ tự động hóa (như Dataflow) và API chuyên về bảo mật (DLP).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Dataflow and the Cloud Data Loss Prevention API to mask sensitive data. Write the processed data in BigQuery.

Lý do (🛠️ Phân tích chi tiết):

  • Cloud DLP API chuyên phát hiện (inspect) và mask (che giấu/mask) dữ liệu nhạy cảm (như tên, email, số điện thoại) bằng các phương pháp như pseudonymization (thay thế bằng hash hoặc fake data), giữ nguyên cấu trúc và toàn bộ dữ liệu mà không xóa.
  • Dataflow (Apache Beam trên GCP) xử lý dữ liệu lớn theo batch/streaming, tích hợp DLP để transform dữ liệu từ GCS một cách scalable.
  • Kết quả ghi vào BigQuery để phân tích dễ dàng, tuân thủ privacy (GDPR/CCPA) và retain all data cho future use (vì chỉ mask, không remove).
  • Đây là best practice GCP đến 2026: DLP v2 hỗ trợ masking động, Dataflow templates sẵn cho DLP integration (cập nhật Q1/2026 với AI-based de-identification).

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn (giữ nguyên văn bản gốc tiếng Anh). Tôi đánh dấu ✅ đúng / ❌ sai dựa trên yêu cầu câu hỏi, với lý do bằng tiếng Việt rõ ràng:

  • ❌ [SAI] Use the Cloud Data Loss Prevention API and Dataflow to detect and remove sensitive fields from the data in Cloud Storage. Write the filtered data in BigQuery.
    Lý do sai: Phương án dùng DLP để detect và REMOVE (xóa) fields nhạy cảm → mất dữ liệu vĩnh viễn, vi phạm yêu cầu "retaining all of the data for potential future use cases". Chỉ phù hợp nếu không cần giữ structure đầy đủ; BigQuery chỉ nhận data đã filter (không scalable cho retain).

  • ❌ [SAI] Use customer-managed encryption keys (CMEK) to directly encrypt the data in Cloud Storage. Use federated queries from BigQuery. Share the encryption key by following the principle of least privilege.
    Lý do sai: CMEK chỉ encrypt toàn bộ object trong GCS (không target specific sensitive fields), không preprocess/protect chi tiết. Federated queries từ BigQuery cần access bucket (phức tạp với restricted bucket), và share key vẫn rủi ro exposure. Không mask/remove PII, không retain "processed data" cho analyses trực tiếp.

  • ✅ [ĐÚNG] Use Dataflow and the Cloud Data Loss Prevention API to mask sensitive data. Write the processed data in BigQuery.
    Lý do đúng (🧩 Chi tiết): Như đã giải thích ở trên – DLP mask (e.g., redact, replace, crypto-hash) bảo vệ PII mà giữ toàn bộ dữ liệu, Dataflow xử lý end-to-end từ GCS → BigQuery. Scalable, cost-effective, tuân thủ privacy laws. Best practice cho ETL với de-identification.

  • ❌ [SAI] Use Dataflow and Cloud KMS to encrypt sensitive fields and write the encrypted data in BigQuery. Share the encryption key by following the principle of least privilege.
    Lý do sai: Cloud KMS dùng cho encrypt/decrypt symmetric/asymmetric keys (không tự detect sensitive fields như DLP). Phải manual identify fields → không tự động/scalable. Encrypt fields rồi share key vẫn cần quyền truy cập (rủi ro), và data encrypted khó query/analyze trực tiếp trong BigQuery mà không decrypt (vi phạm least privilege thực tế).

📘 Tài liệu tham khảo (cập nhật đến 2026)

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần ví dụ code Dataflow, hãy hỏi thêm.

Câu 420
You need to load a dataset with multiple terabytes of clickstream data into BigQuery. The data arrives each day as compressed JSON files in a Cloud Storage bucket. You need a low-cost, programmatic, and scalable solution to load the data into BigQuery. What should you do?
  1. A Create an external table in BigQuery pointing to the Cloud Storage bucket and run the INSERT INTO ... SELECT * FROM external_table command.
  2. B Use the BigQuery Data Transfer Service from Cloud Storage.
  3. C Create a Cloud Run function to run a Python script to read and parse each JSON file, and use the BigQuery streaming insert API.
  4. D Use Cloud Data Fusion to create a pipeline to load the JSON files into BigQuery.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi yêu cầu một giải pháp thấp chi phí (low-cost), có thể lập trình tự động (programmatic) và có khả năng mở rộng (scalable) để tải dữ liệu clickstream với dung lượng hàng terabytes (multiple terabytes) vào BigQuery. Dữ liệu đến hàng ngày dưới dạng file JSON nén (compressed JSON files) trong Cloud Storage bucket.

🔍 Yêu cầu chính:

  • Xử lý dữ liệu lớn hàng TB/ngày → Cần tránh chi phí cao và độ trễ lớn.
  • Dữ liệu JSON nén → Phù hợp với các công cụ hỗ trợ định dạng này mà không cần giải nén thủ công.
  • Giải pháp phải tự động hóa qua code/SQL, không phụ thuộc giao diện web, và mở rộng theo quy mô dữ liệu lớn.

📘 Kiến thức cập nhật (đến 2026): BigQuery hỗ trợ external tables cho JSON (từ phiên bản mới nhất BigQuery 2024+), cho phép query trực tiếp từ GCS mà không tốn storage. Load dữ liệu lớn qua batch (như COPY/INSERT SELECT) hiệu quả hơn streaming cho TB-scale.

✅ Đáp án đúng

Create an external table in BigQuery pointing to the Cloud Storage bucket and run the INSERT INTO ... SELECT * FROM external_table command.

Lý do chọn đáp án này:

  • 🛠️ Low-cost: External table không tốn storage (chỉ metadata), query on-demand chỉ tính phí scan dữ liệu (rẻ hơn ~10x so với streaming). Sau đó, INSERT INTO... SELECT* chỉ load dữ liệu một lần vào native table, tự động partition/clustering.
  • 🎯 Programmatic: Chạy qua SQL script (bq command-line, API, hoặc Cloud Composer), dễ tự động hóa hàng ngày qua Cloud Scheduler.
  • ⚡ Scalable: BigQuery xử lý TB-scale query song song, hỗ trợ JSON nén trực tiếp (auto-detect schema). Không cần parse thủ công, lý tưởng cho dữ liệu daily mới.
  • ✅ Hoàn hảo cho kịch bản: Query external trước để validate, rồi materialize nếu cần.

📋 Giải thích tất cả các phương án

  • ✅ Create an external table in BigQuery pointing to the Cloud Storage bucket and run the INSERT INTO ... SELECT * FROM external_table command.
    🟢 Đúng: Như giải thích trên, đây là cách tối ưu nhất cho dữ liệu lớn JSON nén. BigQuery external table hỗ trợ GCS URI pattern (e.g., gs://bucket/*.json.gz), tự parse JSON. INSERT SELECT chạy batch, chi phí ~$5/TB (on-demand), scalable 100%.

  • ❌ Use the BigQuery Data Transfer Service from Cloud Storage.
    🔴 Sai: Data Transfer Service (nay là BigQuery Transfer Service) hỗ trợ scheduled load từ GCS, nhưng không programmatic thuần túy (chủ yếu UI/schedule, API hạn chế tùy chỉnh). Với TB JSON nén daily, có thể chậm (throttling), và chi phí tương đương nhưng thiếu linh hoạt validate schema trước load. Không lý tưởng cho "programmatic".

  • ❌ Create a Cloud Run function to run a Python script to read and parse each JSON file, and use the BigQuery streaming insert API.
    🔴 Sai: Chi phí cao (streaming ~$0.01/200MB → ~$50/TB, đắt 10x batch). Không scalable cho TB daily (throttle 100MB/s/account, cần shard phức tạp). Parse JSON thủ công tốn CPU/RAM, Cloud Run scale giới hạn (1k instances), dễ timeout/fail với file lớn.

  • ❌ Use Cloud Data Fusion to create a pipeline to load the JSON files into BigQuery.
    🔴 Sai: Data Fusion (dựa CDAP) là ETL no-code/low-code, overkill và đắt đỏ (~$0.40/vCore/giờ + storage). Setup pipeline phức tạp cho simple load, không "low-cost" (chi phí instance luôn chạy), kém programmatic so với SQL thuần. Phù hợp complex transform, không phải raw load JSON.

📚 Tài liệu tham khảo

Hy vọng phân tích giúp bạn ôn thi hiệu quả! 🚀 Nếu cần ví dụ code SQL, hỏi thêm nhé!