Ngân hàng đề — Google Cloud Professional Data Engineer

Tìm thấy 429 câu.

Câu 241
Data Analysts in your company have the Cloud IAM Owner role assigned to them in their projects to allow them to work with multiple GCP products in their projects. Your organization requires that all BigQuery data access logs be retained for 6 months. You need to ensure that only audit personnel in your company can access the data access logs for all projects. What should you do?
  1. A Enable data access logs in each Data Analyst's project. Restrict access to Observability Logging via Cloud IAM roles.
  2. B Export the data access logs via a project-level export sink to a Cloud Storage bucket in the Data Analysts' projects. Restrict access to the Cloud Storage bucket.
  3. C Export the data access logs via a project-level export sink to a Cloud Storage bucket in a newly created projects for audit logs. Restrict access to the project with the exported logs.
  4. D Export the data access logs via an aggregated export sink to a Cloud Storage bucket in a newly created project for audit logs. Restrict access to the project that contains the exported logs.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi

Câu hỏi tập trung vào quản lý và bảo mật Cloud Audit Logs trong Google Cloud Platform (GCP), cụ thể là data access logs của BigQuery.

  • Bối cảnh vấn đề 📊: Các Data Analyst trong công ty được cấp quyền Cloud IAM Owner tại các project của họ, cho phép họ truy cập và quản lý hầu hết tài nguyên GCP (bao gồm logs). Tổ chức yêu cầu lưu trữ data access logs của BigQuery ít nhất 6 tháng (mặc định chỉ 400 ngày cho một số logs, nhưng data access logs cần cấu hình riêng để enable và export nhằm retain lâu dài). Quan trọng nhất, chỉ nhân viên audit mới được phép truy cập data access logs từ TẤT CẢ các project, tránh Data Analyst (với quyền Owner) có thể xem logs nhạy cảm.

  • Mục tiêu chính 🔒: Cần một giải pháp tập trung hóa logs từ nhiều project, lưu trữ an toàn, và hạn chế truy cập chỉ cho audit personnel, đồng thời đảm bảo logs được retain đúng thời hạn (qua export đến Cloud Storage với lifecycle rules).

  • Kiến thức cốt lõi (cập nhật đến 2026 theo GCP docs mới nhất): Data access logs là loại Admin Activity audit logs cho BigQuery, phải enable thủ công vì mặc định tắt để tránh chi phí. Để retain lâu, sử dụng Cloud Logging export sinks đẩy logs đến Cloud Storage. Aggregated sinks (tại organization/folder level) cho phép thu thập logs từ tất cả child projects, khác với project-level sinks chỉ thu thập trong project đó. Retain 6 tháng có thể cấu hình qua bucket lifecycle.

Đáp án đúng ✅: Export the data access logs via an aggregated export sink to a Cloud Storage bucket in a newly created project for audit logs. Restrict access to the project that contains the exported logs.

Lý do chọn đáp án đúng 🟢:

  • Aggregated export sink (tạo tại organization hoặc folder level) tự động thu thập data access logs từ TẤT CẢ projects con (bao gồm projects của Data Analyst), không cần cấu hình từng project.
  • Logs được export đến Cloud Storage bucket trong project audit riêng biệt (tạo mới), dễ dàng restrict IAM roles chỉ cho audit personnel (ví dụ: grant roles/logging.viewer hoặc roles/storage.objectViewer chỉ cho họ).
  • Đảm bảo tập trung hóa và bảo mật: Data Analyst không thể truy cập project audit (vì họ chỉ Owner ở project riêng). Bucket lifecycle rule giữ logs 6 tháng.
  • Phù hợp best practice GCP 2026: Aggregated sinks là cách scale cho multi-project logging mà không phụ thuộc quyền Owner ở từng project.

❌ Phân tích tất cả các phương án

  • [SAI] Enable data access logs in each Data Analyst's project. Restrict access to Observability Logging via Cloud IAM roles.
    ❌ Sai vì: Chỉ enable logs ở từng project riêng lẻ (project-level), Data Analyst với quyền Owner vẫn có thể truy cập logs qua Cloud Console/Logging API (Owner bao gồm roles/logging.admin). Không giải quyết tập trung logs từ tất cả projects, và restrict IAM không hiệu quả vì Owner quyền cao. Không export nên không retain lâu dài.

  • [SAI] Export the data access logs via a project-level export sink to a Cloud Storage bucket in the Data Analysts' projects. Restrict access to the Cloud Storage bucket.
    ❌ Sai vì: Project-level sink chỉ export logs trong project đó, không thu thập từ tất cả projects. Bucket nằm trong project Data Analyst, họ (Owner) vẫn kiểm soát và truy cập bucket dễ dàng, không đảm bảo chỉ audit xem. Phải cấu hình sink ở từng project – không scale.

  • [SAI] Export the data access logs via a project-level export sink to a Cloud Storage bucket in a newly created projects for audit logs. Restrict access to the project with the exported logs.
    ❌ Sai vì: Vẫn dùng project-level sink (phải tạo ở từng project Data Analyst), chỉ export logs của project đó, không thu thập từ tất cả projects một cách tự động. Dù bucket ở project audit mới, việc cấu hình nhiều sink thủ công phức tạp và Data Analyst có thể chỉnh sửa sink trong project họ.

📘 Tài liệu tham khảo (GCP docs cập nhật 2026)

Giải pháp này đảm bảo tuân thủ nguyên tắc least privilege và centralized auditing theo GCP! 🚀

Câu 242
Each analytics team in your organization is running BigQuery jobs in their own projects. You want to enable each team to monitor slot usage within their projects.
What should you do?
  1. A Create a Cloud Monitoring dashboard based on the BigQuery metric query/scanned_bytes
  2. B Create a Cloud Monitoring dashboard based on the BigQuery metric slots/allocated_for_project
  3. C Create a log export for each project, capture the BigQuery job execution logs, create a custom metric based on the totalSlotMs, and create a Cloud Monitoring dashboard based on the custom metric
  4. D Create an aggregated log export at the organization level, capture the BigQuery job execution logs, create a custom metric based on the totalSlotMs, and create a Cloud Monitoring dashboard based on the custom metric
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào BigQuery (dịch vụ kho dữ liệu và phân tích dữ liệu lớn của Google Cloud Platform - GCP). Tình huống: Mỗi đội ngũ phân tích (analytics team) trong tổ chức đang chạy các job BigQuery trong project riêng biệt của họ. Yêu cầu là cho phép mỗi team tự monitor (giám sát) việc sử dụng slot (slot là đơn vị tài nguyên tính toán trong BigQuery, quyết định hiệu suất và chi phí job) trong project của chính họ.

Mục tiêu chính: Cần một giải pháp đơn giản, trực tiếp để theo dõi slot usage per project, sử dụng Cloud Monitoring (trước đây là Stackdriver Monitoring) – công cụ giám sát metrics của GCP. Điều này giúp mỗi team tự quản lý mà không cần can thiệp tổ chức cấp cao, phù hợp với mô hình multi-project. 📊

Bối cảnh cập nhật đến 2026: BigQuery sử dụng slot-based pricing (từ năm 2020, flex slots và flat-rate slots), với các metric Monitoring được tối ưu hóa. Metric slots/allocated_for_project là metric chuẩn per-project, phản ánh slots được phân bổ thực tế cho project tại thời điểm query. Không cần log export phức tạp vì GCP cung cấp metrics sẵn sàng. 🛠️

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create a Cloud Monitoring dashboard based on the BigQuery metric slots/allocated_for_project

Lý do:

  • Metric slots/allocated_for_project chính xác đo lường slot usage được phân bổ cho từng project cụ thể, cập nhật real-time (mỗi phút).
  • Mỗi team có thể tạo dashboard riêng trong project của họ, filter theo project ID, không ảnh hưởng cross-project.
  • Đơn giản nhất: Không cần config log export hay custom metric, chỉ cần chọn metric từ BigQuery API trong Cloud Monitoring. Tiết kiệm thời gian và chi phí. 🚀
  • Phù hợp best practice GCP: Sử dụng managed metrics trước khi tự build custom.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ SAI: Create a Cloud Monitoring dashboard based on the BigQuery metric query/scanned_bytes
    Giải thích: Metric query/scanned_bytes chỉ đo lượng dữ liệu quét (bytes) trong query, không liên quan đến slot usage. Slot usage phụ thuộc vào độ phức tạp query, parallelism, và on-demand/flex slots, chứ không phải bytes scanned. Sử dụng cái này sẽ gây nhầm lẫn, không giúp monitor tài nguyên tính toán thực tế. 📉

  • ✅ ĐÚNG: Create a Cloud Monitoring dashboard based on the BigQuery metric slots/allocated_for_project
    Giải thích: Như đã nêu ở trên, đây là metric lý tưởng cho per-project slot monitoring. Nó hiển thị slots trung bình được allocate (ví dụ: 0-2000 slots tùy tier), dễ visualize qua chart/line graph trong dashboard. Mỗi team chỉ cần quyền Monitoring Viewer trong project là xem được. ⭐

  • ❌ SAI: Create a log export for each project, capture the BigQuery job execution logs, create a custom metric based on the totalSlotMs, and create a Cloud Monitoring dashboard based on the custom metric
    Giải thích: Phương án này quá phức tạp và không cần thiết. totalSlotMs (từ BigQuery logs) đo tổng slot-milliseconds đã dùng per job, nhưng phải export logs → parse → tạo custom metric (qua Cloud Logging → Metrics Explorer). Tốn tài nguyên, delay (logs không real-time), và mỗi team phải config riêng – vi phạm yêu cầu "enable each team to monitor" đơn giản. GCP khuyến nghị dùng metric sẵn. ⏳

  • ❌ SAI: Create an aggregated log export at the organization level, capture the BigQuery job execution logs, create a custom metric based on the totalSlotMs, and create a Cloud Monitoring dashboard based on the custom metric
    Giải thích: Tương tự phương án trước, nhưng tệ hơn vì aggregated tại organization level → logs mix tất cả project, khó filter per-project cho từng team. Central admin phải quản lý, team không tự monitor được. Không scalable cho multi-team, và vẫn phức tạp với totalSlotMs (chỉ retrospective, không predictive slot usage). 🚫

📘 Tài liệu tham khảo (cập nhật GCP 2026)

Giải pháp này đảm bảo tuân thủ nguyên tắc least privilege và self-service cho teams! 🌟

Câu 243
You are operating a streaming Cloud Dataflow pipeline. Your engineers have a new version of the pipeline with a different windowing algorithm and triggering strategy. You want to update the running pipeline with the new version. You want to ensure that no data is lost during the update. What should you do?
  1. A Update the Cloud Dataflow pipeline inflight by passing the --update option with the --jobName set to the existing job name
  2. B Update the Cloud Dataflow pipeline inflight by passing the --update option with the --jobName set to a new unique job name
  3. C Stop the Cloud Dataflow pipeline with the Cancel option. Create a new Cloud Dataflow job with the updated code
  4. D Stop the Cloud Dataflow pipeline with the Drain option. Create a new Cloud Dataflow job with the updated code
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi xoay quanh việc cập nhật một pipeline Cloud Dataflow đang chạy ở chế độ streaming (xử lý dữ liệu thời gian thực). Các kỹ sư đã phát triển phiên bản mới với thuật toán windowing (cách chia dữ liệu thành các cửa sổ thời gian) và chiến lược triggering (cách kích hoạt xử lý dữ liệu) khác biệt. Mục tiêu là cập nhật pipeline mà không mất dữ liệu nào trong quá trình chuyển đổi.
🛠️ Vấn đề cốt lõi: Trong Dataflow streaming, việc thay đổi windowing hoặc triggering có thể gây xung đột state (trạng thái dữ liệu), nên cần phương pháp an toàn để xử lý dữ liệu đang "bay" (in-flight data) trước khi dừng job cũ và khởi chạy job mới. Theo tài liệu GCP mới nhất (cập nhật đến 2026), Drain là cách khuyến nghị để đảm bảo watermark tiến triển và không bỏ lỡ dữ liệu.

📘 Tài liệu tham khảo:

✅ Đáp án đúng

Stop the Cloud Dataflow pipeline with the Drain option. Create a new Cloud Dataflow job with the updated code
Lý do chọn:

  • Drain cho phép pipeline tiếp tục xử lý toàn bộ dữ liệu đang pending (dữ liệu chưa được xử lý hoàn chỉnh theo watermark) trước khi dừng job cũ một cách graceful (mượt mà). Sau đó, tạo job mới với code cập nhật sẽ đảm bảo không mất dữ liệu, đặc biệt phù hợp với thay đổi windowing và triggering (có thể không tương thích với update inflight).
  • Đây là best practice của GCP cho streaming jobs với thay đổi lớn, tránh tình trạng duplicate hoặc loss data. Thời gian drain phụ thuộc vào lượng dữ liệu backlog, nhưng watermark sẽ tiến triển đầy đủ.
    🛠️ Lợi ích: Zero data loss, state migration tự nhiên qua sink (như Pub/Sub hoặc BigQuery).

🔍 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn giữ nguyên văn bản gốc bằng tiếng Anh, với giải thích sai/đúng bằng tiếng Việt dựa trên tài liệu GCP mới nhất:

  • Update the Cloud Dataflow pipeline inflight by passing the --update option with the --jobName set to the existing job name
    ❌ Sai: Lệnh --update với --jobName hiện tại chỉ hoạt động nếu code mới tương thích hoàn toàn với job cũ (ví dụ: cùng windowing, trigger, và schema). Với thay đổi windowing/triggering, Dataflow không thể migrate state đúng cách, dẫn đến data loss hoặc duplicate. GCP docs khuyến cáo tránh update inflight cho trường hợp này.

  • Update the Cloud Dataflow pipeline inflight by passing the --update option with the --jobName set to a new unique job name
    ❌ Sai: --update bắt buộc phải dùng --jobName của job hiện tại để overwrite/migrate. Nếu dùng jobName mới, nó sẽ tạo job riêng biệt (không update), gây hai job chạy song song, dẫn đến duplicate data và lãng phí tài nguyên. Không giải quyết được vấn đề loss data.

  • Stop the Cloud Dataflow pipeline with the Cancel option. Create a new Cloud Dataflow job with the updated code
    ❌ Sai: Cancel dừng job ngay lập tức, bỏ qua toàn bộ dữ liệu in-flight (pending theo watermark), gây data loss nghiêm trọng trong streaming. GCP cảnh báo Cancel chỉ dùng cho test/debug, không cho production update.

  • Stop the Cloud Dataflow pipeline with the Drain option. Create a new Cloud Dataflow job with the updated code
    ✅ Đúng: Như giải thích ở trên, Drain xử lý hết dữ liệu backlog trước khi dừng, đảm bảo no data loss. Job mới sẽ tiếp nhận từ watermark cuối cùng của job cũ (nếu dùng cùng source như Pub/Sub). Hoàn hảo cho thay đổi windowing/triggering theo GCP best practices 2026.

🧩 Tóm tắt nhanh: Chọn Drain + new job là cách an toàn nhất cho streaming Dataflow với thay đổi lớn, tránh rủi ro từ update inflight! 🚀

Câu 244
You need to move 2 PB of historical data from an on-premises storage appliance to Cloud Storage within six months, and your outbound network capacity is constrained to 20 Mb/sec. How should you migrate this data to Cloud Storage?
  1. A Use Transfer Appliance to copy the data to Cloud Storage
  2. B Use gsutil cp ג€"J to compress the content being uploaded to Cloud Storage
  3. C Create a private URL for the historical data, and then use Storage Transfer Service to copy the data to Cloud Storage
  4. D Use trickle or ionice along with gsutil cp to limit the amount of bandwidth gsutil utilizes to less than 20 Mb/sec so it does not interfere with the production traffic
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi này xoay quanh việc di chuyển một lượng dữ liệu lớn (2 PB dữ liệu lịch sử từ thiết bị lưu trữ on-premises sang Cloud Storage của Google Cloud) trong thời hạn 6 tháng, với giới hạn băng thông outbound chỉ 20 Mb/sec.

📊 Tính toán thời gian ước lượng để hiểu vấn đề:

  • 2 PB = 2.000.000 GB (khoảng 16.000.000.000.000 bytes).
  • 20 Mb/sec = 2.5 MB/sec (sau khi quy đổi).
  • Thời gian upload online: Khoảng 2.500-3.000 ngày (hơn 6-8 năm), vượt xa hạn 6 tháng! Vậy, cần giải pháp offline (vận chuyển vật lý) thay vì online để tránh bottleneck mạng. Đây là tình huống điển hình cho large-scale data migration trong Google Cloud (cập nhật đến 2026, Transfer Appliance vẫn là lựa chọn hàng đầu cho >1 PB với mạng yếu).

🛠️ Mục tiêu: Chọn phương pháp migrate hiệu quả, an toàn, tuân thủ SLA thời gian.

✅ Đáp án đúng

Use Transfer Appliance to copy the data to Cloud Storage

Lý do chọn đáp án này (🟢 Đúng tuyệt đối):
Transfer Appliance là dịch vụ offline migration của Google Cloud, sử dụng thiết bị phần cứng (như Octavia) ship đi on-premises. Bạn copy dữ liệu trực tiếp vào thiết bị (hỗ trợ 100-480 PB tùy model mới nhất 2026), ship ngược về Google data center, rồi tự động upload vào Cloud Storage.
✅ Ưu điểm nổi bật:

  • Bỏ qua hoàn toàn giới hạn mạng 20 Mb/sec.
  • Thời gian chỉ vài tuần (copy on-prem + ship).
  • An toàn dữ liệu (encryption, checksum), hỗ trợ resume nếu lỗi.
  • Phù hợp chính xác với 2 PB và hạn 6 tháng.
    📘 Nguồn tham khảo: Google Cloud Transfer Appliance Docs (2026 update) – Khuyến nghị cho datasets >150 TB với mạng <100 Mbps.

🔍 Giải thích tất cả các phương án (đúng & sai)

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên tính khả thi, hiệu suất và best practices Google Cloud (phiên bản mới nhất 2026).

  • ✅ Use Transfer Appliance to copy the data to Cloud Storage
    🟢 Đúng vì: Như giải thích trên, đây là giải pháp offline lý tưởng cho dữ liệu lớn + mạng yếu. Không phụ thuộc internet, tốc độ copy local lên đến 10-40 GB/sec trên thiết bị. Đã được chứng minh qua hàng nghìn case thực tế (Google re:Invent 2025 case studies).

  • ❌ Use gsutil cp -J to compress the content being uploaded to Cloud Storage
    🔴 Sai vì: gsutil cp -J chỉ nén dữ liệu trước khi upload online (parallel compression), giảm kích thước ~30-50% tùy content, nhưng vẫn phải truyền qua mạng 20 Mb/sec. Thời gian vẫn hàng nghìn ngày, không đáp ứng 6 tháng. Ngoài ra, CPU on-prem có thể bottleneck khi nén 2 PB. Không giải quyết gốc rễ vấn đề bandwidth.

  • ❌ Create a private URL for the historical data, and then use Storage Transfer Service to copy the data to Cloud Storage
    🔴 Sai vì: Storage Transfer Service (STS) dùng cho transfer giữa các cloud storage (GCS-to-GCS, S3-to-GCS, HTTP/HTTPS URLs), KHÔNG hỗ trợ trực tiếp on-premises storage. "Private URL" từ appliance on-prem không public/reachable qua internet ổn định, và STS vẫn cần bandwidth outbound cao (dù hỗ trợ POS/ resumable). Với 2 PB, STS chỉ hiệu quả cho <100 TB online. (Cập nhật 2026: STS vẫn chưa hỗ trợ native on-prem appliances).

  • ❌ Use trickle or ionice along with gsutil cp to limit the amount of bandwidth gsutil utilizes to less than 20 Mb/sec so it does not interfere with the production traffic
    🔴 Sai vì: trickle/ionice chỉ giới hạn bandwidth của gsutil để tránh ảnh hưởng production traffic, nhưng vẫn upload online với tốc độ tối đa 20 Mb/sec → thời gian vượt 6 tháng rất xa. Không tăng tốc độ tổng thể, chỉ "throttle" thêm, lãng phí thời gian. gsutil tốt cho small-scale, không scale cho PB-level.

📚 Tài liệu tham khảo bổ sung

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần demo hoặc lab, hỏi thêm nhé!

Câu 245
You receive data files in CSV format monthly from a third party. You need to cleanse this data, but every third month the schema of the files changes. Your requirements for implementing these transformations include:
✑ Executing the transformations on a schedule
✑ Enabling non-developer analysts to modify transformations
✑ Providing a graphical tool for designing transformations
What should you do?
  1. A Use Dataprep by Trifacta to build and maintain the transformation recipes, and execute them on a scheduled basis
  2. B Load each month's CSV data into BigQuery, and write a SQL query to transform the data to a standard schema. Merge the transformed tables together with a SQL query
  3. C Help the analysts write a Dataflow pipeline in Python to perform the transformation. The Python code should be stored in a revision control system and modified as the incoming data's schema changes
  4. D Use Apache Spark on Dataproc to infer the schema of the CSV file before creating a Dataframe. Then implement the transformations in Spark SQL before writing the data out to Cloud Storage and loading into BigQuery
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống: Bạn nhận dữ liệu dưới dạng file CSV hàng tháng từ một bên thứ ba. Nhiệm vụ chính là làm sạch (cleanse) dữ liệu, nhưng mỗi ba tháng một lần, schema của file thay đổi. Các yêu cầu cụ thể cho việc triển khai transformation bao gồm:

  • Thực thi transformation theo lịch trình (Executing on a schedule) 📅.
  • Cho phép các analyst không phải lập trình viên (non-developer analysts) chỉnh sửa transformation 👥.
  • Cung cấp công cụ giao diện đồ họa (graphical tool) để thiết kế transformation 🎨.

Mục tiêu là chọn giải pháp phù hợp nhất trên Google Cloud Platform (GCP) để xử lý dữ liệu linh hoạt, dễ sử dụng mà không cần code phức tạp. Đây là câu hỏi điển hình trong kỳ thi Google Cloud Professional Data Engineer, nhấn mạnh vào các dịch vụ ETL/ELT no-code/low-code. (Kiến thức cập nhật đến 2026: Dataprep by Trifacta vẫn là lựa chọn hàng đầu cho data wrangling với giao diện kéo-thả, tích hợp schedule qua Cloud Composer hoặc trực tiếp).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Dataprep by Trifacta to build and maintain the transformation recipes, and execute them on a scheduled basis

Lý do chi tiết:

  • 🛠️ Dataprep by Trifacta là công cụ no-code/low-code với giao diện đồ họa trực quan (visual wrangling), cho phép analysts kéo-thả để thiết kế recipe transformation (như clean, join, schema mapping) mà không cần code.
  • 📅 Hỗ trợ schedule tự động qua tích hợp với Cloud Scheduler hoặc Airflow/Cloud Composer.
  • 🔄 Xử lý schema thay đổi linh hoạt: Trifacta tự động infer schema CSV và cho phép chỉnh sửa recipe dễ dàng mỗi 3 tháng.
  • 👥 Hoàn hảo cho non-developer: Analysts có thể maintain recipe mà không cần dev skills.
  • Đây là giải pháp best practice của GCP cho data prep hàng loạt (batch).
    📘 Tài liệu tham khảo: Dataprep by Trifacta Docs & GCP Data Engineer Exam Guide 2024-2026.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên các yêu cầu chính: schedule, non-dev friendly, graphical tool.

  • Use Dataprep by Trifacta to build and maintain the transformation recipes, and execute them on a scheduled basis
    ✅ Đúng hoàn toàn 🏆: Như đã giải thích ở trên, đáp ứng 100% yêu cầu với giao diện đồ họa, schedule dễ dàng, và linh hoạt cho schema thay đổi. Không cần code, lý tưởng cho analysts.

  • Load each month's CSV data into BigQuery, and write a SQL query to transform the data to a standard schema. Merge the transformed tables together with a SQL query
    ❌ Sai: BigQuery mạnh về SQL analytics nhưng không có graphical tool – analysts phải viết SQL thủ công, không phù hợp non-developer. Schema thay đổi yêu cầu edit query lặp lại, khó maintain. Schedule có thể qua Scheduled Queries nhưng thiếu visual design. Không linh hoạt cho cleanse phức tạp.

  • Help the analysts write a Dataflow pipeline in Python to perform the transformation. The Python code should be stored in a revision control system and modified as the incoming data's schema changes
    ❌ Sai: Dataflow (Apache Beam) là code-heavy (Python/Java), yêu cầu dev skills cao – analysts non-dev không thể tự modify. Có schedule qua templates, nhưng không graphical. Schema thay đổi cần edit code và CI/CD, phức tạp và không thân thiện.

  • Use Apache Spark on Dataproc to infer the schema of the CSV file before creating a Dataframe. Then implement the transformations in Spark SQL before writing the data out to Cloud Storage and loading into BigQuery
    ❌ Sai: Dataproc + Spark infer schema tốt và hỗ trợ Spark SQL, nhưng giao diện là code-based (Scala/Python/SQL), không có graphical tool visual. Analysts non-dev khó chỉnh sửa. Schedule qua Dataproc workflows, nhưng tổng thể quá technical và tốn tài nguyên cho job hàng tháng đơn giản.

Kết luận tổng quát 🎯: Chỉ Dataprep by Trifacta mới đáp ứng đầy đủ 3 yêu cầu cốt lõi, giúp quy trình scalable và user-friendly trên GCP. Các lựa chọn khác phù hợp dev team hơn! 🚀

Câu 246 Chọn nhiều đáp án
You want to migrate an on-premises Hadoop system to Cloud Dataproc. Hive is the primary tool in use, and the data format is Optimized Row Columnar (ORC).
All ORC files have been successfully copied to a Cloud Storage bucket. You need to replicate some data to the cluster's local Hadoop Distributed File System
(HDFS) to maximize performance. What are two ways to start using Hive in Cloud Dataproc? (Choose two.)
  1. A Run the gsutil utility to transfer all ORC files from the Cloud Storage bucket to HDFS. Mount the Hive tables locally.
  2. B Run the gsutil utility to transfer all ORC files from the Cloud Storage bucket to any node of the Dataproc cluster. Mount the Hive tables locally.
  3. C Run the gsutil utility to transfer all ORC files from the Cloud Storage bucket to the master node of the Dataproc cluster. Then run the Hadoop utility to copy them do HDFS. Mount the Hive tables from HDFS.
  4. D Leverage Cloud Storage connector for Hadoop to mount the ORC files as external Hive tables. Replicate external Hive tables to the native ones.
  5. E Load the ORC files into BigQuery. Leverage BigQuery connector for Hadoop to mount the BigQuery tables as external Hive tables. Replicate external Hive tables to the native ones.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi thuộc chủ đề di chuyển hệ thống Hadoop on-premises sang Google Cloud Dataproc (dịch vụ quản lý Hadoop/Spark trên Google Cloud).

  • Bối cảnh chính: Hệ thống đang sử dụng Hive làm công cụ chính để xử lý dữ liệu ở định dạng ORC (Optimized Row Columnar) – một định dạng columnar hiệu suất cao cho Hive. Tất cả file ORC đã được copy thành công lên Cloud Storage bucket (GCS).
  • Yêu cầu cụ thể: Cần replicate (sao chép một phần dữ liệu) sang HDFS cục bộ của cluster Dataproc để tối ưu hiệu suất (vì truy cập HDFS local nhanh hơn đọc từ GCS qua network).
  • Mục tiêu: Tìm 2 cách để bắt đầu sử dụng Hive trên Dataproc, đảm bảo dữ liệu được replicate vào HDFS và Hive có thể query hiệu quả.
  • Lưu ý kỹ thuật (cập nhật đến 2026): Dataproc phiên bản mới nhất (2.1+) tích hợp sẵn Cloud Storage connector (hadoop3-gcs-connector), hỗ trợ Hive đọc trực tiếp gs:// paths với định dạng ORC. Managed Hive tables lưu dữ liệu trên HDFS (/user/hive/warehouse/), external tables có thể point đến GCS hoặc HDFS. Để replicate, có thể dùng INSERT INTO native_table SELECT * FROM external_table hoặc công cụ copy như hadoop distcp.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Lựa chọn thứ 1 và lựa chọn thứ 4.

  • 🛠️ Lý do: Hai cách này trực tiếp hỗ trợ replicate dữ liệu ORC từ GCS sang HDFS cục bộ của Dataproc cluster, sau đó cấu hình Hive tables để query hiệu suất cao.
    • Lựa chọn 1: Copy file trực tiếp vào HDFS rồi tạo Hive table local (trên HDFS).
    • Lựa chọn 4: Sử dụng connector GCS để tạo external table (không copy ngay), sau đó replicate dữ liệu sang native table (tự động copy vào HDFS).
  • Những cách khác không hiệu quả, không chuẩn hoặc không khả thi (chi tiết bên dưới). Điều này phù hợp best practice Dataproc 2026: Ưu tiên connector cho access nhanh, replicate chọn lọc cho performance.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn giữ nguyên văn bản gốc tiếng Anh, với giải thích hoàn toàn bằng tiếng Việt sử dụng kiến thức Dataproc mới nhất:

  • Run the gsutil utility to transfer all ORC files from the Cloud Storage bucket to HDFS. Mount the Hive tables locally.
    ✅ Đúng.
    🧩 Giải thích: gsutil (utility chính thức của Google) có thể dùng để copy file ORC từ gs://bucket trực tiếp vào HDFS của cluster (qua scheme hdfs:// trong môi trường Dataproc với connector đã enable). Sau khi copy, tạo Hive table local (managed hoặc external) pointing đến path HDFS (CREATE TABLE ... LOCATION 'hdfs://path/to/orc'). Cách này replicate toàn bộ dữ liệu vào HDFS cục bộ, tối ưu performance cho query lớn. Trong Dataproc 2.0+, gsutil tích hợp tốt với Hadoop filesystem, tránh network latency từ GCS. Best cho full migration.

  • Run the gsutil utility to transfer all ORC files from the Cloud Storage bucket to any node of the Dataproc cluster. Mount the Hive tables locally.
    ❌ Sai.
    🧩 Giải thích: Copy bằng gsutil vào local disk của bất kỳ node nào (ví dụ /tmp/ trên worker/master) không đáng tin cậy vì local disk không replicated (dữ liệu mất khi node fail hoặc cluster scale). Hive không khuyến khích "mount locally" trên local FS (chỉ hỗ trợ tốt HDFS/GCS). Không replicate đúng vào HDFS cluster, vi phạm yêu cầu performance và durability.

  • Run the gsutil utility to transfer all ORC files from the Cloud Storage bucket to the master node of the Dataproc cluster. Then run the Hadoop utility to copy them do HDFS. Mount the Hive tables from HDFS.
    ❌ Sai.
    🧩 Giải thích: Bước 1 copy gsutil vào master node local disk (single point of failure, master không dùng lưu data lớn). Bước 2 dùng "Hadoop utility" (như hdfs dfs -put) để copy sang HDFS khả thi nhưng không efficient cho dữ liệu lớn (ORC files), tốn thời gian double copy và rủi ro overload master. Dataproc khuyến dùng hadoop distcp gs:// hdfs:// trực tiếp thay vì qua master. "do HDFS" có lẽ lỗi đánh máy của "to HDFS".

  • Leverage Cloud Storage connector for Hadoop to mount the ORC files as external Hive tables. Replicate external Hive tables to the native ones.
    ✅ Đúng.
    🧩 Giải thích: Sử dụng GCS connector (mặc định trong Dataproc) tạo external Hive table (CREATE EXTERNAL TABLE ... LOCATION 'gs://bucket/path' STORED AS ORC). Sau đó replicate bằng INSERT INTO native_table SELECT * FROM external_table – tự động copy dữ liệu ORC sang native/managed table trên HDFS (/user/hive/warehouse/). Cách này selectively replicate (chỉ dữ liệu cần performance cao), hỗ trợ ORC schema inference tốt, không cần copy thủ công ban đầu. Best practice 2026 cho hybrid access (query GCS trước, replicate sau).

  • Load the ORC files into BigQuery. Leverage BigQuery connector for Hadoop to mount the BigQuery tables as external Hive tables. Replicate external Hive tables to the native ones.
    ❌ Sai.
    🧩 Giải thích: Load ORC vào BigQuery (dùng bq load --source_format=ORC) thừa thãi, tốn chi phí và thời gian transform (BigQuery không native ORC như Hive). BigQuery connector (cho Hadoop/Hive) tồn tại nhưng chỉ hỗ trợ Hive đọc BQ tables như external (qua LOCATION 'bq://project.dataset.table'), không phù hợp migrate Hadoop-Hive thuần. Replicate sang native vẫn ok nhưng toàn bộ flow không cần thiết, phức tạp và không tối ưu performance so với GCS direct.

🏆 Kết luận & Lời khuyên

✅ Chọn lựa chọn 1 và 4 để bắt đầu Hive nhanh, an toàn trên Dataproc.
🛠️ Thực hành: Tạo cluster Dataproc với Hive (init action nếu cần), test hadoop distcp bổ sung cho large-scale replicate. Theo dõi cost HDFS storage (local nhưng cluster terminate mất data trừ khi persistent).
📘 Nâng cao: Sử dụng Dataproc Serverless (2025+) cho Hive workflows không cần HDFS full.

Câu 247
You are implementing several batch jobs that must be executed on a schedule. These jobs have many interdependent steps that must be executed in a specific order. Portions of the jobs involve executing shell scripts, running Hadoop jobs, and running queries in BigQuery. The jobs are expected to run for many minutes up to several hours. If the steps fail, they must be retried a fixed number of times. Which service should you use to manage the execution of these jobs?
  1. A Cloud Scheduler
  2. B Cloud Dataflow
  3. C Cloud Functions
  4. D Cloud Composer
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống triển khai nhiều batch jobs (công việc xử lý hàng loạt) cần chạy theo lịch trình cố định. Các jobs này có nhiều bước phụ thuộc lẫn nhau (interdependent steps), phải thực hiện theo thứ tự cụ thể. Các bước bao gồm:

  • Chạy shell scripts (script dòng lệnh).
  • Chạy Hadoop jobs (xử lý dữ liệu lớn kiểu MapReduce).
  • Chạy queries trên BigQuery (kho dữ liệu phân tích của Google Cloud).

Jobs chạy lâu từ vài phút đến vài giờ, và nếu bước nào thất bại, cần retry (thử lại) một số lần cố định.

Mục tiêu: Chọn dịch vụ Google Cloud phù hợp để quản lý và orchestrate (điều phối) việc thực thi toàn bộ quy trình này một cách tự động, đáng tin cậy.

🛠️ Yêu cầu chính của giải pháp:

  • Hỗ trợ lập lịch (scheduling).
  • Xử lý dependencies (phụ thuộc thứ tự).
  • Tích hợp đa dạng tasks (shell, Hadoop, BigQuery).
  • Hỗ trợ retry tự động.
  • Phù hợp jobs dài hạn (không phải serverless ngắn).

📘 Kiến thức cập nhật (tính đến 2026): Dựa trên Google Cloud docs mới nhất (Cloud Composer 3.x dựa trên Apache Airflow 2.9+), dịch vụ này hỗ trợ đầy đủ orchestration workflows phức tạp với DAGs (Directed Acyclic Graphs), sensors cho dependencies, và operators cho shell/Hadoop/BigQuery. (Nguồn: Google Cloud Composer Documentation, Apache Airflow Providers).

✅ Đáp án đúng: Cloud Composer

Lý do chọn: Cloud Composer là dịch vụ managed Apache Airflow trên Google Cloud, lý tưởng cho orchestration workflows phức tạp. Nó hỗ trợ:

  • DAGs để định nghĩa thứ tự steps với dependencies rõ ràng (ví dụ: task shell script → Hadoop job → BigQuery query).
  • Lập lịch cron cho batch jobs.
  • Operators tích hợp sẵn cho Bash (shell scripts), Hadoop, BigQuery (qua BigQueryOperator/SQL).
  • Retry tự động với retries cố định và backoff.
  • Chạy jobs dài (giờ) trên môi trường Kubernetes, scale tự động. Không dịch vụ nào khác xử lý đầy đủ dependencies + đa dạng tasks + retry + scheduling như Composer.

📋 Giải thích tất cả các phương án (đúng/sai)

  • Cloud Scheduler ❌ [SAI]
    Cloud Scheduler chỉ dùng để lập lịch kích hoạt các jobs đơn giản (gửi HTTP requests hoặc Pub/Sub messages theo cron). Nó không hỗ trợ orchestration phức tạp, không quản lý dependencies giữa steps, không chạy shell/Hadoop/BigQuery trực tiếp, và không có retry logic cho workflows dài. Phù hợp jobs độc lập ngắn, không phải batch jobs phụ thuộc. (Nguồn: Cloud Scheduler Docs).

  • Cloud Dataflow ❌ [SAI]
    Cloud Dataflow là dịch vụ batch/stream processing dựa trên Apache Beam, chuyên pipeline dữ liệu lớn (ETL). Nó không phải orchestration tool, khó định nghĩa thứ tự phụ thuộc tùy chỉnh với shell scripts/Hadoop, retry chỉ cho data processing chứ không cho workflows đa dạng. Jobs Dataflow chạy parallel, không phù hợp sequential steps với BigQuery queries riêng lẻ. (Nguồn: Dataflow Docs).

  • Cloud Functions ❌ [SAI]
    Cloud Functions là serverless functions chạy code ngắn (tối đa 60 phút/execution, khuyến nghị <9 phút). Nó không hỗ trợ jobs dài (giờ), không orchestrate dependencies phức tạp, không chạy Hadoop/shell trực tiếp dễ dàng, và retry chỉ cơ bản cho invocations. Không phù hợp batch jobs nặng. (Nguồn: Cloud Functions Docs, cập nhật 2nd gen đến 2026).

  • Cloud Composer ✅ [ĐÚNG]
    Như đã giải thích trên: Hoàn hảo cho orchestration batch jobs phức tạp với scheduling, DAGs dependencies, đa operators (shell/BashOperator, Hadoop via KubernetesPodOperator, BigQueryOperator), retry retries, và scale cho jobs dài. Đây là lựa chọn chuẩn theo best practices Google Cloud cho workflows dữ liệu. (Nguồn: Composer Best Practices).

🛠️ Khuyến nghị triển khai: Sử dụng DAG Python trong Composer với BashOperator cho shell, KubernetesPodOperator cho Hadoop, BigQueryOperator cho queries. Set retries=3 và retry_delay cho mỗi task. Test trên môi trường dev trước khi prod!

Câu 248
You work for a shipping company that has distribution centers where packages move on delivery lines to route them properly. The company wants to add cameras to the delivery lines to detect and track any visual damage to the packages in transit. You need to create a way to automate the detection of damaged packages and flag them for human review in real time while the packages are in transit. Which solution should you choose?
  1. A Use BigQuery machine learning to be able to train the model at scale, so you can analyze the packages in batches.
  2. B Train an AutoML model on your corpus of images, and build an API around that model to integrate with the package tracking applications.
  3. C Use the Cloud Vision API to detect for damage, and raise an alert through Cloud Functions. Integrate the package tracking applications with this function.
  4. D Use TensorFlow to create a model that is trained on your corpus of images. Create a Python notebook in Cloud Datalab that uses this model so you can analyze for damaged packages.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một công ty vận chuyển hàng hóa có các trung tâm phân phối, nơi các gói hàng di chuyển trên dây chuyền giao hàng. Công ty muốn lắp đặt camera trên dây chuyền để phát hiện và theo dõi hư hỏng trực quan (visual damage) trên gói hàng trong quá trình vận chuyển. Yêu cầu chính là tự động hóa việc phát hiện gói hàng hỏng và đánh dấu để con người xem xét ngay lập tức (real-time) trong khi gói hàng vẫn đang di chuyển.

🛠️ Yêu cầu cốt lõi:

  • Real-time processing: Phân tích hình ảnh từ camera liên tục, không batch.
  • Custom detection: Phát hiện "damage" cụ thể, cần huấn luyện mô hình trên dữ liệu hình ảnh (corpus of images) của công ty.
  • Tích hợp: Kết nối với ứng dụng theo dõi gói hàng để flag và alert ngay lập tức.
  • Giải pháp Google Cloud: Sử dụng các dịch vụ ML/AI như AutoML, Vision API để xử lý hình ảnh thời gian thực.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Train an AutoML model on your corpus of images, and build an API around that model to integrate with the package tracking applications.

Lý do 🏆:

  • AutoML Vision (nay là phần của Vertex AI) cho phép huấn luyện mô hình tùy chỉnh (custom model) trên bộ dữ liệu hình ảnh của công ty mà không cần chuyên gia ML sâu, phù hợp cho object detection/classification damage.
  • Xây dựng API (qua Vertex AI Prediction/Endpoints) cho phép inference real-time từ camera stream, tích hợp dễ dàng với app tracking.
  • Đáp ứng real-time và automation hoàn hảo: Camera gửi ảnh → API predict → Flag nếu damage → Human review ngay.
  • Cập nhật 2026: Vertex AI AutoML vẫn hỗ trợ, với cải tiến như Edge deployment cho low-latency (theo docs GCP 2024-2026).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phần giải thích hoàn toàn bằng tiếng Việt với lý do đúng/sai dựa trên kiến thức GCP mới nhất (Vertex AI, BigQuery ML v2026).

  • ❌ [SAI] Use BigQuery machine learning to be able to train the model at scale, so you can analyze the packages in batches.
    Lý do sai 🚫: BigQuery ML chỉ phù hợp cho batch processing và phân tích dữ liệu lớn (SQL-based ML), không hỗ trợ real-time inference từ camera stream. Huấn luyện "at scale" nhưng analyze "in batches" vi phạm yêu cầu real-time trong transit. Không tối ưu cho computer vision (hình ảnh).

  • ✅ [ĐÚNG] Train an AutoML model on your corpus of images, and build an API around that model to integrate with the package tracking applications.
    Lý do đúng 🎯: Như đã giải thích ở trên. AutoML/Vertex AI Prediction API hỗ trợ custom training trên ảnh damage, deploy endpoint real-time (low-latency <1s), tích hợp seamless với apps. Hoàn hảo cho scenario camera real-time.

  • ❌ [SAI] Use the Cloud Vision API to detect for damage, and raise an alert through Cloud Functions. Integrate the package tracking applications with this function.
    Lý do sai 🚫: Cloud Vision API là pre-trained cho general tasks (label detection, object recognition), không customize cho "damage" cụ thể mà không train thêm. Không dùng corpus images của bạn. Cloud Functions chỉ trigger alert, nhưng accuracy thấp cho custom damage → Không reliable real-time.

  • ❌ [SAI] Use TensorFlow to create a model that is trained on your corpus of images. Create a Python notebook in Cloud Datalab that uses this model so you can analyze for damaged packages.
    Lý do sai 🚫: TensorFlow tốt cho custom model, nhưng Cloud Datalab (nay Vertex AI Workbench) chỉ dùng cho development/notebook prototyping, không deploy real-time API. Phân tích thủ công/batch, không automate/integrate với camera tracking apps. Không scalable cho production real-time (cập nhật 2026: Datalab deprecated, chuyển sang Vertex Workbench).

📘 Tài liệu tham khảo (cập nhật mới nhất GCP đến 2026)

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm chi tiết, hỏi nhé!

Câu 249
You are migrating your data warehouse to BigQuery. You have migrated all of your data into tables in a dataset. Multiple users from your organization will be using the data. They should only see certain tables based on their team membership. How should you set user permissions?
  1. A Assign the users/groups data viewer access at the table level for each table
  2. B Create SQL views for each team in the same dataset in which the data resides, and assign the users/groups data viewer access to the SQL views
  3. C Create authorized views for each team in the same dataset in which the data resides, and assign the users/groups data viewer access to the authorized views
  4. D Create authorized views for each team in datasets created for each team. Assign the authorized views data viewer access to the dataset in which the data resides. Assign the users/groups data viewer access to the datasets in which the authorized views reside
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc thiết lập quyền truy cập (permissions) cho người dùng trong BigQuery sau khi migrate data warehouse. Cụ thể:

  • Dữ liệu đã được migrate vào các bảng (tables) trong một dataset.
  • Nhiều người dùng từ tổ chức sẽ sử dụng dữ liệu, nhưng mỗi người chỉ nên thấy các bảng cụ thể dựa trên membership team của họ (ví dụ: team A chỉ thấy bảng của team A, team B chỉ thấy bảng của team B).
  • Mục tiêu: Kiểm soát granular (chi tiết) ở mức bảng (table-level) để tránh người dùng thấy toàn bộ dataset hoặc các bảng không liên quan.
  • Đây là tính năng IAM (Identity and Access Management) của BigQuery (Google Cloud), không liên quan đến AWS (có thể là nhầm lẫn trong mô tả). BigQuery hỗ trợ IAM ở nhiều mức: project, dataset, table, view (từ phiên bản cập nhật 2019 và ổn định đến 2026).

Vấn đề cốt lõi: Cần giải pháp đơn giản, an toàn, granular nhất mà không cần tạo thêm views phức tạp, tận dụng table-level IAM (đã hỗ trợ đầy đủ từ BigQuery v1.0+).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Assign the users/groups data viewer access at the table level for each table

Lý do:

  • 🛠️ BigQuery hỗ trợ IAM permissions ở mức table-level (từ năm 2019, cập nhật đầy đủ đến 2026), cho phép gán quyền roles/bigquery.dataViewer trực tiếp cho từng bảng cụ thể.
  • ✅ Điều này chính xác đáp ứng yêu cầu: Mỗi team/group chỉ được gán quyền viewer cho các bảng liên quan, họ không thấy bảng khác dù cùng dataset. Không cần tạo thêm objects (views), tránh overhead.
  • An toàn & hiệu quả: Granular nhất, dễ quản lý qua Google Cloud Console, gcloud CLI hoặc Terraform. User chỉ query được bảng được phép, không truy cập dataset-wide.
  • Phiên bản mới nhất (2026): Table-level IAM vẫn là best practice cho row-of-access control cơ bản (theo Google Cloud docs).

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm lý do bằng tiếng Việt:

  • ✅ Assign the users/groups data viewer access at the table level for each table
    🛠️ Đúng hoàn toàn: Như giải thích ở trên, đây là cách trực tiếp, granular, không cần workaround. User/group chỉ thấy và query được bảng được gán quyền, lý tưởng cho "certain tables based on team".

  • ❌ Create SQL views for each team in the same dataset in which the data resides, and assign the users/groups data viewer access to the SQL views
    🧩 Sai vì: SQL views thông thường (routine views) không che giấu dữ liệu gốc. Nếu user có quyền dataViewer trên dataset, họ vẫn query trực tiếp tables gốc qua INFORMATION_SCHEMA hoặc SQL, bỏ qua views. Không restrict "only see certain tables".

  • ❌ Create authorized views for each team in the same dataset in which the data resides, and assign the users/groups data viewer access to the authorized views
    🧩 Sai vì: Authorized views (views được ủy quyền) chủ yếu dùng để cross-dataset/project sharing an toàn, che giấu base tables. Nhưng trong cùng dataset, nếu user có quyền dataset-level, họ vẫn thấy và query tables gốc. Không giải quyết "only see certain tables" hiệu quả, phức tạp hóa không cần thiết.

  • ❌ Create authorized views for each team in datasets created for each team. Assign the authorized views data viewer access to the dataset in which the data resides. Assign the users/groups data viewer access to the datasets in which the authorized views reside
    🧩 Sai vì: Logic rối loạn – gán dataViewer cho authorized views vào dataset gốc (nơi data resides) sẽ cho views quyền truy cập tables, nhưng user chỉ cần quyền trên dataset views. Vấn đề: Tạo nhiều dataset thừa (overhead chi phí), không granular ở table-level, và không ngăn user thấy tables nếu có quyền gián tiếp. Phức tạp, không scale tốt.

📘 Tài liệu tham khảo (cập nhật đến 2026)

  • BigQuery IAM Documentation: Cloud IAM for BigQuery – Chi tiết table-level permissions.
  • Authorized Views Guide: Authorized Views – Giải thích hạn chế của views trong same dataset.
  • Best Practices: BigQuery Security Whitepaper (2025 update) – Khuyến nghị table-level cho team-based access.
  • gcloud CLI: gcloud bigquery tables iam add cho table-level IAM.

Kết luận 💡: Table-level IAM là giải pháp tối ưu, native của BigQuery! Nếu cần row-level, kết hợp Authorized Views hoặc Column-level security (mới 2024+).

Câu 250
You want to build a managed Hadoop system as your data lake. The data transformation process is composed of a series of Hadoop jobs executed in sequence.
To accomplish the design of separating storage from compute, you decided to use the Cloud Storage connector to store all input data, output data, and intermediary data. However, you noticed that one Hadoop job runs very slowly with Cloud Dataproc, when compared with the on-premises bare-metal Hadoop environment (8-core nodes with 100-GB RAM). Analysis shows that this particular Hadoop job is disk I/O intensive. You want to resolve the issue. What should you do?
  1. A Allocate sufficient memory to the Hadoop cluster, so that the intermediary data of that particular Hadoop job can be held in memory
  2. B Allocate sufficient persistent disk space to the Hadoop cluster, and store the intermediate data of that particular Hadoop job on native HDFS
  3. C Allocate more CPU cores of the virtual machine instances of the Hadoop cluster so that the networking bandwidth for each instance can scale up
  4. D Allocate additional network interface card (NIC), and configure link aggregation in the operating system to use the combined throughput when working with Cloud Storage
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh việc xây dựng một hệ thống Hadoop được quản lý (managed Hadoop) trên Google Cloud Dataproc làm data lake, với thiết kế tách biệt storage và compute bằng cách sử dụng Cloud Storage connector để lưu trữ tất cả dữ liệu đầu vào (input data), đầu ra (output data) và dữ liệu trung gian (intermediary data). 📁

Quá trình biến đổi dữ liệu (data transformation) gồm chuỗi các Hadoop jobs chạy tuần tự. Tuy nhiên, một Hadoop job cụ thể chạy rất chậm trên Dataproc so với môi trường on-premises bare-metal Hadoop (máy chủ 8-core, 100GB RAM). Phân tích cho thấy job này disk I/O intensive (tốn nhiều I/O đĩa). 🎯

Vấn đề cốt lõi: Cloud Storage là object storage, truy cập qua network I/O (chậm hơn local disk), dẫn đến bottleneck cho job I/O-intensive. Mục tiêu là giải quyết vấn đề mà không thay đổi thiết kế tổng thể (vẫn tách storage-compute cho input/output). 🛠️

(Kiến thức cập nhật: Theo tài liệu Google Cloud Dataproc phiên bản mới nhất 2024-2026, Cloud Storage connector tối ưu cho throughput lớn nhưng latency cao cho small/random I/O. Best practice: Sử dụng local HDFS cho intermediate data trong workload I/O-intensive. Nguồn: Dataproc Documentation - Performance Tuning, HDFS on Dataproc.)

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Allocate sufficient persistent disk space to the Hadoop cluster, and store the intermediate data of that particular Hadoop job on native HDFS.

Lý do:

  • Job disk I/O intensive cần local disk access nhanh (low-latency), không phù hợp với Cloud Storage (network-bound).
  • Persistent Disk (PD) trên Dataproc cho phép mount native HDFS trên local SSD/HDD, tăng tốc intermediate data (spill từ memory sang disk).
  • Giữ nguyên input/output trên Cloud Storage (tách storage-compute), chỉ chuyển intermediate data sang HDFS để tránh network overhead.
  • Hiệu quả: Tăng performance lên gấp nhiều lần cho I/O-intensive jobs, tương đương on-prem. 🚀 (Nguồn: Dataproc SSD Persistent Disk, cập nhật 2025 hỗ trợ balanced/extreme PD cho high IOPS).

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Allocate sufficient memory to the Hadoop cluster, so that the intermediary data of that particular Hadoop job can be held in memory
    Giải thích sai: Tăng memory chỉ giúp giữ intermediate data trong RAM (shuffle/spill), nhưng job disk I/O intensive vẫn phải đọc/ghi đĩa nhiều (không chỉ memory-bound). Nếu data vượt memory, vẫn spill ra Cloud Storage chậm. Không giải quyết root cause network I/O. 💥 (Không phải best practice cho I/O-heavy; chỉ tối ưu memory-bound jobs).

  • ✅ [ĐÚNG] Allocate sufficient persistent disk space to the Hadoop cluster, and store the intermediate data of that particular Hadoop job on native HDFS
    Giải thích đúng: Như phần trên, PD + native HDFS cung cấp local disk I/O nhanh (high IOPS, low latency) cho intermediate data, bypass Cloud Storage cho phần bottleneck. Giữ tách storage-compute tổng thể. Hiệu suất gần on-prem. 🏆 (Best practice chính thức GCP).

  • ❌ [SAI] Allocate more CPU cores of the virtual machine instances of the Hadoop cluster so that the networking bandwidth for each instance can scale up
    Giải thích sai: Tăng CPU cores không scale networking bandwidth trực tiếp (bandwidth GCP VM scale theo machine type, không phải cores). Vấn đề là disk I/O, không phải CPU/network capacity. Chỉ làm tốn chi phí vô ích. 🚫 (Nguồn: GCP Compute Engine Networking Limits).

  • ❌ [SAI] Allocate additional network interface card (NIC), and configure link aggregation in the operating system to use the combined throughput when working with Cloud Storage
    Giải thích sai: Multi-NIC + link aggregation tăng throughput network, nhưng không giảm latency cho random I/O của Cloud Storage (vẫn network-bound). Phức tạp config, không giải quyết disk I/O root cause (Cloud Storage vẫn chậm cho small reads/writes). Không khuyến nghị cho Hadoop. 🔌 (GCP hỗ trợ multi-NIC premium tier, nhưng không fix I/O issue; nguồn: Dataproc Networking).