Ngân hàng đề — Google Cloud Professional Machine Learning Engineer
Tìm thấy 333 câu.
- A Use BigQuery’s scheduling service to run the model retraining query periodically.
- B Create a pipeline in Vertex AI Pipelines that executes the retraining query, and use the Cloud Scheduler API to run the query weekly.
- C Use Cloud Scheduler to trigger a Cloud Function every week that runs the query for retraining the model.
- D Use the BigQuery API Connector and Cloud Scheduler to trigger Workflows every week that retrains the model.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc retrain (huấn luyện lại) một mô hình Linear Regression được tạo bằng BigQuery ML trên dữ liệu tích lũy hàng tuần. Yêu cầu chính là tối ưu hóa nỗ lực phát triển (development effort) và chi phí lập lịch (scheduling cost).
- Bối cảnh: BigQuery ML cho phép huấn luyện mô hình ML trực tiếp trong BigQuery mà không cần di chuyển dữ liệu. Để retrain, bạn chạy lại câu query
CREATE OR REPLACE MODELtrên dữ liệu mới (cumulative data). - Thách thức: Cần tự động hóa quy trình hàng tuần, nhưng phải đơn giản nhất (ít code, ít dịch vụ bên ngoài) để giảm effort và cost.
- Kiến thức cập nhật (đến 2026): BigQuery hỗ trợ Scheduled Queries (tính năng native từ 2018, cải tiến liên tục qua BigQuery Reservation API và Dataform cho orchestration). Đây là cách tối ưu nhất cho query-based ML như BigQuery ML, không cần pipeline phức tạp. (Nguồn: BigQuery Scheduled Queries Documentation - phiên bản mới nhất 2026).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use BigQuery’s scheduling service to run the model retraining query periodically.
Lý do 🛠️:
- BigQuery có Scheduled Queries tích hợp sẵn, cho phép lập lịch chạy query (như
CREATE OR REPLACE MODEL) hàng tuần chỉ qua UI hoặc SQL (gõscheduletrong query editor). - Minimize development effort: Không cần code thêm, không deploy service nào – chỉ click vài nút hoặc dùng
bq queryCLI. - Minimize scheduling cost: Miễn phí (chỉ tính phí query execution), không tốn tài nguyên Cloud Functions/Workflows/Pipelines.
- Hoàn hảo cho BigQuery ML vì query retrain đơn giản, dữ liệu cumulative sẵn trong BigQuery.
📋 Giải thích tất cả các phương án (đúng/sai)
-
✅ Use BigQuery’s scheduling service to run the model retraining query periodically.
Đúng 🏆: Như phân tích trên, đây là giải pháp native, zero-code, low-cost nhất. Scheduled Queries hỗ trợ cron-like scheduling (e.g., weekly), tự động retrain model bằngCREATE OR REPLACE. (Nguồn: BigQuery ML Training Docs). -
❌ Create a pipeline in Vertex AI Pipelines that executes the retraining query, and use the Cloud Scheduler API to run the query weekly.
Sai 🚫: Vertex AI Pipelines dành cho ML workflow phức tạp (nhiều bước: data prep, training, deployment). Ở đây chỉ cần query đơn giản → overkill, tốn effort build pipeline YAML/Kubeflow, tốn cost Vertex AI (Pipeline runs ~$0.03/giờ + query fees). Cloud Scheduler thêm layer không cần. -
❌ Use Cloud Scheduler to trigger a Cloud Function every week that runs the query for retraining the model.
Sai ⚠️: Cloud Functions cần code Python/Node.js để gọi BigQuery API (bq_client.query()), deploy function → tăng development effort cao. Cost: Functions invocations (~$0.0000025/lần) + query + cold start latency. Phức tạp hơn Scheduled Queries native. -
❌ Use the BigQuery API Connector and Cloud Scheduler to trigger Workflows every week that retrains the model.
Sai 🔧: Workflows (Cloud Workflows) dùng cho orchestration đa bước, cần YAML workflow definition + BigQuery API calls. Tốn effort thiết kế workflow, cost (~$0.000025/step) + Scheduler. Không đơn giản bằng Scheduled Queries trực tiếp.
📘 Tài liệu tham khảo chính (cập nhật 2026)
- BigQuery Scheduled Queries – Hướng dẫn lập lịch query ML.
- BigQuery ML Retraining Best Practices – Khuyến nghị dùng Scheduled Queries.
- Vertex AI vs BigQuery ML Comparison – Xác nhận Scheduled Queries cho simple retraining.
- AWS? Câu hỏi là Google Cloud thuần túy (BigQuery ML), không liên quan AWS (như SageMaker). Nếu nhầm, SageMaker tương đương dùng EventBridge + Lambda, nhưng kém optimal hơn.
Hy vọng phân tích giúp bạn ôn thi hiệu quả! 🚀
- A Use the aiplatform.log_classification_metrics function to log the F1 score, and use the aiplatform.log_metrics function to log the confusion matrix.
- B Use the aiplatform.log_classification_metrics function to log the F1 score and the confusion matrix.
- C Use the aiplatform.log_metrics function to log the F1 score and the confusion matrix.
- D Use the aiplatform.log_metrics function to log the F1 score: and use the aiplatform.log_classification_metrics function to log the confusion matrix.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào việc di chuyển (migrate) một mô hình phân loại scikit-learn sang TensorFlow trên Google Cloud Vertex AI. Bạn sẽ huấn luyện mô hình TensorFlow mới bằng cùng bộ dữ liệu huấn luyện (training set) như mô hình scikit-learn cũ, sau đó so sánh hiệu suất của hai mô hình trên một bộ test chung dựa trên F1 score (một chỉ số scalar đo độ chính xác cân bằng giữa precision và recall) và confusion matrix (ma trận nhầm lẫn, hiển thị chi tiết phân loại đúng/sai theo từng lớp).
Để thực hiện, bạn sử dụng Vertex AI Python SDK để ghi log thủ công (manually log) các metrics này vào Vertex AI Experiments, nhằm dễ dàng so sánh trong giao diện Vertex AI UI (như bảng experiment, biểu đồ, visualization). Câu hỏi yêu cầu cách log đúng để hỗ trợ so sánh F1 score (dùng cho ranking/sort) và confusion matrix (dùng cho visualization riêng).
📘 Lưu ý kiến thức cập nhật (đến 2026): Theo Vertex AI v2024+ (tương tự v2025-2026), aiplatform.log_metrics dùng cho metrics scalar (float), trong khi aiplatform.log_classification_metrics hỗ trợ cả scalar classification (như F1) lẫn confusion matrix (ma trận 2D/list/dict, được visualize đặc biệt). Các scalar từ classification metrics được log nội bộ qua log_metrics để so sánh.
✅ Đáp án đúng
Use the aiplatform.log_metrics function to log the F1 score: and use the aiplatform.log_classification_metrics function to log the confusion matrix.
Lý do lựa chọn:
- F1 score là metric scalar (float), nên dùng
log_metrics({'f1_score': value})để log độc lập, dễ so sánh trực tiếp trong experiment comparison UI (hỗ trợ sorting, charting theo F1 của các run/model). - Confusion matrix là ma trận phức tạp (không phải scalar), chỉ hỗ trợ log qua
log_classification_metrics(confusion_matrix=cm_matrix)để Vertex AI tự động visualize (như heatmap). - Cách này cho phép log riêng biệt, linh hoạt so sánh cả hai metrics mà không lẫn lộn context, phù hợp "manually log" và "compare based on their F1 scores and confusion matrices". Nếu log F1 qua classification, nó vẫn OK nhưng không tách biệt tối ưu cho comparison scalar thuần.
🔍 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Tôi đánh dấu ✅/❌ dựa trên tính đúng đắn theo docs Vertex AI mới nhất.
-
❌ [SAI] Use the aiplatform.log_classification_metrics function to log the F1 score, and use the aiplatform.log_metrics function to log the confusion matrix.
Giải thích sai:log_classification_metricscó thể log F1 score (qua paramf1_score), nhưnglog_metricschỉ chấp nhận Dict[str, float] (scalar), không hỗ trợ confusion matrix (ma trận 2D). Log confusion matrix vào đây sẽ báo lỗi hoặc không visualize đúng. -
❌ [SAI] Use the aiplatform.log_classification_metrics function to log the F1 score and the confusion matrix.
Giải thích sai: Mặc dù function này hỗ trợ cả F1 (f1_score) và confusion matrix trong một lệnh gọi duy nhất, nhưng câu hỏi yêu cầu log riêng biệt để so sánh thủ công (F1 như scalar độc lập, confusion matrix như viz riêng). Log chung sẽ nhóm chúng vào "classification metrics context", làm khó compare F1 trực tiếp như metric chính (không tối ưu cho sorting F1 across models). -
❌ [SAI] Use the aiplatform.log_metrics function to log the F1 score and the confusion matrix.
Giải thích sai:log_metricschỉ log scalar metrics (float), không hỗ trợ confusion matrix (yêu cầu cấu trúc ma trận). Thử log confusion matrix sẽ thất bại vì không khớp Dict[str, float], dẫn đến mất visualization. -
✅ [ĐÚNG] Use the aiplatform.log_metrics function to log the F1 score: and use the aiplatform.log_classification_metrics function to log the confusion matrix.
Giải thích đúng (như phần trên): Phân tách rõ ràng – scalar F1 qualog_metricscho comparison số học, ma trận qualog_classification_metricscho viz. Hoàn hảo cho migrate/compare hai models trên Vertex AI Experiments.
🛠️ Ví dụ code minh họa (dùng trong job/train custom)
import aiplatform
from sklearn.metrics import f1_score, confusion_matrix
# Log F1 score (scalar)
aiplatform.log_metrics({'f1_score': f1_value}) # e.g., 0.85
# Log confusion matrix
cm = confusion_matrix(y_true, y_pred, labels=classes)
aiplatform.log_classification_metrics(confusion_matrix=cm.tolist())
📘 Tài liệu tham khảo
- Chính thức Google Cloud Vertex AI Experiments: Log metrics Python SDK & Classification metrics (cập nhật Q4 2024, áp dụng đến 2026).
- Comparison UI: Vertex AI Experiments docs – Hỗ trợ sort bằng scalar metrics như F1, viz confusion matrix riêng.
- SDK Reference: aiplatform Python client v1.40+ (2025-2026 tương thích).
Hy vọng phân tích giúp bạn nắm vững! 🚀 Nếu cần code đầy đủ hoặc lab Vertex AI, hỏi thêm nhé!
- A Include a comprehensive set of demographic features
- B Include only the demographic groups that most frequently interact with advertisements
- C Collect a random sample of production traffic to build the training dataset
- D Collect a stratified sample of production traffic to build the training dataset
- E Conduct fairness tests across sensitive categories and demographics on the trained model
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc phát triển một mô hình machine learning (ML) để hỗ trợ công ty tạo các chiến dịch quảng cáo trực tuyến nhắm mục tiêu (targeted online advertising campaigns) hiệu quả hơn. 🔍 Nhiệm vụ chính là tạo dataset huấn luyện mô hình mà tránh tạo ra hoặc củng cố bias không công bằng (unfair bias). Bias ở đây ám chỉ sự thiên vị trong dữ liệu dẫn đến mô hình phân biệt đối xử không mong muốn với các nhóm dân số nhạy cảm (sensitive groups như giới tính, độ tuổi, dân tộc...).
Câu hỏi yêu cầu chọn hai phương án đúng từ các lựa chọn, dựa trên các best practices trong ML engineering để đảm bảo tính công bằng (fairness). 🛡️️ Chủ đề liên quan đến nguyên tắc Responsible AI/ML trên AWS (cập nhật đến phiên bản mới nhất 2026, như tích hợp trong Amazon SageMaker Clarify cho fairness evaluation). Mục tiêu là xây dựng dataset đại diện (representative) và kiểm tra fairness sau huấn luyện.
✅ Đáp án đúng (Chọn hai phương án)
Hai lựa chọn đúng là:
- Collect a stratified sample of production traffic to build the training dataset
Lý do: Phân tầng mẫu (stratified sampling) đảm bảo dataset bao quát đầy đủ các nhóm dân số (strata) một cách cân bằng, tránh underrepresentation của nhóm thiểu số → Giảm bias từ dữ liệu gốc. 🏗️ - Conduct fairness tests across sensitive categories and demographics on the trained model
Lý do: Kiểm tra fairness trên các hạng mục nhạy cảm (như tuổi, giới tính) sau huấn luyện giúp phát hiện và khắc phục bias còn sót lại, tuân thủ AWS SageMaker Clarify metrics (như demographic parity). 📊
📋 Giải thích tất cả các phương án (Đúng và Sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên nguyên tắc tránh bias theo AWS ML best practices (2026). ❌ Sai vì có thể làm tăng bias; ✅ Đúng vì thúc đẩy fairness.
-
Include a comprehensive set of demographic features
❌ Sai: Việc bao gồm đầy đủ đặc trưng nhân khẩu học (demographic features như tuổi, giới tính, dân tộc) có thể củng cố bias hiện có trong dữ liệu, vì mô hình học theo các pattern thiên vị từ lịch sử (ví dụ: quảng cáo chỉ nhắm nhóm trẻ tuổi). AWS khuyến cáo không sử dụng trực tiếp sensitive features mà nên engineer proxy features hoặc loại bỏ để tránh proxy discrimination. 🛑 -
Include only the demographic groups that most frequently interact with advertisements
❌ Sai: Chỉ lấy nhóm tương tác nhiều nhất (majority groups) sẽ tạo imbalance nghiêm trọng, dẫn đến mô hình bỏ qua nhóm thiểu số → Bias cao, vi phạm nguyên tắc representative data. Điều này giống under-sampling minorities, không phù hợp với targeted advertising cần đa dạng. 🚫 -
Collect a random sample of production traffic to build the training dataset
❌ Sai: Mẫu ngẫu nhiên (random sample) từ traffic sản xuất có thể không đại diện cho các nhóm nhỏ (long-tail distributions), dẫn đến underrepresentation → Bias tự nhiên từ dữ liệu thực tế. AWS khuyên dùng stratified/random stratified thay vì pure random cho fairness. 🎲 -
Collect a stratified sample of production traffic to build the training dataset
✅ Đúng: Phân tầng theo các nhóm (stratified) đảm bảo tỷ lệ cân bằng giữa các strata nhạy cảm từ traffic thực tế → Dataset representative, giảm selection bias. Đây là best practice trong AWS SageMaker Data Wrangler (2026) cho ML pipelines. 🎯 -
Conduct fairness tests across sensitive categories and demographics on the trained model
✅ Đúng: Thực hiện kiểm tra fairness (fairness tests) trên mô hình đã huấn luyện với metrics như Equalized Odds/Opportunity giúp phát hiện và mitigate bias. AWS SageMaker Clarify hỗ trợ tự động hóa điều này với built-in fairness analyzers (cập nhật 2026). 🔬
📘 Tài liệu tham khảo
- AWS Documentation: Amazon SageMaker Clarify - Fairness Measures (v1.5+, 2026 updates on stratified sampling).
- AWS ML Best Practices: Responsible AI on AWS – Nhấn mạnh stratified sampling & post-training audits.
- Google Cloud ML Fairness (tương đương): Vertex AI Fairness Indicators, nhưng áp dụng AWS context. 🌐
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần ví dụ code SageMaker, hãy hỏi thêm.
-
A
1. Initialize the Vertex SDK with the name of your experiment. Log parameters and metrics for each experiment, and attach dataset and model artifacts as inputs and outputs to each execution.
2. After a successful experiment create a Vertex AI pipeline. -
B
1. Initialize the Vertex SDK with the name of your experiment. Log parameters and metrics for each experiment, save your dataset to a Cloud Storage bucket, and upload the models to Vertex AI Model Registry.
2. After a successful experiment, create a Vertex AI pipeline. -
C
1. Create a Vertex AI pipeline with parameters you want to track as arguments to your PipelineJob. Use the Metrics, Model, and Dataset artifact types from the Kubeflow Pipelines DSL as the inputs and outputs of the components in your pipeline.
2. Associate the pipeline with your experiment when you submit the job. -
D
1. Create a Vertex AI pipeline. Use the Dataset and Model artifact types from the Kubeflow Pipelines DSL as the inputs and outputs of the components in your pipeline.
2. In your training component, use the Vertex AI SDK to create an experiment run. Configure the log_params and log_metrics functions to track parameters and metrics of your experiment.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào việc phát triển mô hình ML trong Vertex AI Workbench notebook (một môi trường notebook dựa trên Jupyter trên Google Cloud). Yêu cầu chính là:
- Theo dõi artifacts (như dataset, model) và so sánh các mô hình trong quá trình thử nghiệm với các cách tiếp cận khác nhau.
- Chuyển tiếp nhanh chóng và dễ dàng các thử nghiệm thành công sang production khi lặp lại (iterate) trên việc triển khai mô hình. Mục tiêu là sử dụng các công cụ của Vertex AI Experiments và pipelines để quản lý thí nghiệm linh hoạt, theo dõi metrics/params/artifacts, và dễ dàng scale lên production mà không cần viết lại code nhiều. Đây là tính năng cốt lõi của Vertex AI (cập nhật đến 2026, hỗ trợ experiment tracking tích hợp với SDK và pipelines).
📘 Tài liệu tham khảo: Vertex AI Experiments, Vertex AI Pipelines.
✅ Đáp án đúng: Phương án 1
Lý do chọn: Phương án này sử dụng Vertex AI Experiments SDK một cách chính xác nhất trong notebook để khởi tạo experiment, log params/metrics, và attach artifacts (dataset/model) trực tiếp vào executions – giúp dễ dàng so sánh experiments và track lineage. Sau experiment thành công, tạo Vertex AI pipeline để chuyển sang production nhanh chóng (pipeline hỗ trợ compile từ notebook code và deploy tự động). Điều này phù hợp hoàn hảo với yêu cầu "track artifacts, compare models during experimentation" và "rapidly transition to production". Không cần pipeline từ đầu, giữ sự linh hoạt cho notebook experimentation.
🛠️ Ưu điểm: Tích hợp native với Workbench, hỗ trợ comparison qua UI, và pipeline là bước tự nhiên để productionize.
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng phương án. Tôi giữ nguyên nội dung văn bản gốc bằng tiếng Anh, chỉ giải thích lý do đúng/sai bằng tiếng Việt với emoji nổi bật.
-
Phương án 1 (ĐÚNG):
- Initialize the Vertex SDK with the name of your experiment. Log parameters and metrics for each experiment, and attach dataset and model artifacts as inputs and outputs to each execution.
- After a successful experiment create a Vertex AI pipeline.
✅ Giải thích đúng: Như đã nêu ở trên, đây là quy trình chuẩn cho experimentation trong notebook: SDK cho phép attach artifacts trực tiếp vào experiment runs, hỗ trợ comparison và visualization. Pipeline sau đó dùng để automate production, tận dụng code/notebook hiện có.
-
Phương án 2 (SAI):
- Initialize the Vertex SDK with the name of your experiment. Log parameters and metrics for each experiment, save your dataset to a Cloud Storage bucket, and upload the models to Vertex AI Model Registry.
- After a successful experiment, create a Vertex AI pipeline.
❌ Giải thích sai: Phần 1 sai vì chỉ save dataset vào Cloud Storage và upload model vào Model Registry không tạo lineage tracking đầy đủ cho artifacts trong experiments (không attach trực tiếp vào executions). Điều này làm khó so sánh models/artifacts giữa các approaches, thiếu tính năng core của Experiments SDK. Phần 2 đúng nhưng không cứu vãn được lỗi chính.
-
Phương án 3 (SAI):
- Create a Vertex AI pipeline with parameters you want to track as arguments to your PipelineJob. Use the Metrics, Model, and Dataset artifact types from the Kubeflow Pipelines DSL as the inputs and outputs of the components in your pipeline.
- Associate the pipeline with your experiment when you submit the job.
❌ Giải thích sai: Bắt đầu bằng pipeline từ đầu không phù hợp cho notebook experimentation linh hoạt (quá rigid, khó iterate nhanh trong Workbench). Kubeflow DSL artifacts tốt cho pipeline nhưng không hỗ trợ comparison experiments dễ dàng như SDK trực tiếp. "Associate pipeline with experiment" chỉ là tracking cơ bản, không đáp ứng "track artifacts and compare models during different approaches" một cách rapidly.
-
Phương án 4 (SAI):
- Create a Vertex AI pipeline. Use the Dataset and Model artifact types from the Kubeflow Pipelines DSL as the inputs and outputs of the components in your pipeline.
- In your training component, use the Vertex AI SDK to create an experiment run. Configure the log_params and log_metrics functions to track parameters and metrics of your experiment.
❌ Giải thích sai: Tương tự phương án 3, tạo pipeline trước làm phức tạp hóa notebook dev (không rapid cho iteration). Phần 2 dùng SDK bên trong training component của pipeline là lộn xộn, không phải cách chuẩn – SDK dành cho notebook experiments độc lập, không integrate mượt mà như attach trực tiếp vào executions. Thiếu hỗ trợ đầy đủ cho artifacts tracking và production transition.
🏆 Kết luận
Phương án 1 là lựa chọn tối ưu, tận dụng Vertex AI Experiments + Pipelines một cách native và hiệu quả nhất (cập nhật 2026: hỗ trợ AI Platform -> Vertex AI full migration). Sử dụng nó để tránh vendor lock-in và scale dễ dàng! 🚀
📘 Nguồn bổ sung: Vertex AI SDK Experiments, Best Practices for ML Experimentation.
- A Ensure that the Workbench instance that you created is in the same region of the Vertex AI Pipelines resources you will use.
- B Ensure that the Vertex AI Workbench instance is on the same subnetwork of the Vertex AI Pipeline resources that you will use.
- C Ensure that the Vertex AI Workbench instance is assigned the Identity and Access Management (IAM) Vertex AI User role.
- D Ensure that the Vertex AI Workbench instance is assigned the Identity and Access Management (IAM) Notebooks Runner role.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả tình huống: Bạn vừa tạo một dự án Google Cloud mới. Bạn đã thử nghiệm thành công việc submit một Vertex AI Pipeline job từ Cloud Shell (môi trường shell trên trình duyệt với quyền của user hiện tại). Tuy nhiên, khi sử dụng Vertex AI Workbench user-managed notebook instance (một instance notebook do user tự quản lý) để chạy code tương tự, job thất bại với lỗi "insufficient permissions" (quyền không đủ).
Vấn đề cốt lõi: Cloud Shell sử dụng credentials của user IAM cá nhân (thường có quyền đầy đủ nếu bạn là owner/project editor), nhưng Vertex AI Workbench instance chạy dưới service account mặc định của instance (thường là Compute Engine default service account), nên cần cấp quyền IAM cụ thể cho service account đó để gọi API Vertex AI Pipelines (như aiplatform.pipelines.create).
📘 Tài liệu tham khảo:
- Vertex AI Workbench documentation (cập nhật 2024-2026).
- IAM roles for Vertex AI (roles/aiplatform.user cho pipelines).
- Troubleshoot Vertex AI Pipelines.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Ensure that the Vertex AI Workbench instance is assigned the Identity and Access Management (IAM) Vertex AI User role.
Lý do:
- Vertex AI Workbench user-managed notebook sử dụng service account của instance (thường là
project-number-compute@developer.gserviceaccount.com) để authenticate các API calls. - Để submit Vertex AI Pipeline job, service account cần quyền Vertex AI User role (roles/aiplatform.user), bao gồm các permissions như
aiplatform.pipelines.create,aiplatform.pipelines.get,aiplatform.runs.create. - Cloud Shell OK vì dùng user credentials (có quyền project-level), nhưng notebook cần cấp role này trực tiếp cho instance/service account qua IAM policy.
- 🛠️ Cách thực hiện: Vào IAM & Admin > IAM, chọn service account của instance, add role
Vertex AI User.
✅ Điều này giải quyết chính xác lỗi permissions mà không ảnh hưởng region/network (vì đã test từ Cloud Shell cùng project).
📋 Giải thích tất cả các phương án
-
❌ Phương án SAI: Ensure that the Workbench instance that you created is in the same region of the Vertex AI Pipelines resources you will use.
Lý do sai: Region matching không phải nguyên nhân permissions error. Vertex AI Pipelines hỗ trợ multi-region access qua API (dùng endpoint global nhưus-central1-aiplatform.googleapis.com), và Cloud Shell đã submit thành công từ region khác nếu có. Permissions là vấn đề IAM, không liên quan region (theo docs Vertex AI multi-regional support, cập nhật 2025). -
❌ Phương án SAI: Ensure that the Vertex AI Workbench instance is on the same subnetwork of the Vertex AI Pipeline resources that you will use.
Lý do sai: Subnetwork/VPC peering chỉ cần cho private endpoints hoặc custom networking (Vertex AI hỗ trợ public API mặc định). Lỗi "insufficient permissions" là IAM/auth error (403 Forbidden), không phải network connectivity (như timeout/DNS). Cloud Shell dùng public internet, nên nếu network issue thì Cloud Shell cũng fail. -
✅ Phương án ĐÚNG: Ensure that the Vertex AI Workbench instance is assigned the Identity and Access Management (IAM) Vertex AI User role.
Lý do đúng: Như giải thích ở phần đáp án đúng. Role này cung cấp chính xác permissions cần thiết cho pipelines mà không over-privileged (principle of least privilege). Đã được confirm trong troubleshooting docs Vertex AI (2026 updates không thay đổi core IAM). -
❌ Phương án SAI: Ensure that the Vertex AI Workbench instance is assigned the Identity and Access Management (IAM) Notebooks Runner role.
Lý do sai: Không tồn tại role chính thức tên "Notebooks Runner" trong IAM Vertex AI (có thể nhầm vớiroles/aiplatform.notebookUserhoặc legacy roles). Các role notebook-focused nhưVertex AI Notebook Userchỉ cho quản lý notebooks (create/start instance), không bao gồm permissions cho Pipelines (aiplatform.pipelines.*). Sử dụng sẽ vẫn fail permissions.
🧠 Lưu ý bổ sung: Sau khi add role, restart kernel notebook hoặc regenerate access token nếu cần. Test bằng gcloud auth list trong notebook để verify service account.
- A Use Vertex AI Data Labeling Service to label the images, and tram an AutoML image classification model. Deploy the model, and configure Pub/Sub to publish a message when an image is categorized into the failing class.
- B Use Vertex AI Data Labeling Service to label the images, and train an AutoML image classification model. Schedule a daily batch prediction job that publishes a Pub/Sub message when the job completes.
- C Convert the images into an embedding representation. Import this data into BigQuery, and train a BigQuery ML K-means clustering model with two clusters. Deploy the model and configure Pub/Sub to publish a message when a semiconductor’s data is categorized into the failing cluster.
- D Import the tabular data into BigQuery, use Vertex AI Data Labeling Service to label the data and train an AutoML tabular classification model. Deploy the model, and configure Pub/Sub to publish a message when a semiconductor’s data is categorized into the failing class.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một tình huống thực tế trong ngành sản xuất bán dẫn:
- 📸 Dữ liệu đầu vào: Ảnh độ phân giải cao (high-definition images) của từng bán dẫn được chụp real-time ở cuối dây chuyền lắp ráp.
- ☁️ Lưu trữ: Ảnh và dữ liệu bảng (tabular data: batch number, serial number, dimensions, weight) được upload lên Cloud Storage bucket.
- 🎯 Yêu cầu chính: Xây dựng ứng dụng real-time tự động hóa kiểm soát chất lượng (quality control), cấu hình model training và serving để tối đa hóa độ chính xác model (maximizing model accuracy).
- 🚀 Ngữ cảnh: Cần xử lý real-time, phát hiện sản phẩm lỗi (failing class/cluster), và tích hợp thông báo (Pub/Sub) để kích hoạt hành động ngay lập tức.
Đây là bài toán image classification supervised learning, ưu tiên độ chính xác cao với dữ liệu ảnh chính (vì chất lượng bán dẫn thường kiểm tra qua hình ảnh vi mô), kết hợp metadata tabular.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Vertex AI Data Labeling Service to label the images, and train an AutoML image classification model. Deploy the model, and configure Pub/Sub to publish a message when an image is categorized into the failing class.
🛠️ Lý do chọn đáp án này (toàn diện và phù hợp nhất):
- Sử dụng Vertex AI Data Labeling Service để gắn nhãn ảnh chính xác, hỗ trợ human-in-the-loop cho dữ liệu chất lượng cao.
- AutoML Image Classification tự động train model với độ chính xác cao (state-of-the-art trên Vertex AI đến 2026), tối ưu cho ảnh HD.
- Deploy model cho online prediction real-time, tích hợp Pub/Sub trigger ngay khi phân loại vào lớp "failing" (lỗi) – đảm bảo xử lý real-time.
- Kết hợp metadata tabular làm context (nhưng focus ảnh). Đây là giải pháp end-to-end trên Google Cloud, max accuracy mà không cần code phức tạp.
📘 Tài liệu tham khảo:
- Vertex AI AutoML Image (cập nhật 2025).
- Vertex AI Data Labeling.
- Real-time prediction with Pub/Sub.
🔍 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên yêu cầu real-time, max accuracy, và sử dụng dữ liệu ảnh chính.
-
Use Vertex AI Data Labeling Service to label the images, and train an AutoML image classification model. Deploy the model, and configure Pub/Sub to publish a message when an image is categorized into the failing class.
✅ Đúng: Như giải thích trên, đây là giải pháp hoàn hảo – labeling ảnh, AutoML cho accuracy cao, deploy real-time với Pub/Sub trigger tức thì khi detect failing class. Phù hợp 100% với real-time quality control trên ảnh HD. -
Use Vertex AI Data Labeling Service to label the images, and train an AutoML image classification model. Schedule a daily batch prediction job that publishes a Pub/Sub message when the job completes.
❌ Sai: Labeling và train AutoML tốt, nhưng schedule daily batch prediction chỉ chạy hàng ngày – KHÔNG real-time (chậm trễ 24h), không phù hợp với yêu cầu "end of the assembly line in real time". Pub/Sub chỉ trigger sau job hoàn thành, miss mục tiêu tự động hóa tức thì. -
Convert the images into an embedding representation. Import this data into BigQuery, and train a BigQuery ML K-means clustering model with two clusters. Deploy the model and configure Pub/Sub to publish a message when a semiconductor’s data is categorized into the failing cluster.
❌ Sai: Chuyển ảnh thành embedding (có thể dùng Vertex AI Embeddings) rồi train BigQuery ML K-means (unsupervised clustering) chỉ phân 2 cụm (pass/fail) – accuracy thấp vì không supervised, không tận dụng labeling, và embedding mất chi tiết ảnh HD. Real-time deploy OK nhưng không max accuracy; tabular data bị bỏ qua một phần. -
Import the tabular data into BigQuery, use Vertex AI Data Labeling Service to label the data and train an AutoML tabular classification model. Deploy the model, and configure Pub/Sub to publish a message when a semiconductor’s data is categorized into the failing class.
❌ Sai: Chỉ dùng tabular data (batch, serial, dimensions, weight) bỏ qua ảnh HD chính – KHÔNG max accuracy cho quality control (chất lượng bán dẫn cần inspect hình ảnh vi mô). AutoML Tabular OK cho dữ liệu số nhưng không thay thế image analysis; labeling tabular kém hiệu quả hơn ảnh.
🧮 Tóm tắt so sánh nhanh:
| Tiêu chí | Đúng | B | C | D |
|----------|------|---|----|---|
| Real-time | ✅ | ❌ | ✅ | ✅ |
| Max accuracy (ảnh focus) | ✅ | ✅ | ❌ | ❌ |
| Supervised + Labeling | ✅ | ✅ | ❌ | ✅ (nhưng tabular) |
Giải pháp đúng cân bằng hoàn hảo cho Google Cloud Vertex AI ecosystem! 🚀
- A Deploy the training jobs by using TPU VMs with TPUv3 Pod slices, and use the TPUEmbeading API
- B Deploy the training jobs in an autoscaling Google Kubernetes Engine cluster with CPUs
- C Deploy a matrix factorization model training job by using BigQuery ML
- D Deploy the training jobs by using Compute Engine instances with A100 GPUs, and use the tf.nn.embedding_lookup API
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một công ty mạng xã hội đang phát triển nhanh chóng, đội ngũ đang xây dựng các mô hình khuyến nghị (recommender models) bằng TensorFlow trên cụm CPU on-premises. Dữ liệu bao gồm hàng tỷ sự kiện lịch sử người dùng (billions of historical user events) và 100.000 đặc trưng phân loại (categorical features). Vấn đề là thời gian huấn luyện mô hình tăng theo dữ liệu. Nhiệm vụ là di chuyển mô hình lên Google Cloud, chọn cách có khả năng mở rộng nhất (most scalable) và giảm thiểu thời gian huấn luyện (minimizes training time) nhất.
🔍 Yêu cầu chính:
- Xử lý dữ liệu khổng lồ với embedding lớn (do categorical features nhiều).
- Sử dụng TensorFlow → ưu tiên hạ tầng tối ưu cho deep learning lớn.
- Tập trung vào scalability (mở rộng theo dữ liệu tăng) và tốc độ huấn luyện nhanh.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Deploy the training jobs by using TPU VMs with TPUv3 Pod slices, and use the TPUEmbeading API (lưu ý: "TPUEmbeading API" có thể là lỗi chính tả của TPU Embedding API – API chuyên xử lý embedding cho TPU).
Lý do chọn đáp án này 🏆:
- TPU VMs với TPUv3 Pod slices là hạ tầng siêu mở rộng (scalable) nhất trên Google Cloud cho huấn luyện mô hình lớn, hỗ trợ hàng nghìn chip TPU kết nối qua Pod (lên đến 4096 chip TPUv3), lý tưởng cho dữ liệu billions events và 100k categorical features.
- TPU Embedding API (tích hợp TensorFlow) tối ưu hóa sparse embedding – vấn đề cốt lõi của recommender models với categorical features lớn, giảm thời gian huấn luyện đáng kể so với CPU/GPU thông thường (có thể nhanh gấp 10-100x).
- Phù hợp mới nhất 2026: TPUv3 Pod vẫn là lựa chọn scalable cao cho embedding-heavy workloads (Vertex AI hỗ trợ TPUv5e/Pod mới hơn, nhưng TPUv3 Pod slice vẫn chuẩn cho TensorFlow large-scale).
- 📘 Nguồn tham khảo:
- Google Cloud TPU Embedding API Docs (cập nhật 2025).
- Vertex AI Training with TPUs (phiên bản mới nhất hỗ trợ Pod slices).
🛠️ Giải thích tất cả các phương án (đúng/sai)
-
Phương án đúng ✅: Deploy the training jobs by using TPU VMs with TPUv3 Pod slices, and use the TPUEmbeading API
Giải thích: Đây là lựa chọn tối ưu nhất vì TPU Pod slices cho phép scale ngang cực lớn (distributed training), kết hợp TPU Embedding API xử lý embedding sparse hiệu quả cho 100k features và billions data. Giảm training time tối đa, phù hợp TensorFlow recommender. Không có lựa chọn nào scalable hơn! -
Phương án sai ❌: Deploy the training jobs in an autoscaling Google Kubernetes Engine cluster with CPUs
Giải thích: GKE autoscaling với CPU chỉ scale theo node số lượng, nhưng CPU chậm hơn TPU/GPU rất nhiều cho deep learning (đặc biệt embedding). Không giảm training time hiệu quả với dữ liệu lớn, vẫn gặp bottleneck như on-premises. -
Phương án sai ❌: Deploy a matrix factorization model training job by using BigQuery ML
Giải thích: BigQuery ML chỉ hỗ trợ matrix factorization đơn giản (như ALS), không phù hợp TensorFlow recommender phức tạp với 100k features và billions events. Không scalable cho custom deep learning models, training time vẫn lâu và thiếu flexibility. -
Phương án sai ❌: Deploy the training jobs by using Compute Engine instances with A100 GPUs, and use the tf.nn.embedding_lookup API
Giải thích: A100 GPUs tốt cho training, nhưng ít scalable hơn TPU Pod (scale giới hạn ~thousands GPUs vs. TPU Pod).tf.nn.embedding_lookupkhông tối ưu sparse embedding như TPU Embedding API, dẫn đến training time lâu hơn với categorical features lớn.
Kết luận 🎯: Chọn TPU để maximize scalability & minimize time – chuẩn cho Google Cloud ML Engineer! Nếu cần code sample, hỏi thêm nhé! 🚀
- A Use Vertex Al Model Monitoring. Enable prediction drift monitoring on the endpoint, and specify a notification email.
- B In Cloud Logging, create a logs-based alert using the logs in the Vertex Al endpoint. Configure Cloud Logging to send an email when the alert is triggered.
- C In Cloud Monitoring create a logs-based metric and a threshold alert for the metric. Configure Cloud Monitoring to send an email when the alert is triggered.
- D Export the container logs of the endpoint to BigQuery. Create a Cloud Function to run a SQL query over the exported logs and send an email. Use Cloud Scheduler to trigger the Cloud Function.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm
📖 Nội dung câu hỏi:
Câu hỏi mô tả tình huống bạn đang huấn luyện và triển khai các phiên bản cập nhật của một mô hình hồi quy (regression model) sử dụng dữ liệu bảng (tabular data) trên các dịch vụ Google Cloud Vertex AI, bao gồm Vertex AI Pipelines (để tự động hóa quy trình huấn luyện), Vertex AI Training (huấn luyện mô hình), Vertex AI Experiments (quản lý thí nghiệm), và Vertex AI Endpoints (triển khai mô hình để inference). Mô hình đã được triển khai trên Vertex AI Endpoint, và người dùng gọi mô hình qua endpoint này.
Yêu cầu chính: Nhận email thông báo khi phân phối dữ liệu đặc trưng (feature data distribution) thay đổi đáng kể, để có thể kích hoạt lại pipeline huấn luyện và triển khai phiên bản mô hình mới.
🛠️ Vấn đề cốt lõi: Phát hiện sự thay đổi (drift) trong phân phối dữ liệu đầu vào (feature drift), một phần của model monitoring trong MLOps, để đảm bảo mô hình không bị lỗi thời (model staleness). Đây là tính năng native của Vertex AI, cập nhật đến năm 2026 với hỗ trợ monitoring drift toàn diện hơn (prediction drift, feature attribution drift).
✅ Đáp án đúng:
Use Vertex Al Model Monitoring. Enable prediction drift monitoring on the endpoint, and specify a notification email.
Lý do lựa chọn (chi tiết):
Vertex AI Model Monitoring (nay là Vertex AI Model Monitoring với tích hợp nâng cao từ 2023-2026) chính là giải pháp native và tối ưu nhất. Bạn có thể kích hoạt prediction drift monitoring trực tiếp trên endpoint, so sánh phân phối dự đoán thực tế với baseline (từ training data), gián tiếp phát hiện feature drift qua skew trong prediction. Hệ thống tự động gửi email alert qua Cloud Monitoring notifications khi drift vượt ngưỡng (threshold). Điều này tích hợp liền mạch với Vertex AI Pipelines để retrigger huấn luyện tự động. Không cần code custom, chi phí thấp, và hỗ trợ alerting qua email/Cloud Functions.
📘 Tài liệu tham khảo:
- Vertex AI Model Monitoring Documentation (cập nhật 2026: hỗ trợ skew detection cho tabular data).
- Prediction Drift Monitoring Guide.
🔍 Giải thích tất cả các phương án (đúng/sai)
-
✅ [ĐÚNG] Use Vertex Al Model Monitoring. Enable prediction drift monitoring on the endpoint, and specify a notification email.
🟢 Đúng vì: Đây là tính năng chuyên dụng của Vertex AI, hỗ trợ phát hiện drift (bao gồm feature/prediction skew) trên endpoint deployed model. Cấu hình đơn giản qua Console/CLI: enable monitoring job, set baseline từ training dataset, threshold (e.g., Jensen-Shannon divergence > 0.1), và tích hợp notification channel gửi email. Tự động hóa MLOps end-to-end. Hoàn hảo cho tabular regression models. -
❌ [SAI] In Cloud Logging, create a logs-based alert using the logs in the Vertex Al endpoint. Configure Cloud Logging to send an email when the alert is triggered.
🔴 Sai vì: Cloud Logging chỉ ghi logs thô từ endpoint (e.g., request/response), không có khả năng phân tích thống kê drift (như KS test hoặc PSI cho distribution shift). Logs-based alert chỉ phù hợp cho error logs hoặc simple patterns, không detect "feature data distribution changes significantly". Phải tự parse logs thủ công, không scalable và không native cho ML monitoring. -
❌ [SAI] In Cloud Monitoring create a logs-based metric and a threshold alert for the metric. Configure Cloud Monitoring to send an email when the alert is triggered.
🔴 Sai vì: Cloud Monitoring (trước là Stackdriver) hỗ trợ logs-based metrics (e.g., count errors), nhưng không built-in cho ML-specific metrics như drift detection. Bạn phải tự định nghĩa metric từ logs (e.g., custom feature stats), phức tạp và thiếu accuracy so với Vertex AI's statistical tests. Không giải quyết trực tiếp feature distribution shift mà chỉ threshold đơn giản, dễ false positive/negative. -
❌ [SAI] Export the container logs of the endpoint to BigQuery. Create a Cloud Function to run a SQL query over the exported logs and send an email. Use Cloud Scheduler to trigger the Cloud Function.
🔴 Sai vì: Quá phức tạp, tốn kém (BigQuery storage/query costs + Cloud Function invocations), và không realtime (Cloud Scheduler chạy định kỳ, e.g., hàng giờ). Phải tự implement drift detection qua SQL (e.g., aggregate features rồi tính KS statistic), thiếu tích hợp với Vertex AI ecosystem. Không khuyến khích vì Vertex AI Model Monitoring làm tốt hơn, native hơn từ 2023+.
🛡️ Khuyến nghị thực tế (MLOps best practices 2026): Sử dụng Vertex AI Feature Store kết hợp Model Monitoring để baseline features chính xác hơn, và tích hợp Eventarc để auto-retrigger pipelines khi alert fire. Điều này đảm bảo model freshness cho production tabular ML workloads! 🚀
-
A
1. Specify sampled Shapley as the explanation method with a path count of 5.
2. Deploy the model to Vertex AI Endpoints.
3. Create a Model Monitoring job that uses prediction drift as the monitoring objective. -
B
1. Specify Integrated Gradients as the explanation method with a path count of 5.
2. Deploy the model to Vertex AI Endpoints.
3. Create a Model Monitoring job that uses prediction drift as the monitoring objective. -
C
1. Specify sampled Shapley as the explanation method with a path count of 50.
2. Deploy the model to Vertex AI Endpoints.
3. Create a Model Monitoring job that uses training-serving skew as the monitoring objective. -
D
1. Specify Integrated Gradients as the explanation method with a path count of 50.
2. Deploy the model to Vertex AI Endpoints.
3. Create a Model Monitoring job that uses training-serving skew as the monitoring objective.
Xem giải thích
🧩 Giải thích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc triển khai mô hình XGBoost trên Vertex AI (dịch vụ Machine Learning của Google Cloud) để phục vụ online prediction (dự đoán thời gian thực). Các yêu cầu chính bao gồm:
- Upload mô hình vào Vertex AI Model Registry và cấu hình explanation method (phương pháp giải thích mô hình) sao cho minimal latency (độ trễ thấp nhất) khi trả về kết quả dự đoán kèm giải thích.
- Alert (cảnh báo) khi feature attributions (độ quan trọng của các đặc trưng) của mô hình thay đổi có ý nghĩa theo thời gian (ví dụ: do data drift hoặc skew).
- Đây là tình huống thực tế với mô hình tabular như XGBoost, nơi cần cân bằng giữa tốc độ giải thích và độ chính xác, đồng thời giám sát mô hình sau khi deploy lên Vertex AI Endpoints.
Vertex AI hỗ trợ hai explanation methods chính cho online predictions: Sampled Shapley (nhanh, phù hợp low-latency) và Integrated Gradients (chính xác hơn nhưng chậm). Ngoài ra, Model Monitoring jobs giúp phát hiện drift/skew để alert qua Cloud Monitoring/Logging. (Kiến thức cập nhật Vertex AI phiên bản mới nhất 2024-2026, theo docs GCP).
📘 Tài liệu tham khảo:
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Lựa chọn đầu tiên
- Specify sampled Shapley as the explanation method with a path count of 5.
- Deploy the model to Vertex AI Endpoints.
- Create a Model Monitoring job that uses prediction drift as the monitoring objective.
Lý do chi tiết 🛠️:
- Bước 1: Sampled Shapley là phương pháp nhanh nhất cho tabular models như XGBoost, sử dụng path count = 5 (thấp, giúp giảm số lượng đường dẫn tính toán → minimal latency dưới 100ms/request). Integrated Gradients chậm hơn vì cần nhiều steps integral. Path count thấp đảm bảo tốc độ cao mà vẫn đủ chính xác cho production.
- Bước 2: Bắt buộc phải deploy lên Vertex AI Endpoints để phục vụ online prediction với explanations.
- Bước 3: Prediction drift monitoring objective theo dõi sự thay đổi trong predictions và feature attributions over time (khi explanations enabled), tự động alert nếu attributions drift có ý nghĩa (threshold configurable). Điều này phù hợp để phát hiện thay đổi attributions do data distribution shift.
📋 Phân tích tất cả các phương án
Dưới đây là phân tích từng lựa chọn một cách chi tiết. Mỗi phương án gồm 3 bước, tôi giữ nguyên văn bản gốc tiếng Anh và giải thích tại sao đúng/sai bằng tiếng Việt:
-
✅ Phương án ĐÚNG (như đã giải thích ở trên):
- Specify sampled Shapley as the explanation method with a path count of 5.
- Deploy the model to Vertex AI Endpoints.
- Create a Model Monitoring job that uses prediction drift as the monitoring objective.
→ Hoàn hảo khớp yêu cầu: Low latency (Sampled Shapley + path=5), deploy đúng, và monitoring prediction drift phát hiện attribution changes hiệu quả.
-
❌ Phương án SAI 1:
- Specify Integrated Gradients as the explanation method with a path count of 5.
- Deploy the model to Vertex AI Endpoints.
- Create a Model Monitoring job that uses prediction drift as the monitoring objective.
→ Sai ở bước 1: Integrated Gradients chậm hơn đáng kể so với Sampled Shapley (cần nhiều integral steps, latency cao ngay cả path=5), không đáp ứng minimal latency cho online predictions trên XGBoost. Các bước còn lại đúng nhưng không cứu vãn.
-
❌ Phương án SAI 2:
- Specify sampled Shapley as the explanation method with a path count of 50.
- Deploy the model to Vertex AI Endpoints.
- Create a Model Monitoring job that uses training-serving skew as the monitoring objective.
→ Sai kép: Bước 1 dùng path count=50 quá cao → tăng số đường dẫn tính toán gấp 10 lần, latency cao (không minimal). Bước 3 dùng training-serving skew chỉ so sánh training vs serving data (phát hiện skew ban đầu), không monitor changes over time của attributions như prediction drift.
-
❌ Phương án SAI 3:
- Specify Integrated Gradients as the explanation method with a path count of 50.
- Deploy the model to Vertex AI Endpoints.
- Create a Model Monitoring job that uses training-serving skew as the monitoring objective.
→ Sai toàn diện: Bước 1 kết hợp Integrated Gradients + path=50 → latency cực cao (không phù hợp online). Bước 3 sai như trên (training-serving skew không monitor attribution drift over time). Chỉ bước 2 đúng.
Tóm tắt nhanh 🎯: Chọn phương án cân bằng tốc độ (Sampled Shapley + low path count) và giám sát đúng (prediction drift) để đáp ứng đầy đủ yêu cầu production trên Vertex AI!
- A Load the data in BigQuery. Use BigQuery ML to train an Autoencoder model.
- B Load the data in BigQuery. Use BigQuery ML to train a matrix factorization model.
- C Read data to a Vertex AI Workbench notebook. Use TensorFlow to train a two-tower model.
- D Read data to a Vertex AI Workbench notebook. Use TensorFlow to train a matrix factorization model.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi này thuộc lĩnh vực Machine Learning trên Google Cloud Platform (GCP), cụ thể là xây dựng mô hình recommendation system (hệ thống gợi ý) cho một startup game. Dữ liệu đầu vào là vài terabytes dữ liệu có cấu trúc lưu trữ trong Cloud Storage (GCS), bao gồm:
- Gameplay time data: Thời gian chơi game của người dùng.
- User metadata: Thông tin metadata về người dùng (như sở thích, lịch sử).
- Game metadata: Thông tin metadata về game (như thể loại, độ khó).
Mục tiêu: Xây dựng mô hình gợi ý game mới cho người dùng, với yêu cầu ít code nhất (least amount of coding) – nghĩa là ưu tiên các giải pháp no-code/low-code như SQL-based ML thay vì viết code Python/TensorFlow phức tạp.
Đây là bài toán collaborative filtering điển hình trong recommendation systems, nơi mô hình học từ tương tác user-item (người dùng - game) để dự đoán sở thích. Dữ liệu lớn (TB scale) phù hợp với BigQuery để xử lý hiệu quả mà không cần ETL phức tạp.
📘 Dẫn nguồn:
- BigQuery ML documentation (cập nhật 2024-2026): cloud.google.com/bigquery/docs/bigquery-ml-introduction.
- Vertex AI & Recommendation Models: cloud.google.com/vertex-ai/docs/generative-ai/recommendations.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Load the data in BigQuery. Use BigQuery ML to train a matrix factorization model.
Lý do 🛠️:
- Matrix factorization là thuật toán chuẩn cho recommendation systems (phân tích ma trận user-item để dự đoán rating/gameplay time), hỗ trợ trực tiếp collaborative filtering – phù hợp hoàn hảo với dữ liệu user-gameplay-game.
- BigQuery ML cho phép train mô hình chỉ bằng SQL query đơn giản (CREATE MODEL), không cần code Python, xử lý TB data scale dễ dàng sau khi load từ GCS vào BigQuery.
- Đây là giải pháp ít code nhất, chỉ cần 1-2 lệnh SQL để load data và train (ví dụ:
CREATE OR REPLACE MODEL USING MATRIX_FACTORIZATION). - Hiệu suất cao trên BigQuery, tự động scale, tích hợp ML.INFERENCE() để predict ngay.
- Phù hợp phiên bản mới nhất (2026): BigQuery ML vẫn là lựa chọn top cho low-code recsys.
❌ Giải thích tất cả các phương án (đúng/sai)
-
Phương án 1 [SAI]: Load the data in BigQuery. Use BigQuery ML to train an Autoencoder model.
Lý do sai ❌: Autoencoder dùng cho dimensionality reduction hoặc anomaly detection (như nén dữ liệu), KHÔNG phù hợp cho recommendation (không học user-item interactions trực tiếp). BigQuery ML hỗ trợ autoencoder nhưng không phải best-fit cho bài toán gợi ý game. Vẫn low-code nhưng thuật toán sai → mô hình kém hiệu quả. -
Phương án 2 [ĐÚNG]: Load the data in BigQuery. Use BigQuery ML to train a matrix factorization model.
Lý do đúng ✅: Như đã giải thích ở trên – low-code tối ưu, thuật toán chuẩn, scale TB data, chỉ SQL. Best practice cho recsys trên GCP. -
Phương án 3 [SAI]: Read data to a Vertex AI Workbench notebook. Use TensorFlow to train a two-tower model.
Lý do sai ❌: Two-tower model (Retrieval + Ranking) mạnh cho recsys lớn nhưng yêu cầu code TensorFlow phức tạp (hàng trăm dòng notebook), ETL data từ GCS sang Workbench, train/distribute – vi phạm "least coding". Vertex AI Workbench tốt cho custom ML nhưng overkill và tốn công code hơn BigQuery ML. -
Phương án 4 [SAI]: Read data to a Vertex AI Workbench notebook. Use TensorFlow to train a matrix factorization model.
Lý do sai ❌: Matrix factorization đúng thuật toán nhưng implement bằng TensorFlow trong notebook đòi hỏi code nhiều (define model, loss function, train loop), ETL data thủ công – KHÔNG phải least coding. BigQuery ML làm việc này chỉ bằng SQL, nhanh hơn và ít lỗi hơn.
🧠 Kết luận: Chọn BigQuery ML matrix factorization là optimal cho low-code, scale lớn, phù hợp startup. Nếu cần advanced hơn (embedding), mới chuyển Vertex AI! 🚀