Ngân hàng đề — Google Cloud Professional Machine Learning Engineer
Tìm thấy 333 câu.
- A Write a query that preprocesses the data by using BigQuery and creates a new table. Create a Vertex AI managed dataset with the new table as the data source.
- B Use Dataflow to preprocess the data. Write the output in TFRecord format to a Cloud Storage bucket.
- C Write a query that preprocesses the data by using BigQuery. Export the query results as CSV files, and use those files to create a Vertex AI managed dataset.
- D Use a Vertex AI Workbench notebook instance to preprocess the data by using the pandas library. Export the data as CSV files, and use those files to create a Vertex AI managed dataset.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc chuẩn bị dữ liệu đơn giản và hiệu quả nhất để huấn luyện mô hình AutoML dự đoán giá nhà, sử dụng bộ dữ liệu công khai nhỏ lưu trữ trong BigQuery.
✅ Mục tiêu chính: Sử dụng Vertex AI (dịch vụ ML của Google Cloud) để tạo managed dataset từ dữ liệu BigQuery, với cách tiếp cận đơn giản nhất (simplest) và hiệu quả nhất (most efficient).
🛠️ Bối cảnh: AutoML Tables (nay là Vertex AI AutoML) hỗ trợ trực tiếp dữ liệu từ BigQuery mà không cần export, giúp tiết kiệm thời gian, chi phí và tránh các bước trung gian phức tạp. Kiến thức dựa trên tài liệu Vertex AI cập nhật đến 2026 (phiên bản Vertex AI v1.50+ và BigQuery ML integrations).
📘 Tài liệu tham khảo:
- Vertex AI Documentation: Create a dataset from BigQuery
- Google Cloud ML Engineer Exam Guide (2024-2026)
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Write a query that preprocesses the data by using BigQuery and creates a new table. Create a Vertex AI managed dataset with the new table as the data source.
Lý do:
🟢 Đây là cách đơn giản và hiệu quả nhất vì:
- BigQuery hỗ trợ query SQL trực tiếp để preprocess (làm sạch, transform dữ liệu) và tạo bảng mới ngay lập tức, không cần tool ngoài.
- Vertex AI managed dataset có thể import trực tiếp từ BigQuery table, tự động schema inference, không export file trung gian → Tiết kiệm thời gian, chi phí lưu trữ và xử lý.
- Phù hợp dataset nhỏ, công khai → Không cần scaling lớn. Đây là best practice cho AutoML Tabular theo docs Google Cloud 2026.
🔍 Giải thích tất cả các phương án (đúng/sai)
-
✅ [ĐÚNG] Write a query that preprocesses the data by using BigQuery and creates a new table. Create a Vertex AI managed dataset with the new table as the data source.
🟢 Đúng vì: Như phân tích trên, kết hợp BigQuery query (preprocess inline) + Vertex AI direct import → 0 bước export, nhanh nhất cho dataset nhỏ. Hỗ trợ full AutoML features (schema auto-detect, versioning). -
❌ [SAI] Use Dataflow to preprocess the data. Write the output in TFRecord format to a Cloud Storage bucket.
❌ Sai vì: Dataflow (Apache Beam) quá phức tạp cho dataset nhỏ, yêu cầu code pipeline toàn diện. TFRecord phù hợp deep learning lớn-scale (TensorFlow), không cần cho AutoML Tabular → Không "simplest", tốn chi phí compute/storage. Vertex AI không ưu tiên format này cho tabular data. -
❌ [SAI] Write a query that preprocesses the data by using BigQuery. Export the query results as CSV files, and use those files to create a Vertex AI managed dataset.
❌ Sai vì: Bước export CSV trung gian không cần thiết → Tăng thời gian (download/upload), chi phí storage (Cloud Storage), và rủi ro lỗi format. Vertex AI hỗ trợ direct BigQuery → Export làm chậm hơn, không "most efficient". -
❌ [SAI] Use a Vertex AI Workbench notebook instance to preprocess the data by using the pandas library. Export the data as CSV files, and use those files to create a Vertex AI managed dataset.
❌ Sai vì: Workbench + pandas yêu cầu setup instance, code Python thủ công → Phức tạp, tốn resource cho dataset nhỏ. Export CSV lại thừa → Không tận dụng native BigQuery SQL, vi phạm "simplest approach". Phù hợp custom ML, không AutoML.
🧠 Kết luận: Phương án đúng tận dụng native integrations Google Cloud (BigQuery ↔ Vertex AI), giảm latency và cost tối đa! 🚀
- A Trigger a Cloud Build workflow to run tests, build custom Docker images, push the images to Artifact Registry, and launch the pipeline in Vertex AI Pipelines.
- B Trigger GitHub Actions to run the tests, launch a job on Cloud Run to build custom Docker images, push the images to Artifact Registry, and launch the pipeline in Vertex AI Pipelines.
- C Trigger GitHub Actions to run the tests, build custom Docker images, push the images to Artifact Registry, and launch the pipeline in Vertex AI Pipelines.
- D Trigger GitHub Actions to run the tests, launch a Cloud Build workflow to build custom Docker images, push the images to Artifact Registry, and launch the pipeline in Vertex AI Pipelines.
Xem giải thích
🧩 Giải thích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc tối ưu hóa quy trình CI/CD cho Vertex AI ML pipeline trên Google Cloud Platform (GCP). Cụ thể:
- Bối cảnh: Bạn đã xây dựng một pipeline ML trên Vertex AI bao gồm các bước preprocessing và training, mỗi nhóm bước chạy trên custom Docker image riêng biệt (tức là cần build và push images tùy chỉnh).
- Môi trường: Tổ chức sử dụng GitHub và GitHub Actions làm CI/CD để chạy unit tests và integration tests.
- Yêu cầu chính:
- Tự động hóa retraining model: Có thể kích hoạt thủ công (manual) hoặc tự động khi code mới merge vào main branch.
- Mục tiêu: Giảm thiểu số bước cần thiết để thiết lập workflow (minimize steps), đồng thời tối đa hóa tính linh hoạt (maximum flexibility).
- Thách thức: Cần kết hợp GitHub Actions (nhẹ, nhanh cho tests) với các dịch vụ GCP mạnh mẽ như Cloud Build (build/push Docker scalable), Artifact Registry (lưu trữ images), và Vertex AI Pipelines (chạy pipeline).
Câu hỏi kiểm tra kiến thức về tích hợp GitHub Actions với GCP services (cập nhật đến 2026: Vertex AI Pipelines hỗ trợ triggers linh hoạt, Cloud Build v2 với GitHub App integration native). 📘 Tài liệu tham khảo:
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Trigger GitHub Actions to run the tests, launch a Cloud Build workflow to build custom Docker images, push the images to Artifact Registry, and launch the pipeline in Vertex AI Pipelines.
Lý do 🛠️:
- Phân công nhiệm vụ tối ưu: GitHub Actions chạy tests nhanh/linh hoạt (native của org), sau đó trigger Cloud Build để xử lý heavy tasks như build custom Docker images (scalable, parallel builds lên đến 2026 với Cloud Build Remote). Cloud Build tự động push vào Artifact Registry và launch Vertex AI Pipeline qua
gcloudhoặc API. - Minimize steps: Chỉ cần 1 GitHub workflow YAML đơn giản (runs-on Ubuntu, dùng
google-github-actions/authvàgcloud builds submit/triggers run). Tích hợp native qua GitHub App, không cần setup phức tạp. - Max flexibility: Hỗ trợ manual dispatch (workflow_dispatch), push/merge main trigger, và parameterized inputs (e.g., branch name). Cloud Build cho phép custom
cloudbuild.yamllinh hoạt cho multi-step pipelines. - Đây là best practice GCP 2026: Hybrid CI/CD giảm chi phí, tận dụng strengths (GH tests nhẹ → CB builds nặng). ✅
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ Phương án SAI 1: Trigger a Cloud Build workflow to run tests, build custom Docker images, push the images to Artifact Registry, and launch the pipeline in Vertex AI Pipelines.
Giải thích: Cloud Build không phù hợp chạy unit/integration tests phức tạp từ GitHub repo (tests cần môi trường dev nhanh, dependencies Node/Python linh hoạt). Cloud Build tập trung build/deploy, chạy tests kém hiệu quả, tốn kém hơn (billing theo build minutes). Không tận dụng GitHub Actions sẵn có, vi phạm "minimize steps" vì phải migrate toàn bộ tests sang Cloud Build triggers → thiếu flexibility. -
❌ Phương án SAI 2: Trigger GitHub Actions to run the tests, launch a job on Cloud Run to build custom Docker images, push the images to Artifact Registry, and launch the pipeline in Vertex AI Pipelines.
Giải thích: Cloud Run KHÔNG dùng để build Docker images (nó là serverless runtime cho chạy containers đã build sẵn, không hỗ trợ Docker build context). Build trên Cloud Run sẽ fail (no Docker daemon), tốn kém và không scalable cho ML custom images lớn. Phải dùng Cloud Build hoặc Buildpacks → vi phạm kiến trúc GCP chuẩn. -
❌ Phương án SAI 3: Trigger GitHub Actions to run the tests, build custom Docker images, push the images to Artifact Registry, and launch the pipeline in Vertex AI Pipelines.
Giải thích: GitHub Actions có thể build Docker (quadocker/build-push-action), nhưng giới hạn nghiêm ngặt (runner time 6h max, CPU/RAM thấp, không parallel tốt cho multi-image ML pipelines). Với custom images phức tạp (preprocessing/training), dễ timeout/OOM, kém scalable so Cloud Build (unlimited concurrency 2026). Không "maximum flexibility" cho large-scale retraining → không minimize steps dài hạn. -
✅ Phương án ĐÚNG: Trigger GitHub Actions to run the tests, launch a Cloud Build workflow to build custom Docker images, push the images to Artifact Registry, and launch the pipeline in Vertex AI Pipelines.
Giải thích: Như đã phân tích ở trên, đây là cách tối ưu nhất, hybrid model tận dụng GitHub (tests/manual trigger) + Cloud Build (build/push/launch). Dễ setup chỉ với 1 GitHub Actions YAML + 1 cloudbuild.yaml. Hỗ trợ full automation/flexibility theo yêu cầu. 🏆
Kết luận 🚀: Phương án đúng đảm bảo quy trình end-to-end tự động, scalable trên GCP, phù hợp certification Google Cloud Professional Machine Learning Engineer. Nếu implement, dùng GitHub Actions marketplace actions như google-github-actions/cloudbuild!
- A Use the TRANSFORM clause with the ML.ONE_HOT_ENCODER function on the categorical features at model creation and select the categorical and non-categorical features.
- B Use the ML.ONE_HOT_ENCODER function on the categorical features and select the encoded categorical features and non-categorical features as inputs to create your model.
- C Use the CREATE MODEL statement and select the categorical and non-categorical features.
- D Use the ML.MULTI_HOT_ENCODER function on the categorical features, and select the encoded categorical features and non-categorical features as inputs to create your model.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc xây dựng mô hình Machine Learning (ML) trong BigQuery ML (một dịch vụ của Google Cloud) để dự đoán hành vi mua hàng của khách hàng dựa trên dataset chứa giao dịch. Các bước chính bao gồm:
- Phát triển mô hình trong BigQuery ML.
- Xuất mô hình ra Cloud Storage để phục vụ dự đoán online (online prediction).
- Dataset có các đặc trưng phân loại (categorical features) như product category và payment method, cùng các đặc trưng số (non-categorical).
- Mục tiêu: Triển khai mô hình nhanh nhất có thể (deploy as quickly as possible).
🛠️ Vấn đề cốt lõi: BigQuery ML hỗ trợ tự động xử lý encoding cho các categorical features (như one-hot encoding) khi tạo mô hình, giúp đơn giản hóa quy trình mà không cần can thiệp thủ công. Điều này đặc biệt phù hợp với các mô hình như logistic regression, boosted trees hoặc AutoML Tables, giúp tiết kiệm thời gian deploy. Kiến thức dựa trên tài liệu BigQuery ML cập nhật đến năm 2026 (phiên bản GA mới nhất từ Google Cloud, hỗ trợ tự động encoding trong CREATE MODEL cho hầu hết các thuật toán).
📘 Tài liệu tham khảo:
- BigQuery ML Syntax - CREATE MODEL (Google Cloud Docs, cập nhật 2025).
- Handling Categorical Data in BigQuery ML (hướng dẫn tự động one-hot encoding).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng:
Use the CREATE MODEL statement and select the categorical and non-categorical features.
Lý do:
BigQuery ML tự động xử lý categorical features bằng cách áp dụng one-hot encoding nội bộ khi bạn chỉ định trực tiếp chúng trong câu lệnh CREATE MODEL (không cần TRANSFORM clause thủ công). Điều này giúp quy trình tạo mô hình và export sang Cloud Storage diễn ra nhanh nhất, vì không tốn thêm bước encoding riêng lẻ. Sau khi tạo model, bạn có thể export dễ dàng bằng ML.EXPORT_MODEL và sử dụng cho online prediction (ví dụ: tích hợp với Vertex AI hoặc custom endpoint). Đây là cách đơn giản, hiệu quả nhất theo best practices của Google Cloud đến 2026.
📋 Giải thích tất cả các phương án (đúng và sai)
-
Use the TRANSFORM clause with the ML.ONE_HOT_ENCODER function on the categorical features at model creation and select the categorical and non-categorical features.
❌ Sai: Việc sử dụng TRANSFORM clause với ML.ONE_HOT_ENCODER là thủ công và dư thừa, vì BigQuery ML đã tự động encode. Hơn nữa, chọn cả categorical gốc + non-categorical sẽ gây duplicate features (lặp dữ liệu), dẫn đến lỗi model training hoặc kết quả sai lệch. Cách này làm chậm deploy vì phức tạp hơn, không phải "nhanh nhất". -
Use the ML.ONE_HOT_ENCODER function on the categorical features and select the encoded categorical features and non-categorical features as inputs to create your model.
❌ Sai: ML.ONE_HOT_ENCODER chỉ dùng trong TRANSFORM cho custom logic (ví dụ: high-cardinality features), nhưng ở đây không cần thiết. Chọn encoded features + non-categorical vẫn yêu cầu bước tiền xử lý riêng, tăng thời gian so với tự động encoding. Không phù hợp cho deploy nhanh, và có thể gây vấn đề khi export model nếu encoding không khớp. -
Use the CREATE MODEL statement and select the categorical and non-categorical features.
✅ Đúng: Như đã giải thích ở trên, đây là cách tối ưu và nhanh nhất. BigQuery ML tự động one-hot encode categorical features (giới hạn cardinality < 100 unique values theo docs 2026), hỗ trợ trực tiếp cho online prediction sau export. Ví dụ SQL:CREATE MODEL my_model OPTIONS(...) AS SELECT categorical_col, numeric_col FROM table;. -
Use the ML.MULTI_HOT_ENCODER function on the categorical features, and select the encoded categorical features and non-categorical features as inputs to create your model.
❌ Sai: ML.MULTI_HOT_ENCODER dùng cho multi-label hoặc array-based categorical (như tags), không phù hợp với categorical đơn giản như product category/payment method. Áp dụng thủ công làm phức tạp hóa, chậm deploy, và có thể không tương thích với online prediction nếu encoding không chuẩn.
🧠 Lời khuyên từ Google Cloud Professional ML Engineer: Luôn ưu tiên tự động hóa của BigQuery ML để scale nhanh. Nếu categorical có cardinality cao (>1000), mới dùng custom encoder! 🚀
- A Use Vertex AI Pipelines with the Kubeflow Pipelines SDK to create a pipeline that reads the images from Cloud Storage and trains the model.
- B Use Vertex AI Pipelines with TensorFlow Extended (TFX) to create a pipeline that reads the images from Cloud Storage and trains the model.
- C Import the labeled images as a managed dataset in Vertex AI and use AutoML to train the model.
- D Convert the image dataset to a tabular format using Dataflow Load the data into BigQuery and use BigQuery ML to train the model.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi yêu cầu phát triển một mô hình phân loại hình ảnh (image classification model) sử dụng tập dữ liệu lớn chứa hình ảnh đã được gắn nhãn (labeled images) lưu trữ trong Cloud Storage bucket trên Google Cloud Platform (GCP).
📌 Yêu cầu chính: Tìm cách xử lý dữ liệu từ Cloud Storage và huấn luyện mô hình một cách hiệu quả, phù hợp với quy mô lớn. Đây là tình huống điển hình trong Vertex AI, nơi cần chọn phương pháp đơn giản, managed và scalable nhất cho nhiệm vụ vision (phân loại hình ảnh).
🛠️ Bối cảnh: Vertex AI cung cấp các công cụ end-to-end cho ML, từ dữ liệu đến deployment. Với dữ liệu hình ảnh đã gắn nhãn sẵn, ưu tiên AutoML để tránh custom code phức tạp, đặc biệt với dataset lớn (không cần viết pipeline thủ công).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Import the labeled images as a managed dataset in Vertex AI and use AutoML to train the model.
Lý do:
- Vertex AI hỗ trợ import trực tiếp dữ liệu hình ảnh từ Cloud Storage vào managed dataset (Dataset resource) dành riêng cho Image data type, tự động xử lý metadata nhãn (labels).
- AutoML Vision (nay tích hợp trong Vertex AI) là giải pháp no-code/low-code lý tưởng cho image classification: tự động train high-quality model với dataset lớn mà không cần custom pipeline hay code.
- ✅ Ưu điểm: Scalable (hàng triệu images), managed (Google lo infra), nhanh chóng (train chỉ vài giờ), và đạt accuracy cao (state-of-the-art). Phù hợp nhất cho "develop a model" mà không chỉ định custom requirements.
- Theo docs GCP 2026: Vertex AI AutoML Image vẫn là recommended first choice cho vision tasks (cập nhật với Imagen models và multimodal support).
📋 Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên tính phù hợp, simplicity và best practice của Vertex AI (phiên bản mới nhất 2026).
-
❌ SAI: Use Vertex AI Pipelines with the Kubeflow Pipelines SDK to create a pipeline that reads the images from Cloud Storage and trains the model.
Giải thích: Kubeflow Pipelines (KFP SDK) dùng để xây custom pipeline orchestration cho workflows phức tạp (ví dụ: data prep + train + serve). Tuy có thể đọc Cloud Storage và train image model (qua custom components), nhưng quá phức tạp và overkill cho task đơn giản như image classification với labeled data. Không phải "should you do" (best practice), vì đòi hỏi code nhiều, debug lâu, không managed như AutoML. -
❌ SAI: Use Vertex AI Pipelines with TensorFlow Extended (TFX) to create a pipeline that reads the images from Cloud Storage and trains the model.
Giải thích: TFX là framework end-to-end cho production ML pipelines (data validation, transformation, tuning), hỗ trợ image data. Có thể build pipeline đọc Cloud Storage và train TF model, nhưng không phù hợp cho beginner/task đơn giản – yêu cầu kiến thức sâu về TFX components (ExampleGen, Trainer). Vertex AI ưu tiên AutoML trước TFX cho non-experts; TFX dành cho custom scalable production. -
✅ ĐÚNG: Import the labeled images as a managed dataset in Vertex AI and use AutoML to train the model.
Giải thích: Như đã nêu ở phần đáp án đúng – managed dataset hỗ trợ CSV/JSON manifest từ Cloud Storage (gồm image URI + labels). AutoML tự động train, evaluate và deploy model. Best practice cho large labeled image datasets, scalable đến petabytes. -
❌ SAI: Convert the image dataset to a tabular format using Dataflow Load the data into BigQuery and use BigQuery ML to train the model.
Giải thích: BigQuery ML (BQML) chỉ hỗ trợ tabular/structured data (SQL-based models như logistic regression, boosting), KHÔNG hỗ trợ native image classification (images là unstructured binary). Convert image sang tabular (feature extraction?) qua Dataflow là impractical và inefficient cho large dataset (mất metadata, accuracy thấp). BQML 2026 vẫn chưa có vision support đầy đủ; dùng cho tabular ML, không phải images.
📘 Tài liệu tham khảo (cập nhật 2026)
- Vertex AI AutoML Vision Docs ✅ (Managed datasets & import từ GCS).
- Vertex AI Datasets Guide 🛠️ (Image import specifics).
- BigQuery ML Limitations ❌ (No image support).
- Kubeflow vs AutoML Comparison (Best practices cho pipelines).
- GCP Professional ML Engineer Exam Guide (2026): Nhấn mạnh AutoML cho vision tasks đầu tiên.
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần ví dụ code, hỏi thêm nhé!
- A Increase the score threshold
- B Decrease the score threshold.
- C Add more positive examples to the training set
- D Add more negative examples to the training set
- E Reduce the maximum number of node hours for training
Xem giải thích
🔍 Phân tích chi tiết câu hỏi trắc nghiệm
🧩 Nội dung câu hỏi:
Câu hỏi xoay quanh việc phát triển một mô hình học máy để phát hiện giao dịch thẻ tín dụng gian lận (fraudulent credit card transactions). Ưu tiên hàng đầu là tối ưu hóa việc phát hiện, vì bỏ sót chỉ một giao dịch gian lận cũng có thể gây thiệt hại nghiêm trọng cho chủ thẻ. Bạn đã sử dụng AutoML (cụ thể là Vertex AI AutoML Tables trong Google Cloud) để huấn luyện mô hình ban đầu dựa trên dữ liệu thông tin hồ sơ người dùng (users' profile information) và dữ liệu giao dịch thẻ tín dụng (credit card transaction data). Sau khi huấn luyện, mô hình thất bại trong việc phát hiện nhiều giao dịch gian lận (failing to detect many fraudulent transactions). Câu hỏi yêu cầu điều chỉnh các tham số huấn luyện trong AutoML để cải thiện hiệu suất mô hình, và cần chọn hai lựa chọn đúng.
📘 Bối cảnh kỹ thuật:
- Đây là bài toán phân loại nhị phân (binary classification) với lớp positive (gian lận - fraud, thường hiếm gặp) và lớp negative (bình thường). Mô hình hiện tại có recall thấp (bỏ sót nhiều positive), cần ưu tiên recall cao hơn precision để tránh false negatives.
- AutoML cho phép điều chỉnh score threshold (ngưỡng điểm số để phân loại) và cân bằng dữ liệu huấn luyện (thêm ví dụ positive/negative). Kiến thức dựa trên Vertex AI phiên bản mới nhất 2026, nơi AutoML hỗ trợ tự động hóa nhưng vẫn cho phép tinh chỉnh threshold và dataset imbalance handling (xem tài liệu: Vertex AI AutoML Tabular).
✅ Đáp án đúng (chọn hai):
- Decrease the score threshold: Giảm ngưỡng điểm số giúp mô hình dễ dàng hơn trong việc dự đoán "gian lận", tăng recall bằng cách chấp nhận nhiều false positives hơn, phù hợp ưu tiên phát hiện.
- Add more positive examples to the training set: Thêm ví dụ gian lận (positive) giúp cân bằng dataset, cải thiện khả năng học lớp hiếm gặp, giảm bias về lớp negative đông đảo.
🛠️ Giải thích chi tiết tất cả các phương án
-
❌ Increase the score threshold
Sai: Tăng ngưỡng điểm số làm mô hình kén chọn hơn (conservative), chỉ dự đoán "gian lận" khi độ tin cậy rất cao. Điều này giảm recall, dẫn đến bỏ sót nhiều giao dịch gian lận hơn – trái ngược yêu cầu ưu tiên phát hiện. Trong Vertex AI, threshold mặc định ~0.5, tăng lên sẽ làm precision cao nhưng recall thấp. -
✅ Decrease the score threshold
Đúng: Giảm ngưỡng (ví dụ từ 0.5 xuống 0.3) làm mô hình linh hoạt hơn, dễ dự đoán positive (gian lận) ngay cả với điểm số thấp. Tăng recall mạnh mẽ, phù hợp tình huống "không bỏ sót một trường hợp nào". Vertex AI hỗ trợ điều chỉnh này qua API hoặc UI sau training (tài liệu: Score Thresholds). -
✅ Add more positive examples to the training set
Đúng: Giao dịch gian lận thường imbalanced (positive hiếm, <1-5%). Thêm positive examples (qua oversampling, synthetic data như SMOTE, hoặc thu thập thêm) giúp mô hình học tốt lớp hiếm, cải thiện recall và F1-score. AutoML Vertex AI tự động xử lý imbalance nhưng thêm manual positive là best practice (tài liệu: Handle Imbalanced Data). -
❌ Add more negative examples to the training set
Sai: Thêm negative (giao dịch bình thường, đã đông) làm imbalance tệ hơn, mô hình càng bias về negative, giảm recall cho fraud. Không khuyến khích trong AutoML vì làm dataset phình to vô ích, tốn tài nguyên. -
❌ Reduce the maximum number of node hours for training
Sai: Giảm node hours (thời gian huấn luyện) làm training kém sâu, mô hình chưa hội tụ, hiệu suất tệ hơn. Nên tăng node hours (lên 1000+ nếu cần) để AutoML khám phá không gian mô hình tốt hơn (tài liệu: Training Budgets).
Tài liệu tham khảo chính (cập nhật 2026):
- 📘 Google Cloud Vertex AI AutoML Documentation
- 📘 Best Practices for Fraud Detection
- 🎯 Kết quả: Hai lựa chọn đúng giúp tối ưu recall mà không cần custom code, phù hợp AutoML! 🚀
- A Deploy an online Vertex AI prediction endpoint. Set the max replica count to 1
- B Deploy an online Vertex AI prediction endpoint. Set the max replica count to 100
- C Deploy an online Vertex AI prediction endpoint with one GPU per replica. Set the max replica count to 1
- D Deploy an online Vertex AI prediction endpoint with one GPU per replica. Set the max replica count to 100
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi yêu cầu triển khai một mô hình phân loại scikit-learn (thư viện machine learning phổ biến dựa trên CPU) vào môi trường production. Các yêu cầu chính bao gồm:
- Phục vụ liên tục 24/7 (endpoint phải luôn sẵn sàng).
- Xử lý hàng triệu requests/giây (millions of requests per second) trong khoảng thời gian cao điểm từ 8 giờ sáng đến 7 giờ tối.
- Tối ưu hóa chi phí (minimize the cost of deployment) – đây là yếu tố quyết định, vì vậy cần tránh các tài nguyên thừa như GPU không cần thiết.
🛠️ Vertex AI Prediction Endpoint là dịch vụ của Google Cloud dùng để deploy mô hình ML vào production với khả năng auto-scaling dựa trên traffic. Mô hình scikit-learn là lightweight, chạy tốt trên CPU (không cần GPU), nên ưu tiên cấu hình CPU để tiết kiệm chi phí. Auto-scaling replicas (số lượng instance song song) giúp xử lý tải cao mà không cần over-provisioning.
📘 Nguồn tham khảo:
- Vertex AI Prediction Documentation (cập nhật 2024-2026)
- Vertex AI Pricing (phiên bản mới nhất) – CPU rẻ hơn GPU gấp nhiều lần.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Deploy an online Vertex AI prediction endpoint. Set the max replica count to 100
Lý do:
- Vertex AI endpoint hỗ trợ online prediction (real-time, 24/7) với auto-scaling dựa trên traffic, phù hợp cho hàng triệu RPS cao điểm.
- Max replica count = 100 cho phép scale lên tối đa 100 instances CPU để xử lý tải lớn, nhưng chỉ scale khi cần (min replica thường là 0-1 để tiết kiệm).
- Không dùng GPU (mặc định CPU cho scikit-learn), giúp minimize cost vì GPU đắt hơn (khoảng 3-10x CPU theo pricing 2026).
- Nếu max replica =1 thì không đủ tải; thêm GPU thì lãng phí chi phí cho model CPU-based.
📋 Giải thích chi tiết tất cả các phương án
-
Deploy an online Vertex AI prediction endpoint. Set the max replica count to 1
❌ Sai: Max replica chỉ 1 không đủ xử lý millions RPS (một instance CPU chỉ handle ~100-1000 RPS tùy model). Sẽ gây bottleneck, latency cao, không đáp ứng 24/7 high-traffic. Dù rẻ nhưng vi phạm yêu cầu scale. -
Deploy an online Vertex AI prediction endpoint. Set the max replica count to 100
✅ Đúng: Cấu hình lý tưởng! Auto-scale lên 100 CPU replicas xử lý tải cao điểm hiệu quả, idle thì scale down về min (tiết kiệm ~80-90% chi phí so với always-on). Phù hợp scikit-learn (CPU-only), 24/7 ready, min cost tối ưu. -
Deploy an online Vertex AI prediction endpoint with one GPU per replica. Set the max replica count to 1
❌ Sai: GPU/replica thừa thãi cho scikit-learn (không tận dụng GPU acceleration), chi phí cao gấp bội (GPU như A100/T4 ~$1-3/giờ vs CPU ~$0.1/giờ). Max=1 vẫn không scale, latency cao – vi phạm cả scale và cost. -
Deploy an online Vertex AI prediction endpoint with one GPU per replica. Set the max replica count to 100
❌ Sai: Scale tốt (100 replicas) nhưng GPU không cần thiết, đẩy chi phí lên cao (100 GPU replicas có thể tốn hàng nghìn USD/ngày). Scikit-learn chạy chậm hơn trên GPU do không optimize, không minimize cost – lãng phí lớn!
🧩 Tóm tắt insight: Chọn CPU + high max replicas là chìa khóa cho workload CPU-intensive như scikit-learn với traffic bursty, theo best practices Vertex AI 2026 (enable autoscaling, dedicated resources chỉ khi cần).
- A Configure a v3-8 TPU VM. SSH into the VM to train and debug the model.
- B Configure a v3-8 TPU node. Use Cloud Shell to SSH into the Host VM to train and debug the model.
- C Configure a n1 -standard-4 VM with 4 NVIDIA P100 GPUs. SSH into the VM and use ParameterServerStraregv to train the model.
- D Configure a n1-standard-4 VM with 4 NVIDIA P100 GPUs. SSH into the VM and use MultiWorkerMirroredStrategy to train the model.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc thiết lập môi trường huấn luyện mô hình TensorFlow cho đội ngũ nghiên cứu phát triển thuật toán phân tích tài chính tiên tiến. Mục tiêu chính là:
- Duy trì dễ dàng debug (gỡ lỗi) mô hình phức tạp.
- Giảm thời gian huấn luyện mô hình.
Người dùng làm việc với TensorFlow, cần môi trường hỗ trợ huấn luyện phân tán (distributed training) trên phần cứng mạnh như TPU hoặc GPU, nhưng phải dễ truy cập qua SSH để debug trực tiếp. Đây là tình huống điển hình trên Google Cloud Platform (GCP), sử dụng các instance Compute Engine với TPU hoặc NVIDIA GPUs. Kiến thức dựa trên TensorFlow 2.x (cập nhật đến 2026, với các strategy phân tán được tối ưu hóa cho multi-GPU/TPU) và GCP docs mới nhất (TPU v3/v4, GPU P100/A100/H100).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Configure a n1-standard-4 VM with 4 NVIDIA P100 GPUs. SSH into the VM and use MultiWorkerMirroredStrategy to train the model.
Lý do:
- n1-standard-4 VM với 4 NVIDIA P100 GPUs 🛠️: Đây là instance Compute Engine hỗ trợ multi-GPU (P100 là GPU mạnh cho TensorFlow), cho phép huấn luyện nhanh hơn nhờ song song hóa. Bạn có thể SSH trực tiếp vào VM để debug dễ dàng (chạy Jupyter, TensorBoard, pdb), không bị hạn chế như TPU.
- MultiWorkerMirroredStrategy 📈: Strategy này dành cho multi-GPU trên single VM (mirrored synchronous training), tự động phân phối dữ liệu và gradient, giảm thời gian training đáng kể (scale linear với số GPU). Nó dễ debug vì code chạy trên cùng một process space, hỗ trợ TensorFlow Debugger đầy đủ (cập nhật TF 2.15+ đến 2026). Không cần parameter server phức tạp.
Kết hợp hoàn hảo: Debug dễ + Training nhanh!
Nguồn tham khảo:
- TensorFlow Distributed Training Guide (MirroredStrategy cho multi-GPU).
- GCP Compute Engine GPU docs (n1-standard với P100/A100).
- GCP ML Best Practices (cập nhật 2025-2026).
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên khả năng debug dễ dàng và giảm training time.
-
❌ [SAI] Configure a v3-8 TPU VM. SSH into the VM to train and debug the model.
Giải thích sai: TPU v3-8 VM (Cloud TPU Pod slice) rất mạnh cho training nhanh (hàng trăm TFLOPS), nhưng không hỗ trợ SSH trực tiếp vào TPU worker để debug. Bạn chỉ SSH vào control plane VM, TPU core chạy isolated (JAX/TPU code), khó gắn debugger như pdb/TensorBoard. Phù hợp production nhưng không dễ debug như yêu cầu. (Cập nhật 2026: TPU v5e cải thiện nhưng vẫn hạn chế SSH debug). -
❌ [SAI] Configure a v3-8 TPU node. Use Cloud Shell to SSH into the Host VM to train and debug the model.
Giải thích sai: TPU v3-8 node là single-host TPU (không phải Pod), nhanh cho training nhưng Cloud Shell SSH chỉ vào Host VM (không phải TPU worker). Debug TensorFlow trên TPU yêu cầu TPU-specific tools (như XLA compiler), phức tạp và chậm (không trực tiếp như GPU VM). Không tối ưu cho debug team research. -
❌ [SAI] Configure a n1 -standard-4 VM with 4 NVIDIA P100 GPUs. SSH into the VM and use ParameterServerStraregv to train the model.
Giải thích sai: VM và 4 P100 GPUs ✅ tốt (SSH dễ, training nhanh), nhưng ParameterServerStrategy (lỗi chính tả "Straregv" → Strategy) là strategy cũ và phức tạp (parameter servers async), khó debug vì multi-process/multi-node, dễ lỗi coordination. Không khuyến nghị cho single VM multi-GPU (TF docs 2026 ưu tiên MirroredStrategy). Làm chậm debug và có thể tăng thời gian training do overhead. -
✅ [ĐÚNG] Configure a n1-standard-4 VM with 4 NVIDIA P100 GPUs. SSH into the VM and use MultiWorkerMirroredStrategy to train the model.
Giải thích đúng: Như phần trên, hoàn hảo cân bằng: GPU scale training nhanh (4x P100 ~ linear speedup), SSH VM debug mượt, MirroredStrategy đơn giản synchronous cho single/multi-worker. Lý tưởng cho nghiên cứu TensorFlow! 🚀
•Input dataset
•Max tree depth of the boosted tree regressor
•Optimizer learning rate
You need to compare the pipeline performance of the different parameter combinations measured in F1 score, time to train, and model complexity. You want your approach to be reproducible, and track all pipeline runs on the same platform. What should you do?
-
A
1. Use BigQueryML to create a boosted tree regressor, and use the hyperparameter tuning capability.
2. Configure the hyperparameter syntax to select different input datasets: max tree depths, and optimizer learning rates. Choose the grid search option. -
B
1. Create a Vertex AI pipeline with a custom model training job as part of the pipeline. Configure the pipeline’s parameters to include those you are investigating.
2. In the custom training step, use the Bayesian optimization method with F1 score as the target to maximize. -
C
1. Create a Vertex AI Workbench notebook for each of the different input datasets.
2. In each notebook, run different local training jobs with different combinations of the max tree depth and optimizer learning rate parameters.
3. After each notebook finishes, append the results to a BigQuery table. -
D
1. Create an experiment in Vertex AI Experiments.
2. Create a Vertex AI pipeline with a custom model training job as part of the pipeline. Configure the pipeline’s parameters to include those you are investigating.
3. Submit multiple runs to the same experiment, using different values for the parameters.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào việc xây dựng và đánh giá một pipeline ML với nhiều tham số đầu vào (input dataset, max tree depth của boosted tree regressor, optimizer learning rate). Mục tiêu là khảo sát các trade-off giữa các tổ hợp tham số khác nhau, đo lường qua F1 score (hiệu suất mô hình), thời gian huấn luyện (time to train), và độ phức tạp mô hình (model complexity). Yêu cầu chính là cách tiếp cận phải có tính tái lập (reproducible) và theo dõi tất cả các lần chạy pipeline trên cùng một nền tảng (same platform).
🛠️ Bối cảnh kỹ thuật: Đây là tình huống thực tế trong Google Cloud Vertex AI, nơi bạn cần một công cụ quản lý thí nghiệm (experiment tracking) để chạy nhiều biến thể của pipeline, so sánh metrics đa chiều, và đảm bảo tính nhất quán. Kiến thức cập nhật đến 2026: Vertex AI Experiments (trước đây là Vertex AI Experiments SDK) hỗ trợ đầy đủ tracking pipeline runs với parameters, metrics (bao gồm custom như F1, time, complexity), và reproducibility qua versioning.
📘 Tài liệu tham khảo:
- Vertex AI Experiments Documentation (cập nhật 2024-2026).
- Vertex AI Pipelines.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng là lựa chọn cuối cùng:
- Create an experiment in Vertex AI Experiments.
- Create a Vertex AI pipeline with a custom model training job as part of the pipeline. Configure the pipeline’s parameters to include those you are investigating.
- Submit multiple runs to the same experiment, using different values for the parameters.
Lý do chọn đáp án này 🏆:
- Vertex AI Experiments cho phép tạo một experiment duy nhất để track nhiều pipeline runs với các tổ hợp tham số khác nhau (dataset, tree depth, learning rate), đảm bảo reproducible qua versioning pipeline và parameters.
- Theo dõi đa metrics: Tự động log F1 score, time to train (qua pipeline metadata), model complexity (qua custom metrics), và visualize trade-offs trên cùng một dashboard.
- Same platform: Tất cả runs nằm trong một experiment trên Vertex AI, dễ so sánh, scale, và tái chạy.
- Phù hợp phiên bản mới nhất (2026): Hỗ trợ custom jobs trong pipelines + Experiments cho hyperparameter sweeps thủ công/grid.
📋 Phân tích tất cả các phương án (đúng/sai)
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm giải thích lý do bằng tiếng Việt.
-
❌ Phương án 1 (SAI):
- Use BigQueryML to create a boosted tree regressor, and use the hyperparameter tuning capability.
- Configure the hyperparameter syntax to select different input datasets: max tree depths, and optimizer learning rates. Choose the grid search option.
Giải thích sai: BigQuery ML (BQML) hỗ trợ boosted tree regressor và hyperparameter tuning (grid/random search từ 2023), nhưng không linh hoạt với input dataset đa dạng (chỉ query-based, khó thay đổi dataset động trong pipeline). Không track time to train/model complexity đầy đủ, thiếu pipeline reproducibility thực thụ, và không phải "same platform" cho custom metrics như F1 chi tiết. Không phù hợp cho boosted tree regressor phức tạp với custom params.
-
❌ Phương án 2 (SAI):
- Create a Vertex AI pipeline with a custom model training job as part of the pipeline. Configure the pipeline’s parameters to include those you are investigating.
- In the custom training step, use the Bayesian optimization method with F1 score as the target to maximize.
Giải thích sai: Vertex AI Pipelines hỗ trợ custom jobs và params, nhưng Bayesian optimization là auto hyperparameter tuning (tối ưu tự động F1), không phải manual investigation trade-offs (thời gian, complexity). Không có experiment tracking để so sánh tất cả runs trên same platform; chỉ một tuning job, thiếu reproducibility cho multiple combinations thủ công.
-
❌ Phương án 3 (SAI):
- Create a Vertex AI Workbench notebook for each of the different input datasets.
- In each notebook, run different local training jobs with different combinations of the max tree depth and optimizer learning rate parameters.
- After each notebook finishes, append the results to a BigQuery table.
Giải thích sai: Vertex AI Workbench (notebooks) cho phép training local, nhưng không reproducible (manual per dataset/notebook, dễ lỗi versioning). Không "same platform" thống nhất (nhiều notebooks + manual append BQ), khó track pipeline đầy đủ, và tốn công so sánh metrics đa chiều. Không scale cho nhiều combinations.
-
✅ Phương án 4 (ĐÚNG):
- Create an experiment in Vertex AI Experiments.
- Create a Vertex AI pipeline with a custom model training job as part of the pipeline. Configure the pipeline’s parameters to include those you are investigating.
- Submit multiple runs to the same experiment, using different values for the parameters.
Giải thích đúng: Hoàn hảo khớp yêu cầu! Experiments track tất cả runs của cùng pipeline với params khác nhau, log metrics (F1, time, complexity) tự động, visualize trade-offs (plots, comparisons). Reproducible 100% qua experiment artifacts, versioning. Cập nhật 2026: Tích hợp sâu với Pipelines cho grid-like manual sweeps.
🧠 Lời khuyên thực hành: Sử dụng Kubeflow Pipelines trong Vertex AI + Experiments SDK để implement nhanh. Ví dụ code: aiplatform.Experiment(run_name=...) để submit runs!
- A Update the model monitoring job to use a lower sampling rate.
- B Update the model monitoring job to use the more recent training data that was used to retrain the model.
- C Temporarily disable the alert. Enable the alert again after a sufficient amount of new production traffic has passed through the Vertex AI endpoint.
- D Temporarily disable the alert until the model can be retrained again on newer training data. Retrain the model again after a sufficient amount of new production traffic has passed through the Vertex AI endpoint.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh tình huống training-serving skew alert (cảnh báo lệch dữ liệu giữa huấn luyện và phục vụ) từ Vertex AI Model Monitoring job đang chạy trên môi trường production (sản xuất).
-
Training-serving skew là hiện tượng dữ liệu dùng để huấn luyện mô hình (training data) khác biệt đáng kể so với dữ liệu thực tế khi mô hình phục vụ (serving data từ production traffic). Vertex AI Model Monitoring tự động phát hiện điều này bằng cách so sánh phân phối thống kê giữa training dataset baseline (dữ liệu huấn luyện làm chuẩn) và serving data (dữ liệu đầu vào thực tế từ endpoint).
-
Bạn đã retrain mô hình với dữ liệu huấn luyện mới hơn (more recent training data), deploy lại lên Vertex AI endpoint, nhưng vẫn nhận alert. Lý do tiềm ẩn: Model Monitoring job vẫn sử dụng training dataset cũ làm baseline, nên khi so sánh với serving data mới (có thể đã thay đổi), skew vẫn được phát hiện dù mô hình đã được cập nhật.
Mục tiêu: Xử lý alert một cách hiệu quả và đúng quy trình, đảm bảo monitoring phản ánh đúng trạng thái mô hình mới nhất. Đây là kiến thức cốt lõi trong Vertex AI Model Monitoring (cập nhật đến 2026, theo tài liệu Google Cloud Vertex AI v1.5+).
📘 Tài liệu tham khảo:
✅ Đáp án đúng và lý do lựa chọn
Update the model monitoring job to use the more recent training data that was used to retrain the model.
🛠️ Lý do chi tiết:
- Khi retrain và deploy mô hình mới, training dataset baseline trong monitoring job cần được cập nhật tương ứng để làm chuẩn mới. Nếu không, job vẫn dùng baseline cũ → so sánh lệch lạc → alert sai (false positive).
- Quy trình chuẩn: Sử dụng lệnh
gcloud ai model-monitoring-jobs updatehoặc Console để upload/update training dataset URI mới nhất. Sau đó, job sẽ baseline trên data huấn luyện mới → skew alert sẽ biến mất nếu serving data phù hợp. - Đây là giải pháp ngay lập tức, chính xác, tránh downtime và đảm bảo monitoring liên tục (theo best practices Vertex AI 2026).
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] Update the model monitoring job to use a lower sampling rate.
🧩 Phân tích: Giảm sampling rate (tỷ lệ lấy mẫu serving data) chỉ làm giảm độ chính xác phát hiện skew, không giải quyết gốc rễ (baseline training data cũ). Có thể che giấu vấn đề tạm thời nhưng dẫn đến alert muộn hoặc sai lệch thống kê lớn hơn, vi phạm nguyên tắc monitoring đáng tin cậy. -
✅ [ĐÚNG] Update the model monitoring job to use the more recent training data that was used to retrain the model.
🛠️ Phân tích: Như đã giải thích ở phần đáp án đúng. Đây là hành động trực tiếp khắc phục skew bằng cách đồng bộ baseline với mô hình mới, đảm bảo phát hiện chính xác drift/skew thực tế trong tương lai. -
❌ [SAI] Temporarily disable the alert. Enable the alert again after a sufficient amount of new production traffic has passed through the Vertex AI endpoint.
🧩 Phân tích: Tắt alert tạm thời (quadisable_alerttrong job config) bỏ qua vấn đề baseline cũ, không giải quyết skew. Serving traffic mới vẫn so với baseline cũ → alert sẽ tái phát ngay khi bật lại. Không khuyến khích vì mất monitoring production quan trọng. -
❌ [SAI] Temporarily disable the alert until the model can be retrain again on newer training data. Retrain the model again after a sufficient amount of new production traffic has passed through the Vertex AI endpoint.
🧩 Phân tích: Sai kép: (1) Tắt alert và chờ retrain lại là lãng phí (mô hình mới đã deploy tốt); (2) Retrain dựa trên production traffic thay vì thu thập training data riêng biệt có thể gây data leakage hoặc bias. Không khớp quy trình Vertex AI, chỉ làm phức tạp hóa vấn đề.
- A Use the features for monitoring. Set a monitoring-frequency value that is higher than the default.
- B Use the features for monitoring. Set a prediction-sampling-rate value that is closer to 1 than 0.
- C Use the features and the feature attributions for monitoring. Set a monitoring-frequency value that is lower than the default.
- D Use the features and the feature attributions for monitoring. Set a prediction-sampling-rate value that is closer to 0 than 1.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh việc triển khai Vertex AI Model Monitoring trên Google Cloud để phát hiện drift (sự thay đổi) trong mô hình dự báo doanh số bán hàng được xây dựng bằng Vertex AI. Các yếu tố chính cần xem xét:
- Dữ liệu đầu vào: Dự báo dựa trên dữ liệu giao dịch lịch sử, nhưng sắp tới sẽ có thay đổi trong phân phối đặc trưng (feature distributions) và tương quan giữa các đặc trưng (correlations between features).
- Tải lượng: Dự kiến lượng request dự đoán lớn (large volume of prediction requests).
- Mục tiêu: Sử dụng Model Monitoring cho drift detection, đồng thời giảm thiểu chi phí (minimize the cost).
- Các thông số chính trong Vertex AI Model Monitoring (theo tài liệu cập nhật mới nhất đến 2026):
- Features monitoring: Theo dõi sự thay đổi phân phối đặc trưng (distribution drift).
- Feature attributions (giải thích mô hình, như XAI): Theo dõi sự thay đổi trong tầm quan trọng của đặc trưng, giúp phát hiện gián tiếp sự thay đổi tương quan giữa các đặc trưng.
- monitoring-frequency: Tần suất kiểm tra (mặc định 1 ngày/lần); giá trị cao hơn = kiểm tra thường xuyên hơn → chi phí cao hơn.
- prediction-sampling-rate: Tỷ lệ lấy mẫu dự đoán (0-1); gần 1 = lấy mẫu nhiều → chi phí cao; gần 0 = lấy mẫu ít → chi phí thấp, phù hợp với volume lớn.
📘 Tài liệu tham khảo:
- Vertex AI Model Monitoring Overview (cập nhật 2024-2026).
- Configuring Model Monitoring – Chi tiết về sampling rate và frequency để tối ưu chi phí.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use the features and the feature attributions for monitoring. Set a prediction-sampling-rate value that is closer to 0 than 1.
Lý do:
- 🛠️ Phát hiện drift toàn diện: Sử dụng features để theo dõi thay đổi phân phối, kết hợp feature attributions để phát hiện thay đổi tương quan (vì attributions phản ánh cách mô hình đánh giá tương tác giữa features).
- 💰 Tối ưu chi phí: Với volume dự đoán lớn, đặt prediction-sampling-rate gần 0 (ví dụ: 0.01) giúp chỉ lấy mẫu ít request → giảm đáng kể chi phí tính theo số lượng predictions được monitor, mà vẫn đảm bảo phát hiện drift hiệu quả.
- Đây là cách cân bằng giữa độ chính xác phát hiện và chi phí thấp nhất theo best practices của Google Cloud (không cần tăng frequency).
❌ Giải thích tất cả các phương án (đúng/sai)
-
Use the features for monitoring. Set a monitoring-frequency value that is higher than the default.
❌ Sai: Chỉ dùng features không đủ để phát hiện thay đổi tương quan giữa features (cần attributions). Đặt frequency cao hơn mặc định (ví dụ: hàng giờ thay vì hàng ngày) sẽ tăng số lần kiểm tra → chi phí cao hơn, trái với mục tiêu minimize cost. -
Use the features for monitoring. Set a prediction-sampling-rate value that is closer to 1 than 0.
❌ Sai: Chỉ dùng features thiếu sót phần tương quan. Đặt sampling-rate gần 1 (ví dụ: 0.9) nghĩa là lấy mẫu gần như tất cả predictions → chi phí rất cao với volume lớn, không tối ưu. -
Use the features and the feature attributions for monitoring. Set a monitoring-frequency value that is lower than the default.
❌ Sai: Kết hợp features + attributions là đúng cho drift và tương quan. Tuy nhiên, frequency thấp hơn mặc định (ví dụ: hàng tuần) có thể bỏ lỡ thay đổi nhanh chóng trong feature distributions/correlations → không đảm bảo phát hiện kịp thời, dù tiết kiệm chi phí một phần nhưng không phải cách min cost tối ưu (vẫn tốn kém hơn nếu không điều chỉnh sampling rate). -
Use the features and the feature attributions for monitoring. Set a prediction-sampling-rate value that is closer to 0 than 1.
✅ Đúng: Như đã giải thích ở trên – đầy đủ phát hiện drift/tương quan và sampling-rate thấp là chìa khóa min cost với high-volume traffic.