Ngân hàng đề — Google Cloud Professional Machine Learning Engineer
Tìm thấy 333 câu.
•Customer_id
•Product_id
•Date
•Days_since_last_purchase (measured in days)
•Average_purchase_frequency (measured in 1/days)
•Purchase (binary class, if customer purchased product on the Date)
You need to interpret your model’s results for each individual prediction. What should you do?
- A Create a BigQuery table. Use BigQuery ML to build a boosted tree classifier. Inspect the partition rules of the trees to understand how each prediction flows through the trees.
- B Create a Vertex AI tabular dataset. Train an AutoML model to predict customer purchases. Deploy the model to a Vertex AI endpoint and enable feature attributions. Use the “explain” method to get feature attribution values for each individual prediction.
- C Create a BigQuery table. Use BigQuery ML to build a logistic regression classification model. Use the values of the coefficients of the model to interpret the feature importance, with higher values corresponding to more importance
- D Create a Vertex AI tabular dataset. Train an AutoML model to predict customer purchases. Deploy the model to a Vertex AI endpoint. At each prediction, enable L1 regularization to detect non-informative features.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm
📖 Nội dung câu hỏi:
Câu hỏi mô tả tình huống bạn làm việc cho một công ty bán lẻ, cần xây dựng mô hình dự đoán xem khách hàng có mua sản phẩm vào một ngày cụ thể hay không (nhiệm vụ phân loại nhị phân dựa trên trường Purchase). Dữ liệu đã được xử lý thành bảng với các trường:
Customer_id: ID khách hàng.Product_id: ID sản phẩm.Date: Ngày cụ thể.Days_since_last_purchase: Số ngày kể từ lần mua cuối cùng (đơn vị ngày).Average_purchase_frequency: Tần suất mua trung bình (đơn vị 1/ngày).Purchase: Nhãn nhị phân (1 nếu mua, 0 nếu không).
Yêu cầu chính là interpret model's results for each individual prediction, nghĩa là cần giải thích kết quả mô hình cho từng dự đoán riêng lẻ (không phải tổng quát toàn bộ mô hình), để hiểu đóng góp của từng feature vào quyết định cuối cùng cho một khách hàng cụ thể. Đây là nhu cầu về explainability ở mức instance-level, thường sử dụng kỹ thuật như feature attribution (ví dụ: SHAP values).
🎯 Đáp án đúng:
[ĐÚNG] Create a Vertex AI tabular dataset. Train an AutoML model to predict customer purchases. Deploy the model to a Vertex AI endpoint and enable feature attributions. Use the “explain” method to get feature attribution values for each individual prediction.
✅ Lý do lựa chọn đáp án đúng:
Phương án này hoàn hảo vì Vertex AI AutoML Tabular (cập nhật đến 2026) hỗ trợ feature attributions dựa trên SHAP (SHapley Additive exPlanations), cho phép kích hoạt khi deploy model lên endpoint. Sau đó, sử dụng phương thức explain trong Vertex AI Prediction API để lấy giá trị attribution cho từng prediction riêng lẻ, hiển thị mức độ ảnh hưởng tích cực/tiêu cực của từng feature (như Days_since_last_purchase làm giảm xác suất mua nếu lớn). Điều này trực tiếp đáp ứng yêu cầu interpret individual predictions. Vertex AI tabular dataset phù hợp với dữ liệu dạng bảng CSV/JSON/BigQuery.
🛠️ Giải thích tất cả các phương án
-
[SAI] Create a BigQuery table. Use BigQuery ML to build a boosted tree classifier. Inspect the partition rules of the trees to understand how each prediction flows through the trees.
❌ Phân tích sai: BigQuery ML hỗ trợ boosted tree (ARIMA_PLUS hoặc BOOSTED_TREE_CLASSIFIER), nhưng việc inspect partition rules chỉ cung cấp global tree structure (luật phân chia chung), không hỗ trợ individual prediction paths một cách dễ dàng hoặc tự động cho từng instance. Không có API trực tiếp để trace flow cho từng prediction riêng lẻ, dẫn đến khó interpret chi tiết. -
[ĐÚNG] Create a Vertex AI tabular dataset. Train an AutoML model to predict customer purchases. Deploy the model to a Vertex AI endpoint and enable feature attributions. Use the “explain” method to get feature attribution values for each individual prediction.
✅ Phân tích đúng: Như đã giải thích ở trên, đây là cách chuẩn xác nhất với Vertex AI Explainable AI (XAI), hỗ trợ AutoML Tabular từ 2021 và cập nhật đến 2026 với tích hợp SHAP integral gradients.explainmethod trả về attribution scores cho từng feature per prediction, dễ visualize (bar charts, etc.). Hoàn toàn phù hợp dữ liệu tabular. -
[SAI] Create a BigQuery table. Use BigQuery ML to build a logistic regression classification model. Use the values of the coefficients of the model to interpret the feature importance, with higher values corresponding to more importance.
❌ Phân tích sai: Logistic regression trong BigQuery ML cho coefficients (quaML.WEIGHTS), nhưng đây chỉ là global feature importance (tổng quát cho toàn mô hình), không áp dụng cho individual predictions. Coefficients không giải thích tại sao một prediction cụ thể lại là 1/0 cho một khách hàng, vì thiếu instance-level attribution. -
[SAI] Create a Vertex AI tabular dataset. Train an AutoML model to predict customer purchases. Deploy the model to a Vertex AI endpoint. At each prediction, enable L1 regularization to detect non-informative features.
❌ Phân tích sai: L1 regularization (lasso) dùng trong training phase để chọn feature (sparsity), không áp dụng "at each prediction" (runtime). Vertex AI AutoML không hỗ trợ enable L1 per prediction; nó là hyperparameter cố định lúc train. Phương án này không cung cấp interpret cho individual predictions mà chỉ detect feature kém tầm quan trọng toàn cục.
📘 Tài liệu tham khảo (cập nhật mới nhất đến 2026)
- Vertex AI Feature Attributions: Vertex AI Explainable AI Documentation (hỗ trợ AutoML Tabular, SHAP từ 2023+).
- BigQuery ML Limitations: BigQuery ML Docs - Model Insights (chỉ global explain, không instance-level cho trees).
- AutoML Tabular Workbench: Vertex AI AutoML Tabular (deploy với explainable AI enabled).
- Google Cloud ML Certification Guide (Professional ML Engineer): Xác nhận yêu cầu instance-level explain thường dùng Vertex AI XAI (Exam topics: Design explainable solutions).
Hy vọng phân tích này giúp bạn nắm vững! 🚀
- A Use the Vertex AI Vision Occupancy Analytics model.
- B Use the Vertex AI Vision Person/vehicle detector model.
- C Train an AutoML object detection model on an annotated dataset by using Vertex AutoML.
- D Train a Seq2Seq+ object detection model on an annotated dataset by using Vertex AutoML.
Xem giải thích
🧩 Giải thích chi tiết nội dung câu hỏi
Câu hỏi mô tả một công ty thu thập video trực tiếp (live video footage) từ khu vực checkout trong các cửa hàng bán lẻ. Nhiệm vụ là xây dựng một mô hình (model) để phát hiện số lượng khách hàng đang chờ dịch vụ (waiting for service) theo thời gian thực gần (near real time) dựa trên video này. Yêu cầu chính là triển khai giải pháp nhanh chóng (quickly) và với nỗ lực tối thiểu (minimal effort).
📌 Mục tiêu cốt lõi: Không cần huấn luyện mô hình từ đầu, ưu tiên sử dụng các công cụ sẵn có của Google Cloud Vertex AI để xử lý video, đếm occupancy (số lượng người) một cách tự động và chính xác cho kịch bản bán lẻ.
✅ Đáp án đúng: Use the Vertex AI Vision Occupancy Analytics model
Lý do lựa chọn:
Đây là lựa chọn tối ưu vì Vertex AI Vision Occupancy Analytics là mô hình được huấn luyện sẵn (pre-trained model) chuyên biệt cho việc phân tích occupancy từ video trực tiếp, bao gồm đếm số lượng khách hàng đang chờ đợi (queue length), thời gian chờ trung bình, và các chỉ số liên quan đến dòng người trong khu vực như checkout. Nó hỗ trợ xử lý near real-time qua streaming video, không yêu cầu annotation dữ liệu hay huấn luyện thêm, giúp triển khai nhanh chóng và minimal effort. Theo tài liệu cập nhật mới nhất (2024-2026), mô hình này được thiết kế chính xác cho use case bán lẻ như đếm khách chờ thanh toán.
🛠️ Cách triển khai: Upload video stream vào Vertex AI Vision, kích hoạt Occupancy Analytics để nhận kết quả ngay lập tức.
📘 Nguồn tham khảo:
- Vertex AI Vision Occupancy Analytics Overview (Google Cloud Docs, cập nhật 2025).
- Vertex AI Vision Models.
📋 Phân tích tất cả các phương án
-
✅ Use the Vertex AI Vision Occupancy Analytics model
Giải thích đúng: Như đã nêu ở trên, mô hình này được tối ưu hóa cho việc đếm occupancy và queue waiting trong video bán lẻ, hỗ trợ real-time analytics mà không cần custom training. Hoàn hảo cho yêu cầu "quickly and minimal effort". -
❌ Use the Vertex AI Vision Person/vehicle detector model
Giải thích sai: Mô hình này chỉ phát hiện vị trí và số lượng người/xe (person/vehicle detection) trong khung hình, nhưng không chuyên phân tích occupancy, queue waiting hay thời gian chờ. Nó không cung cấp metrics như số khách đang chờ thanh toán, dẫn đến thiếu độ chính xác cho use case cụ thể và vẫn cần post-processing thêm, không "minimal effort". -
❌ Train an AutoML object detection model on an annotated dataset by using Vertex AutoML
Giải thích sai: Lựa chọn này yêu cầu chuẩn bị dataset đã annotate (labeled) và huấn luyện mô hình object detection từ đầu qua Vertex AutoML. Quá trình này tốn thời gian (collect data, label, train ~ hours/days), không đáp ứng "implement quickly" và "minimal effort". Phù hợp hơn cho custom objects, không phải occupancy chuẩn. -
❌ Train a Seq2Seq+ object detection model on an annotated dataset by using Vertex AutoML
Giải thích sai: Seq2Seq+ là mô hình nâng cao cho object detection với sequence processing (phù hợp video dài), nhưng vẫn cần dataset annotated và huấn luyện custom, phức tạp hơn AutoML thông thường. Không nhanh chóng, effort cao, và không tận dụng pre-trained models sẵn có cho occupancy analytics.
Kết luận tổng quát 🏆: Vertex AI Vision cung cấp các mô hình chuyên biệt như Occupancy Analytics để giải quyết nhanh vấn đề real-time counting trong retail, tiết kiệm chi phí và thời gian so với training từ đầu. Nếu cần tùy chỉnh sâu hơn, mới xem xét AutoML!
- A Use Tabular Workflow for Wide & Deep through Vertex AI Pipelines to jointly train wide linear models and deep neural networks
- B Use Google Kubernetes Engine to build a custom training pipeline for XGBoost-based models
- C Use Tabular Workflow for TabNet through Vertex AI Pipelines to train attention-based models
- D Use Cloud Composer to build the training pipelines for custom deep learning-based models
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả tình huống bạn là một nhà phân tích tại một ngân hàng lớn, đang phát triển một pipeline ML mạnh mẽ, có khả năng mở rộng để huấn luyện nhiều mô hình regression và classification trên dữ liệu tabular (dữ liệu bảng). Ưu tiên hàng đầu là model interpretability (khả năng giải thích mô hình, rất quan trọng trong lĩnh vực ngân hàng để tuân thủ quy định như GDPR hoặc kiểm toán). Bạn muốn productionize pipeline nhanh nhất có thể (triển khai sản xuất nhanh chóng mà không cần xây dựng từ đầu).
📌 Yêu cầu chính: Chọn giải pháp trên Google Cloud (Vertex AI) hỗ trợ pipeline tự động, scalable, tập trung interpretability cho tabular data, và nhanh triển khai.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Tabular Workflow for TabNet through Vertex AI Pipelines to train attention-based models
🛠️ Lý do chi tiết:
- TabNet là mô hình attention-based chuyên cho dữ liệu tabular, sử dụng cơ chế self-attention để tự động tập trung vào các feature quan trọng, mang lại interpretability cao (có thể visualize attention masks để giải thích quyết định mô hình – lý tưởng cho banking).
- Tabular Workflow trong Vertex AI Pipelines cho phép tự động hóa toàn bộ pipeline (preprocessing, training, evaluation) chỉ với vài cú click hoặc code đơn giản, hỗ trợ cả regression và classification.
- Nhanh productionize: Không cần code custom, scalable tự động trên Vertex AI, phù hợp kiến trúc MLOps mới nhất (cập nhật 2024-2026 với Vertex AI v2).
- So với các option khác, đây là lựa chọn optimized cho interpretability + tốc độ.
📘 Nguồn tham khảo: Vertex AI Documentation - Tabular Workflows & TabNet Paper & Integration (cập nhật GA năm 2023, enhanced 2025).
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên tiêu chí: interpretability, scalability, hỗ trợ regression/classification, và tốc độ productionize.
-
Use Tabular Workflow for Wide & Deep through Vertex AI Pipelines to jointly train wide linear models and deep neural networks
❌ Sai vì: Wide & Deep kết hợp linear (interpretable) và DNN (deep neural networks – black-box, interpretability thấp), không ưu tiên giải thích mô hình như attention mechanisms. Tuy dùng Vertex AI Pipelines nhanh, nhưng không focus interpretability cho banking (DNN phần khó audit). Không phải lựa chọn tối ưu nhất. -
Use Google Kubernetes Engine to build a custom training pipeline for XGBoost-based models
❌ Sai vì: GKE mạnh scalable cho custom pipeline, XGBoost tốt cho tabular (hỗ trợ regression/classification), nhưng phải build từ đầu (code Kubeflow hoặc custom scripts) → chậm productionize. XGBoost interpretable (feature importance), nhưng không "nhanh nhất" và thiếu automation đầy đủ như Vertex AI Workflows. -
Use Tabular Workflow for TabNet through Vertex AI Pipelines to train attention-based models
✅ Đúng vì: Như giải thích trên – TabNet attention-based siêu interpretable (visualize feature attention), Tabular Workflow tự động hóa pipeline end-to-end trên Vertex AI (preprocess → train → deploy), hỗ trợ đa mô hình tabular, scalable, và nhanh nhất (no-code/low-code). Hoàn hảo cho yêu cầu! -
Use Cloud Composer to build the training pipelines for custom deep learning-based models
❌ Sai vì: Cloud Composer (Apache Airflow) tốt cho orchestrate pipeline phức tạp, nhưng custom DL models thường black-box (interpretability kém). Phải build thủ công DAGs → chậm productionize, không optimized cho tabular/ML interpretability. Không phải lựa chọn nhanh/scalable cho MLOps Vertex AI.
🧠 Tóm tắt insight: Trong Vertex AI (cập nhật 2026), Tabular Workflows là "silver bullet" cho tabular ML với interpretability cao, giúp ngân hàng deploy nhanh mà vẫn compliant! Nếu cần customize thêm, kết hợp với Vertex AI Feature Store. 🚀
- A Create a Vertex AI custom training job with GPU accelerators for the second worker pool. Use tf.distribute.MultiWorkerMirroredStrategy for distribution.
- B Create a Vertex AI custom distributed training job with Reduction Server. Use N1 high-memory machine type instances for the first and second pools, and use N1 high-CPU machine type instances for the third worker pool.
- C Create a training job that uses Cloud TPU VMs. Use tf.distribute.TPUStrategy for distribution.
- D Create a Vertex AI custom training job with a single worker pool of A2 GPU machine type instances. Use tf.distribute.MirroredStrategv for distribution.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc cấu hình một công việc huấn luyện (training job) cho mô hình Transformer được phát triển bằng TensorFlow, dùng để dịch văn bản. Dữ liệu huấn luyện gồm hàng triệu tài liệu lưu trong Cloud Storage bucket (Google Cloud Storage). Mục tiêu chính là:
- Sử dụng distributed training (huấn luyện phân tán) để giảm thời gian huấn luyện.
- Tối ưu hóa nỗ lực sửa đổi code (minimize effort to modify code) và quản lý cấu hình cluster (manage cluster’s configuration).
🛠️ Bối cảnh kỹ thuật:
- Vertex AI (dịch vụ ML managed trên Google Cloud) cung cấp custom training jobs hỗ trợ distributed training tự động quản lý cluster (scale workers, accelerators như GPU/TPU).
- TensorFlow hỗ trợ các strategy phân tán như
MultiWorkerMirroredStrategy(cho multi-node GPU/CPU, dễ tích hợp với ít thay đổi code). - Yêu cầu phiên bản mới nhất đến 2026: Vertex AI hỗ trợ multi-worker pools với GPU accelerators (A100/A2 VMs), TPU v5e/v5p, và tích hợp TensorFlow 2.15+ (cập nhật 2024-2026 nhấn mạnh managed distributed training với zero-config cluster management).
📘 Nguồn tham khảo:
- Vertex AI Custom Training (Google Cloud Docs, cập nhật 2025).
- TensorFlow Distributed Training (TensorFlow 2.16+, 2025).
- Vertex AI Multi-Worker Pools (hỗ trợ GPU/TPU pools linh hoạt).
✅ Đáp án đúng
Create a Vertex AI custom training job with GPU accelerators for the second worker pool. Use tf.distribute.MultiWorkerMirroredStrategy for distribution.
Lý do lựa chọn:
- Vertex AI custom training job tự động quản lý cluster (không cần thủ công setup Kubernetes/GKE), hỗ trợ multiple worker pools (pool 1 CPU, pool 2 GPU) để tối ưu chi phí/thời gian.
tf.distribute.MultiWorkerMirroredStrategylà strategy chuẩn cho distributed training multi-node/multi-GPU, chỉ cần thay đổi ít code (wrap model trong strategy.scope()), Vertex AI tự detect và config workers qua environment variables (TF_CONFIG).- Phù hợp Transformer (tốn compute cao), GPU accelerators (như A100/NVIDIA) trên pool 2 tăng tốc sync gradients hiệu quả.
- Minimize effort: Không cần code phức tạp, Vertex AI handle scaling, data từ GCS tự mount.
📋 Giải thích tất cả các phương án
-
✅ Create a Vertex AI custom training job with GPU accelerators for the second worker pool. Use tf.distribute.MultiWorkerMirroredStrategy for distribution.
Đúng vì: Như giải thích trên, đây là cách managed nhất trên Vertex AI, hỗ trợ multi-pool (pool 2 GPU cho heavy compute), strategy dễ tích hợp mà không sửa code nhiều. Vertex AI tự config TF_CONFIG cho MultiWorkerMirroredStrategy, giảm effort quản lý cluster xuống mức thấp nhất. -
❌ Create a Vertex AI custom distributed training job with Reduction Server. Use N1 high-memory machine type instances for the first and second pools, and use N1 high-CPU machine type instances for the third worker pool.
Sai vì: Reduction Server là tính năng của AWS SageMaker (không tồn tại trên Vertex AI/Google Cloud). Vertex AI không hỗ trợ Reduction Server; thay vào đó dùng built-in all-reduce. N1 machine types (high-mem/CPU) lỗi thời (2025+ ưu tiên C4/A3/N4), và config 3 pools thủ công tăng effort quản lý, không minimize code/cluster config. -
❌ Create a training job that uses Cloud TPU VMs. Use tf.distribute.TPUStrategy for distribution.
Sai vì: Cloud TPU VMs yêu cầu sửa code nhiều hơn (TPUStrategy cần compile graph, data pipeline TPU-specific như tf.data.experimental.ops`), không "minimize effort to modify code". Vertex AI hỗ trợ TPU, nhưng MultiWorkerMirroredStrategy linh hoạt hơn cho GPU/CPU mix. TPU tốt cho Transformer nhưng không phải lựa chọn managed đơn giản nhất cho multi-GPU setup. -
❌ Create a Vertex AI custom training job with a single worker pool of A2 GPU machine type instances. Use tf.distribute.MirroredStrategv for distribution.
Sai vì:MirroredStrategy(lưu ý lỗi chính tả "Strategv") chỉ hỗ trợ single-node multi-GPU, không phải distributed multi-node (không scale workers). Single pool A2 GPU (A100 VMs) không đủ cho "distributed training to reduce time" với millions docs. Tăng effort scale thủ công nếu cần multi-node.
🧠 Kết luận: Lựa chọn đúng tận dụng Vertex AI managed features + TensorFlow native strategy để zero-effort cluster management, lý tưởng cho production ML workflows trên Google Cloud (cập nhật 2026).
-
A
1. Create a Vertex AI managed dataset.
2. Use a Vertex AI training pipeline to train your model.
3. Generate batch predictions in Vertex AI. -
B
1. Use a Vertex AI Pipelines custom training job component to tram your model.
2. Generate predictions by using a Vertex AI Pipelines model batch predict component. -
C
1. Upload your dataset to BigQuery.
2. Use a Vertex AI custom training job to train your model.
3. Generate predictions by using Vertex Al SDK custom prediction routines. -
D
1. Use Vertex AI Experiments to train your model.
2. Register your model in Vertex AI Model Registry.
3. Generate batch predictions in Vertex AI.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi
Câu hỏi tập trung vào việc phát triển quy trình huấn luyện (training) và triển khai (running) mô hình tùy chỉnh (custom model) trong môi trường production. Yêu cầu chính là hiển thị lineage (dòng dõi, traceability) cho mô hình và các dự đoán (predictions).
- Lineage ở đây đề cập đến khả năng theo dõi toàn bộ chuỗi sự kiện từ dữ liệu đầu vào, qua quá trình huấn luyện, đến mô hình cuối cùng và các dự đoán được tạo ra. Điều này rất quan trọng trong ML Ops để đảm bảo tính minh bạch, tái tạo (reproducibility), và tuân thủ quy định.
- Trong Vertex AI (Google Cloud) – nền tảng ML chính thức của Google – lineage được hỗ trợ tốt nhất qua Vertex AI Pipelines, cho phép theo dõi metadata, artifacts, và dependencies một cách tự động qua các component trong pipeline.
- Mục tiêu: Chọn phương án tích hợp đầy đủ lineage cho cả training và predictions, không chỉ riêng lẻ từng bước.
📘 Tài liệu tham khảo:
- Vertex AI Pipelines Documentation (cập nhật 2024-2026: hỗ trợ Kubeflow-based pipelines với metadata tracking).
- Vertex AI Model Lineage (mới nhất: tích hợp Artifact và Metadata Store cho end-to-end tracing).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Lựa chọn thứ 2 – 1. Use a Vertex AI Pipelines custom training job component to train your model. 2. Generate predictions by using a Vertex AI Pipelines model batch predict component.
Lý do 🛠️:
- Phương án này sử dụng Vertex AI Pipelines làm nền tảng chính, với custom training job component cho training và model batch predict component cho predictions.
- Vertex AI Pipelines tự động ghi nhận full lineage (bao gồm inputs/outputs, parameters, artifacts) qua Metadata Store, cho phép visualize và query lineage graph từ training đến predictions.
- Đây là cách chuẩn và được khuyến nghị cho production ML workflows, đảm bảo traceability end-to-end mà không cần code thủ công. Phù hợp phiên bản mới nhất (2026): hỗ trợ managed pipelines với auto-lineage.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng phương án một cách chi tiết, với đánh giá đúng/sai dựa trên khả năng cung cấp lineage đầy đủ cho model và predictions:
-
❌ Phương án 1 (SAI):
- Create a Vertex AI managed dataset.
- Use a Vertex AI training pipeline to train your model.
- Generate batch predictions in Vertex AI.
Giải thích sai: Mặc dù Vertex AI managed dataset và training pipeline hỗ trợ một phần lineage (qua metadata), nhưng batch predictions riêng lẻ không liên kết tự động với pipeline trước đó. Không tạo thành end-to-end lineage graph, dẫn đến khó trace predictions về nguồn gốc model/dataset. Thiếu tích hợp pipelines cho toàn bộ flow.
-
✅ Phương án 2 (ĐÚNG):
- Use a Vertex AI Pipelines custom training job component to train your model.
- Generate predictions by using a Vertex AI Pipelines model batch predict component.
Giải thích đúng: Như đã nêu ở trên, toàn bộ quy trình nằm trong một Vertex AI Pipeline, tự động capture lineage qua components. Có thể query lineage bằng API/UI, visualize DAG với đầy đủ inputs/outputs. Hoàn hảo cho production traceability (xác nhận qua docs 2026).
-
❌ Phương án 3 (SAI):
- Upload your dataset to BigQuery.
- Use a Vertex AI custom training job to train your model.
- Generate predictions by using Vertex AI SDK custom prediction routines.
Giải thích sai: BigQuery chỉ lưu dataset (có lineage riêng nhưng không tích hợp Vertex AI), custom training job có lineage cơ bản, nhưng SDK custom prediction routines là code thủ công – không tự động ghi lineage. Không có pipeline thống nhất, khó trace predictions về model/dataset.
-
❌ Phương án 4 (SAI):
- Use Vertex AI Experiments to train your model.
- Register your model in Vertex AI Model Registry.
- Generate batch predictions in Vertex AI.
Giải thích sai: Vertex AI Experiments tốt cho tracking trials (metrics/lineage per experiment), Model Registry lưu model metadata, nhưng batch predictions không liên kết tự động với experiment/registry. Thiếu pipeline orchestration, lineage bị fragmented – không end-to-end cho predictions.
Kết luận 🎯: Chỉ phương án 2 đảm bảo lineage hoàn chỉnh, phù hợp best practices Vertex AI Pipelines cho ML production! Nếu cần code sample hoặc demo, hãy cho tôi biết nhé! 🚀
- A Use the Vision API to parse the text from each PDF file. Use the Natural Language API analyzeSentiment feature to infer overall satisfaction scores.
- B Use the Vision API to parse the text from each PDF file. Use the Natural Language API analyzeEntitySentiment feature to infer overall satisfaction scores.
- C Uptrain a Document AI custom extractor to parse the text in the comments section of each PDF file. Use the Natural Language API analyzeSentiment feature to infer overall satisfaction scores.
- D Uptrain a Document AI custom extractor to parse the text in the comments section of each PDF file. Use the Natural Language API analyzeEntitySentiment feature to infer overall satisfaction scores.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một tình huống thực tế trong ngành khách sạn: Bạn có bộ dữ liệu chứa các bình luận viết tay của khách hàng được quét từ biểu mẫu phản hồi trên giấy, lưu trữ dưới dạng file PDF. Mỗi biểu mẫu có bố cục giống hệt nhau (same layout). Nhiệm vụ là dự đoán nhanh điểm hài lòng tổng thể (overall satisfaction score) từ phần bình luận của khách hàng trên mỗi biểu mẫu.
🛠️ Yêu cầu chính cần giải quyết:
- Trích xuất văn bản từ PDF quét (scanned), đặc biệt là phần bình luận cụ thể (comments section), vì là handwritten và có layout cố định.
- Phân tích cảm xúc để suy ra điểm hài lòng tổng thể (không phải chi tiết từng thực thể).
- Cần giải pháp nhanh chóng (quickly), tận dụng các dịch vụ Google Cloud để xử lý tự động quy mô lớn.
📘 Kiến thức liên quan (cập nhật đến 2026):
- Sử dụng Document AI (phiên bản mới nhất với Custom Extractors hỗ trợ handwritten text và fixed-layout documents hiệu quả hơn Vision API cho structured extraction).
- Vision API (Document OCR) phù hợp cho OCR general nhưng kém chính xác hơn với layout cụ thể.
- Natural Language API:
analyzeSentimentcho sentiment tổng thể document-level;analyzeEntitySentimentcho sentiment theo entity (không phù hợp overall score). - Nguồn tham khảo:
- Document AI Custom Extractors (Google Cloud Docs, 2024-2026 updates).
- Natural Language API Sentiment Analysis (v4 API hỗ trợ multilingual sentiment).
- Vision API OCR Limitations.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Uptrain a Document AI custom extractor to parse the text in the comments section of each PDF file. Use the Natural Language API analyzeSentiment feature to infer overall satisfaction scores.
Lý do:
- Document AI custom extractor được huấn luyện (uptrain) cho layout cố định, trích xuất chính xác phần comments section từ PDF scanned/handwritten, nhanh hơn và chính xác hơn Vision API (giảm lỗi OCR ở vùng cụ thể).
- analyzeSentiment tính toán sentiment tổng thể (magnitude/score từ -1.0 đến 1.0), lý tưởng để map thành overall satisfaction score.
- Giải pháp end-to-end nhanh chóng, scalable cho production (batch processing).
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt với emoji đánh dấu đúng/sai.
-
❌ [SAI] Use the Vision API to parse the text from each PDF file. Use the Natural Language API analyzeSentiment feature to infer overall satisfaction scores.
Giải thích sai: Vision API (Document OCR) có thể parse text từ PDF, nhưng không tối ưu cho layout cố định và vùng comments cụ thể → dễ lỗi với handwritten text, tốn thời gian post-processing. Phần analyzeSentiment đúng nhưng bước đầu kém hiệu quả, không "quickly" như yêu cầu. -
❌ [SAI] Use the Vision API to parse the text from each PDF file. Use the Natural Language API analyzeEntitySentiment feature to infer overall satisfaction scores.
Giải thích sai: Vision API như trên (không phù hợp layout cụ thể). analyzeEntitySentiment chỉ phân tích sentiment theo từng entity (e.g., "room" positive, "service" negative), không cho overall score → không match yêu cầu tổng thể. -
✅ [ĐÚNG] Uptrain a Document AI custom extractor to parse the text in the comments section of each PDF file. Use the Natural Language API analyzeSentiment feature to infer overall satisfaction scores.
Giải thích đúng: Như phần trên, custom extractor uptrain cho fixed layout → extract chính xác vùng comments. analyzeSentiment hoàn hảo cho overall satisfaction. Giải pháp tối ưu nhất theo best practices Google Cloud 2026. -
❌ [SAI] Uptrain a Document AI custom extractor to parse the text in the comments section of each PDF file. Use the Natural Language API analyzeEntitySentiment feature to infer overall satisfaction scores.
Giải thích sai: Document AI custom extractor đúng (tốt nhất cho task). Nhưng analyzeEntitySentiment sai vì chỉ cho sentiment entity-level, không suy ra overall score → thiếu tổng hợp toàn bộ comments.
🧠 Kết luận: Chọn giải pháp kết hợp Document AI cho extraction chính xác + NL API sentiment tổng thể để đạt hiệu suất cao nhất! Nếu triển khai, khuyến nghị test với sample PDFs trên Google Cloud Console. 🚀
dt=datetmie.now().strftime("%Y%m%d%H%M%S")
f"export-{dt}.yaml", f"preprocess-{dt}.yaml", f"train-{dt}.yaml", f"calibrate-{dt}.yaml"
You launch your Vertex AI pipeline as the following:
job = aip.PipelineJob(
display_name="my-awesome-pipeline",
template_path="pipeline.json",
job_id=f"my-awesome-pipeline-{dt}",
parameter_values=params,
enable_caching=True,
location="europe-west1"
)
You perform many model iterations by adjusting the code and parameters of the training step. You observe high costs associated with the development, particularly the data export and preprocessing steps. You need to reduce model development costs. What should you do?
-
A
Change the components’ YAML filenames to export.yaml, preprocess,yaml, f "train-
{dt}.yaml", f"calibrate-{dt).vaml". - B Add the {"kubeflow.v1.caching": True} parameter to the set of params provided to your PipelineJob.
- C Move the first step of your pipeline to a separate step, and provide a cached path to Cloud Storage as an input to the main pipeline.
- D Change the name of the pipeline to f"my-awesome-pipeline-{dt}".
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh việc tối ưu hóa chi phí phát triển mô hình trong Vertex AI Pipelines (Google Cloud). Bạn đã xây dựng một pipeline với 4 bước chính sử dụng Kubeflow v2 API (qua các file YAML component):
export-{dt}.yaml: Xuất dữ liệu từ BigQuery lớn.preprocess-{dt}.yaml: Tiền xử lý dữ liệu.train-{dt}.yaml: Huấn luyện mô hình phân loại.calibrate-{dt}.yaml: Hiệu chỉnh mô hình.
Pipeline được khởi chạy với PipelineJob có enable_caching=True, nhưng mỗi lần lặp lại mô hình (chỉ chỉnh sửa code/params ở bước train), các bước export và preprocess vẫn chạy lại, gây chi phí cao (đặc biệt với BigQuery lớn).
Vấn đề cốt lõi: Tên file YAML chứa timestamp dt (ví dụ: export-20240101120000.yaml), làm cache không khớp giữa các lần chạy vì Kubeflow caching dựa trên hash của component spec (bao gồm tên file YAML). Kết quả: Các bước không thay đổi vẫn bị tính phí lại.
Mục tiêu: Giảm chi phí bằng cách tận dụng caching hiệu quả trong Vertex AI Pipelines (cập nhật đến 2026: Hỗ trợ Kubeflow Pipelines v2 với caching dựa trên input/output hash và component metadata, ưu tiên tên component cố định).
📘 Tài liệu tham khảo:
- Vertex AI Pipelines Caching (Google Cloud Docs, phiên bản mới nhất 2026).
- Kubeflow Pipelines v2 Caching (chính thức Kubeflow).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Change the components’ YAML filenames to export.yaml, preprocess.yaml, f"train-{dt}.yaml", f"calibrate-{dt}.yaml".
Lý do 🛠️:
- Giữ tên file YAML cố định (
export.yaml,preprocess.yaml) cho các bước không thay đổi → Kubeflow v2 sẽ cache thành công dựa trên hash giống nhau (input data từ BigQuery không đổi). - Chỉ dynamic tên cho
train-{dt}.yamlvàcalibrate-{dt}.yaml(các bước bị chỉnh sửa thường xuyên) → Tránh cache sai, chỉ rerun các bước cần thiết. - Với
enable_caching=Trueđã set, chi phí export/preprocess giảm mạnh (có thể >90% nếu data ổn định). Đây là best practice cho iterative development.
📋 Giải thích tất cả các phương án
-
[ĐÚNG] Change the components’ YAML filenames to export.yaml, preprocess.yaml, f"train-{dt}.yaml", f"calibrate-{dt}.yaml".
✅ Đúng vì như phân tích trên: Tên YAML cố định kích hoạt cache hiệu quả cho export/preprocess, giảm chi phí rerun. Phù hợp Kubeflow v2 (hash bao gồm filename). -
[SAI] Add the {"kubeflow.v1.caching": True} parameter to the set of params provided to your PipelineJob.
❌ Sai vìenable_caching=Trueđã được set ở PipelineJob (tương đương Kubeflow caching level pipeline). Thêm paramkubeflow.v1.caching(cũ, v1) không cần thiết và không giải quyết vấn đề tên YAML dynamic – cache vẫn fail do hash thay đổi. -
[SAI] Move the first step of your pipeline to a separate step, and provide a cached path to Cloud Storage as an input to the main pipeline.
❌ Sai vì tách export thành pipeline riêng vẫn yêu cầu chạy thủ công mỗi lần (không tự động cache), tăng complexity quản lý. Không tận dụng native Vertex AI caching; chỉ là workaround kém hiệu quả, không giảm chi phí iterative dev. -
[SAI] Change the name of the pipeline to f"my-awesome-pipeline-{dt}".
❌ Sai vì tên pipeline (job_id) chỉ ảnh hưởng metadata/job tracking, không liên quan đến component caching. Cache dựa trên component spec (YAML hash/input), không phải job name. Job name đã códt, nhưng vấn đề ở component names.
- A Create a n2-standard-4 VM instance and install Java, Scala, and Apache Spark dependencies on it.
- B Create a Google Kubernetes Engine cluster with a basic node pool configuration, install Java, Scala, and Apache Spark dependencies on it.
- C Create a Standard (1 master, 3 workers) Dataproc cluster, and run a Vertex AI Workbench notebook instance on it.
- D Create a Vertex AI Workbench notebook with instance type n2-standard-4.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc xây dựng một proof of concept (POC) để di chuyển một công việc data science (data science job) từ hạ tầng on-premises sang Google Cloud. Các workload hiện tại sử dụng PySpark (một thư viện Python trên nền Apache Spark), và mục tiêu là giảm thiểu chi phí (minimal cost) và công sức (minimal effort).
- Bối cảnh: Startup có nhiều workload data science trên PySpark chạy on-premises. Nhóm cần migrate sang Google Cloud một cách nhanh chóng, dễ dàng cho POC.
- Yêu cầu chính: Đề xuất bước đầu tiên (what should you do first?) để thực hiện migration POC, tận dụng dịch vụ managed của Google Cloud nhằm tránh cài đặt thủ công phức tạp.
- Kiến thức liên quan (cập nhật đến 2026): Google Cloud cung cấp Dataproc như dịch vụ managed Spark/Hadoop, hỗ trợ PySpark native mà không cần cài đặt dependencies. Vertex AI Workbench (trước đây là AI Platform Notebooks) cho phép chạy notebooks Jupyter và kết nối với Dataproc để thực thi PySpark jobs. Đây là cách tối ưu cho migration PySpark với chi phí thấp (Dataproc cluster nhỏ, autoscaling).
📘 Tài liệu tham khảo:
- Dataproc Documentation (phiên bản mới nhất 2026: hỗ trợ Spark 3.5+ và PySpark kernels).
- Vertex AI Workbench & Dataproc Integration.
- Migrate Spark workloads to Dataproc.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create a Standard (1 master, 3 workers) Dataproc cluster, and run a Vertex AI Workbench notebook instance on it.
Lý do:
- 🛠️ Dataproc là dịch vụ fully managed cho Spark/Hadoop trên Google Cloud, hỗ trợ PySpark native ngay từ đầu (không cần install Java/Scala/Spark thủ công). Cluster Standard (1 master + 3 workers) là cấu hình nhỏ gọn, chi phí thấp (~$0.1-0.2/giờ, autoscaling), phù hợp POC.
- 📓 Vertex AI Workbench notebook có thể attach trực tiếp vào Dataproc cluster, cho phép chạy PySpark code từ notebook mà không thay đổi code gốc – migration effort tối thiểu.
- ✅ Đây là bước đầu tiên lý tưởng: Tạo cluster nhanh (5-10 phút), test job PySpark ngay, scale sau. Phù hợp best practice migration từ on-premises Spark.
❌ Phân tích tất cả các phương án
-
[SAI] Create a n2-standard-4 VM instance and install Java, Scala, and Apache Spark dependencies on it.
❌ Sai vì: VM Compute Engine yêu cầu cài đặt thủ công toàn bộ stack (Java 11+, Scala 2.12+, Spark 3.5+), tốn effort cao (scripting, troubleshooting), không managed, dễ lỗi và chi phí vận hành cao hơn Dataproc (không autoscaling). Không phù hợp "minimal effort". -
[SAI] Create a Google Kubernetes Engine cluster with a basic node pool configuration, install Java, Scala, and Apache Spark dependencies on it.
❌ Sai vì: GKE là managed Kubernetes, nhưng vẫn phải install Spark thủ công (qua Helm charts hoặc operators), phức tạp hơn Dataproc (cần config pods, storage). Chi phí cao hơn cho POC nhỏ, effort lớn (DevOps skills cần thiết), không tối ưu cho PySpark native. -
[ĐÚNG] Create a Standard (1 master, 3 workers) Dataproc cluster, and run a Vertex AI Workbench notebook instance on it.
✅ Đúng như đã giải thích ở trên: Managed Spark + notebook integration = minimal cost/effort cho PySpark POC. -
[SAI] Create a Vertex AI Workbench notebook with instance type n2-standard-4.
❌ Sai vì: Workbench chỉ là notebook instance (JupyterLab) với runtime Python/TensorFlow, không có Spark cluster built-in để chạy PySpark jobs lớn. PySpark cần distributed compute (như Dataproc), nếu không sẽ fallback single-node (chậm, không scale), không migrate được workload gốc.
- A Vertex ML Metadata, Vertex AI Feature Store, and Vertex AI Vizier
- B Vertex AI Pipelines, Vertex AI Experiments, and Vertex AI Vizier
- C Vertex ML Metadata, Vertex AI Experiments, and Vertex AI TensorBoard
- D Vertex AI Pipelines, Vertex AI Feature Store, and Vertex AI TensorBoard
Xem giải thích
🧩 Giải thích chi tiết nội dung câu hỏi
Câu hỏi yêu cầu xây dựng một mô hình Machine Learning (ML) để hỗ trợ quyết định phê duyệt đơn vay vốn tại ngân hàng. Bạn cần chọn các dịch vụ Vertex AI phù hợp trong quy trình làm việc (workflow) để:
- Theo dõi (track) các tham số huấn luyện (training parameters) và chỉ số hiệu suất (metrics) theo từng epoch huấn luyện.
- So sánh hiệu suất giữa các phiên bản mô hình (model versions) để chọn mô hình tốt nhất dựa trên các metrics đã chọn.
Mục tiêu chính là quản lý thí nghiệm (experiments), theo dõi metadata huấn luyện và trực quan hóa metrics chi tiết theo epoch. Vertex AI là nền tảng ML end-to-end của Google Cloud, giúp tích hợp các công cụ này một cách liền mạch. (Kiến thức cập nhật đến 2026: Vertex AI Experiments và TensorBoard hỗ trợ theo dõi real-time metrics với tích hợp sâu vào AutoML và custom training jobs).
📘 Tài liệu tham khảo:
✅ Đáp án đúng: Vertex ML Metadata, Vertex AI Experiments, and Vertex AI TensorBoard
Lý do lựa chọn:
- Vertex ML Metadata 🛠️: Quản lý metadata toàn diện, bao gồm tracking training parameters, lineage (dòng dõi dữ liệu/mô hình), và artifacts – giúp ghi nhận chi tiết mọi thay đổi trong huấn luyện.
- Vertex AI Experiments 📊: Cho phép tạo experiments để track và so sánh metrics, parameters giữa các runs/versions mô hình một cách tự động, với giao diện dashboard để chọn best model dựa trên metrics (như accuracy, loss).
- Vertex AI TensorBoard 🔍: Chuyên visualize metrics theo từng epoch (curves loss/accuracy), hỗ trợ so sánh nhiều training runs – lý tưởng cho việc theo dõi huấn luyện sâu.
Bộ ba này bao quát đầy đủ yêu cầu: metadata tracking + experiment comparison + epoch-level visualization, phù hợp workflow ML tiêu chuẩn trên Vertex AI (custom training jobs hoặc AutoML).
❌ Phân tích tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên tính phù hợp với yêu cầu tracking parameters/metrics per epoch và so sánh model versions.
-
Vertex ML Metadata, Vertex AI Feature Store, and Vertex AI Vizier ❌
Giải thích sai: Vertex ML Metadata đúng cho tracking metadata/parameters, nhưng Vertex AI Feature Store chỉ dùng lưu trữ và phục vụ features (dữ liệu đầu vào), không liên quan đến tracking huấn luyện hay metrics epoch. Vertex AI Vizier (nay là Hyperparameter Tuning) chỉ tối ưu siêu tham số, không hỗ trợ so sánh versions đầy đủ hay visualize epoch. Thiếu công cụ experiment và visualization chính. -
Vertex AI Pipelines, Vertex AI Experiments, and Vertex AI Vizier ❌
Giải thích sai: Vertex AI Experiments đúng cho so sánh models, nhưng Vertex AI Pipelines chỉ orchestrate workflow (MLOps pipelines), không track metrics epoch chi tiết. Vertex AI Vizier chỉ tuning hyperparameters, không visualize training curves. Không có metadata tracking đầy đủ hoặc TensorBoard cho epoch metrics. -
Vertex ML Metadata, Vertex AI Experiments, and Vertex AI TensorBoard ✅
(Đã giải thích chi tiết ở phần đáp án đúng – bộ công cụ hoàn hảo!) -
Vertex AI Pipelines, Vertex AI Feature Store, and Vertex AI TensorBoard ❌
Giải thích sai: Vertex AI TensorBoard đúng cho metrics epoch, nhưng Vertex AI Pipelines chỉ quản lý pipeline orchestration, không track experiments/parameters. Vertex AI Feature Store chỉ xử lý features, không hỗ trợ so sánh model versions. Thiếu Experiments và Metadata cho tracking/so sánh toàn diện.
🛡️ Lưu ý cuối: Trong thực tế Vertex AI (2026), bạn có thể tích hợp thêm Model Registry để deploy best model sau Experiments. Khuyến nghị thử nghiệm trên Google Cloud Console để verify!
- A Download a pre-trained object detection model from TensorFlow Hub. Fine-tune the model in Vertex AI Workbench by using the annotated image data.
- B Train an object detection model in AutoML by using the annotated image data.
- C Create a pipeline in Vertex AI Pipelines and configure the AutoMLTrainingJobRunOp component to train a custom object detection model by using the annotated image data.
- D Train an object detection model in Vertex AI custom training by using the annotated image data.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả tình huống bạn làm việc cho một công ty bảo hiểm ô tô, đang chuẩn bị proof-of-concept (POC) cho ứng dụng ML sử dụng hình ảnh xe bị hỏng để suy luận các phần bị hỏng. Đội ngũ đã thu thập bộ dữ liệu hình ảnh được chú thích (annotated images) từ cơ sở dữ liệu yêu cầu bồi thường, với mỗi hình ảnh có bounding box (hộp giới hạn) cho từng phần hỏng và tên phần tương ứng. Bạn có ngân sách đủ để huấn luyện mô hình trên Google Cloud, và mục tiêu là tạo mô hình ban đầu một cách nhanh chóng.
🔑 Yêu cầu chính: Chọn phương pháp nhanh nhất để train object detection model (phát hiện đối tượng với bounding box và nhãn tùy chỉnh), phù hợp cho POC mà không cần code phức tạp. Đây là bài kiểm tra kiến thức về Vertex AI và AutoML Vision Object Detection trên Google Cloud (cập nhật đến phiên bản mới nhất 2026, hỗ trợ import CSV/JSON annotations với bounding boxes).
📘 Tài liệu tham khảo:
- Vertex AI AutoML Vision Object Detection (Google Cloud Docs, 2026).
- Quickstart: Train an object detection model.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Train an object detection model in AutoML by using the annotated image data.
Lý do:
🛠️ AutoML Vision Object Detection là giải pháp nhanh nhất cho POC, chỉ cần upload dữ liệu annotated (hỗ trợ định dạng CSV với bounding box coordinates và labels), tự động train model mà không cần code. Thời gian train chỉ vài giờ, phù hợp ngân sách và yêu cầu "quickly create an initial model". Vertex AI AutoML xử lý toàn bộ pipeline (preprocessing, training, evaluation) với độ chính xác cao cho custom objects như "damaged parts". Đây là lựa chọn tối ưu theo best practices Google Cloud cho non-expert teams.
🔍 Giải thích tất cả các phương án (đúng/sai)
-
[SAI] Download a pre-trained object detection model from TensorFlow Hub. Fine-tune the model in Vertex AI Workbench by using the annotated image data.
❌ Sai vì: Phương pháp này yêu cầu code custom (TensorFlow/Keras scripts) để load pre-trained model (như SSD hoặc EfficientDet từ TF Hub), convert annotations sang TFRecord, và fine-tune trong Vertex AI Workbench (Jupyter notebook). Quá phức tạp và tốn thời gian cho POC nhanh (cần dev expertise, debug data pipeline). Không phù hợp "quickly" so với AutoML no-code. (📘 Xem: Vertex AI Workbench docs). -
[ĐÚNG] Train an object detection model in AutoML by using the annotated image data.
✅ Đúng vì: Như giải thích trên, AutoML tự động hóa toàn bộ, hỗ trợ trực tiếp bounding box annotations (normalized coordinates + labels). Upload dataset qua Console/UI, train chỉ với vài cú click, deploy ngay. Hoàn hảo cho POC với dữ liệu company-specific. (🛠️ Best practice cho <1000 images). -
[SAI] Create a pipeline in Vertex AI Pipelines and configure the AutoMLTrainingJobRunOp component to train a custom object detection model by using the annotated image data.
❌ Sai vì: Vertex AI Pipelines dùng cho orchestration phức tạp (MLOps workflows), và AutoMLTrainingJobRunOp là Kubeflow component để automate AutoML jobs – nhưng quá overkill cho POC đơn giản. Cần thiết kế pipeline YAML/code, không "quickly" (thêm overhead setup). Dùng trực tiếp AutoML Console nhanh hơn nhiều. -
[SAI] Train an object detection model in Vertex AI custom training by using the annotated image data.
❌ Sai vì: Custom training yêu cầu container hóa code (Docker với TF/PyTorch), viết training script xử lý annotations, và submit job qua Vertex AI. Tốn thời gian setup (data conversion, hyperparam tuning), cần ML engineer expertise. Phù hợp production-scale, không phải POC nhanh. (📘 So sánh: Custom vs AutoML in Vertex AI docs).
Tóm tắt khuyến nghị 🚀: Với POC, ưu tiên AutoML để iterate nhanh → evaluate → scale sang custom nếu cần. Nếu dữ liệu lớn hơn, migrate sang Vertex AI Training sau!