Ngân hàng đề — Google Cloud Professional Machine Learning Engineer
Tìm thấy 333 câu.
- A Embed the augmentation functions dynamically in the tf.Data pipeline.
- B Embed the augmentation functions dynamically as part of Keras generators.
- C Use Dataflow to create all possible augmentations, and store them as TFRecords.
- D Use Dataflow to create the augmentations dynamically per training run, and stage them as TFRecords.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc tối ưu hóa pipeline xử lý dữ liệu cho mô hình phát hiện lỗi hình ảnh (visual defect detection) sử dụng TensorFlow và Keras trong một đội ngũ AI của công ty ô tô.
- Vấn đề chính: Bạn cần áp dụng các hàm tăng cường dữ liệu hình ảnh (image augmentation) như dịch chuyển (translation), cắt (cropping), và điều chỉnh độ tương phản (contrast tweaking). Những hàm này được áp dụng ngẫu nhiên (randomly) cho từng batch huấn luyện (each training batch).
- Mục tiêu: Tối ưu thời gian chạy (run time) và sử dụng tài nguyên tính toán (compute resources utilization), nghĩa là pipeline phải nhanh, hiệu quả, hỗ trợ song song hóa (parallelism), và tận dụng GPU/TPU mà không lãng phí lưu trữ hoặc I/O.
- Bối cảnh: Đây là kỹ thuật ML tiêu chuẩn để cải thiện hiệu suất mô hình bằng cách tạo dữ liệu đa dạng on-the-fly, tránh overfitting. Câu hỏi nhấn mạnh dynamic augmentation (áp dụng động theo batch) thay vì pre-generate (tạo trước).
Kiến thức cập nhật đến 2026: Theo TensorFlow 2.15+ và tf.data API mới nhất (hỗ trợ AutoGraph, prefetching, caching nâng cao), tf.data pipeline là lựa chọn hàng đầu cho high-performance data loading, đặc biệt trên Google Cloud Vertex AI hoặc TPU.
📘 Tài liệu tham khảo:
- TensorFlow tf.data Guide (cập nhật 2024-2026).
- Image Augmentation in tf.data.
- Google Cloud Dataflow docs cho TFRecords (không khuyến nghị cho dynamic aug).
✅ Đáp án đúng
Embed the augmentation functions dynamically in the tf.data pipeline.
Lý do lựa chọn:
- 🛠️ tf.data pipeline cho phép nhúng các hàm augmentation động (dynamically) trực tiếp vào luồng dữ liệu, áp dụng ngẫu nhiên per batch với hiệu suất cao nhờ song song hóa (parallelism), prefetching, và GPU/TPU fusion.
- Nó tối ưu run time (giảm latency I/O) và compute resources (không cần lưu trữ dữ liệu augmented lớn), phù hợp hoàn hảo với yêu cầu "randomly apply to each training batch".
- So với các cách khác, tf.data nhanh hơn 10-100x nhờ compiled graph execution (không Python loop).
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên best practices TensorFlow/GCP.
-
✅ Embed the augmentation functions dynamically in the tf.data pipeline.
Đúng 🏆: Như đã giải thích ở trên, đây là cách tối ưu nhất theo TensorFlow docs. Sử dụngtf.image(ví dụ:tf.image.random_crop,tf.image.random_brightness) trongmap()hoặcexperimental.map_parallel()của tf.data.Dataset. Hỗ trợ distributed training trên Vertex AI, giảm bottleneck I/O và tận dụng TPU hiệu quả. -
❌ Embed the augmentation functions dynamically as part of Keras generators.
Sai: Keras generators (nhưtf.keras.utils.image_dataset_from_directoryhoặc customSequence) dựa trên Python generator, chậm hơn tf.data vì single-threaded và không fuse với graph. Dẫn đến run time cao (CPU bottleneck) và compute waste khi scale batch lớn, không khuyến nghị cho production (TensorFlow khuyến cáo migrate sang tf.data từ 2020+). -
❌ Use Dataflow to create all possible augmentations, and store them as TFRecords.
Sai: Dataflow (Apache Beam trên GCP) dùng để pre-generate tất cả augmentations và lưu TFRecords – tốn storage khổng lồ (vì "all possible" tạo dữ liệu bùng nổ), không dynamic/random per batch, và run time chậm do I/O đọc file. Phù hợp offline ETL, không cho training real-time (vi phạm yêu cầu "randomly apply to each training batch"). -
❌ Use Dataflow to create the augmentations dynamically per training run, and stage them as TFRecords.
Sai: Dù "dynamically per training run", vẫn yêu cầu Dataflow job tạo và stage TFRecords trước khi train – gây latency cao (spin-up job ~phút), compute overhead (Dataflow workers), và không per batch (chỉ per run). Không tối ưu so với tf.data on-the-fly, dễ lỗi scale (GCP quota limits).
Kết luận 🎯: Chọn tf.data để đạt performance cao nhất, đặc biệt trên Google Cloud AI Platform/Vertex AI với TPU v5p (2025+). Tránh pre-processing với Dataflow trừ khi dataset cực lớn và augmentation deterministic!
All the information needed to compute the success metric is available in BigQuery and is updated hourly. The model is trained on eight weeks of data, on average its performance degrades below the acceptable baseline after five weeks, and training time is 12 hours. You want to ensure that the model’s performance is above the acceptable baseline while minimizing cost. How should you monitor the model to determine when retraining is necessary?
- A Use Vertex AI Model Monitoring to detect skew of the input features with a sample rate of 100% and a monitoring frequency of two days.
- B Schedule a cron job in Cloud Tasks to retrain the model every week before the newsletter is created.
- C Schedule a weekly query in BigQuery to compute the success metric.
- D Schedule a daily Dataflow job in Cloud Composer to compute the success metric.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi thuộc lĩnh vực Machine Learning Operations (MLOps) trên Google Cloud Platform (GCP), tập trung vào việc giám sát và retraining mô hình AI để đảm bảo hiệu suất ổn định với chi phí tối thiểu.
- Bối cảnh: Bạn làm việc cho nhà xuất bản trực tuyến với >50 triệu người đọc. Mô hình AI recommend nội dung cho bản tin hàng tuần.
- Metric thành công: Bài viết được mở trong vòng 2 ngày kể từ ngày xuất bản bản tin VÀ người dùng ở lại trang ít nhất 1 phút.
- Dữ liệu: Tất cả thông tin cần tính metric nằm trong BigQuery, cập nhật hàng giờ.
- Đặc tính mô hình: Train trên 8 tuần dữ liệu, hiệu suất giảm dưới baseline sau 5 tuần, thời gian train 12 giờ.
- Mục tiêu: Giám sát mô hình để phát hiện khi nào cần retrain, đảm bảo performance > baseline, tối ưu chi phí.
- Thách thức chính: Model drift nhanh (sau 5 tuần), cần monitor metric business (không phải chỉ features), dữ liệu sẵn có, tránh retrain không cần thiết (12h/train tốn kém).
Vấn đề cốt lõi: Không cần monitor skew features hay retrain định kỳ cứng nhắc, mà cần tính metric success định kỳ từ BigQuery để quyết định retrain linh hoạt, tiết kiệm chi phí. ✅ (Dựa trên best practices MLOps GCP 2024-2026: Vertex AI Pipelines & BigQuery ML cho monitoring metric-based).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Schedule a weekly query in BigQuery to compute the success metric.
Lý do:
- 🛠️ Phù hợp nhất: Metric success là business metric (open rate + dwell time), dữ liệu sẵn trong BigQuery (update hourly). Query hàng tuần cho phép tính chính xác trên dữ liệu 2 ngày + lịch sử, phát hiện drift sau ~5 tuần sớm (trước khi degrade nghiêm trọng).
- 💰 Tối ưu chi phí: BigQuery query rẻ (pay-per-query, ~$5/TB scanned), không cần infra liên tục. Chạy weekly tránh over-monitoring.
- 📈 Hiệu quả: Kết quả query trigger retrain thủ công/tự động (qua Cloud Functions/Eventarc), đảm bảo performance > baseline mà không retrain thừa (tiết kiệm 12h/train).
- 🔄 Cập nhật 2026: BigQuery hỗ trợ Scheduled Queries (UI/CLI), tích hợp BigQuery ML cho metric computation realtime-ish, phù hợp scale 50M users.
❌ Giải thích tất cả các phương án (đúng/sai)
-
SAI - Use Vertex AI Model Monitoring to detect skew of the input features with a sample rate of 100% and a monitoring frequency of two days.
Lý do sai: Vertex AI Model Monitoring (cập nhật 2026: Vertex AI Feature Store + Model Monitoring) chuyên detect feature skew/drift (thay đổi phân phối input), không phải business metric (success rate). Sample 100% + freq 2 ngày tốn kém (costly alerting, storage), không giải quyết drift prediction/output. Không minimize cost cho scale lớn. 🛑 -
SAI - Schedule a cron job in Cloud Tasks to retrain the model every week before the newsletter is created.
Lý do sai: Retrain định kỳ hàng tuần (cron via Cloud Tasks/Cloud Scheduler) không thông minh, vì model chỉ degrade sau 5 tuần → lãng phí (12h/train x 4 lần/tháng, chi phí cao). Không monitor metric thực tế, chỉ "blind retrain". Không đảm bảo performance > baseline nếu drift không đều. 💸 -
ĐÚNG - Schedule a weekly query in BigQuery to compute the success metric.
Lý do đúng: Như phần trên. 🟢 Đây là cách metric-driven monitoring chuẩn GCP MLOps: Query SQL đơn giản trên BigQuery (ví dụ:SELECT AVG(dwell_time) WHERE open_date <= publish_date + 2d), threshold baseline → alert/retrain. Tiết kiệm, scalable. -
SAI - Schedule a daily Dataflow job in Cloud Composer to compute the success metric.
Lý do sai: Dataflow (Apache Beam) + Cloud Composer (Airflow) cho batch/streaming ETL phức tạp, thừa thãi cho query đơn giản trên BigQuery (SQL native rẻ hơn). Daily → overkill (cost Dataflow vCPU/GB cao), không cần cho metric 2-day window. Composer overhead quản lý DAG. ❌
📘 Tài liệu tham khảo (GCP cập nhật 2026)
- BigQuery Scheduled Queries: docs.cloud.google.com/bigquery/docs/scheduled-queries – Hướng dẫn setup weekly metric jobs.
- Vertex AI MLOps Monitoring: cloud.google.com/vertex-ai/docs/model-monitoring/overview – Phân biệt feature vs. business metrics.
- Best Practices MLOps: cloud.google.com/architecture/ml-on-gcp-best-practices – Metric-based retraining.
- BigQuery Pricing (2026): ~$6.25/TB query → Rẻ cho 50M rows/weekly.
Hy vọng phân tích giúp bạn ôn thi Google Cloud Professional ML Engineer! 🚀
- A Train an anomaly detection model on the training dataset, and run all incoming requests through this model. If an anomaly is detected, send the most recent serving data to the labeling service.
- B Identify temporal patterns in your model’s performance over the previous year. Based on these patterns, create a schedule for sending serving data to the labeling service for the next year.
- C Compare the cost of the labeling service with the lost revenue due to model performance degradation over the past year. If the lost revenue is greater than the cost of the labeling service, increase the frequency of model retraining; otherwise, decrease the model retraining frequency.
- D Run training-serving skew detection batch jobs every few days to compare the aggregate statistics of the features in the training dataset with recent serving data. If skew is detected, send the most recent serving data to the labeling service.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi mô tả một tình huống thực tế trong MLOps trên nền tảng AWS (cụ thể liên quan đến Amazon SageMaker): Bạn đã triển khai một mô hình ML vào production cách đây một năm. Mỗi tháng, bạn thu thập toàn bộ raw requests gửi đến dịch vụ prediction của mô hình, sau đó gửi một phần nhỏ đến dịch vụ labeling thủ công (human labeling service) để đánh giá hiệu suất. Sau một năm, bạn nhận thấy hiệu suất mô hình đôi khi suy giảm mạnh chỉ sau một tháng, nhưng đôi khi mất vài tháng mới phát hiện. Dịch vụ labeling rất tốn kém, nhưng bạn cần tránh tình trạng suy giảm hiệu suất lớn (performance degradation). Mục tiêu chính: Xác định tần suất retrain mô hình sao cho duy trì hiệu suất cao mà tối ưu hóa chi phí (minimize cost).
Vấn đề cốt lõi là data drift hoặc model degradation không đều đặn, cần phương pháp proactive monitoring để phát hiện sớm thay vì chỉ dựa vào labeling định kỳ hàng tháng. Đây là thách thức điển hình trong AWS SageMaker, nơi cần sử dụng các công cụ tự động như Model Monitor để giảm chi phí labeling thủ công. (Kiến thức cập nhật đến 2026: SageMaker Model Monitor v2 hỗ trợ skew detection realtime/batch với tích hợp Lambda và CloudWatch).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng:
Run training-serving skew detection batch jobs every few days to compare the aggregate statistics of the features in the training dataset with recent serving data. If skew is detected, send the most recent serving data to the labeling service.
Lý do chọn đáp án này 🛠️:
Phương án này sử dụng training-serving skew detection (phát hiện sự lệch lạc giữa dữ liệu training và serving data) qua các batch jobs định kỳ (mỗi vài ngày), giúp phát hiện sớm data drift một cách tự động và rẻ tiền. Chỉ khi skew được detect (qua thống kê aggregate features như mean, std, quantile), mới trigger gửi serving data gần nhất đến labeling service – tránh labeling không cần thiết, giảm chi phí đáng kể. Đây là best practice của Amazon SageMaker Model Monitor (Capture và Monitoring jobs), cho phép schedule batch jobs hàng ngày/tuần, tích hợp với alerting để retrain kịp thời. Không phụ thuộc lịch cố định, mà data-driven, phù hợp với degradation không đều (đôi khi nhanh, đôi khi chậm). Tiết kiệm hơn so với labeling hàng tháng toàn bộ.
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, với ✅ đúng hoặc ❌ sai, giữ nguyên văn bản gốc tiếng Anh. Mỗi phân tích dựa trên logic MLOps AWS và hạn chế thực tế:
-
[SAI] Train an anomaly detection model on the training dataset, and run all incoming requests through this model. If an anomaly is detected, send the most recent serving data to the labeling service.
❌ Sai vì: Phương án này train một anomaly detection model riêng trên training dataset, sau đó run toàn bộ incoming requests realtime – tốn kém compute (inference overhead cao, không scale tốt cho production traffic lớn). Anomaly detection chỉ detect outliers trong serving data so với training, không trực tiếp đo lường skew giữa training-serving (có thể miss concept drift dần dần). SageMaker hỗ trợ anomaly detection nhưng không khuyến khích cho skew chính, và chạy "all requests" tăng latency/cost không cần thiết. Không giải quyết degradation chậm (vài tháng). -
[SAI] Identify temporal patterns in your model’s performance over the previous year. Based on these patterns, create a schedule for sending serving data to the labeling service for the next year.
❌ Sai vì: Phân tích temporal patterns từ dữ liệu quá khứ để tạo lịch cố định (schedule) labeling – không linh hoạt, vì degradation không theo pattern đều đặn (câu hỏi nhấn mạnh "sometimes after a month, other times several months"). Dựa lịch cố định sẽ miss drift đột ngột hoặc lãng phí nếu pattern thay đổi (non-stationary data). Labeling vẫn tốn kém định kỳ, không proactive như skew detection tự động. SageMaker Clarify chỉ hỗ trợ bias analysis, không phải temporal scheduling tự động. -
[SAI] Compare the cost of the labeling service with the lost revenue due to model performance degradation over the past year. If the lost revenue is greater than the cost of the labeling service, increase the frequency of model retraining; otherwise, decrease the model retraining frequency.
❌ Sai vì: Đây là cách cost-revenue trade-off đơn giản, nhưng reactive (dựa quá khứ), không predict/prevent degradation tương lai. Không chỉ rõ tần suất cụ thể hay cách detect (chỉ so sánh lost revenue mơ hồ), dẫn đến retrain mù quáng – có thể overtrain (tốn kém) hoặc undertrain (miss degradation). SageMaker không có built-in ROI calculator như vậy; thiếu monitoring tự động, không giải quyết vấn đề "degradation không đều". -
[ĐÚNG] Run training-serving skew detection batch jobs every few days to compare the aggregate statistics of the features in the training dataset with recent serving data. If skew is detected, send the most recent serving data to the labeling service.
✅ Đúng vì: Như giải thích ở trên – tự động, chi phí thấp, phát hiện sớm qua batch jobs (SageMaker Model Monitor: baseline từ training data, monitor serving data stats). Chỉ trigger labeling khi cần, tối ưu cost. Hỗ trợ retrain kịp thời mà không over-label.
📘 Tài liệu tham khảo
- Amazon SageMaker Model Monitor Documentation (cập nhật 2026): Training-Serving Skew Detection – Hướng dẫn batch monitoring jobs với CloudWatch alarms.
- SageMaker Best Practices for Model Monitoring (2025 update): Model Quality Monitoring – Case studies về skew detection giảm cost 70%.
- MLOps on AWS Whitepaper (2026): Nhấn mạnh proactive skew > reactive labeling.
Phương án đúng giúp xây dựng pipeline MLOps bền vững trên SageMaker! 🚀
1. Check for availability of the movie tickets at the selected cinema.
2. Assign the ticket price and accept payment.
3. Reserve the tickets at the selected cinema.
4. Send successful purchases to your database.
Each step in this process has low latency requirements (less than 50 milliseconds). You have developed a logistic regression model with BigQuery ML that predicts whether offering a promo code for free popcorn increases the chance of a ticket purchase, and this prediction should be added to the ticket purchase process. You want to identify the simplest way to deploy this model to production while adding minimal latency. What should you do?
- A Run batch inference with BigQuery ML every five minutes on each new set of tickets issued.
- B Export your model in TensorFlow format, and add a tfx_bsl.public.beam.RunInference step to the Dataflow pipeline.
- C Export your model in TensorFlow format, deploy it on Vertex AI, and query the prediction endpoint from your streaming pipeline.
- D Convert your model with TensorFlow Lite (TFLite), and add it to the mobile app so that the promo code and the incoming request arrive together in Pub/Sub.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một nền tảng bán vé xem phim cho chuỗi rạp chiếu phim lớn, nơi khách hàng sử dụng ứng dụng di động để tìm phim và mua vé. Các yêu cầu mua vé được gửi đến Pub/Sub (dịch vụ message queue của Google Cloud), sau đó được xử lý bởi Dataflow streaming pipeline với 4 bước chính:
- Kiểm tra tính sẵn có của vé phim tại rạp được chọn.
- Gán giá vé và chấp nhận thanh toán.
- Đặt chỗ vé tại rạp.
- Gửi thông tin mua vé thành công đến cơ sở dữ liệu.
📌 Yêu cầu quan trọng: Mỗi bước phải có độ trễ thấp dưới 50 milliseconds (ms). Bạn đã xây dựng mô hình logistic regression bằng BigQuery ML để dự đoán liệu việc cung cấp mã khuyến mãi popcorn miễn phí có tăng cơ hội mua vé không. Nhiệm vụ là triển khai mô hình này vào production một cách đơn giản nhất, thêm độ trễ tối thiểu vào quy trình mua vé.
🛠️ Mục tiêu chính: Tích hợp dự đoán vào pipeline streaming mà không làm chậm quy trình (low latency), ưu tiên giải pháp đơn giản và hiệu quả nhất trên Google Cloud (cập nhật đến 2026, với BigQuery ML hỗ trợ export mô hình dễ dàng sang TensorFlow).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Convert your model with TensorFlow Lite (TFLite), and add it to the mobile app so that the promo code and the incoming request arrive together in Pub/Sub.
Lý do:
- Giải pháp này đơn giản nhất vì chuyển đổi mô hình logistic regression từ BigQuery ML sang TensorFlow Lite (TFLite) – định dạng nhẹ, tối ưu cho thiết bị di động (on-device inference).
- Dự đoán được thực hiện trên app di động trước khi gửi yêu cầu đến Pub/Sub, nên không thêm độ trễ nào vào Dataflow pipeline (prediction và request đến cùng lúc). Độ trễ inference trên mobile <50ms dễ dàng đạt được với TFLite (nhẹ, không cần network call).
- Phù hợp low-latency streaming: Pipeline chỉ nhận dữ liệu đã có prediction sẵn, tránh overhead từ server-side inference.
- Cập nhật 2026: BigQuery ML hỗ trợ export trực tiếp sang TF SavedModel rồi convert TFLite qua TensorFlow tools (tf.lite.TFLiteConverter).
📘 Tài liệu tham khảo:
- BigQuery ML Export Models (Google Cloud Docs, 2026).
- TensorFlow Lite for on-device ML (TensorFlow.org).
❌ Giải thích tất cả các phương án (đúng/sai)
-
[SAI] Run batch inference with BigQuery ML every five minutes on each new set of tickets issued.
❌ Sai vì: Đây là batch processing (chạy mỗi 5 phút), không phải real-time streaming. Không đáp ứng low-latency (<50ms) – prediction chỉ có sau vài phút, làm gián đoạn quy trình mua vé ngay lập tức. BigQuery ML batch inference phù hợp phân tích lịch sử, không cho production streaming. -
[SAI] Export your model in TensorFlow format, and add a tfx_bsl.public.beam.RunInference step to the Dataflow pipeline.
❌ Sai vì: Export sang TensorFlow và thêm RunInference (từ TensorFlow Beam Serving Library) vào Dataflow sẽ thêm độ trễ đáng kể (model loading, inference trên worker nodes ~100-200ms+). Dataflow streaming overhead cao cho low-latency; không "minimal latency" so với on-device. -
[SAI] Export your model in TensorFlow format, deploy it on Vertex AI, and query the prediction endpoint from your streaming pipeline.
❌ Sai vì: Deploy lên Vertex AI Prediction (Vertex AI Endpoints) yêu cầu network call (HTTP/gRPC) từ Dataflow – độ trễ ~50-100ms+ (latency endpoint + queueing). Không đơn giản (cần deploy, scale endpoint), vi phạm yêu cầu <50ms và "simplest way". -
[ĐÚNG] Convert your model with TensorFlow Lite (TFLite), and add it to the mobile app so that the promo code and the incoming request arrive together in Pub/Sub.
✅ Đúng vì: Như giải thích ở trên – zero added latency cho pipeline, inference on-device siêu nhanh với TFLite (optimized cho mobile CPU/GPU). Đơn giản: Export BigQuery ML → Convert TFLite → Embed vào app (Flutter/Android/iOS SDK hỗ trợ). Hoàn hảo cho streaming low-latency!
🧩 Tóm tắt lợi ích giải pháp đúng: Giảm chi phí (không server), bảo mật (dữ liệu không rời device), và scale vô hạn (mỗi user tự inference). Đây là best practice cho mobile ML trên Google Cloud!
- A Train a time-series model to predict the machines’ performance values. Configure an alert if a machine’s actual performance values significantly differ from the predicted performance values.
- B Develop a simple heuristic (e.g., based on z-score) to label the machines’ historical performance data. Use this heuristic to monitor server performance in real time.
- C Develop a simple heuristic (e.g., based on z-score) to label the machines’ historical performance data. Train a model to predict anomalies based on this labeled dataset.
- D Hire a team of qualified analysts to review and label the machines’ historical performance data. Train a model based on this manually labeled dataset.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc xây dựng một giải pháp bảo trì dự đoán (predictive maintenance) cho các máy chủ trong trung tâm dữ liệu. Nhóm của bạn sử dụng dữ liệu giám sát (monitoring data) để phát hiện sớm các sự cố máy chủ tiềm ẩn. Vấn đề chính: Dữ liệu sự cố (incident data) chưa được gắn nhãn (unlabeled). Câu hỏi yêu cầu bước đầu tiên (what should you do first?) để triển khai giải pháp này một cách hiệu quả, nhanh chóng và khả thi.
Bối cảnh quan trọng:
- Predictive maintenance thường dựa vào ML để dự đoán hỏng hóc từ dữ liệu thời gian thực (như CPU, memory, temperature).
- Vì dữ liệu chưa labeled, đây là tình huống unsupervised learning hoặc rule-based anomaly detection – không thể train supervised model ngay lập tức.
- Ưu tiên: Giải pháp phải nhanh chóng triển khai, chi phí thấp, và có thể monitor real-time mà không cần label thủ công tốn kém.
- Theo kiến thức AWS/ML best practices (cập nhật đến 2026, ví dụ Amazon Lookout for Equipment hoặc SageMaker Anomaly Detection), bước đầu tiên thường là unsupervised heuristics như z-score để detect outliers trước khi scale lên ML models phức tạp. (📘 Nguồn: AWS SageMaker Documentation - Anomaly Detection (2025 update); Google Cloud ML Engineer Exam Guide - Unsupervised Techniques).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng:
Develop a simple heuristic (e.g., based on z-score) to label the machines’ historical performance data. Use this heuristic to monitor server performance in real time.
Lý do 🛠️:
- Đây là bước đầu tiên lý tưởng vì dữ liệu chưa labeled, heuristic đơn giản như z-score (đo độ lệch chuẩn từ mean) là phương pháp unsupervised anomaly detection nhanh chóng, không cần label thủ công.
- Label historical data bằng heuristic để tạo pseudo-labels, sau đó áp dụng real-time monitoring ngay lập tức – tiết kiệm thời gian, chi phí, và có thể deploy trong vài ngày.
- Phù hợp với nguyên tắc MVP (Minimum Viable Product) trong ML ops: Bắt đầu bằng rule-based trước khi train model phức tạp. AWS khuyến nghị tương tự cho industrial IoT (Amazon Monitron/Lookout).
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] Train a time-series model to predict the machines’ performance values. Configure an alert if a machine’s actual performance values significantly differ from the predicted performance values.
Phân tích: Phương án này dùng time-series forecasting (như Prophet hoặc LSTM) để predict "normal" performance và detect anomaly qua residual (sai lệch). Tuy khả thi unsupervised, nhưng KHÔNG phải bước đầu tiên vì cần train model phức tạp trên historical data – tốn thời gian (tuần/tháng) và compute, trong khi incident data chưa labeled làm model kém chính xác. Không ưu tiên "first step" nhanh chóng. -
✅ [ĐÚNG] Develop a simple heuristic (e.g., based on z-score) to label the machines’ historical performance data. Use this heuristic to monitor server performance in real time.
Phân tích: Như đã giải thích ở trên. Hoàn hảo cho first step: Z-score = (x - mean)/std_dev > threshold (e.g., 3) để flag anomaly. Dễ implement (Python/Pandas), real-time (stream qua Kinesis/SageMaker), và tạo baseline để iterate sau. (🛠️ Ưu điểm: Zero-shot deployment, low cost). -
❌ [SAI] Develop a simple heuristic (e.g., based on z-score) to label the machines’ historical performance data. Train a model to predict anomalies based on this labeled dataset.
Phân tích: Heuristic để label là tốt, nhưng train model supervised ở bước đầu là thừa thãi – tốn tài nguyên (train/test split, tuning) và chậm hơn dùng heuristic real-time trực tiếp. First step nên là monitor ngay, không phải train model (có thể làm sau khi có feedback). -
❌ [SAI] Hire a team of qualified analysts to review and label the machines’ historical performance data. Train a model based on this manually labeled dataset.
Phân tích: Label thủ công tốn kém (nhân sự, thời gian hàng tháng), không scale cho big data monitoring, và không cần thiết cho first step. AWS/Google ML best practices ưu tiên automated labeling (heuristic/self-supervised) trước human-in-loop. Chỉ dùng khi heuristic fail cao.
Kết luận 🚀: Chọn heuristic real-time để triển khai nhanh, thu thập feedback tự nhiên, rồi scale lên ML models (e.g., AWS SageMaker Canvas hoặc AutoML). (📘 Tài liệu tham khảo thêm: AWS Well-Architected ML Lens (2026); "Hands-On Machine Learning with Scikit-Learn" - Ch.9 Anomaly Detection).
- A Tokenize all of the fields using hashed dummy values to replace the real values.
- B Use principal component analysis (PCA) to reduce the four sensitive fields to one PCA vector.
- C Coarsen the data by putting AGE into quantiles and rounding LATITUDE_LONGTTUDE into single precision. The other two fields are already as coarse as possible.
- D Remove all sensitive data fields, and ask the data science team to build their models using non-sensitive data.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi này thuộc chủ đề bảo mật dữ liệu nhạy cảm (sensitive data) trong quy trình xây dựng mô hình Machine Learning (ML), đặc biệt là bảo vệ thông tin cá nhân của khách hàng (PII - Personally Identifiable Information). Bạn là nhân viên của một nhà bán lẻ quần áo toàn cầu, nhiệm vụ là đảm bảo dữ liệu được xử lý an toàn trước khi cung cấp cho đội ngũ data science dùng để huấn luyện mô hình.
Các trường dữ liệu nhạy cảm được xác định: AGE (tuổi), IS_EXISTING_CUSTOMER (khách hàng hiện tại hay không), LATITUDE_LONGITUDE (vị trí địa lý), và SHIRT_SIZE (kích cỡ áo). Những trường này có thể tiết lộ thông tin cá nhân, dẫn đến rủi ro vi phạm quy định như GDPR, CCPA hoặc AWS Data Privacy best practices.
Mục tiêu: Áp dụng kỹ thuật anonymization hoặc pseudonymization để bảo vệ dữ liệu mà vẫn giữ tính hữu ích cho ML (theo hướng dẫn AWS SageMaker Security best practices, cập nhật đến 2026 với SageMaker Pipelines và Data Protection features). Câu hỏi kiểm tra kiến thức về các phương pháp bảo mật dữ liệu trong AWS ML workflows, tránh re-identification attacks.
📘 Tài liệu tham khảo:
- AWS SageMaker Documentation: "Protecting Data Privacy in Amazon SageMaker" (https://docs.aws.amazon.com/sagemaker/latest/dg/data-protection.html).
- AWS ML Best Practices: "Data Anonymization Techniques" trong SageMaker Processing Jobs (cập nhật 2025 với hỗ trợ tokenization pipelines).
- NIST Privacy Framework (áp dụng trong AWS): Tokenization là phương pháp được khuyến nghị cho PII.
✅ Đáp án đúng
Tokenize all of the fields using hashed dummy values to replace the real values.
Lý do lựa chọn:
- Tokenization với hashed dummy values (thay thế giá trị thực bằng mã hash giả) là kỹ thuật pseudonymization chuẩn mực trên AWS, đảm bảo dữ liệu không thể đảo ngược (irreversible) mà vẫn giữ cấu trúc và utility cho training ML.
- Hashing (ví dụ SHA-256) làm mất thông tin gốc, ngăn chặn re-identification, phù hợp với tất cả 4 trường: AGE → hash tuổi; IS_EXISTING_CUSTOMER → hash boolean; LATITUDE_LONGITUDE → hash tọa độ; SHIRT_SIZE → hash kích cỡ.
- Trong AWS SageMaker (2026), hỗ trợ qua Processing Jobs hoặc Data Wrangler với tokenizers tích hợp, tuân thủ HIPAA/GDPR. Không ảnh hưởng performance model vì giữ cardinality gốc. 🛡️
🛠️ Giải thích tất cả các phương án
-
✅ Tokenize all of the fields using hashed dummy values to replace the real values.
Đúng vì: Như giải thích trên, đây là best practice AWS cho anonymization, bảo vệ toàn diện 4 trường mà không mất utility. Hỗ trợ salt hashing để tránh rainbow table attacks. (AWS SageMaker Clarify & Security Hub khuyến nghị). -
❌ Use principal component analysis (PCA) to reduce the four sensitive fields to one PCA vector.
Sai vì: PCA là kỹ thuật giảm chiều (dimensionality reduction) để giữ variance, nhưng không bảo vệ privacy - vector PCA vẫn chứa thông tin gốc có thể bị reverse-engineered (re-identification risk cao). Không phù hợp PII theo AWS guidelines; chỉ dùng cho feature engineering, không anonymization. (Ví dụ: Location data vẫn traceable). -
❌ Coarsen the data by putting AGE into quantiles and rounding LATITUDE_LONGTTUDE into single precision. The other two fields are already as coarse as possible.
Sai vì: Coarsening (làm thô dữ liệu) như quantiles cho AGE hoặc rounding LATITUDE_LONGITUDE chỉ giảm độ chính xác, không loại bỏ rủi ro privacy (vẫn identifiable qua combination, ví dụ vị trí ~1km vẫn lộ khu vực). IS_EXISTING_CUSTOMER/SHIRT_SIZE không "coarse" đủ; vi phạm k-anonymity threshold. AWS không khuyến nghị cho PII đầy đủ (chỉ perturbation nhẹ). -
❌ Remove all sensitive data fields, and ask the data science team to build their models using non-sensitive data.
Sai vì: Xóa hoàn toàn (feature removal) làm mất utility ML lớn (AGE/SHIRT_SIZE quan trọng cho recommendation models), dẫn đến model kém hiệu quả. AWS ưu tiên privacy-preserving ML (như tokenization/Federated Learning) thay vì discard data. Không scalable cho real-world retailer use cases. 🚫
- A This is not a good result because the model should have a higher accuracy for those who renew their subscription than for those who cancel their subscription.
- B This is not a good result because the model is performing worse than predicting that people will always renew their subscription.
- C This is a good result because predicting those who cancel their subscription is more difficult, since there is less data for this group.
- D This is a good result because the accuracy across both groups is greater than 80%.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi xoay quanh việc diễn giải kết quả của một mô hình Neural Network (NN) Classifier trong bài toán dự đoán khách hàng hủy đăng ký (churn prediction) cho một nhà xuất bản tạp chí. Dữ liệu thăm dò (EDA) cho thấy sự mất cân bằng lớp (class imbalance) nghiêm trọng: 90% khách hàng gia hạn (renew) và chỉ 10% hủy (cancel).
Mô hình sau khi huấn luyện đạt:
- 99% accuracy khi dự đoán những người hủy đăng ký (lớp thiểu số - minority class).
- 82% accuracy khi dự đoán những người gia hạn (lớp đa số - majority class).
"Accuracy" ở đây ám chỉ per-class accuracy (có thể là recall hoặc precision cho từng lớp riêng lẻ, nhưng thường trong ngữ cảnh imbalance được hiểu là tỷ lệ dự đoán đúng cho lớp đó). Câu hỏi yêu cầu diễn giải kết quả này một cách hợp lý, xem xét vấn đề imbalance dữ liệu – một thách thức phổ biến trong ML thực tế (theo kiến thức cập nhật AWS SageMaker đến 2026, nơi khuyến nghị sử dụng baseline models và metrics như F1-score, AUC-ROC thay vì accuracy thuần).
Tính toán weighted accuracy của mô hình (giả sử recall per class):
(0.9 × 82%) + (0.1 × 99%) = 83.7% – thấp hơn baseline đơn giản luôn dự đoán "renew" đạt 90%.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: "This is not a good result because the model is performing worse than predicting that people will always renew their subscription."
🛠️ Lý do chi tiết:
Trong bài toán imbalanced classification, baseline model đơn giản nhất là luôn dự đoán lớp đa số ("renew"), đạt 90% accuracy tổng. Mô hình NN này chỉ đạt weighted accuracy ~83.7%, thấp hơn baseline, chứng tỏ chưa tốt (có thể overfit lớp thiểu số hoặc không học được pattern thực). AWS (SageMaker Clarify & Model Monitor, phiên bản 2026) nhấn mạnh luôn so sánh với baseline để tránh misleading metrics như accuracy. Đây là bài học cốt lõi trong AWS Certified Machine Learning - Specialty (MLS-C01).
📋 Giải thích tất cả các phương án
-
❌ "This is not a good result because the model should have a higher accuracy for those who renew their subscription than for those who cancel their subscription."
Sai vì với dữ liệu imbalance (90/10), lớp thiểu số (cancel) khó dự đoán hơn do ít sample, dễ overfit/underfit. Không có quy tắc cứng nhắc "majority phải acc cao hơn minority". Model có acc renew 82% (thấp) nhưng cancel 99% (cao bất thường) cho thấy vấn đề khác, không phải lý do này. AWS docs khuyên dùng stratified metrics thay vì so acc trực tiếp. -
✅ "This is not a good result because the model is performing worse than predicting that people will always renew their subscription."
Đúng (như đã giải thích ở trên). Baseline "always renew" đạt 90% total acc, model chỉ ~83.7% – underperform rõ rệt. Đây là dấu hiệu cần cải thiện (resampling, class weights, hoặc metrics như Precision-Recall AUC). Phù hợp best practice AWS SageMaker (Hyperparameter Tuning với imbalance handling). -
❌ "This is a good result because predicting those who cancel their subscription is more difficult, since there is less data for this group."
Sai dù cancel khó hơn do ít data (10%), nhưng acc tổng thể thấp hơn baseline chứng tỏ model chưa hiệu quả. Acc 99% cho minority có thể là "ảo" (high recall nhưng low precision, dẫn FP cao). Không nên coi là "good" mà không check total performance hoặc business metrics (ví dụ recall@minority quan trọng cho churn, nhưng vẫn cần baseline). -
❌ "This is a good result because the accuracy across both groups is greater than 80%."
Sai vì ngưỡng 80% arbitrary, không có cơ sở khoa học. Với imbalance, accuracy >80% vẫn kém baseline 90%. AWS khuyến cáo tránh accuracy thuần cho imbalance, dùng F1-minority hoặc ROC-AUC (theo SageMaker Model Evaluation tools 2026).
📘 Tài liệu tham khảo
- AWS SageMaker Documentation (2026): "Handling Imbalanced Datasets" – https://docs.aws.amazon.com/sagemaker/latest/dg/imbalanced-data.html (nhấn baseline comparison).
- AWS Certified Machine Learning - Specialty Exam Guide (MLS-C01 v2.0+): Domain 3: Modeling (Exploratory Data Analysis & Model Evaluation).
- Google Cloud ML Best Practices (tương đương, vì role): "Imbalanced datasets" trong Vertex AI – https://cloud.google.com/vertex-ai/docs/general/ml-best-practices.
- Sách: "Evaluating Machine Learning Models" by Alice Zheng (O'Reilly, cập nhật 2024).
🔍 Khuyến nghị: Trong thực tế AWS/GCP, dùng SMOTE/resampling hoặc class_weight='balanced' trong XGBoost/NN để fix imbalance!
- A Remove the data transformation step from your pipeline.
- B Containerize the PySpark transformation step, and add it to your pipeline.
- C Add a ContainerOp to your pipeline that spins a Dataproc cluster, runs a transformation, and then saves the transformed data in Cloud Storage.
- D Deploy Apache Spark at a separate node pool in a Google Kubernetes Engine cluster. Add a ContainerOp to your pipeline that invokes a corresponding transformation job for this Spark instance.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi này thuộc chủ đề Machine Learning Pipelines trên Google Cloud Platform (GCP), cụ thể là cách tích hợp quy trình tiền xử lý dữ liệu (preprocessing) vào Kubeflow Pipelines để parametrize (tham số hóa) toàn bộ pipeline huấn luyện mô hình.
- Bối cảnh: Bạn đã xây dựng mô hình huấn luyện trên dữ liệu Parquet lưu trong Hive table trên Google Cloud (có lẽ là Dataproc Metastore hoặc tương tự). Dữ liệu được truy cập qua Hive, sau đó tiền xử lý bằng PySpark và xuất ra file CSV lưu vào Cloud Storage. Tiếp theo là các bước huấn luyện và đánh giá mô hình.
- Yêu cầu chính: Làm cho toàn bộ quy trình (preprocessing + train + evaluate) có thể tham số hóa trong Kubeflow Pipelines, nghĩa là pipeline phải linh hoạt, có thể chạy lại với các tham số khác nhau (như kích thước dữ liệu, hyperparameters).
- Thách thức: PySpark yêu cầu môi trường Spark phân tán để xử lý dữ liệu lớn hiệu quả. Kubeflow Pipelines chạy trên Kubernetes (GKE), nên cần tích hợp ContainerOp (Kubernetes operator) để thực thi bước transformation mà không làm phức tạp pipeline.
- Phiên bản cập nhật (đến 2026): Kubeflow 2.x+ hỗ trợ tốt Dataproc operators qua Kubeflow Pipelines SDK (kfp.google), tích hợp mượt mà với Cloud Dataproc serverless (ra mắt 2023-2024), cho phép spin cluster tạm thời mà không cần quản lý thủ công.
📘 Tài liệu tham khảo:
- Kubeflow Pipelines SDK - Google Operators
- Dataproc Serverless for Spark Jobs
- GCP ML Engineer Exam Guide
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Add a ContainerOp to your pipeline that spins a Dataproc cluster, runs a transformation, and then saves the transformed data in Cloud Storage.
Lý do 🛠️:
- Đây là cách tối ưu và được khuyến nghị trong Kubeflow trên GCP. Sử dụng DataprocCreateClusterOp và DataprocSubmitJobOp (qua kfp.google) để tự động tạo cluster Dataproc tạm thời, submit PySpark job (transformation), lưu CSV vào Cloud Storage, rồi xóa cluster.
- Ưu điểm: Scalable cho big data, chi phí thấp (pay-per-use với serverless), tích hợp IAM tự động, dễ parametrize (cluster size, job params qua pipeline arguments). Không cần quản lý infra thủ công, phù hợp production ML workflows.
- Hoàn hảo cho dữ liệu Hive/Parquet lớn, vì Dataproc hỗ trợ Hive Metastore native.
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] Remove the data transformation step from your pipeline.
Giải thích sai: Việc loại bỏ bước transformation sẽ phá hủy pipeline, vì dữ liệu gốc (Parquet từ Hive) cần được tiền xử lý bằng PySpark thành CSV trước khi train/evaluate. Kubeflow Pipelines yêu cầu toàn bộ workflow phải được encapsulate, không thể bỏ qua preprocessing – điều này vi phạm nguyên tắc reproducibility và parametrization. -
❌ [SAI] Containerize the PySpark transformation step, and add it to your pipeline.
Giải thích sai: Containerize PySpark đơn thuần (Docker image với Spark) chỉ phù hợp small data trên single node, không scalable cho dữ liệu lớn từ Hive (cần distributed computing). Spark yêu cầu cluster manager (YARN/Kubernetes), container đơn lẻ sẽ OOM hoặc chậm. Kubeflow khuyến cáo dùng managed services như Dataproc thay vì tự containerize phức tạp. -
✅ [ĐÚNG] Add a ContainerOp to your pipeline that spins a Dataproc cluster, runs a transformation, and then saves the transformed data in Cloud Storage.
Giải thích đúng (như phần trên): 🏆 Cách chuẩn GCP, sử dụng Kubeflow Google Operators để automate Dataproc lifecycle. Ví dụ code:dataproc_create = dataproc.create_cluster(...); submit_pyspark_job(dataproc_create.outputs.cluster_name, ...);. Tiết kiệm 70-80% thời gian setup so với manual. -
❌ [SAI] Deploy Apache Spark at a separate node pool in a Google Kubernetes Engine cluster. Add a ContainerOp to your pipeline that invokes a corresponding transformation job for this Spark instance.
Giải thích sai: Triển khai Spark standalone trên GKE node pool riêng rất phức tạp và tốn kém (cần config Spark-on-K8s, autoscaling, persistent volumes). Không linh hoạt như Dataproc (không serverless, quản lý thủ công cluster), dễ bottleneck GKE quota. GCP docs khuyên dùng Dataproc cho Spark jobs trong ML pipelines thay vì tự host trên GKE.
- A Deploy the models to a Vertex AI endpoint using the traffic-split=0=80, PREVIOUS_MODEL_ID=20 configuration.
- B Wrap the models inside an App Engine application using the --splits PREVIOUS_VERSION=0.2, NEW_VERSION=0.8 configuration
- C Wrap the models inside a Cloud Run container using the REVISION1=20, REVISION2=80 revision configuration.
- D Implement random splitting in Dataflow using beam.Partition() with a partition function calling a Vertex AI endpoint.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi này xoay quanh việc triển khai một mô hình Machine Learning (ML) để phát hiện cảm xúc (sentiment) từ các bài đăng của người dùng trên trang mạng xã hội của công ty, nhằm xác định các sự cố như outages hoặc bugs.
- Quy trình hiện tại: Sử dụng Dataflow để thực hiện dự đoán thời gian thực (real-time predictions) trên dữ liệu được ingest từ Pub/Sub (một hệ thống messaging streaming của Google Cloud).
- Yêu cầu chính:
- Có nhiều lần huấn luyện (training iterations) cho mô hình.
- Giữ hai phiên bản mới nhất (latest two versions) luôn live sau mỗi lần chạy.
- Phân bổ traffic giữa hai phiên bản theo tỷ lệ 80:20, với phiên bản mới nhất nhận 80% traffic.
- Pipeline phải đơn giản nhất có thể, với quản lý tối thiểu (minimal management).
Mục tiêu là chọn giải pháp tích hợp tốt với Dataflow/Pub/Sub, hỗ trợ model versioning, traffic splitting tự động, và không cần code phức tạp. Đây là kịch bản điển hình trong Vertex AI cho serving ML models trong môi trường production với A/B testing hoặc canary deployments. (Kiến thức cập nhật đến 2026: Vertex AI tiếp tục hỗ trợ traffic splitting linh hoạt qua endpoints, tích hợp seamless với Dataflow cho online predictions.)
📘 Tài liệu tham khảo:
- Vertex AI Endpoints - Traffic Splitting
- Dataflow with Vertex AI Batch/Online Predictions
- Pub/Sub to Dataflow Streaming Pipeline
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Deploy the models to a Vertex AI endpoint using the traffic-split=0=80, PREVIOUS_MODEL_ID=20 configuration.
Lý do:
- 🛠️ Vertex AI endpoints là giải pháp native và đơn giản nhất cho ML serving trên Google Cloud, hỗ trợ deploy multiple model versions đến cùng một endpoint.
- Traffic splitting được cấu hình chính xác qua tham số
traffic-split(ví dụ:0=80nghĩa là version mới nhất nhận 80%,PREVIOUS_MODEL_ID=20cho version cũ 20%). - Tích hợp hoàn hảo với Dataflow: Dataflow có thể gọi online predictions đến Vertex AI endpoint qua Apache Beam transforms (như
InvokeVertexAIPrediction). - Minimal management: Vertex AI tự động quản lý scaling, versioning, monitoring; giữ đúng hai versions live, và dễ update sau mỗi training iteration.
- Đáp ứng real-time từ Pub/Sub mà không cần wrapper code phức tạp.
📋 Giải thích tất cả các phương án (đúng/sai)
-
✅ Deploy the models to a Vertex AI endpoint using the traffic-split=0=80, PREVIOUS_MODEL_ID=20 configuration.
Đúng vì: Như giải thích ở trên, đây là cách chuẩn và tối ưu cho traffic splitting giữa model versions trong Vertex AI. Cú pháptraffic-splitchính xác theo API mới nhất (2026), dễ integrate với Dataflow cho streaming predictions từ Pub/Sub. Không cần quản lý thủ công, auto-scaling và monitoring built-in. 🏆 -
❌ Wrap the models inside an App Engine application using the --splits PREVIOUS_VERSION=0.2, NEW_VERSION=0.8 configuration.
Sai vì: App Engine dùng cho web apps, không phải ML serving realtime.--splitschỉ hỗ trợ version splitting cho app instances, không native cho ML models. Việc wrap models vào App Engine sẽ phức tạp hóa pipeline (cần custom code cho predictions), không tích hợp tốt với Dataflow/Pub/Sub, và quản lý versions thủ công hơn Vertex AI. 🚫 -
❌ Wrap the models inside a Cloud Run container using the REVISION1=20, REVISION2=80 revision configuration.
Sai vì: Cloud Run dùng cho containerized services, hỗ trợ revision traffic splitting nhưng dành cho services, không phải ML models chuyên dụng. Phải wrap models thủ công (custom container với TensorFlow Serving hoặc tương tự), dẫn đến management overhead cao (scaling, monitoring models riêng). Không đơn giản bằng Vertex AI, và kém optimize cho Dataflow streaming. 🛑 -
❌ Implement random splitting in Dataflow using beam.Partition() with a partition function calling a Vertex AI endpoint.
Sai vì:beam.Partition()là low-level transform trong Dataflow, yêu cầu code custom phức tạp để random split traffic và gọi Vertex AI. Vi phạm yêu cầu pipeline đơn giản, minimal management (phải maintain partition logic, handle errors, scaling). Không tự động quản lý two versions live, dễ lỗi trong production streaming. 😵
Kết luận: Giải pháp Vertex AI là best practice cho MLOps trên Google Cloud, đảm bảo scalability và simplicity! 🚀
- A Create a Google Kubernetes Engine cluster with a node pool that has 4 V100 GPUs. Prepare and submit a TFJob operator to this node pool.
- B Create a Vertex AI Workbench user-managed notebooks instance with 4 V100 GPUs, and use it to train your model.
- C Package your code with Setuptools, and use a pre-built container. Train your model with Vertex AI using a custom tier that contains the required GPUs.
- D Configure a Compute Engine VM with all the dependencies that launches the training. Train your model with Vertex AI using a custom tier that contains the required GPUs.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm
📖 Nội dung câu hỏi:
Câu hỏi xoay quanh việc phát triển một mô hình nhận diện hình ảnh (image recognition) sử dụng framework PyTorch dựa trên kiến trúc ResNet50. Code đã chạy tốt trên laptop cá nhân với subsample nhỏ, nhưng dataset đầy đủ có 200k ảnh đã gắn nhãn. Mục tiêu là scale nhanh chóng workload huấn luyện (training) trong khi tối ưu hóa chi phí, sử dụng đúng 4 GPU V100.
🛤️ Yêu cầu chính: Chọn giải pháp phù hợp trên Google Cloud (Vertex AI ecosystem) để huấn luyện phân tán (distributed training) hiệu quả, hỗ trợ PyTorch, tự động scale và quản lý tài nguyên GPU mà không lãng phí.
Lưu ý cập nhật 2026: Vertex AI (trước đây là AI Platform) đã hỗ trợ mạnh mẽ custom containers cho PyTorch với distributed training qua Horovod hoặc PyTorch DDP, tích hợp Accelerator Tier tùy chỉnh cho V100/A100/H100 GPUs (theo docs Vertex AI v2026.01).
✅ Đáp án đúng:
Package your code with Setuptools, and use a pre-built container. Train your model with Vertex AI using a custom tier that contains the required GPUs.
Lý do chọn: Giải pháp này tối ưu nhất vì sử dụng pre-built container (như gcr.io/cloud-aiplatform/training/pytorch-gpu.1-13 hoặc tương tự từ NGC/PyTorch hub), đóng gói code bằng Setuptools để dễ deploy. Vertex AI Custom Training Job với custom tier (scale lên 4 V100 GPUs) tự động xử lý distributed training (multi-GPU/node), checkpointing, logging (TensorBoard), và preemptible VMs để giảm chi phí ~70%. Phù hợp PyTorch, scale nhanh mà không cần quản lý infra thủ công.
📘 Nguồn: Vertex AI Custom Training Docs & PyTorch Containers (cập nhật 2026).
🛠️ Giải thích tất cả các phương án (Đúng/Sai)
-
❌ [SAI] Create a Google Kubernetes Engine cluster with a node pool that has 4 V100 GPUs. Prepare and submit a TFJob operator to this node pool.
Lý do sai: TFJob chỉ dành cho TensorFlow (Kubernetes operator chuyên biệt), không hỗ trợ PyTorch native. GKE yêu cầu tự quản lý cluster, Kubeflow pipelines phức tạp, tốn thời gian setup (không "quickly scale") và chi phí cao hơn do luôn-on nodes. Không tối ưu cho single-job training như Vertex AI. -
❌ [SAI] Create a Vertex AI Workbench user-managed notebooks instance with 4 V100 GPUs, and use it to train your model.
Lý do sai: Vertex AI Workbench (Managed Notebooks) phù hợp dev/test nhỏ, nhưng với 200k images và 4 V100, nó không scale distributed tự động (chỉ single-instance), dễ hết RAM/disk, và user-managed nghĩa là tự attach GPU + deps (PyTorch, DDP). Chi phí cao vì notebooks chạy liên tục, không auto-scale/shutdown như Training Jobs. Không "minimizing cost". -
✅ [ĐÚNG] Package your code with Setuptools, and use a pre-built container. Train your model with Vertex AI using a custom tier that contains the required GPUs.
Lý do đúng: Như đã giải thích ở trên – pre-built container (PyTorch-optimized) + Setuptools cho entrypoint script đơn giản. Vertex AI Custom Tier (e.g.,ACCELERATOR_TYPE=NVIDIA_TESLA_V100,ACCELERATOR_COUNT=4) tự động provision VMs, multi-replica training, hyperparameter tuning, và spot/preemptible để tiết kiệm. Scale nhanh (submit job <5 phút), theo dõi qua Console/CLI. Hoàn hảo cho workload này!
📘 Nguồn: Vertex AI GPU Training (2026 updates hỗ trợ V100 legacy). -
❌ [SAI] Configure a Compute Engine VM with all the dependencies that launches the training. Train your model with Vertex AI using a custom tier that contains the required GPUs.
Lý do sai: Compute Engine VM yêu cầu tự install tất cả deps (PyTorch, CUDA, NCCL cho multi-GPU), phức tạp và lỗi-prone. Kết hợp với Vertex AI Custom Tier nghe hay nhưng VM riêng lẻ không integrate tốt với Vertex AI (phải dùng container thay vì raw VM). Tốn thời gian setup, không auto-scale/distributed, chi phí cao hơn do manual management.
🎯 Kết luận: Chọn đáp án đúng giúp scale nhanh + tiết kiệm nhất trên Vertex AI, phù hợp best practices Google Cloud ML (2026). Nếu deploy thực tế, dùng lệnh gcloud ai custom-jobs create với YAML config! 🚀