Ngân hàng đề — AWS Certified Machine Learning Specialty

Tìm thấy 371 câu.

Câu 301
A company deployed a machine learning (ML) model on the company website to predict real estate prices. Several months after deployment, an ML engineer notices that the accuracy of the model has gradually decreased.

The ML engineer needs to improve the accuracy of the model. The engineer also needs to receive notifications for any future performance issues.

Which solution will meet these requirements?
  1. A Perform incremental training to update the model. Activate Amazon SageMaker Model Monitor to detect model performance issues and to send notifications.
  2. B Use Amazon SageMaker Model Governance. Configure Model Governance to automatically adjust model hyperparameters. Create a performance threshold alarm in Amazon CloudWatch to send notifications.
  3. C Use Amazon SageMaker Debugger with appropriate thresholds. Configure Debugger to send Amazon CloudWatch alarms to alert the team. Retrain the model by using only data from the previous several months.
  4. D Use only data from the previous several months to perform incremental training to update the model. Use Amazon SageMaker Model Monitor to detect model performance issues and to send notifications.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong AWS SageMaker: Một công ty đã triển khai mô hình Machine Learning (ML) trên website để dự đoán giá bất động sản. Sau vài tháng, độ chính xác của mô hình giảm dần (model drift – hiện tượng phổ biến do dữ liệu mới thay đổi theo thời gian, như biến động thị trường bất động sản).

Nhiệm vụ của ML engineer là:

  • Cải thiện độ chính xác: Cần cập nhật mô hình với dữ liệu mới mà không mất kiến thức cũ.
  • Nhận thông báo cho vấn đề tương lai: Giám sát liên tục và cảnh báo tự động.

🛠️ Vấn đề cốt lõi: Model drift (data drift hoặc concept drift) cần giải pháp incremental training (huấn luyện tăng dần) kết hợp giám sát tự động như Amazon SageMaker Model Monitor (phiên bản mới nhất 2026 hỗ trợ monitoring baseline, drift detection, và tích hợp CloudWatch/SNS cho alerts).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Perform incremental training to update the model. Activate Amazon SageMaker Model Monitor to detect model performance issues and to send notifications.

Lý do:

  • Incremental training (huấn luyện tăng dần) cho phép cập nhật mô hình với dữ liệu mới mà vẫn giữ trọng số cũ, tránh train từ đầu (tiết kiệm chi phí và thời gian). SageMaker hỗ trợ native cho hầu hết algorithms (như XGBoost, Linear Learner) từ phiên bản 2023+.
  • Amazon SageMaker Model Monitor (cập nhật 2026: hỗ trợ JSONLines, CSV, Parquet; detect data quality, bias, drift; baseline từ training data) giám sát production traffic, phát hiện accuracy drop, và gửi notifications qua CloudWatch Alarms + SNS/Email.
  • Giải pháp toàn diện, hiệu quả nhất, phù hợp best practices AWS ML lifecycle (MLOps).

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc:

  • ✅ Perform incremental training to update the model. Activate Amazon SageMaker Model Monitor to detect model performance issues and to send notifications.
    Đúng vì: Như giải thích trên, incremental training khắc phục drift hiệu quả, Model Monitor là công cụ chuẩn cho post-deployment monitoring (viết reports, alerts tự động). Hoàn hảo cho yêu cầu.

  • ❌ Use Amazon SageMaker Model Governance. Configure Model Governance to automatically adjust model hyperparameters. Create a performance threshold alarm in Amazon CloudWatch to send notifications.
    Sai vì: Không tồn tại service "Amazon SageMaker Model Governance" (cập nhật 2026: SageMaker có Model Registry, Clarify cho bias, nhưng không tự động adjust hyperparameters). Hyperparameter tuning dùng SageMaker Hyperparameter Optimization (HPO), không phải governance. CloudWatch alarm ok nhưng không giải quyết root cause (drift).

  • ❌ Use Amazon SageMaker Debugger with appropriate thresholds. Configure Debugger to send Amazon CloudWatch alarms to alert the team. Retrain the model by using only data from the previous several months.
    Sai vì: SageMaker Debugger chỉ dùng cho debug quá trình training (monitor tensors, gradients), không phải monitoring production inference. Retrain chỉ data recent gây catastrophic forgetting (mô hình quên kiến thức cũ, accuracy tệ hơn với data lịch sử). Không phù hợp production monitoring.

  • ❌ Use only data from the previous several months to perform incremental training to update the model. Use Amazon SageMaker Model Monitor to detect model performance issues and to send notifications.
    Sai vì: Dù Model Monitor đúng, nhưng incremental training chỉ dùng data recent vi phạm nguyên tắc (phải kết hợp data cũ + mới để tránh bias). SageMaker incremental training yêu cầu data mới phù hợp schema cũ, không nên loại bỏ historical data.

📘 Tài liệu tham khảo (cập nhật AWS 2026)

🛠️ Lời khuyên: Áp dụng SageMaker Pipelines cho full MLOps để automate retraining + monitoring!

Câu 302 Chọn nhiều đáp án
A university wants to develop a targeted recruitment strategy to increase new student enrollment. A data scientist gathers information about the academic performance history of students. The data scientist wants to use the data to build student profiles. The university will use the profiles to direct resources to recruit students who are likely to enroll in the university.

Which combination of steps should the data scientist take to predict whether a particular student applicant is likely to enroll in the university? (Choose two.)
  1. A Use Amazon SageMaker Ground Truth to sort the data into two groups named "enrolled" or "not enrolled."
  2. B Use a forecasting algorithm to run predictions.
  3. C Use a regression algorithm to run predictions.
  4. D Use a classification algorithm to run predictions.
  5. E Use the built-in Amazon SageMaker k-means algorithm to cluster the data into two groups named "enrolled" or "not enrolled."
Xem giải thích

🧠 Phân tích câu hỏi trắc nghiệm AWS SageMaker bởi AWS Certified DevOps Engineer Professional

Chào bạn! 🛠️ Tôi là chuyên gia AWS Certified DevOps Engineer Professional với kiến thức cập nhật đến năm 2026 (dựa trên các tính năng mới nhất của Amazon SageMaker như SageMaker Studio 2.0, hỗ trợ mô hình ML tiên tiến và tích hợp Ground Truth Plus với active learning). Hãy cùng phân tích kỹ lưỡng câu hỏi này về Machine Learning trên AWS để xây dựng mô hình dự đoán khả năng đăng ký của sinh viên (enrollment prediction).

🧩 1. Giải thích nội dung câu hỏi một cách chi tiết

Câu hỏi mô tả tình huống: Một trường đại học muốn tăng tuyển sinh mới bằng cách sử dụng dữ liệu lịch sử về hiệu suất học tập của sinh viên để xây dựng profile sinh viên. Mục tiêu là dự đoán xem một ứng viên cụ thể có khả năng đăng ký (enroll) vào trường hay không, từ đó hướng nguồn lực tuyển sinh hiệu quả.

  • Đây là bài toán supervised learning kiểu binary classification (phân loại nhị phân): Input là dữ liệu lịch sử (features như điểm số, lịch sử học tập), Output là nhãn enrolled (đăng ký) hoặc not enrolled (không đăng ký).
  • Data scientist cần chuẩn bị dữ liệu (labeling) và chạy mô hình dự đoán. Câu hỏi yêu cầu chọn TWO steps phù hợp nhất trong SageMaker để thực hiện.
    SageMaker là dịch vụ ML end-to-end của AWS, hỗ trợ từ data labeling (Ground Truth) đến training/deploy models (bao gồm classification algorithms).

✅ 2. Đáp án đúng và lý do lựa chọn

Hai đáp án đúng là:

  • Use Amazon SageMaker Ground Truth to sort the data into two groups named "enrolled" or "not enrolled."
  • Use a classification algorithm to run predictions.

Lý do chọn (chi tiết):
✅ SageMaker Ground Truth là công cụ labeling dữ liệu chính xác cho supervised learning. Nó giúp "sort" (phân loại/đánh nhãn) dữ liệu lịch sử thành hai nhóm enrolled/not enrolled một cách hiệu quả, hỗ trợ human-in-the-loop, active learning (tính năng mới 2024-2026), và tích hợp trực tiếp với SageMaker training jobs. Không có labeling, không thể train mô hình supervised!
✅ Classification algorithm phù hợp hoàn hảo cho binary prediction (enroll hay không), vì output là discrete classes (nhị phân). SageMaker built-in algorithms như XGBoost, Linear Learner hỗ trợ binary classification với metrics như accuracy, F1-score (cập nhật SageMaker 2026 với AutoML cho classification nhanh hơn).
Kết hợp hai bước: Labeling trước (Ground Truth) → Train & predict sau (classification) → Hoàn hảo cho pipeline ML!

🔍 3. Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một, giữ nguyên nội dung tiếng Anh gốc. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm giải thích chi tiết bằng tiếng Việt dựa trên best practices AWS SageMaker mới nhất.

  • ✅ Use Amazon SageMaker Ground Truth to sort the data into two groups named "enrolled" or "not enrolled."
    Đúng! 🏆 Ground Truth là dịch vụ labeling dữ liệu hàng đầu của SageMaker, chuyên "sort/group" dữ liệu thành nhãn supervised (như binary labels enrolled/not enrolled). Nó hỗ trợ workforce management, quality checks, và tích hợp S3/SageMaker Pipelines (cập nhật 2025 với Ground Truth Plus cho auto-labeling). Bước đầu tiên cần thiết cho bất kỳ mô hình classification nào.

  • ❌ Use a forecasting algorithm to run predictions.
    Sai! 🚫 Forecasting (như DeepAR/Prophet trong SageMaker) dùng cho time-series prediction (dự đoán xu hướng thời gian liên tục, ví dụ doanh số tương lai). Bài toán này không có yếu tố thời gian, mà là classification tĩnh dựa trên profile sinh viên → Không phù hợp, sẽ cho kết quả sai lệch.

  • ❌ Use a regression algorithm to run predictions.
    Sai! 📉 Regression (như Linear Learner/XGBoost regression trong SageMaker) dự đoán giá trị liên tục (continuous, ví dụ dự đoán điểm GPA = 3.5). Nhưng output ở đây là nhị phân (enroll/không) → Sẽ không chính xác, metrics như RMSE không phù hợp cho classification.

  • ✅ Use a classification algorithm to run predictions.
    Đúng! 🎯 Đây là lựa chọn cốt lõi cho bài toán binary/multiclass classification. SageMaker hỗ trợ hàng loạt algorithms như XGBoost Classifier, BlazingText, Object2Vec (cập nhật 2026 với JumpStart Models cho pre-trained classifiers). Train trên dữ liệu đã label → Predict probability enroll (threshold 0.5).

  • ❌ Use the built-in Amazon SageMaker k-means algorithm to cluster the data into two groups named "enrolled" or "not enrolled."
    Sai! 🤔 K-means là unsupervised clustering (không cần labels, tự group dữ liệu thành k clusters dựa trên similarity). Dù có thể force 2 clusters, nó không đảm bảo clusters khớp với "enrolled/not enrolled" (có thể lẫn lộn), thiếu interpretability và không dùng dữ liệu lịch sử có nhãn. Phải dùng supervised cho prediction chính xác!

📘 6. Tài liệu tham khảo (AWS Official Docs - cập nhật 2026)

Kết luận: 🎉 Kết hợp Ground Truth + Classification là pipeline chuẩn AWS cho bài toán này. Nếu cần code sample hoặc pipeline CDK/Serverless, hỏi tôi nhé! 🚀

Câu 303 Chọn nhiều đáp án
A machine learning (ML) specialist is using the Amazon SageMaker DeepAR forecasting algorithm to train a model on CPU-based Amazon EC2 On-Demand instances. The model currently takes multiple hours to train. The ML specialist wants to decrease the training time of the model.

Which approaches will meet this requirement? (Choose two.)
  1. A Replace On-Demand Instances with Spot Instances.
  2. B Configure model auto scaling dynamically to adjust the number of instances automatically.
  3. C Replace CPU-based EC2 instances with GPU-based EC2 instances.
  4. D Use multiple training instances.
  5. E Use a pre-trained version of the model. Run incremental training.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc tối ưu hóa thời gian huấn luyện (training time) mô hình Machine Learning sử dụng thuật toán Amazon SageMaker DeepAR – một thuật toán dự báo thời gian (time-series forecasting) dựa trên mạng nơ-ron hồi quy (RNN).

  • Bối cảnh: Chuyên gia ML đang huấn luyện mô hình trên EC2 On-Demand instances dựa trên CPU, và quá trình này mất nhiều giờ (multiple hours). Mục tiêu là giảm thời gian huấn luyện.
  • Yêu cầu chọn 2 phương án đúng từ các lựa chọn, phù hợp với SageMaker training (không phải inference).
  • Kiến thức cập nhật 2026: Theo tài liệu AWS SageMaker mới nhất (phiên bản 2024-2026), DeepAR hỗ trợ GPU acceleration và distributed training trên nhiều instances để tăng tốc độ tính toán song song, đặc biệt hiệu quả cho dữ liệu lớn. Không có thay đổi lớn về hỗ trợ Spot cho training time reduction trực tiếp.

📘 Tài liệu tham khảo:

✅ Đáp án đúng (Chọn 2)

Hai phương án đúng là những cách tăng tốc độ tính toán trực tiếp cho quá trình training DeepAR:

  1. Replace CPU-based EC2 instances with GPU-based EC2 instances: ✅ Thay CPU bằng GPU (như ml.p3 hoặc ml.g5 instances) giúp tăng tốc đáng kể vì DeepAR sử dụng mạng nơ-ron sâu, tận dụng GPU cho phép tính toán song song matrix nhanh hơn hàng chục lần so với CPU.
  2. Use multiple training instances: ✅ Sử dụng nhiều instances (distributed training) cho phép phân phối dữ liệu và tính toán song song qua SageMaker's built-in sharding, giảm thời gian từ linear scaling (ví dụ: 4 instances giảm ~75% thời gian).

🛠️ Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn một cách rõ ràng, giữ nguyên văn bản gốc tiếng Anh:

  • Replace On-Demand Instances with Spot Instances.
    ❌ Sai: Spot Instances chỉ giúp giảm chi phí bằng cách sử dụng dung lượng dư thừa, nhưng không giảm thời gian training. Chúng có thể bị gián đoạn (interruption), thậm chí làm tăng thời gian tổng thể nếu checkpoint không hoàn hảo. DeepAR hỗ trợ Spot nhưng ưu tiên cho cost-saving, không phải speed.

  • Configure model auto scaling dynamically to adjust the number of instances automatically.
    ❌ Sai: Auto scaling dành cho inference endpoints (hosting), không áp dụng cho training jobs. Training là job batch một lần, không dynamic scale như serving.

  • Replace CPU-based EC2 instances with GPU-based EC2 instances.
    ✅ Đúng: DeepAR hỗ trợ GPU instances (như ml.p4d, ml.g5), và GPU vượt trội cho workload neural network training nhờ CUDA cores. Thời gian giảm từ hours xuống minutes tùy scale dữ liệu (AWS benchmarks: lên đến 10x faster).

  • Use multiple training instances.
    ✅ Đúng: SageMaker hỗ trợ distributed training cho DeepAR với parameter distribution (e.g., smdistributed dataparallel), phân phối batch data qua nhiều instances, đạt near-linear speedup (scale-out).

  • Use a pre-trained version of the model. Run incremental training.
    ❌ Sai: DeepAR không có pre-trained models công khai từ AWS (khác với BlazingText hoặc built-in như Forecast). Incremental training (warm start) dùng cho fine-tuning nhưng không giảm thời gian training ban đầu từ scratch; nó chỉ tiếp tục từ checkpoint, không giải quyết vấn đề "multiple hours" cho full train.

Câu 304
A chemical company has developed several machine learning (ML) solutions to identify chemical process abnormalities. The time series values of independent variables and the labels are available for the past 2 years and are sufficient to accurately model the problem.

The regular operation label is marked as 0 The abnormal operation label is marked as 1. Process abnormalities have a significant negative effect on the company’s profits. The company must avoid these abnormalities.

Which metrics will indicate an ML solution that will provide the GREATEST probability of detecting an abnormality?
  1. A Precision = 0.91 -
    Recall = 0.6
  2. B Precision = 0.61 -
    Recall = 0.98
  3. C Precision = 0.7 -
    Recall = 0.9
  4. D Precision = 0.98 -
    Recall = 0.8
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi thuộc chủ đề Machine Learning trên AWS, cụ thể liên quan đến việc chọn metrics đánh giá model ML để phát hiện bất thường (abnormality) trong quy trình hóa học của một công ty. Dữ liệu thời gian thực (time series) từ 2 năm trước bao gồm các biến độc lập và nhãn (label): 0 cho hoạt động bình thường và 1 cho hoạt động bất thường. Bất thường gây thiệt hại lớn về lợi nhuận, nên công ty phải ưu tiên tránh bỏ lỡ bất kỳ trường hợp bất thường nào (tức là giảm thiểu False Negative - FN).

Mục tiêu là chọn bộ metrics (Precision và Recall) cho thấy model có xác suất cao nhất phát hiện bất thường (tương đương label 1). Đây là bài toán binary classification với dataset có thể imbalanced (bất thường hiếm hơn bình thường), thường dùng trên Amazon SageMaker để train/deploy model. Theo best practices AWS (cập nhật đến 2026, SageMaker phiên bản mới nhất hỗ trợ AutoML và metrics tùy chỉnh), ưu tiên Recall cao để đảm bảo detect hết positive cases (bất thường), chấp nhận false positive nếu cần.

✅ Đáp án đúng: Precision = 0.61 - Recall = 0.98

Lý do chọn đáp án này: Trong ngữ cảnh phát hiện bất thường (anomaly detection), Recall (độ nhạy - sensitivity) là metric quan trọng nhất vì nó đo lường tỷ lệ True Positive (TP) / (TP + FN) – tức là khả năng model phát hiện đúng tất cả các trường hợp bất thường thực tế. Với Recall = 0.98 (98%), model chỉ bỏ lỡ 2% bất thường, mang lại xác suất cao nhất tránh thiệt hại kinh doanh. Precision = 0.61 (thấp) nghĩa là có nhiều false alarm (báo động giả), nhưng điều này chấp nhận được vì chi phí xử lý false positive thấp hơn false negative (bỏ lỡ bất thường gây lỗ lớn). Các lựa chọn khác có Recall thấp hơn, không đảm bảo detect tối đa bất thường. Đây phù hợp với AWS ML best practices cho high-stakes detection (như SageMaker Clarify và Model Monitor).

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng phương án, dựa trên công thức:

  • Precision = TP / (TP + FP): Tỷ lệ dự đoán bất thường đúng (giảm false alarm).

  • Recall = TP / (TP + FN): Tỷ lệ phát hiện hết bất thường thực tế (ưu tiên ở đây). Sử dụng F1-score tham chiếu (harmonic mean của Precision & Recall) để so sánh, nhưng ưu tiên Recall cao nhất.

  • ❌ Precision = 0.91 - Recall = 0.6
    Phương án này sai vì Recall chỉ 0.6 (60%), nghĩa là model bỏ lỡ 40% bất thường thực tế (FN cao), dẫn đến rủi ro kinh doanh lớn dù Precision cao (ít false alarm). Không phù hợp mục tiêu "GREATEST probability of detecting an abnormality". F1-score ≈ 0.74 (thấp).

  • ✅ Precision = 0.61 - Recall = 0.98
    Phương án đúng như đã giải thích: Recall cao nhất (0.98), đảm bảo detect gần như toàn bộ bất thường, ưu tiên tránh thiệt hại. Precision thấp chấp nhận được. F1-score ≈ 0.76 (cao nhất trong các lựa chọn về detection).

  • ❌ Precision = 0.7 - Recall = 0.9
    Phương án này sai vì Recall = 0.9 (90%) vẫn bỏ lỡ 10% bất thường, kém hơn 0.98. Precision cân bằng hơn nhưng không bù đắp được rủi ro FN. F1-score ≈ 0.79 (tốt nhưng Recall chưa tối ưu).

  • ❌ Precision = 0.98 - Recall = 0.8
    Phương án này sai vì Recall chỉ 0.8 (80%), bỏ lỡ 20% bất thường – rủi ro cao nhất dù Precision rất cao (ít false alarm). Ưu tiên Precision quá mức, không phù hợp với yêu cầu detect tối đa. F1-score ≈ 0.88 (cao nhưng thiên về accuracy chung).

🛠️ Tài liệu tham khảo

  • 📘 AWS Documentation (SageMaker, cập nhật 2026): Evaluating Machine Learning Models – Nhấn mạnh Recall cho anomaly detection trong imbalanced data.
  • 📘 AWS ML Specialty Exam Guide (DOP-C02 & MLS-C01): Metrics prioritization cho high-cost FN scenarios.
  • 🧪 SageMaker Built-in Algorithms: Time-series forecasting/anomaly detection (DeepAR, Random Cut Forest) ưu tiên Recall metrics.
  • 🔗 Blog AWS: "Handling Imbalanced Data in Amazon SageMaker" (aws.amazon.com/blogs/machine-learning) – Best practices đến 2026.

Hy vọng phân tích này giúp bạn ôn thi AWS hiệu quả! 🚀 Nếu cần thêm ví dụ code SageMaker, hãy hỏi nhé.

Câu 305
An online delivery company wants to choose the fastest courier for each delivery at the moment an order is placed. The company wants to implement this feature for existing users and new users of its application. Data scientists have trained separate models with XGBoost for this purpose, and the models are stored in Amazon S3. There is one model for each city where the company operates.

Operation engineers are hosting these models in Amazon EC2 for responding to the web client requests, with one instance for each model, but the instances have only a 5% utilization in CPU and memory. The operation engineers want to avoid managing unnecessary resources.

Which solution will enable the company to achieve its goal with the LEAST operational overhead?
  1. A Create an Amazon SageMaker notebook instance for pulling all the models from Amazon S3 using the boto3 library. Remove the existing instances and use the notebook to perform a SageMaker batch transform for performing inferences offline for all the possible users in all the cities. Store the results in different files in Amazon S3. Point the web client to the files.
  2. B Prepare an Amazon SageMaker Docker container based on the open-source multi-model server. Remove the existing instances and create a multi-model endpoint in SageMaker instead, pointing to the S3 bucket containing all the models. Invoke the endpoint from the web client at runtime, specifying the TargetModel parameter according to the city of each request.
  3. C Keep only a single EC2 instance for hosting all the models. Install a model server in the instance and load each model by pulling it from Amazon S3. Integrate the instance with the web client using Amazon API Gateway for responding to the requests in real time, specifying the target resource according to the city of each request.
  4. D Prepare a Docker container based on the prebuilt images in Amazon SageMaker. Replace the existing instances with separate SageMaker endpoints, one for each city where the company operates. Invoke the endpoints from the web client, specifying the URL and EndpointName parameter according to the city of each request.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty giao hàng trực tuyến muốn tối ưu hóa việc chọn courier nhanh nhất cho mỗi đơn hàng ngay khi khách hàng đặt hàng. Họ đã huấn luyện các mô hình XGBoost riêng biệt cho từng thành phố (mỗi model lưu trữ trong Amazon S3), và cần hỗ trợ cả người dùng hiện tại lẫn mới.

Hiện tại, các kỹ sư vận hành đang host các model này trên Amazon EC2 với một instance riêng cho mỗi model, dẫn đến tình trạng utilization CPU và memory chỉ 5% – tức là lãng phí tài nguyên nghiêm trọng. Mục tiêu là giảm thiểu operational overhead (chi phí quản lý, bảo trì instance thủ công như scaling, patching, monitoring) trong khi vẫn đảm bảo inference thời gian thực từ web client.

Yêu cầu cốt lõi:

  • 🚀 Real-time inference: Phải invoke model ngay khi có request từ web client, dựa trên thành phố của đơn hàng.
  • 💰 Tiết kiệm tài nguyên: Tránh chạy nhiều instance riêng lẻ, tận dụng sharing.
  • 🛠️ Least operational overhead: Giải pháp managed service, tự động scale, không cần quản lý server.

Kiến thức AWS cập nhật đến 2026: Amazon SageMaker là dịch vụ ML managed hàng đầu, hỗ trợ Multi-Model Endpoints (MME) với Multi-Model Server (MMS) open-source, cho phép host nhiều model trên một endpoint duy nhất, chỉ load model khi cần (cold-start nhanh), auto-scaling, và tích hợp S3 seamless. EC2 yêu cầu quản lý thủ công cao hơn.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Prepare an Amazon SageMaker Docker container based on the open-source multi-model server. Remove the existing instances and create a multi-model endpoint in SageMaker instead, pointing to the S3 bucket containing all the models. Invoke the endpoint from the web client at runtime, specifying the TargetModel parameter according to the city of each request.

Lý do chọn 🏆:

  • Giải pháp này sử dụng SageMaker Multi-Model Endpoint (MME) với open-source MMS (dựa trên Docker container), cho phép host tất cả model từ một S3 bucket trên MỘT endpoint duy nhất.
  • Model chỉ load on-demand (khi invoke với TargetModel parameter chỉ định tên model/thành phố), giảm lãng phí tài nguyên xuống mức tối thiểu (không idle như EC2).
  • Least operational overhead: SageMaker managed hoàn toàn (auto-scaling, monitoring via CloudWatch, security, patching), xóa instance EC2 cũ. Web client invoke real-time qua API với tham số TargetModel.
  • Phù hợp real-time inference cho mọi user, scale theo traffic. Theo AWS 2026, MME hỗ trợ XGBoost native và MMS v1.5+ tối ưu latency <100ms.

🔍 Phân tích tất cả các phương án (đúng/sai)

Dưới đây là phân tích chi tiết từng lựa chọn. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, đánh dấu ✅/❌, và giải thích hoàn toàn bằng tiếng Việt với lý do dựa trên best practices AWS.

  • Phương án 1 ❌:
    Create an Amazon SageMaker notebook instance for pulling all the models from Amazon S3 using the boto3 library. Remove the existing instances and use the notebook to perform a SageMaker batch transform for performing inferences offline for all the possible users in all the cities. Store the results in different files in Amazon S3. Point the web client to the files.
    Tại sao sai? 🛑: SageMaker Notebook chỉ dùng cho development/exploration, KHÔNG phù hợp production real-time inference. Batch Transform là offline processing (chạy batch trước, lưu kết quả S3), không đáp ứng "at the moment an order is placed" (real-time). Web client phải query file S3 thay vì invoke model – chậm, không scale cho user mới/dynamic data, và overhead cao vì cần schedule batch thường xuyên. Overhead lớn vì notebook không managed cho inference.

  • Phương án 2 ✅:
    Prepare an Amazon SageMaker Docker container based on the open-source multi-model server. Remove the existing instances and create a multi-model endpoint in SageMaker instead, pointing to the S3 bucket containing all the models. Invoke the endpoint from the web client at runtime, specifying the TargetModel parameter according to the city of each request.
    Tại sao đúng? 🎯: Như đã giải thích ở phần đáp án đúng. Đây là best practice AWS cho multi-model low-utilization: MMS + MME tiết kiệm 70-90% chi phí so EC2, zero-management, hỗ trợ XGBoost/S3 direct. TargetModel enable dynamic routing theo city. Cập nhật 2026: SageMaker Inference Recommender recommend MME cho case này.

  • Phương án 3 ❌:
    Keep only a single EC2 instance for hosting all the models. Install a model server in the instance and load each model by pulling it from Amazon S3. Integrate the instance with the web client using Amazon API Gateway for responding to the requests in real time, specifying the target resource according to the city of each request.
    Tại sao sai? ⚠️: Vẫn dùng EC2 tự quản lý (cài model server thủ công, pull S3), chỉ consolidate 1 instance nhưng operational overhead cao: Phải handle scaling, patching OS, monitoring, high availability (ALB/ASG), và load tất cả model có thể gây memory bloat nếu nhiều city. API Gateway chỉ là proxy, không giải quyết root vấn đề idle resources. Không "least overhead" so với SageMaker managed.

  • Phương án 4 ❌:
    Prepare a Docker container based on the prebuilt images in Amazon SageMaker. Replace the existing instances with separate SageMaker endpoints, one for each city where the company operates. Invoke the endpoints from the web client, specifying the URL and EndpointName parameter according to the city of each request.
    Tại sao sai? 🔄: Tạo nhiều endpoint riêng (một/city) tương tự EC2 cũ – vẫn lãng phí (mỗi endpoint provision instance riêng, idle 95%). Overhead cao vì quản lý nhiều endpoint (deploy/update từng cái, cost cao theo số city). SageMaker prebuilt images tốt nhưng không optimize multi-model. Web client cần logic routing URL phức tạp, không efficient như single endpoint + TargetModel.

📘 Tài liệu tham khảo AWS (cập nhật 2026)

Giải pháp này giúp công ty scale toàn cầu với chi phí thấp nhất! 🚀 Nếu cần demo code boto3, hỏi thêm nhé!

Câu 306
A company builds computer-vision models that use deep learning for the autonomous vehicle industry. A machine learning (ML) specialist uses an Amazon EC2 instance that has a CPU:GPU ratio of 12:1 to train the models.

The ML specialist examines the instance metric logs and notices that the GPU is idle half of the time. The ML specialist must reduce training costs without increasing the duration of the training jobs.

Which solution will meet these requirements?
  1. A Switch to an instance type that has only CPUs.
  2. B Use a heterogeneous cluster that has two different instances groups.
  3. C Use memory-optimized EC2 Spot Instances for the training jobs.
  4. D Switch to an instance type that has a CPU:GPU ratio of 6:1.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty phát triển mô hình computer-vision sử dụng deep learning cho ngành xe tự lái (autonomous vehicle). Một chuyên gia ML đang sử dụng Amazon EC2 instance với tỷ lệ CPU:GPU = 12:1 để huấn luyện (train) các mô hình này.

Từ log metrics, chuyên gia nhận thấy GPU đang idle (không hoạt động) 50% thời gian, nghĩa là GPU bị lãng phí một nửa thời lượng do chờ đợi các tác vụ khác (thường là CPU xử lý dữ liệu đầu vào như data preprocessing, loading data từ storage).

Yêu cầu chính: Giảm chi phí huấn luyện (training costs) mà KHÔNG tăng thời gian huấn luyện (duration of training jobs). Điều này đòi hỏi giải pháp phải tối ưu hóa việc sử dụng tài nguyên hiện tại, làm GPU bận rộn hơn mà không kéo dài tổng thời gian train.

🛠️ Vấn đề cốt lõi: Trong deep learning cho computer-vision (như object detection, segmentation), GPU xử lý compute-intensive tasks nhanh, nhưng CPU thường bottleneck ở data pipeline (preprocessing images/videos). Tỷ lệ CPU:GPU 12:1 nghĩa là quá nhiều CPU so với GPU → GPU chờ CPU → idle cao → lãng phí tiền (vì instance tính phí theo giờ sử dụng).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Switch to an instance type that has a CPU:GPU ratio of 6:1.

Lý do:

  • Giảm tỷ lệ CPU:GPU từ 12:1 xuống 6:1 sẽ cân bằng tải giữa CPU và GPU. CPU ít hơn → giảm thời gian chờ của GPU (idle từ 50% xuống thấp hơn) → GPU hoạt động hiệu quả hơn, tăng throughput training mà không tăng tổng thời gian (thậm chí có thể giảm nhẹ do parallelism tốt hơn).
  • Giảm chi phí: Instance với ít vCPU hơn thường rẻ hơn (ví dụ: AWS p4d.24xlarge có 96 vCPU:8 GPU = 12:1, chi phí cao; chuyển sang loại như p4de hoặc g5.48xlarge với tỷ lệ cân bằng hơn ~6:1, giá thấp hơn do ít CPU). Không cần thay đổi code hoặc kiến trúc.
  • Phù hợp phiên bản AWS 2026: AWS khuyến nghị chọn instance GPU với tỷ lệ CPU phù hợp workload (AWS Deep Learning AMIs và SageMaker hỗ trợ tuning này).

📋 Phân tích tất cả các phương án (đúng/sai)

  • Phương án SAI: Switch to an instance type that has only CPUs.
    ❌ Lý do sai: Deep learning computer-vision cần GPU mạnh (như NVIDIA A100/H100) cho matrix multiplications và convolutions. Chỉ CPU sẽ làm training chậm gấp 10-100 lần, tăng duration (vi phạm yêu cầu). Chi phí có thể rẻ hơn ban đầu nhưng tổng chi phí cao hơn do thời gian dài. Không giải quyết GPU idle vì loại bỏ GPU hoàn toàn.

  • Phương án SAI: Use a heterogeneous cluster that has two different instances groups.
    ❌ Lý do sai: Heterogeneous cluster (như EC2 với instance groups khác nhau trong EMR hoặc EKS) phức tạp, cần thay đổi code (phân phối workload qua NCCL/MPI), dễ gây imbalance và overhead communication. Không trực tiếp giảm GPU idle trên single instance, có thể tăng duration do sync giữa nodes. Không tối ưu cho single-job training trên EC2.

  • Phương án SAI: Use memory-optimized EC2 Spot Instances for the training jobs.
    ❌ Lý do sai: Memory-optimized (R-family như r7g) tập trung RAM cao, không có GPU mạnh phù hợp deep learning (thiếu NVIDIA Tesla/A100). Spot Instances rẻ (giảm 90% chi phí) nhưng rủi ro bị interrupt cao ở training dài (checkpointing phức tạp), dễ tăng duration nếu restart. Không giải quyết GPU idle vì sai loại instance.

  • Phương án ĐÚNG: Switch to an instance type that has a CPU:GPU ratio of 6:1.
    ✅ Lý do đúng (như đã giải thích ở trên): Tối ưu bottleneck CPU → GPU busy 100%, giảm lãng phí → chi phí thấp hơn mà duration không tăng. Dễ implement chỉ bằng thay instance type trong launch template.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

  • AWS EC2 Instance Types Guide: https://aws.amazon.com/ec2/instance-types/ (p4d/p5/g5/p5 series với tỷ lệ CPU:GPU được optimize cho ML training).
  • AWS Deep Learning Best Practices: https://docs.aws.amazon.com/sagemaker/latest/dg/train-ml-best-practices.html (phần "Choosing the Right Instance Type" nhấn mạnh balance CPU/GPU cho data loading).
  • Amazon SageMaker/EC2 ML Workloads Whitepaper (2025 update): Khuyến nghị monitor GPU utilization via CloudWatch, điều chỉnh ratio cho vision models.
  • Nguồn metrics: CloudWatch GPU metrics (GPUUtilization, CPUUtilization) để detect idle.

🛠️ Lời khuyên DevOps: Sử dụng AWS Compute Optimizer để recommend instance, kết hợp Spot + Savings Plans cho scale, và TensorBoard hoặc SageMaker Debugger theo dõi real-time!

Câu 307
A company wants to forecast the daily price of newly launched products based on 3 years of data for older product prices, sales, and rebates. The time-series data has irregular timestamps and is missing some values.

Data scientist must build a dataset to replace the missing values. The data scientist needs a solution that resamples the data daily and exports the data for further modeling.

Which solution will meet these requirements with the LEAST implementation effort?
  1. A Use Amazon EMR Serverless with PySpark.
  2. B Use AWS Glue DataBrew.
  3. C Use Amazon SageMaker Studio Data Wrangler.
  4. D Use Amazon SageMaker Studio Notebook with Pandas.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc xử lý dữ liệu time-series cho mục đích dự báo giá sản phẩm hàng ngày. Cụ thể:

  • Dữ liệu đầu vào: 3 năm dữ liệu về giá sản phẩm cũ, doanh số (sales) và chiết khấu (rebates), với timestamps không đều (irregular) và giá trị thiếu (missing values).
  • Yêu cầu chính: Data scientist cần xây dựng dataset mới bằng cách thay thế missing values (imputation), resample dữ liệu theo ngày (daily resampling), rồi export để sử dụng trong modeling (dự báo).
  • Tiêu chí chọn giải pháp: LEAST implementation effort (ít nỗ lực triển khai nhất), nghĩa là ưu tiên công cụ no-code/low-code, trực quan, không cần viết code phức tạp.
    Đây là bài toán data preparation điển hình trong AWS SageMaker ecosystem, phù hợp với các tool xử lý ETL/ETL cho ML với dữ liệu time-series. (Kiến thức cập nhật AWS 2024-2026: SageMaker Data Wrangler hỗ trợ mạnh mẽ time-series transforms như resample, forward-fill/back-fill cho missing values).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Amazon SageMaker Studio Data Wrangler.

🛠️ Lý do chi tiết:

  • SageMaker Data Wrangler là công cụ visual, no-code/low-code chuyên biệt cho data preparation trong SageMaker Studio, được thiết kế để xử lý nhanh các bước như imputation missing values (sử dụng methods như mean/median/forward-fill), resample time-series theo daily (tích hợp Pandas-like transforms với giao diện drag-and-drop), và export trực tiếp sang CSV/Parquet/SageMaker Feature Store cho modeling.
  • Least effort: Người dùng chỉ cần import data → chọn flow template → apply transforms visually → export (1-2 click), không cần code. Hỗ trợ preview real-time và tích hợp liền mạch với SageMaker Pipelines cho production.
  • So với các option khác, nó giảm effort xuống mức tối thiểu cho data scientists.

📘 Tài liệu tham khảo:

  • AWS Docs: Amazon SageMaker Data Wrangler (cập nhật 2025: hỗ trợ time-series resampling v2.0+).
  • AWS Blog: "Prepare Time-Series Data with Data Wrangler" (2024).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên effort triển khai và khả năng đáp ứng yêu cầu (resample daily + imputation + export).

  • Use Amazon EMR Serverless with PySpark.
    ❌ Sai: EMR Serverless mạnh cho big data Spark jobs, nhưng yêu cầu viết code PySpark thủ công (custom UDFs cho resample với window functions và imputation via fillna), setup clusterless job, và export qua S3. Effort cao (code + debug), không visual, phù hợp scale lớn chứ không phải quick prep.

  • Use AWS Glue DataBrew.
    ❌ Sai: DataBrew là visual data prep tool (no-code recipes cho cleaning/imputation), hỗ trợ resample cơ bản nhưng không chuyên sâu cho irregular time-series (thiếu Pandas-level resample như asfreq('D') hoặc interpolation advanced). Export OK nhưng cần custom steps nhiều hơn, và ít tích hợp ML so với SageMaker. Effort trung bình-cao cho time-series phức tạp.

  • Use Amazon SageMaker Studio Data Wrangler.
    ✅ Đúng (như đã giải thích ở trên): Visual flow hoàn hảo cho yêu cầu, zero-code cho resample + imputation, export 1-click. Least effort nhất!

  • Use Amazon SageMaker Studio Notebook with Pandas.
    ❌ Sai: Notebook với Pandas linh hoạt (dùng pd.resample('D') + interpolate()), nhưng yêu cầu viết code đầy đủ (import, handle timezone, loop imputation), debug manual, và export thủ công. Effort cao hơn Data Wrangler (cùng Studio nhưng thiếu visual automation). Phù hợp custom logic chứ không least effort.

🧠 Kết luận nổi bật: SageMaker Data Wrangler là "one-stop-shop" cho data wrangling time-series với least effort, giúp data scientist focus vào modeling thay vì code ETL. Nếu scale production, tích hợp trực tiếp SageMaker Processing Jobs! 🚀

Câu 308
A data scientist is building a forecasting model for a retail company by using the most recent 5 years of sales records that are stored in a data warehouse. The dataset contains sales records for each of the company’s stores across five commercial regions. The data scientist creates a working dataset with StoreID. Region. Date, and Sales Amount as columns. The data scientist wants to analyze yearly average sales for each region. The scientist also wants to compare how each region performed compared to average sales across all commercial regions.

Which visualization will help the data scientist better understand the data trend?
  1. A Create an aggregated dataset by using the Pandas GroupBy function to get average sales for each year for each store. Create a bar plot, faceted by year, of average sales for each store. Add an extra bar in each facet to represent average sales.
  2. B Create an aggregated dataset by using the Pandas GroupBy function to get average sales for each year for each store. Create a bar plot, colored by region and faceted by year, of average sales for each store. Add a horizontal line in each facet to represent average sales.
  3. C Create an aggregated dataset by using the Pandas GroupBy function to get average sales for each year for each region. Create a bar plot of average sales for each region. Add an extra bar in each facet to represent average sales.
  4. D Create an aggregated dataset by using the Pandas GroupBy function to get average sales for each year for each region. Create a bar plot, faceted by year, of average sales for each region. Add a horizontal line in each facet to represent average sales.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc chọn visualization phù hợp nhất để một nhà khoa học dữ liệu (data scientist) phân tích dữ liệu bán hàng từ data warehouse (có thể là Amazon Redshift hoặc Amazon EMR trên AWS). Dataset bao gồm các cột: StoreID (ID cửa hàng), Region (5 vùng thương mại), Date (ngày), và Sales Amount (số tiền bán hàng). Mục tiêu chính là:

  • Phân tích average sales hàng năm cho từng region (🛤️ trend theo năm và vùng).
  • So sánh hiệu suất từng region so với average sales toàn bộ các regions (📊 so sánh tương đối).

Visualization cần giúp hiểu rõ trend dữ liệu (như sự biến động theo năm và so sánh dễ dàng). Đây là kỹ năng cốt lõi trong AWS SageMaker Data Wrangler hoặc Amazon QuickSight (cập nhật đến 2026, hỗ trợ Pandas integration và faceting trong QuickSight ML insights). Sử dụng Pandas GroupBy để aggregate dữ liệu trước khi plot (thường với Matplotlib/Seaborn).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng:
Create an aggregated dataset by using the Pandas GroupBy function to get average sales for each year for each region. Create a bar plot, faceted by year, of average sales for each region. Add a horizontal line in each facet to represent average sales.

Lý do chọn (chi tiết):

  • 🛠️ Aggregate đúng mức độ: GroupBy theo year và region → tính avg sales per region per year, phù hợp chính xác với yêu cầu "yearly average sales for each region".
  • 📈 Bar plot faceted by year: Mỗi facet (ô nhỏ) là một năm, hiển thị bar cho từng region → dễ thấy trend theo năm và so sánh giữa regions trong cùng năm.
  • ➡️ Horizontal line cho average sales: Đại diện overall average across all regions, giúp so sánh trực quan (region nào trên/dưới đường kẻ → outperform/underperform).
  • 🎯 Tối ưu hiểu trend: Faceting theo năm tránh overcrowd chart, line reference dễ nhận diện pattern (ví dụ: region nào tăng trưởng tốt hơn average). Đây là best practice trong AWS QuickSight (SPICE engine hỗ trợ faceting đến 2026) và Pandas plotting.

📘 Tài liệu tham khảo:

🔍 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên mức độ aggregate, loại plot, và khả năng so sánh trend.

  • ❌ Phương án SAI:
    Create an aggregated dataset by using the Pandas GroupBy function to get average sales for each year for each store. Create a bar plot, faceted by year, of average sales for each store. Add an extra bar in each facet to represent average sales.
    Giải thích sai: Aggregate theo store (không phải region) → quá chi tiết (hàng trăm store?), không tập trung vào "yearly average per region". Bar per store faceted by year gây overcrowd (quá nhiều bar), khó so sánh region. Extra bar cho average không hiệu quả bằng line (bar cạnh tranh không gian).

  • ❌ Phương án SAI:
    Create an aggregated dataset by using the Pandas GroupBy function to get average sales for each year for each store. Create a bar plot, colored by region and faceted by year, of average sales for each store. Add a horizontal line in each facet to represent average sales.
    Giải thích sai: Vẫn aggregate theo store → không khớp yêu cầu region-level. Color by region trên bar per store làm chart lộn xộn (nhiều store cùng region chồng color), faceted by year vẫn khó tổng hợp trend region. Line tốt nhưng aggregate sai gốc rễ.

  • ❌ Phương án SAI:
    Create an aggregated dataset by using the Pandas GroupBy function to get average sales for each year for each region. Create a bar plot of average sales for each region. Add an extra bar in each facet to represent average sales.
    Giải thích sai: Aggregate đúng (year + region) nhưng không facet by year → chỉ một plot duy nhất gộp tất cả năm, mất trend theo thời gian (khó thấy yearly change). "Extra bar in each facet" vô nghĩa vì không có facet, gây nhầm lẫn và kém trực quan.

  • ✅ Phương án ĐÚNG (đã phân tích chi tiết ở trên):
    Create an aggregated dataset by using the Pandas GroupBy function to get average sales for each year for each region. Create a bar plot, faceted by year, of average sales for each region. Add a horizontal line in each facet to represent average sales.
    Giải thích đúng (tóm tắt): Hoàn hảo khớp yêu cầu: region-level avg, yearly trend qua faceting, so sánh dễ dàng với reference line.

💡 Lời khuyên thực hành trên AWS: Sử dụng SageMaker Studio để code Pandas + plot, export sang QuickSight cho dashboard. Test với sample data để verify trend! 🚀

Câu 309
A company uses sensors on devices such as motor engines and factory machines to measure parameters, temperature and pressure. The company wants to use the sensor data to predict equipment malfunctions and reduce services outages.

Machine learning (ML) specialist needs to gather the sensors data to train a model to predict device malfunctions. The ML specialist must ensure that the data does not contain outliers before training the model.

How can the ML specialist meet these requirements with the LEAST operational overhead?
  1. A Load the data into an Amazon SageMaker Studio notebook. Calculate the first and third quartile. Use a SageMaker Data Wrangler data flow to remove only values that are outside of those quartiles.
  2. B Use an Amazon SageMaker Data Wrangler bias report to find outliers in the dataset. Use a Data Wrangler data flow to remove outliers based on the bias report.
  3. C Use an Amazon SageMaker Data Wrangler anomaly detection visualization to find outliers in the dataset. Add a transformation to a Data Wrangler data flow to remove outliers.
  4. D Use Amazon Lookout for Equipment to find and remove outliers from the dataset.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào một công ty sử dụng dữ liệu từ cảm biến (sensors) trên các thiết bị như động cơ và máy móc nhà máy để đo lường các thông số như nhiệt độ và áp suất. Mục tiêu là sử dụng dữ liệu này để dự đoán sự cố thiết bị (equipment malfunctions) và giảm thời gian gián đoạn dịch vụ (service outages).
Chuyên gia Machine Learning (ML specialist) cần thu thập dữ liệu cảm biến để huấn luyện mô hình dự đoán sự cố, nhưng phải đảm bảo dữ liệu không chứa outliers (giá trị ngoại lai) trước khi training.
Yêu cầu chính: Thực hiện với LEAST operational overhead (ít nỗ lực vận hành nhất), nghĩa là ưu tiên giải pháp tự động hóa cao, dễ sử dụng, không cần code thủ công phức tạp.
📘 Bối cảnh AWS cập nhật đến 2026: SageMaker Data Wrangler (phiên bản mới nhất tích hợp mạnh mẽ các công cụ visualization và transformation tự động cho data prep, bao gồm anomaly detection để xử lý outliers mà không cần code nhiều).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use an Amazon SageMaker Data Wrangler anomaly detection visualization to find outliers in the dataset. Add a transformation to a Data Wrangler data flow to remove outliers.

Lý do chọn:
🛠️ SageMaker Data Wrangler cung cấp anomaly detection visualization tích hợp sẵn (dựa trên thuật toán unsupervised ML như isolation forest hoặc statistical methods), giúp tự động phát hiện outliers qua giao diện trực quan chỉ với vài cú click. Sau đó, thêm transformation vào data flow để loại bỏ outliers một cách liền mạch, không cần code thủ công hay tính toán thủ công.
✅ Đây là giải pháp least operational overhead vì toàn bộ quy trình nằm trong Data Wrangler (no-code/low-code), phù hợp cho data prep trước training model trên SageMaker. Không yêu cầu kiến thức sâu về ML để implement.
Nguồn tham khảo: AWS SageMaker Data Wrangler Documentation - Anomaly Detection (cập nhật 2025-2026 với enhancements cho IoT/sensor data).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích lý do đúng/sai bằng tiếng Việt:

  • ❌ Phương án SAI: Load the data into an Amazon SageMaker Studio notebook. Calculate the first and third quartile. Use a SageMaker Data Wrangler data flow to remove only values that are outside of those quartiles.
    Giải thích sai: Phương án này yêu cầu tính toán thủ công Q1 và Q3 (IQR method) trong SageMaker Studio notebook (dùng Python/Pandas), sau đó mới apply Data Wrangler. Điều này tạo operational overhead cao vì cần code, debug và chuyển dữ liệu giữa notebook/Data Wrangler. Không tự động, không phải least effort. Data Wrangler có cách tốt hơn qua visualization.

  • ❌ Phương án SAI: Use an Amazon SageMaker Data Wrangler bias report to find outliers in the dataset. Use a Data Wrangler data flow to remove outliers based on the bias report.
    Giải thích sai: Bias report trong Data Wrangler dùng để phát hiện bias trong dữ liệu (như gender/race bias), không phải để detect outliers (giá trị ngoại lai thống kê). Sử dụng sai công cụ dẫn đến kết quả không chính xác, và vẫn cần manual interpretation. Không phải giải pháp chuẩn cho outliers, overhead cao hơn anomaly detection.

  • ✅ Phương án ĐÚNG (đã giải thích ở trên): Use an Amazon SageMaker Data Wrangler anomaly detection visualization to find outliers in the dataset. Add a transformation to a Data Wrangler data flow to remove outliers.
    Giải thích đúng: Tích hợp sẵn, trực quan, tự động hóa cao nhất cho sensor data prep. Least overhead!

  • ❌ Phương án SAI: Use Amazon Lookout for Equipment to find and remove outliers from the dataset.
    Giải thích sai: Amazon Lookout for Equipment là dịch vụ predictive maintenance chuyên biệt cho equipment/sensor data, detect anomalies để dự đoán failures chứ không dùng để remove outliers từ dataset trước training. Nó không export/transform data dễ dàng cho SageMaker training, yêu cầu setup model riêng (overhead cao: labeling, training detector). Không phù hợp cho data prep đơn giản.
    Nguồn tham khảo: AWS Lookout for Equipment Docs (2026: focus on inference, không thay thế data cleaning).

🧩 Kết luận: Giải pháp đúng tận dụng sức mạnh no-code của SageMaker Data Wrangler, lý tưởng cho ML specialist xử lý IoT/sensor data với ít effort nhất! Nếu apply thực tế, bắt đầu từ Data Wrangler trong SageMaker Studio. 🚀

Câu 310
A data scientist obtains a tabular dataset that contains 150 correlated features with different ranges to build a regression model. The data scientist needs to achieve more efficient model training by implementing a solution that minimizes impact on the model’s performance. The data scientist decides to perform a principal component analysis (PCA) preprocessing step to reduce the number of features to a smaller set of independent features before the data scientist uses the new features in the regression model.

Which preprocessing step will meet these requirements?
  1. A Use the Amazon SageMaker built-in algorithm for PCA on the dataset to transform the data.
  2. B Load the data into Amazon SageMaker Data Wrangler. Scale the data with a Min Max Scaler transformation step. Use the SageMaker built-in algorithm for PCA on the scaled dataset to transform the data.
  3. C Reduce the dimensionality of the dataset by removing the features that have the highest correlation. Load the data into Amazon SageMaker Data Wrangler. Perform a Standard Scaler transformation step to scale the data. Use the SageMaker built-in algorithm for PCA on the scaled dataset to transform the data.
  4. D Reduce the dimensionality of the dataset by removing the features that have the lowest correlation. Load the data into Amazon SageMaker Data Wrangler. Perform a Min Max Scaler transformation step to scale the data. Use the SageMaker built-in algorithm for PCA on the scaled dataset to transform the data.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc xử lý dữ liệu tiền xử lý (preprocessing) cho một bộ dữ liệu dạng bảng (tabular dataset) có 150 đặc trưng (features) tương quan lẫn nhau (correlated) và có phạm vi giá trị khác nhau (different ranges). Mục tiêu là xây dựng mô hình hồi quy (regression model) một cách hiệu quả hơn về thời gian huấn luyện (efficient model training), đồng thời giảm thiểu tác động đến hiệu suất mô hình (minimizes impact on performance).

Data scientist chọn Principal Component Analysis (PCA) làm bước tiền xử lý để giảm số lượng features xuống tập hợp nhỏ hơn các features độc lập (independent features) trước khi đưa vào mô hình hồi quy.

Vấn đề cốt lõi:

  • Features correlated → PCA giúp loại bỏ tương quan bằng cách tạo ra các principal components độc lập.
  • Different ranges → PCA rất nhạy cảm với scale (phương sai của features có scale lớn sẽ dominate, làm méo mó kết quả). Do đó, bắt buộc phải scale dữ liệu trước PCA để đảm bảo tính công bằng.
  • Sử dụng Amazon SageMaker (cụ thể Data Wrangler và built-in PCA algorithm) để thực hiện, phù hợp với best practice AWS ML workflow (cập nhật đến 2026).

🛠️ Yêu cầu chính: Chọn bước preprocessing chính xác nhất, sử dụng SageMaker tools, scale đúng cách trước PCA, không làm thay đổi bản chất PCA thuần túy.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng:
Load the data into Amazon SageMaker Data Wrangler. Scale the data with a Min Max Scaler transformation step. Use the SageMaker built-in algorithm for PCA on the scaled dataset to transform the data.

Lý do:

  • SageMaker Data Wrangler là công cụ lý tưởng cho data preparation (flow editor visual), hỗ trợ import data, transformations như scaling và PCA trực tiếp trong pipeline.
  • Min Max Scaler (scale về [0,1]) là lựa chọn phù hợp cho PCA khi features có ranges khác nhau, giúp normalize mà không làm mất thông tin variance (PCA dựa trên covariance matrix).
  • Quy trình: Load → Scale → PCA → Transform → Ready for training. Điều này tối ưu hiệu quả, không ảnh hưởng performance vì PCA sau scaling giữ nguyên thông tin chính.
  • Theo best practice AWS (2026), PCA built-in algorithm trong SageMaker yêu cầu scaling trước để tránh bias từ scales khác nhau.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên tính chính xác, tuân thủ PCA best practice và yêu cầu câu hỏi (scale trước PCA, không can thiệp thủ công vào features).

  • ❌ Phương án SAI: Use the Amazon SageMaker built-in algorithm for PCA on the dataset to transform the data.
    Lý do sai: PCA built-in của SageMaker KHÔNG scale tự động và rất nhạy cảm với features có different ranges. Features scale lớn sẽ dominate variance → PCA biased, ảnh hưởng nghiêm trọng đến model performance. Thiếu bước scale → Không đạt yêu cầu "minimizes impact".

  • ✅ Phương án ĐÚNG: Load the data into Amazon SageMaker Data Wrangler. Scale the data with a Min Max Scaler transformation step. Use the SageMaker built-in algorithm for PCA on the scaled dataset to transform the data.
    Lý do đúng: Quy trình hoàn hảo: Data Wrangler xử lý end-to-end, Min Max Scaler normalize ranges → PCA chính xác trên dữ liệu scaled → Giảm features correlated hiệu quả, training nhanh hơn. Hoàn toàn khớp yêu cầu.

  • ❌ Phương án SAI: Reduce the dimensionality of the dataset by removing the features that have the highest correlation. Load the data into Amazon SageMaker Data Wrangler. Perform a Standard Scaler transformation step to scale the data. Use the SageMaker built-in algorithm for PCA on the scaled dataset to transform the data.
    Lý do sai: Remove features highest correlation là phương pháp thủ công (feature selection), KHÔNG phải PCA thuần túy (PCA transform linear combinations, không remove trực tiếp). Điều này thay đổi dữ liệu gốc, có thể mất thông tin → Tác động performance. Standard Scaler tốt nhưng bước remove thừa và sai hướng.

  • ❌ Phương án SAI: Reduce the dimensionality of the dataset by removing the features that have the lowest correlation. Load the data into Amazon SageMaker Data Wrangler. Perform a Min Max Scaler transformation step to scale the data. Use the SageMaker built-in algorithm for PCA on the scaled dataset to transform the data.
    Lý do sai: Remove lowest correlation còn tệ hơn – giữ lại highly correlated features → Không giải quyết vấn đề tương quan, PCA sau đó vẫn kém hiệu quả. Min Max Scaler đúng nhưng bước remove sai logic (PCA cần giữ tất cả để compute components), làm phức tạp hóa và ảnh hưởng performance.

📘 Tài liệu tham khảo (AWS cập nhật mới nhất đến 2026)

  • SageMaker PCA Algorithm: docs.aws.amazon.com/sagemaker/latest/dg/pca.html – Nhấn mạnh "Scale features trước PCA vì sensitivity to variance".
  • SageMaker Data Wrangler: docs.aws.amazon.com/sagemaker/latest/dg/data-wrangler.html – Hỗ trợ MinMaxScaler, StandardScaler và PCA transforms trong visual flow.
  • Best Practices ML Preprocessing: AWS ML Specialty Exam Guide & SageMaker Processing Jobs (2026 updates: Tích hợp tự động scaling trong Data Wrangler pipelines).
  • Thực hành: AWS re:Post & Workshops – PCA với scaling cho tabular data.

🛠️ Lời khuyên DevOps: Trong pipeline CI/CD, tích hợp Data Wrangler export thành Processing Job để automate preprocessing trước training! 🚀