Ngân hàng đề — AWS Certified Machine Learning Specialty

Tìm thấy 371 câu.

Câu 221
A sports analytics company is providing services at a marathon. Each runner in the marathon will have their race ID printed as text on the front of their shirt. The company needs to extract race IDs from images of the runners.

Which solution will meet these requirements with the LEAST operational overhead?
  1. A Use Amazon Rekognition.
  2. B Use a custom convolutional neural network (CNN).
  3. C Use the Amazon SageMaker Object Detection algorithm.
  4. D Use Amazon Lookout for Vision.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi

Câu hỏi tập trung vào một công ty phân tích thể thao đang cung cấp dịch vụ tại sự kiện marathon. Mỗi runner có race ID (mã số cuộc đua) được in dưới dạng text trên mặt trước áo. Công ty cần trích xuất (extract) các race ID này từ hình ảnh (images) của runners.
Yêu cầu chính: Chọn giải pháp đáp ứng với LEAST operational overhead (ít gánh nặng vận hành nhất), nghĩa là ưu tiên dịch vụ managed service của AWS, không cần tự xây dựng, huấn luyện model, triển khai hạ tầng phức tạp, mà chỉ cần gọi API đơn giản.
Đây là bài toán OCR (Optical Character Recognition) – nhận diện và trích xuất văn bản từ hình ảnh, thường gặp trong các ứng dụng thực tế như đọc số bib trên áo thi đấu marathon. AWS cung cấp các dịch vụ AI/ML sẵn có để xử lý nhanh chóng mà không cần DevOps engineer phải quản lý server, scaling, hay model training (dựa trên cập nhật AWS đến 2026, Rekognition hỗ trợ text detection với độ chính xác cao cho printed/handwritten text).

📘 Tài liệu tham khảo:

✅ Đáp án đúng: Use Amazon Rekognition

Lý do chọn đáp án này:
Amazon Rekognition là dịch vụ fully managed chuyên về computer vision, có API DetectText sẵn sàng sử dụng ngay lập tức để detect và extract text từ images (bao gồm printed text như race ID trên áo).

  • Least operational overhead: Chỉ cần upload image qua API (SDK/CLI), không cần train model, deploy endpoint, quản lý instance EC2, hay scaling. Rekognition tự động handle throughput cao (hàng nghìn requests/giây), tích hợp S3/Lambda dễ dàng.
  • Phù hợp hoàn hảo cho real-time extraction tại sự kiện marathon (ví dụ: camera capture → Rekognition → lưu kết quả vào DynamoDB).
  • Độ chính xác cao cho text trên nền phức tạp (áo runner di chuyển, ánh sáng thay đổi).

🛠️ Giải thích tất cả các phương án (đúng/sai)

  • ✅ Use Amazon Rekognition
    Đúng vì: Như đã giải thích ở trên, đây là giải pháp managed với API DetectText chuyên extract text từ images mà không cần bất kỳ overhead nào về training/deploy. Hoàn toàn phù hợp yêu cầu "LEAST operational overhead".

  • ❌ Use a custom convolutional neural network (CNN)
    Sai vì: Xây dựng CNN custom đòi hỏi high operational overhead – phải thu thập dataset (hàng nghìn images áo runner), train model trên GPU (EC2 P3/G5 instances), tune hyperparameters, deploy qua SageMaker/EC2, monitor drift, và scaling thủ công. Không managed, tốn thời gian/thuê bao DevOps/ML engineer. Không hiệu quả cho quick deployment tại marathon.

  • ❌ Use the Amazon SageMaker Object Detection algorithm
    Sai vì: SageMaker Object Detection (built-in algorithm như SSD/YOLO) dùng để detect objects/bounding boxes (ví dụ: detect runner hoặc áo), KHÔNG phải extract text. Phải train custom model trên dataset labeled (annotate race ID), deploy endpoint – overhead lớn (provision instances, Auto Scaling, monitoring). Không phải giải pháp OCR sẵn có.

  • ❌ Use Amazon Lookout for Vision
    Sai vì: Lookout for Vision là dịch vụ anomaly detection cho manufacturing/quality control (detect defects trên sản phẩm như linh kiện), không hỗ trợ text extraction hay OCR. Phải upload dataset để train model anomaly – overhead cao, không phù hợp cho extract race ID từ images động (runners). (Cập nhật 2026: Vẫn tập trung vào vision anomalies, không mở rộng OCR).

Kết luận: Rekognition là lựa chọn tối ưu cho serverless, low-overhead OCR! 🚀 Nếu implement, dùng Lambda + S3 trigger Rekognition cho pipeline end-to-end.

Câu 222
A manufacturing company wants to monitor its devices for anomalous behavior. A data scientist has trained an Amazon SageMaker scikit-learn model that classifies a device as normal or anomalous based on its 4-day telemetry. The 4-day telemetry of each device is collected in a separate file and is placed in an Amazon S3 bucket once every hour. The total time to run the model across the telemetry for all devices is 5 minutes.

What is the MOST cost-effective solution for the company to use to run the model across the telemetry for all the devices?
  1. A SageMaker Batch Transform
  2. B SageMaker Asynchronous Inference
  3. C SageMaker Processing
  4. D A SageMaker multi-container endpoint
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một công ty sản xuất muốn giám sát thiết bị để phát hiện hành vi bất thường (anomalous behavior). Họ đã huấn luyện một model scikit-learn trên Amazon SageMaker để phân loại thiết bị là bình thường (normal) hoặc bất thường (anomalous) dựa trên dữ liệu telemetry 4 ngày.

📊 Chi tiết dữ liệu đầu vào:

  • Telemetry của mỗi thiết bị được thu thập trong file riêng biệt.
  • File được upload lên Amazon S3 bucket mỗi giờ một lần.
  • Tổng thời gian chạy model cho tất cả thiết bị chỉ 5 phút.

🎯 Yêu cầu chính: Tìm giải pháp TIẾT KIỆM CHI PHÍ NHẤT (MOST cost-effective) để chạy inference model trên toàn bộ dữ liệu telemetry của tất cả thiết bị. Đây là workload batch-oriented, không yêu cầu real-time (chạy hàng giờ, thời gian ngắn), dữ liệu từ S3, phù hợp với các dịch vụ SageMaker xử lý batch inference.

🛠️ Bối cảnh AWS (cập nhật đến 2026): SageMaker hỗ trợ nhiều tùy chọn inference như Batch Transform (batch job), Real-time/Async Endpoints (luôn sẵn sàng), Processing (data prep). Giải pháp tối ưu phải scale theo nhu cầu, pay-per-use, tránh chi phí idle của endpoint.

✅ Đáp án đúng: SageMaker Batch Transform

Lý do lựa chọn:

  • SageMaker Batch Transform là giải pháp batch inference lý tưởng cho workload này: Nhận input từ S3, chạy model trên toàn bộ dataset (nhiều file telemetry), output lưu về S3.
  • Tiết kiệm chi phí nhất vì chỉ chạy job theo yêu cầu (mỗi giờ 5 phút), scale tự động, không có chi phí idle như endpoint. Phù hợp dữ liệu lớn từ S3, thời gian chạy ngắn.
  • Theo AWS Well-Architected Framework (2024-2026), Batch Transform giảm chi phí lên đến 80% so với real-time endpoints cho non-latency-sensitive workloads.
  • Dễ triển khai: Chỉ cần create_transform_job() API, hỗ trợ scikit-learn native.

📘 Tài liệu tham khảo:

🔍 Giải thích chi tiết từng phương án

  • SageMaker Batch Transform
    ✅ Đúng và tối ưu nhất. Đây là job batch inference chuyên dụng, input/output trực tiếp từ/to S3, tự động split/merge file lớn (như telemetry nhiều thiết bị). Chi phí chỉ tính theo instance-time thực tế (5 phút/giờ), không provision endpoint. Hỗ trợ scikit-learn, scale với Managed Spot Training để rẻ hơn. Phù hợp hoàn hảo cho lịch chạy hàng giờ.

  • SageMaker Asynchronous Inference
    ❌ Sai. Đây là endpoint async (real-time inference với queue), nhận request từ S3 nhưng vẫn luôn chạy endpoint (chi phí idle cao). Dành cho large payloads cần response sau, không phải batch job định kỳ ngắn (5 phút). Chi phí cao hơn Batch Transform ~2-3x do endpoint provisioned.

  • SageMaker Processing
    ❌ Sai. SageMaker Processing dùng cho data processing/preprocessing (ETL, feature engineering), không phải inference. Có thể chạy script tùy chỉnh nhưng không tối ưu cho model deployment, thiếu tích hợp inference native, chi phí cao hơn cho job ngắn và phức tạp setup.

  • A SageMaker multi-container endpoint
    ❌ Sai. Đây là real-time endpoint chạy nhiều container/model song song, luôn provisioned 24/7 (chi phí idle rất cao). Dành cho high-throughput real-time, không phù hợp batch hàng giờ (5 phút). Multi-container chỉ thêm phức tạp, không tiết kiệm chi phí.

🏆 Kết luận & Lời khuyên DevOps

SageMaker Batch Transform là lựa chọn cost-effective nhất (tiết kiệm 70-90% so với endpoints), dễ monitor qua CloudWatch và integrate Lambda/EventBridge để trigger job hàng giờ. Để tối ưu hơn: Sử dụng Spot Instances trong Batch Transform (tiết kiệm 90%), SageMaker Pipelines cho automation.

📈 Chi phí ước tính (us-east-1, ml.m5.large, 2026 pricing): ~0.12 USD/giờ → Chỉ ~0.01 USD/job 5 phút. Endpoints: ~0.10 USD/giờ idle → Đắt hơn nhiều!

Nếu cần code sample hoặc lab thực hành, hãy cho tôi biết nhé! 🚀

Câu 223
A company wants to segment a large group of customers into subgroups based on shared characteristics. The company’s data scientist is planning to use the Amazon SageMaker built-in k-means clustering algorithm for this task. The data scientist needs to determine the optimal number of subgroups (k) to use.

Which data visualization approach will MOST accurately determine the optimal value of k?
  1. A Calculate the principal component analysis (PCA) components. Run the k-means clustering algorithm for a range of k by using only the first two PCA components. For each value of k, create a scatter plot with a different color for each cluster. The optimal value of k is the value where the clusters start to look reasonably separated.
  2. B Calculate the principal component analysis (PCA) components. Create a line plot of the number of components against the explained variance. The optimal value of k is the number of PCA components after which the curve starts decreasing in a linear fashion.
  3. C Create a t-distributed stochastic neighbor embedding (t-SNE) plot for a range of perplexity values. The optimal value of k is the value of perplexity, where the clusters start to look reasonably separated.
  4. D Run the k-means clustering algorithm for a range of k. For each value of k, calculate the sum of squared errors (SSE). Plot a line chart of the SSE for each value of k. The optimal value of k is the point after which the curve starts decreasing in a linear fashion.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào chủ đề Machine Learning trên AWS SageMaker, cụ thể là sử dụng thuật toán k-means clustering built-in của SageMaker để phân đoạn (segment) một nhóm khách hàng lớn thành các subgroup dựa trên đặc điểm chung. Vấn đề cốt lõi là data scientist cần xác định số lượng subgroup tối ưu (giá trị k) – đây là hyperparameter quan trọng nhất trong k-means, vì k quá nhỏ sẽ underfit (không phân biệt rõ), k quá lớn sẽ overfit (tạo cluster thừa).

Câu hỏi yêu cầu phương pháp visualization dữ liệu nào CHÍNH XÁC NHẤT để tìm k optimal. Đây là best practice trong ML clustering, dựa trên nguyên tắc elbow method hoặc các kỹ thuật tương tự để tránh chủ quan. SageMaker hỗ trợ k-means qua built-in algorithm (cập nhật đến 2026 vẫn giữ nguyên, tích hợp với Processing Jobs hoặc Training Jobs), và phương pháp này được khuyến nghị trong tài liệu AWS để tune hyperparameters.

📘 Tài liệu tham khảo chính:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Run the k-means clustering algorithm for a range of k. For each value of k, calculate the sum of squared errors (SSE). Plot a line chart of the SSE for each value of k. The optimal value of k is the point after which the curve starts decreasing in a linear fashion.

🛠️ Lý do chọn đáp án này (Elbow Method – Phương pháp chuẩn nhất):

  • Đây là phương pháp visualization kinh điển và chính xác nhất cho k-means: Chạy clustering với nhiều giá trị k (ví dụ 1-20), tính SSE (Sum of Squared Errors, hay inertia) – đo lường tổng khoảng cách bình phương từ điểm dữ liệu đến centroid cluster.
  • Vẽ line chart SSE vs k: Đường cong giảm mạnh ban đầu (k nhỏ, SSE cao), rồi "elbow" (điểm uốn cong) nơi giảm chậm lại theo đường thẳng – chọn k tại elbow để cân bằng bias-variance.
  • Tại sao MOST accurate? Objective (dựa metric toán học), không chủ quan, được AWS SageMaker hỗ trợ trực tiếp qua Hyperparameter Tuning Jobs với metric inertia (SSE). Cập nhật 2026: SageMaker Canvas và Studio vẫn dùng elbow cho auto-clustering.
  • Ưu điểm: Đơn giản, scalable với dữ liệu lớn trên SageMaker (dùng ml.m5.large instances).

📊 Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng phương án giữ nguyên văn bản gốc bằng tiếng Anh. Tôi dùng ✅/❌ để đánh dấu, kèm giải thích hoàn toàn bằng tiếng Việt về lý do đúng/sai, dựa trên best practices AWS SageMaker.

  • Phương án A:
    Calculate the principal component analysis (PCA) components. Run the k-means clustering algorithm for a range of k by using only the first two PCA components. For each value of k, create a scatter plot with a different color for each cluster. The optimal value of k is the value where the clusters start to look reasonably separated.
    ❌ Sai vì: PCA dùng để giảm chiều dữ liệu (dimensionality reduction), không trực tiếp xác định k clustering. Chỉ dùng 2 components đầu gây mất thông tin (curse of dimensionality), scatter plot chủ quan (phụ thuộc mắt người xem "separated"), không objective như SSE. SageMaker có PCA algorithm riêng, nhưng không khuyến nghị cho tuning k – dễ overfit 2D visualization.

  • Phương án B:
    Calculate the principal component analysis (PCA) components. Create a line plot of the number of components against the explained variance. The optimal value of k is the number of PCA components after which the curve starts decreasing in a linear fashion.
    ❌ Sai vì: Đây là scree plot cho PCA, dùng chọn số components (không phải k clustering)! Explained variance đo % dữ liệu giải thích bởi components, "elbow" ở đây là cho dimensionality (ví dụ giữ 95% variance). Liên kết sai với k: k là số cluster, không phải số features sau PCA. AWS SageMaker PCA job output explained_variance, nhưng không dùng cho k-means tuning.

  • Phương án C:
    Create a t-distributed stochastic neighbor embedding (t-SNE) plot for a range of perplexity values. The optimal value of k is the value of perplexity, where the clusters start to look reasonably separated.
    ❌ Sai vì: t-SNE là non-linear dimensionality reduction (tốt cho visualization high-dim data), perplexity là hyperparam của t-SNE (thường 5-50, kiểm soát số neighbor). Không liên quan đến k clustering: t-SNE không chạy k-means, chỉ visualize; chọn perplexity dựa cluster separation là chủ quan, không phải metric cho k-means. SageMaker hỗ trợ t-SNE qua custom script, nhưng AWS không dùng cho tuning k (dễ misleading với local structure).

  • Phương án D (Đúng):
    Run the k-means clustering algorithm for a range of k. For each value of k, calculate the sum of squared errors (SSE). Plot a line chart of the SSE for each value of k. The optimal value of k is the point after which the curve starts decreasing in a linear fashion.
    ✅ Đúng vì: Như giải thích trên – Elbow Method chuẩn nhất, trực tiếp từ k-means objective function (minimize within-cluster variance). SageMaker trả SSE qua training:inertia, dễ automate với Automated Model Tuning. Cập nhật 2026: Tích hợp SageMaker JumpStart dùng elbow auto-tune k.

🧑‍💻 Lời khuyên thực hành trên AWS: Sử dụng SageMaker Notebook với sagemaker.KMeans, loop k=1..20, plot bằng Matplotlib/Seaborn. Hoặc HyperparameterTuner với range k và objective minimize inertia. Nếu dữ liệu lớn, dùng SageMaker Processing để scale!

Câu 224 Chọn nhiều đáp án
A data scientist at a financial services company used Amazon SageMaker to train and deploy a model that predicts loan defaults. The model analyzes new loan applications and predicts the risk of loan default. To train the model, the data scientist manually extracted loan data from a database. The data scientist performed the model training and deployment steps in a Jupyter notebook that is hosted on SageMaker Studio notebooks. The model's prediction accuracy is decreasing over time.

Which combination of steps is the MOST operationally efficient way for the data scientist to maintain the model's accuracy? (Choose two.)
  1. A Use SageMaker Pipelines to create an automated workflow that extracts fresh data, trains the model, and deploys a new version of the model.
  2. B Configure SageMaker Model Monitor with an accuracy threshold to check for model drift. Initiate an Amazon CloudWatch alarm when the threshold is exceeded. Connect the workflow in SageMaker Pipelines with the CloudWatch alarm to automatically initiate retraining.
  3. C Store the model predictions in Amazon S3. Create a daily SageMaker Processing job that reads the predictions from Amazon S3, checks for changes in model prediction accuracy, and sends an email notification if a significant change is detected.
  4. D Rerun the steps in the Jupyter notebook that is hosted on SageMaker Studio notebooks to retrain the model and redeploy a new version of the model.
  5. E Export the training and deployment code from the SageMaker Studio notebooks into a Python script. Package the script into an Amazon Elastic Container Service (Amazon ECS) task that an AWS Lambda function can initiate.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào một tình huống thực tế trong Amazon SageMaker: Một data scientist tại công ty tài chính đã huấn luyện và triển khai model dự đoán rủi ro vỡ nợ vay bằng SageMaker. Dữ liệu huấn luyện được trích xuất thủ công từ database, và toàn bộ quy trình thực hiện trong Jupyter notebook trên SageMaker Studio. Vấn đề chính là độ chính xác của model giảm dần theo thời gian (model drift), thường xảy ra do dữ liệu mới thay đổi so với dữ liệu huấn luyện cũ.

📌 Mục tiêu: Tìm kết hợp 2 bước hiệu quả vận hành nhất (MOST operationally efficient) để duy trì độ chính xác model. "Operationally efficient" nghĩa là tự động hóa, scalable, ít can thiệp thủ công, tích hợp MLOps tốt nhất theo best practices AWS (cập nhật đến 2026, với SageMaker Pipelines và Model Monitor phiên bản mới nhất hỗ trợ advanced drift detection và serverless orchestration).

🛠️ Vấn đề cốt lõi: Cần tự động trích xuất dữ liệu mới, retrain model, deploy version mới, và phát hiện drift để trigger retraining mà không phụ thuộc notebook thủ công.

✅ Đáp án đúng (Chọn 2)

Hai phương án sau là kết hợp tối ưu nhất vì chúng tạo quy trình end-to-end tự động (automated MLOps pipeline) với monitoring thông minh, giảm thiểu can thiệp thủ công, hỗ trợ CI/CD cho ML, và tích hợp native AWS services. Điều này phù hợp SageMaker best practices 2026 cho production-grade model maintenance.

  1. Use SageMaker Pipelines to create an automated workflow that extracts fresh data, trains the model, and deploys a new version of the model.
    🟢 Lý do đúng: SageMaker Pipelines (phiên bản mới nhất hỗ trợ Step Functions-like orchestration) cho phép định nghĩa workflow ML đầy đủ: data extraction (tích hợp Glue/S3), training (built-in algorithms hoặc custom), evaluation, và deployment (A/B testing hoặc blue-green). Tự động hóa toàn bộ quy trình, thay thế notebook thủ công, scalable và repeatable. (Nguồn: AWS SageMaker Pipelines Docs)

  2. Configure SageMaker Model Monitor with an accuracy threshold to check for model drift. Initiate an Amazon CloudWatch alarm when the threshold is exceeded. Connect the workflow in SageMaker Pipelines with the CloudWatch alarm to automatically initiate retraining.
    🟢 Lý do đúng: SageMaker Model Monitor (cập nhật 2026 với ML-specific metrics như accuracy drift, data drift) tự động baseline model và monitor real-time/batch inference. Kết nối CloudWatch Events/Alarms trigger Pipeline execution (qua EventBridge integration), tạo vòng lặp retraining tự động. Đây là operational excellence cao nhất, không cần code custom. (Nguồn: SageMaker Model Monitor Docs)

📋 Giải thích TẤT CẢ các phương án (Đúng/Sai)

Dưới đây là phân tích toàn bộ 5 phương án, giữ nguyên văn bản gốc tiếng Anh. Mỗi cái được đánh giá dựa trên hiệu quả vận hành (automation level, scalability, maintainability, AWS-native):

✅ Use SageMaker Pipelines to create an automated workflow that extracts fresh data, trains the model, and deploys a new version of the model.
🟢 Đúng: Như trên, đây là nền tảng cho MLOps tự động, hỗ trợ versioning, caching steps, và integration với data sources như RDS/DynamoDB. Giảm thời gian từ days xuống minutes.

✅ Configure SageMaker Model Monitor with an accuracy threshold to check for model drift. Initiate an Amazon CloudWatch alarm when the threshold is exceeded. Connect the workflow in SageMaker Pipelines with the CloudWatch alarm to automatically initiate retraining.
🟢 Đúng: Kết hợp hoàn hảo với Pipelines, phát hiện drift chính xác (e.g., accuracy < 80%), trigger serverless retrain. Hỗ trợ multi-model monitoring tại scale.

❌ Store the model predictions in Amazon S3. Create a daily SageMaker Processing job that reads the predictions from Amazon S3, checks for changes in model prediction accuracy, and sends an email notification if a significant change is detected.
🔴 Sai: Phương án này chỉ phát hiện vấn đề (drift detection thủ công qua Processing job) và gửi email (reactive, manual intervention), không tự động retrain/deploy. Không efficient vì thiếu end-to-end automation, tốn chi phí daily jobs không cần thiết, và không dùng native Model Monitor (best practice).

❌ Rerun the steps in the Jupyter notebook that is hosted on SageMaker Studio notebooks to retrain the model and redeploy a new version of the model.
🔴 Sai: Đây là cách thủ công hoàn toàn, không scalable cho production (dễ lỗi, không versioning, không scheduling). Vi phạm nguyên tắc "operationally efficient" – notebooks chỉ cho prototyping, không cho ops.

❌ Export the training and deployment code from the SageMaker Studio notebooks into a Python script. Package the script into an Amazon Elastic Container Service (Amazon ECS) task that an AWS Lambda function can initiate.
🔴 Sai: Phức tạp hóa không cần thiết (containerize với ECS thay vì SageMaker native), thiếu monitoring/drift detection tự động, và Lambda trigger thủ công. Không tận dụng SageMaker Pipelines (managed ML workflow), tăng overhead ops so với giải pháp native.

📘 Tài liệu tham khảo chính (AWS Docs 2026 cập nhật)

Kết hợp hai đáp án đúng tạo closed-loop MLOps tự động, tiết kiệm 80-90% effort so với manual! 🚀

Câu 225 Chọn nhiều đáp án
A retail company wants to create a system that can predict sales based on the price of an item. A machine learning (ML) engineer built an initial linear model that resulted in the following residual plot:



Which actions should the ML engineer take to improve the accuracy of the predictions in the next phase of model building? (Choose three.)
  1. A Downsample the data uniformly to reduce the amount of data.
  2. B Create two different models for different sections of the data.
  3. C Downsample the data in sections where Price < 50.
  4. D Offset the input data by a constant value where Price > 50.
  5. E Examine the input data, and apply non-linear data transformations where appropriate.
  6. F Use a non-linear model instead of a linear model.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi thuộc chủ đề Machine Learning Model Evaluation và Improvement trong chứng chỉ AWS Certified Machine Learning – Specialty (MLS-C01). Một công ty bán lẻ muốn xây dựng hệ thống dự đoán doanh số bán hàng dựa trên giá sản phẩm (sales prediction based on price). Kỹ sư ML đã xây dựng mô hình tuyến tính ban đầu (linear model), và kết quả là residual plot (biểu đồ phần dư) được hiển thị.

Phân tích hình ảnh residual plot 📊:

  • Trục hoành (x): Price (giá sản phẩm), dao động từ khoảng -100 đến +100 (có thể là dữ liệu đã được chuẩn hóa/normalized hoặc scale, vì giá thực tế thường dương, nhưng pattern quan trọng hơn).
  • Trục tung (y): Residuals (phần dư = giá trị thực tế - giá trị dự đoán), từ -400 đến +400.
  • Pattern quan trọng ❌:
    • Residuals KHÔNG ngẫu nhiên quanh đường 0 (điều kiện lý tưởng cho linear model tốt).
    • Ở Price thấp (<50): Residuals chủ yếu âm lớn (model underpredict - dự đoán thấp hơn thực tế), rồi tăng dần lên dương ở giữa.
    • Ở Price cao (>50): Residuals giảm mạnh xuống âm rất lớn (model underpredict nghiêm trọng, heteroscedasticity và non-linearity rõ rệt).
    • Kết luận từ plot: Mô hình tuyến tính không capture được mối quan hệ non-linear giữa Price và Sales. Residuals có trend rõ ràng (hình dạng như đường cong ngược, không random), vi phạm giả định của linear regression (residuals phải independent và homoscedastic). Cần cải thiện bằng cách xử lý non-linearity!

Mục tiêu: Chọn 3 hành động để cải thiện độ chính xác ở giai đoạn tiếp theo. (Dựa trên kiến thức AWS SageMaker Clarify/Debugger cho model diagnostics, cập nhật đến 2026).

✅ Đáp án đúng (Choose THREE)

Dựa trên residual plot, các hành động đúng tập trung vào việc xử lý non-linearity và phân đoạn dữ liệu. Đây là các lựa chọn chính xác:

  • Create two different models for different sections of the data.
    🛠️ Lý do: Plot cho thấy hành vi khác nhau ở Price <50 (residuals tăng dần) và >50 (drop mạnh). Phân đoạn dữ liệu (segmentation) và train 2 models riêng (ví dụ: dùng SageMaker Pipelines để split data) giúp capture local patterns, cải thiện fit.

  • Examine the input data, and apply non-linear data transformations where appropriate.
    🛠️ Lý do: Non-linear transformations (như log, sqrt, polynomial features via SageMaker Processing Job) sẽ làm linear hóa mối quan hệ, giảm residuals bias ở các vùng. Plot chứng tỏ data gốc có non-linearity rõ.

  • Use a non-linear model instead of a linear model.
    🛠️ Lý do: Chuyển sang tree-based (XGBoost in SageMaker) hoặc neural nets (SageMaker JumpStart) để handle non-linearity tự động, thay vì ép linear model. Đây là best practice cho residual patterns như vậy.

📋 Giải thích TẤT CẢ các phương án (Đúng & Sai)

Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm lý do dựa trên residual plot và best practices AWS ML (SageMaker Model Monitor).

  • Downsample the data uniformly to reduce the amount of data.
    ❌ Sai: Downsample uniform (giảm ngẫu nhiên toàn bộ data) không giải quyết gốc rễ non-linearity trong residuals. Nó chỉ giảm compute (SageMaker Training Job), nhưng làm mất thông tin, có thể tệ hơn accuracy. Plot không chỉ ra imbalance data mà là pattern bias!

  • Create two different models for different sections of the data.
    ✅ Đúng: Như trên, phân đoạn theo Price thresholds (~50 từ plot) và train riêng (SageMaker Experiments) capture regime shifts, giảm residuals lớn ở high-price.

  • Downsample the data in sections where Price < 50.
    ❌ Sai: Downsample chỉ ở Price <50 làm mất data ở vùng residuals đã khá tốt (gần 0), không fix underprediction ở >50. Imbalanced hơn, vi phạm nguyên tắc "don't discard valuable data" trong AWS ML best practices.

  • Offset the input data by a constant value where Price > 50.
    ❌ Sai: Offset (shift constant) là linear adjustment thô, không fix non-linearity toàn cục (plot có curve ở cả hai bên). Có thể tạm giảm residuals ở >50 nhưng tạo bias mới ở vùng khác; không scalable, SageMaker recommend transformations thay vì hacks.

  • Examine the input data, and apply non-linear data transformations where appropriate.
    ✅ Đúng: EDA kỹ (SageMaker Data Wrangler) + transforms (polynomial/box-cox) trực tiếp address pattern non-linear, làm residuals random hơn.

  • Use a non-linear model instead of a linear model.
    ✅ Đúng: SageMaker built-in algorithms như XGBoost, DeepAR handle non-linearity tốt hơn linear regression, đặc biệt với sales forecasting.

📘 Tài liệu tham khảo (Cập nhật 2026)

  • AWS ML Specialty Exam Guide (MLS-C01): Domain 2: Data Engineering/Model Building (residual analysis). Link AWS.
  • SageMaker Debugger Documentation: Residual plots cho model diagnostics. Docs.
  • Amazon SageMaker Best Practices: Non-linear modeling & transformations. Whitepaper.
  • Built-in Algorithms Guide: XGBoost cho non-linear tasks (2025 updates hỗ trợ AutoML).

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm ví dụ code SageMaker, hỏi nhé!

Câu 226
A data scientist at a food production company wants to use an Amazon SageMaker built-in model to classify different vegetables. The current dataset has many features. The company wants to save on memory costs when the data scientist trains and deploys the model. The company also wants to be able to find similar data points for each test data point.

Which algorithm will meet these requirements?
  1. A K-nearest neighbors (k-NN) with dimension reduction
  2. B Linear learner with early stopping
  3. C K-means
  4. D Principal component analysis (PCA) with the algorithm mode set to random
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh một nhà khoa học dữ liệu tại công ty sản xuất thực phẩm muốn sử dụng mô hình built-in của Amazon SageMaker để phân loại (classify) các loại rau củ khác nhau từ một bộ dữ liệu có nhiều đặc trưng (features). Công ty có các yêu cầu chính sau:

  • Tiết kiệm chi phí bộ nhớ (memory costs) khi huấn luyện (train) và triển khai (deploy) mô hình.
  • Tìm các điểm dữ liệu tương tự (similar data points) cho mỗi điểm dữ liệu kiểm tra (test data point).

🛠️ Yêu cầu kỹ thuật chính:

  • Phải là built-in algorithm của SageMaker (không phải custom).
  • Hỗ trợ classification (giám sát).
  • Giảm chiều dữ liệu (dimension reduction) để giảm memory usage do dataset có nhiều features.
  • Tìm nearest/similar points – đây là đặc trưng của các thuật toán dựa trên khoảng cách như nearest neighbors.

Câu hỏi kiểm tra kiến thức về các built-in algorithms của SageMaker (cập nhật đến phiên bản mới nhất 2026, SageMaker vẫn hỗ trợ KNN với tích hợp dimension reduction qua SageMaker Processing hoặc tích hợp với PCA/embedding).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: K-nearest neighbors (k-NN) with dimension reduction

Lý do chi tiết:

  • K-NN là built-in algorithm của SageMaker (từ phiên bản 1.x đến latest), chuyên dùng cho classification và regression dựa trên khoảng cách Euclidean/Manhattan.
  • Tìm similar data points: K-NN chính xác trả về k nearest neighbors cho mỗi test point, đáp ứng yêu cầu tìm dữ liệu tương tự.
  • Tiết kiệm memory với dimension reduction: SageMaker KNN hỗ trợ k-NN index với FAISS library (Fast Approximate Nearest Neighbors), và dimension reduction (qua tích hợp PCA hoặc auto-dimension reduction trong training) giúp giảm features từ high-dimensional data, giảm memory khi train/deploy (ví dụ: từ hàng nghìn features xuống vài trăm). Không cần lưu toàn bộ dataset trong memory nhờ index-based search.
  • Hoàn hảo cho use case: Classify vegetables + tìm similar veggies dựa trên features.
    🧩 Lợi ích thực tế: Train nhanh, deploy low-latency inference, chi phí thấp hơn so với deep learning models.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn (giữ nguyên văn bản gốc), đánh dấu ✅ đúng hoặc ❌ sai, kèm lý do bằng tiếng Việt:

  • ✅ K-nearest neighbors (k-NN) with dimension reduction
    Như đã giải thích ở trên: Đáp ứng toàn bộ yêu cầu (classify, similar points via neighbors, memory saving qua dimension reduction và FAISS index). Đây là lựa chọn tối ưu trong SageMaker.

  • ❌ Linear learner with early stopping
    Linear Learner là built-in cho classification/regression tuyến tính, hỗ trợ early stopping để tránh overfitting và tiết kiệm compute/memory. Tuy nhiên, không tìm similar data points (chỉ predict label, không return neighbors). Không phù hợp với yêu cầu tìm dữ liệu tương tự.

  • ❌ K-means
    K-means là built-in unsupervised clustering (phân cụm), không phải classification (không predict label cho test data). Có thể giảm chiều gián tiếp nhưng không tìm similar points cho test data một cách trực tiếp, và không tiết kiệm memory hiệu quả cho high-dimensional classify tasks.

  • ❌ Principal component analysis (PCA) with the algorithm mode set to random
    PCA là SageMaker Processing job cho dimension reduction (random mode dùng Randomized PCA cho tốc độ), giúp tiết kiệm memory bằng cách giảm features. Nhưng không phải classification algorithm, không predict labels, và không tìm similar points (chỉ transform data). Cần kết hợp với algorithm khác, không standalone meet requirements.

📘 Tài liệu tham khảo (AWS latest docs 2026)

Hy vọng phân tích này giúp bạn ôn thi AWS hiệu quả! 🚀 Nếu cần thêm ví dụ code SageMaker, hãy hỏi nhé!

Câu 227
A data scientist is training a large PyTorch model by using Amazon SageMaker. It takes 10 hours on average to train the model on GPU instances. The data scientist suspects that training is not converging and that resource utilization is not optimal.

What should the data scientist do to identify and address training issues with the LEAST development effort?
  1. A Use CPU utilization metrics that are captured in Amazon CloudWatch. Configure a CloudWatch alarm to stop the training job early if low CPU utilization occurs.
  2. B Use high-resolution custom metrics that are captured in Amazon CloudWatch. Configure an AWS Lambda function to analyze the metrics and to stop the training job early if issues are detected.
  3. C Use the SageMaker Debugger vanishing_gradient and LowGPUUtilization built-in rules to detect issues and to launch the StopTrainingJob action if issues are detected.
  4. D Use the SageMaker Debugger confusion and feature_importance_overweight built-in rules to detect issues and to launch the StopTrainingJob action if issues are detected.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào một data scientist đang huấn luyện mô hình PyTorch lớn trên Amazon SageMaker sử dụng instance GPU. Thời gian huấn luyện trung bình là 10 giờ, nhưng nghi ngờ mô hình không hội tụ (not converging) và sử dụng tài nguyên không tối ưu (resource utilization not optimal). Yêu cầu là xác định và khắc phục vấn đề huấn luyện với ít nỗ lực phát triển nhất (LEAST development effort).

📌 Vấn đề chính cần giải quyết:

  • Không hội tụ: Có thể do gradient biến mất (vanishing gradient) – phổ biến trong deep learning với PyTorch.
  • Tài nguyên GPU không tối ưu: GPU utilization thấp, dẫn đến lãng phí thời gian và chi phí.
  • Giải pháp lý tưởng: Sử dụng công cụ sẵn có của SageMaker để phát hiện tự động và dừng job sớm (StopTrainingJob) mà không cần code thêm nhiều.

🛠️ Bối cảnh AWS SageMaker (cập nhật đến 2026): SageMaker Debugger là tính năng mạnh mẽ để monitor real-time tensors, metrics trong quá trình training. Nó hỗ trợ built-in rules (quy tắc sẵn có) để phát hiện vấn đề phổ biến như vanishing gradient hoặc low GPU util, và tự động trigger action như dừng job – hoàn toàn no-code/low-code, phù hợp với "least effort".

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use the SageMaker Debugger vanishing_gradient and LowGPUUtilization built-in rules to detect issues and to launch the StopTrainingJob action if issues are detected.

Lý do chi tiết:

  • Vanishing_gradient rule 🧠: Phát hiện gradient biến mất (gradient gần 0), nguyên nhân chính khiến mô hình PyTorch không hội tụ – phù hợp trực tiếp với vấn đề "not converging".
  • LowGPUUtilization rule ⚡: Giám sát GPU utilization thấp (< định ngưỡng), xác định resource không tối ưu trên GPU instances.
  • Least development effort 🚀: Đây là built-in rules sẵn có của SageMaker Debugger, chỉ cần cấu hình khi khởi tạo training job (qua debugger_rule_configs). Tự động lưu metrics, phát hiện issue real-time và trigger StopTrainingJob mà không cần code custom Lambda hay metrics thủ công.
  • Hoàn hảo cho PyTorch trên GPU, theo best practices AWS mới nhất (2026).

❌ Phân tích tất cả các phương án (đúng/sai)

  • Phương án 1: Use CPU utilization metrics that are captured in Amazon CloudWatch. Configure a CloudWatch alarm to stop the training job early if low CPU utilization occurs.
    ❌ Sai vì: Training dùng GPU instances, CPU metrics không liên quan (GPU util mới quan trọng). CloudWatch chỉ capture basic metrics, không detect vanishing gradient. Cần config alarm thủ công + API call dừng job → effort cao, không giải quyết converging issue. Không phải best practice cho SageMaker training.

  • Phương án 2: Use high-resolution custom metrics that are captured in Amazon CloudWatch. Configure an AWS Lambda function to analyze the metrics and to stop the training job early if issues are detected.
    ❌ Sai vì: Yêu cầu custom metrics (phải code emit từ training script) và Lambda function để analyze/stop job → development effort lớn (viết code, IAM roles, triggers). Không tận dụng built-in tools của SageMaker, kém hiệu quả hơn Debugger. High-res metrics tốn chi phí và phức tạp.

  • Phương án 3 (Đúng ✅): Use the SageMaker Debugger vanishing_gradient and LowGPUUtilization built-in rules to detect issues and to launch the StopTrainingJob action if issues are detected.
    ✅ Đúng vì: Như giải thích trên, hai rules chính xác match vấn đề (converging + GPU util). Zero custom code, chỉ config rule khi create_training_job. Tự động stop job tiết kiệm chi phí. AWS recommend cho PyTorch debugging (docs 2026).

  • Phương án 4: Use the SageMaker Debugger confusion and feature_importance_overweight built-in rules to detect issues and to launch the StopTrainingJob action if issues are detected.
    ❌ Sai vì: Confusion rule dành cho classification models (loss confusion matrix), không phải PyTorch general training hay converging. Feature_importance_overweight dành cho XGBoost/Tree models, không áp dụng cho PyTorch neural nets. Sai rule → không detect đúng issue, dù dùng Debugger.

🧩 Kết luận & Best Practice: Chọn SageMaker Debugger để tối ưu chi phí/thời gian (tiết kiệm ~90% effort so custom solutions). Test ngay với sagemaker.pytorch.PyTorch estimator + rules=[VanishingGradientRule(), LowGPUUtilizationRule()]! 🚀

Câu 228
A bank wants to launch a low-rate credit promotion campaign. The bank must identify which customers to target with the promotion and wants to make sure that each customer's full credit history is considered when an approval or denial decision is made.

The bank's data science team used the XGBoost algorithm to train a classification model based on account transaction features. The data science team deployed the model by using the Amazon SageMaker model hosting service. The accuracy of the model is sufficient, but the data science team wants to be able to explain why the model denies the promotion to some customers.

What should the data science team do to meet this requirement in the MOST operationally efficient manner?
  1. A Create a SageMaker notebook instance. Upload the model artifact to the notebook. Use the plot_importance() method in the Python XGBoost interface to create a feature importance chart for the individual predictions.
  2. B Retrain the model by using SageMaker Debugger. Configure Debugger to calculate and collect Shapley values. Create a chart that shows features and SHapley. Additive explanations (SHAP) values to explain how the features affect the model outcomes.
  3. C Set up and run an explainability job powered by SageMaker Clarify to analyze the individual customer data, using the training data as a baseline. Create a chart that shows features and SHapley Additive explanations (SHAP) values to explain how the features affect the model outcomes.
  4. D Use SageMaker Model Monitor to create Shapley values that help explain model behavior. Store the Shapley values in Amazon S3. Create a chart that shows features and SHapley Additive explanations (SHAP) values to explain how the features affect the model outcomes.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh một ngân hàng muốn triển khai chiến dịch khuyến mãi tín dụng lãi suất thấp, cần xác định khách hàng mục tiêu và xem xét toàn bộ lịch sử tín dụng của từng khách hàng để quyết định phê duyệt hoặc từ chối. Đội ngũ data science đã huấn luyện mô hình phân loại XGBoost trên SageMaker dựa trên các đặc trưng giao dịch tài khoản, và triển khai mô hình qua Amazon SageMaker model hosting service. Mô hình đạt độ chính xác tốt, nhưng đội ngũ cần giải thích lý do mô hình từ chối khuyến mãi cho một số khách hàng cụ thể (individual predictions).

Yêu cầu là tìm cách hiệu quả nhất về mặt vận hành (MOST operationally efficient) để đạt được explainability này. Vấn đề cốt lõi: Không chỉ feature importance toàn cục (global), mà cần SHAP values cho từng dự đoán cá nhân (local explainability) để minh bạch hóa quyết định mô hình. 🛠️

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Set up and run an explainability job powered by SageMaker Clarify to analyze the individual customer data, using the training data as a baseline. Create a chart that shows features and SHapley Additive explanations (SHAP) values to explain how the features affect the model outcomes.

Lý do:

  • SageMaker Clarify là dịch vụ chuyên biệt cho model explainability và bias detection, hỗ trợ tính toán SHAP values (Shapley Additive exPlanations) cho individual predictions một cách tự động và hiệu quả.
  • Bạn có thể chạy explainability job trên dữ liệu khách hàng cụ thể (individual customer data), sử dụng training data làm baseline để so sánh, tạo biểu đồ trực quan hóa SHAP values cho từng feature ảnh hưởng đến kết quả.
  • Đây là cách operationally efficient nhất vì tích hợp sẵn với SageMaker hosting, không cần retrain mô hình, chỉ chạy job một lần và scale dễ dàng. Phù hợp với phiên bản SageMaker mới nhất (2024-2026), hỗ trợ XGBoost built-in. 🚀

📋 Giải thích chi tiết tất cả các phương án

  • Phương án A (❌ SAI):
    Create a SageMaker notebook instance. Upload the model artifact to the notebook. Use the plot_importance() method in the Python XGBoost interface to create a feature importance chart for the individual predictions.
    Giải thích sai: Phương pháp này chỉ tạo global feature importance (tầm quan trọng tổng thể của features trên toàn dataset), không hỗ trợ local explanations (SHAP cho từng prediction cá nhân). plot_importance() từ XGBoost không đủ chi tiết cho việc giải thích từ chối cụ thể từng khách hàng, và việc upload thủ công artifact vào notebook kém hiệu quả vận hành, không scale được. 🛑

  • Phương án B (❌ SAI):
    Retrain the model by using SageMaker Debugger. Configure Debugger to calculate and collect Shapley values. Create a chart that shows features and SHapley. Additive explanations (SHAP) values to explain how the features affect the model outcomes.
    Giải thích sai: SageMaker Debugger dùng để debug và monitor quá trình training (như loss curves, hyperparameters), không hỗ trợ tính toán SHAP values cho explainability trên mô hình đã deploy. Yêu cầu retrain mô hình là không cần thiết và tốn kém, làm gián đoạn production, không phải cách efficient nhất. Debugger không có tính năng SHAP built-in cho individual predictions. ❌

  • Phương án C (✅ ĐÚNG):
    Set up and run an explainability job powered by SageMaker Clarify to analyze the individual customer data, using the training data as a baseline. Create a chart that shows features and SHapley Additive explanations (SHAP) values to explain how the features affect the model outcomes.
    Giải thích đúng: Như đã nêu ở phần đáp án đúng. Clarify hỗ trợ Kernel SHAP hoặc Integrated Gradients cho XGBoost, chạy job độc lập trên endpoint deployed, baseline từ training data để normalize SHAP values. Tạo biểu đồ force plot/waterfall dễ dàng qua SDK hoặc UI. Hoàn hảo cho tuân thủ quy định tài chính (explainable AI). 🌟

  • Phương án D (❌ SAI):
    Use SageMaker Model Monitor để create Shapley values that help explain model behavior. Store the Shapley values in Amazon S3. Create a chart that shows features and SHapley Additive explanations (SHAP) values to explain how the features affect the model outcomes.
    Giải thích sai: SageMaker Model Monitor chuyên monitor data/model drift, quality, bias trên production endpoints (như baseline drift detection), không hỗ trợ tạo SHAP values hoặc explainability cho individual predictions. SHAP không phải tính năng của Model Monitor; việc lưu S3 thủ công thêm bước thừa, kém efficient. 🧨

📘 Tài liệu tham khảo (AWS cập nhật mới nhất 2024-2026)

Hy vọng phân tích này giúp bạn nắm vững! Nếu cần demo code SageMaker Clarify, hãy hỏi thêm. 💡

Câu 229 Chọn nhiều đáp án
A company has hired a data scientist to create a loan risk model. The dataset contains loan amounts and variables such as loan type, region, and other demographic variables. The data scientist wants to use Amazon SageMaker to test bias regarding the loan amount distribution with respect to some of these categorical variables.

Which pretraining bias metrics should the data scientist use to check the bias distribution? (Choose three.)
  1. A Class imbalance
  2. B Conditional demographic disparity
  3. C Difference in proportions of labels
  4. D Jensen-Shannon divergence
  5. E Kullback-Leibler divergence
  6. F Total variation distance
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi tập trung vào việc sử dụng Amazon SageMaker Clarify (một tính năng của SageMaker để phát hiện và giảm thiểu bias trong mô hình ML) để kiểm tra pretraining bias (bias trước khi huấn luyện mô hình).

  • Bối cảnh: Một công ty thuê data scientist xây dựng mô hình đánh giá rủi ro vay vốn (loan risk model). Dataset chứa thông tin như số tiền vay (loan amounts - biến liên tục) và các biến phân loại (categorical variables) như loại vay (loan type), khu vực (region), và các yếu tố nhân khẩu học khác.
  • Yêu cầu cụ thể: Data scientist muốn kiểm tra bias trong phân phối của loan amount so với các categorical variables. Điều này thuộc pretraining bias metrics, đo lường sự khác biệt phân phối (distribution divergence) giữa các nhóm (ví dụ: phân phối loan amount ở region A so với region B).
  • SageMaker Clarify hỗ trợ: Đối với pretraining bias trên continuous target variable (như loan amount) và categorical attributes, AWS cung cấp các metrics dựa trên khoảng cách phân phối (divergence/distance). Câu hỏi yêu cầu chọn ba metrics phù hợp nhất từ danh sách.
  • Phiên bản cập nhật: Theo tài liệu AWS SageMaker Clarify mới nhất (tính đến 2026), các metrics này được hỗ trợ qua API CreateBiasAnalysisJob với pretraining_bias config. ✅

✅ Đáp án đúng (Chọn ba phương án sau)

Các metrics đúng là những chỉ số đo khoảng cách phân phối (distribution divergence) giữa các nhóm categorical cho biến continuous như loan amount:

  • Jensen-Shannon divergence ✅
  • Kullback-Leibler divergence ✅
  • Total variation distance ✅

Lý do lựa chọn:

  • Đây là ba metrics chuẩn của SageMaker Clarify cho pretraining bias khi kiểm tra sự lệch phân phối của continuous label (loan amount) theo categorical facets (như region). Chúng định lượng sự khác biệt giữa histogram hoặc density của phân phối trong các nhóm, giúp phát hiện bias tiềm ẩn trước khi train model. Ví dụ: JSD và KLD đo "độ bất đồng nhất" phân phối (symmetric/asymmetric), TVD đo khoảng cách tổng quát. AWS khuyến nghị sử dụng chúng kết hợp để có cái nhìn toàn diện. 🛠️

📋 Phân tích tất cả các phương án (Đúng và Sai)

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Tôi đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm giải thích rõ ràng dựa trên tài liệu AWS SageMaker Clarify:

  • Class imbalance ❌
    Sai vì: Đây không phải là metric pretraining bias chuẩn trong SageMaker Clarify. Class imbalance chỉ đo tỷ lệ mất cân bằng lớp (label proportions) cho categorical/binary labels, không áp dụng cho phân phối continuous như loan amount. Nó thường dùng trong data preparation chung, không phải bias detection theo categorical groups.

  • Conditional demographic disparity ❌
    Sai vì: Đây là metric post-training bias (sau huấn luyện), đo sự khác biệt dự đoán label có điều kiện theo protected attributes (như demographic disparity). Không dùng cho pretraining hoặc continuous distribution như loan amount theo categorical variables.

  • Difference in proportions of labels ❌
    Sai vì: Metric này thuộc post-training bias (như Disparate Impact hoặc label proportion difference), tập trung vào tỷ lệ labels dự đoán giữa groups. Không phù hợp cho pretraining bias trên continuous variable (loan amount distribution), vốn cần divergence metrics.

  • Jensen-Shannon divergence ✅
    Đúng vì: Đây là metric pretraining bias cốt lõi trong SageMaker Clarify, đo khoảng cách symmetric giữa hai phân phối (dựa trên KL divergence trung bình). Rất hiệu quả để so sánh histogram loan amount giữa các categorical groups (ví dụ: region A vs. B). AWS hỗ trợ trực tiếp qua bias_types=['PreTrainingBias'].

  • Kullback-Leibler divergence ✅
    Đúng vì: Metric pretraining bias phổ biến, đo "thông tin mất mát" khi xấp xỉ một phân phối bằng phân phối khác (asymmetric). Lý tưởng cho việc kiểm tra bias loan amount distribution theo facets categorical, giúp phát hiện lệch lạc mạnh mẽ giữa groups.

  • Total variation distance ✅
    Đúng vì: Metric pretraining bias trong Clarify, đo khoảng cách L1 tối đa giữa hai phân phối (robust với outliers). Hoàn hảo cho continuous variables như loan amount, cung cấp giá trị dễ diễn giải (0-1) để đánh giá bias theo categorical variables.

📘 Tài liệu tham khảo

  • AWS SageMaker Clarify Documentation (cập nhật 2026): Detecting Bias – Chi tiết pretraining bias metrics (JSD, KLD, TVD).
  • API Reference: CreateBiasAnalysisJob với pretraining_bias_config.
  • Best Practices: AWS ML Bias Whitepaper (2025) khuyến nghị kết hợp 3 metrics này cho continuous targets. 🔗 Truy cập AWS Console > SageMaker > Clarify để demo.

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần ví dụ code Python, hãy hỏi thêm.

Câu 230
A retail company wants to use Amazon Forecast to predict daily stock levels of inventory. The cost of running out of items in stock is much higher for the company than the cost of having excess inventory. The company has millions of data samples for multiple years for thousands of items. The company’s purchasing department needs to predict demand for 30-day cycles for each item to ensure that restocking occurs.

A machine learning (ML) specialist wants to use item-related features such as "category," "brand," and "safety stock count." The ML specialist also wants to use a binary time series feature that has "promotion applied?" as its name. Future promotion information is available only for the next 5 days.

The ML specialist must choose an algorithm and an evaluation metric for a solution to produce prediction results that will maximize company profit.

Which solution will meet these requirements?
  1. A Train a model by using the Autoregressive Integrated Moving Average (ARIMA) algorithm. Evaluate the model by using the Weighted Quantile Loss (wQL) metric at 0.75 (P75).
  2. B Train a model by using the Autoregressive Integrated Moving Average (ARIMA) algorithm. Evaluate the model by using the Weighted Absolute Percentage Error (WAPE) metric.
  3. C Train a model by using the Convolutional Neural Network - Quantile Regression (CNN-QR) algorithm. Evaluate the model by using the Weighted Quantile Loss (wQL) metric at 0.75 (P75).
  4. D Train a model by using the Convolutional Neural Network - Quantile Regression (CNN-QR) algorithm. Evaluate the model by using the Weighted Absolute Percentage Error (WAPE) metric.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh việc sử dụng Amazon Forecast để dự đoán nhu cầu hàng tồn kho hàng ngày cho một công ty bán lẻ. 📈

  • Yêu cầu kinh doanh chính: Chi phí hết hàng (stockout) cao hơn nhiều so với dư thừa hàng tồn kho. Do đó, mô hình cần ưu tiên dự đoán cao hơn thực tế (over-forecast) để tránh mất doanh thu, thay vì dự đoán chính xác trung bình. Dự đoán cho chu kỳ 30 ngày đối với hàng nghìn items, dựa trên dữ liệu lịch sử lớn (millions samples qua nhiều năm).
  • Features được sử dụng:
    🛠️ Item metadata (related time series không phụ thuộc thời gian): "category", "brand", "safety stock count".
    🛠️ Related time series (binary): "promotion applied?" – chỉ biết thông tin tương lai cho 5 ngày tới (Forecast hỗ trợ future known values cho related time series đến hết horizon dự đoán).
  • Mục tiêu: Chọn algorithm và evaluation metric phù hợp để tối đa hóa lợi nhuận (maximize profit), nghĩa là metric phải phản ánh chi phí bất đối xứng (asymmetric cost).
  • Bối cảnh AWS mới nhất (2026): Amazon Forecast hỗ trợ các algorithm như ARIMA (classical), CNN-QR (deep learning quantile regression), DeepAR+, ETS, v.v. Forecast cho phép quantile forecasts và custom metrics như Weighted Quantile Loss (wQL) để handle business costs.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Train a model by using the Convolutional Neural Network - Quantile Regression (CNN-QR) algorithm. Evaluate the model by using the Weighted Quantile Loss (wQL) metric at 0.75 (P75).

Lý do chi tiết:

  • 🧠 CNN-QR là algorithm deep learning chuyên cho quantile regression, hỗ trợ đầy đủ item metadata (category, brand, safety stock) và related time series (promotion). Nó excel trong dự đoán probabilistic với horizon dài (30 ngày) và dữ liệu lớn. Phù hợp vì promotion chỉ biết 5 ngày nhưng Forecast tự interpolate/extend.
  • 📊 wQL tại P75 (75th percentile): Metric này weight loss để phạt nặng hơn cho under-forecast (dự đoán thấp dẫn đến hết hàng), khớp với chi phí bất đối xứng. P75 nghĩa là dự đoán mức nhu cầu mà 75% trường hợp thực tế ≤ dự đoán → ưu tiên over-forecast để tối đa hóa lợi nhuận.
  • Kết hợp này đảm bảo mô hình tối ưu profit theo best practices AWS Forecast.

🔍 Giải thích tất cả các phương án (đúng/sai)

  • ❌ SAI: Train a model by using the Autoregressive Integrated Moving Average (ARIMA) algorithm. Evaluate the model by using the Weighted Quantile Loss (wQL) metric at 0.75 (P75).
    Lý do sai: ARIMA là classical time series KHÔNG hỗ trợ item metadata (category, brand, safety stock) hoặc related time series (promotion). Nó chỉ dùng target time series thuần túy, không tận dụng features → kém hiệu quả với dữ liệu phức tạp và hàng nghìn items. Dù wQL tốt, nhưng ARIMA không phải quantile model native trong Forecast, dẫn đến kết quả kém.

  • ❌ SAI: Train a model by using the Autoregressive Integrated Moving Average (ARIMA) algorithm. Evaluate the model by using the Weighted Absolute Percentage Error (WAPE) metric.
    Lý do sai: ARIMA không hỗ trợ metadata/related TS như trên. WAPE là metric symmetric (penalize over/under-forecast ngang nhau), không phù hợp với chi phí bất đối xứng (hết hàng đắt hơn) → không tối đa hóa profit.

  • ✅ ĐÚNG: Train a model by using the Convolutional Neural Network - Quantile Regression (CNN-QR) algorithm. Evaluate the model by using the Weighted Quantile Loss (wQL) metric at 0.75 (P75).
    Lý do đúng: Như giải thích ở phần đáp án đúng: CNN-QR hỗ trợ đầy đủ features, wQL@P75 ưu tiên tránh under-forecast → lý tưởng cho business case này.

  • ❌ SAI: Train a model by using the Convolutional Neural Network - Quantile Regression (CNN-QR) algorithm. Evaluate the model by using the Weighted Absolute Percentage Error (WAPE) metric.
    Lý do sai: CNN-QR tốt cho features và quantile, nhưng WAPE symmetric nên không handle chi phí bất đối xứng → mô hình có thể under-forecast, dẫn đến mất profit do hết hàng.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

  • Amazon Forecast Developer Guide: CNN-QR và Quantile Forecasts – Xác nhận hỗ trợ metadata/related TS và P10/P50/P90.
  • Evaluation Metrics: Weighted Quantile Loss (wQL) – Best cho asymmetric loss như stockout costs.
  • Best Practices: Forecast cho Retail Demand – Khuyến nghị CNN-QR + wQL cho inventory với promotion features.
  • AWS re:Post & Workshops: Các case study về P75 cho over-forecast priority (cập nhật Q1/2026).

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ code boto3, hãy hỏi nhé!