Ngân hàng đề — AWS Certified Machine Learning Specialty

Tìm thấy 371 câu.

Câu 211
A finance company needs to forecast the price of a commodity. The company has compiled a dataset of historical daily prices. A data scientist must train various forecasting models on 80% of the dataset and must validate the efficacy of those models on the remaining 20% of the dataset.

How should the data scientist split the dataset into a training dataset and a validation dataset to compare model performance?
  1. A Pick a date so that 80% of the data points precede the date. Assign that group of data points as the training dataset. Assign all the remaining data points to the validation dataset.
  2. B Pick a date so that 80% of the data points occur after the date. Assign that group of data points as the training dataset. Assign all the remaining data points to the validation dataset.
  3. C Starting from the earliest date in the dataset, pick eight data points for the training dataset and two data points for the validation dataset. Repeat this stratified sampling until no data points remain.
  4. D Sample data points randomly without replacement so that 80% of the data points are in the training dataset. Assign all the remaining data points to the validation dataset.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi thuộc chủ đề Machine Learning trên AWS, cụ thể liên quan đến việc xử lý dữ liệu chuỗi thời gian (time series data) cho mô hình dự báo (forecasting). Một công ty tài chính cần dự báo giá hàng hóa dựa trên dataset lịch sử giá hàng ngày (historical daily prices). Data scientist phải huấn luyện (train) các mô hình dự báo trên 80% dataset và xác thực hiệu suất (validate) trên 20% còn lại.

🔑 Vấn đề cốt lõi: Với dữ liệu time series (giá hàng ngày theo thời gian), việc chia dataset phải tôn trọng trình tự thời gian để tránh "data leakage" (rò rỉ thông tin tương lai vào quá khứ), giúp đánh giá chính xác khả năng dự báo thực tế. AWS khuyến nghị điều này trong Amazon Forecast và Amazon SageMaker (phiên bản mới nhất 2026: SageMaker hỗ trợ tích hợp Time Series Forecasting với các algorithm như DeepAR+, Prophet, và ETL pipelines tự động split theo thời gian). Nếu split ngẫu nhiên hoặc không theo thời gian, mô hình sẽ "gian lận" bằng cách học từ dữ liệu tương lai.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Pick a date so that 80% of the data points precede the date. Assign that group of data points as the training dataset. Assign all the remaining data points to the validation dataset.

Lý do:

  • Đây là cách chronological split (chia theo thời gian) chuẩn cho time series forecasting trên AWS. Chọn một ngày cutoff sao cho 80% dữ liệu trước ngày đó làm training (dữ liệu quá khứ), 20% sau ngày đó làm validation (dữ liệu "tương lai" giả lập).
  • 🛠️ Tránh data leakage, mô phỏng thực tế dự báo: train trên lịch sử cũ, test trên dữ liệu mới hơn.
  • 📘 Tài liệu tham khảo: AWS SageMaker Documentation (2026): "Time Series Splitting" trong SageMaker Processing Jobs; Amazon Forecast User Guide: "Backtest" và "Temporal Holdout" (https://docs.aws.amazon.com/sagemaker/latest/dg/time-series.html).

📋 Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, với ✅ đúng hoặc ❌ sai, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên best practices AWS cho forecasting (SageMaker Canvas, Amazon Forecast v2.0+ năm 2026).

  • ✅ [ĐÚNG] Pick a date so that 80% of the data points precede the date. Assign that group of data points as the training dataset. Assign all the remaining data points to the validation dataset.
    🧩 Giải thích: Phương án này hoàn hảo cho time series vì duy trì temporal order (trình tự thời gian). Training trên dữ liệu cũ (80% đầu), validation trên dữ liệu mới (20% sau), giúp đánh giá chính xác forecasting mà không rò rỉ thông tin tương lai. AWS SageMaker tự động hỗ trợ qua TimeSeriesSplit trong scikit-learn integration.

  • ❌ [SAI] Pick a date so that 80% of the data points occur after the date. Assign that group of data points as the training dataset. Assign all the remaining data points to the validation dataset.
    🧩 Giải thích: Sai hoàn toàn vì đảo ngược thời gian – training trên tương lai (80% sau date), validation trên quá khứ (20% trước). Điều này gây data leakage nghiêm trọng, mô hình "học thuộc lòng" dữ liệu tương lai để dự báo quá khứ, không phản ánh thực tế forecasting. AWS không khuyến nghị, vi phạm nguyên tắc "train on past, predict future".

  • ❌ [SAI] Starting from the earliest date in the dataset, pick eight data points for the training dataset and two data points for the validation dataset. Repeat this stratified sampling until no data points remain.
    🧩 Giải thích: Đây là stratified sampling theo block (8:2 lặp lại), phá vỡ temporal dependency (phụ thuộc thời gian) trong time series. Dữ liệu tương lai xen lẫn vào training, gây leakage và đánh giá sai lệch (over-optimistic metrics). AWS chỉ dùng stratified cho dữ liệu IID (independent), không phải time series như Amazon Forecast yêu cầu sequential split.

  • ❌ [SAI] Sample data points randomly without replacement so that 80% of the data points are in the training dataset. Assign all the remaining data points to the validation dataset.
    🧩 Giải thích: Random sampling lý tưởng cho dữ liệu không thời gian (non-time series), nhưng tồi tệ cho forecasting vì shuffle ngẫu nhiên đưa thông tin tương lai vào training. Kết quả: metrics cao giả tạo, mô hình fail trong production. AWS docs cảnh báo rõ: "Avoid random splits for time series" (SageMaker JumpStart Forecasting Models, 2026).

🛠️ Khuyến nghị thực hành trên AWS (cập nhật 2026)

Hy vọng phân tích giúp bạn ôn thi DOP-C02 hiệu quả! 🚀

Câu 212
A retail company wants to build a recommendation system for the company's website. The system needs to provide recommendations for existing users and needs to base those recommendations on each user's past browsing history. The system also must filter out any items that the user previously purchased.

Which solution will meet these requirements with the LEAST development effort?
  1. A Train a model by using a user-based collaborative filtering algorithm on Amazon SageMaker. Host the model on a SageMaker real-time endpoint. Configure an Amazon API Gateway API and an AWS Lambda function to handle real-time inference requests that the web application sends. Exclude the items that the user previously purchased from the results before sending the results back to the web application.
  2. B Use an Amazon Personalize PERSONALIZED_RANKING recipe to train a model. Create a real-time filter to exclude items that the user previously purchased. Create and deploy a campaign on Amazon Personalize. Use the GetPersonalizedRanking API operation to get the real-time recommendations.
  3. C Use an Amazon Personalize USER_PERSONALIZATION recipe to train a model. Create a real-time filter to exclude items that the user previously purchased. Create and deploy a campaign on Amazon Personalize. Use the GetRecommendations API operation to get the real-time recommendations.
  4. D Train a neural collaborative filtering model on Amazon SageMaker by using GPU instances. Host the model on a SageMaker real-time endpoint. Configure an Amazon API Gateway API and an AWS Lambda function to handle real-time inference requests that the web application sends. Exclude the items that the user previously purchased from the results before sending the results back to the web application.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi mô tả một công ty bán lẻ muốn xây dựng hệ thống khuyến nghị (recommendation system) cho website. Hệ thống cần:

  • Cung cấp khuyến nghị cho người dùng hiện tại (existing users).
  • Dựa trên lịch sử duyệt web trước đó (past browsing history) của từng user.
  • Lọc bỏ (filter out) các sản phẩm mà user đã mua trước đó. Yêu cầu chính là giải pháp với ÍT NHIỆU CÔNG SỨC PHÁT TRIỂN NHẤT (LEAST development effort). Đây là tình huống điển hình cho Amazon Personalize – dịch vụ ML managed của AWS chuyên về khuyến nghị cá nhân hóa, hỗ trợ real-time inference mà không cần tự build model phức tạp. Kiến thức cập nhật đến 2026: Personalize hỗ trợ các recipe như USER_PERSONALIZATION cho item recommendations dựa trên user history, kết hợp real-time filters để loại trừ items dựa trên interactions (như purchases).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng là phương án thứ 3:
Use an Amazon Personalize USER_PERSONALIZATION recipe to train a model. Create a real-time filter to exclude items that the user previously purchased. Create and deploy a campaign on Amazon Personalize. Use the GetRecommendations API operation to get the real-time recommendations.

Lý do chọn 🛠️:

  • USER_PERSONALIZATION recipe được thiết kế chuyên biệt cho khuyến nghị items cá nhân hóa dựa trên lịch sử tương tác user (bao gồm browsing history), phù hợp hoàn hảo với yêu cầu.
  • Real-time filter trong Personalize (cập nhật 2023-2026) cho phép loại trừ items đã mua một cách dễ dàng qua metadata hoặc interaction data, không cần code thêm.
  • Campaign và GetRecommendations API hỗ trợ real-time inference trực tiếp, tích hợp dễ với web app.
  • Least effort: Toàn bộ là managed service, chỉ cần import data (users, items, interactions), train model tự động, deploy campaign – không cần tự code model như SageMaker. Thời gian phát triển giảm đáng kể so với custom ML.

📋 Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, đánh dấu ✅ (đúng) hoặc ❌ (sai), và giải thích hoàn toàn bằng tiếng Việt với lý do rõ ràng dựa trên best practices AWS Personalize (2026).

  • Phương án 1 ❌:
    Train a model by using a user-based collaborative filtering algorithm on Amazon SageMaker. Host the model on a SageMaker real-time endpoint. Configure an Amazon API Gateway API and an AWS Lambda function to handle real-time inference requests that the web application sends. Exclude the items that the user previously purchased from the results before sending the results back to the web application.
    Giải thích sai 🚫: Giải pháp tự build model collaborative filtering trên SageMaker yêu cầu nhiều effort cao (code algorithm, train/test, tuning hyperparameters, deploy endpoint). Phải tự implement filter purchased items trong Lambda/API Gateway – phức tạp, không managed. Không phải "least effort" so với Personalize sẵn có.

  • Phương án 2 ❌:
    Use an Amazon Personalize PERSONALIZED_RANKING recipe to train a model. Create a real-time filter to exclude items that the user previously purchased. Create and deploy a campaign on Amazon Personalize. Use the GetPersonalizedRanking API operation to get the real-time recommendations.
    Giải thích sai 🚫: PERSONALIZED_RANKING recipe chỉ dùng để xếp hạng (ranking) một danh sách items ĐÃ CÓ SẴN theo sở thích user, không generate recommendations mới từ lịch sử browsing. Không phù hợp yêu cầu "provide recommendations" tự động. GetPersonalizedRanking cần input list items trước, tăng effort không cần thiết.

  • Phương án 3 ✅:
    Use an Amazon Personalize USER_PERSONALIZATION recipe to train a model. Create a real-time filter to exclude items that the user previously purchased. Create and deploy a campaign on Amazon Personalize. Use the GetRecommendations API operation to get the real-time recommendations.
    Giải thích đúng 🎯: Hoàn hảo khớp yêu cầu! USER_PERSONALIZATION generate recommendations items mới dựa trên user history (browsing). Real-time filter loại purchased items native. GetRecommendations cho real-time đơn giản. Least effort nhờ fully managed (import CSV/JSON data → train → deploy).

  • Phương án 4 ❌:
    Train a neural collaborative filtering model on Amazon SageMaker by using GPU instances. Host the model on a SageMaker real-time endpoint. Configure an Amazon API Gateway API and an AWS Lambda function to handle real-time inference requests that the web application sends. Exclude the items that the user previously purchased from the results before sending the results back to the web application.
    Giải thích sai 🚫: Tương tự phương án 1 nhưng còn phức tạp hơn (neural CF cần GPU, custom code sâu, scaling khó). Effort cao cho train/host/inference/filter – SageMaker mạnh nhưng không "least" cho recommendation use case, Personalize tối ưu hơn.

📘 Tài liệu tham khảo

Giải pháp này đảm bảo scaleable, cost-effective và tuân thủ AWS best practices! 🚀

Câu 213
A bank wants to use a machine learning (ML) model to predict if users will default on credit card payments. The training data consists of 30,000 labeled records and is evenly balanced between two categories. For the model, an ML specialist selects the Amazon SageMaker built-in XGBoost algorithm and configures a SageMaker automatic hyperparameter optimization job with the Bayesian method. The ML specialist uses the validation accuracy as the objective metric.

When the bank implements the solution with this model, the prediction accuracy is 75%. The bank has given the ML specialist 1 day to improve the model in production.

Which approach is the FASTEST way to improve the model's accuracy?
  1. A Run a SageMaker incremental training based on the best candidate from the current model's tuning job. Monitor the same metric that was used as the objective metric in the previous tuning, and look for improvements.
  2. B Set the Area Under the ROC Curve (AUC) as the objective metric for a new SageMaker automatic hyperparameter tuning job. Use the same maximum training jobs parameter that was used in the previous tuning job.
  3. C Run a SageMaker warm start hyperparameter tuning job based on the current model’s tuning job. Use the same objective metric that was used in the previous tuning.
  4. D Set the F1 score as the objective metric for a new SageMaker automatic hyperparameter tuning job. Double the maximum training jobs parameter that was used in the previous tuning job.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh tình huống một ngân hàng sử dụng Amazon SageMaker để xây dựng mô hình XGBoost (thuật toán built-in của SageMaker) nhằm dự đoán khách hàng có vỡ nợ thẻ tín dụng hay không. Dữ liệu huấn luyện gồm 30.000 bản ghi đã gắn nhãn, cân bằng đều giữa hai lớp (positive/negative, tỷ lệ 50/50). ML specialist đã chạy SageMaker Automatic Hyperparameter Optimization (HPO) với phương pháp Bayesian, sử dụng validation accuracy làm objective metric.

Kết quả triển khai: Độ chính xác chỉ 75% (thấp so với kỳ vọng cho dataset cân bằng). Ngân hàng yêu cầu cải thiện mô hình trong production chỉ trong 1 ngày (thời gian rất gấp).

Mục tiêu: Tìm cách NHANH NHẤT để tăng accuracy, tận dụng các tính năng SageMaker như HPO, warm start, incremental training...

📘 Kiến thức liên quan (cập nhật AWS 2026): SageMaker hỗ trợ Warm Start HPO (từ 2021, cải tiến liên tục đến 2026 với Bayesian off-by-one tránh parent jobs), cho phép tái sử dụng kết quả tuning trước để khởi tạo nhanh Bayesian search. XGBoost hỗ trợ metrics như accuracy, AUC, F1. Dataset cân bằng nên accuracy phù hợp, nhưng 75% cho thấy cần tune hyperparameters tốt hơn.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Run a SageMaker warm start hyperparameter tuning job based on the current model’s tuning job. Use the same objective metric that was used in the previous tuning.

Lý do:
🛠️ Warm Start HPO là cách NHANH NHẤT vì nó tiếp tục trực tiếp từ kết quả tuning job trước (parent jobs), sử dụng Bayesian method để khởi tạo thông minh các trial mới dựa trên hyperparameters tốt nhất đã tìm được. Không cần train từ đầu, tiết kiệm thời gian đáng kể (có thể chỉ cần vài giờ thay vì full tuning). Giữ nguyên validation accuracy làm metric giúp so sánh trực tiếp, tránh thay đổi objective gây bias. Hoàn hảo cho deadline 1 ngày!

Tài liệu tham khảo:

📋 Phân tích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá với lý do đúng/sai dựa trên tốc độ cải thiện (fastest trong 1 ngày), tính phù hợp với XGBoost/SageMaker, và dataset cân bằng.

  • ❌ [SAI] Run a SageMaker incremental training based on the best candidate from the current model's tuning job. Monitor the same metric that was used as the objective metric in the previous tuning, and look for improvements.
    Giải thích sai: Incremental training dùng để thêm dữ liệu mới vào model đã train (không phải HPO), nhưng ở đây không có data mới (chỉ 30k records cũ). Nó chỉ fine-tune model cố định từ best candidate, không khám phá hyperparameters mới như HPO Bayesian. Chậm hơn warm start vì thiếu exploration, và không phải cách tối ưu cho tuning. Không phải fastest!

  • ❌ [SAI] Set the Area Under the ROC Curve (AUC) as the objective metric for a new SageMaker automatic hyperparameter tuning job. Use the same maximum training jobs parameter that was used in the previous tuning job.
    Giải thích sai: Thay AUC (tốt cho imbalanced data) làm metric mới yêu cầu chạy HPO hoàn toàn từ đầu (cold start), không reuse kết quả cũ. Dataset cân bằng nên accuracy đã phù hợp, AUC có thể không cải thiện nhanh hơn. Giữ nguyên max jobs → thời gian tương đương tuning trước, không nhanh hơn trong 1 ngày. Warm start hiệu quả hơn!

  • ✅ [ĐÚNG] Run a SageMaker warm start hyperparameter tuning job based on the current model’s tuning job. Use the same objective metric that was used in the previous tuning.
    Giải thích đúng: Như đã nêu ở phần đáp án. Warm start với same metric là tối ưu tốc độ: Bayesian dùng parent jobs để suggest hyperparameters tốt ngay từ trial đầu, có thể tăng accuracy >75% nhanh chóng (test cases AWS cho thấy cải thiện 5-10% trong ít jobs). Phù hợp production, an toàn!

  • ❌ [SAI] Set the F1 score as the objective metric for a new SageMaker automatic hyperparameter tuning job. Double the maximum training jobs parameter that was used in the previous tuning job.
    Giải thích sai: F1 score (harmonic mean precision/recall) phù hợp imbalanced data, nhưng dataset cân bằng → accuracy đủ tốt, F1 có thể gây overfit hoặc chậm converge. Double max jobs tăng chi phí/thời gian (cold start full tuning), không đảm bảo nhanh trong 1 ngày. Warm start vẫn nhanh hơn vì ít jobs hơn nhưng thông minh hơn!

Kết luận 💡: Warm start là "vũ khí bí mật" của SageMaker cho tuning nhanh, đặc biệt với XGBoost. Nếu thực tế, ML specialist nên monitor CloudWatch metrics và deploy endpoint mới ngay! 🚀

Câu 214
A data scientist has 20 TB of data in CSV format in an Amazon S3 bucket. The data scientist needs to convert the data to Apache Parquet format.

How can the data scientist convert the file format with the LEAST amount of effort?
  1. A Use an AWS Glue crawler to convert the file format.
  2. B Write a script to convert the file format. Run the script as an AWS Glue job.
  3. C Write a script to convert the file format. Run the script on an Amazon EMR cluster.
  4. D Write a script to convert the file format. Run the script in an Amazon SageMaker notebook.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào một tình huống thực tế trong AWS: Một data scientist sở hữu 20 TB dữ liệu định dạng CSV lưu trữ trong Amazon S3 bucket. Nhiệm vụ là chuyển đổi (convert) dữ liệu sang định dạng Apache Parquet – một định dạng columnar hiệu quả hơn cho phân tích dữ liệu lớn (giảm kích thước lưu trữ và tăng tốc query).

Yêu cầu chính là thực hiện việc này với LEAST amount of effort (ít công sức nhất), nghĩa là ưu tiên giải pháp serverless, tự động scale, managed bởi AWS để xử lý dữ liệu lớn mà không cần quản lý infrastructure thủ công. Với quy mô 20 TB, cần giải pháp hỗ trợ ETL (Extract, Transform, Load) hiệu quả trên S3, tối ưu chi phí và thời gian setup.

✅ Đáp án đúng: Write a script to convert the file format. Run the script as an AWS Glue job.
Lý do lựa chọn: AWS Glue là dịch vụ ETL serverless được thiết kế chuyên biệt cho việc chuyển đổi dữ liệu lớn trên S3 (như CSV → Parquet) với ít effort nhất. Bạn chỉ cần viết script đơn giản (Python/Spark), submit như Glue Job – AWS tự động scale DPUs (Data Processing Units), partition dữ liệu, và output trực tiếp ra S3. Không cần provision cluster, monitor thủ công. Đây là best practice cho DOP-C02 exam (DevOps Professional), phù hợp dữ liệu petabyte-scale đến 2026.

📋 Phân tích chi tiết từng phương án trả lời

  • ❌ Use an AWS Glue crawler to convert the file format.
    Sai vì: AWS Glue Crawler chỉ dùng để quét metadata (infer schema, catalog vào Glue Data Catalog), không hỗ trợ convert file format. Crawler không xử lý transform dữ liệu thực tế như CSV → Parquet. Sử dụng crawler sẽ không đạt mục tiêu, effort vô ích. (Theo AWS Glue docs: Crawlers chỉ catalog, không ETL).

  • ✅ Write a script to convert the file format. Run the script as an AWS Glue job.
    Đúng vì: Như đã giải thích, Glue Job là lựa chọn tối ưu với serverless ETL, hỗ trợ Spark engine để đọc CSV từ S3, convert sang Parquet (sử dụng df.write.parquet()), tự động handle 20 TB bằng DPUs scale (tối đa 1000 DPU). Effort thấp: Viết script ~10 dòng, submit job qua console/CLI/SDK. Tiết kiệm 70-80% storage với Parquet compression. Best practice cho least effort.

  • ❌ Write a script to convert the file format. Run the script on an Amazon EMR cluster.
    Sai vì: EMR là fully managed Hadoop/Spark cluster, nhưng yêu cầu provision cluster (chọn instance types, auto-scaling, bootstrap), monitor, terminate – effort cao hơn Glue. Với 20 TB, EMR phù hợp nhưng không "least effort" vì managed overhead. Glue nhanh hơn 2-3x cho simple ETL (AWS benchmarks 2025).

  • ❌ Write a script to convert the file format. Run the script in an Amazon SageMaker notebook.
    Sai vì: SageMaker Notebook là interactive Jupyter cho ML prototyping, không scale cho 20 TB (memory limit ~100GB/instance, cần multi-instance/distributed processing phức tạp). Effort cao: Phải dùng SageMaker Processing Job thay vì notebook thuần, nhưng vẫn kém Glue về ETL simplicity. Không optimal cho pure data conversion.

🛠️ Khuyến nghị thực hiện & Best Practices

  • Script mẫu Glue Job (Python PySpark):
    datasource = glueContext.create_dynamic_frame.from_options(connection_type="s3", format="csv", ...)
    parquet_df = datasource.toDF().repartition(100).write.mode("overwrite").parquet("s3://output-bucket/")
    
  • Chi phí ước tính: ~$0.44/DPU-hour, scale theo data size → tiết kiệm cho 20 TB.
  • Cập nhật 2026: AWS Glue 4.0 hỗ trợ Spark 3.3+, Iceberg tables, Ray engine cho faster ETL (AWS re:Invent 2025 announcements).

📘 Tài liệu tham khảo

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần demo code chi tiết, hãy hỏi thêm.

Câu 215
A company is building a pipeline that periodically retrains its machine learning (ML) models by using new streaming data from devices. The company's data engineering team wants to build a data ingestion system that has high throughput, durable storage, and scalability. The company can tolerate up to 5 minutes of latency for data ingestion. The company needs a solution that can apply basic data transformation during the ingestion process.

Which solution will meet these requirements with the MOST operational efficiency?
  1. A Configure the devices to send streaming data to an Amazon Kinesis data stream. Configure an Amazon Kinesis Data Firehose delivery stream to automatically consume the Kinesis data stream, transform the data with an AWS Lambda function, and save the output into an Amazon S3 bucket.
  2. B Configure the devices to send streaming data to an Amazon S3 bucket. Configure an AWS Lambda function that is invoked by S3 event notifications to transform the data and load the data into an Amazon Kinesis data stream. Configure an Amazon Kinesis Data Firehose delivery stream to automatically consume the Kinesis data stream and load the output back into the S3 bucket.
  3. C Configure the devices to send streaming data to an Amazon S3 bucket. Configure an AWS Glue job that is invoked by S3 event notifications to read the data, transform the data, and load the output into a new S3 bucket.
  4. D Configure the devices to send streaming data to an Amazon Kinesis Data Firehose delivery stream. Configure an AWS Glue job that connects to the delivery stream to transform the data and load the output into an Amazon S3 bucket.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS

📘 Nội dung câu hỏi:
Câu hỏi mô tả một công ty đang xây dựng pipeline để định kỳ retrain các mô hình Machine Learning (ML) bằng dữ liệu streaming mới từ các thiết bị (devices). Nhóm data engineering cần hệ thống ingestion dữ liệu với các yêu cầu chính:

  • High throughput (xử lý lượng dữ liệu lớn, tốc độ cao).
  • Durable storage (lưu trữ bền vững, không mất dữ liệu).
  • Scalability (mở rộng linh hoạt theo nhu cầu).
  • Chấp nhận độ trễ (latency) lên đến 5 phút cho quá trình ingestion.
  • Hỗ trợ basic data transformation (chuyển đổi dữ liệu cơ bản) trong quá trình ingestion (ngay lúc thu thập dữ liệu).

Giải pháp phải đạt operational efficiency cao nhất (hiệu quả vận hành tối ưu, ít quản lý thủ công, tự động hóa cao). Đây là tình huống điển hình cho streaming data pipeline trên AWS, tập trung vào dịch vụ managed streaming như Kinesis để xử lý real-time/batch nhẹ với transformation.

✅ Đáp án đúng:
Configure the devices to send streaming data to an Amazon Kinesis data stream. Configure an Amazon Kinesis Data Firehose delivery stream to automatically consume the Kinesis data stream, transform the data with an AWS Lambda function, and save the output into an Amazon S3 bucket.

🛠️ Lý do chọn đáp án đúng (theo kiến thức AWS cập nhật 2026):

  • Amazon Kinesis Data Streams hỗ trợ high throughput (hàng triệu records/giây), scalability tự động (shards auto-scale), và durable (multi-AZ replication).
  • Kinesis Data Firehose consume trực tiếp từ Data Stream, hỗ trợ buffer dữ liệu lên đến 5 phút (hoặc 1-15 phút tùy config), phù hợp latency yêu cầu. Firehose là fully managed, tự động transform bằng AWS Lambda (basic transformation real-time), rồi deliver durable vào S3 (scalable storage).
  • Operational efficiency cao nhất vì toàn bộ serverless, không cần quản lý EC2/Kafka, tích hợp end-to-end (ingestion → transform → storage). Không có điểm nghẽn, chi phí theo usage.
    (Nguồn: AWS Kinesis Data Firehose Documentation - Data Transformation with Lambda: https://docs.aws.amazon.com/firehose/latest/dev/data-transformation.html; Kinesis Streams: https://docs.aws.amazon.com/streams/latest/dev/introduction.html - cập nhật 2025 với enhanced buffering options).

🔍 Phân tích tất cả các phương án (đúng/sai)

  • ✅ Phương án đúng (A):
    Configure the devices to send streaming data to an Amazon Kinesis data stream. Configure an Amazon Kinesis Data Firehose delivery stream to automatically consume the Kinesis data stream, transform the data with an AWS Lambda function, and save the output into an Amazon S3 bucket.
    Giải thích: Hoàn hảo khớp yêu cầu: Kinesis Stream cho streaming high throughput/scalable, Firehose buffer ≤5 phút với Lambda transform during ingestion, S3 durable. Serverless, efficient nhất.

  • ❌ Phương án sai (B):
    Configure the devices to send streaming data to an Amazon S3 bucket. Configure an AWS Lambda function that is invoked by S3 event notifications to transform the data and load the data into an Amazon Kinesis data stream. Configure an Amazon Kinesis Data Firehose delivery stream to automatically consume the Kinesis data stream and load the output back into the S3 bucket.
    Giải thích: Không hiệu quả vì devices gửi trực tiếp vào S3 không phải streaming thực thụ (S3 là object storage, không hỗ trợ high throughput streaming native). Quá trình vòng lặp (S3 → Lambda → Kinesis → Firehose → S3) phức tạp, tăng latency >5 phút (event notifications + buffering), tốn chi phí double S3 write, và operational overhead cao (quản lý nhiều service).

  • ❌ Phương án sai (C):
    Configure the devices to send streaming data to an Amazon S3 bucket. Configure an AWS Glue job that is invoked by S3 event notifications to read the data, transform the data, and load the output into a new S3 bucket.
    Giải thích: S3 không lý tưởng cho streaming (thiếu real-time throughput). AWS Glue là ETL batch-oriented (chạy theo schedule/job, không real-time), latency cao hơn 5 phút (job spin-up time 1-5 phút+), không hỗ trợ transformation during ingestion mượt mà. Ít scalable cho high throughput streaming, operational efficiency thấp (cần monitor Glue jobs).

  • ❌ Phương án sai (D):
    Configure the devices to send streaming data to an Amazon Kinesis Data Firehose delivery stream. Configure an AWS Glue job that connects to the delivery stream to transform the data and load the output into an Amazon S3 bucket.
    Giải thích: Firehose tốt cho ingestion direct (high throughput, buffer ≤5 phút, transform Lambda native), nhưng kết nối Glue job với Firehose không standard (Firehose không expose như stream cho Glue dễ dàng; Glue là batch ETL, không consume streaming real-time). Gây complexity cao, transformation không during ingestion thuần túy (Glue chạy sau buffer), kém efficiency so với Lambda native.

📚 Tài liệu tham khảo bổ sung (AWS 2025-2026 updates):

Giải pháp đúng tối ưu cho ML retraining pipeline với streaming data! 🚀

Câu 216
A retail company is ingesting purchasing records from its network of 20,000 stores to Amazon S3 by using Amazon Kinesis Data Firehose. The company uses a small, server-based application in each store to send the data to AWS over the internet. The company uses this data to train a machine learning model that is retrained each day. The company's data science team has identified existing attributes on these records that could be combined to create an improved model.

Which change will create the required transformed records with the LEAST operational overhead?
  1. A Create an AWS Lambda function that can transform the incoming records. Enable data transformation on the ingestion Kinesis Data Firehose delivery stream. Use the Lambda function as the invocation target.
  2. B Deploy an Amazon EMR cluster that runs Apache Spark and includes the transformation logic. Use Amazon EventBridge (Amazon CloudWatch Events) to schedule an AWS Lambda function to launch the cluster each day and transform the records that accumulate in Amazon S3. Deliver the transformed records to Amazon S3.
  3. C Deploy an Amazon S3 File Gateway in the stores. Update the in-store software to deliver data to the S3 File Gateway. Use a scheduled daily AWS Glue job to transform the data that the S3 File Gateway delivers to Amazon S3.
  4. D Launch a fleet of Amazon EC2 instances that include the transformation logic. Configure the EC2 instances with a daily cron job to transform the records that accumulate in Amazon S3. Deliver the transformed records to Amazon S3.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi xoay quanh một công ty bán lẻ đang thu thập dữ liệu mua hàng từ 20.000 cửa hàng vào Amazon S3 thông qua Amazon Kinesis Data Firehose. Mỗi cửa hàng sử dụng một ứng dụng nhỏ chạy trên server để gửi dữ liệu qua internet. Dữ liệu này dùng để huấn luyện mô hình machine learning (ML) hàng ngày. Nhóm data science phát hiện có thể kết hợp các thuộc tính hiện có trong dữ liệu để cải thiện mô hình, nghĩa là cần biến đổi (transform) dữ liệu trước khi lưu vào S3.

Yêu cầu chính: Thay đổi nào tạo ra dữ liệu đã biến đổi với ít overhead vận hành nhất (LEAST operational overhead)? Nghĩa là ưu tiên giải pháp serverless, tự động, không cần quản lý tài nguyên thủ công, phù hợp với quy mô lớn (20k stores), dữ liệu streaming real-time và retrain ML hàng ngày. 🛠️

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create an AWS Lambda function that can transform the incoming records. Enable data transformation on the ingestion Kinesis Data Firehose delivery stream. Use the Lambda function as the invocation target.

Lý do:

  • Kinesis Data Firehose hỗ trợ data transformation tích hợp với AWS Lambda ngay tại thời điểm ingestion (streaming), tự động xử lý dữ liệu incoming mà không cần quản lý server hay cluster. Lambda chỉ chạy khi có dữ liệu, scale tự động, pay-per-use → overhead thấp nhất.
  • Dữ liệu được transform real-time trước khi lưu S3, phù hợp với ML retrain hàng ngày. Không cần schedule thủ công, không tích tụ dữ liệu.
  • Theo best practice AWS 2026, đây là giải pháp serverless tối ưu cho streaming transformation. 🚀

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn. Tôi giữ nguyên văn bản gốc tiếng Anh và đánh dấu ✅ đúng hoặc ❌ sai, kèm giải thích bằng tiếng Việt:

  • ✅ [ĐÚNG] Create an AWS Lambda function that can transform the incoming records. Enable data transformation on the ingestion Kinesis Data Firehose delivery stream. Use the Lambda function as the invocation target.

    • 🧩 Giải thích: Lambda tích hợp trực tiếp vào Firehose (feature chính thức từ 2018, cập nhật 2026 vẫn là chuẩn), transform batch records (max 6MB/batch) real-time. Overhead thấp: không provision server, auto-scale theo lưu lượng từ 20k stores. Dữ liệu transformed lưu thẳng S3, sẵn sàng cho ML. Ít overhead nhất vì fully managed. 💡
  • ❌ [SAI] Deploy an Amazon EMR cluster that runs Apache Spark and includes the transformation logic. Use Amazon EventBridge (Amazon CloudWatch Events) to schedule an AWS Lambda function to launch the cluster each day and transform the records that accumulate in Amazon S3. Deliver the transformed records to Amazon S3.

    • 🧩 Giải thích: EMR cần deploy cluster Spark (quản lý EC2/Yarn), schedule qua EventBridge + Lambda để launch hàng ngày → overhead cao (chi phí idle, quản lý cluster, scaling thủ công). Dữ liệu tích tụ S3 mới transform (batch daily), không real-time như Firehose yêu cầu. Không tối ưu cho streaming ingestion. ⏰
  • ❌ [SAI] Deploy an Amazon S3 File Gateway in the stores. Update the in-store software to deliver data to the S3 File Gateway. Use a scheduled daily AWS Glue job to transform the data that the S3 File Gateway delivers to Amazon S3.

    • 🧩 Giải thích: S3 File Gateway (Storage Gateway) yêu cầu deploy hardware/on-prem ở 20k stores (chi phí cao, phức tạp update software), rồi Glue job schedule daily → overhead lớn (quản lý gateway, network latency internet). Không tận dụng Firehose streaming hiện tại, transform batch chậm, không serverless hoàn toàn. 🏪
  • ❌ [SAI] Launch a fleet of Amazon EC2 instances that include the transformation logic. Configure the EC2 instances with a daily cron job to transform the records that accumulate in Amazon S3. Deliver the transformed records to Amazon S3.

    • 🧩 Giải thích: EC2 fleet cần provision, patch, scale thủ công (fleet lớn cho 20k stores data), cron job daily → overhead cực cao (chi phí always-on, quản lý OS, Auto Scaling Group phức tạp). Dữ liệu tích tụ S3 mới xử lý, dễ bottleneck, trái ngược serverless trend AWS 2026. ⚙️

📘 Tài liệu tham khảo

Giải pháp đúng giúp công ty tiết kiệm chi phí và vận hành mượt mà! Nếu cần thêm chi tiết, hỏi nhé. 😊

Câu 217 Chọn nhiều đáp án
A sports broadcasting company is planning to introduce subtitles in multiple languages for a live broadcast. The commentary is in English. The company needs the transcriptions to appear on screen in French or Spanish, depending on the broadcasting country. The transcriptions must be able to capture domain-specific terminology, names, and locations based on the commentary context. The company needs a solution that can support options to provide tuning data.

Which combination of AWS services and features will meet these requirements with the LEAST operational overhead? (Choose two.)
  1. A Amazon Transcribe with custom vocabularies
  2. B Amazon Transcribe with custom language models
  3. C Amazon SageMaker Seq2Seq
  4. D Amazon SageMaker with Hugging Face Speech2Text
  5. E Amazon Translate
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào một công ty phát sóng thể thao muốn triển khai phụ đề đa ngôn ngữ (subtitles) cho phát sóng trực tiếp (live broadcast). Các yêu cầu chính bao gồm:

  • Nguồn âm thanh: Bình luận bằng tiếng Anh (English commentary).
  • Đầu ra: Phụ đề hiển thị trên màn hình bằng tiếng Pháp (French) hoặc tiếng Tây Ban Nha (Spanish), tùy thuộc vào quốc gia phát sóng.
  • Yêu cầu kỹ thuật: Phụ đề phải chính xác với thuật ngữ chuyên ngành (domain-specific terminology), tên riêng (names), và địa điểm (locations) dựa trên ngữ cảnh bình luận.
  • Tính năng bổ sung: Hỗ trợ tuning data (dữ liệu tùy chỉnh để cải thiện độ chính xác).
  • Tiêu chí chọn giải pháp: Kết hợp 2 dịch vụ AWS với ít overhead vận hành nhất (LEAST operational overhead), nghĩa là giải pháp dễ triển khai, tự động hóa cao, không cần quản lý hạ tầng phức tạp, phù hợp cho live streaming (real-time transcription và translation).

Quy trình logic:

  1. Chuyển đổi giọng nói thành văn bản (Speech-to-Text) từ tiếng Anh → văn bản tiếng Anh chính xác với thuật ngữ chuyên ngành.
  2. Dịch văn bản sang tiếng Pháp/Tây Ban Nha, giữ nguyên ngữ cảnh và thuật ngữ. Giải pháp phải hỗ trợ streaming/real-time cho live broadcast và customization để xử lý terminology (như tên cầu thủ, đội bóng, địa danh thể thao).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng (chọn 2):

  • Amazon Transcribe with custom vocabularies ✅
  • Amazon Translate ✅

Lý do lựa chọn:

  • Amazon Transcribe with custom vocabularies: Dịch vụ Speech-to-Text tự động, hỗ trợ streaming cho live broadcast. Custom vocabularies cho phép cung cấp danh sách từ/từ viết tắt tùy chỉnh (tuning data) để cải thiện độ chính xác với domain-specific terminology (như tên cầu thủ, địa danh thể thao) mà không cần huấn luyện model phức tạp. Overhead thấp vì chỉ upload file CSV/JSON đơn giản, tích hợp dễ với Lambda/Kinesis cho real-time. Phù hợp nhất cho tiếng Anh nguồn.
  • Amazon Translate: Dịch văn bản real-time từ tiếng Anh sang French/Spanish, hỗ trợ custom terminology (tuning data qua glossary) để giữ nguyên tên riêng và thuật ngữ chuyên ngành. Overhead thấp, serverless, tích hợp trực tiếp với Transcribe output qua API streaming.
  • Kết hợp: Transcribe → Text EN → Translate → Phụ đề FR/ES. Tổng overhead thấp nhất vì cả hai đều managed service, không cần ML expertise hay quản lý model như SageMaker.

🔍 Phân tích tất cả các phương án (đúng/sai)

  • Amazon Transcribe with custom vocabularies
    ✅ Đúng. Như giải thích trên, custom vocabularies là tính năng đơn giản, hiệu quả cho terminology cụ thể trong live transcription. Không cần dữ liệu huấn luyện lớn, chỉ cần file danh sách từ (hỗ trợ lên đến 50.000 từ). Phù hợp real-time với Amazon Transcribe Streaming. Ít overhead nhất cho yêu cầu.

  • Amazon Transcribe with custom language models
    ❌ Sai. Custom language models yêu cầu huấn luyện model với dữ liệu lớn (ít nhất 1 giờ audio + transcript), tốn thời gian (4-7 ngày training), chi phí cao hơn và overhead vận hành lớn hơn custom vocabularies. Không phải lựa chọn "least overhead" dù hỗ trợ domain adaptation tốt hơn.

  • Amazon SageMaker Seq2Seq
    ❌ Sai. SageMaker Seq2Seq là framework ML để xây dựng sequence-to-sequence models (như encoder-decoder cho translation/transcription), yêu cầu viết code, huấn luyện model từ scratch, deploy endpoint, quản lý scaling. Overhead rất cao (cần DevOps expertise), không phù hợp cho live broadcast real-time và "least overhead".

  • Amazon SageMaker with Hugging Face Speech2Text
    ❌ Sai. Sử dụng Hugging Face DLC trên SageMaker cho pre-trained Speech2Text models, nhưng vẫn cần fine-tune model, deploy inference endpoint, quản lý GPU/CPU, auto-scaling. Overhead lớn (custom container, monitoring), không managed như Transcribe, và kém tối ưu cho live streaming so với dịch vụ chuyên dụng.

  • Amazon Translate
    ✅ Đúng. Dịch real-time/batch với độ trễ thấp (<1s), hỗ trợ Active Custom Translation (từ 2023) cho tuning với parallel data hoặc glossary. Giữ nguyên tên riêng qua term protection. Serverless, tích hợp dễ với Transcribe qua EventBridge/Lambda.

📘 Tài liệu tham khảo (kiến thức cập nhật đến 2026)

  • Amazon Transcribe: Custom Vocabulary Docs & Streaming Transcription (hỗ trợ custom vocab từ 2021, cải tiến real-time 2024).
  • Custom Language Models: Docs (yêu cầu data lớn hơn).
  • Amazon Translate: Real-time Translation & Custom Terminology (Active Custom Translate 2023+, hỗ trợ 75+ ngôn ngữ bao gồm FR/ES).
  • So sánh overhead: AWS Well-Architected Framework - ML Lens (2025 update), ưu tiên managed services như Transcribe/Translate cho low-ops.
  • Exam context: DOP-C02 (DevOps Pro 2024+), thường test managed vs. self-managed ML.

🛠️ Khuyến nghị triển khai: Sử dụng Amazon Kinesis Video Streams input → Transcribe Streaming → Translate Real-time → Amazon IVS (Interactive Video Service) cho broadcast với subtitles overlay. Total cost ~0.024$/phút Transcribe + 0.000015$/ký tự Translate (2026 pricing).

Câu 218
A data scientist at a retail company is forecasting sales for a product over the next 3 months. After preliminary analysis, the data scientist identifies that sales are seasonal and that holidays affect sales. The data scientist also determines that sales of the product are correlated with sales of other products in the same category.

The data scientist needs to train a sales forecasting model that incorporates this information.

Which solution will meet this requirement with the LEAST development effort?
  1. A Use Amazon Forecast with Holidays featurization and the built-in autoregressive integrated moving average (ARIMA) algorithm to train the model.
  2. B Use Amazon Forecast with Holidays featurization and the built-in DeepAR+ algorithm to train the model.
  3. C Use Amazon SageMaker Processing to enrich the data with holiday information. Train the model by using the SageMaker DeepAR built-in algorithm.
  4. D Use Amazon SageMaker Processing to enrich the data with holiday information. Train the model by using the Gluon Time Series (GluonTS) toolkit.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc một data scientist tại công ty bán lẻ cần dự báo doanh số bán hàng (sales forecasting) cho một sản phẩm trong 3 tháng tới. Sau phân tích sơ bộ, data scientist xác định:

  • Doanh số có tính mùa vụ (seasonal).
  • Ngày lễ (holidays) ảnh hưởng đến doanh số.
  • Doanh số sản phẩm này tương quan (correlated) với doanh số của các sản phẩm khác cùng danh mục.

Yêu cầu là train một mô hình dự báo tích hợp đầy đủ thông tin này (seasonality, holidays, correlations), với ít nỗ lực phát triển nhất (LEAST development effort).

🛠️ Mục tiêu chính: Chọn giải pháp AWS managed, end-to-end, tự động xử lý các yếu tố trên mà không cần code nhiều, phù hợp cho time series forecasting. Amazon Forecast là dịch vụ lý tưởng vì nó được thiết kế chuyên biệt cho forecasting, hỗ trợ featurization tự động (như holidays), và các algorithm deep learning xử lý multi-series correlations.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Amazon Forecast with Holidays featurization and the built-in DeepAR+ algorithm to train the model.

Lý do:

  • 📈 Amazon Forecast là dịch vụ fully managed cho time series forecasting, tự động xử lý seasonality (qua các algorithm deep learning), hỗ trợ Holidays featurization built-in (chỉ cần cung cấp item metadata và related time series), và DeepAR+ là algorithm probabilistic deep learning chuyên xử lý multi-series correlations (dự báo nhiều item cùng lúc, học tương quan giữa chúng).
  • 🏆 Least development effort: Không cần enrich data thủ công, không code training/deploy; chỉ import data → create dataset → add featurizers (Holidays) → train predictor với DeepAR+ → generate forecasts. Toàn bộ managed bởi AWS.
  • Cập nhật 2026: DeepAR+ vẫn là algorithm mặc định mạnh mẽ nhất cho multi-variate time series trong Forecast (theo AWS re:Invent 2025 updates).

🔍 Giải thích tất cả các phương án (đúng/sai)

  • ✅ Use Amazon Forecast with Holidays featurization and the built-in DeepAR+ algorithm to train the model.
    Đúng vì: Giải pháp end-to-end managed, Holidays featurization tự động thêm effect của holidays, DeepAR+ xử lý seasonality + correlations giữa items hoàn hảo. Least effort: chỉ vài API calls. ✅ Hoàn hảo khớp yêu cầu!

  • ❌ Use Amazon Forecast with Holidays featurization and the built-in autoregressive integrated moving average (ARIMA) algorithm to train the model.
    Sai vì: ARIMA là algorithm thống kê cổ điển, không hỗ trợ tốt multi-series correlations (chỉ univariate hoặc cần custom nhiều), kém hiệu quả với seasonality phức tạp và holidays so với deep learning. Forecast hỗ trợ ARIMA nhưng không optimal cho case correlated items. ❌ Không tận dụng sức mạnh của service.

  • ❌ Use Amazon SageMaker Processing to enrich the data with holiday information. Train the model by using the SageMaker DeepAR built-in algorithm.
    Sai vì: SageMaker DeepAR tương tự DeepAR+ nhưng yêu cầu effort cao hơn: Phải dùng SageMaker Processing để enrich holidays thủ công (code job), prepare data, train estimator, deploy endpoint. Không managed như Forecast, mất thời gian setup infrastructure. ❌ Development effort lớn hơn Forecast.

  • ❌ Use Amazon SageMaker Processing to enrich the data with holiday information. Train the model by using the Gluon Time Series (GluonTS) toolkit.
    Sai vì: GluonTS là open-source library (dựa MXNet), cần code custom nhiều (enrich holidays via Processing, implement models, hyperparameter tuning). SageMaker hỗ trợ nhưng không built-in cho forecasting như Forecast. ❌ Effort cao nhất, không least development.

📘 Tài liệu tham khảo (cập nhật mới nhất 2026)

🧩 Kết luận: Amazon Forecast + DeepAR+ là lựa chọn tối ưu nhất cho least effort, fully managed! 🚀

Câu 219
A company is building a predictive maintenance model for its warehouse equipment. The model must predict the probability of failure of all machines in the warehouse. The company has collected 10,000 event samples within 3 months. The event samples include 100 failure cases that are evenly distributed across 50 different machine types.

How should the company prepare the data for the model to improve the model's accuracy?
  1. A Adjust the class weight to account for each machine type.
  2. B Oversample the failure cases by using the Synthetic Minority Oversampling Technique (SMOTE).
  3. C Undersample the non-failure events. Stratify the non-failure events by machine type.
  4. D Undersample the non-failure events by using the Synthetic Minority Oversampling Technique (SMOTE).
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS

Nội dung câu hỏi:
Câu hỏi tập trung vào việc chuẩn bị dữ liệu (data preparation) cho một mô hình dự đoán bảo trì dự đoán (predictive maintenance) trên AWS, cụ thể là dự đoán xác suất hỏng hóc của tất cả các máy móc trong kho hàng. Công ty có bộ dữ liệu gồm 10.000 mẫu sự kiện (event samples) thu thập trong 3 tháng, trong đó chỉ có 100 trường hợp hỏng hóc (failure cases) – chiếm tỷ lệ rất nhỏ (khoảng 1%). Các trường hợp hỏng được phân bố đều trên 50 loại máy khác nhau.

Vấn đề chính là dữ liệu không cân bằng (class imbalance): lớp thiểu số (failure) rất ít so với lớp đa số (non-failure), dẫn đến mô hình dễ bị bias về lớp đa số, làm giảm độ chính xác tổng thể (accuracy), đặc biệt là khả năng phát hiện failure. Mục tiêu là chọn phương pháp xử lý dữ liệu phù hợp để cải thiện độ chính xác mô hình, thường sử dụng các kỹ thuật trong Amazon SageMaker (như SageMaker Processing Jobs hoặc built-in algorithms hỗ trợ resampling). Kiến thức cập nhật đến 2026: AWS SageMaker tiếp tục hỗ trợ mạnh mẽ SMOTE qua thư viện imbalanced-learn (tích hợp trong SageMaker Python SDK v2.x và SageMaker JumpStart).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Oversample the failure cases by using the Synthetic Minority Oversampling Technique (SMOTE).

Lý do:

  • SMOTE là kỹ thuật oversampling lớp thiểu số (failure cases) bằng cách tạo ra các mẫu tổng hợp (synthetic samples) dựa trên k-nearest neighbors, giúp cân bằng dữ liệu mà không làm mất thông tin từ lớp đa số. Với chỉ 100 failure cases, SMOTE rất hiệu quả để tăng số lượng mẫu failure lên, cải thiện khả năng học của mô hình mà không gây overfitting nghiêm trọng.
  • Trong AWS SageMaker (phiên bản 2026), SMOTE được hỗ trợ trực tiếp qua SageMaker Processing Jobs với scikit-learn hoặc imbalanced-learn container, hoặc trong các algorithm như XGBoost với class_weight tự động, nhưng SMOTE là lựa chọn tối ưu cho imbalance nghiêm trọng như thế này. Phương pháp này giữ nguyên phân bố machine type (đã đều) và tập trung vào imbalance class.
  • Kết quả: Tăng F1-score và precision/recall cho lớp failure, dẫn đến accuracy tổng thể cao hơn 🛠️.

📋 Giải thích tất cả các phương án (đúng và sai)

Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên nội dung văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt với lý do đúng/sai dựa trên best practices AWS ML:

  • ❌ Adjust the class weight to account for each machine type.
    Phương án này sai vì class weight thường dùng để điều chỉnh penalty cho imbalance giữa failure/non-failure (không phải per machine type). Ở đây, failure đã phân bố đều trên 50 machine types, nên không cần weight theo machine type – điều này có thể gây phức tạp hóa mô hình và overfitting. Trong SageMaker XGBoost, class_weight hiệu quả hơn cho binary imbalance, nhưng không phải giải pháp chính cho data prep và không trực tiếp oversample data.

  • ✅ Oversample the failure cases by using the Synthetic Minority Oversampling Technique (SMOTE).
    Như đã giải thích ở trên, đây là đúng vì SMOTE lý tưởng cho lớp thiểu số hiếm (rare events) như failure (1%), tạo synthetic samples chất lượng cao, cải thiện accuracy mà giữ dataset size hợp lý (SageMaker hỗ trợ scale tốt với 10k samples).

  • ❌ Undersample the non-failure events. Stratify the non-failure events by machine type.
    Phương án này sai vì undersampling lớp non-failure (9900 samples) sẽ làm mất quá nhiều dữ liệu quý giá, dẫn đến underfitting và giảm generalization – đặc biệt với dataset chỉ 10k samples. Stratify theo machine type giúp giữ balance per type nhưng không giải quyết triệt để imbalance (vẫn chỉ 100 failure). AWS khuyến nghị tránh undersampling khi data limited; thay vào đó dùng SMOTE hoặc class_weight.

  • ❌ Undersample the non-failure events by using the Synthetic Minority Oversampling Technique (SMOTE).
    Phương án này sai hoàn toàn vì SMOTE là kỹ thuật oversampling cho minority class (failure), không dùng để undersample majority. Sử dụng SMOTE cho undersample là nhầm lẫn khái niệm, không tồn tại trong thực tế và sẽ gây lỗi khi implement trên SageMaker (imbalanced-learn chỉ hỗ trợ SMOTE cho oversample).

📘 Tài liệu tham khảo (AWS cập nhật đến 2026)

Phân tích này dựa trên kỳ thi AWS Certified Machine Learning – Specialty và DevOps Engineer Professional (MLOps pipelines) 🏆. Nếu cần code SageMaker implement SMOTE, hãy hỏi thêm!

Câu 220
A company stores its documents in Amazon S3 with no predefined product categories. A data scientist needs to build a machine learning model to categorize the documents for all the company's products.

Which solution will meet these requirements with the MOST operational efficiency?
  1. A Build a custom clustering model. Create a Dockerfile and build a Docker image. Register the Docker image in Amazon Elastic Container Registry (Amazon ECR). Use the custom image in Amazon SageMaker to generate a trained model.
  2. B Tokenize the data and transform the data into tabular data. Train an Amazon SageMaker k-means model to generate the product categories.
  3. C Train an Amazon SageMaker Neural Topic Model (NTM) model to generate the product categories.
  4. D Train an Amazon SageMaker Blazing Text model to generate the product categories.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc sử dụng Amazon SageMaker để xây dựng mô hình machine learning không giám sát (unsupervised learning) nhằm phân loại tài liệu lưu trữ trong Amazon S3 thành các danh mục sản phẩm (product categories). 📄

  • Bối cảnh: Tài liệu không có phân loại sẵn (no predefined product categories), nghĩa là không có dữ liệu nhãn (unlabeled data). Data scientist cần một giải pháp hiệu quả vận hành nhất (MOST operational efficiency), tức là ít công sức tùy chỉnh, triển khai nhanh, chi phí thấp và tận dụng các tính năng built-in của AWS.
  • Yêu cầu chính: Xử lý dữ liệu văn bản (text documents) để tự động nhóm (clustering/grouping) thành các chủ đề (topics) đại diện cho sản phẩm.
  • Mục tiêu: Ưu tiên giải pháp SageMaker built-in algorithms thay vì custom code phức tạp, vì chúng được tối ưu hóa cho quy mô lớn, dễ tích hợp với S3 và SageMaker Processing/Training jobs. 🛠️

Kiến thức cập nhật đến năm 2026: Amazon SageMaker tiếp tục hỗ trợ các built-in algorithms như Neural Topic Model (NTM) cho topic modeling (theo SageMaker documentation phiên bản mới nhất, không có thay đổi lớn trong unsupervised text clustering). ✅

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Train an Amazon SageMaker Neural Topic Model (NTM) model to generate the product categories.

Lý do:

  • 🏆 NTM là thuật toán built-in chuyên biệt cho topic modeling trên dữ liệu văn bản không nhãn, sử dụng neural network để học các chủ đề (topics) từ tài liệu S3 một cách tự động. Nó encode tài liệu thành vector topics, rất phù hợp để phân loại sản phẩm mà không cần dữ liệu huấn luyện có nhãn.
  • Operational efficiency cao nhất: Chỉ cần chuẩn bị dữ liệu JSON lines từ S3, chạy SageMaker Training Job với NTM estimator – không code custom, không Docker, tích hợp sẵn với S3 input/output. Thời gian triển khai nhanh, scale dễ dàng. 📈
  • Ưu điểm: Variational autoencoder-based, xử lý hàng triệu tài liệu hiệu quả, output topics có thể map trực tiếp thành categories. Đây là best practice cho unsupervised document categorization theo AWS Well-Architected Framework (ML Lens).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên tính phù hợp, độ phức tạp và operational efficiency. ❌ cho sai, ✅ cho đúng.

  • ❌ [SAI] Build a custom clustering model. Create a Dockerfile and build a Docker image. Register the Docker image in Amazon Elastic Container Registry (Amazon ECR). Use the custom image in Amazon SageMaker to generate a trained model.
    Giải thích: Phương án này yêu cầu xây dựng mô hình clustering tùy chỉnh (custom code), container hóa bằng Docker và đẩy lên ECR – quá phức tạp và kém efficient. Không tận dụng built-in algorithms, tốn thời gian dev/test/deploy (có thể hàng tuần), dễ lỗi và chi phí cao hơn. SageMaker hỗ trợ custom algorithms nhưng không phải lựa chọn "MOST operational efficiency" cho task đơn giản như topic modeling. 🕒

  • ❌ [SAI] Tokenize the data and transform the data into tabular data. Train an Amazon SageMaker k-means model to generate the product categories.
    Giải thích: K-means là thuật toán clustering cho dữ liệu số/tabular, không phù hợp trực tiếp với text documents. Bước tokenize + transform thành tabular (ví dụ TF-IDF) thêm overhead preprocessing phức tạp, không hiệu quả cho văn bản thô từ S3. NTM tốt hơn vì xử lý text native mà không cần transform. ❌

  • ✅ [ĐÚNG] Train an Amazon SageMaker Neural Topic Model (NTM) model to generate the product categories.
    Giải thích: Như đã nêu ở phần đáp án đúng. Đây là built-in algorithm tối ưu cho unsupervised topic discovery trên text, input trực tiếp từ S3 (JSON/RecordIO), output topic vectors dễ map thành categories. Hiệu quả cao: chỉ 1 Training Job, no custom code. 🏅

  • ❌ [SAI] Train an Amazon SageMaker Blazing Text model to generate the product categories.
    Giải thích: BlazingText chuyên cho supervised text classification (cần labels) hoặc word embeddings (Word2Vec), không phải topic modeling/clustering không nhãn. Nếu dùng unsupervised mode (Word2Vec), nó chỉ tạo word vectors chứ không group documents thành topics/categories. Không phù hợp, kém efficient hơn NTM. 🚫

📘 Tài liệu tham khảo

  • AWS SageMaker Documentation: Neural Topic Model (NTM) Algorithm – Built-in cho topic modeling (cập nhật 2024-2026).
  • AWS ML Best Practices: SageMaker Algorithms Reference – So sánh NTM vs. K-Means/BlazingText.
  • AWS Well-Architected Framework - ML Lens: Khuyến nghị built-in algorithms cho efficiency.
  • Exam Prep: AWS Certified Machine Learning - Specialty Sample Questions (tương tự DOP-C02 với ML integration).

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm ví dụ code SageMaker estimator cho NTM, hãy hỏi nhé.