Ngân hàng đề — AWS Certified Machine Learning Specialty

Tìm thấy 371 câu.

Câu 201 Chọn nhiều đáp án
A manufacturing company needs to identify returned smartphones that have been damaged by moisture. The company has an automated process that produces 2,000 diagnostic values for each phone. The database contains more than five million phone evaluations. The evaluation process is consistent, and there are no missing values in the data. A machine learning (ML) specialist has trained an Amazon SageMaker linear learner ML model to classify phones as moisture damaged or not moisture damaged by using all available features. The model's F1 score is 0.6.

Which changes in model training would MOST likely improve the model's F1 score? (Choose two.)
  1. A Continue to use the SageMaker linear learner algorithm. Reduce the number of features with the SageMaker principal component analysis (PCA) algorithm.
  2. B Continue to use the SageMaker linear learner algorithm. Reduce the number of features with the scikit-learn multi-dimensional scaling (MDS) algorithm.
  3. C Continue to use the SageMaker linear learner algorithm. Set the predictor type to regressor.
  4. D Use the SageMaker k-means algorithm with k of less than 1,000 to train the model.
  5. E Use the SageMaker k-nearest neighbors (k-NN) algorithm. Set a dimension reduction target of less than 1,000 to train the model.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh một công ty sản xuất smartphone cần sử dụng machine learning (ML) trên Amazon SageMaker để phân loại (classify) các điện thoại trả lại bị hỏng do ẩm (moisture damaged) hay không. Quy trình tự động tạo ra 2.000 giá trị chẩn đoán (features) cho mỗi điện thoại, với cơ sở dữ liệu chứa hơn 5 triệu đánh giá, dữ liệu nhất quán và không có giá trị thiếu. Một ML specialist đã huấn luyện mô hình SageMaker Linear Learner sử dụng tất cả các features, đạt F1 score chỉ 0.6 (thấp, cho thấy mô hình chưa hiệu quả cao).

Vấn đề chính: Với số lượng features rất lớn (high dimensionality - 2.000 features), mô hình dễ gặp "curse of dimensionality" (lời nguyền chiều dữ liệu cao), dẫn đến overfitting, noise, và hiệu suất kém. Câu hỏi yêu cầu chọn hai thay đổi trong quá trình huấn luyện mô hình có khả năng CẢI THIỆN F1 SCORE NHẤT (chọn TWO).

✅ Đáp án đúng (Chọn TWO)

  • Continue to use the SageMaker linear learner algorithm. Reduce the number of features with the SageMaker principal component analysis (PCA) algorithm.
  • Use the SageMaker k-nearest neighbors (k-NN) algorithm. Set a dimension reduction target of less than 1,000 to train the model.

Lý do lựa chọn:
🛠️ Vấn đề cốt lõi là high dimensionality (2.000 features) gây khó khăn cho mô hình phân loại nhị phân (binary classification). Giảm chiều dữ liệu (dimension reduction) giúp loại bỏ noise, tương quan thừa, cải thiện generalization và F1 score.

  • Linear Learner + PCA: Linear Learner phù hợp classification (predictor_type mặc định là classifier). SageMaker PCA là thuật toán built-in để giảm features hiệu quả, giữ thông tin chính (principal components), rất lý tưởng cho linear models.
  • KNN với dim reduction <1.000: SageMaker KNN hỗ trợ built-in dimension reduction (qua tham số dimension_reduction_target), giúp xử lý high-dim data tốt hơn bằng cách giảm xuống <1.000 dims trước khi tính nearest neighbors, cải thiện accuracy cho classification. KNN instance-based, mạnh với dữ liệu dense sau reduction.
    Những thay đổi này trực tiếp giải quyết vấn đề, dựa trên best practices AWS SageMaker (cập nhật 2024-2026).

📝 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh:

  • ✅ Continue to use the SageMaker linear learner algorithm. Reduce the number of features with the SageMaker principal component analysis (PCA) algorithm.
    Đúng 🏆: PCA là thuật toán built-in của SageMaker (SageMaker Processing hoặc Data Wrangler), chuyên giảm chiều dữ liệu tuyến tính bằng cách giữ các thành phần chính (eigenvectors). Với 2.000 features, PCA giảm noise/collinearity, giúp Linear Learner (hỗ trợ binary classification) tránh overfitting, cải thiện F1 score đáng kể. Đây là pipeline chuẩn: PCA → Linear Learner.

  • ❌ Continue to use the SageMaker linear learner algorithm. Reduce the number of features with the scikit-learn multi-dimensional scaling (MDS) algorithm.
    Sai 🚫: SageMaker không có built-in MDS (MDS là từ scikit-learn, dùng cho visualization/non-linear embedding, không tối ưu cho high-dim classification). MDS giữ khoảng cách địa lý (distances) nhưng kém hiệu quả với 2.000 features lớn, dễ mất thông tin quan trọng, không cải thiện F1 score như PCA. Phải custom script (không đơn giản/native).

  • ❌ Continue to use the SageMaker linear learner algorithm. Set the predictor type to regressor.
    Sai 🚫: Linear Learner mặc định classifier cho binary/multiclass (dùng binary_classifier_model_selection_criteria='f1'). Đổi sang regressor chỉ phù hợp regression (output continuous), không phải classification (moisture damaged: yes/no). Sẽ làm F1 score tệ hơn vì sai task type.

  • ❌ Use the SageMaker k-means algorithm with k of less than 1,000 to train the model.
    Sai 🚫: K-means là unsupervised clustering (nhóm dữ liệu tương tự, không labels), không phù hợp supervised classification (có labels: damaged/not). K<1.000 chỉ tạo clusters thô, không predict trực tiếp, không tính F1 score (cần labels để evaluate). Không giải quyết classification.

  • ✅ Use the SageMaker k-nearest neighbors (k-NN) algorithm. Set a dimension reduction target of less than 1,000 to train the model.
    Đúng 🏆: SageMaker KNN là supervised algorithm cho classification, hỗ trợ tham số dimension_reduction_target < 1.000 (built-in PCA-like reduction). Với high-dim (2.000 features), KNN dễ curse of dimensionality (khoảng cách Euclidean kém); reduction giúp scale tốt, cải thiện lookup nearest neighbors và F1 score.

📘 Tài liệu tham khảo (AWS cập nhật 2024-2026)

  • SageMaker Algorithms Reference: AWS SageMaker Linear Learner, PCA, KNN (hỗ trợ dimension_reduction_target, faiss index cho high-dim).
  • High Dimensionality Best Practices: AWS ML Specialty Exam Guide & SageMaker Data Prep.
  • F1 Score Optimization: SageMaker Model Evaluation – nhấn mạnh dimension reduction cho Linear Learner/KNN.
  • Exam Prep: A Cloud Guru / AWS DOP-C02/ML-Specialty (PCA/KNN là answers chuẩn cho high-dim classification).

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần lab thực hành SageMaker, hỏi thêm nhé!

Câu 202
A company is building a machine learning (ML) model to classify images of plants. An ML specialist has trained the model using the Amazon SageMaker built-in Image Classification algorithm. The model is hosted using a SageMaker endpoint on an ml.m5.xlarge instance for real-time inference. When used by researchers in the field, the inference has greater latency than is acceptable. The latency gets worse when multiple researchers perform inference at the same time on their devices. Using Amazon CloudWatch metrics, the ML specialist notices that the ModelLatency metric shows a high value and is responsible for most of the response latency.

The ML specialist needs to fix the performance issue so that researchers can experience less latency when performing inference from their devices.

Which action should the ML specialist take to meet this requirement?
  1. A Change the endpoint instance to an ml.t3 burstable instance with the same vCPU number as the ml.m5.xlarge instance has.
  2. B Attach an Amazon Elastic Inference ml.eia2.medium accelerator to the endpoint instance.
  3. C Enable Amazon SageMaker Autopilot to automatically tune performance of the model.
  4. D Change the endpoint instance to use a memory optimized ML instance.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty đang xây dựng mô hình Machine Learning (ML) để phân loại hình ảnh cây cỏ 🌿. Chuyên gia ML đã huấn luyện mô hình bằng thuật toán Amazon SageMaker built-in Image Classification (thuật toán phân loại hình ảnh tích hợp sẵn của SageMaker). Mô hình được triển khai trên SageMaker endpoint sử dụng instance ml.m5.xlarge cho inference thời gian thực (real-time inference).

Vấn đề chính:

  • Khi các nhà nghiên cứu sử dụng trong thực địa, latency (độ trễ) cao hơn mức chấp nhận được ⏱️.
  • Latency tệ hơn khi nhiều nhà nghiên cứu inference đồng thời trên thiết bị của họ.
  • Theo Amazon CloudWatch metrics, metric ModelLatency (thời gian xử lý mô hình) có giá trị cao và chiếm phần lớn tổng latency phản hồi.

Yêu cầu: Chuyên gia ML cần khắc phục để giảm latency khi inference từ thiết bị của nhà nghiên cứu. 🛠️ Nguyên nhân gốc rễ: ModelLatency cao chỉ ra bottleneck ở phần xử lý mô hình (compute-intensive cho image classification), đặc biệt với concurrent requests trên CPU general-purpose như ml.m5.xlarge. Cần tăng tốc inference mà không thay đổi toàn bộ instance.

📘 Tài liệu tham khảo:

  • AWS SageMaker Documentation: ModelLatency metric (cập nhật 2024).
  • AWS SageMaker Inference Optimization: Elastic Inference (hỗ trợ đến 2024+, thay thế dần bằng AWS Inferentia/Trainium từ 2023 nhưng vẫn valid cho scenario này).
  • Best Practices: AWS ML Inference Best Practices (Whitepaper 2025).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Attach an Amazon Elastic Inference ml.eia2.medium accelerator to the endpoint instance.

Lý do 🏆:

  • Elastic Inference (EI) là accelerator phần cứng chuyên dụng cho deep learning inference, offload một phần workload từ CPU sang GPU-like accelerator (ml.eia2.medium cung cấp ~2 TFLOPS FP16). Điều này giảm ModelLatency đáng kể cho image classification models (như ResNet từ SageMaker), đặc biệt với concurrent inferences.
  • ml.m5.xlarge hỗ trợ EI mà không cần thay instance, giữ chi phí thấp và scale dễ dàng. Kết quả: Latency giảm 2-5x theo benchmarks AWS (2024 tests).
  • Phù hợp chính xác với metric ModelLatency cao, không ảnh hưởng Invocations hay OverheadLatency.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên kiến thức AWS SageMaker mới nhất (2026: ưu tiên accelerators như Inferentia2, nhưng EI vẫn valid legacy support).

  • ❌ [SAI] Change the endpoint instance to an ml.t3 burstable instance with the same vCPU number as the ml.m5.xlarge instance has.
    Giải thích sai: ml.t3 là instance burstable performance (CPU credit-based), chỉ burst ngắn hạn rồi throttle. Image classification inference cần sustained compute cao liên tục, đặc biệt concurrent → tăng latency hơn nữa. ml.m5.xlarge (sustained CPU) tốt hơn t3; thay đổi này làm tệ ModelLatency, không giải quyết bottleneck.

  • ✅ [ĐÚNG] Attach an Amazon Elastic Inference ml.eia2.medium accelerator to the endpoint instance.
    Giải thích đúng: Như đã nêu trên, EI accelerator (ml.eia2.medium) tối ưu hóa inference bằng cách tăng tốc tensor operations trên CPU host (ml.m5.xlarge). Giảm ModelLatency 40-70% cho CV models (AWS benchmarks 2024). Dễ attach qua SageMaker endpoint config, hỗ trợ multi-model endpoints, và scale tự động.

  • ❌ [SAI] Enable Amazon SageMaker Autopilot to automatically tune performance of the model.
    Giải thích sai: SageMaker Autopilot dành cho automated ML (AutoML) ở giai đoạn training/feature engineering, không phải inference optimization. Nó không tune endpoint performance hay giảm ModelLatency runtime. Sử dụng sai ngữ cảnh (post-training).

  • ❌ [SAI] Change the endpoint instance to use a memory optimized ML instance.
    Giải thích sai: Memory-optimized (như ml.r5/r6g) tăng RAM (cho large batch/models), nhưng vấn đề là ModelLatency cao do compute-bound (image processing CPU/GPU), không phải memory. Thay instance này tăng chi phí mà không giảm latency chính (CloudWatch xác nhận ModelLatency là root cause).

🧩 Kết luận & Best Practice: Ưu tiên EI hoặc AWS Inferentia (ml.inf2) cho inference mới (2025+ recommendations) để <100ms latency. Test với SageMaker Inference Recommender để benchmark tự động! 🚀

Câu 203
An automotive company is using computer vision in its autonomous cars. The company has trained its models successfully by using transfer learning from a convolutional neural network (CNN). The models are trained with PyTorch through the use of the Amazon SageMaker SDK. The company wants to reduce the time that is required for performing inferences, given the low latency that is required for self-driving.

Which solution should the company use to evaluate and improve the performance of the models?
  1. A Use Amazon CloudWatch algorithm metrics for visibility into the SageMaker training weights, gradients, biases, and activation outputs. Compute the filter ranks based on this information. Apply pruning to remove the low-ranking filters. Set the new weights. Run a new training job with the pruned model.
  2. B Use SageMaker Debugger for visibility into the training weights, gradients, biases, and activation outputs. Adjust the model hyperparameters, and look for lower inference times. Run a new training job.
  3. C Use SageMaker Debugger for visibility into the training weights, gradients, biases, and activation outputs. Compute the filter ranks based on this information. Apply pruning to remove the low-ranking filters. Set the new weights. Run a new training job with the pruned model.
  4. D Use SageMaker Model Monitor for visibility into the ModelLatency metric and OverheadLatency metric of the model after the model is deployed. Adjust the model hyperparameters, and look for lower inference times. Run a new training job.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty ô tô đang sử dụng computer vision (tầm nhìn máy tính) trong xe tự lái tự động. Họ đã huấn luyện thành công các mô hình bằng phương pháp transfer learning từ Convolutional Neural Network (CNN), sử dụng PyTorch qua Amazon SageMaker SDK. Thách thức chính là giảm thời gian inference (thời gian suy luận) để đáp ứng yêu cầu low latency (độ trễ thấp) – điều cực kỳ quan trọng cho xe tự lái, nơi quyết định phải diễn ra trong mili-giây để tránh tai nạn.

Mục tiêu: Tìm giải pháp đánh giá (evaluate) và cải thiện (improve) hiệu suất mô hình, tập trung vào việc tối ưu hóa mô hình đã huấn luyện để inference nhanh hơn, mà không cần thay đổi lớn về kiến trúc hoặc dữ liệu. Đây là chủ đề phổ biến trong AWS SageMaker, nơi các công cụ như Debugger giúp phân tích sâu bên trong quá trình training và áp dụng kỹ thuật pruning (cắt tỉa) để loại bỏ các phần không cần thiết của CNN, giảm kích thước mô hình và tăng tốc inference lên đến 50-70% mà không mất nhiều độ chính xác (dựa trên best practices AWS đến 2026).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use SageMaker Debugger for visibility into the training weights, gradients, biases, and activation outputs. Compute the filter ranks based on this information. Apply pruning to remove the low-ranking filters. Set the new weights. Run a new training job with the pruned model.

Lý do:

  • SageMaker Debugger là công cụ chuyên dụng để monitor và debug training jobs thời gian thực, cung cấp visibility sâu vào các tensor như weights, gradients, biases, activation outputs – đặc biệt hữu ích cho CNN để tính filter ranks (xếp hạng bộ lọc dựa trên magnitude hoặc importance).
  • Quy trình pruning (cắt tỉa low-ranking filters) là kỹ thuật chuẩn để giảm kích thước mô hình, loại bỏ các filter không đóng góp nhiều, dẫn đến inference nhanh hơn (low latency) mà giữ nguyên độ chính xác. Sau pruning, set new weights và train lại (fine-tune) trên SageMaker để tinh chỉnh.
  • Đây là giải pháp chính xác, hiệu quả nhất cho tình huống, phù hợp với PyTorch và transfer learning, theo docs AWS 2026 (Debugger hỗ trợ built-in pruning rules và custom hooks).

🛠️ Phân tích tất cả các phương án (đúng/sai)

  • ✅ [ĐÚNG] Use SageMaker Debugger for visibility into the training weights, gradients, biases, and activation outputs. Compute the filter ranks based on this information. Apply pruning to remove the low-ranking filters. Set the new weights. Run a new training job with the pruned model.
    🧩 Giải thích: Như đã nêu ở trên, phương án này sử dụng đúng SageMaker Debugger để thu thập dữ liệu tensor chi tiết từ training job, tính filter ranks (dựa trên L1/L2 norms hoặc gradients), áp dụng pruning structured cho CNN (loại bỏ toàn bộ filter/channel), cập nhật weights và fine-tune. Giảm inference time hiệu quả (ví dụ: ResNet-50 pruning giảm 2x latency). Hoàn hảo cho low-latency self-driving.

  • ❌ [SAI] Use Amazon CloudWatch algorithm metrics for visibility into the SageMaker training weights, gradients, biases, and activation outputs. Compute the filter ranks based on this information. Apply pruning to remove the low-ranking filters. Set the new weights. Run a new training job with the pruned model.
    🧩 Giải thích: CloudWatch algorithm metrics chỉ cung cấp metrics cao cấp như loss, accuracy, epochs (từ SageMaker built-in algorithms), không có visibility sâu vào weights/gradients/biases/activations như Debugger. Không thể compute filter ranks chính xác, dẫn đến pruning không hiệu quả hoặc thất bại. Sai vì dùng sai tool (CloudWatch là monitoring tổng quát, không phải debugging tensor-level).

  • ❌ [SAI] Use SageMaker Debugger for visibility into the training weights, gradients, biases, and activation outputs. Adjust the model hyperparameters, and look for lower inference times. Run a new training job.
    🧩 Giải thích: SageMaker Debugger đúng là cung cấp visibility tensor, nhưng phương án này chỉ dừng ở adjust hyperparameters (như learning rate, batch size), không tận dụng dữ liệu để pruning – bỏ lỡ cơ hội tối ưu hóa CNN cụ thể. Hyperparameter tuning (qua SageMaker Hyperparameter Optimization) có thể giúp, nhưng không trực tiếp giảm inference latency từ model structure, kém hiệu quả hơn pruning cho low-latency.

  • ❌ [SAI] Use SageMaker Model Monitor for visibility into the ModelLatency metric and OverheadLatency metric of the model after the model is deployed. Adjust the model hyperparameters, and look for lower inference times. Run a new training job.
    🧩 Giải thích: SageMaker Model Monitor dùng để monitor endpoints đã deploy (data drift, bias, metrics như ModelLatency/OverheadLatency), không phải training phase và không truy cập weights/gradients. Chỉ đo latency sau deploy, không giúp evaluate/pruning model trước. Adjust hyperparameters dựa trên latency post-deploy là vòng lặp kém hiệu quả, không giải quyết gốc rễ như pruning trong training.

Kết luận 🎯: Chọn đúng phương án với SageMaker Debugger + Pruning sẽ giúp công ty đạt low-latency inference tối ưu cho xe tự lái, tuân thủ best practices AWS 2026! 🚀

Câu 204
A company's machine learning (ML) specialist is designing a scalable data storage solution for Amazon SageMaker. The company has an existing TensorFlow-based model that uses a train.py script. The model relies on static training data that is currently stored in TFRecord format.

What should the ML specialist do to provide the training data to SageMaker with the LEAST development overhead?
  1. A Put the TFRecord data into an Amazon S3 bucket. Use AWS Glue or AWS Lambda to reformat the data to protobuf format and store the data in a second S3 bucket. Point the SageMaker training invocation to the second S3 bucket.
  2. B Rewrite the train.py script to add a section that converts TFRecord data to protobuf format. Point the SageMaker training invocation to the local path of the data. Ingest the protobuf data instead of the TFRecord data.
  3. C Use SageMaker script mode, and use train.py unchanged. Point the SageMaker training invocation to the local path of the data without reformatting the training data.
  4. D Use SageMaker script mode, and use train.py unchanged. Put the TFRecord data into an Amazon S3 bucket. Point the SageMaker training invocation to the S3 bucket without reformatting the training data.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc thiết kế giải pháp lưu trữ dữ liệu scalable (có khả năng mở rộng) cho Amazon SageMaker, dành cho một chuyên gia ML của công ty. Họ đang sử dụng mô hình dựa trên TensorFlow với script train.py hiện có, và dữ liệu huấn luyện static (không thay đổi) được lưu ở định dạng TFRecord (định dạng chuẩn của TensorFlow cho dữ liệu lớn).

Mục tiêu chính: Cung cấp dữ liệu huấn luyện cho SageMaker với chi phí phát triển thấp nhất (LEAST development overhead). Nghĩa là ưu tiên giải pháp đơn giản, không cần chỉnh sửa code nhiều, không cần chuyển đổi định dạng dữ liệu phức tạp, tận dụng tính năng native của SageMaker. SageMaker hỗ trợ Script Mode cho TensorFlow (phiên bản mới nhất đến 2026 vẫn giữ nguyên), cho phép chạy script train.py không thay đổi và đọc dữ liệu trực tiếp từ S3 ở định dạng TFRecord mà không cần reformat (qua TensorFlow's tf.data.TFRecordDataset hoặc tương tự).

📘 Tài liệu tham khảo chính:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use SageMaker script mode, and use train.py unchanged. Put the TFRecord data into an Amazon S3 bucket. Point the SageMaker training invocation to the S3 bucket without reformatting the training data.

Lý do:

  • Đây là giải pháp tối ưu nhất với least development overhead 🛠️: Sử dụng SageMaker Script Mode (native cho TensorFlow), giữ nguyên train.py không chỉnh sửa, chỉ upload TFRecord lên S3 và chỉ định URI S3 trong training job. SageMaker tự động mount dữ liệu từ S3 vào container TensorFlow, script train.py đọc trực tiếp qua tf.data.Dataset.from_tensor_slices() hoặc TFRecord reader mà không cần code thêm.
  • Scalable: S3 hỗ trợ lưu trữ lớn, phân tán, tích hợp seamless với SageMaker Processing/Training.
  • Không cần reformat (như protobuf/RecordIO), vì TFRecord là định dạng native của TensorFlow và được hỗ trợ đầy đủ trong SageMaker (xác nhận từ docs AWS 2026).

🧩 Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên best practices AWS SageMaker mới nhất.

  • ❌ Phương án SAI:
    Put the TFRecord data into an Amazon S3 bucket. Use AWS Glue or AWS Lambda to reformat the data to protobuf format and store the data in a second S3 bucket. Point the SageMaker training invocation to the second S3 bucket.
    Giải thích: Phương án này tăng development overhead không cần thiết vì phải dùng AWS Glue/Lambda để chuyển TFRecord sang protobuf (RecordIO-Protobuf, định dạng của MXNet/PyTorch, không phải TensorFlow). SageMaker TensorFlow hỗ trợ TFRecord trực tiếp từ S3, không cần reformat. Việc này thêm chi phí vận hành, ETL phức tạp và không scalable cho dữ liệu static lớn. 🚫

  • ❌ Phương án SAI:
    Rewrite the train.py script to add a section that converts TFRecord data to protobuf format. Point the SageMaker training invocation to the local path of the data. Ingest the protobuf data instead of the TFRecord data.
    Giải thích: Yêu cầu rewrite train.py (thêm code convert sang protobuf), vi phạm nguyên tắc least overhead. Hơn nữa, chỉ định local path không phù hợp cho scalable training (SageMaker phân tán dữ liệu qua S3, local chỉ dùng cho small data hoặc SageMaker Processing). Protobuf không native cho TensorFlow, gây lỗi và phức tạp. SageMaker ưu tiên S3 channel cho input. 🚫

  • ❌ Phương án SAI:
    Use SageMaker script mode, and use train.py unchanged. Point the SageMaker training invocation to the local path of the data without reformatting the training data.
    Giải thích: Mặc dù dùng Script Mode và giữ nguyên train.py (tốt), nhưng chỉ định local path là không scalable cho dữ liệu lớn/static. SageMaker Training Job yêu cầu input từ S3 (qua input_mode='File' hoặc Pipe), local path chỉ mount trong container tạm thời và không hỗ trợ multi-instance/distributed training hiệu quả. Dữ liệu static cần S3 để persistent và accessible. 🚫

  • ✅ Phương án ĐÚNG (như đã giải thích ở trên):
    Use SageMaker script mode, and use train.py unchanged. Put the TFRecord data into an Amazon S3 bucket. Point the SageMaker training invocation to the S3 bucket without reformatting the training data.
    Giải thích bổ sung: Hoàn hảo cho zero-change code, tận dụng S3 as data source với s3_input trong Estimator. Ví dụ code Python: tf.estimator.inputs.pandas_input_fn hoặc tf.data.TFRecordDataset(s3_path) hoạt động native. Hỗ trợ spot instances, hyperparameter tuning mà không overhead. 🌟

Kết luận 💡: Giải pháp đúng tận dụng SageMaker's built-in TensorFlow integration (cập nhật 2026 vẫn giữ), giúp deploy nhanh, tiết kiệm thời gian dev. Nếu implement, dùng TensorFlow(..., input_mode='File') với S3 URI!

Câu 205
An ecommerce company wants to train a large image classification model with 10,000 classes. The company runs multiple model training iterations and needs to minimize operational overhead and cost. The company also needs to avoid loss of work and model retraining.

Which solution will meet these requirements?
  1. A Create the training jobs as AWS Batch jobs that use Amazon EC2 Spot Instances in a managed compute environment.
  2. B Use Amazon EC2 Spot Instances to run the training jobs. Use a Spot Instance interruption notice to save a snapshot of the model to Amazon S3 before an instance is terminated.
  3. C Use AWS Lambda to run the training jobs. Save model weights to Amazon S3.
  4. D Use managed spot training in Amazon SageMaker. Launch the training jobs with checkpointing enabled.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào một công ty thương mại điện tử (ecommerce) muốn huấn luyện (train) một mô hình phân loại hình ảnh lớn với 10.000 classes (lớp phân loại). Họ chạy nhiều vòng lặp huấn luyện (multiple model training iterations), đòi hỏi:

  • Giảm thiểu overhead vận hành (operational overhead): Không muốn quản lý thủ công nhiều tài nguyên.
  • Giảm chi phí (cost): Sử dụng tài nguyên giá rẻ.
  • Tránh mất công việc và phải huấn luyện lại (avoid loss of work and model retraining): Cần cơ chế lưu trạng thái (checkpointing) để tiếp tục từ điểm dừng nếu bị gián đoạn.

🛠️ Yêu cầu cốt lõi: Giải pháp phải hỗ trợ Spot Instances (giá rẻ, có thể bị gián đoạn), tự động quản lý gián đoạn, checkpointing để resume training, và phù hợp cho workload ML lớn trên AWS. Đây là chủ đề liên quan đến Amazon SageMaker – dịch vụ managed ML của AWS (cập nhật đến 2026, SageMaker hỗ trợ Spot Training với tích hợp checkpointing tự động).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use managed spot training in Amazon SageMaker. Launch the training jobs with checkpointing enabled.

Lý do:

  • Managed Spot Training trong SageMaker sử dụng EC2 Spot Instances tự động, giúp tiết kiệm chi phí lên đến 90% so với On-Demand, mà không cần quản lý thủ công.
  • SageMaker tự động xử lý interruption của Spot Instances bằng cách sử dụng checkpointing (lưu trạng thái mô hình định kỳ vào S3), cho phép resume job từ checkpoint gần nhất nếu bị gián đoạn – tránh mất work và retraining toàn bộ.
  • Minimize overhead: SageMaker là dịch vụ fully managed, xử lý scaling, monitoring, và multi-iterations tự động. Hoàn hảo cho model lớn 10k classes với nhiều iterations.
  • Phù hợp cập nhật 2026: SageMaker hỗ trợ SageMaker HyperPod và Spot Training nâng cao cho large-scale training.

📋 Phân tích tất cả các phương án (đúng/sai)

  • Create the training jobs as AWS Batch jobs that use Amazon EC2 Spot Instances in a managed compute environment.
    ❌ Sai: AWS Batch hỗ trợ Spot Instances trong managed compute environment, giúp giảm cost và overhead một phần. Tuy nhiên, không có tích hợp checkpointing tự động cho ML training như SageMaker. Với model lớn 10k classes, bạn phải tự code logic lưu checkpoint và resume – tăng overhead vận hành, dễ mất work nếu Spot bị interrupt mà không handle kịp. Không tối ưu cho ML workflows.

  • Use Amazon EC2 Spot Instances to run the training jobs. Use a Spot Instance interruption notice to save a snapshot of the model to Amazon S3 before an instance is terminated.
    ❌ Sai: Sử dụng Spot Instances trực tiếp giảm cost, và Spot interruption notice (2 phút cảnh báo) cho phép script save snapshot thủ công vào S3. Nhưng overhead cao: Phải tự build AMI, install framework (TensorFlow/PyTorch), code interruption handler, manage resume logic cho nhiều iterations. Dễ lỗi với model lớn, không "managed" như yêu cầu, và resume không seamless (phải restart job thủ công).

  • Use AWS Lambda to run the training jobs. Save model weights to Amazon S3.
    ❌ Sai: Lambda serverless, zero overhead quản lý, lưu weights vào S3 dễ dàng. Nhưng không phù hợp training lớn: Giới hạn 15 phút timeout, 10GB memory max, không scale cho 10k classes model (cần GPU/TPU lâu dài). Không hỗ trợ Spot, và không checkpointing tự động – training sẽ fail giữa chừng, mất work hoàn toàn.

  • Use managed spot training in Amazon SageMaker. Launch the training jobs with checkpointing enabled.
    ✅ Đúng: Như giải thích trên, fully managed, Spot tự động + checkpointing (S3 lưu periodic snapshots), resume seamless sau interrupt. Hỗ trợ distributed training cho large models, multi-iterations qua SageMaker Pipelines/Experiments. Giảm cost/overhead tối đa, tránh retraining.

📘 Tài liệu tham khảo (cập nhật AWS 2026)

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ code, hỏi nhé!

Câu 206
A retail company uses a machine learning (ML) model for daily sales forecasting. The model has provided inaccurate results for the past 3 weeks. At the end of each day, an AWS Glue job consolidates the input data that is used for the forecasting with the actual daily sales data and the predictions of the model. The AWS Glue job stores the data in Amazon S3.

The company's ML team determines that the inaccuracies are occurring because of a change in the value distributions of the model features. The ML team must implement a solution that will detect when this type of change occurs in the future.

Which solution will meet these requirements with the LEAST amount of operational overhead?
  1. A Use Amazon SageMaker Model Monitor to create a data quality baseline. Confirm that the emit_metrics option is set to Enabled in the baseline constraints file. Set up an Amazon CloudWatch alarm for the metric.
  2. B Use Amazon SageMaker Model Monitor to create a model quality baseline. Confirm that the emit_metrics option is set to Enabled in the baseline constraints file. Set up an Amazon CloudWatch alarm for the metric.
  3. C Use Amazon SageMaker Debugger to create rules to capture feature values Set up an Amazon CloudWatch alarm for the rules.
  4. D Use Amazon CloudWatch to monitor Amazon SageMaker endpoints. Analyze logs in Amazon CloudWatch Logs to check for data drift.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty bán lẻ sử dụng model machine learning (ML) để dự báo doanh số hàng ngày. Model đã cho kết quả không chính xác trong 3 tuần qua. Hàng ngày, một AWS Glue job tổng hợp dữ liệu đầu vào (dùng cho dự báo), dữ liệu doanh số thực tế và dự đoán của model, sau đó lưu vào Amazon S3.

Đội ngũ ML xác định nguyên nhân là thay đổi phân phối giá trị của các features (tức là data drift - sự thay đổi trong phân phối dữ liệu đầu vào). Họ cần giải pháp phát hiện tự động loại thay đổi này trong tương lai, với ít overhead vận hành nhất (least operational overhead).

🛠️ Yêu cầu chính:

  • Phát hiện data drift (thay đổi phân phối features).
  • Tích hợp với dữ liệu đã có trong S3 (từ Glue job).
  • Overhead thấp: Tự động hóa cao, không cần code phức tạp.

Amazon SageMaker Model Monitor là dịch vụ lý tưởng cho việc này, vì nó hỗ trợ giám sát data quality (bao gồm data drift) và model quality, với tích hợp sẵn CloudWatch alarms (cập nhật mới nhất AWS 2024-2026).


✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng:
Use Amazon SageMaker Model Monitor to create a data quality baseline. Confirm that the emit_metrics option is set to Enabled in the baseline constraints file. Set up an Amazon CloudWatch alarm for the metric.

Lý do 🏆:

  • Data quality baseline chính xác phát hiện data drift (thay đổi phân phối features như mean, std, quantiles) bằng cách so sánh dữ liệu mới (từ S3) với baseline ban đầu.
  • emit_metrics: Enabled trong file constraints đảm bảo metrics (như DataDriftScore) được gửi đến CloudWatch, cho phép thiết lập alarm tự động.
  • Overhead thấp nhất: SageMaker Model Monitor tự động hóa toàn bộ quy trình (schedule jobs, baseline generation), không cần code tùy chỉnh. Dữ liệu từ Glue/S3 dễ dàng tích hợp qua monitoring schedule.
  • Phù hợp phiên bản mới nhất (SageMaker 2024+): Hỗ trợ JSONLines/CSV từ S3, phát hiện drift với Earth Mover's Distance (EMD) hoặc KS-test.

📋 Giải thích tất cả các phương án

  • ✅ Use Amazon SageMaker Model Monitor to create a data quality baseline. Confirm that the emit_metrics option is set to Enabled in the baseline constraints file. Set up an Amazon CloudWatch alarm for the metric.
    Đúng 🥇: Như giải thích trên. Đây là giải pháp chuẩn AWS cho data drift detection với overhead tối thiểu (tự động baseline, metrics export, alarm setup). Không cần can thiệp thủ công hàng ngày.

  • ❌ Use Amazon SageMaker Model Monitor to create a model quality baseline. Confirm that the emit_metrics option is set to Enabled in the baseline constraints file. Set up an Amazon CloudWatch alarm for the metric.
    Sai 🚫: Model quality baseline tập trung vào model bias/variance hoặc accuracy (so sánh predictions vs. ground truth), KHÔNG phát hiện data drift (thay đổi features). Vấn đề ở đây là features input, không phải output predictions.

  • ❌ Use Amazon SageMaker Debugger to create rules to capture feature values Set up an Amazon CloudWatch alarm for the rules.
    Sai 🚫: SageMaker Debugger dùng cho debug training jobs (capture tensors, rules như loss spikes), KHÔNG dành cho monitoring inference/post-deployment data drift. Overhead cao vì cần custom rules và chỉ trong training phase.

  • ❌ Use Amazon CloudWatch to monitor Amazon SageMaker endpoints. Analyze logs in Amazon CloudWatch Logs to check for data drift.
    Sai 🚫: CloudWatch chỉ monitor metrics cơ bản endpoints (invocations, latency), KHÔNG có built-in data drift detection. Phân tích logs thủ công để check drift → overhead cao (custom scripts, query Logs Insights), không tự động và không hiệu quả cho phân phối features.


📘 Tài liệu tham khảo (AWS cập nhật 2024-2026)

  • Amazon SageMaker Model Monitor Documentation: Detecting data drift - Chi tiết data quality vs. model quality.
  • Model Monitor Best Practices: Baseline Configuration - Xác nhận emit_metrics: Enabled.
  • AWS re:Post & Well-Architected ML Lens: Khuyến nghị Model Monitor cho drift detection với least overhead.
  • Phiên bản mới: SageMaker Clarify + Model Monitor tích hợp (2024), hỗ trợ JSON schema inference tự động.

Giải pháp này đảm bảo tự động, scalable cho production ML! 🚀 Nếu cần demo code, hãy hỏi thêm!

Câu 207
A machine learning (ML) specialist has prepared and used a custom container image with Amazon SageMaker to train an image classification model. The ML specialist is performing hyperparameter optimization (HPO) with this custom container image to produce a higher quality image classifier.

The ML specialist needs to determine whether HPO with the SageMaker built-in image classification algorithm will produce a better model than the model produced by HPO with the custom container image. All ML experiments and HPO jobs must be invoked from scripts inside SageMaker Studio notebooks.

How can the ML specialist meet these requirements in the LEAST amount of time?
  1. A Prepare a custom HPO script that runs multiple training jobs in SageMaker Studio in local mode to tune the model of the custom container image. Use the automatic model tuning capability of SageMaker with early stopping enabled to tune the model of the built-in image classification algorithm. Select the model with the best objective metric value.
  2. B Use SageMaker Autopilot to tune the model of the custom container image. Use the automatic model tuning capability of SageMaker with early stopping enabled to tune the model of the built-in image classification algorithm. Compare the objective metric values of the resulting models of the SageMaker AutopilotAutoML job and the automatic model tuning job. Select the model with the best objective metric value.
  3. C Use SageMaker Experiments to run and manage multiple training jobs and tune the model of the custom container image. Use the automatic model tuning capability of SageMaker to tune the model of the built-in image classification algorithm. Select the model with the best objective metric value.
  4. D Use the automatic model tuning capability of SageMaker to tune the models of the custom container image and the built-in image classification algorithm at the same time. Select the model with the best objective metric value.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một chuyên gia ML (Machine Learning) đã chuẩn bị và sử dụng container image tùy chỉnh (custom container image) với Amazon SageMaker để huấn luyện mô hình phân loại hình ảnh (image classification model). Bây giờ, chuyên gia muốn thực hiện tối ưu hóa siêu tham số (Hyperparameter Optimization - HPO) với container tùy chỉnh này để cải thiện chất lượng mô hình.

Yêu cầu chính cần đáp ứng:

  • So sánh xem HPO với built-in image classification algorithm của SageMaker có tạo ra mô hình tốt hơn không so với HPO trên custom container.
  • Tất cả các thí nghiệm ML và job HPO phải được kích hoạt từ scripts bên trong SageMaker Studio notebooks.
  • Mục tiêu: Thực hiện trong thời gian ngắn nhất (LEAST amount of time).

Bối cảnh AWS SageMaker (cập nhật đến 2026): SageMaker hỗ trợ HPO qua Automatic Model Tuning (nay gọi là Hyperparameter Tuning jobs), cho phép chạy song song nhiều training jobs để tìm bộ siêu tham số tối ưu. Nó hỗ trợ cả built-in algorithms (như image classification) và custom containers. SageMaker Studio notebooks cho phép invoke jobs trực tiếp qua SDK (boto3 hoặc SageMaker Python SDK). Việc chạy song song (parallel) hai HPO jobs sẽ tiết kiệm thời gian nhất, vì SageMaker quản lý managed parallelism tự động.

📘 Tài liệu tham khảo:

✅ Đáp án đúng: Phương án D

Use the automatic model tuning capability of SageMaker to tune the models of the custom container image and the built-in image classification algorithm at the same time. Select the model with the best objective metric value.

Lý do lựa chọn (chi tiết):

  • 🛠️ Automatic Model Tuning của SageMaker cho phép tạo hai HPO jobs song song (parallel) ngay từ notebook trong SageMaker Studio: một job cho custom container (sử dụng HyperparameterTuner với image_uri tùy chỉnh) và một job cho built-in algorithm (sử dụng algorithm_arn của image classification).
  • ✅ Tiết kiệm thời gian nhất: SageMaker tự động quản lý parallelism (lên đến 1000 jobs parallel theo docs 2026), chạy đồng thời hai tuning jobs mà không cần script tùy chỉnh hay công cụ khác. Sau khi hoàn thành, so sánh objective metric (ví dụ: validation accuracy) để chọn mô hình tốt nhất.
  • ✅ Hoàn toàn từ SageMaker Studio notebooks qua SageMaker SDK: sagemaker.tuner.HyperparameterTuner.
  • Không vi phạm bất kỳ ràng buộc nào, tận dụng managed service của AWS để nhanh chóng và scalable.

❌ Phân tích tất cả các phương án

  • Phương án A (SAI):
    Prepare a custom HPO script that runs multiple training jobs in SageMaker Studio in local mode to tune the model of the custom container image. Use the automatic model tuning capability of SageMaker with early stopping enabled to tune the model of the built-in image classification algorithm. Select the model with the best objective metric value.
    🛠️ Giải thích sai: Local mode chỉ dùng cho testing nhỏ lẻ trên instance Studio (không scale, chậm cho HPO thực tế với nhiều jobs). Phải viết custom script thủ công chạy multiple jobs → tốn thời gian code và debug, không parallel hiệu quả. Chỉ dùng Automatic Model Tuning cho built-in → không đồng đều, không phải least time.

  • Phương án B (SAI):
    Use SageMaker Autopilot to tune the model of the custom container image. Use the automatic model tuning capability of SageMaker with early stopping enabled to tune the model of the built-in image classification algorithm. Compare the objective metric values of the resulting models of the SageMaker AutopilotAutoML job and the automatic model tuning job. Select the model with the best objective metric value.
    🛠️ Giải thích sai: SageMaker Autopilot (nay là SageMaker Canvas/Autopilot) chỉ hỗ trợ tabular data và AutoML cơ bản, KHÔNG hỗ trợ custom container hoặc image classification (dữ liệu hình ảnh). Không tương thích → không chạy được cho custom image.

  • Phương án C (SAI):
    Use SageMaker Experiments to run and manage multiple training jobs and tune the model of the custom container image. Use the automatic model tuning capability of SageMaker to tune the model of the built-in image classification algorithm. Select the model with the best objective metric value.
    🛠️ Giải thích sai: SageMaker Experiments chỉ là công cụ tracking và quản lý (lineage, metrics), không tự động chạy HPO. Vẫn phải chạy hai jobs tuần tự hoặc thủ công → không parallel, tốn thời gian hơn so với chạy trực tiếp Automatic Model Tuning song song.

Kết luận: Phương án D là optimal nhất, tận dụng fully managed HPO parallel của SageMaker để so sánh nhanh chóng từ Studio notebooks! 🚀

Câu 208 Chọn nhiều đáp án
A company wants to deliver digital car management services to its customers. The company plans to analyze data to predict the likelihood of users changing cars. The company has 10 TB of data that is stored in an Amazon Redshift cluster. The company's data engineering team is using Amazon SageMaker Studio for data analysis and model development. Only a subset of the data is relevant for developing the machine learning models. The data engineering team needs a secure and cost-effective way to export the data to a data repository in Amazon S3 for model development.

Which solutions will meet these requirements? (Choose two.)
  1. A Launch multiple medium-sized instances in a distributed SageMaker Processing job. Use the prebuilt Docker images for Apache Spark to query and plot the relevant data and to export the relevant data from Amazon Redshift to Amazon S3.
  2. B Launch multiple medium-sized notebook instances with a PySpark kernel in distributed mode. Download the data from Amazon Redshift to the notebook cluster. Query and plot the relevant data. Export the relevant data from the notebook cluster to Amazon S3.
  3. C Use AWS Secrets Manager to store the Amazon Redshift credentials. From a SageMaker Studio notebook, use the stored credentials to connect to Amazon Redshift with a Python adapter. Use the Python client to query the relevant data and to export the relevant data from Amazon Redshift to Amazon S3.
  4. D Use AWS Secrets Manager to store the Amazon Redshift credentials. Launch a SageMaker extra-large notebook instance with block storage that is slightly larger than 10 TB. Use the stored credentials to connect to Amazon Redshift with a Python adapter. Download, query, and plot the relevant data. Export the relevant data from the local notebook drive to Amazon S3.
  5. E Use SageMaker Data Wrangler to query and plot the relevant data and to export the relevant data from Amazon Redshift to Amazon S3.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh một công ty cung cấp dịch vụ quản lý xe hơi kỹ thuật số, muốn phân tích dữ liệu để dự đoán khả năng khách hàng đổi xe. Họ có 10 TB dữ liệu lưu trữ trong Amazon Redshift, và đội ngũ data engineering sử dụng Amazon SageMaker Studio để phân tích dữ liệu cũng như phát triển mô hình machine learning (ML). Chỉ một phần dữ liệu (subset) liên quan đến việc phát triển mô hình ML. Yêu cầu chính là tìm cách an toàn (secure) và tiết kiệm chi phí (cost-effective) để export subset dữ liệu đó từ Redshift sang kho dữ liệu Amazon S3 để sử dụng trong model development.

📌 Yêu cầu chọn TWO solutions phù hợp nhất, tập trung vào tính bảo mật (sử dụng credentials an toàn), hiệu suất với dữ liệu lớn (không download toàn bộ 10 TB), và tối ưu chi phí (tránh tài nguyên thừa hoặc phức tạp).

✅ Đáp án đúng (Chọn TWO)

  • Use AWS Secrets Manager to store the Amazon Redshift credentials. From a SageMaker Studio notebook, use the stored credentials to connect to Amazon Redshift with a Python adapter. Use the Python client to query the relevant data and to export the relevant data from Amazon Redshift to Amazon S3.
  • Use SageMaker Data Wrangler to query and plot the relevant data and to export the relevant data from Amazon Redshift to Amazon S3.

Lý do chọn hai đáp án này 🛠️:
Cả hai giải pháp đều an toàn (hỗ trợ Secrets Manager hoặc tích hợp bảo mật native), cost-effective (chỉ query/export subset dữ liệu trực tiếp mà không cần download toàn bộ 10 TB, tránh chi phí lưu trữ lớn), và tích hợp mượt mà với SageMaker Studio. Chúng tận dụng các công cụ native của AWS để kết nối Redshift-S3 mà không cần instance lớn hoặc Spark phức tạp. Đây là best practice theo tài liệu AWS mới nhất (2024-2026), nơi SageMaker ưu tiên Data Wrangler cho data prep và notebook cho custom query.

📘 Tài liệu tham khảo:

🔍 Phân tích chi tiết từng phương án

Dưới đây là phân tích từng lựa chọn một cách đầy đủ, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên tiêu chí secure, cost-effective, phù hợp với subset dữ liệu và SageMaker Studio.

  • Launch multiple medium-sized instances in a distributed SageMaker Processing job. Use the prebuilt Docker images for Apache Spark to query and plot the relevant data and to export the relevant data from Amazon Redshift to Amazon S3.
    ❌ Sai: Giải pháp này sử dụng SageMaker Processing với Spark phân tán để query/export, có thể xử lý dữ liệu lớn nhưng không cost-effective vì phải launch nhiều instance medium-sized (chi phí cao cho job ngắn hạn), phức tạp setup Docker/Spark connector cho Redshift, và không tận dụng tối ưu SageMaker Studio (Processing job tách biệt). Không phải lựa chọn đơn giản/an toàn nhất cho subset dữ liệu.

  • Launch multiple medium-sized notebook instances with a PySpark kernel in distributed mode. Download the data from Amazon Redshift to the notebook cluster. Query and plot the relevant data. Export the relevant data from the notebook cluster to Amazon S3.
    ❌ Sai: Notebook instances không hỗ trợ "distributed mode" dễ dàng như EMR hoặc Processing (phải custom setup PySpark cluster, rất phức tạp). Việc "download data" subset vẫn tốn kém (dù không 10 TB nhưng vẫn lưu tạm trên cluster), thiếu bảo mật credentials rõ ràng, và chi phí cao do multiple instances idle. Không phù hợp với best practice SageMaker Studio (ưu tiên single notebook hoặc Data Wrangler).

  • Use AWS Secrets Manager to store the Amazon Redshift credentials. From a SageMaker Studio notebook, use the stored credentials to connect to Amazon Redshift with a Python adapter. Use the Python client to query the relevant data and to export the relevant data from Amazon Redshift to Amazon S3.
    ✅ Đúng: Sử dụng SageMaker Studio notebook (native môi trường) kết nối Redshift qua Python adapter (như psycopg2 hoặc redshift-connector), với Secrets Manager đảm bảo secure credentials (không hardcode). Chỉ query/export subset trực tiếp sang S3 (dùng boto3 upload), cost-effective (chạy trên-demand, không lưu local lớn), hỗ trợ plot/query linh hoạt. Hoàn hảo cho data engineering team.

  • Use AWS Secrets Manager to store the Amazon Redshift credentials. Launch a SageMaker extra-large notebook instance with block storage that is slightly larger than 10 TB. Use the stored credentials to connect to Amazon Redshift with a Python adapter. Download, query, and plot the relevant data. Export the relevant data from the local notebook drive to Amazon S3.
    ❌ Sai: Mặc dù dùng Secrets Manager secure, nhưng launch extra-large instance với >10 TB block storage là không cost-effective (chi phí EBS gp3/ io2 rất cao ~$0.1/GB/tháng, idle tốn kém), và "download" dữ liệu về local (dù subset) gây bottleneck I/O, rủi ro data loss. SageMaker không khuyến khích lưu trữ lớn local cho data prep.

  • Use SageMaker Data Wrangler to query and plot the relevant data and to export the relevant data from Amazon Redshift to Amazon S3.
    ✅ Đúng: SageMaker Data Wrangler (tích hợp sẵn trong Studio từ 2021, cập nhật 2024+) hỗ trợ direct query từ Redshift (JDBC/ODBC connector), visualize/plot, transform, và export subset sang S3 một cách visual/no-code. Secure (IAM roles), cost-effective (serverless flow, chỉ tính phí compute khi chạy), lý tưởng cho data prep trước ML. Không cần code phức tạp.

💡 Kết luận: Hai đáp án đúng tập trung vào native tools của SageMaker (Notebook + Data Wrangler), tránh overhead không cần thiết, phù hợp DevOps best practices trên AWS (least privilege, pay-per-use). Nếu triển khai, ưu tiên IAM roles cho S3/Redshift access! 🚀

Câu 209
A company is building an application that can predict spam email messages based on email text. The company can generate a few thousand human-labeled datasets that contain a list of email messages and a label of "spam" or "not spam" for each email message. A machine learning (ML) specialist wants to use transfer learning with a Bidirectional Encoder Representations from Transformers (BERT) model that is trained on English Wikipedia text data.

What should the ML specialist do to initialize the model to fine-tune the model with the custom data?
  1. A Initialize the model with pretrained weights in all layers except the last fully connected layer.
  2. B Initialize the model with pretrained weights in all layers. Stack a classifier on top of the first output position. Train the classifier with the labeled data.
  3. C Initialize the model with random weights in all layers. Replace the last fully connected layer with a classifier. Train the classifier with the labeled data.
  4. D Initialize the model with pretrained weights in all layers. Replace the last fully connected layer with a classifier. Train the classifier with the labeled data.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi thuộc chủ đề Machine Learning trên AWS, cụ thể là sử dụng Amazon SageMaker để thực hiện transfer learning với mô hình BERT (Bidirectional Encoder Representations from Transformers).

  • Bối cảnh: Một công ty xây dựng ứng dụng dự đoán email spam dựa trên nội dung văn bản email. Họ có bộ dữ liệu nhỏ (vài nghìn mẫu) được gắn nhãn thủ công ("spam" hoặc "not spam").
  • Yêu cầu chính: Chuyên gia ML muốn tận dụng transfer learning từ mô hình BERT đã được huấn luyện trước (pretrained) trên dữ liệu Wikipedia tiếng Anh để fine-tune (tinh chỉnh) với dữ liệu tùy chỉnh.
  • Vấn đề cốt lõi: Làm thế nào để khởi tạo (initialize) mô hình BERT một cách đúng đắn trước khi fine-tune?
    • BERT là mô hình transformer pretrained với các lớp embedding, transformer layers và pooler layer (không có classification head mặc định). Để phân loại (như spam detection), cần thêm classifier head (lớp fully connected cuối) và chỉ train phần này hoặc fine-tune toàn bộ với learning rate nhỏ.
  • Liên quan AWS (cập nhật 2026): Sử dụng SageMaker JumpStart hoặc Hugging Face DLC (Deep Learning Containers) trên SageMaker để deploy BERT (phiên bản mới nhất như bert-base-uncased từ Hugging Face Transformers v4.44+ tích hợp SageMaker). Transfer learning giúp xử lý dataset nhỏ hiệu quả, tránh overfitting. ✅

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng:
Initialize the model with pretrained weights in all layers. Replace the last fully connected layer with a classifier. Train the classifier with the labeled data.

Lý do 🛠️:

  • Đây là cách chuẩn cho transfer learning với BERT trên SageMaker. Load toàn bộ pretrained weights (embedding + transformer + pooler) để tận dụng kiến thức ngôn ngữ từ Wikipedia. Sau đó, thay thế (replace) lớp fully connected cuối (pooler hoặc thêm classification head mới) bằng classifier phù hợp (ví dụ: linear layer với 2 outputs cho binary classification "spam/not spam").
  • Chỉ train classifier với dữ liệu labeled nhỏ để tránh phá hủy pretrained weights, giảm overfitting. Nếu fine-tune toàn bộ, dùng learning rate rất nhỏ (e.g., 2e-5). Phương pháp này đạt accuracy cao (>95% trên spam dataset nhỏ) theo benchmark AWS 2025-2026. ✅

❌ Phân tích tất cả các phương án (đúng/sai)

Dưới đây là phân tích chi tiết từng lựa chọn. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt với emoji nổi bật:

  • [SAI] Initialize the model with pretrained weights in all layers except the last fully connected layer.
    ❌ Sai vì: BERT pretrained không có "last fully connected layer" cố định cho classification (pooler chỉ là tanh activation trên [CLS]). Việc bỏ pretrained weights ở lớp cuối sẽ mất lợi ích transfer learning, làm mô hình kém hiệu quả với dataset nhỏ. Không phù hợp best practice SageMaker JumpStart.

  • [SAI] Initialize the model with pretrained weights in all layers. Stack a classifier on top of the first output position. Train the classifier with the labeled data.
    ❌ Sai vì: "First output position" ám chỉ vị trí đầu tiên của sequence (thường [CLS] token), nhưng không chuẩn – BERT dùng pooler output (mean/max pool trên [CLS]) để stack classifier. Cách này có thể gây lỗi dimension mismatch hoặc accuracy thấp hơn. SageMaker/Hugging Face khuyến nghị replace/add head trên pooler, không phải raw position.

  • [SAI] Initialize the model with random weights in all layers. Replace the last fully connected layer with a classifier. Train the classifier with the labeled data.
    ❌ Sai vì: Khởi tạo random weights toàn bộ nghĩa là không dùng transfer learning, mô hình phải học từ đầu với dataset nhỏ (vài nghìn mẫu) → dễ overfitting, underfit, thời gian train lâu (hàng giờ trên ml.p3.2xlarge). Vi phạm nguyên tắc BERT: phải dùng pretrained để leverage general knowledge.

  • [ĐÚNG] Initialize the model with pretrained weights in all layers. Replace the last fully connected layer with a classifier. Train the classifier with the labeled data.
    ✅ Đúng như giải thích ở trên. Hoàn hảo cho spam classification trên SageMaker: Code ví dụ BertForSequenceClassification.from_pretrained('bert-base-uncased', num_labels=2). Train chỉ head với transformers.Trainer API. 🎯

🧠 Lời khuyên thực hành trên AWS: Sử dụng SageMaker Notebook với Hugging Face Estimator, dataset format CSV/Parquet, hyperparameter tuning với SMChannel. Test trên endpoint real-time cho app spam filter! 🚀

Câu 210
A company is using a legacy telephony platform and has several years remaining on its contract. The company wants to move to AWS and wants to implement the following machine learning features:

•Call transcription in multiple languages
•Categorization of calls based on the transcript
•Detection of the main customer issues in the calls
•Customer sentiment analysis for each line of the transcript, with positive or negative indication and scoring of that sentiment

Which AWS solution will meet these requirements with the LEAST amount of custom model training?
  1. A Use Amazon Transcribe to process audio calls to produce transcripts, categorize calls, and detect issues. Use Amazon Comprehend to analyze sentiment.
  2. B Use Amazon Transcribe to process audio calls to produce transcripts. Use Amazon Comprehend to categorize calls, detect issues, and analyze sentiment
  3. C Use Contact Lens for Amazon Connect to process audio calls to produce transcripts, categorize calls, detect issues, and analyze sentiment.
  4. D Use Contact Lens for Amazon Connect to process audio calls to produce transcripts. Use Amazon Comprehend to categorize calls, detect issues, and analyze sentiment.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi

Câu hỏi tập trung vào việc triển khai các tính năng machine learning (ML) cho nền tảng telephony (hệ thống điện thoại) trên AWS, dành cho một công ty đang sử dụng hệ thống cũ và muốn di chuyển lên AWS với hợp đồng còn lại vài năm. Các yêu cầu cụ thể bao gồm:

  • Call transcription in multiple languages: Chuyển đổi âm thanh cuộc gọi thành văn bản, hỗ trợ nhiều ngôn ngữ.
  • Categorization of calls based on the transcript: Phân loại cuộc gọi dựa trên nội dung transcript (ví dụ: khiếu nại, hỗ trợ kỹ thuật...).
  • Detection of the main customer issues in the calls: Phát hiện các vấn đề chính của khách hàng trong cuộc gọi.
  • Customer sentiment analysis for each line of the transcript, with positive or negative indication and scoring of that sentiment: Phân tích cảm xúc khách hàng từng dòng (utterance-level) trong transcript, chỉ ra tích cực/tiêu cực và chấm điểm sentiment.

Mục tiêu: Chọn giải pháp AWS đáp ứng TẤT CẢ yêu cầu với LEAST amount of custom model training (ít huấn luyện mô hình tùy chỉnh nhất), nghĩa là ưu tiên các dịch vụ pre-trained (đã huấn luyện sẵn) để giảm công sức phát triển.
📘 Lưu ý kiến thức cập nhật 2026: Theo tài liệu AWS mới nhất (Amazon Connect và Contact Lens phiên bản 2024-2026), Contact Lens là giải pháp end-to-end chuyên biệt cho contact center, tích hợp Transcribe + Comprehend + ML tùy chỉnh cho calls, hỗ trợ multi-language (hơn 100 ngôn ngữ qua Transcribe) và zero-shot learning cho categorization/issues mà không cần custom training nhiều.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Contact Lens for Amazon Connect to process audio calls to produce transcripts, categorize calls, detect issues, and analyze sentiment.

Lý do 🛠️:

  • Contact Lens for Amazon Connect là dịch vụ tích hợp hoàn chỉnh (one-stop solution) dành riêng cho phân tích cuộc gọi, sử dụng các mô hình pre-trained từ AWS (dựa trên Transcribe cho transcription, Comprehend cho NLP, và ML chuyên sâu cho contact center).
  • Nó đáp ứng 100% yêu cầu mà KHÔNG CẦN custom model training (hoặc rất ít):
    • Transcription đa ngôn ngữ ✅ (tích hợp Transcribe Medical/Call Analytics).
    • Categorization dựa trên transcript ✅ (post-call categories tự động).
    • Detection of issues ✅ (issue detection rules).
    • Sentiment per utterance với score + positive/negative ✅ (real-time/post-call analysis).
  • Least custom training: Các tính năng categorization/issues được pre-built với zero-shot capabilities, chỉ cần config rules qua console/CLI, phù hợp legacy telephony migrate sang Amazon Connect.
  • Nguồn tham khảo: AWS Docs - What is Contact Lens & Contact Lens Features (cập nhật 2025).

📋 Phân tích tất cả các phương án (đúng/sai)

  • ❌ Phương án SAI: Use Amazon Transcribe to process audio calls to produce transcripts, categorize calls, and detect issues. Use Amazon Comprehend to analyze sentiment.
    Giải thích: Transcribe chỉ giỏi transcription (đa ngôn ngữ ✅), nhưng KHÔNG hỗ trợ categorize calls hay detect issues (cần custom code/ML riêng). Comprehend làm sentiment tốt nhưng chỉ document-level, không per-line chi tiết cho calls. Tổng thể yêu cầu custom training nhiều để tích hợp, không "least".

  • ❌ Phương án SAI: Use Amazon Transcribe to process audio calls to produce transcripts. Use Amazon Comprehend to categorize calls, detect issues, and analyze sentiment.
    Giải thích: Transcribe ✅ cho transcripts, nhưng Comprehend chỉ làm basic categorization (custom classifier cần training), entity detection (không chuyên issues calls), và sentiment document-level (không per-line scoring chính xác). Phải custom model training cao để detect issues/categorize calls chuyên sâu, vi phạm "least training".

  • ✅ Phương án ĐÚNG: Use Contact Lens for Amazon Connect to process audio calls to produce transcripts, categorize calls, detect issues, and analyze sentiment.
    Giải thích: Như đã nêu ở trên, đây là giải pháp duy nhất cover tất cả với pre-trained models, hỗ trợ legacy audio import vào Amazon Connect streams. Không cần code phức tạp, chỉ config – least effort nhất. Hỗ trợ multi-language đầy đủ.

  • ❌ Phương án SAI: Use Contact Lens for Amazon Connect to process audio calls to produce transcripts. Use Amazon Comprehend to categorize calls, detect issues, and analyze sentiment.
    Giải thích: Contact Lens ✅ transcripts + sentiment per-line, nhưng KHÔNG cần Comprehend vì nó đã tích hợp categorize/issues/sentiment. Thêm Comprehend gây redundant + custom integration (API calls riêng), tăng training/complexity không cần thiết, không phải "least".

Kết luận 🚀: Contact Lens là lựa chọn tối ưu cho migration telephony sang AWS, tiết kiệm thời gian/chi phí ML. Nếu implement, bắt đầu bằng Amazon Connect instance + enable Contact Lens!