Ngân hàng đề — AWS Certified Machine Learning Specialty
Tìm thấy 371 câu.
A data scientist will use statistical modeling to discover abstract topics and to provide a list of the top words for each category to help the auditors assess the relevance of the topic.
Which algorithms are best suited to this scenario? (Choose two.)
- A Latent Dirichlet allocation (LDA)
- B Random forest classifier
- C Neural topic modeling (NTM)
- D Linear support vector machine
- E Linear regression
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một tình huống thực tế trong lĩnh vực dược phẩm trên nền tảng AWS (liên quan đến dịch vụ Amazon SageMaker cho machine learning trên dữ liệu text). Công ty lưu trữ tài liệu kiểm toán dạng text, cần phân tích nhanh để khám phá 10 chủ đề chính (topics) trong tài liệu, giúp phân công công việc kiểm toán. Đặc biệt, ưu tiên cao nhất cho tài liệu liên quan đến adverse events (sự kiện bất lợi). Data scientist sử dụng statistical modeling để phát hiện các chủ đề trừu tượng (abstract topics) và cung cấp danh sách top words cho từng chủ đề, giúp auditors đánh giá tính liên quan.
Mục tiêu chính: Chọn 2 thuật toán phù hợp nhất cho topic modeling (mô hình hóa chủ đề) trên dữ liệu text lớn, không giám sát (unsupervised), tập trung vào việc tự động trích xuất topics và từ khóa đại diện. Đây là ứng dụng điển hình của Amazon SageMaker BlazingText hoặc SageMaker Processing với các built-in algorithms (cập nhật đến AWS 2026: SageMaker hỗ trợ LDA và NTM qua SageMaker JumpStart và custom containers).
📘 Nguồn tham khảo:
- AWS SageMaker Documentation: "Built-in Algorithms" (LDA & Neural Topic Models) - https://docs.aws.amazon.com/sagemaker/latest/dg/topic-models.html
- AWS ML Specialty Exam Guide (2024-2026 updates): Topic modeling trong SageMaker.
✅ Đáp án đúng (Chọn 2)
Latent Dirichlet allocation (LDA) và Neural topic modeling (NTM) là hai thuật toán phù hợp nhất.
Lý do lựa chọn:
- Cả hai đều là thuật toán topic modeling unsupervised, chuyên dùng để tự động phát hiện các chủ đề ẩn trong corpus text lớn, cung cấp phân phối topics trên documents và top words per topic. Điều này khớp hoàn hảo với nhu cầu "discover abstract topics" và "top words for each category" để ưu tiên adverse events (ví dụ: topics chứa từ như "adverse", "event", "side effect" sẽ được ưu tiên).
- Trong AWS SageMaker (phiên bản mới nhất 2026), LDA là built-in algorithm cổ điển (probabilistic generative model), còn NTM là deep learning-based (sử dụng neural networks cho scalability cao hơn trên dữ liệu lớn). Chúng hỗ trợ xử lý batch text, phù hợp audits định kỳ. 🛠️
📋 Giải thích chi tiết từng phương án
-
Latent Dirichlet allocation (LDA) ✅ Đúng.
LDA là thuật toán probabilistic topic model kinh điển (unsupervised), giả định documents là hỗn hợp các topics, mỗi topic là phân phối từ trên từ vựng. Nó tự động khám phá K topics (ở đây K=10), output top words/topic và topic proportions/document. Hoàn hảo cho scenario, hỗ trợ trực tiếp trong SageMaker (training job vớisagemaker.lda.LDA). Lý tưởng cho text audits để prioritize adverse events topics. 🚀 -
Random forest classifier ❌ Sai.
Random Forest là ensemble supervised classifier (dùng cho classification tasks với labeled data), không phải topic modeling. Nó dự đoán class labels (ví dụ: classify documents vào categories có sẵn), chứ không tự khám phá abstract topics hay top words. Không phù hợp unsupervised discovery trên text audits. Phải cần labels trước (không có ở đây). 🌳 -
Neural topic modeling (NTM) ✅ Đúng.
NTM là neural network-based topic model (variational autoencoder), cải tiến LDA với deep learning cho hiệu suất cao hơn trên dữ liệu lớn/sparse. Nó học topics qua neural embeddings, cung cấp top words và topic distributions mượt mà. SageMaker hỗ trợ NTM qua custom algorithms hoặc JumpStart (2026: tích hợp tốt với SageMaker Canvas cho non-coders). Lý tưởng cho scale audits lớn, prioritize adverse events. 🧠 -
Linear support vector machine ❌ Sai.
Linear SVM là supervised classifier tối ưu hóa hyperplane để phân loại (text classification với TF-IDF features), cần labeled training data. Không hỗ trợ topic discovery unsupervised hay top words per abstract topic. Chỉ dùng nếu đã có labels cho adverse events (không khớp scenario). ⚡ -
Linear regression ❌ Sai.
Linear Regression là supervised regression model dự đoán continuous values (ví dụ: sentiment score), không liên quan topic modeling hay classification. Hoàn toàn không phù hợp với text analysis để khám phá topics/top words trên audits. 📈
Which solution will meet these requirements with the LEAST development effort?
- A Index company documents by using Amazon Kendra. Integrate the chatbot with Amazon Kendra by using the Amazon Kendra Query API operation to answer customer questions.
- B Train a Bidirectional Attention Flow (BiDAF) network based on past customer questions and company documents. Deploy the model as a real-time Amazon SageMaker endpoint. Integrate the model with the chatbot by using the SageMaker Runtime InvokeEndpoint API operation to answer customer questions.
- C Train an Amazon SageMaker Blazing Text model based on past customer questions and company documents. Deploy the model as a real-time SageMaker endpoint. Integrate the model with the chatbot by using the SageMaker Runtime InvokeEndpoint API operation to answer customer questions.
- D Index company documents by using Amazon OpenSearch Service. Integrate the chatbot with OpenSearch Service by using the OpenSearch Service k-nearest neighbors (k-NN) Query API operation to answer customer questions.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc triển khai một chatbot để trả lời các câu hỏi phổ biến từ khách hàng, dựa hoàn toàn trên tài liệu nội bộ của công ty. Yêu cầu chính là chọn giải pháp ít nỗ lực phát triển nhất (LEAST development effort) trên AWS.
✅ Mục tiêu cốt lõi: Chatbot cần "hiểu" và trích xuất thông tin từ tài liệu công ty (như PDF, Word, web pages) để trả lời tự nhiên bằng ngôn ngữ người dùng, mà không cần xây dựng mô hình AI phức tạp từ đầu. AWS cung cấp các dịch vụ managed để giảm thiểu code custom, training dữ liệu và quản lý hạ tầng.
🛠️ Bối cảnh AWS (cập nhật đến 2026): Với các tính năng mới nhất như Amazon Kendra's generative AI integration (qua Amazon Bedrock), dịch vụ này hỗ trợ index tài liệu tự động, semantic search và Q&A mà không cần training model. Điều này phù hợp với yêu cầu "least effort" vì chỉ cần upload tài liệu và gọi API.
📘 Tài liệu tham khảo:
- AWS Kendra Documentation: Amazon Kendra Developer Guide (cập nhật 2025-2026 với hỗ trợ RAG - Retrieval-Augmented Generation).
- AWS Well-Architected Framework - Machine Learning Lens: Nhấn mạnh managed services như Kendra để giảm operational overhead.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Index company documents by using Amazon Kendra. Integrate the chatbot with Amazon Kendra by using the Amazon Kendra Query API operation to answer customer questions.
Lý do chọn 🏆:
- Amazon Kendra là dịch vụ fully managed chuyên cho enterprise search và Q&A, tự động index tài liệu (hỗ trợ 20+ định dạng), sử dụng ML để hiểu ngữ nghĩa và trả lời chính xác mà không cần training model hay code phức tạp.
- Chỉ cần: Tạo index → Upload tài liệu → Gọi Query API (hỗ trợ natural language queries). Ít effort nhất vì AWS xử lý indexing, ranking, security (ACL, IAM).
- So với các option khác, không yêu cầu custom ML training hay vector search thủ công, phù hợp "least development effort". Với cập nhật 2026, Kendra tích hợp Bedrock cho generative answers, tăng độ chính xác.
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể:
-
Index company documents by using Amazon Kendra. Integrate the chatbot with Amazon Kendra by using the Amazon Kendra Query API operation to answer customer questions.
✅ Đúng vì: Kendra là giải pháp managed lý tưởng cho Q&A trên tài liệu, tự động xử lý semantic search và retrieval. Query API đơn giản (POST request với text query), trả về top answers với confidence score. Không cần dev ML pipeline, chỉ config index (sync S3, FSx, etc.). Least effort, scalable đến hàng triệu docs. -
Train a Bidirectional Attention Flow (BiDAF) network based on past customer questions and company documents. Deploy the model as a real-time Amazon SageMaker endpoint. Integrate the model with the chatbot by using the SageMaker Runtime InvokeEndpoint API operation to answer customer questions.
❌ Sai vì: BiDAF là mô hình QA cổ điển (2017), yêu cầu training từ đầu trên dữ liệu lớn (past questions + docs), tune hyperparameters, xử lý data preprocessing. Deploy SageMaker endpoint tốn effort cao (script, container, monitoring). Không phải least effort, dễ overkill và kém chính xác nếu data ít. -
Train an Amazon SageMaker Blazing Text model based on past customer questions and company documents. Deploy the model as a real-time SageMaker endpoint. Integrate the model with the chatbot by using the SageMaker Runtime InvokeEndpoint API operation to answer customer questions.
❌ Sai vì: BlazingText chuyên cho text classification/word2vec, không tối ưu cho Q&A/retrieval trên docs. Vẫn cần full training pipeline (labeling data, hyperparameter tuning), deploy endpoint và invoke API – effort cao gấp nhiều lần Kendra. Phù hợp NLP tasks khác, không phải document-based Q&A. -
Index company documents by using Amazon OpenSearch Service. Integrate the chatbot with OpenSearch Service by using the OpenSearch Service k-nearest neighbors (k-NN) Query API operation to answer customer questions.
❌ Sai vì: OpenSearch (fork của Elasticsearch) tốt cho search, nhưng k-NN yêu cầu vector embeddings tự tạo (dùng model như BERT), index vectors thủ công, config plugin k-NN. Effort cao: Xử lý embedding pipeline, similarity search, reranking. Không managed như Kendra cho Q&A semantic, dễ kém chính xác với natural language mà không thêm custom logic.
🏁 Kết luận & Best Practices
Giải pháp Kendra là optimal cho scenario này, giảm TCO và time-to-market. Nếu scale lớn, kết hợp với Amazon Lex cho chatbot UI và Bedrock cho generative responses (cập nhật 2026).
📘 Nguồn bổ sung:
- AWS re:Invent 2025 sessions on Kendra + Bedrock.
- Sample code Kendra Query API: AWS SDK Docs.
Nếu cần demo code hoặc architecture diagram, hãy cho tôi biết! 🚀
The company has a small internal team that is working on the project. The internal team has no ML expertise and no ML experience.
Which solution will meet these requirements with the LEAST amount of effort from the internal team?
- A Set up a private workforce that consists of the internal team. Use the private workforce and the SageMaker Ground Truth active learning feature to label the data. Use Amazon Rekognition Custom Labels for model training and hosting.
- B Set up a private workforce that consists of the internal team. Use the private workforce to label the data. Use Amazon Rekognition Custom Labels for model training and hosting.
- C Set up a private workforce that consists of the internal team. Use the private workforce and the SageMaker Ground Truth active learning feature to label the data. Use the SageMaker Object Detection algorithm to train a model. Use SageMaker batch transform for inference.
- D Set up a public workforce. Use the public workforce to label the data. Use the SageMaker Object Detection algorithm to train a model. Use SageMaker batch transform for inference.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh một công ty muốn sử dụng machine learning (ML) trên AWS để xác định các ngôi nhà đã lắp đặt tấm pin mặt trời (solar panels) từ hình ảnh vệ tinh (satellite images), nhằm mục đích tiếp thị nhắm mục tiêu. Họ có 8.000 hình ảnh làm dữ liệu huấn luyện và sẽ dùng Amazon SageMaker Ground Truth để gán nhãn (labeling) dữ liệu. Đội ngũ nội bộ nhỏ, không có kinh nghiệm ML, nên giải pháp cần giảm thiểu tối đa nỗ lực (LEAST amount of effort) từ đội ngũ này.
Yêu cầu chính: Giải pháp phải tận dụng SageMaker Ground Truth, dễ dàng cho người mới, tự động hóa cao, và phù hợp với dữ liệu object detection (phát hiện vật thể pin mặt trời trên ảnh).
📘 Nguồn tham khảo: AWS SageMaker Ground Truth Documentation (cập nhật 2024-2026), Amazon Rekognition Custom Labels.
✅ Đáp án đúng
Set up a private workforce that consists of the internal team. Use the private workforce and the SageMaker Ground Truth active learning feature to label the data. Use Amazon Rekognition Custom Labels for model training and hosting.
Lý do lựa chọn:
- Private workforce sử dụng đội ngũ nội bộ để gán nhãn, phù hợp với dữ liệu nhạy cảm (hình ảnh nhà dân).
- SageMaker Ground Truth active learning tự động chọn các hình ảnh khó nhất để gán nhãn, giảm đáng kể số lượng task (từ 8.000 xuống chỉ ~10-20% nhờ ML loop), tiết kiệm effort cho team không chuyên.
- Amazon Rekognition Custom Labels là giải pháp no-code/low-code lý tưởng cho người mới: tự động huấn luyện, deploy model object detection từ dữ liệu đã label, tích hợp sẵn Ground Truth, không cần kiến thức ML sâu (như tuning hyperparameters). Đây là lựa chọn least effort tổng thể.
🛠️ Ưu điểm cập nhật 2026: Rekognition CL hỗ trợ active learning native, inference real-time/batch, và auto-scaling hosting.
📋 Giải thích tất cả các phương án
-
✅ Set up a private workforce that consists of the internal team. Use the private workforce and the SageMaker Ground Truth active learning feature to label the data. Use Amazon Rekognition Custom Labels for model training and hosting.
🟢 Đúng vì: Kết hợp active learning giảm effort labeling (team chỉ label dữ liệu "khó"), Rekognition Custom Labels xử lý toàn bộ training/hosting tự động, không cần code ML. Least effort cho team nội bộ không expertise. -
❌ Set up a private workforce that consists of the internal team. Use the private workforce to label the data. Use Amazon Rekognition Custom Labels for model training and hosting.
🔴 Sai vì: Thiếu active learning, team phải label toàn bộ 8.000 ảnh thủ công → effort cao gấp nhiều lần (hàng nghìn giờ làm việc). Không tối ưu least effort, dù Rekognition CL vẫn tốt cho training. -
❌ Set up a private workforce that consists of the internal team. Use the private workforce and the SageMaker Ground Truth active learning feature to label the data. Use the SageMaker Object Detection algorithm to train a model. Use SageMaker batch transform for inference.
🔴 Sai vì: Dù có active learning giảm labeling, SageMaker Object Detection algorithm yêu cầu kiến thức ML chuyên sâu (chọn algorithm, tuning hyperparameters, data preprocessing, custom script) → team không expertise sẽ gặp khó khăn lớn. Batch transform chỉ cho inference offline, không linh hoạt như hosting real-time. -
❌ Set up a public workforce. Use the public workforce to label the data. Use the SageMaker Object Detection algorithm to train a model. Use SageMaker batch transform for inference.
🔴 Sai vì: Public workforce (external như Mechanical Turk) giảm labeling effort nhưng dữ liệu nhạy cảm (nhà dân) có rủi ro bảo mật cao; hơn nữa, SageMaker Object Detection vẫn đòi hỏi effort ML cao (training custom), không phù hợp team không kinh nghiệm. Không phải least effort tổng thể vì thiếu tự động hóa end-to-end.
📘 Lưu ý: Public workforce rẻ nhưng cần quản lý quality control, không lý tưởng cho dữ liệu private.
Which solution will meet these requirements with the LEAST development effort?
- A Use Amazon SageMaker Data Wrangler with a custom transformation to identify and redact the PII.
- B Create a custom AWS Lambda function to read the files, identify the PII. and redact the PII
- C Use AWS Glue DataBrew to identity and redact the PII
- D Use an AWS Glue development endpoint to implement the PII redaction from within a notebook
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào một công ty lưu trữ kho dữ liệu machine learning (ML) trên Amazon S3. Một data scientist đang chuẩn bị dữ liệu để huấn luyện mô hình ML và cần xóa bỏ (redact) thông tin cá nhân có thể nhận diện (PII - Personally Identifiable Information) từ bộ dữ liệu này. Yêu cầu chính là chọn giải pháp với ÍT NHIỆU CÔNG SỨC PHÁT TRIỂN NHẤT (LEAST development effort).
🛠️ Vấn đề cốt lõi: PII bao gồm các thông tin nhạy cảm như tên, địa chỉ email, số điện thoại, SSN, v.v. Giải pháp phải xử lý dữ liệu trên S3 một cách tự động, dễ dàng, không yêu cầu viết code tùy chỉnh phức tạp. AWS cung cấp các công cụ ETL (Extract, Transform, Load) và data preparation để hỗ trợ việc này, với ưu tiên cho dịch vụ có tính năng built-in sẵn cho PII detection và redaction (dựa trên cập nhật AWS Glue DataBrew phiên bản mới nhất đến 2026, hỗ trợ hơn 300 recipe transformations bao gồm PII handling).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use AWS Glue DataBrew to identity and redact the PII
Lý do: AWS Glue DataBrew là dịch vụ visual data preparation (chuẩn bị dữ liệu trực quan) được thiết kế dành riêng cho data scientists, cho phép phát hiện (detect) và xóa bỏ PII tự động mà không cần viết code tùy chỉnh. Bạn chỉ cần tạo một DataBrew project, kết nối với S3, chọn built-in recipe cho PII (hỗ trợ detect các loại như email, phone, name, address, credit card, v.v.), và apply transformation trực tiếp qua giao diện kéo-thả. Kết quả xuất ra S3 bucket mới. Điều này đáp ứng LEAST development effort vì toàn bộ quy trình là low-code/no-code, nhanh chóng (chỉ vài phút setup), và scale tự động với Spark engine. Theo tài liệu AWS mới nhất (2026), DataBrew tích hợp machine learning-based PII detection chính xác cao, giảm thiểu effort so với các giải pháp custom.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, với giữ nguyên văn bản gốc tiếng Anh và đánh giá đúng/sai dựa trên tiêu chí LEAST development effort:
-
❌ Use Amazon SageMaker Data Wrangler with a custom transformation to identify and redact the PII
Sai vì: SageMaker Data Wrangler là công cụ mạnh cho data prep trong ML workflow, nhưng yêu cầu custom transformation (viết script Python hoặc Spark để detect/redact PII). Điều này đòi hỏi development effort cao (code logic detect PII bằng regex/ML models), debug, và integrate với S3. Không phải giải pháp ít effort nhất, dù hỗ trợ import từ S3. -
❌ Create a custom AWS Lambda function to read the files, identify the PII. and redact the PII
Sai vì: Lambda là serverless compute, nhưng yêu cầu viết full custom code (Python/Node.js) để đọc S3 objects, scan PII (sử dụng thư viện như presidio hoặc regex), redact, và ghi lại S3. Phải handle scale (large datasets), error handling, permissions – effort cao, không visual/no-code, dễ vượt Lambda limits (15 phút runtime, 10GB memory). -
✅ Use AWS Glue DataBrew to identity and redact the PII
Đúng vì: Như đã giải thích ở trên, built-in PII detection/redaction qua giao diện visual, hỗ trợ S3 trực tiếp, zero custom code. Ít effort nhất cho data scientists (drag-and-drop recipes), chi phí pay-per-job, và output sạch sẵn sàng cho ML training. -
❌ Use an AWS Glue development endpoint to implement the PII redaction from within a notebook
Sai vì: Glue Development Endpoint (nay là Glue Interactive Sessions/Studio notebooks) cho phép viết Spark code trong Jupyter notebook, nhưng vẫn yêu cầu implement custom logic cho PII (code PySpark/SQL để scan/redact). Effort tương đương coding từ đầu, không low-code như DataBrew, và phức tạp hơn cho non-engineers.
📘 Tài liệu tham khảo
- AWS Glue DataBrew Documentation (cập nhật 2026): AWS Docs - Detect and Redact PII with DataBrew – Chi tiết built-in PII transformations.
- AWS re:Invent 2025/2026 Announcements: DataBrew enhancements với ML-powered PII (xem AWS Blog: "Simplifying PII Handling in Data Prep").
- Exam Topic DOP-C02: Phần Data Preparation & ML Ops trên S3 (AWS Certified DevOps Engineer Professional blueprint).
🛠️ Lời khuyên: Trong thực tế DevOps, ưu tiên DataBrew cho quick data cleaning trước khi push vào SageMaker training jobs để optimize pipeline!
Four times a year, the company samples the data from the previous 90 days to check the ML model for drift. After the 90-day period, the company must keep the files for compliance reasons.
The company needs to use S3 storage classes to minimize costs. The company wants to maintain the same storage durability of the data.
Which solution will meet these requirements?
- A Store the daily objects in the S3 Standard-InfrequentAccess (S3 Standard-IA) storage class. Configure an S3 Lifecycle rule to move the objects to S3 Glacier Flexible Retrieval after 90 days.
- B Store the daily objects in the S3 One Zone-Infrequent Access (S3 One Zone-IA) storage class. Configure an S3 Lifecycle rule to move the objects to S3 Glacier Flexible Retrieval after 90 days.
- C Store the daily objects in the S3 Standard-InfrequentAccess (S3 Standard-IA) storage class. Configure an S3 Lifecycle rule to move the objects to S3 Glacier Deep Archive after 90 days.
- D Store the daily objects in the S3 One Zone-Infrequent Access (S3 One Zone-IA) storage class. Configure an S3 Lifecycle rule to move the objects to S3 Glacier Deep Archive after 90 days.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh việc tối ưu hóa chi phí lưu trữ dữ liệu trên Amazon S3 cho một công ty triển khai mô hình Machine Learning (ML) trong môi trường production. 🔍
-
Yêu cầu chính:
- Dữ liệu: Các file hàng ngày (100 GB/file, kích thước tăng dần) chứa inputs và predictions của ML model, lưu dưới dạng object trong S3 bucket.
- Truy cập dữ liệu: 4 lần/năm, công ty sample dữ liệu từ 90 ngày trước để kiểm tra model drift (sự lệch lạc của mô hình theo thời gian).
- Giữ dữ liệu: Sau 90 ngày, phải giữ vĩnh viễn vì lý do compliance (tuân thủ quy định pháp lý).
- Mục tiêu: Sử dụng S3 storage classes để giảm thiểu chi phí tối đa, nhưng duy trì độ bền (durability) giống nhau (99.999999999% - 11 9's).
-
Thách thức: Dữ liệu trong 90 ngày cần infrequent access (truy cập không thường xuyên), sau đó là archival lâu dài với truy cập rất hiếm. Phải chọn storage class cân bằng giữa chi phí thấp và độ bền cao. 🛠️
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Store the daily objects in the S3 Standard-InfrequentAccess (S3 Standard-IA) storage class. Configure an S3 Lifecycle rule để move the objects to S3 Glacier Deep Archive after 90 days.
Lý do chi tiết:
- S3 Standard-IA: Phù hợp cho giai đoạn 90 ngày đầu vì hỗ trợ infrequent access (truy cập không thường xuyên, chỉ 4 lần/năm), chi phí lưu trữ thấp hơn S3 Standard (khoảng 1/3-1/4), độ bền 11 9's (multi-AZ). Minimum storage duration 30 ngày phù hợp vì file >100GB.
- S3 Glacier Deep Archive: Sau 90 ngày, đây là storage class rẻ nhất cho archival lâu dài (compliance), chi phí chỉ ~1/10 so với Standard-IA, độ bền 11 9's, retrieval time 12 giờ (phù hợp vì access rất hiếm).
- S3 Lifecycle rule: Tự động transition sau 90 ngày, zero downtime, giảm chi phí mà không mất dữ liệu.
- Tối ưu hoàn hảo: Giữ durability cao, chi phí thấp nhất theo best practices AWS 2024-2026 (S3 Intelligent-Tiering có thể thay thế nhưng câu hỏi yêu cầu explicit classes). 📉💰
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn (giữ nguyên văn bản gốc bằng tiếng Anh). Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm lý do bằng tiếng Việt rõ ràng dựa trên kiến thức AWS mới nhất (2026).
-
❌ [SAI] Store the daily objects in the S3 Standard-InfrequentAccess (S3 Standard-IA) storage class. Configure an S3 Lifecycle rule to move the objects to S3 Glacier Flexible Retrieval after 90 days.
- Lý do sai: S3 Standard-IA tốt cho 90 ngày đầu (infrequent access, durability 11 9's). Tuy nhiên, S3 Glacier Flexible Retrieval (retrieval 1-5 phút, chi phí cao hơn ~5-10 lần so với Deep Archive) không tối ưu chi phí cho archival compliance (access rất hiếm). Nên dùng Deep Archive rẻ hơn để "minimize costs". 🤑
-
❌ [SAI] Store the daily objects in the S3 One Zone-Infrequent Access (S3 One Zone-IA) storage class. Configure an S3 Lifecycle rule to move the objects to S3 Glacier Flexible Retrieval after 90 days.
- Lý do sai: S3 One Zone-IA chỉ lưu ở 1 Availability Zone (AZ), durability chỉ 9 9's (thấp hơn 11 9's), không duy trì durability giống nhau như yêu cầu (rủi ro mất dữ liệu nếu AZ fail). Glacier Flexible Retrieval cũng đắt đỏ không cần thiết. Không phù hợp compliance. ⚠️
-
✅ [ĐÚNG] Store the daily objects in the S3 Standard-InfrequentAccess (S3 Standard-IA) storage class. Configure an S3 Lifecycle rule to move the objects to S3 Glacier Deep Archive after 90 days.
- Lý do đúng: Như phân tích ở phần trên – Standard-IA cho access 90 ngày (durability 11 9's, chi phí thấp), Deep Archive cho archival (rẻ nhất, durability 11 9's, retrieval 12h phù hợp). Hoàn hảo minimize costs mà giữ durability. 🎯
-
❌ [SAI] Store the daily objects in the S3 One Zone-Infrequent Access (S3 One Zone-IA) storage class. Configure an S3 Lifecycle rule to move the objects to S3 Glacier Deep Archive after 90 days.
- Lý do sai: S3 One Zone-IA có durability chỉ 9 9's (single AZ), vi phạm yêu cầu "maintain the same storage durability" (phải 11 9's như Standard-IA). Deep Archive tốt nhưng giai đoạn đầu sai. Rủi ro cao cho dữ liệu ML quan trọng. 🚫
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- AWS S3 Storage Classes: https://docs.aws.amazon.com/AmazonS3/latest/userguide/storage-class-intro.html (so sánh chi phí/durability).
- S3 Lifecycle Policies: https://docs.aws.amazon.com/AmazonS3/latest/userguide/object-lifecycle-mgmt.html (transition rules).
- S3 Features by Storage Class (Table): https://aws.amazon.com/s3/storage-classes/ (Deep Archive vs Flexible: chi phí Deep Archive thấp hơn 75%).
- AWS Well-Architected Framework - Cost Optimization: Pillar khuyến nghị IA + Deep Archive cho infrequent/archival data.
- Exam Prep DOP-C02: Topic S3 Optimization (phiên bản 2024+).
Giải pháp này production-ready, dễ implement qua AWS Console/CLI/Terraform! 🚀 Nếu cần code sample Lifecycle policy, hỏi thêm nhé! 😊
Which solution will meet these requirements with the LEAST development effort?
- A Use Amazon SageMaker Feature Store to select the features. Create a data flow to perform feature-level metadata analysis. Create an Amazon DynamoDB table to store feature-level metadata. Use Amazon QuickSight to analyze the metadata.
- B Use Amazon SageMaker Feature Store to set feature groups for the current features that the ML models use. Assign the required metadata for each feature. Use SageMaker Studio to analyze the metadata.
- C Use Amazon SageMaker Features Store to apply custom algorithms to analyze the feature-level metadata that the company requires. Create an Amazon DynamoDB table to store feature-level metadata. Use Amazon QuickSight to analyze the metadata.
- D Use Amazon SageMaker Feature Store to set feature groups for the current features that the ML models use. Assign the required metadata for each feature. Use Amazon QuickSight to analyze the metadata.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc xây dựng hệ thống kiểm toán (auditing) cho các mô hình Machine Learning (ML) trên AWS, với các yêu cầu chính sau:
- Phân tích metadata của các features (đặc trưng dữ liệu) mà mô hình ML sử dụng.
- Tạo báo cáo phân tích metadata này.
- Thiết lập (set) data sensitivity (độ nhạy cảm dữ liệu) và authorship (tác giả sở hữu) cho từng feature.
- Giải pháp phải đáp ứng với ÍT NHẤT nỗ lực phát triển (least development effort), nghĩa là ưu tiên các tính năng native của AWS, tránh custom code phức tạp.
Chủ đề chính là Amazon SageMaker Feature Store – dịch vụ lưu trữ và quản lý features cho ML, hỗ trợ metadata tại mức feature-level (bao gồm custom parameters như sensitivity và authorship). Giải pháp lý tưởng phải tận dụng trực tiếp khả năng của Feature Store để assign metadata và visualize/report qua công cụ tích hợp sẵn.
📘 Kiến thức cập nhật (2026): SageMaker Feature Store (phiên bản mới nhất) hỗ trợ feature definitions với custom_metadata (dict tùy chỉnh cho sensitivity, ownership), tích hợp QuickSight cho visualization metadata mà không cần code thêm. (Nguồn: AWS SageMaker Feature Store Docs).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Amazon SageMaker Feature Store to set feature groups for the current features that the ML models use. Assign the required metadata for each feature. Use Amazon QuickSight to analyze the metadata.
Lý do:
- Giải pháp này sử dụng native features của SageMaker Feature Store: Tạo Feature Groups để quản lý features hiện tại, assign metadata trực tiếp qua
FeatureDefinition(hỗ trợ custom_metadata cho sensitivity và authorship) mà không cần code custom. - QuickSight tích hợp sẵn với Feature Store (qua Athena hoặc direct connector) để query/visualize metadata, tạo báo cáo tự động – đây là cách least effort nhất cho auditing/reporting.
- Không cần DynamoDB, data flow hay algorithms tùy chỉnh, giảm thiểu phát triển.
🛠️ Ưu điểm: Fully managed, scalable, phù hợp audits ML theo best practices AWS (MLOps).
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể bằng tiếng Việt:
-
❌ Phương án SAI: Use Amazon SageMaker Feature Store to select the features. Create a data flow to perform feature-level metadata analysis. Create an Amazon DynamoDB table to store feature-level metadata. Use Amazon QuickSight to analyze the metadata.
Lý do sai: Yêu cầu tạo data flow tùy chỉnh (custom ETL) và DynamoDB để lưu metadata – tốn nhiều effort phát triển (code Lambda/Glue, schema design), không tận dụng native metadata storage của Feature Store. QuickSight chỉ là phần cuối, không giải quyết gốc rễ. -
❌ Phương án SAI: Use Amazon SageMaker Feature Store to set feature groups for the current features that the ML models use. Assign the required metadata for each feature. Use SageMaker Studio to analyze the metadata.
Lý do sai: Phần set Feature Groups và assign metadata là đúng (native), nhưng SageMaker Studio chỉ là IDE cho data scientists (explore notebooks), không phải tool reporting/auditing chuyên dụng. Phân tích metadata ở đây cần custom notebooks/SQL – effort cao hơn QuickSight (thiếu dashboard tự động, sharing reports). -
❌ Phương án SAI: Use Amazon SageMaker Features Store to apply custom algorithms to analyze the feature-level metadata that the company requires. Create an Amazon DynamoDB table to store feature-level metadata. Use Amazon QuickSight to analyze the metadata.
Lý do sai: Sai chính tả ("Features Store" thay vì "Feature Store"), và yêu cầu custom algorithms (SageMaker Processing/Endpoints) + DynamoDB – hoàn toàn tốn effort cao (viết code phân tích, migrate data). Không dùng native metadata của Feature Store, vi phạm "least effort". -
✅ Phương án ĐÚNG: Use Amazon SageMaker Feature Store to set feature groups for the current features that the ML models use. Assign the required metadata for each feature. Use Amazon QuickSight to analyze the metadata.
Lý do đúng: Hoàn hảo khớp yêu cầu – Feature Groups native để quản lý features/metadata (sensitivity/authorship quaparameters), QuickSight cho báo cáo phân tích (direct integration via Athena/S3 export). Zero custom code, fully managed.
🛠️ Best practice: Theo AWS Well-Architected Framework for ML (Lens: Operational Excellence).
Tài liệu tham khảo thêm:
Which solution will meet these requirements?
- A Define security groups to allow all HTTP inbound and outbound traffic. Assign the security groups to the SageMaker notebook instance.
- B Configure the SageMaker notebook instance to have access to the VPC. Grant permission in the AWS Key Management Service (AWS KMS) key policy to the notebook’s VPC.
- C Assign an IAM role that provides S3 read access for the dataset to the SageMaker notebook. Grant permission in the KMS key policy to the IAM role.
- D Assign the same KMS key that encrypts the data in Amazon S3 to the SageMaker notebook instance.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào bảo mật dữ liệu trong AWS khi một chuyên gia Machine Learning (ML) tải dataset lên Amazon S3 bucket được bảo vệ bởi server-side encryption với AWS KMS keys (SSE-KMS). Yêu cầu chính là đảm bảo Amazon SageMaker notebook instance có thể đọc dataset từ S3 mà không gặp lỗi quyền truy cập.
🔍 Các yếu tố cốt lõi cần lưu ý:
- S3 bucket sử dụng SSE-KMS, nghĩa là dữ liệu được mã hóa bằng KMS customer-managed key (CMK) hoặc AWS-managed key. Để đọc dữ liệu, cần quyền S3:GetObject và KMS:Decrypt trên key policy.
- SageMaker notebook instance chạy trong môi trường AWS, sử dụng IAM role để truy cập tài nguyên bên ngoài (như S3).
- Giải pháp phải an toàn, tuân thủ nguyên tắc least privilege, tránh mở rộng quyền không cần thiết (ví dụ: không mở port HTTP rộng rãi).
- Kiến thức cập nhật đến 2026: Theo tài liệu AWS mới nhất (SageMaker Execution Role và KMS integration), IAM role là cách chuẩn để SageMaker truy cập S3 SSE-KMS (không thay đổi từ 2023-2026).
📘 Tài liệu tham khảo:
- Amazon SageMaker Developer Guide: IAM Roles for SageMaker
- AWS KMS Key Policies for S3 SSE-KMS
- S3 Server-Side Encryption
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Assign an IAM role that provides S3 read access for the dataset to the SageMaker notebook. Grant permission in the KMS key policy to the IAM role.
🛠️ Lý do chi tiết:
- SageMaker notebook instance sử dụng IAM execution role (gắn trực tiếp khi tạo notebook) để truy cập S3. Role này cần policy cho phép s3:GetObject trên bucket/object cụ thể.
- Với SSE-KMS, role còn cần quyền kms:Decrypt, kms:DescribeKey trong KMS key policy (external policy của KMS key, cho phép principal là IAM role).
- Đây là best practice của AWS: Tuân thủ least privilege, dễ quản lý, không cần thay đổi VPC hay security groups. SageMaker tự động sử dụng role để gọi S3 API khi đọc dataset (ví dụ:
pd.read_csv('s3://bucket/dataset.csv')).
❌ Phân tích tất cả các phương án (đúng/sai)
-
Define security groups to allow all HTTP inbound and outbound traffic. Assign the security groups to the SageMaker notebook instance.
❌ Sai: Security groups kiểm soát network traffic (Layer 4), không liên quan đến IAM/KMS permissions cho S3 API calls. Mở tất cả HTTP traffic là rủi ro bảo mật cao (vi phạm nguyên tắc least privilege), và S3 truy cập qua HTTPS endpoint (không cần inbound/outbound tùy chỉnh). SageMaker notebook mặc định có network access đến S3 public endpoints. -
Configure the SageMaker notebook instance to have access to the VPC. Grant permission in the AWS Key Management Service (AWS KMS) key policy to the notebook’s VPC.
❌ Sai: VPC chỉ cần cho private S3 endpoints (VPC Endpoint), nhưng câu hỏi không đề cập private bucket. KMS key policy không hỗ trợ principal là VPC (chỉ hỗ trợ IAM roles/users/accounts). Phải grant cho IAM role của notebook, không phải VPC endpoint (VPC endpoint dùng IAM policy riêng). -
Assign an IAM role that provides S3 read access for the dataset to the SageMaker notebook. Grant permission in the KMS key policy to the IAM role.
✅ Đúng: Như giải thích ở trên. Đây là giải pháp chuẩn, kết hợp IAM policy cho S3 + KMS key policy cho decrypt. SageMaker notebook sử dụng role này để đọc dữ liệu mã hóa một cách an toàn. -
Assign the same KMS key that encrypts the data in Amazon S3 to the SageMaker notebook instance.
❌ Sai: SageMaker notebook không hỗ trợ assign KMS key trực tiếp như một resource (khác với EBS encryption). KMS key chỉ dùng cho encryption/decryption qua API calls, không phải attach vào instance. Phải dùng IAM role + key policy để cấp quyền.
🧠 Kết luận: Giải pháp đúng nhấn mạnh IAM-centric security – cốt lõi của AWS Well-Architected Framework (Security Pillar). Nếu triển khai, kiểm tra CloudTrail logs để verify quyền truy cập! 🚀
How should the ML specialist design the transformation step to meet these requirements with the LEAST operational effort?
- A Use an Amazon Managed Streaming for Apache Kafka (Amazon MSK) cluster to ingest event data. Use Amazon Kinesis Data Analytics to transform the most recent 10 minutes of data before inference.
- B Use Amazon Kinesis Data Streams to ingest event data. Store the data in Amazon S3 by using Amazon Kinesis Data Firehose. Use AWS Lambda to transform the most recent 10 minutes of data before inference.
- C Use Amazon Kinesis Data Streams to ingest event data. Use Amazon Kinesis Data Analytics to transform the most recent 10 minutes of data before inference.
- D Use an Amazon Managed Streaming for Apache Kafka (Amazon MSK) cluster to ingest event data. Use AWS Lambda to transform the most recent 10 minutes of data before inference.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi xoay quanh việc thiết kế quy trình ingestion và transformation dữ liệu thời gian thực cho một nền tảng podcast có hàng ngàn người dùng. Thuật toán ML phát hiện engagement thấp dựa trên cửa sổ trượt 10 phút (10-minute running window) từ các sự kiện người dùng như nghe (listening), tạm dừng (pausing), và đóng (closing) podcast. ML specialist cần biến đổi dữ liệu để chuẩn bị cho inference ML, với yêu cầu ít nỗ lực vận hành nhất (LEAST operational effort).
🔑 Yêu cầu cốt lõi:
- Xử lý dữ liệu streaming thời gian thực (real-time events).
- Áp dụng windowing 10 phút để tính toán engagement.
- Giải pháp phải serverless hoặc managed cao, giảm thiểu quản lý infra (như cluster, state management thủ công).
- Phù hợp với quy mô lớn (thousands of users), kiến thức AWS mới nhất đến 2026: Sử dụng Amazon Kinesis Data Streams cho ingestion, Amazon Managed Service for Apache Flink (trước đây là Kinesis Data Analytics) cho transformation windowing native.
✅ Đáp án đúng
Use Amazon Kinesis Data Streams to ingest event data. Use Amazon Kinesis Data Analytics to transform the most recent 10 minutes of data before inference.
Lý do lựa chọn:
- Kinesis Data Streams (KDS) là dịch vụ fully managed streaming lý tưởng cho ingestion real-time với throughput cao, không cần quản lý server.
- Kinesis Data Analytics (KDA) hỗ trợ windowing tumbling/sliding native qua SQL/Flink, dễ dàng xử lý "most recent 10 minutes" mà không cần code stateful phức tạp. Đây là giải pháp serverless hoàn toàn, tự động scale, least effort (chỉ config query/window).
- Tổng thể: End-to-end streaming pipeline với zero operational overhead cho windowing, phù hợp inference real-time. AWS khuyến nghị pattern này cho ML streaming (Feature Store + SageMaker inference).
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc. Mỗi phương án được đánh giá dựa trên operational effort, khả năng real-time windowing 10 phút, và tính managed/serverless.
-
❌ Use an Amazon Managed Streaming for Apache Kafka (Amazon MSK) cluster to ingest event data. Use Amazon Kinesis Data Analytics to transform the most recent 10 minutes of data before inference.
Sai vì: MSK yêu cầu quản lý cluster (provisioning, scaling, monitoring, patching), tăng operational effort cao so với KDS serverless. KDA có thể kết nối MSK nhưng không optimal, vi phạm "LEAST effort". (MSK phù hợp enterprise Kafka, không phải minimal effort). -
❌ Use Amazon Kinesis Data Streams to ingest event data. Store the data in Amazon S3 by using Amazon Kinesis Data Firehose. Use AWS Lambda để transform the most recent 10 minutes of data before inference.
Sai vì: Firehose batch dữ liệu vào S3 (không real-time thuần), mất tính streaming window 10 phút liên tục. Lambda phải track state thủ công (ví dụ dùng DynamoDB cho window), phức tạp, dễ lỗi, và cold start với high volume → high effort. Không native cho running window. -
✅ Use Amazon Kinesis Data Streams to ingest event data. Use Amazon Kinesis Data Analytics to transform the most recent 10 minutes of data before inference.
Đúng vì: Như giải thích trên, KDS + KDA là streaming-native, windowing built-in (Flink/SQL), fully managed, scale tự động. Least effort: Deploy query đơn giản, output trực tiếp cho inference (SageMaker/Kinesis). -
❌ Use an Amazon Managed Streaming for Apache Kafka (Amazon MSK) cluster to ingest event data. Use AWS Lambda to transform the most recent 10 minutes of data before inference.
Sai vì: MSK đã tốn effort quản lý cluster. Lambda với MSK cần consumer groups + state management thủ công cho 10-min window (Kinesis Client Library hoặc custom), rất phức tạp, không scalable real-time → highest effort, dễ fail với thousands users.
🛠️ Khuyến nghị triển khai thực tế (dựa AWS best practices 2026)
- Config KDA với Apache Flink cho tumbling window 10 phút:
TUMBLE(event_time, INTERVAL '10' MINUTE). - Output từ KDA → Amazon SageMaker Endpoint cho inference real-time.
- Monitoring: CloudWatch + KDA metrics (MillisBehindLatest).
📘 Tài liệu tham khảo
- AWS Docs: Amazon Kinesis Data Analytics for Windowed Aggregations (cập nhật Flink 1.18+).
- AWS Well-Architected: Streaming Data Lake on AWS (pattern KDS + KDA).
- Exam Guide DOP-C02: Real-time ML processing với least ops (2024-2026 blueprint).
- Blog: Real-time Podcast Analytics with Kinesis (tương tự case study).
Which solution will improve recall in the LEAST amount of time?
- A Add class weights to the MLP's loss function, and then retrain.
- B Gather more data by using Amazon Mechanical Turk, and then retrain.
- C Train a k-means algorithm instead of an MLP.
- D Train an anomaly detection model instead of an MLP.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào một chuyên gia Machine Learning (ML) đang huấn luyện mô hình Multilayer Perceptron (MLP) – một loại mạng nơ-ron nhân tạo cơ bản – trên bộ dữ liệu đa lớp (multi-class). Lớp mục tiêu (target class) độc đáo và khác biệt so với các lớp khác, nhưng recall (tỷ lệ phát hiện đúng các mẫu thuộc lớp mục tiêu) không đạt yêu cầu. Chuyên gia đã thử thay đổi số lượng và kích thước các lớp ẩn (hidden layers) của MLP, nhưng kết quả không cải thiện đáng kể.
📌 Vấn đề cốt lõi: Bộ dữ liệu có sự mất cân bằng lớp (class imbalance), lớp mục tiêu hiếm gặp (minority class), dẫn đến recall thấp cho lớp đó. Câu hỏi yêu cầu giải pháp cải thiện recall nhanh nhất (LEAST amount of time), nghĩa là ưu tiên phương pháp ít tốn thời gian triển khai và huấn luyện nhất, phù hợp với môi trường AWS như SageMaker (phiên bản mới nhất 2026 hỗ trợ class weighting tự động qua các estimator như PyTorch hoặc TensorFlow).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Add class weights to the MLP's loss function, and then retrain.
🛠️ Lý do:
- Đây là giải pháp nhanh nhất vì chỉ cần thêm trọng số lớp (class weights) vào hàm mất mát (loss function) như CrossEntropyLoss trong PyTorch hoặc TensorFlow trên AWS SageMaker, sau đó retrain ngay lập tức mà không cần thu thập dữ liệu mới hay thay đổi mô hình.
- Class weights ưu tiên lớp thiểu số (target class) bằng cách tăng penalty cho lỗi dự đoán lớp đó, giúp cải thiện recall nhanh chóng (thường chỉ mất vài phút đến giờ tùy kích thước data).
- Trong AWS SageMaker (cập nhật 2026), hỗ trợ trực tiếp qua
class_weightparameter trong các built-in algorithms hoặc custom scripts, không cần code phức tạp. - ✅ Hiệu quả cao với class imbalance, đặc biệt khi target class "unique" (hiếm), và không ảnh hưởng đến thời gian huấn luyện nhiều.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh, với giải thích bằng tiếng Việt sử dụng emoji để nổi bật:
-
✅ Add class weights to the MLP's loss function, and then retrain.
🟢 Đúng: Như đã giải thích ở trên, đây là cách tối ưu thời gian nhất (chỉ chỉnh loss function và retrain), trực tiếp giải quyết imbalance mà không cần data mới hay model khác. Phù hợp AWS SageMaker Training Jobs (phiên bản 2026 hỗ trợ automatic weighting qua SageMaker Debugger). -
❌ Gather more data by using Amazon Mechanical Turk, and then retrain.
🔴 Sai: Thu thập dữ liệu mới qua Amazon Mechanical Turk (dịch vụ crowdsourcing của AWS) tốn rất nhiều thời gian (ngày đến tuần để label data chất lượng cao), chi phí cao, và không đảm bảo cải thiện recall ngay (data mới có thể không cân bằng). Không phải giải pháp "LEAST time". -
❌ Train a k-means algorithm instead of an MLP.
🔴 Sai: K-means là thuật toán clustering không giám sát (unsupervised), không phù hợp cho bài toán phân loại đa lớp có nhãn (supervised multi-class) như MLP. Nó không xử lý class imbalance hay cải thiện recall cho target class cụ thể, và phải retrain từ đầu tốn thời gian hơn. -
❌ Train an anomaly detection model instead of an MLP.
🔴 Sai: Anomaly detection (như Isolation Forest hoặc Autoencoder trên SageMaker) phù hợp phát hiện outlier không có nhãn lớp cụ thể, nhưng ở đây là multi-class supervised với target class "unique". Chuyển model đòi hỏi retrain hoàn toàn, thiết kế lại pipeline, tốn thời gian nhiều hơn class weighting.
📘 Tài liệu tham khảo (cập nhật AWS 2026)
- AWS SageMaker Documentation: Handle Imbalanced Datasets – Hướng dẫn class weights trong loss functions cho PyTorch/TensorFlow estimators.
- AWS ML Best Practices: Class Imbalance in SageMaker – Ví dụ code thêm weights để boost recall nhanh.
- Mechanical Turk: MTurk for Data Labeling – Xác nhận tốn thời gian so với weighting.
- SageMaker Built-in Algorithms: Hỗ trợ anomaly detection/k-means, nhưng không optimal cho multi-class recall (AWS re:Invent 2025 updates).
🧠 Kết luận: Giải pháp class weights là best practice nhanh nhất cho imbalance trên AWS, giúp đạt DOP-C01 certification level! 🚀
Which combination of actions will meet these requirements with the LEAST operational overhead? (Choose two.)
- A Use SageMaker Clarify to automatically detect data bias
- B Turn on the bias detection option in SageMaker Ground Truth to automatically analyze data features.
- C Use SageMaker Model Monitor to generate a bias drift report.
- D Configure SageMaker Data Wrangler to generate a bias report.
- E Use SageMaker Experiments to perform a data check
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào một chuyên gia Machine Learning (ML) đã tải lên 5 TB dữ liệu vào môi trường Amazon SageMaker Studio. Sau khi thực hiện làm sạch dữ liệu ban đầu (initial data cleansing), chuyên gia cần tạo và xem báo cáo phân tích chi tiết về bias tiềm ẩn (potential bias) trong dữ liệu đã tải lên trước khi bắt đầu huấn luyện mô hình.
Yêu cầu chính là chọn kết hợp 2 hành động (choose two) để đáp ứng với ít overhead vận hành nhất (LEAST operational overhead).
- Bias ở đây đề cập đến sự thiên lệch trong dữ liệu (data bias), như phân bố không đồng đều giữa các nhóm (ví dụ: giới tính, chủng tộc), có thể ảnh hưởng đến mô hình ML.
- SageMaker Studio là IDE tích hợp cho ML, hỗ trợ các công cụ như Clarify, Data Wrangler để xử lý bias tự động, giảm thiểu công sức thủ công.
- Overhead thấp nghĩa là ưu tiên các tính năng tích hợp sẵn (built-in), tự động, không cần code phức tạp hoặc dịch vụ bên ngoài (dựa trên AWS SageMaker phiên bản mới nhất 2024-2026, với Clarify và Data Wrangler được tối ưu cho quy mô lớn như 5TB).
✅ Đáp án đúng (Chọn 2)
Hai lựa chọn đúng là:
- Use SageMaker Clarify to automatically detect data bias
- Configure SageMaker Data Wrangler to generate a bias report
Lý do chọn:
- Cả hai đều là các công cụ tích hợp sẵn trong SageMaker Studio, hỗ trợ phát hiện bias tự động trên dữ liệu lớn (5TB) mà không cần code thủ công, chỉ cần cấu hình đơn giản.
- Clarify quét bias trên dataset gốc (pre-training), tạo báo cáo chi tiết về các metric như Class Imbalance, Intersectional Bias.
- Data Wrangler (flow tích hợp trong Studio) sinh báo cáo bias qua vài cú click, hỗ trợ transform data và visualize bias metrics ngay lập tức.
- Kết hợp chúng mang lại báo cáo toàn diện với overhead thấp nhất, phù hợp quy trình ML end-to-end trong Studio (theo AWS Well-Architected ML Lens 2024).
🔍 Phân tích tất cả các phương án
Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên văn bản gốc bằng tiếng Anh cho phương án, và sử dụng ✅ (đúng) hoặc ❌ (sai) kèm giải thích bằng tiếng Việt:
-
✅ Use SageMaker Clarify to automatically detect data bias
Đúng vì: SageMaker Clarify là công cụ chuyên biệt để phát hiện và giảm thiểu bias trong dữ liệu và mô hình. Nó hỗ trợ tự động quét dataset lớn (như 5TB), tạo báo cáo bias metrics (ví dụ: Demographic Parity, Conditional Demographic Disparity) chỉ với vài dòng code hoặc qua UI Studio. Hoàn hảo cho giai đoạn pre-training, overhead thấp nhờ tích hợp serverless. (Cập nhật 2024: Hỗ trợ multi-modal data và bias facet mới). -
❌ Turn on the bias detection option in SageMaker Ground Truth to automatically analyze data features.
Sai vì: SageMaker Ground Truth là dịch vụ gán nhãn dữ liệu (labeling service), bias detection chỉ áp dụng cho labeling jobs (nhãn do con người tạo), không phải dữ liệu thô đã upload và cleansed. Nó không tự động phân tích features bias trên dataset gốc, và yêu cầu setup labeling workflow phức tạp, tăng overhead không cần thiết. -
❌ Use SageMaker Model Monitor to generate a bias drift report.
Sai vì: SageMaker Model Monitor dùng để giám sát mô hình sau triển khai (post-deployment), phát hiện drift/bias trong predictions (inference time), không phải kiểm tra bias trên dữ liệu pre-training. Nó cần endpoint đã deploy, không phù hợp giai đoạn trước train, và overhead cao hơn do yêu cầu baseline model. -
✅ Configure SageMaker Data Wrangler to generate a bias report.
Đúng vì: Data Wrangler là công cụ visual data prep trong SageMaker Studio, cho phép tạo bias report chỉ qua UI (drag-and-drop) trên dataset lớn. Nó tính toán bias metrics (như class imbalance, intersectional) và visualize ngay, tích hợp liền mạch với cleansing flow. Overhead thấp nhất cho exploratory analysis (cập nhật 2025: Tối ưu cho petabyte-scale với Spark integration). -
❌ Use SageMaker Experiments to perform a data check
Sai vì: SageMaker Experiments dùng để track và quản lý thí nghiệm ML (metadata, params, metrics), không có tính năng built-in phát hiện bias hoặc sinh báo cáo bias. Nó chỉ log data check thủ công (nếu code thêm), overhead cao vì cần tự implement, không tự động như Clarify/Data Wrangler.
📘 Tài liệu tham khảo (AWS chính thức, cập nhật 2024-2026)
- SageMaker Clarify: AWS Docs - Detect Bias – Hướng dẫn detect data bias tự động.
- SageMaker Data Wrangler: AWS Blog - Bias Report in Data Wrangler – Chi tiết generate bias report.
- SageMaker Studio Features: AWS Well-Architected ML Lens – Best practices cho bias detection với least overhead.
- Exam Guide DOP-C02: Phần SageMaker/ML Ops nhấn mạnh Clarify & Data Wrangler cho data quality checks.
🛠️ Lời khuyên DevOps: Trong pipeline CI/CD, tích hợp Clarify/Data Wrangler qua Step Functions để automate bias checks, đảm bảo MLOps scalable cho production! Nếu cần demo code, hãy hỏi thêm nhé! 🚀