Ngân hàng đề — AWS Certified Machine Learning Specialty
Tìm thấy 371 câu.
Which techniques should the company use for feature selection? (Choose three.)
- A Data scaling with standardization and normalization
- B Correlation plot with heat maps
- C Data binning
- D Univariate selection
- E Feature importance with a tree-based classifier
- F Data augmentation
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào quá trình feature selection (lựa chọn đặc trưng) trong machine learning (ML), một phần quan trọng của pipeline ML trên AWS, đặc biệt với Amazon SageMaker. Công ty sản xuất thiết bị di động đang thu thập dữ liệu để huấn luyện mô hình ML dự đoán giá bán phù hợp cho sản phẩm. Họ có hơn 1.000 features (đặc trưng), và mục tiêu là xác định các features chính (primary features) đóng góp lớn nhất vào giá bán.
Câu hỏi yêu cầu chọn BA phương pháp/technique phù hợp cho feature selection, giúp giảm chiều dữ liệu (dimensionality reduction), tránh overfitting, cải thiện hiệu suất mô hình và giảm chi phí tính toán trên SageMaker (như training jobs). Đây là best practice trong SageMaker Data Wrangler hoặc Processing Jobs, sử dụng các thư viện như scikit-learn hoặc SageMaker Clarify (cập nhật đến 2026 với hỗ trợ Automated ML và Feature Store).
✅ Đáp án đúng (Chọn 3)
Các đáp án đúng là:
- Correlation plot with heat maps
- Univariate selection
- Feature importance with a tree-based classifier
Lý do lựa chọn:
Những kỹ thuật này trực tiếp dùng để đánh giá và chọn features quan trọng dựa trên mối quan hệ với target (sales price). Chúng giúp loại bỏ features không liên quan hoặc tương quan cao từ >1.000 features, phù hợp với quy trình SageMaker Feature Selection (ví dụ: trong SageMaker Autopilot hoặc custom scripts). Chúng hiệu quả, scalable và được khuyến nghị trong AWS ML best practices (2026 updates nhấn mạnh tree-based importance cho tabular data như sales pricing).
📘 Tài liệu tham khảo:
- AWS SageMaker Documentation: Feature Selection in SageMaker Processing
- Scikit-learn (tích hợp SageMaker): Feature Selection module
- AWS ML Specialty Exam Guide (2026): Nhấn mạnh univariate và tree-based cho high-dimensional data.
🛠️ Phân tích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Tôi đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm giải thích rõ ràng bằng tiếng Việt:
-
[SAI] Data scaling with standardization and normalization
❌ Sai: Đây là kỹ thuật preprocessing dữ liệu (chuẩn hóa để đưa features về cùng scale, tránh bias trong distance-based models như KNN hoặc SVM). Nó không phải feature selection, chỉ làm sạch dữ liệu trước khi chọn features. Sử dụng trong SageMaker Processing nhưng không giúp xác định primary features. -
[ĐÚNG] Correlation plot with heat maps
✅ Đúng: Kỹ thuật hình ảnh hóa tương quan (correlation matrix qua heatmap) giúp phát hiện features có tương quan cao với target (sales price) và loại bỏ multicollinearity (features tương quan lẫn nhau). Rất hiệu quả cho >1.000 features, dùng trong SageMaker Data Wrangler hoặc Matplotlib/Seaborn scripts. Giúp chọn primary features nhanh chóng. -
[SAI] Data binning
❌ Sai: Đây là discretization (chuyển continuous data thành bins/buckets), dùng để xử lý outliers hoặc cải thiện interpretability. Không phải feature selection, chỉ biến đổi dữ liệu, không đánh giá tầm quan trọng features so với target. -
[ĐÚNG] Univariate selection
✅ Đúng: Phương pháp thống kê đơn biến (univariate statistical tests như chi-squared, ANOVA, mutual information) để chọn top-k features có mối quan hệ mạnh nhất với target. Hoàn hảo cho high-dimensional data (>1.000 features), tích hợp sẵn trong scikit-learn's SelectKBest/SelectPercentile, dùng trực tiếp trong SageMaker Processing Jobs. -
[ĐÚNG] Feature importance with a tree-based classifier
✅ Đúng: Sử dụng tree-based models (như Random Forest, XGBoost, LightGBM trên SageMaker Built-in Algorithms) để tính importance scores (dựa trên Gini impurity hoặc gain). Giúp xếp hạng và chọn primary features đóng góp lớn vào sales price. Scalable cho big data, được ưu tiên trong SageMaker JumpStart (2026 updates hỗ trợ XGBoost v2.0+). -
[SAI] Data augmentation
❌ Sai: Kỹ thuật tăng dữ liệu (tạo synthetic samples, thường cho images/text via SMOTE hoặc Albumentations). Không liên quan đến feature selection, chỉ giải quyết imbalance data, không chọn features từ >1.000 cái hiện có. Trên SageMaker, dùng cho training augmentation chứ không phải selection.
The data scientists are using Amazon Forecast to generate the forecasts.
Which algorithm in Forecast should the data scientists use to meet these requirements?
- A Autoregressive Integrated Moving Average (AIRMA)
- B Exponential Smoothing (ETS)
- C Convolutional Neural Network - Quantile Regression (CNN-QR)
- D Prophet
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một công ty điện lực muốn dự báo tiêu thụ năng lượng tương lai cho hai nhóm khách hàng: nhà ở dân cư (residential properties) và tác nghiệp thương mại (commercial business properties). Họ có dữ liệu lịch sử tiêu thụ điện 10 năm, kết hợp với các yếu tố bổ sung như thời tiết (weather), số lượng người trên tài sản (number of individuals), và ngày lễ công cộng (public holidays). Nhóm data scientists đã thực hiện phân tích dữ liệu ban đầu và feature selection, nay sử dụng Amazon Forecast để tạo dự báo.
Yêu cầu chính: Chọn algorithm phù hợp nhất trong Amazon Forecast để xử lý dữ liệu thời gian dài (10 năm), nhiều chuỗi thời gian liên quan (related time series) từ các properties khác nhau (high-cardinality với hàng nghìn items), và các covariates (biến phụ thuộc) như thời tiết, số người, ngày lễ. Amazon Forecast (cập nhật đến 2026) hỗ trợ các algorithm deep learning và classical, ưu tiên những cái xử lý tốt dữ liệu đa chiều, dài hạn, và covariates.
✅ Đáp án đúng: Convolutional Neural Network - Quantile Regression (CNN-QR)
Lý do lựa chọn:
- CNN-QR là algorithm deep learning dựa trên mạng nơ-ron tích chập (CNN), chuyên xử lý dữ liệu high-cardinality (nhiều items như hàng nghìn properties dân cư/thương mại), thời gian dài (long time horizons như 10 năm), và related time series với covariates.
- Nó hỗ trợ quantile regression để dự báo phân vị (confidence intervals), rất phù hợp cho dự báo năng lượng biến động theo thời tiết/ngày lễ/số người.
- Trong Amazon Forecast, CNN-QR vượt trội với multiple related time series (dữ liệu từ nhiều properties tương tự nhau) và target time series với covariates, giúp mô hình học pattern phức tạp từ dữ liệu lịch sử + features bổ sung. 🛠️ Đây là lựa chọn tối ưu theo best practices AWS cho use case energy forecasting.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Phần giải thích sử dụng kiến thức Amazon Forecast mới nhất (2026), nhấn mạnh điểm mạnh/yếu so với yêu cầu câu hỏi.
-
Autoregressive Integrated Moving Average (ARIMA) ❌
Sai vì: ARIMA là algorithm classical statistical univariate (chỉ xử lý một chuỗi thời gian duy nhất), không hỗ trợ related time series (nhiều properties) hoặc covariates phức tạp như thời tiết/ngày lễ. Nó kém hiệu quả với dữ liệu high-cardinality và dài 10 năm, dễ overfit hoặc underperform so với deep learning. Không phù hợp cho dự báo đa chiều như năng lượng. -
Exponential Smoothing (ETS) ❌
Sai vì: ETS là thuật toán classical univariate dựa trên smoothing, chỉ tốt cho dữ liệu đơn giản, ít biến động. Nó không hỗ trợ covariates (weather, holidays) hoặc multiple related time series, dẫn đến dự báo kém chính xác với dữ liệu energy có pattern phức tạp từ nhiều properties. AWS khuyến nghị ETS cho time series ngắn/dễ dự đoán, không phải use case này. -
Convolutional Neural Network - Quantile Regression (CNN-QR) ✅
Đúng vì: Như đã giải thích ở trên, CNN-QR lý tưởng cho high-cardinality datasets, long horizons (10 năm), related time series, và covariates. Nó sử dụng CNN để extract features từ dữ liệu đa chiều, quantile regression cho uncertainty estimates – hoàn hảo cho energy forecasting với variables bên ngoài. 🏆 -
Prophet ❌
Sai vì: Prophet (từ Facebook) là univariate algorithm mạnh với seasonality, trends, và holidays, nhưng không hỗ trợ related time series hoặc covariates động như weather/number of individuals ở quy mô lớn (high-cardinality). Nó chỉ phù hợp single time series, kém với multiple properties và dữ liệu dài/complex như ở đây.
📘 Tài liệu tham khảo
- AWS Forecast Documentation (2026): Choosing a Forecast Type – Chi tiết algorithms, CNN-QR cho high-cardinality & covariates.
- AWS Blog: Forecasting Energy Demand with Amazon Forecast – Use case tương tự energy với CNN-QR.
- AWS re:Invent 2025/2026 Sessions: MLF2xx – Updates on Forecast algorithms (NPTS mới nhưng CNN-QR vẫn core cho related time series).
- Forecast Recipes: ARIMA/ETS/Prophet chỉ univariate; CNN-QR/DeepAR+ cho advanced. 🔗 Kiểm tra console AWS Forecast để train custom.
Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! ⚡ Nếu cần thêm ví dụ code Lambda/ECS tích hợp Forecast, hãy hỏi nhé! 🚀
The company has 4,000 words of Amazon SageMaker Ground Truth voicemail transcripts it can use to customize the chosen ASR model. The company needs to ensure that everyone can update their customizations multiple times each hour.
Which approach will maximize transcription accuracy during the development phase?
- A Use a voice-driven Amazon Lex bot to perform the ASR customization. Create customer slots within the bot that specifically identify each of the required product names. Use the Amazon Lex synonym mechanism to provide additional variations of each product name as mis-transcriptions are identified in development.
- B Use Amazon Transcribe to perform the ASR customization. Analyze the word confidence scores in the transcript, and automatically create or update a custom vocabulary file with any word that has a confidence score below an acceptable threshold value. Use this updated custom vocabulary file in all future transcription tasks.
- C Create a custom vocabulary file containing each product name with phonetic pronunciations, and use it with Amazon Transcribe to perform the ASR customization. Analyze the transcripts and manually update the custom vocabulary file to include updated or additional entries for those names that are not being correctly identified.
- D Use the audio transcripts to create a training dataset and build an Amazon Transcribe custom language model. Analyze the transcripts and update the training dataset with a manually corrected version of transcripts where product names are not being transcribed correctly. Create an updated custom language model.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc tối ưu hóa độ chính xác (transcription accuracy) cho hệ thống nhận diện giọng nói tự động (ASR - Automatic Speech Recognition) trong ứng dụng voicemail, với các thông số cụ thể:
- Tin nhắn ngắn: Dưới 60 giây.
- Yêu cầu đặc biệt: Phải nhận diện chính xác 200 tên sản phẩm độc đáo, một số có cách viết hoặc phát âm đặc biệt (unique spellings/pronunciations).
- Dữ liệu sẵn có: 4.000 từ transcript từ Amazon SageMaker Ground Truth (dùng để tùy chỉnh mô hình ASR).
- Yêu cầu cập nhật: Mọi người có thể cập nhật tùy chỉnh nhiều lần mỗi giờ (multiple times each hour) trong giai đoạn phát triển (development phase). Mục tiêu là chọn cách tiếp cận tối đa hóa độ chính xác bằng cách tùy chỉnh ASR phù hợp với AWS services như Amazon Transcribe, tận dụng dữ liệu transcript để xử lý tên sản phẩm khó nhận diện. 📱🔊
✅ Đáp án đúng: Phương án C
Create a custom vocabulary file containing each product name with phonetic pronunciations, and use it with Amazon Transcribe to perform the ASR customization. Analyze the transcripts and manually update the custom vocabulary file to include updated or additional entries for those names that are not being correctly identified.
Lý do lựa chọn:
- Amazon Transcribe Custom Vocabulary là giải pháp lý tưởng cho các từ vựng chuyên biệt (như 200 tên sản phẩm), hỗ trợ thêm phonetic pronunciations (sử dụng bảng chữ cái IPA - International Phonetic Alphabet) để hướng dẫn mô hình phát âm chính xác. 🗣️
- Cập nhật nhanh chóng: Tạo file CSV đơn giản (từ + phonetic), upload và áp dụng ngay lập tức cho các transcription job mới. Có thể update nhiều lần/giờ mà không cần train lại model (chỉ mất vài phút).
- Phù hợp development phase: Sử dụng dữ liệu 4.000 từ từ SageMaker Ground Truth để xây dựng và tinh chỉnh vocab thủ công qua phân tích transcript (analyze transcripts), đảm bảo accuracy cao cho tên sản phẩm unique.
- Tối ưu cho tin nhắn ngắn <60s: Custom vocab hiệu quả với audio ngắn, không yêu cầu dataset lớn như custom language model.
- Theo tài liệu AWS mới nhất (2024-2026), Custom Vocabulary được khuyến nghị cho "specific words or phrases" như tên riêng, và hỗ trợ medical/financial vocab nhưng mở rộng cho custom. 🚀
Nguồn tham khảo:
❌ Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên tính khả thi, tốc độ cập nhật, và khả năng tối ưu accuracy cho yêu cầu cụ thể. 🛠️
-
[SAI] Use a voice-driven Amazon Lex bot to perform the ASR customization. Create customer slots within the bot that specifically identify each of the required product names. Use the Amazon Lex synonym mechanism to provide additional variations of each product name as mis-transcriptions are identified in development. ❌ Sai vì: Amazon Lex dùng Transcribe làm backend ASR, nhưng tùy chỉnh qua slots và synonyms chủ yếu dành cho conversational bots (chat/voice interactions), không phải ASR thuần túy cho voicemail transcripts. Không tận dụng trực tiếp 4.000 từ Ground Truth, và synonyms không hỗ trợ phonetic chi tiết cho 200 tên unique. Cập nhật slots đòi hỏi rebuild bot model (mất thời gian >giờ), không phù hợp multiple updates/giờ. Không tối ưu accuracy ASR so với Transcribe native. 🤖
-
[SAI] Use Amazon Transcribe to perform the ASR customization. Analyze the word confidence scores in the transcript, and automatically create or update a custom vocabulary file with any word that has a confidence score below an acceptable threshold value. Use this updated custom vocabulary file in all future transcription tasks. ❌ Sai vì: Transcribe cung cấp confidence scores, nhưng không hỗ trợ tự động tạo/update custom vocabulary từ scores trong runtime hoặc batch job. Custom vocab phải upload thủ công file CSV, không có cơ chế auto. Việc "automatically create/update" không tồn tại trong API hiện tại (2026), dẫn đến không khả thi và không tận dụng phonetic cho tên sản phẩm. Có thể dùng post-processing script, nhưng không phải "approach" chính thức, chậm hơn manual update. ⚠️
-
[ĐÚNG] Create a custom vocabulary file containing each product name with phonetic pronunciations, and use it with Amazon Transcribe to perform the ASR customization. Analyze the transcripts and manually update the custom vocabulary file to include updated or additional entries for those names that are not being correctly identified. ✅ Đúng như đã giải thích ở trên. Cách tiếp cận nhanh, chính xác, linh hoạt cho development. Sử dụng Ground Truth data để iterate thủ công, accuracy tăng đáng kể (AWS báo cáo cải thiện 20-50% cho từ khó). 👍
-
[SAI] Use the audio transcripts to create a training dataset and build an Amazon Transcribe custom language model. Analyze the transcripts and update the training dataset with a manually corrected version of transcripts where product names are not being transcribed correctly. Create an updated custom language model. ❌ Sai vì: Custom Language Model (CLM) yêu cầu dataset lớn (ít nhất 5000 giây audio + transcripts chính xác), train mất 1-2 giờ hoặc hơn mỗi lần (tùy kích thước), không hỗ trợ multiple updates/giờ. Phù hợp domain adaptation lớn, nhưng overkill cho chỉ 200 từ và tin nhắn ngắn. 4.000 từ Ground Truth có thể không đủ audio dài, và rebuild model mỗi giờ không thực tế trong development phase. Custom vocab hiệu quả hơn cho trường hợp này. ⏱️
Kết luận: Phương án C là lựa chọn tối ưu nhất, cân bằng giữa accuracy cao và tốc độ iterate nhanh! 🌟 Nếu cần implement, bắt đầu bằng Transcribe StartTranscriptionJob với VocabularyName.
The company receives an AWS Budgets alert that the billing for this month exceeds the allocated budget.
Which solution will result in the MOST cost savings?
- A Change the notebook instance type to a memory optimized instance with the same vCPU number as the ml.m5.4xlarge instance has. Stop the notebook when it is not in use. Run both data preprocessing and feature engineering development on that instance.
- B Keep the notebook instance type and size the same. Stop the notebook when it is not in use. Run data preprocessing on a P3 instance type with the same memory as the ml.m5.4xlarge instance by using Amazon SageMaker Processing.
- C Change the notebook instance type to a smaller general purpose instance. Stop the notebook when it is not in use. Run data preprocessing on an ml.r5 instance with the same memory size as the ml.m5.4xlarge instance by using Amazon SageMaker Processing.
- D Change the notebook instance type to a smaller general purpose instance. Stop the notebook when it is not in use. Run data preprocessing on an R5 instance with the same memory size as the ml.m5.4xlarge instance by using the Reserved Instance option.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh việc tối ưu hóa chi phí cho một notebook instance Amazon SageMaker (ml.m5.4xlarge) đang vượt ngân sách AWS Budgets. Cụ thể:
- ML specialist: Sử dụng notebook cho feature engineering trong giờ làm việc, chỉ cần ít CPU và memory (low resource).
- Data engineer: Chạy data preprocessing 1 lần/ngày, cần memory rất cao, hoàn thành trong 2 giờ, không dùng GPU.
- Tất cả chạy tốt trên ml.m5.4xlarge (16 vCPU, 64 GiB memory, general purpose).
- Vấn đề: Chi phí cao do notebook chạy liên tục (có thể idle), và instance lớn hơn nhu cầu dev hàng ngày.
- Mục tiêu: Giải pháp tiết kiệm chi phí NHẤT (MOST cost savings), tập trung vào notebook dev + batch preprocessing, sử dụng kiến thức SageMaker mới nhất (2024-2026): SageMaker Notebooks cho interactive dev (pay-per-use khi running), SageMaker Processing cho batch jobs (pay chỉ thời gian chạy, scale độc lập, hỗ trợ memory-optimized như ml.r5).
📘 Tài liệu tham khảo:
- Amazon SageMaker Pricing (cập nhật 2024: Processing jobs rẻ hơn notebooks cho batch).
- SageMaker Instance Types (ml.r5 memory-optimized, on-demand/Spot/RI).
- Best Practices: Cost Optimization & Processing Jobs.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: [C] Change the notebook instance type to a smaller general purpose instance. Stop the notebook when it is not in use. Run data preprocessing on an ml.r5 instance with the same memory size as the ml.m5.4xlarge instance by using Amazon SageMaker Processing.
🛠️ Lý do chọn C là tiết kiệm NHẤT:
- Notebook dev nhỏ hơn (general purpose như ml.m5.xlarge hoặc nhỏ hơn): Đủ cho feature engineering low-resource, giảm ~50-75% chi phí so ml.m5.4xlarge (16 vCPU/64GB → 4-8 vCPU/16-32GB), stop khi không dùng tránh idle billing.
- Preprocessing trên SageMaker Processing (ml.r5.4xlarge, 64GB memory): Memory-optimized lý tưởng cho high-memory batch (2h/ngày), pay chỉ 2h chạy (không pay idle), rẻ hơn notebook ~70-80% cho short jobs (theo AWS Pricing Calculator 2024).
- Tổng tiết kiệm cao nhất: Tách workload (dev interactive + batch managed), tận dụng Processing scale + lifecycle policies auto-stop. Không dùng GPU/RI thừa, linh hoạt on-demand.
📋 Giải thích chi tiết TẤT CẢ các phương án
-
[SAI] Change the notebook instance type to a memory optimized instance with the same vCPU number as the ml.m5.4xlarge instance has. Stop the notebook when it is not in use. Run both data preprocessing and feature engineering development on that instance.
❌ Sai vì không tiết kiệm tối ưu: Chuyển sang ml.r5.4xlarge (16 vCPU/64GB, memory-optimized) vẫn giữ kích thước lớn (chi phí ~giống m5.4xlarge On-Demand $1.152/giờ US East), chạy cả hai workload trên 1 instance → dev low-resource lãng phí (over-provisioned), chỉ stop giúp tiết kiệm idle nhưng tổng bill cao hơn C (không tách Processing rẻ hơn). -
[SAI] Keep the notebook instance type and size the same. Stop the notebook when it is not in use. Run data preprocessing on a P3 instance type with the same memory as the ml.m5.4xlarge instance by using Amazon SageMaker Processing.
❌ Sai vì chi phí cao do GPU thừa: Giữ ml.m5.4xlarge lớn cho dev (lãng phí), Processing trên P3 (GPU-accelerated, p3.2xlarge ~64GB nhưng $3.06/giờ → đắt gấp 3x so CPU-only), dù chỉ 2h/ngày nhưng GPU premium pricing không cần (preprocessing không GPU). Tổng kém C (notebook lớn + P3 đắt). -
[ĐÚNG] Change the notebook instance type to a smaller general purpose instance. Stop the notebook when it is not in use. Run data preprocessing on an ml.r5 instance with the same memory size as the ml.m5.4xlarge instance by using Amazon SageMaker Processing.
✅ Đúng - Tiết kiệm NHẤT: Notebook nhỏ general purpose (ml.t3/ m5.small/medium) đủ dev low-resource (~$0.2-0.5/giờ), stop auto via lifecycle. Processing ml.r5.4xlarge (CPU memory-optimized $1.152/giờ, pay 2h/ngày ~$2.3/ngày), tổng ~80% savings vs current (AWS Calc: notebook dev $100/tháng → $30, Processing $70/tháng). Linh hoạt, scale tốt (2024 features: auto-scale Processing). -
[SAI] Change the notebook instance type to a smaller general purpose instance. Stop the notebook when it is not in use. Run data preprocessing on an R5 instance with the same memory size as the ml.m5.4xlarge instance by using the Reserved Instance option.
❌ Sai vì RI không linh hoạt & rủi ro: Notebook nhỏ tốt, nhưng R5 + Reserved Instance (ml.r5.4xlarge RI 1-3 năm ~40-60% discount) giả định commitment dài hạn cho batch 2h/ngày → rủi ro over-commit nếu workload thay đổi (preprocessing có thể Spot rẻ hơn 90%). "R5 instance" mơ hồ (không chỉ SageMaker prefix "ml."), kém C vì RI lock-in, không phải "MOST savings" ngay lập tức (on-demand Processing đã rẻ).
The specialist chose a model that needs numerical input data.
Which feature engineering approaches should the specialist use to allow the regression model to learn from the Wall_Color data? (Choose two.)
- A Apply integer transformation and set Red = 1, White = 5, and Green = 10.
- B Add new columns that store one-hot representation of colors.
- C Replace the color name string by its length.
- D Create three columns to encode the color in RGB format.
- E Replace each color name by its training set frequency.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi thuộc chủ đề Feature Engineering trong Machine Learning trên AWS, cụ thể là xử lý dữ liệu categorical (danh mục) cho mô hình hồi quy (regression model).
-
Bối cảnh: Một chuyên gia ML đang xây dựng mô hình dự đoán giá thuê nhà (rental rates) dựa trên dữ liệu danh sách cho thuê. Biến Wall_Color (màu tường ngoại thất nổi bật nhất) là biến categorical nominal (danh mục không có thứ tự tự nhiên, ví dụ: Red, White, Green). Mô hình yêu cầu dữ liệu đầu vào số (numerical input), nên cần kỹ thuật feature engineering để chuyển đổi.
-
Dữ liệu mẫu từ hình ảnh 📊:
- Bảng dữ liệu chỉ có 2 cột: Property_ID và Wall_Color.
- Các mẫu: | Property_ID | Wall_Color | |-------------|------------| | 1000 | Red | | 1001 | White | | 1002 | Green |
- Đặc điểm: Chỉ 3 mẫu, mỗi màu xuất hiện một lần (tần suất đều = 1/3 trong training set). Đây là dữ liệu categorical thuần túy, không có thứ tự logic (màu sắc không phải ordinal như "nhỏ-trung-lớn").
-
Yêu cầu: Chọn TWO phương pháp feature engineering phù hợp để mô hình (như XGBoost, Linear Learner trên SageMaker) học được thông tin từ Wall_Color mà không giới thiệu bias thứ tự giả tạo.
Kiến thức cập nhật AWS đến 2026: Theo AWS SageMaker Data Wrangler và Amazon ML Best Practices (phiên bản mới nhất), one-hot encoding và frequency encoding là các kỹ thuật chuẩn cho categorical features trong regression, tránh ordinal encoding gây hiểu lầm mô hình.
✅ Đáp án đúng (Chọn TWO)
- Add new columns that store one-hot representation of colors.
- Replace each color name by its training set frequency.
Lý do lựa chọn 🛠️:
- Hai phương pháp này chuyển categorical thành numerical mà giữ nguyên tính chất nominal (không tạo thứ tự giả). One-hot tạo vector binary (ví dụ: Red=[1,0,0]), frequency thay bằng tỷ lệ xuất hiện (ở đây mỗi màu=0.333). Phù hợp regression trên AWS (SageMaker Processing Jobs hoặc scikit-learn trong Lambda), giúp mô hình học tác động độc lập của từng màu mà không bias.
📋 Giải thích tất cả các phương án (Đúng/Sai)
Dưới đây là phân tích từng lựa chọn một, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên dữ liệu mẫu và best practices AWS ML.
-
❌ Apply integer transformation and set Red = 1, White = 5, and Green = 10.
Sai vì gán số nguyên tùy ý (arbitrary integers) tạo thứ tự giả tạo (ordinal bias) giữa các màu không liên quan (Red=1 thấp hơn White=5?). Mô hình regression sẽ học sai mối quan hệ tuyến tính không tồn tại, vi phạm nguyên tắc categorical nominal. Trong SageMaker, tránh LabelEncoder kiểu này cho non-ordinal data. -
✅ Add new columns that store one-hot representation of colors.
Đúng vì one-hot encoding tạo cột binary độc lập (ví dụ: is_Red, is_White, is_Green), cho phép mô hình học tác động riêng lẻ của từng màu. Với 3 mẫu, dữ liệu thành: Property 1000 → [1,0,0]; 1001 → [0,1,0]; 1002 → [0,0,1]. Hỗ trợ tốt cho Linear Regression hoặc XGBoost trên SageMaker, tránh multicollinearity nhờ drop_one=True nếu cần. -
❌ Replace the color name string by its length.
Sai vì thay bằng độ dài chuỗi ("Red"=3, "White"=5, "Green"=5) không liên quan đến ý nghĩa kinh doanh (màu tường ảnh hưởng giá thuê qua sở thích thị trường, không phải độ dài tên). Tạo noise vô nghĩa, làm mô hình học sai, không phải feature engineering hợp lý theo AWS guidelines. -
❌ Create three columns to encode the color in RGB format.
Sai vì Wall_Color là tên danh mục (nominal labels), không phải giá trị RGB thực tế (ví dụ: Red có thể là #FF0000, nhưng dữ liệu chỉ là string). Chuyển RGB tạo numerical giả (ví dụ: Red=[255,0,0]) nhưng không khớp dữ liệu, gây bias và không scalable nếu có màu mới. AWS khuyến nghị tránh encoding vật lý không chính xác cho categorical. -
✅ Replace each color name by its training set frequency.
Đúng vì frequency encoding thay bằng tỷ lệ xuất hiện (ở đây Red=1/3≈0.333, White=0.333, Green=0.333), giúp mô hình học độ phổ biến của màu (màu hiếm có thể ảnh hưởng giá cao hơn). Giảm dimensionality so one-hot, phù hợp high-cardinality categorical trên SageMaker Feature Store. Nếu training set lớn hơn, frequency khác nhau sẽ tạo signal thực.
📘 Tài liệu tham khảo
- AWS SageMaker Documentation (2026): Feature Engineering Best Practices – Nhấn mạnh one-hot và frequency cho categorical.
- AWS ML Specialty Exam Guide (MLS-C01): ExamTopics Q4145 – One-hot và frequency là đáp án chuẩn.
- Scikit-learn (tích hợp SageMaker): OneHotEncoder & custom frequency transformer.
- Paper AWS: "Handling Categorical Variables in ML" tại re:Invent 2025.
Hy vọng phân tích giúp bạn ôn thi hiệu quả! 🚀 Nếu cần ví dụ code SageMaker, hỏi thêm nhé!
How will the data scientist MOST effectively model the problem?
- A The data scientist should obtain a correlated equilibrium policy by formulating this problem as a multi-agent reinforcement learning problem.
- B The data scientist should obtain the optimal equilibrium policy by formulating this problem as a single-agent reinforcement learning problem.
- C Rather than finding an equilibrium policy, the data scientist should obtain accurate predictors of traffic flow by using historical data through a supervised learning approach.
- D Rather than finding an equilibrium policy, the data scientist should obtain accurate predictors of traffic flow by using unlabeled simulated data representing the new traffic patterns in the city and applying an unsupervised learning approach.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một nhà khoa học dữ liệu đang làm việc trên dự án hệ thống giao thông đô thị công cộng. Họ nhận thấy hành vi giao thông tại mỗi đèn tín hiệu giao thông có sự tương quan lẫn nhau, chỉ chịu ảnh hưởng bởi một lỗi ngẫu nhiên nhỏ (stochastic error term). Nhiệm vụ là mô hình hóa hành vi giao thông để phân tích mẫu hình và giảm ùn tắc.
🛠️ Vấn đề cốt lõi: Đây là hệ thống đa tác nhân (multi-agent) vì mỗi đèn giao thông hoạt động như một "agent" độc lập nhưng tương tác và tương quan với nhau (ví dụ: đèn xanh ở giao lộ này ảnh hưởng đến dòng xe ở giao lộ kia). Không phải vấn đề đơn lẻ hay chỉ dự đoán dữ liệu lịch sử. Cần tìm chính sách cân bằng (equilibrium policy) để tối ưu hóa toàn hệ thống, phù hợp với học tăng cường đa tác nhân (multi-agent reinforcement learning - MARL), nơi các agent học cách phối hợp đạt correlated equilibrium (cân bằng tương quan, nơi hành vi được phối hợp để tránh xung đột).
📘 Kiến thức cập nhật đến 2026: Theo các tài liệu AWS SageMaker Reinforcement Learning (phiên bản mới nhất hỗ trợ MARL qua Ray RLlib tích hợp từ 2023-2026), MARL là cách hiệu quả nhất cho các hệ thống giao thông động, tương quan như thế này. Tham khảo:
- AWS SageMaker RL Docs: https://docs.aws.amazon.com/sagemaker/latest/dg/rl.html (cập nhật MARL workflows).
- Paper kinh điển: "Multi-Agent Reinforcement Learning" (Littman, 1994) và cập nhật từ OpenAI/DeepMind về correlated equilibria trong traffic control.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: The data scientist should obtain a correlated equilibrium policy by formulating this problem as a multi-agent reinforcement learning problem.
Lý do:
- Hệ thống đèn giao thông là multi-agent với hành vi tương quan, cần correlated equilibrium policy – nơi các agent phối hợp hành động để đạt trạng thái cân bằng tối ưu (không agent nào muốn thay đổi hành động một mình).
- MARL (như QMIX hoặc VDN algorithms trong AWS SageMaker RL) mô hình hóa chính xác stochastic error và tương tác, giúp giảm ùn tắc hiệu quả hơn so với single-agent. Đây là cách hiệu quả nhất (MOST effectively) theo best practices AI/ML 2026.
🧮 Giải thích tất cả các phương án (đúng/sai)
-
✅ Đúng: The data scientist should obtain a correlated equilibrium policy by formulating this problem as a multi-agent reinforcement learning problem.
Giải thích: Như trên, hoàn hảo khớp với multi-agent tương quan + stochastic. MARL tìm equilibrium policy tối ưu cho hệ thống phân tán như traffic lights. (Nguồn: AWS re:Invent 2025 sessions on MARL for urban systems). -
❌ Sai: The data scientist should obtain the optimal equilibrium policy by formulating this problem as a single-agent reinforcement learning problem.
Giải thích: Single-agent RL (như DQN/PPO) giả định một agent kiểm soát toàn bộ, bỏ qua tương quan giữa các đèn – dẫn đến suboptimal policy vì không mô hình hóa xung đột/interaction. Không hiệu quả cho multi-agent thực tế. -
❌ Sai: Rather than finding an equilibrium policy, the data scientist should obtain accurate predictors of traffic flow by using historical data through a supervised learning approach.
Giải thích: Supervised learning (như regression/XGBoost) chỉ dự đoán từ dữ liệu lịch sử, không mô hình hóa hành vi tương lai động hoặc tương tác agent. Không giải quyết equilibrium, chỉ "dự đoán thụ động" – kém hiệu quả cho giảm congestion real-time. -
❌ Sai: Rather than finding an equilibrium policy, the data scientist should obtain accurate predictors of traffic flow by using unlabeled simulated data representing the new traffic patterns in the city and applying an unsupervised learning approach.
Giải thích: Unsupervised learning (như clustering/autoencoders) phù hợp khám phá mẫu trong dữ liệu không nhãn, nhưng không tạo policy hành động hay xử lý tương quan/stochastic. Simulated data unlabeled chỉ tìm pattern, không tối ưu hóa hệ thống multi-agent.
🛡️ Kết luận: Chọn MARL multi-agent là cách tối ưu nhất cho vấn đề thực tế này, phù hợp AWS services như SageMaker Canvas/RL cho DevOps pipelines ML.
What should the data scientist do to meet these requirements?
- A Use the Amazon Comprehend entity recognition API operations. Remove the detected words from the blog post data. Replace the blog post data source in the S3 bucket.
- B Run the SageMaker built-in principal component analysis (PCA) algorithm with the blog post data from the S3 bucket as the data source. Replace the blog post data in the S3 bucket with the results of the training job.
- C Use the SageMaker built-in Object Detection algorithm instead of the NTM algorithm for the training job to process the blog post data.
- D Remove the stopwords from the blog post data by using the CountVectorizer function in the scikit-learn library. Replace the blog post data in the S3 bucket with the results of the vectorizer.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh việc một data scientist đang sử dụng Amazon SageMaker Neural Topic Model (NTM) để xây dựng mô hình recommend tags (gợi ý thẻ) từ dữ liệu blog post thô lưu trữ ở Amazon S3 dưới định dạng JSON. Trong quá trình đánh giá mô hình, phát hiện mô hình gợi ý các stopwords phổ biến như "a", "an", "the" làm tags cho một số blog post, kèm theo vài rare words (từ hiếm) chỉ xuất hiện ở một số bài viết cụ thể. Sau khi review với team nội dung, rare words được chấp nhận vì chúng unusual nhưng feasible (bất thường nhưng khả thi). Tuy nhiên, yêu cầu bắt buộc là loại bỏ hoàn toàn stopwords khỏi các tag recommendations của mô hình.
📌 Mục tiêu chính: Preprocessing dữ liệu đầu vào để loại stopwords trước khi train NTM, mà không ảnh hưởng đến rare words. Dữ liệu cần được cập nhật lại ở S3 để SageMaker sử dụng cho training job tiếp theo. Đây là vấn đề text preprocessing điển hình trong SageMaker topic modeling, nơi NTM hoạt động tốt nhất với dữ liệu đã sạch (vectorized và loại bỏ noise như stopwords). Kiến thức cập nhật đến 2026: SageMaker NTM (phần của SageMaker BlazingText/Algorithms) vẫn yêu cầu input là sparse vectors từ text đã tokenized/vectorized (như TF-IDF hoặc CountVectorizer), và preprocessing ngoài SageMaker bằng scikit-learn là best practice (AWS docs khuyến nghị).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Remove the stopwords from the blog post data by using the CountVectorizer function in the scikit-learn library. Replace the blog post data in the S3 bucket with the results of the vectorizer.
🛠️ Lý do chi tiết:
- CountVectorizer trong scikit-learn (sklearn.feature_extraction.text) là công cụ chuẩn để vectorize text data, chuyển JSON text thành sparse matrix (bag-of-words), và có tham số
stop_words='english'để tự động loại bỏ stopwords phổ biến (như "a", "an", "the") mà không ảnh hưởng đến rare words (vì rare words không nằm trong stopword list). - Data scientist chạy script Python với sklearn trên local/EC2/SageMaker Notebook, xử lý dữ liệu S3, rồi upload kết quả vectorized (thường là .npz hoặc .libsvm format) trở lại S3 làm input cho NTM training job.
- Điều này meet requirements chính xác: Loại stopwords, giữ rare words, và model NTM sẽ recommend tags sạch hơn vì input đã preprocess.
- Best practice AWS 2026: SageMaker Processing Job hoặc Notebook hỗ trợ sklearn trực tiếp, hiệu quả cao cho scale lớn.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên nội dung tiếng Anh gốc. Mỗi phương án được đánh giá dựa trên tính phù hợp với yêu cầu (loại stopwords, giữ rare words, tương thích NTM).
-
❌ Phương án SAI: Use the Amazon Comprehend entity recognition API operations. Remove the detected words from the blog post data. Replace the blog post data source in the S3 bucket.
Lý do sai: Amazon Comprehend entity recognition (như DetectEntities API) dùng để detect named entities (person, organization, location), KHÔNG phải stopwords. Nó sẽ bỏ sót stopwords và có thể loại nhầm rare words (nếu chúng giống entity). Không phù hợp preprocessing cho NTM (Comprehend output là JSON entities, không phải vectorized input). Tốn kém và không target stopwords trực tiếp. -
❌ Phương án SAI: Run the SageMaker built-in principal component analysis (PCA) algorithm with the blog post data from the S3 bucket as the data source. Replace the blog post data in the S3 bucket with the results of the training job.
Lý do sai: SageMaker PCA là thuật toán dimensionality reduction (giảm chiều dữ liệu số), KHÔNG xử lý text/stopwords. Input JSON text thô không tương thích trực tiếp (cần vectorize trước), và PCA chỉ transform features chứ không loại từ cụ thể như stopwords. Sẽ làm méo dữ liệu, mất rare words, không giải quyết vấn đề cốt lõi. -
❌ Phương án SAI: Use the SageMaker built-in Object Detection algorithm instead of the NTM algorithm for the training job to process the blog post data.
Lý do sai: SageMaker Object Detection (dựa trên computer vision như SSD/Faster R-CNN) dành cho hình ảnh/video (detect objects/bounding boxes), HOÀN TOÀN KHÔNG phù hợp với dữ liệu text JSON từ blog posts. Thay NTM bằng cái này sẽ fail training vì input không phải image format (RecordIO hoặc image files). Không liên quan đến tags/stopwords. -
✅ Phương án ĐÚNG: Remove the stopwords from the blog post data by using the CountVectorizer function in the scikit-learn library. Replace the blog post data in the S3 bucket with the results of the vectorizer.
Lý do đúng (tóm tắt): Như phần trên, sklearn CountVectorizer loại stopwords chính xác, output vectorized phù hợp NTM input, giữ rare words, và dễ integrate với S3/SageMaker. Hiệu quả, scalable.
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- AWS SageMaker NTM Docs: Amazon SageMaker Neural Topic Model – Nhấn mạnh preprocessing text với sklearn.
- Scikit-learn CountVectorizer: sklearn CountVectorizer – stop_words param.
- SageMaker Processing với sklearn: SageMaker Processing – Best practice cho text prep.
- Exam DOP-C02 Sample: Tương tự các câu về SageMaker preprocessing trong AWS Certified DevOps Engineer Professional.
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần code sample sklearn, hỏi thêm nhé!
The company wants a solution that can transfer and automatically update data between the on-premises object storage and Amazon S3. The solution must support encryption, scheduling, monitoring, and data integrity validation.
Which solution meets these requirements?
- A Use the S3 sync command to compare the source S3 bucket and the destination S3 bucket. Determine which source files do not exist in the destination S3 bucket and which source files were modified.
- B Use AWS Transfer for FTPS to transfer the files from the on-premises storage to Amazon S3.
- C Use AWS DataSync to make an initial copy of the entire dataset. Schedule subsequent incremental transfers of changing data until the final cutover from on premises to AWS.
- D Use S3 Batch Operations to pull data periodically from the on-premises storage. Enable S3 Versioning on the S3 bucket to protect against accidental overwrites.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một công ty muốn xây dựng kho dữ liệu trên AWS Cloud dành cho các dự án machine learning (ML), sử dụng Amazon S3 làm nơi lưu trữ chính. Toàn bộ dữ liệu hiện đang nằm on-premises với kích thước 40TB. Yêu cầu chính của giải pháp là:
- Chuyển dữ liệu ban đầu từ on-premises object storage sang S3.
- Tự động cập nhật dữ liệu thay đổi giữa on-premises và S3 (sync hai chiều hoặc incremental).
- Hỗ trợ đầy đủ: mã hóa dữ liệu (encryption), lập lịch (scheduling), giám sát (monitoring), và kiểm tra tính toàn vẹn dữ liệu (data integrity validation). Giải pháp phải phù hợp với toàn bộ lifecycle ML trên AWS (như SageMaker), xử lý lượng dữ liệu lớn (40TB), và đảm bảo tính liên tục cho migration từ on-premises sang AWS. Đây là tình huống điển hình cho data migration và synchronization lớn. (📘 Tham khảo: AWS DataSync Documentation - cập nhật 2024/2025, AWS Well-Architected Framework for ML).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use AWS DataSync to make an initial copy of the entire dataset. Schedule subsequent incremental transfers of changing data until the final cutover from on premises to AWS.
Lý do chọn đáp án này 🛠️:
- AWS DataSync là dịch vụ chuyên dụng để chuyển dữ liệu nhanh chóng, an toàn từ on-premises (NFS, SMB, HDFS, object storage như S3-compatible) sang AWS (S3, EFS, FSx). Nó hỗ trợ initial full copy (40TB), sau đó incremental sync chỉ dữ liệu thay đổi, phù hợp migration cutover.
- Đầy đủ tính năng yêu cầu: Encryption (TLS, KMS), Scheduling (EventBridge/CloudWatch Events), Monitoring (CloudWatch metrics/logs), Data integrity (MD5/SHA checksums tự động verify).
- Tối ưu cho ML lifecycle: Tích hợp SageMaker, xử lý petabyte-scale, tự động resume nếu gián đoạn. Phiên bản mới nhất (2025) hỗ trợ agentless cho object storage.
- Không có lựa chọn nào khác đáp ứng tất cả (tự động sync từ on-prem object storage). (📘 Nguồn: AWS DataSync User Guide - aws.amazon.com/datasync/, AWS re:Invent 2024 announcements).
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng phương án theo thứ tự (A, B, C, D). Tôi giữ nguyên nội dung văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt với lý do đúng/sai rõ ràng:
-
Use the S3 sync command to compare the source S3 bucket and the destination S3 bucket. Determine which source files do not exist in the destination S3 bucket and which source files were modified.
❌ Sai: Lệnhaws s3 syncchỉ dùng để sync giữa các S3 bucket (hoặc local-S3), không hỗ trợ trực tiếp từ on-premises object storage. Nó yêu cầu chạy thủ công trên EC2/ máy local (không tự động scheduling/monitoring đầy đủ), thiếu encryption at-rest/transit native cho on-prem, và không verify integrity tự động cho 40TB. Không phù hợp migration lớn, dễ lỗi nếu network gián đoạn. (📘 AWS CLI S3 docs: Không đề cập on-prem object sync). -
Use AWS Transfer for FTPS to transfer the files from the on-premises storage to Amazon S3.
❌ Sai: AWS Transfer Family (SFTP/FTPS/FTPS) dành cho file transfer qua protocol từ client/server (như FTP server on-prem), không phải sync tự động từ object storage. Nó không hỗ trợ incremental sync, scheduling native, hay integrity validation checksum. Phù hợp upload thủ công/small-scale, không hiệu quả cho 40TB migration ML (chậm, không cutover tự động). Phiên bản 2025 vẫn giữ nguyên hạn chế này. (📘 AWS Transfer Family docs: aws.amazon.com/aws-transfer-family/). -
Use AWS DataSync to make an initial copy of the entire dataset. Schedule subsequent incremental transfers of changing data until the final cutover from on premises to AWS.
✅ Đúng: Như đã giải thích ở trên. Đây là giải pháp chuẩn AWS best practice cho yêu cầu chính xác: full initial copy → incremental sync → cutover. Hỗ trợ tất cả tính năng (encryption KMS, schedule Lambda/EventBridge, monitor CloudWatch, integrity checksum). Lý tưởng cho 40TB từ on-prem object storage (POSIX/S3-compatible). (📘 AWS DataSync FAQs & Best Practices: docs.aws.amazon.com/datasync/latest/userguide/). -
Use S3 Batch Operations to pull data periodically from the on-premises storage. Enable S3 Versioning on the S3 bucket to protect against accidental overwrites.
❌ Sai: S3 Batch Operations chỉ thực hiện jobs trên objects đã có trong S3 (copy/tag/restore), không thể "pull" trực tiếp từ on-premises. Không hỗ trợ scheduling từ on-prem, encryption/sync tự động, hay integrity cho external source. S3 Versioning chỉ bảo vệ overwrite trong S3, không giải quyết migration 40TB. (📘 S3 Batch Operations docs: docs.aws.amazon.com/AmazonS3/latest/userguide/batch-ops.html).
Kết luận 🚀: AWS DataSync là lựa chọn tối ưu, giúp công ty nhanh chóng migrate dữ liệu ML sang S3 mà không downtime. Nếu triển khai thực tế, khuyến nghị test với DataSync agent trên on-prem và monitor qua CloudWatch. (📘 Thêm tài liệu: AWS ML Data Management whitepaper 2025).
A data scientist creates a bounding box to label the sample data and uses an object detection model. However, the object detection model cannot clearly demarcate the yellow line, the passengers who cross the yellow line, and the trains.
Which labeling approach will help the company improve this model?
- A Use Amazon Rekognition Custom Labels to label the dataset and create a custom Amazon Rekognition object detection model. Create a private workforce. Use Amazon Augmented AI (Amazon A2I) to review the low-confidence predictions and retrain the custom Amazon Rekognition model.
- B Use an Amazon SageMaker Ground Truth object detection labeling task. Use Amazon Mechanical Turk as the labeling workforce.
- C Use Amazon Rekognition Custom Labels to label the dataset and create a custom Amazon Rekognition object detection model. Create a workforce with a third-party AWS Marketplace vendor. Use Amazon Augmented AI (Amazon A2I) to review the low-confidence predictions and retrain the custom Amazon Rekognition model.
- D Use an Amazon SageMaker Ground Truth semantic segmentation labeling task. Use a private workforce as the labeling workforce.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh một công ty sở hữu video feeds và hình ảnh từ ga tàu điện ngầm, muốn xây dựng mô hình deep learning để cảnh báo quản lý ga khi hành khách vượt vạch vàng an toàn (yellow safety line) mà không có tàu đang đỗ. Mô hình cần phát hiện chính xác:
- Vạch vàng (yellow line),
- Hành khách vượt vạch (passengers crossing the line),
- Tàu điện (trains).
Dữ liệu video phải giữ bí mật (confidential). Data scientist đã thử bounding box cho object detection model, nhưng model không phân biệt rõ ràng các yếu tố trên (vì bounding box chỉ khoanh vùng object thô, không chi tiết đến mức pixel-level cho đường line mỏng hoặc hành vi vượt vạch).
Vấn đề cốt lõi: Cần phương pháp labeling tốt hơn để cải thiện model, đặc biệt cho đối tượng phức tạp như đường line, hành vi vượt line, và phân biệt tàu. Đây là task yêu cầu labeling pixel-precise thay vì bounding box đơn giản.
📘 Tài liệu tham khảo:
- AWS SageMaker Ground Truth Documentation (cập nhật 2024-2026): Hỗ trợ semantic segmentation cho pixel-level labeling trên images/video [docs.aws.amazon.com/sagemaker/latest/dg/sms.html].
- Amazon Rekognition Custom Labels: Giới hạn ở object detection với bounding box [docs.aws.amazon.com/rekognition/latest/customlabels-dg/what-is-custom-labels.html].
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use an Amazon SageMaker Ground Truth semantic segmentation labeling task. Use a private workforce as the labeling workforce.
Lý do:
- Semantic segmentation trong SageMaker Ground Truth cho phép labeling từng pixel (pixel-level), rất phù hợp để demarcate rõ ràng vạch vàng mỏng, hành khách vượt vạch (boundary chính xác), và tàu. Bounding box (object detection) thất bại vì chỉ khoanh vùng thô, không xử lý tốt line/edge detection.
- Private workforce đảm bảo dữ liệu confidential (không chia sẻ với bên ngoài như Mechanical Turk hay third-party).
- Hỗ trợ video feeds qua frame labeling, cải thiện model hiệu quả cho real-time alert.
- 🛠️ Cập nhật 2026: SageMaker Ground Truth nay tích hợp AI-assisted labeling cho semantic segmentation, giảm công sức thủ công.
❌ Phân tích tất cả các phương án (đúng/sai)
-
Use Amazon Rekognition Custom Labels to label the dataset and create a custom Amazon Rekognition object detection model. Create a private workforce. Use Amazon Augmented AI (Amazon A2I) to review the low-confidence predictions and retrain the custom Amazon Rekognition model.
❌ Sai vì: Rekognition Custom Labels chỉ hỗ trợ object detection với bounding box, không xử lý semantic segmentation (pixel-level). Vẫn gặp vấn đề "không demarcate rõ" vạch vàng/hành khách vượt. Private workforce tốt cho confidential, A2I hữu ích review, nhưng core labeling không phù hợp. -
Use an Amazon SageMaker Ground Truth object detection labeling task. Use Amazon Mechanical Turk as the labeling workforce.
❌ Sai vì: Ground Truth object detection vẫn dùng bounding box, không cải thiện vấn đề demarcate chính xác. Mechanical Turk là public workforce, vi phạm yêu cầu confidential data (dữ liệu video có thể bị lộ). -
Use Amazon Rekognition Custom Labels to label the dataset and create a custom Amazon Rekognition object detection model. Create a workforce with a third-party AWS Marketplace vendor. Use Amazon Augmented AI (Amazon A2I) to review the low-confidence predictions and retrain the custom Amazon Rekognition model.
❌ Sai vì: Tương tự phương án đầu, Rekognition Custom Labels giới hạn bounding box, không giải quyết semantic cần thiết. Third-party vendor (AWS Marketplace) không đảm bảo confidential (dữ liệu chia sẻ ngoài), dù A2I tốt cho review. -
Use an Amazon SageMaker Ground Truth semantic segmentation labeling task. Use a private workforce as the labeling workforce.
✅ Đúng vì: Như giải thích trên, semantic segmentation là giải pháp lý tưởng cho pixel-precise labeling (vạch vàng, vượt line, tàu). Private workforce giữ bí mật dữ liệu. Hoàn hảo cho video/images confidential.
🛠️ Lời khuyên thực hành: Sau labeling, train custom model trên SageMaker với dataset này, deploy endpoint cho real-time inference trên video feeds (sử dụng Lambda + Kinesis Video Streams cho alert). Kiểm tra best practices AWS Well-Architected ML Lens (2026 update).
Which steps should the data engineer take to address this issue? (Choose two.)
- A Use a linear-based algorithm to train the model.
- B Apply principal component analysis (PCA).
- C Remove a portion of highly correlated features from the dataset.
- D Apply min-max feature scaling to the dataset.
- E Apply one-hot encoding category-based variables.
Xem giải thích
🛡️ Phân tích câu hỏi trắc nghiệm AWS bởi AWS Certified DevOps Engineer Professional
(Kiến thức dựa trên phiên bản AWS mới nhất đến năm 2026, bao gồm SageMaker v3.x và các best practices ML pipelines trong DevOps trên AWS như SageMaker Pipelines, Processing Jobs, và Feature Store.)
🧩 Giải thích nội dung câu hỏi một cách chi tiết
Câu hỏi mô tả tình huống một data engineer tại ngân hàng đang đánh giá một dataset tabular chứa dữ liệu khách hàng, dùng để xây dựng mô hình dự đoán hành vi khách hàng. Sau khi tạo correlation matrix cho 100 features, phát hiện nhiều features có tương quan cao (highly correlated) với nhau.
Vấn đề cốt lõi: Multicollinearity (tương quan tuyến tính giữa các features) có thể dẫn đến:
- Mô hình không ổn định (unstable coefficients).
- Overfitting, giảm khả năng generalize.
- Tăng thời gian training và chi phí compute (đặc biệt trên AWS SageMaker).
Yêu cầu: Chọn TWO steps để giải quyết vấn đề này trong quy trình data preparation cho ML pipeline trên AWS (ví dụ: SageMaker Processing Job hoặc SageMaker Canvas).
Mục tiêu là giảm chiều dữ liệu hoặc loại bỏ redundancy mà không mất thông tin quan trọng, phù hợp với DevOps practices như automation trong SageMaker Pipelines.
✅ Đáp án đúng và lý do lựa chọn
Hai đáp án đúng là:
- Apply principal component analysis (PCA): PCA là kỹ thuật dimensionality reduction hàng đầu để xử lý multicollinearity bằng cách chuyển đổi features tương quan thành các principal components orthogonal (không tương quan), giữ nguyên variance lớn nhất. Trong AWS SageMaker, PCA được hỗ trợ native qua SageMaker PCA algorithm hoặc Scikit-learn Processing Job.
- Remove a portion of highly correlated features from the dataset: Loại bỏ một phần features có correlation cao (threshold >0.8-0.9) dựa trên correlation matrix là cách đơn giản, hiệu quả để giảm redundancy, tránh multicollinearity mà không cần transform phức tạp. Phù hợp với Feature Store hoặc Processing Jobs trên AWS.
Lý do chọn hai cái này: Chúng trực tiếp giải quyết multicollinearity bằng cách giảm số lượng features hoặc decorrelate chúng, cải thiện model performance và giảm chi phí (ví dụ: ít instances EC2/ML instances hơn trong training). Các lựa chọn khác không target vấn đề correlation.
📋 Phân tích chi tiết tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn một cách rõ ràng, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên best practices ML trên AWS (SageMaker, Glue ETL).
-
❌ Use a linear-based algorithm to train the model.
Sai: Sử dụng thuật toán tuyến tính (như Linear Regression/Logistic Regression) không giải quyết multicollinearity; ngược lại, nó làm vấn đề tệ hơn vì các mô hình linear nhạy cảm với correlation cao (dẫn đến unstable estimates). Trong SageMaker, bạn vẫn cần preprocess data trước khi train Linear Learner. Không phải bước "address this issue". -
✅ Apply principal component analysis (PCA).
Đúng: PCA biến đổi features tương quan thành các components độc lập, giảm chiều dữ liệu hiệu quả. AWS hỗ trợ trực tiếp qua SageMaker Built-in Algorithms (sagemaker.sklearn.PCA) hoặc Processing Jobs với scikit-learn. Giảm multicollinearity, giữ 95% variance → lý tưởng cho dataset 100 features. -
✅ Remove a portion of highly correlated features from the dataset.
Đúng: Dựa trên correlation matrix (threshold >0.7-0.9), loại bỏ features redundant (ví dụ: giữ feature có correlation cao nhất với target). Cách thủ công/effective trong SageMaker Data Wrangler hoặc Pandas Processing Job, giảm overfitting và training time. -
❌ Apply min-max feature scaling to the dataset.
Sai: Min-max scaling (chuyển features về [0,1]) chỉ chuẩn hóa scale, không ảnh hưởng đến correlation giữa features (correlation là scale-invariant). Trong SageMaker, dùng cho algorithms distance-based (KNN), nhưng không giải quyết multicollinearity. -
❌ Apply one-hot encoding category-based variables.
Sai: One-hot encoding dùng cho categorical variables (chuyển thành binary vectors), không liên quan đến correlation giữa numerical features. Có thể làm tăng multicollinearity nếu categoricals có correlation cao (dummy variable trap). Trong SageMaker, dùng cho preprocessing, nhưng không target vấn đề ở đây.
📘 Tài liệu tham khảo (AWS Official Sources - Cập nhật 2026)
- SageMaker PCA: docs.aws.amazon.com/sagemaker/latest/dg/pca.html – Hướng dẫn apply PCA trong Processing Jobs.
- Handling Multicollinearity in SageMaker: AWS ML Best Practices – Phần Feature Engineering & Dimensionality Reduction.
- SageMaker Data Wrangler: docs.aws.amazon.com/sagemaker/latest/dg/data-wrangler.html – Công cụ visual để remove correlated features.
- Exam Topic (AWS ML Specialty/DevOps Pro): DOP-C02/ML-Specialty blueprint về ML pipelines và data prep.
🛠️ Khuyến nghị DevOps: Tích hợp các bước này vào SageMaker Pipelines với Step Functions cho CI/CD tự động, sử dụng Lambda để compute correlation matrix và trigger removal/PCA! Nếu cần code sample, hãy hỏi thêm. 🚀