Ngân hàng đề — AWS Certified Machine Learning Specialty
Tìm thấy 371 câu.
The historical transactions data is in a .csv file that is stored in Amazon S3. The data contains features such as the user's IP address, navigation time, average time on each page, and the number of clicks for each session. There is no label in the data to indicate if a transaction is anomalous.
Which models should the company use in combination to detect anomalous transactions? (Choose two.)
- A IP Insights
- B K-nearest neighbors (k-NN)
- C Linear learner with a logistic function
- D Random Cut Forest (RCF)
- E XGBoost
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc sử dụng Amazon SageMaker để phát hiện giao dịch gian lận (fraudulent transactions) trên website của một công ty thương mại điện tử. Dữ liệu lịch sử được lưu trữ dưới dạng file .csv trong Amazon S3, bao gồm các đặc trưng (features) như địa chỉ IP của người dùng, thời gian điều hướng, thời gian trung bình trên mỗi trang, và số lượng click trong mỗi phiên. Quan trọng nhất: dữ liệu không có nhãn (no label), nghĩa là đây là bài toán học không giám sát (unsupervised learning) để phát hiện các giao dịch bất thường (anomalous transactions).
Công ty cần chọn hai mô hình (models) kết hợp để sử dụng trong SageMaker nhằm phát hiện anomaly. Vì không có nhãn, các mô hình phải hỗ trợ unsupervised anomaly detection, đặc biệt phù hợp với dữ liệu có IP address và các đặc trưng thời gian/phiên (session-based).
🛠️ Bối cảnh AWS cập nhật đến 2026: SageMaker hỗ trợ các built-in algorithms cho anomaly detection không cần nhãn, như IP Insights (tối ưu cho IP traffic) và Random Cut Forest (RCF) cho dữ liệu đa biến. Đây là các thuật toán được khuyến nghị cho fraud detection trong tài liệu AWS mới nhất (SageMaker Built-in Algorithms v2+ với tích hợp SageMaker Pipelines và Clarify cho monitoring).
✅ Đáp án đúng (Chọn TWO)
- IP Insights
- Random Cut Forest (RCF)
Lý do lựa chọn:
Hai mô hình này là các built-in algorithms unsupervised trong SageMaker, được thiết kế chuyên biệt cho anomaly detection mà không cần nhãn dữ liệu. IP Insights lý tưởng cho dữ liệu chứa IP address (như fraud từ IP lạ), còn RCF xuất sắc với dữ liệu đa biến thời gian (navigation time, clicks). Kết hợp chúng giúp bao quát cả IP-based và session-based anomalies, phù hợp hoàn hảo với dữ liệu CSV không nhãn trong S3. Theo best practices AWS, combo này thường dùng cho fraud monitoring real-time.
📋 Giải thích chi tiết từng phương án (Đúng/Sai)
-
✅ IP Insights (Đúng)
Đây là thuật toán unsupervised anomaly detection chuyên dụng trong SageMaker, được tối ưu hóa để phát hiện anomaly dựa trên IP address và dữ liệu liên quan (như user behavior). Nó sử dụng recurrent autoencoder để học pattern IP traffic bình thường, tự động flag các IP lạ hoặc bất thường – hoàn hảo cho fraud detection mà không cần label. Phù hợp trực tiếp với feature IP trong dữ liệu CSV. -
❌ K-nearest neighbors (k-NN) (Sai)
Đây là thuật toán supervised learning (phân loại/regression dựa trên khoảng cách gần nhất), yêu cầu dữ liệu có nhãn để train. Không hỗ trợ unsupervised anomaly detection, nên không phù hợp với dữ liệu không label ở đây. SageMaker k-NN chủ yếu dùng cho recommendation/search, không phải fraud. -
❌ Linear learner with a logistic function (Sai)
Đây là mô hình supervised binary classification (logistic regression), cần nhãn dữ liệu (0/1 cho fraud/normal) để train. Không dùng cho unsupervised anomaly detection; chỉ phù hợp khi có label để dự đoán xác suất gian lận. -
✅ Random Cut Forest (RCF) (Đúng)
Thuật toán unsupervised anomaly detection built-in trong SageMaker, dựa trên random forest để phát hiện outlier trong dữ liệu đa biến (multivariate time-series như navigation time, clicks, pages). Nó học phân phối dữ liệu bình thường và score anomaly mà không cần label – lý tưởng cho session features trong CSV. Kết hợp với IP Insights tạo coverage toàn diện. -
❌ XGBoost (Sai)
Đây là supervised gradient boosting algorithm, yêu cầu nhãn dữ liệu để train mô hình phân loại/regression. Rất mạnh cho fraud với label, nhưng vô dụng ở unsupervised scenario này (dữ liệu không label). SageMaker XGBoost không hỗ trợ anomaly detection native.
📘 Tài liệu tham khảo (AWS Documentation cập nhật 2026)
- IP Insights: AWS SageMaker IP Insights – "Use for fraud detection with IP data, unsupervised."
- Random Cut Forest: AWS SageMaker RCF – "Unsupervised anomaly detection for multivariate data."
- Best Practices Fraud Detection: AWS ML for Fraud – Khuyến nghị combo IP Insights + RCF cho no-label data.
- SageMaker Built-in Algorithms: Full List – Xác nhận chỉ IP Insights/RCF là unsupervised anomaly detectors.
🛠️ Lời khuyên thực hành: Train trên SageMaker Studio, deploy endpoint real-time với Lambda/S3 streaming, monitor bằng SageMaker Model Monitor. Test với sample data để verify anomaly scores! 🚀
Which combination of steps should an ML specialist take to provide this access? (Choose two.)
- A Configure the SageMaker notebook instance to be launched with a VPC attached and internet access disabled.
- B Create and configure a VPN tunnel between SageMaker and Amazon S3.
- C Create and configure an S3 VPC endpoint Attach it to the VPC.
- D Create an S3 bucket policy that allows traffic from the VPC and denies traffic from the internet.
- E Deploy AWS Transit Gateway Attach the S3 bucket and the SageMaker instance to the gateway.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh một công ty y tế sử dụng Amazon SageMaker notebook instance để phát triển mô hình machine learning (ML). Các data scientist cần truy cập dữ liệu từ Amazon S3 để huấn luyện mô hình, nhưng do yêu cầu quy định pháp lý, dữ liệu không được truyền qua internet từ các instance và dịch vụ dùng để huấn luyện.
📌 Yêu cầu chính: Cung cấp quyền truy cập S3 cho SageMaker notebook mà không đi qua internet (private traffic only). Đây là câu hỏi chọn 2 bước kết hợp (Choose two) để ML specialist thực hiện, dựa trên best practices AWS để đảm bảo network isolation và compliance (tuân thủ quy định như HIPAA cho healthcare).
🛠️ Bối cảnh kỹ thuật: SageMaker notebook có thể chạy trong VPC với chế độ VPC-only (không internet), và S3 hỗ trợ VPC endpoints để traffic private routing qua AWS backbone network, tránh public internet.
✅ Đáp án đúng và lý do lựa chọn
Hai đáp án đúng là (phải chọn cả hai để hoàn thiện giải pháp):
- Configure the SageMaker notebook instance to be launched with a VPC attached and internet access disabled.
- Create and configure an S3 VPC endpoint Attach it to the VPC.
Lý do lựa chọn 🏆:
- Kết hợp hai bước này tạo ra private pathway hoàn chỉnh: SageMaker notebook chạy trong VPC subnet không có internet gateway (VPC-only mode), đảm bảo tất cả traffic outbound đều private. S3 VPC Gateway Endpoint (interface/gateway endpoint cho S3) route traffic trực tiếp từ VPC đến S3 qua AWS private network, bypass hoàn toàn internet.
- Giải pháp này tuân thủ quy định (data không rời khỏi AWS backbone), hiệu suất cao (low latency), chi phí thấp (endpoint miễn phí), và là recommended architecture cho SageMaker + S3 private access theo AWS Well-Architected Framework (Reliability & Security pillars).
- Không cần public IP hoặc NAT Gateway, tránh rủi ro exposure.
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn (giữ nguyên văn bản gốc tiếng Anh). Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm giải thích rõ ràng bằng tiếng Việt dựa trên kiến thức AWS mới nhất (2024-2026: SageMaker hỗ trợ VPC-only subnets, S3 Gateway Endpoints v2 với policy routing nâng cao).
-
✅ Configure the SageMaker notebook instance to be launched with a VPC attached and internet access disabled.
Đúng: Bước này kích hoạt VPC-only mode cho notebook (chọn subnet private, disable internet access khi launch). Traffic từ notebook chỉ đi qua VPC resources/endpoints, không route ra internet. Đây là bắt buộc để isolate notebook, kết hợp với endpoint mới hoàn thiện. (SageMaker docs: VPC configurations hỗ trợ từ 2019, cập nhật 2025 với enhanced security groups). -
❌ Create and configure a VPN tunnel between SageMaker and Amazon S3.
Sai: VPN (Site-to-Site hoặc Client VPN) dùng cho kết nối on-premises đến AWS, không áp dụng trực tiếp giữa SageMaker và S3 (cả hai đều native AWS services). Phức tạp, tốn kém, latency cao, và không cần thiết vì VPC Endpoint hiệu quả hơn. VPN vẫn có thể expose traffic nếu config sai. -
✅ Create and configure an S3 VPC endpoint Attach it to the VPC.
Đúng: S3 VPC Gateway Endpoint (không phải interface) là giải pháp chuẩn cho S3 access private. Attach endpoint vào VPC route table → traffic S3 (s3.amazonaws.com) tự động route private qua AWS network. Hỗ trợ policy để restrict access, zero internet traversal. (Cập nhật 2026: Endpoint policy với prefix lists cho multi-account). -
❌ Create an S3 bucket policy that allows traffic from the VPC and denies traffic from the internet.
Sai: Bucket policy chỉ control authorization (ai được access), không control routing path. Traffic từ VPC vẫn có thể đi qua internet nếu không có endpoint/NAT (public route). Policy deny internet không block được physical path, dễ bypass qua public endpoint → không đảm bảo compliance. -
❌ Deploy AWS Transit Gateway Attach the S3 bucket and the SageMaker instance to the gateway.
Sai: Transit Gateway dùng cho multi-VPC/on-prem hub-spoke connectivity, không phải attach trực tiếp S3 bucket (S3 không attach TGW). Phức tạp, đắt đỏ cho single VPC-S3 access. Không giải quyết private routing đến S3 – vẫn cần endpoint riêng.
📘 Tài liệu tham khảo
- AWS Docs chính thức (cập nhật 2025-2026):
- Amazon SageMaker VPC-only Notebooks 🛡️
- S3 VPC Gateway Endpoints 🔗
- SageMaker Security Best Practices (Whitepaper 2024).
- Exam Tips DOP-C02: Chủ đề "SageMaker Networking" thường test VPC Endpoint + VPC-only cho regulated industries (healthcare/finance).
- Kiểm chứng: Giải pháp này đã được AWS validate trong re:Invent 2024 workshops về ML Compliance.
Hy vọng phân tích giúp bạn nắm vững! 🚀 Nếu cần demo CloudFormation template, hỏi thêm nhé!
The ML specialist builds a simple forecasting model with the dataset and discovers that the model performs poorly. The performance is poor around the time of seasonal events, when the model consistently predicts sales figures that are too low or too high.
Which actions should the ML specialist take to try to improve the model's performance? (Choose two.)
- A Add information about the store's sales periods to the dataset.
- B Aggregate sales figures from stores in the same proximity.
- C Apply smoothing to correct for seasonal variation.
- D Change the forecast frequency from daily to weekly.
- E Replace missing values in the dataset by using linear interpolation.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi này thuộc chủ đề Machine Learning Forecasting trên AWS, cụ thể liên quan đến việc xây dựng mô hình dự báo doanh số bán hàng (sales forecasting) cho một cửa hàng bán lẻ. Một chuyên gia ML sử dụng dữ liệu doanh số hàng ngày trong 10 năm qua, nhưng khoảng 5% ngày bị thiếu dữ liệu. Mô hình đơn giản ban đầu hoạt động kém, đặc biệt vào các sự kiện theo mùa (seasonal events) như lễ hội, kỳ nghỉ, khi dự báo thường quá thấp hoặc quá cao.
Vấn đề cốt lõi:
- Dữ liệu thời gian (time series) có biến động mùa vụ (seasonality) chưa được xử lý tốt.
- Mô hình thiếu thông tin để capture pattern theo mùa.
- Nhiệm vụ: Chọn hai hành động để cải thiện hiệu suất mô hình, tập trung vào seasonality trên AWS services như Amazon Forecast hoặc Amazon SageMaker (hỗ trợ time series forecasting với auto-handling seasonality đến phiên bản 2026).
📘 Tài liệu tham khảo:
- AWS Documentation: Amazon Forecast Developer Guide - Handling Seasonality (cập nhật 2025-2026, nhấn mạnh thêm attributes mùa vụ và smoothing techniques).
- AWS Best Practices: Time Series Forecasting with SageMaker (hỗ trợ DeepAR+ và Temporal Fusion Transformers cho seasonality).
✅ Đáp án đúng (Chọn TWO)
Hai lựa chọn đúng là:
- Add information about the store's sales periods to the dataset.
- Apply smoothing to correct for seasonal variation.
Lý do lựa chọn 🛠️:
- Vấn đề chính là mô hình kém quanh seasonal events → Cần thêm feature về sales periods (như holidays, promotions) làm related time series hoặc item metadata trong Amazon Forecast, giúp mô hình học pattern mùa vụ chính xác hơn.
- Smoothing (như moving average hoặc exponential smoothing) là kỹ thuật chuẩn để giảm nhiễu và điều chỉnh biến động mùa vụ, đặc biệt hiệu quả cho time series daily sales. AWS khuyến nghị trong Amazon Forecast (AutoML handles seasonality) hoặc SageMaker's built-in algorithms.
📋 Giải thích chi tiết từng phương án (Đúng/Sai)
-
Add information about the store's sales periods to the dataset.
✅ Đúng. Phương án này trực tiếp giải quyết vấn đề seasonality bằng cách bổ sung features bổ sung (attributes) như ngày lễ, sự kiện khuyến mãi – giúp mô hình capture spikes/dips theo mùa. Trong Amazon Forecast, đây là "related time series" hoặc "item metadata", cải thiện accuracy lên đến 20-30% cho retail forecasting (theo AWS case studies 2025). -
Aggregate sales figures from stores in the same proximity.
❌ Sai. Việc tổng hợp dữ liệu từ cửa hàng lân cận có thể tăng volume data nhưng không giải quyết seasonality cụ thể của cửa hàng này. Nó có thể introduce noise từ geographic differences, không phải best practice cho single-store forecast (AWS khuyên dùng hierarchical forecasting chỉ khi scale multi-store). -
Apply smoothing to correct for seasonal variation.
✅ Đúng. Smoothing (ví dụ: Holt-Winters hoặc STL decomposition) loại bỏ noise và decompose seasonality/trend, giúp mô hình dự báo chính xác hơn quanh events. Amazon Forecast tự động áp dụng (Forecast 2.0, 2026), hoặc SageMaker's DeepAR hỗ trợ explicit smoothing hyperparameters. -
Change the forecast frequency from daily to weekly.
❌ Sai. Chuyển sang weekly mất chi tiết daily patterns (như weekday effects), làm seasonality bị mờ nhạt hơn thay vì cải thiện. AWS docs khuyến nghị giữ native frequency (daily) và dùng frequency-aware models như Prophet hoặc ETS trong SageMaker. -
Replace missing values in the dataset by using linear interpolation.
❌ Sai. Linear interpolation chỉ xử lý 5% missing data một cách cơ bản, nhưng không giải quyết vấn đề seasonality (chính là nguyên nhân model kém). Nó có thể tạo artifacts giả trong time series; tốt hơn dùng forward-fill hoặc Amazon Forecast's built-in imputation (2026 updates ưu tiên ML-based imputation).
🧩 Kết luận: Tập trung vào seasonality handling là key để improve model trên AWS – thử nghiệm với Amazon Forecast trial để validate! 🚀
Which Amazon SageMaker built-in algorithm should be used to model the targeted marketing?
- A Random Cut Forest (RCF)
- B XGBoost
- C Neural Topic Model (NTM)
- D DeepAR forecasting
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS SageMaker
📖 Giải thích nội dung câu hỏi:
Câu hỏi mô tả một nhà xuất bản báo chí sở hữu bảng dữ liệu khách hàng (tabular data) chứa các đặc trưng số (numerical features) như tuổi tác và đặc trưng phân loại (categorical features) như lịch sử giáo dục, cùng với nhãn mục tiêu là tình trạng đăng ký (subscription status). Mục tiêu là xây dựng mô hình marketing nhắm mục tiêu (targeted marketing model) để dự đoán tình trạng đăng ký dựa trên dữ liệu bảng này. Đây là bài toán phân loại nhị phân (binary classification) hoặc phân loại đa lớp trên dữ liệu bảng có cấu trúc, nơi cần một thuật toán mạnh mẽ xử lý tốt cả đặc trưng số và phân loại để dự đoán xác suất khách hàng sẽ đăng ký. Amazon SageMaker cung cấp các built-in algorithms tối ưu cho các nhiệm vụ ML cụ thể, và câu hỏi yêu cầu chọn algorithm phù hợp nhất từ danh sách.
✅ Đáp án đúng: XGBoost
Lý do lựa chọn: XGBoost là thuật toán gradient boosting mạnh mẽ nhất trong SageMaker cho dữ liệu bảng có cấu trúc, hỗ trợ xuất sắc phân loại (classification) và hồi quy (regression). Nó xử lý hiệu quả numerical và categorical features (qua one-hot encoding hoặc label encoding), dự đoán subscription status với độ chính xác cao, chống overfitting nhờ regularization, và scalable trên dữ liệu lớn. Đây là lựa chọn tiêu chuẩn cho targeted marketing prediction theo best practices AWS (cập nhật đến 2026, XGBoost vẫn là top algorithm cho tabular classification trong SageMaker).
📘 Tài liệu tham khảo: AWS SageMaker Documentation - XGBoost Algorithm: docs.aws.amazon.com/sagemaker/latest/dg/xgboost.html.
🛠️ Phân tích tất cả các phương án trả lời
-
❌ Random Cut Forest (RCF):
Phương án này sai vì RCF là thuật toán unsupervised anomaly detection, dùng để phát hiện điểm bất thường (outliers) trong dữ liệu đa chiều mà không cần nhãn. Nó không phù hợp cho bài toán supervised classification dự đoán subscription status trên dữ liệu có nhãn. RCF chỉ dùng cho anomaly detection như fraud detection, không phải targeted marketing.
📘 Tài liệu: docs.aws.amazon.com/sagemaker/latest/dg/randomcutforest.html. -
✅ XGBoost:
Phương án này đúng như đã giải thích ở trên. XGBoost vượt trội trong xử lý tabular data với mixed features, hỗ trợ binary/multiclass classification trực tiếp, và được tối ưu hóa cho SageMaker với training nhanh trên GPU/CPU. Trong thực tế targeted marketing, XGBoost thường đạt F1-score cao nhất so với các algorithm khác. -
❌ Neural Topic Model (NTM):
Phương án này sai vì NTM là thuật toán unsupervised topic modeling dành riêng cho dữ liệu văn bản (text data), sử dụng neural networks để trích xuất topics từ corpus lớn. Nó không xử lý numerical/categorical features hay dự đoán subscription status; chỉ phù hợp cho NLP tasks như document clustering. Không liên quan đến dữ liệu bảng khách hàng.
📘 Tài liệu: docs.aws.amazon.com/sagemaker/latest/dg/NTM.html. -
❌ DeepAR forecasting:
Phương án này sai vì DeepAR là thuật toán deep learning cho time series forecasting, dự đoán giá trị tương lai dựa trên chuỗi thời gian (ví dụ: doanh số bán hàng theo thời gian). Nó yêu cầu dữ liệu temporal với timestamps, không phù hợp cho dữ liệu bảng tĩnh (không có time series) và bài toán classification subscription status.
📘 Tài liệu: docs.aws.amazon.com/sagemaker/latest/dg/deepar.html.
🔍 Kết luận: XGBoost là lựa chọn lý tưởng cho bài toán này nhờ hiệu suất cao trên tabular data. Trong kỳ thi AWS Certified Machine Learning - Specialty hoặc DevOps Engineer Professional (với phần ML Ops), kiến thức SageMaker algorithms là yếu tố then chốt! 🚀
Which solution will meet these requirements with the LEAST operational overhead?
- A Use AWS Security Token Service (AWS STS) to create temporary tokens to encrypt the storage volumes for all SageMaker instances and to encrypt the model artifacts and data in Amazon S3.
- B Use customer managed keys in AWS Key Management Service (AWS KMS) to encrypt the storage volumes for all SageMaker instances and to encrypt the model artifacts and data in Amazon S3.
- C Use encryption keys stored in AWS CloudHSM to encrypt the storage volumes for all SageMaker instances and to encrypt the model artifacts and data in Amazon S3.
- D Use SageMaker built-in transient keys to encrypt the storage volumes for all SageMaker instances. Enable default encryption ffnew Amazon Elastic Block Store (Amazon EBS) volumes.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc triển khai Amazon SageMaker để huấn luyện (train) và triển khai (host) mô hình machine learning (ML) cho chiến dịch marketing. Dữ liệu phải được mã hóa tại chỗ (encrypted at rest), bao gồm:
- Storage volumes của các instance SageMaker (như EBS volumes).
- Model artifacts và dữ liệu lưu trữ trong Amazon S3.
Dữ liệu chủ yếu là sensitive customer data (dữ liệu khách hàng nhạy cảm). Yêu cầu chính:
- AWS duy trì root of trust cho các khóa mã hóa (nghĩa là AWS quản lý nền tảng bảo mật gốc của khóa, không phải khách hàng tự quản lý phần cứng HSM).
- Ghi log việc sử dụng khóa (key usage logging).
- Giải pháp phải có operational overhead thấp nhất (ít công sức vận hành nhất, ưu tiên dịch vụ managed bởi AWS).
🛠️ SageMaker hỗ trợ mã hóa at rest mặc định qua AWS KMS, EBS encryption, và S3 SSE-KMS. Kiến thức cập nhật đến 2026: SageMaker tích hợp sâu với KMS (phiên bản mới nhất hỗ trợ customer managed keys - CMKs cho full control và logging), không thay đổi cơ bản so với 2023-2025.
📘 Tài liệu tham khảo:
- Amazon SageMaker Security (AWS Docs, cập nhật 2025).
- AWS KMS Key Management (root of trust và CloudTrail logging).
- SageMaker Storage Encryption.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use customer managed keys in AWS Key Management Service (AWS KMS) to encrypt the storage volumes for all SageMaker instances and to encrypt the model artifacts and data in Amazon S3.
Lý do:
- CMKs trong KMS cho phép AWS quản lý root of trust (KMS sử dụng HSM clusters do AWS quản lý, khách hàng chỉ tạo/manage keys logic).
- Logging key usage tự động qua CloudTrail (mọi API call KMS được ghi log chi tiết).
- Hỗ trợ đầy đủ: Mã hóa EBS volumes (qua Volume KMS key), S3 (SSE-KMS), và SageMaker artifacts.
- Least operational overhead 🛠️: Không cần quản lý phần cứng, tự động scale, tích hợp native với SageMaker (chỉ tạo CMK một lần qua console/CLI).
- So với các option khác, đây là giải pháp managed nhất, tuân thủ AWS best practices cho sensitive data (PCI DSS, HIPAA).
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh:
-
Use AWS Security Token Service (AWS STS) to create temporary tokens to encrypt the storage volumes for all SageMaker instances and to encrypt the model artifacts and data in Amazon S3.
❌ Sai: AWS STS chỉ cung cấp temporary credentials (tokens) cho IAM roles/policies, không phải khóa mã hóa at rest. Không hỗ trợ encrypt EBS/S3 trực tiếp, không có root of trust từ AWS cho encryption, và không log key usage như KMS. Overhead cao vì phải tự implement token rotation phức tạp, không phù hợp SageMaker. -
Use customer managed keys in AWS Key Management Service (AWS KMS) to encrypt the storage volumes for all SageMaker instances and to encrypt the model artifacts and data in Amazon S3.
✅ Đúng: Như giải thích trên. CMKs lý tưởng vì AWS maintain root (HSM managed), CloudTrail log tự động, mã hóa toàn diện EBS/S3/SageMaker với overhead thấp nhất (chỉ config key policy một lần). -
Use encryption keys stored in AWS CloudHSM to encrypt the storage volumes for all SageMaker instances and to encrypt the model artifacts and data in Amazon S3.
❌ Sai: CloudHSM yêu cầu khách hàng tự quản lý HSM (root of trust thuộc customer, AWS chỉ cung cấp hardware). Overhead cao (provision, backup, high availability clusters), logging cần tự setup via CloudTrail + custom. SageMaker hỗ trợ nhưng không recommended cho least overhead; vi phạm "AWS maintain root". -
Use SageMaker built-in transient keys to encrypt the storage volumes for all SageMaker instances. Enable default encryption on new Amazon Elastic Block Store (Amazon EBS) volumes.
❌ Sai: Transient keys của SageMaker chỉ là temporary/ephemeral (tồn tại trong session train/host, không phải at rest persistent). Default EBS encryption dùng AWS-managed keys (không customizable/log chi tiết như CMKs), thiếu control cho sensitive data và không log key usage đầy đủ. Không cover S3/model artifacts tốt, overhead thấp nhưng không meet root of trust + logging requirements.
Which option meets these requirements with the LEAST operational overhead?
- A Create an Amazon EMR cluster. Create external tables in the Apache Hive metastore, referencing the data that is stored in the S3 bucket. Explore the data from the Hive console.
- B Use AWS Glue to crawl the S3 bucket and create tables in the AWS Glue Data Catalog. Use Amazon Athena to explore the data.
- C Create an Amazon Redshift cluster. Use the COPY command to ingest the data from Amazon S3. Explore the data from the Amazon Redshift query editor GUI.
- D Create an Amazon Redshift cluster. Create external tables in an external schema, referencing the S3 bucket that contains the data. Explore the data from the Amazon Redshift query editor GUI.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả tình huống thực tế trong AWS:
Một data scientist đang xây dựng mô hình dự đoán mức tồn kho hàng hóa cho công ty. Toàn bộ dữ liệu lịch sử (khoảng 500 GB dưới dạng file .csv) được lưu trữ trong data lake trên Amazon S3. Data scientist muốn sử dụng SQL để khám phá (explore) dữ liệu trước khi train model. Yêu cầu chính:
- Minimize costs (giảm thiểu chi phí).
- LEAST operational overhead (ít nhất công sức vận hành/quản lý hạ tầng).
🛠️ Yêu cầu kỹ thuật chính: Giải pháp phải hỗ trợ query SQL trực tiếp trên dữ liệu S3 mà không cần di chuyển dữ liệu lớn, serverless để tránh quản lý cluster/server, và chi phí thấp (pay-per-use). Đây là kịch bản điển hình cho serverless analytics trên AWS, phù hợp với kiến thức cập nhật đến 2026 (Athena hỗ trợ federated queries, Glue crawlers tự động schema discovery với ML inference).
📘 Tài liệu tham khảo:
- AWS Athena Documentation (serverless query service).
- AWS Glue Data Catalog (2024+ updates: improved crawler performance).
- AWS Well-Architected Framework: Data Analytics Lens (minimize overhead với serverless).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use AWS Glue to crawl the S3 bucket and create tables in the AWS Glue Data Catalog. Use Amazon Athena to explore the data.
Lý do chi tiết:
🟢 Glue Crawler tự động scan S3 bucket, infer schema từ .csv (hỗ trợ partitioning, compression), tạo metadata tables trong Glue Data Catalog (central catalog serverless). Không cần code ETL phức tạp.
🟢 Amazon Athena là serverless query engine (dựa trên Presto/Trino 400+), query SQL trực tiếp trên S3 mà không di chuyển dữ liệu, scan chỉ dữ liệu cần thiết ( columnar format tối ưu).
✅ LEAST operational overhead: Zero management (no provisioning, scaling tự động), pay-per-query (khoảng $5/TB scanned, minimize costs với 500GB). Phù hợp data lake exploration, tích hợp SageMaker cho ML sau. Không có downtime, scale infinite.
🔍 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên operational overhead, costs, và phù hợp yêu cầu (query SQL trên S3 mà không load data).
-
❌ Phương án SAI: Create an Amazon EMR cluster. Create external tables in the Apache Hive metastore, referencing the data that is stored in the S3 bucket. Explore the data from the Hive console.
Giải thích sai: EMR yêu cầu provision và quản lý cluster (EC2 instances, auto-scaling, termination), overhead cao (setup Hive metastore, config security). Chi phí fixed (cluster runtime ngay cả idle), không serverless. Phù hợp big data processing, nhưng thừa thãi cho chỉ explore SQL 500GB → vi phạm "LEAST overhead" và "minimize costs". -
✅ Phương án ĐÚNG: Use AWS Glue to crawl the S3 bucket and create tables in the AWS Glue Data Catalog. Use Amazon Athena to explore the data.
Giải thích đúng: (Như phần trên) Serverless hoàn toàn, Glue crawler chạy on-demand (1-2 phút), Athena query tức thì qua console/workgroup. Tích hợp Glue Catalog làm hive metastore cho nhiều service. Tối ưu nhất cho data lake query, chi phí thấp (~$0.005/scan 1MB sau tối ưu). -
❌ Phương án SAI: Create an Amazon Redshift cluster. Use the COPY command to ingest the data from Amazon S3. Explore the data from the Amazon Redshift query editor GUI.
Giải thích sai: COPY load toàn bộ 500GB vào Redshift storage (RA3 nodes với managed storage, nhưng vẫn tốn chi phí lưu trữ ~$0.024/GB/tháng). Yêu cầu provision cluster (DC2/DRA nodes, scaling manual), overhead cao (backup, maintenance). Không giữ data tại S3 gốc, tăng costs → không minimize và không least overhead. -
❌ Phương án SAI: Create an Amazon Redshift cluster. Create external tables in an external schema, referencing the S3 bucket that contains the data. Explore the data from the Amazon Redshift query editor GUI.
Giải thích sai: Sử dụng Redshift Spectrum (external tables query S3), nhưng vẫn cần Redshift cluster chạy liên tục (main cluster costs ~$0.25/giờ/node + Spectrum $5/TB scanned). Overhead quản lý cluster cao, không serverless thuần. Phù hợp data warehouse hybrid, nhưng thừa cho chỉ exploration → không least overhead so với Athena.
🏆 Kết luận khuyến nghị
Giải pháp Glue + Athena là best practice cho data lake analytics (S3-centric), dễ mở rộng sang ML (SageMaker Studio tích hợp Athena). Để optimize hơn: Sử dụng S3 Partitioning + columnar formats (Parquet/ORC) để giảm scan costs 90%. Test qua AWS Free Tier Athena! 🚀
The company has configured an Amazon SageMaker training job to use a single ml.p2.xlarge instance with File input mode to train the built-in Object Detection algorithm. The training process was successful last month but is now failing because of a lack of storage. Aside from the addition of training data, nothing has changed in the model training process.
A machine learning (ML) specialist needs to change the training configuration to fix the problem. The solution must optimize performance and must minimize the cost of training.
Which solution will meet these requirements?
- A Modify the training configuration to use two ml.p2.xlarge instances.
- B Modify the training configuration to use Pipe input mode.
- C Modify the training configuration to use a single ml.p3.2xlarge instance.
- D Modify the training configuration to use Amazon Elastic File System (Amazon EFS) instead of Amazon S3 to store the input training data.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS SageMaker
📖 Nội dung câu hỏi:
Câu hỏi mô tả một công ty phân tích địa không gian xử lý hàng nghìn hình ảnh vệ tinh mới mỗi ngày để phát hiện tàu bè cho vận tải thương mại. Dữ liệu huấn luyện (training data) được lưu trữ trên Amazon S3 và kích thước tăng dần hàng ngày với dữ liệu mới.
Công ty sử dụng Amazon SageMaker training job với:
- Một instance ml.p2.xlarge (có 1 GPU NVIDIA K80, dung lượng lưu trữ EBS khoảng 100-150 GB tùy config).
- File input mode để huấn luyện thuật toán built-in Object Detection.
Quá trình huấn luyện thành công tháng trước nhưng bây giờ thất bại do thiếu dung lượng lưu trữ (lack of storage). Không có thay đổi gì ngoài việc thêm dữ liệu huấn luyện.
Nhiệm vụ của ML specialist: Thay đổi cấu hình huấn luyện để sửa lỗi, tối ưu hiệu suất (performance) và giảm thiểu chi phí (minimize cost).
🛠️ Vấn đề cốt lõi:
- File input mode tải toàn bộ dữ liệu từ S3 vào lưu trữ cục bộ của instance (EBS volume), dẫn đến hết dung lượng khi dữ liệu tăng (ml.p2.xlarge có giới hạn storage ~100GB).
- Cần giải pháp xử lý dữ liệu lớn hơn mà không tăng instance hoặc cost không cần thiết, tận dụng streaming data hiệu quả.
✅ Đáp án đúng:
Modify the training configuration to use Pipe input mode.
Lý do chọn đáp án đúng (🟢):
- Pipe input mode cho phép stream dữ liệu trực tiếp từ S3 vào GPU memory mà không tải toàn bộ vào EBS storage, giải quyết triệt để vấn đề thiếu dung lượng.
- Thuật toán built-in Object Detection của SageMaker hỗ trợ Pipe mode (theo docs AWS mới nhất 2024-2026), giúp tăng throughput lên đến 10x so với File mode cho dữ liệu lớn, tối ưu performance bằng cách giảm I/O bottleneck.
- Tiết kiệm chi phí: Giữ nguyên single instance ml.p2.xlarge (không tăng số lượng hoặc loại instance đắt hơn), chỉ thay input mode → chi phí thấp nhất, phù hợp yêu cầu "minimize cost".
- Không ảnh hưởng đến quy trình hiện tại, chỉ cần chỉnh
channelconfig trong training job (ví dụ:InputMode: 'Pipe').
📘 Tài liệu tham khảo:
- AWS SageMaker Data Input Modes (cập nhật 2024): Pipe mode lý tưởng cho datasets >100GB, built-in algos như Object Detection.
- SageMaker Object Detection Algorithm: Xác nhận hỗ trợ Pipe mode.
- Best practices: Optimizing SageMaker Training.
🔍 Giải thích tất cả các phương án (Đúng/Sai)
-
❌ [SAI] Modify the training configuration to use two ml.p2.xlarge instances.
Phương án này sử dụng distributed training (2 instances), tăng khả năng xử lý song song nhưng không giải quyết gốc rễ thiếu storage trên single instance. File mode vẫn tải full data vào mỗi instance → vẫn fail nếu data quá lớn. Tăng cost gấp đôi (2x instance), không tối ưu performance/cost vì vấn đề chỉ là storage, không phải compute. -
✅ [ĐÚNG] Modify the training configuration to use Pipe input mode.
(Đã giải thích chi tiết ở trên: Stream data → fix storage, optimize perf, min cost). -
❌ [SAI] Modify the training configuration to use a single ml.p3.2xlarge instance.
ml.p3.2xlarge (V100 GPU, storage lớn hơn ~250GB) có thể chứa data lớn hơn, nhưng cost cao hơn ~3-4x so với p2.xlarge (theo pricing 2024-2026). Vẫn dùng File mode → không tận dụng tối ưu, và không cần GPU mạnh hơn vì Object Detection trên p2 đã đủ. Vi phạm "minimize cost" và không fix hiệu quả như Pipe mode. -
❌ [SAI] Modify the training configuration to use Amazon Elastic File System (Amazon EFS) instead of Amazon S3 to store the input training data.
EFS dùng cho shared storage multi-instance, nhưng ở đây single instance và input data từ S3. SageMaker không khuyến khích EFS cho input channels (chậm hơn S3, tăng latency I/O). Cost cao (EFS ~$0.3/GB/tháng), không fix File mode issue (vẫn tải data cục bộ), và phức tạp setup → không optimize perf/cost.
💡 Lời khuyên thực tế:
Sử dụng Pipe mode là best practice cho incremental large datasets trên SageMaker (2026 updates vẫn giữ nguyên). Test với sagemaker.pytorch estimator và monitor qua CloudWatch Logs/Metrics để verify throughput tăng. 🚀
The company has an algorithm for transcribing customer calls that requires GPUs for inference. The company wants to store these transcriptions in an Amazon S3 bucket in the AWS Cloud for model development.
Which solution should an ML specialist use to deliver the transcriptions to the S3 bucket as quickly as possible?
- A Order and use an AWS Snowball Edge Compute Optimized device with an NVIDIA Tesla module to run the transcription algorithm. Use AWS DataSync to send the resulting transcriptions to the transcription S3 bucket.
- B Order and use an AWS Snowcone device with Amazon EC2 Inf1 instances to run the transcription algorithm. Use AWS DataSync to send the resulting transcriptions to the transcription S3 bucket.
- C Order and use AWS Outposts to run the transcription algorithm on GPU-based Amazon EC2 instances. Store the resulting transcriptions in the transcription S3 bucket.
- D Use AWS DataSync to ingest the audio files to Amazon S3. Create an AWS Lambda function to run the transcription algorithm on the audio files when they are uploaded to Amazon S3. Configure the function to write the resulting transcriptions to the transcription S3 bucket.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh một công ty đang sử dụng Amazon SageMaker để xây dựng mô hình máy học (ML) dự đoán customer churn dựa trên dữ liệu transcript (bản ghi âm) từ các cuộc gọi khách hàng. Dữ liệu audio gốc nằm ở hệ thống VoIP on-premises với dung lượng petabytes (hàng nghìn terabytes), kết nối với AWS qua VPN 100 Mbps – một đường truyền chậm so với khối lượng dữ liệu khổng lồ.
Công ty có thuật toán transcription (chuyển audio thành text) yêu cầu GPU để inference (suy luận), và họ muốn lưu trữ các bản transcript này vào Amazon S3 bucket trên AWS một cách nhanh nhất có thể để phát triển mô hình ML.
🔑 Thách thức chính:
- Khối lượng dữ liệu lớn (petabytes) → Không thể truyền trực tiếp qua VPN chậm (100 Mbps chỉ khoảng 12.5 MB/s, mất hàng tháng/năm).
- Cần xử lý transcription on-premises để tận dụng GPU địa phương, chỉ gửi file transcript (nhỏ gọn hơn audio gốc rất nhiều).
- Giải pháp phải tối ưu tốc độ, phù hợp với edge computing và hybrid cloud AWS.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Order and use an AWS Snowball Edge Compute Optimized device with an NVIDIA Tesla module to run the transcription algorithm. Use AWS DataSync to send the resulting transcriptions to the transcription S3 bucket.
🛠️ Lý do chi tiết:
- AWS Snowball Edge Compute Optimized là thiết bị edge computing mạnh mẽ, hỗ trợ NVIDIA Tesla GPU (như V100 hoặc tương đương theo cập nhật 2023-2026), lý tưởng để chạy inference GPU-intensive on-premises.
- Thiết bị có dung lượng storage lớn (lên đến 210TB SSD + HDD tùy model), đủ xử lý petabytes bằng cách ship nhiều device.
- Chạy algorithm transcription ngay trên Snowball → Tạo transcript text nhỏ gọn → Sử dụng AWS DataSync (tích hợp sẵn trên Snowball Edge) để sync nhanh chóng lên S3 qua kết nối hiện có hoặc ship về AWS.
- Nhanh nhất: Giảm dữ liệu truyền (text << audio), tránh bottleneck VPN, phù hợp dữ liệu lớn/high-velocity networking on-premises.
- Cập nhật AWS 2026: Snowball Edge hỗ trợ ML workloads với GPU modules, tích hợp SageMaker inference (AWS docs: Snow Family Compute Optimized).
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm giải thích chi tiết bằng tiếng Việt:
-
✅ [ĐÚNG] Order and use an AWS Snowball Edge Compute Optimized device with an NVIDIA Tesla module to run the transcription algorithm. Use AWS DataSync to send the resulting transcriptions to the transcription S3 bucket.
🏆 Phương án tối ưu nhất: Snowball Edge Compute Optimized có GPU NVIDIA Tesla chuyên cho ML inference, xử lý on-premises hiệu quả với petabytes dữ liệu. DataSync đảm bảo truyền transcript nhanh, an toàn lên S3. Không phụ thuộc VPN chậm cho audio gốc. -
❌ [SAI] Order and use an AWS Snowcone device with Amazon EC2 Inf1 instances to run the transcription algorithm. Use AWS DataSync to send the resulting transcriptions to the transcription S3 bucket.
🚫 Snowcone là thiết bị siêu nhỏ gọn (8TB storage, CPU-only, không hỗ trợ GPU NVIDIA Tesla hay EC2 Inf1 – Inf1 là instance Inferentia/Trainium chỉ có trên cloud AWS, không trên edge device). Không đủ sức mạnh cho petabytes và GPU inference, dẫn đến chậm và không khả thi. -
❌ [SAI] Order and use AWS Outposts to run the transcription algorithm on GPU-based Amazon EC2 instances. Store the resulting transcriptions in the transcription S3 bucket.
🔄 AWS Outposts là rack server on-premises chạy full AWS services (bao gồm EC2 GPU như G4/G5), nhưng: (1) Cài đặt phức tạp/dài hạn (tuần/tháng), không "nhanh nhất"; (2) Storage local trên Outposts phải sync với S3 qua VPN 100Mbps chậm; (3) Không tối ưu cho petabytes di động như Snowball (ship được). Phù hợp enterprise lớn nhưng overkill và chậm hơn. -
❌ [SAI] Use AWS DataSync to ingest the audio files to Amazon S3. Create an AWS Lambda function to run the transcription algorithm on the audio files when they are uploaded to Amazon S3. Configure the function to write the resulting transcriptions to the transcription S3 bucket.
⏳ Chậm nhất: DataSync truyền toàn bộ petabytes audio qua VPN 100Mbps → Mất hàng năm! Lambda serverless không hỗ trợ GPU (chỉ CPU, giới hạn 15 phút), không chạy được algorithm GPU inference. Phải dùng SageMaker/Inf1 thay thế, nhưng vẫn bottleneck truyền dữ liệu gốc.
📘 Tài liệu tham khảo (cập nhật AWS 2023-2026)
- AWS Snowball Edge: docs.aws.amazon.com/snowball/latest/developer-guide/sbe.html – Chi tiết GPU modules cho ML.
- DataSync với Snowball: docs.aws.amazon.com/datasync/latest/userguide/create-snowball-job.html.
- AWS ML on Edge: aws.amazon.com/machine-learning/edge/ – SageMaker Edge + Snowball.
- Exam DOP-C02/SAP-C02: Chủ đề Hybrid/ML Data Transfer (Well-Architected Framework: Operational Excellence).
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần thêm ví dụ thực tế, hỏi nhé!
How should the ML specialist design the data ingestion to meet these requirements with the LEAST operational overhead?
- A Ingest event data by using a GraphQLAPI in AWS AppSync. Store the data in an Amazon DynamoDB table. Use DynamoDB Streams to call an AWS Lambda function to transform the most recent 10 minutes of data before inference.
- B Ingest event data by using Amazon Kinesis Data Streams. Store the data in Amazon S3 by using Amazon Kinesis Data Firehose. Use AWS Glue to transform the most recent 10 minutes of data before inference.
- C Ingest event data by using Amazon Kinesis Data Streams. Use an Amazon Kinesis Data Analytics for Apache Flink application to transform the most recent 10 minutes of data before inference.
- D Ingest event data by using Amazon Managed Streaming for Apache Kafka (Amazon MSK). Use an AWS Lambda function to transform the most recent 10 minutes of data before inference.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi xoay quanh việc thiết kế hệ thống ingestion dữ liệu sự kiện (event data) cho một nền tảng podcast có hàng nghìn người dùng. Các sự kiện bao gồm listening, pausing, exiting podcast, được sử dụng để phát hiện bất thường (anomaly detection) dựa trên cửa sổ thời gian chạy liên tục 10 phút (10-minute running window). Dữ liệu sự kiện cần chuyển đổi nhỏ (small transformations) trước khi đưa vào inference ML. Yêu cầu chính là thiết kế với operational overhead thấp nhất (LEAST operational overhead), nghĩa là ưu tiên giải pháp serverless, real-time streaming, hỗ trợ windowing hiệu quả mà không cần quản lý nhiều tài nguyên thủ công.
📱 Bối cảnh AWS mới nhất (2026): AWS nhấn mạnh vào các dịch vụ managed streaming như Kinesis và MSK, với Kinesis Data Analytics hỗ trợ Apache Flink cho real-time analytics, windowing (tumbling/sliding windows), và transformations on-the-fly. Điều này phù hợp với workload high-throughput, low-latency cho ML pipelines.
(Nguồn: AWS Documentation - Amazon Kinesis Data Analytics for Apache Flink, cập nhật 2025: docs.aws.amazon.com/kinesisanalytics)
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Ingest event data by using Amazon Kinesis Data Streams. Use an Amazon Kinesis Data Analytics for Apache Flink application to transform the most recent 10 minutes of data before inference.
🛠️ Lý do chi tiết:
- Amazon Kinesis Data Streams là dịch vụ streaming real-time, chịu tải hàng nghìn sự kiện/giây từ podcast users, hỗ trợ ingestion scalable mà không cần quản lý server.
- Kinesis Data Analytics (KDA) for Apache Flink là serverless, tự động xử lý running window 10 phút (sử dụng sliding/tumbling windows trong Flink SQL/Stateful Streams), áp dụng transformations nhỏ ngay trên stream trước inference.
- Least operational overhead: Toàn bộ serverless (auto-scale, no provisioning shards/topics thủ công), tích hợp trực tiếp với ML services như SageMaker cho inference. Không cần lưu trữ trung gian hay batch processing.
🎯 Ưu điểm vượt trội: Hỗ trợ exactly-once semantics, fault-tolerant, và low-latency (~seconds) cho anomaly detection.
(Nguồn: AWS Well-Architected Framework - ML Lens 2025: aws.amazon.com/architecture/well-architected)
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên yêu cầu real-time windowing, transformations, và least overhead. Sử dụng kiến thức AWS 2026 (Flink 1.18+ trong KDA hỗ trợ advanced windows).
-
❌ Phương án SAI: Ingest event data by using a GraphQLAPI in AWS AppSync. Store the data in an Amazon DynamoDB table. Use DynamoDB Streams to call an AWS Lambda function to transform the most recent 10 minutes of data before inference.
🧨 Lý do sai: AppSync + DynamoDB phù hợp API-driven apps nhưng không phải streaming real-time cho high-volume events (podcast thousands users). DynamoDB Streams trigger Lambda chỉ là near-real-time (~seconds delay), khó maintain 10-min running window (cần query lịch sử thủ công trong Lambda, dễ timeout/memory issues). Overhead cao: Quản lý GSI, Lambda scaling, state management. Không least overhead cho ML streaming. -
❌ Phương án SAI: Ingest event data by using Amazon Kinesis Data Streams. Store the data in Amazon S3 by using Amazon Kinesis Data Firehose. Use AWS Glue to transform the most recent 10 minutes of data before inference.
🧨 Lý do sai: Kinesis Data Streams + Firehose tốt cho ingestion -> S3 (batch durable storage), nhưng AWS Glue là ETL batch-oriented (chạy jobs hàng giờ/ngày), không hỗ trợ real-time 10-min window. Phải crawl S3 thủ công để lấy "most recent 10 minutes" → latency cao, không phù hợp anomaly detection. Overhead: Quản lý Glue jobs, partitions S3, không serverless cho streaming transforms. -
✅ Phương án ĐÚNG: Ingest event data by using Amazon Kinesis Data Streams. Use an Amazon Kinesis Data Analytics for Apache Flink application to transform the most recent 10 minutes of data before inference.
🛠️ Xác nhận đúng: Như giải thích trên, Flink trong KDA xử lý in-stream transformations + windowing native (e.g.,TUMBLEhoặcSLIDEwindow 10 phút), output trực tiếp cho inference. Serverless hoàn toàn, auto-scale theo traffic podcast peaks.
(Nguồn: AWS Blog - Real-time Anomaly Detection with KDA Flink, 2024: aws.amazon.com/blogs/big-data) -
❌ Phương án SAI: Ingest event data by using Amazon Managed Streaming for Apache Kafka (Amazon MSK). Use an AWS Lambda function to transform the most recent 10 minutes of data before inference.
🧨 Lý do sai: Amazon MSK (managed Kafka) mạnh cho streaming nhưng Lambda event source mapping chỉ poll batches, khó maintain stateful 10-min window (Lambda stateless, cần external state như DynamoDB → phức tạp). Overhead cao: Provision MSK clusters (EC2-based, dù managed), manage consumer groups, Lambda concurrency. Không least overhead so với KDA Flink (fully serverless).
🚀 Kết luận & Best Practices
Giải pháp đúng tận dụng Kinesis ecosystem serverless cho ML streaming pipelines, giảm chi phí ~70% so với MSK/Lambda (theo AWS Pricing Calculator 2026). Khuyến nghị: Kết hợp với Amazon SageMaker cho inference end-to-end. Nếu scale lớn hơn, xem xét Kinesis Data Streams enhanced fan-out.
📘 Tài liệu tham khảo chính:
Which approach will meet these requirements with the LEAST operational overhead?
- A Configure an S3 event notification that invokes an AWS Lambda function when new documents are created. Configure the Lambda function to create three SageMaker batch transform jobs, one batch transform job for each model for each document.
- B Deploy all the models to a single SageMaker endpoint. Treat each model as a production variant. Configure an S3 event notification that invokes an AWS Lambda function when new documents are created. Configure the Lambda function to call each production variant and return the results of each model.
- C Deploy each model to its own SageMaker endpoint Configure an S3 event notification that invokes an AWS Lambda function when new documents are created. Configure the Lambda function to call each endpoint and return the results of each model.
- D Deploy each model to its own SageMaker endpoint. Create three AWS Lambda functions. Configure each Lambda function to call a different endpoint and return the results. Configure three S3 event notifications to invoke the Lambda functions when new documents are created.
Xem giải thích
🧩 Giải thích chi tiết nội dung câu hỏi
Câu hỏi này thuộc chủ đề Amazon SageMaker trong AWS, tập trung vào việc triển khai (deploy) nhiều phiên bản mô hình machine learning (ML) để dự đoán phân loại tài liệu thời gian thực (real-time inference), với tần suất cao: tài liệu mới được lưu vào Amazon S3 bucket mỗi 3 giây.
Công ty đã phát triển ba phiên bản mô hình ML trên SageMaker để phân loại văn bản tài liệu. Yêu cầu chính là deploy ba mô hình này để dự đoán cho từng tài liệu mới, đồng thời chọn giải pháp có operational overhead thấp nhất (ít chi phí quản lý, vận hành nhất).
🛠️ Các yếu tố cần xem xét:
- Tần suất cao (mỗi 3 giây): Cần giải pháp real-time (như SageMaker endpoints cho inference online), không phải batch processing.
- Ba mô hình: Nên tận dụng tính năng multi-model endpoint hoặc production variants của SageMaker để giảm số lượng endpoint cần quản lý.
- Trigger: Sử dụng S3 event notification để kích hoạt xử lý tự động khi file mới upload.
- Least overhead: Giảm thiểu số endpoint, Lambda, event notifications; tối ưu chi phí và quản lý (scaling, monitoring).
Mục tiêu là real-time prediction từ ba mô hình với chi phí vận hành thấp nhất theo các best practices AWS SageMaker (cập nhật đến 2026, hỗ trợ production variants cho multi-model hosting).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Deploy all the models to a single SageMaker endpoint. Treat each model as a production variant. Configure an S3 event notification that invokes an AWS Lambda function when new documents are created. Configure the Lambda function to call each production variant and return the results of each model.
Lý do chọn 🏆:
- Giải pháp này sử dụng một SageMaker endpoint duy nhất với production variants (tính năng SageMaker cho phép host nhiều mô hình trên cùng endpoint, chia sẻ tài nguyên compute). Lambda chỉ gọi InvokeEndpoint với target variant khác nhau trên cùng một endpoint, giảm tối đa overhead: chỉ 1 endpoint (dễ scale, monitor), 1 S3 event, 1 Lambda.
- Phù hợp real-time (latency thấp <1s), tần suất cao (3s/document), và least overhead so với các option khác (không tạo nhiều endpoint/Lambda/event).
- Theo AWS best practices 2026: Production variants hỗ trợ traffic splitting, A/B testing, và auto-scaling cho multi-model.
📋 Phân tích tất cả các phương án
-
Phương án 1 ❌:
Configure an S3 event notification that invokes an AWS Lambda function when new documents are created. Configure the Lambda function to create three SageMaker batch transform jobs, one batch transform job for each model for each document.
Giải thích sai: Batch transform jobs dành cho batch inference (xử lý hàng loạt, không real-time), phải chờ job hoàn thành (phù hợp dataset lớn, không phải single doc mỗi 3s). Tạo 3 jobs/document gây overhead cực cao: queue jobs, polling status, chi phí compute lặp lại, không scale tốt. Không đáp ứng real-time. -
Phương án 2 ✅:
Deploy all the models to a single SageMaker endpoint. Treat each model as a production variant. Configure an S3 event notification that invokes an AWS Lambda function when new documents are created. Configure the Lambda function to call each production variant and return the results of each model.
Giải thích đúng: Như đã phân tích ở đáp án đúng. Least overhead nhờ 1 endpoint multi-variant (SageMaker tự quản lý models), 1 Lambda gọi TargetModel param để route đến variant cụ thể. Hiệu quả cho high-frequency inference. -
Phương án 3 ❌:
Deploy each model to its own SageMaker endpoint Configure an S3 event notification that invokes an AWS Lambda function when new documents are created. Configure the Lambda function to call each endpoint and return the results of each model.
Giải thích sai: Tạo 3 endpoints riêng biệt (mỗi endpoint cần instance riêng, scaling độc lập) tăng overhead quản lý (provisioning, monitoring, chi phí gấp 3). Lambda gọi 3 endpoints tăng latency/network calls, dù chỉ 1 S3 event + 1 Lambda. -
Phương án 4 ❌:
Deploy each model to its own SageMaker endpoint. Create three AWS Lambda functions. Configure each Lambda function to call a different endpoint and return the results. Configure three S3 event notifications to invoke the Lambda functions when new documents are created.
Giải thích sai: Overhead cao nhất: 3 endpoints + 3 Lambdas + 3 S3 events (duplicate triggers, concurrency issues). Phức tạp quản lý, chi phí cao, không hiệu quả cho cùng document (có thể race conditions).
📘 Tài liệu tham khảo
- AWS SageMaker Documentation (2026): Host multiple models with production variants và Production Variants – Giải thích multi-model hosting giảm chi phí 50-90%.
- AWS Well-Architected Framework - ML Lens: Khuyến nghị single endpoint cho multi-variant để least overhead real-time inference.
- S3 Event Notifications: docs.aws.amazon.com/AmazonS3/latest/userguide/NotificationHowTo.html.
- SageMaker Inference Best Practices: Whitepaper AWS re:Invent 2025/2026 về optimized endpoints.
Giải pháp này đảm bảo scalable, cost-effective theo tiêu chuẩn DevOps Engineer Professional! 🚀