Ngân hàng đề — AWS Certified Machine Learning Specialty
Tìm thấy 371 câu.
The company will use this data to train a model for real-time defect detection in new parts as the parts move on a conveyor belt in the facilities. The company needs a solution that minimizes costs for compute infrastructure and that maximizes the scalability of resources for training. The solution also must facilitate the company's use of an ML model in the low-connectivity environments.
Which solution will meet these requirements?
- A Move the training data to an Amazon S3 bucket. Train and evaluate the model by using Amazon SageMaker. Optimize the model by using SageMaker Neo. Deploy the model on a SageMaker hosting services endpoint.
- B Train and evaluate the model on premises. Upload the model to an Amazon S3 bucket. Deploy the model on an Amazon SageMaker hosting services endpoint.
- C Move the training data to an Amazon S3 bucket. Train and evaluate the model by using Amazon SageMaker. Optimize the model by using SageMaker Neo. Set up an edge device in the manufacturing facilities with AWS IoT Greengrass. Deploy the model on the edge device.
- D Train the model on premises. Upload the model to an Amazon S3 bucket. Set up an edge device in the manufacturing facilities with AWS IoT Greengrass. Deploy the model on the edge device.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh một công ty sản xuất muốn áp dụng machine learning (ML) để tự động hóa kiểm soát chất lượng sản phẩm tại các cơ sở sản xuất ở vị trí xa xôi với kết nối internet hạn chế. Họ có 20TB dữ liệu huấn luyện (gồm hình ảnh sản phẩm bị lỗi đã được gắn nhãn) lưu trữ tại data center on-premises của công ty. Mục tiêu là huấn luyện mô hình ML để phát hiện lỗi thời gian thực trên các bộ phận sản phẩm mới khi chúng di chuyển trên băng chuyền.
Yêu cầu chính của giải pháp (theo phiên bản AWS mới nhất đến 2026):
- Giảm thiểu chi phí hạ tầng tính toán (compute infrastructure): Tránh đầu tư lớn vào on-premises hardware.
- Tối đa hóa khả năng mở rộng (scalability) cho việc huấn luyện mô hình: Cần xử lý lượng dữ liệu lớn (20TB) một cách linh hoạt.
- Hỗ trợ sử dụng mô hình ở môi trường kết nối thấp (low-connectivity): Mô hình phải chạy cục bộ tại cơ sở sản xuất mà không phụ thuộc internet liên tục.
Giải pháp lý tưởng phải kết hợp cloud training scalable (như SageMaker), lưu trữ dữ liệu (S3), tối ưu hóa mô hình cho edge (SageMaker Neo/Edge), và triển khai edge computing (IoT Greengrass) để chạy offline. SageMaker cung cấp managed training với spot instances để tiết kiệm chi phí, và Greengrass hỗ trợ inference tại edge mà không cần cloud connectivity (cập nhật Greengrass v2.10+ năm 2024-2026 hỗ trợ ML inference tốt hơn).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng là phương án thứ 3:
Move the training data to an Amazon S3 bucket. Train and evaluate the model by using Amazon SageMaker. Optimize the model by using SageMaker Neo. Set up an edge device in the manufacturing facilities with AWS IoT Greengrass. Deploy the model on the edge device.
Lý do chọn đáp án này 🛠️:
- Chuyển dữ liệu sang S3: Dễ dàng, chi phí thấp, tích hợp trực tiếp với SageMaker cho training.
- Huấn luyện và đánh giá trên SageMaker: SageMaker managed service, maximize scalability với distributed training (có thể dùng managed spot training để minimize costs lên đến 90%). Xử lý 20TB dữ liệu hiệu quả với Processing Jobs hoặc Training Jobs.
- Tối ưu hóa bằng SageMaker Neo: Neo (nay tích hợp SageMaker Edge Manager từ 2023+) compile mô hình thành định dạng tối ưu cho edge devices (như TensorFlow Lite, ONNX), giảm kích thước và tăng tốc inference.
- Triển khai trên edge với AWS IoT Greengrass: Greengrass v2 (cập nhật 2026) cho phép deploy mô hình ML components chạy offline hoàn toàn tại facilities, hỗ trợ real-time inference trên conveyor belt mà không cần internet. Hoàn hảo cho low-connectivity.
- Tổng thể: Đáp ứng tất cả yêu cầu – training scalable/cloud-based tiết kiệm chi phí, edge deployment cho môi trường xa xôi.
📋 Phân tích tất cả các phương án (đúng/sai)
Dưới đây là phân tích chi tiết từng phương án. Nội dung phương án giữ nguyên bản tiếng Anh, giải thích hoàn toàn bằng tiếng Việt với lý do đúng/sai dựa trên AWS best practices.
-
Phương án 1 ❌ (SAI):
Move the training data to an Amazon S3 bucket. Train and evaluate the model by using Amazon SageMaker. Optimize the model by using SageMaker Neo. Deploy the model on a SageMaker hosting services endpoint.
Giải thích sai: Phần training trên SageMaker tốt (scalable, low-cost), nhưng deploy trên SageMaker endpoint yêu cầu kết nối internet liên tục để inference (real-time calls qua API). Không phù hợp low-connectivity tại facilities xa xôi. Chi phí endpoint cao hơn edge nếu traffic lớn. -
Phương án 2 ❌ (SAI):
Train and evaluate the model on premises. Upload the model to an Amazon S3 bucket. Deploy the model on an Amazon SageMaker hosting services endpoint.
Giải thích sai: Train on-premises không scalable (cần mua GPU/TPU lớn cho 20TB data, chi phí cao, khó mở rộng), vi phạm minimize costs và maximize scalability. Endpoint vẫn cần connectivity, không giải quyết low-connectivity. -
Phương án 3 ✅ (ĐÚNG):
Move the training data to an Amazon S3 bucket. Train and evaluate the model by using Amazon SageMaker. Optimize the model by using SageMaker Neo. Set up an edge device in the manufacturing facilities with AWS IoT Greengrass. Deploy the model on the edge device.
Giải thích đúng: Như phân tích trên, kết hợp hoàn hảo cloud training scalable/low-cost + edge deployment offline với Greengrass. SageMaker Neo đảm bảo mô hình chạy mượt trên edge hardware (như NVIDIA Jetson hoặc Raspberry Pi). -
Phương án 4 ❌ (SAI):
Train the model on premises. Upload the model to an Amazon S3 bucket. Set up an edge device in the manufacturing facilities with AWS IoT Greengrass. Deploy the model on the edge device.
Giải thích sai: Train on-premises vẫn kém scalable và tốn kém (không tận dụng SageMaker managed resources). Phần edge với Greengrass tốt, nhưng thiếu tối ưu hóa (Neo) và training cloud làm giải pháp không tối ưu chi phí/scalability.
📘 Tài liệu tham khảo (AWS cập nhật đến 2026)
- Amazon SageMaker Documentation: SageMaker Training & SageMaker Edge Manager (thay thế Neo) – Hỗ trợ spot training và edge optimization.
- AWS IoT Greengrass ML Inference: Greengrass ML Components – Offline inference với SageMaker models (v2.12+ năm 2025 hỗ trợ multi-model).
- AWS Well-Architected Framework - ML Lens: ML Lens – Khuyến nghị hybrid cloud-edge cho low-connectivity.
- Exam Prep DOP-C02: Câu hỏi tương tự trong AWS Certified DevOps Engineer Professional (2024-2026 blueprint).
Giải pháp này là best practice cho ML tại edge! 🚀 Nếu cần demo code hoặc architecture diagram, hãy cho tôi biết nhé!
SageMaker. Three compute-optimized instances support the expected peak load of the website.
Response times on the product recommendation page are increasing at the beginning of each month. Some users are encountering errors. The website receives the majority of its traffic between 8 AM and 6 PM on weekdays in a single time zone.
Which of the following options are the MOST effective in solving the issue while keeping costs to a minimum? (Choose two.)
- A Configure the endpoint to use Amazon Elastic Inference (EI) accelerators.
- B Create a new endpoint configuration with two production variants.
- C Configure the endpoint to automatically scale with the InvocationsPerInstance metric.
- D Deploy a second instance pool to support a blue/green deployment of models.
- E Reconfigure the endpoint to use burstable instances.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi mô tả một website thương mại điện tử sử dụng Amazon SageMaker để host endpoint recommendation engine dựa trên TensorFlow. Hệ thống hiện dùng 3 instances compute-optimized (như ml.c5) để chịu tải peak. Vấn đề chính:
- Thời gian phản hồi (response times) tăng dần đầu mỗi tháng, dẫn đến lỗi cho một số user.
- Traffic tập trung cao 8 AM - 6 PM các ngày weekday trong một múi giờ duy nhất (dễ dự đoán).
Mục tiêu: Chọn 2 giải pháp hiệu quả NHẤT để giảm latency và lỗi, đồng thời giữ chi phí tối thiểu.
🛠️ Phân tích vấn đề cốt lõi: Tải tăng đột biến đầu tháng (có thể do hành vi user như mua sắm đầu kỳ), nhưng traffic có pattern rõ ràng → cần scaling linh hoạt và tối ưu inference mà không lãng phí resources 24/7. SageMaker endpoints hỗ trợ auto-scaling, accelerators để tăng performance mà không cần scale instance lớn.
✅ Đáp án đúng (Chọn TWO)
Configure the endpoint to use Amazon Elastic Inference (EI) accelerators.
Configure the endpoint to automatically scale with the InvocationsPerInstance metric.
Lý do lựa chọn:
- Hai giải pháp này trực tiếp giải quyết latency cao và lỗi bằng cách tăng tốc inference (EI accelerators giảm thời gian xử lý TensorFlow mà không cần GPU full) và scale động theo tải (InvocationsPerInstance đo invocations/instance, scale out/up chỉ khi cần, thu hẹp off-peak → min cost).
- Phù hợp pattern traffic dự đoán được (weekday daytime), tránh over-provisioning 3 instances fixed.
- Theo AWS best practices (SageMaker Inference docs 2024-2026), đây là cách cost-effective nhất cho ML endpoints với traffic biến động.
📘 Tài liệu tham khảo: - Amazon SageMaker Endpoints Auto Scaling (InvocationsPerInstance là metric mặc định hiệu quả).
- Amazon Elastic Inference (tăng throughput 50-100% với chi phí thấp hơn GPU, dù deprecated 2024 nhưng vẫn valid cho exam patterns cũ).
🔍 Phân tích chi tiết TẤT CẢ các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên hiệu quả giải quyết vấn đề + min cost.
-
Configure the endpoint to use Amazon Elastic Inference (EI) accelerators.
✅ ĐÚNG. EI gắn accelerators (như eia1.medium) vào instances hiện tại, tăng tốc inference TensorFlow (giảm latency 2-5x cho recommendation models) mà không cần thay instance lớn hơn hoặc scale nhiều. Cost thấp vì chỉ trả thêm cho EI usage thực tế, lý tưởng cho peak đầu tháng. Không ảnh hưởng traffic pattern. -
Create a new endpoint configuration with two production variants.
❌ SAI. Production variants dùng cho A/B testing hoặc multi-model endpoints, không giải quyết scaling/load. Tạo thêm variant chỉ tăng complexity và cost (split traffic), không giảm latency peak. Không liên quan đến traffic tăng đột biến. -
Configure the endpoint to automatically scale with the InvocationsPerInstance metric.
✅ ĐÚNG. SageMaker Application Auto Scaling dùng metric InvocationsPerInstance (số requests/instance/phút) để scale out instances tự động khi tải cao (đầu tháng), scale in off-peak (night/weekend). Min cost nhờ min/max capacity settings phù hợp traffic 8AM-6PM. Metric này tối ưu nhất cho inference endpoints (AWS recommend). -
Deploy a second instance pool to support a blue/green deployment of models.
❌ SAI. Blue/green dùng cho zero-downtime deployment models mới, không phải scaling production load. Tạo pool thứ 2 tăng cost gấp đôi (chạy song song), lãng phí vì traffic không liên tục 24/7. Không giải quyết latency hiện tại. -
Reconfigure the endpoint to use burstable instances.
❌ SAI. Burstable instances (như ml.t3) có CPU credit cho burst, nhưng không phù hợp ML inference nặng (TensorFlow cần compute sustained). Compute-optimized (ml.c5) đã tốt hơn; burstable có thể gây throttling khi peak dài (đầu tháng), tăng latency thay vì giảm, và cost không tiết kiệm hơn auto-scaling. AWS recommend c5/m5 cho SageMaker inference.
🛠️ Tóm tắt: Chỉ 2 giải pháp đúng kết hợp hoàn hảo (tăng tốc + scale thông minh), giữ 3 instances base min capacity để cost thấp!
The accuracy of the predictions with the current model is below 50%. The company wants to improve the model performance and launch the new product as soon as possible.
Which solution will meet these requirements with the LEAST operational overhead?
- A Create a service-linked role for Amazon Elastic Container Service (Amazon ECS) with access to the S3 bucket. Create an ECS cluster that is based on an AWS Deep Learning Containers image. Write the code to perform the feature engineering. Train a logistic regression model for predicting the price, pointing to the bucket with the dataset. Wait for the training job to complete. Perform the inferences.
- B Create an Amazon SageMaker notebook with a new IAM role that is associated with the notebook. Pull the dataset from the S3 bucket. Explore different combinations of feature engineering transformations, regression algorithms, and hyperparameters. Compare all the results in the notebook, and deploy the most accurate configuration in an endpoint for predictions.
- C Create an IAM role with access to Amazon S3, Amazon SageMaker, and AWS Lambda. Create a training job with the SageMaker built-in XGBoost model pointing to the bucket with the dataset. Specify the price as the target feature. Wait for the job to complete. Load the model artifact to a Lambda function for inference on prices of new houses.
- D Create an IAM role for Amazon SageMaker with access to the S3 bucket. Create a SageMaker AutoML job with SageMaker Autopilot pointing to the bucket with the dataset. Specify the price as the target attribute. Wait for the job to complete. Deploy the best model for predictions.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh một công ty bất động sản đang phát triển sản phẩm dự đoán giá nhà mới dựa trên dữ liệu lịch sử lưu trữ dưới dạng file CSV trong Amazon S3. Dữ liệu có header, chứa các trường categorical (dữ liệu phân loại) và missing values (giá trị thiếu). Các data scientist đã sử dụng Python với thư viện open-source để fill missing values bằng 0, loại bỏ tất cả categorical fields, và huấn luyện mô hình bằng linear regression với tham số mặc định. Kết quả accuracy chỉ dưới 50%, cần cải thiện hiệu suất mô hình nhanh chóng với LEAST operational overhead (ít nỗ lực vận hành nhất).
Mục tiêu: Chọn giải pháp tối ưu hóa quy trình ML (feature engineering, xử lý dữ liệu, chọn thuật toán, tuning hyperparameters) để đạt accuracy cao hơn, triển khai nhanh, và giảm thiểu công việc thủ công. 🔍 Đây là bài kiểm tra kiến thức về Amazon SageMaker và các dịch vụ ML tự động hóa trên AWS (cập nhật đến 2026, SageMaker Autopilot vẫn là công cụ hàng đầu cho AutoML).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create an IAM role for Amazon SageMaker with access to the S3 bucket. Create a SageMaker AutoML job with SageMaker Autopilot pointing to the bucket with the dataset. Specify the price as the target attribute. Wait for the job to complete. Deploy the best model for predictions.
Lý do chọn đáp án này (LEAST operational overhead):
🛠️ SageMaker Autopilot (hay SageMaker AutoML) tự động hóa toàn bộ pipeline ML: tự động xử lý missing values (không cần fill thủ công bằng 0), chuyển đổi categorical fields (one-hot encoding, embedding), feature engineering (tạo features mới), chọn thuật toán tốt nhất (linear regression, XGBoost, Deep Learning,...), tuning hyperparameters, và đánh giá mô hình. Chỉ cần chỉ định target column (giá nhà) và S3 bucket, chờ job hoàn thành rồi deploy endpoint. Không cần code Python, notebook, hay kiến thức sâu về ML.
📈 Accuracy cải thiện đáng kể vì Autopilot thử nghiệm hàng trăm mô hình và chọn best candidate. Thời gian nhanh (giờ thay vì ngày), phù hợp "launch as soon as possible".
Tài liệu tham khảo: AWS SageMaker Autopilot Documentation (cập nhật 2024-2026, hỗ trợ tabular data CSV với categorical/missing values).
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên tiêu chí LEAST operational overhead và khả năng cải thiện accuracy từ dữ liệu thô (categorical + missing values).
-
Phương án A (SAI ❌):
Create a service-linked role for Amazon Elastic Container Service (Amazon ECS) with access to the S3 bucket. Create an ECS cluster that is based on an AWS Deep Learning Containers image. Write the code to perform the feature engineering. Train a logistic regression model for predicting the price, pointing to the bucket with the dataset. Wait for the training job to complete. Perform the inferences.
Giải thích sai: Phương án này yêu cầu viết code thủ công cho feature engineering, setup ECS cluster với Deep Learning Containers, và chỉ dùng logistic regression (phù hợp classification, không phải regression cho giá nhà). Overhead cao: quản lý container, scaling ECS, debug code. Không tự động hóa, accuracy khó cải thiện từ linear regression hiện tại. 🏗️ Phức tạp hơn SageMaker managed service. -
Phương án B (SAI ❌):
Create an Amazon SageMaker notebook with a new IAM role that is associated with the notebook. Pull the dataset from the S3 bucket. Explore different combinations of feature engineering transformations, regression algorithms, and hyperparameters. Compare all the results in the notebook, and deploy the most accurate configuration in an endpoint for predictions.
Giải thích sai: Dùng SageMaker Notebook để thử nghiệm thủ công (feature engineering, algorithms, hyperparameters). Overhead lớn: data scientist phải explore combinations (thời gian dài, cần expertise). Không tự động, dễ bias như fill 0/miss categorical. Tuy linh hoạt nhưng vi phạm "LEAST operational overhead" vì phải code và so sánh manual. 📓 Tốn công hơn AutoML. -
Phương án C (SAI ❌):
Create an IAM role with access to Amazon S3, Amazon SageMaker, and AWS Lambda. Create a training job with the SageMaker built-in XGBoost model pointing to the S3 bucket with the dataset. Specify the price as the target feature. Wait for the job to complete. Load the model artifact to a Lambda function for inference on prices of new houses.
Giải thích sai: Dùng SageMaker XGBoost built-in (tốt cho tabular data), nhưng không xử lý tự động categorical/missing values (vẫn cần preprocessing thủ công ngoài job). Chuyển inference sang Lambda thêm overhead (cold start, memory limit cho ML model). Chỉ 1 algorithm (XGBoost), không tuning tự động toàn diện. ⚡ Ít overhead hơn A/B nhưng vẫn cần feature engineering riêng, không optimal cho "improve performance" nhanh. -
Phương án D (ĐÚNG ✅):
Create an IAM role for Amazon SageMaker with access to the S3 bucket. Create a SageMaker AutoML job with SageMaker Autopilot pointing to the bucket with the dataset. Specify the price as the target attribute. Wait for the job to complete. Deploy the best model for predictions.
Giải thích đúng: Như đã nêu ở phần ✅, Autopilot tự động toàn bộ từ dữ liệu thô CSV (header, categorical, missing). Overhead thấp nhất: chỉ tạo IAM role, chạy job, deploy best model. Hỗ trợ regression tasks, accuracy cao nhờ leaderboards. 🚀 Hoàn hảo cho non-expert data scientists, phù hợp yêu cầu.
Tài liệu: SageMaker Autopilot Best Practices (xử lý categorical/missing tự động).
Kết luận 💡: SageMaker Autopilot là giải pháp serverless AutoML lý tưởng cho doanh nghiệp muốn nhanh chóng, ít code. Nếu cần customize sâu hơn, dùng Notebook sau. DevOps Engineer nên ưu tiên managed services để scale!
Which combination of feature engineering techniques should the data scientist use to meet these requirements? (Choose two.)
- A Named entity recognition
- B Coreference
- C Stemming
- D Term frequency-inverse document frequency (TF-IDF)
- E Sentiment analysis
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào quá trình feature engineering (kỹ thuật xây dựng đặc trưng) trong phân tích dữ liệu văn bản (text data) cho một data scientist đang xử lý bình luận khách hàng (customer comments) về sản phẩm của công ty. Mục tiêu là trình bày phân tích khám phá ban đầu (initial exploratory analysis) sử dụng biểu đồ (charts) và word cloud trước khi triển khai mô hình xử lý ngôn ngữ tự nhiên (NLP model).
📌 Yêu cầu chính: Chọn hai kỹ thuật feature engineering phù hợp để chuẩn bị dữ liệu (preprocessing), giúp tạo ra các biểu đồ và word cloud hiệu quả. Đây là bước tiền xử lý cơ bản trong pipeline NLP trên AWS, thường được thực hiện qua Amazon SageMaker Processing Jobs, AWS Glue, hoặc tích hợp với thư viện như NLTK/scikit-learn trong SageMaker Notebook. Phiên bản cập nhật 2026: AWS tiếp tục hỗ trợ các kỹ thuật này qua SageMaker JumpStart và Amazon Bedrock cho generative AI, nhưng feature engineering cơ bản vẫn dựa trên stemming và TF-IDF cho exploratory analysis (theo AWS ML Best Practices 2025+).
✅ Đáp án đúng (Chọn TWO)
Stemming và Term frequency-inverse document frequency (TF-IDF).
Lý do lựa chọn:
- 🛠️ Stemming: Giảm các từ biến thể về dạng gốc (ví dụ: "running", "runs" → "run"), giúp loại bỏ noise, chuẩn hóa dữ liệu để word cloud và charts hiển thị rõ ràng các từ chính mà không bị phân tán. Rất phù hợp cho exploratory analysis ban đầu trên dữ liệu khách hàng không cấu trúc.
- 📊 TF-IDF: Tính toán trọng số từ dựa trên tần suất xuất hiện (TF) và độ hiếm trong toàn bộ corpus (IDF), giúp ưu tiên các từ quan trọng (như từ khóa sản phẩm) trong word cloud và biểu đồ. Điều này làm nổi bật insight từ bình luận khách hàng, chuẩn bị tốt cho mô hình NLP sau.
Kết hợp hai kỹ thuật này tạo vector hóa dữ liệu hiệu quả cho visualization trên AWS SageMaker Canvas hoặc QuickSight.
🧠 Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh:
-
❌ Named entity recognition
Sai: Đây là kỹ thuật trích xuất thực thể có tên (như tên người, tổ chức, địa điểm) từ văn bản, thuộc giai đoạn post-processing hoặc model inference (sử dụng Amazon Comprehend NER), không phải feature engineering cơ bản cho exploratory analysis. Nó không giúp tạo word cloud hay charts ban đầu mà cần mô hình đã train sẵn. -
❌ Coreference
Sai: Kỹ thuật giải quyết tham chiếu đại từ (coreference resolution, ví dụ: "Apple" chỉ công ty hay quả táo?), là bước advanced NLP phức tạp, thường dùng trong model sâu (như BERT trên SageMaker), không phù hợp cho initial feature engineering đơn giản với charts/word cloud. Quá nặng cho exploratory phase. -
✅ Stemming
Đúng: Như đã giải thích, đây là text normalization cơ bản, giảm dạng từ về gốc, giúp dữ liệu sạch cho visualization. Hỗ trợ trực tiếp word cloud bằng cách gộp từ đồng nghĩa, phổ biến trong AWS SageMaker scripts với NLTK. -
✅ Term frequency-inverse document frequency (TF-IDF)
Đúng: Vectorization technique chuẩn cho text features, tính trọng số từ quan trọng, lý tưởng cho word cloud (kích thước từ tỷ lệ trọng số) và charts (top terms). Tích hợp sẵn trong scikit-learn trên SageMaker Processing (phiên bản 2026 hỗ trợ JumpStart models). -
❌ Sentiment analysis
Sai: Đây là model output phân tích cảm xúc (positive/negative), không phải feature engineering để chuẩn bị dữ liệu. Amazon Comprehend Sentiment là service riêng, dùng sau exploratory, không tạo features cho word cloud/charts ban đầu.
📘 Tài liệu tham khảo
- AWS Documentation: Amazon SageMaker Processing for Text Feature Engineering (cập nhật 2025, hỗ trợ NLTK/TF-IDF).
- AWS ML Best Practices: Text Processing in SageMaker (2024-2026 editions nhấn mạnh stemming/TF-IDF cho EDA).
- NLTK Guide (tích hợp AWS): Stemming & TF-IDF chapters, dùng trong SageMaker Notebooks.
Hy vọng phân tích này giúp bạn ôn thi AWS Certified Machine Learning hoặc DevOps Pro! 🚀 Nếu cần code ví dụ SageMaker, hãy hỏi thêm.
What can the data scientist reasonably conclude about the distributional forecast related to the test set?
- A The coverage scores indicate that the distributional forecast is poorly calibrated. These scores should be approximately equal to each other at all quantiles.
- B The coverage scores indicate that the distributional forecast is poorly calibrated. These scores should peak at the median and be lower at the tails.
- C The coverage scores indicate that the distributional forecast is correctly calibrated. These scores should always fall below the quantile itself.
- D The coverage scores indicate that the distributional forecast is correctly calibrated. These scores should be approximately equal to the quantile itself.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc đánh giá mô hình dự báo phân phối (distributional forecast) của GluonTS trên Amazon SageMaker DeepAR.
- Bối cảnh: Một data scientist đang đánh giá mô hình DeepAR (một thuật toán dự báo chuỗi thời gian probabilistic trên SageMaker). Các chỉ số đánh giá trên tập test là coverage score (độ bao phủ) đạt 0.489 tại quantile 0.5 (median) và 0.889 tại quantile 0.9.
- Yêu cầu kết luận: Data scientist có thể suy luận gì về việc calibration (hiệu chỉnh) của distributional forecast trên tập test?
- Coverage score là metric đo lường tỷ lệ các điểm dữ liệu thực tế nằm trong khoảng dự báo tương ứng với quantile q (ví dụ: quantile 0.5 là khoảng chứa 50% xác suất ở giữa phân phối dự báo).
- Trong probabilistic forecasting (dự báo xác suất), một mô hình well-calibrated (hiệu chỉnh tốt) sẽ có coverage score xấp xỉ bằng chính giá trị quantile (Coverage(q) ≈ q). Điều này có nghĩa là khoảng dự báo tại quantile 0.5 nên bao phủ khoảng 50% dữ liệu thực tế, và tại 0.9 là khoảng 90%.
- Kiến thức AWS cập nhật 2026: SageMaker DeepAR (phiên bản mới nhất tích hợp GluonTS) hỗ trợ evaluation metrics như coverage qua SageMaker Processing Jobs hoặc GluonTS library. Metric này được tính tự động trong
gluonts.evaluation(GluonTS v0.15+), đảm bảo calibration cho time-series forecasting. 📘 Tài liệu tham khảo:- GluonTS Documentation - Coverage Metric
- AWS SageMaker DeepAR Guide (cập nhật 2025: hỗ trợ quantile regression với calibration checks).
🛠️ Dữ liệu cụ thể: Coverage 0.489 ≈ 0.5 (rất gần) và 0.889 ≈ 0.9 (rất gần) → Mô hình calibrated tốt!
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: The coverage scores indicate that the distributional forecast is correctly calibrated. These scores should be approximately equal to the quantile itself.
Lý do (🧩 Phân tích sâu):
- Coverage score phải xấp xỉ bằng quantile để chứng tỏ mô hình calibrated đúng: khoảng dự báo "đáng tin cậy" (reliable intervals).
- Ở đây, 0.489 ~ 0.5 và 0.889 ~ 0.9 → Hoàn hảo! Mô hình dự báo phân phối khớp với dữ liệu thực tế trên test set, không under/over-confident.
- Theo GluonTS/SageMaker best practices, sai lệch <5% (như 1.1% và 1.2% ở đây) là chấp nhận được cho production.
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] The coverage scores indicate that the distributional forecast is poorly calibrated. These scores should be approximately equal to each other at all quantiles.
Giải thích: Sai vì coverage KHÔNG phải bằng nhau giữa các quantile. Mỗi quantile q có coverage riêng (tăng dần theo q: 0.5 < 0.9). Nếu bằng nhau, mô hình sẽ poorly calibrated (ví dụ: quá hẹp ở tail quantiles). 🧩 Metric đúng phải theo quy luật Coverage(q) ≈ q. -
❌ [SAI] The coverage scores indicate that the distributional forecast is poorly calibrated. These scores should peak at the median and be lower at the tails.
Giải thích: Sai hoàn toàn! Coverage tăng tuyến tính theo quantile (cao hơn ở tails), không "peak tại median" (0.5). Nếu peak tại median, mô hình sẽ overestimate uncertainty ở giữa và underestimate ở tails → poorly calibrated. 📘 GluonTS docs xác nhận: coverage monotonic increasing. -
❌ [SAI] The coverage scores indicate that the distributional forecast is correctly calibrated. These scores should always fall below the quantile itself.
Giải thích: Sai vì coverage KHÔNG luôn "dưới quantile" (fall below). Phải xấp xỉ bằng (≈ q) cho calibration tốt. "Always below" ám chỉ mô hình conservative (quá rộng intervals), dẫn đến low sharpness nhưng không calibrated chính xác. -
✅ [ĐÚNG] The coverage scores indicate that the distributional forecast is correctly calibrated. These scores should be approximately equal to the quantile itself.
Giải thích: Đúng 100%! Dữ liệu khớp hoàn hảo (0.489≈0.5, 0.889≈0.9). Đây là tiêu chuẩn vàng cho proper scoring rules trong probabilistic forecasts trên SageMaker/GluonTS. 🛠️ Trong thực tế, data scientist có thể deploy model này với confidence cao!
Kết luận tổng quát 🎯: Câu hỏi kiểm tra hiểu biết sâu về evaluation metrics trong SageMaker DeepAR. Chọn đúng giúp tránh overfit/underfit trong time-series production. Nếu implement, dùng sagemaker.session.Estimator với metric_definitions cho coverage tracking!
A team of data scientists is using the telemetry data to perform machine learning (ML) to conduct anomaly detection and predict maintenance before the devices start to deteriorate. The team needs a scalable, secure, high-velocity data ingestion mechanism. The team has decided to use Amazon S3 as the data storage location.
Which approach meets these requirements?
- A Ingest the data by using an HTTP API call to a web server that is hosted on Amazon EC2. Set up EC2 instances in an Auto Scaling configuration behind an Elastic Load Balancer to load the data into Amazon S3.
- B Ingest the data over Message Queuing Telemetry Transport (MQTT) to AWS IoT Core. Set up a rule in AWS IoT Core to use Amazon Kinesis Data Firehose to send data to an Amazon Kinesis data stream that is configured to write to an S3 bucket.
- C Ingest the data over Message Queuing Telemetry Transport (MQTT) to AWS IoT Core. Set up a rule in AWS IoT Core to direct all MQTT data to an Amazon Kinesis Data Firehose delivery stream that is configured to write to an S3 bucket.
- D Ingest the data over Message Queuing Telemetry Transport (MQTT) to Amazon Kinesis data stream that is configured to write to an S3 bucket.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một công ty năng lượng sở hữu các thiết bị IoT như turbine gió, trạm thời tiết và panel năng lượng mặt trời, chúng liên tục tạo ra dữ liệu telemetry (dữ liệu đo lường thời gian thực). Các thiết bị này nằm ở nhiều vị trí địa lý khác nhau và có kết nối internet không ổn định, đòi hỏi giải pháp ingestion dữ liệu phải chịu được ngắt kết nối tạm thời, hỗ trợ giao thức nhẹ cho thiết bị edge.
Mục tiêu chính là thực hiện bảo trì dự đoán (predictive maintenance) bằng cách sử dụng dữ liệu này cho machine learning (ML): phát hiện bất thường (anomaly detection) và dự đoán hỏng hóc trước khi thiết bị suy giảm. Đội ngũ data scientists cần cơ chế ingestion dữ liệu scalable (mở rộng), secure (bảo mật), high-velocity (tốc độ cao, xử lý lượng lớn dữ liệu real-time). Amazon S3 được chọn làm nơi lưu trữ cuối cùng.
Yêu cầu cốt lõi (theo best practices AWS mới nhất 2026):
- Hỗ trợ MQTT (giao thức nhẹ, pub/sub lý tưởng cho IoT với kết nối kém ổn định).
- Scalable & high-velocity: Xử lý hàng triệu thiết bị, batch dữ liệu để tối ưu chi phí.
- Secure: Xác thực thiết bị, mã hóa dữ liệu.
- Direct to S3: Không trung gian phức tạp, tận dụng AWS IoT Core làm gateway IoT.
🛠️ Giải pháp lý tưởng: Sử dụng AWS IoT Core làm điểm đầu vào cho MQTT, kết hợp AWS IoT Rules để route dữ liệu trực tiếp đến Amazon Kinesis Data Firehose (hỗ trợ batching, transformation, và delivery trực tiếp đến S3 với buffer để chịu tải cao và kết nối kém).
✅ Đáp án đúng
Ingest the data over Message Queuing Telemetry Transport (MQTT) to AWS IoT Core. Set up a rule in AWS IoT Core to direct all MQTT data to an Amazon Kinesis Data Firehose delivery stream that is configured to write to an S3 bucket.
Lý do lựa chọn:
- ✅ MQTT hoàn hảo cho IoT: Hỗ trợ QoS (Quality of Service) 0/1/2 để đảm bảo delivery ngay cả khi kết nối gián đoạn (reconnect tự động).
- ✅ AWS IoT Core là dịch vụ managed IoT, secure với X.509 certs, device shadows, fleet provisioning (cập nhật 2026 hỗ trợ IoT FleetWise cho edge ML).
- ✅ IoT Rule direct republish MQTT topics đến Kinesis Data Firehose (không cần stream trung gian), Firehose tự động batch, compress, encrypt và deliver to S3 với S3 bucket partitioning (theo thời gian/thiết bị).
- ✅ Scalable & high-velocity: Firehose xử lý >10,000 records/sec, buffer up to 128MB/24h, VPC endpoints cho security. Phù hợp predictive maintenance với SageMaker integration từ S3.
- ❌ Không có điểm yếu: Tránh overhead của EC2/stream phức tạp.
📋 Phân tích tất cả các phương án
-
❌ Phương án A: Ingest the data by using an HTTP API call to a web server that is hosted on Amazon EC2. Set up EC2 instances in an Auto Scaling configuration behind an Elastic Load Balancer to load the data into Amazon S3.
- Lý do sai: HTTP không phù hợp cho thiết bị IoT kết nối kém (stateless, không retry tự động như MQTT). EC2 + ELB + Auto Scaling tốn kém quản lý (patching, scaling manual), không high-velocity cho telemetry real-time (latency cao). Không secure native cho IoT (cần tự implement auth). Theo AWS best practices 2026, dùng managed services như IoT Core thay vì self-managed EC2.
-
❌ Phương án B: Ingest the data over Message Queuing Telemetry Transport (MQTT) to AWS IoT Core. Set up a rule in AWS IoT Core to use Amazon Kinesis Data Firehose to send data to an Amazon Kinesis data stream that is configured to write to an S3 bucket.
- Lý do sai: Phức tạp thừa: IoT Rule gửi đến Firehose, rồi Firehose lại push đến Kinesis Data Streams (KDS) trước khi KDS to S3. Điều này tạo hai lớp streaming (double buffering), tăng latency/cost (KDS shards $0.015/giờ), không cần thiết vì Firehose direct to S3 đã đủ (batch 1-128MB). AWS docs 2026 khuyến nghị Firehose standalone cho S3 sink, tránh KDS trừ khi cần exactly-once/low-latency consumer.
-
✅ Phương án C: Ingest the data over Message Queuing Telemetry Transport (MQTT) to AWS IoT Core. Set up a rule in AWS IoT Core to direct all MQTT data to an Amazon Kinesis Data Firehose delivery stream that is configured to write to an S3 bucket.
- Lý do đúng: Như phần ✅ trên. Direct & efficient: IoT Core → Rule → Firehose → S3 (1 hop). Hỗ trợ data transformation (Lambda), error handling (S3 backup), và enhanced fan-out (2026 updates). Scalable cho hàng triệu devices, tích hợp ML (S3 → SageMaker).
-
❌ Phương án D: Ingest the data over Message Queuing Telemetry Transport (MQTT) to Amazon Kinesis data stream that is configured to write to an S3 bucket.
- Lý do sai: Kinesis Data Streams (KDS) không hỗ trợ MQTT native – cần custom producer/app dùng Kinesis Producer Library (KPL) trên thiết bị/edge, phức tạp cho IoT unstable (không device management). Không secure như IoT Core (thiếu certs/fleet indexing). KDS yêu cầu consumer (KCL/Lambda) để pull to S3, tăng operational overhead so với Firehose direct.
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- AWS IoT Core Developer Guide: docs.aws.amazon.com/iot/latest/developerguide/iot-rules.html – Rules actions cho Firehose.
- Kinesis Data Firehose: docs.aws.amazon.com/firehose/latest/dev/what-is-this-service.html – Direct to S3, IoT integration.
- AWS IoT Best Practices: aws.amazon.com/blogs/iot/iot-data-ingestion-to-s3-using-kinesis-firehose – Case study tương tự.
- DOP-C02 Exam Guide (2026): Nhấn mạnh managed IoT + streaming cho high-velocity edge data.
🛠️ Khuyến nghị triển khai: Enable IoT Device Defender cho anomaly detection real-time, kết nối S3 với Athena/SageMaker cho ML pipeline!
Each product can be classified into multiple categories that the company defines. These categories are related but are not mutually exclusive. For example, if there is mention of "Sample Yogurt" in the document of customer comments, then "Sample Yogurt" should be classified as "yogurt," "snack," and "dairy product."
The team is using Amazon Comprehend to train the model and must complete the project as soon as possible.
Which functionality of Amazon Comprehend should the team use to meet these requirements?
- A Custom classification with multi-class mode
- B Custom classification with multi-label mode
- C Custom entity recognition
- D Built-in models
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một công ty bán lẻ thu thập bình luận khách hàng từ mạng xã hội, website công ty và nhật ký cuộc gọi. Nhóm data scientists và engineers muốn sử dụng Natural Language Processing (NLP) trên Amazon Comprehend để xây dựng mô hình phân loại, nhằm xác định chủ đề chung và sản phẩm mà khách hàng đề cập.
Điểm quan trọng:
- Mỗi sản phẩm có thể thuộc nhiều categories (danh mục) do công ty tự định nghĩa, và các categories này không loại trừ lẫn nhau (related but not mutually exclusive).
- Ví dụ: Nếu bình luận đề cập "Sample Yogurt", nó phải được phân loại đồng thời vào "yogurt", "snack" và "dairy product".
- Nhóm cần train mô hình tùy chỉnh trên Amazon Comprehend và hoàn thành dự án nhanh nhất có thể.
Mục tiêu chính: Chọn chức năng phù hợp của Amazon Comprehend để hỗ trợ phân loại multi-label (nhiều nhãn không độc quyền), cập nhật theo phiên bản mới nhất của AWS (hỗ trợ Custom Classification multi-label từ năm 2021 và ổn định đến 2026).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Custom classification with multi-label mode
Lý do:
- Amazon Comprehend hỗ trợ Custom Classification với chế độ multi-label mode, cho phép mô hình dự đoán nhiều nhãn cùng lúc cho một tài liệu mà không yêu cầu các nhãn phải loại trừ lẫn nhau. Điều này khớp hoàn hảo với yêu cầu: một bình luận có thể gắn nhiều categories (như yogurt + snack + dairy).
- Chế độ này được thiết kế để train nhanh trên dữ liệu tùy chỉnh, giúp hoàn thành dự án "as soon as possible" 🛠️. AWS khuyến nghị sử dụng multi-label cho các trường hợp categories liên quan và chồng chéo.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh:
-
Custom classification with multi-class mode ❌
Sai vì: Multi-class mode chỉ hỗ trợ một nhãn duy nhất cho mỗi tài liệu (mutually exclusive), không phù hợp với yêu cầu multi-categories không độc quyền. Nếu dùng, mô hình sẽ buộc chọn chỉ một category (ví dụ chỉ "yogurt" mà bỏ qua "snack" và "dairy"), dẫn đến kết quả không chính xác. -
Custom classification with multi-label mode ✅
Đúng vì: Như đã giải thích ở trên, chế độ này cho phép nhiều nhãn đồng thời (multi-label), train trên dữ liệu tùy chỉnh nhanh chóng, khớp ví dụ "Sample Yogurt" với nhiều categories. Đây là tính năng tối ưu của Comprehend cho NLP classification chồng chéo 🏆. -
Custom entity recognition ❌
Sai vì: Custom Entity Recognition dùng để nhận diện và trích xuất entity cụ thể (như tên sản phẩm "Sample Yogurt"), chứ không phải phân loại vào categories. Nó không hỗ trợ multi-label classification cho chủ đề/categories, nên không đáp ứng yêu cầu phân loại sản phẩm vào nhiều danh mục. -
Built-in models ❌
Sai vì: Built-in models (như Detect Entities, Keyphrases, Sentiment) là mô hình sẵn có, không tùy chỉnh được cho categories cụ thể của công ty. Chúng không hỗ trợ train multi-label cho dữ liệu riêng, và không giải quyết được việc phân loại sản phẩm vào nhiều categories tùy chỉnh một cách chính xác.
📘 Tài liệu tham khảo (cập nhật AWS 2026)
- AWS Documentation - Amazon Comprehend Custom Classification: Custom Classification – Chi tiết multi-label vs multi-class.
- AWS Blog - Multi-label Custom Classification: Introducing multi-label mode (ra mắt 2021, ổn định đến 2026).
- Exam Guide DOP-C02: Phần ML Services trên Comprehend, nhấn mạnh custom multi-label cho non-exclusive categories.
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm ví dụ code hoặc demo, hãy hỏi nhé!
SageMaker training job.
Which combination of steps should the data engineer take to meet these requirements? (Choose three.)
- A Create a SageMaker development endpoint in the data science team's VPC.
- B Create an AWS Glue development endpoint in the data science team's VPC.
- C Create SageMaker notebooks by using the AWS Glue development endpoint.
- D Create SageMaker notebooks by using the SageMaker console.
- E Attach a decryption policy to the SageMaker notebooks.
- F Create an IAM policy and an IAM role for the SageMaker notebooks.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào việc thiết lập môi trường để data science team có thể truy cập trực tiếp các ETL scripts từ AWS Glue thông qua Amazon SageMaker notebooks nằm trong VPC (Virtual Private Cloud). Cụ thể:
- Data engineer đang sử dụng AWS Glue để tạo datasets tối ưu hóa và bảo mật trong Amazon S3.
- Data science team cần:
- Truy cập ETL scripts trực tiếp từ SageMaker notebooks trong VPC.
- Sau khi setup, chạy AWS Glue job và kích hoạt SageMaker training job từ notebooks đó.
- Yêu cầu chọn 3 bước kết hợp để đáp ứng (dựa trên kiến thức AWS mới nhất đến 2026, AWS Glue và SageMaker hỗ trợ tích hợp chặt chẽ qua Glue Development Endpoints để phát triển ETL code an toàn trong VPC).
Mục tiêu chính là tích hợp Glue Development Endpoint với SageMaker notebooks để dev/test ETL scripts, đồng thời cấp quyền IAM để thực thi jobs. Điều này đảm bảo bảo mật (VPC), tối ưu hóa (access trực tiếp scripts), và khả năng thực thi (run Glue job + SageMaker training).
✅ Đáp án đúng (Chọn 3 phương án sau)
Các bước đúng là sự kết hợp hoàn hảo để:
- Tạo môi trường dev endpoint trong VPC cho Glue ETL.
- Kết nối SageMaker notebooks với endpoint đó để access scripts.
- Cấp quyền IAM để notebooks có thể run Glue job và invoke SageMaker training.
Lý do lựa chọn:
- Theo tài liệu AWS (cập nhật 2024-2026), Glue Development Endpoint là công cụ chính để dev ETL scripts từ SageMaker notebooks trong VPC, hỗ trợ Jupyter kernel và tích hợp trực tiếp. Kết hợp IAM role đảm bảo notebooks có quyền
glue:StartJobRunvàsagemaker:CreateTrainingJob. Không cần SageMaker dev endpoint (không tồn tại) hay các bước thừa khác.
🛠️ Phân tích tất cả các phương án (Đúng/Sai)
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên tính khả thi, phù hợp yêu cầu và docs AWS mới nhất:
-
❌ Create a SageMaker development endpoint in the data science team's VPC.
Sai: SageMaker không có khái niệm "development endpoint" (dev endpoint chỉ dành cho AWS Glue để dev ETL scripts). SageMaker sử dụng Notebook Instances hoặc Studio cho dev, nhưng không hỗ trợ dev endpoint kiểu Glue. Sử dụng cái này sẽ không cho phép access ETL scripts từ Glue trực tiếp trong VPC, dẫn đến lỗi tích hợp. -
✅ Create an AWS Glue development endpoint in the data science team's VPC.
Đúng: Đây là bước đầu tiên cần thiết. Glue Development Endpoint (cập nhật Glue 4.0+) cho phép tạo môi trường Jupyter trong VPC để dev/test ETL scripts an toàn, kết nối với S3 datasets. Endpoint này hỗ trợ VPC endpoints và security groups, đáp ứng yêu cầu access scripts từ notebooks. -
✅ Create SageMaker notebooks by using the AWS Glue development endpoint.
Đúng: Khi tạo SageMaker Notebook Instance, bạn chọn Glue Development Endpoint trong console để attach (tích hợp kernel Glue Python/Scala). Điều này cho phép notebooks access trực tiếp ETL scripts từ endpoint trong VPC, hỗ trợ iterative dev mà không cần copy code thủ công. -
❌ Create SageMaker notebooks by using the SageMaker console.
Sai: Có thể tạo notebooks qua console, nhưng cách này không attach với Glue Dev Endpoint, nên data science team không access được ETL scripts trực tiếp từ Glue. Notebooks sẽ cô lập, không đáp ứng yêu cầu VPC-integrated dev. -
❌ Attach a decryption policy to the SageMaker notebooks.
Sai: Không có "decryption policy" chuẩn cho SageMaker notebooks (có thể nhầm với KMS decryption cho encrypted data). Yêu cầu tập trung vào access scripts và run jobs, không liên quan decrypt. IAM policy cho Glue/SageMaker jobs mới cần, không phải decrypt-specific. -
✅ Create an IAM policy and an IAM role for the SageMaker notebooks.
Đúng: Bước cuối cùng để thực thi. Role cho Notebook Instance cần policy với quyền nhưglue:StartJobRun,glue:GetJobRun,sagemaker:CreateTrainingJob,s3:GetObject(cho S3 scripts/datasets). Attach role qua ExecutionRole khi tạo notebook, đảm bảo team có thể run Glue job và invoke SageMaker training sau setup.
📘 Tài liệu tham khảo (AWS Docs cập nhật 2024-2026)
- AWS Glue Development Endpoints: Developing AWS Glue ETL scripts using SageMaker notebooks – Hướng dẫn tạo endpoint trong VPC và attach SageMaker.
- SageMaker Notebook + Glue Integration: Use AWS Glue interactive sessions in SageMaker Studio (Glue 4.0 hỗ trợ Studio kernels).
- IAM for SageMaker Notebooks: IAM roles for Amazon SageMaker – Ví dụ policy cho Glue jobs.
- AWS Exam DOP-C02 Guide: Topic "Automation" nhấn mạnh Glue-SageMaker hybrid workflows (AWS Certified DevOps Engineer Professional).
Setup này đảm bảo bảo mật cao (VPC), tối ưu (truy cập trực tiếp), và scalability (run jobs on-demand)! 🚀 Nếu cần demo code, hãy hỏi thêm nhé!
TransactionTimestamp (Timestamp)
CardName (Varchar)
CardNo (Varchar)
The data engineer must provide the data so that any row with a CardNo value of NULL is removed. Also, the TransactionTimestamp column must be separated into a TransactionDate column and a TransactionTime column. Finally, the CardName column must be renamed to NameOnCard.
The data will be extracted on a monthly basis and will be loaded into an S3 bucket. The solution must minimize the effort that is needed to set up infrastructure for the ingestion and transformation. The solution also must be automated and must minimize the load on the Amazon Redshift cluster.
Which solution meets these requirements?
- A Set up an Amazon EMR cluster. Create an Apache Spark job to read the data from the Amazon Redshift cluster and transform the data. Load the data into the S3 bucket. Schedule the job to run monthly.
- B Set up an Amazon EC2 instance with a SQL client tool, such as SQL Workbench/J, to query the data from the Amazon Redshift cluster directly Export the resulting dataset into a file. Upload the file into the S3 bucket. Perform these tasks monthly.
- C Set up an AWS Glue job that has the Amazon Redshift cluster as the source and the S3 bucket as the destination. Use the built-in transforms Filter, Map, and RenameField to perform the required transformations. Schedule the job to run monthly.
- D Use Amazon Redshift Spectrum to run a query that writes the data directly to the S3 bucket. Create an AWS Lambda function to run the query monthly.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc một data engineer cần trích xuất và biến đổi dữ liệu từ Amazon Redshift để cung cấp cho đội ngũ data scientists chạy các job huấn luyện machine learning. Dữ liệu sẽ được lưu trữ cuối cùng trong Amazon S3.
📋 Yêu cầu cụ thể về dữ liệu và biến đổi:
- Nguồn dữ liệu: Amazon Redshift với schema mẫu bao gồm các cột
TransactionTimestamp(Timestamp),CardName(Varchar),CardNo(Varchar). - Biến đổi bắt buộc:
- Loại bỏ bất kỳ hàng nào có giá trị
CardNolà NULL (sử dụng filter). - Tách cột
TransactionTimestampthành hai cột riêng:TransactionDate(ngày) vàTransactionTime(giờ). - Đổi tên cột
CardNamethànhNameOnCard.
- Loại bỏ bất kỳ hàng nào có giá trị
- Tần suất: Trích xuất hàng tháng (monthly basis).
- Yêu cầu giải pháp (rất quan trọng - phải đáp ứng TẤT CẢ):
- ✅ Giảm thiểu công sức thiết lập infrastructure (minimize effort for setup).
- ✅ Tự động hóa (automated).
- ✅ Giảm tải cho Redshift cluster (minimize load on Redshift).
🛠️ Giải pháp lý tưởng: Phải là dịch vụ serverless ETL, hỗ trợ kết nối trực tiếp Redshift → S3, có built-in transforms cho filter/map/rename, và scheduler tự động, không cần quản lý server/cluster.
📘 Dẫn nguồn tham khảo (cập nhật AWS 2026):
- AWS Glue ETL Documentation: https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-transforms.html (Filter, Map, RenameField transforms).
- AWS Glue Connectors: https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-connect.html#aws-glue-programming-etl-connect-redshift.
- AWS Redshift UNLOAD vs. Glue: https://docs.aws.amazon.com/redshift/latest/dg/r_UNLOAD.html (so sánh load).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Set up an AWS Glue job that has the Amazon Redshift cluster as the source and the S3 bucket as the destination. Use the built-in transforms Filter, Map, and RenameField to perform the required transformations. Schedule the job to run monthly.
Lý do chọn đáp án này 🏆:
- AWS Glue là dịch vụ serverless ETL hoàn hảo: Kết nối trực tiếp Redshift (source) → S3 (sink) qua JDBC, không cần setup infrastructure thủ công.
- Built-in transforms chính xác:
Filter: Loại bỏ rows cóCardNo = NULL.Map: TáchTransactionTimestampthànhTransactionDate(e.g., DATE(transaction_timestamp)) vàTransactionTime(e.g., TIME(transaction_timestamp)).RenameField: ĐổiCardName→NameOnCard.
- Tự động hóa: Schedule job hàng tháng qua AWS Glue Triggers hoặc Amazon EventBridge.
- Giảm tải Redshift: Glue đọc dữ liệu qua JDBC với parallelism, không unload trực tiếp từ cluster (tránh heavy query load).
- Minimize effort: Visual ETL job generator, code-free hoặc PySpark script đơn giản, deploy trong phút.
❌ Phân tích tất cả các phương án (đúng/sai)
-
Phương án 1: Set up an Amazon EMR cluster. Create an Apache Spark job to read the data from the Amazon Redshift cluster and transform the data. Load the data into the S3 bucket. Schedule the job to run monthly.
❌ Sai vì: Yêu cầu setup EMR cluster (managed Hadoop/Spark), tốn effort cao (provision instances, configure security, scaling). Không serverless, vẫn load Redshift qua JDBC/Spark connector. Dù có schedule (Step Functions/Airflow), không minimize infrastructure như Glue. (Không phù hợp yêu cầu "minimize effort"). -
Phương án 2: Set up an Amazon EC2 instance with a SQL client tool, such as SQL Workbench/J, to query the data from the Amazon Redshift cluster directly Export the resulting dataset into a file. Upload the file into the S3 bucket. Perform these tasks monthly.
❌ Sai vì: Manual hoàn toàn (EC2 + SQL client như SQL Workbench/J), không automated (phải "perform monthly" thủ công). Setup EC2 tốn effort (AMI, security groups, patching). Load nặng Redshift do query trực tiếp, export file rồi upload S3 – không scale, không serverless. -
Phương án 3 (ĐÚNG): Set up an AWS Glue job that has the Amazon Redshift cluster as the source and the S3 bucket as the destination. Use the built-in transforms Filter, Map, and RenameField to perform the required transformations. Schedule the job to run monthly.
✅ Đúng như phân tích trên: Serverless, zero-infra, transforms built-in chính xác, schedule dễ dàng, low-load Redshift. Hoàn hảo match tất cả requirements. -
Phương án 4: Use Amazon Redshift Spectrum to run a query that writes the data directly to the S3 bucket. Create an AWS Lambda function to run the query monthly.
❌ Sai vì: Redshift Spectrum chỉ đọc dữ liệu từ S3 (external tables), KHÔNG hỗ trợ write trực tiếp từ Redshift ra S3 qua Spectrum. Để write dùngUNLOADcommand (nhưng vẫn load nặng cluster, không transform phức tạp như split/rename). Lambda trigger query ok cho automation, nhưng không handle transforms và vi phạm "minimize load" (query join trực tiếp trên Redshift). (Cập nhật 2026: Spectrum vẫn read-only từ S3).
🧠 Kết luận: AWS Glue là lựa chọn best practice cho ETL Redshift → S3 với transforms nhẹ, serverless 100%! Nếu implement, dùng Glue Studio cho visual dev. 🚀
How should the ML specialist package the Docker container so that SageMaker can launch the training correctly?
- A Specify the server argument in the ENTRYPOINT instruction in the Dockerfile.
- B Specify the training program in the ENTRYPOINT instruction in the Dockerfile.
- C Include the path to the training data in the docker build command when packaging the container.
- D Use a COPY instruction in the Dockerfile to copy the training program to the /opt/ml/train directory.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc đóng gói Docker container cho thuật toán huấn luyện tùy chỉnh (custom training algorithm) trên Amazon SageMaker. Một chuyên gia ML muốn triển khai thuật toán tùy chỉnh bằng Docker container, và SageMaker hỗ trợ điều này. Vấn đề chính là cách cấu hình Dockerfile để SageMaker có thể khởi chạy training job một cách chính xác.
SageMaker yêu cầu container phải tuân thủ cấu trúc chuẩn (SageMaker container structure):
- Code huấn luyện nằm ở
/opt/ml/code/. - Dữ liệu đầu vào ở
/opt/ml/input/data/. - Output model ở
/opt/ml/model/. - ENTRYPOINT trong Dockerfile phải chỉ định chương trình huấn luyện chính (training program), thường là script Python nhận arguments từ SageMaker (như hyperparameters, channels qua environment variables).
Điều này đảm bảo SageMaker có thể mount volumes, inject data và hyperparameters tự động tại runtime. Kiến thức dựa trên AWS SageMaker phiên bản mới nhất (2026), không thay đổi cơ bản từ các bản trước.
✅ Đáp án đúng
Specify the training program in the ENTRYPOINT instruction in the Dockerfile.
Lý do chọn đáp án này:
- Theo tài liệu chính thức AWS, ENTRYPOINT trong Dockerfile phải chỉ định trực tiếp chương trình huấn luyện (ví dụ:
ENTRYPOINT ["python3", "/opt/ml/code/train.py"]). - SageMaker sẽ gọi container với arguments chuẩn (hyperparameters, input paths), và ENTRYPOINT đảm bảo script chạy đúng, xử lý env vars như
SM_CHANNEL_TRAINING,SM_HP_LEARNING_RATE. - Đây là yêu cầu bắt buộc cho bring-your-own-container (BYOC) trong training jobs, giúp SageMaker khởi chạy seamless mà không cần chỉnh sửa thêm.
📋 Giải thích chi tiết từng phương án
-
Specify the server argument in the ENTRYPOINT instruction in the Dockerfile.
❌ Sai. "Server argument" thường dùng cho hosting endpoint (inference server như MMS hoặc Triton), không phải training. Training chỉ cần training program, không phải server logic. Nếu dùng server ở đây, container sẽ không xử lý training data đúng cách. -
Specify the training program in the ENTRYPOINT instruction in the Dockerfile.
✅ Đúng. Như giải thích ở trên, đây là cách chuẩn để SageMaker gọi training script. Container sẽ nhận input từ SageMaker volumes và env vars, chạy script và output model vào/opt/ml/model/. -
Include the path to the training data in the docker build command when packaging the container.
❌ Sai. Docker build chỉ build image tĩnh, không biết data path tại runtime. SageMaker tự động mount training data vào/opt/ml/input/data/training/qua S3 channels. Đưa data path vào build command sẽ làm image không portable và lỗi khi deploy. -
Use a COPY instruction in the Dockerfile to copy the training program to the /opt/ml/train directory.
❌ Sai./opt/ml/trainkhông tồn tại trong cấu trúc SageMaker (thường là nhầm với/opt/ml/input/data/training/dành cho data). Code phải copy vào/opt/ml/code/(ví dụ:COPY train.py /opt/ml/code/), rồi ENTRYPOINT chạy từ đó. Copy sai path sẽ khiến SageMaker không tìm thấy script.
📘 Tài liệu tham khảo
- AWS SageMaker Documentation: Adapt Your Own Training Container – Xác nhận ENTRYPOINT cho training program.
- Your Algorithms – Training Container: docs.aws.amazon.com/sagemaker/latest/dg/your-algorithms-training-algo.html – Chi tiết cấu trúc
/opt/ml/và Dockerfile requirements (cập nhật 2026). - SageMaker Best Practices: AWS re:Post và GitHub samples (sagemaker-training-toolkit).
🛠️ Mẹo thực hành: Khi build, dùng docker build -t my-algo . và push lên ECR, rồi tạo Training Job với AlgorithmSpecification.TrainingImage. Test local bằng docker run với mock env vars!