Ngân hàng đề — AWS Certified Machine Learning Engineer Associate
Tìm thấy 635 câu.
Which solution will meet these requirements?
- A Use a vertical bar chart to visualize the outliers. Use a calculated field in QuickSight to take the square roots of the outlier prices to generate the chart. Configure a custom AWS Lambda function to scan the data for anomalies.
- B Use AWS Glue DataBrew to preprocess the data. Set the REMOVE_OUTLIERS operation to eliminate data rows that include unusually high prices. Invoke an AWS Lambda function to store the removed rows in Amazon DynamoDB.
- C Use a vertical bar chart to visualize the outliers. Use a calculated field in QuickSight to square the outlier prices to generate the chart. Use QuickSight anomaly detection insights to determine which prices are unusually high.
- D Use a QuickSight filter to find the lowest 10 values for sneaker price. Assign a specific color to the 10 lowest values.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi
Câu hỏi tập trung vào việc sử dụng Amazon QuickSight để xử lý và trực quan hóa dữ liệu giá bán giày sneaker (sneakers) được tổng hợp từ nhiều website bán lẻ. 🛍️
- Yêu cầu chính:
- Xác định giá cao bất thường (unusually high outliers) trong dữ liệu thời gian. 📈
- Hiển thị trực quan (visually display) các outliers này trên dashboard QuickSight.
- Bối cảnh: Dữ liệu là giá bán scraped (thu thập tự động) từ nhiều nguồn, cần phân tích outlier mà không làm thay đổi dữ liệu gốc, ưu tiên giải pháp native của QuickSight để đơn giản và hiệu quả.
- Mục tiêu: Giải pháp phải detect outliers tự động và visualize chúng một cách rõ ràng, phù hợp với tính năng ML built-in của QuickSight (cập nhật đến 2026: QuickSight hỗ trợ Anomaly Detection Insights mạnh mẽ hơn với ML tự động).
📘 Tài liệu tham khảo:
- Amazon QuickSight Anomaly Detection (phiên bản mới nhất 2026: Tích hợp ML insights để detect outliers dựa trên thống kê và pattern thời gian).
- QuickSight Visual Types & Calculated Fields.
✅ Đáp án đúng
Use a vertical bar chart to visualize the outliers. Use a calculated field in QuickSight to square the outlier prices to generate the chart. Use QuickSight anomaly detection insights to determine which prices are unusually high.
Lý do lựa chọn:
- ✅ QuickSight anomaly detection insights là tính năng ML native (built-in) tự động detect outliers cao/thấp dựa trên dữ liệu lịch sử và pattern, hoàn hảo cho yêu cầu "determine which prices are unusually high outliers". Không cần code custom.
- ✅ Vertical bar chart lý tưởng để visualize outliers (cột cao nổi bật).
- ✅ Calculated field để square (bình phương) outlier prices: Làm scale up giá trị outliers để chúng nổi bật hơn trên chart (ví dụ: giá 100$ thành 10,000, dễ thấy), kết hợp anomaly insights để filter/select outliers.
- Giải pháp đơn giản, serverless, chi phí thấp, phù hợp best practice AWS 2026.
❌ Phân tích tất cả các phương án
-
Use a vertical bar chart to visualize the outliers. Use a calculated field in QuickSight to take the square roots of the outlier prices to generate the chart. Configure a custom AWS Lambda function to scan the data for anomalies.
❌ Sai:- Square roots (căn bậc hai) làm giảm giá trị (ví dụ: 100$ → ~10$), khiến outliers ít nổi bật hơn trên chart, ngược với yêu cầu visualize.
- Lambda custom để scan anomalies phức tạp, không cần thiết vì QuickSight có anomaly detection built-in (tiết kiệm thời gian phát triển). Không khớp best practice native tools.
-
Use AWS Glue DataBrew to preprocess the data. Set the REMOVE_OUTLIERS operation to eliminate data rows that include unusually high prices. Invoke an AWS Lambda function to store the removed rows in Amazon DynamoDB.
❌ Sai:- REMOVE_OUTLIERS trong DataBrew xóa dữ liệu outliers, nhưng yêu cầu là detect và display chúng, không phải loại bỏ.
- Lưu vào DynamoDB qua Lambda thêm layer phức tạp, không trực tiếp visualize trên QuickSight dashboard. Không phù hợp luồng end-to-end.
-
Use a vertical bar chart to visualize the outliers. Use a calculated field in QuickSight to square the outlier prices to generate the chart. Use QuickSight anomaly detection insights to determine which prices are unusually high.
✅ Đúng (như phân tích ở phần trên): Kết hợp hoàn hảo ML anomaly detection + calculated field scale up + bar chart để detect và visualize chính xác. -
Use a QuickSight filter to find the lowest 10 values for sneaker price. Assign a specific color to the 10 lowest values.
❌ Sai:- Filter lowest 10 values (giá thấp nhất) ngược hoàn toàn với "unusually high outliers" (giá cao bất thường).
- Chỉ color hóa top 10 thấp, không detect outliers động bằng ML, thiếu tính tự động và chính xác cho dữ liệu thời gian lớn.
Kết luận: Giải pháp đúng tận dụng QuickSight ML-native để scalable và dễ maintain. 🏆 Nếu implement, bắt đầu bằng dataset → Add insight → Anomaly detection → Custom field + Bar chart!
What effect will the ML engineer observe in the anomaly detection results if the ML engineer changes the Direction parameter to Lower than expected?
- A Increased anomaly identification frequency and increased recall
- B Decreased anomaly identification frequency and decreased recall
- C Increased anomaly identification frequency and decreased recall
- D Decreased anomaly identification frequency and increased recall
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào tính năng Anomaly Detection (phát hiện bất thường) trong Amazon QuickSight, một dịch vụ BI và visualization của AWS sử dụng ML để tự động phát hiện các giá trị bất thường trong dữ liệu thời gian thực, chẳng hạn như nhiệt độ hoạt động của máy móc.
-
Bối cảnh: Kỹ sư ML đang theo dõi nhiệt độ máy cao bất thường hoặc thấp bất thường so với mức bình thường. Họ thiết lập:
- Severity parameter: "Low and above" → Phát hiện các bất thường từ mức độ Low trở lên (bao gồm Low, Medium, High).
- Direction parameter: "All" → Phát hiện cả hai hướng bất thường (cao hơn mong đợi - Higher than expected, và thấp hơn mong đợi - Lower than expected).
-
Thay đổi: Kỹ sư thay Direction từ "All" sang "Lower than expected" → Chỉ phát hiện bất thường thấp hơn mong đợi (ví dụ: nhiệt độ máy thấp bất thường).
-
Hiệu ứng cần phân tích: Sự thay đổi này ảnh hưởng như thế nào đến tần suất phát hiện bất thường (anomaly identification frequency) và recall (tỷ lệ thu hồi - khả năng phát hiện đúng các bất thường thực sự, tính bằng True Positives / (True Positives + False Negatives)).
Theo tài liệu AWS QuickSight mới nhất (cập nhật đến 2026), việc giới hạn Direction sẽ giảm phạm vi phát hiện, dẫn đến ít bất thường được xác định hơn và recall giảm vì bỏ lỡ các bất thường ở hướng ngược lại. Điều này phù hợp với mô hình ML của QuickSight dựa trên các thuật toán như Random Cut Forest (RCF) để tính toán deviation từ baseline.
✅ Đáp án đúng
Decreased anomaly identification frequency and decreased recall
Lý do lựa chọn (🛠️ Giải thích chi tiết):
- Ban đầu với Direction = "All", QuickSight phát hiện cả high và low anomalies (ví dụ: nhiệt độ quá cao hoặc quá thấp).
- Sau thay đổi sang "Lower than expected", chỉ phát hiện low anomalies → Tần suất phát hiện giảm (frequency decreased) vì loại bỏ hoàn toàn high anomalies.
- Recall giảm vì hệ thống bỏ lỡ các true positives ở hướng high (false negatives tăng), làm tỷ lệ thu hồi tổng thể thấp hơn. Severity vẫn "Low and above" không thay đổi, nhưng Direction là yếu tố quyết định phạm vi.
Điều này giúp tinh chỉnh alert chính xác hơn cho low temperatures, nhưng đánh đổi bằng việc miss high risks.
📋 Phân tích tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên hành vi của QuickSight Anomaly Detection (phiên bản 2026: hỗ trợ granular control qua Direction để giảm noise).
-
Increased anomaly identification frequency and increased recall
❌ Sai: Thay đổi Direction sang "Lower than expected" không tăng tần suất (vì thu hẹp phạm vi từ All xuống chỉ low), và recall cũng không tăng (miss high anomalies). Điều này ngược với logic filter của QuickSight. -
Decreased anomaly identification frequency and decreased recall
✅ Đúng: Như giải thích trên, tần suất giảm (ít anomalies hơn), recall giảm (bỏ lỡ một phần true positives). Phù hợp với best practices AWS: Sử dụng Direction để prioritize specific deviations. -
Increased anomaly identification frequency and decreased recall
❌ Sai: Tần suất không tăng (filter chặt hơn), dù recall có giảm. QuickSight không amplify detections khi restrict Direction; nó chỉ focus subset. -
Decreased anomaly identification frequency and increased recall
❌ Sai: Tần suất đúng là giảm, nhưng recall không tăng (vì precision có thể tăng cho low anomalies, nhưng recall tổng thể giảm do miss high ones). Recall đo lường coverage toàn bộ anomalies.
📘 Tài liệu tham khảo
- AWS QuickSight Documentation (2026): Anomaly Detection Parameters - Chi tiết Severity, Direction (All/Higher/Lower), và metrics như frequency/recall.
- AWS re:Post & Blogs: QuickSight ML Insights - Ví dụ thực tế về impact của Direction trên detection rates.
- Exam Prep: AWS Certified Machine Learning - Specialty (MLS-C01) & DevOps Pro (DOP-C02) đề cập QuickSight trong monitoring/ML ops.
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần ví dụ code hoặc demo, hãy hỏi thêm.
The company needs to perform a lift and shift from the on-premises Kubernetes cluster to an Amazon Elastic Kubernetes Service (Amazon EKS) cluster.
Which solution will meet this requirement with the LEAST operational overhead?
- A Redesign the ML services to be configured in Kubeflow. Deploy the new Kubeflow managed ML services to the EKS cluster.
- B Upload the Docker images to an Amazon Elastic Container Registry (Amazon ECR) repository. Configure a deployment pipeline to deploy the images to the EKS cluster.
- C Migrate the training data to an Amazon Redshift cluster. Retrain the models from the migrated training data by using Amazon Redshift ML. Deploy the retrained models to the EKS cluster.
- D Configure an Amazon SageMaker AI notebook. Retrain the models with the same code. Deploy the retrained models to the EKS cluster.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc lift and shift (di chuyển trực tiếp mà không thay đổi lớn) các workflow ML đang chạy trên Kubernetes cluster on-premises sang Amazon Elastic Kubernetes Service (Amazon EKS).
- Bối cảnh: Công ty có các ML services thực hiện training (huấn luyện mô hình) và inference (dự đoán) cho các ML models. Mỗi service chạy từ Docker image độc lập (standalone).
- Yêu cầu chính: Giải pháp phải đáp ứng với LEAST operational overhead (ít nhất gánh nặng vận hành), nghĩa là ưu tiên cách di chuyển đơn giản, nhanh chóng, không cần redesign code, retrain models hay thay đổi kiến trúc lớn.
- Mục tiêu: Di chuyển nguyên trạng (lift and shift) từ on-premises K8s sang EKS, tận dụng tính tương thích của Kubernetes manifests và container images.
Đây là kịch bản điển hình trong AWS DevOps khi migrate containerized workloads, nhấn mạnh vào ECR (container registry) và EKS (managed K8s). Kiến thức cập nhật đến 2026: EKS hỗ trợ Kubernetes phiên bản 1.30+ với Fargate và EC2, tích hợp seamless với ECR qua IAM roles và deployment pipelines (như GitHub Actions, CodePipeline).
📘 Tài liệu tham khảo:
- AWS EKS Documentation: Lift workloads to EKS.
- Amazon ECR Best Practices: Push Docker images to ECR.
- AWS Well-Architected Framework - Migration: Phần Operational Excellence pillar.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Upload the Docker images to an Amazon Elastic Container Registry (Amazon ECR) repository. Configure a deployment pipeline to deploy the images to the EKS cluster.
Lý do 🛠️:
- Đây là giải pháp lift and shift thuần túy với least operational overhead: Chỉ cần push Docker images hiện có lên ECR (registry managed của AWS, tích hợp trực tiếp với EKS), sau đó dùng deployment pipeline (như AWS CodePipeline, ArgoCD hoặc kubectl apply) để deploy manifests Kubernetes lên EKS.
- Không cần thay đổi code, retrain models hay redesign services → Tiết kiệm thời gian, chi phí và rủi ro.
- Tương thích 100%: EKS chạy Kubernetes chuẩn, hỗ trợ Helm charts và images từ ECR qua Image Pull Secrets hoặc IRSA (IAM Roles for Service Accounts).
- Hiệu suất cao: ECR scanning tự động, lifecycle policies, replication cross-region (cập nhật 2024-2026).
📋 Giải thích chi tiết tất cả các phương án
-
Redesign the ML services to be configured in Kubeflow. Deploy the new Kubeflow managed ML services to the EKS cluster.
❌ Sai vì yêu cầu high operational overhead: Phải redesign toàn bộ services để phù hợp Kubeflow (framework ML trên K8s với pipelines, notebooks). Kubeflow on EKS (qua eks-kubeflow addon) tốt cho ML native nhưng không phải lift and shift – cần refactor code, configs, và test lại → Tăng thời gian deploy, phức tạp hóa migration. -
Upload the Docker images to an Amazon Elastic Container Registry (Amazon ECR) repository. Configure a deployment pipeline to deploy the images to the EKS cluster.
✅ Đúng như đã giải thích ở trên: Đơn giản nhất, tận dụng Docker images sẵn có, ECR + pipeline (CI/CD) đảm bảo deploy tự động, scalable trên EKS mà không thay đổi gì lớn. -
Migrate the training data to an Amazon Redshift cluster. Retrain the models from the migrated training data by using Amazon Redshift ML. Deploy the retrained models to the EKS cluster.
❌ Sai vì không phải lift and shift: Phải migrate data sang Redshift (data warehouse), retrain models bằng Redshift ML (tích hợp SageMaker dưới hood từ 2023+). Overhead cao: ETL data, refactor training logic, không giữ nguyên Docker services → Phù hợp greenfield, không cho migration nhanh. -
Configure an Amazon SageMaker AI notebook. Retrain the models with the same code. Deploy the retrained models to the EKS cluster.
❌ Sai vì thay đổi lớn: Dùng SageMaker Notebook (Studio hoặc JupyterLab) để retrain (dù code giống), sau deploy endpoints sang EKS. Overhead: Chuyển từ containerized K8s sang managed SageMaker → Mất tính độc lập của services, cần tích hợp SageMaker SDK, không least effort cho lift and shift. (Cập nhật 2026: SageMaker hỗ trợ K8s inference nhưng vẫn require refactor).
Kết luận 🚀: Giải pháp đúng ưu tiên container portability – core principle của AWS container services! Nếu implement, khuyến nghị dùng eksctl để tạo cluster và AWS Load Balancer Controller cho ingress.
Which solution will meet these requirements?
- A Use Amazon Q Business to develop the responses. Configure a document attribute filter so that responses about prices use only the documents from the past month.
- B Use Amazon Q Business to develop the responses. Configure the source attribution citation so that responses about prices use only the documents from the past month.
- C Segment the documents into folders based on the month of document creation. Configure Amazon Q Developer to use only the documents from the past month to develop responses about prices.
- D Segment the documents into folders based on the month of document creation. Grant the assistant access to only the documents from the past month for responses about prices. Use Amazon Q Developer to develop the responses.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh một công ty bán lẻ đang xây dựng trợ lý AI hỗ trợ khách hàng (AI-powered assistant), sử dụng một tập tài liệu lớn (large body of documentation) để trả lời các câu hỏi chung. Yêu cầu đặc biệt là các phản hồi liên quan đến giá cả (prices) phải chỉ sử dụng tài liệu mới hơn 1 tháng (less than 1 month old).
📌 Mục tiêu chính: Tìm giải pháp tích hợp AI AWS để xử lý RAG (Retrieval-Augmented Generation), đảm bảo tính chính xác thời gian thực cho thông tin giá, tránh sử dụng dữ liệu cũ dẫn đến sai sót kinh doanh. Giải pháp phải linh hoạt, không cần phân đoạn thủ công tài liệu, và hỗ trợ filter động dựa trên nội dung truy vấn (query về giá).
🛠️ Bối cảnh AWS cập nhật 2026: Sử dụng Amazon Q Business (dịch vụ generative AI cho doanh nghiệp, hỗ trợ tùy chỉnh index, metadata filtering cho semantic search). Không dùng Amazon Q Developer (chỉ dành cho code gen trong IDE như VS Code).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Amazon Q Business to develop the responses. Configure a document attribute filter so that responses about prices use only the documents from the past month.
Lý do chọn 🏆:
- Amazon Q Business hỗ trợ document attribute filters (lọc dựa trên metadata như ngày tạo/update), áp dụng động theo truy vấn (ví dụ: chỉ filter cho query về "prices" qua filter expressions hoặc custom plugins).
- Tài liệu có thể được ingest với attributes (như
document_agehoặclast_updated), sau đó configure retriever để filter< 1 monthchỉ khi query match keyword "price". - Đảm bảo tuân thủ yêu cầu mà không cần phân đoạn folder thủ công, scalable cho large docs.
- Nguồn tham khảo: AWS Amazon Q Business Docs - Filters & Metadata Filtering in Q Business (cập nhật 2025-2026).
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên tính năng AWS mới nhất.
-
Use Amazon Q Business to develop the responses. Configure a document attribute filter so that responses about prices use only the documents from the past month.
✅ Đúng 🥇: Như giải thích trên, attribute filter cho phép lọc metadata động (e.g.,lastModified < now() - 30 days) chỉ áp dụng cho query về prices qua query-time filtering hoặc conversational routing. Hoàn hảo cho RAG enterprise, hỗ trợ hybrid search (semantic + keyword). Không cần thay đổi kiến trúc docs. -
Use Amazon Q Business to develop the responses. Configure the source attribution citation so that responses about prices use only the documents from the past month.
❌ Sai 🚫: Source attribution citation chỉ dùng để trích dẫn nguồn (citations) trong response (e.g., link docs), không filter retrieval. Nó không kiểm soát docs được dùng để generate, dẫn đến vẫn pull docs cũ cho prices. Tính năng này chỉ transparency, không phải filtering (xem Q Business Citations Docs). -
Segment the documents into folders based on the month of document creation. Configure Amazon Q Developer to use only the documents from the past month to develop responses about prices.
❌ Sai 🔒: Amazon Q Developer là tool cho developers (code suggestions, chat in IDE), không hỗ trợ business docs hoặc RAG cho customer assistant. Phân đoạn folder thủ công (S3 folders) phức tạp, không scalable, và Q Developer không index/connect S3 như Q Business. (Nguồn: Amazon Q Developer vs Q Business). -
Segment the documents into folders based on the month of document creation. Grant the assistant access to only the documents from the past month for responses about prices. Use Amazon Q Developer to develop the responses.
❌ Sai 🗂️: Tương tự trên, phân đoạn folder + IAM access (grant access) không linh hoạt cho dynamic query (không tự động switch theo nội dung câu hỏi). Amazon Q Developer sai tool (không phải cho end-user assistant). Q Business mới là lựa chọn đúng cho enterprise RAG với folder connectors, nhưng cách này vẫn fail vì tool sai.
🏅 Kết luận & Lời khuyên DevOps
Giải pháp tối ưu dùng Amazon Q Business với metadata-driven filtering để zero-downtime updates docs giá mới hàng tháng qua data sources sync (S3/Confluence). Triển khai với CloudFormation cho IaC. Test với Q Business console để verify filter accuracy! 🚀
Tài liệu bổ sung: Amazon Q Business Best Practices (2026 edition).
The company needs to use Amazon SageMaker AI built-in algorithms for the model. An ML engineer converts the categorical features by using one-hot encoding.
Which algorithm should the ML engineer implement to meet these requirements?
- A Use the CatBoost algorithm to recommend the next airport destination.
- B Use the DeepAR forecasting algorithm to recommend the next airport destination.
- C Use the Factorization Machines algorithm to recommend the next airport destination.
- D Use the k-means algorithm to cluster users into groups. Map each group to the next airport destination based on user search history.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh việc một công ty du lịch muốn xây dựng mô hình Machine Learning (ML) để gợi ý sân bay đích đến tiếp theo cho người dùng. Dữ liệu đầu vào bao gồm hàng triệu bản ghi về vị trí người dùng, lịch sử tìm kiếm gần đây trên website công ty, và 2.000 sân bay khả dụng. Đặc điểm nổi bật:
- Nhiều tính năng phân loại (categorical features).
- Cột mục tiêu (target column) dự kiến tạo ra ma trận thưa thớt chiều cao (high-dimensional sparse matrix) – thường xảy ra sau khi áp dụng one-hot encoding cho categorical features, dẫn đến số lượng cột lớn (ví dụ: 2.000 sân bay có thể tạo 2.000 cột binary).
- Yêu cầu sử dụng các thuật toán built-in của Amazon SageMaker (không phải third-party).
- ML engineer đã chuyển đổi categorical features bằng one-hot encoding.
Mục tiêu: Chọn thuật toán SageMaker phù hợp nhất để xử lý dữ liệu thưa thớt cao chiều cho hệ thống gợi ý (recommendation system). 📘 Tài liệu tham khảo: Amazon SageMaker Built-in Algorithms (cập nhật AWS 2024-2026, Factorization Machines vẫn là lựa chọn chuẩn cho recommendation với sparse data).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use the Factorization Machines algorithm to recommend the next airport destination.
Lý do:
- 🛠️ Factorization Machines (FM) là thuật toán built-in của SageMaker, được thiết kế đặc biệt cho dữ liệu thưa thớt cao chiều (high-dimensional sparse data) như sau one-hot encoding.
- Nó mô hình hóa tương tác giữa các tính năng (feature interactions) một cách hiệu quả, lý tưởng cho recommendation systems (gợi ý sân bay dựa trên lịch sử tìm kiếm và vị trí).
- Xử lý tốt multi-class classification hoặc ranking cho target như sân bay (2000 lớp), với độ chính xác cao trên dữ liệu lớn. SageMaker hỗ trợ FM từ phiên bản 2018 và vẫn tối ưu đến 2026.
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] Use the CatBoost algorithm to recommend the next airport destination.
Lý do sai: CatBoost là thư viện third-party (từ Yandex), không phải built-in algorithm của SageMaker. SageMaker hỗ trợ XGBoost hoặc Linear Learner cho gradient boosting, nhưng CatBoost yêu cầu custom container hoặc script riêng, vi phạm yêu cầu "built-in algorithms". Không tối ưu cho sparse recommendation. -
❌ [SAI] Use the DeepAR forecasting algorithm to recommend the next airport destination.
Lý do sai: DeepAR là thuật toán dự báo chuỗi thời gian (time-series forecasting) built-in của SageMaker, dùng cho dữ liệu liên tục theo thời gian (như doanh số bán hàng). Không phù hợp cho recommendation discrete (sân bay cụ thể) với sparse categorical data; sẽ kém hiệu quả và không xử lý high-dimensional target tốt. -
✅ [ĐÚNG] Use the Factorization Machines algorithm to recommend the next airport destination.
Lý do đúng: Như đã giải thích ở trên – hoàn hảo cho sparse high-dimensional data sau one-hot encoding, chuyên cho recommendation/ranking. SageMaker FM hỗ trợ training nhanh trên hàng triệu records với GPU/CPU. -
❌ [SAI] Use the k-means algorithm to cluster users into groups. Map each group to the next airport destination based on user search history.
Lý do sai: k-means là thuật toán clustering unsupervised built-in, dùng để nhóm dữ liệu tương đồng (cluster users), không phải recommendation trực tiếp. Phải thêm bước "map group to airport" thủ công dựa trên search history – phức tạp, không tận dụng sparse matrix tốt, và kém chính xác so với FM cho personalized recommendation.
🔗 Tài liệu tham khảo chính thức (cập nhật AWS 2026)
- 📘 Factorization Machines in Amazon SageMaker – Chi tiết về xử lý sparse data.
- 📘 Built-in Algorithms Overview – Danh sách đầy đủ, xác nhận FM cho recommendation.
- 🛠️ SageMaker Examples: Recommendation with FM.
Phân tích này dựa trên best practices AWS DevOps Engineer Professional DOP-C02 (2024-2026). 🚀
The ML engineer must ensure that the SageMaker AI endpoint can handle incoming requests at the start of each business day.
Which solution will meet this requirement?
- A Reduce the SageMaker AI auto scaling cooldown period to the minimum supported value. Add an auto scaling lifecycle hook to scale the SageMaker AI instances.
- B Change the target metric to CPU utilization.
- C Modify the scaling policy target value to one.
- D Apply a step scaling policy that scales based on an Amazon CloudWatch alarm. Apply a second CloudWatch alarm and scaling policy to scale the minimum number of instances from zero to one at the start of each business day.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh việc cấu hình auto scaling cho endpoint SageMaker AI trong một ứng dụng ML inference. Cụ thể:
- Một kỹ sư ML thiết lập SageMaker AI auto scaling sử dụng target tracking scaling policy với mục tiêu 100 invocations per model per minute (100 lời gọi mô hình mỗi phút).
- Hệ thống scale tốt trong giờ làm việc bình thường (normal business hours), nghĩa là nó tự động tăng/giảm instance dựa trên traffic.
- Vấn đề: Vào đầu mỗi ngày kinh doanh (start of each business day), không có instance nào sẵn sàng (zero instances available), dẫn đến delay xử lý requests vì phải chờ scale up từ 0.
- Yêu cầu: Đảm bảo endpoint SageMaker có thể xử lý requests ngay từ đầu ngày kinh doanh, tránh tình trạng cold start (khởi động lạnh từ 0 instance).
🛠️ Nguyên nhân gốc rễ: SageMaker cho phép scale down về 0 instance khi không có traffic (ví dụ: ban đêm hoặc cuối tuần) để tiết kiệm chi phí. Với target tracking policy, scaling chỉ kích hoạt khi có invocations vượt ngưỡng, nên đầu ngày (khi traffic đột ngột tăng) sẽ mất thời gian provision instance mới (cold start có thể mất vài phút).
📘 Kiến thức AWS cập nhật 2026: SageMaker hỗ trợ Target Tracking, Step Scaling, và Scheduled Scaling qua Application Auto Scaling. Từ 2023-2026, AWS khuyến nghị kết hợp CloudWatch Alarms với Step Scaling cho các pattern predictable như business day start, và Minimum Capacity để tránh cold start (xem SageMaker docs).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Apply a step scaling policy that scales based on an Amazon CloudWatch alarm. Apply a second CloudWatch alarm and scaling policy to scale the minimum number of instances from zero to one at the start of each business day.
Lý do 🟢:
- Giải pháp này chủ động warm up endpoint bằng cách sử dụng Step Scaling Policy kết hợp CloudWatch Alarm để scale dựa trên metric invocations.
- Alarm thứ nhất: Theo dõi invocations và trigger step scaling (tăng/giảm theo bước định sẵn, linh hoạt hơn target tracking cho spike traffic).
- Alarm thứ hai + scaling policy: Scheduled hoặc time-based alarm để tự động scale min instances từ 0 lên 1 ngay đầu ngày kinh doanh (sử dụng CloudWatch Scheduled Actions hoặc Composite Alarms). Điều này tránh cold start, giữ ít nhất 1 instance "warm" sẵn sàng.
- Hoàn hảo cho pattern predictable daily traffic, tiết kiệm chi phí (chỉ 1 instance min ban đêm) và scale nhanh.
❌ Phân tích tất cả các phương án (đúng/sai)
-
Phương án SAI: Reduce the SageMaker AI auto scaling cooldown period to the minimum supported value. Add an auto scaling lifecycle hook to scale the SageMaker AI instances.
Giải thích sai 🔴: Giảm cooldown period (thời gian chờ giữa các scaling action, min là 60s) chỉ giúp scale nhanh hơn sau khi đã có traffic, không giải quyết cold start từ 0 instance. Lifecycle hook (từ EC2 ASG) dùng để pause scaling và chạy script, nhưng SageMaker không hỗ trợ lifecycle hooks trực tiếp (chỉ qua Lambda integration phức tạp). Không chủ động warm up đầu ngày. -
Phương án SAI: Change the target metric to CPU utilization.
Giải thích sai 🔴: Chuyển metric sang CPU utilization không liên quan vì vấn đề là zero instances, không phải overload CPU. Invocations per model/min là metric phù hợp nhất cho SageMaker inference (proxy cho load). CPU chỉ gián tiếp và kém chính xác cho serverless-like scaling của SageMaker. -
Phương án SAI: Modify the scaling policy target value to one.
Giải thích sai 🔴: Đặt target value = 1 (1 invocation/phút) sẽ khiến hệ thống luôn scale up sớm, dẫn đến chi phí cao không cần thiết (instance chạy 24/7 dù traffic thấp). Không giải quyết cold start mà còn lãng phí, vì min capacity vẫn có thể về 0 nếu policy cho phép. -
Phương án ĐÚNG (như trên): Apply a step scaling policy that scales based on an Amazon CloudWatch alarm. Apply a second CloudWatch alarm and scaling policy to scale the minimum number of instances from zero to one at the start of each business day.
Giải thích đúng 🟢: (Đã phân tích ở phần đáp án đúng). Linh hoạt, chi phí tối ưu, phù hợp best practice.
📚 Tài liệu tham khảo (AWS cập nhật 2026)
- SageMaker Endpoint Auto Scaling: https://docs.aws.amazon.com/sagemaker/latest/dg/endpoint-auto-scaling.html (hỗ trợ Step Scaling + Min Capacity).
- Application Auto Scaling với CloudWatch: https://docs.aws.amazon.com/autoscaling/application/userguide/what-is-application-auto-scaling.html.
- Scheduled Scaling: https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/ConfigureScheduledScaling.html.
- Exam Tip DOP-C02: Kết hợp alarms cho predictable workloads (AWS re:Post & Practice Exams 2025-2026).
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần thêm ví dụ code CloudFormation, hãy hỏi nhé!
The ML engineer discovers that the training script is underutilizing GPU resources. The ML engineer must identify the point in the training script where resource utilization can be optimized.
Which solution will meet this requirement?
- A Use Amazon CloudWatch metrics to create a report that describes GPU utilization over time.
- B Add SageMaker Profiler annotations to the training script. Run the script and generate a report from the results.
- C Use AWS CloudTrail to create a report that describes GPU utilization and GPU memory utilization over time.
- D Create a default monitor in Amazon SageMaker Model Monitor and suggest a baseline. Generate a report based on the constraints and statistics the monitor generates.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh tình huống một kỹ sư ML đang sử dụng Amazon SageMaker Studio notebook để huấn luyện mạng nơ-ron bằng cách tạo một estimator. Estimator này chạy một script huấn luyện Python sử dụng Distributed Data Parallel (DDP) trên một instance duy nhất có nhiều GPU hơn một.
🔍 Vấn đề chính: Script huấn luyện đang không sử dụng hết tài nguyên GPU (underutilizing GPU resources). Kỹ sư cần xác định chính xác điểm (point) trong script huấn luyện nơi có thể tối ưu hóa tài nguyên để cải thiện hiệu suất.
📌 Yêu cầu giải pháp: Phải là cách xác định điểm cụ thể trong code gây ra tình trạng lãng phí GPU, đặc biệt trong môi trường DDP trên multi-GPU single-instance. Giải pháp cần hỗ trợ phân tích chi tiết mã nguồn, không chỉ báo cáo tổng quát.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Add SageMaker Profiler annotations to the training script. Run the script and generate a report from the results.
🛠️ Lý do chi tiết:
- SageMaker Profiler là công cụ chuyên dụng của AWS để phân tích hiệu suất training job thời gian thực, đặc biệt tối ưu cho các workload multi-GPU như DDP (Distributed Data Parallel).
- Bằng cách thêm annotations (chú thích) vào script Python (ví dụ:
sm_profiler.profiler.annotate_python_function()), Profiler sẽ thu thập dữ liệu chi tiết về CPU/GPU utilization, memory, I/O tại từng hàm cụ thể trong code. - Sau khi chạy job, generate ProfilerReport sẽ hiển thị biểu đồ tương tác (System Metrics, Python Profiler, Framework Profiler) giúp pinpoint chính xác điểm bottleneck (ví dụ: hàm nào gây idle GPU do data loading chậm hoặc synchronization issue trong DDP).
- Đây là giải pháp chuẩn và cập nhật nhất (SageMaker Profiler hỗ trợ đầy đủ đến phiên bản 2026, tích hợp seamless với estimator và Studio notebooks).
- Không cần config phức tạp, chỉ thêm vài dòng code vào script là có report tự động trong SageMaker console.
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ Use Amazon CloudWatch metrics to create a report that describes GPU utilization over time.
Sai vì: CloudWatch cung cấp metrics tổng quát nhưGPUUtilizationhoặcGPUMemoryUtilizationtheo thời gian (qua dashboard hoặc alarms), nhưng không xác định được điểm cụ thể trong script gây underutilization. Nó chỉ cho biết "GPU idle bao nhiêu %", không drill-down vào code (ví dụ: hàm nào trong DDP gây bottleneck). Phù hợp monitor tổng thể, không phải debug chi tiết. -
✅ Add SageMaker Profiler annotations to the training script. Run the script and generate a report from the results.
Đúng vì: Như giải thích ở trên, đây là giải pháp chính xác nhất để annotate và profile từng phần code, hỗ trợ DDP/multi-GPU, generate report với insights actionable (cập nhật đầy đủ đến 2026). -
❌ Use AWS CloudTrail to create a report that describes GPU utilization and GPU memory utilization over time.
Sai vì: CloudTrail là dịch vụ audit log API calls (ghi lại actions như create training job), không thu thập metrics performance như GPU utilization. Không có dữ liệu về resource usage trong script, chỉ log management events – hoàn toàn không liên quan đến việc debug training efficiency. -
❌ Create a default monitor in Amazon SageMaker Model Monitor and suggest a baseline. Generate a report based on the constraints and statistics the monitor generates.
Sai vì: SageMaker Model Monitor dùng để monitor chất lượng model sau training (drift, bias, data quality trên endpoints), không phải performance training job. Nó tập trung vào inference metrics (như accuracy), không profile GPU trong script huấn luyện hoặc DDP – không giúp identify điểm optimize code.
📘 Tài liệu tham khảo (cập nhật mới nhất đến 2026)
- AWS SageMaker Profiler Documentation: https://docs.aws.amazon.com/sagemaker/latest/dg/profiler.html – Hướng dẫn chi tiết annotations và reports cho multi-GPU DDP.
- SageMaker Training Best Practices: https://docs.aws.amazon.com/sagemaker/latest/dg/train-best-practices.html – Nhấn mạnh Profiler cho GPU optimization.
- AWS re:Post & Blogs: Các case study về DDP profiling (tìm "SageMaker Profiler DDP" trên AWS console docs, cập nhật Q1/2026).
🎯 Kết luận: Sử dụng SageMaker Profiler là cách hiệu quả nhất để giải quyết underutilization GPU trong DDP, giúp tiết kiệm thời gian debug! Nếu cần code sample, hãy hỏi thêm nhé! 🚀
The model must process each image and predict the cost of the object in the image. The model also must notify the user when processing is complete.
Which solution will meet these requirements?
- A Store the images in an Amazon S3 bucket. Deploy the model on SageMaker AI. Use batch transform jobs for model inference. Use an Amazon Simple Queue Service (Amazon SQS) queue to notify users.
- B Store the images in an Amazon S3 bucket. Deploy the model on SageMaker AI. Use an asynchronous inference strategy for model inference. Use an Amazon Simple Notification Service (Amazon SNS) topic to notify users.
- C Store the images in an Amazon Elastic File System (Amazon EFS) file system. Deploy the model on SageMaker AI. Use batch transform jobs for model inference. Use an Amazon Simple Queue Service (Amazon SQS) queue to notify users.
- D Store the images in an Amazon Elastic File System (Amazon EFS) file system. Deploy the model on SageMaker AI. Use an asynchronous inference strategy for model inference. Use an Amazon Simple Notification Service (Amazon SNS) topic to notify users.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi mô tả một công ty đang phát triển công cụ ước tính chi phí nội bộ sử dụng mô hình ML trên Amazon SageMaker AI. Người dùng sẽ upload các ảnh độ phân giải cao (high-resolution images, thường có kích thước lớn). Mô hình cần:
- Xử lý từng ảnh riêng lẻ để dự đoán chi phí của vật thể trong ảnh.
- Thông báo cho người dùng khi quá trình xử lý hoàn tất.
Yêu cầu giải pháp phải bao gồm: lưu trữ ảnh, deploy mô hình trên SageMaker, chiến lược inference phù hợp (xử lý dự đoán), và cơ chế thông báo.
🔑 Thách thức chính:
- Ảnh lớn → cần xử lý không đồng bộ (asynchronous) để tránh timeout (SageMaker real-time endpoints chỉ hỗ trợ payload <6MB).
- Thông báo hoàn thành → cần dịch vụ như SNS/SQS tích hợp với SageMaker.
- Theo kiến thức AWS cập nhật đến 2026 (SageMaker version mới nhất hỗ trợ Serverless Inference và Async Inference với notification built-in), giải pháp phải tối ưu chi phí, scalable và hỗ trợ S3 làm input/output storage cho large payloads.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng:
Store the images in an Amazon S3 bucket. Deploy the model on SageMaker AI. Use an asynchronous inference strategy for model inference. Use an Amazon Simple Notification Service (Amazon SNS) topic to notify users.
Lý do 🛠️:
- S3 bucket: Lý tưởng lưu ảnh lớn, tích hợp trực tiếp với SageMaker async inference (input từ S3, output lưu S3).
- Asynchronous inference: Phù hợp xử lý từng ảnh riêng lẻ với payload lớn/long-running jobs. SageMaker tự động queue requests, xử lý async, và notify qua SNS khi output sẵn sàng (built-in feature từ 2021, cập nhật 2026 với serverless support).
- SNS topic: Hoàn hảo cho push notification đến user (email/SMS/mobile), scalable và low-latency. SageMaker async endpoint config
NotificationConfigvới SNS để thông báo output location. - Giải pháp này meet all requirements: Per-image processing, notification, cost-effective (pay-per-inference).
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm giải thích cụ thể:
-
❌ Store the images in an Amazon S3 bucket. Deploy the model on SageMaker AI. Use batch transform jobs for model inference. Use an Amazon Simple Queue Service (Amazon SQS) queue to notify users.
Sai vì: Batch transform chỉ phù hợp batch lớn (nhiều ảnh cùng lúc), không phải per-image với notification realtime. Không có built-in SNS/SQS notification (phải custom Lambda trigger). SQS là pull-based, kém hiệu quả cho user notification so với SNS push. Không async per-request như yêu cầu. -
✅ Store the images in an Amazon S3 bucket. Deploy the model on SageMaker AI. Use an asynchronous inference strategy for model inference. Use an Amazon Simple Notification Service (Amazon SNS) topic to notify users.
Đúng vì: Như giải thích ở trên. Async inference xử lý large payloads (ảnh high-res), auto-notify qua SNS khi complete (configMaxConcurrentInvocationsPerInstance,OutputConfig). S3 perfect integration. Đây là best practice AWS cho use case này. -
❌ Store the images in an Amazon Elastic File System (Amazon EFS) file system. Deploy the model on SageMaker AI. Use batch transform jobs for model inference. Use an Amazon Simple Queue Service (Amazon SQS) queue to notify users.
Sai vì: EFS không phù hợp lưu ảnh lớn (expensive, throughput limit cho high-res images; SageMaker batch transform ưu tiên S3). Batch transform + SQS thiếu notification tự động. EFS chỉ dùng cho multi-container/shared FS, không cần thiết ở đây. -
❌ Store the images in an Amazon Elastic File System (Amazon EFS) file system. Deploy the model on SageMaker AI. Use an asynchronous inference strategy for model inference. Use an Amazon Simple Notification Service (Amazon SNS) topic to notify users.
Sai vì: Async inference hỗ trợ S3 input/output (EFS không được recommend cho async endpoints do access pattern). SageMaker async yêu cầu S3 URIs cho input (s3://bucket/key), EFS phức tạp hóa và tăng chi phí không cần thiết.
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- Amazon SageMaker Asynchronous Inference: docs.aws.amazon.com/sagemaker/latest/dg/async-inference.html – Chi tiết
NotificationConfigvới SNS. - SageMaker Inference Modes: docs.aws.amazon.com/sagemaker/latest/dg/inference-modes-normal-async.html – So sánh async vs batch.
- Serverless Inference (new 2025+): aws.amazon.com/blogs/machine-learning/amazon-sagemaker-serverless-inference-updates/ – Hỗ trợ async cho large payloads.
- AWS Well-Architected ML Lens: Khuyến nghị S3 + async cho image processing.
Giải pháp này đảm bảo DevOps best practices: Scalable, observable (CloudWatch metrics), và secure (IAM roles cho S3/SNS). 🚀
The company trains and deploys a new model as a shadow variant for testing on live traffic from hospitals. The company monitors the performance of the new model for a month. During the month of testing, the shadow variant has a higher recall than the existing model but has a lower precision.
What should the company do next?
- A Promote the shadow variant to full production.
- B Extend the shadow testing period to capture more data. Monitor the new model to determine whether precision improves.
- C Use a blue/green deployment strategy to allocate a small percentage of traffic to the shadow variant to reduce model errors.
- D Disable the shadow variant and roll back to the main variant.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi này thuộc chủ đề Amazon SageMaker trong AWS, tập trung vào việc quản lý và triển khai mô hình machine learning (ML) endpoint với shadow variants (biến thể bóng) để testing trên traffic thực tế.
-
Bối cảnh: Một công ty y tế sử dụng SageMaker endpoint để host model dự đoán rủi ro bệnh nhân tái nhập viện (patient readmission risk). Họ ưu tiên độ chính xác cao (high accuracy) nhưng chịu được false positives (dự đoán sai rằng bệnh nhân sẽ tái nhập viện, tức chấp nhận cảnh báo thừa thay vì bỏ sót). Model hiện tại đã giảm hiệu suất (degraded) trong năm qua.
-
Hành động đã thực hiện: Công ty train và deploy model mới dưới dạng shadow variant (chạy song song, nhận traffic thực từ bệnh viện để test mà không ảnh hưởng production variant chính). Sau 1 tháng monitoring, shadow variant có recall cao hơn (tỷ lệ phát hiện đúng các trường hợp tái nhập viện thực sự tốt hơn, ít bỏ sót hơn) nhưng precision thấp hơn (nhiều false positives hơn).
-
Vấn đề cần quyết định: Bước tiếp theo là gì? (What should the company do next?)
🛠️ Shadow testing trong SageMaker cho phép so sánh performance mà không rủi ro, và nếu model mới tốt hơn theo metrics ưu tiên (ở đây là recall cao vì tolerate FP), thì promote lên production là bước logic tiếp theo. Đây là best practice trong SageMaker Production Variants (cập nhật đến 2026, SageMaker hỗ trợ shadow mode với auto-scaling và monitoring qua CloudWatch).
📘 Tài liệu tham khảo:
- AWS SageMaker Documentation: Production variants và Shadow testing (phiên bản mới nhất 2026 hỗ trợ enhanced monitoring với SageMaker Model Monitor).
- AWS Well-Architected Framework: ML Lens - Deploying models safely.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Promote the shadow variant to full production.
🧩 Lý do:
- Công ty tolerate false positives → Ưu tiên recall cao (giảm false negatives, tránh bỏ sót bệnh nhân cần can thiệp). Shadow variant đã chứng minh higher recall sau 1 tháng test thực tế, vượt trội model cũ (đã degrade). Precision thấp hơn là chấp nhận được vì FP không gây hại nghiêm trọng (chỉ cảnh báo thừa).
- Shadow testing chính là để validate trước khi promote. Sau testing thành công, promote variant là bước chuẩn theo SageMaker (chuyển shadow thành production variant chính, scale traffic 100%). Không cần test thêm vì metrics đã khớp yêu cầu.
- Theo best practice AWS (2026), promote đảm bảo high availability với zero-downtime deployment qua variant weights.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên metrics, yêu cầu business và SageMaker workflow:
-
Promote the shadow variant to full production.
✅ Đúng. Như giải thích trên, higher recall khớp ưu tiên (tolerate FP), test 1 tháng đủ dữ liệu live traffic. Promote là hành động tiếp theo chuẩn, chuyển shadow thành production mà không gián đoạn (variant management tự động scale). -
Extend the shadow testing period to capture more data. Monitor the new model to determine whether precision improves.
❌ Sai. Precision thấp không phải vấn đề (vì tolerate FP), và recall đã cao hơn – mục tiêu chính đạt được. Test thêm 1 tháng đã đủ (live traffic từ hospitals), kéo dài chỉ tốn kém không cần thiết. SageMaker Model Monitor đã cung cấp data đầy đủ; không nên chờ precision cải thiện vì không phải KPI chính. -
Use a blue/green deployment strategy to allocate a small percentage of traffic to the shadow variant to reduce model errors.
❌ Sai. Shadow variant đã nhận full live traffic để test (không ảnh hưởng production), nên không cần blue/green (strategy cho canary release với % traffic nhỏ). Blue/green phù hợp deploy mới hoàn toàn, nhưng ở đây shadow đã validate; allocate nhỏ % sẽ giảm hiệu quả test và không giải quyết degrade của model cũ. -
Disable the shadow variant and roll back to the main variant.
❌ Sai. Model mới tốt hơn về recall (ưu tiên hàng đầu), model cũ đã degrade → Rollback sẽ giữ performance kém. Disable shadow là hành động tiêu cực, bỏ qua lợi ích test (higher recall). Không khớp business need: high accuracy với tolerate FP.
🛠️ Lời khuyên DevOps: Sử dụng SageMaker Pipelines + CloudWatch alarms để automate promote khi recall > threshold. Test A/B variants nếu cần precision sau. Theo DOP-C02 exam (2026), shadow → promote là pattern chuẩn cho ML ops! 🚀
An ML engineer needs to configure auto scaling for the SageMaker endpoints to respond rapidly to traffic changes. The solution must use target tracking scaling policies.
Which configuration will be MOST responsive to sudden changes in traffic?
- A Configure auto scaling based on the SageMaker AI InvocationsPerInstance standard metric. Configure 10-second interval resolution, and set the default 300-second scale-in cooldown period.
- B Configure auto scaling based on the SageMaker AI InvocationsPerInstance metric. Configure high-resolution 10-second intervals, and set a 600-second scale-in cooldown period.
- C Configure auto scaling based on the SageMaker InvocationsPerInstance standard metric. Configure 10-second intervals resolution, and set a 600-second scale-in cooldown period.
- D Configure auto scaling based on the SageMaker InvocationsPerInstance metric. Configure high-resolution 10-second intervals, and set the default 300-second scale-in cooldown period.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc cấu hình auto scaling cho Amazon SageMaker real-time endpoints sử dụng target tracking scaling policies để xử lý tình huống traffic tăng đột ngột (như website của hãng hàng không có lượng truy cập cao, dẫn đến nhu cầu điều chỉnh giá vé nhanh chóng). Vấn đề trước đây là model không scale kịp thời, gây chậm trễ.
Yêu cầu chính: Tìm cấu hình PHẢN ỨNG NHANH NHẤT (MOST responsive) với sự thay đổi đột ngột của traffic. Các yếu tố cần tối ưu bao gồm:
- Metric phù hợp: Phải dùng metric chuẩn của SageMaker để đo invocations trên mỗi instance.
- Khoảng thời gian kiểm tra (intervals): Ngắn nhất có thể (10 giây với high-resolution) để phát hiện thay đổi nhanh.
- Scale-in cooldown: Thời gian chờ trước khi scale down sau khi traffic giảm; giá trị mặc định 300 giây linh hoạt hơn 600 giây cho việc điều chỉnh kịp thời.
Kiến thức cập nhật (AWS 2026): SageMaker hỗ trợ Application Auto Scaling với high-resolution CloudWatch alarms (10 giây) cho endpoints, giúp scale-out nhanh chóng dựa trên metric InvocationsPerInstance. Scale-out cooldown mặc định 60 giây, scale-in mặc định 300 giây. (Không thay đổi lớn từ 2023-2026).
✅ Đáp án đúng: Lựa chọn D
Configure auto scaling based on the SageMaker InvocationsPerInstance metric. Configure high-resolution 10-second intervals, and set the default 300-second scale-in cooldown period.
Lý do chọn:
- Metric InvocationsPerInstance là chuẩn chính xác (không có tiền tố "AI"), đo số lượng invocations/inference requests trên mỗi instance – lý tưởng cho traffic đột ngột.
- High-resolution 10-second intervals cho phép CloudWatch đánh giá metric chỉ trong 10 giây, phát hiện và scale-out cực nhanh (nhanh hơn 1 phút chuẩn).
- Default 300-second scale-in cooldown cân bằng, tránh scale down quá sớm sau peak nhưng vẫn linh hoạt hơn 600 giây, giúp hệ thống responsive với thay đổi liên tục. 🛠️ Kết hợp này đảm bảo scale-out chỉ trong ~60-70 giây sau traffic spike, phù hợp nhất cho real-time pricing.
📋 Phân tích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc:
-
❌ Lựa chọn A: Configure auto scaling based on the SageMaker AI InvocationsPerInstance standard metric. Configure 10-second interval resolution, and set the default 300-second scale-in cooldown period.
Sai vì metric không tồn tại ("AI InvocationsPerInstance" – AWS chỉ dùng "InvocationsPerInstance", tiền tố "AI" không chuẩn). "10-second interval resolution" không chỉ rõ high-resolution, nên mặc định dùng 60 giây CloudWatch, chậm hơn. Cooldown 300 giây tốt nhưng metric sai làm toàn bộ policy thất bại. -
❌ Lựa chọn B: Configure auto scaling based on the SageMaker AI InvocationsPerInstance metric. Configure high-resolution 10-second intervals, and set a 600-second scale-in cooldown period.
Sai vì metric sai (có "AI", không phải metric chuẩn của SageMaker). High-resolution 10 giây tốt cho detect nhanh, nhưng scale-in cooldown 600 giây quá dài – giữ instances lâu sau peak, làm hệ thống kém responsive với traffic dao động nhanh (khó scale up/down kịp). -
❌ Lựa chọn C: Configure auto scaling based on the SageMaker InvocationsPerInstance standard metric. Configure 10-second intervals resolution, and set a 600-second scale-in cooldown period.
Metric đúng ("InvocationsPerInstance"), nhưng "10-second intervals resolution" không phải high-resolution (chỉ là interval thường, CloudWatch vẫn dùng 60 giây đánh giá, chậm detect spike). Cooldown 600 giây làm scale-in chậm, không responsive với sudden changes. -
✅ Lựa chọn D: Configure auto scaling based on the SageMaker InvocationsPerInstance metric. Configure high-resolution 10-second intervals, and set the default 300-second scale-in cooldown period.
Hoàn hảo: Metric chuẩn, high-resolution 10 giây detect siêu nhanh, cooldown mặc định 300 giây cân bằng. Đây là config tối ưu theo best practices AWS cho high-traffic endpoints.
📘 Tài liệu tham khảo
- AWS SageMaker Docs (Endpoint Auto Scaling): https://docs.aws.amazon.com/sagemaker/latest/dg/endpoint-auto-scaling.html (Cập nhật 2024-2026: Nhấn mạnh high-res 10s alarms cho InvocationsPerInstance).
- Application Auto Scaling: https://docs.aws.amazon.com/autoscaling/application/userguide/target-tracking-scaling-policy.html (High-res metrics & cooldowns).
- CloudWatch High-Resolution Alarms: https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/creating-high-res-alarm.html.
🛠️ Lời khuyên thực tế: Test config qua AWS Console > SageMaker > Endpoints > Scaling, monitor metric InvocationsPerInstance để verify!