Ngân hàng đề — AWS Certified Machine Learning Specialty

Tìm thấy 371 câu.

Câu 111
A Data Scientist received a set of insurance records, each consisting of a record ID, the final outcome among 200 categories, and the date of the final outcome.
Some partial information on claim contents is also provided, but only for a few of the 200 categories. For each outcome category, there are hundreds of records distributed over the past 3 years. The Data Scientist wants to predict how many claims to expect in each category from month to month, a few months in advance.
What type of machine learning model should be used?
  1. A Classification month-to-month using supervised learning of the 200 categories based on claim contents.
  2. B Reinforcement learning using claim IDs and timestamps where the agent will identify how many claims in each category to expect from month to month.
  3. C Forecasting using claim IDs and timestamps to identify how many claims in each category to expect from month to month.
  4. D Classification with supervised learning of the categories for which partial information on claim contents is provided, and forecasting using claim IDs and timestamps for all other categories.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong lĩnh vực dữ liệu bảo hiểm trên AWS, nơi một Data Scientist nhận bộ dữ liệu gồm các hồ sơ bảo hiểm (insurance records). Mỗi hồ sơ bao gồm:

  • Record ID: Mã định danh duy nhất cho hồ sơ.
  • Final outcome: Kết quả cuối cùng thuộc 1 trong 200 categories (nhãn phân loại).
  • Date of the final outcome: Ngày xảy ra kết quả.
  • Partial information on claim contents: Thông tin một phần về nội dung khiếu nại, nhưng chỉ có cho một vài categories (không đầy đủ cho tất cả).

Dữ liệu có quy mô: Mỗi category có hàng trăm records phân bố qua 3 năm qua.
Mục tiêu: Dự đoán số lượng claims (hồ sơ) mong đợi cho mỗi category, theo tháng (month-to-month), với khả năng dự báo vài tháng trước (forecasting ahead).

🛠️ Vấn đề cốt lõi: Đây là bài toán dự báo chuỗi thời gian (time series forecasting) đa chuỗi (multi-series forecasting) cho 200 series riêng biệt (mỗi category là một series số lượng theo thời gian). Dữ liệu chính là timestamps (ngày/tháng) và claim IDs để đếm số lượng. Không phải dự đoán category cho hồ sơ mới (classification), vì không có features mới để classify, mà là dự báo số lượng tương lai dựa trên lịch sử thời gian.
Trên AWS, điều này phù hợp với Amazon Forecast (dịch vụ chuyên forecasting không cần ML engineer) hoặc SageMaker Forecasting algorithms như DeepAR+, Temporal Fusion Transformer (TFT) – cập nhật mới nhất đến 2026 hỗ trợ multi-horizon forecasting và cold-start handling cho categories ít dữ liệu.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Forecasting using claim IDs and timestamps to identify how many claims in each category to expect from month to month.

Lý do:

  • Bài toán yêu cầu dự báo số lượng (count) theo thời gian (số claims/category/tháng, vài tháng trước) → Đây là time series forecasting kinh điển, sử dụng claim IDs để aggregate count và timestamps làm trục thời gian.
  • Với 200 categories, mỗi cái là một related time series (item metadata có thể dùng category làm group). Amazon Forecast/SageMaker xử lý hoàn hảo multi-series forecasting, không cần features chi tiết (partial info không bắt buộc).
  • ✅ Ưu điểm: Xử lý seasonality (mùa vụ bảo hiểm), trend, và noise tự động; hỗ trợ hierarchical forecasting nếu cần tổng hợp categories. Không vi phạm vì dữ liệu lịch sử đủ (hàng trăm records/3 năm).

🧩 Giải thích tất cả các phương án (đúng và sai)

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên tính phù hợp với bài toán forecasting số lượng thời gian, theo best practices AWS ML (SageMaker/Amazon Forecast đến 2026).

  • ❌ [SAI] Classification month-to-month using supervised learning of the 200 categories based on claim contents.
    Giải thích sai: Classification dùng để dự đoán nhãn category cho một instance mới (ví dụ: classify hồ sơ mới vào category nào), không phải dự báo số lượng tương lai. Dữ liệu chỉ có partial claim contents (không đầy đủ cho 200 categories), thiếu features mới hàng tháng → Không khả thi. Supervised learning cần labeled data mới, nhưng đây là predict count từ lịch sử, không classify "month-to-month".

  • ❌ [SAI] Reinforcement learning using claim IDs and timestamps where the agent will identify how many claims in each category to expect from month to month.
    Giải thích sai: Reinforcement learning (RL) dùng cho quyết định sequential với reward (như game/robot), không phù hợp forecasting số lượng cố định từ lịch sử. RL phức tạp, cần environment/action space (không có ở đây), tốn kém train (không scale cho 200 series). AWS SageMaker RL (như Coach/RLlib) dành cho optimization, không phải pure forecasting → Sai hoàn toàn.

  • ✅ [ĐÚNG] Forecasting using claim IDs and timestamps to identify how many claims in each category to expect from month to month.
    Giải thích đúng: Như đã phân tích ở phần đáp án. Sử dụng claim IDs đếm quantity, timestamps làm time index → Amazon Forecast import dữ liệu dạng target_time_series (count/category/thời gian), tự động train models như Prophet/DeepAR. Hỗ trợ month-to-month granularity và few months ahead với accuracy cao (MAE/WAPE metrics).

  • ❌ [SAI] Classification with supervised learning of the categories for which partial information on claim contents is provided, and forecasting using claim IDs and timestamps for all other categories.
    Giải thích sai: Hybrid approach phức tạp hóa không cần thiết. Classification chỉ cho partial categories vẫn sai vì mục tiêu là count forecasting, không classify. Phân tách 200 categories (vài cái classify + còn lại forecast) gây inconsistency, khó maintain pipeline trên AWS (SageMaker Processing Jobs). Best practice: Dùng pure forecasting cho tất cả, vì timestamps đủ cho mọi category.

🛠️ Khuyến nghị triển khai trên AWS: Sử dụng Amazon Forecast cho no-code (dataset import → train → predict), hoặc SageMaker Canvas/Autopilot cho low-code. DevOps: CI/CD với CodePipeline + SageMaker Pipelines để automate retrain hàng tháng. 🚀

Câu 112
A company that promotes healthy sleep patterns by providing cloud-connected devices currently hosts a sleep tracking application on AWS. The application collects device usage information from device users. The company's Data Science team is building a machine learning model to predict if and when a user will stop utilizing the company's devices. Predictions from this model are used by a downstream application that determines the best approach for contacting users.
The Data Science team is building multiple versions of the machine learning model to evaluate each version against the company's business goals. To measure long-term effectiveness, the team wants to run multiple versions of the model in parallel for long periods of time, with the ability to control the portion of inferences served by the models.
Which solution satisfies these requirements with MINIMAL effort?
  1. A Build and host multiple models in Amazon SageMaker. Create multiple Amazon SageMaker endpoints, one for each model. Programmatically control invoking different models for inference at the application layer.
  2. B Build and host multiple models in Amazon SageMaker. Create an Amazon SageMaker endpoint configuration with multiple production variants. Programmatically control the portion of the inferences served by the multiple models by updating the endpoint configuration.
  3. C Build and host multiple models in Amazon SageMaker Neo to take into account different types of medical devices. Programmatically control which model is invoked for inference based on the medical device type.
  4. D Build and host multiple models in Amazon SageMaker. Create a single endpoint that accesses multiple models. Use Amazon SageMaker batch transform to control invoking the different models through the single endpoint.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh một công ty cung cấp thiết bị theo dõi giấc ngủ kết nối cloud, đang host ứng dụng thu thập dữ liệu trên AWS. Nhóm Data Science đang xây dựng nhiều phiên bản mô hình machine learning (ML) để dự đoán khi nào người dùng ngừng sử dụng thiết bị (churn prediction). Các dự đoán này được dùng bởi ứng dụng downstream để quyết định cách liên hệ người dùng.
Yêu cầu chính:

  • Chạy nhiều phiên bản mô hình song song trong thời gian dài để đánh giá hiệu quả dài hạn so với mục tiêu kinh doanh.
  • Kiểm soát tỷ lệ inference (phần trăm dự đoán) được phục vụ bởi từng mô hình.
  • Giải pháp phải có effort tối thiểu (MINIMAL effort).

🛠️ Mục tiêu kỹ thuật: Sử dụng Amazon SageMaker để host và deploy models với khả năng A/B testing hoặc traffic splitting mà không cần quản lý phức tạp. Theo tài liệu AWS SageMaker mới nhất (2024-2026), tính năng Production Variants trong endpoint configuration cho phép deploy nhiều models trên một endpoint duy nhất, dễ dàng điều chỉnh tỷ lệ traffic qua API update.

✅ Đáp án đúng

Đáp án đúng là lựa chọn thứ hai:
Build and host multiple models in Amazon SageMaker. Create an Amazon SageMaker endpoint configuration with multiple production variants. Programmatically control the portion of the inferences served by the multiple models by updating the endpoint configuration.

Lý do chọn:

  • Phương án này sử dụng Production Variants (tính năng cốt lõi của SageMaker endpoints từ 2019 và được tối ưu hóa đến 2026), cho phép deploy nhiều models trên một endpoint duy nhất.
  • Có thể programmatically update endpoint config qua AWS SDK/CLI để thay đổi tỷ lệ traffic (ví dụ: 70% variant A, 30% variant B) mà không downtime, hỗ trợ shadow testing hoặc canary deployments.
  • Minimal effort: Chỉ cần một endpoint, tự động load balancing, tích hợp monitoring qua CloudWatch/SageMaker Model Monitor. Hoàn hảo cho long-term parallel runs và business goal evaluation.
    ✅ Ưu điểm nổi bật: Tiết kiệm chi phí, dễ scale, và phù hợp real-time inference.

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên best practices AWS SageMaker (cập nhật 2026).

  • Phương án A (SAI):
    Build and host multiple models in Amazon SageMaker. Create multiple Amazon SageMaker endpoints, one for each model. Programmatically control invoking different models for inference at the application layer.
    ❌ Lý do sai: Tạo nhiều endpoints riêng biệt đòi hỏi effort cao (quản lý nhiều tài nguyên, routing logic ở app layer phức tạp, tăng chi phí). Không hỗ trợ traffic splitting tự động, khó kiểm soát tỷ lệ inference chính xác và đồng đều cho long-term testing. Phù hợp short-term nhưng không minimal effort.

  • Phương án B (ĐÚNG):
    Build and host multiple models in Amazon SageMaker. Create an Amazon SageMaker endpoint configuration with multiple production variants. Programmatically control the portion of the inferences served by the multiple models by updating the endpoint configuration.
    ✅ Lý do đúng: Như đã giải thích ở trên, Production Variants là giải pháp tối ưu nhất của SageMaker cho multi-model parallel inference. Update config qua UpdateEndpoint API chỉ mất giây, hỗ trợ traffic distribution (InitialVariantWeight), tích hợp A/B testing. Minimal operational overhead, lý tưởng cho yêu cầu.

  • Phương án C (SAI):
    Build and host multiple models in Amazon SageMaker Neo to take into account different types of medical devices. Programmatically control which model is invoked for inference based on the medical device type.
    ❌ Lý do sai: SageMaker Neo dành cho edge deployment (optimize models cho thiết bị IoT/edge như compile ONNX/TensorFlow cho low-latency trên hardware cụ thể), không phải cloud inference parallel. Câu hỏi tập trung vào cloud-hosted app, không liên quan "medical devices" (chỉ sleep devices), và không hỗ trợ traffic splitting cho long-term evaluation.

  • Phương án D (SAI):
    Build and host multiple models in Amazon SageMaker. Create a single endpoint that accesses multiple models. Use Amazon SageMaker batch transform jobs to control invoking the different models through the single endpoint.
    ❌ Lý do sai: Batch Transform dùng cho offline/batch inference (xử lý dữ liệu lớn một lần, không real-time), không phù hợp downstream app cần predictions liên tục. Không hỗ trợ parallel long-term với traffic control; single endpoint multi-model chỉ cho hosting, không routing động qua batch.

📘 Tài liệu tham khảo

  • AWS SageMaker Documentation (2024-2026): Production Variants – Hướng dẫn deploy multi-variants và update traffic split.
  • SageMaker Best Practices: A/B Testing with Production Variants – Case study tương tự churn prediction.
  • DevOps Pro Exam Guide: Domain 4 (Automation), nhấn mạnh minimal effort deployments với SageMaker APIs.
    🛠️ Lời khuyên: Sử dụng SageMaker Studio để prototype nhanh, tích hợp Lambda/EventBridge cho auto-update variants dựa trên metrics.
Câu 113
An agricultural company is interested in using machine learning to detect specific types of weeds in a 100-acre grassland field. Currently, the company uses tractor-mounted cameras to capture multiple images of the field as 10 ֳ— 10 grids. The company also has a large training dataset that consists of annotated images of popular weed classes like broadleaf and non-broadleaf docks.
The company wants to build a weed detection model that will detect specific types of weeds and the location of each type within the field. Once the model is ready, it will be hosted on Amazon SageMaker endpoints. The model will perform real-time inferencing using the images captured by the cameras.
Which approach should a Machine Learning Specialist take to obtain accurate predictions?
  1. A Prepare the images in RecordIO format and upload them to Amazon S3. Use Amazon SageMaker to train, test, and validate the model using an image classification algorithm to categorize images into various weed classes.
  2. B Prepare the images in Apache Parquet format and upload them to Amazon S3. Use Amazon SageMaker to train, test, and validate the model using an object- detection single-shot multibox detector (SSD) algorithm.
  3. C Prepare the images in RecordIO format and upload them to Amazon S3. Use Amazon SageMaker to train, test, and validate the model using an object- detection single-shot multibox detector (SSD) algorithm.
  4. D Prepare the images in Apache Parquet format and upload them to Amazon S3. Use Amazon SageMaker to train, test, and validate the model using an image classification algorithm to categorize images into various weed classes.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một công ty nông nghiệp muốn sử dụng machine learning (ML) để phát hiện các loại cỏ dại cụ thể (như broadleaf và non-broadleaf docks) trong cánh đồng cỏ rộng 100 acre. Họ sử dụng camera gắn trên máy kéo để chụp ảnh theo lưới 10x10 grids, và đã có dataset lớn với ảnh được chú thích (annotated images). Mục tiêu là xây dựng model phát hiện loại cỏ dại và vị trí chính xác của từng loại trong ảnh (weed detection model that detects specific types of weeds and the location of each type). Model sẽ được triển khai trên Amazon SageMaker endpoints để thực hiện real-time inferencing từ ảnh camera.

Vấn đề cốt lõi: Cần chọn cách tiếp cận để đạt dự đoán chính xác (accurate predictions), tập trung vào object detection (phát hiện đối tượng và vị trí bounding box) thay vì chỉ phân loại ảnh thông thường. SageMaker hỗ trợ các built-in algorithms cho object detection như SSD (Single Shot MultiBox Detector), và format dữ liệu chuẩn cho image training là RecordIO để tối ưu hiệu suất (theo tài liệu AWS SageMaker mới nhất 2024-2026).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Prepare the images in RecordIO format and upload them to Amazon S3. Use Amazon SageMaker to train, test, and validate the model using an object-detection single-shot multibox detector (SSD) algorithm.

Lý do 🛠️:

  • Đây là nhiệm vụ object detection (phát hiện loại cỏ dại và vị trí trong ảnh), SSD là algorithm built-in của SageMaker chuyên xử lý bounding box và class labels – phù hợp hoàn hảo với annotated dataset.
  • RecordIO format là định dạng tối ưu cho image data trong SageMaker (nhỏ gọn, hiệu quả training trên GPU, hỗ trợ trực tiếp bởi SSD algorithm), giúp đạt accuracy cao và tốc độ nhanh cho real-time inferencing.
  • Quy trình chuẩn: Upload S3 → Train/validate/test trên SageMaker → Deploy endpoint.

📋 Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên tính phù hợp với yêu cầu object detection + location và best practices SageMaker (cập nhật 2026).

  • ❌ Phương án SAI: Prepare the images in RecordIO format and upload them to Amazon S3. Use Amazon SageMaker to train, test, and validate the model using an image classification algorithm to categorize images into various weed classes.
    Giải thích: RecordIO đúng format cho images, nhưng image classification chỉ phân loại toàn bộ ảnh (ví dụ: ảnh chứa broadleaf hay không), không detect vị trí cụ thể (location). Không phù hợp với yêu cầu bounding box, dẫn đến predictions kém chính xác cho real-time weed location.

  • ❌ Phương án SAI: Prepare the images in Apache Parquet format and upload them to Amazon S3. Use Amazon SageMaker to train, test, and validate the model using an object-detection single-shot multibox detector (SSD) algorithm.
    Giải thích: SSD đúng cho object detection (hỗ trợ location), nhưng Apache Parquet là format tabular (dành cho structured data như CSV), không hỗ trợ trực tiếp image pixels trong SageMaker image algorithms. Sẽ gây lỗi training hoặc hiệu suất kém; RecordIO mới là chuẩn cho SSD.

  • ✅ Phương án ĐÚNG: Prepare the images in RecordIO format and upload them to Amazon S3. Use Amazon SageMaker to train, test, and validate the model using an object-detection single-shot multibox detector (SSD) algorithm.
    Giải thích: Kết hợp hoàn hảo RecordIO (tối ưu image data cho SageMaker) + SSD (chuyên object detection với bounding boxes). Đảm bảo accuracy cao, hỗ trợ annotated dataset, và deploy real-time endpoint mượt mà.

  • ❌ Phương án SAI: Prepare the images in Apache Parquet format and upload them to Amazon S3. Use Amazon SageMaker to train, test, and validate the model using an image classification algorithm to categorize images into various weed classes.
    Giải thích: Cả hai đều sai: Parquet không phù hợp images + image classification không detect location. Kết hợp này sẽ thất bại hoàn toàn trong việc xác định vị trí cỏ dại, không đạt accurate predictions.

Kết luận 🚀: Phương án đúng tận dụng built-in capabilities của SageMaker để xử lý object detection hiệu quả, giúp công ty triển khai nhanh chóng trên cánh đồng lớn!

Câu 114
A manufacturer is operating a large number of factories with a complex supply chain relationship where unexpected downtime of a machine can cause production to stop at several factories. A data scientist wants to analyze sensor data from the factories to identify equipment in need of preemptive maintenance and then dispatch a service team to prevent unplanned downtime. The sensor readings from a single machine can include up to 200 data points including temperatures, voltages, vibrations, RPMs, and pressure readings.
To collect this sensor data, the manufacturer deployed Wi-Fi and LANs across the factories. Even though many factory locations do not have reliable or high- speed internet connectivity, the manufacturer would like to maintain near-real-time inference capabilities.
Which deployment architecture for the model will address these business requirements?
  1. A Deploy the model in Amazon SageMaker. Run sensor data through this model to predict which machines need maintenance.
  2. B Deploy the model on AWS IoT Greengrass in each factory. Run sensor data through this model to infer which machines need maintenance.
  3. C Deploy the model to an Amazon SageMaker batch transformation job. Generate inferences in a daily batch report to identify machines that need maintenance.
  4. D Deploy the model in Amazon SageMaker and use an IoT rule to write data to an Amazon DynamoDB table. Consume a DynamoDB stream from the table with an AWS Lambda function to invoke the endpoint.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một nhà sản xuất lớn với nhiều nhà máy có chuỗi cung ứng phức tạp, nơi downtime của một máy có thể làm ngừng sản xuất ở nhiều nơi. 🏭 Một data scientist muốn phân tích dữ liệu sensor (lên đến 200 điểm dữ liệu/máy như nhiệt độ, điện áp, rung động, RPM, áp suất) để dự đoán bảo trì trước (preemptive maintenance) và cử đội ngũ dịch vụ tránh downtime không kế hoạch.

Các thách thức chính:

  • Thu thập dữ liệu qua Wi-Fi/LAN tại nhà máy.
  • Nhiều địa điểm không có kết nối internet đáng tin cậy hoặc tốc độ cao. 🌐❌
  • Yêu cầu near-real-time inference (suy luận gần thời gian thực) để xử lý nhanh chóng.

Câu hỏi yêu cầu chọn kiến trúc triển khai mô hình ML phù hợp nhất với yêu cầu kinh doanh này. 🛠️ Mục tiêu là edge computing để giảm phụ thuộc internet, đảm bảo inference nhanh tại chỗ.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Deploy the model on AWS IoT Greengrass in each factory. Run sensor data through this model to infer which machines need maintenance.

Lý do:

  • AWS IoT Greengrass cho phép triển khai mô hình ML tại edge (local) trên thiết bị IoT ở mỗi nhà máy, hỗ trợ near-real-time inference mà không cần kết nối internet liên tục. 📱🔗
  • Dữ liệu sensor được xử lý ngay tại chỗ, chỉ sync kết quả lên cloud khi cần (ví dụ: khi có kết nối). Điều này giải quyết hoàn hảo vấn đề kết nối kém, tránh latency cao từ cloud.
  • Greengrass hỗ trợ SageMaker models qua ML inference tại edge (tính năng cập nhật đến 2026: Greengrass V2 với SageMaker Neo cho optimization). ⚡
  • Phù hợp với quy mô lớn (nhiều nhà máy), giảm chi phí truyền dữ liệu lớn (200 data points/máy).

📋 Phân tích tất cả các phương án (đúng/sai)

  • ✅ Deploy the model on AWS IoT Greengrass in each factory. Run sensor data through this model to infer which machines need maintenance.
    Giải thích đúng: Phương án này lý tưởng vì Greengrass chạy ML models cục bộ tại edge, hỗ trợ near-real-time mà độc lập với internet. Dữ liệu sensor xử lý ngay, chỉ publish kết quả lên AWS IoT Core khi kết nối ổn định. Hoàn hảo cho môi trường nhà máy kết nối kém! 🏭🚀

  • ❌ Deploy the model in Amazon SageMaker. Run sensor data through this model to predict which machines need maintenance.
    Giải thích sai: SageMaker là dịch vụ cloud thuần túy, yêu cầu gửi dữ liệu lên endpoint qua internet ổn định và tốc độ cao. Với kết nối kém, sẽ gây delay lớn, không đạt near-real-time. Không giải quyết vấn đề edge computing. 🌐⏳

  • ❌ Deploy the model to an Amazon SageMaker batch transformation job. Generate inferences in a daily batch report to identify machines that need maintenance.
    Giải thích sai: Batch transformation chỉ xử lý dữ liệu theo lô (batch), tạo báo cáo hàng ngày – hoàn toàn không phải near-real-time. Downtime có thể xảy ra trước khi báo cáo sẵn sàng, vi phạm yêu cầu kinh doanh khẩn cấp. 📊🕒

  • ❌ Deploy the model in Amazon SageMaker và use an IoT rule to write data to an Amazon DynamoDB table. Consume a DynamoDB stream from the table with an AWS Lambda function to invoke the endpoint.
    Giải thích sai: Quy trình này phụ thuộc hoàn toàn vào cloud: IoT rule → DynamoDB → Stream → Lambda → SageMaker endpoint. Với internet kém, dữ liệu bị ùn tắc, latency cao (thêm nhiều hop), không đảm bảo near-real-time. Quá phức tạp và không hiệu quả cho edge. 🔄⚠️

📘 Tài liệu tham khảo (cập nhật đến 2026)

  • AWS IoT Greengrass: AWS IoT Greengrass ML Inference – Hỗ trợ deploy SageMaker models tại edge.
  • Amazon SageMaker Edge: SageMaker Edge Manager – Tích hợp với Greengrass cho inference local.
  • AWS Well-Architected Framework - IoT Lens: Nhấn mạnh edge cho near-real-time với kết nối kém.
  • AWS re:Invent 2025 updates: Greengrass V2.15+ cải thiện ML runtime cho sensor-heavy workloads.

Kiến trúc này tối ưu DevOps: Auto-deploy qua Greengrass Groups, monitor qua CloudWatch, scale dễ dàng! 🛡️✨

Câu 115
A Machine Learning Specialist is designing a scalable data storage solution for Amazon SageMaker. There is an existing TensorFlow-based model implemented as a train.py script that relies on static training data that is currently stored as TFRecords.
Which method of providing training data to Amazon SageMaker would meet the business requirements with the LEAST development overhead?
  1. A Use Amazon SageMaker script mode and use train.py unchanged. Point the Amazon SageMaker training invocation to the local path of the data without reformatting the training data.
  2. B Use Amazon SageMaker script mode and use train.py unchanged. Put the TFRecord data into an Amazon S3 bucket. Point the Amazon SageMaker training invocation to the S3 bucket without reformatting the training data.
  3. C Rewrite the train.py script to add a section that converts TFRecords to protobuf and ingests the protobuf data instead of TFRecords.
  4. D Prepare the data in the format accepted by Amazon SageMaker. Use AWS Glue or AWS Lambda to reformat and store the data in an Amazon S3 bucket.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi tập trung vào việc thiết kế giải pháp lưu trữ dữ liệu scalable (có khả năng mở rộng) cho Amazon SageMaker, dành cho một Machine Learning Specialist. Có một mô hình TensorFlow hiện có dưới dạng script train.py, phụ thuộc vào dữ liệu huấn luyện static (không thay đổi) lưu trữ dưới định dạng TFRecords. Yêu cầu là chọn phương pháp cung cấp dữ liệu huấn luyện cho SageMaker với LEAST development overhead (ít nỗ lực phát triển nhất), nghĩa là giữ nguyên script train.py mà không cần chỉnh sửa lớn, đồng thời đảm bảo tính scalable (dữ liệu phải ở nơi AWS hỗ trợ mở rộng tự động như S3).

Mục tiêu chính: Sử dụng SageMaker để huấn luyện mô hình TensorFlow mà không cần reformat dữ liệu TFRecords, ưu tiên script mode để giảm thiểu thay đổi code. 📘

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng:
Use Amazon SageMaker script mode and use train.py unchanged. Put the TFRecord data into an Amazon S3 bucket. Point the Amazon SageMaker training invocation to the S3 bucket without reformatting the training data.

Lý do:

  • SageMaker script mode cho TensorFlow (phiên bản mới nhất đến 2026) hỗ trợ trực tiếp TFRecords mà không cần thay đổi train.py. Dữ liệu chỉ cần upload lên Amazon S3 (vị trí scalable chuẩn cho training jobs), và SageMaker sẽ tự động mount dữ liệu vào đường dẫn /opt/ml/input/data/training/ trong container.
  • Đây là cách least overhead: Giữ nguyên script, không reformat, chỉ cần chỉ định S3 URI trong Estimator. Phù hợp với dữ liệu static và scalable nhờ S3's durability/infinite storage. 🛠️

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, với lý do đúng/sai dựa trên tài liệu AWS SageMaker mới nhất (TensorFlow script mode v2.15+ hỗ trợ TFRecords native):

  • ❌ Phương án SAI:
    Use Amazon SageMaker script mode and use train.py unchanged. Point the Amazon SageMaker training invocation to the local path of the data without reformatting the training data.
    Giải thích: SageMaker training jobs KHÔNG hỗ trợ local path trực tiếp cho dữ liệu đầu vào. Dữ liệu phải nằm trên S3 để đảm bảo phân tán và scalable (multi-instance training). Sử dụng local path chỉ khả dụng trong local mode (SageMaker Studio/Notebook), không phải production training jobs. Điều này vi phạm yêu cầu scalable và gây lỗi invocation. 🚫

  • ✅ Phương án ĐÚNG (như đã giải thích ở trên):
    Use Amazon SageMaker script mode and use train.py unchanged. Put the TFRecord data into an Amazon S3 bucket. Point the Amazon SageMaker training invocation to the S3 bucket without reformatting the training data.
    Giải thích: Hoàn hảo vì SageMaker TensorFlow container tự xử lý TFRecords từ S3 qua input_mode='Pipe' hoặc File, script train.py truy cập trực tiếp mà không cần code thêm. Least overhead, scalable 100%. 🌟

  • ❌ Phương án SAI:
    Rewrite the train.py script to add a section that converts TFRecords to protobuf and ingests the protobuf data instead of TFRecords.
    Giải thích: Yêu cầu rewrite script (thêm code convert TFRecords sang protobuf), tạo high development overhead – trái ngược yêu cầu "least overhead". SageMaker không bắt buộc protobuf cho TensorFlow; TFRecords đã native support. Không cần thiết và tốn công. 🔄

  • ❌ Phương án SAI:
    Prepare the data in the format accepted by Amazon SageMaker. Use AWS Glue or AWS Lambda to reformat and store the data in an Amazon S3 bucket.
    Giải thích: TFRecords ĐÃ là format chấp nhận được trong SageMaker script mode cho TensorFlow, không cần reformat. Sử dụng Glue/Lambda thêm ETL layer phức tạp, tăng chi phí/độ trễ/overhead (viết job, trigger, monitor). Phù hợp hơn cho data khác (như CSV/Parquet), không phải TFRecords static. 🛑

📚 Tài liệu tham khảo

Phân tích dựa trên kinh nghiệm AWS Certified DevOps Engineer Professional & ML Specialty. Nếu cần demo code, hỏi thêm nhé! 🚀

Câu 116
The chief editor for a product catalog wants the research and development team to build a machine learning system that can be used to detect whether or not individuals in a collection of images are wearing the company's retail brand. The team has a set of training data.
Which machine learning algorithm should the researchers use that BEST meets their requirements?
  1. A Latent Dirichlet Allocation (LDA)
  2. B Recurrent neural network (RNN)
  3. C K-means
  4. D Convolutional neural network (CNN)
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này thuộc lĩnh vực Machine Learning trên AWS, cụ thể là lựa chọn thuật toán phù hợp nhất để xây dựng hệ thống phát hiện (detection) xem các cá nhân trong bộ sưu tập hình ảnh có đang mặc trang phục thuộc thương hiệu bán lẻ của công ty hay không. Nhóm nghiên cứu và phát triển (R&D) đã có bộ dữ liệu huấn luyện (training data) sẵn sàng.

📌 Yêu cầu chính:

  • Đây là nhiệm vụ phân loại hình ảnh (image classification) hoặc phát hiện đối tượng (object detection) trong computer vision, tập trung vào việc nhận diện logo/thương hiệu trên quần áo của con người trong ảnh.
  • Thuật toán phải tối ưu cho xử lý hình ảnh, tận dụng đặc trưng không gian (spatial features) như cạnh, hình dạng, texture – đặc biệt với dữ liệu hình ảnh lớn.
  • Trong AWS, nhiệm vụ này thường được triển khai qua Amazon SageMaker (với built-in algorithms hoặc custom models), Amazon Rekognition (cho custom labels), hoặc EC2 với frameworks như TensorFlow/PyTorch. Kiến thức cập nhật đến 2026: SageMaker hỗ trợ SageMaker JumpStart với pre-trained CNN models (như ResNet, EfficientNet) cho transfer learning, giúp huấn luyện nhanh trên training data.

Mục tiêu là chọn thuật toán ML tốt nhất (BEST) phù hợp với yêu cầu, ưu tiên supervised learning vì có training data (ảnh labeled: mặc brand/không mặc).

✅ Đáp án đúng: Convolutional neural network (CNN)

Lý do lựa chọn:

  • CNN là thuật toán chuyên biệt cho computer vision, vượt trội trong việc trích xuất đặc trưng hình ảnh qua các lớp convolutional layers (nhận diện pattern như logo trên quần áo).
  • Với training data, sử dụng transfer learning (fine-tune pre-trained models như ImageNet) để đạt accuracy cao nhanh chóng.
  • 🛠️ Ứng dụng AWS: Trong SageMaker, CNN được hỗ trợ qua built-in Image Classification algorithm hoặc custom PyTorch/TensorFlow, tích hợp AutoML cho tối ưu. Đến 2026, SageMaker Canvas hỗ trợ no-code CNN cho business users như chief editor.
  • Kết quả: Hệ thống chính xác cao, scalable cho production catalog.

📋 Giải thích TẤT CẢ các phương án (đúng/sai)

  • ❌ Latent Dirichlet Allocation (LDA):
    Phương án này SAI vì LDA là thuật toán unsupervised topic modeling dành cho dữ liệu văn bản (text), phân tích chủ đề từ documents (như bag-of-words). Không xử lý hình ảnh, không trích xuất visual features → Không phù hợp cho detection thương hiệu trên ảnh.

  • ❌ Recurrent neural network (RNN):
    Phương án này SAI vì RNN (và variants như LSTM/GRU) chuyên cho dữ liệu chuỗi (sequential data) như time series, speech, hoặc text generation. Mặc dù có thể dùng cho video (sequence of frames), nhưng kém hiệu quả cho hình ảnh tĩnh so với CNN, dễ gặp vanishing gradient → Không BEST cho image detection.

  • ❌ K-means:
    Phương án này SAI vì K-means là thuật toán unsupervised clustering, nhóm dữ liệu tương đồng dựa trên khoảng cách (distance metrics) mà không cần labels. Không học được đặc trưng phức tạp như logo trên quần áo, chỉ phù hợp segment đơn giản → Không tận dụng training data labeled.

  • ✅ Convolutional neural network (CNN):
    (Như đã giải thích ở trên) ĐÚNG vì lý tưởng cho image-based tasks, với khả năng hierarchical feature extraction (edges → shapes → objects). Hiệu suất cao trên AWS với GPU acceleration.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

🛠️ Lời khuyên DevOps: Deploy model qua SageMaker Endpoints với Auto Scaling, CI/CD via CodePipeline, monitoring bằng CloudWatch → Production-ready! Nếu cần code sample, hãy hỏi thêm. 🚀

Câu 117
A retail company is using Amazon Personalize to provide personalized product recommendations for its customers during a marketing campaign. The company sees a significant increase in sales of recommended items to existing customers immediately after deploying a new solution version, but these sales decrease a short time after deployment. Only historical data from before the marketing campaign is available for training.
How should a data scientist adjust the solution?
  1. A Use the event tracker in Amazon Personalize to include real-time user interactions.
  2. B Add user metadata and use the HRNN-Metadata recipe in Amazon Personalize.
  3. C Implement a new solution using the built-in factorization machines (FM) algorithm in Amazon SageMaker.
  4. D Add event type and event value fields to the interactions dataset in Amazon Personalize.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một công ty bán lẻ đang sử dụng Amazon Personalize để cung cấp các gợi ý sản phẩm cá nhân hóa cho khách hàng trong một chiến dịch marketing. Sau khi triển khai solution version mới, doanh số bán hàng của các sản phẩm được gợi ý tăng mạnh ngay lập tức đối với khách hàng hiện tại, nhưng sau đó giảm dần. Dữ liệu huấn luyện chỉ có lịch sử trước chiến dịch marketing, dẫn đến model không thể thích ứng với hành vi người dùng thay đổi (như tăng tương tác do campaign).

📌 Vấn đề cốt lõi: Model được train trên dữ liệu cũ, nên ban đầu recommend tốt (dựa trên pattern cũ), nhưng không capture được real-time interactions mới từ campaign, khiến hiệu suất giảm. Data scientist cần điều chỉnh để model học liên tục từ dữ liệu mới mà không phải retrain toàn bộ thường xuyên.

🛠️ Bối cảnh AWS Personalize (cập nhật đến 2026): Personalize hỗ trợ event tracker để ghi nhận tương tác thời gian thực, giúp cập nhật model động mà không cần retrain ngay lập tức. Điều này phù hợp với scenario có traffic cao và thay đổi hành vi đột ngột.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use the event tracker in Amazon Personalize to include real-time user interactions.

Lý do (🧩 Phân tích sâu):

  • Event tracker cho phép gửi PUT events thời gian thực (như xem sản phẩm, click, mua hàng) qua SDK/API, giúp Personalize cập nhật user history động cho từng user.
  • Model sẽ thích ứng nhanh với thay đổi hành vi từ campaign (tăng sales ban đầu rồi ổn định), cải thiện recommendation mà không phụ thuộc dữ liệu lịch sử cũ.
  • Đây là best practice cho cold start hoặc shift in user behavior, theo docs AWS mới nhất (Personalize v2+ hỗ trợ batch/real-time inference với event ingestion scale cao).

📋 Giải thích tất cả các phương án (đúng/sai)

  • ✅ Use the event tracker in Amazon Personalize to include real-time user interactions.
    Đúng vì: Như phân tích trên, đây là cách trực tiếp giải quyết thiếu real-time data. Event tracker tích hợp sẵn, scale tự động, giúp model personalize realtime mà không retrain (tiết kiệm chi phí). ✅ Hoàn hảo cho scenario traffic spike từ campaign.

  • ❌ Add user metadata and use the HRNN-Metadata recipe in Amazon Personalize.
    Sai vì: User metadata (như tuổi, giới tính) là static data, phải import trước khi train và không capture tương tác động. HRNN-Metadata recipe chỉ cải thiện accuracy với metadata có sẵn, nhưng không giải quyết vấn đề data cũ vs. real-time shift. Retrain với metadata vẫn dùng historical data trước campaign. ❌ Không liên quan đến real-time.

  • ❌ Implement a new solution using the built-in factorization machines (FM) algorithm in Amazon SageMaker.
    Sai vì: FM trong SageMaker là algorithm cơ bản cho recommendation, nhưng yêu cầu build từ đầu (data prep, train, deploy endpoint), phức tạp và tốn kém hơn Personalize (managed service). Không tận dụng solution hiện tại, vi phạm nguyên tắc "adjust solution" trong Personalize. ❌ Quá heavy-lift, không phù hợp.

  • ❌ Add event type and event value fields to the interactions dataset in Amazon Personalize.
    Sai vì: Interactions dataset là historical/batch data, thêm event type/value chỉ cải thiện dataset cũ khi retrain (mất thời gian 1-2 ngày/solution version). Không giải quyết real-time interactions từ campaign đang diễn ra. Event fields đã hỗ trợ cơ bản, nhưng vấn đề là timing (data cũ). ❌ Chỉ fix dataset static, không dynamic.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

🛠️ Lời khuyên DevOps: Tích hợp Event Tracker với Lambda/API Gateway cho high-throughput, monitor qua CloudWatch/Personalize metrics (Recommendation Latency, Coverage). Retrain periodic nếu cần! 🚀

Câu 118 Chọn nhiều đáp án
A machine learning (ML) specialist wants to secure calls to the Amazon SageMaker Service API. The specialist has configured Amazon VPC with a VPC interface endpoint for the Amazon SageMaker Service API and is attempting to secure traffic from specific sets of instances and IAM users. The VPC is configured with a single public subnet.
Which combination of steps should the ML specialist take to secure the traffic? (Choose two.)
  1. A Add a VPC endpoint policy to allow access to the IAM users.
  2. B Modify the users' IAM policy to allow access to Amazon SageMaker Service API calls only.
  3. C Modify the security group on the endpoint network interface to restrict access to the instances.
  4. D Modify the ACL on the endpoint network interface to restrict access to the instances.
  5. E Add a SageMaker Runtime VPC endpoint interface to the VPC.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS SageMaker

📘 Nội dung câu hỏi được giải thích rõ ràng:
Câu hỏi xoay quanh việc một chuyên gia Machine Learning (ML specialist) muốn bảo mật các cuộc gọi đến Amazon SageMaker Service API (các API dịch vụ chính của SageMaker như tạo training job, endpoint, model, v.v.). Họ đã cấu hình Amazon VPC với VPC interface endpoint dành riêng cho SageMaker Service API để traffic đi private mà không qua internet công khai. VPC chỉ có một public subnet duy nhất. Mục tiêu là bảo mật traffic từ các nhóm instances cụ thể và IAM users cụ thể.
Câu hỏi yêu cầu chọn TWO (2) bước kết hợp để đạt được điều này.
🛠️ Ngữ cảnh quan trọng: Interface endpoint tạo ra Elastic Network Interface (ENI) trong subnet, cho phép kiểm soát truy cập qua VPC Endpoint Policy (chính sách tài nguyên cho IAM principals) và Security Groups (kiểm soát traffic mạng từ instances). VPC public subnet không ảnh hưởng vì endpoint đảm bảo private connectivity.

✅ Đáp án đúng (Chọn TWO):
Hai bước đúng là:

  1. Add a VPC endpoint policy to allow access to the IAM users. (Cho phép IAM users truy cập qua endpoint policy).
  2. Modify the security group on the endpoint network interface to restrict access to the instances. (Sửa Security Group trên ENI của endpoint để giới hạn instances).

🔍 Lý do chọn hai đáp án đúng (theo best practices AWS mới nhất 2026):

  • VPC interface endpoint hỗ trợ endpoint policy (resource-based policy) để kiểm soát IAM users/roles nào được phép gọi API qua endpoint đó, thay vì chỉ dựa IAM policy thông thường (vì IAM policy không đủ chi tiết cho endpoint-specific access).
  • Security Group trên ENI endpoint kiểm soát traffic từ instances bằng cách allow inbound từ SG/IP của instances cụ thể, đảm bảo chỉ instances mong muốn mới kết nối được.
    Kết hợp hai bước này secure hoàn chỉnh: Endpoint policy cho authorization (IAM level), SG cho network-level restriction (instance level). Đây là cách AWS khuyến nghị trong tài liệu VPC Endpoints (cập nhật 2024-2026).

📋 Phân tích tất cả các phương án (Đúng/Sai chi tiết)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (Đúng) hoặc ❌ (Sai), kèm giải thích đầy đủ bằng tiếng Việt dựa trên docs AWS mới nhất:

  • Add a VPC endpoint policy to allow access to the IAM users.
    ✅ ĐÚNG 🏆: VPC endpoint policy là chính sách tài nguyên (JSON) gắn trực tiếp vào endpoint, cho phép chỉ định IAM users/roles cụ thể được phép thực hiện sagemaker:* actions qua endpoint này. Điều này secure traffic từ IAM users bằng cách deny-all mặc định và allow explicit, phù hợp yêu cầu "specific sets of IAM users". Không cần thay đổi IAM policy riêng.

  • Modify the users' IAM policy to allow access to Amazon SageMaker Service API calls only.
    ❌ SAI 🚫: IAM policy của users chỉ kiểm soát authorization chung đến SageMaker API (identity-based), nhưng không secure traffic qua endpoint cụ thể. Endpoint traffic vẫn cần endpoint policy riêng để kiểm soát principal access. Nếu chỉ sửa IAM policy, traffic từ users không mong muốn vẫn có thể đi qua endpoint nếu không bị chặn.

  • Modify the security group on the endpoint network interface to restrict access to the instances.
    ✅ ĐÚNG 🛡️: Interface endpoint có ENI trong public subnet, và Security Group trên ENI kiểm soát inbound traffic (allow từ SG/IP của instances cụ thể). Điều này restrict chính xác "specific sets of instances", ngăn instances ngoài không kết nối được đến SageMaker API qua endpoint.

  • Modify the ACL on the endpoint network interface to restrict access to the instances.
    ❌ SAI 🔒: Network ACL (NACL) áp dụng cho subnet level, không trực tiếp cho ENI của endpoint. Interface endpoints sử dụng Security Groups (stateful, instance-level) thay vì NACL (stateless, subnet-level). Sửa NACL trên subnet chỉ affect toàn bộ traffic subnet, không targeted cho endpoint ENI và kém hiệu quả hơn SG.

  • Add a SageMaker Runtime VPC endpoint interface to the VPC.
    ❌ SAI ⚠️: Câu hỏi chỉ về SageMaker Service API (management APIs như CreateTrainingJob). SageMaker Runtime là endpoint riêng cho inference (InvokeEndpoint), không liên quan. Thêm nó không secure Service API và tạo endpoint thừa, vi phạm nguyên tắc least privilege.

📚 Tài liệu tham khảo (AWS docs cập nhật mới nhất đến 2026)

  • VPC Interface Endpoints for SageMaker: AWS VPC Endpoints docs – Chi tiết endpoint policy và SG config.
  • SageMaker VPC Endpoints: Interface VPC Endpoints (AWS PrivateLink) – Endpoint policies và security groups.
  • Best Practices: AWS Well-Architected Framework - Security Pillar (2024 update), khuyến nghị kết hợp endpoint policy + SG cho private API access.
  • Exam Topic: DOP-C02 (DevOps Pro 2024 syllabus), phần SageMaker Networking & Security.

💡 Lời khuyên thi chứng chỉ: Tập trung vào sự khác biệt giữa IAM policy (identity), endpoint policy (resource), SG (network L4), NACL (subnet L3). Thực hành trên AWS Console để config endpoint policy JSON! 🚀

Câu 119
An e commerce company wants to launch a new cloud-based product recommendation feature for its web application. Due to data localization regulations, any sensitive data must not leave its on-premises data center, and the product recommendation model must be trained and tested using nonsensitive data only. Data transfer to the cloud must use IPsec. The web application is hosted on premises with a PostgreSQL database that contains all the data. The company wants the data to be uploaded securely to Amazon S3 each day for model retraining.
How should a machine learning specialist meet these requirements?
  1. A Create an AWS Glue job to connect to the PostgreSQL DB instance. Ingest tables without sensitive data through an AWS Site-to-Site VPN connection directly into Amazon S3.
  2. B Create an AWS Glue job to connect to the PostgreSQL DB instance. Ingest all data through an AWS Site-to-Site VPN connection into Amazon S3 while removing sensitive data using a PySpark job.
  3. C Use AWS Database Migration Service (AWS DMS) with table mapping to select PostgreSQL tables with no sensitive data through an SSL connection. Replicate data directly into Amazon S3.
  4. D Use PostgreSQL logical replication to replicate all data to PostgreSQL in Amazon EC2 through AWS Direct Connect with a VPN connection. Use AWS Glue to move data from Amazon EC2 to Amazon S3.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty thương mại điện tử (e-commerce) muốn triển khai tính năng gợi ý sản phẩm dựa trên machine learning (product recommendation) cho ứng dụng web trên cloud AWS. Các ràng buộc nghiêm ngặt bao gồm:

  • Dữ liệu nhạy cảm (sensitive data) không được phép rời khỏi data center on-premises do quy định địa phương hóa dữ liệu (data localization regulations).
  • Model ML chỉ được train và test bằng dữ liệu không nhạy cảm (nonsensitive data).
  • Chuyển dữ liệu lên cloud phải sử dụng IPsec (giao thức mã hóa an toàn).
  • Ứng dụng web chạy on-premises với PostgreSQL database chứa toàn bộ dữ liệu.
  • Yêu cầu: Upload dữ liệu an toàn lên Amazon S3 hàng ngày để retrain model ML.

Mục tiêu là thiết kế giải pháp tuân thủ tất cả ràng buộc, sử dụng dịch vụ AWS để extract dữ liệu không nhạy cảm từ PostgreSQL on-premises và đưa trực tiếp vào S3, đảm bảo an toàn qua IPsec mà không làm rò rỉ sensitive data.
📘 Tài liệu tham khảo: AWS Glue JDBC Connector (docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-connect.html#aws-glue-programming-etl-connect-jdbc), AWS Site-to-Site VPN (docs.aws.amazon.com/vpn/latest/s2svpn/VPC_VPN.html) – cập nhật đến 2026 với hỗ trợ IPsec VPN cho Glue crawlers/ETL jobs.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create an AWS Glue job to connect to the PostgreSQL DB instance. Ingest tables without sensitive data through an AWS Site-to-Site VPN connection directly into Amazon S3.

Lý do:

  • 🛠️ AWS Glue hỗ trợ kết nối JDBC trực tiếp đến PostgreSQL on-premises, crawl và ingest chỉ các bảng không chứa sensitive data (bằng cách chọn table mapping thủ công hoặc crawler rules).
  • AWS Site-to-Site VPN sử dụng IPsec để mã hóa toàn bộ traffic từ on-premises đến VPC AWS, đảm bảo an toàn và tuân thủ yêu cầu.
  • Dữ liệu được đưa trực tiếp vào S3 (dạng Parquet/CSV) hàng ngày qua Glue ETL job theo schedule (Glue triggers hoặc EventBridge).
  • ✅ Hoàn toàn tuân thủ: Không sensitive data rời on-premises, chỉ nonsensitive data dùng cho ML retraining trên S3. Đây là giải pháp tối ưu, đơn giản, serverless theo best practices AWS ML workflows (2026 updates hỗ trợ Glue 4.0 với Spark 3.5+ cho PostgreSQL).

📋 Giải thích chi tiết tất cả các phương án

  • Phương án 1: Create an AWS Glue job to connect to the PostgreSQL DB instance. Ingest tables without sensitive data through an AWS Site-to-Site VPN connection directly into Amazon S3.
    ✅ Đúng vì: AWS Glue JDBC connector kết nối an toàn qua Site-to-Site VPN (IPsec), chỉ ingest tables nonsensitive (chọn lọc tại source), trực tiếp lưu S3. Không vi phạm quy định, hiệu quả cho daily batch.

  • Phương án 2: Create an AWS Glue job to connect to the PostgreSQL DB instance. Ingest all data through an AWS Site-to-Site VPN connection into Amazon S3 while removing sensitive data using a PySpark job.
    ❌ Sai vì: Ingest tất cả data (bao gồm sensitive) từ PostgreSQL trước, dù sau đó remove bằng PySpark – điều này vẫn làm sensitive data rời on-premises tạm thời, vi phạm nghiêm ngặt "must not leave its on-premises data center". Không an toàn và không cần thiết.

  • Phương án 3: Use AWS Database Migration Service (AWS DMS) with table mapping to select PostgreSQL tables with no sensitive data through an SSL connection. Replicate data directly into Amazon S3.
    ❌ Sai vì: DMS hỗ trợ S3 sink (Parquet), nhưng chỉ dùng SSL (không phải IPsec yêu cầu). DMS replication thường ongoing/full load, không lý tưởng cho daily batch; table mapping chọn nonsensitive OK nhưng thiếu IPsec làm vi phạm. DMS không phải lựa chọn chính cho Glue/ML pipelines.

  • Phương án 4: Use PostgreSQL logical replication to replicate all data to PostgreSQL in Amazon EC2 through AWS Direct Connect with a VPN connection. Use AWS Glue to move data from Amazon EC2 to Amazon S3.
    ❌ Sai vì: Replicate tất cả data (bao gồm sensitive) lên EC2 PostgreSQL trong cloud – sensitive data rời on-premises vĩnh viễn, vi phạm hoàn toàn. Direct Connect + VPN OK cho IPsec nhưng giải pháp phức tạp, tốn kém (EC2 management), không cần thiết so với direct S3.

🧠 Kết luận: Giải pháp đúng tận dụng Glue + VPN là serverless, chi phí thấp, scalable cho ML on S3 (SageMaker integration). Khuyến nghị test với Glue Studio cho visual ETL! 📘 Nguồn bổ sung: AWS ML Specialty Exam Guide (aws.amazon.com/certification/certified-machine-learning-specialty/), Glue for PostgreSQL (2026: Glue 5.0 preview).

Câu 120 Chọn nhiều đáp án
A logistics company needs a forecast model to predict next month's inventory requirements for a single item in 10 warehouses. A machine learning specialist uses
Amazon Forecast to develop a forecast model from 3 years of monthly data. There is no missing data. The specialist selects the DeepAR+ algorithm to train a predictor. The predictor means absolute percentage error (MAPE) is much larger than the MAPE produced by the current human forecasters.
Which changes to the CreatePredictor API call could improve the MAPE? (Choose two.)
  1. A Set PerformAutoML to true.
  2. B Set ForecastHorizon to 4.
  3. C Set ForecastFrequency to W for weekly.
  4. D Set PerformHPO to true.
  5. E Set FeaturizationMethodName to filling.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh Amazon Forecast – dịch vụ machine learning của AWS chuyên dự báo chuỗi thời gian (time series forecasting). Một công ty logistics cần mô hình dự báo nhu cầu hàng tồn kho cho một mặt hàng duy nhất tại 10 kho hàng trong tháng tới. Chuyên gia ML đã sử dụng DeepAR+ algorithm để huấn luyện predictor từ 3 năm dữ liệu hàng tháng (không có dữ liệu thiếu). Tuy nhiên, MAPE (Mean Absolute Percentage Error) của predictor này cao hơn đáng kể so với dự báo thủ công của con người.

Vấn đề chính: Cần cải thiện MAPE bằng cách thay đổi các tham số trong API CreatePredictor. Chọn 2 thay đổi phù hợp.
Bối cảnh kỹ thuật:

  • Dữ liệu: Monthly (hàng tháng), 36 điểm dữ liệu (3 năm).
  • Dự báo: Next month's inventory → ForecastHorizon nên ngắn (1 tháng).
  • DeepAR+: Thuật toán deep learning cho multi-item forecasting, nhưng có thể chưa tối ưu nếu không dùng AutoML hoặc HPO.

✅ Đáp án đúng (Chọn 2)

  • Set PerformAutoML to true.
  • Set PerformHPO to true.

Lý do chọn:
Hai thay đổi này trực tiếp tối ưu hóa mô hình mà không thay đổi cấu trúc dữ liệu gốc. PerformAutoML = true cho phép Amazon Forecast tự động thử nghiệm và chọn thuật toán tốt nhất (có thể tốt hơn DeepAR+ cho dữ liệu đơn giản này). PerformHPO = true thực hiện Hyperparameter Optimization (tối ưu siêu tham số) để tinh chỉnh DeepAR+, giảm MAPE hiệu quả. Với dữ liệu sạch và ít item (1 item x 10 warehouses), các tính năng này thường cải thiện đáng kể độ chính xác.

🔍 Giải thích chi tiết tất cả các phương án (Đúng/Sai)

  • ✅ Set PerformAutoML to true.
    Đúng 🛠️: Khi đặt PerformAutoML: true, Forecast sẽ tự động đánh giá nhiều thuật toán (như DeepAR+, ETS, NPTS, Prophet) và chọn cái tốt nhất dựa trên backtest. DeepAR+ có thể không phải lựa chọn tối ưu cho dữ liệu monthly đơn giản với ít item → AutoML giúp giảm MAPE bằng cách chọn model phù hợp hơn. (Mặc định là false, phải enable thủ công).

  • ❌ Set ForecastHorizon to 4.
    Sai 📉: ForecastHorizon là số điểm dự báo phía trước (ở đây monthly → 1 tháng cần Horizon=1). Set=4 nghĩa là dự báo 4 tháng tới, làm model tập trung vào dài hạn → tăng lỗi cho dự báo ngắn hạn (next month). Điều này làm MAPE tệ hơn vì dữ liệu chỉ 36 tháng, không đủ train dài hạn tốt.

  • ❌ Set ForecastFrequency to W for weekly.
    Sai ⏰: Dữ liệu gốc là monthly (D hoặc 1M), không phải weekly (W). Thay đổi frequency sẽ làm sai lệch dữ liệu (aggregate hoặc interpolate không chính xác), dẫn đến model kém → MAPE tăng vọt. Frequency phải khớp dữ liệu đầu vào.

  • ✅ Set PerformHPO to true.
    Đúng ⚙️: PerformHPO: true kích hoạt Hyperparameter Optimization (Bayesian Optimization) để tự động tinh chỉnh hyperparameters của DeepAR+ (như learning rate, epochs). Với dữ liệu sạch, HPO giúp model hội tụ tốt hơn, giảm MAPE so với mặc định (false). Rất hiệu quả cho single-item forecasting.

  • ❌ Set FeaturizationMethodName to filling.
    Sai 🧹: FeaturizationMethodName dùng để xử lý missing values hoặc outliers. Dữ liệu không thiếu ("no missing data"), nên method "filling" (điền giá trị) không cần thiết và có thể introduce noise (ví dụ: điền bằng mean/zero) → làm model kém chính xác hơn. Các method khác như "missing" hoặc "integral" phù hợp hơn nếu cần, nhưng ở đây không nên thay đổi.

📘 Tài liệu tham khảo (Cập nhật AWS 2026)

  • AWS Forecast API Reference: CreatePredictor – Chi tiết params PerformAutoML, PerformHPO, ForecastHorizon, ForecastFrequency, FeaturizationMethod.
  • Best Practices Guide: Improving Forecast Accuracy – Khuyến nghị dùng AutoML/HPO cho dữ liệu ít phức tạp.
  • DeepAR+ Docs: Algorithms – HPO cải thiện 5-20% MAPE trung bình.
  • Release Notes 2023-2026: Không thay đổi core params; AutoML/HPO vẫn là best practice cho DOP-C02/DevOps ML workflows.

💡 Lời khuyên DevOps: Trong pipeline CI/CD (CodePipeline + SageMaker/Forecast), luôn enable AutoML/HPO qua CDK/Terraform để tự động hóa tuning, monitor MAPE qua CloudWatch!