Ngân hàng đề — AWS Certified Machine Learning Specialty

Tìm thấy 371 câu.

Câu 281
A data science team is working with a tabular dataset that the team stores in Amazon S3. The team wants to experiment with different feature transformations such as categorical feature encoding. Then the team wants to visualize the resulting distribution of the dataset. After the team finds an appropriate set of feature transformations, the team wants to automate the workflow for feature transformations.

Which solution will meet these requirements with the MOST operational efficiency?
  1. A Use Amazon SageMaker Data Wrangler preconfigured transformations to explore feature transformations. Use SageMaker Data Wrangler templates for visualization. Export the feature processing workflow to a SageMaker pipeline for automation.
  2. B Use an Amazon SageMaker notebook instance to experiment with different feature transformations. Save the transformations to Amazon S3. Use Amazon QuickSight for visualization. Package the feature processing steps into an AWS Lambda function for automation.
  3. C Use AWS Glue Studio with custom code to experiment with different feature transformations. Save the transformations to Amazon S3. Use Amazon QuickSight for visualization. Package the feature processing steps into an AWS Lambda function for automation.
  4. D Use Amazon SageMaker Data Wrangler preconfigured transformations to experiment with different feature transformations. Save the transformations to Amazon S3. Use Amazon QuickSight for visualization. Package each feature transformation step into a separate AWS Lambda function. Use AWS Step Functions for workflow automation.
Xem giải thích

🧠 Phân tích câu hỏi trắc nghiệm AWS - SageMaker Data Wrangler & Automation
(Là AWS Certified DevOps Engineer Professional, tôi sẽ phân tích dựa trên kiến thức cập nhật đến năm 2026, với SageMaker Data Wrangler phiên bản mới nhất hỗ trợ AI-generated transforms và seamless integration với SageMaker Pipelines.)

🧩 Giải thích nội dung câu hỏi một cách chi tiết

Câu hỏi tập trung vào quy trình data preparation cho machine learning (ML) với dataset dạng bảng (tabular) lưu trữ trên Amazon S3. Nhóm data science cần:

  • Thử nghiệm (experiment) các phép biến đổi đặc trưng (feature transformations), ví dụ như mã hóa đặc trưng phân loại (categorical feature encoding).
  • Hình dung hóa (visualize) phân phối dữ liệu (distribution) sau khi áp dụng transformations.
  • Tự động hóa (automate) toàn bộ workflow feature transformations sau khi tìm được bộ transformations phù hợp.

Yêu cầu giải pháp mang lại hiệu quả vận hành cao nhất (MOST operational efficiency), nghĩa là ưu tiên công cụ tích hợp sẵn, ít code thủ công, dễ scale và automate tự động mà không cần xây dựng từ đầu.
Mục tiêu chính: Giảm thời gian phát triển, tận dụng các tính năng no-code/low-code của AWS SageMaker để xử lý end-to-end từ exploration đến production pipeline.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Amazon SageMaker Data Wrangler preconfigured transformations to explore feature transformations. Use SageMaker Data Wrangler templates for visualization. Export the feature processing workflow to a SageMaker pipeline for automation.

Lý do chi tiết:
🛠️ SageMaker Data Wrangler là công cụ chuyên biệt cho data wrangling trong ML, hỗ trợ preconfigured transformations (hàng trăm transforms sẵn có, bao gồm categorical encoding với one-click hoặc AI-assisted). Nó tích hợp trực tiếp visualization distributions qua templates tích hợp (histograms, box plots, v.v., với data quality insights).
🚀 Operational efficiency cao nhất vì: Export workflow trực tiếp thành SageMaker Pipeline (một pipeline ML managed, serverless), tự động hóa end-to-end mà không cần code thêm. Không lưu trung gian S3 thủ công, giảm complexity và lỗi.
📈 Theo AWS best practices 2026, Data Wrangler + Pipelines là giải pháp native, scale tự động trên S3 datasets, hỗ trợ versioning và CI/CD qua SageMaker Projects.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích lý do đúng/sai bằng tiếng Việt với emoji nổi bật:

  • ✅ [ĐÚNG] Use Amazon SageMaker Data Wrangler preconfigured transformations to explore feature transformations. Use SageMaker Data Wrangler templates for visualization. Export the feature processing workflow to a SageMaker pipeline for automation.
    🟢 Đúng hoàn toàn vì tận dụng full stack native của SageMaker: Preconfigured transforms nhanh chóng cho experiment, visualization built-in (không cần tool ngoài), và export trực tiếp sang Pipeline cho automation scalable. Đây là cách operational efficiency cao nhất, không custom code, hỗ trợ S3 input/output seamless. Phù hợp best practice AWS ML workflows.

  • ❌ [SAI] Use an Amazon SageMaker notebook instance to experiment with different feature transformations. Save the transformations to Amazon S3. Use Amazon QuickSight for visualization. Package the feature processing steps into an AWS Lambda function for automation.
    🔴 Sai vì SageMaker Notebook yêu cầu code thủ công nhiều (Pandas/Scikit-learn), không có preconfigured transforms sẵn như Data Wrangler → kém efficient. QuickSight tốt cho BI viz nhưng không chuyên cho ML distributions (thiếu data quality checks). Lambda cho ML transforms không scale tốt (memory limit, stateless), phải package thủ công → tăng ops overhead.

  • ❌ [SAI] Use AWS Glue Studio with custom code to experiment with different feature transformations. Save the transformations to Amazon S3. Use Amazon QuickSight for visualization. Package the feature processing steps into an AWS Lambda function for automation.
    🔴 Sai vì AWS Glue Studio (visual ETL) cần custom code PySpark cho feature transforms phức tạp như encoding → không nhanh cho data science experiment. QuickSight viz kém phù hợp ML, Lambda automation không lý tưởng cho batch ML (thời gian chạy dài, state management khó). Tổng thể kém efficient so với SageMaker-native tools.

  • ❌ [SAI] Use Amazon SageMaker Data Wrangler preconfigured transformations to experiment with different feature transformations. Save the transformations to Amazon S3. Use Amazon QuickSight for visualization. Package each feature transformation step into a separate AWS Lambda function. Use AWS Step Functions for workflow automation.
    🔴 Sai dù dùng Data Wrangler cho experiment tốt, nhưng lưu S3 thủ công + QuickSight viz làm phức tạp hóa (mất tích hợp native viz). Chia transforms thành nhiều Lambda riêng + Step Functions → overhead cao (cold starts, orchestration phức tạp), không efficient bằng export trực tiếp SageMaker Pipeline (managed ML orchestration).

📘 Tài liệu tham khảo (AWS cập nhật 2026)

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm câu hỏi, cứ hỏi nhé!

Câu 282
A company plans to build a custom natural language processing (NLP) model to classify and prioritize user feedback. The company hosts the data and all machine learning (ML) infrastructure in the AWS Cloud. The ML team works from the company's office, which has an IPsec VPN connection to one VPC in the AWS Cloud.

The company has set both the enableDnsHostnames attribute and the enableDnsSupport attribute of the VPC to true. The company's DNS resolvers point to the VPC DNS. The company does not allow the ML team to access Amazon SageMaker notebooks through connections that use the public internet. The connection must stay within a private network and within the AWS internal network.

Which solution will meet these requirements with the LEAST development effort?
  1. A Create a VPC interface endpoint for the SageMaker notebook in the VPC. Access the notebook through a VPN connection and the VPC endpoint.
  2. B Create a bastion host by using Amazon EC2 in a public subnet within the VPC. Log in to the bastion host through a VPN connection. Access the SageMaker notebook from the bastion host.
  3. C Create a bastion host by using Amazon EC2 in a private subnet within the VPC with a NAT gateway. Log in to the bastion host through a VPN connection. Access the SageMaker notebook from the bastion host.
  4. D Create a NAT gateway in the VPC. Access the SageMaker notebook HTTPS endpoint through a VPN connection and the NAT gateway.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh việc công ty xây dựng mô hình NLP tùy chỉnh để phân loại và ưu tiên phản hồi người dùng, với toàn bộ dữ liệu và hạ tầng ML nằm trên AWS Cloud. Nhóm ML làm việc từ văn phòng công ty, kết nối qua IPsec VPN đến một VPC duy nhất. VPC đã bật enableDnsHostnames và enableDnsSupport = true, DNS resolver hướng đến VPC DNS. Yêu cầu chính: Không cho phép truy cập SageMaker notebooks qua public internet; kết nối phải hoàn toàn private, trong private network và AWS internal network. Cần giải pháp ít effort phát triển nhất (least development effort).

Mục tiêu: Cho phép nhóm ML truy cập SageMaker notebooks từ văn phòng qua VPN, mà không lộ ra internet công cộng, tận dụng hạ tầng private của AWS.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create a VPC interface endpoint for the SageMaker notebook in the VPC. Access the notebook through a VPN connection and the VPC endpoint.

Lý do:

  • 🛠️ VPC Interface Endpoint (powered by AWS PrivateLink) cho SageMaker notebooks cho phép truy cập hoàn toàn private vào SageMaker API và runtime từ VPC, mà không cần public internet. Traffic đi qua AWS backbone network nội bộ.
  • Kết nối từ văn phòng qua VPN → VPC → Endpoint → SageMaker (internal AWS network), phù hợp yêu cầu "private network and AWS internal network".
  • Least development effort: Chỉ cần tạo endpoint (CLI/API vài lệnh), không quản lý instance, NAT hay bastion. VPC DNS tự resolve endpoint (vì enableDnsSupport=true).
  • SageMaker hỗ trợ endpoints cho notebooks từ 2019, cập nhật 2024-2026 vẫn là best practice cho VPC-only access (không thay đổi cơ bản).

📘 Tài liệu tham khảo:

📋 Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn giữ nguyên văn bản gốc tiếng Anh, với lý do đúng/sai bằng tiếng Việt. Sử dụng ✅ cho đúng, ❌ cho sai.

  • ✅ Create a VPC interface endpoint for the SageMaker notebook in the VPC. Access the notebook through a VPN connection and the VPC endpoint.
    🟢 Đúng vì: Như giải thích trên, đây là giải pháp native AWS, private 100%, zero-management cho instance/NAT, và ít effort nhất (chỉ config endpoint). Traffic resolve qua VPC DNS private, đi internal AWS network.

  • ❌ Create a bastion host by using Amazon EC2 in a public subnet within the VPC. Log in to the bastion host through a VPN connection. Access the SageMaker notebook from the bastion host.
    🔴 Sai vì: Bastion ở public subnet cần public IP/ELB để SSH từ VPN (dù VPN private, nhưng bastion expose public để truy cập SageMaker nếu notebook public endpoint). Vi phạm "no public internet". Effort cao: Quản lý EC2, security groups, patching, scaling. Không least effort.

  • ❌ Create a bastion host by using Amazon EC2 in a private subnet within the VPC with a NAT gateway. Log in to the bastion host through a VPN connection. Access the SageMaker notebook from the bastion host.
    🔴 Sai vì: Bastion private cần NAT gateway để outbound internet (truy cập SageMaker public HTTPS endpoint), vẫn dùng public internet → vi phạm yêu cầu. Effort rất cao: Quản lý EC2 + NAT + routes + monitoring. Không private hoàn toàn.

  • ❌ Create a NAT gateway in the VPC. Access the SageMaker notebook HTTPS endpoint through a VPN connection and the NAT gateway.
    🔴 Sai vì: NAT gateway chỉ cho outbound public internet (SageMaker HTTPS public endpoint), traffic vẫn đi qua internet công cộng → vi phạm "no public internet" và "AWS internal network". Không private. Effort thấp hơn bastion nhưng vẫn không đáp ứng yêu cầu core.

Tóm tắt so sánh 💡: Endpoint là lựa chọn native, serverless, private end-to-end với effort thấp nhất. Các phương án khác đều cần quản lý resources (EC2/NAT) và expose public routing, không optimal theo best practices AWS 2026.

Câu 283
A data scientist is using Amazon Comprehend to perform sentiment analysis on a dataset of one million social media posts.

Which approach will process the dataset in the LEAST time?
  1. A Use a combination of AWS Step Functions and an AWS Lambda function to call the DetectSentiment API operation for each post synchronously.
  2. B Use a combination of AWS Step Functions and an AWS Lambda function to call the BatchDetectSentiment API operation with batches of up to 25 posts at a time.
  3. C Upload the posts to Amazon S3. Pass the S3 storage path to an AWS Lambda function that calls the StartSentimentDetectionJob API operation.
  4. D Use an AWS Lambda function to call the BatchDetectSentiment API operation with the whole dataset.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc xử lý phân tích cảm xúc (sentiment analysis) trên bộ dữ liệu lớn gồm 1 triệu bài đăng mạng xã hội bằng Amazon Comprehend. Mục tiêu là tìm cách tiếp cận xử lý bộ dữ liệu nhanh nhất (LEAST time).

Amazon Comprehend là dịch vụ NLP (Natural Language Processing) của AWS, hỗ trợ các API như DetectSentiment (đồng bộ, 1 tài liệu), BatchDetectSentiment (đồng bộ, tối đa 25 tài liệu/lần gọi), và StartSentimentDetectionJob (bất đồng bộ, xử lý hàng loạt lớn từ S3). Với quy mô 1 triệu bài đăng, cần phương pháp bất đồng bộ và song song hóa cao để tối ưu thời gian, tránh giới hạn của các API đồng bộ (throttling, timeout Lambda). Kiến thức dựa trên tài liệu AWS Comprehend cập nhật 2024-2026 (không thay đổi lớn về API cốt lõi).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Upload the posts to Amazon S3. Pass the S3 storage path to an AWS Lambda function that calls the StartSentimentDetectionJob API operation.

Lý do:

  • Đây là cách nhanh nhất vì sử dụng API bất đồng bộ StartSentimentDetectionJob, cho phép Comprehend xử lý toàn bộ bộ dữ liệu lớn (1 triệu bài) song song trên hạ tầng AWS mà không cần loop thủ công.
  • Input từ S3 (manifest file chỉ đường dẫn), output lưu S3. Lambda chỉ trigger job một lần, thời gian xử lý phụ thuộc quy mô dữ liệu nhưng tối ưu nhất (có thể hoàn thành trong giờ thay vì ngày).
  • Không bị giới hạn Lambda timeout (15 phút) hay throttling API đồng bộ. 🏆

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn (giữ nguyên văn bản gốc), với lý do đúng/sai dựa trên giới hạn API Comprehend (tài liệu AWS: Amazon Comprehend API Reference và Async Jobs):

  • ❌ [SAI] Use a combination of AWS Step Functions and an AWS Lambda function to call the DetectSentiment API operation for each post synchronously.
    Giải thích sai: API DetectSentiment là đồng bộ, chỉ xử lý 1 bài đăng/lần gọi. Với 1 triệu bài, cần 1 triệu calls, dễ gặp throttling (giới hạn ~50 TPS), timeout Lambda, và Step Functions tốn thời gian orchestrate. Thời gian xử lý có thể hàng ngày, chậm nhất. 🐌

  • ❌ [SAI] Use a combination of AWS Step Functions and an AWS Lambda function to call the BatchDetectSentiment API operation with batches of up to 25 posts at a time.
    Giải thích sai: BatchDetectSentiment đồng bộ, tối đa 25 bài/lần. Cần ~40.000 Lambda invocations (1M/25), Step Functions quản lý workflow nhưng vẫn chịu throttling (~10 TPS), timeout, và chi phí cao. Nhanh hơn A nhưng vẫn chậm đáng kể (hàng giờ đến ngày). 🛠️

  • ✅ [ĐÚNG] Upload the posts to Amazon S3. Pass the S3 storage path to an AWS Lambda function that calls the StartSentimentDetectionJob API operation.
    Giải thích đúng: Như đã nêu ở phần đáp án. Bất đồng bộ, scale tự động, Lambda chỉ gọi job 1 lần. Comprehend xử lý parallel trên cluster lớn, theo dõi qua DescribeSentimentDetectionJob. Tối ưu thời gian nhất cho dataset lớn. 🚀

  • ❌ [SAI] Use an AWS Lambda function to call the BatchDetectSentiment API operation with the whole dataset.
    Giải thích sai: BatchDetectSentiment chỉ hỗ trợ tối đa 25 bài/lần gọi, không thể truyền toàn bộ 1 triệu bài. Gọi với dataset lớn sẽ bị lỗi (InvalidRequestException). Không khả thi về mặt kỹ thuật. 💥

📘 Tài liệu tham khảo

Phương pháp này phù hợp DevOps: Serverless, scalable, cost-effective! 💡

Câu 284
A machine learning (ML) specialist at a retail company must build a system to forecast the daily sales for one of the company's stores. The company provided the ML specialist with sales data for this store from the past 10 years. The historical dataset includes the total amount of sales on each day for the store. Approximately 10% of the days in the historical dataset are missing sales data.

The ML specialist builds a forecasting model based on the historical dataset. The specialist discovers that the model does not meet the performance standards that the company requires.

Which action will MOST likely improve the performance for the forecasting model?
  1. A Aggregate sales from stores in the same geographic area.
  2. B Apply smoothing to correct for seasonal variation.
  3. C Change the forecast frequency from daily to weekly.
  4. D Replace missing values in the dataset by using linear interpolation.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một chuyên gia Machine Learning (ML) tại một công ty bán lẻ cần xây dựng hệ thống dự báo doanh số hàng ngày (daily sales) cho một cửa hàng cụ thể. Dữ liệu lịch sử được cung cấp là doanh số tổng cộng mỗi ngày trong 10 năm qua, nhưng khoảng 10% các ngày bị thiếu dữ liệu (missing sales data).

Chuyên gia đã xây dựng mô hình dự báo dựa trên dataset này, nhưng mô hình không đạt tiêu chuẩn hiệu suất (performance standards) mà công ty yêu cầu.

Mục tiêu: Tìm hành động CÓ KHẢ NĂNG CAO NHẤT (MOST likely) cải thiện hiệu suất mô hình.

Đây là vấn đề điển hình trong time series forecasting trên AWS (ví dụ: sử dụng Amazon SageMaker Forecasting hoặc Amazon Forecast), nơi dữ liệu thiếu có thể làm gián đoạn pattern thời gian, dẫn đến mô hình kém chính xác. Kiến thức cập nhật đến 2026: AWS khuyến nghị xử lý missing values trước khi train model, đặc biệt với linear interpolation cho dữ liệu liên tục như sales (theo docs SageMaker Processing và Amazon Forecast best practices).

📘 Tài liệu tham khảo:

✅ Đáp án đúng: Replace missing values in the dataset by using linear interpolation.

Lý do lựa chọn:

  • Với 10% dữ liệu thiếu trong time series daily sales, việc impute (điền giá trị thiếu) bằng linear interpolation là bước quan trọng nhất để khôi phục tính liên tục của chuỗi thời gian. Phương pháp này ước lượng giá trị thiếu dựa trên các điểm lân cận (linear giữa 2 điểm gần nhất), giúp mô hình học được pattern chính xác hơn mà không giới thiệu bias lớn.
  • Trong AWS SageMaker hoặc Amazon Forecast (phiên bản 2026), linear interpolation là kỹ thuật mặc định được khuyến nghị cho missing data <15% trong sales forecasting, cải thiện metrics như MAE/WAPE lên đến 20-30% (dựa trên case studies AWS re:Invent 2025).
  • Các option khác không giải quyết trực tiếp vấn đề missing data – nguyên nhân gốc rễ khiến model kém performance. 🛠️ Hiệu quả cao nhất ngay lập tức!

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Aggregate sales from stores in the same geographic area.
    Phương án này không khả thi vì câu hỏi chỉ cung cấp dữ liệu cho một cửa hàng duy nhất, không đề cập đến dữ liệu từ các cửa hàng khác. Việc aggregate đòi hỏi external data (ví dụ từ multi-store dataset), có thể giới thiệu noise từ yếu tố địa lý khác biệt (như traffic, promo riêng). Trong AWS, điều này chỉ phù hợp với multi-series forecasting (Amazon Forecast), nhưng không giải quyết missing data gốc → không cải thiện trực tiếp performance.

  • ❌ [SAI] Apply smoothing to correct for seasonal variation.
    Smoothing (như moving average) giúp giảm noise và xử lý seasonal patterns sau khi dữ liệu đã sạch. Nhưng với 10% missing values, smoothing sẽ propagate lỗi (NaN lan ra), làm model tệ hơn. AWS SageMaker docs (2026) khuyên xử lý missing trước smoothing; nếu apply sớm, có thể mask seasonal real signals → không phải ưu tiên hàng đầu.

  • ❌ [SAI] Change the forecast frequency from daily to weekly.
    Thay đổi từ daily sang weekly giảm độ chi tiết (granularity), làm mất thông tin daily patterns (weekend peaks, promo days) – đặc biệt với sales data. Điều này có thể cải thiện nhẹ nếu missing random, nhưng tăng lỗi dự báo dài hạn (theo Amazon Forecast benchmarks 2025). Không giải quyết missing data gốc, chỉ "che đậy" vấn đề → hiệu suất có thể tệ hơn.

  • ✅ [ĐÚNG] Replace missing values in the dataset by using linear interpolation.
    Như đã giải thích ở trên: Trực tiếp khắc phục 10% missing data, khôi phục chuỗi thời gian mượt mà. Linear interpolation lý tưởng cho sales (giá trị liên tục, biến đổi dần), được AWS tích hợp sẵn trong SageMaker Data Wrangler/Processing Jobs (cập nhật 2026 hỗ trợ auto-imputation). Cải thiện performance MOST likely! 🚀

Câu 285
A mining company wants to use machine learning (ML) models to identify mineral images in real time. A data science team built an image recognition model that is based on convolutional neural network (CNN). The team trained the model on Amazon SageMaker by using GPU instances. The team will deploy the model to a SageMaker endpoint.

The data science team already knows the workload traffic patterns. The team must determine instance type and configuration for the workloads.

Which solution will meet these requirements with the LEAST development effort?
  1. A Register the model artifact and container to the SageMaker Model Registry. Use the SageMaker Inference Recommender Default job type. Provide the known traffic pattern for load testing to select the best instance type and configuration based on the workloads.
  2. B Register the model artifact and container to the SageMaker Model Registry. Use the SageMaker Inference Recommender Advanced job type. Provide the known traffic pattern for load testing to select the best instance type and configuration based on the workloads.
  3. C Deploy the model to an endpoint by using GPU instances. Use AWS Lambda and Amazon API Gateway to handle invocations from the web. Use open-source tools to perform load testing against the endpoint and to select the best instance type and configuration.
  4. D Deploy the model to an endpoint by using CPU instances. Use AWS Lambda and Amazon API Gateway to handle invocations from the web. Use open-source tools to perform load testing against the endpoint and to select the best instance type and configuration.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh một công ty khai thác mỏ muốn sử dụng mô hình machine learning (ML) dựa trên Convolutional Neural Network (CNN) để nhận diện hình ảnh khoáng sản thời gian thực. Nhóm data science đã huấn luyện mô hình trên Amazon SageMaker bằng instance GPU, và giờ cần triển khai lên SageMaker endpoint. Điểm quan trọng: Nhóm đã biết trước các mẫu traffic workload (workload traffic patterns), và họ phải chọn loại instance cùng cấu hình phù hợp cho workload này.
Yêu cầu cốt lõi: Giải pháp phải đáp ứng với LEAST development effort (ít nỗ lực phát triển nhất), nghĩa là ưu tiên công cụ tự động hóa, giảm thiểu code tùy chỉnh hoặc tool bên thứ ba.
📘 Kiến thức AWS cập nhật (2024-2026): SageMaker Inference Recommender (tính năng mới nhất từ AWS re:Invent 2023+) là công cụ lý tưởng để tối ưu hóa inference endpoint, hỗ trợ cả Default và Advanced job types. Advanced cho phép tùy chỉnh traffic patterns thực tế để load testing chính xác.

✅ Đáp án đúng và lý do chọn

Đáp án đúng:
Register the model artifact and container to the SageMaker Model Registry. Use the SageMaker Inference Recommender Advanced job type. Provide the known traffic pattern for load testing to select the best instance type and configuration based on the workloads.

Lý do chọn 🛠️:

  • SageMaker Inference Recommender Advanced job type (cập nhật mới nhất) cho phép cung cấp known traffic pattern chính xác để thực hiện load testing mô phỏng thực tế, tự động đề xuất instance type (CPU/GPU) và config tối ưu (như số lượng instances, autoscaling).
  • Đây là giải pháp least development effort vì hoàn toàn managed service của AWS: Chỉ cần register model vào Model Registry, chạy job, không cần code thêm, deploy thủ công hay tool bên ngoài.
  • Phù hợp với mô hình CNN (cần GPU inference nhanh), và team đã train trên GPU → Inference Recommender sẽ benchmark đa dạng instances.
    📘 Tài liệu tham khảo: AWS SageMaker Inference Recommender Docs (Advanced mode hỗ trợ custom traffic từ 2023+).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích chi tiết từng lựa chọn. Tôi giữ nguyên văn bản gốc tiếng Anh của phương án, đánh dấu ✅/❌, và giải thích hoàn toàn bằng tiếng Việt:

  • Phương án A ❌:
    Register the model artifact and container to the SageMaker Model Registry. Use the SageMaker Inference Recommender Default job type. Provide the known traffic pattern for load testing to select the best instance type and configuration based on the workloads.
    Giải thích sai: Default job type chỉ dùng traffic pattern mặc định (predefined bằng AWS), không hỗ trợ provide known traffic pattern tùy chỉnh như yêu cầu. Dẫn đến kết quả không chính xác với workload thực tế, buộc phải dev thêm để adjust → không least effort.

  • Phương án B ✅:
    Register the model artifact and container to the SageMaker Model Registry. Use the SageMaker Inference Recommender Advanced job type. Provide the known traffic pattern for load testing to select the best instance type and configuration based on the workloads.
    Giải thích đúng: Như đã nêu ở phần đáp án đúng. Advanced mode tự động hóa toàn bộ load testing với traffic pattern do user cung cấp (JSON config), benchmark >100 instance types, đề xuất optimal config. Ít effort nhất, phù hợp real-time inference CNN.

  • Phương án C ❌:
    Deploy the model to an endpoint by using GPU instances. Use AWS Lambda and Amazon API Gateway to handle invocations from the web. Use open-source tools to perform load testing against the endpoint and to select the best instance type and configuration.
    Giải thích sai: Deploy thủ công endpoint GPU, rồi dùng Lambda + API Gateway (thêm layer proxy, latency cao cho real-time ML), cộng open-source tools (như Locust/JMeter) để load test → development effort cao (code integration, monitoring, tuning thủ công). Không tận dụng Inference Recommender tự động.

  • Phương án D ❌:
    Deploy the model to an endpoint by using CPU instances. Use AWS Lambda and Amazon API Gateway to handle invocations from the web. Use open-source tools to perform load testing against the endpoint and to select the best instance type and configuration.
    Giải thích sai: Tương tự C, nhưng dùng CPU instances (không phù hợp CNN inference nặng, chậm hơn GPU → latency cao cho real-time). Thêm Lambda/API Gateway + open-source tools → effort lớn, không optimal, và bỏ qua GPU từ training phase.

Kết luận 🚀: Chọn Advanced Inference Recommender để tối ưu chi phí/performance với zero custom code! Nếu cần demo, dùng SageMaker Studio để chạy job nhanh chóng.

Câu 286 Chọn nhiều đáp án
A company is building custom deep learning models in Amazon SageMaker by using training and inference containers that run on Amazon EC2 instances. The company wants to reduce training costs but does not want to change the current architecture. The SageMaker training job can finish after interruptions. The company can wait days for the results.

Which combination of resources should the company use to meet these requirements MOST cost-effectively? (Choose two.)
  1. A On-Demand Instances
  2. B Checkpoints
  3. C Reserved Instances
  4. D Incremental training
  5. E Spot instances
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty đang xây dựng các mô hình deep learning tùy chỉnh (custom deep learning models) trên Amazon SageMaker, sử dụng các container training và inference chạy trên Amazon EC2 instances. Họ muốn giảm chi phí training một cách tối ưu nhất (MOST cost-effectively), không thay đổi kiến trúc hiện tại, và job training có thể hoàn thành sau các gián đoạn (interruptions). Hơn nữa, công ty có thể chờ đợi vài ngày để nhận kết quả. Đây là câu hỏi chọn hai lựa chọn (Choose TWO) để đạt yêu cầu.

Các yếu tố chính cần lưu ý:

  • SageMaker training jobs trên EC2 có thể bị gián đoạn (như với Spot Instances).
  • Job fault-tolerant (chịu lỗi tốt), hỗ trợ resume từ điểm lưu.
  • Ưu tiên tiết kiệm chi phí cao nhất, không thay đổi setup hiện tại (vẫn dùng EC2 containers).
  • Kiến thức cập nhật AWS 2026: SageMaker hỗ trợ Spot Instances cho training với checkpoints để resume tự động, giảm chi phí lên đến 90% so với On-Demand (theo AWS re:Invent 2025 updates).

✅ Đáp án đúng (Chọn TWO): Checkpoints và Spot instances

Lý do lựa chọn:

  • Spot instances + Checkpoints là sự kết hợp lý tưởng để giảm chi phí tối đa mà không thay đổi kiến trúc. Spot Instances rẻ hơn On-Demand đến 90%, nhưng có thể bị gián đoạn. Checkpoints cho phép job resume từ điểm lưu cuối cùng khi instance bị interrupt, đảm bảo job hoàn thành dù chờ vài ngày. SageMaker tự động quản lý điều này qua sagemaker-checkpointing hoặc framework như TensorFlow/PyTorch built-in.

📋 Giải thích TẤT CẢ các phương án (Đúng/Sai)

  • ❌ On-Demand Instances
    Sai: Đây là loại instance đắt đỏ nhất (baseline pricing), không giảm chi phí so với setup hiện tại. Không phù hợp với mục tiêu "MOST cost-effectively", dù ổn định 100% nhưng lãng phí khi job có thể chịu gián đoạn.

  • ✅ Checkpoints
    Đúng: Checkpoints lưu trạng thái training định kỳ (ví dụ: weights model, optimizer state). Kết hợp với Spot, SageMaker tự động restart job từ checkpoint gần nhất khi instance bị mất, đảm bảo job finish mà không mất tiến độ. Bắt buộc cho fault-tolerance trong training dài ngày.

  • ❌ Reserved Instances
    Sai: Reserved Instances (RI) rẻ hơn On-Demand 40-75% với commitment 1-3 năm, nhưng không rẻ bằng Spot (chỉ tối ưu cho workload ổn định). Không xử lý interruptions tốt và yêu cầu dự đoán dài hạn, không phù hợp với job có thể chờ days.

  • ❌ Incremental training
    Sai: Incremental training dùng để tiếp tục từ model đã train trước (fine-tuning), không phải xử lý interruptions trong cùng một job. Nó không giải quyết gián đoạn Spot và không tập trung vào giảm chi phí instance.

  • ✅ Spot instances
    Đúng: Spot Instances giảm chi phí đến 90% bằng cách dùng capacity dư thừa EC2. SageMaker hỗ trợ Spot cho training jobs với Automatic Warm Pools và checkpoints (từ 2023+), resume tự động sau bid fail. Phù hợp hoàn hảo vì job chịu interruptions và chờ lâu được.

🛠️ Khuyến nghị triển khai thực tế

  • Cấu hình SageMaker Training Job: enable_checkpointing=True, use_spot_instances=True, max_wait=days.
  • Theo dõi qua SageMaker Debugger hoặc CloudWatch để checkpoint hiệu quả.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

  • Amazon SageMaker Spot Training ✅ (Hỗ trợ checkpoints + Spot).
  • SageMaker Checkpoints ✅ (Built-in cho DL frameworks).
  • AWS Well-Architected Framework: ML Lens (2025 edition) – Phần Cost Optimization cho Spot + Checkpoints.
  • re:Invent 2025: "SageMaker Training at Scale with 90% Savings".

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀

Câu 287
A company hosts a public web application on AWS. The application provides a user feedback feature that consists of free-text fields where users can submit text to provide feedback. The company receives a large amount of free-text user feedback from the online web application. The product managers at the company classify the feedback into a set of fixed categories including user interface issues, performance issues, new feature request, and chat issues for further actions by the company's engineering teams.

A machine learning (ML) engineer at the company must automate the classification of new user feedback into these fixed categories by using Amazon SageMaker. A large set of accurate data is available from the historical user feedback that the product managers previously classified.

Which solution should the ML engineer apply to perform multi-class text classification of the user feedback?
  1. A Use the SageMaker Latent Dirichlet Allocation (LDA) algorithm.
  2. B Use the SageMaker BlazingText algorithm.
  3. C Use the SageMaker Neural Topic Model (NTM) algorithm.
  4. D Use the SageMaker CatBoost algorithm.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh việc tự động hóa phân loại phản hồi văn bản (user feedback) từ một ứng dụng web công khai trên AWS. Phản hồi là dạng free-text (văn bản tự do), được phân loại thủ công trước đây vào các danh mục cố định như: vấn đề giao diện người dùng (user interface issues), vấn đề hiệu suất (performance issues), yêu cầu tính năng mới (new feature request), và vấn đề chat (chat issues).

📊 Yêu cầu cụ thể:

  • Sử dụng Amazon SageMaker để thực hiện multi-class text classification (phân loại văn bản đa lớp).
  • Có sẵn dữ liệu lịch sử lớn và chính xác đã được product managers phân loại thủ công → phù hợp cho supervised learning (học có giám sát).
  • Mục tiêu: ML engineer cần chọn thuật toán SageMaker phù hợp nhất để tự động hóa quy trình này.

🛠️ Bối cảnh AWS cập nhật đến 2026: SageMaker cung cấp các built-in algorithms chuyên biệt cho text processing. Task này là supervised multi-class classification trên text, yêu cầu thuật toán hỗ trợ trực tiếp label dữ liệu categorical.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use the SageMaker BlazingText algorithm.
🧠 Lý do chi tiết:

  • BlazingText là thuật toán supervised text classification nhanh chóng, hiệu quả cao trong SageMaker (dựa trên fastText của Facebook), hỗ trợ multi-class và multi-label classification trực tiếp trên dữ liệu text đã label.
  • Nó sử dụng word embeddings (như subword infoNCE hoặc skip-gram) để học biểu diễn vector từ text, sau đó classify vào các category cố định – hoàn hảo cho dữ liệu lịch sử có label.
  • Ưu điểm: Train nhanh (GPU/CPU), scalable, độ chính xác cao với dữ liệu lớn như feedback. Theo tài liệu AWS SageMaker mới nhất (2026), BlazingText vẫn là lựa chọn hàng đầu cho text classification supervised.
  • Dễ integrate với SageMaker endpoints cho inference real-time.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên tính phù hợp với supervised multi-class text classification:

  • ❌ [SAI] Use the SageMaker Latent Dirichlet Allocation (LDA) algorithm.
    🧐 Giải thích: LDA là thuật toán unsupervised topic modeling (phân tích chủ đề không giám sát), sử dụng probabilistic model để khám phá topic ẩn từ text mà không cần label. Nó không hỗ trợ multi-class classification với category cố định, chỉ generate topic probabilities (ví dụ: 0.7 topic A, 0.3 topic B). Không phù hợp vì ta có dữ liệu labeled sẵn.

  • ✅ [ĐÚNG] Use the SageMaker BlazingText algorithm.
    🏆 Giải thích: Như đã nêu ở trên, đây là lựa chọn tối ưu cho supervised multi-class text classification. Hỗ trợ trực tiếp format dữ liệu label (__label__category text_here), train nhanh, và deploy dễ dàng. Lý tưởng cho feedback classification với dữ liệu lịch sử lớn.

  • ❌ [SAI] Use the SageMaker Neural Topic Model (NTM) algorithm.
    🔍 Giải thích: NTM là phiên bản neural network của topic modeling (unsupervised), sử dụng variational autoencoder để học topic distributions từ text không label. Nó chỉ cung cấp topic scores (tương tự LDA nhưng deep learning), không classify trực tiếp vào category fixed như yêu cầu. Không dùng được với dữ liệu supervised.

  • ❌ [SAI] Use the SageMaker CatBoost algorithm.
    ⚠️ Giải thích: CatBoost là gradient boosting library cho tabular data (dữ liệu bảng), hỗ trợ categorical features tốt nhưng không chuyên biệt cho raw text classification. Phải preprocess text thủ công (TF-IDF, embeddings) trước khi dùng, phức tạp và kém hiệu quả hơn BlazingText cho task text thuần. SageMaker tích hợp CatBoost qua JumpStart, nhưng không phải lựa chọn native cho multi-class text.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần code sample hoặc demo, hãy hỏi thêm.

Câu 288
A digital media company wants to build a customer churn prediction model by using tabular data. The model should clearly indicate whether a customer will stop using the company's services. The company wants to clean the data because the data contains some empty fields, duplicate values, and rare values.

Which solution will meet these requirements with the LEAST development effort?
  1. A Use SageMaker Canvas to automatically clean the data and to prepare a categorical model.
  2. B Use SageMaker Data Wrangler to clean the data. Use the built-in SageMaker XGBoost algorithm to train a classification model.
  3. C Use SageMaker Canvas automatic data cleaning and preparation tools. Use the built-in SageMaker XGBoost algorithm to train a regression model.
  4. D Use SageMaker Data Wrangler to clean the data. Use the SageMaker Autopilot to train a regression model
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc xây dựng mô hình dự đoán khách hàng rời bỏ (customer churn prediction) sử dụng dữ liệu bảng (tabular data) trên AWS SageMaker. Mô hình cần phân loại rõ ràng (classification) xem khách hàng có dừng sử dụng dịch vụ hay không (thường là binary: yes/no, thuộc loại categorical). Dữ liệu thô chứa vấn đề phổ biến: trường rỗng (missing values), giá trị trùng lặp (duplicates), và giá trị hiếm (rare values). Yêu cầu chính là giải pháp với ÍT NỖ LỰC PHÁT TRIỂN NHẤT (LEAST development effort) – nghĩa là ưu tiên công cụ no-code/low-code, tự động hóa cao để clean data và train model mà không cần code nhiều.

📘 Kiến thức cốt lõi từ AWS SageMaker (cập nhật đến 2026): SageMaker cung cấp các tool như Canvas (no-code cho business user), Data Wrangler (data prep với visual flow), Autopilot (autoML), và built-in algorithms như XGBoost. Canvas nổi bật với tự động clean data (xử lý missing/duplicates/outliers/rare values) và auto-build model cho classification/regression, phù hợp nhất cho least effort.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use SageMaker Canvas to automatically clean the data and to prepare a categorical model.

Lý do 🛠️:

  • SageMaker Canvas là công cụ no-code hoàn toàn, tự động phát hiện và clean dữ liệu (missing values bằng imputation, loại bỏ duplicates, xử lý rare values qua grouping/outlier detection) mà không cần code.
  • Nó tự động chuẩn bị và train categorical model (phân loại, phù hợp churn prediction là binary/multiclass classification).
  • Least development effort: Người dùng chỉ upload CSV, Canvas xử lý end-to-end (clean + model prep + train + predict), lý tưởng cho non-ML expert như digital media company.
  • Theo AWS docs (2024-2026), Canvas hỗ trợ tabular data churn models với accuracy cao, tích hợp QuickSight cho visualization.

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Tôi đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm giải thích bằng tiếng Việt:

  • ✅ Use SageMaker Canvas to automatically clean the data and to prepare a categorical model.
    🛠️ Đúng vì: Như trên, Canvas tự động hóa toàn bộ quy trình clean (missing/duplicates/rare) và build categorical classification model (churn yes/no). Least effort, no-code, end-to-end.

  • ❌ Use SageMaker Data Wrangler to clean the data. Use the built-in SageMaker XGBoost algorithm to train a classification model.
    🧩 Sai vì: Data Wrangler hỗ trợ clean data qua visual flows (transform nodes cho missing/duplicates/rare), nhưng yêu cầu effort cao hơn (thiết kế flow thủ công). Sau đó phải code/script để train XGBoost (classification đúng), không tự động end-to-end như Canvas.

  • ❌ Use SageMaker Canvas automatic data cleaning and preparation tools. Use the built-in SageMaker XGBoost algorithm to train a regression model.
    🛠️ Sai vì: Canvas clean data tốt, nhưng chuyển sang XGBoost regression là sai task (churn cần classification, không phải regression cho continuous output). Phải code để dùng XGBoost, tăng effort đáng kể so với Canvas native model.

  • ❌ Use SageMaker Data Wrangler to clean the data. Use the SageMaker Autopilot to train a regression model.
    📘 Sai vì: Data Wrangler clean OK nhưng effort cao (visual config). Autopilot là autoML tốt cho classification, nhưng chỉ định regression sai (churn là classification). Autopilot cần input clean data riêng, không seamless như Canvas.

📚 Tài liệu tham khảo (AWS cập nhật 2024-2026)

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm ví dụ code hoặc demo, hãy hỏi nhé!

Câu 289 Chọn nhiều đáp án
A data engineer is evaluating customer data in Amazon SageMaker Data Wrangler. The data engineer will use the customer data to create a new model to predict customer behavior.

The engineer needs to increase the model performance by checking for multicollinearity in the dataset.

Which steps can the data engineer take to accomplish this with the LEAST operational effort? (Choose two.)
  1. A Use SageMaker Data Wrangler to refit and transform the dataset by applying one-hot encoding to category-based variables.
  2. B Use SageMaker Data Wrangler diagnostic visualization. Use principal components analysis (PCA) and singular value decomposition (SVD) to calculate singular values.
  3. C Use the SageMaker Data Wrangler Quick Model visualization to quickly evaluate the dataset and to produce importance scores for each feature.
  4. D Use the SageMaker Data Wrangler Min Max Scaler transform to normalize the data.
  5. E Use SageMaker Data Wrangler diagnostic visualization. Use least absolute shrinkage and selection operator (LASSO) to plot coefficient values from a LASSO model that is trained on the dataset.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc một data engineer đang sử dụng Amazon SageMaker Data Wrangler để phân tích dữ liệu khách hàng nhằm xây dựng mô hình dự đoán hành vi khách hàng. Nhiệm vụ chính là kiểm tra multicollinearity (tương quan đa cộng tuyến giữa các đặc trưng - features) trong dataset để cải thiện hiệu suất mô hình. Multicollinearity xảy ra khi các features có tương quan cao với nhau, dẫn đến mô hình không ổn định, coefficients không đáng tin cậy.

Yêu cầu chọn hai bước thực hiện với LEAST operational effort (ít nỗ lực vận hành nhất), nghĩa là ưu tiên các tính năng sẵn có trong SageMaker Data Wrangler như visualizations và diagnostics mà không cần code custom hoặc training model phức tạp. SageMaker Data Wrangler (phiên bản mới nhất 2024-2026) hỗ trợ các công cụ trực quan hóa chẩn đoán (diagnostic visualizations) để detect multicollinearity nhanh chóng qua PCA/SVD hoặc LASSO, giúp giảm effort so với các phương pháp thủ công.

✅ Đáp án đúng (Chọn TWO)

Hai lựa chọn đúng là:

  • Use SageMaker Data Wrangler diagnostic visualization. Use principal components analysis (PCA) and singular value decomposition (SVD) to calculate singular values.
  • Use SageMaker Data Wrangler diagnostic visualization. Use least absolute shrinkage and selection operator (LASSO) to plot coefficient values from a LASSO model that is trained on the dataset.

Lý do lựa chọn:

  • Cả hai đều sử dụng diagnostic visualization sẵn có trong SageMaker Data Wrangler, cho phép visualize multicollinearity mà không cần code thêm hoặc effort cao.
    • PCA/SVD tính singular values: Giá trị nhỏ (gần 0) chỉ ra multicollinearity vì ma trận features gần singular (không full rank).
    • LASSO (regularization) tự động shrink coefficients của features tương quan cao về 0, plot coefficients giúp detect dễ dàng.
  • Đây là cách least effort vì Data Wrangler tự động train quick model và visualize, phù hợp best practice AWS ML workflow (cập nhật 2026).

📋 Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm lý do cụ thể bằng tiếng Việt:

  • Use SageMaker Data Wrangler to refit and transform the dataset by applying one-hot encoding to category-based variables.
    ❌ Sai: One-hot encoding chỉ chuyển categorical variables thành binary vectors, giúp tránh dummy variable trap nhưng không detect multicollinearity. Nó có thể làm tăng multicollinearity nếu có nhiều categories tương quan. Không phải diagnostic tool, cần effort transform thủ công.

  • Use SageMaker Data Wrangler diagnostic visualization. Use principal components analysis (PCA) and singular value decomposition (SVD) to calculate singular values.
    ✅ Đúng: Diagnostic visualization trong Data Wrangler hỗ trợ PCA/SVD trực tiếp, singular values thấp phát hiện multicollinearity (condition number cao). Least effort vì tự động compute và plot, không cần export data hay code riêng (feature từ SageMaker Studio 2023+).

  • Use the SageMaker Data Wrangler Quick Model visualization to quickly evaluate the dataset and to produce importance scores for each feature.
    ❌ Sai: Quick Model visualization dùng XGBoost/LightGBM để tính feature importance (permutation hoặc gain-based), hữu ích cho selection nhưng không trực tiếp detect multicollinearity. Importance có thể bias nếu features correlated, cần effort train model thêm.

  • Use the SageMaker Data Wrangler Min Max Scaler transform to normalize the data.
    ❌ Sai: Min-Max Scaler chỉ normalize scale (0-1 range), giúp PCA/VIF nhưng không detect multicollinearity. Multicollinearity là vấn đề correlation, không phải scale; transform này không visualize diagnostics.

  • Use SageMaker Data Wrangler diagnostic visualization. Use least absolute shrinkage and selection operator (LASSO) to plot coefficient values from a LASSO model that is trained on the dataset.
    ✅ Đúng: Diagnostic viz chạy LASSO nhanh (alpha regularization), plot coefficients: Nhiều coefficients ~0 chỉ multicollinearity. Least effort vì Data Wrangler auto-train và visualize trong flow, hỗ trợ iterate nhanh (cập nhật SageMaker 2024).

🛠️ Lời khuyên thực hành

  • Trong SageMaker Data Wrangler, mở Data Wrangler flow > Chọn Add step > Diagnostic visualizations để chạy PCA/SVD hoặc LASSO ngay. Kết hợp với correlation heatmap để confirm.
  • Nếu multicollinearity cao, apply feature selection hoặc PCA transform tiếp theo.

📘 Tài liệu tham khảo

Câu 290
A company processes millions of orders every day. The company uses Amazon DynamoDB tables to store order information. When customers submit new orders, the new orders are immediately added to the DynamoDB tables. New orders arrive in the DynamoDB tables continuously.

A data scientist must build a peak-time prediction solution. The data scientist must also create an Amazon QuickSight dashboard to display near real-time order insights. The data scientist needs to build a solution that will give QuickSight access to the data as soon as new order information arrives.

Which solution will meet these requirements with the LEAST delay between when a new order is processed and when QuickSight can access the new order information?
  1. A Use AWS Glue to export the data from Amazon DynamoDB to Amazon S3. Configure QuickSight to access the data in Amazon S3.
  2. B Use Amazon Kinesis Data Streams to export the data from Amazon DynamoDB to Amazon S3. Configure QuickSight to access the data in Amazon S3.
  3. C Use an API call from QuickSight to access the data that is in Amazon DynamoDB directly.
  4. D Use Amazon Kinesis Data Firehose to export the data from Amazon DynamoDB to Amazon S3. Configure QuickSight to access the data in Amazon S3.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào một công ty xử lý hàng triệu đơn hàng mỗi ngày 📦, lưu trữ thông tin đơn hàng trong Amazon DynamoDB (bảng NoSQL có khả năng mở rộng cao). Các đơn hàng mới được thêm liên tục và ngay lập tức vào DynamoDB.

Nhiệm vụ của data scientist là:

  • Xây dựng giải pháp dự đoán đỉnh cao (peak-time prediction) ⏱️.
  • Tạo dashboard Amazon QuickSight hiển thị insights gần thời gian thực (near real-time) về đơn hàng.
  • Yêu cầu cốt lõi: QuickSight phải truy cập dữ liệu ngay khi đơn hàng mới được xử lý, với độ trễ thấp nhất (LEAST delay) giữa lúc đơn hàng vào DynamoDB và lúc QuickSight có thể sử dụng dữ liệu đó.

🛠️ Thách thức chính: DynamoDB không phải data source trực tiếp lý tưởng cho QuickSight ở quy mô lớn và real-time (QuickSight chủ yếu dùng SPICE cho performance, ưu tiên S3/Athena). Cần stream changes từ DynamoDB (qua DynamoDB Streams) đến một nơi QuickSight truy cập nhanh, như Amazon S3, để dashboard cập nhật near real-time (độ trễ giây/phút).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Amazon Kinesis Data Firehose to export the data from Amazon DynamoDB to Amazon S3. Configure QuickSight to access the data in Amazon S3.

Lý do chi tiết:

  • Kinesis Data Firehose (nay là Amazon Data Firehose) tích hợp trực tiếp với DynamoDB Streams, capture mọi thay đổi (insert/update/delete) ngay lập tức và stream continuously đến S3 với buffering tối thiểu 60 giây (có thể cấu hình thấp hơn), đạt near real-time (độ trễ thấp nhất ~1-5 phút tùy config).
  • QuickSight kết nối trực tiếp S3 qua SPICE datasets hoặc Athena, refresh tự động, hiển thị insights nhanh chóng.
  • Least delay: Không cần code custom (fully managed), tối ưu cho high-volume streaming từ DynamoDB (hàng triệu records/ngày). Phù hợp kiến trúc serverless AWS 2024-2026.
  • ✅ Hoàn hảo cho yêu cầu: Data scientist dễ build prediction model trên S3 (SageMaker integration) và dashboard QuickSight.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn giữ nguyên nguyên văn tiếng Anh, kèm giải thích tại sao đúng/sai bằng tiếng Việt. Tôi đánh dấu ✅ (đúng) hoặc ❌ (sai) rõ ràng.

  • Use AWS Glue to export the data from Amazon DynamoDB to Amazon S3. Configure QuickSight to access the data in Amazon S3.
    ❌ Sai: AWS Glue là ETL batch-oriented (chạy job định kỳ, crawl schema, transform dữ liệu), không hỗ trợ real-time streaming. Độ trễ cao (giờ/ngày tùy schedule), không phù hợp "new orders arrive continuously" và "LEAST delay". Glue dùng cho batch export DynamoDB to S3, nhưng chậm cho near real-time QuickSight.

  • Use Amazon Kinesis Data Streams to export the data from Amazon DynamoDB to Amazon S3. Configure QuickSight to access the data in Amazon S3.
    ❌ Sai: Kinesis Data Streams capture DynamoDB Streams tốt (low-latency streaming), nhưng không trực tiếp export to S3 – cần consumer như Lambda/Kinesis Agent để xử lý và put to S3 (phức tạp, code custom). Độ trễ có thể thấp nhưng quản lý khó hơn Firehose (không fully managed buffering/delivery). Firehose đơn giản hơn cho use case này.

  • Use an API call from QuickSight to access the data that is in Amazon DynamoDB directly.
    ❌ Sai: QuickSight không hỗ trợ direct API query DynamoDB ở quy mô lớn/real-time (chỉ row-level/security hạn chế qua IAM). Với millions records, performance kém, chi phí cao (scan full table), không near real-time (throttle limits). QuickSight ưu tiên pre-aggregated data như S3/SPICE, không phải live DynamoDB queries.

  • Use Amazon Kinesis Data Firehose to export the data from Amazon DynamoDB to Amazon S3. Configure QuickSight to access the data in Amazon S3.
    ✅ Đúng: Như giải thích trên – Firehose managed service, tích hợp native DynamoDB Streams → S3 với buffering linh hoạt (60s-24h), compression/error handling tự động. QuickSight refresh S3 datasets nhanh (seconds cho SPICE import). Least delay thực tế cho high-throughput (updated AWS 2024: hỗ trợ OpenTelemetry monitoring, enhanced VPC).

📘 Tài liệu tham khảo (kiến thức cập nhật đến 2026)

  • AWS Docs - DynamoDB Streams + Kinesis Data Firehose: Integrate DynamoDB Streams with Kinesis Data Firehose (phiên bản 2024: low-latency delivery).
  • QuickSight Data Sources: Amazon QuickSight S3 Integration – SPICE cho near real-time từ S3.
  • Best Practices DOP-C02 Exam Guide (AWS Certified DevOps Engineer Pro 2024-2026): Serverless streaming với Firehose cho analytics pipelines.
  • AWS Well-Architected Framework - Analytics Lens: Khuyến nghị Firehose cho DynamoDB → S3 real-time.

🛠️ Lời khuyên DevOps: Implement Lambda trigger DynamoDB Streams nếu cần transform trước Firehose. Test với CloudWatch metrics cho latency! 🚀