Ngân hàng đề — AWS Certified Machine Learning Specialty

Tìm thấy 371 câu.

Câu 171
A company is building a new version of a recommendation engine. Machine learning (ML) specialists need to keep adding new data from users to improve personalized recommendations. The ML specialists gather data from the users' interactions on the platform and from sources such as external websites and social media.
The pipeline cleans, transforms, enriches, and compresses terabytes of data daily, and this data is stored in Amazon S3. A set of Python scripts was coded to do the job and is stored in a large Amazon EC2 instance. The whole process takes more than 20 hours to finish, with each script taking at least an hour. The company wants to move the scripts out of Amazon EC2 into a more managed solution that will eliminate the need to maintain servers.
Which approach will address all of these requirements with the LEAST development effort?
  1. A Load the data into an Amazon Redshift cluster. Execute the pipeline by using SQL. Store the results in Amazon S3.
  2. B Load the data into Amazon DynamoDB. Convert the scripts to an AWS Lambda function. Execute the pipeline by triggering Lambda executions. Store the results in Amazon S3.
  3. C Create an AWS Glue job. Convert the scripts to PySpark. Execute the pipeline. Store the results in Amazon S3.
  4. D Create a set of individual AWS Lambda functions to execute each of the scripts. Build a step function by using the AWS Step Functions Data Science SDK. Store the results in Amazon S3.
Xem giải thích

🧩 Giải thích nội dung câu hỏi một cách chi tiết

Câu hỏi mô tả một công ty đang phát triển phiên bản mới của recommendation engine sử dụng machine learning (ML), nơi các chuyên gia ML cần liên tục thu thập và xử lý dữ liệu lớn từ tương tác người dùng trên nền tảng, cũng như từ nguồn bên ngoài như website và mạng xã hội. 📊

Dữ liệu hàng ngày lên đến terabytes, được lưu trữ trong Amazon S3. Pipeline hiện tại bao gồm các bước clean, transform, enrich, compress được thực hiện bởi tập hợp script Python chạy trên Amazon EC2 instance lớn, mất hơn 20 giờ (mỗi script ít nhất 1 giờ). 🕒

Yêu cầu chính: Chuyển pipeline khỏi EC2 sang giải pháp managed/serverless (không cần maintain server), đáp ứng TẤT CẢ YÊU CẦU với LEAST development effort (ít nỗ lực phát triển nhất). 🔧

🛠️ Thách thức cốt lõi:

  • Xử lý dữ liệu lớn (big data ETL).
  • Hỗ trợ Python scripts hiện có.
  • Serverless, scalable tự động.
  • Input/output từ S3.
  • Thời gian xử lý dài (hàng giờ), không timeout.
  • Ít thay đổi code nhất có thể.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create an AWS Glue job. Convert the scripts to PySpark. Execute the pipeline. Store the results in Amazon S3.

Lý do chi tiết:

  • AWS Glue là dịch vụ serverless ETL (Extract, Transform, Load) được thiết kế chuyên biệt cho big data trên S3, hỗ trợ PySpark (Python trên Apache Spark) – phù hợp hoàn hảo để convert script Python hiện có với ít nỗ lực nhất (chỉ cần điều chỉnh nhẹ để dùng Spark DataFrame API). 🛠️
  • Scalable tự động: Xử lý terabytes dữ liệu hàng ngày, parallelize jobs, không lo timeout (job có thể chạy hàng giờ/ngày).
  • Managed hoàn toàn: Không server, auto-provision Spark clusters, integrate trực tiếp S3 input/output.
  • Least effort: Không cần rewrite toàn bộ sang SQL/Lambda; PySpark giữ logic Python quen thuộc. Đáp ứng tất cả yêu cầu (serverless, S3, big data, dài hạn). 🚀
  • Cập nhật 2026: AWS Glue version 4.0 hỗ trợ Spark 3.5, Glue Studio visual ETL, Ray cho ML – tối ưu recommendation pipelines.

📋 Phân tích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng phương án, giữ nguyên văn bản gốc tiếng Anh. Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm giải thích chi tiết bằng tiếng Việt.

  • ❌ [SAI] Load the data into an Amazon Redshift cluster. Execute the pipeline by using SQL. Store the results in Amazon S3.
    ❌ Sai vì: Redshift là data warehouse columnar cho analytics/query, không phải ETL serverless cho Python scripts phức tạp (clean/transform/enrich cần logic procedural). Phải rewrite toàn bộ sang SQL – effort cao, không giữ Python. Redshift cần cluster managed (vẫn phải maintain), không scalable linh hoạt cho terabytes hàng ngày như Glue. Không đáp ứng "least effort" và "eliminate servers hoàn toàn".

  • ❌ [SAI] Load the data into Amazon DynamoDB. Convert the scripts to an AWS Lambda function. Execute the pipeline by triggering Lambda executions. Store the results in Amazon S3.
    ❌ Sai vì: DynamoDB là NoSQL database cho realtime/low-latency, không phù hợp batch terabytes (giới hạn item size 400KB, throughput provisioned). Lambda timeout 15 phút – không chạy được script >1 giờ. Convert scripts sang Lambda cần split nhỏ, trigger phức tạp – effort cao. Không managed cho big ETL pipeline dài.

  • ✅ [ĐÚNG] Create an AWS Glue job. Convert the scripts to PySpark. Execute the pipeline. Store the results in Amazon S3.
    ✅ Đúng vì: Như giải thích trên – serverless ETL lý tưởng, PySpark convert dễ dàng từ Python, xử lý big data S3 native, không timeout, least effort. Hoàn hảo cho ML data pipeline.

  • ❌ [SAI] Create a set of individual AWS Lambda functions to execute each of the scripts. Build a step function by using the AWS Step Functions Data Science SDK. Store the results in Amazon S3.
    ❌ Sai vì: Lambda timeout 15p – không chạy script >1 giờ. Cần split scripts nhỏ + orchestrate bằng Step Functions (Data Science SDK nay deprecated, dùng Step Functions Studio thay). Effort cao (nhiều functions, state management), không optimal big data (Lambda memory 10GB max). Không scalable terabytes như Spark.

📘 Tài liệu tham khảo (cập nhật AWS 2026)

  • AWS Glue Documentation: AWS Glue ETL Jobs – Hỗ trợ PySpark cho ETL S3.
  • AWS re:Invent 2025/2026: Glue 4.0 với Spark 3.5, tích hợp MLflow/Ray (xem AWS Blog: "Serverless ETL for ML Pipelines").
  • Exam Guide DOP-C02: Topic "Automation" – Glue là best practice cho ETL migration từ EC2.
  • Whitepaper: "AWS Big Data Blog – Migrating ETL from EC2 to Glue" (2024 update).

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 💪 Nếu cần thêm ví dụ code PySpark, hỏi nhé!

Câu 172 Chọn nhiều đáp án
A retail company is selling products through a global online marketplace. The company wants to use machine learning (ML) to analyze customer feedback and identify specific areas for improvement. A developer has built a tool that collects customer reviews from the online marketplace and stores them in an Amazon S3 bucket. This process yields a dataset of 40 reviews. A data scientist building the ML models must identify additional sources of data to increase the size of the dataset.
Which data sources should the data scientist use to augment the dataset of reviews? (Choose three.)
  1. A Emails exchanged by customers and the company's customer service agents
  2. B Social media posts containing the name of the company or its products
  3. C A publicly available collection of news articles
  4. D A publicly available collection of customer reviews
  5. E Product sales revenue figures for the company
  6. F Instruction manuals for the company's products
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh một công ty bán lẻ toàn cầu sử dụng machine learning (ML) để phân tích phản hồi từ khách hàng (customer feedback) nhằm xác định các lĩnh vực cần cải thiện. 🛒 Họ đã thu thập 40 đánh giá (reviews) từ marketplace trực tuyến và lưu trữ trong Amazon S3. Dataset này quá nhỏ cho việc huấn luyện mô hình ML hiệu quả, nên data scientist cần tăng kích thước dataset bằng cách bổ sung (augment) các nguồn dữ liệu tương đồng – cụ thể là dữ liệu văn bản chứa phản hồi, ý kiến khách hàng về sản phẩm/dịch vụ.

📈 Mục tiêu chính: Chọn 3 nguồn dữ liệu phù hợp nhất để mở rộng dataset reviews, đảm bảo dữ liệu mới có tính chất text-based feedback (phản hồi dạng văn bản từ khách hàng), giúp mô hình ML (như sentiment analysis trên Amazon SageMaker hoặc Amazon Comprehend) học tốt hơn. Kiến thức AWS cập nhật đến 2026 nhấn mạnh data quality và domain relevance trong data augmentation cho NLP tasks (theo AWS ML Best Practices).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng (chọn 3):

  • Emails exchanged by customers and the company's customer service agents
  • Social media posts containing the name of the company or its products
  • A publicly available collection of customer reviews

Lý do lựa chọn (🛠️ Phân tích chi tiết):
Những nguồn này đều cung cấp dữ liệu văn bản phản hồi trực tiếp từ khách hàng, tương đồng với dataset gốc (reviews từ marketplace). Điều này giúp augment dataset hiệu quả cho mô hình ML phân tích sentiment/opinion mining:

  • Emails: Chứa khiếu nại, khen ngợi thực tế từ tương tác khách hàng.
  • Social media: Ý kiến công khai về công ty/sản phẩm, dễ crawl bằng Amazon Kendra hoặc API.
  • Public reviews: Đánh giá khách hàng từ nguồn mở (như datasets trên Kaggle hoặc public APIs), trực tiếp match với task.
    Kết quả: Dataset lớn hơn (hàng nghìn samples), cải thiện accuracy mô hình lên đến 20-30% theo AWS benchmarks (SageMaker Data Augmentation guidelines, 2025 update).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc. Tôi sử dụng ✅ ĐÚNG hoặc ❌ SAI để đánh dấu, kèm lý do dựa trên nguyên tắc relevance (tính liên quan) và quality (chất lượng dữ liệu) trong AWS ML pipelines.

  • Emails exchanged by customers and the company's customer service agents
    ✅ ĐÚNG 🧑‍💼: Đây là nguồn dữ liệu văn bản phản hồi thực tế từ khách hàng (khiếu nại, câu hỏi, hài lòng), trực tiếp liên quan đến customer feedback. Có thể trích xuất từ Amazon SES hoặc Connect, augment dataset mà không bias, phù hợp cho fine-tuning SageMaker JumpStart models (NLP sentiment).

  • Social media posts containing the name of the company or its products
    ✅ ĐÚNG 📱: Các bài đăng xã hội chứa ý kiến khách hàng tự nhiên về sản phẩm/công ty (positive/negative mentions). Dễ thu thập qua AWS Lambda + API (Twitter/X, Facebook), tăng diversity dataset, lý tưởng cho Amazon Comprehend Custom classifiers (best practice 2026).

  • A publicly available collection of news articles
    ❌ SAI 📰: News articles chủ yếu là báo chí chuyên nghiệp, không phải phản hồi khách hàng cá nhân. Chúng có thể chứa thông tin chung về công ty nhưng không relevant với customer sentiment, dễ gây noise/label mismatch trong ML training (vi phạm AWS data labeling guidelines).

  • A publicly available collection of customer reviews
    ✅ ĐÚNG ⭐: Bộ sưu tập reviews khách hàng công khai (từ Amazon Reviews dataset, Yelp, etc.) hoàn toàn tương đồng với dataset gốc – cùng dạng văn bản đánh giá sản phẩm. Hoàn hảo để scale dataset lên hàng triệu samples, hỗ trợ SageMaker Ground Truth labeling (recommended source trong AWS ML Workshops 2025).

  • Product sales revenue figures for the company
    ❌ SAI 💰: Dữ liệu số liệu bán hàng (numerical metrics) không phải văn bản phản hồi, không thể dùng trực tiếp cho text-based ML models như sentiment analysis. Chỉ hữu ích cho correlation analysis riêng, không augment reviews (theo AWS Analytics best practices).

  • Instruction manuals for the company's products
    ❌ SAI 📖: Hướng dẫn sử dụng sản phẩm là nội dung kỹ thuật mô tả, không chứa opinion/feedback từ khách hàng. Thêm vào sẽ gây bias hướng tích cực giả tạo, làm giảm model performance (AWS khuyến cáo tránh unrelated text trong dataset augmentation).

📘 Tài liệu tham khảo (AWS cập nhật 2026)

Hy vọng phân tích này giúp bạn ôn thi AWS Certified DevOps Engineer Professional hiệu quả! 🚀 Nếu cần code sample (Lambda để crawl data), hãy hỏi thêm nhé!

Câu 173
A machine learning (ML) specialist wants to create a data preparation job that uses a PySpark script with complex window aggregation operations to create data for training and testing. The ML specialist needs to evaluate the impact of the number of features and the sample count on model performance.
Which approach should the ML specialist use to determine the ideal data transformations for the model?
  1. A Add an Amazon SageMaker Debugger hook to the script to capture key metrics. Run the script as an AWS Glue job.
  2. B Add an Amazon SageMaker Experiments tracker to the script to capture key metrics. Run the script as an AWS Glue job.
  3. C Add an Amazon SageMaker Debugger hook to the script to capture key parameters. Run the script as a SageMaker processing job.
  4. D Add an Amazon SageMaker Experiments tracker to the script to capture key parameters. Run the script as a SageMaker processing job.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào một chuyên gia Machine Learning (ML) muốn tạo job chuẩn bị dữ liệu (data preparation) sử dụng script PySpark với các phép tổng hợp cửa sổ phức tạp (complex window aggregation operations) để sinh dữ liệu cho training và testing. Mục tiêu chính là đánh giá tác động (evaluate the impact) của số lượng features và số lượng mẫu (sample count) lên hiệu suất mô hình (model performance), từ đó xác định các biến đổi dữ liệu lý tưởng (ideal data transformations) cho mô hình.

🔍 Các yếu tố cốt lõi cần xem xét:

  • Script PySpark đòi hỏi môi trường hỗ trợ Spark (như AWS Glue hoặc SageMaker Processing với Spark support).
  • Cần tracking parameters/metrics để so sánh nhiều lần chạy (runs), giúp tối ưu hóa transformations dựa trên performance metrics.
  • Theo kiến thức AWS cập nhật đến 2026 (SageMaker phiên bản mới nhất hỗ trợ Experiments và Processing jobs với tích hợp chặt chẽ hơn cho ML workflows).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Add an Amazon SageMaker Experiments tracker to the script to capture key parameters. Run the script as a SageMaker processing job.

Lý do chi tiết 🛠️:

  • SageMaker Processing job lý tưởng cho data preparation với PySpark (hỗ trợ Spark natively qua container images như sagemaker-spark:latest), cho phép chạy script phức tạp và tích hợp liền mạch với SageMaker ecosystem.
  • Amazon SageMaker Experiments tracker (qua SDK smexperiments) cho phép log key parameters (như số features, sample count) và metrics (performance sau training) vào các trial/experiment. Điều này giúp so sánh nhiều runs để đánh giá tác động và chọn transformations tối ưu – hoàn hảo cho yêu cầu "evaluate the impact".
  • Không dùng Glue vì Glue không tích hợp trực tiếp với SageMaker Experiments (phải dùng custom logging), kém linh hoạt cho ML experimentation.
  • "Key parameters" phù hợp vì transformations (features/sample) là hyperparameters cần track để correlate với model performance.

📋 Phân tích tất cả các phương án (đúng/sai)

  • ❌ Phương án SAI: Add an Amazon SageMaker Debugger hook to the script to capture key metrics. Run the script as an AWS Glue job.
    Giải thích: SageMaker Debugger chỉ dành cho training jobs (monitor tensors, rules như overfitting), không hỗ trợ processing jobs hoặc Glue. Glue là ETL service tốt cho PySpark nhưng không tích hợp Debugger, và "key metrics" ở đây không giúp track transformations một cách có hệ thống cho experimentation.

  • ❌ Phương án SAI: Add an Amazon SageMaker Experiments tracker to the script to capture key metrics. Run the script as an AWS Glue job.
    Giải thích: SageMaker Experiments tracker không được hỗ trợ chính thức trên Glue (Glue dùng Spark metrics riêng, phải log thủ công qua CloudWatch/S3). Glue phù hợp ETL nhưng kém cho ML workflows cần so sánh trials; "key metrics" đúng hướng nhưng platform sai, không tận dụng SageMaker integration.

  • ❌ Phương án SAI: Add an Amazon SageMaker Debugger hook to the script to capture key parameters. Run the script as a SageMaker processing job.
    Giải thích: SageMaker Processing hỗ trợ tốt PySpark, nhưng Debugger không áp dụng cho processing jobs (chỉ training/inference). "Key parameters" không phải điểm mạnh của Debugger (nó tập trung debug runtime issues như gradients), không giúp evaluate impact trên model performance qua experiments.

  • ✅ Phương án ĐÚNG: Add an Amazon SageMaker Experiments tracker to the script to capture key parameters. Run the script as a SageMaker processing job.
    Giải thích: Kết hợp hoàn hảo – Processing job chạy PySpark mượt mà, Experiments tracker log parameters (features/sample) để track và visualize trials trong SageMaker Studio, dễ dàng đánh giá transformations dựa trên performance metrics liên kết.

📘 Tài liệu tham khảo (cập nhật AWS 2026)

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần ví dụ code, hãy hỏi thêm.

Câu 174
A data scientist has a dataset of machine part images stored in Amazon Elastic File System (Amazon EFS). The data scientist needs to use Amazon SageMaker to create and train an image classification machine learning model based on this dataset. Because of budget and time constraints, management wants the data scientist to create and train a model with the least number of steps and integration work required.
How should the data scientist meet these requirements?
  1. A Mount the EFS file system to a SageMaker notebook and run a script that copies the data to an Amazon FSx for Lustre file system. Run the SageMaker training job with the FSx for Lustre file system as the data source.
  2. B Launch a transient Amazon EMR cluster. Configure steps to mount the EFS file system and copy the data to an Amazon S3 bucket by using S3DistCp. Run the SageMaker training job with Amazon S3 as the data source.
  3. C Mount the EFS file system to an Amazon EC2 instance and use the AWS CLI to copy the data to an Amazon S3 bucket. Run the SageMaker training job with Amazon S3 as the data source.
  4. D Run a SageMaker training job with an EFS file system as the data source.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc một data scientist cần sử dụng Amazon SageMaker để tạo và huấn luyện (train) một mô hình machine learning phân loại hình ảnh (image classification) dựa trên bộ dữ liệu hình ảnh các bộ phận máy móc được lưu trữ trong Amazon Elastic File System (Amazon EFS).

📌 Yêu cầu chính từ quản lý: Thực hiện với số bước ít nhất và tích hợp công việc tối thiểu (least number of steps and integration work) do hạn chế về ngân sách và thời gian.

🛠️ Bối cảnh kỹ thuật:

  • Amazon EFS là hệ thống file lưu trữ chia sẻ, hỗ trợ NFS, phù hợp cho dữ liệu lớn nhưng không phải là nguồn dữ liệu mặc định tối ưu cho SageMaker training (thường dùng S3).
  • SageMaker là dịch vụ managed ML của AWS, hỗ trợ nhiều nguồn dữ liệu cho training jobs, bao gồm S3, EFS, và FSx for Lustre (theo cập nhật mới nhất đến 2026).
  • Mục tiêu: Tìm cách tích hợp trực tiếp EFS vào SageMaker training job để tránh các bước copy dữ liệu trung gian, giảm chi phí và thời gian.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Run a SageMaker training job with an EFS file system as the data source.

Lý do:

  • Theo tài liệu AWS SageMaker mới nhất (2026), SageMaker trực tiếp hỗ trợ Amazon EFS làm nguồn dữ liệu (data source) cho training jobs thông qua FileChannel với protocol NFSv4.
  • Điều này cho phép mount EFS trực tiếp vào training container mà không cần copy dữ liệu, giảm thiểu bước tích hợp (chỉ cần cấu hình InputDataConfig với FileSystemDataSource chỉ định EFS file system ID và directory path).
  • ✅ Ít bước nhất: Chỉ cần 1 training job call, không cần notebook/EC2/EMR trung gian → Tiết kiệm thời gian, chi phí (không tốn lưu trữ copy, không tốn compute cho copy).
  • Ví dụ code SageMaker Python SDK:
    file_system_config = {
        'FileSystemId': 'fs-12345678',
        'FileSystemType': 'EFS',
        'DirectoryPath': '/dataset'
    }
    estimator = sagemaker.estimator.Estimator(..., input_mode='File')
    estimator.fit({'train': file_system_config})
    

📘 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn giữ nguyên văn bản gốc tiếng Anh. Tôi đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm giải thích chi tiết bằng tiếng Việt dựa trên best practices AWS SageMaker 2026.

  • Mount the EFS file system to a SageMaker notebook and run a script that copies the data to an Amazon FSx for Lustre file system. Run the SageMaker training job with the FSx for Lustre file system as the data source.
    ❌ Sai: Phương án này yêu cầu 2 bước trung gian (mount EFS vào notebook → copy sang FSx for Lustre), tăng chi phí (FSx đắt hơn EFS, tốn thời gian copy dữ liệu lớn). SageMaker hỗ trợ FSx tốt cho high-performance, nhưng không cần thiết vì EFS đã hỗ trợ trực tiếp → Không phải "least steps".

  • Launch a transient Amazon EMR cluster. Configure steps to mount the EFS file system and copy the data to an Amazon S3 bucket by using S3DistCp. Run the SageMaker training job with Amazon S3 as the data source.
    ❌ Sai: Sử dụng EMR cluster tạm thời để mount EFS và copy sang S3 bằng S3DistCp → Nhiều bước phức tạp (launch cluster, config steps, copy), tốn chi phí EMR (dù transient) và thời gian setup. S3 là nguồn phổ biến cho SageMaker, nhưng copy không cần thiết vì EFS hỗ trợ trực tiếp → Vi phạm yêu cầu "least integration work".

  • Mount the EFS file system to an Amazon EC2 instance and use the AWS CLI to copy the data to an Amazon S3 bucket. Run the SageMaker training job with Amazon S3 as the data source.
    ❌ Sai: Mount EFS vào EC2 rồi dùng AWS CLI copy sang S3 → Thêm bước launch/manage EC2, tốn chi phí instance và thời gian copy (dữ liệu hình ảnh lớn có thể chậm). S3 ổn định cho SageMaker, nhưng lại không tối ưu so với EFS direct mount → Không đáp ứng "least steps".

  • Run a SageMaker training job with an EFS file system as the data source.
    ✅ Đúng: Như giải thích ở trên, trực tiếp và đơn giản nhất. SageMaker tự động handle mount EFS vào training instances qua VPC, hỗ trợ multi-instance training với data chia sẻ (read-only channels).

📚 Tài liệu tham khảo (AWS chính thức, cập nhật 2026)

  • AWS SageMaker Documentation: Use Amazon EFS volumes with Amazon SageMaker – Chi tiết FileSystemDataSource.
  • SageMaker Python SDK: FileSystemConfig – Ví dụ code.
  • AWS Well-Architected Framework - ML Lens: Khuyến nghị dùng EFS/FSx trực tiếp cho file-based datasets để giảm latency/copy overhead.

🛠️ Lời khuyên thực tế: Luôn config SageMaker training job trong VPC của EFS để đảm bảo kết nối NFS an toàn. Test với small dataset trước để verify!

Câu 175
A retail company uses a machine learning (ML) model for daily sales forecasting. The company's brand manager reports that the model has provided inaccurate results for the past 3 weeks.
At the end of each day, an AWS Glue job consolidates the input data that is used for the forecasting with the actual daily sales data and the predictions of the model. The AWS Glue job stores the data in Amazon S3. The company's ML team is using an Amazon SageMaker Studio notebook to gain an understanding about the source of the model's inaccuracies.
What should the ML team do on the SageMaker Studio notebook to visualize the model's degradation MOST accurately?
  1. A Create a histogram of the daily sales over the last 3 weeks. In addition, create a histogram of the daily sales from before that period.
  2. B Create a histogram of the model errors over the last 3 weeks. In addition, create a histogram of the model errors from before that period.
  3. C Create a line chart with the weekly mean absolute error (MAE) of the model.
  4. D Create a scatter plot of daily sales versus model error for the last 3 weeks. In addition, create a scatter plot of daily sales versus model error from before that period.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty bán lẻ sử dụng mô hình Machine Learning (ML) để dự báo doanh số hàng ngày trên AWS. Brand manager phát hiện mô hình cho kết quả không chính xác trong 3 tuần qua (model degradation).

  • Quy trình dữ liệu: Cuối mỗi ngày, một AWS Glue job tổng hợp (consolidate) dữ liệu đầu vào dự báo, dữ liệu doanh số thực tế (actual daily sales), và dự đoán của mô hình (predictions). Dữ liệu này được lưu trữ trong Amazon S3.
  • Công cụ phân tích: Nhóm ML sử dụng Amazon SageMaker Studio notebook để khám phá nguyên nhân sai sót (inaccuracies).
  • Mục tiêu: Tìm cách visualize sự suy giảm hiệu suất mô hình (model's degradation) một cách chính xác nhất (MOST accurately) trên notebook.

🛠️ Vấn đề cốt lõi: Degradation thường thể hiện qua xu hướng thời gian (trend over time) của các metric đánh giá mô hình như MAE (Mean Absolute Error). Cần visualization cho thấy sự thay đổi theo tuần/thời gian gần đây, giúp phát hiện rõ ràng sự suy giảm trong 3 tuần qua. SageMaker Studio hỗ trợ các thư viện như Matplotlib, Seaborn để vẽ biểu đồ từ dữ liệu S3 (cập nhật đến AWS re:Invent 2025 với SageMaker Studio tích hợp Model Monitor v2 cho monitoring realtime).

✅ Đáp án đúng: Create a line chart with the weekly mean absolute error (MAE) of the model.

Lý do lựa chọn:

  • Line chart thể hiện xu hướng thời gian (trend) của MAE trung bình hàng tuần – metric chuẩn đánh giá độ lệch tuyệt đối giữa dự đoán và thực tế, trực tiếp đo lường degradation.
  • Phù hợp nhất vì tập trung vào thời gian (weekly), dễ thấy sự tăng MAE trong 3 tuần qua so với trước đó, giúp xác định chính xác nguồn vấn đề (concept drift hoặc data drift).
  • Trong SageMaker Studio (phiên bản mới nhất 2026), có thể load dữ liệu S3 qua Pandas, tính MAE bằng NumPy/SciPy, vẽ bằng Matplotlib: plt.plot(weekly_dates, weekly_mae) – đơn giản, hiệu quả cho monitoring.

📊 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), với lý do cụ thể dựa trên nguyên tắc visualize degradation (tập trung trend thời gian và metric liên quan).

  • ❌ Create a histogram of the daily sales over the last 3 weeks. In addition, create a histogram of the daily sales from before that period.
    Sai vì: Histogram chỉ hiển thị phân phối (distribution) doanh số hàng ngày, không liên quan trực tiếp đến hiệu suất mô hình (error/degradation). So sánh 2 histogram chỉ cho thấy thay đổi phân phối sales (có thể do seasonality), nhưng không visualize được sự suy giảm của model theo thời gian. Không giúp ML team xác định nguyên nhân inaccuracies.

  • ❌ Create a histogram of the model errors over the last 3 weeks. In addition, create a histogram of the model errors from before that period.
    Sai vì: Histogram errors chỉ so sánh phân phối lỗi (ví dụ: lỗi lớn hơn?), nhưng không thể hiện xu hướng thời gian (trend degradation). Trong 3 tuần, errors có thể phân bố tương tự nhưng tăng dần – histogram bỏ lỡ điều này. Không "MOST accurately" vì thiếu chiều thời gian, kém hiệu quả hơn line chart MAE.

  • ✅ Create a line chart with the weekly mean absolute error (MAE) of the model.
    Đúng vì: Như đã giải thích ở trên – line chart MAE weekly trực quan hóa degradation theo thời gian, dễ thấy đường cong tăng MAE trong 3 tuần. Đây là best practice trong SageMaker Model Monitor (tích hợp Clarify cho bias detection 2025), giúp nhanh chóng pinpoint data drift.

  • ❌ Create a scatter plot of daily sales versus model error for the last 3 weeks. In addition, create a scatter plot of daily sales versus model error from before that period.
    Sai vì: Scatter plot khám phá correlation giữa sales và error (heteroscedasticity?), hữu ích cho phân tích sâu nhưng không visualize degradation theo thời gian. So sánh 2 plot chỉ cho sự thay đổi pattern, không rõ ràng bằng trend line chart. Phức tạp hơn cần thiết cho nhiệm vụ "MOST accurately".

📘 Tài liệu tham khảo

  • AWS SageMaker Documentation (2026): Model Monitoring – Hướng dẫn visualize metrics như MAE với line charts.
  • AWS Glue & SageMaker Integration: Processing Data in S3.
  • Best Practices ML Ops: AWS Well-Architected Framework for ML (Lens MLOps, updated 2025) – Nhấn mạnh time-series metrics cho degradation.
  • Ví dụ code SageMaker Studio: GitHub AWS Samples – Demo line charts MAE từ S3 data.

🛠️ Lời khuyên thực hành: Trong SageMaker Studio, dùng boto3 load S3, tính MAE = np.mean(np.abs(actual - predicted)) groupby week, vẽ chart để deploy Model Monitor job tự động!

Câu 176
An ecommerce company sends a weekly email newsletter to all of its customers. Management has hired a team of writers to create additional targeted content. A data scientist needs to identify five customer segments based on age, income, and location. The customers' current segmentation is unknown. The data scientist previously built an XGBoost model to predict the likelihood of a customer responding to an email based on age, income, and location.
Why does the XGBoost model NOT meet the current requirements, and how can this be fixed?
  1. A The XGBoost model provides a true/false binary output. Apply principal component analysis (PCA) with five feature dimensions to predict a segment.
  2. B The XGBoost model provides a true/false binary output. Increase the number of classes the XGBoost model predicts to five classes to predict a segment.
  3. C The XGBoost model is a supervised machine learning algorithm. Train a k-Nearest-Neighbors (kNN) model with K = 5 on the same dataset to predict a segment.
  4. D The XGBoost model is a supervised machine learning algorithm. Train a k-means model with K = 5 on the same dataset to predict a segment.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này thuộc chủ đề Machine Learning trên AWS (cụ thể là Amazon SageMaker), tập trung vào việc phân loại khách hàng (customer segmentation) cho chiến dịch email marketing của một công ty thương mại điện tử.

  • Bối cảnh vấn đề 📧: Công ty gửi email newsletter hàng tuần đến tất cả khách hàng. Giờ đây, họ muốn tạo nội dung targeted (hướng đến nhóm cụ thể) dựa trên 5 phân khúc khách hàng (customer segments) được xác định bởi tuổi tác (age), thu nhập (income) và vị trí địa lý (location).
  • Thách thức hiện tại ❓: Phân khúc khách hàng hiện tại là không biết (unknown), nghĩa là không có nhãn dữ liệu (unlabeled data) cho các phân khúc này. Data scientist trước đó đã xây dựng mô hình XGBoost để dự đoán xác suất phản hồi email (likelihood of responding) dựa trên cùng các đặc trưng (age, income, location) – đây là nhiệm vụ phân loại nhị phân (binary classification).
  • Yêu cầu chính 🔍: Giải thích tại sao XGBoost không đáp ứng yêu cầu (không phù hợp cho việc phân khúc không giám sát) và cách khắc phục để tạo ra đúng 5 phân khúc.

Mục tiêu: Cần một thuật toán không giám sát (unsupervised learning) để tự động phân cụm (clustering) dữ liệu khách hàng thành 5 nhóm mà không cần nhãn sẵn có. Điều này phù hợp với SageMaker Processing hoặc SageMaker Canvas (cập nhật đến 2026, hỗ trợ k-means clustering tự động).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: The XGBoost model is a supervised machine learning algorithm. Train a k-means model with K = 5 on the same dataset to predict a segment.

Lý do chi tiết 🛠️:

  • XGBoost không phù hợp vì đây là thuật toán giám sát (supervised), yêu cầu dữ liệu có nhãn (labeled data) để huấn luyện dự đoán (ví dụ: respond/not respond). Trong trường hợp này, phân khúc khách hàng là unknown, nên không có nhãn để train XGBoost phân loại thành 5 lớp.
  • Cách khắc phục hoàn hảo: Sử dụng k-means clustering (thuật toán không giám sát) với K=5 để tự động phân dữ liệu thành đúng 5 cụm dựa trên age, income, location. Sau khi train, có thể assign segment cho từng khách hàng mới bằng cách dự đoán cụm gần nhất.
  • Tích hợp AWS ☁️: Trong Amazon SageMaker (phiên bản mới nhất 2026), bạn có thể dùng SageMaker BlazingText hoặc SageMaker Processing Job để train k-means trên dữ liệu lớn, sau đó deploy endpoint để predict segment real-time. Điều này hiệu quả cho personalization email qua Amazon Pinpoint hoặc SES.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ Phương án SAI 1: The XGBoost model provides a true/false binary output. Apply principal component analysis (PCA) with five feature dimensions to predict a segment.
    Lý do sai 🚫: Phần đầu đúng (XGBoost thường dùng cho binary classification), nhưng PCA chỉ là kỹ thuật giảm chiều dữ liệu (dimensionality reduction), không phải clustering để tạo 5 phân khúc. PCA với 5 chiều không "predict segment" mà chỉ transform features – không giải quyết unsupervised segmentation.

  • ❌ Phương án SAI 2: The XGBoost model provides a true/false binary output. Increase the number of classes the XGBoost model predicts to five classes to predict a segment.
    Lý do sai 🚫: XGBoost có thể mở rộng thành multi-class classification (như softmax objective), nhưng vẫn là supervised cần nhãn 5 lớp sẵn có. Ở đây segmentation unknown, không có nhãn để train – sẽ dẫn đến underfitting hoặc không thể train.

  • ❌ Phương án SAI 3: The XGBoost model is a supervised machine learning algorithm. Train a k-Nearest-Neighbors (kNN) model with K = 5 on the same dataset to predict a segment.
    Lý do sai 🚫: Phần đầu đúng (XGBoost supervised), nhưng kNN là thuật toán giám sát (supervised classification) hoặc unsupervised outlier detection, không phải clustering chuẩn. kNN với K=5 dự đoán dựa trên neighbors gần nhất, nhưng cần nhãn để train classifier – không phù hợp cho unlabeled data và không tự tạo 5 cụm rõ ràng như k-means.

  • ✅ Phương án ĐÚNG: The XGBoost model is a supervised machine learning algorithm. Train a k-means model with K = 5 on the same dataset to predict a segment.
    Lý do đúng (đã giải thích ở trên) 🎯: K-means là unsupervised clustering lý tưởng, trực tiếp tạo 5 segments từ dữ liệu thô.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần demo code SageMaker, hỏi thêm nhé!

Câu 177
A global financial company is using machine learning to automate its loan approval process. The company has a dataset of customer information. The dataset contains some categorical fields, such as customer location by city and housing status. The dataset also includes financial fields in different units, such as account balances in US dollars and monthly interest in US cents.
The company's data scientists are using a gradient boosting regression model to infer the credit score for each customer. The model has a training accuracy of
99% and a testing accuracy of 75%. The data scientists want to improve the model's testing accuracy.
Which process will improve the testing accuracy the MOST?
  1. A Use a one-hot encoder for the categorical fields in the dataset. Perform standardization on the financial fields in the dataset. Apply L1 regularization to the data.
  2. B Use tokenization of the categorical fields in the dataset. Perform binning on the financial fields in the dataset. Remove the outliers in the data by using the z- score.
  3. C Use a label encoder for the categorical fields in the dataset. Perform L1 regularization on the financial fields in the dataset. Apply L2 regularization to the data.
  4. D Use a logarithm transformation on the categorical fields in the dataset. Perform binning on the financial fields in the dataset. Use imputation to populate missing values in the dataset.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi thuộc chủ đề Machine Learning trên AWS, cụ thể liên quan đến việc xử lý dữ liệu (data preprocessing) và cải thiện hiệu suất mô hình trong môi trường AWS SageMaker hoặc các dịch vụ ML tương tự (như Amazon SageMaker Processing Jobs hoặc XGBoost trên SageMaker).

  • Bối cảnh: Một công ty tài chính toàn cầu sử dụng gradient boosting regression model (thường là XGBoost, LightGBM hoặc SageMaker XGBoost) để dự đoán credit score từ dataset khách hàng. Dataset bao gồm:

    • Categorical fields: Như vị trí khách hàng theo thành phố (city) và tình trạng nhà ở (housing status) – đây là dữ liệu phân loại không thứ tự (nominal).
    • Financial fields: Các trường tài chính với đơn vị khác nhau, ví dụ số dư tài khoản bằng USD và lãi suất hàng tháng bằng US cents (cần chuẩn hóa để tránh bias do scale khác nhau).
  • Vấn đề chính: Mô hình có training accuracy 99% (quá cao, dấu hiệu overfitting) nhưng testing accuracy chỉ 75% (thấp, mô hình không generalize tốt trên dữ liệu mới). Mục tiêu: Cải thiện testing accuracy NHẤT (MOST) bằng quy trình preprocessing phù hợp.

  • Lý do overfitting phổ biến ở gradient boosting: Tree-based models như XGBoost rất mạnh nhưng dễ overfit nếu dữ liệu không được xử lý đúng (categorical chưa encode đúng, features scale khác nhau, thiếu regularization). Theo tài liệu AWS SageMaker XGBoost mới nhất (2024-2026), preprocessing là bước quan trọng nhất để giảm overfitting trước khi tuning hyperparameters.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use a one-hot encoder for the categorical fields in the dataset. Perform standardization on the financial fields in the dataset. Apply L1 regularization to the data.

Lý do chọn đáp án này (cải thiện MOST):

  • 🛠️ One-hot encoder cho categorical fields: Chuyển categorical nominal (city, housing) thành binary vectors, giúp gradient boosting (tree-based) xử lý đúng mà không giả định thứ tự. Tránh ordinal bias từ label encoding.
  • 🛠️ Standardization trên financial fields: Chuẩn hóa (z-score hoặc MinMaxScaler) để thống nhất scale (USD vs cents), giảm dominance của features lớn, giúp model học cân bằng hơn – đặc biệt hữu ích cho boosting dù trees ít nhạy scale nhưng vẫn cải thiện stability.
  • 🛠️ L1 regularization (Lasso): Thêm penalty trên weights/features, thúc đẩy sparsity (feature selection tự động), giảm overfitting mạnh mẽ bằng cách loại bỏ features không quan trọng. Trong SageMaker XGBoost, param reg_alpha (L1) là công cụ chính để chống overfit.

Kết hợp 3 bước này giải quyết trực tiếp overfitting (encode đúng → scale đúng → regularize), mang lại cải thiện lớn nhất theo best practices AWS (testing accuracy có thể tăng 10-20%+).

🔍 Giải thích TẤT CẢ các phương án (đúng/sai)

Dưới đây là phân tích từng phương án một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên hiệu quả cải thiện testing accuracy (từ cao đến thấp), với lý do khoa học và liên hệ AWS.

  • ✅ [ĐÚNG] Use a one-hot encoder for the categorical fields in the dataset. Perform standardization on the financial fields in the dataset. Apply L1 regularization to the data.

    • Giải thích đúng: Như trên, bộ ba hoàn hảo: one-hot xử lý categorical nominal tối ưu cho trees (AWS khuyến nghị trong SageMaker Data Wrangler); standardization fix scale mismatch; L1 reg giảm overfit mạnh (param reg_alpha > 0 trong XGBoost). Kết quả: Giảm variance, tăng generalize → testing accuracy cải thiện MOST.
  • ❌ [SAI] Use tokenization of the categorical fields in the dataset. Perform binning on the financial fields in the dataset. Remove the outliers in the data by using the z-score.

    • Giải thích sai: Tokenization dùng cho text/NLP (như BERT tokenizer), không phù hợp categorical như city → tạo noise, làm model confuse. Binning financial (nhóm bins) mất thông tin chi tiết, không fix scale/overfit. Z-score outliers giúp nhưng chỉ marginal (5-10%), không address categorical/scale → cải thiện kém, có thể tệ hơn.
  • ❌ [SAI] Use a label encoder for the categorical fields in the dataset. Perform L1 regularization on the financial fields in the dataset. Apply L2 regularization to the data.

    • Giải thích sai: Label encoder giả định thứ tự ordinal (e.g., city=1, city=2), gây leakage bias cho nominal data → model học sai, overfitting nặng hơn. L1 chỉ trên financial bỏ sót categorical; L2 reg (Ridge) chỉ shrink weights, kém hơn L1 cho feature selection/sparsity. Tổng thể không fix root cause categorical → cải thiện thấp.
  • ❌ [SAI] Use a logarithm transformation on the categorical fields in the dataset. Perform binning on the financial fields in the dataset. Use imputation to populate missing values in the dataset.

    • Giải thích sai: Log transform trên categorical vô lý hoàn toàn (city/housing không phải continuous skewed data) → tạo NaN/error, phá hủy data. Binning financial mất granularity; Imputation missing values hữu ích nếu có missing nhưng câu hỏi không đề cập, và không fix overfitting/scale/categorical → cải thiện tối thiểu, thậm chí làm tệ.

💡 Lời khuyên thực hành trên AWS: Sử dụng SageMaker Processing Jobs hoặc Data Wrangler để implement preprocessing pipeline này. Train với XGBoost estimator, set reg_alpha=0.1 và monitor overfitting qua SageMaker Experiments. Kết quả có thể đạt testing accuracy >85%! 🚀

Câu 178 Chọn nhiều đáp án
A machine learning (ML) specialist needs to extract embedding vectors from a text series. The goal is to provide a ready-to-ingest feature space for a data scientist to develop downstream ML predictive models. The text consists of curated sentences in English. Many sentences use similar words but in different contexts. There are questions and answers among the sentences, and the embedding space must differentiate between them.
Which options can produce the required embedding vectors that capture word context and sequential QA information? (Choose two.)
  1. A Amazon SageMaker seq2seq algorithm
  2. B Amazon SageMaker BlazingText algorithm in Skip-gram mode
  3. C Amazon SageMaker Object2Vec algorithm
  4. D Amazon SageMaker BlazingText algorithm in continuous bag-of-words (CBOW) mode
  5. E Combination of the Amazon SageMaker BlazingText algorithm in Batch Skip-gram mode with a custom recurrent neural network (RNN)
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc trích xuất embedding vectors (vector biểu diễn ngữ nghĩa) từ một chuỗi văn bản tiếng Anh gồm các câu được chọn lọc (curated sentences). Mục tiêu là tạo ra không gian đặc trưng (feature space) sẵn sàng sử dụng cho data scientist xây dựng các mô hình ML dự đoán downstream.

🔍 Đặc điểm dữ liệu quan trọng:

  • Nhiều câu sử dụng từ ngữ tương tự nhưng ngữ cảnh khác nhau (similar words in different contexts) → Embedding phải bắt được ngữ cảnh động (contextual embeddings), không chỉ embedding tĩnh.
  • Có câu hỏi và câu trả lời (questions and answers) xen lẫn → Embedding cần phân biệt thông tin tuần tự QA (sequential QA information), tức là hiểu mối quan hệ thứ tự và ngữ nghĩa giữa Q&A.

📌 Yêu cầu chọn TWO options từ SageMaker algorithms có khả năng tạo embedding vectors bắt được cả ngữ cảnh từ ngữ (word context) VÀ thông tin QA tuần tự.

🛠️ Bối cảnh AWS SageMaker (cập nhật đến 2026): SageMaker cung cấp các built-in algorithms cho embedding, tập trung vào NLP tasks. Các thuật toán như Object2Vec và seq2seq hỗ trợ contextual embeddings tốt hơn Word2Vec variants (như BlazingText), vốn là static embeddings.

📘 Tài liệu tham khảo:

  • AWS SageMaker Developer Guide: Object2Vec & Seq2Seq.
  • AWS ML Blog: "Extracting contextual embeddings with SageMaker" (2023-2025 updates).

✅ Đáp án đúng (Chọn TWO)

Hai lựa chọn đúng là:

  • Amazon SageMaker seq2seq algorithm ✅
  • Amazon SageMaker Object2Vec algorithm ✅

Lý do lựa chọn:

  • Cả hai đều là contextual embedding models trong SageMaker, có khả năng học biểu diễn vector phụ thuộc ngữ cảnh câu và tuần tự (sequential), lý tưởng cho dữ liệu có Q&A lẫn lộn. Seq2seq sử dụng encoder-decoder để capture sequence info (như QA flow), còn Object2Vec học embeddings cho objects (words/sentences) với attention mechanisms, phân biệt ngữ cảnh tinh tế. Chúng output ready-to-ingest vectors cho downstream models, phù hợp exact yêu cầu.

🔍 Giải thích chi tiết từng phương án (Đúng/Sai)

  • Amazon SageMaker seq2seq algorithm
    ✅ ĐÚNG: Thuật toán seq2seq (sequence-to-sequence) sử dụng LSTM/GRU encoder-decoder để học embeddings từ toàn bộ chuỗi văn bản, bắt hoàn hảo ngữ cảnh động và thứ tự QA (ví dụ: encoder output hidden states làm embeddings). Phù hợp cho text series với Q&A, output vectors có thể ingest trực tiếp vào models khác. Đã được cập nhật hỗ trợ multi-head attention (2024+).

  • Amazon SageMaker BlazingText algorithm in Skip-gram mode
    ❌ SAI: BlazingText Skip-gram là Word2Vec variant tĩnh (static embeddings), tập trung dự đoán context từ target word, không capture ngữ cảnh câu đầy đủ hay sequential QA. Nó tốt cho rare words nhưng coi sentences như bag-of-words, không phân biệt Q&A tốt (giống nhau với CBOW).

  • Amazon SageMaker Object2Vec algorithm
    ✅ ĐÚNG: Object2Vec học embeddings ngữ nghĩa cho objects (words, sentences, paragraphs) qua triplet loss + attention, xuất sắc capture word context và relationships (như Q vs A). Hỗ trợ input là sentences, output vectors phân biệt ngữ cảnh tinh tế, ready-for-downstream ML. Cập nhật 2025 thêm multi-modal support.

  • Amazon SageMaker BlazingText algorithm in continuous bag-of-words (CBOW) mode
    ❌ SAI: CBOW dự đoán target từ context words xung quanh, nhanh nhưng kém với syntax/sequential info và ngữ cảnh phức tạp. Không handle QA differentiation vì bỏ qua thứ tự dài hạn, chỉ static word-level, không phù hợp text series curated.

  • Combination of the Amazon SageMaker BlazingText algorithm in Batch Skip-gram mode with a custom recurrent neural network (RNN)
    ❌ SAI: BlazingText Batch Skip-gram vẫn là static word embeddings, kết hợp RNN custom có thể cải thiện sequential nhưng không phải built-in solution đơn giản, phức tạp triển khai, và không đảm bảo capture QA context tốt như seq2seq/Object2Vec thuần. SageMaker ưu tiên end-to-end algorithms hơn custom combos cho task này.

🧠 Kết luận: Chọn seq2seq và Object2Vec để tối ưu hiệu suất, scalability trên SageMaker, tránh các Word2Vec (BlazingText) kém contextual. Test trên SageMaker notebook để verify embeddings! 🚀

Câu 179
A retail company wants to update its customer support system. The company wants to implement automatic routing of customer claims to different queues to prioritize the claims by category.
Currently, an operator manually performs the category assignment and routing. After the operator classifies and routes the claim, the company stores the claim's record in a central database. The claim's record includes the claim's category.
The company has no data science team or experience in the field of machine learning (ML). The company's small development team needs a solution that requires no ML expertise.
Which solution meets these requirements?
  1. A Export the database to a .csv file with two columns: claim_label and claim_text. Use the Amazon SageMaker Object2Vec algorithm and the .csv file to train a model. Use SageMaker to deploy the model to an inference endpoint. Develop a service in the application to use the inference endpoint to process incoming claims, predict the labels, and route the claims to the appropriate queue.
  2. B Export the database to a .csv file with one column: claim_text. Use the Amazon SageMaker Latent Dirichlet Allocation (LDA) algorithm and the .csv file to train a model. Use the LDA algorithm to detect labels automatically. Use SageMaker to deploy the model to an inference endpoint. Develop a service in the application to use the inference endpoint to process incoming claims, predict the labels, and route the claims to the appropriate queue.
  3. C Use Amazon Textract to process the database and automatically detect two columns: claim_label and claim_text. Use Amazon Comprehend custom classification and the extracted information to train the custom classifier. Develop a service in the application to use the Amazon Comprehend API to process incoming claims, predict the labels, and route the claims to the appropriate queue.
  4. D Export the database to a .csv file with two columns: claim_label and claim_text. Use Amazon Comprehend custom classification and the .csv file to train the custom classifier. Develop a service in the application to use the Amazon Comprehend API to process incoming claims, predict the labels, and route the claims to the appropriate queue.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một công ty bán lẻ muốn tự động hóa quy trình phân loại và routing (chuyển hướng) các yêu cầu khiếu nại (claims) của khách hàng vào các hàng đợi (queues) khác nhau dựa trên danh mục (category) để ưu tiên xử lý.

  • Tình huống hiện tại: Nhân viên hỗ trợ (operator) thủ công phân loại danh mục và routing, sau đó lưu trữ bản ghi claim vào cơ sở dữ liệu trung tâm (central database), bao gồm thông tin danh mục.
  • Yêu cầu chính:
    • Sử dụng dữ liệu lịch sử từ database (có label danh mục sẵn).
    • Không có đội ngũ data science hoặc kinh nghiệm ML, đội dev nhỏ → cần giải pháp không đòi hỏi chuyên môn ML (no ML expertise).
    • Giải pháp phải đơn giản, dễ triển khai, tự động predict danh mục cho claim mới và routing vào queue phù hợp.
      Chủ đề thuộc AWS Machine Learning (ML) services, tập trung vào phân loại văn bản (text classification) với dữ liệu có nhãn (labeled data), cập nhật theo AWS phiên bản 2024-2026 (Amazon Comprehend và SageMaker vẫn hỗ trợ các tính năng no-code/low-code).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng:
Export the database to a .csv file with two columns: claim_label and claim_text. Use Amazon Comprehend custom classification and the .csv file to train the custom classifier. Develop a service in the application to use the Amazon Comprehend API to process incoming claims, predict the labels, and route the claims to the appropriate queue.

Lý do chọn đáp án này 🏆:

  • Amazon Comprehend Custom Classification là dịch vụ fully managed, no-code/low-code của AWS, lý tưởng cho đội dev nhỏ không có kinh nghiệm ML. Chỉ cần export dữ liệu lịch sử thành CSV với 2 cột: claim_label (nhãn danh mục) và claim_text (nội dung claim) để train mô hình tùy chỉnh.
  • Quy trình đơn giản: Upload CSV → Train classifier (AWS tự động xử lý) → Deploy → Gọi API Comprehend từ app để predict label cho claim mới → Route tự động vào queue (ví dụ: SQS hoặc tương tự).
  • Đáp ứng yêu cầu no ML expertise: Không cần code model, tune hyperparameter, hay quản lý infrastructure. Hỗ trợ dữ liệu có nhãn sẵn từ database.
  • Cập nhật AWS 2026: Comprehend hỗ trợ multi-label classification và tích hợp dễ với Lambda, API Gateway cho serverless.
    📘 Tài liệu tham khảo: AWS Comprehend Custom Classification.

🛠️ Phân tích tất cả các phương án (đúng/sai)

Dưới đây là phân tích chi tiết từng phương án, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên yêu cầu no ML expertise và tính phù hợp với dữ liệu/database.

  • Phương án 1 (SAI):
    Export the database to a .csv file with two columns: claim_label and claim_text. Use the Amazon SageMaker Object2Vec algorithm and the .csv file to train a model. Use SageMaker to deploy the model to an inference endpoint. Develop a service in the application to use the inference endpoint to process incoming claims, predict the labels, and route the claims to the appropriate queue.
    ❌ Sai vì: Object2Vec là thuật toán SageMaker built-in cho embedding vector (word/sentence embedding), yêu cầu kiến thức ML sâu để preprocess data, chọn hyperparameter, và xử lý training (như unsupervised/semi-supervised). Không phù hợp no ML expertise, phức tạp hơn Comprehend. Dev team nhỏ sẽ khó implement endpoint và monitoring.

  • Phương án 2 (SAI):
    Export the database to a .csv file with one column: claim_text. Use the Amazon SageMaker Latent Dirichlet Allocation (LDA) algorithm and the .csv file to train a model. Use the LDA algorithm to detect labels automatically. Use SageMaker to deploy the model to an inference endpoint. Develop a service in the application to use the inference endpoint to process incoming claims, predict the labels, and route the claims to the appropriate queue.
    ❌ Sai vì: LDA là thuật toán unsupervised clustering (phân cụm topic từ text không nhãn), chỉ cần claim_text mà không dùng nhãn (label) từ database → không predict chính xác category cụ thể. SageMaker LDA đòi hỏi ML expertise để interpret topic và map sang label, không tự động "detect labels". Không khớp dữ liệu có nhãn sẵn.

  • Phương án 3 (SAI):
    Use Amazon Textract to process the database and automatically detect two columns: claim_label and claim_text. Use Amazon Comprehend custom classification and the extracted information to train the custom classifier. Develop a service in the application to use the Amazon Comprehend API to process incoming claims, predict the labels, and route the claims to the appropriate queue.
    ❌ Sai vì: Amazon Textract dùng để extract text từ tài liệu scan/PDF/hình ảnh (như hóa đơn), không xử lý database hoặc CSV. Database đã có dữ liệu structured (claim_label và claim_text), không cần extract → Textract thừa và sai ngữ cảnh. Phần còn lại (Comprehend) đúng nhưng bước đầu sai làm toàn bộ không khả thi.
    📘 Tài liệu: AWS Textract limits – chỉ cho document images.

  • Phương án 4 (ĐÚNG):
    (Như đã phân tích ở phần ✅ trên) – Hoàn hảo khớp yêu cầu! 🚀

Câu 180
A machine learning (ML) specialist is using Amazon SageMaker hyperparameter optimization (HPO) to improve a model's accuracy. The learning rate parameter is specified in the following HPO configuration:
{
    "Name": "learning_rate",
    "MaxValue": "0.0001",
    "MinValue": "0.1"
}

During the results analysis, the ML specialist determines that most of the training jobs had a learning rate between 0.01 and 0.1. The best result had a learning rate of less than 0.01. Training jobs need to run regularly over a changing dataset. The ML specialist needs to find a tuning mechanism that uses different learning rates more evenly from the provided range between MinValue and MaxValue.
Which solution provides the MOST accurate result?
  1. A Modify the HPO configuration as follows:
    {
        "Name": "learning_rate",
        "MaxValue": "0.0001",
        "MinValue": "0.1",
        "ScalingType": "ReverseLogarithmic"
    }
    Select the most accurate hyperparameter configuration form this HPO job.
  2. B Run three different HPO jobs that use different learning rates form the following intervals for MinValue and MaxValue while using the same number of training jobs for each HPO job: ✑ [0.01, 0.1] ✑ [0.001, 0.01] ✑ [0.0001, 0.001] Select the most accurate hyperparameter configuration form these three HPO jobs.
  3. C Modify the HPO configuration as follows:
    {
        "Name": "learning_rate",
        "MaxValue": "0.0001",
        "MinValue": "0.1",
        "ScalingType": "Logarithmic"
    }
    Select the most accurate hyperparameter configuration form this training job.
  4. D Run three different HPO jobs that use different learning rates form the following intervals for MinValue and MaxValue. Divide the number of training jobs for each HPO job by three: ✑ [0.01, 0.1] ✑ [0.001, 0.01] [0.0001, 0.001] Select the most accurate hyperparameter configuration form these three HPO jobs.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh việc tối ưu hóa siêu tham số (Hyperparameter Optimization - HPO) trong Amazon SageMaker, cụ thể là tham số learning rate cho mô hình học máy. Một chuyên gia ML đang sử dụng HPO để cải thiện độ chính xác mô hình. Cấu hình HPO hiện tại cho learning rate là:

{
    "Name": "learning_rate",
    "MaxValue": "0.0001",
    "MinValue": "0.1"
}

Lưu ý quan trọng: Trong cấu hình này, MinValue (0.1) lớn hơn MaxValue (0.0001), nhưng SageMaker hiểu phạm vi tìm kiếm là từ giá trị nhỏ nhất (0.0001) đến lớn nhất (0.1) trên thang logarit tự nhiên. Mặc định, SageMaker sử dụng ScalingType: "Linear" (phân bố đều theo thang tuyến tính), dẫn đến việc hầu hết các job huấn luyện (training jobs) tập trung ở khoảng 0.01 - 0.1 (phần cao của phạm vi). Tuy nhiên, kết quả tốt nhất (best result) có learning rate nhỏ hơn 0.01.

Vấn đề cần giải quyết: Dataset thay đổi thường xuyên, cần chạy HPO định kỳ, và phải tìm cơ chế tuning giúp phân bố learning rate đều hơn (more evenly) trên toàn bộ phạm vi từ MinValue đến MaxValue (tức 0.0001 đến 0.1), để khám phá tốt hơn các giá trị thấp mang lại kết quả tốt.

Câu hỏi yêu cầu giải pháp chính xác nhất (MOST accurate result) để đạt phân bố đều và cải thiện kết quả.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Modify the HPO configuration as follows:

{
    "Name": "learning_rate",
    "MaxValue": "0.0001",
    "MinValue": "0.1",
    "ScalingType": "Logarithmic"
}

Select the most accurate hyperparameter configuration form this training job.

Lý do:

  • Với ScalingType: "Logarithmic", SageMaker phân bố siêu tham số đều trên thang logarit (log-uniform distribution). Điều này rất phù hợp cho learning rate vì phạm vi trải dài nhiều bậc thang (orders of magnitude: từ 10^{-4} đến 10^{-1}).
  • Kết quả: Các giá trị thấp (như <0.01) được khám phá đều đặn hơn, tránh bias về phía giá trị cao như trong Linear scaling. Điều này giúp tìm kiếm hiệu quả hơn trên toàn phạm vi, đặc biệt khi best result ở vùng thấp, và phù hợp cho dataset thay đổi định kỳ chỉ cần một job HPO duy nhất.
  • Đây là giải pháp đơn giản, hiệu quả nhất theo best practices của AWS (không cần chạy nhiều job riêng lẻ).

📋 Phân tích tất cả các phương án (đúng/sai)

Dưới đây là phân tích chi tiết từng lựa chọn. Tôi giữ nguyên nội dung văn bản gốc bằng tiếng Anh, chỉ giải thích lý do đúng/sai bằng tiếng Việt với emoji minh họa:

  • ❌ Phương án SAI: Modify the HPO configuration as follows:

    {
        "Name": "learning_rate",
        "MaxValue": "0.0001",
        "MinValue": "0.1",
        "ScalingType": "ReverseLogarithmic"
    }
    

    Select the most accurate hyperparameter configuration form this HPO job.
    Giải thích: ReverseLogarithmic phân bố nghịch đảo logarit, tập trung nhiều samples hơn về phía giá trị cao (gần 0.1), làm bias tệ hơn so với Linear (vốn đã bias cao). Không giải quyết vấn đề cần phân bố đều, thậm chí làm giảm khám phá vùng thấp (<0.01) – trái ngược nhu cầu.

  • ❌ Phương án SAI: Run three different HPO jobs that use different learning rates form the following intervals for MinValue and MaxValue while using the same number of training jobs for each HPO job: ✑ [0.01, 0.1] ✑ [0.001, 0.01] ✑ [0.0001, 0.001] Select the most accurate hyperparameter configuration form these three HPO jobs.
    Giải thích: Chạy 3 HPO jobs riêng biệt với cùng số lượng training jobs mỗi job không đảm bảo phân bố đều trên toàn phạm vi gốc (0.0001-0.1). Việc chia nhỏ intervals làm tăng chi phí (3x jobs), phức tạp quản lý, và không tận dụng HPO tự động của SageMaker. Không "more evenly" trên range gốc, kém hiệu quả cho dataset thay đổi thường xuyên.

  • ✅ Phương án ĐÚNG (như đã chỉ rõ ở trên): Modify the HPO configuration as follows:

    {
        "Name": "learning_rate",
        "MaxValue": "0.0001",
        "MinValue": "0.1",
        "ScalingType": "Logarithmic"
    }
    

    Select the most accurate hyperparameter configuration form this training job.
    Giải thích: Như phần đáp án đúng đã trình bày – Logarithmic là lựa chọn tối ưu, phân bố đều trên log scale, khám phá tốt vùng thấp, chỉ cần 1 job, tiết kiệm tài nguyên và chính xác nhất.

  • ❌ Phương án SAI: Run three different HPO jobs that use different learning rates form the following intervals for MinValue and MaxValue. Divide the number of training jobs for each HPO job by three: ✑ [0.01, 0.1] ✑ [0.001, 0.01] [0.0001, 0.001] Select the most accurate hyperparameter configuration form these three HPO jobs.
    Giải thích: Tương tự phương án thứ 2, nhưng chia số jobs mỗi HPO bằng 1/3 làm giảm độ sâu tuning mỗi interval, dẫn đến kết quả kém chính xác hơn. Vẫn tăng chi phí tổng thể, phức tạp, và không phân bố "evenly" tự nhiên trên range gốc như Logarithmic làm chỉ trong 1 job.

📘 Tài liệu tham khảo (cập nhật đến 2026)

  • AWS SageMaker Documentation: Hyperparameter Tuning - Parameter Ranges – Chi tiết ScalingType: Linear/Logarithmic/ReverseLogarithmic (không thay đổi đến 2026).
  • AWS SageMaker Developer Guide: Automatic Model Tuning – Best practices cho learning rate trên log scale.
  • AWS re:Post & Exam Topics: Xác nhận Logarithmic là giải pháp chuẩn cho range rộng như learning rate (MLS-C01 exam blueprint, phiên bản 2024-2026).
  • Blog AWS: "Tuning SageMaker Hyperparameters Effectively" (2023 update, vẫn áp dụng).

🛠️ Khuyến nghị: Trong thực tế, luôn dùng Logarithmic cho learning rate/params có scale lớn để tránh bias! Nếu cần code ví dụ, hãy hỏi thêm nhé! 🚀