Ngân hàng đề — AWS Certified Machine Learning Specialty

Tìm thấy 371 câu.

Câu 51
An office security agency conducted a successful pilot using 100 cameras installed at key locations within the main office. Images from the cameras were uploaded to Amazon S3 and tagged using Amazon Rekognition, and the results were stored in Amazon ES. The agency is now looking to expand the pilot into a full production system using thousands of video cameras in its office locations globally. The goal is to identify activities performed by non-employees in real time
Which solution should the agency consider?
  1. A Use a proxy server at each local office and for each camera, and stream the RTSP feed to a unique Amazon Kinesis Video Streams video stream. On each stream, use Amazon Rekognition Video and create a stream processor to detect faces from a collection of known employees, and alert when non-employees are detected.
  2. B Use a proxy server at each local office and for each camera, and stream the RTSP feed to a unique Amazon Kinesis Video Streams video stream. On each stream, use Amazon Rekognition Image to detect faces from a collection of known employees and alert when non-employees are detected.
  3. C Install AWS DeepLens cameras and use the DeepLens_Kinesis_Video module to stream video to Amazon Kinesis Video Streams for each camera. On each stream, use Amazon Rekognition Video and create a stream processor to detect faces from a collection on each stream, and alert when non-employees are detected.
  4. D Install AWS DeepLens cameras and use the DeepLens_Kinesis_Video module to stream video to Amazon Kinesis Video Streams for each camera. On each stream, run an AWS Lambda function to capture image fragments and then call Amazon Rekognition Image to detect faces from a collection of known employees, and alert when non-employees are detected.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một cơ quan an ninh văn phòng đã thử nghiệm thành công pilot với 100 camera tại các vị trí chính, nơi hình ảnh được upload lên Amazon S3, gắn tag bằng Amazon Rekognition (phiên bản Image), và kết quả lưu vào Amazon ES (nay là Amazon OpenSearch Service). Bây giờ, họ muốn mở rộng quy mô lớn lên hàng nghìn camera video tại các văn phòng toàn cầu, với mục tiêu nhận diện real-time các hoạt động của non-employees (nhân viên không phải công ty) thông qua phân tích khuôn mặt từ bộ sưu tập nhân viên đã biết.

🔑 Thách thức chính:

  • Quy mô lớn: Hàng nghìn stream video real-time, cần giải pháp scalable, low-latency.
  • Real-time detection: Phân tích video liên tục, không phải ảnh tĩnh.
  • Tích hợp: Sử dụng camera hiện có (RTSP feed phổ biến), proxy local để xử lý stream.
  • Công nghệ cốt lõi: Cần dịch vụ AWS hỗ trợ video streaming và AI video analysis như Amazon Kinesis Video Streams kết hợp Amazon Rekognition Video với Stream Processor để detect faces real-time từ collection known faces.

Giải pháp phải tối ưu chi phí, scalable globally, tận dụng edge computing tại office và serverless processing trên AWS (cập nhật đến 2026: Rekognition Video Stream Processors hỗ trợ face detection real-time với độ chính xác cao, tích hợp Kinesis Video Streams fragments).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use a proxy server at each local office and for each camera, and stream the RTSP feed to a unique Amazon Kinesis Video Streams video stream. On each stream, use Amazon Rekognition Video and create a stream processor to detect faces from a collection of known employees, and alert when non-employees are detected.

Lý do 🛠️:

  • Hoàn hảo cho real-time video analysis với hàng nghìn camera: Proxy local (như GStreamer hoặc MediaLive) convert RTSP sang Kinesis Video Streams (KVS) unique stream/camera.
  • Amazon Rekognition Video + Stream Processor (tính năng native từ 2019, cập nhật 2026 hỗ trợ face search real-time) tự động detect faces từ face collection (lưu trữ known employees), so sánh và alert non-match ngay lập tức.
  • Scalable globally: KVS ingest unlimited streams, edge proxy giảm latency/bandwidth.
  • Không yêu cầu hardware mới, tận dụng camera RTSP hiện có.

📋 Phân tích tất cả các phương án (đúng/sai)

  • ✅ Use a proxy server at each local office and for each camera, and stream the RTSP feed to a unique Amazon Kinesis Video Streams video stream. On each stream, use Amazon Rekognition Video and create a stream processor to detect faces from a collection of known employees, and alert when non-employees are detected.
    Đúng vì: Như giải thích trên, đây là best practice AWS cho video surveillance real-time tại scale lớn. Rekognition Video Stream Processor xử lý video fragments tự động, face search từ collection với độ chính xác >99%, tích hợp alert via SNS/SQS. Hoàn toàn serverless, chi phí theo usage.

  • ❌ Use a proxy server at each local office and for each camera, and stream the RTSP feed to a unique Amazon Kinesis Video Streams video stream. On each stream, use Amazon Rekognition Image to detect faces from a collection of known employees and alert when non-employees are detected.
    Sai vì: Amazon Rekognition Image chỉ dành cho ảnh tĩnh (JPEG/PNG), không hỗ trợ video streams real-time. Không thể apply trực tiếp lên KVS stream mà không extract frames thủ công (tốn kém, không real-time). Rekognition Video mới là lựa chọn đúng cho video.

  • ❌ Install AWS DeepLens cameras and use the DeepLens_Kinesis_Video module to stream video to Amazon Kinesis Video Streams for each camera. On each stream, use Amazon Rekognition Video and create a stream processor to detect faces from a collection on each stream, and alert when non-employees are detected.
    Sai vì: AWS DeepLens đã bị deprecated từ 2020 (end-of-support chính thức, không khuyến nghị 2026). Yêu cầu thay thế toàn bộ hàng nghìn camera hiện có bằng DeepLens (không feasible, tốn kém). Dù KVS + Rekognition Video đúng, nhưng hardware cũ kỹ không scale globally.

  • ❌ Install AWS DeepLens cameras and use the DeepLens_Kinesis_Video module to stream video to Amazon Kinesis Video Streams for each camera. On each stream, run an AWS Lambda function to capture image fragments and then call Amazon Rekognition Image to detect faces from a collection of known employees, and alert when non-employees are detected.
    Sai vì: Kết hợp 2 vấn đề: (1) DeepLens deprecated, không thay thế camera; (2) Lambda extract fragments + Rekognition Image không real-time (cold starts, polling overhead), kém scalable với thousands streams (Lambda concurrency limits), tốn chi phí cao hơn Stream Processor native. Rekognition Image không tối ưu cho video.

📘 Tài liệu tham khảo (AWS Documentation cập nhật 2026)

Giải pháp đúng đảm bảo high availability, low latency cho production! 🚀 Nếu cần demo code Terraform/ CDK, hỏi thêm nhé!

Câu 52
A Marketing Manager at a pet insurance company plans to launch a targeted marketing campaign on social media to acquire new customers. Currently, the company has the following data in Amazon Aurora:
✑ Profiles for all past and existing customers
✑ Profiles for all past and existing insured pets
✑ Policy-level information
✑ Premiums received
✑ Claims paid
What steps should be taken to implement a machine learning model to identify potential new customers on social media?
  1. A Use regression on customer profile data to understand key characteristics of consumer segments. Find similar profiles on social media
  2. B Use clustering on customer profile data to understand key characteristics of consumer segments. Find similar profiles on social media
  3. C Use a recommendation engine on customer profile data to understand key characteristics of consumer segments. Find similar profiles on social media.
  4. D Use a decision tree classifier engine on customer profile data to understand key characteristics of consumer segments. Find similar profiles on social media.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh việc một Marketing Manager tại công ty bảo hiểm thú cưng muốn triển khai mô hình machine learning (ML) để xác định khách hàng tiềm năng mới trên social media. Dữ liệu hiện có lưu trữ trong Amazon Aurora bao gồm:

  • Profiles của tất cả khách hàng cũ và hiện tại (past and existing customers).
  • Profiles của thú cưng được bảo hiểm (insured pets).
  • Thông tin cấp policy (policy-level information).
  • Premiums đã nhận (tiền phí bảo hiểm).
  • Claims đã trả (tiền bồi thường).

Mục tiêu là hiểu đặc trưng chính (key characteristics) của các phân khúc khách hàng (consumer segments) từ dữ liệu này, sau đó tìm profiles tương tự trên social media để chạy chiến dịch marketing nhắm mục tiêu.

Đây là bài toán unsupervised learning điển hình, vì không có nhãn (label) sẵn để phân loại khách hàng tiềm năng. Chúng ta cần phân nhóm (clustering) khách hàng hiện tại dựa trên profiles để xác định "look-alike audiences", rồi áp dụng để tìm kiếm tương tự trên social media. AWS hỗ trợ qua Amazon SageMaker với các thuật toán clustering như K-Means, phù hợp triển khai end-to-end từ dữ liệu Aurora (export qua S3) 🛠️.

📘 Tài liệu tham khảo:

  • AWS SageMaker Documentation: Clustering Algorithms (cập nhật 2024-2026).
  • AWS ML Specialty Exam Guide (MLS-C01/DOP-C02): Phần "Exploratory Data Analysis & Unsupervised Learning".
  • AWS re:Post & Well-Architected Framework: ML Workloads (Lens: Secure & Operational Excellence).

✅ Đáp án đúng: Use clustering on customer profile data to understand key characteristics of consumer segments. Find similar profiles on social media

Lý do lựa chọn:

  • Clustering là thuật toán unsupervised learning lý tưởng để phân nhóm khách hàng dựa trên đặc trưng profiles (tuổi, thu nhập, loại thú cưng, lịch sử claims, v.v.) mà không cần label.
  • Kết quả tạo ra các consumer segments (nhóm khách hàng tương đồng), giúp hiểu key characteristics (ví dụ: nhóm trẻ tuổi nuôi chó lớn, hay claim cao).
  • Sau đó, sử dụng segments này làm embedding hoặc features để tìm profiles tương tự (look-alike) trên social media qua API như Facebook Custom Audiences hoặc SageMaker Feature Store kết hợp Amazon Personalize/SageMaker Canvas 🧩.
  • Trong AWS (2026), SageMaker BlazingText hoặc K-Means hỗ trợ scale lớn từ dữ liệu Aurora, tích hợp Lambda/ECS cho pipeline tự động. Đây là best practice cho customer segmentation trong marketing ML.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ✅ [ĐÚNG] Use clustering on customer profile data to understand key characteristics of consumer segments. Find similar profiles on social media
    Như đã giải thích ở trên: Clustering (K-Means/HDBSCAN trong SageMaker) tự động nhóm dữ liệu đa chiều từ profiles, premiums, claims → tạo segments → match look-alike trên social. Hoàn hảo cho bài toán không label! 🚀

  • ❌ [SAI] Use regression on customer profile data to understand key characteristics of consumer segments. Find similar profiles on social media
    Regression (Linear/XGBoost trong SageMaker) là supervised learning dùng để dự đoán giá trị liên tục (ví dụ: dự đoán premiums tương lai), không phù hợp phân nhóm segments. Nó cần target variable (label) và không tạo clusters tự nhiên, dẫn đến sai lệch khi tìm profiles tương tự.

  • ❌ [SAI] Use a recommendation engine on customer profile data to understand key characteristics of consumer segments. Find similar profiles on social media
    Recommendation engine (Amazon Personalize) tập trung gợi ý items/users dựa trên collaborative filtering (user-item interactions), không phải phân tích segments từ profiles tĩnh. Nó cần lịch sử tương tác lớn (ratings), không hiệu quả cho dữ liệu Aurora thuần túy và khó "hiểu key characteristics" mà không có behavior data.

  • ❌ [SAI] Use a decision tree classifier engine on customer profile data to understand key characteristics of consumer segments. Find similar profiles on social media
    Decision tree classifier (SageMaker XGBoost/Built-in) là supervised classification cần labels sẵn (ví dụ: "churn yes/no") để phân loại. Không có label ở đây, nên không thể dùng để khám phá segments unsupervised; chỉ tạo cây quyết định dựa trên giả định label, dẫn đến underfitting hoặc bias cao.

Kết luận: Chọn clustering để scale ML pipeline trên AWS một cách chính xác và hiệu quả! 🌟 Nếu cần code SageMaker notebook, hãy hỏi thêm nhé 🛠️.

Câu 53
A manufacturing company has a large set of labeled historical sales data. The manufacturer would like to predict how many units of a particular part should be produced each quarter.
Which machine learning approach should be used to solve this problem?
  1. A Logistic regression
  2. B Random Cut Forest (RCF)
  3. C Principal component analysis (PCA)
  4. D Linear regression
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một công ty sản xuất có bộ dữ liệu lịch sử bán hàng lớn đã được gắn nhãn (labeled), và họ muốn dự đoán số lượng đơn vị (units) của một bộ phận cụ thể cần sản xuất mỗi quý.
✅ Đây là bài toán regression (dự đoán giá trị liên tục - continuous value) dựa trên dữ liệu lịch sử, thường liên quan đến time series forecasting hoặc dự báo nhu cầu sản xuất.
🛠️ Trong AWS, vấn đề này có thể được giải quyết bằng Amazon SageMaker với các built-in algorithms cho regression tasks, sử dụng dữ liệu labeled để huấn luyện mô hình dự đoán số lượng sản xuất (ví dụ: qua Linear Learner algorithm).
📘 Nguồn tham khảo: AWS SageMaker Documentation - Built-in Algorithms (cập nhật 2024-2026): SageMaker Linear Learner.

✅ Đáp án đúng: Linear regression

Linear regression là lựa chọn đúng vì:

  • Bài toán yêu cầu dự đoán giá trị số liên tục (số lượng units sản xuất mỗi quý), đây là nhiệm vụ regression điển hình.
  • Linear regression mô hình hóa mối quan hệ tuyến tính giữa các biến đầu vào (dữ liệu lịch sử bán hàng labeled) và biến đầu ra (số lượng dự đoán).
  • Trong AWS SageMaker, Linear Learner hỗ trợ linear regression hiệu quả cho dữ liệu lớn, scalable trên cloud, và phù hợp với dữ liệu labeled.
    🧩 Nó đơn giản, nhanh, và là baseline tốt cho forecasting trước khi thử các mô hình phức tạp hơn như XGBoost hoặc DeepAR (cho time series).

📋 Giải thích chi tiết tất cả các phương án

  • ❌ [SAI] Logistic regression
    Logistic regression dùng cho phân loại (classification), dự đoán xác suất thuộc lớp rời rạc (binary/multiclass, ví dụ: 0/1). Không phù hợp vì output là số lượng units liên tục, không phải lớp. Trong SageMaker, nó là Linear Learner với binary classification mode.

  • ❌ [SAI] Random Cut Forest (RCF)
    RCF là thuật toán phát hiện bất thường (anomaly detection) unsupervised, dùng để tìm outliers trong dữ liệu (như doanh số bất thường). Không dùng cho dự đoán regression với dữ liệu labeled. Trong SageMaker, RCF dành cho monitoring, không phải forecasting.

  • ❌ [SAI] Principal component analysis (PCA)
    PCA là kỹ thuật giảm chiều dữ liệu (dimensionality reduction) unsupervised, giúp visualize hoặc preprocess dữ liệu bằng cách giữ các thành phần chính. Không phải mô hình dự đoán, chỉ hỗ trợ bước tiền xử lý trước regression. SageMaker hỗ trợ PCA cho feature engineering.

  • ✅ [ĐÚNG] Linear regression
    Như đã giải thích ở trên: Hoàn hảo cho regression task với dữ liệu labeled, dự đoán continuous values. SageMaker Linear Learner tối ưu hóa cho linear regression trên dữ liệu lớn, hỗ trợ GPU/CPU scaling (cập nhật 2026).

🛠️ Lời khuyên thực hành: Sử dụng SageMaker Studio để train Linear Learner, kết hợp với Amazon Forecast nếu cần time series nâng cao. Test trên dataset sales để validate RMSE thấp!
📘 Nguồn bổ sung: AWS ML Specialty Exam Guide (2024), Amazon SageMaker Algorithms.

Câu 54
A financial services company is building a robust serverless data lake on Amazon S3. The data lake should be flexible and meet the following requirements:
✑ Support querying old and new data on Amazon S3 through Amazon Athena and Amazon Redshift Spectrum.
✑ Support event-driven ETL pipelines
✑ Provide a quick and easy way to understand metadata
Which approach meets these requirements?
  1. A Use an AWS Glue crawler to crawl S3 data, an AWS Lambda function to trigger an AWS Glue ETL job, and an AWS Glue Data catalog to search and discover metadata.
  2. B Use an AWS Glue crawler to crawl S3 data, an AWS Lambda function to trigger an AWS Batch job, and an external Apache Hive metastore to search and discover metadata.
  3. C Use an AWS Glue crawler to crawl S3 data, an Amazon CloudWatch alarm to trigger an AWS Batch job, and an AWS Glue Data Catalog to search and discover metadata.
  4. D Use an AWS Glue crawler to crawl S3 data, an Amazon CloudWatch alarm to trigger an AWS Glue ETL job, and an external Apache Hive metastore to search and discover metadata.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc xây dựng một data lake serverless mạnh mẽ trên Amazon S3 cho công ty dịch vụ tài chính. Data lake cần linh hoạt và đáp ứng 3 yêu cầu chính:

  • Hỗ trợ truy vấn dữ liệu cũ và mới trên S3 thông qua Amazon Athena (query serverless trực tiếp trên S3) và Amazon Redshift Spectrum (query external data từ Redshift cluster).
  • Hỗ trợ ETL pipelines theo sự kiện (event-driven): Nghĩa là tự động kích hoạt xử lý dữ liệu khi có sự kiện xảy ra (ví dụ: file mới upload lên S3).
  • Cung cấp cách nhanh chóng và dễ dàng để hiểu metadata: Cần một catalog metadata tích hợp để khám phá, tìm kiếm schema dữ liệu.

🛠️ Giải pháp cốt lõi: Sử dụng AWS Glue (dịch vụ ETL serverless) làm trung tâm, vì nó tích hợp hoàn hảo với S3, Athena, Redshift Spectrum. Glue Crawler quét dữ liệu S3 để tự động tạo schema và lưu vào Glue Data Catalog (metastore chuẩn cho Athena/Redshift Spectrum). Event-driven dùng Lambda trigger từ S3 events để chạy Glue ETL jobs. Đây là kiến trúc serverless chuẩn theo best practices AWS đến năm 2026 (AWS Glue v4.0 hỗ trợ Spark 3.5, tối ưu hơn cho data lake).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use an AWS Glue crawler to crawl S3 data, an AWS Lambda function to trigger an AWS Glue ETL job, and an AWS Glue Data catalog to search and discover metadata.

Lý do:

  • ✅ Glue Crawler quét S3 → populate Glue Data Catalog → Athena/Redshift Spectrum query dễ dàng (hỗ trợ old/new data).
  • ✅ Lambda function nhận event từ S3 (qua EventBridge/S3 notifications) → trigger Glue ETL job serverless, event-driven hoàn hảo.
  • ✅ Glue Data Catalog là metastore native, dễ search/discover metadata qua console/CLI/API.
  • Toàn bộ serverless, scalable, không cần quản lý infra. Phù hợp 100% yêu cầu!

🔍 Phân tích tất cả các phương án (đúng/sai)

  • Phương án ĐÚNG ✅:
    Use an AWS Glue crawler to crawl S3 data, an AWS Lambda function to trigger an AWS Glue ETL job, and an AWS Glue Data catalog to search and discover metadata.
    Giải thích: Hoàn hảo như trên. Lambda + Glue ETL = event-driven chuẩn (S3 Event → Lambda → Glue Job). Glue Catalog tích hợp native với Athena/Redshift Spectrum, dễ khám phá metadata qua Glue Studio/Data Catalog UI. Best practice AWS Lake Formation/Glue.

  • Phương án SAI ❌:
    Use an AWS Glue crawler to crawl S3 data, an AWS Lambda function to trigger an AWS Batch job, and an external Apache Hive metastore to search and discover metadata.
    Giải thích: AWS Batch không phải ETL serverless cho data lake (cần EC2/Fargate, không event-driven mượt như Glue). External Hive metastore không tích hợp tốt với Athena/Redshift Spectrum (phải config thủ công, kém linh hoạt, không hỗ trợ auto-crawl như Glue Catalog).

  • Phương án SAI ❌:
    Use an AWS Glue crawler to crawl S3 data, an Amazon CloudWatch alarm to trigger an AWS Batch job, and an AWS Glue Data Catalog to search and discover metadata.
    Giải thích: CloudWatch Alarm chỉ trigger khi metric threshold (ví dụ CPU cao), không phải event-driven thực sự (không phản ứng ngay với S3 file mới). AWS Batch không phù hợp ETL data lake serverless. Glue Catalog OK nhưng tổng thể không khớp.

  • Phương án SAI ❌:
    Use an AWS Glue crawler to crawl S3 data, an Amazon CloudWatch alarm to trigger an AWS Glue ETL job, and an external Apache Hive metastore to search and discover metadata.
    Giải thích: CloudWatch Alarm không hỗ trợ event-driven từ S3 events (chỉ metric-based). External Hive metastore không native với Athena/Redshift Spectrum (cần EMR/ custom setup phức tạp, kém "quick/easy" cho metadata).

🧩 Kết luận: Chỉ phương án đầu tiên đáp ứng serverless, event-driven, metadata discovery một cách tối ưu. Sử dụng AWS Glue Lake Formation để nâng cao governance nếu scale lớn hơn! 🚀

Câu 55
A company's Machine Learning Specialist needs to improve the training speed of a time-series forecasting model using TensorFlow. The training is currently implemented on a single-GPU machine and takes approximately 23 hours to complete. The training needs to be run daily.
The model accuracy is acceptable, but the company anticipates a continuous increase in the size of the training data and a need to update the model on an hourly, rather than a daily, basis. The company also wants to minimize coding effort and infrastructure changes.
What should the Machine Learning Specialist do to the training solution to allow it to scale for future demand?
  1. A Do not change the TensorFlow code. Change the machine to one with a more powerful GPU to speed up the training.
  2. B Change the TensorFlow code to implement a Horovod distributed framework supported by Amazon SageMaker. Parallelize the training to as many machines as needed to achieve the business goals.
  3. C Switch to using a built-in AWS SageMaker DeepAR model. Parallelize the training to as many machines as needed to achieve the business goals.
  4. D Move the training to Amazon EMR and distribute the workload to as many machines as needed to achieve the business goals.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi tập trung vào việc cải thiện tốc độ huấn luyện (training) mô hình dự báo chuỗi thời gian (time-series forecasting) được xây dựng bằng TensorFlow trên một máy single-GPU, hiện mất khoảng 23 giờ mỗi lần chạy và cần chạy hàng ngày. 🔄 Công ty dự kiến dữ liệu huấn luyện sẽ tăng liên tục, cần cập nhật mô hình hàng giờ thay vì hàng ngày, đồng thời tối thiểu hóa nỗ lực code và thay đổi hạ tầng.

Mục tiêu chính là scale giải pháp training để đáp ứng nhu cầu tương lai: nhanh hơn, xử lý dữ liệu lớn hơn, chạy thường xuyên hơn, mà không cần viết lại code nhiều hoặc quản lý infra phức tạp. 🛠️ Đây là tình huống điển hình trong Amazon SageMaker, nơi hỗ trợ distributed training để phân tán workload lên nhiều GPU/máy.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Change the TensorFlow code to implement a Horovod distributed framework supported by Amazon SageMaker. Parallelize the training to as many machines as needed to achieve the business goals.

Lý do:
Horovod là framework phân tán huấn luyện (distributed training) được Amazon SageMaker hỗ trợ chính thức cho TensorFlow, cho phép parallelize training lên nhiều instance/GPU mà chỉ cần thay đổi ít code (thêm wrapper Horovod vào code TensorFlow hiện tại). Điều này giúp giảm thời gian từ 23 giờ xuống đáng kể (scale linearly với số máy), xử lý dữ liệu lớn, chạy hourly dễ dàng. SageMaker tự động quản lý infra (như Elastic Fabric Adapter - EFA cho multi-node), minimize coding effort và infra changes. Phù hợp hoàn hảo với yêu cầu "minimize coding effort and infrastructure changes". 🚀

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh:

  • ❌ [SAI] Do not change the TensorFlow code. Change the machine to one with a more powerful GPU to speed up the training.
    Phương án này chỉ nâng cấp GPU đơn lẻ (ví dụ p4d.24xlarge), giúp tăng tốc một phần nhưng không scale cho dữ liệu tăng lớn hoặc chạy hourly vì vẫn là single-instance. Không đáp ứng "continuous increase in data size" và "scale for future demand". Không cần distributed, nhưng giới hạn về tốc độ và chi phí cao cho GPU mạnh.

  • ✅ [ĐÚNG] Change the TensorFlow code to implement a Horovod distributed framework supported by Amazon SageMaker. Parallelize the training to as many machines as needed to achieve the business goals.
    Như đã giải thích ở trên: Horovod + SageMaker là giải pháp tối ưu, hỗ trợ multi-GPU/multi-node với code changes minimum (chỉ thêm hvd.DistributedOptimizer), tự động scale theo nhu cầu kinh doanh. SageMaker xử lý orchestration, checkpointing, và fault-tolerance. Hoàn hảo cho TensorFlow time-series models.

  • ❌ [SAI] Switch to using a built-in AWS SageMaker DeepAR model. Parallelize the training to as many machines as needed to achieve the business goals.
    DeepAR là algorithm built-in của SageMaker cho time-series forecasting, hỗ trợ distributed training, nhưng yêu cầu switch hoàn toàn code TensorFlow custom sang DeepAR (khác architecture và input format). Vi phạm "minimize coding effort" vì phải rewrite model, không giữ nguyên TensorFlow code hiện tại (accuracy đã acceptable).

  • ❌ [SAI] Move the training to Amazon EMR and distribute the workload to as many machines as needed to achieve the business goals.
    Amazon EMR dành cho big data processing (Spark/Hadoop), không optimized cho ML training TensorFlow như SageMaker. Cần setup cluster thủ công, install TensorFlow/Horovod, quản lý infra phức tạp (EC2, Spark MLlib kém cho deep learning). Không minimize infra changes, và kém hiệu quả hơn SageMaker cho distributed DL.

📘 Tài liệu tham khảo (cập nhật đến 2026)

Giải pháp này đảm bảo scale bền vững mà không phức tạp! 🌟

Câu 56
Which of the following metrics should a Machine Learning Specialist generally use to compare/evaluate machine learning classification models against each other?
  1. A Recall
  2. B Misclassification rate
  3. C Mean absolute percentage error (MAPE)
  4. D Area Under the ROC Curve (AUC)
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc chọn metric phù hợp nhất để một Machine Learning Specialist sử dụng khi so sánh và đánh giá các mô hình phân loại (classification models) trong machine learning trên AWS, đặc biệt với các dịch vụ như Amazon SageMaker.

  • Bối cảnh chính: Trong phân loại ML (ví dụ: binary hoặc multi-class classification), việc so sánh models cần một metric tổng quát, độc lập với ngưỡng phân loại (threshold), hoạt động tốt trên dữ liệu không cân bằng (imbalanced data), và phản ánh khả năng phân biệt giữa các lớp (classes). AWS khuyến nghị sử dụng các metric tiêu chuẩn trong SageMaker để đánh giá mô hình như BinaryClassification hoặc MulticlassClassification metrics (cập nhật theo AWS ML best practices đến 2026, với SageMaker Processing Jobs và Clarify hỗ trợ AUC làm metric chính cho model comparison).
  • Mục tiêu: Tìm metric tốt nhất để so sánh giữa các models khác nhau, không phải metric tốt nhất cho một task cụ thể.

📘 Tài liệu tham khảo:

✅ Đáp án đúng: Area Under the ROC Curve (AUC)

Lý do lựa chọn:

  • AUC (hay AUC-ROC) là metric tổng quát nhất để so sánh các mô hình phân loại, đo lường khả năng phân biệt giữa positive và negative classes trên toàn bộ ngưỡng (threshold-independent). Giá trị AUC từ 0 đến 1: 0.5 = random, 1.0 = perfect model.
  • 🛠️ Ưu điểm trên AWS: Trong SageMaker, AUC được tích hợp sẵn trong Model Monitor, Clarify, và Processing Jobs để so sánh hyperparameters tuning (Hyperparameter Tuning Jobs). Nó hoạt động xuất sắc với imbalanced datasets (phổ biến trong real-world AWS ML workloads như fraud detection).
  • Không bị ảnh hưởng bởi class imbalance, khác với accuracy, nên là lựa chọn hàng đầu để "compare/evaluate models against each other".

📊 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh, với giải thích chi tiết bằng tiếng Việt:

  • Recall ❌
    Sai vì: Recall (True Positive Rate - TPR) chỉ tập trung vào tỷ lệ phát hiện đúng positive cases, phụ thuộc vào threshold cụ thể và không đánh giá tổng thể model. Không phù hợp để so sánh models vì bỏ qua False Positives và Precision. Trong SageMaker, Recall chỉ dùng cho task cụ thể (như anomaly detection), không phải comparison chung.

  • Misclassification rate ❌
    Sai vì: Đây là 1 - Accuracy (tỷ lệ dự đoán sai), bị ảnh hưởng nặng bởi class imbalance (ví dụ: nếu 99% negative, accuracy cao nhưng model kém). Không độc lập threshold và kém trong so sánh models trên AWS datasets đa dạng. SageMaker khuyên tránh dùng cho imbalanced classification.

  • Mean absolute percentage error (MAPE) ❌
    Sai vì: MAPE là metric cho regression tasks (dự đoán số liên tục), tính lỗi phần trăm tuyệt đối. Hoàn toàn không áp dụng cho classification (output là categories). Trong SageMaker, MAPE dùng cho Regression algorithms như Linear Learner, không liên quan đến classification models.

  • Area Under the ROC Curve (AUC) ✅
    Đúng vì: Như đã giải thích ở trên, là metric chuẩn mực và phổ biến nhất cho việc so sánh classification models trên AWS (SageMaker, Forecast, Fraud Detector). Độc lập threshold, robust với imbalance, và được AWS cập nhật hỗ trợ đầy đủ đến 2026 với tích hợp MLflow và Model Registry.

🧩 Kết luận: AUC là lựa chọn tối ưu cho Machine Learning Specialist trên AWS khi cần so sánh models một cách công bằng và đáng tin cậy! Nếu deploy trên SageMaker, hãy dùng BinaryClassificationMetrics để track AUC tự động.

Câu 57
A company is running a machine learning prediction service that generates 100 TB of predictions every day. A Machine Learning Specialist must generate a visualization of the daily precision-recall curve from the predictions, and forward a read-only version to the Business team.
Which solution requires the LEAST coding effort?
  1. A Run a daily Amazon EMR workflow to generate precision-recall data, and save the results in Amazon S3. Give the Business team read-only access to S3.
  2. B Generate daily precision-recall data in Amazon QuickSight, and publish the results in a dashboard shared with the Business team.
  3. C Run a daily Amazon EMR workflow to generate precision-recall data, and save the results in Amazon S3. Visualize the arrays in Amazon QuickSight, and publish them in a dashboard shared with the Business team.
  4. D Generate daily precision-recall data in Amazon ES, and publish the results in a dashboard shared with the Business team.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào một dịch vụ dự đoán machine learning (ML) của công ty, tạo ra 100 TB dữ liệu dự đoán mỗi ngày 📊. Nhiệm vụ của Machine Learning Specialist là:

  • Tạo biểu đồ đường cong precision-recall hàng ngày (precision-recall curve) từ dữ liệu dự đoán này. Đây là một metric quan trọng trong ML để đánh giá hiệu suất mô hình phân loại, yêu cầu tính toán phức tạp trên lượng dữ liệu lớn (cần so sánh predictions với ground truth labels).
  • Chia sẻ phiên bản read-only (chỉ đọc) cho đội ngũ Business, dưới dạng visualization dễ hiểu, không phải raw data.
  • Yêu cầu chính: Giải pháp nào đòi hỏi ít nỗ lực coding nhất (LEAST coding effort) 🛠️, nghĩa là ưu tiên các dịch vụ AWS tự động hóa cao, ít phải viết script tùy chỉnh.

Với quy mô 100 TB/ngày, cần công cụ xử lý big data mạnh mẽ như Amazon EMR (dựa trên Spark/Hadoop) để tính toán precision-recall 🧮, sau đó visualize bằng tool BI không code như Amazon QuickSight. Kiến thức cập nhật đến 2026: QuickSight hỗ trợ trực tiếp visualize dữ liệu lớn từ S3 (bao gồm arrays/JSON), EMR tích hợp Step Functions/Lambda cho workflow tự động, và Amazon ES đã chuyển sang Amazon OpenSearch Service (từ 2021).

Đáp án đúng ✅:
Run a daily Amazon EMR workflow to generate precision-recall data, and save the results in Amazon S3. Visualize the arrays in Amazon QuickSight, and publish them in a dashboard shared with the Business team.

Lý do chọn đáp án đúng 💡:
Giải pháp này ít coding nhất vì:

  • EMR tự động hóa workflow hàng ngày (qua Spark jobs) để tính precision-recall trên 100 TB, lưu kết quả dưới dạng arrays (mảng dữ liệu precision/recall values) vào S3 – chỉ cần script EMR đơn giản hoặc SageMaker Processing nếu tích hợp.
  • QuickSight kết nối trực tiếp S3 (SPICE engine xử lý dữ liệu lớn), tự động vẽ precision-recall curve từ arrays mà không cần code (drag-and-drop charts) 📈, publish dashboard read-only chia sẻ với Business team qua email/link.
  • Tổng effort: EMR setup 1 lần + QuickSight no-code viz → tối ưu cho DevOps/ML ops.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên coding effort, khả năng xử lý 100 TB, và tính phù hợp visualize/read-only.

  • ❌ Run a daily Amazon EMR workflow to generate precision-recall data, and save the results in Amazon S3. Give the Business team read-only access to S3.
    Sai vì: EMR phù hợp tính toán data lớn và lưu S3 ✅, nhưng Business team chỉ có raw data (precision-recall values) ở S3 mà không có visualization 📊. Họ phải tự download, dùng Excel/Python để vẽ curve → tăng coding effort cao cho Business (vi phạm read-only viz). Không ít effort nhất.

  • ❌ Generate daily precision-recall data in Amazon QuickSight, and publish the results in a dashboard shared with the Business team.
    Sai vì: QuickSight là tool BI no-code viz tuyệt vời cho dashboard 🖥️, nhưng không thể generate precision-recall data trực tiếp trên 100 TB raw predictions. QuickSight chỉ query/analyze data đã có (từ S3/ Athena), không xử lý compute ML metrics phức tạp/big data → cần custom code ETL trước, tăng effort coding nhiều so với dùng EMR.

  • ✅ Run a daily Amazon EMR workflow to generate precision-recall data, and save the results in Amazon S3. Visualize the arrays in Amazon QuickSight, and publish them in a dashboard shared with the Business team.
    Đúng vì: Kết hợp hoàn hảo EMR cho compute heavy (ít code Spark job) + QuickSight no-code viz arrays từ S3 → curve tự động, dashboard read-only share dễ dàng. Ít effort nhất, scalable đến 2026 với QuickSight ML Insights mới.

  • ❌ Generate daily precision-recall data in Amazon ES, and publish the results in a dashboard shared with the Business team.
    Sai vì: Amazon ES (nay OpenSearch Service) dành cho search/log analytics 🔍, không tối ưu compute precision-recall trên 100 TB (ingest data tốn kém, cần Lambda/PPL queries phức tạp). Viz dashboard qua Kibana có, nhưng coding effort cao để transform data → không phải least effort.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

Giải pháp này phù hợp DevOps best practices: Automate với EventBridge + EMR Steps, monitor bằng CloudWatch! 🚀

Câu 58
A Machine Learning Specialist is preparing data for training on Amazon SageMaker. The Specialist is using one of the SageMaker built-in algorithms for the training. The dataset is stored in .CSV format and is transformed into a numpy.array, which appears to be negatively affecting the speed of the training.
What should the Specialist do to optimize the data for training on SageMaker?
  1. A Use the SageMaker batch transform feature to transform the training data into a DataFrame.
  2. B Use AWS Glue to compress the data into the Apache Parquet format.
  3. C Transform the dataset into the RecordIO protobuf format.
  4. D Use the SageMaker hyperparameter optimization feature to automatically optimize the data.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS SageMaker

📘 Nội dung câu hỏi:
Câu hỏi mô tả tình huống một Machine Learning Specialist đang chuẩn bị dữ liệu huấn luyện (training data) cho mô hình trên Amazon SageMaker, sử dụng một trong những built-in algorithms của SageMaker (như XGBoost, Linear Learner, v.v.). Dữ liệu gốc ở định dạng .CSV và đã được chuyển đổi thành numpy.array, dẫn đến tình trạng tốc độ huấn luyện bị ảnh hưởng tiêu cực (chậm hơn).
🛠️ Vấn đề cốt lõi: SageMaker yêu cầu dữ liệu đầu vào ở định dạng tối ưu để xử lý hiệu quả, đặc biệt với built-in algorithms. Định dạng numpy.array không phải là chuẩn input cho training job (SageMaker mong đợi dữ liệu từ S3 dưới dạng file), và việc chuyển từ CSV sang array có thể gây overhead về memory/load time. Câu hỏi yêu cầu giải pháp tối ưu hóa dữ liệu để tăng tốc độ training, tận dụng các tính năng native của SageMaker như Pipe mode (streaming data từ S3 mà không cần load toàn bộ vào memory).

✅ Đáp án đúng:
Transform the dataset into the RecordIO protobuf format.
Lý do lựa chọn: Đây là định dạng được AWS khuyến nghị chính thức cho built-in algorithms của SageMaker để tối ưu tốc độ và memory usage. RecordIO protobuf (Protocol Buffers wrapped in RecordIO) cho phép Pipe mode – streaming dữ liệu trực tiếp từ S3 vào training instance mà không cần tải toàn bộ dataset vào disk/memory trước, giảm thời gian khởi động training lên đến 40-50% so với File mode (CSV/numpy). Numpy.array không tương thích trực tiếp; phải convert sang protobuf trước khi upload S3. Điều này đặc biệt hiệu quả với dataset lớn từ CSV. (Cập nhật đến 2026: Vẫn là best practice theo SageMaker docs).

📚 Tài liệu tham khảo:

🔍 Phân tích tất cả các phương án (đúng/sai)

  • ❌ Use the SageMaker batch transform feature to transform the training data into a DataFrame.
    Phương án này sai vì batch transform dùng để inference (dự đoán trên dữ liệu đã trained), không phải chuẩn bị training data. Nó transform input thành output predictions, không liên quan đến việc optimize training speed từ CSV/numpy. DataFrame (Pandas) còn gây overhead memory cao hơn, không giải quyết vấn đề chậm training.

  • ❌ Use AWS Glue to compress the data into the Apache Parquet format.
    Phương án này sai một phần đúng nhưng không tối ưu. AWS Glue có thể ETL và compress CSV thành Parquet ( columnar format, tiết kiệm storage ~75%), nhưng SageMaker built-in algorithms không hỗ trợ Parquet làm input trực tiếp cho training (chỉ hỗ trợ với custom algos hoặc vài built-in như BlazingText qua Data Wrangler). Parquet tốt cho query/analytics (Athena/Glue), nhưng không accelerate training như RecordIO protobuf với Pipe mode.

  • ✅ Transform the dataset into the RecordIO protobuf format.
    Phương án này đúng như đã giải thích ở trên. Sử dụng công cụ như numpy.array_to protobuf (SageMaker Python SDK) để convert, upload S3, và enable Pipe mode trong training job. Giảm thời gian data loading đáng kể, đặc biệt dataset lớn >10GB. Best practice cho built-in algos.

  • ❌ Use the SageMaker hyperparameter optimization feature to automatically optimize the data.
    Phương án này sai hoàn toàn vì Hyperparameter Optimization (HPO) dùng để tối ưu hyperparameters của model (như learning rate, epochs), không phải optimize dữ liệu input. Nó chạy nhiều training jobs parallel để tìm best config model, nhưng không fix vấn đề format data chậm (numpy.array vẫn gây bottleneck).

Câu 59
A Machine Learning Specialist is required to build a supervised image-recognition model to identify a cat. The ML Specialist performs some tests and records the following results for a neural network-based image classifier:
Total number of images available = 1,000
Test set images = 100 (constant test set)
The ML Specialist notices that, in over 75% of the misclassified images, the cats were held upside down by their owners.
Which techniques can be used by the ML Specialist to improve this specific test error?
  1. A Increase the training data by adding variation in rotation for training images.
  2. B Increase the number of epochs for model training
  3. C Increase the number of layers for the neural network.
  4. D Increase the dropout rate for the second-to-last layer.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi thuộc chủ đề Machine Learning trên AWS, cụ thể là xây dựng mô hình nhận diện hình ảnh (image classification) sử dụng neural network trên dịch vụ Amazon SageMaker (phiên bản cập nhật mới nhất 2026, hỗ trợ tích hợp data augmentation qua SageMaker Processing và Training Jobs).

Một Machine Learning Specialist cần xây dựng mô hình supervised learning để nhận diện mèo (cat identification). Dữ liệu: Tổng 1.000 ảnh, test set cố định 100 ảnh. Kết quả test cho thấy hơn 75% lỗi phân loại sai xảy ra khi mèo bị chủ cầm ngược đầu (upside down).

Vấn đề cốt lõi: Mô hình thiếu khả năng tổng quát hóa (generalization) với biến thể xoay ảnh (rotation variation), dẫn đến test error cụ thể cao ở trường hợp này. Câu hỏi yêu cầu kỹ thuật cải thiện lỗi test cụ thể này, không phải lỗi tổng quát. 🛠️ Đây là ví dụ kinh điển về data augmentation để tăng tính rotation invariance trong computer vision.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Increase the training data by adding variation in rotation for training images.

Lý do:

  • Vấn đề chính là 75% misclassified images do mèo bị cầm ngược (xoay 180 độ). Giải pháp trực tiếp là data augmentation bằng cách thêm biến thể xoay (rotation) vào dữ liệu huấn luyện, giúp mô hình học được invariance với góc xoay.
  • Trên SageMaker (2026), sử dụng SageMaker Data Wrangler hoặc Processing Jobs với thư viện như Albumentations/TensorFlow/Keras để tự động rotate ảnh (ví dụ: RandomRotation(180)). Điều này tăng kích thước dữ liệu hiệu quả mà không cần thu thập ảnh mới, giảm test error cụ thể đáng kể mà không gây overfit.
  • Kết quả: Mô hình robust hơn với real-world variations. 📈

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên best practices AWS SageMaker cho image classification (không giải quyết trực tiếp vấn đề rotation-specific error).

  • ✅ Increase the training data by adding variation in rotation for training images.
    Đúng: Như giải thích trên, đây là kỹ thuật data augmentation chuẩn (AWS khuyến nghị trong SageMaker JumpStart models và Processing). Trực tiếp target lỗi upside-down bằng cách simulate rotation trong training set, cải thiện accuracy trên test set cố định. Không làm phức tạp mô hình.

  • ❌ Increase the number of epochs for model training
    Sai: Tăng epochs chỉ giúp mô hình hội tụ tốt hơn với dữ liệu hiện tại, nhưng không giải quyết rotation variation. Có nguy cơ overfitting (học thuộc lòng training data), làm test error tệ hơn trên biến thể mới (upside-down cats). SageMaker Training Jobs theo dõi qua Early Stopping để tránh.

  • ❌ Increase the number of layers for the neural network.
    Sai: Thêm layers làm mô hình sâu hơn (deeper network), tăng capacity học feature phức tạp, nhưng không target rotation-specific error. Có thể dẫn đến vanishing gradients hoặc cần tuning hyperparams nhiều hơn (qua SageMaker Hyperparameter Tuning). Không hiệu quả cho vấn đề data distribution mismatch.

  • ❌ Increase the dropout rate for the second-to-last layer.
    Sai: Dropout giúp regularization chống overfitting bằng cách random drop neurons, nhưng không tạo rotation invariance. Nó chỉ cải thiện generalization tổng quát, không simulate biến thể upside-down. Trong SageMaker, dropout là hyperparam tuning, nhưng không phải giải pháp chính cho data bias này.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

  • Amazon SageMaker Documentation: Data Augmentation for Image Classification – Hướng dẫn rotation augmentation với Keras/TensorFlow.
  • AWS Best Practices: Improving Model Robustness with Augmentation – Case study tương tự với rotation.
  • SageMaker JumpStart: Models pre-trained với built-in augmentation (e.g., ResNet với RandomRotation).
  • AWS ML Specialty Exam Guide: Nhấn mạnh data augmentation cho test error do variations (domain shift).

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần demo code SageMaker, hãy hỏi thêm.

Câu 60
A Machine Learning Specialist needs to be able to ingest streaming data and store it in Apache Parquet files for exploration and analysis.
Which of the following services would both ingest and store this data in the correct format?
  1. A AWS DMS
  2. B Amazon Kinesis Data Streams
  3. C Amazon Kinesis Data Firehose
  4. D Amazon Kinesis Data Analytics
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào một Chuyên gia Machine Learning (Machine Learning Specialist) cần thu thập dữ liệu streaming (ingest streaming data) và lưu trữ nó dưới dạng file Apache Parquet để phục vụ cho việc khám phá và phân tích dữ liệu (exploration and analysis).

  • Yêu cầu chính: Dịch vụ AWS phải vừa ingest (thu thập dữ liệu thời gian thực) vừa store (lưu trữ) dữ liệu đúng định dạng Parquet (một định dạng columnar storage hiệu quả cho phân tích lớn, thường dùng với Athena hoặc Glue).
  • Bối cảnh: Dữ liệu streaming thường đến từ nguồn như IoT, logs, metrics... và cần xử lý để lưu vào S3 hoặc tương tự, hỗ trợ ML workflows như SageMaker hoặc EMR.
  • Phiên bản AWS mới nhất (2026): AWS vẫn ưu tiên Kinesis family cho streaming, với Firehose hỗ trợ native Parquet conversion qua data transformation (Lambda) và delivery trực tiếp vào S3.

🛠️ Đáp án đúng: Amazon Kinesis Data Firehose ✅
Lý do lựa chọn: Kinesis Data Firehose là dịch vụ serverless chuyên ingest streaming data từ nguồn như Kinesis Streams, agents, hoặc HTTP, sau đó tự động transform (qua Lambda) và lưu trữ trực tiếp vào S3 dưới dạng Parquet mà không cần code phức tạp. Nó hỗ trợ compression, buffering, và format conversion (bao gồm Parquet) ngay từ 2020 và cập nhật đến 2026 với tích hợp tốt hơn với Glue Schema Registry. Điều này lý tưởng cho ML exploration vì Parquet dễ query bằng Athena.

📋 Giải thích tất cả các phương án (Đúng/Sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên khả năng vừa ingest vừa store đúng định dạng Parquet:

  • [SAI] AWS DMS ❌
    Giải thích: AWS DMS (Database Migration Service) dùng để migrate dữ liệu giữa databases (on-prem sang AWS hoặc giữa RDS/DynamoDB), không hỗ trợ streaming real-time hay chuyển đổi định dạng Parquet. Nó chỉ copy dữ liệu thô (CSV/JSON), phù hợp cho batch migration chứ không phải ingest streaming cho ML analysis.

  • [SAI] Amazon Kinesis Data Streams ❌
    Giải thích: Kinesis Data Streams chỉ ingest và lưu trữ streaming data thô (raw records) trong shards với retention 24h-365 ngày, nhưng không tự động store vào Parquet hay transform định dạng. Người dùng phải dùng consumer riêng (như Lambda/Flink) để xử lý và đẩy sang S3, không "both ingest and store" một cách native.

  • [ĐÚNG] Amazon Kinesis Data Firehose ✅
    Giải thích: Như đã nêu, Firehose hoàn hảo cho yêu cầu: Ingest từ nhiều nguồn, transform record-to-record qua Lambda (convert sang Parquet), buffer data, và delivery trực tiếp vào S3 dưới Parquet format với partitioning tự động. Hỗ trợ error handling và monitoring qua CloudWatch, lý tưởng cho ML pipelines.

  • [SAI] Amazon Kinesis Data Analytics ❌
    Giải thích: Kinesis Data Analytics (nay là Amazon Managed Service for Apache Flink) dùng để xử lý streaming SQL/Flink applications (aggregate, filter), nhưng không store dữ liệu trực tiếp vào Parquet. Output chỉ gửi sang sinks như Kinesis Streams/Firehose/S3 (raw), yêu cầu thêm bước transform riêng.

📘 Tài liệu tham khảo (AWS Documentation mới nhất - 2026)

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm ví dụ code hoặc architecture diagram, hãy hỏi nhé!