Ngân hàng đề — AWS Certified Machine Learning Engineer Associate
Tìm thấy 635 câu.
A company is concerned about potential Distributed Denial of Service (DDoS) attacks on its application, which has been targeted in the past. They require advanced protection, including automatic mitigation for complex DDoS threats and access to AWS’s DDoS Response Team (DRT) for real-time support during incidents.
Which AWS service should they use to meet these requirements?
-
A
AWS Shield Standard
-
B
AWS Shield Advanced
-
C
AWS WAF
-
D
Amazon CloudWatch
Xem giải thích
Đáp án
B — AWS Shield Advanced
Vì sao đúng
Đề nêu hai điều kiện: đã từng bị tấn công và cần bảo vệ nâng cao. Shield Advanced cho phát hiện và giảm thiểu tấn công ở tầng ứng dụng, đội ứng cứu DDoS trực 24/7, báo cáo chi tiết về từng đợt tấn công, và bảo vệ chi phí — hoàn tiền phần phí tăng thêm do tấn công gây ra. Nó cũng bao gồm AWS WAF không tính phí riêng.
Vì sao các phương án khác sai
- A. Shield Standard — bật sẵn miễn phí cho mọi tài khoản, nhưng chỉ chống tấn công phổ biến ở tầng mạng và tầng vận chuyển.
- C. AWS WAF — lọc yêu cầu HTTP độc hại theo luật; hữu ích và thường dùng cùng Shield, nhưng tự nó không phải giải pháp chống DDoS toàn diện.
- D. CloudWatch — cho biết đang bị tấn công, không ngăn được gì.
An organization needs to ensure that its AWS resources comply with regulatory standards. They decide to implement AWS Config to monitor and manage the configuration of these resources. The organization also requires automated alerts for non-compliant resources and periodic compliance checks.
Which combination of AWS Config features should they enable to meet these requirements? (Choose TWO)
-
A
Configuration Items
-
B
AWS Config Rules with periodic evaluation
-
C
Configuration Stream
-
D
Custom rules created with AWS Lambda
-
E
Real-time proactive evaluation mode
Xem giải thích
Đáp án
B và D — quy tắc Config đánh giá định kỳ, và quy tắc tự viết bằng Lambda
Vì sao đúng
Config kiểm tra tuân thủ bằng rule, và đề cần cả hai loại:
- B. Quy tắc đánh giá định kỳ — chạy theo chu kỳ bạn đặt để rà soát toàn bộ tài nguyên, kể cả những thứ lâu rồi không ai đụng tới. Có sẵn hàng trăm quy tắc quản lý.
- D. Quy tắc tự viết bằng Lambda — khi tiêu chuẩn của công ty không khớp quy tắc sẵn có, bạn viết hàm Lambda nhận cấu hình tài nguyên và trả về tuân thủ hay không.
Vì sao các phương án khác sai
- A. Configuration Items — là bản ghi trạng thái của một tài nguyên tại một thời điểm; chúng là dữ liệu để đánh giá, không phải cơ chế đánh giá.
- C. Configuration Stream — luồng thông báo khi có thay đổi, cũng là dữ liệu chứ không đánh giá.
- E. Chế độ đánh giá chủ động thời gian thực — Config có đánh giá proactive cho một số loại tài nguyên trước khi tạo, nhưng cách mô tả này không phải một chế độ có thật.
A company wants to use AWS Config to maintain a detailed history of configuration changes across all of its resources for auditing purposes. They need to ensure the history of these changes is available long-term for regulatory requirements.
Which approach will allow them to retain and access this history over an extended period?
-
A
Enable configuration items and configure periodic evaluations.
-
B
Use configuration snapshots and store them in Amazon S3.
-
C
Enable AWS Config Rules and use CloudTrail to track changes.
-
D
Use a configuration stream and set up an Amazon RDS database for storage.
Xem giải thích
Đáp án
B — Dùng configuration snapshot và lưu vào Amazon S3
Vì sao đúng
Đề cần lịch sử chi tiết phục vụ kiểm toán, tức là dữ liệu phải giữ lâu và lấy ra được. Config tạo snapshot — ảnh chụp toàn bộ cấu hình tài nguyên tại một thời điểm — và đẩy vào bucket S3 mà bạn chỉ định. Từ đó áp được chính sách vòng đời, khoá ghi đè, và truy vấn bằng Athena khi cần dựng lại hiện trạng của một ngày cụ thể.
Vì sao các phương án khác sai
- A. Bật configuration item kèm đánh giá định kỳ — configuration item là dữ liệu nền, nhưng không tự lo phần lưu trữ lâu dài.
- C. Config Rules cộng CloudTrail — rule kiểm tra tuân thủ, CloudTrail ghi lời gọi API; không cái nào là kho lịch sử cấu hình.
- D. Configuration stream lưu vào RDS — stream chỉ phát thông báo; và dựng một CSDL quan hệ để chứa là tự chuốc việc vận hành.
A company is building a web-based AI application by using Amazon SageMaker. The application will provide the following capabilities and features: ML experimentation, training, a central model registry, model deployment, and model monitoring.
The application must ensure secure and isolated use of training data during the ML lifecycle. The training data is stored in Amazon S3.
The company needs to use the central model registry to manage different versions of models in the application.
Which action will meet this requirement with the LEAST operational overhead?
- A Create a separate Amazon Elastic Container Registry (Amazon ECR) repository for each model.
- B Use Amazon Elastic Container Registry (Amazon ECR) and unique tags for each model version.
- C Use the SageMaker Model Registry and model groups to catalog the models.
- D Use the SageMaker Model Registry and unique tags for each model version.
Xem giải thích
🧩 Giải thích chi tiết nội dung câu hỏi
Câu hỏi thuộc case study về một công ty đang xây dựng ứng dụng web-based AI sử dụng Amazon SageMaker. Ứng dụng cần hỗ trợ các tính năng: thử nghiệm ML (experimentation), huấn luyện mô hình (training), model registry trung tâm, triển khai mô hình (deployment), và giám sát mô hình (monitoring). 📊
Yêu cầu chính:
- Đảm bảo dữ liệu huấn luyện an toàn và cô lập trong suốt vòng đời ML (lưu trữ ở Amazon S3). 🔒
- Sử dụng central model registry để quản lý các phiên bản khác nhau của mô hình (model versions).
- Mục tiêu: Chọn hành động đáp ứng yêu cầu với operational overhead thấp nhất (LEAST operational overhead) – nghĩa là giải pháp đơn giản, tự động hóa cao, không cần quản lý thủ công nhiều. ⚡
Bối cảnh SageMaker: SageMaker cung cấp đầy đủ lifecycle ML end-to-end, với Model Registry là tính năng native để catalog, version, và quản lý models một cách tập trung. Câu hỏi tập trung vào việc chọn công cụ phù hợp nhất cho registry mà không cần custom setup phức tạp. 🛠️
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use the SageMaker Model Registry and model groups to catalog the models.
Lý do:
- SageMaker Model Registry là tính năng built-in của SageMaker (từ năm 2020, cập nhật liên tục đến 2026), cho phép tạo model groups để nhóm các versions của cùng một model. Nó tự động catalog, track lineage, approve/stage models, và tích hợp seamless với S3 cho dữ liệu an toàn (qua VPC, IAM policies). 📘
- Least operational overhead: Không cần quản lý infrastructure thủ công (như ECR repos hay tags), chỉ cần API calls hoặc Studio UI. Hỗ trợ bias detection, explainability, monitoring tự động.
- Phù hợp hoàn hảo với yêu cầu central registry và secure isolated data (S3 integration với encryption, access controls). Tiết kiệm chi phí và thời gian so với custom solutions. 🚀
❌ Phân tích tất cả các phương án (đúng/sai)
-
❌ Create a separate Amazon Elastic Container Registry (Amazon ECR) repository for each model.
Phương án này sai vì ECR chỉ lưu Docker images (containerized models), không phải central model registry native cho ML lifecycle. Tạo repo riêng cho mỗi model yêu cầu operational overhead cao: quản lý hàng trăm repos thủ công, tagging versions, scan vulnerabilities, lifecycle policies. Không hỗ trợ trực tiếp model lineage hay S3 data isolation trong SageMaker context. Phù hợp cho container deploy, không phải ML registry. 🗑️ -
❌ Use Amazon Elastic Container Registry (Amazon ECR) and unique tags for each model version.
Phương án này sai vì ECR với tags chỉ quản lý image versions cơ bản, thiếu model-specific features như approval workflows, lineage tracking, hay integration với SageMaker experiments. Overhead cao do cần script tự động tag/push/pull, không central hóa (vẫn phân tán theo tags), và không đảm bảo secure isolation cho S3 training data. Không phải giải pháp ML-native. 🔄 -
✅ Use the SageMaker Model Registry and model groups to catalog the models.
Phương án này đúng như đã giải thích ở trên. Đây là best practice AWS recommend, fully managed, scalable đến hàng triệu models/versions mà không cần ops effort. Tích hợp trực tiếp với SageMaker Studio, Pipelines, và Monitoring. 💯 -
❌ Use the SageMaker Model Registry and unique tags for each model version.
Phương án này sai vì SageMaker Model Registry không dùng tags cho versioning chính thức (tags chỉ metadata phụ). Thay vào đó, dùng model groups với built-in versioning (auto-generated). Dùng tags sẽ tạo overhead tự quản lý (custom scripts), mất lợi ích native như automatic staging/approval. Không "least overhead". 🏷️
📚 Tài liệu tham khảo (cập nhật mới nhất AWS đến 2026)
- AWS SageMaker Developer Guide: Model Registry – Chi tiết model groups và versioning.
- AWS Well-Architected Framework - ML Lens: Khuyến nghị dùng Model Registry cho least overhead.
- SageMaker Updates (re:Invent 2024/2025): Tích hợp AI governance, vẫn giữ Model Registry core (không thay đổi lớn).
- Exam DOP-C02 Blueprint: Topic SageMaker operations và ML lifecycle management.
Hy vọng phân tích giúp bạn nắm vững! Nếu cần đào sâu thêm, hỏi nhé. 🌟
A company is building a web-based AI application by using Amazon SageMaker. The application will provide the following capabilities and features: ML experimentation, training, a central model registry, model deployment, and model monitoring.
The application must ensure secure and isolated use of training data during the ML lifecycle. The training data is stored in Amazon S3.
The company is experimenting with consecutive training jobs.
How can the company MINIMIZE infrastructure startup times for these jobs?
- A Use Managed Spot Training.
- B Use SageMaker managed warm pools.
- C Use SageMaker Training Compiler.
- D Use the SageMaker distributed data parallelism (SMDDP) library.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi thuộc Case Study về một công ty đang xây dựng ứng dụng AI dựa trên web sử dụng Amazon SageMaker. Ứng dụng cung cấp các tính năng chính: thử nghiệm ML (ML experimentation), huấn luyện mô hình (training), kho lưu trữ mô hình trung tâm (central model registry), triển khai mô hình (model deployment), và giám sát mô hình (model monitoring).
🛡️ Yêu cầu bảo mật: Ứng dụng phải đảm bảo sử dụng dữ liệu huấn luyện an toàn và cô lập trong toàn bộ vòng đời ML (ML lifecycle). Dữ liệu huấn luyện được lưu trữ trong Amazon S3.
🚀 Vấn đề cốt lõi: Công ty đang thử nghiệm với các công việc huấn luyện liên tiếp (consecutive training jobs), và họ cần GIẢM THỦI GIAN KHỞI ĐỘNG CƠ SỞ HẠ TẦNG (MINIMIZE infrastructure startup times) cho những job này.
Mục tiêu chính: Tập trung vào việc tối ưu hóa thời gian khởi động tài nguyên (như EC2 instances) cho các training job chạy liên tục, giúp giảm độ trễ khi bắt đầu job mới mà không ảnh hưởng đến bảo mật dữ liệu S3 (vì SageMaker tự động xử lý isolation qua VPC, IAM roles, v.v.).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use SageMaker managed warm pools.
🛠️ Lý do chi tiết:
SageMaker managed warm pools là tính năng cho phép giữ lại các instance huấn luyện (training instances) ở trạng thái "ấm" (warm) sau khi job hoàn thành, thay vì shutdown hoàn toàn. Khi job tiếp theo khởi động, hệ thống reuse các instance này ngay lập tức, giảm đáng kể thời gian khởi động cơ sở hạ tầng (cold start time có thể giảm từ 10-15 phút xuống chỉ vài giây). Điều này lý tưởng cho consecutive training jobs trong giai đoạn thử nghiệm. Tính năng được SageMaker quản lý hoàn toàn (managed), hỗ trợ tích hợp với S3 data và đảm bảo isolation. Đây là giải pháp chính thức và tối ưu nhất từ AWS để minimize startup times (cập nhật mới nhất AWS SageMaker 2024-2026).
📋 Giải thích tất cả các phương án (đúng và sai)
-
Use Managed Spot Training ❌
Sai vì: Managed Spot Training sử dụng Spot Instances để tiết kiệm chi phí (lên đến 90%), nhưng không giảm thời gian khởi động cơ sở hạ tầng. Thực tế, Spot Instances có thể tăng startup time do phải bidding và chờ availability, đặc biệt với consecutive jobs. Không phù hợp cho minimize infrastructure startup. -
Use SageMaker managed warm pools ✅
Đúng vì: Như đã giải thích ở trên, warm pools giữ instance ở trạng thái sẵn sàng, trực tiếp giảm cold start time cho các job liên tiếp. Tính năng này được thiết kế dành riêng cho experimentation và iterative training, tích hợp native với SageMaker Training Jobs. -
Use SageMaker Training Compiler ❌
Sai vì: Training Compiler tối ưu hóa mã mô hình (model graph optimization) để train nhanh hơn trên hardware (như tăng throughput 1.5-2x), nhưng không ảnh hưởng đến thời gian khởi động infrastructure. Nó chỉ can thiệp vào giai đoạn runtime của job, không giải quyết vấn đề startup instances. -
Use the SageMaker distributed data parallelism (SMDDP) library ❌
Sai vì: SMDDP là thư viện hỗ trợ distributed training trên GPU (dùng Horovod hoặc MPI backend) để tăng tốc độ huấn luyện phân tán, nhưng không giảm startup time. Nó tập trung vào parallelism dữ liệu trong quá trình train, không liên quan đến việc khởi động cluster instances.
📘 Tài liệu tham khảo (AWS cập nhật mới nhất đến 2026)
- SageMaker Warm Pools: AWS Docs - Amazon SageMaker Warm Pools – Chi tiết về giảm startup time lên đến 90% cho iterative jobs.
- Training Jobs Best Practices: AWS SageMaker Developer Guide - Training Optimization – So sánh warm pools vs. các tính năng khác.
- Exam Prep: AWS Certified Machine Learning - Specialty & DevOps Engineer Professional (DOPE-C02) blueprints nhấn mạnh warm pools cho consecutive experimentation.
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần thêm case study SageMaker, cứ hỏi nhé! 😊
A company is building a web-based AI application by using Amazon SageMaker. The application will provide the following capabilities and features: ML experimentation, training, a central model registry, model deployment, and model monitoring.
The application must ensure secure and isolated use of training data during the ML lifecycle. The training data is stored in Amazon S3.
The company must implement a manual approval-based workflow to ensure that only approved models can be deployed to production endpoints.
Which solution will meet this requirement?
- A Use SageMaker Experiments to facilitate the approval process during model registration.
- B Use SageMaker ML Lineage Tracking on the central model registry. Create tracking entities for the approval process.
- C Use SageMaker Model Monitor to evaluate the performance of the model and to manage the approval.
- D Use SageMaker Pipelines. When a model version is registered, use the AWS SDK to change the approval status to "Approved."
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi thuộc case study về Amazon SageMaker trong bối cảnh xây dựng ứng dụng AI web-based. Ứng dụng cần hỗ trợ đầy đủ vòng đời ML: thử nghiệm ML (experimentation), huấn luyện (training), model registry trung tâm, triển khai model (deployment) và giám sát model (monitoring).
🔒 Yêu cầu bảo mật chính: Đảm bảo dữ liệu huấn luyện từ Amazon S3 được sử dụng an toàn và cách ly (secure and isolated) trong toàn bộ lifecycle ML.
✅ Yêu cầu cốt lõi: Triển khai workflow thủ công dựa trên phê duyệt (manual approval-based workflow) để chỉ các model được phê duyệt mới deploy lên production endpoints.
Mục tiêu là chọn giải pháp tích hợp tốt nhất với SageMaker để kiểm soát quy trình phê duyệt model trước khi triển khai sản xuất, phù hợp với phiên bản SageMaker mới nhất (cập nhật đến 2026) hỗ trợ Model Registry với status phê duyệt (Pending/Approved/Rejected) và Pipelines cho automation workflow.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use SageMaker Pipelines. When a model version is registered, use the AWS SDK to change the approval status to "Approved."
Lý do chọn 🛠️:
- SageMaker Pipelines là dịch vụ orchestration workflow end-to-end cho ML, hỗ trợ RegisterModel step để đăng ký model vào Model Registry.
- Sau khi đăng ký, sử dụng AWS SDK (boto3) để gọi API
update_model_packagevà thay đổi approval status từ "Pending" sang "Approved" một cách thủ công (manual), đảm bảo chỉ model được phê duyệt mới deploy qua CreateEndpoint hoặc TransformJob. - Giải pháp này hoàn hảo khớp yêu cầu: Tích hợp manual approval vào pipeline, hỗ trợ isolated data via S3 (với IAM/Security Groups), và phù hợp production workflow. Đây là best practice theo AWS Well-Architected Framework cho ML (2024-2026 updates).
📋 Phân tích tất cả các phương án
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên tính năng SageMaker mới nhất.
-
❌ Use SageMaker Experiments to facilitate the approval process during model registration.
Giải thích sai 🚫: SageMaker Experiments chỉ dùng để track và so sánh trials thí nghiệm (như hyperparameters, metrics), không hỗ trợ approval workflow hay thay đổi status trong Model Registry. Nó không integrate trực tiếp với deployment/approval, dẫn đến không kiểm soát được model production. -
❌ Use SageMaker ML Lineage Tracking on the central model registry. Create tracking entities for the approval process.
Giải thích sai 🚫: ML Lineage Tracking (nay là SageMaker Lineage trong Clarify/ML Insights) dùng để trace artifacts/lineage (data, models, params) trong registry, nhưng không có cơ chế approval manual. Tạo tracking entities chỉ theo dõi lineage, không thay đổi status phê duyệt hay gate deployment. -
❌ Use SageMaker Model Monitor để evaluate the performance of the model and to manage the approval.
Giải thích sai 🚫: SageMaker Model Monitor (nay tích hợp Model Quality Monitor v2.0+) dùng để giám sát drift/bias/performance post-deployment (CPU/GPU/JSON logs), không dùng cho pre-deployment approval. Nó evaluate sau khi model live, không gate registry status hay manual workflow. -
✅ Use SageMaker Pipelines. When a model version is registered, use the AWS SDK to change the approval status to "Approved."
Giải thích đúng 🟢: Như đã phân tích ở trên, Pipelines + SDK enable manual approval step chính xác sau RegisterModel, tích hợp đầy đủ với S3 data isolation (qua Pipeline parameters/IAM), và hỗ trợ production endpoints. Đây là giải pháp scalable nhất theo AWS re:Invent 2025 guidelines.
📘 Tài liệu tham khảo (Cập nhật mới nhất 2026)
- AWS SageMaker Developer Guide - Model Registry: https://docs.aws.amazon.com/sagemaker/latest/dg/model-registry.html (Approval statuses & SDK update_model_package).
- SageMaker Pipelines: https://docs.aws.amazon.com/sagemaker/latest/dg/pipelines.html (RegisterModel step + manual gates).
- AWS Well-Architected ML Lens (2024): https://docs.aws.amazon.com/wellarchitected/latest/machine-learning-lens/welcome.html (Section: Model Approval Workflows).
- Boto3 SageMaker API: https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/sagemaker/client/update_model_package.html (Change approval status).
Hy vọng phân tích này giúp bạn nắm vững DevOps cho SageMaker! 🚀 Nếu cần demo code Pipeline, hãy hỏi thêm nhé!
A company is building a web-based AI application by using Amazon SageMaker. The application will provide the following capabilities and features: ML experimentation, training, a central model registry, model deployment, and model monitoring.
The application must ensure secure and isolated use of training data during the ML lifecycle. The training data is stored in Amazon S3.
The company needs to run an on-demand workflow to monitor bias drift for models that are deployed to real-time endpoints from the application.
Which action will meet this requirement?
- A Configure the application to invoke an AWS Lambda function that runs a SageMaker Clarify job.
- B Invoke an AWS Lambda function to pull the sagemaker-model-monitor-analyzer built-in SageMaker image.
- C Use AWS Glue Data Quality to monitor bias.
- D Use SageMaker notebooks to compare the bias.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi thuộc case study về một công ty đang xây dựng ứng dụng web-based AI sử dụng Amazon SageMaker. Ứng dụng cung cấp các tính năng: thí nghiệm ML (ML experimentation), huấn luyện mô hình (training), registry mô hình trung tâm (central model registry), triển khai mô hình (model deployment), và giám sát mô hình (model monitoring). 📊
Yêu cầu chính:
- Đảm bảo sử dụng dữ liệu huấn luyện lưu trữ trên Amazon S3 một cách an toàn và cô lập (secure and isolated) trong toàn bộ vòng đời ML. 🔒
- Chạy workflow on-demand (theo yêu cầu, không liên tục) để giám sát bias drift (sự thay đổi bias theo thời gian) cho các mô hình đã triển khai trên real-time endpoints. 🎯
Bias drift ở đây đề cập đến sự thay đổi trong bias của mô hình khi dữ liệu đầu vào thực tế khác biệt so với dữ liệu huấn luyện, cần giải pháp tự động hóa và kích hoạt theo nhu cầu. Câu hỏi tập trung vào hành động phù hợp nhất để đáp ứng yêu cầu này, dựa trên các tính năng SageMaker mới nhất (cập nhật đến 2026). 🛠️
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Configure the application to invoke an AWS Lambda function that runs a SageMaker Clarify job.
Lý do:
- Amazon SageMaker Clarify là dịch vụ chuyên dụng để phát hiện và giám sát bias (thiên kiến) cũng như drift (sự thay đổi) trong mô hình ML, bao gồm bias drift cho real-time endpoints. Nó hỗ trợ chạy processing jobs on-demand, có thể kích hoạt qua AWS Lambda để tạo workflow tự động.
- Workflow này đảm bảo secure và isolated: Clarify sử dụng execution role riêng biệt, truy cập S3 qua IAM policies, và chạy trong môi trường cô lập của SageMaker. Lambda làm trigger linh hoạt, không cần infrastructure liên tục.
- Phù hợp hoàn hảo với yêu cầu on-demand và tích hợp SageMaker endpoints. 🚀
📋 Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, với ✅ cho đúng và ❌ cho sai. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt:
-
✅ Configure the application to invoke an AWS Lambda function that runs a SageMaker Clarify job.
Phương án này đúng vì SageMaker Clarify được thiết kế chuyên biệt cho việc phân tích bias và bias drift trên endpoints real-time. Lambda trigger job Clarify processing một cách on-demand, hỗ trợ baseline từ S3, và tự động hóa toàn bộ quy trình mà không vi phạm isolation. Đây là best practice theo AWS Well-Architected Framework cho ML. 🏆 -
❌ Invoke an AWS Lambda function to pull the sagemaker-model-monitor-analyzer built-in SageMaker image.
Phương án sai vì SageMaker Model Monitor (sử dụng imagesagemaker-model-monitor-analyzer) chủ yếu giám sát data drift, quality drift, và feature drift, chứ không hỗ trợ trực tiếp bias drift. Image này dành cho monitoring liên tục (scheduled), không tối ưu cho bias analysis on-demand. Sử dụng nó sẽ không đáp ứng yêu cầu bias cụ thể. 🚫 -
❌ Use AWS Glue Data Quality to monitor bias.
Phương án sai vì AWS Glue Data Quality chỉ kiểm tra chất lượng dữ liệu thô (data quality rules như null checks, schema validation), không phải bias trong mô hình ML hoặc drift trên endpoints. Nó không tích hợp với SageMaker endpoints và không hỗ trợ ML-specific metrics như bias drift. Không phù hợp với lifecycle ML. 📉 -
❌ Use SageMaker notebooks to compare the bias.
Phương án sai vì SageMaker Notebooks là môi trường Jupyter tương tác cho experimentation thủ công, không phải workflow on-demand tự động. Việc so sánh bias thủ công không scalable, không secure cho production (dễ lộ dữ liệu), và không tích hợp trực tiếp với real-time endpoints mà không cần code phức tạp thêm. ❌
📘 Tài liệu tham khảo
- AWS SageMaker Clarify Documentation: Amazon SageMaker Clarify – Chi tiết về bias detection và drift monitoring (cập nhật 2024-2026, hỗ trợ real-time endpoints).
- SageMaker Model Monitor: Model Monitor – Phân biệt với Clarify.
- AWS Well-Architected Machine Learning Lens: ML Lens – Best practices cho monitoring bias drift.
- AWS re:Post và Blogs: Tìm "SageMaker Clarify bias drift Lambda" cho examples thực tế. 🌐
Phân tích dựa trên kiến thức AWS DOP-C02 (DevOps Engineer Professional) phiên bản mới nhất 2026! Nếu cần thêm chi tiết, hỏi nhé! 💡
An ML engineer is developing a fraud detection model on AWS. The training dataset includes transaction logs, customer profiles, and tables from an on-premises MySQL database. The transaction logs and customer profiles are stored in Amazon S3.
The dataset has a class imbalance that affects the learning of the model's algorithm. Additionally, many of the features have interdependencies. The algorithm is not capturing all the desired underlying patterns in the data.
Which AWS service or feature can aggregate the data from the various data sources?
- A Amazon EMR Spark jobs
- B Amazon Kinesis Data Streams
- C Amazon DynamoDB
- D AWS Lake Formation
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi thuộc case study về một ML engineer đang phát triển mô hình phát hiện gian lận (fraud detection) trên AWS. Dữ liệu huấn luyện bao gồm:
- Transaction logs và customer profiles lưu trữ trong Amazon S3 (dữ liệu đám mây).
- Tables từ on-premises MySQL database (dữ liệu tại chỗ, cần di chuyển hoặc tổng hợp). Ngoài ra, dataset gặp vấn đề class imbalance (mất cân bằng lớp), interdependencies giữa các features (tương quan lẫn nhau), và thuật toán chưa capture hết underlying patterns (mẫu hình ngầm).
Câu hỏi chính: "Which AWS service or feature can aggregate the data from the various data sources?"
📌 Mục tiêu: Tìm dịch vụ AWS có khả năng tổng hợp (aggregate) dữ liệu từ nhiều nguồn khác nhau (S3, on-premises MySQL) để tạo dataset thống nhất cho ML, hỗ trợ xử lý vấn đề imbalance và interdependencies trong data lake/ML workflow.
✅ Đáp án đúng: AWS Lake Formation
Lý do lựa chọn 🛠️:
AWS Lake Formation là dịch vụ chuyên xây dựng, quản lý và bảo mật data lake trên S3. Nó tự động tổng hợp dữ liệu từ nhiều nguồn (S3, databases như MySQL qua AWS Glue crawlers), tạo centralized catalog (Glue Data Catalog), hỗ trợ data ingestion, transformation và governance. Trong context ML fraud detection:
- Crawl và ingest dữ liệu on-premises MySQL vào S3 data lake.
- Xử lý class imbalance/interdependencies qua ML transforms (như SageMaker integration).
- Phù hợp phiên bản mới nhất (2024-2026): Lake Formation hỗ trợ GovCloud, zero-ETL integrations với RDS/MySQL, và fine-grained access control cho ML teams.
📘 Tài liệu tham khảo: AWS Lake Formation Documentation & Lake Formation for ML.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn giữ nguyên văn bản gốc bằng tiếng Anh, kèm giải thích chi tiết bằng tiếng Việt tại sao đúng/sai:
-
❌ Amazon EMR Spark jobs
Sai vì EMR (Elastic MapReduce) dùng Spark để xử lý batch/streaming big data (như ETL jobs), nhưng không phải dịch vụ aggregate tự động từ nhiều nguồn như on-premises DB + S3. EMR cần code thủ công (Spark jobs) để đọc dữ liệu, không có governance/catalog sẵn như data lake. Không giải quyết trực tiếp class imbalance/interdependencies mà chỉ process raw data. -
❌ Amazon Kinesis Data Streams
Sai vì Kinesis Data Streams là dịch vụ streaming real-time (ingest cao throughput), phù hợp dữ liệu liên tục như transaction logs thời gian thực, nhưng không aggregate dữ liệu tĩnh từ S3 hoặc on-premises MySQL. Nó thiếu catalog/transform cho ML dataset, chỉ buffer streams chứ không build data lake. -
❌ Amazon DynamoDB
Sai vì DynamoDB là NoSQL database managed, tối ưu cho operational workloads (low-latency reads/writes), nhưng không hỗ trợ aggregate dữ liệu lớn từ S3/MySQL. Nó không phải data lake tool, thiếu ETL/crawling cho historical data, và không scale cho ML training datasets lớn với interdependencies. -
✅ AWS Lake Formation
Đúng vì đây là dịch vụ cốt lõi để aggregate data từ various sources (S3, JDBC như MySQL) vào data lake trên S3. Tích hợp Glue để crawl metadata, transform data, và hỗ trợ ML (SageMaker) xử lý imbalance/patterns. Hoàn hảo cho case study này với dữ liệu hybrid (cloud + on-premises).
🧠 Kết luận nổi bật: AWS Lake Formation là lựa chọn tối ưu cho data aggregation trong ML pipelines trên AWS (theo best practices DOP-C02 exam đến 2026), giúp ML engineer nhanh chóng build dataset sạch cho fraud detection mà không cần code phức tạp! 🚀
An ML engineer is developing a fraud detection model on AWS. The training dataset includes transaction logs, customer profiles, and tables from an on-premises MySQL database. The transaction logs and customer profiles are stored in Amazon S3.
The dataset has a class imbalance that affects the learning of the model's algorithm. Additionally, many of the features have interdependencies. The algorithm is not capturing all the desired underlying patterns in the data.
After the data is aggregated, the ML engineer must implement a solution to automatically detect anomalies in the data and to visualize the result.
Which solution will meet these requirements?
- A Use Amazon Athena to automatically detect the anomalies and to visualize the result.
- B Use Amazon Redshift Spectrum to automatically detect the anomalies. Use Amazon QuickSight to visualize the result.
- C Use Amazon SageMaker Data Wrangler to automatically detect the anomalies and to visualize the result.
- D Use AWS Batch to automatically detect the anomalies. Use Amazon QuickSight to visualize the result.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi thuộc chủ đề Machine Learning trên AWS, cụ thể là quy trình chuẩn bị dữ liệu (data preparation) cho mô hình phát hiện gian lận (fraud detection).
-
Bối cảnh: Một kỹ sư ML đang xây dựng mô hình trên AWS. Dữ liệu huấn luyện bao gồm:
- Transaction logs và customer profiles lưu trữ trong Amazon S3.
- Bảng từ cơ sở dữ liệu MySQL on-premises.
-
Vấn đề dữ liệu:
- Class imbalance: Dữ liệu không cân bằng giữa các lớp (ví dụ: ít giao dịch gian lận so với bình thường), ảnh hưởng đến việc học của thuật toán.
- Interdependencies giữa các features: Các đặc trưng có mối quan hệ phụ thuộc lẫn nhau.
- Thuật toán không capture hết các pattern ẩn trong dữ liệu.
-
Yêu cầu chính (sau khi aggregate dữ liệu):
- Automatically detect anomalies (tự động phát hiện bất thường trong dữ liệu, như outliers, missing values, hoặc patterns bất thường).
- Visualize the result (hiển thị kết quả trực quan).
Giải pháp phải phù hợp với dữ liệu từ S3 và on-premises DB, hỗ trợ xử lý imbalance/interdependencies gián tiếp qua data prep, và tập trung vào anomaly detection + visualization tự động. Đây là bước tiền xử lý dữ liệu trước khi train model trên SageMaker.
📘 Dẫn nguồn: AWS Documentation - SageMaker Data Wrangler (cập nhật 2024-2026): docs.aws.amazon.com/sagemaker/latest/dg/data-wrangler.html. Tính năng anomaly detection được tích hợp sẵn trong flow data preparation.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Amazon SageMaker Data Wrangler to automatically detect the anomalies and to visualize the result.
Lý do 🛠️:
- SageMaker Data Wrangler là công cụ chuyên dụng cho data preparation trong SageMaker, hỗ trợ import dữ liệu từ S3, Athena, JDBC (như MySQL on-premises).
- Tự động detect anomalies: Có built-in transforms như anomaly detection (sử dụng algorithms phát hiện outliers, custom rules), xử lý class imbalance (qua sampling, SMOTE), và visualize interdependencies (correlation plots, histograms).
- Visualization tích hợp: Giao diện no-code/low-code với interactive visualizations (graphs, charts) ngay trong flow, export trực tiếp sang SageMaker Processing/Training jobs.
- Phù hợp hoàn hảo với case study: Xử lý dữ liệu aggregate từ nhiều nguồn, giải quyết vấn đề pattern capture bằng data quality checks trước training.
- Cập nhật mới nhất (2026): Data Wrangler hỗ trợ MLflow integration và advanced anomaly models (như Isolation Forest).
❌ Phân tích tất cả các phương án (đúng/sai)
-
Phương án SAI: Use Amazon Athena to automatically detect the anomalies and to visualize the result.
Giải thích: Athena là dịch vụ query serverless trên S3 (SQL-based), giỏi aggregate data nhưng KHÔNG có tính năng tự động detect anomalies hay visualization built-in. Phải viết query thủ công (ví dụ: Z-score), không visualize trực tiếp (cần QuickSight riêng). Không xử lý on-premises MySQL dễ dàng mà không qua connector phức tạp. ❌ Không đáp ứng "automatically detect + visualize". -
Phương án SAI: Use Amazon Redshift Spectrum to automatically detect the anomalies. Use Amazon QuickSight to visualize the result.
Giải thích: Redshift Spectrum query dữ liệu external (S3) từ Redshift cluster, nhưng KHÔNG tự động detect anomalies (chỉ query SQL, cần code custom). QuickSight visualize tốt nhưng tách biệt, không tích hợp anomaly detection. Chi phí cao cho data warehouse không cần thiết ở data prep phase; khó kết nối on-premises MySQL trực tiếp. ❌ Không "automatically" và không liền mạch. -
Phương án ĐÚNG: Use Amazon SageMaker Data Wrangler to automatically detect the anomalies and to visualize the result.
Giải thích: Như đã phân tích ở trên. ✅ Hoàn toàn khớp yêu cầu, end-to-end data prep với anomaly detection (built-in analyzers) và visualization interactive. Hỗ trợ fraud detection workflows chuẩn AWS. -
Phương án SAI: Use AWS Batch to automatically detect the anomalies. Use Amazon QuickSight to visualize the result.
Giải thích: AWS Batch là batch computing cho jobs containerized (như run scripts Python/ML), KHÔNG có anomaly detection tự động (phải code thủ công với libraries như scikit-learn). QuickSight visualize riêng lẻ, không tích hợp. Không phù hợp data prep từ S3/MySQL; overhead cao cho simple anomaly tasks. ❌ Không "automatically" và thiếu integration ML.
🧠 Lời khuyên DevOps: Trong pipeline CI/CD (CodePipeline + SageMaker), Data Wrangler export flows thành Processing Jobs tự động hóa toàn bộ. Kết hợp EventBridge cho monitoring anomalies!
An ML engineer is developing a fraud detection model on AWS. The training dataset includes transaction logs, customer profiles, and tables from an on-premises MySQL database. The transaction logs and customer profiles are stored in Amazon S3.
The dataset has a class imbalance that affects the learning of the model's algorithm. Additionally, many of the features have interdependencies. The algorithm is not capturing all the desired underlying patterns in the data.
The training dataset includes categorical data and numerical data. The ML engineer must prepare the training dataset to maximize the accuracy of the model.
Which action will meet this requirement with the LEAST operational overhead?
- A Use AWS Glue to transform the categorical data into numerical data.
- B Use AWS Glue to transform the numerical data into categorical data.
- C Use Amazon SageMaker Data Wrangler to transform the categorical data into numerical data.
- D Use Amazon SageMaker Data Wrangler to transform the numerical data into categorical data.
Xem giải thích
🧩 Giải thích chi tiết nội dung câu hỏi
Câu hỏi thuộc case study về một kỹ sư ML đang phát triển mô hình phát hiện gian lận (fraud detection) trên AWS. Dataset huấn luyện bao gồm:
- Transaction logs và customer profiles lưu trữ trong Amazon S3.
- Tables từ on-premises MySQL database.
Các vấn đề chính của dataset:
- Class imbalance: Phân bố lớp không cân bằng, ảnh hưởng đến việc học của thuật toán.
- Interdependencies giữa các features: Các đặc trưng có sự phụ thuộc lẫn nhau.
- Thuật toán không capture hết patterns: Mô hình chưa nắm bắt đầy đủ các mẫu ngầm trong dữ liệu.
- Dataset chứa cả categorical data (dữ liệu phân loại, như chuỗi ký tự) và numerical data (dữ liệu số).
Yêu cầu chính: Chuẩn bị dataset để tối đa hóa độ chính xác (accuracy) của mô hình, với LEAST operational overhead (ít nỗ lực vận hành nhất).
📘 Bối cảnh AWS mới nhất (2026): Trong SageMaker (phiên bản mới nhất), việc chuẩn bị dữ liệu ML cần xử lý encoding categorical sang numerical (ví dụ: one-hot encoding, label encoding) để hầu hết các thuật toán ML có thể học hiệu quả. Đồng thời, cần xử lý imbalance (như SMOTE) và feature engineering để giảm interdependencies (như PCA). Tool lý tưởng phải low-code/no-code để giảm overhead.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Amazon SageMaker Data Wrangler to transform the categorical data into numerical data.
Lý do:
- 🛠️ SageMaker Data Wrangler là công cụ chuyên biệt cho data preparation trong ML workflows, hỗ trợ giao diện visual (flow-based UI) để transform categorical data thành numerical (qua các built-in transforms như one-hot encoding, target encoding). Điều này giúp model capture tốt hơn patterns, xử lý imbalance (custom flows cho resampling), và interdependencies (feature importance, dimensionality reduction).
- LEAST operational overhead: Không cần viết code phức tạp như Glue jobs; chỉ drag-and-drop, export trực tiếp sang SageMaker Processing/Training jobs. Tích hợp liền mạch với S3 và SageMaker, hỗ trợ import từ MySQL qua JDBC.
- Giải quyết trực tiếp vấn đề: Categorical phải encode sang numerical để model học (hầu hết algorithms như XGBoost, neural nets yêu cầu input numerical).
📋 Phân tích tất cả các phương án (đúng/sai)
-
❌ Use AWS Glue to transform the categorical data into numerical data.
Sai vì: AWS Glue là ETL service cho big data (Spark-based), có thể transform categorical sang numerical qua custom scripts (ETL jobs). Tuy nhiên, operational overhead cao: Phải viết PySpark/Scala code, debug jobs, manage crawlers cho schema. Không chuyên cho ML data prep, thiếu visual tools cho encoding/imbalance handling. Phù hợp scale lớn nhưng không "least overhead" so với Data Wrangler. -
❌ Use AWS Glue to transform the numerical data into categorical data.
Sai vì: Transform numerical sang categorical (binning/discretization) là không hợp lý ở đây. Vấn đề chính là categorical chưa numerical hóa để model học; làm ngược lại sẽ làm mất thông tin số (precision), giảm accuracy. Glue dù làm được nhưng vẫn overhead cao và không giải quyết vấn đề cốt lõi. -
✅ Use Amazon SageMaker Data Wrangler to transform the categorical data into numerical data.
Đúng vì: Như giải thích ở trên. Data Wrangler (updated 2024-2026) hỗ trợ >300 transforms built-in, bao gồm categorical encoding, imbalance fixing (undersample/oversample), feature engineering cho interdependencies. Export flow thành code (Python/PyTorch), zero-code cho ML engineers. Least overhead! -
❌ Use Amazon SageMaker Data Wrangler to transform the numerical data into categorical data.
Sai vì: Tương tự phương án 2, transform numerical sang categorical không giúp capture patterns tốt hơn (làm dữ liệu thô hơn). Data Wrangler làm được binning nhưng không phải giải pháp tối ưu cho accuracy; vấn đề là categorical cần numerical hóa trước.
📚 Tài liệu tham khảo (AWS docs cập nhật 2026)
- SageMaker Data Wrangler: AWS SageMaker Data Wrangler Documentation – Chi tiết transforms cho categorical encoding và ML prep.
- So sánh Glue vs. SageMaker cho ML: AWS re:Post - Data Prep Best Practices (Glue cho ETL lớn, Wrangler cho ML low-overhead).
- Fraud Detection ML Patterns: AWS ML Best Practices - Fraud Detection – Nhấn mạnh data prep với Wrangler cho imbalance/categorical.
- Exam Topic DOP-C02: Phần SageMaker/Glue trong DevOps Pro (2024 blueprint).
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần case study khác, hỏi nhé!