Ngân hàng đề — AWS Certified Machine Learning Specialty

Tìm thấy 371 câu.

Câu 141
A telecommunications company is developing a mobile app for its customers. The company is using an Amazon SageMaker hosted endpoint for machine learning model inferences.
Developers want to introduce a new version of the model for a limited number of users who subscribed to a preview feature of the app. After the new version of the model is tested as a preview, developers will evaluate its accuracy. If a new version of the model has better accuracy, developers need to be able to gradually release the new version for all users over a fixed period of time.
How can the company implement the testing model with the LEAST amount of operational overhead?
  1. A Update the ProductionVariant data type with the new version of the model by using the CreateEndpointConfig operation with the InitialVariantWeight parameter set to 0. Specify the TargetVariant parameter for InvokeEndpoint calls for users who subscribed to the preview feature. When the new version of the model is ready for release, gradually increase InitialVariantWeight until all users have the updated version.
  2. B Configure two SageMaker hosted endpoints that serve the different versions of the model. Create an Application Load Balancer (ALB) to route traffic to both endpoints based on the TargetVariant query string parameter. Reconfigure the app to send the TargetVariant query string parameter for users who subscribed to the preview feature. When the new version of the model is ready for release, change the ALB's routing algorithm to weighted until all users have the updated version.
  3. C Update the DesiredWeightsAndCapacity data type with the new version of the model by using the UpdateEndpointWeightsAndCapacities operation with the DesiredWeight parameter set to 0. Specify the TargetVariant parameter for InvokeEndpoint calls for users who subscribed to the preview feature. When the new version of the model is ready for release, gradually increase DesiredWeight until all users have the updated version.
  4. D Configure two SageMaker hosted endpoints that serve the different versions of the model. Create an Amazon Route 53 record that is configured with a simple routing policy and that points to the current version of the model. Configure the mobile app to use the endpoint URL for users who subscribed to the preview feature and to use the Route 53 record for other users. When the new version of the model is ready for release, add a new model version endpoint to Route 53, and switch the policy to weighted until all users have the updated version.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh việc triển khai một ứng dụng mobile sử dụng Amazon SageMaker hosted endpoint để thực hiện suy luận mô hình machine learning (ML). 🛠️ Công ty telecom muốn:

  • Giới thiệu phiên bản model mới chỉ cho một số user đăng ký preview feature (kiểm tra accuracy).
  • Sau khi test, nếu model mới tốt hơn, gradually release (phát hành dần dần) cho tất cả user trong một khoảng thời gian cố định.
  • Yêu cầu: Least operational overhead (ít overhead vận hành nhất), nghĩa là tránh tạo nhiều endpoint phức tạp, giảm thiểu downtime, quản lý traffic đơn giản.

📘 Bối cảnh AWS SageMaker (cập nhật 2026): SageMaker hỗ trợ multi-model/multi-variant endpoints với traffic splitting qua weights (trọng số). Có thể thêm variant mới vào endpoint hiện tại, set weight=0 ban đầu, dùng TargetVariant parameter trong InvokeEndpoint để route traffic cho preview users, rồi dần tăng weight để shift traffic tự động. Điều này không cần tạo endpoint mới, chỉ update weights in-place qua API UpdateEndpointWeightsAndCapacities.

Nguồn tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng:
Update the DesiredWeightsAndCapacity data type with the new version of the model by using the UpdateEndpointWeightsAndCapacities operation with the DesiredWeight parameter set to 0. Specify the TargetVariant parameter for InvokeEndpoint calls for users who subscribed to the preview feature. When the new version of the model is ready for release, gradually increase DesiredWeight until all users have the updated version.

Lý do 🏆:

  • Đây là cách chuẩn AWS cho traffic shifting trên single SageMaker endpoint với multi-production variants.
  • Thêm variant mới với DesiredWeight=0 (không nhận traffic tự động), dùng TargetVariant trong InvokeEndpoint để route chỉ cho preview users (không ảnh hưởng user khác).
  • Khi ready, gọi UpdateEndpointWeightsAndCapacities để tăng dần DesiredWeight (ví dụ: 0 → 0.1 → 0.5 → 1.0), SageMaker tự động shift traffic không downtime, hỗ trợ canary/blue-green deployment.
  • Least overhead: Chỉ 1 endpoint, update API nhanh (giây), auto-scale, monitor qua CloudWatch. Phù hợp DevOps best practice DOP-C02.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn (giữ nguyên văn bản gốc tiếng Anh). Tôi đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm giải thích chi tiết bằng tiếng Việt.

  • Update the ProductionVariant data type with the new version of the model by using the CreateEndpointConfig operation with the InitialVariantWeight parameter set to 0. Specify the TargetVariant parameter for InvokeEndpoint calls for users who subscribed to the preview feature. When the new version of the model is ready for release, gradually increase InitialVariantWeight until all users have the updated version.
    ❌ Sai: CreateEndpointConfig chỉ tạo config mới, không update endpoint hiện tại (phải gọi UpdateEndpoint sau, gây overhead). InitialVariantWeight chỉ dùng cho config mới, không hỗ trợ update weights dần dần trên endpoint đang chạy. Không phải cách chính thức, dễ lỗi deployment.

  • Configure two SageMaker hosted endpoints that serve the different versions of the model. Create an Application Load Balancer (ALB) to route traffic to both endpoints based on the TargetVariant query string parameter. Reconfigure the app to send the TargetVariant query string parameter for users who subscribed to the preview feature. When the new version of the model is ready for release, change the ALB's routing algorithm to weighted until all users have the updated version.
    ❌ Sai: Tạo 2 endpoints riêng tăng chi phí và overhead (quản lý scale, monitor riêng). TargetVariant là parameter của SageMaker InvokeEndpoint, không phải query cho ALB (ALB route host/path/query, nhưng phức tạp code app + Lambda/Proxy). Weighted routing ALB không tự động như SageMaker native, overhead cao hơn nhiều.

  • Update the DesiredWeightsAndCapacity data type with the new version of the model by using the UpdateEndpointWeightsAndCapacities operation with the DesiredWeight parameter set to 0. Specify the TargetVariant parameter for InvokeEndpoint calls for users who subscribed to the preview feature. When the new version of the model is ready for release, gradually increase DesiredWeight until all users have the updated version.
    ✅ Đúng: Như giải thích trên. Native SageMaker feature, zero-downtime, single endpoint, hỗ trợ preview qua TargetVariant và gradual shift weights. Overhead thấp nhất (chỉ API calls).

  • Configure two SageMaker hosted endpoints that serve the different versions of the model. Create an Amazon Route 53 record that is configured with a simple routing policy and that points to the current version of the model. Configure the mobile app to use the endpoint URL for users who subscribed to the preview feature and to use the Route 53 record for other users. When the new version of the model is ready for release, add a new model version endpoint to Route 53, and switch the policy to weighted until all users have the updated version.
    ❌ Sai: 2 endpoints + Route53 weighted policy gây overhead lớn (DNS propagation chậm 1-5 phút, không real-time; chi phí endpoints đôi). App phải hardcode 2 URLs (phức tạp logic), không dùng TargetVariant native. Route53 weighted phù hợp DNS failover, không ideal cho ML inference traffic shifting.

Kết luận 🎯: Lựa chọn đúng tận dụng SageMaker production variants để A/B testing + canary release hiệu quả nhất, phù hợp chứng chỉ AWS Certified DevOps Engineer - Professional (DOP-C02). Nếu deploy thực tế, kết hợp CloudWatch + X-Ray để monitor accuracy/traffic! 🚀

Câu 142
A company offers an online shopping service to its customers. The company wants to enhance the site's security by requesting additional information when customers access the site from locations that are different from their normal location. The company wants to update the process to call a machine learning (ML) model to determine when additional information should be requested.
The company has several terabytes of data from its existing ecommerce web servers containing the source IP addresses for each request made to the web server. For authenticated requests, the records also contain the login name of the requesting user.
Which approach should an ML specialist take to implement the new security feature in the web application?
  1. A Use Amazon SageMaker Ground Truth to label each record as either a successful or failed access attempt. Use Amazon SageMaker to train a binary classification model using the factorization machines (FM) algorithm.
  2. B Use Amazon SageMaker to train a model using the IP Insights algorithm. Schedule updates and retraining of the model using new log data nightly.
  3. C Use Amazon SageMaker Ground Truth to label each record as either a successful or failed access attempt. Use Amazon SageMaker to train a binary classification model using the IP Insights algorithm.
  4. D Use Amazon SageMaker to train a model using the Object2Vec algorithm. Schedule updates and retraining of the model using new log data nightly.
Xem giải thích

🧩 Giải thích nội dung câu hỏi
Câu hỏi xoay quanh một công ty cung cấp dịch vụ mua sắm trực tuyến muốn tăng cường bảo mật bằng cách yêu cầu thông tin bổ sung khi khách hàng truy cập từ vị trí (IP) khác với vị trí thông thường. Họ muốn tích hợp mô hình Machine Learning (ML) để quyết định khi nào cần yêu cầu thêm thông tin.
Dữ liệu có sẵn: Hàng terabytes log từ web server ecommerce, chứa source IP addresses cho mọi request, và đối với authenticated requests còn có login name của user.
Mục tiêu: Triển khai tính năng bảo mật mới trên web app bằng cách sử dụng ML specialist.
📘 Yêu cầu chính: Phát hiện anomaly dựa trên IP-user association mà không cần dữ liệu labeled (vì log chỉ có IP và user, không có nhãn thành công/thất bại). Đây là vấn đề unsupervised anomaly detection chuyên biệt cho IP insights.

✅ Đáp án đúng:
Use Amazon SageMaker to train a model using the IP Insights algorithm. Schedule updates and retraining of the model using new log data nightly.
Lý do lựa chọn:

  • IP Insights là thuật toán built-in của Amazon SageMaker (cập nhật đến 2026), được thiết kế chuyên biệt cho anomaly detection dựa trên mối liên hệ giữa IP addresses và users. Nó sử dụng unsupervised learning (không cần labeling), học từ dữ liệu lịch sử để phát hiện IP lạ so với hành vi bình thường của user.
  • Phù hợp hoàn hảo với dữ liệu log (IP + user login), giúp dự đoán xác suất một IP-user pair là "bình thường" hay "anomalous".
  • Schedule updates nightly: SageMaker hỗ trợ retraining tự động qua SageMaker Pipelines hoặc Processing Jobs, sử dụng log mới để model luôn cập nhật (best practice cho fraud detection).
    🛠️ Lợi ích: Triển khai nhanh, scalable trên terabytes data, tích hợp dễ với web app qua SageMaker endpoints real-time inference.

📋 Phân tích tất cả các phương án (giữ nguyên văn bản gốc):

  • ❌ Use Amazon SageMaker Ground Truth to label each record as either a successful or failed access attempt. Use Amazon SageMaker to train a binary classification model using the factorization machines (FM) algorithm.
    Sai vì: Ground Truth dùng để label dữ liệu thủ công (supervised), nhưng dữ liệu log không có nhãn "successful/failed" sẵn, và việc label terabytes data tốn kém/không khả thi. FM (Factorization Machines) là cho recommendation systems (như CTR prediction), không chuyên cho IP anomaly detection. Không phù hợp với unsupervised nature của vấn đề.

  • ✅ Use Amazon SageMaker to train a model using the IP Insights algorithm. Schedule updates and retraining of the model using new log data nightly.
    Đúng vì: Như giải thích ở trên – thuật toán lý tưởng, unsupervised, hỗ trợ retraining định kỳ.

  • ❌ Use Amazon SageMaker Ground Truth to label each record as either a successful or failed access attempt. Use Amazon SageMaker to train a binary classification model using the IP Insights algorithm.
    Sai vì: IP Insights là unsupervised algorithm (dựa trên matrix factorization cho IP-user associations), không dùng cho binary classification (supervised). Việc dùng Ground Truth để label là thừa và sai, vì thuật toán không cần nhãn – nó tự học pattern từ dữ liệu lịch sử.

  • ❌ Use Amazon SageMaker to train a model using the Object2Vec algorithm. Schedule updates and retraining of the model using new log data nightly.
    Sai vì: Object2Vec dùng cho embedding objects (như text/images thành vectors), không chuyên cho IP anomaly. Không tận dụng được cấu trúc IP-user pair, kém hiệu quả hơn IP Insights (thiết kế riêng cho fraud/IP detection). Retraining nightly là tốt nhưng thuật toán sai.

📘 Tài liệu tham khảo (AWS docs cập nhật 2026):

Câu 143
A retail company wants to combine its customer orders with the product description data from its product catalog. The structure and format of the records in each dataset is different. A data analyst tried to use a spreadsheet to combine the datasets, but the effort resulted in duplicate records and records that were not properly combined. The company needs a solution that it can use to combine similar records from the two datasets and remove any duplicates.
Which solution will meet these requirements?
  1. A Use an AWS Lambda function to process the data. Use two arrays to compare equal strings in the fields from the two datasets and remove any duplicates.
  2. B Create AWS Glue crawlers for reading and populating the AWS Glue Data Catalog. Call the AWS Glue SearchTables API operation to perform a fuzzy- matching search on the two datasets, and cleanse the data accordingly.
  3. C Create AWS Glue crawlers for reading and populating the AWS Glue Data Catalog. Use the FindMatches transform to cleanse the data.
  4. D Create an AWS Lake Formation custom transform. Run a transformation for matching products from the Lake Formation console to cleanse the data automatically.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi mô tả một công ty bán lẻ cần kết hợp (combine) dữ liệu đơn hàng khách hàng (customer orders) với dữ liệu mô tả sản phẩm từ catalog sản phẩm. Hai bộ dữ liệu có cấu trúc và định dạng khác nhau, dẫn đến khó khăn khi xử lý thủ công bằng spreadsheet (kết quả là duplicate records và không combine đúng). Yêu cầu chính là giải pháp tự động để khớp các records tương tự (similar records) từ hai bộ dữ liệu và loại bỏ duplicates.

🔍 Vấn đề cốt lõi: Cần fuzzy matching (khớp gần đúng) vì dữ liệu không hoàn hảo khớp 100%, thường gặp trong data cleansing trên AWS với công cụ ETL như Glue. Giải pháp phải scalable, không thủ công, phù hợp với big data.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create AWS Glue crawlers for reading and populating the AWS Glue Data Catalog. Use the FindMatches transform to cleanse the data.

Lý do:

  • 🛠️ AWS Glue crawlers tự động crawl và catalog dữ liệu từ S3 hoặc nguồn khác vào AWS Glue Data Catalog, giúp khám phá schema động.
  • FindMatches transform (dựa trên ML) là tính năng chuyên biệt của AWS Glue để phát hiện duplicates và fuzzy matching giữa records tương tự, ngay cả khi cấu trúc khác nhau. Nó sử dụng ML models (như supervised hoặc unsupervised) để gán confidence score, tự động cleanse dữ liệu mà không cần code phức tạp.
  • ✅ Hoàn hảo cho yêu cầu: Kết hợp orders với product catalog, loại bỏ duplicates hiệu quả. Đây là best practice theo AWS Well-Architected Framework cho data analytics (cập nhật 2024-2026).

📘 Tài liệu tham khảo:

  • AWS Glue Developer Guide: FindMatches transform (phiên bản mới nhất 2026 hỗ trợ Glue 4.0 với cải tiến ML).
  • AWS re:Post và Well-Architected Labs: Data Analytics Lens.

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, đánh dấu ✅ (đúng) hoặc ❌ (sai), với lý do cụ thể dựa trên tính năng AWS cập nhật đến 2026:

  • ❌ [SAI] Use an AWS Lambda function to process the data. Use two arrays to compare equal strings in the fields from the two datasets and remove any duplicates.
    Phương án này không phù hợp vì Lambda chỉ xử lý serverless code, nhưng so sánh strings bằng arrays chỉ làm exact matching (khớp chính xác), không xử lý fuzzy matching hoặc cấu trúc khác nhau. Dễ gây lỗi với big data (memory limit 10GB), không scalable cho datasets lớn, và phải code thủ công – trái với yêu cầu tự động cleanse. Không dùng ML để detect similar records.

  • ❌ [SAI] Create AWS Glue crawlers for reading and populating the AWS Glue Data Catalog. Call the AWS Glue SearchTables API operation to perform a fuzzy- matching search on the two datasets, and cleanse the data accordingly.
    Crawlers đúng bước đầu, nhưng SearchTables API chỉ dùng để tìm kiếm metadata tables trong Catalog (như keyword search), không hỗ trợ fuzzy-matching trên dữ liệu records. Không có tính năng cleanse dữ liệu thực tế, chỉ query catalog. Sai lầm lớn về API usage (cập nhật Glue API 2026 vẫn giữ nguyên).

  • ✅ [ĐÚNG] Create AWS Glue crawlers for reading and populating the AWS Glue Data Catalog. Use the FindMatches transform to cleanse the data.
    Như đã giải thích ở phần đáp án đúng: Crawlers + FindMatches là combo hoàn chỉnh, ML-based, tự động detect/merge duplicates với accuracy cao (hỗ trợ custom tuning models trong Glue Studio 2026). Best fit cho retail data integration.

  • ❌ [SAI] Create an AWS Lake Formation custom transform. Run a transformation for matching products from the Lake Formation console to cleanse the data automatically.
    AWS Lake Formation tập trung governance và access control cho data lake, hỗ trợ custom transforms nhưng chỉ cho data quality rules cơ bản (như column validation), không có built-in fuzzy matching hoặc duplicate detection như FindMatches. Phải code custom Spark job phức tạp, không tự động cho product matching. Lake Formation 2026 ưu tiên security hơn ETL matching (xem docs Lake Formation transforms).

🛡️ Lưu ý cuối: Giải pháp đúng tận dụng AWS Glue ETL/ML – scalable, serverless, chi phí thấp cho data pipelines. Nếu implement, bắt đầu bằng Glue Studio visual editor để test FindMatches nhanh chóng!

Câu 144
A company provisions Amazon SageMaker notebook instances for its data science team and creates Amazon VPC interface endpoints to ensure communication between the VPC and the notebook instances. All connections to the Amazon SageMaker API are contained entirely and securely using the AWS network.
However, the data science team realizes that individuals outside the VPC can still connect to the notebook instances across the internet.
Which set of actions should the data science team take to fix the issue?
  1. A Modify the notebook instances' security group to allow traffic only from the CIDR ranges of the VPC. Apply this security group to all of the notebook instances' VPC interfaces.
  2. B Create an IAM policy that allows the sagemaker:CreatePresignedNotebooklnstanceUrl and sagemaker:DescribeNotebooklnstance actions from only the VPC endpoints. Apply this policy to all IAM users, groups, and roles used to access the notebook instances.
  3. C Add a NAT gateway to the VPC. Convert all of the subnets where the Amazon SageMaker notebook instances are hosted to private subnets. Stop and start all of the notebook instances to reassign only private IP addresses.
  4. D Change the network ACL of the subnet the notebook is hosted in to restrict access to anyone outside the VPC.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong AWS: Công ty đã triển khai Amazon SageMaker notebook instances cho đội ngũ data science bên trong một VPC (Virtual Private Cloud). Họ đã tạo VPC interface endpoints (cụ thể là endpoints cho SageMaker API) để đảm bảo mọi giao tiếp giữa VPC và các notebook instances diễn ra hoàn toàn qua mạng AWS nội bộ, an toàn và không qua internet công khai. Điều này giúp các lệnh gọi API đến SageMaker (như tạo hoặc quản lý notebook) được bảo vệ.

Vấn đề chính (issue): Mặc dù API được bảo vệ, đội ngũ data science phát hiện rằng các cá nhân bên ngoài VPC vẫn có thể kết nối trực tiếp đến notebook instances qua internet. Lý do là SageMaker notebook instances mặc định có thể truy cập qua presigned URL (URL tạm thời được tạo bởi SageMaker API), và URL này có thể được chia sẻ hoặc sử dụng từ bất kỳ đâu trên internet nếu không có kiểm soát thêm. Mục tiêu là ngăn chặn truy cập từ bên ngoài VPC, chỉ cho phép từ nội bộ VPC.

Câu hỏi yêu cầu tập hợp các hành động (set of actions) để khắc phục, tập trung vào bảo mật truy cập notebook instances một cách tối ưu, không ảnh hưởng đến hoạt động bên trong VPC. Đây là chủ đề liên quan đến bảo mật SageMaker với VPC endpoints và IAM policies (cập nhật theo AWS Well-Architected Framework và SageMaker security best practices đến năm 2026).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng là phương án thứ 2:
[ĐÚNG] Create an IAM policy that allows the sagemaker:CreatePresignedNotebooklnstanceUrl and sagemaker:DescribeNotebooklnstance actions from only the VPC endpoints. Apply this policy to all IAM users, groups, and roles used to access the notebook instances.

Lý do chi tiết:
🛠️ SageMaker notebook instances được truy cập chủ yếu qua presigned URLs được tạo bởi action sagemaker:CreatePresignedNotebookInstanceUrl (lưu ý: tên action đúng là CreatePresignedNotebookInstanceUrl – có thể là lỗi đánh máy nhỏ trong câu hỏi) và cần sagemaker:DescribeNotebookInstance để lấy thông tin. Những URL này mặc định cho phép truy cập public qua internet.

🧩 Giải pháp chính xác: Tạo IAM policy sử dụng điều kiện aws:SourceVpce (VPC Endpoint ID) để chỉ cho phép các action này khi gọi từ VPC interface endpoint của SageMaker. Điều này ngăn IAM users/roles bên ngoài VPC tạo presigned URL, từ đó chặn truy cập internet. Policy phải apply cho tất cả IAM entities dùng để access notebook. Đây là best practice AWS khuyến nghị (không cần thay đổi network config, tận dụng IAM fine-grained control). Hoạt động bình thường bên trong VPC vì endpoints vẫn cho phép gọi API nội bộ.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng phương án một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên kiến thức AWS mới nhất (SageMaker VPC-only mode, IAM conditions, VPC endpoints – cập nhật 2024-2026).

  • ❌ Phương án SAI 1:
    [SAI] Modify the notebook instances' security group to allow traffic only from the CIDR ranges of the VPC. Apply this security group to all of the notebook instances' VPC interfaces.
    Giải thích sai: Security group (SG) kiểm soát traffic inbound/outbound đến ENI (Elastic Network Interface) của notebook instances, nhưng presigned URL hoạt động ở layer ứng dụng (HTTPS port 443) và bypass SG nếu gọi trực tiếp từ SageMaker domain. SG chỉ giới hạn IP source (CIDR VPC), nhưng người ngoài VPC có thể dùng presigned URL mà không cần IP từ VPC, dẫn đến không chặn hiệu quả. Hơn nữa, notebook instances trong VPC-only mode đã dùng private IP, nhưng truy cập URL vẫn public. Không phải giải pháp gốc rễ (root cause là IAM-generated URL).

  • ✅ Phương án ĐÚNG 2:
    [ĐÚNG] Create an IAM policy that allows the sagemaker:CreatePresignedNotebooklnstanceUrl and sagemaker:DescribeNotebooklnstance actions from only the VPC endpoints. Apply this policy to all IAM users, groups, and roles used to access the notebook instances.
    Giải thích đúng (tóm tắt lại): Như phần trên, policy dùng aws:SourceVpce condition chặn tạo presigned URL từ ngoài endpoint, đảm bảo chỉ VPC nội bộ mới generate URL hợp lệ. An toàn, scalable, không downtime, phù hợp DevOps best practices. Action names đúng theo AWS IAM (CreatePresignedNotebookInstanceUrl và DescribeNotebookInstance).

  • ❌ Phương án SAI 3:
    [SAI] Add a NAT gateway to the VPC. Convert all of the subnets where the Amazon SageMaker notebook instances are hosted to private subnets. Stop and start all of the notebook instances to reassign only private IP addresses.
    Giải thích sai: NAT gateway dùng cho outbound internet từ private subnet, không chặn inbound truy cập đến notebook. Chuyển sang private subnet và restart chỉ assign private IP (đã có sẵn nếu dùng VPC), nhưng presigned URL vẫn public và cho phép kết nối HTTPS từ internet. Gây downtime (stop/start instances), tốn chi phí NAT, và không giải quyết root cause (URL generation).

  • ❌ Phương án SAI 4:
    [SAI] Change the network ACL of the subnet the notebook is hosted in to restrict access to anyone outside the VPC.
    Giải thích sai: Network ACL (NACL) là stateless firewall ở subnet level, có thể deny traffic ngoài VPC CIDR, nhưng không hiệu quả với presigned URL vì URL route qua SageMaker edge locations/public endpoints (không qua NACL của subnet). NACL phức tạp hơn SG, dễ misconfig (allow/deny both directions), và không chặn application-layer access. SageMaker khuyến nghị dùng IAM thay vì network controls cho trường hợp này.

📘 Tài liệu tham khảo (cập nhật mới nhất AWS 2026)

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ policy JSON, hãy hỏi nhé!

Câu 145
A company will use Amazon SageMaker to train and host a machine learning (ML) model for a marketing campaign. The majority of data is sensitive customer data. The data must be encrypted at rest. The company wants AWS to maintain the root of trust for the master keys and wants encryption key usage to be logged.
Which implementation will meet these requirements?
  1. A Use encryption keys that are stored in AWS Cloud HSM to encrypt the ML data volumes, and to encrypt the model artifacts and data in Amazon S3.
  2. B Use SageMaker built-in transient keys to encrypt the ML data volumes. Enable default encryption for new Amazon Elastic Block Store (Amazon EBS) volumes.
  3. C Use customer managed keys in AWS Key Management Service (AWS KMS) to encrypt the ML data volumes, and to encrypt the model artifacts and data in Amazon S3.
  4. D Use AWS Security Token Service (AWS STS) to create temporary tokens to encrypt the ML storage volumes, and to encrypt the model artifacts and data in Amazon S3.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc triển khai mã hóa dữ liệu at rest (khi dữ liệu đang lưu trữ) cho một công ty sử dụng Amazon SageMaker để huấn luyện (train) và triển khai (host) mô hình machine learning (ML) trong chiến dịch marketing. Dữ liệu chủ yếu là dữ liệu khách hàng nhạy cảm, đòi hỏi:

  • Mã hóa at rest cho các volumes dữ liệu ML (thường là EBS volumes trong SageMaker training jobs).
  • Mã hóa model artifacts và dữ liệu trong Amazon S3.
  • AWS duy trì root of trust cho master keys (nghĩa là AWS quản lý gốc tin cậy của các khóa chính, như qua KMS với HSM do AWS bảo vệ).
  • Ghi log việc sử dụng khóa mã hóa (key usage logging), thường qua CloudTrail.

🛠️ Yêu cầu chính: Phương án phải đảm bảo mã hóa toàn diện, AWS kiểm soát root of trust (không phải customer-managed HSM), và hỗ trợ audit log. SageMaker tích hợp chặt chẽ với AWS KMS cho mã hóa EBS và S3 (SSE-KMS).

📘 Kiến thức cập nhật đến 2026: Theo tài liệu AWS mới nhất (SageMaker v2.x và KMS 2024+), SageMaker hỗ trợ customer managed keys (CMKs) trong KMS để mã hóa EBS volumes (qua VolumeKmsKeyId), S3 bucket (SSE-KMS), và model artifacts. AWS quản lý root of trust qua FIPS 140-2/3 HSM clusters, với key usage logged tự động qua CloudTrail.

✅ Đáp án đúng

Use customer managed keys in AWS Key Management Service (AWS KMS) to encrypt the ML data volumes, and to encrypt the model artifacts and data in Amazon S3.

Lý do lựa chọn:

  • Customer managed keys (CMKs) trong KMS cho phép mã hóa EBS volumes (SageMaker training storage), model artifacts, và S3 objects qua SSE-KMS.
  • AWS duy trì root of trust: AWS quản lý HSM backing KMS keys (FIPS-validated), khách hàng chỉ quản lý policy/key rotation.
  • Log key usage: Tự động ghi log qua CloudTrail (events như Decrypt, Encrypt), hỗ trợ audit compliance.
  • Hoàn hảo cho SageMaker: Sử dụng KmsKeyId trong training job config và S3 bucket policy. ✅

❌ Giải thích tất cả các phương án

  • [SAI] Use encryption keys that are stored in AWS Cloud HSM to encrypt the ML data volumes, and to encrypt the model artifacts and data in Amazon S3.
    ❌ Sai vì: AWS CloudHSM là HSM do customer quản lý hoàn toàn (self-managed), AWS không duy trì root of trust (customer chịu trách nhiệm HSM appliances, backups). Không tích hợp native với SageMaker EBS/S3 encryption, khó log key usage (phải dùng custom integration với CloudTrail, không tự động). Không phù hợp yêu cầu "AWS maintain root of trust".

  • [SAI] Use SageMaker built-in transient keys to encrypt the ML data volumes. Enable default encryption for new Amazon Elastic Block Store (Amazon EBS) volumes.
    ❌ Sai vì: Transient keys của SageMaker chỉ dùng cho in-transit hoặc temporary (ephemeral) trong training job, không mã hóa at rest vĩnh viễn cho EBS volumes. Default EBS encryption dùng AWS-managed keys (không customer-managed, không log chi tiết key usage). Không đáp ứng mã hóa model artifacts/S3 và root of trust rõ ràng.

  • [ĐÚNG] Use customer managed keys in AWS Key Management Service (AWS KMS) to encrypt the ML data volumes, and to encrypt the model artifacts and data in Amazon S3.
    ✅ Đúng vì: Như giải thích trên, CMKs KMS mã hóa EBS (VolumeKmsKeyId), S3 (SSE-KMS), AWS root of trust, full CloudTrail logging. Tích hợp seamless với SageMaker (2024+ features hỗ trợ key rotation tự động).

  • [SAI] Use AWS Security Token Service (AWS STS) to create temporary tokens to encrypt the ML storage volumes, and to encrypt the model artifacts and data in Amazon S3.
    ❌ Sai vì: AWS STS chỉ tạo temporary credentials (access keys/tokens) cho IAM roles, không phải encryption keys. Không dùng để mã hóa EBS/S3 at rest, không có root of trust cho keys, và không log như KMS. Hoàn toàn không liên quan đến encryption.

📘 Tài liệu tham khảo

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần ví dụ code Terraform/CLI, hãy hỏi thêm.

Câu 146
A machine learning specialist stores IoT soil sensor data in Amazon DynamoDB table and stores weather event data as JSON files in Amazon S3. The dataset in
DynamoDB is 10 GB in size and the dataset in Amazon S3 is 5 GB in size. The specialist wants to train a model on this data to help predict soil moisture levels as a function of weather events using Amazon SageMaker.
Which solution will accomplish the necessary transformation to train the Amazon SageMaker model with the LEAST amount of administrative overhead?
  1. A Launch an Amazon EMR cluster. Create an Apache Hive external table for the DynamoDB table and S3 data. Join the Hive tables and write the results out to Amazon S3.
  2. B Crawl the data using AWS Glue crawlers. Write an AWS Glue ETL job that merges the two tables and writes the output to an Amazon Redshift cluster.
  3. C Enable Amazon DynamoDB Streams on the sensor table. Write an AWS Lambda function that consumes the stream and appends the results to the existing weather files in Amazon S3.
  4. D Crawl the data using AWS Glue crawlers. Write an AWS Glue ETL job that merges the two tables and writes the output in CSV format to Amazon S3.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một machine learning specialist đang xử lý dữ liệu từ hai nguồn:

  • Dữ liệu cảm biến đất IoT (soil sensor data) lưu trong Amazon DynamoDB với kích thước 10 GB.
  • Dữ liệu sự kiện thời tiết (weather event data) lưu dưới dạng file JSON trong Amazon S3 với kích thước 5 GB.

Mục tiêu là train một mô hình học máy trên SageMaker để dự đoán mức độ ẩm đất (soil moisture levels) dựa trên các sự kiện thời tiết. Nhiệm vụ chính là transform dữ liệu (chuyển đổi và kết hợp hai nguồn dữ liệu này) sao cho phù hợp để SageMaker sử dụng, với tiêu chí LEAST amount of administrative overhead (ít chi phí quản trị nhất, nghĩa là serverless, tự động hóa cao, không cần quản lý cluster hay infrastructure thủ công).

🛠️ SageMaker yêu cầu dữ liệu training thường ở định dạng như CSV/Parquet trên S3 để dễ dàng đọc trực tiếp qua các built-in algorithm hoặc custom script. Giải pháp cần tập trung vào ETL (Extract-Transform-Load) hiệu quả, tận dụng các dịch vụ serverless của AWS để giảm thiểu overhead (như quản lý scaling, patching, v.v.).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Crawl the data using AWS Glue crawlers. Write an AWS Glue ETL job that merges the two tables and writes the output in CSV format to Amazon S3.

Lý do chọn đáp án này (theo kiến thức AWS cập nhật đến 2026):

  • AWS Glue là dịch vụ serverless ETL hoàn hảo cho việc crawl metadata (schema discovery) từ DynamoDB và S3 JSON, sau đó tạo ETL job để join/merge hai dataset và output trực tiếp ra CSV trên S3.
  • SageMaker có thể đọc trực tiếp từ S3 (channel input), không cần thêm bước trung gian.
  • Least administrative overhead: Glue tự động scale, pay-per-use, không cần quản lý cluster. Crawler chạy on-demand, ETL job có thể schedule hoặc trigger bằng EventBridge.
  • Hỗ trợ đầy đủ DynamoDB (qua JDBC connector) và S3 JSON (Glue Data Catalog). Định dạng CSV lý tưởng cho ML training.
    📘 Tài liệu tham khảo: AWS Glue Documentation - Crawling DynamoDB/S3 | SageMaker Data Sources.

📋 Giải thích tất cả các phương án (đúng/sai)

  • Phương án A ❌: Launch an Amazon EMR cluster. Create an Apache Hive external table for the DynamoDB table and S3 data. Join the Hive tables and write the results out to Amazon S3.
    Giải thích sai: EMR yêu cầu launch và quản lý cluster (EC2 instances, scaling, termination), tạo overhead cao (admin phải config Hive, bootstrap, monitor). Không serverless, tốn kém cho batch job 15GB. Glue/ Athena hiệu quả hơn cho ETL đơn giản.

  • Phương án B ❌: Crawl the data using AWS Glue crawlers. Write an AWS Glue ETL job that merges the two tables and writes the output to an Amazon Redshift cluster.
    Giải thích sai: Mặc dù dùng Glue crawl/ETL (tốt), nhưng output ra Redshift (data warehouse) là thừa thãi. Redshift cần quản lý cluster (provisioning, scaling, vacuuming), chi phí cao, và SageMaker không cần warehouse – chỉ cần file trên S3. Overhead lớn hơn so với output trực tiếp S3.

  • Phương án C ❌: Enable Amazon DynamoDB Streams on the sensor table. Write an AWS Lambda function that consumes the stream and appends the results to the existing weather files in Amazon S3.
    Giải thích sai: DynamoDB Streams + Lambda dành cho real-time streaming dữ liệu mới (changes), không phù hợp cho historical batch data (10GB cũ). Không merge được với S3 weather data một cách toàn diện, thiếu transform/join, và append JSON thủ công dễ lỗi schema.

  • Phương án D ✅: Crawl the data using AWS Glue crawlers. Write an AWS Glue ETL job that merges the two tables and writes the output in CSV format to Amazon S3.
    Giải thích đúng: Hoàn hảo serverless end-to-end: Crawler tự infer schema từ DynamoDB/S3, ETL job (Spark-based) join/merge dễ dàng (PySpark/SQL), output CSV chuẩn ML trên S3. Zero admin (Glue Data Catalog tích hợp SageMaker), scale tự động cho 15GB data. Cập nhật 2026: Glue 4.0 hỗ trợ Spark 3.3+ với performance tốt hơn.

🧩 Tóm tắt lợi ích tổng thể: Giải pháp đúng tận dụng Glue Data Catalog làm unified metadata layer, cho phép SageMaker training channel chỉ định s3://bucket/path/*.csv. Tiết kiệm thời gian từ crawl (5-10 phút) đến ETL (dưới 1 giờ cho 15GB). Tránh các dịch vụ managed cluster như EMR/Redshift để minimize overhead!

Câu 147
A company sells thousands of products on a public website and wants to automatically identify products with potential durability problems. The company has
1.000 reviews with date, star rating, review text, review summary, and customer email fields, but many reviews are incomplete and have empty fields. Each review has already been labeled with the correct durability result.
A machine learning specialist must train a model to identify reviews expressing concerns over product durability. The first model needs to be trained and ready to review in 2 days.
What is the MOST direct approach to solve this problem within 2 days?
  1. A Train a custom classifier by using Amazon Comprehend.
  2. B Build a recurrent neural network (RNN) in Amazon SageMaker by using Gluon and Apache MXNet.
  3. C Train a built-in BlazingText model using Word2Vec mode in Amazon SageMaker.
  4. D Use a built-in seq2seq model in Amazon SageMaker.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi mô tả một công ty bán hàng trực tuyến với hàng nghìn sản phẩm, muốn tự động phát hiện sản phẩm có vấn đề về độ bền dựa trên dữ liệu đánh giá (reviews). Họ có 1.000 reviews đã được labeled sẵn (gắn nhãn kết quả độ bền đúng), bao gồm các trường: ngày đánh giá, số sao, nội dung đánh giá, tóm tắt đánh giá, và email khách hàng. Tuy nhiên, nhiều reviews không đầy đủ (có trường trống).
Nhiệm vụ: Train một mô hình ML để phân loại reviews thể hiện lo ngại về độ bền sản phẩm. Yêu cầu nhanh nhất (model sẵn sàng review trong 2 ngày), nên cần cách tiếp cận trực tiếp nhất (MOST direct approach).
🔑 Vấn đề cốt lõi: Đây là bài toán text classification (phân loại văn bản) trên dữ liệu labeled nhỏ (1.000 mẫu), dữ liệu thô (incomplete), thời gian eo hẹp → ưu tiên dịch vụ managed, low-code để train nhanh mà không cần code phức tạp.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Train a custom classifier by using Amazon Comprehend.
🛠️ Lý do:

  • Amazon Comprehend Custom Classifier là dịch vụ fully managed chuyên cho text classification, hỗ trợ train trên dữ liệu labeled (CSV/JSON với labels). Chỉ cần upload dữ liệu labeled → train model trong vài giờ (thường <1 ngày với 1.000 mẫu), phù hợp deadline 2 ngày.
  • Xử lý tốt dữ liệu incomplete (tự động ignore trường trống, focus vào text như review text/summary).
  • Không cần code custom, low-code: Console/API upload → train → deploy endpoint inference ngay.
  • Cập nhật 2026: Comprehend Custom Classifier vẫn là lựa chọn tối ưu cho classification nhanh (hỗ trợ multi-class, multilingual).
    📘 Tài liệu tham khảo: AWS Comprehend Custom Classification (train time ~hours cho small dataset).

📋 Giải thích tất cả các phương án

  • Train a custom classifier by using Amazon Comprehend.
    ✅ Đúng 🏆: Như trên, trực tiếp nhất (direct) vì managed service, train nhanh (upload labeled data → done), không cần infra/code. Hoàn hảo cho 1.000 reviews labeled, xử lý incomplete data tự động.

  • Build a recurrent neural network (RNN) in Amazon SageMaker by using Gluon and Apache MXNet.
    ❌ Sai: RNN (Gluon/MXNet) yêu cầu code custom từ scratch (preprocess data, define model, train script), setup SageMaker notebook/estimator. Với 2 ngày, quá phức tạp/tốn thời gian (data cleaning + hyperparam tuning). MXNet deprecated 2024, không khuyến khích 2026.

  • Train a built-in BlazingText model using Word2Vec mode in Amazon SageMaker.
    ❌ Sai: BlazingText Word2Vec mode chỉ tạo word embeddings (unsupervised, không classify). Không dùng cho classification labeled. (BlazingText có classification mode supervised, nhưng option chỉ định "Word2Vec mode" → sai). Thời gian preprocess cũng lâu hơn Comprehend.

  • Use a built-in seq2seq model in Amazon SageMaker.
    ❌ Sai: Seq2seq (như trong JumpStart/HuggingFace) dùng cho sequence-to-sequence tasks (translation, summarization), không phải classification. Cần fine-tune custom, tốn thời gian >2 ngày với data nhỏ/incomplete.

🔍 Tóm tắt so sánh: Comprehend thắng vì zero-code, fastest time-to-model (🕐<2 ngày), các option SageMaker yêu cầu engineering effort cao (code/data prep). Theo best practice AWS ML (2026): Dùng Comprehend cho simple text classification trước SageMaker custom.
📘 Nguồn bổ sung: AWS SageMaker BlazingText Docs, Comprehend vs SageMaker Comparison.

Câu 148
A company that runs an online library is implementing a chatbot using Amazon Lex to provide book recommendations based on category. This intent is fulfilled by an AWS Lambda function that queries an Amazon DynamoDB table for a list of book titles, given a particular category. For testing, there are only three categories implemented as the custom slot types: "comedy," "adventure,` and "documentary.`
A machine learning (ML) specialist notices that sometimes the request cannot be fulfilled because Amazon Lex cannot understand the category spoken by users with utterances such as "funny," "fun," and "humor." The ML specialist needs to fix the problem without changing the Lambda code or data in DynamoDB.
How should the ML specialist fix the problem?
  1. A Add the unrecognized words in the enumeration values list as new values in the slot type.
  2. B Create a new custom slot type, add the unrecognized words to this slot type as enumeration values, and use this slot type for the slot.
  3. C Use the AMAZON.SearchQuery built-in slot types for custom searches in the database.
  4. D Add the unrecognized words as synonyms in the custom slot type.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống thực tế trong AWS Amazon Lex:
Một công ty vận hành thư viện trực tuyến đang triển khai chatbot bằng Amazon Lex để gợi ý sách dựa trên category (danh mục). Intent này được fulfill bởi AWS Lambda function, Lambda sẽ query bảng Amazon DynamoDB để lấy danh sách sách theo category cụ thể.
Hiện tại, slot type tùy chỉnh (custom slot type) chỉ có 3 giá trị enumeration: "comedy", "adventure", và "documentary".
🛠️ Vấn đề: Machine Learning (ML) specialist nhận thấy Lex đôi khi không hiểu các utterance (câu nói) như "funny", "fun", "humor" (đều ám chỉ "comedy"). Điều này dẫn đến request không được fulfill.
Yêu cầu fix: Phải giải quyết mà KHÔNG thay đổi code Lambda (vì Lambda query dựa trên giá trị chính xác trong DynamoDB) và KHÔNG thay đổi data trong DynamoDB (dữ liệu sách vẫn giữ nguyên category gốc như "comedy").
📘 Kiến thức cốt lõi (cập nhật AWS Lex V2 đến 2026): Amazon Lex sử dụng slot types để extract giá trị từ user input. Custom slot types hỗ trợ enumeration values (giá trị liệt kê chính) và synonyms (từ đồng nghĩa map về giá trị chính), giúp Lex hiểu biến thể ngôn ngữ tự nhiên mà không làm thay đổi giá trị slot được gửi đến Lambda.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Add the unrecognized words as synonyms in the custom slot type.

Lý do chi tiết:
🛠️ Trong Amazon Lex, cơ chế synonyms cho phép thêm các từ như "funny", "fun", "humor" làm synonyms của giá trị enumeration "comedy" trong custom slot type hiện tại. Khi user nói "funny", Lex sẽ match và fill slot với giá trị gốc "comedy", gửi chính xác đến Lambda để query DynamoDB.
✅ Ưu điểm: Không thay đổi code Lambda hay data DynamoDB, tận dụng ML của Lex để xử lý ngôn ngữ tự nhiên (NLU). Đây là best practice theo AWS Lex documentation (V2 models hỗ trợ synonyms động hơn từ 2023-2026).
❌ Không cần tạo slot mới hay thay enum, giữ nguyên thiết kế hiện tại.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên nội dung phương án gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm lý do bằng tiếng Việt rõ ràng:

  • Add the unrecognized words in the enumeration values list as new values in the slot type.
    ❌ Sai hoàn toàn. Thêm "funny", "fun", "humor" trực tiếp vào enumeration values sẽ làm slot type có giá trị mới (như "funny"). Khi Lex fill slot với "funny", Lambda sẽ query DynamoDB với "funny" → không tìm thấy sách (vì data chỉ có "comedy"). Vi phạm yêu cầu không thay đổi Lambda/DynamoDB, gây lỗi runtime. 🧩 Không tận dụng synonyms, làm slot type phình to không cần thiết.

  • Create a new custom slot type, add the unrecognized words to this slot type as enumeration values, and use this slot type for the slot.
    ❌ Sai. Tạo slot type mới với "funny" làm enum values, rồi assign cho slot → Lex fill slot với "funny" gửi đến Lambda → query DynamoDB thất bại tương tự phương án trên. 🛠️ Phức tạp hóa thiết kế (phải update intent/slot), không giải quyết root cause ngôn ngữ tự nhiên, và vẫn vi phạm không thay đổi backend. AWS khuyến nghị dùng synonyms thay vì slot mới.

  • Use the AMAZON.SearchQuery built-in slot types for custom searches in the database.
    ❌ Sai. AMAZON.SearchQuery là built-in slot type cho search queries tự do (như tìm kiếm văn bản mở), không phù hợp cho category enum cố định (chỉ 3 giá trị). Nó sẽ extract chuỗi tự do như "funny books" → Lambda phải parse phức tạp (cần thay code), không chính xác cho match exact DynamoDB. 📘 Theo AWS docs (2026), built-in này dành cho search engine, không thay thế custom enum + synonyms.

  • Add the unrecognized words as synonyms in the custom slot type.
    ✅ Đúng 100%. Như đã giải thích ở phần đáp án, synonyms map "funny" → "comedy" (giá trị slot giữ nguyên), Lex xử lý NLU tự động. Lambda nhận "comedy" → query thành công. 🛠️ Giải pháp tối ưu, scalable, không ảnh hưởng backend.

📘 Tài liệu tham khảo (AWS cập nhật mới nhất đến 2026)

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần demo code Lex bot, hãy hỏi thêm.

Câu 149
A manufacturing company uses machine learning (ML) models to detect quality issues. The models use images that are taken of the company's product at the end of each production step. The company has thousands of machines at the production site that generate one image per second on average.
The company ran a successful pilot with a single manufacturing machine. For the pilot, ML specialists used an industrial PC that ran AWS IoT Greengrass with a long-running AWS Lambda function that uploaded the images to Amazon S3. The uploaded images invoked a Lambda function that was written in Python to perform inference by using an Amazon SageMaker endpoint that ran a custom model. The inference results were forwarded back to a web service that was hosted at the production site to prevent faulty products from being shipped.
The company scaled the solution out to all manufacturing machines by installing similarly configured industrial PCs on each production machine. However, latency for predictions increased beyond acceptable limits. Analysis shows that the internet connection is at its capacity limit.
How can the company resolve this issue MOST cost-effectively?
  1. A Set up a 10 Gbps AWS Direct Connect connection between the production site and the nearest AWS Region. Use the Direct Connect connection to upload the images. Increase the size of the instances and the number of instances that are used by the SageMaker endpoint.
  2. B Extend the long-running Lambda function that runs on AWS IoT Greengrass to compress the images and upload the compressed files to Amazon S3. Decompress the files by using a separate Lambda function that invokes the existing Lambda function to run the inference pipeline.
  3. C Use auto scaling for SageMaker. Set up an AWS Direct Connect connection between the production site and the nearest AWS Region. Use the Direct Connect connection to upload the images.
  4. D Deploy the Lambda function and the ML models onto the AWS IoT Greengrass core that is running on the industrial PCs that are installed on each machine. Extend the long-running Lambda function that runs on AWS IoT Greengrass to invoke the Lambda function with the captured images and run the inference on the edge component that forwards the results directly to the web service.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty sản xuất sử dụng mô hình Machine Learning (ML) để phát hiện lỗi chất lượng sản phẩm qua hình ảnh chụp ở cuối mỗi bước sản xuất. Họ có hàng ngàn máy sản xuất, mỗi máy tạo ra 1 hình ảnh/giây trung bình, dẫn đến lượng dữ liệu khổng lồ (khoảng 86 triệu ảnh/ngày nếu tính trung bình).

  • Pilot thành công: Sử dụng AWS IoT Greengrass trên một PC công nghiệp với Lambda long-running để upload ảnh lên Amazon S3. Ảnh trigger Lambda Python chạy inference qua Amazon SageMaker endpoint (mô hình custom). Kết quả trả về web service tại chỗ để ngăn sản phẩm lỗi xuất xưởng.
  • Vấn đề khi scale: Cài PC tương tự trên tất cả máy, nhưng latency dự đoán tăng vượt ngưỡng chấp nhận do kết nối internet đạt giới hạn dung lượng (bandwidth nghẽn vì upload hàng triệu ảnh/giây).

Mục tiêu: Giải quyết MOST cost-effectively (tiết kiệm chi phí nhất), tập trung vào việc giảm tải internet mà không làm phức tạp hạ tầng. Đây là tình huống điển hình cho edge computing với AWS IoT Greengrass Version 2 (cập nhật đến 2026, hỗ trợ ML inference local qua SageMaker Neo, containers, và Lambda). 📘 Tài liệu tham khảo: AWS IoT Greengrass ML Inference, SageMaker Edge Manager.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Deploy the Lambda function and the ML models onto the AWS IoT Greengrass core that is running on the industrial PCs that are installed on each machine. Extend the long-running Lambda function that runs on AWS IoT Greengrass to invoke the Lambda function with the captured images and run the inference on the edge component that forwards the results directly to the web service.

Lý do chọn 🛠️:

  • Giải quyết gốc rễ vấn đề: Chạy inference hoàn toàn trên edge (local tại PC công nghiệp), không upload ảnh lên cloud → loại bỏ hoàn toàn bottleneck internet, giảm latency xuống mức pilot (dưới 1s).
  • Cost-effective nhất: Tận dụng hạ tầng hiện có (PC + Greengrass core), chỉ deploy Lambda + ML models local qua SageMaker Neo hoặc Docker containers trên Greengrass v2. Không tốn phí bandwidth S3/Direct Connect, không scale cloud resources. Chi phí chỉ là Greengrass (miễn phí core, tính theo component usage).
  • Cập nhật AWS 2026: Greengrass v2 hỗ trợ ML inference offline với models từ SageMaker (export qua Neo compiler cho TensorFlow/PyTorch), Lambda long-running invoke local functions seamless. Kết quả forward trực tiếp web service tại chỗ. ✅ Hoàn hảo cho IoT industrial scale lớn!

❌ Phân tích tất cả các phương án

  • Phương án 1 (SAI): Set up a 10 Gbps AWS Direct Connect connection between the production site and the nearest AWS Region. Use the Direct Connect connection to upload the images. Increase the size of the instances and the number of instances that are used by the SageMaker endpoint.

    • Lý do sai ❌: Không cost-effective – Direct Connect 10Gbps rất đắt (port hour + data transfer), cộng thêm scale SageMaker instances (ml.* lớn hơn → chi phí EC2 cao). Vẫn upload toàn bộ ảnh (hàng TB/ngày) → latency cloud-bound vẫn cao do queue S3/Lambda. Không tận dụng edge, chỉ "băng rộng hóa" vấn đề!
  • Phương án 2 (SAI): Extend the long-running Lambda function that runs on AWS IoT Greengrass to compress the images and upload the compressed files to Amazon S3. Decompress the files by using a separate Lambda function that invokes the existing Lambda function to run the inference pipeline.

    • Lý do sai ❌: Compress giúp giảm kích thước file, nhưng vẫn phải upload → bandwidth vẫn nghẽn (dù ít hơn), cộng thêm overhead decompress + 2 Lambda tiers → tăng latency và chi phí invocation/data transfer S3. Không giải quyết scale 1000+ máy, chỉ "vá víu" tạm thời!
  • Phương án 3 (SAI): Use auto scaling for SageMaker. Set up an AWS Direct Connect connection between the production site and the nearest AWS Region. Use the Direct Connect connection to upload the images.

    • Lý do sai ❌: Tương tự phương án 1, Direct Connect đắt đỏ + autoscaling SageMaker (dựa traffic) vẫn tốn kém cho real-time inference (hàng triệu request/giây). Upload ảnh vẫn là bottleneck chính, auto scaling chỉ xử lý backend cloud chứ không giảm latency edge-to-cloud. Không tiết kiệm!
  • Phương án 4 (ĐÚNG): Deploy the Lambda function and the ML models onto the AWS IoT Greengrass core that is running on the industrial PCs that are installed on each machine. Extend the long-running Lambda function that runs on AWS IoT Greengrass to invoke the Lambda function with the captured images and run the inference on the edge component that forwards the results directly to the web service.

    • Lý do đúng ✅ (như phần trên): Edge-native, zero-upload, low-latency, low-cost. Hoàn toàn phù hợp best practice AWS cho industrial IoT/ML (xem Greengrass ML Components). 🏆

Kết luận 🚀: Giải pháp edge computing với Greengrass là optimal cho high-volume IoT, tiết kiệm >90% chi phí so với cloud-only. Nếu triển khai, test với Greengrass ML Compile để optimize models! 📘

Câu 150
A data scientist is using an Amazon SageMaker notebook instance and needs to securely access data stored in a specific Amazon S3 bucket.
How should the data scientist accomplish this?
  1. A Add an S3 bucket policy allowing GetObject, PutObject, and ListBucket permissions to the Amazon SageMaker notebook ARN as principal.
  2. B Encrypt the objects in the S3 bucket with a custom AWS Key Management Service (AWS KMS) key that only the notebook owner has access to.
  3. C Attach the policy to the IAM role associated with the notebook that allows GetObject, PutObject, and ListBucket operations to the specific S3 bucket.
  4. D Use a script in a lifecycle configuration to configure the AWS CLI on the instance with an access key ID and secret.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS

📘 Nội dung câu hỏi:
Câu hỏi tập trung vào việc một data scientist đang sử dụng Amazon SageMaker notebook instance và cần truy cập an toàn vào dữ liệu lưu trữ trong một Amazon S3 bucket cụ thể. SageMaker notebook instance là môi trường Jupyter Notebook được quản lý bởi AWS, chạy trên EC2 instances, và nó kế thừa quyền truy cập từ IAM role được gắn vào instance. Mục tiêu là đảm bảo quyền truy cập GetObject, PutObject, và ListBucket một cách bảo mật, tuân thủ nguyên tắc least privilege và best practices của AWS (không sử dụng access keys tĩnh). Điều này liên quan đến IAM policies, S3 bucket policies, và security best practices trong SageMaker (cập nhật theo AWS Well-Architected Framework và SageMaker Security Guide đến năm 2026).

✅ Đáp án đúng:
Attach the policy to the IAM role associated with the notebook that allows GetObject, PutObject, and ListBucket operations to the specific S3 bucket.

Lý do chọn đáp án này:
🛠️ Đây là phương pháp tốt nhất và an toàn nhất theo khuyến nghị của AWS. SageMaker notebook instance sử dụng execution role (IAM role) để truy cập các dịch vụ AWS như S3 mà không cần credentials tĩnh. Bạn tạo một IAM policy cho phép các action cụ thể trên S3 bucket, sau đó attach vào role của notebook instance (qua console SageMaker hoặc API CreateNotebookInstance). Điều này đảm bảo quyền truy cập tạm thời qua instance metadata, tuân thủ zero-trust model, và dễ quản lý/audit qua CloudTrail. Không cần thay đổi bucket policy hoặc credentials thủ công.

🔍 Giải thích chi tiết từng phương án trả lời

  • ❌ [SAI] Add an S3 bucket policy allowing GetObject, PutObject, and ListBucket permissions to the Amazon SageMaker notebook ARN as principal.
    Phương án này sai vì SageMaker notebook instance không có ARN có thể dùng làm principal trong S3 bucket policy. ARN của notebook là dạng arn:aws:sagemaker:region:account:notebook-instance/notebook-name, nhưng nó không phải là entity xác thực (như IAM user/role). Bucket policy yêu cầu principal là IAM principal (user/role), không phải resource ARN. Sử dụng cách này sẽ bị từ chối và không grant quyền đúng cách, vi phạm best practices (IAM role mới là cách chính thức).

  • ❌ [SAI] Encrypt the objects in the S3 bucket with a custom AWS Key Management Service (AWS KMS) key that only the notebook owner has access to.
    Phương án này không giải quyết vấn đề truy cập, chỉ tập trung vào encryption at rest. KMS key policy có thể grant quyền decrypt cho IAM role, nhưng vẫn cần IAM policy riêng cho S3 actions (GetObject, etc.). Nếu không có quyền S3 cơ bản, data scientist vẫn không đọc/ghi được dù object đã encrypt. Đây là phụ trợ, không phải giải pháp chính cho access control.

  • ✅ [ĐÚNG] Attach the policy to the IAM role associated with the notebook that allows GetObject, PutObject, and ListBucket operations to the specific S3 bucket.
    Như đã giải thích ở trên, đây là best practice của AWS: Sử dụng IAM role policy với resource-specific permissions (arn:aws:s3:::bucket/*). Ví dụ policy JSON:

    {
      "Version": "2012-10-17",
      "Statement": [{
        "Effect": "Allow",
        "Action": ["s3:GetObject", "s3:PutObject", "s3:ListBucket"],
        "Resource": ["arn:aws:s3:::your-bucket", "arn:aws:s3:::your-bucket/*"]
      }]
    }
    

    Attach vào role qua IAM console, sau đó chỉ định role khi tạo notebook instance.

  • ❌ [SAI] Use a script in a lifecycle configuration to configure the AWS CLI on the instance with an access key ID and secret.
    Phương án này rất không an toàn và vi phạm AWS security best practices (không bao giờ lưu access keys tĩnh trên instance). Lifecycle configuration chỉ dùng để install packages/setup môi trường, nhưng hardcode credentials có nguy cơ leak qua logs/code, dễ bị hack. SageMaker khuyến cáo dùng IAM roles thay thế, và tính năng này đã bị deprecated trong các hướng dẫn mới (2023-2026) để tránh shared credentials.

📚 Tài liệu tham khảo (cập nhật mới nhất AWS 2026)

Hy vọng phân tích này giúp bạn ôn thi AWS Certified DevOps Engineer Professional hiệu quả! 🚀 Nếu cần ví dụ code Terraform/ CDK, hãy hỏi thêm.