Ngân hàng đề — AWS Certified Machine Learning Specialty

Tìm thấy 371 câu.

Câu 231 Chọn nhiều đáp án
An online retail company wants to develop a natural language processing (NLP) model to improve customer service. A machine learning (ML) specialist is setting up distributed training of a Bidirectional Encoder Representations from Transformers (BERT) model on Amazon SageMaker. SageMaker will use eight compute instances for the distributed training.

The ML specialist wants to ensure the security of the data during the distributed training. The data is stored in an Amazon S3 bucket.

Which combination of steps should the ML specialist take to protect the data during the distributed training? (Choose three.)
  1. A Run distributed training jobs in a private VPC. Enable inter-container traffic encryption.
  2. B Run distributed training jobs across multiple VPCs. Enable VPC peering.
  3. C Create an S3 VPC endpoint. Then configure network routes, endpoint policies, and S3 bucket policies.
  4. D Grant read-only access to SageMaker resources by using an IAM role.
  5. E Create a NAT gateway. Assign an Elastic IP address for the NAT gateway.
  6. F Configure an inbound rule to allow traffic from a security group that is associated with the training instances.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi tập trung vào việc bảo mật dữ liệu trong quá trình distributed training một mô hình BERT (một loại mô hình NLP phổ biến) trên Amazon SageMaker sử dụng 8 compute instances. Dữ liệu huấn luyện được lưu trữ trong Amazon S3 bucket.

  • Bối cảnh: Công ty bán lẻ trực tuyến muốn xây dựng mô hình NLP để cải thiện dịch vụ khách hàng. ML specialist cần thiết lập distributed training an toàn, tránh rò rỉ dữ liệu qua internet hoặc các lỗ hổng mạng/IAM.
  • Yêu cầu chính: Chọn 3 bước kết hợp để bảo vệ dữ liệu trong quá trình training phân tán (distributed training), bao gồm bảo mật mạng (network security), mã hóa traffic giữa các container, và kiểm soát quyền truy cập.
  • Kiến thức AWS cập nhật (đến 2026): SageMaker hỗ trợ distributed training với các tính năng như SageMaker Training sử dụng TensorFlow/ Hugging Face containers, VPC-only mode, inter-container traffic encryption (mặc định từ SageMaker 2023+), S3 VPC endpoints (Gateway endpoints miễn phí), và IAM roles với least privilege (theo AWS Well-Architected Framework for ML).

Mục tiêu là đảm bảo dữ liệu không đi qua public internet, traffic nội bộ được mã hóa, và quyền truy cập hạn chế chỉ read-only.

✅ Đáp án đúng và lý do lựa chọn

Các đáp án đúng là 3 lựa chọn sau, tạo thành sự kết hợp hoàn chỉnh để bảo mật dữ liệu end-to-end trong distributed training trên SageMaker:

  1. Run distributed training jobs in a private VPC. Enable inter-container traffic encryption.
    ✅ Lý do: Chạy job trong private VPC (VPC-only mode của SageMaker) ngăn dữ liệu ra public internet. Inter-container traffic encryption (sử dụng TLS 1.2+) mã hóa giao tiếp giữa 8 instances, bảo vệ dữ liệu nhạy cảm trong quá trình all-reduce hoặc gradient sync (theo SageMaker Distributed Training best practices).

  2. Create an S3 VPC endpoint. Then configure network routes, endpoint policies, and S3 bucket policies.
    ✅ Lý do: S3 VPC Gateway Endpoint cho phép SageMaker instances truy cập S3 privately mà không qua internet/NAT. Cấu hình route tables, endpoint policies (full access hoặc resource-specific), và S3 bucket policies (deny public access) đảm bảo dữ liệu chỉ di chuyển nội bộ VPC, tuân thủ zero-trust model.

  3. Grant read-only access to SageMaker resources by using an IAM role.
    ✅ Lý do: Sử dụng IAM execution role cho SageMaker với least privilege (chỉ s3:GetObject, s3:ListBucket read-only) ngăn chặn write/delete hoặc escalation. SageMaker tự động assume role này, tích hợp SageMaker ExecutionRoleForTraining template, giảm rủi ro nếu instances bị compromise.

Kết hợp này tạo lớp bảo mật đa tầng: Network isolation + Encryption in transit + Access control, phù hợp với AWS Shared Responsibility Model cho ML workloads.

🛠️ Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng phương án một cách đầy đủ, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm giải thích rõ ràng dựa trên best practices AWS SageMaker security (cập nhật 2026).

  • ✅ Run distributed training jobs in a private VPC. Enable inter-container traffic encryption.
    Giải thích đúng: Đây là bước cốt lõi để isolate training environment khỏi public internet (VPC-only mode chỉ cho phép traffic nội bộ). Inter-container encryption được enable qua EnableInterContainerTrafficEncryption: true trong training job config, bảo vệ dữ liệu NLP nhạy cảm (customer data) giữa các instances trong cluster 8-node. Không có bước này, traffic có thể bị sniff nếu dùng public subnet.

  • ❌ Run distributed training jobs across multiple VPCs. Enable VPC peering.
    Giải thích sai: SageMaker distributed training không hỗ trợ cross-VPC trực tiếp cho single job (phải cùng VPC/Subnet). VPC peering thêm complexity/overhead (latency cao cho all-reduce), không cần thiết và tăng attack surface (peering config dễ misconfig). Best practice là single private VPC thay vì multi-VPC.

  • ✅ Create an S3 VPC endpoint. Then configure network routes, endpoint policies, and S3 bucket policies.
    Giải thích đúng: S3 Gateway Endpoint (interface/gateway) route traffic S3 qua private IP, bypass internet hoàn toàn. Routes (add endpoint to route table), endpoint policies (IAM-like policy cho endpoint), và bucket policies (condition on vpc-endpoint) chặn access ngoài VPC, đảm bảo dữ liệu chỉ tải về instances an toàn. Tiết kiệm chi phí so với NAT.

  • ✅ Grant read-only access to SageMaker resources by using an IAM role.
    Giải thích đúng: IAM role cho SageMaker (attach policy với s3:GetObject* read-only) áp dụng principle of least privilege. SageMaker sử dụng role này để access S3, ngăn write operations hoặc lateral movement. Tích hợp AmazonSageMakerFullAccess nhưng customize read-only, kiểm tra qua CloudTrail.

  • ❌ Create a NAT gateway. Assign an Elastic IP address for the NAT gateway.
    Giải thích sai: NAT Gateway + EIP dùng cho outbound internet access từ private subnet (ví dụ download models từ public ECR). Nhưng câu hỏi tập trung protect data during training, NAT expose traffic ra internet (dù chỉ outbound), không an toàn cho dữ liệu S3 private. Thay vào đó, dùng VPC endpoints để tránh NAT hoàn toàn.

  • ❌ Configure an inbound rule to allow traffic from a security group that is associated with the training instances.
    Giải thích sai: Inbound rule từ self-referencing SG (allow traffic từ SG của chính instances) chỉ cần cho inter-instance comm (SageMaker auto-config trong VPC). Nhưng expose port (như 8080/8888 cho training) tăng rủi ro nếu misconfig NACL/SG, không phải bước bảo mật chính. SageMaker quản lý traffic nội bộ tự động, không khuyến khích manual inbound mở rộng.

📘 Tài liệu tham khảo

  • AWS SageMaker Documentation (2026): Secure SageMaker Training Jobs – VPC mode & inter-container encryption.
  • Amazon S3 Security Best Practices: S3 VPC Endpoints.
  • IAM for SageMaker: Execution Roles.
  • AWS Well-Architected ML Lens: Security Pillar – Multi-layer protection cho distributed training.
  • Exam Prep: AWS Certified Machine Learning Specialty (MLS-C01) & DevOps Pro (DOP-C02) – Topic: SageMaker Networking & Security.

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần thêm ví dụ code Terraform/CLI, hãy hỏi nhé!

Câu 232
An analytics company has an Amazon SageMaker hosted endpoint for an image classification model. The model is a custom-built convolutional neural network (CNN) and uses the PyTorch deep learning framework. The company wants to increase throughput and decrease latency for customers that use the model.

Which solution will meet these requirements MOST cost-effectively?
  1. A Use Amazon Elastic Inference on the SageMaker hosted endpoint.
  2. B Retrain the CNN with more layers and a larger dataset.
  3. C Retrain the CNN with more layers and a smaller dataset.
  4. D Choose a SageMaker instance type that has multiple GPUs.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc tối ưu hóa hiệu suất cho một endpoint Amazon SageMaker hosted, sử dụng mô hình phân loại hình ảnh là mạng nơ-ron tích chập (CNN) tùy chỉnh dựa trên framework PyTorch. 🖼️ Công ty muốn tăng throughput (số lượng yêu cầu xử lý mỗi giây) và giảm latency (thời gian phản hồi) cho khách hàng sử dụng mô hình này, đồng thời phải chọn giải pháp tiết kiệm chi phí nhất (MOST cost-effectively).

🔍 Bối cảnh kỹ thuật: SageMaker hosted endpoint là dịch vụ managed để deploy và scale mô hình ML cho inference thời gian thực. Mô hình CNN PyTorch thường yêu cầu tài nguyên GPU để xử lý nhanh các phép tính ma trận phức tạp trong inference. Vấn đề là cần cải thiện performance mà không tốn kém quá mức, tránh nâng cấp instance lớn hoặc retrain mô hình (vốn tốn thời gian và chi phí cao). 📈

✅ Đáp án đúng: Use Amazon Elastic Inference on the SageMaker hosted endpoint

Lý do lựa chọn:

  • Amazon Elastic Inference (EI) là giải pháp tối ưu nhất về chi phí vì nó cho phép gắn thêm accelerator inference chuyên dụng (như GPU nhỏ) vào instance hiện tại mà không cần thay thế toàn bộ instance bằng loại có GPU đầy đủ. 🛠️
  • EI tăng tốc inference cho các framework như PyTorch, đặc biệt hiệu quả với CNN, giúp tăng throughput lên đến 10x và giảm latency đáng kể mà chỉ tốn thêm ~20-50% chi phí so với instance gốc (thay vì gấp đôi hoặc hơn với multi-GPU).
  • SageMaker hỗ trợ EI trực tiếp trên hosted endpoint, dễ tích hợp qua sagemaker.pytorch.PyTorchModel với tham số elastic_inference=True. 🚀
  • Đây là cách cost-effective nhất vì linh hoạt scale theo nhu cầu, pay-per-use, và không yêu cầu retrain mô hình. Theo cập nhật AWS đến 2026, EI vẫn là best practice cho inference acceleration trên SageMaker (hỗ trợ PyTorch 2.x+).

📋 Giải thích tất cả các phương án (đúng/sai)

  • Use Amazon Elastic Inference on the SageMaker hosted endpoint.
    ✅ Đúng 🏆: Như đã giải thích, EI cung cấp acceleration inference hiệu quả, tăng throughput/giảm latency mà chi phí thấp nhất (chỉ thêm inference chip cần thiết). Không ảnh hưởng đến training, deploy nhanh chóng. Lý tưởng cho CNN PyTorch trên SageMaker.

  • Retrain the CNN with more layers and a larger dataset.
    ❌ Sai 💸: Retrain với nhiều layers hơn và dataset lớn hơn sẽ tăng độ phức tạp mô hình, dẫn đến inference chậm hơn (cao hơn latency) và cần tài nguyên lớn hơn. Quá trình retrain tốn kém (thời gian, compute), không giải quyết trực tiếp vấn đề endpoint hiện tại, và không cost-effective.

  • Retrain the CNN with more layers and a smaller dataset.
    ❌ Sai 🚫: Thêm layers nhưng dataset nhỏ hơn dễ gây overfitting (mô hình kém generalize), làm giảm accuracy và performance inference. Retrain vẫn tốn kém, không cải thiện throughput/latency mà có thể tệ hơn.

  • Choose a SageMaker instance type that has multiple GPUs.
    ❌ Sai 📈: Instance multi-GPU (như ml.p3.8xlarge) tăng performance mạnh nhưng chi phí cao gấp nhiều lần (GPU full > EI), không linh hoạt (overprovision nếu traffic thấp). Không phải MOST cost-effectively vì EI rẻ hơn cho cùng hiệu suất inference.

📘 Tài liệu tham khảo (cập nhật AWS đến 2026)

  • AWS SageMaker Documentation: Elastic Inference – Hướng dẫn tích hợp EI với PyTorch endpoints.
  • AWS Blog: Accelerate inference with Elastic Inference (vẫn active, hỗ trợ PyTorch 2.0+).
  • Exam Guide DOP-C02: Phần SageMaker Optimization (Best Practices 2024-2026).
    🔗 Kiểm tra AWS Console > SageMaker > Endpoints > Elastic Inference để test thực tế!
Câu 233 Chọn nhiều đáp án
An ecommerce company is collecting structured data and unstructured data from its website, mobile apps, and IoT devices. The data is stored in several databases and Amazon S3 buckets. The company is implementing a scalable repository to store structured data and unstructured data. The company must implement a solution that provides a central data catalog, self-service access to the data, and granular data access policies and encryption to protect the data.

Which combination of actions will meet these requirements with the LEAST amount of setup? (Choose three.)
  1. A Identify the existing data in the databases and S3 buckets. Link the data to AWS Lake Formation.
  2. B Identify the existing data in the databases and S3 buckets. Link the data to AWS Glue.
  3. C Run AWS Glue crawlers on the linked data sources to create a central data catalog.
  4. D Apply granular access policies by using AWS Identity and Access Management (1AM). Configure server-side encryption on each data source.
  5. E Apply granular access policies and encryption by using AWS Lake Formation.
  6. F Apply granular access policies and encryption by using AWS Glue.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi mô tả một công ty thương mại điện tử đang thu thập dữ liệu có cấu trúc (structured data) và dữ liệu không cấu trúc (unstructured data) từ website, ứng dụng di động và thiết bị IoT. Dữ liệu hiện được lưu trữ ở nhiều cơ sở dữ liệu (databases) và Amazon S3 buckets. Công ty cần triển khai một kho lưu trữ có khả năng mở rộng (scalable repository) để lưu trữ cả hai loại dữ liệu này, đồng thời đáp ứng các yêu cầu sau:

  • Central data catalog: Danh mục dữ liệu trung tâm để dễ dàng khám phá và quản lý.
  • Self-service access: Truy cập tự phục vụ cho người dùng.
  • Granular data access policies và encryption: Chính sách kiểm soát truy cập chi tiết (fine-grained) và mã hóa để bảo vệ dữ liệu.

Yêu cầu chọn kết hợp 3 hành động (combination of actions) giúp đáp ứng với ít thiết lập nhất (LEAST amount of setup). Đây là câu hỏi kiểu "chọn 3" trong kỳ thi AWS Certified DevOps Engineer Professional, tập trung vào AWS Lake Formation – dịch vụ xây dựng data lake hiện đại, tích hợp sâu với AWS Glue cho catalog và permissions (cập nhật đến 2026, Lake Formation hỗ trợ governed tables, row/column-level security, và tự động encryption với SSE-KMS/SSE-S3).

✅ Đáp án đúng và lý do lựa chọn

Các đáp án đúng là 3 lựa chọn sau, tạo thành quy trình tối ưu với ít setup nhất nhờ tận dụng AWS Lake Formation làm trung tâm (data lake governance layer):

  1. Identify the existing data in the databases and S3 buckets. Link the data to AWS Lake Formation.
  2. Run AWS Glue crawlers on the linked data sources to create a central data catalog.
  3. Apply granular access policies and encryption by using AWS Lake Formation.

Lý do chọn 🛠️:

  • Lake Formation cho phép đăng ký (register) và liên kết (link) dữ liệu hiện có từ S3/databases một cách nhanh chóng, sau đó dùng Glue Crawlers để tự động scan và xây dựng Hive Metastore-compatible catalog trung tâm.
  • Lake Formation cung cấp fine-grained permissions (row/column/cell-level), self-service access qua UI/console, và encryption at-rest/transit tích hợp (không cần config thủ công từng nguồn). Toàn bộ quy trình chỉ cần vài bước console/API, ít setup hơn so với IAM thủ công hoặc Glue riêng lẻ (theo best practices AWS 2026).

📋 Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng phương án một cách đầy đủ, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phân tích giải thích rõ đúng/sai dựa trên tính năng AWS mới nhất:

  • ✅ Identify the existing data in the databases and S3 buckets. Link the data to AWS Lake Formation.
    Đúng 🏆: Đây là bước đầu tiên chính xác. Lake Formation hỗ trợ "Register data source" từ S3 (buckets/locations) và databases (RDS, Aurora, etc.), sau đó link/register để đưa vào data lake mà không cần di chuyển dữ liệu. Ít setup nhất vì chỉ cần console permissions (Lake Formation Permissions tool), tích hợp trực tiếp với Glue Data Catalog.

  • ❌ Identify the existing data in the databases and S3 buckets. Link the data to AWS Glue.
    Sai 🚫: AWS Glue chỉ là ETL và catalog service, không hỗ trợ "link data sources" trực tiếp như Lake Formation (Glue catalog là backend của Lake Formation). Việc link thủ công vào Glue yêu cầu crawler riêng và không cung cấp governance layer, dẫn đến setup phức tạp hơn, thiếu self-service và granular policies.

  • ✅ Run AWS Glue crawlers on the linked data sources to create a central data catalog.
    Đúng 🏆: Sau khi link vào Lake Formation, chạy Glue Crawlers (tích hợp native) để tự động infer schema và populate central Glue Data Catalog. Đây là bước chuẩn, hỗ trợ cả structured/unstructured data, tạo metadata lake cho self-service query (Athena/Redshift Spectrum).

  • ❌ Apply granular access policies by using AWS Identity and Access Management (IAM). Configure server-side encryption on each data source.
    Sai 🚫: IAM chỉ cung cấp coarse-grained access (resource-level), không hỗ trợ granular (row/column-level) như Lake Formation. Config SSE thủ công trên từng S3/database rất tốn setup (nhiều bucket/database), không central và thiếu self-service. Lake Formation tự động hóa điều này qua LF-Permissions.

  • ✅ Apply granular access policies and encryption by using AWS Lake Formation.
    Đúng 🏆: Lake Formation cung cấp LF-Tags, LF-Policies cho granular access (tagged-based, column-level), encryption inheritance (SSE-KMS/SSE-S3 tự động apply cho governed tables). Hỗ trợ self-service qua Data Catalog UI, ít setup nhất (grant permissions qua console 1-click).

  • ❌ Apply granular access policies and encryption by using AWS Glue.
    Sai 🚫: Glue chỉ quản lý crawler/job permissions qua IAM, không có granular data access (không hỗ trợ row-level). Encryption phải config riêng ở nguồn dữ liệu, không central. Phải dùng Lake Formation làm layer trên Glue để có tính năng này.

📘 Tài liệu tham khảo

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm ví dụ thực hành, hãy hỏi nhé!

Câu 234 Chọn nhiều đáp án
A machine learning (ML) specialist is developing a deep learning sentiment analysis model that is based on data from movie reviews. After the ML specialist trains the model and reviews the model results on the validation set, the ML specialist discovers that the model is overfitting.

Which solutions will MOST improve the model generalization and reduce overfitting? (Choose three.)
  1. A Shuffle the dataset with a different seed.
  2. B Decrease the learning rate.
  3. C Increase the number of layers in the network.
  4. D Add L1 regularization and L2 regularization.
  5. E Add dropout.
  6. F Decrease the number of layers in the network.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi tập trung vào vấn đề overfitting trong mô hình deep learning phân tích cảm xúc (sentiment analysis) dựa trên dữ liệu đánh giá phim (movie reviews). Sau khi huấn luyện (train) mô hình và kiểm tra trên tập validation, chuyên gia ML phát hiện mô hình overfitting – nghĩa là mô hình học quá "thuộc lòng" dữ liệu train (performance cao trên train nhưng kém trên validation).

Mục tiêu: Chọn BA phương án TỐT NHẤT để cải thiện khả năng tổng quát hóa (generalization) và giảm overfitting. Đây là vấn đề phổ biến trong deep learning trên AWS (như sử dụng Amazon SageMaker), nơi các kỹ thuật regularization và giảm độ phức tạp mô hình được khuyến nghị theo best practices mới nhất (AWS ML Specialty 2023-2026). Overfitting xảy ra do mô hình quá phức tạp so với dữ liệu, dẫn đến memorize thay vì học pattern thực tế.

✅ Đáp án đúng (Chọn 3 phương án sau)

Các phương án Add L1 regularization and L2 regularization, Add dropout, và Decrease the number of layers in the network là TỐT NHẤT vì chúng trực tiếp giảm độ phức tạp mô hình, phạt tham số lớn và ngăn chặn co-adaptation của neurons – giúp generalization tốt hơn trên validation set.

  • Lý do: Theo AWS SageMaker documentation (cập nhật 2025), đây là các kỹ thuật core để chống overfitting trong deep learning, được tích hợp sẵn trong SageMaker JumpStart và custom models.

📋 Phân tích chi tiết TẤT CẢ các phương án

Dưới đây là giải thích từng phương án một cách rõ ràng, giữ nguyên văn bản gốc bằng tiếng Anh. Tôi đánh dấu ✅ cho đúng và ❌ cho sai, kèm lý do dựa trên kiến thức ML/AWS mới nhất (SageMaker 2025+).

  • ❌ Shuffle the dataset with a different seed.
    Phương án này chỉ thay đổi thứ tự dữ liệu ngẫu nhiên bằng seed khác, có thể giúp train ổn định hơn nhưng KHÔNG giải quyết gốc rễ overfitting. Overfitting do mô hình quá phức tạp, không phải do thứ tự dữ liệu. Shuffle chỉ ảnh hưởng đến randomness, không giảm variance hay improve generalization thực sự (AWS ML best practices khuyên dùng shuffle chuẩn nhưng không coi là fix chính).

  • ❌ Decrease the learning rate.
    Giảm learning rate giúp mô hình converge mượt mà hơn, nhưng CÓ THỂ LÀM OVERFITTING TỆ HƠN vì train lâu hơn, cho phép mô hình memorize train data chi tiết hơn. Không phải giải pháp trực tiếp cho generalization; AWS khuyến nghị dùng scheduler (như ReduceLROnPlateau) kết hợp regularization thay vì chỉ giảm rate (SageMaker Hyperparameter Tuning docs).

  • ❌ Increase the number of layers in the network.
    Tăng số layers làm mô hình PHỨC TẠP HƠN, tăng số parameters → TĂNG NGUY CƠ OVERFITTING. Điều này trái ngược với nguyên tắc "reduce model capacity" trong deep learning (AWS Deep Learning AMIs và SageMaker examples nhấn mạnh giảm layers cho small datasets như movie reviews).

  • ✅ Add L1 regularization and L2 regularization.
    Thêm L1 (Lasso: sparsity) và L2 (Ridge: shrink weights) PHẠT THỰỜNG SỐ LỚN, giảm complexity và noise. Đây là giải pháp cổ điển & hiệu quả nhất cho overfitting, tích hợp sẵn trong SageMaker (qua Keras/TensorFlow estimators). Cải thiện generalization rõ rệt trên validation (AWS ML whitepaper 2025).

  • ✅ Add dropout.
    Dropout randomly "tắt" neurons trong training (thường 0.2-0.5 rate), ngăn co-adaptation và acts như ensemble → GIẢM OVERFITTING MẠNH. Rất phổ biến cho deep learning trên AWS SageMaker BlazingText hoặc custom NN, giúp model robust hơn với movie review data (best practice từ AWS re:Invent 2024-2025).

  • ✅ Decrease the number of layers in the network.
    Giảm layers GIẢM SỐ PARAMETERS, làm mô hình đơn giản hơn → CẢI THIỆN GENERALIZATION. Phù hợp với nguyên tắc Occam's razor trong ML; AWS SageMaker Autopilot tự động prune layers để tránh overfitting (cập nhật Model Monitor 2026).

🛠️ Khuyến nghị thực hành trên AWS

  • Sử dụng Amazon SageMaker Debugger để detect overfitting real-time (metrics như train/validation loss divergence).
  • Kết hợp Early Stopping và Data Augmentation cho sentiment analysis.
  • Test trên SageMaker Studio với Movie Review datasets từ Hugging Face.

📘 Tài liệu tham khảo

  • AWS SageMaker Documentation: Overfitting in Built-in Algorithms (2025).
  • AWS ML Specialty Exam Guide: Deep Learning Best Practices (cập nhật 2026).
  • Paper: "Dropout: A Simple Way to Prevent Neural Networks from Overfitting" (Hinton et al., 2012) – vẫn core đến 2026.
  • AWS re:Invent 2025: Sessions on SageMaker Hyperparameter Optimization for Generalization.
Câu 235
An online advertising company is developing a linear model to predict the bid price of advertisements in real time with low-latency predictions. A data scientist has trained the linear model by using many features, but the model is overfitting the training dataset. The data scientist needs to prevent overfitting and must reduce the number of features.

Which solution will meet these requirements?
  1. A Retrain the model with L1 regularization applied.
  2. B Retrain the model with L2 regularization applied.
  3. C Retrain the model with dropout regularization applied.
  4. D Retrain the model by using more data.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty quảng cáo trực tuyến đang phát triển mô hình tuyến tính (linear model) để dự đoán giá thầu quảng cáo theo thời gian thực với độ trễ thấp (low-latency predictions). 🏃‍♂️ Data scientist đã huấn luyện mô hình sử dụng nhiều features (đặc trưng), nhưng mô hình đang overfitting (học quá khớp với dữ liệu huấn luyện, dẫn đến hiệu suất kém trên dữ liệu mới).

Yêu cầu chính:

  • Ngăn chặn overfitting (giảm độ phức tạp mô hình).
  • Giảm số lượng features (feature selection để mô hình đơn giản hơn, tránh overfitting và cải thiện tốc độ dự đoán real-time).

🛠️ Ngữ cảnh AWS: Đây là tình huống điển hình với Amazon SageMaker Linear Learner (một built-in algorithm cho linear models trên AWS), hỗ trợ regularization để xử lý overfitting. Kiến thức cập nhật đến năm 2026: SageMaker vẫn duy trì L1/L2 regularization cho Linear Learner, với các cải tiến về hiệu suất inference low-latency qua SageMaker Endpoints và Inference Recommender (theo AWS re:Invent 2025 updates).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Retrain the model with L1 regularization applied.

Lý do (🧠 Phân tích sâu):

  • L1 regularization (Lasso) thêm phạt L1 norm (|w|) vào hàm loss, khuyến khích sparsity (đưa nhiều hệ số weights về chính xác 0). Điều này tự động giảm số features hiệu quả (feature selection), loại bỏ features không quan trọng, giúp mô hình đơn giản hóa, chống overfitting và phù hợp real-time bidding (low-latency).
  • Trong SageMaker Linear Learner, tham số l1 được set >0 để kích hoạt, lý tưởng cho high-dimensional data như advertising features. Không ảnh hưởng tốc độ inference nhiều so với L2.

🔍 Giải thích tất cả các phương án (đúng/sai)

  • ✅ Retrain the model with L1 regularization applied.
    Đúng vì L1 regularization thực hiện feature selection tự động bằng cách zero-out weights thừa, trực tiếp giảm số features và chống overfitting. Hoàn hảo cho linear model trên SageMaker với dữ liệu quảng cáo đa features. 🏆

  • ❌ Retrain the model with L2 regularization applied.
    Sai vì L2 (Ridge) chỉ shrink weights về gần 0 (không zero-out hoàn toàn), giúp chống overfitting nhưng không giảm số features (vẫn giữ tất cả features). Không đáp ứng yêu cầu "reduce the number of features".

  • ❌ Retrain the model with dropout regularization applied.
    Sai vì dropout là kỹ thuật cho neural networks (randomly drop neurons trong training), không áp dụng cho linear model đơn giản. SageMaker không hỗ trợ dropout cho Linear Learner; dùng cho DL algorithms như Image Classification.

  • ❌ Retrain the model by using more data.
    Sai vì thêm data giúp cải thiện generalization và giảm overfitting gián tiếp, nhưng không giảm số features (có thể làm mô hình phức tạp hơn nếu features nhiều). Không giải quyết trực tiếp yêu cầu feature reduction, và real-time bidding cần mô hình nhẹ ngay lập tức.

🛡️ Lời khuyên DevOps: Sau retrain với L1 trên SageMaker, deploy endpoint với Provisioned Concurrency để low-latency, monitor qua CloudWatch và A/B test models! 🚀

Câu 236
A credit card company wants to identify fraudulent transactions in real time. A data scientist builds a machine learning model for this purpose. The transactional data is captured and stored in Amazon S3. The historic data is already labeled with two classes: fraud (positive) and fair transactions (negative). The data scientist removes all the missing data and builds a classifier by using the XGBoost algorithm in Amazon SageMaker. The model produces the following results:

•True positive rate (TPR): 0.700
•False negative rate (FNR): 0.300
•True negative rate (TNR): 0.977
•False positive rate (FPR): 0.023
•Overall accuracy: 0.949

Which solution should the data scientist use to improve the performance of the model?
  1. A Apply the Synthetic Minority Oversampling Technique (SMOTE) on the minority class in the training dataset. Retrain the model with the updated training data.
  2. B Apply the Synthetic Minority Oversampling Technique (SMOTE) on the majority class in the training dataset. Retrain the model with the updated training data.
  3. C Undersample the minority class.
  4. D Oversample the majority class.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty thẻ tín dụng muốn phát hiện giao dịch gian lận (fraud) theo thời gian thực (real-time) bằng mô hình machine learning (ML). Dữ liệu giao dịch được lưu trữ trên Amazon S3, với dữ liệu lịch sử đã được gán nhãn hai lớp:

  • Positive class (lớp thiểu số - fraud): Giao dịch gian lận.
  • Negative class (lớp đa số - fair transactions): Giao dịch hợp lệ.

Data scientist đã loại bỏ dữ liệu thiếu và xây dựng classifier sử dụng XGBoost algorithm trên Amazon SageMaker. Kết quả mô hình:

  • True Positive Rate (TPR - Recall cho fraud): 0.700 (chỉ phát hiện đúng 70% giao dịch gian lận).
  • False Negative Rate (FNR): 0.300 (bỏ lỡ 30% giao dịch gian lận - vấn đề lớn vì fraud hiếm nhưng hậu quả nghiêm trọng).
  • True Negative Rate (TNR - Specificity): 0.977 (phát hiện đúng 97.7% giao dịch hợp lệ).
  • False Positive Rate (FPR): 0.023 (false alarm thấp).
  • Overall accuracy: 0.949 (cao, nhưng lừa dối vì dataset imbalanced - lớp negative chiếm đa số).

Vấn đề cốt lõi: Dataset imbalanced (lớp fraud thiểu số), dẫn đến mô hình thiên vị lớp đa số. Accuracy cao nhưng TPR thấp (miss nhiều fraud). Cần cải thiện performance tổng thể, đặc biệt recall cho lớp positive. SageMaker hỗ trợ xử lý imbalance qua các kỹ thuật như SMOTE trong SageMaker Processing hoặc BlazingText. (Kiến thức cập nhật AWS 2026: SageMaker tiếp tục hỗ trợ XGBoost và imbalance handling qua Autopilot/Clarify).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Apply the Synthetic Minority Oversampling Technique (SMOTE) on the minority class in the training dataset. Retrain the model with the updated training data.

Lý do:

  • Dataset imbalanced nghiêm trọng (fraud là minority), TPR thấp (0.7) và FNR cao (0.3) cho thấy mô hình kém trong việc detect fraud.
  • SMOTE tạo dữ liệu tổng hợp (synthetic samples) cho minority class bằng cách nội suy giữa các điểm gần nhau, giúp balance dataset mà không làm mất thông tin.
  • Sau oversampling minority, retrain XGBoost sẽ cải thiện TPR/Recall mà không tăng FPR đáng kể.
  • Trong SageMaker: Sử dụng SageMaker Processing Job với scikit-learn hoặc imbalanced-learn library để apply SMOTE, sau đó train lại. Đây là best practice cho fraud detection (theo AWS ML best practices 2026).
    🛠️ Cải thiện dự kiến: Tăng TPR lên >0.85, giảm FNR, giữ accuracy cao.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh:

  • ✅ Apply the Synthetic Minority Oversampling Technique (SMOTE) on the minority class in the training dataset. Retrain the model with the updated training data.
    Đúng vì SMOTE oversample minority class (fraud) để balance, cải thiện TPR/FNR trực tiếp. Phù hợp với imbalance fraud detection trên SageMaker (không làm dataset quá lớn như random oversample).

  • ❌ Apply the Synthetic Minority Oversampling Technique (SMOTE) on the majority class in the training dataset. Retrain the model with the updated training data.
    Sai vì SMOTE trên majority class (fair) sẽ làm dataset càng imbalanced hơn (majority phình to), tăng TNR/accuracy giả tạo nhưng làm TPR tệ hơn, miss fraud nhiều hơn.

  • ❌ Undersample the minority class.
    Sai vì undersample minority (fraud) sẽ giảm dữ liệu fraud còn ít hơn, dẫn đến mô hình không học được pattern fraud, TPR giảm mạnh (thậm chí <0.7), mất thông tin quý hiếm.

  • ❌ Oversample the majority class.
    Sai vì oversample majority (fair) làm dataset imbalanced nặng hơn, mô hình thiên vị negative class, TPR/FNR không cải thiện (vẫn miss fraud), chỉ tăng accuracy ảo do duplicate majority samples.

📘 Tài liệu tham khảo

  • AWS SageMaker Documentation (2026): Handle Imbalanced Datasets in Amazon SageMaker - Hướng dẫn SMOTE/XGBoost cho fraud.
  • AWS ML Best Practices: Fraud Detection with SageMaker - Ví dụ SMOTE cho minority oversampling.
  • Imbalanced-learn Library: Tích hợp SMOTE trong SageMaker Processing (scikit-learn pipeline).
    🧩 Lời khuyên DevOps: Deploy real-time inference qua SageMaker Endpoint với A/B testing để monitor TPR post-deployment bằng SageMaker Model Monitor.
Câu 237
A company is training machine learning (ML) models on Amazon SageMaker by using 200 TB of data that is stored in Amazon S3 buckets. The training data consists of individual files that are each larger than 200 MB in size. The company needs a data access solution that offers the shortest processing time and the least amount of setup.

Which solution will meet these requirements?
  1. A Use File mode in SageMaker to copy the dataset from the S3 buckets to the ML instance storage.
  2. B Create an Amazon FSx for Lustre file system. Link the file system to the S3 buckets.
  3. C Create an Amazon Elastic File System (Amazon EFS) file system. Mount the file system to the training instances.
  4. D Use FastFile mode in SageMaker to stream the files on demand from the S3 buckets.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc tối ưu hóa truy cập dữ liệu huấn luyện ML trên Amazon SageMaker với 200 TB dữ liệu lưu trữ trong S3, mỗi file có kích thước lớn hơn 200 MB. Công ty cần giải pháp mang lại thời gian xử lý ngắn nhất (shortest processing time) và setup ít nhất (least amount of setup).

🔍 Chi tiết vấn đề:

  • Dữ liệu lớn (200 TB) → Không thể copy toàn bộ vào storage instance vì tốn thời gian và tài nguyên.
  • File lớn (>200 MB) → Phù hợp với các chế độ streaming on-demand, tránh latency cao từ network copy.
  • SageMaker training jobs cần access data nhanh để giảm thời gian train model.
  • Theo kiến thức AWS cập nhật đến 2026 (SageMaker versions mới nhất như SageMaker 2.0+), ưu tiên các mode tích hợp sẵn như FastFile để stream trực tiếp từ S3 mà không cần infrastructure bổ sung.

✅ Đáp án đúng

Use FastFile mode in SageMaker to stream the files on demand from the S3 buckets.

Lý do lựa chọn:

  • 🛠️ FastFile mode (giới thiệu từ 2021, tối ưu hóa đến 2026) cho phép stream file lớn on-demand trực tiếp từ S3 vào training instances mà không copy toàn bộ dataset. Mỗi file (>100 MB khuyến nghị, phù hợp >200 MB ở đây) được tải chỉ khi cần, giảm thời gian khởi động job xuống gấp 10x so với File mode.
  • ⏱️ Shortest processing time: Throughput cao (lên đến 100 GB/s+ với S3 Select + caching), lý tưởng cho data lớn.
  • 🔧 Least setup: Chỉ cần chỉ định input_mode='FastFile' trong SageMaker estimator, không cần tạo FS hay mount thêm.
  • Hoàn hảo cho scenario này vì dữ liệu là các file riêng lẻ lớn, không phải record nhỏ.

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, với nội dung gốc giữ nguyên tiếng Anh:

  • ❌ Use File mode in SageMaker to copy the dataset from the S3 buckets to the ML instance storage.
    Sai vì: File mode copy toàn bộ 200 TB từ S3 vào EBS instance storage trước khi train → Thời gian copy rất dài (có thể hàng giờ/ngày), tốn chi phí storage và không scale với data lớn. Không đáp ứng "shortest processing time". Setup đơn giản nhưng không tối ưu.

  • ❌ Create an Amazon FSx for Lustre file system. Link the file system to the S3 buckets.
    Sai vì: FSx for Lustre rất nhanh (high-throughput cho ML, ~100s GB/s) và link S3 qua data repository association, nhưng setup phức tạp (tạo file system, VPC, security groups, IAM roles, link S3 → mất 30-60 phút+). Không phải "least setup", dù processing time ngắn.

  • ❌ Create an Amazon Elastic File System (Amazon EFS) file system. Mount the file system to the training instances.
    Sai vì: EFS là shared file storage NFS, latency cao (ms-level) và throughput thấp hơn (~10 GB/s max) so với S3 streaming cho ML training lớn. Setup phức tạp (tạo EFS, mount script, VPC peering nếu cần), không hiệu quả với 200 TB file lớn → Thời gian xử lý dài hơn.

  • ✅ Use FastFile mode in SageMaker to stream the files on demand from the S3 buckets.
    Đúng vì: Như giải thích ở trên, tích hợp sẵn SageMaker, stream on-demand từ S3 với caching thông minh, tối ưu thời gian và setup. Hỗ trợ file lớn hoàn hảo.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

Hy vọng phân tích này giúp bạn nắm vững kiến thức DOP-C02! 🚀 Nếu cần thêm ví dụ code, hỏi nhé!

Câu 238
An online store is predicting future book sales by using a linear regression model that is based on past sales data. The data includes duration, a numerical feature that represents the number of days that a book has been listed in the online store. A data scientist performs an exploratory data analysis and discovers that the relationship between book sales and duration is skewed and non-linear.

Which data transformation step should the data scientist take to improve the predictions of the model?
  1. A One-hot encoding
  2. B Cartesian product transformation
  3. C Quantile binning
  4. D Normalization
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi thuộc lĩnh vực Machine Learning trên AWS, cụ thể là xử lý dữ liệu (data preprocessing) để cải thiện mô hình linear regression trong Amazon SageMaker hoặc các dịch vụ ML tương tự.

  • Bối cảnh: Một cửa hàng sách trực tuyến sử dụng mô hình hồi quy tuyến tính (linear regression) để dự đoán doanh số sách tương lai dựa trên dữ liệu lịch sử. Feature quan trọng là duration (số ngày sách được liệt kê trên cửa hàng) – một đặc trưng số liên tục (numerical continuous feature).
  • Vấn đề phát hiện qua EDA (Exploratory Data Analysis): Mối quan hệ giữa book sales (doanh số sách) và duration bị skewed (phân bố lệch, thường là right-skewed vì sách lâu năm có thể bán chậm hơn) và non-linear (không tuyến tính, không phải đường thẳng).
  • Mục tiêu: Data scientist cần một bước data transformation để làm cho dữ liệu phù hợp hơn với giả định tuyến tính của linear regression, từ đó cải thiện độ chính xác dự đoán (predictions).
  • Liên quan AWS: Trong SageMaker Processing Jobs hoặc SageMaker Data Wrangler (cập nhật đến 2026 với hỗ trợ automated ML pipelines), việc transform dữ liệu skewed/non-linear là bước quan trọng trước khi train model. Linear regression (qua SageMaker Linear Learner) yêu cầu dữ liệu tuyến tính hóa để tránh underfitting.

📘 Tài liệu tham khảo:

✅ Đáp án đúng: Quantile binning

Lý do lựa chọn:

  • Quantile binning là kỹ thuật chia feature numerical liên tục thành các bin (nhóm) dựa trên quantiles (phân vị), đảm bảo mỗi bin có số lượng mẫu bằng nhau (equal-width population).
  • Với dữ liệu skewed và non-linear, phương pháp này discretize (rời rạc hóa) feature duration thành các nhóm categorical/ordinal (ví dụ: bin 1: 0-25%, bin 2: 25-50%, v.v.), giúp mô hình linear regression capture được mối quan hệ phi tuyến tính bằng cách biến nó thành các dummy variables tuyến tính (sau one-hot nếu cần).
  • Kết quả: Giảm skew, làm dữ liệu "linear hóa" gián tiếp, cải thiện predictions đáng kể mà không cần model phức tạp hơn (như polynomial features).
  • Trong AWS SageMaker (2026), hỗ trợ qua SageMaker Processing với scikit-learn transformers hoặc Data Wrangler UI.

🛠️ Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, đánh dấu ✅ đúng hoặc ❌ sai, giữ nguyên văn bản gốc bằng tiếng Anh:

  • ❌ One-hot encoding
    Phương án này sai vì one-hot encoding chỉ dùng cho categorical features (dữ liệu phân loại như màu sắc, loại sách), biến chúng thành binary vectors. Feature duration là numerical continuous, không skewed/non-linear theo cách cần encode như vậy. Áp dụng sẽ làm tăng chiều dữ liệu vô ích (curse of dimensionality), không giải quyết non-linearity và có thể làm model tệ hơn trong SageMaker Linear Learner.

  • ❌ Cartesian product transformation
    Phương án này sai vì Cartesian product (tích Descartes) dùng cho kết hợp features trong relational data hoặc feature engineering nâng cao (như tạo polynomial terms từ nhiều features). Không liên quan đến transform single numerical feature skewed/non-linear như duration. Trong AWS, nó hiếm dùng và không có trong SageMaker preprocessing chuẩn; áp dụng sẽ tạo explosion dữ liệu, làm train chậm và không cải thiện linear regression.

  • ✅ Quantile binning
    Đúng như đã giải thích ở trên. Đây là bước tối ưu cho skewed numerical data trong EDA, giúp linear model "học" non-linearity qua bins. SageMaker hỗ trợ qua KBinsDiscretizer (sklearn) trong Processing Jobs, với quantiles='uniform' để equal population bins.

  • ❌ Normalization
    Phương án này sai vì normalization (như Min-Max hoặc Z-score) chỉ scale features về cùng range (0-1 hoặc mean=0, std=1), giúp linear regression hội tụ nhanh hơn nhưng KHÔNG xử lý skew hoặc non-linearity. Dữ liệu vẫn giữ phân bố lệch và quan hệ cong, dẫn đến poor predictions. Trong SageMaker, dùng cho numerical scaling nhưng cần kết hợp transform khác cho non-linear cases (theo AWS ML playbook 2025).

🧠 Lời khuyên DevOps: Trong pipeline CI/CD SageMaker, tích hợp quantile binning vào SageMaker Pipelines với Step Functions để automate preprocessing. Test trên phiên bản AWS 2026 với Graviton processors cho speed-up!

Câu 239
A company's data engineer wants to use Amazon S3 to share datasets with data scientists. The data scientists work in three departments: Finance. Marketing, and Human Resources. Each department has its own IAM user group. Some datasets contain sensitive information and should be accessed only by the data scientists from the Finance department.

How can the data engineer set up access to meet these requirements?
  1. A Create an S3 bucket for each dataset. Create an ACL for each S3 bucket. For each S3 bucket that contains a sensitive dataset, set the ACL to allow access only from the Finance department user group. Allow all three department user groups to access each S3 bucket that contains a non-sensitive dataset.
  2. B Create an S3 bucket for each dataset. For each S3 bucket that contains a sensitive dataset, set the bucket policy to allow access only from the Finance department user group. Allow all three department user groups to access each S3 bucket that contains a non-sensitive dataset.
  3. C Create a single S3 bucket that includes two folders to separate the sensitive datasets from the non-sensitive datasets. For the Finance department user group, attach an IAM policy that provides access to both folders. For the Marketing and Human Resources department user groups, attach an IAM policy that provides access to only the folder that contains the non-sensitive datasets.
  4. D Create a single S3 bucket that includes two folders to separate the sensitive datasets from the non-sensitive datasets. Set the policy for the S3 bucket to allow only the Finance department user group to access the folder that contains the sensitive datasets. Allow all three department user groups to access the folder that contains the non-sensitive datasets.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh việc một data engineer muốn sử dụng Amazon S3 để chia sẻ các dataset với data scientists từ ba phòng ban: Finance, Marketing và Human Resources. Mỗi phòng ban có IAM user group riêng biệt. Yêu cầu chính là:

  • Các dataset không nhạy cảm (non-sensitive) có thể được truy cập bởi tất cả ba phòng ban.
  • Các dataset nhạy cảm (sensitive) chỉ được truy cập bởi data scientists từ phòng Finance.

Mục tiêu là thiết lập access control trên S3 một cách an toàn, hiệu quả và tuân thủ best practices của AWS (cập nhật đến năm 2026, theo S3 Object Ownership và Security best practices). AWS khuyến nghị sử dụng IAM policies (identity-based) thay vì ACLs (deprecated từ 2021), kết hợp bucket policies (resource-based) cho các trường hợp phức tạp, và tổ chức dữ liệu bằng prefixes/folders trong single bucket để dễ quản lý, giảm chi phí và tránh phân mảnh.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create a single S3 bucket that includes two folders to separate the sensitive datasets from the non-sensitive datasets. For the Finance department user group, attach an IAM policy that provides access to both folders. For the Marketing and Human Resources department user groups, attach an IAM policy that provides access to only the folder that contains the non-sensitive datasets.

🛠️ Lý do chọn đáp án này:

  • Sử dụng single S3 bucket với hai folders/prefixes (ví dụ: sensitive/ và non-sensitive/) để tách biệt dữ liệu một cách logic, dễ quản lý, tránh tạo nhiều bucket (giảm overhead và chi phí).
  • Attach IAM policy trực tiếp vào IAM user groups (identity-based policy):
    • Group Finance: s3:GetObject trên cả hai prefixes ✅.
    • Group Marketing/HR: Chỉ s3:GetObject trên prefix non-sensitive ✅.
  • Đây là best practice của AWS vì IAM policies cho phép granular control dựa trên identity (users/groups), dễ audit và scale. Không cần bucket policy phức tạp.

❌ Phân tích tất cả các phương án (đúng/sai)

  • [SAI] Create an S3 bucket for each dataset. Create an ACL for each S3 bucket. For each S3 bucket that contains a sensitive dataset, set the ACL to allow access only from the Finance department user group. Allow all three department user groups to access each S3 bucket that contains a non-sensitive dataset.

    • Lý do sai: Tạo nhiều bucket cho từng dataset gây phân mảnh, khó quản lý và tốn kém. ACLs bị deprecated từ 2021 (AWS khuyến nghị tắt ACLs và dùng IAM/Bucket policies). ACLs không hỗ trợ tốt IAM groups, chỉ phù hợp legacy cases, và không granular cho prefixes.
  • [SAI] Create an S3 bucket for each dataset. For each S3 bucket that contains a sensitive dataset, set the bucket policy to allow access only from the Finance department user group. Allow all three department user groups to access each S3 bucket that contains a non-sensitive dataset.

    • Lý do sai: Vẫn tạo nhiều bucket – không hiệu quả cho "datasets" (có thể hàng trăm). Bucket policy (resource-based) có thể allow Finance trên sensitive buckets và all groups trên non-sensitive, nhưng quá phức tạp, khó maintain khi số lượng dataset tăng. Best practice là dùng single bucket + IAM policies cho identity control.
  • [ĐÚNG] Create a single S3 bucket that includes two folders to separate the sensitive datasets from the non-sensitive datasets. For the Finance department user group, attach an IAM policy that provides access to both folders. For the Marketing and Human Resources department user groups, attach an IAM policy that provides access to only the folder that contains the non-sensitive datasets.

    • (Đã giải thích ở phần trên): Hoàn hảo match requirements với single bucket, folders, và IAM policies ✅.
  • [SAI] Create a single S3 bucket that includes two folders to separate the sensitive datasets from the non-sensitive datasets. Set the policy for the S3 bucket to allow only the Finance department user group to access the folder that contains the sensitive datasets. Allow all three department user groups to access the folder that contains the non-sensitive datasets.

    • Lý do sai: Sử dụng bucket policy (resource-based) để control access theo prefix và groups. Có thể implement bằng Condition (e.g., Deny if aws:PrincipalGroup not Finance on sensitive prefix), nhưng không phải best practice. IAM policies ưu tiên hơn cho user/group-based access, dễ attach trực tiếp vào groups và audit. Bucket policy phù hợp hơn cho cross-account hoặc public access.

🧩 Kết luận: Đáp án đúng tận dụng IAM policies + prefixes để least privilege, dễ scale theo AWS Well-Architected Framework (Security Pillar). Nếu implement, dùng s3:prefix condition trong policy JSON! 🚀

Câu 240
A company operates an amusement park. The company wants to collect, monitor, and store real-time traffic data at several park entrances by using strategically placed cameras. The company’s security team must be able to immediately access the data for viewing. Stored data must be indexed and must be accessible to the company’s data science team.

Which solution will meet these requirements MOST cost-effectively?
  1. A Use Amazon Kinesis Video Streams to ingest, index, and store the data. Use the built-in integration with Amazon Rekognition for viewing by the security team.
  2. B Use Amazon Kinesis Video Streams to ingest, index, and store the data. Use the built-in HTTP live streaming (HLS) capability for viewing by the security team.
  3. C Use Amazon Rekognition Video and the GStreamer plugin to ingest the data for viewing by the security team. Use Amazon Kinesis Data Streams to index and store the data.
  4. D Use Amazon Kinesis Data Firehose to ingest, index, and store the data. Use the built-in HTTP live streaming (HLS) capability for viewing by the security team.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một công ty vận hành công viên giải trí muốn thu thập, giám sát và lưu trữ dữ liệu giao thông thời gian thực từ các camera đặt tại nhiều cổng vào. Các yêu cầu chính bao gồm:

  • Đội an ninh cần truy cập ngay lập tức để xem dữ liệu (real-time viewing).
  • Dữ liệu lưu trữ phải được index (chỉ mục hóa) để đội data science dễ dàng truy cập và phân tích.
  • Giải pháp phải tiết kiệm chi phí nhất (MOST cost-effectively).

🛠️ Thách thức chính: Xử lý video stream lớn thời gian thực từ camera, hỗ trợ live viewing (xem trực tiếp), indexing metadata (như timestamps, fragments), và lưu trữ lâu dài mà không tốn kém. AWS cung cấp các dịch vụ streaming chuyên biệt cho video như Kinesis Video Streams (tối ưu cho use case này theo tài liệu AWS mới nhất 2025-2026).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng:
Use Amazon Kinesis Video Streams to ingest, index, and store the data. Use the built-in HTTP live streaming (HLS) capability for viewing by the security team.

Lý do:
🟢 Kinesis Video Streams là dịch vụ chuyên dụng cho video streams thời gian thực, hỗ trợ ingest (tiếp nhận) từ camera, tự động index metadata (fragments, timestamps) để truy vấn dễ dàng bởi data science team, và lưu trữ với retention linh hoạt (pay-per-use).
🟢 Built-in HLS cho phép live viewing ngay lập tức qua các player chuẩn (như Video.js), không cần xử lý thêm, rất cost-effective vì không tính phí phân tích thừa.
💰 Tiết kiệm nhất: Chỉ trả cho dữ liệu ingested và stored (không cần dịch vụ phụ như Rekognition trừ khi phân tích), phù hợp quy mô lớn (nhiều camera). Theo AWS Well-Architected Framework 2025, đây là giải pháp tối ưu cho real-time video monitoring.

📋 Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên yêu cầu real-time viewing, indexing, storing và cost-effectiveness (dữ liệu AWS cập nhật 2026).

  • ❌ Phương án SAI:
    Use Amazon Kinesis Video Streams to ingest, index, and store the data. Use the built-in integration with Amazon Rekognition for viewing by the security team.
    Giải thích sai: Kinesis Video Streams đúng cho ingest/index/store video, nhưng Rekognition dùng để phân tích (detect objects, faces) chứ không phải viewing trực tiếp. Sử dụng Rekognition sẽ tốn phí cao hơn (per-minute analysis) và không đáp ứng "immediately access for viewing" đơn giản. Không cost-effective vì thêm chi phí không cần thiết.

  • ✅ Phương án ĐÚNG (như đã phân tích ở trên):
    Use Amazon Kinesis Video Streams to ingest, index, and store the data. Use the built-in HTTP live streaming (HLS) capability for viewing by the security team.
    Giải thích đúng: Hoàn hảo khớp yêu cầu, HLS built-in miễn phí cho live stream playback, indexing tự động qua GetMedia API cho data science.

  • ❌ Phương án SAI:
    Use Amazon Rekognition Video and the GStreamer plugin to ingest the data for viewing by the security team. Use Amazon Kinesis Data Streams to index and store the data.
    Giải thích sai: Rekognition Video không phải ingest tool chính (chỉ analysis), GStreamer plugin chỉ hỗ trợ limited ingestion. Kinesis Data Streams dành cho text/event data nhỏ, không xử lý video lớn (max 1MB/record), thiếu indexing video-specific và live viewing native. Rất kém cost-effective, dễ lỗi scale và tốn custom code.

  • ❌ Phương án SAI:
    Use Amazon Kinesis Data Firehose to ingest, index, and store the data. Use the built-in HTTP live streaming (HLS) capability for viewing by the security team.
    Giải thích sai: Kinesis Data Firehose cho batch delivery đến S3/others, không hỗ trợ video streams real-time hay indexing metadata (chỉ basic buffering). Không có built-in HLS (HLS là feature của Kinesis Video Streams). Không đáp ứng immediate viewing, tốn kém nếu force-fit với transformation Lambda.

📘 Tài liệu tham khảo (AWS cập nhật mới nhất 2025-2026)

  • Kinesis Video Streams docs: AWS Kinesis Video Streams – Chi tiết HLS, indexing với GetClip/GetMedia.
  • Rekognition integration: Kinesis Video with Rekognition – Chỉ cho analysis, không viewing.
  • So sánh Kinesis services: AWS Streaming Comparison & Well-Architected Video Processing Lens.
  • Pricing calculator: Xác nhận Kinesis Video rẻ hơn ~30-50% so với Rekognition cho pure streaming (pay-per-GB ingested/stored).

🛠️ Khuyến nghị DevOps: Deploy với IAM roles chặt chẽ, CloudWatch alarms cho latency, và S3 deep archive cho long-term storage để tối ưu chi phí hơn nữa!