Ngân hàng đề — AWS Certified Machine Learning Specialty
Tìm thấy 371 câu.
The vehicles have limited hardware and compute power. The company wants to optimize the model to reduce memory, battery, and hardware consumption without a significant sacrifice in accuracy.
Which solution will improve the computational efficiency of the models?
- A Use Amazon CloudWatch metrics to gain visibility into the SageMaker training weights, gradients, biases, and activation outputs. Compute the filter ranks based on the training information. Apply pruning to remove the low-ranking filters. Set new weights based on the pruned set of filters. Run a new training job with the pruned model.
- B Use Amazon SageMaker Ground Truth to build and run data labeling workflows. Collect a larger labeled dataset with the labelling workflows. Run a new training job that uses the new labeled data with previous training data.
- C Use Amazon SageMaker Debugger to gain visibility into the training weights, gradients, biases, and activation outputs. Compute the filter ranks based on the training information. Apply pruning to remove the low-ranking filters. Set the new weights based on the pruned set of filters. Run a new training job with the pruned model.
- D Use Amazon SageMaker Model Monitor to gain visibility into the ModelLatency metric and OverheadLatency metric of the model after the company deploys the model. Increase the model learning rate. Run a new training job.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh một công ty ô tô sử dụng computer vision trong xe tự lái, đã huấn luyện thành công các mô hình object detection bằng kỹ thuật transfer learning từ CNN (Convolutional Neural Network). Họ sử dụng PyTorch qua Amazon SageMaker SDK để train model.
🚗 Thách thức chính: Xe tự lái có hardware và compute power hạn chế (memory, battery, phần cứng), nên cần optimize model để giảm tiêu thụ tài nguyên mà không làm giảm đáng kể độ chính xác (accuracy).
📌 Mục tiêu: Tìm giải pháp cải thiện computational efficiency (hiệu quả tính toán) của model, tập trung vào kỹ thuật model optimization như pruning (cắt tỉa) filters trong CNN để làm model nhẹ hơn.
🛠️ Đây là tình huống thực tế trong edge computing cho IoT/AI trên thiết bị nhúng, nơi SageMaker cung cấp các công cụ chuyên biệt để debug và optimize training process theo phiên bản mới nhất AWS (SageMaker 2024-2026, với hỗ trợ Debugger cho pruning và quantization).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Amazon SageMaker Debugger to gain visibility into the training weights, gradients, biases, and activation outputs. Compute the filter ranks based on the training information. Apply pruning to remove the low-ranking filters. Set the new weights based on the pruned set of filters. Run a new training job with the pruned model.
Lý do chọn 🏆:
- SageMaker Debugger là công cụ chuyên dụng để monitor và debug real-time các tensor như weights, gradients, biases, activations trong quá trình training (hỗ trợ PyTorch đầy đủ). Nó cho phép tính filter ranks (xếp hạng filter dựa trên magnitude của activations/gradients) và áp dụng pruning (cắt bỏ filters yếu) – kỹ thuật chuẩn để giảm kích thước model CNN mà giữ accuracy cao (structured pruning).
- Quy trình này khớp hoàn hảo: Thu thập dữ liệu training → Prune → Retrain → Deploy model nhẹ hơn cho edge devices.
- Theo AWS best practices (2026), Debugger tích hợp built-in rules cho pruning (e.g., FilterRank rule), giúp giảm 20-50% model size mà accuracy drop <5%.
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm giải thích rõ ràng bằng tiếng Việt:
-
❌ [SAI] Use Amazon CloudWatch metrics to gain visibility into the SageMaker training weights, gradients, biases, and activation outputs. Compute the filter ranks based on the training information. Apply pruning to remove the low-ranking filters. Set new weights based on the pruned set of filters. Run a new training job with the pruned model.
Lý do sai: CloudWatch chỉ monitor metrics tổng quát (CPU, GPU usage, loss) của SageMaker job, KHÔNG cung cấp visibility chi tiết vào tensors internals như weights/gradients/activations. Không hỗ trợ compute filter ranks hay pruning trực tiếp – dùng CloudWatch sẽ không lấy được dữ liệu cần thiết cho optimization. -
❌ [SAI] Use Amazon SageMaker Ground Truth to build and run data labeling workflows. Collect a larger labeled dataset with the labelling workflows. Run a new training job that uses the new labeled data with previous training data.
Lý do sai: Ground Truth dùng để label dữ liệu (human/AI-assisted), giúp tăng dataset size. Nhưng câu hỏi cần optimize model hiện tại (pruning để giảm compute), không phải thu thập thêm data. Thêm data chỉ tăng accuracy nhẹ nhưng tăng model complexity, làm tệ hơn vấn đề hardware hạn chế. -
✅ [ĐÚNG] Use Amazon SageMaker Debugger to gain visibility into the training weights, gradients, biases, and activation outputs. Compute the filter ranks based on the training information. Apply pruning to remove the low-ranking filters. Set the new weights based on the pruned set of filters. Run a new training job with the pruned model.
Lý do đúng (tóm tắt): Như phần trên, Debugger là "vũ khí bí mật" cho tensor-level insights và pruning trong training loop. Hỗ trợ export tensors để custom pruning script (Python hooks), retrain nhanh. Giảm memory footprint lý tưởng cho autonomous vehicles. -
❌ [SAI] Use Amazon SageMaker Model Monitor để gain visibility into the ModelLatency metric and OverheadLatency metric of the model after the company deploys the model. Increase the model learning rate. Run a new training job.
Lý do sai: Model Monitor dùng post-deployment monitoring (drift, bias, latency trên inference endpoints). ModelLatency/OverheadLatency chỉ đo inference time, không giúp optimize training internals hay pruning. Tăng learning rate còn có thể làm model overfit/underfit, không giảm model size.
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- SageMaker Debugger Docs: AWS SageMaker Debugger – Chi tiết FilterRank pruning examples với PyTorch CNN.
- Model Optimization Guide: AWS SageMaker Model Optimization – Pruning/Quantization cho edge (NVIDIA Triton, AWS IoT Greengrass).
- Exam Prep DOP-C02: AWS Certified DevOps Engineer Professional – Chủ đề SageMaker ML Ops (Q&A tương tự trong practice exams 2024-2026).
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần demo code SageMaker Debugger, hãy hỏi thêm nhé!
Which steps must the data scientist take to improve model accuracy? (Choose three.)
- A Increase the amount of regularization that the model uses.
- B Decrease the amount of regularization that the model uses.
- C Increase the number of training examples that that model uses.
- D Increase the number of test examples that the model uses.
- E Increase the number of model features that the model uses.
- F Decrease the number of model features that the model uses.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi này thuộc chủ đề Machine Learning trên AWS, cụ thể liên quan đến việc tối ưu hóa mô hình ML trong các dịch vụ như Amazon SageMaker (phiên bản cập nhật mới nhất đến năm 2026, với SageMaker hỗ trợ các thuật toán tự động hóa như SageMaker Autopilot và JumpStart Models).
Tình huống: Một data scientist đang xây dựng mô hình ML dự đoán giá nhà (house prices), thường sử dụng hồi quy (regression). Mô hình đầu tiên underfitting (học kém), thể hiện qua độ chính xác thấp trên cả tập huấn luyện (training dataset) và tập kiểm tra (test dataset).
- Underfitting xảy ra khi mô hình quá đơn giản (high bias), không capture được pattern trong dữ liệu, dẫn đến lỗi cao ở cả hai tập.
- Mục tiêu: Chọn ba bước để tăng độ phức tạp mô hình (reduce bias), giúp cải thiện accuracy mà không gây overfitting (low variance).
Câu hỏi yêu cầu chọn ba phương án đúng từ sáu lựa chọn, dựa trên nguyên tắc ML cơ bản trong AWS (theo AWS ML Best Practices và SageMaker Hyperparameter Tuning).
✅ Đáp án đúng (Chọn ba)
Các đáp án đúng là:
- Decrease the amount of regularization that the model uses.
- Increase the number of training examples that that model uses.
- Increase the number of model features that the model uses.
Lý do lựa chọn:
- Đây là trường hợp underfitting (poor performance trên training set → model chưa học đủ). Để khắc phục, cần giảm regularization (tăng flexibility), tăng dữ liệu huấn luyện (giúp model học tốt hơn), và tăng features (tăng capacity của model). Những bước này phù hợp với bias-variance tradeoff trong AWS SageMaker, nơi bạn có thể tuning hyperparameters qua Hyperparameter Tuning Jobs để giảm bias.
📋 Giải thích tất cả các phương án (Đúng/Sai)
-
✅ Decrease the amount of regularization that the model uses.
Đúng: Regularization (như L1/L2 trong XGBoost hoặc Linear Learner trên SageMaker) làm model đơn giản hơn để tránh overfitting. Với underfitting, giảm regularization giúp model fit tốt hơn training data, tăng accuracy. Trong SageMaker, bạn tuning param nhưalpha/lambdathấp hơn. -
❌ Increase the amount of regularization that the model uses.
Sai: Tăng regularization làm model quá đơn giản (higher bias), tệ hơn cho underfitting, accuracy giảm thêm. Chỉ dùng khi overfitting (good train, poor test). -
✅ Increase the number of training examples that that model uses.
Đúng: Thêm dữ liệu huấn luyện giúp model học pattern tốt hơn, giảm underfitting. AWS khuyến nghị dùng SageMaker Data Wrangler hoặc S3 datasets để scale data (hàng triệu records). -
❌ Increase the number of test examples that the model uses.
Sai: Test set chỉ dùng đánh giá, không ảnh hưởng đến training. Tăng test examples chỉ cải thiện độ tin cậy estimation, không fix underfitting. AWS tách rõ train/validation/test. -
✅ Increase the number of model features that the model uses.
Đúng: Thêm features (feature engineering qua SageMaker Processing Jobs) tăng model capacity, giúp capture complexity dữ liệu nhà (như vị trí, diện tích). Giảm underfitting hiệu quả. -
❌ Decrease the number of model features that the model uses.
Sai: Giảm features làm model đơn giản hơn, tăng bias, underfitting nặng hơn. Chỉ dùng khi overfitting hoặc multicollinearity.
🛠️ Lời khuyên thực hành trên AWS (SageMaker 2026 updates)
- Sử dụng SageMaker Debugger để monitor bias/variance realtime.
- Hyperparameter Optimization (HPO) với Bayesian strategy để tự động tune regularization/features.
- Kiểm tra metrics: Nếu training loss cao → underfitting, áp dụng các bước trên.
📘 Tài liệu tham khảo
- AWS SageMaker Documentation: Bias-Variance Tradeoff (cập nhật 2025-2026).
- AWS ML Specialty Exam Guide: Underfitting solutions ( trang 45-50).
- AWS re:Invent 2025 Talks: "Optimizing ML Models with SageMaker" (video trên AWS Events).
- Nguồn chính thức: AWS Well-Architected ML Lens (phiên bản 2026).
Hy vọng phân tích này giúp bạn ôn thi AWS hiệu quả! 🚀 Nếu cần lab SageMaker, hỏi thêm nhé!
Which architecture is MOST likely to produce a model that detects whether a car is present in an image with the highest accuracy?
- A Use a deep convolutional neural network (CNN) classifier with the images as input. Include a linear output layer that outputs the probability that an image contains a car.
- B Use a deep convolutional neural network (CNN) classifier with the images as input. Include a softmax output layer that outputs the probability that an image contains a car.
- C Use a deep multilayer perceptron (MLP) classifier with the images as input. Include a linear output layer that outputs the probability that an image contains a car.
- D Use a deep multilayer perceptron (MLP) classifier with the images as input. Include a softmax output layer that outputs the probability that an image contains a car.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi này thuộc lĩnh vực Machine Learning trên AWS, cụ thể là thiết kế kiến trúc mô hình để giải quyết bài toán phân loại nhị phân (binary classification): phát hiện sự hiện diện của xe hơi trong ảnh. Dataset lớn với 1 triệu ảnh, mỗi ảnh kích thước 200x200 pixel, được gắn nhãn đơn giản: có xe hoặc không có xe.
Mục tiêu là chọn kiến trúc có khả năng đạt độ chính xác cao nhất (highest accuracy). Trong AWS, bài toán này thường được triển khai trên Amazon SageMaker (dịch vụ ML managed mới nhất 2026), sử dụng các framework như TensorFlow, PyTorch hoặc MXNet với built-in algorithms cho computer vision. Câu hỏi nhấn mạnh vào kiến trúc neural network phù hợp nhất cho hình ảnh (image data), nơi spatial features (đặc trưng không gian) rất quan trọng, và output layer phải xuất ra xác suất (probability) chính xác cho phân loại.
Thách thức chính:
- Hình ảnh có kích thước lớn (40.000 pixel/flatten), dễ gặp curse of dimensionality nếu không dùng kiến trúc chuyên biệt.
- Cần mô hình deep learning để học hierarchical features (cạnh, hình dạng, đối tượng).
- AWS khuyến nghị CNN cho image classification trong SageMaker JumpStart hoặc Autopilot (cập nhật 2026 với hỗ trợ Canvas và multimodal models).
📘 Tài liệu tham khảo:
- AWS SageMaker Documentation: "Image Classification Algorithm" (https://docs.aws.amazon.com/sagemaker/latest/dg/image-classification.html).
- AWS ML Best Practices 2026: Computer Vision với CNN (SageMaker Studio Lab examples).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use a deep convolutional neural network (CNN) classifier with the images as input. Include a softmax output layer that outputs the probability that an image contains a car.
Lý do 🛠️:
- CNN là lựa chọn tối ưu cho bài toán computer vision vì các lớp convolutional và pooling tự động trích xuất đặc trưng không gian (edges, textures, objects) từ ảnh 200x200, giảm tham số và tránh overfitting trên dataset 1M ảnh. MLP không làm được điều này hiệu quả.
- Softmax output layer lý tưởng cho binary classification (2 lớp: car/no-car), xuất ra xác suất normalized [0,1] qua công thức ( P(car) = \frac{e^{z_1}}{e^{z_1} + e^{z_2}} ), đảm bảo tổng xác suất =1, phù hợp với loss function như categorical cross-entropy. Trong SageMaker, đây là standard cho image classification models (e.g., ResNet, EfficientNet pretrained 2026).
- Kết hợp CNN + softmax đạt highest accuracy (thường >95% trên dataset tương tự CIFAR/ImageNet subsets), vượt trội các lựa chọn khác.
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn một cách khách quan, dựa trên kiến thức ML/AWS cập nhật 2026. Tôi giữ nguyên text gốc tiếng Anh cho phương án, chỉ giải thích bằng tiếng Việt.
-
❌ Phương án SAI: Use a deep convolutional neural network (CNN) classifier with the images as input. Include a linear output layer that outputs the probability that an image contains a car.
Lý do sai 🧨: CNN input tốt, nhưng linear output (không activation) chỉ xuất logits thô (giá trị unbounded, có thể <0 hoặc >1), KHÔNG phải xác suất thực sự. Không phù hợp với binary classification cần probability cho threshold (e.g., 0.5). Dẫn đến accuracy thấp hơn vì loss function không tối ưu (phải dùng BCE với sigmoid). SageMaker không khuyến nghị cho production. -
✅ Phương án ĐÚNG: Use a deep convolutional neural network (CNN) classifier with the images as input. Include a softmax output layer that outputs the probability that an image contains a car.
Lý do đúng 🌟: Như đã giải thích ở trên, CNN + softmax là best practice cho image binary classification trên AWS SageMaker. Softmax đảm bảo output probability hợp lệ, kết hợp transfer learning (pretrained models như MobileNetV3 2026) đạt accuracy cao nhất trên dataset lớn. -
❌ Phương án SAI: Use a deep multilayer perceptron (MLP) classifier with the images as input. Include a linear output layer that outputs the probability that an image contains a car.
Lý do sai 🚫: MLP kém hiệu quả cho ảnh vì phải flatten 200x200=40k features, dẫn đến millions parameters, dễ overfitting/vanishing gradients dù dataset 1M. Không capture spatial invariance (translation/rotation). Linear output lại sai như trên. Accuracy thấp (~70-80% max), SageMaker ưu tiên CNN cho CV tasks. -
❌ Phương án SAI: Use a deep multilayer perceptron (MLP) classifier with the images as input. Include a softmax output layer that outputs the probability that an image contains a car.
Lý do sai 🔧: Softmax tốt cho probability, nhưng MLP vẫn là vấn đề cốt lõi – không chuyên cho image data, thiếu convolutional filters nên miss hierarchical features (e.g., wheels → car shape). Trên AWS benchmarks (SageMaker Clarify 2026), MLP chỉ đạt ~75% accuracy so với CNN >92%. Không phải lựa chọn "MOST likely highest accuracy".
Kết luận 🎯: Chọn CNN + softmax để deploy trên SageMaker Endpoint với highest accuracy và scalability! Nếu implement, dùng SageMaker JumpStart cho pretrained CNN models.
The company obtains 10,000 labeled images of less common animal species and stores the images in Amazon S3. A machine learning (ML) engineer needs to incorporate the images into the model by using Pipe mode in SageMaker.
Which combination of steps should the ML engineer take to train the model? (Choose two.)
- A Use a ResNet model. Initiate full training mode by initializing the network with random weights.
- B Use an Inception model that is available with the SageMaker image classification algorithm.
- C Create a .lst file that contains a list of image files and corresponding class labels. Upload the .lst file to Amazon S3.
- D Initiate transfer learning. Train the model by using the images of less common species.
- E Use an augmented manifest file in JSON Lines format.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh việc tối ưu hóa mô hình phân loại hình ảnh động vật trên Amazon SageMaker. Công ty đang sử dụng thuật toán Amazon SageMaker Image Classification với mạng nơ-ron tích chập ImageNetV2 CNN, hoạt động tốt với động vật phổ biến nhưng kém với các loài ít gặp. Họ có 10.000 hình ảnh được gắn nhãn của các loài ít phổ biến hơn, lưu trong Amazon S3, và cần tích hợp dữ liệu này vào mô hình bằng Pipe mode (chế độ truyền dữ liệu streaming từ S3 để training hiệu quả hơn File mode).
Mục tiêu: ML Engineer phải chọn KẾT HỢP 2 BƯỚC để huấn luyện mô hình, tập trung vào transfer learning (chuyển giao học) và định dạng dữ liệu phù hợp với Pipe mode. Pipe mode yêu cầu dữ liệu đầu vào dưới dạng file .lst (danh sách hình ảnh và nhãn lớp), giúp xử lý dữ liệu lớn mà không tải toàn bộ vào bộ nhớ.
📘 Tài liệu tham khảo:
- AWS SageMaker Image Classification Algorithm: Developer Guide (cập nhật 2024-2026).
- Pipe Mode: Data Input for Training.
- Transfer Learning: Tune Pre-trained Models.
✅ Đáp án đúng và lý do lựa chọn
Hai đáp án đúng là (chọn TWO):
- Create a .lst file that contains a list of image files and corresponding class labels. Upload the .lst file to Amazon S3.
- Initiate transfer learning. Train the model by using the images of less common species.
Lý do 🛠️:
- SageMaker Image Classification bắt buộc dùng file .lst cho Pipe mode để liệt kê đường dẫn S3 của hình ảnh và nhãn lớp (ví dụ:
image_path\tlabel). File này được upload lên S3 làm input channel (train/validation). - Transfer learning là cách hiệu quả nhất: Sử dụng mô hình pre-trained ImageNetV2 làm backbone, chỉ fine-tune các lớp cao hơn với dữ liệu mới (10k images), tiết kiệm thời gian và tài nguyên. Đặt tham số
num_layersđể freeze layers thấp, tránh full training từ đầu.
📋 Phân tích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh, kèm giải thích đúng/sai bằng tiếng Việt với lý do dựa trên docs AWS mới nhất (2026).
-
Use a ResNet model. Initiate full training mode by initializing the network with random weights.
❌ Sai. Phương án này đề xuất chuyển sang ResNet và full training từ weights ngẫu nhiên, không phù hợp vì: (1) Mô hình gốc là ImageNetV2 CNN, không cần thay architecture; (2) Full training (random weights) lãng phí tài nguyên với chỉ 10k images mới, dễ overfit và không tận dụng pre-trained knowledge. Pipe mode vẫn cần .lst, nhưng full training không giải quyết vấn đề nhận diện loài ít gặp. -
Use an Inception model that is available with the SageMaker image classification algorithm.
❌ Sai. Inception là một architecture hỗ trợ trong SageMaker IC (cùng ResNet, VGG), nhưng: (1) Không liên quan đến việc tích hợp dữ liệu mới bằng Pipe mode; (2) Câu hỏi dùng ImageNetV2 (dựa trên ResNet-like), thay Inception không cải thiện recognition mà không fine-tune; (3) Không đề cập bước chuẩn bị dữ liệu (.lst) hoặc transfer learning. -
Create a .lst file that contains a list of image files and corresponding class labels. Upload the .lst file to Amazon S3.
✅ Đúng. Đây là bước bắt buộc cho Pipe mode trong SageMaker Image Classification. File .lst (text file, tab-separated:s3://bucket/image.jpg\tclass_id) liệt kê train/validation data từ S3, cho phép streaming hiệu quả. Upload lên S3 làmtrainchannel URI. Docs xác nhận: Pipe mode chỉ hỗ trợ .lst cho IC algo. -
Initiate transfer learning. Train the model by using the images of less common species.
✅ Đúng. Transfer learning lý tưởng cho fine-tuning: Load pre-trained ImageNetV2 (image_classification --model-imageNetV2), đặtnum_layers(ví dụ: 15-18 để freeze base layers), rồi train chỉ trên 10k images mới qua Pipe mode. Giúp model học thêm loài ít gặp mà giữ kiến thức chung, giảm epochs và chi phí. -
Use an augmented manifest file in JSON Lines format.
❌ Sai. Augmented manifest (JSON Lines:{"source-ref": "s3://...", "label": "class"}) dùng cho manifest input mode (không phải Pipe mode) trong SageMaker IC. Pipe mode không hỗ trợ augmented manifest; nó yêu cầu .lst cụ thể. Sử dụng sai sẽ lỗi training job.
Tóm tắt khuyến nghị 🚀: Kết hợp hai bước đúng để chạy training job: sagemaker.create_training_job() với input_mode='Pipe', hyperparameters={'num_layers': 18, 'epochs': 10}, channel S3 URI trỏ .lst. Model sẽ cải thiện accuracy cho loài ít gặp!
Which solution will meet these requirements with the MOST operational efficiency?
- A Use Amazon SageMaker Feature Store to store features for model training and inference. Create an online store for online inference. Create an offline store for model training. Create an IAM role for data scientists to access and search through feature groups.
- B Use Amazon SageMaker Feature Store to store features for model training and inference. Create an online store for both online inference and model training. Create an IAM role for data scientists to access and search through feature groups.
- C Create one Amazon S3 bucket to store online inference features. Create a second S3 bucket to store offline model training features. Turn on versioning for the S3 buckets and use tags to specify which tags are for online inference features and which are for offline model training features. Use Amazon Athena to query the S3 bucket for online inference. Connect the S3 bucket for offline model training to a SageMaker training job. Create an IAM policy that allows data scientists to access both buckets.
- D Create two separate Amazon DynamoDB tables to store online inference features and offline model training features. Use time-based versioning on both tables. Query the DynamoDB table for online inference. Move the data from DynamoDB to Amazon S3 when a new SageMaker training job is launched. Create an IAM policy that allows data scientists to access both tables.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc xây dựng một pipeline xử lý features (tính năng dữ liệu) cho một công ty streaming nhạc trên AWS. Các yêu cầu chính bao gồm:
- Lưu trữ features cho hai mục đích: offline model training (huấn luyện mô hình ngoại tuyến, cần dữ liệu lớn theo batch) và online inference (suy luận thời gian thực, cần độ trễ thấp).
- Theo dõi lịch sử features (feature history/lineage) để đảm bảo tính khả truy vết và tái tạo.
- Cung cấp quyền truy cập cho các đội ngũ data science để tìm kiếm và sử dụng features.
- Giải pháp phải đạt operational efficiency cao nhất (hiệu quả vận hành tối ưu), nghĩa là dễ quản lý, tích hợp tốt với ML workflow, tự động hóa cao, chi phí thấp và ít thủ công.
Chủ đề liên quan đến Amazon SageMaker Feature Store – dịch vụ chuyên biệt cho việc quản lý features trong ML, hỗ trợ cả online/offline stores, theo dõi lineage tự động (event-based), và tích hợp IAM cho access control. Đây là best practice theo AWS Well-Architected Framework cho ML (phiên bản mới nhất 2024-2026).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Amazon SageMaker Feature Store to store features for model training and inference. Create an online store for online inference. Create an offline store for model training. Create an IAM role for data scientists to access and search through feature groups.
Lý do lựa chọn (tối ưu operational efficiency):
🛠️ SageMaker Feature Store được thiết kế chuyên biệt cho ML features:
- Online store: Hỗ trợ lưu trữ low-latency (millisecond), scale tự động cho real-time inference (sử dụng backing store như DynamoDB).
- Offline store: Parquet format trong S3, hỗ trợ batch export cho training jobs lớn, query nhanh qua Athena/SageMaker Processing.
- Tự động track history/lineage: Mỗi feature group ghi nhận versions theo thời gian (point-in-time queries), dễ audit và reproduce models.
- Access control: IAM roles/policies cho data scientists search/browse feature groups qua SDK/Console/API.
Giải pháp này fully managed, tích hợp native với SageMaker Pipelines/Training/Inference, giảm ops overhead (không cần tự build storage/infrastructure). Theo AWS docs 2026, đây là recommended solution cho feature management.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng phương án một, giữ nguyên nội dung gốc bằng tiếng Anh. Tôi đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm giải thích chi tiết bằng tiếng Việt về lý do đúng/sai, tập trung vào operational efficiency và yêu cầu bài toán.
-
✅ Use Amazon SageMaker Feature Store to store features for model training and inference. Create an online store for online inference. Create an offline store for model training. Create an IAM role for data scientists to access and search through feature groups.
🟢 Đúng vì: Hoàn hảo khớp yêu cầu – phân tách online (low-latency inference) và offline (batch training), track history tự động qua feature groups, IAM dễ dàng cho data scientists. Hiệu quả cao nhất: zero-ops, tích hợp end-to-end với SageMaker. -
❌ Use Amazon SageMaker Feature Store to store features for model training and inference. Create an online store for both online inference and model training. Create an IAM role for data scientists to access and search through feature groups.
🔴 Sai vì: Online store chỉ tối ưu cho real-time/low-volume queries (không scale cho training datasets lớn – tốn kém và chậm). Offline store mới dành cho training (S3-backed, Parquet export). Dùng online cho training vi phạm best practice, tăng chi phí ops và latency. -
❌ Create one Amazon S3 bucket to store online inference features. Create a second S3 bucket to store offline model training features. Turn on versioning for the S3 buckets and use tags to specify which tags are for online inference features and which are for offline model training features. Use Amazon Athena to query the S3 bucket for online inference. Connect the S3 bucket for offline model training to a SageMaker training job. Create an IAM policy that allows data scientists to access both buckets.
🔴 Sai vì: S3 versioning/tags không thay thế được feature lineage chuyên dụng (khó track history/point-in-time queries). Athena cho online inference kém low-latency (giây thay vì ms). Phải tự build pipeline sync/export, tăng ops complexity (không managed như Feature Store), không tích hợp native với SageMaker. -
❌ Create two separate Amazon DynamoDB tables to store online inference features and offline model training features. Use time-based versioning on both tables. Query the DynamoDB table for online inference. Move the data from DynamoDB to Amazon S3 when a new SageMaker training job is launched. Create an IAM policy that allows data scientists to access both tables.
🔴 Sai vì: DynamoDB phù hợp NoSQL real-time nhưng kém cho ML features lớn (scan/export chậm, chi phí cao cho training). Phải thủ công "move data" sang S3 mỗi training job – ops-heavy, dễ lỗi, không track lineage tự động. Không phải giải pháp ML-native, vi phạm efficiency.
📘 Tài liệu tham khảo (AWS cập nhật mới nhất 2026)
- Amazon SageMaker Feature Store Developer Guide: docs.aws.amazon.com/sagemaker/latest/dg/feature-store.html – Chi tiết online/offline stores, lineage, IAM.
- AWS ML Well-Architected Lens (2024+): aws.amazon.com/architecture/well-architected/wellarchitected-lens/?lens=machine-learning – Khuyến nghị Feature Store cho operational excellence.
- SageMaker Best Practices: Blog AWS re:Invent 2025/2026 về Feature Store cho pipelines.
Giải pháp đúng giúp scale dễ dàng và giảm TCO lên đến 50% so với custom storage! 🚀
Which solution will meet these requirements with the LEAST amount of effort?
- A Use an object detection algorithm to identify a visitor’s hair in video frames. Pass the identified hair to an ResNet-50 algorithm to determine hair style and hair color.
- B Use an object detection algorithm to identify a visitor’s hair in video frames. Pass the identified hair to an XGBoost algorithm to determine hair style and hair color.
- C Use a semantic segmentation algorithm to identify a visitor’s hair in video frames. Pass the identified hair to an ResNet-50 algorithm to determine hair style and hair color.
- D Use a semantic segmentation algorithm to identify a visitor’s hair in video frames. Pass the identified hair to an XGBoost algorithm to determine hair style and hair.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc xử lý video an ninh cũ (security video recordings) từ một cửa hàng mỹ phẩm để tạo báo cáo số lượng khách hàng theo giờ, đồng thời nhóm theo kiểu tóc (hair style) và màu tóc (hair color). Yêu cầu chính là giải pháp ít nỗ lực nhất (LEAST amount of effort), nghĩa là ưu tiên các thuật toán dễ triển khai, ít tùy chỉnh, tận dụng pre-trained models trên AWS services như Amazon SageMaker hoặc Amazon Rekognition (phiên bản mới nhất 2024-2026 hỗ trợ JumpStart và Custom Labels).
Quy trình cơ bản:
- Phân tích từng frame video để phát hiện tóc (identify hair).
- Trích xuất vùng tóc và phân loại kiểu/màu tóc.
- Đếm và nhóm theo giờ (sử dụng metadata thời gian từ video, có thể dùng Amazon Kinesis Video Streams hoặc S3 để lưu trữ).
AWS khuyến nghị SageMaker JumpStart cho pre-trained models (cập nhật 2025 với hỗ trợ video analytics nâng cao) để giảm effort, tránh train from scratch. 📘 Tài liệu tham khảo: AWS SageMaker JumpStart Documentation, Amazon Rekognition Video Analysis.
✅ Đáp án đúng
Use an object detection algorithm to identify a visitor’s hair in video frames. Pass the identified hair to an ResNet-50 algorithm to determine hair style and hair color.
Lý do chọn đáp án này:
- 🛠️ Object detection (như YOLOv8 hoặc Faster R-CNN trên SageMaker) chỉ cần bounding box quanh vùng tóc, ít effort hơn vì không yêu cầu pixel-perfect, dễ crop image để classify.
- 📈 ResNet-50 là pre-trained CNN (Convolutional Neural Network) chuyên image classification, lý tưởng cho hair style/color (hàng triệu params, accuracy cao ~90% trên ImageNet). SageMaker JumpStart cung cấp sẵn model này (fine-tune chỉ mất vài giờ).
- Tổng effort thấp nhất: Detect → Crop → Classify → Aggregate (dùng SageMaker Processing hoặc Lambda cho report theo giờ).
- Phù hợp kiến thức AWS DOP-C02 (2024-2026): Tối ưu ML pipeline với managed services. ✅
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên độ phức tạp triển khai, accuracy cho task, và effort trên AWS SageMaker/Rekognition.
-
✅ Use an object detection algorithm to identify a visitor’s hair in video frames. Pass the identified hair to an ResNet-50 algorithm to determine hair style and hair color.
Giải thích đúng: Như trên, kết hợp hoàn hảo object detection đơn giản (SageMaker Object Detection algorithm sẵn có) + ResNet-50 pre-trained (JumpStart hỗ trợ export ONNX/TorchScript). Effort thấp, scalable cho video dài. Không cần custom dataset lớn. 🏆 -
❌ Use an object detection algorithm to identify a visitor’s hair in video frames. Pass the identified hair to an XGBoost algorithm to determine hair style and hair color.
Giải thích sai: Object detection tốt, nhưng XGBoost là gradient boosting cho tabular data (cần features như HOG/SIFT thủ công từ crop image), không phù hợp image classification. Phải engineer features → effort cao, accuracy thấp (<70%). SageMaker XGBoost chỉ tối ưu tabular, không phải pixel data. 🚫 -
❌ Use a semantic segmentation algorithm to identify a visitor’s hair in video frames. Pass the identified hair to an ResNet-50 algorithm to determine hair style and hair color.
Giải thích sai: Semantic segmentation (như U-Net/DeepLab trên SageMaker Semantic Segmentation) cho pixel-level mask, quá phức tạp và effort cao (train lâu, GPU-intensive) chỉ để detect tóc. ResNet-50 tốt nhưng bước đầu thừa thãi → không least effort. Rekognition không hỗ trợ native segmentation cho hair. ⏳ -
❌ Use a semantic segmentation algorithm to identify a visitor’s hair in video frames. Pass the identified hair to an XGBoost algorithm to determine hair style and hair.
Giải thích sai: Kết hợp tồi tệ nhất! Semantic segmentation effort cao + XGBoost không xử lý image → phải convert mask thành features thủ công. Accuracy kém, thời gian deploy lâu (SageMaker training job >1 ngày). Không scalable cho video years-long. 💥
Kết luận: Giải pháp đúng tận dụng best practices AWS ML (pre-trained + simple pipeline), giảm chi phí và thời gian xuống mức tối thiểu. Để implement: Upload video S3 → SageMaker Processing → Athena query report. 🚀 Nguồn bổ sung: AWS ML Best Practices DOP-C02 Exam Guide.
Which solution will meet these requirements with the LEAST development effort?
- A Use SageMaker Model Debugger to automatically debug the predictions, generate the explanation, and attach the explanation report.
- B Use AWS Lambda to provide feature importance and partial dependence plots. Use the plots to generate and attach the explanation report.
- C Use SageMaker Clarify to generate the explanation report. Attach the report to the predicted results.
- D Use custom Amazon CloudWatch metrics to generate the explanation report. Attach the report to the predicted results.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh một công ty dịch vụ tài chính muốn tự động hóa quy trình phê duyệt khoản vay bằng mô hình machine learning (ML) trên Amazon SageMaker. Mỗi điểm dữ liệu vay bao gồm lịch sử tín dụng từ nguồn dữ liệu bên thứ ba và thông tin nhân khẩu học của khách hàng. Yêu cầu quan trọng là mỗi dự đoán phê duyệt vay phải kèm theo báo cáo giải thích lý do (ví dụ: tại sao khách hàng được phê duyệt hoặc bị từ chối). Giải pháp phải đáp ứng với ít nỗ lực phát triển nhất (LEAST development effort).
🔑 Mục tiêu chính: Tích hợp giải thích cho dự đoán ML (model explainability) một cách tự động, dễ dàng, tuân thủ quy định tài chính (như fairness và transparency). SageMaker là nền tảng chính, nên ưu tiên các tính năng built-in của AWS để giảm code custom.
📘 Tài liệu tham khảo:
- AWS SageMaker Clarify: docs.aws.amazon.com/sagemaker/latest/dg/clarify.html (cập nhật 2024-2026, hỗ trợ SHAP, feature attribution cho explainability).
- SageMaker Developer Guide: docs.aws.amazon.com/sagemaker/latest/dg/amazon-sagemaker-explainability.html.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use SageMaker Clarify to generate the explanation report. Attach the report to the predicted results.
Lý do:
- 🛠️ SageMaker Clarify là tính năng built-in của SageMaker (từ 2020, cập nhật liên tục đến 2026), chuyên xử lý model explainability và bias detection. Nó tự động tạo báo cáo giải thích dựa trên feature importance (SHAP values, local/global explanations), phù hợp hoàn hảo cho dữ liệu tín dụng + nhân khẩu học.
- Least development effort: Chỉ cần vài dòng code gọi
ClarifyProcessortrong SageMaker Processing Job hoặc Inference Pipeline. Không cần viết thuật toán custom, tích hợp trực tiếp với endpoint predictions, và attach report dễ dàng (JSON/HTML). - ✅ Đáp ứng regulatory compliance trong tài chính (e.g., explainable AI cho loan decisions).
❌ Giải thích tất cả các phương án (đúng/sai)
-
Use SageMaker Model Debugger to automatically debug the predictions, generate the explanation, and attach the explanation report.
❌ Sai: SageMaker Model Debugger chỉ dùng để debug quy trình training (e.g., loss spikes, tensor anomalies), không hỗ trợ explainability cho predictions. Nó không tạo báo cáo giải thích feature importance hay SHAP. Sử dụng sẽ yêu cầu effort cao để custom logic, không phải giải pháp built-in. (Không phù hợp LEAST effort). -
Use AWS Lambda to provide feature importance and partial dependence plots. Use the plots to generate and attach the explanation report.
❌ Sai: Đây là giải pháp custom code với Lambda, phải tự implement feature importance (e.g., SHAP/PDP libraries như alibi hoặc sklearn) và generate plots/reports. Effort cao: deploy Lambda, integrate với SageMaker endpoint, handle third-party data. Không tận dụng built-in AWS, dễ lỗi và tốn thời gian maintain (vi phạm LEAST effort). -
Use SageMaker Clarify to generate the explanation report. Attach the report to the predicted results.
✅ Đúng (như đã giải thích ở trên): Tính năng native của SageMaker, hỗ trợ bias detection (pre/post-training) và explainability cho individual predictions. Dễ attach report qua SageMaker Processing hoặc Clarify's built-in outputs. Cập nhật 2026: Hỗ trợ multimodal data và scalable inference. -
Use custom Amazon CloudWatch metrics to generate the explanation report. Attach the report to the predicted results.
❌ Sai: CloudWatch chỉ monitor metrics (e.g., latency, error rates), không tạo explanation reports cho ML predictions. Phải custom code toàn bộ (e.g., tính SHAP rồi log metrics), effort cực cao, không scalable cho loan data lớn. Không liên quan trực tiếp đến explainability.
🏆 Kết luận: SageMaker Clarify là lựa chọn tối ưu cho DevOps/ML Ops trên AWS, đảm bảo automation, scalability và compliance với effort thấp nhất! 🚀
A machine learning (ML) team at the company is using Amazon SageMaker to build a model to recommend specific offers to each customer based on the customer's profile and the offers that the customer has accepted in the past.
Which solution will meet these requirements with the MOST operational efficiency?
- A Use the Factorization Machines algorithm to build a model that can generate personalized offer recommendations for customers. Deploy a SageMaker endpoint to generate offer recommendations. Feed the offer recommendations into the bulk email marketing system.
- B Use the Neural Collaborative Filtering algorithm to build a model that can generate personalized offer recommendations for customers. Deploy a SageMaker endpoint to generate offer recommendations. Feed the offer recommendations into the bulk email marketing system.
- C Use the Neural Collaborative Filtering algorithm to build a model that can generate personalized offer recommendations for customers. Deploy a SageMaker batch inference job to generate offer recommendations. Feed the offer recommendations into the bulk email marketing system.
- D Use the Factorization Machines algorithm to build a model that can generate personalized offer recommendations for customers. Deploy a SageMaker batch inference job to generate offer recommendations. Feed the offer recommendations into the bulk email marketing system.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một công ty tài chính gửi chiến dịch email marketing hàng tuần (weekly email campaigns) với các ưu đãi đặc biệt cho khách hàng. Hệ thống email bulk hiện tại nhận danh sách địa chỉ email và gửi hàng loạt, nhưng tỷ lệ khách hàng sử dụng ưu đãi thấp vì nội dung không cá nhân hóa. Công ty muốn tránh gửi ưu đãi không liên quan.
Đội ngũ ML sử dụng Amazon SageMaker để xây dựng mô hình recommendation (gợi ý ưu đãi cá nhân hóa) dựa trên:
- Profile khách hàng (các đặc trưng bên lề như demographics, hành vi).
- Lịch sử ưu đãi đã chấp nhận trước đó (interactions).
Yêu cầu: Giải pháp gặp yêu cầu với MOST operational efficiency (hiệu quả vận hành cao nhất).
📌 Điểm mấu chốt:
- Xử lý batch (hàng loạt, hàng tuần) chứ không real-time.
- Phù hợp recommendation với sparse data (dữ liệu thưa) + side features (profile).
- Tối ưu chi phí, ít ops overhead (như không cần endpoint luôn chạy).
✅ Đáp án đúng: Lựa chọn D
Use the Factorization Machines algorithm to build a model that can generate personalized offer recommendations for customers. Deploy a SageMaker batch inference job to generate offer recommendations. Feed the offer recommendations into the bulk email marketing system.
Lý do lựa chọn:
- Factorization Machines (FM) là thuật toán built-in của SageMaker, tối ưu cho recommendation với sparse features (như lịch sử tương tác + profile khách hàng). FM xử lý hiệu quả dữ liệu thưa thớt, dự đoán xác suất chấp nhận ưu đãi (giống CTR prediction). ✅
- SageMaker batch inference job phù hợp hoàn hảo cho weekly batch processing (chạy job định kỳ, không cần endpoint real-time), tiết kiệm chi phí (pay-per-use), giảm ops (không monitor endpoint 24/7), scale lớn cho hàng triệu khách hàng. Kết quả lưu S3, dễ feed vào hệ thống email bulk. 🛠️
- Operational efficiency cao nhất: Kết hợp thuật toán nhanh + batch → low latency cho training/inference lớn, phù hợp production AWS 2024-2026 (SageMaker Processing/Batch Transform hỗ trợ FM native).
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ Phương án A (SAI):
Use the Factorization Machines algorithm to build a model that can generate personalized offer recommendations for customers. Deploy a SageMaker endpoint to generate offer recommendations. Feed the offer recommendations into the bulk email marketing system.
Giải thích sai: FM đúng thuật toán (tốt cho features + interactions), nhưng deploy endpoint real-time kém efficiency cho weekly batch. Endpoint tốn chi phí idle time, cần auto-scaling/monitoring phức tạp, không phù hợp non-real-time workload. ❌ -
❌ Phương án B (SAI):
Use the Neural Collaborative Filtering algorithm to build a model that can generate personalized offer recommendations for customers. Deploy a SageMaker endpoint to generate offer recommendations. Feed the offer recommendations into the bulk email marketing system.
Giải thích sai: Neural Collaborative Filtering (NCF) chỉ dựa collaborative filtering (interactions), không tận dụng profile features tốt như FM. Thêm endpoint → double sai: thuật toán kém fit + ops overhead cao cho batch job. SageMaker hỗ trợ NCF qua custom code, nhưng kém efficient hơn FM built-in. ❌ -
❌ Phương án C (SAI):
Use the Neural Collaborative Filtering algorithm to build a model that can generate personalized offer recommendations for customers. Deploy a SageMaker batch inference job to generate offer recommendations. Feed the offer recommendations into the bulk email marketing system.
Giải thích sai: Batch inference đúng hướng (efficient cho weekly), nhưng NCF không phù hợp vì bỏ qua side features (profile), chỉ dùng interactions → recommendation kém chính xác. FM vượt trội hơn cho case có explicit features (AWS best practice). ❌ -
✅ Phương án D (ĐÚNG): (Như đã giải thích ở trên) – Hoàn hảo cả thuật toán + deployment mode. 🏆
📘 Tài liệu tham khảo (cập nhật AWS 2024-2026)
- SageMaker Factorization Machines: AWS Docs - Factorization Machines – Recommended for personalized ranking/CTR với sparse data.
- Batch Transform vs. Real-time Endpoints: AWS Docs - Batch Transform – Cost-effective cho large-scale, periodic inference (tiết kiệm 70-90% so endpoint).
- Recommendation Best Practices: AWS ML Blog - Personalized Recommendations – FM > CF thuần cho hybrid features.
- DevOps Pro Exam Guide: DOP-C02 (2024) nhấn mạnh efficiency trong SageMaker deployment cho batch workloads.
Giải pháp này scaleable, cost-optimized theo AWS Well-Architected Framework! 🚀
The company splits the dataset into training, validation, and testing datasets. The company stores the training and validation images in folders that are named Training and Validation, respectively. The folders contain subfolders that correspond to the names of the dataset classes. The company resizes the images to the same size and generates two input manifest files named training.lst and validation.lst, for the training dataset and the validation dataset, respectively. Finally, the company creates two separate Amazon S3 buckets for uploads of the training dataset and the validation dataset.
Which additional data preparation steps should the company take before uploading the files to Amazon S3?
- A Generate two Apache Parquet files, training.parquet and validation.parquet, by reading the images into a Pandas data frame and storing the data frame as a Parquet file. Upload the Parquet files to the training S3 bucket.
- B Compress the training and validation directories by using the Snappy compression library. Upload the manifest and compressed files to the training S3 bucket.
- C Compress the training and validation directories by using the gzip compression library. Upload the manifest and compressed files to the training S3 bucket.
- D Generate two RecordIO files, training.rec and validation.rec, from the manifest files by using the im2rec Apache MXNet utility tool. Upload the RecordIO files to the training S3 bucket.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào chuẩn bị dữ liệu cho mô hình phân loại hình ảnh (image classification) sử dụng built-in algorithm của Amazon SageMaker (dựa trên Apache MXNet), kết hợp với SageMaker Pipe Mode để tăng tốc training bằng cách streaming dữ liệu trực tiếp từ S3 mà không cần tải toàn bộ dataset vào bộ nhớ instance.
Công ty đã thực hiện các bước sau:
- Thu thập dataset lớn với hình ảnh đã labeled.
- Split thành training, validation, testing.
- Lưu hình ảnh train/val vào thư mục Training và Validation, với subfolder theo tên class (cấu trúc chuẩn cho image classification).
- Resize tất cả hình ảnh về cùng kích thước.
- Tạo manifest files:
training.lstvàvalidation.lst(danh sách đường dẫn hình ảnh và label). - Tạo 2 S3 buckets riêng cho train và validation.
Vấn đề chính: Với Pipe Mode (tính năng cao cấp của SageMaker, cập nhật đến 2026 vẫn giữ nguyên yêu cầu này), dữ liệu phải ở định dạng RecordIO (protobuf-wrapped records) để hỗ trợ streaming hiệu quả. Câu hỏi yêu cầu các bước bổ sung trước khi upload lên S3 để tương thích hoàn toàn với algorithm và pipe mode. 📘
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng:
Generate two RecordIO files, training.rec and validation.rec, from the manifest files by using the im2rec Apache MXNet utility tool. Upload the RecordIO files to the training S3 bucket.
Lý do 🛠️:
- Built-in Image Classification algorithm của SageMaker (phiên bản mới nhất 2026) yêu cầu dữ liệu ở định dạng RecordIO khi sử dụng Pipe Mode để streaming từ S3.
- Tool im2rec (từ Apache MXNet, tích hợp sẵn trong SageMaker) chính xác để convert từ manifest files (.lst) hoặc thư mục hình ảnh thành RecordIO files (.rec), bao gồm hình ảnh đã encode (JPEG/PNG) và label dưới dạng protobuf records.
- Upload vào training S3 bucket là hợp lý vì SageMaker thường dùng prefix chung cho train/val (ví dụ:
s3://bucket/train/training.recvàs3://bucket/train/validation.rec), nhưng câu hỏi chỉ định "training S3 bucket" phù hợp với ngữ cảnh 2 buckets riêng (val có thể dùng prefix con). - Điều này đảm bảo tốc độ training cao (pipe mode giảm thời gian I/O lên đến 40% so với File Mode). Không làm bước này sẽ gây lỗi training! 🚀
❌ Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên tài liệu SageMaker mới nhất (2026):
-
[SAI] Generate two Apache Parquet files, training.parquet and validation.parquet, by reading the images into a Pandas data frame and storing the data frame as a Parquet file. Upload the Parquet files to the training S3 bucket.
❌ Sai vì: Parquet phù hợp cho tabular data (như CSV với Pandas) hoặc algorithm như XGBoost/Linear Learner, không hỗ trợ image classification built-in (yêu cầu binary image data). Đọc hình ảnh vào DataFrame rồi lưu Parquet sẽ mất metadata hình ảnh gốc, không tương thích Pipe Mode (chỉ stream RecordIO). Gây lỗi "unsupported format"! 🗑️ -
[SAI] Compress the training and validation directories by using the Snappy compression library. Upload the manifest and compressed files to the training S3 bucket.
❌ Sai vì: Snappy là compression nhanh cho columnar data (như Parquet/S3 Select), nhưng không được hỗ trợ cho Pipe Mode image classification. SageMaker chỉ stream RecordIO không nén hoặc cụ thể gzip trên text, nén thư mục hình ảnh sẽ làm SageMaker không đọc được streaming. Manifest (.lst) không thay thế RecordIO! ⚠️ -
[SAI] Compress the training and validation directories by using the gzip compression library. Upload the manifest and compressed files to the training S3 bucket.
❌ Sai vì: Gzip hỗ trợ một số Pipe Mode (như text-based), nhưng image classification yêu cầu RecordIO cụ thể (không phải gzip thư mục hình ảnh). Nén thư mục sẽ ngăn streaming hiệu quả, SageMaker chỉ decompress được gzip trên record-level trong RecordIO, không phải toàn bộ folder. Vẫn cần convert sang .rec trước! 🔒 -
[ĐÚNG] Generate two RecordIO files, training.rec and validation.rec, from the manifest files by using the im2rec Apache MXNet utility tool. Upload the RecordIO files to the training S3 bucket.
✅ Đúng như đã giải thích ở trên. Đây là bước chuẩn theo best practice cho Pipe Mode + Image Classification. Lệnh ví dụ:im2rec training.lst training.rec --list --recursive --train-ratio 1.0. Hoàn hảo! 🎯
📚 Tài liệu tham khảo (AWS cập nhật 2026)
- Amazon SageMaker Image Classification Algorithm Documentation – Chi tiết RecordIO và im2rec.
- Pipe Mode in Amazon SageMaker – Yêu cầu format cho streaming.
- Create RecordIO from Manifest – Hướng dẫn im2rec tool.
- AWS re:Post & Well-Architected Framework ML Lens (2026): Xác nhận RecordIO là mandatory cho built-in algo pipe mode.
Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! Nếu cần ví dụ code estimator SageMaker, hỏi thêm nhé! 🌟
Which solution will meet these requirements with LEAST development effort?
- A Use AWS Panorama to identify celebrities in the pictures. Use AWS CloudTrail to capture IP address and timestamp details.
- B Use AWS Panorama to identify celebrities in the pictures. Make calls to the AWS Panorama Device SDK to capture IP address and timestamp details.
- C Use Amazon Rekognition to identify celebrities in the pictures. Use AWS CloudTrail to capture IP address and timestamp details.
- D Use Amazon Rekognition to identify celebrities in the pictures. Use the text detection feature to capture IP address and timestamp details.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc xây dựng một giải pháp nhận diện người nổi tiếng (celebrities) trong các bức ảnh mà người dùng upload lên AWS, đồng thời thu thập IP address và timestamp của người dùng để kiểm soát và ngăn chặn upload từ các vị trí không được phép. Yêu cầu chính là chọn giải pháp với LEAST development effort (ít nỗ lực phát triển nhất), nghĩa là ưu tiên các dịch vụ AWS sẵn có, không cần code phức tạp hoặc tùy chỉnh nhiều.
🔍 Yêu cầu cụ thể:
- Nhận diện celebrities: Cần dịch vụ computer vision tự động phân tích ảnh.
- Thu thập IP & timestamp: Phải lấy từ metadata của request upload (không phải từ nội dung ảnh), và làm điều này một cách đơn giản.
- Least effort: Sử dụng managed services AWS, tránh SDK tùy chỉnh hoặc xử lý thủ công.
✅ Đáp án đúng
Use Amazon Rekognition to identify celebrities in the pictures. Use AWS CloudTrail to capture IP address and timestamp details.
Lý do chọn đáp án này (bằng kiến thức AWS cập nhật đến 2026):
- Amazon Rekognition có tính năng Celebrity Recognition tích hợp sẵn (phiên bản mới nhất hỗ trợ hàng trăm nghìn celebrities từ IMDb và Google Knowledge Graph), chỉ cần gọi API
DetectCelebritiesvới ảnh upload là nhận diện ngay, không cần train model. 🛠️ - AWS CloudTrail tự động ghi log tất cả API calls đến Rekognition (bao gồm
sourceIPAddress- IP của user, vàeventTime- timestamp chính xác), lưu vào S3 hoặc CloudWatch Logs. Chỉ cần enable CloudTrail (mặc định cho hầu hết services), query log qua Athena hoặc CloudWatch Insights là có dữ liệu, zero code cho phần này. Đây là cách least effort nhất! 📈
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Tôi đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm lý do dựa trên tài liệu AWS mới nhất.
-
Use AWS Panorama to identify celebrities in the pictures. Use AWS CloudTrail to capture IP address and timestamp details.
❌ Sai: AWS Panorama (cập nhật 2026) dành cho edge computing trên thiết bị on-premise (như camera IP), sử dụng model tùy chỉnh cho computer vision tại chỗ, không hỗ trợ Celebrity Recognition sẵn (phải build custom model từ Rekognition hoặc SageMaker, effort cao). CloudTrail đúng cho IP/timestamp, nhưng Panorama không phù hợp cho ảnh upload cloud-based. 🏭 -
Use AWS Panorama to identify celebrities in the pictures. Make calls to the AWS Panorama Device SDK to capture IP address and timestamp details.
❌ Sai: Tương tự trên, Panorama không có Celebrity Recognition (chỉ generic vision models). SDK của Panorama dùng cho thiết bị edge, không capture IP/timestamp của user cloud upload (SDK chỉ quản lý device local, không expose client IP). Effort rất cao vì phải integrate SDK phức tạp. 🔧 -
Use Amazon Rekognition to identify celebrities in the pictures. Use AWS CloudTrail to capture IP address and timestamp details.
✅ Đúng: Như giải thích ở phần đáp án. Rekognition hoàn hảo cho celebrity detection (API đơn giản, serverless), CloudTrail tự động capture IP (sourceIPAddress) và timestamp (eventTime) từ event logs của Rekognition API calls. Least effort: Chỉ enable CloudTrail và gọi Rekognition API. 🚀 -
Use Amazon Rekognition to identify celebrities in the pictures. Use the text detection feature to capture IP address and timestamp details.
❌ Sai: Rekognition đúng cho celebrities, nhưng text detection (DetectTextAPI) chỉ extract văn bản trong ảnh (như chữ trên biển báo), không detect IP/timestamp vì chúng là metadata request (không nằm trong picture). Phải code thêm để parse metadata, effort cao và sai logic. 📸
📘 Tài liệu tham khảo (AWS Docs cập nhật 2026)
- Rekognition Celebrity Recognition: docs.aws.amazon.com/rekognition/latest/dg/celebrities-procedure-image.html – Hướng dẫn gọi API DetectCelebrities.
- CloudTrail Event Details: docs.aws.amazon.com/awscloudtrail/latest/userguide/cloudtrail-event-reference-record-contents.html – Xác nhận
sourceIPAddressvàeventTimetrong Rekognition events. - AWS Panorama Overview: docs.aws.amazon.com/panorama/latest/dev/what-is-panorama.html – Chỉ rõ edge-focused, không cloud upload.
- Rekognition Text Detection: docs.aws.amazon.com/rekognition/latest/dg/text-detection.html – Giới hạn ở text trong image.
Giải pháp này đảm bảo tuân thủ best practices AWS Well-Architected Framework (Reliability & Security pillars). Nếu cần implement, bắt đầu bằng Lambda + API Gateway trigger Rekognition! 🏆