Ngân hàng đề — AWS Certified Machine Learning Engineer Associate
Tìm thấy 635 câu.
The ML engineer must perform preprocessing on the dataset to ensure that the model produces accurate predictions for the typical house or apartment.
Which solution will meet these requirements?
- A Remove the outliers and perform a log transformation on the Square Meters variable.
- B Keep the outliers and perform normalization on the Square Meters variable.
- C Remove the outliers and perform one-hot encoding on the Square Meters variable.
- D Keep the outliers and perform one-hot encoding on the Square Meters variable.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào preprocessing dữ liệu trong quy trình Machine Learning (ML) trên AWS, cụ thể là xử lý outliers (giá trị ngoại lai) để đảm bảo model dự đoán giá nhà/căn hộ chính xác cho các trường hợp điển hình (typical house or apartment).
-
Bối cảnh dữ liệu: Dataset có 10.000 rows, sử dụng 3 features: Square Meters (diện tích, biến số liên tục), Price (giá, có thể là target), Age of Building (tuổi nhà). Có 2 outliers rõ ràng: một mansion lớn (diện tích cực lớn) và một apartment nhỏ cực kỳ (diện tích cực nhỏ). Những outliers này có thể làm lệch model, vì dataset chủ yếu đại diện cho nhà/căn hộ thông thường.
-
Yêu cầu chính: Preprocessing để loại bỏ ảnh hưởng của outliers, giúp model tập trung vào dữ liệu điển hình, tránh overfitting/underfitting hoặc bias trong predictions. Trong AWS SageMaker (dịch vụ ML chính), preprocessing thường dùng SageMaker Processing Jobs, Data Wrangler, hoặc Scikit-learn Processing để detect/remove outliers và apply transformations.
-
Vấn đề cốt lõi: Outliers trên Square Meters (biến skewed/right-skewed do mansion lớn) cần xử lý bằng remove outliers (ví dụ: IQR method, Z-score) và log transformation (giảm skewness, làm phân phối gần normal hơn, phù hợp cho linear models hoặc neural nets).
Kiến thức cập nhật AWS 2026: SageMaker hỗ trợ Automated Data Processing với outlier detection qua Clarify hoặc Autopilot, và transformations qua Feature Store hoặc Processing Containers (phiên bản mới nhất tích hợp Canvas cho no-code preprocessing).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Remove the outliers and perform a log transformation on the Square Meters variable.
Lý do chi tiết 🛠️:
- Remove outliers: Loại bỏ mansion lớn và apartment nhỏ để dataset đại diện cho typical cases, tránh model học nhiễu (noise) và cải thiện accuracy/generalization. Phương pháp phổ biến: IQR (Interquartile Range) hoặc Z-score (>3σ).
- Log transformation trên Square Meters: Biến này thường right-skewed (tail dài bên phải do outliers lớn), log(x+1) giúp normalize phân phối, giảm variance, phù hợp cho algorithms như Linear Regression, XGBoost trong SageMaker. Kết hợp hai bước này meet requirements hoàn hảo.
- Lợi ích trên AWS: Trong SageMaker Processing, dùng Pandas/Scikit-learn script để implement, sau đó train model trên cleaned dataset.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn giữ nguyên text gốc bằng tiếng Anh, với đánh giá đúng/sai và lý do bằng tiếng Việt:
-
✅ [ĐÚNG] Remove the outliers and perform a log transformation on the Square Meters variable.
🟢 Đúng hoàn toàn: Như giải thích trên, remove outliers loại bỏ nhiễu từ mansion/small apartment, log transform xử lý skewness của Square Meters (phù hợp real estate data). Kết quả: Model accurate cho typical houses. Best practice trong SageMaker ML pipelines. -
❌ [SAI] Keep the outliers and perform normalization on the Square Meters variable.
🔴 Sai vì: Giữ outliers vẫn để model học bias từ extremes, normalization (MinMaxScaler/StandardScaler) chỉ scale range [0-1] hoặc mean=0/std=1 nhưng không xử lý skewness/outliers, dẫn đến predictions kém cho typical cases. -
❌ [SAI] Remove the outliers and perform one-hot encoding on the Square Meters variable.
🔴 Sai vì: Remove outliers tốt, nhưng one-hot encoding dành cho categorical variables (như màu sắc), không áp dụng cho Square Meters (continuous numerical). Sẽ tạo curse of dimensionality (hàng nghìn categories), làm model explode và vô nghĩa. -
❌ [SAI] Keep the outliers and perform one-hot encoding on the Square Meters variable.
🔴 Sai kép: Giữ outliers gây bias, one-hot encoding sai loại biến (numerical → categorical giả tạo). Hoàn toàn không meet requirements, lãng phí compute trên SageMaker.
📘 Tài liệu tham khảo
- AWS SageMaker Documentation (2026 update): Preprocessing Data – Hướng dẫn outlier removal và transformations trong Processing Jobs.
- SageMaker Data Wrangler: Outlier Detection & Transformations – No-code tool với log transform và IQR.
- AWS ML Best Practices: Handling Skewed Data – Whitepaper về log transform cho real estate ML.
- Scikit-learn (tích hợp SageMaker):
sklearn.preprocessing.PowerTransformercho log-like transforms.
Hy vọng phân tích này giúp bạn ôn thi AWS hiệu quả! 🚀 Nếu cần code sample SageMaker Processing, hãy hỏi thêm.
How should the ML engineer capture bias metrics to display on the dashboard?
- A Capture AWS CloudTrail metrics from SageMaker Clarify.
- B Capture Amazon CloudWatch metrics from SageMaker Clarify.
- C Capture SageMaker Model Monitor metrics from Amazon EventBridge.
- D Capture SageMaker Model Monitor metrics from Amazon Simple Notification Service (Amazon SNS).
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS
📘 Nội dung câu hỏi được giải thích rõ ràng:
Câu hỏi xoay quanh việc phát triển mô hình Machine Learning (ML) trên Amazon SageMaker AI. Công ty cần giám sát bias (thiên kiến) trong mô hình và hiển thị kết quả trên dashboard. Một ML engineer đã tạo bias monitoring job (công việc giám sát bias). Vấn đề cốt lõi là: Làm thế nào để capture (thu thập) các metrics về bias để hiển thị trên dashboard?
🛠️ Bối cảnh kỹ thuật (cập nhật đến 2026):
- SageMaker Clarify là tính năng chuyên dụng của SageMaker để phát hiện bias, fairness và explainability trong mô hình ML (theo AWS re:Invent 2024 và docs mới nhất). Bias monitoring job trong Clarify tự động thu thập metrics như pre-training bias, post-training bias (ví dụ: ClassImbalance, DemographicParity).
- Các metrics này được publish trực tiếp đến Amazon CloudWatch để monitoring và visualize trên dashboard (CloudWatch dashboards hoặc SageMaker Studio).
- Không liên quan đến SageMaker Model Monitor (dùng cho data drift, model quality), CloudTrail (audit logs), EventBridge/SNS (notifications).
✅ Đáp án đúng: Capture Amazon CloudWatch metrics from SageMaker Clarify.
Lý do lựa chọn (chi tiết):
Khi chạy bias monitoring job trong SageMaker Clarify, AWS tự động emit (gửi) các CloudWatch metrics cụ thể như BiasMetric.PreTrainingBias.CIP, BiasMetric.PostTrainingBias.ClassImbalance... ML engineer chỉ cần query CloudWatch để capture metrics và build dashboard (sử dụng CloudWatch dashboards hoặc Grafana integration). Đây là cách chuẩn và hiệu quả nhất theo best practices AWS (không cần code custom, real-time monitoring).
🔍 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên docs AWS SageMaker Clarify (2026 updates).
-
❌ Capture AWS CloudTrail metrics from SageMaker Clarify.
Phân tích sai: CloudTrail chỉ ghi audit logs (API calls, events) của SageMaker Clarify, không capture metrics số lượng như bias scores. CloudTrail không dùng để hiển thị dashboard metrics thời gian thực (chỉ logs text-based). Không phù hợp cho monitoring bias. -
✅ Capture Amazon CloudWatch metrics from SageMaker Clarify.
Phân tích đúng: Đây là phương pháp chính thức. Clarify bias jobs publish metrics trực tiếp đến CloudWatch Metrics namespaceAWS/SageMaker/Clarify(dimensions: JobName, MetricGroup). Dễ dàng visualize trên CloudWatch dashboard hoặc SageMaker Model Dashboard. Hỗ trợ alarms, retention dài hạn. -
❌ Capture SageMaker Model Monitor metrics from Amazon EventBridge.
Phân tích sai: SageMaker Model Monitor dùng cho model quality/drift, không phải bias (bias thuộc Clarify). EventBridge chỉ route events/notifications (như job completion), không capture metrics để dashboard. Không emit bias metrics. -
❌ Capture SageMaker Model Monitor metrics from Amazon Simple Notification Service (Amazon SNS).
Phân tích sai: Tương tự trên, Model Monitor không xử lý bias. SNS chỉ gửi notifications (email/SMS) dựa trên events, không lưu trữ hay capture metrics để hiển thị dashboard. Không có integration trực tiếp cho bias metrics.
📚 Tài liệu tham khảo (AWS chính thức - cập nhật 2026)
- SageMaker Clarify Bias Monitoring: docs.aws.amazon.com/sagemaker/latest/dg/clarify-detect-bias.html – Chi tiết CloudWatch metrics.
- CloudWatch Integration: docs.aws.amazon.com/sagemaker/latest/dg/model-monitor-cloudwatch.html (mở rộng cho Clarify).
- Best Practices: AWS Well-Architected ML Lens (2025): Nhấn mạnh Clarify + CloudWatch cho bias dashboard.
- Demo Code: SageMaker Examples repo trên GitHub:
sagemaker-clarify-bias-monitor-cloudwatch.
Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ code, hỏi nhé!
What should the ML engineer do to resolve this issue?
- A Retrain the model by using a new SageMaker AI training job. Check for errors by using SageMaker Debugger.
- B Retrain the model with new training data. Reuse the original baseline in Model Monitor.
- C Retrain the model with new training data. Use the new baseline in Model Monitor.
- D Rerun the SageMaker AI pipeline after enabling the emit_metrics option in the baseline constraints file.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh vấn đề SageMaker Model Monitor trong AWS SageMaker – một công cụ giám sát chất lượng mô hình ML tự động.
- Bối cảnh: Công ty sử dụng mô hình ML trên SageMaker để dự đoán tai nạn giao thông do ổ gà gây ra. ML engineer đã cấu hình Model Monitor chạy như một phần của SageMaker Pipeline (pipeline tự động hóa quy trình ML).
- Vấn đề cụ thể: Trong output của MonitoringExecution (quá trình thực thi giám sát), phát hiện nhiều vi phạm baseline_drift_check (kiểm tra sự lệch drift so với baseline ban đầu), dẫn đến pipeline bị fail (thất bại).
- Mục tiêu: ML engineer cần hành động để resolve (giải quyết) vấn đề này, đảm bảo pipeline tiếp tục chạy mà không fail do drift.
Lý do vấn đề xảy ra 🛠️: Data drift (sự thay đổi dữ liệu đầu vào theo thời gian) làm dữ liệu mới lệch khỏi baseline (dữ liệu tham chiếu ban đầu được tạo từ training job đầu tiên). Model Monitor so sánh dữ liệu inference thời gian thực với baseline, nếu drift quá ngưỡng thì báo violation và fail pipeline.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Retrain the model with new training data. Use the new baseline in Model Monitor.
Lý do chi tiết 🎯:
- Khi phát hiện baseline_drift_check violations, điều này cho thấy dữ liệu mới đã thay đổi đáng kể so với baseline cũ (do data drift thực tế như thay đổi mẫu tai nạn giao thông).
- Giải pháp chuẩn AWS: Retrain mô hình với dữ liệu training mới (đại diện cho dữ liệu hiện tại), sau đó tạo và sử dụng baseline mới từ training job này cho Model Monitor.
- Baseline mới sẽ được capture từ dữ liệu training/inference gần nhất, giúp Model Monitor so sánh chính xác hơn, tránh false positive violations.
- Điều này tuân thủ best practice của SageMaker Pipeline: Tự động hóa retrain và update baseline để duy trì model quality. Không cần can thiệp thủ công, pipeline sẽ pass sau update.
📋 Phân tích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên tài liệu AWS SageMaker mới nhất (2024-2026), nơi Model Monitor hỗ trợ drift detection nâng cao với custom baselines và tích hợp chặt chẽ với Pipelines.
-
Retrain the model by using a new SageMaker AI training job. Check for errors by using SageMaker Debugger.
❌ Sai vì: SageMaker Debugger dùng để debug training process (lỗi code, hyperparameters, tensor anomalies), không liên quan đến drift violations ở Model Monitor (là vấn đề runtime data quality). Retrain thôi chưa đủ, phải update baseline mới; Debugger chỉ làm phức tạp hóa mà không giải quyết gốc rễ drift. -
Retrain the model with new training data. Reuse the original baseline in Model Monitor.
❌ Sai vì: Reuse baseline cũ sẽ tiếp tục detect drift (vi phạm vẫn xảy ra vì dữ liệu mới đã thay đổi). AWS khuyến cáo không reuse baseline cũ khi data drift rõ rệt; phải tạo baseline mới từ retrain để reflect data distribution hiện tại, tránh pipeline fail liên tục. -
Retrain the model with new training data. Use the new baseline in Model Monitor.
✅ Đúng như đã giải thích ở trên. Đây là cách chính thức để remediate drift: Retrain → Capture new baseline (tự động quamodel-monitorstep trong Pipeline) → Update constraints. Pipeline sẽ pass sau đó. -
Rerun the SageMaker AI pipeline after enabling the emit_metrics option in the baseline constraints file.
❌ Sai vì:emit_metricschỉ publish metrics đến CloudWatch (cho alerting/visualization), không fix drift violations. Baseline constraints định nghĩa ngưỡng drift (e.g., PS/KS stats), nhưng enabling emit_metrics không thay đổi logic check – violations vẫn fail pipeline khi rerun.
📘 Tài liệu tham khảo (AWS cập nhật 2024-2026)
- AWS SageMaker Model Monitor Documentation: Monitor model quality – Chi tiết về baseline_drift_check và remediation (retrain + new baseline).
- SageMaker Pipelines: Model monitoring in pipelines – Hướng dẫn xử lý violations qua QualityCheck steps.
- Best Practices: Handle data drift – Blog AWS nhấn mạnh update baseline sau retrain.
- Release Notes 2024: SageMaker hỗ trợ automated baselines trong Pipelines v2, không thay đổi logic drift handling cơ bản.
Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ code Pipeline, hãy hỏi nhé!
Which evaluation metric should the company use to evaluate the models to meet this requirement?
- A F1 score
- B Area Under the ROC Curve (AUC)
- C Precision
- D Recall
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc đánh giá mô hình học máy (ML models) được sử dụng để dự đoán giao dịch gian lận (fraudulent transactions) trên nền tảng AWS, cụ thể là các dịch vụ như Amazon SageMaker (phiên bản cập nhật mới nhất đến 2026 hỗ trợ tích hợp Model Monitor và Clarify cho bias detection trong fraud detection).
Công ty muốn xác định càng nhiều giao dịch gian lận càng tốt, nghĩa là ưu tiên giảm thiểu số lượng giao dịch gian lận bị bỏ lỡ (false negatives - FN). Đây là tình huống điển hình của dữ liệu không cân bằng (imbalanced dataset) trong fraud detection, nơi giao dịch gian lận rất hiếm (thường <1%). Do đó, metric đánh giá cần tập trung vào khả năng phát hiện toàn bộ các trường hợp dương tính thực tế (true positives), ngay cả khi chấp nhận một số báo động giả (false positives).
📘 Tài liệu tham khảo:
- AWS SageMaker Documentation: "Evaluating Binary Classification Models" (https://docs.aws.amazon.com/sagemaker/latest/dg/model-monitor-interpreting-results.html) – Nhấn mạnh Recall cho use case fraud detection.
- AWS ML Best Practices (2024-2026 updates): SageMaker Canvas và Autopilot hỗ trợ metrics như Recall cho imbalanced classification.
✅ Đáp án đúng: Recall
Lý do lựa chọn:
Recall (hay True Positive Rate - TPR) = TP / (TP + FN), đo lường tỷ lệ giao dịch gian lận thực tế được mô hình phát hiện chính xác. Trong fraud detection, ưu tiên tối đa hóa việc "bắt" hết gian lận (giảm FN), ngay cả khi có FP cao (báo động giả, có thể kiểm tra thủ công). Điều này phù hợp hoàn hảo với yêu cầu "identify as many fraudulent transactions as possible". AWS khuyến nghị Recall cho các use case tương tự trong SageMaker Processing Jobs và Model Monitor (cập nhật 2026 với tích hợp generative AI metrics).
🛠️ Ví dụ thực tế trên AWS: Sử dụng SageMaker Debugger để track Recall threshold >0.95 cho fraud models.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn, với giữ nguyên văn bản gốc bằng tiếng Anh và đánh giá đúng/sai dựa trên yêu cầu câu hỏi:
-
F1 score ❌ SAI:
F1 score là trung bình điều hòa (harmonic mean) giữa Precision và Recall, phù hợp cho trường hợp cần cân bằng cả hai (ví dụ: khi FP và FN đều costly). Tuy nhiên, ở đây ưu tiên chỉ Recall cao, không cần cân bằng Precision, nên F1 không tối ưu. AWS SageMaker dùng F1 cho balanced datasets, không phải fraud imbalanced. -
Area Under the ROC Curve (AUC) ❌ SAI:
AUC đo lường khả năng phân biệt lớp tổng thể qua các ngưỡng khác nhau (trade-off TPR/FPR). Nó tốt cho đánh giá tổng quát nhưng không tập trung cụ thể vào việc tối đa hóa phát hiện fraud (có thể ưu tiên Precision ở một số ngưỡng). AWS Model Monitor hỗ trợ AUC, nhưng docs khuyến nghị Recall riêng cho "high recall use cases" như fraud. -
Precision ❌ SAI:
Precision = TP / (TP + FP), đo tỷ lệ dự đoán dương tính đúng. Nó ưu tiên giảm báo động giả (ít FP), phù hợp khi FP costly (ví dụ: khóa tài khoản nhầm). Nhưng câu hỏi yêu cầu bắt hết fraud, chấp nhận FP, nên Precision không phù hợp. Trong SageMaker, Precision dùng cho low-FP scenarios, không phải high-recall fraud. -
Recall ✅ ĐÚNG:
Như đã giải thích ở trên, Recall trực tiếp đáp ứng yêu cầu bằng cách tối ưu hóa TP, giảm FN. AWS cập nhật 2026 (SageMaker JumpStart models) tích hợp Recall-based optimization cho fraud pipelines với Amazon Fraud Detector (hybrid ML rules).
🧠 Lưu ý bổ sung: Trong AWS DevOps pipeline, tích hợp metric này vào CI/CD với SageMaker Pipelines và CodePipeline để tự động retrain models nếu Recall < threshold. Nếu cần code sample, có thể dùng sklearn.metrics.recall_score trong SageMaker Processing.
Which solution will meet these requirements?
- A Increase the number of training iterations. Retrain the model.
- B Apply L1 regularization to the training data. Retrain the model.
- C Decrease the number of training iterations. Retrain the model.
- D Use SageMaker Debugger to apply L1 regularization to the running model.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi
Câu hỏi tập trung vào vấn đề phổ biến trong Machine Learning trên Amazon SageMaker AI: Một kỹ sư ML đã huấn luyện mô hình nhưng gặp tình trạng overfitting (mô hình học quá tốt trên dữ liệu huấn luyện, dẫn đến hiệu suất kém trên dữ liệu mới) và dữ liệu huấn luyện chứa nhiều đặc trưng không cần thiết (unnecessary features), làm tăng độ phức tạp mô hình. Nhiệm vụ là tìm giải pháp giảm overfitting đồng thời giảm tác động của các đặc trưng thừa.
✅ Giải pháp cần đáp ứng: Phải xử lý cả hai vấn đề một cách hiệu quả, tận dụng các tính năng của SageMaker (như built-in algorithms hoặc custom training scripts hỗ trợ regularization). Theo tài liệu AWS SageMaker mới nhất (cập nhật đến 2026), regularization là kỹ thuật chuẩn để chống overfitting và thực hiện feature selection tự động.
📘 Tài liệu tham khảo:
- AWS SageMaker Documentation: Hyperparameter Tuning for Built-in Algorithms (hỗ trợ L1/L2 regularization).
- Preventing Overfitting in SageMaker.
- SageMaker Examples: Regularization with XGBoost.
✅ Đáp án đúng: Apply L1 regularization to the training data. Retrain the model.
Lý do lựa chọn 🛠️:
- L1 regularization (Lasso) thêm penalty dựa trên giá trị tuyệt đối của các trọng số mô hình vào hàm loss, đẩy nhiều trọng số về 0, tự động loại bỏ (hoặc giảm mạnh) tác động của unnecessary features (feature selection tự động).
- Đồng thời, nó làm mịn mô hình, giảm overfitting bằng cách hạn chế độ phức tạp.
- Trong SageMaker, bạn có thể áp dụng L1 qua hyperparameters (ví dụ:
alphacho XGBoost hoặc custom script với TensorFlow/Keras:kernel_regularizer=l1(0.01)), sau đó retrain job. Đây là giải pháp trực tiếp, hiệu quả nhất theo best practices AWS 2026.
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] Increase the number of training iterations. Retrain the model.
Giải thích: Tăng số lượng iterations (epochs) sẽ khiến mô hình học sâu hơn trên dữ liệu huấn luyện, làm tăng overfitting thay vì giảm (vì mô hình ghi nhớ noise và unnecessary features nhiều hơn). Không giải quyết được đặc trưng thừa, chỉ làm vấn đề tệ hơn. SageMaker Training Jobs hỗ trợ epochs hyperparam, nhưng tăng nó không phải cách chống overfitting chuẩn. -
✅ [ĐÚNG] Apply L1 regularization to the training data. Retrain the model.
Giải thích: Như đã phân tích ở trên, L1 regularization lý tưởng cho cả hai vấn đề: feature selection (giảm unnecessary features) và regularization chống overfitting. Áp dụng vào training data qua script hoặc built-in algos (như Linear Learner:l1_reg=0.01), rồi retrain – hoàn hảo cho SageMaker. -
❌ [SAI] Decrease the number of training iterations. Retrain the model.
Giải thích: Giảm iterations có thể giúp early stopping để tránh overfitting (dừng sớm trước khi học noise), nhưng không xử lý unnecessary features (mô hình vẫn dùng hết features, chỉ học ít hơn). Đây là giải pháp tạm thời, không triệt để; SageMaker khuyến nghị kết hợp với regularization thay vì chỉ giảm epochs. -
❌ [SAI] Use SageMaker Debugger to apply L1 regularization to the running model.
Giải thích: SageMaker Debugger chỉ dùng để monitor và debug training process (theo dõi metrics như loss, gradients realtime), không hỗ trợ apply regularization trực tiếp vào running model. L1 phải set trước trong training script/hyperparameters, không phải "apply on-the-fly". Sai lầm phổ biến – Debugger là công cụ phân tích (tích hợp TensorBoard), không thay đổi model architecture.
🛠️ Khuyến nghị thực hành: Sử dụng SageMaker Hyperparameter Tuning Jobs kết hợp L1 với early stopping để tối ưu. Test trên validation set để đo overfitting (gap giữa train/validation accuracy). Nếu custom model, tích hợp vào estimator.fit()!
Which solution will meet these requirements?
- A Perform ordinal encoding to represent categories of the feature.
- B Perform similarity encoding to represent categories of the feature.
- C Perform one-hot encoding to represent categories of the feature.
- D Perform target encoding to represent categories of the feature.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào quy trình preprocessing dữ liệu trong Amazon SageMaker Data Wrangler – một công cụ trực quan của AWS SageMaker giúp xây dựng pipeline xử lý dữ liệu cho machine learning (ML). Cụ thể:
- Một ML engineer đang sử dụng SageMaker Data Wrangler để preprocess dataset trước khi train một classification model (mô hình phân loại).
- Vấn đề chính: Một text feature (cột đặc trưng văn bản) có hàng nghìn giá trị khác nhau, nhưng chúng chỉ khác biệt do lỗi chính tả (spelling errors), dẫn đến high cardinality (số lượng category quá lớn, gây khó khăn cho encoding thông thường).
- Yêu cầu: Áp dụng encoding method phù hợp để sau preprocessing, feature này có thể được sử dụng hiệu quả để train model, mà không bị ảnh hưởng bởi noise từ lỗi chính tả. Phương pháp phải giảm số lượng category tương tự bằng cách nhóm chúng dựa trên semantic similarity (độ tương đồng ý nghĩa).
SageMaker Data Wrangler (cập nhật đến phiên bản mới nhất năm 2026) hỗ trợ các built-in transforms như encoding cho categorical/text features, đặc biệt xử lý high-cardinality và noisy data. ✅ Mục tiêu là chọn encoding giúp model học được pattern thực sự, tránh curse of dimensionality.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Perform similarity encoding to represent categories of the feature.
🛠️ Lý do chi tiết:
- Similarity encoding là transform chuyên biệt trong SageMaker Data Wrangler (dùng sentence transformers hoặc embedding models như BERT-based) để nhóm các category text tương tự dựa trên semantic similarity (độ giống về ý nghĩa).
- Với hàng nghìn values chỉ khác spelling errors, nó sẽ tự động cluster/group chúng thành ít category hơn (ví dụ: "NewYork", "NewYorkk", "new york" → cùng một embedding vector đại diện).
- Kết quả: Feature trở thành low-dimensional embedding (vector dense), dễ dàng dùng cho classification model (như XGBoost, neural nets), giảm overfitting và cải thiện performance.
- Đây là giải pháp tối ưu theo best practices AWS cho noisy text/high-cardinality categorical data trong Data Wrangler flows. 📈 Hiệu quả cao hơn one-hot vì tránh sparse matrices khổng lồ.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn giữ nguyên nội dung gốc bằng tiếng Anh, kèm giải thích sai/đúng bằng tiếng Việt dựa trên tính phù hợp với vấn đề spelling errors + high cardinality trong SageMaker Data Wrangler:
-
❌ Perform ordinal encoding to represent categories of the feature.
🧨 Sai vì: Ordinal encoding giả sử thứ tự tự nhiên giữa categories (gán số 1,2,3...), không phù hợp cho text feature không có thứ tự như tên địa điểm hay lỗi chính tả. Nó sẽ tạo ra noise lớn, làm model học sai pattern, dẫn đến poor accuracy. -
✅ Perform similarity encoding to represent categories of the feature.
🏆 Đúng vì: Như giải thích trên, đây là transform lý tưởng trong Data Wrangler cho noisy text với spelling variants. Nó dùng pre-trained embeddings để map tương tự → cùng representation, giảm cardinality hiệu quả mà giữ semantic. Hoàn hảo cho classification training. -
❌ Perform one-hot encoding to represent categories of the feature.
🚫 Sai vì: One-hot tạo cột riêng cho từng unique value, với hàng nghìn values sẽ sinh ra ma trận sparse cực lớn (curse of dimensionality), tốn memory/CPU và gây overfitting. Không xử lý được spelling errors (mỗi variant vẫn là category riêng). -
❌ Perform target encoding to represent categories of the feature.
⚠️ Sai vì: Target encoding thay category bằng mean target value tương ứng, dễ gây data leakage (leak thông tin target vào features) nếu không cross-validate cẩn thận. Không giải quyết spelling errors (variants vẫn riêng biệt), và kém hiệu quả với high-cardinality noisy text.
📘 Tài liệu tham khảo (cập nhật AWS 2026)
- AWS SageMaker Data Wrangler Documentation: Transforms for categorical data – Chi tiết về Similarity Encoding với ví dụ spelling correction.
- AWS ML Best Practices: Processing High-Cardinality Categorical Features – Hướng dẫn dùng embedding cho noisy text.
- SageMaker Updates 2025-2026: Tích hợp sâu hơn với Sentence Transformers và zero-shot similarity trong Data Wrangler v2.x.
Hy vọng phân tích này giúp bạn ôn thi AWS Certified DevOps Engineer Professional hiệu quả! 🚀 Nếu cần thêm ví dụ code hoặc flow Data Wrangler, hãy hỏi nhé!
The ML engineer needs to ensure that the model achieves generalization.
Which solution will meet this requirement?
- A Increase the learning rate and decrease the mini-batch size.
- B Increase the learning rate as the number of epochs increases.
- C Decrease the learning rate and increase the mini-batch size.
- D Decrease the learning rate and decrease the mini-batch size.
Xem giải thích
🧠 Phân tích câu hỏi trắc nghiệm AWS SageMaker AI
📖 Nội dung câu hỏi được giải thích chi tiết:
Câu hỏi mô tả tình huống một kỹ sư ML đang huấn luyện mô hình text generation (tạo văn bản) trên Amazon SageMaker AI. Sau vài epoch (vòng lặp huấn luyện), hàm loss không hội tụ (không giảm dần ổn định) và độ chính xác trên tập validation bắt đầu dao động (oscillating results). Vấn đề này thường chỉ ra mô hình đang gặp khó khăn trong việc generalization (khái quát hóa), có thể do gradient không ổn định, overfitting nhẹ hoặc underfitting do noise từ mini-batch. Mục tiêu là chọn giải pháp giúp mô hình học ổn định hơn, giảm dao động và cải thiện khả năng khái quát hóa trên dữ liệu mới.
🛠️ Đây là vấn đề phổ biến trong training deep learning trên SageMaker, liên quan đến hyperparameter tuning như learning rate (tốc độ học) và mini-batch size (kích thước lô dữ liệu).
✅ Đáp án đúng:
Decrease the learning rate and increase the mini-batch size.
Lý do lựa chọn (chi tiết):
- Giảm learning rate (LR): LR quá cao gây ra dao động loss và accuracy vì cập nhật weight quá mạnh, dẫn đến "nhảy qua" điểm tối ưu. Giảm LR giúp gradient descent mượt mà hơn, loss hội tụ ổn định, giảm oscillate – đặc biệt hiệu quả cho text generation models như Transformer trên SageMaker (theo best practices 2024-2026).
- Tăng mini-batch size: Batch nhỏ tạo gradient noisy (biến động cao), gây oscillate. Batch lớn hơn làm trung bình gradient ổn định, giảm variance, cải thiện generalization mà không cần thay đổi architecture.
Kết hợp hai thay đổi này là giải pháp chuẩn để fix non-convergence và oscillation, giúp model generalize tốt trên validation set. SageMaker hỗ trợ tuning tự động qua Hyperparameter Tuning Jobs hoặc SageMaker Automatic Model Tuning (phiên bản mới nhất 2026).
🔍 Giải thích tất cả các phương án (đúng/sai):
-
❌ [SAI] Increase the learning rate and decrease the mini-batch size.
Tăng LR làm gradient update mạnh hơn, tăng dao động loss/accuracy (worse oscillation). Giảm batch size tăng noise gradient, làm vấn đề tồi tệ hơn – hoàn toàn ngược với nhu cầu stabilize và generalize. -
❌ [SAI] Increase the learning rate as the number of epochs increases.
Tăng LR theo epoch (learning rate scheduling tăng dần) rất hiếm dùng, thường chỉ cho warm-up phase ngắn; ở đây loss đã không converge, tăng LR sẽ đẩy xa minima, tăng oscillate – không giúp generalization. -
✅ [ĐÚNG] Decrease the learning rate and increase the mini-batch size.
Như giải thích trên: Giảm LR ổn định descent, tăng batch giảm noise – combo lý tưởng cho convergence và generalization trên SageMaker (xác nhận từ best practices). -
❌ [SAI] Decrease the learning rate and decrease the mini-batch size.
Giảm LR tốt cho ổn định, nhưng giảm batch size tăng noise (stochastic gradient variance cao), vẫn gây oscillate dù chậm hơn – không giải quyết gốc rễ generalization.
📘 Tài liệu tham khảo (cập nhật AWS 2026):
- Amazon SageMaker Developer Guide: "Troubleshoot training jobs" & "Hyperparameter tuning best practices" (https://docs.aws.amazon.com/sagemaker/latest/dg/train-debugger.html).
- AWS ML Best Practices: "Optimize deep learning hyperparameters" (SageMaker Automatic Tuning với LR/batch size – cập nhật Q1/2026).
- Blog AWS: "Stabilizing training in SageMaker for LLMs" (2025 series on text generation models).
🧩 Mẹo DevOps: Sử dụng SageMaker Debugger để monitor loss/accuracy real-time và Profiler để detect oscillation sớm!
The company needs a hybrid system to make the shared data store accessible to on-premises servers and Amazon SageMaker AI notebooks that will consume the data. File locking is required for the data producers.
Which AWS storage solution will meet these requirements?
- A Use an Amazon S3 bucket to store the data. Use Mountpoint for Amazon S3 to mount the S3 bucket to the on-premises servers and the SageMaker AI notebooks.
- B Use an Amazon Elastic File System (Amazon EFS) file system to store the data. Mount the file system to the on-premises servers and the SageMaker AI notebooks.
- C Use an Amazon FSx for Lustre file system to store the data. Mount the file system to the on-premises servers and the SageMaker AI notebooks.
- D Use an Amazon Elastic Block Store (Amazon EBS) volume to store the data. Mount the volume to the on-premises servers and the SageMaker AI notebooks.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một tình huống thực tế trong môi trường AWS liên quan đến lưu trữ dữ liệu cho huấn luyện ML (Machine Learning):
- Công ty đang sử dụng NFS-based data store (hệ thống lưu trữ dựa trên NFS) để lưu dữ liệu, và các hệ thống Linux-based truy cập vào đó.
- Yêu cầu xây dựng hệ thống hybrid (kết hợp on-premises và cloud), cho phép shared data store (lưu trữ chia sẻ) có thể truy cập từ:
- On-premises servers (máy chủ tại chỗ).
- Amazon SageMaker AI notebooks (các notebook AI trên SageMaker).
- Yêu cầu quan trọng nhất: File locking (khóa file) bắt buộc dành cho data producers (các bên sản xuất dữ liệu), để tránh xung đột khi nhiều bên ghi/đọc đồng thời.
📘 Tóm tắt vấn đề cốt lõi: Cần giải pháp lưu trữ NFS-compatible, hỗ trợ hybrid access (qua Direct Connect/VPN cho on-premises), multi-mount (nhiều máy mount cùng lúc), và file locking chuẩn NFSv4. Đây là yêu cầu phổ biến trong DevOps cho ML workflows hybrid đến năm 2026 (AWS EFS đã hỗ trợ NFSv4.1/4.2 đầy đủ).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use an Amazon Elastic File System (Amazon EFS) file system to store the data. Mount the file system to the on-premises servers and the SageMaker AI notebooks.
🛠️ Lý do chi tiết:
- Amazon EFS là dịch vụ file storage managed NFS (NFSv4.1), hoàn hảo thay thế NFS-based data store hiện tại.
- Hỗ trợ hybrid connectivity: Mount từ on-premises qua AWS Direct Connect hoặc Site-to-Site VPN, và từ SageMaker notebooks (SageMaker hỗ trợ EFS mounting native từ 2023+).
- File locking đầy đủ: EFS hỗ trợ NFSv4 file locking (mandatory cho data producers), tránh race conditions trong ML training.
- Scalability: Tự động scale cho ML workloads lớn, multi-AZ, POSIX-compliant. Phù hợp phiên bản AWS 2026 với EFS IA (Infrequent Access) cho cost-optimize.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt với emoji nổi bật:
-
✅ Use an Amazon Elastic File System (Amazon EFS) file system to store the data. Mount the file system to the on-premises servers and the SageMaker AI notebooks.
🟢 Đúng vì: Như đã giải thích ở trên. EFS là lựa chọn lý tưởng cho NFS hybrid với file locking chuẩn. SageMaker notebooks mount EFS seamless qua VPC. -
❌ Use an Amazon S3 bucket to store the data. Use Mountpoint for Amazon S3 to mount the S3 bucket to the on-premises servers and the SageMaker AI notebooks.
🔴 Sai vì: S3 là object storage, không phải file system thực thụ. Mountpoint for S3 (ra mắt 2023, open-source) chỉ hỗ trợ read-heavy workloads, KHÔNG hỗ trợ file locking (S3 atomic operations không tương đương NFS locking). Data producers sẽ gặp xung đột. Hybrid access có thể qua VPC endpoints, nhưng không meet yêu cầu locking. -
❌ Use an Amazon FSx for Lustre file system to store the data. Mount the file system to the on-premises servers and the SageMaker AI notebooks.
🔴 Sai vì: FSx for Lustre là parallel distributed file system (dựa Lustre protocol), KHÔNG phải NFS chuẩn (chỉ mount qua Lustre clients đặc biệt). Hybrid với on-premises NFS khó khăn, yêu cầu custom networking phức tạp. File locking Lustre mạnh nhưng không POSIX/NFSv4 compatible trực tiếp cho SageMaker notebooks thông thường. Phù hợp HPC hơn ML hybrid. -
❌ Use an Amazon Elastic Block Store (Amazon EBS) volume to store the data. Mount the volume to the on-premises servers and the SageMaker AI notebooks.
🔴 Sai vì: EBS là block storage (như ổ cứng ảo), chỉ attach vào MỘT EC2 instance duy nhất (single-AZ hoặc multi-attach giới hạn). KHÔNG thể share/mount với on-premises servers hoặc nhiều SageMaker notebooks cùng lúc. Không hỗ trợ hybrid access và file locking multi-client.
📚 Tài liệu tham khảo (AWS cập nhật đến 2026)
- Amazon EFS Documentation: EFS Features - NFS Locking & Hybrid Access (xác nhận NFSv4.1 locking, VPC/SageMaker integration).
- SageMaker EFS Mounting: AWS Blog - SageMaker with EFS (hướng dẫn mount hybrid).
- Mountpoint for S3 Limitations: AWS Announcement - No Locking (rõ ràng không hỗ trợ write locking).
- FSx for Lustre: AWS FSx Docs - Protocol (Lustre-only, không NFS).
- DevOps Pro Exam Guide 2026: AWS Certified DevOps Engineer Professional blueprint nhấn mạnh EFS cho shared file systems hybrid ML (domain EKS/SageMaker storage).
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm case study, hỏi nhé!
The company needs a no-code solution to transform the data. The solution must store the transformed data back to the same S3 bucket for model training.
Which solution will meet these requirements?
- A Configure an AWS Glue DataBrew project that connects to the data. Use the DataBrew interactive interface to create a recipe that performs the one-hot encoding transformation. Create a job to apply the transformation and to write the output back to an S3 bucket.
- B Configure an AWS Glue Data Catalog table that points to the data. Use Amazon Athena to write SQL commands to perform the one-hot encoding transformation. Configure Athena to write the query results back to an S3 bucket.
- C Configure an AWS Glue Data Catalog table that points to the data. Create an AWS Glue ETL interactive notebook. Use the notebook to perform the one-hot encoding transformation. Run the configured cells and write the results back to an S3 bucket.
- D Configure an Amazon Redshift cluster to access the data by using Redshift Spectrum. Use SQL commands to perform the one-hot encoding transformation within Amazon Redshift. Configure Amazon Redshift to write the results back to an S3 bucket.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào yêu cầu xử lý dữ liệu lớn lưu trữ trong Amazon S3 dưới định dạng Apache Parquet. Công ty cần áp dụng one-hot encoding (một kỹ thuật biến đổi dữ liệu phân loại thành dạng số 0/1 để phù hợp với machine learning) cho một số cột cụ thể.
Yêu cầu chính:
- No-code solution (giải pháp không cần viết code, tức là giao diện kéo-thả hoặc visual, dễ sử dụng cho người không phải lập trình viên).
- Biến đổi dữ liệu và lưu kết quả trở lại cùng S3 bucket để phục vụ model training (huấn luyện mô hình ML).
✅ Mục tiêu: Tìm giải pháp AWS no-code, hỗ trợ Parquet/S3, one-hot encoding, và output về S3 mà không cần code thủ công. Dựa trên kiến thức AWS cập nhật đến 2026 (AWS Glue DataBrew phiên bản mới nhất hỗ trợ đầy đủ transformations visual cho ML prep).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Configure an AWS Glue DataBrew project that connects to the data. Use the DataBrew interactive interface to create a recipe that performs the one-hot encoding transformation. Create a job to apply the transformation and to write the output back to an S3 bucket.
Lý do chi tiết:
- 🛠️ AWS Glue DataBrew là dịch vụ no-code visual data preparation (chuẩn bị dữ liệu trực quan không code) của AWS, chuyên xử lý dataset lớn từ S3 (hỗ trợ Parquet native).
- 📱 Giao diện interactive cho phép tạo recipe (công thức biến đổi) bằng kéo-thả, bao gồm one-hot encoding sẵn có (transform "Encode" > One-hot).
- 🚀 Tạo job để chạy recipe scale lớn, output trực tiếp về S3 bucket cùng vị trí, tối ưu cho ML training (tích hợp SageMaker).
- 💡 Hoàn hảo match "no-code" và lưu S3, không cần code/SQL/notebook.
❌ Giải thích tất cả các phương án (đúng/sai)
-
✅ Phương án ĐÚNG (như trên):
Configure an AWS Glue DataBrew project that connects to the data. Use the DataBrew interactive interface to create a recipe that performs the one-hot encoding transformation. Create a job to apply the transformation and to write the output back to an S3 bucket.
Giải thích: Đây là giải pháp no-code thuần túy với giao diện visual, hỗ trợ one-hot encoding trực tiếp, scale job tự động, và output S3 seamless. Phù hợp 100% yêu cầu. (Cập nhật 2026: DataBrew hỗ trợ >300 transforms, bao gồm ML-specific như one-hot). -
❌ Phương án SAI 1:
Configure an AWS Glue Data Catalog table that points to the data. Use Amazon Athena to write SQL commands to perform the one-hot encoding transformation. Configure Athena to write the query results back to S3 bucket.
Giải thích: Amazon Athena yêu cầu viết SQL commands (code-based), không phải no-code. One-hot encoding phức tạp trong SQL (cần UNPIVOT/MULTIPLE CASE), không visual. Output S3 OK nhưng vi phạm "no-code". -
❌ Phương án SAI 2:
Configure an AWS Glue Data Catalog table that points to the data. Create an AWS Glue ETL interactive notebook. Use the notebook to perform the one-hot encoding transformation. Run the configured cells and write the results back to an S3 bucket.
Giải thích: AWS Glue ETL interactive notebook là code-based (PySpark/Scala notebook), yêu cầu viết code để implement one-hot (sử dụng pandas/spark.ml). Không phải no-code, dù output S3 dễ dàng. -
❌ Phương án SAI 3:
Configure an Amazon Redshift cluster to access the data by using Redshift Spectrum. Use SQL commands to perform the one-hot encoding transformation within Amazon Redshift. Configure Amazon Redshift to write the results back to an S3 bucket.
Giải thích: Redshift Spectrum + Redshift cluster yêu cầu SQL commands (code), one-hot phức tạp (cần stored proc hoặc multiple queries). Cần provision cluster (chi phí cao, không serverless/no-code), dù output S3 qua UNLOAD OK.
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- AWS Glue DataBrew: docs.aws.amazon.com/databrew – Transformations guide (one-hot encoding).
- AWS Exam DOP-C02: Data preparation no-code (DataBrew vs. Glue ETL).
- AWS re:Post & Well-Architected ML Lens: Khuyến nghị DataBrew cho visual prep trước SageMaker training.
🧠 Kết luận: DataBrew là lựa chọn tối ưu cho no-code ML data prep trên S3! 🚀
Which feature of SageMaker AI should the company use to meet these requirements?
- A SageMaker AI built-in algorithms
- B SageMaker Canvas
- C SageMaker JumpStart
- D SageMaker AI script mode
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc migrate các mô hình Machine Learning (ML) từ môi trường on-premises sang Amazon SageMaker AI, với các mô hình dựa trên framework PyTorch. Yêu cầu chính là tái sử dụng tối đa các custom scripts hiện có của công ty trên AWS.
- Bối cảnh: SageMaker là dịch vụ managed ML của AWS, hỗ trợ toàn bộ lifecycle từ training đến deployment. Khi migrate từ on-premises, công ty không muốn viết lại code mà muốn giữ nguyên logic custom (ví dụ: data processing, model training scripts trong PyTorch).
- Thách thức chính: Cần một tính năng cho phép bring your own code/scripts (BYOC) mà vẫn tận dụng infrastructure managed của SageMaker như distributed training, hyperparameter tuning.
- Phiên bản cập nhật (2026): SageMaker hỗ trợ PyTorch qua SageMaker Python SDK và Script Mode (còn gọi là Framework Mode), cho phép wrap custom PyTorch scripts mà không cần thay đổi nhiều (chỉ cần định nghĩa entry_point và source_dir).
📘 Tài liệu tham khảo chính:
- AWS SageMaker Documentation: Train with PyTorch (cập nhật 2025-2026).
- Bring Your Own Algorithm.
✅ Đáp án đúng: SageMaker AI script mode
Lý do lựa chọn:
- Script Mode (hay Framework Mode) trong SageMaker cho phép tái sử dụng trực tiếp custom PyTorch scripts từ on-premises. Bạn chỉ cần cung cấp file entry_point (ví dụ:
train.py) và thư mục source chứa các scripts phụ, SageMaker sẽ tự động wrap chúng vào Docker container chuẩn (sử dụngsagemaker.pytorch.PyTorchestimator). - ✅ Ưu điểm nổi bật: Hỗ trợ minimal code changes – chỉ cần thêm hooks như
model_fn,input_fnnếu cần; tự động xử lý data channels, checkpointing, và scaling. Hoàn hảo cho migration PyTorch models. - 🛠️ Ví dụ code đơn giản (Python SDK):
from sagemaker.pytorch import PyTorch pytorch_estimator = PyTorch( entry_point='train.py', # Custom script từ on-premises source_dir='code/', # Thư mục chứa scripts role=role, instance_count=1, instance_type='ml.m5.large', framework_version='2.1', py_version='py310' ) - Đây là giải pháp official best practice cho custom frameworks như PyTorch (không phải built-in algo).
📋 Giải thích chi tiết TẤT CẢ các phương án
-
SageMaker AI built-in algorithms ❌
Sai vì: Đây là các thuật toán sẵn có của AWS (như XGBoost, Linear Learner, không có PyTorch built-in). Chúng yêu cầu sử dụng API chuẩn của AWS, không hỗ trợ custom scripts. Phù hợp cho quick prototyping nhưng không reuse được code on-premises. (Tham khảo: Built-in Algorithms). -
SageMaker Canvas ❌
Sai vì: SageMaker Canvas là công cụ no-code/low-code dành cho non-technical users (business analysts), sử dụng point-and-click để build models từ CSV data. Không hỗ trợ PyTorch custom scripts hay code migration; chỉ dùng autoML nội bộ. (Tham khảo: SageMaker Canvas). -
SageMaker JumpStart ❌
Sai vì: JumpStart cung cấp pre-trained models và solutions sẵn (từ Hugging Face, TensorFlow Hub), dễ deploy 1-click. Tuy hỗ trợ một số PyTorch models nhưng chỉ cho fine-tuning limited, không cho phép full custom scripts migration. Không phải cho reuse toàn bộ code on-premises. (Tham khảo: SageMaker JumpStart, cập nhật 2026 với thêm foundation models). -
SageMaker AI script mode ✅
Đúng vì: Như đã giải thích trên, đây là tính năng lý tưởng cho custom PyTorch scripts, hỗ trợ training/inference với zero-to-minimal refactoring. SageMaker tự handle orchestration, giúp migrate seamless. (Best practice từ AWS re:Invent 2025 sessions về SageMaker Frameworks).
🧠 Lời khuyên DevOps: Khi implement, dùng SageMaker Pipelines để automate migration pipeline, kết hợp Model Registry cho versioning. Test với SageMaker Debugger để monitor custom scripts! 🚀