Ngân hàng đề — AWS Certified Machine Learning Specialty
Tìm thấy 371 câu.
A data science team wants to use Amazon Redshift machine learning (Amazon Redshift ML) to build a model and run predictions for new data directly within the data warehouse.
Which combination of steps should the company take to use Amazon Redshift ML to meet these requirements? (Choose three.)
- A Define the feature variables and target variable for the churn prediction model.
- B Use the SOL EXPLAIN_MODEL function to run predictions.
- C Write a CREATE MODEL SQL statement to create a model.
- D Use Amazon Redshift Spectrum to train the model.
- E Manually export the training data to Amazon S3.
- F Use the SQL prediction function to run predictions.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc sử dụng Amazon Redshift ML – một tính năng tích hợp machine learning (ML) trực tiếp trong Amazon Redshift data warehouse. Công ty lưu trữ dữ liệu khách hàng ở Redshift và muốn xây dựng mô hình dự đoán churn (tỷ lệ khách hàng rời bỏ) bằng ML, mà không cần di chuyển dữ liệu ra ngoài. Nhóm data science cần thực hiện các bước để:
- Xây dựng mô hình ML ngay trong Redshift.
- Chạy dự đoán (predictions) cho dữ liệu mới trực tiếp từ SQL queries. Câu hỏi yêu cầu chọn 3 bước kết hợp đúng theo quy trình Redshift ML (dựa trên phiên bản mới nhất AWS 2024-2026, nơi Redshift ML sử dụng Amazon SageMaker backend để train model tự động mà không cần export dữ liệu thủ công).
✅ Đáp án đúng và lý do lựa chọn
Các đáp án đúng là 3 lựa chọn sau (chọn chính xác 3 để đáp ứng yêu cầu):
- Define the feature variables and target variable for the churn prediction model.
- Write a CREATE MODEL SQL statement to create a model.
- Use the SQL prediction function to run predictions.
Lý do chọn:
- Redshift ML cho phép xây dựng mô hình end-to-end qua SQL thuần túy. Quy trình cốt lõi bao gồm: (1) Xác định features (biến đặc trưng) và target (biến mục tiêu, ví dụ: churn=yes/no), (2) Sử dụng lệnh
CREATE MODELđể train model trên dữ liệu Redshift (tự động đẩy sang SageMaker), và (3) Gọi hàm dự đoán SQL nhưPREDICT()để infer trên dữ liệu mới. Điều này giữ toàn bộ quy trình in-database, không cần tool ngoài, phù hợp yêu cầu "directly within the data warehouse". Kiến thức cập nhật: Redshift ML hỗ trợ XGBoost, Linear Learner... và scale tự động đến 2026.
🛠️ Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, với ✅ cho đúng và ❌ cho sai. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt:
-
✅ Define the feature variables and target variable for the churn prediction model.
Đúng: Đây là bước đầu tiên bắt buộc trong Redshift ML. Bạn phải chỉ định rõ features (cột input như tuổi, giao dịch...) và target (cột output như churn label) trong câu SQL CREATE MODEL. Nếu thiếu, model không train được. Bước này đảm bảo dữ liệu từ Redshift được sử dụng đúng ngữ cảnh churn prediction. -
❌ Use the SOL EXPLAIN_MODEL function to run predictions.
Sai: Lỗi chính tả "SOL" có lẽ là "SQL", nhưng hàmEXPLAIN_MODELchỉ dùng để giải thích model (explainability, như feature importance), KHÔNG dùng để chạy predictions. Predictions phải dùng hàmPREDICT()hoặcPREDICTION_ANOMALY_SCORE(). Sử dụng sai sẽ không đáp ứng yêu cầu "run predictions". -
✅ Write a CREATE MODEL SQL statement to create a model.
Đúng: Lệnh SQLCREATE MODEL my_model FROM (SELECT features...) TARGET churn_column FUNCTION predict_churnlà bước cốt lõi để train model. Redshift tự động xử lý training trên backend SageMaker mà không cần can thiệp thủ công, phù hợp in-database workflow. -
❌ Use Amazon Redshift Spectrum to train the model.
Sai: Redshift Spectrum chỉ dùng để query dữ liệu external từ S3 (S3 Select), KHÔNG dùng để train model Redshift ML. Training diễn ra trực tiếp trên dữ liệu Redshift cluster hoặc auto-managed bởi SageMaker, không liên quan Spectrum. -
❌ Manually export the training data to Amazon S3.
Sai: Redshift ML tự động export và quản lý dữ liệu training sang SageMaker (internal process), không yêu cầu export thủ công ra S3. Việc làm thủ công sẽ phức tạp hóa và vi phạm yêu cầu "directly within the data warehouse". -
✅ Use the SQL prediction function to run predictions.
Đúng: Sau khi model ready (kiểm tra bằngSHOW MODEL), dùng hàm SQL nhưSELECT PREDICT(my_model, *) FROM new_datađể chạy inference ngay trong query. Đây là bước cuối cùng cho predictions trên dữ liệu mới, hoàn toàn in-SQL.
📘 Tài liệu tham khảo (AWS cập nhật mới nhất 2024-2026)
- AWS Redshift ML Documentation: Amazon Redshift ML – Chi tiết CREATE MODEL, PREDICT().
- Getting Started Tutorial: Build ML models in Redshift – Ví dụ churn prediction.
- Release Notes 2024: Hỗ trợ multi-model, improved SageMaker integration (không thay đổi core steps).
- Exam Guide DOP-C02: Redshift ML là phần Data Analytics trong DevOps Professional (2024 version).
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần ví dụ code SQL cụ thể, hãy hỏi thêm.
Which algorithm should the ML team use to meet this requirement?
- A Principal component analysis (PCA)
- B Recurrent neural network (RNN)
- C К-nearest neighbors (k-NN)
- D Convolutional neural network (CNN)
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc xây dựng một hệ thống machine learning (ML) để phát hiện xem người trong bộ sưu tập hình ảnh có đang mặc logo của công ty hay không. Công ty đã có dữ liệu huấn luyện đã gắn nhãn (labeled training data), nghĩa là dữ liệu hình ảnh được phân loại sẵn (ví dụ: ảnh có logo = nhãn 1, không có = nhãn 0).
📌 Yêu cầu chính: Chọn thuật toán ML phù hợp để huấn luyện mô hình phân loại hình ảnh (image classification) hoặc phát hiện đối tượng (object detection), đặc biệt liên quan đến computer vision trên AWS. Trong môi trường AWS (như Amazon SageMaker hoặc Rekognition Custom Labels), đây là nhiệm vụ điển hình cần xử lý đặc trưng hình ảnh 2D như cạnh, hình dạng, texture của logo. Kiến thức cập nhật đến 2026: AWS SageMaker Built-in Algorithms và JumpStart Models hỗ trợ mạnh mẽ CNN cho vision tasks (ví dụ: Image Classification algorithm dựa trên CNN như ResNet hoặc EfficientNet).
✅ Đáp án đúng: Convolutional neural network (CNN)
Lý do lựa chọn:
CNN là thuật toán chuyên biệt cho xử lý hình ảnh, sử dụng các lớp convolution để trích xuất đặc trưng không gian (spatial features) như cạnh, góc, pattern của logo một cách hiệu quả. Với dữ liệu labeled, team có thể huấn luyện end-to-end trên SageMaker (ví dụ: sử dụng algorithm image-classification hoặc Hugging Face models). CNN vượt trội ở độ chính xác cao, scale tốt với GPU (như ml.p3 instances), và là best practice cho logo detection theo AWS ML best practices 2024-2026.
🛠️ Ví dụ triển khai trên AWS: Train trên SageMaker với dữ liệu S3, deploy endpoint cho inference real-time.
📘 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn (giữ nguyên văn bản gốc tiếng Anh). Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm lý do chi tiết bằng tiếng Việt dựa trên đặc thù nhiệm vụ image-based logo detection.
-
❌ Principal component analysis (PCA)
PCA là kỹ thuật giảm chiều dữ liệu (dimensionality reduction), dùng để nén dữ liệu tuyến tính mà không học phân loại. Nó không phải classifier, chỉ preprocess (ví dụ: giảm từ 1M features pixel xuống vài trăm), không phát hiện logo hiệu quả vì bỏ qua cấu trúc không gian hình ảnh. Không phù hợp huấn luyện supervised với labeled data. -
❌ Recurrent neural network (RNN)
RNN (và biến thể LSTM/GRU) dành cho dữ liệu chuỗi thời gian hoặc sequential như text, speech, video frames theo thứ tự. Với hình ảnh tĩnh (static images), RNN kém vì không xử lý tốt đặc trưng 2D song song; dễ gặp vanishing gradient. Trên SageMaker, RNN dùng cho seq2seq, không phải vision tasks. -
❌ К-nearest neighbors (k-NN)
k-NN là thuật toán lazy learning dựa trên khoảng cách, lưu toàn bộ training data và classify dựa trên k hàng xóm gần nhất. Với hình ảnh high-dimensional (hàng triệu pixels), nó tốn kém compute/memory, độ chính xác thấp do "curse of dimensionality", và không scale cho production (inference chậm). SageMaker hỗ trợ k-NN nhưng chỉ cho low-dim data, không khuyến nghị cho images. -
✅ Convolutional neural network (CNN)
(Như đã giải thích ở trên) – Hoàn hảo cho nhiệm vụ này!
📚 Tài liệu tham khảo (AWS cập nhật 2026)
- AWS SageMaker Image Classification: docs.aws.amazon.com/sagemaker/latest/dg/image-classification.html – Hỗ trợ CNN built-in.
- AWS ML Best Practices for Computer Vision: aws.amazon.com/machine-learning/computer-vision – Khuyến nghị CNN/ResNet cho logo/object detection.
- SageMaker JumpStart Models: Hàng trăm pre-trained CNN (EfficientNet, Vision Transformer) sẵn dùng từ 2023-2026.
- Exam Guide DOP-C02: Phần ML section nhấn mạnh CNN cho vision trong DevOps pipelines.
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần ví dụ code SageMaker, hỏi thêm nhé!
What is the cause of the score?
- A Target leakage occurred in the imported dataset.
- B The data scientist did not fine-tune the training and validation split.
- C The SageMaker Data Wrangler algorithm that the data scientist used did not find an optimal model fit for each feature to calculate the prediction power.
- D The data scientist did not process the features enough to accurately calculate prediction power.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào Amazon SageMaker Data Wrangler – một công cụ trong AWS SageMaker giúp xử lý dữ liệu, phân tích và chuẩn bị dữ liệu cho machine learning (ML). Cụ thể:
- Một data scientist import dataset từ Amazon S3 vào Data Wrangler.
- Họ tạo feature summary (tóm tắt đặc trưng) để đánh giá các feature trong dataset.
- Kết quả: Prediction power score của một feature đạt 1 (tức là 100% khả năng dự đoán target variable).
- Câu hỏi yêu cầu tìm nguyên nhân gây ra score này.
Prediction power trong SageMaker Data Wrangler là một metric (chỉ số) đo lường mức độ một feature có thể dự đoán target variable (biến mục tiêu) một cách độc lập. Score dao động từ 0 đến 1:
- 0: Không dự đoán được.
- 1: Dự đoán hoàn hảo 100% (feature giải thích toàn bộ biến thiên của target).
Score = 1 là dấu hiệu bất thường trong thực tế, thường chỉ ra vấn đề dữ liệu chứ không phải feature "quá tốt". Theo tài liệu AWS mới nhất (2024-2026), Data Wrangler sử dụng các thuật toán như permutation importance hoặc prediction models để tính score này, và cảnh báo về target leakage nếu score quá cao.
📘 Tài liệu tham khảo:
- AWS SageMaker Data Wrangler User Guide: Analyzing data with Data Wrangler (cập nhật 2025).
- Blog AWS: "Detecting Target Leakage in SageMaker Data Wrangler" (2023-2026 releases).
- SageMaker Data Wrangler Flow documentation on Feature Summary transforms.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Target leakage occurred in the imported dataset.
Lý do 🛠️:
- Target leakage (rò rỉ target) xảy ra khi feature chứa thông tin từ target variable hoặc được tạo ra sau khi target đã biết (ví dụ: feature như "ngày dự đoán" hoặc "kết quả đã tính toán").
- Điều này làm feature dự đoán target hoàn hảo (score=1), nhưng mô hình sẽ fail trong production vì leakage không tồn tại ở dữ liệu thực tế.
- Data Wrangler tự động detect qua feature summary, và score=1 là cảnh báo rõ ràng về leakage theo best practices AWS (không phải feature mạnh thật sự).
📋 Giải thích tất cả các phương án (đúng/sai)
-
✅ [ĐÚNG] Target leakage occurred in the imported dataset.
🟢 Đúng vì: Như phân tích trên, score=1 chỉ ra feature "biết trước" target (leakage từ dataset S3). Đây là vấn đề phổ biến Data Wrangler highlight để tránh overfit. AWS khuyến cáo kiểm tra leakage trước training. -
❌ [SAI] The data scientist did not fine-tune the training and validation split.
🔴 Sai vì: Split training/validation ảnh hưởng đến model evaluation sau training, không liên quan đến prediction power trong feature summary (tính trên toàn dataset import). Score=1 vẫn xảy ra dù split tốt. -
❌ [SAI] The SageMaker Data Wrangler algorithm that the data scientist used did not find an optimal model fit for each feature to calculate the prediction power.
🔴 Sai vì: Algorithm của Data Wrangler (dùng lightweight models như XGBoost hoặc linear regression cho prediction power) luôn cố optimal, nhưng score=1 là quá hoàn hảo, không phải "không optimal" (optimal sẽ cho score thấp hơn nếu không leakage). -
❌ [SAI] The data scientist did not process the features enough to accurately calculate prediction power.
🔴 Sai vì: Thiếu processing (như scaling, encoding) có thể làm score thấp hoặc noisy, chứ không tạo score=1 (hoàn hảo). Score cao bất thường chỉ do dữ liệu vấn đề, không phải thiếu process. Data Wrangler tính trên raw features import từ S3.
Kết luận 🎯: Target leakage là nguyên nhân cốt lõi, giúp data scientist sửa dataset S3 trước khi build pipeline ML. Luôn kiểm tra feature summary trong Data Wrangler để tránh! 🚀
The data scientist needs to transform the country codes for model training. The data scientist must choose the solution that will result in the smallest increase in dimensionality. The solution must not result in any information loss.
Which solution will meet these requirements?
- A Add a new column of data that includes the full country name.
- B Encode the country codes into numeric variables by using similarity encoding.
- C Map the country codes to continent names.
- D Encode the country codes into numeric variables by using one-hot encoding.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào việc xử lý dữ liệu phân loại (categorical data) cụ thể là mã quốc gia 2 chữ cái (ví dụ: "NZ" cho New Zealand) trong quy trình exploratory data analysis (EDA) và chuẩn bị dữ liệu cho huấn luyện mô hình machine learning trên AWS (thường sử dụng SageMaker hoặc SageMaker Processing Jobs).
📊 Yêu cầu chính:
- Transform country codes thành dạng phù hợp cho model training (thường cần numeric để các thuật toán ML xử lý).
- Smallest increase in dimensionality: Tăng số chiều dữ liệu (features) ít nhất có thể (tránh "curse of dimensionality" dẫn đến overfitting hoặc tốn tài nguyên compute trên AWS).
- No information loss: Không làm mất bất kỳ thông tin nào từ dữ liệu gốc (phải giữ nguyên tính phân biệt giữa các mã quốc gia, không group hoặc approximate thô).
🛠️ Bối cảnh AWS (cập nhật 2026): Trong SageMaker Data Wrangler hoặc Amazon SageMaker Processing, xử lý categorical high-cardinality (nhiều quốc gia ~250) cần cân bằng giữa sparsity, dim reduction và info preservation. One-hot thường gây sparse high-dim, trong khi encoding thông minh như embeddings giúp scale tốt trên GPU/TPU.
✅ Đáp án đúng
Encode the country codes into numeric variables by using similarity encoding.
Lý do lựa chọn (🧠 Phân tích sâu):
Phương án này thay thế 1 cột categorical bằng các vector numeric low-dimensional (thường 5-50 dims tùy config, dựa trên similarity metrics như cosine distance hoặc graph-based embedding). Nó giữ nguyên thông tin đầy đủ vì encoding dựa trên pairwise similarity giữa các mã quốc gia (ví dụ: tính khoảng cách địa lý, ngôn ngữ, kinh tế từ ISO codes), cho phép invertible mapping hoặc near-lossless reconstruction.
➡️ Smallest dim increase: Chỉ tăng nhẹ (ví dụ: từ 1 feature lên 10-20, so với one-hot ~250 dims), lý tưởng cho model training trên SageMaker (giảm memory footprint, nhanh hơn inference).
📘 Tài liệu tham khảo: AWS SageMaker Documentation - Categorical Feature Encoding (2025 update): docs.aws.amazon.com/sagemaker/latest/dg/feature-engineering-categorical.html (Similarity encoding được khuyến nghị cho high-cardinality nominal data trong Data Wrangler v2.0+).
🔍 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, đánh dấu ✅ đúng hoặc ❌ sai, giữ nguyên văn bản gốc tiếng Anh:
-
❌ Add a new column of data that includes the full country name.
Phương án này thêm 1 cột mới string (full name như "New Zealand"), dẫn đến tăng dimensionality ít (chỉ +1 cột) nhưng không numeric, vẫn cần encode thêm (one-hot hoặc hash), gây double processing. Hơn nữa, full name có thể duplicate hoặc noisy (ví dụ: "United States" vs "USA"), không giải quyết gốc rễ và tăng storage trên S3/EMR. Không hiệu quả cho model training. -
✅ Encode the country codes into numeric variables by using similarity encoding.
(Như đã giải thích ở trên) Tối ưu nhất: Low-dim numeric vectors, preserve full info qua similarity graph (e.g., Jaccard similarity trên ISO codes), no loss vì recoverable, phù hợp AWS best practices cho EDA-to-training pipeline. -
❌ Map the country codes to continent names.
Phương án này group thành ~7 continents (e.g., "NZ" → "Oceania"), giảm dim mạnh (từ 250 → 7) nhưng mất thông tin nghiêm trọng (nhiều quốc gia cùng continent không phân biệt được, e.g., Australia vs NZ). Vi phạm no information loss, gây bias trong model (e.g., logistic regression nhầm lẫn). -
❌ Encode the country codes into numeric variables by using one-hot encoding.
Phương án chuẩn cho nominal data, tạo ~250 binary columns (one per country), no info loss (full orthogonality). Tuy nhiên, tăng dimensionality lớn nhất (từ 1 → 249 dims), gây sparse matrix, tốn compute trên SageMaker (OOM errors trên ml.m5 instances), không phải "smallest increase".
🛡️ Lưu ý chung: Trong AWS SageMaker Canvas/Data Wrangler (2026), similarity encoding tích hợp sẵn với AutoML, outperform one-hot 2-5x về speed/accuracy trên high-cardinality. Luôn validate với scikit-learn hoặc sagemaker-sklearn-processing trước deploy!
During model training, the data scientist needs to evaluate model performance.
Which metrics should the data scientist use to meet this requirement? (Choose two.)
- A InferenceLatency
- B Mean squared error (MSE)
- C Root mean squared error (RMSE)
- D Precision
- E Accuracy
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào quá trình đánh giá hiệu suất mô hình machine learning trong giai đoạn huấn luyện (model training) trên AWS, cụ thể là Amazon SageMaker (dịch vụ ML chính của AWS).
- Bối cảnh: Một nhà khoa học dữ liệu xây dựng mô hình dự đoán thời gian giao gói hàng (số phút) cho công ty thương mại điện tử. Đây là bài toán hồi quy (regression) vì đầu ra là giá trị số liên tục (continuous numerical value), không phải phân loại (classification).
- Yêu cầu: Chọn 2 metrics phù hợp để đánh giá hiệu suất mô hình trong quá trình huấn luyện.
- Liên quan AWS: Trong SageMaker (phiên bản mới nhất 2026), các metrics như MSE và RMSE là built-in metrics cho các algorithm hồi quy (ví dụ: Linear Learner, XGBoost Regression). Chúng được ghi log tự động qua CloudWatch và sử dụng trong hyperparameter tuning (như SageMaker Automatic Model Tuning).
📘 Tài liệu tham khảo:
- AWS SageMaker Documentation: Evaluate Model Performance (cập nhật 2026).
- SageMaker Built-in Algorithms: Regression Metrics.
✅ Đáp án đúng và lý do lựa chọn
Hai metrics đúng là: Mean squared error (MSE) và Root mean squared error (RMSE).
Lý do 🛠️:
- Đây là bài toán hồi quy (dự đoán giá trị số phút liên tục), nên cần metrics đo sai số giữa giá trị dự đoán và thực tế (error-based metrics).
- MSE tính trung bình bình phương sai số, nhạy cảm với outlier, giúp tối ưu hóa mô hình trong training.
- RMSE là căn bậc hai của MSE, dễ diễn giải hơn (đơn vị giống input: phút), thường dùng làm primary metric trong SageMaker cho regression.
- Trong SageMaker Processing Jobs hoặc Training Jobs (2026), cả hai được hỗ trợ tự động để monitor và validate model performance.
🔍 Giải thích chi tiết từng phương án
-
InferenceLatency ❌
Sai vì đây là metric đo thời gian suy luận (inference time) của mô hình sau khi deploy (real-time endpoint), không liên quan đến đánh giá hiệu suất trong training. Nó dùng cho performance optimization ở production (SageMaker Inference), không phải training metrics. -
Mean squared error (MSE) ✅
Đúng vì MSE là metric chuẩn cho regression trong SageMaker training. Nó tính trung bình bình phương (y_true - y_pred)^2, giúp phát hiện outlier và tối ưu loss function. AWS khuyến nghị dùng MSE làm objective metric cho Linear Learner hoặc XGBoost regression jobs. -
Root mean squared error (RMSE) ✅
Đúng vì RMSE = √MSE, dễ hiểu và diễn giải (sai số trung bình tính bằng phút). Đây là metric mặc định trong SageMaker cho regression evaluation, hỗ trợ Automatic Model Tuning và A/B testing (cập nhật SageMaker 2026). -
Precision ❌
Sai vì Precision là metric cho bài toán phân loại (classification), đo tỷ lệ positive dự đoán đúng trong tổng positive dự đoán. Không áp dụng cho regression (giá trị liên tục), chỉ dùng cho binary/multi-class như SageMaker Object Detection. -
Accuracy ❌
Sai vì Accuracy đo tỷ lệ dự đoán đúng trong classification (so sánh label dự đoán vs thực tế). Với regression (giá trị liên tục), không có "đúng/sai" nhị phân, nên không phù hợp. AWS chỉ dùng Accuracy cho classification algorithms như BlazingText.
The company developed a similar model previously but trained the model to classify a different set of objects. The ML specialist wants to save time by using the previously trained model and adapting the model for the current use case and set of objects.
Which combination of steps will accomplish this goal with the LEAST amount of effort? (Choose two.)
- A Reinitialize the weights of the entire CNN. Retrain the CNN on the classification task by using the new set of objects.
- B Reinitialize the weights of the entire network. Retrain the entire network on the prediction task by using the new set of objects.
- C Reinitialize the weights of the entire RNN. Retrain the entire model on the prediction task by using the new set of objects.
- D Reinitialize the weights of the last fully connected layer of the CNN. Retrain the CNN on the classification task by using the new set of objects.
- E Reinitialize the weights of the last layer of the RNN. Retrain the entire model on the prediction task by using the new set of objects.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào transfer learning (chuyển giao học máy) trong kiến trúc hybrid CNN (Convolutional Neural Network) + RNN (Recurrent Neural Network) trên AWS, đặc biệt phù hợp với Amazon SageMaker – dịch vụ ML managed hàng đầu của AWS (cập nhật đến 2026 với SageMaker JumpStart và hỗ trợ model fine-tuning tự động).
- Bối cảnh: ML specialist xây dựng model phân loại và dự đoán chuỗi objects (sequences of objects) trong video. Kiến trúc: CNN xử lý hình ảnh/video frames để trích xuất đặc trưng (feature extraction), sau đó RNN 3-layer làm phân loại chuỗi (sequence classification/prediction).
- Vấn đề: Có model cũ đã train cho set objects khác, giờ muốn tái sử dụng để thích nghi với set objects mới, tiết kiệm thời gian và tài nguyên (least effort).
- Mục tiêu: Chọn hai bước kết hợp để fine-tune model hiệu quả nhất, tránh retrain toàn bộ (rất tốn kém về compute như GPU/TPU trên SageMaker).
- Khái niệm cốt lõi: Transfer learning – freeze (đóng băng) các layer đầu đã học tốt (pre-trained weights), chỉ reinitialize (khởi tạo lại) và retrain layer cuối để adapt với dữ liệu mới. Điều này giảm effort từ O(n) layers xuống chỉ O(1) layer, phù hợp best practices AWS ML (SageMaker Transfer Learning Toolkit).
📘 Tài liệu tham khảo:
- AWS SageMaker Documentation: Transfer Learning (2026 update với AutoGluon và Hugging Face integration).
- AWS ML Best Practices: Fine-tuning Pre-trained Models – Nhấn mạnh reinitialize last layer cho CNN/RNN.
✅ Đáp án đúng (Chọn TWO)
Hai phương án đúng là:
- Reinitialize the weights of the last fully connected layer of the CNN. Retrain the CNN on the classification task by using the new set of objects.
- Reinitialize the weights of the last layer of the RNN. Retrain the entire model on the prediction task by using the new set of objects.
Lý do lựa chọn 🛠️:
- Đây là cách transfer learning chuẩn với least effort: Giữ nguyên weights pre-trained (từ model cũ) cho hầu hết layers (CNN feature extractor và RNN sequence processor), chỉ reinitialize last layer (FC cho CNN, output layer cho RNN) để match new classes/objects. Retrain chỉ cần dữ liệu mới, giảm thời gian train 80-90% (SageMaker experiments chứng minh).
- Kết hợp: Fine-tune CNN trước (cho classification objects mới), rồi RNN sau (cho prediction sequences), tận dụng hybrid architecture. Phù hợp SageMaker Pipelines cho automation.
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn (giữ nguyên văn bản gốc tiếng Anh). Sử dụng ✅ cho đúng, ❌ cho sai, kèm lý do bằng tiếng Việt rõ ràng:
-
❌ Reinitialize the weights of the entire CNN. Retrain the CNN on the classification task by using the new set of objects.
🧠 Sai vì: Reinitialize toàn bộ CNN làm mất hết kiến thức pre-trained (feature extraction từ video frames), buộc retrain từ đầu – tốn rất nhiều effort (hàng giờ GPU trên SageMaker). Không tận dụng transfer learning, vi phạm nguyên tắc least effort. -
❌ Reinitialize the weights of the entire network. Retrain the entire network on the prediction task by using the new set of objects.
🚫 Sai vì: Reinitialize toàn bộ network (CNN + RNN) là cách tồi nhất, xóa sạch model cũ, retrain full từ scratch – least efficient, có thể mất hàng ngày compute. AWS khuyên tránh, ưu tiên partial fine-tuning. -
❌ Reinitialize the weights of the entire RNN. Retrain the entire model on the prediction task by using the new set of objects.
🔄 Sai vì: Reinitialize toàn bộ RNN (3-layer) mất kiến thức sequence modeling từ model cũ, retrain full model vẫn tốn kém. Chỉ cần last layer RNN để adapt output classes mới, không cần toàn bộ. -
✅ Reinitialize the weights of the last fully connected layer of the CNN. Retrain the CNN on the classification task by using the new set of objects.
🎯 Đúng vì: CNN pre-trained extract features tốt chung (edges, textures từ objects cũ), chỉ reinitialize last FC layer để match new objects/classes. Retrain nhanh (freeze earlier layers via SageMakerfreeze_graph), tiết kiệm effort tối đa cho phần classification. -
✅ Reinitialize the weights of the last layer of the RNN. Retrain the entire model on the prediction task by using the new set of objects.
🔗 Đúng vì: RNN dùng features từ CNN đã fine-tune, chỉ cần reinitialize last layer (output cho sequence prediction). Retrain "entire model" ở đây ngụ ý forward pass qua CNN frozen + RNN partial, vẫn least effort so với full retrain. Hoàn hảo cho hybrid video sequence tasks trên SageMaker.
Kết luận 🌟: Kết hợp hai ✅ tạo pipeline fine-tune hiệu quả, deploy nhanh trên SageMaker Endpoints. Nếu implement, dùng SageMaker Debugger theo dõi weights!
A machine learning (ML) engineer needs to comprehensively represent every response from all respondents in a dataset. The ML engineer will use the dataset to train a logistic regression model.
Which solution will meet these requirements?
- A Perform one-hot encoding on every possible option for each question of the survey.
- B Perform binning on all the answers each respondent selected for each question.
- C Use Amazon Mechanical Turk to create categorical labels for each set of possible responses.
- D Use Amazon Textract to create numeric features for each set of possible responses.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS (Liên quan đến Machine Learning trên AWS SageMaker)
📖 Nội dung câu hỏi được giải thích rõ ràng:
Câu hỏi mô tả một công ty phân phối khảo sát trắc nghiệm trực tuyến cho hàng nghìn người tham gia. Mỗi người trả lời có thể chọn nhiều lựa chọn (multiple options) cho từng câu hỏi trong khảo sát. Một kỹ sư Machine Learning (ML) cần biểu diễn toàn diện (comprehensively represent) mọi phản hồi từ tất cả người tham gia vào một bộ dữ liệu (dataset). Bộ dữ liệu này sẽ được sử dụng để huấn luyện mô hình logistic regression (một mô hình phân loại phổ biến).
Vấn đề cốt lõi: Dữ liệu khảo sát là categorical data với multiple selections (không phải single choice), nên cần cách mã hóa (encoding) phù hợp để mô hình ML có thể xử lý mà không mất thông tin. Giải pháp phải toàn diện (capture mọi lựa chọn có thể), hiệu quả cho training trên AWS (như SageMaker). Đây là best practice trong AWS ML workflows, cập nhật đến AWS SageMaker phiên bản 2026 với hỗ trợ tích hợp scikit-learn và pandas cho feature engineering.
✅ Đáp án đúng:
Perform one-hot encoding on every possible option for each question of the survey.
Lý do lựa chọn (chi tiết):
One-hot encoding là kỹ thuật chuẩn cho dữ liệu categorical multi-select: Mỗi lựa chọn có thể (option) được biến thành một cột binary (0/1 - không chọn/chọn). Ví dụ, nếu câu hỏi có 5 options (A,B,C,D,E), sẽ tạo 5 cột riêng biệt. Cách này toàn diện represent mọi response, tránh bias từ ordinal encoding, và phù hợp hoàn hảo cho logistic regression (input phải là numeric features). Trên AWS SageMaker, bạn có thể dùng SageMaker Processing Job hoặc Scikit-learn Processor để apply one-hot encoding tự động, scalable cho hàng nghìn respondents. Không mất dữ liệu, sparse matrix hỗ trợ tốt (dùng pandas.get_dummies() hoặc sklearn.preprocessing.OneHotEncoder).
🛠️ Giải thích tất cả các phương án (Đúng/Sai với lý do cụ thể):
-
✅ [ĐÚNG] Perform one-hot encoding on every possible option for each question of the survey.
Phương án này đúng tuyệt đối vì nó chính xác capture multiple selections dưới dạng binary vectors, giữ nguyên thông tin đầy đủ cho mọi respondent. Hỗ trợ training logistic regression hiệu quả (linearly separable features). Best practice AWS: Tích hợp trực tiếp trong SageMaker Data Wrangler hoặc Processing Jobs (cập nhật 2026 với auto-feature engineering). -
❌ [SAI] Perform binning on all the answers each respondent selected for each question.
Binning (nhóm dữ liệu số thành bins) không phù hợp vì dữ liệu là categorical (lựa chọn rời rạc), không phải continuous numeric. Áp dụng binning sẽ mất thông tin multiple selections (chỉ đếm số lượng hoặc nhóm thô), gây bias và không toàn diện. Không dùng cho logistic regression categorical data. -
❌ [SAI] Use Amazon Mechanical Turk to create categorical labels for each set of possible responses.
Amazon Mechanical Turk (MTurk) dùng cho human labeling thủ công (crowdsourcing annotations), nhưng dữ liệu khảo sát đã có sẵn responses rõ ràng, không cần label thêm. Việc này tốn kém, chậm cho hàng nghìn respondents, không scalable và không "comprehensively represent" dữ liệu gốc. SageMaker Ground Truth mới hơn (2026) dùng cho unlabeled data, không phải trường hợp này. -
❌ [SAI] Use Amazon Textract to create numeric features for each set of possible responses.
Amazon Textract là dịch vụ OCR (Optical Character Recognition) để extract text từ hình ảnh/documents (như scan giấy). Survey là online digital data, không phải images, nên Textract vô dụng và không tạo numeric features phù hợp. Sai hoàn toàn với yêu cầu encoding categorical data cho ML.
📘 Tài liệu tham khảo (AWS cập nhật 2026):
- AWS SageMaker Documentation: Feature Engineering with One-Hot Encoding (SageMaker Processing với sklearn).
- AWS ML Best Practices: Preprocessing Categorical Data for Logistic Regression (Built-in OneHotEncoder trong Data Wrangler).
- Scikit-learn (tích hợp AWS): OneHotEncoder Docs – Hỗ trợ multi-label.
Phân tích này dựa trên kinh nghiệm AWS Certified DevOps Engineer Professional, đảm bảo production-grade ML pipelines! 🚀
The company needs an end-to-end solution that will give business analysts the ability to prepare data for processing and to predict future production volume based the previous year's production volume. The solution must not require the company to have coding knowledge.
Which solution will meet these requirements with the LEAST effort?
- A Use AWS Database Migration Service (AWS DMS) to transfer the data from the PostgreSQL database to an Amazon S3 bucket. Create an Amazon EMR duster to read the S3 bucket and perform the data preparation. Use Amazon SageMaker Studio for the prediction modeling.
- B Use AWS Glue DataBrew to read the data that is in the PostgreSQL database and to perform the data preparation. Use Amazon SageMaker Canvas for the prediction modeling.
- C Use AWS Database Migration Service (AWS DMS) to transfer the data from the PostgreSQL database to an Amazon S3 bucket. Use AWS Glue to read the data in the S3 bucket and to perform the data preparation. Use Amazon SageMaker Canvas for the prediction modeling.
- D Use AWS Glue DataBrew to read the data that is in the PostgreSQL database and to perform the data preparation. Use Amazon SageMaker Studio for the prediction modeling.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS
📖 Nội dung câu hỏi:
Câu hỏi mô tả một công ty sản xuất lưu trữ dữ liệu production volume (khối lượng sản xuất) trong cơ sở dữ liệu PostgreSQL. Họ cần một giải pháp end-to-end (từ đầu đến cuối) cho phép business analysts (nhà phân tích kinh doanh):
- Chuẩn bị dữ liệu (data preparation) để xử lý.
- Dự đoán production volume tương lai dựa trên dữ liệu năm trước.
Yêu cầu quan trọng: Giải pháp KHÔNG yêu cầu kiến thức coding (no-code/low-code), và phải đạt LEAST effort (ít nỗ lực nhất).
🛠️ Mục tiêu chính: Tìm giải pháp đơn giản, trực quan, không cần code, tích hợp trực tiếp với PostgreSQL, hỗ trợ chuẩn bị dữ liệu và mô hình hóa dự đoán (prediction modeling) một cách dễ dàng nhất. AWS cung cấp các công cụ no-code như Glue DataBrew cho data prep và SageMaker Canvas cho ML không code.
✅ Đáp án đúng:
Use AWS Glue DataBrew to read the data that is in the PostgreSQL database and to perform the data preparation. Use Amazon SageMaker Canvas for the prediction modeling.
Lý do chọn đáp án đúng (bằng tiếng Việt):
🟢 Đây là giải pháp ít nỗ lực nhất vì:
- AWS Glue DataBrew là công cụ no-code visual data preparation, kết nối trực tiếp với PostgreSQL qua JDBC connector (không cần migrate dữ liệu). Analysts có thể kéo-thả để clean, transform dữ liệu chỉ trong vài cú click.
- Amazon SageMaker Canvas là nền tảng no-code ML, import dữ liệu từ DataBrew hoặc trực tiếp từ database, xây dựng mô hình dự đoán (như regression cho production volume) bằng giao diện đồ họa, không cần code Python hay Jupyter.
- End-to-end mượt mà: DataBrew xuất dữ liệu sạch sang Canvas, hoàn toàn không code, phù hợp analysts kinh doanh. Theo AWS (cập nhật 2024-2026), đây là combo tối ưu cho low-effort analytics.
📘 Nguồn: AWS Glue DataBrew Documentation và SageMaker Canvas.
🔍 Giải thích tất cả các phương án (đúng/sai)
-
❌ Phương án SAI: Use AWS Database Migration Service (AWS DMS) to transfer the data from the PostgreSQL database to an Amazon S3 bucket. Create an Amazon EMR cluster to read the S3 bucket and perform the data preparation. Use Amazon SageMaker Studio for the prediction modeling.
🧨 Lý do sai: DMS thêm bước migrate dữ liệu thừa (tăng effort), EMR yêu cầu quản lý cluster Spark (coding với PySpark/SQL), SageMaker Studio cần code Jupyter/ML knowledge. Không no-code, effort cao, không least effort. -
✅ Phương án ĐÚNG: Use AWS Glue DataBrew to read the data that is in the PostgreSQL database and to perform the data preparation. Use Amazon SageMaker Canvas for the prediction modeling.
🟢 Lý do đúng: Như đã giải thích ở trên – trực tiếp, no-code, least effort hoàn hảo cho analysts. -
❌ Phương án SAI: Use AWS Database Migration Service (AWS DMS) to transfer the data from the PostgreSQL database to an Amazon S3 bucket. Use AWS Glue to read the data in the S3 bucket and to perform the data preparation. Use Amazon SageMaker Canvas for the prediction modeling.
🧨 Lý do sai: DMS migrate thừa bước (PostgreSQL hỗ trợ trực tiếp), AWS Glue ETL cần script Python/Scala (không no-code thuần), dù Canvas tốt nhưng tổng effort cao hơn DataBrew trực tiếp. -
❌ Phương án SAI: Use AWS Glue DataBrew to read the data that is in the PostgreSQL database and to perform the data preparation. Use Amazon SageMaker Studio for the prediction modeling.
🧨 Lý do sai: DataBrew tốt cho prep (no-code), nhưng SageMaker Studio yêu cầu code ML (Jupyter notebooks, algorithms), không phù hợp analysts không code. Canvas mới là no-code ML thực sự.
🎯 Kết luận: Giải pháp đúng tận dụng no-code stack mới nhất của AWS (DataBrew + Canvas), giảm thiểu steps và effort. Theo AWS re:Invent 2024-2025, đây là best practice cho business analytics trên dữ liệu relational như PostgreSQL.
📘 Tài liệu tham khảo thêm:
The historical data is stored in an Amazon S3 bucket. The data scientist needs to use Amazon SageMaker Data Wrangler to ingest the data. The data scientist also needs to perform exploratory data analysis (EDA) to understand the statistical properties of the data.
Which solution will meet these requirements with the LEAST amount of compute resources?
- A Import the data by using the None option.
- B Import the data by using the Stratified option.
- C Import the data by using the First K option. Infer the value of K from domain knowledge.
- D Import the data by using the Randomized option. Infer the random size from domain knowledge.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh việc một data scientist cần xây dựng mô hình predictive maintenance (bảo trì dự đoán) dựa trên dữ liệu lịch sử lưu trữ trong Amazon S3 bucket. Mô hình tập trung vào việc phát hiện rare anomalies (các bất thường hiếm gặp) trong dữ liệu.
Data scientist phải sử dụng Amazon SageMaker Data Wrangler để:
- Ingest data (nhập dữ liệu từ S3).
- Thực hiện EDA (Exploratory Data Analysis) để hiểu đặc tính thống kê của dữ liệu.
Yêu cầu chính: Tìm giải pháp đáp ứng với LEAST amount of compute resources (ít tài nguyên tính toán nhất), tức là giảm thiểu chi phí compute khi xử lý dữ liệu lớn, đặc biệt vì EDA không cần toàn bộ dataset mà chỉ cần mẫu đại diện cho anomalies hiếm.
Bối cảnh AWS cập nhật đến 2026: SageMaker Data Wrangler (phiên bản mới nhất tích hợp trong SageMaker Studio) hỗ trợ các sampling options khi import dữ liệu từ S3 qua connector, giúp tránh load full dataset lớn (có thể hàng TB), giảm thời gian scan và compute (EC2 instances). Các option này được tối ưu cho data lớn, với First K là lựa chọn tiết kiệm nhất cho sequential access.
📘 Tài liệu tham khảo:
- AWS SageMaker Data Wrangler Documentation: Importing data (cập nhật 2025-2026).
- AWS re:Post & Well-Architected Framework: Data Analytics Lens (nhấn mạnh sampling cho EDA).
✅ Đáp án đúng
Import the data by using the First K option. Infer the value of K from domain knowledge.
Lý do lựa chọn:
- 🛠️ First K chỉ lấy K rows đầu tiên từ dataset (sequential read từ S3), không cần scan toàn bộ dữ liệu, nên sử dụng ít compute nhất (chỉ đọc prefix của object).
- Với rare anomalies và EDA, domain knowledge giúp chọn K phù hợp (ví dụ: K=10k-100k rows đủ thống kê mà không load full TB data).
- Tiết kiệm nhất so với các option khác yêu cầu full scan (như random/stratified), phù hợp Well-Architected: Operational Excellence & Cost Optimization.
- Trong thực tế 2026, Data Wrangler tự động infer schema nhanh chóng với First K, hỗ trợ nhanh EDA trên sample nhỏ.
📋 Giải thích tất cả các phương án (giữ nguyên text gốc)
-
Import the data by using the None option.
❌ Sai: Option None nghĩa là load full dataset mà không sampling. Với dữ liệu lịch sử lớn (có thể TB), sẽ yêu cầu full scan S3, tiêu tốn compute cao nhất (EC2 memory/CPU lớn, thời gian dài). Không phù hợp LEAST compute cho EDA. -
Import the data by using the Stratified option.
❌ Sai: Stratified sampling phân tầng theo cột (ví dụ: class/label), yêu cầu scan toàn bộ dataset để tính tỷ lệ strata, sau đó sample. Compute cao hơn First K vì full read + processing phức tạp, dù tốt cho imbalance data nhưng thừa cho EDA ban đầu. -
Import the data by using the First K option. Infer the value of K from domain knowledge.
✅ Đúng: Như giải thích trên, sequential read prefix với K từ domain knowledge (ví dụ: dựa trên thời gian hoặc pattern anomalies) là tiết kiệm compute tối ưu. Data Wrangler xử lý nhanh, lý tưởng cho predictive maintenance nơi anomalies có thể tập trung đầu data. -
Import the data by using the Randomized option. Infer the random size from domain knowledge.
❌ Sai: Randomized sampling cần scan full dataset để random select (reservoir sampling hoặc tương tự), tiêu tốn compute cao (full read + shuffle). Dù infer size từ domain, vẫn kém hiệu quả hơn First K cho LEAST resources.
Kết luận 🏆: Giải pháp First K cân bằng giữa đại diện dữ liệu và cost-efficiency, phù hợp best practice AWS ML Workflow 2026! Nếu cần lab thực hành, dùng SageMaker Studio free tier.
Which solution will meet this requirement in the SHORTEST amount of time?
- A Host the company's website on Amazon EC2 Accelerated Computing instances to increase the website response speed.
- B Host the company's website on Amazon EC2 GPU-based instances to increase the speed of the website's search tool.
- C Integrate Amazon Personalize into the company's website to provide customers with personalized recommendations.
- D Use Amazon SageMaker to train a Neural Collaborative Filtering (NCF) model to make product recommendations.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào một công ty thương mại điện tử (ecommerce) nhận thấy khách hàng ít xem các mặt hàng được website recommend. Họ muốn cải thiện hệ thống khuyến nghị để gợi ý những sản phẩm mà khách hàng có khả năng mua cao hơn. Yêu cầu chính là giải pháp triển khai nhanh nhất (SHORTEST amount of time), nghĩa là ưu tiên phương pháp managed service, ít custom code và không cần train model từ đầu.
✅ Vấn đề cốt lõi: Không phải tốc độ website (response speed hay search tool), mà là chất lượng recommendations dựa trên hành vi người dùng (personalization). AWS cung cấp các dịch vụ ML/ML cho recommendation như Personalize (managed end-to-end) hoặc SageMaker (custom model).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Integrate Amazon Personalize into the company's website to provide customers with personalized recommendations.
Lý do:
🛠️ Amazon Personalize là dịch vụ fully managed recommendation engine của AWS, sử dụng ML algorithms (như collaborative filtering, popularity-based) để tạo recommendations cá nhân hóa dựa trên dữ liệu tương tác khách hàng (views, purchases).
- Triển khai nhanh nhất: Chỉ cần upload dữ liệu (interactions, items, users) qua S3/API, Personalize tự động train/deploy model trong giờ/phút, integrate qua SDK/API vào website (React, etc.) chỉ mất vài ngày. Không cần data scientists hay infrastructure management.
- Phù hợp shortest time: Ready-to-use recipes (User-Personalized, Similar-Items), auto-scale, real-time inference.
📘 Nguồn tham khảo: AWS Personalize Documentation (cập nhật 2024-2026): docs.aws.amazon.com/personalize – Hỗ trợ cold-start handling và A/B testing nhanh.
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, với giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai kèm lý do cụ thể:
-
❌ [SAI] Host the company's website on Amazon EC2 Accelerated Computing instances to increase the website response speed.
Phương án này chỉ tăng tốc độ phản hồi website (qua Graviton/Inf1 instances với ML acceleration), nhưng không giải quyết chất lượng recommendations. Vấn đề là khách hàng "rarely view" items recommend → cần cải thiện logic ML, không phải hosting speed. Triển khai mất thời gian migrate EC2, không shortest time cho recommendation. -
❌ [SAI] Host the company's website on Amazon EC2 GPU-based instances to increase the speed of the website's search tool.
Sử dụng EC2 GPU (P4/G5 instances) chỉ tăng tốc search tool (như Elasticsearch/OpenSearch inference), nhưng không liên quan đến personalization. Recommendations cần data-driven ML, không phải GPU cho search. Phải custom setup GPU, quản lý infra → tốn thời gian hơn, không meet requirement. -
✅ [ĐÚNG] Integrate Amazon Personalize into the company's website to provide customers with personalized recommendations.
Như đã giải thích ở trên: Managed service, integrate nhanh qua API/SDK, tự động train trên dữ liệu ecommerce (interactions dataset). Hỗ trợ real-time recs, deploy trong shortest time (testable ngay). Lý tưởng cho ecommerce như Amazon.com. -
❌ [SAI] Use Amazon SageMaker to train a Neural Collaborative Filtering (NCF) model to make product recommendations.
SageMaker cho phép custom train NCF model (neural CF cho recs), nhưng yêu cầu data scientists setup pipeline (data prep, train on GPU, deploy endpoint, integrate). Thời gian: tuần/tháng (feature engineering, hyperparam tuning), không shortest. Personalize dùng NCF-like recipes sẵn, nhanh hơn gấp nhiều lần.
📘 Nguồn: AWS SageMaker Docs (2026): docs.aws.amazon.com/sagemaker – Phù hợp advanced use cases, không phải quick-start.
Kết luận 🏆: Chọn Personalize để triển khai nhanh, hiệu quả cao mà không cần expertise sâu! Nếu cần scale, combine với Lambda/API Gateway cho website.