Ngân hàng đề — AWS Certified Machine Learning Specialty
Tìm thấy 371 câu.
The developer wants to verify that an autoregressive integrated moving average (ARIMA) approach will be a suitable model for the use case.
How should the developer verify the suitability of an ARIMA approach?
- A Use Amazon SageMaker Data Wrangler. Import the data from Amazon S3. Impute hourly missing data. Perform a Seasonal Trend decomposition.
- B Use Amazon SageMaker Autopilot. Create a new experiment that specifies the S3 data location. Choose ARIMA as the machine learning (ML) problem. Check the model performance.
- C Use Amazon SageMaker Data Wrangler. Import the data from Amazon S3. Resample data by using the aggregate daily total. Perform a Seasonal Trend decomposition.
- D Use Amazon SageMaker Autopilot. Create a new experiment that specifies the S3 data location. Impute missing hourly values. Choose ARIMA as the machine learning (ML) problem. Check the model performance.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc xác thực tính phù hợp của mô hình ARIMA (Autoregressive Integrated Moving Average) cho mô hình dự báo nhu cầu hàng ngày (daily demand forecasting).
- Bối cảnh: Một lập trình viên tại công ty bán lẻ lưu trữ dữ liệu nhu cầu lịch sử theo giờ (hourly demand data) trong bucket Amazon S3. Dữ liệu này thiếu một số giờ (missing data for some hours), gây ra khoảng trống trong chuỗi thời gian (time series).
- Mục tiêu: Verify ARIMA có phù hợp không. ARIMA là mô hình time series cổ điển, yêu cầu:
- Dữ liệu liên tục, không có khoảng trống (cần impute missing values).
- Kiểm tra stationarity (tính dừng): Phân tích phân rã xu hướng mùa vụ (Seasonal Trend decomposition, thường dùng STL - Seasonal and Trend decomposition using Loess) để xem thành phần trend, seasonal, và residual có stationary không. Nếu residual stationary sau differencing, ARIMA phù hợp.
- Thách thức: Dữ liệu hourly nhưng dự báo daily → Cần giữ granularity hourly để phân tích chính xác seasonality (ví dụ: peak giờ cao điểm). Không resample sớm để tránh mất thông tin.
✅ Cách tiếp cận đúng: Sử dụng SageMaker Data Wrangler để chuẩn bị dữ liệu (import từ S3, impute missing hourly, rồi decompose) – đây là bước exploratory data analysis (EDA) chuẩn để verify model trước khi train.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Amazon SageMaker Data Wrangler. Import the data from Amazon S3. Impute hourly missing data. Perform a Seasonal Trend decomposition.
Lý do:
- SageMaker Data Wrangler (phiên bản mới nhất 2024-2026) có Time Series Canvas chuyên biệt cho time series prep: Import trực tiếp từ S3, impute missing values (hỗ trợ forward-fill, interpolation, mean/median theo hourly), rồi Seasonal Trend decomposition (STL) để visualize trend/seasonal/residual.
- Điều này verify ARIMA suitability bằng cách kiểm tra: Nếu residual sau decompose là stationary (ACF/PACF không có pattern mạnh), ARIMA phù hợp. Không train model ngay, chỉ EDA.
- Giữ nguyên hourly granularity, phù hợp với missing hourly data. 🛠️ Hoàn hảo cho use case!
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt với lý do đúng/sai dựa trên tính năng AWS SageMaker mới nhất (2026).
-
Use Amazon SageMaker Data Wrangler. Import the data from Amazon S3. Impute hourly missing data. Perform a Seasonal Trend decomposition.
✅ ĐÚNG (như đã giải thích ở trên). Data Wrangler lý tưởng cho EDA time series: Impute giữ hourly freq, decompose kiểm tra seasonality → Verify ARIMA trực tiếp mà không train model. 🧩 -
Use Amazon SageMaker Autopilot. Create a new experiment that specifies the S3 data location. Choose ARIMA as the machine learning (ML) problem. Check the model performance.
❌ SAI: SageMaker Autopilot (AutML) dành cho tabular classification/regression, KHÔNG hỗ trợ chọn ARIMA cụ thể như "ML problem type". Autopilot tự động thử XGBoost/Linear Learner, không chuyên time series ARIMA. Không impute missing → Dữ liệu lỗi, không verify suitability mà chỉ train blindly. 📉 -
Use Amazon SageMaker Data Wrangler. Import the data from Amazon S3. Resample data by using the aggregate daily total. Perform a Seasonal Trend decomposition.
❌ SAI: Resample aggregate daily total làm mất granularity hourly, gây sai lệch (ví dụ: missing hours làm daily total thấp). ARIMA cần chuỗi hourly đầy đủ để detect intra-day seasonality. Decompose trên daily data không verify chính xác cho hourly-based daily forecast. Data Wrangler hỗ trợ resample nhưng không phù hợp ở đây. 🔄 -
Use Amazon SageMaker Autopilot. Create a new experiment that specifies the S3 data location. Impute missing hourly values. Check the model performance.
❌ SAI: Autopilot KHÔNG có built-in impute missing hourly tự động cho time series (chỉ basic preprocessing tabular). Không chọn ARIMA được, và "check performance" là train+eval → KHÔNG verify suitability (chỉ test model, không EDA decompose). Missing data gây fail experiment. 🚫
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- SageMaker Data Wrangler Time Series: AWS Docs - Data Wrangler Time Series Canvas – Hỗ trợ impute, decompose STL từ 2023+.
- ARIMA Suitability Check: AWS ML Time Series Blog – Khuyến nghị decompose trước ARIMA/SARIMA.
- Autopilot Limitations: AWS Docs - Autopilot – Không hỗ trợ ARIMA native.
🛡️ Lời khuyên: Luôn dùng Data Wrangler cho prep trước khi train với SageMaker Canvas/Autopilot/Forecast!
Which solution will meet the company’s security requirements?
- A Connect the SageMaker notebook instances that are in the VPC by using AWS Site-to-Site VPN to encrypt all internet-bound traffic. Configure VPC flow logs. Monitor all network traffic to detect and prevent any malicious activity.
- B Configure the VPC that contains the SageMaker notebook instances to use VPC interface endpoints to establish connections for training and hosting. Modify any existing security groups that are associated with the VPC interface endpoint to allow only outbound connections for training and hosting.
- C Create an IAM policy that prevents access the internet. Apply the IAM policy to an IAM role. Assign the IAM role to the SageMaker notebook instances in addition to any IAM roles that are already assigned to the instances.
- D Create VPC security groups to prevent all incoming and outgoing traffic. Assign the security groups to the SageMaker notebook instances.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh việc triển khai Amazon SageMaker để phát triển các mô hình machine learning (ML) trong môi trường AWS. Công ty lưu trữ dữ liệu huấn luyện (training data) trong Amazon S3 bucket, và các SageMaker notebook instances được host bên trong một VPC (Virtual Private Cloud). Yêu cầu bảo mật nghiêm ngặt từ công ty: SageMaker notebook instances tuyệt đối không được có kết nối internet (no internet connectivity).
Mục tiêu là tìm giải pháp đảm bảo notebook instances có thể truy cập S3 để lấy dữ liệu training, đồng thời thực hiện các hoạt động huấn luyện (training) và hosting mô hình mà không cần kết nối internet, tuân thủ policy bảo mật. Điều này đòi hỏi sử dụng các kết nối private routing trong VPC, tránh public internet hoàn toàn. 🛡️
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Configure the VPC that contains the SageMaker notebook instances to use VPC interface endpoints to establish connections for training and hosting. Modify any existing security groups that are associated with the VPC interface endpoint to allow only outbound connections for training and hosting.
Lý do:
- Giải pháp này sử dụng VPC interface endpoints (powered by AWS PrivateLink) để kết nối private từ VPC đến các dịch vụ SageMaker (như API training, hosting) và S3 mà không đi qua internet.
- Cụ thể: Tạo interface endpoints cho
com.amazonaws.region.sagemaker(API),com.amazonaws.region.sagemaker.runtime(hosting), vàcom.amazonaws.region.s3(cho dữ liệu). - Sau đó, chỉnh sửa security groups của endpoint để chỉ cho phép outbound traffic cần thiết (training/hosting), đảm bảo kiểm soát chặt chẽ và không có internet access.
- Đây là best practice theo AWS (cập nhật 2024-2026), giúp notebook instances hoạt động isolated hoàn toàn trong VPC. ✅ Hoàn hảo cho yêu cầu "no internet connectivity"!
📋 Phân tích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn. Tôi giữ nguyên văn bản gốc tiếng Anh của phương án, đánh dấu ✅ (đúng) hoặc ❌ (sai), và giải thích rõ ràng bằng tiếng Việt:
-
❌ Connect the SageMaker notebook instances that are in the VPC by using AWS Site-to-Site VPN to encrypt all internet-bound traffic. Configure VPC flow logs. Monitor all network traffic to detect and prevent any malicious activity.
Phân tích sai: Site-to-Site VPN dùng để kết nối on-premises với AWS VPC qua đường hầm mã hóa, nhưng không ngăn chặn internet access từ notebook instances. Traffic vẫn có thể route ra internet nếu NAT gateway hoặc public subnet tồn tại. VPC Flow Logs chỉ monitor, không block internet. Giải pháp này thừa thãi và không meet yêu cầu "no internet connectivity". 🛑 -
✅ Configure the VPC that contains the SageMaker notebook instances to use VPC interface endpoints to establish connections for training and hosting. Modify any existing security groups that are associated with the VPC interface endpoint to allow only outbound connections for training and hosting.
Phân tích đúng: Như đã giải thích ở trên, interface endpoints cung cấp kết nối private đến S3 và SageMaker services (training/hosting) mà không cần internet. Security groups tinh chỉnh chỉ cho phép outbound cần thiết, đảm bảo an toàn tối đa. Đây là giải pháp chuẩn AWS cho VPC-only access. 🚀 -
❌ Create an IAM policy that prevents access the internet. Apply the IAM policy to an IAM role. Assign the IAM role to the SageMaker notebook instances in addition to any IAM roles that are already assigned to the instances.
Phân tích sai: IAM chỉ kiểm soát quyền truy cập tài nguyên AWS (như S3 buckets, SageMaker jobs), không kiểm soát network connectivity như internet access. IAM role không block được route table hoặc NAT gateway dẫn ra internet. Giải pháp này vô hiệu! 🔒❌ -
❌ Create VPC security groups to prevent all incoming and outgoing traffic. Assign the security groups to the SageMaker notebook instances.
Phân tích sai: Block tất cả incoming/outgoing traffic sẽ ngăn notebook instances truy cập S3 hoặc SageMaker services (cần outbound đến endpoints). Notebook không thể training/hosting dữ liệu từ S3. Security groups chỉ là stateful firewall, không thay thế được VPC endpoints cho private access. Quá cực đoan và phá hỏng chức năng! 🚫
📘 Tài liệu tham khảo (AWS docs cập nhật mới nhất 2024-2026)
- Amazon SageMaker Notebook Instances in a VPC – Hướng dẫn VPC endpoints cho no internet access.
- VPC Interface Endpoints for SageMaker – Chi tiết endpoints cho training/hosting.
- AWS PrivateLink for Amazon S3 – Kết nối S3 private.
- Best Practices: SageMaker Security in VPC – Blog AWS chính thức.
Giải pháp đúng đảm bảo zero internet trust cho SageMaker, phù hợp DevOps Professional! Nếu cần demo code Terraform/CloudFormation, hỏi thêm nhé! 🛠️
The ML engineer wants to use recall as the objective metric. The ML engineer also wants to expand the hyperparameter range for a new hyperparameter tuning job. The new hyperparameter range will include the range of the previously performed tuning job.
Which approach will run the new hyperparameter tuning job in the LEAST amount of time?
- A Use a warm start hyperparameter tuning job.
- B Use a checkpointing hyperparameter tuning job.
- C Use the same random seed for the hyperparameter tuning job.
- D Use multiple jobs in parallel for the hyperparameter tuning job.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào Amazon SageMaker Hyperparameter Tuning Job sử dụng Bayesian optimization để tối ưu hóa siêu tham số (hyperparameters). Kỹ sư ML ban đầu sử dụng precision làm chỉ số mục tiêu (objective metric). Bây giờ, họ muốn:
- Chuyển sang recall làm objective metric mới.
- Mở rộng phạm vi hyperparameter (hyperparameter range) cho job tuning mới, với phạm vi mới bao gồm toàn bộ phạm vi cũ (tức là mở rộng thêm nhưng vẫn giữ nguyên phần cũ).
- Mục tiêu: Chạy job tuning mới với thời gian ngắn nhất (LEAST amount of time).
🛠️ Lý do câu hỏi quan trọng: Trong SageMaker, hyperparameter tuning có thể tốn kém về thời gian và chi phí vì phải chạy nhiều training job để thử nghiệm các tổ hợp hyperparameters. Khi thay đổi metric hoặc range, việc tái sử dụng kết quả từ job cũ là chìa khóa để giảm thời gian, đặc biệt với Bayesian optimization (sử dụng mô hình surrogate để dự đoán điểm tốt dựa trên lịch sử trials).
📘 Kiến thức cập nhật (AWS 2026): SageMaker hỗ trợ Warm Start cho Bayesian optimization từ phiên bản mới nhất (re:Post và Developer Guide 2025-2026), cho phép job mới kế thừa trials từ parent jobs ngay cả khi metric hoặc range thay đổi (miễn range mới bao gồm range cũ).
✅ Đáp án đúng: Use a warm start hyperparameter tuning job
Lý do chọn đáp án này:
- Warm Start cho phép tạo một hyperparameter tuning job con (child job) từ parent job cũ, tái sử dụng toàn bộ lịch sử trials (bao gồm kết quả training, metrics) để khởi tạo mô hình Bayesian surrogate ngay lập tức.
- Với Bayesian optimization, SageMaker sẽ fit lại surrogate model dựa trên data cũ, chỉ chạy thêm trials mới cho phần range mở rộng hoặc metric mới (recall). Điều này giảm đáng kể số lượng trials cần chạy, dẫn đến thời gian ngắn nhất.
- Điều kiện phù hợp: Range mới bao gồm range cũ → trials cũ vẫn valid; metric thay đổi cũng được hỗ trợ (SageMaker remap metrics nếu cần).
- Lợi ích thời gian: Thay vì bắt đầu từ zero (cold start), warm start có thể tiết kiệm 50-80% thời gian tùy quy mô (dựa trên AWS benchmarks).
Dẫn nguồn:
- AWS SageMaker Developer Guide: Warm start hyperparameter tuning (cập nhật 2025).
- AWS re:Post: Best practices for warm start with Bayesian (2026 examples).
📋 Giải thích tất cả các phương án (đúng/sai)
-
✅ Use a warm start hyperparameter tuning job
(Đã giải thích chi tiết ở trên - Đây là cách tối ưu nhất để tái sử dụng lịch sử job cũ, giảm thời gian chạy đáng kể với Bayesian optimization). -
❌ Use a checkpointing hyperparameter tuning job
Checkpointing chỉ lưu trạng thái của một training job cá nhân (ví dụ: model weights giữa các epoch) để resume nếu job bị gián đoạn. Nó không áp dụng cho hyperparameter tuning job (không tái sử dụng trials giữa các tuning jobs khác nhau). Sử dụng cách này sẽ không tận dụng lịch sử cũ, dẫn đến thời gian chạy đầy đủ như cold start → KHÔNG giảm thời gian. -
❌ Use the same random seed for the hyperparameter tuning job
Random seed làm cho quá trình sampling hyperparameters deterministic (luôn chọn cùng trials nếu cùng seed). Tuy nhiên, nó không tái sử dụng kết quả training từ job cũ, vẫn phải chạy lại tất cả trials từ đầu với metric/range mới → Thời gian không giảm, thậm chí có thể lâu hơn do không học từ lịch sử. -
❌ Use multiple jobs in parallel for the hyperparameter tuning job
Parallel jobs tăng max parallel jobs (ví dụ: chạy 10 trials cùng lúc thay vì sequential), giúp giảm thời gian tổng thể cho cold start. Nhưng ở đây, nó không tái sử dụng lịch sử job cũ, vẫn phải khám phá toàn bộ space (kể cả phần cũ) → Không phải cách LEAST time so với warm start, vì warm start thông minh hơn với Bayesian.
🧠 Tóm tắt so sánh thời gian (dựa trên AWS docs):
Warm Start << Parallel Jobs < Same Seed/Checkpoint (gần cold start).
Nếu cần thực hành, dùng SageMaker Console → Create Tuning Job → Warm start mode! 🚀
The editors test the first version of the tool and report that the tool seems to look for word matches in general. The editors have to spend additional time to filter the results to look for the articles where the queried words are most important. A group of data scientists must redesign the tool so that it isolates the most frequently used words in a document. The tool also must capture the relevance and importance of words for each document in the corpus.
Which solution meets these requirements?
- A Extract the topics from each article by using Latent Dirichlet Allocation (LDA) topic modeling. Create a topic table by assigning the sum of the topic counts as a score for each word in the articles. Configure the tool to retrieve the articles where this topic count score is higher for the queried words.
- B Build a term frequency for each word in the articles that is weighted with the article's length. Build an inverse document frequency for each word that is weighted with all articles in the corpus. Define a final highlight score as the product of both of these frequencies. Configure the tool to retrieve the articles where this highlight score is higher for the queried words.
- C Download a pretrained word-embedding lookup table. Create a titles-embedding table by averaging the title's word embedding for each article in the corpus. Define a highlight score for each word as inversely proportional to the distance between its embedding and the title embedding. Configure the tool to retrieve the articles where this highlight score is higher for the queried words.
- D Build a term frequency score table for each word in each article of the corpus. Assign a score of zero to all stop words. For any other words, assign a score as the word’s frequency in the article. Configure the tool to retrieve the articles where this frequency score is higher for the queried words.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một công ty tin tức đang phát triển công cụ tìm kiếm bài báo dành cho biên tập viên. Công cụ cần tìm kiếm các bài báo phù hợp và đại diện nhất cho các từ khóa được query trong bộ sưu tập (corpus) các tài liệu tin tức lịch sử.
- Vấn đề hiện tại: Phiên bản đầu chỉ tìm kiếm dựa trên sự khớp từ thông thường (word matches), khiến biên tập viên phải lọc thủ công để tìm bài báo mà từ khóa thực sự quan trọng (most important).
- Yêu cầu redesign:
- Cách ly các từ được sử dụng thường xuyên nhất trong từng tài liệu (isolate the most frequently used words in a document).
- Bắt lấy sự liên quan và tầm quan trọng của từ đối với từng tài liệu trong corpus (capture the relevance and importance of words for each document).
Mục tiêu là xây dựng giải pháp tính toán điểm số nổi bật (highlight score) cho từ khóa trong từng bài báo, ưu tiên bài có từ khóa quan trọng cao (cao về tần suất trong bài nhưng hiếm trong corpus). Đây là bài toán IR (Information Retrieval) kinh điển, thường dùng TF-IDF trên AWS (ví dụ: Amazon OpenSearch Service, SageMaker, hoặc Kendra với custom ranking). Kiến thức cập nhật đến 2026: TF-IDF vẫn là nền tảng trong AWS OpenSearch (phiên bản 2.x) và SageMaker Processing cho feature engineering.
📘 Tài liệu tham khảo:
- AWS Docs: Amazon OpenSearch Service - TF-IDF Scoring (BM25 là variant của TF-IDF).
- SageMaker: Text Processing with TF-IDF.
- Kendra: Custom Document Enrichment hỗ trợ custom scoring.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng:
Build a term frequency for each word in the articles that is weighted with the article's length. Build an inverse document frequency for each word that is weighted with all articles in the corpus. Define a final highlight score as the product of both of these frequencies. Configure the tool to retrieve the articles where this highlight score is higher for the queried words.
Lý do:
- Đây chính là TF-IDF (Term Frequency - Inverse Document Frequency) 🛠️:
- TF (Term Frequency): Tần suất từ trong bài, điều chỉnh theo độ dài bài (tránh ưu tiên bài dài).
- IDF (Inverse Document Frequency): Đo độ hiếm của từ trong toàn corpus (từ phổ biến như "the" có IDF thấp).
- TF-IDF score = TF × IDF: Điểm cao khi từ thường xuyên trong bài NHƯNG hiếm trong corpus → chính xác đáp ứng yêu cầu "most important" và "relevance".
- Triển khai dễ trên AWS SageMaker (BlazingText/Processing) hoặc OpenSearch với script scoring. Phù hợp nhất cho search tool.
📋 Giải thích tất cả các phương án (Đúng/Sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên việc có đáp ứng cách ly từ quan trọng nhất và tầm quan trọng so với corpus không.
-
Phương án 1 ❌ (SAI):
Extract the topics from each article by using Latent Dirichlet Allocation (LDA) topic modeling. Create a topic table by assigning the sum of the topic counts as a score for each word in the articles. Configure the tool to retrieve the articles where this topic count score is higher for the queried words.
Giải thích sai: LDA là topic modeling (nhóm từ thành chủ đề), không tập trung vào tần suất từ đơn lẻ hay tầm quan trọng cụ thể của từ query. Sum topic counts chỉ cho điểm chủ đề chung, không isolate từ quan trọng nhất trong document và bỏ qua context corpus. Không hiệu quả cho word-level search, phức tạp hơn cần thiết (SageMaker LDA tồn tại nhưng overkill). -
Phương án 2 ✅ (ĐÚNG):
Build a term frequency for each word in the articles that is weighted with the article's length. Build an inverse document frequency for each word that is weighted with all articles in the corpus. Define a final highlight score as the product of both of these frequencies. Configure the tool to retrieve the articles where this highlight score is higher for the queried words.
Giải thích đúng: Như trên, TF-IDF chuẩn xác, cân bằng tần suất local (TF, weighted by length) và global rarity (IDF). Trực tiếp "capture relevance" và ưu tiên bài có từ query most representative. -
Phương án 3 ❌ (SAI):
Download a pretrained word-embedding lookup table. Create a titles-embedding table by averaging the title's word embedding for each article in the corpus. Define a highlight score for each word as inversely proportional to the distance between its embedding and the title embedding. Configure the tool to retrieve the articles where this highlight score is higher for the queried words.
Giải thích sai: Word embedding (như Word2Vec/BERT từ Hugging Face trên SageMaker) đo semantic similarity, chỉ so sánh từ query với title embedding (không dùng nội dung bài). Bỏ qua tần suất trong document và corpus-wide importance, chỉ tốt cho title matching chứ không "isolate most frequently used words" hay relevance toàn bài. -
Phương án 4 ❌ (SAI):
Build a term frequency score table for each word in each article of the corpus. Assign a score of zero to all stop words. For any other words, assign a score as the word’s frequency in the article. Configure the tool to retrieve the articles where this frequency score is higher for the queried words.
Giải thích sai: Chỉ là TF đơn thuần (bỏ stop words), không có IDF → ưu tiên bài có từ query xuất hiện nhiều, nhưng không phạt từ phổ biến toàn corpus (ví dụ: "news" xuất hiện khắp nơi vẫn top). Không "capture importance relative to corpus", giống vấn đề ban đầu (chỉ word matches cao tần suất).
Kết luận 🎯: TF-IDF là giải pháp tối ưu, dễ scale trên AWS với chi phí thấp. Nếu triển khai, dùng Lambda + OpenSearch cho real-time query!
A machine learning (ML) specialist must develop a solution to achieve high availability. The solution must have a recovery time objective (RTO) of 5 minutes.
Which solution will meet these requirements with the LEAST effort?
- A Deploy multiple instances for each endpoint in a VPC that spans at least two Regions.
- B Use the SageMaker auto scaling feature for the hosted recommendation models.
- C Deploy multiple instances for each production endpoint in a VPC that spans least two subnets that are in a second Availability Zone.
- D Frequently generate backups of the production recommendation model. Deploy the backups in a second Region.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh một công ty đang phát triển hệ thống khuyến nghị ML (machine learning recommendation system) sử dụng Amazon SageMaker hosting services, hiện đang triển khai model trong một Availability Zone (AZ) duy nhất thuộc một Region AWS. Họ có KPI uptime quan trọng (business-critical), và cần giải pháp high availability (HA) với Recovery Time Objective (RTO) chỉ 5 phút.
Yêu cầu chính: Giải pháp phải đạt HA với ít nỗ lực nhất (LEAST effort).
🛠️ Bối cảnh kỹ thuật: SageMaker endpoints cho phép host model inference với multi-instance để scale và HA. Theo tài liệu AWS mới nhất (2024-2026), SageMaker hỗ trợ multi-AZ deployment tự động failover nhanh (RTO <5 phút) bằng cách phân bố instances qua các subnets ở nhiều AZ trong cùng Region, mà không cần code phức tạp hay multi-region.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Deploy multiple instances for each production endpoint in a VPC that spans least two subnets that are in a second Availability Zone.
Lý do:
- Phương án này tận dụng tính năng multi-instance endpoints của SageMaker, triển khai instances qua ít nhất 2 subnets ở 2 AZ khác nhau trong cùng VPC và Region. SageMaker tự động load balance và failover giữa các AZ nếu một AZ fail, đạt RTO dưới 5 phút (thường chỉ vài giây đến phút).
- Least effort: Chỉ cần cấu hình endpoint với subnets multi-AZ khi create/update (qua Console, CLI hoặc SDK), không cần backup, multi-region hay custom code. Đây là best practice AWS cho SageMaker HA theo AWS Well-Architected Framework (Reliability pillar).
- ✅ Hiệu quả cao: Đảm bảo uptime >99.99% mà không di chuyển data/model cross-region.
📋 Phân tích tất cả các phương án (đúng/sai)
-
❌ Phương án SAI: Deploy multiple instances for each endpoint in a VPC that spans at least two Regions.
Giải thích sai: VPC không thể span qua nhiều Region (VPC là tài nguyên regional). Multi-region yêu cầu multi-account hoặc cross-region replication phức tạp, tăng effort lớn (setup VPC peering/Global Accelerator), và RTO vượt 5 phút do latency replication model artifacts. Không phải least effort và không khả thi trực tiếp. -
❌ Phương án SAI: Use the SageMaker auto scaling feature for the hosted recommendation models.
Giải thích sai: Auto scaling chỉ scale instances horizontally trong cùng AZ dựa trên metrics (CPU/Memory/Invocations), không đảm bảo HA cross-AZ. Nếu AZ fail, toàn bộ endpoint down → RTO cao. Phải kết hợp multi-AZ mới HA, nên không đủ yêu cầu và không least effort cho uptime KPI. -
✅ Phương án ĐÚNG: Deploy multiple instances for each production endpoint in a VPC that spans least two subnets that are in a second Availability Zone.
Giải thích đúng: Như đã nêu ở phần đáp án. SageMaker tự động phân bố instances đều qua subnets multi-AZ, hỗ trợ Production Variants với min/max instances. Failover seamless qua Application Load Balancer (ALB) nội bộ, RTO <5 phút. Least effort: Chỉ update endpoint config (e.g.,create_endpoint_configvớiSubnetsmulti-AZ). -
❌ Phương án SAI: Frequently generate backups of the production recommendation model. Deploy the backups in a second Region.
Giải thích sai: Backup model (qua S3 artifacts) và deploy second Region yêu cầu manual failover (update endpoint DNS/traffic), replication data, và cold start instances → RTO >5 phút (thường 10-30 phút). Effort cao (scripting, monitoring, cross-region costs), không tự động như multi-AZ. Phù hợp disaster recovery (RPO/RTO dài), không phải HA intra-region.
📘 Tài liệu tham khảo (kiến thức AWS cập nhật 2024-2026)
- Amazon SageMaker Documentation: Host models with multi-instance endpoints for high availability & Deploy models across multiple AZs → Xác nhận multi-AZ subnets cho RTO thấp.
- AWS Well-Architected Framework (Reliability Pillar): SageMaker HA best practices.
- AWS re:Post & Blogs: Achieving high availability with SageMaker (2024 updates hỗ trợ Inference Recommender cho multi-AZ).
- Exam Tips (DOPE-C01): Câu hỏi kiểu này thường test multi-AZ vs multi-region cho least effort HA.
🛠️ Khuyến nghị triển khai nhanh: Sử dụng AWS Console → SageMaker → Endpoints → Create endpoint config → Chọn subnets ở 2+ AZ → Deploy production variant với MultiInstanceConfig. Test failover bằng Chaos Engineering (AWS Fault Injection Simulator)!
A machine learning (ML) specialist wants to build an automated document processing workflow to extract text from specific fields from the documents and to classify the documents. The ML specialist wants a solution that requires low maintenance.
Which solution will meet these requirements with the LEAST operational effort?
- A Use a PaddleOCR model in Amazon SageMaker to detect and extract the required text and fields. Use a SageMaker text classification model to classify the document.
- B Use a PaddleOCR model in Amazon SageMaker to detect and extract the required text and fields. Use Amazon Comprehend to classify the document.
- C Use Amazon Textract to detect and extract the required text and fields. Use Amazon Rekognition to classify the document.
- D Use Amazon Textract to detect and extract the required text and fields. Use Amazon Comprehend to classify the document.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi
📘 Nội dung câu hỏi:
Câu hỏi mô tả một công ty toàn cầu nhận và xử lý hàng trăm tài liệu hàng ngày dưới dạng PDF in ấn hoặc JPG. Chuyên gia ML muốn xây dựng workflow tự động để trích xuất văn bản từ các trường cụ thể (như form fields) và phân loại tài liệu. Yêu cầu chính là giải pháp ít nỗ lực vận hành nhất (low maintenance), nghĩa là ưu tiên các dịch vụ managed hoàn toàn, không cần tự train model hay quản lý infrastructure.
✅ Đây là tình huống điển hình cho OCR (Optical Character Recognition) kết hợp NLP classification, phù hợp với các dịch vụ AWS serverless như Textract (extract text/forms) và Comprehend (classify documents).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Amazon Textract to detect and extract the required text and fields. Use Amazon Comprehend to classify the document.
🛠️ Lý do chi tiết:
- Amazon Textract là dịch vụ managed chuyên trích xuất text, forms, tables từ tài liệu scanned (PDF/JPG), tự động detect fields cụ thể mà không cần training model. Nó xử lý hàng trăm documents dễ dàng với API serverless, scale tự động.
- Amazon Comprehend là dịch vụ NLP managed cho phân loại tài liệu (custom classification), hỗ trợ train nhanh trên dữ liệu labeled với ít effort, và inference serverless.
- Least operational effort: Cả hai đều fully managed, không cần deploy endpoint, monitor model, hay update PaddleOCR/SageMaker. Phù hợp workflow low-maintenance cho high-volume documents.
📘 Tài liệu tham khảo: AWS Textract Documentation (2024-2026) & Amazon Comprehend Custom Classification.
📋 Phân tích tất cả các phương án
-
❌ Use a PaddleOCR model in Amazon SageMaker to detect and extract the required text and fields. Use a SageMaker text classification model to classify the document.
Giải thích sai: PaddleOCR (model OCR open-source) cần deploy/train trong SageMaker, đòi hỏi effort cao như quản lý endpoint, versioning model, auto-scaling, và monitoring. SageMaker text classification cũng yêu cầu training từ scratch. Không phải low-maintenance so với Textract/Comprehend managed. -
❌ Use a PaddleOCR model in Amazon SageMaker to detect and extract the required text and fields. Use Amazon Comprehend to classify the document.
Giải thích sai: Phần OCR dùng PaddleOCR trên SageMaker vẫn tốn effort deploy/maintain model (custom container, inference endpoint). Comprehend tốt cho classify nhưng tổng thể không least effort vì SageMaker overhead. -
❌ Use Amazon Textract to detect and extract the required text and fields. Use Amazon Rekognition to classify the document.
Giải thích sai: Textract đúng cho extract text/fields. Nhưng Amazon Rekognition chuyên computer vision (detect objects, labels, faces trong images/videos), không phù hợp classify text-based documents (nó không xử lý semantic classification như Comprehend). Sử dụng Rekognition sẽ kém chính xác và cần custom logic phức tạp. -
✅ Use Amazon Textract to detect and extract the required text and fields. Use Amazon Comprehend to classify the document.
Giải thích đúng: Như đã phân tích ở trên, đây là combo managed services lý tưởng, serverless, auto-scale, hỗ trợ high-volume với zero maintenance cho model deployment. Hoàn hảo cho yêu cầu low operational effort đến 2026 (không thay đổi lớn trong AWS ML services).
🎯 Kết luận: Giải pháp đúng tận dụng fully managed AWS AI services để giảm thiểu ops burden, phù hợp chứng chỉ DevOps Engineer Pro (tập trung automation & least effort).
Which metrics should the data scientist use to optimize the classifier? (Choose two.)
- A Specificity
- B False positive rate
- C Accuracy
- D F1 score
- E True positive rate
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi này thuộc lĩnh vực Machine Learning trên AWS, cụ thể là đánh giá mô hình phân loại (classifier) trong bối cảnh phát hiện gian lận thẻ tín dụng (credit card fraud detection). 🛡️️
- Bối cảnh vấn đề: Công ty có dữ liệu giao dịch thẻ tín dụng với tỷ lệ gian lận trung bình chỉ 2% (dữ liệu mất cân bằng - imbalanced dataset). Mô hình được huấn luyện trên dữ liệu một năm, và mục tiêu chính là xác định chính xác càng nhiều giao dịch gian lận càng tốt (capture as many fraudulent transactions as possible).
- Yêu cầu: Chọn hai metrics để data scientist sử dụng nhằm tối ưu hóa mô hình. Điều này nhấn mạnh vào việc ưu tiên giảm thiểu bỏ sót gian lận (false negatives) vì tỷ lệ gian lận thấp, không phải ưu tiên độ chính xác tổng quát.
- Liên quan AWS: Trong Amazon SageMaker (phiên bản cập nhật 2026 với SageMaker Studio và Clarify cho bias detection), các metrics này được sử dụng trong model evaluation qua SageMaker Model Monitor hoặc Clarify để xử lý imbalanced data. Fraud detection thường dùng binary classification với threshold tuning. 📘
Tài liệu tham khảo:
- AWS SageMaker Documentation: Model evaluation metrics (cập nhật 2025-2026).
- AWS ML Specialty Exam Guide (MLS-C01/DOP-C02): Nhấn mạnh recall/F1 cho fraud detection.
✅ Đáp án đúng và lý do lựa chọn
Hai metrics đúng là: F1 score và True positive rate.
- Lý do chọn True positive rate (hay Recall): 🏆 Với tỷ lệ gian lận chỉ 2%, mô hình cần tối đa hóa việc phát hiện true positives (các giao dịch gian lận thực sự). True positive rate = TP / (TP + FN), giúp giảm false negatives (bỏ sót gian lận). Đây là metric ưu tiên cho "capture as many fraudulent transactions as possible".
- Lý do chọn F1 score: 🏆 F1 = 2 * (Precision * Recall) / (Precision + Recall), cân bằng giữa Precision (giảm false positives) và Recall. Hoàn hảo cho imbalanced data, tránh overfit vào majority class (99% non-fraud). SageMaker hỗ trợ F1 trực tiếp trong hyperparameter tuning.
🔍 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm lý do bằng tiếng Việt rõ ràng:
-
Specificity ❌
Sai: Specificity = TN / (TN + FP), đo tỷ lệ phát hiện đúng giao dịch hợp lệ (non-fraud). Với 98% dữ liệu là non-fraud, metric này ưu tiên majority class, không giúp "capture fraudulent transactions". Không phù hợp mục tiêu chính. -
False positive rate ❌
Sai: False positive rate = FP / (FP + TN) = 1 - Specificity, đo tỷ lệ nhầm lẫn giao dịch hợp lệ thành gian lận. Giảm FPR có thể làm giảm recall (bỏ sót gian lận thực), trái với yêu cầu capture tối đa fraud. Không ưu tiên cho imbalanced fraud detection. -
Accuracy ❌
Sai: Accuracy = (TP + TN) / Total, đo độ chính xác tổng quát. Với 2% fraud, mô hình đoán tất cả là non-fraud vẫn đạt ~98% accuracy – lừa dối (misleading). AWS khuyến cáo tránh accuracy cho imbalanced data trong SageMaker best practices. -
F1 score ✅
Đúng: Harmonic mean của Precision và Recall, lý tưởng cho imbalanced binary classification. Giúp tối ưu threshold để capture fraud mà không flood false positives. SageMaker tích hợp F1 trong built-in algorithms như XGBoost cho fraud detection. -
True positive rate ✅
Đúng: Còn gọi là Recall/Sensitivity, trực tiếp đo khả năng bắt được tất cả fraud (TP / (TP + FN)). Ưu tiên cao nhất cho use case này, AWS dùng trong ROC curve và PR curve evaluation trên SageMaker Processing Jobs.
Lời khuyên thực hành trên AWS 🛠️: Sử dụng SageMaker Debugger hoặc Model Monitor để track metrics thời gian thực, kết hợp SMOTE hoặc class weights cho imbalanced data. Threshold tuning qua ROC-AUC để max Recall/F1! 🚀
Which solution will meet these requirements?
- A Amazon S3 with S3 Cross-Region Replication (CRR)
- B Amazon Elastic Block Store (Amazon EBS) with snapshots that are shared in a secondary Region
- C Amazon Elastic File System (Amazon EFS) Standard storage that is configured with Regional availability
- D AWS Storage Gateway Volume Gateway
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một data scientist đang thiết kế một repository (kho lưu trữ) chứa nhiều hình ảnh xe hơi. Yêu cầu chính bao gồm:
- Tự động scale kích thước để lưu trữ hình ảnh mới hàng ngày (không cần can thiệp thủ công).
- Hỗ trợ versioning (phiên bản hóa) cho các hình ảnh.
- Duy trì nhiều bản sao accessible ngay lập tức (immediately accessible) của dữ liệu ở nhiều AWS Regions khác nhau (để đảm bảo tính sẵn sàng cao và disaster recovery).
📘 Bối cảnh AWS: Đây là bài toán về object storage phù hợp với lưu trữ không cấu trúc như hình ảnh, cần scale vô hạn, versioning native, và replication cross-region. Kiến thức cập nhật đến 2026: AWS S3 vẫn là lựa chọn chuẩn với CRR hỗ trợ replication asynchronous/sync (qua S3 Replication Time Control - RTC) cho độ trễ thấp (<15 phút SLA).
✅ Đáp án đúng: Amazon S3 with S3 Cross-Region Replication (CRR)
Lý do lựa chọn:
- Amazon S3 tự động scale vô hạn (không giới hạn kích thước), lý tưởng cho repository hình ảnh tăng hàng ngày 🛠️.
- S3 Versioning native hỗ trợ versioning đầy đủ (giữ tất cả phiên bản object).
- CRR replicate objects (bao gồm versions, metadata) sang bucket khác ở Region khác, đảm bảo multiple copies immediately accessible (sau replication, objects có thể đọc ngay từ bucket đích).
- Hoàn hảo cho yêu cầu: Scale auto, versioning, multi-region durability (99.999999999% theo AWS 2026 specs).
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh:
-
✅ Amazon S3 with S3 Cross-Region Replication (CRR)
🟢 Đúng vì: S3 là object storage scale tự động, versioning built-in (enable per bucket). CRR replicate real-time/asynchronously cross-region, hỗ trợ versioned objects, delete markers. Với S3 RTC (2026 update), đảm bảo 99.99% objects replicate trong 15 phút → immediately accessible. Phù hợp 100% cho image repo. -
❌ Amazon Elastic Block Store (Amazon EBS) with snapshots that are shared in a secondary Region
🔴 Sai vì: EBS là block storage gắn với EC2 instance (không phải repository object), không scale tự động cho hàng ngàn images mới hàng ngày (cần resize volume thủ công). Snapshots là point-in-time backup, phải copy thủ công sang Region khác (không immediate, mất thời gian), không hỗ trợ versioning native cho objects, và không accessible ngay như live data. -
❌ Amazon Elastic File System (Amazon EFS) Standard storage that is configured with Regional availability
🔴 Sai vì: EFS là file storage (NFS), scale auto nhưng chỉ multi-AZ trong 1 Region (Regional availability ≠ cross-region). Không hỗ trợ versioning objects, replication cross-region cần DMS/EC2 thủ công (không immediate). Không tối ưu cho image repo lớn (expensive hơn S3 cho objects). -
❌ AWS Storage Gateway Volume Gateway
🔴 Sai vì: Đây là hybrid storage (cache local on-prem, upload S3/FSx), không phải pure cloud repository scale auto hàng ngày. Volume mode là iSCSI block, không versioning native, replication cross-region không immediate (phụ thuộc gateway). Không phù hợp cho data scientist cloud-native.
📚 Tài liệu tham khảo (AWS cập nhật 2026)
- Amazon S3 Cross-Region Replication → Chi tiết CRR + Versioning.
- S3 Replication Time Control (RTC) → SLA immediate access.
- AWS Storage Services Whitepaper → So sánh S3/EBS/EFS/Gateway.
- Exam guide DOP-C02: S3 cho scalable object storage với multi-region DR.
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm case study, hỏi nhé!
Which solution will meet these requirements with the LEAST operational overhead?
- A Create a new SageMaker endpoint for the new model. Configure an Application Load Balancer (ALB) to distribute traffic between the old model and the new model.
- B Modify the existing endpoint to use SageMaker production variants to distribute traffic between the old model and the new model.
- C Modify the existing endpoint to use SageMaker batch transform to distribute traffic between the old model and the new model.
- D Create a new SageMaker endpoint for the new model. Configure a Network Load Balancer (NLB) to distribute traffic between the old model and the new model.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi xoay quanh một công ty thương mại điện tử muốn cập nhật API engine khuyến nghị ML thời gian thực (real-time) đang chạy trên Amazon SageMaker trong môi trường production. Các yêu cầu chính bao gồm:
- Triển khai model mới mà không cần thay đổi ứng dụng đang gọi API (tức là giữ nguyên endpoint URL).
- Đánh giá hiệu suất model mới bằng cách sử dụng traffic production thực tế trước khi rollout toàn bộ (gợi ý về A/B testing hoặc canary deployment).
- Giải pháp phải có operational overhead thấp nhất (least operational overhead), nghĩa là giảm thiểu công sức quản lý, cấu hình và tài nguyên.
📘 Bối cảnh AWS SageMaker (cập nhật đến 2026): SageMaker endpoints hỗ trợ hosting model thời gian thực với tính năng Production Variants (từ phiên bản SageMaker 2019 và vẫn là best practice mới nhất), cho phép chạy nhiều phiên bản model song song trên cùng một endpoint, phân bổ traffic linh hoạt (traffic splitting) mà không cần load balancer bên ngoài. Điều này lý tưởng cho blue-green hoặc canary deployments với zero-downtime và minimal changes.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Modify the existing endpoint to use SageMaker production variants to distribute traffic between the old model and the new model.
Lý do:
🛠️ Phương án này sử dụng Production Variants của SageMaker – tính năng native cho phép host nhiều model variants trên cùng một endpoint, tự động split traffic (ví dụ: 90% old model, 10% new model) qua tham số VariantWeight. Không cần tạo endpoint mới, không cần load balancer ngoài, và không thay đổi URL endpoint → ứng dụng client không cần chỉnh sửa.
✅ Least operational overhead: Chỉ cần update endpoint config (qua AWS Console, CLI hoặc SDK), SageMaker tự scale và monitor metrics (latency, error rate) qua CloudWatch. Hỗ trợ rollback nhanh bằng cách điều chỉnh weight. Phù hợp hoàn hảo cho đánh giá production traffic với real-time inference.
(Đây là best practice theo AWS Well-Architected Framework cho ML workloads.)
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Tôi đánh dấu ✅ (đúng) hoặc ❌ (sai) kèm giải thích bằng tiếng Việt:
-
✅ Modify the existing endpoint to use SageMaker production variants to distribute traffic between the old model and the new model.
🛠️ Đúng vì: Như đã giải thích ở trên, đây là giải pháp native, zero-change cho client apps, hỗ trợ traffic splitting tự động và monitoring tích hợp. Overhead thấp nhất, không cần infra ngoài. -
❌ Create a new SageMaker endpoint for the new model. Configure an Application Load Balancer (ALB) to distribute traffic between the old model and the new model.
🚫 Sai vì: Tạo endpoint mới yêu cầu thay đổi client apps để gọi ALB DNS thay vì SageMaker endpoint trực tiếp. ALB không native hỗ trợ SageMaker endpoints (cần custom target groups với IP/hostname), tăng overhead: quản lý ALB rules, health checks, scaling policies riêng. Không hiệu quả cho real-time ML traffic và vi phạm yêu cầu "không thay đổi applications". -
❌ Modify the existing endpoint to use SageMaker batch transform to distribute traffic between the old model and the new model.
🚫 Sai vì: Batch Transform chỉ dùng cho batch inference (xử lý dữ liệu lớn offline, không real-time). Không hỗ trợ traffic distribution cho production API thời gian thực. Sử dụng sai tính năng sẽ gây latency cao, không phù hợp đánh giá real-time performance. -
❌ Create a new SageMaker endpoint for the new model. Configure a Network Load Balancer (NLB) to distribute traffic between the old model and the new model.
🚫 Sai vì: Tương tự ALB, tạo endpoint mới buộc thay đổi client apps để route qua NLB. NLB chỉ layer 4 (TCP/UDP), kém linh hoạt cho HTTP API so với ALB, và vẫn cần quản lý target groups phức tạp cho SageMaker endpoints. Overhead cao hơn Production Variants rất nhiều.
📚 Tài liệu tham khảo (AWS cập nhật mới nhất 2026)
- SageMaker Production Variants: AWS Docs - Host Multiple Model Versions – Chi tiết traffic splitting và A/B testing.
- SageMaker Endpoints Best Practices: AWS ML Best Practices – Khuyến nghị dùng variants cho low-overhead deployments.
- CloudWatch Metrics cho SageMaker: Monitoring Endpoints.
🧑💼 Lời khuyên từ AWS Certified DevOps Engineer Professional: Ưu tiên serverless/native features như Production Variants để giảm toil, tích hợp với SageMaker Pipelines cho CI/CD full ML lifecycle!
The ML specialist develop a solution by using Amazon SageMaker DeepAR to account for the missing values in the training dataset.
Which approach will meet these requirements with the LEAST development effort?
- A Impute the missing values by using the linear regression method. Use the entire dataset and the imputed values to train the DeepAR model.
- B Replace the missing values with not a number (NaN). Use the entire dataset and the encoded missing values to train the DeepAR model.
- C Impute the missing values by using a forward fill. Use the entire dataset and the imputed values to train the DeepAR model.
- D Impute the missing values by using the mean value. Use the entire dataset and the imputed values to train the DeepAR model.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi
📘 Nội dung câu hỏi:
Câu hỏi xoay quanh một chuyên gia ML tại công ty sản xuất sử dụng Amazon SageMaker DeepAR để dự báo nhu cầu nguyên liệu đầu vào và năng lượng. Dataset huấn luyện được lưu dưới dạng file JSON, và hầu hết dữ liệu trong target variable (biến mục tiêu) bị thiếu (missing values). Nhiệm vụ là phát triển giải pháp xử lý missing values cho DeepAR với ít nỗ lực phát triển nhất (LEAST development effort).
DeepAR là thuật toán thời gian chuỗi (time-series forecasting) trong SageMaker, chuyên dự báo dựa trên dữ liệu lịch sử có thể có missing values. Vấn đề chính là cách encode missing values trong dataset JSON để DeepAR xử lý tự động mà không cần tiền xử lý phức tạp. AWS khuyến nghị sử dụng định dạng JSON Lines với các trường như start, target, cat (categorical), và dynamic_feat (features động), nơi target có thể chứa NaN để biểu thị missing values.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Replace the missing values with not a number (NaN). Use the entire dataset and the encoded missing values to train the DeepAR model.
Lý do:
🛠️ DeepAR hỗ trợ native (tích hợp sẵn) xử lý missing values trong target variable bằng cách sử dụng NaN (Not a Number) trực tiếp trong dataset JSON. Bạn chỉ cần thay thế missing values bằng NaN và train model trên toàn bộ dataset mà không cần code thêm imputation. Điều này đảm bảo LEAST development effort vì không yêu cầu thư viện bên ngoài (như scikit-learn, pandas) hay logic tùy chỉnh. SageMaker DeepAR tự động bỏ qua NaN trong quá trình huấn luyện và dự báo, giữ nguyên tính toàn vẹn dữ liệu thời gian chuỗi.
✅ Phương pháp này phù hợp với best practices AWS (cập nhật đến 2026), tránh bias từ imputation thủ công.
🔍 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên tính khả thi, effort và tương thích với DeepAR:
-
❌ [SAI] Impute the missing values by using the linear regression method. Use the entire dataset and the imputed values to train the DeepAR model.
Phương án này yêu cầu phát triển thêm code phức tạp để train một model linear regression riêng (sử dụng scikit-learn hoặc SageMaker Linear Learner) nhằm điền missing values. Effort cao vì cần xử lý dữ liệu thời gian chuỗi, có thể gây bias (linear không phù hợp với pattern phi tuyến của DeepAR), và DeepAR không cần imputation vì hỗ trợ NaN native. Không phải LEAST effort. -
✅ [ĐÚNG] Replace the missing values with not a number (NaN). Use the entire dataset and the encoded missing values to train the DeepAR model.
Như đã giải thích ở trên: Tích hợp sẵn, zero-effort thêm, DeepAR tự handle NaN trong target (xem docs AWS). Sử dụng toàn bộ dataset mà không mất dữ liệu. -
❌ [SAI] Impute the missing values by using a forward fill. Use the entire dataset and the imputed values to train the DeepAR model.
Forward fill (ffill trong pandas) giả định giá trị trước đó tiếp tục, phù hợp dữ liệu liên tục nhưng yêu cầu code pandas preprocessing trước khi upload S3. Effort trung bình-cao, có thể distort pattern thời gian (ví dụ: missing do sự cố sản xuất không nên "kéo dài" giá trị cũ). DeepAR không cần vì NaN tốt hơn. -
❌ [SAI] Impute the missing values by using the mean value. Use the entire dataset and the imputed values to train the DeepAR model.
Sử dụng mean imputation đơn giản nhưng vẫn cần code (pandas fillna(mean)), gây underestimation variance trong time-series (mean không capture seasonality/trend). Effort không thấp, và DeepAR khuyến cáo tránh vì NaN cho kết quả chính xác hơn.
📚 Tài liệu tham khảo
- AWS SageMaker Developer Guide - DeepAR Algorithm: DeepAR Forecasting Algorithm (cập nhật 2024-2026): Xác nhận "The target values can contain missing values, which are indicated by NaN."
- SageMaker Data Formats: JSON Lines Input Format – Target hỗ trợ NaN.
- Best Practices Time-Series: AWS re:Post và Well-Architected ML Lens (2025 edition).
🛡️ Lời khuyên DevOps: Trong pipeline CI/CD (CodePipeline + SageMaker Pipelines), tích hợp bước encode NaN tự động để scale production!