Ngân hàng đề — AWS Certified Machine Learning Engineer Associate

Tìm thấy 635 câu.

Câu 221 Deployment and Orchestration of ML Workflows

A financial services company is using Amazon SageMaker to develop a machine learning model for fraud detection. The dataset contains transaction records, customer profiles, and behavioral data, stored in Amazon S3. The data is updated in real-time, and new fraudulent patterns emerge frequently, requiring regular model updates. To ensure that the model remains effective, the company wants to automate the model retraining process based on new data, detect data drift in real-time, and perform hyperparameter tuning for optimal performance. Additionally, the company needs to ensure that model predictions are explainable to regulators.

Which is the most appropriate solution to meet these requirements?

  1. A

    Use SageMaker Model Monitor to detect data drift, SageMaker Pipelines to automate the retraining process, and SageMaker Clarify to explain model predictions.

  2. B

    Use SageMaker Clarify to detect data drift, SageMaker Feature Store to manage real-time feature updates, and SageMaker Autopilot to retrain the model based on new data.

  3. C

    Use SageMaker Feature Store to automatically update features as new data comes in, SageMaker Model Monitor to automate model retraining, and SageMaker Autopilot to handle hyperparameter tuning.

  4. D

    Use SageMaker Data Wrangler to preprocess new data, SageMaker Pipelines to automate retraining and hyperparameter tuning, and SageMaker Ground Truth to explain model predictions.

Xem giải thích

Đáp án

A — Model Monitor phát hiện trôi dữ liệu, Pipelines tự động hoá việc huấn luyện lại

Vì sao đúng

Cặp này tạo thành vòng khép kín cho mô hình chạy lâu dài: Model Monitor so dữ liệu thật với nền và phát hiện khi phân bố trôi đi; Pipelines đóng gói toàn bộ chuỗi tiền xử lý – huấn luyện – đánh giá – đăng ký thành một quy trình chạy lại được, kích hoạt được từ cảnh báo. Không có Pipelines thì mỗi lần huấn luyện lại là một lần làm tay.

Vì sao các phương án khác sai

  • B. Clarify để phát hiện trôi dữ liệu — Clarify lo thiên lệch và khả năng giải thích; phần phát hiện trôi là của Model Monitor.
  • C. Feature Store tự cập nhật đặc trưng — nó lưu và phục vụ đặc trưng, không tự nhận ra mô hình đang kém đi.
  • D. Data Wrangler tiền xử lý — công cụ chuẩn bị dữ liệu, thiếu hẳn phần phát hiện trôi.
Câu 222 Chọn nhiều đáp án Data Preparation for Machine Learning

A machine learning engineer is using SageMaker Data Wrangler to understand relationships between features in the dataset. The engineer wants to visualize the distribution of numerical data and identify relationships between features. Which of the following visualization methods should be used in Data Wrangler to achieve these objectives? (Choose TWO)

  1. A

    Partial Dependence Plots (PDP) for feature impact

  2. B

    Confusion matrix for evaluating model predictions

  3. C

    ROC curves for model evaluation

  4. D

    Histograms for understanding data distribution

  5. E

    Line plots for analyzing trends over time

Xem giải thích

Đáp án

D và E — biểu đồ tần suất, và biểu đồ đường

Vì sao đúng

Đề hỏi về khám phá dữ liệu, không phải đánh giá mô hình:

  • D. Histogram cho thấy phân bố của một biến số — lệch trái hay phải, có bao nhiêu đỉnh, giá trị dị thường nằm đâu. Đây là biểu đồ đầu tiên nên xem với biến số.
  • E. Biểu đồ đường cho thấy xu hướng theo thời gian, thứ mà histogram không thể hiện được.

Vì sao các phương án khác sai

  • A. Partial Dependence Plot — cho biết mô hình phản ứng ra sao với một đặc trưng; cần có mô hình đã huấn luyện.
  • B. Ma trận nhầm lẫn và C. Đường ROC — đều là công cụ đánh giá mô hình phân loại, dùng sau khi huấn luyện chứ không phải lúc tìm hiểu dữ liệu.
Câu 223 ML Solution Monitoring, Maintenance, and Security

A company wants to ensure that all its EC2 instances are monitored for both CPU utilization and disk read/write operations. They also need to track configuration changes to these EC2 instances over time, such as changes to security groups, instance type, and AMI IDs. Which combination of AWS services will best meet these requirements?

  1. A

    Enable AWS CloudTrail to monitor CPU and disk usage and use AWS Config to capture configuration changes.

  2. B

    Set up CloudWatch Alarms for CPU and disk operations and AWS CloudTrail for tracking configuration changes on EC2 instances.

  3. C

    Use AWS CloudWatch to monitor CPU utilization and disk read/write operations, and AWS Config to track configuration changes to the EC2 instances.

  4. D

    Use CloudWatch Logs to capture both CPU utilization and configuration changes and create custom log filters for disk operations.

Xem giải thích

Đáp án

C — CloudWatch cho chỉ số CPU và đọc ghi đĩa, AWS Config cho thay đổi cấu hình

Vì sao đúng

Hai loại câu hỏi khác nhau nên cần hai dịch vụ: CloudWatch thu chỉ số vận hành theo thời gian và đặt cảnh báo khi vượt ngưỡng. Config chụp lại trạng thái cấu hình của instance theo thời gian, nên trả lời được "cấu hình này đã đổi lúc nào, từ gì sang gì".

Vì sao các phương án khác sai

  • A. CloudTrail theo dõi CPU và đĩa — CloudTrail ghi lời gọi API, không thu chỉ số hiệu năng.
  • B. CloudTrail theo dõi thay đổi cấu hình — nó ghi lời gọi đã tạo ra thay đổi, nhưng không giữ trạng thái cấu hình để so trước sau như Config.
  • D. CloudWatch Logs bắt cả thay đổi cấu hình — Logs chứa log ứng dụng, không tự biết cấu hình tài nguyên.
Câu 224 Deployment and Orchestration of ML Workflows

An ML engineer is using SageMaker Debugger's profiling feature to optimize resource usage during training. They notice that the GPU utilization is very low, causing longer training times. Which of the following actions could help improve GPU utilization and optimize the training process?

  1. A

    Use SageMaker Debugger to monitor for vanishing gradients, which will improve GPU utilization automatically.

  2. B

    Adjust the batch size and data loading pipeline to ensure that the GPU is being used efficiently throughout the training.

  3. C

    Migrate the model to CPU-only instances to avoid GPU underutilization.

  4. D

    Use SageMaker Neo to optimize the model for better GPU performance during training.

Xem giải thích

Đáp án

B — Chỉnh kích thước lô và đường ống nạp dữ liệu để GPU không phải chờ

Vì sao đúng

GPU dùng thấp gần như luôn có nghĩa là GPU đang đói dữ liệu: khâu đọc và tiền xử lý trên CPU không theo kịp tốc độ tính toán, nên GPU ngồi không giữa các lô. Hai nút chỉnh trực tiếp là tăng kích thước lô để mỗi lần tính được nhiều hơn, và làm nhanh đường nạp — tăng số luồng đọc, nạp trước, hoặc đổi sang định dạng đọc nhanh hơn.

Vì sao các phương án khác sai

  • A. Theo dõi gradient biến mất — vấn đề về hội tụ của mô hình, không liên quan tới mức dùng GPU.
  • C. Chuyển sang máy chỉ có CPU — bỏ cuộc thay vì sửa; huấn luyện sẽ chậm hơn nhiều.
  • D. Dùng SageMaker Neo — tối ưu mô hình sau khi huấn luyện xong để triển khai, không ảnh hưởng gì tới lúc huấn luyện.
Câu 225 Chọn nhiều đáp án ML Model Development

You are working on a machine learning project where your model frequently encounters issues such as vanishing gradients and overfitting during training. You decide to use Amazon SageMaker Debugger to monitor and debug the model in real-time to resolve these issues before deployment.

Which of the following actions should you take when configuring SageMaker Debugger for this task? (Select Two)

  1. A

    Set up built-in or custom debug rules to monitor training metrics and detect issues such as vanishing gradients or overfitting.

  2. B

    Enable SageMaker Model Monitor to visualize and track resource utilization, such as CPU and GPU, during training.

  3. C

    Configure a Debugger hook in the training script to capture and log tensors, which provide insights into the model's performance.

  4. D

    Use SageMaker Profiler to automatically correct hyperparameters such as learning rate and batch size during training.

Xem giải thích

Đáp án

A và C — đặt luật gỡ lỗi, và gắn hook trong mã huấn luyện để ghi lại tensor

Vì sao đúng

Hai phương án này là hai nửa của cách Debugger hoạt động:

  • C. Hook gắn vào mã huấn luyện để lấy mẫu tensor — trọng số, gradient, giá trị mất mát — ra nơi lưu. Không có hook thì chẳng có dữ liệu nào để phân tích.
  • A. Luật chạy trên dữ liệu đó và tự phát hiện vấn đề: có sẵn luật cho gradient biến mất và quá khớp, đúng hai thứ đề nêu, hoặc bạn tự viết luật riêng.

Vì sao các phương án khác sai

  • B. Model Monitor để xem mức dùng tài nguyên — Model Monitor theo dõi mô hình đã triển khai; phần tài nguyên lúc huấn luyện là của Profiler.
  • D. Profiler tự sửa siêu tham số — Profiler chỉ đo mức dùng tài nguyên; không có tính năng tự sửa tốc độ học nào.
Câu 226 Deployment and Orchestration of ML Workflows

A law firm wants to create an intelligent search system for internal legal documents. These documents are stored in different formats (PDFs, Word documents, and scanned images), and the firm needs to extract and index key information like case names, legal terms, and dates. The system should also allow the legal team to perform natural language searches to quickly retrieve relevant documents. Which combination of AWS services should the firm use to achieve this goal?

  1. A

    Amazon Rekognition to analyze the scanned images, Amazon Comprehend to detect entities and extract text from documents, and Amazon Kendra to provide natural language search capabilities.

  2. B

    Amazon Textract for extracting text from scanned images, Amazon Comprehend for entity recognition, and Amazon Kendra for building an intelligent search interface.

  3. C

    Amazon Kendra to store the documents, Amazon Rekognition for facial recognition in the scanned images, and Amazon Textract for analyzing legal terms.

  4. D

    Amazon Comprehend to categorize and index the documents, Amazon Lex for voice interaction with the system, and Amazon S3 for storing the document metadata.

Xem giải thích

Đáp án

B — Textract trích văn bản từ ảnh quét, Comprehend nhận thực thể

Vì sao đúng

Tài liệu ở nhiều định dạng, trong đó có ảnh quét, nên bước đầu bắt buộc là chuyển ảnh thành văn bản — đó là việc của Textract, và nó giữ được cả bảng lẫn cặp khoá–giá trị. Sau đó Comprehend rút ra thực thể để đánh chỉ mục. Muốn phần tìm kiếm mạnh hơn thì thêm Kendra ở trên cùng.

Vì sao các phương án khác sai

  • A. Rekognition để phân tích ảnh quét — Rekognition nhận diện vật thể và chữ ngắn trên ảnh, không dựng cho tài liệu nhiều trang.
  • C. Rekognition nhận diện khuôn mặt — hoàn toàn lệch nhu cầu tìm kiếm văn bản pháp lý.
  • D. Lex cho tương tác giọng nói — dựng chatbot, không liên quan tới việc trích và đánh chỉ mục tài liệu.
Câu 227 ML Model Development

A data scientist is using SageMaker Script Mode to train a machine learning model. They need to modularize their code by separating model definition, training logic, and inference logic into different Python files. They also need to include several external dependencies. How should the data scientist proceed to ensure SageMaker can correctly run their custom model?

  1. A

    Build a custom Docker container that includes all the Python files and dependencies, and then push it to Amazon ECR for SageMaker to use.

  2. B

    Use SageMaker’s built-in algorithms and ignore the need for custom modularized code.

  3. C

    Write all code into a single Python file and include dependencies directly within the script.

  4. D

    Split the code into multiple Python files, package them, and specify a single entry point script in the SageMaker estimator. Use a requirements.txt file to install dependencies.

Xem giải thích

Đáp án

D — Tách mã thành nhiều tệp Python, đóng gói lại, và khai một entry point

Vì sao đúng

Script Mode của SageMaker được dựng cho đúng cách làm này: bạn giữ mã theo cấu trúc mô-đun bình thường, khai source_dir trỏ tới cả thư mục và entry_point là tệp khởi động. SageMaker đóng gói cả thư mục, chép vào container và chạy tệp đó — nên tách được phần định nghĩa mô hình, phần huấn luyện và phần suy luận mà không phải dựng container riêng.

Vì sao các phương án khác sai

  • A. Tự dựng container Docker — làm được nhưng nặng hơn hẳn; Script Mode sinh ra để khỏi phải làm vậy.
  • B. Dùng thuật toán dựng sẵn và bỏ nhu cầu mô-đun hoá — bỏ luôn yêu cầu của đề.
  • C. Nhồi hết vào một tệp — ngược thẳng với mục tiêu tách mã.
Câu 228 Chọn nhiều đáp án Data Preparation for Machine Learning

An ML engineer is using Amazon SageMaker Data Wrangler to preprocess data for a machine learning project. They need to import customer transaction data from various sources, clean and transform the data, and engineer features for model training. Which TWO of the following steps should the engineer take to prepare the data using Data Wrangler? (Choose TWO.)

  1. A

    Export the transformed data directly to Amazon Athena for further analysis.

  2. B

    Use Data Wrangler to apply one-hot encoding on categorical variables.

  3. C

    Perform data transformations using SageMaker Studio's in-built Jupyter notebooks.

  4. D

    Manually write custom Python scripts to handle missing data and normalization tasks.

  5. E

    Import data from Amazon S3, Redshift, or SageMaker Feature Store into Data Wrangler.

Xem giải thích

Đáp án

B và E — mã hoá một-nóng cho biến phân loại, và nhập dữ liệu từ S3, Redshift hoặc Feature Store

Vì sao đúng

Hai phương án này đúng là hai việc Data Wrangler làm và đề đang cần:

  • E. Nhập dữ liệu — Data Wrangler nối trực tiếp tới S3, Athena, Redshift và Feature Store, nên gom được dữ liệu giao dịch từ nhiều nguồn mà không phải chép tay.
  • B. Mã hoá một-nóng — biến cột phân loại thành các cột số 0/1, bước bắt buộc trước khi đưa vào phần lớn thuật toán, và có sẵn trong danh sách phép biến đổi.

Vì sao các phương án khác sai

  • A. Xuất thẳng sang Athena — Athena là công cụ truy vấn, không phải đích xuất của Data Wrangler.
  • C. Dùng notebook trong Studio — làm được nhưng phải viết mã, trái mục đích của Data Wrangler.
  • D. Tự viết script Python — cũng vậy, đó là thứ công cụ này thay thế.
Câu 229 ML Solution Monitoring, Maintenance, and Security

A company deployed a binary classification model on a real-time Amazon SageMaker endpoint. To ensure long-term performance, they set up a schedule for SageMaker Model Monitor to evaluate model quality based on actual predictions and outcomes over time.
Which of the following steps is required to effectively monitor model quality using SageMaker Model Monitor?

  1. A

    Enable scheduled retraining on SageMaker to automatically improve model accuracy without monitoring metrics.

  2. B

    Set up custom metrics for feature attribution and monitor the bias drift for predictions from the real-time endpoint.

  3. C

    Enable automatic hyperparameter tuning for the endpoint and log model predictions to Amazon CloudWatch.

  4. D

    Create baseline statistics from the model’s training data and enable data capture to record requests and predictions.

Xem giải thích

Đáp án

D — Tạo thống kê nền từ dữ liệu huấn luyện và bật ghi lại dữ liệu ở endpoint

Vì sao đúng

Model Monitor làm việc bằng cách so sánh với một mốc. Hai điều kiện bắt buộc: có baseline tính từ dữ liệu huấn luyện (khoảng giá trị, phân bố, tỷ lệ thiếu), và bật data capture để endpoint ghi lại dữ liệu vào cùng dự đoán ra xuống S3. Thiếu một trong hai thì lịch giám sát chạy mà không có gì để so.

Vì sao các phương án khác sai

  • A. Bật huấn luyện lại theo lịch — huấn luyện lại vô điều kiện, không phải giám sát; và không biết vì sao mà huấn luyện lại.
  • B. Chỉ số riêng cho đóng góp đặc trưng — đó là một loại giám sát khác (trôi đóng góp), không phải điều kiện nền để bắt đầu.
  • C. Tinh chỉnh siêu tham số cho endpoint — không có khái niệm này; tinh chỉnh diễn ra lúc huấn luyện.
Câu 230 Deployment and Orchestration of ML Workflows

A company wants to integrate generative AI capabilities into their applications using foundation models from multiple AI providers. They need a solution that offers a serverless environment, flexibility in model customization, and ensures that data remains within a specific AWS region for compliance reasons.
Which AWS service would best meet these requirements?

  1. A

    Amazon Bedrock

  2. B

    Amazon Lex

  3. C

    Amazon SageMaker

  4. D

    Amazon Rekognition

Xem giải thích

Đáp án

A — Amazon Bedrock

Vì sao đúng

Đề nêu đúng ba đặc điểm của Bedrock: nhiều nhà cung cấp mô hình nền truy cập qua một API thống nhất, không máy chủ nên không phải cấp phát hay quản lý hạ tầng, và tích hợp sẵn với phần còn lại của AWS. Bedrock cũng lo phần tuỳ biến bằng fine-tune và truy xuất tăng cường mà không phải tự dựng gì.

Vì sao các phương án khác sai

  • B. Amazon Lex — dựng chatbot theo ý định và khe thông tin, không phải nền tảng mô hình nền.
  • C. Amazon SageMaker — làm được nhưng bạn phải chọn máy, triển khai endpoint và tự vận hành; trái yêu cầu "không máy chủ".
  • D. Amazon Rekognition — phân tích ảnh và video, không phải AI sinh nội dung.