Ngân hàng đề — AWS Certified Machine Learning Engineer Associate

Tìm thấy 635 câu.

Câu 391 Deployment and Orchestration of ML Workflows

You are working with SageMaker built-in algorithms and need to perform image classification. Which of the following statements correctly describes the process of using a SageMaker built-in algorithm for training and deployment?

  1. A

    SageMaker built-in algorithms require the use of custom containers where you manually install the necessary libraries.

  2. B

    You need to bring your dataset and specify hyperparameters, but the algorithm is pre-packaged and managed by AWS, including the necessary infrastructure setup.

  3. C

    The built-in algorithms come fully trained and do not require further training on your specific dataset, ensuring immediate deployment.

  4. D

    SageMaker built-in algorithms cannot be used for image classification tasks, and custom algorithms must be used for these purposes.

Xem giải thích

Đáp án

B — Bạn đưa dữ liệu và khai siêu tham số, phần còn lại thuật toán lo

Vì sao đúng

Đây là toàn bộ ý nghĩa của thuật toán dựng sẵn: mã mô hình, container và phần tối ưu đã có sẵn. Việc của bạn gọn lại thành ba thứ — dữ liệu đã gán nhãn ở định dạng thuật toán yêu cầu, bộ siêu tham số, và loại máy để chạy. Với phân loại ảnh, thuật toán Image Classification còn hỗ trợ học chuyển giao từ mô hình đã huấn luyện trên tập lớn.

Vì sao các phương án khác sai

  • A. Phải dùng container tự dựng — sai; container đã có sẵn, đó chính là điểm mạnh.
  • C. Đã huấn luyện sẵn, không cần huấn luyện thêm — sai; chúng là thuật toán, không phải mô hình đã huấn luyện. (Mô hình đã huấn luyện sẵn nằm ở JumpStart.)
  • D. Không dùng được cho phân loại ảnh — sai; có thuật toán riêng cho việc này.
Câu 392 ML Model Development

A data scientist at a financial institution wants to train a machine learning model using Amazon SageMaker's built-in algorithms. They need to optimize for ease of use and infrastructure management, while still having the flexibility to define their hyperparameters and perform automatic hyperparameter tuning. The data scientist is considering the built-in "XGBoost" algorithm for a classification task. Which of the following steps are necessary to properly set up and train the model using the built-in algorithm?

  1. A

    Customize the XGBoost algorithm container and manage infrastructure manually before training.

  2. B

    Use the pre-packaged XGBoost algorithm container, specify the data location, define the hyperparameters, and select an appropriate compute instance.

  3. C

    Develop a custom XGBoost container to allow for proprietary data processing pipelines.

  4. D

    Install the XGBoost framework locally, develop the model, and upload the trained model to SageMaker for inference.

Xem giải thích

Đáp án

B — Dùng container XGBoost đóng gói sẵn, khai vị trí dữ liệu và siêu tham số

Vì sao đúng

Đề nêu tiêu chí dễ dùng, và đây là con đường ngắn nhất: trỏ tới ảnh container XGBoost do AWS cung cấp, khai đường dẫn dữ liệu trên S3 cùng siêu tham số, rồi chạy job. Không viết mã mô hình, không dựng ảnh Docker, và container đã được tối ưu sẵn cho hiệu năng.

Vì sao các phương án khác sai

  • A. Tuỳ biến container và tự quản hạ tầng — trái thẳng tiêu chí dễ dùng.
  • C. Tự dựng container XGBoost — chỉ cần khi bạn có bước xử lý riêng không nhét được vào container chuẩn; ở đây không có nhu cầu đó.
  • D. Cài XGBoost trên máy cá nhân rồi tải mô hình lên — mất hết lợi thế huấn luyện trên đám mây, và không mở rộng được khi dữ liệu lớn.
Câu 393 Chọn nhiều đáp án Deployment and Orchestration of ML Workflows

When using Amazon SageMaker's automatic hyperparameter tuning, which of the following must be specified to start the tuning job? (Select TWO)

  1. A

    Performance metric

  2. B

    Training algorithm

  3. C

    Batch size for model training

  4. D

    Number of EC2 spot instances

  5. E

    A fixed set of hyperparameter values

Xem giải thích

Đáp án

A và B — chỉ số đánh giá, và thuật toán huấn luyện

Vì sao đúng

Job tinh chỉnh cần biết huấn luyện bằng gì và thế nào là tốt hơn. Thuật toán (hoặc ảnh container) cho biết cái thứ nhất; chỉ số mục tiêu cùng hướng tối ưu cho biết cái thứ hai — thiếu nó thì dịch vụ không so được hai lần chạy và không biết chọn cấu hình tiếp theo ra sao.

Vì sao các phương án khác sai

  • E. Tập giá trị siêu tham số cố định — sai bản chất: bạn khai khoảng để dịch vụ dò trong đó; cố định thì chẳng còn gì để tinh chỉnh.
  • C. Kích thước lô — có thể là một trong những siêu tham số cần dò, nhưng không bắt buộc.
  • D. Số máy spot — tuỳ chọn tiết kiệm chi phí, không phải điều kiện để job chạy.

Ghi chú

Câu này trùng nội dung với một câu khác trong cùng ngân hàng đề, chỉ khác thứ tự phương án. Bộ đề nguồn có một số câu lặp như vậy.

Câu 394 ML Model Development

Which of the following scenarios would be most appropriate for using Amazon SageMaker Script Mode over creating a custom Docker container?

  1. A

    The machine learning team needs to use a custom version of TensorFlow that has specialized libraries not supported by SageMaker’s built-in containers.

  2. B

    A data scientist wants to focus on model development using a pre-configured PyTorch environment while avoiding the complexity of managing container infrastructure.

  3. C

    A developer requires full control over the operating system, libraries, and dependencies for an advanced C++ model.

  4. D

    The model needs to be trained and deployed in a completely different programming language, such as Rust, which SageMaker does not support natively.

Xem giải thích

Đáp án

B — Khi muốn tập trung phát triển mô hình trên container PyTorch đã cấu hình sẵn

Vì sao đúng

Script Mode đúng chỗ khi khung học máy bạn dùng đã có container dựng sẵn và bạn chỉ cần chạy mã của mình bên trong đó. Viết script, khai entry point, thêm thư viện qua requirements.txt — không phải dựng và bảo trì ảnh Docker nào. Đó là con đường ít việc nhất mà vẫn dùng được mã riêng.

Vì sao các phương án khác sai

  • A. Cần bản TensorFlow tuỳ chỉnh — khi phiên bản hoặc bản vá không có trong container chuẩn thì phải dựng container riêng.
  • C. Cần toàn quyền với hệ điều hành và thư viện hệ thống — vượt quá những gì Script Mode cho phép.
  • D. Cần ngôn ngữ lập trình hoàn toàn khác — container dựng sẵn gắn với hệ sinh thái Python.
Câu 395 ML Model Development

A machine learning team is using Amazon SageMaker Experiments to run multiple training jobs with different hyperparameter configurations and feature sets. They want to organize, track, and compare the results to identify the best-performing model. Which of the following components in SageMaker Experiments allows the team to track and compare each individual training run within the experiment?

  1. A

    SageMaker Pipelines

  2. B

    Trials

  3. C

    SageMaker Autopilot

  4. D

    Trial Components

Xem giải thích

Đáp án

B — Trials

Vì sao đúng

Trial là đơn vị tổ chức đúng tầng mà đề cần: mỗi trial ứng với một lần chạy có một cấu hình cụ thể — một bộ siêu tham số và một tập đặc trưng. Nhiều trial nằm chung dưới một experiment, nên xếp cạnh nhau và so được ngay trong Studio.

Vì sao các phương án khác sai

  • D. Trial Components — là các bước bên trong một trial (tiền xử lý, huấn luyện, đánh giá), tầng thấp hơn thứ đề hỏi.
  • A. Pipelines — điều phối quy trình, không phải cấu trúc tổ chức thí nghiệm.
  • C. Autopilot — tự dựng mô hình thay bạn; ở đây nhóm đang chủ động thiết kế các cấu hình.
Câu 396 Data Preparation for Machine Learning

A startup working on autonomous driving technology needs to label a large dataset of road images for training their object detection model. The team wants to minimize the cost of labeling while maintaining high-quality results. They aim to use a combination of human labelers and automation to label straightforward objects while having human reviewers handle more complex cases.

Which of the following strategies BEST fits their needs?

  1. A

    Use SageMaker Ground Truth with active learning to automate labeling of simple cases while sending complex cases to human labelers for review.

  2. B

    Use SageMaker Feature Store to preprocess and label the images before passing them to the machine learning model for training.

  3. C

    Use a fully automated labeling approach without human intervention to reduce labeling costs.

  4. D

    Use SageMaker Experiments to run multiple trials of different labeling approaches to determine which strategy reduces costs the most.

Xem giải thích

Đáp án

A — Ground Truth kèm học chủ động để tự động gán nhãn những ca đơn giản

Vì sao đúng

Với hàng loạt ảnh đường phố, phần lớn khung hình có vật thể rõ ràng và dễ nhận. Học chủ động để mô hình gán nhãn những ca đó, và chỉ đẩy sang người những ảnh mô hình không chắc — trời tối, vật thể bị che, tình huống hiếm. Đó đúng là những ca quyết định chất lượng của mô hình phát hiện vật thể, nên tiền được tiêu vào đúng chỗ.

Vì sao các phương án khác sai

  • B. Feature Store để gán nhãn — Feature Store lưu đặc trưng, không gán nhãn.
  • C. Tự động hoàn toàn, không có người — chất lượng tụt ngay ở những ca khó, mà với xe tự lái thì đó là những ca nguy hiểm nhất.
  • D. Experiments để thử các cách gán nhãn — Experiments theo dõi các lần huấn luyện, không phải công cụ gán nhãn.
Câu 397 Deployment and Orchestration of ML Workflows

During a model training session, you notice that your training job is taking significantly longer than expected. To identify the root cause, you decide to use the SageMaker Profiler to analyze resource bottlenecks, such as CPU and GPU underutilization. Which of the following actions should you take to correctly set up SageMaker Profiler for this task?

  1. A

    Add SageMaker Profiler start and stop commands in the training script to collect system resource metrics like CPU and GPU utilization.

  2. B

    Define custom debug rules to monitor vanishing gradients and early stopping criteria during training.

  3. C

    Use SageMaker Model Monitor to track and visualize the training job's input data distribution changes.

  4. D

    Automatically correct detected bottlenecks by reducing the training batch size using SageMaker Debugger’s real-time insights.

Xem giải thích

Đáp án

A — Thêm lệnh bắt đầu và kết thúc của Profiler vào mã huấn luyện để thu dữ liệu

Vì sao đúng

Huấn luyện lâu hơn dự kiến thường là nút thắt tài nguyên chứ không phải lỗi mô hình. Profiler thu số liệu về mức dùng GPU, CPU, bộ nhớ và I/O; đánh dấu đoạn cần đo bằng lệnh bắt đầu và kết thúc thì dữ liệu gắn được với đúng giai đoạn, nên nhìn ra ngay nút thắt nằm ở khâu nạp dữ liệu hay ở phần tính toán.

Vì sao các phương án khác sai

  • B. Đặt luật gỡ lỗi cho gradient biến mất — đó là Debugger và nhắm vào chất lượng mô hình, không phải tốc độ.
  • C. Dùng Model Monitor cho job huấn luyện — Model Monitor chỉ hoạt động với mô hình đã triển khai.
  • D. Tự sửa nút thắt bằng cách giảm kích thước lô — Profiler chỉ đo và báo; nó không tự thay đổi cấu hình, và giảm lô có khi còn làm GPU nhàn rỗi thêm.
Câu 398 ML Solution Monitoring, Maintenance, and Security

A machine learning team has deployed a model to production using Amazon SageMaker. They set up SageMaker Model Monitor to detect any data drift. After several weeks, they notice that the predictions of the model are becoming less accurate compared to the actual outcomes, and SageMaker Model Monitor has flagged significant differences in the input data distribution over time.

What is the MOST likely cause of the declining model performance, and what is the recommended action?

  1. A

    Data drift has occurred, causing the model to receive input data that is different from the training data. The team should update the model's hyperparameters to resolve the issue.

  2. B

    Data drift has occurred, causing the input data distribution to deviate from the training data. The team should retrain the model with updated training data to resolve the issue.

  3. C

    Model quality drift has occurred because the model’s predictions are becoming less accurate. The team should switch to a different algorithm to improve performance.

  4. D

    Bias drift has occurred, leading the model to favor certain outcomes. The team should retrain the model with a new algorithm and set stricter constraints in SageMaker Model Monitor.

Xem giải thích

Đáp án

B — Trôi dữ liệu: phân bố dữ liệu vào đã lệch khỏi dữ liệu lúc huấn luyện

Vì sao đúng

Model Monitor phát hiện trôi dữ liệu bằng cách so thống kê của dữ liệu đi vào endpoint với mốc lấy từ tập huấn luyện — khoảng giá trị, trung bình, tỷ lệ thiếu, tần suất các hạng mục. Khi báo động nổi lên sau vài tuần, đó là dấu hiệu thế giới thực đã đổi chứ không phải mô hình hỏng.

Ưu điểm của loại giám sát này: nó chỉ cần dữ liệu vào, phát hiện được ngay, trong khi đo độ chính xác phải chờ có nhãn thật.

Vì sao các phương án khác sai

  • A — mô tả gần đúng nhưng thiếu điểm cốt lõi là so với phân bố lúc huấn luyện.
  • C. Trôi chất lượng mô hình — cần biết kết quả thực tế để đo; đó là loại giám sát khác.
  • D. Trôi thiên lệch — nói về việc ưu ái nhóm nào đó, không phải chuyện phân bố dữ liệu.
Câu 399 Chọn nhiều đáp án ML Solution Monitoring, Maintenance, and Security

A data science team deployed a model using Amazon SageMaker real-time endpoints. They are concerned that the distribution of data in production may change over time, potentially leading to data drift. To mitigate this risk, they want to set up continuous monitoring of the model in production. Which of the following does Amazon SageMaker Model Monitor allow the team to track to address their concerns? (Select TWO)

  1. A

    Model Latency Drift

  2. B

    Data Drift

  3. C

    Feature Attribution Drift

  4. D

    AWS CloudFormation Drift

Xem giải thích

Đáp án

B và C — trôi dữ liệu, và trôi đóng góp đặc trưng

Vì sao đúng

Đề lo rằng phân bố dữ liệu sản xuất sẽ đổi theo thời gian, và Model Monitor có đúng hai loại giám sát cho mối lo đó:

  • B. Trôi dữ liệu — so phân bố dữ liệu vào với mốc từ tập huấn luyện.
  • C. Trôi đóng góp đặc trưng — theo dõi mô hình đang dựa vào đặc trưng nào để quyết định. Đôi khi phân bố từng cột vẫn y nguyên nhưng quan hệ giữa chúng đã đổi, và chỉ loại này bắt được.

Vì sao các phương án khác sai

  • A. Model Latency Drift — không phải một loại giám sát của Model Monitor; độ trễ là chỉ số vận hành, theo dõi bằng CloudWatch.
  • D. CloudFormation Drift — là khái niệm của CloudFormation (tài nguyên thật lệch khỏi mẫu), hoàn toàn không liên quan tới học máy.
Câu 400 ML Solution Monitoring, Maintenance, and Security

A machine learning engineer is tasked with automating the registration, versioning, and deployment of models as part of a workflow using SageMaker Pipelines. The goal is to track model development, including data preprocessing, training runs, and version history. The engineer needs to ensure that the pipeline can log the necessary details for auditing and compliance.
Which SageMaker feature should the engineer use to track the history of the model’s development?

  1. A

    Model Group

  2. B

    Approval Status

  3. C

    Model Lineage

  4. D

    Model Metrics

Xem giải thích

Đáp án

C — Model Lineage (nguồn gốc mô hình)

Vì sao đúng

Lineage ghi lại chuỗi quan hệ tạo ra một mô hình: dữ liệu nào, job xử lý nào, job huấn luyện nào, siêu tham số nào, rồi ra model package nào. Khi tự động hoá bằng Pipelines, SageMaker dựng đồ thị này giúp bạn — nhờ vậy trả lời được câu "bản đang chạy sản xuất sinh ra từ đâu", thứ bắt buộc khi bị kiểm toán hoặc khi cần dựng lại.

Vì sao các phương án khác sai

  • A. Model Group — chỗ gom các phiên bản của cùng một mô hình; là cách sắp xếp, không phải dấu vết nguồn gốc.
  • B. Approval Status — trạng thái phê duyệt, quyết định bản nào được ra sản xuất.
  • D. Model Metrics — chỉ số đánh giá của một bản, chỉ là một mảnh trong bức tranh lineage.