Ngân hàng đề — AWS Certified Machine Learning Engineer Associate

Tìm thấy 635 câu.

Câu 241 Chọn nhiều đáp án Deployment and Orchestration of ML Workflows

You are training a machine learning model using PyTorch on Amazon SageMaker and want to ensure that the system resources, such as CPU and GPU, are being effectively utilized. You plan to use SageMaker Profiler to track these metrics and improve the efficiency of your training job. Which of the following steps should you follow to set up and use SageMaker Profiler correctly? (Select Two)

  1. A

    Manually enable automatic SageMaker Profiler optimization in the SageMaker console to reduce memory overhead during training.

  2. B

    Configure the SageMaker Profiler to automatically tune your hyperparameters, such as learning rate, based on the resource utilization data.

  3. C

    Define the profiler configuration in your SageMaker Estimator, specifying the duration of the profiling in seconds for the CPU and GPU.

  4. D

    Use SageMaker Neo to compile your model for optimized CPU and GPU performance after training is completed.

  5. E

    Import the necessary SageMaker Profiler modules and add start and stop profiling commands in your PyTorch training script.

Xem giải thích

Đáp án

C và E — khai cấu hình profiler trong Estimator, và thêm lệnh bắt đầu/kết thúc trong mã

Vì sao đúng

SageMaker Profiler cần cấu hình ở cả hai phía:

  • C. Phía Estimator — khai profiler_config để nói cho SageMaker biết lấy mẫu trong bao lâu và với tần suất nào, và bật thu thập chỉ số hệ thống.
  • E. Phía mã huấn luyện — nhập thư viện profiler và đánh dấu đoạn cần đo bằng lệnh bắt đầu và kết thúc, nhờ vậy dữ liệu gắn được với đúng giai đoạn của quá trình huấn luyện.

Vì sao các phương án khác sai

  • A. Bật tối ưu tự động trên console — không có tính năng như vậy; Profiler chỉ đo.
  • B. Profiler tự tinh chỉnh siêu tham số — không; đó là việc của Automatic Model Tuning.
  • D. Dùng SageMaker Neo — biên dịch tối ưu mô hình sau khi huấn luyện xong để triển khai.
Câu 242 Data Preparation for Machine Learning

A financial services company is building a fraud detection system to monitor transactions in real-time. They plan to use Amazon Kinesis Data Streams to ingest transaction data and want the system to scale automatically based on demand. They require low-latency for fast fraud detection responses.

Which feature of Amazon Kinesis Data Streams should the company implement to meet these requirements?

  1. A

    Manually scale shards based on throughput requirements.

  2. B

    Use Kinesis Firehose to automatically scale the stream based on traffic.

  3. C

    Enable provisioned mode and dynamically add shards.

  4. D

    Use enhanced fan-out to reduce consumer latency and ensure dedicated throughput per consumer.

Xem giải thích

Đáp án

D — Dùng enhanced fan-out để giảm độ trễ và cấp thông lượng riêng cho mỗi consumer

Vì sao đúng

Với phát hiện gian lận thời gian thực, độ trễ là tất cả. Không bật enhanced fan-out thì mọi consumer chia nhau 2 MB/giây mỗi shard và phải hỏi liên tục, nên độ trễ tăng theo số ứng dụng đọc. Enhanced fan-out cấp đường riêng cho từng consumer và đẩy dữ liệu sang, đưa độ trễ về khoảng 70 mili giây.

Vì sao các phương án khác sai

  • A. Tự thêm shard bằng tay — vẫn là chia sẻ băng thông giữa các consumer, cộng thêm việc vận hành.
  • B. Firehose tự co giãn — Firehose gom lô, không đạt yêu cầu thời gian thực.
  • C. Provisioned mode rồi thêm shard động — tăng tổng thông lượng nhưng không giải quyết chuyện nhiều consumer giành nhau phần của từng shard.
Câu 243 ML Model Development

A data scientist is setting up an environment to develop, train, and deploy machine learning models using AWS SageMaker. The scientist wants a fully managed Jupyter notebook environment that simplifies infrastructure setup and provides pre-installed libraries such as TensorFlow and PyTorch. Additionally, they want to ensure that their work is automatically saved and can be accessed later.
Which AWS SageMaker feature would BEST meet these requirements?

  1. A

    SageMaker Studio with Jupyter Lab spaces

  2. B

    Amazon EC2 with Jupyter notebook manually installed

  3. C

    AWS Glue with integrated Jupyter environment

  4. D

    SageMaker Notebook Instances

Xem giải thích

Đáp án

D — SageMaker Notebook Instances

Vì sao đúng

Đề mô tả nhu cầu của một người: một môi trường Jupyter được quản lý để phát triển, huấn luyện và triển khai. Notebook instance là đúng thứ đó — một máy EC2 do AWS quản lý, cài sẵn Jupyter và các thư viện ML, bật lên là dùng được ngay.

Vì sao các phương án khác sai

  • A. Studio với Jupyter Lab spaces — mạnh hơn và hợp khi làm theo nhóm, nhưng đề ở đây chỉ nói tới một nhà khoa học dữ liệu và một môi trường notebook được quản lý.
  • B. EC2 tự cài Jupyter — phải tự cài, tự vá, tự bảo mật; trái yêu cầu "được quản lý".
  • C. AWS Glue kèm Jupyter — Glue có notebook nhưng để phát triển job ETL, không phải môi trường phát triển ML.

Ghi chú

So với câu về làm việc nhóm và chia sẻ notebook thì đáp án lại là Studio. Hai câu rất giống nhau, khác nhau đúng ở chỗ có yêu cầu nhiều người dùng hay không.

Câu 244 Chọn nhiều đáp án Data Preparation for Machine Learning

A machine learning engineer is using Amazon SageMaker Data Wrangler to prepare data for training a model. The dataset has missing values, inconsistent scales across features, and categorical data. Which of the following actions can the engineer perform using Data Wrangler to resolve these issues? (Choose THREE)

  1. A

    Manually code all the transformations using Python and upload the final dataset to SageMaker.

  2. B

    Use Amazon EMR to pre-process the data for missing values and scaling.

  3. C

    Perform imputation to fill in missing values using strategies like mean or median.

  4. D

    One-hot encode categorical data to convert it into a numerical format.

  5. E

    Normalize feature values to ensure all features are on a similar scale.

Xem giải thích

Đáp án

C, D và E — điền giá trị thiếu, mã hoá một-nóng, và chuẩn hoá thang đo

Vì sao đúng

Ba vấn đề đề nêu ứng với ba phép biến đổi có sẵn trong Data Wrangler:

  • C. Điền giá trị thiếu bằng trung bình hoặc trung vị, thay vì bỏ cả dòng và mất dữ liệu.
  • D. Mã hoá một-nóng biến cột phân loại thành các cột số, bước bắt buộc cho phần lớn thuật toán.
  • E. Chuẩn hoá đưa các đặc trưng về cùng thang; không làm thì đặc trưng có đơn vị lớn sẽ áp đảo hàm mất mát ở những thuật toán nhạy với khoảng cách.

Vì sao các phương án khác sai

  • A. Tự viết toàn bộ bằng Python — chính là thứ Data Wrangler sinh ra để thay thế.
  • B. Dùng EMR để tiền xử lý — dựng cả một cụm cho việc đã có công cụ trực quan làm sẵn.
Câu 245 ML Solution Monitoring, Maintenance, and Security

A machine learning model deployed into production has started underperforming due to the introduction of new data, resulting in the model making less accurate predictions. The team is concerned about both data drift (changes in input data distribution) and model drift (declining model accuracy).
Which AWS service combination is most suitable for addressing these issues?

  1. A

    Amazon S3 to store new data and SageMaker Model Monitor to manage model drift.

  2. B

    SageMaker Feature Store to track data distribution and Amazon Kinesis to monitor model predictions.

  3. C

    AWS CloudTrail to monitor model drift and AWS Lambda to retrain the model.

  4. D

    SageMaker Model Monitor to track data drift and SageMaker Pipelines to automate model retraining.

Xem giải thích

Đáp án

D — Model Monitor theo dõi trôi dữ liệu, Pipelines tự động hoá việc huấn luyện lại

Vì sao đúng

Đề mô tả đúng hiện tượng trôi: dữ liệu mới khác dữ liệu lúc huấn luyện nên mô hình kém dần. Model Monitor so dữ liệu thật với thống kê nền và cảnh báo khi lệch. Pipelines đóng gói chuỗi tiền xử lý – huấn luyện – đánh giá – đăng ký thành quy trình chạy lại được, nên phản ứng với cảnh báo là chạy một pipeline chứ không phải làm tay từng bước.

Vì sao các phương án khác sai

  • A. S3 lưu dữ liệu mới, Model Monitor quản lý trôi — S3 chỉ là nơi lưu, thiếu hẳn phần tự động huấn luyện lại.
  • B. Feature Store theo dõi phân bố — kho đặc trưng, không có phần phát hiện trôi.
  • C. CloudTrail theo dõi trôi mô hình — CloudTrail ghi lời gọi API, hoàn toàn không biết gì về chất lượng mô hình.
Câu 246 ML Model Development

A company is training a large neural network model with billions of parameters on Amazon SageMaker. The model is too large to fit into the memory of a single GPU. What technique should the company use to distribute the training across multiple instances in this scenario?

  1. A

    Gradient Parallelism

  2. B

    Hybrid Model and Data Parallelism

  3. C

    Data Parallelism

  4. D

    Model Parallelism

Xem giải thích

Đáp án

D — Song song hoá theo mô hình (Model Parallelism)

Vì sao đúng

Điều kiện quyết định trong đề là mô hình không vừa bộ nhớ một GPU. Song song theo mô hình chia chính mô hình — theo lớp hoặc theo tensor — ra nhiều GPU, mỗi GPU giữ một phần tham số. Đó là cách duy nhất trong danh sách giải được vấn đề bộ nhớ.

Vì sao các phương án khác sai

  • C. Song song theo dữ liệu — giữ bản sao đầy đủ của mô hình trên mỗi GPU, nên nếu một GPU không chứa nổi mô hình thì cách này không chạy được.
  • B. Kết hợp cả hai — dùng thật trong các mô hình rất lớn, nhưng đề chỉ hỏi kỹ thuật giải đúng vấn đề bộ nhớ, và câu trả lời cốt lõi là song song theo mô hình.
  • A. Song song theo gradient — không phải tên gọi của kỹ thuật nào trong SageMaker.
Câu 247 Data Preparation for Machine Learning

A retail company uses Amazon SageMaker to train machine learning models for product recommendation. The team stores raw customer behavior data, such as clicks and purchases, in an Amazon S3 bucket. They want to preprocess this data to remove duplicates, normalize numerical features, and encode categorical variables before training their model. Additionally, the preprocessing needs to scale to handle increasing data volumes.

Which solution will most effectively meet their requirements?

  1. A

    Leverage AWS Glue for ETL to preprocess the data, and then load the transformed data into Amazon SageMaker.

  2. B

    Use Amazon Athena to query the data and perform all preprocessing using SQL, then export the processed data back to Amazon S3.

  3. C

    Write a preprocessing script in Python and run it on Amazon EC2 instances with auto-scaling enabled.

  4. D

    Use Amazon SageMaker Processing Jobs with a pre-built container for data preprocessing and store the processed data back in S3.

Xem giải thích

Đáp án

D — Dùng SageMaker Processing Jobs với container dựng sẵn

Vì sao đúng

Processing Jobs sinh ra cho đúng khâu này: chạy tiền xử lý trên máy do SageMaker cấp phát, xong việc tự trả máy về nên chỉ trả tiền cho thời gian chạy thật. Có container dựng sẵn cho scikit-learn và Spark nên không phải tự đóng gói. Kết quả ghi thẳng ra S3, nối liền mạch sang bước huấn luyện trong cùng một pipeline.

Vì sao các phương án khác sai

  • A. Glue ETL — làm được, nhưng nằm ngoài hệ SageMaker nên phải nối thêm và quản lý riêng; Processing Jobs gần với quy trình ML hơn.
  • B. Làm hết bằng SQL trong Athena — nhiều phép tiền xử lý cho ML (chuẩn hoá, mã hoá, xử lý ảnh) rất khó hoặc không diễn đạt được bằng SQL.
  • C. Tự chạy script trên EC2 với auto scaling — phải tự dựng và tự vận hành, đúng thứ dịch vụ được quản lý giúp bạn tránh.
Câu 248 Data Preparation for Machine Learning

An e-commerce company needs to build a recommendation engine using a machine learning model trained on customer activity data, including purchase history, browsing patterns, and demographics. The team wants to streamline their model development process by creating a centralized repository of preprocessed features that can be reused across different projects. The system needs to support real-time recommendations and batch training.

What is the MOST appropriate setup for the team to ensure consistent feature reuse, while also optimizing for real-time and batch processing?

  1. A

    Store all features in Amazon S3 and manually preprocess them for each model training run.

  2. B

    Use the SageMaker Feature Store to create an offline store for batch training and an online store for real-time inference.

  3. C

    Use SageMaker Ground Truth to label the data and store the features in an offline store for training only.

  4. D

    Use the SageMaker Feature Store to create an online store for batch processing and an offline store for real-time recommendations.

Xem giải thích

Đáp án

B — Dùng Feature Store: offline store cho huấn luyện theo lô, online store cho gợi ý thời gian thực

Vì sao đúng

Hệ gợi ý cần đặc trưng ở hai nhịp khác nhau. Offline store giữ toàn bộ lịch sử trên S3 kèm dấu thời gian, dùng để huấn luyện. Online store giữ giá trị mới nhất của từng khách và trả về trong vài mili giây khi trang web đang chờ. Điều quan trọng: cả hai lấy từ cùng một định nghĩa đặc trưng, nên tránh được lỗi kinh điển là đặc trưng lúc huấn luyện khác lúc chạy thật.

Vì sao các phương án khác sai

  • A. Để hết trên S3 rồi tiền xử lý tay mỗi lần — lặp công, và mỗi lần làm lại là một cơ hội để hai bên lệch nhau.
  • C. Ground Truth để gán nhãn — dịch vụ gán nhãn dữ liệu, không phải kho đặc trưng.
  • D. Dùng online store cho xử lý theo lô — đảo vai: online tối ưu cho đọc từng khoá độ trễ thấp, quét cả tập để huấn luyện thì phải dùng offline.
Câu 249 Chọn nhiều đáp án ML Model Development

A machine learning engineer is using Amazon SageMaker Studio to perform exploratory data analysis on a large dataset stored in Amazon S3. The analyst needs to create visualizations, run data preprocessing workflows, and share the results with team members. The analyst also wants to avoid setting up infrastructure manually. Considering these requirements, which SageMaker features should the analyst leverage to achieve these tasks efficiently? (Select THREE)

  1. A

    Take advantage of the built-in integrations with popular data visualization libraries such as Matplotlib and Seaborn to create interactive data visualizations.

  2. B

    Store the dataset in the SageMaker notebook instance's local storage to avoid the complexity of using external services like Amazon S3.

  3. C

    Use SageMaker Automatic Model Tuning to automatically handle hyperparameter optimization during the preprocessing and analysis steps.

  4. D

    Share SageMaker Studio notebooks directly with team members for collaboration, ensuring that they can access the same environment and results.

  5. E

    Use Amazon SageMaker's Real-Time Inference Endpoints to deploy the model and generate immediate predictions on the dataset.

  6. F

    Utilize SageMaker Studio's JupyterLab environment to interactively run code cells, visualize data, and tweak preprocessing workflows without worrying about underlying infrastructure.

Xem giải thích

Đáp án

A, D và F

Vì sao đúng

Ba phương án này đúng là những gì Studio cung cấp cho khâu khám phá dữ liệu:

  • F. Môi trường JupyterLab — chạy từng ô mã tương tác, xem kết quả ngay, đúng nhịp làm việc khi đang dò dữ liệu.
  • A. Tích hợp sẵn thư viện vẽ biểu đồ — matplotlib, seaborn và các thư viện quen thuộc đã có trong ảnh dựng sẵn, không phải cài.
  • D. Chia sẻ notebook cho đồng đội — Studio chia sẻ được kèm cả ngữ cảnh chạy, nên người khác mở ra là thấy đúng thứ bạn thấy.

Vì sao các phương án khác sai

  • B. Chép dữ liệu vào ổ cục bộ của notebook — tập dữ liệu lớn thì không vừa, và mất luôn lợi thế đọc thẳng từ S3.
  • C. Automatic Model Tuning — tối ưu siêu tham số, thuộc giai đoạn huấn luyện.
  • E. Endpoint suy luận thời gian thực — để phục vụ dự đoán, không phải để khám phá dữ liệu.
Câu 250 ML Model Development

A robotics company is using SageMaker to train a reinforcement learning agent to navigate through a dynamic warehouse environment. During training, the agent sometimes moves farther away from the goal instead of closer. What is the most likely cause of this behavior?

  1. A

    Negative rewards or penalties

  2. B

    Adjustments to the policy function

  3. C

    Adjustments to the action space

  4. D

    Positive rewards

Xem giải thích

Đáp án

D — Phần thưởng dương

Vì sao đúng

Trong học tăng cường, tác nhân học bằng cách tối đa hoá tổng phần thưởng. Muốn nó lặp lại một hành vi thì thưởng dương cho hành vi đó — ở đây là những bước đi đưa robot lại gần đích hơn. Đây là phần "củng cố" trong chính tên gọi của phương pháp: hành vi được thưởng thì xác suất được chọn lại tăng lên.

Vì sao các phương án khác sai

  • A. Phần thưởng âm hoặc hình phạt — dùng để ngăn hành vi xấu, không phải để khuyến khích hành vi tốt. Thực tế người ta dùng cả hai, nhưng câu hỏi hỏi cái nào khuyến khích.
  • B. Chỉnh hàm chính sách — chính sách là thứ học ra được từ phần thưởng, không phải nút bạn vặn tay để dạy hành vi.
  • C. Chỉnh không gian hành động — thay đổi tập lựa chọn robot có, không nói gì về việc lựa chọn nào là tốt.