Ngân hàng đề — AWS Certified Machine Learning Engineer Associate

Tìm thấy 635 câu.

Câu 291 Deployment and Orchestration of ML Workflows

A machine learning team is using SageMaker Model Registry to manage the lifecycle of their models. After training a new version of a model, they want to automate the registration of the model in the registry and deploy it to a real-time endpoint. What features of SageMaker Model Registry and Pipelines will help achieve this?

  1. A

    Model versions in the registry cannot be automatically deployed; manual steps are always required.

  2. B

    SageMaker Pipelines can automate the registration of the model in the registry and then trigger the deployment to a real-time endpoint based on predefined steps.

  3. C

    The team must manually register the model in the Model Registry and deploy it using the AWS Management Console.

  4. D

    SageMaker Model Registry automatically deploys the model to a real-time endpoint without any further configuration.

Xem giải thích

Đáp án

B — SageMaker Pipelines tự động hoá việc đăng ký mô hình vào registry

Vì sao đúng

Pipelines có sẵn bước RegisterModel, nên sau khi bước đánh giá đạt ngưỡng, mô hình được đăng ký thành một phiên bản mới trong Model Registry mà không ai phải thao tác. Từ đó bạn đặt trạng thái phê duyệt, và trạng thái đó lại kích hoạt được quy trình triển khai — thành một chuỗi khép kín có kiểm soát.

Vì sao các phương án khác sai

  • A. Không thể tự động triển khai — sai; đó chính là mô hình làm việc mà Pipelines cộng Registry hướng tới.
  • C. Phải đăng ký bằng tay — sai, và đánh mất lợi ích chính của Pipelines.
  • D. Registry tự triển khai lên endpoint — sai theo hướng ngược lại: Registry lưu và quản lý phiên bản, việc triển khai vẫn phải được kích hoạt, thường sau bước phê duyệt.
Câu 292 ML Model Development

You are tasked with building a reinforcement learning model in Amazon SageMaker to optimize warehouse robot navigation. You need to define the reinforcement learning (RL) environment, the agent's actions, and the reward function. How should you formulate the environment and rewards to ensure the robot reaches its destination efficiently?

  1. A

    The environment should simulate a fixed path for the robot, with a reward given only at the destination.

  2. B

    The environment should simulate real-world warehouse conditions with dynamic obstacles, and rewards should be given when the robot makes progress toward the destination, with penalties for collisions or delays.

  3. C

    The environment should simulate fixed navigation tasks, and the reward should be given for maintaining a steady speed throughout the navigation.

  4. D

    The environment should simulate dynamic obstacles, and rewards should be given only if the robot avoids all obstacles.

Xem giải thích

Đáp án

B — Môi trường mô phỏng điều kiện kho thật với vật cản động, và thưởng theo hành vi điều hướng

Vì sao đúng

Chất lượng của tác nhân học tăng cường bị chặn bởi chất lượng môi trường mô phỏng. Kho hàng thật có người đi lại, xe nâng và hàng đặt tạm, nên môi trường phải có vật cản động thì tác nhân mới học được cách né. Huấn luyện trong môi trường tĩnh rồi thả vào kho thật là kịch bản thất bại kinh điển — tác nhân chưa từng gặp tình huống nó sẽ gặp mỗi ngày.

Vì sao các phương án khác sai

  • A và C. Mô phỏng đường đi cố định — tác nhân chỉ học thuộc một lộ trình, không học được cách điều hướng.
  • D. Có vật cản động nhưng chỉ thưởng khi tới đích — phần thưởng quá thưa: robot đi hàng nghìn bước mới nhận một tín hiệu, nên gần như không học được. Cần thêm thưởng trung gian cho tiến bộ từng bước.
Câu 293 Chọn nhiều đáp án ML Model Development

A retail company is using Amazon Personalize to provide personalized product recommendations based on real-time customer behavior on their e-commerce platform. They want to integrate Amazon DynamoDB to store customer session data, use Amazon Kinesis for real-time event streaming, and leverage Amazon S3 for historical data storage. Additionally, they need the solution to scale during peak shopping periods, such as holidays. Which of the following strategies should they implement? (Choose TWO.)

  1. A

    Use Amazon Personalize to directly process real-time streaming data from Amazon Kinesis and Amazon DynamoDB.

  2. B

    Use Amazon EC2 Auto Scaling to manage the increasing number of personalized recommendation requests during peak traffic.

  3. C

    Use AWS Lambda to trigger real-time updates to Amazon Personalize when new customer events are detected in Amazon DynamoDB.

  4. D

    Store session data in Amazon RDS for high availability and scalability during peak traffic.

  5. E

    Store historical customer interaction data in Amazon S3 and periodically feed it into Amazon Personalize for batch training.

Xem giải thích

Đáp án

C và E — Lambda cập nhật thời gian thực, và S3 lưu dữ liệu lịch sử để huấn luyện lại định kỳ

Vì sao đúng

Amazon Personalize làm việc ở hai nhịp và đề cần cả hai:

  • C. Lambda gửi sự kiện thời gian thực qua event tracker khi khách vừa xem hay vừa mua, nhờ vậy gợi ý phản ánh được phiên hiện tại chứ không chỉ hành vi tuần trước.
  • E. Dữ liệu lịch sử trên S3 dùng để huấn luyện lại giải pháp theo định kỳ — đó là nơi mô hình học ra các quy luật dài hạn.

Vì sao các phương án khác sai

  • A. Personalize tự đọc luồng thời gian thực — nó không tự nối vào luồng; phải có bên đẩy sự kiện vào, và đó là vai của Lambda.
  • B. EC2 Auto Scaling cho số gợi ý tăng — Personalize là dịch vụ được quản lý, tự co giãn.
  • D. Lưu dữ liệu phiên trong RDS — thêm một CSDL phải vận hành mà không giải quyết gì.
Câu 294 Deployment and Orchestration of ML Workflows

A data engineer is tasked with building a real-time streaming data pipeline where data from IoT devices is continuously ingested and needs to be distributed to multiple consumers for further processing and analysis. However, the engineer is concerned about potential bottlenecks due to multiple consumers sharing the same read throughput. What feature of Kinesis Data Streams can address this concern?

  1. A

    Shard Splitting

  2. B

    Enhanced Fan-out

  3. C

    Partition Key

  4. D

    Kinesis Producer Library

Xem giải thích

Đáp án

B — Enhanced Fan-out

Vì sao đúng

Đề nêu đúng hai điều kiện: nhiều consumer cùng đọc, và cần độ trễ thấp. Không bật fan-out thì các consumer chia nhau 2 MB/giây của mỗi shard và phải hỏi liên tục. Enhanced fan-out cấp cho mỗi consumer đường riêng 2 MB/giây trên mỗi shard và đẩy dữ liệu sang, đưa độ trễ về khoảng 70 mili giây.

Vì sao các phương án khác sai

  • A. Tách shard — tăng tổng thông lượng nhưng các consumer vẫn giành nhau phần của từng shard.
  • C. Partition key — quyết định bản ghi rơi vào shard nào, không liên quan tới phía đọc.
  • D. Kinesis Producer Library — thư viện cho bên ghi vào, không phải bên đọc.
Câu 295 Deployment and Orchestration of ML Workflows

A data scientist is using Amazon SageMaker for a machine learning project. They are deciding between two options: using SageMaker notebook instances or SageMaker Studio notebooks. The project involves running multiple notebooks and requires easy collaboration with other team members.

Which of the following should the data scientist choose?

  1. A

    Launch Jupyter notebooks on EC2 instances for better customization and control.

  2. B

    Use local Jupyter notebooks and store results in an Amazon S3 bucket for collaboration.

  3. C

    SageMaker notebook instances, as they operate as standalone environments with flexible configurations.

  4. D

    SageMaker Studio notebooks, as they provide a collaborative environment with persistent storage and integrated tools.

Xem giải thích

Đáp án

D — Studio notebook, vì có môi trường cộng tác và lưu trữ bền

Vì sao đúng

Khi đề nhấn vào cộng tác và lưu trữ giữ nguyên qua các phiên, Studio là câu trả lời: nhiều người dùng trong một domain, chia sẻ notebook kèm ngữ cảnh, đổi loại máy mà không mất công việc đang làm, và mở thẳng các công cụ khác trong cùng giao diện.

Vì sao các phương án khác sai

  • C. Notebook instance — là máy đơn lẻ gắn với một người; hợp khi cần toàn quyền tuỳ biến máy, nhưng không có mô hình nhiều người dùng.
  • A. EC2 tự cài Jupyter — phải tự vận hành mọi thứ.
  • B. Jupyter cục bộ, kết quả để trên S3 — không dùng được tài nguyên đám mây cho huấn luyện, và "cộng tác" chỉ dừng ở việc chép tệp qua lại.

Ghi chú

Có câu khác trong bộ đề hỏi môi trường Jupyter được quản lý cho một người, và đáp án ở đó lại là notebook instance. Hai câu rất giống nhau; điều phân biệt là có yêu cầu nhiều người dùng hay không.

Câu 296 Deployment and Orchestration of ML Workflows

A company wants to automate the processing of onboarding documents submitted by new employees. These documents include scanned images of identification cards and PDF files containing contracts and employee information. The company needs to extract relevant details such as employee names, addresses, and job roles from both images and text documents, and categorize the documents based on content (e.g., ID card, contract, tax form). The system should also enable easy searching of these documents using natural language queries.

Which combination of AWS services would BEST meet the company’s needs?

  1. A

    Amazon Textract for text extraction from images and documents, Amazon Comprehend for entity recognition and document categorization, and Amazon Kendra for search functionality.

  2. B

    Amazon Comprehend for extracting text from images, Amazon Kendra for categorizing documents, and Amazon Elasticsearch for indexing and searching the documents.

  3. C

    Amazon Rekognition for image analysis, Amazon SageMaker for building a custom entity recognition model, and Amazon DynamoDB for storing the document metadata.

  4. D

    Amazon Rekognition for text and image analysis, Amazon Comprehend for entity extraction, and Amazon S3 for storing the categorized documents.

Xem giải thích

Đáp án

A — Textract trích văn bản từ ảnh và tài liệu, Comprehend nhận thực thể

Vì sao đúng

Hồ sơ gồm ảnh quét giấy tờ tuỳ thân và tệp PDF, nên bước đầu bắt buộc là chuyển thành văn bản. Textract đọc được cả chữ in lẫn chữ viết tay và giữ được cấu trúc biểu mẫu — cặp khoá–giá trị và bảng, thứ rất quan trọng với giấy tờ hành chính. Sau đó Comprehend rút ra thực thể như tên, ngày tháng, số hiệu.

Vì sao các phương án khác sai

  • B. Comprehend để trích văn bản từ ảnh — Comprehend làm việc trên văn bản, không đọc ảnh.
  • C và D. Rekognition để phân tích tài liệu — Rekognition nhận diện vật thể, khuôn mặt và chữ ngắn trên ảnh; nó không dựng cho tài liệu nhiều trang có cấu trúc.
Câu 297 Deployment and Orchestration of ML Workflows

A machine learning engineer is tasked with automating the machine learning workflow, from data ingestion to model deployment. Which AWS services should the engineer use to automate the end-to-end ML lifecycle efficiently?


  1. A

    Manually write Python scripts to manage data ingestion, model development, and deployment on EC2 instances.

  2. B

    Use SageMaker Feature Store for all tasks, including data collection, model training, and deployment.

  3. C

    Use AWS Glue for model development and SageMaker Pipelines for monitoring.

  4. D

    Use SageMaker Data Wrangler for data preparation, SageMaker Model Registry for model versioning, and SageMaker Pipelines to orchestrate the workflow.

Xem giải thích

Đáp án

D — Data Wrangler chuẩn bị dữ liệu, Model Registry quản lý mô hình, Pipelines điều phối

Vì sao đúng

Đề cần tự động hoá toàn bộ chuỗi từ nhận dữ liệu tới triển khai, và ba thành phần này phủ đúng ba chặng: Data Wrangler cho khâu chuẩn bị và xuất ra được thành bước chạy lại được; Model Registry giữ phiên bản mô hình cùng trạng thái phê duyệt; Pipelines nối tất cả thành một DAG chạy tự động.

Vì sao các phương án khác sai

  • A. Tự viết script Python quản lý mọi thứ — chính là thứ cần tự động hoá, và không có phiên bản hay lịch sử.
  • B. Dùng Feature Store cho mọi việc — Feature Store chỉ quản đặc trưng, không huấn luyện và không triển khai.
  • C. Glue để phát triển mô hình — Glue là ETL, không phải nền tảng học máy.
Câu 298 ML Model Development

A machine learning team is training a deep learning model using Amazon SageMaker and wants to monitor the training process in real time to identify issues like vanishing gradients and overfitting. They also need to automatically track and compare different training experiments, including hyperparameter configurations and model performance, to optimize the model.

Which combination of SageMaker features should the team use to achieve these goals?

  1. A

    Use SageMaker Debugger for real-time training monitoring and SageMaker Experiments to track and compare training runs

  2. B

    Use SageMaker Neo to optimize model performance and SageMaker Feature Store to manage training data features

  3. C

    Use SageMaker Ground Truth for data labeling and SageMaker Autopilot for model optimization

  4. D

    Use SageMaker Studio to manually track training metrics and SageMaker Data Wrangler to handle training data

Xem giải thích

Đáp án

A — SageMaker Debugger để theo dõi lúc huấn luyện, Experiments để so các lần chạy

Vì sao đúng

Hai công cụ ở hai tầng khác nhau và đề cần cả hai. Debugger nhìn vào bên trong một lần huấn luyện: lấy mẫu tensor và gradient theo thời gian thực, có luật dựng sẵn bắt gradient biến mất hay quá khớp, và dừng job khi luật kích hoạt. Experiments nhìn giữa các lần chạy: ghi lại siêu tham số và chỉ số để so xem cấu hình nào tốt hơn.

Vì sao các phương án khác sai

  • B. Neo và Feature Store — Neo biên dịch mô hình đã xong cho phần cứng đích; Feature Store quản đặc trưng. Không cái nào theo dõi lúc huấn luyện.
  • C. Ground Truth và Autopilot — gán nhãn và tự dựng mô hình, thuộc giai đoạn trước.
  • D. Theo dõi thủ công trong Studio — Studio hiển thị chỉ số, nhưng "thủ công" là bỏ mất phần tự phát hiện mà Debugger cho.
Câu 299 Data Preparation for Machine Learning

Which type of machine learning is most suitable for grouping products into categories based on customer reviews without knowing the product categories beforehand?

  1. A

    Reinforcement learning

  2. B

    Supervised learning

  3. C

    Unsupervised learning

  4. D

    Regression

Xem giải thích

Đáp án

C — Học không giám sát

Vì sao đúng

Chi tiết quyết định nằm ở cụm "chưa biết trước các nhóm sản phẩm". Không có nhãn thì không có gì để mô hình học theo, nên bài toán thuộc nhóm không giám sát — cụ thể là phân cụm. Thuật toán tự tìm ra cấu trúc trong dữ liệu đánh giá và gom những sản phẩm được nói tới theo cách tương tự nhau.

Vì sao các phương án khác sai

  • B. Học có giám sát — cần cặp đầu vào và nhãn đúng cho từng mẫu; ở đây không có.
  • D. Hồi quy — là một dạng học có giám sát, dự đoán giá trị số liên tục.
  • A. Học tăng cường — học từ phần thưởng qua tương tác với môi trường; không có môi trường nào ở đây.

Nhớ nhanh

Đề nói "không biết trước nhãn" hoặc "tự tìm nhóm" thì gần như luôn là học không giám sát.

Câu 300 Deployment and Orchestration of ML Workflows

You are tasked with building an AI application using Amazon Bedrock. The application will experience variable traffic, with periods of very high usage followed by long idle times. Which pricing option is the MOST cost-effective for this use case?

  1. A

    Reserved instance pricing

  2. B

    On-demand pricing

  3. C

    Spot instance pricing

  4. D

    Provisioned throughput mode

Xem giải thích

Đáp án

B — Giá theo nhu cầu (on-demand)

Vì sao đúng

Đề mô tả tải lên xuống thất thường: cao điểm rồi nghỉ dài. On-demand của Bedrock tính tiền theo số token thật sự dùng, nên lúc không ai gọi thì gần như không tốn. Đó là mô hình đúng khi không đoán trước được lưu lượng.

Vì sao các phương án khác sai

  • D. Provisioned throughput — mua sẵn một mức thông lượng và trả tiền theo giờ dù có dùng hay không; chỉ đáng khi tải cao và ổn định, ngược hẳn tình huống đề nêu.
  • A. Reserved instance và C. Spot instance — là mô hình giá của EC2, Bedrock không có hai lựa chọn này.