Ngân hàng đề — AWS Certified Machine Learning Engineer Associate

Tìm thấy 635 câu.

Câu 371 Chọn nhiều đáp án Deployment and Orchestration of ML Workflows

A Machine Learning engineer is building an e-commerce chatbot using Amazon Bedrock. The chatbot must handle real-time customer queries, generate personalized recommendations, and operate at scale without infrastructure management. The engineer also needs to fine-tune the foundational model with historical customer data and ensure the security of this data.
Which combination of services should the engineer use to achieve these requirements? (Choose TWO)

  1. A

    Use Amazon Bedrock for deploying foundation models with serverless infrastructure and fine-tuning capabilities

  2. B

    Use AWS Lambda for model deployment, which enables the scaling of real-time inference for customer queries

  3. C

    Use Amazon S3 for storing historical customer data and generated model outputs

  4. D

    Use Amazon SageMaker to deploy and fine-tune foundation models, ensuring customization of machine learning models

  5. E

    Use AWS CloudTrail to monitor model inference processes and ensure fine-grained data access control

Xem giải thích

Đáp án

A và C — Bedrock triển khai mô hình nền không máy chủ, và S3 lưu dữ liệu

Vì sao đúng

  • A. Bedrock cho gọi mô hình nền qua API mà không cấp phát hạ tầng nào, tự co giãn theo lưu lượng — đúng thứ chatbot cần khi lượng khách lên xuống thất thường.
  • C. S3 lưu lịch sử tương tác và kết quả sinh ra: rẻ, bền, và là nguồn dữ liệu cho việc tinh chỉnh hoặc phân tích sau này.

Vì sao các phương án khác sai

  • B. Dùng Lambda để triển khai mô hình — Lambda gọi được Bedrock, nhưng nó không "triển khai mô hình"; mô hình nằm ở phía Bedrock.
  • D. Dùng SageMaker để triển khai và tinh chỉnh — làm được, nhưng đề đã chọn Bedrock và SageMaker kéo theo việc quản lý endpoint, trái yêu cầu không máy chủ.
  • E. CloudTrail để giám sát suy luận — CloudTrail ghi lời gọi API, hữu ích cho kiểm toán nhưng không phải mảnh ghép cốt lõi.
Câu 372 ML Solution Monitoring, Maintenance, and Security

A machine learning team wants to track the history of model versions, including training data, hyperparameters, and performance metrics over time. Which feature of SageMaker should they use?

  1. A

    SageMaker Model Registry

  2. B

    SageMaker Notebook Instances

  3. C

    SageMaker Training Jobs

  4. D

    SageMaker Pipelines

Xem giải thích

Đáp án

A — SageMaker Model Registry

Vì sao đúng

Model Registry là nơi quản lý vòng đời của mô hình theo phiên bản: mỗi lần đăng ký tạo ra một model package trong một model group, kèm siêu dữ liệu về dữ liệu huấn luyện, siêu tham số, chỉ số đánh giá, và trạng thái phê duyệt. Nhờ vậy trả lời được câu "bản đang chạy sản xuất được huấn luyện từ dữ liệu nào" — thứ bắt buộc khi bị kiểm toán.

Vì sao các phương án khác sai

  • B. Notebook Instances — môi trường phát triển.
  • C. Training Jobs — mỗi job lưu thông tin của một lần chạy, nhưng không có khái niệm phiên bản và phê duyệt xuyên suốt.
  • D. Pipelines — điều phối quy trình; nó ghi vào registry chứ bản thân không phải nơi lưu.
Câu 373 Data Preparation for Machine Learning

A financial institution is using Amazon Athena to run SQL queries on its transaction logs stored in Amazon S3. They need a scalable way to keep track of schema changes in these logs, as the structure of the logs may change over time. Additionally, they want to ensure the queries are fast and cost-effective.

Which of the following is the most effective approach to ensure efficient schema management and optimized query performance?

  1. A

    Use AWS Glue Crawlers to automatically detect schema changes and store the metadata in the Glue Data Catalog for Amazon Athena to query.

  2. B

    Enable S3 Select to query only the necessary parts of the logs, ensuring that schema changes are handled automatically.

  3. C

    Create an Amazon Athena table with a hardcoded schema and manually update it whenever the log structure changes.

  4. D

    Store the transaction logs in Amazon RDS and create a pipeline to replicate them to Amazon Athena for better querying capabilities.

Xem giải thích

Đáp án

A — Dùng Glue Crawler tự phát hiện thay đổi lược đồ và lưu siêu dữ liệu vào catalog

Vì sao đúng

Log giao dịch thường thêm trường mới theo thời gian. Crawler chạy theo lịch, so với lược đồ đã đăng ký, và cập nhật Data Catalog khi phát hiện khác biệt — có cả tuỳ chọn quyết định xử lý cột mới ra sao. Athena đọc lược đồ từ catalog nên truy vấn tự nhìn thấy trường mới mà không ai phải sửa gì.

Vì sao các phương án khác sai

  • B. Bật S3 Select — đọc một phần của một đối tượng; không quản lý lược đồ và không thay được Athena cho truy vấn trên toàn bộ tập dữ liệu.
  • C. Khai lược đồ cứng rồi sửa tay — chính là việc thủ công mà đề muốn tránh.
  • D. Chuyển log sang RDS — đổi hẳn kiến trúc và mất lợi thế truy vấn tại chỗ trên S3.
Câu 374 Data Preparation for Machine Learning

A machine learning team is using Amazon SageMaker Experiments to evaluate different model configurations. They want to track and compare the performance of several training runs with varying hyperparameters, data features, and algorithms to identify the best performing model. What components in SageMaker Experiments should they use to organize and track these training runs?

  1. A

    Pipelines, Trials, Stages, Logs

  2. B

    Experiments, Trials, Trial Components, Trackers

  3. C

    Hyperparameters, Experiments, Trackers, Notebooks

  4. D

    Trackers, Algorithms, Metrics, Pipelines

Xem giải thích

Đáp án

B — Experiments, Trials, Trial Components, Trackers

Vì sao đúng

Bốn khái niệm xếp thành ba tầng cộng một công cụ ghi:

  • Experiment là cả dự án, ví dụ "dự báo rời bỏ khách hàng".
  • Trial là một lần chạy với một cấu hình cụ thể.
  • Trial component là một bước trong trial đó — tiền xử lý, huấn luyện, đánh giá.
  • Tracker là thứ bạn gọi trong mã để ghi siêu tham số và chỉ số vào các thành phần trên.

Vì sao các phương án khác sai

  • A — "Stages" và "Logs" không phải khái niệm của Experiments.
  • C và D — trộn lẫn siêu tham số, thuật toán và notebook vào; chúng là thứ được ghi lại chứ không phải cấu trúc tổ chức.
Câu 375 ML Model Development

Which of the following features of SageMaker Jupyter Notebooks helps ensure that data is not lost during the machine learning workflow?

  1. A

    Persistent storage

  2. B

    On-demand instance provisioning

  3. C

    Real-time inference endpoints

  4. D

    Pre-installed libraries like TensorFlow and PyTorch

Xem giải thích

Đáp án

A — Lưu trữ bền (persistent storage)

Vì sao đúng

Notebook của SageMaker gắn với một ổ EBS giữ nguyên nội dung khi dừng và khởi động lại máy. Nhờ vậy mã, dữ liệu và kết quả trung gian không mất khi bạn tắt máy để tiết kiệm chi phí — thói quen rất phổ biến vì máy huấn luyện đắt.

Vì sao các phương án khác sai

  • B. Cấp phát máy theo yêu cầu — nói về tính linh hoạt, không bảo vệ dữ liệu.
  • C. Endpoint suy luận thời gian thực — thuộc giai đoạn triển khai.
  • D. Thư viện cài sẵn — tiện lợi, nhưng không liên quan tới việc giữ dữ liệu.
Câu 376 ML Model Development

A data scientist is working with Amazon SageMaker Studio and needs to set up a JupyterLab environment for collaboration with their team. Which of the following best describes how SageMaker Studio notebooks are different from standalone SageMaker notebook instances?

  1. A

    SageMaker Studio notebooks are integrated within SageMaker Studio, allowing persistent storage and multiple notebooks per project, while standalone notebook instances operate independently.

  2. B

    SageMaker Studio notebooks are pre-configured only for data preparation tasks and cannot be used for model training or deployment.

  3. C

    SageMaker Studio notebooks are for lightweight tasks, whereas standalone SageMaker notebook instances are used for heavy machine learning workloads.

  4. D

    Standalone SageMaker notebook instances have built-in collaboration tools, whereas SageMaker Studio notebooks do not support collaboration.

Xem giải thích

Đáp án

A — Studio notebook nằm trong Studio, cho phép chia sẻ và cộng tác

Vì sao đúng

Studio notebook không phải một máy riêng lẻ mà là không gian bên trong domain của Studio. Nhờ vậy chia sẻ được cho đồng đội kèm cả ngữ cảnh chạy, đổi loại máy tính toán mà không mất công việc đang làm, và mở thẳng các công cụ khác trong cùng giao diện.

Vì sao các phương án khác sai

  • B. Chỉ dựng cho chuẩn bị dữ liệu — sai; dùng cho cả vòng đời.
  • C. Studio chỉ hợp việc nhẹ, notebook instance mới mạnh — sai; Studio đổi được sang máy rất mạnh khi cần.
  • D. Notebook instance có sẵn công cụ cộng tác — sai theo hướng ngược lại: đó chính là điểm Studio hơn hẳn.
Câu 377 Chọn nhiều đáp án ML Solution Monitoring, Maintenance, and Security

A financial institution's machine learning model in production shows a significant drop in prediction accuracy over the last month. Upon investigation, the data science team suspects that the real-world data used for inference has shifted from the training dataset. They want to continuously monitor the quality of the input data to ensure it meets the expected constraints.

Which steps should the team take to monitor the data quality using Amazon SageMaker Model Monitor? (Select TWO)

  1. A

    Manually review each data point for drift after the model has made predictions.

  2. B

    Use AWS CloudTrail to log and evaluate prediction errors.

  3. C

    Define a baseline of the training data to set expected statistical properties.

  4. D

    Enable data capture on the SageMaker endpoint to store incoming requests and predictions.

  5. E

    Set up a batch transform job to retrain the model when data drift is detected.

Xem giải thích

Đáp án

C và D — định nghĩa mốc từ dữ liệu huấn luyện, và bật ghi lại dữ liệu ở endpoint

Vì sao đúng

Model Monitor làm việc bằng cách so sánh với một mốc, nên cần đúng hai điều kiện:

  • C. Mốc (baseline) tính từ dữ liệu huấn luyện: khoảng giá trị, phân bố, tỷ lệ thiếu. Không có mốc thì không có gì để so.
  • D. Data capture để endpoint ghi lại dữ liệu vào và dự đoán ra xuống S3. Không có dữ liệu thật thì job giám sát chạy trong chân không.

Vì sao các phương án khác sai

  • A. Xem tay từng điểm dữ liệu — không mở rộng nổi và luôn muộn.
  • B. Dùng CloudTrail để đánh giá lỗi dự đoán — CloudTrail ghi lời gọi API, không biết gì về chất lượng dự đoán.
  • E. Cứ phát hiện trôi là huấn luyện lại — là phản ứng sau khi giám sát, không phải bước thiết lập giám sát; và huấn luyện lại vô điều kiện là vội vàng.
Câu 378 ML Model Development

You are a data scientist at a healthcare company tasked with building a predictive model using SageMaker. Your team wants to avoid the complexity of managing infrastructure and prefers a solution that minimizes setup time.

Which of the following SageMaker features would be MOST appropriate for your use case?

  1. A

    Use SageMaker’s built-in algorithms, which come pre-packaged in containers, requiring only data input and hyperparameter setup.

  2. B

    Use SageMaker Ground Truth to automatically handle infrastructure setup and dataset labeling.

  3. C

    Build a custom container with proprietary algorithms and manage the infrastructure manually for optimized performance.

  4. D

    Build your own deep learning model using TensorFlow or PyTorch and manually deploy it on an Amazon EC2 instance.

Xem giải thích

Đáp án

A — Dùng thuật toán dựng sẵn, chúng đã đóng gói trong container

Vì sao đúng

Đề nêu rõ muốn tránh phức tạp về hạ tầng. Thuật toán dựng sẵn của SageMaker đi kèm container đã tối ưu: bạn chỉ trỏ tới dữ liệu, khai siêu tham số và loại máy, còn lại dịch vụ lo. Chúng cũng hỗ trợ sẵn huấn luyện phân tán và đọc dữ liệu theo luồng từ S3.

Vì sao các phương án khác sai

  • B. Ground Truth để lo hạ tầng — Ground Truth là dịch vụ gán nhãn, không liên quan tới hạ tầng huấn luyện.
  • C. Tự dựng container và tự quản hạ tầng — trái thẳng yêu cầu.
  • D. Tự viết mô hình rồi tự triển khai — cũng trái yêu cầu, và tốn thời gian hơn nhiều.
Câu 379 Data Preparation for Machine Learning

A machine learning team is using Amazon SageMaker Ground Truth to label a dataset for an object detection task. They want to reduce labeling costs without sacrificing quality. What strategy provided by Ground Truth would BEST help the team achieve this?

  1. A

    Use a completely automated machine learning model to handle all labeling tasks.

  2. B

    Utilize a third-party vendor for all labeling tasks to ensure the highest accuracy.

  3. C

    Implement active learning, where a model labels simple cases, and human labelers handle complex cases.

  4. D

    Fully manual labeling by a private workforce within the organization.

Xem giải thích

Đáp án

C — Dùng học chủ động: mô hình gán nhãn ca dễ, người xử lý ca khó

Vì sao đúng

Học chủ động của Ground Truth cân bằng đúng hai mục tiêu đề nêu. Một mô hình được huấn luyện dần trên chính nhãn đang thu, tự gán nhãn những ảnh nó đã tự tin, và chỉ chuyển cho người những ảnh có độ tin cậy thấp. Tỷ lệ cần người giảm dần theo thời gian, nên chi phí giảm mà chất lượng ở những ca khó vẫn do người quyết định.

Vì sao các phương án khác sai

  • A. Để máy làm hết — chất lượng tụt ở đúng những ca khó, mà với phát hiện vật thể thì ca khó là phần quan trọng nhất.
  • B. Thuê ngoài toàn bộ — đắt, và không tận dụng được phần tự động hoá.
  • D. Người làm hết — chính xác nhưng bỏ qua mục tiêu giảm chi phí.
Câu 380 Data Preparation for Machine Learning

A global retail company collects customer feedback in various languages through email and web forms. The company wants to automatically detect customer sentiment, categorize the feedback into relevant product categories, and translate the feedback into English for further analysis by the product team. The solution should support scalability and automate the entire process.

Which combination of services would BEST meet the company’s needs?

  1. A

    Amazon Comprehend for sentiment analysis, Amazon Translate for translating feedback to English, and Amazon SageMaker for classifying feedback into product categories.

  2. B

    Amazon Rekognition for analyzing the text within images, Amazon Comprehend for categorizing feedback, and Amazon Translate for language translation.

  3. C

    Amazon Translate for translating feedback, Amazon Comprehend for sentiment analysis and entity recognition, and AWS Lambda for automating the process.

  4. D

    Amazon Comprehend for both sentiment analysis and categorization, Amazon Translate for translation, and Amazon Kendra for providing search capabilities across the feedback.

Xem giải thích

Đáp án

C — Translate dịch phản hồi, Comprehend phân tích cảm xúc và phân loại

Vì sao đúng

Thứ tự quan trọng ở đây. Phản hồi đến bằng nhiều ngôn ngữ, nên dịch về một ngôn ngữ chung trước rồi mới phân tích sẽ cho kết quả nhất quán và chỉ phải nuôi một bộ phân loại. Comprehend sau đó lo cả hai việc: chấm điểm cảm xúc và phân loại chủ đề, với bản tuỳ chỉnh thì học được nhóm chủ đề riêng của công ty.

Vì sao các phương án khác sai

  • A — đảo thứ tự: phân tích cảm xúc trước rồi mới dịch thì mỗi ngôn ngữ là một bài toán riêng.
  • B. Rekognition để đọc chữ trong ảnh — phản hồi đến qua email và biểu mẫu web, đã là văn bản.
  • D. Comprehend làm cả hai rồi mới dịch — Comprehend có hỗ trợ vài ngôn ngữ, nhưng dịch sau thì việc phân tích đã diễn ra trên nhiều ngôn ngữ khác nhau, kém nhất quán.