Ngân hàng đề — AWS Certified Machine Learning Engineer Associate
Tìm thấy 635 câu.
A machine learning engineer is training a deep learning model on SageMaker and notices that the model performs well on the training data but poorly on the validation set, indicating overfitting. The engineer wants to use SageMaker Debugger to automatically detect overfitting and take corrective actions.
Which of the following SageMaker Debugger features can help identify and resolve overfitting issues during training?
-
A
Set up SageMaker Profiler to track GPU memory utilization and adjust the batch size to prevent overfitting.
-
B
Use built-in SageMaker Debugger rules to monitor the loss curve and alert when overfitting is detected.
-
C
Monitor I/O operations using SageMaker Profiler to detect overfitting and dynamically adjust the learning rate.
-
D
Enable real-time alerts in SageMaker Debugger when CPU utilization exceeds a threshold to detect overfitting.
Xem giải thích
Đáp án
B — Dùng luật dựng sẵn của SageMaker Debugger để theo dõi đường mất mát và cảnh báo khi quá khớp
Vì sao đúng
Triệu chứng đề nêu — tốt trên tập huấn luyện, kém trên tập kiểm định — là quá khớp, và Debugger có sẵn luật overfit để bắt đúng nó. Luật này so đường mất mát của hai tập: khi mất mát huấn luyện tiếp tục giảm mà mất mát kiểm định bắt đầu tăng, nó kích hoạt. Gắn thêm hành động thì job tự dừng, khỏi trả tiền cho những giờ huấn luyện chỉ làm mô hình tệ đi.
Vì sao các phương án khác sai
- A và C. Dùng Profiler — Profiler đo mức dùng tài nguyên (GPU, I/O); nó không nhìn vào chất lượng mô hình nên không phát hiện quá khớp.
- D. Cảnh báo khi CPU vượt ngưỡng — chỉ số hạ tầng, hoàn toàn không liên quan tới quá khớp.
A machine learning team is running several trials using Amazon SageMaker Experiments to identify the best model configuration for their project. After multiple trials, they need to visualize and compare the accuracy and precision of each trial to select the optimal model.
What feature of SageMaker Experiments will allow them to easily perform this comparison?
-
A
SageMaker Studio's graphical interface
-
B
SageMaker Pipelines
-
C
Trackers in SageMaker CLI
-
D
CloudWatch Metrics Dashboard
Xem giải thích
Đáp án
A — Giao diện đồ hoạ của SageMaker Studio
Vì sao đúng
Studio có sẵn màn hình cho Experiments: liệt kê mọi trial, hiện siêu tham số và chỉ số thành bảng, cho sắp xếp và lọc, và vẽ biểu đồ so sánh nhiều lần chạy cạnh nhau. Đó là cách nhanh nhất để nhìn ra cấu hình nào thắng, không phải viết dòng mã nào.
Vì sao các phương án khác sai
- B. Pipelines — điều phối quy trình, không phải công cụ so sánh kết quả.
- C. Tracker trong SageMaker CLI — tracker dùng để ghi dữ liệu vào lúc chạy, không phải để xem lại và so sánh.
- D. Bảng điều khiển CloudWatch Metrics — xem được chỉ số thô nhưng không hiểu cấu trúc experiment và trial, nên so sánh rất bất tiện.
A retail company has deployed a demand forecasting model using Amazon SageMaker. They want to ensure the model remains accurate over time by monitoring both data quality and model performance. They also need to automate the retraining of the model if significant data drift or performance degradation is detected. Which combination of AWS services and features would best help the company achieve this?
-
A
Use SageMaker Ground Truth for monitoring data drift, Amazon CloudWatch for logging, and AWS Step Functions for model deployment
-
B
Use SageMaker Data Wrangler for detecting data drift, AWS Glue for data processing, and SageMaker Debugger for retraining the model
-
C
Use SageMaker Neo for monitoring model performance, AWS Lambda for data drift detection, and Amazon Kinesis for real-time predictions
-
D
Use SageMaker Model Monitor for detecting data drift, Amazon CloudWatch for setting alerts, and SageMaker Pipelines for automating model retraining
Xem giải thích
Đáp án
D — Model Monitor phát hiện trôi dữ liệu, CloudWatch đặt cảnh báo
Vì sao đúng
Cặp này là cách chuẩn để giám sát mô hình đã triển khai. Model Monitor chạy theo lịch, so dữ liệu thật với thống kê nền và ghi kết quả vi phạm ra CloudWatch; ở đó bạn đặt alarm để báo cho người trực hoặc kích hoạt quy trình huấn luyện lại. Đúng cả hai vế đề nêu: chất lượng dữ liệu và hiệu năng mô hình.
Vì sao các phương án khác sai
- A. Ground Truth để theo dõi trôi dữ liệu — Ground Truth là dịch vụ gán nhãn.
- B. Data Wrangler để phát hiện trôi — công cụ chuẩn bị dữ liệu trước khi huấn luyện.
- C. SageMaker Neo để theo dõi hiệu năng mô hình — Neo biên dịch tối ưu mô hình cho phần cứng đích; nó không giám sát gì.
An e-commerce company has deployed a web application on AWS using multiple EC2 instances behind an Elastic Load Balancer (ELB). The company wants to monitor real-time performance metrics of its application, such as CPU utilization, memory usage, and request latency, and receive alerts if any thresholds are exceeded. Which AWS services and features should be used to meet these requirements?
-
A
Set up AWS Lambda functions to periodically pull CPU and memory metrics from EC2 and push them to CloudWatch for monitoring and alerting.
-
B
Use AWS CloudTrail to capture EC2 instance activity and automatically alert on high CPU or memory usage.
-
C
Use AWS CloudWatch to monitor CPU utilization and set CloudWatch Alarms for threshold breaches. Install the CloudWatch Agent on EC2 instances to monitor memory usage and other custom metrics.
-
D
Enable AWS Config to track the CPU and memory metrics, and set up SNS notifications for threshold alerts.
Xem giải thích
Đáp án
C — Dùng CloudWatch theo dõi mức dùng CPU và đặt CloudWatch Alarm theo ngưỡng
Vì sao đúng
CloudWatch là dịch vụ chỉ số của AWS: EC2 đẩy sẵn mức dùng CPU vào đó, và alarm kích hoạt khi vượt ngưỡng trong khoảng thời gian bạn đặt, rồi bắn thông báo qua SNS hoặc kích hoạt hành động co giãn. Đây là con đường ngắn nhất và cũng là cách AWS thiết kế sẵn.
Lưu ý thực tế: bộ nhớ không nằm trong chỉ số mặc định của EC2 — phải cài CloudWatch agent mới thu được.
Vì sao các phương án khác sai
- A. Lambda định kỳ hỏi rồi đẩy chỉ số — tự viết lại thứ đã có sẵn, và trễ hơn.
- B. CloudTrail để cảnh báo CPU cao — CloudTrail ghi lời gọi API, không thu chỉ số hiệu năng.
- D. AWS Config theo dõi CPU — Config theo dõi cấu hình tài nguyên, không phải hiệu năng.
A machine learning engineer is setting up a Jupyter notebook environment to streamline an end-to-end machine learning workflow, including model training and deployment. The engineer wants to minimize manual infrastructure management and use a fully integrated environment within Amazon SageMaker.
Which option should the engineer choose to achieve this?
-
A
Set up a SageMaker notebook instance from the AWS Management Console and attach an IAM role.
-
B
Manually provision an EC2 instance with Jupyter notebooks installed and manage the environment through AWS CLI.
-
C
Install SageMaker libraries on a local machine and run Jupyter notebooks locally.
-
D
Use SageMaker Studio to set up a domain and configure multiple Jupyter Lab notebooks within the same workspace.
Xem giải thích
Đáp án
D — Dùng SageMaker Studio, tạo domain và cấu hình nhiều không gian JupyterLab
Vì sao đúng
Đề cần một môi trường liền mạch từ huấn luyện tới triển khai. Studio dựng quanh khái niệm domain và user profile: mỗi người có không gian riêng, lưu trữ bền qua các phiên, đổi loại máy tính toán mà không mất công việc đang làm, và mở thẳng được các công cụ khác — Experiments, Pipelines, Model Registry — trong cùng giao diện.
Vì sao các phương án khác sai
- A. Notebook instance từ console — chạy được, nhưng là máy đơn lẻ, không có mô hình nhiều không gian và phải tự nối tới các công cụ khác.
- B. Tự dựng EC2 với Jupyter — phải tự cài, tự vá, tự bảo mật.
- C. Chạy Jupyter trên máy cá nhân — không dùng được tài nguyên đám mây cho huấn luyện, và không triển khai liền mạch được.
Which of the following types of drift is best described as a systematic favoring of certain groups in the model's predictions over time?
-
A
Data drift
-
B
Model quality drift
-
C
Feature attribution drift
-
D
Bias drift
Xem giải thích
Đáp án
D — Bias drift (trôi thiên lệch)
Vì sao đúng
Đề mô tả đúng định nghĩa: mô hình ưu ái một cách có hệ thống một số nhóm nhất định, và mức ưu ái đó thay đổi theo thời gian. Nguyên nhân thường là dân số thực tế đã đổi so với dữ liệu huấn luyện. SageMaker Clarify tính các chỉ số thiên lệch theo nhóm, còn Model Monitor theo dõi chúng theo lịch và cảnh báo khi vượt ngưỡng.
Vì sao các phương án khác sai
- A. Data drift — phân bố dữ liệu vào thay đổi, không nói gì về sự công bằng giữa các nhóm.
- B. Model quality drift — độ chính xác tổng thể giảm; mô hình có thể vẫn chính xác cao mà vẫn thiên lệch nặng với một nhóm nhỏ.
- C. Feature attribution drift — mức đóng góp của từng đặc trưng thay đổi.
A company is preparing a machine learning pipeline using Amazon SageMaker Data Wrangler. The data engineer needs to integrate multiple data sources, clean and transform the data, and output it to an Amazon S3 bucket for training. What steps should the engineer take to ensure the data is properly prepared for training in SageMaker?
-
A
Export the transformed dataset from Data Wrangler directly to SageMaker Feature Store for training.
-
B
Use SageMaker Data Wrangler to create a pipeline, automatically export the transformed data to S3 in Apache Parquet format, and then use SageMaker training jobs to access the data.
-
C
Clean and transform the data using SageMaker Processing jobs, and then manually upload it to an Amazon S3 bucket.
-
D
Use Data Wrangler to import, clean, and transform the data, and then export the final dataset to Amazon S3 in CSV format.
Xem giải thích
Đáp án
B — Dùng Data Wrangler tạo pipeline và tự động xuất dữ liệu đã biến đổi
Vì sao đúng
Data Wrangler không dừng ở việc làm sạch một lần: luồng biến đổi bạn dựng bằng chuột xuất ra được thành pipeline hoặc notebook, nối thẳng sang Pipelines hay Feature Store. Nhờ vậy công việc khám phá bằng tay biến thành quy trình lặp lại được, không có bước chép dữ liệu thủ công nào ở giữa.
Vì sao các phương án khác sai
- A. Xuất thẳng sang Feature Store — làm được, nhưng thiếu phần tự động hoá mà đề nhấn mạnh.
- C. Dùng Processing job rồi nối tay — bỏ qua đúng thế mạnh của Data Wrangler và thêm bước thủ công.
- D. Xuất ra rồi tự xử lý tiếp — cũng để lại khoảng trống thủ công giữa hai bước.
An organization needs to develop and deploy a machine learning model for real-time customer behavior prediction using SageMaker. They want a solution that provides persistent storage, allows collaborative development, and makes it easy to track and manage multiple experiments in an integrated interface.
Which of the following best meets the organization's needs?
-
A
SageMaker Studio
-
B
SageMaker Neo
-
C
SageMaker Autopilot
-
D
SageMaker Ground Truth
Xem giải thích
Đáp án
A — SageMaker Studio
Vì sao đúng
Từ khoá của đề là lưu trữ bền và một môi trường phát triển đầy đủ cho cả vòng đời. Studio cho mỗi người dùng một không gian lưu trữ giữ nguyên qua các phiên, đổi loại máy mà không mất công việc, và mở thẳng được Experiments, Pipelines, Model Registry cùng endpoint suy luận thời gian thực trong cùng một chỗ.
Vì sao các phương án khác sai
- B. SageMaker Neo — biên dịch tối ưu mô hình cho phần cứng đích, không phải môi trường phát triển.
- C. Autopilot — tự dựng mô hình từ dữ liệu bảng; là một tính năng, không phải nơi làm việc.
- D. Ground Truth — dịch vụ gán nhãn dữ liệu.
A customer support system uses both Amazon Kendra for retrieving relevant support documents and Amazon Lex for handling voice and text interactions. The company wants to ensure that their Kendra index is always up to date with new documents from an S3 bucket, and that Lex handles fallback intents when the user's question is unclear.
What solution ensures that both services are working together optimally?
-
A
Use AWS Glue to continuously update the Kendra index and rely on Lex's speech recognition for fallback intent handling.
-
B
Use Amazon EventBridge to trigger Lambda functions that update the Kendra index and implement Lex's fallback intent for handling ambiguous queries.
-
C
Implement AWS CloudTrail to monitor the Kendra index and integrate it with Lex’s natural language processing for ambiguous questions.
-
D
Use Amazon S3 Event Notifications to directly update the Kendra index and set up manual intents in Lex for ambiguous queries.
Xem giải thích
Đáp án
B — Dùng EventBridge kích hoạt Lambda để cập nhật chỉ mục Kendra
Vì sao đúng
Vấn đề là giữ chỉ mục Kendra luôn mới khi tài liệu hỗ trợ thay đổi. EventBridge nhận sự kiện từ nhiều nguồn theo mẫu bạn khai, rồi gọi Lambda để đồng bộ đúng phần đã đổi. Cách này tách bạch: nguồn phát sự kiện không cần biết gì về Kendra, và thêm nguồn mới chỉ là thêm một quy tắc.
Vì sao các phương án khác sai
- A. Dùng Glue để cập nhật liên tục chỉ mục — Glue là ETL theo mẻ, không phải cơ chế phản ứng theo sự kiện.
- C. CloudTrail giám sát chỉ mục — CloudTrail ghi lời gọi API, không cập nhật gì.
- D. S3 Event Notification cập nhật thẳng Kendra — S3 Event Notification không nhận Kendra làm đích; vẫn phải qua Lambda, và nó chỉ bắt được thay đổi trên S3.
A customer service company is looking to enhance its chatbot by providing real-time voice interaction with users. The bot needs to handle user queries in both text and speech formats. It must convert incoming speech to text, process the text using a chatbot engine, and then convert the bot’s text responses back to speech for the user.
Which combination of services should the company implement?
-
A
Use Amazon Transcribe to convert user speech to text, Amazon Lex to handle the chatbot processing, and Amazon Polly to convert text responses to speech.
-
B
Use Amazon Polly to convert user speech to text, Amazon Lex to process the text queries, and Amazon Transcribe to convert text responses to speech.
-
C
Use Amazon Transcribe to convert user speech to text, AWS Lambda to process the queries, and Amazon Polly to convert text responses to speech.
-
D
Use Amazon Rekognition for speech-to-text conversion, Amazon Lex for chatbot processing, and Amazon Polly for text-to-speech conversion.
Xem giải thích
Đáp án
A — Transcribe chuyển giọng nói thành văn bản, Lex xử lý hội thoại, Polly trả lời bằng giọng nói
Vì sao đúng
Ba dịch vụ ghép thành vòng tương tác giọng nói hoàn chỉnh, mỗi cái đúng vai của nó: Transcribe nhận dạng lời nói; Lex hiểu ý định và quản lý luồng hội thoại — đây là phần "chatbot" thật sự; Polly đọc câu trả lời ra thành tiếng. Người dùng gõ chữ thì bỏ qua bước đầu và bước cuối.
Vì sao các phương án khác sai
- B. Polly để chuyển giọng nói thành văn bản — ngược chiều.
- C. Lambda xử lý hội thoại thay Lex — Lambda chạy được logic nghiệp vụ, nhưng phần hiểu ý định và quản lý luồng phải tự viết lại từ đầu.
- D. Rekognition để nhận dạng giọng nói — Rekognition phân tích ảnh và video, không xử lý âm thanh.