Ngân hàng đề — AWS Certified Machine Learning Engineer Associate
Tìm thấy 635 câu.
A retail company wants to process large volumes of real-time customer interaction data to personalize user experience on their website. They plan to use Amazon Kinesis to ingest, analyze, and store the streaming data. They need to perform complex analytics on the data in real-time and store it for future analysis. Which combination of services should they use to achieve these objectives?
-
A
Use Kinesis Firehose for data ingestion, Kinesis Data Analytics for real-time analysis, and AWS Glue for data storage in Amazon S3.
-
B
Use Kinesis Data Streams for data ingestion, AWS Lambda for real-time analysis, and Kinesis Firehose to deliver the processed data to Amazon Redshift.
-
C
Use Kinesis Data Streams for data ingestion, Kinesis Firehose for real-time analysis, and Kinesis Data Analytics to store the processed data in Amazon RDS.
-
D
Use Kinesis Data Streams for data ingestion, Kinesis Data Analytics for real-time analysis, and Kinesis Firehose to deliver the processed data to Amazon S3.
Xem giải thích
Đáp án
D — Kinesis Data Streams để nhận dữ liệu, Kinesis Data Analytics để phân tích thời gian thực
Vì sao đúng
Hai vai tách bạch: Data Streams là ống nhận dữ liệu tương tác của khách, giữ lại để đọc lại được và cho nhiều consumer cùng đọc. Data Analytics (nay chạy trên Apache Flink) phân tích trên dòng đang chảy — tính theo cửa sổ thời gian, phát hiện mẫu hành vi, và bắn kết quả ra ngay để cá nhân hoá trang web.
Vì sao các phương án khác sai
- A. Firehose để nhận dữ liệu — Firehose là dịch vụ giao nhận, không giữ dữ liệu để đọc lại và không cho nhiều consumer đọc song song.
- B. Lambda để phân tích thời gian thực — Lambda xử lý theo từng bản ghi, rất khó tính toán theo cửa sổ thời gian vì bản thân nó không giữ trạng thái.
- C. Firehose để phân tích — Firehose không phân tích gì; nó chỉ chuyển dữ liệu tới đích.
A machine learning engineer is training a deep learning model on SageMaker using TensorFlow. During training, the model's performance stops improving, and the engineer suspects that the model may be experiencing vanishing gradients. They want to use SageMaker Debugger to monitor the training and resolve this issue. Which steps should the engineer take to identify and address the vanishing gradients using SageMaker Debugger?
-
A
Utilize SageMaker Profiler to collect GPU utilization data and implement early stopping when vanishing gradients are detected.
-
B
Enable SageMaker Debugger to monitor resource utilization and set up alarms for underutilized GPUs.
-
C
Use SageMaker Debugger’s built-in rules to monitor the gradient values during training and adjust the learning rate if vanishing gradients are detected.
-
D
Run a custom SageMaker Debugger rule to visualize CPU and memory usage, then adjust the batch size to optimize resource utilization.
Xem giải thích
Đáp án
C — Dùng luật dựng sẵn của SageMaker Debugger để theo dõi giá trị gradient
Vì sao đúng
Triệu chứng "mô hình ngừng cải thiện" cộng nghi ngờ về gradient chỉ thẳng tới luật VanishingGradient của Debugger. Nó lấy mẫu tensor gradient trong lúc huấn luyện và kích hoạt khi các giá trị nhỏ dần tới mức trọng số gần như không đổi nữa. Gắn thêm hành động thì job tự dừng, khỏi trả tiền cho những giờ huấn luyện vô ích.
Vì sao các phương án khác sai
- A và D. Dùng Profiler — Profiler đo mức dùng tài nguyên (GPU, CPU, I/O); nó không nhìn vào giá trị bên trong mô hình nên không phát hiện được vấn đề gradient.
- B. Debugger nhưng đặt cảnh báo cho tài nguyên — đúng công cụ, sai thứ cần theo dõi.
Nhớ nhanh
Debugger nhìn vào bên trong mô hình, Profiler nhìn vào tài nguyên máy.
A data scientist wants to optimize a machine learning model for deployment across multiple edge devices using SageMaker Neo. After training the model using TensorFlow in SageMaker, what are the next steps the data scientist should take to optimize the model for edge deployment while maintaining performance?
-
A
Use SageMaker Neo to compile the model for the target hardware platform and deploy it using AWS IoT Greengrass.
-
B
Set up SageMaker Debugger to capture real-time metrics during inference on edge devices.
-
C
Use SageMaker Debugger to profile the edge device resource usage and optimize the model based on profiling results.
-
D
Manually rewrite the model’s code to ensure it is compatible with the target edge hardware.
Xem giải thích
Đáp án
A — Dùng SageMaker Neo biên dịch mô hình cho nền tảng phần cứng đích rồi triển khai
Vì sao đúng
Neo sinh ra cho đúng việc này: nhận mô hình đã huấn luyện và biên dịch lại cho kiến trúc cụ thể của thiết bị biên. Kết quả thường chạy nhanh hơn nhiều lần và chiếm ít bộ nhớ hơn, vì Neo tối ưu đồ thị tính toán và sinh mã riêng cho chip đó. Một mô hình biên dịch được cho nhiều nền tảng khác nhau mà không phải sửa mã nguồn.
Vì sao các phương án khác sai
- B và C. Dùng Debugger — Debugger làm việc trong lúc huấn luyện, không phải công cụ tối ưu cho triển khai.
- D. Viết lại mã mô hình bằng tay cho từng thiết bị — tốn công khủng khiếp và phải làm lại cho mỗi loại phần cứng; đó chính là thứ Neo thay bạn làm.
A machine learning engineer is tasked with building a predictive model for an e-commerce platform using Amazon SageMaker. The team is considering whether to use a SageMaker Notebook Instance or SageMaker Studio. The engineer needs the following features:
- The ability to run multiple notebooks in parallel within the same environment.
- A centralized space for persistent storage of datasets, notebooks, and logs.
- Seamless integration with existing AWS services like S3 and IAM.
Which SageMaker environment should the engineer choose, and why?
-
A
SageMaker Studio because it is specifically designed for large-scale deep learning workloads, offering GPU-based instances for parallel notebook execution.
-
B
SageMaker Notebook Instance because it allows running multiple notebooks in isolated environments, ensuring that tasks remain independent.
-
C
SageMaker Notebook Instance because it provides access to the full AWS CLI, enabling tighter control over resources and configurations.
-
D
SageMaker Studio because it offers a persistent Jupyter Lab space where multiple notebooks can be managed, allowing centralized data access and integration with AWS services.
Xem giải thích
Đáp án
D — SageMaker Studio, vì có không gian JupyterLab lưu trữ bền và nhiều người dùng chung
Vì sao đúng
Hai đặc điểm quyết định của Studio: lưu trữ bền giữ nguyên công việc qua các phiên và cả khi đổi loại máy tính toán, và mô hình nhiều user profile trong một domain cho cả nhóm làm việc với cấu hình chung. Các công cụ khác — Experiments, Pipelines, Model Registry — mở thẳng trong cùng giao diện.
Vì sao các phương án khác sai
- A. Studio vì dựng riêng cho học sâu quy mô lớn — kết luận đúng nhưng lý do sai; Studio là môi trường phát triển tổng quát, không chuyên cho học sâu.
- B và C. Notebook Instance — mỗi instance là máy đơn lẻ gắn với một người; chạy được nhiều notebook nhưng không có mô hình cộng tác nhiều người dùng.
Which of the following is the primary advantage of using the Hyperband algorithm for hyperparameter tuning over other methods?
-
A
It guarantees that all hyperparameter combinations are tested
-
B
It uses prior training results to suggest the next set of hyperparameters to test
-
C
It allocates more resources to promising configurations while stopping underperforming configurations early
-
D
It randomly selects hyperparameters to reduce training time
Xem giải thích
Đáp án
C — Dồn tài nguyên cho cấu hình có triển vọng và dừng những cấu hình kém
Vì sao đúng
Đây chính là ý tưởng cốt lõi của Hyperband. Nó chạy nhiều cấu hình với ngân sách nhỏ, đánh giá giữa chừng, loại bớt nửa dưới, rồi cấp thêm tài nguyên cho số còn lại và lặp lại. Nhờ vậy không tốn thời gian chạy tới cùng những cấu hình rõ ràng đang tụt lại.
Vì sao các phương án khác sai
- A. Đảm bảo thử hết mọi tổ hợp — đó là grid search.
- B. Dùng kết quả trước để gợi ý lần thử tiếp — đó là tối ưu Bayes; nó thông minh ở khâu chọn cấu hình, còn Hyperband thông minh ở khâu phân bổ tài nguyên.
- D. Chọn ngẫu nhiên để giảm thời gian — đó là random search, và ngẫu nhiên tự nó không giảm thời gian.
A retail company is developing a machine learning model to forecast product demand using Amazon SageMaker. The dataset includes sales history, customer demographics, and seasonal trends. The data is stored in S3 and updated frequently. The company wants to automate the workflow for data preprocessing, feature engineering, and model training. They also need to ensure that the model is continuously retrained as new data arrives, and performance degradation is monitored. Additionally, the company wants to track different versions of the model and hyperparameter configurations during experimentation.
Which solution should the company implement?
-
A
Use SageMaker Feature Store for real-time feature updates, SageMaker Clarify to monitor data bias, and SageMaker Pipelines to retrain the model based on new data.
-
B
Use SageMaker Pipelines to automate data preprocessing, model training, and retraining, SageMaker Model Monitor to detect performance degradation, and SageMaker Experiments to track different model versions and hyperparameter settings.
-
C
Use SageMaker Data Wrangler for data preprocessing, SageMaker Feature Store for feature engineering, and SageMaker Autopilot to automate the model training and hyperparameter tuning process.
-
D
Use SageMaker Ground Truth to label data, SageMaker Model Monitor to automate model retraining, and SageMaker Autopilot for hyperparameter optimization.
Xem giải thích
Đáp án
B — Dùng SageMaker Pipelines tự động hoá tiền xử lý, huấn luyện và huấn luyện lại
Vì sao đúng
Dự báo nhu cầu là bài toán phải chạy đi chạy lại khi có dữ liệu bán hàng mới. Pipelines đóng gói cả chuỗi thành một DAG chạy tự động và có phiên bản, nên huấn luyện lại chỉ là một lần chạy quy trình chứ không phải chuỗi thao tác tay. Đó là điều quan trọng nhất với mô hình sống lâu.
Vì sao các phương án khác sai
- A. Feature Store cập nhật thời gian thực cộng Clarify — cả hai đều hữu ích nhưng không giải bài toán tự động hoá quy trình.
- C. Data Wrangler cộng Feature Store — lo khâu chuẩn bị và lưu đặc trưng; thiếu phần điều phối toàn bộ chuỗi.
- D. Ground Truth để gán nhãn — dữ liệu bán hàng đã có nhãn sẵn là chính doanh số; không cần gán nhãn.
A media company wants to automate the content moderation process for user-submitted images and accompanying text comments. They need to identify inappropriate content such as violent imagery in photos and offensive language in the text. The solution should automatically flag inappropriate submissions and store the flagged data for further review.
Which combination of services should the company use to implement this automated system?
-
A
Amazon Rekognition for content moderation of images, Amazon Comprehend for identifying offensive language in the text, and Amazon S3 for storing flagged data.
-
B
Amazon SageMaker to build custom models for detecting inappropriate images, Amazon Comprehend for text analysis, and Amazon RDS for storing flagged data.
-
C
Amazon Rekognition to detect inappropriate content in images, Amazon Comprehend for sentiment analysis of the text, and Amazon DynamoDB for storing flagged data.
-
D
Amazon Comprehend to detect sentiment and entity extraction from text, Amazon Lex for conversational moderation, and Amazon Elasticsearch for storing flagged data.
Xem giải thích
Đáp án
A — Rekognition kiểm duyệt ảnh, Comprehend phân tích văn bản
Vì sao đúng
Hai loại nội dung nên cần hai dịch vụ. Rekognition có API kiểm duyệt riêng, trả về nhãn phân cấp cho nội dung không phù hợp trong ảnh và video kèm điểm tin cậy. Comprehend đọc phần bình luận: nhận diện thực thể, phân tích cảm xúc, và với bản tuỳ chỉnh thì phân loại được nội dung độc hại theo tiêu chí riêng của công ty.
Vì sao các phương án khác sai
- B. Tự dựng mô hình bằng SageMaker — huấn luyện lại thứ đã có sẵn dưới dạng API.
- C — dùng đúng hai dịch vụ nhưng mô tả sai vai ở vế sau.
- D. Dùng Comprehend cho ảnh — Comprehend làm việc trên văn bản, không đọc được ảnh.
A machine learning engineer is building a model to predict rare diseases based on a dataset where the majority of patients do not have the disease. To handle the imbalance in the dataset, the engineer is considering using SageMaker Data Wrangler to preprocess the data.
Which technique provided by SageMaker Data Wrangler is the BEST option to address class imbalance without increasing the risk of overfitting?
-
A
Simple oversampling to duplicate samples from the minority class.
-
B
Random undersampling to reduce the number of samples in the majority class.
-
C
Synthetic Minority Oversampling Technique (SMOTE) to generate new synthetic samples in the minority class.
-
D
Feature scaling to normalize the feature values in the dataset.
Xem giải thích
Đáp án
C — SMOTE, sinh mẫu tổng hợp cho lớp thiểu số
Vì sao đúng
Với bệnh hiếm, số ca dương tính rất ít nên nhân bản y hệt chúng khiến mô hình học thuộc đúng mấy mẫu đó thay vì học ra quy luật. SMOTE khác ở chỗ nó nội suy giữa các mẫu thật gần nhau trong không gian đặc trưng để tạo ra mẫu mới chưa từng có, nên mô hình thấy được vùng lân cận của lớp hiếm chứ không chỉ vài điểm rời rạc.
Vì sao các phương án khác sai
- A. Nhân bản đơn giản — làm được nhưng dễ gây quá khớp đúng như trên; SMOTE là bản cải tiến của chính nó.
- B. Giảm mẫu lớp đa số — vứt bỏ dữ liệu thật, và với dữ liệu y tế thì mất mát đó rất đắt.
- D. Chuẩn hoá thang đo — cần thiết cho nhiều thuật toán nhưng không liên quan tới mất cân bằng lớp.
A machine learning engineer is working on a binary classification problem for detecting fraudulent transactions. The model's performance is being evaluated using multiple metrics. The model has a precision of 90% and a recall of 60%. The engineer is concerned about missing too many fraudulent transactions.
Which evaluation metric should the engineer focus on to improve the model’s ability to identify more fraud cases?
-
A
F1 Score
-
B
Accuracy
-
C
Recall
-
D
Precision
Xem giải thích
Đáp án
C — Recall
Vì sao đúng
Với phát hiện gian lận, cái giá của hai loại sai rất khác nhau. Bỏ sót một giao dịch gian lận nghĩa là mất tiền thật. Báo nhầm một giao dịch sạch chỉ gây phiền — thêm một bước xác minh. Recall đo đúng phần "bắt được bao nhiêu trong số ca gian lận thật", nên đây là chỉ số cần tối đa hoá.
Vì sao các phương án khác sai
- B. Accuracy — với dữ liệu mất cân bằng nặng thì vô dụng: đoán "không gian lận" cho tất cả đã đạt 98%.
- D. Precision — quan trọng để hạn chế báo động giả, nhưng tối ưu nó thường hy sinh recall, tức là bỏ sót nhiều hơn.
- A. F1 — cân bằng hai cái trên; hợp lý khi hai loại sai có giá ngang nhau, mà ở đây thì không.
A machine learning engineer is using SageMaker Data Wrangler to preprocess data from multiple sources. They want to analyze the relationships between different features in the dataset before selecting which features to include in model training. The engineer also needs to understand which features might have the greatest impact on the target variable. Which two functionalities of Data Wrangler will help the engineer achieve this? (Choose TWO)
-
A
Feature selection through automatic feature scaling.
-
B
Feature importance visualization using the Quick Model.
-
C
Exporting data directly to Amazon SageMaker for model deployment.
-
D
Recursive feature elimination.
-
E
Data visualization tools such as histograms and bar charts.
Xem giải thích
Đáp án
B và E — xem mức quan trọng của đặc trưng bằng Quick Model, và dùng biểu đồ tần suất, biểu đồ cột
Vì sao đúng
Hai công cụ phân tích của Data Wrangler ở hai mức độ sâu khác nhau:
- E. Biểu đồ tần suất và biểu đồ cột cho thấy phân bố của từng đặc trưng — lệch trái hay phải, có bao nhiêu đỉnh, giá trị dị thường nằm đâu. Bước nhìn đầu tiên.
- B. Quick Model đi xa hơn: huấn luyện nhanh một mô hình cây rồi trả về điểm quan trọng của từng đặc trưng, nên biết ngay cái nào đáng giữ.
Vì sao các phương án khác sai
- A. Chọn đặc trưng bằng cách chuẩn hoá tự động — chuẩn hoá là phép biến đổi, không chọn gì.
- C. Xuất dữ liệu sang SageMaker để triển khai — thuộc giai đoạn sau, không phải phân tích.
- D. Recursive Feature Elimination — kỹ thuật có thật nhưng không phải tính năng của Data Wrangler; đây là bẫy lặp lại nhiều lần trong bộ đề.