Ngân hàng đề — AWS Certified Machine Learning Engineer Associate
Tìm thấy 635 câu.
A healthcare organization needs to analyze large volumes of patient medical forms and identify important entities such as patient names, diagnoses, and medications. These forms include both text and handwritten content. The organization also needs to extract and categorize information to automate the documentation process and store the extracted metadata for further querying. Which combination of services would BEST meet the organization’s requirements?
-
A
Amazon SageMaker for handwriting recognition, Amazon Comprehend for sentiment analysis, and Amazon Kendra for querying the extracted metadata.
-
B
Amazon Rekognition for detecting text and images, Amazon Comprehend for analyzing the extracted text, and Amazon RDS for storing structured data.
-
C
Amazon Textract for extracting both printed and handwritten text, Amazon Comprehend for entity recognition, and Amazon S3 for storing the extracted metadata.
-
D
Amazon Comprehend for entity extraction, Amazon Textract for analyzing images and text, and Amazon Elasticsearch for indexing and querying the results.
Xem giải thích
Đáp án
C — Amazon Textract để trích văn bản (kể cả chữ viết tay), Amazon Comprehend để nhận thực thể
Vì sao đúng
Đề có hai bước rõ ràng. Textract đọc được cả chữ in lẫn chữ viết tay trên biểu mẫu, và giữ được cấu trúc bảng và cặp khoá–giá trị. Comprehend (bản Medical hợp nhất với dữ liệu y tế) nhận ra tên bệnh nhân, chẩn đoán và thuốc từ văn bản đã trích.
Vì sao các phương án khác sai
- A. SageMaker để đọc chữ viết tay — tự huấn luyện lại thứ Textract đã làm sẵn.
- B. Rekognition để phát hiện văn bản — Rekognition đọc được chữ ngắn trên ảnh (biển hiệu, biển số), không phải công cụ cho biểu mẫu nhiều trang.
- D. Đảo thứ tự hai dịch vụ — Comprehend làm việc trên văn bản, nên phải có Textract trước.
A company is using Amazon Kinesis to monitor IoT sensor data for real-time security threats. The system needs to detect anomalies in the data stream and trigger alerts immediately, minimizing latency. They plan to have multiple consumers process the data streams concurrently.
What approach should the company take to ensure optimal scalability and low latency?
-
A
Use Kinesis Firehose to scale automatically and reduce latency for each consumer.
-
B
Enable enhanced fan-out for each consumer to achieve dedicated read throughput and lower latency.
-
C
Use AWS Glue for ETL and real-time data analysis of the streaming data.
-
D
Use standard Kinesis consumers and scale the number of shards as needed.
Xem giải thích
Đáp án
B — Bật enhanced fan-out cho từng consumer
Vì sao đúng
Không bật thì mọi consumer chia nhau 2 MB/giây của mỗi shard và phải hỏi liên tục, nên càng thêm ứng dụng đọc thì độ trễ càng tăng — trái yêu cầu "cảnh báo ngay lập tức". Enhanced fan-out cấp cho mỗi consumer một đường riêng 2 MB/giây trên mỗi shard và đẩy dữ liệu sang, kéo độ trễ xuống khoảng 70 mili giây.
Vì sao các phương án khác sai
- A. Firehose để giảm độ trễ — Firehose gom lô, tối thiểu cũng trễ vài chục giây.
- C. Glue cho phân tích thời gian thực — Glue chạy theo mẻ.
- D. Consumer thường, tăng số shard — tăng tổng thông lượng nhưng các consumer vẫn giành nhau phần của từng shard.
A machine learning engineer at a healthcare company has deployed a model on Amazon SageMaker and wants to monitor the data quality to ensure that the data used in production maintains high integrity.
Which of the following actions is MOST appropriate for monitoring data quality using Amazon SageMaker?
-
A
Use a custom SageMaker algorithm to filter out low-quality data before making predictions in the production environment.
-
B
Enable SageMaker Feature Store to log and validate all incoming data used for predictions and correct data quality issues in real-time.
-
C
Use SageMaker Model Monitor to evaluate incoming data for missing values, outliers, and deviations from the baseline training data.
-
D
Set up Amazon CloudWatch to automatically retrain the model whenever the incoming data has quality issues.
Xem giải thích
Đáp án
C — Dùng SageMaker Model Monitor để kiểm dữ liệu vào
Vì sao đúng
Model Monitor sinh ra cho đúng việc này: nó lấy thống kê nền từ dữ liệu huấn luyện (khoảng giá trị, tỷ lệ thiếu, phân bố), rồi so dữ liệu thật đi vào endpoint với nền đó theo lịch. Thiếu giá trị, giá trị lạ, hay phân bố trôi đi đều bị phát hiện và báo qua CloudWatch — không phải viết dòng kiểm tra nào.
Vì sao các phương án khác sai
- A. Tự viết thuật toán lọc dữ liệu xấu — che triệu chứng và tự dựng lại thứ có sẵn.
- B. Feature Store để ghi và kiểm dữ liệu — kho đặc trưng, không có phần đánh giá chất lượng so với nền.
- D. CloudWatch tự huấn luyện lại — CloudWatch không huấn luyện gì; và huấn luyện lại vô điều kiện là phản ứng sai khi chưa biết vì sao dữ liệu đổi.
A retail company uses Amazon SageMaker to forecast sales by training a machine learning model on historical data. The data includes columns such as product_id, date, price, promotion, and units_sold. The company wants to improve its forecast accuracy by incorporating new features, such as holidays and competitor pricing.
What is the most effective way to include these new features in the model training workflow?
-
A
Add the new features to the dataset, retrain the model, and compare the new model’s performance.
-
B
Enable Amazon SageMaker’s AutoPilot feature to automatically detect and add holiday and competitor pricing features.
-
C
Use Amazon Forecast, which automatically incorporates external variables, such as holidays, into the forecasting model.
-
D
Use Amazon SageMaker’s hyperparameter tuning to adjust the model based on new features.
Xem giải thích
Đáp án
A — Thêm đặc trưng mới vào dữ liệu, huấn luyện lại, rồi so với mô hình cũ
Vì sao đúng
Đây là quy trình đúng và cũng là quy trình duy nhất trả lời được câu hỏi "đặc trưng mới có giúp không". Thêm cột, huấn luyện lại, rồi so chỉ số trên cùng một tập kiểm tra. Không có bước so sánh thì bạn chỉ có một mô hình khác chứ không biết nó tốt hơn hay tệ hơn.
Vì sao các phương án khác sai
- B. AutoPilot tự phát hiện ngày lễ — AutoPilot tự chọn thuật toán và tiền xử lý, nhưng không tự sinh ra đặc trưng ngoại cảnh mà bạn chưa cung cấp.
- C. Chuyển sang Amazon Forecast — Forecast có hỗ trợ chuỗi thời gian liên quan, nhưng đổi hẳn dịch vụ là thay đổi lớn hơn nhiều so với việc đề đang hỏi.
- D. Tinh chỉnh siêu tham số — tối ưu mô hình trên đúng bộ đặc trưng cũ; không đánh giá được đặc trưng mới.
Suppose you have a relatively small model that trains very quickly, but you want to systematically guarantee coverage of the hyperparameter space. Which method is most suitable in this case?
-
A
Grid Search
-
B
Bayesian Optimization
-
C
Random Search
-
D
Hyperband
Xem giải thích
Đáp án
A — Grid Search
Vì sao đúng
Điều kiện của đề rất cụ thể: mô hình nhỏ, huấn luyện rất nhanh, và cần đảm bảo phủ hết không gian siêu tham số một cách có hệ thống. Grid search thử mọi tổ hợp trên lưới đã khai, nên đó là phương pháp duy nhất cho lời đảm bảo về độ phủ. Nhược điểm của nó — số tổ hợp bùng nổ theo số chiều — không thành vấn đề khi mỗi lần huấn luyện rất rẻ.
Vì sao các phương án khác sai
- B. Bayesian Optimization — thông minh hơn và ít lần thử hơn, nhưng nó đi theo hướng có triển vọng nên cố ý bỏ qua nhiều vùng; không đảm bảo phủ.
- C. Random Search — hiệu quả khi không gian lớn, nhưng ngẫu nhiên thì không có gì đảm bảo.
- D. Hyperband — cắt sớm những cấu hình kém để tiết kiệm thời gian; lợi thế đó vô nghĩa khi huấn luyện vốn đã nhanh.
You are a data scientist tasked with developing a model for a large e-commerce platform. The company needs a solution that minimizes development time and operational complexity. The team is considering using Amazon SageMaker's built-in algorithms instead of building custom models.
Which of the following advantages do built-in SageMaker algorithms offer for this use case? (Choose TWO)
-
A
Custom algorithms provide faster training times and more efficient resource usage compared to built-in algorithms.
-
B
Custom frameworks such as PyTorch and TensorFlow must be used with SageMaker to enable hyperparameter tuning.
-
C
Built-in algorithms can only be used for simple tasks like linear regression and are not suitable for complex machine learning tasks.
-
D
Pre-configured environments and infrastructure management are handled by AWS, eliminating the need for custom containers.
-
E
Built-in algorithms like XGBoost and Linear Learner abstract away infrastructure setup and allow quick model deployment.
Xem giải thích
Đáp án
D và E — AWS lo sẵn môi trường và hạ tầng, và thuật toán dựng sẵn giấu đi phần hạ tầng
Vì sao đúng
Yêu cầu của đề là giảm thời gian phát triển và độ phức tạp vận hành, và hai phương án này nói đúng cách SageMaker đạt được điều đó: môi trường và container đã dựng sẵn nên không phải cài đặt gì, còn thuật toán dựng sẵn như XGBoost hay Linear Learner dùng được ngay bằng vài dòng cấu hình, không cần viết mã mô hình.
Vì sao các phương án khác sai
- A. Thuật toán tự viết huấn luyện nhanh hơn — sai; thuật toán dựng sẵn của AWS đã được tối ưu kỹ, thường nhanh hơn bản tự viết.
- B. Bắt buộc phải dùng PyTorch hoặc TensorFlow — sai; đó là lựa chọn, không phải bắt buộc.
- C. Thuật toán dựng sẵn chỉ làm được việc đơn giản — sai; chúng phủ cả thị giác máy tính, xử lý ngôn ngữ và dự báo chuỗi thời gian.
A company is building a multi-lingual customer support system that allows users to communicate both via text and voice. The system should transcribe customer support calls into text, translate the text into the user’s preferred language, and generate audio responses in the same language. Which combination of AWS services is MOST appropriate for this use case?
-
A
Use Amazon Polly to convert speech to text, AWS Glue to manage translations, and Amazon Transcribe to generate speech responses.
-
B
Use AWS Lambda to convert speech to text, Amazon Translate to handle text translations, and Amazon SageMaker to generate speech responses.
-
C
Use Amazon Transcribe to convert speech to text, Amazon Translate to translate the text, and Amazon Polly to convert the translated text to speech.
-
D
Use Amazon Transcribe to convert speech to text, Amazon Lex to handle translations, and Amazon Polly to generate the voice responses.
Xem giải thích
Đáp án
C — Amazon Transcribe chuyển giọng nói thành văn bản, Amazon Translate để dịch
Vì sao đúng
Ba dịch vụ đúng vai cho ba việc: Transcribe làm nhận dạng giọng nói, hỗ trợ nhiều ngôn ngữ và tách được người nói. Translate dịch máy dựa trên mạng nơ-ron giữa hàng chục cặp ngôn ngữ. Nếu cần trả lời bằng giọng nói thì thêm Polly ở cuối.
Vì sao các phương án khác sai
- A. Polly để chuyển giọng nói thành văn bản — ngược chiều: Polly biến văn bản thành giọng nói.
- B. Lambda để nhận dạng giọng nói — Lambda chỉ chạy mã, tự nó không có năng lực nhận dạng nào.
- D. Lex để dịch — Lex dựng chatbot theo ý định người dùng, không phải công cụ dịch.
A financial services company is using Amazon SageMaker Model Monitor to track the feature attribution drift of their credit risk prediction model. They notice that the importance of a particular feature, "annual income," has shifted significantly over time, making it a more dominant factor in the model's predictions.
What is the BEST course of action to address this shift in feature importance?
-
A
Switch to a different model algorithm that automatically reduces the importance of "annual income" in predictions.
-
B
Fine-tune the hyperparameters of the existing model to lower the weight of the "annual income" feature.
-
C
Retrain the model with a larger dataset that excludes the "annual income" feature to prevent bias in predictions.
-
D
Use SageMaker Model Monitor to alert the team about this shift and retrain the model with updated data to ensure feature importance remains balanced.
Xem giải thích
Đáp án
D — Để Model Monitor cảnh báo về thay đổi đó, rồi huấn luyện lại mô hình
Vì sao đúng
Trôi đóng góp đặc trưng nghĩa là quan hệ giữa dữ liệu và kết quả đã đổi trong thực tế — ví dụ thu nhập hằng năm không còn dự báo rủi ro tín dụng tốt như trước. Phản ứng đúng là để hệ thống báo, xem lại dữ liệu mới, rồi huấn luyện lại trên dữ liệu phản ánh thực tế hiện tại. Mô hình phải bám theo thế giới, không phải ngược lại.
Vì sao các phương án khác sai
- A. Đổi thuật toán để tự giảm trọng số đặc trưng đó — không thuật toán nào "tự biết" nên giảm cái gì; đây là hiểu sai về cách mô hình học.
- B. Tinh chỉnh siêu tham số để hạ trọng số — siêu tham số điều khiển cách học, không chỉ định trọng số của một đặc trưng cụ thể.
- C. Bỏ hẳn đặc trưng đó ra — vội vàng: đóng góp thay đổi không có nghĩa đặc trưng vô dụng; bỏ đi có thể làm mô hình tệ hơn.
A transportation company is analyzing GPS data from its fleet of vehicles to optimize delivery routes. The company needs to preprocess and clean this data, create relevant features like average speed and distance traveled, and store these features in a scalable feature store. The team wants to ensure the features are available for real-time predictions on delivery times and wants to train a machine learning model using these features in an easily managed notebook environment.
Which services should the company use to preprocess the data, store the features, and train the model?
-
A
Use SageMaker Data Wrangler for preprocessing, Amazon Redshift for feature storage, and SageMaker Notebooks to train and deploy the machine learning model.
-
B
Use SageMaker Notebooks for data preprocessing, Amazon DynamoDB for feature storage, and SageMaker Feature Store to train the machine learning model.
-
C
Use SageMaker Data Wrangler for data preprocessing, SageMaker Feature Store to store features, and SageMaker Notebooks to train and deploy the machine learning model.
-
D
Use Amazon QuickSight to preprocess and clean the data, SageMaker Feature Store for feature storage, and SageMaker Autopilot to train the machine learning model.
Xem giải thích
Đáp án
C — Data Wrangler để tiền xử lý, Feature Store để lưu đặc trưng
Vì sao đúng
Đúng hai bước đề nêu, mỗi bước một công cụ chuyên trách. Data Wrangler làm sạch dữ liệu GPS bằng giao diện trực quan và xuất ra pipeline chạy lại được. Feature Store lưu đặc trưng đã tính kèm dấu thời gian, dùng chung cho cả huấn luyện lẫn suy luận — nhờ vậy tránh được chuyện đặc trưng lúc huấn luyện khác lúc chạy thật, lỗi kinh điển của hệ thống ML.
Vì sao các phương án khác sai
- A. Redshift để lưu đặc trưng — kho dữ liệu phân tích, không có khái niệm nhóm đặc trưng hay truy xuất độ trễ thấp lúc suy luận.
- B. DynamoDB để lưu đặc trưng — nhanh nhưng phải tự dựng toàn bộ phần quản lý phiên bản và dấu thời gian mà Feature Store làm sẵn.
- D. QuickSight để làm sạch dữ liệu — công cụ trực quan hoá cho người xem báo cáo, không phải công cụ tiền xử lý.
An ML engineer wants to automate the hyperparameter tuning of their machine learning model using Amazon SageMaker. To minimize costs, they also want to utilize spot instances during the tuning process. Which combination of factors must the engineer specify to initiate the automatic tuning process in SageMaker?
-
A
Dataset, algorithm, performance metric
-
B
Algorithm, hyperparameter range, performance metric
-
C
Hyperparameter range, dataset, algorithm
-
D
Algorithm, performance metric, model parameters
Xem giải thích
Đáp án
C — Khoảng siêu tham số, tập dữ liệu, và thuật toán
Vì sao đúng
Một job tinh chỉnh siêu tham số của SageMaker cần tối thiểu ba thứ: thuật toán (hoặc ảnh container) để biết huấn luyện cái gì, tập dữ liệu để huấn luyện, và khoảng giá trị của từng siêu tham số để biết được phép thử trong phạm vi nào. Có ba thứ này là job chạy được; việc dùng spot instance chỉ là một cờ cấu hình thêm.
Vì sao các phương án khác sai
- A và B — thiếu một trong ba thành phần bắt buộc.
- D. Tham số của mô hình — tham số mô hình là thứ học ra được trong lúc huấn luyện (trọng số), khác hẳn siêu tham số do bạn đặt trước. Đây là chỗ hay nhầm nhất.