Ngân hàng đề — AWS Certified Machine Learning Engineer Associate
Tìm thấy 635 câu.
A data scientist is using AWS Lambda to preprocess incoming JSON data from an IoT device stream. The Lambda function needs to validate the data, enrich it by calling an external API, and forward the processed data to an Amazon Kinesis stream for further machine learning processing. The solution should be cost-effective, minimize latency, and ensure the Lambda function does not exceed its timeout limit during the API call.
Which of the following approaches should the data scientist take?
-
A
Use AWS Lambda with asynchronous invocation to call the external API, validate the data, and forward the processed data to Kinesis.
-
B
Use AWS Lambda with the built-in AWS SDK to call the external API synchronously, validate the data, and send the result to Kinesis.
-
C
Use AWS Step Functions to manage the workflow, including the API call and data validation, and then use a separate Lambda function to forward data to Kinesis.
-
D
Use Lambda to invoke an AWS Fargate container for handling the external API call, and forward the enriched data to Kinesis from the Lambda function after validation.
Xem giải thích
Đáp án
C — Dùng AWS Step Functions điều phối cả quy trình
Vì sao đúng
Đề mô tả một chuỗi nhiều bước có gọi API bên ngoài — thứ có thể chậm hoặc hỏng. Nhét cả chuỗi vào một hàm Lambda nghĩa là bạn phải tự viết thử lại, tự xử lý lỗi, và trả tiền cho thời gian hàm nằm chờ API. Step Functions tách từng bước thành một trạng thái, có thử lại và bắt lỗi khai báo sẵn, giữ trạng thái giữa các bước, và không tính tiền lúc chờ.
Vì sao các phương án khác sai
- A. Gọi API bất đồng bộ trong Lambda — Lambda kết thúc trước khi có kết quả trả về, nên bước làm giàu dữ liệu không hoàn thành.
- B. Gọi API đồng bộ trong Lambda — chạy được nhưng hàm phải nằm chờ, tốn tiền và dễ chạm trần thời gian khi API chậm.
- D. Lambda gọi container Fargate — thêm một tầng hạ tầng cho việc Step Functions làm gọn hơn.
A healthcare company needs to ingest and store real-time health monitoring data from wearable devices. The data must be compressed and transformed into a format that supports fast querying in Amazon Athena. They also need to ensure the processed data is securely stored in Amazon S3.
Which of the following solutions would meet the company's requirements using Amazon Kinesis Data Firehose?
-
A
Ingest data using Amazon Kinesis Data Streams, transform the data using AWS Lambda, compress it with Gzip, and store it in Amazon S3 in CSV format.
-
B
Ingest data using Amazon Kinesis Data Firehose, compress the data using Gzip, convert it to Parquet format using the built-in data transformation feature, and store it in Amazon S3 with server-side encryption.
-
C
Ingest data using Amazon Kinesis Data Firehose, use a Lambda function for data transformation, compress the data with Snappy, and store it in Amazon Redshift.
-
D
Ingest data using Amazon Kinesis Data Streams, compress the data with Zlib, transform it to JSON format, and store it in Amazon S3.
Xem giải thích
Đáp án
B — Nhận dữ liệu bằng Kinesis Data Firehose và nén bằng Gzip
Vì sao đúng
Firehose là dịch vụ giao nhận được quản lý hoàn toàn: nó gom lô, nén và ghi xuống S3 mà không cần viết consumer nào. Nén Gzip ngay trong Firehose giảm mạnh dung lượng lưu trữ và chi phí truyền — đúng yêu cầu đề. Với dữ liệu thiết bị đeo đổ về liên tục, đây là con đường ít việc nhất.
Vì sao các phương án khác sai
- A và D. Dùng Data Streams — Data Streams là ống chứa; muốn xuống S3 kèm nén thì vẫn phải tự viết và tự vận hành consumer.
- C. Firehose kèm Lambda để biến đổi — Lambda dùng khi cần đổi định dạng hoặc làm giàu dữ liệu; ở đây chỉ cần nén, mà Firehose nén sẵn nên thêm Lambda là thừa.
A Machine Learning Engineer is preparing to train a deep learning model on Amazon SageMaker using a custom algorithm. The dataset is very large and stored across multiple S3 buckets. Due to the data size, it needs to be streamed directly into SageMaker for training rather than loaded all at once.
What is the recommended approach to efficiently manage and stream this data during training in Amazon SageMaker?
-
A
Use SageMaker Pipe mode to stream the data directly from S3 to the training instances.
-
B
Use AWS Glue to join the data into a single file before loading it into SageMaker.
-
C
Manually download and preprocess the data on an Amazon EC2 instance before uploading it to SageMaker.
-
D
Use the SageMaker batch transform feature to prepare the data before training.
Xem giải thích
Đáp án
A — Dùng Pipe mode để truyền dữ liệu thẳng từ S3 vào máy huấn luyện
Vì sao đúng
Ở chế độ File mode mặc định, SageMaker tải toàn bộ dữ liệu xuống ổ đĩa trước khi bắt đầu — với tập dữ liệu rất lớn thì đó là hàng chục phút chờ và một ổ đĩa đủ to phải trả tiền. Pipe mode truyền dữ liệu thành luồng thẳng vào tiến trình huấn luyện, nên bắt đầu học gần như ngay lập tức và không bị giới hạn bởi dung lượng ổ.
Vì sao các phương án khác sai
- B. Gộp dữ liệu thành một tệp bằng Glue — thêm một bước xử lý tốn kém, và tệp khổng lồ còn khó đọc song song hơn.
- C. Tải về EC2 xử lý trước — thêm hạ tầng và thêm một bản sao dữ liệu.
- D. Dùng batch transform để chuẩn bị dữ liệu — batch transform là để chạy suy luận trên tập dữ liệu, không phải công cụ chuẩn bị dữ liệu huấn luyện.
A data scientist is using Amazon SageMaker to develop a machine learning model for customer behavior prediction. The scientist is working in a SageMaker Notebook Instance and wants to collaborate with a team of data analysts on this project. Additionally, they need to ensure that their work, including data visualizations and intermediate results, is saved and can be accessed even after the notebook instance is stopped.
Which configuration should the data scientist consider to achieve these goals? (Select TWO)
-
A
Use SageMaker Studio to enable collaboration, as it allows multiple users to access shared resources within a Jupyter Lab environment.
-
B
Enable persistent storage in the SageMaker Notebook Instance to automatically save data and code, ensuring that all work is retained even after stopping the instance.
-
C
Install all necessary libraries manually in the SageMaker Notebook Instance to ensure the team has access to the correct dependencies for their machine learning project.
-
D
Use SageMaker notebook instances to share code directly, which allows for real-time collaboration between multiple users.
-
E
Configure a custom EC2 instance with an EBS volume for persistent storage, ensuring data is saved when the instance is stopped.
Xem giải thích
Đáp án
A và B — dùng Studio để cộng tác, và bật lưu trữ bền
Vì sao đúng
Hai vấn đề của đề được giải bằng hai thứ khác nhau:
- A. Studio giải phần cộng tác: nhiều người dùng trong một domain, chia sẻ notebook kèm ngữ cảnh, quản trị viên cấp quyền tập trung.
- B. Lưu trữ bền giải phần không mất việc: nội dung giữ nguyên khi dừng và khởi động lại máy, thói quen phổ biến vì máy huấn luyện đắt.
Vì sao các phương án khác sai
- C. Cài thư viện bằng tay ở mỗi máy — không mở rộng được và mỗi người một môi trường khác nhau.
- D. Chia sẻ mã trực tiếp giữa các notebook instance — instance là máy riêng lẻ, không có cơ chế chia sẻ như vậy.
- E. Tự dựng EC2 kèm EBS — được lưu trữ bền nhưng mất hết phần được quản lý và phần cộng tác.
Machine learning team is tasked with preparing a dataset for a model that will predict customer churn. They decide to use SageMaker Data Wrangler to streamline the data preparation. The dataset contains time-based information, and they want to engineer new features such as "day of the week" from timestamps and assess the importance of these features.
What TWO features of SageMaker Data Wrangler will assist the team in this process? (Choose TWO)
-
A
Time series analysis and forecasting.
-
B
Feature engineering options to create new features like "day of the week" from timestamps.
-
C
The Quick Model feature to automatically compute feature importance.
-
D
Integrating data directly from Amazon RDS for real-time predictions.
-
E
Data export to Amazon Athena for advanced SQL-based transformations.
Xem giải thích
Đáp án
B và C — các phép tạo đặc trưng mới, và Quick Model để tính mức quan trọng
Vì sao đúng
Hai thứ này là phần Data Wrangler đóng góp nhiều nhất cho bài toán dự báo rời bỏ khách hàng:
- B. Tạo đặc trưng mới — ví dụ tách "thứ trong tuần" từ cột ngày. Đặc trưng dạng này thường mang nhiều tín hiệu hơn hẳn cột gốc, và làm bằng chuột chứ không phải viết mã.
- C. Quick Model — huấn luyện nhanh một mô hình cây rồi trả về điểm quan trọng của từng đặc trưng, nên biết ngay cái nào đáng giữ trước khi dựng quy trình đầy đủ.
Vì sao các phương án khác sai
- A. Phân tích và dự báo chuỗi thời gian — không phải trọng tâm của Data Wrangler.
- D. Nối thẳng RDS cho dự đoán thời gian thực — Data Wrangler làm việc trên dữ liệu tĩnh, không phục vụ suy luận.
- E. Xuất sang Athena để biến đổi bằng SQL — Athena là nguồn đọc vào, không phải đích xuất ra.
You are tasked with training a machine learning model for image classification using Amazon SageMaker's built-in algorithms.
Which of the following steps accurately describe the process of using SageMaker's built-in algorithms?
-
A
Set up a custom container, upload the algorithm, and configure infrastructure.
-
B
Select the algorithm from the SageMaker console, upload your dataset, configure hyperparameters, and launch the training job.
-
C
Create a custom script for model training, manually configure hyperparameter tuning, and deploy the model on an EC2 instance.
-
D
Use SageMaker built-in image generation algorithms and skip hyperparameter configuration to deploy a pre-trained model.
Xem giải thích
Đáp án
B — Chọn thuật toán, tải dữ liệu lên, khai siêu tham số
Vì sao đúng
Đó là toàn bộ quy trình khi dùng thuật toán dựng sẵn: mã mô hình và container đã có, việc của bạn gọn lại thành ba thứ — dữ liệu đã gán nhãn ở định dạng thuật toán yêu cầu, bộ siêu tham số, và loại máy để chạy. Với phân loại ảnh còn dùng được học chuyển giao từ mô hình đã huấn luyện trên tập lớn, nên cần ít dữ liệu hơn nhiều.
Vì sao các phương án khác sai
- A. Dựng container riêng và tự cấu hình hạ tầng — chính là thứ thuật toán dựng sẵn giúp tránh.
- C. Viết script huấn luyện riêng — đó là Script Mode, dùng khi cần thuật toán tự viết.
- D. Bỏ qua bước khai siêu tham số — siêu tham số vẫn phải khai; và "thuật toán sinh ảnh" là loại bài toán khác hẳn phân loại ảnh.
A machine learning engineer is training a decision tree model and is tasked with tuning the model’s hyperparameters. Which of the following is an example of a hyperparameter that the engineer can configure before training begins?
-
A
The weights assigned to each node in the decision tree
-
B
The depth of the tree
-
C
The data points used for training
-
D
The split criteria for each node learned from the data
Xem giải thích
Đáp án
B — Độ sâu của cây
Vì sao đúng
Độ sâu tối đa là giá trị bạn đặt trước khi huấn luyện để giới hạn mức độ phức tạp của cây, nên nó đúng là siêu tham số. Nó cũng là nút chỉnh quan trọng nhất với cây quyết định: cây quá sâu học thuộc nhiễu (quá khớp), cây quá nông không nắm được quy luật (thiếu khớp).
Vì sao các phương án khác sai
- A. Trọng số ở mỗi nút và D. Tiêu chí chia tại mỗi nút học từ dữ liệu — đều là thứ thuật toán tự tìm ra trong lúc huấn luyện, tức là tham số mô hình.
- C. Các điểm dữ liệu dùng để huấn luyện — đó là dữ liệu đầu vào, không phải tham số nào cả.
A machine learning engineer is working with Amazon SageMaker to train a decision tree model and needs to optimize the hyperparameters. The model's performance needs to be optimized by tuning the depth of the tree and the minimum samples required for a split.
Which of the following describes the difference between hyperparameters and model parameters?
-
A
Hyperparameters are learned from the training data, while model parameters are set before training.
-
B
Hyperparameters are external configurations set before training begins, while model parameters are learned during training.
-
C
Hyperparameters control the size of the data input, while model parameters control the training speed.
-
D
Model parameters influence how quickly the model learns, while hyperparameters control the training algorithm.
Xem giải thích
Đáp án
B — Siêu tham số là cấu hình bên ngoài đặt trước khi huấn luyện, tham số mô hình học từ dữ liệu
Vì sao đúng
Ranh giới này quyết định cách bạn làm việc. Siêu tham số điều khiển cách học — tốc độ học, độ sâu cây, hệ số chính quy hoá — và bạn chọn chúng trước. Tham số mô hình là kết quả của quá trình học — trọng số, điểm cắt tại mỗi nút.
Hệ quả thực tế: "tinh chỉnh siêu tham số" nghĩa là thử nhiều bộ giá trị của nhóm thứ nhất, và mỗi lần thử lại chạy huấn luyện để tìm ra nhóm thứ hai.
Vì sao các phương án khác sai
- A — đảo ngược hoàn toàn hai khái niệm.
- C. Siêu tham số điều khiển kích thước dữ liệu đầu vào — mô tả sai vai trò của cả hai.
- D — cũng gán sai vai: chính siêu tham số (như tốc độ học) mới quyết định mô hình học nhanh hay chậm.
A machine learning team needs to train a model on a large dataset that can fit into memory, but the training would be too slow on a single machine. Which distributed training method should they use in Amazon SageMaker to speed up the process?
-
A
Model parallelism, where the model is split across multiple machines.
-
B
Use data parallelism to split the dataset across several machines, with each machine training a copy of the model.
-
C
Manually manage data distribution and gradient synchronization across multiple machines.
-
D
Train the model on a single large instance with more compute power.
Xem giải thích
Đáp án
B — Song song theo dữ liệu: chia tập dữ liệu cho nhiều máy
Vì sao đúng
Đề nói rõ dữ liệu vừa bộ nhớ, vấn đề chỉ là chậm. Đó đúng là tình huống của song song theo dữ liệu: mỗi máy giữ một bản sao đầy đủ của mô hình và xử lý một phần dữ liệu, gradient được gộp lại rồi mọi bản sao cùng cập nhật. Thời gian huấn luyện giảm gần tuyến tính theo số máy.
Vì sao các phương án khác sai
- A. Song song theo mô hình — dùng khi mô hình không vừa một GPU; ở đây không phải vấn đề đó.
- C. Tự quản lý phân phối dữ liệu và đồng bộ gradient — SageMaker có thư viện làm sẵn; tự viết là chuốc lấy một lớp lỗi rất khó gỡ.
- D. Dùng một máy mạnh hơn — có giới hạn, và không mở rộng được khi dữ liệu tiếp tục lớn.
A media company is developing a machine learning model to categorize a large library of video content. Since manually labeling thousands of videos would be costly, the team wants to reduce labeling costs by automating part of the labeling process while ensuring that complex cases are reviewed by human annotators. They also want the labeling process to improve over time based on the labeled data.
What is the BEST strategy for optimizing labeling costs while ensuring high-quality results?
-
A
Use SageMaker Ground Truth’s active learning to automate simple labeling tasks and send more complex labeling tasks to human annotators.
-
B
Rely solely on human workers to manually label all video content, ensuring maximum accuracy but at a higher cost.
-
C
Use SageMaker Ground Truth to automate all labeling tasks without human involvement to reduce costs.
-
D
Use SageMaker Experiments to track labeling performance and automatically adjust labeling strategies based on accuracy metrics.
Xem giải thích
Đáp án
A — Ground Truth với học chủ động, tự động gán nhãn những ca đơn giản
Vì sao đúng
Học chủ động cân bằng đúng hai mục tiêu đề nêu. Một mô hình được huấn luyện dần trên chính nhãn đang thu, tự gán nhãn những video nó đã tự tin, và chỉ chuyển cho người những ca khó. Tỷ lệ cần người giảm dần theo thời gian, nên chi phí giảm mà chất lượng ở những ca quyết định vẫn do người kiểm soát.
Vì sao các phương án khác sai
- B. Người làm hết — chất lượng cao nhất nhưng đắt và chậm; đề nói rõ muốn giảm chi phí.
- C. Máy làm hết, không có người — chất lượng tụt ở đúng những ca khó, và không có ai phát hiện ra khi mô hình gán nhãn sai một cách hệ thống.
- D. Experiments để theo dõi hiệu quả gán nhãn — Experiments theo dõi các lần huấn luyện, không phải công cụ gán nhãn.