Ngân hàng đề — AWS Certified Machine Learning Engineer Associate
Tìm thấy 635 câu.
Which of the following is an example of unsupervised learning in machine learning?
-
A
Predicting the price of a stock based on historical data
-
B
Predicting whether an email is spam or not
-
C
Forecasting sales for the next quarter
-
D
Grouping customers based on purchasing behavior without labeled categories
Xem giải thích
Đáp án
D — Gom nhóm khách hàng theo hành vi mua sắm mà không có nhãn sẵn
Vì sao đúng
Dấu hiệu nhận biết học không giám sát là không có biến mục tiêu. Ở đây không ai nói trước "khách này thuộc nhóm A", thuật toán tự tìm ra cấu trúc và gom những người có hành vi tương tự. Kết quả là các cụm, và việc đặt tên cho từng cụm là do con người diễn giải sau.
Vì sao các phương án khác sai
Cả ba phương án còn lại đều có đáp án đúng để học theo, nên đều là học có giám sát:
- A. Dự đoán giá cổ phiếu — hồi quy, mục tiêu là giá.
- B. Phân loại email có phải thư rác không — phân loại nhị phân, mục tiêu là nhãn spam.
- C. Dự báo doanh số quý tới — hồi quy trên chuỗi thời gian.
A manufacturing company is using machine learning to predict equipment failures. The company has data on equipment usage, temperature, and maintenance records stored in an Amazon S3 bucket. They need to clean and analyze this data, engineer features such as average machine temperature over time, and store these features for both batch and real-time predictions. The company also wants to utilize a pre-built machine learning model to accelerate the deployment process.
Which combination of services would BEST meet the company's needs?
-
A
Use SageMaker Data Wrangler to clean and preprocess the data, store the features in Amazon Redshift, and use SageMaker JumpStart to deploy a pre-built model.
-
B
Use SageMaker Data Wrangler to clean and transform the data, store the features in SageMaker Feature Store, and use SageMaker JumpStart to deploy a pre-built predictive maintenance model.
-
C
Use Amazon EMR to preprocess the data, store the features in Amazon DynamoDB, and use SageMaker to train and deploy the model.
-
D
Use SageMaker Notebooks to preprocess the data, store the features in Amazon S3, and use SageMaker Autopilot to train and deploy a predictive model.
Xem giải thích
Đáp án
B — Data Wrangler làm sạch và biến đổi, Feature Store lưu đặc trưng
Vì sao đúng
Hai bước đúng vai. Data Wrangler nối tới nguồn dữ liệu, hiện ngay hồ sơ chất lượng, và cho làm sạch bằng giao diện rồi xuất thành pipeline chạy lại được. Feature Store lưu đặc trưng đã tính kèm dấu thời gian, dùng chung cho huấn luyện và cho dự đoán — nhờ vậy mô hình dự báo hỏng hóc thiết bị luôn dùng cùng một cách tính đặc trưng ở cả hai giai đoạn.
Vì sao các phương án khác sai
- A — đúng phần Data Wrangler nhưng sai nơi lưu đặc trưng.
- C. EMR và DynamoDB — làm được nhưng phải dựng cụm và tự viết toàn bộ phần quản lý phiên bản đặc trưng.
- D. Notebook và S3 — S3 lưu được tệp nhưng không có khái niệm nhóm đặc trưng, dấu thời gian hay truy xuất độ trễ thấp.
An organization needs to train a machine learning model with a custom algorithm that is not natively supported by SageMaker's built-in algorithms. The organization prefers to use SageMaker’s pre-built containers for framework support, but with custom training and inference code.
Which approach should the organization take?
-
A
Use SageMaker's built-in algorithms and modify them to include custom training code.
-
B
Write custom training code directly in the SageMaker notebook instance and run it using the notebook’s execution kernel.
-
C
Utilize SageMaker’s Script Mode, write the custom training logic in a Python script, and specify the entry point in a SageMaker estimator.
-
D
Build a custom Docker container from scratch, package all dependencies, and upload it to Amazon ECR for use in SageMaker.
Xem giải thích
Đáp án
C — Dùng Script Mode, viết logic huấn luyện trong một script Python
Vì sao đúng
Đề nêu rõ hai điều kiện: thuật toán không có sẵn và không muốn tự dựng container. Script Mode nằm đúng giữa hai thái cực đó: bạn viết mã huấn luyện bằng khung học máy quen thuộc, còn SageMaker chạy nó trong container dựng sẵn của khung đó. Thư viện thêm thì khai trong requirements.txt.
Vì sao các phương án khác sai
- A. Sửa thuật toán dựng sẵn — không sửa được; chúng là container đóng.
- B. Chạy thẳng trong notebook instance — chạy được lúc thử nghiệm, nhưng không dùng được tài nguyên huấn luyện phân tán và không đóng gói lại được thành job.
- D. Tự dựng container từ đầu — chính là điều đề nói muốn tránh.
A pharmaceutical company is using machine learning to analyze clinical trial data. They need to ingest raw data from multiple sources, such as CSV files and relational databases, and clean and join this data. Afterward, the data should be used to create a set of features, including patient age, treatment history, and genetic markers. The company also wants to maintain a version-controlled feature repository that can be reused across different machine learning models.
Which AWS services should the company use to preprocess the data, create features, and maintain a version-controlled feature repository?
-
A
Use Amazon QuickSight for data preprocessing, SageMaker Feature Store to store features, and SageMaker JumpStart to deploy a pre-built model.
-
B
Use AWS Glue for data cleaning, store features in Amazon DynamoDB, and use SageMaker to train models.
-
C
Use Amazon Athena to preprocess the data, store the features in Amazon S3, and use SageMaker Autopilot to automatically train models.
-
D
Use SageMaker Data Wrangler for data preprocessing, SageMaker Feature Store to store and manage the features, and SageMaker Notebooks to train the models.
Xem giải thích
Đáp án
D — Data Wrangler để tiền xử lý, Feature Store để lưu đặc trưng
Vì sao đúng
Đề nêu nhiều nguồn dữ liệu — tệp CSV và CSDL quan hệ — nên cần một công cụ nối được tới nhiều nguồn và gộp lại. Data Wrangler làm đúng việc đó: nối tới S3, Athena, Redshift, làm sạch bằng giao diện, rồi xuất thành pipeline. Feature Store sau đó giữ đặc trưng đã tính cho cả huấn luyện lẫn suy luận, kèm quản lý phiên bản — điều bắt buộc với dữ liệu thử nghiệm lâm sàng.
Vì sao các phương án khác sai
- A. QuickSight để tiền xử lý — công cụ trực quan hoá cho người xem báo cáo.
- B. Lưu đặc trưng trong DynamoDB — nhanh nhưng phải tự dựng phần quản lý phiên bản và dấu thời gian.
- C. Athena để tiền xử lý — SQL làm được nhiều việc, nhưng các phép biến đổi cho ML như chuẩn hoá hay mã hoá thì rất vụng.
A data scientist is setting up a new SageMaker pipeline for automating an ML workflow. They need to define the steps in the pipeline, which include training a model, registering the model in the model registry, and deploying it. What structure should they follow to create a functional SageMaker pipeline?
-
A
Only define the steps, without specifying any parameters or names, since SageMaker Pipelines automatically handles these.
-
B
Create a notebook in SageMaker Studio, then manually execute each step in the pipeline.
-
C
Use AWS CloudFormation to define and automate all steps of the pipeline without utilizing SageMaker's pipeline features.
-
D
Define the name of the pipeline, the steps for each action, and parameters for step configuration.
Xem giải thích
Đáp án
D — Khai tên pipeline, các bước cho từng hành động, và tham số
Vì sao đúng
Một định nghĩa pipeline hoàn chỉnh cần đúng ba thứ: tên để nhận diện và quản lý phiên bản, danh sách bước mô tả từng hành động cùng phụ thuộc giữa chúng, và tham số để truyền giá trị vào lúc chạy thay vì viết cứng. Có tham số thì cùng một pipeline dùng lại được cho nhiều bộ dữ liệu hay nhiều môi trường.
Vì sao các phương án khác sai
- A. Chỉ khai bước, bỏ tên và tham số — thiếu tên thì không quản lý được, thiếu tham số thì mọi giá trị bị viết cứng.
- B. Chạy tay từng bước trong notebook — trái thẳng mục đích tự động hoá.
- C. Dùng CloudFormation thay Pipelines — CloudFormation dựng hạ tầng, không điều phối quy trình học máy.
A healthcare company wants to transcribe doctor-patient conversations into text for further analysis and store the results in Amazon S3. They need the solution to automatically remove any personally identifiable information (PII) to ensure compliance with HIPAA regulations. Which combination of AWS services should be used to achieve this goal?
-
A
Use Amazon Transcribe Medical to convert speech to text, utilize AWS Lambda to trigger Amazon Comprehend Medical for PII detection and redaction, and store the processed text in Amazon S3.
-
B
Use Amazon Transcribe Medical to convert speech to text, store the transcriptions in Amazon S3, and enable Amazon Macie to scan and redact PII.
-
C
Use Amazon Transcribe to convert speech to text, store the transcriptions in Amazon DynamoDB, and use AWS Glue to detect and redact PII from the data.
-
D
Use Amazon Transcribe to convert speech to text, store the transcriptions in Amazon RDS, and manually review the data for PII.
Xem giải thích
Đáp án
A — Amazon Transcribe Medical chuyển giọng nói thành văn bản, Lambda lo phần tự động hoá
Vì sao đúng
Điểm quyết định là Transcribe Medical, bản chuyên cho y tế: nó được huấn luyện trên từ vựng lâm sàng nên nhận đúng tên thuốc, thuật ngữ chẩn đoán và cách nói của bác sĩ — thứ mà bản phổ thông hay nghe nhầm. Lambda gắn vào sự kiện S3 để tự khởi động phiên chép lại khi có bản ghi âm mới, và ghi kết quả trở lại S3.
Vì sao các phương án khác sai
- B — dùng đúng Transcribe Medical nhưng thiếu phần tự động hoá mà đề yêu cầu.
- C và D. Dùng Transcribe bản thường — sẽ sai thuật ngữ y khoa; với hồ sơ bệnh án thì đó là lỗi nghiêm trọng.
A data scientist is tasked with ensuring that their machine learning model is trustworthy, transparent, and avoids potential risks, including bias and poor generalization. They want to use a technique that assesses the model's fairness and explainability, both before and after training, while also meeting regulatory requirements.
Which AWS service can help them achieve this?
-
A
Amazon SageMaker Clarify
-
B
Amazon Personalize
-
C
AWS Glue
-
D
Amazon SageMaker Ground Truth
Xem giải thích
Đáp án
A — Amazon SageMaker Clarify
Vì sao đúng
Clarify là công cụ chuyên cho AI có trách nhiệm và phủ đúng những gì đề nêu: đo thiên lệch trước khi huấn luyện (dữ liệu có mất cân đối giữa các nhóm không) và sau khi huấn luyện (mô hình có đối xử khác nhau không), cộng với giải thích dự đoán bằng giá trị SHAP để biết đặc trưng nào đẩy kết quả về đâu. Đó là phần "minh bạch và đáng tin".
Vì sao các phương án khác sai
- B. Amazon Personalize — dịch vụ gợi ý, không liên quan.
- C. AWS Glue — ETL.
- D. Ground Truth — gán nhãn dữ liệu; có giúp giảm thiên lệch ở khâu nhãn nhờ người rà soát, nhưng nó không đo thiên lệch của mô hình và không giải thích dự đoán.
What is the main difference between identity-based policies and resource-based policies in AWS IAM?
-
A
Identity-based policies are applied to IAM users, groups, or roles, whereas resource-based policies are directly attached to AWS resources like S3 buckets.
-
B
Identity-based policies can only be applied to groups, while resource-based policies can only be applied to individual users.
-
C
Resource-based policies must always be AWS managed policies, while identity-based policies must be inline policies.
-
D
Identity-based policies are applied directly to AWS resources, while resource-based policies are attached to users, groups, or roles.
Xem giải thích
Đáp án
A — Chính sách theo danh tính gắn vào user, group hoặc role; chính sách theo tài nguyên gắn vào chính tài nguyên
Vì sao đúng
Khác biệt nằm ở chỗ gắn và ở câu hỏi mà mỗi loại trả lời. Chính sách theo danh tính đi theo người gọi: "danh tính này được làm gì". Chính sách theo tài nguyên đi theo thứ bị gọi: "ai được đụng vào tài nguyên này" — ví dụ bucket policy, key policy của KMS, trust policy của vai.
Hệ quả thực tế quan trọng: chính sách theo tài nguyên có trường Principal, còn chính sách theo danh tính thì không (vì principal chính là thứ nó được gắn vào). Chính sách theo tài nguyên cũng là cách cấp quyền liên tài khoản mà không cần đóng vai.
Vì sao các phương án khác sai
- B. Chỉ gắn được vào group — sai; gắn được cho cả user, group và role.
- C. Chính sách theo tài nguyên phải là chính sách do AWS quản lý — sai, chúng do bạn viết.
- D — đảo ngược hoàn toàn hai khái niệm.
You are tasked with training a deep neural network model with billions of parameters using Amazon SageMaker. The model is too large to fit into a single GPU, and the dataset is also substantial. Which SageMaker technique should you use to ensure efficient training?
-
A
Use only a single machine with larger memory to handle both the data and the model.
-
B
Data parallelism, where the dataset is split across multiple machines but the model remains on one GPU.
-
C
Model parallelism, where the model itself is split across multiple GPUs, and each GPU handles a portion of the model.
-
D
Use pre-built SageMaker containers without any adjustments for distributed training.
Xem giải thích
Đáp án
C — Song song theo mô hình: chia chính mô hình ra nhiều GPU
Vì sao đúng
Đề nói rõ mô hình không vừa một GPU. Chỉ song song theo mô hình giải được ràng buộc bộ nhớ đó: chia theo lớp hoặc theo tensor, mỗi GPU giữ một phần tham số và truyền kết quả trung gian cho nhau. Thư viện của SageMaker tự phân tích đồ thị và chia sao cho cân bằng.
Vì sao các phương án khác sai
- B. Song song theo dữ liệu — mỗi máy vẫn phải giữ bản sao đầy đủ của mô hình, nên không giải quyết được vấn đề bộ nhớ.
- A. Dùng một máy bộ nhớ lớn hơn — có trần: mô hình hàng tỷ tham số vượt quá mọi loại máy đơn.
- D. Dùng container dựng sẵn mà không chỉnh gì — huấn luyện phân tán không tự bật; phải khai cấu hình.
An organization is working on a text classification problem and decides to use Amazon SageMaker's built-in algorithms. The team is concerned about optimizing model performance without manually tuning hyperparameters. What features of SageMaker can help address this requirement?
-
A
Use SageMaker's automatic hyperparameter tuning with the built-in linear learner algorithm.
-
B
Manually adjust hyperparameters for the built-in algorithm by trial and error.
-
C
Utilize custom algorithms to bypass hyperparameter tuning altogether.
-
D
Use SageMaker JumpStart to automatically deploy pre-trained text generation models without any customization.
Xem giải thích
Đáp án
A — Dùng tinh chỉnh siêu tham số tự động với thuật toán dựng sẵn
Vì sao đúng
Kết hợp này cho kết quả tốt nhất trên công sức bỏ ra: thuật toán dựng sẵn đã được tối ưu và có sẵn container, còn Automatic Model Tuning dò khoảng siêu tham số bằng tối ưu Bayes nên tìm ra cấu hình tốt với ít lần thử. Chạy song song được và có thể dừng sớm nhánh kém.
Vì sao các phương án khác sai
- B. Chỉnh tay bằng cách thử dần — chậm, không hệ thống, và không lặp lại được.
- C. Dùng thuật toán tự viết để khỏi phải tinh chỉnh — tự viết không loại bỏ nhu cầu tinh chỉnh; mô hình nào cũng có siêu tham số.
- D. Dùng JumpStart triển khai mô hình sinh văn bản — sai dạng bài toán: đây là phân loại văn bản, không phải sinh văn bản.