Ngân hàng đề — AWS Certified AI Practitioner
Tìm thấy 623 câu.
A media company is exploring cutting-edge AI models to automate tasks such as content generation and language translation. The development team is particularly interested in using Transformer models due to their efficiency and performance in natural language processing tasks. To make an informed decision, the team needs to identify which models belong to the Transformer architecture and how they can be applied to their use cases.
Which of the following is an example of a Transformer model?
-
A
DALL-E
-
B
Adobe Firefly
-
C
Stable Diffusion
-
D
ChatGPT
Xem giải thích
Đáp án
D — ChatGPT
Vì sao đúng
Chữ GPT viết tắt của Generative Pre-trained Transformer — tên gọi đã nói thẳng kiến trúc. ChatGPT dùng cơ chế self-attention đặc trưng của Transformer để xử lý văn bản.
Vì sao các phương án khác sai
Ba phương án còn lại đều là mô hình sinh ảnh dựa trên diffusion, một kiến trúc khác hẳn:
- A. DALL-E — sinh ảnh từ mô tả chữ.
- B. Adobe Firefly — sinh ảnh của Adobe.
- C. Stable Diffusion — tên đã có sẵn chữ "Diffusion".
Phân biệt hai kiến trúc
| Transformer | Diffusion | |
|---|---|---|
| Cơ chế | Self-attention trên chuỗi token | Khử nhiễu dần từ nhiễu ngẫu nhiên |
| Chủ yếu dùng cho | Văn bản | Ảnh |
| Ví dụ | GPT, BERT, Claude, Llama | DALL-E, Stable Diffusion, Firefly |
Cần lưu ý là ranh giới này đang mờ dần — nhiều mô hình sinh ảnh hiện nay dùng Transformer ở phần hiểu prompt. Nhưng theo cách phân loại mà đề dùng thì ChatGPT là câu trả lời.
A healthcare technology company is developing machine learning models to analyze both structured data, such as patient records, and unstructured data, such as medical images and clinical notes. The data science team is working on feature engineering to extract the most relevant information for the models but is aware that the process differs depending on whether the data is structured or unstructured. To ensure they approach each data type correctly, they need to understand the key differences in feature engineering tasks for structured versus unstructured data in machine learning.
What is a key difference in feature engineering tasks for structured data compared to unstructured data in the context of machine learning?
-
A
Feature engineering for structured data is not necessary as the data is already in a usable format, whereas for unstructured data, extensive preprocessing is always required
-
B
Feature engineering for structured data focuses on image recognition, whereas for unstructured data, it focuses on numerical data analysis
-
C
Feature engineering tasks for structured data and unstructured data are identical and do not vary based on data type
-
D
Feature engineering for structured data often involves tasks such as normalization and handling missing values, while for unstructured data, it involves tasks such as tokenization and vectorization
Xem giải thích
Đáp án
D — Với dữ liệu có cấu trúc thì làm chuẩn hoá và xử lý giá trị thiếu; với dữ liệu phi cấu trúc thì tách token và vector hoá
Vì sao đúng
Kỹ thuật kỹ nghệ đặc trưng khác nhau vì điểm xuất phát khác nhau:
| Dữ liệu có cấu trúc (hồ sơ bệnh nhân) | Dữ liệu phi cấu trúc (ảnh y khoa, ghi chú lâm sàng) | |
|---|---|---|
| Đã có | Các cột rõ ràng, kiểu dữ liệu xác định | Chỉ là văn bản thô hoặc điểm ảnh |
| Việc chính | Chuẩn hoá thang đo, xử lý giá trị thiếu, mã hoá biến phân loại | Tách token, vector hoá (embedding), trích đặc trưng ảnh |
| Bản chất | Tinh chỉnh đặc trưng đã có | Tạo ra đặc trưng từ con số không |
Nói gọn: dữ liệu có cấu trúc đã sẵn dạng bảng để mô hình đọc, chỉ cần dọn dẹp; dữ liệu phi cấu trúc phải biến thành số trước đã.
Vì sao các phương án khác sai
- A. Dữ liệu có cấu trúc không cần kỹ nghệ đặc trưng — sai: chuẩn hoá thang đo và xử lý giá trị thiếu là bắt buộc, bỏ qua thì mô hình lệch hẳn.
- B — đảo ngược: gán nhận dạng ảnh cho dữ liệu có cấu trúc.
- C. Giống hệt nhau, không khác gì theo loại dữ liệu — phủ nhận đúng điều câu hỏi kiểm tra.
An app developer is building an educational application to help high-school students understand fundamental concepts in mathematics, such as calculating the probability of drawing a spade from a deck of cards.
Which approach would be the most suitable for this purpose?
-
A
The developer should apply supervised learning, a machine learning approach where the model learns from labeled datasets to make predictions
-
B
The developer should utilize unsupervised learning, a machine learning method that identifies patterns and structures in data without using labeled datasets
-
C
The developer should create a rule-based application that uses predefined mathematical rules and formulas to answer probability questions accurately
-
D
The developer should leverage reinforcement learning (RL), a type of machine learning where an agent learns to make decisions by receiving rewards or penalties
Xem giải thích
Đáp án
C — Xây ứng dụng dựa trên luật, dùng các quy tắc và công thức toán học định sẵn
Vì sao đúng
Đây là câu kiểm tra một điều quan trọng hơn kỹ thuật: biết khi nào KHÔNG nên dùng học máy.
Xác suất rút được quân bích từ bộ bài là 13/52 = 25% — một phép tính xác định, có công thức, luôn cho cùng kết quả. Học máy giải bài toán khi không có công thức và phải suy ra quy luật từ dữ liệu. Ở đây công thức đã có sẵn từ vài trăm năm.
Ứng dụng dựa trên luật còn hơn hẳn ở ba điểm mà bối cảnh giáo dục rất cần:
- Luôn đúng 100% — mô hình học máy thì không bảo đảm được điều đó, mà dạy sai học sinh thì tệ.
- Giải thích được từng bước — chính là thứ ứng dụng dạy học phải làm.
- Không cần dữ liệu huấn luyện, không cần huấn luyện, không cần hạ tầng.
Vì sao các phương án khác sai
- A. Học có giám sát — cần hàng nghìn mẫu có nhãn để học lại một công thức đã biết, và vẫn có sai số.
- B. Học không giám sát — dùng để tìm nhóm ẩn trong dữ liệu, không trả lời được câu hỏi xác suất.
- D. Học tăng cường — cần môi trường và cơ chế thưởng phạt; hoàn toàn lệch bài toán.
A legal firm is looking to implement an AI solution that can generate detailed, accurate responses to client queries by retrieving relevant information from its extensive database of legal documents. The firm is considering the use of Retrieval Augmented Generation (RAG) through Amazon Bedrock to enhance the quality and relevance of the generated content. The team wants to understand the best-fit use cases for RAG to determine if it aligns with their needs for knowledge retrieval and content generation.
Which of the following represents the best-fit use cases for utilizing Retrieval Augmented Generation (RAG) in Amazon Bedrock? (Select two)
-
A
Medical queries chatbot
-
B
Product recommendations that match shopper preferences
-
C
Customer service chatbot
-
D
Image generation from text prompt
-
E
Original content creation
Xem giải thích
Đáp án
A và C — chatbot y khoa và chatbot chăm sóc khách hàng
Vì sao đúng
RAG hợp nhất với những trường hợp cần trả lời chính xác dựa trên kho tài liệu có sẵn:
- A. Chatbot y khoa — câu trả lời phải bám vào tài liệu y khoa đã thẩm định, và phải dẫn được nguồn. Đây là lĩnh vực mà ảo giác gây hậu quả nặng nhất, nên việc buộc mô hình trích từ tài liệu thật là bắt buộc.
- C. Chatbot chăm sóc khách hàng — trả lời dựa trên tài liệu sản phẩm và chính sách của công ty, vốn thay đổi liên tục. RAG cập nhật bằng cách đánh chỉ mục lại, không phải huấn luyện lại.
Điểm chung: có kho tri thức xác định để truy xuất, và độ chính xác quan trọng hơn sáng tạo.
Vì sao các phương án khác sai
- B. Gợi ý sản phẩm theo sở thích người mua — đó là bài toán hệ gợi ý dựa trên hành vi người dùng (Amazon Personalize), không phải truy xuất tài liệu.
- D. Sinh ảnh từ mô tả chữ — RAG làm việc với văn bản để truy xuất, không liên quan tới sinh ảnh.
- E. Sáng tác nội dung nguyên bản — ngược hẳn mục đích của RAG: RAG ràng buộc mô hình vào tài liệu có sẵn, còn sáng tác cần tự do.
A retail company is developing machine learning models on AWS to improve product recommendations and customer insights. To ensure consistency and collaboration among its data science team, the company needs a solution for storing, sharing, and managing the inputs used during the model training and inference phases. The company is evaluating AWS services that can help streamline this process.
What do you suggest
-
A
Amazon SageMaker Ground Truth
-
B
Amazon SageMaker Clarify
-
C
Amazon SageMaker Feature Store
-
D
Amazon SageMaker Data Wrangler
Xem giải thích
Đáp án
C — Amazon SageMaker Feature Store
Vì sao đúng
Đề nêu đúng ba việc mà Feature Store sinh ra để làm: lưu trữ, chia sẻ và quản lý các đầu vào dùng trong cả huấn luyện lẫn suy luận.
Vế cuối mới là điểm quyết định. Feature Store có hai kho song song:
| Offline store | Online store | |
|---|---|---|
| Dùng cho | Huấn luyện — đọc cả khối lịch sử | Suy luận — tra cứu độ trễ mili giây |
| Lưu ở | S3 | Kho khoá–giá trị độ trễ thấp |
Điều này giải một vấn đề rất hay gặp và rất khó tìm ra: lệch giữa huấn luyện và phục vụ. Nếu đội huấn luyện tính đặc trưng bằng một đoạn mã, còn hệ thống sản xuất tính lại bằng đoạn mã khác, hai bên sẽ lệch nhau âm thầm và mô hình kém đi mà không ai biết vì sao. Feature Store bảo đảm cùng một định nghĩa đặc trưng cho cả hai phía, đồng thời cho cả đội dùng lại đặc trưng của nhau.
Vì sao các phương án khác sai
- D. Data Wrangler — tạo ra đặc trưng; Feature Store lưu chúng lại. Hai bước nối tiếp nhau, và đề hỏi về lưu trữ và chia sẻ.
- A. Ground Truth — gán nhãn dữ liệu.
- B. Clarify — phát hiện sai lệch và giải thích dự đoán.
A financial services company is exploring Amazon Bedrock to streamline its AI development for use cases such as fraud detection, personalized customer service, and automated reporting. The company is particularly interested in understanding the key features and benefits of Amazon Bedrock, including its ability to simplify access to powerful foundation models, support customizations, and integrate with existing AWS services.
To make an informed decision, the company needs to identify which of the following accurately applies to Amazon Bedrock and its capabilities? (Select two)
-
A
You can use a customized model in the Provisioned Throughput or On-Demand mode
-
B
Smaller models are cheaper to use than larger models
-
C
You can use a customized model only in the Provisioned Throughput mode
-
D
Larger models are cheaper to use than smaller models
-
E
You can use the On-Demand mode only with time-based term commitments
Xem giải thích
Đáp án
B và C
Vì sao đúng
- C. Mô hình đã tuỳ chỉnh chỉ dùng được ở chế độ Provisioned Throughput — đây là ràng buộc quan trọng nhất về chi phí của Bedrock. Mô hình tuỳ chỉnh là của riêng bạn, nên AWS phải dành riêng năng lực tính toán thay vì gộp chung như mô hình nền. Hệ quả: trả tiền cho năng lực đã đặt trước dù có gọi hay không, nên tinh chỉnh chỉ đáng khi dùng đủ nhiều.
- B. Mô hình nhỏ rẻ hơn mô hình lớn — đúng và hợp trực giác: mô hình lớn cần nhiều tài nguyên tính toán hơn cho mỗi token.
Vì sao các phương án khác sai
- A. Mô hình tuỳ chỉnh dùng được ở cả Provisioned Throughput lẫn On-Demand — mâu thuẫn trực tiếp với C, và C mới là cái đúng.
- D. Mô hình lớn rẻ hơn mô hình nhỏ — đảo ngược.
- E. On-Demand chỉ dùng được kèm cam kết theo thời hạn — sai và ngược hẳn ý nghĩa của tên gọi: On-Demand nghĩa là KHÔNG cam kết gì, trả theo token đã dùng. Cam kết theo thời hạn là đặc điểm của Provisioned Throughput.
A software company is developing a generative AI model for language translation and needs to optimize the way the model processes and understands text. The development team is focusing on improving the model’s ability to convert words into a form that the AI can effectively interpret and generate accurate translations. To achieve this, they need to clarify the roles of tokens and embeddings in the model’s language processing.
Which of the following summarizes the differences between a token and an embedding in the context of generative AI?
-
A
A token is a sequence of characters that a model can interpret or predict as a single unit of meaning, whereas, an embedding is a vector of numerical values that represents condensed information obtained by transforming input into that vector
-
B
An embedding is a sequence of characters that a model can interpret or predict as a single unit of meaning, whereas, a token is a vector of numerical values that represents condensed information obtained by transforming input into that vector
-
C
Both token and embedding refer to a sequence of characters that a model can interpret or predict as a single unit of meaning
-
D
Both token and embedding refer to a vector of numerical values that represents condensed information obtained by transforming input into that vector
Xem giải thích
Đáp án
A — Token là chuỗi ký tự mà mô hình hiểu như một đơn vị nghĩa; embedding là vector số biểu diễn thông tin đã nén của đầu vào đó
Vì sao đúng
Hai khái niệm nằm ở hai bước liên tiếp trong đường đi của văn bản qua mô hình:
"Học máy rất thú vị"
│ tách token
▼
["Học", "máy", "rất", "thú", "vị"] ← TOKEN: vẫn là ký tự
│ tra bảng embedding
▼
[[0.21, -0.45, 0.88, ...], ...] ← EMBEDDING: đã thành số
│
▼
mô hình xử lý
Mô hình không tính toán trên chữ, nên bước chuyển sang vector là bắt buộc. Và vì là vector, hai từ gần nghĩa sẽ nằm gần nhau trong không gian đó — đó là nền tảng của tìm kiếm ngữ nghĩa và RAG.
Về mặt thực dụng: token là thứ bạn bị tính tiền, còn embedding là thứ được lưu trong kho vector.
Vì sao các phương án khác sai
- B — đảo ngược hai định nghĩa, đây là phương án nhiễu chính.
- C — nói cả hai đều là chuỗi ký tự; sai ở embedding.
- D — nói cả hai đều là vector số; sai ở token.
A healthcare technology company is developing AI-driven applications to assist doctors in diagnosing diseases. As part of its commitment to ethical standards, the company wants to ensure that its AI models are fair, transparent, and free from bias. To achieve this, the data science team is exploring AWS services and tools that can help implement Responsible AI practices, as understanding which AWS services support these practices is critical for the company’s AI development strategy.
Which AWS services/tools can be used to implement Responsible AI practices? (Select two)
-
A
Amazon SageMaker JumpStart
-
B
Amazon SageMaker Model Monitor
-
C
Amazon Inspector
-
D
Amazon SageMaker Clarify
-
E
AWS Audit Manager
Xem giải thích
Đáp án
B và D — SageMaker Model Monitor và SageMaker Clarify
Vì sao đúng
Đề nêu ba mục tiêu — công bằng, minh bạch, không thiên lệch — và hai dịch vụ này phủ chúng ở hai thời điểm khác nhau:
- D. Clarify — trước và trong lúc phát triển: phát hiện sai lệch trong dữ liệu huấn luyện, đo sai lệch trong dự đoán, và giải thích đặc trưng nào ảnh hưởng bao nhiêu (dựa trên giá trị SHAP). Với chẩn đoán bệnh thì khả năng giải thích là điều kiện bắt buộc, không phải tuỳ chọn.
- B. Model Monitor — sau khi triển khai: theo dõi trôi dữ liệu, trôi chất lượng mô hình và trôi sai lệch theo thời gian. Điều này cần thiết vì một mô hình công bằng lúc ra mắt vẫn có thể lệch dần khi dân số bệnh nhân thay đổi.
Vì sao các phương án khác sai
- E. AWS Audit Manager — thu thập bằng chứng cho kiểm toán tuân thủ theo chuẩn như HIPAA; liên quan tới quản trị nói chung nhưng không đo sai lệch hay giải thích mô hình. Đây là phương án nhiễu gần nhất trong bối cảnh y tế.
- C. Amazon Inspector — quét lỗ hổng bảo mật của workload; là bảo mật hạ tầng, không phải AI có trách nhiệm.
- A. SageMaker JumpStart — kho mô hình dựng sẵn.
A retail company is exploring AI technologies to improve its inventory management by analyzing images from store cameras and shelves. The development team is considering both computer vision and image processing for different tasks but wants to understand the key differences between the two. Knowing how these technologies differ in terms of their capabilities — whether for recognizing objects, making predictions, or simply manipulating images — will help the team choose the right approach for each task.
Given this context, how would you highlight the differences between computer vision and image processing?
-
A
Computer vision focuses on enhancing and manipulating images for visual quality, whereas image processing involves interpreting and understanding the content of images to make decisions
-
B
Image processing focuses on enhancing and manipulating images for visual quality, whereas computer vision involves interpreting and understanding the content of images to make decisions
-
C
Computer vision and image processing are identical fields with no distinct differences in their applications or techniques
-
D
Image processing uses machine learning algorithms, while computer vision relies solely on pre-programmed rules
Xem giải thích
Đáp án
B — Xử lý ảnh lo cải thiện và biến đổi ảnh; thị giác máy tính lo diễn giải và hiểu nội dung ảnh để ra quyết định
Vì sao đúng
Phân biệt gọn nhất là nhìn vào đầu ra:
| Image processing | Computer vision | |
|---|---|---|
| Vào → Ra | Ảnh → Ảnh | Ảnh → Hiểu biết (nhãn, toạ độ, quyết định) |
| Câu hỏi | Làm ảnh này rõ hơn thế nào? | Trong ảnh này có gì? |
| Ví dụ | Khử nhiễu, chỉnh sáng, cắt cúp, xoay | Nhận diện vật thể, đếm số lượng, phát hiện khuôn mặt |
Với công ty bán lẻ trong đề thì cả hai đều dùng và thường nối tiếp nhau: xử lý ảnh làm sạch ảnh camera thiếu sáng trước (tiền xử lý), rồi thị giác máy tính đếm số hộp còn trên kệ.
Vì sao các phương án khác sai
- A — đảo ngược hai định nghĩa; đây là phương án nhiễu chính.
- D — nói xử lý ảnh dùng học máy còn thị giác máy tính chỉ dựa vào luật lập trình sẵn; ngược hẳn thực tế, thị giác máy tính hiện đại gần như hoàn toàn dựa trên học sâu.
- C — nói hai lĩnh vực giống hệt nhau; phủ nhận đúng điều đang được hỏi.
A company developing AI-powered customer service chatbots is exploring ways to improve the quality and accuracy of responses using Reinforcement Learning from Human Feedback (RLHF). The data science team is considering using Amazon SageMaker Ground Truth to assist with gathering and processing human feedback during model training. To ensure this solution aligns with their needs, they want to understand how SageMaker Ground Truth supports the key capabilities required for implementing RLHF, such as collecting, labeling, and managing human input effectively.
What do you suggest?
-
A
SageMaker Ground Truth enables the creation of high-quality labeled datasets by incorporating human feedback in the labeling process, which can be used to improve reinforcement learning models
-
B
SageMaker Ground Truth automatically generates synthetic data for training reinforcement learning models without any human intervention
-
C
SageMaker Ground Truth uses pre-trained models to eliminate the need for human feedback in the reinforcement learning process
-
D
SageMaker Ground Truth is specifically designed for real-time decision-making in autonomous systems, bypassing the need for any data labeling
Xem giải thích
Đáp án
A — Ground Truth tạo ra tập dữ liệu có nhãn chất lượng cao bằng cách đưa phản hồi của con người vào quy trình gán nhãn
Vì sao đúng
RLHF cần một thứ mà máy không tự sinh ra được: đánh giá của con người về việc câu trả lời nào tốt hơn. Ground Truth là dịch vụ AWS cung cấp đúng hạ tầng đó — quản lý đội gán nhãn, giao diện làm việc, cơ chế kiểm soát chất lượng và đồng thuận giữa nhiều người.
Trong quy trình RLHF, dữ liệu đó dùng để huấn luyện mô hình phần thưởng — mô hình học cách chấm điểm câu trả lời thay cho con người, rồi mô hình chính được tối ưu theo điểm đó.
Ground Truth cũng có gán nhãn tự động cho phần dễ, để con người tập trung vào phần khó — cách cắt giảm chi phí thực tế nhất của khâu này.
Vì sao các phương án khác sai
- B. Tự sinh dữ liệu tổng hợp mà không cần con người — sai, và mâu thuẫn thẳng với chữ HF (human feedback) trong tên phương pháp.
- C. Dùng mô hình huấn luyện sẵn để loại bỏ nhu cầu phản hồi của con người — cũng phủ nhận chính thành phần cốt lõi.
- D. Dựng cho việc ra quyết định thời gian thực trong hệ thống tự hành, bỏ qua gán nhãn — mô tả hoàn toàn sai về dịch vụ.