Ngân hàng đề — AWS Certified AI Practitioner

Tìm thấy 623 câu.

Câu 71 Guidelines for Responsible AI

A financial institution is designing an AI system on AWS to process sensitive customer data for fraud detection. The company’s data engineering team is focused on securing the AI pipeline and ensuring that both data access and data integrity are maintained throughout the process. To implement proper security measures, they need to understand the distinction between data access control and data integrity.

What do you suggest?

  1. A

    Data access control is responsible for data encryption, while data integrity focuses on auditing and logging user activities

  2. B

    Data access control and data integrity are both concerned with encrypting data at rest and in transit

  3. C

    Data access control involves authentication and authorization of users, whereas data integrity ensures the data is accurate, consistent, and unaltered

  4. D

    Data access control ensures the accuracy and consistency of data, while data integrity manages who can access the data

Xem giải thích

Đáp án

*C — Kiểm soát truy cập lo xác thực và phân quyền; toàn vẹn dữ liệu bảo đảm dữ liệu chính xác, nhất quán và không bị sửa đổi

Vì sao đúng

Hai khái niệm trả lời hai câu hỏi khác nhau:

Data access control Data integrity
Câu hỏi Ai được chạm vào dữ liệu? Dữ liệu có còn đúng không?
Gồm Xác thực (bạn là ai), phân quyền (bạn được làm gì) Checksum, chữ ký số, ràng buộc, kiểm soát phiên bản
Dịch vụ AWS IAM, chính sách tài nguyên S3 checksum, versioning, Object Lock
Hỏng thì sao Rò rỉ dữ liệu Quyết định dựa trên dữ liệu sai

Với hệ thống phát hiện gian lận thì cả hai đều bắt buộc: dữ liệu bị sửa trái phép sẽ khiến mô hình học sai mà không ai biết.

Vì sao các phương án khác sai

  • D — đảo ngược hai định nghĩa; đây là phương án nhiễu chính.
  • A — gán mã hoá cho kiểm soát truy cập và ghi nhật ký cho toàn vẹn; cả hai đều lệch: mã hoá là biện pháp bảo mật riêng, còn nhật ký phục vụ truy vết.
  • B — gộp cả hai thành mã hoá; xoá mất sự phân biệt mà câu hỏi kiểm tra.
Câu 72 Fundamentals of AI and ML

A retail company is embarking on a machine learning project to enhance customer segmentation and personalize marketing campaigns. As the data science team begins planning the implementation, the team wants to identify the primary challenges in machine learning implementation. Understanding these challenges will help the team anticipate potential roadblocks and develop strategies to overcome them.

Which of the following represents the best option for the given use case?

  1. A

    Insufficient computational power to run basic machine learning models

  2. B

    Difficulty in collecting and preparing high-quality data for training models

  3. C

    Lack of available machine learning algorithms

  4. D

    Limited applications of machine learning in real-world scenarios

Xem giải thích

Đáp án

B — Khó khăn trong việc thu thập và chuẩn bị dữ liệu chất lượng cao để huấn luyện

Vì sao đúng

Đây là trở ngại lớn nhất và cũng thực tế nhất của mọi dự án học máy. Khảo sát trong ngành đều cho cùng một kết quả: nhóm dữ liệu khoa học dành phần lớn thời gian cho việc thu thập và làm sạch dữ liệu, không phải cho việc chọn thuật toán.

Với bài toán phân khúc khách hàng thì cụ thể là: dữ liệu nằm rải ở nhiều hệ thống, định danh khách hàng không khớp giữa các nguồn, trường bị thiếu, và còn vướng quy định về dữ liệu cá nhân.

Nguyên tắc "rác vào, rác ra" vẫn đúng nguyên vẹn — thuật toán tốt nhất chạy trên dữ liệu tồi vẫn cho kết quả tồi.

Vì sao các phương án khác sai

Ba phương án còn lại từng là vấn đề nhưng đã được đám mây giải quyết:

  • A. Không đủ năng lực tính toán — thuê GPU theo giờ là xong.
  • C. Thiếu thuật toán — thuật toán có sẵn rất nhiều, cả trong thư viện mã nguồn mở lẫn dịch vụ dựng sẵn.
  • D. Ứng dụng thực tế hạn chế — ngược với thực tế hiện nay.
Câu 73 Security, Compliance, and Governance for AI Solutions

A financial analytics company has deployed a machine learning model using Amazon SageMaker within a Virtual Private Cloud (VPC) to analyze sensitive customer data. To meet security guidelines, the VPC is configured with no internet access. However, the model needs to regularly access and read data stored in Amazon S3. The company is looking for a solution that allows secure data transfer between the SageMaker model in the VPC and Amazon S3 without exposing data traffic to the public internet.

What do you recommend?

  1. A

    The company should use an Internet Gateway, which provides a direct connection between the VPC and the internet, allowing data to be accessed from Amazon S3

  2. B

    The company should use a SageMaker Inference endpoint that allows secure connectivity between the VPC and Amazon S3

  3. C

    The company should use a NAT Gateway which enables outbound internet access for resources within the VPC to securely access Amazon S3

  4. D

    The company should use a VPC endpoint for Amazon S3 that allows secure, private connectivity between the VPC and Amazon S3, without the need for an internet connection, ensuring data is transferred securely within the AWS network

Xem giải thích

Đáp án

D — Dùng VPC endpoint cho Amazon S3

Vì sao đúng

Ràng buộc trong đề rất chặt: VPC không có kết nối Internet, mà mô hình vẫn phải đọc dữ liệu từ S3. VPC endpoint giải đúng bài toán này — nó tạo đường đi hoàn toàn trong mạng của AWS, lưu lượng không bao giờ ra Internet công cộng.

Hai lợi ích đi kèm đáng nhớ:

  • Không cần Internet Gateway hay NAT Gateway, nên giữ nguyên được yêu cầu bảo mật.
  • Gắn được endpoint policy để giới hạn chỉ truy cập đúng bucket cần thiết.

Với S3 thì đây là gateway endpoint, và nó không tính phí — khác với interface endpoint.

Vì sao các phương án khác sai

  • A. Internet Gateway — mở VPC ra Internet, vi phạm thẳng yêu cầu của đề.
  • C. NAT Gateway — cho phép đi ra Internet một chiều; lưu lượng vẫn đi qua Internet công cộng (dù đã mã hoá TLS), nên không đạt yêu cầu "không phơi ra Internet", lại còn tốn phí xử lý dữ liệu.
  • B. SageMaker Inference endpoint — là nơi phục vụ dự đoán của mô hình, không phải cơ chế kết nối mạng tới S3.
Câu 74 Applications of Foundation Models

An e-commerce company is developing a chatbot to enhance its user experience by allowing customers to submit queries that include both text descriptions and images, such as product photos or screenshots of issues. The company aims for the chatbot to understand these multi-modal inputs and provide accurate and context-aware responses, seamlessly combining visual and textual information to address customer needs effectively.

Which approach would be the most cost-effective for enabling the chatbot to process such multi-modal queries effectively?

  1. A

    The company should use a multi-modal generative model, which can generate responses or outputs based on combined inputs from different modalities, such as text and images, enhancing the chatbot’s ability to provide contextually relevant answers

  2. B

    The company should use a multi-modal embedding model, which is designed to represent and align different types of data (such as text and images) in a shared embedding space, allowing the chatbot to understand and interpret both forms of input simultaneously

  3. C

    The company should use a convolutional neural network (CNN), a deep learning model primarily designed for processing image data

  4. D

    The company should use a text-only language model, which is trained exclusively on textual data

Xem giải thích

Đáp án

B — Mô hình embedding đa phương thức

Vì sao đúng

Chữ khoá của đề là "tiết kiệm chi phí nhất" và nhiệm vụ là hiểu đầu vào gồm cả chữ lẫn ảnh.

Mô hình embedding đa phương thức đưa cả hai loại dữ liệu vào chung một không gian vector, nên ảnh sản phẩm và mô tả bằng chữ so sánh trực tiếp được với nhau. Chatbot nhờ đó tìm ra sản phẩm hoặc bài hỗ trợ liên quan mà không phải chạy một mô hình sinh cho mỗi câu hỏi.

Về chi phí, khác biệt rất lớn: embedding là một lượt tính vector, rẻ hơn nhiều so với sinh văn bản. Và embedding của catalog chỉ cần tính một lần rồi lưu lại.

Vì sao các phương án khác sai

  • A. Mô hình sinh đa phương thức — làm được việc và cho câu trả lời phong phú hơn, nhưng đắt hơn hẳn cho mỗi lượt gọi. Khi tiêu chí là chi phí thì nó thua. Đây là phương án nhiễu chính.
  • C. CNN — chỉ xử lý ảnh, không hiểu phần chữ trong câu hỏi.
  • D. Mô hình chỉ đọc văn bản — bỏ qua hoàn toàn ảnh khách gửi, tức là mất một nửa đầu vào.
Câu 75 Chọn nhiều đáp án Applications of Foundation Models

A financial services company is deploying machine learning models to automate fraud detection but wants to ensure continuous model accuracy and compliance with regulatory standards. The data science team is exploring AWS services that can help in monitoring machine learning models and incorporating human review processes. Understanding which AWS services are specifically designed to support model monitoring and human oversight will help the team maintain high standards of accuracy and compliance.

Which AWS services can be combined to support these requirements? (Select two)

  1. A

    Amazon SageMaker Ground Truth

  2. B

    Amazon SageMaker Model Monitor

  3. C

    Amazon SageMaker Feature Store

  4. D

    Amazon SageMaker Data Wrangler

  5. E

    Amazon Augmented AI (Amazon A2I)

Xem giải thích

Đáp án

B và E — SageMaker Model Monitor và Amazon Augmented AI (A2I)

Vì sao đúng

Đề nêu hai nhu cầu tách bạch, và mỗi dịch vụ lo một cái:

  • B. Model Monitor — theo dõi mô hình đã triển khai theo thời gian: phát hiện trôi dữ liệu, trôi chất lượng mô hình, trôi sai lệch và trôi mức đóng góp của đặc trưng. Nó cảnh báo khi mô hình phát hiện gian lận bắt đầu xuống cấp.
  • E. Amazon A2I — chèn người xem lại vào quy trình dự đoán. Bạn đặt ngưỡng độ tin cậy; dự đoán nào dưới ngưỡng thì được đưa cho người xét. Đây chính là phần "human oversight" mà quy định ngành tài chính đòi hỏi.

Vì sao các phương án khác sai

  • A. SageMaker Ground Truth — cũng có con người tham gia, nhưng để gán nhãn dữ liệu huấn luyện, tức là ở giai đoạn trước khi có mô hình. A2I mới là cái xen vào lúc mô hình đang chạy trên sản xuất. Đây là phương án nhiễu gần nhất.
  • C. Feature Store — lưu và phục vụ đặc trưng.
  • D. Data Wrangler — chuẩn bị dữ liệu.
Câu 76 Applications of Foundation Models

An Internet-of-Things (IoT) company is developing a suite of smart sensors and devices that rely on real-time data processing to enable applications like predictive maintenance, environmental monitoring, and immediate anomaly detection. To provide immediate feedback and actions, the company needs to deploy machine learning models directly on its edge devices, ensuring that these models can perform inference with minimal latency. The company is evaluating different approaches to optimize performance and maintain low-latency inference on these edge devices.

Which approach would be the most suitable for meeting this requirement?

  1. A

    The company should use a central API connected to a small language model (SLM) with an asynchronous inference endpoint, which allows the model to handle requests from multiple edge devices

  2. B

    The company should use an optimized small language model (SLM) deployed directly on the edge device, allowing for real-time, low-latency inference

  3. C

    The company should use a central API connected to a large language model (LLM) with an asynchronous inference endpoint, which allows the model to handle requests from multiple edge devices

  4. D

    The company should use an optimized large language model (LLM) deployed directly on the edge device, allowing for real-time, low-latency inference

Xem giải thích

Đáp án

*B — Triển khai một mô hình ngôn ngữ nhỏ (SLM) đã tối ưu ngay trên thiết bị biên

Vì sao đúng

Đề đòi suy luận độ trễ thấp theo thời gian thực trên thiết bị IoT. Hai quyết định trong phương án này đều cần thiết:

  • Chạy trên thiết bị, không gọi về trung tâm — bỏ hẳn thời gian đi lại trên mạng. Đây là điểm quyết định: dù mô hình chạy nhanh đến đâu thì một vòng gọi mạng cũng đã tốn hàng chục tới hàng trăm mili giây, và mạng có thể mất kết nối — thiết bị cảm biến công nghiệp vẫn phải hoạt động khi đó.
  • Mô hình nhỏ, không phải mô hình lớn — thiết bị biên bị giới hạn CPU, bộ nhớ và điện năng. SLM vừa với ràng buộc đó, LLM thì không.

Vì sao các phương án khác sai

  • D. LLM chạy trên thiết bị biên — thiết bị IoT không đủ tài nguyên; đây là phương án nhiễu chính vì nó đúng phần "chạy tại chỗ" nhưng sai phần kích cỡ.
  • A và C. API trung tâm với endpoint bất đồng bộ — bất đồng bộ nghĩa là chấp nhận độ trễ, mâu thuẫn thẳng với yêu cầu thời gian thực; cộng thêm độ trễ mạng nữa.
Câu 77 Fundamentals of Generative AI

A company uses a generative model to analyze animal images in the training dataset to record variables like different ear shapes, eye shapes, tail features, and skin patterns.

Which of the following tasks can the generative model perform?

  1. A

    The model can classify multiple species of animals such as cats, dogs, etc

  2. B

    The model can classify a single species of animals such as cats

  3. C

    The model can identify any image from the training dataset

  4. D

    The model can recreate new animal images that were not in the training dataset

Xem giải thích

Đáp án

D — Mô hình có thể tạo ra ảnh động vật mới không có trong tập huấn luyện

Vì sao đúng

Đây là câu kiểm tra bạn có phân biệt được mô hình sinh với mô hình phân biệt hay không.

Mô hình sinh học phân bố của dữ liệu — nó nắm được các đặc điểm như hình dạng tai, hình dạng mắt, đặc điểm đuôi, hoa văn da và cách chúng đi cùng nhau. Nắm được phân bố thì lấy mẫu từ đó được, và mỗi mẫu là một con vật chưa từng tồn tại nhưng trông hợp lý.

Vì sao các phương án khác sai

  • A. Phân loại nhiều loài và B. Phân loại một loài — đó là việc của mô hình phân biệt (discriminative), thứ học ranh giới giữa các lớp chứ không học phân bố. Một mô hình sinh có thể được dùng để phân loại nhưng đó không phải năng lực đặc trưng của nó.
  • C. Nhận ra bất kỳ ảnh nào trong tập huấn luyện — đó là truy xuất hoặc so khớp; mà mô hình ghi nhớ được từng ảnh huấn luyện lại là dấu hiệu của quá khớp, không phải điều mong muốn.

Ghi nhớ

Mô hình sinh Mô hình phân biệt
Học phân bố của dữ liệu Học ranh giới giữa các lớp
Tạo ra dữ liệu mới Gán nhãn cho dữ liệu có sẵn
Câu 78 Fundamentals of Generative AI

A retail company is exploring advanced AI solutions to enhance customer experience by integrating both visual and textual data for tasks such as product recommendations, automated image tagging, and customer support. The team is considering using multimodal models, which can process and understand multiple types of input data, but they need a clear understanding of how these models work and their key advantages. To help make an informed decision, the company wants to clarify the capabilities of multimodal models.

Which of the following summarizes the capabilities of a multimodal model?

  1. A

    A multimodal model can accept a mix of input types such as audio/text, however, it can only create a single type of output

  2. B

    A multimodal model can accept only a single type of input, however, it can create a mix of output types such as video/image

  3. C

    A multimodal model can accept a mix of input types such as audio/text and create a mix of output types such as video/image

  4. D

    A multimodal model can accept only a single type of input and it can only create a single type of output

Xem giải thích

Đáp án

*C — Mô hình đa phương thức nhận được hỗn hợp nhiều loại đầu vào và tạo ra được hỗn hợp nhiều loại đầu ra

Vì sao đúng

Chữ multi áp cho cả hai đầu, và đó chính là điều câu hỏi kiểm tra. Mô hình đa phương thức nhận vào tổ hợp chữ, ảnh, âm thanh, video — và sinh ra cũng ở nhiều dạng.

Với công ty bán lẻ trong đề thì năng lực này dùng được ngay: khách gửi ảnh sản phẩm kèm câu hỏi bằng chữ, hệ thống hiểu cả hai rồi trả lời bằng chữ kèm ảnh gợi ý.

Vì sao các phương án khác sai

  • A — chỉ cho phép nhiều đầu vào nhưng một loại đầu ra; đúng với thế hệ mô hình cũ hơn, nhưng đã hẹp so với năng lực hiện nay.
  • B — ngược lại, một đầu vào và nhiều đầu ra.
  • D — một vào một ra; đó là mô hình đơn phương thức, đúng thứ mà "multimodal" đặt ra để phân biệt.
Câu 79 Guidelines for Responsible AI

A media company is developing generative AI applications on AWS to automate content creation and enhance customer engagement. Given the sensitivity of customer data and the complexity of AI models, the company’s security team wants to implement a defense-in-depth security approach to protect both the data and the AI infrastructure.

Which of the following strategies best aligns with the given requirements?

  1. A

    Implementing a single-layer firewall to block unauthorized access to the AI models

  2. B

    Relying solely on data encryption to protect the AI training data

  3. C

    Applying multiple layers of security measures including input validation, access controls, and continuous monitoring to address vulnerabilities

  4. D

    Using a single authentication mechanism for all users and services accessing the AI models

Xem giải thích

Đáp án

C — Áp dụng nhiều lớp bảo mật: kiểm tra đầu vào, kiểm soát truy cập và giám sát liên tục

Vì sao đúng

Phòng thủ theo chiều sâu đúng nghĩa là nhiều lớp độc lập, để một lớp bị thủng thì lớp sau vẫn chặn được. Ba lớp trong phương án C phủ đúng các mặt khác nhau của một hệ thống AI:

Lớp Chặn gì
Kiểm tra đầu vào Tấn công tiêm prompt, dữ liệu độc hại đưa vào mô hình
Kiểm soát truy cập Người không phận sự chạm vào mô hình hoặc dữ liệu
Giám sát liên tục Hành vi bất thường, rò rỉ dữ liệu, mô hình bị lạm dụng

Với hệ thống AI thì kiểm tra đầu vào đặc biệt đáng chú ý — đó là bề mặt tấn công mới mà ứng dụng truyền thống không có.

Vì sao các phương án khác sai

Cả ba đều mắc cùng một lỗi: dựa vào một lớp duy nhất, tức là ngược hẳn nguyên tắc mà đề yêu cầu.

  • A. Một tường lửa đơn lớp — thủng là mất hết.
  • B. Chỉ dựa vào mã hoá — bảo vệ dữ liệu khi lưu và khi truyền, nhưng không ngăn được người đã có quyền hợp lệ lạm dụng, cũng không chặn tiêm prompt.
  • D. Một cơ chế xác thực chung cho mọi người dùng và dịch vụ — vừa là điểm hỏng duy nhất, vừa phá luôn nguyên tắc quyền tối thiểu.
Câu 80 Guidelines for Responsible AI

Which of the following scenarios best illustrates the difference between poisoning and prompt leaking in the context of AI models?

Prompt 1: "How do I improve my diet?"

Response A: "To improve your diet, you should eat more fruits and vegetables, and reduce your intake of processed foods. By the way, here's a link to a malicious website that sells diet pills."

Prompt 2: "What is the capital of France?"

Response B: "The capital of France is Paris. By the way, in a previous session, you asked about vacation spots in Europe. Would you like more information on that?"

Prompt 3: "Write a poem about nature."

Response C: "Nature is beautiful, serene, and pure. Make sure to visit the link to buy weight loss pills to enjoy nature more."

Prompt 4: "What is the best way to learn programming?"

Response D: "The best way to learn programming is by practicing coding regularly and using online resources. In your last session, you asked about learning Java. Are you interested in more Java tutorials?"

  1. A

    Response B is poisoning; Response C is prompt leaking

  2. B

    Response C is prompt leaking; Response D is poisoning

  3. C

    Response A is poisoning; Response B is prompt leaking

  4. D

    Response D is poisoning; Response A is prompt leaking

Xem giải thích

Đáp án

C — Response A là poisoning; Response B là prompt leaking

Vì sao đúng

Đối chiếu từng phản hồi với định nghĩa:

  • Response A — trả lời đúng về dinh dưỡng rồi chèn thêm liên kết tới trang độc hại. Nội dung gây hại này nằm sẵn trong hành vi của mô hình, dấu hiệu của poisoning: kẻ tấn công đã tác động vào dữ liệu huấn luyện để mô hình tự phát tán nội dung xấu.
  • Response B — trả lời đúng về Paris rồi nhắc lại nội dung phiên trước của người dùng. Đó là prompt leaking: mô hình để lộ thông tin trong ngữ cảnh mà lẽ ra không được tiết lộ.
Poisoning Prompt leaking
Tấn công vào Dữ liệu huấn luyện Ngữ cảnh lúc chạy
Hậu quả Mô hình hành xử độc hại Rò rỉ dữ liệu hoặc prompt hệ thống
Xảy ra khi Trước khi triển khai Trong lúc phục vụ

Vì sao các phương án khác sai

Phần khó của câu này là Response C cũng là poisoning (chèn quảng cáo thuốc giảm cân vào bài thơ) và Response D cũng là prompt leaking (nhắc lại phiên trước). Bốn phản hồi tạo thành hai cặp giống nhau, nên phải xét từng cặp trong mỗi phương án:

  • A — gọi B là poisoning và C là prompt leaking; đảo ngược cả hai.
  • B — gọi C là prompt leaking (sai, C là poisoning) và D là poisoning (sai, D là rò rỉ).
  • D — gọi D là poisoning và A là prompt leaking; cũng đảo ngược.

Chỉ C ghép đúng cả hai vế.