Ngân hàng đề — AWS Certified Data Engineer Associate

Tìm thấy 867 câu.

Câu 241 Data Ingestion and Transformation

A company is building a real-time recommendation system using SageMaker. The system needs to quickly retrieve input data (features) for low-latency model predictions. The features are stored in SageMaker Feature Store. Which feature store option should the company use to meet these requirements?

  1. A

    Use S3 to store features and perform real-time lookups from SageMaker models

  2. B

    Offline store, as it provides scalable data access for batch predictions

  3. C

    Online store, optimized for real-time applications with low-latency access

  4. D

    Use SageMaker Autopilot to automate the feature retrieval process for predictions

Xem giải thích

Đáp án

C — Online store, tối ưu cho truy cập độ trễ thấp

Vì sao đúng

Feature Store có hai kho cho hai mục đích khác nhau. Online store giữ giá trị đặc trưng mới nhất của từng thực thể và trả về trong vài mili giây theo khoá — đúng thứ hệ thống gợi ý cần khi phải trả lời ngay lúc người dùng đang xem trang.

Vì sao các phương án khác sai

  • B. Offline store — lưu toàn bộ lịch sử trên S3, dùng để huấn luyện và dự đoán theo lô; độ trễ hoàn toàn không hợp cho thời gian thực.
  • A. Tự lưu đặc trưng trên S3 — chính là dựng lại offline store bằng tay, và vẫn chậm.
  • D. SageMaker Autopilot — tự chọn mô hình giúp bạn, không liên quan tới việc lấy đặc trưng lúc suy luận.
Câu 242 Data Operations and Support

A company ingests streaming data from Amazon Kinesis Data Streams and wants to analyze this data in real-time using Amazon Redshift. They want to periodically store the results of an aggregation query for faster retrieval. Which feature of Amazon Redshift will best meet this requirement?

  1. A

    Use Redshift Spectrum to query the Kinesis stream as external data.

  2. B

    Use a Federated Query to directly query the Kinesis stream.

  3. C

    Use a materialized view with automatic refresh to store and update the pre-computed results.

  4. D

    Set up an ETL job using AWS Glue to store the Kinesis data in Redshift tables.

Xem giải thích

Đáp án

C — Dùng materialized view có tự làm mới

Vì sao đúng

Redshift nạp được thẳng từ Kinesis bằng streaming ingestion, và materialized view đặt trên đó sẽ tính trước kết quả tổng hợp rồi lưu lại. Bật tự làm mới thì Redshift tự cập nhật khi có dữ liệu mới, nên bảng điều khiển đọc kết quả đã tính sẵn thay vì gộp lại từ đầu mỗi lần mở.

Vì sao các phương án khác sai

  • A. Spectrum truy vấn luồng Kinesis — Spectrum chỉ đọc tệp trên S3, không đọc được luồng.
  • B. Federated Query tới Kinesis — Federated Query dành cho CSDL quan hệ như RDS, không phải luồng.
  • D. Glue ETL theo lô — chạy theo mẻ nên trễ hơn hẳn, và phải vận hành thêm một job nữa.
Câu 243 Data Operations and Support

A team of data scientists is collaborating on a machine learning project using Amazon SageMaker. They want to track and compare various model iterations, including hyperparameter settings, data processing steps, and evaluation metrics. Which feature of SageMaker is best suited for this task?

  1. A

    SageMaker Experiments

  2. B

    SageMaker Autopilot

  3. C

    SageMaker Studio

  4. D

    SageMaker Feature Store

Xem giải thích

Đáp án

A — SageMaker Experiments

Vì sao đúng

Experiments ghi lại từng lần chạy kèm siêu tham số đã dùng, dữ liệu đầu vào, mã nguồn và chỉ số kết quả, rồi cho xếp cạnh nhau để so. Đó đúng là vấn đề của một nhóm cùng thử nhiều hướng: không có nó thì sau vài chục lần chạy chẳng ai nhớ bản tốt nhất đã dùng cấu hình gì.

Vì sao các phương án khác sai

  • B. Autopilot — tự thử nhiều mô hình giúp bạn, nhưng đây là nhóm tự thiết kế thí nghiệm của họ.
  • C. Studio — môi trường làm việc; nó hiển thị Experiments chứ bản thân không phải cơ chế theo dõi.
  • D. Feature Store — quản lý đặc trưng dùng chung, không ghi lại lịch sử huấn luyện.
Câu 244 Data Ingestion and Transformation

A company is ingesting real-time data from IoT sensors and wants to perform real-time analytics. They need a solution that can handle large volumes of streaming data and provide the ability to process, store, and analyze the data with minimal manual intervention. Which combination of AWS services should the company use?

  1. A

    Amazon MSK, Kinesis Data Firehose, Amazon EMR

  2. B

    Kinesis Data Streams, Amazon S3, Amazon Athena

  3. C

    Kinesis Data Streams, Amazon Redshift, AWS Lambda

  4. D

    Kinesis Data Streams, Kinesis Data Firehose, Managed Apache Flink

Xem giải thích

Đáp án

D — Kinesis Data Streams, Kinesis Data Firehose, và Managed Apache Flink

Vì sao đúng

Ba thành phần, ba vai rõ ràng và không chồng lấn:

  • Data Streams nhận dữ liệu cảm biến, giữ lại để đọc lại được khi xử lý hỏng.
  • Managed Apache Flink phân tích trên dòng đang chảy: tính theo cửa sổ thời gian, so ngưỡng, phát hiện bất thường ngay.
  • Firehose đưa dữ liệu thô hoặc kết quả xuống nơi lưu lâu dài mà không phải viết consumer.

Vì sao các phương án khác sai

  • A. MSK, Firehose, EMR — Kafka tự quản nặng hơn cần thiết, và EMR xử lý theo lô.
  • B. Streams, S3, Athena — lưu rồi mới phân tích, không đạt "thời gian thực".
  • C. Streams, Redshift, Lambda — Lambda xử lý từng bản ghi, khó tính theo cửa sổ thời gian vì bản thân nó không giữ trạng thái.
Câu 245 Data Operations and Support

An organization uses Amazon RDS for their database needs and wants to monitor database performance in near real-time. Additionally, they need to track changes to the database configuration and ensure compliance with security best practices, such as encryption at rest. Which combination of AWS services should the organization use to meet these requirements?

  1. A

    AWS CloudWatch for monitoring RDS performance metrics, AWS CloudTrail for logging API actions on the RDS instance, and AWS Config for compliance checks on database encryption.

  2. B

    AWS CloudTrail for monitoring RDS performance metrics, AWS CloudWatch for logging API actions, and AWS Config for tracking compliance.

  3. C

    AWS CloudWatch for compliance checks, AWS Config for tracking RDS performance metrics, and AWS CloudTrail for logging API actions.

  4. D

    AWS Config for monitoring RDS performance metrics, AWS CloudTrail for logging API actions, and AWS CloudWatch for compliance checks.

Xem giải thích

Đáp án

A — CloudWatch cho chỉ số hiệu năng RDS, CloudTrail cho nhật ký lời gọi API

Vì sao đúng

Đúng phân vai mà AWS thiết kế: CloudWatch thu chỉ số của RDS như CPU, kết nối, độ trễ đọc ghi và cho đặt cảnh báo — phần "theo dõi hiệu năng gần thời gian thực". CloudTrail ghi ai đã gọi lời gọi nào lên instance, phần truy nguyên. Còn Config giữ lịch sử thay đổi cấu hình để chứng minh tuân thủ.

Vì sao các phương án khác sai

  • B — đảo vai CloudTrail và CloudWatch; CloudTrail không thu chỉ số hiệu năng.
  • C — Config không đo hiệu năng, còn CloudWatch không phải công cụ kiểm tra tuân thủ.
  • D — Config không thu chỉ số hiệu năng.

Nhớ nhanh

CloudWatch đo máy chạy thế nào, CloudTrail ghi ai đã làm gì, Config giữ cấu hình đã đổi ra sao.

Câu 246 Data Store Management

In AWS OpenSearch, which of the following best describes the role of replica shards?

  1. A

    Replica shards can only store data in cold storage for cost-effective redundancy.

  2. B

    Replica shards are backups of primary shards and improve fault tolerance by handling read requests.

  3. C

    Replica shards store the original document data, while primary shards store indexed data for fast searching.

  4. D

    Replica shards handle both read and write operations, providing redundancy and load balancing for queries.

Xem giải thích

Đáp án

B — Replica là bản sao của primary shard, tăng khả năng chịu lỗi và gánh thêm truy vấn đọc

Vì sao đúng

Replica là bản sao đầy đủ của một primary shard và luôn nằm trên node khác. Nhờ đó nó phục vụ hai việc: mất node chứa primary thì OpenSearch đưa một replica lên thay, và trong lúc bình thường thì truy vấn đọc được chia cho cả primary lẫn replica nên thông lượng đọc tăng theo số bản sao.

Vì sao các phương án khác sai

  • A. Replica chỉ nằm ở cold storage — sai; replica nằm cùng tầng với primary của nó.
  • C. Replica giữ tài liệu gốc còn primary giữ dữ liệu đã đánh chỉ mục — sai hoàn toàn, cả hai chứa cùng một thứ.
  • D. Replica xử lý cả đọc lẫn ghi — sai: ghi luôn đi qua primary trước rồi mới nhân bản sang replica; replica chỉ phục vụ đọc.
Câu 247 Chọn nhiều đáp án Data Ingestion and Transformation

A data engineer needs to catalog semi-structured data stored in Amazon S3 for querying using AWS Athena. The data includes multiple file formats like CSV and Parquet. The team also wants to ensure that only new data added since the last scan is cataloged. Which combination of AWS Glue features should they use to meet this requirement? (Choose Two)

  1. A

    Bookmarks for Incremental Crawling

  2. B

    Schema Inference

  3. C

    AWS Glue Jobs

  4. D

    AWS Glue Crawler

  5. E

    AWS Glue DataBrew

Xem giải thích

Đáp án

A và D — Glue Crawler, và bookmark để quét tăng dần

Vì sao đúng

  • D. Crawler quét tệp trên S3, tự nhận ra lược đồ của cả CSV lẫn Parquet, rồi ghi bảng vào Glue Data Catalog để Athena truy vấn được.
  • A. Bookmark cho quét tăng dần giải quyết vế "dữ liệu mới thêm vào": lần chạy sau crawler chỉ nhìn phần chưa quét thay vì duyệt lại toàn bộ bucket, nên vừa nhanh vừa rẻ khi hồ dữ liệu lớn dần.

Vì sao các phương án khác sai

  • B. Schema Inference — mô tả việc crawler làm chứ không phải một tính năng riêng để bật.
  • C. Glue Jobs — biến đổi dữ liệu, không phải cataloging.
  • E. DataBrew — làm sạch dữ liệu bằng giao diện trực quan, không đăng ký bảng vào catalog.
Câu 248 Data Store Management

A healthcare company needs to maintain sensitive patient data in an Amazon Redshift cluster. To comply with HIPAA regulations, they must ensure that this data is encrypted both in transit and at rest. They also need to restrict access to the cluster to only authorized users within their organization. The company also needs to optimize costs as they scale their data warehouse operations. Which of the following configurations will best meet their security and cost optimization requirements?

  1. A

    Enable Amazon Redshift Enhanced VPC Routing, use AWS KMS for encryption, and resize the cluster automatically using Amazon Redshift RA3 nodes.

  2. B

    Use SSL encryption for data in transit, encrypt data at rest with an AWS-managed key, and enable Auto Pause to automatically stop the cluster when not in use.

  3. C

    Use Amazon Redshift Spectrum to query encrypted S3 data, enable multi-factor authentication (MFA) for all users, and resize the cluster manually as needed.

  4. D

    Enable Amazon S3 Server-Side Encryption (SSE) for the storage of data, use AWS CloudTrail for logging access, and resize the cluster based on reserved instance pricing.

Xem giải thích

Đáp án

A — Bật Enhanced VPC Routing và dùng AWS KMS để mã hoá

Vì sao đúng

Hai yêu cầu HIPAA trong đề được giải bằng hai cơ chế:

  • Enhanced VPC Routing ép mọi lưu lượng giữa cụm và S3 đi qua VPC của bạn thay vì ra Internet, nên áp được security group, VPC endpoint và ghi lại bằng VPC Flow Logs.
  • AWS KMS mã hoá khi lưu bằng khoá bạn kiểm soát, và mọi lần dùng khoá đều vào CloudTrail — đúng thứ hồ sơ kiểm toán y tế đòi.

Vì sao các phương án khác sai

  • B — SSL cho đường truyền thì đúng, nhưng dùng khoá do AWS quản lý thì bạn không sửa được key policy và không kiểm soát vòng đời khoá, yếu hơn hẳn cho yêu cầu tuân thủ.
  • C — Spectrum là công cụ truy vấn, còn MFA bảo vệ việc đăng nhập; cả hai không phải cơ chế mã hoá dữ liệu trong cụm.
  • D — mã hoá bucket S3 không mã hoá dữ liệu nằm trong cụm Redshift.
Câu 249 Data Ingestion and Transformation

A data engineer needs to execute SQL queries on an Amazon Redshift cluster from a serverless AWS Lambda function. The queries must be run asynchronously without maintaining a persistent connection to the Redshift database. Which of the following solutions would be the most efficient for this use case?

  1. A

    Use a JDBC connection between the Lambda function and the Redshift cluster to execute queries.

  2. B

    Use the Amazon Redshift Data API to run queries against the Redshift cluster asynchronously.

  3. C

    Use Amazon Redshift Query Editor to manually run queries from the Lambda function.

  4. D

    Use Amazon Redshift Spectrum to run the queries directly on data stored in Amazon S3.

Xem giải thích

Đáp án

B — Dùng Amazon Redshift Data API

Vì sao đúng

Data API nhận truy vấn qua HTTP rồi trả về ngay một mã định danh, không giữ kết nối nào. Đó là điều kiện đề nêu, và cũng là cách đúng để Lambda nói chuyện với Redshift: hàm gửi truy vấn rồi kết thúc, không phải nằm chờ và không tốn thời gian chạy. Xác thực bằng IAM nên không phải nhúng mật khẩu vào hàm.

Vì sao các phương án khác sai

  • A. Kết nối JDBC — Lambda co giãn theo số lời gọi, mỗi bản sao lại mở một kết nối nên rất dễ làm cạn số kết nối của cụm; lại phải giữ hàm chạy trong lúc chờ truy vấn.
  • C. Query Editor — công cụ trên giao diện web cho người dùng, không gọi được từ mã.
  • D. Redshift Spectrum — truy vấn dữ liệu trên S3, không giải quyết chuyện kết nối.
Câu 250 Data Security and Governance

A company has adopted a multi-account AWS strategy and uses AWS Organizations to manage access between accounts. The company wants to ensure that administrators across various accounts have the necessary permissions to manage Amazon EC2 instances, while adhering to the principle of least privilege. Which AWS service should the company use to define and enforce permissions at the organizational level?

  1. A

    AWS IAM Access Analyzer

  2. B

    AWS Service Control Policies (SCPs)

  3. C

    AWS Systems Manager

  4. D

    AWS Identity and Access Management (IAM) Roles

Xem giải thích

Đáp án

B — AWS Service Control Policies (SCP)

Vì sao đúng

SCP đặt trần quyền cho cả tài khoản hoặc cả nhánh trong AWS Organizations. Dù quản trị viên của tài khoản con có tự cấp cho mình chính sách IAM rộng tới đâu, họ cũng không vượt được trần đó. Đây là công cụ duy nhất trong danh sách áp được từ tổ chức xuống nhiều tài khoản cùng lúc.

Vì sao các phương án khác sai

  • A. IAM Access Analyzer — chỉ ra quyền thừa và chia sẻ ra ngoài, nhưng không cưỡng chế gì.
  • C. Systems Manager — vận hành máy chủ, không liên quan tới phân quyền tài khoản.
  • D. IAM Roles — cấp quyền trong một tài khoản; chính quản trị viên tài khoản đó lại sửa được, nên không đặt được trần từ trên xuống.

Nhớ nhanh

SCP không cấp quyền cho ai, nó chỉ giới hạn quyền tối đa mà tài khoản có thể có.