Ngân hàng đề — AWS Certified Data Engineer Associate

Tìm thấy 867 câu.

Câu 211 Data Ingestion and Transformation

A data engineer is using AWS Glue to process a large volume of streaming data in near-real-time. To optimize the ETL job for both cost and efficiency, the engineer needs to ensure that only new data is processed incrementally, without reprocessing previously ingested data. Which AWS Glue feature should the engineer enable?

  1. A

    AWS Glue Bookmarks

  2. B

    AWS Glue Crawlers

  3. C

    Amazon Kinesis Data Streams

  4. D

    AWS Glue Job Triggers

Xem giải thích

Đáp án

A — AWS Glue Bookmarks

Vì sao đúng

Bookmark là dấu vị trí Glue tự lưu sau mỗi lần job chạy xong, ghi lại đã xử lý tới đâu. Lần chạy kế tiếp nó bắt đầu từ đúng chỗ đó nên chỉ đọc dữ liệu mới. Không có bookmark thì mỗi lần chạy là quét lại từ đầu, vừa lâu vừa tốn — đúng hai thứ đề muốn tránh.

Vì sao các phương án khác sai

  • B. Glue Crawlers — suy ra lược đồ và cập nhật catalog, không theo dõi đã xử lý tới đâu.
  • C. Kinesis Data Streams — dịch vụ nhận luồng dữ liệu, nằm trước Glue trong đường ống.
  • D. Glue Job Triggers — quyết định khi nào job chạy, không quyết định job đọc từ đâu.
Câu 212 Data Operations and Support

A company has implemented a decoupled architecture where multiple services communicate via message queues. They use Amazon SQS to handle message queues and Amazon EventBridge to route specific application events to different AWS services. The company also needs to implement a feature where certain events from EventBridge should trigger a notification to various team members using Amazon SNS. Which of the following is the correct way to implement this feature?

  1. A

    Use an Amazon SQS queue with a redrive policy to send failed events to SNS for notifications.

  2. B

    Configure an Amazon SQS queue as an EventBridge target and use Lambda to forward messages to Amazon SNS.

  3. C

    Create an EventBridge rule to forward events directly to Amazon SNS, which will send notifications to the subscribed team members.

  4. D

    Create an SNS topic for each EventBridge rule and configure SNS to poll EventBridge for new events.

Xem giải thích

Đáp án

C — Tạo quy tắc EventBridge chuyển sự kiện thẳng tới SNS

Vì sao đúng

SNS là đích đến có sẵn của EventBridge. Khai quy tắc khớp mẫu sự kiện rồi đặt SNS topic làm target là xong — không hàm nào ở giữa, không gì phải vận hành, không thêm độ trễ. Đây là cách ít mảnh ghép nhất, nên cũng ít chỗ hỏng nhất.

Vì sao các phương án khác sai

  • A. Hàng đợi SQS với redrive gửi sang SNS — redrive đưa thông điệp hỏng sang hàng đợi chết, không phải sang SNS; hiểu sai cơ chế.
  • B. Qua SQS rồi Lambda chuyển tiếp — thêm hai thành phần để làm việc mà EventBridge đã làm sẵn.
  • D. SNS hỏi EventBridge lấy sự kiện — sai chiều: EventBridge đẩy sự kiện đi, SNS không đi hỏi ai cả.
Câu 213 Data Ingestion and Transformation

A data engineer needs to query data from multiple sources, including an Amazon RDS PostgreSQL database and an on-premises PostgreSQL database. The engineer wants to avoid duplicating data into the Amazon Redshift cluster and minimize data transmission. Which solution should the data engineer use?

  1. A

    Use AWS Glue to run ETL jobs that extract, transform, and load the data into Redshift tables.

  2. B

    Use Amazon Athena to query the external databases and join the results with Redshift tables.

  3. C

    Use Redshift Spectrum to query the data directly from the external PostgreSQL databases.

  4. D

    Set up Redshift Federated Queries and create external schema definitions pointing to the PostgreSQL databases.

Xem giải thích

Đáp án

D — Dùng Redshift Federated Query và khai external schema trỏ tới PostgreSQL

Vì sao đúng

Federated Query cho Redshift đọc trực tiếp từ RDS hoặc Aurora PostgreSQL ngay trong câu truy vấn, ghép được với bảng nội bộ. Redshift đẩy phần lọc xuống tận CSDL nguồn nên chỉ kéo về những dòng cần thiết — thoả cả hai điều kiện "không nhân bản dữ liệu" và "ít truyền dữ liệu".

Vì sao các phương án khác sai

  • A. Glue ETL nạp vào Redshift — chính là nhân bản dữ liệu, thứ đề cấm.
  • B. Athena truy vấn CSDL ngoài — Athena đọc dữ liệu trên S3; muốn nối tới CSDL quan hệ thì phải dựng thêm connector và cũng không ghép trực tiếp với bảng Redshift.
  • C. Redshift Spectrum — chỉ đọc dữ liệu trên S3, không nối tới PostgreSQL.
Câu 214 Data Store Management

A company needs to deploy Amazon Redshift in a multi-AZ setup to ensure higher availability and fault tolerance for their data warehouse workloads. They also want to separate compute and storage scaling to handle varying query loads and data growth. Which node type should they use for this setup?

  1. A

    RA3 nodes

  2. B

    T3 nodes

  3. C

    R5 nodes

  4. D

    DC2 nodes

Xem giải thích

Đáp án

A — Node RA3

Vì sao đúng

Chỉ RA3 đáp ứng cả hai yêu cầu: đây là loại node duy nhất tách tính toán khỏi lưu trữ — dữ liệu nằm trên lớp managed storage dùng S3, còn node chỉ lo tính toán, nên hai thứ mở rộng độc lập. Và cụm Redshift Multi-AZ chỉ hỗ trợ RA3.

Vì sao các phương án khác sai

  • B. T3 và C. R5 — là họ máy EC2, Redshift không có loại node nào tên như vậy.
  • D. DC2 — lưu trữ nằm ngay trên SSD của node nên dung lượng gắn chặt với số node, và không dùng được cho cụm Multi-AZ.
Câu 215 Data Security and Governance

A healthcare company is using DynamoDB to store patient records and requires a secure, consistent way to track changes to these records for compliance auditing. The data changes should include both the old and new values of the records before and after modification. Additionally, these changes must be sent to an Elasticsearch cluster to enable full-text search capabilities for the patient records. Which solution will meet these requirements?

  1. A

    Enable DynamoDB Streams with the "Old Image" option. Use AWS Lambda to process the changes and store the old records in Amazon RDS for auditing.

  2. B

    Enable DynamoDB Streams with the "New and Old Image" option. Use AWS Lambda to process the stream data and send the changes to an Elasticsearch domain.

  3. C

    Enable DynamoDB Streams with the "Keys Only" option. Use AWS Lambda to process the stream data and send it to an Amazon S3 bucket for archiving. Use Amazon S3 Select to query the changes.

  4. D

    Enable DynamoDB Streams with the "New Image" option. Use Amazon Kinesis Data Streams to capture the changes and send them to an Elasticsearch domain.

Xem giải thích

Đáp án

B — Bật DynamoDB Streams kiểu "New and Old Image", dùng Lambda xử lý

Vì sao đúng

Đề nói rõ cần cả giá trị trước và sau mỗi thay đổi cho mục đích tuân thủ. Chỉ NEW_AND_OLD_IMAGES mang đủ hai phần. Lambda gắn thẳng vào stream nên mọi thay đổi được xử lý ngay và ghi sang nơi lưu kiểm toán, không cần dịch vụ trung gian nào.

Vì sao các phương án khác sai

  • A. Old Image — chỉ có trạng thái trước, không biết đã đổi thành gì.
  • C. Keys Only — chỉ có khoá chính; giá trị đã đổi thì mất hẳn, không truy lại được nữa.
  • D. New Image dùng Kinesis Data Streams — thiếu giá trị cũ, và thêm một dịch vụ trung gian không cần thiết.
Câu 216 Data Security and Governance

A DevOps team wants to set up an alerting mechanism for their production environment. They need to be notified when specific AWS CloudTrail events, such as unauthorized API requests or actions that modify IAM policies, occur. Which combination of services should be used to meet this requirement?

  1. A

    AWS Config and Amazon CloudWatch Logs

  2. B

    AWS Config and AWS Lambda

  3. C

    AWS CloudTrail and Amazon CloudWatch Alarms

  4. D

    AWS CloudTrail and AWS Lambda

Xem giải thích

Đáp án

C — AWS CloudTrail cộng Amazon CloudWatch Alarms

Vì sao đúng

CloudTrail ghi lại mọi lời gọi API, kể cả những lời gọi bị từ chối vì thiếu quyền và những lời gọi sửa chính sách IAM. Đẩy nhật ký đó sang CloudWatch Logs, đặt metric filter khớp với mẫu cần cảnh giác, rồi gắn CloudWatch Alarm lên chỉ số đó. Đây là khuôn mẫu AWS khuyến nghị sẵn cho việc cảnh báo bảo mật.

Vì sao các phương án khác sai

  • A. Config và CloudWatch Logs — Config không ghi lời gọi API, và chỉ có Logs thì không có phần bắn cảnh báo.
  • B. Config và Lambda — thiếu luôn nguồn dữ liệu về lời gọi API.
  • D. CloudTrail và Lambda — Lambda chạy được mã phản ứng, nhưng đề hỏi cơ chế cảnh báo, và tự viết lại thứ Alarm đã làm sẵn là thừa.
Câu 217 Data Operations and Support

A data engineer is building a data pipeline using Kinesis Data Streams for real-time data ingestion. The system needs to scale automatically based on the incoming data volume. What capacity mode should the engineer configure to meet this requirement?

  1. A

    On-demand mode

  2. B

    Scheduled scaling

  3. C

    Auto-scaling mode

  4. D

    Provisioned mode

Xem giải thích

Đáp án

A — On-demand mode

Vì sao đúng

On-demand là chế độ duy nhất của Kinesis Data Streams tự điều chỉnh số shard theo lưu lượng thật, chịu được mức tăng gấp đôi so với đỉnh của 30 ngày trước mà không ai phải can thiệp. Bạn trả tiền theo lượng dữ liệu đi qua.

Vì sao các phương án khác sai

  • B. Scheduled scaling và C. Auto-scaling mode — Kinesis Data Streams không có hai chế độ mang tên này; chỉ có on-demand và provisioned.
  • D. Provisioned mode — bạn tự khai số shard và tự đổi khi cần, đúng thứ đề muốn tránh.
Câu 218 Chọn nhiều đáp án Data Security and Governance

A data engineer is tasked with implementing a Time to Live (TTL) feature for DynamoDB tables to automatically delete expired items after a set period. However, the engineer needs to ensure that these deletions do not incur additional costs for write throughput and that updates to TTL-queued items are possible before deletion. Which of the following statements about TTL are correct? (Choose TWO.)

  1. A

    TTL must be manually triggered through the API once an item’s expiration time has passed.

  2. B

    TTL deletions consume write capacity units (WCUs) from the table’s provisioned throughput.

  3. C

    Once TTL is enabled, the deletions occur immediately after the expiration time is reached.

  4. D

    Deletions due to TTL can be captured in DynamoDB Streams for downstream processing.

  5. E

    Items queued for TTL deletion can still be updated before the deletion occurs.

Xem giải thích

Đáp án

D và E

Vì sao đúng

  • D. Xoá do TTL vẫn hiện trong DynamoDB Streams — bản ghi kèm cờ đánh dấu đây là xoá của hệ thống, nên vẫn lưu trữ hoặc xử lý tiếp được trước khi mục biến mất hẳn.
  • E. Mục đang chờ xoá vẫn cập nhật được — TTL không xoá đúng khoảnh khắc hết hạn; trong khoảng chờ đó bạn vẫn ghi đè được, kể cả đẩy thời hạn ra xa để giữ mục lại.

Vì sao các phương án khác sai

  • A. Phải gọi API để kích hoạt xoá — sai, TTL chạy nền hoàn toàn tự động.
  • B. Xoá do TTL tiêu tốn đơn vị ghi — sai, và đây chính là điều đề quan tâm: TTL không tính phí ghi, đó là lý do người ta dùng nó thay vì tự xoá.
  • C. Xoá ngay khi hết hạn — sai; AWS chỉ cam kết thường trong vòng vài ngày, nên truy vấn vẫn phải tự lọc mục đã hết hạn.
Câu 219 Data Operations and Support

A company is using AWS Glue for an ETL job that runs on Apache Spark. The job is set to the default DPU (Data Processing Unit) configuration, but the team notices that the job is using more computing power than necessary, leading to higher costs. What can the team do to minimize costs without affecting job performance?

  1. A

    Increase the number of DPUs to speed up the job completion.

  2. B

    Switch the ETL job to a Python Shell job to reduce the DPU requirement.

  3. C

    Use AWS Glue DataBrew to handle ETL transformations.

  4. D

    Reduce the number of DPUs to the minimum required for the job.

Xem giải thích

Đáp án

D — Giảm số DPU xuống mức tối thiểu mà job thật sự cần

Vì sao đúng

Glue tính tiền theo DPU nhân với thời gian chạy. Đề nói job đang dùng nhiều tài nguyên hơn mức cần, nghĩa là phần lớn DPU nằm không mà vẫn tính tiền. Cách chữa đúng là nhìn vào chỉ số sử dụng thật của job rồi hạ số DPU xuống sát nhu cầu. Job có thể chạy lâu hơn một chút nhưng tổng chi phí giảm.

Vì sao các phương án khác sai

  • A. Tăng DPU — làm nặng thêm đúng vấn đề đề đang than phiền.
  • B. Đổi sang Python Shell job — chỉ hợp với việc nhẹ chạy trên một máy; job Spark xử lý dữ liệu lớn thì không chuyển sang được.
  • C. Đổi sang DataBrew — công cụ trực quan cho người không viết mã, không phải cách tối ưu một job Spark đã có.
Câu 220 Data Ingestion and Transformation

A data engineer is building a data pipeline that includes transforming large datasets stored in Amazon S3 using AWS Glue. The engineer wants to automate the deployment of this pipeline, including the S3 bucket, Glue jobs, and triggers that start the job when new data arrives. How can the engineer accomplish this using AWS best practices?

  1. A

    Manually create the S3 bucket and Glue jobs through the AWS Management Console, then configure a CloudWatch Event to trigger the Glue job.

  2. B

    Create an AWS CloudFormation template that defines the S3 bucket, Glue jobs, and Lambda triggers. Deploy the template using the AWS CLI.

  3. C

    Define the S3 bucket, Glue jobs, and Lambda triggers in an AWS SAM template. Use the AWS SAM CLI to deploy and manage the application.

  4. D

    Write a Python script to create the S3 bucket, Glue jobs, and Lambda triggers using the Boto3 SDK. Run the script manually.

Xem giải thích

Đáp án

C — Khai bucket S3, Glue job và trigger Lambda trong một mẫu AWS SAM rồi triển khai bằng SAM CLI

Vì sao đúng

Đề cần hạ tầng dưới dạng mã cộng khả năng triển khai lặp lại. SAM là phần mở rộng của CloudFormation cho ứng dụng không máy chủ: khai một AWS::Serverless::Function là có sẵn hàm Lambda, vai IAM và trigger từ sự kiện S3. Tài nguyên nào SAM không có kiểu riêng — như Glue job — thì viết thẳng cú pháp CloudFormation trong cùng mẫu.

Vì sao các phương án khác sai

  • A. Bấm tay trên Console — không lặp lại được, không đưa vào kiểm soát phiên bản.
  • B. Mẫu CloudFormation thuần — làm được, và SAM chạy trên chính nó. Ở đây SAM sát hơn vì phần trigger Lambda từ S3 là chỗ nó viết gọn nhất và có SAM CLI chạy thử tại máy.
  • D. Script Boto3 — bạn phải tự lo trạng thái, tự xử lý chạy lại và tự dọn khi hỏng — đó đúng là những việc CloudFormation sinh ra để làm hộ.