Ngân hàng đề — AWS Certified Data Engineer Associate

Tìm thấy 867 câu.

Câu 291 Data Security and Governance

A company is building a security monitoring system using AWS OpenSearch. They want to implement role-based access control (RBAC) to manage permissions for different teams accessing the system. What security measure should the company implement to restrict access appropriately?

  1. A

    Native authentication only with no role-based controls

  2. B

    TLS encryption for data at rest

  3. C

    Attribute-based access control (ABAC) without considering roles

  4. D

    Role-based access control (RBAC) with predefined roles and permissions

Xem giải thích

Đáp án

D — RBAC với các vai và quyền định nghĩa sẵn

Vì sao đúng

Chính đề đã nêu yêu cầu là kiểm soát truy cập theo vai, và OpenSearch có sẵn cơ chế đó: security plugin đi kèm nhiều vai dựng sẵn, cho phép cấp quyền tới mức chỉ mục, tới từng tài liệu và từng trường. Ánh xạ nhóm người dùng vào vai thì mỗi nhóm chỉ thấy phần dữ liệu của mình, và thêm người mới chỉ là gán vai.

Vì sao các phương án khác sai

  • A. Chỉ xác thực, không phân vai — xác thực trả lời "bạn là ai", không trả lời "bạn được làm gì"; mọi người sẽ có cùng quyền.
  • B. TLS cho dữ liệu khi lưu — mô tả sai: TLS bảo vệ đường truyền, còn khi lưu thì dùng mã hoã đĩa. Dù sao mã hoá cũng không phải phân quyền.
  • C. ABAC mà bỏ qua vai — kiểm soát theo thuộc tính là mô hình hợp lệ, nhưng trái thẳng với yêu cầu RBAC của đề.
Câu 292 Data Operations and Support

A media company wants to automate their data ingestion pipeline to process raw log files stored in Amazon S3, transform the data, and load it into Amazon Redshift on a daily schedule. They also want to trigger the process automatically when new data arrives. Which AWS Glue feature should they use to build and orchestrate this data pipeline?

  1. A

    Glue Data Quality

  2. B

    Glue Crawlers

  3. C

    Glue Workflows

  4. D

    Glue DataBrew

Xem giải thích

Đáp án

C — Glue Workflows

Vì sao đúng

Đường ống của đề gồm nhiều bước phải chạy đúng thứ tự: crawl log thô, biến đổi, rồi nạp vào Redshift, và lặp lại hằng ngày. Workflow gói cả chuỗi đó thành một thực thể, dùng trigger nối các bước và trigger theo lịch để chạy mỗi ngày. Khi hỏng, bạn thấy ngay bước nào gãy và chạy lại từ đúng chỗ đó.

Vì sao các phương án khác sai

  • A. Glue Data Quality — đặt luật kiểm tra chất lượng dữ liệu, không điều phối các bước.
  • B. Glue Crawlers — chỉ suy ra lược đồ; nó là một bước trong workflow chứ không phải thứ điều phối.
  • D. Glue DataBrew — làm sạch dữ liệu bằng giao diện trực quan, không dựng đường ống nhiều bước.
Câu 293 Data Store Management

A data engineer wants to improve the performance of a complex query that is frequently executed. The query takes significant time to run due to complex aggregations on large datasets stored in Redshift. Which solution provides the MOST performance improvement?

  1. A

    Create a materialized view in Redshift with auto-refresh enabled to pre-compute and store query results.

  2. B

    Use Amazon Redshift Spectrum to query data directly from S3, bypassing Redshift.

  3. C

    Create a regular view to store the query logic and run the view as needed.

  4. D

    Schedule the query to run periodically using an Amazon CloudWatch Events rule.

Xem giải thích

Đáp án

A — Tạo materialized view trong Redshift và bật tự làm mới

Vì sao đúng

Materialized view tính trước rồi lưu lại kết quả, nên truy vấn phức tạp chạy thường xuyên chỉ còn là đọc bảng đã tính sẵn. Bật tự làm mới thì Redshift tự cập nhật khi dữ liệu nguồn đổi, và với nhiều trường hợp nó chỉ cập nhật phần thay đổi chứ không tính lại từ đầu. Redshift còn tự dùng materialized view cho những truy vấn khớp, kể cả khi câu truy vấn không nhắc tới nó.

Vì sao các phương án khác sai

  • B. Spectrum truy vấn thẳng S3 — đưa dữ liệu ra khỏi cụm thường chậm hơn, và không tính trước gì cả.
  • C. View thường — chỉ là câu truy vấn được đặt tên; mỗi lần gọi vẫn chạy lại toàn bộ phép tổng hợp.
  • D. Hẹn giờ chạy truy vấn — chạy sẵn không có nghĩa kết quả được lưu lại; người dùng vẫn phải chờ khi họ mở báo cáo.
Câu 294 Chọn nhiều đáp án Data Operations and Support

A company is building a serverless image processing pipeline using AWS Lambda and Amazon S3. When an image is uploaded to an S3 bucket, it triggers a Lambda function that resizes the image into different resolutions (small, medium, and large) and stores the resized images back into the S3 bucket under different folder paths. However, the company has encountered issues with Lambda's memory limits and timeout errors when processing large image files.

To optimize this image processing pipeline while maintaining a serverless architecture, which of the following steps should the company take? (Choose THREE.)

  1. A

    Implement AWS Step Functions to orchestrate multiple Lambda invocations for handling each image resolution.

  2. B

    Use AWS Fargate instead of Lambda to handle large image files and process them outside of the serverless architecture.

  3. C

    Increase the memory allocation for the Lambda function to speed up processing.

  4. D

    Use Amazon S3 multipart uploads to split large images and process them in chunks within the Lambda function.

  5. E

    Split the Lambda function into multiple smaller functions, each handling a specific image resolution (e.g., small, medium, large).

  6. F

    Enable Amazon S3 Transfer Acceleration to reduce the upload and processing time for large images.

Xem giải thích

Đáp án

A, C và E — điều phối bằng Step Functions, tăng bộ nhớ cho hàm, và tách thành nhiều hàm nhỏ

Vì sao đúng

Cả ba đều tấn công đúng giới hạn của Lambda khi ảnh lớn dần:

  • C. Tăng bộ nhớ — trong Lambda, CPU được cấp tỷ lệ thuận với bộ nhớ. Nâng bộ nhớ là nâng luôn sức tính, nên việc đổi kích thước ảnh chạy nhanh hơn hẳn. Đây là nút chỉnh dễ nhất.
  • E. Tách thành nhiều hàm nhỏ, mỗi hàm một khâu — mỗi hàm chạy trong trần 15 phút của riêng nó, và chỉnh bộ nhớ riêng cho khâu nặng.
  • A. Step Functions — điều phối chuỗi hàm đó, tự thử lại từng bước và giữ trạng thái, thay vì nhét cả quy trình vào một lần gọi.

Vì sao các phương án khác sai

  • B. Đổi sang Fargate — giải quyết được nhưng bỏ hẳn kiến trúc không máy chủ mà đề đang xây.
  • D. Multipart upload để chia ảnh — multipart chia việc tải lên, không chia được nội dung ảnh thành phần xử lý riêng.
  • F. Transfer Acceleration — tăng tốc tải lên từ xa, không liên quan tới thời gian xử lý.
Câu 295 Data Operations and Support

A financial services company needs to automate the deployment of its Amazon EC2 instances while ensuring that they are launched with the correct configurations and security settings. The company also requires an audit trail to track all configuration changes and user activity for compliance purposes. Which combination of AWS services should the company use to meet these requirements?

  1. A

    AWS Systems Manager for deploying EC2 instances, AWS IAM for applying

  2. B

    AWS IAM for deploying EC2 instances, AWS Systems Manager for applying configuration settings, and AWS Config for tracking configuration changes.

  3. C

    AWS CloudFormation for deploying EC2 instances, AWS Config for applying configuration settings, and AWS CloudTrail for tracking user activity.

  4. D

    AWS CloudFormation for deploying EC2 instances, AWS Systems Manager for applying configuration settings, and AWS CloudTrail for tracking user activity.

Xem giải thích

Đáp án

D — CloudFormation để triển khai EC2, Systems Manager để áp cấu hình

Vì sao đúng

Hai công cụ, hai giai đoạn: CloudFormation dựng hạ tầng từ mã, nên mỗi lần triển khai đều ra đúng cùng một kết quả và có thể huỷ gọn khi hỏng. Systems Manager lo phần sau khi máy đã chạy: áp cấu hình, cài phần mềm, vá lỗi và giữ máy không trôi khỏi chuẩn.

Vì sao các phương án khác sai

  • A và B — IAM là dịch vụ phân quyền, nó không triển khai hay cấu hình máy nào.
  • C. Config để áp cấu hình — hiểu sai vai: Config quan sát và đánh giá cấu hình, nó không thực hiện thay đổi.
Câu 296 Data Operations and Support

A data engineer is tasked with setting up an Amazon Redshift cluster that can handle high-performance queries while allowing the automatic recovery of failed nodes. The company also needs to scale the cluster as demand increases. Which combination of Redshift features best addresses these requirements?

  1. A

    Amazon Redshift serverless to handle unpredictable workloads and avoid node failure.

  2. B

    Redshift managed storage with RA3 nodes for independent scaling of compute and storage.

  3. C

    Provisioned clusters with DC2 nodes and manual resizing to handle scaling.

  4. D

    Use Amazon RDS and Redshift Spectrum for automatic node recovery.

Xem giải thích

Đáp án

B — Redshift managed storage với node RA3

Vì sao đúng

RA3 tách tính toán khỏi lưu trữ: dữ liệu nằm trên lớp managed storage dùng S3, còn node chỉ lo tính toán, nên hai thứ mở rộng độc lập. Điểm quan trọng cho yêu cầu "tự khôi phục khi node hỏng": vì dữ liệu không nằm cứng trên node, Redshift thay node hỏng mà không phải chép lại dữ liệu, nên khôi phục nhanh hơn hẳn.

Vì sao các phương án khác sai

  • A. Redshift Serverless — hợp với tải thất thường, nhưng đề nói rõ cần cụm với node, và ở đây yêu cầu là hiệu năng cao ổn định.
  • C. DC2 kèm đổi cỡ bằng tay — lưu trữ gắn cứng với node nên không mở rộng độc lập, và "bằng tay" trái yêu cầu tự động.
  • D. RDS cộng Spectrum — RDS là CSDL giao dịch, không phải kho dữ liệu; ghép này không giải bài toán nào của đề.
Câu 297 Data Ingestion and Transformation

A data engineer needs to perform real-time analytics on streaming data to detect fraudulent transactions as they occur. The analytics must allow for SQL-based queries to detect specific patterns or anomalies in the data stream. Which service should be used to achieve this?

  1. A

    Amazon Kinesis Data Firehose

  2. B

    Amazon Kinesis Data Streams

  3. C

    Amazon CloudWatch Logs

  4. D

    Amazon Managed Service for Apache Flink

Xem giải thích

Đáp án

D — Amazon Managed Service for Apache Flink

Vì sao đúng

Chi tiết quyết định trong đề là truy vấn bằng SQL trên dòng đang chảy. Flink hỗ trợ Flink SQL, cho viết truy vấn có cửa sổ thời gian để bắt mẫu gian lận — ví dụ đếm số giao dịch của cùng một thẻ trong năm phút. Nó giữ được trạng thái giữa các bản ghi, thứ mà xử lý từng bản ghi rời rạc không làm được.

Vì sao các phương án khác sai

  • A. Kinesis Data Firehose — giao nhận dữ liệu tới đích, không có phần truy vấn.
  • B. Kinesis Data Streams — cái ống chứa dữ liệu; nó là nguồn cho Flink chứ tự nó không phân tích.
  • C. CloudWatch Logs — lưu và tra cứu log vận hành, không phải nền tảng xử lý luồng.
Câu 298 Data Ingestion and Transformation

A company is designing a data ingestion system for real-time analytics. They want to ensure that the system can remember which data has already been processed, so it only ingests new data during each subsequent run. Which AWS service and configuration should they choose to achieve this?

  1. A

    Use AWS Lambda with stateless data ingestion to process incoming data events.

  2. B

    Set up Amazon Kinesis Data Streams with stateful ingestion, tracking the sequence numbers of processed data.

  3. C

    Use Amazon S3 to store all ingested data and manually track which files have been processed.

  4. D

    Configure AWS Glue with stateless ingestion to always reprocess all available data.

Xem giải thích

Đáp án

B — Kinesis Data Streams với nạp có trạng thái, ghi nhớ sequence number đã xử lý

Vì sao đúng

Mỗi bản ghi trong Kinesis có một sequence number tăng dần trong shard. Lưu lại số của bản ghi cuối cùng xử lý xong — gọi là checkpoint — thì lần chạy sau đọc tiếp đúng từ vị trí đó, nên chỉ nạp dữ liệu mới. Thư viện KCL làm sẵn việc này và giữ checkpoint trong một bảng DynamoDB.

Vì sao các phương án khác sai

  • A. Lambda không trạng thái — không nhớ đã xử lý tới đâu, đúng thứ đề muốn tránh.
  • C. Lưu hết vào S3 rồi tự theo dõi bằng tay — làm thủ công, không mở rộng nổi và rất dễ sót.
  • D. Glue luôn xử lý lại toàn bộ — trái thẳng yêu cầu "chỉ nạp dữ liệu mới".
Câu 299 Data Security and Governance

A company uses AWS Lake Formation to manage its data lake, but now needs to share a subset of data with another AWS account. They want to ensure that users in the recipient account can only access specific rows and columns of the shared data. What should the data engineer do to securely share this data?

  1. A

    Use Lake Formation to create data filters for row and column-level security and share the data using AWS Resource Access Manager (RAM). Additionally, create resource links in the recipient account.

  2. B

    Use AWS Glue Data Catalog to share the data and configure AWS KMS for encryption of the shared data.

  3. C

    Use AWS IAM policies to limit access to the specific rows and columns and share the data using Amazon S3 Cross-Region Replication.

  4. D

    Share the data using Amazon Redshift Spectrum and set up VPC Peering between the two AWS accounts for secure data access.

Xem giải thích

Đáp án

A — Dùng Lake Formation tạo data filter cho quyền theo dòng và cột, rồi chia sẻ

Vì sao đúng

Data filter của Lake Formation cho khai điều kiện lọc dòng và danh sách cột được phép, rồi gắn bộ lọc đó vào quyền cấp cho tài khoản nhận. Bên nhận truy vấn bảng như bình thường nhưng chỉ thấy phần dữ liệu đã lọc — dữ liệu không phải sao chép hay cắt ra thành bản riêng, và quyền vẫn do bên sở hữu kiểm soát.

Vì sao các phương án khác sai

  • B. Glue Data Catalog cộng KMS — catalog là nơi lưu siêu dữ liệu và KMS lo mã hoá; không cái nào lọc theo dòng hay cột.
  • C. Chính sách IAM giới hạn dòng và cột — IAM cấp quyền tới mức tài nguyên, không hiểu khái niệm dòng và cột trong một bảng.
  • D. Spectrum cộng VPC Peering — công cụ truy vấn và kết nối mạng, không phải cơ chế phân quyền chi tiết.
Câu 300 Data Ingestion and Transformation

A company wants to share live data from its Amazon Redshift cluster with another department's Redshift cluster in a different AWS region. The consumer cluster will handle its own compute requirements, while the original data remains in the producer cluster. What is the most efficient way to achieve this?

  1. A

    Copy the data from the producer cluster to the consumer cluster using AWS Glue.

  2. B

    Export the data to Amazon S3 and have the consumer cluster query the data using Redshift Spectrum.

  3. C

    Create a Redshift Data Share to share the data with the consumer cluster.

  4. D

    Set up an ETL pipeline to continuously replicate the data across both clusters.

Xem giải thích

Đáp án

C — Tạo Redshift Data Share

Vì sao đúng

Data sharing cho cụm sản xuất chia sẻ dữ liệu sống sang cụm khác — kể cả khác tài khoản và khác khu vực — mà không sao chép và không di chuyển dữ liệu. Cụm tiêu thụ đọc thẳng từ nơi lưu gốc và tự trả chi phí tính toán cho truy vấn của mình, đúng điều kiện đề nêu.

Vì sao các phương án khác sai

  • A. Chép dữ liệu bằng Glue — tạo bản sao, nên dữ liệu lệch nhau ngay khi bên nguồn thay đổi.
  • B. Xuất ra S3 rồi truy vấn bằng Spectrum — thêm một chặng và dữ liệu chỉ mới tới thời điểm xuất gần nhất, không còn là "dữ liệu sống".
  • D. Đường ống ETL nhân bản liên tục — phải dựng và vận hành, luôn có độ trễ, và nhân đôi chi phí lưu trữ.