Ngân hàng đề — AWS Certified Data Engineer Associate

Tìm thấy 867 câu.

Câu 321 Data Ingestion and Transformation

A data engineer wants to prepare a large dataset for a machine learning model using Amazon SageMaker. The data contains missing values and inconsistent formatting. The engineer prefers a visual interface to clean and transform the data before training the model. Which SageMaker tool should the engineer use?

  1. A

    SageMaker Data Wrangler

  2. B

    SageMaker Experiments

  3. C

    SageMaker Feature Store

  4. D

    SageMaker Autopilot

Xem giải thích

Đáp án

A — SageMaker Data Wrangler

Vì sao đúng

Data Wrangler dựng cho đúng khâu chuẩn bị dữ liệu bằng giao diện: nối tới nguồn, xem ngay phân bố và tỷ lệ ô trống, rồi chọn trong hơn 300 phép biến đổi có sẵn để điền giá trị thiếu và chuẩn hoá định dạng. Luồng biến đổi xuất được thành pipeline chạy lại, nên việc làm một lần biến thành quy trình.

Vì sao các phương án khác sai

  • B. Experiments — theo dõi và so sánh các lần huấn luyện, không xử lý dữ liệu.
  • C. Feature Store — nơi lưu và phục vụ đặc trưng sau khi đã tính xong.
  • D. Autopilot — tự dựng mô hình từ dữ liệu; nó có bước tiền xử lý nhưng làm ẩn bên trong, không phải công cụ để kỹ sư chủ động làm sạch.
Câu 322 Data Ingestion and Transformation

A data engineer is setting up a data lake using AWS Lake Formation. The data lake will ingest data from Amazon S3 and an on-premises database. The engineer needs to ensure the data is searchable across AWS services, cataloged with relevant metadata, and securely accessed by various departments within the organization. Which combination of features in AWS Lake Formation should the engineer use?

  1. A

    AWS IAM Policies for access control, AWS Glue ETL for data ingestion, and AWS Resource Access Manager (RAM) for sharing the catalog across AWS accounts.

  2. B

    AWS Glue Data Catalog for organizing metadata, Lake Formation Tag-Based Access Control (TBAC) for fine-grained security, and Lake Formation Blueprints for automating data ingestion from both S3 and on-premises databases.

  3. C

    Amazon S3 for storage, AWS KMS for data encryption, and AWS Athena for querying the data lake.

  4. D

    AWS Glue Crawlers for cataloging, Lake Formation Blueprints for data ingestion, and AWS IAM Roles for department-based access control.

Xem giải thích

Đáp án

B — Glue Data Catalog cho siêu dữ liệu, cộng Tag-Based Access Control của Lake Formation

Vì sao đúng

Hai mảnh này khớp đúng hai yêu cầu của đề. Data Catalog là nơi đăng ký bảng và lược đồ, nhờ đó dữ liệu tìm kiếm được và Athena hay Redshift Spectrum truy vấn được. TBAC gắn nhãn lên cơ sở dữ liệu, bảng và cột rồi cấp quyền theo nhãn — nên phân quyền không phải viết lại mỗi khi thêm bảng mới, và làm được tới mức dòng và cột.

Vì sao các phương án khác sai

  • A. Chính sách IAM để phân quyền — IAM cấp quyền tới mức tài nguyên, không hiểu khái niệm dòng và cột trong bảng.
  • C. S3, KMS và Athena — lưu trữ, mã hoá và truy vấn; thiếu hẳn phần catalog lẫn phân quyền chi tiết.
  • D. Crawler cộng Blueprints — đều là công cụ hợp lệ của Lake Formation nhưng đây là phần nạp và cataloging, thiếu phần kiểm soát truy cập mà đề nhấn mạnh.
Câu 323 Chọn nhiều đáp án Data Security and Governance

A data engineer at an e-commerce company is configuring an Amazon Redshift cluster to store customer transaction data. The company requires the data to be encrypted at rest and secured from unauthorized network access. Which combination of the following actions should the data engineer take to meet these requirements? (Choose TWO.)

  1. A

    Set up a VPC endpoint to restrict access to the cluster from within the company’s on-premises network.

  2. B

    Associate the Redshift cluster with a security group that allows inbound traffic from specific IP addresses.

  3. C

    Use IAM users only to control access to the Redshift cluster.

  4. D

    Deploy the Redshift cluster in a publicly accessible subnet.

  5. E

    Enable Amazon Redshift encryption using AWS Key Management Service (KMS) or customer-managed keys (CMKs).

Xem giải thích

Đáp án

B và E — gắn security group cho phép đúng nguồn truy cập, và bật mã hoá bằng KMS

Vì sao đúng

Đề đòi hai lớp và mỗi phương án lo một lớp:

  • E. Mã hoá bằng KMS lo dữ liệu khi lưu. Dùng khoá do khách hàng quản lý thì kiểm soát được key policy, bật xoay vòng, và mọi lần dùng khoá đều vào CloudTrail.
  • B. Security group lo lớp mạng: chỉ nguồn được liệt kê mới mở được kết nối tới cổng của cụm, mọi nơi khác bị chặn ngay ở tầng dưới cùng.

Vì sao các phương án khác sai

  • A. VPC endpoint chỉ cho mạng nội bộ công ty — endpoint hữu ích, nhưng mô tả này lẫn giữa VPC endpoint và Direct Connect; nó không phải cơ chế giới hạn nguồn truy cập tới cụm.
  • C. Chỉ dùng IAM user — quản danh tính chứ không mã hoá và không chặn đường mạng.
  • D. Đặt cụm ở subnet công khai — làm yếu bảo mật đi, ngược hẳn yêu cầu.
Câu 324 Data Ingestion and Transformation

A company is migrating their on-premises data warehouse to Amazon Redshift. They need to ensure minimal downtime during the migration and plan to execute an ongoing replication of the data from their on-premises source to Redshift. The data is updated frequently in the source system, and the company requires near-real-time updates in Redshift for analytics. Which service or feature combination should the data engineer use to meet this requirement?

  1. A

    AWS DataSync for ongoing replication and Amazon Redshift Spectrum to query data.

  2. B

    AWS Database Migration Service (DMS) in full load mode only.

  3. C

    AWS Glue for ETL jobs and Redshift Spectrum for real-time queries.

  4. D

    AWS Database Migration Service (DMS) in change data capture (CDC) mode and Redshift COPY command.

Xem giải thích

Đáp án

D — Dùng AWS DMS ở chế độ change data capture (CDC)

Vì sao đúng

Từ khoá quyết định là ngừng hoạt động tối thiểu. DMS chạy hai pha: nạp toàn bộ dữ liệu hiện có, rồi chuyển sang CDC để liên tục áp mọi thay đổi phát sinh từ hệ thống cũ sang Redshift. Nhờ vậy hai bên gần như đồng bộ, và lúc chuyển đổi chỉ cần dừng ghi vài phút thay vì cả cửa sổ bảo trì dài.

Vì sao các phương án khác sai

  • A. DataSync — đồng bộ tệp giữa các nơi lưu trữ, không hiểu cấu trúc bảng của CSDL.
  • B. DMS chỉ chạy full load — nạp xong là dừng, mọi thay đổi sau đó bị bỏ lỡ, nên phải khoá hệ thống cũ suốt quá trình.
  • C. Glue ETL cộng Spectrum — chạy theo mẻ, luôn có độ trễ và phải tự viết phần bắt thay đổi.
Câu 325 Data Ingestion and Transformation

A retail company uses AWS Glue DataBrew to clean and prepare customer data for analytics. The data is stored in various formats, including CSV and JSON, in Amazon S3. The company wants to ensure consistent formatting across the datasets by removing null values, standardizing column names, and converting dates into a uniform format. What is the best approach to achieve this with AWS Glue DataBrew?

  1. A

    Use Glue Data Catalog to create new metadata definitions

  2. B

    Use DataBrew Recipes to apply transformation steps

  3. C

    Use Glue Crawlers to infer the schema of the data

  4. D

    Use AWS Glue ETL Jobs to perform data transformations

Xem giải thích

Đáp án

B — Dùng DataBrew Recipe để áp các bước biến đổi

Vì sao đúng

Recipe là chuỗi các bước làm sạch được ghi lại thành một thực thể dùng lại được: áp cho bộ dữ liệu khác, gắn vào job chạy theo lịch, và có phiên bản để quay lại bản trước. Đó chính là thứ biến việc làm sạch một lần thành quy trình lặp lại — và nó xử lý được cả CSV lẫn JSON.

Vì sao các phương án khác sai

  • A. Tạo định nghĩa siêu dữ liệu trong Data Catalog — mô tả dữ liệu, không biến đổi gì.
  • C. Glue Crawler — suy ra lược đồ, cũng không biến đổi.
  • D. Glue ETL Jobs — làm được nhưng phải viết mã Spark; đề nói rõ nhóm đang dùng DataBrew, tức là muốn cách không viết mã.
Câu 326 Data Ingestion and Transformation

A financial company wants to enable its business users to perform interactive data analysis without needing in-depth knowledge of querying languages. The users should be able to ask questions in natural language and receive visual responses. Which feature of Amazon QuickSight should the company utilize?

  1. A

    Redshift Spectrum

  2. B

    QuickSight Q

  3. C

    Dashboards

  4. D

    SPICE Engine

Xem giải thích

Đáp án

B — QuickSight Q

Vì sao đúng

Q cho người dùng gõ câu hỏi bằng ngôn ngữ thường rồi tự dựng biểu đồ trả lời — đúng yêu cầu "không cần biết ngôn ngữ truy vấn". Nó dựa trên một topic do quản trị viên chuẩn bị: đặt tên thân thiện cho các trường, khai từ đồng nghĩa, chỉ rõ đâu là số đo đâu là chiều.

Vì sao các phương án khác sai

  • A. Redshift Spectrum — truy vấn dữ liệu trên S3 bằng SQL, tức là vẫn phải biết SQL.
  • C. Dashboards — báo cáo dựng sẵn; người xem chỉ lọc trong khuôn có sẵn, không hỏi câu mới được.
  • D. SPICE — bộ nhớ đệm giúp truy vấn nhanh; chuyện tốc độ, không phải giao diện hỏi.
Câu 327 Data Ingestion and Transformation

A company is planning to migrate several petabytes of data from its on-premises data center to AWS. The company has limited internet bandwidth, and the data transfer needs to be performed securely. The company is also concerned about potential damage to hardware during transportation due to harsh environmental conditions. Which solution from the AWS Snow Family would best suit this scenario?

  1. A

    AWS Snowcone

  2. B

    AWS Snowball Edge (Storage Optimized)

  3. C

    AWS Snowmobile

  4. D

    AWS DataSync

Xem giải thích

Đáp án

C — AWS Snowmobile

Vì sao đúng

Quy mô đề nêu là vài petabyte trở lên với băng thông hạn chế. Snowmobile là container 45 foot kéo bằng xe tải, chở tới 100 petabyte một chuyến. Dữ liệu được mã hoá bằng khoá KMS của bạn, có giám sát và hộ tống trên đường.

Vì sao các phương án khác sai

  • A. Snowcone — chỉ vài terabyte, sai quy mô cả nghìn lần.
  • B. Snowball Edge Storage Optimized — vài chục terabyte mỗi thiết bị; vài petabyte thì cần hàng trăm chiếc, tốn công điều phối tới mức không thực tế.
  • D. DataSync — chuyển qua mạng, mà băng thông chính là thứ đang thiếu.
Câu 328 Chọn nhiều đáp án Data Operations and Support

A company wants to implement an efficient backup strategy for their EC2 instances' EBS volumes. They need a solution that minimizes storage costs while ensuring data durability. Which TWO solutions should the company implement? (Choose Two)

  1. A

    Use Amazon S3 Glacier to store long-term EBS snapshots.

  2. B

    Use multi-attach for the EBS volumes to ensure durability.

  3. C

    Enable automated EBS snapshot lifecycle policies for retention and deletion.

  4. D

    Store EBS snapshots in a separate EC2 instance for redundancy.

  5. E

    Implement incremental snapshots to minimize storage costs.

Xem giải thích

Đáp án

C và E — bật chính sách vòng đời tự động cho snapshot, và dùng snapshot tăng dần

Vì sao đúng

  • E. Snapshot tăng dần — EBS snapshot vốn chỉ lưu khối đã thay đổi kể từ lần chụp trước, nên chi phí lưu trữ thấp hơn nhiều so với chụp toàn phần mỗi lần. Đây là phần "giảm chi phí".
  • C. Chính sách vòng đời (Data Lifecycle Manager) — tự chụp theo lịch và tự xoá bản quá hạn. Không có nó thì snapshot cứ dồn lại mãi và chi phí lớn dần, hoặc ai đó phải nhớ dọn tay.

Vì sao các phương án khác sai

  • A. Đưa snapshot vào S3 Glacier — snapshot đã nằm trên S3 do AWS quản lý; có tầng lưu trữ riêng cho snapshot nhưng không phải bằng cách tự đẩy sang Glacier.
  • B. Multi-attach — cho một volume gắn vào nhiều instance trong cùng một AZ; đó là chuyện chia sẻ, không phải sao lưu.
  • D. Cất snapshot trong một EC2 khác — không phải cách snapshot hoạt động, và một instance thì kém bền hơn S3 rất nhiều.
Câu 329 Data Operations and Support

A company has multiple microservices deployed in AWS, where each microservice must process events sent from different producers. The processing needs to happen asynchronously, and some events are required to be delivered to multiple subscribers, while others should be delivered to a single processing service. The solution must ensure fault-tolerant and highly available communication between these services. Which combination of AWS services should the company use to meet these requirements?

  1. A

    Use Amazon SNS to broadcast events to multiple subscribers, and Amazon SQS for single-service event processing.

  2. B

    Use Amazon SQS to broadcast events to multiple subscribers and Amazon EventBridge for single-service event processing.

  3. C

    Use Amazon SNS to broadcast events to multiple subscribers and AWS Step Functions for single-service event processing.

  4. D

    Use Amazon EventBridge for broadcasting events to multiple subscribers and Amazon SNS for single-service event processing.

Xem giải thích

Đáp án

A — SNS để phát tới nhiều bên đăng ký, SQS cho phần xử lý đơn lẻ

Vì sao đúng

Đây là khuôn mẫu fan-out kinh điển và mỗi dịch vụ đúng vai của nó: SNS là mô hình xuất bản – đăng ký, một thông điệp tới mọi bên quan tâm, hợp với "nhiều microservice cùng cần một sự kiện". SQS là hàng đợi, mỗi thông điệp được đúng một consumer xử lý, có thử lại và có hàng đợi chết. Ghép SNS trước SQS thì mỗi dịch vụ có hàng đợi đệm riêng, xử lý theo nhịp của mình.

Vì sao các phương án khác sai

  • B. SQS để phát tới nhiều bên — sai bản chất: thông điệp trong SQS bị một consumer lấy đi, các bên khác không thấy nữa.
  • C. Step Functions cho phần xử lý đơn lẻ — Step Functions điều phối quy trình nhiều bước, không phải hàng đợi.
  • D. EventBridge phát còn SNS xử lý đơn lẻ — EventBridge phát tin được, nhưng SNS không phải hàng đợi cho xử lý đơn lẻ; hai vai bị đảo.
Câu 330 Data Ingestion and Transformation

A social media platform stores user activity logs in DynamoDB, where data is being written at a high velocity, and rapid retrieval of recent user actions is critical. The activity logs need to be queried by user ID to monitor individual user behavior over time. Which design pattern would optimize both write throughput and read performance for this use case?

  1. A

    Use a composite key with the user ID as the partition key and the timestamp of the action as the sort key.

  2. B

    Use a simple partition key based on the user ID.

  3. C

    Use a partition key based on action type and a sort key based on user ID.

  4. D

    Use a global secondary index with the action type as the partition key and the timestamp as the sort key.

Xem giải thích

Đáp án

A — Khoá tổng hợp: user ID làm partition key, timestamp làm sort key

Vì sao đúng

Kiểu truy cập của đề là "lấy nhanh các hành động gần đây của một người". Lấy user ID làm partition key thì mọi bản ghi của người đó nằm chung một phân vùng, gom bằng một lời gọi Query. Timestamp làm sort key khiến chúng tự sắp theo thời gian, nên lấy N hành động mới nhất chỉ là đọc ngược từ cuối rồi dừng — không quét, không sắp lại.

Vì sao các phương án khác sai

  • B. Chỉ partition key là user ID — khoá chính phải duy nhất, nên mỗi người chỉ giữ được một bản ghi; bản sau đè bản trước.
  • C. Partition key là loại hành động — số loại hành động ít nên phân vùng bị dồn nóng, và vẫn không gom được theo người.
  • D. GSI theo loại hành động — mở đường tra cứu khác, không giải quyết kiểu truy cập chính.