Ngân hàng đề — AWS Certified Data Engineer Associate

Tìm thấy 867 câu.

Câu 311 Chọn nhiều đáp án Data Security and Governance

A data engineering team at a healthcare company uses Amazon Athena to query sensitive patient data stored in Amazon S3. To comply with HIPAA regulations, the team must ensure that all Athena query results and data interactions are encrypted both at rest and in transit. Which combination of security settings should the team implement? (Choose TWO.)

  1. A

    Use AWS Identity and Access Management (IAM) policies to allow Athena access only to specific S3 buckets.

  2. B

    Configure Athena to encrypt query results with SSE-S3 (Server-Side Encryption with Amazon S3-managed keys).

  3. C

    Use AWS Key Management Service (KMS) to encrypt data stored in Amazon S3.

  4. D

    Enable Athena to use client-side encryption for all query results.

  5. E

    Enable Amazon S3 bucket policies to block public access and enforce encryption of objects at rest.

Xem giải thích

Đáp án

B và C — mã hoá kết quả truy vấn của Athena, và mã hoá dữ liệu trên S3 bằng KMS

Vì sao đúng

Điểm dễ bỏ sót ở đây là Athena ghi kết quả truy vấn ra một bucket S3, và kết quả đó chứa đúng dữ liệu bệnh nhân vừa truy vấn. Mã hoá dữ liệu nguồn mà quên mã hoá kết quả là để lộ một bản sao không được bảo vệ.

  • C lo dữ liệu gốc trên S3, dùng KMS nên kiểm soát và kiểm toán được việc dùng khoá.
  • B lo phần kết quả, cấu hình ngay trong workgroup của Athena.

Vì sao các phương án khác sai

  • A. Chính sách IAM giới hạn Athena — cần thiết cho phân quyền, nhưng đề hỏi về mã hoá.
  • D. Mã hoá phía máy khách cho mọi kết quả — Athena không hỗ trợ kiểu này cho kết quả truy vấn.
  • E. Chặn truy cập công khai và bắt buộc mã hoá — việc nên làm, nhưng là biện pháp phòng vệ chung chứ không phải cơ chế mã hoá mà đề hỏi.
Câu 312 Data Security and Governance

A company is running a production workload on Amazon ECS using the EC2 launch type. They store their Docker container images in Amazon ECR. To improve security, the company wants to ensure that only authorized ECS tasks can pull images from ECR. What should they do to secure the ECS-ECR interaction?

  1. A

    Use AWS Shield to protect the ECS tasks from unauthorized access to ECR.

  2. B

    Use an IAM role with permissions to pull images from ECR, and attach it to the ECS task execution role.

  3. C

    Create an Amazon S3 bucket policy to limit access to the container images.

  4. D

    Configure a VPC endpoint for ECS to directly access ECR without an internet connection.

Xem giải thích

Đáp án

B — Dùng IAM role có quyền kéo ảnh từ ECR, gắn vào task

Vì sao đúng

ECS xác thực với ECR bằng vai IAM, không bằng mật khẩu hay khoá nào nhúng trong ảnh. Cụ thể, ecsTaskExecutionRole cần quyền ecr:GetAuthorizationToken, ecr:BatchGetImage và ecr:GetDownloadUrlForLayer. Nhờ vậy thông tin đăng nhập là tạm thời, tự xoay vòng, và không bao giờ nằm trong mã hay biến môi trường.

Vì sao các phương án khác sai

  • A. AWS Shield — chống tấn công từ chối dịch vụ, không liên quan tới phân quyền.
  • C. Bucket policy của S3 — ảnh container nằm trong ECR, không phải trong bucket bạn quản.
  • D. VPC endpoint cho ECS — endpoint giúp lưu lượng không ra Internet, tốt cho bảo mật mạng nhưng không cấp quyền; thiếu vai IAM thì vẫn không kéo được ảnh.
Câu 313 Data Store Management

A data engineer is optimizing an Amazon Redshift cluster that stores several years of transaction data. The company has noticed that older data is queried infrequently but needs to be retained for compliance purposes. The newer data is queried frequently for real-time analytics. Which strategy should the engineer implement to balance query performance and cost-effectiveness?

  1. A

    Archive older data to an Amazon RDS database and use Amazon Redshift for the latest data.

  2. B

    Use Amazon S3 Glacier for older data and Amazon Redshift DS2 nodes for newer data.

  3. C

    Use Amazon Redshift RA3 nodes with managed storage, move older data to Amazon S3 using Redshift Spectrum, and keep newer data in the Redshift cluster.

  4. D

    Use Dense Compute (DC2) nodes to store both older and newer data within the Redshift cluster.

Xem giải thích

Đáp án

C — RA3 với managed storage, đưa dữ liệu cũ sang S3 và truy vấn bằng Spectrum

Vì sao đúng

Đề mô tả đúng bài toán phân tầng: dữ liệu cũ vẫn phải giữ nhưng ít khi truy vấn. RA3 tách tính toán khỏi lưu trữ nên bạn chọn số node theo nhu cầu tính toán, còn dữ liệu nguội tự chuyển xuống S3 do Redshift quản lý. Phần thật sự cũ thì đẩy hẳn ra S3 và truy vấn qua Spectrum khi cần — vẫn nằm trong cùng một câu SQL, ghép được với bảng nóng.

Vì sao các phương án khác sai

  • A. Lưu trữ sang RDS — RDS là CSDL giao dịch, không hợp cho truy vấn phân tích trên dữ liệu lớn.
  • B. Glacier cộng node DS2 — Glacier phải khôi phục trước khi đọc nên không truy vấn tại chỗ được, và DS2 là thế hệ cũ đã bị RA3 thay thế.
  • D. DC2 chứa cả cũ lẫn mới — lưu trữ gắn cứng với node, nên giữ nhiều năm dữ liệu là phải mua thêm cả CPU không dùng tới.
Câu 314 Data Security and Governance

A company needs to ensure that their S3 buckets do not have public access enabled due to strict compliance regulations. They want to continuously monitor their environment for any S3 bucket policy changes and automatically trigger an alert if public access is detected. Which combination of services would allow them to meet this requirement?

  1. A

    AWS CloudTrail and AWS Lambda

  2. B

    AWS Config and Amazon CloudWatch

  3. C

    Amazon CloudWatch Logs and AWS Lambda

  4. D

    AWS Config and AWS CloudTrail

Xem giải thích

Đáp án

B — AWS Config kết hợp Amazon CloudWatch

Vì sao đúng

Config có sẵn quy tắc quản lý s3-bucket-public-read-prohibited và s3-bucket-public-write-prohibited. Nó liên tục đánh giá mọi bucket và đánh dấu bucket nào lệch chuẩn. Khi trạng thái tuân thủ đổi, Config phát sự kiện; CloudWatch bắt sự kiện đó rồi bắn cảnh báo. Config còn giữ dòng thời gian thay đổi — đúng thứ cần khi bị kiểm toán.

Vì sao các phương án khác sai

  • A và C. CloudTrail hoặc CloudWatch Logs kèm Lambda — bắt được lúc có ai đổi, nhưng không đánh giá được trạng thái hiện tại của những bucket chưa ai đụng tới từ lâu.
  • D. Config cộng CloudTrail — Config đúng, nhưng CloudTrail chỉ ghi log; thiếu hẳn phần bắn cảnh báo mà đề yêu cầu.
Câu 315 Data Store Management

A company needs to create a serverless ETL pipeline to transform data stored in Amazon S3. The pipeline must be deployed as code, and the company prefers to use a serverless framework. Additionally, the pipeline should automatically scale based on data volume. Which combination of services and tools should the company use?

  1. A

    AWS Glue for ETL, AWS SAM for deployment, and Amazon EC2 Auto Scaling for scalability

  2. B

    AWS Glue for ETL, AWS SAM for deployment, and AWS Lambda for scalability

  3. C

    AWS Glue for ETL, AWS CloudFormation for deployment, and AWS Lambda for scalability

  4. D

    AWS Glue for ETL, AWS Elastic Beanstalk for deployment, and AWS Lambda for scalability

Xem giải thích

Đáp án

B — Glue cho ETL, SAM để triển khai, Lambda cho phần co giãn

Vì sao đúng

Cả ba mảnh đều thoả điều kiện "không máy chủ" và "triển khai dưới dạng mã": Glue chạy ETL không cần cụm; SAM khai hạ tầng thành mã với cú pháp gọn cho Lambda và trigger; Lambda lo các bước phụ như kích hoạt job hay xử lý sự kiện, tự co giãn theo số lời gọi.

Vì sao các phương án khác sai

  • A. EC2 Auto Scaling — máy chủ phải vận hành, trái thẳng yêu cầu không máy chủ.
  • C. CloudFormation để triển khai — làm được, nhưng đề nói rõ ưu tiên một framework không máy chủ, mà SAM chính là framework đó (nó chạy trên CloudFormation).
  • D. Elastic Beanstalk — nền tảng cho ứng dụng web chạy trên máy chủ, không dựng cho ETL không máy chủ.
Câu 316 Data Operations and Support

A data engineer is managing a highly available web application running on Amazon EC2 instances behind an Elastic Load Balancer (ELB). The traffic is dynamic, with high peaks during business hours and minimal traffic during off-hours. The engineer needs to optimize both cost and performance while maintaining availability. Which of the following solutions will BEST meet these requirements?

  1. A

    Use On-Demand Instances and manually stop the instances during off-hours.

  2. B

    Use Reserved Instances to handle the entire workload and save costs over time.

  3. C

    Use Spot Instances exclusively to handle all traffic and achieve cost savings.

  4. D

    Use Auto Scaling with On-Demand Instances and configure a scaling policy to add or remove instances based on CPU utilization.

Xem giải thích

Đáp án

D — Auto Scaling với On-Demand Instance và chính sách co giãn theo tải

Vì sao đúng

Lưu lượng có đỉnh trong giờ làm việc và xuống thấp ngoài giờ, nên thứ cần là năng lực bám theo nhu cầu. Auto Scaling thêm máy khi chỉ số vượt ngưỡng và thu bớt khi tải giảm, nên bạn chỉ trả tiền cho phần đang thật sự phục vụ. Nó cũng tự thay máy hỏng, tức là được luôn phần sẵn sàng cao mà đề nêu.

Vì sao các phương án khác sai

  • A. Tự tay dừng máy ngoài giờ — thủ công, không phản ứng được với đỉnh bất thường trong ngày.
  • B. Reserved Instance cho toàn bộ tải — phải cam kết theo mức đỉnh, nên trả tiền cho phần dư suốt thời gian còn lại. Reserved chỉ nên phủ phần nền.
  • C. Chỉ dùng Spot Instance — rẻ nhất nhưng có thể bị thu hồi với hai phút báo trước; dùng một mình cho ứng dụng web sản xuất là chấp nhận rủi ro mất dịch vụ.
Câu 317 Data Security and Governance

A company wants to ensure their customer data adheres to strict data quality rules before loading it into an Amazon Redshift data warehouse. The team uses AWS Glue ETL jobs to process the data and wants to validate certain rules, such as ensuring all customer email addresses follow a standard format and there are no null values in required fields. How can they implement these data quality checks in AWS Glue?

  1. A

    Use AWS Glue Data Quality to enforce data validation rules

  2. B

    Use AWS Glue Workflows to orchestrate a data quality job before loading the data

  3. C

    Use AWS Glue DataBrew to manually inspect the data before processing

  4. D

    Use AWS Glue Crawler to identify any schema mismatches

Xem giải thích

Đáp án

A — Dùng AWS Glue Data Quality để áp luật kiểm tra dữ liệu

Vì sao đúng

Glue Data Quality là thành phần sinh ra đúng cho việc này: bạn khai luật bằng ngôn ngữ DQDL (ví dụ cột nào không được rỗng, giá trị phải nằm trong khoảng nào, khoá phải duy nhất), gắn thẳng vào job ETL, và chọn hành động khi luật gãy — cho job dừng hoặc tách riêng những dòng hỏng. Nhờ vậy dữ liệu bẩn bị chặn trước khi vào Redshift.

Vì sao các phương án khác sai

  • B. Glue Workflows — điều phối thứ tự các bước; nó quyết định khi nào chạy chứ không biết luật chất lượng nào.
  • C. DataBrew để soi bằng mắt — "thủ công" là chỗ hỏng: không tự động và không chặn được gì.
  • D. Glue Crawler — chỉ phát hiện lệch lược đồ, không kiểm tra giá trị bên trong.
Câu 318 Data Ingestion and Transformation

A company uses Amazon Redshift for its data warehouse and has operational data stored in Amazon Aurora with PostgreSQL compatibility. They need to run real-time analytics without moving data into Redshift. How can the company MOST efficiently query the data in both Redshift and Aurora?

  1. A

    Set up an ETL process to periodically move the data from Aurora to Redshift and then query the combined data in Redshift.

  2. B

    Use Amazon Redshift Federated Queries to query data directly from the Aurora PostgreSQL database without moving the data.

  3. C

    Configure Amazon RDS Data API to directly pull data into Redshift for querying.

  4. D

    Export the data from Aurora into Amazon S3 and use Redshift Spectrum to query both datasets.

Xem giải thích

Đáp án

B — Dùng Redshift Federated Query truy vấn thẳng Aurora PostgreSQL

Vì sao đúng

Federated Query cho Redshift đọc trực tiếp dữ liệu đang nằm trong Aurora ngay trong câu truy vấn, ghép được với bảng trong kho. Dữ liệu không phải di chuyển nên luôn là bản mới nhất — đúng nghĩa phân tích thời gian thực. Redshift còn đẩy điều kiện lọc xuống tận Aurora nên chỉ kéo về những dòng cần.

Vì sao các phương án khác sai

  • A. ETL định kỳ — dữ liệu chỉ mới tới lần chạy gần nhất, không còn là thời gian thực, và trái yêu cầu "không di chuyển dữ liệu".
  • C. RDS Data API kéo dữ liệu vào Redshift — Data API là cách gửi câu truy vấn qua HTTP, không phải công cụ nạp dữ liệu.
  • D. Xuất ra S3 rồi dùng Spectrum — thêm một chặng và một bản sao, vẫn trễ theo nhịp xuất.
Câu 319 Data Security and Governance

A company is auditing its AWS account for compliance purposes. They need a service that allows them to review the historical activity of AWS API calls and track resource configuration changes over time to ensure they are aligned with compliance rules. Which combination of services will best meet this requirement?

  1. A

    Amazon CloudWatch and AWS Config

  2. B

    AWS Config and Amazon CloudWatch Logs

  3. C

    AWS Config and AWS CloudTrail

  4. D

    AWS CloudTrail and Amazon GuardDuty

Xem giải thích

Đáp án

C — AWS Config kết hợp AWS CloudTrail

Vì sao đúng

Đề hỏi hai thứ và mỗi dịch vụ trả lời một:

  • CloudTrail giữ lịch sử lời gọi API — ai gọi gì, lúc nào, từ địa chỉ nào. Đây là phần "hoạt động API trong quá khứ".
  • AWS Config giữ lịch sử cấu hình tài nguyên — tại một thời điểm bất kỳ, security group đó mở những cổng nào. Đây là phần "thay đổi cấu hình".

Kiểm toán cần cả hai: một cái nói ai đã làm, cái kia nói kết quả ra sao.

Vì sao các phương án khác sai

  • A và B — CloudWatch (và CloudWatch Logs) lo chỉ số và log vận hành, không giữ lịch sử lời gọi API dưới dạng kiểm toán được.
  • D. CloudTrail cộng GuardDuty — GuardDuty phát hiện mối đe doạ theo hành vi, hữu ích nhưng không giữ lịch sử cấu hình mà đề yêu cầu.
Câu 320 Data Operations and Support

A company is using an AWS Lambda function to process incoming orders from an Amazon SQS queue. However, the company notices that during peak hours, the Lambda function is unable to process messages from the queue fast enough, causing a backlog of unprocessed orders. What is the most efficient way to handle this backlog in AWS Lambda?

  1. A

    Increase the timeout setting for the Lambda function.

  2. B

    Increase the batch size for the Lambda function’s SQS trigger.

  3. C

    Implement Lambda event source scaling with Amazon Kinesis Data Streams.

  4. D

    Enable Lambda reserved concurrency to limit the number of concurrent executions.

  5. E

    Set up Amazon EC2 instances to handle the queue processing instead of AWS Lambda.

Xem giải thích

Đáp án

B — Tăng batch size cho trigger SQS của hàm Lambda

Vì sao đúng

Mỗi lần gọi Lambda đều có phần chi phí cố định: khởi tạo, kết nối, ghi log. Khi hàng đợi dồn lại vào giờ cao điểm mà mỗi lần gọi chỉ xử lý vài thông điệp thì phần chi phí đó chiếm phần lớn thời gian. Nâng batch size (tối đa 10.000 với hàng đợi chuẩn) cho mỗi lần gọi xử lý nhiều đơn hàng hơn, nên thông lượng tăng rõ mà không đổi gì khác.

Vì sao các phương án khác sai

  • A. Tăng timeout — chỉ có tác dụng nếu hàm đang bị cắt giữa chừng; đề nói tồn đọng, không nói hàm chạy quá hạn.
  • C. Chuyển sang Kinesis — thay hẳn kiến trúc để giải một vấn đề chỉnh tham số là xong.
  • D. Reserved concurrency — giới hạn số bản sao chạy song song, tức là làm chậm thêm.
  • E. Chuyển sang EC2 — bỏ kiến trúc không máy chủ và ôm việc vận hành.