Ngân hàng đề — AWS Certified Data Engineer Associate
Tìm thấy 867 câu.
A company is using Amazon EC2 instances with Amazon EBS volumes for high-performance applications. They want to ensure maximum security for their data, both when stored on EBS volumes and while being transmitted to their EC2 instances. Additionally, the company needs to be able to manage the encryption keys for compliance purposes. Which combination of security features and services should the company implement?
-
A
Use unencrypted EBS volumes for better performance and enable VPC encryption for secure data transfer.
-
B
Use Amazon S3 for data storage with customer-managed keys (CMKs) from AWS KMS for encryption, and EC2 instance storage for temporary data.
-
C
Use EBS encryption at rest with customer-managed keys (CMKs) from AWS KMS and enable VPC endpoint encryption for data in transit.
-
D
Use EBS encryption at rest with AWS-managed keys and enable SSL for data in transit.
Xem giải thích
Đáp án
C — Mã hoá EBS khi lưu bằng khoá do khách hàng quản lý (CMK) của KMS, kèm mã hoá đường truyền
Vì sao đúng
Đề đòi bảo vệ dữ liệu ở cả hai trạng thái. Mã hoá EBS bằng CMK lo phần khi lưu, và điểm quan trọng là dùng khoá của chính bạn nên kiểm soát được key policy, bật xoay vòng khoá, và mọi lần dùng khoá đều vào CloudTrail để kiểm toán. Phần trên đường truyền do mã hoá ở tầng mạng lo. Mã hoá EBS còn bao luôn dữ liệu giữa instance và volume, cùng mọi snapshot sinh ra sau đó.
Vì sao các phương án khác sai
- A. Không mã hoá EBS cho nhanh — đánh đổi sai; chi phí hiệu năng của mã hoá EBS gần như không đo được.
- B. Chuyển sang S3 — đổi hẳn kiến trúc lưu trữ; ứng dụng cần ổ khối chứ không phải kho đối tượng.
- D. Dùng khoá do AWS quản lý — vẫn mã hoá, nhưng bạn không sửa được key policy và không kiểm soát vòng đời khoá, yếu hơn cho yêu cầu "bảo mật tối đa".
A retail company is using an Amazon DynamoDB table to store its product data. The table has a partition key based on the ProductID attribute. The company now needs to query the table by Category and sort the results by Price. However, the table’s current partition key does not allow this query pattern, and the company wants a flexible, scalable solution that can be modified later if needed. Which solution will allow this query pattern while maintaining flexibility?
-
A
Use a Global Secondary Index (GSI) with Category as the partition key and Price as the sort key.
-
B
Modify the base table’s partition key to be Category and create a GSI with ProductID as the partition key.
-
C
Use an LSI with Category as the partition key and no sort key.
-
D
Use a Local Secondary Index (LSI) with Category as the partition key and Price as the sort key.
Xem giải thích
Đáp án
A — Global Secondary Index với Category làm partition key và Price làm sort key
Vì sao đúng
Bảng gốc đã dùng ProductID làm partition key, mà DynamoDB chỉ tra cứu hiệu quả theo khoá. Muốn truy vấn theo Category thì phải có chỉ mục lấy Category làm partition key — và chỉ GSI mới cho phép đổi partition key. Đặt Price làm sort key thì lọc theo khoảng giá hoặc sắp theo giá trong mỗi danh mục trở thành thao tác đọc liên tiếp, rất rẻ.
Vì sao các phương án khác sai
- B. Đổi partition key của bảng gốc — không sửa được sau khi tạo bảng; muốn đổi phải dựng bảng mới và di chuyển toàn bộ dữ liệu.
- C và D. Dùng LSI — LSI bắt buộc dùng chung partition key với bảng gốc, tức là vẫn phải biết ProductID. Ngoài ra LSI chỉ tạo được ngay lúc tạo bảng.
A data engineering team is running a batch data processing workload on Amazon EC2 that requires high read and write performance. The workload is temporary, but during its run, it generates a large volume of intermediate data. The team also needs low-latency access to data stored locally on the instance during processing. Cost optimization is a priority. Which of the following Amazon EC2 storage options is BEST suited for this workload?
-
A
Instance Store Volumes
-
B
Amazon EFS (Elastic File System)
-
C
Amazon S3 Standard
-
D
Amazon EBS General Purpose SSD (gp3)
Xem giải thích
Đáp án
A — Instance Store Volumes
Vì sao đúng
Hai điều kiện của đề khớp đúng với instance store: cần đọc ghi rất nhanh, và dữ liệu chỉ tạm thời trong lúc job chạy. Instance store là ổ NVMe gắn thẳng vào máy vật lý nên không đi qua mạng như EBS, cho độ trễ thấp nhất và thông lượng cao nhất. Nó cũng đã nằm trong giá instance, không tính tiền riêng.
Điểm phải nhớ: dữ liệu mất khi instance dừng hoặc bị huỷ. Ở đây không sao vì đề nói rõ khối lượng công việc là tạm thời.
Vì sao các phương án khác sai
- B. EFS — hệ tệp chia sẻ qua mạng, độ trễ cao hơn hẳn ổ cục bộ.
- C. S3 Standard — kho đối tượng, không gắn thành ổ đĩa cho tính toán nóng.
- D. gp3 — bền và tiện, nhưng vẫn là ổ qua mạng và phải trả thêm tiền cho thứ không cần giữ.
A company requires a highly scalable database solution for a real-time analytics system that handles massive traffic and diverse data formats. The data needs to be written and read with millisecond latency, but complex queries and joins are not required. Which database service is most suitable for this use case?
-
A
Amazon Redshift
-
B
Amazon Aurora
-
C
Amazon RDS for PostgreSQL
-
D
Amazon DynamoDB
Xem giải thích
Đáp án
D — Amazon DynamoDB
Vì sao đúng
Ba yêu cầu của đề — lưu lượng rất lớn, dữ liệu đa dạng không cùng cấu trúc, và đọc ghi ở mức mili giây một chữ số — đều là điểm mạnh cốt lõi của DynamoDB. Nó chia dữ liệu qua nhiều phân vùng nên mở rộng theo chiều ngang gần như không giới hạn, mỗi mục chỉ bắt buộc có khoá chính nên chứa được nhiều dạng dữ liệu khác nhau trong cùng một bảng.
Vì sao các phương án khác sai
- A. Redshift — kho dữ liệu phân tích, tối ưu cho truy vấn quét lớn chứ không phải đọc ghi từng bản ghi ở mili giây.
- B. Aurora và C. RDS for PostgreSQL — CSDL quan hệ cần lược đồ cố định, và mở rộng ghi chủ yếu bằng cách nâng cấu hình máy nên có trần.
A company wants to automate their data processing pipeline using AWS Glue. The pipeline involves several AWS Glue jobs that need to be executed sequentially. The pipeline includes tasks like data extraction from Amazon S3, data transformation, and data loading into Amazon Redshift. The company also needs to ensure that if any of the tasks fail, the workflow should be able to detect the failure and stop the pipeline. Which feature of AWS Glue should the company use to build and manage this data pipeline?
-
A
AWS Lambda
-
B
AWS Glue Workflows
-
C
AWS Glue Triggers
-
D
AWS Glue Crawler
Xem giải thích
Đáp án
B — AWS Glue Workflows
Vì sao đúng
Workflow là thứ gói nhiều job và crawler thành một quy trình có thứ tự, với trigger nối các bước lại: job này xong mới chạy job kia, hoặc chờ nhiều nhánh cùng xong. Bạn theo dõi cả quy trình như một thực thể, thấy ngay bước nào hỏng và chạy lại từ đó.
Vì sao các phương án khác sai
- A. AWS Lambda — tự viết phần điều phối, phải tự lo trạng thái và thử lại; nếu cần điều phối phức tạp thì Step Functions mới là công cụ đúng, không phải Lambda trần.
- C. Glue Triggers — là mảnh ghép bên trong workflow, dùng riêng thì chỉ nối được từng cặp bước chứ không có cái nhìn tổng thể.
- D. Glue Crawler — suy ra lược đồ, không điều phối gì cả.
An AWS Glue ETL job is processing data stored in an S3 bucket. The data engineer wants to ensure that only new files added to the bucket since the last ETL run are processed. How can the engineer configure this to avoid reprocessing previously loaded files?
-
A
Configure AWS Glue to reprocess all files each time the job is triggered.
-
B
Set up a stateless ingestion mechanism to process all data again.
-
C
Use AWS Lambda to monitor the S3 bucket and trigger the ETL job whenever a new file is added.
-
D
Use AWS Glue Bookmarks to enable incremental data processing.
Xem giải thích
Đáp án
D — Dùng AWS Glue Bookmarks để xử lý tăng dần
Vì sao đúng
Bookmark là dấu vị trí Glue tự lưu sau mỗi lần job chạy xong. Lần chạy kế tiếp nó so với dấu đó và chỉ đọc tệp mới thêm vào. Đây là tính năng dựng sẵn cho đúng bài toán, chỉ cần bật cho job và cấp quyền là xong.
Vì sao các phương án khác sai
- A. Xử lý lại toàn bộ mỗi lần — chính là điều đề muốn tránh; càng ngày càng lâu và càng đắt.
- B. Cơ chế không trạng thái — không nhớ đã xử lý tới đâu, nên vẫn phải đọc lại tất cả.
- C. Lambda theo dõi bucket rồi kích hoạt job — giải quyết chuyện khi nào chạy, nhưng job chạy xong vẫn quét lại toàn bộ nếu không có bookmark. Hai việc khác nhau.
A retail company has product catalogs from two different sources: its own internal system and a competitor's database. The data is not structured identically, and there are no common unique identifiers such as primary keys. The company needs to identify matching or duplicate products between these catalogs. Which AWS Glue transformation should be used to solve this problem?
-
A
Detect PII
-
B
Find Matches
-
C
Data Catalog Update
-
D
Change File Format
Xem giải thích
Đáp án
B — Find Matches
Vì sao đúng
Đề mô tả đúng bài toán mà Find Matches sinh ra để giải: hai nguồn dữ liệu không có khoá chung và cấu trúc khác nhau, cần biết bản ghi nào chỉ cùng một sản phẩm. Find Matches dùng học máy — bạn gán nhãn một ít cặp mẫu để dạy nó, rồi nó tự tìm các cặp còn lại dựa trên độ tương đồng chứ không so khớp chính xác.
Vì sao các phương án khác sai
- A. Detect PII — nhận diện thông tin cá nhân, việc hoàn toàn khác.
- C. Data Catalog Update — cập nhật siêu dữ liệu vào catalog.
- D. Change File Format — đổi định dạng tệp, ví dụ CSV sang Parquet.
A retail company uses Amazon Athena to query customer transaction logs stored in Amazon S3 for both ad-hoc analysis and report generation. The data engineer wants to isolate queries run by different teams (e.g., business analysts and data scientists) and control the cost for each team while monitoring their query activity. Additionally, the data scientist team uses Apache Spark for advanced analytics. What is the BEST approach to achieve this?
-
A
Enable query result caching for frequently run queries and set a spending limit on each IAM user.
-
B
Create separate Amazon Athena workgroups for each team and set up Spark-enabled workgroups for the data scientists.
-
C
Set up Amazon EMR clusters for both teams and schedule queries via Amazon CloudWatch Events.
-
D
Use AWS Glue to partition data into different buckets for each team and control access with IAM roles.
Xem giải thích
Đáp án
B — Tạo workgroup Athena riêng cho từng nhóm
Vì sao đúng
Workgroup là đơn vị cô lập của Athena. Mỗi workgroup có nơi lưu kết quả riêng, chỉ số và lịch sử truy vấn riêng, và quan trọng nhất là đặt được hạn mức dữ liệu quét riêng cho từng nhóm. Nhờ vậy nhóm phân tích tuỳ hứng không thể vô tình quét hết ngân sách của nhóm làm báo cáo, và bạn biết chính xác chi phí thuộc về ai.
Vì sao các phương án khác sai
- A. Bộ đệm kết quả kèm hạn mức — hạn mức thì đúng hướng nhưng không có cơ chế đặt riêng cho từng nhóm nếu không dùng workgroup; bộ đệm chỉ giúp truy vấn lặp lại.
- C. Dựng cụm EMR cho cả hai nhóm — bỏ hẳn Athena để đổi lấy hạ tầng phải vận hành.
- D. Chia dữ liệu ra nhiều bucket — nhân bản dữ liệu và vẫn không tách được chi phí truy vấn.
A retail company uses DynamoDB to store product information and wants to optimize query performance for different access patterns. The company needs to retrieve data by both product category and manufacturer in addition to the product ID. The data must remain consistent across all queries. Which of the following combinations of secondary indexes should the company implement to achieve this requirement?
-
A
Local Secondary Index (LSI) with productID as the partition key and manufacturer as the sort key; GSI with category as the partition key.
-
B
Local Secondary Index (LSI) with productID as the partition key and category as the sort key; Global Secondary Index (GSI) with manufacturer as the partition key.
-
C
Two Global Secondary Indexes (GSIs), both using productID as the partition key but with different sort keys for category and manufacturer.
-
D
Global Secondary Index (GSI) with category as the partition key and another GSI with manufacturer as the partition key.
Xem giải thích
Đáp án
D — Một GSI lấy category làm partition key, và một GSI khác lấy manufacturer
Vì sao đúng
Đây là hai kiểu truy cập độc lập nhau, nên cần hai chỉ mục và mỗi cái lấy đúng thuộc tính cần tra làm partition key. Chỉ GSI mới cho phép đổi partition key khác với bảng gốc. Nhờ vậy truy vấn theo danh mục và truy vấn theo nhà sản xuất đều là thao tác Query rẻ, không phải quét bảng.
Vì sao các phương án khác sai
- A và B. LSI với productID làm partition key — LSI bắt buộc dùng chung partition key với bảng gốc, nên vẫn phải biết productID trước; không giải quyết được gì.
- C. Hai GSI đều lấy productID làm partition key — chép lại đúng kiểu truy cập đã có, không mở thêm đường tra cứu nào.
A financial institution stores logs in Amazon S3. They need to automate the extraction and transformation of this data to perform hourly updates and load it into Amazon Redshift for analytics. The logs are stored in CSV format, and the institution wants to convert them to a columnar format to optimize the query performance in Redshift. How should they implement this ETL pipeline using AWS Glue?
-
A
Use a Glue ETL Job to convert the data from CSV to Apache Parquet format, and load the data directly into Amazon Redshift.
-
B
Set up AWS Lambda functions to monitor S3, convert CSV to Parquet, and upload the files into Redshift using the Redshift COPY command.
-
C
Copy the data from S3 to Amazon RDS, run transformations manually using Python scripts on EC2, and export the data back to Amazon Redshift.
-
D
Use a Glue Crawler to crawl the data in Amazon S3 and convert it to Parquet, then load it into Amazon Redshift.
Xem giải thích
Đáp án
A — Dùng Glue ETL job đổi CSV sang Apache Parquet rồi nạp vào Redshift
Vì sao đúng
Đề cần tự động, theo giờ, và Glue ETL job đáp ứng cả hai: chạy theo lịch, biến đổi bằng Spark rồi ghi vào Redshift qua connector dựng sẵn, không máy chủ nào phải trông. Bước đổi sang Parquet không thừa — định dạng cột nén tốt hơn nhiều nên vừa nhanh vừa rẻ ở khâu nạp và ở mọi truy vấn sau đó qua Spectrum.
Vì sao các phương án khác sai
- B. Lambda theo dõi rồi tự đổi định dạng — trần 15 phút và trần bộ nhớ sẽ gãy khi log lớn dần.
- C. Chép sang RDS rồi chạy script tay — thủ công, đúng thứ đề muốn tự động hoá.
- D. Crawler đổi sang Parquet — hiểu sai vai trò: crawler chỉ suy ra lược đồ, nó không biến đổi dữ liệu.