Ngân hàng đề — AWS Certified Data Engineer Associate
Tìm thấy 867 câu.
A data engineering team needs to process large volumes of transaction logs stored in Amazon S3 and load the transformed data into an Amazon Aurora PostgreSQL database for further analysis. The team wants to automate the transformation and loading process while ensuring minimal impact on the Aurora database's performance during peak hours. Which of the following solutions will best achieve these objectives?
-
A
Create an AWS Lambda function that reads and transforms the S3 data. Configure the function to write the data directly to Aurora PostgreSQL using a batch process.
-
B
Use AWS Glue to transform the S3 data and store the transformed data temporarily in Amazon S3. Then, use Amazon Aurora’s native integration with Amazon S3 to load the data.
-
C
Use AWS Glue to run an ETL job that transforms the S3 data and writes it directly into Aurora using the JDBC connection.
-
D
Use Amazon EMR to process and transform the S3 data. Then, schedule a batch process to load the transformed data into Amazon Aurora using Amazon RDS Data API.
Xem giải thích
Đáp án
B — Dùng Glue biến đổi dữ liệu rồi tạm lưu trên S3 trước khi nạp vào Aurora
Vì sao đúng
Với khối lượng lớn, ghi thẳng từng dòng vào CSDL quan hệ là chỗ nghẽn: mỗi lần chèn là một vòng gọi qua mạng, và Aurora phải ghi nhật ký giao dịch cho từng dòng. Cách chuẩn là biến đổi xong thì ghi ra tệp trên S3, rồi dùng lệnh nạp hàng loạt của PostgreSQL để đưa vào một lần — nhanh hơn nhiều bậc và không giữ giao dịch mở quá lâu.
Vì sao các phương án khác sai
- A. Lambda đọc và biến đổi — trần 15 phút và trần bộ nhớ không kham nổi "khối lượng lớn".
- C. Glue ghi thẳng vào Aurora — chạy được, và đây là phương án dễ nhầm nhất; nhưng chính là kiểu ghi từng dòng chậm nêu ở trên.
- D. EMR rồi nạp theo lô — kết quả tương tự nhưng phải dựng và vận hành cụm, nặng hơn Glue.
A data engineer is designing a DynamoDB table for an e-commerce application. The primary use case involves accessing product information based on category and manufacturer. The product's unique identifier is the product ID, and queries need to support listing all products within a category and retrieving products from a specific manufacturer. What is the best approach to design the table for efficient querying?
-
A
Use a Global Secondary Index (GSI) with product ID as the partition key and category as the sort key.
-
B
Use a Local Secondary Index (LSI) with product ID as the partition key and category as the sort key.
-
C
Use a Global Secondary Index (GSI) with category as the partition key and manufacturer as the sort key.
-
D
Use a Local Secondary Index (LSI) with category as the partition key and manufacturer as the sort key.
Xem giải thích
Đáp án
C — Global Secondary Index với category làm partition key và manufacturer làm sort key
Vì sao đúng
Kiểu truy cập chính là tra theo danh mục và nhà sản xuất, còn khoá chính của bảng lại là mã sản phẩm. Chỉ GSI mới cho đổi partition key sang thuộc tính khác. Đặt category làm partition key và manufacturer làm sort key thì cả hai kiểu truy vấn — lấy hết một danh mục, hoặc lọc thêm theo nhà sản xuất trong danh mục đó — đều thành thao tác Query rẻ.
Vì sao các phương án khác sai
- A. GSI với product ID làm partition key — chép lại đúng khoá đã có, không mở thêm đường tra cứu.
- B và D. Dùng LSI — LSI bắt buộc dùng chung partition key với bảng gốc, nên vẫn phải biết product ID trước; ngoài ra LSI chỉ tạo được ngay lúc tạo bảng.
A company is using Amazon S3 to store large datasets for long-term archival and compliance purposes. To reduce costs, they want to transition this data to a more cost-effective storage class while maintaining accessibility for compliance audits. The data will be accessed infrequently but must be retrieved within hours if necessary. Which S3 storage class is the most appropriate for this use case?
-
A
S3 Intelligent-Tiering
-
B
S3 Glacier
-
C
S3 Glacier Deep Archive
-
D
S3 Standard
Xem giải thích
Đáp án
B — S3 Glacier
Vì sao đúng
Đề cần lưu trữ lâu dài và tuân thủ với chi phí thấp, nhưng vẫn phải lấy ra được trong thời gian hợp lý. Glacier rẻ hơn S3 Standard nhiều lần mà thời gian khôi phục tính bằng phút nếu chọn chế độ nhanh, hoặc vài giờ với chế độ thường. Đó là cân bằng đúng cho dữ liệu lưu trữ có thể bị yêu cầu xuất trình khi kiểm toán.
Vì sao các phương án khác sai
- A. Intelligent-Tiering — hợp khi không đoán được kiểu truy cập; ở đây đã biết chắc là ít đụng tới, nên trả thêm phí theo dõi mà chẳng được gì.
- C. Glacier Deep Archive — rẻ nhất nhưng thời gian lấy ra tới 12 giờ; chỉ hợp khi gần như chắc chắn không bao giờ cần gấp.
- D. S3 Standard — chính là tầng đắt mà công ty đang muốn rời khỏi.
An e-commerce platform uses Amazon Kinesis to detect fraudulent transactions in real-time. The data must be processed immediately as it is received, and the platform has multiple consumer applications that need access to the data. What should the data engineer use to ensure each consumer has dedicated read throughput while minimizing latency?
-
A
Kinesis Firehose
-
B
AWS Lambda with Synchronous Invocation
-
C
Enhanced Fan-Out
-
D
Standard Consumer with Polling Interval Optimization
Xem giải thích
Đáp án
C — Enhanced Fan-Out
Vì sao đúng
Hai điều kiện của đề — xử lý ngay khi nhận được và có nhiều ứng dụng tiêu thụ — chính là bài toán enhanced fan-out giải. Mỗi consumer được cấp đường riêng 2 MB/giây trên mỗi shard và dữ liệu được đẩy sang thay vì phải hỏi liên tục, nên độ trễ xuống khoảng 70 mili giây và các ứng dụng không giành băng thông của nhau.
Vì sao các phương án khác sai
- A. Kinesis Firehose — gom lô trước khi giao nên tối thiểu cũng trễ vài chục giây.
- B. Lambda gọi đồng bộ — cách gọi hàm, không giải quyết chuyện nhiều consumer chia nhau băng thông của shard.
- D. Consumer thường có tối ưu chu kỳ hỏi — vẫn là mô hình hỏi và vẫn chia sẻ 2 MB/giây; hỏi dày hơn chỉ tốn thêm lời gọi.
A financial institution needs to process large datasets containing personally identifiable information (PII), such as customer names, social security numbers, and credit card details. They want to automatically identify and protect this sensitive data before loading it into their data warehouse. Which AWS Glue transformation should they use to achieve this?
-
A
Filter Transformation
-
B
Change File Format to Parquet
-
C
Detect PII
-
D
Find Matches Transformation
Xem giải thích
Đáp án
C — Detect PII
Vì sao đúng
Detect PII của AWS Glue quét dữ liệu và nhận ra các loại thông tin cá nhân bằng bộ nhận dạng dựng sẵn — tên, số bảo hiểm xã hội, số thẻ tín dụng đều nằm trong đó. Sau khi phát hiện, bạn chọn cách xử lý ngay trong job: che, băm, hoặc chỉ đánh dấu để rà lại.
Vì sao các phương án khác sai
- A. Filter Transformation — lọc dòng theo điều kiện bạn tự viết; muốn dùng thì phải đã biết cột nào nhạy cảm, tức là chưa giải quyết phần khó.
- B. Đổi định dạng sang Parquet — chuyện hiệu năng và dung lượng, không liên quan tới bảo vệ dữ liệu.
- D. Find Matches — tìm bản ghi trùng lặp giữa các tập dữ liệu, việc hoàn toàn khác.
A company uses AWS Lambda to process streaming data from IoT sensors. Each Lambda invocation processes batches of 100 records from the stream, and the function is invoked whenever enough data is collected. The processed data is then sent to an Amazon S3 bucket. The company wants to ensure that Lambda function invocations remain stateless and that each invocation is independent. Which of the following statements best describes the behavior of AWS Lambda in this scenario?
-
A
AWS Lambda functions are stateless, but state can be managed by using services like Amazon S3 or DynamoDB.
-
B
AWS Lambda functions maintain state across invocations.
-
C
AWS Lambda requires manual intervention to scale in response to an increased data load.
-
D
AWS Lambda functions automatically store state between invocations.
Xem giải thích
Đáp án
A — Lambda không giữ trạng thái, muốn có trạng thái thì dùng dịch vụ bên ngoài
Vì sao đúng
Mỗi lần gọi Lambda là một môi trường thực thi có thể hoàn toàn mới, và AWS không đảm bảo hai lần gọi liên tiếp dùng chung một bản sao. Vì vậy không được dựa vào biến toàn cục để nhớ gì giữa các lần gọi. Cần trạng thái thì đưa ra ngoài: DynamoDB cho bộ đếm và checkpoint, ElastiCache cho dữ liệu nóng, S3 cho kết quả trung gian.
Lưu ý thực tế: môi trường có thể được dùng lại, và người ta khai thác điều đó để giữ kết nối CSDL — nhưng đó là tối ưu, không phải thứ được phép dựa vào cho tính đúng đắn.
Vì sao các phương án khác sai
- B và D. Lambda giữ trạng thái giữa các lần gọi — sai; đây chính là hiểu nhầm phổ biến nhất về Lambda.
- C. Phải can thiệp tay để co giãn — sai; Lambda tự tăng số bản sao theo tải, với Kinesis thì theo số shard.
A data engineer is using DynamoDB with DAX enabled for a high-throughput application. However, during peak times, they notice some requests are being throttled. What is the most likely cause of this throttling, and how can it be mitigated?
-
A
The DAX cache is full; clear the cache to improve performance.
-
B
The DAX cluster has reached its request capacity; add more nodes to the DAX cluster.
-
C
The DynamoDB table’s provisioned throughput is too low; enable auto-scaling for the table.
-
D
The application is making too many write operations; increase the write capacity of the DynamoDB table.
Xem giải thích
Đáp án
B — Cụm DAX đã chạm trần năng lực; thêm node vào cụm
Vì sao đúng
Đề nói rõ DAX đang bật và việc chặn xảy ra vào giờ cao điểm. Mỗi node DAX có trần số yêu cầu mỗi giây tuỳ loại máy; vượt trần đó thì chính DAX trả lỗi chặn, dù bảng DynamoDB phía sau vẫn thừa năng lực. Cách chữa là thêm node để chia tải đọc, hoặc đổi sang loại node lớn hơn.
Vì sao các phương án khác sai
- A. Bộ đệm đầy nên phải xoá — DAX tự loại mục cũ khi đầy, đầy bộ đệm không gây lỗi chặn.
- C. Năng lực bảng quá thấp — có thể xảy ra, nhưng khi DAX đang phục vụ phần lớn lượt đọc từ bộ đệm thì bảng ít khi là nút thắt; đề cũng chỉ nói tới đọc.
- D. Quá nhiều thao tác ghi — DAX ghi xuyên qua xuống bảng, nhưng đề nói vấn đề nằm ở lượt đọc lúc cao điểm.
A media production company needs to transfer several terabytes of high-resolution video footage from a remote location with limited bandwidth to Amazon S3. They also want to process the videos locally for compression before transferring them to reduce data size. Which AWS service and device combination is the most suitable?
-
A
AWS Snowmobile
-
B
AWS DataSync with a direct S3 transfer
-
C
AWS Snowcone with AWS DataSync
-
D
AWS Snowball Edge (Compute Optimized) with local processing
Xem giải thích
Đáp án
D — AWS Snowball Edge (Compute Optimized) kèm xử lý tại chỗ
Vì sao đúng
Đề có ba ràng buộc: vài terabyte dữ liệu, băng thông hạn chế, và cần xử lý video ngay tại chỗ. Snowball Edge Compute Optimized đáp ứng cả ba: dung lượng đủ cho cỡ terabyte, chạy được EC2 và Lambda ngay trên thiết bị để xử lý trước, và chuyển đi bằng cách gửi thiết bị nên không phụ thuộc đường truyền.
Vì sao các phương án khác sai
- A. Snowmobile — dành cho quy mô hàng chục petabyte, quá lớn, và không có năng lực điện toán biên.
- B. DataSync qua mạng — chính băng thông là thứ đang thiếu.
- C. Snowcone kèm DataSync — Snowcone chỉ vài terabyte nhưng năng lực tính toán rất hạn chế, không kham nổi xử lý video; phần DataSync lại quay về phụ thuộc mạng.
A data engineering team is building a real-time data processing pipeline to analyze clickstream data from their e-commerce website. The data needs to be ingested quickly, stored for long-term analysis, and occasionally queried for real-time insights. The team wants to optimize both the storage cost and the performance of the pipeline. Which TWO strategies should the team implement to achieve these goals? (Choose TWO.)
-
A
Use Amazon S3 Intelligent-Tiering to store the clickstream data, automatically optimizing for cost as access patterns change.
-
B
Use Amazon Kinesis Data Streams for real-time ingestion and Amazon S3 Glacier for long-term storage of processed data.
-
C
Use Amazon EBS General Purpose SSD (gp3) volumes for storing clickstream data to provide low-latency and high-throughput access.
-
D
Use Amazon Kinesis Data Firehose to ingest data into Amazon S3, and then query the data with Amazon Athena for real-time analysis.
-
E
Use Amazon S3 Standard for storing the data and configure Amazon S3 Lifecycle policies to transition data to S3 Glacier Deep Archive for cost savings.
Xem giải thích
Đáp án
A và D — Firehose nạp vào S3 rồi truy vấn tại chỗ, và dùng S3 Intelligent-Tiering
Vì sao đúng
Ba yêu cầu của đề — nạp nhanh, lưu lâu dài, truy vấn được — được giải bằng hai mảnh:
- D. Firehose vào S3 — không máy chủ, tự co giãn theo lưu lượng clickstream, tự gom lô và nén; sau đó Athena truy vấn thẳng trên S3 mà không phải nạp đi đâu nữa.
- A. Intelligent-Tiering — dữ liệu clickstream có kiểu truy cập khó đoán (dữ liệu mới hay dùng, cũ thì thưa dần). Tầng này tự chuyển theo thực tế sử dụng và không tính phí lấy ra.
Vì sao các phương án khác sai
- B. Glacier cho lưu trữ lâu dài — rẻ nhưng phải khôi phục trước khi đọc, nên "truy vấn được" không còn đúng.
- C. Ổ EBS gp3 — ổ khối gắn với một máy, không phải nơi lưu cho hồ dữ liệu.
- E. S3 Standard kèm lifecycle — làm được, nhưng bạn phải tự đoán và tự khai ngưỡng chuyển tầng; Intelligent-Tiering làm việc đó theo dữ liệu thật.
A social networking company is building a recommendation engine based on its user relationship data, which follows a graph structure. The company has chosen Amazon Neptune for storing and querying this data. The system must handle frequent updates as users add friends, likes, and followers. The company also needs to perform real-time queries such as “find mutual friends” or “suggest friends of friends.” What is the most effective approach to scale Amazon Neptune for this use case?
-
A
Scale Neptune by adding additional Read Replicas to distribute the query load.
-
B
Implement DynamoDB Streams with Neptune for managing frequent data updates.
-
C
Scale Neptune by partitioning (sharding) the graph data based on user IDs.
-
D
Use Neptune’s Multi-Master replication to scale both read and write operations.
Xem giải thích
Đáp án
A — Mở rộng Neptune bằng cách thêm Read Replica để chia tải truy vấn
Vì sao đúng
Neptune tách tính toán khỏi lưu trữ: một instance ghi chính, cùng tối đa 15 read replica dùng chung một lớp lưu trữ. Vì vậy thêm replica là thêm sức đọc gần như tức thì, không phải chép dữ liệu và không phải chia lại đồ thị. Với hệ gợi ý — nơi lượt đọc áp đảo lượt ghi — đây đúng là hướng mở rộng.
Vì sao các phương án khác sai
- B. DynamoDB Streams với Neptune — hai dịch vụ khác nhau; Streams không phải cơ chế mở rộng cho Neptune.
- C. Chia nhỏ đồ thị theo user ID — Neptune không hỗ trợ sharding, và chia đồ thị vốn là bài toán rất khó vì truy vấn hay đi xuyên qua các phần.
- D. Multi-Master để mở rộng cả ghi — Neptune không có chế độ nhiều instance ghi; chỉ có một instance ghi chính.