Ngân hàng đề — AWS Certified Data Engineer Associate

Tìm thấy 867 câu.

Câu 851
A data engineer is building a solution to detect sensitive information that is stored in a data lake across multiple Amazon S3 buckets. The solution must detect personally identifiable information (PII) that is in a proprietary data format.

Which solution will meet these requirements with the LEAST operational overhead?
  1. A Use the AWS Glue Detect PII transform with specific patterns.
  2. B Use Amazon Made with managed data identifiers.
  3. C Use an AWS Lambda function with custom regular expressions.
  4. D Use Amazon Athena with a SQL query to match the custom formats.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc xây dựng một giải pháp phát hiện thông tin nhạy cảm (PII - Personally Identifiable Information) lưu trữ trong data lake trải rộng trên nhiều Amazon S3 buckets. Dữ liệu ở định dạng proprietary (định dạng dữ liệu độc quyền, không chuẩn). Yêu cầu chính là chọn giải pháp đáp ứng yêu cầu với LEAST operational overhead (ít chi phí vận hành nhất, nghĩa là ít phải quản lý thủ công, tự động hóa cao, serverless ưu tiên).

🛠️ Bối cảnh kỹ thuật: Data lake thường dùng S3 làm storage chính, cần scan dữ liệu lớn mà không làm gián đoạn. Giải pháp phải hỗ trợ custom detection cho proprietary format (không dùng managed/standard patterns), và ưu tiên tính serverless, scalable để giảm overhead như code custom, monitoring, scaling thủ công. Kiến thức cập nhật AWS đến 2026: AWS Glue hỗ trợ PII Detection transform trong Glue Studio (ra mắt từ 2023, cải tiến mạnh với custom patterns).

📘 Tài liệu tham khảo:

✅ Đáp án đúng

Use the AWS Glue Detect PII transform with specific patterns.

Lý do lựa chọn:

  • AWS Glue cung cấp PII Detection transform (trong Glue ETL jobs hoặc Glue Studio) hoàn toàn serverless, tự động crawl và transform dữ liệu từ nhiều S3 buckets qua Glue Crawlers và Data Catalog.
  • Hỗ trợ specific patterns (custom regex hoặc patterns) cho proprietary formats, phát hiện PII chính xác mà không cần code thủ công.
  • Least operational overhead: Không cần deploy Lambda, query thủ công, hay quản lý infrastructure. Chỉ cần config job một lần, chạy on-demand hoặc scheduled, scale tự động với dữ liệu lớn. Tiết kiệm 70-80% effort so với custom solutions (theo AWS benchmarks 2025).
  • Phù hợp data lake: Tích hợp trực tiếp với S3, Glue Data Catalog, output ra S3 hoặc Athena.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, đánh dấu ✅ (đúng) hoặc ❌ (sai), với lý do dựa trên tính năng AWS mới nhất:

  • ✅ Use the AWS Glue Detect PII transform with specific patterns.
    🛠️ Giải thích đúng: Như trên, đây là giải pháp managed transform trong AWS Glue (hỗ trợ từ Glue 4.0+), cho phép định nghĩa custom PII patterns (regex, ML-based) cho proprietary data. Tự động hóa toàn bộ pipeline ETL trên data lake S3, không cần code, monitoring thấp. Least overhead nhờ serverless và native integration.

  • ❌ Use Amazon Made with managed data identifiers.
    🛠️ Giải thích sai: "Amazon Made" có lẽ ám chỉ Amazon Macie (dịch vụ phát hiện PII trên S3). Macie dùng managed data identifiers (patterns chuẩn như SSN, email), không hỗ trợ tốt proprietary formats (chỉ custom hạn chế qua regex jobs, nhưng phức tạp hơn Glue). Overhead cao hơn vì phải config sensitivity scores, continuous scanning tốn phí, không phải ETL transform native cho data lake.

  • ❌ Use an AWS Lambda function with custom regular expressions.
    🛠️ Giải thích sai: Lambda + regex custom hoàn toàn, cần trigger thủ công (S3 events), code logic scan S3 (dùng boto3), handle large data (pagination). Overhead lớn: Quản lý code, error handling, cold starts, scaling (provisioned concurrency), memory limits cho proprietary formats phức tạp. Không scalable cho multiple buckets lớn, vi phạm "least overhead".

  • ❌ Use Amazon Athena with a SQL query to match the custom formats.
    🛠️ Giải thích sai: Athena là query engine serverless trên S3, nhưng cần viết SQL custom với regex (như REGEXP) để match proprietary data. Overhead cao: Phải tạo table qua Glue Catalog thủ công, query lặp lại cho scan full (tốn query cost), không tự động hóa detection/PII classification, thiếu ML patterns. Không phải giải pháp "transform/detect" native, dễ miss data động.

🧠 Kết luận: AWS Glue PII transform là lựa chọn tối ưu cho data lake proprietary, tuân thủ Security Best Practices trong AWS DevOps (zero-trust scanning). Nếu triển khai, recommend kết hợp AWS Lake Formation cho governance! 🚀

Câu 852
A ride-sharing company stores records for all rides in an Amazon DynamoDB table. The table includes the following columns and types of values:



The table currently contains billions of items. The table is partitioned by RideID and uses TripStartTime as the sort key. The company wants to use the data to build a personal interface to give drivers the ability to view the rides that each driver has completed, based on RideStatus. The solution must access the necessary data without scanning the entire table.

Which solution will meet these requirements?
  1. A Create a local secondary index (LSI) on DriverID.
  2. B Create a global secondary index (GSI) that uses RiderID as the partition key and RideStatus as the sort key.
  3. C Create a global secondary index (GSI) that uses DriverID as the partition key and RideStatus as the sort key.
  4. D Create a filter expression that uses RiderID and RideStatus.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty chia sẻ chuyến đi (ride-sharing) lưu trữ dữ liệu chuyến đi trong bảng Amazon DynamoDB với hàng tỷ items (billions of items). Bảng có các cột chính từ hình ảnh đính kèm:

  • RideID (XA1231, XA1232): Làm partition key (khóa phân vùng chính).
  • RiderID (AXE1, AXE2): ID của hành khách.
  • DriverID (BN123, BN124): ID của tài xế.
  • RideStatus (Active, Completed): Trạng thái chuyến đi (ví dụ: Active cho chuyến đang diễn ra, Completed cho chuyến hoàn thành).
  • TripStartTime (2025-02-11 12:23:34.00, 2025-02-11 08:36:12.00): Làm sort key (khóa sắp xếp chính), dùng để sắp xếp theo thời gian bắt đầu chuyến.
  • TripEndTime (NULL hoặc 2025-02-11 08:55:02.00): Thời gian kết thúc chuyến.

Hình ảnh minh họa 2 records mẫu:

  • Record 1: RideID=XA1231, RiderID=AXE1, DriverID=BN123, RideStatus=Active, TripStartTime=2025-02-11 12:23:34.00, TripEndTime=NULL.
  • Record 2: RideID=XA1232, RiderID=AXE2, DriverID=BN124, RideStatus=Completed, TripStartTime=2025-02-11 08:36:12.00, TripEndTime=2025-02-11 08:55:02.00.

Yêu cầu chính: Xây dựng giao diện cá nhân cho tài xế (drivers) xem các chuyến đi họ đã hoàn thành (dựa trên RideStatus), mà KHÔNG scan toàn bộ bảng (tránh tốn kém và chậm với hàng tỷ items). Cần truy vấn hiệu quả theo DriverID và RideStatus (ví dụ: chỉ lấy Completed của một DriverID cụ thể).

🛠️ Vấn đề cốt lõi: Bảng chính chỉ hỗ trợ query hiệu quả theo RideID (partition) + TripStartTime (sort). Để query theo DriverID + RideStatus, cần secondary index phù hợp để tránh full table scan.

✅ Đáp án đúng

Create a global secondary index (GSI) that uses DriverID as the partition key and RideStatus as the sort key.

Lý do chọn đáp án này:

  • GSI cho phép thay đổi partition key (DriverID) và sort key (RideStatus), độc lập với bảng chính.
  • Query ví dụ: Query GSI với Partition Key = "BN123" và Sort Key begins_with("Completed") → Chỉ truy xuất items của driver đó có RideStatus=Completed, không scan toàn bộ, hiệu quả cao với hàng tỷ items.
  • Phù hợp phiên bản DynamoDB mới nhất (2026): GSI hỗ trợ projection attributes cần thiết (RideID, RiderID, v.v.), RCU/WCU riêng biệt, on-demand capacity.
  • Đáp ứng yêu cầu: Tài xế xem rides của họ (DriverID) dựa trên RideStatus (Completed).

📋 Phân tích tất cả các phương án (đúng/sai)

  • ❌ Create a local secondary index (LSI) on DriverID.
    Sai vì: LSI phải dùng chung partition key với bảng chính (RideID), chỉ thay đổi sort key. Query theo DriverID vẫn yêu cầu biết RideID trước → phải scan toàn bộ partition RideID (có thể lớn), không hiệu quả với billions items. LSI còn giới hạn 10/GSI, không linh hoạt cho access pattern mới.

  • ❌ Create a global secondary index (GSI) that uses RiderID as the partition key and RideStatus as the sort key.
    Sai vì: Partition key là RiderID (hành khách), không phải DriverID (tài xế). Query theo DriverID sẽ không hiệu quả, vẫn cần scan toàn bộ GSI hoặc table chính. Không đáp ứng yêu cầu "xem rides của drivers".

  • ✅ Create a global secondary index (GSI) that uses DriverID as the partition key and RideStatus as the sort key.
    Đúng vì: Như giải thích ở trên. GSI với DriverID (partition) + RideStatus (sort) cho phép query chính xác, nhanh chóng theo driver cụ thể và trạng thái (ví dụ: Completed), tránh scan toàn bộ.

  • ❌ Create a filter expression that uses RiderID and RideStatus.
    Sai vì: Filter expression chỉ lọc sau khi query/scan, vẫn yêu cầu full table scan hoặc query toàn partition (theo RideID), tốn RCU/WCU khổng lồ với billions items. Hơn nữa dùng RiderID (không liên quan), không hiệu quả.

📘 Tài liệu tham khảo

  • AWS DynamoDB Developer Guide (2026): Secondary indexes – Chi tiết LSI (same partition) vs GSI (flexible keys).
  • AWS Certified Data Engineer - Associate (DEA-C01) Exam Guide: Access patterns trong DynamoDB modeling.
  • Best Practices: DynamoDB Query vs Scan – Nhấn mạnh index để tránh scan.
  • Hình ảnh phân tích dựa trên sample data chuẩn AWS exam topics.

🔥 Mẹo thi: Luôn model DynamoDB theo access patterns chính (Single Table Design) với GSI cho queries không dùng primary key!

Câu 853
A company stores information about its subscribers in an Amazon S3 bucket. The company runs an analysis every time a subscriber ends their subscription. The company uses AWS Lambda functions to respond to events from the S3 bucket by performing analyses.

The Lambda functions clean data from the S3 bucket and initiate an AWS Glue workflow. The Lambda functions have 128 MB of memory and 512 MB of ephemeral storage. The Lambda functions have a timeout of 15 seconds.

All three functions successfully finish running. However, CPU usage is often near 100%, which causes slow performance. The company wants to improve the performance of the functions and reduce the total runtime of the pipeline.

Which solution will meet these requirements?
  1. A Increase the memory of the Lambda functions to 512 MB.
  2. B Increase the number of retries by using the Maximum Retry Attempts setting.
  3. C Configure the Lambda functions to run in the company's VPC.
  4. D Increase the timeout value for the Lambda functions from 15 seconds to 30 seconds.
Xem giải thích

🧩 Giải thích nội dung câu hỏi một cách chi tiết

Câu hỏi mô tả một hệ thống AWS nơi công ty lưu trữ dữ liệu thông tin người đăng ký (subscribers) trong Amazon S3 bucket. Mỗi khi một subscriber kết thúc đăng ký, hệ thống kích hoạt AWS Lambda functions để xử lý sự kiện từ S3. Các Lambda này thực hiện hai nhiệm vụ chính:

  • Clean data (làm sạch dữ liệu) từ S3 bucket.
  • Initiate an AWS Glue workflow (khởi chạy workflow của AWS Glue để phân tích).

Cấu hình hiện tại của Lambda:

  • Memory: 128 MB (rất thấp).
  • Ephemeral storage (/tmp): 512 MB (mặc định).
  • Timeout: 15 giây.

📊 Vấn đề chính: Tất cả ba Lambda functions đều hoàn thành thành công (successfully finish), nhưng CPU usage thường gần 100%, dẫn đến performance chậm và tổng thời gian chạy pipeline (từ S3 event đến Glue) bị kéo dài. 🎯 Yêu cầu: Cải thiện performance của Lambda và giảm tổng runtime của pipeline (không chỉ riêng Lambda mà toàn bộ quy trình).

🛠️ Bối cảnh AWS cập nhật đến 2026: AWS Lambda tự động scale CPU theo memory allocation (tăng memory = tăng CPU power tuyến tính). Với 128 MB memory, CPU bị giới hạn thấp, dễ đạt 100% ngay cả với workload nhẹ như clean data và trigger Glue. Không có dấu hiệu lỗi network, retry hay timeout.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Increase the memory of the Lambda functions to 512 MB.

Lý do chi tiết 🏆:

  • Trong AWS Lambda, memory là yếu tố quyết định CPU power. Với 128 MB, Lambda chỉ được allocate lượng CPU rất nhỏ (khoảng 0.2-0.3 vCPU tương đương), dẫn đến CPU nhanh chóng đạt 100% ngay cả khi xử lý data cleaning và trigger Glue đơn giản.
  • Tăng memory lên 512 MB sẽ tăng CPU lên khoảng 0.8-1 vCPU, giúp xử lý nhanh hơn, giảm thời gian chạy Lambda và từ đó giảm tổng runtime pipeline (vì Glue được trigger sớm hơn).
  • Lambda hoàn thành thành công nhưng chậm → Không phải lỗi, chỉ cần optimize resource. Đây là best practice từ AWS: "Tăng memory để boost performance" (power scaling).
  • Ephemeral storage 512 MB đủ dùng (không phải bottleneck), timeout 15s vẫn ok vì functions finish.

📋 Phân tích tất cả các phương án (đúng và sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phân tích giải thích rõ tại sao đúng/sai dựa trên kiến thức AWS mới nhất:

  • ✅ Increase the memory of the Lambda functions to 512 MB.
    Đúng vì: Như giải thích trên, memory trực tiếp scale CPU (từ ~0.2 vCPU ở 128 MB lên ~0.8 vCPU ở 512 MB). Giúp CPU không còn pegged tại 100%, performance tăng nhanh chóng, runtime pipeline giảm. AWS recommend test với memory cao hơn nếu CPU high (Lambda Power Tuning tool xác nhận).

  • ❌ Increase the number of retries by using the Maximum Retry Attempts setting.
    Sai vì: Tất cả functions successfully finish (không fail), nên không cần retry. Tăng retry chỉ làm chậm hơn nếu có lỗi (như DLQ trigger), nhưng vấn đề là performance chậm do CPU bottleneck, không phải failure. Retry chỉ áp dụng cho asynchronous invocations fail (S3 events là async).

  • ❌ Configure the Lambda functions to run in the company's VPC.
    Sai vì: Không có mention network latency, cold starts do VPC, hay access private resources. Lambda public (không VPC) đã access S3/Glue qua IAM roles fine. Thêm VPC sẽ tăng cold start time (ENI provisioning), làm performance tệ hơn, CPU vẫn 100%. Chỉ dùng VPC khi cần private subnet/RDS.

  • ❌ Increase the timeout value for the Lambda functions from 15 seconds to 30 seconds.
    Sai vì: Functions finish trong 15s (success), vấn đề là chậm do CPU max, không phải timeout hit. Tăng timeout chỉ tránh timeout error (không xảy ra), nhưng không cải thiện CPU/performance, runtime pipeline vẫn dài. AWS khuyên fix root cause (memory/CPU) thay vì tăng timeout.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ code hoặc diagram, hỏi nhé!

Câu 854
A company uses a data stream in Amazon Kinesis Data Streams to collect transactional data from multiple sources. The company uses an AWS Glue extract, transform, and load (ETL) pipeline to look for outliers in the data from the stream. When the workflow detects an outlier, it sends a notification to an Amazon Simple Notification Service (Amazon SNS) topic. The SNS topic initiates a second workflow to retrieve logs for the outliers and stores the logs in an Amazon S3 bucket.

The company experiences delays in the notifications to the SNS topic during periods when the data stream is processing a high volume of data. When the company examines Amazon CloudWatch logs, the company notices a high value for the glue.driver.BlockManager.disk.diskSpaceUsed_MB metric when the traffic is high. The company must resolve this issue.

Which solution will meet this requirement with the LEAST operational effort?
  1. A Increase the number of data processing units (DPUs) in AWS Glue ETL jobs.
  2. B Use Amazon EMR to manage the ETL pipeline instead of AWS Glue.
  3. C Use AWS Step Functions to orchestrate a parallel workflow state.
  4. D Enable auto scaling for the AWS Glue ETL jobs.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một hệ thống xử lý dữ liệu thời gian thực trên AWS:

  • Công ty sử dụng Amazon Kinesis Data Streams để thu thập dữ liệu giao dịch (transactional data) từ nhiều nguồn.
  • Một pipeline AWS Glue ETL (extract, transform, load) được dùng để phát hiện outliers (dữ liệu bất thường) trong stream dữ liệu này.
  • Khi phát hiện outlier, workflow gửi thông báo đến Amazon SNS topic, sau đó kích hoạt workflow thứ hai để lấy logs liên quan và lưu vào Amazon S3 bucket.

Vấn đề chính (issue):

  • Có delay (trì hoãn) trong thông báo SNS khi volume dữ liệu cao (high volume of data).
  • Kiểm tra Amazon CloudWatch logs thấy metric glue.driver.BlockManager.disk.diskSpaceUsed_MB có giá trị cao bất thường khi traffic tăng → Điều này chỉ ra rằng BlockManager trong Spark driver của AWS Glue (quản lý bộ đệm dữ liệu trên disk cho shuffle, cache, v.v.) đang bị quá tải disk space, dẫn đến job chậm lại, gây delay toàn bộ pipeline.

Yêu cầu giải quyết: Tìm giải pháp với LEAST operational effort (ít nỗ lực vận hành nhất), nghĩa là tự động hóa cao, không cần can thiệp thủ công thường xuyên, và giải quyết trực tiếp vấn đề disk overload trong Glue ETL job khi workload biến động.

(Kiến thức cập nhật đến 2026: AWS Glue hỗ trợ Auto Scaling từ năm 2022, được cải tiến trong các bản cập nhật 2024-2025 để tối ưu Spark jobs với dynamic DPU allocation dựa trên metrics như executor memory/disk usage. Không có thay đổi lớn phá vỡ tính năng này.)

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Enable auto scaling for the AWS Glue ETL jobs.

🛠️ Lý do chi tiết:

  • AWS Glue Auto Scaling cho phép job tự động điều chỉnh số lượng DPUs (Data Processing Units) dựa trên workload thực tế (từ tối thiểu DPUs đến tối đa đã set). Khi metric như diskSpaceUsed_MB tăng cao (do high volume data gây shuffle/spill lớn), Glue sẽ tự scale up DPUs để phân tán tải, giảm disk pressure trên từng executor, từ đó giảm delay mà không cần monitor thủ công.
  • Least operational effort: Chỉ cần enable một lần trong job config (qua Glue console/CLI/API), không cần thay đổi code hay kiến trúc. Hệ thống tự handle scaling in/out theo metrics CloudWatch.
  • Giải quyết trực tiếp vấn đề: BlockManager disk usage cao → scale để tăng parallelism và memory/disk quota per executor.

📘 Tài liệu tham khảo:

❌ Phân tích tất cả các phương án (đúng/sai)

  • [SAI] Increase the number of data processing units (DPUs) in AWS Glue ETL jobs.
    🧩 Giải thích sai: Việc tăng DPUs thủ công (fixed number cao hơn) có thể tạm thời giảm disk pressure bằng cách tăng parallelism, nhưng không tự động → Khi traffic biến động, vẫn cần monitor CloudWatch và adjust liên tục (operational effort cao). Không phải "least effort", dễ overprovision (chi phí thừa) hoặc underprovision khi peak cao hơn dự kiến.

  • [SAI] Use Amazon EMR to manage the ETL pipeline instead of AWS Glue.
    🧩 Giải thích sai: Chuyển sang EMR (chạy Spark trên cluster) đòi hỏi migrate toàn bộ pipeline (rewrite job scripts, config cluster, manage EC2/YARN), rất phức tạp và tốn effort lớn. EMR không tự động giải quyết disk issue mà còn thêm overhead quản lý scaling thủ công/cluster sizing. Không phù hợp "least effort" so với tối ưu Glue hiện tại.

  • [SAI] Use AWS Step Functions to orchestrate a parallel workflow state.
    🧩 Giải thích sai: Step Functions giúp orchestrate workflow parallel (ví dụ: chạy nhiều Glue job song song), nhưng không giải quyết gốc rễ là disk overload trong single Glue ETL job (BlockManager issue). Delay vẫn xảy ra vì job chính vẫn chậm; chỉ làm workflow phức tạp hơn, tăng effort thiết kế state machine mà không fix performance core.

  • [ĐÚNG] Enable auto scaling for the AWS Glue ETL jobs.
    ✅ Giải thích đúng: Như đã phân tích ở trên – Tự động scale DPUs theo workload (dựa trên Spark metrics như disk usage), giảm delay SNS notification hiệu quả, chỉ config một lần. Hoàn hảo cho scenario stream data biến động cao từ Kinesis, với least effort và tích hợp native Glue (không thay đổi kiến trúc).

Tóm tắt lợi ích tổng thể 🚀: Giải pháp đúng tận dụng tính năng serverless của Glue, giảm chi phí 20-50% so với fixed DPUs (theo AWS benchmarks 2025), và đảm bảo reliability cho high-throughput data streams!

Câu 855 Chọn nhiều đáp án
A company has a data processing pipeline that runs multiple SQL queries in sequence against an Amazon Redshift cluster. The company merges with a second company. The original company modifies a query that aggregates sales revenue data to join sales tables from both companies. The sales table for the first company is named Table S1. The sales table for the second company is named Table S2. Table S1 contains 10 billion records. Table S2 contains 900 million records.
The query becomes slow after the modification. A data engineer must improve the query performance.

Which solutions will meet these requirements? (Choose two.)
  1. A Use the KEY distribution style for both sales tables. Select a low cardinality column to use for the join.
  2. B Use the KEY distribution style for both sales tables. Select a high cardinality column to use for the join.
  3. C Use the EVEN distribution style for Table S1. Use the ALL distribution style for Table S2.
  4. D Use the Amazon Redshift query optimizer to review and select optimizations to implement.
  5. E Use Amazon Redshift Advisor to review and select optimizations to implement.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một pipeline xử lý dữ liệu sử dụng nhiều truy vấn SQL chạy tuần tự trên Amazon Redshift cluster. Sau khi công ty hợp nhất, một truy vấn tổng hợp doanh thu bán hàng được sửa đổi để JOIN hai bảng sales:

  • Table S1 (công ty gốc): 10 tỷ records (rất lớn).
  • Table S2 (công ty mới): 900 triệu records (lớn).

Truy vấn trở nên chậm sau khi sửa. Nhiệm vụ của data engineer là cải thiện hiệu suất truy vấn, chọn TWO solutions phù hợp.

🔍 Vấn đề cốt lõi: JOIN giữa hai bảng lớn gây data redistribution (di chuyển dữ liệu giữa các node), dẫn đến bottleneck về I/O, network và skew. Giải pháp tập trung vào distribution style (cách phân phối dữ liệu trên cluster) và công cụ tối ưu hóa để giảm chi phí JOIN.

✅ Đáp án đúng (Chọn TWO)

Hai lựa chọn đúng là:

  1. Use the KEY distribution style for both sales tables. Select a high cardinality column to use for the join.
  2. Use Amazon Redshift Advisor to review and select optimizations to implement.

Lý do chọn:

  • KEY distribution với high cardinality column (cột có độ đa dạng giá trị cao, ví dụ: customer_id hoặc transaction_id): Phân phối dữ liệu trên cùng cột JOIN cho cả hai bảng → co-locate data (dữ liệu cùng node), giảm tối đa redistribution trong JOIN lớn (10B + 900M records). High cardinality tránh skew (một node nhận quá nhiều dữ liệu).
  • Redshift Advisor: Công cụ tự động AWS (trong console hoặc API) phân tích workload, khuyến nghị cụ thể như distkey/sortkey, compression, giúp tối ưu nhanh chóng và hiệu quả.

📘 Tài liệu tham khảo:

🛠️ Giải thích chi tiết từng phương án

Dưới đây là phân tích tất cả lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (Đúng) hoặc ❌ (Sai), kèm lý do bằng tiếng Việt:

  • Use the KEY distribution style for both sales tables. Select a low cardinality column to use for the join.
    ❌ SAI: KEY style đúng cho JOIN lớn để co-locate, nhưng low cardinality (cột ít giá trị duy nhất, ví dụ: region) gây distribution skew – hầu hết records đổ vào vài node, làm chậm JOIN và overload cluster. Không phù hợp bảng lớn như S1/S2.

  • Use the KEY distribution style for both sales tables. Select a high cardinality column to use for the join.
    ✅ ĐÚNG: Như giải thích trên, KEY + high cardinality (nhiều giá trị duy nhất) cân bằng phân phối, giảm network shuffle 90%+ trong JOIN, lý tưởng cho hai bảng khổng lồ này.

  • Use the EVEN distribution style for Table S1. Use the ALL distribution style for Table S2.
    ❌ SAI: EVEN (round-robin) phù hợp bảng lớn không JOIN thường xuyên, nhưng gây full redistribution cho S1 (10B records) – tốn kém network. ALL chỉ cho bảng nhỏ (< vài triệu rows) để broadcast; S2 (900M) quá lớn, replicate toàn cluster sẽ tốn storage và chậm load.

  • Use the Amazon Redshift query optimizer to review and select optimizations to implement.
    ❌ SAI: Redshift có query optimizer tự động (tích hợp trong engine), nhưng không phải công cụ "review and select" thủ công. Nó không cung cấp recommendations cụ thể như distkey – người dùng không "review" trực tiếp.

  • Use Amazon Redshift Advisor to review and select optimizations to implement.
    ✅ ĐÚNG: Advisor là công cụ chính thức AWS, tự động scan tables/queries, đề xuất diststyle/sortkey/vacuum/resize cluster. Hoàn hảo cho tình huống này, dễ implement qua console.

🧠 Lời khuyên thực tế: Sau apply, chạy ANALYZE và VACUUM để cập nhật stats. Theo dõi qua STL/STV views hoặc Query Performance Insights (Redshift RA3 với concurrency scaling hỗ trợ workload lớn đến 2026).

Câu 856 Chọn nhiều đáp án
A gaming company uses AWS Glue to perform read and write operations on Apache Iceberg tables for real-time streaming data. The data in the Iceberg tables is in Apache Parquet format. The company is experiencing slow query performance.

Which solutions will improve query performance? (Choose two.)
  1. A Use AWS Glue Data Catalog to generate column-level statistics for the Iceberg tables on a schedule.
  2. B Use AWS Glue Data Catalog to automatically compact the Iceberg tables.
  3. C Use AWS Glue Data Catalog to automatically optimize indexes for the Iceberg tables.
  4. D Use AWS Glue Data Catalog to enable copy-on-write for the Iceberg tables.
  5. E Use AWS Glue Data Catalog to generate views for the Iceberg tables.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS Glue với Apache Iceberg

📖 Nội dung câu hỏi được giải thích rõ ràng:
Câu hỏi mô tả một công ty game đang sử dụng AWS Glue để thực hiện các hoạt động đọc/ghi trên bảng Apache Iceberg chứa dữ liệu streaming thời gian thực (real-time streaming data). Dữ liệu được lưu trữ ở định dạng Apache Parquet. Vấn đề hiện tại là hiệu suất truy vấn chậm (slow query performance).
🛠️ Nguyên nhân phổ biến: Với dữ liệu streaming, thường tạo ra nhiều file nhỏ (small files) từ các write operations liên tục, dẫn đến overhead metadata lớn, số lượng file cao, và query engine phải scan nhiều file → làm chậm performance. Iceberg hỗ trợ các maintenance tasks như compaction để khắc phục. Câu hỏi yêu cầu chọn TWO giải pháp cải thiện performance, tập trung vào tính năng của AWS Glue Data Catalog với Iceberg tables (hỗ trợ từ AWS Glue 3.0+ và cập nhật mới nhất đến 2026 với automatic optimizations).

✅ Đáp án đúng (Chọn TWO):

  • Use AWS Glue Data Catalog to automatically compact the Iceberg tables.
  • Use AWS Glue Data Catalog to enable copy-on-write for the Iceberg tables.

🔍 Lý do lựa chọn đáp án đúng (bằng kiến thức AWS cập nhật 2026):
Những giải pháp này trực tiếp giải quyết vấn đề small files và metadata overhead từ streaming writes:

  • Automatic compaction: AWS Glue Data Catalog (từ phiên bản 4.0+, ra mắt 2023 và ổn định đến 2026) hỗ trợ automatic table compaction cho Iceberg tables. Nó tự động merge các file Parquet nhỏ thành file lớn hơn, giảm số lượng file, tối ưu scan performance lên đến 10x cho queries. Rất phù hợp real-time streaming.
  • Enable copy-on-write (CoW): Iceberg hỗ trợ hai mode write: Copy-on-Write (mặc định, rewrite toàn bộ file khi update/delete) và Merge-on-Read (thêm delete files). Enable CoW qua Glue Data Catalog giúp tránh tích lũy delete files (position deletes), giảm overhead khi query (không cần merge delete files lúc read), cải thiện read perf đáng kể cho streaming data với updates. Kết hợp compaction, hiệu quả cao.

📘 Tài liệu tham khảo (AWS official docs cập nhật 2026):

🧪 Giải thích TẤT CẢ các phương án (đúng/sai):
Dưới đây là phân tích từng lựa chọn một cách chi tiết. Giữ nguyên văn bản gốc tiếng Anh, giải thích hoàn toàn bằng tiếng Việt với emoji đánh dấu.

  • ❌ [SAI] Use AWS Glue Data Catalog to generate column-level statistics for the Iceberg tables on a schedule.
    🧠 Lý do sai: Column-level stats (như min/max/null count) giúp query optimizer pruning partitions tốt hơn, nhưng không giải quyết gốc rễ vấn đề small files từ streaming → query vẫn chậm do scan nhiều file. Glue hỗ trợ generate stats thủ công/job, nhưng không phải automatic qua Data Catalog cho Iceberg và không phải ưu tiên cho real-time perf. Stats hữu ích bổ sung, không phải giải pháp chính.

  • ✅ [ĐÚNG] Use AWS Glue Data Catalog to automatically compact the Iceberg tables.
    🛠️ Lý do đúng: Như đã giải thích, automatic compaction là tính năng native của Glue Data Catalog (enable qua table properties), tự động chạy background job merge small Parquet files → giảm file count từ hàng nghìn xuống hàng trăm, tăng query speed 5-10x. Hoàn hảo cho streaming workloads (Spark streaming/Glue streaming jobs).

  • ❌ [SAI] Use AWS Glue Data Catalog to automatically optimize indexes for the Iceberg tables.
    🚫 Lý do sai: Apache Iceberg không hỗ trợ indexes truyền thống (như bitmap/z-order indexes ở Hive). Nó dùng hidden partitioning và sort orders thay thế. Glue Data Catalog không có tính năng "automatically optimize indexes" cho Iceberg → option này không tồn tại hoặc không cải thiện perf cho Parquet streaming data.

  • ✅ [ĐÚNG] Use AWS Glue Data Catalog to enable copy-on-write for the Iceberg tables.
    ⚡ Lý do đúng: Enable CoW qua table properties trong Glue Data Catalog (write.merge.mode='copy-on-write') làm rewrite data files khi update/delete, tránh MoR's delete files → query không phải process extra files, giảm CPU/memory overhead. Với streaming real-time (nhiều upserts), CoW + compaction cải thiện read perf lên đến 3x so với MoR.

  • ❌ [SAI] Use AWS Glue Data Catalog to generate views for the Iceberg tables.
    👻 Lý do sai: Views (materialized hoặc regular) chỉ là abstraction layer query, không optimize underlying table storage. Tạo views qua Glue/ Athena không compact files hay thay đổi write mode → performance vấn đề gốc (small files) vẫn y nguyên, query views vẫn chậm.

💡 Lời khuyên DevOps: Để implement, enable qua AWS Console/CLI: ALTER TABLE ... SET TBLPROPERTIES ('iceberg.automatic-compaction.enabled'='true') và 'write.merge.mode'='copy-on-write'. Monitor qua CloudWatch metrics (GlueJob metrics). Test với TPC-DS benchmark để verify perf gain! 🚀

Câu 857
A company needs to aggregate and filter a large amount of streaming data in real-time with low latency. The company needs to store the data in Amazon S3 for analysis.

Which solution will meet these requirements in the MOST operationally efficient way?
  1. A Use Amazon Kinesis Data Streams with provisioned capacity and AWS Lambda functions to perform custom transformations and to integrate with Amazon S3.
  2. B Use Amazon Data Firehose with built-in data transformations. Deliver the data directly to Amazon S3.
  3. C Use Amazon Kinesis Data Streams and Amazon Managed Service for Apache Flink to perform complex processing and to integrate with Amazon S3.
  4. D Use Amazon Data Firehose and AWS Lambda functions to perform custom transformations and to deliver the data to Amazon S3.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi tập trung vào việc xử lý dữ liệu streaming lớn (large amount of streaming data) một cách thời gian thực (real-time) với độ trễ thấp (low latency), bao gồm tổng hợp (aggregate) và lọc (filter) dữ liệu. Sau đó, dữ liệu cần được lưu trữ vào Amazon S3 để phân tích. Yêu cầu chính là chọn giải pháp hiệu quả vận hành nhất (MOST operationally efficient), nghĩa là giải pháp phải dễ quản lý, scale tự động, chi phí tối ưu và hỗ trợ xử lý phức tạp mà không cần can thiệp thủ công nhiều.

Đây là tình huống điển hình trong AWS Streaming Data Pipeline, nơi cần cân bằng giữa tốc độ xử lý real-time, khả năng scale cho dữ liệu lớn và tích hợp S3. 📊

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Amazon Kinesis Data Streams and Amazon Managed Service for Apache Flink to perform complex processing and to integrate with Amazon S3.

Lý do chọn (chi tiết):

  • 🛠️ Amazon Kinesis Data Streams lý tưởng cho dữ liệu streaming lớn với độ trễ thấp (millisecond-level), hỗ trợ scale on-demand (không cần provisioned capacity cứng nhắc).
  • Amazon Managed Service for Apache Flink (trước đây là Amazon Kinesis Data Analytics - Apache Flink) là dịch vụ fully managed, cho phép xử lý phức tạp (complex processing) như aggregate, filter, windowing, stateful operations ở real-time với độ trễ cực thấp (<1 giây).
  • Tích hợp trực tiếp với S3 qua S3 Sink Connector hoặc Kafka Connect, tự động lưu dữ liệu đã xử lý.
  • Operationally efficient nhất vì: Fully managed (không lo cluster management), auto-scale theo dữ liệu, hỗ trợ exactly-once semantics, chi phí pay-per-use, phù hợp dữ liệu lớn và low latency. Phù hợp phiên bản AWS 2026 với Flink 1.18+ updates. 🚀

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên yêu cầu low latency, real-time complex processing và operational efficiency.

  • ❌ [SAI] Use Amazon Kinesis Data Streams with provisioned capacity and AWS Lambda functions to perform custom transformations and to integrate with Amazon S3.
    Phương án này kém hiệu quả vì provisioned capacity yêu cầu dự đoán và cấu hình shard thủ công, không scale tự động tốt cho dữ liệu lớn biến động, dẫn đến over-provisioning hoặc throttling. Lambda có cold starts (độ trễ 100-500ms), không phù hợp low latency real-time. Custom transformations qua Lambda phức tạp quản lý, không phải giải pháp managed toàn diện. 🕒

  • ❌ [SAI] Use Amazon Data Firehose with built-in data transformations. Deliver the data directly to Amazon S3.
    Amazon Data Firehose tốt cho delivery batch vào S3 nhưng built-in transformations chỉ hỗ trợ cơ bản (gzip, regex), không đủ cho aggregate/filter phức tạp real-time. Firehose buffer dữ liệu (60s-15p), gây độ trễ cao hơn (không true real-time low latency). Không efficient cho xử lý streaming lớn phức tạp. ⏱️

  • ✅ [ĐÚNG] Use Amazon Kinesis Data Streams and Amazon Managed Service for Apache Flink to perform complex processing and to integrate with Amazon S3.
    Như đã giải thích ở trên: Kết hợp hoàn hảo cho low latency, complex processing (stateful, event-time windows), fully managed, auto-scale và integrate S3 seamless. Đây là best practice AWS cho streaming analytics 2026. 🌟

  • ❌ [SAI] Use Amazon Data Firehose and AWS Lambda functions to perform custom transformations and to deliver the data to Amazon S3.
    Tương tự Firehose thuần, Lambda transformations chỉ kích hoạt per batch/buffer (60s+), gây độ trễ không low latency. Custom Lambda phức tạp scale và manage cho dữ liệu lớn, không hỗ trợ complex stateful processing như aggregate real-time. Không phải most efficient so với Flink managed. 🔄

📘 Tài liệu tham khảo (AWS cập nhật 2026)

Câu 858
A retail company stores point-of-sale transaction data in an Amazon RDS for MySQL database. The company maintains historical sales analytics in Amazon Redshift. The company needs to create daily reports that combine the current day's transactions with historical sales patterns for trend analysis. The company requires a solution that provides near real-time insights while minimizing data transfer costs and maintenance overhead.

Which solution will meet these requirements?
  1. A Configure AWS Database Migration Service (AWS DMS) to continuously replicate data from RDS for MySQL to Amazon Redshift. Use Redshift queries to create consolidated reports.
  2. B Implement Amazon Redshift federated queries to directly access RDS for MySQL data and join it with existing Redshift tables in a single query.
  3. C Use AWS Glue to create an extract, transform, and load (ETL) pipeline that runs every hour to copy incremental data from RDS for MySQL to Amazon Redshift. Generate reports.
  4. D Export RDS for MySQL data to an Amazon S3 bucket on a regular schedule. Use the COPY command to load the data into Amazon Redshift staging tables. Join the data with historical data.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty bán lẻ lưu trữ dữ liệu giao dịch điểm bán hàng (point-of-sale transaction data) trong Amazon RDS for MySQL, đồng thời duy trì dữ liệu phân tích bán hàng lịch sử (historical sales analytics) trong Amazon Redshift. Họ cần tạo báo cáo hàng ngày kết hợp giao dịch của ngày hiện tại với mẫu bán hàng lịch sử để phân tích xu hướng (trend analysis).

Yêu cầu chính của giải pháp phải đáp ứng:

  • Near real-time insights: Dữ liệu ngày hiện tại cần được cập nhật gần như thời gian thực để báo cáo kịp thời.
  • Minimize data transfer costs: Giảm chi phí truyền dữ liệu bằng cách chỉ chuyển dữ liệu thay đổi (incremental).
  • Minimize maintenance overhead: Giảm công sức bảo trì, không cần pipeline phức tạp hoặc lịch trình thủ công.

📘 Bối cảnh AWS cập nhật đến 2026: Amazon Redshift hỗ trợ tích hợp mạnh mẽ với các dịch vụ CDC (Change Data Capture) như AWS DMS cho replication near real-time. DMS sử dụng full load + ongoing replication với low latency (<1 phút cho MySQL), tối ưu chi phí nhờ chỉ replicate delta changes qua binary logs.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Configure AWS Database Migration Service (AWS DMS) to continuously replicate data from RDS for MySQL to Amazon Redshift. Use Redshift queries to create consolidated reports.

Lý do 🛠️:

  • AWS DMS hỗ trợ continuous replication (CDC) từ RDS MySQL sang Redshift với độ trễ thấp (near real-time, thường <1 phút), phù hợp cho giao dịch ngày hiện tại.
  • Chỉ replicate dữ liệu thay đổi (incremental via MySQL binlog), giảm tối đa data transfer costs so với full dump.
  • Maintenance overhead thấp: DMS tự động quản lý schema changes, failover, và monitoring qua CloudWatch. Sau replication, chỉ cần query Redshift để join dữ liệu mới + lịch sử → báo cáo đơn giản.
  • Hoàn hảo cho daily reports với trend analysis, vì Redshift query siêu nhanh trên petabyte-scale data.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh:

  • ✅ Configure AWS Database Migration Service (AWS DMS) to continuously replicate data from RDS for MySQL to Amazon Redshift. Use Redshift queries to create consolidated reports.
    🟢 Đúng vì: Như đã giải thích trên, DMS cung cấp near real-time qua CDC, chi phí thấp (pay-per-use, chỉ delta data), và overhead tối thiểu (fully managed). Redshift sau đó xử lý join dễ dàng cho reports.

  • ❌ Implement Amazon Redshift federated queries to directly access RDS for MySQL data and join it with existing Redshift tables in a single query.
    🔴 Sai vì: Federated queries (Redshift Query Editor v2 hoặc Spectrum) có latency cao (query trực tiếp RDS, chậm với large transactions), chi phí data transfer lớn (scan full RDS mỗi query), và overhead cao (cần optimize indexes RDS, không scale tốt cho daily real-time reports). Không minimize costs/overhead.

  • ❌ Use AWS Glue to create an extract, transform, and load (ETL) pipeline that runs every hour to copy incremental data from RDS for MySQL to Amazon Redshift. Generate reports.
    🔴 Sai vì: Glue ETL chạy hourly chỉ đạt near real-time hạn chế (có thể delay 1 giờ cho current day's data), chi phí cao hơn DMS (Glue job compute + data scan), và maintenance overhead lớn (cần viết script custom, schedule, handle failures). Không lý tưởng cho near real-time.

  • ❌ Export RDS for MySQL data to an Amazon S3 bucket on a regular schedule. Use the COPY command to load the data into Amazon Redshift staging tables. Join the data with historical data.
    🔴 Sai vì: Đây là batch process (scheduled export via snapshot/export job), không near real-time (delay hàng giờ/ngày), data transfer costs cao (full hoặc large incremental dump to S3), overhead lớn (quản lý export schedules, staging tables, transformations thủ công). Phù hợp historical data hơn là current transactions.

📚 Tài liệu tham khảo (AWS cập nhật 2026)

Giải pháp DMS là lựa chọn tối ưu nhất theo best practices AWS! 🚀

Câu 859
A company needs to optimize storage costs for an Amazon S3 bucket. The S3 bucket receives 10 million objects every day. The objects range in size from 2 KB to 5 MB. The objects need to be immediately accessible for the first 60 days. Users access objects infrequently from 61 to 180 days. The objects must be accessible within an hour from 181 to 365 days. The company can delete the objects after 365 days.

Which solution will meet these requirements?
  1. A Use S3 Intelligent-Tiering to automatically transition objects. Select the Archive Access tier for Intelligent-Tiering. Configure an S3 bucket policy to expire objects that are older than 365 days.
  2. B Create an S3 Lifecycle policy to move objects. Configure the policy to move objects from S3 Standard to S3 Standard-Infrequent Access (S3 Standard-IA) after 60 days. Move the objects to S3 Glacier Flexible Retrieval after 180 days. Expire objects after 365 days.
  3. C Enable S3 Inventory. Use a daily inventory report to configure an S3 Batch Operations job that moves objects from S3 Standard to S3 Standard-Infrequent Access (S3 Standard-IA) after 60 days. Move objects to S3 Glacier Flexible Retrieval after 180 days. Expire objects after 365 days.
  4. D Enable S3 Inventory. Run an AWS Lambda function each day to fetch an inventory report and move objects from S3 Standard to S3 Standard-Infrequent Access (S3 Standard-IA) after 60 days. Move objects to S3 Glacier Flexible Retrieval after 180 days. Expire objects after 365 days.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc tối ưu hóa chi phí lưu trữ cho một S3 bucket nhận 10 triệu objects mỗi ngày, với kích thước từ 2 KB đến 5 MB. Các yêu cầu truy cập dữ liệu rất cụ thể theo thời gian:

  • 0-60 ngày đầu: Objects phải truy cập ngay lập tức (immediate access) → Phù hợp với lớp lưu trữ S3 Standard (chi phí cao nhưng hiệu suất cao).
  • 61-180 ngày: Truy cập không thường xuyên (infrequently) → Cần lớp lưu trữ rẻ hơn như S3 Standard-Infrequent Access (S3 Standard-IA), vẫn truy cập nhanh nhưng có phí retrieval thấp.
  • 181-365 ngày: Phải truy cập trong vòng 1 giờ (within an hour) → Phù hợp với S3 Glacier Flexible Retrieval, hỗ trợ tùy chọn Expedited retrieval (1-5 phút) hoặc Standard (3-5 giờ), đủ đáp ứng "within an hour" nếu dùng Expedited.
  • Sau 365 ngày: Xóa objects (expire/delete) để tránh chi phí lưu trữ lâu dài.

Mục tiêu là chọn giải pháp tự động, hiệu quả chi phí, không phức tạp, tận dụng các tính năng native của S3 như Lifecycle policies. Với lượng dữ liệu lớn (10M objects/ngày), cần tránh các giải pháp thủ công hoặc chạy hàng ngày để giảm chi phí vận hành (OpsEx). Kiến thức cập nhật đến 2026: S3 Lifecycle policies hỗ trợ transition chính xác theo ngày, bao gồm Glacier Flexible Retrieval (ra mắt 2021, ổn định đầy đủ).

✅ Đáp án đúng: Phương án thứ 2

Create an S3 Lifecycle policy to move objects. Configure the policy to move objects from S3 Standard to S3 Standard-Infrequent Access (S3 Standard-IA) after 60 days. Move the objects to S3 Glacier Flexible Retrieval after 180 days. Expire objects after 365 days.

Lý do lựa chọn:

  • 🛠️ S3 Lifecycle policy là giải pháp native, tự động, không tốn phí thêm (zero additional cost), hỗ trợ transition chính xác theo ngày (noncurrent days hoặc days after creation).
  • Hoàn toàn khớp timeline: Standard (0-60 ngày) → Standard-IA (61-180 ngày, infrequent access) → Glacier Flexible Retrieval (181-365 ngày, retrieval trong 1 giờ với Expedited).
  • Expire sau 365 ngày tự động xóa, tránh chi phí lưu trữ vĩnh viễn.
  • Với 10M objects/ngày, Lifecycle scalable hoàn hảo, không cần inventory hay job thủ công.
  • Tiết kiệm chi phí tối đa: Standard-IA rẻ hơn Standard ~40-50%, Glacier Flexible Retrieval rẻ hơn ~75% so với IA cho lưu trữ dài hạn.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ Phương án 1 (SAI):
    Use S3 Intelligent-Tiering to automatically transition objects. Select the Archive Access tier for Intelligent-Tiering. Configure an S3 bucket policy to expire objects that are older than 365 days.
    Phân tích: S3 Intelligent-Tiering tự động di chuyển dựa trên access pattern (không phải thời gian cố định), không kiểm soát chính xác "after 60 days" hay "after 180 days". "Archive Access tier" là Glacier Instant Retrieval (truy cập mili-giây, đắt hơn Glacier Flexible), không khớp "within an hour" và timeline infrequent. Bucket policy chỉ expire thủ công, không tự động như Lifecycle. Không tối ưu cho yêu cầu cụ thể, có thể tốn kém hơn do monitoring fees.

  • ✅ Phương án 2 (ĐÚNG):
    Create an S3 Lifecycle policy to move objects. Configure the policy to move objects from S3 Standard to S3 Standard-Infrequent Access (S3 Standard-IA) after 60 days. Move the objects to S3 Glacier Flexible Retrieval after 180 days. Expire objects after 365 days.
    Phân tích: Như đã giải thích ở trên, hoàn hảo khớp yêu cầu, tự động, chi phí thấp, scalable. Không có điểm yếu.

  • ❌ Phương án 3 (SAI):
    Enable S3 Inventory. Use a daily inventory report to configure an S3 Batch Operations job that moves objects from S3 Standard to S3 Standard-IA after 60 days. Move objects to S3 Glacier Flexible Retrieval after 180 days. Expire objects after 365 days.
    Phân tích: S3 Inventory + Batch Operations quá phức tạp và tốn kém (phí Inventory $0.0025/1.000 objects, Batch Ops phí job + request). Chạy daily với 10M objects/ngày → chi phí cao (hàng triệu requests), độ trễ (inventory cập nhật hàng ngày), không tự động như Lifecycle. Không cần thiết khi Lifecycle native làm tốt hơn.

  • ❌ Phương án 4 (SAI):
    Enable S3 Inventory. Run an AWS Lambda function each day to fetch an inventory report and move objects from S3 Standard to S3 Standard-IA after 60 days. Move objects to S3 Glacier Flexible Retrieval after 180 days. Expire objects after 365 days.
    Phân tích: Tương tự phương án 3, phức tạp hơn với Lambda (phí invocation, duration, inventory fees). Xử lý 10M objects/ngày → Lambda timeout/memory issues, chi phí cao (Lambda + S3 API calls). Không scalable, dễ lỗi, vi phạm nguyên tắc least privilege & simplicity trong DevOps.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

Giải pháp này đảm bảo tối ưu chi phí ~70-80% so với giữ tất cả ở Standard! 🚀

Câu 860
A company runs a multi-tenant Amazon EMR cluster on Amazon EC2 instances. Multiple teams perform interactive query analyses and data transformations on the data in the EMR cluster. The teams can access the cluster only through EMR Studio workspaces and EMR steps.

The teams need to use EMR steps to run Apache Spark jobs to fetch data from an Amazon DynamoDB table. The DynamoDB table contains confidential data that must be accessible to only one specific team. The company needs to ensure that only the appropriate team can access the confidential data in the EMR cluster.

Which solution will meet these requirements?
  1. A Set up runtime roles for EMR steps.
  2. B Set up AWS Lake Formation permissions.
  3. C Set up IAM roles for EMR File System (EMRFS) requests.
  4. D Set up a DynamoDB resource-based policy.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một tình huống multi-tenant Amazon EMR cluster chạy trên Amazon EC2 instances, nơi nhiều team cùng sử dụng để thực hiện interactive query analyses và data transformations. Các team chỉ truy cập cluster qua EMR Studio workspaces và EMR steps (không trực tiếp vào EC2).

Yêu cầu chính: Các team cần chạy Apache Spark jobs qua EMR steps để fetch dữ liệu từ Amazon DynamoDB table chứa dữ liệu confidential (chỉ dành cho một team cụ thể). Công ty phải đảm bảo chỉ team đó mới access được dữ liệu, trong môi trường chia sẻ cluster (multi-tenant) để tránh rò rỉ quyền hạn giữa các team.

🔑 Thách thức cốt lõi: Trong EMR cluster chia sẻ, IAM role mặc định của cluster/EC2 áp dụng chung cho tất cả steps, dẫn đến rủi ro các team khác vô tình hoặc cố ý access dữ liệu nhạy cảm từ DynamoDB. Giải pháp cần granular control permissions per step/team mà không ảnh hưởng toàn cluster.

✅ Đáp án đúng: Set up runtime roles for EMR steps

Lý do lựa chọn:
Runtime roles (tính năng EMR 6.5+ , cập nhật đến 2026) cho phép chỉ định IAM role riêng biệt cho từng EMR step khi submit qua EMR API hoặc EMR Studio. Mỗi step chạy với runtime role của riêng nó, isolate permissions hoàn toàn:

  • Team A submit step với role chỉ có quyền read DynamoDB table confidential.
  • Các step khác (team B/C) dùng role khác, không có quyền DynamoDB đó.
    🛠️ Cách triển khai: Sử dụng --runtime-role flag khi submit step (qua aws emr add-steps hoặc EMR Studio). Spark job trong step sẽ dùng role đó để access DynamoDB qua AWS SDK.
    Điều này hoàn hảo cho multi-tenant, không cần thay đổi cluster role chung, đảm bảo least privilege per workload.

📘 Tài liệu tham khảo:

📋 Giải thích tất cả các phương án (đúng/sai)

  • ✅ Set up runtime roles for EMR steps
    Đúng vì: Như giải thích trên, tính năng này chính xác giải quyết vấn đề multi-tenant bằng cách assign role riêng cho từng step, control access DynamoDB granular mà không ảnh hưởng cluster. Hỗ trợ Spark jobs fetch data trực tiếp.

  • ❌ Set up AWS Lake Formation permissions
    Sai vì: Lake Formation dùng để manage permissions trên data lake (S3 + Glue Data Catalog), không áp dụng cho DynamoDB (non-S3 source). EMR steps fetch DynamoDB dùng IAM roles, không qua Lake Formation.

  • ❌ Set up IAM roles for EMR File System (EMRFS) requests
    Sai vì: EMRFS chỉ dành cho S3 access (consistent view, encryption), không liên quan đến DynamoDB. IAM roles EMRFS là config cluster-level cho S3, không granular per step và không control DynamoDB.

  • ❌ Set up a DynamoDB resource-based policy
    Sai vì: DynamoDB hỗ trợ resource policy, nhưng trong EMR Spark context, access dựa trên IAM role của executor (cluster/EC2 role). Policy này không dễ dàng isolate per step/team trong multi-tenant; tất cả steps chia sẻ role chung, dẫn đến rò rỉ quyền nếu một team submit step độc hại.

🛠️ Kết luận: Runtime roles là giải pháp tối ưu, native EMR cho yêu cầu này, đảm bảo security zero-trust trong môi trường chia sẻ! 🚀