Ngân hàng đề — AWS Certified Data Engineer Associate

Tìm thấy 867 câu.

Câu 661
A company stores details about transactions in an Amazon S3 bucket. The company wants to log all writes to the S3 bucket into another S3 bucket that is in the same AWS Region.
Which solution will meet this requirement with the LEAST operational effort?
  1. A Configure an S3 Event Notifications rule for all activities on the transactions S3 bucket to invoke an AWS Lambda function. Program the Lambda function to write the event to Amazon Kinesis Data Firehose. Configure Kinesis Data Firehose to write the event to the logs S3 bucket.
  2. B Create a trail of management events in AWS CloudTraiL. Configure the trail to receive data from the transactions S3 bucket. Specify an empty prefix and write-only events. Specify the logs S3 bucket as the destination bucket.
  3. C Configure an S3 Event Notifications rule for all activities on the transactions S3 bucket to invoke an AWS Lambda function. Program the Lambda function to write the events to the logs S3 bucket.
  4. D Create a trail of data events in AWS CloudTraiL. Configure the trail to receive data from the transactions S3 bucket. Specify an empty prefix and write-only events. Specify the logs S3 bucket as the destination bucket.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc ghi log tất cả các hoạt động ghi (writes) vào một S3 bucket chứa dữ liệu giao dịch (transactions S3 bucket), và lưu log này vào một S3 bucket khác cùng Region. Yêu cầu chính là giải pháp có ít nỗ lực vận hành nhất (LEAST operational effort).

  • Chi tiết vấn đề: "Writes" ở đây ám chỉ các hoạt động mức object như PutObject, CopyObject (tạo hoặc ghi dữ liệu vào object trong S3). Không phải management events (như CreateBucket). Giải pháp phải tự động, đáng tin cậy, và tối ưu hóa về quản lý (không cần code tùy chỉnh, ít components).
  • Yêu cầu then chốt: Least operational effort nghĩa là ưu tiên dịch vụ native AWS, không code, không Lambda hay streaming phức tạp. Phải cùng Region để tránh chi phí chuyển dữ liệu.

✅ Đáp án đúng

Create a trail of data events in AWS CloudTraiL. Configure the trail to receive data from the transactions S3 bucket. Specify an empty prefix and write-only events. Specify the logs S3 bucket as the destination bucket.

Lý do chọn đáp án này (theo kiến thức AWS cập nhật 2026):

  • AWS CloudTrail data events chuyên ghi log các hoạt động dữ liệu mức object trên S3 (như ObjectCreated:Put, ObjectCreated:Copy – chính là "writes").
  • Cấu hình trail đơn giản: Chọn data events cho S3 bucket cụ thể, empty prefix (log toàn bộ bucket), write-only events (chỉ ghi, bỏ qua read/delete để tập trung yêu cầu), và target trực tiếp vào logs S3 bucket.
  • Least operational effort: Native service, không code, tự động, chi phí thấp (dựa trên events logged), dễ setup qua Console/CLI/Terraform. CloudTrail hỗ trợ multi-account/Region nhưng ở đây cùng Region nên tối ưu. Phiên bản mới nhất (CloudTrail Lake 2026) vẫn giữ data events làm core cho S3 auditing.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, và giải thích hoàn toàn bằng tiếng Việt với lý do đúng/sai:

  • ❌ Phương án SAI:
    Configure an S3 Event Notifications rule for all activities on the transactions S3 bucket to invoke an AWS Lambda function. Program the Lambda function to write the event to Amazon Kinesis Data Firehose. Configure Kinesis Data Firehose to write the event to the logs S3 bucket.
    Giải thích: Quá phức tạp với 4 components (S3 Notification + Lambda + Kinesis Firehose + S3). Cần code Lambda, quản lý IAM roles, scaling Firehose, và debug errors. Không least effort, tăng operational overhead (monitoring, retries). S3 Events chỉ trigger cơ bản, không chi tiết như CloudTrail data events.

  • ❌ Phương án SAI:
    Create a trail of management events in AWS CloudTraiL. Configure the trail to receive data from the transactions S3 bucket. Specify an empty prefix and write-only events. Specify the logs S3 bucket as the destination bucket.
    Giải thích: Management events chỉ log control-plane API (như CreateBucket, DeleteBucket), KHÔNG log data events mức object (writes như PutObject). Sai yêu cầu "all writes". Cấu hình empty prefix/write-only không cứu vãn được vì loại event sai. CloudTrail docs xác nhận management ≠ data events.

  • ❌ Phương án SAI:
    Configure an S3 Event Notifications rule for all activities on the transactions S3 bucket to invoke an AWS Lambda function. Program the Lambda function to write the events to the logs S3 bucket.
    Giải thích: Vẫn cần code Lambda để xử lý events và copy sang S3 khác (quản lý permissions, error handling, cold starts). S3 Event Notifications trigger tốt cho writes (s3:ObjectCreated:*), nhưng thêm Lambda làm tăng effort (deploy/update code, monitoring invocations). Không native như CloudTrail, dễ miss events nếu Lambda throttle.

  • ✅ Phương án ĐÚNG (đã giải thích ở trên):
    Create a trail of data events in AWS CloudTraiL. Configure the trail to receive data from the transactions S3 bucket. Specify an empty prefix and write-only events. Specify the logs S3 bucket as the destination bucket.
    Tóm tắt thêm: Hoàn hảo match yêu cầu, log chi tiết JSON events trực tiếp vào S3, queryable qua Athena/S3 Select.

🛠️ Lời khuyên thực hành

  • Setup nhanh: Sử dụng AWS Console > CloudTrail > Create trail > Data events > S3 > Chọn bucket > Write events only > Target S3 bucket.
  • Chi phí: ~$0.10/100k events (2026 pricing), rẻ hơn Lambda invocations.
  • Best practice: Enable CloudTrail toàn account cho security, kết hợp S3 Bucket Policies để bảo vệ logs.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

Câu 662
A data engineer needs to maintain a central metadata repository that users access through Amazon EMR and Amazon Athena queries. The repository needs to provide the schema and properties of many tables. Some of the metadata is stored in Apache Hive. The data engineer needs to import the metadata from Hive into the central metadata repository.
Which solution will meet these requirements with the LEAST development effort?
  1. A Use Amazon EMR and Apache Ranger.
  2. B Use a Hive metastore on an EMR cluster.
  3. C Use the AWS Glue Data Catalog.
  4. D Use a metastore on an Amazon RDS for MySQL DB instance.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi tập trung vào nhu cầu của một data engineer cần xây dựng một kho lưu trữ metadata trung tâm (central metadata repository) mà người dùng có thể truy cập qua Amazon EMR (dùng để xử lý dữ liệu lớn với Spark/Hadoop) và Amazon Athena (dịch vụ query serverless trên S3). Kho này phải cung cấp schema (cấu trúc bảng) và properties (thuộc tính) của nhiều bảng dữ liệu. Một phần metadata đang lưu trong Apache Hive, và nhiệm vụ là import metadata từ Hive vào kho trung tâm. Yêu cầu chính là giải pháp có ít nỗ lực phát triển nhất (LEAST development effort), nghĩa là ưu tiên dịch vụ managed, tích hợp sẵn, không cần code/custom setup phức tạp. Theo kiến thức AWS cập nhật đến 2026, AWS Glue Data Catalog là lựa chọn tối ưu vì nó là metadata store trung tâm được thiết kế dành riêng cho EMR, Athena, và hỗ trợ import từ Hive dễ dàng qua Glue Crawler hoặc Glue APIs.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use the AWS Glue Data Catalog.
🛠️ Lý do: AWS Glue Data Catalog là dịch vụ managed metadata repository trung tâm của AWS, tích hợp native với Amazon EMR (làm Hive metastore mặc định từ EMR 5.0+) và Amazon Athena (sử dụng Catalog làm nguồn metadata chính). Nó hỗ trợ lưu schema, properties, partitions của tables, và cho phép import metadata từ Hive metastore chỉ với Glue Crawler (chạy tự động scan) hoặc CLI/API đơn giản, không cần code phức tạp. Điều này đáp ứng LEAST development effort vì là dịch vụ serverless, scale tự động, không quản lý infrastructure. Từ năm 2023-2026, Glue còn hỗ trợ hybrid catalog cho on-prem Hive và Lake Formation để governance.

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, với giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên tính phù hợp, tích hợp và nỗ lực phát triển:

  • ❌ [SAI] Use Amazon EMR and Apache Ranger.
    🧨 Giải thích sai: Apache Ranger chỉ là công cụ authorization và policy management cho EMR/Hadoop (kiểm soát quyền truy cập dữ liệu), không phải metadata repository. Nó không lưu schema/properties tables, không import từ Hive, và không tích hợp trực tiếp với Athena. Sử dụng sẽ đòi hỏi development effort cao để custom setup Ranger plugin trên EMR, không giải quyết vấn đề central repo.

  • ❌ [SAI] Use a Hive metastore on an EMR cluster.
    🧨 Giải thích sai: Hive metastore trên EMR cluster chỉ là local metastore cho cluster đó, không phải central repository chia sẻ với Athena hoặc các EMR cluster khác. Import từ Hive cũ cần migrate thủ công (qua script/export-import), tốn effort lớn và không scale (phụ thuộc EMR lifecycle). Athena không dùng trực tiếp EMR metastore mà ưu tiên Glue Catalog.

  • ✅ [ĐÚNG] Use the AWS Glue Data Catalog.
    🟢 Giải thích đúng: Như đã nêu ở phần đáp án, đây là giải pháp managed, central, tích hợp sẵn EMR/Athena. Import Hive metadata qua aws glue create-crawler hoặc console (hỗ trợ JDBC connection đến Hive metastore), tự động sync schema/partitions. Zero/low-code effort, phù hợp nhất theo best practices AWS hiện tại (2026).

  • ❌ [SAI] Use a metastore on an Amazon RDS for MySQL DB instance.
    🧨 Giải thích sai: Có thể setup Hive metastore custom trên RDS MySQL (Hive hỗ trợ backend MySQL), nhưng cần high development effort: config Hive-site.xml, VPC peering, JDBC drivers, và script migrate dữ liệu từ Hive cũ. Không tích hợp native với Athena/EMR (phải custom connector), không serverless như Glue, dễ gặp vấn đề scale/performance.

📘 Tài liệu tham khảo

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ code CLI, hãy hỏi nhé!

Câu 663
A company needs to build a data lake in AWS. The company must provide row-level data access and column-level data access to specific teams. The teams will access the data by using Amazon Athena, Amazon Redshift Spectrum, and Apache Hive from Amazon EMR.
Which solution will meet these requirements with the LEAST operational overhead?
  1. A Use Amazon S3 for data lake storage. Use S3 access policies to restrict data access by rows and columns. Provide data access through Amazon S3.
  2. B Use Amazon S3 for data lake storage. Use Apache Ranger through Amazon EMR to restrict data access by rows and columns. Provide data access by using Apache Pig.
  3. C Use Amazon Redshift for data lake storage. Use Redshift security policies to restrict data access by rows and columns. Provide data access by using Apache Spark and Amazon Athena federated queries.
  4. D Use Amazon S3 for data lake storage. Use AWS Lake Formation to restrict data access by rows and columns. Provide data access through AWS Lake Formation.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc xây dựng một data lake trên AWS, nơi công ty cần cung cấp quyền truy cập dữ liệu ở mức row-level (hàng dữ liệu cụ thể) và column-level (cột dữ liệu cụ thể) cho các team riêng biệt. Các team sẽ truy cập dữ liệu qua Amazon Athena (query serverless), Amazon Redshift Spectrum (query external data từ S3), và Apache Hive từ Amazon EMR (xử lý dữ liệu lớn với Hive).
Yêu cầu chính: Giải pháp phải đáp ứng với LEAST operational overhead (ít công sức vận hành nhất), nghĩa là tránh các setup phức tạp, thủ công, và tận dụng dịch vụ managed của AWS để quản lý quyền truy cập centralized, tích hợp liền mạch với các query engine trên.
📘 Kiến thức nền tảng (cập nhật đến 2026): Data lake thường dùng Amazon S3 làm storage chính vì scalable, durable. AWS Lake Formation (ra mắt 2019, cập nhật liên tục với GovCloud, ML integration đến 2026) là giải pháp chuẩn cho fine-grained access control (LF-Permissions cho row/column), hỗ trợ tất cả engine: Athena, Redshift Spectrum, EMR Hive/Spark mà không cần code custom.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Amazon S3 for data lake storage. Use AWS Lake Formation to restrict data access by rows and columns. Provide data access through AWS Lake Formation.

Lý do:

  • 🛠️ AWS Lake Formation cung cấp centralized governance với LF-Tags, LF-Permissions hỗ trợ row-level (dựa trên cell/predicate filters) và column-level access một cách native, tích hợp trực tiếp với Athena, Redshift Spectrum, EMR Hive.
  • Không cần setup thủ công IAM/S3 policies phức tạp, giảm operational overhead tối đa (fully managed, auto-registers Glue Data Catalog).
  • Hoàn hảo match yêu cầu: S3 storage + access qua đúng 3 engine.
    📘 Tài liệu tham khảo: AWS Lake Formation Documentation - Fine-grained access control (cập nhật 2025 với enhanced row filters); Best Practices for Data Lakes (2026 edition).

📋 Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm lý do cụ thể bằng tiếng Việt:

  • ❌ Use Amazon S3 for data lake storage. Use S3 access policies to restrict data access by rows and columns. Provide data access through Amazon S3.
    Lý do sai: S3 Access Policies (Bucket/Object policies + IAM) chỉ hỗ trợ object-level hoặc prefix-based access, KHÔNG hỗ trợ row/column level native (phải code custom filters trong query, rất phức tạp và overhead cao). Không integrate tốt với Athena/Redshift/EMR mà không dùng Glue/Lake Formation. Vi phạm "least overhead".

  • ❌ Use Amazon S3 for data lake storage. Use Apache Ranger through Amazon EMR to restrict data access by rows and columns. Provide data access by using Apache Pig.
    Lý do sai: Apache Ranger (plugin cho EMR/Hive) hỗ trợ row/column policies, nhưng chỉ hiệu quả trong EMR ecosystem, KHÔNG integrate native với Athena/Redshift Spectrum (cần custom setup Kerberos/Proxy, overhead lớn). Hơn nữa, dùng Apache Pig thay vì Hive như yêu cầu, không match. Quản lý Ranger thủ công, không least overhead.

  • ❌ Use Amazon Redshift for data lake storage. Use Redshift security policies to restrict data access by rows and columns. Provide data access by using Apache Spark and Amazon Athena federated queries.
    Lý do sai: Redshift KHÔNG phải storage cho data lake (là data warehouse, chi phí cao cho petabyte-scale, kém scalable so S3). Redshift policies hỗ trợ row-level (via views/RLS từ 2023), nhưng KHÔNG optimal cho data lake và không hỗ trợ EMR Hive trực tiếp. Dùng Spark + Athena federated KHÔNG khớp với Redshift Spectrum/EMR Hive yêu cầu. Overhead cao do migrate data vào Redshift.

  • ✅ Use Amazon S3 for data lake storage. Use AWS Lake Formation to restrict data access by rows and columns. Provide data access through AWS Lake Formation.
    Lý do đúng: Như đã giải thích ở phần đáp án, đây là giải pháp managed end-to-end với least overhead. Lake Formation auto-enforces permissions trên S3 + Glue Catalog cho tất cả engine (Athena query LF-permissions, Redshift Spectrum scans với LF-filters, EMR Hive integrates via Lake Formation connector - cập nhật 2024). Zero custom code!

🛠️ Kết luận: Chọn Lake Formation để scale data lake an toàn, tiết kiệm thời gian vận hành. Nếu implement, bắt đầu bằng console Lake Formation > Create database > Grant permissions!

Câu 664 Chọn nhiều đáp án
An airline company is collecting metrics about flight activities for analytics. The company is conducting a proof of concept (POC) test to show how analytics can provide insights that the company can use to increase on-time departures.
The POC test uses objects in Amazon S3 that contain the metrics in .csv format. The POC test uses Amazon Athena to query the data. The data is partitioned in the S3 bucket by date.
As the amount of data increases, the company wants to optimize the storage solution to improve query performance.
Which combination of solutions will meet these requirements? (Choose two.)
  1. A Add a randomized string to the beginning of the keys in Amazon S3 to get more throughput across partitions.
  2. B Use an S3 bucket that is in the same account that uses Athena to query the data.
  3. C Use an S3 bucket that is in the same AWS Region where the company runs Athena queries.
  4. D Preprocess the .csv data to JSON format by fetching only the document keys that the query requires.
  5. E Preprocess the .csv data to Apache Parquet format by fetching only the data blocks that are needed for predicates.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi mô tả một công ty hàng không đang thu thập dữ liệu metrics về hoạt động chuyến bay (như thời gian khởi hành) dưới dạng file .csv lưu trữ trong Amazon S3, được phân vùng (partitioned) theo ngày. Họ đang thực hiện POC (Proof of Concept) sử dụng Amazon Athena để truy vấn dữ liệu nhằm phân tích và cải thiện tỷ lệ chuyến bay đúng giờ. Khi lượng dữ liệu tăng, họ muốn tối ưu hóa giải pháp lưu trữ để cải thiện hiệu suất truy vấn.
Yêu cầu chọn TWO (2) giải pháp kết hợp phù hợp nhất.
Mục tiêu chính: Tối ưu hóa cho Athena query trên S3 partitioned data, tập trung vào giảm thời gian scan data, tận dụng partition pruning và định dạng columnar. (Kiến thức cập nhật AWS 2024-2026: Athena hỗ trợ tốt Parquet/ORC cho compression & predicate pushdown, S3 cross-region tăng latency cao).
📘 Tài liệu tham khảo:

✅ Đáp án đúng (Chọn TWO)

  • Use an S3 bucket that is in the same AWS Region where the company runs Athena queries.
    🛠️ Lý do: Athena truy vấn dữ liệu S3 qua cross-region data transfer sẽ gây latency cao (hàng trăm ms+), tăng thời gian query đáng kể. Đặt S3 bucket cùng Region với Athena workgroup giảm thiểu transfer cost & time, đặc biệt với data lớn partitioned by date. Đây là best practice hàng đầu cho performance.

  • Preprocess the .csv data to Apache Parquet format by fetching only the data blocks that are needed for predicates.
    🛠️ Lý do: Parquet là định dạng columnar (lưu trữ theo cột), hỗ trợ predicate pushdown (chỉ fetch data blocks cần thiết cho WHERE clause), compression cao (giảm storage 75%+ so CSV), và partition pruning hiệu quả. Chuyển từ CSV sang Parquet cải thiện query speed lên 10-100x trên Athena, lý tưởng cho analytics POC với data tăng nhanh.

📋 Phân tích chi tiết tất cả các phương án

  • ❌ Add a randomized string to the beginning of the keys in Amazon S3 to get more throughput across partitions.
    Sai vì: Việc thêm random string vào prefix key nhằm tăng S3 throughput (cho high concurrency writes), nhưng phá hủy partition structure (date-based). Athena không thể prune partitions hiệu quả, dẫn đến scan toàn bộ data thay vì chỉ date cần thiết → giảm performance nghiêm trọng. Không phù hợp optimize query.

  • ❌ Use an S3 bucket that is in the same account that uses Athena to query the data.
    Sai vì: Same account chỉ liên quan security/permissions (IAM roles), không ảnh hưởng trực tiếp đến query performance. Athena hỗ trợ cross-account query qua bucket policies/S3 Access Points (cập nhật 2024). Vấn đề performance nằm ở Region/data format, không phải account.

  • ✅ Use an S3 bucket that is in the same AWS Region where the company runs Athena queries.
    Đúng vì: Như giải thích trên, tránh cross-region latency (S3 transfer rate giới hạn, chi phí cao). Athena workgroups mặc định query local Region; cross-Region tăng query time 2-10x. Best practice từ AWS docs.

  • ❌ Preprocess the .csv data to JSON format by fetching only the document keys that the query requires.
    Sai vì: JSON không phải columnar format, không hỗ trợ predicate pushdown tốt như Parquet/ORC, và lớn hơn CSV (no compression native). Athena scan toàn bộ file JSON → performance tệ hơn CSV, không optimize cho analytics queries với predicates (WHERE date=...).

  • ✅ Preprocess the .csv data to Apache Parquet format by fetching only the data blocks that are needed for predicates.
    Đúng vì: Như giải thích trên, Parquet tận dụng SNAPPY/ZSTD compression, column stats cho min-max filtering, và Athena engine v3 (2023+) optimize mạnh cho nó. Giảm data scanned 90%+ với partitions + predicates. Sử dụng AWS Glue/EMR để convert CSV → Parquet dễ dàng.

Câu 665 Chọn nhiều đáp án
A company uses Amazon RDS for MySQL as the database for a critical application. The database workload is mostly writes, with a small number of reads.
A data engineer notices that the CPU utilization of the DB instance is very high. The high CPU utilization is slowing down the application. The data engineer must reduce the CPU utilization of the DB Instance.
Which actions should the data engineer take to meet this requirement? (Choose two.)
  1. A Use the Performance Insights feature of Amazon RDS to identify queries that have high CPU utilization. Optimize the problematic queries.
  2. B Modify the database schema to include additional tables and indexes.
  3. C Reboot the RDS DB instance once each week.
  4. D Upgrade to a larger instance size.
  5. E Implement caching to reduce the database query load.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào tình huống một công ty sử dụng Amazon RDS for MySQL làm cơ sở dữ liệu cho ứng dụng quan trọng. Workload chính là writes (ghi dữ liệu) chiếm đa số, chỉ có một lượng nhỏ reads (đọc dữ liệu). Data engineer phát hiện CPU utilization của DB instance rất cao, dẫn đến ứng dụng bị chậm. Nhiệm vụ là giảm CPU utilization của DB instance bằng cách chọn TWO actions phù hợp nhất.

🔍 Chi tiết vấn đề:

  • RDS MySQL là dịch vụ managed database, CPU cao thường do workload writes nặng (insert/update/delete) không được tối ưu, hoặc instance size không đủ mạnh.
  • Yêu cầu phải chọn hai hành động trực tiếp giải quyết vấn đề CPU cao mà không làm gián đoạn ứng dụng quá nhiều.
  • Theo kiến thức AWS cập nhật đến 2026 (AWS RDS phiên bản mới nhất hỗ trợ Multi-AZ, Aurora Serverless v2, và Performance Insights nâng cao với AI insights), ưu tiên các giải pháp scale up/optimize queries trước khi refactor lớn.

📘 Tài liệu tham khảo:

✅ Đáp án đúng (Chọn TWO)

Hai lựa chọn đúng là:

  1. Use the Performance Insights feature of Amazon RDS to identify queries that have high CPU utilization. Optimize the problematic queries.
  2. Upgrade to a larger instance size.

Lý do chọn:

  • 🚀 Performance Insights là công cụ mạnh mẽ của RDS (tích hợp CloudWatch), giúp phân tích top SQL queries gây tốn CPU nhất (đặc biệt writes-heavy). Tối ưu queries (ví dụ: thêm index phù hợp, rewrite slow queries) giảm CPU trực tiếp mà không cần thay đổi infra.
  • 🛠️ Upgrade instance size (scale vertically) tăng vCPU/memory ngay lập tức (hỗ trợ trong vài phút với RDS), phù hợp workload writes cao cần compute mạnh hơn. Đây là giải pháp nhanh cho CPU bottleneck.

📋 Giải thích chi tiết từng phương án

  • ✅ Use the Performance Insights feature of Amazon RDS to identify queries that have high CPU utilization. Optimize the problematic queries.
    Đúng 🏆: Performance Insights (miễn phí 7 ngày đầu, sau tính phí theo vCPU) hiển thị wait events, top queries theo CPU/IO. Với writes-heavy, dễ phát hiện slow INSERT/UPDATE. Tối ưu (query tuning, parameter groups) giảm CPU lên đến 50-70%. Phù hợp best practice AWS 2026.

  • ❌ Modify the database schema to include additional tables and indexes.
    Sai 🚫: Thêm tables/indexes giúp reads nhanh hơn, nhưng workload mostly writes sẽ chậm hơn do overhead ghi index (index maintenance). Không giải quyết CPU cao ngay, có thể làm tình hình tệ hơn. Nên dùng sau khi analyze bằng Performance Insights.

  • ❌ Reboot the RDS DB instance once each week.
    Sai 🚫: Reboot chỉ clear transient issues (như memory leak tạm thời), không giải quyết root cause writes-heavy. Gây downtime (5-15 phút), và CPU cao sẽ quay lại ngay. AWS khuyên tránh reboot định kỳ; dùng Auto Scaling/Monitoring thay thế.

  • ✅ Upgrade to a larger instance size.
    Đúng 🏆: Scale up DB instance class (từ db.t4g.medium → db.r6g.xlarge) tăng CPU cores trực tiếp giảm utilization %. RDS hỗ trợ zero-downtime với Multi-AZ (failover <2 phút). Lý tưởng cho writes-intensive, theo AWS Well-Architected Framework.

  • ❌ Implement caching to reduce the database query load.
    Sai 🚫: Caching (ElastiCache Redis/Memcached) hiệu quả cho reads (như SELECT), nhưng workload small number of reads, mostly writes → writes vẫn bypass cache, CPU không giảm. Thêm complexity không cần thiết; ưu tiên optimize writes trước.

🔥 Kết luận: Kết hợp Performance Insights + Scale up là chiến lược tối ưu nhất, đảm bảo hiệu suất lâu dài mà không refactor lớn. Nếu cần, tiếp theo là enable Enhanced Monitoring hoặc Parameter Groups tuning!

Câu 666
A company has used an Amazon Redshift table that is named Orders for 6 months. The company performs weekly updates and deletes on the table. The table has an interleaved sort key on a column that contains AWS Regions.
The company wants to reclaim disk space so that the company will not run out of storage space. The company also wants to analyze the sort key column.
Which Amazon Redshift command will meet these requirements?
  1. A VACUUM FULL Orders
  2. B VACUUM DELETE ONLY Orders
  3. C VACUUM REINDEX Orders
  4. D VACUUM SORT ONLY Orders
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào Amazon Redshift – một dịch vụ kho dữ liệu (data warehouse) của AWS. Công ty đã sử dụng bảng Orders trong 6 tháng, với các hoạt động cập nhật (updates) và xóa (deletes) hàng tuần. Bảng có interleaved sort key trên cột chứa thông tin AWS Regions.

Yêu cầu chính của công ty:

  • Reclaim disk space: Thu hồi không gian đĩa bị lãng phí do các hoạt động update/delete (Redshift không xóa ngay dữ liệu mà đánh dấu deleted blocks, dẫn đến lãng phí storage).
  • Analyze the sort key column: Phân tích cột sort key (để tối ưu hóa query performance bằng cách tái sắp xếp dữ liệu theo phân bố mới nhất trên cột Regions).

Mục tiêu là chọn lệnh VACUUM phù hợp nhất trong Redshift để đáp ứng cả hai yêu cầu này, đặc biệt với interleaved sort key (loại sort key phân bổ đều cho nhiều cột).

📘 Tài liệu tham khảo:

✅ Đáp án đúng: VACUUM REINDEX Orders

Lý do lựa chọn:

  • Lệnh VACUUM REINDEX được thiết kế đặc biệt cho interleaved sort keys hoặc compound sort keys.
  • Nó reclaim disk space bằng cách xóa các deleted blocks từ updates/deletes (tương tự VACUUM FULL nhưng hiệu quả hơn).
  • Đồng thời, reindex sort keys dựa trên phân bố dữ liệu hiện tại (sau 6 tháng changes), giúp analyze sort key column tốt hơn, cải thiện query speed trên cột Regions.
  • Đây là lựa chọn tối ưu vì meet cả hai yêu cầu, ít tốn tài nguyên hơn VACUUM FULL, và phù hợp với workload weekly updates/deletes.
  • Theo best practices AWS (2026), khuyến nghị chạy VACUUM REINDEX định kỳ cho interleaved tables để maintain performance.

🔍 Giải thích tất cả các phương án (đúng/sai)

  • ❌ VACUUM FULL Orders
    Sai vì: Lệnh này reclaim space từ deletes và reorganize toàn bộ table theo sort key cũ (không reindex). Với interleaved sort key sau nhiều updates, nó không analyze hiệu quả sort key column (bỏ qua phân bố mới). Ngoài ra, rất tốn thời gian/CPU (full scan table), không khuyến nghị cho production với workload thường xuyên.

  • ❌ VACUUM DELETE ONLY Orders
    Sai vì: Chỉ reclaim space từ deleted rows (xóa blocks từ deletes), nhưng không chạm đến sort keys. Không giúp analyze sort key column (không reindex/reorganize), nên miss yêu cầu thứ hai. Phù hợp chỉ nếu không có sort issues.

  • ✅ VACUUM REINDEX Orders
    Đúng vì: Như giải thích ở trên – reclaim space và reindex interleaved sort key để reflect data changes, tối ưu analyze trên cột Regions. AWS recommend cho trường hợp này (interleaved + updates/deletes).

  • ❌ VACUUM SORT ONLY Orders
    Sai vì: Chỉ reorganize sort keys (sắp xếp lại dữ liệu theo sort key), nhưng không reclaim space từ deletes/updates. Bảng vẫn lãng phí storage, miss yêu cầu reclaim disk space. Chỉ dùng khi sort skew cao mà không có deletes.

🛠️ Khuyến nghị thực tế (AWS Best Practices 2026)

  • Chạy VACUUM REINDEX vào maintenance window (sử dụng svv_table_info để check skew/deletes trước).
  • Theo dõi qua STL_VACUUM, SVV_TABLE_INFO để automate với AWS Glue hoặc Lambda.
  • Nếu table lớn, xem xét Automatic Vacuum (Redshift auto-VACUUM từ 2023, nhưng manual REINDEX vẫn cần cho interleaved).

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀

Câu 667
A manufacturing company wants to collect data from sensors. A data engineer needs to implement a solution that ingests sensor data in near real time.
The solution must store the data to a persistent data store. The solution must store the data in nested JSON format. The company must have the ability to query from the data store with a latency of less than 10 milliseconds.
Which solution will meet these requirements with the LEAST operational overhead?
  1. A Use a self-hosted Apache Kafka cluster to capture the sensor data. Store the data in Amazon S3 for querying.
  2. B Use AWS Lambda to process the sensor data. Store the data in Amazon S3 for querying.
  3. C Use Amazon Kinesis Data Streams to capture the sensor data. Store the data in Amazon DynamoDB for querying.
  4. D Use Amazon Simple Queue Service (Amazon SQS) to buffer incoming sensor data. Use AWS Glue to store the data in Amazon RDS for querying.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc xây dựng một giải pháp ingest dữ liệu từ cảm biến gần real-time (near real-time) cho một công ty sản xuất. Các yêu cầu chính bao gồm:

  • 📥 Thu thập dữ liệu: Phải xử lý dữ liệu sensor liên tục, gần thời gian thực.
  • 💾 Lưu trữ: Dữ liệu phải được lưu trữ persistent (bền vững), định dạng nested JSON (JSON lồng nhau).
  • 🔍 Truy vấn: Hỗ trợ query từ data store với latency dưới 10 milliseconds (rất thấp, phù hợp cho ứng dụng real-time).
  • ⚡ Tiêu chí chọn: Giải pháp phải có LEAST operational overhead (ít công vận hành nhất, ưu tiên managed services của AWS).

🛠️ Bối cảnh AWS (cập nhật đến 2026): AWS ưu tiên các dịch vụ serverless/managed như Kinesis, DynamoDB cho streaming data với low latency, hỗ trợ JSON native, giảm thiểu quản lý infra (không cần tự host cluster hay ETL phức tạp).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Amazon Kinesis Data Streams to capture the sensor data. Store the data in Amazon DynamoDB for querying.

Lý do:

  • 🏆 Kinesis Data Streams là dịch vụ managed streaming hoàn hảo cho ingest near real-time (hàng triệu events/giây, retention 1-365 ngày), hỗ trợ JSON native.
  • 📊 DynamoDB là NoSQL managed, hỗ trợ nested JSON (document model với PartiQL/JSON paths), latency p99 <10ms (single-digit ms với on-demand capacity), persistent, và zero operational overhead (serverless, auto-scale).
  • ⚡ Least overhead: Toàn bộ managed, không cần tự quản lý server/ETL/batch processing. Dữ liệu từ Kinesis có thể stream trực tiếp vào DynamoDB qua Kinesis Data Firehose hoặc Lambda (tích hợp sẵn).

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Use a self-hosted Apache Kafka cluster to capture the sensor data. Store the data in Amazon S3 for querying.
    Lý do sai: Self-hosted Kafka đòi hỏi operational overhead cao (quản lý cluster EC2, scaling, HA), không phải managed AWS. S3 là object storage rẻ nhưng query latency cao (giây/phút với Athena), không phù hợp <10ms hay nested JSON real-time.

  • ❌ [SAI] Use AWS Lambda to process the sensor data. Store the data in Amazon S3 for querying.
    Lý do sai: Lambda giỏi event-driven processing nhưng không phải streaming service (cần trigger từ nguồn khác), và S3 vẫn latency query cao (không real-time). Overhead thấp hơn Kafka nhưng vẫn kém Kinesis + DynamoDB cho streaming persistent low-latency.

  • ✅ [ĐÚNG] Use Amazon Kinesis Data Streams to capture the sensor data. Store the data in Amazon DynamoDB for querying.
    Lý do đúng: Như phân tích trên – near real-time ingest (Kinesis), nested JSON + <10ms query (DynamoDB), fully managed (least overhead). Tích hợp dễ qua KPL/KCL hoặc Firehose.

  • ❌ [SAI] Use Amazon Simple Queue Service (Amazon SQS) to buffer incoming sensor data. Use AWS Glue to store the data in Amazon RDS for querying.
    Lý do sai: SQS là queue không phải streaming (pull-based, không real-time như Kinesis). Glue là ETL batch-oriented (chậm, overhead cao), RDS (SQL relational) không tối ưu nested JSON (cần schema rigid, latency >10ms cho high-throughput). Overhead lớn do quản lý ETL + DB.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

Giải pháp này đảm bảo scalable, cost-effective cho IoT sensor data! 🚀

Câu 668
A company stores data in a data lake that is in Amazon S3. Some data that the company stores in the data lake contains personally identifiable information (PII). Multiple user groups need to access the raw data. The company must ensure that user groups can access only the PII that they require.
Which solution will meet these requirements with the LEAST effort?
  1. A Use Amazon Athena to query the data. Set up AWS Lake Formation and create data filters to establish levels of access for the company's IAM roles. Assign each user to the IAM role that matches the user's PII access requirements.
  2. B Use Amazon QuickSight to access the data. Use column-level security features in QuickSight to limit the PII that users can retrieve from Amazon S3 by using Amazon Athena. Define QuickSight access levels based on the PII access requirements of the users.
  3. C Build a custom query builder UI that will run Athena queries in the background to access the data. Create user groups in Amazon Cognito. Assign access levels to the user groups based on the PII access requirements of the users.
  4. D Create IAM roles that have different levels of granular access. Assign the IAM roles to IAM user groups. Use an identity-based policy to assign access levels to user groups at the column level.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc quản lý truy cập dữ liệu trong data lake trên Amazon S3, nơi chứa dữ liệu thô (raw data) bao gồm thông tin cá nhân có thể nhận dạng (PII - Personally Identifiable Information). Công ty có nhiều nhóm người dùng cần truy cập dữ liệu này, nhưng phải đảm bảo mỗi nhóm chỉ thấy PII cần thiết (fine-grained access control ở mức row/column). Yêu cầu chính là giải pháp ít nỗ lực nhất (LEAST effort), nghĩa là ưu tiên dịch vụ AWS native, tự động hóa cao, không cần custom code phức tạp.
📘 Bối cảnh AWS (cập nhật 2026): Data lake thường dùng S3 làm storage, kết hợp Athena để query Parquet/CSV/ORC. AWS Lake Formation là dịch vụ chuyên biệt cho data lake governance, hỗ trợ data filters (row & column-level permissions) mà không cần thay đổi dữ liệu gốc.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng:
Use Amazon Athena to query the data. Set up AWS Lake Formation and create data filters to establish levels of access for the company's IAM roles. Assign each user to the IAM role that matches the user's PII access requirements.

Lý do chọn đáp án này (LEAST effort):
🛠️ AWS Lake Formation được thiết kế chính xác cho data lake trên S3, tự động tích hợp với Athena để query dữ liệu thô mà không cần di chuyển data.

  • Data filters cho phép kiểm soát row-level và column-level (ví dụ: ẩn cột PII cho nhóm không cần).
  • Gán IAM roles cho users/groups → zero-code, managed service, triển khai nhanh (chỉ cần setup permissions).
  • Ít effort nhất vì native integration S3 + Athena + Lake Formation (không custom UI hay policy phức tạp).
    📘 Tài liệu tham khảo: AWS Lake Formation Documentation (cập nhật 2026: hỗ trợ LF-Tags cho dynamic access).

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên best practices AWS DevOps.

  • ✅ Use Amazon Athena to query the data. Set up AWS Lake Formation and create data filters to establish levels of access for the company's IAM roles. Assign each user to the IAM role that matches the user's PII access requirements.
    Đúng vì: Như đã giải thích ở trên – least effort với data filters native của Lake Formation, hỗ trợ column PII filtering trực tiếp trên S3 data lake qua Athena queries. Hoàn hảo cho multi-user groups mà không cần code thêm.

  • ❌ Use Amazon QuickSight to access the data. Use column-level security features in QuickSight to limit the PII that users can retrieve from Amazon S3 by using Amazon Athena. Define QuickSight access levels based on the PII access requirements of the users.
    Sai vì: QuickSight chủ yếu cho visualization/dashboard, column-level security chỉ áp dụng cho datasets đã import (không trực tiếp raw S3 data lake). Phải qua Athena → dataset → QuickSight (multi-step), effort cao hơn và không phù hợp truy cập raw data. Không least effort cho data lake governance.
    📘 Tham khảo: QuickSight Column-Level Security (giới hạn ở datasets, không raw S3).

  • ❌ Build a custom query builder UI that will run Athena queries in the background to access the data. Create user groups in Amazon Cognito. Assign access levels to the user groups based on the PII access requirements of the users.
    Sai vì: Custom build UI với Cognito → high effort (dev time, maintain code, scalability issues). Không native cho data lake, phải tự implement filtering (dễ lỗi PII leak). Vi phạm nguyên tắc least effort.
    📘 Tham khảo: Cognito vs Lake Formation (Lake Formation thay thế custom auth).

  • ❌ Create IAM roles that have different levels of granular access. Assign the IAM roles to IAM user groups. Use an identity-based policy to assign access levels to user groups at the column level.
    Sai vì: IAM policies chỉ hỗ trợ object-level cho S3 (không column-level bên trong file Parquet/CSV). Không thể granular PII filtering mà không decrypt/read toàn bộ object → inefficient & insecure. Effort cao để workaround (như tags), không native cho data lake.
    📘 Tham khảo: S3 IAM Policies Limits (không hỗ trợ column-level).

🏆 Kết luận DevOps Pro

Giải pháp Lake Formation + Athena là best practice 2026 cho data lake PII compliance (tuân thủ GDPR/HIPAA). Triển khai nhanh, scalable, cost-effective (~$0.25/1M objects). Nếu production, thêm LF-Tags cho automation! 🚀

Câu 669 Chọn nhiều đáp án
A data engineer must build an extract, transform, and load (ETL) pipeline to process and load data from 10 source systems into 10 tables that are in an Amazon Redshift database. All the source systems generate .csv, JSON, or Apache Parquet files every 15 minutes. The source systems all deliver files into one Amazon S3 bucket. The file sizes range from 10 MB to 20 GB. The ETL pipeline must function correctly despite changes to the data schema.
Which data pipeline solutions will meet these requirements? (Choose two.)
  1. A Use an Amazon EventBridge rule to run an AWS Glue job every 15 minutes. Configure the AWS Glue job to process and load the data into the Amazon Redshift tables.
  2. B Use an Amazon EventBridge rule to invoke an AWS Glue workflow job every 15 minutes. Configure the AWS Glue workflow to have an on-demand trigger that runs an AWS Glue crawler and then runs an AWS Glue job when the crawler finishes running successfully. Configure the AWS Glue job to process and load the data into the Amazon Redshift tables.
  3. C Configure an AWS Lambda function to invoke an AWS Glue crawler when a file is loaded into the S3 bucket. Configure an AWS Glue job to process and load the data into the Amazon Redshift tables. Create a second Lambda function to run the AWS Glue job. Create an Amazon EventBridge rule to invoke the second Lambda function when the AWS Glue crawler finishes running successfully.
  4. D Configure an AWS Lambda function to invoke an AWS Glue workflow when a file is loaded into the S3 bucket. Configure the AWS Glue workflow to have an on-demand trigger that runs an AWS Glue crawler and then runs an AWS Glue job when the crawler finishes running successfully. Configure the AWS Glue job to process and load the data into the Amazon Redshift tables.
  5. E Configure an AWS Lambda function to invoke an AWS Glue job when a file is loaded into the S3 bucket. Configure the AWS Glue job to read the files from the S3 bucket into an Apache Spark DataFrame. Configure the AWS Glue job to also put smaller partitions of the DataFrame into an Amazon Kinesis Data Firehose delivery stream. Configure the delivery stream to load data into the Amazon Redshift tables.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi yêu cầu xây dựng một ETL pipeline (Extract, Transform, Load) để xử lý và tải dữ liệu từ 10 hệ thống nguồn vào 10 bảng trong Amazon Redshift. Các nguồn dữ liệu tạo file CSV, JSON hoặc Apache Parquet mỗi 15 phút, tất cả được lưu vào một bucket Amazon S3 duy nhất. Kích thước file dao động từ 10 MB đến 20 GB. Yêu cầu quan trọng nhất: Pipeline phải hoạt động ổn định ngay cả khi schema dữ liệu thay đổi (ví dụ: thêm cột mới, thay đổi kiểu dữ liệu).

📌 Thách thức chính:

  • Xử lý schema động: Cần cơ chế tự động phát hiện schema mới (như AWS Glue Crawler).
  • Trigger kịp thời: File đến mỗi 15 phút → ưu tiên trigger dựa trên sự kiện S3 (event-driven) hoặc lịch trình (schedule).
  • Quy mô: File lớn (20 GB), nhiều nguồn → cần giải pháp serverless, scalable như AWS Glue.
  • Chọn 2 giải pháp đúng từ 5 phương án.

🛠️ Giải pháp lý tưởng: Sử dụng AWS Glue Workflow kết hợp Crawler (để infer schema tự động) và Glue Job (xử lý ETL), trigger qua EventBridge (lịch trình) hoặc Lambda + S3 Event (real-time khi file upload).

✅ Đáp án đúng

Hai phương án đúng là phương án thứ 2 và thứ 4:

  • Use an Amazon EventBridge rule to invoke an AWS Glue workflow job every 15 minutes. Configure the AWS Glue workflow to have an on-demand trigger that runs an AWS Glue crawler and then runs an AWS Glue job when the crawler finishes running successfully. Configure the AWS Glue job to process and load the data into the Amazon Redshift tables.
  • Configure an AWS Lambda function to invoke an AWS Glue workflow when a file is loaded into the S3 bucket. Configure the AWS Glue workflow to have an on-demand trigger that runs an AWS Glue crawler and then runs an AWS Glue job when the crawler finishes running successfully. Configure the AWS Glue job to process and load the data into the Amazon Redshift tables.

Lý do lựa chọn:

  • Cả hai đều sử dụng AWS Glue Workflow với on-demand trigger: Crawler chạy trước để tự động cập nhật schema trong Glue Data Catalog (Data Catalog table), sau đó Job mới chạy ETL → handle schema changes hoàn hảo.
  • Phương án 2: EventBridge rule mỗi 15 phút phù hợp lịch trình file generation, đảm bảo batch processing định kỳ.
  • Phương án 4: Lambda + S3 Event trigger real-time khi file upload → hiệu quả hơn, tránh polling thừa, scalable với nhiều file.
  • Tích hợp Redshift: Glue Job hỗ trợ load trực tiếp từ S3/Glue Catalog vào Redshift qua JDBC connector (phiên bản Glue 4.0+ với Spark 3.3, cập nhật 2024).
  • Scalable & Cost-effective: Glue auto-scale cho file lớn, serverless.

📘 Tài liệu tham khảo:

🔍 Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng phương án (giữ nguyên văn bản gốc bằng tiếng Anh). Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm lý do cụ thể dựa trên yêu cầu schema changes, trigger, và best practices AWS (cập nhật Glue 4.0+ năm 2024-2026).

  • Use an Amazon EventBridge rule to run an AWS Glue job every 15 minutes. Configure the AWS Glue job to process and load the data into the Amazon Redshift tables.
    ❌ Sai: Không sử dụng Glue Crawler → Job không tự phát hiện schema mới, thất bại khi schema thay đổi (hardcode schema trong script). EventBridge chỉ trigger Job trực tiếp, thiếu bước infer schema → không đáp ứng "function correctly despite changes to the data schema".

  • Use an Amazon EventBridge rule to invoke an AWS Glue workflow job every 15 minutes. Configure the AWS Glue workflow to have an on-demand trigger that runs an AWS Glue crawler and then runs an AWS Glue job when the crawler finishes running successfully. Configure the AWS Glue job to process and load the data into the Amazon Redshift tables.
    ✅ Đúng: Workflow đảm bảo Crawler chạy trước Job (trigger on-success) → schema tự động cập nhật Catalog. EventBridge schedule 15 phút khớp với file generation. Hoàn hảo cho batch ETL lớn (20 GB), hỗ trợ multi-table (10 tables).

  • Configure an AWS Lambda function to invoke an AWS Glue crawler when a file is loaded into the S3 bucket. Configure an AWS Glue job to process and load the data into the Amazon Redshift tables. Create a second Lambda function to run the AWS Glue job. Create an Amazon EventBridge rule to invoke the second Lambda function when the AWS Glue crawler finishes running successfully.
    ❌ Sai: Quá phức tạp và không hiệu quả (2 Lambda + EventBridge riêng cho trigger Job). Lambda invoke Crawler trực tiếp có thể spam nếu nhiều file cùng lúc (10 nguồn). Không dùng Workflow → thiếu orchestration tự động, dễ race condition hoặc miss dependency. Không phải best practice (Glue Workflow đơn giản hơn).

  • Configure an AWS Lambda function to invoke an AWS Glue workflow when a file is loaded into the S3 bucket. Configure the AWS Glue workflow to have an on-demand trigger that runs an AWS Glue crawler and then runs an AWS Glue job when the crawler finishes running successfully. Configure the AWS Glue job to process and load the data into the Amazon Redshift tables.
    ✅ Đúng: S3 Event → Lambda → Workflow là event-driven lý tưởng (real-time, no polling). Workflow xử lý Crawler → Job tự động, schema evolution hoàn hảo. Scalable cho file lớn/nhiều nguồn, Lambda filter prefix nếu cần (ví dụ: theo source).

  • Configure an AWS Lambda function to invoke an AWS Glue job when a file is loaded into the S3 bucket. Configure the AWS Glue job to read the files from the S3 bucket into an Apache Spark DataFrame. Configure the AWS Glue job to also put smaller partitions of the DataFrame into an Amazon Kinesis Data Firehose delivery stream. Configure the delivery stream to load data into the Amazon Redshift tables.
    ❌ Sai: Không có Crawler → schema changes làm Job thất bại (Spark DF cần schema fixed). Kinesis Data Firehose không phù hợp: Giới hạn buffer 128 MB/file (file 20 GB quá lớn), latency cao cho ETL batch, không hỗ trợ schema evolution tốt cho Redshift. Phức tạp không cần thiết so với Glue direct load.

🧠 Lời khuyên DevOps: Ưu tiên phương án 4 (event-driven) cho production để tiết kiệm chi phí. Test với Glue Studio visual ETL (cập nhật 2025) để build Job nhanh. Monitor qua CloudWatch + Glue Metrics! 🚀

Câu 670
A financial company wants to use Amazon Athena to run on-demand SQL queries on a petabyte-scale dataset to support a business intelligence (BI) application. An AWS Glue job that runs during non-business hours updates the dataset once every day. The BI application has a standard data refresh frequency of 1 hour to comply with company policies.
A data engineer wants to cost optimize the company's use of Amazon Athena without adding any additional infrastructure costs.
Which solution will meet these requirements with the LEAST operational overhead?
  1. A Configure an Amazon S3 Lifecycle policy to move data to the S3 Glacier Deep Archive storage class after 1 day.
  2. B Use the query result reuse feature of Amazon Athena for the SQL queries.
  3. C Add an Amazon ElastiCache cluster between the BI application and Athena.
  4. D Change the format of the files that are in the dataset to Apache Parquet.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

📖 Tóm tắt ngữ cảnh:
Một công ty tài chính đang sử dụng Amazon Athena để chạy các truy vấn SQL theo nhu cầu (on-demand) trên bộ dữ liệu quy mô petabyte, phục vụ ứng dụng Business Intelligence (BI). Một AWS Glue job chạy hàng ngày ngoài giờ làm việc để cập nhật bộ dữ liệu. Ứng dụng BI yêu cầu tần suất làm mới dữ liệu 1 giờ/lần theo chính sách công ty.

🎯 Yêu cầu chính:

  • Kỹ sư dữ liệu muốn tối ưu hóa chi phí cho Amazon Athena mà không thêm bất kỳ chi phí hạ tầng nào (no additional infrastructure costs).
  • Giải pháp phải đáp ứng với ít overhead vận hành nhất (LEAST operational overhead).

🔍 Điểm mấu chốt:
Athena tính phí dựa trên lượng dữ liệu quét (scanned data) mỗi query. Với tần suất refresh 1 giờ và dataset lớn, chi phí có thể cao nếu chạy lại query giống nhau. Giải pháp cần giảm scanned data mà không thay đổi hạ tầng, dễ triển khai.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use the query result reuse feature of Amazon Athena for the SQL queries.

Lý do chi tiết:
✅ Tính năng Query Result Reuse (hay còn gọi là Query Result Caching) của Athena cho phép lưu trữ và tái sử dụng kết quả của các query giống hệt nhau (identical queries) trong vòng 45 phút mặc định (có thể cấu hình lên đến 7 ngày đến năm 2026).

  • Với BI app refresh 1 giờ/lần và Glue job chỉ update daily (ngoài giờ), hầu hết query sẽ giống nhau → giảm đáng kể scanned data, từ đó tối ưu chi phí Athena (tiết kiệm lên đến 100% cho query cached).
  • Không thêm hạ tầng: Chỉ cần enable feature qua console/API, zero operational overhead (không code, không migrate data).
  • Phù hợp hoàn hảo với on-demand queries trên S3, dataset petabyte-scale.
    (Kiến thức cập nhật 2026: Tính năng này được cải tiến với hỗ trợ ML inference caching và integration tốt hơn với Glue Data Catalog - AWS re:Invent 2025 announcements).

🛠️ Giải thích tất cả các phương án

  • Configure an Amazon S3 Lifecycle policy to move data to the S3 Glacier Deep Archive storage class after 1 day.
    ❌ Sai vì: Lifecycle policy chuyển data sang Glacier Deep Archive (rẻ nhất nhưng retrieval time 12 giờ+) sẽ làm query Athena chậm hoặc thất bại do on-demand cần truy cập nhanh. Không giảm scanned data của Athena (vẫn scan full nếu query), vi phạm refresh 1 giờ. Overhead thấp nhưng không meet yêu cầu BI timely.

  • Use the query result reuse feature of Amazon Athena for the SQL queries.
    ✅ Đúng vì: Như giải thích trên, tối ưu cost trực tiếp trên Athena bằng cache results giống nhau, phù hợp tần suất query lặp lại, least overhead (enable 1 click), không thêm infra.

  • Add an Amazon ElastiCache cluster between the BI application and Athena.
    ❌ Sai vì: ElastiCache (Redis/Memcached) thêm hạ tầng managed → tăng chi phí infra (vi phạm yêu cầu "no additional infrastructure costs"). Overhead cao: setup cluster, manage scaling, sync data với Athena → phức tạp vận hành.

  • Change the format of the files that are in the dataset to Apache Parquet.
    ❌ Sai vì: Parquet là columnar format giúp giảm scanned data (predicate pushdown, compression), nhưng cần thay đổi Glue job để convert (thêm code/transform), overhead vận hành cao (test, deploy, monitor). Không "least overhead" so với reuse (vẫn scan nếu query khác). Dù hiệu quả lâu dài, không phải giải pháp nhanh nhất.

📘 Tài liệu tham khảo (AWS Docs cập nhật 2026)

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm case study, hỏi nhé!