Ngân hàng đề — AWS Certified Data Engineer Associate
Tìm thấy 867 câu.
The company runs a daily report on the S3 data. Some days, the company runs the report before all the daily data has been uploaded to the S3 bucket. A data engineer must be able to send a message that identifies any incomplete data to an existing Amazon Simple Notification Service (Amazon SNS) topic.
Which solution will meet this requirement with the LEAST operational overhead?
- A Create data quality checks for the source datasets that the daily reports use. Create a new AWS managed Apache Airflow cluster. Run the data quality checks by using Airflow tasks that run data quality queries on the columns data type and the presence of null values. Configure Airflow Directed Acyclic Graphs (DAGs) to send an email notification that informs the data engineer about the incomplete datasets to the SNS topic.
- B Create data quality checks on the source datasets that the daily reports use. Create a new Amazon EMR cluster. Use Apache Spark SQL to create Apache Spark jobs in the EMR cluster that run data quality queries on the columns data type and the presence of null values. Orchestrate the ETL pipeline by using an AWS Step Functions workflow. Configure the workflow to send an email notification that informs the data engineer about the incomplete datasets to the SNS topic.
- C Create data quality checks on the source datasets that the daily reports use. Create data quality actions by using AWS Glue workflows to confirm the completeness and consistency of the datasets. Configure the data quality actions to create an event in Amazon EventBridge if a dataset is incomplete. Configure EventBridge to send the event that informs the data engineer about the incomplete datasets to the Amazon SNS topic.
- D Create AWS Lambda functions that run data quality queries on the columns data type and the presence of null values. Orchestrate the ETL pipeline by using an AWS Step Functions workflow that runs the Lambda functions. Configure the Step Functions workflow to send an email notification that informs the data engineer about the incomplete datasets to the SNS topic.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh một công ty sử dụng AWS Glue Data Catalog để lập chỉ mục (index) dữ liệu được tải lên Amazon S3 hàng ngày. Họ có quy trình ETL batch hàng ngày (extract, transform, load) để tải dữ liệu từ nguồn bên ngoài vào bucket S3. Công ty chạy báo cáo hàng ngày trên dữ liệu S3, nhưng đôi khi báo cáo được chạy trước khi toàn bộ dữ liệu hàng ngày được tải lên đầy đủ, dẫn đến dữ liệu không hoàn chỉnh (incomplete data). Yêu cầu là data engineer phải gửi thông báo (message) xác định dữ liệu không hoàn chỉnh đến một Amazon SNS topic đã tồn tại. Giải pháp cần ít overhead vận hành nhất (LEAST operational overhead), nghĩa là ưu tiên các dịch vụ managed, tự động hóa cao, không cần quản lý hạ tầng thủ công nhiều.
📘 Tài liệu tham khảo: AWS Glue Data Catalog & ETL (https://docs.aws.amazon.com/glue/latest/dg/what-is-glue.html), Amazon S3 & SNS integration (https://docs.aws.amazon.com/sns/latest/dgg/what-is.html) – cập nhật đến AWS re:Invent 2025.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create data quality checks on the source datasets that the daily reports use. Create data quality actions by using AWS Glue workflows to confirm the completeness and consistency of the datasets. Configure the data quality actions to create an event in Amazon EventBridge if a dataset is incomplete. Configure EventBridge to send the event that informs the data engineer about the incomplete datasets to the Amazon SNS topic.
Lý do chọn đáp án này 🛠️:
- Đây là giải pháp fully managed từ AWS Glue Data Quality (tính năng ra mắt 2022 và cập nhật mạnh mẽ đến 2026), cho phép tạo data quality checks (kiểm tra độ hoàn chỉnh, nhất quán dữ liệu) trực tiếp trên datasets trong Glue Data Catalog mà không cần code thủ công hay quản lý cluster.
- Sử dụng Glue workflows để orchestrate, kết hợp data quality actions tự động trigger Amazon EventBridge event nếu dữ liệu incomplete → EventBridge route event đến SNS topic hiện có.
- Least operational overhead: Không cần provision cluster, code Lambda, hoặc DAGs phức tạp; chỉ config rules và actions qua console/CLI/API. Tích hợp native với S3/Glue Catalog, scale tự động, chi phí thấp.
📘 Nguồn: AWS Glue Data Quality (https://docs.aws.amazon.com/glue/latest/dg/data-quality.html), Glue Workflows & EventBridge (https://docs.aws.amazon.com/glue/latest/dg/aws-glue-api-workflow.html) – phiên bản 2026 hỗ trợ advanced rules như row count, null checks.
📋 Phân tích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể bằng tiếng Việt, dựa trên overhead vận hành và tính phù hợp.
-
Phương án 1: Create data quality checks for the source datasets that the daily reports use. Create a new AWS managed Apache Airflow cluster. Run the data quality checks by using Airflow tasks that run data quality queries on the columns data type and the presence of null values. Configure Airflow Directed Acyclic Graphs (DAGs) to send an email notification that informs the data engineer about the incomplete datasets to the SNS topic.
❌ Sai: Giải pháp này tạo MWAA (Managed Workflows for Apache Airflow) cluster mới, yêu cầu quản lý DAGs, tasks, scheduling thủ công (code Python/SQL queries). Overhead cao vì phải maintain Airflow environment, debug DAGs, và integrate SNS (email notification không trực tiếp là SNS message). Không native với Glue Catalog, không least overhead.
🛠️ Vấn đề chính: Airflow tốt cho orchestration phức tạp nhưng overkill cho simple data quality check. -
Phương án 2: Create data quality checks on the source datasets that the daily reports use. Create a new Amazon EMR cluster. Use Apache Spark SQL to create Apache Spark jobs in the EMR cluster that run data quality queries on the columns data type and the presence of null values. Orchestrate the ETL pipeline by using an AWS Step Functions workflow. Configure the workflow to send an email notification that informs the data engineer about the incomplete datasets to the SNS topic.
❌ Sai: Tạo EMR cluster mới (provision EC2 instances, Spark jobs) để chạy queries → overhead cao (quản lý cluster lifecycle, scaling, cost). Step Functions orchestrate tốt nhưng kết hợp EMR làm phức tạp, không fully managed cho data quality. Notification qua email thay vì SNS trực tiếp, không tối ưu. Không tận dụng Glue native features.
🛠️ Vấn đề chính: EMR phù hợp big data processing nhưng thừa thãi cho daily checks trên S3/Glue, vi phạm "least overhead". -
Phương án 3 (Đúng): Create data quality checks on the source datasets that the daily reports use. Create data quality actions by using AWS Glue workflows to confirm the completeness and consistency of the datasets. Configure the data quality actions to create an event in Amazon EventBridge if a dataset is incomplete. Configure EventBridge to send the event that informs the data engineer about the incomplete datasets to the Amazon SNS topic.
✅ Đúng: Như đã giải thích ở trên, serverless, zero-management với Glue Data Quality rules (ví dụ:rowCount > expectedcho completeness). Actions trigger EventBridge → SNS seamless, no code needed. Hoàn hảo cho scenario detect incomplete data trước report.
🛠️ Ưu điểm nổi bật: Tích hợp sâu với S3/Glue Catalog, auto-scale, monitor qua CloudWatch. -
Phương án 4: Create AWS Lambda functions that run data quality queries on the columns data type and the presence of null values. Orchestrate the ETL pipeline by using an AWS Step Functions workflow that runs the Lambda functions. Configure the Step Functions workflow to send an email notification that informs the data engineer about the incomplete datasets to the SNS topic.
❌ Sai: Yêu cầu code Lambda functions thủ công (queries SQL/Athena?), deploy/maintain code → overhead trung bình (debug, permissions S3/Glue). Step Functions orchestrate tốt nhưng không native data quality như Glue; notification email không khớp SNS message. Phù hợp hơn cho custom logic, nhưng không least overhead so với Glue.
🛠️ Vấn đề chính: Custom code tăng maintenance, thiếu built-in checks cho datasets.
Kết luận 🎯: Giải pháp đúng tận dụng AWS-native managed services (Glue DQ + EventBridge + SNS) để minimize ops, phù hợp best practices DevOps trên AWS đến 2026!
The marketing team should have access to obfuscated claim information but should have full access to customer contact information. The claims team should have access to customer information for each claim that the team processes. The analytics team should have access only to obfuscated PII data.
Which solution will enforce these data access requirements with the LEAST administrative overhead?
- A Create a separate Redshift cluster for each team. Load only the required data for each team. Restrict access to clusters based on the teams.
- B Create views that include required fields for each of the data requirements. Grant the teams access only to the view that each team requires.
- C Create a separate Amazon Redshift database role for each team. Define masking policies that apply for each team separately. Attach appropriate masking policies to each team role.
- D Move the customer data to an Amazon S3 bucket. Use AWS Lake Formation to create a data lake. Use fine-grained security capabilities to grant each team appropriate permissions to access the data.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào bảo mật dữ liệu PII (Personally Identifiable Information) trong Amazon Redshift cluster, nơi lưu trữ dữ liệu khách hàng. Công ty có ba team với quyền truy cập khác nhau:
- Marketing team 🏢: Truy cập obfuscated (ẩn danh) thông tin claim, nhưng full access vào thông tin liên lạc khách hàng.
- Claims team ⚖️: Truy cập full thông tin khách hàng chỉ cho các claim mà team xử lý.
- Analytics team 📊: Chỉ truy cập obfuscated PII data.
Yêu cầu giải pháp enforce data access requirements với LEAST administrative overhead (ít công quản trị nhất), nghĩa là tránh duplicate data, di chuyển lớn hoặc cấu hình phức tạp, ưu tiên tính năng native của Redshift để mask/hide dữ liệu động mà không copy dữ liệu.
Mục tiêu chính: Row/Column-level security + dynamic masking, phù hợp với Redshift Data Masking Policies (tính năng mới từ AWS 2023, cập nhật đến 2026 vẫn là best practice cho multi-tenant access trong single cluster).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create a separate Amazon Redshift database role for each team. Define masking policies that apply for each team separately. Attach appropriate masking policies to each team role.
Lý do:
- Giải pháp này sử dụng Redshift database roles và masking policies native (ra mắt 2023, fully supported đến 2026) để dynamic mask dữ liệu dựa trên role của user/team.
- LEAST overhead 🛠️: Không cần tạo cluster/view riêng, không di chuyển data; chỉ define policy một lần và attach role (attach/detach dễ dàng).
- Phù hợp chính xác: Marketing thấy claim obfuscated nhưng contact full; Claims thấy full cho claim của họ (kết hợp row-level filter nếu cần); Analytics chỉ obfuscated PII.
- Overhead thấp: Quản lý central trong 1 cluster, scale tự động, audit dễ qua IAM/Redshift logs.
🔍 Phân tích tất cả các phương án (đúng/sai)
-
❌ Phương án SAI: Create a separate Redshift cluster for each team. Load only the required data for each team. Restrict access to clusters based on the teams.
Giải thích: Tạo 3 cluster riêng dẫn đến high overhead (duplicate storage/cost, ETL pipeline phức tạp để load data filtered, sync data giữa clusters). Không scale tốt, vi phạm LEAST admin effort; Redshift không khuyến khích multi-cluster cho access control (dùng single cluster + policies tốt hơn theo AWS best practices 2026). -
❌ Phương án SAI: Create views that include required fields for each of the data requirements. Grant the teams access only to the view that each team requires.
Giải thích: Views chỉ filter columns/rows tĩnh, không hỗ trợ dynamic masking/obfuscation (ví dụ: marketing cần full contact nhưng obfuscated claim – views khó enforce linh hoạt mà không duplicate logic). Overhead cao do maintain nhiều views + grants; không handle row-level per claim cho Claims team hiệu quả. Redshift ưu tiên masking policies over views cho PII (cập nhật 2024+). -
✅ Phương án ĐÚNG: Create a separate Amazon Redshift database role for each team. Define masking policies that apply for each team separately. Attach appropriate masking policies to each team role.
Giải thích: Như đã nêu ở trên, native Redshift feature với zero-copy masking (mask at query time). Hỗ trợ regex/replace/partial mask cho PII; kết hợp row-level security (RLS) cho Claims team. Least overhead: 1 cluster, policy reusable, RBAC đơn giản. Best practice AWS DOP-C02 (2026). -
❌ Phương án SAI: Move the customer data to an Amazon S3 bucket. Use AWS Lake Formation to create a data lake. Use fine-grained security capabilities to grant each team appropriate permissions to access the data.
Giải thích: Di chuyển data sang S3 + Lake Formation là migration lớn (ETL từ Redshift → S3, refactor queries analytics). Overhead cao: Setup Lake Formation governance, LF Tags/Permissions phức tạp hơn Redshift native; không phù hợp nếu workload chính vẫn là OLAP queries trên Redshift (Lake Formation tốt cho raw data lake, không thay thế Redshift cluster).
📘 Tài liệu tham khảo (AWS cập nhật mới nhất 2026)
- Redshift Data Masking Policies: AWS Docs - Manage data masking policies (ra mắt GA 2023, enhanced 2025 với AI-driven mask).
- Redshift Security Best Practices: AWS DOP-C02 Exam Guide – Section: Redshift Fine-Grained Access Control.
- Blog AWS: Introducing dynamic data masking in Amazon Redshift (2023, vẫn core feature 2026).
- Well-Architected Framework: Security Pillar – Use native masking over custom solutions for PII.
Giải pháp này đảm bảo compliance GDPR/CCPA với least privilege! 🚀
A few days after the company added the new topic, Amazon CloudWatch raised an alarm on the RootDiskUsed metric for the MSK cluster.
How should the company address the CloudWatch alarm?
- A Expand the storage of the MSK broker. Configure the MSK cluster storage to expand automatically.
- B Expand the storage of the Apache ZooKeeper nodes.
- C Update the MSK broker instance to a larger instance type. Restart the MSK cluster.
- D Specify the Target Volume-in-GiB parameter for the existing topic.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả tình huống một công ty tài chính thêm tính năng mới vào ứng dụng di động, dẫn đến việc tạo một topic mới trong cluster Amazon Managed Streaming for Apache Kafka (Amazon MSK) hiện có. Sau vài ngày, Amazon CloudWatch kích hoạt báo động trên metric RootDiskUsed của cluster MSK. Metric này đo lường mức độ sử dụng dung lượng đĩa gốc (root disk) trên các broker nodes trong cluster Kafka.
Vấn đề cốt lõi: Topic mới có thể làm tăng lượng dữ liệu lưu trữ (log segments, indexes) trên broker nodes, dẫn đến hết dung lượng đĩa (disk space). Câu hỏi yêu cầu cách xử lý báo động CloudWatch một cách hiệu quả, tập trung vào việc mở rộng storage mà không gây gián đoạn lớn cho dịch vụ streaming thời gian thực. Đây là tình huống phổ biến trong MSK khi dữ liệu tăng đột biến, và AWS khuyến nghị sử dụng tính năng storage auto-scaling (cập nhật từ năm 2023, ổn định đến 2026).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Expand the storage of the MSK broker. Configure the MSK cluster storage to expand automatically.
Lý do:
- Metric RootDiskUsed trực tiếp liên quan đến dung lượng đĩa trên broker nodes của MSK (nơi lưu trữ dữ liệu topics). Khi thêm topic mới, dữ liệu tích tụ nhanh chóng gây đầy đĩa.
- AWS MSK hỗ trợ mở rộng storage thủ công hoặc tự động (auto-scaling) cho broker nodes từ dung lượng cơ bản (1000 GiB) lên đến 16 TiB/broker (cập nhật mới nhất 2026). Tính năng này không yêu cầu downtime, cluster vẫn hoạt động bình thường trong quá trình mở rộng.
- Cấu hình auto-expand đảm bảo storage tự động tăng khi disk usage vượt ngưỡng (ví dụ: 80%), ngăn ngừa báo động tương lai. Đây là giải pháp tối ưu, scalable theo best practices của AWS cho MSK clusters.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, với lý do đúng/sai dựa trên tài liệu AWS MSK mới nhất:
-
✅ Expand the storage of the MSK broker. Configure the MSK cluster storage to expand automatically.
🛠️ Đúng vì: Như đã giải thích, đây là cách chính thức để xử lý RootDiskUsed. MSK cho phép resize EBS volumes của broker mà không restart, và auto-scaling dựa trên CloudWatch alarms tự động mở rộng (tăng 5% mỗi lần, max 16 TiB). Giải quyết gốc rễ vấn đề disk space từ topic mới. -
❌ Expand the storage of the Apache ZooKeeper nodes.
🧩 Sai vì: ZooKeeper nodes trong MSK chỉ quản lý metadata và coordination (như topic configs, partition leaders), không lưu trữ dữ liệu messages/topics. RootDiskUsed là metric của broker nodes, không phải ZooKeeper. Mở rộng ZooKeeper không ảnh hưởng đến disk usage của brokers và là lãng phí tài nguyên. -
❌ Update the MSK broker instance to a larger instance type. Restart the MSK cluster.
🛠️ Sai vì: MSK không hỗ trợ thay đổi instance type sau khi tạo cluster (immutable field). Việc này yêu cầu tạo cluster mới và migrate dữ liệu, gây downtime lớn (hàng giờ/ngày) và rủi ro mất dữ liệu. RootDiskUsed chủ yếu do storage size, không phải CPU/RAM (instance type), nên resize storage hiệu quả hơn mà không restart. -
❌ Specify the Target Volume-in-GiB parameter for the existing topic.
📘 Sai vì: Target Volume-in-GiB là tham số của Amazon EBS volumes trong EC2 (cho auto-provisioned IOPS), không áp dụng cho MSK topics. Topics trong Kafka/MSK không có tham số storage riêng; storage được quản lý ở mức cluster/broker. Thao tác này sẽ lỗi hoặc không có hiệu lực.
📘 Tài liệu tham khảo
- AWS MSK Documentation - Cluster Storage Management (cập nhật 2026): https://docs.aws.amazon.com/msk/latest/developerguide/msk-storage.html – Chi tiết auto-scaling và resize broker storage.
- CloudWatch Metrics for MSK: https://docs.aws.amazon.com/msk/latest/developerguide/cloudwatch-metrics.html – Giải thích RootDiskUsed và alarms.
- MSK Best Practices: AWS Well-Architected Framework for Streaming (2026 edition) khuyến nghị auto-scaling storage cho workloads tăng trưởng như mobile apps.
Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ thực tế, hãy hỏi nhé!
Which solution will meet these requirements with the LEAST effort?
- A Use an AWS Glue crawler to scan the S3 buckets and RDS databases and build a data catalog. Use data stewards to inspect the data and update the data catalog with the data format.
- B Use an AWS Glue crawler to build a data catalog. Use AWS Glue crawler classifiers to recognize the format of data and store the format in the catalog.
- C Use Amazon Macie to build a data catalog and to identify sensitive data elements. Collect the data format information from Macie.
- D Use scripts to scan data elements and to assign data classifications based on the format of the data.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi yêu cầu xây dựng một enterprise data catalog (danh mục dữ liệu doanh nghiệp) dựa trên các Amazon S3 buckets (nơi lưu trữ dữ liệu phi cấu trúc) và Amazon RDS databases (cơ sở dữ liệu quan hệ). Danh mục này phải bao gồm metadata về storage format (định dạng lưu trữ dữ liệu, ví dụ: CSV, Parquet, JSON, Avro...). Giải pháp cần LEAST effort (ít nỗ lực nhất), nghĩa là ưu tiên tự động hóa cao, không cần can thiệp thủ công nhiều.
🛠️ Bối cảnh AWS: AWS Glue Data Catalog là kho metadata trung tâm cho dữ liệu trên S3 và RDS. Crawler của Glue có thể tự động quét dữ liệu từ S3 (hỗ trợ nhiều format) và RDS (qua JDBC connector), đồng thời sử dụng classifiers để nhận diện định dạng tự động. Điều này phù hợp với kiến thức cập nhật đến 2026 (AWS Glue version 4.0+ hỗ trợ classifiers nâng cao cho ML/structured data).
📘 Tài liệu tham khảo:
- AWS Glue Crawlers & Classifiers (cập nhật 2025).
- AWS Glue Data Catalog Best Practices.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use an AWS Glue crawler to build a data catalog. Use AWS Glue crawler classifiers to recognize the format of data and store the format in the catalog.
Lý do 🏆:
- AWS Glue Crawler tự động quét S3 buckets (hỗ trợ Parquet, ORC, JSON...) và RDS (qua JDBC), tạo schema/table trong Glue Data Catalog với least effort (chỉ cần config IAM role và schedule).
- Classifiers (built-in hoặc custom) tự động nhận diện storage format (groks, XML, CSV...) và lưu metadata trực tiếp vào catalog → Không cần code/script thủ công, hoàn toàn tự động hóa.
- Đây là giải pháp chuẩn AWS cho data catalog enterprise, tích hợp Lake Formation/Athena/EMR, tiết kiệm chi phí và thời gian nhất (re:Post AWS 2025 xác nhận classifiers hỗ trợ >90% format phổ biến).
📋 Phân tích tất cả các phương án (đúng/sai)
-
Phương án A: Use an AWS Glue crawler to scan the S3 buckets and RDS databases and build a data catalog. Use data stewards to inspect the data and update the data catalog with the data format.
❌ Sai: Crawler quét tốt, nhưng yêu cầu data stewards thủ công inspect/update format → Tăng effort lớn (nhân sự, thời gian), không "least effort". Không tận dụng classifiers tự động. -
Phương án B: Use an AWS Glue crawler to build a data catalog. Use AWS Glue crawler classifiers to recognize the format of data and store the format in the catalog.
✅ Đúng: Như phân tích trên, tự động 100% từ crawler + classifiers → Least effort, hỗ trợ S3/RDS đầy đủ, metadata format lưu sẵn trong catalog (docs AWS Glue 2026 xác nhận). -
Phương án C: Use Amazon Macie to build a data catalog and to identify sensitive data elements. Collect the data format information from Macie.
❌ Sai: Macie chuyên phát hiện sensitive data (PII, PHI) trên S3, không build data catalog hay metadata format (chỉ classification nhạy cảm). Không hỗ trợ RDS, effort cao vì phải custom collect → Không phù hợp. -
Phương án D: Use scripts to scan data elements and to assign data classifications based on the format of the data.
❌ Sai: Custom scripts (Lambda/EC2?) để quét → Effort cao (code, maintain, scale cho enterprise), không tận dụng managed service như Glue. Không "least effort", dễ lỗi và tốn kém hơn crawler tự động.
The data engineer needs to modify the current process to scan for the custom PII categories across multiple datasets within the data lake.
Which solution will meet these requirements with the LEAST operational overhead?
- A Manually review the data for custom PII categories.
- B Implement custom data quality rules in DataBrew. Apply the custom rules across datasets.
- C Develop custom Python scripts to detect the custom PII categories. Call the scripts from DataBrew.
- D Implement regex patterns to extract PII information from fields during extract transform, and load (ETL) operations into the data lake.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh một công ty sử dụng data lake trên AWS để phân tích dữ liệu hàng quý nhằm đánh giá hàng tồn kho. Một data engineer đang dùng AWS Glue DataBrew để phát hiện thông tin PII (Personally Identifiable Information) của khách hàng trong dữ liệu. Tuy nhiên, chính sách bảo mật của công ty coi một số danh mục tùy chỉnh (custom categories) là PII, nhưng những danh mục này không nằm trong các quy tắc data quality chuẩn của DataBrew.
Nhiệm vụ là sửa đổi quy trình hiện tại để quét (scan) các danh mục PII tùy chỉnh này trên nhiều bộ dữ liệu (multiple datasets) trong data lake, với ít gánh nặng vận hành nhất (LEAST operational overhead).
🔑 Yêu cầu cốt lõi: Giải pháp phải tích hợp mượt mà với DataBrew, dễ mở rộng cho nhiều datasets, tự động hóa cao, và giảm thiểu công sức bảo trì thủ công (như phát triển code phức tạp hoặc kiểm tra tay). Đây là tình huống thực tế trong AWS, nơi DataBrew được thiết kế để xử lý data profiling và quality rules một cách scalable cho data lakes (S3-based).
✅ Đáp án đúng: Implement custom data quality rules in DataBrew. Apply the custom rules across datasets.
Lý do lựa chọn:
- AWS Glue DataBrew (cập nhật đến 2026) hỗ trợ tạo custom data quality rules một cách native, cho phép định nghĩa quy tắc tùy chỉnh để detect PII dựa trên regex, patterns, hoặc logic cụ thể (như custom categories).
- Bạn có thể apply rules này lên nhiều datasets chỉ qua giao diện console hoặc API, mà không cần code thêm. Điều này giảm thiểu operational overhead vì:
- Tự động hóa hoàn toàn: Chạy jobs định kỳ trên data lake (S3), integrate với Glue crawlers cho discovery.
- Scalable: Xử lý petabyte-scale data mà không cần quản lý infrastructure.
- Least effort: Không cần dev scripts, ETL riêng, hay manual review – chỉ cần define rules một lần và reuse.
- So với các option khác, đây là native feature, phù hợp với best practices AWS cho data governance và PII detection trong data lakes.
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] Manually review the data for custom PII categories.
Phương án này yêu cầu kiểm tra thủ công, dẫn đến operational overhead cao (thời gian, lỗi con người, không scalable cho multiple datasets và quarterly runs). Không phù hợp với automation của AWS DataBrew, vi phạm nguyên tắc least overhead. -
✅ [ĐÚNG] Implement custom data quality rules in DataBrew. Apply the custom rules across datasets.
Như đã giải thích ở trên: Native support cho custom rules (regex-based hoặc custom logic) trong DataBrew, apply dễ dàng qua projects/recipes. Ít overhead nhất vì no-code/low-code, reusable, và integrate trực tiếp với data lake. -
❌ [SAI] Develop custom Python scripts to detect the custom PII categories. Call the scripts from DataBrew.
Yêu cầu phát triển và maintain Python scripts (dùng AWS Lambda hoặc Glue jobs), sau đó integrate vào DataBrew qua custom recipes. Overhead cao hơn do: code dev, testing, error handling, và dependency management – không native như custom rules. -
❌ [SAI] Implement regex patterns to extract PII information from fields during extract transform, and load (ETL) operations into the data lake.
Phương án này chỉ áp dụng trong quá trình ingest ETL (ví dụ Glue ETL jobs), không scan dữ liệu hiện có trong data lake. Phải rebuild pipeline, tăng overhead cho retrospective scans trên multiple datasets, và không tận dụng DataBrew trực tiếp.
📘 Tài liệu tham khảo (cập nhật AWS 2026)
- AWS Glue DataBrew Documentation: Data quality rules – Chi tiết custom rules cho PII detection, hỗ trợ regex/custom patterns từ 2022+.
- AWS re:Post & Best Practices: DataBrew for PII governance (updated 2025).
- Exam Guide DOP-C02: Phần DataBrew & Glue trong Data Analytics (least overhead solutions).
- AWS Well-Architected Framework: Data Lake pillar – Recommend native tools như DataBrew cho quality rules để minimize ops.
🛠️ Lời khuyên thực hành: Test custom rules trên DataBrew console với sample data lake (S3 bucket). Schedule jobs qua EventBridge cho quarterly runs! Nếu cần scale hơn, combine với AWS Lake Formation cho governance.
Occasionally, the daily data file is empty or is missing values for required fields. When the file is missing data, the company can use the previous day’s CSV file.
A data engineer needs to ensure that the previous day's data file is overwritten only if the new daily file is complete and valid.
Which solution will meet these requirements with the LEAST effort?
- A Invoke an AWS Lambda function to check the file for missing data and to fill in missing values in required fields.
- B Configure the AWS Glue ETL pipeline to use AWS Glue Data Quality rules. Develop rules in Data Quality Definition Language (DQDL) to check for missing values in required fields and empty files.
- C Use AWS Glue Studio to change the code in the ETL pipeline to fill in any missing values in the required fields with the most common values for each field.
- D Run a SQL query in Amazon Athena to read the CSV file and drop missing rows. Copy the corrected CSV file to the second S3 bucket.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh một tình huống thực tế trong AWS:
✅ Một công ty nhận file dữ liệu hàng ngày từ đối tác, lưu vào Amazon S3 bucket đầu tiên.
✅ Họ sử dụng AWS Glue ETL pipeline chạy hàng ngày để extract, transform, load (ETL): làm sạch và biến đổi dữ liệu, sau đó xuất ra file Daily.csv vào S3 bucket thứ hai.
🛑 Vấn đề: Thỉnh thoảng file dữ liệu mới rỗng (empty) hoặc thiếu giá trị ở các trường bắt buộc (missing values for required fields).
📈 Yêu cầu: Chỉ ghi đè (overwrite) file Daily.csv của ngày trước nếu file mới hoàn chỉnh và hợp lệ (complete and valid); ngược lại, giữ nguyên file cũ.
🎯 Mục tiêu: Tìm giải pháp ít công sức nhất (LEAST effort), tận dụng tối ưu các dịch vụ AWS hiện có (kiến thức cập nhật đến 2026, AWS Glue Data Quality là tính năng mạnh mẽ trong Glue 4.0+).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Configure the AWS Glue ETL pipeline to use AWS Glue Data Quality rules. Develop rules in Data Quality Definition Language (DQDL) to check for missing values in required fields and empty files.
🟢 Lý do chọn đáp án này (least effort nhất):
- AWS Glue Data Quality (tính năng built-in từ Glue 3.0+, cập nhật mạnh mẽ ở Glue 4.0/5.0 năm 2024-2026) cho phép định nghĩa quy tắc chất lượng dữ liệu (Data Quality rules) ngay trong pipeline ETL bằng ngôn ngữ DQDL (Data Quality Definition Language).
- Có thể dễ dàng viết rules kiểm tra file rỗng (row_count > 0) và missing values (is_complete) cho required fields.
- Nếu rule pass → pipeline tiếp tục và overwrite Daily.csv. Nếu fail → pipeline dừng hoặc fallback (không overwrite), hoàn toàn tự động, không cần code thêm hay dịch vụ ngoài.
- Least effort: Tích hợp trực tiếp vào Glue job hiện tại, chỉ config rules qua console/API/Script, không thay đổi architecture.
📘 Tài liệu tham khảo: - AWS Glue Data Quality Documentation (cập nhật 2026: Hỗ trợ DQDL v2 với advanced rules như completeness, validity).
- DQDL Syntax Examples (rules như
row_count() > 0 and is_complete("required_column")).
🔍 Phân tích tất cả các phương án (đúng/sai)
Dưới đây là phân tích chi tiết từng lựa chọn. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt với lý do đúng/sai. Sử dụng ✅/❌ để nổi bật.
-
❌ Invoke an AWS Lambda function to check the file for missing data and to fill in missing values in required fields.
Sai vì: Phải thêm Lambda function mới, trigger qua S3 Event hoặc Step Functions, code custom để check/fill data → tăng effort cao (deploy Lambda, IAM roles, error handling). Không tận dụng Glue pipeline hiện tại, vi phạm "least effort". Không xử lý overwrite logic tự động. -
✅ Configure the AWS Glue ETL pipeline to use AWS Glue Data Quality rules. Develop rules in Data Quality Definition Language (DQDL) to check for missing values in required fields and empty files.
Đúng vì: Như đã giải thích ở trên – tích hợp native vào Glue ETL, rules DQDL đơn giản (ví dụ:is_complete(["field1", "field2"]) and row_count() > 0), pipeline tự động stop nếu fail → giữ file cũ. Least effort, scalable, serverless. -
❌ Use AWS Glue Studio to change the code in the ETL pipeline to fill in any missing values in the required fields with the most common values for each field.
Sai vì: Thay đổi code ETL qua Glue Studio (visual editor) để fill missing bằng "most common values" → không xử lý file empty và không đảm bảo "complete/valid" (fill giả có thể invalid). Phải refactor job, test lại → effort cao, không fallback tự nhiên đến file cũ. -
❌ Run a SQL query in Amazon Athena to read the CSV file and drop missing rows. Copy the corrected CSV file to the second S3 bucket.
Sai vì: Thêm step ngoài pipeline (Athena query + S3 copy via Lambda/CLI) → phức tạp hóa workflow (schedule riêng, handle empty result). Athena giỏi query nhưng không tích hợp ETL tự động, effort cao với multi-step orchestration, không least effort so với Glue DQ.
🎯 Kết luận: Giải pháp đúng tận dụng built-in feature của AWS Glue (Data Quality), phù hợp DevOps best practices: minimal changes, high reliability. Nếu implement, test DQDL rules trên dev job trước khi prod! 🛠️
To help cost-optimize its storage, the company wants to gather information about incomplete multipart uploads and outdated versions that are present in the S3 buckets.
Which solution will meet these requirements with the LEAST operational effort?
- A Use AWS CLI to gather the information.
- B Use Amazon S3 Inventory configurations reports to gather the information.
- C Use the Amazon S3 Storage Lens dashboard to gather the information.
- D Use AWS usage reports for Amazon S3 to gather the information.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh một công ty marketing sử dụng Amazon S3 để lưu trữ dữ liệu marketing, với một số bucket kích hoạt versioning (phiên bản hóa). Công ty chạy nhiều job để đọc và tải dữ liệu lên bucket. Để tối ưu hóa chi phí lưu trữ (cost-optimize), họ cần thu thập thông tin về:
- Incomplete multipart uploads (các upload đa phần chưa hoàn thành – tốn phí lưu trữ nhưng không có dữ liệu hữu ích).
- Outdated versions (các phiên hóa cũ/noncurrent versions – do versioning, chúng vẫn tồn tại và tốn phí trừ khi xóa hoặc áp dụng lifecycle rules).
Yêu cầu giải pháp với LEAST operational effort (ít nỗ lực vận hành nhất), nghĩa là ưu tiên công cụ tự động, dashboard sẵn có, không cần script thủ công hay xử lý phức tạp. 📈
✅ Đáp án đúng: Use the Amazon S3 Storage Lens dashboard to gather the information.
Lý do lựa chọn:
Amazon S3 Storage Lens là dashboard phân tích lưu trữ S3 toàn diện (ra mắt 2020, cập nhật liên tục đến 2026), cung cấp metrics và insights tự động về incomplete multipart uploads và noncurrent versions (outdated versions) trên tất cả bucket/accounts. Nó hỗ trợ:
- Top 1,000 incomplete multipart uploads lớn nhất.
- Metrics về số lượng và kích thước noncurrent versions.
- Không cần cấu hình phức tạp, chỉ enable dashboard là có báo cáo real-time, recommendations tự động để xóa (giảm chi phí).
Least operational effort vì zero-code, visual dashboard, phù hợp cross-account/organization. 🛠️
Nguồn tham khảo: AWS S3 Storage Lens Documentation (Metrics cho IncompleteMultipartBytes và NoncurrentVersionBytes).
📋 Giải thích tất cả các phương án
-
Use AWS CLI to gather the information.
❌ Sai: AWS CLI (lệnh nhưaws s3api list-multipart-uploadshoặclist-object-versions) có thể liệt kê thông tin, nhưng yêu cầu script tự động, loop qua tất cả bucket, xử lý thủ công – operational effort cao (phải code, schedule Lambda/EC2). Không scale tốt cho nhiều bucket/jobs. Không phải giải pháp least effort. -
Use Amazon S3 Inventory configurations reports to gather the information.
❌ Sai: S3 Inventory báo cáo object list (CSV/Parquet), bao gồm version status (current/noncurrent), nhưng KHÔNG hỗ trợ incomplete multipart uploads (chỉ objects hoàn chỉnh). Cần cấu hình reports (daily/weekly), download/process thủ công – effort trung bình, không real-time dashboard. 🧐
Nguồn: S3 Inventory Docs. -
Use the Amazon S3 Storage Lens dashboard to gather the information.
✅ Đúng (như giải thích trên): Tự động, dashboard trực quan, metrics chính xác cho cả hai yêu cầu, least effort. 💯 -
Use AWS usage reports for Amazon S3 to gather the information.
❌ Sai: AWS Usage Reports (Cost Explorer/Billing) chỉ báo tổng chi phí S3 (storage classes, requests), KHÔNG chi tiết incomplete uploads hay outdated versions. Không dùng để "gather information" cụ thể cho cleanup. Chỉ hỗ trợ high-level cost analysis. 📊
Nguồn: AWS Billing Docs.
The company wants to reduce Athena costs but does not want to recreate the data pipeline.
Which solution will meet these requirements with the LEAST management effort?
- A Change the Firehose output format to Apache Parquet. Provide a custom S3 object YYYYMMDD prefix expression and specify a large buffer size. For the existing data, create an AWS Glue extract, transform, and load (ETL) job. Configure the ETL job to combine small JSON files, convert the JSON files to large Parquet files, and add the YYYYMMDD prefix. Use the ALTER TABLE ADD PARTITION statement to reflect the partition on the existing Athena table.
- B Create an Apache Spark job that combines JSON files and converts the JSON files to Apache Parquet files. Launch an Amazon EMR ephemeral cluster every day to run the Spark job to create new Parquet files in a different S3 location. Use the ALTER TABLE SET LOCATION statement to reflect the new S3 location on the existing Athena table.
- C Create a Kinesis data stream as a delivery destination for Firehose. Use Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) to run Apache Flink on the Kinesis data stream. Use Flink to aggregate the data and save the data to Amazon S3 in Apache Parquet format with a custom S3 object YYYYMMDD prefix. Use the ALTER TABLE ADD PARTITION statement to reflect the partition on the existing Athena table.
- D Integrate an AWS Lambda function with Firehose to convert source records to Apache Parquet and write them to Amazon S3. In parallel, run an AWS Glue extract, transform, and load (ETL) job to combine the JSON files and convert the JSON files to large Parquet files. Create a custom S3 object YYYYMMDD prefix. Use the ALTER TABLE ADD PARTITION statement to reflect the partition on the existing Athena table.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi mô tả một công ty game đang sử dụng Amazon Kinesis Data Streams để thu thập dữ liệu clickstream (dữ liệu về hành vi click của người dùng). Dữ liệu sau đó được Amazon Data Firehose chuyển tiếp và lưu trữ dưới định dạng JSON vào Amazon S3. Các data scientist sử dụng Amazon Athena để query dữ liệu mới nhất nhằm lấy insights kinh doanh.
📈 Vấn đề chính: Công ty muốn giảm chi phí Athena (vì Athena tính phí dựa trên lượng dữ liệu scan, JSON nhỏ lẻ không partition làm scan nhiều, tốn kém) mà không recreate (tái tạo) data pipeline (giữ nguyên Kinesis -> Firehose -> S3).
🎯 Yêu cầu giải pháp: Phải có LEAST management effort (ít nỗ lực quản lý nhất), nghĩa là config đơn giản, tự động hóa cao, không thêm nhiều dịch vụ phức tạp hoặc chạy job lặp lại hàng ngày.
Giải pháp lý tưởng: Chuyển sang định dạng Apache Parquet (columnar, nén tốt, scan nhanh) + partitioning (ví dụ theo YYYYMMDD) để Athena chỉ scan partition cần thiết, giảm chi phí đáng kể. Sử dụng kiến thức AWS cập nhật 2026: Firehose hỗ trợ trực tiếp Parquet output với buffer lớn và dynamic partitioning (S3 prefix).
📘 Tài liệu tham khảo:
- Amazon Data Firehose Documentation (Parquet conversion).
- Amazon Athena Best Practices (Partitioning & Parquet).
- AWS Glue ETL for S3 Data.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Change the Firehose output format to Apache Parquet. Provide a custom S3 object YYYYMMDD prefix expression and specify a large buffer size. For the existing data, create an AWS Glue extract, transform, and load (ETL) job. Configure the ETL job to combine small JSON files, convert the JSON files to large Parquet files, and add the YYYYMMDD prefix. Use the ALTER TABLE ADD PARTITION statement to reflect the partition on the existing Athena table.
Lý do chọn (chi tiết):
- 🛠️ Không recreate pipeline: Chỉ config Firehose (thay output format sang Parquet, set prefix YYYYMMDD cho partitioning động, buffer lớn để tạo file Parquet lớn → giảm small files). Dữ liệu mới tự động tối ưu ngay lập tức.
- 🔄 Xử lý existing data: Glue ETL job chạy one-time (không lặp lại), combine JSON nhỏ → Parquet lớn + partition → ALTER TABLE ADD PARTITION (Athena nhận diện partition ngay).
- 💡 Least management effort: Config Firehose là serverless (0 management), Glue ETL one-time + schedule nếu cần (crawl tự động). Giảm chi phí Athena lên đến 75% nhờ Parquet + partition (theo AWS best practices 2026).
- 🚀 Hiệu quả cao: Firehose native hỗ trợ Parquet conversion từ 2021, cập nhật 2026 thêm dynamic partitioning tốt hơn.
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt với lý do đúng/sai.
✅ Phương án ĐÚNG (như trên):
Change the Firehose output format to Apache Parquet. Provide a custom S3 object YYYYMMDD prefix expression and specify a large buffer size. For the existing data, create an AWS Glue extract, transform, and load (ETL) job. Configure the ETL job to combine small JSON files, convert the JSON files to large Parquet files, and add the YYYYMMDD prefix. Use the ALTER TABLE ADD PARTITION statement to reflect the partition on the existing Athena table.
(Giải thích như phần trên – tối ưu nhất, least effort).
❌ Phương án SAI 1:
Create an Apache Spark job that combines JSON files and converts the JSON files to Apache Parquet files. Launch an Amazon EMR ephemeral cluster every day to run the Spark job to create new Parquet files in a different S3 location. Use the ALTER TABLE SET LOCATION statement to reflect the new S3 location on the existing Athena table.
Lý do sai:
- 🕒 Management effort cao: Phải launch EMR cluster hàng ngày (ephemeral nhưng vẫn config, monitor, scale), Spark job custom → tốn dev & ops.
- 🔄 Vi phạm không recreate pipeline: Tạo S3 location mới + ALTER TABLE SET LOCATION (thay đổi toàn bộ table location) → gần như rebuild pipeline.
- 💰 Tốn kém hơn: EMR không serverless hoàn toàn, chi phí cluster + daily run > Firehose native.
❌ Phương án SAI 2:
Create a Kinesis data stream as a delivery destination for Firehose. Use Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) to run Apache Flink on the Kinesis data stream. Use Flink to aggregate the data and save the data to Amazon S3 in Apache Parquet format with a custom S3 object YYYYMMDD prefix. Use the ALTER TABLE ADD PARTITION statement to reflect the partition on the existing Athena table.
Lý do sai:
- 🔄 Recreate pipeline: Thêm Kinesis stream mới làm destination cho Firehose gốc → Firehose -> Kinesis -> Flink -> S3 (thêm layer phức tạp, không giữ nguyên pipeline).
- 🛠️ Management effort cao: Flink (Managed Service for Apache Flink, cập nhật 2023 từ KDA) cần config application, scaling, monitoring → không least effort. Aggregate custom → dev nặng.
- ⚠️ Không cần thiết: Firehose đã hỗ trợ Parquet trực tiếp, không cần Flink trung gian.
❌ Phương án SAI 3:
Integrate an AWS Lambda function with Firehose to convert source records to Apache Parquet and write them to Amazon S3. In parallel, run an AWS Glue extract, transform, and load (ETL) job to combine the JSON files and convert the JSON files to large Parquet files. Create a custom S3 object YYYYMMDD prefix. Use the ALTER TABLE ADD PARTITION statement to reflect the partition on the existing Athena table.
Lý do sai:
- 🛠️ Management effort cao: Lambda integration với Firehose cần code custom (record transformation), debug payload limits (15MB), timeout → dev & test nhiều.
- 🔄 Parallel runs phức tạp: Glue ETL cho existing + Lambda cho new → 2 processes riêng biệt, dễ lỗi sync partition/prefix. Firehose transformation chỉ cho một số format, Parquet phức tạp hơn JSON.
- 💡 Không optimal: Firehose native Parquet tốt hơn Lambda (serverless thực thụ, buffer tự động), Lambda tăng latency & chi phí invocation.
🎉 Kết luận: Phương án đúng là lựa chọn serverless-native nhất, phù hợp DevOps best practices AWS 2026!
Which solution will meet these requirements with the LEAST ongoing maintenance?
- A Use the DynamoDB TTL feature to automatically expire data based on timestamps.
- B Configure a scheduled Amazon EventBridge rule to invoke an AWS Lambda function to check for data that is older than 1 month. Configure the Lambda function to delete old data.
- C Configure a stream on the DynamoDB table to invoke an AWS Lambda function. Configure the Lambda function to delete data in the table that is older than 1 month.
- D Use an AWS Lambda function to periodically scan the DynamoDB table for data that is older than 1 month. Configure the Lambda function to delete old data.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc quản lý chi phí và kích thước của một bảng Amazon DynamoDB hiện có 📊. Công ty yêu cầu:
- Giải pháp không làm gián đoạn các hoạt động đọc/ghi đang diễn ra (không downtime hoặc impact performance).
- Tự động xóa dữ liệu sau đúng 1 tháng ⏰.
- Ưu tiên giải pháp có ÍT BẢO TRÌ LIÊN TỤC NHẤT (least ongoing maintenance) 🛠️, nghĩa là không cần can thiệp thủ công, monitoring phức tạp hay chi phí vận hành cao.
Đây là yêu cầu điển hình trong DevOps trên AWS, đặc biệt với DynamoDB – dịch vụ NoSQL serverless, nơi chi phí tỷ lệ với lưu trữ và throughput. Giải pháp phải tích hợp sẵn (native) để giảm maintenance, tránh custom code dễ lỗi.
✅ Đáp án đúng: Use the DynamoDB TTL feature to automatically expire data based on timestamps.
Lý do chọn:
- DynamoDB TTL (Time to Live) là tính năng native, serverless của AWS (cập nhật ổn định đến 2026), tự động xóa items dựa trên attribute timestamp (ví dụ:
expirationTimevới giá trị epoch time sau 1 tháng). - Không gián đoạn read/write: Xóa diễn ra asynchronous, background, không ảnh hưởng RCU/WCU.
- Least maintenance: Zero code, zero scheduling, AWS tự quản lý – chỉ enable TTL trên attribute một lần duy nhất 🛡️.
- Tiết kiệm chi phí: Xóa miễn phí (không tốn RCU/WCU cho delete), giảm storage billing ngay lập tức 💰.
- Phù hợp bảng existing vì có thể enable TTL mà không recreate table.
📋 Giải thích tất cả các phương án (sử dụng kiến thức AWS mới nhất 2026)
-
✅ Use the DynamoDB TTL feature to automatically expire data based on timestamps.
Đúng vì: Đây là giải pháp tối ưu nhất theo best practice AWS. TTL được thiết kế chính xác cho trường hợp này: thêm attribute TTL (number type, epoch seconds), enable qua console/CLI/API. AWS tự scan và xóa soft-delete (items vẫn readable đến khi xóa hoàn toàn, thường <48h). Không cần code, không phí invoke/scan, scale tự động với bảng lớn. Least maintenance thực sự! -
❌ Configure a scheduled Amazon EventBridge rule to invoke an AWS Lambda function to check for data that is older than 1 month. Configure the Lambda function to delete old data.
Sai vì: Dù dùng EventBridge (trước là CloudWatch Events, ổn định 2026) để schedule Lambda, giải pháp này yêu cầu viết code scan/query/delete → tốn RCU cho scan (chi phí cao với bảng lớn), phí Lambda invocations và maintenance cao (debug code, handle errors, permissions IAM, monitor logs CloudWatch). Có thể gián đoạn nếu Lambda throttle hoặc delete batch lớn gây hot partitions. -
❌ Configure a stream on the DynamoDB table to invoke an AWS Lambda function. Configure the Lambda function to delete data in the table that is older than 1 month.
Sai vì: DynamoDB Streams chỉ capture thay đổi mới (insert/update/delete), không scan dữ liệu cũ → Lambda không thể detect "data older than 1 month" từ stream. Phải kết hợp scan riêng, dẫn đến logic phức tạp, maintenance cao, phí stream + Lambda liên tục. Không phù hợp cho delete retroactive, dễ loop delete vô tận nếu code kém. -
❌ Use an AWS Lambda function to periodically scan the DynamoDB table for data that is older than 1 month. Configure the Lambda function to delete old data.
Sai vì: Scan toàn bộ bảng rất tốn kém (1 RCU/1KB scanned, bảng lớn → bill cao), đặc biệt periodic (dùng cron/EventBridge). Maintenance cao: code xử lý pagination, batch delete (Limit 25 items/batch), handle throttling, state management. Có thể throttle table gây gián đoạn read/write thực tế, vi phạm yêu cầu.
📘 Tài liệu tham khảo (AWS chính thức, cập nhật 2026)
- DynamoDB TTL Guide: https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/ttl-howitworks.html (Best practice cho auto-expire).
- DynamoDB Pricing: https://aws.amazon.com/dynamodb/pricing/ (Xác nhận TTL free delete, scan tốn RCU).
- AWS Well-Architected Framework - Cost Optimization: TTL được recommend cho data lifecycle.
- Exam Prep DOP-C02: Chủ đề DynamoDB management thường test TTL vs custom solutions.
Giải pháp này giúp DevOps Engineer đạt zero-touch operations 🚀! Nếu cần demo code enable TTL, hỏi thêm nhé!
The company has an S3 bucket in an AWS account named Hub-Account. The S3 bucket is encrypted by an AWS Key Management Service (AWS KMS) key. The company's QuickSight instance is in a separate account named BI-Account.
The company updates the S3 bucket policy to grant access to the QuickSight service role. The company wants to enable cross-account access to allow QuickSight to interact with the S3 bucket.
Which combination of steps will meet this requirement? (Choose two.)
- A Use the existing AWS KMS key to encrypt connections from QuickSight to the S3 bucket.
- B Add the S3 bucket as a resource that the QuickSight service role can access.
- C Use AWS Resource Access Manager (AWS RAM) to share the S3 bucket with the BI-Account account.
- D Add an IAM policy to the QuickSight service role to give QuickSight access to the KMS key that encrypts the S3 bucket.
- E Add the KMS key as a resource that the QuickSight service role can access.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc cho phép truy cập cross-account giữa Amazon QuickSight (nằm trong tài khoản BI-Account) và một S3 bucket (nằm trong tài khoản Hub-Account). S3 bucket này được mã hóa bằng AWS KMS key. Công ty đã cập nhật S3 bucket policy để cấp quyền cho QuickSight service role, nhưng cần thêm các bước để QuickSight có thể tương tác đầy đủ (đọc dữ liệu, visualize).
Yêu cầu chính: Chọn TWO bước kết hợp để enable cross-account access.
- QuickSight sử dụng service role (thường là
aws-quicksight-service-role-v0) để truy cập tài nguyên bên ngoài. - Với S3 cross-account: Cần bucket policy cho phép service role truy cập (s3:GetObject, ListBucket).
- Với KMS encryption: Service role cần quyền KMS (kms:Decrypt, kms:DescribeKey) qua KMS key policy (không phải IAM policy trên role).
- Quy trình AWS: Trong QuickSight console, khi tạo dataset từ S3 cross-account, hệ thống hướng dẫn grant permissions trực tiếp lên resource policies (S3 bucket và KMS key).
📘 Tài liệu tham khảo:
- AWS Docs: Granting Amazon QuickSight permissions to access AWS resources (cập nhật 2024-2026).
- Cross-account access for QuickSight and S3.
- AWS Well-Architected Framework: Security Pillar (DevOps Professional exam topics).
✅ Đáp án đúng (Chọn TWO)
Hai lựa chọn đúng là:
- Add the S3 bucket as a resource that the QuickSight service role can access.
- Add the KMS key as a resource that the QuickSight service role can access.
Lý do lựa chọn 🛠️:
- Để QuickSight đọc dữ liệu từ S3 cross-account, phải grant permissions trực tiếp lên S3 bucket policy (Principal là QuickSight service role ARN:
arn:aws:quicksight:region:BI-Account-ID:service-role/...). - Vì bucket encrypted bằng KMS (customer-managed key), QuickSight service role cần quyền KMS qua KMS key policy (không modify IAM role). QuickSight console tự động generate policy statements để add vào resource policies. Đây là cách chuẩn theo best practices AWS, tránh modify managed service role.
📋 Giải thích TẤT CẢ các phương án (Đúng/Sai)
-
Use the existing AWS KMS key to encrypt connections from QuickSight to the S3 bucket.
❌ SAI: KMS key dùng để encrypt dữ liệu tại rest trong S3, không phải encrypt connections (kết nối). QuickSight-S3 dùng HTTPS/TLS mặc định, không cần KMS cho transport encryption. Lựa chọn này nhầm lẫn khái niệm. -
Add the S3 bucket as a resource that the QuickSight service role can access.
✅ ĐÚNG: Phải add S3 bucket vào permissions của QuickSight service role qua bucket policy update. QuickSight console cung cấp policy template (Allow s3:GetObject*, s3:ListBucket*) với Principal là service role ARN. Đáp ứng cross-account access cơ bản. -
Use AWS Resource Access Manager (AWS RAM) to share the S3 bucket with the BI-Account account.
❌ SAI: AWS RAM hỗ trợ share resources như VPC, Transit Gateway, License Configurations, KHÔNG hỗ trợ S3 buckets. S3 cross-account dùng bucket policies hoặc IAM roles, không phải RAM (xem AWS RAM docs). -
Add an IAM policy to the QuickSight service role to give QuickSight access to the KMS key that encrypts the S3 bucket.
❌ SAI: QuickSight service role là AWS managed role, không được phép attach custom IAM policies trực tiếp (vi phạm least privilege). Quyền KMS phải grant qua KMS key policy (resource-based), không phải identity-based policy trên role. -
Add the KMS key as a resource that the QuickSight service role can access.
✅ ĐÚNG: Với S3 encrypted bằng KMS, cần add KMS key vào permissions qua KMS key policy (cho phép kms:Decrypt, kms:GenerateDataKey*). QuickSight console generate statement với Principal là service role ARN, enable decrypt dữ liệu cross-account. Bắt buộc cho mã hóa!
🛠️ Lời khuyên thực hành: Test bằng QuickSight > Datasets > New dataset > S3 > Chọn cross-account bucket > Follow grant permissions. Luôn dùng IAM Access Analyzer kiểm tra policies. (DevOps Pro tip! 🚀)