Ngân hàng đề — AWS Certified Data Engineer Associate
Tìm thấy 867 câu.
The company wants to ensure the API's Lambda function operate without being affected by other Lambda functions.
Which solution will meet this requirement MOST cost-effectively?
- A Increase the number of read capacity unit (RCU) in DynamoDB.
- B Configure provisioned concurrency for the Lambda function.
- C Configure reserved concurrency for the Lambda function.
- D Increase the Lambda function timeout and allocated memory.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một ứng dụng sử dụng Amazon API Gateway REST API kết nối với AWS Lambda function để lấy dữ liệu từ Amazon DynamoDB. Người dùng gặp vấn đề độ trễ cao ngắt quãng (intermittent high latency), và kỹ sư dữ liệu phát hiện Lambda function bị throttling thường xuyên khi các Lambda function khác của công ty tăng đột biến số lượng invocations.
Yêu cầu chính: Đảm bảo Lambda function của API không bị ảnh hưởng bởi các Lambda function khác, và giải pháp phải cost-effective nhất (tiết kiệm chi phí nhất).
Vấn đề cốt lõi là throttling do giới hạn concurrency (số lượng invocation đồng thời) của Lambda ở mức account-wide (mặc định 1.000 concurrency cho tài khoản). Khi các function khác "chiếm hết" concurrency, function này bị throttle dù DynamoDB ổn. Giải pháp cần cách ly concurrency cho function cụ thể mà không tốn kém. (Kiến thức cập nhật AWS 2026: Lambda concurrency vẫn quản lý theo model này, với tùy chọn burst và reserved để scale.)
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Configure reserved concurrency for the Lambda function.
🛠️ Lý do chi tiết:
- Reserved concurrency dành riêng một phần quota concurrency (ví dụ: 200) chỉ cho function này, ngăn các function khác sử dụng, tránh throttling. Tổng concurrency account không thay đổi, chỉ phân bổ lại.
- Cost-effective nhất vì hoàn toàn miễn phí (không tính phí như provisioned concurrency), chỉ cần cấu hình qua Console/CLI/Terraform. Phù hợp với yêu cầu "không bị ảnh hưởng bởi function khác".
- Trong AWS 2026, đây là best practice cho production workloads cần đảm bảo (SLA), ví dụ reserve 80% cho critical functions.
📋 Giải thích tất cả các phương án
-
Increase the number of read capacity unit (RCU) in DynamoDB.
❌ Sai: Tăng RCU chỉ cải thiện throughput đọc DynamoDB, không giải quyết throttling Lambda (lỗi 429 do concurrency, không phải DynamoDB provisioned mode). Vấn đề là Lambda invocations bị chặn trước khi query DB, nên vô ích và tốn phí (RCU tính theo giờ). -
Configure provisioned concurrency for the Lambda function.
❌ Sai: Provisioned concurrency giữ warm instances để giảm cold start latency, nhưng vẫn dùng chung account concurrency quota. Nếu function khác throttle account quota, function này vẫn bị ảnh hưởng. Không cost-effective vì tốn phí cao (~0.00000467 USD/GB-giây per provisioned request, AWS 2026 pricing). -
Configure reserved concurrency for the Lambda function.
✅ Đúng: Như giải thích trên, cách ly hoàn toàn concurrency, miễn phí, trực tiếp giải quyết throttling từ function khác. Config dễ dàng: Lambda Console > Configuration > Concurrency > Reserve (ví dụ: 500 cho function này). -
Increase the Lambda function timeout and allocated memory.
❌ Sai: Tăng timeout/memory chỉ giúp function chạy lâu hơn/nhanh hơn (nhờ CPU scale), nhưng không tránh throttling concurrency. Throttling xảy ra trước khi function execute, nên vô hiệu.
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- AWS Lambda Concurrency: docs.aws.amazon.com/lambda/latest/dg/lambda-concurrency.html – Chi tiết reserved vs provisioned.
- Lambda Best Practices: docs.aws.amazon.com/lambda/latest/dg/best-practices.html – Khuyến nghị reserved cho isolation.
- API Gateway + Lambda Throttling: docs.aws.amazon.com/apigateway/latest/developerguide/limits.html.
- DynamoDB Capacity: docs.aws.amazon.com/amazondynamodb/latest/developerguide/HowItWorks.ReadWriteCapacityMode.html (xác nhận không liên quan throttling Lambda).
Giải pháp này đảm bảo high availability mà tối ưu chi phí! 🚀 Nếu cần demo code Terraform, hỏi thêm nhé!
The non-PII data must be available to everyone in the company. The PII data must be available only to a limited group of employees.
Which solution will meet these requirements with the LEAST operational overhead?
- A Store the JSON file in an Amazon S3 bucket. Configure AWS Glue to split the file into one file that contains the PII data and one file that contains the non-PII data. Store the output files in separate S3 buckets. Grant the required access to the buckets based on the type of user.
- B Store the JSON file in an Amazon S3 bucket. Use Amazon Macie to identify PII data and to grant access based on the type of user.
- C Store the JSON file in an Amazon S3 bucket. Catalog the file schema in AWS Lake Formation. Use Lake Formation permissions to provide access to the required data based on the type of user.
- D Create two Amazon RDS PostgreSQL databases. Load the PII data and the non-PII data into the separate databases. Grant access to the databases based on the type of user.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh một công ty sở hữu file JSON chứa dữ liệu cá nhân có thể nhận dạng (PII - Personally Identifiable Information) và dữ liệu không phải PII (non-PII). Yêu cầu chính là:
- Làm cho dữ liệu có thể truy vấn và phân tích (querying and analysis).
- Non-PII: Có sẵn cho toàn bộ nhân viên công ty.
- PII: Chỉ dành cho nhóm nhân viên hạn chế.
- Giải pháp phải có operational overhead thấp nhất (LEAST operational overhead), nghĩa là giảm thiểu công sức quản lý, xử lý dữ liệu thủ công, di chuyển file hoặc thiết lập phức tạp.
Vấn đề cốt lõi là phân quyền truy cập chi tiết (fine-grained access control) trên cùng một file JSON lưu trong S3, mà không cần tách file vật lý hoặc sử dụng cơ sở dữ liệu truyền thống. Đây là tình huống điển hình cho data lake trên AWS, nơi dữ liệu thô được lưu trữ và quản lý quyền qua metadata. ✅ Giải pháp lý tưởng phải tận dụng S3 làm storage gốc và công cụ quản lý quyền ở mức schema/dữ liệu logic.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Store the JSON file in an Amazon S3 bucket. Catalog the file schema in AWS Lake Formation. Use Lake Formation permissions to provide access to the required data based on the type of user.
Lý do chọn đáp án này 🛠️:
- AWS Lake Formation (tích hợp với S3 và Glue Data Catalog) cho phép catalog schema của file JSON mà không cần di chuyển hoặc biến đổi dữ liệu vật lý. Nó hỗ trợ phân quyền granular (ở mức cột/dữ liệu cụ thể) dựa trên user/group, lý tưởng để tách PII/non-PII logic mà giữ nguyên file gốc.
- Querying và analysis dễ dàng qua Athena, EMR hoặc Redshift Spectrum mà không overhead cao.
- Least operational overhead: Chỉ cần thiết lập permissions một lần qua UI/CLI/API, tự động hóa với IAM/LF policies. Không cần ETL job định kỳ hay quản lý DB.
- Cập nhật 2026: Lake Formation v3+ hỗ trợ row/column-level security nâng cao cho JSON/Parquet, tích hợp seamless với GovCloud cho PII compliance (GDPR/HIPAA).
📋 Phân tích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể bằng tiếng Việt:
-
Store the JSON file in an Amazon S3 bucket. Configure AWS Glue to split the file into one file that contains the PII data and one file that contains the non-PII data. Store the output files in separate S3 buckets. Grant the required access to the buckets based on the type of user.
❌ Sai. Phương án này yêu cầu Glue job để split file, tạo overhead cao: phải viết ETL script, chạy job định kỳ (nếu file update), quản lý 2 S3 buckets riêng, và bucket-level ACL/IAM policy chỉ coarse-grained (không granular). Không phù hợp "least overhead" vì tốn công maintain pipeline. -
Store the JSON file in an Amazon S3 bucket. Use Amazon Macie to identify PII data and to grant access based on the type of user.
❌ Sai. Amazon Macie giỏi phát hiện PII (scan S3), nhưng không hỗ trợ grant access granular dựa trên type user cho querying. Nó chỉ alert/remediate, không thay thế permissions logic trên schema. Overhead vẫn cao nếu kết hợp thêm tool khác; không native cho analysis. -
Store the JSON file in an Amazon S3 bucket. Catalog the file schema in AWS Lake Formation. Use Lake Formation permissions to provide access to the required data based on the type of user.
✅ Đúng. Như đã giải thích ở trên: Lake Formation catalog schema JSON qua Glue, áp dụng permissions ở mức dữ liệu (data filtering) cho PII/non-PII mà không touch file gốc. Hỗ trợ query qua Athena với zero-ETL. Overhead thấp nhất nhờ managed service. -
Create two Amazon RDS PostgreSQL databases. Load the PII data and the non-PII data into the separate databases. Grant access to the databases based on the type of user.
❌ Sai. RDS PostgreSQL là relational DB, yêu cầu load/extract dữ liệu từ JSON (ETL phức tạp), quản lý 2 instances riêng (scaling, backup, patching). Overhead cực cao: chi phí, không scale cho big data lake, không "least operational" so với S3-native solution.
📘 Tài liệu tham khảo (Cập nhật AWS 2026)
- AWS Lake Formation Documentation: Lake Formation Permissions & Data Access – Chi tiết row/column filtering cho JSON.
- AWS Well-Architected Framework - Data Lake Lens: Building Data Lakes – Khuyến nghị Lake Formation cho PII governance.
- AWS re:Invent 2025/2026 Updates: Lake Formation enhancements cho hybrid/multi-cloud PII (xem AWS Blog: "Zero-ETL for Data Lakes").
- Exam Prep DOP-C02: Topic "Data Security & Governance" nhấn mạnh Lake Formation cho least-overhead access control.
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần thêm ví dụ code Terraform/CLI, hãy hỏi nhé!
A data engineer needs to use the AWS CLI to create the cross-Region snapshot.
Which combination of steps will meet these requirements? (Choose two.)
- A Create a KMS key and configure a snapshot copy grant in the source AWS Region.
- B In the source AWS Region, enable snapshot copying. Specify the name of the snapshot copy grant that is created in the destination AWS Region.
- C In the source AWS Region, enable snapshot copying. Specify the name of the snapshot copy grant that is created in the source AWS Region.
- D Create a KMS key and configure a snapshot copy grant in the destination AWS Region.
- E Convert the cluster to a Multi-AZ deployment.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào chiến lược phục hồi thảm họa (Disaster Recovery - DR) cho một Amazon Redshift cluster được mã hóa bằng AWS Key Management Service (AWS KMS). Công ty muốn tạo snapshot cross-Region (bản sao lưu giữa các Region khác nhau) để đảm bảo tính sẵn sàng cao.
- Yêu cầu cụ thể: Sử dụng AWS CLI để tạo snapshot cross-Region từ source Region sang destination Region.
- Thách thức chính: Vì cluster được mã hóa bằng KMS, việc copy snapshot cross-Region cần xử lý khóa mã hóa một cách an toàn. Redshift yêu cầu KMS key và snapshot copy grant phải được cấu hình đúng để cho phép copy dữ liệu mã hóa giữa các Region (theo tài liệu AWS mới nhất 2025-2026, quy trình này vẫn giữ nguyên với hỗ trợ KMS multi-Region keys nếu áp dụng, nhưng câu hỏi ngụ ý key riêng biệt).
- Định dạng: Chọn 2 bước kết hợp để đáp ứng yêu cầu.
Quy trình cơ bản (dựa trên AWS CLI cho Redshift):
- Tạo KMS key ở destination Region.
- Tạo snapshot copy grant ở destination Region, grant quyền cho KMS key source.
- Ở source Region, enable snapshot copying và chỉ định tên snapshot copy grant từ destination.
✅ Đáp án đúng (Chọn 2)
Hai phương án đúng là:
- In the source AWS Region, enable snapshot copying. Specify the name of the snapshot copy grant that is created in the destination AWS Region.
- Create a KMS key and configure a snapshot copy grant in the destination AWS Region.
Lý do lựa chọn:
- Để copy snapshot cross-Region với dữ liệu mã hóa KMS, AWS Redshift bắt buộc phải có snapshot copy grant ở destination Region (tạo bằng
aws redshift create-snapshot-copy-grant), liên kết với KMS key ở destination (tạo bằngaws kms create-key). - Sau đó, ở source Region, enable snapshot copying (bằng
modify-clusterhoặcenable-snapshot-copy) và chỉ định tên grant từ destination để Redshift tự động copy snapshot khi tạo. - Sự kết hợp này đảm bảo bảo mật khóa mã hóa (khóa source không rời Region gốc) và hỗ trợ DR tự động. Sử dụng AWS CLI như
aws redshift create-cluster-snapshotvới--destination-regionsau khi cấu hình.
🛠️ Giải thích tất cả các phương án (Đúng/Sai)
-
❌ Create a KMS key and configure a snapshot copy grant in the source AWS Region.
Sai: Snapshot copy grant phải được tạo ở destination Region, không phải source. Nếu tạo ở source, Redshift không thể sử dụng để authorize copy cross-Region, dẫn đến lỗi "InvalidSnapshotCopyGrant" khi thực hiện CLI command. -
✅ In the source AWS Region, enable snapshot copying. Specify the name of the snapshot copy grant that is created in the destination AWS Region.
Đúng: Đây là bước cuối cùng ở source để kích hoạt copy tự động. Chỉ định tên grant từ destination (ARN) qua tham số--snapshot-copy-grant-nametrong lệnhaws redshift enable-snapshot-copy. Đảm bảo snapshot được replicate an toàn sang Region khác. -
❌ In the source AWS Region, enable snapshot copying. Specify the name of the snapshot copy grant that is created in the source AWS Region.
Sai: Tương tự phương án đầu, grant ở source không hợp lệ cho cross-Region. Redshift yêu cầu grant cross-Region reference, nếu chỉ định grant source sẽ thất bại khi copy. -
✅ Create a KMS key and configure a snapshot copy grant in the destination AWS Region.
Đúng: Bước đầu tiên thiết yếu. Tạo KMS key ở destination (aws kms create-key), sau đó tạo grant (aws redshift create-snapshot-copy-grant --snapshot-copy-grant-name <name> --kms-key-id <destination-key-arn> --tags ...). Grant này cho phép KMS source decrypt và re-encrypt bằng key destination. -
❌ Convert the cluster to a Multi-AZ deployment.
Sai: Multi-AZ chỉ tăng high availability trong cùng Region (tự động failover compute nodes), không liên quan đến cross-Region snapshot hay DR. Không hỗ trợ copy dữ liệu sang Region khác.
📘 Tài liệu tham khảo (Cập nhật AWS 2025-2026)
- AWS Redshift Documentation: Copying snapshots across AWS Regions & Snapshot copy grants for encrypted snapshots.
- AWS CLI Reference:
aws redshift create-snapshot-copy-grant&aws redshift enable-snapshot-copy. - KMS Cross-Region Support: Multi-Region Keys (tùy chọn, nhưng câu hỏi dùng key riêng).
- Exam Prep: AWS Certified DevOps Engineer Professional DOP-C02 (phiên bản 2025), phần Redshift DR.
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần ví dụ AWS CLI cụ thể, hãy hỏi thêm.
Most of the source databases are hosted on Amazon RDS. However, one source database is an on-premises Microsoft SQL Server Enterprise instance. The company needs to implement a solution to replicate existing data from all source databases and all future changes to the target S3 data lake.
Which solution will meet these requirements MOST cost-effectively?
- A Use one AWS Glue job to replicate existing data. Use a second AWS Glue job to replicate future changes.
- B Use AWS Database Migration Service (AWS DMS) to replicate existing data. Use AWS Glue jobs to replicate future changes.
- C Use AWS Database Migration Service (AWS DMS) to replicate existing data and future changes.
- D Use AWS Glue jobs to replicate existing data. Use Amazon Kinesis Data Streams to replicate future changes.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc xây dựng một data lake trên Amazon S3 📦, nơi cần replicate (sao chép) dữ liệu từ nhiều nguồn cơ sở dữ liệu (databases) vào định dạng Apache Parquet (một định dạng columnar hiệu quả cho phân tích dữ liệu lớn). Các nguồn bao gồm:
- Hầu hết là Amazon RDS (các instance RDS như MySQL, PostgreSQL, SQL Server trên AWS).
- Một nguồn on-premises Microsoft SQL Server Enterprise (chạy ngoài AWS, cần kết nối qua mạng).
Yêu cầu chính:
- Sao chép dữ liệu hiện có (existing data) – tức là full load ban đầu.
- Sao chép tất cả thay đổi tương lai (future changes) – tức là ongoing replication, thường qua CDC (Change Data Capture) để capture insert/update/delete real-time hoặc near real-time.
- Giải pháp phải cost-effective nhất (tiết kiệm chi phí nhất) 💰.
Thách thức: Hỗ trợ đa nguồn (cloud + on-prem), output Parquet vào S3, và xử lý cả batch + streaming mà không phức tạp hóa kiến trúc. AWS DMS là lựa chọn lý tưởng vì hỗ trợ tất cả (cập nhật đến 2026, DMS version 3.4+ hỗ trợ Parquet native cho S3 target với CDC đầy đủ).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use AWS Database Migration Service (AWS DMS) to replicate existing data and future changes.
Lý do chi tiết:
- AWS DMS hỗ trợ full load (existing data) và CDC (future changes) trong một task duy nhất 🛠️, giúp đơn giản hóa, giảm chi phí vận hành (không cần nhiều service).
- Hỗ trợ tất cả nguồn: RDS (MySQL, PostgreSQL, SQL Server, Oracle...) và on-premises SQL Server Enterprise (qua DMS endpoint với replication instance).
- Target là S3: DMS xuất trực tiếp vào S3 bucket dưới dạng Parquet (hỗ trợ từ 2019, tối ưu hóa đến 2026 với columnar storage và compression tự động).
- Cost-effective nhất: Giá DMS dựa trên replication instance hours (khoảng $0.018/giờ cho t3.micro), không cần thêm ETL hay streaming service. Hỗ trợ multi-source qua multi-task, scale tự động.
- So với các option khác, tránh overhead của Glue (batch-only) hay Kinesis (streaming đắt đỏ).
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn giữ nguyên văn bản gốc bằng tiếng Anh, kèm giải thích bằng tiếng Việt với lý do đúng/sai:
-
❌ [SAI] Use one AWS Glue job to replicate existing data. Use a second AWS Glue job to replicate future changes.
Glue là ETL batch-oriented 🐌, giỏi transform dữ liệu lớn nhưng không hỗ trợ CDC real-time cho future changes (chỉ polling thủ công, kém hiệu quả). Không kết nối on-prem SQL Server dễ dàng (cần VPC peering phức tạp). Cần 2 jobs riêng → tăng chi phí dev + run (Glue DPUs ~$0.44/giờ). Không cost-effective. -
❌ [SAI] Use AWS Database Migration Service (AWS DMS) to replicate existing data. Use AWS Glue jobs to replicate future changes.
DMS làm tốt existing data, nhưng dùng Glue cho future changes là thừa thãi vì DMS đã hỗ trợ CDC đầy đủ (log-based cho SQL Server). Tách ra → tăng chi phí (DMS + Glue), phức tạp monitoring. Không tối ưu nhất. -
✅ [ĐÚNG] Use AWS Database Migration Service (AWS DMS) to replicate existing data and future changes.
Như đã giải thích ở trên: Một service duy nhất xử lý full load + CDC, hỗ trợ RDS + on-prem SQL Server, output Parquet trực tiếp vào S3. Tiết kiệm chi phí, scale dễ (tăng instance size nếu cần). Phù hợp best practice AWS cho data lake replication (2026). -
❌ [SAI] Use AWS Glue jobs to replicate existing data. Use Amazon Kinesis Data Streams to replicate future changes.
Glue cho existing OK nhưng kém cho CDC; Kinesis Data Streams là streaming real-time (shard-based, ~$0.015/shard/giờ) nhưng không native hỗ trợ CDC database (cần custom producer từ DB logs, phức tạp với on-prem). Cần thêm Kinesis Data Firehose để sink Parquet vào S3 → chi phí cao (provisioned throughput), overkill cho use case này.
📘 Tài liệu tham khảo (AWS docs cập nhật 2026)
- AWS DMS for S3 Target (Parquet support): docs.aws.amazon.com/dms/latest/userguide/CHAP_Target.S3.html – Hỗ trợ full load + CDC với Parquet.
- DMS Source: SQL Server (on-prem): docs.aws.amazon.com/dms/latest/userguide/CHAP_Source.SQLServer.html – CDC via MSSQL log.
- RDS as DMS Source: docs.aws.amazon.com/dms/latest/userguide/CHAP_Source.RDS.html.
- Best Practices Data Lake: AWS Well-Architected Framework - Data Lake Lens (2025 update).
- Pricing DMS: aws.amazon.com/dms/pricing – Xác nhận cost-effective so với Glue/Kinesis.
Giải pháp này đảm bảo zero-downtime replication và dễ integrate với Athena/Redshift cho query data lake! 🚀
The data engineer runs queries once each week to extract metrics from the orders data based the order date for multiple date ranges. The data engineer needs an optimization solution that ensures the query performance will not degrade when the volume of data increases.
Which solution will meet this requirement MOST cost-effectively?
- A Partition the data based on order date. Use Amazon Athena to query the data.
- B Partition the data based on order date. Use Amazon Redshift to query the data.
- C Partition the data based on load date. Use Amazon EMR to query the data.
- D Partition the data based on load date. Use Amazon Aurora to query the data.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc tối ưu hóa hiệu suất (performance) của một data pipeline xử lý dữ liệu đơn hàng bán lẻ (retail orders). Dữ liệu được ingest hàng ngày vào Amazon S3 bucket. Data engineer chạy query hàng tuần để trích xuất metrics từ dữ liệu đơn hàng, dựa trên order date (ngày đặt hàng) cho nhiều khoảng ngày (multiple date ranges). Yêu cầu chính là giải pháp đảm bảo performance không suy giảm khi volume dữ liệu tăng, và phải cost-effective nhất (tiết kiệm chi phí nhất).
🔍 Thách thức chính:
- Dữ liệu lớn tăng dần theo thời gian → cần cơ chế partitioning để query nhanh (partition pruning giúp Athena/others skip dữ liệu không liên quan).
- Query infrequent (hàng tuần), không cần always-on cluster.
- Phải query trực tiếp trên S3 mà không cần di chuyển dữ liệu lớn.
- Theo kiến thức AWS cập nhật đến 2026 (Athena engine version 3 hỗ trợ columnar formats tốt hơn, partitioning tự động với Glue Crawler), giải pháp phải serverless, pay-per-use để cost-effective.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Partition the data based on order date. Use Amazon Athena to query the data.
Lý do 🛠️:
- Partition theo order date: Hoàn hảo vì query filter theo order date → Athena sử dụng partition pruning (bỏ qua partitions không khớp), scan ít dữ liệu hơn, performance ổn định dù data volume tăng (hàng TB/PB).
- Amazon Athena: Serverless query engine cho S3, không cần quản lý infra, pay-per-query (TB scanned), phù hợp query infrequent. Cost thấp nhất (khoảng $5/TB scanned năm 2026), tích hợp Glue Catalog cho partitioning tự động. Không degrade performance nhờ columnar formats như Parquet/ORC + workgroups cho concurrency.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, đánh dấu ✅ đúng hoặc ❌ sai, giữ nguyên văn bản gốc:
-
✅ Partition the data based on order date. Use Amazon Athena to query the data.
🟢 Đúng vì: Partition theo đúng trường query (order date) → pruning hiệu quả. Athena là lựa chọn serverless, cost-effective nhất cho ad-hoc queries trên S3 (không cần provision cluster). Theo AWS best practices 2026, Athena với federated queries và ML integrations đảm bảo scale tự động, chi phí chỉ tính theo scan data thực tế. -
❌ Partition the data based on order date. Use Amazon Redshift to query the data.
🔴 Sai vì: Mặc dù partition order date tốt, nhưng Redshift là managed data warehouse yêu cầu COPY data từ S3 vào cluster (thêm chi phí storage/load, O&D cluster luôn chạy ~$0.25/giờ/node). Query weekly không justify cost cao (concurrency scaling đắt), performance degrade nếu không optimize Spectrum đúng. Không cost-effective bằng Athena. -
❌ Partition the data based on load date. Use Amazon EMR to query the data.
🔴 Sai vì: Partition theo load date (ngày load) không khớp query filter (order date) → không pruning hiệu quả, scan toàn bộ data → performance degrade khi volume tăng. EMR là cluster-based Hadoop/Spark (transient/phsi), cần provision/manage cluster mỗi lần query (thời gian setup 5-10p), chi phí cao hơn Athena (~2-3x cho small jobs), không serverless thuần. -
❌ Partition the data based on load date. Use Amazon Aurora to query the data.
🔴 Sai vì: Partition load date không phù hợp (order date mới đúng cho pruning). Aurora là relational OLTP DB (MySQL/PostgreSQL-compatible), không query trực tiếp S3 (cần ETL vào DB trước, thêm latency/cost). Scale kém với semi-structured data lớn, chi phí provisioned IOPS cao (~$0.10/GB/tháng), không cost-effective cho batch analytics.
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- Amazon Athena Documentation: Partitioning data in Athena – Best practices partitioning cho S3 queries.
- AWS Big Data Blog: Optimize Amazon Athena performance – Nhấn mạnh partition pruning với order date.
- AWS Well-Architected Framework - Analytics Lens: Khuyến nghị Athena cho cost-effective S3 querying (vs Redshift/EMR).
- Redshift Spectrum vs Athena Comparison: AWS re:Post – Athena rẻ hơn 5-10x cho infrequent queries.
Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ code Glue/Athena, hãy hỏi nhé!
The data engineer needs a solution to determine whether a specific set of values in the city and state columns of the primary dataset exactly match the same specific values in the reference dataset. The data engineer wants to use Data Quality Definition Language (DQDL) rules in an AWS Glue Data Quality job.
Which rule will meet these requirements?
- A DatasetMatch "reference” “city->ref_city, state->ref_state” = 1.0
- B Referentiallntegrity “city,state” “reference.{ref_city,ref_state}” = 1.0
- C DatasetMatch “reference” “city->ref_city, state->ref_state” = 100
- D Referentialintegrity “city,state” "reference.{ref_city,ref_state}” = 100
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào AWS Glue Data Quality (một tính năng của AWS Glue cho phép định nghĩa và thực thi các quy tắc chất lượng dữ liệu bằng ngôn ngữ DQDL - Data Quality Definition Language).
-
Bối cảnh: Có hai dataset chứa thông tin doanh số bán hàng theo thành phố (city) và bang (state):
- Primary dataset: Dataset chính cần kiểm tra.
- Reference dataset: Dataset tham chiếu làm "nguồn chuẩn".
-
Yêu cầu cụ thể: Kiểm tra xem tất cả các giá trị cụ thể trong cột
cityvàstatecủa primary dataset có khớp chính xác 100% (exactly match) với các giá trị tương ứng trong reference dataset không. Nghĩa là, mọi cặp giá trị (city, state) trong primary phải tồn tại trong reference (không được phép có giá trị "lạc loài" hoặc không khớp). -
Giải pháp mong muốn: Sử dụng quy tắc DQDL trong AWS Glue Data Quality job (cập nhật mới nhất đến 2026: AWS Glue Data Quality hỗ trợ DQDL v1.0+ với các rule như ReferentialIntegrity cho kiểm tra toàn vẹn tham chiếu giữa hai dataset).
Mục tiêu là đảm bảo tính toàn vẹn tham chiếu (referential integrity): Không có giá trị NULL hoặc không khớp, đạt tỷ lệ = 1.0 (100%).
📘 Tài liệu tham khảo:
- AWS Glue Data Quality DQDL Rules (cập nhật 2024-2026).
- AWS Glue Data Quality Developer Guide.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Referentiallntegrity “city,state” “reference.{ref_city,ref_state}” = 1.0
Lý do 🛠️:
- Quy tắc ReferentialIntegrity (chú ý: có thể viết là
ReferentialIntegrityvới chữ 'I' hoa chuẩn) được thiết kế chính xác để kiểm tra tính toàn vẹn tham chiếu giữa primary dataset và reference dataset. - Cú pháp chuẩn DQDL:
"city,state": Chỉ định các cột cần kiểm tra trong primary dataset (composite key: cặp city-state)."reference.{ref_city,ref_state}": Tham chiếu đến cột tương ứng trong reference dataset (dấu{}dùng để map multi-column).
- = 1.0: Threshold = 1.0 nghĩa là 100% các giá trị non-NULL trong primary phải tồn tại exactly trong reference. Nếu <1.0, job sẽ báo lỗi hoặc cảnh báo.
- Đây là rule hiệu quả nhất cho yêu cầu "exactly match" multi-column, chạy trên Spark engine của Glue, hỗ trợ large-scale data (cập nhật 2026: tối ưu hóa cho Glue 4.0 với Spark 3.3+).
🔍 Phân tích tất cả các phương án (đúng/sai)
Dưới đây là phân tích chi tiết từng lựa chọn. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt với lý do đúng/sai dựa trên DQDL spec mới nhất.
-
Phương án 1:
DatasetMatch "reference” “city->ref_city, state->ref_state” = 1.0
❌ Sai.
Lý do:DatasetMatchdùng để so sánh độ tương đồng (similarity) giữa hai dataset dựa trên cosine similarity hoặc fuzzy matching (không phải exact match). Cú pháp dùng->cho column mapping, nhưng nó tính score probabilistic (0-1), không đảm bảo referential integrity exactly (có thể khớp gần đúng). Không phù hợp cho "exactly match" strict như yêu cầu. Ngoài ra, dấu ngoặc kép không chuẩn (thiếu khớp). -
Phương án 2:
Referentiallntegrity “city,state” “reference.{ref_city,ref_state}” = 1.0
✅ Đúng.
Lý do: Như đã giải thích ở trên. Đây là rule chuẩn xác nhất cho kiểm tra exact existence của multi-column values từ primary trong reference. Threshold1.0đảm bảo 0% violation. Hỗ trợ full DQDL syntax (dấu{}cho multi-ref columns). -
Phương án 3:
DatasetMatch “reference” “city->ref_city, state->ref_state” = 100
❌ Sai.
Lý do: Tương tự phương án 1,DatasetMatchkhông dành cho exact referential check mà chỉ đo similarity (dù =1.0 cũng không strict). Hơn nữa, threshold=100không hợp lệ trong DQDL (phải là 0.0-1.0, không phải % kiểu 100). Sẽ gây syntax error khi chạy Glue job. -
Phương án 4:
Referentialintegrity “city,state” "reference.{ref_city,ref_state}” = 100
❌ Sai.
Lý do: RuleReferentialIntegrityđúng về chức năng (exact match referential), cú pháp cơ bản OK, nhưng threshold =100 không hợp lệ (DQDL chỉ chấp nhận 0.0-1.0). Phải dùng=1.0tương đương 100%. Ngoài ra, chữ thườngintegritycó thể gây issue (chuẩn làReferentialIntegrity). Job sẽ fail validation.
💡 Lời khuyên thực hành
- Để implement: Tạo Glue Data Quality job qua Console/CLI, attach rule này vào DQDL script, chạy trên Glue 4.0 (Spark 3.3+, hỗ trợ DQ tốt hơn). Test với sample data để verify.
- Nếu scale lớn: Kết hợp với Glue Crawler và Lake Formation cho governance (cập nhật 2026).
Hy vọng phân tích này giúp bạn ôn thi AWS Certified DevOps Engineer Professional hiệu quả! 🚀
The on-premises database is continuously updated. The company must ensure that the data in Amazon Redshift is updated as quickly as possible.
Which solution will meet these requirements?
- A Use the pg_dump utility to generate a backup of the PostgreSQL database. Use the AWS Schema Conversion Tool (AWS SCT) to upload the backup to Amazon Redshift. Set up a cron job to perform a backup. Upload the backup to Amazon Redshift every night.
- B Create an AWS Database Migration Service (AWS DMS) full-load task. Set Amazon Redshift as the target. Configure the task to use the change data capture (CDC) feature.
- C Use the pg_dump utility to generate a backup of the PostgreSQL database. Upload the backup to an Amazon S3 bucket. Use the COPY command to import the data into Amazon Redshift.
- D Create an AWS Database Migration Service (AWS DMS) full-load task. Set Amazon Redshift as the target. Configure the task to perform a full load of the database to Amazon Redshift every night.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc di chuyển dữ liệu khách hàng từ cơ sở dữ liệu PostgreSQL on-premises sang Amazon Redshift (một data warehouse của AWS). 🛤️ Công ty đã thiết lập kết nối VPN giữa on-premises và AWS, đảm bảo kết nối an toàn và ổn định. Điểm quan trọng nhất: cơ sở dữ liệu on-premises được cập nhật liên tục (continuously updated), và yêu cầu là dữ liệu trong Redshift phải được cập nhật nhanh nhất có thể (as quickly as possible).
📌 Mục tiêu chính: Không chỉ migrate ban đầu mà còn cần replication thời gian thực hoặc gần thực tế (near real-time) để đồng bộ thay đổi liên tục, tránh độ trễ lớn. Giải pháp phải tận dụng kết nối VPN sẵn có, hỗ trợ PostgreSQL làm nguồn và Redshift làm đích đến.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create an AWS Database Migration Service (AWS DMS) full-load task. Set Amazon Redshift as the target. Configure the task to use the change data capture (CDC) feature.
Lý do chọn đáp án này 🏆:
- AWS DMS (Database Migration Service) hỗ trợ full load (di chuyển toàn bộ dữ liệu ban đầu) kết hợp CDC (Change Data Capture) để capture và replicate các thay đổi liên tục từ PostgreSQL (nguồn hỗ trợ CDC tốt) sang Redshift (target endpoint chính thức).
- CDC đảm bảo cập nhật gần real-time (thường trong vài giây đến phút), phù hợp với yêu cầu "updated as quickly as possible" cho database continuously updated.
- Sử dụng VPN sẵn có để kết nối source on-premises. Đây là best practice theo AWS cho migration + ongoing replication đến Redshift (cập nhật đến 2026, DMS v3.4+ hỗ trợ Redshift streaming ingestion cho hiệu suất cao hơn).
🔍 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt với lý do đúng/sai rõ ràng:
-
❌ SAI: Use the pg_dump utility to generate a backup of the PostgreSQL database. Use the AWS Schema Conversion Tool (AWS SCT) to upload the backup to Amazon Redshift. Set up a cron job to perform a backup. Upload the backup to Amazon Redshift every night.
Giải thích sai: pg_dump chỉ tạo backup tĩnh (snapshot), SCT chủ yếu dùng cho schema conversion, không hỗ trợ replication liên tục. Cron job chạy hàng đêm gây độ trễ lớn (24h), không đáp ứng "updated as quickly as possible" cho dữ liệu continuously updated. Không hiệu quả cho production. -
✅ ĐÚNG: Create an AWS Database Migration Service (AWS DMS) full-load task. Set Amazon Redshift as the target. Configure the task to use the change data capture (CDC) feature.
Giải thích đúng (như phần trên): Kết hợp full load + CDC cho migration ban đầu và đồng bộ ongoing, tốc độ nhanh (near real-time), tận dụng VPN. Hỗ trợ PostgreSQL source với logical replication (WAL-based CDC). -
❌ SAI: Use the pg_dump utility to generate a backup of the PostgreSQL database. Upload the backup to an Amazon S3 bucket. Use the COPY command to import the data into Amazon Redshift.
Giải thích sai: Đây chỉ là one-time migration (backup → S3 → COPY), không có cơ chế đồng bộ thay đổi liên tục. pg_dump không capture delta changes, dẫn đến dữ liệu Redshift lạc hậu so với source continuously updated. Không phù hợp yêu cầu. -
❌ SAI: Create an AWS Database Migration Service (AWS DMS) full-load task. Set Amazon Redshift as the target. Configure the task to perform a full load of the database to Amazon Redshift every night.
Giải thích sai: DMS full load lặp lại hàng đêm chỉ reload toàn bộ dữ liệu (rất tốn tài nguyên, downtime cao), không dùng CDC nên bỏ lỡ thay đổi intraday. Độ trễ 24h, không "as quickly as possible". DMS yêu cầu CDC cho ongoing sync.
📘 Tài liệu tham khảo (cập nhật mới nhất đến 2026)
- AWS DMS User Guide: Using a PostgreSQL database as a source for AWS DMS và Using Amazon Redshift as a target – Xác nhận CDC support cho PostgreSQL → Redshift.
- AWS Redshift Docs: Migrate data using AWS DMS – Nhấn mạnh full load + CDC cho continuous replication.
- AWS Blog (2024-2026 updates): DMS enhancements với Redshift Streaming cho low-latency CDC (re:Post và re:Invent 2025 announcements).
- Best Practices: AWS Well-Architected Framework – Data Transfer pillar khuyến nghị DMS cho hybrid replication.
🛠️ Lời khuyên DevOps: Sau setup DMS, monitor CloudWatch metrics (CDC lag, full load duration) và enable Multi-AZ cho DMS replication instance để HA. Test CDC với workload giả lập trước production! 🚀
Which solution will meet these requirements in the MOST cost-effective way?
- A Create an Amazon RDS MySQL cluster. Use AWS Glue to transform and load the CSV and JSON files into database tables. Provide the data analysts access to the MySQL cluster.
- B Create an AWS Glue DataBrew project that contains the new data. Make the DataBrew project available to the data analysts.
- C Store the data in an Amazon S3 bucket. Use an AWS Glue crawler to catalog the S3 bucket as tables. Create an Amazon Athena workgroup that has a data usage threshold. Grant the data analysts access to the Athena workgroup.
- D Load the data into Super-fast, Parallel, In-memory Calculation Engine (SPICE) in Amazon QuickSight. Allow the data analysts to create analyses and dashboards in QuickSight.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi mô tả một công ty có các bộ dữ liệu mới ở định dạng CSV và JSON. Một data engineer cần làm cho dữ liệu này có sẵn (available) cho nhóm data analysts, những người sẽ phân tích dữ liệu bằng các truy vấn SQL. Yêu cầu chính là chọn giải pháp tiết kiệm chi phí nhất (MOST cost-effective).
🛠️ Phân tích yêu cầu chính:
- Dữ liệu thô (CSV/JSON) → Không cần xử lý phức tạp ban đầu.
- Người dùng cuối: Data analysts → Cần truy vấn SQL (không phải visualization hay data prep).
- Tiêu chí: Cost-effective → Ưu tiên serverless, pay-per-use, tránh tài nguyên always-on đắt đỏ như RDS.
- AWS services liên quan: Tập trung vào lưu trữ rẻ (S3), catalog (Glue), query engine serverless (Athena).
📈 Mục tiêu: Làm dữ liệu queryable bằng SQL với chi phí thấp nhất, không cần quản lý infrastructure.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Store the data in an Amazon S3 bucket. Use an AWS Glue crawler to catalog the S3 bucket as tables. Create an Amazon Athena workgroup that has a data usage threshold. Grant the data analysts access to the Athena workgroup.
Lý do chọn 🏆:
- Tiết kiệm chi phí nhất (pay-per-query với Athena serverless, S3 storage rẻ ~$0.023/GB/tháng).
- Hỗ trợ SQL trực tiếp: Athena cho phép query CSV/JSON trên S3 bằng SQL chuẩn, không cần load dữ liệu vào DB.
- Glue Crawler: Tự động infer schema, catalog dữ liệu thành tables trong AWS Glue Data Catalog (miễn phí crawl cơ bản).
- Athena Workgroup với data usage threshold: Giới hạn chi phí (threshold) để tránh bill bất ngờ, phù hợp team analysts (tính năng mới nhất AWS 2023-2026).
- Scalable & NoOps: Không cần provision server, chỉ scan dữ liệu khi query (~$5/TB scanned).
🔍 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên kiến thức AWS mới nhất (2026).
-
❌ Phương án SAI: Create an Amazon RDS MySQL cluster. Use AWS Glue to transform and load the CSV and JSON files into database tables. Provide the data analysts access to the MySQL cluster.
Lý do sai: RDS MySQL là managed relational DB always-on, chi phí cao (~$0.1/giờ/instance + storage), không cost-effective cho dữ liệu lớn/ít query. Phải dùng Glue ETL để transform/load (thêm chi phí compute). Không phù hợp dữ liệu semi-structured (JSON), dễ schema mismatch. Athena rẻ hơn gấp nhiều lần cho query ad-hoc. -
❌ Phương án SAI: Create an AWS Glue DataBrew project that contains the new data. Make the DataBrew project available to the data analysts.
Lý do sai: AWS Glue DataBrew là tool data preparation/cleaning/visual profiling (ra mắt 2021, cập nhật 2026), không hỗ trợ SQL queries trực tiếp. Analysts chỉ clean data qua UI recipes, không query SQL như yêu cầu. Không cost-effective vì tính phí per-minute job (~$1/30 job-hours), và không làm dữ liệu "available" cho SQL analysis. -
✅ Phương án ĐÚNG: Store the data in an Amazon S3 bucket. Use an AWS Glue crawler to catalog the S3 bucket as tables. Create an Amazon Athena workgroup that has a data usage threshold. Grant the data analysts access to the Athena workgroup.
Lý do đúng (tóm tắt lại): Kết hợp S3 (lưu trữ rẻ) + Glue Crawler (catalog schema tự động) + Athena (serverless SQL query, workgroup control cost). Hỗ trợ CSV/JSON native, query federated, tích hợp Lake Formation cho governance. Cost: ~$5/TB scanned + S3 storage, thấp nhất cho workload này (xác nhận AWS Well-Architected Data Analytics). -
❌ Phương án SAI: Load the data into Super-fast, Parallel, In-memory Calculation Engine (SPICE) in Amazon QuickSight. Allow the data analysts to create analyses and dashboards in QuickSight.
Lý do sai: SPICE là in-memory engine của QuickSight (BI tool), dùng cho visualization/dashboards, không hỗ trợ SQL queries trực tiếp (chỉ import data vào dataset). Phải load dữ liệu vào SPICE (chi phí ~$0.45/1M rows + capacity units), không cost-effective cho raw data lớn. Analysts cần SQL thuần, không phải viz.
📘 Tài liệu tham khảo
- AWS Athena Documentation (2026): https://docs.aws.amazon.com/athena/latest/ug/what-is.html → Serverless querying on S3.
- AWS Glue Crawlers (2026): https://docs.aws.amazon.com/glue/latest/dg/crawler-s3.html → Catalog S3 data.
- Athena Workgroups & Cost Controls: https://docs.aws.amazon.com/athena/latest/ug/workgroups.html#workgroup-data-usage-control → Threshold feature.
- AWS Well-Architected Framework - Data Analytics Lens: https://docs.aws.amazon.com/wellarchitected/latest/analytics-lens → Cost-effective patterns cho SQL on S3.
- Pricing Calculator: https://calculator.aws → So sánh Athena vs RDS (Athena rẻ hơn 10x cho ad-hoc queries).
🛡️ Kết luận: Giải pháp Athena + S3 + Glue là best practice cho data lake querying, đảm bảo scalable và tiết kiệm! Nếu cần demo code hoặc deep-dive, hỏi thêm nhé! 🚀
A marketing team needs to join the Orders data with an Amazon Redshift table named Campaigns in the marketing team's data warehouse. The operational Aurora database must not be affected.
Which solution will meet these requirements with the LEAST operational effort?
- A Use AWS Database Migration Service (AWS DMS) Serverless to replicate the Orders table to Amazon Redshift. Create a materialized view in Amazon Redshift to join with the Campaigns table.
- B Use the Aurora zero-ETL integration with Amazon Redshift to replicate the Orders table. Create a materialized view in Amazon Redshift to join with the Campaigns table.
- C Use AWS Glue to replicate the Orders table to Amazon Redshift. Create a materialized view in Amazon Redshift to join with the Campaigns table.
- D Use federated queries to query the Orders table directly from Aurora. Create a materialized view in Amazon Redshift to join with the Campaigns table.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một tình huống thực tế của công ty bán lẻ sử dụng Amazon Aurora (một cơ sở dữ liệu quan hệ tương thích MySQL/PostgreSQL, hiệu suất cao) để lưu trữ bảng Orders chứa hàng tỷ rows và hơn 100.000 giao dịch/giây.
- Yêu cầu chính 1: Tạo báo cáo operational (báo cáo hoạt động thời gian thực) từ bảng Orders với minimal latency (độ trễ thấp nhất), nghĩa là cần truy vấn nhanh chóng mà không làm chậm hệ thống chính.
- Yêu cầu chính 2: Nhóm marketing cần join dữ liệu Orders với bảng Campaigns trong Amazon Redshift (kho dữ liệu phân tích), nhưng không được ảnh hưởng đến Aurora database hoạt động (operational DB phải ổn định, không bị overload bởi query join).
- Mục tiêu: Giải pháp LEAST operational effort (ít nỗ lực vận hành nhất), tức là tự động hóa cao, ít cấu hình thủ công, không cần quản lý server/ETL job phức tạp.
Vấn đề cốt lõi: Aurora phù hợp cho OLTP (transactional workload cao), nhưng Redshift cho OLAP (analytics/join lớn). Cần replicate dữ liệu từ Aurora sang Redshift một cách near-real-time (gần thời gian thực) để join mà không query trực tiếp Aurora (tránh ảnh hưởng performance).
📘 Kiến thức AWS cập nhật 2026: Tính năng Aurora zero-ETL integration with Amazon Redshift (ra mắt 2023, stable đến 2026) là giải pháp serverless, tự động replicate dữ liệu từ Aurora PostgreSQL sang Redshift với zero transformation và minimal latency (<1 phút), lý tưởng cho workload lớn như billions rows + high TPS.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use the Aurora zero-ETL integration with Amazon Redshift to replicate the Orders table. Create a materialized view in Amazon Redshift to join with the Campaigns table.
Lý do 🛠️:
- Zero-ETL tự động replicate dữ liệu từ Aurora sang Redshift near-real-time (sub-minute latency), hỗ trợ billions rows và high TPS mà không cần ETL job thủ công.
- Không ảnh hưởng Aurora: Replication read-only từ Aurora, dùng Amazon Aurora PostgreSQL (hỗ trợ zero-ETL từ 2023).
- Materialized view trong Redshift lưu kết quả join Orders + Campaigns, refresh tự động, cho phép query nhanh cho báo cáo operational.
- LEAST operational effort: Chỉ enable zero-ETL qua console/API (1-click), serverless, auto-scale, không manage DMS/Glue tasks.
- Phù hợp minimal latency cho operational reports và analytics join.
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ Phương án SAI: Use AWS Database Migration Service (AWS DMS) Serverless to replicate the Orders table to Amazon Redshift. Create a materialized view in Amazon Redshift to join with the Campaigns table.
Giải thích: DMS Serverless hỗ trợ replication continuous, nhưng yêu cầu tạo endpoints, tasks, rules thủ công, monitor lag, handle schema changes → operational effort cao (không least). Latency cao hơn zero-ETL (có thể >1 phút), không optimized cho Aurora-to-Redshift như zero-ETL. DMS phù hợp migration hơn real-time analytics. -
✅ Phương án ĐÚNG: Use the Aurora zero-ETL integration with Amazon Redshift to replicate the Orders table. Create a materialized view in Amazon Redshift to join with the Campaigns table.
Giải thích: Như trên, zero-ETL là native integration (Aurora PostgreSQL → Redshift), serverless + automatic (enable via AWS Console/SQL), latency thấp nhất (<60s), scale tự động cho high TPS. Materialized view tối ưu join mà không query Aurora trực tiếp → least effort, không ảnh hưởng operational DB. -
❌ Phương án SAI: Use AWS Glue to replicate the Orders table to Amazon Redshift. Create a materialized view in Amazon Redshift to join with the Campaigns table.
Giải thích: AWS Glue là ETL service (serverless Spark jobs), cần viết script, schedule jobs (cron/Glue triggers), handle CDC (change data capture) thủ công → operational effort lớn (develop/maintain code, monitor failures). Latency cao (batch-based, không real-time), không phù hợp billions rows/high TPS so với zero-ETL native. -
❌ Phương án SAI: Use federated queries to query the Orders table directly from Aurora. Create a materialized view in Amazon Redshift to join with the Campaigns table.
Giải thích: Redshift Federated Query cho phép query Aurora trực tiếp (PUSH-down predicates), nhưng với billions rows + 100k TPS, gây latency cao (network overhead), ảnh hưởng Aurora performance (query lớn từ Redshift overload OLTP DB). Không replicate → materialized view chỉ join một phần, không lưu trữ dữ liệu bền vững cho reports. High effort monitor throttling/limits.
📚 Tài liệu tham khảo (AWS Docs cập nhật 2026)
- Aurora zero-ETL to Redshift: AWS Docs - Zero-ETL integrations with Amazon Aurora and Amazon Redshift ✅ (Hỗ trợ PostgreSQL 15+, auto-scale).
- Redshift Materialized Views: AWS Redshift Docs - Materialized Views 🛠️.
- So sánh DMS/Glue/Federated: AWS re:Invent 2023/2024 Sessions - DOP204 (Zero-ETL least effort cho use case này).
- Exam Tip (DOP-C02): Zero-ETL là câu hỏi hot trong DevOps Pro 2024-2026 cho hybrid OLTP/OLAP workloads.
Giải pháp này đảm bảo scalable, cost-effective với pay-per-use! 🚀
The files are stored in an Amazon S3 bucket. Files are no larger than 5 MB.
A data engineer is developing the extract, transform, and load (ETL) pipeline for the CSV files. The data engineer configured a Redshift cluster and an AWS Lambda function that copies the data out of the files into the Redshift cluster.
Which additional steps should the data engineer perform to meet these requirements?
- A Configure the bucket to send S3 event notifications to Amazon EventBridge. Configure an EventBridge rule that matches S3 new object created events. Set the Lambda function as the target.
- B Configure the $3 bucket to send S3 event notifications to an Amazon Simple Queue Service (Amazon SQS) queue. Configure the Lambda function to process the queue.
- C Configure AWS Database Migration Service (AWS DMS) to stream new S3 objects to a data stream in Amazon Kinesis Data Streams. Set the Lambda function as the target of the data stream.
- D Configure an Amazon EventBridge rule that matches S3 new object created events. Set an Amazon Simple Queue Service (Amazon SQS) queue as the target of the rule. Configure the Lambda function to process the queue.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một công ty đang xây dựng ứng dụng mới để ingest (hấp thụ dữ liệu) các file CSV vào Amazon Redshift (kho dữ liệu phân tích). Ứng dụng đã có frontend hoàn thiện. Các file CSV được lưu trữ trong Amazon S3 bucket, với kích thước không lớn hơn 5 MB (rất nhỏ, phù hợp xử lý batch nhanh).
Một data engineer đang phát triển ETL pipeline (Extract, Transform, Load):
- Đã config Redshift cluster.
- Đã config AWS Lambda function để copy dữ liệu từ file CSV ra Redshift.
Yêu cầu chính: Data engineer cần thực hiện các bước bổ sung để trigger (kích hoạt) Lambda tự động khi có file mới upload vào S3, đảm bảo pipeline ETL hoạt động đáng tin cậy, scalable và decoupled (tách biệt).
Mục tiêu: Xây dựng luồng event-driven từ S3 → Lambda → Redshift, tận dụng kích thước file nhỏ để xử lý nhanh, tránh overload Lambda nếu có nhiều file đồng thời. Theo best practice AWS (cập nhật 2024-2026), ưu tiên SQS làm buffer để retry và reliability cao.
📘 Tài liệu tham khảo:
- AWS Well-Architected Framework: Reliability Pillar (S3 Event Notifications với SQS/Lambda).
- AWS Docs: S3 Event Notifications & Lambda with S3/SQS (phiên bản mới nhất 2026 hỗ trợ EventBridge Pipes cho advanced routing, nhưng không bắt buộc ở đây).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Configure the S3 bucket to send S3 event notifications to an Amazon Simple Queue Service (Amazon SQS) queue. Configure the Lambda function to process the queue.
Lý do:
🛠️ Đây là best practice cho ETL pipeline với S3 events:
- S3 Event Notifications gửi trực tiếp đến SQS queue (hỗ trợ fan-out, durable queueing).
- Lambda được trigger bởi SQS (event source mapping), tự động poll và process messages → copy CSV vào Redshift.
- Ưu điểm: Decoupling (S3 không direct gọi Lambda, tránh throttling nếu nhiều file), at-least-once delivery, retry tự động, scalable với file nhỏ <5MB. Phù hợp serverless ETL (Glue thay thế nếu complex hơn, nhưng ở đây Lambda đủ).
- Không cần layer thừa như EventBridge/DMS, giữ đơn giản và cost-effective.
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn giữ nguyên văn bản gốc bằng tiếng Anh, với lý do đúng/sai bằng tiếng Việt. Sử dụng pattern S3 → Trigger → Lambda làm cơ sở đánh giá.
-
❌ Phương án SAI: Configure the bucket to send S3 event notifications to Amazon EventBridge. Configure an EventBridge rule that matches S3 new object created events. Set the Lambda function as the target.
Giải thích: Sai vì thêm layer EventBridge không cần thiết. S3 có thể gửi events trực tiếp đến EventBridge (từ 2020), nhưng config bucket → EventBridge → rule → Lambda tạo overhead (latency cao hơn, cost thêm). Không optimal cho simple ETL; EventBridge phù hợp routing complex hơn (nhiều sources). -
✅ Phương án ĐÚNG: Configure the $3 bucket to send S3 event notifications to an Amazon Simple Queue Service (Amazon SQS) queue. Configure the Lambda function to process the queue.
Giải thích: (Như phần trên) Hoàn hảo cho reliability: S3 → SQS (queue events), Lambda poll SQS → process CSV → Redshift. Xử lý concurrent files tốt, dead-letter queue cho error handling. Lưu ý: "$3" là lỗi typo của "S3". -
❌ Phương án SAI: Configure AWS Database Migration Service (AWS DMS) to stream new S3 objects to a data stream in Amazon Kinesis Data Streams. Set the Lambda function as the target of the data stream.
Giải thích: Hoàn toàn sai vì AWS DMS không hỗ trợ S3 objects. DMS dùng cho database migration/replication (e.g., RDS → Redshift), không stream file S3 sang Kinesis. Kinesis phù hợp streaming real-time lớn, thừa thãi cho file <5MB static CSV. -
❌ Phương án SAI: Configure an Amazon EventBridge rule that matches S3 new object created events. Set an Amazon Simple Queue Service (Amazon SQS) queue as the target of the rule. Configure the Lambda function to process the queue.
Giải thích: Sai vì EventBridge rule cần source từ S3 events, nhưng config này gián tiếp: EventBridge "kéo" events từ S3 (không phải bucket push trực tiếp). Thêm latency/complexity so với S3 → SQS direct. EventBridge Pipes (mới 2023+) có thể simplify, nhưng vẫn không phải lựa chọn tối ưu nhất.
Kết luận 💡: Chọn phương án dùng S3 + SQS + Lambda để build resilient ETL – scalable đến hàng nghìn file/ngày mà không lo mất events! Nếu production, thêm IAM roles, VPC cho Redshift, và monitoring CloudWatch.