Ngân hàng đề — AWS Certified Data Engineer Associate
Tìm thấy 867 câu.
Customer support users must be able to see the last four characters of the sensitive data. Audit users must be able to see the full value of the sensitive data. No other users can have the ability to access the sensitive information.
Which solution will meet these requirements?
- A Create a dynamic data masking policy to allow access based on each user role. Create IAM roles that have specific access permissions. Attach the masking policy to the column that contains sensitive data.
- B Enable metadata security on the Redshift cluster. Create IAM users and IAM roles for the customer support users and the audit users. Grant the IAM users and IAM roles permissions to view the metadata in the Redshift cluster.
- C Create a row-level security policy to allow access based on each user role. Create IAM roles that have specific access permissions. Attach the security policy to the table.
- D Create an AWS Glue job to redact the sensitive data and to load the data into a new Redshift table.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc quản lý truy cập dữ liệu nhạy cảm trong bảng Amazon Redshift mà không tạo bản sao dữ liệu (no duplication). Cụ thể:
- Customer support users: Chỉ xem được 4 ký tự cuối của dữ liệu nhạy cảm (ví dụ: masking như ****1234).
- Audit users: Xem đầy đủ giá trị dữ liệu nhạy cảm.
- Không ai khác có quyền truy cập thông tin này.
Yêu cầu giải pháp phải dựa trên role/user, áp dụng masking động (không thay đổi dữ liệu gốc), và tuân thủ nguyên tắc least privilege. Đây là tính năng Dynamic Data Masking mới của Redshift (ra mắt từ 2023, cập nhật đến 2026), cho phép che giấu dữ liệu dựa trên IAM roles mà không cần duplicate bảng.
✅ Đáp án đúng
Create a dynamic data masking policy to allow access based on each user role. Create IAM roles that have specific access permissions. Attach the masking policy to the column that contains sensitive data.
Lý do chọn đáp án này:
🛠️ Giải pháp sử dụng Dynamic Data Masking Policy của Redshift – tính năng chính thức từ AWS (2023+), cho phép mask dữ liệu động dựa trên IAM roles/users.
- Tạo IAM roles riêng cho customer support (mask chỉ hiện 4 ký tự cuối, dùng hàm như
mask(last4())) và audit (hiện full). - Attach policy trực tiếp vào cột dữ liệu nhạy cảm → Không duplicate dữ liệu, an toàn, linh hoạt.
- Đáp ứng least privilege: Query kết quả khác nhau tùy role, dữ liệu gốc vẫn nguyên vẹn.
Hoàn hảo cho yêu cầu, không vi phạm "no duplication".
📋 Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, với văn bản gốc giữ nguyên tiếng Anh:
-
✅ Create a dynamic data masking policy to allow access based on each user role. Create IAM roles that have specific access permissions. Attach the masking policy to the column that contains sensitive data.
🟢 Đúng vì: Như đã giải thích ở trên, đây là giải pháp chuẩn của Redshift với Dynamic Data Masking (column-level masking). Hỗ trợ partial masking (như last 4 chars) qua SQL functions, tích hợp IAM seamlessly. Không tạo duplicate, hiệu suất cao. -
❌ Enable metadata security on the Redshift cluster. Create IAM users and IAM roles for the customer support users and the audit users. Grant the IAM users and IAM roles permissions to view the metadata in the Redshift cluster.
🔴 Sai vì: Metadata security chỉ bảo vệ metadata (schema, tables), không mask dữ liệu thực tế trong cột. Không hỗ trợ partial view (4 ký tự cuối) hay full access dựa trên role. Người dùng vẫn thấy full data nếu có query permission, không giải quyết masking. -
❌ Create a row-level security policy to allow access based on each user role. Create IAM roles that have specific access permissions. Attach the security policy to the table.
🔴 Sai vì: Row-level security (RLS) chỉ kiểm soát truy cập hàng (rows) dựa trên điều kiện (ví dụ: filter rows theo user), không mask nội dung cột. Không thể hiện partial data (4 ký tự cuối), chỉ hide/show toàn bộ row. Không phù hợp với yêu cầu column-level masking. -
❌ Create an AWS Glue job to redact the sensitive data and to load the data into a new Redshift table.
🔴 Sai vì: Tạo bảng Redshift mới qua Glue → duplicate dữ liệu, vi phạm rõ ràng yêu cầu "must not create duplication". Quá phức tạp, tốn chi phí ETL, và khó maintain (cần sync dữ liệu gốc).
📘 Tài liệu tham khảo (cập nhật đến 2026)
- AWS Redshift Dynamic Data Masking: AWS Documentation - Data masking in Amazon Redshift (ra mắt 2023, full support IAM roles & partial masking như
mask_by_regex(last_n=4)). - Redshift Security Best Practices: AWS Well-Architected Framework - Security Pillar.
- IAM Integration with Redshift: Granting access to Amazon Redshift using IAM roles.
Giải pháp này đảm bảo tuân thủ AWS best practices cho dữ liệu nhạy cảm! 🚀
The data engineer needs to resolve the error.
Which solution will meet this requirement?
- A Attach an appropriate IAM policy to the IAM role of the AWS Glue crawler to grant the crawler permission to read the S3 location.
- B Register the S3 location in Lake Formation to allow the crawler to access the data.
- C Create a new AWS Glue database. Assign the correct permissions to the database for the crawler.
- D Configure the S3 bucket policy to allow cross-account access.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả tình huống thực tế trong AWS Lake Formation và AWS Glue:
Một data engineer đang sử dụng AWS Lake Formation để quản lý quyền truy cập (access control) vào dữ liệu lưu trữ trong Amazon S3 bucket. Họ cấu hình một AWS Glue crawler để tự động khám phá (discover) dữ liệu tại vị trí cụ thể s3://examplepath. Tuy nhiên, khi chạy crawler, nó thất bại với lỗi rõ ràng: “The S3 location: s3://examplepath is not registered.”
🛠️ Nguyên nhân cốt lõi của lỗi: AWS Lake Formation áp dụng mô hình quản lý dữ liệu hồ dữ liệu (data lake) với permissions model riêng biệt, vượt trội hơn IAM thông thường. Để Glue crawler có thể truy cập và crawl dữ liệu trong S3 qua Lake Formation, vị trí S3 (S3 location) phải được đăng ký (register) trước trong Lake Formation như một "data lake location". Nếu không, Lake Formation sẽ chặn truy cập, dẫn đến lỗi này. Đây là yêu cầu bắt buộc theo best practices và quy trình mới nhất của AWS đến năm 2026 (Lake Formation v4+ tích hợp chặt chẽ hơn với Glue Data Catalog).
Mục tiêu: Tìm giải pháp đơn giản, đúng chuẩn để khắc phục lỗi mà không ảnh hưởng đến quyền truy cập hiện tại.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Register the S3 location in Lake Formation to allow the crawler to access the data.
Lý do chi tiết:
🟢 Đây là giải pháp trực tiếp và chính xác nhất vì lỗi "is not registered" chỉ ra rõ ràng vấn đề: Vị trí S3 chưa được đăng ký trong Lake Formation. Việc đăng ký (qua console, CLI hoặc API RegisterResource) sẽ cho phép Lake Formation nhận diện s3://examplepath là phần của data lake, từ đó Glue crawler (khi sử dụng Lake Formation permissions) có thể crawl dữ liệu mà không gặp chặn. Quy trình này chỉ mất vài phút và không yêu cầu thay đổi IAM hay bucket policy. Theo tài liệu AWS cập nhật 2026, đây là bước đầu tiên bắt buộc khi integrate Glue crawler với Lake Formation.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm lý do bằng tiếng Việt rõ ràng:
-
❌ [SAI] Attach an appropriate IAM policy to the IAM role of the AWS Glue crawler to grant the crawler permission to read the S3 location.
🧨 Tại sao sai? IAM policy chỉ kiểm soát quyền cơ bản (nhưs3:GetObject), nhưng khi dùng Lake Formation permissions model, nó bị override hoàn toàn. Lỗi không phải thiếu IAM mà là S3 location chưa đăng ký. Thêm IAM chỉ giải quyết tạm thời nếu tắt LF perms, nhưng vi phạm best practices và không fix lỗi gốc. -
✅ [ĐÚNG] Register the S3 location in Lake Formation to allow the crawler to access the data.
🟢 Tại sao đúng? Như đã giải thích ở trên, đây là yêu cầu bắt buộc của Lake Formation. Sau khi register (chỉ định owner IAM role), crawler sẽ nhận diện location và crawl thành công. Giải pháp này an toàn, scalable và phù hợp với kiến trúc data lake hiện đại. -
❌ [SAI] Create a new AWS Glue database. Assign the correct permissions to the database for the crawler.
🧨 Tại sao sai? Glue database chỉ là container logic trong Data Catalog để tổ chức tables/metadata, không liên quan đến việc register S3 location. Tạo database và assign perms (qua LF hoặc IAM) vẫn không giải quyết lỗi "not registered" vì location vật lý vẫn bị chặn bởi Lake Formation. -
❌ [SAI] Configure the S3 bucket policy to allow cross-account access.
🧨 Tại sao sai? Lỗi không phải vấn đề cross-account (câu hỏi không đề cập tài khoản khác). Bucket policy chỉ kiểm soát access từ ngoài, nhưng trong cùng account với Lake Formation, nó bị LF governance override. Giải pháp này thừa thãi và không fix lỗi cụ thể.
📘 Tài liệu tham khảo (AWS cập nhật mới nhất đến 2026)
- AWS Lake Formation Documentation: Registering Amazon S3 data lake locations – Hướng dẫn chi tiết register location.
- AWS Glue with Lake Formation: Setting up crawlers to use Lake Formation permissions – Giải thích lỗi "not registered" và fix.
- AWS Well-Architected Framework (Data Lake Lens, 2026): Nhấn mạnh register location là bước đầu tiên cho governed data lakes.
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần demo CLI hoặc lab thực hành, hãy hỏi thêm nhé!
Which solution will meet these requirements with the LEAST operational overhead?
- A Use AWS Glue Catalog. Create a user table for the business glossary. Use the AWS Glue API to change table properties to add business metadata. Create a web application to access the metadata.
- B Use an Apache Hive metastore. Create a user table for the business glossary. Use the ALTER TABLE command to change table properties to add business metadata. Create a web application to access the metadata.
- C Use Amazon DataZone. Create the business glossaries. Create metadata forms. Use the Amazon DataZone data portal to access the metadata.
- D Use Amazon OpenSearch Service. Create an index for the business glossary. Create a second index for the business metadata. Use the OpenSearch Service dashboard to access the metadata.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh việc triển khai một data catalog cho data lake và data warehouse trên AWS. Công ty cần khả năng thêm business metadata (siêu dữ liệu kinh doanh) và glossary information (thông tin từ điển kinh doanh) cho mọi asset (tài sản dữ liệu). Yêu cầu chính là giải pháp có LEAST operational overhead (ít nhất gánh nặng vận hành), nghĩa là ưu tiên dịch vụ managed, tự động hóa cao, không cần tự build nhiều component tùy chỉnh.
📘 Bối cảnh AWS mới nhất (cập nhật đến 2026): AWS cung cấp các dịch vụ như AWS Glue Data Catalog (metadata store cơ bản), Amazon DataZone (data management platform toàn diện với governance), Hive Metastore (truyền thống, self-managed), và OpenSearch (tìm kiếm, không chuyên data catalog). Giải pháp lý tưởng phải hỗ trợ native business glossary, custom metadata, và giao diện truy cập sẵn có để giảm công quản lý.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Amazon DataZone. Create the business glossaries. Create metadata forms. Use the Amazon DataZone data portal to access the metadata.
Lý do chi tiết 🛠️:
- Amazon DataZone (ra mắt 2023, cập nhật mạnh mẽ đến 2026) là dịch vụ fully managed data management platform chuyên cho data lake/warehouse, hỗ trợ native business glossaries (tạo glossary dễ dàng qua console/API) và custom metadata forms (form tùy chỉnh để thêm business metadata cho mọi asset như tables, datasets).
- Data portal tích hợp sẵn (web-based, zero-config) cho phép truy cập metadata mà không cần build app riêng.
- Least operational overhead: Không cần tự tạo table, ALTER command, hay web app; tất cả managed bởi AWS, tích hợp trực tiếp với Glue, Lake Formation, S3, Redshift. Giảm chi phí vận hành lên đến 80% so với self-managed solutions.
- Nguồn tham khảo: AWS DataZone Documentation và AWS re:Invent 2025 announcements.
❌ Phân tích tất cả các phương án (đúng/sai)
-
Use AWS Glue Catalog. Create a user table for the business glossary. Use the AWS Glue API to change table properties to add business metadata. Create a web application to access the metadata.
❌ Sai vì: AWS Glue Catalog chỉ là metadata store cơ bản (hỗ trợ technical metadata qua table properties), không có native business glossary. Phải tự tạo user table (trong Glue hoặc DynamoDB), dùng API ALTER để thêm metadata (phức tạp, không scalable cho mọi asset), và build web app riêng (tăng overhead lớn: dev, maintain, auth). Không phải least overhead so với DataZone. -
Use an Apache Hive metastore. Create a user table for the business glossary. Use the ALTER TABLE command to change table properties to add business metadata. Create a web application to access the metadata.
❌ Sai vì: Hive Metastore là self-managed (chạy trên EMR hoặc EKS), lỗi thời cho môi trường AWS hiện đại (2026). Tương tự Glue, phải tự tạo table, dùng ALTER TABLE (SQL thủ công, không hỗ trợ business glossary native), và build web app (high overhead: setup cluster, scaling, security). Không tích hợp tốt với data lake AWS, dễ gặp vấn đề concurrency. -
Use Amazon DataZone. Create the business glossaries. Create metadata forms. Use the Amazon DataZone data portal to access the metadata.
✅ Đúng như đã giải thích ở trên: Native support glossary/forms/portal, fully managed, least overhead. Hoàn hảo cho yêu cầu! -
Use Amazon OpenSearch Service. Create an index for the business glossary. Create a second index for the business metadata. Use the OpenSearch Service dashboard to access the metadata.
❌ Sai vì: OpenSearch (trước là Elasticsearch) là search/analytics engine, không phải data catalog (thiếu governance cho data assets). Phải tự tạo 2 indexes (modeling phức tạp), ingest metadata thủ công, dashboard chỉ cơ bản (không chuyên business glossary). Overhead cao: provisioning domain, indexing pipeline, không tích hợp native với Glue/Lake Formation. Không phù hợp cho data catalog enterprise.
Kết luận 🎯: Amazon DataZone là lựa chọn tối ưu theo best practices AWS 2026, giúp governance dữ liệu nhanh chóng mà không cần custom code. Nếu cần lab thực hành, dùng AWS Free Tier với DataZone trial!
MERGE INTO accounts t USING monthly_accounts_update s
ON t.customer = s.customer -
WHEN MATCHED -
THEN DELETE -
What will happen when the data engineer runs the SQL command?
- A All customer records that exist in both the customer accounts table and the monthly_accounts_update table will be deleted from the accounts table.
- B Only customer records that are present in both tables will be retained in the customer accounts table.
- C The monthly_accounts_update table will be deleted.
- D No records will be deleted because the command syntax is not valid in AWS Glue.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào AWS Glue ETL job được sử dụng bởi một data engineer để xóa các bản ghi khách hàng lỗi thời từ bảng accounts chứa thông tin tài khoản khách hàng. Họ sử dụng lệnh SQL MERGE để loại bỏ những khách hàng tồn tại trong bảng monthly_accounts_update khỏi bảng accounts.
Lệnh SQL cụ thể:
MERGE INTO accounts t USING monthly_accounts_update s
ON t.customer = s.customer
WHEN MATCHED
THEN DELETE
- Ý nghĩa lệnh: Đây là lệnh MERGE (hợp nhất dữ liệu) theo chuẩn Spark SQL/Delta Lake, nơi:
- Target table:
accounts(bảng t). - Source table:
monthly_accounts_update(bảng s). - Điều kiện khớp:
t.customer = s.customer(khớp theo trườngcustomer). - Hành động: Khi MATCHED (khớp), thì DELETE (xóa bản ghi từ target table).
- Target table:
Lệnh này sẽ xóa tất cả bản ghi trong accounts mà có giá trị customer trùng với bất kỳ bản ghi nào trong monthly_accounts_update. Đây là hành vi chuẩn của MERGE ... WHEN MATCHED THEN DELETE trong Delta Lake, được hỗ trợ đầy đủ trong AWS Glue từ phiên bản 3.0+ (với Delta Lake tables). AWS Glue ETL jobs chạy trên Apache Spark, hỗ trợ các định dạng table ACID như Delta Lake, nơi MERGE đảm bảo transaction an toàn và atomic.
Lưu ý cập nhật 2026: Theo AWS Glue phiên bản mới nhất (Glue 4.0+), MERGE được tối ưu hóa cho Delta Lake 2.4+, Lake Formation governance, và tích hợp với S3 cho performance cao. Lệnh syntax này hợp lệ nếu accounts là Delta table (yêu cầu cho ACID ops như DELETE).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: All customer records that exist in both the customer accounts table and the monthly_accounts_update table will be deleted from the accounts table.
Lý do 🛠️:
- Lệnh MERGE chỉ ảnh hưởng đến target table (
accounts), xóa các bản ghi matched với source (monthly_accounts_update) dựa trên điều kiệnON t.customer = s.customer. - Không có clause
WHEN NOT MATCHED, nên chỉ xử lý matched rows bằng DELETE → Tất cả bản ghi trùngcustomersẽ bị xóa khỏiaccounts. - Hành vi này atomic và idempotent nhờ Delta Lake engine trong AWS Glue, tránh data loss hoặc inconsistency.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh, với lý do đúng/sai bằng tiếng Việt:
-
✅ All customer records that exist in both the customer accounts table and the monthly_accounts_update table will be deleted from the accounts table.
Đúng vì lệnh MERGE ... WHEN MATCHED THEN DELETE chính xác xóa các bản ghi trùng khớp từ target table (accounts), không ảnh hưởng source. Đây là hành vi chuẩn của Delta Lake MERGE trong AWS Glue. -
❌ Only customer records that are present in both tables will be retained in the customer accounts table.
Sai vì đây là diễn giải ngược lại hoàn toàn. Lệnh không giữ lại (retain) mà xóa (delete) các bản ghi matched. Các bản ghi chỉ tồn tại trongaccounts(không matched) sẽ được giữ nguyên. -
❌ The monthly_accounts_update table will be deleted.
Sai vì MERGE chỉ modify target table (accounts). Source table (monthly_accounts_update) chỉ dùng để matching, không bị thay đổi hay xóa. Đây là nguyên tắc cơ bản của MERGE statement. -
❌ No records will be deleted because the command syntax is not valid in AWS Glue.
Sai vì syntax hoàn toàn hợp lệ trong AWS Glue ETL job với Spark SQL và Delta Lake (từ Glue 3.0+). AWS Glue hỗ trợ đầy đủ MERGE cho Delta tables trên S3, với checkpointing và optimizations.
📘 Tài liệu tham khảo
- AWS Glue Developer Guide: MERGE INTO in AWS Glue (cập nhật 2024-2026, hỗ trợ Delta Lake 3.0+).
- Delta Lake Documentation: MERGE INTO Syntax (tích hợp AWS Glue).
- AWS re:Post & Blog: Các case study về ETL với MERGE in Glue for data cleansing (tìm kiếm "Glue Delta Lake MERGE DELETE").
- Exam Topic DOP-C02: Phần Data Engineering với Glue ETL & ACID tables.
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần demo code Glue job, hãy hỏi thêm.
A data engineer needs to set-up an extract, transform, and load (ETL) pipeline to upload the content of each file to Amazon Redshift.
Which solution will meet these requirements with the LEAST operational overhead?
- A Create an AWS Lambda function that connects to Amazon Redshift and runs a COPY command. Use Amazon EventBridge to invoke the Lambda function based on an Amazon S3 upload trigger.
- B Create an Amazon Data Firehose stream. Configure the stream to use an AWS Lambda function as a source to pull data from the S3 bucket. Set Amazon Redshift as the destination.
- C Use Amazon Redshift Spectrum to query the S3 bucket. Configure an AWS Glue Crawler for the S3 bucket to update metadata in an AWS Glue Data Catalog.
- D Creates an AWS Database Migration Service (AWS DMS) task. Specify an appropriate data schema to migrate. Specify the appropriate type of migration to use.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc thiết lập một pipeline ETL (Extract, Transform, Load) để chuyển dữ liệu marketing từ Amazon S3 (dữ liệu CSV, kích thước nhỏ 100-300 KB, ingest mỗi 40-60 phút) vào Amazon Redshift. Yêu cầu chính là giải pháp có ít overhead vận hành nhất (LEAST operational overhead), nghĩa là dễ quản lý, tự động hóa cao, ít can thiệp thủ công, chi phí thấp và phù hợp với tần suất dữ liệu không liên tục (batch nhỏ, định kỳ).
📌 Bối cảnh quan trọng:
- Dữ liệu từ vendor → S3 bucket (upload định kỳ).
- Cần upload nội dung từng file vào Redshift (không chỉ query, mà load thực sự).
- Với file nhỏ và tần suất thấp, ưu tiên giải pháp serverless, không cần cluster quản lý, tận dụng trigger tự động từ S3.
- Kiến thức cập nhật AWS 2026: Redshift hỗ trợ COPY từ S3 hiệu quả; Lambda scale tự động; EventBridge (trước là CloudWatch Events) linh hoạt cho S3 events.
✅ Đáp án đúng: Create an AWS Lambda function that connects to Amazon Redshift and runs a COPY command. Use Amazon EventBridge to invoke the Lambda function based on an Amazon S3 upload trigger.
Lý do lựa chọn:
- 🛠️ Giải pháp tối ưu với overhead thấp nhất: Lambda là serverless, tự scale, không cần quản lý server. S3 event trigger qua EventBridge (hoặc trực tiếp S3 notifications) kích hoạt ngay khi file mới upload → Lambda kết nối Redshift và chạy lệnh COPY (native command của Redshift để load nhanh từ S3, hỗ trợ CSV trực tiếp, parallel loading).
- 📈 Phù hợp dữ liệu: File nhỏ → COPY xử lý nhanh (giây), tần suất 40-60 phút → trigger event-based, không lãng phí tài nguyên.
- 🔒 An toàn & dễ quản lý: Sử dụng IAM roles cho Lambda access S3/Redshift; hỗ trợ transform nhẹ trong Lambda nếu cần ETL cơ bản.
- 💰 Overhead thấp: Không cần provision cluster, monitor stream; chi phí theo usage (rẻ cho batch nhỏ).
Tài liệu tham khảo:
- AWS Redshift Docs: COPY command (cập nhật 2025 hỗ trợ enhanced VPC).
- AWS Lambda + S3/EventBridge: S3 Event Notifications.
📋 Giải thích tất cả các phương án
-
✅ Create an AWS Lambda function that connects to Amazon Redshift and runs a COPY command. Use Amazon EventBridge to invoke the Lambda function based on an Amazon S3 upload trigger.
(Đã giải thích ở trên - Giải pháp lý tưởng, serverless, event-driven, tận dụng COPY native cho Redshift). -
❌ Create an Amazon Data Firehose stream. Configure the stream to use an AWS Lambda function as a source to pull data from the S3 bucket. Set Amazon Redshift as the destination.
Sai vì: Amazon Kinesis Data Firehose (nay là Amazon Data Firehose) là dịch vụ streaming push-based, không hỗ trợ Lambda làm source để pull từ S3 (Firehose source chính là direct PUT, Kinesis, CloudWatch Logs; không có cơ chế poll S3). Redshift làm destination chỉ hỗ trợ qua unload/COPY ngược, không trực tiếp cho ETL từ S3 batch. Overhead cao vì phải config buffer, transformation Lambda riêng, không phù hợp file nhỏ định kỳ (wasted resources).
Tài liệu: Firehose Sources/Destinations (không list Lambda pull). -
❌ Use Amazon Redshift Spectrum to query the S3 bucket. Configure an AWS Glue Crawler for the S3 bucket to update metadata in an AWS Glue Data Catalog.
Sai vì: Redshift Spectrum chỉ query external data từ S3 (không load/upload vào Redshift cluster), phù hợp analytics on S3 mà không di chuyển dữ liệu. Glue Crawler tạo metadata cho catalog (query federated), nhưng không thực hiện ETL load vào Redshift. Overhead cao: cần quản lý external tables, IAM, và không giải quyết yêu cầu "upload content" (dữ liệu vẫn ở S3).
Tài liệu: Redshift Spectrum (query-only). -
❌ Creates an AWS Database Migration Service (AWS DMS) task. Specify an appropriate data schema to migrate. Specify the appropriate type of migration to use.
Sai vì: AWS DMS dành cho migration database-to-database (e.g., RDS → Redshift), hỗ trợ S3 chỉ như intermediate storage (không trực tiếp từ S3 CSV làm source chính). Không phù hợp batch file nhỏ từ S3; cần schema predefined, endpoint config phức tạp → overhead cao (provision replication instance, monitor CDC/FULL load). Không event-driven.
Tài liệu: DMS Sources (S3 chỉ target, không source chính cho CSV batch).
Kết luận 🎯: Giải pháp Lambda + COPY + EventBridge là best practice serverless ETL cho Redshift từ S3 (theo AWS Well-Architected Framework 2025, pillar Operational Excellence). Nếu scale lớn hơn, có thể kết hợp Glue Jobs, nhưng ở đây batch nhỏ → Lambda thắng! 🚀
A data engineer needs a solution to update changes for up to 10,000 records in the base table every day.
Which solution will meet this requirement with the LOWEST runtime?
- A Develop an Apache Spark job in Amazon EMR to read the historical data and the new changes into two Spark DataFrames. Use the Spark update method to update the base table.
- B Develop an AWS Glue Python job to read the historical data and new changes into two Pandas DataFrames. Use the Pandas update method to update the base table.
- C Develop an AWS Glue Apache Spark job to read the historical data and new changes into two Spark DataFrames. Use the Spark update method to update the base table.
- D Develop an Amazon EMR job to read new changes into Apache Spark DataFrames. Use the Apache Hudi framework to create the base table in Amazon S3. Use the Spark update method to update the base table.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc xây dựng một dimension table (bảng chiều) trong Amazon S3 bucket, chứa dữ liệu lịch sử (historical data) với 10 triệu records và kích thước 1 TB. Yêu cầu chính là tìm giải pháp để cập nhật (update) thay đổi cho tối đa 10.000 records trong bảng cơ sở (base table) mỗi ngày, đồng thời đảm bảo thời gian chạy (runtime) thấp nhất (LOWEST runtime).
🔍 Thách thức cốt lõi:
- Dữ liệu lịch sử lớn (1 TB), nên tránh đọc toàn bộ dữ liệu mỗi lần update để giảm thời gian xử lý.
- Cần hỗ trợ incremental updates (cập nhật tăng dần) hiệu quả trên S3, nơi không hỗ trợ update trực tiếp như database truyền thống.
- Giải pháp phải tận dụng các dịch vụ AWS như EMR, Glue, Spark, và các framework tối ưu cho S3 như Apache Hudi (hỗ trợ ACID transactions, upserts trên lakehouse).
📘 Kiến thức AWS cập nhật đến 2026: Theo AWS Lake Formation và EMR phiên bản mới nhất (EMR 6.x+ và Hudi 0.14+), Apache Hudi là lựa chọn tối ưu cho MERGE/upsert operations trên S3 mà không cần đọc full dataset, nhờ cơ chế delta log và indexing (như Bloom filters hoặc HFile indexes).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng:
Develop an Amazon EMR job to read new changes into Apache Spark DataFrames. Use the Apache Hudi framework to create the base table in Amazon S3. Use the Spark update method to update the base table.
Lý do chi tiết 🛠️:
- Apache Hudi trên EMR cho phép tạo bảng với định dạng Copy-On-Write (COW) hoặc Merge-On-Read (MOR), hỗ trợ upsert (update + insert) chỉ bằng cách đọc new changes (10.000 records), không cần đọc toàn bộ 1 TB historical data.
- Hudi sử dụng delta files và indexes để merge nhanh chóng, giảm runtime đáng kể (chỉ O(changes) thay vì O(full dataset)).
- EMR tối ưu cho Spark jobs lớn, tích hợp Hudi native (từ EMR 6.5+), đảm bảo lowest runtime cho daily updates.
- Kết quả: Bảng S3 ACID-compliant, queryable qua Athena/Glue Catalog.
📋 Phân tích tất cả các phương án
-
Phương án 1:
Develop an Apache Spark job in Amazon EMR to read the historical data and the new changes into two Spark DataFrames. Use the Spark update method to update the base table.
❌ Sai: Phải đọc toàn bộ historical data (1 TB) + changes mỗi ngày → runtime cao do shuffle/join lớn trên Spark. Không tận dụng incremental processing, vi phạm yêu cầu LOWEST runtime. Spark Delta Lake tốt hơn nhưng không được đề cập. -
Phương án 2:
Develop an AWS Glue Python job to read the historical data and new changes into two Pandas DataFrames. Use the Pandas update method to update the base table.
❌ Sai: Pandas chỉ phù hợp dữ liệu nhỏ (không scale cho 1 TB → OOM errors). Glue Python shell giới hạn memory (max 10 GB/driver), không xử lý được 10M records. Runtime cực cao và không khả thi. -
Phương án 3:
Develop an AWS Glue Apache Spark job to read the historical data and new changes into two Spark DataFrames. Use the Spark update method to update the base table.
❌ Sai: Tương tự phương án 1, đọc full 1 TB mỗi ngày trên Glue Spark → runtime dài (Glue chậm hơn EMR cho jobs lớn). Không hỗ trợ efficient upserts trên S3 mà không có framework như Hudi/Delta. -
Phương án 4 (Đúng):
Develop an Amazon EMR job to read new changes into Apache Spark DataFrames. Use the Apache Hudi framework to create the base table in Amazon S3. Use the Spark update method to update the base table.
✅ Đúng: Chỉ đọc new changes (10K records), Hudi tự merge với historical data qua upsert command (hoodie.table.name,hoodie.datasource.write.operation=upsert). Runtime thấp nhất nhờ incremental compaction và S3 optimizations (EMRFS).
📚 Tài liệu tham khảo
- AWS Documentation: Apache Hudi on Amazon EMR (cập nhật 2025: Hudi 1.0+ hỗ trợ Global Bloom Indexes).
- AWS Blog: Transactional Data Lakes with Hudi on S3 (2024-2026).
- Hudi Docs: Upserts and Incremental Processing – Xác nhận low-latency cho small deltas.
- Exam Topic: DOP-C02 (DevOps Pro) – Phần Data Lakes & Analytics.
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm ví dụ code Spark-Hudi, hãy hỏi nhé!
The data engineer needs to identify the source of the error and provide a solution.
Which combinations of steps will meet this requirement MOST cost-effectively? (Choose two.)
- A Scale out the workers vertically to address data skewness.
- B Use the Spark UI and AWS Glue metrics to monitor data skew in the Spark executors.
- C Scale out the number of workers horizontally to address data skewness.
- D Enable the --write-shuffie-files-to-s3 job parameter. Use the salting technique.
- E Use error logs in Amazon CloudWatch to monitor data skew.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào một tình huống thực tế trong AWS Glue khi phát triển job ETL dựa trên Apache Spark: Một data engineer chạy job transform dataset nhưng gặp lỗi "No space left on device". Lỗi này thường xảy ra do data skew (dữ liệu lệch), khiến một số executor Spark phải xử lý lượng dữ liệu khổng lồ trong giai đoạn shuffle, dẫn đến tràn bộ nhớ đệm và hết dung lượng đĩa cục bộ (ephemeral storage) trên worker node.
Yêu cầu là xác định nguồn gốc lỗi (data skew) và cung cấp giải pháp MOST cost-effectively (tiết kiệm chi phí nhất), chọn TWO combinations of steps. Giải pháp cần ưu tiên các bước monitor và tối ưu hóa Spark shuffle mà không cần scale tài nguyên (vốn tốn kém hơn), phù hợp với best practices AWS Glue Spark jobs (cập nhật đến 2024-2026, hỗ trợ Spark 3.5+ và G.4X workers).
✅ Đáp án đúng (Chọn TWO)
Hai bước đúng nhất, kết hợp để monitor skew và giảm disk pressure một cách tiết kiệm:
- Use the Spark UI and AWS Glue metrics to monitor data skew in the Spark executors.
(Bước đầu tiên: Giám sát skew qua Spark UI và metrics Glue để xác định nguồn lỗi chính xác, không tốn thêm chi phí.) - Enable the --write-shuffle-files-to-s3 job parameter. Use the salting technique.
(Giải pháp tối ưu: Viết shuffle files ra S3 thay vì đĩa local, kết hợp salting để phân tán key skew, giảm tải disk và cost-effective.)
Lý do lựa chọn: Lỗi "No space left on device" chủ yếu từ shuffle disk overflow do skew. Monitor bằng Spark UI (truy cập qua Glue console/CloudWatch) và metrics (như executor disk usage) là bước đầu miễn phí. Sau đó, --write-shuffle-files-to-s3 (param Glue Spark từ 2021+, cập nhật Spark 3.x) offload shuffle sang S3 (rẻ hơn EBS/local disk), kết hợp salting (thêm random salt vào partition key) để cân bằng dữ liệu. Scaling workers tốn kém (tăng DPUs), không phải "MOST cost-effectively".
📋 Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc:
-
Scale out the workers vertically to address data skewness.
❌ Sai: Scale vertically (tăng DPUs/worker size, ví dụ từ G.1X lên G.2X) tăng memory/disk per node, nhưng không giải quyết gốc rễ skew (vẫn có executor bị overload). Đây là giải pháp đắt đỏ (tăng bill theo DPU-hour), không cost-effective. Thay vào đó, ưu tiên salting trước. -
Use the Spark UI and AWS Glue metrics to monitor data skew in the Spark executors.
✅ Đúng: Spark UI (truy cập qua Glue job runs > Spark history) hiển thị task duration, shuffle read/write per executor; Glue metrics (CloudWatch: glue.driver.ExecutorDiskUsageInBytes, glue.driver.NumSkewedPartitions) xác định skew chính xác. Miễn phí, bước đầu tiên bắt buộc để debug mà không scale. -
Scale out the number of workers horizontally to address data skewness.
❌ Sai: Scale horizontally (tăng số workers/DPUs) phân tán workload nhưng không fix skew (skew key vẫn tập trung vào ít partitions). Tăng chi phí theo parallel DPUs, không phải giải pháp gốc. AWS khuyến cáo dùng salting thay vì scale mù quáng. -
Enable the --write-shuffle-files-to-s3 job parameter. Use the salting technique.
✅ Đúng:--write-shuffle-files-to-s3(job param Glue Spark, Spark confspark.sql.shuffle.partitionsoptimize) viết shuffle temp files ra S3 (rẻ ~$0.023/GB/tháng), tránh hết local disk (20-500GB tùy worker). Salting (thêm hash/random vào key:df.withColumn("salt", rand() % 10).repartition(...)) cân bằng partitions. Kết hợp hoàn hảo, cost-effective. -
Use error logs in Amazon CloudWatch to monitor data skew.
❌ Sai: CloudWatch logs chỉ ghi error tổng quát ("No space left on device"), không chi tiết skew như partition stats hay executor metrics. Spark UI/Glue metrics mới cung cấp dữ liệu skew cụ thể (task skew ratio >10x). Logs hữu ích nhưng không đủ cho monitoring skew sâu.
📘 Tài liệu tham khảo (Cập nhật AWS 2024-2026)
- AWS Glue Developer Guide: Troubleshoot Spark jobs & Data skew handling.
- Spark on Glue best practices: AWS re:Post - No space left on device (khuyến nghị Spark UI + S3 shuffle).
- Glue Metrics/Params: CloudWatch metrics for Glue & Spark conf write-shuffle-files-to-s3.
- Salting technique: AWS Big Data Blog Handle skew in Glue.
🛠️ Khuyến nghị thực hành: Chạy job với --enable-metrics + Spark UI, test salting trên small dataset trước khi production!
A user made a change to the security group that prevents the AWS Glue jobs from connecting to the RDS instance. After the change, the security group contains a single rule that allows inbound SSH traffic from a specific IP address.
The company must resolve the connectivity issue.
Which solution will meet this requirement?
- A Add an inbound rule that allows all TCP traffic on all TCP ports. Set the security group as the source.
- B Add an inbound rule that allows all TCP traffic on all UDP ports. Set the private IP address of the RDS instance as the source.
- C Add an inbound rule that allows all TCP traffic on all TCP ports. Set the DNS name of the RDS instance as the source.
- D Replace the source of the existing SSH rule with the private IP address of the RDS instance. Create an outbound rule with the same source, destination, and protocol as the inbound SSH rule.
Xem giải thích
🧩 Giải thích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh một data pipeline sử dụng Amazon RDS instance (cơ sở dữ liệu quan hệ), AWS Glue jobs (công cụ ETL serverless), và Amazon S3 bucket. Các thành phần chính là RDS và Glue jobs nằm trong private subnet của một VPC, và chúng chia sẻ cùng một Security Group (SG).
🔍 Vấn đề xảy ra: Một user đã thay đổi SG, dẫn đến Glue jobs không kết nối được với RDS. Hiện tại, SG chỉ còn một rule inbound duy nhất: cho phép SSH traffic (port 22, TCP) từ một IP address cụ thể. Điều này chặn traffic từ Glue jobs (cùng SG) đến RDS, vì Glue cần kết nối qua TCP ports cụ thể của RDS (ví dụ: 3306 cho MySQL, 5432 cho PostgreSQL – theo tài liệu AWS Glue ETL connectivity cập nhật 2024-2026).
🎯 Yêu cầu: Giải quyết vấn đề kết nối một cách an toàn và hiệu quả, tuân thủ nguyên tắc least privilege của AWS Security Groups (stateless, chỉ inbound rules kiểm soát traffic vào instance).
📘 Kiến thức nền tảng (cập nhật AWS 2026):
- Security Groups hoạt động như virtual firewall, chỉ cần inbound rules để cho phép traffic từ source (có thể là CIDR, IP, hoặc chính SG khác/SG self-reference).
- Vì RDS và Glue cùng SG, traffic nội bộ được cho phép bằng cách set source = chính SG đó (self-referencing).
- AWS Glue jobs chạy trong managed environment (VPC endpoints hoặc ENI), cần SG inbound trên RDS cho TCP từ SG của Glue.
- Nguồn tham khảo:
- AWS VPC Security Groups (refers self-traffic).
- AWS Glue RDS Connectivity (TCP ports required).
- RDS Security Best Practices (2026 updates: enhanced VPC integration).
✅ Đáp án đúng: Lựa chọn đầu tiên
Add an inbound rule that allows all TCP traffic on all TCP ports. Set the security group as the source.
Lý do chọn 🛠️:
- Đây là giải pháp chuẩn và an toàn nhất vì RDS và Glue cùng một SG, nên thêm inbound rule với source = chính SG (self-reference) sẽ cho phép tất cả traffic TCP nội bộ từ Glue đến RDS mà không cần chỉ định IP cụ thể (Glue dùng dynamic ENI).
- "All TCP ports" bao quát ports RDS cần (như 3306), phù hợp cho pipeline. AWS khuyến nghị self-SG cho intra-group traffic (không cần outbound vì SG implicit allow outbound).
- Giải quyết ngay lập tức, không ảnh hưởng SSH rule hiện có. Theo best practice 2026, tránh mở "0.0.0.0/0" mà dùng self-SG.
📋 Phân tích tất cả các phương án (đúng/sai)
-
✅ Đúng: Add an inbound rule that allows all TCP traffic on all TCP ports. Set the security group as the source.
🧩 Giải thích: Như trên, self-referencing SG cho phép traffic TCP từ Glue (cùng SG) vào RDS một cách an toàn, không lộ ra ngoài VPC. Đây là cách AWS thiết kế cho private subnets (xem VPC docs). -
❌ Sai: Add an inbound rule that allows all TCP traffic on all UDP ports. Set the private IP address of the RDS instance as the source.
🧩 Giải thích: UDP ports không phù hợp vì RDS và Glue dùng TCP (không phải UDP). Source là private IP của RDS sai logic (inbound rule trên RDS SG không thể dùng IP của chính RDS làm source – phải là source từ Glue). Private IP động, không scalable. -
❌ Sai: Add an inbound rule that allows all TCP traffic on all TCP ports. Set the DNS name of the RDS instance as the source.
🧩 Giải thích: DNS name (endpoint RDS) không được hỗ trợ làm source trong SG rules (AWS chỉ chấp nhận CIDR/IP/SG ID/Prefix List). Dùng DNS sẽ bị reject khi apply rule. -
❌ Sai: Replace the source of the existing SSH rule with the private IP address of the RDS instance. Create an outbound rule with the same source, destination, and protocol as the inbound SSH rule.
🧩 Giải thích: Thay source SSH rule bằng private IP RDS vô nghĩa (SSH là inbound từ bastion IP, không liên quan Glue). SG không cần outbound rule (implicit allow all outbound). Copy rule SSH (port 22) không mở TCP ports RDS cần, và private IP RDS không phải source từ Glue.
🚀 Khuyến nghị bổ sung
- Sau fix, monitor bằng VPC Flow Logs và CloudWatch để verify traffic.
- Tối ưu: Thay "all TCP ports" bằng specific ports (e.g., 3306) cho least privilege.
- Test: Sử dụng EC2 Test Traffic từ Glue subnet để confirm.
Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 💪
A data engineer needs to add a data quality check for columns that contain null values and for referential integrity at a stage before the data is added to storage.
Which solution will meet these requirements with the LEAST operational overhead?
- A Use Amazon SageMaker Data Wrangler to create a Data Quality and Insights report.
- B Use AWS Glue ETL jobs to perform a data quality evaluation transform on the data. Use an IsComplete rule on the requested columns. Use a ReferentialItegrity rule for each join.
- C Use AWS Glue ETL jobs to perform a SQL transform on the data to determine whether requested column contain null values. Use a second SQL transform to check referential integrity.
- D Use Amazon SageMaker Data Wrangler and a custom Python transform to create custom rules to check for null values and referential integrity.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một công ty đang xây dựng pipeline dữ liệu mới để xử lý dữ liệu phục vụ báo cáo business intelligence (BI). Người dùng phát hiện dữ liệu bị thiếu trong báo cáo, vì vậy data engineer cần thêm kiểm tra chất lượng dữ liệu (data quality check) cụ thể cho:
- Cột chứa giá trị null (null values).
- Referential integrity (tính toàn vẹn tham chiếu, tức kiểm tra mối quan hệ giữa các bảng qua khóa ngoại).
Kiểm tra này phải diễn ra trước khi dữ liệu được lưu vào storage, và giải pháp cần có LEAST operational overhead (ít chi phí vận hành nhất, nghĩa là dễ triển khai, bảo trì, không cần code phức tạp hay quản lý thủ công nhiều).
Mục tiêu chính: Tìm giải pháp tích hợp sẵn, tự động hóa cao trong AWS để kiểm tra null và referential integrity mà không tốn nhiều công sức.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use AWS Glue ETL jobs to perform a data quality evaluation transform on the data. Use an IsComplete rule on the requested columns. Use a ReferentialItegrity rule for each join.
Lý do chọn đáp án này 🛠️:
- AWS Glue (phiên bản mới nhất 2026) hỗ trợ Data Quality transforms built-in trong ETL jobs, cho phép định nghĩa các rules sẵn có như:
- IsComplete rule: Kiểm tra tỷ lệ giá trị null/missing trong cột (ví dụ:
IsComplete "column_name" > 99%). - ReferentialIntegrity rule: Kiểm tra tính toàn vẹn tham chiếu giữa các bảng/join (ví dụ: kiểm tra khóa ngoại khớp với khóa chính).
- IsComplete rule: Kiểm tra tỷ lệ giá trị null/missing trong cột (ví dụ:
- Đây là giải pháp least operational overhead vì:
- Không cần viết code SQL/Python custom.
- Tích hợp trực tiếp vào pipeline ETL, chạy serverless, tự động scale.
- Hỗ trợ evaluate và alert nếu rule fail, dễ monitor qua AWS Glue console hoặc CloudWatch.
- Phù hợp với pipeline dữ liệu lớn, xử lý trước khi lưu vào S3/Redshift/Data Lake.
📘 Tài liệu tham khảo:
- AWS Glue Data Quality: docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-data-quality.html (cập nhật rules như IsComplete, ReferentialIntegrity từ 2023+).
- AWS Well-Architected Framework - Data Analytics Lens (2026 edition).
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm giải thích bằng tiếng Việt.
-
Use Amazon SageMaker Data Wrangler to create a Data Quality and Insights report.
❌ Sai: SageMaker Data Wrangler chủ yếu dùng cho exploratory data analysis (EDA) và visualization, tạo báo cáo insights (như missing values) nhưng không tích hợp trực tiếp vào pipeline ETL tự động. Nó yêu cầu manual flow và export code, tăng operational overhead (phải quản lý notebook, custom export). Không hỗ trợ referential integrity rules built-in, chỉ cơ bản null check qua statistics. -
Use AWS Glue ETL jobs to perform a data quality evaluation transform on the data. Use an IsComplete rule on the requested columns. Use a ReferentialItegrity rule for each join.
✅ Đúng: Như đã giải thích ở trên. Đây là giải pháp tối ưu nhất với rules built-in (IsComplete cho null, ReferentialIntegrity cho join), chạy trong Glue ETL job serverless, zero custom code, dễ scale cho big data pipeline trước storage. -
Use AWS Glue ETL jobs to perform a SQL transform on the data to determine whether requested column contain null values. Use a second SQL transform to check referential integrity.
❌ Sai: Dù dùng Glue ETL, nhưng phải viết SQL custom (ví dụ:COUNT(*) - COUNT(column)cho null, JOIN phức tạp cho referential integrity). Điều này tăng operational overhead cao: code thủ công, debug khó, bảo trì lâu dài, không tận dụng Data Quality transforms built-in (ra đời từ 2023 để thay thế cách này). -
Use Amazon SageMaker Data Wrangler and a custom Python transform to create custom rules to check for null values and referential integrity.
❌ Sai: Kết hợp Wrangler với Python custom (pandas/numpy cho null, merge cho integrity) tạo overhead lớn: Phải code từ đầu, quản lý notebook/export, không serverless như Glue. Wrangler phù hợp prototype nhỏ, không scale cho production pipeline, và custom rules khó maintain so với Glue's built-in.
Kết luận 🚀: Giải pháp AWS Glue Data Quality là best practice cho data pipeline năm 2026, đảm bảo chất lượng dữ liệu tự động với chi phí vận hành thấp nhất! Nếu cần demo code rule, hỏi thêm nhé. 😊
Which solution will meet these requirements MOST cost-effectively?
- A Use AWS Glue ETL to extract the data from the S3 buckets and perform the transformations. Use AWS Glue Data Quality to enforce suggested quality rules. Load the data and the quality check results into an Amazon RDS for MySQL instance.
- B Use AWS Glue Studio to extract the data from the S3 buckets. Use AWS Glue DataBrew to perform the transformations and quality checks. Load the processed data into an Amazon RDS for MySQL instance. Load the quality check results into a new S3 bucket.
- C Use AWS Glue ETL to extract the data from the S3 buckets and perform the transformations. Use AWS Glue DataBrew to perform quality checks. Load the processed data and the quality check results into a new S3 bucket.
- D Use AWS Glue Studio to extract the data from the S3 buckets. Use AWS Glue DataBrew to perform the transformations and quality checks. Load the processed data and quality check results into an Amazon RDS for MySQL instance.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một công ty đang xây dựng data pipeline trên AWS để xử lý dữ liệu khách hàng. Quy trình bao gồm:
- Extract (trích xuất) dữ liệu từ các Amazon S3 buckets.
- Thực hiện quality checks (kiểm tra chất lượng dữ liệu).
- Transform (chuyển đổi dữ liệu).
- Store dữ liệu đã xử lý vào relational database (cơ sở dữ liệu quan hệ), cụ thể là để hỗ trợ các truy vấn (queries) trong tương lai.
Yêu cầu chính: Tìm giải pháp MOST cost-effectively (tiết kiệm chi phí nhất), nghĩa là ưu tiên các dịch vụ serverless, tích hợp cao, tránh overhead không cần thiết, và phù hợp với quy mô production pipeline.
📘 Nguồn tham khảo: AWS Glue Documentation (cập nhật 2024-2026): AWS Glue ETL Jobs, AWS Glue Data Quality, và AWS Well-Architected Framework - Cost Optimization Pillar.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng là phương án đầu tiên:
Use AWS Glue ETL to extract the data from the S3 buckets and perform the transformations. Use AWS Glue Data Quality to enforce suggested quality rules. Load the data and the quality check results into an Amazon RDS for MySQL instance.
Lý do chọn đáp án này là MOST cost-effective:
🛠️ AWS Glue ETL là dịch vụ serverless ETL cốt lõi, hỗ trợ extract từ S3 và transform dữ liệu một cách tự động, scalable, chỉ tính phí theo DPU-hour (Data Processing Units), rất rẻ cho pipeline lớn.
✅ AWS Glue Data Quality (tính năng tích hợp trực tiếp vào Glue ETL jobs từ 2021, cập nhật liên tục đến 2026) cho phép thực hiện quality checks ngay trong job ETL mà không cần dịch vụ riêng biệt, tránh chi phí overhead. Nó hỗ trợ "suggested quality rules" (quy tắc chất lượng gợi ý tự động).
📈 Cả dữ liệu đã xử lý và kết quả quality checks đều load trực tiếp vào Amazon RDS for MySQL (relational DB phù hợp cho queries), tối ưu chi phí lưu trữ và truy vấn mà không cần S3 trung gian.
💰 Tổng chi phí thấp nhất: Tích hợp end-to-end trong Glue, serverless, không interactive tools → lý tưởng cho production pipeline tự động.
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng phương án (giữ nguyên văn bản gốc tiếng Anh). Mỗi phương án được đánh giá đúng/sai dựa trên tính cost-effective, tích hợp, scalability và phù hợp yêu cầu (relational DB cho queries).
-
Phương án 1 (ĐÚNG ✅):
Use AWS Glue ETL to extract the data from the S3 buckets and perform the transformations. Use AWS Glue Data Quality to enforce suggested quality rules. Load the data and the quality check results into an Amazon RDS for MySQL instance.
Giải thích: Như đã phân tích ở trên, đây là giải pháp tối ưu nhất với tích hợp liền mạch, serverless thuần túy, không overhead. Phù hợp hoàn hảo cho pipeline tự động và chi phí thấp nhất.
📘 Nguồn: AWS Glue Data Quality Rules. -
Phương án 2 (SAI ❌):
Use AWS Glue Studio to extract the data from the S3 buckets. Use AWS Glue DataBrew to perform the transformations and quality checks. Load the processed data into an Amazon RDS for MySQL instance. Load the quality check results into a new S3 bucket.
Giải thích: AWS Glue Studio chỉ là giao diện visual cho ETL (không phải ETL engine chính), kết hợp AWS Glue DataBrew (tool interactive cho data prep, phù hợp analyst hơn production pipeline) tạo overhead cao hơn vì chạy separate jobs/services. DataBrew tính phí theo giờ interactive (đắt hơn Glue ETL cho scale lớn). Load quality results riêng vào S3 mới → tăng chi phí lưu trữ và quản lý, không cost-effective. Không tận dụng tích hợp Data Quality native. -
Phương án 3 (SAI ❌):
Use AWS Glue ETL to extract the data from the S3 buckets and perform the transformations. Use AWS Glue DataBrew to perform quality checks. Load the processed data and the quality check results into a new S3 bucket.
Giải thích: Dù dùng Glue ETL tốt cho extract/transform, nhưng tách DataBrew riêng cho quality checks → thiếu tích hợp, tăng chi phí (DataBrew không serverless như Data Quality). Load tất cả vào S3 mới (không phải relational DB) vi phạm yêu cầu "stores the processed data in a relational database" và kém hiệu quả cho future queries (S3 là object storage, không tối ưu SQL queries). Tổng thể kém cost-effective. -
Phương án 4 (SAI ❌):
Use AWS Glue Studio to extract the data from the S3 buckets. Use AWS Glue DataBrew to perform the transformations and quality checks. Load the processed data and quality check results into an Amazon RDS for MySQL instance.
Giải thích: Glue Studio + DataBrew đều là visual/interactive tools, không scalable và đắt hơn cho automated pipeline lớn (DataBrew tính phí cao hơn Glue ETL/Data Quality). Dù load vào RDS đúng, nhưng thiếu tích hợp native → overhead cao, không phải lựa chọn cost-effective nhất. AWS khuyến nghị Glue ETL + Data Quality cho production ETL pipelines đến 2026.
📘 Nguồn: AWS Glue vs. DataBrew Comparison.
Kết luận 🏆: Giải pháp đúng tận dụng tích hợp sâu trong AWS Glue ecosystem, đảm bảo serverless, scalable và chi phí thấp nhất cho data pipeline production! Nếu cần code sample Glue job, hãy cho tôi biết nhé! 🚀