Ngân hàng đề — AWS Certified Data Engineer Associate
Tìm thấy 867 câu.
Which solution will meet these requirements in the MOST operationally efficient way?
-
A
Create a small Amazon EC2 instance that polls the S3 bucket for new files. Run transformation code on a schedule to generate the output. Use operating system commands to send email messages.
- B Run an Amazon Elastic Container Service (Amazon ECS) task to poll the S3 bucket for new files. Run transformation code on a schedule to generate the output. Use operating system commands to send email messages.
- C Create an AWS Lambda function to transform the data. Use Amazon S3 Event Notifications to invoke the Lambda function when a new object is created. Publish the output to an Amazon Simple Notification Service (Amazon SNS) topic. Subscribe the data engineer’s email account to the topic.
- D Deploy an Amazon EMR cluster. Use EMR File System (EMRFS) to access the files in the S3 bucket. Run transformation code on a schedule to generate the output to a second S3 bucket. Create an Amazon Simple Notification Service (Amazon SNS) topic. Configure Amazon S3 Event Notifications to notify the topic when a new object is created.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi yêu cầu tìm giải pháp hiệu quả vận hành nhất (MOST operationally efficient) để một data engineer chạy job biến đổi dữ liệu (data transformation) tự động mỗi khi người dùng thêm file mới vào S3 bucket. Các yêu cầu cụ thể:
- Job chạy dưới 1 phút.
- Gửi output qua email cho data engineer.
- Tần suất: 1 file/giờ (thấp, không cần xử lý cao tải).
🛠️ Mục tiêu chính: Giải pháp phải serverless, event-driven (không polling), chi phí thấp, không quản lý hạ tầng, và tích hợp email tự động. AWS ưu tiên các dịch vụ như Lambda, S3 Events, SNS để đạt hiệu quả vận hành cao nhất (theo nguyên tắc Well-Architected Framework: Operational Excellence pillar).
📘 Dẫn nguồn tham khảo:
- AWS Well-Architected Framework (2024+): docs.aws.amazon.com/wellarchitected/latest/framework/welcome.html
- Amazon S3 Event Notifications: docs.aws.amazon.com/AmazonS3/latest/userguide/NotificationHowTo.html
- AWS Lambda với S3: docs.aws.amazon.com/lambda/latest/dg/with-s3.html
- Amazon SNS Email Subscriptions: docs.aws.amazon.com/sns/latest/dg/sns-email-subscriptions.html
✅ Đáp án đúng và lý do lựa chọn
Create an AWS Lambda function to transform the data. Use Amazon S3 Event Notifications to invoke the Lambda function when a new object is created. Publish the output to an Amazon Simple Notification Service (Amazon SNS) topic. Subscribe the data engineer’s email account to the topic.
Lý do chọn 🏆:
- Đây là giải pháp serverless hoàn toàn, event-driven (S3 Event Notifications kích hoạt Lambda ngay lập tức khi có file mới, không polling lãng phí).
- Lambda chạy <1 phút lý tưởng (hỗ trợ tối đa 15 phút), scale tự động, chi phí pay-per-use (rẻ cho 1 job/giờ).
- SNS publish output và subscribe email tự động, không cần code OS commands.
- Hiệu quả vận hành cao nhất: Không quản lý server/container/cluster, zero downtime, monitoring tự động qua CloudWatch. Phù hợp kiến trúc AWS hiện đại đến 2026 (Lambda SnapStart/ARM cho tốc độ cao hơn).
🔍 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, với giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên tính operationally efficient (chi phí, quản lý, scalability, event-driven).
-
❌ Create a small Amazon EC2 instance that polls the S3 bucket for new files. Run transformation code on a schedule to generate the output. Use operating system commands to send email messages.
Sai vì: Polling S3 liên tục (dù tần suất thấp 1/giờ) gây lãng phí CPU/idle time, EC2 luôn chạy (chi phí ~$5-10/tháng dù job ngắn). Scheduled code không trigger realtime. OS commands gửi email không đáng tin cậy (spam filter, no retry). Phải quản lý patching/security → không efficient. -
❌ Run an Amazon Elastic Container Service (Amazon ECS) task to poll the S3 bucket for new files. Run transformation code on a schedule to generate the output. Use operating system commands to send email messages.
Sai vì: ECS vẫn cần quản lý container/cluster (EC2/Fargate), polling lãng phí tương tự EC2. Scheduled task không event-driven. OS email không native AWS, phức tạp scale/retry. Dù Fargate serverless hơn EC2 nhưng vẫn overhead cao cho job <1 phút/giờ → kém efficient hơn Lambda. -
✅ Create an AWS Lambda function to transform the data. Use Amazon S3 Event Notifications to invoke the Lambda function when a new object is created. Publish the output to an Amazon Simple Notification Service (Amazon SNS) topic. Subscribe the data engineer’s email account to the topic.
Đúng vì: Như đã giải thích ở trên – serverless, event-driven, native integration, chi phí ~$0.000001/job (rẻ nhất), zero management. Hỗ trợ ARM/Graviton3 (2024+) cho hiệu suất cao. -
❌ Deploy an Amazon EMR cluster. Use EMR File System (EMRFS) to access the files in the S3 bucket. Run transformation code on a schedule to generate the output to a second S3 bucket. Create an Amazon Simple Notification Service (Amazon SNS) topic. Configure Amazon S3 Event Notifications to notify the topic when a new object is created.
Sai vì: EMR dành cho big data/Hadoop/Spark (overkill cho job <1 phút). Cluster luôn-on hoặc auto-scale chi phí cao (~$0.1-1/giờ), scheduled code + S3 event chỉ notify (không transform trực tiếp). Output sang S3 thứ 2 rồi SNS phức tạp thừa, phải quản lý cluster lifecycle → không efficient cho workload nhỏ.
🛠️ Kết luận: Giải pháp Lambda + S3 Events + SNS là best practice AWS cho event-driven micro-jobs, giúp đạt Operational Excellence tối ưu! 🚀
A data engineer notices that the workflow is generating errors as a result of how customer postal codes are stored in the data lake. Some postal codes include unnecessary numbers or invalid characters.
The data engineer needs a solution to address the errors and correct the postal codes in the data lake.
- A Create a schema definition for PySpark that matches the format the processing workflow requires for postal codes. Pass the schema to the DynamicFrame during processing.
- B Use AWS Glue workflow properties to allow job state sharing. Configure the AWS Glue jobs to read values from the postal code column by using the properties from a previously successful run of the jobs.
- C Configure the column.push_down_predicate setting and the catalogPartitionPredicate settings for the postal code column in the DynamicFrame.
- D Set the DynamicFrame additional_options parameter ‘useS3ListImplementation’ to True.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh một data lake được quản lý bởi Amazon S3 và AWS Glue Data Catalog, lưu trữ thông tin liên lạc khách hàng. Công ty sử dụng PySpark kết hợp AWS Glue jobs với DynamicFrame để xử lý dữ liệu qua một workflow.
🚨 Vấn đề chính: Workflow gặp lỗi do cột postal codes (mã bưu điện) chứa dữ liệu không chuẩn – có số thừa hoặc ký tự invalid (không hợp lệ).
🎯 Yêu cầu giải pháp: Data engineer cần cách sửa lỗi và chuẩn hóa postal codes ngay trong data lake, đảm bảo workflow chạy mượt mà mà không thay đổi dữ liệu gốc một cách thủ công.
Đây là tình huống phổ biến trong ETL pipelines trên AWS Glue (phiên bản mới nhất 2024-2026 hỗ trợ DynamicFrame với schema enforcement tốt hơn), nơi dữ liệu "dirty" cần được cast/enforce qua schema để tránh lỗi Spark.
✅ Đáp án đúng
Create a schema definition for PySpark that matches the format the processing workflow requires for postal codes. Pass the schema to the DynamicFrame during processing.
Lý do chọn đáp án này 🛠️:
- Trong AWS Glue PySpark, DynamicFrame (dựa trên Apache Spark) cho phép define schema tùy chỉnh để ép dữ liệu khớp format mong muốn. Bằng cách tạo schema PySpark (ví dụ: định nghĩa cột postal_code là string với pattern regex hoặc length cụ thể), rồi pass trực tiếp vào DynamicFrame (qua
DynamicFrame.fromDFhoặcapply_mapping), bạn có thể chuẩn hóa dữ liệu tự động – loại bỏ ký tự invalid, trim số thừa ngay lúc đọc/processing. - Giải pháp này không thay đổi dữ liệu gốc trên S3, mà chỉ enforce tại runtime, tránh lỗi schema mismatch. Đây là best practice theo AWS Glue 4.0+ (hỗ trợ schema evolution và casting tốt hơn).
📘 Tài liệu tham khảo: - AWS Glue Developer Guide: DynamicFrame Schema Handling (cập nhật 2024).
- PySpark StructType docs: Schema Enforcement in Glue.
📋 Giải thích chi tiết từng phương án
Dưới đây là phân tích tất cả 4 phương án, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm lý do cụ thể dựa trên tính năng AWS Glue mới nhất.
-
Create a schema definition for PySpark that matches the format the processing workflow requires for postal codes. Pass the schema to the DynamicFrame during processing.
✅ Đúng 🏆: Như đã giải thích ở trên, đây là cách trực tiếp và hiệu quả nhất để enforce schema lên DynamicFrame, tự động clean/sửa postal codes (ví dụ: cast sang string clean hoặc apply regex transform). Phù hợp hoàn hảo với PySpark jobs trên Glue, tránh lỗi mà không cần rewrite toàn bộ workflow. -
Use AWS Glue workflow properties to allow job state sharing. Configure the AWS Glue jobs to read values from the postal code column by using the properties from a previously successful run of the jobs.
❌ Sai 🚫: Job state sharing (qua Glue Workflow properties) chỉ dùng để chia sẻ trạng thái giữa các job (như bookmark hoặc metrics), giúp resume workflow sau failure. Nó không clean dữ liệu hay đọc lại postal codes từ run cũ – sẽ không sửa được invalid characters, thậm chí gây loop lỗi nếu run trước fail. Không liên quan đến data quality. -
Configure the column.push_down_predicate setting and the catalogPartitionPredicate settings for the postal code column in the DynamicFrame.
❌ Sai 🚫: push_down_predicate và catalogPartitionPredicate là các tùy chọn filtering partitions (partition pruning) trên Glue Data Catalog/S3, dùng để lọc dữ liệu trước khi load dựa trên predicate (ví dụ: postal_code = 'valid'). Chúng không sửa/clean dữ liệu – chỉ skip partitions invalid, dẫn đến mất data thay vì fix postal codes. -
Set the DynamicFrame additional_options parameter ‘useS3ListImplementation’ to True.
❌ Sai 🚫: useS3ListImplementation=True là tùy chọn optimize listing objects trên S3 (dùng S3 List API thay vì Spark's file listing), giúp tăng tốc độ scan large buckets trong Glue jobs (cải thiện từ Glue 3.0+). Hoàn toàn không liên quan đến việc clean postal codes – chỉ cải thiện performance I/O, không xử lý data quality.
🛠️ Khuyến nghị thực hiện
- Code mẫu PySpark:
from pyspark.sql.types import StructType, StructField, StringType schema = StructType([StructField("postal_code", StringType(), True)]) df = spark.read.schema(schema).parquet("s3://path/") dynamic_frame = DynamicFrame.fromDF(df, glueContext, "dyf") - Test trên Glue Studio hoặc SageMaker để verify trước production.
🎉 Giải pháp này đảm bảo data lake sạch, scalable cho volume lớn!
Which solution will meet this requirement?
- A Create an Amazon Simple Notification Service (Amazon SNS) FIFO topic. Subscribe the team’s email account to the SNS topic. Create an AWS Lambda function that initiates when the AWS Glue job state changes to FAILED. Set the SNS topic as the target.
- B Create an Amazon Simple Notification Service (Amazon SNS) standard topic. Subscribe the team’s email account to the SNS topic. Create an Amazon EventBridge rule that triggers when the AWS Glue job state changes to FAILED. Set the SNS topic as the target.
- C Create an Amazon Simple Queue Service (Amazon SQS) FIFO queue. Subscribe the team’s email account to the SQS queue. Create an AWS Config rule that triggers when the AWS Glue job state changes to FAILED. Set the SQS queue as the target.
- D Create an Amazon Simple Queue Service (Amazon SQS) standard queue. Subscribe the team’s email account to the SQS queue. Create an Amazon EventBridge rule that triggers when the AWS Glue job state changes to FAILESet the SQS queue as the target.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi mô tả tình huống một data engineer đang khắc phục sự cố cho AWS Glue workflow (luồng công việc xử lý dữ liệu ETL) thường xuyên thất bại do vấn đề chất lượng dữ liệu (data quality issues). Nhóm business reporting team cần nhận thông báo email mỗi khi workflow thất bại trong tương lai.
Yêu cầu giải pháp: Thiết lập hệ thống thông báo tự động, đáng tin cậy, tận dụng các dịch vụ AWS để phát hiện trạng thái FAILED của AWS Glue job và gửi email ngay lập tức.
📘 Kiến thức nền tảng (cập nhật AWS 2024-2026): AWS Glue jobs tự động phát ra events đến Amazon EventBridge khi trạng thái thay đổi (như FAILED, SUCCEEDED). EventBridge có thể trigger các target như SNS để gửi thông báo, và SNS hỗ trợ subscription email trực tiếp mà không cần code phức tạp.
✅ Đáp án đúng
Create an Amazon Simple Notification Service (Amazon SNS) standard topic. Subscribe the team’s email account to the SNS topic. Create an Amazon EventBridge rule that triggers when the AWS Glue job state changes to FAILED. Set the SNS topic as the target.
Lý do chọn đáp án này:
🛠️ Giải pháp sử dụng EventBridge rule để lắng nghe event Glue job state = FAILED (từ namespace aws.glue), sau đó target trực tiếp SNS standard topic – dịch vụ lý tưởng cho notifications realtime.
✅ SNS standard topic hỗ trợ subscription email trực tiếp (confirm qua email), không cần FIFO vì không yêu cầu thứ tự nghiêm ngặt. Không cần Lambda hay queue trung gian, đơn giản, chi phí thấp và scale tốt.
📘 Tài liệu tham khảo: AWS Glue Events in EventBridge & SNS Email Subscriptions (cập nhật 2024).
❌ Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, với đánh giá đúng/sai dựa trên tính khả thi, best practice AWS:
-
Create an Amazon Simple Notification Service (Amazon SNS) FIFO topic. Subscribe the team’s email account to the SNS topic. Create an AWS Lambda function that initiates when the AWS Glue job state changes to FAILED. Set the SNS topic as the target.
❌ Sai: SNS FIFO chỉ cần khi yêu cầu ordering nghiêm ngặt (exactly-once delivery), nhưng ở đây notifications email không cần – standard topic đủ và rẻ hơn. Thêm Lambda trung gian (trigger từ Glue state) là thừa thãi, phức tạp hóa (cần IAM roles, code), trong khi EventBridge target SNS trực tiếp. Không phải best practice. -
Create an Amazon Simple Notification Service (Amazon SNS) standard topic. Subscribe the team’s email account to the SNS topic. Create an Amazon EventBridge rule that triggers when the AWS Glue job state changes to FAILED. Set the SNS topic as the target.
✅ Đúng: Như đã giải thích ở phần đáp án đúng. Hoàn hảo, native integration giữa Glue → EventBridge → SNS → Email. Đơn giản, serverless, chi phí tối ưu. -
Create an Amazon Simple Queue Service (Amazon SQS) FIFO queue. Subscribe the team’s email account to the SQS queue. Create an AWS Config rule that triggers when the AWS Glue job state changes to FAILED. Set the SQS queue as the target.
❌ Sai: SQS không hỗ trợ subscription email trực tiếp (email không poll queue). AWS Config dùng để theo dõi configuration changes (như resource updates), không phải runtime states như Glue job FAILED (Config không capture events này). FIFO thừa, giải pháp không hoạt động. -
Create an Amazon Simple Queue Service (Amazon SQS) standard queue. Subscribe the team’s email account to the SQS queue. Create an Amazon EventBridge rule that triggers when the AWS Glue job state changes to FAILESet the SQS queue as the target.
❌ Sai: Tương tự, SQS không subscribe email trực tiếp (cần Lambda/SQS trigger để process và gửi SNS/email). EventBridge có thể target SQS, nhưng thêm queue làm chậm/lằng nhằng cho simple notification. Không hiệu quả so với SNS trực tiếp. (Lưu ý: Có lỗi typo "FAILESet" trong option gốc, nhưng không ảnh hưởng phân tích).
🛠️ Best practice khuyến nghị: Sử dụng CloudWatch Events/EventBridge + SNS cho monitoring Glue failures là pattern chuẩn trong AWS Well-Architected Framework (Reliability pillar). Test bằng cách chạy Glue job fail và kiểm tra email sau 1-2 phút! 🚀
The company needs to implement a monitoring mechanism that will alert stakeholders if the pipelines fail.
Which solution will meet these requirements with the LEAST operational overhead?
- A Create an Amazon EventBridge rule to match AWS Glue job failure events. Configure the rule to target an AWS Lambda function to process events. Configure the function to send notifications to an Amazon Simple Notification Service (Amazon SNS) topic.
- B Configure an Amazon CloudWatch Logs log group for the AWS Glue jobs. Create an Amazon EventBridge rule to match new log creation events in the log group. Configure the rule to target an AWS Lambda function that reads the logs and sends notifications to an Amazon Simple Notification Service (Amazon SNS) topic if AWS Glue job failure logs are present.
- C Create an Amazon EventBridge rule to match AWS Glue job failure events. Define an Amazon CloudWatch metric based on the EventBridge rule. Set up a CloudWatch alarm based on the metric to send notifications to an Amazon Simple Notification Service (Amazon SNS) topic.
- D Configure an Amazon CloudWatch Logs log group for the AWS Glue jobs. Create an Amazon EventBridge rule to match new log creation events in the log group. Configure the rule to send notifications to an Amazon Simple Notification Service (Amazon SNS) topic.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc triển khai cơ chế giám sát (monitoring) cho các AWS Glue jobs – những công việc xử lý dữ liệu ETL quan trọng trong data pipelines của công ty. ✅ Yêu cầu chính là thông báo (alert) cho stakeholders nếu pipeline thất bại (fail), đồng thời phải chọn giải pháp có operational overhead thấp nhất (LEAST operational overhead).
🛠️ Bối cảnh AWS Glue: AWS Glue là dịch vụ serverless ETL, tự động tạo logs trong CloudWatch Logs và phát ra các sự kiện (events) qua Amazon EventBridge khi job thay đổi trạng thái (như RUNNING, SUCCEEDED, FAILED). Mục tiêu là sử dụng các dịch vụ managed/serverless để tránh quản lý code tùy chỉnh, scaling thủ công hoặc parsing logs phức tạp, đảm bảo độ tin cậy cao với chi phí vận hành thấp. Kiến thức cập nhật đến 2026: AWS Glue hỗ trợ EventBridge events chi tiết hơn (bao gồm JobFailed event với metadata như jobRunId, error message), và tích hợp sâu với CloudWatch metrics/alarms mà không cần Lambda.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create an Amazon EventBridge rule to match AWS Glue job failure events. Define an Amazon CloudWatch metric based on the EventBridge rule. Set up a CloudWatch alarm based on the metric to send notifications to an Amazon Simple Notification Service (Amazon SNS) topic.
Lý do chi tiết 🏆:
- Giải pháp này serverless hoàn toàn, không cần viết code Lambda hay parsing logs thủ công, giảm thiểu operational overhead tối đa.
- EventBridge rule khớp trực tiếp với event
AWS::Glue::Jobtrạng tháiFAILED(có sẵn từ AWS Glue). - Từ rule, tạo CloudWatch metric (high-resolution metric) đếm số event failure → CloudWatch alarm tự động trigger khi metric > 0 → gửi thông báo qua SNS (hỗ trợ email/SMS/Slack).
- Least overhead: Chỉ cấu hình rule + metric + alarm (console/CLI/IaC), tự động scale, không maintain code/logic xử lý failure.
- Hiệu quả cao: Metric filter trên EventBridge events chính xác, realtime alerting.
📋 Giải thích tất cả các phương án (đúng/sai)
-
✅ Phương án ĐÚNG (như trên):
Create an Amazon EventBridge rule to match AWS Glue job failure events. Define an Amazon CloudWatch metric based on the EventBridge rule. Set up a CloudWatch alarm based on the metric to send notifications to an Amazon Simple Notification Service (Amazon SNS) topic.
Giải thích: Hoàn hảo vì tận dụng native integration Glue → EventBridge → CloudWatch metric/alarm → SNS. Không code, không logs parsing, overhead thấp nhất. ✅ -
❌ Phương án SAI:
Create an Amazon EventBridge rule to match AWS Glue job failure events. Configure the rule to target an AWS Lambda function to process events. Configure the function to send notifications to an Amazon Simple Notification Service (Amazon SNS) topic.
Giải thích: Mặc dù dùng EventBridge khớp failure events đúng, nhưng thêm Lambda function yêu cầu viết/deploy code (process events → SNS), phải manage permissions, cold starts, error handling → operational overhead cao hơn so với metric/alarm native. ❌ -
❌ Phương án SAI:
Configure an Amazon CloudWatch Logs log group for the AWS Glue jobs. Create an Amazon EventBridge rule to match new log creation events in the log group. Configure the rule to target an AWS Lambda function that reads the logs and sends notifications to an Amazon Simple Notification Service (Amazon SNS) topic if AWS Glue job failure logs are present.
Giải thích: Phức tạp thừa: Phải config Logs Insights, rule trigger trên log creation (không phải failure-specific), Lambda phải read/parse logs tìm keyword failure (như "Job failed") → tốn tài nguyên, chậm (log latency), dễ false positive/negative, overhead cao (code + log management). ❌ -
❌ Phương án SAI:
Configure an Amazon CloudWatch Logs log group for the AWS Glue jobs. Create an Amazon EventBridge rule to match new log creation events in the log group. Configure the rule to send notifications to an Amazon Simple Notification Service (Amazon SNS) topic.
Giải thích: Trigger trên new log events (bất kỳ log nào, không filter failure) → alert spam liên tục (mỗi job run đều có logs), không thông minh/intelligent. Không parse nội dung logs → không detect failure chính xác, overhead thấp nhưng không meet yêu cầu "if pipelines fail". ❌
📘 Tài liệu tham khảo (cập nhật AWS 2026)
- AWS Glue Monitoring: AWS Glue Developer Guide - Monitoring Jobs – Chi tiết EventBridge events cho Job states (FAILED).
- EventBridge + CloudWatch Metrics: Amazon EventBridge - Metrics from Events – Tạo custom metrics từ rules.
- CloudWatch Alarms for Glue: CloudWatch - Alarm on Glue Metrics – Native metrics + custom từ events.
- Best Practices DevOps: AWS Well-Architected Framework - Reliability Pillar (2024+): Khuyến nghị serverless alerting cho ETL jobs để least overhead.
Giải pháp này đảm bảo high availability, cost-effective cho critical pipelines! 🚀
One of the AWS Glue jobs begins to fail. A data engineer investigates the error and wants to examine metrics for all individual stages within the job.
How can the data engineer access the stage metrics?
- A Examine the AWS Glue job and stage details in the Spark UI.
- B Examine the AWS Glue job and stage metrics in Amazon CloudWatch.
- C Examine the AWS Glue job and stage logs in AWS CloudTrail logs.
- D Examine the AWS Glue job and stage details by using the run insights feature on the job.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào AWS Glue – một dịch vụ ETL serverless sử dụng Apache Spark để xử lý dữ liệu lớn. Công ty đang chạy các AWS Glue Apache Spark jobs cho workload ETL (Extract, Transform, Load), và đã kích hoạt logging và monitoring cho tất cả job.
📌 Tình huống cụ thể: Một job Glue bắt đầu fail. Data engineer cần xem metrics chi tiết cho từng individual stage (giai đoạn) trong job Spark đó.
- Stage ở đây đề cập đến các giai đoạn thực thi trong Spark (như shuffle, compute tasks), nơi metrics bao gồm thời gian chạy, CPU/memory usage, shuffle read/write, v.v.
- Mục tiêu: Tìm cách access stage metrics một cách chính xác và hiệu quả nhất, dựa trên tính năng native của AWS Glue (cập nhật đến 2026: AWS Glue version 4.0 hỗ trợ Spark 3.3+ với monitoring nâng cao).
🛠️ Bối cảnh kỹ thuật: AWS Glue tích hợp chặt chẽ với Spark UI (Spark Web UI) để debug sâu, đặc biệt cho stage-level metrics, thay vì chỉ metrics tổng quát.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Examine the AWS Glue job and stage details in the Spark UI.
Lý do 🏆:
- AWS Glue cung cấp Spark UI tích hợp trực tiếp trong AWS Management Console cho mỗi job run. Từ trang Jobs > Runs, bạn click vào job run cụ thể → Spark UI tab để xem chi tiết toàn bộ job: Executors, Stages, Storage, Environment, SQL, v.v.
- Stage metrics được hiển thị chi tiết ở tab Stages: Bao gồm metrics như Input/Output records, Shuffle read/write, Task duration, GC time, memory usage... Đây là cách chính thức và chi tiết nhất để debug failure ở mức stage (theo AWS best practices).
- Không cần tool bên ngoài; Spark UI lưu trữ history qua Spark History Server trong Glue, hỗ trợ lên đến 30 ngày (có thể config lâu hơn).
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên tính năng AWS Glue cập nhật 2026:
-
✅ [ĐÚNG] Examine the AWS Glue job and stage details in the Spark UI.
🛠️ Giải thích đúng: Như trên, đây là cách trực tiếp và đầy đủ nhất. Spark UI native trong Glue console cho phép drill-down metrics từng stage/task. Lý tưởng cho troubleshooting failure (ví dụ: stage fail do OOM hoặc skew data). -
❌ [SAI] Examine the AWS Glue job and stage metrics in Amazon CloudWatch.
🛠️ Giải thích sai: CloudWatch cung cấp metrics tổng quát cho Glue job nhưglue.driver.ExecutorUtilization,glue.driver.JVMHeapMemoryUtilization,glue.driver.CPUUtilization(ở mức driver/executor). Không có stage-level metrics chi tiết như Spark stages. CloudWatch Logs chỉ lưu driver/executor logs, không phải stage metrics. -
❌ [SAI] Examine the AWS Glue job and stage logs in AWS CloudTrail logs.
🛠️ Giải thích sai: CloudTrail ghi API calls và audit trail (như CreateJob, StartJobRun), không phải logs hay metrics runtime của Spark job/stage. CloudTrail không lưu nội dung ETL processing hay stage details – chỉ metadata quản trị. -
❌ [SAI] Examine the AWS Glue job and stage details by using the run insights feature on the job.
🛠️ Giải thích sai: Run insights (hay Glue Insights) là feature mới (từ 2023+) để phân tích tổng quan job runs như duration, cost, data processed qua ML-based recommendations. Không hỗ trợ stage-level metrics chi tiết; chỉ overview ở mức job, không drill-down như Spark UI.
📘 Tài liệu tham khảo (AWS chính thức, cập nhật 2026)
- AWS Glue Monitoring Spark UI: docs.aws.amazon.com/glue/latest/dg/monitor-spark-ui.html – Hướng dẫn truy cập Spark UI cho stage details.
- AWS Glue Metrics in CloudWatch: docs.aws.amazon.com/glue/latest/dg/monitor-cloudwatch-metrics.html – Xác nhận chỉ metrics tổng quát.
- AWS Glue Run Insights: docs.aws.amazon.com/glue/latest/dg/aws-glue-insights.html – Overview, không stage-level.
- Spark UI in AWS Glue (Video Guide): AWS re:Post & Blogs 2024-2026 về Glue 4.0 Spark 3.3+ monitoring.
Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ thực hành, hãy hỏi nhé.
The data engineer wants to improve query performance and to automate partition management.
Which solution will meet these requirements?
- A Use an AWS Lambda function that runs daily. Configure the function to manually create new partitions in AWS Glue for each day’s data.
- B Use partition projection in Athena. Configure the table properties by using a date range from 5 years ago to the present.
- C Reduce the number of partitions by changing the partitioning schema from daily to monthly granularity.
- D Increase the processing capacity of Athena queries by allocating more compute resources.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào vấn đề hiệu suất query chậm trên một bảng được phân vùng (partitioned table) cao trong Amazon Athena. Bảng này chứa dữ liệu hàng ngày (daily data) trong 5 năm qua, với phân vùng theo ngày (partitioned by date). Data engineer muốn cải thiện hiệu suất query và tự động hóa quản lý phân vùng (automate partition management).
🔍 Phân tích vấn đề chính:
- Với hàng ngàn phân vùng (khoảng 5 năm × 365 ngày ≈ 1.825 partitions), Athena phải quét metadata từ AWS Glue Catalog để prune partitions, dẫn đến query chậm do overhead metadata lớn.
- Giải pháp cần tự động hóa (không manual) và tối ưu performance mà không thay đổi dữ liệu gốc.
- Đây là tình huống phổ biến với dữ liệu time-series lớn, và AWS khuyến nghị sử dụng các tính năng như partition projection để xử lý.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use partition projection in Athena. Configure the table properties by using a date range from 5 years ago to the present.
Lý do chi tiết:
- Partition projection là tính năng của Athena (cập nhật đến 2026, hỗ trợ đầy đủ trong Athena engine version 3) cho phép tự động dự đoán và tạo ảo partitions dựa trên cấu hình (như date range), mà không cần AWS Glue crawler, Lambda hay manual MSCK REPAIR.
- Với date range từ 5 năm trước đến hiện tại, Athena sẽ prune partitions hiệu quả chỉ dựa trên query predicate (ví dụ: WHERE date = '2024-01-01'), giảm thời gian scan metadata xuống gần zero, cải thiện query performance lên đến 10x cho bảng partitioned lớn.
- Tự động hóa hoàn hảo: Không cần quản lý partitions thủ công, phù hợp dữ liệu daily liên tục.
- Đây là best practice từ AWS cho time-series data.
📋 Phân tích tất cả các phương án
🛠️ Phương án 1 ❌: Use an AWS Lambda function that runs daily. Configure the function to manually create new partitions in AWS Glue for each day’s data.
Giải thích sai: Phương án này yêu cầu Lambda chạy hàng ngày để manually thêm partitions vào Glue Catalog (sử dụng create_partition API). Mặc dù cải thiện prune partitions, nhưng không tự động hóa thực sự (phụ thuộc schedule, code custom, error-prone với dữ liệu lớn 5 năm), và tăng overhead vận hành (monitoring Lambda, cost). Không giải quyết partitions cũ tồn tại, chỉ thêm mới. Không phải giải pháp tối ưu theo AWS.
🛠️ Phương án 2 ✅: Use partition projection in Athena. Configure the table properties by using a date range from 5 years ago to the present.
Giải thích đúng: Như đã phân tích ở trên, partition projection tự động generate partitions ảo dựa trên config (storage.location.template cho S3 path như s3://bucket/YYYY/MM/DD/). Athena engine version 3 (2026) hỗ trợ interval, range, int cho date, giảm query time đáng kể mà không chạm metadata Glue. Hoàn hảo cho daily partitions liên tục.
🛠️ Phương án 3 ❌: Reduce the number of partitions by changing the partitioning schema from daily to monthly granularity.
Giải thích sai: Chuyển sang monthly partitions giảm số lượng (từ ~1.825 xuống ~60), giúp prune nhanh hơn, nhưng yêu cầu re-partition dữ liệu S3 (repurposing lớn, downtime, cost cao với 5 năm data). Không tự động hóa management (vẫn cần handle monthly), và làm query chậm hơn nếu cần daily granularity (scan full month). Không khuyến khích cho existing data theo AWS best practices.
🛠️ Phương án 4 ❌: Increase the processing capacity of Athena queries by allocating more compute resources.
Giải thích sai: Athena là serverless, tự scale compute dựa trên data scanned (qua Workgroups với query timeout/enforcement). Tăng compute (như set bytes scanned limit cao hơn) không giải quyết root cause là metadata overhead từ partitions lớn (vẫn scan Glue Catalog chậm). Chỉ tốn kém hơn (pay-per-TB scanned), không automate partitions. Athena docs khuyên optimize partitions trước compute.
📘 Tài liệu tham khảo
- AWS Athena Documentation: Partition Projection (cập nhật 2024-2026, engine v3 hỗ trợ advanced config).
- AWS Glue & Athena Best Practices: Improving Athena Performance (partition projection là #1 cho high-partitioned tables).
- AWS Well-Architected Framework - Analytics Lens: Khuyến nghị partition projection cho time-series >1000 partitions.
Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ code config projection, hãy hỏi nhé!
The workflow must either process or reject each transaction within 24-hours. The workflow must run for less than 24 hours total.
Which solution will meet these requirements with the LEAST operational cost?
- A Create a standard workflow in AWS Step Functions. Implement a Wait for Callback pattern to wait for the validation steps to finish.
- B Create an express workflow in AWS Step Functions. Implement a Wait for Callback pattern to wait for the validation steps to finish.
- C Use AWS Lambda functions to implement the workflow. Use Amazon EventBridge to invoke the validation steps.
- D Use Amazon Managed Workflows for Apache Airflow (Amazon MWAA) to implement the workflow.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc triển khai một workflow xử lý giao dịch (transactions) trên AWS, với các yêu cầu chính sau:
- Mỗi giao dịch phải trải qua nhiều cấp độ validation, trong đó mỗi cấp phụ thuộc vào cấp trước đó (sequential dependency).
- Workflow phải xử lý hoặc từ chối mỗi giao dịch trong vòng 24 giờ.
- Tổng thời gian chạy workflow phải dưới 24 giờ.
- Giải pháp cần chi phí vận hành thấp nhất (LEAST operational cost).
📘 Bối cảnh kỹ thuật: Workflow có tính chất long-running do các bước validation có thể mất thời gian (cần chờ callback từ hệ thống bên ngoài), nhưng tổng execution time <24h. AWS Step Functions là dịch vụ orchestration lý tưởng cho serverless workflows. Kiến thức cập nhật đến 2026: Standard Workflows hỗ trợ execution timeout lên đến 1 năm, Express Workflows tối đa 30 phút (từ update 2023), và chỉ Standard hỗ trợ Wait for Callback pattern (.waitForTaskToken).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create a standard workflow in AWS Step Functions. Implement a Wait for Callback pattern to wait for the validation steps to finish.
🛠️ Lý do chi tiết:
- Standard Workflows trong AWS Step Functions hỗ trợ Wait for Callback pattern (sử dụng
.waitForTaskToken), cho phép workflow tạm dừng chờ callback từ dịch vụ bên ngoài (như validation steps) mà không timeout ngay. Workflow có thể chạy tổng <24h, phù hợp yêu cầu. - Chi phí thấp nhất: Chỉ tính phí dựa trên state transitions (0.000025 USD/1.000 transitions), không tính thời gian chờ (idle time miễn phí). Không cần quản lý server, operational overhead thấp.
- Đáp ứng sequential dependency qua ASL (Amazon States Language).
📘 Tài liệu tham khảo:
- AWS Step Functions Developer Guide - Standard vs Express (cập nhật 2025).
- Callback Patterns.
- Pricing (Standard rẻ hơn cho long-running).
📋 Giải thích tất cả các phương án (đúng/sai)
-
✅ Create a standard workflow in AWS Step Functions. Implement a Wait for Callback pattern to wait for the validation steps to finish.
🟢 Đúng vì: Hỗ trợ đầy đủ callback cho validation dài hơi, execution <24h OK, chi phí chỉ state transitions (rẻ nhất cho low-volume transactions). Không cần quản lý infra. -
❌ Create an express workflow in AWS Step Functions. Implement a Wait for Callback pattern to wait for the validation steps to finish.
🔴 Sai vì: Express Workflows KHÔNG hỗ trợ Wait for Callback (.waitForTaskToken chỉ dành cho Standard). Execution timeout tối đa 30 phút (không đủ cho validation 24h). Dù rẻ hơn cho high-volume short-duration, nhưng không đáp ứng yêu cầu callback và thời gian. -
❌ Use AWS Lambda functions to implement the workflow. Use Amazon EventBridge to invoke the validation steps.
🔴 Sai vì: Lambda + EventBridge chỉ là event-driven invocation, không phải workflow orchestration tự nhiên (khó quản lý state, retry, dependency sequential). Operational cost cao: Phải tự code error handling, monitoring; Lambda timeout 15 phút/step. Không hiệu quả, chi phí tích lũy từ invocations cao hơn Step Functions. -
❌ Use Amazon Managed Workflows for Apache Airflow (Amazon MWAA) to implement the workflow.
🔴 Sai vì: MWAA dành cho data pipelines batch-oriented (DAGs), không phù hợp transaction realtime/long-running. Chi phí cao: Managed cluster EC2 luôn chạy (min 0.98 USD/giờ/environment), scheduler + workers. Overhead lớn, không serverless, execution >24h dễ vượt cost so với Step Functions.
🧠 Kết luận: Standard Step Functions là lựa chọn tối ưu về cost/effort cho workflow sequential với callbacks! 🚀
The solution requires a long-running workflow with 15 GiB memory capacity to process the data concurrently, followed by a correlation process that begins only after the first two processes complete.
Which solution will meet these requirements with the LEAST operational overhead?
- A Use Amazon Managed Workflows for Apache Airflow (Amazon MWAA) to orchestrate the workflow by using AWS Glue. Configure AWS Glue to begin the third process after the first two processes have finished.
- B Use Amazon EMR to run each process in the workflow. Create an Amazon Simple Queue Service (Amazon SQS) queue to handle messages that indicate the completion of the first two processes. Configure an AWS Lambda function to process the SQS queue by running the third process.
- C Use AWS Glue workflows to run the first two processes in parallel. Ensure that the third process starts after the first two processes have finished.
- D Use AWS Step Functions to orchestrate a workflow that uses multiple AWS Lambda functions. Ensure that the third process starts after the first two processes have finished.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc thiết kế một workflow xử lý dữ liệu ETL (Extract, Transform, Load) trên AWS với yêu cầu cụ thể sau:
- Dữ liệu đầu vào: 500 GB dữ liệu audience và advertising hàng ngày, lưu dưới dạng CSV files trong Amazon S3, với schemas đã đăng ký trong AWS Glue Data Catalog.
- Yêu cầu xử lý: Chuyển đổi các file CSV này sang định dạng Apache Parquet và lưu vào một S3 bucket khác.
- Yêu cầu workflow:
🛠️ Workflow phải dài hạn (long-running).
🧠 Mỗi process cần 15 GiB memory capacity để xử lý dữ liệu concurrently (song song).
📊 Có 3 processes: Hai processes đầu chạy song song (parallel), sau đó process thứ ba (correlation process) chỉ bắt đầu sau khi hai processes đầu hoàn thành. - Tiêu chí chọn giải pháp: LEAST operational overhead (ít overhead vận hành nhất), nghĩa là ưu tiên giải pháp serverless, tự động scale, không cần quản lý infrastructure thủ công.
Đây là tình huống điển hình cho ETL jobs lớn trên AWS, nơi AWS Glue là dịch vụ chuyên biệt cho data catalog và ETL serverless với Spark/Parquet support. Workflow cần orchestration đơn giản với dependencies (parallel + sequential).
📘 Tài liệu tham khảo:
- AWS Glue Documentation: AWS Glue Workflows (cập nhật 2024-2026).
- AWS Well-Architected Framework - Data Analytics Lens (2025 edition).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use AWS Glue workflows to run the first two processes in parallel. Ensure that the third process starts after the first two processes have finished.
Lý do:
🛠️ AWS Glue Workflows là tính năng built-in của AWS Glue, cho phép orchestrate các Glue ETL jobs một cách serverless, hỗ trợ parallel execution (hai jobs đầu chạy song song) và dependencies (job thứ ba chỉ trigger sau khi hai job đầu hoàn thành 100%).
💡 Phù hợp hoàn hảo:
- Hỗ trợ 15 GiB memory (Glue Spark jobs scale lên đến 100+ GB, DPU-based).
- Long-running workflows với auto-scaling cho 500 GB data.
- Least overhead: Không cần quản lý cluster, Airflow, hay queue; chỉ định nghĩa jobs và triggers trong console/CLI.
- Tích hợp native với S3 + Glue Data Catalog + Parquet output.
✅ Đây là giải pháp AWS-native, serverless ETL orchestration với chi phí thấp nhất và vận hành đơn giản nhất (zero infrastructure management).
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt với emoji để nổi bật.
-
Use Amazon Managed Workflows for Apache Airflow (Amazon MWAA) to orchestrate the workflow by using AWS Glue. Configure AWS Glue to begin the third process after the first two processes have finished.
❌ Sai: MWAA là managed Airflow tốt cho complex DAGs, nhưng tạo overhead cao (cần config environment, VPC, scaling, monitoring Airflow). Không "least overhead" vì phải quản lý DAGs, scheduler, và dependencies thủ công qua code Python. Glue workflows đơn giản hơn cho ETL thuần túy. -
Use Amazon EMR to run each process in the workflow. Create an Amazon Simple Queue Service (Amazon SQS) queue to handle messages that indicate the completion of the first two processes. Configure an AWS Lambda function to process the SQS queue by running the third process.
❌ Sai: EMR yêu cầu quản lý cluster (provisioning, scaling, termination), overhead lớn cho long-running jobs. SQS + Lambda thêm complexity (polling, error handling), không serverless thuần. Không hiệu quả cho ETL S3-Parquet với Glue Catalog; EMR phù hợp hơn cho custom Hadoop/Spark lớn nhưng overhead cao. -
Use AWS Glue workflows to run the first two processes in parallel. Ensure that the third process starts after the first two processes have finished.
✅ Đúng: Như đã giải thích ở trên. Glue Workflows hỗ trợ parallel branches (hai jobs concurrent với 15 GiB/GPU) và All-to-One join trigger cho job thứ ba. Serverless 100%, tích hợp Data Catalog, tối ưu cho Parquet conversion. Least overhead! -
Use AWS Step Functions to orchestrate a workflow that uses multiple AWS Lambda functions. Ensure that the third process starts after the first two processes have finished.
❌ Sai: Step Functions tốt cho orchestration, nhưng Lambda chỉ hỗ trợ max 10.24 GB memory (không đủ 15 GiB). Không phù hợp cho data-intensive ETL 500 GB (Lambda timeout 15 phút, không long-running). Phải dùng Container/ECS thay Lambda → tăng overhead so với Glue native.
🏆 Kết luận & Best Practices
Giải pháp AWS Glue Workflows là optimal cho ETL serverless với dependencies. Để implement: Tạo 3 Glue Jobs (Spark script cho CSV→Parquet), rồi trigger workflow với parallel triggers.
🔍 Test tip: Trong exam DOP-C02, ưu tiên "least overhead" → serverless + native service (Glue > MWAA/EMR/Step Functions). Theo AWS 2026 updates, Glue Workflows vẫn là gold standard cho Data Catalog-based ETL.
What additional step must the company take to meet this requirement?
- A Create a service control policy (SCP) to grant the data stream read access to the cross-account Lambda execution role. Attach the SCP to Account A.
- B Add a resource-based policy to the data stream to allow read access for the cross-account Lambda execution role.
- C Create a service control policy (SCP) to grant the data stream read access to the cross-account Lambda execution role. Attach the SCP to Account B.
- D Add a resource-based policy to the cross-account Lambda function to grant the data stream read access to the function.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh tình huống cross-account access trong AWS, cụ thể là việc sử dụng Amazon Kinesis Data Streams với chế độ enhanced fanout để nhận dữ liệu streaming từ nhiều producers ở Account A. Công ty muốn AWS Lambda ở Account B xử lý dữ liệu từ stream này. Họ đã tạo Lambda execution role ở Account B với các quyền cần thiết (như kinesis:DescribeStream, kinesis:GetRecords, v.v.) để truy cập stream. Tuy nhiên, để Lambda ở Account B có thể đọc dữ liệu từ stream ở Account A, cần một bước bổ sung để cho phép truy cập cross-account.
Vấn đề cốt lõi: AWS yêu cầu hai chiều quyền hạn cho cross-account:
- Lambda role ở Account B phải có IAM policy cho phép hành động trên resource (stream ARN ở Account A).
- Resource (Kinesis stream ở Account A) phải có resource-based policy tin cậy (trust) Lambda service và role ARN từ Account B.
Điều này dựa trên least privilege principle và cập nhật AWS đến 2026: Kinesis Data Streams hỗ trợ resource policies cho cross-account consumers như Lambda (xem AWS Well-Architected Framework - Reliability Pillar).
📘 Tài liệu tham khảo:
- AWS Docs: Control access to Kinesis Data Streams resources using resource-based policies (cập nhật 2025).
- AWS Docs: Using Lambda with Kinesis cross-account (hướng dẫn event source mapping cross-account).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Add a resource-based policy to the data stream to allow read access for the cross-account Lambda execution role.
Lý do 🛠️:
- Kinesis Data Streams hỗ trợ resource-based policies (bucket policy tương tự S3) để grant quyền cross-account trực tiếp trên stream ARN.
- Policy này phải cho phép principal là Lambda service (
lambda.amazonaws.com) với conditionArnLikekhớp Lambda role ARN ở Account B, và actions nhưkinesis:GetRecords*, kinesis:DescribeStream. - Đây là bước bắt buộc bổ sung vì Lambda execution role chỉ cấp quyền từ phía consumer; resource policy từ phía producer (stream) mới mở cửa cross-account.
- Không cần thay đổi Organizations SCP vì SCP chỉ hạn chế (deny), không grant permissions.
📋 Phân tích tất cả các phương án (đúng/sai)
-
❌ Phương án SAI: Create a service control policy (SCP) to grant the data stream read access to the cross-account Lambda execution role. Attach the SCP to Account A.
Giải thích sai: SCP trong AWS Organizations chỉ dùng để hạn chế quyền (whitelist/deny) cho các tài khoản con, không thể grant quyền cụ thể cross-account cho resource như Kinesis stream. SCP không hỗ trợ resource-level grants; nó áp dụng ở organizational level và không thay thế IAM/resource policies. Attach vào Account A vô ích vì SCP không grant đọc stream. -
✅ Phương án ĐÚNG: Add a resource-based policy to the data stream to allow read access for the cross-account Lambda execution role.
Giải thích đúng: Như đã nêu ở phần đáp án, đây là cách chuẩn AWS cho Kinesis streams (enhanced fanout). Ví dụ policy JSON:{ "Statement": [{ "Effect": "Allow", "Principal": {"Service": "lambda.amazonaws.com"}, "Action": ["kinesis:DescribeStream", "kinesis:GetShardIterator", "kinesis:GetRecords"], "Resource": "arn:aws:kinesis:region:AccountA:stream/stream-name", "Condition": {"ArnLike": {"AWS:SourceArn": "arn:aws:lambda:region:AccountB:function:function-name"}} }] }Áp dụng qua AWS Console/CLI:
aws kinesis put-stream-resource-policy. -
❌ Phương án SAI: Create a service control policy (SCP) to grant the data stream read access to the cross-account Lambda execution role. Attach the SCP to Account B.
Giải thích sai: Tương tự phương án đầu, SCP không grant quyền mà chỉ restrict. Attach vào Account B (consumer) không ảnh hưởng đến stream ở Account A (producer). SCP không xử lý cross-account resource access cho Kinesis. -
❌ Phương án SAI: Add a resource-based policy to the cross-account Lambda function to grant the data stream read access to the function.
Giải thích sai: Lambda functions không hỗ trợ resource-based policies cho việc grant quyền từ resource khác (như Kinesis). Resource policy phải đặt trên Kinesis stream (provider), không phải Lambda (consumer). Lambda chỉ dùng execution role IAM policy cho outbound access.
Kết luận 🎯: Bước này đảm bảo enhanced fanout hoạt động cross-account mượt mà, hỗ trợ up to 20 consumers/shard với low latency (<70ms). Test bằng CloudWatch metrics như GetRecords.IteratorAgeMilliseconds.
Which solution will meet these requirements?
- A Use AWS Glue DataBrew to create the partitions for the AWS Glue table.
- B Use an AWS Lambda function to create the partitions for the AWS Glue table.
- C Set partition projection properties for the AWS Glue table.
- D Configure an AWS Glue crawler to run on a set schedule.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh việc sử dụng Amazon Athena để phân tích dữ liệu lưu trữ trong Amazon S3 bucket. Một data engineer cần cấu hình phân vùng (partitions) cho bảng AWS Glue Data Catalog theo các cấp độ year (năm), month (tháng) và day (ngày). Yêu cầu đặc biệt là phải tạo partitions mỗi ngày để thích ứng với thay đổi schema (schema changes) trong dữ liệu mới.
Mục tiêu là tìm giải pháp tự động, hiệu quả, không cần can thiệp thủ công hàng ngày, vì dữ liệu thay đổi liên tục và partitions cần được cập nhật kịp thời để Athena query nhanh chóng, tiết kiệm chi phí (nhờ partition pruning). Đây là tình huống phổ biến với dữ liệu time-series lớn trong S3 (ví dụ: s3://bucket/year=2024/month=10/day=15/data.parquet). Giải pháp phải phù hợp với best practices AWS năm 2024-2026, ưu tiên serverless và no-maintenance.
✅ Đáp án đúng: Set partition projection properties for the AWS Glue table.
Lý do lựa chọn:
- Partition Projection là tính năng của AWS Glue Data Catalog (cập nhật mới nhất đến 2026) cho phép tự động "chiếu" (project) partitions dựa trên pattern thời gian (year/month/day) mà KHÔNG cần tạo partitions thủ công, crawl hay thêm metadata.
- Nó sử dụng các thuộc tính cấu hình (như
projection.enabled=true,projection.year.type=enum,projection.month.type=int:1:12,projection.day.type=int:1:31) để Glue/Athena tự suy luận partitions từ đường dẫn S3 khi query. - Hoàn hảo cho schema changes hàng ngày: Dữ liệu mới tự động được nhận diện mà không cần update catalog thủ công, giảm latency và chi phí (không scan toàn bộ dữ liệu).
- Ưu điểm: Serverless, scale tự động, hỗ trợ Athena query nhanh với partition pruning. Đây là recommended solution cho S3 partitioned data theo time-based (AWS Well-Architected Framework - Reliability pillar).
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh:
-
❌ Use AWS Glue DataBrew to create the partitions for the AWS Glue table.
Sai vì: AWS Glue DataBrew là công cụ data preparation và transformation (giống Pandas trên cloud), dùng để clean/visualize data chứ KHÔNG hỗ trợ tạo partitions cho Glue table. Nó không tích hợp trực tiếp với Glue Catalog để quản lý partitions động hàng ngày. Sử dụng sẽ phức tạp, không tự động và không giải quyết schema changes. -
❌ Use an AWS Lambda function to create the partitions for the AWS Glue table.
Sai vì: Lambda có thể dùng AWS SDK (Glue client.add_partition()) để tạo partitions thủ công, nhưng yêu cầu code custom, schedule (EventBridge), và maintain logic xử lý schema changes mỗi ngày. Không phải best practice: Tốn công phát triển, dễ lỗi nếu dữ liệu lớn/schema thay đổi bất ngờ, và chi phí invoke cao hơn Partition Projection (no-maintenance). -
✅ Set partition projection properties for the AWS Glue table.
Đúng vì: Như đã giải thích ở trên, đây là giải pháp native, zero-maintenance của AWS Glue. Cấu hình một lần qua Console/CLI/Terraform (ví dụ: ALTER TABLE SET projection.enabled=true), Athena tự handle partitions mới hàng ngày từ S3 path, hỗ trợ schema evolution tự động. Phù hợp 100% yêu cầu. -
❌ Configure an AWS Glue crawler to run on a set schedule.
Sai vì: Crawler có thể detect partitions và schema từ S3, nhưng chạy schedule hàng ngày sẽ chậm (scan metadata, mất 15-60 phút), tốn chi phí (DPUs), và KHÔNG lý tưởng cho schema changes thường xuyên (có thể overwrite schema cũ gây query lỗi). Crawler phù hợp one-time discovery, không phải daily ops.
🛠️ Khuyến nghị triển khai thực tế
- Bước cấu hình Partition Projection (CLI ví dụ):
aws glue update-table --database-name mydb --table-input '{"Name":"mytable","StorageDescriptor":{"Location":"s3://bucket/","PartitionKeys":[{"Name":"year","Type":"string"},{"Name":"month","Type":"string"},{"Name":"day","Type":"string"}],"Projection":{"Enabled":true,"ProjectionType":"enum:int:string","Parameters":{"year":"2020/2021/2022/...","month":"01/02/.../12","day":"01/02/.../31"}}}' - Test query Athena:
SELECT * FROM table WHERE year='2024' AND month='10' AND day='15';→ Tự prune!
📘 Tài liệu tham khảo (AWS Docs cập nhật 2026)
- AWS Glue Partition Projection – Hướng dẫn chính thức.
- Athena Partition Projection – Tích hợp Athena.
- AWS re:Post - Best Practices for S3 Partitions – So sánh với crawler/Lambda.
- AWS Well-Architected: Data Analytics Lens (2024 edition).
Giải pháp này đảm bảo high availability, cost-optimized cho workload lớn! 🚀