Ngân hàng đề — AWS Certified Data Engineer Associate
Tìm thấy 867 câu.
Which combination of tasks will meet these requirements with the LEAST operational overhead? (Choose two.)
- A Configure AWS Glue triggers to run the ETL jobs every hour.
- B Use AWS Glue DataBrew to clean and prepare the data for analytics.
- C Use AWS Lambda functions to schedule and run the ETL jobs every hour.
- D Use AWS Glue connections to establish connectivity between the data sources and Amazon Redshift.
- E Use the Redshift Data API to load transformed data into Amazon Redshift.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS
📘 Nội dung câu hỏi:
Câu hỏi mô tả một data engineer đang xây dựng data pipeline trên AWS bằng cách sử dụng AWS Glue ETL jobs (extract, transform, load). Nguồn dữ liệu đầu vào là Amazon RDS (cơ sở dữ liệu quan hệ) và MongoDB (cơ sở dữ liệu NoSQL). Sau khi xử lý biến đổi (transformations), dữ liệu được load vào Amazon Redshift để phục vụ phân tích (analytics). Yêu cầu quan trọng là pipeline phải cập nhật dữ liệu mỗi giờ (hourly basis).
Mục tiêu là chọn KẾT HỢP HAI nhiệm vụ (choose two) giúp đáp ứng yêu cầu với ít overhead vận hành nhất (LEAST operational overhead), nghĩa là giải pháp tự động hóa cao, dễ quản lý, không cần can thiệp thủ công nhiều, tận dụng các tính năng native của AWS Glue để giảm chi phí và độ phức tạp.
✅ Đáp án đúng (chọn TWO):
- Configure AWS Glue triggers to run the ETL jobs every hour.
- Use AWS Glue connections to establish connectivity between the data sources and Amazon Redshift.
🛠️ Lý do chọn đáp án đúng:
Hai lựa chọn này kết hợp hoàn hảo để tạo pipeline ETL tự động, serverless với Glue:
- Glue Triggers cho phép lập lịch chạy job định kỳ (cron-like, ví dụ mỗi giờ) mà không cần dịch vụ bên ngoài như Lambda hay cron jobs thủ công, giảm thiểu quản lý infrastructure.
- Glue Connections là tính năng native hỗ trợ kết nối JDBC cho RDS/Redshift và MongoDB connector (hỗ trợ từ Glue 2.0+), tự động quản lý credentials, VPC, security groups, giúp ETL job đọc/ghi dữ liệu seamless mà không cần code tùy chỉnh phức tạp.
Kết hợp chúng mang lại least operational overhead vì toàn bộ pipeline chạy tự động trên managed service của AWS, scale theo nhu cầu, không cần monitor server hay custom scheduler (cập nhật đến Glue version 4.0 năm 2024-2026).
🔍 Giải thích chi tiết TẤT CẢ các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh, với lý do đúng/sai bằng tiếng Việt. Tôi sử dụng ✅ cho đúng và ❌ cho sai để dễ theo dõi.
-
✅ Configure AWS Glue triggers to run the ETL jobs every hour.
Phương án này ĐÚNG vì AWS Glue hỗ trợ triggers (schedule, event-based, hoặc conditional) để tự động kích hoạt ETL jobs theo lịch cố định như mỗi giờ (sử dụng cron expression ví dụ0 * ? * * *). Đây là cách native, serverless, không cần thêm dịch vụ nào khác, giảm overhead tối đa so với custom scheduler. Glue triggers tích hợp trực tiếp với job bookmarks để xử lý incremental data, phù hợp hourly updates từ RDS/MongoDB. -
❌ Use AWS Glue DataBrew to clean and prepare the data for analytics.
Phương án này SAI vì AWS Glue DataBrew là công cụ interactive, visual-based cho data preparation/cleaning (như profiling, transformations đơn giản), dành cho analyst/explorer chứ không phải ETL pipeline tự động hourly. DataBrew không hỗ trợ scheduling jobs định kỳ như Glue ETL, và không tối ưu cho bulk load vào Redshift từ RDS/MongoDB, dẫn đến overhead cao hơn khi phải kết hợp thủ công. -
❌ Use AWS Lambda functions to schedule and run the ETL jobs every hour.
Phương án này SAI dù Lambda có thể dùng CloudWatch Events để trigger Glue jobs, nhưng nó tạo thêm layer phức tạp (viết code Lambda, quản lý IAM roles, timeout limits), tăng operational overhead so với Glue triggers native. Glue được thiết kế để tự quản lý scheduling, nên dùng Lambda là redundant và không "least overhead". -
✅ Use AWS Glue connections to establish connectivity between the data sources and Amazon Redshift.
Phương án này ĐÚNG vì Glue Connections lưu trữ và quản lý kết nối dữ liệu một lần dùng cho tất cả jobs: hỗ trợ JDBC cho RDS/Redshift (MySQL/PostgreSQL), MongoDB connector (từ Glue 3.0+ với VPC peering), tự động handle SSL/VPC/Secrets Manager. Không cần hard-code credentials trong script PySpark/Scala, giảm lỗi và overhead bảo trì, lý tưởng cho pipeline multi-source như RDS + MongoDB → Redshift. -
❌ Use the Redshift Data API to load transformed data into Amazon Redshift.
Phương án này SAI vì Redshift Data API dành cho ad-hoc queries/inserts từ ứng dụng/Lambda (REST API, batch nhỏ), không phù hợp bulk load từ Glue ETL (dữ liệu lớn, hourly). Nó yêu cầu code tùy chỉnh xử lý JSON payloads, gặp limits (1MB/request, concurrency), và overhead cao hơn so với Glue's native Redshift sink (S3 staging + COPY command tự động).
📚 Tài liệu tham khảo (cập nhật mới nhất AWS 2024-2026)
- AWS Glue Triggers: AWS Glue Developer Guide - Triggers (hỗ trợ cron scheduling từ Glue 1.0+, tối ưu hóa version 4.0).
- AWS Glue Connections: AWS Glue Programming ETL Connectors (MongoDB connector docs: here).
- So sánh DataBrew vs Glue: AWS Glue vs DataBrew.
- Redshift Data API limits: Amazon Redshift Data API (không khuyến nghị cho ETL bulk).
Các tính năng này ổn định đến 2026 theo AWS Well-Architected Framework for Data Analytics.
Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ code Glue script, hãy hỏi nhé!
Which solution will meet this requirement?
- A Turn on concurrency scaling in workload management (WLM) for Redshift Serverless workgroups.
- B Turn on concurrency scaling at the workload management (WLM) queue level in the Redshift cluster.
- C Turn on concurrency scaling in the settings during the creation of any new Redshift cluster.
- D Turn on concurrency scaling for the daily usage quota for the Redshift cluster.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào Amazon Redshift cluster chạy trên RA3 nodes (một loại node hỗ trợ lưu trữ tách biệt, cho phép scale compute độc lập). Công ty muốn scale read và write capacity để đáp ứng nhu cầu tăng đột biến. Nhiệm vụ của data engineer là bật tính năng Concurrency Scaling – một cơ chế tự động thêm cluster tạm thời (concurrency clusters) để xử lý các truy vấn đồng thời cao, hỗ trợ cả read và write queries mà không làm gián đoạn workload chính.
Yêu cầu chính: Tìm giải pháp bật Concurrency Scaling phù hợp cho Redshift cluster provisioned (không phải Serverless).
📘 Kiến thức nền: Concurrency Scaling chỉ khả dụng cho provisioned clusters (như RA3), giới hạn tối đa 10 concurrency clusters, mỗi cluster chạy tối đa 1 giờ/query, và được quản lý qua Workload Management (WLM). Tính năng này miễn phí cho 1 giờ sử dụng mỗi ngày (tính đến 2024-2026, theo AWS Well-Architected Framework).
🛠️ Nguồn tham khảo:
- AWS Redshift Docs: Concurrency Scaling
- AWS Redshift User Guide: WLM Queue Configuration (cập nhật 2024).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Turn on concurrency scaling at the workload management (WLM) queue level in the Redshift cluster.
Lý do:
- Concurrency Scaling được kích hoạt chính xác tại mức WLM queue trong Redshift provisioned cluster (như RA3 nodes).
- Bạn truy cập qua Query Editor v2 hoặc AWS Console > Clusters > Workload management, chọn queue cụ thể (ví dụ: queue mặc định hoặc custom), rồi bật "Concurrency scaling mode" thành Auto.
- Điều này cho phép cluster tự động scale thêm compute cho queries spill over từ queue chính, hỗ trợ read/write scaling hiệu quả.
- Phù hợp hoàn hảo với cluster hiện có (không cần tạo mới), và RA3 nodes hỗ trợ đầy đủ tính năng này.
🧩 Lợi ích: Giảm latency truy vấn lên đến 97% trong peak load, theo AWS benchmarks 2024.
❌ Giải thích tất cả các phương án
-
[SAI] Turn on concurrency scaling in workload management (WLM) for Redshift Serverless workgroups.
❌ Sai vì: Câu hỏi đề cập Redshift cluster RA3 nodes (provisioned cluster), không phải Redshift Serverless. Serverless có cơ chế auto-scaling riêng (base và concurrency capacity), và Concurrency Scaling chỉ được hỗ trợ ở Serverless từ 2023 nhưng không cấu hình qua WLM queue giống provisioned. Áp dụng sai môi trường sẽ không work. -
[ĐÚNG] Turn on concurrency scaling at the workload management (WLM) queue level in the Redshift cluster.
✅ Đúng như đã giải thích ở trên: Đây là cách chuẩn và chính xác nhất cho provisioned clusters, bật tại queue level để tự động handle overflow queries. -
[SAI] Turn on concurrency scaling in the settings during the creation of any new Redshift cluster.
❌ Sai vì: Concurrency Scaling không được bật lúc tạo cluster mới (cluster parameter group hoặc launch settings). Nó phải cấu hình sau khi cluster running, qua WLM queue. Bạn có thể enable mặc định qua parameter group (concurrency_scaling_cluster), nhưng không phải "during creation settings", và câu hỏi là cho cluster đã tồn tại. -
[SAI] Turn on concurrency scaling for the daily usage quota for the Redshift cluster.
❌ Sai vì: Daily usage quota (1 giờ miễn phí/ngày) là giới hạn billing, không phải nơi bật tính năng. Concurrency Scaling tự động sử dụng quota này khi bật qua WLM, nhưng không cấu hình quota để bật scaling. Quota được quản lý riêng qua AWS Support nếu cần tăng.
Which combination of steps will meet these requirements MOST cost-effectively? (Choose two.)
- A Use an AWS Lambda function and the Athena Boto3 client start_query_execution API call to invoke the Athena queries programmatically.
- B Create an AWS Step Functions workflow and add two states. Add the first state before the Lambda function. Configure the second state as a Wait state to periodically check whether the Athena query has finished using the Athena Boto3 get_query_execution API call. Configure the workflow to invoke the next query when the current query has finished running.
- C Use an AWS Glue Python shell job and the Athena Boto3 client start_query_execution API call to invoke the Athena queries programmatically.
- D Use an AWS Glue Python shell script to run a sleep timer that checks every 5 minutes to determine whether the current Athena query has finished running successfully. Configure the Python shell script to invoke the next query when the current query has finished running.
- E Use Amazon Managed Workflows for Apache Airflow (Amazon MWAA) to orchestrate the Athena queries in AWS Batch.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi yêu cầu một data engineer cần orchestrate (điều phối) một chuỗi các truy vấn Amazon Athena chạy hàng ngày, với mỗi truy vấn có thể chạy hơn 15 phút. Yêu cầu chọn TWO (hai) bước kết hợp MOST cost-effectively (tiết kiệm chi phí nhất).
- Amazon Athena là dịch vụ serverless query dữ liệu trên S3 bằng SQL, tính phí theo dữ liệu scan (TB scanned), không tính phí thời gian chạy.
- Thách thức chính: Các query chạy lâu (>15 phút), nên không thể chạy trực tiếp trong AWS Lambda (timeout tối đa 15 phút). Cần cơ chế start query và poll (kiểm tra trạng thái) để chờ hoàn thành trước khi chạy query tiếp theo.
- Mục tiêu: Giải pháp rẻ nhất (cost-effective), phù hợp orchestrate hàng ngày, serverless, không lãng phí tài nguyên. Sử dụng kiến thức AWS cập nhật đến 2026: Athena hỗ trợ Boto3 SDK (start_query_execution và get_query_execution), Step Functions hỗ trợ Wait state cho polling dài hạn mà không timeout.
✅ Đáp án đúng (Chọn TWO)
Hai lựa chọn đúng là:
- Use an AWS Lambda function and the Athena Boto3 client start_query_execution API call to invoke the Athena queries programmatically.
- Create an AWS Step Functions workflow and add two states. Add the first state before the Lambda function. Configure the second state as a Wait state to periodically check whether the Athena query has finished using the Athena Boto3 get_query_execution API call. Configure the workflow to invoke the next query when the current query has finished running.
Lý do lựa chọn:
- Kết hợp Lambda + Step Functions là giải pháp serverless, rẻ nhất 🛠️:
- Lambda chỉ start_query_execution (chạy <1 giây, không bị timeout vì không chờ query hoàn thành).
- Step Functions orchestrate workflow với Wait state (hỗ trợ chờ vô hạn, poll get_query_execution mỗi 30s-1p), tự động chạy query tiếp theo khi hoàn thành.
- Chi phí thấp: Lambda ~$0.20/1M requests, Step Functions ~$0.025/1K state transitions (rẻ hơn Glue), Athena chỉ tính dữ liệu scan. Phù hợp chạy hàng ngày, scale tự động.
- Không cần quản lý server, tích hợp EventBridge cho schedule hàng ngày.
📋 Phân tích tất cả các phương án
Dưới đây là giải thích chi tiết từng lựa chọn. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ phân tích bằng tiếng Việt với đánh giá đúng/sai:
✅ Use an AWS Lambda function and the Athena Boto3 client start_query_execution API call to invoke the Athena queries programmatically.
Đúng 🟢: Lambda lý tưởng để khởi chạy query qua Boto3 (start_query_execution trả về query ID ngay lập tức). Lambda không cần chờ kết quả (tránh timeout 15p), chỉ invoke và return. Kết hợp với Step Functions để poll, đây là phần cốt lõi của giải pháp cost-effective.
✅ Create an AWS Step Functions workflow and add two states. Add the first state before the Lambda function. Configure the second state as a Wait state to periodically check whether the Athena query has finished using the Athena Boto3 get_query_execution API call. Configure the workflow to invoke the next query when the current query has finished running.
Đúng 🟢: Step Functions Express/Standard Workflow hoàn hảo cho orchestrate long-running tasks (hỗ trợ Wait state lên đến 1 năm). Poll get_query_execution định kỳ (mỗi 30s), retry tự động nếu fail. Workflow: Lambda start query → Wait poll → Lambda query tiếp theo. Rẻ hơn alternatives, tích hợp CloudWatch cho monitoring.
❌ Use an AWS Glue Python shell job and the Athena Boto3 client start_query_execution API call to invoke the Athena queries programmatically.
Sai 🔴: Glue Python Shell job (DPU-minimum 0.0625, timeout 5-60p tùy config) có thể start query nhưng đắt hơn nhiều (~$0.44/DPU-hour). Không cần thiết cho task đơn giản như start Athena (Lambda rẻ hơn 10x). Glue phù hợp ETL phức tạp với Spark, không cost-effective cho orchestrate query thuần.
❌ Use an AWS Glue Python shell script to run a sleep timer that checks every 5 minutes to determine whether the current Athena query has finished running successfully. Configure the Python shell script to invoke the next query when the current query has finished running.
Sai 🔴: Sleep timer mỗi 5 phút kém hiệu quả (waste thời gian chờ, tăng chi phí Glue). Glue Python Shell không tối ưu cho polling dài (có thể timeout nếu query >60p), và chi phí cao vì chạy liên tục DPU. Step Functions Wait state poll thông minh hơn, rẻ hơn, không lãng phí.
❌ Use Amazon Managed Workflows for Apache Airflow (Amazon MWAA) to orchestrate the Athena queries in AWS Batch.
Sai 🔴: MWAA (Airflow managed) + AWS Batch quá phức tạp và đắt (MWAA ~$0.49/hour/environment + Batch compute). MWAA không native cho Athena orchestrate đơn giản (cần DAG custom operators), Batch phù hợp batch container jobs chứ không phải query serverless. Không cost-effective so với Step Functions (rẻ hơn 5-10x cho workflow đơn giản).
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- AWS Athena Developer Guide: Boto3 Athena Client (start_query_execution & get_query_execution).
- AWS Step Functions: Orchestrating Athena Queries – Best practice cho long-running queries.
- AWS Well-Architected Data Analytics Lens: Khuyến nghị Lambda + Step Functions cho cost-effective orchestration (whitepaper 2025).
- Pricing Calculator: So sánh Lambda/Step Functions vs Glue/MWAA: AWS Pricing.
Giải pháp này đảm bảo serverless, scalable, rẻ nhất cho orchestrate Athena hàng ngày! 🚀
The company's current workloads use Apache Pig, Apache Oozie, Apache Spark, Apache Hbase, and Apache Flink. The on-premises workloads process petabytes of data in seconds. The company must maintain similar or better performance after the migration to AWS.
Which extract, transform, and load (ETL) service will meet these requirements?
- A AWS Glue
- B Amazon EMR
- C AWS Lambda
- D Amazon Redshift
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc chọn dịch vụ ETL (Extract, Transform, Load) phù hợp cho công ty đang di chuyển workloads từ on-premises sang AWS. Các yêu cầu chính bao gồm:
- Giảm operational overhead (giảm chi phí vận hành, quản lý hạ tầng thủ công).
- Khám phá serverless options (ưu tiên các lựa chọn không cần quản lý server).
- Workloads hiện tại sử dụng Apache Pig, Apache Oozie, Apache Spark, Apache Hbase, và Apache Flink – đây là các framework big data phổ biến cho xử lý dữ liệu lớn.
- Xử lý petabytes dữ liệu chỉ trong vài giây (yêu cầu hiệu suất cao, low-latency).
- Sau migration, phải duy trì performance tương đương hoặc tốt hơn.
📘 Bối cảnh AWS (cập nhật đến 2026): AWS cung cấp nhiều dịch vụ ETL/big data như EMR (hỗ trợ đầy đủ các Apache frameworks với managed clusters hoặc serverless), Glue (serverless ETL dựa trên Spark), Lambda (serverless compute nhỏ lẻ), và Redshift (data warehouse). Dịch vụ phải hỗ trợ chính xác các công cụ on-premises để tránh rewrite code lớn.
✅ Đáp án đúng: Amazon EMR
Lý do lựa chọn:
- Amazon EMR là dịch vụ quản lý Hadoop/Spark ecosystem đầy đủ, hỗ trợ native tất cả các frameworks trong câu hỏi: Apache Pig (cho scripting), Oozie (workflow scheduler), Spark (batch/streaming), Hbase (NoSQL database), và Flink (stream processing).
- Hiệu suất cao: Xử lý petabytes dữ liệu trong seconds nhờ spot instances, auto-scaling clusters, và EMR Serverless (ra mắt 2022, cập nhật 2026 hỗ trợ Spark/Flink serverless hoàn chỉnh) – đáp ứng low-latency và scale lớn.
- Giảm overhead: Managed service (không quản lý underlying Hadoop), serverless mode loại bỏ provision clusters thủ công.
- Không cần rewrite code lớn từ on-premises, migration mượt mà với EMR on EKS hoặc Graviton processors cho performance tốt hơn.
🛠️ Tài liệu tham khảo:
- AWS EMR Documentation (hỗ trợ Pig, Oozie, Spark, HBase, Flink).
- EMR Serverless (serverless cho big data, cập nhật 2026 với Flink improvements).
📋 Phân tích tất cả các phương án (đúng/sai)
-
AWS Glue ❌
Sai vì: AWS Glue là dịch vụ serverless ETL thuần túy dựa trên Spark (PySpark/Scala), hỗ trợ crawl data và job scheduling, nhưng KHÔNG hỗ trợ native Apache Pig, Oozie, Hbase, Flink. Glue phù hợp ETL nhẹ (terabytes), không xử lý petabytes trong seconds (latency cao hơn EMR). Không đáp ứng performance và frameworks cụ thể, dù giảm overhead tốt. -
Amazon EMR ✅
Đúng vì: Như giải thích trên, hỗ trợ đầy đủ tất cả frameworks (Pig/Oozie/Spark/Hbase/Flink), scale petabytes với transient clusters/auto-termination, EMR Serverless cho serverless true. Performance vượt trội nhờ optimized runtimes (Spark 3.x, Flink 1.18+ cập nhật 2026), giảm overhead 80% so on-premises. -
AWS Lambda ❌
Sai vì: Lambda là serverless compute cho functions nhỏ (15 phút timeout, 10GB RAM), KHÔNG thể xử lý petabytes dữ liệu hay các Apache frameworks phức tạp. Chỉ dùng cho micro-ETL, performance kém xa (không parallel big data), không phù hợp migration workloads lớn. -
Amazon Redshift ❌
Sai vì: Redshift là data warehouse columnar cho analytics/OLAP, KHÔNG phải ETL service hỗ trợ Pig/Oozie/Spark/Hbase/Flink. Dùng cho query petabytes nhưng latency cao (phút thay vì giây), overhead cao với cluster management, không serverless thuần và không migrate trực tiếp các frameworks on-premises.
🧩 Kết luận: Amazon EMR là lựa chọn tối ưu, cân bằng serverless + performance + compatibility. Nếu cần tư vấn migration cụ thể, hãy cung cấp thêm chi tiết workloads! 🚀
Which solution will meet this requirement with the LEAST operational effort?
- A Use an Amazon Kinesis Data Firehose delivery stream to process the dataset. Create an AWS Lambda transform function to identify the PII. Use an AWS SDK to obfuscate the PII. Set the S3 data lake as the target for the delivery stream.
- B Use the Detect PII transform in AWS Glue Studio to identify the PII. Obfuscate the PII. Use an AWS Step Functions state machine to orchestrate a data pipeline to ingest the data into the S3 data lake.
- C Use the Detect PII transform in AWS Glue Studio to identify the PII. Create a rule in AWS Glue Data Quality to obfuscate the PII. Use an AWS Step Functions state machine to orchestrate a data pipeline to ingest the data into the S3 data lake.
- D Ingest the dataset into Amazon DynamoDB. Create an AWS Lambda function to identify and obfuscate the PII in the DynamoDB table and to transform the data. Use the same Lambda function to ingest the data into the S3 data lake.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh việc một data engineer cần sử dụng các dịch vụ AWS để ingest (tiêu thụ dữ liệu) một dataset vào Amazon S3 data lake. Dataset được profile (phân tích cấu trúc và nội dung) và phát hiện chứa PII (Personally Identifiable Information - thông tin nhận dạng cá nhân). Yêu cầu chính là triển khai giải pháp để profile dataset và obfuscate (che giấu/mã hóa) PII, với tiêu chí LEAST operational effort (ít nỗ lực vận hành nhất).
🔍 Yêu cầu cốt lõi:
- Profile và detect PII: Xác định thông tin nhạy cảm như tên, email, số điện thoại...
- Obfuscate PII: Che giấu dữ liệu (ví dụ: thay thế bằng hash, mask như XXXX).
- Ingest vào S3 data lake: Đưa dữ liệu đã xử lý vào S3.
- Least effort: Ưu tiên giải pháp low-code/no-code, sử dụng built-in features của AWS để giảm custom code và quản lý.
📘 Kiến thức AWS cập nhật 2026: AWS Glue Studio (phiên bản mới nhất) hỗ trợ Detect PII transform (ra mắt từ 2023, cải tiến liên tục) để tự động detect PII mà không cần code. Kết hợp với Step Functions để orchestrate pipeline ETL tự động, scalable. Không cần dịch vụ trung gian phức tạp.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use the Detect PII transform in AWS Glue Studio to identify the PII. Obfuscate the PII. Use an AWS Step Functions state machine to orchestrate a data pipeline to ingest the data into the S3 data lake.
Lý do chọn (bằng tiếng Việt):
Giải pháp này sử dụng AWS Glue Studio với Detect PII transform built-in để detect PII một cách visual/no-code, sau đó obfuscate qua các transform đơn giản (như Replace hoặc Custom Transform với ít code). AWS Step Functions orchestrate toàn bộ pipeline (Glue job → S3), tự động retry/error handling. Đây là least operational effort vì:
- ✅ Low-code: Glue Studio drag-and-drop, không cần viết Lambda phức tạp.
- ✅ Scalable: Step Functions quản lý workflow serverless.
- ✅ Tích hợp S3 native: Ingest trực tiếp vào data lake.
🛠️ Dẫn nguồn: AWS Glue Detect PII Transform & Glue Studio Visual ETL (cập nhật 2026).
📋 Phân tích tất cả các phương án (đúng/sai)
-
Phương án 1: Use an Amazon Kinesis Data Firehose delivery stream to process the dataset. Create an AWS Lambda transform function to identify the PII. Use an AWS SDK to obfuscate the PII. Set the S3 data lake as the target for the delivery stream.
❌ Sai vì: Yêu cầu custom Lambda để detect/obfuscate PII (sử dụng AWS SDK như Comprehend hoặc regex), dẫn đến operational effort cao (viết code, test, monitor Lambda). Kinesis Firehose phù hợp streaming real-time nhưng không optimal cho batch dataset profiling. Không tận dụng built-in PII detection. -
Phương án 2 (ĐÚNG): Use the Detect PII transform in AWS Glue Studio to identify the PII. Obfuscate the PII. Use an AWS Step Functions state machine to orchestrate a data pipeline to ingest the data into the S3 data lake.
✅ Đúng vì: Như giải thích trên, Glue Studio Detect PII là built-in (hỗ trợ 50+ PII types, ML-powered), obfuscate dễ dàng qua transforms (e.g., masking). Step Functions orchestrate end-to-end với least effort (visual workflow, no server management). Phù hợp batch ingest vào S3. -
Phương án 3: Use the Detect PII transform in AWS Glue Studio to identify the PII. Create a rule in AWS Glue Data Quality to obfuscate the PII. Use an AWS Step Functions state machine to orchestrate a data pipeline to ingest the data into the S3 data lake.
❌ Sai vì: AWS Glue Data Quality chỉ dùng để check/validate data quality (rules như completeness, freshness), KHÔNG hỗ trợ obfuscate/transform data (chỉ recommend/drop bad records). Không thể dùng rule để thay đổi dữ liệu PII, dẫn đến giải pháp không khả thi. -
Phương án 4: Ingest the dataset into Amazon DynamoDB. Create an AWS Lambda function to identify and obfuscate the PII in the DynamoDB table and to transform the data. Use the same Lambda function to ingest the data into the S3 data lake.
❌ Sai vì: DynamoDB không phải data lake (là NoSQL DB, chi phí cao cho large dataset, không scalable như S3). Yêu cầu custom Lambda đa nhiệm (detect + obfuscate + export), tăng operational effort (provisioning, error-prone). Không trực tiếp ingest vào S3, vi phạm least effort.
🧩 Tóm tắt: Giải pháp đúng tận dụng Glue Studio + Step Functions cho ETL pipeline serverless, low-effort. Tránh custom code hoặc dịch vụ không phù hợp! 🚀
The company wants to improve the existing architecture to provide automated orchestration and to require minimal manual effort.
Which solution will meet these requirements with the LEAST operational overhead?
- A AWS Glue workflows
- B AWS Step Functions tasks
- C AWS Lambda functions
- D Amazon Managed Workflows for Apache Airflow (Amazon MWAA) workflows
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc cải thiện kiến trúc ETL (Extract, Transform, Load) cho một công ty đang sử dụng AWS Glue và Amazon EMR để xử lý dữ liệu từ các cơ sở dữ liệu vận hành (operational databases) vào data lake trên Amazon S3.
- Vấn đề hiện tại: Các workflow ETL hiện tại thiếu orchestration tự động (tự động điều phối các bước), dẫn đến cần nhiều nỗ lực thủ công để quản lý luồng công việc.
- Yêu cầu chính:
- Cung cấp orchestration tự động (tự động hóa việc sắp xếp, thực thi và xử lý lỗi giữa các bước ETL).
- Minimal manual effort (ít can thiệp thủ công nhất).
- LEAST operational overhead (chi phí vận hành thấp nhất, nghĩa là serverless, không cần quản lý hạ tầng, tích hợp sẵn với Glue và EMR).
Kiến trúc mong muốn phải tích hợp mượt mà với Glue (cho ETL jobs) và EMR (cho Spark/Hadoop processing), đồng thời tận dụng các dịch vụ AWS mới nhất đến năm 2026, nơi AWS Step Functions được ưu tiên cho orchestration serverless với hỗ trợ native cho distributed Map states và Express Workflows (cập nhật 2024-2025).
📘 Tài liệu tham khảo:
- AWS Step Functions Developer Guide: Integrations with AWS Glue và EMR và EMR Steps.
- AWS Well-Architected Framework - Data Lake: Orchestration pillar (2025 edition).
- Announcement: AWS Glue Workflows retirement (Dec 31, 2024) → Migrate to Step Functions.
✅ Đáp án đúng: AWS Step Functions tasks
Lý do lựa chọn:
- AWS Step Functions là dịch vụ serverless workflow orchestrator được thiết kế chuyên biệt để điều phối tự động các bước ETL giữa Glue jobs và EMR clusters/steps mà không cần quản lý hạ tầng (pay-per-use, no servers).
- ✅ Tích hợp native: Hỗ trợ trực tiếp Glue Job states và EMR AddStep/RunJobFlow actions, cho phép retry, branching, error handling tự động (ví dụ: Parallel/Choice states cho multi-ETL pipelines).
- ✅ Least operational overhead: Fully managed, visual workflow studio (2025 updates với AI-assisted design), scale tự động với Express Workflows (high-throughput ETL). Không cần DAG setup phức tạp như Airflow.
- ✅ Phù hợp 2026: Hỗ trợ distributed processing với Map states (cho fan-out Glue jobs) và integration với S3 EventBridge để trigger tự động từ data ingestion.
- So với các option khác, Step Functions giảm 90% manual effort theo AWS case studies (ví dụ: ETL pipelines cho data lakes).
🔍 Giải thích tất cả các phương án
-
❌ AWS Glue workflows
Phương án này sai vì AWS Glue Workflows sẽ bị retire hoàn toàn vào 31/12/2024 (theo thông báo AWS 2023, cập nhật 2025 yêu cầu migrate). Nó chỉ hỗ trợ orchestration cơ bản cho Glue jobs, không tích hợp tốt với EMR, thiếu advanced features như parallelism hay custom error handling. Overhead cao hơn do sắp deprecated và cần migrate thủ công. -
✅ AWS Step Functions tasks
Phương án này đúng như đã giải thích ở trên: Serverless, native integration với Glue/EMR/S3, automated orchestration với visual designer và zero infrastructure management – lý tưởng cho LEAST overhead trong data lake ETL (2026 best practice). -
❌ AWS Lambda functions
Phương án này sai vì Lambda chỉ là serverless compute cho individual functions, không có orchestration built-in (phải tự code retry/branching via SDK). Không hỗ trợ native Glue/EMR states, dẫn đến high manual effort (custom logic cho ETL coordination) và overhead lớn khi scale cho multi-step workflows. -
❌ Amazon Managed Workflows for Apache Airflow (Amazon MWAA) workflows
Phương án này sai dù là managed Airflow (DAG-based orchestration), nhưng operational overhead cao hơn vì cần setup Airflow environment (VPC, scaling groups), manage plugins/operators cho Glue/EMR, và monitoring riêng. Không serverless hoàn toàn như Step Functions; phù hợp hơn cho complex ML pipelines chứ không phải simple ETL với LEAST effort (AWS khuyến nghị Step Functions cho lightweight orchestration 2025+).
🛠️ Khuyến nghị triển khai: Sử dụng Step Functions với ASL (Amazon States Language) để define workflow: Start → Glue Crawler → Glue Job → EMR Step → S3 Sink, trigger via EventBridge. Test với AWS Console cho zero-downtime migration!
A data engineer examined data access patterns to identify trends. During the first 6 months, most data files are accessed several times each day. Between 6 months and 2 years, most data files are accessed once or twice each month. After 2 years, data files are accessed only once or twice each year.
The data engineer needs to use an S3 Lifecycle policy to develop new data storage rules. The new storage solution must continue to provide high availability.
Which solution will meet these requirements in the MOST cost-effective way?
- A Transition objects to S3 One Zone-Infrequent Access (S3 One Zone-IA) after 6 months. Transfer objects to S3 Glacier Flexible Retrieval after 2 years.
- B Transition objects to S3 Standard-Infrequent Access (S3 Standard-IA) after 6 months. Transfer objects to S3 Glacier Flexible Retrieval after 2 years.
- C Transition objects to S3 Standard-Infrequent Access (S3 Standard-IA) after 6 months. Transfer objects to S3 Glacier Deep Archive after 2 years.
- D Transition objects to S3 One Zone-Infrequent Access (S3 One Zone-IA) after 6 months. Transfer objects to S3 Glacier Deep Archive after 2 years.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc tối ưu hóa chi phí lưu trữ dữ liệu trên Amazon S3 bằng cách sử dụng S3 Lifecycle policy, dựa trên mô hình truy cập dữ liệu thực tế của một công ty.
- Tình huống hiện tại: Tất cả dữ liệu được lưu ở S3 Standard (lớp lưu trữ thường dùng cho dữ liệu truy cập thường xuyên, chi phí cao nhưng độ khả dụng cao 99.99% và độ bền 99.999999999%).
- Mô hình truy cập dữ liệu (dựa trên phân tích của data engineer):
- 0-6 tháng đầu: Truy cập nhiều lần mỗi ngày → Phù hợp với lớp lưu trữ frequent access (như S3 Standard).
- 6 tháng - 2 năm: Truy cập 1-2 lần/tháng → Infrequent access (ít truy cập, cần lớp rẻ hơn nhưng vẫn nhanh).
- Sau 2 năm: Truy cập 1-2 lần/năm → Rare access (rất ít, cần lớp lưu trữ archive giá rẻ nhất).
- Yêu cầu chính:
- Áp dụng S3 Lifecycle policy để tự động chuyển lớp lưu trữ theo thời gian.
- Tiếp tục đảm bảo high availability (độ khả dụng cao, tức là dữ liệu phải được sao chép qua nhiều Availability Zones - AZs, không chỉ một AZ duy nhất).
- MOST cost-effective (tiết kiệm chi phí nhất) mà vẫn đáp ứng mô hình truy cập.
Mục tiêu: Chuyển dữ liệu sang các lớp lưu trữ rẻ hơn theo giai đoạn, nhưng ưu tiên high availability (loại bỏ lớp chỉ dùng 1 AZ) và chọn lớp archive rẻ nhất cho dữ liệu ít truy cập nhất.
📘 Tài liệu tham khảo:
- AWS S3 Storage Classes: Amazon S3 Storage Classes (cập nhật 2024-2026, không thay đổi lớn).
- S3 Lifecycle Policies: Managing your storage lifecycle.
- So sánh chi phí: AWS Pricing Calculator và S3 Pricing (Deep Archive rẻ nhất ~$0.00099/GB/tháng, Flexible Retrieval ~$0.004/GB/tháng).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Transition objects to S3 Standard-Infrequent Access (S3 Standard-IA) after 6 months. Transfer objects to S3 Glacier Deep Archive after 2 years.
🛠️ Lý do chi tiết:
- Sau 6 tháng → Chuyển sang S3 Standard-IA: Phù hợp với truy cập infrequent (1-2 lần/tháng), chi phí lưu trữ rẻ hơn S3 Standard (~80% tiết kiệm), high availability (99.9%, sao chép qua 3 AZs), thời gian truy cập mili-giây, phí retrieval thấp. Minimum storage duration 30 ngày phù hợp.
- Sau 2 năm → Chuyển sang S3 Glacier Deep Archive: Rẻ nhất cho dữ liệu rarely accessed (1-2 lần/năm), chi phí lưu trữ chỉ ~$0.00099/GB/tháng, retrieval time 12 giờ (standard) hoặc nhanh hơn với phí cao. Vẫn đảm bảo high availability và độ bền 99.999999999%.
- Cost-effective nhất: Kết hợp IA đa AZ + archive rẻ nhất, tiết kiệm tối đa mà không hy sinh HA. Lifecycle policy hỗ trợ chuyển trực tiếp từ IA sang Deep Archive.
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ Phương án SAI: Transition objects to S3 One Zone-Infrequent Access (S3 One Zone-IA) after 6 months. Transfer objects to S3 Glacier Flexible Retrieval after 2 years.
🧩 Lý do sai: S3 One Zone-IA chỉ lưu ở 1 AZ duy nhất (không high availability, độ khả dụng 99%, dễ mất dữ liệu nếu AZ fail) → Vi phạm yêu cầu "high availability". Glacier Flexible Retrieval đắt hơn Deep Archive (~4x chi phí lưu trữ), không cost-effective cho truy cập 1-2 lần/năm (retrieval hours-days, phí cao hơn). -
❌ Phương án SAI: Transition objects to S3 Standard-Infrequent Access (S3 Standard-IA) after 6 months. Transfer objects to S3 Glacier Flexible Retrieval after 2 years.
🧩 Lý do sai: Phần sau 6 tháng đúng (Standard-IA đảm bảo high availability), nhưng Glacier Flexible Retrieval đắt hơn Deep Archive (chi phí lưu trữ cao gấp 4 lần, phù hợp hơn cho truy cập hàng tháng chứ không phải yearly). Không phải "MOST cost-effective". -
✅ Phương án ĐÚNG: Transition objects to S3 Standard-Infrequent Access (S3 Standard-IA) after 6 months. Transfer objects to S3 Glacier Deep Archive after 2 years.
🛠️ Lý do đúng: Như đã giải thích ở phần đáp án đúng. Hoàn hảo khớp mô hình access, high HA (Standard-IA multi-AZ, Deep Archive multi-AZ), và rẻ nhất tổng thể. -
❌ Phương án SAI: Transition objects to S3 One Zone-Infrequent Access (S3 One Zone-IA) after 6 months. Transfer objects to S3 Glacier Deep Archive after 2 years.
🧩 Lý do sai: Dù Deep Archive rẻ và phù hợp sau 2 năm, nhưng S3 One Zone-IA không high availability (single AZ) → Không đáp ứng yêu cầu chính. Rủi ro cao hơn Standard-IA mà tiết kiệm ít hơn.
Kết luận 💡: Giải pháp đúng cân bằng hoàn hảo giữa chi phí thấp, high availability và Lifecycle policy tự động. Trong thực tế DevOps, nên test với S3 Storage Lens để xác nhận pattern trước khi deploy! 🚀
The sales team recently requested access to the data that is in the ETL Redshift cluster so the team can perform weekly summary analysis tasks. The sales team needs to join data from the ETL cluster with data that is in the sales team's BI cluster.
The company needs a solution that will share the ETL cluster data with the sales team without interrupting the critical analysis tasks. The solution must minimize usage of the computing resources of the ETL cluster.
Which solution will meet these requirements?
- A Set up the sales team BI cluster as a consumer of the ETL cluster by using Redshift data sharing.
- B Create materialized views based on the sales team's requirements. Grant the sales team direct access to the ETL cluster.
- C Create database views based on the sales team's requirements. Grant the sales team direct access to the ETL cluster.
- D Unload a copy of the data from the ETL cluster to an Amazon S3 bucket every week. Create an Amazon Redshift Spectrum table based on the content of the ETL cluster.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh tình huống một công ty sử dụng hai Amazon Redshift provisioned cluster riêng biệt:
- ETL cluster: Dùng cho các hoạt động extract, transform, load (ETL) để hỗ trợ các nhiệm vụ phân tích quan trọng (critical analysis tasks). Cluster này cần được bảo vệ để tránh gián đoạn.
- BI cluster của sales team: Dùng cho các nhiệm vụ business intelligence (BI), và sales team cần truy cập dữ liệu từ ETL cluster để thực hiện weekly summary analysis tasks, cụ thể là join data giữa hai cluster.
Yêu cầu chính của giải pháp:
- Chia sẻ dữ liệu từ ETL cluster cho sales team mà không làm gián đoạn critical analysis tasks.
- Tối thiểu hóa sử dụng tài nguyên compute của ETL cluster (producer cluster).
- Giải pháp phải hiệu quả, an toàn và tuân thủ best practices AWS Redshift mới nhất (cập nhật đến 2026, với Redshift ra mắt nhiều cải tiến như serverless và data sharing nâng cao).
Vấn đề cốt lõi: Cần zero-ETL data sharing cross-cluster, tránh copy data thủ công hoặc tăng tải compute trên ETL cluster.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Set up the sales team BI cluster as a consumer of the ETL cluster by using Redshift data sharing.
Lý do:
- Redshift Data Sharing (tính năng cross-account/cluster data sharing) cho phép chia sẻ dữ liệu zero-copy từ producer cluster (ETL) sang consumer cluster (BI sales) mà không sao chép dữ liệu, không tăng tải compute trên producer.
- Sales team có thể query và join data trực tiếp từ ETL cluster như dữ liệu local, hỗ trợ weekly analysis mà không ảnh hưởng critical tasks.
- Hoàn toàn phù hợp yêu cầu: Minimize compute resources (chỉ metadata được chia sẻ, compute chạy trên consumer), và không gián đoạn ETL cluster.
- Tính năng này hỗ trợ across regions/accounts, với governance qua namespaces/databases, và được khuyến nghị chính thức cho scenarios như thế này (cập nhật 2026 với tích hợp Redshift Serverless).
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên việc có đáp ứng không gián đoạn critical tasks và minimize compute ETL cluster hay không.
-
Set up the sales team BI cluster as a consumer of the ETL cluster by using Redshift data sharing.
✅ Đúng. Như đã giải thích ở trên, đây là giải pháp zero-copy, low-impact lý tưởng. Producer (ETL) chỉ share metadata, consumer (BI) chịu toàn bộ compute cho join/query. Không cần unload/copy data, hỗ trợ real-time sharing. -
Create materialized views based on the sales team's requirements. Grant the sales team direct access to the ETL cluster.
❌ Sai. Materialized views yêu cầu refresh định kỳ (manual hoặc auto), tăng compute/load cao trên ETL cluster (refresh có thể chiếm 20-50% resources tùy size data). Direct access từ sales team còn làm gián đoạn critical tasks do concurrent queries từ BI workload. -
Create database views based on the sales team's requirements. Grant the sales team direct access to the ETL cluster.
❌ Sai. Database views chỉ là logical abstraction (không lưu data), nhưng direct access vẫn yêu cầu sales team query trực tiếp ETL cluster → tăng compute/load từ BI queries, dễ gián đoạn critical ETL tasks. Không minimize resources như data sharing. -
Unload a copy of the data from the ETL cluster to an Amazon S3 bucket every week. Create an Amazon Redshift Spectrum table based on the content of the ETL cluster.
❌ Sai. Unload weekly → tăng compute ETL (UNLOAD command tốn resources), copy data sang S3 (storage cost + latency), rồi Spectrum query từ BI cluster. Không real-time (chỉ weekly), không join seamless như data sharing, và vẫn gián đoạn do unload process.
🛠️ Khuyến nghị triển khai thực tế
- Bước 1: Producer (ETL owner) authorize consumer (sales BI) qua Redshift console/CLI/API.
- Bước 2: Create datashare trên ETL namespace → Consumer tạo database từ share.
- Bước 3: Grant permissions chi tiết (column-level) để an toàn.
- Monitoring: Sử dụng Redshift Query Editor V2 hoặc CloudWatch để track usage (consumer-side compute).
📘 Tài liệu tham khảo (AWS chính thức, cập nhật 2026)
- Amazon Redshift Data Sharing Documentation – Chi tiết zero-copy sharing.
- Redshift Best Practices for Cross-Cluster Sharing – Blog AWS 2023+, vẫn valid.
- AWS Redshift Exam Guide DOP-C02 – Topic Redshift trong DevOps Professional.
Giải pháp này đảm bảo tuân thủ AWS Well-Architected Framework (Reliability & Cost Optimization)! 🚀
Which solution will meet this requirement MOST cost-effectively?
- A Use an Amazon EMR provisioned cluster to read from all sources. Use Apache Spark to join the data and perform the analysis.
- B Copy the data from DynamoDB, Amazon RDS, and Amazon Redshift into Amazon S3. Run Amazon Athena queries directly on the S3 files.
- C Use Amazon Athena Federated Query to join the data from all data sources.
- D Use Redshift Spectrum to query data from DynamoDB, Amazon RDS, and Amazon S3 directly from Redshift.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào một data engineer cần join dữ liệu từ nhiều nguồn khác nhau (Amazon DynamoDB, Amazon RDS, Amazon Redshift, và Amazon S3) để thực hiện một công việc phân tích một lần duy nhất (one-time analysis job). Yêu cầu chính là chọn giải pháp tiết kiệm chi phí nhất (MOST cost-effectively).
🛠️ Các yếu tố quan trọng cần xem xét:
- Đây là job một lần, nên ưu tiên giải pháp serverless (không cần quản lý cluster), pay-per-use (chỉ trả tiền cho query thực tế), và không cần copy dữ liệu (tránh chi phí export/import lớn).
- Nguồn dữ liệu đa dạng: DynamoDB (NoSQL), RDS (relational DB), Redshift (data warehouse), S3 (object storage).
- Kiến thức cập nhật đến 2026: Amazon Athena Federated Query (ra mắt từ 2020 và được cải tiến liên tục) hỗ trợ query liên kết (federated join) trực tiếp từ các nguồn này mà không di chuyển data, với connectors cho DynamoDB, RDS (qua Lambda), Redshift, và S3 native.
📘 Tài liệu tham khảo:
- AWS Athena Federated Query: docs.aws.amazon.com/athena/latest/ug/querying-federated-data-sources.html (cập nhật 2024-2026, hỗ trợ 20+ connectors).
- AWS Well-Architected Framework - Cost Optimization Pillar: Nhấn mạnh serverless cho one-time workloads.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Amazon Athena Federated Query to join the data from all data sources.
Lý do 🏆:
- Athena Federated Query cho phép query và join dữ liệu trực tiếp từ tất cả các nguồn (DynamoDB, RDS, Redshift, S3) mà không cần copy data, sử dụng connectors (như Lambda cho RDS/DynamoDB/Redshift và native cho S3).
- Tiết kiệm chi phí nhất cho one-time job: Serverless, pay-per-TB scanned (khoảng $5/TB), không tốn phí cluster hay storage di chuyển. Thời gian setup nhanh (chỉ deploy connector).
- Hỗ trợ SQL chuẩn (Presto/Trino engine), dễ join cross-source.
📋 Giải thích tất cả các phương án
-
Use an Amazon EMR provisioned cluster to read from all sources. Use Apache Spark to join the data and perform the analysis.
❌ Sai vì không tiết kiệm chi phí: EMR provisioned cluster yêu cầu provision và quản lý EC2 instances (tối thiểu 1 giờ/billing), tốn phí idle time dù chỉ one-time job. Spark mạnh cho big data nhưng overkill và đắt hơn (khoảng $0.07-$5/giờ/node tùy loại), không serverless. -
Copy the data from DynamoDB, Amazon RDS, and Amazon Redshift into Amazon S3. Run Amazon Athena queries directly on the S3 files.
❌ Sai vì tốn kém và phức tạp: Việc copy data từ DynamoDB (export via DMS/Glue), RDS (dump/export), Redshift (UNLOAD) sang S3 phát sinh chi phí lớn (storage + transfer + thời gian ETL). Athena trên S3 chỉ query file-based (Parquet/CSV), không hiệu quả cho relational join trực tiếp, và one-time job không đáng copy hàng TB data. -
Use Amazon Athena Federated Query to join the data from all data sources.
✅ Đúng như đã giải thích ở trên: Giải pháp tối ưu nhất về chi phí và hiệu suất cho one-time analysis, federated pushdown query giảm data movement. -
Use Redshift Spectrum to query data from DynamoDB, Amazon RDS, and Amazon S3 directly from Redshift.
❌ Sai vì không khả thi và tốn kém: Redshift Spectrum chỉ query dữ liệu external trên S3 (file-based như Parquet), không hỗ trợ trực tiếp DynamoDB/RDS (cần ETL trước). Yêu cầu Redshift cluster đang chạy (tốn $0.25-$13/giờ/node), không serverless, kém hiệu quả cho one-time job.
🧠 Kết luận: Athena Federated Query là lựa chọn serverless, no-ETL, cost-optimized lý tưởng cho workload này theo best practices AWS 2026! 🚀
Which combination of resources will meet these requirements MOST cost-effectively? (Choose two.)
- A Use Hadoop Distributed File System (HDFS) as a persistent data store.
- B Use Amazon S3 as a persistent data store.
- C Use x86-based instances for core nodes and task nodes.
- D Use Graviton instances for core nodes and task nodes.
- E Use Spot Instances for all primary nodes.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc lập kế hoạch sử dụng một EMR cluster provisioned (cụm được cung cấp sẵn) chạy các job Apache Spark để phân tích big data. Công ty yêu cầu độ tin cậy cao (high reliability), đội ngũ big data phải tuân thủ best practices cho workload cost-optimized (tối ưu chi phí) và long-running (chạy lâu dài) trên Amazon EMR. Giải pháp phải duy trì hiệu suất hiện tại (maintain current performance level).
📌 Yêu cầu chọn TWO (2) lựa chọn kết hợp tài nguyên MOST cost-effectively (tiết kiệm chi phí nhất).
- Provisioned EMR cluster: Không phải serverless, nên cần tối ưu instance types và storage.
- Long-running workloads: Cần persistent storage bền vững, tránh mất dữ liệu khi cluster terminate.
- High reliability + cost-optimized: Ưu tiên giải pháp ổn định, giá rẻ hơn mà không giảm performance (ví dụ: Graviton instances cho Spark hiệu suất tốt hơn 40% giá/trình suất).
- Cập nhật AWS 2026: EMR phiên bản 7.x+ hỗ trợ Graviton4 (Graviton instances mới nhất), S3 với S3 Express One Zone cho EMR I/O nhanh, best practices khuyến nghị S3 thay HDFS và Graviton cho workloads Spark.
✅ Đáp án đúng (Chọn TWO)
Đáp án đúng là:
- Use Amazon S3 as a persistent data store. ✅
- Use Graviton instances for core nodes and task nodes. ✅
Lý do lựa chọn (kết hợp MOST cost-effectively): 🛠️ S3 + Graviton là best practices AWS cho EMR long-running Spark:
- S3 làm persistent store rẻ hơn HDFS (không tốn chi phí lưu trữ local), bền vững qua cluster lifecycle, tích hợp EMRFS cho consistency cao → high reliability mà cost-optimized (tiết kiệm 70-80% so HDFS).
- Graviton (Arm-based như r8g/m8g) cho core/task nodes: Giá rẻ hơn 40% so x86, hiệu suất Spark cao hơn (do native Arm optimization), hỗ trợ full EMR apps đến 2026 → maintain performance mà tiết kiệm lớn cho long-running. Kết hợp này giảm TCO (Total Cost of Ownership) lên đến 50% cho workloads tương tự, theo AWS benchmarks.
📋 Phân tích TẤT CẢ các phương án (Đúng/Sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên văn bản gốc tiếng Anh, chỉ giải thích bằng tiếng Việt với lý do đúng/sai dựa trên best practices EMR cập nhật nhất.
-
❌ Use Hadoop Distributed File System (HDFS) as a persistent data store.
Sai vì: HDFS chỉ lưu trữ local trên instance (ephemeral), dữ liệu mất khi cluster terminate/restart → không high reliability cho long-running. Best practices AWS khuyến nghị tránh HDFS làm persistent store, dùng S3 thay thế để bền vững và cost-optimized (HDFS tốn chi phí instance storage cao). -
✅ Use Amazon S3 as a persistent data store.
Đúng vì: S3 là persistent object storage rẻ (khoảng 0.023$/GB/tháng), tích hợp EMRFS cho read/write nhanh như HDFS mà không mất dữ liệu. Hoàn hảo cho Spark jobs long-running, high reliability (99.999999999% durability), và cost-optimized (pay-per-use, no provisioning). AWS docs: "Use S3 for input/output data". -
❌ Use x86-based instances for core nodes and task nodes.
Sai vì: x86 (như m5/r5) là standard nhưng đắt hơn Graviton 40% cho cùng performance. Với Spark workloads, Graviton (r7g/r8g) nhanh hơn do Arm optimization, best practices 2026 ưu tiên Graviton cho EMR để cost-optimized mà maintain performance. -
✅ Use Graviton instances for core nodes and task nodes.
Đúng vì: Graviton2/4 (r7g/m7g) cho core/task nodes tiết kiệm chi phí cao (up to 40% rẻ hơn x86), hiệu suất Spark tốt hơn 30-50% (theo AWS benchmarks EMR 7.x). High reliability (stable availability), phù hợp long-running, AWS khuyến nghị chính thức cho big data analysis. -
❌ Use Spot Instances for all primary nodes.
Sai vì: Primary nodes (master + core) cần high reliability, Spot Instances dễ bị interrupt (giá rẻ nhưng không ổn định) → rủi ro downtime cho Spark jobs. Best practices: Chỉ dùng Spot cho task nodes (non-critical), master/core phải On-Demand/Reserved cho long-running.
📘 Tài liệu tham khảo (AWS Official - Cập nhật 2026)
- Amazon EMR Best Practices: Khuyến nghị S3 + Graviton.
- EMR Graviton Instances: Benchmarks Spark performance.
- EMR Cluster Storage: S3 vs HDFS.
- Spot on EMR: Chỉ task nodes.
- AWS Well-Architected Framework: Big Data Lens (2026 edition) – Cost Optimization pillar.
Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm câu hỏi, cứ hỏi nhé!