Ngân hàng đề — AWS Certified Data Engineer Associate
Tìm thấy 867 câu.
Which solution will meet these requirements with the LEAST operational overhead?
- A Use Kinesis Data Streams to stage data in Amazon S3. Use the COPY command to load data from Amazon S3 directly into Amazon Redshift to make the data immediately available for real-time analysis.
- B Access the data from Kinesis Data Streams by using SQL queries. Create materialized views directly on top of the stream. Refresh the materialized views regularly to query the most recent stream data.
- C Create an external schema in Amazon Redshift to map the data from Kinesis Data Streams to an Amazon Redshift object. Create a materialized view to read data from the stream. Set the materialized view to auto refresh.
- D Connect Kinesis Data Streams to Amazon Kinesis Data Firehose. Use Kinesis Data Firehose to stage the data in Amazon S3. Use the COPY command to load the data from Amazon S3 to a table in Amazon Redshift.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi mô tả một công ty muốn triển khai khả năng phân tích thời gian thực (real-time analytics) với lượng dữ liệu streaming lớn (vài gigabytes/giây). Họ sử dụng Amazon Kinesis Data Streams để thu thập dữ liệu streaming và Amazon Redshift để xử lý, lưu trữ. Mục tiêu là tạo insights gần thời gian thực (near real-time) bằng các công cụ BI và analytics hiện có.
Yêu cầu chính: Giải pháp nào đáp ứng với ít overhead vận hành nhất (LEAST operational overhead)?
- Nghĩa là ưu tiên giải pháp tự động hóa cao, ít can thiệp thủ công, không cần ETL phức tạp, và hỗ trợ query trực tiếp từ BI tools lên dữ liệu streaming mà không làm gián đoạn hiệu suất Redshift.
- Thách thức: Kinesis Data Streams xử lý high-throughput streaming, Redshift là data warehouse columnar, cần tích hợp mượt mà để tránh latency cao hoặc quản lý nhiều service riêng lẻ. ✅ Giải pháp lý tưởng phải tận dụng tính năng native mới nhất của AWS (cập nhật đến 2026), như Redshift Streaming Ingestion cho phép materialized views tự động refresh từ Kinesis Streams.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create an external schema in Amazon Redshift to map the data from Kinesis Data Streams to an Amazon Redshift object. Create a materialized view to read data from the stream. Set the materialized view to auto refresh.
Lý do:
- Đây là tính năng Redshift Streaming (ra mắt 2022, cập nhật liên tục đến 2026) cho phép Redshift trực tiếp ingest dữ liệu từ Kinesis Data Streams qua external schema và materialized view với auto-refresh (tự động refresh mỗi 1 phút hoặc theo cấu hình).
- Overhead thấp nhất: Không cần staging S3, không ETL thủ công, query BI tools (như Tableau, QuickSight) trực tiếp trên MV như bảng thông thường. Hỗ trợ scale đến hàng GB/s mà không ảnh hưởng cluster Redshift.
- Đáp ứng near real-time: Dữ liệu sẵn sàng query ngay sau refresh, latency thấp (~1 phút). 🛠️ Ưu điểm: Tích hợp native, tự động hóa cao, chi phí thấp, dễ quản lý.
📋 Phân tích tất cả các phương án
-
❌ Use Kinesis Data Streams to stage data in Amazon S3. Use the COPY command to load data from Amazon S3 directly into Amazon Redshift to make the data immediately available for real-time analysis.
Sai vì: Không phải real-time thực sự – phải stage dữ liệu vào S3 trước (cần consumer riêng để write S3), rồi chạy COPY command định kỳ. Latency cao (phút đến giờ), overhead lớn (quản lý shard, consumer, schedule COPY). Không tận dụng streaming native, dễ miss dữ liệu nếu COPY fail. Không "immediately available". -
❌ Access the data from Kinesis Data Streams by using SQL queries. Create materialized views directly on top of the stream. Refresh the materialized views regularly to query the most recent stream data.
Sai vì: Redshift KHÔNG hỗ trợ materialized views trực tiếp trên Kinesis Streams mà không qua external schema. Không có SQL query native trực tiếp từ Redshift lên Kinesis (cần Lambda/consumer trung gian). Refresh "regularly" thủ công → overhead cao, không auto, không scale GB/s. -
✅ Create an external schema in Amazon Redshift to map the data from Kinesis Data Streams to an Amazon Redshift object. Create a materialized view to read data from the stream. Set the materialized view to auto refresh.
Đúng vì: Sử dụng Redshift external schema (Fedration + Streaming) để map Kinesis stream như object Redshift. Materialized view auto-refresh (mặc định 1 phút, configurable) ingest dữ liệu liên tục. Overhead thấp: Native, serverless, hỗ trợ BI tools query trực tiếp. Scale tự động theo Kinesis throughput, cập nhật AWS 2026 vẫn là best practice cho near real-time. -
❌ Connect Kinesis Data Streams to Amazon Kinesis Data Firehose. Use Kinesis Data Firehose to stage the data in Amazon S3. Use the COPY command to load the data from Amazon S3 to a table in Amazon Redshift.
Sai vì: Tương tự phương án 1, dùng Firehose buffer vào S3 (batch 1-15 phút) rồi COPY → latency không near real-time (5-60 phút). Overhead cao: Quản lý Firehose delivery stream, S3 lifecycle, COPY schedule. Không tận dụng streaming trực tiếp, kém hiệu quả hơn native Redshift Streaming.
📘 Tài liệu tham khảo
- AWS Documentation: Amazon Redshift Streaming Ingestion for Amazon Kinesis Data Streams (cập nhật 2025-2026: Hỗ trợ auto-refresh, multi-stream, RA3 nodes).
- AWS Blog: Near real-time Analytics with Redshift and Kinesis (2022, vẫn valid 2026).
- Exam Prep: AWS Certified DevOps Engineer Professional DOP-C02 blueprint – Domain 4: Automation (Streaming & Analytics).
- Console Guide: Redshift > Query Editor > CREATE EXTERNAL SCHEMA / MATERIALIZED VIEW STREAMING.
🛠️ Lời khuyên: Trong thực tế, test với CREATE MATERIALIZED VIEW mv_stream AUTO REFRESH YES trên Redshift RA3 cluster để verify GB/s throughput!
A data engineer discovers that dashboard queries are becoming slower over time. The data engineer determines that the root cause of the slowing queries is long-running AWS Glue jobs.
Which actions should the data engineer take to improve the performance of the AWS Glue jobs? (Choose two.)
- A Partition the data that is in the S3 bucket. Organize the data by year, month, and day.
- B Increase the AWS Glue instance size by scaling up the worker type.
- C Convert the AWS Glue schema to the DynamicFrame schema class.
- D Adjust AWS Glue job scheduling frequency so the jobs run half as many times each day.
- E Modify the IAM role that grants access to AWS glue to grant access to all S3 features.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi mô tả một công ty sử dụng Amazon QuickSight để theo dõi dashboard ứng dụng, với dữ liệu được xử lý bởi AWS Glue jobs và lưu trữ trong một bucket Amazon S3 duy nhất. Dữ liệu mới được thêm hàng ngày, dẫn đến tình trạng queries trên dashboard chậm dần theo thời gian. Kỹ sư dữ liệu xác định nguyên nhân gốc rễ là AWS Glue jobs chạy lâu hơn.
Nhiệm vụ: Chọn hai hành động để cải thiện hiệu suất của AWS Glue jobs.
Vấn đề cốt lõi là dữ liệu tích tụ không được tối ưu hóa, khiến Glue phải quét toàn bộ dữ liệu lớn trong S3, dẫn đến job chậm (theo nguyên tắc partition pruning và scan optimization trong AWS Glue phiên bản mới nhất 2024-2026). QuickSight phụ thuộc vào dữ liệu sạch từ Glue, nên tối ưu Glue là chìa khóa. 🛠️
✅ Đáp án đúng (Chọn TWO)
- Partition the data that is in the S3 bucket. Organize the data by year, month, and day.
- Increase the AWS Glue instance size by scaling up the worker type.
Lý do chọn:
Hai hành động này trực tiếp giải quyết vấn đề dữ liệu lớn tích tụ hàng ngày, giúp giảm thời gian quét dữ liệu và tăng công suất xử lý. Partitioning giảm lượng dữ liệu scan (lên đến 90% theo best practices AWS), còn scale up worker tăng DPUs (Data Processing Units) để xử lý song song nhanh hơn. Đây là các giải pháp chuẩn trong AWS Glue Performance Tuning Guide (cập nhật 2025). 📈
📋 Phân tích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, với nội dung gốc giữ nguyên tiếng Anh. Tôi đánh dấu ✅ cho đúng, ❌ cho sai, và giải thích rõ ràng dựa trên kiến thức AWS mới nhất (Glue 4.0+, Spark 3.3+).
-
✅ Partition the data that is in the S3 bucket. Organize the data by year, month, and day.
Giải thích đúng: Phân vùng (partition) dữ liệu theo năm/tháng/ngày (ví dụ:s3://bucket/year=2026/month=01/day=15/) kích hoạt partition pruning trong Glue ETL, giúp chỉ quét dữ liệu cần thiết thay vì toàn bộ bucket. Với dữ liệu thêm hàng ngày, điều này giảm thời gian job từ hàng giờ xuống phút, đặc biệt hiệu quả cho QuickSight queries. Theo AWS docs 2025, partitioning là bước đầu tiên trong performance optimization cho S3 data lakes. 🏆 -
✅ Increase the AWS Glue instance size by scaling up the worker type.
Giải thích đúng: Tăng kích thước instance bằng cách scale up worker type (từ Standard → G.1X, G.2X, hoặc G.025X với Glue 4.0) tăng số DPUs (ví dụ: G.2X = 4.5 DPU/worker), cải thiện xử lý song song và memory cho dữ liệu lớn. Lý tưởng cho job chậm do volume tăng dần. AWS khuyến nghị scale vertically trước khi horizontally (2026 best practices). 🚀 -
❌ Convert the AWS Glue schema to the DynamicFrame schema class.
Giải thích sai: DynamicFrame là lớp mặc định trong AWS Glue (dựa trên Spark DataFrame nhưng schema-on-read linh hoạt). Không cần "convert schema to DynamicFrame" vì Glue jobs đã dùng nó; việc này không cải thiện performance mà chỉ liên quan đến schema evolution. Thay vào đó, dùngfrom_cataloghoặcpush_down_predicateđể tối ưu. Không giải quyết root cause scan chậm. 🤏 -
❌ Adjust AWS Glue job scheduling frequency so the jobs run half as many times each day.
Giải thích sai: Giảm tần suất chạy (ví dụ: từ 24 lần/ngày xuống 12) chỉ giảm tổng workload, không cải thiện thời gian chạy của từng job (vẫn chậm do scan full S3). Dashboard QuickSight cần dữ liệu tươi mới hàng ngày, nên giảm frequency có thể làm dữ liệu cũ hơn, tệ hơn cho user experience. Không phải giải pháp performance tuning. ⏰ -
❌ Modify the IAM role that grants access to AWS glue to grant access to all S3 features.
Giải thích sai: IAM role chỉ ảnh hưởng đến quyền truy cập, không liên quan đến hiệu suất job (Glue đã có quyền s3:GetObject/ListBucket cơ bản). "All S3 features" có thể gây over-privileged (vi phạm least privilege), nhưng không giảm thời gian scan dữ liệu lớn. Vấn đề là dữ liệu structure, không phải permission. 🔒
📘 Tài liệu tham khảo
- AWS Glue Developer Guide - Performance Tuning (2025): https://docs.aws.amazon.com/glue/latest/dg/monitor-performance.html (Partitioning & Worker types).
- Amazon S3 Best Practices for ETL: https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-partitions.html.
- QuickSight + Glue Integration: https://docs.aws.amazon.com/quicksight/latest/user/connecting-to-datasets-glue.html.
- AWS Well-Architected Framework - Data Analytics Lens (2026): Nhấn mạnh partitioning cho data lakes.
Tất cả dựa trên cập nhật AWS re:Invent 2025 và Glue 4.0. Nếu cần lab thực hành, dùng AWS Glue Studio! 🌟
Which Step Functions state should the data engineer use to meet these requirements?
- A Parallel state
- B Choice state
- C Map state
- D Wait state
Xem giải thích
🧩 Phân tích chi tiết câu hỏi
Câu hỏi tập trung vào AWS Step Functions, một dịch vụ orchestration serverless dùng để xây dựng workflow phức tạp, đáng tin cậy. Một data engineer cần thiết kế workflow để parallel process (xử lý song song) một large collection of data files (bộ sưu tập lớn các file dữ liệu) và apply a specific transformation (áp dụng một phép biến đổi cụ thể) cho mỗi file.
Yêu cầu chính: Workflow phải lặp qua từng file một cách song song (không tuần tự), tận dụng tính năng iteration để xử lý hàng loạt items từ input array, đảm bảo scalability cho dữ liệu lớn. Điều này phù hợp với các scenario ETL (Extract-Transform-Load) hoặc data pipeline trên AWS. 📘
Đáp án đúng: Map state ✅
Lý do chọn:
Map state là state chuyên dụng để iterate song song qua một mảng input (array), chạy một sub-workflow (Iterator) cho mỗi item một cách parallel. Nó lý tưởng cho việc xử lý large collection files, vì:
- Hỗ trợ MaxConcurrency để kiểm soát số lượng parallel executions (mặc định không giới hạn, nhưng có thể set để tránh throttling).
- ResultPath để thu thập output từ tất cả iterations thành mảng mới.
- Tích hợp ItemProcessor (từ phiên bản 2020+) cho optimized processing với Lambda hoặc nested workflows.
Điều này khớp hoàn hảo với yêu cầu parallel transformation từng file, giúp scale tự động mà không cần code thủ công. Theo docs AWS mới nhất (2026), Map state là best practice cho dynamic parallelism trong data workflows. 🛠️
Tài liệu tham khảo:
❌ Giải thích tất cả các phương án (đúng/sai)
-
Parallel state ❌
Sai vì: Parallel state chỉ chạy nhiều branches fixed (cố định) song song theo cấu trúc static (như 3 tasks A, B, C cùng lúc), không hỗ trợ iteration qua array động như collection files. Nó không lặp qua từng item riêng lẻ, nên không phù thể large collection cần process từng file. -
Choice state ❌
Sai vì: Choice state dùng cho conditional branching (phân nhánh dựa điều kiện), như if-else để chọn path dựa rules (ví dụ: file size > 1GB thì path A). Nó không hỗ trợ parallel processing hay iteration, chỉ sequential decision-making. -
Map state ✅
(Đã giải thích chi tiết ở trên - lựa chọn tối ưu cho parallel iteration). -
Wait state ❌
Sai vì: Wait state chỉ delay execution (chờ fixed time, timestamp hoặc until signal), dùng cho timing control như rate limiting. Nó không xử lý parallel hay transformation, chỉ pause workflow.
The data engineer must identify and remove duplicate information from the legacy application data.
Which solution will meet these requirements with the LEAST operational overhead?
- A Write a custom extract, transform, and load (ETL) job in Python. Use the DataFrame.drop_duplicates() function by importing the Pandas library to perform data deduplication.
- B Write an AWS Glue extract, transform, and load (ETL) job. Use the FindMatches machine learning (ML) transform to transform the data to perform data deduplication.
- C Write a custom extract, transform, and load (ETL) job in Python. Import the Python dedupe library. Use the dedupe library to perform data deduplication.
- D Write an AWS Glue extract, transform, and load (ETL) job. Import the Python dedupe library. Use the dedupe library to perform data deduplication.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc migrate một ứng dụng legacy sang Amazon S3 data lake, nơi dữ liệu legacy chứa thông tin trùng lặp (duplicate information). Nhiệm vụ của data engineer là xác định và loại bỏ duplicate từ dữ liệu này, với yêu cầu LEAST operational overhead (ít nhất công sức vận hành, quản lý).
🔍 Bối cảnh chính:
- Dữ liệu đang được di chuyển vào S3 data lake – một kiến trúc lưu trữ dữ liệu lớn, phân tích trên AWS.
- Vấn đề: Duplicate data có thể làm giảm hiệu suất phân tích, tăng chi phí lưu trữ và xử lý.
- Mục tiêu: Giải pháp tối ưu hóa vận hành, nghĩa là ưu tiên dịch vụ serverless/managed, không cần code custom phức tạp, tự động scale, và ít bảo trì (theo best practices AWS DevOps đến 2026).
🛠️ Yêu cầu cốt lõi: Sử dụng công cụ ETL (Extract, Transform, Load) để deduplication, nhưng phải ít overhead nhất – tránh custom code, thư viện bên thứ ba, ưu tiên tính năng built-in ML của AWS.
✅ Đáp án đúng
Write an AWS Glue extract, transform, and load (ETL) job. Use the FindMatches machine learning (ML) transform to transform the data to perform data deduplication.
Lý do chọn đáp án này:
- AWS Glue FindMatches là ML transform built-in (từ AWS Glue 3.0+, cập nhật đến 2026) chuyên dùng để deduplication và record linkage trên dữ liệu lớn trong data lake. Nó sử dụng ML tự động (dựa trên transformer models) để phát hiện duplicate gần giống (fuzzy matching), không cần code custom.
- Least operational overhead: Glue là serverless ETL, tự động scale, integrate trực tiếp với S3, không cần quản lý server/infra. Chỉ cần define job qua console/CLI, ML train tự động trên sample data.
- Phù hợp migrate legacy data: Xử lý terabytes dữ liệu duplicate mà không cần expertise ML sâu.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên operational overhead (code custom, thư viện, quản lý infra) và hiệu quả deduplication trên S3 data lake (kiến thức AWS Glue/DataBrew/ML mới nhất 2026).
-
❌ [SAI] Write a custom extract, transform, and load (ETL) job in Python. Use the DataFrame.drop_duplicates() function by importing the Pandas library to perform data deduplication.
Phương án này yêu cầu code custom Python + Pandas, chạy trên EC2/EMR/Lambda. Overhead cao: Pandas chỉ exact matching (không fuzzy), không scale tốt cho data lake lớn (memory-intensive), cần tự quản lý job scheduling, error handling, scaling. Không phải giải pháp managed, vi phạm "least overhead". -
✅ [ĐÚNG] Write an AWS Glue extract, transform, and load (ETL) job. Use the FindMatches machine learning (ML) transform to transform the data to perform data deduplication.
Như đã giải thích ở trên: Built-in ML transform của Glue, serverless, fuzzy dedup tự động. Overhead thấp nhất: No custom code, integrate S3 native, ML labeling chỉ cần vài phút. -
❌ [SAI] Write a custom extract, transform, and load (ETL) job in Python. Import the Python dedupe library. Use the dedupe library to perform data deduplication.
Sử dụng thư viện dedupe (third-party) trong custom Python ETL. Overhead cao: Cần train model thủ công, active learning loop, chạy trên infra tự quản (EC2/EMR), không serverless. Dedupe tốt cho fuzzy matching nhưng không integrate native với S3/Glue, tăng complexity migrate. -
❌ [SAI] Write an AWS Glue extract, transform, and load (ETL) job. Import the Python dedupe library. Use the dedupe library to perform data deduplication.
Kết hợp Glue + custom dedupe library. Overhead trung bình-cao: Glue managed tốt nhưng import third-party library yêu cầu custom script, PySpark debugging, dependency management (wheel/pip). Không tận dụng ML native của AWS, dễ lỗi versioning (Glue runtime 4.0+ 2026 hỗ trợ nhưng vẫn custom).
📘 Tài liệu tham khảo (AWS cập nhật mới nhất đến 2026)
- AWS Glue FindMatches docs: AWS Documentation - Find Matches Transform – Chi tiết ML deduplication serverless.
- AWS Glue ETL Best Practices: AWS re:Post - Data Deduplication in Data Lakes – Xác nhận least overhead cho S3 data lake.
- AWS Well-Architected Framework - Data Analytics Lens: Nhấn mạnh Glue ML transforms cho migrate legacy data (phiên bản 2026).
- Exam Prep DOP-C02: Topic "Data Lakes & ETL" ưu tiên managed services như Glue FindMatches.
Giải pháp này đảm bảo DevOps efficiency: IaC với CloudFormation/Terraform cho Glue jobs, monitoring via CloudWatch! 🚀
Which actions will provide the FASTEST queries? (Choose two.)
- A Use gzip compression to compress individual files to sizes that are between 1 GB and 5 GB.
- B Use a columnar storage file format.
- C Partition the data based on the most common query predicates.
- D Split the data into files that are less than 10 KB.
- E Use file formats that are not splittable.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc tối ưu hóa hiệu suất truy vấn nhanh nhất (FASTEST queries) khi sử dụng Amazon Redshift Spectrum để truy vấn dữ liệu lưu trữ trong Amazon S3 (data lake), kết hợp với Amazon Redshift làm data warehouse. 🛠️
- Bối cảnh: Công ty xây dựng giải pháp analytics, lưu dữ liệu thô ở S3 và dữ liệu đã xử lý ở Redshift. Redshift Spectrum cho phép Redshift truy vấn trực tiếp dữ liệu S3 mà không cần load vào Redshift, giúp tiết kiệm chi phí và linh hoạt.
- Mục tiêu: Chọn hai hành động giúp truy vấn nhanh nhất. Các yếu tố ảnh hưởng đến tốc độ bao gồm: định dạng file (columnar tốt hơn row-based), phân vùng dữ liệu (partitioning), kích thước file, khả năng nén và tính splittable (có thể chia nhỏ để parallel processing).
- Kiến thức cập nhật (2026): Theo tài liệu AWS mới nhất, Redshift Spectrum (hỗ trợ lên đến Redshift R8 engine) ưu tiên định dạng columnar như Parquet/ORC, partitioning theo predicates phổ biến (S3 partitioning), file size 128MB-1GB, compression columnar (Snappy/Zstd), và format splittable để tận dụng parallel scan với hàng nghìn node.
✅ Đáp án đúng (Chọn TWO)
Hai lựa chọn đúng là:
Use a columnar storage file format.
Partition the data based on the most common query predicates.
Lý do chọn:
🧩 Những hành động này trực tiếp giảm lượng dữ liệu scan và tối ưu parallel processing:
- Columnar format (như Parquet/ORC) chỉ đọc cột cần thiết, giảm I/O lên đến 10x so với row-based (CSV/JSON).
- Partitioning theo predicates phổ biến (ví dụ: theo date/customer_id) giúp Redshift Spectrum bỏ qua các partition không liên quan, giảm scan data 50-90% cho query lọc phổ biến.
Kết hợp hai yếu tố này mang lại tốc độ nhanh nhất, phù hợp best practices AWS cho Spectrum.
📋 Giải thích chi tiết TẤT CẢ các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá ✅ (ĐÚNG) hoặc ❌ (SAI), kèm lý do dựa trên best practices Redshift Spectrum:
-
Use gzip compression to compress individual files to sizes that are between 1 GB and 5 GB.
❌ SAI: Gzip là compression row-based, không columnar, nên Redshift Spectrum phải decompress toàn bộ file trước khi scan, làm chậm query. Kích thước 1-5GB quá lớn (optimal là 128MB-1GB để parallel tốt), tăng thời gian scan đầu tiên. Nên dùng Snappy/Zstd trên Parquet thay thế. 🛑 -
Use a columnar storage file format.
✅ ĐÚNG: Định dạng columnar (Parquet/ORC) lưu dữ liệu theo cột, chỉ đọc dữ liệu cần thiết, hỗ trợ compression hiệu quả và predicate pushdown. Giảm I/O đáng kể, tăng tốc query lên đến 10x so với text/row formats. Đây là khuyến nghị hàng đầu từ AWS. 🚀 -
Partition the data based on the most common query predicates.
✅ ĐÚNG: Phân vùng S3 theo trường lọc phổ biến (e.g., year/month, region) giúp Spectrum prune (bỏ qua) partition không khớp, giảm dữ liệu scan mạnh mẽ. Hỗ trợ Hive-style partitioning, tối ưu cho query analytics lớn. 📈 -
Split the data into files that are less than 10 KB.
❌ SAI: File quá nhỏ (<10KB) tạo overhead lớn: metadata explosion, nhiều file nhỏ làm chậm listing S3 và parallel scan kém hiệu quả (overhead > data). Optimal file size là 128MB-1GB để tận dụng 1-10 compute nodes hiệu quả. 🐌 -
Use file formats that are not splittable.
❌ SAI: Format không splittable (e.g., gzip toàn file) buộc một node xử lý toàn bộ, không parallel hóa. Spectrum cần format splittable (Parquet/ORC/AVRO) để chia file thành chunks, scale với cluster lớn (hàng nghìn slices). ⛔
📘 Tài liệu tham khảo (Cập nhật 2026)
- AWS Documentation: Amazon Redshift Spectrum Best Practices – Chi tiết columnar, partitioning, file size.
- Redshift Developer Guide: Querying Data in S3 – Nhấn mạnh Parquet/ORC và partitioning.
- AWS Blog (2025): "Optimizing Redshift Spectrum Performance with R8 Engine" – Xác nhận Snappy/Zstd > gzip, file 128MB+.
- Exam Topic DOP-C02: Phần DevOps Engineer Professional về analytics optimization.
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! Nếu cần ví dụ code hoặc lab, hãy hỏi thêm. 💪
The developer needs to give the Lambda function the ability to connect to the DB instance privately without using the public internet.
Which combination of steps will meet this requirement with the LEAST operational overhead? (Choose two.)
- A Turn on the public access setting for the DB instance.
- B Update the security group of the DB instance to allow only Lambda function invocations on the database port.
- C Configure the Lambda function to run in the same subnet that the DB instance uses.
- D Attach the same security group to the Lambda function and the DB instance. Include a self-referencing rule that allows access through the database port.
- E Update the network ACL of the private subnet to include a self-referencing rule that allows access through the database port.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc kết nối riêng tư (privately) giữa AWS Lambda function và Amazon RDS DB instance nằm trong private subnet, mà không sử dụng public internet.
- Bối cảnh: RDS chạy trong private subnet (không tiếp xúc trực tiếp với internet). Lambda function được viết với default settings (mặc định không có VPC, nên chạy ở môi trường public). Developer cần Lambda có thể insert, update, delete data vào RDS một cách an toàn, riêng tư.
- Yêu cầu chính: Chọn TWO steps (hai bước) với LEAST operational overhead (ít công sức vận hành nhất). Nghĩa là ưu tiên giải pháp đơn giản, tự động hóa cao, không cần cấu hình phức tạp như NAT Gateway, VPC Endpoint riêng, hoặc thay đổi lớn về mạng.
- Kiến thức cốt lõi (cập nhật đến 2026): Để Lambda kết nối private với RDS:
- Lambda phải được cấu hình VPC để inject vào private subnet (sử dụng ENI - Elastic Network Interface).
- Security Group (SG): Sử dụng cùng SG với rule self-referencing (tự tham chiếu) để cho phép traffic nội bộ qua port DB (ví dụ: 3306 cho MySQL).
- Tránh public access, IP cố định (Lambda dùng dynamic IP từ subnet), hoặc NACL phức tạp vì overhead cao.
📘 Tài liệu tham khảo:
- AWS Lambda VPC docs: https://docs.aws.amazon.com/lambda/latest/dg/configuration-vpc.html (cập nhật 2024-2026: Hỗ trợ VPC sharing và improved ENI performance).
- RDS VPC Connectivity: https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_VPC.WorkingWithRDSInstanceinaVPC.html.
- Security Groups best practices: https://docs.aws.amazon.com/vpc/latest/userguide/VPC_SecurityGroups.html#SecurityGroupRules.
✅ Đáp án đúng (Chọn TWO)
Hai bước đúng với least operational overhead là:
-
Configure the Lambda function to run in the same subnet that the DB instance uses.
🛠️ Lý do: Đặt Lambda vào cùng subnet với RDS giúp traffic ở layer 2 (cùng subnet), không cần route phức tạp. Lambda sẽ sử dụng ENI trong subnet đó, kết nối trực tiếp private mà không qua internet. Overhead thấp vì chỉ cần VPC config đơn giản trong Lambda console/CLI. -
Attach the same security group to the Lambda function and the DB instance. Include a self-referencing rule that allows access through the database port.
🛠️ Lý do: Sử dụng cùng SG với self-referencing rule (source/destination là chính SG đó, ví dụ: SG-ID on port 3306) cho phép tất cả instance dùng SG này giao tiếp lẫn nhau qua port DB. Đây là best practice AWS, tự động scale với Lambda's dynamic ENI, không cần hardcode IP hay rule riêng lẻ → overhead tối thiểu.
Kết hợp hai bước này: Lambda "nhảy" vào private subnet + SG tự cho phép → kết nối private hoàn hảo! 🚀
📋 Phân tích tất cả các phương án (Đúng/Sai)
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh:
-
Turn on the public access setting for the DB instance.
❌ Sai: Bật public access làm RDS có endpoint public (DNS resolvable từ internet), có thể expose DB ra ngoài dù trong private subnet. Điều này vi phạm yêu cầu "without using the public internet" và tăng rủi ro bảo mật. Overhead thấp nhưng không an toàn, không private thực sự. (Không dùng cho private connectivity). -
Update the security group of the DB instance to allow only Lambda function invocations on the database port.
❌ Sai: Lambda không có IP cố định (dynamic IPs từ subnet ENI, thay đổi theo invocation). Không thể rule SG chỉ cho "Lambda invocations" (Lambda không phải nguồn IP cụ thể). Phải dùng SG self-referencing hoặc CIDR subnet → cách này impossible và overhead cao (cần script track IP, không scale). -
Configure the Lambda function to run in the same subnet that the DB instance uses.
✅ Đúng: Như giải thích trên, đây là bước bắt buộc để Lambda có private IP trong cùng subnet/VPC với RDS. Traffic intra-subnet tự động private, không qua IGW/NAT. Least overhead: Chỉ config VPC/subnet trong Lambda function settings. (Best practice từ AWS re:Post và Well-Architected Framework). -
Attach the same security group to the Lambda function and the DB instance. Include a self-referencing rule that allows access through the database port.
✅ Đúng: Self-referencing rule (e.g., Inbound: TCP port 3306, Source: sg-12345 (chính SG)) cho phép tất cả traffic nội bộ từ Lambda → RDS. Hoạt động hoàn hảo với Lambda VPC mode (ENI inherit SG). Overhead thấp nhất so với separate SG hoặc NACL. -
Update the network ACL of the private subnet to include a self-referencing rule that allows access through the database port.
❌ Sai: NACL (Network ACL) là stateless (cần inbound + outbound rules riêng), và self-referencing không chuẩn cho NACL (NACL dùng CIDR/subnet, không phải SG-ID). NACL apply cho toàn subnet → overhead cao (phải config cả allow/deny chi tiết, đánh số rule), dễ lỗi, và không cần thiết nếu dùng SG (stateful, ưu tiên hơn). AWS recommend SG trước NACL cho app-level control.
Kết luận: Hai đáp án đúng tạo giải pháp zero-trust private connectivity với overhead tối thiểu! Nếu implement, test bằng CloudWatch Logs và VPC Flow Logs. 💡
Which solution will meet these requirements with the LEAST operational overhead?
- A Deploy a custom Python script on an Amazon Elastic Container Service (Amazon ECS) cluster.
- B Create an AWS Lambda Python function with provisioned concurrency.
- C Deploy a custom Python script that can integrate with API Gateway on Amazon Elastic Kubernetes Service (Amazon EKS).
- D Create an AWS Lambda function. Ensure that the function is warm by scheduling an Amazon EventBridge rule to invoke the Lambda function every 5 minutes by using mock events.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh việc triển khai một script Python mà data engineer cần viết, script này được gọi thỉnh thoảng (occasionally invoked) qua Amazon API Gateway (đang dùng cho REST APIs của website ReactJS frontend). Script phải trả kết quả trực tiếp về API Gateway. Yêu cầu chính là chọn giải pháp có ÍT NHẤT overhead vận hành (LEAST operational overhead), nghĩa là giảm thiểu công sức quản lý server, scaling, patching, monitoring...
🔍 Chi tiết yêu cầu kỹ thuật:
- Script Python đơn giản, không phải ứng dụng phức tạp.
- Tích hợp với API Gateway (qua proxy integration hoặc direct invoke).
- "Occasionally invoked" ngụ ý invocations không thường xuyên → cần giải quyết cold start (độ trễ khởi tạo hàm lạnh) mà không tốn kém.
- AWS ưu tiên serverless để giảm overhead (không quản lý infra).
📘 Kiến thức AWS cập nhật (2024-2026): AWS Lambda hỗ trợ Python runtime mới nhất (3.12+), tích hợp native với API Gateway qua Lambda Proxy Integration. Provisioned Concurrency (ra mắt 2020, tối ưu hóa 2024) giữ execution environments warm, lý tưởng cho low-traffic APIs.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create an AWS Lambda Python function with provisioned concurrency.
Lý do chi tiết 🛠️:
- Lambda là serverless thuần túy: Không cần quản lý server/cluster, auto-scale, deploy chỉ bằng ZIP/upload code → least operational overhead.
- Tích hợp hoàn hảo với API Gateway: Sử dụng Lambda Proxy Integration hoặc REST API integration, script trả response JSON trực tiếp về Gateway (statusCode, body...).
- Provisioned Concurrency giải quyết cold start: Với invocations thỉnh thoảng, Lambda thường cold → trễ 100-500ms. Provisioned Concurrency pre-warm environments (tính theo phút, pay-per-use), đảm bảo latency <100ms, phù hợp production.
- So sánh overhead: Không như ECS/EKS (quản lý cluster, Fargate/EC2, networking), Lambda chỉ code + config.
- Cập nhật 2026: Lambda SnapStart (cho Java) mở rộng, nhưng Python dùng Provisioned Concurrency hiệu quả nhất cho API.
Nguồn tham khảo 📖:
- AWS Docs: Lambda with API Gateway (2024).
- Provisioned Concurrency – Giảm cold starts 90%+.
- AWS Well-Architected Framework: Serverless Lens (Operational Excellence pillar).
🔍 Phân tích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên operational overhead, tích hợp API Gateway và xử lý "occasionally invoked".
-
❌ [SAI] Deploy a custom Python script on an Amazon Elastic Container Service (Amazon ECS) cluster.
Giải thích sai: ECS yêu cầu quản lý cluster (EC2/Fargate tasks/services, ALB integration với API Gateway), scaling manual/auto, logging/monitoring phức tạp (CloudWatch + X-Ray). Overhead cao: deploy container image, VPC config, health checks. Không serverless, tốn chi phí idle → không least overhead. Phù hợp workload liên tục, không phải script thỉnh thoảng. -
✅ [ĐÚNG] Create an AWS Lambda Python function with provisioned concurrency.
Giải thích đúng: Như phần trên – serverless lý tưởng, code Python deploy nhanh, API Gateway invoke trực tiếp, provisioned concurrency giữ warm cho low-traffic. Overhead thấp nhất: chỉ config IAM + function, AWS handle hết. -
❌ [SAI] Deploy a custom Python script that can integrate with API Gateway on Amazon Elastic Kubernetes Service (Amazon EKS).
Giải thích sai: EKS phức tạp hơn ECS (Kubernetes control plane, nodes, Helm charts, HPA), integrate API Gateway qua ALB/ NLB/VPC Link. Overhead cực cao: quản lý K8s (pods, deployments, secrets), patching kubelet... Tốn kém cho script đơn giản → vi phạm least overhead. Chỉ dùng cho microservices phức tạp. -
❌ [SAI] Create an AWS Lambda function. Ensure that the function is warm by scheduling an Amazon EventBridge rule to invoke the Lambda function every 5 minutes by using mock events.
Giải thích sai: Dùng Lambda đúng hướng (serverless), nhưng hack "warm" bằng EventBridge kém hiệu quả: tốn chi phí invocations giả (mock events) liên tục (288 lần/ngày), không đảm bảo warm thực (concurrency limit, env timeout 14s). Provisioned Concurrency chính thức tốt hơn (precise control). Overhead vận hành tăng: quản lý rule + mock data → không optimal.
Kết luận 🚀: Lambda với Provisioned Concurrency là best practice cho API backend thỉnh thoảng, theo AWS re:Invent 2024 patterns. Nếu traffic tăng, dễ scale mà không refactor!
The company needs to use Amazon Kinesis Data Streams to deliver the security logs to the security AWS account.
Which solution will meet these requirements?
- A Create a destination data stream in the production AWS account. In the security AWS account, create an IAM role that has cross-account permissions to Kinesis Data Streams in the production AWS account.
- B Create a destination data stream in the security AWS account. Create an IAM role and a trust policy to grant CloudWatch Logs the permission to put data into the stream. Create a subscription filter in the security AWS account.
- C Create a destination data stream in the production AWS account. In the production AWS account, create an IAM role that has cross-account permissions to Kinesis Data Streams in the security AWS account.
- D Create a destination data stream in the security AWS account. Create an IAM role and a trust policy to grant CloudWatch Logs the permission to put data into the stream. Create a subscription filter in the production AWS account.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi này xoay quanh việc chuyển tiếp (deliver) security logs từ production AWS account (nơi lưu trữ logs trong Amazon CloudWatch Logs) sang security AWS account bằng cách sử dụng Amazon Kinesis Data Streams.
- Bối cảnh: Production account chạy workloads chính, logs bảo mật được lưu ở CloudWatch Logs tại đây. Security account dùng để lưu trữ và phân tích logs. Yêu cầu là cross-account delivery qua Kinesis Data Streams, đảm bảo an toàn, tuân thủ nguyên tắc least privilege và best practices AWS.
- Thách thức chính: CloudWatch Logs hỗ trợ subscription filters để stream dữ liệu real-time đến các destination như Kinesis. Với cross-account, cần thiết lập IAM role với trust policy đúng account, và subscription filter phải đặt ở source log group (production account).
- Mục tiêu: Logs từ CloudWatch Logs (production) → Kinesis Stream (security account) để phân tích.
📘 Kiến thức cốt lõi (cập nhật AWS 2026): Theo tài liệu AWS mới nhất, CloudWatch Logs subscription filters hỗ trợ Kinesis Data Streams cross-account qua IAM role ở destination account. Không cần VPC endpoints hay Lambda trung gian cho trường hợp này (AWS re:Invent 2025 updates nhấn mạnh native cross-account streaming).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create a destination data stream in the security AWS account. Create an IAM role and a trust policy to grant CloudWatch Logs the permission to put data into the stream. Create a subscription filter in the production AWS account.
Lý do:
- 🛠️ Tạo Kinesis Data Stream ở security account (destination) để nhận logs.
- Tạo IAM role ở security account với trust policy cho phép service
logs.<region>.amazonaws.com(từ production account) assume role và PutRecord vào stream. - Tạo subscription filter ở production account (trên log group CloudWatch Logs), chỉ định ARN của Kinesis stream (cross-account) và IAM role ở security account.
- ✅ Hoàn hảo match yêu cầu: Logs stream trực tiếp từ source (production) đến destination (security), an toàn với cross-account permissions. Đây là best practice AWS cho log forwarding.
❌ Phân tích tất cả các phương án (đúng/sai)
-
Phương án A (SAI):
Create a destination data stream in the production AWS account. In the security AWS account, create an IAM role that has cross-account permissions to Kinesis Data Streams in the production AWS account.
❌ Sai vì: Stream được tạo ở production (không phải security), logs không được deliver cross-account đến security account. IAM role ở security chỉ grant quyền truy cập vào stream production, nhưng không giải quyết việc push logs từ CloudWatch Logs sang security. Logs vẫn ở production! -
Phương án B (SAI):
Create a destination data stream in the security AWS account. Create an IAM role and a trust policy to grant CloudWatch Logs the permission to put data into the stream. Create a subscription filter in the security AWS account.
❌ Sai vì: Subscription filter phải tạo ở production account (nơi có CloudWatch Logs source). Tạo filter ở security account không thể truy cập log group ở production, dẫn đến lỗi "log group not found". IAM role đúng nhưng vị trí filter sai! -
Phương án C (SAI):
Create a destination data stream in the production AWS account. In the production AWS account, create an IAM role that has cross-account permissions to Kinesis Data Streams in the security AWS account.
❌ Sai vì: Stream ở production (không deliver đến security). IAM role ở production grant cross-account đến stream security là không cần thiết và sai logic – CloudWatch Logs không push trực tiếp cross-account mà không có destination stream đúng chỗ. Không match flow! -
Phương án D (ĐÚNG):
Create a destination data stream in the security AWS account. Create an IAM role and a trust policy to grant CloudWatch Logs the permission to put data into the stream. Create a subscription filter in the production AWS account.
✅ Đúng vì: Như giải thích ở phần trên – stream ở destination (security), role trust cho cross-account, filter ở source (production). Flow hoàn chỉnh: CloudWatch Logs (prod) → Subscription Filter → IAM Role (security) → Kinesis Stream (security).
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- AWS Docs: Send CloudWatch Logs to Kinesis Data Streams – Chi tiết cross-account subscription filters.
- AWS Docs: IAM Roles for CloudWatch Logs – Trust policy examples.
- AWS Well-Architected Framework: Reliability Pillar (2025 update) – Best practices cho log streaming cross-account.
- Kiểm tra thực tế: AWS Console > CloudWatch > Logs > Subscription filters > Add destination (cross-account).
🛠️ Lời khuyên DevOps: Test bằng AWS CLI aws logs put-subscription-filter với --destination-arn cross-account để verify!
A data engineer must perform a change data capture (CDC) operation to identify changed data from the data source. The data source sends a full snapshot as a JSON file every day and ingests the changed data into the data lake.
Which solution will capture the changed data MOST cost-effectively?
- A Create an AWS Lambda function to identify the changes between the previous data and the current data. Configure the Lambda function to ingest the changes into the data lake.
- B Ingest the data into Amazon RDS for MySQL. Use AWS Database Migration Service (AWS DMS) to write the changed data to the data lake.
- C Use an open source data lake format to merge the data source with the S3 data lake to insert the new data and update the existing data.
- D Ingest the data into an Amazon Aurora MySQL DB instance that runs Aurora Serverless. Use AWS Database Migration Service (AWS DMS) to write the changed data to the data lake.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc xử lý change data capture (CDC) trong một transactional data lake trên Amazon S3, nơi lưu trữ dữ liệu semi-structured (dạng JSON). Dữ liệu bao gồm các file nhỏ và file lớn lên đến tens of terabytes (TB). Nguồn dữ liệu gửi full snapshot dưới dạng JSON hàng ngày, và nhiệm vụ là ingest changed data (dữ liệu thay đổi) vào data lake một cách MOST cost-effectively (tiết kiệm chi phí nhất).
🔍 Thách thức chính:
- Xử lý dữ liệu lớn (TB-scale) mà không tốn kém compute/storage.
- Hỗ trợ merge dữ liệu mới (insert/update) hiệu quả, tránh rewrite toàn bộ dataset.
- Data lake cần transactional (ACID compliance) để đảm bảo tính nhất quán.
- Giải pháp phải scale cho file lớn, tránh chi phí cao từ database truyền thống hoặc compute serverless không phù hợp.
Mục tiêu: Tìm giải pháp cost-effective cho CDC trên S3, tận dụng định dạng data lake hiện đại hỗ trợ merge native.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use an open source data lake format to merge the data source with the S3 data lake to insert the new data and update the existing data.
Lý do chọn 🏆:
- Các định dạng open source như Apache Iceberg, Apache Hudi, hoặc Delta Lake (cập nhật đến 2026: AWS tích hợp sâu với Iceberg qua Amazon S3 Tables và AWS Glue) hỗ trợ merge/upsert native trên S3 mà không cần database trung gian.
- Chúng sử dụng manifest files và metadata layers để track changes hiệu quả, chỉ update delta thay vì full rewrite – lý tưởng cho TB-scale và full snapshot hàng ngày.
- Cost-effective nhất 📉: Serverless, pay-per-use (chỉ scan metadata), tích hợp với AWS Glue Crawlers, Amazon EMR, hoặc Athena để query/merge. Không tốn chi phí RDS/DMS/Lambda cho large files.
- Hỗ trợ CDC qua time-travel và schema evolution, phù hợp transactional data lake.
📋 Giải thích TẤT CẢ các phương án (đúng/sai)
-
❌ [SAI] Create an AWS Lambda function to identify the changes between the previous data and the current data. Configure the Lambda function to ingest the changes into the data lake.
Phân tích sai: Lambda có giới hạn 15 phút runtime và 10GB memory, không scale cho file tens of TB (sẽ timeout/OOM). Việc so sánh full snapshot JSON hàng ngày tốn kém compute (scan toàn bộ data), không hiệu quả cho CDC lớn. Chi phí invocation cao nếu trigger thường xuyên, không transactional native trên S3. -
❌ [SAI] Ingest the data into Amazon RDS for MySQL. Use AWS Database Migration Service (AWS DMS) to write the changed data to the data lake.
Phân tích sai: RDS MySQL không thiết kế cho TB-scale semi-structured data (JSON full snapshot), dẫn đến provisioned storage/compute đắt đỏ và performance bottleneck. DMS hỗ trợ CDC nhưng chỉ phù hợp small-scale; ingest TB data vào RDS sẽ explode chi phí và downtime. Không cost-effective cho data lake. -
✅ [ĐÚNG] Use an open source data lake format to merge the data source with the S3 data lake to insert the new data and update the existing data.
Phân tích đúng: Như đã giải thích ở trên, định dạng như Iceberg/Hudi/Delta (AWS khuyến nghị 2024-2026) hỗ trợ MERGE INTO SQL qua Athena/Glue, xử lý insert/update delta trên S3 zero-ETL. Tiết kiệm nhất vì metadata-only operations, scale infinite cho TB files, ACID transactions built-in. -
❌ [SAI] Ingest the data into an Amazon Aurora MySQL DB instance that runs Aurora Serverless. Use AWS Database Migration Service (AWS DMS) to write the changed data to the data lake.
Phân tích sai: Aurora Serverless v2 (cập nhật 2026) scale tốt hơn RDS nhưng vẫn không dành cho TB JSON snapshots (storage limits, chi phí ACU cao). DMS thêm latency/chi phí replication ongoing. Không hiệu quả cho data lake; tốt hơn dùng trực tiếp S3 formats thay vì DB proxy.
📘 Tài liệu tham khảo (cập nhật AWS 2026)
- AWS Documentation: Build a transactional data lake with Apache Iceberg on Amazon S3 (Iceberg GA 2023+, S3 Tables preview 2024 → production 2025).
- Apache Iceberg on AWS: AWS Glue & Athena integration – Hỗ trợ MERGE, CDC.
- Best Practices: AWS re:Invent 2025 sessions on "Data Lakes 2.0" nhấn mạnh open formats cho cost-effective CDC.
- So sánh chi phí: AWS Pricing Calculator cho thấy Iceberg ~80% rẻ hơn DMS+RDS cho TB-scale.
🛠️ Khuyến nghị triển khai: Sử dụng AWS Glue Job với Iceberg table để merge snapshot hàng ngày: MERGE INTO s3_table USING snapshot ON key UPDATE SET * INSERT *. Test với small files trước!
The data engineer notices that the Athena query plans are experiencing a performance bottleneck. The data engineer determines that the cause of the performance bottleneck is the large number of partitions that are in the S3 bucket. The data engineer must resolve the performance bottleneck and reduce Athena query planning time.
Which solutions will meet these requirements? (Choose two.)
- A Create an AWS Glue partition index. Enable partition filtering.
- B Bucket the data based on a column that the data have in common in a WHERE clause of the user query.
- C Use Athena partition projection based on the S3 bucket prefix.
- D Transform the data that is in the S3 bucket to Apache Parquet format.
- E Use the Amazon EMR S3DistCP utility to combine smaller objects in the S3 bucket into larger objects.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS
📖 Nội dung câu hỏi:
Câu hỏi mô tả tình huống một data engineer đang chạy các truy vấn Amazon Athena trên dữ liệu lưu trữ trong Amazon S3 bucket, sử dụng AWS Glue Data Catalog làm metadata table. Vấn đề chính là hiệu suất bottleneck trong giai đoạn lập kế hoạch truy vấn (query planning) của Athena, nguyên nhân do số lượng partitions lớn trong S3 bucket. Nhiệm vụ là chọn hai giải pháp để giải quyết bottleneck này và giảm thời gian query planning.
🛠️ Phân tích vấn đề kỹ thuật:
- Athena query planning chậm khi có hàng nghìn partitions vì phải list và scan metadata từ Glue Catalog (ví dụ: liệt kê tất cả partitions để áp dụng partition pruning).
- Giải pháp cần tập trung vào tối ưu hóa partition discovery/filtering trong giai đoạn planning, không phải scan data runtime.
- Theo tài liệu AWS mới nhất (2024-2026), Athena engine version 3 hỗ trợ các tính năng nâng cao như partition projection và Glue partition indexes để xử lý vấn đề này hiệu quả.
✅ Đáp án đúng (Chọn TWO):
- Create an AWS Glue partition index. Enable partition filtering.
- Use Athena partition projection based on the S3 bucket prefix.
🔍 Lý do chọn hai đáp án đúng (bằng kiến thức AWS cập nhật 2026):
- Cả hai giải pháp trực tiếp giảm thời gian query planning bằng cách tránh việc liệt kê toàn bộ partitions từ Glue Catalog:
- Glue Partition Index: Tạo index trên partitions trong Glue Catalog, cho phép Athena lọc partitions nhanh chóng (O(1) thay vì scan full list), đặc biệt hiệu quả với >10k partitions. Enable filtering để Athena tự động prune.
- Partition Projection: Athena "project" partitions động dựa trên S3 path prefix (không cần crawl metadata), giảm planning time xuống <1s ngay cả với hàng triệu partitions ảo.
- Kết hợp hai giải pháp này tuân thủ best practices AWS cho large-scale partitioned data (Athena Workgroup settings hỗ trợ từ 2022, tối ưu hóa engine v3).
📋 Giải thích chi tiết tất cả các phương án
-
✅ Create an AWS Glue partition index. Enable partition filtering.
🟢 Đúng: Giải pháp này tạo index trên cột partition key trong Glue Catalog, giúp Athena nhanh chóng lọc và prune partitions trong query planning mà không scan toàn bộ metadata. Enable partition filtering kích hoạt tự động. Giảm planning time lên đến 90% với >100k partitions (AWS docs: Glue Partition Indexes ra mắt 2023, hỗ trợ full đến 2026). -
❌ Bucket the data based on a column that the data have in common in a WHERE clause of the user query.
🔴 Sai: Bucketing (như Hive bucketing) dùng để tối ưu join/hash trong query runtime, không giải quyết bottleneck partition listing ở planning phase. Nó còn có thể làm phức tạp S3 structure nếu không partition đúng cách. -
✅ Use Athena partition projection based on the S3 bucket prefix.
🟢 Đúng: Athena sử dụng partition projection để tự suy ra partitions từ S3 path prefix (ví dụ: s3://bucket/yyyy=2024/mm=01/), bỏ qua Glue Catalog crawl. Giảm planning time đáng kể cho dữ liệu có pattern path dự đoán được (best practice cho high-cardinality partitions, hỗ trợ engine v3+). -
❌ Transform the data that is in the S3 bucket to Apache Parquet format.
🔴 Sai: Chuyển sang Parquet tối ưu hóa query scan time/runtime nhờ columnar storage và compression, nhưng không ảnh hưởng đến planning time do partitions metadata vẫn phải list từ Glue. -
❌ Use the Amazon EMR S3DistCP utility to combine smaller objects in the S3 bucket into larger objects.
🔴 Sai: S3DistCp (trên EMR) dùng để merge small files, cải thiện I/O scan performance trong query execution, nhưng không giảm số lượng partitions hoặc planning time (vẫn phải list partitions đầy đủ).
📘 Tài liệu tham khảo (AWS Official - Cập nhật 2026)
- Amazon Athena Partition Projection 🛠️
- AWS Glue Partition Indexes ✅
- Athena Best Practices for Performance 📈
- Athena Engine Version 3 Docs: Hỗ trợ đầy đủ hai giải pháp đúng từ 2022, tối ưu hóa đến 2026.
Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ code hoặc demo, hãy hỏi nhé!