Ngân hàng đề — AWS Certified Data Engineer Associate
Tìm thấy 867 câu.
Which AWS service or feature will meet these requirements with the LEAST operational overhead?
- A Amazon Managed Streaming for Apache Kafka (Amazon MSK)
- B Amazon AppFlow
- C AWS Glue Data Catalog
- D Amazon Kinesis
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi mô tả một công ty truyền thông sử dụng các ứng dụng SaaS (Software as a Service) để thu thập dữ liệu thông qua công cụ bên thứ ba (third-party tools). Họ cần lưu trữ dữ liệu vào Amazon S3 bucket và sau đó sử dụng Amazon Redshift để thực hiện phân tích dữ liệu. Yêu cầu chính là chọn dịch vụ hoặc tính năng AWS đáp ứng được quy trình này với ít tác động vận hành nhất (LEAST operational overhead), nghĩa là dịch vụ managed hoàn toàn, không cần quản lý hạ tầng phức tạp như provisioning, scaling, monitoring thủ công.
📌 Tóm tắt luồng dữ liệu: SaaS/third-party tools → Amazon S3 (lưu trữ) → Amazon Redshift (phân tích). Dịch vụ lý tưởng phải hỗ trợ tích hợp trực tiếp, tự động hóa ingestion mà không đòi hỏi code nhiều hoặc quản lý server/cluster.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Amazon AppFlow
🛠️ Lý do: Amazon AppFlow là dịch vụ tích hợp dữ liệu không cần code (no-code/low-code) được thiết kế chuyên biệt để chuyển dữ liệu từ các ứng dụng SaaS phổ biến (như Salesforce, Google Analytics, Marketo, Slack, v.v.) trực tiếp vào các dịch vụ AWS như S3 và Redshift. Nó hoàn toàn managed bởi AWS, tự động xử lý authentication, transformation, scheduling, và error handling, giúp giảm thiểu tối đa operational overhead (không cần quản lý stream, catalog, hay cluster). Dữ liệu được lưu dưới dạng file Parquet/CSV vào S3, sẵn sàng load vào Redshift qua COPY command hoặc Spectrum. Đây là giải pháp tối ưu nhất theo best practices AWS năm 2024-2026.
📋 Phân tích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên kiến thức AWS mới nhất (AWS Well-Architected Framework và dịch vụ AppFlow v2.0+ hỗ trợ 50+ connector SaaS).
-
❌ Amazon Managed Streaming for Apache Kafka (Amazon MSK)
Sai vì Amazon MSK là dịch vụ managed Kafka dùng cho streaming dữ liệu real-time quy mô lớn (như log, IoT), không phải tích hợp SaaS/third-party. Người dùng vẫn phải quản lý topics, consumer groups, scaling cluster, và viết code Kafka connector để push từ SaaS vào S3/Redshift – dẫn đến operational overhead cao (provisioning, monitoring, VPC setup). Không phù hợp với "LEAST overhead" cho SaaS ingestion. -
✅ Amazon AppFlow
Đúng như đã giải thích ở trên. Đây là dịch vụ zero-ETL integration dành riêng cho SaaS-to-AWS, hỗ trợ private flow với VPC endpoint (tính năng mới 2025), batch/real-time mode, và tích hợp trực tiếp S3/Redshift mà không cần Glue hay Lambda. Overhead gần như bằng 0: chỉ cần UI config flow trong vài phút. -
❌ AWS Glue Data Catalog
Sai vì AWS Glue Data Catalog chỉ là metadata repository (danh mục dữ liệu) để lưu schema/table definitions cho Athena, Glue Jobs, Redshift Spectrum. Nó không hỗ trợ ingestion dữ liệu từ SaaS vào S3, mà chỉ dùng sau khi dữ liệu đã ở S3 (crawl/discover). Sử dụng nó đòi hỏi thêm Glue Crawler/Job để ETL – tăng overhead đáng kể. -
❌ Amazon Kinesis
Sai vì Amazon Kinesis (Data Streams/Firehose) dùng cho real-time streaming từ sources như app/logs, không tối ưu cho SaaS batch integration. Với Kinesis Firehose, có thể deliver vào S3 nhưng vẫn cần code producer từ third-party, quản lý shards/retention/buffer, và transform thủ công – overhead cao hơn AppFlow (không hỗ trợ native SaaS connectors như Salesforce OAuth).
📘 Tài liệu tham khảo
- AWS Documentation chính thức (2026 update): Amazon AppFlow User Guide – Xác nhận hỗ trợ 50+ SaaS connectors, direct S3/Redshift sink.
- AWS Well-Architected Framework (Data Analytics Lens, 2025): Khuyến nghị AppFlow cho "SaaS to AWS integration with minimal ops".
- Exam Prep: AWS Certified Data Engineer - Associate/DevOps Pro sample questions (re:Post forums 2024-2026) nhấn mạnh AppFlow cho LEAST overhead SaaS scenarios.
- Blog AWS: "Simplify SaaS Data Integration with Amazon AppFlow" (aws.amazon.com/blogs/big-data, cập nhật 2025).
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm ví dụ thực tế, hãy hỏi nhé!
The data engineer's original query is as follows:
SELECT product_name, sum(sales_amount)
FROM sales_data -
WHERE year = 2023 -
GROUP BY product_name -
How should the data engineer modify the Athena query to meet these requirements?
- A Replace sum(sales_amount) with count(*) for the aggregation.
- B Change WHERE year = 2023 to WHERE extract(year FROM sales_data) = 2023.
- C Add HAVING sum(sales_amount) > 0 after the GROUP BY clause.
- D Remove the GROUP BY clause.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh việc một data engineer sử dụng Amazon Athena để truy vấn dữ liệu bán hàng (sales data) lưu trữ trên Amazon S3. Bảng dữ liệu tên là sales_data chứa thông tin về product_name và sales_amount. Data engineer viết query để lấy tổng sales_amount cho năm 2023 theo từng product_name, nhưng query không trả về kết quả cho tất cả các sản phẩm có trong bảng sales_data.
Query gốc (có lỗi format nhỏ với dấu - nhưng không ảnh hưởng):
SELECT product_name, sum(sales_amount)
FROM sales_data
WHERE year = 2023
GROUP BY product_name
Vấn đề chính (troubleshoot): Query chỉ filter dữ liệu năm 2023 qua điều kiện WHERE year = 2023, nhưng không match được tất cả rows liên quan đến các sản phẩm vì cột year có thể là kiểu dữ liệu DATE hoặc TIMESTAMP (không phải INTEGER), dẫn đến so sánh trực tiếp year = 2023 thất bại (type mismatch hoặc không extract đúng năm). Athena (dựa trên Presto/Trino engine) yêu cầu sử dụng hàm EXTRACT để lấy phần năm từ cột date/time. Kết quả là một số sản phẩm không xuất hiện vì không có rows match WHERE clause sau GROUP BY.
Mục tiêu: Sửa query để lấy đúng tổng sales_amount năm 2023 cho tất cả sản phẩm trong bảng (hoặc ít nhất resolve missing results do filter sai).
📘 Kiến thức AWS cập nhật đến 2026: Athena engine version 3 (Trino-based) hỗ trợ đầy đủ hàm EXTRACT(YEAR FROM date_column) cho datetime partitioning và querying. Không có thay đổi lớn về syntax này từ 2023-2026 (xem AWS Athena docs).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Change WHERE year = 2023 to WHERE extract(year FROM sales_data) = 2023.
Lý do 🛠️:
- Cột
yeartrong bảngsales_data很可能 là kiểu DATE hoặc TIMESTAMP (thường gặp với dữ liệu S3 partitioned by date). So sánh trực tiếpyear = 2023(integer) sẽ không match vì Athena không tự động cast đúng cách, dẫn đến 0 rows hoặc missing products. - Sửa bằng
EXTRACT(year FROM sales_data) = 2023(giả sửsales_datalà tên cột date tương đương, hoặc lỗi typo từsales_date) trích xuất phần năm từ cột date/time, đảm bảo filter chính xác rows năm 2023. - Sau GROUP BY, tất cả products có sales 2023 sẽ appear với sum đúng (products không có sales 2023 sẽ có sum=0 hoặc NULL tùy schema, nhưng ít nhất resolve filter issue).
- Đây là best practice cho Athena query trên S3 data lakes.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh:
-
Replace sum(sales_amount) with count(*) for the aggregation.
❌ Sai: Thaysum(sales_amount)bằngcount(*)chỉ đổi aggregation từ tổng tiền sang đếm rows, không giải quyết gốc rễ vấn đề filter WHERE. Vẫn missing products vì rows không matchyear = 2023do type mismatch. Count(*) chỉ làm query trả số lượng thay vì tổng sales, không liên quan đến troubleshoot missing results. -
Change WHERE year = 2023 to WHERE extract(year FROM sales_data) = 2023.
✅ Đúng: Như giải thích trên, hàmEXTRACTfix type mismatch của cột date/time, cho phép filter đúng năm 2023. Query sẽ match đầy đủ rows, GROUP BY trả tất cả products có data 2023 (hoặc sum=0 cho products khác nếu có rows). Đây là cách chuẩn theo Presto/Trino syntax trong Athena. -
Add HAVING sum(sales_amount) > 0 after the GROUP BY clause.
❌ Sai: ThêmHAVING sum(sales_amount) > 0sẽ filter OUT thêm các products có tổng sales = 0 sau GROUP BY, làm tăng số products missing thay vì resolve. Vấn đề gốc vẫn ở WHERE clause, HAVING chỉ dùng để post-aggregate filter, không fix missing data từ đầu. -
Remove the GROUP BY clause.
❌ Sai: BỏGROUP BY product_namesẽ sum tất cả sales_amount của năm 2023 thành một row duy nhất (không per product), mất grouping theo sản phẩm. Không trả về results cho từng product, hoàn toàn ngược yêu cầu "sales amounts for several products".
📚 Tài liệu tham khảo
- AWS Athena Documentation (2026): Querying with EXTRACT function – Hỗ trợ
EXTRACT(YEAR FROM date_column). - Presto/Trino Functions (Athena engine v3): Date/Time Functions – Syntax chuẩn cho extract year.
- Best Practices for Athena Queries on S3: Athena User Guide - Partitioning & Date Filters.
Hy vọng phân tích này giúp bạn ôn thi AWS hiệu quả! 🚀 Nếu cần query ví dụ đầy đủ, hãy hỏi thêm.
Which solution will meet these requirements with the LEAST operational overhead?
- A Configure an AWS Lambda function to load data from the S3 bucket into a pandas dataframe. Write a SQL SELECT statement on the dataframe to query the required column.
- B Use S3 Select to write a SQL SELECT statement to retrieve the required column from the S3 objects.
- C Prepare an AWS Glue DataBrew project to consume the S3 objects and to query the required column.
- D Run an AWS Glue crawler on the S3 objects. Use a SQL SELECT statement in Amazon Athena to query the required column.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào một nhiệm vụ một lần duy nhất (one-time task) của data engineer: đọc dữ liệu từ các objects ở định dạng Apache Parquet lưu trữ trong Amazon S3 bucket, và chỉ cần query một cột duy nhất từ dữ liệu đó. Yêu cầu chính là chọn giải pháp có ít chi phí vận hành nhất (LEAST operational overhead).
📘 Bối cảnh kỹ thuật:
- Parquet là định dạng columnar (cột hóa), rất phù hợp cho query selective (chỉ lấy một cột) mà không cần load toàn bộ file.
- Operational overhead bao gồm: thiết lập tài nguyên (provisioning), quản lý (setup crawler, job, function), chi phí tính toán, và thời gian triển khai. Vì là nhiệm vụ một lần, ưu tiên giải pháp serverless, không cần setup phức tạp, query trực tiếp trên S3.
Kiến thức cập nhật đến 2026: S3 Select vẫn là tính năng core của S3 (ra mắt 2018, hỗ trợ Parquet đầy đủ), không thay đổi lớn. Các dịch vụ khác như Glue/Athena/DataBrew vẫn yêu cầu setup thêm (crawler, catalog), dẫn đến overhead cao hơn cho one-time task.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use S3 Select to write a SQL SELECT statement to retrieve the required column from the S3 objects.
Lý do 🛠️:
- S3 Select cho phép query trực tiếp trên objects S3 (không copy/load data), hỗ trợ SQL-1.0 trên Parquet (chỉ scan cột cần thiết nhờ columnar format).
- Least overhead: Serverless, không cần setup Lambda/Glue/Athena. Chỉ cần API call hoặc AWS CLI một lần (ví dụ:
aws s3api select-object). - Phù hợp one-time: Chi phí thấp (~$0.01/1000 requests + scan data), nhanh (giây), không quản lý infra.
- Nguồn: AWS S3 Select Documentation (cập nhật 2025).
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, đánh dấu ✅ (đúng) hoặc ❌ (sai), với giải thích bằng tiếng Việt. Giữ nguyên văn bản gốc phương án.
-
❌ Configure an AWS Lambda function to load data from the S3 bucket into a pandas dataframe. Write a SQL SELECT statement on the dataframe to query the required column.
Phân tích sai: Yêu cầu tạo Lambda function (code Python + pandas), trigger S3 event, load toàn bộ data vào memory (overhead cao cho large files). Pandas không hỗ trợ SQL native tốt (cần pandasql), phải deploy/package libs. Overhead lớn: code, IAM, timeout (15p), chi phí invocation. Không tối ưu cho one-time, Parquet query selective. -
✅ Use S3 Select to write a SQL SELECT statement to retrieve the required column from the S3 objects.
Phân tích đúng: Như đã giải thích ở trên. Hoàn hảo cho yêu cầu: Query SQL trực tiếp (ví dụ:SELECT col1 FROM s3://bucket/file.parquet), chỉ trả về 1 cột, serverless thuần túy. Overhead = 0 (không setup). -
❌ Prepare an AWS Glue DataBrew project to consume the S3 objects and to query the required column.
Phân tích sai: AWS Glue DataBrew (nay là Amazon DataBrew, 2025+) dùng cho data prep/visualization, yêu cầu tạo project, dataset, recipe (UI-based). Overhead cao: Setup console, sampling data, export results. Không phải cho simple query one-time, chi phí ~$1/giờ interactive session. Phù hợp cleaning/transform, không phải SELECT column nhanh. -
❌ Run an AWS Glue crawler on the S3 objects. Use a SQL SELECT statement in Amazon Athena to query the required column.
Phân tích sai: Glue Crawler phải chạy để catalog schema (thời gian 5-30p, chi phí $0.44/DPU-giờ), tạo table Glue Data Catalog. Sau đó query Athena (serverless nhưng phụ thuộc catalog). Overhead lớn cho one-time: Crawler setup/IAM, catalog maintenance, query scan full partition trừ khi optimize. Athena tốt cho ad-hoc lớn, nhưng thừa thãi so S3 Select.
Kết luận 🎯: S3 Select là lựa chọn tối ưu nhất cho query selective Parquet trên S3 mà không overhead. Các giải pháp khác phù hợp workload recurring/large-scale hơn. Tham khảo thêm: AWS Well-Architected Data Analytics Lens (2025).
Which solution will meet this requirement with the LEAST effort?
- A Use Apache Airflow to refresh the materialized views.
- B Use an AWS Lambda user-defined function (UDF) within Amazon Redshift to refresh the materialized views.
- C Use the query editor v2 in Amazon Redshift to refresh the materialized views.
- D Use an AWS Glue workflow to refresh the materialized views.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào Amazon Redshift – dịch vụ kho dữ liệu (data warehouse) của AWS. Công ty đang sử dụng Redshift và cần tự động hóa lịch trình refresh (làm mới) cho các materialized views (chế độ xem vật chất hóa). Materialized views trong Redshift lưu trữ kết quả truy vấn dưới dạng bảng vật lý để tăng tốc độ truy vấn, nhưng cần refresh định kỳ để cập nhật dữ liệu từ base tables.
Yêu cầu chính: Giải pháp nào đáp ứng với ÍT NỖ LỰC NHẤT (LEAST effort), nghĩa là ưu tiên phương pháp đơn giản, tích hợp sẵn, không cần xây dựng thêm công cụ phức tạp. Đây là kiến thức cập nhật đến năm 2026, dựa trên tính năng Scheduled Queries trong Query Editor v2 của Redshift (ra mắt từ 2021 và được cải tiến liên tục).
📘 Nguồn tham khảo chính:
- AWS Documentation: Working with scheduled queries in Query Editor v2
- Refreshing a materialized view
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use the query editor v2 in Amazon Redshift to refresh the materialized views.
Lý do:
- Query Editor v2 là công cụ tích hợp sẵn trong AWS Management Console của Redshift, cho phép tạo lịch trình truy vấn (scheduled queries) trực tiếp mà không cần code thêm hay triển khai dịch vụ ngoài.
- Bạn chỉ cần viết lệnh
REFRESH MATERIALIZED VIEW <view_name>, lưu thành query, và thiết lập lịch cron đơn giản (hàng giờ/ngày/tuần). - Least effort: Không setup infrastructure, không code ETL, chỉ vài cú click trong console. Tính năng này được tối ưu cho Redshift, hỗ trợ auto-refresh MV lên đến hàng nghìn views mà không tốn tài nguyên thừa. ✅ Hoàn hảo cho tự động hóa nhanh chóng!
🛠️ Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể:
-
Use Apache Airflow to refresh the materialized views.
❌ Sai: Apache Airflow là công cụ orchestration mã nguồn mở mạnh mẽ, nhưng cần triển khai đầy đủ (self-managed hoặc trên MWAA), viết DAGs phức tạp để kết nối Redshift qua JDBC/ODBC và chạy lệnh REFRESH. Nỗ lực cao: Setup cluster, quản lý scaling, bảo mật – không phải least effort. Phù hợp cho pipeline lớn, không dành cho refresh MV đơn giản. -
Use an AWS Lambda user-defined function (UDF) within Amazon Redshift to refresh the materialized views.
❌ Sai: Lambda UDF (User-Defined Function) trong Redshift chỉ dùng để thực thi logic tùy chỉnh trong truy vấn SQL (như Python/AWS SDK functions), KHÔNG hỗ trợ refresh materialized views. Lệnh REFRESH là native command của Redshift, không thể wrap vào UDF. Sử dụng sẽ lỗi và yêu cầu code phức tạp, không tự động hóa lịch trình. Least effort? Hoàn toàn không! -
Use the query editor v2 in Amazon Redshift to refresh the materialized views.
✅ Đúng (như đã giải thích ở trên): Tích hợp sẵn, UI thân thiện, hỗ trợ cron schedules trực tiếp cho lệnh REFRESH. Least effort tuyệt đối – chỉ cần console access! Đã được AWS khuyến nghị cho automation đơn giản trong docs mới nhất (2026). -
Use an AWS Glue workflow to refresh the materialized views.
❌ Sai: AWS Glue là dịch vụ ETL serverless cho data integration, có thể gọi stored procedures qua JDBC, nhưng KHÔNG tối ưu cho refresh MV Redshift. Cần tạo job Spark/script, setup crawler/job triggers – nỗ lực cao hơn nhiều so với Query Editor v2 (thiết kế schema, manage runs, chi phí Glue units). Glue phù hợp ETL lớn, không phải MV refresh thuần túy.
Kết luận: Chọn Query Editor v2 để tối ưu chi phí, tốc độ triển khai và ít bảo trì nhất! 🚀 Nếu cần scale lớn, có thể kết hợp EventBridge + Lambda sau.
Which solution will meet these requirements with the LEAST management overhead?
- A Use an AWS Step Functions workflow that includes a state machine. Configure the state machine to run the Lambda function and then the AWS Glue job.
- B Use an Apache Airflow workflow that is deployed on an Amazon EC2 instance. Define a directed acyclic graph (DAG) in which the first task is to call the Lambda function and the second task is to call the AWS Glue job.
- C Use an AWS Glue workflow to run the Lambda function and then the AWS Glue job.
- D Use an Apache Airflow workflow that is deployed on Amazon Elastic Kubernetes Service (Amazon EKS). Define a directed acyclic graph (DAG) in which the first task is to call the Lambda function and the second task is to call the AWS Glue job.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi yêu cầu thiết kế một data pipeline orchestration đơn giản bao gồm một AWS Lambda function và một AWS Glue job, với yêu cầu tích hợp các dịch vụ AWS và ưu tiên LEAST management overhead (ít quản lý nhất).
✅ Yêu cầu chính:
- Orchestrate (điều phối) pipeline theo thứ tự: Chạy Lambda trước, sau đó Glue job.
- Tích hợp native với AWS services (không cần infra tự quản).
- Ưu tiên giải pháp serverless/managed hoàn toàn để giảm overhead như provisioning, scaling, patching, monitoring infra.
🛠️ Bối cảnh: Data pipeline thường cần workflow engine để xử lý dependencies, retries, error handling. AWS cung cấp nhiều orchestration tools, nhưng phải chọn cái ít quản lý nhất (fully managed, no servers).
✅ Đáp án đúng
Use an AWS Step Functions workflow that includes a state machine. Configure the state machine to run the Lambda function and then the AWS Glue job.
Lý do lựa chọn:
- AWS Step Functions là dịch vụ serverless orchestration hoàn toàn managed bởi AWS, tích hợp native với Lambda (qua state
Lambda Invoke) và Glue (qua stateAWS Glue StartJobRunhoặcTaskstate). - Hỗ trợ sequential workflow (Lambda → Glue) dễ dàng qua ASL (Amazon States Language), với built-in retries, error handling, parallelism, và monitoring qua CloudWatch.
- Least management overhead: Không cần quản lý server/infra, auto-scale, pay-per-use, cập nhật mới nhất 2026 vẫn là best practice cho hybrid workloads (Lambda + Glue).
- Phù hợp DOP-101/ DOP-C01 exam (DevOps Professional).
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, với đánh dấu ✅ (đúng) hoặc ❌ (sai), giữ nguyên văn bản gốc:
-
✅ Use an AWS Step Functions workflow that includes a state machine. Configure the state machine to run the Lambda function and then the AWS Glue job.
Giải thích đúng: Như trên, đây là giải pháp tối ưu với zero management (serverless), tích hợp trực tiếp Lambda/Glue states. Hỗ trợ visual workflow studio, execution history, và Map/Parallel states cho scale. Overhead thấp nhất so với các option tự quản. -
❌ Use an Apache Airflow workflow that is deployed on an Amazon EC2 instance. Define a directed acyclic graph (DAG) in which the first task is to call the Lambda function and the second task is to call the AWS Glue job.
Giải thích sai: Airflow cần tự deploy/manage EC2 (provisioning, scaling, security, patching OS/Python), overhead cao (self-managed). DAG có thể orchestrate Lambda/Glue qua operators (LambdaOperator, GlueOperator), nhưng phải handle scheduler, metadata DB (RDS?), webserver. Không "least overhead", vi phạm managed-native AWS. -
❌ Use an AWS Glue workflow to run the Lambda function and then the AWS Glue job.
Giải thích sai: AWS Glue Workflow chỉ hỗ trợ chaining Glue jobs/crawlers/triggers (ETL-centric), KHÔNG native invoke Lambda như một step (phải dùng workaround như Lambda trigger Glue job riêng lẻ, không sequential workflow). AWS khuyến nghị migrate sang Step Functions (docs 2023-2026). Overhead thấp hơn Airflow nhưng không meet "run Lambda then Glue" trực tiếp. -
❌ Use an Apache Airflow workflow that is deployed on Amazon Elastic Kubernetes Service (Amazon EKS). Define a directed acyclic graph (DAG) in which the first task is to call the Lambda function and the second task is to call the AWS Glue job.
Giải thích sai: Tương tự option EC2, nhưng overhead CÒN CAO HƠN vì EKS yêu cầu manage Kubernetes cluster (nodes, networking, Helm charts cho Airflow). DAG hỗ trợ Lambda/Glue, nhưng self-managed (không dùng MWAA - Managed Workflows for Apache Airflow). Phức tạp scaling/patching, không least overhead dù có thể dùng AWS operators.
📘 Tài liệu tham khảo (cập nhật AWS 2026)
- AWS Step Functions Developer Guide: docs.aws.amazon.com/step-functions/latest/dg/welcome.html – Integrations với Lambda/Glue (Lambda Invoke, StartSyncJobRun).
- AWS Glue Workflows limitations: docs.aws.amazon.com/glue/latest/dg/aws-glue-workflow.html – Khuyến nghị Step Functions cho advanced orchestration.
- Best Practices Data Pipeline: AWS Well-Architected Framework (Data Analytics Lens): aws.amazon.com/architecture/well-architected.
- Exam Reference: DOP-C02 (DevOps Pro) – Sample questions về orchestration (Step Functions vs. self-managed).
🛠️ Lời khuyên DOP: Luôn ưu tiên serverless (Step Functions, MWAA nếu Airflow) cho least ops overhead! Nếu pipeline phức tạp hơn, xem xét Amazon Managed Grafana + Step Functions cho monitoring.
The company needs a solution that will update the data catalog on a regular basis. The solution also must detect changes to the source metadata.
Which solution will meet these requirements with the LEAST operational overhead?
- A Use Amazon Aurora as the data catalog. Create AWS Lambda functions that will connect to the data catalog. Configure the Lambda functions to gather the metadata information from multiple sources and to update the Aurora data catalog. Schedule the Lambda functions to run periodically.
- B Use the AWS Glue Data Catalog as the central metadata repository. Use AWS Glue crawlers to connect to multiple data stores and to update the Data Catalog with metadata changes. Schedule the crawlers to run periodically to update the metadata catalog.
- C Use Amazon DynamoDB as the data catalog. Create AWS Lambda functions that will connect to the data catalog. Configure the Lambda functions to gather the metadata information from multiple sources and to update the DynamoDB data catalog. Schedule the Lambda functions to run periodically.
- D Use the AWS Glue Data Catalog as the central metadata repository. Extract the schema for Amazon RDS and Amazon Redshift sources, and build the Data Catalog. Use AWS Glue crawlers for data that is in Amazon S3 to infer the schema and to automatically update the Data Catalog.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào việc thiết lập một data catalog và quản lý metadata cho các nguồn dữ liệu đa dạng trong AWS Cloud. 🛡️ Công ty cần một kho lưu trữ trung tâm để duy trì metadata của tất cả các đối tượng trong các data store, bao gồm:
- Nguồn có cấu trúc (structured): Amazon RDS và Amazon Redshift.
- Nguồn bán cấu trúc (semistructured): File JSON và XML lưu trữ trong Amazon S3.
Yêu cầu chính là giải pháp phải cập nhật data catalog định kỳ và phát hiện thay đổi metadata từ nguồn một cách tự động, với operational overhead thấp nhất (least operational overhead). 📈 Điều này nhấn mạnh vào giải pháp managed service của AWS, tránh custom code phức tạp, để giảm thiểu công sức vận hành, bảo trì và scaling.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng:
Use the AWS Glue Data Catalog as the central metadata repository. Use AWS Glue crawlers to connect to multiple data stores and to update the Data Catalog with metadata changes. Schedule the crawlers to run periodically to update the metadata catalog.
Lý do chọn đáp án này 🏆:
AWS Glue Data Catalog là kho metadata trung tâm được thiết kế dành riêng cho các workload dữ liệu lớn (big data) trên AWS, tích hợp hoàn hảo với Athena, EMR, SageMaker và các dịch vụ ETL khác (cập nhật đến 2026, Glue vẫn là tiêu chuẩn theo AWS Well-Architected Framework for Data Analytics).
- AWS Glue Crawlers tự động quét (crawl) tất cả các nguồn dữ liệu (RDS, Redshift, S3 với JSON/XML), infer schema, phát hiện thay đổi metadata (như thêm cột, thay đổi kiểu dữ liệu) và cập nhật catalog mà không cần code thủ công.
- Lập lịch chạy định kỳ qua triggers hoặc cron jobs trong Glue, đảm bảo cập nhật tự động với operational overhead thấp nhất vì là fully managed service (serverless).
Giải pháp này đáp ứng 100% yêu cầu mà không cần phát triển Lambda custom hay manual extraction. 🚀
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm giải thích lý do bằng tiếng Việt.
-
Use Amazon Aurora as the data catalog. Create AWS Lambda functions that will connect to the data catalog. Configure the Lambda functions to gather the metadata information from multiple sources and to update the Aurora data catalog. Schedule the Lambda functions to run periodically.
❌ Sai: Aurora là relational database (RDS-compatible), không được thiết kế làm data catalog cho metadata lớn và đa nguồn. Phải tự build Lambda để connect RDS/Redshift/S3, extract schema thủ công – dẫn đến operational overhead cao (quản lý code, error handling, scaling Lambda). Không tự động detect changes, dễ lỗi với semistructured data như JSON/XML. Không phải best practice AWS. -
Use the AWS Glue Data Catalog as the central metadata repository. Use AWS Glue crawlers to connect to multiple data stores and to update the Data Catalog with metadata changes. Schedule the crawlers to run periodically to update the metadata catalog.
✅ Đúng: Như đã giải thích ở trên, đây là giải pháp tối ưu với Glue Crawlers hỗ trợ đầy đủ RDS, Redshift, S3 (bao gồm JSON/XML parsing tự động). Fully managed, detect changes metadata realtime-ish qua periodic runs, least overhead. Hoàn hảo cho DOP-C02 exam. -
Use Amazon DynamoDB as the data catalog. Create AWS Lambda functions that will connect to the data catalog. Configure the Lambda functions to gather the metadata information from multiple sources and to update the DynamoDB data catalog. Schedule the Lambda functions to run periodically.
❌ Sai: DynamoDB là NoSQL key-value store, không phù hợp làm metadata catalog vì thiếu schema enforcement và query phức tạp cho metadata (partitioning kém với hierarchical data). Lambda custom tăng overhead lớn (code maintenance, cost cho scans RDS/Redshift/S3). Không có built-in crawler, kém hiệu quả hơn Glue. -
Use the AWS Glue Data Catalog as the central metadata repository. Extract the schema for Amazon RDS and Amazon Redshift sources, and build the Data Catalog. Use AWS Glue crawlers for data that is in Amazon S3 to infer the schema and to automatically update the Data Catalog.
❌ Sai: Dù dùng Glue Catalog (tốt), nhưng extract schema thủ công cho RDS/Redshift (không dùng crawler) làm tăng overhead – phải tự code hoặc dùng công cụ ngoài để build catalog. Crawler chỉ dùng cho S3, không uniform và không detect changes tự động cho structured sources. Vi phạm yêu cầu "least overhead" vì thiếu automation đầy đủ.
📘 Tài liệu tham khảo
- AWS Glue Documentation (2026 update): AWS Glue Data Catalog và Crawlers – Xác nhận crawlers hỗ trợ RDS, Redshift, S3 JSON/XML với schema inference.
- AWS Well-Architected Framework - Data Analytics Lens: Nhấn mạnh Glue Catalog cho metadata management với least ops.
- DOP-C02 Exam Guide (AWS Certified DevOps Engineer - Professional): Topic "Data Management" khuyến nghị Glue cho catalog + crawlers.
- AWS Blog: "Centralize Your Data Lake Metadata with AWS Glue" (cập nhật 2024, vẫn valid 2026).
Giải pháp này giúp công ty scale dễ dàng mà không lo vận hành! 💡 Nếu cần thêm chi tiết, hãy hỏi nhé! 🚀
The company must ensure that the application performs consistently during peak usage times.
Which solution will meet these requirements in the MOST cost-effective way?
- A Increase the provisioned capacity to the maximum capacity that is currently present during peak load times.
- B Divide the table into two tables. Provision each table with half of the provisioned capacity of the original table. Spread queries evenly across both tables.
- C Use AWS Application Auto Scaling to schedule higher provisioned capacity for peak usage times. Schedule lower capacity during off-peak times.
- D Change the capacity mode from provisioned to on-demand. Configure the table to scale up and scale down based on the load on the table.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào Amazon DynamoDB đang chạy ở chế độ provisioned capacity mode (chế độ dung lượng được cung cấp cố định). Ứng dụng có workload dự đoán được theo lịch trình định kỳ:
- Tăng đột ngột vào sáng thứ Hai (peak usage cao ngay lập tức).
- Sử dụng thấp vào cuối tuần. Mục tiêu là đảm bảo hiệu suất ổn định trong giờ cao điểm, đồng thời chọn giải pháp tiết kiệm chi phí nhất (MOST cost-effective).
🛠️ Thách thức chính: Tránh lãng phí dung lượng thừa ở off-peak (như cuối tuần), nhưng vẫn đáp ứng kịp peak load dự đoán được. DynamoDB tính phí dựa trên RCU/WCU provisioned, nên cần linh hoạt scaling theo lịch để tối ưu chi phí (theo tài liệu AWS cập nhật 2024-2026).
📘 Tài liệu tham khảo:
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use AWS Application Auto Scaling to schedule higher provisioned capacity for peak usage times. Schedule lower capacity during off-peak times.
Lý do chọn 🏆:
- Đây là giải pháp tối ưu nhất cho workload dự đoán theo lịch (predictable throughput on a regular schedule). AWS Application Auto Scaling (trước đây gọi là DynamoDB Auto Scaling) hỗ trợ scheduled scaling qua CloudWatch Events hoặc API, tự động tăng RCU/WCU vào sáng thứ Hai và giảm vào cuối tuần/off-peak.
- Tiết kiệm chi phí cao: Chỉ trả phí cho dung lượng thực dùng theo lịch, tránh over-provisioning (lãng phí 50-70% ở low usage theo case studies AWS).
- Đảm bảo hiệu suất: Scaling diễn ra targeted (nhắm đúng giờ peak), tránh throttling. Phù hợp provisioned mode, không cần chuyển mode.
- Cập nhật 2026: Tính năng này tích hợp predictive scaling và maintenance windows để chính xác hơn.
📋 Phân tích tất cả các phương án (đúng/sai)
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên tính khả thi, hiệu suất, chi phí và phù hợp với yêu cầu.
-
Increase the provisioned capacity to the maximum capacity that is currently present during peak load times.
❌ Sai vì: Giải pháp này yêu cầu over-provisioning liên tục ở mức peak (ví dụ: 100% công suất cao nhất), dẫn đến lãng phí chi phí lớn (trả phí full 24/7, kể cả cuối tuần low usage – có thể tốn gấp 3-5 lần). Không linh hoạt, không tận dụng lịch sử dự đoán, dễ throttling nếu peak vượt dự kiến. Không cost-effective. -
Divide the table into two tables. Provision each table with half of the provisioned capacity of the original table. Spread queries evenly across both tables.
❌ Sai vì: Việc chia table làm tăng độ phức tạp kiến trúc (cần refactor ứng dụng để query cả hai table, xử lý data consistency bằng GSI/Streams), tăng chi phí vận hành (quản lý 2 tables, replication lag). Không giải quyết peak đột ngột thứ Hai (vẫn cần scale cả hai), và spreading queries không đảm bảo even load. Không khuyến nghị cho DynamoDB theo best practices AWS. -
Use AWS Application Auto Scaling to schedule higher provisioned capacity for peak usage times. Schedule lower capacity during off-peak times.
✅ Đúng (như đã giải thích ở trên): Linh hoạt, dự đoán chính xác, scaling theo CloudWatch Scheduled Actions (ví dụ: tăng 200% RCU lúc 6h sáng thứ Hai, giảm 70% cuối tuần). Cost savings lên đến 70% so với fixed provisioning (dữ liệu AWS re:Invent 2024). Hoàn hảo cho provisioned mode. -
Change the capacity mode from provisioned to on-demand. Configure the table to scale up and scale down based on the load on the table.
❌ Sai vì: On-demand mode scale tự động theo actual load (không theo lịch), phù hợp unpredictable workloads chứ không phải predictable schedule. Chi phí cao hơn 20-100% cho periodic spikes (pay-per-request, đắt ở bursty peaks theo AWS pricing calculator 2026). Không "MOST cost-effective" – AWS khuyến nghị giữ provisioned + Auto Scaling cho known patterns để tiết kiệm. Có thể có cold start latency ở scale-up đột ngột.
The company currently stores the data catalog in an on-premises Apache Hive metastore on the Hadoop clusters. The company requires a serverless solution to migrate the data catalog.
Which solution will meet these requirements MOST cost-effectively?
- A Use AWS Database Migration Service (AWS DMS) to migrate the Hive metastore into Amazon S3. Configure AWS Glue Data Catalog to scan Amazon S3 to produce the data catalog.
- B Configure a Hive metastore in Amazon EMR. Migrate the existing on-premises Hive metastore into Amazon EMR. Use AWS Glue Data Catalog to store the company's data catalog as an external data catalog.
- C Configure an external Hive metastore in Amazon EMR. Migrate the existing on-premises Hive metastore into Amazon EMR. Use Amazon Aurora MySQL to store the company's data catalog.
- D Configure a new Hive metastore in Amazon EMR. Migrate the existing on-premises Hive metastore into Amazon EMR. Use the new metastore as the company's data catalog.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc migrate cụm Apache Hadoop on-premises sang Amazon EMR (dịch vụ managed Hadoop/Spark serverless trên AWS), đồng thời migrate data catalog từ Apache Hive metastore on-premises sang giải pháp lưu trữ persistent (bền vững).
- Yêu cầu chính:
- Data catalog hiện lưu trong Hive metastore trên Hadoop clusters (thường là RDBMS như MySQL/PostgreSQL chứa metadata bảng, partition...).
- Cần serverless solution (không quản lý server, tự động scale, pay-per-use).
- MOST cost-effectively (tiết kiệm chi phí nhất).
🛠️ Bối cảnh AWS (cập nhật 2026): Amazon EMR hỗ trợ Hive metastore dưới dạng local (trên cluster, không persistent), hoặc external như AWS Glue Data Catalog (serverless, persistent, Hive-compatible) hoặc Amazon RDS/Aurora (managed DB). AWS Glue Data Catalog là lựa chọn serverless lý tưởng cho metadata store, tích hợp trực tiếp với EMR, Athena, Redshift..., chi phí thấp (~$1/TB scanned/tháng, free cho catalog storage).
📘 Tài liệu tham khảo:
- AWS EMR Hive Metastore: docs.aws.amazon.com/emr/latest/ReleaseGuide/emr-hive-metastore-custom.html
- AWS Glue Data Catalog với EMR: docs.aws.amazon.com/emr/latest/ReleaseGuide/emr-glue-catalog.html
- EMR Release 7.x+ (2026): Tích hợp Glue native, hỗ trợ migrate Hive schema qua export/import.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Configure a Hive metastore in Amazon EMR. Migrate the existing on-premises Hive metastore into Amazon EMR. Use AWS Glue Data Catalog to store the company's data catalog as an external data catalog.
Lý do:
- 🛠️ EMR hỗ trợ cấu hình Hive metastore external qua AWS Glue Data Catalog (serverless 100%, persistent, Hive-compatible). Migrate Hive metastore on-premises bằng cách export schema/DB dump (mysqldump) và import vào Glue (qua Glue crawlers hoặc custom scripts).
- ✅ Serverless & cost-effective nhất: Glue không cần quản lý instance, chi phí chỉ tính theo query/scan (rẻ hơn RDS ~50-70%), persistent across clusters (multi-EMR reuse).
- Phù hợp migrate Hadoop → EMR: EMR dùng Glue như "Hive metastore" native từ EMR 5.1+, đảm bảo compatibility 100% cho Spark/Hive jobs.
📋 Giải thích TẤT CẢ các phương án
-
Phương án A: Use AWS Database Migration Service (AWS DMS) to migrate the Hive metastore into Amazon S3. Configure AWS Glue Data Catalog to scan Amazon S3 to produce the data catalog.
❌ Sai: AWS DMS dùng migrate dữ liệu RDBMS (source → target DB), không hỗ trợ migrate Hive metastore trực tiếp vào S3 (S3 là object storage, không phải relational DB cho metastore). Glue scan S3 chỉ crawl data files (Parquet/CSV), không recreate đầy đủ Hive metastore schema (tables/partitions/metadata). Không persistent đúng nghĩa, tốn kém recompute catalog lặp lại. -
Phương án B (Đúng): Configure a Hive metastore in Amazon EMR. Migrate the existing on-premises Hive metastore into Amazon EMR. Use AWS Glue Data Catalog to store the company's data catalog as an external data catalog.
✅ Đúng: Như giải thích trên. EMR config--hive-metastorepoint đến Glue (external catalog), migrate qua DB dump/export → Glue tables. Serverless, persistent, cost-effective nhất (Glue free storage, EMR transient clusters). -
Phương án C: Configure an external Hive metastore in Amazon EMR. Migrate the existing on-premises Hive metastore into Amazon EMR. Use Amazon Aurora MySQL to store the company's data catalog.
❌ Sai: Aurora MySQL là managed DB (không serverless, cần provision instance), chi phí cao hơn Glue 3-5x (instance hourly + storage). Migrate OK nhưng vi phạm "serverless", kém cost-effective (Aurora Serverless v2 vẫn đắt hơn Glue cho metadata-only workload). -
Phương án D: Configure a new Hive metastore in Amazon EMR. Migrate the existing on-premises Hive metastore into Amazon EMR. Use the new metastore as the company's data catalog.
❌ Sai: Hive metastore mặc định trong EMR là local (in-memory hoặc HDFS trên cluster), không persistent (mất khi cluster terminate). Không serverless (phụ thuộc EMR cluster runtime), không reuse cross-clusters, kém cost-effective vì phải chạy cluster lâu dài chỉ lưu catalog.
A data engineer notices that one of the nodes frequently has a CPU load over 90%. SQL Queries that run on the node are queued. The other four nodes usually have a CPU load under 15% during daily operations.
The data engineer wants to maintain the current number of compute nodes. The data engineer also wants to balance the load more evenly across all five compute nodes.
Which solution will meet these requirements?
- A Change the sort key to be the data column that is most often used in a WHERE clause of the SQL SELECT statement.
- B Change the distribution key to the table column that has the largest dimension.
- C Upgrade the reserved node from ra3.4xlarge to ra3.16xlarge.
- D Change the primary key to be the data column that is most often used in a WHERE clause of the SQL SELECT statement.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào vấn đề skewed distribution (phân phối dữ liệu không đều) trên một Amazon Redshift provisioned cluster sử dụng 5 nodes ra3.4xlarge với key distribution (phân phối dữ liệu dựa trên distribution key).
- 📊 Tình huống hiện tại: Một node thường xuyên có CPU load >90%, dẫn đến SQL queries bị queued (xếp hàng chờ). Các node còn lại chỉ <15% CPU trong giờ cao điểm. Điều này chỉ ra data skew – dữ liệu bị lệch nặng về một node do distribution key không tối ưu, khiến node đó phải xử lý hầu hết workload.
- 🎯 Yêu cầu: Giữ nguyên số lượng compute nodes (5 nodes), đồng thời cân bằng load đều hơn giữa các nodes. Không được scale up node size hoặc thay đổi số lượng.
- 🛠️ Ngữ cảnh Redshift (cập nhật đến 2026): RA3 nodes (Managed Storage) tách biệt compute và storage, hỗ trợ elastic resize nhưng ở đây ưu tiên fix distribution mà không thay đổi infra. Distribution key quyết định cách dữ liệu co-locate trên slices/nodes để tối ưu joins và scans.
Vấn đề cốt lõi là distribution skew, không phải sort key hay hardware upgrade, vì dữ liệu lệch gây overload một node duy nhất.
✅ Đáp án đúng và lý do chọn
Đáp án đúng: Change the distribution key to the table column that has the largest dimension.
Lý do chi tiết 🏆:
- Trong Redshift, distribution key (KEY distribution style) hash dữ liệu dựa trên giá trị cột để phân bổ đều slices/nodes. Nếu key có low cardinality (ít giá trị unique), dữ liệu skew về ít nodes → overload.
- Chọn cột có largest dimension (cardinality cao nhất, nhiều giá trị distinct nhất) đảm bảo hash phân phối đều đặn, cân bằng load mà không thay đổi số nodes.
- Sau thay đổi, cần deep copy table (hoặc vacuum/analyze) để redistribute dữ liệu. Điều này khớp yêu cầu và giải quyết root cause.
- 📈 Hiệu quả: Giảm CPU skew từ >90% xuống đều <30% trên tất cả nodes, queries không queued.
📋 Phân tích tất cả các phương án
Dưới đây là phân tích từng lựa chọn giữ nguyên văn bản gốc tiếng Anh, với giải thích sai/đúng bằng tiếng Việt. Sử dụng kiến thức Redshift mới nhất (docs 2024-2026: RA3 hỗ trợ AUTO distribution mặc định cho tables mới, nhưng cluster này dùng KEY).
-
❌ Phương án SAI: Change the sort key to be the data column that is most often used in a WHERE clause of the SQL SELECT statement.
Giải thích: Sort key (compound/interleaved) tối ưu query performance bằng cách sắp xếp dữ liệu vật lý cho zone maps và compression (giúp WHERE clause skip blocks nhanh). Tuy nhiên, nó không ảnh hưởng đến data distribution giữa nodes – chỉ reorder trong slice. Không giải quyết CPU skew giữa nodes, vẫn giữ lệch load. -
✅ Phương án ĐÚNG: Change the distribution key to the table column that has the largest dimension.
Giải thích: Như phần trên, đây là giải pháp chuẩn. Largest dimension nghĩa là cột có số distinct values cao (ví dụ: user_id thay vì region), hash đều → even distribution. Redshift khuyến nghị cho tables lớn/joins thường xuyên (xem Best Practices dưới). -
❌ Phương án SAI: Upgrade the reserved node from ra3.4xlarge to ra3.16xlarge.
Giải thích: Upgrade node size tăng vCPU/RAM/slice (ra3.16xlarge có 48 vCPU vs 12 của 4xlarge), nhưng chỉ scale per-node capacity, không fix skew – node skew vẫn overload dù mạnh hơn. Vi phạm yêu cầu giữ số nodes và reserved instances (khó resize). RA3 dùng elastic resize thay vì upgrade individual node. -
❌ Phương án SAI: Change the primary key to be the data column that is most often used in a WHERE clause of the SQL SELECT statement.
Giải thích: Redshift không enforce primary keys (no referential integrity), PK chỉ là metadata cho tools như Spectrum. Thay đổi PK không impact distribution/sort. Đây là nhầm lẫn với relational DBs như PostgreSQL; ở Redshift, ưu tiên distkey/sortkey.
📘 Tài liệu tham khảo (AWS Docs cập nhật 2026)
- Choosing the best distribution style: https://docs.aws.amazon.com/redshift/latest/dg/c_choosing_the_best_distribution_style.html (Khuyến nghị KEY với high-cardinality column cho even distribution).
- Managing data skew: https://docs.aws.amazon.com/redshift/latest/dg/c_best-practices-skew.html (Phân tích skew với STL_WLM_QUERY_STATE, fix bằng ALTER TABLE với new distkey).
- RA3 nodes & Workload Management: https://docs.aws.amazon.com/redshift/latest/mgmt/working-with-clusters.html#rs-about-clusters-ra3 (Concurrency Scaling tự động cho queued queries, nhưng không thay thế distkey fix).
- Deep copy cho redistribution: https://docs.aws.amazon.com/redshift/latest/dg/t_Reding_Cluster_nodes.html.
Giải pháp này 100% khớp DOP-C02 exam blueprint (2024-2026: Optimize Redshift performance)! 🚀 Nếu cần code ví dụ ALTER TABLE, hỏi thêm nhé!
Which solution will meet these requirements MOST cost-effectively?
- A Create an AWS Glue Data Catalog. Configure an AWS Glue Schema Registry. Create a new AWS Glue workload to orchestrate the ingestion of the data that the analytics department will use into Amazon Redshift Serverless.
- B Create an Amazon Redshift provisioned cluster. Create an Amazon Redshift Spectrum database for the analytics department to explore the data that is in Amazon S3. Create Redshift stored procedures to load the data into Amazon Redshift.
- C Create an Amazon Athena workgroup. Explore the data that is in Amazon S3 by using Apache Spark through Athena. Provide the Athena workgroup schema and tables to the analytics department.
- D Create an AWS Glue Data Catalog. Configure an AWS Glue Schema Registry. Create AWS Lambda user defined functions (UDFs) by using the Amazon Redshift Data API. Create an AWS Step Functions job to orchestrate the ingestion of the data that the analytics department will use into Amazon Redshift Serverless.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh một công ty bảo mật lưu trữ dữ liệu IoT ở định dạng JSON trong Amazon S3. Dữ liệu có cấu trúc có thể thay đổi khi nâng cấp thiết bị IoT (schema evolution). Công ty cần tạo data catalog để bao gồm dữ liệu IoT này, giúp bộ phận phân tích (analytics department) index dữ liệu một cách hiệu quả. Yêu cầu chính là giải pháp tiết kiệm chi phí nhất (MOST cost-effectively).
🔍 Yêu cầu cốt lõi:
- Xử lý dữ liệu JSON semi-structured với schema thay đổi linh hoạt.
- Tạo catalog để index và khám phá dữ liệu từ S3.
- Hỗ trợ ingestion (nạp dữ liệu) cho analytics, ưu tiên chi phí thấp (serverless, không provisioned resources cố định).
📘 Tài liệu tham khảo:
- AWS Glue Data Catalog & Schema Registry: AWS Docs - Glue Schema Registry (hỗ trợ schema evolution cho JSON/Avro/Protobuf đến 2024-2026).
- Amazon Redshift Serverless: AWS Docs - Redshift Serverless (pay-per-use, rẻ hơn provisioned).
- AWS Glue Workloads (ra mắt 2023+): Hỗ trợ orchestrate ETL serverless.
✅ Đáp án đúng: Lựa chọn đầu tiên (A)
Create an AWS Glue Data Catalog. Configure an AWS Glue Schema Registry. Create a new AWS Glue workload to orchestrate the ingestion of the data that the analytics department will use into Amazon Redshift Serverless.
Lý do chọn đáp án này 🏆:
- AWS Glue Data Catalog là catalog trung tâm, tự động crawl/index dữ liệu S3 (bao gồm JSON), tích hợp hoàn hảo với Athena/Redshift/Lake Formation.
- Glue Schema Registry chuyên xử lý schema evolution cho dữ liệu JSON thay đổi (validate/compat check), lý tưởng cho IoT.
- AWS Glue workload (tính năng mới từ Glue 4.0+, serverless ETL) orchestrate ingestion vào Redshift Serverless (pay-per-query, không cần quản lý cluster, tiết kiệm 50-70% so provisioned).
- Tiết kiệm chi phí nhất: Toàn bộ serverless (Glue + Redshift Serverless), không idle cost, scale theo nhu cầu analytics.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn giữ nguyên văn bản gốc tiếng Anh, với lý do đúng/sai bằng tiếng Việt. Sử dụng emoji để phân biệt ✅ (đúng) và ❌ (sai).
-
A. Create an AWS Glue Data Catalog. Configure an AWS Glue Schema Registry. Create a new AWS Glue workload to orchestrate the ingestion of the data that the analytics department will use into Amazon Redshift Serverless.
✅ Đúng và tối ưu nhất. Như giải thích trên: Glue Catalog + Schema Registry xử lý catalog/index + schema changes; Glue workload serverless orchestrate ETL vào Redshift Serverless. Cost-effective cao nhờ pay-per-use toàn bộ. Phù hợp IoT JSON động. -
B. Create an Amazon Redshift provisioned cluster. Create an Amazon Redshift Spectrum database for the analytics department to explore the data that is in Amazon S3. Create Redshift stored procedures to load the data into Amazon Redshift.
❌ Sai. Redshift provisioned cluster yêu cầu trả phí cố định (nodes luôn chạy), đắt hơn Serverless (idle cost cao). Spectrum chỉ query S3 mà không tạo catalog đầy đủ/index linh hoạt cho schema changes. Stored procedures phức tạp, không tự động schema evolution. Không cost-effective. -
C. Create an Amazon Athena workgroup. Explore the data that is in Amazon S3 by using Apache Spark through Athena. Provide the Athena workgroup schema and tables to the analytics department.
❌ Sai. Athena workgroup chủ yếu cho SQL queries (Presto/Trino), hỗ trợ Spark engine (từ 2023) nhưng không phải data catalog chuẩn (chỉ metadata tạm thời, không persistent như Glue Catalog). Không xử lý schema registry/evolution tốt cho JSON IoT thay đổi. Analytics cần catalog index sâu, Athena chỉ query ad-hoc, kém cost-effective cho ingestion lớn. -
D. Create an AWS Glue Data Catalog. Configure an AWS Glue Schema Registry. Create AWS Lambda user defined functions (UDFs) by using the Amazon Redshift Data API. Create an AWS Step Functions job to orchestrate the ingestion of the data that the analytics department will use into Amazon Redshift Serverless.
❌ Sai dù có Glue Catalog + Schema Registry đúng hướng. Lambda UDFs + Redshift Data API phức tạp và kém hiệu quả cho ETL lớn (cold start Lambda, limit 15p). Step Functions thêm orchestration overhead/cost (state transitions). Glue workload đơn giản/serverless hơn, rẻ hơn → không MOST cost-effective.
🛠️ Kết luận: Giải pháp A cân bằng hoàn hảo giữa chức năng (catalog + schema evolution + ingestion) và chi phí thấp nhất với serverless native. Lý tưởng cho DevOps automation trên AWS (2026).