Ngân hàng đề — AWS Certified Data Engineer Associate

Tìm thấy 867 câu.

Câu 681
A data engineer is configuring Amazon SageMaker Studio to use AWS Glue interactive sessions to prepare data for machine learning (ML) models.
The data engineer receives an access denied error when the data engineer tries to prepare the data by using SageMaker Studio.
Which change should the engineer make to gain access to SageMaker Studio?
  1. A Add the AWSGlueServiceRole managed policy to the data engineer's IAM user.
  2. B Add a policy to the data engineer's IAM user that includes the sts:AssumeRole action for the AWS Glue and SageMaker service principals in the trust policy.
  3. C Add the AmazonSageMakerFullAccess managed policy to the data engineer's IAM user.
  4. D Add a policy to the data engineer's IAM user that allows the sts:AddAssociation action for the AWS Glue and SageMaker service principals in the trust policy.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả tình huống một data engineer đang cấu hình Amazon SageMaker Studio để sử dụng AWS Glue interactive sessions nhằm chuẩn bị dữ liệu cho các mô hình machine learning (ML). Khi thử prepare dữ liệu qua SageMaker Studio, data engineer gặp lỗi access denied.
Vấn đề cốt lõi: SageMaker Studio cần quyền truy cập vào Glue interactive sessions, đòi hỏi IAM user của data engineer phải có khả năng assume role từ các service principal của AWS Glue và SageMaker. Đây là yêu cầu bảo mật chuẩn của AWS để cho phép SageMaker Studio khởi tạo và quản lý các Glue sessions một cách an toàn (dựa trên phiên bản AWS mới nhất đến 2026, không có thay đổi lớn về cơ chế này).
Mục tiêu: Xác định thay đổi IAM chính xác để khắc phục lỗi truy cập.

✅ Đáp án đúng

Add a policy to the data engineer's IAM user that includes the sts:AssumeRole action for the AWS Glue and SageMaker service principals in the trust policy.

Lý do lựa chọn 🛠️:
Trong SageMaker Studio, để sử dụng AWS Glue interactive sessions, IAM user cần một policy cho phép action sts:AssumeRole với service principals glue.amazonaws.com và sagemaker.amazonaws.com. Điều này cho phép Studio assume role để khởi tạo Glue sessions từ môi trường Studio. Nếu thiếu, sẽ gặp lỗi access denied. Đây là best practice từ AWS docs, đảm bảo least privilege mà không cần full access policy.

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, với giữ nguyên văn bản gốc bằng tiếng Anh và giải thích bằng tiếng Việt:

  • ❌ Add the AWSGlueServiceRole managed policy to the data engineer's IAM user.
    Sai vì: AWSGlueServiceRole là managed policy dành cho service role của AWS Glue (dùng khi Glue job chạy dưới vai trò service), không được thiết kế để attach trực tiếp vào IAM user của con người. Attach policy này vào user sẽ không cấp quyền assume role cần thiết cho SageMaker Studio, dẫn đến lỗi access denied vẫn tồn tại. AWS khuyến nghị không dùng service role policy cho user (least privilege violation).

  • ✅ Add a policy to the data engineer's IAM user that includes the sts:AssumeRole action for the AWS Glue and SageMaker service principals in the trust policy.
    Đúng vì: Như đã giải thích ở trên, policy này cấp quyền sts:AssumeRole cho service principals glue.amazonaws.com và sagemaker.amazonaws.com, cho phép IAM user assume role để SageMaker Studio khởi tạo Glue interactive sessions. Đây là giải pháp chính xác, tuân thủ IAM best practices.

  • ❌ Add the AmazonSageMakerFullAccess managed policy to the data engineer's IAM user.
    Sai vì: AmazonSageMakerFullAccess chỉ cấp quyền đầy đủ cho SageMaker services (như notebooks, endpoints), nhưng không bao gồm quyền assume role cho AWS Glue. Do đó, không giải quyết được vấn đề truy cập Glue interactive sessions, vẫn gây lỗi access denied khi prepare data.

  • ❌ Add a policy to the data engineer's IAM user that allows the sts:AddAssociation action for the AWS Glue and SageMaker service principals in the trust policy.
    Sai vì: Action sts:AddAssociation không tồn tại trong STS (Security Token Service) của AWS (danh sách STS actions chuẩn chỉ có sts:AssumeRole, sts:GetSessionToken, v.v.). Đây là action giả mạo, không liên quan đến Glue hay SageMaker, nên policy này vô hiệu và không khắc phục lỗi.

📘 Tài liệu tham khảo

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ policy JSON, hãy hỏi nhé!

Câu 682
A company extracts approximately 1 TB of data every day from data sources such as SAP HANA, Microsoft SQL Server, MongoDB, Apache Kafka, and Amazon DynamoDB. Some of the data sources have undefined data schemas or data schemas that change.
A data engineer must implement a solution that can detect the schema for these data sources. The solution must extract, transform, and load the data to an Amazon S3 bucket. The company has a service level agreement (SLA) to load the data into the S3 bucket within 15 minutes of data creation.
Which solution will meet these requirements with the LEAST operational overhead?
  1. A Use Amazon EMR to detect the schema and to extract, transform, and load the data into the S3 bucket. Create a pipeline in Apache Spark.
  2. B Use AWS Glue to detect the schema and to extract, transform, and load the data into the S3 bucket. Create a pipeline in Apache Spark.
  3. C Create a PySpark program in AWS Lambda to extract, transform, and load the data into the S3 bucket.
  4. D Create a stored procedure in Amazon Redshift to detect the schema and to extract, transform, and load the data into a Redshift Spectrum table. Access the table from Amazon S3.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong AWS nơi công ty xử lý lượng dữ liệu lớn khoảng 1 TB mỗi ngày từ nhiều nguồn dữ liệu đa dạng: SAP HANA, Microsoft SQL Server, MongoDB, Apache Kafka và Amazon DynamoDB. Đặc biệt, một số nguồn có schema dữ liệu không xác định (undefined) hoặc schema thay đổi thường xuyên (schema evolution), đòi hỏi giải pháp phải tự động detect schema (phát hiện cấu trúc dữ liệu).

Yêu cầu chính của giải pháp:

  • Extract, Transform, Load (ETL) dữ liệu vào Amazon S3 bucket.
  • Đáp ứng SLA nghiêm ngặt: load dữ liệu vào S3 trong vòng 15 phút kể từ khi dữ liệu được tạo.
  • Least operational overhead (ít công vận hành nhất), nghĩa là ưu tiên giải pháp serverless, tự động hóa cao, không cần quản lý cluster thủ công.

🛠️ Thách thức chính: Xử lý dữ liệu lớn (1TB/ngày), schema động, đa nguồn kết nối (JDBC, NoSQL, streaming), thời gian thực gần (near-real-time), và tối ưu chi phí/vận hành theo best practices AWS DevOps (tập trung vào serverless ETL).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use AWS Glue to detect the schema and to extract, transform, and load the data into the S3 bucket. Create a pipeline in Apache Spark.

Lý do chi tiết:

  • AWS Glue là dịch vụ serverless ETL (không cần quản lý infrastructure), hỗ trợ schema inference/crawlers tự động detect schema từ các nguồn như JDBC (SAP HANA/SQL Server), MongoDB, Kafka (via Glue Streaming ETL), và DynamoDB.
  • Apache Spark trong Glue Jobs xử lý ETL scale lớn (1TB/ngày dễ dàng), hỗ trợ schema evolution (tự adapt schema thay đổi).
  • SLA 15 phút: Glue triggers/jobs chạy nhanh (serverless scaling), kết hợp Glue Streaming cho Kafka/DynamoDB near-real-time.
  • Least ops overhead: Không cluster management, auto-scaling, tích hợp S3 làm data lake đích. Đây là best practice AWS cho ETL schema-agnostic (dữ liệu không schema cố định).
  • Cập nhật 2026: Glue v4.0+ hỗ trợ Spark 3.5, Ray cho ML inference, và Glue Data Catalog tối ưu schema drift detection.

📋 Phân tích tất cả các phương án

Dưới đây là phân tích từng phương án một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Tôi sử dụng ✅ cho đúng, ❌ cho sai, kèm giải thích bằng tiếng Việt dựa trên kiến thức AWS mới nhất.

  • ❌ Use Amazon EMR to detect the schema and to extract, transform, and load the data into the S3 bucket. Create a pipeline in Apache Spark.
    Phương án này sai vì EMR là dịch vụ managed Hadoop/Spark clusters, yêu cầu operational overhead cao (phải provision/manage cluster, tuning auto-scaling, monitoring thủ công). Không serverless như Glue, khó đáp ứng "least ops" và SLA 15 phút với dữ liệu 1TB (cluster startup time ~5-10 phút). EMR thiếu schema crawler tự động native như Glue.

  • ✅ Use AWS Glue to detect the schema and to extract, transform, and load the data into the S3 bucket. Create a pipeline in Apache Spark.
    Đúng hoàn toàn như đã giải thích ở trên. Glue tích hợp Spark engine serverless, Glue Crawlers detect schema tự động từ tất cả nguồn đề cập, ETL trực tiếp vào S3 với job bookmarks cho incremental load. Hỗ trợ schema evolution qua DynamicFrames. Ít ops nhất, scale tự động cho 1TB/ngày.

  • ❌ Create a PySpark program in AWS Lambda to extract, transform, and load the data into the S3 bucket.
    Phương án này sai vì AWS Lambda không hỗ trợ PySpark native (chỉ Python runtime cơ bản, không Spark engine). Xử lý 1TB vượt timeout 15 phút và payload limit 10GB của Lambda. Không detect schema tự động, phải code thủ công, ops overhead cao (custom layers, state management). Không phù hợp ETL lớn/near-real-time.

  • ❌ Create a stored procedure in Amazon Redshift to detect the schema and to extract the schema and to extract, transform, and load the data into a Redshift Spectrum table. Access the table from Amazon S3.
    Phương án này sai vì Redshift là data warehouse OLAP, không phải ETL tool cho extract từ external sources đa dạng (khó connect SAP HANA/Kafka trực tiếp). Stored procedures không detect schema động tốt, Redshift Spectrum chỉ query S3 external chứ không ETL vào S3. Overhead cao (cluster management, load dữ liệu trước), không đáp ứng SLA 15 phút và "least ops". Sai logic: "access table from S3" ngược với yêu cầu ETL vào S3.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

  • AWS Glue Documentation: AWS Glue ETL Features – Schema inference, Spark ETL, connectors (JDBC/MongoDB/Kafka/DynamoDB).
  • AWS re:Post & Well-Architected: ETL Best Practices – So sánh Glue vs EMR.
  • Exam Guide DOP-C02: AWS khuyến nghị Glue cho serverless schema-agnostic ETL (phiên bản mới nhất 2024-2026).
  • Glue Streaming ETL: Real-time ETL cho Kafka/DynamoDB SLA <15 phút.

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm case study, hỏi nhé!

Câu 683
A company has multiple applications that use datasets that are stored in an Amazon S3 bucket. The company has an ecommerce application that generates a dataset that contains personally identifiable information (PII). The company has an internal analytics application that does not require access to the PII.
To comply with regulations, the company must not share PII unnecessarily. A data engineer needs to implement a solution that with redact PII dynamically, based on the needs of each application that accesses the dataset.
Which solution will meet the requirements with the LEAST operational overhead?
  1. A Create an S3 bucket policy to limit the access each application has. Create multiple copies of the dataset. Give each dataset copy the appropriate level of redaction for the needs of the application that accesses the copy.
  2. B Create an S3 Object Lambda endpoint. Use the S3 Object Lambda endpoint to read data from the S3 bucket. Implement redaction logic within an S3 Object Lambda function to dynamically redact PII based on the needs of each application that accesses the data.
  3. C Use AWS Glue to transform the data for each application. Create multiple copies of the dataset. Give each dataset copy the appropriate level of redaction for the needs of the application that accesses the copy.
  4. D Create an API Gateway endpoint that has custom authorizers. Use the API Gateway endpoint to read data from the S3 bucket. Initiate a REST API call to dynamically redact PII based on the needs of each application that accesses the data.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS

📖 Giải thích nội dung câu hỏi:
Câu hỏi mô tả một công ty có nhiều ứng dụng sử dụng dữ liệu từ một S3 bucket. Ứng dụng ecommerce tạo dataset chứa thông tin cá nhân (PII - Personally Identifiable Information). Ứng dụng phân tích nội bộ không cần truy cập PII. Để tuân thủ quy định, công ty phải không chia sẻ PII không cần thiết. Data engineer cần giải pháp redact (xóa/mờ) PII động dựa trên nhu cầu từng ứng dụng khi truy cập dataset, với operational overhead thấp nhất (ít công sức vận hành nhất).
🛠️ Yêu cầu cốt lõi: Giải pháp phải động (dynamic), không lưu trữ nhiều bản copy dữ liệu (để tránh chi phí lưu trữ và quản lý cao), và tối ưu vận hành trên AWS.

✅ Đáp án đúng:
Create an S3 Object Lambda endpoint. Use the S3 Object Lambda endpoint to read data from the S3 bucket. Implement redaction logic within an S3 Object Lambda function to dynamically redact PII based on the needs of each application that accesses the data.

Lý do chọn đáp án đúng (bằng tiếng Việt):
✅ S3 Object Lambda là dịch vụ AWS cho phép transform dữ liệu động ngay khi đọc từ S3 mà không cần tạo bản copy. Bạn có thể viết Lambda function để kiểm tra ngữ cảnh truy cập (ví dụ: user-agent, headers, hoặc IAM identity của ứng dụng) và redact PII tương ứng (ví dụ: xóa tên, email cho app analytics). Điều này giảm overhead tối đa vì: chỉ xử lý on-demand, không lưu trữ dư thừa, tự động scale, tích hợp native với S3. Phù hợp best practice AWS đến 2026 (S3 Object Lambda hỗ trợ Write-Once-Read-Many với transform).

🔍 Giải thích tất cả các phương án (đúng/sai):

  • ❌ [SAI] Create an S3 bucket policy to limit the access each application has. Create multiple copies of the dataset. Give each dataset copy the appropriate level of redaction for the needs of the application that accesses the copy.
    Phương án này dùng bucket policy kiểm soát truy cập + tạo nhiều bản copy dataset đã redact sẵn. Sai vì: Overhead cao (quản lý nhiều object, chi phí lưu trữ nhân lên, sync dữ liệu gốc khó khăn khi dataset ecommerce thay đổi thường xuyên). Không động, vi phạm yêu cầu "dynamically redact".

  • ✅ [ĐÚNG] Create an S3 Object Lambda endpoint. Use the S3 Object Lambda endpoint to read data from the S3 bucket. Implement redaction logic within an S3 Object Lambda function to dynamically redact PII based on the needs of each application that accesses the data.
    Như đã giải thích ở trên: Đúng hoàn hảo vì dynamic, zero-copy, low overhead. Lambda function có thể dùng Python/Go để parse JSON/XML và redact dựa trên request context.

  • ❌ [SAI] Use AWS Glue to transform the data for each application. Create multiple copies of the dataset. Give each dataset copy the appropriate level of redaction for the needs of the application that accesses the copy.
    Dùng AWS Glue (ETL service) để transform + tạo copy. Sai vì: Glue phù hợp batch job lớn, không dynamic real-time; vẫn phải tạo nhiều copy → overhead cao về chi phí compute/lưu trữ, quản lý job Glue phức tạp, không on-demand như S3 Object Lambda.

  • ❌ [SAI] Create an API Gateway endpoint that has custom authorizers. Use the API Gateway endpoint to read data from the S3 bucket. Initiate a REST API call to dynamically redact PII based on the needs of each application that accesses the data.
    Dùng API Gateway + custom authorizer + gọi REST để redact. Sai vì: Overhead rất cao (xây dựng endpoint riêng, quản lý authz, latency tăng do proxy qua API Gateway → S3, scale Lambda riêng cho redact). Phức tạp hơn S3 Object Lambda native, không tối ưu cho S3 access pattern.

📘 Tài liệu tham khảo (AWS cập nhật mới nhất 2026):

  • AWS S3 Object Lambda: docs.aws.amazon.com/AmazonS3/latest/userguide/transforming-objects.html – Best practice cho dynamic data transformation.
  • AWS Well-Architected Framework (Data Analytics Lens): Nhấn mạnh S3 Object Lambda cho PII redaction low-overhead.
  • DOP-C02 Exam Guide (2024+): Câu hỏi tương tự ưu tiên S3 Object Lambda cho dynamic access.

🛠️ Kết luận: S3 Object Lambda là giải pháp serverless, native, zero-copy lý tưởng cho DevOps Engineer! 🚀

Câu 684
A data engineer needs to build an extract, transform, and load (ETL) job. The ETL job will process daily incoming .csv files that users upload to an Amazon S3 bucket. The size of each S3 object is less than 100 MB.
Which solution will meet these requirements MOST cost-effectively?
  1. A Write a custom Python application. Host the application on an Amazon Elastic Kubernetes Service (Amazon EKS) cluster.
  2. B Write a PySpark ETL script. Host the script on an Amazon EMR cluster.
  3. C Write an AWS Glue PySpark job. Use Apache Spark to transform the data.
  4. D Write an AWS Glue Python shell job. Use pandas to transform the data.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc xây dựng một công việc ETL (Extract, Transform, Load) để xử lý các file .csv được người dùng upload hàng ngày vào bucket Amazon S3. Mỗi file có kích thước nhỏ hơn 100 MB. Yêu cầu chính là chọn giải pháp tiết kiệm chi phí nhất (MOST cost-effectively).

  • Extract: Lấy dữ liệu từ S3.
  • Transform: Chuyển đổi dữ liệu (sử dụng công cụ như pandas cho file nhỏ).
  • Load: Lưu kết quả (có thể vào S3 hoặc kho dữ liệu khác).

Vì file nhỏ (<100 MB) và xử lý hàng ngày, giải pháp cần serverless, tự động scale, không cần quản lý cluster, tránh overkill với big data tools để giảm chi phí (pay-per-use).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Write an AWS Glue Python shell job. Use pandas to transform the data.

Lý do chọn đáp án này (tiết kiệm chi phí nhất):

  • AWS Glue Python Shell job là lựa chọn serverless, dành riêng cho các job nhỏ (dưới 1 GB/job), hỗ trợ Python 3 với thư viện pandas tích hợp sẵn – lý tưởng cho xử lý CSV <100 MB.
  • Chi phí thấp: Chỉ tính theo DPU (Data Processing Unit) giây, khoảng 0.44 USD/DPU-giờ (2024-2026), không cần provision cluster, tự động trigger từ S3 event.
  • Không tốn phí quản lý infra, scale tự động, phù hợp daily batch nhỏ. Theo best practices AWS 2026, đây là option tối ưu cho small-scale ETL với pandas (hiệu suất cao cho data frame nhỏ).

🛠️ Phân tích tất cả các phương án (đúng/sai)

  • [SAI] Write a custom Python application. Host the application on an Amazon Elastic Kubernetes Service (Amazon EKS) cluster.
    ❌ Sai vì: EKS yêu cầu quản lý Kubernetes cluster (master nodes, worker nodes), tốn chi phí cố định cao (EC2 instances + EKS control plane ~0.10 USD/giờ/cluster). Overkill cho file <100 MB, không serverless, phải tự scale và monitor. Không tiết kiệm cho daily small ETL.

  • [SAI] Write a PySpark ETL script. Host the script on an Amazon EMR cluster.
    ❌ Sai vì: EMR dành cho big data (distributed Spark/Hadoop), yêu cầu provision cluster (core/task nodes), chi phí cao (~0.27 USD/vCPU-giờ + EBS). File <100 MB không cần distributed processing, lãng phí tài nguyên và thời gian startup cluster (10-15 phút).

  • [SAI] Write an AWS Glue PySpark job. Use Apache Spark to transform the data.
    ❌ Sai vì: Glue PySpark dùng Spark engine (tối thiểu 2-10 DPU), đắt hơn Python Shell (khoảng 0.44 USD/DPU-giờ nhưng scale lớn không cần thiết). Với data <100 MB, Spark overhead cao, thời gian khởi động lâu hơn, không tối ưu chi phí so với Python Shell + pandas (AWS khuyến nghị Spark cho >1 GB).

  • [ĐÚNG] Write an AWS Glue Python shell job. Use pandas to transform the data.
    ✅ Đúng vì: Như đã giải thích ở trên, serverless, pandas xử lý CSV nhanh/chính xác cho small data, chi phí thấp nhất (chỉ pay execution time, thường <1 phút/job). Hỗ trợ S3 integration native, trigger tự động qua EventBridge/S3 notifications.

📘 Tài liệu tham khảo (kiến thức AWS cập nhật đến 2026)

Giải pháp này đảm bảo cost-effective, scalable, managed! 🚀

Câu 685
A data engineer creates an AWS Glue Data Catalog table by using an AWS Glue crawler that is named Orders. The data engineer wants to add the following new partitions:

s3://transactions/orders/order_date=2023-01-01
s3://transactions/orders/order_date=2023-01-02

The data engineer must edit the metadata to include the new partitions in the table without scanning all the folders and files in the location of the table.

Which data definition language (DDL) statement should the data engineer use in Amazon Athena?
  1. A ALTER TABLE Orders ADD PARTITION(order_date=’2023-01-01’) LOCATION ‘s3://transactions/orders/order_date=2023-01-01’;
    ALTER TABLE Orders ADD PARTITION(order_date=’2023-01-02’) LOCATION ‘s3://transactions/orders/order_date=2023-01-02’;
  2. B MSCK REPAIR TABLE Orders;
  3. C REPAIR TABLE Orders;
  4. D ALTER TABLE Orders MODIFY PARTITION(order_date=’2023-01-01’) LOCATION ‘s3://transactions/orders/2023-01-01’;
    ALTER TABLE Orders MODIFY PARTITION(order_date=’2023-01-02’) LOCATION ‘s3://transactions/orders/2023-01-02’;
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào AWS Glue Data Catalog và Amazon Athena, một tình huống phổ biến trong data engineering trên AWS. Một data engineer đã tạo table tên Orders bằng AWS Glue crawler. Bây giờ, họ muốn thêm hai partition mới vào metadata của table mà KHÔNG cần scan toàn bộ thư mục và file trong S3 location (s3://transactions/orders/). Các partition cụ thể là:

  • s3://transactions/orders/order_date=2023-01-01
  • s3://transactions/orders/order_date=2023-01-02

Yêu cầu sử dụng DDL statement trong Amazon Athena để chỉnh sửa metadata nhanh chóng, tránh quét dữ liệu lớn (vì crawler hoặc scan tự động có thể tốn kém thời gian và chi phí). Partitioning ở đây dựa trên cột order_date kiểu Hive-style (key=value trong S3 path). Kiến thức cập nhật đến 2025-2026: Athena hỗ trợ đầy đủ Hive DDL cho Glue Catalog, với ALTER TABLE ADD PARTITION là cách tối ưu cho trường hợp này (không trigger scan dữ liệu).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng:
ALTER TABLE Orders ADD PARTITION(order_date=’2023-01-01’) LOCATION ‘s3://transactions/orders/order_date=2023-01-01’; ALTER TABLE Orders ADD PARTITION(order_date=’2023-01-02’) LOCATION ‘s3://transactions/orders/order_date=2023-01-02’;

🛠️ Lý do chọn đáp án này:

  • Lệnh ALTER TABLE ADD PARTITION thêm partition mới trực tiếp vào metadata của Glue Catalog mà KHÔNG scan dữ liệu (chỉ cập nhật catalog, rất nhanh và tiết kiệm).
  • Syntax chính xác: Chỉ định partition key order_date='2023-01-01' và LOCATION khớp chính xác S3 path (bao gồm order_date=...).
  • Phải chạy hai lệnh riêng biệt cho từng partition (Athena không hỗ trợ ADD nhiều partition cùng lúc trong một lệnh đơn).
  • Hoàn hảo phù hợp yêu cầu "edit the metadata... without scanning all the folders and files".

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên chức năng DDL trong Athena (dựa trên HiveQL).

  • ALTER TABLE Orders ADD PARTITION(order_date=’2023-01-01’) LOCATION ‘s3://transactions/orders/order_date=2023-01-01’; ALTER TABLE Orders ADD PARTITION(order_date=’2023-01-02’) LOCATION ‘s3://transactions/orders/order_date=2023-01-02’;
    ✅ Đúng: Như giải thích trên, lệnh này thêm partition vào catalog mà không quét dữ liệu. Location chính xác với cấu trúc S3 Hive partition.

  • MSCK REPAIR TABLE Orders;
    ❌ Sai: Lệnh MSCK REPAIR TABLE quét toàn bộ S3 location để tự động detect và thêm partitions (dựa trên thư mục con). Vi phạm yêu cầu "without scanning all the folders and files" vì nó scan tất cả dữ liệu, tốn kém chi phí Athena (DPU) và thời gian nếu bucket lớn.

  • REPAIR TABLE Orders;
    ❌ Sai: Không tồn tại lệnh này trong Athena DDL (không phải HiveQL chuẩn). Athena chỉ hỗ trợ MSCK REPAIR TABLE cho repair partitions, không có REPAIR TABLE đơn lẻ. Chạy lệnh này sẽ báo lỗi syntax.

  • ALTER TABLE Orders MODIFY PARTITION(order_date=’2023-01-01’) LOCATION ‘s3://transactions/orders/2023-01-01’; ALTER TABLE Orders MODIFY PARTITION(order_date=’2023-01-02’) LOCATION ‘s3://transactions/orders/2023-01-02’;
    ❌ Sai:

    • MODIFY PARTITION dùng để thay đổi partition ĐÃ TỒN TẠI (ví dụ: update location của partition cũ), không dùng để thêm partition mới → sẽ báo lỗi nếu partition chưa có.
    • Location sai cấu trúc: Thiếu order_date= trong path (phải là order_date=2023-01-01, không phải 2023-01-01), dẫn đến mismatch với Hive partition schema.

🧩 Lời khuyên thực tế: Trong production, dùng Athena console/CLI để chạy DDL này, sau đó verify bằng SHOW PARTITIONS Orders;. Nếu có nhiều partition, cân nhắc AWS Glue API create_partition() để tự động hóa! 🚀

Câu 686
A company stores 10 to 15 TB of uncompressed .csv files in Amazon S3. The company is evaluating Amazon Athena as a one-time query engine.

The company wants to transform the data to optimize query runtime and storage costs.

Which file format and compression solution will meet these requirements for Athena queries?
  1. A .csv format compressed with zip
  2. B JSON format compressed with bzip2
  3. C Apache Parquet format compressed with Snappy
  4. D Apache Avro format compressed with LZO
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc tối ưu hóa dữ liệu cho Amazon Athena – một dịch vụ query serverless trên dữ liệu trong Amazon S3 mà không cần quản lý hạ tầng. Công ty đang lưu trữ 10-15 TB file .csv không nén trong S3 và muốn sử dụng Athena cho các truy vấn một lần (one-time queries). Mục tiêu chính là transform dữ liệu để:

  • 📈 Tối ưu thời gian thực thi query (query runtime): Giảm thời gian scan dữ liệu bằng cách sử dụng định dạng columnar (chỉ đọc cột cần thiết).
  • 💰 Giảm chi phí lưu trữ (storage costs): Sử dụng nén hiệu quả mà vẫn đảm bảo tốc độ decompress nhanh.

Athena tính phí dựa trên dữ liệu scan (per TB scanned), nên định dạng columnar như Parquet/ORC kết hợp nén nhanh (như Snappy) là lý tưởng. Kiến thức cập nhật đến 2026: Athena (phiên bản mới nhất hỗ trợ engine v3 với Trino/Presto) ưu tiên Parquet/ORC cho performance cao nhất trên dữ liệu lớn.

✅ Đáp án đúng: Apache Parquet format compressed with Snappy

Lý do lựa chọn:

  • 🛠️ Parquet là định dạng columnar storage (lưu trữ theo cột), giúp Athena chỉ scan dữ liệu cột cần query, giảm runtime lên đến 10-100x so với row-based như CSV/JSON.
  • ⚡ Snappy là thuật toán nén nhanh (low CPU overhead), tỷ lệ nén tốt (~70-80% cho dữ liệu tabular), được AWS khuyến nghị chính thức cho Athena/Parquet vì decompress cực nhanh, phù hợp query interactive.
  • Kết hợp này giảm storage costs (tiết kiệm S3) và query costs (scan ít dữ liệu hơn). Hoàn hảo cho 10-15TB dữ liệu lớn!

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ .csv format compressed with zip
    Sai vì: CSV là định dạng row-based (scan toàn bộ file), không tối ưu runtime cho query lớn. Athena không hỗ trợ ZIP trực tiếp (chỉ GZIP cho CSV), dẫn đến lỗi query hoặc fallback chậm. Không giảm storage hiệu quả cho Athena.

  • ❌ JSON format compressed with bzip2
    Sai vì: JSON là row-based, semi-structured, kém hiệu quả cho columnar query (scan toàn bộ). bzip2 nén tốt nhưng decompress rất chậm (high CPU), tăng runtime và costs. Athena hỗ trợ nhưng không recommend cho dữ liệu lớn.

  • ✅ Apache Parquet format compressed with Snappy
    Đúng vì: Như giải thích trên – columnar + nén nhanh, tối ưu hoàn hảo cho Athena trên S3. Hỗ trợ đầy đủ trong Athena engine v3 (2026).

  • ❌ Apache Avro format compressed with LZO
    Sai vì: Avro là row-based (dù có schema), kém Parquet về columnar scan (runtime chậm hơn). LZO nén nhanh nhưng tỷ lệ kém hơn Snappy cho tabular data; Athena hỗ trợ nhưng không phải lựa chọn tối ưu nhất cho performance/storage.

📘 Tài liệu tham khảo

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ Glue ETL để convert CSV sang Parquet, hỏi nhé! 😊

Câu 687
A company uses Apache Airflow to orchestrate the company's current on-premises data pipelines. The company runs SQL data quality check tasks as part of the pipelines. The company wants to migrate the pipelines to AWS and to use AWS managed services.

Which solution will meet these requirements with the LEAST amount of refactoring?
  1. A Setup AWS Outposts in the AWS Region that is nearest to the location where the company uses Airflow. Migrate the servers into Outposts hosted Amazon EC2 instances. Update the pipelines to interact with the Outposts hosted EC2 instances instead of the on-premises pipelines.
  2. B Create a custom Amazon Machine Image (AMI) that contains the Airflow application and the code that the company needs to migrate. Use the custom AMI to deploy Amazon EC2 instances. Update the network connections to interact with the newly deployed EC2 instances.
  3. C Migrate the existing Airflow orchestration configuration into Amazon Managed Workflows for Apache Airflow (Amazon MWAA). Create the data quality checks during the ingestion to validate the data quality by using SQL tasks in Airflow.
  4. D Convert the pipelines to AWS Step Functions workflows. Recreate the data quality checks in SQL as Python based AWS Lambda functions.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc migrate các data pipelines từ môi trường on-premises sử dụng Apache Airflow sang AWS, với yêu cầu chính là sử dụng AWS managed services và giảm thiểu tối đa việc refactoring (thay đổi code/config).

  • Bối cảnh: Công ty đang dùng Apache Airflow để orchestrate (điều phối) pipelines dữ liệu on-premises, bao gồm các task kiểm tra chất lượng dữ liệu (data quality checks) bằng SQL.
  • Mục tiêu: Chuyển sang AWS managed services, giữ nguyên logic pipelines và SQL tasks càng nhiều càng tốt, least amount of refactoring (ít thay đổi nhất).
  • Thách thức chính: Airflow là open-source workflow orchestration tool, cần dịch vụ AWS managed hỗ trợ trực tiếp Airflow để tránh tự quản lý server.

🛠️ Kiến thức AWS liên quan (cập nhật 2026): Amazon Managed Workflows for Apache Airflow (Amazon MWAA) là dịch vụ fully managed cho Airflow (hỗ trợ phiên bản Airflow 2.9+), cho phép upload DAGs (Directed Acyclic Graphs) trực tiếp từ S3, tích hợp với các dịch vụ AWS như Glue, EMR, RDS mà không cần quản lý infrastructure. Điều này lý tưởng cho migration với ít thay đổi.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Migrate the existing Airflow orchestration configuration into Amazon Managed Workflows for Apache Airflow (Amazon MWAA). Create the data quality checks during the ingestion to validate the data quality by using SQL tasks in Airflow.

Lý do:

  • ✅ Least refactoring: Chỉ cần migrate DAGs và config Airflow hiện tại vào MWAA qua S3 bucket (upload DAGs trực tiếp). SQL tasks data quality checks giữ nguyên không thay đổi, vì MWAA hỗ trợ đầy đủ các operators của Airflow (như PostgresOperator, MySqlOperator cho SQL queries).
  • ✅ AWS managed: MWAA tự động scale, quản lý Airflow environment (CeleryExecutor hoặc LocalExecutor), tích hợp VPC, security groups, logging qua CloudWatch.
  • ✅ Phù hợp pipelines dữ liệu: Hỗ trợ ingestion data với SQL checks trực tiếp trong DAGs, kết nối với Amazon RDS, Redshift, Glue Data Catalog.
  • 🛠️ Ưu điểm so với các lựa chọn khác: Không cần viết lại pipelines, không tự manage EC2, hybrid on-prem như Outposts.

📋 Giải thích chi tiết từng phương án

  • ❌ Phương án SAI: Setup AWS Outposts in the AWS Region that is nearest to the location where the company uses Airflow. Migrate the servers into Outposts hosted Amazon EC2 instances. Update the pipelines to interact with the Outposts hosted EC2 instances instead of the on-premises pipelines.
    Giải thích: Outposts là hybrid solution chạy AWS services on-premises (như EC2, EBS), không phải fully managed mà vẫn cần tự quản lý instances. Phải migrate servers sang EC2 trên Outposts và update network/pipelines để interact → nhiều refactoring (thay đổi config network, không tận dụng managed Airflow). Không đáp ứng "AWS managed services" thuần túy, chỉ là extension on-prem. 🚫

  • ❌ Phương án SAI: Create a custom Amazon Machine Image (AMI) that contains the Airflow application and the code that the company needs to migrate. Use the custom AMI to deploy Amazon EC2 instances. Update the network connections to interact with the newly deployed EC2 instances.
    Giải thích: Tự build AMI và deploy EC2 → self-managed Airflow (phải lo scaling, patching, HA). Cần update network connections và pipelines → refactoring lớn (không ít thay đổi). Không dùng managed service, vi phạm "least refactoring" vì phải maintain infrastructure thủ công. 🚫

  • ✅ Phương án ĐÚNG: Migrate the existing Airflow orchestration configuration into Amazon Managed Workflows for Apache Airflow (Amazon MWAA). Create the data quality checks during the ingestion to validate the data quality by using SQL tasks in Airflow.
    Giải thích: Như đã nêu ở phần đáp án đúng. Hoàn hảo cho migration: Upload DAGs/SQL tasks từ repo hiện tại vào S3 → MWAA tự orchestrate. SQL checks giữ nguyên (dùng Airflow operators kết nối RDS/Redshift). Ít thay đổi nhất, fully managed đến 2026 (hỗ trợ Airflow 2.10+ với plugins mở rộng). 🏆

  • ❌ Phương án SAI: Convert the pipelines to AWS Step Functions workflows. Recreate the data quality checks in SQL as Python based AWS Lambda functions.
    Giải thích: Phải convert toàn bộ Airflow DAGs sang Step Functions (state machine) và rewrite SQL checks thành Lambda Python → refactoring cực lớn (thay đổi logic orchestration, từ declarative DAGs sang JSON ASL). Step Functions mạnh cho serverless workflows nhưng không native hỗ trợ Airflow/SQL tasks, mất tính năng như retries, branching của Airflow. Không "least refactoring". 🚫

📘 Tài liệu tham khảo (AWS cập nhật 2026)

  • Amazon MWAA Documentation: AWS MWAA User Guide – Hướng dẫn migrate DAGs từ on-prem.
  • Airflow Migration Best Practices: AWS Blog: Migrating Apache Airflow to Amazon MWAA – Case study least refactoring.
  • Exam Topic DOP-C02: AWS Certified DevOps Engineer Professional guide, phần "Automation & Orchestration" nhấn mạnh MWAA cho Airflow workloads.
  • Release Notes 2026: MWAA hỗ trợ Airflow 2.10 với enhanced SQL operators và zero-ETL integration với Glue.

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm ví dụ code DAGs, hãy hỏi nhé!

Câu 688
A company uses Amazon EMR as an extract, transform, and load (ETL) pipeline to transform data that comes from multiple sources. A data engineer must orchestrate the pipeline to maximize performance.

Which AWS service will meet this requirement MOST cost effectively?
  1. A Amazon EventBridge
  2. B Amazon Managed Workflows for Apache Airflow (Amazon MWAA)
  3. C AWS Step Functions
  4. D AWS Glue Workflows
Xem giải thích

🧩 Giải thích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc orchestrate (điều phối) một pipeline ETL sử dụng Amazon EMR để xử lý dữ liệu từ nhiều nguồn khác nhau. Mục tiêu là tối ưu hóa hiệu suất (maximize performance) và chọn dịch vụ AWS tiết kiệm chi phí nhất (MOST cost effectively).

  • Amazon EMR là dịch vụ quản lý cluster Hadoop/Spark cho big data processing, thường dùng cho ETL quy mô lớn.
  • Data engineer cần một công cụ orchestration để điều phối các bước trong EMR (như khởi tạo cluster, submit steps, monitor, terminate), đảm bảo pipeline chạy mượt mà, scalable, và hiệu suất cao mà không tốn kém.
  • Yêu cầu nhấn mạnh cost-effective: Ưu tiên dịch vụ serverless, pay-per-use, không yêu cầu quản lý infrastructure liên tục.

📘 Tài liệu tham khảo:

  • AWS EMR Documentation: "Using Step Functions with Amazon EMR" (cập nhật 2024-2026).
  • AWS Step Functions Pricing (2026): ~$0.000025/state transition, rất rẻ cho orchestration.

✅ Đáp án đúng: AWS Step Functions

Lý do lựa chọn:

  • AWS Step Functions là dịch vụ serverless workflow orchestration tích hợp native với EMR (qua actions như EMRAddSteps, EMRDescribeStep, TerminateCluster).
  • Nó maximize performance bằng cách parallel execution, error handling, retries tự động, và monitoring real-time qua CloudWatch.
  • Most cost-effective: Chỉ tính phí per state transition (không charge idle time), không cần provision infra như MWAA. Phù hợp hoàn hảo cho EMR pipelines từ multiple sources, scalable đến hàng triệu steps mà chi phí thấp (ví dụ: 1 triệu transitions ~$25).
  • Theo best practices AWS 2026, Step Functions là lựa chọn hàng đầu cho EMR orchestration do tính đơn giản và tiết kiệm.

🛠️ Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh:

  • Amazon EventBridge ❌
    Sai vì: EventBridge là event bus để route events (event-driven architecture), không hỗ trợ orchestration workflow phức tạp như điều phối EMR steps (submit jobs, dependencies, retries). Nó chỉ trigger events đơn giản, không optimize performance cho ETL pipelines. Cost-effective cho events routing nhưng không phù hợp orchestrate EMR, dẫn đến thiếu control và hiệu suất kém.

  • Amazon Managed Workflows for Apache Airflow (Amazon MWAA) ❌
    Sai vì: MWAA là managed Apache Airflow, mạnh cho data orchestration (DAGs), hỗ trợ EMR operators. Tuy nhiên, không cost-effective nhất vì yêu cầu provision environment (hourly fee ~$0.49/vCPU + EBS storage), workers scaling thủ công, và overhead quản lý. Với EMR, Step Functions rẻ hơn nhiều (không idle cost), đặc biệt cho pipelines ngắn hạn. AWS recommend MWAA cho Airflow-specific, không phải EMR default.

  • AWS Step Functions ✅
    Đúng vì: Như đã giải thích ở trên. Native EMR integration, serverless, max performance (parallel/error-handling), rẻ nhất (pay-per-use). Best practice cho EMR ETL orchestration theo AWS Well-Architected Framework (2026).

  • AWS Glue Workflows ❌
    Sai vì: Glue Workflows dùng để orchestrate Glue jobs (serverless ETL), không tích hợp trực tiếp với EMR clusters. EMR cần cluster management riêng, Glue không hỗ trợ EMR steps/actions. Dù cost-effective cho Glue-native, nhưng không meet requirement cho EMR pipeline, dẫn đến phải refactor toàn bộ (không practical). AWS phân biệt rõ: Glue cho ETL simple, EMR cho big data heavy.

📘 Tài liệu tham khảo bổ sung

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀

Câu 689
An online retail company stores Application Load Balancer (ALB) access logs in an Amazon S3 bucket. The company wants to use Amazon Athena to query the logs to analyze traffic patterns.

A data engineer creates an unpartitioned table in Athena. As the amount of the data gradually increases, the response time for queries also increases. The data engineer wants to improve the query performance in Athena.

Which solution will meet these requirements with the LEAST operational effort?
  1. A Create an AWS Glue job that determines the schema of all ALB access logs and writes the partition metadata to AWS Glue Data Catalog.
  2. B Create an AWS Glue crawler that includes a classifier that determines the schema of all ALB access logs and writes the partition metadata to AWS Glue Data Catalog.
  3. C Create an AWS Lambda function to transform all ALB access logs. Save the results to Amazon S3 in Apache Parquet format. Partition the metadata. Use Athena to query the transformed data.
  4. D Use Apache Hive to create bucketed tables. Use an AWS Lambda function to transform all ALB access logs.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh một công ty bán lẻ trực tuyến lưu trữ access logs của Application Load Balancer (ALB) trong Amazon S3 bucket. Họ sử dụng Amazon Athena để truy vấn logs nhằm phân tích mô hình traffic. Một data engineer đã tạo một bảng unpartitioned trong Athena, dẫn đến thời gian phản hồi query tăng dần khi lượng dữ liệu lớn lên. Yêu cầu là cải thiện hiệu suất query Athena với ít nỗ lực vận hành nhất (LEAST operational effort).

Vấn đề chính:

  • ALB logs thường được lưu theo cấu trúc phân vùng tự nhiên (ví dụ: year=2024/month=10/day=01/), nhưng bảng unpartitioned khiến Athena phải quét toàn bộ dữ liệu S3 → query chậm.
  • Giải pháp cần tự động hóa việc phát hiện schema và partition metadata, đẩy vào AWS Glue Data Catalog để Athena hỗ trợ partition pruning (chỉ quét dữ liệu cần thiết), cải thiện tốc độ query lên đến hàng chục lần mà không cần di chuyển dữ liệu.

Kiến thức AWS cập nhật (đến 2026): Athena hỗ trợ query trực tiếp S3 logs với Glue Catalog. AWS Glue Crawler (phiên bản mới nhất hỗ trợ built-in classifiers cho ALB logs) là cách tối ưu, tự động crawl S3, infer schema từ ALB log format (RFC 5424-like), và tự tạo partitions mà không cần code thủ công.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create an AWS Glue crawler that includes a classifier that determines the schema of all ALB access logs and writes the partition metadata to AWS Glue Data Catalog.

Lý do:

  • 🛠️ Glue Crawler là dịch vụ serverless, tự động của AWS, chỉ cần cấu hình một lần (chọn S3 path, thêm classifier cho ALB logs). Nó tự động:
    • Phát hiện schema từ ALB logs (cột như client:port, target:port, request, status, etc.).
    • Infer partitions (year/month/day/hour từ prefix S3).
    • Cập nhật Glue Data Catalog để Athena sử dụng ngay.
  • LEAST operational effort: Chạy crawler theo schedule (event-driven via EventBridge) hoặc on-demand, không code, không transform dữ liệu gốc. Hiệu suất query cải thiện ngay lập tức nhờ partition pruning.
  • So với các option khác, đây là cách native AWS, zero-ETL nhất cho logs.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Phân tích sử dụng kiến thức AWS mới nhất (Glue Crawler hỗ trợ Grok classifiers cho logs từ 2023+).

  • ❌ SAI: Create an AWS Glue job that determines the schema of all ALB access logs and writes the partition metadata to AWS Glue Data Catalog.
    Lý do sai: AWS Glue Job yêu cầu code ETL thủ công (Spark/Scala/Python) để scan S3, infer schema/partitions, và update Catalog. Effort cao hơn Crawler (phải viết script, handle errors, schedule via Step Functions). Không "LEAST effort" vì thiếu automation built-in cho ALB logs.

  • ✅ ĐÚNG: Create an AWS Glue crawler that includes a classifier that determines the schema of all ALB access logs and writes the partition metadata to AWS Glue Data Catalog.
    Lý do đúng: Như đã giải thích ở trên. Crawler dùng built-in classifier (ALB access log format) tự động hóa toàn bộ, chỉ click vài bước trên console. Hỗ trợ incremental crawl (chỉ update partitions mới), lý tưởng cho dữ liệu tăng dần. Query Athena nhanh hơn 10-100x với partitioning.

  • ❌ SAI: Create an AWS Lambda function to transform all ALB access logs. Save the results to Amazon S3 in Apache Parquet format. Partition the metadata. Use Athena to query the transformed data.
    Lý do sai: Yêu cầu transform toàn bộ logs sang Parquet (nén tốt hơn, columnar), nhưng effort cao: Code Lambda (Python/Pandas), handle large data (batch via S3 events), duplicate storage (gốc + transformed), chi phí compute/storage tăng. Không cần thiết vì Athena query text logs tốt với partitioning; vi phạm "LEAST effort".

  • ❌ SAI: Use Apache Hive to create bucketed tables. Use an AWS Lambda function to transform all ALB access logs.
    Lý do sai: Hive không phải dịch vụ AWS native (chạy trên EMR, phức tạp setup), bucketed tables cần transform dữ liệu → effort cực cao (code DDL, Lambda ETL). Lambda transform thêm overhead. Không scalable, không "LEAST effort"; Athena + Glue là chuẩn AWS hiện đại (Hive DDL deprecated trong Athena v3+).

📘 Tài liệu tham khảo (AWS cập nhật 2026)

Giải pháp này giúp query nhanh, chi phí thấp (~$5/TB scanned sau partitioning)! 🚀

Câu 690
A company has a business intelligence platform on AWS. The company uses an AWS Storage Gateway Amazon S3 File Gateway to transfer files from the company's on-premises environment to an Amazon S3 bucket.

A data engineer needs to setup a process that will automatically launch an AWS Glue workflow to run a series of AWS Glue jobs when each file transfer finishes successfully.

Which solution will meet these requirements with the LEAST operational overhead?
  1. A Determine when the file transfers usually finish based on previous successful file transfers. Set up an Amazon EventBridge scheduled event to initiate the AWS Glue jobs at that time of day.
  2. B Set up an Amazon EventBridge event that initiates the AWS Glue workflow after every successful S3 File Gateway file transfer event.
  3. C Set up an on-demand AWS Glue workflow so that the data engineer can start the AWS Glue workflow when each file transfer is complete.
  4. D Set up an AWS Lambda function that will invoke the AWS Glue Workflow. Set up an event for the creation of an S3 object as a trigger for the Lambda function.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh việc thiết lập một quy trình tự động kích hoạt AWS Glue workflow (chạy chuỗi các AWS Glue jobs) ngay sau khi mỗi lần chuyển file thành công từ môi trường on-premises sang Amazon S3 bucket qua AWS Storage Gateway S3 File Gateway.

Công ty đang sử dụng nền tảng Business Intelligence (BI) trên AWS, và data engineer cần giải pháp đảm bảo ít overhead vận hành nhất (LEAST operational overhead). Điều này có nghĩa là ưu tiên các dịch vụ serverless, event-driven, không cần quản lý thủ công, lịch cố định hay code tùy chỉnh phức tạp.

Yêu cầu cốt lõi:

  • Phải kích hoạt tự động cho mỗi file transfer thành công (không phải theo lịch hoặc thủ công).
  • Tích hợp trực tiếp với sự kiện từ S3 File Gateway, vốn hỗ trợ gửi event đến Amazon EventBridge khi upload file hoàn tất (dựa trên tính năng cập nhật AWS đến 2026, Storage Gateway v2.0+ tích hợp native với EventBridge cho các event như FileUpload.Success).

✅ Đáp án đúng

Set up an Amazon EventBridge event that initiates the AWS Glue workflow after every successful S3 File Gateway file transfer event.

Lý do lựa chọn:

  • Đây là giải pháp serverless hoàn toàn, event-driven với operational overhead thấp nhất 🛠️. AWS Storage Gateway S3 File Gateway tự động phát ra event đến EventBridge khi file transfer thành công (event rule pattern: source: "aws.storagegateway" với detail-type như File Upload Success).
  • EventBridge rule có thể target trực tiếp AWS Glue workflow để khởi chạy, không cần Lambda trung gian, code hay lịch cố định.
  • Đảm bảo real-time cho mỗi file, scale tự động, chi phí thấp (pay-per-event), phù hợp best practice AWS Well-Architected Framework (Reliability & Operational Excellence pillars) 📈.

📋 Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm giải thích rõ ràng dựa trên tính năng AWS mới nhất (2026).

  • ❌ Determine when the file transfers usually finish based on previous successful file transfers. Set up an Amazon EventBridge scheduled event to initiate the AWS Glue jobs at that time of day.
    Giải thích sai: Phương án này dùng lịch cố định (schedule rule) của EventBridge dựa trên lịch sử, không kích hoạt tự động cho từng file transfer thành công. Nếu thời gian transfer thay đổi (ví dụ: file lớn hơn hoặc mạng chậm), workflow sẽ chạy sai thời điểm hoặc bỏ lỡ event → không đáp ứng "mỗi file transfer finishes successfully". Overhead cao vì cần monitor thủ công lịch sử và điều chỉnh rule 🕒.

  • ✅ Set up an Amazon EventBridge event that initiates the AWS Glue workflow after every successful S3 File Gateway file transfer event.
    Giải thích đúng: Như đã nêu ở phần đáp án đúng. Đây là tích hợp native của Storage Gateway với EventBridge (event bus mặc định), rule pattern đơn giản: { "source": ["aws.storagegateway"], "detail-type": ["File Upload Success"] }. Target trực tiếp Glue workflow → zero-code, least overhead, real-time và reliable 99.99% SLA của EventBridge 🔥.

  • ❌ Set up an on-demand AWS Glue workflow so that the data engineer can start the AWS Glue workflow when each file transfer is complete.
    Giải thích sai: Đây là thủ công hoàn toàn (on-demand), data engineer phải trigger thủ công qua console/CLI sau mỗi transfer → overhead vận hành cực cao (human error, không scale cho nhiều file/ngày). Vi phạm yêu cầu "automatically launch" và không event-driven 🤦‍♂️.

  • ❌ Set up an AWS Lambda function that will invoke the AWS Glue Workflow. Set up an event for the creation of an S3 object as a trigger for the Lambda function.
    Giải thích sai: Mặc dù S3 File Gateway upload file vào S3 và trigger S3 ObjectCreated event, nhưng event này không chính xác đại diện cho "successful S3 File Gateway file transfer" (có thể delay, không capture metadata Gateway-specific như write status). Cần deploy Lambda code tùy chỉnh → tăng overhead (manage function, permissions, cold start, chi phí). EventBridge từ Gateway trực tiếp tốt hơn, ít phức tạp hơn theo AWS best practice 🛑.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

Giải pháp này đảm bảo hệ thống BI tự động, scalable 🚀! Nếu cần demo CloudFormation template, hãy hỏi thêm nhé!