Ngân hàng đề — AWS Certified Data Engineer Associate
Tìm thấy 867 câu.
Which solution will meet these requirements with the LEAST operational overhead?
- A Use an AWS Lambda function that includes both the business and the analytics logic to perform time-based aggregations over a window of up to 30 minutes for the data in Amazon Kinesis Data Streams.
- B Use Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) to analyze the data that might occasionally contain duplicates by using multiple types of aggregations.
- C Use an AWS Lambda function that includes both the business and the analytics logic to perform aggregations for a tumbling window of up to 30 minutes, based on the event timestamp.
- D Use Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) to analyze the data by using multiple types of aggregations to perform time-based analytics over a window of up to 30 minutes.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc quản lý việc ingestion dữ liệu streaming thời gian thực vào AWS, với yêu cầu chính là thực hiện phân tích thời gian thực (real-time analytics) sử dụng tổng hợp dựa trên thời gian (time-based aggregations) trong một cửa sổ thời gian lên đến 30 phút. Giải pháp phải chịu lỗi cao (highly fault tolerant) và có chi phí vận hành thấp nhất (LEAST operational overhead).
Dữ liệu streaming thường đến từ nguồn như Amazon Kinesis Data Streams (dù không chỉ rõ, nhưng ngầm định phổ biến). Thách thức lớn là xử lý windowing dài 30 phút một cách tự động, chịu lỗi (checkpointing, state management), mà không cần quản lý thủ công cluster hay server. AWS cung cấp các dịch vụ managed để giảm overhead, đặc biệt với streaming analytics đến năm 2026, nơi Amazon Managed Service for Apache Flink (tên mới của Kinesis Data Analytics for Apache Flink) là lựa chọn tối ưu cho windowed aggregations.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) to analyze the data by using multiple types of aggregations to perform time-based analytics over a window of up to 30 minutes.
Lý do lựa chọn (dựa trên phiên bản AWS mới nhất 2026):
- 🛠️ Dịch vụ này fully managed, hỗ trợ windowing linh hoạt (tumbling, sliding, session windows) lên đến hàng giờ mà không lo timeout.
- 📈 Hỗ trợ multiple aggregations (sum, count, avg...) và time-based analytics chính xác dựa trên event time.
- 🔄 Highly fault tolerant nhờ stateful processing với checkpointing tự động vào S3, exactly-once semantics.
- ⚡ Least operational overhead: Không cần provision cluster, auto-scale, chỉ config SQL hoặc Java/Scala/PyFlink app.
- So với Lambda (timeout 15 phút max), đây là giải pháp chuẩn cho streaming dài hạn.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng phương án một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên yêu cầu window 30 phút, fault tolerant, và least overhead.
-
Use an AWS Lambda function that includes both the business and the analytics logic to perform time-based aggregations over a window of up to 30 minutes for the data in Amazon Kinesis Data Streams.
❌ Sai: Lambda chỉ hỗ trợ runtime tối đa 15 phút (theo giới hạn AWS Lambda 2026), không thể giữ state cho window 30 phút. Việc tự implement windowing bằng Kinesis Streams yêu cầu custom buffering (DynamoDB/S3), tăng overhead lớn, không fault tolerant tự nhiên (retry manual). Không phù hợp real-time analytics phức tạp. -
Use Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) to analyze the data that might occasionally contain duplicates by using multiple types of aggregations.
❌ Sai: Mặc dù MSAF hỗ trợ aggregations và deduplication, phương án này đề cập duplicates không liên quan đến yêu cầu (câu hỏi không nhắc duplicates). Nó thiếu chi tiết về time-based window 30 phút, làm phương án không khớp chính xác. Overhead thấp nhưng không full-match requirements. -
Use an AWS Lambda function that includes both the business and the analytics logic to perform aggregations for a tumbling window of up to 30 minutes, based on the event timestamp.
❌ Sai: Tương tự phương án A, Lambda timeout 15 phút không hỗ trợ tumbling window 30 phút (cần state ngoài như DynamoDB). Phải tự code event timestamp buffering, dễ lỗi, không fault tolerant (không auto-checkpoint), overhead cao do quản lý state thủ công. Không khuyến nghị cho streaming dài. -
Use Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) to analyze the data by using multiple types of aggregations to perform time-based analytics over a window of up to 30 minutes.
✅ Đúng: Hoàn hảo khớp tất cả: Windowing lên 30 phút (hỗ trợ tumbling/sliding), multiple aggregations, real-time analytics, fault tolerant (checkpoints, auto-restart), least overhead (serverless, pay-per-use). MSAF xử lý event time watermarking tự động.
📘 Tài liệu tham khảo (cập nhật AWS 2026)
- Amazon Managed Service for Apache Flink: docs.aws.amazon.com/managed-flink/latest/java/windowing.html (windowing chi tiết) & aws.amazon.com/managed-service-apache-flink/ (fault tolerance).
- AWS Lambda Limits: docs.aws.amazon.com/lambda/latest/dg/configuration-function-common.html (timeout 15 phút max).
- AWS Streaming Best Practices: AWS Well-Architected Framework - Stream Processing Lens (2026 edition).
Giải pháp này đảm bảo serverless, scalable cho DevOps! 🚀
Which solution will meet these requirements with the LEAST operational overhead?
- A Create snapshots of the gp2 volumes. Create new gp3 volumes from the snapshots. Attach the new gp3 volumes to the EC2 instances.
- B Create new gp3 volumes. Gradually transfer the data to the new gp3 volumes. When the transfer is complete, mount the new gp3 volumes to the EC2 instances to replace the gp2 volumes.
- C Change the volume type of the existing gp2 volumes to gp3. Enter new values for volume size, IOPS, and throughput.
- D Use AWS DataSync to create new gp3 volumes. Transfer the data from the original gp2 volumes to the new gp3 volumes.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc nâng cấp Amazon EBS General Purpose SSD từ gp2 lên gp3 cho các EC2 instances đang hoạt động. 🔄
Công ty yêu cầu:
- Không gây gián đoạn (no interruptions) cho EC2 instances.
- Tránh mất dữ liệu (no data loss) trong quá trình migration.
- Least operational overhead (ít nỗ lực vận hành nhất, ưu tiên giải pháp đơn giản, tự động hóa cao).
Bối cảnh AWS cập nhật đến 2026: gp3 là thế hệ mới hơn gp2 (ra mắt 2020), cung cấp hiệu suất tốt hơn với IOPS/throughput độc lập, chi phí thấp hơn. AWS hỗ trợ in-place modification (thay đổi trực tiếp trên volume hiện tại) mà không detach volume, không downtime cho instances đang chạy. Điều này là tính năng chuẩn, không cần snapshot hay tool bên thứ ba. 🛠️
✅ Đáp án đúng và lý do chọn
Đáp án đúng: Change the volume type of the existing gp2 volumes to gp3. Enter new values for volume size, IOPS, and throughput.
Lý do:
- Phương án này sử dụng ModifyVolume API/Console để thay đổi type từ gp2 sang gp3 trực tiếp trên volume hiện tại (in-place), không detach, không downtime, không mất dữ liệu. ✅
- Có thể điều chỉnh size, IOPS (3,000-16,000 baseline), throughput (125-1,000 MB/s) cùng lúc.
- Least operational overhead: Chỉ 1 bước đơn giản qua AWS Console, CLI (
aws ec2 modify-volume), hoặc SDK. Hoàn tất trong vài phút, tự động. - Phù hợp best practice AWS cho migration gp2 → gp3 (không khuyến khích tạo volume mới để tránh phức tạp).
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt. Sử dụng ✅ cho đúng, ❌ cho sai.
-
Phương án A: Create snapshots of the gp2 volumes. Create new gp3 volumes from the snapshots. Attach the new gp3 volumes to the EC2 instances.
❌ Sai vì: Yêu cầu detach volume cũ trước khi attach mới → gây downtime/gián đoạn cho EC2 (dù ngắn). Tạo snapshot + volume mới → operational overhead cao (nhiều bước: snapshot, create volume, detach/attach, test data). Không phải least overhead, có rủi ro mất dữ liệu nếu snapshot fail. -
Phương án B: Create new gp3 volumes. Gradually transfer the data to the new gp3 volumes. When the transfer is complete, mount the new gp3 volumes to the EC2 instances to replace the gp2 volumes.
❌ Sai vì: Phải tạo volume mới + copy dữ liệu thủ công (gradually transfer) → overhead cực cao (thời gian dài, monitor sync, xử lý inconsistency). Detach/mount mới → downtime, rủi ro data loss nếu sync không hoàn hảo. Không hiệu quả so với in-place mod. -
Phương án C (Đúng): Change the volume type of the existing gp2 volumes to gp3. Enter new values for volume size, IOPS, and throughput.
✅ Đúng như đã giải thích ở trên: In-place, no downtime, least overhead. Hỗ trợ đầy đủ cho gp2 → gp3 (AWS xác nhận volumes attached to running instances vẫn mod được). -
Phương án D: Use AWS DataSync to create new gp3 volumes. Transfer the data from the original gp2 volumes to the new gp3 volumes.
❌ Sai vì: DataSync dùng cho transfer lớn giữa storage (S3/EFS/FSx), nhưng với EBS → không trực tiếp hỗ trợ, phải mount cả hai volume → phức tạp, overhead cao (setup agent, task, monitor). Vẫn cần detach/swap → downtime, không least overhead.
📘 Tài liệu tham khảo (AWS cập nhật mới nhất 2026)
- AWS EBS User Guide - Modify an EBS volume: docs.aws.amazon.com/ebs/latest/userguide/modify-volume.html → Xác nhận in-place gp2→gp3 no downtime.
- gp3 FAQs: aws.amazon.com/ebs/gp3/faqs → "Change gp2 to gp3 using ModifyVolume without detaching."
- AWS Well-Architected Framework - Reliability Pillar: Khuyến nghị in-place mod để minimize disruption.
- CLI Example:
aws ec2 modify-volume --volume-id vol-123 --volume-type gp3 --iops 3000 --throughput 125. 🛠️
Giải pháp này giúp migration mượt mà, tiết kiệm! Nếu cần demo CLI hoặc lab, hỏi thêm nhé. 🚀
Which solution will meet these requirements in the MOST operationally efficient way?
- A Create a view in the EC2 instance-based SQL Server databases that contains the required data elements. Create an AWS Glue job that selects the data directly from the view and transfers the data in Parquet format to an S3 bucket. Schedule the AWS Glue job to run every day.
- B Schedule SQL Server Agent to run a daily SQL query that selects the desired data elements from the EC2 instance-based SQL Server databases. Configure the query to direct the output .csv objects to an S3 bucket. Create an S3 event that invokes an AWS Lambda function to transform the output format from .csv to Parquet.
- C Use a SQL query to create a view in the EC2 instance-based SQL Server databases that contains the required data elements. Create and run an AWS Glue crawler to read the view. Create an AWS Glue job that retrieves the data and transfers the data in Parquet format to an S3 bucket. Schedule the AWS Glue job to run every day.
- D Create an AWS Lambda function that queries the EC2 instance-based databases by using Java Database Connectivity (JDBC). Configure the Lambda function to retrieve the required data, transform the data into Parquet format, and transfer the data into an S3 bucket. Use Amazon EventBridge to schedule the Lambda function to run every day.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào một tình huống di chuyển cơ sở dữ liệu (migration) từ các instance Amazon EC2 chạy Microsoft SQL Server sang Amazon RDS for SQL Server. Trong quá trình này, đội ngũ phân tích (analytics team) cần export dữ liệu lớn hàng ngày dưới dạng kết quả từ các phép SQL JOIN trên nhiều bảng, và lưu trữ dữ liệu đó vào Amazon S3 với định dạng Apache Parquet (định dạng columnar hiệu quả cho big data analytics). Yêu cầu chính là tìm giải pháp MOST operationally efficient (hiệu quả vận hành nhất), nghĩa là giải pháp phải serverless, scalable, ít quản lý thủ công, chi phí thấp và xử lý tốt dữ liệu lớn mà không cần can thiệp nhiều.
🛠️ Thách thức chính:
- Dữ liệu lớn từ SQL JOIN phức tạp trên EC2 SQL Server (nguồn dữ liệu tạm thời).
- Export hàng ngày đến khi migration hoàn tất.
- Chuyển sang Parquet (hỗ trợ compression tốt, columnar storage phù hợp với Athena/Glue).
- Tối ưu vận hành: Tránh self-managed servers, ưu tiên AWS managed services như Glue cho ETL.
📘 Kiến thức AWS cập nhật 2026: AWS Glue (với Spark 3.5+ và connectors JDBC mới nhất) hỗ trợ query trực tiếp từ SQL Server views qua JDBC, convert sang Parquet một cách native, serverless và scalable cho petabyte-scale data. Không cần VPC peering phức tạp nhờ Glue Connection.
✅ Đáp án đúng
Đáp án đúng là lựa chọn đầu tiên:
Create a view in the EC2 instance-based SQL Server databases that contains the required data elements. Create an AWS Glue job that selects the data directly from the view and transfers the data in Parquet format to an S3 bucket. Schedule the AWS Glue job to run every day.
Lý do chọn đáp án này 🏆:
Giải pháp này hiệu quả vận hành nhất vì:
- Tạo VIEW trong SQL Server trên EC2 để encapsulate logic JOIN phức tạp, giúp query đơn giản hóa (Glue chỉ cần SELECT * FROM VIEW).
- AWS Glue Job (ETL serverless trên Apache Spark) query trực tiếp từ VIEW qua JDBC connector (hỗ trợ SQL Server native), transform và write Parquet vào S3 một bước duy nhất, scalable cho dữ liệu lớn (auto-scale DPUs).
- Schedule hàng ngày qua Glue Triggers hoặc EventBridge, fully managed, không lo scaling, monitoring tự động qua CloudWatch.
- Tiết kiệm chi phí/ops: Không cần Lambda timeout/memory limit, không crawler thừa, chỉ pay-per-job. Phù hợp migration scenario.
Nguồn tham khảo:
- AWS Glue ETL Docs (2026): https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-connect.html#glue-pyspark-jdbc-connect
- AWS Glue JDBC Connections: https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-connect-jdbc.html
📋 Phân tích tất cả các phương án
-
Phương án 1 (✅ ĐÚNG):
Create a view in the EC2 instance-based SQL Server databases that contains the required data elements. Create an AWS Glue job that selects the data directly from the view and transfers the data in Parquet format to an S3 bucket. Schedule the AWS Glue job to run every day.
Giải thích: Như phần trên, đây là giải pháp tối ưu nhất với Glue ETL native cho relational-to-Parquet pipeline. Hỗ trợ pushdown predicates, partitioning tự động, và dynamic frames cho dữ liệu lớn. Không overhead từ intermediate steps. Operationally efficient cao nhất cho daily exports. -
Phương án 2 (❌ SAI):
Schedule SQL Server Agent to run a daily SQL query that selects the desired data elements from the EC2 instance-based SQL Server databases. Configure the query to direct the output .csv objects to an S3 bucket. Create an S3 event that invokes an AWS Lambda function to transform the output format from .csv to Parquet.
Giải thích: Không efficient vì SQL Server Agent trên EC2 là self-managed (cần maintain EC2, patching, scaling), export CSV kém hiệu quả (row-based, không compress tốt như Parquet). Lambda convert CSV→Parquet không scale cho dữ liệu lớn (15min timeout, 10GB memory limit 2026), dễ fail với joins phức tạp. Thêm S3 event overhead, tăng chi phí và latency. -
Phương án 3 (❌ SAI):
Use a SQL query to create a view in the EC2 instance-based SQL Server databases that contains the required data elements. Create and run an AWS Glue crawler to read the view. Create an AWS Glue job that retrieves the data and transfers the data in Parquet format to an S3 bucket. Schedule the AWS Glue job to run every day.
Giải thích: Tạo VIEW tốt, nhưng Glue Crawler không cần thiết và không efficient cho scenario này. Crawler dùng để infer schema từ data lakes/files (S3/JDBC tables), không phải daily dynamic views (chỉ snapshot metadata một lần). Job vẫn phải query lại DB, gây double-work (crawl + job), tăng chi phí DPUs và thời gian. AWS recommend direct query cho scheduled ETL thay vì crawler. -
Phương án 4 (❌ SAI):
Create an AWS Lambda function that queries the EC2 instance-based databases by using Java Database Connectivity (JDBC). Configure the Lambda function to retrieve the required data, transform the data into Parquet format, and transfer the data into an S3 bucket. Use Amazon EventBridge to schedule the Lambda function to run every day.
Giải thích: Lambda không phù hợp cho dữ liệu lớn từ SQL JOIN (JDBC pull toàn bộ data vào memory, dễ OOM với big datasets). Transform Parquet cần thư viện như Arrow/Parquet libs phức tạp trong Java/Python runtime. Timeout 15p, concurrency limit, không parallelize tốt như Spark. EventBridge schedule OK nhưng toàn bộ pipeline không scalable/operationally efficient, phải custom error handling.
Kết luận 🎯: Phương án 1 tận dụng AWS Glue – dịch vụ ETL managed tốt nhất cho relational extract + Parquet write, phù hợp DevOps best practices (IaC với Glue Studio/CloudFormation). Các phương án khác thêm complexity hoặc limits, vi phạm "MOST operationally efficient".
Which table views should the data engineer use to meet this requirement?
- A STL_USAGE_CONTROL
- B STL_ALERT_EVENT_LOG
- C STL_QUERY_METRICS
- D STL_PLAN_INFO
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào Amazon Redshift – một kho dữ liệu (data warehouse) được đội ngũ data engineering sử dụng cho báo cáo hoạt động (operational reporting). Vấn đề chính là ngăn chặn các vấn đề hiệu suất (performance issues) do các truy vấn chạy lâu (long-running queries). Data engineer cần chọn một system table (bảng hệ thống) trong Redshift để ghi nhận các bất thường (anomalies) khi query optimizer (tối ưu hóa truy vấn) phát hiện các điều kiện có thể dẫn đến vấn đề hiệu suất.
Mục tiêu chính: Theo dõi và ghi log các cảnh báo từ query optimizer để phát hiện sớm các vấn đề tiềm ẩn, giúp tối ưu hóa hiệu suất cluster Redshift. Đây là tính năng query alert events trong Redshift, giúp phát hiện anomalies như skew, spill to disk, hoặc các vấn đề phân phối dữ liệu.
📘 Kiến thức cập nhật (đến 2026): Theo tài liệu AWS Redshift mới nhất (Redshift RA3 nodes, concurrency scaling, và query insights trong Amazon Redshift), các system tables như STL_* vẫn được sử dụng để monitor performance. Query optimizer tự động generate alerts cho anomalies.
✅ Đáp án đúng: STL_ALERT_EVENT_LOG
Lý do lựa chọn:
- Bảng này chính xác ghi nhận các sự kiện cảnh báo (alert events) được tạo bởi query optimizer khi phát hiện các điều kiện bất thường có thể gây vấn đề hiệu suất, chẳng hạn như: distribution skew, sortkey skew, broadcast skew, hoặc spill to disk.
- Nó lưu trữ thông tin chi tiết về query ID, event time, alert type, và recommendations, giúp data engineer phân tích và khắc phục nhanh chóng.
- Phù hợp hoàn hảo với yêu cầu "record anomalies when a query optimizer identifies conditions that might indicate performance issues".
- Dẫn nguồn: AWS Documentation - STL_ALERT_EVENT_LOG (cập nhật 2024-2026, vẫn là standard table cho query alerts).
🛠️ Phân tích tất cả các phương án
-
❌ STL_USAGE_CONTROL
Bảng này ghi nhận các hành động liên quan đến giới hạn sử dụng (usage limits) như query queue timeouts hoặc concurrency limits, không phải anomalies từ query optimizer. Nó theo dõi control actions (ví dụ: cancel query do vượt quota), chứ không phải performance anomalies như skew hoặc optimizer alerts. Không phù hợp vì không tập trung vào query optimization issues. -
✅ STL_ALERT_EVENT_LOG
Đúng 100% như đã giải thích ở trên. Đây là bảng chuyên biệt cho alert events từ query optimizer, bao gồm các loại alert như Compaction, Corrupted, Skew, Spill,... Giúp detect và resolve performance issues proactively. Hoàn hảo cho yêu cầu. -
❌ STL_QUERY_METRICS
Bảng này cung cấp metrics tổng hợp cho queries (như CPU time, I/O, execution time theo slice), dùng để phân tích performance sau khi query chạy. Không ghi nhận anomalies hoặc alerts từ optimizer, mà chỉ là dữ liệu thống kê. Không đáp ứng yêu cầu về "query optimizer identifies conditions". -
❌ STL_PLAN_INFO
Bảng này lưu trữ thông tin query execution plan (như steps, operators, và plan nodes), giúp debug query plans phức tạp. Không phải là log cho anomalies hoặc performance alerts từ optimizer, mà chỉ mô tả plan tree. Không liên quan trực tiếp đến việc record optimizer-detected issues.
📘 Tài liệu tham khảo bổ sung
- Amazon Redshift System Tables Reference – Tổng hợp tất cả STL_* tables.
- Monitoring Query Performance in Redshift – Hướng dẫn sử dụng STL_ALERT_EVENT_LOG cho alerts (cập nhật với Redshift Serverless 2026).
- Mẹo thực tế: Sử dụng query
SELECT * FROM stl_alert_event_log WHERE event = 'skew';để check anomalies cụ thể! 🚀
Which solution will meet these requirements MOST cost-effectively?
- A Use an AWS Glue PySpark job to ingest the source data into the data lake in .csv format.
- B Create an AWS Glue extract, transform, and load (ETL) job to read from the .csv structured data source. Configure the job to ingest the data into the data lake in JSON format.
- C Use an AWS Glue PySpark job to ingest the source data into the data lake in Apache Avro format.
- D Create an AWS Glue extract, transform, and load (ETL) job to read from the .csv structured data source. Configure the job to write the data into the data lake in Apache Parquet format.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào việc ingest dữ liệu structured từ file .csv (có 15 cột) vào Amazon S3 data lake, sao cho data analysts có thể chạy query trên Amazon Athena chỉ với 1-2 cột, và họ hiếm khi query toàn bộ file. Yêu cầu chính là giải pháp cost-effective nhất (tiết kiệm chi phí nhất).
📘 Bối cảnh quan trọng:
- Amazon Athena tính phí dựa trên lượng dữ liệu scan (per TB scanned), nên cần giảm thiểu dữ liệu scan không cần thiết.
- File .csv gốc là row-based (dữ liệu theo hàng), dẫn đến scan toàn bộ file dù chỉ query 1 cột → tốn kém.
- Giải pháp lý tưởng phải chuyển đổi dữ liệu sang định dạng columnar (như Parquet) để Athena chỉ scan cột cần thiết, kết hợp compression và partitioning để tối ưu chi phí (theo best practices AWS đến 2026, với Athena engine version 3 hỗ trợ columnar formats tốt hơn).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create an AWS Glue extract, transform, and load (ETL) job to read from the .csv structured data source. Configure the job to write the data into the data lake in Apache Parquet format.
Lý do 🛠️:
- Apache Parquet là định dạng columnar storage (lưu trữ theo cột), hỗ trợ compression cao (Snappy/GZIP) và column pruning (Athena tự động bỏ qua cột không query).
- Với query chỉ 1-2/15 cột, Athena scan ít dữ liệu hơn đáng kể so với .csv (giảm đến 90% chi phí scan theo AWS benchmarks 2025-2026).
- AWS Glue ETL job dễ dàng read .csv từ source và write Parquet vào S3, hỗ trợ partitioning (ví dụ theo ngày/cột phổ biến) để tối ưu thêm.
- Đây là best practice cho S3 data lake với Athena: cost-effective, scalable, không cần infrastructure thủ công (Glue serverless đến 2026).
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, đánh dấu ✅ đúng hoặc ❌ sai. Giữ nguyên văn bản gốc, chỉ giải thích bằng tiếng Việt:
-
❌ Use an AWS Glue PySpark job to ingest the source data into the data lake in .csv format.
Phương án này chỉ copy .csv gốc vào S3 mà không transform. Athena vẫn phải scan toàn bộ file (15 cột) dù query 1-2 cột → chi phí cao, không tận dụng columnar pruning. Không cost-effective, vi phạm yêu cầu "rarely query entire file". -
❌ Create an AWS Glue extract, transform, and load (ETL) job to read from the .csv structured data source. Configure the job to ingest the data into the data lake in JSON format.
JSON là row-based (tương tự .csv), không hỗ trợ columnar storage tốt. Athena vẫn scan gần hết file → chi phí tương đương .csv, không giảm scan dữ liệu thừa. Glue ETL thừa thãi nếu chỉ convert sang JSON kém hiệu quả hơn. -
❌ Use an AWS Glue PySpark job to ingest the source data into the data lake in Apache Avro format.
Avro là row-based (dù hỗ trợ schema evolution), Athena không prune cột hiệu quả như Parquet → scan nhiều dữ liệu không cần, chi phí cao. PySpark job có thể dùng nhưng định dạng Avro không tối ưu cho query selective columns trên Athena. -
✅ Create an AWS Glue extract, transform, and load (ETL) job to read from the .csv structured data source. Configure the job to write the data into the data lake in Apache Parquet format.
Như đã giải thích ở trên: Columnar format lý tưởng, Glue ETL đơn giản, giảm chi phí Athena tối đa nhờ column pruning + compression. Hoàn hảo cho data lake với query partial columns.
📘 Tài liệu tham khảo (cập nhật đến 2026)
- 🛠️ AWS Athena Best Practices: Use Columnar Formats like Parquet/ORC – Giải thích column pruning và chi phí scan.
- 🧩 AWS Glue ETL for Data Lakes: CSV to Parquet – Hướng dẫn transform sang Parquet.
- 📘 Amazon S3 Data Lakes with Athena: Cost Optimization (2025 Update) – Benchmarks giảm 75-90% chi phí với Parquet.
- ✅ AWS Certified Data Engineer - Associate Exam Guide (2026) – Bao gồm DOP-C02 với data lake patterns.
Giải pháp này đảm bảo tuân thủ AWS Well-Architected Framework: Cost Optimization Pillar! 🚀
A data engineering team needs to limit access to the records. Each HR department should be able to access records for only employees who are within the HR department's Region.
Which combination of steps should the data engineering team take to meet this requirement with the LEAST operational overhead? (Choose two.)
- A Use data filters for each Region to register the S3 paths as data locations.
- B Register the S3 path as an AWS Lake Formation location.
- C Modify the IAM roles of the HR departments to add a data filter for each department's Region.
- D Enable fine-grained access control in AWS Lake Formation. Add a data filter for each Region.
- E Create a separate S3 bucket for each Region. Configure an IAM policy to allow S3 access. Restrict access based on Region.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào việc hạn chế truy cập dữ liệu nhân sự (employee records) lưu trữ trong data lake dựa trên Amazon S3, nơi có 5 văn phòng ở các AWS Regions khác nhau. Mỗi văn phòng có bộ phận HR sử dụng IAM role riêng biệt, và yêu cầu là HR chỉ truy cập được records của nhân viên trong Region của mình. Nhóm data engineering cần giải pháp với ít overhead vận hành nhất (LEAST operational overhead), chọn TWO steps kết hợp.
🛠️ Bối cảnh chính: Sử dụng AWS Lake Formation để quản lý truy cập fine-grained (chi tiết) trên S3 data lake, thay vì IAM policies thô hoặc bucket riêng lẻ. Lake Formation giúp phân quyền theo metadata (như Region), tránh copy dữ liệu và giảm chi phí quản lý cross-Region.
✅ Đáp án đúng (Chọn TWO)
Hai bước đúng là:
Register the S3 path as an AWS Lake Formation location.
Enable fine-grained access control in AWS Lake Formation. Add a data filter for each Region.
Lý do lựa chọn 📈:
- Đây là cách chuẩn và ít overhead nhất theo best practices AWS Lake Formation (cập nhật đến 2026). Đầu tiên, register S3 path làm Lake Formation location để Lake Formation quản lý metadata và quyền trên dữ liệu S3 gốc (không cần di chuyển dữ liệu). Sau đó, enable FGAC (fine-grained access control) và thêm data filters cho từng Region để HR role chỉ thấy records khớp filter (ví dụ: filter WHERE region = 'us-east-1').
- Giải pháp này tự động hóa permissions, hỗ trợ cross-Region mà không cần sửa IAM roles trực tiếp hoặc tạo bucket riêng, giảm rủi ro và effort bảo trì. Lake Formation tích hợp chặt chẽ với S3, Glue Catalog, hỗ trợ query engines như Athena/Redshift Spectrum.
🔍 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, với ✅ đúng hoặc ❌ sai, giữ nguyên văn bản gốc tiếng Anh:
-
❌ Use data filters for each Region to register the S3 paths as data locations.
Sai vì data filters chỉ áp dụng sau khi S3 path đã được register làm Lake Formation location. Không thể dùng data filters để "register" location – thứ tự sai. Data filters dùng để hạn chế quyền xem dữ liệu (row/column-level), không thay thế bước register ban đầu. Overhead cao nếu bỏ qua register. -
✅ Register the S3 path as an AWS Lake Formation location.
Đúng! Bước đầu tiên bắt buộc để Lake Formation nhận diện S3 path làm "data location" và áp dụng governance. Sau register, Lake Formation tạo permissions blueprint trên Glue Data Catalog, cho phép FGAC mà không di chuyển dữ liệu. Ít overhead vì chỉ cần làm một lần cho data lake. -
❌ Modify the IAM roles of the HR departments to add a data filter for each department's Region.
Sai vì data filters thuộc Lake Formation, không attach trực tiếp vào IAM roles. Phải grant permissions qua Lake Formation console/CLI trước (ví dụ: Grant data filter to IAM role). Sửa IAM roles thủ công sẽ tạo policies phức tạp (như s3:prefix conditions), không scale tốt cho 5 Regions và vi phạm least overhead. -
✅ Enable fine-grained access control in AWS Lake Formation. Add a data filter for each Region.
Đúng! Sau register location, enable FGAC trên database/table để kích hoạt row-level security. Data filters (ví dụ: filter "region_column = 'ap-southeast-1'") restrict HR role chỉ thấy dữ liệu Region cụ thể. Tích hợp với Athena/Glue jobs, tự động enforce khi query. -
❌ Create a separate S3 bucket for each Region. Configure an IAM policy to allow S3 access. Restrict access based on Region.
Sai vì tạo 5 buckets riêng tốn kém và overhead cao (quản lý replication, versioning, lifecycle cross-Region). IAM policies chỉ restrict bucket/prefix thô, không fine-grained như data filters (không filter theo employee region metadata). Không tận dụng Lake Formation, tăng complexity vận hành.
📘 Tài liệu tham khảo (Cập nhật AWS 2026)
- AWS Lake Formation Documentation: Register locations & Data filters – Giải thích thứ tự register → enable FGAC → add filters.
- Fine-Grained Access Control: Enable FGAC – Hỗ trợ row/column filters từ 2020+, ổn định đến 2026.
- Best Practices Data Lakes: AWS Well-Architected Framework – Data Lake Lens: Sử dụng Lake Formation cho governance cross-Region.
- Exam Prep: AWS Certified Data Engineer/DevOps Pro – Sample questions về Lake Formation permissions (re:Post & A Cloud Guru).
Giải pháp này đảm bảo zero-copy access, secure và scalable! 🚀 Nếu cần demo CLI/script, hỏi thêm nhé!
The company's cloud infrastructure team manually built a Step Functions state machine. The cloud infrastructure team launched an EMR cluster into a VPC to support the EMR jobs. However, the deployed Step Functions state machine is not able to run the EMR jobs.
Which combination of steps should the company take to identify the reason the Step Functions state machine is not able to run the EMR jobs? (Choose two.)
- A Use AWS CloudFormation to automate the Step Functions state machine deployment. Create a step to pause the state machine during the EMR jobs that fail. Configure the step to wait for a human user to send approval through an email message. Include details of the EMR task in the email message for further analysis.
- B Verify that the Step Functions state machine code has all IAM permissions that are necessary to create and run the EMR jobs. Verify that the Step Functions state machine code also includes IAM permissions to access the Amazon S3 buckets that the EMR jobs use. Use Access Analyzer for S3 to check the S3 access properties.
- C Check for entries in Amazon CloudWatch for the newly created EMR cluster. Change the AWS Step Functions state machine code to use Amazon EMR on EKS. Change the IAM access policies and the security group configuration for the Step Functions state machine code to reflect inclusion of Amazon Elastic Kubernetes Service (Amazon EKS).
- D Query the flow logs for the VPC. Determine whether the traffic that originates from the EMR cluster can successfully reach the data providers. Determine whether any security group that might be attached to the Amazon EMR cluster allows connections to the data source servers on the informed ports.
- E Check the retry scenarios that the company configured for the EMR jobs. Increase the number of seconds in the interval between each EMR task. Validate that each fallback state has the appropriate catch for each decision state. Configure an Amazon Simple Notification Service (Amazon SNS) topic to store the error messages.
Xem giải thích
🧩 Giải thích nội dung câu hỏi một cách chi tiết
Câu hỏi mô tả một tình huống thực tế trong AWS: Một công ty sử dụng AWS Step Functions để điều phối (orchestrate) một data pipeline bao gồm các công việc Amazon EMR (Elastic MapReduce). Pipeline này có hai giai đoạn chính:
- Ingest data: Các EMR jobs lấy dữ liệu từ các data sources (nguồn dữ liệu bên ngoài) và lưu vào Amazon S3 bucket.
- Load data: Các EMR jobs khác tải dữ liệu từ S3 vào Amazon Redshift.
Đội ngũ cloud infrastructure đã xây dựng thủ công một Step Functions state machine, và khởi chạy EMR cluster trong một VPC (Virtual Private Cloud) để hỗ trợ các EMR jobs. Vấn đề: State machine không thể chạy được các EMR jobs.
Mục tiêu câu hỏi: Xác định kết hợp 2 bước (choose two) để phân tích nguyên nhân (identify the reason) tại sao Step Functions không chạy được EMR jobs. Đây là vấn đề troubleshooting phổ biến liên quan đến quyền IAM, mạng VPC, security groups, và kết nối. Theo kiến thức AWS cập nhật đến 2026 (AWS Well-Architected Framework DevOps Pillar và EMR best practices), các nguyên nhân chính có thể là:
- Thiếu IAM permissions cho Step Functions gọi API EMR và truy cập S3.
- Vấn đề mạng: EMR cluster trong VPC không kết nối được với data sources do flow logs, security groups, hoặc NACLs.
📘 Tài liệu tham khảo:
- AWS Step Functions IAM Permissions (cập nhật 2025).
- Amazon EMR Networking in VPC (2026).
- VPC Flow Logs Troubleshooting.
- Access Analyzer for S3.
✅ Đáp án đúng và lý do lựa chọn
Hai đáp án đúng là phương án thứ 2 và phương án thứ 4.
Lý do chọn:
- 🛠️ Phương án 2: Step Functions cần execution role IAM với đầy đủ quyền để gọi các EMR API (như
emr:RunJobFlow,emr:AddJobFlowSteps) và truy cập S3 (s3:PutObject,s3:GetObject). Access Analyzer for S3 (tính năng mới từ 2023, cập nhật 2026) giúp kiểm tra chính xác quyền S3 access. Đây là bước bắt buộc đầu tiên trong troubleshooting vì thiếu IAM thường gây lỗi "Access Denied" ngay khi state machine cố chạy EMR. - 🛠️ Phương án 4: EMR cluster trong VPC có thể gặp vấn đề network connectivity đến data sources (ví dụ: ports bị chặn). VPC Flow Logs cho phép query traffic (REJECT/ACCEPT) từ EMR để data providers, kết hợp kiểm tra security groups trên EMR cluster (inbound/outbound rules). Đây là bước chẩn đoán mạng chuẩn theo AWS VPC troubleshooting guide.
Kết hợp hai bước này sẽ cover hai nguyên nhân phổ biến nhất: quyền truy cập và kết nối mạng.
📋 Giải thích tất cả các phương án (đúng và sai)
-
❌ Phương án SAI:
Use AWS CloudFormation to automate the Step Functions state machine deployment. Create a step to pause the state machine during the EMR jobs that fail. Configure the step to wait for a human user to send approval through an email message. Include details of the EMR task in the email message for further analysis.
Giải thích: Phương án này tập trung vào tự động hóa deployment (CloudFormation) và human approval (pause state với email), không phải identify nguyên nhân thất bại. Nó chỉ là workaround sau khi biết lỗi, không giúp troubleshoot IAM hay network ngay lập tức. Không phù hợp với câu hỏi "identify the reason". -
✅ Phương án ĐÚNG:
Verify that the Step Functions state machine code has all IAM permissions that are necessary to create and run the EMR jobs. Verify that the Step Functions state machine code also includes IAM permissions to access the Amazon S3 buckets that the EMR jobs use. Use Access Analyzer for S3 to check the S3 access properties.
Giải thích: Hoàn toàn chính xác! Step Functions execution role phải có policy nhưAmazonEMRFullAccess(hoặc custom) để tạo/run EMR clusters/jobs, vàAmazonS3FullAccesscho bucket. Access Analyzer for S3 (policy-based analysis, hỗ trợ query unused access đến 2026) giúp phát hiện chính xác quyền thiếu, tránh over-privileged. Đây là bước security troubleshooting đầu tiên theo AWS IAM best practices. -
❌ Phương án SAI:
Check for entries in Amazon CloudWatch for the newly created EMR cluster. Change the AWS Step Functions state machine code to use Amazon EMR on EKS. Change the IAM access policies and the security group configuration for the Step Functions state machine code to reflect inclusion of Amazon Elastic Kubernetes Service (Amazon EKS).
Giải thích: Kiểm tra CloudWatch Logs hữu ích nhưng không đủ sâu; phần còn lại gợi ý migrate sang EMR on EKS (một giải pháp thay thế từ 2022, cập nhật 2026), thay đổi IAM/SG cho EKS – điều này không identify nguyên nhân hiện tại (EMR classic trong VPC), mà là giải pháp thay đổi architecture. Không liên quan trực tiếp đến troubleshooting. -
✅ Phương án ĐÚNG:
Query the flow logs for the VPC. Determine whether the traffic that originates from the EMR cluster can successfully reach the data providers. Determine whether any security group that might be attached to the Amazon EMR cluster allows connections to the data source servers on the informed ports.
Giải thích: VPC Flow Logs (enable trên VPC/subnet/ENI của EMR) là công cụ chuẩn để query traffic (sử dụng Athena hoặc CloudWatch Logs Insights). Kiểm tra ACCEPT/REJECT từ EMR đến data sources (ports cụ thể như 443 cho HTTPS), và security groups (inbound từ data providers hoặc outbound). EMR trong VPC thường gặp vấn đề này nếu SG/NACL chặn, phù hợp 100% với mô tả "launched into a VPC". -
❌ Phương án SAI:
Check the retry scenarios that the company configured for the EMR jobs. Increase the number of seconds in the interval between each EMR task. Validate that each fallback state has the appropriate catch for each decision state. Configure an Amazon Simple Notification Service (Amazon SNS) topic to store the error messages.
Giải thích: Phương án này xử lý retry/catch logic trong Step Functions (Callback/Wait states) và SNS notifications, hữu ích cho error handling sau khi biết lỗi, nhưng không identify nguyên nhân gốc (như IAM hay network). Tăng interval hay SNS chỉ là optimization, không phải troubleshooting bước đầu.
🏆 Kết luận và lời khuyên DevOps
Kết hợp IAM verification + VPC Flow Logs là cách hiệu quả nhất để nhanh chóng pinpoint vấn đề. Trong thực tế DevOps, hãy dùng AWS X-Ray để trace Step Functions execution và CloudTrail cho API calls thất bại. Áp dụng Infrastructure as Code (CloudFormation/Terraform) để tránh build thủ công như trường hợp này! 🚀
A data engineer must launch new EC2 instances from an Amazon Machine Image (AMI) and configure the instances to preserve the data.
Which solution will meet this requirement?
- A Launch new EC2 instances by using an AMI that is backed by an EC2 instance store volume that contains the application data. Apply the default settings to the EC2 instances.
- B Launch new EC2 instances by using an AMI that is backed by a root Amazon Elastic Block Store (Amazon EBS) volume that contains the application data. Apply the default settings to the EC2 instances.
- C Launch new EC2 instances by using an AMI that is backed by an EC2 instance store volume. Attach an Amazon Elastic Block Store (Amazon EBS) volume to contain the application data. Apply the default settings to the EC2 instances.
- D Launch new EC2 instances by using an AMI that is backed by an Amazon Elastic Block Store (Amazon EBS) volume. Attach an additional EC2 instance store volume to contain the application data. Apply the default settings to the EC2 instances.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh một ứng dụng chạy trên Amazon EC2 instances, nơi dữ liệu được tạo ra hiện tại chỉ là tạm thời (temporary), có nghĩa là dữ liệu có thể mất khi instance bị terminate (dừng hoặc xóa). Công ty cần persist dữ liệu (lưu trữ lâu dài) ngay cả khi các instance EC2 bị terminate.
Một data engineer phải:
- Launch (khởi chạy) các EC2 instance mới từ một Amazon Machine Image (AMI).
- Cấu hình instances để bảo toàn dữ liệu (preserve the data).
Mục tiêu chính: Tìm giải pháp đảm bảo dữ liệu ứng dụng tồn tại độc lập với vòng đời của instance, sử dụng AMI và cấu hình mặc định (default settings).
🛠️ Kiến thức cốt lõi từ AWS (cập nhật đến 2026):
- EC2 Instance Store: Lưu trữ tạm thời (ephemeral), dữ liệu mất hoàn toàn khi instance stop/terminate hoặc hardware failure.
- Amazon EBS (Elastic Block Store): Lưu trữ bền vững (persistent), dữ liệu giữ nguyên khi instance terminate (trừ khi cấu hình delete). Delete on Termination mặc định:
- Root volume:
true(xóa khi terminate). - Additional volumes (không phải root):
false(giữ lại).
- Root volume:
- AMI types:
- Instance store-backed AMI: Root volume từ instance store (ephemeral).
- EBS-backed AMI: Root volume từ EBS snapshot (persistent, nhưng root vẫn delete mặc định).
Giải pháp phải di chuyển dữ liệu sang EBS riêng biệt để tránh mất dữ liệu khi terminate.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Launch new EC2 instances by using an AMI that is backed by an EC2 instance store volume. Attach an Amazon Elastic Block Store (Amazon EBS) volume to contain the application data. Apply the default settings to the EC2 instances.
Lý do:
- AMI backed by instance store → Root volume là ephemeral (không lưu dữ liệu quan trọng ở đây).
- Attach thêm EBS volume riêng cho dữ liệu ứng dụng → EBS là persistent, default Delete on Termination = false → Dữ liệu giữ nguyên khi instance terminate.
- Áp dụng default settings → Không cần chỉnh sửa, phù hợp yêu cầu.
- Data engineer có thể mount EBS vào instance mới, copy dữ liệu từ temporary storage sang EBS trước khi terminate instance cũ. Khi launch mới từ AMI, attach EBS tương tự → Dữ liệu persist.
🛠️ Ưu điểm: Chi phí thấp (EBS gp3/io2 hiện đại, hỗ trợ Multi-Attach đến 2026), scalable.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn giữ nguyên văn bản gốc bằng tiếng Anh, kèm giải thích sai/đúng bằng tiếng Việt dựa trên hành vi lưu trữ AWS:
-
❌ [SAI] Launch new EC2 instances by using an AMI that is backed by an EC2 instance store volume that contains the application data. Apply the default settings to the EC2 instances.
Lý do sai: AMI backed by instance store chứa dữ liệu ứng dụng → Toàn bộ dữ liệu (bao gồm app data) nằm trên root volume ephemeral. Khi launch instance mới và terminate → Dữ liệu mất hoàn toàn (instance store xóa dữ liệu). Không persist được. -
❌ [SAI] Launch new EC2 instances by using an AMI that is backed by a root Amazon Elastic Block Store (Amazon EBS) volume that contains the application data. Apply the default settings to the EC2 instances.
Lý do sai: AMI backed by EBS root volume chứa dữ liệu → Root EBS persistent qua snapshot, nhưng default Delete on Termination = true cho root → Khi terminate instance → Root volume bị xóa, dữ liệu mất. Không đáp ứng yêu cầu persist. -
✅ [ĐÚNG] Launch new EC2 instances by using an AMI that is backed by an EC2 instance store volume. Attach an Amazon Elastic Block Store (Amazon EBS) volume to contain the application data. Apply the default settings to the EC2 instances.
Lý do đúng: Như phần trên – Root từ instance store (ephemeral, chỉ cho OS), dữ liệu app trên EBS additional volume (persistent, Delete on Termination = false mặc định) → Dữ liệu bảo toàn khi terminate. Hoàn hảo với default settings. -
❌ [SAI] Launch new EC2 instances by using an AMI that is backed by an Amazon Elastic Block Store (Amazon EBS) volume. Attach an additional EC2 instance store volume to contain the application data. Apply the default settings to the EC2 instances.
Lý do sai: AMI backed by EBS root, nhưng dữ liệu app trên additional instance store → Instance store ephemeral → Khi terminate → Dữ liệu app mất ngay, dù root EBS an toàn.
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- AWS EC2 User Guide - Storage: EC2 Instance Store vs. EBS & Add EBS Volume.
- AMI Types: Create AMI from Instance Store.
- Delete on Termination: Modify Delete on Termination (default root=true, data=false).
- DevOps Best Practice: Sử dụng EBS cho persistent data trong DOP-C02 exam blueprint (2024-2026).
🛠️ Lời khuyên DevOps: Luôn dùng EBS CSI Driver cho EKS/EC2, hoặc EFS nếu shared storage cần. Test bằng AWS Console/CLI: aws ec2 describe-volumes.
Which solution will give the company the ability to use Spark to access Athena?
- A Athena query settings
- B Athena workgroup
- C Athena data source
- D Athena query editor
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào việc chuyển đổi từ việc sử dụng Amazon Athena với các truy vấn SQL (sử dụng lệnh Create Table As Select - CTAS) cho các tác vụ ETL (Extract, Transform, Load) sang sử dụng Apache Spark để tạo ra các phân tích dữ liệu (analytics). Công ty cần một giải pháp cho phép Spark truy cập vào Athena một cách trực tiếp, thay vì chỉ dùng SQL thuần túy.
📌 Bối cảnh chính: Athena là dịch vụ serverless query trên dữ liệu S3, hỗ trợ cả SQL và Spark (từ phiên bản engine mới nhất như Athena engine version 3 với Spark hỗ trợ từ năm 2023 và cập nhật đến 2026). Để kích hoạt Spark, cần cấu hình đúng cơ chế cho phép Spark chạy trên dữ liệu Athena mà không cần SQL truyền thống, tận dụng khả năng ETL mạnh mẽ hơn của Spark (như Spark DataFrames, MLlib, v.v.).
🛠️ Mục tiêu: Tìm giải pháp cho phép Spark access Athena để xử lý dữ liệu, tạo bảng, và analytics.
✅ Đáp án đúng: Athena workgroup
Lý do lựa chọn:
- Athena workgroup là cơ chế chính để cấu hình và chạy Apache Spark trên Athena (theo tài liệu AWS mới nhất 2024-2026). Khi tạo một workgroup với Spark engine (ví dụ: engine version 2023.10 hoặc mới hơn), bạn có thể sử dụng Spark notebooks, Spark SQL, hoặc Spark APIs trực tiếp để truy cập dữ liệu Athena trên S3.
- Điều này thay thế hoàn hảo cho CTAS SQL: Spark có thể tạo bảng, transform dữ liệu, và lưu kết quả vào S3 hoặc catalog (Glue Data Catalog). Workgroup quản lý tài nguyên, quyền truy cập, và engine type (SQL hoặc Spark), cho phép chuyển seamless từ SQL sang Spark mà không cần thay đổi hạ tầng.
- ✅ Lợi ích nổi bật: Hỗ trợ serverless Spark scaling tự động, tích hợp với SageMaker Studio hoặc Jupyter notebooks, và tối ưu chi phí theo query.
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên tính khả thi để Spark access Athena.
-
❌ Athena query settings
Giải thích sai: Athena query settings chỉ dùng để cấu hình các tham số chung cho query như timeout, kết quả lưu trữ (query result location trên S3), hoặc encryption. Nó không hỗ trợ kích hoạt Spark engine hay cho phép Spark truy cập trực tiếp. Settings chỉ áp dụng cho SQL queries, không thay thế CTAS bằng Spark. -
✅ Athena workgroup
Giải thích đúng: Như đã nêu ở trên, workgroup là cách duy nhất để chọn Spark engine trong Athena (tạo workgroup mới với "Spark" làm engine type). Spark sau đó access dữ liệu qua workgroup này, hỗ trợ ETL analytics đầy đủ, thay thế SQL/CTAS. Đây là tính năng core được AWS cập nhật liên tục (Spark 3.3+). -
❌ Athena data source
Giải thích sai: Athena data source thường đề cập đến việc đăng ký nguồn dữ liệu trong Glue Data Catalog hoặc connector cho federated queries (như từ RDS, DynamoDB). Nó không liên quan đến Spark access Athena, mà chỉ dùng để Athena query dữ liệu ngoài S3. Không hỗ trợ chuyển sang Spark ETL. -
❌ Athena query editor
Giải thích sai: Athena query editor là giao diện web (trong AWS Console) để viết và chạy SQL queries thủ công. Nó không hỗ trợ Spark code (như PySpark hoặc Scala), chỉ dành cho SQL/CTAS. Để dùng Spark, cần notebooks riêng qua workgroup, không phải editor này.
📘 Tài liệu tham khảo
- AWS Documentation chính thức (cập nhật 2024-2026):
- Using Apache Spark in Amazon Athena – Hướng dẫn tạo workgroup Spark.
- Athena Workgroups – Chi tiết engine selection (SQL vs. Spark).
- Athena Engine Versions – Xác nhận Spark support từ engine 3.
- AWS Well-Architected Framework - Analytics Lens: Khuyến nghị workgroup cho Spark ETL.
- Exam Prep DOP-C02: Chủ đề Athena Spark thường xuất hiện trong phần serverless analytics (Blueprints AWS 2024).
🧠 Lời khuyên DevOps: Trong thực tế, hãy dùng Terraform/IaC để provision workgroup Spark, kết hợp IAM roles cho secure access, và monitor qua CloudWatch/CloudTrail để scale ETL jobs hiệu quả! 🚀
A data engineer must ensure that the AWS Glue Data Catalog synchronizes with the S3 storage when the company adds new partitions to the bucket.
Which solution will meet these requirements with the LEAST latency?
- A Schedule an AWS Glue crawler to run every morning.
- B Manually run the AWS Glue CreatePartition API twice each day.
- C Use code that writes data to Amazon S3 to invoke the Boto3 AWS Glue create_partition API call.
- D Run the MSCK REPAIR TABLE command from the AWS Glue console.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc partitioning dữ liệu trong Amazon S3 cho data lake theo định dạng đường dẫn Hive-style: s3://bucket/prefix/year=2023/month=01/day=01.
📊 Yêu cầu chính: Khi công ty thêm partition mới vào S3 bucket (ví dụ: dữ liệu mới theo ngày), AWS Glue Data Catalog phải đồng bộ (synchronize) ngay lập tức với S3 storage để các query engine như Athena hoặc Spark có thể truy vấn partition mới mà không bị delay.
⚡ Tiêu chí quan trọng: Giải pháp phải có LEAST latency (độ trễ thấp nhất), nghĩa là đồng bộ gần như real-time khi thêm partition, tránh các phương pháp scan toàn bộ hoặc chạy theo lịch.
🛠️ Bối cảnh AWS: Sử dụng AWS Glue để quản lý metadata partitions trong Data Catalog, kết hợp với S3 partitioning để tối ưu performance và cost cho data lake lớn (theo best practices AWS Lake Formation và Glue đến 2026).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use code that writes data to Amazon S3 to invoke the Boto3 AWS Glue create_partition API call.
Lý do chi tiết:
- Phương án này đồng bộ ngay lập tức (real-time) khi code viết dữ liệu vào S3, bằng cách gọi trực tiếp API
create_partitionqua Boto3 (SDK Python cho AWS). - 🕒 Least latency: Không cần chờ scheduler, crawler scan, hay manual intervention – partition metadata được thêm vào Glue Catalog chỉ trong vài giây sau khi write data.
- 🧩 Phù hợp partitioning Hive-style: API hỗ trợ chính xác format
year=2023/month=01/day=01, tự động infer schema nếu cần. - 📈 Scalable & efficient: Lý tưởng cho data lake lớn, tránh overhead scan toàn bộ table (như crawler hoặc MSCK REPAIR). Đây là best practice từ AWS Glue docs (cập nhật 2026), đặc biệt khi tích hợp với ETL jobs (Glue Jobs, EMR, Lambda).
Tài liệu tham khảo:
- 📘 AWS Glue Data Catalog Partitions (API
create_partition). - 📘 Managing Partitions for ETL Processing (Best practices low-latency sync).
- 📘 Boto3 Glue Client.
🔍 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] Schedule an AWS Glue crawler to run every morning.
Phương án này chạy AWS Glue Crawler theo lịch hàng sáng, quét S3 để detect partitions mới và cập nhật Catalog.
❌ Lý do sai: Latency cao (chờ đến sáng hôm sau, có thể delay 24h+), không real-time. Crawler scan toàn bộ prefix, tốn cost và thời gian cho data lake lớn (không efficient theo AWS best practices 2026). Chỉ phù hợp initial crawl, không phải incremental sync. -
❌ [SAI] Manually run the AWS Glue CreatePartition API twice each day.
Phương án yêu cầu chạy thủ công API CreatePartition 2 lần/ngày.
❌ Lý do sai: Manual operation không tự động, dễ lỗi con người, và vẫn có latency (chỉ 2 lần/ngày, miss real-time additions). Không scalable cho production data lake với partitions thêm liên tục. -
✅ [ĐÚNG] Use code that writes data to Amazon S3 to invoke the Boto3 AWS Glue create_partition API call.
Như đã giải thích ở trên: Tích hợp trực tiếp vào code write data, gọi Boto3 API để add partition metadata ngay lập tức.
✅ Ưu điểm nổi bật: Zero manual effort, least latency (<1 phút), hỗ trợ idempotent (tránh duplicate partitions), tích hợp dễ với Spark/EMR/Glue Jobs. -
❌ [SAI] Run the MSCK REPAIR TABLE command from the AWS Glue console.
Phương án chạy lệnh MSCK REPAIR TABLE từ Glue console (tương đương Athena/Hive command).
❌ Lý do sai: Lệnh này scan toàn bộ S3 table để repair/add partitions, gây latency cao (giờ/phút cho TB data), tốn cost Glue/Athena. AWS khuyến cáo tránh dùng thường xuyên (deprecated cho large tables từ Glue 4.0+ 2026), thay bằng API programmatic.
💡 Lời khuyên thực tế: Trong production, kết hợp với AWS Glue Workflows hoặc EventBridge + Lambda để tự động hóa thêm, đảm bảo 100% real-time sync cho data lake! 🚀