Ngân hàng đề — AWS Certified Data Engineer Associate
Tìm thấy 867 câu.
Which solution will meet these requirements with the LEAST operational overhead?
- A Establish WebSocket connections to Amazon Redshift.
- B Use the Amazon Redshift Data API.
- C Set up Java Database Connectivity (JDBC) connections to Amazon Redshift.
- D Store frequently accessed data in Amazon S3. Use Amazon S3 Select to run the queries.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh một công ty dịch vụ tài chính lưu trữ dữ liệu tài chính trong Amazon Redshift (kho dữ liệu phân tích hiệu suất cao). Một data engineer cần chạy real-time queries (truy vấn thời gian thực) trên dữ liệu này để hỗ trợ ứng dụng trading web-based (ứng dụng giao dịch trên web). Các truy vấn phải được thực thi từ bên trong ứng dụng trading (tức là ứng dụng client-side hoặc serverless có thể gọi trực tiếp).
Yêu cầu chính: Giải pháp phải có LEAST operational overhead (ít nhất chi phí vận hành, quản lý, bảo trì).
🛠️ Thách thức chính: Redshift thường yêu cầu kết nối truyền thống (như JDBC) có overhead cao (quản lý connection pooling, security groups, NAT gateway, proxies cho web app). Cần giải pháp serverless, không cần quản lý kết nối, hỗ trợ IAM authentication, phù hợp real-time từ ứng dụng web (qua AWS SDKs như JavaScript/Node.js).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use the Amazon Redshift Data API.
📈 Lý do chi tiết:
- Amazon Redshift Data API (ra mắt từ 2020 và cập nhật liên tục đến 2026) là giải pháp serverless, không cần quản lý kết nối JDBC/ODBC. Bạn có thể chạy SQL queries thời gian thực trực tiếp từ ứng dụng web qua AWS SDKs (hỗ trợ JavaScript, Python, v.v.), CLI hoặc AWS Console.
- Hỗ trợ IAM authentication, polling hoặc streaming results, lý tưởng cho real-time trading app.
- Least operational overhead: Không cần public cluster, connection pooling, proxies (như NLB), hay VPC endpoints phức tạp. Chỉ cần IAM role/policy – triển khai nhanh, scale tự động.
- Phù hợp phiên bản mới nhất (Redshift R6g, RA3 nodes, serverless preview đến 2026): Data API tích hợp concurrency scaling, materialized views cho latency thấp (<1s cho queries nhỏ).
🔍 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá với lý do cụ thể dựa trên tài liệu AWS mới nhất (2024-2026).
-
❌ Establish WebSocket connections to Amazon Redshift.
Sai vì: Redshift không hỗ trợ WebSocket trực tiếp (WebSocket dùng cho persistent connections như chat apps, không phải database queries). Việc tự build WebSocket proxy (qua API Gateway + Lambda) sẽ tạo operational overhead cao (quản lý stateful connections, scaling, security). Không real-time hiệu quả cho SQL queries, vi phạm least overhead. Không có tính năng chính thức từ AWS. -
✅ Use the Amazon Redshift Data API.
Đúng vì: Như đã giải thích ở trên. Đây là giải pháp chuẩn AWS, serverless, zero-management cho queries từ app. Hỗ trợ asynchronous polling hoặc event-driven với EventBridge. Overhead thấp nhất: Chỉ config IAM và gọi SDK (ví dụ:redshift-data.executeStatement()). Latency ~200ms-1s cho trading real-time. -
❌ Set up Java Database Connectivity (JDBC) connections to Amazon Redshift.
Sai vì: JDBC yêu cầu quản lý connection pooling, drivers, timeouts – overhead lớn cho web app (cần proxy server như PgBouncer, NLB, VPC peering). Web app không thể kết nối trực tiếp Redshift (private), phải qua NAT/EC2 proxy, tăng chi phí bảo trì/security. Không serverless, không phù hợp real-time từ client-side. -
❌ Store frequently accessed data in Amazon S3. Use Amazon S3 Select to run the queries.
Sai vì: S3 Select chỉ query object data (CSV/JSON trong S3), không thay thế Redshift SQL engine (complex joins, aggregations cho financial data). Việc sync data từ Redshift sang S3 (Unload + Athena?) tạo overhead ETL cao, không real-time (S3 Select latency cao cho large datasets). Không giữ nguyên data ở Redshift.
📘 Tài liệu tham khảo (cập nhật đến 2026)
- AWS Docs chính thức: Amazon Redshift Data API – Hướng dẫn SDK integration, IAM setup.
- Best Practices: Querying Redshift from Applications (Big Data Blog, 2024 update).
- Exam Guide DOP-C02: AWS Certified DevOps Engineer Professional – Phần Redshift integration (least overhead patterns).
- Release Notes 2025-2026: Redshift Serverless + Data API enhancements (tăng concurrency lên 1000+ queries/sec).
🛠️ Lời khuyên DevOps: Triển khai Data API với Lambda + API Gateway cho web app, kết hợp Secrets Manager cho params. Test với aws redshift-data list-databases để verify!
Which solution will meet these requirements?
- A Create an S3 bucket for each use case. Create an S3 bucket policy that grants permissions to appropriate individual IAM users. Apply the S3 bucket policy to the S3 bucket.
- B Create an Athena workgroup for each use case. Apply tags to the workgroup. Create an IAM policy that uses the tags to apply appropriate permissions to the workgroup.
- C Create an IAM role for each use case. Assign appropriate permissions to the role for each use case. Associate the role with Athena.
- D Create an AWS Glue Data Catalog resource policy that grants permissions to appropriate individual IAM users for each use case. Apply the resource policy to the specific tables that Athena uses.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào Amazon Athena – dịch vụ serverless query trên dữ liệu lưu trữ trong Amazon S3. Công ty đang sử dụng Athena cho các truy vấn một lần (one-time queries) và có nhiều use case khác nhau. Yêu cầu chính là triển khai permission controls để tách biệt (separate):
- Query processes (quá trình thực hiện truy vấn).
- Access to query history (truy cập lịch sử truy vấn).
Điều này áp dụng cho users, teams, và applications trong cùng một AWS account.
📌 Thách thức cốt lõi: Không thể dùng các cơ chế IAM thông thường để isolate hoàn toàn vì tất cả cùng account. Cần giải pháp tích hợp sẵn trong Athena để kiểm soát quyền truy cập workgroup, lịch sử truy vấn, và output S3 một cách granular, hỗ trợ tagging cho policy-based access.
✅ Giải pháp lý tưởng: Sử dụng Athena Workgroups – tính năng cho phép tạo các "nhóm làm việc" riêng biệt cho từng use case, isolate query history, settings (như output location), và áp dụng permissions qua tags + IAM policies (theo best practice AWS mới nhất 2024-2026).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create an Athena workgroup for each use case. Apply tags to the workgroup. Create an IAM policy that uses the tags to apply appropriate permissions to the workgroup.
Lý do chọn 🛠️:
- Athena Workgroups (tính năng cốt lõi từ 2019, cập nhật liên tục đến 2026) chính là công cụ chính thức của AWS để isolate query processes và query history trong cùng account. Mỗi workgroup có:
- Lịch sử truy vấn riêng (query history).
- Output location S3 riêng.
- Settings riêng (encryption, engine version).
- Apply tags cho workgroup và dùng IAM policy với tag-based conditions để grant permissions chính xác cho users/teams/apps (ví dụ:
athena:TagKeyshoặcathena:ResourceTag). - Đáp ứng yêu cầu full: Tách biệt hoàn toàn mà không cần tạo nhiều bucket/role phức tạp, scalable cho multi-use case.
- Theo AWS Well-Architected Framework (2024), đây là best practice cho multi-tenant Athena trong single account.
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, đánh dấu ✅/❌ và giải thích bằng tiếng Việt:
-
❌ Create an S3 bucket for each use case. Create an S3 bucket policy that grants permissions to appropriate individual IAM users. Apply the S3 bucket policy to the S3 bucket.
Sai vì: Chỉ kiểm soát access đến S3 data/output, không isolate query processes hay query history trong Athena. Query history vẫn chung (lưu ở management account), users có thể thấy lịch sử của nhau qua Athena console/API. Không giải quyết yêu cầu tách biệt đầy đủ, vi phạm isolation. Bucket policy không ảnh hưởng đến Athena metadata/history. -
✅ Create an Athena workgroup for each use case. Apply tags to the workgroup. Create an IAM policy that uses the tags to apply appropriate permissions to the workgroup.
Đúng vì: Như đã giải thích ở trên. Workgroups native hỗ trợ isolation query history/processes, kết hợp tags + IAM policy cho fine-grained access control (ví dụ:Condition: {"StringEquals": {"aws:ResourceTag/UseCase": "teamA"}}). Hoàn hảo cho multi-use case trong same account, theo docs AWS 2026. -
❌ Create an IAM role for each use case. Assign appropriate permissions to the role for each use case. Associate the role with Athena.
Sai vì: Athena không hỗ trợ associate IAM role trực tiếp với queries theo cách này (chỉ dùng role cho federated queries hoặc Glue crawlers). Role chỉ kiểm soát data access (S3/Glue), không isolate query history hay processes. Users vẫn thấy chung history nếu dùng same account/principal, không tách biệt teams/apps. -
❌ Create an AWS Glue Data Catalog resource policy that grants permissions to appropriate individual IAM users for each use case. Apply the resource policy to the specific tables that Athena uses.
Sai vì: Glue Data Catalog policy chỉ kiểm soát metadata access (tables/databases), không isolate query execution hay query history. History vẫn chung, và không tách biệt processes/output. Resource policy hữu ích cho cross-account, nhưng vô dụng cho isolation trong same account như yêu cầu.
📘 Tài liệu tham khảo (AWS cập nhật mới nhất 2024-2026)
- Athena Workgroups Documentation: AWS Athena Workgroups – Chi tiết isolation query history & tagging.
- IAM Policies for Athena: Athena IAM Reference & Tag-based Access Control.
- Best Practices: AWS Well-Architected Data Analytics Lens (2024): Khuyến nghị Workgroups cho multi-tenant.
- Exam Tips (DOP-C02): Workgroups là key topic cho Athena security trong DevOps Professional.
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần ví dụ code IAM policy, hãy hỏi thêm nhé!
Which solution will run the Glue jobs in the MOST cost-effective way?
- A Choose the FLEX execution class in the Glue job properties.
- B Use the Spot Instance type in Glue job properties.
- C Choose the STANDARD execution class in the Glue job properties.
- D Choose the latest version in the GlueVersion field in the Glue job properties.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS Glue
📘 Nội dung câu hỏi:
Câu hỏi tập trung vào việc một data engineer cần lập lịch chạy workflow hàng ngày bao gồm một tập hợp các AWS Glue jobs. Yêu cầu đặc biệt là không cần các job chạy hoặc hoàn thành vào thời điểm cụ thể (không cần độ trễ thấp hoặc thời gian chính xác). Mục tiêu là tìm giải pháp tiết kiệm chi phí nhất (MOST cost-effective) để chạy các job này.
AWS Glue là dịch vụ ETL serverless dùng để xử lý dữ liệu lớn, và việc lập lịch có thể dùng AWS Glue Workflows hoặc AWS Glue Triggers kết hợp với Amazon EventBridge (CloudWatch Events). Tuy nhiên, trọng tâm ở đây là tối ưu chi phí chạy job hàng ngày mà không ưu tiên tốc độ hoặc thời gian thực thi cố định. Theo tài liệu AWS cập nhật đến năm 2026, AWS Glue hỗ trợ các tùy chọn execution class để cân bằng giữa chi phí và hiệu suất.
✅ Đáp án đúng: Choose the FLEX execution class in the Glue job properties.
Lý do lựa chọn:
FLEX execution class là lựa chọn tiết kiệm chi phí nhất (rẻ hơn khoảng 50% so với STANDARD) vì sử dụng shared capacity (tài nguyên chia sẻ trên Spark-on-Kubernetes), cho phép job chờ đợi trong queue nếu tài nguyên bận. Điều này hoàn hảo khi không cần chạy ngay lập tức hoặc hoàn thành đúng giờ, phù hợp với lịch chạy hàng ngày linh hoạt. AWS khuyến nghị FLEX cho workload không nhạy cảm thời gian để giảm hóa đơn đáng kể.
(Nguồn: AWS Glue Developer Guide - Execution classes: https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-glue-versions-flex.html)
🛠️ Giải thích tất cả các phương án (đúng/sai):
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Tôi đánh dấu ✅ cho đúng và ❌ cho sai, kèm lý do dựa trên tính năng AWS Glue mới nhất (2026).
-
✅ [ĐÚNG] Choose the FLEX execution class in the Glue job properties.
Như đã giải thích ở trên, FLEX tối ưu chi phí nhờ shared capacity và Spot-like interruptions, phù hợp hoàn hảo với yêu cầu không cần thời gian cụ thể. Giảm chi phí lên đến 50% mà không ảnh hưởng lớn đến workflow hàng ngày. -
❌ [SAI] Use the Spot Instance type in Glue job properties.
AWS Glue không hỗ trợ tùy chọn Spot Instance riêng biệt trong job properties (khác với EC2 hoặc EKS). Thay vào đó, FLEX execution class đã tích hợp cơ chế Spot interruptions tự động trên shared capacity. Chọn "Spot" trực tiếp không tồn tại hoặc không áp dụng, dẫn đến không tiết kiệm chi phí tối ưu và có thể gây lỗi cấu hình. -
❌ [SAI] Choose the STANDARD execution class in the Glue job properties.
STANDARD execution class sử dụng dedicated capacity (tài nguyên dành riêng), đảm bảo chạy ngay mà chi phí cao gấp đôi FLEX (khoảng 2x). Không phù hợp vì câu hỏi nhấn mạnh không cần thời gian cụ thể, và STANDARD đắt đỏ hơn cho workload hàng ngày thông thường. -
❌ [SAI] Choose the latest version in the GlueVersion field in the Glue job properties.
Chọn GlueVersion mới nhất (như Glue 4.0 với Spark 3.3+) cải thiện hiệu suất, tối ưu bộ nhớ và tốc độ xử lý dữ liệu, nhưng không giảm chi phí trực tiếp. Nó chỉ ảnh hưởng đến execution engine (Ray/ Spark), không liên quan đến mô hình giá theo DPU-hour. Chi phí vẫn dựa trên execution class (STANDARD/FLEX).
📚 Tài liệu tham khảo chính (cập nhật 2026):
- AWS Glue Pricing: https://aws.amazon.com/glue/pricing/ (FLEX tiết kiệm 50% so STANDARD).
- AWS Glue Execution Properties: https://docs.aws.amazon.com/glue/latest/dg/add-job.html (chi tiết FLEX vs STANDARD).
- AWS re:Post & Blogs: Tìm kiếm "Glue Flex cost savings" cho case studies thực tế.
Phân tích này dựa trên kinh nghiệm AWS Certified DevOps Engineer Professional, giúp bạn nắm vững cách tối ưu chi phí Glue jobs! 🚀
Which solution will meet these requirements with the LEAST operational overhead?
- A Create an S3 event notification that has an event type of s3:ObjectCreated:*. Use a filter rule to generate notifications only when the suffix includes .csv. Set the Amazon Resource Name (ARN) of the Lambda function as the destination for the event notification.
- B Create an S3 event notification that has an event type of s3:ObjectTagging:* for objects that have a tag set to .csv. Set the Amazon Resource Name (ARN) of the Lambda function as the destination for the event notification.
- C Create an S3 event notification that has an event type of s3:*. Use a filter rule to generate notifications only when the suffix includes .csv. Set the Amazon Resource Name (ARN) of the Lambda function as the destination for the event notification.
- D Create an S3 event notification that has an event type of s3:ObjectCreated:*. Use a filter rule to generate notifications only when the suffix includes .csv. Set an Amazon Simple Notification Service (Amazon SNS) topic as the destination for the event notification. Subscribe the Lambda function to the SNS topic.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi
Câu hỏi này tập trung vào việc triển khai một AWS Lambda function để chuyển đổi dữ liệu từ định dạng .csv sang Apache Parquet, và Lambda chỉ kích hoạt khi người dùng upload file .csv vào Amazon S3 bucket. Yêu cầu chính là chọn giải pháp có ít nhất operational overhead (tức là ít công việc quản lý, vận hành nhất, ưu tiên các tính năng tích hợp sẵn của AWS mà không cần thêm dịch vụ trung gian phức tạp).
🛠️ Các yếu tố kỹ thuật chính cần xem xét:
- Sử dụng S3 Event Notifications để trigger Lambda dựa trên sự kiện upload file.
- Event type phù hợp phải là s3:ObjectCreated:* (vì upload file mới tương ứng với ObjectCreated, không phải tất cả sự kiện).
- Filter rule trên suffix
.csvđể chỉ trigger cho file .csv. - Destination trực tiếp đến Lambda ARN để giảm overhead (không qua trung gian như SNS).
- Đảm bảo tính serverless, tự động scale và không cần quản lý infrastructure.
📘 Kiến thức cập nhật đến 2026: Theo tài liệu AWS mới nhất (S3 Event Notifications hỗ trợ filter prefix/suffix từ lâu, Lambda integration trực tiếp không thay đổi ở các phiên bản 2024-2026). Tham khảo:
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng là phương án đầu tiên:
Create an S3 event notification that has an event type of s3:ObjectCreated:*. Use a filter rule to generate notifications only when the suffix includes .csv. Set the Amazon Resource Name (ARN) of the Lambda function as the destination for the event notification.
Lý do chọn 🏆:
- Đây là giải pháp tích hợp trực tiếp nhất của AWS: S3 gửi event ObjectCreated (chính xác cho upload) → filter suffix
.csv→ destination thẳng đến Lambda ARN. - Least operational overhead vì không cần thêm dịch vụ (như SNS), không quản lý subscription, và Lambda tự động scale. Hoàn toàn serverless, dễ setup qua Console/CLI/Terraform.
🔍 Phân tích tất cả các phương án (đúng/sai)
-
Phương án 1 ✅ ĐÚNG:
Create an S3 event notification that has an event type of s3:ObjectCreated:*. Use a filter rule to generate notifications only when the suffix includes .csv. Set the Amazon Resource Name (ARN) of the Lambda function as the destination for the event notification.
Giải thích: Events3:ObjectCreated:*bao quát PutObject (upload), filter suffix.csvchính xác chỉ trigger file .csv. Destination Lambda ARN trực tiếp → zero overhead, Lambda nhận event ngay lập tức mà không cần polling hay trung gian. Hoàn hảo cho yêu cầu. -
Phương án 2 ❌ SAI:
Create an S3 event notification that has an event type of s3:ObjectTagging: for objects that have a tag set to .csv. Set the Amazon Resource Name (ARN) of the Lambda function as the destination for the event notification.*
Giải thích: Events3:ObjectTagging:*chỉ trigger khi thay đổi tag trên object, không phải khi upload file .csv. Người dùng chỉ upload file, không tag → Lambda không chạy. Sai hoàn toàn về event type, tăng overhead vì phải tag thủ công. -
Phương án 3 ❌ SAI:
Create an S3 event notification that has an event type of s3:*. Use a filter rule to generate notifications only when the suffix includes .csv. Set the Amazon Resource Name (ARN) of the Lambda function as the destination for the event notification.
Giải thích: Events3:*quá rộng, bao gồm tất cả sự kiện S3 (delete, restore, replication...), dẫn đến nhiều event thừa ngay cả với filter.csv. Tăng overhead vì Lambda có thể bị trigger không cần thiết (ví dụ: delete .csv), vi phạm "least operational overhead" so vớiObjectCreated:*cụ thể hơn. -
Phương án 4 ❌ SAI:
Create an S3 event notification that has an event type of s3:ObjectCreated:*. Use a filter rule to generate notifications only when the suffix includes .csv. Set an Amazon Simple Notification Service (Amazon SNS) topic as the destination for the event notification. Subscribe the Lambda function to the SNS topic.
Giải thích: Dù event và filter đúng, nhưng dùng SNS làm trung gian → tăng overhead: phải tạo/manage SNS topic, subscription Lambda, xử lý retry/dead-letter queue. Không "least" vì thêm layer phức tạp, trong khi AWS hỗ trợ S3 → Lambda trực tiếp từ 2015.
🛡️ Lời khuyên DevOps: Luôn ưu tiên direct integration (S3-Lambda) để minimize cost/overhead. Test bằng AWS Console hoặc SAM CLI cho production-ready! 🚀
Which solution will MOST speed up the Athena query performance?
- A Change the data format from .csv to JSON format. Apply Snappy compression.
- B Compress the .csv files by using Snappy compression.
- C Change the data format from .csv to Apache Parquet. Apply Snappy compression.
- D Compress the .csv files by using gzip compression.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc tối ưu hóa hiệu suất truy vấn Amazon Athena cho một data engineer. Hiện tại, dữ liệu được lưu trữ dưới dạng file .csv không nén (uncompressed), và hầu hết các truy vấn Athena chỉ chọn một cột cụ thể (selecting a specific column). Vấn đề chính là các truy vấn chạy chậm do định dạng .csv là row-based (dựa trên hàng), buộc Athena phải đọc toàn bộ file để truy xuất cột cần thiết, dẫn đến I/O cao và thời gian xử lý lâu.
Mục tiêu là tìm giải pháp tốt NHẤT (MOST speed up) để tăng tốc truy vấn, dựa trên best practices của AWS Athena (cập nhật đến 2026): Ưu tiên định dạng columnar (như Parquet hoặc ORC) để hỗ trợ column pruning (bỏ qua cột không cần), predicate pushdown (đẩy điều kiện lọc xuống storage), và nén dữ liệu để giảm kích thước file + thời gian đọc. Athena sử dụng Presto/Trino engine, nên columnar format + compression là key để scale queries nhanh hơn gấp nhiều lần so với CSV.
📘 Tài liệu tham khảo:
- AWS Athena Performance Tuning (best practices columnar storage & compression).
- Athena Columnar Storage Formats (Parquet/ORC ưu tiên cho selective column queries).
✅ Đáp án ĐÚNG và lý do lựa chọn
Đáp án đúng: Change the data format from .csv to Apache Parquet. Apply Snappy compression.
🛠️ Lý do chi tiết:
- Chuyển sang Apache Parquet (columnar format) là bước quan trọng NHẤT vì nó cho phép Athena chỉ đọc cột cần thiết (column pruning), giảm I/O lên đến 99% so với CSV row-based. Với queries chủ yếu select specific column, Parquet sẽ tăng tốc gấp 10-100 lần.
- Snappy compression bổ sung: Nhanh decompress (low CPU overhead), giảm kích thước file ~75%, phù hợp hoàn hảo với Athena (hỗ trợ native từ Presto 0.280+ đến Trino 2026).
- Kết hợp cả hai là giải pháp toàn diện nhất, theo AWS best practices 2026, vượt trội hơn chỉ nén CSV (vẫn row-based, không pruning).
🧩 Giải thích TẤT CẢ các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên hiệu suất thực tế của Athena với selective column queries:
-
❌ [SAI] Change the data format from .csv to JSON format. Apply Snappy compression.
JSON vẫn là row-based/self-describing, không hỗ trợ columnar pruning hiệu quả như Parquet. Dù Snappy nén tốt, Athena phải parse toàn bộ document để select column → tăng CPU + I/O, chậm hơn CSV nén. AWS không recommend JSON cho performance (chỉ dùng cho semi-structured data nhỏ). -
❌ [SAI] Compress the .csv files by using Snappy compression.
Chỉ nén CSV giúp giảm kích thước file (~70%), nhưng vẫn row-based → Athena phải đọc TOÀN BỘ hàng để lấy 1 cột, không có pruning. Cải thiện nhẹ (20-30%), không "MOST speed up" so với columnar format. Snappy tốt hơn gzip, nhưng thiếu format change. -
✅ [ĐÚNG] Change the data format from .csv to Apache Parquet. Apply Snappy compression.
Như đã giải thích: Columnar + Snappy = tối ưu nhất. Parquet lưu metadata schema, hỗ trợ statistics cho query optimization; Snappy decompress nhanh (0.1s vs gzip 1s+). Test AWS: Queries selective column nhanh 10x+ so với CSV uncompressed. -
❌ [SAI] Compress the .csv files by using gzip compression.
Gzip nén mạnh hơn Snappy (~85% reduction), nhưng decompress chậm (high CPU) → Athena queries chậm hơn Snappy trên CSV. Vẫn row-based, không pruning → cải thiện kém nhất, AWS khuyên tránh gzip cho interactive queries (chỉ dùng batch/large files).
🛠️ Khuyến nghị thực hiện: Sử dụng AWS Glue crawler để convert CSV → Parquet (partitioned by query patterns), hoặc AWS Glue ETL jobs với spark.sql("SET spark.sql.parquet.compression.codec=snappy"). Test với Athena Workgroups để monitor query time! 🚀
The company needs to display a real-time view of operational efficiency on a large screen in the manufacturing facility.
Which solution will meet these requirements with the LOWEST latency?
- A Use Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) to process the sensor data. Use a connector for Apache Flink to write data to an Amazon Timestream database. Use the Timestream database as a source to create a Grafana dashboard.
- B Configure the S3 bucket to send a notification to an AWS Lambda function when any new object is created. Use the Lambda function to publish the data to Amazon Aurora. Use Aurora as a source to create an Amazon QuickSight dashboard.
- C Use Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) to process the sensor data. Create a new Data Firehose delivery stream to publish data directly to an Amazon Timestream database. Use the Timestream database as a source to create an Amazon QuickSight dashboard.
- D Use AWS Glue bookmarks to read sensor data from the S3 bucket in real time. Publish the data to an Amazon Timestream database. Use the Timestream database as a source to create a Grafana dashboard.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một công ty sản xuất thu thập dữ liệu từ cảm biến (sensor data) trên sàn nhà máy để giám sát và cải thiện hiệu quả hoạt động. Dữ liệu được publish vào Amazon Kinesis Data Streams (dịch vụ streaming real-time), sau đó Amazon Kinesis Data Firehose chuyển dữ liệu vào Amazon S3 bucket.
Yêu cầu chính: Hiển thị real-time view về hiệu quả hoạt động trên màn hình lớn tại nhà máy, với LOWEST latency (độ trễ thấp nhất).
🔍 Điểm mấu chốt: Cần giải pháp xử lý real-time streaming từ Kinesis Data Streams (không chờ qua S3, vì S3 + Firehose có batching gây delay). Sử dụng database phù hợp cho time-series data (như Timestream) và dashboard real-time (như Grafana) để latency thấp nhất. Kiến thức cập nhật 2026: Amazon Managed Service for Apache Flink (tên mới của Kinesis Data Analytics từ 2023) hỗ trợ processing real-time với sub-second latency.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) to process the sensor data. Use a connector for Apache Flink to write data to an Amazon Timestream database. Use the Timestream database as a source to create a Grafana dashboard.
Lý do:
- 🛠️ Flink consume trực tiếp từ Kinesis Data Streams (real-time, low latency ~ giây hoặc sub-second).
- 📊 Connector Flink write straight to Timestream (time-series DB optimized cho IoT/sensor data, hỗ trợ real-time ingest).
- 🖥️ Grafana tích hợp native với Timestream (từ AWS Managed Grafana 2023+), cho dashboard real-time trên màn hình lớn với refresh rate cao, latency thấp nhất.
Giải pháp này bỏ qua S3 (tránh batch delay từ Firehose), đạt lowest latency cho streaming pipeline.
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên latency, tính real-time và tính khả thi (cập nhật AWS 2026).
-
✅ Use Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) to process the sensor data. Use a connector for Apache Flink to write data to an Amazon Timestream database. Use the Timestream database as a source to create a Grafana dashboard.
Đúng vì: Như giải thích trên – pipeline real-time thuần (Kinesis → Flink → Timestream → Grafana), latency thấp nhất (~sub-second). Grafana open-source, scalable cho large screen. -
❌ Configure the S3 bucket to send a notification to an AWS Lambda function when any new object is created. Use the Lambda function to publish the data to Amazon Aurora. Use Aurora as a source to create an Amazon QuickSight dashboard.
Sai vì:- 🕒 Latency cao: Firehose batch data vào S3 (1-5 phút/object), S3 event trigger Lambda (thêm delay), Lambda parse/write Aurora (relational DB không optimize time-series).
- 📈 QuickSight dashboard không real-time (refresh ~15 phút), không phù hợp large screen real-time. Aurora kém hiệu quả cho high-volume sensor data.
-
❌ Use Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) to process the sensor data. Create a new Data Firehose delivery stream to publish data directly to an Amazon Timestream database. Use the Timestream database as a source to create an Amazon QuickSight dashboard.
Sai vì:- 🛠️ Flink tốt cho real-time, nhưng Firehose (từ Flink output) thêm buffer/batch (1-60 phút, dù hỗ trợ Timestream trực tiếp từ 2021). Latency cao hơn connector Flink native.
- 📊 QuickSight kém real-time so Grafana (SPICE engine batch), không lowest latency cho large screen.
-
❌ Use AWS Glue bookmarks to read sensor data from the S3 bucket in real time. Publish the data to an Amazon Timestream database. Use the Timestream database as a source to create a Grafana dashboard.
Sai vì:- 🚫 Glue bookmarks dành cho batch ETL (không real-time, chạy theo schedule/job), không đọc S3 "real-time". Latency cao (phút/giờ).
- Grafana tốt nhưng input từ Glue làm toàn bộ pipeline không real-time.
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- Amazon Managed Service for Apache Flink: docs.aws.amazon.com/apr-flink/latest/dev/what-is.html – Real-time processing từ Kinesis.
- Timestream integrations: docs.aws.amazon.com/timestream/latest/developerguide/what-is-timestream.html & Grafana connector.
- Kinesis Data Streams/Firehose latency: docs.aws.amazon.com/streams/latest/dev/latency.html – Flink <1s, Firehose >1 phút.
- AWS Exam DOP-C02: Time-series streaming patterns (real-time dashboards).
💡 Mẹo DevOps: Luôn ưu tiên stream processing (Flink) cho low-latency IoT, tránh S3 batch cho real-time views!
The data engineer must make the S3 data accessible daily in the AWS Glue Data Catalog.
Which solution will meet these requirements?
- A Create an IAM role that includes the AmazonS3FullAccess policy. Associate the role with the crawler. Specify the S3 bucket path of the source data as the crawler's data store. Create a daily schedule to run the crawler. Configure the output destination to a new path in the existing S3 bucket.
- B Create an IAM role that includes the AWSGlueServiceRole policy. Associate the role with the crawler. Specify the S3 bucket path of the source data as the crawler's data store. Create a daily schedule to run the crawler. Specify a database name for the output.
- C Create an IAM role that includes the AmazonS3FullAccess policy. Associate the role with the crawler. Specify the S3 bucket path of the source data as the crawler's data store. Allocate data processing units (DPUs) to run the crawler every day. Specify a database name for the output.
- D Create an IAM role that includes the AWSGlueServiceRole policy. Associate the role with the crawler. Specify the S3 bucket path of the source data as the crawler's data store. Allocate data processing units (DPUs) to run the crawler every day. Configure the output destination to a new path in the existing S3 bucket.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả tình huống thực tế trong AWS Glue:
Một công ty lưu trữ dữ liệu hàng ngày về hiệu suất danh mục đầu tư dưới dạng file .csv trong Amazon S3 bucket. Data engineer sử dụng AWS Glue crawlers để quét (crawl) dữ liệu từ S3.
Yêu cầu chính: Làm cho dữ liệu S3 này có thể truy cập hàng ngày trong AWS Glue Data Catalog (nơi lưu trữ metadata schema của dữ liệu).
📌 Mục tiêu: Crawler phải chạy hàng ngày (daily schedule) để tự động cập nhật schema vào Data Catalog, giúp các dịch vụ như Athena, Glue Jobs có thể query dữ liệu mà không cần crawl thủ công mỗi lần.
🛠️ Công nghệ liên quan (cập nhật đến 2026): AWS Glue Crawler (phiên bản mới nhất hỗ trợ S3, schema inference tự động cho CSV, tích hợp Lake Formation cho governance), IAM roles managed policies, và Data Catalog làm trung tâm metadata store. Không cần export data ra S3 khác vì crawler chỉ lưu metadata (schema, partitions), không phải raw data.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng:
Create an IAM role that includes the AWSGlueServiceRole policy. Associate the role with the crawler. Specify the S3 bucket path of the source data as the crawler's data store. Create a daily schedule to run the crawler. Specify a database name for the output.
Lý do chi tiết (best practice AWS):
- 🛡️ IAM role với AWSGlueServiceRole: Đây là managed policy chuẩn của AWS cho Glue services (bao gồm quyền crawl S3, write metadata vào Data Catalog). Policy này least privilege, an toàn hơn so với full access.
- 🔗 Associate role với crawler + S3 path: Crawler sẽ scan đúng bucket/path nguồn.
- ⏰ Daily schedule: Đảm bảo chạy tự động hàng ngày, cập nhật schema mới cho dữ liệu CSV daily.
- 📊 Specify database name for output: Metadata được lưu trực tiếp vào Glue Data Catalog database (ví dụ: "portfolio_db"), làm dữ liệu accessible ngay cho query.
✅ Hoàn hảo khớp yêu cầu, không thừa không thiếu. (Không cần DPU thủ công vì crawler dùng config mặc định 2 DPU, schedule tự scale).
🔍 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn một, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên docs AWS Glue mới nhất (2026: vẫn dùng AWSGlueServiceRole làm default, crawler output chỉ đến Catalog).
-
❌ Phương án SAI:
Create an IAM role that includes the AmazonS3FullAccess policy. Associate the role with the crawler. Specify the S3 bucket path of the source data as the crawler's data store. Create a daily schedule to run the crawler. Configure the output destination to a new path in the existing S3 bucket.
Giải thích sai: Policy AmazonS3FullAccess quá rộng (full control S3 toàn account), vi phạm least privilege principle – AWS khuyến cáo dùng AWSGlueServiceRole thay thế. Hơn nữa, output destination to S3 path hoàn toàn sai vì crawler KHÔNG lưu data/metadata ra S3 mà chỉ vào Data Catalog. Schedule đúng nhưng không cứu vãn được. -
✅ Phương án ĐÚNG:
Create an IAM role that includes the AWSGlueServiceRole policy. Associate the role with the crawler. Specify the S3 bucket path of the source data as the crawler's data store. Create a daily schedule to run the crawler. Specify a database name for the output.
Giải thích đúng: Như đã phân tích ở trên – full match best practice: policy chuẩn, schedule daily, output trực tiếp đến database trong Data Catalog để accessible ngay. -
❌ Phương án SAI:
Create an IAM role that includes the AmazonS3FullAccess policy. Associate the role with the crawler. Specify the S3 bucket path of the source data as the crawler's data store. Allocate data processing units (DPUs) to run the crawler every day. Specify a database name for the output.
Giải thích sai: Policy AmazonS3FullAccess sai (quá rộng, không cần thiết). Allocate DPUs every day thừa thãi – crawler có config DPU mặc định (min 2), schedule tự chạy với config đó, không cần allocate thủ công cho daily run. Output database đúng nhưng policy và DPU làm sai toàn bộ. -
❌ Phương án SAI:
Create an IAM role that includes the AWSGlueServiceRole policy. Associate the role with the crawler. Specify the S3 bucket path of the source data as the crawler's data store. Allocate data processing units (DPUs) to run the crawler every day. Configure the output destination to a new path in the existing S3 bucket.
Giải thích sai: Policy đúng nhưng allocate DPUs every day không chuẩn (schedule dùng crawler config sẵn). Output to S3 path sai cơ bản – crawler chỉ output metadata đến Catalog, không phải S3 (S3 chỉ là input).
📘 Tài liệu tham khảo (AWS Docs chính thức, cập nhật 2026)
- 🛡️ IAM Policies for AWS Glue Crawlers – Xác nhận AWSGlueServiceRole là default cho crawler.
- ⏰ Scheduling Crawlers – Hỗ trợ cron daily, không cần DPU manual.
- 📊 Crawler Output to Data Catalog – Chỉ định database name, metadata lưu Catalog (không S3).
- 🔗 Create Crawler Tutorial – Full steps khớp đáp án đúng.
Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ code Terraform/CloudFormation, hỏi nhé!
A data engineer wants to store the load statuses of Redshift tables in an Amazon DynamoDB table. The data engineer creates an AWS Lambda function to publish the details of the load statuses to DynamoDB.
How should the data engineer invoke the Lambda function to write load statuses to the DynamoDB table?
- A Use a second Lambda function to invoke the first Lambda function based on Amazon CloudWatch events.
- B Use the Amazon Redshift Data API to publish an event to Amazon EventBridge. Configure an EventBridge rule to invoke the Lambda function.
- C Use the Amazon Redshift Data API to publish a message to an Amazon Simple Queue Service (Amazon SQS) queue. Configure the SQS queue to invoke the Lambda function.
- D Use a second Lambda function to invoke the first Lambda function based on AWS CloudTrail events.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh một công ty đang load dữ liệu giao dịch hàng ngày vào các bảng Amazon Redshift vào cuối mỗi ngày. Họ cần theo dõi trạng thái load (đã load hay chưa load) của các bảng. Một data engineer muốn lưu trạng thái này vào bảng Amazon DynamoDB bằng cách sử dụng AWS Lambda function để publish chi tiết trạng thái.
Vấn đề cốt lõi: Làm thế nào để invoke (kích hoạt) Lambda function một cách hiệu quả để ghi trạng thái load vào DynamoDB? Câu hỏi tập trung vào tích hợp tự động giữa Redshift và các dịch vụ AWS khác, đảm bảo quy trình serverless, scalable và không cần quản lý infrastructure thủ công. Đây là kịch bản điển hình trong ETL pipelines trên AWS, nơi Redshift thường dùng cho data warehouse, và cần notify trạng thái sau khi load hoàn tất (ví dụ: sau lệnh COPY hoặc UNLOAD).
Bối cảnh kỹ thuật (cập nhật đến 2026): Amazon Redshift hỗ trợ Data API (ra mắt 2019, cập nhật liên tục) cho phép thực thi query/load mà không cần kết nối trực tiếp, và tích hợp native với Amazon EventBridge để publish events khi các hoạt động như COPY (load data) hoàn tất. Điều này giúp trigger Lambda mà không cần polling hay intermediate services phức tạp.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use the Amazon Redshift Data API to publish an event to Amazon EventBridge. Configure an EventBridge rule to invoke the Lambda function.
Lý do:
- Redshift Data API hỗ trợ publish events tự động đến Amazon EventBridge khi các hoạt động như load data (COPY command) hoàn tất hoặc thất bại. Event này chứa metadata như table name, status (success/failure), timestamp – hoàn hảo để track load statuses.
- EventBridge rule có thể filter event cụ thể (ví dụ: source từ Redshift, detail-type là "Redshift Data API Execution") và invoke Lambda trực tiếp, đảm bảo event-driven architecture serverless, low-latency, và scalable.
- Ưu điểm: Không cần code thêm, chi phí thấp (pay-per-event), tích hợp native từ AWS (không deprecated). Đây là best practice theo AWS Well-Architected Framework cho data pipelines (Reliability pillar).
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên văn bản gốc tiếng Anh của phương án, chỉ dịch giải thích bằng tiếng Việt để dễ hiểu:
-
✅ Use the Amazon Redshift Data API to publish an event to Amazon EventBridge. Configure an EventBridge rule to invoke the Lambda function.
🛠️ Đúng vì: Như giải thích trên, đây là tích hợp native và tự động của Redshift Data API với EventBridge (từ phiên bản Redshift ra mắt 2021, cập nhật 2025 với enhanced event schemas). Lambda nhận event → query status nếu cần → update DynamoDB. Hiệu quả cao, không polling. -
❌ Use a second Lambda function to invoke the first Lambda function based on Amazon CloudWatch events.
🧩 Sai vì: CloudWatch Events (nay là EventBridge) không có native trigger từ Redshift load completion mà không qua Data API. Sử dụng Lambda thứ hai để "bridge" sẽ tạo circular dependency, tăng complexity, chi phí (2 Lambdas), và không scalable. Không phải best practice cho real-time status tracking. -
❌ Use the Amazon Redshift Data API to publish a message to an Amazon Simple Queue Service (Amazon SQS) queue. Configure the SQS queue to invoke the Lambda function.
🛠️ Sai vì: Redshift Data API không hỗ trợ publish trực tiếp đến SQS (chỉ EventBridge hoặc custom code). SQS phù hợp cho decoupling nhưng ở đây thêm latency (polling queue), overhead (dead-letter queue), và không tận dụng event schema phong phú của EventBridge. Không hiệu quả cho near-real-time tracking. -
❌ Use a second Lambda function to invoke the first Lambda function based on AWS CloudTrail events.
❌ Sai vì: CloudTrail ghi audit logs (API calls như DescribeClusters), không phải data load events từ Redshift tables. Load status (table-level) không có trong CloudTrail events, dẫn đến không chính xác và delay cao (CloudTrail batch ~15 phút). Phù hợp audit/security, không phải operational tracking.
📘 Tài liệu tham khảo (cập nhật mới nhất 2026)
- AWS Redshift Data API Documentation: Executing queries with Data API – Phần "Event integration with EventBridge".
- Amazon EventBridge + Redshift Integration: Redshift Events in EventBridge (schema events cho COPY/UNLOAD status).
- AWS Well-Architected: Data Analytics Lens (2025 edition): Khuyến nghị event-driven cho ETL status tracking.
- Sample Code: AWS Samples GitHub – Redshift Data API to EventBridge Lambda.
Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần ví dụ code CloudFormation, hãy hỏi thêm nhé!
Which AWS service should the data engineer use to transfer the data in the MOST operationally efficient way?
- A AWS DataSync
- B AWS Glue
- C AWS Direct Connect
- D Amazon S3 Transfer Acceleration
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một data engineer cần chuyển 5 TB dữ liệu từ data center on-premises sang Amazon S3 bucket một cách bảo mật. Đặc điểm nổi bật:
- Khoảng 5% dữ liệu thay đổi hàng ngày ➡️ Cần hỗ trợ chuyển dữ liệu delta (chỉ phần thay đổi) để tránh chuyển toàn bộ lặp lại.
- Cập nhật định kỳ và tự động hóa quy trình với lịch chạy (scheduling).
- Dữ liệu đa dạng nhiều định dạng file (không phải chỉ một loại).
- Yêu cầu hiệu quả hoạt động cao nhất (MOST operationally efficient): Nghĩa là tối ưu chi phí, thời gian, dễ quản lý, tự động và bảo mật.
📘 Tài liệu tham khảo chính:
- AWS DataSync Documentation (cập nhật 2024-2026): https://docs.aws.amazon.com/datasync/latest/userguide/what-is-datasync.html
- AWS Well-Architected Framework - Operational Excellence Pillar (2025 edition).
✅ Đáp án đúng: AWS DataSync
Lý do lựa chọn 🛠️:
- AWS DataSync là dịch vụ chuyên biệt cho việc chuyển dữ liệu lớn giữa on-premises và AWS (hỗ trợ S3 làm đích đến).
- Hỗ trợ incremental sync (chỉ chuyển 5% thay đổi hàng ngày) qua tính năng delta transfer, giảm băng thông và thời gian.
- Tự động hóa và scheduling: Tích hợp với Amazon EventBridge (trước là CloudWatch Events) để chạy lịch định kỳ (daily/hourly).
- Bảo mật cao: Sử dụng AWS PrivateLink, mã hóa dữ liệu (TLS 1.2+), IAM roles, và VPC endpoints.
- Đa định dạng: Hỗ trợ bất kỳ loại file nào mà không cần ETL.
- Hiệu quả nhất: Quản lý qua console/CLI/API, monitoring qua CloudWatch, chi phí theo GB chuyển (rẻ hơn các lựa chọn khác cho trường hợp này). Phù hợp với dữ liệu 5 TB + thay đổi liên tục.
📋 Giải thích tất cả các phương án (sử dụng emoji đánh dấu đúng/sai)
-
✅ AWS DataSync
Đúng vì đây là giải pháp tối ưu nhất cho yêu cầu: Chuyển dữ liệu on-prem sang S3 với agent cài trên server on-prem, hỗ trợ sync liên tục/delta, lập lịch tự động, và xử lý đa định dạng mà không cần code. Đáp ứng đầy đủ "MOST operationally efficient" với monitoring tích hợp và bảo mật end-to-end. (Nguồn: AWS DataSync User Guide - Features 2026). -
❌ AWS Glue
Sai vì AWS Glue là dịch vụ ETL (Extract, Transform, Load) serverless, dùng cho xử lý dữ liệu lớn trên S3/Data Lake với Spark/SQL, không phải để chuyển dữ liệu on-prem. Nó không hỗ trợ delta sync on-prem trực tiếp, cần connector phức tạp (như JDBC), và không tối ưu cho file-based transfer định kỳ mà không transform. Phù hợp hơn cho analytics, không phải sync file. (Nguồn: AWS Glue Developer Guide - Crawlers & Jobs). -
❌ AWS Direct Connect
Sai vì AWS Direct Connect là kết nối mạng riêng dedicated (fiber optic) từ on-prem sang AWS VPC/Direct S3, chỉ cung cấp băng thông ổn định cao chứ không tự động hóa transfer hay delta sync. Không có scheduling tích hợp, không xử lý file/format, và yêu cầu setup phần cứng phức tạp/đắt đỏ cho 5 TB + daily updates. Chỉ là "ống dẫn" mạng, không phải dịch vụ transfer. (Nguồn: AWS Direct Connect User Guide - Dedicated Connections). -
❌ Amazon S3 Transfer Acceleration
Sai vì đây là tính năng tăng tốc upload/download qua edge locations (CloudFront), chỉ tối ưu Internet public transfer từ client/end-user sang S3, không hỗ trợ on-prem agent hay delta sync. Không tự động hóa scheduling, không xử lý thay đổi hàng ngày hiệu quả, và kém bảo mật hơn cho dữ liệu lớn/private (dùng public endpoints). Phù hợp upload ad-hoc, không phải sync định kỳ. (Nguồn: Amazon S3 User Guide - Transfer Acceleration).
Kết luận 🎯: AWS DataSync là lựa chọn hiệu quả vận hành cao nhất (operational excellence), giúp data engineer tiết kiệm thời gian và chi phí lâu dài! Nếu triển khai, khuyến nghị dùng DataSync agent v3+ (2025) cho hiệu suất tốt hơn.
The company requires a cost-effective solution to migrate the data to AWS. The solution must cause minimal downtown for the applications that access the database.
Which AWS service should the company use to meet these requirements?
- A AWS Lambda
- B AWS Database Migration Service (AWS DMS)
- C AWS Direct Connect
- D AWS DataSync
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh một công ty đang sử dụng Microsoft SQL Server on-premises để lưu trữ dữ liệu giao dịch tài chính. Họ di chuyển dữ liệu lên AWS (cụ thể là Amazon RDS for SQL Server) vào cuối mỗi tháng. Gần đây, chi phí di chuyển dữ liệu tăng cao, nên cần một giải pháp tiết kiệm chi phí, đồng thời giảm thiểu thời gian downtime cho các ứng dụng truy cập cơ sở dữ liệu.
🛠️ Yêu cầu chính:
- Cost-effective: Giảm chi phí truyền dữ liệu (data transfer costs).
- Minimal downtime: Không làm gián đoạn ứng dụng, hỗ trợ migration liên tục hoặc gần thời gian thực.
- Homogeneous migration: Từ on-premises SQL Server sang RDS SQL Server (cùng engine).
📘 Bối cảnh AWS: Đây là kịch bản migration database điển hình, nơi cần công cụ hỗ trợ replication dữ liệu với chi phí thấp và downtime thấp (AWS cập nhật DMS đến 2026 vẫn là lựa chọn hàng đầu cho DB migration).
✅ Đáp án đúng: AWS Database Migration Service (AWS DMS)
Lý do chọn:
AWS DMS là dịch vụ chuyên dụng cho di chuyển và replicate dữ liệu database giữa các nguồn on-premises và AWS (như RDS). Nó hỗ trợ full load + ongoing replication (di chuyển toàn bộ + đồng bộ liên tục), giúp minimal downtime (downtime chỉ vài phút). Với SQL Server → RDS SQL Server, DMS sử dụng native CDC (Change Data Capture) để replicate thay đổi thời gian thực, giảm chi phí data transfer so với export/import thủ công (chỉ chuyển delta changes). DMS miễn phí engine, chỉ tính phí storage/instance, rất cost-effective cho migration hàng tháng.
🛠️ Lợi ích nổi bật: Hỗ trợ schema conversion nếu cần, multi-region, và tích hợp VPC/Direct Connect để tối ưu network costs (cập nhật 2026: DMS hỗ trợ DMS Fleet Advisor cho assessment tự động).
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh:
-
❌ AWS Lambda
Sai vì: AWS Lambda là dịch vụ serverless compute chạy code theo sự kiện, không thiết kế cho database migration lớn hoặc replication liên tục. Sử dụng Lambda để script export/import dữ liệu sẽ tốn kém (thời gian chạy dài, memory cao) và downtime cao (không hỗ trợ CDC realtime). Không phù hợp cho dữ liệu tài chính lớn hàng tháng, dễ vượt giới hạn execution time (15 phút). -
✅ AWS Database Migration Service (AWS DMS)
Đúng vì: Như giải thích ở trên, DMS tối ưu cho homogeneous DB migration với minimal downtime qua ongoing replication. Giảm chi phí bằng cách chỉ sync changes (không full dump hàng tháng), hỗ trợ SQL Server native. Cost-effective nhất cho kịch bản này (chỉ ~$0.018/giờ/instance + storage). -
❌ AWS Direct Connect
Sai vì: AWS Direct Connect là dịch vụ kết nối mạng riêng tư (dedicated fiber) từ on-premises đến AWS, giúp giảm latency và data transfer costs qua public internet. Tuy nhiên, nó không phải công cụ migration dữ liệu – chỉ là "đường ống" truyền dữ liệu. Vẫn cần tool khác (như DMS) để extract/replicate DB, không giải quyết downtime hoặc migration logic. -
❌ AWS DataSync
Sai vì: AWS DataSync dành cho di chuyển file/object storage (NFS/SMB/S3), không hỗ trợ relational database như SQL Server (không hiểu schema, transactions, CDC). Dùng cho DataSync sẽ yêu cầu export DB thành file (dump), gây downtime cao và chi phí data transfer đầy đủ (không delta sync), không cost-effective cho DB transactional.
📘 Tài liệu tham khảo (cập nhật AWS 2026)
- AWS DMS Documentation: https://docs.aws.amazon.com/dms/latest/userguide/Welcome.html (Hướng dẫn migration SQL Server → RDS, CDC support).
- AWS Database Migration Best Practices: https://aws.amazon.com/blogs/database/best-practices-for-migrating-microsoft-sql-server-to-amazon-rds/ (Tiết kiệm chi phí với DMS replication).
- Pricing DMS: https://aws.amazon.com/dms/pricing/ (Xác nhận cost model miễn phí engine).
- AWS Well-Architected Framework - Migration: Lens cập nhật 2026 nhấn mạnh DMS cho minimal downtime DB workloads.
🛠️ Khuyến nghị thực tế: Kết hợp DMS với AWS Schema Conversion Tool (SCT) nếu có schema changes, và Direct Connect để tối ưu network cho migration lớn!