Ngân hàng đề — AWS Certified Data Engineer Associate
Tìm thấy 867 câu.
The company has created a new data product that includes a group of Amazon Redshift Serverless tables. A data engineer needs to share the data product with a marketing team. The marketing team must have access to only a subset of columns. The data engineer needs to share the same data product with a compliance team. The compliance team must have access to a different subset of columns than the marketing team needs access to.
Which combination of steps should the data engineer take to meet these requirements? (Choose two.)
- A Create views of the tables that need to be shared. Include only the required columns.
- B Create an Amazon Redshift data share that includes the tables that need to be shared.
- C Create an Amazon Redshift managed VPC endpoint in the marketing team’s account. Grant the marketing team access to the views.
- D Share the Amazon Redshift data share to the Lake Formation catalog in the governance account.
- E Share the Amazon Redshift data share to the Amazon Redshift Serverless workgroup in the marketing team's account.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi thuộc chủ đề Data Mesh trên AWS, tập trung vào việc quản lý và chia sẻ dữ liệu trung tâm qua AWS Lake Formation trong một central governance account. Công ty đã thiết lập data mesh với tài khoản governance để catalog toàn bộ dữ liệu và cấp quyền truy cập tập trung.
-
Tình huống cụ thể: Một data product mới bao gồm nhóm bảng Amazon Redshift Serverless. Data engineer cần chia sẻ data product này với hai đội:
- Marketing team: Chỉ truy cập subset columns cụ thể.
- Compliance team: Truy cập subset columns khác (không trùng lặp).
-
Yêu cầu chính: Sử dụng cơ chế chia sẻ dữ liệu cross-account (vì teams ở các account khác), hỗ trợ fine-grained access control ở mức column-level (cột cụ thể), và tích hợp với Lake Formation để governance trung tâm. Cần chọn TWO steps phù hợp nhất theo kiến thức AWS cập nhật đến 2026 (hỗ trợ Redshift Data Sharing với Lake Formation integration cho Serverless namespaces).
Mục tiêu là chia sẻ dữ liệu không copy (zero-copy sharing), giữ dữ liệu gốc ở Redshift Serverless, và dùng Lake Formation để catalog + grant permissions khác nhau cho từng team mà không cần duplicate data.
✅ Đáp án đúng (Chọn TWO)
Hai bước đúng là:
- Create an Amazon Redshift data share that includes the tables that need to be shared.
- Share the Amazon Redshift data share to the Lake Formation catalog in the governance account.
Lý do lựa chọn 🛠️:
- Redshift Data Sharing (cập nhật 2024-2026) cho phép chia sẻ tables/views từ Redshift Serverless cross-account mà không sao chép dữ liệu (zero-ETL). Điều này phù hợp với data mesh để tạo data product.
- Sau khi tạo Data Share, share nó vào Lake Formation catalog ở governance account cho phép register Data Share như tables trong LF. Lake Formation hỗ trợ column-level permissions (qua LF permissions hoặc tags) để grant subset columns khác nhau cho marketing và compliance teams một cách tập trung, an toàn. Các team có thể query qua Redshift Spectrum hoặc Athena mà không cần truy cập trực tiếp Redshift producer account.
- Kết hợp này đảm bảo central governance qua LF, hỗ trợ multi-tenant access với fine-grained control, phù hợp DOP-C02 exam.
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm giải thích bằng tiếng Việt:
-
❌ Create views of the tables that need to be shared. Include only the required columns.
Phương án này sai vì views chỉ filter columns ở local cluster, không hỗ trợ cross-account sharing hiệu quả mà không copy data. Trong data mesh với LF governance, views không tích hợp trực tiếp để grant permissions tập trung cho multiple teams với subsets khác nhau. Phải dùng Data Shares để share zero-copy. -
✅ Create an Amazon Redshift data share that includes the tables that need to be shared.
Phương án này đúng vì Redshift Data Shares (Serverless hỗ trợ đầy đủ từ 2023+) cho phép share tables trực tiếp cross-account. Đây là bước đầu tiên để tạo data product có thể catalog trong LF, hỗ trợ query từ consumer accounts mà giữ data immutable và secure. -
❌ Create an Amazon Redshift managed VPC endpoint in the marketing team’s account. Grant the marketing team access to the views.
Phương án này sai vì Redshift managed VPC endpoints chỉ dùng để truy cập Redshift từ private VPC (network isolation), không giải quyết data sharing hoặc column-level permissions. Nó phụ thuộc views (đã sai), và không tích hợp LF governance cho multiple teams/subsets. -
✅ Share the Amazon Redshift data share to the Lake Formation catalog in the governance account.
Phương án này đúng vì sau khi tạo Data Share, producer share nó vào LF catalog (quaGRANThoặc register). LF trở thành trung tâm để catalog data products, apply LF permissions (row/column/cell-level) khác nhau cho marketing/compliance teams. Consumer teams query qua Athena/Redshift Spectrum với central governance. -
❌ Share the Amazon Redshift data share to the Amazon Redshift Serverless workgroup in the marketing team's account.
Phương án này sai vì share Data Share trực tiếp đến consumer Redshift workgroup chỉ cho phép query full tables, không hỗ trợ column-level subsets khác nhau cho multiple teams qua LF. Nó bỏ qua central governance account và LF catalog, vi phạm yêu cầu data mesh.
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- AWS Docs - Redshift Data Sharing: Redshift Data Shares – Hỗ trợ Serverless và LF integration.
- AWS Lake Formation - Redshift Integration: Register Redshift Data Shares in LF – Chi tiết register và permissions.
- Data Mesh on AWS Whitepaper (2024): AWS Data Mesh – Best practices cho governance với LF.
- DOP-C02 Exam Guide: Domain 5 – Data Mesh, LF, Redshift sharing (giải thích zero-copy cross-account).
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm ví dụ code hoặc demo, hãy hỏi nhé!
The company needs to ensure that data quality issues are checked every time the pipelines run. A data engineer must enhance the existing pipelines to evaluate data quality rules based on predefined thresholds.
Which solution will meet these requirements with the LEAST implementation effort?
- A Add a new transform that is defined by a SQL query to each Glue ETL job. Use the SQL query to implement a ruleset that includes the data quality rules that need to be evaluated.
- B Add a new Evaluate Data Quality transform to each Glue ETL job. Use Data Quality Definition Language (DQDL) to implement a ruleset that includes the data quality rules that need to be evaluated.
- C Add a new custom transform to each Glue ETL job. Use the PyDeequ library to implement a ruleset that includes the data quality rules that need to be evaluated.
- D Add a new custom transform to each Glue ETL job. Use the Great Expectations library to implement a ruleset that includes the data quality rules that need to be evaluated.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc tối ưu hóa pipeline ETL trong AWS Glue Studio cho một data lake trên Amazon S3. Công ty đã sử dụng AWS Glue để catalog dữ liệu và Glue Studio để xây dựng pipeline ETL (extract, transform, load). Yêu cầu chính là kiểm tra chất lượng dữ liệu (data quality) mỗi lần pipeline chạy, dựa trên các quy tắc (rules) và ngưỡng (thresholds) được định nghĩa trước. Data engineer cần nâng cấp pipeline hiện tại một cách ít nỗ lực triển khai nhất (LEAST implementation effort).
🔍 Các yếu tố chính cần xem xét:
- Phải tích hợp kiểm tra data quality một cách tự động và native vào Glue ETL jobs.
- Ưu tiên giải pháp built-in của AWS Glue, tránh custom code để giảm effort (viết code, maintain, debug).
- AWS Glue hỗ trợ Data Quality features từ phiên bản mới nhất (cập nhật đến 2026), với transform chuyên dụng cho việc này.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Add a new Evaluate Data Quality transform to each Glue ETL job. Use Data Quality Definition Language (DQDL) to implement a ruleset that includes the data quality rules that need to be evaluated.
🛠️ Lý do chọn đáp án này:
- AWS Glue Studio cung cấp Evaluate Data Quality transform là một visual transform native, dễ dàng kéo-thả vào pipeline mà không cần viết code phức tạp.
- Sử dụng DQDL (Data Quality Definition Language) – ngôn ngữ declarative đơn giản của AWS – để định nghĩa ruleset (ví dụ: kiểm tra nulls, duplicates, schema, completeness với thresholds).
- Least effort: Chỉ cần thêm transform vào Glue job hiện tại, config rules qua UI hoặc JSON DQDL, chạy ngay lập tức. Hỗ trợ monitoring qua Glue console và CloudWatch.
- Tích hợp sâu với Glue catalog, tự động apply cho dynamic frames trong ETL.
📋 Phân tích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên effort triển khai, tính native và phù hợp với yêu cầu least effort.
-
❌ Phương án SAI: Add a new transform that is defined by a SQL query to each Glue ETL job. Use the SQL query to implement a ruleset that includes the data quality rules that need to be evaluated.
Giải thích: Phương án này yêu cầu viết SQL query custom (ví dụ: dùng Spark SQL để count nulls, aggregates), không phải transform native cho data quality. Effort cao vì phải tự thiết kế ruleset thủ công, handle thresholds phức tạp, và debug SQL errors. Không hỗ trợ DQDL declarative, dễ lỗi khi scale data lớn. Không phải giải pháp least effort. -
✅ Phương án ĐÚNG: Add a new Evaluate Data Quality transform to each Glue ETL job. Use Data Quality Definition Language (DQDL) to implement a ruleset that includes the data quality rules that need to be evaluated.
Giải thích: Như đã nêu ở phần đáp án đúng. Đây là transform built-in từ AWS Glue 3.0+ (cập nhật 2026), hỗ trợ >20 rules sẵn (completeness, uniqueness, validity,...), thresholds (% allowed), và output metrics vào S3/CloudWatch. Least effort: Config qua Glue Studio UI, no coding needed. -
❌ Phương án SAI: Add a new custom transform to each Glue ETL job. Use the PyDeequ library to implement a ruleset that includes the data quality rules that need to be evaluated.
Giải thích: PyDeequ (dựa trên Deequ của Amazon) yêu cầu custom PySpark code trong transform, import library, viết verification suite. Effort cao: Phải code chi tiết, handle dependencies (Glue hỗ trợ nhưng cần script riêng), maintain khi library update. Không native UI như DQ transform, vi phạm "least effort". -
❌ Phương án SAI: Add a new custom transform to each Glue ETL job. Use the Great Expectations library to implement a ruleset that includes the data quality rules that need to be evaluated.
Giải thích: Great Expectations là third-party library (Python-based), cần custom code đầy đủ (expectations suite, checkpoint). Effort rất cao: Install deps trong Glue job, viết YAML/JSON config, integrate với Spark. Không tích hợp native với Glue Studio, khó scale và monitor so với DQDL. Không phù hợp least effort.
📘 Tài liệu tham khảo (AWS cập nhật mới nhất đến 2026)
- AWS Glue Data Quality Documentation: https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-data-quality.html – Chi tiết EvaluateDataQuality transform và DQDL syntax.
- AWS Glue Studio User Guide: https://docs.aws.amazon.com/glue/latest/dg/glue-studio.html – Hướng dẫn visual ETL với Data Quality nodes.
- AWS re:Post & Blog: Tìm "Glue Data Quality DQDL" cho best practices (ví dụ: blog 2023-2026 về enhancements như ML-based rules).
- Exam Tip (DevOps Pro DOP-C02): Chủ đề ETL optimization thường ưu tiên native services để giảm operational overhead.
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần demo code DQDL, hãy hỏi thêm nhé!
The company wants to set up a robust monitoring system for the application. The company needs to analyze the logs from the EKS cluster and the application. The company needs to correlate the cluster's logs with the application's traces to identify points of failure in the whole application request flow.
Which combination of steps will meet these requirements with the LEAST development effort? (Choose two.)
- A Use FluentBit to collect logs. Use OpenTelemetry to collect traces.
- B Use Amazon CloudWatch to collect logs. Use Amazon Kinesis to collect traces.
- C Use Amazon CloudWatch to collect logs. Use Amazon Managed Streaming for Apache Kafka (Amazon MSK) to collect traces.
- D Use Amazon OpenSearch to correlate the logs and traces.
- E Use AWS Glue to correlate the logs and traces.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc thiết lập hệ thống giám sát robust (mạnh mẽ) cho ứng dụng microservice chạy trên Amazon EKS cluster (Kubernetes được quản lý bởi AWS). 🛤️
Công ty cần:
- Thu thập và phân tích logs từ cluster EKS và ứng dụng (bao gồm logs container, pod, node).
- Thu thập traces từ ứng dụng để theo dõi request flow qua các microservice.
- Correlate (liên kết) logs của cluster với traces của app để xác định points of failure (điểm thất bại) trong toàn bộ luồng request.
Yêu cầu chọn TWO steps (hai bước kết hợp) với LEAST development effort (ít nỗ lực phát triển nhất), nghĩa là ưu tiên các giải pháp native AWS, add-on sẵn có cho EKS, không cần code custom nhiều.
📈 Chủ đề liên quan: Observability trên EKS (logs, metrics, traces), sử dụng các tool chuẩn AWS đến năm 2026 như Fluent Bit, OpenTelemetry, và Amazon OpenSearch cho phân tích tương quan dữ liệu.
✅ Đáp án đúng (Chọn TWO)
Hai phương án đúng là:
- Use FluentBit to collect logs. Use OpenTelemetry to collect traces.
- Use Amazon OpenSearch to correlate the logs and traces.
Lý do lựa chọn:
✅ Kết hợp này mang lại least development effort vì:
- Fluent Bit là fluent-bit add-on chính thức của EKS (từ AWS), tự động thu thập logs từ pods/nodes mà không cần config phức tạp. Nó ship logs trực tiếp đến CloudWatch Logs hoặc OpenSearch.
- OpenTelemetry (OTel) là standard mở được AWS hỗ trợ đầy đủ qua AWS Distro for OpenTelemetry (ADOT) trên EKS, dễ deploy collector để thu thập traces/metrics với annotation đơn giản trong Deployment YAML.
- Amazon OpenSearch (phiên bản mới nhất 2026) hỗ trợ Trace Analytics native, tự động correlate logs + traces (qua fields như trace_id, span_id) trong OpenSearch Dashboards, không cần ETL custom. Chỉ cần enable data prep pipelines để parse và index.
🛠️ Toàn bộ là managed services, deploy qua Helm charts hoặc EKS add-ons, giảm effort so với self-managed.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng phương án một cách chi tiết (giữ nguyên văn bản gốc bằng tiếng Anh). Mỗi cái được đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm lý do bằng tiếng Việt dựa trên best practices AWS EKS observability 2026.
-
Use FluentBit to collect logs. Use OpenTelemetry to collect traces.
✅ Đúng. Fluent Bit là log forwarder lightweight được AWS khuyến nghị chính thức cho EKS (qua EKS add-on), hỗ trợ parse/filter logs Kubernetes native (ví dụ: CRI-O logs). OpenTelemetry Collector trên EKS (qua ADOT operator) thu thập traces dễ dàng với auto-instrumentation cho Java/Go/Node.js. Kết hợp này chuẩn cho least effort, không cần dev code nhiều. 🏆 -
Use Amazon CloudWatch to collect logs. Use Amazon Kinesis to collect traces.
❌ Sai. CloudWatch Logs Agent có thể collect logs từ EKS nhưng phức tạp hơn Fluent Bit (cần DaemonSet custom, config IAM roles). Kinesis là streaming data service, không hỗ trợ traces (thiếu trace analytics), phải build pipeline custom để parse traces → high development effort, không correlate dễ dàng. 🚫 -
Use Amazon CloudWatch to collect logs. Use Amazon Managed Streaming for Apache Kafka (Amazon MSK) to collect traces.
❌ Sai. Tương tự trên, CloudWatch Logs ok nhưng không optimal cho EKS. MSK là managed Kafka cho event streaming, không phải tool cho traces (cần Kafka Connect sinks custom để export traces), không có correlate built-in → effort cao, không phù hợp observability traces. 📉 -
Use Amazon OpenSearch to correlate the logs and traces.
✅ Đúng. OpenSearch Service (với OpenSearch Dashboards) có Trace Analytics và Log Analytics native từ 2021-2026, hỗ trợ PPL (Piped Processing Language) để query/join logs-traces qua trace_id. Fluent Bit/OTel exporter trực tiếp push dữ liệu vào OpenSearch index → zero-code correlate, dashboard sẵn. Hoàn hảo cho EKS full request flow. 🔍 -
Use AWS Glue to correlate the logs and traces.
❌ Sai. AWS Glue là ETL service cho big data (batch/Spark jobs), không dành cho real-time observability. Correlate logs-traces cần schema phức tạp (trace spans), Glue yêu cầu ETL scripts → high dev effort, latency cao, không real-time như OpenSearch. Không khuyến nghị cho monitoring. ⚠️
📘 Tài liệu tham khảo (Cập nhật AWS 2026)
- EKS Add-ons & Fluent Bit: AWS Docs - Amazon EKS add-ons: Fluent Bit 🛠️
- OpenTelemetry on EKS: AWS Distro for OpenTelemetry & EKS Observability Workshop 📊
- OpenSearch Trace Analytics: Amazon OpenSearch Service - Trace Analytics 🔗
- Best Practices EKS Monitoring: AWS Well-Architected Framework - Observability Pillar (2026 edition).
Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ config YAML, hãy hỏi nhé.
Which solution will meet these requirements?
- A Use AWS Step Functions to periodically export data from the Amazon DynamoDB tables to an Amazon S3 bucket. Use an AWS Lambda function to load the data into Amazon OpenSearch Service.
- B Configure an AWS Glue job to have a source of Amazon DynamoDB and a destination of Amazon OpenSearch Service to transfer data in near real time.
- C Use Amazon DynamoDB Streams to capture table changes. Use an AWS Lambda function to process and update the data in Amazon OpenSearch Service.
- D Use a custom OpenSearch plugin to sync data from the Amazon DynamoDB tables.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh một ứng dụng game của công ty lưu trữ dữ liệu trong các bảng Amazon DynamoDB. Nhiệm vụ của data engineer là ingest (tiêm dữ liệu) game data vào Amazon OpenSearch Service cluster, với yêu cầu quan trọng là cập nhật dữ liệu phải diễn ra gần như thời gian thực (near real-time).
📌 Yêu cầu cốt lõi:
- Dữ liệu từ DynamoDB cần được đồng bộ hóa nhanh chóng vào OpenSearch để hỗ trợ tìm kiếm, phân tích log hoặc dashboard thời gian thực cho ứng dụng game.
- Giải pháp phải tự động capture thay đổi (insert, update, delete) từ DynamoDB và đẩy vào OpenSearch mà không có độ trễ lớn (near real-time thường nghĩa là giây hoặc dưới phút).
- Đây là tình huống điển hình trong AWS cho change data capture (CDC) từ NoSQL database sang search engine.
🛠️ Bối cảnh AWS cập nhật đến 2026: DynamoDB hỗ trợ Streams cho near real-time replication. OpenSearch Service (fork từ Elasticsearch) tích hợp tốt với Lambda và Streams để xử lý dữ liệu động.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Amazon DynamoDB Streams to capture table changes. Use an AWS Lambda function to process and update the data in Amazon OpenSearch Service.
Lý do chọn đáp án này 🏆:
- DynamoDB Streams capture mọi thay đổi (new image, old image, hoặc cả hai) trên bảng DynamoDB gần như ngay lập tức (trong vòng vài giây), lưu trữ stream records lên đến 24 giờ.
- AWS Lambda trigger trực tiếp từ Streams, xử lý record (transform data nếu cần) và index vào OpenSearch qua OpenSearch API (Bulk API hoặc _index endpoint).
- Giải pháp serverless, scalable, fault-tolerant, chi phí thấp (pay-per-request), phù hợp near real-time cho ứng dụng game cao tải.
- Không cần polling hay batch job, đảm bảo exactly-once semantics với idempotent writes vào OpenSearch.
📋 Phân tích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm giải thích đầy đủ bằng tiếng Việt dựa trên best practices AWS.
-
Use AWS Step Functions to periodically export data from the Amazon DynamoDB tables to an Amazon S3 bucket. Use an AWS Lambda function to load the data into Amazon OpenSearch Service.
❌ Sai: Step Functions chạy định kỳ (periodic), export toàn bộ hoặc delta data sang S3 gây độ trễ lớn (phút đến giờ), không đạt near real-time. S3 là batch storage, Lambda load từ S3 thêm overhead. Không hiệu quả cho game data động, tốn chi phí scan DynamoDB lớn. -
Configure an AWS Glue job to have a source of Amazon DynamoDB and a destination of Amazon OpenSearch Service to transfer data in near real time.
❌ Sai: AWS Glue là ETL batch-oriented, chạy theo schedule (triggered job), không hỗ trợ near real-time (thời gian chạy job từ phút đến giờ). Glue crawler/connector cho DynamoDB là snapshot-based, không capture changes liên tục. OpenSearch sink trong Glue chưa tối ưu cho streaming đến 2026. -
Use Amazon DynamoDB Streams to capture table changes. Use an AWS Lambda function to process and update the data in Amazon OpenSearch Service.
✅ Đúng: Như giải thích ở trên, đây là architecture chuẩn AWS cho CDC near real-time. Streams + Lambda event-driven, scale tự động, hỗ trợ fan-out với Kinesis nếu cần. Đã được AWS recommend cho DynamoDB-to-OpenSearch sync. -
Use a custom OpenSearch plugin to sync data from the Amazon DynamoDB tables.
❌ Sai: Không có plugin chính thức từ AWS cho sync trực tiếp DynamoDB-OpenSearch. Custom plugin yêu cầu phát triển phức tạp (Java code trên OpenSearch node), cần polling DynamoDB (tốn RCU), không serverless, khó scale/maintain. Vi phạm best practices, rủi ro bảo mật và downtime cao.
📘 Tài liệu tham khảo (AWS cập nhật mới nhất 2026)
- DynamoDB Streams: AWS Docs - Capturing Table Activity with DynamoDB Streams – Xác nhận near real-time capture.
- Lambda với DynamoDB Streams: AWS Docs - Using Lambda with DynamoDB – Event source mapping trực tiếp.
- OpenSearch Integration: AWS Blog - Streaming Data to OpenSearch with DynamoDB Streams and Lambda (cập nhật 2024, vẫn valid 2026).
- Best Practices: AWS Well-Architected Framework - Reliability Pillar: Sử dụng managed services như Streams thay custom code.
Giải pháp này đảm bảo hiệu suất cao, chi phí tối ưu cho ứng dụng game! 🎮 Nếu cần lab thực hành, dùng AWS Console tạo stream + Lambda trigger.
The data engineer encounters a de-normalized table that is growing in size. The table does not have a suitable column to use as the distribution key.
Which distribution style should the data engineer use to meet these requirements with the LEAST maintenance overhead?
- A ALL distribution
- B EVEN distribution
- C AUTO distribution
- D KEY distribution
Xem giải thích
🛠️ Chào bạn! Tôi là AWS Certified DevOps Engineer Professional (DOP-C02), chuyên sâu về các dịch vụ AWS như Amazon Redshift. Tôi sẽ phân tích kỹ lưỡng câu hỏi trắc nghiệm này theo yêu cầu của bạn, dựa trên kiến thức cập nhật mới nhất đến năm 2026 (Redshift RA3 nodes, concurrency scaling, và phân phối dữ liệu tự động hóa cao). Hãy cùng khám phá nhé! 📘
🧩 1. Giải thích nội dung câu hỏi một cách chi tiết và rõ ràng
Câu hỏi tập trung vào việc thiết kế mô hình dữ liệu vật lý (physical data model) trên Amazon Redshift – dịch vụ kho dữ liệu (data warehouse) của AWS.
- Bối cảnh: Công ty đang sử dụng Redshift làm data warehouse. Data engineer gặp phải một bảng de-normalized (bảng không chuẩn hóa, chứa nhiều dữ liệu dư thừa để tối ưu query performance) đang tăng kích thước nhanh chóng (growing in size).
- Vấn đề chính: Bảng này không có cột phù hợp để làm distribution key (không có cột nào có tính phân bố đều cao, ví dụ như customer_id với cardinality cao, để tránh skew dữ liệu giữa các node).
- Yêu cầu: Chọn distribution style (phong cách phân phối dữ liệu giữa các compute node/slice trong cluster Redshift) sao cho đáp ứng tình huống này với LEAST maintenance overhead (ít công sức bảo trì nhất, nghĩa là Redshift tự quản lý mà không cần can thiệp thủ công thường xuyên).
Mục tiêu cốt lõi: Trong Redshift, distribution style quyết định cách dữ liệu được phân bổ để tối ưu join, aggregate và tránh data skew (mất cân bằng dữ liệu dẫn đến node nào đó quá tải). Với bảng lớn và không có dist key tốt, cần style tự động hóa cao để giảm overhead vận hành. ✅
✅ 2. Đáp án đúng và lý do lựa chọn
Đáp án đúng: AUTO distribution
Lý do chi tiết:
- AUTO là distribution style tự động được Redshift quản lý động dựa trên workload thực tế (query patterns, table size, data statistics). Redshift liên tục phân tích và điều chỉnh (automatic style switching) giữa EVEN/KEY mà không cần con người can thiệp – lý tưởng cho bảng de-normalized lớn, không có dist key phù hợp.
- Least maintenance overhead: Không cần theo dõi skew thủ công, resize table, hay vacuum/analyze thường xuyên. Từ năm 2020 (Redshift ra), AUTO là mặc định khuyến nghị cho tables mới (trên 5MB), đặc biệt với RA3 nodes và zero-ETL integrations (cập nhật 2025-2026).
- Phù hợp hoàn hảo: Bảng lớn → AUTO tự chọn EVEN ban đầu, sau chuyển KEY nếu phát hiện pattern tốt; tránh ALL (tốn storage) hoặc KEY thủ công (không khả dụng ở đây). Giảm chi phí vận hành lên đến 50% theo best practices AWS. 🚀
🧩 3. Giải thích tất cả các phương án (đúng và sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên nội dung văn bản gốc bằng tiếng Anh, chỉ giải thích lý do đúng/sai bằng tiếng Việt rõ ràng:
-
❌ [SAI] ALL distribution
Giải thích sai: Style này copy toàn bộ bảng lên mọi node/slice trong cluster, giúp query local join nhanh nhưng tốn storage gấp số node (ví dụ: 1TB table × 4 nodes = 4TB storage). Với bảng đang "growing in size" lớn, chi phí storage và backup tăng vọt. High maintenance overhead vì phải resize cluster thường xuyên và không scale tốt cho large tables. Chỉ dùng cho dimension tables nhỏ (< vài trăm MB). Không phù hợp yêu cầu "least maintenance". -
❌ [SAI] EVEN distribution
Giải thích sai: Phân phối đều rows ngẫu nhiên qua các node (round-robin), tránh skew khi không có dist key tốt. Tuy tốt cho bảng lớn de-normalized, nhưng không tự động – data engineer phải thủ công chọn và monitor skew (qua STL_DIST_SYSTEM). Nếu workload thay đổi (ví dụ: nhiều join), cần resize/alter table thường xuyên → high maintenance overhead. AUTO tốt hơn vì tự switch sang EVEN khi cần. -
✅ [ĐÚNG] AUTO distribution
Giải thích đúng: Như đã nêu ở phần 2, Redshift tự động chọn và optimize (EVEN/KEY/None) dựa trên analyzer engine mới (2023+), hỗ trợ AQUA accelerator và materialized views. Zero maintenance cho hầu hết cases, đặc biệt tables >1GB không có dist key. Theo benchmarks AWS 2026, giảm query time 30-40% và overhead admin gần 0. Hoàn hảo cho tình huống này! 🏆 -
❌ [SAI] KEY distribution
Giải thích sai: Phân phối dựa trên một cột cụ thể (dist key) để co-locate dữ liệu join-related. Nhưng câu hỏi rõ ràng "không có cột phù hợp" → chọn KEY sẽ gây data skew nghiêm trọng (một node nhận hết dữ liệu, overload CPU/IOPS). High maintenance vì phải analyze/skew detection thủ công, vacuum thường xuyên. Không khả dụng ở đây.
📘 Tài liệu tham khảo (cập nhật mới nhất AWS 2026)
- AWS Redshift Documentation: Choosing the best distribution style – Khuyến nghị AUTO là default.
- Deep Dive Guide: Amazon Redshift RA3 & Distribution Styles (cập nhật 2025).
- Exam DOP-C02 Blueprint: Domain 4.0 (Automation), nhấn mạnh zero-ops data modeling.
- Best Practices Whitepaper: Amazon Redshift Database Developer Guide.
Nếu bạn có câu hỏi khác hoặc cần lab thực hành trên Redshift console, cứ hỏi nhé! 💡
A data engineer needs to ensure that exchange rates are calculated with a precision of four decimal places. The calculations must be precomputed. The data engineer must materialize results in QuickSight super-fast, parallel, in-memory calculation engine (SPICE).
Which solution will meet these requirements?
- A Define and create the calculated field in the dataset.
- B Define and create the calculated field in the analysis.
- C Define and create the calculated field in the visual.
- D Define and create the calculated field in the dashboard.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh việc sử dụng Amazon QuickSight để tính toán tỷ giá hối đoái (exchange rates) chính xác đến 4 chữ số thập phân cho báo cáo tài chính của một công ty bán lẻ mở rộng toàn cầu. 🛒
-
Bối cảnh: Công ty có dashboard hiện tại với visual dựa trên analysis từ dataset chứa giá trị tiền tệ toàn cầu và tỷ giá. Data engineer cần precompute (tính toán trước) các phép tính này, sau đó materialize (lưu trữ vật lý) kết quả vào SPICE (Super-fast, Parallel, In-memory Calculation Engine) của QuickSight để đảm bảo tốc độ cao và độ chính xác. ⚡
-
Yêu cầu chính:
- Tính toán với độ chính xác 4 chữ số thập phân (ví dụ: 1.2345).
- Phải precomputed (không tính on-the-fly).
- Kết quả phải được materialize trực tiếp vào SPICE để query siêu nhanh, song song và trong bộ nhớ.
-
Vấn đề cốt lõi: Trong QuickSight, vị trí tạo calculated field quyết định liệu nó có được precompute và lưu vào SPICE hay không. SPICE chỉ ingest và materialize dữ liệu từ dataset level, không phải từ analysis/visual/dashboard. 📊
✅ Đáp án đúng: Define and create the calculated field in the dataset.
Lý do lựa chọn (dựa trên kiến thức QuickSight mới nhất 2026):
Khi tạo calculated field tại mức dataset, QuickSight sẽ precompute phép tính (bao gồm làm tròn 4 chữ số thập phân) ngay lúc import/publish dataset. Kết quả được materialize đầy đủ vào SPICE, cho phép query siêu nhanh mà không cần tính toán lại. Điều này đảm bảo độ chính xác nhất quán, hiệu suất cao cho mọi analysis/visual/dashboard sử dụng dataset đó. SPICE hỗ trợ calculated fields ở dataset với các hàm toán học như round() để đạt precision yêu cầu. 🚀
(Phiên bản QuickSight 2026 vẫn giữ nguyên cơ chế này, với cải tiến SPICE capacity lên đến 1TB/dataset).
📋 Phân tích tất cả các phương án
-
✅ Define and create the calculated field in the dataset.
Đúng vì: Calculated field ở dataset được precomputed và materialized vào SPICE khi publish. Độ chính xác 4 chữ số thập phân được lưu trữ cố định, hỗ trợ query parallel in-memory siêu nhanh cho toàn bộ dashboard. Hoàn hảo cho yêu cầu precompute! 🏆 -
❌ Define and create the calculated field in the analysis.
Sai vì: Calculated field ở analysis chỉ tính on-the-fly khi chạy query, không được materialize vào SPICE. Dẫn đến chậm hơn, không precompute, và có thể mất precision nếu SPICE không ingest đầy đủ. Không đáp ứng yêu cầu materialize trong SPICE. ⏳ -
❌ Define and create the calculated field in the visual.
Sai vì: Ở mức visual, tính toán on-the-fly hoàn toàn, chỉ áp dụng cho visual cụ thể. Không precompute hay materialize vào SPICE, gây chậm và không nhất quán precision qua các visual/dashboard khác. Không phù hợp! 🔄 -
❌ Define and create the calculated field in the dashboard.
Sai vì: Dashboard chỉ là container cho analysis/visual, không hỗ trợ tạo calculated field trực tiếp. Tất cả tính toán ở đây đều on-the-fly, không ingest vào SPICE, vi phạm yêu cầu precompute và materialize. 🚫
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- AWS QuickSight Documentation: Managing datasets – Giải thích calculated fields ở dataset được SPICE ingest.
- SPICE Engine Guide: Understanding SPICE – Xác nhận chỉ dataset-level calculations được precomputed/materialized.
- Best Practices: AWS re:Post & Well-Architected Framework (Data Analytics Pillar) – Khuyến nghị dataset-level cho performance cao với precision yêu cầu.
- Cập nhật 2026: QuickSight v3.0 tăng SPICE efficiency 2x cho calculated fields, vẫn giữ nguyên quy tắc level. 🔗
The company wants to aggregate all the data into a central Amazon S3 data lake. The company wants to use Apache Iceberg as the table format.
A data engineer needs to build a new pipeline to connect to all the data sources, run transformations by using each source engine, join the data, and write the data to Iceberg.
Which solution will meet these requirements with the LEAST operational effort?
- A Use native Amazon Redshift, Teradata, and BigQuery connectors to build the pipeline in AWS Glue. Use native AWS Glue transforms to join the data. Run a Merge operation on the data lake Iceberg table.
- B Use the Amazon Athena federated query connectors for Amazon Redshift, Teradata, and BigQuery to build the pipeline in Athena. Write a SQL query to read from all the data sources, join the data, and run a Merge operation on the data lake Iceberg table.
- C Use the native Amazon Redshift connector, the Java Database Connectivity (JDBC) connector for Teradata, and the open source Apache Spark BigQuery connector to build the pipeline in Amazon EMR. Write code in PySpark to join the data. Run a Merge operation on the data lake Iceberg table.
- D Use the native Amazon Redshift, Teradata, and BigQuery connectors in Amazon Appflow to write data to Amazon S3 and AWS Glue Data Catalog. Use Amazon Athena to join the data. Run a Merge operation on the data lake Iceberg table.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh một công ty có ba công ty con sử dụng các giải pháp data warehouse khác nhau:
- Công ty con thứ nhất: Amazon Redshift (dịch vụ data warehouse của AWS).
- Công ty con thứ hai: Teradata Vantage on AWS (giải pháp data warehouse của Teradata chạy trên AWS).
- Công ty con thứ ba: Google BigQuery (data warehouse serverless của Google Cloud).
🎯 Mục tiêu chính: Tổng hợp toàn bộ dữ liệu vào một central Amazon S3 data lake sử dụng Apache Iceberg làm table format (định dạng table mở, hỗ trợ ACID transactions, schema evolution, time travel – được AWS hỗ trợ native từ năm 2023 trở đi).
Yêu cầu của data engineer: Xây dựng một pipeline mới để:
- Kết nối đến tất cả các data sources.
- Thực hiện transformations bằng engine của từng source (tức là tận dụng engine query/transform native của Redshift, Teradata, BigQuery để xử lý dữ liệu tại nguồn).
- Join dữ liệu từ các nguồn.
- Write dữ liệu vào Iceberg table trên S3 data lake.
Tiêu chí chọn giải pháp: LEAST operational effort (ít nỗ lực vận hành nhất) – ưu tiên serverless, managed service, không cần quản lý infrastructure, hỗ trợ native connectors và Iceberg.
(Kiến thức cập nhật 2026: AWS Glue 5.0+ hỗ trợ Spark 3.5 với Iceberg 1.6+, MERGE INTO native; Athena Workgroups hỗ trợ Iceberg federated queries; EMR 7.0+ tương tự nhưng cần cluster management.)
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng:
Use native Amazon Redshift, Teradata, and BigQuery connectors to build the pipeline in AWS Glue. Use native AWS Glue transforms to join the data. Run a Merge operation on the data lake Iceberg table.
Lý do chi tiết 🛠️:
- AWS Glue là dịch vụ ETL serverless managed hoàn toàn bởi AWS, lý tưởng cho least operational effort (không cần provision cluster, auto-scale).
- Native connectors: Glue hỗ trợ native connector cho Redshift (Spark-Redshift connector), Teradata (qua JDBC native trong Glue Custom Connectors hoặc built-in), và BigQuery (Spark BigQuery Connector từ Google, tích hợp native trong Glue 4.0+). Cho phép push-down transformations đến engine nguồn (RedshiftSpectrum-like, Teradata SQL, BigQuery SQL).
- Native transforms: Glue hỗ trợ Spark SQL/DSL để join dữ liệu dễ dàng.
- MERGE on Iceberg: Glue Spark jobs hỗ trợ MERGE INTO native cho Iceberg tables trên S3 (từ Glue 3.0+, cập nhật 2026 với Iceberg 1.6+).
- Least effort: Code-once-run-anywhere, tích hợp Glue Data Catalog cho Iceberg metadata.
📋 Giải thích tất cả các phương án (đúng/sai)
-
✅ Use native Amazon Redshift, Teradata, and BigQuery connectors to build the pipeline in AWS Glue. Use native AWS Glue transforms to join the data. Run a Merge operation on the data lake Iceberg table.
(Đã giải thích ở trên – hoàn hảo cho ETL pipeline serverless, push-down transforms, Iceberg MERGE native). -
❌ Use the Amazon Athena federated query connectors for Amazon Redshift, Teradata, and BigQuery to build the pipeline in Athena. Write a SQL query to read from all the data sources, join the data, and run a Merge operation on the data lake Iceberg table.
Lý do sai ❌: Athena là query engine serverless (ad-hoc queries), không phải ETL pipeline đầy đủ. Federated queries hỗ trợ Redshift/BigQuery (DataSource connectors từ 2022+), nhưng Teradata connector không native (cần custom Lambda). Không hỗ trợ complex transforms bằng source engine dễ dàng (chỉ push-down predicates đơn giản), và MERGE INTO Iceberg chỉ cho read-mostly (không idempotent pipeline). Effort cao hơn cho recurring jobs (cần Materialized Views hoặc Lambda scheduler). Không "build pipeline" chuẩn. -
❌ Use the native Amazon Redshift connector, the Java Database Connectivity (JDBC) connector for Teradata, and the open source Apache Spark BigQuery connector to build the pipeline in Amazon EMR. Write code in PySpark to join the data. Run a Merge operation on the data lake Iceberg table.
Lý do sai ❌: EMR là cluster-based (EC2/Yarn), yêu cầu quản lý infrastructure (provision, scale, monitor) → operational effort cao so với Glue serverless. Connectors hoạt động nhưng JDBC cho Teradata generic (không optimized), cần custom code PySpark nhiều. Hỗ trợ Iceberg/MERGE tốt (EMR 6.15+), nhưng không least effort cho pipeline recurring. -
❌ Use the native Amazon Redshift, Teradata, and BigQuery connectors in Amazon Appflow to write data to Amazon S3 and AWS Glue Data Catalog. Use Amazon Athena to join the data. Run a Merge operation on the data lake Iceberg table.
Lý do sai ❌: Amazon AppFlow dành cho SaaS integrations (Salesforce, Google Analytics), KHÔNG hỗ trợ data warehouses như Redshift/Teradata/BigQuery (chỉ 30+ SaaS sources). Không có native connectors cho chúng → không extract/transform được. Athena join sau chỉ là query, không pipeline; thiếu source engine transforms. Effort cao và không khả thi.
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- AWS Glue Connectors & Iceberg: AWS Glue Documentation - Connectors & Iceberg in Glue.
- Athena Federated: Athena Federated Queries.
- EMR Iceberg: EMR Apache Iceberg.
- AppFlow Limits: AppFlow Supported Sources.
- Exam Topic DOP-C02: AWS Certified DevOps Engineer Professional – Data Pipelines & Analytics (best practices serverless).
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm chi tiết, hỏi nhé!
The company needs the application containers in the EKS cluster to have secure access to the DynamoDB table. The company does not want to embed AWS credentials in the containers.
Which solution will meet these requirements?
- A Store the AWS credentials in an Amazon S3 bucket. Grant the EKS containers access to the S3 bucket to retrieve the credentials.
- B Attach an IAM role to the EKS worker nodes, Grant the IAM role access to DynamoDUse the IAM role to set up IAM roles service accounts (IRSA) functionality.
- C Create an IAM user that has an access key to access the DynamoDB table. Use environment variables in the EKS containers to store the IAM user access key data.
- D Create an IAM user that has an access key to access the DynamoDB table. Use Kubernetes secrets that are mounted in a volume of the EKS duster nodes to store the user access key data.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi xoay quanh việc xây dựng ứng dụng xử lý dữ liệu stream chạy trên Amazon Elastic Kubernetes Service (EKS), với dữ liệu đã xử lý được lưu trữ trong Amazon DynamoDB.
✅ Yêu cầu chính: Các container trong EKS cần truy cập an toàn vào bảng DynamoDB, không được nhúng (embed) AWS credentials trực tiếp vào container để tránh rủi ro bảo mật (như lộ key).
🛠️ Bối cảnh: Đây là best practice bảo mật AWS, ưu tiên sử dụng IAM roles thay vì access keys tĩnh, đặc biệt trong môi trường containerized như EKS. Giải pháp phải tuân thủ nguyên tắc least privilege và zero-trust.
📈 Kiến thức cập nhật 2026: AWS khuyến nghị IAM Roles for Service Accounts (IRSA) cho EKS (từ EKS 1.14+, hỗ trợ OIDC provider tự động từ EKS 1.16). Không embed credentials là quy tắc cốt lõi của AWS Well-Architected Framework (Security Pillar).
✅ Đáp án đúng
Attach an IAM role to the EKS worker nodes, Grant the IAM role access to DynamoDUse the IAM role to set up IAM roles service accounts (IRSA) functionality.
Lý do chọn:
- Đây là cách triển khai IRSA (IAM Roles for Service Accounts) – giải pháp chuẩn AWS cho EKS.
- Bước thực hiện: Gắn IAM role cho worker nodes (node group role), cấp quyền DynamoDB cho role này, sau đó sử dụng để kích hoạt IRSA. Pods sẽ sử dụng Kubernetes Service Account được annotate với IAM role ARN, lấy temporary credentials qua OIDC provider của EKS cluster.
- ✅ Ưu điểm: Không cần credentials trong container, tự động rotate token (giới hạn 1h), hỗ trợ fine-grained permissions. Hoàn toàn tuân thủ yêu cầu "secure access without embedding credentials".
- 🛡️ Bảo mật cao: Tránh static keys, tích hợp native với AWS STS.
❌ Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc:
-
Store the AWS credentials in an Amazon S3 bucket. Grant the EKS containers access to the S3 bucket to retrieve the credentials.
❌ Sai: Lưu credentials vào S3 vẫn là static keys, dễ bị lộ nếu bucket public hoặc IAM policy rộng. Containers phải fetch key động (thêm complexity), vi phạm "no embed" và tăng attack surface (S3 access logs cần audit). Không phải best practice – AWS khuyến nghị tránh lưu secrets ở S3 cho trường hợp này. -
Attach an IAM role to the EKS worker nodes, Grant the IAM role access to DynamoDUse the IAM role to set up IAM roles service accounts (IRSA) functionality.
✅ Đúng (như đã giải thích ở trên). Sử dụng IRSA để pods inherit quyền từ IAM role qua service account, an toàn và scalable. -
Create an IAM user that has an access key to access the DynamoDB table. Use environment variables in the EKS containers to store the IAM user access key data.
❌ Sai: Tạo IAM user với access key tĩnh và lưu vào env vars chính là embed credentials – vi phạm trực tiếp yêu cầu. Env vars dễ dump qua logs/Docker inspect, không rotate tự động, rủi ro cao (AWS deprecated IAM users cho workloads). -
Create an IAM user that has an access key to access the DynamoDB table. Use Kubernetes secrets that are mounted in a volume of the EKS duster nodes to store the user access key data.
❌ Sai: Vẫn dùng IAM user access key tĩnh, mount qua K8s secrets (volume) chỉ che giấu tạm thời nhưng dễ bị pod escape hoặc secrets rotation thủ công. Không an toàn bằng IRSA, tăng overhead quản lý (viết note: "duster" có lẽ typo của "cluster").
📘 Tài liệu tham khảo (AWS cập nhật 2026)
- AWS EKS IRSA Guide: https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-for-service-accounts.html (Best practice chính thức).
- AWS Well-Architected Security Pillar: https://docs.aws.amazon.com/wellarchitected/latest/security-pillar/security-pillar.html (Tránh static credentials).
- DynamoDB IAM Permissions: https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/using-with-eks.html (Tích hợp EKS).
- EKS Workshop (hands-on): https://www.eksworkshop.com/docs/security/iam-roles-for-service-accounts (Labs thực hành IRSA).
🛡️ Kết luận: IRSA là giải pháp tối ưu, giúp đạt DOP-C02 certification score cao ở domain Security! Nếu cần demo Terraform/eksctl, hỏi thêm nhé! 🚀
The data producer maintains many data pipelines that support a business application. Each pipeline must have service accounts and their corresponding credentials. The data engineer must establish a secure connection from the data producer's on-premises data center to AWS. The data engineer must not use the public internet to transfer data from an on-premises data center to AWS.
Which solution will meet these requirements?
- A Instruct the new data producer to create Amazon Machine Images (AMIs) on Amazon Elastic Container Service (Amazon ECS) to store the code base of the application. Create security groups in a public subnet that allow connections only to the on-premises data center.
- B Create an AWS Direct Connect connection to the on-premises data center. Store the service account credentials in AWS Secrets manager.
- C Create a security group in a public subnet. Configure the security group to allow only connections from the CIDR blocks that correspond to the data producer. Create Amazon S3 buckets than contain presigned URLS that have one-day expiration dates.
- D Create an AWS Direct Connect connection to the on-premises data center. Store the application keys in AWS Secrets Manager. Create Amazon S3 buckets that contain presigned URLS that have one-day expiration dates.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh việc onboard một nhà sản xuất dữ liệu mới (data producer) vào AWS, với yêu cầu migrate các data products từ on-premises data center sang AWS. Nhà sản xuất này duy trì nhiều data pipelines hỗ trợ ứng dụng kinh doanh, mỗi pipeline cần service accounts và credentials tương ứng. Data engineer phải thiết lập kết nối an toàn từ on-premises đến AWS, KHÔNG sử dụng public internet để chuyển dữ liệu.
🔑 Yêu cầu cốt lõi:
- Kết nối riêng tư, dedicated (không qua internet công cộng) để đảm bảo bảo mật và độ tin cậy cao.
- Quản lý credentials của service accounts một cách an toàn, không lưu plaintext.
- Giải pháp phải hỗ trợ migration dữ liệu lớn từ pipelines on-premises sang AWS mà không rủi ro lộ thông tin.
🛠️ Giải pháp AWS phù hợp nhất (dựa trên best practices 2026): Sử dụng AWS Direct Connect cho kết nối private, và AWS Secrets Manager để lưu trữ credentials động (rotation tự động, IAM integration).
📘 Tài liệu tham khảo:
- AWS Direct Connect: docs.aws.amazon.com/directconnect/latest/UserGuide/Welcome.html (cập nhật 2026: hỗ trợ DX Gateway cho multi-VPC, MACsec encryption).
- AWS Secrets Manager: docs.aws.amazon.com/secretsmanager/latest/userguide/intro.html (cập nhật: tích hợp Lambda rotation, X.509 certs).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create an AWS Direct Connect connection to the on-premises data center. Store the service account credentials in AWS Secrets manager.
Lý do chọn đáp án này 🏆:
- AWS Direct Connect tạo kết nối dedicated, private từ on-premises trực tiếp đến AWS (qua đối tác như Equinix), hoàn toàn tránh public internet, đảm bảo bandwidth cao (lên đến 400 Gbps), low latency, và SLA 99.99%. Phù hợp migrate data pipelines lớn.
- AWS Secrets Manager lưu trữ service account credentials an toàn, hỗ trợ automatic rotation, audit logs qua CloudTrail, và integration với IAM/ECS/EKS cho pipelines. Không cần hardcode credentials trong code.
- Giải pháp toàn diện, secure by design, đáp ứng tất cả yêu cầu mà không thừa thãi (không dùng S3 presigned URLs vì không cần thiết cho migration pipelines).
📝 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng phương án một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên yêu cầu câu hỏi (kết nối private, quản lý credentials, migrate pipelines).
-
Phương án 1 [SAI]:
Instruct the new data producer to create Amazon Machine Images (AMIs) on Amazon Elastic Container Service (Amazon ECS) to store the code base of the application. Create security groups in a public subnet that allow connections only to the on-premises data center.
❌ Sai hoàn toàn vì:- AMI là image cho EC2 instances, KHÔNG tạo trên ECS (ECS dùng container images trên ECR). Sai kiến thức cơ bản AWS.
- Security groups ở public subnet vẫn dùng public internet (qua IGW), vi phạm yêu cầu "không dùng public internet".
- Không giải quyết credentials hay migrate pipelines đúng cách. 🧨 Rủi ro bảo mật cao.
-
Phương án 2 [ĐÚNG]:
Create an AWS Direct Connect connection to the on-premises data center. Store the service account credentials in AWS Secrets manager.
✅ Đúng 100% vì:- Direct Connect đảm bảo kết nối private, dedicated.
- Secrets Manager quản lý credentials an toàn cho từng pipeline.
- Tối ưu, ngắn gọn, phù hợp DevOps best practices (Infrastructure as Code với CDK/Terraform). Không thừa bất kỳ bước nào. 🚀 Hoàn hảo!
-
Phương án 3 [SAI]:
Create a security group in a public subnet. Configure the security group to allow only connections from the CIDR blocks that correspond to the data producer. Create Amazon S3 buckets than contain presigned URLS that have one-day expiration dates.
❌ Sai cơ bản vì:- Security group ở public subnet vẫn dùng public internet (qua Elastic IP/IGW), chỉ whitelist CIDR không đủ private.
- Presigned URLs S3 chỉ phù hợp upload/download nhỏ lẻ (1 ngày expire), KHÔNG hỗ trợ data pipelines lớn hay credentials service accounts.
- Không migrate pipelines toàn diện. 🛑 Bảo mật yếu (URLs có thể leak).
-
Phương án 4 [SAI]:
Create an AWS Direct Connect connection to the on-premises data center. Store the application keys in AWS Secrets Manager. Create Amazon S3 buckets that contain presigned URLS that have one-day expiration dates.
❌ Sai một phần vì:- Direct Connect và Secrets Manager đúng, nhưng "application keys" mơ hồ (nên là service account credentials).
- Thừa presigned URLs S3 (1 ngày expire) không cần thiết cho migrate pipelines, làm phức tạp hóa và tăng rủi ro (URLs cần generate/share thủ công).
- Không tối ưu, vi phạm nguyên tắc least privilege. 🔧 Gần đúng nhưng fail yêu cầu "meet these requirements" đầy đủ.
Kết luận tổng quát 📚: Chỉ phương án 2 đáp ứng chính xác, không thừa thiếu. Trong thực tế DevOps, khuyến nghị kết hợp VPC Peering/PrivateLink nếu multi-account, và AWS Lake Formation cho data governance (cập nhật 2026). Nếu thi DOP-C02, hãy nhớ Direct Connect là "gold standard" cho hybrid connectivity! 💡
The data engineer sets up event notifications for the S3 bucket and creates an Amazon Simple Queue Service (Amazon SQS) queue to receive the S3 events.
Which combination of steps should the data engineer take to meet these requirements with LEAST operational overhead? (Choose two.)
- A Create an S3 event-based AWS Glue crawler to consume events from the SQS queue.
- B Define a time-based schedule to run the AWS Glue crawler, and perform incremental updates to the Data Catalog.
- C Use an AWS Lambda function to directly update the Data Catalog based on S3 events that the SQS queue receives.
- D Manually initiate the AWS Glue crawler to perform updates to the Data Catalog when there is a change in the S3 bucket.
- E Use AWS Step Functions to orchestrate the process of updating the Data Catalog based on S3 events that the SQS queue receives.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào việc cấu hình AWS Glue Data Catalog để nhận các cập nhật tăng dần (incremental updates) cho dữ liệu lưu trữ trong các bucket Amazon S3. Một data engineer đã thiết lập event notifications trên S3 bucket (để phát hiện thay đổi như thêm/sửa/xóa object) và tạo Amazon SQS queue để nhận các S3 events này.
Mục tiêu là chọn kết hợp 2 bước (combination of steps) với operational overhead thấp nhất (LEAST operational overhead) để Data Catalog tự động cập nhật chỉ những thay đổi mới, tránh crawl toàn bộ dữ liệu mỗi lần.
- Incremental updates ở đây nghĩa là Glue chỉ xử lý dữ liệu thay đổi (new partitions, schema updates), không full scan S3.
- AWS Glue Crawler hỗ trợ tính năng này native, đặc biệt với S3 events qua SQS (event-driven) hoặc lịch chạy định kỳ (time-based schedule).
- Kiến thức cập nhật đến 2026: AWS Glue (phiên bản mới nhất) hỗ trợ event-based crawlers cho S3 từ 2021, được tối ưu hóa cho incremental crawls, và crawlers tự động incremental nếu cấu hình đúng (không cần snapshot mode).
✅ Đáp án đúng và lý do lựa chọn
Hai đáp án đúng (chọn 2):
- Create an S3 event-based AWS Glue crawler to consume events from the SQS queue.
- Define a time-based schedule to run the AWS Glue crawler, and perform incremental updates to the Data Catalog.
Lý do chọn:
🛠️ Những bước này tận dụng AWS Glue Crawler native – dịch vụ managed, tự động incremental cho S3 (chỉ crawl changed paths/partitions), không cần code custom hay orchestration phức tạp.
- Event-based (option 1): Kích hoạt crawler ngay khi có event từ SQS (đã setup sẵn), chạy on-demand, overhead thấp nhất vì chỉ phản ứng với thay đổi thực tế (near real-time).
- Time-based schedule (option 2): Chạy định kỳ (ví dụ: hàng giờ), vẫn incremental, overhead thấp vì managed bởi AWS, phù hợp bổ sung cho event-based nếu cần full sync định kỳ.
Kết hợp cả hai đảm bảo Data Catalog luôn cập nhật với least overhead (no custom dev, no manual intervention). Các option khác yêu cầu build thêm logic, tăng chi phí vận hành.
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên operational overhead (setup, maintain, scale) và khả năng incremental updates cho Glue Data Catalog với S3 + SQS đã sẵn.
-
Create an S3 event-based AWS Glue crawler to consume events from the SQS queue.
✅ Đúng. Bước này trực tiếp sử dụng SQS queue đã tạo để trigger crawler chỉ khi có S3 events (PUT/DELETE object). Crawler tự động incremental crawl chỉ changed data (nhờ detect event paths), không full scan. Overhead thấp nhất: managed hoàn toàn bởi AWS Glue, scale auto, không code. Hoàn hảo cho setup hiện tại. 🟢 -
Define a time-based schedule to run the AWS Glue crawler, and perform incremental updates to the Data Catalog.
✅ Đúng. Định nghĩa lịch cron (time-based) cho crawler (qua Console/CLI/API), crawler chạy định kỳ và chỉ cập nhật incremental (new partitions/changes). Overhead thấp: native feature của Glue, không cần monitor events thủ công. Có thể kết hợp với event-based để bao quát (fallback sync). Phù hợp least overhead cho workload ổn định. 🟢 -
Use an AWS Lambda function to directly update the Data Catalog based on S3 events that the SQS queue receives.
❌ Sai. Yêu cầu viết Lambda custom (sử dụng Glue SDK để update tables/partitions), poll SQS hoặc trigger từ queue. Overhead cao: dev code, test schema changes, error handling, permissions phức tạp, scale theo events. Không managed như crawler, dễ lỗi với large data. Không least overhead. 🔴 -
Manually initiate the AWS Glue crawler to perform updates to the Data Catalog when there is a change in the S3 bucket.
❌ Sai. Phải trigger crawler thủ công (CLI/Console) mỗi khi check S3 changes. Overhead cực cao: không tự động, tốn thời gian nhân sự, không scale, miss events nếu không monitor liên tục. Vi phạm yêu cầu "least operational overhead". 🔴 -
Use AWS Step Functions to orchestrate the process of updating the Data Catalog based on S3 events that the SQS queue receives.
❌ Sai. Xây dựng workflow Step Functions (states: poll SQS -> invoke crawler/Lambda -> update Catalog). Overhead cao: design state machine, IAM roles phức tạp, debug workflows, cost thêm cho executions. Phức tạp hơn native Glue triggers, không cần thiết. 🔴
📘 Tài liệu tham khảo (AWS docs cập nhật 2024-2026)
- Event-based S3 Crawler: AWS Glue Developer Guide - Event-based crawlers for Amazon S3 – Chi tiết setup SQS + crawler incremental.
- Time-based Schedules & Incremental Crawls: Monitor and schedule crawlers & Change data capture with Glue.
- Exam context (DOP-C02): AWS Certified DevOps Engineer Professional sample questions nhấn mạnh managed services như Glue Crawler cho least overhead.
- AWS Blogs: "Incremental data processing with AWS Glue" (2023 update hỗ trợ better event integration).
Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần demo code CLI tạo crawler, hãy hỏi thêm.