Ngân hàng đề — AWS Certified Data Engineer Associate

Tìm thấy 867 câu.

Câu 741
A data engineer has implemented data quality rules in 1,000 AWS Glue Data Catalog tables. Because of a recent change in business requirements, the data engineer must edit the data quality rules.

How should the data engineer meet this requirement with the LEAST operational overhead?
  1. A Create a pipeline in AWS Glue ETL to edit the rules for each of the 1,000 Data Catalog tables. Use an AWS Lambda function to call the corresponding AWS Glue job for each Data Catalog table.
  2. B Create an AWS Lambda function that makes an API call to AWS Glue Data Quality to make the edits.
  3. C Create an Amazon EMR cluster. Run a pipeline on Amazon EMR that edits the rules for each Data Catalog table. Use an AWS Lambda function to run the EMR pipeline.
  4. D Use the AWS Management Console to edit the rules within the Data Catalog.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc cập nhật data quality rules cho 1.000 bảng trong AWS Glue Data Catalog sau khi có thay đổi yêu cầu kinh doanh. Data engineer cần chọn phương pháp ít tốn kém vận hành nhất (LEAST operational overhead), nghĩa là ưu tiên giải pháp tự động hóa cao, serverless, scale dễ dàng mà không cần quản lý infrastructure phức tạp.

AWS Glue Data Quality (DQ) là tính năng cho phép định nghĩa và chạy rules kiểm tra chất lượng dữ liệu trực tiếp trên Catalog tables (từ năm 2021 và cập nhật liên tục đến 2026 với hỗ trợ ML recommendations và API enhancements). Việc edit rules cho hàng nghìn tables đòi hỏi API calls batch hoặc loop để tránh manual work, vì console không scale cho quy mô lớn.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create an AWS Lambda function that makes an API call to AWS Glue Data Quality to make the edits.

Lý do chọn đáp án này (bằng tiếng Việt chi tiết):
🛠️ Phương pháp này serverless hoàn toàn, Lambda tự động scale và chỉ tính phí theo execution time (rẻ cho 1.000 calls). AWS Glue DQ hỗ trợ API trực tiếp như UpdateTable hoặc BatchCreatePartition kết hợp với DQ rules (rules lưu trong table properties). Lambda có thể loop qua GetTables API để lấy list tables, rồi gọi UpdateTable với DQ rules mới – chỉ vài phút setup, không cần quản lý cluster hay job. Đây là least overhead vì không provision resources, tự động retry, và tích hợp IAM policies đơn giản. Phù hợp best practice DevOps cho automation trên AWS (cập nhật 2026: Glue DQ APIs hỗ trợ bulk updates qua SDK).

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn (giữ nguyên văn bản gốc), đánh dấu ✅ đúng hoặc ❌ sai, kèm lý do bằng tiếng Việt:

  • ❌ Create a pipeline in AWS Glue ETL to edit the rules for each of the 1,000 Data Catalog tables. Use an AWS Lambda function to call the corresponding AWS Glue job for each Data Catalog table.
    🧩 Sai vì overhead cao: Glue ETL jobs dùng để transform data, không phải edit metadata như DQ rules (phải custom script gọi Glue APIs bên trong job – phức tạp). Với 1.000 tables, tạo 1.000 jobs riêng lẻ gây quản lý khó, chi phí DPU cao, Lambda trigger chỉ tăng thêm layer không cần thiết. Không phải cách native cho metadata updates.

  • ✅ Create an AWS Lambda function that makes an API call to AWS Glue Data Quality to make the edits.
    🛠️ Đúng – lý do chi tiết ở phần trên: Serverless, trực tiếp dùng Glue DQ APIs (như boto3 Glue client), scale cho 1.000 tables dễ dàng với concurrency. Overhead thấp nhất: code ngắn gọn, deploy 1 function.

  • ❌ Create an Amazon EMR cluster. Run a pipeline on Amazon EMR that edits the rules for each Data Catalog table. Use an AWS Lambda function to run the EMR pipeline.
    🧩 Sai vì overhead cực cao: EMR là cho big data processing (Spark/Hive), không phù hợp edit Catalog metadata (phải dùng EMR steps gọi Glue SDK – thừa thãi). Provision cluster tốn thời gian/boot time, chi phí EC2 cao, quản lý scaling phức tạp cho task đơn giản. Lambda trigger EMR chỉ làm tình hình tệ hơn.

  • ❌ Use the AWS Management Console to edit the rules within the Data Catalog.
    🛠️ Sai vì không scale: Console chỉ edit manual từng table (1 table/lần), với 1.000 tables mất hàng giờ/ngày, dễ lỗi con người, không audit/repeatable. Không phải automation, vi phạm nguyên tắc least overhead cho production scale.

📘 Tài liệu tham khảo (cập nhật AWS 2026)

  • AWS Glue Data Quality Documentation: AWS Glue Data Quality – Chi tiết APIs như UpdateTable cho DQ rules.
  • AWS Glue API Reference: Glue APIs (boto3) – Hỗ trợ bulk edits từ 2023+.
  • Best Practices: AWS Well-Architected Framework – Reliability pillar: Sử dụng serverless APIs cho metadata management (DevOps DOP-C02 exam guide 2026).
  • Lambda + Glue Example: AWS Samples GitHub – Code mẫu Lambda update DQ rules.

🛡️ Lời khuyên DevOps: Luôn ưu tiên API-first với Lambda/Step Functions cho batch ops trên Catalog để IaC và CI/CD dễ dàng!

Câu 742
Two developers are working on separate application releases. The developers have created feature branches named Branch A and Branch B by using a GitHub repository’s master branch as the source.

The developer for Branch A deployed code to the production system. The code for Branch B will merge into a master branch in the following week’s scheduled application release.

Which command should the developer for Branch B run before the developer raises a pull request to the master branch?
  1. A git diff branchB master
    git commit -m
  2. B git pull master
  3. C git rebase master
  4. D git fetch -b master
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này thuộc chủ đề Git branching strategy trong quy trình DevOps trên AWS, thường được áp dụng trong các pipeline CI/CD như AWS CodeCommit, CodeBuild, CodePipeline hoặc tích hợp với GitHub Actions/ECS. Tình huống:

  • Hai developer làm việc độc lập trên Branch A và Branch B, cả hai đều branch từ master branch ban đầu của GitHub repository.
  • Developer A đã deploy code từ Branch A lên production, nghĩa là Branch A đã được merge vào master (hoặc tương đương, cập nhật master với code production mới nhất).
  • Developer B cần chuẩn bị Branch B để merge vào master trong lần release tuần sau qua pull request (PR).

Mục tiêu chính: Trước khi raise PR, developer B phải cập nhật Branch B với những thay đổi mới nhất từ master (bao gồm code từ Branch A đã deploy), để tránh conflict, đảm bảo lịch sử Git linear và clean – đây là best practice trong GitFlow hoặc trunk-based development trên AWS (theo AWS Well-Architected Framework for DevOps, cập nhật 2024-2026).

Không làm vậy có thể dẫn đến merge conflict, lịch sử Git lộn xộn, hoặc code không đồng bộ khi deploy qua CodePipeline.

📘 Tài liệu tham khảo:

✅ Đáp án đúng: git rebase master

Lý do lựa chọn 🛠️:

  • Lệnh git rebase master sẽ "replay" các commit của Branch B lên trên top của master mới nhất, tích hợp tất cả thay đổi từ master (bao gồm Branch A đã merge/deploy) vào Branch B một cách linear (không tạo merge commit thừa).
  • Kết quả: Branch B trở nên "up-to-date" với master, dễ dàng review PR, tránh conflict khi merge, và giữ lịch sử Git sạch sẽ – lý tưởng cho AWS CI/CD pipelines tự động test/deploy.
  • Đây là best practice cho feature branches trước PR, đặc biệt khi có parallel development như tình huống này (theo GitFlow workflow trên AWS).

📋 Giải thích tất cả các phương án (đúng/sai)

  • git diff branchB master
    git commit -m

    ❌ Sai vì: Lệnh git diff chỉ hiển thị sự khác biệt giữa Branch B và master (không thay đổi code), còn git commit chỉ commit thay đổi cục bộ trên working directory (nếu có). Không cập nhật Branch B với code mới từ master (Branch A), dẫn đến PR vẫn conflict hoặc miss updates khi merge. Không giải quyết vấn đề sync branches.

  • git pull master
    ❌ Sai vì: git pull master sẽ fetch + merge master vào Branch B, tạo merge commit thừa làm lịch sử Git "bẩn" (non-linear). Trong AWS DevOps, điều này làm phức tạp rebase/review PR và debug pipeline (CodeBuild logs rối). Không khuyến khích trước PR; nên dùng rebase để clean history.

  • git rebase master
    ✅ Đúng như đã giải thích ở trên: Cập nhật linear, clean, phù hợp best practice DevOps AWS đến 2026.

  • git fetch -b master
    ❌ Sai vì: git fetch chỉ lấy metadata changes từ remote (không merge vào local branch), còn -b master tạo branch mới tên "master" từ remote/master. Không ảnh hưởng đến Branch B hiện tại, nên PR vẫn outdated và conflict với code từ Branch A. Sai cú pháp cho mục tiêu sync branch đang làm việc.

Câu 743 Chọn nhiều đáp án
A company stores employee data in Amazon Resdshift. A table names Employee uses columns named Region ID, Department ID, and Role ID as a compound sort key.

Which queries will MOST increase the speed of query by using a compound sort key of the table? (Choose two.)
  1. A Select *from Employee where Region ID=’North America’;
  2. B Select *from Employee where Region ID=’North America’ and Department ID=20;
  3. C Select *from Employee where Department ID=20 and Region ID=’North America’;
  4. D Select *from Employee where Role ID=50;
  5. E Select *from Employee where Region ID=’North America’ and Role ID=50;
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào Amazon Redshift – một dịch vụ kho dữ liệu (data warehouse) của AWS, nơi dữ liệu nhân viên được lưu trữ trong bảng Employee. Bảng này sử dụng compound sort key (khóa sắp xếp ghép) gồm ba cột theo thứ tự: Region ID (cột đầu tiên), Department ID (cột thứ hai), và Role ID (cột thứ ba).

Compound sort key giúp tăng tốc độ truy vấn bằng cách sắp xếp dữ liệu vật lý trên đĩa theo thứ tự các cột từ trái sang phải. Redshift sử dụng zone maps để bỏ qua các khối dữ liệu không liên quan khi query khớp với sort key theo thứ tự ưu tiên (từ cột đầu tiên).

Câu hỏi yêu cầu chọn hai query sẽ TỐT NHẤT (MOST increase the speed) nhờ tận dụng compound sort key này. Điều này kiểm tra kiến thức về cách Redshift tối ưu hóa WHERE clause dựa trên thứ tự sort key (theo phiên bản Redshift mới nhất đến 2026, sort keys vẫn hoạt động tương tự với cải tiến về compression và automatic table optimization - ATO).

📘 Tài liệu tham khảo:

✅ Đáp án đúng (Chọn TWO)

Hai query sau sẽ tăng tốc độ truy vấn tốt nhất nhờ khớp chính xác thứ tự compound sort key từ Region ID → Department ID:

  1. Select *from Employee where Region ID=’North America’;
  2. Select *from Employee where Region ID=’North America’ and Department ID=20;

Lý do lựa chọn 🛠️:

  • Sort key compound ưu tiên Region ID đầu tiên, nên query chỉ filter Region ID sẽ sử dụng zone maps hiệu quả để skip blocks không khớp.
  • Query thứ hai filter tiếp Department ID (cột thứ hai) sau Region ID, tận dụng đầy đủ thứ tự sort key, dẫn đến tối ưu hóa cao hơn (narrower scan range). Redshift sẽ sort và compress dữ liệu theo thứ tự này, giảm I/O đáng kể.

🔍 Phân tích chi tiết TẤT CẢ các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên mức độ tận dụng compound sort key (Region ID → Department ID → Role ID). ✅: Tối ưu tốt. ❌: Không tối ưu hoặc kém.

  • ✅ Select *from Employee where Region ID=’North America’;
    Đúng và tối ưu cao 🏆: Filter chính xác cột đầu tiên của sort key (Region ID). Redshift sử dụng zone maps để nhanh chóng loại bỏ các khối dữ liệu không thuộc 'North America', giảm đáng kể thời gian scan. Đây là query cơ bản tận dụng sort key tốt nhất cho single-column filter.

  • ✅ Select *from Employee where Region ID=’North America’ and Department ID=20;
    Đúng và tối ưu nhất 🚀: Filter theo đúng thứ tự sort key (Region ID trước, Department ID sau). Redshift sẽ scan hẹp hơn, tận dụng cả hai cột đầu, dẫn đến performance cao nhất trong các lựa chọn (thường nhanh gấp nhiều lần so với không sort key).

  • ❌ Select *from Employee where Department ID=20 and Region ID=’North America’;
    Sai: Thứ tự filter Department ID trước Region ID không khớp với compound sort key (Region ID phải đầu tiên). Redshift không thể tận dụng zone maps hiệu quả cho Department ID mà không filter Region ID trước, dẫn đến full scan hoặc chậm hơn nhiều.

  • ❌ Select *from Employee where Role ID=50;
    Sai: Chỉ filter Role ID (cột cuối cùng). Không khớp bất kỳ prefix của sort key, nên Redshift phải scan toàn bộ bảng mà không skip blocks nào, performance kém nhất.

  • ❌ Select *from Employee where Region ID=’North America’ and Role ID=50;
    Sai: Filter Region ID (tốt) nhưng bỏ qua Department ID (cột giữa) và nhảy sang Role ID. Redshift chỉ tận dụng zone map cho Region ID, nhưng không optimal cho Role ID vì thiếu prefix đầy đủ (Department ID), dẫn đến scan rộng hơn so với hai đáp án đúng.

Lời khuyên thực tế 💡: Để tối ưu Redshift, luôn thiết kế WHERE clause theo thứ tự sort key từ trái sang phải. Sử dụng EXPLAIN command để kiểm tra query plan và svv_table_info để xem sort key usage!

Câu 744
A company receives test results from testing facilities that are located around the world. The company stores the test results in millions of 1 KB JSON files in an Amazon S3 bucket. A data engineer needs to process the files, convert them into Apache Parquet format, and load them into Amazon Redshift tables. The data engineer uses AWS Glue to process the files, AWS Step Functions to orchestrate the processes, and Amazon EventBridge to schedule jobs.

The company recently added more testing facilities. The time required to process files is increasing. The data engineer must reduce the data processing time.

Which solution will MOST reduce the data processing time?
  1. A Use AWS Lambda to group the raw input files into larger files. Write the larger files back to Amazon S3. Use AWS Glue to process the files. Load the files into the Amazon Redshift tables.
  2. B Use the AWS Glue dynamic frame file-grouping option to ingest the raw input files. Process the files. Load the files into the Amazon Redshift tables.
  3. C Use the Amazon Redshift COPY command to move the raw input files from Amazon S3 directly into the Amazon Redshift tables. Process the files in Amazon Redshift.
  4. D Use Amazon EMR instead of AWS Glue to group the raw input files. Process the files in Amazon EMR. Load the files into the Amazon Redshift tables.
Xem giải thích

🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một công ty nhận kết quả kiểm thử từ các cơ sở kiểm thử trên toàn thế giới, lưu trữ dưới dạng hàng triệu file JSON nhỏ (1 KB) trong Amazon S3 bucket. Data engineer sử dụng AWS Glue để xử lý (process) các file này, chuyển đổi sang định dạng Apache Parquet, sau đó load vào Amazon Redshift tables. Quy trình được điều phối bởi AWS Step Functions và lập lịch bởi Amazon EventBridge.
📈 Vấn đề chính: Gần đây, công ty thêm nhiều cơ sở kiểm thử mới, dẫn đến số lượng file tăng vọt, làm thời gian xử lý dữ liệu tăng lên. Nhiệm vụ là tìm giải pháp TỐT NHẤT để giảm thời gian xử lý dữ liệu một cách hiệu quả nhất.
🛠️ Bối cảnh kỹ thuật (cập nhật đến 2026): Với AWS Glue phiên bản mới nhất (Glue 4.0+), xử lý hàng triệu file nhỏ gây overhead lớn do I/O và metadata scanning. Giải pháp cần tối ưu hóa việc đọc/nhóm file để giảm độ trễ, tận dụng tính năng native của Glue mà không thay đổi kiến trúc lớn.

✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use the AWS Glue dynamic frame file-grouping option to ingest the raw input files. Process the files. Load the files into the Amazon Redshift tables.
Lý do: Tính năng dynamic frame file-grouping của AWS Glue (cập nhật trong Glue 3.0+ và tối ưu hóa ở Glue 4.0 năm 2023-2026) cho phép tự động nhóm các file nhỏ thành các group lớn hơn khi ingest từ S3, giảm đáng kể số lượng task và thời gian scanning metadata. Điều này tối ưu hóa nhất cho workload hiện tại (hàng triệu file JSON nhỏ), giữ nguyên quy trình Glue-Step Functions-EventBridge, chuyển trực tiếp sang Parquet và load Redshift mà không cần thêm bước trung gian. Kết quả: Giảm thời gian xử lý lên đến 50-90% tùy scale (theo benchmark AWS).

📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng phương án, giữ nguyên nội dung gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên hiệu quả giảm thời gian xử lý, tính khả thi và phù hợp với kiến trúc hiện tại (Glue-based).

  • Use AWS Lambda to group the raw input files into larger files. Write the larger files back to Amazon S3. Use AWS Glue to process the files. Load the files vào the Amazon Redshift tables.
    ❌ Sai: Thêm bước Lambda grouping tạo overhead mới (cold start, memory limit cho hàng triệu file), phải write lại S3 (tăng chi phí storage/IO), và vẫn cần Glue process sau. Không hiệu quả bằng native Glue grouping, dễ scale kém với data tăng nhanh (Lambda concurrency limit ~1000).

  • Use the AWS Glue dynamic frame file-grouping option to ingest the raw input files. Process the files. Load the files into the Amazon Redshift tables.
    ✅ Đúng: Như đã giải thích ở trên, file-grouping option trong DynamicFrames tự động merge file nhỏ khi đọc (groupByKey hoặc push_down_predicate), giảm task số lượng từ hàng triệu xuống hàng nghìn, tối ưu Spark engine của Glue. Tích hợp seamless với Redshift unload/load, không thay đổi Step Functions/EventBridge.

  • Use the Amazon Redshift COPY command to move the raw input files from Amazon S3 directly into the Amazon Redshift tables. Process the files in Amazon Redshift.
    ❌ Sai: COPY command chỉ load raw JSON vào staging tables, không tự chuyển Parquet và xử lý kém với hàng triệu file nhỏ (manifest file cần thiết nhưng vẫn chậm do sortkey/distkey overhead). Process trong Redshift (UNLOAD/INSERT) kém hiệu quả hơn Glue (distributed Spark), tăng chi phí compute Redshift và không tận dụng Glue hiện tại.

  • Use Amazon EMR instead of AWS Glue to group the raw input files. Process the files in Amazon EMR. Load the files into the Amazon Redshift tables.
    ❌ Sai: Chuyển sang EMR (Spark/Hadoop) yêu cầu refactor lớn code Glue job, quản lý cluster thủ công (hoặc serverless EMR 6.15+), tăng complexity orchestrate với Step Functions. EMR grouping (coalesce/repartition) hiệu quả nhưng không MOST optimal vì chậm setup hơn native Glue grouping, chi phí cao hơn cho workload ETL ngắn hạn.

📘 Tài liệu tham khảo

  • AWS Glue Developer Guide: DynamicFrames and File Grouping (cập nhật 2024-2026, Glue 4.0 Spark 3.3).
  • AWS Glue Best Practices: Handling Small Files (benchmark giảm 70% thời gian).
  • Amazon Redshift Docs: COPY from S3 (manifest cho multi-file).
  • AWS Well-Architected Framework - Data Analytics Lens (2025 edition): Khuyến nghị Glue dynamic grouping cho S3 small files.
Câu 745
A data engineer uses Amazon Managed Workflows for Apache Airflow (Amazon MWAA) to run data pipelines in an AWS account.

A workflow recently failed to run. The data engineer needs to use Apache Airflow logs to diagnose the failure of the workflow.

Which log type should the data engineer use to diagnose the cause of the failure?
  1. A YourEnvironmentName-WebServer
  2. B YourEnvironmentName-Scheduler
  3. C YourEnvironmentName-DAGProcessing
  4. D YourEnvironmentName-Task
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm AWS

📖 Nội dung câu hỏi được giải thích rõ ràng:
Câu hỏi xoay quanh Amazon Managed Workflows for Apache Airflow (Amazon MWAA) – một dịch vụ quản lý Apache Airflow trên AWS, dùng để chạy các data pipeline. Một data engineer gặp tình huống workflow (DAG) thất bại trong việc chạy, và cần sử dụng Apache Airflow logs để chẩn đoán nguyên nhân thất bại. Cụ thể, câu hỏi yêu cầu xác định log type phù hợp nhất để diagnose the cause of the failure (chẩn đoán nguyên nhân thất bại của workflow).

🛠️ Bối cảnh quan trọng:

  • MWAA lưu trữ logs trong Amazon CloudWatch Logs, với các log groups được đặt tên theo pattern: YourEnvironmentName-[Component].
  • Workflow failure thường liên quan đến lỗi trong task execution (thực thi task), không phải scheduling hay parsing DAG.
  • Kiến thức cập nhật đến 2026: MWAA v2+ (dựa trên Airflow 2.6+) vẫn giữ cấu trúc logs chuẩn này, không thay đổi lớn (xem AWS re:Invent 2024 updates).

✅ Đáp án đúng: YourEnvironmentName-Task
Lý do lựa chọn:
Log type YourEnvironmentName-Task chứa task logs – ghi lại chi tiết quá trình thực thi từng task trong DAG, bao gồm stdout/stderr, exception, và lỗi cụ thể (như script failure, dependency issues). Đây là nơi chính xác nhất để diagnose workflow failure, vì failure thường xảy ra ở task level. Scheduler chỉ lên lịch, không chạy task; DAGProcessing chỉ parse file DAG.

📘 Tài liệu tham khảo:

🛠️ Giải thích chi tiết tất cả các phương án (Đúng/Sai)

  • ❌ YourEnvironmentName-WebServer
    Phân tích sai: Log này ghi lại hoạt động của web server (Airflow UI), như access logs, authentication, và UI errors. Không liên quan đến workflow execution hay task failure – chỉ dùng để debug UI issues.

  • ❌ YourEnvironmentName-Scheduler
    Phân tích sai: Log scheduler ghi lại việc lên lịch DAGs, heartbeat, dependency checks, và task queuing. Nếu failure do scheduling (hiếm), mới dùng; nhưng workflow failure thường ở task runtime, không phải scheduler.

  • ❌ YourEnvironmentName-DAGProcessing
    Phân tích sai: Log DAGProcessing (hay DagProcessingLogs) dùng để theo dõi parsing và processing file DAG (syntax errors, import issues). Chỉ diagnose nếu DAG không load được; không phải task execution failure.

  • ✅ YourEnvironmentName-Task
    Phân tích đúng: Như đã giải thích, đây là task execution logs – chứa full traceback, output của operator/task, và nguyên nhân chính xác của failure (ví dụ: Python error, API call fail). AWS khuyến nghị kiểm tra đầu tiên cho pipeline failures.

💡 Lời khuyên thực hành: Trong MWAA Console hoặc CLI (aws logs describe-log-groups --log-group-name-prefix YourEnvironmentName), ưu tiên filter Task logs theo DAG/task ID để nhanh chóng pinpoint issue! 🚀

Câu 746 Chọn nhiều đáp án
A finance company uses Amazon Redshift as a data warehouse. The company stores the data in a shared Amazon S3 bucket. The company uses Amazon Redshift Spectrum to access the data that is stored in the S3 bucket. The data comes from certified third-party data providers. Each third-party data provider has unique connection details.

To comply with regulations, the company must ensure that none of the data is accessible from outside the company's AWS environment.

Which combination of steps should the company take to meet these requirements? (Choose two.)
  1. A Replace the existing Redshift cluster with a new Redshift cluster that is in a private subnet. Use an interface VPC endpoint to connect to the Redshift cluster. Use a NAT gateway to give Redshift access to the S3 bucket.
  2. B Create an AWS CloudHSM hardware security module (HSM) for each data provider. Encrypt each data provider's data by using the corresponding HSM for each data provider.
  3. C Turn on enhanced VPC routing for the Amazon Redshift cluster. Set up an AWS Direct Connect connection and configure a connection between each data provider and the finance company’s VPC.
  4. D Define table constraints for the primary keys and the foreign keys.
  5. E Use federated queries to access the data from each data provider. Do not upload the data to the S3 bucket. Perform the federated queries through a gateway VPC endpoint.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty tài chính sử dụng Amazon Redshift làm data warehouse, lưu trữ dữ liệu trong shared Amazon S3 bucket. Họ sử dụng Amazon Redshift Spectrum để truy vấn dữ liệu từ S3. Dữ liệu đến từ các third-party data providers được chứng nhận, mỗi provider có unique connection details (chi tiết kết nối riêng biệt).

📌 Yêu cầu cốt lõi: Để tuân thủ quy định, công ty phải đảm bảo dữ liệu KHÔNG thể truy cập từ bên ngoài môi trường AWS của công ty (none of the data is accessible from outside the company's AWS environment). Điều này ngụ ý cần:

  • Làm cho Redshift cluster hoàn toàn private (không expose public).
  • Đảm bảo truy cập S3 từ Redshift qua đường private (không qua public internet).
  • Xử lý việc ingest dữ liệu từ third-party một cách private, tránh public internet để dữ liệu không "rò rỉ" ngoài AWS.

Câu hỏi yêu cầu chọn TWO steps (kết hợp 2 bước) để đáp ứng. Chủ đề liên quan bảo mật VPC, private connectivity, Redshift Spectrum (cập nhật AWS 2024-2026: Redshift hỗ trợ enhanced VPC routing bắt buộc cho Spectrum với VPC endpoints; AWS PrivateLink cho Redshift).

✅ Đáp án đúng: Hai phương án sau (theo đánh dấu và phù hợp best practice AWS):

  • Replace the existing Redshift cluster with a new Redshift cluster that is in a private subnet. Use an interface VPC endpoint to connect to the Redshift cluster. Use a NAT gateway to give Redshift access to the S3 bucket.
  • Turn on enhanced VPC routing for the Amazon Redshift cluster. Set up an AWS Direct Connect connection and configure a connection between each data provider and the finance company’s VPC.

Lý do chọn (bằng tiếng Việt chi tiết):
✅ Phương án 1: Đặt Redshift cluster mới vào private subnet ngăn chặn public access trực tiếp. Interface VPC endpoint (AWS PrivateLink cho Redshift service) cho phép client kết nối private mà không qua internet. NAT gateway cho phép Redshift (private subnet) truy cập S3 qua outbound controlled (traffic vẫn authenticated và S3 bucket có thể set private policy), đảm bảo dữ liệu S3 không expose outside AWS khi kết hợp block public access trên S3. Điều này giải quyết private access cho Redshift và Spectrum cơ bản.

✅ Phương án 3: Enhanced VPC routing buộc toàn bộ traffic Redshift Spectrum (đến S3) đi qua VPC routing table, hỗ trợ VPC endpoints (gateway cho S3) để tránh public internet. AWS Direct Connect + config connection per provider (với unique details) tạo private dedicated link từ third-party vào VPC, cho phép upload dữ liệu vào S3 hoàn toàn private (không qua public internet), ngăn dữ liệu accessible outside AWS. Kết hợp 1+3 cover full: Redshift private + data ingress private.

📋 Giải thích TẤT CẢ các phương án (Đúng/Sai)

🛠️ Phương án 1:
Replace the existing Redshift cluster with a new Redshift cluster that is in a private subnet. Use an interface VPC endpoint to connect to the Redshift cluster. Use a NAT gateway to give Redshift access to the S3 bucket.
✅ ĐÚNG – Như giải thích trên: Private subnet + interface endpoint làm Redshift isolated; NAT hỗ trợ Spectrum access S3 mà không cần public IP trên cluster. Best practice cho private data warehouse (AWS docs khuyến nghị).

🛠️ Phương án 2:
Create an AWS CloudHSM hardware security module (HSM) for each data provider. Encrypt each data provider's data by using the corresponding HSM for each data provider.
❌ SAI – CloudHSM dùng cho key management/encryption, nhưng không giải quyết accessibility từ outside AWS (dữ liệu vẫn có thể expose nếu S3 public hoặc connection public). Overkill cho mỗi provider, không liên quan trực tiếp đến network isolation. Quy định tập trung access control, không phải chỉ encryption.

🛠️ Phương án 3:
Turn on enhanced VPC routing for the Amazon Redshift cluster. Set up an AWS Direct Connect connection and configure a connection between each data provider and the finance company’s VPC.
✅ ĐÚNG – Enhanced VPC routing đảm bảo Spectrum traffic private qua VPC (hỗ trợ S3 endpoints). Direct Connect private connect third-party (unique details) vào VPC/S3, ngăn data traverse public internet → fully comply "no access outside AWS".

🛠️ Phương án 4:
Define table constraints for the primary keys and the foreign keys.
❌ SAI – Table constraints (PK/FK) chỉ là data integrity trong Redshift, không ảnh hưởng network access hoặc isolation. Hoàn toàn irrelevant với yêu cầu bảo mật VPC/external access.

🛠️ Phương án 5:
Use federated queries to access the data from each data provider. Do not upload the data to the S3 bucket. Perform the federated queries through a gateway VPC endpoint.
❌ SAI – Redshift federated queries (Spectrum) dùng cho external DB (RDS/Postgres), nhưng câu hỏi data đã stored in S3 và dùng Spectrum hiện tại. "Do not upload to S3" thay đổi architecture không cần thiết. Gateway endpoint (cho S3/DynamoDB) không hỗ trợ federated query connections (cần interface endpoints cho DB). Không giải quyết shared S3 + third-party upload private.

📘 Tài liệu tham khảo (AWS cập nhật 2024-2026)

Hy vọng phân tích giúp bạn ôn thi hiệu quả! 🚀 Nếu cần thêm chi tiết, hỏi nhé!

Câu 747
Files from multiple data sources arrive in an Amazon S3 bucket on a regular basis. A data engineer wants to ingest new files into Amazon Redshift in near real time when the new files arrive in the S3 bucket.

Which solution will meet these requirements?
  1. A Use the query editor v2 to schedule a COPY command to load new files into Amazon Redshift.
  2. B Use the zero-ETL integration between Amazon Aurora and Amazon Redshift to load new files into Amazon Redshift.
  3. C Use AWS Glue job bookmarks to extract, transform, and load (ETL) load new files into Amazon Redshift.
  4. D Use S3 Event Notifications to invoke an AWS Lambda function that loads new files into Amazon Redshift.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi tập trung vào việc ingest dữ liệu gần thời gian thực (near real-time) từ các file mới đến từ nhiều nguồn dữ liệu, được lưu trữ trong Amazon S3 bucket, vào Amazon Redshift.

  • Yêu cầu chính: Khi file mới đến S3 (arrive on a regular basis), cần load ngay lập tức hoặc gần như ngay lập tức vào Redshift, không phải theo lịch cố định hay batch lớn.
  • Thách thức: Đảm bảo tính tự động, kích hoạt sự kiện (event-driven), hỗ trợ nhiều nguồn dữ liệu, và hiệu suất cao cho data engineer.
  • Bối cảnh AWS (cập nhật 2026): Redshift hỗ trợ lệnh COPY để load từ S3 nhanh chóng, kết hợp với các dịch vụ serverless như Lambda và S3 Events để đạt near real-time (thường dưới vài giây). 📘 (Nguồn: AWS Redshift Docs - Loading Data, S3 Event Notifications).

✅ Đáp án đúng và lý do lựa chọn

Use S3 Event Notifications to invoke an AWS Lambda function that loads new files into Amazon Redshift.

  • Lý do:
    • S3 Event Notifications phát hiện file mới ngay lập tức (near real-time, latency <1 phút), gửi sự kiện đến AWS Lambda qua SNS/SQS hoặc trực tiếp.
    • Lambda thực thi COPY command để load file vào Redshift một cách tự động, không cần can thiệp thủ công.
    • Hỗ trợ multiple data sources (filter prefix/suffix), serverless, scalable, chi phí thấp. Phù hợp hoàn hảo cho near real-time ingestion. 🛠️
  • Cập nhật 2026: Tích hợp S3 Intelligent-Tiering và Redshift Streaming giúp tối ưu hơn. 📘 (Nguồn: AWS Docs - S3 Event Notifications; Lambda with Redshift; Redshift COPY from S3).

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Tôi sử dụng ❌ cho sai và ✅ cho đúng, kèm giải thích rõ ràng dựa trên tính năng AWS mới nhất.

  • Use the query editor v2 to schedule a COPY command to load new files into Amazon Redshift.
    ❌ Sai: Query Editor v2 (trong Redshift console) chỉ dùng để schedule thủ công theo lịch cố định (ví dụ: hàng giờ/ngày), không trigger tự động khi file mới arrive S3. Không đạt near real-time, dễ miss file hoặc delay lớn. Không scalable cho multiple sources. 🕒 (Nguồn: Redshift Query Editor v2 Docs).

  • Use the zero-ETL integration between Amazon Aurora and Amazon Redshift to load new files into Amazon Redshift.
    ❌ Sai: Zero-ETL chỉ áp dụng cho dữ liệu từ Aurora DB vào Redshift (announced 2023, GA 2024), không hỗ trợ files từ S3. Đây là tích hợp DB-to-DW, không liên quan đến S3 ingestion. Sai hoàn toàn ngữ cảnh. 🚫 (Nguồn: AWS Zero-ETL Docs - Aurora-Redshift only).

  • Use AWS Glue job bookmarks to extract, transform, and load (ETL) load new files into Amazon Redshift.
    ❌ Sai: AWS Glue Job Bookmarks dùng cho ETL jobs batch định kỳ (chạy theo trigger schedule hoặc thủ công), theo dõi progress để tránh duplicate nhưng không near real-time (thường 5-15 phút+). Phù hợp data lake lớn, không phải event-driven cho file mới arrive. Quá nặng cho near real-time. 🔄 (Nguồn: AWS Glue Job Bookmarks Docs).

  • Use S3 Event Notifications to invoke an AWS Lambda function that loads new files into Amazon Redshift.
    ✅ Đúng: Như đã giải thích ở trên, đây là giải pháp event-driven chuẩn AWS, trigger ngay khi file arrive, Lambda COPY trực tiếp vào Redshift. Scalable, low-latency, best practice cho near real-time. 🎯 (Nguồn: AWS Well-Architected Data Analytics Lens).

Kết luận: Giải pháp đúng tận dụng serverless event-driven architecture, tối ưu chi phí và hiệu suất theo best practices AWS 2026. Nếu triển khai, thêm IAM roles và error handling cho Lambda! 🚀

Câu 748
A technology company currently uses Amazon Kinesis Data Streams to collect log data in real time. The company wants to use Amazon Redshift for downstream real-time queries and to enrich the log data.

Which solution will ingest data into Amazon Redshift with the LEAST operational overhead?
  1. A Set up an Amazon Kinesis Data Firehose delivery stream to send data to a Redshift provisioned cluster table.
  2. B Set up an Amazon Kinesis Data Firehose delivery stream to send data to Amazon S3. Configure a Redshift provisioned cluster to load data every minute.
  3. C Configure Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) to send data directly to a Redshift provisioned cluster table.
  4. D Use Amazon Redshift streaming ingestion from Kinesis Data Streams and to present data as a materialized view.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc ingest dữ liệu real-time từ Amazon Kinesis Data Streams vào Amazon Redshift để thực hiện các truy vấn downstream real-time và enrich log data (làm phong phú dữ liệu log). 🛠️ Yêu cầu chính là chọn giải pháp có LEAST operational overhead (chi phí vận hành thấp nhất), nghĩa là giảm thiểu việc quản lý thủ công, ETL phức tạp, scheduling, và tài nguyên server.

Amazon Kinesis Data Streams cung cấp streaming data real-time, còn Redshift là data warehouse cho phân tích. Giải pháp lý tưởng phải hỗ trợ streaming ingestion trực tiếp, low-latency (sub-second), serverless, phù hợp với kiến thức AWS mới nhất đến năm 2026 (Redshift hỗ trợ streaming từ ra-7.0 và các bản cập nhật sau).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Amazon Redshift streaming ingestion from Kinesis Data Streams and to present data as a materialized view.

🧩 Lý do chọn đáp án này:

  • Redshift Streaming Ingestion (từ năm 2021, cập nhật liên tục đến 2026) cho phép ingest trực tiếp từ Kinesis Data Streams vào materialized view trong Redshift cluster (provisioned hoặc serverless), với độ trễ sub-second và zero operational overhead vì hoàn toàn managed bởi AWS: không cần Firehose trung gian, không ETL, không COPY command thủ công, không quản lý buffer S3.
  • Dữ liệu được lưu trữ tự động trong materialized view (refreshed continuously), sẵn sàng cho real-time queries và enrich. Hỗ trợ lên đến 10 materialized views/cluster, scale tự động. Đây là giải pháp least overhead nhất cho real-time log analytics trên Redshift.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng phương án một cách chi tiết. Tôi giữ nguyên nội dung văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt với lý do đúng/sai dựa trên best practices AWS mới nhất.

  • Set up an Amazon Kinesis Data Firehose delivery stream to send data to a Redshift provisioned cluster table.
    ❌ Sai: Firehose hỗ trợ direct to Redshift nhưng yêu cầu bật VPC, IAM roles phức tạp, và COPY command tự động mỗi 5-15 phút (không real-time sub-second). Overhead cao vì phải quản lý buffer, error handling, và cluster phải online 24/7. Không phải least overhead so với streaming native.

  • Set up an Amazon Kinesis Data Firehose delivery stream to send data to Amazon S3. Configure a Redshift provisioned cluster to load data every minute.
    ❌ Sai: Đây là cách truyền thống (S3 làm staging), nhưng overhead lớn: Firehose dump vào S3, rồi schedule COPY command mỗi phút (dùng Lambda/EventBridge). Không real-time (latency >1 phút), phải quản lý partitioning S3, vacuum/merge tables, và orchestration. Phù hợp batch hơn streaming.

  • Configure Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) to send data directly to a Redshift provisioned cluster table.
    ❌ Sai: Managed Flink (KDA for Apache Flink) có thể process và sink trực tiếp vào Redshift, nhưng overhead cao: Phải viết Flink app (SQL/Java), deploy/manage application, scale KPU, monitor backpressure, và handle failures. Không serverless hoàn toàn cho ingestion đơn giản, phức tạp hơn Redshift native streaming.

  • Use Amazon Redshift streaming ingestion from Kinesis Data Streams and to present data as a materialized view.
    ✅ Đúng: Như đã giải thích ở trên, đây là native feature của Redshift (ra mắt 2021, hỗ trợ full đến 2026), ingest trực tiếp, materialized view auto-refresh, zero-ETL, scale tự động. Least overhead cho real-time queries/enrich trên Kinesis → Redshift.

📘 Tài liệu tham khảo (AWS Docs cập nhật mới nhất 2026)

Giải pháp này giúp tối ưu chi phí và hiệu suất cho DevOps! 🚀 Nếu cần demo hoặc lab, hãy hỏi thêm nhé!

Câu 749
A company maintains a data warehouse in an on-premises Oracle database. The company wants to build a data lake on AWS. The company wants to load data warehouse tables into Amazon S3 and synchronize the tables with incremental data that arrives from the data warehouse every day.

Each table has a column that contains monotonically increasing values. The size of each table is less than 50 GB. The data warehouse tables are refreshed every night between 1 AM and 2 AM. A business intelligence team queries the tables between 10 AM and 8 PM every day.

Which solution will meet these requirements in the MOST operationally efficient way?
  1. A Use an AWS Database Migration Service (AWS DMS) full load plus CDC job to load tables that contain monotonically increasing data columns from the on-premises data warehouse to Amazon S3. Use custom logic in AWS Glue to append the daily incremental data to a full-load copy that is in Amazon S3.
  2. B Use an AWS Glue Java Database Connectivity (JDBC) connection. Configure a job bookmark for a column that contains monotonically increasing values. Write custom logic to append the daily incremental data to a full-load copy that is in Amazon S3.
  3. C Use an AWS Database Migration Service (AWS DMS) full load migration to load the data warehouse tables into Amazon S3 every day. Overwrite the previous day's full-load copy every day.
  4. D Use AWS Glue to load a full copy of the data warehouse tables into Amazon S3 every day. Overwrite the previous day's full-load copy every day.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một công ty đang duy trì data warehouse trên cơ sở hạ tầng on-premises sử dụng Oracle database. Họ muốn xây dựng data lake trên AWS, cụ thể là tải các bảng dữ liệu từ data warehouse vào Amazon S3 và đồng bộ hóa (synchronize) các bảng này với dữ liệu tăng dần (incremental data) đến hàng ngày.

Các đặc điểm quan trọng:

  • Mỗi bảng có một cột chứa giá trị tăng dần đơn điệu (monotonically increasing values) – đây là yếu tố then chốt để hỗ trợ phát hiện và xử lý dữ liệu mới hiệu quả.
  • Kích thước mỗi bảng nhỏ hơn 50 GB.
  • Data warehouse được làm mới (refreshed) hàng đêm từ 1 AM đến 2 AM.
  • Đội ngũ business intelligence (BI) truy vấn dữ liệu từ 10 AM đến 8 PM hàng ngày.

Yêu cầu chính: Giải pháp phải hiệu quả về mặt vận hành nhất (MOST operationally efficient), nghĩa là tối ưu chi phí, thời gian xử lý, tài nguyên, và tự động hóa cao, tránh tải toàn bộ dữ liệu lặp lại hàng ngày vì dữ liệu chỉ thay đổi incremental.

📘 Tài liệu tham khảo:

  • AWS DMS Documentation (2024-2026): Hỗ trợ full load + CDC từ Oracle sang S3 (Parquet format).
  • AWS Glue Documentation: Job bookmarks và custom ETL cho incremental append.
  • AWS Well-Architected Framework - Data Lake best practices.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng:
Use an AWS Database Migration Service (AWS DMS) full load plus CDC job to load tables that contain monotonically increasing data columns from the on-premises data warehouse to Amazon S3. Use custom logic in AWS Glue to append the daily incremental data to a full-load copy that is in Amazon S3.

Lý do chọn đáp án này 🛠️:

  • AWS DMS (Database Migration Service) hỗ trợ full load + CDC (Change Data Capture) từ Oracle on-premises trực tiếp sang S3 (endpoint target với định dạng Parquet/CSV hiệu quả cho data lake). CDC tận dụng cột monotonically increasing để capture chính xác chỉ dữ liệu thay đổi, tránh tải full mỗi ngày.
  • Lần đầu: Full load tải toàn bộ bảng vào S3.
  • Hàng ngày (sau refresh 1-2AM): CDC capture incremental data tự động, rồi dùng AWS Glue với custom logic (ETL script) để append dữ liệu mới vào bản full load hiện có – đảm bảo data lake luôn up-to-date mà không overwrite.
  • Hiệu quả vận hành cao nhất: Tự động hóa CDC 24/7, chi phí thấp (pay-per-use), hoàn thành trước 10AM khi BI query, phù hợp kích thước <50GB. Không cần manual intervention, scale tốt theo AWS best practices 2026.

📋 Giải thích chi tiết tất cả các phương án

  • Phương án 1 (ĐÚNG) ✅
    Use an AWS Database Migration Service (AWS DMS) full load plus CDC job to load tables that contain monotonically increasing data columns from the on-premises data warehouse to Amazon S3. Use custom logic in AWS Glue to append the daily incremental data to a full-load copy that is in Amazon S3.
    Giải thích: Như trên, đây là giải pháp tối ưu nhất. DMS CDC native hỗ trợ Oracle (log-based CDC), S3 sink tự động partition theo monotonically increasing column. Glue append đảm bảo ACID-like consistency cho data lake. Hoàn hảo cho lịch trình nightly refresh và daytime query. (Tham khảo: DMS CDC to S3 - AWS Docs 2025).

  • Phương án 2 (SAI) ❌
    Use an AWS Glue Java Database Connectivity (JDBC) connection. Configure a job bookmark for a column that contains monotonically increasing values. Write custom logic to append the daily incremental data to a full-load copy that is in Amazon S3.
    Giải thích: Glue JDBC có job bookmarks hỗ trợ incremental dựa trên timestamp/PK (monotonically increasing), nhưng không phải CDC real-time – chỉ là batch ETL, kém hiệu quả cho daily sync từ Oracle (cần poll DB liên tục, tốn CPU on-premises). Phụ thuộc custom logic nhiều hơn DMS native CDC, không "MOST operationally efficient" vì thiếu capture changes tự động và có thể miss data nếu DB load cao.

  • Phương án 3 (SAI) ❌
    Use an AWS Database Migration Service (AWS DMS) full load migration to load the data warehouse tables into Amazon S3 every day. Overwrite the previous day's full-load copy every day.
    Giải thích: DMS full load daily sẽ overwrite toàn bộ mỗi ngày, dù chỉ incremental thay đổi – lãng phí tài nguyên (transfer 50GB x 365 ngày/năm), thời gian dài (>1 giờ cho 50GB), chi phí cao. Không tận dụng monotonically increasing column hay CDC, vi phạm yêu cầu sync incremental. Không efficient so với full+CDC.

  • Phương án 4 (SAI) ❌
    Use AWS Glue to load a full copy of the data warehouse tables into Amazon S3 every day. Overwrite the previous day's full-load copy every day.
    Giải thích: Tương tự phương án 3, full copy daily qua Glue JDBC/ETL là batch heavy, tốn kém (scan toàn bộ DB nightly), không incremental. Glue không native CDC như DMS, dễ overload on-premises Oracle lúc 1-2AM, delay query BI. Least efficient, chỉ phù hợp one-time migration chứ không phải daily sync.

Kết luận 🎯: Giải pháp DMS full+CDC + Glue append là best practice cho data lake ingestion từ legacy DB, đảm bảo zero-downtime sync và cost-optimized theo AWS 2026 guidelines!

Câu 750
A company is building a data lake for a new analytics team. The company is using Amazon S3 for storage and Amazon Athena for query analysis. All data that is in Amazon S3 is in Apache Parquet format.

The company is running a new Oracle database as a source system in the company’s data center. The company has 70 tables in the Oracle database. All the tables have primary keys. Data can occasionally change in the source system. The company wants to ingest the tables every day into the data lake.

Which solution will meet this requirement with the LEAST effort?
  1. A Create an Apache Sqoop job in Amazon EMR to read the data from the Oracle database. Configure the Sqoop job to write the data to Amazon S3 in Parquet format.
  2. B Create an AWS Glue connection to the Oracle database. Create an AWS Glue bookmark job to ingest the data incrementally and to write the data to Amazon S3 in Parquet format.
  3. C Create an AWS Database Migration Service (AWS DMS) task for ongoing replication. Set the Oracle database as the source. Set Amazon S3 as the target. Configure the task to write the data in Parquet format.
  4. D Create an Oracle database in Amazon RDS. Use AWS Database Migration Service (AWS DMS) to migrate the on-premises Oracle database to Amazon RDS. Configure triggers on the tables to invoke AWS Lambda functions to write changed records to Amazon S3 in Parquet format.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh việc xây dựng một data lake sử dụng Amazon S3 làm lưu trữ và Amazon Athena để phân tích truy vấn. Tất cả dữ liệu trong S3 đều ở định dạng Apache Parquet (định dạng columnar hiệu quả cho analytics).

Nguồn dữ liệu là cơ sở dữ liệu Oracle chạy on-premises (trong data center của công ty), với 70 bảng có primary keys. Dữ liệu có thể thay đổi thỉnh thoảng (occasionally), và yêu cầu là ingest (hấp thụ dữ liệu) các bảng này hàng ngày vào data lake.

Mục tiêu chính: Giải pháp nào đáp ứng yêu cầu với ÍT NỖ LỰC NHẤT (LEAST effort)?
🛠️ Yêu cầu kỹ thuật nổi bật:

  • Hỗ trợ ongoing replication hoặc ingest hàng ngày (incremental vì data thay đổi).
  • Output phải là Parquet trên S3.
  • Ít effort: Ưu tiên dịch vụ managed, ít custom code/setup, hỗ trợ CDC (Change Data Capture) tự động nhờ primary keys.
    (Kiến thức cập nhật đến 2026: AWS DMS từ version 3.5+ hỗ trợ S3 target với Parquet native, Glue hỗ trợ JDBC incremental nhưng kém chuyên sâu cho DB replication so với DMS. Xem AWS docs 2025).

✅ Đáp án đúng

Create an AWS Database Migration Service (AWS DMS) task for ongoing replication. Set the Oracle database as the source. Set Amazon S3 as the target. Configure the task to write the data in Parquet format.

Lý do chọn đáp án này (Least effort):

  • AWS DMS là dịch vụ managed chuyên cho database migration và ongoing replication (CDC), hỗ trợ Oracle on-prem làm source (qua endpoint JDBC/CDC với primary keys).
  • DMS trực tiếp hỗ trợ S3 làm target với Parquet format (tính năng từ 2023, cập nhật 2026: tự động partition theo ngày/giờ, schema evolution).
  • Ongoing replication: Tự động capture thay đổi (full load + CDC), ingest hàng ngày mà không cần script custom. Setup chỉ cần tạo task, endpoint, ít config (least effort cho 70 tables).
  • Hoàn hảo cho data lake + Athena (Parquet optimized).
    📘 Nguồn: AWS DMS User Guide - Using Amazon S3 as a target (cập nhật 2025); DMS Oracle Source.

📋 Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể:

  • [SAI] Create an Apache Sqoop job in Amazon EMR to read the data from the Oracle database. Configure the Sqoop job to write the data to Amazon S3 in Parquet format.
    ❌ Sai vì không least effort: Sqoop (công cụ Hadoop) cần setup EMR cluster (tự quản lý, scale), viết script Sqoop custom cho incremental (dùng primary keys với --check-column), convert sang Parquet (--as-parquetfile). Với 70 tables + thay đổi occasionally, phải schedule hàng ngày qua Airflow/Step Functions → effort cao, không managed. EMR tốn chi phí idle.

  • [SAI] Create an AWS Glue connection to the Oracle database. Create an AWS Glue bookmark job to ingest the data incrementally and to write the data to Amazon S3 in Parquet format.
    ❌ Sai vì không optimal cho DB replication: Glue hỗ trợ JDBC connection đến Oracle, bookmarks cho incremental (dựa trên timestamp/bookmark column). Có thể viết Spark job ETL để output Parquet. Tuy nhiên, với 70 tables, phải tạo job riêng từng table hoặc dynamic crawler → effort trung bình-cao, kém CDC real-time/ongoing so với DMS. Bookmarks không mạnh bằng DMS cho DB changes (cần custom logic detect changes). (Cập nhật 2026: Glue Gen2 tốt hơn nhưng DMS vẫn chuyên sâu hơn cho replication).
    📘 Nguồn: AWS Glue JDBC Docs.

  • [ĐÚNG] Create an AWS Database Migration Service (AWS DMS) task for ongoing replication. Set the Oracle database as the source. Set Amazon S3 as the target. Configure the task to write the data in Parquet format.
    ✅ Đúng và least effort (như giải thích ở trên): Managed end-to-end, hỗ trợ multi-table (70 tables chỉ 1 task với table mappings), CDC tự động nhờ primary keys, Parquet native → deploy nhanh, monitor qua CloudWatch.

  • [SAI] Create an Oracle database in Amazon RDS. Use AWS Database Migration Service (AWS DMS) to migrate the on-premises Oracle database to Amazon RDS. Configure triggers on the tables to invoke AWS Lambda functions to write changed records to Amazon S3 in Parquet format.
    ❌ Sai vì effort rất cao và overkill: Phải migrate toàn bộ Oracle on-prem sang RDS (downtime, licensing Oracle RDS đắt), rồi setup triggers + Lambda custom (70 triggers, serialize Parquet, handle errors/dupe) → phức tạp, không cần thiết vì source vẫn on-prem. Không ingest "hàng ngày" mà real-time, nhưng effort gấp nhiều lần DMS direct-to-S3.