Ngân hàng đề — AWS Certified Data Engineer Associate

Tìm thấy 867 câu.

Câu 831
A company uses Amazon Redshift as a data warehouse solution. One of the datasets that the company stores in Amazon Redshift contains data for a vendor.

Recently, the vendor asked the company to transfer the vendor’s data into the vendor’s Amazon S3 bucket once each week.

Which solution will meet this requirement?
  1. A Create an AWS Lambda function to connect to the Redshift data warehouse. Configure the Lambda function to use the Redshift COPY command to copy the required data to the vendor’s S3 bucket on a schedule.
  2. B Create an AWS Glue job to connect to the Redshift data warehouse. Configure the AWS Glue job to use the Redshift UNLOAD command to load the required data to the vendor’s S3 bucket on a schedule.
  3. C Use the Amazon Redshift data sharing feature. Set the vendor’s S3 bucket as the destination. Configure the source to be as a custom SQL query that selects the required data.
  4. D Configure Amazon Redshift Spectrum to use the vendor’s S3 bucket a destination, Enable data querying in both directions.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh việc xuất dữ liệu từ Amazon Redshift (data warehouse) sang S3 bucket của vendor bên thứ ba theo lịch hàng tuần. 🛤️

  • Bối cảnh: Công ty lưu trữ dữ liệu vendor trong Redshift. Vendor yêu cầu chuyển dữ liệu này ra S3 bucket của họ (không phải S3 của công ty), thực hiện mỗi tuần một lần (cần cơ chế lập lịch tự động). 📅
  • Yêu cầu kỹ thuật: Giải pháp phải kết nối Redshift, trích xuất dữ liệu cụ thể, xuất ra S3 bên ngoài, và chạy theo lịch mà không làm gián đoạn hoạt động Redshift. ⚙️
  • Thách thức chính: Redshift không hỗ trợ xuất trực tiếp ra S3 của account khác một cách đơn giản; cần quyền IAM cross-account và lệnh phù hợp như UNLOAD. Không dùng COPY (vì COPY dùng để load VÀO Redshift từ S3). 🔍
  • Kiến thức cập nhật 2026: AWS khuyến nghị dùng AWS Glue cho ETL job với Redshift integration (hỗ trợ UNLOAD từ Glue 3.0+), kết hợp Amazon EventBridge cho scheduling. Redshift RA3 nodes hỗ trợ UNLOAD hiệu quả hơn với concurrency scaling. 📈

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create an AWS Glue job to connect to the Redshift data warehouse. Configure the AWS Glue job to use the Redshift UNLOAD command to load the required data to the vendor’s S3 bucket on a schedule.

Lý do chi tiết:

  • AWS Glue là dịch vụ ETL serverless, hỗ trợ kết nối Redshift qua JDBC và chạy UNLOAD command để xuất dữ liệu song song ra S3 (hiệu suất cao, hỗ trợ định dạng Parquet/CSV). 🛠️
  • UNLOAD là lệnh chuẩn của Redshift để export dữ liệu từ table/query ra S3, hỗ trợ cross-account với IAM role (vendor cấp bucket policy). Lập lịch qua Glue Triggers hoặc EventBridge. ⏰
  • Ưu điểm: Tự động, scalable, chi phí thấp (pay-per-use), không cần EC2. Phù hợp với DevOps best practices (IaC via CDK/Terraform). 🚀
  • Không vi phạm: Vendor kiểm soát S3 của họ, công ty chỉ cần quyền write tạm thời.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn. Giữ nguyên văn bản gốc (tiếng Anh), chỉ giải thích bằng tiếng Việt với emoji đánh dấu. ❌ cho sai, ✅ cho đúng.

  • Create an AWS Lambda function to connect to the Redshift data warehouse. Configure the Lambda function to use the Redshift COPY command to copy the required data to the vendor’s S3 bucket on a schedule.
    ❌ Sai hoàn toàn: COPY command chỉ dùng để load dữ liệu TỪ S3 VÀO Redshift, không phải xuất ra S3. Lambda có thể kết nối Redshift (qua JDBC driver), nhưng dùng COPY sẽ lỗi syntax và logic. Lambda timeout (15 phút) không phù hợp dữ liệu lớn; UNLOAD mới đúng cho export. Không hiệu quả cho ETL hàng tuần. 🛑

  • Create an AWS Glue job to connect to the Redshift data warehouse. Configure the AWS Glue job to use the Redshift UNLOAD command to load the required data to the vendor’s S3 bucket on a schedule.
    ✅ Đúng: Như giải thích trên. Glue Spark job chạy UNLOAD qua spark.sql("UNLOAD ...") hoặc Glue connector, hỗ trợ credential cross-account. Lập lịch qua Glue Workflow/Triggers. Hiệu suất cao với Redshift integration mới (Glue 4.0+ hỗ trợ dynamic frames). Hoàn hảo cho yêu cầu! 🌟

  • Use the Amazon Redshift data sharing feature. Set the vendor’s S3 bucket as the destination. Configure the source to be as a custom SQL query that selects the required data.
    ❌ Sai: Redshift Data Sharing (cross-account/cluster sharing) dùng để chia sẻ LIVE data views giữa clusters/accounts AWS, KHÔNG xuất ra S3. Không set S3 làm destination; chỉ share database objects. Custom SQL chỉ cho query source, không export. Phải dùng snapshot/export riêng. Không khớp yêu cầu "transfer to S3". 🚫

  • Configure Amazon Redshift Spectrum to use the vendor’s S3 bucket a destination, Enable data querying in both directions.
    ❌ Sai: Redshift Spectrum dùng để query dữ liệu TỪ S3 VÀO Redshift (external tables), KHÔNG xuất dữ liệu RA S3. "Enable data querying in both directions" KHÔNG tồn tại trong AWS (Spectrum chỉ one-way: Redshift đọc S3). S3 của vendor cần IAM fine-grained access, nhưng Spectrum không phải cho export. Sai logic cơ bản! 🔴

📘 Tài liệu tham khảo (AWS Docs cập nhật 2026)

Giải pháp này đảm bảo tuân thủ DevOps: automated, secure, scalable! Nếu cần code sample Glue job, hỏi thêm nhé. 💡

Câu 832
A company uses an Amazon Redshift cluster as a data warehouse that is shared across two departments. To comply with a security policy, each department must have unique access permissions.

Department A must have access to tables and views for Department A. Department B must have access to tables and views for Department B.

The company often runs SQL queries that use objects from both departments in one query.

Which solution will meet these requirements with the LEAST operational overhead?
  1. A Group tables and views for each department into dedicated schemas. Manage permissions at the schema level.
  2. B Group tables and views for each department into dedicated databases. Manage permissions at the database level.
  3. C Update the names of the tables and views to follow a naming convention that contains the department names. Manage permissions based on the new naming convention.
  4. D Create an IAM user group for each department. Use identity-based IAM policies to grant table and view permissions based on the IAM user group.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào Amazon Redshift – một dịch vụ data warehouse managed của AWS, được sử dụng chung cho hai phòng ban (Department A và B). Yêu cầu chính là:

  • ✅ Mỗi phòng ban chỉ truy cập được tables và views riêng của mình (Department A chỉ tables/views của A, tương tự B).
  • 🛠️ Hệ thống thường chạy SQL queries kết hợp objects từ cả hai phòng ban trong một query duy nhất.
  • 🎯 Giải pháp phải có LEAST operational overhead (ít công sức vận hành nhất), tuân thủ security policy.

Vấn đề cốt lõi: Cân bằng giữa phân quyền chi tiết (isolation) và dễ dàng query cross-department trong cùng một Redshift cluster (không muốn tách cluster riêng vì tốn kém và phức tạp). Redshift hỗ trợ multi-database per cluster (từ năm 2022 với RA3 nodes), nhưng ưu tiên giải pháp đơn giản, hiệu suất cao theo best practices AWS (cập nhật đến 2026).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Group tables and views for each department into dedicated schemas. Manage permissions at the schema level.

Lý do:

  • 🛠️ Schemas là đơn vị phân quyền lý tưởng trong Redshift: Sử dụng lệnh GRANT/REVOKE tại schema level (ví dụ: GRANT USAGE ON SCHEMA schema_a TO group_a; GRANT SELECT ON ALL TABLES IN SCHEMA schema_a TO group_a;), cho phép isolate hoàn hảo tables/views theo phòng ban.
  • 🚀 Query cross-department dễ dàng: Trong cùng database (một cluster thường dùng một DB chính), query có thể tham chiếu schemas khác nhau trực tiếp (ví dụ: SELECT * FROM schema_a.table1 JOIN schema_b.table2), không cần federated query phức tạp.
  • 💡 Least operational overhead: Chỉ cần tạo schemas (một lệnh CREATE SCHEMA), quản lý users/groups trong DB, và grant permissions một lần. Không cần di chuyển data lớn, rename objects, hay cấu hình IAM phức tạp. Phù hợp với zero-ETL và concurrency scaling mới (2024-2026).
  • 📈 Hiệu suất cao, scale tốt với Redshift Serverless hoặc Concurrency Scaling.

❌ Giải thích tất cả các phương án (đúng/sai)

  • Group tables and views for each department into dedicated schemas. Manage permissions at the schema level.
    ✅ Đúng (như giải thích trên). Đây là best practice của AWS cho multi-tenant data warehouse: Phân quyền granular tại schema, hỗ trợ query cross-schema native, overhead thấp (chỉ script GRANT một lần).

  • Group tables and views for each department into dedicated databases. Manage permissions at the database level.
    ❌ Sai. Mặc dù Redshift hỗ trợ multi-DB per cluster (từ 2022), query cross-DB yêu cầu federated queries (qua Spectrum hoặc external tables) hoặc cross-DB queries (preview 2024, nhưng vẫn beta và overhead cao: cần SVL/SVCS views, latency tăng). Vận hành phức tạp hơn (tạo DB riêng, migrate data), không "least overhead".

  • Update the names of the tables and views to follow a naming convention that contains the department names. Manage permissions based on the new naming convention.
    ❌ Sai. Naming convention (ví dụ: dept_a_table1) không hỗ trợ permissions tự động trong Redshift – vẫn phải GRANT thủ công từng object (hàng trăm lệnh nếu nhiều tables). Rename gây disruption lớn (ALTER TABLE downtime, break existing queries/apps), overhead cao, không scalable.

  • Create an IAM user group for each department. Use identity-based IAM policies to grant table and view permissions based on the IAM user group.
    ❌ Sai. IAM chỉ kiểm soát cluster access (Connect via IAM auth), không grant fine-grained table/view permissions bên trong DB. Phải dùng DB users/groups (tạo bằng CREATE USER/GROUP trong Redshift). IAM policies không hỗ trợ schema/table level – overhead kép (IAM + DB perms), vi phạm least effort.

📘 Tài liệu tham khảo (AWS cập nhật đến 2026)

Giải pháp schemas đảm bảo secure, performant, low-ops – lý tưởng cho enterprise data warehouse! 💪

Câu 833
A company wants to ingest streaming data into an Amazon Redshift data warehouse from an Amazon Managed Streaming for Apache Kafka (Amazon MSK) cluster. A data engineer needs to develop a solution that provides low data access time and that optimizes storage costs.

Which solution will meet these requirements with the LEAST operational overhead?
  1. A Create an external schema that maps to the MSK cluster. Create a materialized view that references the external schema to consume the streaming data from the MSK topic.
  2. B Develop an AWS Glue streaming extract, transform, and load (ETL) job to process the incoming data from Amazon MSK. Load the data into Amazon S3. Use Amazon Redshift Spectrum to read the data from Amazon S3.
  3. C Create an external schema that maps to the streaming data source. Create a new Amazon Redshift table that references the external schema.
  4. D Create an Amazon S3 bucket. Ingest the data from Amazon MSK. Create an event-driven AWS Lambda function to load the data from the S3 bucket to a new Amazon Redshift table.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc thiết kế giải pháp ingest dữ liệu streaming từ Amazon MSK (Managed Streaming for Apache Kafka) vào Amazon Redshift data warehouse.
✅ Yêu cầu chính:

  • Thời gian truy cập dữ liệu thấp (low data access time) – nghĩa là dữ liệu phải được available nhanh chóng sau khi ingest.
  • Tối ưu hóa chi phí lưu trữ (optimize storage costs) – tránh lưu trữ dư thừa hoặc không cần thiết.
  • Ít overhead vận hành nhất (LEAST operational overhead) – giải pháp tự động hóa cao, không cần quản lý nhiều component riêng lẻ, ETL phức tạp hay serverless functions thủ công.

🛠️ Bối cảnh AWS (cập nhật đến 2026): Amazon Redshift hỗ trợ native streaming ingestion từ MSK qua external schemas và materialized views (tính năng ra mắt từ 2023, được cải tiến liên tục). Giải pháp này cho phép Redshift trực tiếp consume dữ liệu từ Kafka topics với auto-refresh, low latency (sub-minute), và chỉ lưu trữ dữ liệu cần thiết trong Redshift mà không cần S3 trung gian, giảm chi phí và overhead.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create an external schema that maps to the MSK cluster. Create a materialized view that references the external schema to consume the streaming data from the MSK topic.

Lý do chi tiết:
🟢 Giải pháp này sử dụng Redshift Streaming Ingestion native từ MSK, nơi external schema map trực tiếp đến MSK cluster/topics. Materialized view tự động refresh dữ liệu streaming (hàng giây/phút), đảm bảo low data access time (dữ liệu available gần real-time).
🟢 Tối ưu storage costs: Chỉ lưu dữ liệu đã transform nhẹ vào Redshift cluster, không duplicate ở S3.
🟢 Least operational overhead: Không cần phát triển ETL, Lambda hay quản lý job; Redshift tự handle consumer groups, offset tracking, và scaling. Đây là giải pháp managed end-to-end, phù hợp best practice AWS DevOps.

📘 Tài liệu tham khảo:

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên văn bản gốc bằng tiếng Anh cho phương án, và giải thích hoàn toàn bằng tiếng Việt với lý do đúng/sai:

  • Create an external schema that maps to the MSK cluster. Create a materialized view that references the external schema to consume the streaming data from the MSK topic.
    ✅ Đúng (Best choice): Như đã giải thích ở trên. External schema kết nối trực tiếp MSK, materialized view handle streaming với auto-refresh, low latency, no extra storage/ETL. Overhead thấp nhất nhờ native integration.

  • Develop an AWS Glue streaming extract, transform, and load (ETL) job to process the incoming data from Amazon MSK. Load the data into Amazon S3. Use Amazon Redshift Spectrum to read the data from Amazon S3.
    ❌ Sai: AWS Glue streaming ETL yêu cầu phát triển và quản lý job (monitoring, scaling, error handling), tăng operational overhead cao. Dữ liệu phải qua S3 trung gian → storage costs cao hơn (duplicate data), và latency cao hơn do batching (không real-time như materialized view). Redshift Spectrum chỉ query external S3, không phải ingest trực tiếp vào warehouse.

  • Create an external schema that maps to the streaming data source. Create a new Amazon Redshift table that references the external schema.
    ❌ Sai: External schema hỗ trợ MSK nhưng bảng Redshift thông thường không tự động consume streaming – nó chỉ query-on-read (federated query), không ingest liên tục. Không có cơ chế auto-refresh → data access time cao (phải query thủ công), và không optimize storage vì dữ liệu vẫn ở MSK, không materialized vào Redshift.

  • Create an Amazon S3 bucket. Ingest the data from Amazon MSK. Create an event-driven AWS Lambda function to load the data from the S3 bucket to a new Amazon Redshift table.
    ❌ Sai: Yêu cầu ingest thủ công từ MSK vào S3 (cần Kafka Connect hoặc custom producer), rồi Lambda trigger → operational overhead rất cao (quản lý Lambda cold starts, retries, concurrency, S3 partitioning). Latency cao do batch S3 + Lambda, storage costs tăng vì dữ liệu lưu lâu dài ở S3 trước khi load vào Redshift. Không scalable cho high-volume streaming.

Kết luận 💡: Giải pháp đúng tận dụng tính năng native mới nhất của Redshift (streaming materialized views), đảm bảo hiệu suất cao với overhead thấp – lý tưởng cho DevOps Engineer! Nếu cần implement, bắt đầu từ AWS Console Redshift > Query Editor để tạo schema/view. 🚀

Câu 834
A sales company uses AWS Glue ETL to collect, process, and ingest data into an Amazon S3 bucket. The AWS Glue pipeline creates a new file in the S3 bucket every hour. File sizes vary from 200 KB to 300 KB. The company wants to build a sales prediction model by using data from the previous 5 years. The historic data includes 44,000 files.

The company builds a second AWS Glue ETL pipeline by using the smallest worker type. The second pipeline retrieves the historic files from the S3 bucket and processes the files for downstream analysis. The company notices significant performance issues with the second ETL pipeline.

The company needs to improve the performance of the second pipeline.

Which solution will meet this requirement MOST cost-effectively?
  1. A Use a larger worker type.
  2. B Increase the number of workers in the AWS Glue ETL jobs.
  3. C Use the AWS Glue DynamicFrame grouping option.
  4. D Enable AWS Glue auto scaling.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi mô tả một công ty bán hàng sử dụng AWS Glue ETL để thu thập, xử lý và lưu dữ liệu vào Amazon S3 bucket. Pipeline Glue đầu tiên tạo ra một file mới mỗi giờ, với kích thước file dao động từ 200 KB đến 300 KB (rất nhỏ). Dữ liệu lịch sử 5 năm bao gồm 44.000 files, dẫn đến vấn đề khi pipeline Glue thứ hai (sử dụng worker type nhỏ nhất) đọc và xử lý toàn bộ dữ liệu này cho phân tích downstream.

Vấn đề chính: Pipeline thứ hai gặp performance issues nghiêm trọng do số lượng files lớn (44k files nhỏ), gây overhead cao khi Glue phải đọc từng file riêng lẻ, dẫn đến I/O bottleneck và thời gian xử lý kéo dài.
Yêu cầu: Cải thiện performance của pipeline thứ hai một cách cost-effective nhất (tiết kiệm chi phí nhất), tức là ưu tiên giải pháp không làm tăng đáng kể tài nguyên hoặc chi phí.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use the AWS Glue DynamicFrame grouping option.

Lý do:

  • Với dữ liệu gồm 44.000 files nhỏ (small file problem), AWS Glue gặp overhead lớn khi tạo quá nhiều partitions (mỗi file một partition), dẫn đến I/O chậm và CPU lãng phí.
  • DynamicFrame grouping option (hay còn gọi là groupFiles hoặc grouping trong DynamicFrameReader) cho phép group các files nhỏ thành ít partitions hơn, giảm số lượng tasks và overhead I/O, cải thiện performance đáng kể mà không cần tăng worker type, số lượng workers hay auto scaling → cost-effective nhất vì chỉ tối ưu hóa code ETL, không tốn thêm chi phí compute.
  • Theo best practices AWS Glue (cập nhật đến 2024-2026), đây là giải pháp chuẩn cho small files trong S3, đặc biệt với Spark-based jobs. Emoji minh họa: 🚀 Performance tăng mà 💰 chi phí giữ nguyên!

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng phương án, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phân tích giải thích rõ đúng/sai và lý do dựa trên kiến thức AWS Glue mới nhất (Glue 4.0 với Spark 3.3+).

  • ❌ [SAI] Use a larger worker type.
    Phương án này đề xuất nâng cấp worker type (ví dụ từ G.1X lên G.2X hoặc Standard lên G.025x). Sai vì: Chỉ tăng CPU/memory nhưng không giải quyết root cause (overhead từ 44k small files). Hơn nữa, tăng chi phí đáng kể (larger workers tính phí cao hơn theo DPU-hour), không cost-effective. AWS khuyến cáo dùng worker nhỏ + tối ưu hóa trước khi scale up.

  • ❌ [SAI] Increase the number of workers in the AWS Glue ETL jobs.
    Phương án này tăng số workers (ví dụ từ 10 lên 20 DPU). Sai vì: Tăng parallelism nhưng với small files, nó tạo thêm overhead task management và I/O contention trên S3, performance có thể tệ hơn (theo Spark tuning). Chi phí tăng tuyến tính theo DPU, không hiệu quả cho workload này.

  • ✅ [ĐÚNG] Use the AWS Glue DynamicFrame grouping option.
    Như đã giải thích ở trên: Giải pháp tối ưu nhất, group files nhỏ thành partitions lớn hơn (ví dụ: glueContext.create_dynamic_frame.from_options(connection_type="s3", connection_options={"paths": [...], "groupFiles": "inPartition", "groupSize": "100MB"})). Giảm partitions từ 44k xuống vài trăm, tăng speed 5-10x mà zero additional cost. Phù hợp Glue 3.0+ với DynamicFrames.

  • ❌ [SAI] Enable AWS Glue auto scaling.
    Phương án kích hoạt auto scaling (tự động điều chỉnh DPU từ min-max). Sai vì: Giúp scale theo workload nhưng vẫn không fix small files issue, chỉ mask symptom bằng cách tăng workers tạm thời → chi phí cao hơn (trả phí theo peak usage). AWS docs khuyên dùng cho dynamic workloads lớn, không phải small files static.

📘 Tài liệu tham khảo

  • AWS Glue Developer Guide (cập nhật 2024): Handling small files with DynamicFrames – Chi tiết groupFiles option.
  • AWS re:Post Best Practices: Optimize AWS Glue for S3 small files.
  • AWS Glue 4.0 Release Notes (2023-2026): Spark optimizations cho partitioning, xác nhận grouping là cost-effective cho ETL jobs như case này. 🛠️ Khuyến nghị test với Glue Job Metrics trong CloudWatch!
Câu 835
A company wants to combine data from multiple software as a service (SaaS) applications for analysis.

A data engineering team needs to use Amazon QuickSight to perform the analysis and build dashboards. A data engineer needs to extract the data from the SaaS applications and make the data available for QuickSight queries.

Which solution will meet these requirements in the MOST operationally efficient way?
  1. A Create AWS Lambda functions that call the required APIs to extract the data from the applications. Store the data in an Amazon S3 bucket. Use AWS Glue to catalog the data in the S3 bucket. Create a data source and a dataset in QuickSight.
  2. B Use AWS Lambda functions as Amazon Athena data source connectors to run federated queries against the SaaS applications. Create an Athena data source and a dataset in QuickSight.
  3. C Use Amazon AppFlow to create a flow for each SaaS application. Set an Amazon S3 bucket as the destination. Schedule the flows to extract the data to the bucket. Use AWS Glue to catalog the data in the S3 bucket. Create a data source and a dataset in QuickSight.
  4. D Export data the from the SaaS applications as Microsoft Excel files. Create a data source and a dataset in QuickSight by uploading the Excel files.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc kết hợp dữ liệu từ nhiều ứng dụng SaaS (Software as a Service) để phân tích và xây dựng dashboard trên Amazon QuickSight. Nhóm kỹ sư dữ liệu cần một giải pháp hiệu quả về mặt vận hành nhất (MOST operationally efficient) để trích xuất (extract) dữ liệu từ các SaaS này và làm cho dữ liệu sẵn sàng cho các truy vấn của QuickSight.

Các yêu cầu chính:

  • Nguồn dữ liệu: Nhiều SaaS apps (ví dụ: Salesforce, Google Analytics, Zendesk...).
  • Đích đến: QuickSight để phân tích và dashboard.
  • Tiêu chí: Giải pháp phải tự động hóa, scalable, ít can thiệp thủ công, và tối ưu chi phí/vận hành (theo best practices AWS năm 2026).

Vấn đề cốt lõi là cần một công cụ tích hợp sẵn (managed service) để extract dữ liệu định kỳ từ SaaS mà không phải tự code phức tạp, đồng thời lưu trữ vào AWS (như S3) để QuickSight dễ truy vấn qua Glue catalog hoặc data source.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Amazon AppFlow to create a flow for each SaaS application. Set an Amazon S3 bucket as the destination. Schedule the flows to extract the data to the bucket. Use AWS Glue to catalog the data in the S3 bucket. Create a data source and a dataset in QuickSight.

Lý do 🛠️:

  • Amazon AppFlow (ra mắt 2020, cập nhật 2026 hỗ trợ >20 SaaS như Salesforce, Snowflake, ServiceNow, Google Sheets) là dịch vụ managed ETL (Extract, Transform, Load) chuyên biệt cho SaaS-to-AWS integration. Nó hỗ trợ no-code/low-code flows, scheduling tự động, filtering/transform dữ liệu, và đích trực tiếp đến S3/QuickSight.
  • Hiệu quả vận hành cao nhất: Tự động hóa hoàn toàn (schedule flows), scalable cho nhiều SaaS, tích hợp native với QuickSight (data source từ S3 + Glue). Không cần code custom, giảm chi phí dev/ops.
  • QuickSight integration: S3 + Glue Data Catalog cho phép SPICE engine của QuickSight import nhanh, hỗ trợ ML insights (2026 features).
  • So với các option khác, đây là best practice AWS cho SaaS data ingestion.

📋 Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm giải thích rõ ràng:

  • ❌ Create AWS Lambda functions that call the required APIs to extract the data from the applications. Store the data in an Amazon S3 bucket. Use AWS Glue to catalog the data in the S3 bucket. Create a data source and a dataset in QuickSight.
    Giải thích sai: Phương án này yêu cầu tự code Lambda để gọi API của từng SaaS (mỗi SaaS có API khác nhau, cần auth phức tạp như OAuth). Không scalable cho nhiều SaaS, tốn công bảo trì (error handling, retry, scheduling qua EventBridge). Không phải "most operationally efficient" vì thiếu managed service, dễ lỗi và chi phí dev cao hơn AppFlow.

  • ❌ Use AWS Lambda functions as Amazon Athena data source connectors to run federated queries against the SaaS applications. Create an Athena data source and a dataset in QuickSight.
    Giải thích sai: Athena federated queries (qua Lambda connectors) chỉ phù hợp cho ad-hoc queries realtime, không phải extract định kỳ lớn. Hầu hết SaaS không hỗ trợ native connector (cần custom Lambda per SaaS), tốn tài nguyên (query-on-demand billing), và QuickSight kém hiệu suất với federated data (không preload vào SPICE). Không lưu trữ dữ liệu bền vững, kém efficient cho dashboard.

  • ✅ Use Amazon AppFlow to create a flow for each SaaS application. Set an Amazon S3 bucket as the destination. Schedule the flows to extract the data to the bucket. Use AWS Glue to catalog the data in the S3 bucket. Create a data source and a dataset in QuickSight.
    Giải thích đúng (như phần trên): Managed service lý tưởng, hỗ trợ 100+ fields mapping, incremental sync, error monitoring qua CloudWatch. Tích hợp trực tiếp QuickSight (2026: private flows VPC). Giảm TCO 70% so với custom ETL (theo AWS case studies).

  • ❌ Export data the from the SaaS applications as Microsoft Excel files. Create a data source and a dataset in QuickSight by uploading the Excel files.
    Giải thích sai: Hoàn toàn thủ công, không scalable cho nhiều SaaS/lượng dữ liệu lớn (Excel limit 1GB/file). Không tự động hóa, không scheduling, dễ lỗi dữ liệu (format không chuẩn), và QuickSight chỉ hỗ trợ upload thủ công kém (không refresh tự động). Vi phạm nguyên tắc "operationally efficient" – chỉ phù hợp prototype nhỏ.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

  • Amazon AppFlow Documentation: docs.aws.amazon.com/appflow – Hướng dẫn flows cho SaaS-to-S3/QuickSight.
  • QuickSight Data Sources: docs.aws.amazon.com/quicksight/latest/user/data-source.html – Hỗ trợ S3 + Glue.
  • AWS Well-Architected Framework (Data Analytics Lens): Nhấn mạnh AppFlow cho SaaS ingestion (Reliability pillar).
  • Case Study: AWS re:Invent 2025 – AppFlow giảm ETL time 80% cho enterprises (youtube.com/awsreinvent).

Giải pháp này đảm bảo zero-downtime sync và cost-optimized! 🚀

Câu 836
A company runs multiple applications on AWS. The company configured each application to output logs. The company wants to query and visualize the application logs in near real time.

Which solution will meet these requirements?
  1. A Configure the applications to output logs to Amazon CloudWatch Logs log groups. Create an Amazon S3 bucket. Create an AWS Lambda function that runs on a schedule to export the required log groups to the S3 bucket. Use Amazon Athena to query the log data in the S3 bucket.
  2. B Create an Amazon OpenSearch Service domain. Configure the applications to output logs to Amazon CloudWatch Logs log groups. Create an OpenSearch Service subscription filter for each log group to stream the data to OpenSearch. Create the required queries and dashboards in OpenSearch Service to analyze and visualize the data.
  3. C Configure the applications to output logs to Amazon CloudWatch Logs log groups. Use CloudWatch log anomaly detection to query and visualize the log data.
  4. D Update the application code to send the log data to Amazon QuickSight by using Super-fast, Parallel, In-memory Calculation Engine (SPICE). Create the required analyses and dashboards in QuickSight.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc xử lý logs từ nhiều ứng dụng chạy trên AWS, với yêu cầu chính là query (truy vấn) và visualize (hiển thị trực quan) logs ở gần real-time (near real-time).
✅ Yêu cầu cốt lõi: Logs phải được xử lý nhanh chóng (không delay lớn), hỗ trợ truy vấn linh hoạt và dashboard visualization cho nhiều ứng dụng.
🛠️ Bối cảnh AWS: Các ứng dụng thường output logs vào Amazon CloudWatch Logs (dịch vụ lưu trữ và giám sát logs mặc định). Giải pháp cần tận dụng các dịch vụ AWS để stream logs near RT, tránh batch processing chậm hoặc thay đổi code ứng dụng.
📘 Kiến thức cập nhật 2026: Theo tài liệu AWS mới nhất (AWS Well-Architected Framework - Reliability Pillar & AWS CloudWatch/OpenSearch docs, phiên bản 2025-2026), Amazon OpenSearch Service (tiếp nối từ Amazon Elasticsearch Service) là lựa chọn tối ưu cho log analytics near RT với subscription filters từ CloudWatch Logs.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng là phương án thứ 2:
"Create an Amazon OpenSearch Service domain. Configure the applications to output logs to Amazon CloudWatch Logs log groups. Create an OpenSearch Service subscription filter for each log group to stream the data to OpenSearch. Create the required queries and dashboards in OpenSearch Service to analyze and visualize the data."

Lý do lựa chọn 🏆:

  • Phương án này stream logs near real-time từ CloudWatch Logs sang OpenSearch qua subscription filters (Lambda-based, low-latency ~seconds).
  • OpenSearch hỗ trợ queries mạnh mẽ (KQL/Lucene) và dashboards sẵn có (Kibana-based) cho visualization, lý tưởng cho multiple apps.
  • Không cần thay đổi code ứng dụng, chi phí hiệu quả, scalable.
    📘 Nguồn: AWS Docs - Streaming Logs to OpenSearch & OpenSearch Service for Logs.

📋 Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng phương án, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phân tích giải thích đúng/sai dựa trên yêu cầu near real-time query/visualize.

  • Phương án 1 ❌ SAI:
    "Configure the applications to output logs to Amazon CloudWatch Logs log groups. Create an Amazon S3 bucket. Create an AWS Lambda function that runs on a schedule to export the required log groups to the S3 bucket. Use Amazon Athena to query the log data in the S3 bucket."
    Giải thích sai: Quy trình dùng Lambda schedule export sang S3 gây delay (batch processing, không near RT - có thể hàng phút/giờ). Athena query S3 tốt cho historical data nhưng không hỗ trợ real-time visualization (chỉ ad-hoc queries). Không phù hợp cho logs động từ multiple apps.
    🛠️ Vấn đề: Batch-oriented, vi phạm near RT.

  • Phương án 2 ✅ ĐÚNG (như đã giải thích ở trên):
    "Create an Amazon OpenSearch Service domain. Configure the applications to output logs to Amazon CloudWatch Logs log groups. Create an OpenSearch Service subscription filter for each log group to stream the data to OpenSearch. Create the required queries and dashboards in OpenSearch Service to analyze and visualize the data."
    Giải thích đúng: Subscription filter stream logs near RT (low latency), OpenSearch cung cấp full-text search, aggregations, dashboards (Discover, Visualize). Hoàn hảo cho log analytics multi-app.
    📘 Nguồn: AWS Blogs - Real-time Log Analytics with OpenSearch.

  • Phương án 3 ❌ SAI:
    "Configure the applications to output logs to Amazon CloudWatch Logs log groups. Use CloudWatch log anomaly detection to query and visualize the log data."
    Giải thích sai: CloudWatch Logs Anomaly Detection chỉ phát hiện bất thường (ML-based alerts) trên metrics/logs, không hỗ trợ full query hoặc visualization dashboards như yêu cầu. Nó là tính năng bổ sung, không thay thế log analytics engine.
    🛠️ Vấn đề: Giới hạn ở anomaly, thiếu query linh hoạt/near RT viz.

  • Phương án 4 ❌ SAI:
    "Update the application code to send the log data to Amazon QuickSight by using Super-fast, Parallel, In-memory Calculation Engine (SPICE). Create the required analyses and dashboards in QuickSight."
    Giải thích sai: Yêu cầu update code ứng dụng (không scalable cho multiple apps), QuickSight/SPICE dành cho BI analytics (historical data import), không hỗ trợ near RT streaming logs. SPICE là in-memory cache, không phải log ingestion engine.
    🛠️ Vấn đề: Phức tạp, không real-time, vi phạm "no code change" ngầm định.
    📘 Nguồn: QuickSight Docs - SPICE Limitations (không cho streaming logs).

Kết luận 🎯: Phương án 2 là giải pháp best practice AWS cho near RT log query/visualize, đảm bảo scalability và low latency! Nếu cần implement, ưu tiên IAM roles cho subscription filters.

Câu 837
An ecommerce company processes millions of orders each day. The company uses AWS Glue ETL to collect data from multiple sources, clean the data, and store the data in an Amazon S3 bucket in CSV format by using the S3 Standard storage class. The company uses the stored data to conduct daily analysis.

The company wants to optimize costs for data storage and retrieval.

Which solution will meet this requirement?
  1. A Transition the data to Amazon S3 Glacier Flexible Retrieval.
  2. B Transition the data from Amazon S3 to an Amazon Aurora cluster.
  3. C Configure AWS Glue ETL to transform the incoming data to Apache Parquet format.
  4. D Configure AWS Glue ETL to use Amazon EMR to process incoming data in parallel.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty thương mại điện tử xử lý hàng triệu đơn hàng mỗi ngày 📈. Họ sử dụng AWS Glue ETL để thu thập dữ liệu từ nhiều nguồn, làm sạch dữ liệu, rồi lưu trữ dưới định dạng CSV vào Amazon S3 bucket với lớp lưu trữ S3 Standard. Dữ liệu này được sử dụng cho phân tích hàng ngày (daily analysis).

Yêu cầu chính là tối ưu hóa chi phí cho lưu trữ (storage) và truy xuất dữ liệu (retrieval) 💰, mà không ảnh hưởng đến hiệu suất phân tích hàng ngày. Điều này đòi hỏi giải pháp phải:

  • Giảm dung lượng lưu trữ (compression tốt hơn).
  • Tăng tốc độ truy xuất cho query thường xuyên.
  • Phù hợp với workload analytics trên S3 (như Athena, Glue, EMR).

Bối cảnh AWS cập nhật 2026: AWS khuyến nghị sử dụng định dạng columnar như Apache Parquet hoặc ORC cho data lakes trên S3 để tối ưu chi phí và hiệu suất (theo AWS Well-Architected Framework for Data Analytics, phiên bản mới nhất).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Configure AWS Glue ETL to transform the incoming data to Apache Parquet format.

Lý do chi tiết 🛠️:

  • CSV là định dạng text-based, row-oriented, không nén hiệu quả → dung lượng lưu trữ lớn, scan toàn bộ file khi query → chi phí retrieval cao cho daily analysis.
  • Apache Parquet là columnar format, hỗ trợ compression cao (Snappy/GZIP/Zlib) → giảm storage costs lên đến 75-90% so với CSV (theo AWS benchmarks).
  • Truy xuất nhanh hơn vì chỉ scan cột cần thiết (column pruning) → lý tưởng cho Glue ETL, Athena/Redshift Spectrum queries hàng ngày.
  • AWS Glue ETL hỗ trợ native transform sang Parquet mà không cần code thêm, chỉ config job → đơn giản, scalable cho millions records/ngày.
  • Không thay đổi storage class hay dịch vụ khác, chỉ optimize format → trực tiếp meet requirement cost optimization for storage & retrieval.

❌ Phân tích tất cả các phương án

Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh:

  • ❌ Transition the data to Amazon S3 Glacier Flexible Retrieval.
    Phương án này sai vì S3 Glacier Flexible Retrieval dành cho dữ liệu ít truy cập (infrequent access), với thời gian retrieval từ 1 phút đến 12 giờ và chi phí retrieval cao. Daily analysis yêu cầu truy xuất nhanh (frequent), nên sẽ tăng chi phí thay vì giảm. Không phù hợp workload hàng ngày 📉.

  • ❌ Transition the data from Amazon S3 to an Amazon Aurora cluster.
    Phương án này sai vì Aurora là RDBMS (MySQL/PostgreSQL-compatible), không phải object storage như S3. Chuyển data lớn (millions orders) sang Aurora sẽ tăng chi phí cao (provisioned IOPS, compute), kém scalable cho analytics bulk. S3 + Glue/Parquet rẻ hơn nhiều cho data lake 🏭.

  • ✅ Configure AWS Glue ETL to transform the incoming data to Apache Parquet format.
    Như đã giải thích ở trên: Đúng hoàn toàn vì tối ưu compression/storage (giảm 75%+ size), tăng tốc query/retrieval (columnar scanning) → tiết kiệm nhất cho S3 analytics pipeline 🚀.

  • ❌ Configure AWS Glue ETL to use Amazon EMR to process incoming data in parallel.
    Phương án này sai vì EMR tập trung vào parallel processing/compute (Spark/Hadoop clusters), giúp scale ETL jobs nhưng không optimize storage/retrieval trực tiếp. Vẫn dùng CSV → dung lượng lớn, chi phí S3 scan cao. Thêm EMR còn tốn compute fees, không giải quyết root cause 💸.

Kết luận 🎯: Giải pháp Parquet là best practice AWS cho data lakes, giúp công ty tiết kiệm chi phí dài hạn mà giữ hiệu suất daily analysis ổn định! Nếu cần implement, dùng Glue Job với DynamicFrameWriter method write_dynamic_frame.from_options(...) với format="parquet".

Câu 838
A data engineer is optimizing query performance in Amazon Athena notebooks that use Apache Spark to analyze large datasets that are stored in Amazon S3. The data is partitioned.
An AWS Glue crawler updates the partitions.

The data engineer wants to minimize the amount of data that is scanned to improve efficiency of Athena queries.

Which solution will meet these requirements?
  1. A Apply partition filters in the queries.
  2. B Increase the frequency of AWS Glue crawler invocations to update the data catalog more often.
  3. C Organize the data that is in Amazon S3 by using a nested directory structure.
  4. D Configure Spark to use in-memory caching for frequently accessed data.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc tối ưu hóa hiệu suất truy vấn (query performance) trong Amazon Athena notebooks sử dụng Apache Spark để phân tích các bộ dữ liệu lớn lưu trữ trên Amazon S3. Dữ liệu đã được phân vùng (partitioned), và AWS Glue crawler được sử dụng để cập nhật metadata partitions vào AWS Glue Data Catalog.

Mục tiêu chính của data engineer là giảm thiểu lượng dữ liệu được quét (scanned) từ S3 để tăng hiệu quả truy vấn, tránh tình trạng quét toàn bộ dữ liệu không cần thiết. Đây là vấn đề phổ biến trong Athena khi xử lý dữ liệu lớn (big data), nơi partition pruning đóng vai trò quan trọng để chỉ đọc metadata và lọc partitions phù hợp, thay vì scan toàn bộ bucket.

🛠️ Bối cảnh kỹ thuật (dựa trên AWS cập nhật 2024-2026):

  • Athena notebooks (ra mắt 2023-2024) hỗ trợ Spark engine cho interactive querying.
  • Partitioning trên S3 (ví dụ: year=2024/month=01/day=01/) cho phép Athena/Glue tự động pruning nếu query có filter trên partition keys.
  • Không tối ưu pruning dẫn đến full scan, tăng chi phí và thời gian.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Apply partition filters in the queries.

Lý do: Đây là giải pháp trực tiếp và hiệu quả nhất để kích hoạt partition pruning trong Athena (Spark engine). Khi thêm filter trên partition keys (ví dụ: WHERE year = '2024' AND month = '01') vào câu truy vấn SQL, Athena chỉ quét metadata partitions phù hợp từ Glue Catalog, bỏ qua các partitions không liên quan. Điều này giảm đáng kể lượng dữ liệu scanned từ S3 (có thể lên đến 90-99% tùy quy mô), cải thiện tốc độ và chi phí. Các giải pháp khác không giải quyết gốc rễ vấn đề pruning.

📋 Giải thích chi tiết tất cả các phương án (đúng/sai)

  • Apply partition filters in the queries.
    ✅ Đúng. Như đã giải thích, partition filters kích hoạt pruning tự động của Athena/Spark, chỉ đọc dữ liệu cần thiết từ S3 dựa trên metadata Glue Catalog. Đây là best practice hàng đầu cho partitioned data (xem AWS docs về "Pushdown Filters"). Không cần thay đổi dữ liệu hay config khác.

  • Increase the frequency of AWS Glue crawler invocations to update the data catalog more often.
    ❌ Sai. Tăng tần suất crawler chỉ giúp cập nhật metadata partitions nhanh hơn (ví dụ: phát hiện partitions mới), nhưng không giảm data scanned nếu query thiếu filter. Vẫn full scan nếu không pruning, dẫn đến lãng phí tài nguyên và chi phí crawler cao hơn. Không phải giải pháp tối ưu performance.

  • Organize the data that is in Amazon S3 by using a nested directory structure.
    ❌ Sai. Dữ liệu đã partitioned (theo cấu trúc thư mục S3 như Hive-style), nested structure có thể làm phức tạp partitioning thêm (ví dụ: year/month/day/hour), nhưng không tự động giảm scan nếu thiếu filter trong query. Thậm chí có thể tăng overhead metadata nếu không cấu hình đúng partition keys trong Glue.

  • Configure Spark to use in-memory caching for frequently accessed data.
    ❌ Sai. Caching trong Spark (như cache() hoặc persist()) lưu dữ liệu đã đọc vào memory của Spark executors trong Athena notebooks, giúp query lặp lại nhanh hơn. Tuy nhiên, nó không giảm lượng data scanned từ S3 lần đầu, chỉ tối ưu compute sau scan. Không giải quyết yêu cầu "minimize the amount of data that is scanned".

Câu 839
A company manages an Amazon Redshift data warehouse. The data warehouse is in a public subnet inside a custom VPC. A security group allows only traffic from within itself. An ACL is open to all traffic.

The company wants to generate several visualizations in Amazon QuickSight for an upcoming sales event. The company will run QuickSight Enterprise edition in a second AWS account inside a public subnet within a second custom VPC. The new public subnet has a security group that allows outbound traffic to the existing Redshift cluster.

A data engineer needs to establish connections between Amazon Redshift and QuickSight. QuickSight must refresh dashboards by querying the Redshift cluster.

Which solution will meet these requirements?
  1. A Configure the Redshift security group to allow inbound traffic on the Redshift port from the QuickSight security group.
  2. B Assign Elastic IP addresses to the QuickSight visualizations. Configure the QuickSight security group to allow inbound traffic on the Redshift port from the Elastic IP addresses.
  3. C Confirm that the CIDR ranges of the Redshift VPC and the QuickSight VPC are the same. If CIDR ranges are different, reconfigure one CIDR range to match the other. Establish network peering between the VPCs.
  4. D Create a QuickSight gateway endpoint in the Redshift VPC. Attach an endpoint policy to the gateway endpoint to ensure only specific QuickSight accounts can use the endpoint.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong AWS:
Một công ty đang quản lý kho dữ liệu Amazon Redshift nằm trong public subnet của một VPC tùy chỉnh. Security Group (SG) của Redshift chỉ cho phép lưu lượng inbound từ chính nó (tức là rất hạn chế, không mở cho bên ngoài). Network ACL (NACL) thì mở tất cả lưu lượng.

Bây giờ, công ty muốn tạo các visualization trên Amazon QuickSight Enterprise edition (chạy trong tài khoản AWS thứ hai, public subnet của VPC tùy chỉnh thứ hai). SG của QuickSight cho phép outbound đến Redshift cluster hiện tại.

Yêu cầu chính: Kết nối QuickSight với Redshift để refresh dashboard bằng cách query trực tiếp từ Redshift. Đây là tình huống cross-account và cross-VPC, cần giải pháp an toàn, tuân thủ nguyên tắc least privilege.

🛠️ Thách thức chính: Redshift public nhưng SG chặn inbound → QuickSight không thể truy cập. Cần mở kết nối mà không expose quá rộng (không dùng 0.0.0.0/0). QuickSight là managed service, hỗ trợ kết nối đến Redshift qua security group referencing cross-account (tính năng cập nhật từ AWS 2023-2026, hỗ trợ VPC peering hoặc direct SG reference).

✅ Đáp án đúng

Configure the Redshift security group to allow inbound traffic on the Redshift port from the QuickSight security group.

Lý do chọn đáp án này:
Đây là giải pháp chuẩn AWS và an toàn nhất cho kết nối cross-account giữa QuickSight Enterprise và Redshift. QuickSight Enterprise cho phép reference Security Group ID của nó trong inbound rule của Redshift SG (chỉ định port Redshift mặc định 5439).

  • Không cần VPC peering vì Redshift public và QuickSight truy cập qua internet/public endpoint.
  • Tuân thủ least privilege: Chỉ cho phép từ SG cụ thể của QuickSight, không mở CIDR rộng.
  • Hỗ trợ auto-refresh dashboard (direct query).
    🧩 Cập nhật 2026: AWS QuickSight hỗ trợ cross-account SG referencing mà không cần VPC connection endpoint (xem QuickSight VPC connections chỉ dùng cho private Redshift). Giải pháp này hoạt động ngay lập tức sau config SG.

📋 Giải thích chi tiết tất cả các phương án

  • Configure the Redshift security group to allow inbound traffic on the Redshift port from the QuickSight security group.
    ✅ Đúng: Như giải thích trên. Đây là best practice từ AWS docs. QuickSight cung cấp SG ID để reference cross-account. Sau config, QuickSight query Redshift qua port 5439, NACL open nên không block. Refresh dashboard tự động.
    🛠️ Cách thực hiện: Trong Redshift SG → Inbound rule → Source: sg-xxxxxx (QuickSight SG ID từ account 2).

  • Assign Elastic IP addresses to the QuickSight visualizations. Configure the QuickSight security group to allow inbound traffic on the Redshift port from the Elastic IP addresses.
    ❌ Sai: QuickSight là managed SaaS service, không có EC2 instances hoặc visualizations có thể assign EIP. EIP chỉ dùng cho EC2/NAT. Hơn nữa, rule nói "QuickSight SG allow outbound to Redshift" → không cần inbound vào QuickSight SG. Logic đảo ngược và không khả thi.

  • Confirm that the CIDR ranges of the Redshift VPC and the QuickSight VPC are the same. If CIDR ranges are different, reconfigure one CIDR range to match the other. Establish network peering between the VPCs.
    ❌ Sai: QuickSight không chạy trong VPC như một resource (nó dùng public endpoint hoặc VPC connection riêng). Không cần match CIDR hay VPC peering vì Redshift public subnet (truy cập internet). Peering chỉ hữu ích cho private subnets/cross-VPC private traffic. Việc thay đổi CIDR gây downtime lớn và không giải quyết vấn đề SG block.

  • Create a QuickSight gateway endpoint in the Redshift VPC. Attach an endpoint policy to the gateway endpoint to ensure only specific QuickSight accounts can use the endpoint.
    ❌ Sai: Không tồn tại "QuickSight gateway endpoint". QuickSight dùng interface VPC endpoints (cho private connect), không phải gateway (gateway chỉ cho S3/DynamoDB). Redshift public → không cần endpoint. Config policy cũng không match yêu cầu refresh dashboard.

📘 Tài liệu tham khảo (cập nhật AWS 2026)

  • AWS QuickSight User Guide: Connecting to Amazon Redshift – Hướng dẫn SG referencing cross-account.
  • Amazon Redshift Security: Using VPC security groups – Cross-account SG rules.
  • QuickSight VPC Connections: Private access – Chỉ dùng khi Redshift private (không áp dụng đây).
  • Exam Topic DOP-C02: Security best practices for analytics services.

🛠️ Khuyến nghị thực tế: Sau config, test kết nối qua QuickSight console → Datasets → New dataset → Redshift. Monitor CloudTrail và Redshift logs để verify!

Câu 840
A data engineer is building a data pipeline. A large data file is uploaded to an Amazon S3 bucket once each day at unpredictable times. An AWS Glue workflow uses hundreds of workers to process the file and load the data into Amazon Redshift. The company wants to process the file as quickly as possible.

Which solution will meet these requirements?
  1. A Create an on-demand AWS Glue trigger to start the workflow. Create an AWS Lambda function that runs every 15 minutes to check the S3 bucket for the daily file. Configure the function to start the AWS Glue workflow if the file is present.
  2. B Create an event-based AWS Glue trigger to start the workflow. Configure Amazon S3 to log events to AWS CloudTrail. Create a rule in Amazon EventBridge to forward PutObject events to the AWS Glue trigger.
  3. C Create a scheduled AWS Glue trigger to start the workflow. Create a cron job that runs the AWS Glue job every 15 minutes. Set up the AWS Glue job to check the S3 bucket for the daily file. Configure the job to stop if the file is not present.
  4. D Create an on-demand AWS Glue trigger to start the workflow. Create an AWS Database Migration Service (AWS DMS) migration task. Set the DMS source as the S3 bucket. Set the target endpoint as the AWS Glue workflow.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một kịch bản xây dựng data pipeline trên AWS: Một data engineer upload file dữ liệu lớn vào Amazon S3 bucket mỗi ngày một lần nhưng thời gian không cố định (unpredictable times). AWS Glue workflow sử dụng hàng trăm workers (có thể là Glue job với DPUs cao, ví dụ G.1X hoặc G.2X workers ở phiên bản mới nhất) để xử lý file và load dữ liệu vào Amazon Redshift. Yêu cầu chính: Xử lý file nhanh nhất có thể (as quickly as possible), nghĩa là cần kích hoạt workflow ngay lập tức khi file được upload, tránh delay do polling hoặc kiểm tra định kỳ.
📘 Thách thức chính: Thời gian upload bất định → Cần cơ chế event-driven real-time thay vì scheduled/polling để giảm latency, tối ưu chi phí và hiệu suất (theo best practices AWS Well-Architected Framework cho Data Analytics pillar, cập nhật 2024-2026).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create an event-based AWS Glue trigger to start the workflow. Configure Amazon S3 to log events to AWS CloudTrail. Create a rule in Amazon EventBridge to forward PutObject events to the AWS Glue trigger.

Lý do chọn ✅:
Giải pháp này sử dụng event-driven architecture hoàn hảo cho kịch bản file upload bất định. Khi enable S3 data events trong CloudTrail (management/data events cho PutObject), các event được gửi đến EventBridge (trước đây là CloudWatch Events). Rule EventBridge forward event PutObject trực tiếp kích hoạt event-based trigger của AWS Glue workflow/job.
🛠️ Ưu điểm:

  • Real-time/low-latency: Event kích hoạt gần như ngay lập tức (thường <1 phút, tùy region).
  • Scalable: Hỗ trợ hundreds workers (Glue Spark 4.0+ với Ray engine từ 2024).
  • Tối ưu chi phí: Không polling, chỉ chạy khi cần.
    Phù hợp phiên bản AWS 2026: EventBridge hỗ trợ Glue triggers native (docs AWS Glue Developer Guide).

📘 Tài liệu tham khảo:

🛠️ Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên best practices AWS mới nhất.

  • Create an on-demand AWS Glue trigger to start the workflow. Create an AWS Lambda function that runs every 15 minutes to check the S3 bucket for the daily file. Configure the function to start the AWS Glue workflow if the file is present.
    ❌ Sai: Giải pháp dùng Lambda polling mỗi 15 phút (qua CloudWatch Events schedule) để kiểm tra S3 → Không đáp ứng "xử lý nhanh nhất" vì delay tối đa 15 phút (thậm chí lâu hơn nếu upload ngay sau check). Polling tốn chi phí Lambda invocations không cần thiết, vi phạm nguyên tắc event-driven (AWS Well-Architected). Không real-time, kém hiệu quả cho file lớn/hàng trăm workers.

  • Create an event-based AWS Glue trigger to start the workflow. Configure Amazon S3 to log events to AWS CloudTrail. Create a rule in Amazon EventBridge to forward PutObject events to the AWS Glue trigger.
    ✅ Đúng: Như giải thích ở phần trên. Đây là event-based real-time chuẩn AWS: CloudTrail capture PutObject data event → EventBridge rule match pattern (ví dụ: {"source":["aws.s3"],"detail-type":["AWS API Call via CloudTrail"],"detail":{"eventSource":["s3.amazonaws.com"],"eventName":["PutObject"]}}) → Trigger Glue workflow ngay. Hỗ trợ scale lớn, low-latency (<60s thường xuyên ở regions chính, theo AWS benchmarks 2025).

  • Create a scheduled AWS Glue trigger to start the workflow. Create a cron job that runs the AWS Glue job every 15 minutes. Set up the AWS Glue job to check the S3 bucket for the daily file. Configure the job to stop if the file is not present.
    ❌ Sai: Scheduled trigger + cron every 15 phút → Polling trong Glue job (job check S3, exit early nếu không có file). Delay lên đến 15 phút, tốn kém vì mỗi lần chạy full job (hàng trăm workers khởi động vô ích 96 lần/ngày nếu file chỉ 1 lần). Glue job không nên dùng làm poller (không idempotent, chi phí DPUs cao). Không event-driven.

  • Create an on-demand AWS Glue trigger to start the workflow. Create an AWS Database Migration Service (AWS DMS) migration task. Set the DMS source as the S3 bucket. Set the target endpoint as the AWS Glue workflow.
    ❌ Sai: AWS DMS dùng cho database migration (CDC, full load từ RDBMS/NoSQL/S3 as source), KHÔNG hỗ trợ target là Glue workflow (DMS targets chỉ endpoints như Redshift/S3/Kafka, không trigger Glue). DMS từ S3 chỉ hỗ trợ unstructured data replication, không phù hợp xử lý pipeline phức tạp với hundreds workers. Sai kiến trúc, thêm complexity/latency không cần (DMS có overhead setup tasks).