Ngân hàng đề — AWS Certified Data Engineer Associate

Tìm thấy 867 câu.

Câu 781
A car sales company maintains data about cars that are listed for sale in an area. The company receives data about new car listings from vendors who upload the data daily as compressed files into Amazon S3. The compressed files are up to 5 KB in size. The company wants to see the most up-to-date listings as soon as the data is uploaded to Amazon S3.

A data engineer must automate and orchestrate the data processing workflow of the listings to feed a dashboard. The data engineer must also provide the ability to perform one-time queries and analytical reporting. The query solution must be scalable.

Which solution will meet these requirements MOST cost-effectively?
  1. A Use an Amazon EMR cluster to process incoming data. Use AWS Step Functions to orchestrate workflows. Use Apache Hive for one-time queries and analytical reporting. Use Amazon OpenSearch Service to bulk ingest the data into compute optimized instances. Use OpenSearch Dashboards in OpenSearch Service for the dashboard.
  2. B Use a provisioned Amazon EMR cluster to process incoming data. Use AWS Step Functions to orchestrate workflows. Use Amazon Athena for one-time queries and analytical reporting. Use Amazon QuickSight for the dashboard.
  3. C Use AWS Glue to process incoming data. Use AWS Step Functions to orchestrate workflows. Use Amazon Redshift Spectrum for one-time queries and analytical reporting. Use OpenSearch Dashboards in Amazon OpenSearch Service for the dashboard.
  4. D Use AWS Glue to process incoming data. Use AWS Lambda and S3 Event Notifications to orchestrate workflows. Use Amazon Athena for one-time queries and analytical reporting. Use Amazon QuickSight for the dashboard.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi mô tả một công ty bán xe lưu trữ dữ liệu danh sách xe (listings) mới từ các nhà cung cấp, được upload hàng ngày dưới dạng file nén nhỏ (tối đa 5 KB) vào Amazon S3. Yêu cầu chính là xem listings mới nhất ngay lập tức khi dữ liệu được upload, đồng thời tự động hóa và điều phối (orchestrate) quy trình xử lý dữ liệu để cung cấp cho dashboard. Ngoài ra, cần hỗ trợ one-time queries (truy vấn một lần) và analytical reporting (báo cáo phân tích), với giải pháp phải scalable (mở rộng linh hoạt) và MOST cost-effectively (tiết kiệm chi phí nhất).

🛠️ Phân tích yêu cầu chính:

  • Trigger ngay lập tức: Sử dụng sự kiện S3 (S3 Event Notifications) để xử lý real-time.
  • Xử lý dữ liệu nhỏ: Không cần cluster lớn, ưu tiên serverless để tránh chi phí idle.
  • Orchestration: Tự động hóa workflow mà không tốn kém.
  • Queries & Reporting: Serverless query engine, pay-per-use, tích hợp dashboard.
  • Cost-effective: Tránh provisioned resources (như EMR cluster), ưu tiên serverless (Glue, Lambda, Athena).

Dựa trên kiến thức AWS cập nhật đến 2026 (AWS re:Invent 2025 updates: Glue 4.0 với Spark 3.5, Athena engine v3, QuickSight ML insights), giải pháp serverless là tối ưu nhất cho workload nhỏ, intermittent.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use AWS Glue to process incoming data. Use AWS Lambda and S3 Event Notifications to orchestrate workflows. Use Amazon Athena for one-time queries and analytical reporting. Use Amazon QuickSight for the dashboard.

Lý do:

  • 🛠️ AWS Glue: Serverless ETL, xử lý file nhỏ từ S3 hiệu quả, tự động scale, chỉ tính phí execution time (rẻ cho 5KB files).
  • 🚀 AWS Lambda + S3 Event Notifications: Trigger real-time khi upload S3 (near real-time <1 phút), orchestrate serverless mà không cần Step Functions (tiết kiệm ~70% chi phí orchestration so với provisioned).
  • 📊 Amazon Athena: Serverless query S3/Glue output, pay-per-query (TB scanned), scalable cho ad-hoc queries & analytics, tích hợp trực tiếp QuickSight.
  • 🖥️ Amazon QuickSight: Dashboard serverless, kết nối Athena/Glue, real-time viz listings mới, ML-powered insights (2025 update).
  • 💰 Cost-effective nhất: Toàn serverless → zero idle cost, phù hợp data nhỏ/sporadic. Tổng chi phí thấp hơn EMR/Redshift 5-10x (theo AWS Cost Calculator).

❌ Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc. Sử dụng kiến thức AWS 2026 (Glue Crawlers v2.0, Athena federated queries).

  • Phương án 1 [SAI] ❌: Use an Amazon EMR cluster to process incoming data. Use AWS Step Functions to orchestrate workflows. Use Apache Hive for one-time queries and analytical reporting. Use Amazon OpenSearch Service to bulk ingest the data into compute optimized instances. Use OpenSearch Dashboards in OpenSearch Service for the dashboard.

    • Lý do sai: EMR cluster (on-demand/provisioned) tốn kém cho file 5KB (idle cost cao), không real-time. Hive chậm cho ad-hoc queries. OpenSearch (compute-optimized) overkill cho listings đơn giản, bulk ingest không cần thiết → chi phí cao gấp 10x so serverless.
  • Phương án 2 [SAI] ❌: Use a provisioned Amazon EMR cluster to process incoming data. Use AWS Step Functions to orchestrate workflows. Use Amazon Athena for one-time queries and analytical reporting. Use Amazon QuickSight for the dashboard.

    • Lý do sai: EMR provisioned lãng phí cho data nhỏ (phải chạy cluster liên tục), Step Functions thêm chi phí state transitions (~$0.025/1k transitions). Athena/QuickSight tốt nhưng orchestration kém hiệu quả → không cost-effective.
  • Phương án 3 [SAI] ❌: Use AWS Glue to process incoming data. Use AWS Step Functions to orchestrate workflows. Use Amazon Redshift Spectrum for one-time queries and analytical reporting. Use OpenSearch Dashboards in Amazon OpenSearch Service for the dashboard.

    • Lý do sai: Glue tốt, nhưng Step Functions thừa (Lambda rẻ hơn). Redshift Spectrum yêu cầu Redshift cluster (chi phí cao ~$0.25/giờ/node + Spectrum query fees), không pay-per-query thuần như Athena. OpenSearch Dashboards không phù hợp dashboard listings (complex setup) → kém scalable & đắt.
  • Phương án 4 [ĐÚNG] ✅: Use AWS Glue to process incoming data. Use AWS Lambda and S3 Event Notifications to orchestrate workflows. Use Amazon Athena for one-time queries and analytical reporting. Use Amazon QuickSight for the dashboard.

    • Lý do đúng: Như đã giải thích ở phần ✅, toàn serverless, real-time, scalable, chi phí thấp nhất (AWS khuyến nghị cho S3 event-driven ETL - pattern "S3 → Lambda → Glue → Athena → QuickSight"). Hoàn hảo cho "most up-to-date" và ad-hoc queries.

🛡️ Kết luận: Giải pháp serverless là best practice AWS 2026 cho workload này, tránh over-provisioning!

Câu 782
A company has AWS resources in multiple AWS Regions. The company has an Amazon EFS file system in each Region where the company operates. The company’s data science team operates within only a single Region. The data that the data science team works with must remain within the team's Region.

A data engineer needs to create a single dataset by processing files that are in each of the company's Regional EFS file systems. The data engineer wants to use an AWS Step Functions state machine to orchestrate AWS Lambda functions to process the data.

Which solution will meet these requirements with the LEAST effort?
  1. A Peer the VPCs that host the EFS file systems in each Region with the VPC that is in the data science team’s Region. Enable EFS file locking. Configure the Lambda functions in the data science team's Region to mount each of the Region specific file systems. Use the Lambda functions to process the data.
  2. B Configure each of the Regional EFS file systems to replicate data to the data science team's Region. In the data science team’s Region, configure the Lambda functions to mount the replica file systems. Use the Lambda functions to process the data.
  3. C Deploy the Lambda functions to each Region. Mount the Regional EFS file systems to the Lambda functions. Use the Lambda functions to process the data. Store the output in an Amazon S3 bucket in the data science team’s Region.
  4. D Use AWS DataSync to transfer files from each of the Regional EFS files systems to the file system that is in the data science team's Region. Configure the Lambda functions in the data science team's Region to mount the file system that is in the same Region. Use the Lambda functions to process the data.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh một công ty có tài nguyên AWS phân bố ở nhiều Region, với Amazon EFS file system ở mỗi Region nơi họ hoạt động. Đội data science chỉ làm việc trong một Region duy nhất, và dữ liệu họ xử lý phải giữ nguyên trong Region đó (không được di chuyển ra ngoài). Một data engineer cần tạo một dataset duy nhất bằng cách xử lý các file từ tất cả EFS ở các Region khác nhau. Họ muốn sử dụng AWS Step Functions để orchestrate (điều phối) các AWS Lambda functions thực hiện việc xử lý dữ liệu.

Yêu cầu chính: Giải pháp phải đáp ứng với LEAST effort (ít công sức nhất), đảm bảo dữ liệu cuối cùng ở Region của đội data science, và tuân thủ quy tắc dữ liệu không rời Region đội ngũ.

Thách thức kỹ thuật (dựa trên kiến thức AWS cập nhật đến 2026):

  • EFS là dịch vụ Regional (chỉ mount được trong cùng Region/VPC), không hỗ trợ mount cross-Region trực tiếp.
  • Cần chuyển dữ liệu từ các EFS nguồn (multi-Region) về EFS đích ở Region data science một cách managed, hiệu quả.
  • Sử dụng Step Functions + Lambda để orchestrate, nhưng trọng tâm là cách đưa dữ liệu về một nơi để Lambda process.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use AWS DataSync to transfer files from each of the Regional EFS files systems to the file system that is in the data science team's Region. Configure the Lambda functions in the data science team's Region to mount the file system that is in the same Region. Use the Lambda functions to process the data.

Lý do 🛠️:

  • AWS DataSync là dịch vụ managed chuyên chuyển dữ liệu cross-Region giữa EFS (EFS-to-EFS), hỗ trợ incremental sync, compression, và scheduling tự động – least effort vì không cần code custom, chỉ configure task.
  • Dữ liệu được copy về EFS ở Region data science → Lambda mount local EFS → process → dataset duy nhất ở đúng Region, tuân thủ yêu cầu.
  • Step Functions orchestrate DataSync tasks (qua Lambda invoke) + Lambda process → workflow mượt mà.
  • Tiết kiệm effort: Không cần VPC peering phức tạp, deploy multi-Region, hay replication thủ công. DataSync handle security (VPC endpoints), bandwidth optimization (tối ưu 2025+).

📋 Phân tích tất cả các phương án (đúng/sai)

  • ❌ Phương án SAI: Peer the VPCs that host the EFS file systems in each Region with the VPC that is in the data science team’s Region. Enable EFS file locking. Configure the Lambda functions in the data science team's Region to mount each of the Region specific file systems. Use the Lambda functions to process the data.
    Giải thích: VPC peering không hỗ trợ mount EFS cross-Region vì EFS chỉ accessible trong cùng Region (Mount targets là Regional endpoints). File locking cũng vô ích vì không mount được. Effort cao (setup peering multi-Region phức tạp, security groups, routes), và vi phạm "data remain in team's Region" do Lambda cố mount remote EFS (không khả thi).

  • ❌ Phương án SAI: Configure each of the Regional EFS file systems to replicate data to the data science team's Region. In the data science team’s Region, configure the Lambda functions to mount the replica file systems. Use the Lambda functions to process the data.
    Giải thích: EFS không có built-in replication cross-Region (chỉ intra-Region với EFS Replication cho disaster recovery, không phải one-way copy). Phải dùng custom solution (như DataSync hoặc script), làm effort cao hơn DataSync managed. Replica file systems vẫn cần setup thủ công, không least effort.

  • ❌ Phương án SAI: Deploy the Lambda functions to each Region. Mount the Regional EFS file systems to the Lambda functions. Use the Lambda functions to process the data. Store the output in an Amazon S3 bucket in the data science team’s Region.
    Giải thích: Deploy Lambda multi-Region → effort cao (quản lý code, IAM, Step Functions cross-Region sync phức tạp). Dữ liệu process ở nhiều Region → vi phạm yêu cầu "data must remain within the team's Region" (process xảy ra ngoài). Output S3 cross-Region chỉ là partial, không tạo dataset từ EFS unified.

  • ✅ Phương án ĐÚNG (như đã giải thích ở trên): Use AWS DataSync to transfer files from each of the Regional EFS files systems to the file system that is in the data science team's Region. Configure the Lambda functions in the data science team's Region to mount the file system that is in the same Region. Use the Lambda functions to process the data.
    Giải thích bổ sung: Least effort nhờ DataSync agentless (2025+ fully serverless cho EFS), integrate trực tiếp Step Functions. Hiệu suất cao (up to 10 GB/s cross-Region), chi phí thấp (pay-per-transfer).

Kết luận 🚀: Giải pháp DataSync là optimal cho DevOps workflow, scalable đến 2026 với AI-optimized transfers trong DataSync. Nếu implement, dùng Step Functions với Task state gọi DataSync StartTaskExecution.

Câu 783
A company hosts its applications on Amazon EC2 instances. The company must use SSL/TLS connections that encrypt data in transit to communicate securely with AWS infrastructure that is managed by a customer.

A data engineer needs to implement a solution to simplify the generation, distribution, and rotation of digital certificates. The solution must automatically renew and deploy SSL/TLS certificates.

Which solution will meet these requirements with the LEAST operational overhead?
  1. A Store self-managed certificates on the EC2 instances.
  2. B Use AWS Certificate Manager (ACM).
  3. C Implement custom automation scripts in AWS Secrets Manager.
  4. D Use Amazon Elastic Container Service (Amazon ECS) Service Connect.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào một công ty đang chạy ứng dụng trên Amazon EC2 instances, và họ bắt buộc phải sử dụng SSL/TLS để mã hóa dữ liệu trong quá trình truyền (data in transit) khi giao tiếp an toàn với các infrastructure AWS do khách hàng quản lý (customer-managed AWS infrastructure).

Một data engineer cần triển khai giải pháp đơn giản hóa việc tạo (generation), phân phối (distribution), và xoay vòng (rotation) các digital certificates. Giải pháp phải tự động renew (tái cấp mới) và deploy SSL/TLS certificates, đồng thời đạt operational overhead thấp nhất (LEAST operational overhead).

🛠️ Mục tiêu chính: Tìm giải pháp AWS managed service giúp tự động hóa toàn bộ lifecycle của certificates mà không cần quản lý thủ công, phù hợp với best practices DevOps trên AWS (theo kiến thức cập nhật đến 2026, ACM vẫn là lựa chọn chuẩn cho việc này với tích hợp sâu vào EC2, ELB, CloudFront, v.v.).

📘 Tài liệu tham khảo:

✅ Đáp án đúng: Use AWS Certificate Manager (ACM)

Lý do lựa chọn (theo phiên bản AWS mới nhất 2026):
ACM là dịch vụ fully managed của AWS, chuyên xử lý toàn bộ lifecycle của public và private certificates (bao gồm generation qua ACM PCA, distribution tự động, rotation, và auto-renewal miễn phí cho public certs từ public CA như Amazon Trust Services). Nó tích hợp seamless với EC2 qua Application Load Balancer (ALB), Network Load Balancer (NLB), API Gateway, CloudFront, giúp deploy certificates chỉ với vài cú click mà không cần code hay script. Operational overhead thấp nhất vì AWS lo hết renew (tự động 30-60 ngày trước hết hạn) và deploy. Hoàn hảo cho EC2 apps cần SSL/TLS encrypt transit data đến customer-managed infra.

🔍 Giải thích tất cả các phương án (đúng/sai)

  • Store self-managed certificates on the EC2 instances.
    ❌ Sai: Phương án này yêu cầu tự quản lý thủ công (self-signed hoặc từ CA ngoài), lưu certs trực tiếp trên EC2 (ví dụ via IAM roles hoặc EBS). Không hỗ trợ auto-renew/rotation, phải script cron job thủ công để update, dẫn đến operational overhead cao (rủi ro downtime, lỗi config). Không phù hợp với yêu cầu "simplify" và "least overhead". ACM thay thế hoàn hảo mà không cần lưu trên instance.

  • Use AWS Certificate Manager (ACM).
    ✅ Đúng (như đã giải thích ở trên): Fully automated, zero-touch renew/deploy, tích hợp native với AWS services cho EC2 traffic. Theo AWS re:Invent 2025 updates, ACM hỗ trợ imported certs rotation và zero-downtime deployment cho ALB/EC2, đạt least overhead thực sự.

  • Implement custom automation scripts in AWS Secrets Manager.
    ❌ Sai: Secrets Manager chỉ lưu trữ và rotate secrets (như API keys, passwords), không phải certificates (certs cần private key + chain đầy đủ). Phải viết custom Lambda/scripts để generate/distribute/renew certs (ví dụ dùng certbot), rồi sync với Secrets Manager – overhead rất cao (code maintain, error-prone). ACM làm thay hết mà không cần script.

  • Use Amazon Elastic Container Service (Amazon ECS) Service Connect.
    ❌ Sai: ECS Service Connect (ra mắt 2023, update 2026) dùng cho service discovery và mTLS trong containerized workloads trên ECS/EKS, không quản lý certificates (chỉ proxy traffic internal với IAM auth). Không hỗ trợ generation/renew/deploy public SSL/TLS certs cho EC2 apps giao tiếp external AWS infra. Overhead cao vì phải containerize app trước, không phù hợp EC2 bare-metal.

🛡️ Kết luận DevOps Pro: Chọn ACM để tuân thủ zero-trust security và automation-first trên AWS, giảm MTTR (Mean Time To Recovery) xuống mức thấp nhất! 🚀

Câu 784 Chọn nhiều đáp án
A company saves customer data to an Amazon S3 bucket. The company uses server-side encryption with AWS KMS keys (SSE-KMS) to encrypt the bucket. The dataset includes personally identifiable information (PII) such as social security numbers and account details.

Data that is tagged as PII must be masked before the company uses customer data for analysis. Some users must have secure access to the PII data during the pre-processing phase. The company needs a low-maintenance solution to mask and secure the PII data throughout the entire engineering pipeline.

Which combination of solutions will meet these requirements? (Choose two.)
  1. A Use AWS Glue DataBrew to perform extract, transform, and load (ETL) tasks that mask the PII data before analysis.
  2. B Use Amazon GuardDuty to monitor access patterns for the PII data that is used in the engineering pipeline.
  3. C Configure an Amazon Macie discovery job for the S3 bucket.
  4. D Use AWS Identity and Access Management (IAM) to manage permissions and to control access to the PII data.
  5. E Write custom scripts in an application to mask the PII data and to control access.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh một công ty lưu trữ dữ liệu khách hàng trong Amazon S3 bucket, sử dụng mã hóa server-side encryption với AWS KMS keys (SSE-KMS) để bảo vệ dữ liệu. Dataset chứa PII (Personally Identifiable Information) nhạy cảm như số an sinh xã hội và chi tiết tài khoản.

Yêu cầu chính:

  • Dữ liệu được đánh dấu là PII phải được mask (che giấu) trước khi sử dụng cho phân tích.
  • Một số người dùng cần truy cập an toàn vào dữ liệu PII trong giai đoạn pre-processing.
  • Giải pháp phải low-maintenance (bảo trì thấp) cho toàn bộ engineering pipeline.

Câu hỏi yêu cầu chọn TWO (hai) giải pháp kết hợp để đáp ứng mask PII và secure access một cách hiệu quả, tuân thủ nguyên tắc AWS best practices về bảo mật dữ liệu và data pipeline.

📘 Tài liệu tham khảo:

✅ Đáp án đúng (Chọn TWO)

Hai phương án đúng là:

  1. Use AWS Glue DataBrew to perform extract, transform, and load (ETL) tasks that mask the PII data before analysis.
  2. Use AWS Identity and Access Management (IAM) to manage permissions and to control access to the PII data.

Lý do lựa chọn:

  • AWS Glue DataBrew là công cụ visual ETL low-code/no-code chuyên cho data preparation, hỗ trợ masking PII (như hash, redact) qua các recipe sẵn có, tích hợp trực tiếp với S3 và pipeline phân tích (Glue Jobs, Athena). Nó đảm bảo low-maintenance vì không cần code phức tạp, tự động scale và serverless.
  • IAM là giải pháp chuẩn của AWS để quản lý quyền truy cập chi tiết (least privilege), hỗ trợ secure access cho người dùng cụ thể trong pre-processing qua policies, roles, và điều kiện (e.g., MFA, tags). Kết hợp hai cái này tạo pipeline an toàn, tuân thủ GDPR/CCPA mà không cần bảo trì cao.

🛠️ Phân tích chi tiết tất cả các phương án

  • ✅ Use AWS Glue DataBrew to perform extract, transform, and load (ETL) tasks that mask the PII data before analysis.
    Đúng: DataBrew cung cấp giao diện kéo-thả để tạo ETL jobs masking PII (e.g., anonymize SSNs), tích hợp S3/KMS, chạy serverless trên pipeline. Đây là giải pháp low-maintenance lý tưởng cho engineering pipeline, tránh custom code. (Cập nhật 2026: Hỗ trợ ML-based PII detection trong recipes.)

  • ❌ Use Amazon GuardDuty to monitor access patterns for the PII data that is used in the engineering pipeline.
    Sai: GuardDuty là dịch vụ threat detection (phát hiện bất thường như truy cập lạ từ IP đáng ngờ), không masking PII hay kiểm soát access trực tiếp. Nó chỉ monitor sau sự kiện, không giải quyết yêu cầu mask/secure trong pipeline, và không low-maintenance cho mục đích này.

  • ❌ Configure an Amazon Macie discovery job for the S3 bucket.
    Sai: Macie chuyên discovery và classify PII tự động trong S3 (tìm SSNs, tài khoản), tạo findings. Tuy nhiên, nó không masking dữ liệu hay cung cấp secure access; chỉ hỗ trợ detect, cần kết hợp tool khác. Không đủ cho toàn pipeline low-maintenance.

  • ✅ Use AWS Identity and Access Management (IAM) to manage permissions and to control access to the PII data.
    Đúng: IAM quản lý permissions granular (roles, policies với conditions như resource tags=PII), đảm bảo chỉ user cần thiết truy cập pre-processing. Serverless, zero-maintenance sau setup, tích hợp KMS/S3. Best practice cho secure access mà không cần tool ngoài.

  • ❌ Write custom scripts in an application to mask the PII data and to control access.
    Sai: Custom scripts yêu cầu development, testing, maintenance cao (vi phạm low-maintenance), dễ lỗi bảo mật, không scale tự động như AWS managed services. Không khuyến nghị cho production pipeline với PII nhạy cảm.

Kết luận 🏆: Kết hợp DataBrew (masking ETL) + IAM (access control) là giải pháp tối ưu, serverless, tuân thủ AWS Security best practices đến 2026!

Câu 785
A data engineer is launching an Amazon EMR cluster. The data that the data engineer needs to load into the new cluster is currently in an Amazon S3 bucket. The data engineer needs to ensure that data is encrypted both at rest and in transit.

The data that is in the S3 bucket is encrypted by an AWS Key Management Service (AWS KMS) key. The data engineer has an Amazon S3 path that has a Privacy Enhanced Mail (PEM) file.

Which solution will meet these requirements?
  1. A Create an Amazon EMR security configuration. Specify the appropriate AWS KMS key for at-rest encryption for the S3 bucket. Create a second security configuration. Specify the Amazon S3 path of the PEM file for in-transit encryption. Create the EMR cluster, and attach both security configurations to the cluster.
  2. B Create an Amazon EMR security configuration. Specify the appropriate AWS KMS key for local disk encryption for the S3 bucket. Specify the Amazon S3 path of the PEM file for in-transit encryption. Use the security configuration during EMR cluster creation.
  3. C Create an Amazon EMR security configuration. Specify the appropriate AWS KMS key for at-rest encryption for the S3 bucket. Specify the Amazon S3 path of the PEM file for in-transit encryption. Use the security configuration during EMR cluster creation.
  4. D Create an Amazon EMR security configuration. Specify the appropriate AWS KMS key for at-rest encryption for the S3 bucket. Specify the Amazon S3 path of the PEM file for in-transit encryption. Create the EMR cluster, and attach the security configuration to the cluster.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh việc một data engineer đang khởi chạy một Amazon EMR cluster và cần tải dữ liệu từ Amazon S3 bucket vào cluster. Dữ liệu trong S3 đã được mã hóa at rest bằng AWS KMS key. Ngoài ra, có một file PEM (Privacy Enhanced Mail) lưu trữ tại đường dẫn S3 để hỗ trợ mã hóa. Yêu cầu chính: Đảm bảo dữ liệu được mã hóa both at rest (khi lưu trữ trên đĩa local của EMR cluster) và in transit (khi truyền từ S3 đến cluster qua EMRFS).

📘 Chi tiết kỹ thuật:

  • At rest encryption: Sử dụng EMR Security Configuration để mã hóa dữ liệu trên local disk (EBS volumes và instance store) của các node EMR bằng KMS key.
  • In transit encryption: Sử dụng TLS/SSL cho EMRFS (EMR File System) kết nối với S3, với certificate từ file PEM (custom CA cert để verify S3 endpoint).
  • Quy trình: Tạo EMR Security Configuration (JSON file) trước, sau đó áp dụng trong lúc tạo cluster (không attach sau).
  • Lưu ý: S3 đã encrypted riêng, KMS key ở đây dùng cho local disk của EMR, không phải S3 bucket trực tiếp.

🛠️ Kiến thức AWS cập nhật 2026: EMR hỗ trợ security configurations từ phiên bản 5.28+, với TLS 1.2+ cho in-transit và KMS customer-managed keys cho at-rest. Không thay đổi lớn ở EMR 7.x.

📚 Tài liệu tham khảo:

✅ Đáp án đúng: Phương án thứ 3

Create an Amazon EMR security configuration. Specify the appropriate AWS KMS key for at-rest encryption for the S3 bucket. Specify the Amazon S3 path of the PEM file for in-transit encryption. Use the security configuration during EMR cluster creation.

Lý do chọn:

  • ✅ Tạo một security configuration duy nhất (JSON) để cấu hình both at-rest (KMS key cho local disk encryption khi data từ S3 lưu vào cluster) và in-transit (PEM file cho TLS/SSL EMRFS).
  • ✅ Áp dụng during EMR cluster creation (sử dụng --security-configuration trong CLI/API), đảm bảo cluster được config ngay từ đầu.
  • ✅ Phù hợp yêu cầu: Dữ liệu encrypted at rest trên local disk EMR và in transit từ S3. Cụm từ "for the S3 bucket" ám chỉ ngữ cảnh data source, nhưng thực tế KMS apply cho local storage.

📋 Giải thích tất cả các phương án

  • ❌ Phương án 1 (SAI):
    Create an Amazon EMR security configuration. Specify the appropriate AWS KMS key for at-rest encryption for the S3 bucket. Create a second security configuration. Specify the Amazon S3 path of the PEM file for in-transit encryption. Create the EMR cluster, and attach both security configurations to the cluster.
    Lý do sai: EMR chỉ hỗ trợ một security configuration per cluster, không tạo hai cái riêng biệt. Không thể "attach both" – phải chỉ định một cái duy nhất lúc tạo cluster. Quá phức tạp và không theo docs.

  • ❌ Phương án 2 (SAI):
    Create an Amazon EMR security configuration. Specify the appropriate AWS KMS key for local disk encryption for the S3 bucket. Specify the Amazon S3 path of the PEM file for in-transit encryption. Use the security configuration during EMR cluster creation.
    Lý do sai: Cụm "local disk encryption for the S3 bucket" không chính xác – KMS dùng cho local disk của EMR, không phải S3 bucket (S3 đã encrypted riêng). Dù chỉ định "local disk" đúng hơn, nhưng phrasing sai dẫn đến hiểu lầm config nhầm chỗ.

  • ✅ Phương án 3 (ĐÚNG): (Đã giải thích ở trên)
    Hoàn hảo, một config duy nhất, apply đúng thời điểm.

  • ❌ Phương án 4 (SAI):
    Create an Amazon EMR security configuration. Specify the appropriate AWS KMS key for at-rest encryption for the S3 bucket. Specify the Amazon S3 path of the PEM file for in-transit encryption. Create the EMR cluster, and attach the security configuration to the cluster.
    Lý do sai: Không thể "attach sau khi create cluster" – EMR security config phải chỉ định trong lúc launch cluster (immutable sau khi tạo). Nếu tạo cluster trước thì không apply được.

🧩 Kết luận: Phương án 3 là optimal, tuân thủ best practices AWS EMR để secure data pipeline từ S3! 🚀

Câu 786 Chọn nhiều đáp án
A retail company is using an Amazon Redshift cluster to support real-time inventory management. The company has deployed an ML model on a real-time endpoint in Amazon SageMaker.

The company wants to make real-time inventory recommendations. The company also wants to make predictions about future inventory needs.

Which solutions will meet these requirements? (Choose two.)
  1. A Use Amazon Redshift ML to generate inventory recommendations.
  2. B Use SQL to invoke a remote SageMaker endpoint for prediction.
  3. C Use Amazon Redshift ML to schedule regular data exports for offline model training.
  4. D Use SageMaker Autopilot to create inventory management dashboards in Amazon Redshift.
  5. E Use Amazon Redshift as a file storage system to archive old inventory management reports.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào một công ty bán lẻ sử dụng Amazon Redshift cluster để hỗ trợ quản lý hàng tồn kho (inventory management) theo thời gian thực. Họ đã triển khai một mô hình Machine Learning (ML) trên endpoint real-time của Amazon SageMaker.

Yêu cầu chính là:

  • Tạo khuyến nghị hàng tồn kho thời gian thực (real-time inventory recommendations).
  • Dự đoán nhu cầu hàng tồn kho tương lai (predictions about future inventory needs).

Cần chọn HAI giải pháp phù hợp nhất, tận dụng tích hợp giữa Redshift và SageMaker để xử lý dữ liệu real-time mà không cần di chuyển dữ liệu lớn, đảm bảo hiệu suất cao và chi phí tối ưu. Đây là kịch bản điển hình trong AWS, nơi Redshift hỗ trợ ML inference trực tiếp qua SQL (cập nhật mới nhất đến 2026 với Redshift ML enhancements).

✅ Đáp án đúng và lý do lựa chọn

Hai đáp án đúng là:

  1. Use Amazon Redshift ML to generate inventory recommendations.

    • Lý do: Amazon Redshift ML (tính năng ra mắt từ 2020 và cập nhật liên tục) cho phép tạo và chạy mô hình ML trực tiếp từ SQL trong Redshift mà không cần di chuyển dữ liệu. Bạn có thể dùng lệnh CREATE MODEL để huấn luyện mô hình nội bộ (on-cluster) cho khuyến nghị thời gian thực dựa trên dữ liệu inventory hiện tại. Điều này lý tưởng cho real-time recommendations, hỗ trợ cả classification/regression phù hợp với dự đoán nhu cầu.
  2. Use SQL to invoke a remote SageMaker endpoint for prediction.

    • Lý do: Redshift hỗ trợ tích hợp trực tiếp với SageMaker qua lệnh CREATE MODEL với option REMOTE FUNCTION, cho phép gọi SQL query để invoke endpoint SageMaker real-time. Dữ liệu từ Redshift được serialize và gửi đến endpoint để inference, trả về kết quả dự đoán tương lai ngay lập tức, không cần ETL phức tạp. Hoàn hảo cho predictions về future inventory needs.

🛠️ Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên tính phù hợp với yêu cầu real-time recommendations và predictions:

  • Use Amazon Redshift ML to generate inventory recommendations.
    ✅ Đúng: Như đã giải thích, Redshift ML cho phép xây dựng mô hình recommendation trực tiếp từ dữ liệu Redshift qua SQL (ví dụ: CREATE MODEL inventory_recommendations FROM (SELECT ... ) PREDICT ...), hỗ trợ real-time inference với latency thấp, tích hợp hoàn hảo cho inventory management mà không rời khỏi cluster.

  • Use SQL to invoke a remote SageMaker endpoint for prediction.
    ✅ Đúng: Tính năng Redshift ML hỗ trợ remote inference qua SageMaker endpoint bằng lệnh CREATE MODEL ... FUNCTION_NAME AS (SELECT sagemaker_endpoint('model-name', ...)). Điều này cho phép query SQL đơn giản để dự đoán real-time trên dữ liệu inventory, tận dụng mô hình đã deploy sẵn.

  • Use Amazon Redshift ML to schedule regular data exports for offline model training.
    ❌ Sai: Redshift ML tập trung vào in-database ML (huấn luyện/inference nội bộ hoặc remote), không hỗ trợ "schedule regular data exports" cho offline training. Việc export dữ liệu định kỳ sẽ làm chậm real-time processing và không phù hợp với yêu cầu; thay vào đó, dùng UNLOAD hoặc SageMaker Processing Jobs riêng biệt.

  • Use SageMaker Autopilot to create inventory management dashboards in Amazon Redshift.
    ❌ Sai: SageMaker Autopilot chỉ tự động hóa việc xây dựng/tuning mô hình ML (không tạo dashboards). Redshift không phải công cụ dashboard (dùng QuickSight hoặc Tableau thay thế). Phương án này không liên quan đến recommendations hay predictions, chỉ là nhầm lẫn giữa ML modeling và visualization.

  • Use Amazon Redshift as a file storage system to archive old inventory management reports.
    ❌ Sai: Redshift là data warehouse columnar cho analytics/query, không phải file storage (dùng S3 cho archiving). Sử dụng Redshift để lưu reports cũ sẽ tốn kém và kém hiệu quả, không hỗ trợ real-time ML requirements.

📘 Tài liệu tham khảo (cập nhật AWS đến 2026)

Giải pháp này đảm bảo zero-ETL giữa Redshift và SageMaker, tối ưu cho real-time inventory! 🚀

Câu 787
A company stores CSV files in an Amazon S3 bucket. A data engineer needs to process the data in the CSV files and store the processed data in a new S3 bucket.

The process needs to rename a column, remove specific columns, ignore the second row of each file, create a new column based on the values of the first row of the data, and filter the results by a numeric value of a column.

Which solution will meet these requirements with the LEAST development effort?
  1. A Use AWS Glue Python jobs to read and transform the CSV files.
  2. B Use an AWS Glue custom crawler to read and transform the CSV files.
  3. C Use an AWS Glue workflow to build a set of jobs to crawl and transform the CSV files.
  4. D Use AWS Glue DataBrew recipes to read and transform the CSV files.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc xử lý dữ liệu CSV lưu trữ trong Amazon S3 bucket, với các yêu cầu transformation cụ thể:

  • Rename một cột (đổi tên cột).
  • Remove các cột cụ thể (xóa cột không cần).
  • Ignore hàng thứ hai của mỗi file (bỏ qua header hoặc row thứ 2).
  • Tạo cột mới dựa trên giá trị của hàng đầu tiên (tính toán cột mới từ row 1).
  • Filter kết quả theo giá trị số của một cột (lọc dữ liệu dựa trên điều kiện số học).

Sau xử lý, dữ liệu được lưu vào S3 bucket mới. Mục tiêu là chọn giải pháp với LEAST development effort (ít nỗ lực phát triển nhất), nghĩa là ưu tiên công cụ low-code/no-code, dễ sử dụng mà không cần viết code phức tạp. Đây là chủ đề về AWS Glue và các dịch vụ data processing, phù hợp với kiến thức AWS Glue DataBrew (cập nhật đến 2026, DataBrew hỗ trợ visual recipes cho data prep trên S3).

📘 Tài liệu tham khảo:

  • AWS Glue DataBrew Documentation: AWS Glue DataBrew (hỗ trợ 250+ transformations visual).
  • AWS re:Post và Exam Guide DOP-C02 (DevOps Professional, 2024-2026 updates).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use AWS Glue DataBrew recipes to read and transform the CSV files.

🛠️ Lý do: AWS Glue DataBrew là dịch vụ visual data preparation (low-code), cho phép tạo recipes (công thức transformation) qua giao diện drag-and-drop. Nó hỗ trợ tất cả yêu cầu một cách trực quan:

  • Skip rows (ignore row 2).
  • Rename/remove columns.
  • Create custom columns (dựa trên row 1).
  • Filter numeric values.
    Sau đó export trực tiếp ra S3. Least effort vì không cần code Python/Spark, chỉ cần vài cú click (chạy trên serverless, scale tự động). Phù hợp nhất cho data engineer non-coder.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, đánh dấu ✅ (đúng) hoặc ❌ (sai), giữ nguyên nội dung gốc bằng tiếng Anh:

  • ❌ Use AWS Glue Python jobs to read and transform the CSV files.
    Phương án này yêu cầu viết code Python (sử dụng PySpark trong Glue ETL jobs) để đọc CSV từ S3, thực hiện các transformation thủ công (như drop(), withColumn(), filter()). Effort cao vì phải code chi tiết (ví dụ: skip row 2 cần header=False và slice data), debug lỗi, và quản lý job. Không phải "least effort" so với visual tools.

  • ❌ Use an AWS Glue custom crawler to read and transform the CSV files.
    AWS Glue Crawler chỉ dùng để scan metadata và tạo schema trong Glue Data Catalog (classify CSV schema), không hỗ trợ transform data như rename/remove/filter. Custom crawler chỉ customize classification, không làm ETL. Sai hoàn toàn vì không đáp ứng yêu cầu processing.

  • ❌ Use an AWS Glue workflow to build a set of jobs to crawl and transform the CSV files.
    Glue Workflow (nay là Glue Studio Workflows) dùng để orchestrate jobs (crawl + ETL), nhưng vẫn cần build multiple jobs với code Python/Spark cho transformations. Effort lớn: thiết kế workflow, code từng step (crawl trước, transform sau). Phức tạp hơn DataBrew, không "least effort".

  • ✅ Use AWS Glue DataBrew recipes to read and transform the CSV files.
    Như đã giải thích ở trên: Visual recipes hỗ trợ tất cả operations qua UI (Add step > Column > Rename/Remove; Parse > Skip rows; Custom > Create column từ row 1; Filter > Numeric condition). Publish recipe và schedule job tự động export ra S3. Least development effort (no-code, 80-90% faster theo AWS benchmarks 2025).

🧩 Tóm tắt so sánh: DataBrew tối ưu cho data prep nhanh (hàng trăm transformations sẵn), trong khi Glue jobs/workflows cần dev skills cao hơn. Đây là best practice cho DOP-C02 exam! 🚀

Câu 788
A company uses Amazon Redshift as its data warehouse. Data encoding is applied to the existing tables of the data warehouse. A data engineer discovers that the compression encoding applied to some of the tables is not the best fit for the data.

The data engineer needs to improve the data encoding for the tables that have sub-optimal encoding.

Which solution will meet this requirement?
  1. A Run the ANALYZE command against the identified tables. Manually update the compression encoding of columns based on the output of the command.
  2. B Run the ANALYZE COMPRESSION command against the identified tables. Manually update the compression encoding of columns based on the output of the command.
  3. C Run the VACUUM REINDEX command against the identified tables.
  4. D Run the VACUUM RECLUSTER command against the identified tables.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh Amazon Redshift – dịch vụ data warehouse của AWS, nơi công ty đang sử dụng và đã áp dụng data encoding (mã hóa nén dữ liệu) cho các bảng hiện có. Một data engineer phát hiện rằng compression encoding (mã hóa nén) trên một số bảng không tối ưu (sub-optimal), dẫn đến hiệu suất lưu trữ và query kém.
Yêu cầu chính: Cải thiện encoding cho các bảng này một cách hiệu quả nhất.
🛠️ Vấn đề cốt lõi: Redshift hỗ trợ nhiều loại compression encoding như RAW, BYTEDICT, DELTA, LZO, MOSTLY*, v.v. Encoding không phù hợp có thể làm tăng dung lượng lưu trữ và giảm tốc độ query. Giải pháp cần phân tích và đề xuất encoding tối ưu, sau đó cập nhật thủ công cho các cột.
📈 Bối cảnh cập nhật 2026: Theo tài liệu AWS Redshift mới nhất (RA3 nodes, concurrency scaling, AQUA), việc tối ưu compression vẫn dựa vào lệnh ANALYZE COMPRESSION để khuyến nghị encoding dựa trên mẫu dữ liệu thực tế.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Run the ANALYZE COMPRESSION command against the identified tables. Manually update the compression encoding of columns based on the output of the command.

Lý do:

  • Lệnh ANALYZE COMPRESSION quét bảng, phân tích phân bố dữ liệu và đề xuất encoding nén tốt nhất cho từng cột (ví dụ: RUNLENGTH cho dữ liệu lặp, ZSTD cho dữ liệu lớn).
  • Output sẽ liệt kê các encoding tiềm năng với tỷ lệ nén ước tính (savings %). Sau đó, dùng ALTER TABLE để cập nhật encoding thủ công (ví dụ: ALTER TABLE table_name ALTER COLUMN col_name ENCODE auto; hoặc encoding cụ thể).
  • Đây là best practice chính thức của AWS để fix sub-optimal encoding mà không cần rebuild toàn bộ cluster. ✅ Hiệu quả cao, không downtime lớn.

📋 Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên tài liệu Redshift:

  • ❌ SAI: Run the ANALYZE command against the identified tables. Manually update the compression encoding of columns based on the output of the command.
    Giải thích: Lệnh ANALYZE chỉ thu thập statistics (như số lượng distinct values, min/max) để optimizer query planner sử dụng, không phân tích compression encoding. Output không đề xuất encoding nén, nên không giúp cải thiện sub-optimal encoding. Chỉ dùng ANALYZE cho query optimization, không phải compression. 🧮 Không liên quan trực tiếp.

  • ✅ ĐÚNG: Run the ANALYZE COMPRESSION command against the identified tables. Manually update the compression encoding of columns based on the output of the command.
    Giải thích: Như đã nêu ở phần đáp án đúng. Lệnh này chuyên biệt quét 1 triệu dòng mẫu mỗi cột, tính toán encoding tối ưu và dự đoán savings (ví dụ: "ZSTD: 75% savings"). Sau đó ALTER TABLE để apply. Zero-ETL và phù hợp với Redshift serverless/RA3. Đây là giải pháp chuẩn. 🚀 Tối ưu nhất.

  • ❌ SAI: Run the VACUUM REINDEX command against the identified tables.
    Giải thích: VACUUM REINDEX dùng để xây dựng lại sort keys cho các bảng có sort key cũ hoặc fragmented (sau nhiều DELETE/UPDATE). Nó không thay đổi compression encoding mà chỉ reindex data theo sort order. Không giúp fix encoding sub-optimal. 🗑️ Chỉ cho maintenance sort keys.

  • ❌ SAI: Run the VACUUM RECLUSTER command against the identified tables.
    Giải thích: VACUUM RECLUSTER tự động recluster data theo distribution và sort keys để giảm skew và cải thiện query trên large tables. Nó không phân tích hay thay đổi encoding. Chỉ tối ưu physical layout, không chạm đến compression. 🔄 Dùng cho reclustering, không phải encoding.

📘 Tài liệu tham khảo (Cập nhật AWS 2026)

  • AWS Redshift Documentation: ANALYZE COMPRESSION – Chi tiết lệnh và ví dụ output.
  • Best Practices for Compression: Compression Encodings.
  • VACUUM Commands: VACUUM.
  • AWS Well-Architected Framework - Data Warehouse: Phần Performance Efficiency Pillar nhấn mạnh ANALYZE COMPRESSION cho storage optimization.
    🛡️ Lưu ý: Luôn test trên staging cluster trước khi apply ALTER TABLE để tránh downtime. Sử dụng COPY with MAXERROR nếu rebuild table.
Câu 789
The company stores a large volume of customer records in Amazon S3. To comply with regulations, the company must be able to access new customer records immediately for the first 30 days after the records are created. The company accesses records that are older than 30 days infrequently.

The company needs to cost-optimize its Amazon S3 storage.

Which solution will meet these requirements MOST cost-effectively?
  1. A Apply a lifecycle policy to transition records to S3 Standard Infrequent-Access (S3 Standard-IA) storage after 30 days.
  2. B Use S3 Intelligent-Tiering storage.
  3. C Transition records to S3 Glacier Deep Archive storage after 30 days.
  4. D Use S3 Standard-Infrequent Access (S3 Standard-IA) storage for all customer records.
Xem giải thích

🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một công ty lưu trữ lượng lớn hồ sơ khách hàng trong Amazon S3, với yêu cầu tuân thủ quy định pháp lý: phải truy cập ngay lập tức (immediate access) các hồ sơ mới trong 30 ngày đầu tiên sau khi tạo. Tuy nhiên, các hồ sơ cũ hơn 30 ngày chỉ được truy cập thỉnh thoảng (infrequently). Mục tiêu chính là tối ưu hóa chi phí lưu trữ S3 một cách hiệu quả nhất (MOST cost-effectively).

🛠️ Yêu cầu cốt lõi:

  • 30 ngày đầu: Cần độ trễ thấp, chi phí lưu trữ cao hơn (phù hợp S3 Standard).
  • Sau 30 ngày: Truy cập ít, cần lớp lưu trữ rẻ hơn nhưng vẫn hỗ trợ truy cập tương đối nhanh (không phải lưu trữ lâu dài như Glacier).
  • Giải pháp phải tự động, linh hoạt và tiết kiệm chi phí nhất có thể, sử dụng các tính năng S3 như Lifecycle Policies để chuyển lớp lưu trữ (storage classes).

✅ Đáp án đúng và lý do lựa chọn
Apply a lifecycle policy to transition records to S3 Standard Infrequent-Access (S3 Standard-IA) storage after 30 days.

Lý do:

  • Sử dụng S3 Lifecycle Policy (tính năng tự động chuyển đổi storage class) để giữ objects ở S3 Standard (truy cập ngay lập tức, độ trễ mili-giây) trong 30 ngày đầu, đảm bảo tuân thủ quy định.
  • Sau 30 ngày, tự động chuyển sang S3 Standard-IA (rẻ hơn ~40-60% so với Standard, hỗ trợ truy cập infrequent với độ trễ vài mili-giây, phí truy xuất thấp).
  • Đây là giải pháp tối ưu chi phí nhất vì chỉ áp dụng IA cho dữ liệu cũ, tránh phí không cần thiết cho dữ liệu mới, và không có phí chuyển đổi (transition fees thấp). Phù hợp hoàn hảo với mô hình sử dụng (frequent đầu, infrequent sau). Kiến thức cập nhật AWS 2026: S3 Standard-IA vẫn là lựa chọn chuẩn cho dữ liệu infrequent access ngắn hạn.

📋 Giải thích chi tiết tất cả các phương án
✅ Apply a lifecycle policy to transition records to S3 Standard Infrequent-Access (S3 Standard-IA) storage after 30 days.
Giải pháp lý tưởng: Kết hợp S3 Standard ban đầu cho truy cập nhanh + Lifecycle Policy tự động chuyển sang Standard-IA sau đúng 30 ngày, giảm chi phí lưu trữ mà vẫn đảm bảo hiệu suất. Không lãng phí cho dữ liệu mới, tối ưu nhất!

❌ Use S3 Intelligent-Tiering storage.
Sai vì S3 Intelligent-Tiering tự động di chuyển dữ liệu giữa các lớp (Standard, IA, Optional IA, Archive) dựa trên mô hình truy cập thực tế, không đảm bảo chính xác 30 ngày (có thể chuyển sớm hoặc muộn). Chi phí cao hơn do monitoring fee ($0.0025/1.000 objects/tháng), và không khớp yêu cầu "immediately access first 30 days" cố định. Không phải MOST cost-effective cho quy định nghiêm ngặt.

❌ Transition records to S3 Glacier Deep Archive storage after 30 days.
Sai vì S3 Glacier Deep Archive dành cho dữ liệu hiếm khi truy cập (archive lâu dài, retrieval 12-48 giờ, chi phí retrieval cao ~$0.00099/GB). Không phù hợp hồ sơ >30 ngày chỉ "infrequently" (cần truy cập nhanh hơn), vi phạm yêu cầu truy cập và tăng chi phí retrieval nếu cần khôi phục đột ngột.

❌ Use S3 Standard-Infrequent Access (S3 Standard-IA) storage for all customer records.
Sai vì áp dụng Standard-IA cho TẤT CẢ dữ liệu sẽ vi phạm quy định: Dữ liệu mới (30 ngày đầu) cần "immediately access" nhưng IA có độ trễ cao hơn Standard + phí retrieval tối thiểu ($0.01/1.000 requests), làm tăng chi phí không cần thiết cho truy cập frequent. Không tối ưu, lãng phí cho dữ liệu nóng.

📘 Tài liệu tham khảo (AWS cập nhật mới nhất 2026):

  • Amazon S3 Storage Classes – Chi tiết so sánh Standard vs. Standard-IA vs. Intelligent-Tiering vs. Glacier Deep Archive.
  • Object Lifecycle Management – Hướng dẫn Lifecycle Policies để transition sau N ngày.
  • S3 Pricing – Xác nhận Standard-IA rẻ hơn cho infrequent access, không phí monitoring như Intelligent-Tiering.
Câu 790
A data engineer is using Amazon QuickSight to build a dashboard to report a company’s revenue in multiple AWS Regions. The data engineer wants the dashboard to display the total revenue for a Region, regardless of the drill-down levels shown in the visual.

Which solution will meet these requirements?
  1. A Create a table calculation.
  2. B Create a simple calculated field.
  3. C Create a level-aware calculation - aggregate (LAC-A) function.
  4. D Create a level-aware calculation - window (LAC-W) function.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào Amazon QuickSight, một dịch vụ BI (Business Intelligence) của AWS dùng để tạo dashboard trực quan hóa dữ liệu. Một data engineer đang xây dựng dashboard báo cáo doanh thu (revenue) của công ty trên nhiều AWS Regions. Yêu cầu chính là: Dashboard phải hiển thị tổng doanh thu (total revenue) cho một Region cụ thể, bất kể mức độ drill-down (chi tiết hóa) nào đang được hiển thị trong visual (ví dụ: drill-down theo ngày, sản phẩm, hoặc các cấp độ khác).

🛠️ Vấn đề cốt lõi: Trong QuickSight, khi drill-down, các aggregation (tổng hợp) mặc định có thể thay đổi theo level hiện tại, dẫn đến total revenue không giữ nguyên ở level Region. Giải pháp cần cố định aggregation ở level Region mà không bị ảnh hưởng bởi drill-down. Đây là tính năng nâng cao của QuickSight (cập nhật đến phiên bản 2026), sử dụng Level-Aware Calculations (LAC) để kiểm soát level aggregation độc lập.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create a level-aware calculation - aggregate (LAC-A) function.

🧩 Lý do chi tiết: LAC-A (Level-Aware Calculation - Aggregate) cho phép tính toán aggregate (như sum, avg) ở một level cụ thể (ví dụ: Region), và giá trị này giữ nguyên bất kể drill-down xuống level con nào. Ví dụ: sum({revenue}) LAC-A [Region] sẽ luôn hiển thị total revenue của Region đó, ngay cả khi visual drill-down theo ngày hoặc sản phẩm. Đây là giải pháp chính xác theo tài liệu QuickSight mới nhất (2026), phù hợp cho dashboard multi-level visuals như table hoặc pivot table.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn, với giữ nguyên văn bản gốc và đánh dấu ✅/❌:

  • ❌ Create a table calculation.
    🛠️ Sai vì: Table calculations chỉ áp dụng cho bảng cụ thể và tính toán dựa trên hàng/cột hiện tại (như rank, running total). Chúng không kiểm soát level aggregation độc lập với drill-down, nên total revenue sẽ thay đổi khi drill-down. Không phù hợp cho yêu cầu cố định ở level Region.

  • ❌ Create a simple calculated field.
    🛠️ Sai vì: Calculated field thông thường (như sum({revenue})) aggregate theo level hiện tại của visual. Khi drill-down, nó sẽ tính total theo level con (ví dụ: per day), không giữ nguyên total Region. Thiếu khả năng "level-aware" cần thiết.

  • ✅ Create a level-aware calculation - aggregate (LAC-A) function.
    🧩 Đúng vì: Như đã giải thích, LAC-A cố định aggregate ở level chỉ định (Region), đảm bảo total revenue luôn đúng bất kể drill-down. Đây là best practice cho multi-region dashboards trong QuickSight (hỗ trợ từ 2021, tối ưu hóa đến 2026).

  • ❌ Create a level-aware calculation - window (LAC-W) function.
    🛠️ Sai vì: LAC-W dùng cho window functions (như running sum, percentile) trên partition và order cụ thể. Nó phụ thuộc vào thứ tự và partition, không phải aggregate cố định ở level cao hơn. Sẽ thay đổi theo drill-down, không đáp ứng yêu cầu total cố định per Region.

📘 Tài liệu tham khảo

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! 🚀 Nếu cần ví dụ code QuickSight, hãy hỏi thêm.