Ngân hàng đề — AWS Certified Data Engineer Associate
Tìm thấy 867 câu.
A Silicon Valley based healthcare startup uses AWS Cloud for its IT infrastructure. The startup stores patient health records on Amazon Simple Storage Service (Amazon S3). The engineering team needs to implement an archival solution based on Amazon S3 Glacier to enforce regulatory and compliance controls on data access.
Which of the following solutions would you recommend?
-
A
Use Amazon S3 Glacier vault to store the sensitive archived data and then use an Amazon S3 Access Control List to enforce compliance controls
-
B
Use Amazon S3 Glacier Deep Archive to store the sensitive archived data and then use an Amazon S3 Access Control List to enforce compliance controls
-
C
Use Amazon S3 Glacier Deep Archive to store the sensitive archived data and then use an Amazon S3 lifecycle policy to enforce compliance controls
-
D
Use Amazon S3 Glacier vault to store the sensitive archived data and then use a vault lock policy to enforce compliance controls
Xem giải thích
Đáp án
D — Dùng S3 Glacier vault để lưu dữ liệu nhạy cảm, rồi áp vault lock policy để thực thi kiểm soát tuân thủ
Vì sao đúng
Chữ khoá của đề là "thực thi kiểm soát quy định và tuân thủ" đối với việc truy cập dữ liệu, và Vault Lock là cơ chế duy nhất trong danh sách làm được điều đó theo nghĩa mạnh.
Vault Lock policy có một tính chất đặc biệt: sau khi khoá lại (lock), nó trở thành bất biến — không ai sửa hay xoá được nữa, kể cả tài khoản gốc.
Quy trình có hai bước có chủ đích:
1. Khởi tạo lock → chính sách ở trạng thái "in-progress", có 24 giờ để kiểm thử
2. Hoàn tất lock → chính sách trở thành VĨNH VIỄN, không thể huỷ
Ví dụ điển hình: WORM (write once, read many) — cấm xoá bản ghi trong 7 năm. Với hồ sơ sức khoẻ bệnh nhân, đó chính là loại kiểm soát mà quy định đòi hỏi.
Vì sao các phương án khác sai
- *C. Dùng Glacier Deep Archive kèm S3 lifecycle policy — lifecycle policy quản lý vòng đời lưu trữ (chuyển tầng, xoá theo tuổi); nó không kiểm soát truy cập và sửa được bất cứ lúc nào. Đây là phương án nhiễu chính.
- A và *B. Dùng Access Control List — ACL là cơ chế phân quyền cũ và thô, chỉ khai được vài quyền cơ bản, và cũng sửa được tự do nên không đáp ứng yêu cầu bất biến.
An Internet-of-Things (IoT) company needs a solution that can collect near real-time data from all its devices/sensors and store them in nested JSON format. The solution must offer data persistence and support the capability to query the data with a maximum latency of 10 milliseconds.
As a data engineer, how will you implement an optimal solution such that it has the LEAST operational overhead?
-
A
Use Amazon Kinesis Data Streams to capture the sensor data. Define an AWS Lambda function to process the data and write to the DynamoDB table. Store the data in Amazon DynamoDB for querying
-
B
Use Amazon Data Firehose to capture the sensor data. Directly store the data in Amazon DynamoDB for querying
-
C
Configure Amazon Simple Queue Service (Amazon SQS) to capture the real-time sensor data. Define an AWS Lambda function to poll the SQS queues and process the data. Store the data in Amazon DynamoDB for querying
-
D
Use the fully managed Apache Kafka cluster to capture the sensor data in near real-time. Store the data in Amazon S3 for querying
Xem giải thích
Đáp án
A — Kinesis Data Streams thu dữ liệu, Lambda xử lý và ghi vào DynamoDB, rồi truy vấn trên DynamoDB
Vì sao đúng
Đề nêu bốn yêu cầu, và kiến trúc này phủ cả bốn:
| Yêu cầu | Thành phần đáp ứng |
|---|---|
| Thu dữ liệu gần thời gian thực | Kinesis Data Streams |
| Lưu JSON lồng nhau | DynamoDB hỗ trợ map và list lồng nhau |
| Dữ liệu bền vững | DynamoDB nhân bản qua ba AZ |
| Truy vấn độ trễ tối đa 10 mili giây | DynamoDB cho độ trễ mili giây một chữ số |
Yêu cầu về độ trễ là chỗ quyết định: chỉ DynamoDB (và cache) đạt được mức đó cho truy vấn theo khoá.
Vì sao các phương án khác sai
- B. Firehose ghi thẳng vào DynamoDB — không làm được: Firehose không hỗ trợ DynamoDB làm đích. Đích của nó là S3, Redshift, OpenSearch, và một số endpoint HTTP. Đây là bẫy tinh vi vì Firehose đúng là lựa chọn ít công nhất nếu nó hỗ trợ.
- D. Cụm Kafka được quản lý, lưu trên S3 — S3 không đạt được độ trễ truy vấn 10 mili giây; truy cập qua API HTTP mất hàng chục mili giây.
- C. Dùng SQS để thu dữ liệu thời gian thực — SQS là hàng đợi thông điệp để tách rời ứng dụng; nó không dựng cho việc thu luồng dữ liệu cảm biến liên tục, và không giữ được thứ tự theo thiết bị.
A data engineer has created a Data Migration task using an AWS Database Migration Service (AWS DMS). The task seems to migrate all the tables with data but does not migrate the tables that have no data in them.
What is the root cause behind this behavior?
-
A
This is the expected behavior of AWS DMS. Empty tables are not migrated
-
B
There are inadequate resources allocated to the AWS DMS replication instance
-
C
You must include the schema and the tables in the table mapping of your DMS task
-
D
Tables that have secondary indexes are not migrated
Xem giải thích
Đáp án
*A — Đây là hành vi đúng như thiết kế của AWS DMS: bảng rỗng không được migrate
Vì sao đúng
DMS hoạt động theo nguyên tắc di chuyển dữ liệu, không phải di chuyển lược đồ. Ở giai đoạn full load, nó đọc từng bảng nguồn và ghi các dòng sang đích — bảng không có dòng nào thì không có gì để ghi, nên bảng đó không xuất hiện ở đích.
Đây là điểm cần nhớ về phạm vi của DMS:
| DMS làm | DMS không làm |
|---|---|
| Chuyển dữ liệu | Tạo khoá ngoại, ràng buộc, trigger |
| Tạo bảng tối thiểu để chứa dữ liệu | Tạo stored procedure, view, sequence |
| Chuyển thay đổi liên tục (CDC) | Chuyển bảng rỗng |
Cách xử lý: dùng AWS Schema Conversion Tool (SCT) để chuyển lược đồ đầy đủ trước, rồi để DMS lo phần dữ liệu. Hoặc chạy tay câu CREATE TABLE cho các bảng rỗng.
Vì sao các phương án khác sai
- C. Phải đưa schema và bảng vào table mapping của task — nếu thiếu mapping thì không bảng nào được chuyển, chứ không phải chỉ bảng rỗng bị bỏ qua. Triệu chứng không khớp. Đây là phương án nhiễu chính.
- B. Replication instance thiếu tài nguyên — sẽ gây chậm hoặc lỗi, không gây bỏ sót có chọn lọc.
- D. Bảng có secondary index không được chuyển — bịa; DMS không phân biệt theo chỉ mục.
A company utilizes an Amazon S3 bucket to store datasets accessed by various applications, including a financial services application that generates datasets containing personally identifiable information (PII). There is also an internal application that does not need access to this PII. To adhere to regulations, the company must avoid unnecessary sharing of PII. A data engineer is tasked with finding a solution that can dynamically redact PII depending on the specific needs of each application accessing the dataset.
What solution can achieve this with minimal operational overhead?
-
A
Set up an S3 Access Point to read data from the S3 bucket. Leverage the S3 Access Points policy to dynamically redact PII based on the needs of each application that accesses the data
-
B
Set up an S3 gateway endpoint to read data from the S3 bucket. Leverage the S3 bucket policy to control access for each application that accesses the data only via the gateway endpoint
-
C
Set up an S3 Object Lambda endpoint. Use the S3 Object Lambda endpoint to read data from the S3 bucket. Implement redaction logic within an S3 Object Lambda function to dynamically redact PII based on the needs of each application that accesses the data
-
D
Set up an S3 interface endpoint to read data from the S3 bucket. Leverage the S3 bucket policy to control access for each application that accesses the data only via the interface endpoint
Xem giải thích
Đáp án
C — Dựng S3 Object Lambda endpoint và cài logic che PII trong hàm Lambda gắn với nó
Vì sao đúng
Chữ khoá của đề là "che PII một cách động (dynamically redact) tuỳ theo nhu cầu của từng ứng dụng", và S3 Object Lambda là tính năng duy nhất làm được điều đó.
Cách hoạt động: nó chèn một hàm Lambda vào giữa yêu cầu GET và dữ liệu trả về:
Ứng dụng → S3 Object Lambda Access Point → Lambda (che PII) → S3
│
Ứng dụng nhận về dữ liệu ĐÃ ĐƯỢC XỬ LÝ ◀────────────┘
Điểm mấu chốt: dữ liệu gốc trên S3 không đổi — chỉ bản trả về bị biến đổi. Và vì mỗi ứng dụng dùng một access point riêng, mỗi bên nhận được một phiên bản khác nhau của cùng một tệp: ứng dụng tài chính thấy đầy đủ, ứng dụng nội bộ thấy bản đã che.
Vì sao các phương án khác sai
- A. Dùng S3 Access Point và chính sách của nó để che PII — Access Point cho phép hoặc từ chối truy cập ở mức đối tượng và tiền tố; nó không sửa được nội dung trả về. Đây là phương án nhiễu chính, và ranh giới cần nhớ là kiểm soát truy cập khác biến đổi nội dung.
- B và D. Dùng gateway endpoint hoặc interface endpoint kèm bucket policy — đều là cơ chế kiểm soát đường vào mạng, không đụng gì tới nội dung tệp.
A gaming company is developing a mobile game that streams score updates to a backend processor and then publishes results on a leaderboard. The company has hired you as an AWS Certified Data Engineer Associate to design a solution that can handle major traffic spikes, process the mobile game updates in the order of receipt, and store the processed updates in a highly available database. The company wants to minimize the management overhead required to maintain the solution.
Which of the following will you recommend to meet these requirements?
-
A
Push score updates to Amazon Kinesis Data Streams which uses a fleet of Amazon EC2 instances (with Auto Scaling) to process the updates in Amazon Kinesis Data Streams and then store these processed updates in Amazon DynamoDB
-
B
Push score updates to an Amazon Simple Queue Service (Amazon SQS) queue which uses a fleet of Amazon EC2 instances (with Auto Scaling) to process these updates in the Amazon SQS queue and then store these processed updates in an Amazon RDS MySQL database
-
C
Push score updates to Amazon Kinesis Data Streams which uses an AWS Lambda function to process these updates and then store these processed updates in Amazon DynamoDB
-
D
Push score updates to an Amazon Simple Notification Service (Amazon SNS) topic, subscribe an AWS Lambda function to this Amazon SNS topic to process the updates and then store these processed updates in a SQL database running on Amazon EC2 instance
Xem giải thích
Đáp án
*C — Đẩy điểm số vào Kinesis Data Streams, dùng AWS Lambda xử lý, rồi lưu vào DynamoDB
Vì sao đúng
Đề nêu bốn yêu cầu, và kiến trúc này phủ cả bốn:
| Yêu cầu | Thành phần đáp ứng |
|---|---|
| Chịu được đợt tăng đột biến lớn | Kinesis Data Streams đệm dữ liệu |
| Xử lý theo đúng thứ tự nhận | Kinesis bảo đảm thứ tự trong mỗi shard |
| Cơ sở dữ liệu sẵn sàng cao | DynamoDB nhân bản qua ba AZ |
| Ít công quản lý nhất | Cả ba thành phần đều serverless |
Yêu cầu cuối là chỗ phân biệt: Lambda không có máy chủ nào để quản lý và tự co giãn số lần chạy đồng thời theo số shard.
Vì sao các phương án khác sai
- A. Kinesis kèm đội EC2 có Auto Scaling — đúng về thứ tự và khả năng chịu tải, nhưng EC2 đòi bạn quản lý: vá lỗi, giám sát, điều chỉnh quy mô. Vi phạm yêu cầu "ít công quản lý nhất". Đây là phương án nhiễu chính.
- B. SQS kèm đội EC2 — vừa có EC2, vừa dùng standard queue không bảo đảm thứ tự.
- D. SNS kèm Lambda, lưu vào SQL database — SNS không bảo đảm thứ tự và không đệm; nếu Lambda lỗi thì thông điệp mất.
A company maintains its business-critical customer data on an on-premises system in an encrypted format. Over the years, the company has transitioned from using a single encryption key to multiple encryption keys by dividing the data into logical chunks. With the decision to move all the data to an Amazon S3 bucket, the company is now looking for a technique to encrypt each file with a different encryption key to provide maximum security to the migrated on-premises data.
How will you implement this requirement without adding the overhead of splitting the data into logical groups?
-
A
Store the logically divided data into different Amazon S3 buckets. Use server-side encryption with Amazon S3 managed keys (SSE-S3) to encrypt the data
-
B
Configure a single Amazon S3 bucket to hold all data. Use server-side encryption with Amazon S3 managed keys (SSE-S3) to encrypt the data
-
C
Configure a single Amazon S3 bucket to hold all data. Use server-side encryption with AWS KMS (SSE-KMS) and use encryption context to generate a different key for each file/object that you store in the S3 bucket
-
D
Use Multi-Region keys for client-side encryption in the AWS S3 Encryption Client to generate unique keys for each file of data
Xem giải thích
Đáp án
*B — Dùng một bucket S3 duy nhất và mã hoá bằng SSE-S3
Vì sao đúng
Đây là điểm ít người biết nhưng rất đáng nhớ về SSE-S3:
SSE-S3 mã hoá MỖI ĐỐI TƯỢNG bằng một khoá riêng biệt, rồi mã hoá chính khoá đó bằng một khoá gốc mà AWS thường xuyên xoay vòng.
Nghĩa là yêu cầu "mỗi tệp một khoá khác nhau" đã được đáp ứng tự động, không cần bạn làm gì — và đó chính là phần "không thêm gánh nặng chia dữ liệu thành nhóm logic" mà đề nhấn mạnh.
Bạn đổ tất cả vào một bucket, bật SSE-S3, và mỗi tệp có khoá riêng.
Vì sao các phương án khác sai
- C. Dùng SSE-KMS kèm encryption context để sinh khoá khác nhau cho mỗi tệp — hiểu sai vai trò của encryption context: nó là dữ liệu bổ sung để xác thực, gắn vào bản mã và bắt buộc phải khớp khi giải mã. Nó không sinh ra khoá mới. Đây là phương án nhiễu chính, và nghe rất thuyết phục nếu chưa nắm rõ khái niệm.
- A. Chia dữ liệu ra nhiều bucket — chính là việc chia nhóm logic mà đề nói muốn tránh; và SSE-S3 vốn đã cho khoá riêng mỗi đối tượng nên chia bucket chẳng thêm gì.
- D. Dùng Multi-Region key cho mã hoá phía client — Multi-Region key dựng để dùng CÙNG một khoá ở nhiều Region, tức là ngược hẳn mục tiêu "mỗi tệp một khoá".
An e-commerce company wants to develop a click analytics dashboard to see near-real-time user click patterns. The clicks are currently ingested from various devices through Amazon Kinesis Data Streams. The dashboard must be refreshed automatically every ten seconds to display the most updated data. The company is looking for an easy-to-implement solution that can be put into production as soon as possible.
Which solution would you recommend for the given requirements?
-
A
Use Amazon Managed Streaming for Apache Kafka (MSK) to read the data in near-real-time. Develop a custom application for the dashboard by using D3.js
-
B
Use Amazon Kinesis Data Firehose to push data into Amazon S3. Use Amazon QuickSight to build the dashboards from S3 data
-
C
Use Amazon Kinesis Data Firehose to push the data into an Amazon OpenSearch Service. Visualize the data by using OpenSearch (Kibana) dashboards
-
D
Use AWS Glue streaming ETL to store the data into Amazon S3. Use S3 Analytics for analyzing the data
Xem giải thích
Đáp án
*C — Dùng Kinesis Data Firehose đẩy dữ liệu vào Amazon OpenSearch Service, rồi trực quan hoá bằng OpenSearch Dashboards (Kibana)
Vì sao đúng
Đề nêu ba yêu cầu, và OpenSearch là lựa chọn duy nhất thoả cả ba:
- Gần thời gian thực — OpenSearch đánh chỉ mục dữ liệu ngay khi nhận, nên truy vấn được sau vài giây.
- Bảng điều khiển tự làm mới mỗi 10 giây — OpenSearch Dashboards có sẵn tuỳ chọn auto-refresh với chu kỳ ngắn tới vài giây.
- Dễ triển khai, đưa lên sản xuất nhanh — Firehose có tích hợp sẵn với OpenSearch làm đích, và Dashboards đi kèm dịch vụ, không phải viết mã giao diện nào.
Vì sao các phương án khác sai
- B. Firehose đẩy vào S3, dùng QuickSight — hỏng ở độ trễ: Firehose gom lô trước khi ghi ra S3 (tối thiểu 60 giây), và QuickSight không làm mới theo chu kỳ 10 giây. Đây là phương án nhiễu chính vì đường ống này rất phổ biến — chỉ là nó dựng cho phân tích theo lô, không phải bảng điều khiển gần thời gian thực.
- A. Dùng Amazon MSK và tự viết dashboard bằng D3.js — tốn công nhất; đề đòi "dễ triển khai, lên sản xuất càng sớm càng tốt".
- D. Glue streaming ETL rồi dùng "S3 Analytics" — S3 Analytics là công cụ phân tích mẫu truy cập lưu trữ để gợi ý chuyển lớp; nó không phải công cụ phân tích dữ liệu nghiệp vụ.
A data-storage service uses Amazon S3 under the hood to power its storage offerings which allow the customers to upload and view the files immediately. Currently, all the customer files are uploaded directly under a single S3 bucket. The data analytics team has started seeing scalability issues where customer file uploads are failing during peak access hours with more than 5000 requests per second.
Which of the following represents the MOST resource-efficient and cost-optimal way of resolving this issue?
-
A
Change the application architecture to create a new S3 bucket for each day's data and then upload the daily files directly under that day's bucket
-
B
Change the application architecture to create customer-specific custom prefixes within the single bucket and then upload the daily files into those prefixed locations
-
C
Change the application architecture to create a new S3 bucket for each customer and then upload each customer's files directly under the respective buckets
-
D
Change the application architecture to use the S3 Glacier Deep Archive storage class
Xem giải thích
Đáp án
B — Đổi kiến trúc để tạo tiền tố (prefix) riêng cho từng khách hàng trong cùng một bucket
Vì sao đúng
Đây là câu về giới hạn hiệu năng theo tiền tố của S3:
S3 chịu được 3.500 yêu cầu ghi và 5.500 yêu cầu đọc mỗi giây cho mỗi tiền tố phân vùng. Giới hạn này áp cho tiền tố, không phải cho bucket.
Vì tất cả tệp hiện đang nằm trực tiếp dưới gốc bucket, chúng chia nhau một tiền tố duy nhất — nên hệ thống đụng trần ở khoảng 5.000 yêu cầu mỗi giây, đúng con số đề nêu.
Tách theo tiền tố khách hàng thì mỗi tiền tố có hạn mức riêng, và năng lực tổng tăng gần như tuyến tính:
Trước: s3://bucket/file1.dat, file2.dat, ... → 1 tiền tố
Sau: s3://bucket/khach-A/..., s3://bucket/khach-B/... → nhiều tiền tố
Đây cũng là cách tiết kiệm tài nguyên nhất: chỉ đổi quy ước đặt khoá, không tạo thêm bucket nào, không đổi hạ tầng.
Vì sao các phương án khác sai
- C. Tạo một bucket cho mỗi khách hàng — giải quyết được vấn đề nhưng tốn kém về vận hành: mỗi tài khoản mặc định giới hạn 100 bucket (xin tăng lên 1.000), và bạn phải quản lý chính sách, mã hoá, log cho từng cái. Đây là phương án nhiễu chính.
- A. Tạo một bucket cho mỗi ngày — cùng nhược điểm, và còn không giải quyết được vấn đề vì tất cả khách hàng vẫn dồn vào bucket của ngày hôm đó.
- D. Chuyển sang lớp Glacier Deep Archive — làm dữ liệu không truy cập ngay được; đề nói rõ khách hàng phải xem tệp ngay sau khi tải lên.
A global media company uses a fleet of Amazon EC2 instances (behind an Application Load Balancer) to power its video streaming application. To improve the performance of the application, the data engineering team has also created an Amazon CloudFront distribution with the Application Load Balancer as the custom origin. The security team at the company has noticed a spike in the number and types of SQL injection and cross-site scripting attack vectors on the application.
Which of the following solutions would you recommend as the MOST effective in countering these malicious attacks?
-
A
Use security groups with Amazon CloudFront distribution
-
B
Use AWS Config with CloudFront distribution
-
C
Use AWS Web Application Firewall (AWS WAF) with Amazon CloudFront distribution
-
D
Use Amazon Route 53 with Amazon CloudFront distribution
Xem giải thích
Đáp án
C — Dùng AWS WAF gắn với CloudFront distribution
Vì sao đúng
Đề nêu đích danh hai kiểu tấn công — SQL injection và cross-site scripting — và cả hai đều là tấn công tầng 7. AWS WAF là dịch vụ dựng riêng cho chúng.
WAF đọc được nội dung yêu cầu HTTP: đường dẫn, header, cookie, chuỗi truy vấn, thân yêu cầu. Chính khả năng đó cho phép nó nhận ra mã SQL hay đoạn script nhúng trong tham số — thứ mà tầng thấp hơn nhìn vào chỉ thấy một yêu cầu bình thường.
Gắn WAF vào CloudFront (thay vì vào ALB) có lợi thế riêng: yêu cầu độc hại bị chặn ngay tại edge location gần người gửi, không bao giờ chạm tới hạ tầng của bạn.
Bạn cũng không phải tự viết luật: AWS có Managed Rule Groups dựng sẵn cho SQL injection và XSS, được cập nhật liên tục.
Vì sao các phương án khác sai
- A. Dùng security group với CloudFront — không gắn được: CloudFront là dịch vụ biên toàn cầu, không nằm trong VPC nên không có security group. Và dù có thì security group cũng chỉ lọc theo IP và cổng, không đọc được nội dung yêu cầu.
- B. Dùng AWS Config — theo dõi thay đổi cấu hình tài nguyên; nó không lọc lưu lượng.
- D. Dùng Route 53 — dịch vụ DNS; nó phân giải tên miền chứ không kiểm tra nội dung yêu cầu.
An IT company has built a custom data warehousing solution for a large shipping company by using Amazon Redshift. The solution helps the shipping company to analyze the international/domestic cargo transportation details and operational records for the ships. As part of the cost optimizations, the shipping company now wants to move any historical data (any data older than a year) into Amazon S3, as the daily analytical reports consume data for just the last year. However, the data engineers at multiple divisions of the shipping company want to retain the ability to cross-reference this historical data along with the daily reports. The shipping company wants to develop a solution with the LEAST amount of effort and MINIMUM cost.
Which option would you recommend for this requirement?
-
A
Use Redshift Spectrum to create Redshift cluster tables pointing to the underlying historical data in S3. The analytics team can then query this historical data to cross-reference with the daily reports from Redshift
-
B
Set up access to the historical data via Athena. The analytics team can run historical data queries on Athena and continue the daily reporting on Redshift. In case the reports need to be cross-referenced, the analytics team needs to export these in flat files and then do further analysis
-
C
Use the Redshift COPY command to load the S3-based historical data into Redshift. Once the ad-hoc queries are run for the historic data, it can be removed from Redshift
-
D
Use Glue ETL job to load the S3-based historical data into Redshift. Once the ad-hoc queries are run for the historic data, it can be removed from Redshift
Xem giải thích
Đáp án
A — Dùng Redshift Spectrum tạo bảng ngoài trỏ tới dữ liệu lịch sử trên S3
Vì sao đúng
Đề đặt ba ràng buộc, và Redshift Spectrum thoả cả ba:
| Ràng buộc | Vì sao Spectrum đạt |
|---|---|
| Giảm chi phí | Dữ liệu cũ nằm trên S3 rẻ hơn nhiều so với lưu trong cụm Redshift |
| Vẫn đối chiếu chéo được với báo cáo hằng ngày | Spectrum cho JOIN trong cùng một câu SQL giữa bảng ngoài trên S3 và bảng trong Redshift |
| Ít công nhất | Chỉ khai external schema; không cần nạp dữ liệu đi đâu |
Vế thứ hai là chỗ quyết định. Đội phân tích viết một truy vấn duy nhất kiểu:
SELECT ... FROM redshift_bang_nam_nay r
JOIN spectrum_bang_lich_su s ON r.ma_tau = s.ma_tau
mà không cần biết dữ liệu nằm ở hai nơi khác nhau.
Vì sao các phương án khác sai
- B. Dùng Athena cho dữ liệu lịch sử, Redshift cho báo cáo hằng ngày — hai hệ thống tách rời; muốn đối chiếu chéo thì phải xuất dữ liệu từ bên này sang bên kia bằng tay. Đây là phương án nhiễu chính.
- C và D. Nạp dữ liệu lịch sử trở lại Redshift bằng
COPYhoặc Glue rồi xoá sau khi dùng — đi ngược mục tiêu tiết kiệm, và mỗi lần cần đối chiếu lại phải nạp lại; tốn công và tốn tiền.