Ngân hàng đề — AWS Certified Database Specialty

Tìm thấy 358 câu.

Câu 101
A company is using Amazon Aurora PostgreSQL for the backend of its application. The system users are complaining that the responses are slow. A database specialist has determined that the queries to Aurora take longer during peak times. With the Amazon RDS Performance Insights dashboard, the load in the chart for average active sessions is often above the line that denotes maximum CPU usage and the wait state shows that most wait events are IO:XactSync.
What should the company do to resolve these performance issues?
  1. A Add an Aurora Replica to scale the read traffic.
  2. B Scale up the DB instance class.
  3. C Modify applications to commit transactions in batches.
  4. D Modify applications to avoid conflicts by taking locks.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong AWS:
Một công ty sử dụng Amazon Aurora PostgreSQL làm backend cho ứng dụng. Người dùng phàn nàn về phản hồi chậm, đặc biệt vào thời gian cao điểm (peak times). Chuyên gia database sử dụng Amazon RDS Performance Insights để phân tích:

  • Biểu đồ Average Active Sessions (AAS) thường vượt quá đường kẻ chỉ tối đa CPU usage → Chỉ ra workload cao, hệ thống đang bị overload.
  • Wait state chính là IO:XactSync → Đây là wait event đặc trưng trong PostgreSQL/Aurora, đại diện cho thời gian chờ đồng bộ hóa transaction với Write-Ahead Log (WAL) trên storage I/O. Nó xảy ra khi database phải thực hiện fsync (flush dữ liệu synchronous) để đảm bảo tính bền vững dữ liệu, dẫn đến bottleneck I/O writes, đặc biệt với workload có nhiều commit transactions.

Vấn đề cốt lõi: I/O bottleneck do synchronous writes từ commits, không phải CPU thuần túy hay read traffic, dù AAS cao. Giải pháp cần tập trung vào việc tăng capacity I/O và compute để xử lý peak load. (Kiến thức cập nhật AWS 2026: Aurora hỗ trợ Performance Insights với wait events chi tiết hơn, và scaling instance vẫn là best practice cho IO-bound workloads).

📘 Tài liệu tham khảo:

✅ Đáp án đúng: Scale up the DB instance class

Lý do chọn:
Scaling up instance class (ví dụ từ db.r6g.large lên db.r6g.xlarge) sẽ tăng CPU, memory và IOPS provisioned (Aurora sử dụng io2 Block Express với throughput cao hơn theo instance size). IO:XactSync trực tiếp liên quan đến storage I/O limits cho WAL flush, và AAS cao cho thấy cần thêm compute để xử lý concurrent sessions. Đây là giải pháp nhanh nhất, hiệu quả cho write-heavy I/O bound workloads mà không cần thay đổi code. AWS khuyến nghị vertical scaling trước cho Aurora trước khi horizontal.

🛠️ Giải thích tất cả các phương án

  • Add an Aurora Replica to scale the read traffic. ❌ Sai
    Thêm Aurora Replica chỉ scale read traffic (offload SELECT queries), nhưng vấn đề là IO:XactSync từ writes/commits trên primary instance. Replicas không giảm I/O writes trên primary, thậm chí có thể tăng replication lag. Không giải quyết AAS cao trên primary.

  • Scale up the DB instance class. ✅ Đúng
    Như đã giải thích: Tăng instance class cải thiện I/O throughput và CPU, trực tiếp giảm wait IO:XactSync bằng cách cung cấp nhiều IOPS hơn (Aurora tự động provision theo size). Hiệu quả cho peak load, dễ triển khai qua console/CLI.

  • Modify applications to commit transactions in batches. ❌ Sai
    Commit theo batch giảm số lần flush WAL (giúp giảm IO:XactSync), nhưng đây là thay đổi ứng dụng lớn, không đảm bảo giải quyết AAS cao và peak load ngay lập tức. Không phải giải pháp gốc rễ cho hardware limit; AWS ưu tiên infrastructure scaling trước code changes.

  • Modify applications to avoid conflicts by taking locks. ❌ Sai
    Tránh conflicts bằng locks liên quan đến lock contention waits (như ClientLock hoặc LWLock), không phải IO:XactSync (I/O storage). Thay đổi này có thể làm tệ hơn performance do tăng blocking, không liên quan trực tiếp đến wait event được mô tả.

Câu 102
A database specialist deployed an Amazon RDS DB instance in Dev-VPC1 used by their development team. Dev-VPC1 has a peering connection with Dev-VPC2 that belongs to a different development team in the same department. The networking team confirmed that the routing between VPCs is correct; however, the database engineers in Dev-VPC2 are getting a timeout connections error when trying to connect to the database in Dev-VPC1.
What is likely causing the timeouts?
  1. A The database is deployed in a VPC that is in a different Region.
  2. B The database is deployed in a VPC that is in a different Availability Zone.
  3. C The database is deployed with misconfigured security groups.
  4. D The database is deployed with the wrong client connect timeout configuration.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống thực tế trong môi trường AWS:
Một chuyên gia cơ sở dữ liệu (database specialist) đã triển khai một Amazon RDS DB instance nằm trong Dev-VPC1, được đội phát triển sử dụng. Dev-VPC1 có VPC peering connection với Dev-VPC2 (thuộc đội phát triển khác cùng bộ phận).
📡 Nhóm mạng (networking team) đã xác nhận routing giữa hai VPC là đúng (route tables đã được cấu hình peering CIDR blocks chính xác).
Tuy nhiên, kỹ sư cơ sở dữ liệu ở Dev-VPC2 gặp lỗi timeout khi kết nối đến RDS ở Dev-VPC1.

🔍 Vấn đề cốt lõi: Kết nối thất bại dù peering và routing đã OK → Nguyên nhân có thể nằm ở lớp kiểm soát truy cập mạng không phải routing, cụ thể là security groups hoặc NACLs, nhưng câu hỏi tập trung vào timeout kết nối (network-level issue). Đây là tình huống phổ biến khi cross-VPC access RDS qua peering.

✅ Đáp án đúng và lý do lựa chọn

The database is deployed with misconfigured security groups.

🛡️ Lý do chi tiết:

  • Security Groups (SG) của RDS instance hoạt động như stateful firewall kiểm soát inbound/outbound traffic ở mức instance.
  • Khi VPC peering được thiết lập (same Region), traffic từ Dev-VPC2 đến RDS ở Dev-VPC1 sẽ được route đúng, nhưng SG của RDS phải explicitly allow inbound traffic từ CIDR của Dev-VPC2 (ví dụ: port 3306 cho MySQL).
  • Nếu SG chỉ allow từ Dev-VPC1 (self-CIDR) mà không add rule cho peer VPC CIDR, kết nối sẽ timeout ngay lập tức (không phải DNS hay auth error).
  • Đây là nguyên nhân phổ biến nhất theo best practices AWS (updated 2024+ với VPC Flow Logs để troubleshoot). Routing OK loại trừ route tables/NACLs.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ The database is deployed in a VPC that is in a different Region.
    Sai vì: VPC peering chỉ hỗ trợ intra-Region (cùng Region). Nếu khác Region, peering sẽ không thiết lập được ngay từ đầu (AWS báo lỗi). Câu hỏi đã xác nhận peering tồn tại và routing OK → Loại trừ hoàn toàn. Cross-Region cần Transit Gateway hoặc PrivateLink (updated VPC features 2024).

  • ❌ The database is deployed in a VPC that is in a different Availability Zone.
    Sai vì: Availability Zone (AZ) không ảnh hưởng đến kết nối VPC peering. RDS và clients có thể cross-AZ/VPC dễ dàng miễn routing/SG đúng. AZ chỉ liên quan high availability (Multi-AZ RDS), không gây timeout.

  • ✅ The database is deployed with misconfigured security groups.
    Đúng vì: Như giải thích trên, SG là lớp kiểm soát cuối cùng sau routing. Phải add inbound rule: Source = peer VPC CIDR (ví dụ: 10.2.0.0/16), Port = DB port. Không cấu hình → SYN packet bị drop → timeout. AWS khuyến nghị dùng VPC Flow Logs để verify (reject logs).

  • ❌ The database is deployed with the wrong client connect timeout configuration.
    Sai vì: Client timeout (ví dụ: JDBC/ODBC connectTimeout) chỉ là cấu hình ứng dụng-side, điều chỉnh thời gian chờ chứ không gây timeout nếu network block. Vấn đề là network-level block (SG), không phải client config. Nếu là timeout config sai, kết nối vẫn thử nhưng chậm, không phải fail ngay.

📘 Tài liệu tham khảo (AWS cập nhật mới nhất 2024-2026)

🛠️ Khuyến nghị thực tế: Sử dụng AWS Console → RDS → Connectivity & Security → Edit SG → Add inbound rule cho peer VPC. Test bằng telnet/nc từ EC2 ở Dev-VPC2!

Câu 103
A company has a production environment running on Amazon RDS for SQL Server with an in-house web application as the front end. During the last application maintenance window, new functionality was added to the web application to enhance the reporting capabilities for management. Since the update, the application is slow to respond to some reporting queries.
How should the company identify the source of the problem?
  1. A Install and configure Amazon CloudWatch Application Insights for Microsoft .NET and Microsoft SQL Server. Use a CloudWatch dashboard to identify the root cause.
  2. B Enable RDS Performance Insights and determine which query is creating the problem. Request changes to the query to address the problem.
  3. C Use AWS X-Ray deployed with Amazon RDS to track query system traces.
  4. D Create a support request and work with AWS Support to identify the source of the issue.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một tình huống thực tế trong môi trường production: Công ty đang chạy Amazon RDS cho SQL Server làm backend database, kết hợp với web application in-house làm frontend. Trong lần bảo trì gần nhất, họ đã thêm tính năng mới vào web app để cải thiện khả năng reporting cho quản lý. Kết quả là sau update, ứng dụng phản hồi chậm với một số query reporting cụ thể.
📌 Vấn đề cốt lõi: Cần xác định nguồn gốc vấn đề (root cause), đặc biệt liên quan đến hiệu suất query trên RDS SQL Server, vì sự chậm trễ chỉ xảy ra sau khi thêm tính năng reporting mới. Đây là kịch bản điển hình trong DevOps, ưu tiên công cụ monitoring tự động để troubleshoot database performance mà không cần can thiệp thủ công hoặc bên thứ ba.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Enable RDS Performance Insights and determine which query is creating the problem. Request changes to the query to address the problem.

🛠️ Lý do chi tiết:

  • RDS Performance Insights là tính năng chuyên biệt của Amazon RDS (hỗ trợ SQL Server đầy đủ, cập nhật mới nhất đến 2026), giúp monitor và phân tích hiệu suất database ở mức query-level một cách trực quan. Nó cung cấp dashboard với Top SQL queries, wait events, load patterns, và historical data lên đến 2 năm (mặc định 7 ngày miễn phí).
  • Trong trường hợp này, sau update app, query reporting mới có thể là N+1 query, thiếu index, hoặc join kém tối ưu → Performance Insights sẽ xác định chính xác query chậm nhất (qua SQL text, execution time, calls/sec). Sau đó, dev có thể optimize query (ví dụ: thêm index, rewrite SQL).
  • Ưu tiên self-service: Đây là cách nhanh, chi phí thấp (0.01$/i/7ngày sau free tier), không downtime, phù hợp DevOps best practice.
    📘 Tài liệu tham khảo: AWS RDS Performance Insights Docs (cập nhật 2024-2026, hỗ trợ SQL Server Enterprise/Standard).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên tính phù hợp với vấn đề (query chậm trên RDS SQL Server sau app update).

  • ❌ Phương án SAI: Install and configure Amazon CloudWatch Application Insights for Microsoft .NET and Microsoft SQL Server. Use a CloudWatch dashboard to identify the root cause.
    🧠 Giải thích: CloudWatch Application Insights dành cho monitoring toàn bộ ứng dụng .NET/SQL Server (app metrics, logs, anomalies), nhưng không sâu vào database query-level như execution plan hay top SQL chậm. Nó tập trung app-layer (CPU, memory, exceptions), không hiệu quả để pinpoint query reporting cụ thể trên RDS. Phải install agent phức tạp, không phải lựa chọn nhanh cho RDS troubleshooting.

  • ✅ Phương án ĐÚNG: Enable RDS Performance Insights and determine which query is creating the problem. Request changes to the query to address the problem.
    🛠️ Giải thích: Như đã nêu ở phần đáp án đúng, đây là công cụ lý tưởng nhất cho RDS SQL Server, enable chỉ 1-click qua console/CLI/API, dashboard trực quan chỉ ra query chậm (SQL ID, avg latency, waits như IO:LATCH_SH), hỗ trợ drill-down đến query text để fix ngay. Hoàn hảo cho post-update analysis.

  • ❌ Phương án SAI: Use AWS X-Ray deployed with Amazon RDS to track query system traces.
    🧠 Giải thích: AWS X-Ray là service distributed tracing cho ứng dụng (microservices, API calls), không hỗ trợ native deployment trực tiếp trên RDS để trace SQL queries. RDS SQL Server không integrate X-Ray cho database internals (chỉ app-side traces nếu dùng SDK). Không phù hợp identify query chậm, sẽ tốn công setup vô ích và không bao quát DB performance.

  • ❌ Phương án SAI: Create a support request and work with AWS Support to identify the source of the issue.
    🧠 Giải thích: Đây là escalation cuối cùng (Business/Enterprise Support mới hiệu quả), không phải bước đầu tiên theo AWS Well-Architected Framework (Reliability pillar: tự monitor trước). Vấn đề rõ ràng là app query → self-service với Performance Insights nhanh hơn, rẻ hơn, tránh dependency support ticket (có thể mất ngày). Chỉ dùng nếu suspect infrastructure issue (hiếm ở đây).

🔍 Kết luận & Best Practice DevOps

✅ Tóm tắt: Sử dụng RDS Performance Insights là cách tối ưu, nhanh chóng, chi phí thấp để solve vấn đề query chậm sau app update. Trong DevOps Professional, luôn ưu tiên monitoring tự động (CloudWatch + PI) trước khi escalate.
🛤️ Khuyến nghị bổ sung: Kết hợp RDS Parameter Groups (tối ưu SQL Server settings), Query Store (native SQL Server feature trên RDS), và CloudWatch Logs Insights cho query logs nếu cần sâu hơn. Test trên staging trước production!
📘 Nguồn thêm: AWS DevOps Best Practices, RDS Monitoring.

Câu 104
An electric utility company wants to store power plant sensor data in an Amazon DynamoDB table. The utility company has over 100 power plants and each power plant has over 200 sensors that send data every 2 seconds. The sensor data includes time with milliseconds precision, a value, and a fault attribute if the sensor is malfunctioning. Power plants are identified by a globally unique identifier. Sensors are identified by a unique identifier within each power plant. A database specialist needs to design the table to support an efficient method of finding all faulty sensors within a given power plant.
Which schema should the database specialist use when creating the DynamoDB table to achieve the fastest query time when looking for faulty sensors?
  1. A Use the plant identifier as the partition key and the measurement time as the sort key. Create a global secondary index (GSI) with the plant identifier as the partition key and the fault attribute as the sort key.
  2. B Create a composite of the plant identifier and sensor identifier as the partition key. Use the measurement time as the sort key. Create a local secondary index (LSI) on the fault attribute.
  3. C Create a composite of the plant identifier and sensor identifier as the partition key. Use the measurement time as the sort key. Create a global secondary index (GSI) with the plant identifier as the partition key and the fault attribute as the sort key.
  4. D Use the plant identifier as the partition key and the sensor identifier as the sort key. Create a local secondary index (LSI) on the fault attribute.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi tập trung vào việc thiết kế schema cho bảng Amazon DynamoDB để lưu trữ dữ liệu cảm biến từ hơn 100 nhà máy điện (power plants), mỗi nhà máy có hơn 200 cảm biến gửi dữ liệu mỗi 2 giây. Dữ liệu bao gồm: thời gian chính xác đến mili giây, giá trị đo, và thuộc tính fault (nếu cảm biến hỏng).

  • Mục tiêu chính: Tìm tất cả các cảm biến bị lỗi (faulty sensors) trong một nhà máy điện cụ thể một cách hiệu quả nhất (fastest query time).
  • Thách thức: Tốc độ ghi cao (high write throughput), cần query nhanh theo plant_id, chỉ lấy các sensor có fault=true (giả sử fault là boolean hoặc string "true"/"false"), và xác định sensor unique trong plant (plant_id global unique, sensor_id unique within plant).
  • Giả định thiết kế: Để hỗ trợ trạng thái hiện tại (current state) của cảm biến, schema nên ghi đè (overwrite) dữ liệu mới nhất cho mỗi sensor thay vì lưu lịch sử đầy đủ, vì query tập trung vào "faulty sensors" (có lẽ là trạng thái hiện tại bị lỗi). DynamoDB yêu cầu primary key (PK + SK) unique, nên SK cố định per sensor để overwrite.
    ✅ Phiên bản AWS mới nhất (2026): DynamoDB hỗ trợ LSI/GSI với projection tối ưu, adaptive capacity, on-demand mode cho workload cao. Best practice: Sử dụng LSI cho access pattern trong cùng partition để tiết kiệm chi phí và nhanh hơn GSI (không cần provision capacity riêng).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use the plant identifier as the partition key and the sensor identifier as the sort key. Create a local secondary index (LSI) on the fault attribute.

Lý do chọn 🛠️:

  • PK = plant_id: Phân vùng theo nhà máy, dễ query toàn bộ sensors trong 1 plant (partition nhỏ ~200 items).
  • SK = sensor_id (unique trong plant): Cho phép ghi đè dữ liệu mới nhất mỗi 2s (same PK+SK → update value/time/fault), chỉ giữ trạng thái hiện tại per sensor → phù hợp query "faulty sensors hiện tại".
  • LSI với SK = fault: Query LSI bằng PK=plant_id + SK="true" → trả trực tiếp danh sách sensors bị lỗi, sorted theo fault (tất cả "true" cùng nhau), không cần FilterExpression (nhanh hơn, ít RCU). LSI chia sẻ PK với base table, nhanh tương đương base, rẻ hơn GSI.
  • Fastest query: O(1) access pattern, partition cân bằng (100 plants), throughput cao (~10k writes/s toàn hệ thống).
    📘 Tài liệu tham khảo:
  • DynamoDB Best Practices (AWS 2026).
  • LSI vs GSI.

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, đánh dấu ✅/❌, và giải thích bằng tiếng Việt với lý do đúng/sai dựa trên access pattern, hiệu suất query, và workload.

  • ❌ [SAI] Use the plant identifier as the partition key and the measurement time as the sort key. Create a global secondary index (GSI) with the plant identifier as the partition key and the fault attribute as the sort key.
    Giải thích sai 🚫: Base table (PK=plant, SK=time) lưu toàn bộ lịch sử (new item mỗi 2s → hàng triệu items/plant nhanh chóng). GSI (PK=plant, SK=fault) query PK=plant + SK="true" trả tất cả historical faulty measurements (duplicate per sensor), không chỉ sensors hiện tại. GSI tốn kém hơn (extra WCUs/RCUs), chậm hơn do eventually consistent. Không overwrite → query không efficient cho "current faulty sensors".

  • ❌ [SAI] Create a composite of the plant identifier and sensor identifier as the partition key. Use the measurement time as the sort key. Create a local secondary index (LSI) on the fault attribute.
    Giải thích sai 🚫: PK=plant#sensor (composite) + SK=time → lưu time series per sensor (tốt cho history). Nhưng LSI chia sẻ PK=plant#sensor, query LSI bắt buộc chỉ định full PK (plant#sensor cụ thể) → không thể query across all sensors trong plant. Để tìm faulty sensors toàn plant, phải query từng sensor riêng → chậm, không scale với 200+ sensors.

  • ❌ [SAI] Create a composite of the plant identifier and sensor identifier as the partition key. Use the measurement time as the sort key. Create a global secondary index (GSI) with the plant identifier as the partition key and the fault attribute as the sort key.
    Giải thích sai 🚫: Base giống trên (time series per sensor). GSI (PK=plant, SK=fault) query PK=plant + SK="true" trả tất cả historical faults across sensors (nhiều duplicate), không chỉ current. Hot partition nếu nhiều faults, GSI tốn RCU/WCU cao do write amplification (mỗi update base → update GSI). Không overwrite → không fastest cho current status.

  • ✅ [ĐÚNG] Use the plant identifier as the partition key and the sensor identifier as the sort key. Create a local secondary index (LSI) on the fault attribute.
    Giải thích đúng 🎯: Như phần đáp án trên. Hoàn hảo cho overwrite latest state, query LSI trực tiếp lấy danh sách sensors faulty hiện tại mà không filter/scan, nhanh nhất (low latency, minimal RCU). Scale tốt với workload cao đến 2026.

🧠 Tóm tắt insight: Thiết kế DynamoDB ưu tiên access patterns (query faulty per plant) + single-table design với index phù hợp. LSI lý tưởng cho intra-partition queries!

Câu 105
A company is releasing a new mobile game featuring a team play mode. As a group of mobile device users play together, an item containing their statuses is updated in an Amazon DynamoDB table. Periodically, the other users' devices read the latest statuses of their teammates from the table using the BatchGetltemn operation.
Prior to launch, some testers submitted bug reports claiming that the status data they were seeing in the game was not up-to-date. The developers are unable to replicate this issue and have asked a database specialist for a recommendation.
Which recommendation would resolve this issue?
  1. A Ensure the DynamoDB table is configured to be always consistent.
  2. B Ensure the BatchGetltem operation is called with the ConsistentRead parameter set to false.
  3. C Enable a stream on the DynamoDB table and subscribe each device to the stream to ensure all devices receive up-to-date status information.
  4. D Ensure the BatchGetltem operation is called with the ConsistentRead parameter set to true.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một công ty phát hành game di động với chế độ chơi nhóm (team play mode). Khi nhóm người chơi cùng tham gia, một item chứa trạng thái (statuses) của họ được cập nhật vào bảng Amazon DynamoDB. Các thiết bị của người chơi khác định kỳ đọc trạng thái mới nhất của đồng đội từ bảng này bằng hoạt động BatchGetItem.

Trước khi ra mắt, một số tester báo lỗi: dữ liệu trạng thái họ thấy trong game không cập nhật kịp thời (not up-to-date). Nhà phát triển không tái tạo được vấn đề và yêu cầu chuyên gia cơ sở dữ liệu đưa ra khuyến nghị để giải quyết.

Vấn đề cốt lõi: DynamoDB mặc định sử dụng eventually consistent reads (đọc cuối cùng nhất quán), nghĩa là dữ liệu đọc có thể chưa phản ánh ngay lập tức các thay đổi mới nhất (có độ trễ vài giây). Trong game thời gian thực như team play, điều này gây ra tình trạng status "lỗi thời". Giải pháp cần đảm bảo strongly consistent reads để dữ liệu luôn mới nhất tại thời điểm đọc (theo tài liệu AWS DynamoDB cập nhật 2024-2026, strongly consistent reads có latency cao hơn ~2x nhưng đảm bảo tính nhất quán ngay lập tức).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Ensure the BatchGetltem operation is called with the ConsistentRead parameter set to true.

Lý do:

  • Hoạt động BatchGetItem trong DynamoDB hỗ trợ tham số ConsistentRead (mặc định là false → eventually consistent).
  • Đặt ConsistentRead=true kích hoạt strongly consistent read, đảm bảo dữ liệu đọc phản ánh mọi thay đổi đã commit ngay lập tức, giải quyết triệt để vấn đề status không up-to-date.
  • Đây là giải pháp đơn giản, không thay đổi cấu trúc bảng, phù hợp với workload đọc định kỳ cao trong game. AWS khuyến nghị sử dụng cho các trường hợp yêu cầu tính nhất quán cao như real-time status (theo AWS Well-Architected Framework - Reliability Pillar, 2025).

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Tôi đánh dấu ✅ cho đúng, ❌ cho sai, kèm giải thích chi tiết bằng tiếng Việt dựa trên kiến thức DynamoDB mới nhất (2026).

  • ❌ [SAI] Ensure the DynamoDB table is configured to be always consistent.
    Bảng DynamoDB không có tùy chọn cấu hình "always consistent" ở mức table. Tính nhất quán được kiểm soát tại mức operation (như GetItem, BatchGetItem) qua tham số ConsistentRead. Cấu hình table chỉ liên quan RCU/WCU, streams, hoặc global tables – không giải quyết vấn đề đọc dữ liệu lỗi thời.

  • ❌ [SAI] Ensure the BatchGetltem operation is called with the ConsistentRead parameter set to false.
    Tham số ConsistentRead=false là mặc định, dẫn đến eventually consistent reads với độ trễ (stale data lên đến vài giây). Việc đảm bảo đặt false sẽ làm vấn đề tệ hơn, không phải giải pháp – trái ngược hoàn toàn với nhu cầu up-to-date.

  • ❌ [SAI] Enable a stream on the DynamoDB table and subscribe each device to the stream to ensure all devices receive up-to-date status information.
    DynamoDB Streams dùng để capture thay đổi dữ liệu (change data capture - CDC) cho Lambda, Kinesis, hoặc replication, không phải cho direct reads từ thiết bị di động. Subscribe từ device sẽ phức tạp, tốn kém (chi phí stream + polling), và không đảm bảo real-time reads từ table gốc – vi phạm best practice AWS cho mobile apps (dùng cho async processing, không sync reads).

  • ✅ [ĐÚNG] Ensure the BatchGetltem operation is called with the ConsistentRead parameter set to true.
    Như đã giải thích ở trên: Kích hoạt strongly consistent reads trực tiếp trên BatchGetItem, đảm bảo status luôn mới nhất. Hiệu suất: Latency cao hơn nhưng RCU tính phí gấp đôi, phù hợp workload game nhỏ (theo AWS docs 2026).

🛠️ Khuyến nghị bổ sung từ AWS Certified DevOps Engineer Professional

  • Tối ưu hóa: Kết hợp với DynamoDB Accelerator (DAX) cho caching reads nếu latency là vấn đề, nhưng ưu tiên ConsistentRead=true trước.
  • Monitoring: Sử dụng CloudWatch Metrics (ConsumedReadCapacityUnits, ConsistentRead requests) để theo dõi.
  • Test: Trong môi trường dev, dùng DynamoDB Local để replicate issue với ConsistentRead=false.

📘 Tài liệu tham khảo (AWS cập nhật 2024-2026)

Câu 106
A company is running an Amazon RDS for MySQL Multi-AZ DB instance for a business-critical workload. RDS encryption for the DB instance is disabled. A recent security audit concluded that all business-critical applications must encrypt data at rest. The company has asked its database specialist to formulate a plan to accomplish this for the DB instance.
Which process should the database specialist recommend?
  1. A Create an encrypted snapshot of the unencrypted DB instance. Copy the encrypted snapshot to Amazon S3. Restore the DB instance from the encrypted snapshot using Amazon S3.
  2. B Create a new RDS for MySQL DB instance with encryption enabled. Restore the unencrypted snapshot to this DB instance.
  3. C Create a snapshot of the unencrypted DB instance. Create an encrypted copy of the snapshot. Restore the DB instance from the encrypted snapshot.
  4. D Temporarily shut down the unencrypted DB instance. Enable AWS KMS encryption in the AWS Management Console using an AWS managed CMK. Restart the DB instance in an encrypted state.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi xoay quanh một tình huống thực tế trong AWS RDS: Một công ty đang sử dụng Amazon RDS for MySQL Multi-AZ DB instance cho workload kinh doanh quan trọng (business-critical), nhưng RDS encryption (mã hóa dữ liệu tại chỗ - data at rest) bị tắt. Kết quả từ một cuộc kiểm toán bảo mật (security audit) yêu cầu tất cả ứng dụng kinh doanh quan trọng phải mã hóa dữ liệu tại chỗ. Nhiệm vụ của database specialist là đề xuất quy trình (process) để kích hoạt mã hóa cho DB instance này mà không làm gián đoạn dịch vụ quá nhiều.

🛠️ Điểm quan trọng cần lưu ý:

  • RDS không hỗ trợ bật mã hóa trực tiếp trên DB instance đang chạy (running instance), đặc biệt với instance đã tạo mà không có encryption.
  • Multi-AZ chỉ đảm bảo high availability (HA), không ảnh hưởng đến encryption.
  • Giải pháp phải tuân thủ best practice AWS: Sử dụng snapshot để migrate dữ liệu sang instance mới có encryption.
  • Kiến thức cập nhật đến 2026: AWS vẫn giữ nguyên quy trình này (không có tính năng modify encryption in-place cho RDS MySQL), theo docs RDS User Guide mới nhất.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create a snapshot of the unencrypted DB instance. Create an encrypted copy of the snapshot. Restore the DB instance from the encrypted snapshot.

Lý do chọn đáp án này 🟢:

  • Đây là quy trình chuẩn (recommended process) của AWS để kích hoạt encryption cho RDS instance không được mã hóa từ đầu.
  • Bước 1: Tạo snapshot từ instance gốc (unencrypted) → Không ảnh hưởng đến instance đang chạy.
  • Bước 2: Copy snapshot thành encrypted copy (sử dụng KMS key) → Snapshot mới sẽ được mã hóa.
  • Bước 3: Restore từ encrypted snapshot → Tạo DB instance mới hoàn toàn mã hóa, sau đó cutover (chuyển traffic) từ instance cũ sang mới với downtime tối thiểu.
  • Ưu điểm: An toàn, không modify instance gốc, hỗ trợ Multi-AZ, và scale được cho production. Phù hợp với audit yêu cầu mã hóa data at rest (EBS volumes, snapshots).

📋 Phân tích tất cả các phương án (đúng/sai)

Dưới đây là phân tích chi tiết từng lựa chọn. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt với lý do đúng/sai dựa trên docs AWS.

  • ❌ Phương án SAI: Create an encrypted snapshot of the unencrypted DB instance. Copy the encrypted snapshot to Amazon S3. Restore the DB instance from the encrypted snapshot using Amazon S3.
    Giải thích sai: Không thể tạo encrypted snapshot trực tiếp từ unencrypted DB instance (RDS chỉ cho phép copy snapshot để encrypt). Hơn nữa, RDS không hỗ trợ restore từ S3 (S3 chỉ dùng cho export/import data, không phải snapshot restore). Quy trình này không tồn tại và sẽ fail.

  • ❌ Phương án SAI: Create a new RDS for MySQL DB instance with encryption enabled. Restore the unencrypted snapshot to this DB instance.
    Giải thích sai: Không thể restore unencrypted snapshot vào encrypted DB instance. RDS yêu cầu snapshot và target instance phải khớp encryption status (unencrypted snapshot chỉ restore vào unencrypted instance). Điều này vi phạm quy tắc AWS, dẫn đến lỗi "incompatible snapshot".

  • ✅ Phương án ĐÚNG: Create a snapshot of the unencrypted DB instance. Create an encrypted copy of the snapshot. Restore the DB instance from the encrypted snapshot.
    Giải thích đúng: Như đã phân tích ở phần đáp án. Đây là blueprint chính thức từ AWS: Snapshot → Encrypted copy → Restore. Hỗ trợ KMS customer-managed keys, Multi-AZ, và zero-downtime migration qua Read Replica nếu cần.

  • ❌ Phương án SAI: Temporarily shut down the unencrypted DB instance. Enable AWS KMS encryption in the AWS Management Console using an AWS managed CMK. Restart the DB instance in an encrypted state.
    Giải thích sai: Hoàn toàn không thể! RDS không hỗ trợ modify encryption trên existing instance (dù shutdown). Encryption là immutable property khi tạo instance. Thao tác này sẽ báo lỗi trong Console/CLI, gây downtime vô ích và không đạt yêu cầu.

📘 Tài liệu tham khảo (AWS Official Docs - cập nhật 2026)

🛡️ Lời khuyên DevOps: Trong thực tế, kết hợp với AWS DMS hoặc Read Replica để zero-downtime migration cho Multi-AZ. Test trước trên staging!

Câu 107 Chọn nhiều đáp án
A company is migrating its on-premises database workloads to the AWS Cloud. A database specialist performing the move has chosen AWS DMS to migrate an
Oracle database with a large table to Amazon RDS. The database specialist notices that AWS DMS is taking significant time to migrate the data.
Which actions would improve the data migration speed? (Choose three.)
  1. A Create multiple AWS DMS tasks to migrate the large table.
  2. B Configure the AWS DMS replication instance with Multi-AZ.
  3. C Increase the capacity of the AWS DMS replication server.
  4. D Establish an AWS Direct Connect connection between the on-premises data center and AWS.
  5. E Enable an Amazon RDS Multi-AZ configuration.
  6. F Enable full large binary object (LOB) mode to migrate all LOB data for all large tables.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc tối ưu hóa tốc độ di chuyển dữ liệu từ cơ sở dữ liệu Oracle on-premises sang Amazon RDS bằng AWS Database Migration Service (AWS DMS). Một công ty đang migrate workload database lớn, và chuyên viên nhận thấy DMS mất nhiều thời gian với một bảng lớn (large table). Câu hỏi yêu cầu chọn ba hành động để cải thiện tốc độ migrate.

🛠️ Ngữ cảnh chính:

  • AWS DMS sử dụng replication instance (máy chủ trung gian) để đọc dữ liệu từ source (Oracle on-premises) và ghi vào target (RDS).
  • Các yếu tố ảnh hưởng tốc độ: bandwidth mạng, tài nguyên replication instance (CPU/RAM), cách chia nhỏ dữ liệu (như parallel tasks), và chế độ xử lý LOB (nếu có dữ liệu lớn).
  • Không liên quan đến high availability (HA) như Multi-AZ, vì HA ưu tiên độ tin cậy chứ không phải tốc độ.

Dựa trên tài liệu AWS DMS mới nhất (cập nhật đến 2026, phiên bản DMS 3.4+), các best practice nhấn mạnh parallelization, scale instance, và low-latency network.

✅ Đáp án đúng (Chọn THREE)

Các hành động đúng là những phương án cải thiện trực tiếp throughput và parallelism của DMS:

  1. Create multiple AWS DMS tasks to migrate the large table.
    🧩 Lý do: Chia bảng lớn thành nhiều task song song (parallel load) bằng table mapping rules, tăng tốc độ đáng kể (có thể lên 3-5x tùy kích thước).

  2. Increase the capacity of the AWS DMS replication server.
    🧩 Lý do: Replication instance lớn hơn (ví dụ từ c5.large lên c5.4xlarge) cung cấp CPU/RAM cao hơn, xử lý I/O nhanh hơn, phù hợp cho full load lớn.

  3. Establish an AWS Direct Connect connection between the on-premises data center and AWS.
    🧩 Lý do: Direct Connect cung cấp dedicated bandwidth (1-100 Gbps) và latency thấp (<10ms), tránh nghẽn internet VPN/public, tăng tốc transfer dữ liệu lớn.

📋 Giải thích TẤT CẢ các phương án (Đúng/SAI)

Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên văn bản gốc tiếng Anh của phương án, chỉ giải thích bằng tiếng Việt với emoji nổi bật.

  • ✅ Create multiple AWS DMS tasks to migrate the large table.
    Đúng: AWS DMS hỗ trợ parallel tasks qua "table mappings" (ví dụ: partition theo primary key hoặc range). Điều này phân tải dữ liệu lớn thành nhiều luồng song song, giảm thời gian migrate từ hàng giờ xuống phút. Best practice cho large tables (>100GB).

  • ❌ Configure the AWS DMS replication instance with Multi-AZ.
    Sai: Multi-AZ chỉ tăng HA (sync dữ liệu giữa 2 AZ), nhưng gây overhead replication nội bộ, tăng latency và giảm throughput (AWS docs khuyến cáo tắt Multi-AZ cho migrate nhanh). Không cải thiện tốc độ.

  • ✅ Increase the capacity of the AWS DMS replication server.
    Đúng: Replication instance là "động cơ" chính của DMS. Scale up (tăng vCPU/RAM, ví dụ dms.c5.4xlarge) xử lý extract/transform/load nhanh hơn, đặc biệt với bảng lớn. AWS hỗ trợ auto-scale ở một số instance loại mới (2025+).

  • ✅ Establish an AWS Direct Connect connection between the on-premises data center and AWS.
    Đúng: Kết nối internet/VPN thường bottleneck (latency 100-200ms, bandwidth chia sẻ). Direct Connect đảm bảo throughput ổn định cao, lý tưởng cho petabyte-scale migration.

  • ❌ Enable an Amazon RDS Multi-AZ configuration.
    Sai: Multi-AZ trên RDS chỉ replicate standby instance cho failover, không tăng tốc độ write từ DMS (thậm chí chậm hơn do sync). Target RDS chỉ cần Single-AZ cho migrate nhanh.

  • ❌ Enable full large binary object (LOB) mode to migrate all LOB data for all large tables.
    Sai: Full LOB mode migrate từng LOB riêng lẻ, rất chậm với dữ liệu lớn (BLOB/CLOB). Nên dùng limited LOB hoặc segmented LOB mode (mới hơn từ 2024) để chunk và parallel, tăng tốc 10x.

📘 Tài liệu tham khảo (AWS cập nhật 2026)

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm ví dụ config, hãy hỏi nhé.

Câu 108 Chọn nhiều đáp án
A company is migrating a mission-critical 2-TB Oracle database from on premises to Amazon Aurora. The cost for the database migration must be kept to a minimum, and both the on-premises Oracle database and the Aurora DB cluster must remain open for write traffic until the company is ready to completely cut over to Aurora.
Which combination of actions should a database specialist take to accomplish this migration as quickly as possible? (Choose two.)
  1. A Use the AWS Schema Conversion Tool (AWS SCT) to convert the source database schema. Then restore the converted schema to the target Aurora DB cluster.
  2. B Use Oracle's Data Pump tool to export a copy of the source database schema and manually edit the schema in a text editor to make it compatible with Aurora.
  3. C Create an AWS DMS task to migrate data from the Oracle database to the Aurora DB cluster. Select the migration type to replicate ongoing changes to keep the source and target databases in sync until the company is ready to move all user traffic to the Aurora DB cluster.
  4. D Create an AWS DMS task to migrate data from the Oracle database to the Aurora DB cluster. Once the initial load is complete, create an AWS Kinesis Data Firehose stream to perform change data capture (CDC) until the company is ready to move all user traffic to the Aurora DB cluster.
  5. E Create an AWS Glue job and related resources to migrate data from the Oracle database to the Aurora DB cluster. Once the initial load is complete, create an AWS DMS task to perform change data capture (CDC) until the company is ready to move all user traffic to the Aurora DB cluster.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc di chuyển (migrate) một cơ sở dữ liệu Oracle mission-critical dung lượng 2-TB từ on-premises sang Amazon Aurora với các yêu cầu nghiêm ngặt:

  • Chi phí phải tối thiểu (keep the cost to a minimum).
  • Cả hai DB (source Oracle và target Aurora) phải giữ mở cho write traffic cho đến khi công ty sẵn sàng cutover hoàn toàn (chuyển toàn bộ traffic sang Aurora).
  • Cần thực hiện nhanh nhất có thể (as quickly as possible).
  • Chọn TWO actions kết hợp từ database specialist.

🔍 Bối cảnh kỹ thuật:

  • Oracle on-prem là source lớn (2-TB), cần migrate schema + data.
  • Aurora là managed DB (MySQL hoặc PostgreSQL compatible), không native Oracle, nên cần công cụ convert schema và migrate data với CDC (Change Data Capture) để sync ongoing changes (đồng bộ thay đổi liên tục) mà không downtime.
  • Best practice AWS (cập nhật 2024-2026): Sử dụng AWS SCT (Schema Conversion Tool) cho schema conversion và AWS DMS (Database Migration Service) cho data migration full load + CDC, hỗ trợ Oracle → Aurora PostgreSQL/MySQL với chi phí thấp (pay-per-use, no upfront).

✅ Đáp án đúng (Chọn TWO)

Hai phương án đúng là:

  1. Use the AWS Schema Conversion Tool (AWS SCT) to convert the source database schema. Then restore the converted schema to the target Aurora DB cluster.
    Lý do: AWS SCT tự động convert schema Oracle sang Aurora compatible (PostgreSQL/MySQL), nhanh, chính xác, chi phí thấp. Sau convert, restore schema vào Aurora trước khi migrate data.

  2. Create an AWS DMS task to migrate data from the Oracle database to the Aurora DB cluster. Select the migration type to replicate ongoing changes to keep the source and target databases in sync until the company is ready to move all user traffic to the Aurora DB cluster.
    Lý do: DMS hỗ trợ Full Load + CDC cho Oracle → Aurora, sync ongoing changes (writes) realtime, giữ cả hai DB active đến cutover. Nhanh cho 2-TB, chi phí tối thiểu (serverless option từ 2023+).

🛠️ Kết hợp hai actions: SCT convert schema trước → DMS migrate data + CDC → Cutover khi lag thấp (monitor DMS metrics). Thời gian nhanh nhờ parallel load và CDC low-latency.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (Đúng) hoặc ❌ (Sai), với lý do bằng tiếng Việt:

  • ✅ Use the AWS Schema Conversion Tool (AWS SCT) to convert the source database schema. Then restore the converted schema to the target Aurora DB cluster.
    🧩 Đúng vì: SCT là công cụ AWS chính thức, tự động convert schema Oracle phức tạp sang Aurora (hỗ trợ 70-90% tự động, fix manual nếu cần). Nhanh, chi phí thấp (free tool), integrate với DMS. Không dùng manual edit để tránh lỗi.

  • ❌ Use Oracle's Data Pump tool to export a copy of the source database schema and manually edit the schema in a text editor to make it compatible with Aurora.
    🧩 Sai vì: Data Pump chỉ export schema/data Oracle native, không compatible trực tiếp với Aurora (cần edit thủ công lớn cho 2-TB → chậm, lỗi-prone, không scale). Vi phạm "quickly" và "minimum cost" (manual labor cao). Không phải AWS best practice.

  • ✅ Create an AWS DMS task to migrate data from the Oracle database to the Aurora DB cluster. Select the migration type to replicate ongoing changes to keep the source and target databases in sync until the company is ready to move all user traffic to the Aurora DB cluster.
    🧩 Đúng vì: DMS task với Migrate existing data + Ongoing replication (CDC) chính xác sync writes realtime (LogMiner/Debezium cho Oracle). Hỗ trợ Aurora PostgreSQL/MySQL, monitorable via CloudWatch, cutover nhanh (dừng DMS khi lag=0).

  • ❌ Create an AWS DMS task to migrate data from the Oracle database to the Aurora DB cluster. Once the initial load is complete, create an AWS Kinesis Data Firehose stream to perform change data capture (CDC) until the company is ready to move all user traffic to the Aurora DB cluster.
    🧩 Sai vì: DMS chỉ cho initial load, nhưng Kinesis Data Firehose không hỗ trợ CDC từ Oracle (Firehose là streaming cho logs/data lakes, không parse Oracle redo logs). Phức tạp, chi phí cao hơn DMS CDC native, không "quickly".

  • ❌ Create an AWS Glue job and related resources to migrate data from the Oracle database to the Aurora DB cluster. Once the initial load is complete, create an AWS DMS task to perform change data capture (CDC) until the company is ready to move all user traffic to the Aurora DB cluster.
    🧩 Sai vì: AWS Glue là ETL cho big data/batch, không optimize cho DB migration realtime (chậm cho 2-TB transactional DB). DMS CDC sau initial không khớp (Glue không sync schema/data tốt), tăng chi phí/complexity. DMS full lifecycle tốt hơn.

📘 Tài liệu tham khảo (Cập nhật AWS 2024-2026)

Hy vọng phân tích này giúp bạn ôn thi DOP-C02 hiệu quả! 🚀 Nếu cần thêm chi tiết, hỏi nhé!

Câu 109
A company has a 20 TB production Amazon Aurora DB cluster. The company runs a large batch job overnight to load data into the Aurora DB cluster. To ensure the company's development team has the most up-to-date data for testing, a copy of the DB cluster must be available in the shortest possible time after the batch job completes.
How should this be accomplished?
  1. A Use the AWS CLI to schedule a manual snapshot of the DB cluster. Restore the snapshot to a new DB cluster using the AWS CLI.
  2. B Create a dump file from the DB cluster. Load the dump file into a new DB cluster.
  3. C Schedule a job to create a clone of the DB cluster at the end of the overnight batch process.
  4. D Set up a new daily AWS DMS task that will use cloning and change data capture (CDC) on the DB cluster to copy the data to a new DB cluster. Set up a time for the AWS DMS stream to stop when the new cluster is current.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào một cụm Amazon Aurora DB cluster dung lượng lớn 20 TB đang chạy trong môi trường production. Công ty thực hiện một batch job lớn vào ban đêm để tải dữ liệu vào cluster này. Yêu cầu chính là cung cấp một bản copy của DB cluster cho đội ngũ phát triển (dev team) để testing, với thời gian ngắn nhất có thể ngay sau khi batch job hoàn tất.

🔑 Thách thức chính:

  • Dung lượng dữ liệu khổng lồ (20 TB) nên cần phương pháp sao chép nhanh chóng, hiệu quả, tránh downtime hoặc thời gian chờ đợi lâu.
  • Phải đảm bảo dữ liệu up-to-date nhất sau batch job.
  • Giải pháp phải tự động hóa được (schedule) để phù hợp với quy trình overnight.

Mục tiêu là chọn phương pháp tối ưu về tốc độ trên AWS Aurora (phiên bản mới nhất 2024-2026 hỗ trợ cloning native siêu nhanh).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Schedule a job to create a clone of the DB cluster at the end of the overnight batch process.

Lý do chi tiết 🛠️:

  • Aurora Clone là tính năng native của Amazon Aurora (từ năm 2019 và được cải tiến liên tục đến 2026), cho phép tạo bản clone tức thì chỉ bằng cách copy metadata (write-ahead log và storage layer). Dữ liệu thực tế (20 TB) được shared storage giữa source và clone cho đến khi có thay đổi (copy-on-write), nên thời gian tạo clone chỉ vài giây đến vài phút, không phụ thuộc kích thước dữ liệu.
  • Có thể schedule job (qua AWS Lambda, EventBridge, hoặc CLI) ngay cuối batch job để đảm bảo dữ liệu up-to-date.
  • Clone là read-only ban đầu nhưng có thể promote thành writable cho dev team testing ngay lập tức.
  • Tiết kiệm chi phí (không copy full data ngay) và không ảnh hưởng production cluster.
  • Phù hợp nhất với yêu cầu "shortest possible time".

📘 Tài liệu tham khảo:

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm giải thích bằng tiếng Việt.

  • ❌ Use the AWS CLI to schedule a manual snapshot of the DB cluster. Restore the snapshot to a new DB cluster using the AWS CLI.
    Phương án này sai vì snapshot Aurora mất thời gian tạo (snapshot time proportional to size) khoảng 15-30 phút cho 20 TB (tùy IOPS), cộng thêm thời gian restore (có thể hàng giờ). Không phải "shortest time". Snapshot là point-in-time, nhưng manual schedule qua CLI phức tạp và chậm hơn clone. Không khuyến nghị cho large clusters.

  • ❌ Create a dump file from the DB cluster. Load the dump file into a new DB cluster.
    Phương án này sai hoàn toàn vì dump file (qua mysqldump hoặc pg_dump) phải export toàn bộ 20 TB (mất hàng giờ/ngày), rồi import vào cluster mới (thêm hàng giờ nữa). Rất chậm, tốn bandwidth/storage, dễ lỗi với dữ liệu lớn, và không tự động hóa dễ dàng. Không phù hợp production-scale.

  • ✅ Schedule a job to create a clone of the DB cluster at the end of the overnight batch process.
    Như đã giải thích ở trên: Nhanh nhất (seconds-minutes), native Aurora, shared storage, schedule dễ dàng. Hoàn hảo cho yêu cầu!

  • ❌ Set up a new daily AWS DMS task that will use cloning and change data capture (CDC) on the DB cluster to copy the data to a new DB cluster. Set up a time for the AWS DMS stream to stop when the new cluster is current.
    Phương án này sai vì AWS DMS (Database Migration Service) dùng cho migration/replication liên tục, không phải copy nhanh one-time. "Cloning + CDC" không phải tính năng chuẩn DMS cho Aurora (DMS hỗ trợ CDC nhưng setup phức tạp, lag time, overhead cao). Với 20 TB, initial load mất lâu; không "shortest time" so với native clone. DMS phù hợp ongoing sync, không phải post-batch copy.

🧠 Kết luận: Clone là giải pháp DevOps optimal cho Aurora large-scale, giảm thời gian từ giờ xuống phút! 🚀

Câu 110
A company has two separate AWS accounts: one for the business unit and another for corporate analytics. The company wants to replicate the business unit data stored in Amazon RDS for MySQL in us-east-1 to its corporate analytics Amazon Redshift environment in us-west-1. The company wants to use AWS DMS with
Amazon RDS as the source endpoint and Amazon Redshift as the target endpoint.
Which action will allow AVS DMS to perform the replication?
  1. A Configure the AWS DMS replication instance in the same account and Region as Amazon Redshift.
  2. B Configure the AWS DMS replication instance in the same account as Amazon Redshift and in the same Region as Amazon RDS.
  3. C Configure the AWS DMS replication instance in its own account and in the same Region as Amazon Redshift.
  4. D Configure the AWS DMS replication instance in the same account and Region as Amazon RDS.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi xoay quanh việc sử dụng AWS Database Migration Service (DMS) để replicate dữ liệu từ Amazon RDS for MySQL (nằm ở us-east-1, tài khoản business unit) sang Amazon Redshift (nằm ở us-west-1, tài khoản corporate analytics). Hai tài khoản AWS hoàn toàn riêng biệt, và replication cần cross-account + cross-region.
Mục tiêu chính: Xác định vị trí phù hợp để deploy DMS replication instance nhằm đảm bảo DMS có thể kết nối source (RDS) và target (Redshift) thành công.
🛠️ Yếu tố kỹ thuật quan trọng (dựa trên AWS DMS docs cập nhật 2024-2026):

  • DMS replication instance là tài nguyên regional (không thể cross-region trực tiếp).
  • Với Redshift làm target, replication instance PHẢI ở cùng account và cùng Region với Redshift cluster (do Redshift sử dụng IAM role cho lệnh COPY/S3 unload/load, và yêu cầu quyền truy cập trực tiếp từ replication instance's security group/VPC).
  • Cross-account replication hỗ trợ qua IAM roles delegation và endpoint permissions, nhưng replication instance vẫn cần anchor ở account/region của target để tối ưu latency và quyền.
  • DMS hỗ trợ cross-region source-target miễn là network (VPC peering/S3 bucket intermediate) được config đúng.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Configure the AWS DMS replication instance in the same account and Region as Amazon Redshift.
Lý do:

  • Redshift target bắt buộc DMS replication instance phải ở cùng account corporate analytics và Region us-west-1 (theo AWS DMS User Guide - "Amazon Redshift as a target", cập nhật 2025).
  • Từ đây, DMS có thể:
    🧩 Assume IAM role cross-account để đọc RDS source ở us-east-1 (qua endpoint permissions).
    🛠️ Sử dụng S3 bucket làm intermediate storage cross-region (multi-account setup với bucket policy).
    📈 Giảm latency replication vì instance gần target nhất.
  • Đây là best practice cho cross-account/cross-region DMS → Redshift.

📋 Giải thích chi tiết tất cả các phương án

  • Configure the AWS DMS replication instance in the same account and Region as Amazon Redshift.
    ✅ Đúng 🏆: Như giải thích trên, phù hợp hoàn hảo với requirement của Redshift target. DMS instance ở account corporate (us-west-1) có quyền native load data vào Redshift, và kết nối source RDS qua cross-account IAM + network routing.

  • Configure the AWS DMS replication instance in the same account as Amazon Redshift and in the same Region as Amazon RDS.
    ❌ Sai 🚫: Instance ở account corporate nhưng Region us-east-1 (RDS). Redshift KHÔNG hỗ trợ replication instance ở region khác (us-west-1), vì COPY command yêu cầu same-region security group/VPC peering. Gây lỗi "endpoint connection failed" hoặc quyền IAM không match.

  • Configure the AWS DMS replication instance in its own account and in the same Region as Amazon Redshift.
    ❌ Sai 🔒: Instance ở account thứ ba riêng biệt (us-west-1). Redshift target KHÔNG cho phép cross-account replication instance trực tiếp mà không có complex role chaining (không standard). DMS cần native IAM role trong account Redshift để tạo external schema/table → Thiếu quyền, replication fail.

  • Configure the AWS DMS replication instance in the same account and Region as Amazon RDS.
    ❌ Sai 🌐: Instance ở account business unit (us-east-1). Tương tự, Redshift ở account khác + region khác KHÔNG tương thích; DMS không thể grant quyền load vào Redshift cross-account từ replication instance ở account source (vi phạm least-privilege và Redshift cluster policy).

📘 Tài liệu tham khảo

  • AWS DMS User Guide (2025): Using Amazon Redshift as a target → Xác nhận "replication instance and Redshift must be in the same account and Region".
  • AWS Well-Architected Framework - DevOps Pillar: Cross-account DMS patterns.
  • AWS re:Post & Blogs 2024-2026: Case studies về multi-account DMS → Redshift (tìm "DMS cross-account Redshift").
    🆕 Cập nhật mới nhất: DMS v3.4.7+ hỗ trợ enhanced cross-region via DMS Fleet Advisor, nhưng core req cho Redshift vẫn giữ nguyên.