Ngân hàng đề — AWS Certified DevOps Engineer Professional

Tìm thấy 681 câu.

Câu 101 Domain 6: Security and Compliance

A company wants to create an automated monitoring solution to generate real-time customized notifications regarding unrestricted security groups in the company's production AWS account. The notification must contain the name and ID of the noncompliant security group. The DevOps team at the company has already activated the restricted-ssh AWS Config managed rule. The team has also set up an Amazon Simple Notification Service (Amazon SNS) topic and subscribed relevant personnel to it.

Which of the following options represents the BEST solution for the given scenario?

  1. A

    Set up an Amazon EventBridge rule that matches an AWS Config evaluation result of NON_COMPLIANT for all AWS Config managed rules. Create an input transformer for the EventBridge rule. Set up the EventBridge rule to publish a notification to the SNS topic. Set up a filter policy on the SNS topic to send only notifications that contain the text of NON_COMPLIANT in the notification to subscribers

  2. B

    Set up an Amazon EventBridge rule that matches an AWS Config evaluation result of ERROR for the restricted-ssh rule. Create an input transformer for the EventBridge rule. Set up the EventBridge rule to publish a notification to the SNS topic

  3. C

    Set up an Amazon EventBridge rule that matches all AWS Config evaluation results for the restricted-ssh rule. Create an input transformer for the EventBridge rule. Set up the EventBridge rule to publish a notification to the SNS topic. Set up a filter policy on the SNS topic to send only notifications that contain the text of NON_COMPLIANT in the notification to subscribers

  4. D

    Set up an Amazon EventBridge rule that matches an AWS Config evaluation result of NON_COMPLIANT for the restricted-ssh rule. Create an input transformer for the EventBridge rule. Set up the EventBridge rule to publish a notification to the SNS topic

Xem giải thích

Đáp án

D — EventBridge rule khớp kết quả đánh giá NON_COMPLIANT của đúng rule restricted-ssh, dùng input transformer, đẩy sang SNS.

Vì sao đúng

Đề có ba ràng buộc và phương án D là phương án duy nhất thoả cả ba:

  1. Chỉ rule restricted-ssh — lọc theo tên rule trong event pattern.
  2. Chỉ trạng thái NON_COMPLIANT — không phải mọi kết quả.
  3. Thông báo phải có tên và ID của security group — đây là việc của input transformer: bóc trường từ event JSON và ghép thành câu tiếng người.
{
  "source": ["aws.config"],
  "detail-type": ["Config Rules Compliance Change"],
  "detail": {
    "configRuleName": ["restricted-ssh"],
    "newEvaluationResult": { "complianceType": ["NON_COMPLIANT"] }
  }
}

Không có input transformer thì SNS gửi nguyên khối JSON thô — kỹ thuật là đủ thông tin, nhưng đề đòi thông báo tuỳ biến nêu rõ tên và ID.

Vì sao các phương án khác sai

  • A. Khớp NON_COMPLIANT của mọi managed rule — sẽ dội thông báo từ mọi rule đang bật trong tài khoản. Nhiễu nhiều tới mức đội trực bắt đầu bỏ qua, và khi ấy cảnh báo mất hết giá trị.
  • B. Khớp trạng thái ERROR — ERROR nghĩa là rule không đánh giá được (thiếu quyền, throttle). Đó là lỗi vận hành của chính hệ thống giám sát, không phải security group hở.
  • C. Khớp mọi kết quả của restricted-ssh — bao gồm cả COMPLIANT, nên đội sẽ nhận thông báo mỗi lần một security group được sửa đúng. Cảnh báo cho tin tốt là cách nhanh nhất giết một kênh cảnh báo.

Ghi nhớ

Event pattern của EventBridge nên hẹp nhất có thể ngay từ đầu: lọc ở tầng rule không tốn gì, còn lọc ở Lambda phía sau thì vẫn phải trả tiền cho mọi lần gọi. Và khi đề nói "customized notification", hãy tìm chữ input transformer.

Câu 102 Domain 5: Incident and Event Response

A DevOps Engineer has been asked to chalk out a disaster recovery (DR) plan for a workload in production. The workload runs on Amazon EC2 instances behind an Application Load Balancer (ALB). The EC2 instances are configured with an Auto Scaling group across multiple Availability Zones. Amazon Route 53 is configured to point to the ALB using an alias record. Amazon RDS for PostgreSQL DB instance is the database service. The draft DR plan mandates an RTO of three hours and an RPO of around 15 minutes.

Which Disaster Recovery (DR) strategy should the DevOps Engineer opt for a cost-effective solution?

  1. A

    Opt for a pilot light DR strategy. Provision a copy of your entire workload infrastructure to a different AWS Region. Copy the first backup that consists of a full instance backup to the new RDS instance. In case of disaster, apply the incremental backup to the RDS instance in the new AWS Region. Configure Amazon Route 53 health checks to automatically initiate DNS failover to new Region

  2. B

    Opt for a pilot light DR strategy. Provision a copy of your core workload infrastructure to a different AWS Region. Create an RDS read replica in the new Region, and configure the new environment to point to the local RDS PostgreSQL DB instance. Configure Amazon Route 53 health checks to automatically initiate DNS failover to a new Region. Promote the read replica to the primary DB instance in case of a disaster

  3. C

    Configure your workload to simultaneously run in multiple AWS Regions as part of a multi-site active/active DR strategy. Replicate your entire workload to another AWS Region. With this strategy, asynchronous data replication between the regions enables near-zero RPO. Configure Amazon Route 53 with latency-based routing to choose between the active regional endpoint for directing user traffic

  4. D

    Opt for a Warm standby approach by ensuring that there is a scaled-down, but fully functional, copy of your production environment in another AWS Region. Then, deploy enough resources to handle initial traffic, ensuring low RTO, and then rely on Auto Scaling to ramp up for subsequent traffic

Xem giải thích

Đáp án

B — Pilot light: dựng bản sao phần lõi của hạ tầng ở Region khác, tạo RDS read replica cross-Region, khi có sự cố thì promote và mở rộng.

Vì sao đúng

Chọn chiến lược DR luôn bắt đầu từ hai con số. Đề cho RTO = 3 giờ và RPO ≈ 15 phút.

Chiến lược RTO RPO Chi phí
Backup & restore nhiều giờ nhiều giờ thấp nhất
Pilot light hàng chục phút – vài giờ giây tới phút thấp
Warm standby phút giây trung bình
Multi-site active/active gần bằng 0 gần bằng 0 cao nhất

RTO 3 giờ khá rộng rãi, còn RPO 15 phút thì loại backup & restore (snapshot định kỳ không đảm bảo được 15 phút). Pilot light nằm đúng ô: dữ liệu luôn được sao chép liên tục (read replica cross-Region, độ trễ thường vài giây), còn tầng compute thì tắt hoặc chỉ có khuôn và được dựng lên khi cần — hoàn toàn kịp trong 3 giờ.

Điểm cốt lõi phân biệt pilot light: phần dữ liệu luôn sống, phần xử lý thì ngủ.

Vì sao các phương án khác sai

  • A. Pilot light nhưng khôi phục từ full backup — mâu thuẫn nội tại. Khôi phục từ backup đầy đủ đưa RPO về đúng khoảng cách giữa hai lần backup, phá vỡ mục tiêu 15 phút.
  • C. Multi-site active/active — thoả yêu cầu quá dư nhưng đắt hơn nhiều lần và phức tạp hơn nhiều lần. Đề chỉ cần 3 giờ; trả tiền cho RTO gần bằng 0 là chọn sai theo chiều ngược lại.
  • D. Warm standby — cũng thoả, cũng đắt hơn cần thiết: warm standby giữ toàn bộ môi trường chạy thu nhỏ 24/7. Với RTO 3 giờ thì không có lý do trả khoản đó.

Ghi nhớ

Đọc RTO/RPO trước, chọn chiến lược sau. RPO nhỏ ⇒ phải sao chép dữ liệu liên tục; RTO nhỏ ⇒ phải giữ sẵn compute. Câu này RPO chặt còn RTO lỏng, nên đáp án là chiến lược chỉ giữ sẵn dữ liệu.

Câu 103 Domain 1: SDLC Automation

The flagship application at a company is deployed on Amazon EC2 instances running behind an Application Load Balancer (ALB) within an Auto Scaling group. A DevOps Engineer wants to configure a Blue/Green deployment for this application and has already created launch templates and Auto Scaling groups for both blue and green environments, each deploying to their respective target groups. The ALB can direct traffic to either environment's target group, and an Amazon Route 53 record points to the ALB. The goal is to enable an all-at-once transition of traffic from the software running on the blue environment's EC2 instances to the newly deployed software on the green environment's EC2 instances.

What steps should the DevOps Engineer take to fulfill these requirements?

  1. A

    Set up an all-at-once deployment to the blue environment's EC2 instances. Perform a Route 53 DNS update to point to the green environment's endpoint on the ALB

  2. B

    Initiate a rolling restart of the Auto Scaling group for the green environment to deploy the new software on the green environment's EC2 instances. Once the rolling restart is complete, leverage an AWS CLI command to update the ALB and direct traffic to the green environment's target group

  3. C

    Initiate a rolling restart of the Auto Scaling group for the green environment to deploy the new software on the green environment's EC2 instances. Perform a Route 53 DNS update to point to the green environment's endpoint on the ALB

  4. D

    Leverage an AWS CLI command to update the ALB and direct traffic to the green environment's target group. Then initiate a rolling restart of the Auto Scaling group for the green environment to deploy the new software on the green environment's EC2 instances

Xem giải thích

Đáp án

B — Rolling restart ASG xanh để cài phần mềm mới, sau đó dùng lệnh AWS CLI đổi ALB trỏ sang target group của môi trường xanh.

Vì sao đúng

Hạ tầng đã dựng xong: hai ASG, hai target group, một ALB trỏ được vào cả hai, Route 53 alias trỏ vào ALB. Còn lại đúng hai việc, và thứ tự là điều quan trọng nhất:

  1. Deploy trước lên môi trường xanh (đang không nhận traffic) — sai sót ở bước này không ảnh hưởng ai.
  2. Chuyển traffic sau, bằng cách sửa rule của listener trên ALB.
aws elbv2 modify-listener --listener-arn <arn-listener> \
  --default-actions Type=forward,TargetGroupArn=<arn-tg-xanh>

Chuyển ở tầng ALB chứ không phải tầng DNS là điểm mấu chốt của đề: yêu cầu là all-at-once. Sửa listener có hiệu lực ngay và cho mọi client; đổi bản ghi Route 53 thì phải chờ TTL và resolver ngoài tầm kiểm soát, nên trong nhiều phút vẫn có người đi vào môi trường cũ — đó là chuyển dần, không phải all-at-once.

Vì sao các phương án khác sai

  • A. Deploy lên môi trường xanh lam (blue) rồi đổi DNS sang xanh lá (green) — deploy một nơi, chuyển traffic sang nơi khác. Người dùng sẽ được đưa tới môi trường chưa cập nhật.
  • C. Deploy đúng nhưng chuyển bằng Route 53 — chuyển traffic không dứt khoát vì TTL của DNS, trượt yêu cầu all-at-once.
  • D. Chuyển traffic trước rồi mới deploy — nguy hiểm nhất. Traffic thật đổ vào môi trường xanh trong lúc nó đang được rolling restart, nghĩa là cố ý phục vụ khách trên một fleet đang bị xáo trộn.

Ghi nhớ

Blue/Green thì luôn deploy vào môi trường đang nhàn rỗi, chuyển traffic sau cùng, và nếu muốn cắt dứt điểm thì chuyển ở ALB/target group, đừng chuyển ở DNS — DNS không bao giờ tức thời.

Câu 104 Domain 2: Configuration Management and IaC

A company uses AWS CodeDeploy to deploy an AWS Lambda function as the final step of a CI/CD pipeline. The company has developed the Lambda function to handle incoming orders through an order-processing API. The DevOps team has noticed intermittent failures in the API occurring for a brief period after deploying the Lambda function. Upon investigation, the team suspects that the failures are a result of incomplete propagation of database changes before the Lambda function gets invoked.

What measures can help resolve this issue?

  1. A

    Add a BeforeAllowTraffic hook to the AppSpec file that tests and waits for any necessary database changes before traffic can flow to the new version of the Lambda function

  2. B

    Add a BeforeInstall hook to the AppSpec file that tests and waits for any necessary database changes before deploying a new version of the Lambda function

  3. C

    Add a ValidateService hook to the AppSpec file that validates that any necessary database changes are propagated

  4. D

    Add an AfterAllowTestTraffic hook to the AppSpec file that tests and waits for any necessary database changes after allowing a test traffic

Xem giải thích

Đáp án

A — Thêm hook BeforeAllowTraffic vào AppSpec để kiểm tra và chờ thay đổi CSDL lan xong trước khi cho traffic vào phiên bản Lambda mới.

Vì sao đúng

CodeDeploy khi deploy Lambda chỉ có bốn lifecycle event, ít hơn hẳn so với deploy lên EC2:

Start → BeforeAllowTraffic → AllowTraffic → AfterAllowTraffic → End

Chỉ BeforeAllowTraffic và AfterAllowTraffic là hook viết được (mỗi hook là một Lambda validation function). AllowTraffic do chính CodeDeploy thực hiện.

Triệu chứng trong đề — lỗi rải rác ngay sau deploy rồi tự hết — khớp chính xác với "traffic vào trước khi thay đổi CSDL sẵn sàng". Vậy phải chặn ở trước cửa AllowTraffic:

version: 0.0
Resources:
  - myFunction:
      Type: AWS::Lambda::Function
      Properties:
        Name: xu-ly-don
        Alias: prod
        CurrentVersion: 1
        TargetVersion: 2
Hooks:
  - BeforeAllowTraffic: kiem-tra-migration-xong

Hàm kiem-tra-migration-xong chờ tới khi CSDL sẵn sàng rồi gọi PutLifecycleEventHookExecutionStatus với Succeeded; nếu hết giờ thì trả Failed và CodeDeploy tự rollback, không client nào chạm vào phiên bản mới.

Vì sao các phương án khác sai

  • B. BeforeInstall và C. ValidateService — hai event này không tồn tại trong deploy Lambda; chúng thuộc bộ hook của EC2/On-Premises. Khai vào AppSpec của Lambda thì bị bỏ qua, và bạn có một biện pháp bảo vệ chỉ tồn tại trên giấy.
  • D. AfterAllowTestTraffic — có thật nhưng thuộc bộ hook của ECS blue/green, không phải Lambda. Và kể cả có, "sau khi cho test traffic vào" đã là muộn với vấn đề này.

Ghi nhớ

Bộ hook khác nhau theo nền tảng: | Nền tảng | Hook viết được | |---|---| | Lambda | BeforeAllowTraffic, AfterAllowTraffic | | ECS | thêm AfterAllowTestTraffic, BeforeInstall, AfterInstall | | EC2/On-prem | ApplicationStop, BeforeInstall, AfterInstall, ApplicationStart, ValidateService… |

Câu 105 Chọn nhiều đáp án Domain 2: Configuration Management and IaC

During an AWS CloudFormation stack update process, an error occurred in the updated template, causing AWS CloudFormation to initiate an automatic stack rollback. After the rollback attempt, a DevOps engineer noticed that the application remained unavailable, and the stack is now in the UPDATE_ROLLBACK_FAILED state.

To ensure the successful completion of the stack rollback, which actions should the DevOps engineer take? (Select two)

  1. A

    Automatically fix the issue by using AWS CloudFormation stack sets

  2. B

    Automatically fix the issue by using AWS CloudFormation drift detection

  3. C

    Update the existing AWS CloudFormation stack by using the original template

  4. D

    Execute a ContinueUpdateRollback command from the AWS CloudFormation

  5. E

    Manually fix the resources to match the correct state of the stack

Xem giải thích

Đáp án

D và E.

  • D — Gọi ContinueUpdateRollback.
  • E — Sửa tay tài nguyên về đúng trạng thái mà stack mong đợi.

Vì sao đúng

UPDATE_ROLLBACK_FAILED là một trong số ít trạng thái mà CloudFormation không tự thoát được. Nó nghĩa là: bản cập nhật hỏng, CloudFormation đã cố quay về trạng thái trước đó, và chính việc quay về cũng thất bại.

Ở trạng thái này stack chỉ chấp nhận đúng hai thao tác: ContinueUpdateRollback và DeleteStack. Mọi lệnh UpdateStack đều bị từ chối — nên phương án C bất khả thi về mặt cơ học.

Quy trình chuẩn:

  1. Đọc sự kiện của stack tìm tài nguyên nào chặn rollback và vì sao.
  2. Sửa tay nguyên nhân đó (E) — ví dụ trả lại một bản ghi bị xoá, gỡ object khỏi bucket không xoá được, sửa lại tham số bị đổi ngoài CloudFormation.
  3. Gọi ContinueUpdateRollback (D) để CloudFormation thử lại từ chỗ dừng.
aws cloudformation continue-update-rollback --stack-name my-stack
# Bó tay với một tài nguyên cụ thể thì bỏ qua nó:
aws cloudformation continue-update-rollback --stack-name my-stack \
  --resources-to-skip LogicalIdCuaTaiNguyenKetCung

--resources-to-skip đưa stack về UPDATE_ROLLBACK_COMPLETE nhưng để lại drift: tài nguyên bị bỏ qua không còn khớp template.

Vì sao các phương án khác sai

  • A. StackSets — công cụ triển khai một template ra nhiều tài khoản/Region, chẳng liên quan gì tới việc cứu một stack hỏng.
  • B. Drift detection — chỉ báo cáo khác biệt, không sửa gì. Nó có ích ở bước chẩn đoán, nhưng không phải hành động cứu.
  • C. UpdateStack với template gốc — bị API từ chối thẳng ở trạng thái UPDATE_ROLLBACK_FAILED.

Ghi nhớ

Ba trạng thái "kẹt" cần nhớ: UPDATE_ROLLBACK_FAILED → ContinueUpdateRollback; ROLLBACK_FAILED khi tạo mới → chỉ xoá stack được; DELETE_FAILED → xoá lại và bỏ qua tài nguyên cứng đầu.

Câu 106 Domain 3: Resilient Cloud Solutions

A company uses serverless application architecture to process thousands of requests using AWS Lambda with Amazon DynamoDB as the database. The Amazon API Gateway REST API is used to invoke an AWS Lambda function that loads a large amount of data from the Amazon DynamoDB database. This results in cold start latencies of 7-10 seconds. The DynamoDB tables have already been configured with DynamoDB Accelerator (DAX) to reduce latency. Yet, customers report of application latency, especially during peak access hours. The application receives maximum traffic between 2 PM -5 PM every day and gradually reduces thereafter, reporting a minimum traffic post 8 PM.

How should a DevOps engineer configure the AWS Lambda function to reduce its latency at all times?

  1. A

    Configure the AWS Lambda function to use Reserved concurrency. Configure the application Auto Scaling on the Lambda function with reserved concurrency as half of the Lambda function instances invoked during peak traffic hours

  2. B

    Use the Lambda function layers for parallel processing of Lambda runtime requests. Deploy the Lambda function as a .zip file archive to utilize Lambda layer functionality

  3. C

    Configure the Lambda function execution environment to use up to 10,240 MB of ephemeral storage. This transient cache for data between invocations can be accessed from the Lambda function code via /tmp space

  4. D

    Configure the AWS Lambda function to use provisioned concurrency. Configure the application Auto Scaling on the Lambda function with provisioned concurrency values of 1 (minimum) and 100 (maximum) respectively

Xem giải thích

Đáp án

D — Bật provisioned concurrency cho Lambda và dùng Application Auto Scaling với min = 1, max = 100.

Vì sao đúng

Đề mô tả rất rõ cold start 7–10 giây, tập trung vào khung 14h–17h. Provisioned concurrency là cơ chế duy nhất của Lambda loại bỏ cold start: AWS khởi tạo sẵn số môi trường thực thi bạn yêu cầu, chạy xong phần init (nạp runtime, code ngoài handler, mở kết nối), rồi giữ chúng ở trạng thái ấm.

Ghép với Application Auto Scaling thì lượng dự phòng tự lên xuống theo tải, không phải trả tiền giữ 100 môi trường cả ngày cho một khung ba tiếng.

Điểm cần phân biệt cho rõ:

Reserved concurrency Provisioned concurrency
Tác dụng Đặt trần số lần chạy đồng thời Khởi tạo sẵn môi trường
Với cold start không giúp gì loại bỏ
Chi phí miễn phí trả tiền theo lượng giữ sẵn

Vì sao các phương án khác sai

  • A. Reserved concurrency — đây là bẫy trung tâm của câu hỏi. Reserved concurrency chỉ giới hạn trần; nó không giữ môi trường nào ấm cả. Đặt nó xong cold start vẫn nguyên vẹn, mà giờ còn thêm nguy cơ bị throttle. Ngoài ra Application Auto Scaling không nhận reserved concurrency làm scalable target.
  • B. Lambda layer để "xử lý song song" — hiểu sai layer. Layer chỉ là cách đóng gói thư viện dùng chung; nó không liên quan gì tới đồng thời hay hiệu năng khởi động.
  • C. Ephemeral storage /tmp 10.240 MB — /tmp chỉ tồn tại trong vòng đời một môi trường thực thi và không dùng được như cache giữa các lần gọi một cách đảm bảo. Quan trọng hơn: cold start là chi phí khởi tạo môi trường, tăng dung lượng đĩa không rút ngắn được nó. Đề cũng đã có DAX cho phần cache dữ liệu.

Ghi nhớ

Thấy chữ cold start trong đề Lambda thì đáp án gần như chắc chắn là provisioned concurrency. Thấy chữ giới hạn/throttle/bảo vệ downstream thì là reserved concurrency. Hai tên rất giống nhau, tác dụng hoàn toàn khác.

Câu 107 Domain 5: Incident and Event Response

A developer configured an AWS CloudFormation template to create custom resource necessary for the project. The AWS Lambda function for the custom resource executed successfully as seen by the successful creation of the custom resource. But, the CloudFormation stack is not transitioning from in-progress status (CREATE_IN_PROGRESS) to completion status (CREATE_COMPLETE).

Which step did the developer possibly miss for the successful completion of the CloudFormation stack?

  1. A

    Configure the AWS Lambda function to send the response(SUCCESS or FAILED) of the custom resource creation to a pre-signed Amazon Simple Storage Service URL

  2. B

    The AWS CloudFormation resource AWS::CloudFormation::CustomResource should be used to specify a custom resource in the template

  3. C

    If the template developer and custom resource provider are configured to the same person or entity then CloudFormation stack completion fails

  4. D

    After executing the send method in the cfn-response module, the Lambda function terminates, so anything written after this method is ignored

Xem giải thích

Đáp án

A — Lambda phải gửi phản hồi (SUCCESS hoặc FAILED) tới pre-signed S3 URL mà CloudFormation cung cấp.

Vì sao đúng

Custom resource hoạt động theo mô hình gọi lại (callback), và đây là chỗ hầu hết người mới vấp:

CloudFormation → gửi request (kèm ResponseURL) → Lambda
CloudFormation ← ĐỢI ← Lambda PUT phản hồi vào ResponseURL

CloudFormation không nhìn kết quả trả về của hàm Lambda. Nó chỉ chờ một HTTP PUT vào pre-signed S3 URL nằm trong trường ResponseURL của event. Hàm chạy xong, thậm chí chạy đúng, mà quên PUT thì stack đứng ở CREATE_IN_PROGRESS cho tới khi hết giờ chờ — mặc định là 1 giờ. Đúng triệu chứng trong đề.

Trong Python thì cfnresponse lo phần này:

import cfnresponse
def handler(event, context):
    try:
        # ... việc thật ...
        cfnresponse.send(event, context, cfnresponse.SUCCESS, {'KetQua': 'xong'})
    except Exception as e:
        # BẮT BUỘC báo cả khi lỗi, nếu không stack treo nguyên giờ
        cfnresponse.send(event, context, cfnresponse.FAILED, {'Loi': str(e)})

Khối except cũng phải gửi — đó là lỗi phổ biến thứ hai sau việc quên gửi hoàn toàn.

Vì sao các phương án khác sai

  • B. "Phải dùng AWS::CloudFormation::CustomResource" — không bắt buộc. Dạng Custom::TenGiDo hoàn toàn hợp lệ và còn được khuyến nghị vì tên hiện rõ hơn trong console. Cả hai đều cần phản hồi như nhau.
  • C. "Người viết template và người cung cấp custom resource là một thì stack lỗi" — bịa. CloudFormation không có khái niệm nào như vậy.
  • D. "Sau send thì hàm kết thúc, mã phía sau bị bỏ qua" — câu này đúng về mặt sự thật (đó là một lưu ý có thật trong tài liệu AWS) nhưng không phải nguyên nhân của triệu chứng đang mô tả. Đề hỏi bước nào bị bỏ sót, và bước bị bỏ sót là chính lời gọi send. Đây là kiểu nhiễu tinh vi: một câu đúng đặt sai chỗ.

Ghi nhớ

Stack treo ở CREATE_IN_PROGRESS với custom resource ⇒ gần như luôn là thiếu phản hồi. Luôn bọc toàn bộ handler trong try/except và gửi FAILED ở nhánh lỗi — nếu không, mỗi lần deploy hỏng là một tiếng đồng hồ chờ vô ích.

Câu 108 Domain 2: Configuration Management and IaC

An e-commerce company has a serverless application stack that consists of CloudFront, API Gateway and Lambda functions. The company has hired you to improve the current deployment process which creates a new version of the Lambda function and then runs an AWS CLI script for deployment. In case the new version errors out, then another CLI script is invoked to deploy the previous working version of the Lambda function. The company has mandated you to decrease the time to deploy new versions of the Lambda functions and also reduce the time to detect and roll back when errors are identified.

Which of the following solutions would you suggest for the given use case?

  1. A

    Set up and deploy a CloudFormation stack containing a new API Gateway endpoint that points to the new Lambda version. Test the updated CloudFront origin that points to this new API Gateway endpoint and in case errors are detected then revert the CloudFront origin to the previous working API Gateway endpoint

  2. B

    Use Serverless Application Model (SAM) and leverage the built-in traffic-shifting feature of SAM to deploy the new Lambda version via CodeDeploy and use pre-traffic and post-traffic test functions to verify code. Rollback in case CloudWatch alarms are triggered

  3. C

    Set up and deploy nested CloudFormation stacks with the CloudFront distribution as well as the API Gateway in the parent stack. Create and deploy a child stack containing the Lambda functions. To address any changes in a Lambda function, create a CloudFormation change set and deploy. Use pre-traffic and post-traffic test functions of the change set to verify the deployment. Rollback in case CloudWatch alarms are triggered

  4. D

    Set up and deploy nested CloudFormation stacks with the CloudFront distribution as well as the API Gateway in the parent stack. Create and deploy a child stack containing the Lambda functions. To address any changes in a Lambda function, create a CloudFormation change set and deploy. In case the Lambda function errors out, rollback the CloudFormation change set to the previous version

Xem giải thích

Đáp án

B — Dùng AWS SAM với tính năng traffic-shifting sẵn có, deploy qua CodeDeploy, kèm hàm kiểm thử pre-traffic và post-traffic.

Vì sao đúng

Quy trình hiện tại là hai script CLI: một để deploy, một để lùi lại khi hỏng. Vấn đề không phải script xấu, mà là phát hiện lỗi và rollback đều do người khởi động. Đề đòi giảm cả hai: thời gian deploy và thời gian phát hiện + rollback.

SAM khai bốn dòng là có đủ:

Resources:
  HamXuLy:
    Type: AWS::Serverless::Function
    Properties:
      AutoPublishAlias: prod
      DeploymentPreference:
        Type: Linear10PercentEvery1Minute   # hoặc Canary10Percent5Minutes
        Alarms:
          - !Ref AlarmLoiCao
        Hooks:
          PreTraffic:  !Ref KiemTraTruoc
          PostTraffic: !Ref KiemTraSau

Ba thứ có được ngay:

  • Traffic shifting dần qua alias weight — lỗi chỉ chạm vào một phần nhỏ người dùng.
  • Alarms — CloudWatch alarm kêu là CodeDeploy tự rollback, không cần ai bấm gì. Đây chính là vế "giảm thời gian phát hiện".
  • Pre/post-traffic hook — chặn bản hỏng trước khi nó chạm production.

Vì sao các phương án khác sai

  • A. CloudFormation stack với API Gateway endpoint mới cho mỗi phiên bản — mỗi lần deploy lại dựng một endpoint mới rồi sửa origin CloudFront: chậm hơn hẳn (CloudFront lan cấu hình mất vài phút), và rollback vẫn là thao tác tay.
  • C và D. Nested stack — cập nhật nested stack không nhanh hơn, và cũng không có cơ chế tự phát hiện lỗi nào. Rollback vẫn phải người bấm, đúng vấn đề cần bỏ.

Ghi nhớ

Với Lambda, cặp SAM DeploymentPreference + Alarms là cách rẻ nhất để có canary deploy và rollback tự động theo chỉ số. Không có Alarms thì traffic shifting chỉ làm chậm sự cố lại, chứ không tự chữa.

Câu 109 Chọn nhiều đáp án Domain 6: Security and Compliance

A media company extensively uses Amazon S3 buckets for storing images files, documents, and other business-specific data. The company has mandated enabling logging for all Amazon S3 buckets. The audit team publishes the reports of all AWS resources failing company security standards. Until recently, the security team would pick the list of noncompliant Amazon S3 buckets from the audit list and execute remediation actions manually for each resource. This process is not only time-consuming but also leaves noncompliant resources vulnerable for a long duration.

Which combination of steps should a DevOps Engineer take to meet these requirements using an automated solution? (Select two)

  1. A

    Configure AWS Config Auto Remediation for the AWS Config rule s3-bucket-logging-enabled. From the remediation action list choose AWS Lambda to implement a custom function that will enable S3 logging for the S3 bucket ID passed

  2. B

    Configure AWS Config Auto Remediation for the AWS Config rule s3-bucket-logging-enabled. From the remediation action list choose AWS-ConfigureS3BucketLogging

  3. C

    While setting up remediation action, pass the resource ID of non-compliant resources to the remediation action. This configuration is mandatory for auto-remediation to work

  4. D

    Configure AWS Config Auto Remediation for the AWS Config rule s3-logging-enabled. Create your own custom remediation action using AWS Systems Manager Automation documents to enable logging on the S3 bucket

  5. E

    The AutomationAssumeRole in the remediation action parameters should be assumable by SSM. The user must have pass-role permissions for that role when they create the remediation action in AWS Config

Xem giải thích

Đáp án

B và E.

  • B — Bật Auto Remediation cho rule s3-bucket-logging-enabled, chọn sẵn action AWS-ConfigureS3BucketLogging.
  • E — AutomationAssumeRole phải cho SSM assume được, và người tạo remediation phải có quyền iam:PassRole với role đó.

Vì sao đúng

Đề nhấn mạnh quy trình thủ công đang chậm và để lộ tài nguyên lâu. AWS Config Auto Remediation giải quyết trọn vẹn: rule phát hiện NON_COMPLIANT → tự chạy một SSM Automation document để sửa.

Vì sao chọn document dựng sẵn (B). AWS đã có AWS-ConfigureS3BucketLogging; viết Lambda làm lại đúng việc đó là thêm code phải bảo trì mà không được gì.

Vì sao E là bắt buộc. Config không tự sửa bằng danh tính của nó — nó truyền một role (AutomationAssumeRole) cho SSM để SSM dùng role đó. Nên cần đủ hai vế:

// Trust policy của AutomationAssumeRole
{"Effect":"Allow","Principal":{"Service":"ssm.amazonaws.com"},"Action":"sts:AssumeRole"}

và người cấu hình remediation phải có iam:PassRole lên role đó. Thiếu vế nào thì remediation im lặng không chạy, còn rule vẫn báo NON_COMPLIANT mãi.

Vì sao các phương án khác sai

  • A. Chọn Lambda tự viết — về nguyên tắc làm được, nhưng thừa khi đã có document dựng sẵn cho đúng việc này.
  • C. "Truyền resource ID là bắt buộc" — sai ở chữ bắt buộc. Tham số ResourceId là tuỳ chọn; nếu không khai, Config tự truyền ID của tài nguyên vi phạm vào tham số mặc định của document. Có khai thì hữu ích khi tên tham số không khớp, nhưng không phải điều kiện để auto remediation hoạt động.
  • D. Rule tên s3-logging-enabled — không có rule nào tên như vậy. Tên đúng là s3-bucket-logging-enabled. Đề thi AWS rất hay bẫy bằng cách bớt một chữ trong tên managed rule.

Ghi nhớ

Auto Remediation cần đủ bốn mảnh: rule, document, AutomationAssumeRole mà SSM assume được, và iam:PassRole cho người cấu hình. Nó cũng có MaximumAutomaticAttempts — sửa hoài không đạt thì Config dừng, không lặp vô hạn.

Câu 110 Chọn nhiều đáp án Domain 6: Security and Compliance

The DevOps team for an e-commerce company wants to implement a patching plan on AWS Cloud for a large mixed fleet of Windows and Linux servers. The patching plan has to be auditable and must be implemented securely to ensure compliance with the company's business requirements.

Which of the following options would you recommend to address these requirements with MINIMAL effort? (Select two)

  1. A

    Set up an OS-native patching service to manage the update frequency and release approval for all instances. Set up AWS Config to provide audit and compliance reporting

  2. B

    Configure CloudFormation automatic patching support for all applications which will keep the OS up-to-date following the initial installation. Set up AWS Config to provide audit and compliance reporting

  3. C

    Set up Systems Manager Agent on all instances to manage patching. Test patches in pre-production and then deploy as a maintenance window task with the appropriate approval

  4. D

    Apply patch baselines using the AWS-ApplyPatchBaseline SSM document

  5. E

    Apply patch baselines using the AWS-RunPatchBaseline SSM document

Xem giải thích

Đáp án

C và E.

  • C — Cài SSM Agent trên toàn bộ instance, thử bản vá ở pre-production rồi triển khai như maintenance window task có phê duyệt.
  • E — Áp patch baseline bằng document AWS-RunPatchBaseline.

Vì sao đúng

Đề đòi: đội máy lẫn Windows và Linux, phải kiểm toán được, phải an toàn, và ít công sức nhất.

AWS Systems Manager Patch Manager đáp ứng cả bốn bằng một công cụ:

Yêu cầu Patch Manager
Lẫn Windows và Linux Một cơ chế, patch baseline riêng cho từng hệ điều hành
Kiểm toán được Báo cáo tuân thủ vá, tích hợp AWS Config
An toàn Không cần mở SSH/RDP; đi qua SSM Agent
Ít công sức Có sẵn baseline mặc định, có lịch, có phê duyệt

Maintenance window cho phần lịch và phần phê duyệt: vá chỉ chạy trong cửa sổ đã định, sau khi đã thử ở pre-production.

Về tên document (E): đây là mấu chốt phân biệt C/D/E. Document đúng và hiện hành là AWS-RunPatchBaseline, chạy được trên cả Windows lẫn Linux.

Vì sao các phương án khác sai

  • A. Công cụ vá của chính hệ điều hành — nghĩa là WSUS cho Windows và yum-cron/unattended-upgrades cho Linux: hai hệ thống, hai bộ báo cáo, hai chỗ để hỏng. Trái hẳn với "MINIMAL effort".
  • B. "CloudFormation automatic patching" — không tồn tại. CloudFormation cung cấp hạ tầng, không vá hệ điều hành đang chạy.
  • D. AWS-ApplyPatchBaseline — document này đã cũ, chỉ chạy trên Windows và bị AWS-RunPatchBaseline thay thế. Đây đúng là loại bẫy "tên gần giống" mà đề hay dùng; nhớ chữ Run, không phải Apply.

Ghi nhớ

Bộ ba của Patch Manager: patch baseline (vá nào được duyệt) + patch group (áp cho máy nào, gắn bằng tag Patch Group) + maintenance window (chạy lúc nào). Document luôn là AWS-RunPatchBaseline.