Ngân hàng đề — AWS Certified Generative AI Developer Pro

Tìm thấy 100 câu.

Câu 1 Chọn nhiều đáp án AI Safety, Security, and Governance

A digital assistant answers user questions across multiple business functions, including customer support, product usage, and general knowledge. During internal reviews, teams observe that some user inputs contain aggressive or explicit phrasing that varies in severity, while other inputs request guidance in domains, such as investment advice, that the company has decided the assistant should never address, even when the request is phrased in a neutral or professional manner. The platform team wants to enforce consistent safeguards at inference time. The solution must handle both types of inputs predictably, support centralized policy updates, and apply controls automatically during model invocation.

Which design choices should be combined to provide the MOST appropriate solution? (Select two)

  1. A

    Use AWS Step Functions and AWS Lambda to orchestrate a multi stage validation pipeline that classifies prompts by sentiment and topic intent, applies business logic rules, invokes guardrails conditionally, and routes decisions through approval states before model execution

  2. B

    Use Amazon Bedrock Guardrails with the strictest content filter settings across all categories and treat requests related to restricted domains as high severity language violations. Tune thresholds over time based on observed user behavior and false positives. Leverage filter severity to suppress all undesired responses

  3. C

    Use Amazon Bedrock Guardrails content filters to evaluate prompts and responses for unsafe language patterns such as harassment, sexual references, or violent expressions. Configure category specific thresholds to determine when responses should be filtered, masked, or refused. Apply these checks during inference to enforce consistent behavior

  4. D

    Use Amazon Bedrock Guardrails content filters for language analysis and add a custom Lambda based topic classification layer that evaluates intent using predefined domain taxonomies, synchronizes rule updates across environments, and coordinates enforcement decisions through Step Functions prior to invoking the model

  5. E

    Use Amazon Bedrock Guardrails denied topics to define subject areas the assistant must not address under any circumstances. Detect requests related to those domains regardless of tone or wording. Return deterministic refusal responses when such requests are identified

Xem giải thích

Đáp án

C và E.

  • C — Dùng content filter của Bedrock Guardrails cho ngôn ngữ độc hại, với ngưỡng riêng theo từng danh mục
  • E — Dùng denied topics của Bedrock Guardrails cho các chủ đề bị cấm tuyệt đối

Vì sao đúng

Đề mô tả hai loại đầu vào có bản chất khác nhau, và Guardrails có hai cơ chế riêng cho chúng — dùng đúng cơ chế cho đúng loại là toàn bộ ý của câu hỏi.

C — content filter cho ngôn ngữ độc hại. Đề nói ngôn ngữ hung hăng hoặc tục tĩu "có mức độ khác nhau" — nên cần cơ chế phân mức:

{
  "contentPolicyConfig": {
    "filtersConfig": [
      {"type": "HATE",       "inputStrength": "HIGH",   "outputStrength": "HIGH"},
      {"type": "INSULTS",    "inputStrength": "MEDIUM", "outputStrength": "HIGH"},
      {"type": "SEXUAL",     "inputStrength": "HIGH",   "outputStrength": "HIGH"},
      {"type": "VIOLENCE",   "inputStrength": "MEDIUM", "outputStrength": "MEDIUM"},
      {"type": "MISCONDUCT", "inputStrength": "MEDIUM", "outputStrength": "MEDIUM"},
      {"type": "PROMPT_ATTACK", "inputStrength": "HIGH", "outputStrength": "NONE"}
    ]
  }
}

Mỗi danh mục có ngưỡng riêng (NONE, LOW, MEDIUM, HIGH) — đúng nghĩa "category specific thresholds".

E — denied topics cho chủ đề bị cấm. Đề nói tư vấn đầu tư không bao giờ được trả lời, KỂ CẢ khi hỏi bằng giọng lịch sự — đó không phải vấn đề ngôn ngữ, mà là vấn đề chủ đề:

{
  "topicPolicyConfig": {
    "topicsConfig": [{
      "name": "TuVanDauTu",
      "definition": "Bất kỳ khuyến nghị nào về mua bán chứng khoán, phân bổ danh mục đầu tư, hoặc tư vấn tài chính cá nhân",
      "examples": [
        "Tôi nên mua cổ phiếu nào?",
        "Anh nghĩ danh mục của tôi nên phân bổ thế nào?"
      ],
      "type": "DENY"
    }]
  }
}

Guardrails hiểu ý định theo ngữ nghĩa, nên nó bắt được cả những câu hỏi diễn đạt khéo — điều mà content filter (vốn nhìn vào tính độc hại của ngôn ngữ) hoàn toàn không làm được.

Và cả hai đều đáp ứng ba yêu cầu vận hành của đề: áp dụng tự động lúc gọi model, cập nhật chính sách tập trung (một guardrail dùng cho nhiều ứng dụng), và hành vi có thể đoán trước (trả về thông điệp từ chối cố định).

Vì sao các phương án khác sai

  • B. Dùng content filter ở mức nghiêm ngặt nhất cho mọi danh mục, và coi chủ đề bị cấm là vi phạm ngôn ngữ mức cao — sai công cụ cho vấn đề thứ hai: câu hỏi "Anh có thể tư vấn giúp tôi phân bổ danh mục không?" hoàn toàn không có ngôn ngữ độc hại nào — content filter sẽ cho qua dù đặt ngưỡng cao nhất. Và đặt HIGH cho mọi danh mục còn gây chặn nhầm rất nhiều câu hỏi hợp lệ.
  • D. Content filter cộng một lớp phân loại chủ đề tự viết bằng Lambda — tự dựng lại thứ Guardrails làm sẵn: bạn phải viết và bảo trì taxonomy, tự đồng bộ giữa các môi trường, tự lo độ chính xác. Trái yêu cầu "centralized policy updates" và "applied automatically".
  • A. Dựng pipeline nhiều tầng bằng Step Functions và Lambda — phức tạp nhất: thêm độ trễ ở mọi request, thêm nhiều thành phần phải vận hành, cho một việc mà Guardrails làm bằng cấu hình.

Ghi nhớ

Sáu chính sách của Amazon Bedrock Guardrails: | Chính sách | Chặn gì | |---|---| | Content filters | ngôn ngữ độc hại: hate, insults, sexual, violence, misconduct | | Prompt attacks | jailbreak, prompt injection | | Denied topics | chủ đề bị cấm, nhận diện theo NGỮ NGHĨA | | Word filters | từ khoá cụ thể, danh sách tục tĩu | | Sensitive information | PII — che (mask) hoặc chặn | | Contextual grounding | chặn câu trả lời không dựa trên nguồn (chống ảo giác) |

Cách phân biệt hai cơ chế của câu này: | Vấn đề | Cơ chế | |---|---| | "Câu này nói năng thô tục" | content filter — theo mức độ | | "Chủ đề này chúng ta không bao giờ đụng tới" | denied topics — bất kể cách diễn đạt |

Hai đặc điểm vận hành đáng biết:

  1. Guardrail áp cho cả prompt lẫn response, và cấu hình inputStrength với outputStrength riêng biệt.
  2. Guardrail có version — cập nhật ở một chỗ, mọi ứng dụng dùng chung được hưởng; và rollback bằng cách trỏ về version cũ.

Gọi model kèm guardrail:

bedrock.converse(
    modelId='anthropic.claude-3-5-sonnet-20241022-v2:0',
    guardrailConfig={'guardrailIdentifier': 'abc123', 'guardrailVersion': '3'},
    messages=[...]
)

Và ApplyGuardrail API cho phép đánh giá nội dung mà không gọi model — hữu ích khi muốn lọc đầu vào trước, hoặc kiểm tra nội dung từ nguồn không phải Bedrock.

Câu 2 Foundation Model Integration, Data Management and Compliance

An e-commerce company is building a conversational support assistant that searches across product descriptions, troubleshooting guides, and short FAQ entries to answer customer questions. A pure keyword search approach fails to match semantically similar questions, but a pure vector based semantic search sometimes ranks irrelevant results that share general themes but not specific product attributes or model numbers. Stakeholders want higher precision on the first page of results, especially for detailed troubleshooting queries. As a GenAI developer, you must design an advanced search architecture that balances recall and precision and improves the ranking of results used for FM context augmentation.

What is the most effective retrieval and ranking approach to meet these requirements?

  1. A

    Load all troubleshooting documents into Amazon Bedrock Knowledge Bases and rely on default semantic search behavior without applying metadata filters or reranking. Then depend on the FM to synthesize relevance based on the retrieved context

  2. B

    Combine vector similarity search in OpenSearch Serverless with metadata filtering and structured attributes, and leverage the hybrid score that results from this combination for the final ranking. This approach improves the breadth of retrieval while keeping the architecture simple without introducing an additional reranking layer

  3. C

    Combine vector similarity search in Amazon OpenSearch Serverless with metadata filtering and structured attributes such as product IDs or device categories, then apply an Amazon Bedrock reranking model to refine the top results. This approach increases precision by integrating semantic scoring, structured filtering, and learned ranking

  4. D

    Use a high dimensional Titan embedding model for all troubleshooting documents and rely entirely on vector similarity search without using structured filters. Then let the FM interpret the returned text even if it includes multiple product categories or different troubleshooting steps

Xem giải thích

Đáp án

C — Kết hợp vector search trong OpenSearch Serverless với metadata filtering, rồi áp dụng reranking model của Bedrock để tinh chỉnh kết quả đầu.

Vì sao đúng

Đề mô tả hai điểm yếu đối lập của hai phương pháp tìm kiếm: | Phương pháp | Điểm yếu | |---|---| | Keyword search thuần | không khớp được câu hỏi diễn đạt khác nhưng cùng nghĩa | | Vector search thuần | xếp hạng cao những kết quả cùng chủ đề chung nhưng SAI thuộc tính cụ thể (mã model, thông số sản phẩm) |

Đáp án C giải quyết bằng ba tầng bổ sung cho nhau:

① Vector search      → bắt được ngữ nghĩa (recall cao)
        ↓
② Metadata filtering → loại bỏ sai sản phẩm, sai danh mục (precision)
        ↓
③ Reranking model    → sắp xếp lại top-K theo mức liên quan thật (precision cao nhất)

Tầng ② là thứ xử lý đúng điểm yếu đề nêu: câu hỏi về mã model XR-500 sẽ không còn trả về tài liệu của XR-300 chỉ vì hai văn bản "nghe giống nhau":

{
  "query": {
    "bool": {
      "must": [{"knn": {"embedding": {"vector": [...], "k": 50}}}],
      "filter": [
        {"term": {"product_id": "XR-500"}},
        {"term": {"doc_type": "troubleshooting"}}
      ]
    }
  }
}

Tầng ③ là điểm phân biệt với phương án B. Reranker là một mô hình cross-encoder đọc cặp (câu hỏi, tài liệu) cùng lúc rồi chấm điểm liên quan — chính xác hơn hẳn việc so khoảng cách vector, vốn chỉ so hai vector được mã hoá độc lập:

bedrock_agent.rerank(
    queries=[{'textQuery': {'text': cau_hoi}}],
    sources=[{'inlineDocumentSource': {'textDocument': {'text': doc}}} for doc in ket_qua],
    rerankingConfiguration={
        'bedrockRerankingConfiguration': {
            'modelConfiguration': {'modelArn': 'arn:aws:bedrock:...::foundation-model/amazon.rerank-v1:0'},
            'numberOfResults': 5
        }
    }
)

Và đó đúng là điều đề yêu cầu: "higher precision on the first page of results".

Vì sao các phương án khác sai

  • B. Vector search + metadata filtering, dùng điểm hybrid làm xếp hạng cuối — đây là phương án gần nhất, và nó giải quyết được một nửa. Nhưng thiếu bước reranking — nên thứ tự trên trang đầu vẫn dựa vào điểm tương đồng vector, vốn là thứ đang xếp hạng kém. Đề nhấn mạnh "improves the ranking", và đó là việc của reranker.
  • A. Nạp mọi tài liệu vào Bedrock Knowledge Bases và dựa vào hành vi semantic search mặc định — chính là vấn đề đang có: semantic search thuần không phân biệt được mã model và thông số cụ thể. Và "dựa vào FM tự tổng hợp mức liên quan" là giao việc lọc cho model — tốn token và không đáng tin.
  • D. Dùng Titan embedding chiều cao cho mọi tài liệu, chỉ dùng vector similarity — tăng số chiều không sửa được vấn đề bản chất: embedding mã hoá ngữ nghĩa tổng thể, nên hai tài liệu về hai model sản phẩm khác nhau vẫn rất gần nhau trong không gian vector.

Ghi nhớ

Ba tầng của một kiến trúc RAG chất lượng cao: | Tầng | Mục tiêu | Công cụ | |---|---|---| | Retrieval | recall — không bỏ sót | vector search, hybrid search | | Filtering | loại nhiễu có cấu trúc | metadata filter | | Reranking | precision — thứ tự đúng | cross-encoder reranker |

So sánh hai loại mô hình xếp hạng: | | Bi-encoder (embedding) | Cross-encoder (reranker) | |---|---|---| | Cách hoạt động | mã hoá câu hỏi và tài liệu ĐỘC LẬP rồi so khoảng cách | đọc CẢ HAI cùng lúc rồi chấm điểm | | Tốc độ | rất nhanh (vector tính sẵn) | chậm hơn | | Độ chính xác | vừa | cao hơn hẳn | | Quy mô | hàng triệu tài liệu | chỉ top-K (thường 25–100) |

Đó là lý do kiến trúc chuẩn dùng cả hai: bi-encoder thu hẹp từ hàng triệu xuống vài chục, rồi cross-encoder sắp xếp lại vài chục đó.

Và hybrid search là kỹ thuật bổ sung đáng biết: kết hợp BM25 (keyword) với vector similarity trong một truy vấn, thường theo công thức Reciprocal Rank Fusion. Nó đặc biệt hiệu quả với thuật ngữ hiếm và mã sản phẩm — thứ mà embedding hay bỏ sót vì chúng ít xuất hiện trong dữ liệu huấn luyện.

Câu 3 Chọn nhiều đáp án Implementation and Integration

An insurance company is building a GenAI workflow that processes claims descriptions in multiple languages and varying complexity and uses different specialized models for triage, summarization, and risk classification. The company wants to select the appropriate model based on language, token length, and content category and also evolve routing rules over time using performance metrics such as latency, cost, and human quality scores stored in an operational metrics store. You are responsible for designing an orchestration workflow that can apply content based routing, use metric driven rules for model selection, and support confidence based cascading between models.

Which two approaches should you implement to design this intelligent routing workflow so that it can adapt to content characteristics and observed model performance? (Select two)

  1. A

    Use Step Functions with a Lambda pre processing task to compute language, token length, and content attributes and pass them to Choice states for routing to the triage, summarization, or risk classification model. Store performance metrics in DynamoDB and have a Lambda task read these metrics at runtime to adjust routing thresholds dynamically

  2. B

    Use Step Functions with a Lambda pre processing task that computes language and token length, then call a metrics evaluation Lambda function to read latency and quality data from DynamoDB before routing to a Bedrock model. Apply Choice states only after the model inference step so the workflow can examine the generated output and decide whether to invoke a backup model

  3. C

    Use Step Functions Choice states with routing rules fixed by values stored in Parameter Store and update thresholds manually through configuration changes. Write latency and quality metrics to DynamoDB but use them only for offline dashboards rather than feeding them into the real time routing logic

  4. D

    Implement a single Lambda function with hardcoded if else rules to select models based on language and token length and store metrics only as CloudWatch Logs. Rely on manual review of CloudWatch dashboards to determine when to update or redeploy routing rules

  5. E

    Use Step Functions Map or Parallel states to branch decisions and insert a metrics evaluation Lambda task that reads historical performance data to compute confidence scores before selecting the model. Implement a cascading workflow that invokes a primary model first and conditionally calls a backup model when confidence or quality thresholds from the metrics store are not met

Xem giải thích

Đáp án

A và E.

  • A — Step Functions với Lambda tiền xử lý tính ngôn ngữ, độ dài token và thuộc tính nội dung, rồi dùng Choice state để định tuyến; lưu metric hiệu năng để điều chỉnh quy tắc
  • E — Dùng Map hoặc Parallel state để rẽ nhánh, chèn Lambda đánh giá metric đọc dữ liệu hiệu năng lịch sử để tính điểm tin cậy, và cascading giữa các model

Vì sao đúng

Đề yêu cầu ba khả năng, và hai đáp án cùng nhau cung cấp đủ:

  1. Định tuyến theo nội dung (ngôn ngữ, độ dài, danh mục)
  2. Chọn model theo metric (độ trễ, chi phí, điểm chất lượng)
  3. Cascading theo độ tin cậy — thử model rẻ trước, chuyển lên model mạnh nếu kết quả kém

A — nền tảng định tuyến động. Lambda tiền xử lý tính các thuộc tính, Choice state rẽ nhánh dựa trên chúng:

{
  "TienXuLy": {
    "Type": "Task",
    "Resource": "arn:aws:lambda:...:function:phan-tich-noi-dung",
    "Next": "ChonModel"
  },
  "ChonModel": {
    "Type": "Choice",
    "Choices": [
      {"And": [
         {"Variable": "$.ngonNgu", "StringEquals": "vi"},
         {"Variable": "$.soToken", "NumericLessThan": 2000}],
       "Next": "ModelNhe"},
      {"Variable": "$.danhMuc", "StringEquals": "rui-ro-cao", "Next": "ModelManh"}
    ],
    "Default": "ModelTieuChuan"
  }
}

Và vế "evolve routing rules over time" được đáp ứng vì ngưỡng nằm trong dữ liệu (DynamoDB hoặc AppConfig), không nằm trong mã.

E — cascading theo độ tin cậy. Đây là phần mà A một mình không có: Lambda đọc dữ liệu hiệu năng lịch sử để tính điểm tin cậy, rồi quyết định có cần leo thang lên model mạnh hơn không:

ModelNhe chạy → điểm tin cậy 0,45 (thấp)
      ↓ cascading
ModelManh chạy lại → điểm tin cậy 0,92 → trả kết quả

Mẫu này tiết kiệm chi phí đáng kể: phần lớn request được model rẻ xử lý xong, chỉ số ít phải leo thang.

Vì sao các phương án khác sai

  • B. Step Functions với Lambda tiền xử lý, gọi Lambda đánh giá metric đọc DynamoDB trước khi định tuyến — rất gần đúng và đáng bàn. Nó có định tuyến theo nội dung và có đọc metric. Nhưng nó thiếu vế cascading mà đề nêu rõ, và nó mô tả một luồng một chiều (đánh giá rồi chọn rồi xong) chứ không có cơ chế leo thang khi kết quả kém.
  • C. Choice state với ngưỡng cố định trong Parameter Store, cập nhật thủ công; metric chỉ dùng để phân tích offline — trượt vế "adapt over time": metric được ghi nhưng không tham gia vào quyết định lúc chạy. Nó chỉ là báo cáo.
  • D. Một Lambda duy nhất với quy tắc if/else hardcode, metric chỉ ghi vào CloudWatch Logs, xem dashboard thủ công — cứng nhắc nhất: mọi thay đổi quy tắc đều cần sửa mã và triển khai lại, và không có gì tự động.

Ghi nhớ

Ba mẫu định tuyến model trong ứng dụng GenAI: | Mẫu | Cách hoạt động | Lợi ích | |---|---|---| | Content-based routing | chọn model theo đặc điểm đầu vào | model phù hợp cho từng loại việc | | Metric-driven routing | chọn theo độ trễ, chi phí, chất lượng đo được | tự thích nghi theo thời gian | | Confidence cascading | model rẻ trước, leo thang khi kết quả kém | tiết kiệm chi phí lớn nhất |

Các state của Step Functions dùng cho định tuyến: | State | Việc | |---|---| | Choice | rẽ nhánh theo điều kiện | | Map | xử lý song song từng phần tử (nhiều claim cùng lúc) | | Parallel | chạy nhiều model cùng lúc để so sánh | | Task | gọi Bedrock, Lambda, hoặc dịch vụ khác | | Retry / Catch | thử lại và xử lý lỗi khai báo |

Nơi lưu quy tắc định tuyến — chọn theo nhu cầu: | Nơi | Đặc điểm | |---|---| | AWS AppConfig | có validation, gradual rollout, rollback — tốt nhất cho quy tắc thay đổi thường xuyên | | DynamoDB | truy vấn nhanh, hợp với dữ liệu metric lịch sử | | Parameter Store | đơn giản, miễn phí, không có rollout controls | | Biến môi trường | phải triển khai lại — tránh |

Và một lưu ý thực dụng về cascading: đặt trần số lần leo thang (thường 2 tầng) và ghi lại tỷ lệ leo thang như một metric — nếu tỷ lệ đó tăng dần, model tầng dưới đang không còn phù hợp và cần xem lại.

Câu 4 Operational Efficiency and Optimization for GenAI Apps

A multimodal image and text analysis pipeline currently calls a single, large foundation model for every request, regardless of whether the user needs a simple label or a detailed explanation, which increases cost and creates bottlenecks during peak load. Internal analysis shows that many requests could be handled by a smaller, cheaper model without affecting accuracy, while only a subset truly requires the advanced reasoning capabilities of the larger model. The platform team suspects they could reduce costs by splitting the workflow into cheaper classification steps and only invoking the more capable model when necessary.

As a GenAI engineer, how should you implement a solution to optimize cost while preserving output quality for complex workloads?

  1. A

    Set up an Amazon SageMaker inference pipeline where a lightweight classifier runs first and forwards its prediction to the multimodal model so that both models contribute to the final answer. Always invoke the multimodal model to confirm or refine the classifier output to maintain consistent response quality

  2. B

    Precompute simple labels for all images offline using a SageMaker training job and store them in Amazon DynamoDB, but still invoke the large multimodal model online for every incoming request to maintain a consistent experience. Use the stored labels only as metadata in responses rather than as a decision point for bypassing the multimodal model

  3. C

    Configure an Amazon SageMaker inference pipeline where a first, cheaper model hosted on SageMaker performs image or text classification and confidence scoring, and only call an Amazon Bedrock multimodal model in a later stage when the request requires detailed reasoning or the classifier has low confidence. Use routing logic in the pipeline container to short circuit the workflow and return early for simple label requests so that the multimodal model is invoked only when needed

  4. D

    Migrate the large multimodal model to a larger instance type on SageMaker and enable Provisioned Throughput on Amazon Bedrock so that all requests continue to call the same model with lower latency. Rely on the increased capacity to absorb traffic without introducing additional routing or staged inference logic

Xem giải thích

Đáp án

C — Dựng SageMaker inference pipeline với model rẻ hơn chạy trước để phân loại và chấm điểm tin cậy, chỉ gọi model multimodal của Bedrock ở giai đoạn sau khi cần.

Vì sao đúng

Đề mô tả đúng vấn đề: mọi request đều gọi model lớn, kể cả những request chỉ cần một nhãn đơn giản — gây tốn chi phí và nghẽn lúc cao điểm.

Và đề cũng nêu sẵn quan sát then chốt: nhiều request có thể được model nhỏ xử lý mà không ảnh hưởng độ chính xác.

Đáp án C áp dụng mẫu model cascading (còn gọi là routing với confidence threshold):

Request → Model nhẹ (SageMaker) phân loại + chấm điểm tin cậy
              ├─ tin cậy CAO  → trả kết quả ngay        (rẻ, nhanh)
              └─ tin cậy THẤP → gọi model multimodal    (đắt, chỉ khi cần)

Hiệu quả kinh tế rất rõ:

Trước: 100% request × giá model lớn
Sau:    80% request × giá model nhỏ  +  20% × (giá nhỏ + giá lớn)
     ≈ giảm 60–70% chi phí

Và vế "preserving output quality for complex workloads" được giữ nguyên vì những request thật sự phức tạp vẫn được model mạnh xử lý — chỉ khác là chúng được nhận diện trước, thay vì mặc định gửi tất cả.

Điểm tin cậy là cơ chế quyết định:

ket_qua = model_nhe.predict(du_lieu)
if ket_qua['confidence'] >= 0.85:
    return ket_qua['label']              # xong, không gọi model lớn
else:
    return bedrock.converse(modelId=MODEL_MULTIMODAL, messages=[...])

Vì sao các phương án khác sai

  • A. Inference pipeline với classifier chạy trước, nhưng LUÔN gọi model multimodal để cả hai cùng đóng góp — không tiết kiệm gì cả: model đắt vẫn chạy ở mọi request, và giờ còn cộng thêm chi phí của classifier. Cụm "always invoke" là điểm loại.
  • B. Tính sẵn nhãn offline bằng SageMaker training job, lưu vào DynamoDB, nhưng vẫn gọi model lớn cho mọi request online — cùng vấn đề: chi phí online không đổi. Và tính sẵn offline không xử lý được ảnh mới mà người dùng vừa tải lên.
  • D. Chuyển model lớn sang instance to hơn và bật Provisioned Throughput — giải quyết nghẽn nhưng LÀM CHI PHÍ TĂNG: instance lớn hơn đắt hơn, và Provisioned Throughput là cam kết trả tiền cho năng lực giữ sẵn. Đề yêu cầu tối ưu chi phí.

Ghi nhớ

Ba mẫu tối ưu chi phí cho ứng dụng GenAI: | Mẫu | Cách hoạt động | Tiết kiệm | |---|---|---| | Model cascading | model rẻ trước, leo thang khi cần | lớn nhất | | Semantic caching | tái dùng kết quả cho câu hỏi tương tự | lớn với câu hỏi lặp lại | | Prompt compression | rút gọn ngữ cảnh trước khi gửi | vừa | | Batch inference | gom nhiều item vào một request | lớn với workload theo lô |

SageMaker inference pipeline là cơ chế ghép nhiều bước xử lý vào một endpoint:

Request → Container 1 (tiền xử lý) → Container 2 (model) → Container 3 (hậu xử lý) → Response

Ưu điểm: một lời gọi duy nhất, không phải điều phối nhiều endpoint, và dữ liệu không rời khỏi endpoint giữa các bước.

Ba cách chọn ngưỡng tin cậy: | Cách | Đặc điểm | |---|---| | Ngưỡng cố định | đơn giản, phải hiệu chỉnh bằng tay | | Ngưỡng theo danh mục | chính xác hơn — ảnh y tế cần ngưỡng cao hơn ảnh sản phẩm | | Ngưỡng động theo metric | tự điều chỉnh dựa trên tỷ lệ sai đo được |

Và hai chỉ số cần theo dõi khi vận hành cascading: | Chỉ số | Ý nghĩa | |---|---| | Tỷ lệ leo thang | bao nhiêu % phải gọi model lớn — nếu tăng dần thì model nhỏ đang lệch | | Tỷ lệ sai của tầng nhẹ | có bao nhiêu kết quả tin cậy cao nhưng thực ra sai |

Chỉ số thứ hai quan trọng nhất: nó cho biết ngưỡng đang đặt quá thấp — tiết kiệm được tiền nhưng đánh đổi bằng chất lượng, đúng thứ đề bảo phải giữ.

Câu 5 Chọn nhiều đáp án Foundation Model Integration, Data Management and Compliance

A healthcare provider manages a large library of patient education material that is frequently revised based on new clinical guidance, regulatory updates, and quality reviews. Some updates involve minor text edits that should trigger lightweight incremental updates to the vector store, while more urgent corrections must update embeddings immediately to avoid distributing outdated information to patients. Additionally, the provider requires a complete vector index rebuild every quarter to eliminate drift, fragmentation, and outdated embeddings.

What maintenance architecture should the team deploy to keep the vector store current? (Select two)

  1. A

    Use an hourly AWS Batch job to reprocess all documents and regenerate embeddings, with manual triggers for urgent updates

  2. B

    Use an EventBridge scheduled rule and an AWS Step Functions workflow to perform quarterly full index rebuilds

  3. C

    Use Amazon Bedrock Knowledge Bases with a quarterly document sync job for full index rebuild

  4. D

    Use Amazon OpenSearch Service ingestion pipelines for all changes and rely on scheduled cron based reprocessing for full index replacement

  5. E

    Use Amazon EventBridge rules and AWS Lambda to process incremental updates and real-time corrections

Xem giải thích

Đáp án

B và E.

  • E — Dùng EventBridge rule và Lambda để xử lý cập nhật gia tăng và sửa chữa khẩn cấp theo thời gian thực
  • B — Dùng EventBridge scheduled rule và Step Functions để dựng lại toàn bộ index theo quý

Vì sao đúng

Đề mô tả ba loại cập nhật với ba mức khẩn cấp khác nhau, và đáp án phải phủ hết: | Loại cập nhật | Yêu cầu | Đáp án | |---|---|---| | Sửa nhỏ về câu chữ | cập nhật gia tăng, nhẹ nhàng | E | | Sửa khẩn cấp | cập nhật embedding NGAY LẬP TỨC | E | | Dựng lại toàn bộ index | theo quý, loại bỏ drift và phân mảnh | B |

E — luồng sự kiện cho cập nhật gia tăng và khẩn cấp. EventBridge phản ứng ngay khi tài liệu thay đổi, Lambda tính lại embedding cho đúng tài liệu đó:

{
  "source": ["cms.tai-lieu-benh-nhan"],
  "detail-type": ["Tài liệu được cập nhật"],
  "detail": {"mucDoKhan": ["thuong", "khan-cap"]}
}
def lambda_handler(event, context):
    doc_id = event['detail']['docId']
    noi_dung = lay_tai_lieu(doc_id)
    embedding = bedrock.invoke_model(modelId='amazon.titan-embed-text-v2:0',
                                     body=json.dumps({'inputText': noi_dung}))
    opensearch.update(index='tai-lieu', id=doc_id, body={'doc': {'embedding': embedding}})

Vì nó hướng sự kiện, sửa chữa khẩn cấp có hiệu lực trong vài giây — đúng yêu cầu "avoid distributing outdated information to patients".

B — dựng lại toàn bộ theo lịch. Step Functions điều phối một quy trình nhiều bước, chạy theo lịch quý:

EventBridge (cron: quý)
    ↓
Step Functions:
    ① Tạo index mới
    ② Map state: tính lại embedding cho MỌI tài liệu (song song)
    ③ Kiểm tra chất lượng index mới
    ④ Chuyển alias sang index mới (không gián đoạn)
    ⑤ Xoá index cũ

Bước ④ đáng chú ý: dùng alias của OpenSearch để chuyển sang index mới nguyên tử, người dùng không thấy khoảng trống nào.

Vì sao các phương án khác sai

  • A. AWS Batch chạy MỖI GIỜ tính lại embedding cho MỌI tài liệu, kèm trigger thủ công cho cập nhật khẩn — cực kỳ lãng phí: tính lại toàn bộ thư viện mỗi giờ trong khi chỉ vài tài liệu thay đổi. Và "trigger thủ công" trái yêu cầu cập nhật khẩn cấp phải tự động và tức thì.
  • C. Bedrock Knowledge Bases với job đồng bộ theo quý — chỉ phủ được vế dựng lại toàn bộ, thiếu hẳn cơ chế cập nhật gia tăng và khẩn cấp. Sửa chữa khẩn sẽ phải chờ tới quý sau.
  • D. OpenSearch ingestion pipeline cho mọi thay đổi, cộng cron reprocessing để thay index — gần đúng nhưng thiếu vế khẩn cấp có ưu tiên: nó đối xử mọi thay đổi như nhau. Và cron thuần không có khả năng điều phối nhiều bước (kiểm tra chất lượng, chuyển alias, rollback) như Step Functions.

Ghi nhớ

Ba chiến lược bảo trì vector store — thường dùng kết hợp: | Chiến lược | Kích hoạt bởi | Phạm vi | |---|---|---| | Incremental update | sự kiện thay đổi tài liệu | một tài liệu | | Scheduled full rebuild | lịch (tuần, quý) | toàn bộ index | | On-demand rebuild | thủ công hoặc khi đổi embedding model | toàn bộ |

Vì sao cần cả hai — ba lý do khiến rebuild định kỳ không thể bỏ: | Vấn đề | Chi tiết | |---|---| | Drift | tài liệu bị xoá ở nguồn nhưng vector còn sót lại | | Phân mảnh | index bị phân mảnh sau nhiều lần cập nhật, làm chậm truy vấn | | Đổi embedding model | model mới sinh vector không tương thích với vector cũ |

Dòng cuối rất quan trọng: không bao giờ trộn vector từ hai model embedding khác nhau trong một index — khoảng cách giữa chúng vô nghĩa. Đổi model nghĩa là bắt buộc rebuild toàn bộ.

Kỹ thuật blue/green cho index — cách rebuild không gián đoạn:

1. Tạo index mới: tai-lieu-v2
2. Nạp đầy đủ vào tai-lieu-v2
3. Kiểm tra chất lượng
4. Chuyển alias "tai-lieu" từ v1 sang v2  ← nguyên tử
5. Giữ v1 vài ngày rồi xoá

Ứng dụng luôn truy vấn qua alias, nên nó không biết gì về việc chuyển đổi — và rollback chỉ là chuyển alias ngược lại.

Và với dữ liệu y tế như đề mô tả, nhớ thêm hai điều: bật mã hoá at-rest cho vector store, và ghi vết kiểm toán mọi lần cập nhật embedding — vì nội dung hướng dẫn bệnh nhân là dữ liệu chịu quản lý.

Câu 6 Foundation Model Integration, Data Management and Compliance

A product team is building a generative AI powered application that serves multiple business units with different cost, latency, and quality requirements. The application must support using different foundation models and even switching between providers as business priorities evolve. Product managers want to experiment with new models, gradually expose them to a subset of users, and quickly roll back if output quality or latency degrades, all while keeping the core application logic stable.

Engineering leadership has also emphasized that operational changes should not require frequent redeployments, since the system runs across multiple environments and regions. They want configuration changes to be centrally managed, safely rolled out, and observable, so that teams can respond quickly to performance, cost, or compliance signals without introducing unnecessary risk.

Which approach represents the BEST way to design the architecture to enable dynamic model selection and provider switching without requiring code modifications?

  1. A

    Implement an abstraction layer in AWS Lambda that reads model selection rules from AWS AppConfig. Configure provider endpoints by leveraging the Lambda function code. Update AppConfig to toggle feature flags while keeping provider specific logic fixed in code. Redeploy the function when switching providers to ensure consistency

  2. B

    Use Amazon API Gateway to route all inference requests to an AWS Lambda function that reads model and provider selection from AWS AppConfig at runtime. Update routing rules and configuration values in AppConfig to change models or providers without redeploying code. Control rollout and rollback by using AppConfig deployment strategies and validators

  3. C

    Configure Amazon API Gateway with static stage variables that define the model and provider to invoke. Manually update the stage variables during a release window to switch models for all users at once. Rely on API Gateway caching to minimize latency during changes

  4. D

    Store the selected model and provider as environment variables in AWS Lambda and update them whenever a switch is required. Redeploy the Lambda function to propagate the new configuration to all environments. Use Lambda versions and aliases to manage traffic shifting between configurations

Xem giải thích

Đáp án

B — Dùng API Gateway định tuyến mọi request tới Lambda, và Lambda đọc lựa chọn model và provider từ AWS AppConfig NGAY LÚC CHẠY.

Vì sao đúng

Đề nêu bốn yêu cầu, và AppConfig đáp ứng cả bốn:

  1. Đổi model và nhà cung cấp mà giữ nguyên logic ứng dụng
  2. Phát hành dần cho một phần người dùng, rollback nhanh
  3. KHÔNG phải triển khai lại khi đổi cấu hình
  4. Quản lý tập trung, có kiểm soát và quan sát được

AWS AppConfig là dịch vụ quản lý cấu hình động — nó tách cấu hình khỏi mã, và quan trọng nhất là đọc lúc chạy:

import boto3, json
appconfig = boto3.client('appconfigdata')

def lambda_handler(event, context):
    cau_hinh = json.loads(lay_cau_hinh_appconfig())   # đọc lúc chạy, có cache
    model_id = chon_model(cau_hinh, event['nguoiDung'])

    return bedrock.converse(modelId=model_id, messages=event['messages'])

Cấu hình nằm ngoài mã, đổi được bất cứ lúc nào:

{
  "models": {
    "mac_dinh": "anthropic.claude-3-5-haiku-20241022-v1:0",
    "premium": "anthropic.claude-3-5-sonnet-20241022-v2:0",
    "thu_nghiem": "amazon.nova-pro-v1:0"
  },
  "phan_bo": {"premium": 10, "mac_dinh": 90}
}

Và ba tính năng của AppConfig khớp đúng các yêu cầu còn lại: | Tính năng | Đáp ứng yêu cầu | |---|---| | Deployment strategy | phát hành dần theo % — Linear, Canary, AllAtOnce | | Validator (JSON Schema hoặc Lambda) | chặn cấu hình sai trước khi phát hành | | Automatic rollback khi CloudWatch alarm kêu | rollback nhanh khi chất lượng hoặc độ trễ xấu đi |

Tính năng thứ hai đặc biệt quan trọng: cấu hình sai (model ID không tồn tại) sẽ bị chặn ngay lúc phát hành, không phải sau khi đã hỏng production.

Vì sao các phương án khác sai

  • A. Lambda đọc quy tắc từ AppConfig, nhưng cấu hình endpoint của provider nằm trong MÃ Lambda — đây là phương án gần nhất và chỉ sai một chỗ, nhưng chỗ đó quan trọng: thêm nhà cung cấp mới vẫn phải sửa mã và triển khai lại. Đề nói rõ hệ thống phải hỗ trợ "switching between providers" mà không cần redeploy.
  • C. API Gateway với stage variable tĩnh, cập nhật thủ công trong cửa sổ phát hành — không có phát hành dần: stage variable đổi là áp cho 100% người dùng cùng lúc. Và "manually update during a release window" trái hẳn yêu cầu thay đổi nhanh và an toàn.
  • D. Lưu model và provider trong biến môi trường của Lambda, triển khai lại khi cần đổi — vi phạm thẳng yêu cầu "should not require frequent redeployments". Và với nhiều môi trường và nhiều Region, mỗi lần đổi là hàng chục lần triển khai.

Ghi nhớ

So sánh các nơi lưu cấu hình: | Nơi | Đổi không cần deploy | Gradual rollout | Validation | Auto rollback | |---|---|---|---|---| | AppConfig | ✅ | ✅ | ✅ | ✅ | | Parameter Store | ✅ | ❌ | ❌ | ❌ | | DynamoDB | ✅ | tự viết | tự viết | tự viết | | Biến môi trường Lambda | ❌ | ❌ | ❌ | ❌ | | Stage variable của API Gateway | ✅ | ❌ | ❌ | ❌ |

Bốn deployment strategy của AppConfig: | Strategy | Hành vi | |---|---| | AllAtOnce | 100% ngay | | Linear50PercentEvery30Seconds | tăng đều | | Canary10Percent20Minutes | 10% trước, theo dõi, rồi phần còn lại | | Tuỳ chỉnh | tự khai tỷ lệ, khoảng thời gian, thời gian nướng (bake time) |

Ba lưu ý khi dùng AppConfig với Lambda:

  1. Dùng AppConfig Lambda Extension thay vì gọi API trực tiếp — nó cache cấu hình cục bộ và tự làm mới định kỳ, nên không tốn một lời gọi API ở mỗi lần chạy hàm.
  2. Đặt CloudWatch alarm làm điều kiện rollback — ví dụ alarm trên tỷ lệ lỗi hoặc độ trễ p99 của model mới.
  3. Bake time là khoảng chờ sau khi phát hành xong trước khi coi là thành công — đặt đủ dài để alarm kịp phát hiện vấn đề.

Và mẫu kiến trúc đầy đủ cho ứng dụng đa model:

API Gateway → Lambda (abstraction layer)
                 ↓ đọc AppConfig (cache)
              chọn model theo phân khúc người dùng
                 ↓
              Bedrock Converse API (giao diện thống nhất cho mọi model)

Converse API đáng nhắc riêng: nó cho một giao diện chung cho mọi model trên Bedrock, nên đổi từ Claude sang Nova hay Llama không phải sửa cách gọi — chỉ đổi modelId.

Câu 7 Chọn nhiều đáp án Operational Efficiency and Optimization for GenAI Apps

A global e-commerce company operates a conversational shopping assistant that helps customers discover products, understand return policies, and navigate catalog information across multiple Regions. Traffic analysis shows that customers frequently ask nearly identical questions with small wording variations, yet each request triggers a fresh model invocation, driving up token consumption and increasing operational costs. At the same time, users who are geographically distant from the primary hosting Region experience noticeably higher latency during peak shopping periods, especially during seasonal sales and major promotional events.

As the lead GenAI architect, how should you implement a solution so that similar requests reuse prior results and frequent responses are served closer to end users? (Select two)

  1. A

    Build a semantic cache using Amazon OpenSearch Service to store embeddings of user queries and reuse prior responses when a new query is semantically similar. Place the GenAI API behind Amazon CloudFront so that frequently accessed responses are served from edge locations near end users

  2. B

    Use Amazon Bedrock embeddings to compute vector representations of incoming questions and cache the resulting responses in Amazon ElastiCache for low latency retrieval. Configure CloudFront regional edge caches to serve popular API results globally, reducing latency across distant geographies

  3. C

    Build a caching layer in Amazon OpenSearch Service that stores embeddings of entire product descriptions and use keyword boosted vector queries to retrieve potential matches before generating a new answer. Front the GenAI API with Amazon CloudFront so that cached API responses are returned when the same endpoint is accessed again

  4. D

    Store all user queries in Amazon DynamoDB and rely on exact string matching to determine whether a response can be reused. Enable API Gateway response caching to return previously computed answers only when the same request path is invoked

  5. E

    Create an OpenSearch index that stores plain text versions of prompts and rely on keyword matching to identify potential cache hits. Front the API with CloudFront but set very short TTL values so that responses stay fresh for global users

Xem giải thích

Đáp án

A và B.

  • A — Dựng semantic cache bằng OpenSearch lưu embedding của câu hỏi, tái dùng câu trả lời khi câu mới tương tự về ngữ nghĩa; đặt API sau CloudFront
  • B — Dùng embedding của Bedrock tính vector câu hỏi và cache kết quả trong ElastiCache; cấu hình CloudFront regional edge cache

Vì sao đúng

Đề nêu hai vấn đề riêng biệt, và mỗi đáp án phải giải quyết cả hai: | Vấn đề | Giải pháp | |---|---| | Câu hỏi gần giống nhau nhưng khác chữ, mỗi lần đều gọi model | semantic cache (dựa trên embedding) | | Người dùng ở xa Region chính có độ trễ cao | CloudFront (cache ở edge) |

Vì sao phải là semantic cache, không phải cache thông thường:

"Chính sách đổi trả của shop thế nào?"
"Cho mình hỏi về việc đổi hàng"
"Đổi trả hàng có được không?"
        ↓
Ba chuỗi KHÁC NHAU hoàn toàn → cache theo chuỗi luôn MISS
        ↓
Ba vector RẤT GẦN NHAU → semantic cache HIT

Cách hoạt động:

vector = bedrock.invoke_model(modelId='amazon.titan-embed-text-v2:0',
                              body=json.dumps({'inputText': cau_hoi}))

ket_qua = kho_vector.knn_search(vector, k=1)
if ket_qua and ket_qua[0]['score'] >= 0.95:      # ngưỡng tương đồng
    return ket_qua[0]['cau_tra_loi']              # HIT — không gọi model

tra_loi = bedrock.converse(...)                   # MISS
kho_vector.luu(vector, cau_hoi, tra_loi)

A và B khác nhau ở nơi lưu cache, và cả hai đều hợp lệ: | | A: OpenSearch | B: ElastiCache | |---|---|---| | Tìm kiếm vector | k-NN dựng sẵn, quy mô lớn | vector search từ Redis 7.x | | Độ trễ | mili giây | dưới mili giây | | Quy mô | hàng triệu vector | giới hạn bởi RAM | | TTL tự động | qua ILM policy | SETEX — đơn giản hơn |

Và cả hai đều dùng CloudFront cho vế độ trễ địa lý — với regional edge cache (nêu trong B) là lớp cache trung gian giữa edge location và origin, làm tăng tỷ lệ trúng cache cho nội dung ít phổ biến hơn.

Vì sao các phương án khác sai

  • D. Lưu mọi câu hỏi vào DynamoDB và dùng khớp chuỗi CHÍNH XÁC; bật API Gateway response caching — khớp chuỗi chính xác không giải quyết được vấn đề đề nêu: đề nói rõ câu hỏi "khác nhau về cách diễn đạt". Cache này sẽ miss gần như mọi lần.
  • E. Index OpenSearch lưu văn bản thuần của prompt, dùng khớp từ khoá; CloudFront với TTL rất ngắn — cùng vấn đề: khớp từ khoá không phải khớp ngữ nghĩa. Và TTL rất ngắn phá hỏng chính mục đích của CloudFront — cache gần như không bao giờ được dùng.
  • C. Cache embedding của TOÀN BỘ mô tả sản phẩm, dùng truy vấn vector có boost từ khoá — nhầm đối tượng cache: đây là mô tả của retrieval (tìm tài liệu liên quan), không phải caching câu trả lời. Nó không tránh được lần gọi model nào.

Ghi nhớ

Ba tầng cache cho ứng dụng GenAI: | Tầng | Cache gì | Tiết kiệm | |---|---|---| | CloudFront | response HTTP ở edge | độ trễ địa lý + băng thông | | Semantic cache | câu trả lời theo ngữ nghĩa câu hỏi | token của model — lớn nhất | | Prompt caching (Bedrock) | phần prefix cố định của prompt | token của phần ngữ cảnh lặp lại |

Dòng cuối đáng biết riêng: Bedrock prompt caching cache phần đầu cố định của prompt (system prompt, tài liệu ngữ cảnh dài) ở phía dịch vụ — giảm cả chi phí lẫn độ trễ mà không cần hạ tầng nào:

messages=[{"role": "user", "content": [
    {"text": tai_lieu_dai, "cachePoint": {"type": "default"}},   # phần này được cache
    {"text": cau_hoi_moi}                                         # phần này thay đổi
]}]

Ba tham số quan trọng khi thiết kế semantic cache: | Tham số | Ảnh hưởng | |---|---| | Ngưỡng tương đồng | quá thấp ⇒ trả sai câu trả lời; quá cao ⇒ ít trúng | | TTL | dữ liệu nghiệp vụ đổi thì câu trả lời cũ phải hết hạn | | Khoá cache bao gồm ngữ cảnh | người dùng khác nhau, ngôn ngữ khác nhau ⇒ cache riêng |

Dòng cuối là rủi ro lớn nhất: nếu khoá cache không bao gồm danh tính hoặc quyền hạn, người dùng này có thể nhận câu trả lời chứa dữ liệu của người khác. Với trợ lý mua sắm thì rủi ro thấp, nhưng với ứng dụng có dữ liệu riêng tư thì đây là lỗ hổng nghiêm trọng.

Và ngưỡng tương đồng nên hiệu chỉnh bằng dữ liệu thật: lấy vài trăm cặp câu hỏi, đo điểm tương đồng, rồi chọn ngưỡng cân bằng giữa tỷ lệ trúng và tỷ lệ trả sai.

Câu 8 Foundation Model Integration, Data Management and Compliance

A travel planning assistant allows users to ask multi part questions such as "Find me family friendly beach destinations near a warm city in April, then suggest three hotels with kids' clubs and create a day by day itinerary." The current implementation sends the raw user query directly into the retrieval layer, which leads to noisy results and incomplete constraint matching. Product management wants the system to better interpret user intent, break complex questions into smaller retrieval steps, and reformulate queries so that more relevant documents are retrieved before generating a final response.

How should you structure the query handling and transformation pipeline in this scenario?

  1. A

    Implement a multi stage query processing pipeline with AWS Lambda and Step Functions that breaks the user query into intent specific subqueries, performs query expansion using an Amazon Bedrock model, and issues targeted searches through the vector store. Then aggregate the retrieved context and pass the structured results to the foundation model for final synthesis

  2. B

    Store every query and response pair from past users in DynamoDB and use a similarity search over past queries to retrieve the closest match. Then use the matched result as a template for constructing a new answer without performing layered query processing.

  3. C

    Call the foundation model first to generate a preliminary itinerary, extract keywords from that output with Lambda, and run a second retrieval step using those extracted keywords. Then use the reranked results to refine the FM response

  4. D

    Use a Lambda function to rewrite the user query with an Amazon Bedrock model and then pass the rewritten query directly to OpenSearch Serverless for a single vector similarity call. Then route the top K results to the foundation model without performing any additional decomposition

Xem giải thích

Đáp án

A — Dựng pipeline xử lý truy vấn nhiều tầng bằng Lambda và Step Functions: tách câu hỏi thành các subquery theo ý định, mở rộng truy vấn bằng model Bedrock, tìm kiếm có mục tiêu trên vector store, rồi gộp ngữ cảnh đưa vào model để tổng hợp.

Vì sao đúng

Đề đưa ra một câu hỏi nhiều phần, và đó là điểm mấu chốt:

"Tìm điểm đến biển thân thiện với gia đình gần một thành phố ấm áp vào tháng Tư,
 rồi gợi ý ba khách sạn có câu lạc bộ trẻ em,
 và lập lịch trình theo từng ngày."

Đây thực chất là ba nhu cầu truy xuất khác nhau: | Phần | Cần tìm gì | |---|---| | ① | điểm đến biển + thân thiện gia đình + khí hậu tháng Tư | | ② | khách sạn + tiện ích câu lạc bộ trẻ em + tại điểm đến đó | | ③ | thông tin hoạt động để lập lịch trình |

Gửi nguyên câu này vào vector store — như hệ thống hiện tại đang làm — cho ra một vector "trung bình" của cả ba ý, nên kết quả nhiễu và thiếu ràng buộc, đúng như đề mô tả.

Query decomposition giải quyết bằng cách tách ra rồi tìm riêng:

Câu hỏi gốc
    ↓ ① Lambda + Bedrock: TÁCH thành subquery theo ý định
    ["điểm đến biển gia đình khí hậu tháng Tư",
     "khách sạn có câu lạc bộ trẻ em tại {điểm đến}",
     "hoạt động cho gia đình tại {điểm đến}"]
    ↓ ② Query expansion: thêm từ đồng nghĩa, thuật ngữ liên quan
    ↓ ③ Tìm kiếm CÓ MỤC TIÊU cho từng subquery
    ↓ ④ Gộp ngữ cảnh
    ↓ ⑤ FM tổng hợp câu trả lời cuối

Step Functions là công cụ đúng cho luồng này vì nó cho Map state chạy các subquery song song, cộng với Retry và Catch khai báo — và subquery ② phụ thuộc kết quả của ①, nên cần điều phối có thứ tự.

Vì sao các phương án khác sai

  • D. Lambda viết lại truy vấn bằng Bedrock rồi gửi một lần tới OpenSearch, không phân rã — đây là phương án gần nhất, và query rewriting có ích thật. Nhưng nó vẫn là MỘT truy vấn cho BA nhu cầu — vấn đề gốc không được giải quyết. Cụm "without performing any additional decomposition" là điểm loại.
  • C. Gọi FM trước để sinh lịch trình sơ bộ, trích từ khoá, rồi tìm lần hai — thứ tự ngược và rủi ro: FM sinh nội dung khi chưa có dữ liệu thật, nên nó bịa ra điểm đến và khách sạn, rồi bạn đi tìm tài liệu cho những thứ bịa đó. Đây là công thức tạo ảo giác.
  • B. Lưu mọi cặp câu hỏi–câu trả lời cũ vào DynamoDB, tìm câu gần nhất rồi dùng làm mẫu — không phải xử lý truy vấn mà là caching, và nó không đúng cho câu hỏi mới: hai chuyến đi khác nhau cần câu trả lời khác nhau, dùng lại mẫu cũ sẽ cho thông tin sai.

Ghi nhớ

Bốn kỹ thuật biến đổi truy vấn trong RAG: | Kỹ thuật | Việc | |---|---| | Query decomposition | tách câu hỏi phức thành nhiều subquery | | Query expansion | thêm từ đồng nghĩa, thuật ngữ liên quan để tăng recall | | Query rewriting | viết lại cho rõ ràng, bỏ từ thừa | | HyDE | sinh câu trả lời giả định rồi tìm bằng embedding của nó |

Nhận dạng khi nào cần decomposition: | Dấu hiệu trong câu hỏi | Cần tách? | |---|---| | "rồi", "và sau đó", nhiều mệnh đề | ✅ | | Nhiều ràng buộc thuộc các chiều khác nhau | ✅ | | Câu hỏi đơn, một ý | ❌ — tách chỉ thêm độ trễ |

Đó cũng là đánh đổi cần cân nhắc: decomposition tốn thêm lời gọi model và thêm độ trễ. Nên trong thực tế, thường có một bước phân loại độ phức tạp trước — câu đơn giản đi thẳng, câu phức mới qua pipeline đầy đủ.

Và bốn thành phần của một pipeline RAG hoàn chỉnh:

① Query understanding   → phân rã, mở rộng, viết lại
② Retrieval             → vector + keyword + metadata filter
③ Reranking             → cross-encoder sắp xếp lại top-K
④ Generation            → FM tổng hợp từ ngữ cảnh đã chọn

Mỗi tầng cải thiện một khía cạnh khác nhau: ① và ② lo recall, ③ lo precision, ④ lo chất lượng diễn đạt. Bỏ tầng nào thì điểm yếu tương ứng lộ ra — và đề này mô tả đúng triệu chứng của việc thiếu tầng ①.

Câu 9 AI Safety, Security, and Governance

A generative AI analytics application queries a centralized data lake that contains sensitive business data shared across multiple teams. Different users are allowed to see different values within the same dataset depending on attributes such as geography, role, or business unit. During reviews, the security team finds that coarse access controls expose more data than intended for some users. The platform team wants to enforce fine grained data access so that only authorized values within specific rows and specific columns are visible to each request. The solution must integrate cleanly with existing analytics engines and avoid duplicating datasets or building custom filtering logic in application code.

What do you recommend?

  1. A

    Use AWS Lake Formation to apply column-level permissions and combine them with row-level locking semantics to prevent unauthorized access. Configure analytics engines to respect these locks during query execution. Manage exceptions through IAM policies

  2. B

    Store data in Amazon S3 tables and apply IAM policies that restrict access to specific object prefixes representing rows and columns. Enforce access control by reorganizing data layout and controlling S3 permissions. Update application queries to target only allowed prefixes

  3. C

    Use AWS Lake Formation to grant column-level permissions and require application queries to apply WHERE clauses that filter rows dynamically based on user identity. Validate access by reviewing query logs and audit trails. Enforce consistency through development guidelines

  4. D

    Use AWS Lake Formation to define cell-level security policies that restrict access to specific combinations of rows and columns based on user attributes. Grant permissions through Lake Formation so query engines enforce these rules automatically at query time

Xem giải thích

Đáp án

D — Dùng AWS Lake Formation định nghĩa cell-level security policy giới hạn theo tổ hợp dòng và cột cụ thể dựa trên thuộc tính người dùng; cấp quyền qua Lake Formation để query engine tự thực thi lúc chạy.

Vì sao đúng

Đề nêu ba yêu cầu, và cell-level security của Lake Formation đáp ứng cả ba:

  1. Chỉ những GIÁ TRỊ được phép trong dòng và cột cụ thể mới hiển thị
  2. Tích hợp sẵn với các engine phân tích hiện có
  3. Không nhân bản dữ liệu, không tự viết logic lọc trong mã ứng dụng

Cell-level security là giao của hai chiều — chính xác là điều đề mô tả:

              cot_A   cot_B   cot_C(nhạy cảm)
dong_VN         ✓       ✓          ✓        ← người dùng VN thấy đủ
dong_SG         ✓       ✓          ✗        ← thấy dòng nhưng ẩn cột nhạy cảm
dong_US         ✗       ✗          ✗        ← không thấy dòng nào

Cấu hình gồm data filter kết hợp cả hai chiều:

aws lakeformation create-data-cells-filter --table-data '{
  "TableCatalogId": "123456789012",
  "DatabaseName": "kho_du_lieu",
  "TableName": "giao_dich",
  "Name": "loc_theo_khu_vuc",
  "RowFilter": {"FilterExpression": "khu_vuc = '\''APAC'\''"},
  "ColumnWildcard": {"ExcludedColumnNames": ["so_the", "cmnd"]}
}'

aws lakeformation grant-permissions \
  --principal DataLakePrincipalIdentifier=arn:aws:iam::...:role/PhanTichAPAC \
  --resource '{"DataCellsFilter": {"DatabaseName":"kho_du_lieu",
               "TableName":"giao_dich","Name":"loc_theo_khu_vuc"}}' \
  --permissions SELECT

Điểm mấu chốt cho yêu cầu thứ hai và thứ ba: Athena, Redshift Spectrum, EMR và Glue đều tôn trọng quyền Lake Formation tự động. Ứng dụng chạy SELECT * FROM giao_dich và engine tự lọc — không cần WHERE clause nào, không cần bản sao dữ liệu nào.

Vì sao các phương án khác sai

  • C. Lake Formation cấp quyền cột, và yêu cầu truy vấn ứng dụng tự thêm WHERE để lọc dòng theo danh tính — đây là phương án gần nhất và cần phân biệt kỹ. Nó vi phạm thẳng yêu cầu "avoid building custom filtering logic in application code", và tệ hơn: bảo mật phụ thuộc vào việc lập trình viên nhớ thêm WHERE. Quên một chỗ là lộ dữ liệu, và "validate bằng cách xem query log" là phát hiện sau khi việc đã xảy ra.
  • A. Lake Formation với quyền cột kết hợp "row-level locking semantics" — khái niệm không tồn tại: locking là cơ chế điều khiển ĐỒNG THỜI (tránh xung đột khi nhiều giao dịch cùng sửa dữ liệu), không phải cơ chế kiểm soát truy cập. Lake Formation có row-level filter, không có "row-level lock".
  • B. Lưu dữ liệu thành nhiều prefix trên S3 rồi dùng IAM policy giới hạn theo prefix — vi phạm yêu cầu "không nhân bản dữ liệu": bạn phải tổ chức lại toàn bộ layout theo mọi tổ hợp quyền có thể. Với ba thuộc tính (địa lý, vai trò, đơn vị kinh doanh), số tổ hợp bùng nổ. Và nó không lọc được theo CỘT.

Ghi nhớ

Ba mức kiểm soát truy cập của Lake Formation: | Mức | Giới hạn | |---|---| | Table-level | thấy hay không thấy cả bảng | | Column-level | ẩn một số cột | | Row-level | chỉ thấy dòng thoả điều kiện | | Cell-level | GIAO của hai mức trên ← câu này |

Các engine tôn trọng quyền Lake Formation tự động: | Engine | Hỗ trợ | |---|---| | Amazon Athena | ✅ | | Redshift Spectrum | ✅ | | AWS Glue ETL | ✅ | | Amazon EMR | ✅ (với runtime role) | | Amazon QuickSight | ✅ | | SageMaker | ✅ qua Glue catalog |

Đó là lý do đáp án nói "query engines enforce these rules automatically at query time" — bạn cấu hình một lần ở Lake Formation, mọi engine đều tuân theo.

Và tag-based access control (LF-Tags) là cách mở rộng đáng biết khi số lượng bảng lớn: thay vì cấp quyền cho từng bảng, bạn gắn tag cho bảng và cột, rồi cấp quyền theo tag:

Gắn tag: do_nhay_cam = cao  →  cột so_the, cmnd
Cấp quyền: role PhanTich được đọc do_nhay_cam = thap

Cách này giảm số policy phải quản lý từ hàng nghìn xuống vài chục — rất phù hợp với data lake nhiều bảng.

Ba lưu ý khi triển khai:

  1. Phải chuyển sang mô hình quyền Lake Formation — tắt IAM-only access cho catalog.
  2. Ứng dụng phải chạy bằng role được cấp quyền qua Lake Formation, không phải quyền S3 trực tiếp.
  3. Kiểm tra bằng chính engine — quyền có thể trông đúng trong Console nhưng cấu hình sai ở tầng IAM.
Câu 10 Operational Efficiency and Optimization for GenAI Apps

A customer facing application provides real time AI generated responses as users type queries into a chat style interface. User research shows that perceived responsiveness matters more than receiving the entire response at once, especially for longer answers. The current implementation waits for the full model output before returning a response, which leads to noticeable delays and a poor user experience during peak usage. The engineering team wants to redesign the interaction pattern to deliver the first tokens as quickly as possible while still supporting scalable, managed APIs. They must choose an approach that minimizes end user latency without over engineering the frontend integration.

What is the MOST operationally efficient way to improve responsiveness for time sensitive interactions?

  1. A

    Use Amazon Bedrock InvokeModelWithStreaming and push tokens to clients using Amazon API Gateway WebSocket API connections. Maintain persistent bidirectional connections and manage connection state for each client. Close the connection after the full response is generated

  2. B

    Use Amazon Bedrock InvokeModelWithStreaming to stream tokens as they are generated by the foundation model. Forward the streamed output through Amazon API Gateway HTTP API so clients receive partial responses over HTTP immediately. Optimize performance by reducing payload size and using faster HTTP integrations

  3. C

    Use Amazon Bedrock InvokeModelWithStreaming to stream tokens as they are generated by the foundation model. Forward the streamed output through Amazon API Gateway REST API response streaming so clients receive partial responses over HTTP immediately. Avoid buffering responses in the backend to minimize time to first token

  4. D

    Use Amazon Bedrock InvokeModel to generate the full response and return it through Amazon API Gateway HTTP API. Optimize performance by reducing payload size and using faster HTTP integrations. Leverage client side loading indicators to manage perceived latency

Xem giải thích

Đáp án

C — Dùng InvokeModelWithResponseStream của Bedrock để stream token, chuyển tiếp qua API Gateway REST API response streaming, và không buffer ở backend để giảm thời gian tới token đầu tiên.

Vì sao đúng

Đề nêu rõ chỉ số cần tối ưu: "time to first token" — người dùng cảm nhận độ nhanh dựa trên khi nào chữ đầu tiên xuất hiện, không phải khi nào câu trả lời xong.

Kiến trúc streaming đầy đủ có ba mắt xích, và cả ba đều phải stream — chỉ cần một chỗ buffer là mất hết lợi ích:

Bedrock (stream token) → Lambda (KHÔNG buffer) → API Gateway (response streaming) → Client
response = bedrock.invoke_model_with_response_stream(
    modelId='anthropic.claude-3-5-sonnet-20241022-v2:0',
    body=json.dumps({...})
)

for su_kien in response['body']:
    doan = json.loads(su_kien['chunk']['bytes'])
    yield doan['delta']['text']        # đẩy đi NGAY, không tích luỹ

Cụm "Avoid buffering responses in the backend" trong đáp án là điểm quan trọng nhất: nếu Lambda gom hết token rồi mới trả về, người dùng vẫn phải chờ toàn bộ câu trả lời — hệt như không stream.

Và vế "operationally efficient" được đáp ứng vì đây là HTTP thuần, không phải giao thức có trạng thái: | Đặc điểm | Chi tiết | |---|---| | Không quản lý kết nối | client gọi như một request HTTP thường | | Không lưu trạng thái | không có bảng connection nào | | Tích hợp frontend đơn giản | dùng fetch với ReadableStream |

Vì sao các phương án khác sai

  • A. Stream từ Bedrock rồi đẩy token qua API Gateway WebSocket API — hoạt động được nhưng phức tạp hơn hẳn: bạn phải quản lý trạng thái kết nối cho từng client (lưu connectionId vào DynamoDB), xử lý $connect/$disconnect, lo việc client mất mạng, và frontend phải cài đặt WebSocket. Đề nói rõ "without over engineering the frontend integration". (WebSocket đúng khi cần giao tiếp hai chiều liên tục — chat nhiều lượt với server chủ động đẩy tin. Ở đây chỉ cần một chiều.)
  • B. Stream từ Bedrock nhưng chuyển tiếp qua API Gateway HTTP API — HTTP API không hỗ trợ response streaming. Nó sẽ buffer toàn bộ phản hồi rồi mới trả về — nên dù Bedrock có stream, người dùng vẫn chờ hết. Đây là bẫy tinh vi nhất: mọi thứ đúng trừ một chi tiết về khả năng của dịch vụ.
  • D. Dùng InvokeModel (không stream) rồi tối ưu payload và dùng chỉ báo tải ở client — không stream chút nào: chờ toàn bộ câu trả lời rồi mới trả về. Vòng xoay tải chỉ che giấu độ trễ, không giảm nó.

Ghi nhớ

Hai API gọi model của Bedrock: | API | Trả về | |---|---| | InvokeModel / Converse | toàn bộ phản hồi một lần | | InvokeModelWithResponseStream / ConverseStream | từng đoạn token khi model sinh ra |

Khả năng streaming của các dịch vụ trước Lambda: | Dịch vụ | Response streaming | |---|---| | API Gateway REST API | ✅ | | API Gateway HTTP API | ❌ buffer | | Lambda Function URL | ✅ (RESPONSE_STREAM) | | ALB | ✅ | | API Gateway WebSocket | ✅ (nhưng là giao thức khác) |

Dòng Lambda Function URL đáng biết: với ứng dụng không cần các tính năng của API Gateway (authorizer, usage plan, WAF), nó là cách đơn giản và rẻ nhất để stream:

# Cấu hình Function URL với InvokeMode: RESPONSE_STREAM
import awslambdaric
def handler(event, responseStream, context):
    for doan in stream_tu_bedrock():
        responseStream.write(doan)
    responseStream.finish()

Ba chỉ số cần phân biệt khi tối ưu độ trễ GenAI: | Chỉ số | Ý nghĩa | |---|---| | TTFT (time to first token) | thời gian tới chữ đầu tiên — quyết định cảm nhận | | TPOT (time per output token) | tốc độ sinh chữ sau đó | | Tổng độ trễ | TTFT + TPOT × số token |

Streaming không giảm tổng độ trễ — nó giảm TTFT, và đó mới là thứ người dùng cảm nhận. Đề nói đúng điều này: "perceived responsiveness matters more than receiving the entire response at once".

Và ba cách khác để giảm TTFT thật sự: | Cách | Hiệu quả | |---|---| | Prompt caching | cache phần prefix cố định — giảm cả TTFT lẫn chi phí | | Model nhỏ hơn | TTFT thấp hơn đáng kể | | Rút gọn system prompt và ngữ cảnh | ít token đầu vào ⇒ xử lý nhanh hơn |