Ngân hàng đề — Google Cloud Professional Cloud DevOps Engineer
Tìm thấy 269 câu.
- A Enable Container Analysis in Artifact Registry, and check for common vulnerabilities and exposures (CVEs) in your container images
- B Use Binary Authorization to attest images during your CI/CD pipeline
- C Configure Identity and Access Management (IAM) policies to create a least privilege model on your GKE clusters.
- D Deploy Falco or Twistlock on GKE to monitor for vulnerabilities on your running Pods
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào chiến lược shift left on security (dịch chuyển bảo mật sang trái trong quy trình phát triển) trên Google Kubernetes Engine (GKE). Mục tiêu của đội InfoSec là triển khai guard rails (các rào chắn bảo mật) trên tất cả các GKE cluster, nhằm chỉ cho phép triển khai (deploy) các container images đã được tin cậy và phê duyệt.
- Shift left on security nghĩa là tích hợp kiểm tra bảo mật sớm nhất có thể trong CI/CD pipeline, trước khi code hoặc images được deploy vào production, thay vì chỉ phát hiện vấn đề sau khi chạy.
- Vấn đề cụ thể: Đảm bảo chỉ images trusted (đã được ký số, attest hoặc phê duyệt) mới được phép chạy trên GKE, ngăn chặn deploy images độc hại hoặc chưa kiểm tra từ đầu quy trình.
- Yêu cầu: Tìm giải pháp ngăn chặn (enforce) deploy images không approved, phù hợp với Google Cloud best practices cho GKE security (cập nhật đến 2026, dựa trên GKE Enterprise và Binary Authorization v2).
📘 Tài liệu tham khảo:
- Binary Authorization for GKE (Google Cloud Docs, cập nhật 2025).
- Shift Left Security on GKE (Google Cloud Architecture Center).
✅ Đáp án đúng: Use Binary Authorization to attest images during your CI/CD pipeline
Lý do lựa chọn:
- Binary Authorization là tính năng native của GKE (từ Google Cloud), cho phép enforce policy để chỉ deploy images đã được attest (xác thực) bởi các attestor đáng tin cậy (như Cosign hoặc private PKI).
- Nó tích hợp trực tiếp vào CI/CD pipeline (ví dụ: Cloud Build, GitHub Actions), kiểm tra signature/attestation trước khi admission controller của GKE cho phép pod chạy.
- Điều này chính xác là shift left: Ngăn chặn deploy từ nguồn (source), không chỉ scan mà còn block images không approved. Hỗ trợ multi-stage attestation và allowlist/denylist policies trong phiên bản mới nhất (2025-2026).
- Phù hợp 100% với yêu cầu "only allow deployment of trusted and approved images".
🛠️ Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Tôi đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm giải thích bằng tiếng Việt rõ ràng:
-
Enable Container Analysis in Artifact Registry, and check for common vulnerabilities and exposures (CVEs) in your container images
❌ Sai: Container Analysis (trong Artifact Registry) chỉ scan và báo cáo CVEs (lỗ hổng bảo mật) sau khi images được push lên registry. Nó không enforce/block deploy images chưa approved, mà chỉ cung cấp thông báo (alerts/notifications). Không phải shift left thực sự vì không ngăn chặn từ CI/CD pipeline.
📘 Tham khảo: Container Analysis Overview. -
Use Binary Authorization to attest images during your CI/CD pipeline
✅ Đúng: Như đã giải thích ở trên. Đây là giải pháp chính xác và được khuyến nghị cho GKE để enforce trusted images qua attestation trong pipeline. Hỗ trợ PKI-based signing và tích hợp seamless với GKE Autopilot/Standard clusters (cập nhật 2026).
📘 Tham khảo: Configuring Binary Authorization. -
Configure Identity and Access Management (IAM) policies to create a least privilege model on your GKE clusters.
❌ Sai: IAM chỉ kiểm soát quyền truy cập người dùng/dịch vụ (RBAC-like), không liên quan đến verify nội dung images. Nó bảo vệ cluster access nhưng không ngăn deploy images độc hại từ users đã có quyền. Không phải giải pháp cho "trusted images".
📘 Tham khảo: GKE IAM Best Practices. -
Deploy Falco or Twistlock on GKE to monitor for vulnerabilities on your running Pods
❌ Sai: Falco (open-source) hoặc Twistlock (Palo Alto Prisma Cloud) là công cụ runtime security (giám sát Pods đang chạy), phát hiện anomaly sau khi deploy. Đây là shift right (bảo mật sau deploy), không ngăn chặn từ đầu pipeline như yêu cầu "only allow deployment".
📘 Tham khảo: GKE Runtime Security.
Kết luận 💡: Binary Authorization là lựa chọn tối ưu cho GKE, giúp InfoSec team đạt guard rails mạnh mẽ mà không làm chậm CI/CD. Nếu triển khai, kết hợp với Artifact Registry + Container Analysis để scan trước attest! 🚀
- A Configure Binary Authorization in your GKE clusters to enforce deploy-time security policies.
- B Grant the roles/artifactregistry.writer role to the Cloud Build service account. Confirm that no employee has Artifact Registry write permission.
- C Use Cloud Run to write and deploy a custom validator. Enable an Eventarc trigger to perform validations when new images are uploaded.
- D Configure Kritis to run in your GKE clusters to enforce deploy-time security policies.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào một công ty hoạt động trong lĩnh vực được quy định nghiêm ngặt (highly regulated domain), nơi đội ngũ bảo mật yêu cầu chỉ các container images đáng tin cậy (trusted container images) mới được triển khai lên Google Kubernetes Engine (GKE).
📌 Yêu cầu chính: Triển khai giải pháp đáp ứng nhu cầu bảo mật này, đồng thời giảm thiểu overhead quản lý (minimizing management overhead).
🛠️ Bối cảnh: Đây là vấn đề bảo mật deploy-time trong GKE, cần cơ chế kiểm tra và xác thực images trước khi deploy, sử dụng các tính năng native của Google Cloud để tránh tự xây dựng phức tạp. Giải pháp phải tích hợp sẵn, dễ quản lý và tuân thủ các tiêu chuẩn bảo mật cao (như kiểm tra chữ ký số).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Configure Binary Authorization in your GKE clusters to enforce deploy-time security policies.
Lý do:
Binary Authorization là tính năng tích hợp sẵn của GKE (từ phiên bản GKE 1.15+ và cập nhật liên tục đến 2026), cho phép thực thi chính sách bảo mật tại thời điểm deploy (deploy-time) bằng cách kiểm tra chữ ký số (attestation) của container images. Nó đảm bảo chỉ images từ nguồn đáng tin cậy (ví dụ: ký bởi PKI hoặc CI/CD pipeline như Cloud Build) mới được deploy.
🧩 Ưu điểm giảm overhead: Không cần code custom, quản lý qua IAM và policy YAML; tự động tích hợp với GKE clusters. Hoàn hảo cho môi trường regulated.
📘 Nguồn tham khảo: Google Cloud Binary Authorization Docs (cập nhật 2024-2026).
📋 Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên nội dung gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể:
-
Configure Binary Authorization in your GKE clusters to enforce deploy-time security policies.
✅ Đúng: Như đã giải thích ở trên, đây là giải pháp native, chính xác cho deploy-time policy enforcement trên GKE, với overhead thấp nhất (chỉ config policy và attestors). Hỗ trợ đầy đủ cho trusted images qua PKI signatures. -
Grant the roles/artifactregistry.writer role to the Cloud Build service account. Confirm that no employee has Artifact Registry write permission.
❌ Sai: Phương án này chỉ kiểm soát quyền ghi (write permission) vào Artifact Registry (nơi lưu images), không kiểm tra nội dung images tại thời điểm deploy vào GKE. Nó ngăn chặn upload images không được phép nhưng không enforce "trusted images" (không verify chữ ký hoặc source). Overhead thấp nhưng không đáp ứng yêu cầu deploy-time security. -
Use Cloud Run to write and deploy a custom validator. Enable an Eventarc trigger to perform validations when new images are uploaded.
❌ Sai: Đây là giải pháp tự xây dựng (custom) sử dụng Cloud Run + Eventarc (event-driven trigger khi upload images vào registry). Nó kiểm tra tại upload-time, không phải deploy-time vào GKE, và tạo overhead cao (phát triển, maintain code, scaling Cloud Run). Không native, khó scale cho regulated domain. -
Configure Kritis to run in your GKE clusters to enforce deploy-time security policies.
❌ Sai: Kritis là công cụ open-source (nay tích hợp vào Binary Authorization như một backend), nhưng không phải giải pháp chính thức độc lập của GCP. Cấu hình Kritis yêu cầu cài đặt sidecar và custom CRDs, dẫn đến overhead quản lý cao hơn Binary Authorization (phải handle updates, compatibility). Binary Authorization mới là cách recommend từ Google.
🛡️ Kết luận: Binary Authorization là lựa pháp tối ưu, tuân thủ best practices GCP 2026 cho GKE security. Nếu cần demo, có thể thử trên GKE Standard/ Autopilot clusters!
- A Ensure that all postmortems include what caused the incident, identify the person or team responsible for causing the incident, and how to prevent a future occurrence of the incident.
- B Ensure that all postmortems include what caused the incident, how the incident could have been worse, and how to prevent a future occurrence of the incident.
- C Ensure that all postmortems include the severity of the incident, how to prevent a future occurrence of the incident, and what caused the incident without naming internal system components.
- D Ensure that all postmortems include how the incident was resolved and what caused the incident without naming customer information.
- E Ensure that all postmortems include all incident participants in postmortem authoring and share postmortems as widely as possible.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi này thuộc chủ đề postmortem policy (báo cáo hậu sự cố) trong quy trình DevOps và SRE (Site Reliability Engineering), liên quan đến việc triển khai chính sách postmortem cho mọi sự cố (incident) nhằm sử dụng nội bộ tại công ty. CTO yêu cầu bạn định nghĩa một postmortem tốt để đảm bảo chính sách thành công. Câu hỏi yêu cầu chọn hai lựa chọn đúng (Choose two) từ các phương án mô tả nội dung cần có trong postmortem.
Mục tiêu chính: Một postmortem tốt phải không đổ lỗi cá nhân (blameless), tập trung vào nguyên nhân gốc rễ (root cause), tác động (impact/severity), biện pháp khắc phục và phòng ngừa, đồng thời chia sẻ rộng rãi nội bộ để học hỏi tập thể. Đây là nguyên tắc cốt lõi từ Google SRE Book (áp dụng rộng rãi trong AWS qua Well-Architected Framework và AWS Incident Management best practices, cập nhật đến 2026 với các hướng dẫn mới về AI-assisted postmortems trong AWS X-Ray và Amazon DevOps Guru). Postmortem giúp cải thiện hệ thống mà không làm nản lòng đội ngũ.
📘 Tài liệu tham khảo:
- Google SRE Workbook: "Implementing Blameless Postmortems" (chapter 9, edition 2022+).
- AWS Well-Architected Reliability Pillar (2024-2026 updates): "Incident Response and Postmortems".
- AWS Documentation: "Blameless Postmortems" trong AWS Fault Injection Simulator và Amazon CloudWatch.
✅ Đáp án đúng (Chọn hai)
Hai phương án đúng là:
- Ensure that all postmortems include the severity of the incident, how to prevent a future occurrence of the incident, and what caused the incident without naming internal system components.
- Ensure that all postmortems include all incident participants in postmortem authoring and share postmortems as widely as possible.
Lý do lựa chọn 🛠️:
- Những yếu tố này tuân thủ nguyên tắc blameless postmortem chuẩn SRE/AWS: Bao gồm severity (mức độ nghiêm trọng) để đánh giá tác động, root cause (nguyên nhân) mà không tiết lộ chi tiết hệ thống nội bộ (tránh rủi ro bảo mật khi chia sẻ), hành động phòng ngừa (prevention actions) để cải thiện lâu dài. Đồng thời, tham gia toàn bộ người liên quan và chia sẻ rộng rãi nội bộ thúc đẩy văn hóa học hỏi tập thể, tránh lặp lại sự cố. Đây là best practices được AWS khuyến nghị trong Reliability Pillar (2026 updates nhấn mạnh chia sẻ qua AWS Management Console dashboards).
🔍 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng phương án một cách đầy đủ. Tôi giữ nguyên văn bản gốc tiếng Anh của phương án, đánh dấu ✅ (đúng) hoặc ❌ (sai), và giải thích rõ ràng bằng tiếng Việt:
-
Ensure that all postmortems include what caused the incident, identify the person or team responsible for causing the incident, and how to prevent a future occurrence of the incident.
❌ Sai: Phương án này vi phạm nguyên tắc blameless postmortem cốt lõi (không đổ lỗi cá nhân hoặc đội ngũ). Việc "identify the person or team responsible" tạo văn hóa sợ hãi, làm nhân viên ngại báo cáo sự cố, dẫn đến chính sách thất bại. AWS và SRE nhấn mạnh tập trung vào hệ thống, không cá nhân (theo Google SRE Book). -
Ensure that all postmortems include what caused the incident, how the incident could have been worse, and how to prevent a future occurrence of the incident.
❌ Sai: "How the incident could have been worse" (có thể tệ hơn) không phải yếu tố chuẩn trong postmortem. Nó mang tính suy đoán, không giúp phân tích root cause hoặc prevention hiệu quả. Best practices AWS chỉ yêu cầu timeline, impact thực tế, resolution, và actions – không cần "worst-case" trừ khi là risk assessment riêng (không bắt buộc cho mọi postmortem). -
Ensure that all postmortems include the severity of the incident, how to prevent a future occurrence of the incident, and what caused the incident without naming internal system components.
✅ Đúng: Hoàn hảo với SRE principles. Severity đánh giá tác động (SLO/SLI breach), prevention là mục tiêu chính, root cause mà không nêu tên internal components bảo vệ bí mật hệ thống nội bộ khi chia sẻ rộng (AWS khuyến nghị trong internal postmortems để tránh leak qua tools như CloudWatch Logs). -
Ensure that all postmortems include how the incident was resolved and what caused the incident without naming customer information.
❌ Sai: Tuy "how resolved" (cách giải quyết) là tốt và "without naming customer info" đúng về bảo mật (PII compliance theo AWS GDPR/CPRA 2026), nhưng phương án thiếu severity, prevention, và involvement/sharing. Nó không đầy đủ để định nghĩa "good postmortem" – resolution chỉ là phần một, không phải toàn bộ (AWS yêu cầu full timeline + actions). -
Ensure that all postmortems include all incident participants in postmortem authoring and share postmortems as widely as possible.
✅ Đúng: Khuyến khích collaborative authoring (toàn bộ participants tham gia viết) để có góc nhìn đa chiều, tránh bias. Share widely internally lan tỏa lessons learned, cải thiện toàn công ty (AWS DevOps best practices 2026: Sử dụng AWS Knowledge Base hoặc Confluence integration cho sharing).
Kết luận 🎯: Chọn hai ✅ giúp xây dựng chính sách postmortem thành công, thúc đẩy văn hóa DevOps mạnh mẽ trên AWS! Nếu cần ví dụ thực tế hoặc template postmortem, hãy hỏi thêm. 🚀
- A Use a Jenkins server for CI/CD pipelines. Periodically run all tests in the feature branch.
- B Ask the pull request reviewers to run the integration tests before approving the code.
- C Use Cloud Build to run the tests. Trigger all tests to run after a pull request is merged.
- D Use Cloud Build to run tests in a specific folder. Trigger Cloud Build for every GitHub pull request.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi này thuộc lĩnh vực DevOps và CI/CD trên Google Cloud Platform (GCP), tập trung vào việc tự động hóa kiểm thử tích hợp (integration tests) cho các module Infrastructure as Code (IaC) tái sử dụng.
- Bối cảnh: Bạn đang phát triển các module IaC (ví dụ: Terraform hoặc tương tự), mỗi module có các bài kiểm thử tích hợp riêng biệt. Các kiểm thử này sẽ khởi chạy (launch) module trong một dự án kiểm thử (test project) để xác thực tính toàn vẹn.
- Yêu cầu chính: Sử dụng GitHub làm nguồn kiểm soát mã nguồn (source control). Cần kiểm thử liên tục (continuously test) trên feature branch, đảm bảo toàn bộ code được kiểm thử trước khi chấp nhận thay đổi (before changes are accepted), tức là trước khi merge vào main branch.
- Mục tiêu: Triển khai giải pháp tự động hóa (automate) các integration tests, tránh thủ công, để hỗ trợ quy trình CI/CD hiệu quả, phát hiện lỗi sớm trên từng pull request (PR).
🛠️ Yêu cầu cốt lõi: Giải pháp phải trigger tự động trên feature branch/PR, chạy tests trước khi merge, không phải sau hoặc thủ công. Đây là thực hành tốt nhất theo Google Cloud DevOps best practices (dựa trên phiên bản Cloud Build mới nhất 2026, hỗ trợ GitHub App integration và folder-specific triggers).
📘 Tài liệu tham khảo:
- Cloud Build Triggers Documentation (Google Cloud, cập nhật 2026).
- Cloud Build with GitHub – Hỗ trợ trigger trên PR từ feature branch.
- Google Cloud Architecture Framework: DevOps SRE practices.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use Cloud Build to run tests in a specific folder. Trigger Cloud Build for every GitHub pull request.
Lý do:
- 🧩 Phù hợp hoàn hảo: Cloud Build là dịch vụ CI/CD native của GCP, tích hợp sâu với GitHub (qua GitHub App hoặc webhook). Có thể cấu hình trigger trên mọi pull request (PR) từ feature branch, chạy tests trong thư mục cụ thể (specific folder) chứa module IaC và integration tests – tránh chạy toàn bộ repo, tiết kiệm chi phí và thời gian.
- 🚀 Tự động hóa liên tục: Đảm bảo tests chạy trước merge (block merge nếu fail), hỗ trợ "continuously test feature branch" và "tested before changes accepted".
- 📈 Tính năng mới nhất (2026): Cloud Build hỗ trợ folder-level triggers (chạy builds chỉ trên thay đổi trong folder cụ thể), parallel testing, và integration tests với ephemeral test projects (qua Service Account impersonation).
- Lợi ích: Scale tự động, tích hợp Artifact Registry cho IaC outputs, báo cáo kết quả trực tiếp trên GitHub PR (comments, status checks).
📝 Giải thích chi tiết tất cả các phương án
-
❌ [SAI] Use a Jenkins server for CI/CD pipelines. Periodically run all tests in the feature branch.
Lý do sai: Jenkins là công cụ CI/CD bên thứ ba (self-hosted hoặc EKS), không native với GCP, yêu cầu quản lý server phức tạp (không serverless như Cloud Build). "Periodically run" (chạy định kỳ) không đảm bảo liên tục trên feature branch hoặc trước merge – có thể bỏ lỡ thay đổi realtime. Không tự động hóa tối ưu, vi phạm yêu cầu "continuously test" và tăng overhead ops. -
❌ [SAI] Ask the pull request reviewers to run the integration tests before approving the code.
Lý do sai: Đây là quy trình thủ công hoàn toàn, phụ thuộc reviewer – không tự động hóa (automate) như yêu cầu. Dễ lỗi con người, chậm trễ, không "continuously test feature branch". Không scale cho team lớn, vi phạm nguyên tắc CI/CD (shift-left testing). -
❌ [SAI] Use Cloud Build to run the tests. Trigger all tests to run after a pull request is merged.
Lý do sai: Mặc dù dùng Cloud Build (tốt), nhưng trigger sau merge (after merged) nghĩa là tests chạy sau khi code đã accepted – trái ngược yêu cầu "ensure all code is tested before changes are accepted". Không phát hiện lỗi sớm trên feature branch, rủi ro cao (code lỗi đã vào main). -
✅ [ĐÚNG] Use Cloud Build to run tests in a specific folder. Trigger Cloud Build for every GitHub pull request.
Lý do đúng (tóm tắt): Như phân tích trên – trigger PR-specific, folder-targeted, tự động trước merge, phù hợp best practices GCP DevOps 2026. Đảm bảo zero-downtime testing với test projects isolated.
🛠️ Khuyến nghị triển khai: Tạo Cloud Build trigger YAML với github: push/pull_request event, location: dir/module-folder, steps chạy terraform apply/plan + tests. Kết hợp GitHub Actions checks để block merge nếu fail!
- A Use Cloud Monitoring to assess the App Engine CPU utilization metric.
- B Install a continuous profiling tool into Compute Engine. Configure the application to send profiling data to the tool.
- C Periodically run the go tool pprof command against the application instance. Analyze the results by using flame graphs.
- D Configure Cloud Profiler, and initialize the cloud.google.com/go/profiler library in the application.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một tình huống thực tế trong Google Cloud Platform (GCP):
Công ty của bạn xử lý dữ liệu IoT quy mô lớn bằng cách sử dụng Pub/Sub (dịch vụ messaging), App Engine standard environment (nền tảng PaaS để chạy ứng dụng), và một ứng dụng viết bằng ngôn ngữ Go.
📈 Vấn đề: Hiệu suất ứng dụng giảm không nhất quán (inconsistently degrades) tại đỉnh tải (peak load), nhưng không thể tái tạo (could not reproduce) trên máy trạm cá nhân (workstation).
🎯 Yêu cầu: Giám sát liên tục (continuously monitor) ứng dụng trong môi trường production để xác định các đường dẫn code chậm (slow paths in the code), đồng thời giảm thiểu tác động hiệu suất (minimize performance impact) và chi phí quản lý (management overhead).
🛠️ Mục tiêu chính là cần một giải pháp profiling (phân tích hiệu suất code) tự động, low-overhead, phù hợp với App Engine và Go, không yêu cầu can thiệp thủ công thường xuyên.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Configure Cloud Profiler, and initialize the cloud.google.com/go/profiler library in the application.
Lý do:
🟢 Cloud Profiler là dịch vụ GCP chuyên dụng cho continuous profiling (profiling liên tục) với overhead cực thấp (<1-2% CPU), tự động thu thập dữ liệu CPU, heap, wall-clock mà không làm ảnh hưởng hiệu suất production.
- Hỗ trợ App Engine standard và ngôn ngữ Go qua thư viện
cloud.google.com/go/profiler– chỉ cần init một lần trong code là chạy tự động. - Giúp xác định chính xác slow paths qua flame graphs trực quan trên console GCP.
- Không cần reproduce thủ công, giám sát real-time tại peak load, và zero management vì là managed service (cập nhật đến 2026, Cloud Profiler vẫn là best practice theo docs GCP).
📘 Nguồn tham khảo: - Cloud Profiler Overview
- Profiling Go Apps
- App Engine Integration
📋 Giải thích tất cả các phương án (đúng/sai)
-
SAI ❌ Use Cloud Monitoring to assess the App Engine CPU utilization metric.
Giải thích: Cloud Monitoring chỉ theo dõi metrics tổng quát như CPU utilization (sử dụng CPU), không phân tích sâu slow paths trong code. Nó giúp phát hiện vấn đề cao cấp (high-level) nhưng không đủ chi tiết để pinpoint bottlenecks ở level code, không đáp ứng yêu cầu "identify slow paths". Overhead thấp nhưng không giải quyết gốc rễ vấn đề. -
SAI ❌ Install a continuous profiling tool into Compute Engine. Configure the application to send profiling data to the tool.
Giải thích: Ứng dụng chạy trên App Engine standard (PaaS managed), không phải Compute Engine (IaaS VM). Việc install tool trên Compute Engine yêu cầu di chuyển app (không khả thi), tăng management overhead lớn (quản lý VM, scaling), và tác động performance cao hơn so với native GCP service. Không phù hợp môi trường serverless. -
SAI ❌ Periodically run the go tool pprof command against the application instance. Analyze the results by using flame graphs.
Giải thích:go tool pproflà công cụ manual/profiling thủ công của Go, chỉ chạy định kỳ (periodically), không continuous. Khó tái tạo tại peak load, yêu cầu truy cập instance (khó trên App Engine), tăng overhead (stop app để profile), và management cao (phân tích thủ công flame graphs). Không minimize impact như yêu cầu. -
ĐÚNG ✅ Configure Cloud Profiler, and initialize the cloud.google.com/go/profiler library in the application.
Giải thích: Như đã nêu ở phần đáp án đúng – giải pháp tối ưu nhất, low-overhead, continuous, native cho App Engine/Go, tự động visualize slow paths qua UI GCP. Hoàn hảo cho production monitoring! 🚀
- A Run the gcloud container clusters update --logging=SYSTEM command for the development cluster.
- B Run the gcloud container clusters update --logging=WORKLOAD command for the development cluster.
- C Run the gcloud logging sinks update _Default --disabled command in the project associated with the development environment.
- D Add the severity >= DEBUG resource.type = "k8s_container" exclusion filter to the _Default logging sink in the project associated with the development environment.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi
Câu hỏi gốc:
Your company runs services by using Google Kubernetes Engine (GKE). The GKE clusters in the development environment run applications with verbose logging enabled. Developers view logs by using the kubectl logs command and do not use Cloud Logging. Applications do not have a uniform logging structure defined. You need to minimize the costs associated with application logging while still collecting GKE operational logs. What should you do?
📝 Giải thích nội dung câu hỏi:
- Bối cảnh hệ thống: Công ty sử dụng Google Kubernetes Engine (GKE) để chạy các dịch vụ. Các cluster GKE trong môi trường development (dev) đang chạy ứng dụng với verbose logging (ghi log chi tiết, nhiều dữ liệu log, dẫn đến chi phí cao khi lưu trữ).
- Cách developer làm việc: Họ chỉ xem log qua lệnh
kubectl logs(lấy log trực tiếp từ pod/container trên node, không qua Cloud Logging), và không sử dụng Cloud Logging để xem. - Vấn đề với ứng dụng: Các ứng dụng không có cấu trúc log thống nhất (không uniform), nghĩa là log format khác nhau, khó filter chính xác.
- Yêu cầu chính: Giảm thiểu chi phí log ứng dụng (application logging - tức workload/container logs), nhưng vẫn thu thập đầy đủ GKE operational logs (log vận hành GKE như system logs từ node, kubelet, control plane, events).
- Mục tiêu: Tránh gửi log verbose của app lên Cloud Logging (chi phí lưu trữ cao), giữ nguyên khả năng xem log local qua kubectl, và đảm bảo operational logs vẫn được collect để monitor cluster.
- Kiến thức cập nhật (GKE 2024-2026): GKE sử dụng Cloud Logging agent (Fluent Bit) mặc định, gửi tất cả logs (workload, system) lên _Default sink.
--loggingflag là legacy (chỉ ảnh hưởng control plane hoặc cũ), không hiệu quả cho workload logs mới. Giải pháp chuẩn là exclusion filter trên sink để loại log cụ thể, giảm chi phí mà không tắt toàn bộ. (Nguồn: Cloud Logging for GKE, Excluding Logs, cập nhật Q1/2026).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Add the severity >= DEBUG resource.type = "k8s_container" exclusion filter to the _Default logging sink in the project associated with the development environment.
🛠️ Lý do chi tiết:
- Filter exclusion trên _Default sink (sink mặc định của project) sẽ loại bỏ log từ resource.type="k8s_container" (workload logs từ container/app) có severity >= DEBUG (bao gồm DEBUG, INFO, WARNING,... - chính là verbose logs).
- Giữ nguyên operational logs: System logs (k8s_node, k8s_cluster, k8s_control_plane) không bị ảnh hưởng vì resource type khác.
- Phù hợp dev env: Developer vẫn dùng
kubectl logs(local, miễn phí), app không uniform → filter dựa severity linh hoạt, giảm chi phí verbose mà không cần thay đổi app code. - Tối ưu chi phí: Chỉ loại log thừa, không tắt toàn bộ logging agent. Đây là best practice theo docs GKE mới nhất.
(Nguồn: GKE Logging Best Practices, AWS không liên quan - câu hỏi thuần GKE).
📋 Giải thích tất cả các phương án (Đúng/Sai)
-
Run the gcloud container clusters update --logging=SYSTEM command for the development cluster.
❌ Sai: Flag--logging=SYSTEMlà legacy config (dành cho Fluentd cũ, deprecated từ 2022+), chỉ kiểm soát control plane logs, không loại hết workload/container logs (vẫn gửi verbose app logs qua Fluent Bit agent mới). Không giảm chi phí app logging hiệu quả, và có thể gây confuse với config hiện đại. (Nguồn: Legacy Logging). -
Run the gcloud container clusters update --logging=WORKLOAD command for the development cluster.
❌ Sai: Chỉ enable WORKLOAD logs (app/container), loại bỏ SYSTEM/operational logs (node, kubelet,...), vi phạm yêu cầu "still collecting GKE operational logs". Hơn nữa, vẫn là legacy flag, không giải quyết verbose costs. -
Run the gcloud logging sinks update _Default --disabled command in the project associated with the development environment.
❌ Sai: Disable toàn bộ _Default sink sẽ loại tất cả logs (kể cả operational GKE logs), làm mất khả năng monitor cluster. Không phân biệt app vs operational, gây rủi ro cao. -
Add the severity >= DEBUG resource.type = "k8s_container" exclusion filter to the _Default logging sink in the project associated with the development environment.
✅ Đúng: Như giải thích trên, exclusion filter tinh chỉnh chỉ loại verbose app logs (k8s_container >=DEBUG), giữ operational logs, giảm chi phí tối đa mà an toàn. Áp dụng ngay cho project dev env.
🎯 Kết luận: Giải pháp sử dụng exclusion filter là best practice cho GKE hiện đại (2026), linh hoạt và tiết kiệm. Nếu implement: gcloud logging sinks update _Default sink --exclusion=...! 🛡️
- A Grant the logging.logWriter and monitoring.metricWriter roles to the Compute Engine service accounts.
- B Grant the logging.admin and monitoring.editor roles to the Compute Engine service accounts.
- C Grant the logging.editor and monitoring.metricWriter roles to the Compute Engine service accounts.
- D Grant the logging.logWriter and monitoring.editor roles to the Compute Engine service accounts.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi
📘 Nội dung câu hỏi:
Câu hỏi yêu cầu triển khai một nhóm (fleet) các máy ảo Compute Engine trên Google Cloud Platform (GCP). Mục tiêu là đảm bảo rằng các chỉ số giám sát (monitoring metrics) và nhật ký (logs) từ các instance này được hiển thị trong Cloud Logging và Cloud Monitoring, để các đội ngũ vận hành (operations) và an ninh mạng (cyber security) của công ty có thể theo dõi.
🛠️ Yêu cầu cụ thể:
- Sử dụng Identity and Access Management (IAM) để cấp quyền (roles) cho Compute Engine service account (tài khoản dịch vụ mặc định của Compute Engine, thường là
<project-number>-compute@developer.gserviceaccount.com). - Tuân thủ nguyên tắc least privilege (quyền hạn tối thiểu): Chỉ cấp đúng những quyền cần thiết để service account có thể ghi (write) logs và metrics vào các dịch vụ Logging/Monitoring, mà không cấp quyền đọc, chỉnh sửa hoặc quản trị thừa.
✅ Lý do ngữ cảnh quan trọng: Compute Engine instances tự động thu thập metrics/logs cơ bản (như CPU, disk usage) và gửi chúng qua service account. Để dữ liệu này hiển thị trong Cloud Logging/Monitoring, service account cần quyền writer cụ thể, không phải quyền rộng hơn (như editor/admin) để tránh rủi ro bảo mật.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Grant the logging.logWriter and monitoring.metricWriter roles to the Compute Engine service accounts.
🧩 Giải thích lý do:
- logging.logWriter: Cho phép service account ghi logs vào Cloud Logging mà không có quyền đọc, xóa hoặc chỉnh sửa logs của người khác. Đây là quyền tối thiểu cần thiết để logs từ Compute Engine được đẩy lên.
- monitoring.metricWriter: Cho phép ghi custom metrics (bao gồm metrics hệ thống từ Compute Engine) vào Cloud Monitoring, mà không cho phép đọc hoặc chỉnh sửa metrics khác.
- Tuân thủ least privilege: Hai roles này chỉ cấp quyền write-only, phù hợp hoàn hảo với nhu cầu "visible in Cloud Logging and Cloud Monitoring" mà không trao quyền thừa cho ops/cyber sec teams (họ chỉ cần xem dữ liệu, không cần service account quản lý). Theo tài liệu GCP mới nhất (2024-2026), đây là khuyến nghị chính thức cho Compute Engine.
📋 Giải thích tất cả các phương án (đúng/sai)
-
✅ [ĐÚNG] Grant the logging.logWriter and monitoring.metricWriter roles to the Compute Engine service accounts.
🟢 Đúng vì: Như giải thích trên, đây là bộ đôi roles tối thiểu, write-only, đảm bảo logs/metrics được ghi thành công mà không vi phạm least privilege. Compute Engine yêu cầu chính xác những quyền này để agent (như Ops Agent hoặc legacy Stackdriver agent) hoạt động. -
❌ [SAI] Grant the logging.admin and monitoring.editor roles to the Compute Engine service accounts.
🔴 Sai vì:logging.admincấp quyền quản trị đầy đủ (đọc, viết, xóa, chỉnh sửa tất cả logs, bao gồm logs của project/folder/organization).monitoring.editorcho phép chỉnh sửa metrics (bao gồm đọc và thay đổi dashboards/alerts). Những quyền này quá rộng, vi phạm least privilege và tạo rủi ro bảo mật cao (service account có thể xóa logs quan trọng). -
❌ [SAI] Grant the logging.editor and monitoring.metricWriter roles to the Compute Engine service accounts.
🔴 Sai vì:logging.editorcấp quyền chỉnh sửa logs (đọc, viết, cập nhật entries), thừa so với nhu cầu chỉ write logs. Kết hợp vớimonitoring.metricWriter(đúng), nhưng tổng thể vẫn không least privilege vìlogging.editorcho phép đọc/chỉnh sửa logs không cần thiết. -
❌ [SAI] Grant the logging.logWriter and monitoring.editor roles to the Compute Engine service accounts.
🔴 Sai vì:logging.logWriter(đúng, write-only logs), nhưngmonitoring.editorcấp quyền chỉnh sửa đầy đủ metrics (đọc, tạo/updates dashboards, alerts). Thừa quyền so với chỉ cần write metrics, vi phạm nguyên tắc least privilege.
📚 Tài liệu tham khảo (cập nhật mới nhất GCP đến 2026)
- Cloud Logging IAM roles: cloud.google.com/logging/docs/access-control – Xác nhận
logging.logWritercho write-only. - Cloud Monitoring IAM roles: cloud.google.com/monitoring/access-control – Khuyến nghị
monitoring.metricWritercho Compute Engine metrics. - Compute Engine monitoring setup: cloud.google.com/compute/docs/monitoring – Hướng dẫn cấp roles cho service account.
- IAM best practices (least privilege): cloud.google.com/iam/docs/best-practices-service-accounts.
🛠️ Lời khuyên DevOps: Để triển khai, dùng gcloud CLI: gcloud projects add-iam-policy-binding [PROJECT_ID] --member="serviceAccount:[PROJECT_NUMBER]-compute@developer.gserviceaccount.com" --role="roles/logging.logWriter" (tương tự cho metricWriter). Kiểm tra bằng IAM Policy Analyzer! 🚀
- A Deploy the prototype product in a test environment, run a load test, and share the results with the product development team.
- B When the initial product version passes the quality assurance phase and compliance assessments, deploy the product to a staging environment. Share error logs and performance metrics with the product development team.
- C When the new product is used by at least one internal customer in production, share error logs and monitoring metrics with the product development team.
- D Review the design of the product with the product development team to provide feedback early in the design phase.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả vai trò của bạn là Site Reliability Engineer (SRE) chịu trách nhiệm quản lý các dịch vụ dữ liệu và sản phẩm dữ liệu của công ty. Bạn thường gặp phải các thách thức vận hành như khối lượng dữ liệu không dự đoán được và chi phí cao trong quy trình thu thập dữ liệu (data ingestion). Gần đây, bạn biết về một sản phẩm thu thập dữ liệu mới sẽ được phát triển trên Google Cloud. Nhiệm vụ của bạn là hợp tác với đội ngũ phát triển sản phẩm để cung cấp ý kiến vận hành (operational input) cho sản phẩm mới này.
Mục tiêu chính: Tìm cách hợp tác hiệu quả nhất để đảm bảo sản phẩm mới giải quyết được các vấn đề vận hành từ sớm, phù hợp với nguyên tắc SRE (Site Reliability Engineering) của Google Cloud – nhấn mạnh vào việc chuyển trách nhiệm vận hành sang giai đoạn thiết kế sớm (shift left), thay vì chờ đến giai đoạn triển khai hoặc sản xuất. Điều này giúp giảm rủi ro, tối ưu chi phí và đảm bảo độ tin cậy cao theo các thực hành DevOps hiện đại (cập nhật đến 2026, dựa trên Google Cloud SRE Workbook và DevOps best practices).
🛠️ Nguyên tắc cốt lõi: SRE khuyến khích phản hồi sớm trong giai đoạn thiết kế để tránh các vấn đề vận hành phát sinh muộn, thay vì chỉ phản ứng sau khi sản phẩm đã deploy.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Review the design of the product with the product development team to provide feedback early in the design phase.
Lý do: ✅ Phương án này phù hợp nhất với nguyên tắc SRE shift-left trên Google Cloud, nơi SRE tham gia từ giai đoạn thiết kế ban đầu để cung cấp input về vận hành (như scalability, cost optimization cho data ingestion). Việc phản hồi sớm giúp tránh các vấn đề như unpredictable data volume hoặc high cost ngay từ đầu, giảm rework và tăng hiệu quả. Đây là best practice được khuyến nghị trong Google Cloud Professional Cloud DevOps Engineer certification (cập nhật 2026) và SRE principles.
📘 Tài liệu tham khảo:
- Google SRE Workbook (Chapter 2: Implementing SRE): https://sre.google/workbook/
- Google Cloud DevOps Best Practices: https://cloud.google.com/architecture/devops
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng phương án một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên SRE best practices trên Google Cloud (không phải AWS, dù câu hỏi nhấn mạnh Google Cloud data ingestion).
-
Phương án 1: Deploy the prototype product in a test environment, run a load test, and share the results with the product development team.
❌ Sai: Phương án này chờ đến giai đoạn prototype và load test (giai đoạn sau thiết kế), dẫn đến phản hồi muộn. SRE cần input sớm hơn để ảnh hưởng đến thiết kế cốt lõi, tránh lãng phí thời gian test các vấn đề có thể dự đoán (như data volume spikes). Không hiệu quả cho collaboration sớm. -
Phương án 2: When the initial product version passes the quality assurance phase and compliance assessments, deploy the product to a staging environment. Share error logs and performance metrics with the product development team.
❌ Sai: Phương án này chờ QA và compliance hoàn tất mới deploy staging và chia sẻ metrics/logs – quá muộn! SRE phải tham gia trước QA để thiết kế sản phẩm chống chịu lỗi từ đầu. Việc chỉ share logs sau staging không giải quyết root cause vận hành như high cost. -
Phương án 3: When the new product is used by at least one internal customer in production, share error logs and monitoring metrics with the product development team.
❌ Sai: Đây là cách phản ứng thụ động nhất, chờ sản phẩm vào production với khách nội bộ mới share logs/metrics. Vi phạm nguyên tắc SRE (không để production chịu rủi ro đầu tiên), có thể gây downtime hoặc chi phí cao thực tế. SRE ưu tiên proactive design review thay vì reactive monitoring. -
Phương án 4 (Đúng): Review the design of the product with the product development team to provide feedback early in the design phase.
✅ Đúng: Như đã giải thích ở trên, đây là cách hợp tác tối ưu, áp dụng early feedback loop trong SRE lifecycle trên Google Cloud. Giúp tích hợp operational requirements (scalability, cost-efficiency cho data ingestion) ngay từ thiết kế, giảm rủi ro lâu dài.
🧠 Kết luận: Chọn phương án đúng giúp bạn thể hiện vai trò SRE chuyên nghiệp, thúc đẩy DevOps culture trên Google Cloud! Nếu cần thêm ví dụ thực tế (như sử dụng Dataflow cho ingestion), hãy hỏi nhé. 🚀
- A Create a new tag called stable that points to the previously working container, and change the deployment to point to the new tag.
- B Alter the deployment to point to the sha256 digest of the previously working container.
- C Build a new container from a previous Git tag, and do a rolling update on the deployment to the new container.
- D Apply the latest tag to the previous container image, and do a rolling update on the deployment.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh tình huống khắc phục sự cố production trên Google Kubernetes Engine (GKE) – nền tảng quản lý Kubernetes của Google Cloud. Ứng dụng đang gặp vấn đề do container image mới cập nhật gây ra (chưa xác định chính xác thay đổi code), và deployment hiện tại đang trỏ đến tag "latest" (thẻ động, dễ thay đổi). Nhiệm vụ là cập nhật cluster để chạy phiên bản container cũ hoạt động ổn định, đảm bảo tính immutable (không thay đổi) và nhanh chóng, an toàn mà không cần rebuild hay tạo tag mới.
📌 Bối cảnh quan trọng:
- Tag "latest" là mutable (có thể bị ghi đè bất cứ lúc nào), dẫn đến rủi ro rollback không đáng tin cậy.
- Giải pháp lý tưởng phải pin chính xác đến image cụ thể (dùng digest hash), tuân thủ best practices Kubernetes/GKE đến năm 2026 (Kubernetes v1.30+ và GKE 1.30+ khuyến nghị dùng digest cho production để tránh "image drift").
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Alter the deployment to point to the sha256 digest of the previously working container.
Lý do 🛠️:
- Sha256 digest là immutable reference (hash duy nhất của image layer), đảm bảo deployment luôn trỏ đúng image cũ hoạt động tốt, không bị ảnh hưởng bởi tag mutable như "latest".
- Thao tác nhanh: Chỉ edit YAML deployment (thay
image: repo:latestthànhimage: repo@sha256:abc123...), applykubectl apply, Kubernetes tự rolling update. - Tuân thủ best practices GKE/Kubernetes 2026: Tránh tag mutable ở production (theo CIS Kubernetes Benchmark 1.9.0), giảm downtime và tăng tính traceable. Không cần rebuild hay tạo tag mới, tiết kiệm thời gian.
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] Create a new tag called stable that points to the previously working container, and change the deployment to point to the new tag.
Phương án này không an toàn vì tag "stable" vẫn là mutable tag (ai đó có thể push image mới với cùng tag, gây "image drift"). Không immutable như digest, vi phạm nguyên tắc production stability trên GKE (Kubernetes docs khuyến cáo tránh tag-based rollback). -
✅ [ĐÚNG] Alter the deployment to point to the sha256 digest of the previously working container.
Như đã giải thích ở trên: Immutable, chính xác, nhanh chóng. Lấy digest bằngdocker inspecthoặcgcr.io crane digest(Artifact Registry), edit deployment YAML và apply. Hoàn hảo cho hotfix production. -
❌ [SAI] Build a new container from a previous Git tag, and do a rolling update on the deployment to the new container.
Phức tạp và không cần thiết: Rebuild image từ Git cũ tốn thời gian (CI/CD pipeline), có nguy cơ khác biệt môi trường (cache, deps), và tạo image mới vẫn cần tag/digest mới. Không giải quyết nhanh issue "latest tag broken". -
❌ [SAI] Apply the latest tag to the previous container image, and do a rolling update on the deployment.
Nguy hiểm cao: Ghi đè tag "latest" lên image cũ có thể confuse dev khác (họ push mới lại), tag vẫn mutable. Không khuyến khích, dễ gây hỗn loạn registry và vi phạm immutability principle.
📘 Tài liệu tham khảo (cập nhật đến 2026)
- Kubernetes Docs: Container Images – Khuyến nghị dùng digest cho production (v1.30+).
- GKE Docs: Debugging GKE Deployments & Immutable Tags Best Practices (Artifact Registry 2026).
- CIS Benchmark: Phần 5.2.5 – Avoid latest tag in prod.
- Công cụ hỗ trợ:
crane digest(Google's sigstore/cosign tool) để lấy sha256 nhanh.
🛡️ Lời khuyên DevOps: Luôn dùng semantic versioning + digest cho GKE prod (ví dụ: image: gcr.io/project/app:v1.2.3@sha256:...), kết hợp GitOps (ArgoCD/Flux) để tránh tương lai!
- A Select a latency metric for a request-based method of evaluation.
- B Select a latency metric for a window-based method of evaluation.
- C Select an availability metric for a request-based method of evaluation.
- D Select an availability metric for a window-based method of evaluation.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc tạo một Service Level Objective (SLO) trong Google Cloud Monitoring (không phải AWS, mặc dù đề cập chung chung về cloud) cho một dịch vụ sắp được công bố. Mục tiêu cụ thể là xác minh rằng các yêu cầu (requests) đến dịch vụ được xử lý trong thời gian dưới 300 ms ít nhất 90% thời gian trong mỗi tháng lịch.
- Yêu cầu chính: Xác định metric phù hợp (chỉ số đo lường) và phương pháp đánh giá (evaluation method) để thiết lập SLO này.
- Bối cảnh: SLO là một phần của Service Level Indicator (SLI), giúp đo lường hiệu suất dịch vụ. Trong Google Cloud Monitoring (cập nhật đến năm 2026, theo phiên bản Cloud Monitoring v2 và Error Reporting mới nhất), có hai phương pháp đánh giá SLO chính:
- Request-based: Dựa trên từng request riêng lẻ, lý tưởng cho các metric như latency (thời gian phản hồi), tính toán tỷ lệ phần trăm requests đạt ngưỡng (ví dụ: 90% requests < 300ms).
- Window-based: Dựa trên tổng số "good" và "bad" events trong một cửa sổ thời gian cố định (ví dụ: 5 phút), phù hợp hơn cho availability (tính sẵn sàng).
- Mục tiêu SLO: "Fewer than 300 ms at least 90% of the time per calendar month" rõ ràng là mục tiêu latency percentile, không phải availability (uptime).
📘 Tài liệu tham khảo:
- Google Cloud Monitoring SLO Documentation (cập nhật 2025-2026).
- SLO Evaluation Methods – Xác nhận request-based cho latency SLI.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Select a latency metric for a request-based method of evaluation.
🛠️ Lý do chi tiết:
- Latency metric (chỉ số độ trễ) là lựa chọn phù hợp vì mục tiêu là đo thời gian xử lý requests (< 300ms).
- Request-based evaluation tính toán trực tiếp tỷ lệ requests đạt ngưỡng trong khoảng thời gian (ở đây là 90% mỗi tháng), khớp chính xác với yêu cầu "90% of the time per calendar month".
- Phương pháp này sử dụng percentile aggregation (ví dụ: 90th percentile latency ≤ 300ms), được khuyến nghị cho các SLO latency trong Cloud Monitoring.
❌ Phân tích tất cả các phương án (đúng/sai)
-
[ĐÚNG] Select a latency metric for a request-based method of evaluation.
✅ Đúng vì: Kết hợp hoàn hảo metric latency (đo độ trễ requests) với request-based (tính % requests tốt theo từng request riêng lẻ). Đây là cách tiêu chuẩn để đạt SLO 90% requests <300ms/tháng, tránh sai lệch từ window-based. 🏆 -
[SAI] Select a latency metric for a window-based method of evaluation.
❌ Sai vì: Window-based đếm "good/bad" events trong cửa sổ thời gian (ví dụ: 5 phút), không chính xác cho latency percentile. Nó có thể làm méo mó dữ liệu nếu requests không đều, không khớp yêu cầu "per request" <300ms 90% thời gian. 🕒 -
[SAI] Select an availability metric for a request-based method of evaluation.
❌ Sai vì: Availability metric đo uptime (tỷ lệ dịch vụ sẵn sàng, thường 99.9%), không liên quan đến thời gian xử lý requests (latency). Dù request-based đúng method, metric sai nên không verify được <300ms. 📉 -
[SAI] Select an availability metric for a window-based method of evaluation.
❌ Sai vì: Availability + window-based chỉ phù hợp cho SLO uptime (ví dụ: 99% thời gian up), không đo latency. Phương pháp này dùng cho "good minutes" chứ không phải độ trễ requests cụ thể. 🚫
🧠 Lưu ý bổ sung: Trong thực tế triển khai (DevOps best practices), sử dụng Cloud Monitoring dashboards để visualize SLO, kết hợp với Alerting Policies nếu burn rate vượt ngưỡng. Luôn test SLO trước khi publish service!