Ngân hàng đề — Google Cloud Professional Cloud DevOps Engineer

Tìm thấy 269 câu.

Câu 151
Your organization has a containerized web application that runs on-premises. As part of the migration plan to Google Cloud, you need to select a deployment strategy and platform that meets the following acceptance criteria:

1. The platform must be able to direct traffic from Android devices to an Android-specific microservice.
2. The platform must allow for arbitrary percentage-based traffic splitting
3. The deployment strategy must allow for continuous testing of multiple versions of any microservice.

What should you do?
  1. A Deploy the canary release of the application to Cloud Run. Use traffic splitting to direct 10% of user traffic to the canary release based on the revision tag.
  2. B Deploy the canary release of the application to App Engine. Use traffic splitting to direct a subset of user traffic to the new version based on the IP address.
  3. C Deploy the canary release of the application to Compute Engine. Use Anthos Service Mesh with Compute Engine to direct 10% of user traffic to the canary release by configuring the virtual service.
  4. D Deploy the canary release to Google Kubernetes Engine with Anthos Service Mesh. Use traffic splitting to direct 10% of user traffic to the new version based on the user-agent header configured in the virtual service.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một tổ chức có ứng dụng web containerized đang chạy on-premises, và họ đang lập kế hoạch di chuyển lên Google Cloud. Nhiệm vụ là chọn chiến lược triển khai (deployment strategy) và nền tảng (platform) phù hợp với 3 tiêu chí chấp nhận sau:

  1. Nền tảng phải hướng traffic từ thiết bị Android đến microservice dành riêng cho Android 🏠→ Yêu cầu khả năng routing traffic dựa trên đặc tính thiết bị (như user-agent header).
  2. Nền tảng phải hỗ trợ phân chia traffic theo tỷ lệ phần trăm tùy ý (arbitrary percentage-based traffic splitting) 📊→ Có thể split traffic ví dụ 10%, 20%... một cách linh hoạt.
  3. Chiến lược triển khai phải cho phép testing liên tục nhiều phiên bản microservice 🔄→ Phù hợp với canary release để test dần dần nhiều version mà không ảnh hưởng toàn bộ hệ thống.

Mục tiêu là triển khai canary release (phát hành dần dần) để đảm bảo tính ổn định trong migration. Kiến thức dựa trên tài liệu Google Cloud cập nhật đến 2026 (Anthos Service Mesh 1.20+, GKE 1.29+, Cloud Run/App Engine flex mới nhất).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Deploy the canary release to Google Kubernetes Engine with Anthos Service Mesh. Use traffic splitting to direct 10% of user traffic to the new version based on the user-agent header configured in the virtual service.

Lý do:

  • 🛠️ GKE + Anthos Service Mesh hỗ trợ đầy đủ virtual service để route traffic dựa trên user-agent header (tiêu chí 1: hướng Android traffic đến microservice cụ thể).
  • 📊 Hỗ trợ traffic splitting theo % tùy ý (tiêu chí 2: ví dụ 10% như mô tả).
  • 🔄 Canary release trên GKE cho phép continuous testing nhiều version (tiêu chí 3) qua Istio traffic management.
    Đây là giải pháp container-native lý tưởng cho migration từ on-premises, scalable và production-ready đến 2026.

❌ Phân tích tất cả các phương án

  • Deploy the canary release of the application to Cloud Run. Use traffic splitting to direct 10% of user traffic to the canary release based on the revision tag.
    ❌ Sai vì: Cloud Run hỗ trợ traffic splitting theo revision tag (tiêu chí 2 OK với % tùy ý), và canary release cho testing (tiêu chí 3), nhưng KHÔNG hỗ trợ routing dựa trên user-agent header (tiêu chí 1 thất bại). Cloud Run chỉ split theo tag/revision, không có header-based routing như Istio. Không phù hợp cho Android-specific microservice.

  • Deploy the canary release of the application to App Engine. Use traffic splitting to direct a subset of user traffic to the new version based on the IP address.
    ❌ Sai vì: App Engine flex hỗ trợ traffic splitting (tiêu chí 2), canary cho testing (tiêu chí 3), nhưng chỉ split theo IP address (không linh hoạt cho user-agent/Android - tiêu chí 1 thất bại). IP-based không chính xác cho thiết bị di động (Android có thể dùng nhiều IP). App Engine không phải container-first cho microservices phức tạp.

  • Deploy the canary release of the application to Compute Engine. Use Anthos Service Mesh with Compute Engine to direct 10% of user traffic to the canary release by configuring the virtual service.
    ❌ Sai vì: Anthos Service Mesh có thể cấu hình virtual service cho % split (tiêu chí 2) và testing (tiêu chí 3), nhưng Compute Engine (VM-based) KHÔNG phải nền tảng containerized chuẩn (tiêu chí 1 thất bại vì routing Android-specific kém hiệu quả trên VM). Anthos Service Mesh ưu tiên GKE/on-prem clusters, không scale tốt cho containers trên Compute Engine so với GKE.

  • Deploy the canary release to Google Kubernetes Engine with Anthos Service Mesh. Use traffic splitting to direct 10% of user traffic to the new version based on the user-agent header configured in the virtual service.
    ✅ Đúng hoàn toàn như đã giải thích ở phần trên: Đầy đủ 3 tiêu chí, là lựa chọn best practice cho container migration trên Google Cloud! 🚀

Câu 152
Your team is running microservices in Google Kubernetes Engine (GKE). You want to detect consumption of an error budget to protect customers and define release policies. What should you do?
  1. A Create SLIs from metrics. Enable Alert Policies if the services do not pass.
  2. B Use the metrics from Anthos Service Mesh to measure the health of the microservices.
  3. C Create a SLO. Create an Alert Policy on select_slo_burn_rate.
  4. D Create a SLO and configure uptime checks for your services. Enable Alert Policies if the services do not pass.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào quản lý độ tin cậy dịch vụ (SRE - Site Reliability Engineering) trong môi trường Google Kubernetes Engine (GKE), nơi đội ngũ đang chạy các microservices. Mục tiêu chính là phát hiện việc tiêu thụ error budget (ngân sách lỗi - phần dung lượng lỗi được phép trong một khoảng thời gian để đảm bảo SLO) nhằm bảo vệ khách hàng (bằng cách tránh tình trạng dịch vụ kém chất lượng kéo dài) và xác định chính sách phát hành (release policies, ví dụ: dừng release nếu error budget bị đốt quá nhanh).

  • Error budget: Là khái niệm cốt lõi trong SRE, đại diện cho "phần trăm lỗi cho phép" (ví dụ: nếu SLO là 99.9% uptime, error budget là 0.1%). Khi error budget bị tiêu thụ nhanh (burn rate cao), cần cảnh báo để tránh rủi ro.
  • Yêu cầu hành động: Cần thiết lập cơ chế giám sát và cảnh báo dựa trên SLO (Service Level Objective), đặc biệt là theo dõi burn rate của error budget trên GKE (sử dụng Google Cloud Monitoring).
  • Ngữ cảnh cập nhật 2026: Theo tài liệu Google Cloud mới nhất (Cloud Monitoring v2 với SLO monitoring nâng cao, hỗ trợ GKE Autopilot và Anthos), việc sử dụng SLO với Alert Policy trên burn_rate là best practice để tự động hóa phát hiện và tích hợp với release gates (như trong Cloud Build hoặc ArgoCD).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create a SLO. Create an Alert Policy on select_slo_burn_rate.

Lý do 🛠️:

  • Đây là cách chuẩn xác và trực tiếp nhất để phát hiện tiêu thụ error budget. Trước tiên, tạo SLO (dựa trên SLI như request latency/success rate từ metrics GKE). Sau đó, thiết lập Alert Policy cụ thể trên metric select_slo_burn_rate (có sẵn trong Cloud Monitoring), cho phép cảnh báo khi burn rate (tỷ lệ đốt error budget) vượt ngưỡng (ví dụ: >100% trong 6h hoặc >10% trong 28 ngày).
  • Tích hợp hoàn hảo với GKE: Tự động bảo vệ khách hàng (throttle traffic nếu burn cao) và định nghĩa release policies (block deployment qua SLO gates).
  • Phù hợp best practice SRE Google, hỗ trợ multi-burn-rate alerting mới nhất (2025+).

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ Phương án SAI: Create SLIs from metrics. Enable Alert Policies if the services do not pass.

    • Giải thích: Chỉ tạo SLI (Service Level Indicator) từ metrics là chưa đủ, vì SLI chỉ đo lường (như error rate), không trực tiếp quản lý error budget. Alert Policy trên "không pass" quá mơ hồ, không phát hiện burn rate cụ thể, dẫn đến cảnh báo muộn hoặc không chính xác. Không đáp ứng yêu cầu bảo vệ error budget và release policies.
  • ❌ Phương án SAI: Use the metrics from Anthos Service Mesh to measure the health of the microservices.

    • Giải thích: Anthos Service Mesh (nay là GKE Enterprise với Istio) cung cấp metrics tốt cho health (như traffic, latency), nhưng chỉ đo lường sức khỏe chứ không tự động detect error budget consumption hoặc burn rate. Thiếu SLO/alerting layer, không hỗ trợ release policies trực tiếp. Đây chỉ là phần metrics thô, không phải giải pháp toàn diện.
  • ✅ Phương án ĐÚNG: Create a SLO. Create an Alert Policy on select_slo_burn_rate.

    • Giải thích: Như đã nêu ở trên, đây là cách tối ưu và chính xác, sử dụng SLO để định nghĩa error budget, kết hợp Alert Policy trên select_slo_burn_rate (metric built-in của Cloud Monitoring) để cảnh báo realtime. Hỗ trợ GKE native, bảo vệ khách hàng (auto-remediation) và gate releases (tích hợp CI/CD).
  • ❌ Phương án SAI: Create a SLO and configure uptime checks for your services. Enable Alert Policies if the services do not pass.

    • Giải thích: Tạo SLO là đúng bước đầu, nhưng uptime checks (kiểm tra availability từ bên ngoài) chỉ đo uptime cơ bản, không tính error budget đầy đủ (như latency/errors nội bộ microservices). Alert trên "không pass" quá chung chung, bỏ qua burn rate chi tiết. Không phù hợp cho microservices phức tạp trên GKE, dễ miss các vấn đề partial failure.

Kết luận 🚀: Sử dụng SLO + Alert trên select_slo_burn_rate là best practice SRE cho GKE, giúp đội ngũ chủ động quản lý rủi ro và tuân thủ release policies một cách thông minh!

Câu 153
Your organization wants to collect system logs that will be used to generate dashboards in Cloud Operations for their Google Cloud project. You need to configure all current and future Compute Engine instances to collect the system logs, and you must ensure that the Ops Agent remains up to date. What should you do?
  1. A Use the gcloud CLI to install the Ops Agent on each VM listed in the Cloud Asset Inventory,
  2. B Select all VMs with an Agent status of Not detected on the Cloud Operations VMs dashboard. Then select Install agents.
  3. C Use the gcloud CLI to create an Agent Policy.
  4. D Install the Ops Agent on the Compute Engine image by using a startup script
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc thu thập system logs từ các máy ảo Compute Engine (VM) trong Google Cloud Platform (GCP) để tạo dashboard trên Cloud Operations (trước đây là Stackdriver). Yêu cầu chính bao gồm:

  • Cấu hình tất cả VM hiện tại (current) và tương lai (future) để tự động thu thập logs.
  • Đảm bảo Ops Agent (trước là Fluentd/Stackdriver Agent) luôn được cập nhật tự động (up to date). 🛠️ Mục tiêu cốt lõi: Cần một giải pháp tự động, quy mô lớn áp dụng cho toàn bộ project, không thủ công từng VM, và hỗ trợ auto-install/auto-update theo chính sách mới nhất của GCP (Ops Agent policy từ năm 2022, cập nhật đến 2026 với phiên bản agent v2.30+ hỗ trợ policy-based management).

📘 Tài liệu tham khảo:

✅ Đáp án đúng

Use the gcloud CLI to create an Agent Policy.

Lý do lựa chọn:

  • 🛠️ Agent Policy là tính năng chính thức của GCP (ra mắt 2022, tối ưu hóa đến 2026) cho phép tự động cài đặt và cập nhật Ops Agent trên tất cả VM hiện tại và tương lai trong project/organization/folder.
  • Sử dụng lệnh gcloud compute agent-policies create để tạo policy, sau đó assign vào project → Agent sẽ auto-install qua guest agent trên VM mới và auto-update (pull latest version từ repo GCP).
  • Hỗ trợ thu thập system logs (như syslog) và tích hợp trực tiếp với Cloud Operations để tạo dashboard.
  • Ưu điểm: Scale tự động, không cần can thiệp thủ công, đảm bảo compliance và up-to-date (agent version tự sync hàng tuần).

📋 Giải thích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, với giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên yêu cầu "tất cả current/future VM + auto-update".

  • Use the gcloud CLI to install the Ops Agent on each VM listed in the Cloud Asset Inventory
    ❌ Sai: Phương án này yêu cầu cài thủ công từng VM qua gcloud (dùng gcloud compute instances ops-agents install), liệt kê từ Cloud Asset Inventory. Không áp dụng cho VM tương lai (phải lặp lại thủ công), và không auto-update agent. Không scale cho project lớn, vi phạm yêu cầu tự động hóa.

  • Select all VMs with an Agent status of Not detected on the Cloud Operations VMs dashboard. Then select Install agents.
    ❌ Sai: Chỉ cài thủ công trên dashboard Cloud Operations cho VM hiện tại có status "Not detected". Không bao quát VM tương lai (phải monitor và install lại), và update agent phải làm thủ công. Phù hợp troubleshoot nhỏ lẻ, không phải giải pháp quy mô project theo best practice GCP.

  • Use the gcloud CLI to create an Agent Policy.
    ✅ Đúng: Như giải thích trên, đây là giải pháp chuẩn sử dụng gcloud compute agent-policies create --project=PROJECT_ID để định nghĩa policy thu thập logs, auto-deploy/install/update trên toàn bộ VM current + future. Tích hợp native với Cloud Operations, đảm bảo agent luôn latest version.

  • Install the Ops Agent on the Compute Engine image by using a startup script
    ❌ Sai: Chỉ cài trên custom image qua startup script (metadata script chạy lúc boot). Không áp dụng cho VM hiện tại đã chạy (phải rebuild image và recreate VM), và VM từ image public không dùng được. Update agent thủ công (phải rebuild image mới), không tự động theo policy GCP.

Câu 154
Your company has a Google Cloud resource hierarchy with folders for production, test, and development. Your cyber security team needs to review your company's Google Cloud security posture to accelerate security issue identification and resolution. You need to centralize the logs generated by Google Cloud services from all projects only inside your production folder to allow for alerting and near-real time analysis. What should you do?
  1. A Enable the Workflows API and route all the logs to Cloud Logging.
  2. B Create a central Cloud Monitoring workspace and attach all related projects.
  3. C Create an aggregated log sink associated with the production folder that uses a Pub/Sub topic as the destination.
  4. D Create an aggregated log sink associated with the production folder that uses a Cloud Logging bucket as the destination.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một tổ chức có cấu trúc phân cấp tài nguyên Google Cloud với các folder riêng biệt cho môi trường production (sản xuất), test (kiểm thử) và development (phát triển). Nhóm an ninh mạng cần xem xét tình trạng bảo mật Google Cloud để nhanh chóng xác định và khắc phục vấn đề an ninh. Nhiệm vụ cụ thể là tập trung hóa (centralize) các log được tạo bởi các dịch vụ Google Cloud từ TẤT CẢ các project CHỈ TRONG FOLDER PRODUCTION (không bao gồm test hay dev), nhằm hỗ trợ cảnh báo (alerting) và phân tích gần thời gian thực (near-real time analysis).

🔑 Yêu cầu cốt lõi:

  • Log chỉ từ projects bên trong folder production (sử dụng aggregated sink tại mức folder).
  • Không ảnh hưởng đến các folder khác.
  • Hỗ trợ alerting và near-real time → Cần đích đến (destination) phù hợp cho xử lý thời gian thực, không chỉ lưu trữ.

📘 Kiến thức cập nhật (Google Cloud Logging đến 2026): Aggregated log sinks cho phép export log từ hierarchy (folder/org) xuống Pub/Sub (cho real-time streaming/alerting) hoặc log buckets (cho lưu trữ/query). Tài liệu chính thức: Google Cloud Logging Export Sinks và Hierarchical Resource Manager.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create an aggregated log sink associated with the production folder that uses a Pub/Sub topic as the destination.

Lý do 🛠️:

  • Aggregated sink tại folder production: Chỉ thu thập log từ tất cả projects con bên trong folder đó (không lan sang test/dev), phù hợp yêu cầu "only inside your production folder".
  • Destination là Pub/Sub topic: Hỗ trợ near-real time analysis qua streaming (subscribe bằng Cloud Functions, Dataflow hoặc Monitoring log-based metrics). Pub/Sub cho phép alerting tức thì bằng cách publish log entries để xử lý nhanh, lý tưởng cho security posture review.
  • Hoàn hảo cho mục tiêu accelerate security issue identification vì real-time forwarding đến SIEM hoặc tools phân tích.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên tính năng Google Cloud Logging mới nhất.

  • [SAI] Enable the Workflows API and route all the logs to Cloud Logging.
    ❌ Sai hoàn toàn: Workflows API dùng để orchestrate workflows (tự động hóa quy trình), không liên quan đến routing hoặc centralize logs. Cloud Logging đã mặc định thu thập log mà không cần Workflows. Phương án này không giải quyết aggregated export từ folder, không hỗ trợ real-time alerting, và không giới hạn chỉ production folder.

  • [SAI] Create a central Cloud Monitoring workspace and attach all related projects.
    ❌ Sai: Cloud Monitoring (trước là Stackdriver) dùng cho metrics, dashboards và alerting dựa trên metrics/uptime, không phải centralize logs. Logs thuộc Cloud Logging riêng biệt. Attaching projects chỉ chia sẻ metrics, không export log từ folder hierarchy, và không đảm bảo "near-real time analysis" cho security logs.

  • [ĐÚNG] Create an aggregated log sink associated with the production folder that uses a Pub/Sub topic as the destination.
    ✅ Đúng: Như giải thích ở trên. Aggregated sink (tạo tại folder) tự động export log từ tất cả projects con. Pub/Sub là destination lý tưởng cho streaming real-time, kết nối với Eventarc/Functions để alerting. Hoàn toàn khớp yêu cầu, theo best practices Google Cloud Security (CIS benchmarks 2026).

  • [SAI] Create an aggregated log sink associated with the production folder that uses a Cloud Logging bucket as the destination.
    ❌ Sai một phần: Aggregated sink tại production folder là đúng để giới hạn scope, nhưng Cloud Logging bucket chỉ dùng cho lưu trữ/query log lâu dài (retention-based). Bucket không hỗ trợ near-real time alerting/streaming hiệu quả như Pub/Sub (latency cao hơn cho analysis tức thì). Phù hợp storage hơn là security real-time review.

🛡️ Khuyến nghị thực hiện

Hy vọng phân tích này giúp bạn ôn tập hiệu quả! 🚀 Nếu cần demo code Terraform/IaC, hãy hỏi thêm.

Câu 155
You are configuring the frontend tier of an application deployed in Google Cloud. The frontend tier is hosted in nginx and deployed using a managed instance group with an Envoy-based external HTTP(S) load balancer in front. The application is deployed entirely within the europe-west2 region, and only serves users based in the United Kingdom. You need to choose the most cost-effective network tier and load balancing configuration. What should you use?
  1. A Premium Tier with a global load balancer
  2. B Premium Tier with a regional load balancer
  3. C Standard Tier with a global load balancer
  4. D Standard Tier with a regional load balancer
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc cấu hình tier frontend của một ứng dụng được triển khai trên Google Cloud Platform (GCP). Cụ thể:

  • Frontend được host trên nginx, triển khai qua Managed Instance Group (MIG).
  • Phía trước có Envoy-based external HTTP(S) load balancer (một loại load balancer HTTP(S) bên ngoài sử dụng Envoy làm proxy).
  • Toàn bộ ứng dụng nằm hoàn toàn trong region europe-west2 (vùng châu Âu Tây 2, bao gồm các zone ở London, UK).
  • Ứng dụng chỉ phục vụ người dùng tại Vương quốc Anh (UK), không cần phạm vi toàn cầu.
  • Mục tiêu: Chọn network tier và cấu hình load balancing tiết kiệm chi phí nhất (most cost-effective).

🛠️ Bối cảnh kỹ thuật chính:

  • GCP có 2 Network Service Tiers:
    • Premium Tier: Hỗ trợ global anycast routing, latency thấp toàn cầu, nhưng đắt hơn (dựa trên egress traffic).
    • Standard Tier: Chỉ regional routing (giới hạn trong region), rẻ hơn đáng kể (khoảng 30-50% chi phí thấp hơn Premium cho traffic nội vùng).
  • Load balancer có thể là global (Premium) hoặc regional (Standard).
  • Vì app chỉ ở một region duy nhất và users chỉ ở UK (thuộc europe-west2), không cần global LB để tiết kiệm chi phí và tránh phí egress không cần thiết.

📘 Dẫn nguồn tham khảo (cập nhật đến 2026):

✅ Đáp án đúng: Standard Tier with a regional load balancer

Lý do lựa chọn:

  • Tiết kiệm chi phí tối ưu 🤑: Standard Tier chỉ tính phí regional, không có chi phí global anycast. Với app chỉ ở europe-west2 và users UK (cùng region), regional LB đủ đáp ứng, tránh phí Premium (đắt gấp 3-5 lần cho traffic thấp).
  • Phù hợp kỹ thuật 🛠️: Envoy-based external HTTP(S) LB hỗ trợ regional mode với Standard Tier. MIG trong cùng region đảm bảo low latency nội vùng mà không cần global routing.
  • Không lãng phí: Không cần Premium/global vì không có multi-region hoặc global users.

❌ Phân tích tất cả các phương án (đúng/sai)

  • [SAI] Premium Tier with a global load balancer
    ❌ Sai vì: Premium Tier với global LB quá đắt (phí cao cho anycast routing toàn cầu), dù app chỉ ở một region và users địa phương. Không cost-effective, lãng phí cho traffic nội vùng UK.

  • [SAI] Premium Tier with a regional load balancer
    ❌ Sai vì: Premium Tier không hỗ trợ regional LB cho HTTP(S) (chỉ global). Dù cố config regional, vẫn bị charge như Premium (đắt hơn Standard), không tối ưu chi phí.

  • [SAI] Standard Tier with a global load balancer
    ❌ Sai vì: Standard Tier không hỗ trợ global LB (chỉ regional). Global LB yêu cầu Premium Tier, nên lựa chọn này không khả thi về mặt kỹ thuật.

  • [ĐÚNG] Standard Tier with a regional load balancer
    ✅ Đúng vì: Kết hợp hoàn hảo – Standard Tier rẻ, regional LB phù hợp 100% cho single-region app + local users. Latency thấp, scaling tốt với MIG + Envoy, chi phí thấp nhất theo pricing GCP 2026.

Câu 156
You recently deployed your application in Google Kubernetes Engine (GKE) and now need to release a new version of the application. You need the ability to instantly roll back to the previous version of the application in case there are issues with the new version. Which deployment model should you use?
  1. A Perform a rolling deployment, and test your new application after the deployment is complete.
  2. B Perform A/B testing, and test your application periodically after the deployment is complete.
  3. C Perform a canary deployment, and test your new application periodically after the new version is deployed.
  4. D Perform a blue/green deployment, and test your new application after the deployment is complete.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào quy trình triển khai ứng dụng trên Google Kubernetes Engine (GKE) – một dịch vụ quản lý Kubernetes của Google Cloud. Tình huống: Bạn đã triển khai ứng dụng lên GKE và giờ cần phát hành phiên bản mới. Yêu cầu chính là có khả năng rollback ngay lập tức (instantly roll back) về phiên bản cũ nếu phiên bản mới gặp vấn đề.

🛠️ Mục tiêu cốt lõi: Chọn mô hình triển khai (deployment model) hỗ trợ zero-downtime và rollback tức thì mà không làm gián đoạn dịch vụ. Trong GKE, các chiến lược triển khai phổ biến bao gồm rolling update (mặc định), canary, blue/green, v.v., được hỗ trợ qua Deployment, Service, và công cụ như Istio hoặc Knative cho advanced features (cập nhật đến 2026 với GKE Enterprise và Autopilot mode).

📘 Kiến thức cập nhật: Theo tài liệu GKE mới nhất (2026), blue/green deployment cho phép duy trì hai môi trường song song (blue: cũ, green: mới), switch traffic qua Service selector hoặc Istio VirtualService, đảm bảo rollback chỉ bằng một lệnh kubectl hoặc API call.

✅ Đáp án đúng: Perform a blue/green deployment, and test your new application after the deployment is complete.

Lý do lựa chọn:

  • Blue/green deployment tạo hai môi trường riêng biệt: blue (phiên bản cũ, đang live) và green (phiên bản mới). Traffic chỉ switch sang green sau khi test thành công.
  • Rollback tức thì: Nếu green có vấn đề, chỉ cần switch traffic về blue trong vài giây (không cần redeploy), đạt zero-downtime.
  • Phù hợp hoàn hảo với yêu cầu "instantly roll back". Trong GKE, triển khai qua kubectl apply với Service selector hoặc GKE Gateway API (mới 2025+).

Nguồn tham khảo:

📋 Giải thích tất cả các phương án

  • Perform a rolling deployment, and test your new application after the deployment is complete.
    ❌ Sai: Rolling deployment (mặc định của Kubernetes Deployment) thay thế Pods dần dần (ví dụ: maxUnavailable=25%). Rollback yêu cầu kubectl rollout undo, nhưng không instant vì cần thời gian recreate Pods cũ (có thể vài phút, rủi ro downtime nếu surge không đủ). Test sau deploy hoàn tất không đảm bảo rollback nhanh.

  • Perform A/B testing, and test your application periodically after the deployment is complete.
    ❌ Sai: A/B testing dùng để so sánh variants (qua traffic splitting với Istio), không phải mô hình deployment chính cho rollback. Test "periodically" (định kỳ) không hỗ trợ instant rollback; traffic đã mix, khó switch toàn bộ về cũ ngay lập tức. Phù hợp metrics so sánh hơn là release với rollback khẩn cấp.

  • Perform a canary deployment, and test your new application periodically after the new version is deployed.
    ❌ Sai: Canary release đẩy traffic nhỏ dần đến phiên bản mới (qua Istio hoặc manual selector). Rollback bằng giảm traffic về 0%, nhưng không instant vì cần monitor và adjust dần (test "periodically" mất thời gian). Rủi ro blast radius nếu canary fail muộn.

  • Perform a blue/green deployment, and test your new application after the deployment is complete.
    ✅ Đúng: Như giải thích trên, hỗ trợ switch traffic instant giữa blue/green qua Service hoặc Istio, test green riêng trước khi promote. Rollback chỉ là revert selector – hoàn hảo cho GKE Autopilot/Standard clusters (2026).

🛡️ Lời khuyên DevOps: Trong thực tế GKE, kết hợp blue/green với CI/CD (Cloud Build + ArgoCD) và monitoring (Cloud Monitoring/Prometheus) để automate rollback. Tránh rolling cho prod critical apps cần instant recovery!

Câu 157
You are building and deploying a microservice on Cloud Run for your organization. Your service is used by many applications internally. You are deploying a new release, and you need to test the new version extensively in the staging and production environments. You must minimize user and developer impact. What should you do?
  1. A Deploy the new version of the service to the staging environment. Split the traffic, and allow 1% of traffic through to the latest version. Test the latest version. If the test passes, gradually roll out the latest version to the staging and production environments.
  2. B Deploy the new version of the service to the staging environment. Split the traffic, and allow 50% of traffic through to the latest version. Test the latest version. If the test passes, send all traffic to the latest version. Repeat for the production environment.
  3. C Deploy the new version of the service to the staging environment with a new-release tag without serving traffic. Test the new-release version. If the test passes, gradually roll out this tagged version. Repeat for the production environment.
  4. D Deploy a new environment with the green tag to use as the staging environment. Deploy the new version of the service to the green environment and test the new version. If the tests pass, send all traffic to the green environment and delete the existing staging environment. Repeat for the production environment.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm

📘 Nội dung câu hỏi:
Câu hỏi mô tả tình huống bạn đang xây dựng và triển khai một microservice trên Cloud Run (dịch vụ serverless của Google Cloud) cho tổ chức. Microservice này được sử dụng bởi nhiều ứng dụng nội bộ. Bạn đang triển khai một phiên bản mới và cần kiểm tra kỹ lưỡng (test extensively) phiên bản mới này trong môi trường staging và production. Yêu cầu quan trọng là giảm thiểu tối đa tác động đến người dùng và developer (minimize user and developer impact).
🛠️ Mục tiêu chính: Tìm cách deploy phiên bản mới sao cho có thể test đầy đủ mà không ảnh hưởng đến traffic hiện tại, sau đó rollout dần dần nếu test thành công. Cloud Run hỗ trợ revisions (phiên bản container), tags (nhãn để quản lý revisions), và traffic splitting (phân bổ traffic dần dần), theo tài liệu chính thức GCP cập nhật đến 2026 (Cloud Run v2 API với gradual rollout lên đến 100% tự động).

✅ Đáp án đúng:
Phương án thứ 3: Deploy the new version of the service to the staging environment with a new-release tag without serving traffic. Test the new-release version. If the test passes, gradually roll out this tagged version. Repeat for the production environment.

Lý do chọn đáp án đúng (🟢):
Phương án này hoàn hảo vì sử dụng tag "new-release" để deploy revision mới mà KHÔNG phục vụ traffic ngay lập tức (without serving traffic), cho phép test riêng biệt mà zero impact đến người dùng/developer. Sau khi test pass, gradually roll out (rollout dần dần) tag này để kiểm soát traffic mượt mà. Quy trình lặp lại cho production đảm bảo an toàn. Đây là best practice của Cloud Run theo Traffic Management và Revision Tags (GCP docs: cloud.google.com/run/docs/traffic-splitting, cập nhật 2025 với auto-rollout policies).

🔍 Giải thích TẤT CẢ các phương án (đúng/sai)

  • ❌ Phương án 1 (SAI): Deploy the new version of the service to the staging environment. Split the traffic, and allow 1% of traffic through to the latest version. Test the latest version. If the test passes, gradually roll out the latest version to the staging and production environments.
    Lý do SAI: Việc split 1% traffic ngay từ đầu vẫn tạo impact nhỏ đến người dùng thực tế (dù chỉ 1%), vi phạm yêu cầu "minimize user impact". Test phải "extensively" mà không ảnh hưởng traffic staging/production. Cloud Run khuyến nghị test revision riêng trước khi split (không phải split sớm).

  • ❌ Phương án 2 (SAI): Deploy the new version of the service to the staging environment. Split the traffic, and allow 50% of traffic through to the latest version. Test the latest version. If the test passes, send all traffic to the latest version. Repeat for the production environment.
    Lý do SAI: Split 50% traffic tạo impact lớn ngay lập tức (nửa traffic đi qua version mới chưa test đầy đủ), rủi ro cao cho người dùng nội bộ. Không phù hợp với "test extensively" mà minimize impact. Cloud Run hỗ trợ split nhưng chỉ sau khi revision đã sẵn sàng và test riêng.

  • ✅ Phương án 3 (ĐÚNG): Deploy the new version of the service to the staging environment with a new-release tag without serving traffic. Test the new-release version. If the test passes, gradually roll out this tagged version. Repeat for the production environment.
    Lý do ĐÚNG: Deploy với tag riêng (như "new-release") cho phép revision mới tồn tại không nhận traffic (0% traffic), test độc lập qua endpoint tag-specific (ví dụ: url-tag-new-release). Sau test pass, gradual rollout (tăng traffic từ 0% lên 100% theo % hoặc auto). An toàn, zero-downtime, lặp staging → production. Hoàn toàn khớp best practice Cloud Run (revisions immutable, tags mutable).

  • ❌ Phương án 4 (SAI): Deploy a new environment with the green tag to use as the staging environment. Deploy the new version of the service to the green environment and test the new version. If the tests pass, send all traffic to the green environment and delete the existing staging environment. Repeat for the production environment.
    Lý do SAI: Cloud Run KHÔNG có khái niệm "new environment" riêng biệt như blue-green deployment truyền thống (GKE hoặc ECS mới hỗ trợ). "Green tag" chỉ là tag trên cùng service/revision, không tạo environment mới. Việc "delete existing" gây gián đoạn, không gradual, và phức tạp hóa staging/production. Không minimize impact (switch all traffic đột ngột).

📚 Tài liệu tham khảo (cập nhật mới nhất 2026)

  • GCP Cloud Run Docs - Traffic Splitting & Tags: cloud.google.com/run/docs/traffic-splitting (hỗ trợ gradual rollout với 0.1% increments).
  • Revision Management: cloud.google.com/run/docs/managing/revisions (deploy tag without traffic).
  • DevOps Best Practices: Google Cloud Skills Boost - Professional Cloud DevOps Engineer (module Cloud Run deployment strategies, cập nhật Q1/2026).
  • Console/GCLI Example: gcloud run services update --tag new-release --image new-image --region us-central1 (no traffic until update-traffic).

🎯 Kết luận: Phương án 3 là lựa chọn tối ưu cho DevOps trên Cloud Run, đảm bảo canary-like testing với zero-impact ban đầu! 🚀

Câu 158
You work for a global organization and run a service with an availability target of 99% with limited engineering resources.
For the current calendar month, you noticed that the service has 99.5% availability. You must ensure that your service meets the defined availability goals and can react to business changes, including the upcoming launch of new features.
You also need to reduce technical debt while minimizing operational costs. You want to follow Google-recommended practices. What should you do?
  1. A Add N+1 redundancy to your service by adding additional compute resources to the service.
  2. B Identify, measure, and eliminate toil by automating repetitive tasks.
  3. C Define an error budget for your service level availability and minimize the remaining error budget.
  4. D Allocate available engineers to the feature backlog while you ensure that the service remains within the availability target.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả tình huống bạn làm việc cho một tổ chức toàn cầu, quản lý một dịch vụ (service) với mục tiêu availability 99% (tỷ lệ sẵn sàng 99%), nhưng nguồn lực kỹ thuật (engineering resources) hạn chế. Trong tháng hiện tại, dịch vụ đạt 99.5% availability (vượt mục tiêu một chút).
Yêu cầu chính:

  • Đảm bảo dịch vụ đạt mục tiêu availability và phản ứng linh hoạt với thay đổi kinh doanh, bao gồm ra mắt tính năng mới (new features).
  • Giảm nợ kỹ thuật (technical debt) và giảm thiểu chi phí vận hành (operational costs).
  • Tuân thủ các thực hành được Google khuyến nghị (Google-recommended practices), cụ thể là các nguyên tắc Site Reliability Engineering (SRE) từ Google, nhấn mạnh vào error budget, toil elimination và cân bằng giữa reliability & development.

📘 Dẫn nguồn: Site Reliability Engineering Workbook (Google, cập nhật 2023-2026 editions), Chapter 7: Eliminating Toil; SRE on SLOs and Error Budgets. Kiến thức SRE vẫn ổn định đến 2026, không thay đổi lớn trên các nền tảng cloud như AWS/Google Cloud.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Identify, measure, and eliminate toil by automating repetitive tasks.

Lý do:
🛠️ Theo nguyên tắc SRE của Google, toil là các công việc thủ công lặp lại, không sáng tạo (như manual ops, monitoring thủ công), chiếm >50% thời gian engineer là vấn đề lớn. Với nguồn lực hạn chế, việc xác định (identify), đo lường (measure) và loại bỏ toil bằng automation sẽ giải phóng thời gian engineer để:

  • Duy trì availability 99%.
  • Xử lý thay đổi kinh doanh (new features).
  • Giảm technical debt (tự động hóa giảm lỗi thủ công).
  • Giảm costs (ít nhân sự thủ công).
    Hiện tại availability 99.5% > target, nên còn error budget để đầu tư automation mà không rủi ro downtime. Đây là Google-recommended practice cốt lõi để scale với limited resources.

📋 Giải thích chi tiết tất cả các phương án

  • [SAI] Add N+1 redundancy to your service by adding additional compute resources to the service.
    ❌ Sai vì: Thêm redundancy N+1 (dự phòng thừa) chỉ tăng availability tạm thời bằng cách tăng chi phí compute (thêm resources), nhưng không giải quyết gốc rễ: limited resources, technical debt, toil. Không theo SRE – Google ưu tiên reliability qua automation/smart design, không phải "throw hardware" (chi phí cao, không scale). Trên AWS (EC2 Auto Scaling), điều này tốn kém và không giảm operational costs.

  • [ĐÚNG] Identify, measure, and eliminate toil by automating repetitive tasks.
    ✅ Đúng như đã giải thích ở trên: Tập trung automation toil (ví dụ: Terraform/CloudFormation trên AWS cho infra, Lambda cho tasks lặp), giải phóng 20-50% thời gian engineer theo SRE metrics. Phù hợp với availability vượt target, cho phép innovate features mà vẫn reliable.

  • [SAI] Define an error budget for your service level availability and minimize the remaining error budget.
    ❌ Sai vì: SRE dùng error budget (phần "lỗi cho phép" = 1% downtime) để cân bằng dev & reliability – nếu còn budget (như 99.5% > 99%), KHÔNG minimize mà DÙNG NÓ cho features. "Minimize remaining" nghĩa là cố đạt 100% availability, dẫn đến over-engineering, block features, tăng debt/costs. Google khuyến nghị tiêu error budget cho business value, không phải zero-budget.

  • [SAI] Allocate available engineers to the feature backlog while you ensure that the service remains within the availability target.
    ❌ Sai vì: Phân bổ engineer vào feature backlog (tính năng mới) mà không đề cập toil/error budget sẽ rủi ro availability drop với limited resources. SRE không cho phép "ensure manually" – cần automation + SLO monitoring (AWS CloudWatch). Cách này tăng technical debt (features rush), không minimize costs, vi phạm Google practices (phải balance via error budget, không ưu tiên features blindly).

🧩 Kết luận: Chọn automation toil là cách SRE-native nhất, giúp service scale bền vững trên cloud (AWS/GCP), giảm chi phí dài hạn ~30-50% theo case studies Google. Nếu implement trên AWS: Dùng EventBridge + Lambda cho toil automation!

Câu 159
You are developing the deployment and testing strategies for your CI/CD pipeline in Google Cloud. You must be able to:
•Reduce the complexity of release deployments and minimize the duration of deployment rollbacks.
•Test real production traffic with a gradual increase in the number of affected users.

You want to select a deployment and testing strategy that meets your requirements. What should you do?
  1. A Recreate deployment and canary testing
  2. B Blue/green deployment and canary testing
  3. C Rolling update deployment and A/B testing
  4. D Rolling update deployment and shadow testing
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc phát triển chiến lược triển khai (deployment) và kiểm thử (testing) cho pipeline CI/CD trên Google Cloud. Các yêu cầu cụ thể bao gồm:

  • Giảm độ phức tạp (complexity) của các lần triển khai release và giảm thiểu thời gian rollback deployment (tức là khả năng quay lại phiên bản cũ nhanh chóng khi có sự cố).
  • Kiểm thử với lưu lượng thực tế từ production traffic, đồng thời tăng dần số lượng người dùng bị ảnh hưởng (gradual increase in affected users) để đảm bảo an toàn và kiểm soát rủi ro.

Mục tiêu là chọn kết hợp deployment strategy + testing strategy phù hợp nhất trong môi trường Google Cloud, thường áp dụng cho các dịch vụ như Google Kubernetes Engine (GKE), Cloud Run hoặc Anthos. Các chiến lược này giúp triển khai zero-downtime, dễ rollback và test dần dần với traffic thật. 📘 (Tham khảo: Google Cloud Deployment Strategies và Progressive Delivery on GKE, cập nhật đến 2026 với hỗ trợ GKE Enterprise v1.30+).

✅ Đáp án đúng: Blue/green deployment and canary testing

Lý do lựa chọn:

  • Blue/green deployment ✅: Tạo hai môi trường song song (blue: phiên bản cũ ổn định; green: phiên bản mới). Triển khai chỉ bằng cách switch traffic từ blue sang green, giúp giảm complexity (không cần thay đổi dần dần pod/container) và rollback cực nhanh (chỉ switch ngược lại trong vài giây, không downtime). Hoàn hảo cho yêu cầu minimize rollback duration.
  • Canary testing ✅: Gửi traffic production thật đến phiên bản mới với tỷ lệ nhỏ dần dần tăng (ví dụ: 5% → 20% → 100% users), giúp test an toàn mà không ảnh hưởng toàn bộ hệ thống.
  • Kết hợp lý tưởng trên GKE với Progressive Rollouts hoặc Cloud Deploy, hỗ trợ tự động hóa CI/CD. Điều này đáp ứng toàn bộ yêu cầu mà không có điểm yếu. 🛠️

📋 Giải thích tất cả các phương án

  • Recreate deployment and canary testing ❌
    Sai vì: Recreate deployment xóa toàn bộ phiên bản cũ rồi tạo mới, gây downtime lớn và complexity cao (không zero-downtime, rollback khó khăn vì phải recreate lại cũ). Canary testing tốt cho gradual traffic nhưng không bù đắp được nhược điểm của recreate. Không giảm complexity hay rollback nhanh.

  • Blue/green deployment and canary testing ✅
    (Như đã giải thích ở trên: Hoàn hảo khớp yêu cầu, zero-downtime, rollback tức thì và test gradual traffic.)

  • Rolling update deployment and A/B testing ❌
    Sai vì: Rolling update thay thế dần dần pod (ví dụ: 25% mỗi lần), giúp zero-downtime nhưng complexity cao (cấu hình maxSurge/maxUnavailable phức tạp) và rollback chậm (phải roll back từng bước). A/B testing dùng để so sánh feature khác nhau (không phải gradual production traffic thuần túy), thường dựa user segmentation chứ không tăng dần users affected.

  • Rolling update deployment and shadow testing ❌
    Sai vì: Rolling update như trên, không tối ưu complexity/rollback. Shadow testing gửi traffic thật nhưng không ảnh hưởng response đến user (chỉ log/analyze), nên không test "affected users" dần dần (users không thực sự bị ảnh hưởng, chỉ simulate). Không khớp yêu cầu test production traffic với gradual users.

Tóm lại, chỉ blue/green + canary mới cân bằng hoàn hảo! 🚀 (Nguồn bổ sung: AWS tương đương cho so sánh nhưng ưu tiên GCP docs 2026).

Câu 160 Chọn nhiều đáp án
You are creating a CI/CD pipeline to perform Terraform deployments of Google Cloud resources. Your CI/CD tooling is running in Google Kubernetes Engine (GKE) and uses an ephemeral Pod for each pipeline run. You must ensure that the pipelines that run in the Pods have the appropriate Identity and Access Management (IAM) permissions to perform the Terraform deployments. You want to follow Google-recommended practices for identity management. What should you do? (Choose two.)
  1. A Create a new Kubernetes service account, and assign the service account to the Pods. Use Workload Identity to authenticate as the Google service account.
  2. B Create a new JSON service account key for the Google service account, store the key as a Kubernetes secret, inject the key into the Pods, and set the GOOGLE_APPLICATION_CREDENTIALS environment variable.
  3. C Create a new Google service account, and assign the appropriate IAM permissions.
  4. D Create a new JSON service account key for the Google service account, store the key in the secret management store for the CI/CD tool, and configure Terraform to use this key for authentication.
  5. E Assign the appropriate IAM permissions to the Google service account associated with the Compute Engine VM instances that run the Pods.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc thiết lập CI/CD pipeline sử dụng Terraform để triển khai tài nguyên Google Cloud (GCP). Pipeline chạy trên Google Kubernetes Engine (GKE) với các Pod tạm thời (ephemeral Pod) cho mỗi lần chạy. Yêu cầu chính là đảm bảo các Pod có quyền IAM (Identity and Access Management) phù hợp để Terraform thực hiện triển khai, đồng thời tuân thủ thực hành tốt nhất (Google-recommended practices) về quản lý danh tính.

📌 Bối cảnh quan trọng:

  • Ephemeral Pods nghĩa là Pod được tạo mới mỗi lần pipeline chạy và bị xóa sau đó, nên không nên sử dụng key file tĩnh (như JSON keys) vì rủi ro bảo mật cao (dễ leak, khó xoay vòng).
  • Google-recommended practice cho workload trên GKE là sử dụng Workload Identity: Kết nối Kubernetes Service Account (KSA) với Google Service Account (GSA) để Pods tự động authenticate mà không cần key.
  • Câu hỏi yêu cầu chọn hai đáp án đúng (Choose two), phù hợp với quy trình: Tạo GSA với IAM roles → Tạo KSA và bind với GSA qua Workload Identity → Assign KSA cho Pods.

🛠️ Kiến thức cập nhật (đến 2026): Theo tài liệu GCP mới nhất (GKE 1.29+ và Workload Identity Federation), cách này là tiêu chuẩn cho CI/CD tools như Cloud Build, GitLab, Jenkins trên GKE. Tránh service account keys theo nguyên tắc "least privilege" và zero-trust.

✅ Đáp án đúng và lý do lựa chọn

Hai đáp án đúng là phương án đầu tiên và phương án thứ ba (A và C theo thứ tự liệt kê).

  • Lý do chọn:
    • Đây là quy trình chuẩn theo Workload Identity trên GKE: Tạo GSA mới (C) với IAM roles cần thiết (ví dụ: roles/resourcemanager.projectIamAdmin, roles/compute.admin cho Terraform). Sau đó tạo KSA (A), enable Workload Identity trên cluster, bind KSA với GSA bằng annotation iam.gke.io/gcp-service-account, và assign KSA cho Pod spec trong CI/CD. Pods sẽ impersonate GSA mà không cần key, an toàn cho ephemeral workloads.
    • Tuân thủ Google best practices: Giảm bề mặt tấn công, tự động xoay credentials, tích hợp native với Terraform provider google (hỗ trợ Workload Identity từ v4.0+).

📋 Giải thích tất cả các phương án (đúng/sai)

  • ✅ Create a new Kubernetes service account, and assign the service account to the Pods. Use Workload Identity to authenticate as the Google service account.
    Đúng: Đây là bước cuối cùng trong quy trình Workload Identity. KSA được assign vào Pod qua serviceAccountName trong manifest, và annotation bind với GSA cho phép Pod sử dụng short-lived OIDC tokens để gọi GCP APIs. Hoàn hảo cho ephemeral Pods, không lưu key lâu dài. 🛡️

  • ❌ Create a new JSON service account key for the Google service account, store the key as a Kubernetes secret, inject the key into the Pods, and set the GOOGLE_APPLICATION_CREDENTIALS environment variable.
    Sai: Sử dụng JSON key vi phạm best practices vì key tĩnh, dễ bị leak trong Pods tạm thời (qua logs, volumes). Google khuyến cáo không dùng keys cho workloads trên GKE; thay vào đó dùng Workload Identity. Rủi ro cao với CI/CD. 🚫

  • ✅ Create a new Google service account, and assign the appropriate IAM permissions.
    Đúng: Bước đầu tiên cần thiết. Tạo GSA riêng (không dùng default) với principle of least privilege (ví dụ: custom roles cho Terraform actions). GSA này sẽ được bind với KSA để Pods kế thừa quyền. Thiết yếu cho quy trình đầy đủ. 🔑

  • ❌ Create a new JSON service account key for the Google service account, store the key in the secret management store for the CI/CD tool, and configure Terraform to use this key for authentication.
    Sai: Tương tự phương án 2, JSON key không an toàn cho CI/CD trên GKE (dù dùng secret store như Secret Manager). Phải xoay key thủ công, tăng complexity và rủi ro compromise. Google deprecated keys cho workloads managed. 🔒❌

  • ❌ Assign the appropriate IAM permissions to the Google service account associated with the Compute Engine VM instances that run the Pods.
    Sai: GKE Pods chạy trên node pools (Compute Engine VMs), nhưng GSA mặc định của VM (như default-compute-service-account) chỉ cho node-level access, không bind trực tiếp với Pods. Sử dụng sẽ grant quyền quá rộng (toàn cluster), vi phạm least privilege. Phải dùng Workload Identity cho per-Pod isolation. 🖥️🚫

📘 Tài liệu tham khảo (Google Cloud Docs - cập nhật 2026)

Hy vọng phân tích này giúp bạn ôn thi chứng chỉ hiệu quả! 🚀 Nếu cần ví dụ code Terraform/YAML, hãy hỏi thêm.