Ngân hàng đề — Google Cloud Professional Cloud DevOps Engineer

Tìm thấy 269 câu.

Câu 171
A third-party application needs to have a service account key to work properly. When you try to export the key from your cloud project, you receive an error: “The organization policy constraint iam.disableServiceAccounKeyCreation is enforced.” You need to make the third-party application work while following Google-recommended security practices.

What should you do?
  1. A Enable the default service account key, and download the key.
  2. B Remove the iam.disableServiceAccountKeyCreation policy at the organization level, and create a key.
  3. C Disable the service account key creation policy at the project's folder, and download the default key.
  4. D Add a rule to set the iam.disableServiceAccountKeyCreation policy to off in your project, and create a key.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả tình huống: Một ứng dụng bên thứ ba (third-party application) cần service account key để hoạt động bình thường. Khi cố gắng export key từ Google Cloud project, bạn gặp lỗi: “The organization policy constraint iam.disableServiceAccounKeyCreation is enforced.”
✅ Vấn đề cốt lõi: Organization policy iam.disableServiceAccountKeyCreation đang được áp dụng ở mức organization level, ngăn chặn việc tạo service account keys trên toàn tổ chức.
🛠️ Yêu cầu giải quyết: Làm cho ứng dụng hoạt động mà vẫn tuân thủ Google-recommended security practices (thực hành bảo mật được Google khuyến nghị).
📘 Bối cảnh kiến thức (cập nhật đến 2026): Trong Google Cloud, organization policies cho phép quản lý ràng buộc ở các mức (organization > folder > project). Policy iam.disableServiceAccountKeyCreation (boolean policy) mặc định khuyến khích không sử dụng keys để tăng bảo mật (dùng Workload Identity Federation thay thế). Tuy nhiên, nếu cần key cho legacy/third-party apps, có thể override policy ở mức project/folder mà không ảnh hưởng toàn tổ chức. Không nên disable ở org level để tránh rủi ro bảo mật rộng.
🔗 Nguồn tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Add a rule to set the iam.disableServiceAccountKeyCreation policy to off in your project, and create a key.
Lý do:

  • 🛠️ Phương án này override policy ở mức project (mức thấp nhất), chỉ tắt ràng buộc cho project cụ thể mà không ảnh hưởng organization/folder.
  • ✅ Tuân thủ security best practices: Giữ policy chặt chẽ ở org level, chỉ "nới lỏng" cục bộ cho project cần thiết. Sau đó tạo key mới (không dùng default).
  • 📈 Đây là cách Google khuyến nghị cho trường hợp exception (theo docs 2026, hỗ trợ policy inheritance với override).

❌ Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể:

  • [SAI] Enable the default service account key, and download the key.
    ❌ Sai vì: Không tồn tại "default service account key" có thể enable trực tiếp. Service accounts mặc định (như compute engine default) không có key sẵn; policy vẫn chặn tạo/export key mới. Phương án này không giải quyết lỗi policy và vi phạm best practices (không dùng default keys cho third-party).

  • [SAI] Remove the iam.disableServiceAccountKeyCreation policy at the organization level, and create a key.
    ❌ Sai vì: Xóa policy ở organization level sẽ ảnh hưởng toàn bộ tổ chức (hàng nghìn projects), tăng rủi ro bảo mật lớn (dễ leak keys). Google không khuyến nghị disable ở mức cao nhất; chỉ override ở project/folder cho exception.

  • [SAI] Disable the service account key creation policy at the project's folder, and download the default key.
    ❌ Sai vì:

    • Disable ở folder level có thể ảnh hưởng nhiều projects con (không phải chỉ project hiện tại).
    • "Download the default key" sai: Không có default key để download; phải tạo mới. Policy vẫn cần override chính xác ở project để an toàn nhất.
  • [ĐÚNG] Add a rule to set the iam.disableServiceAccountKeyCreation policy to off in your project, and create a key.
    ✅ Đúng vì: Override policy chính xác ở project level (allow inheritance từ org nhưng set off cục bộ). Sau đó tạo key mới cho service account. Hoàn hảo tuân thủ hierarchy override của Google Cloud (org > folder > project), giữ bảo mật tối ưu.
    🛠️ Cách thực hiện thực tế: Sử dụng gcloud org-policies set-policy với rule enforced: false ở project.

Câu 172 Chọn nhiều đáp án
Your team is writing a postmortem after an incident on your external facing application. Your team wants to improve the postmortem policy to include triggers that indicate whether an incident requires a postmortem. Based on Site Reliability Engineering (SRE) practices, what triggers should be defined in the postmortem policy? (Choose two.)
  1. A An external stakeholder asks for a postmortem
  2. B Data is lost due to an incident.
  3. C An internal stakeholder requests a postmortem.
  4. D The monitoring system detects that one of the instances for your application has failed.
  5. E The CD pipeline detects an issue and rolls back a problematic release.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào chính sách postmortem (báo cáo sau sự cố) trong thực hành Site Reliability Engineering (SRE), một phương pháp được Google phát triển và áp dụng rộng rãi trong DevOps trên các nền tảng đám mây như AWS hay Google Cloud.

  • Bối cảnh: Sau một sự cố trên ứng dụng hướng ngoại (external-facing application), đội ngũ muốn cải thiện chính sách postmortem bằng cách định nghĩa các triggers (yếu tố kích hoạt) để quyết định khi nào cần viết postmortem.
  • Yêu cầu chọn: Hai triggers dựa trên nguyên tắc SRE, nhấn mạnh vào tác động thực sự đến khách hàng (customer impact) và mất mát dữ liệu, thay vì các sự cố nội bộ nhỏ lẻ.
  • Mục tiêu SRE: Postmortem không phải viết cho mọi sự cố, mà chỉ dành cho những incident gây tác động lớn (như SLA vi phạm, mất dữ liệu), giúp học hỏi và tránh lặp lại. Theo SRE practices (cập nhật đến 2026), ưu tiên blameless postmortem cho các sự cố có external impact hoặc data loss.

✅ Đáp án đúng (Chọn hai)

Hai lựa chọn đúng là:
An external stakeholder asks for a postmortem
Data is lost due to an incident.

Lý do lựa chọn:
🛠️ Theo SRE principles (Google SRE Book và Workbook, phiên bản cập nhật 2023-2026), postmortem chỉ kích hoạt khi có tác động khách hàng bên ngoài (external customer impact) hoặc mất dữ liệu vĩnh viễn. "External stakeholder" đại diện cho khách hàng thực tế bị ảnh hưởng (như SLA breach), còn "data loss" là rủi ro nghiêm trọng nhất cần phân tích sâu. Những triggers này đảm bảo postmortem tập trung vào vấn đề lớn, tránh lãng phí thời gian cho sự cố nhỏ.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Phần giải thích hoàn toàn bằng tiếng Việt:

  • ✅ An external stakeholder asks for a postmortem
    🟢 Đúng: Đây là trigger chuẩn SRE vì liên quan đến tác động khách hàng bên ngoài. Nếu stakeholder ngoại bộ (khách hàng) yêu cầu, chứng tỏ sự cố ảnh hưởng đến SLA hoặc trải nghiệm người dùng. SRE khuyến khích ưu tiên external impact để cải thiện reliability (theo Google SRE Workbook, Ch. 7: Postmortem Culture).

  • ✅ Data is lost due to an incident.
    🟢 Đúng: Mất dữ liệu là trigger bắt buộc trong SRE, vì nó gây thiệt hại vĩnh viễn và cần phân tích root cause sâu. AWS cũng nhấn mạnh trong Well-Architected Framework (Reliability Pillar, 2024-2026): Data loss yêu cầu postmortem để triển khai backup/DR tốt hơn.

  • ❌ An internal stakeholder requests a postmortem.
    🔴 Sai: Internal stakeholder (như team nội bộ) không đủ để trigger postmortem SRE. SRE tránh overload bằng cách chỉ tập trung external impact; internal issues dùng retrospective nội bộ thay thế (SRE Book: tránh postmortem cho mọi alert nội bộ).

  • ❌ The monitoring system detects that one of the instances for your application has failed.
    🔴 Sai: Một instance fail (trong hệ thống multi-instance) thường không gây downtime nếu có auto-scaling/HA (như EC2 Auto Scaling trên AWS). SRE chỉ trigger postmortem nếu có customer-visible impact, không phải single failure (theo SRE practices: SLO breach mới cần).

  • ❌ The CD pipeline detects an issue and rolls back a problematic release.
    🔴 Sai: Auto-rollback trong CI/CD (như AWS CodePipeline) là self-healing tốt, không gây impact nếu nhanh chóng. SRE coi đây là thành công của automation, không cần postmortem trừ khi có data loss hoặc outage (AWS DevOps Guidance, 2025: Promote canary deployments với rollback tự động).

📘 Tài liệu tham khảo

  • Google SRE Book (2024 edition): Chương 7 - Postmortem Culture: Định nghĩa triggers như customer impact & data loss. Link
  • Google SRE Workbook (2023): Phần Incident Management, ví dụ policy triggers. Link
  • AWS Well-Architected Framework - Reliability Pillar (2026 update): Nhấn mạnh postmortem cho critical incidents với data loss/external impact. Link
  • AWS Certified DevOps Engineer - Professional Exam Guide (2025): Câu hỏi tương tự về SRE integration.

Hy vọng phân tích này giúp bạn chuẩn bị tốt cho kỳ thi! 🚀 Nếu cần thêm ví dụ thực tế trên AWS/GCP, hãy hỏi nhé.

Câu 173
You are implementing a CI/CD pipeline for your application in your company’s multi-cloud environment. Your application is deployed by using custom Compute Engine images and the equivalent in other cloud providers. You need to implement a solution that will enable you to build and deploy the images to your current environment and is adaptable to future changes. Which solution stack should you use?
  1. A Cloud Build with Packer
  2. B Cloud Build with Google Cloud Deploy
  3. C Google Kubernetes Engine with Google Cloud Deploy
  4. D Cloud Build with kpt
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc triển khai CI/CD pipeline cho ứng dụng trong môi trường multi-cloud (đa đám mây, bao gồm Google Cloud và các nhà cung cấp khác). Ứng dụng được triển khai bằng custom Compute Engine images (hình ảnh tùy chỉnh trên Google Compute Engine - dịch vụ VM của Google Cloud) và các tương đương trên các đám mây khác.

Yêu cầu chính:

  • Xây dựng và triển khai images cho môi trường hiện tại.
  • Giải pháp phải linh hoạt, dễ thích ứng với thay đổi tương lai (adaptable to future changes), nghĩa là hỗ trợ multi-cloud, không bị ràng buộc vào một nền tảng cụ thể.

📘 Bối cảnh DevOps: Trong Google Cloud, Cloud Build là công cụ CI/CD mạnh mẽ hỗ trợ tích hợp với các công cụ bên thứ ba. Vì là multi-cloud, cần tool build image cross-platform như Packer (của HashiCorp), giúp tạo images cho GCP Compute Engine, AWS EC2 AMI, Azure Images, v.v. Kiến thức cập nhật đến 2026: Cloud Build vẫn hỗ trợ Packer qua builder chính thức (xem docs GCP 2024-2026).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Cloud Build with Packer

🛠️ Lý do chi tiết:

  • Cloud Build là dịch vụ CI/CD serverless của Google Cloud, hỗ trợ build images nhanh chóng qua các builder tùy chỉnh, tích hợp Git, triggers tự động.
  • Packer là công cụ mã nguồn mở của HashiCorp, chuyên build identical machine images cho multi-cloud (GCP, AWS, Azure, VMware, v.v.) từ một template duy nhất (HCL/JSON). Nó hoàn hảo cho custom Compute Engine images và "equivalent" trên cloud khác.
  • Stack này adaptable: Packer hỗ trợ plugin mở rộng, dễ thêm cloud mới mà không thay đổi pipeline lớn. Cloud Build chạy Packer qua gcr.io/cloud-builders/packer hoặc custom builder.
  • Phù hợp nhất cho yêu cầu: Build/deploy images VM, multi-cloud ready.

Nguồn tham khảo:

📋 Giải thích tất cả các phương án (đúng/sai)

  • ✅ Cloud Build with Packer
    Giải thích đúng: Như trên, đây là stack lý tưởng cho build custom VM images multi-cloud. Packer xử lý việc tạo images giống hệt nhau trên các nền tảng khác nhau, Cloud Build orchestrate pipeline CI/CD. Linh hoạt cao, hỗ trợ tương lai (thêm cloud mới chỉ cần Packer template mới). Hoàn toàn khớp yêu cầu.

  • ❌ Cloud Build with Google Cloud Deploy
    Giải thích sai: Google Cloud Deploy là công cụ delivery/release tập trung vào Kubernetes (GKE/ Anthos), không phải build VM images như Compute Engine. Nó dùng để deploy manifests K8s, không hỗ trợ Packer hay multi-cloud VM images. Không adaptable cho non-K8s environments.

  • ❌ Google Kubernetes Engine with Google Cloud Deploy
    Giải thích sai: GKE (Kubernetes managed) + Cloud Deploy chỉ dành cho containerized workloads trên Kubernetes, không build/deploy VM images (Compute Engine). Hoàn toàn không liên quan đến custom images multi-cloud, thiếu tính linh hoạt cho VM-based apps.

  • ❌ Cloud Build with kpt
    Giải thích sai: kpt (Kubernetes Package Toolkit) dùng quản lý Kubernetes configs (GitOps-style), không build VM images. Cloud Build có thể dùng kpt cho K8s pipelines, nhưng không hỗ trợ Compute Engine images hay multi-cloud VM. Không đáp ứng yêu cầu build/deploy images.

🧠 Tóm tắt insight DevOps: Chọn stack dựa trên workload (VM images → Packer), multi-cloud (→ Packer/ Terraform ecosystem), và GCP-native CI (→ Cloud Build). Tránh Kubernetes tools nếu không phải container! 🚀

Câu 174
Your application's performance in Google Cloud has degraded since the last release. You suspect that downstream dependencies might be causing some requests to take longer to complete. You need to investigate the issue with your application to determine the cause. What should you do?
  1. A Configure Error Reporting in your application.
  2. B Configure Google Cloud Managed Service for Prometheus in your application.
  3. C Configure Cloud Profiler in your application.
  4. D Configure Cloud Trace in your application.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả tình huống: Hiệu suất ứng dụng trên Google Cloud đã suy giảm kể từ bản phát hành gần nhất. Bạn nghi ngờ rằng các dependencies downstream (các dịch vụ phụ thuộc bên dưới) có thể đang gây ra một số request mất nhiều thời gian hơn để hoàn thành. Nhiệm vụ là điều tra vấn đề trong ứng dụng để xác định nguyên nhân.

🔍 Chi tiết vấn đề:

  • Đây là vấn đề về performance degradation (suy giảm hiệu suất), tập trung vào request latency (thời gian xử lý request).
  • Nghi ngờ downstream dependencies (như API calls, microservices khác) gây chậm trễ.
  • Cần công cụ investigate (điều tra) để phân tích distributed tracing (theo dõi chuỗi request qua các dịch vụ), xác định bottleneck (điểm nghẽn) chính xác.

🛠️ Mục tiêu chính: Chọn công cụ Google Cloud phù hợp nhất để trace và phân tích latency từ dependencies, không phải monitoring metrics, error logging hay profiling tài nguyên.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Configure Cloud Trace in your application.

Lý do 📘:

  • Cloud Trace là dịch vụ distributed tracing chuyên dụng của Google Cloud, giúp theo dõi toàn bộ hành trình của request từ đầu đến cuối, bao gồm downstream dependencies. Nó phân tích latency chi tiết (span-level), hiển thị thời gian chờ đợi ở từng service, giúp xác định chính xác điểm nghẽn gây chậm request.
  • Phù hợp hoàn hảo với tình huống: Performance degraded do dependencies → Trace sẽ vẽ trace timeline để visualize bottleneck.
  • Theo tài liệu GCP mới nhất (2024-2026), Cloud Trace hỗ trợ automatic instrumentation cho nhiều ngôn ngữ (Java, Node.js, Go...), tích hợp với Cloud Run, GKE, App Engine, và export dữ liệu sang Cloud Monitoring/Logging 🧩.
  • Nguồn tham khảo: Google Cloud Trace Documentation và Best Practices for Observability.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn (giữ nguyên văn bản gốc tiếng Anh). Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm giải thích chi tiết bằng tiếng Việt:

  • ❌ [SAI] Configure Error Reporting in your application.
    Giải thích: Error Reporting chỉ tập trung vào thu thập và báo cáo lỗi (errors/exceptions), không phân tích latency hay performance của request. Nó hữu ích cho debugging crash, nhưng không trace downstream dependencies hay đo thời gian hoàn thành request. Không giải quyết vấn đề performance degradation ở đây 🛑.

  • ❌ [SAI] Configure Google Cloud Managed Service for Prometheus in your application.
    Giải thích: Managed Service for Prometheus dùng để thu thập metrics thời gian thực (như CPU, memory, custom metrics), phù hợp cho alerting và dashboarding. Tuy nhiên, nó không cung cấp tracing chi tiết cho request flow qua dependencies, chỉ đo tổng quát chứ không phân tích span-level latency. Không phải công cụ investigate bottleneck downstream 🛠️.

  • ❌ [SAI] Configure Cloud Profiler in your application.
    Giải thích: Cloud Profiler chuyên phân tích CPU, memory usage ở mức code-level (sampling profiler), giúp tối ưu hóa code chậm. Nhưng nó không trace request latency từ external dependencies, chỉ tập trung vào tài nguyên nội bộ ứng dụng. Không phù hợp cho vấn đề distributed performance với downstream services ❌.

  • ✅ [ĐÚNG] Configure Cloud Trace in your application.
    Giải thích: Như đã nêu ở trên, đây là lựa chọn tối ưu vì hỗ trợ end-to-end tracing, phân tích chính xác thời gian chờ đợi ở dependencies, giúp pinpoint nguyên nhân degradation. Dễ integrate và scale trên GCP (cập nhật 2026 vẫn là standard cho observability) 🚀.

Kết luận 🌟: Sử dụng Cloud Trace là best practice cho troubleshooting request latency trong môi trường microservices trên Google Cloud. Nếu cần implement, bắt đầu bằng SDK instrumentation hoặc auto-trace cho supported runtimes!

Câu 175 Chọn nhiều đáp án
You are creating a CI/CD pipeline in Cloud Build to build an application container image. The application code is stored in GitHub. Your company requires that production image builds are only run against the main branch and that the change control team approves all pushes to the main branch. You want the image build to be as automated as possible. What should you do? (Choose two.)
  1. A Create a trigger on the Cloud Build job. Set the repository event setting to ‘Pull request’.
  2. B Add the OWNERS file to the Included files filter on the trigger.
  3. C Create a trigger on the Cloud Build job. Set the repository event setting to ‘Push to a branch’
  4. D Configure a branch protection rule for the main branch on the repository.
  5. E Enable the Approval option on the trigger.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi này thuộc chủ đề CI/CD pipeline trên Google Cloud Platform (GCP), cụ thể là sử dụng Cloud Build để xây dựng container image cho ứng dụng. Mã nguồn được lưu trữ trên GitHub. Yêu cầu chính của công ty bao gồm:

  • Chỉ build image production trên nhánh main: Đảm bảo pipeline chỉ kích hoạt khi có push vào nhánh main.
  • Đội ngũ change control phải phê duyệt tất cả push vào nhánh main: Cần cơ chế kiểm soát để tránh push trực tiếp, thường qua Pull Request (PR) và approval.
  • Tự động hóa cao nhất có thể: Tránh các bước thủ công như approval trên trigger của Cloud Build.

Mục tiêu là chọn hai hành động (Choose two) để thiết lập trigger trên Cloud Build sao cho phù hợp, kết hợp với bảo vệ nhánh trên GitHub. Đây là kịch bản thực tế trong DevOps, đảm bảo security và automation theo best practices của GCP đến năm 2026 (Cloud Build hỗ trợ GitHub App integration và branch-specific triggers).

📘 Tài liệu tham khảo:

✅ Đáp án đúng (Chọn hai)

Hai lựa chọn đúng là:

  1. Create a trigger on the Cloud Build job. Set the repository event setting to ‘Push to a branch’
    Lý do: Điều này kích hoạt build chỉ khi có push vào nhánh cụ thể (chỉ định main trong branch filter), đảm bảo image production chỉ build trên main. Kết hợp với branch protection, quy trình tự động sau khi merge PR đã được approve. Đây là cách tối ưu tự động hóa mà không cần manual intervention.

  2. Configure a branch protection rule for the main branch on the repository.
    Lý do: Trên GitHub, rule này yêu cầu PR approval từ change control team trước khi merge vào main, ngăn push trực tiếp. Sau approval và merge, trigger Cloud Build sẽ tự động chạy build – hoàn hảo cho yêu cầu "approves all pushes" mà vẫn automated.

🛠️ Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn một cách đầy đủ, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm lý do cụ thể bằng tiếng Việt:

  • Create a trigger on the Cloud Build job. Set the repository event setting to ‘Pull request’.
    ❌ Sai: Trigger kiểu "Pull request" sẽ kích hoạt build mỗi khi tạo hoặc update PR, dẫn đến build không chỉ trên main mà còn trên các nhánh feature (không phải production). Điều này vi phạm yêu cầu "only run against the main branch" và tạo lãng phí tài nguyên.

  • Add the OWNERS file to the Included files filter on the trigger.
    ❌ Sai: OWNERS file là file config cho GitHub Prow hoặc Kubernetes prow (quản lý code owners cho review), không liên quan đến filter của Cloud Build trigger. Filter "Included files" chỉ kiểm tra thay đổi file cụ thể (ví dụ: **/*.go), không dùng để giới hạn nhánh main hoặc approval. Sử dụng sai ngữ cảnh.

  • Create a trigger on the Cloud Build job. Set the repository event setting to ‘Push to a branch’.
    ✅ Đúng: Như đã giải thích ở phần đáp án, đây là cách chính xác để chỉ build trên push vào main (thiết lập branch name = main trong trigger config). Tự động 100% sau khi merge PR, phù hợp với automation cao.

  • Configure a branch protection rule for the main branch on the repository.
    ✅ Đúng: Như đã giải thích, rule này bắt buộc approval từ change control (require pull request reviews, status checks), ngăn push trực tiếp vào main. Kết hợp với trigger push-to-branch, toàn bộ pipeline automated mà vẫn an toàn.

  • Enable the Approval option on the trigger.
    ❌ Sai: Tùy chọn "Approval required" trên Cloud Build trigger tạo manual gate (ai đó phải click approve thủ công trước khi build), vi phạm yêu cầu "as automated as possible". Approval nên xử lý ở GitHub (branch protection) thay vì Cloud Build để tự động hóa tốt hơn.

Kết luận: Combo hai đáp án đúng tạo quy trình PR → Approval (GitHub) → Merge → Auto Build (Cloud Build), an toàn và tự động! 🚀

Câu 176
You built a serverless application by using Cloud Run and deployed the application to your production environment. You want to identify the resource utilization of the application for cost optimization. What should you do?
  1. A Use Cloud Trace with distributed tracing to monitor the resource utilization of the application.
  2. B Use Cloud Profiler with Ops Agent to monitor the CPU and memory utilization of the application.
  3. C Use Cloud Monitoring to monitor the container CPU and memory utilization of the application.
  4. D Use Cloud Ops to create logs-based metrics to monitor the resource utilization of the application.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc tối ưu hóa chi phí (cost optimization) cho một ứng dụng serverless được xây dựng bằng Cloud Run (dịch vụ chạy container serverless trên Google Cloud) sau khi đã triển khai lên môi trường production.

  • Mục tiêu chính: Xác định sử dụng tài nguyên (resource utilization) của ứng dụng, cụ thể là CPU và memory của container, để giảm chi phí không cần thiết (vì Cloud Run tính phí dựa trên thời gian chạy, CPU và memory sử dụng).
  • Bối cảnh: Cloud Run tự động scale và chỉ tính phí khi có request, nên monitoring metrics như CPU/memory giúp điều chỉnh cấu hình (requests/limits) hoặc scale settings để tiết kiệm.
  • Phiên bản cập nhật: Theo tài liệu Google Cloud mới nhất (2024-2026), Cloud Run hỗ trợ metrics chi tiết qua Cloud Monitoring cho container-level insights.
    📘 Tài liệu tham khảo: Cloud Run metrics in Cloud Monitoring, Cloud Run cost optimization.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Cloud Monitoring to monitor the container CPU and memory utilization of the application.
Lý do:

  • Cloud Monitoring (trước đây là Stackdriver Monitoring) là công cụ chính thức và tích hợp sẵn với Cloud Run, cung cấp metrics tự động như cloud_run.googleapis.com/container/cpu_utilization và cloud_run.googleapis.com/container/memory_utilization.
  • Những metrics này cho phép xem biểu đồ thời gian thực, đặt alert và dashboard để phân tích utilization, từ đó tối ưu chi phí (ví dụ: giảm CPU/memory limits nếu underutilized).
  • 🛠️ Cách thực hiện: Vào Cloud Monitoring > Metrics Explorer, chọn Cloud Run metrics → Tạo chart/alert ngay lập tức mà không cần config thêm. Đây là best practice cho serverless monitoring trên GCP (cập nhật 2026).

❌ Giải thích tất cả các phương án (đúng/sai)

  • [SAI] Use Cloud Trace with distributed tracing to monitor the resource utilization of the application.
    ❌ Sai vì: Cloud Trace chuyên về tracing latency và performance (request spans, bottlenecks), không cung cấp metrics CPU/memory utilization. Nó giúp debug chậm trễ nhưng không đo lường tài nguyên container cho cost optimization. (📘 Tham khảo: Cloud Trace docs).

  • [SAI] Use Cloud Profiler with Ops Agent to monitor the CPU and memory utilization of the application.
    ❌ Sai vì: Cloud Profiler dùng để profile code-level (hotspots CPU/memory trong code), không phải monitor container-level utilization. Ops Agent (cho VM) không cần thiết cho Cloud Run (serverless), và Profiler không thay thế metrics tổng quát cho cost. (📘 Tham khảo: Cloud Profiler for Cloud Run).

  • [ĐÚNG] Use Cloud Monitoring to monitor the container CPU and memory utilization of the application.
    ✅ Đúng như đã giải thích ở trên: Metrics chính xác, dễ dùng, tích hợp sâu với Cloud Run cho cost insights.

  • [SAI] Use Cloud Ops to create logs-based metrics to monitor the resource utilization of the application.
    ❌ Sai vì: "Cloud Ops" không phải dịch vụ chuẩn (có thể ám chỉ Cloud Operations Suite), nhưng logs-based metrics chỉ extract từ logs (như custom logs), không có metrics CPU/memory tự động từ Cloud Run. Phải dùng Cloud Monitoring metrics gốc thay vì tự tạo từ logs (phức tạp và không chính xác). (📘 Tham khảo: Cloud Run logging).

Câu 177
Your company is using HTTPS requests to trigger a public Cloud Run-hosted service accessible at the https://booking-engine-abcdef.a.run.app URL. You need to give developers the ability to test the latest revisions of the service before the service is exposed to customers. What should you do?
  1. A Run the gcloud run deploy booking-engine --no-traffic --tag dev command. Use the https://dev--booking-engine-abcdef.a.run.app URL for testing.
  2. B Run the gcloud run services update-traffic booking-engine --to-revisions LATEST=1 command. Use the https://booking-engine-abcdef.a.run.app URL for testing.
  3. C Pass the curl –H “Authorization:Bearer $(gcloud auth print-identity-token)” auth token. Use the https://booking-engine-abcdef.a.run.app URL to test privately.
  4. D Grant the roles/run.invoker role to the developers testing the booking-engine service. Use the https://booking-engine-abcdef.private.run.app URL for testing.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm về Google Cloud Run

✅ Giải thích nội dung câu hỏi:
Câu hỏi mô tả tình huống công ty đang sử dụng các yêu cầu HTTPS để kích hoạt một dịch vụ Cloud Run được host công khai (publicly accessible) tại URL https://booking-engine-abcdef.a.run.app. Nhiệm vụ là cung cấp cho các lập trình viên (developers) khả năng kiểm tra (test) các phiên bản mới nhất (latest revisions) của dịch vụ trước khi dịch vụ được expose cho khách hàng (customers). Điều này nhấn mạnh nhu cầu về môi trường testing an toàn, không ảnh hưởng đến production traffic, tận dụng tính năng revisions và traffic splitting của Cloud Run. Cloud Run cho phép deploy nhiều revisions (phiên bản) song song, quản lý traffic qua tags hoặc percentages, giúp test riêng biệt mà không gián đoạn dịch vụ chính (theo tài liệu chính thức Google Cloud Run cập nhật đến 2026).

📍 Đáp án đúng:
Run the gcloud run deploy booking-engine --no-traffic --tag dev command. Use the https://dev--booking-engine-abcdef.a.run.app URL for testing.
Lý do lựa chọn: Lệnh gcloud run deploy với --no-traffic sẽ deploy revision mới mà không phân bổ bất kỳ traffic nào cho nó (traffic vẫn giữ nguyên ở revision hiện tại/production). Kết hợp --tag dev tạo một tag riêng biệt cho revision mới, cho phép truy cập qua URL dành riêng: https://dev--booking-engine-abcdef.a.run.app. Developers có thể test latest revision an toàn tại đây mà không ảnh hưởng đến khách hàng. Đây là best practice cho staging/testing trong Cloud Run (xem docs: Cloud Run Traffic Management).

🛠️ Giải thích chi tiết tất cả các phương án

✅ Phương án ĐÚNG (như đã phân tích ở trên):
Run the gcloud run deploy booking-engine --no-traffic --tag dev command. Use the https://dev--booking-engine-abcdef.a.run.app URL for testing.
→ Hoàn hảo cho testing isolated, không route traffic production sang revision mới.

❌ Phương án SAI 1:
Run the gcloud run services update-traffic booking-engine --to-revisions LATEST=1 command. Use the https://booking-engine-abcdef.a.run.app URL for testing.
→ Lệnh update-traffic --to-revisions LATEST=1 sẽ route 100% traffic đến latest revision, ngay lập tức expose cho tất cả khách hàng, vi phạm yêu cầu "trước khi expose to customers". Không tạo môi trường testing riêng biệt.

❌ Phương án SAI 2:
Pass the curl –H “Authorization:Bearer $(gcloud auth print-identity-token)” auth token. Use the https://booking-engine-abcdef.a.run.app URL to test privately.
→ Dịch vụ đã public (không yêu cầu auth), thêm auth token chỉ dùng cho invoke authenticated trên private services. Sử dụng URL production sẽ test trên revision đang live, có nguy cơ ảnh hưởng khách hàng, không isolate latest revision.

❌ Phương án SAI 3:
Grant the roles/run.invoker role to the developers testing the booking-engine service. Use the https://booking-engine-abcdef.private.run.app URL for testing.
→ Role run.invoker chỉ cho phép invoke private services (yêu cầu IAM auth). Dịch vụ gốc là public (URL .a.run.app), không có .private.run.app. Việc grant role không tạo revision testing riêng, và URL sai không tồn tại.

📘 Tài liệu tham khảo (cập nhật mới nhất đến 2026)

Câu 178
You are configuring connectivity across Google Kubernetes Engine (GKE) clusters in different VPCs. You notice that the nodes in Cluster A are unable to access the nodes in Cluster B. You suspect that the workload access issue is due to the network configuration. You need to troubleshoot the issue but do not have execute access to workloads and nodes. You want to identify the layer at which the network connectivity is broken. What should you do?
  1. A Install a toolbox container on the node in Cluster Confirm that the routes to Cluster B are configured appropriately.
  2. B Use Network Connectivity Center to perform a Connectivity Test from Cluster A to Cluster B.
  3. C Use a debug container to run the traceroute command from Cluster A to Cluster B and from Cluster B to Cluster A. Identify the common failure point.
  4. D Enable VPC Flow Logs in both VPCs, and monitor packet drops.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này xoay quanh tình huống khắc phục sự cố kết nối mạng giữa hai cụm Google Kubernetes Engine (GKE) nằm ở các VPC khác nhau trên Google Cloud Platform (GCP). Cụ thể:

  • Tình huống: Các node trong Cluster A không thể truy cập các node trong Cluster B. Vấn đề được nghi ngờ do cấu hình mạng (network configuration).
  • Ràng buộc quan trọng: Bạn không có quyền execute (thực thi lệnh) trên workloads hoặc nodes, nghĩa là không thể SSH vào node, chạy container debug, hoặc thực hiện các lệnh trực tiếp trên chúng.
  • Mục tiêu: Xác định lớp (layer) mạng nơi kết nối bị đứt (ví dụ: L3 routing, firewall rules, L4 ports, v.v.), mà không cần truy cập trực tiếp vào tài nguyên workload/node.
  • Ngữ cảnh cập nhật 2026: GCP đã nâng cấp Network Connectivity Center (NCC) với Connectivity Tests mạnh mẽ hơn, hỗ trợ kiểm tra kết nối end-to-end từ source đến destination mà không yêu cầu agent trên node/workload. Đây là công cụ tiêu chuẩn cho troubleshooting đa-VPC/VPN/Hub-and-Spoke (theo tài liệu GCP mới nhất tại Cloud Network Connectivity documentation).

🛠️ Mẹo DevOps: Trong môi trường GKE multi-cluster/multi-VPC, kết nối thường bị block bởi VPC Peering, Shared VPC, Firewall Rules, hoặc Cloud Router misconfig. Connectivity Test giúp pinpoint layer nhanh chóng.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Network Connectivity Center to perform a Connectivity Test from Cluster A to Cluster B.

Lý do chi tiết:

  • Network Connectivity Center (NCC) là hub trung tâm quản lý kết nối mạng GCP (VPC peering, VPN, Interconnect). Tính năng Connectivity Tests cho phép chạy test từ source endpoint (như GKE cluster subnet CIDR của Cluster A) đến destination endpoint (Cluster B subnet CIDR) mà không cần agent, exec access, hoặc traffic thực tế.
  • Test sẽ phân tích từng layer: L3 (reachability/routing), L4 (ports/protocols), firewall drops, MTU issues, và trả về báo cáo chi tiết với minimum RTT path và failure point chính xác (ví dụ: "Blocked at VPC firewall rule XYZ").
  • Phù hợp hoàn hảo với ràng buộc "no execute access" vì chạy từ control plane NCC, không chạm vào node/workload.
  • Cập nhật 2026: NCC hỗ trợ GKE Autopilot/Standard clusters tự động, tích hợp với Reachability API cho automation (xem Connectivity Tests guide).

📘 Nguồn tham khảo:

🔍 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm lý do dựa trên best practices GCP DevOps.

  • Install a toolbox container on the node in Cluster Confirm that the routes to Cluster B are configured appropriately.
    ❌ Sai vì: Phương án này yêu cầu cài đặt toolbox container (như busybox/netshoot) trực tiếp trên node của Cluster A, đòi hỏi quyền exec vào node/pod (kubectl debug hoặc node SSH). Nhưng câu hỏi rõ ràng nêu "do not have execute access to workloads and nodes", nên không khả thi. Ngoài ra, chỉ check routes thủ công không đủ để identify layer cụ thể (có thể fail ở firewall/L4), và toolbox không scale cho multi-VPC troubleshooting.

  • Use Network Connectivity Center to perform a Connectivity Test from Cluster A to Cluster B.
    ✅ Đúng vì: Như giải thích ở phần trên, đây là cách tối ưu, không xâm lấn sử dụng NCC control plane để test end-to-end, pinpoint layer failure (L3-L7) mà không cần access node/workload. Hỗ trợ GKE CIDR tự động map endpoint, kết quả realtime dashboard.

  • Use a debug container to run the traceroute command from Cluster A to Cluster B and from Cluster B to Cluster A. Identify the common failure point.
    ❌ Sai vì: Yêu cầu chạy debug container (kubectl debug hoặc ephemeral container) trên pod/node ở cả hai cluster, thực thi lệnh traceroute – trực tiếp vi phạm ràng buộc "no execute access". Traceroute chỉ tốt cho L3 hops nhưng không detect firewall drops/L4 issues chính xác bằng NCC, và cần traffic hai chiều thủ công, dễ miss asymmetric failures.

  • Enable VPC Flow Logs in both VPCs, and monitor packet drops.
    ❌ Sai vì: VPC Flow Logs ghi log traffic/drops ở VPC level (L3-L4), hữu ích nhưng không identify layer cụ thể realtime (phải generate traffic trước, chờ log vài phút, rồi filter Cloud Logging – quá chậm và noisy). Không test reachability chủ động từ A đến B, và vẫn cần traffic từ workload (mà bạn không access được). NCC vượt trội hơn vì simulated traffic không cần source thực.

🧰 Khuyến nghị DevOps: Sau khi dùng NCC test, nếu fail ở firewall → check gcloud compute firewall-rules. Scale lên với Network Intelligence Center cho monitoring liên tục! 🚀

Câu 179
You manage an application that runs in Google Kubernetes Engine (GKE) and uses the blue/green deployment methodology. Extracts of the Kubernetes manifests are shown below:

---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: app-green
labels:
  app: my-app
  version: green


---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: app-blue
labels:
  app: my-app
  version: blue


---
apiVersion: v1
kind: Service
metadata:
  name: app-svc
spec:
  selector:
    app: my-app
    version: green


---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: app-ingress
spec:
  defaultBackend:
    service:
      name: app-svc


The Deployment app-green was updated to use the new version of the application. During post-deployment monitoring, you notice that the majority of user requests are failing. You did not observe this behavior in the testing environment. You need to mitigate the incident impact on users and enable the developers to troubleshoot the issue. What should you do?
  1. A Update the Deployment app-blue to use the new version of the application.
  2. B Update the Deployment app-green to use the previous version of the application.
  3. C Change the selector on the Service app-svc to app: my-app.
  4. D Change the selector on the Service app-svc to app: my-app, version: blue.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả một ứng dụng chạy trên Google Kubernetes Engine (GKE) sử dụng phương pháp triển khai blue/green deployment. Đây là kỹ thuật triển khai không gián đoạn, nơi có hai phiên bản Deployment song song:

  • app-blue: Phiên bản cũ (stable).
  • app-green: Phiên bản mới được cập nhật.

Service app-svc hiện tại chỉ định selector app: my-app, version: green, nghĩa là toàn bộ lưu lượng (traffic) từ người dùng được chuyển hướng chỉ đến Deployment app-green qua Ingress app-ingress.

Sau khi cập nhật app-green với phiên bản mới, giám sát post-deployment phát hiện đa số request từ user thất bại (không xảy ra ở môi trường test). Nhiệm vụ:

  • Mitigate impact ngay lập tức (giảm thiểu ảnh hưởng đến user).
  • Enable developers troubleshoot (cho phép dev debug vấn đề trên phiên bản mới).

Mục tiêu là chuyển traffic về phiên bản ổn định (blue) mà vẫn giữ phiên bản mới (green) chạy để phân tích lỗi. Kiến thức dựa trên Kubernetes v1.28+ (cập nhật đến 2026) và GKE best practices cho blue/green (không thay đổi cơ bản từ AWS EKS tương đương, nhưng tập trung GKE).

📘 Tài liệu tham khảo:

✅ Đáp án đúng

Change the selector on the Service app-svc to app: my-app, version: blue.

Lý do lựa chọn:
🛠️ Phương án này chuyển toàn bộ traffic ngay lập tức về Deployment app-blue (phiên bản cũ ổn định), mitigate hoàn toàn impact cho user (không còn failing requests). Đồng thời, Deployment app-green vẫn chạy độc lập với phiên bản mới, cho phép developers logs, metrics, hoặc exec vào pods để troubleshoot mà không ảnh hưởng production traffic. Đây là blue/green rollback chuẩn trên GKE, chỉ chỉnh selector Service (zero-downtime, atomic switch). Không cần redeploy hay scale, phù hợp post-deployment incident.

📋 Giải thích tất cả các phương án

  • Update the Deployment app-blue to use the new version of the application.
    ❌ Sai: Việc cập nhật app-blue sang phiên bản mới sẽ làm cả hai Deployment đều chạy code lỗi, dẫn đến 100% traffic fail. Không mitigate impact mà còn làm tình hình tệ hơn. Không hỗ trợ troubleshoot vì mất luôn phiên bản ổn định. (Phương pháp này giống "big bang" deploy, vi phạm blue/green principle).

  • Update the Deployment app-green to use the previous version of the application.
    ❌ Sai: Đây là rollback trực tiếp Deployment green, mitigate impact tạm thời bằng cách quay về code cũ. Tuy nhiên, không enable troubleshoot vì phiên bản mới bị xóa sạch (pods cũ terminate), dev mất dữ liệu logs/metrics để debug root cause. Không tận dụng lợi thế blue/green (nên ưu tiên switch traffic thay vì thay code).

  • Change the selector on the Service app-svc to app: my-app.
    ❌ Sai: Selector chỉ app: my-app sẽ match cả blue VÀ green pods, traffic được load balance ngẫu nhiên (round-robin theo Kubernetes endpoint slices). Kết quả: ~50% request vẫn fail (từ green), không mitigate đầy đủ impact. Vẫn ảnh hưởng user lớn và khó troubleshoot vì traffic lẫn lộn.

  • Change the selector on the Service app-svc to app: my-app, version: blue.
    ✅ Đúng: Như giải thích trên, switch atomic traffic 100% sang blue (stable), zero-impact cho user. Green pods vẫn live để debug (kubectl logs, port-forward). Hiệu quả cao trên GKE Autopilot/Standard clusters (cập nhật 2026 hỗ trợ selector label matching chính xác).

🧠 Lưu ý DevOps best practice: Sau troubleshoot, scale green xuống 0 nếu cần, hoặc fix và switch lại. Sử dụng GKE Istio/Gateway API cho advanced traffic shifting nếu scale lớn!

Câu 180
You are running a web application deployed to a Compute Engine managed instance group. Ops Agent is installed on all instances. You recently noticed suspicious activity from a specific IP address. You need to configure Cloud Monitoring to view the number of requests from that specific IP address with minimal operational overhead. What should you do?
  1. A Configure the Ops Agent with a logging receiver. Create a logs-based metric.
    B Create a script to scrape the web server log. Export the IP address request metrics to the Cloud Monitoring API.
  2. B Update the application to export the IP address request metrics to the Cloud Monitoring API.
  3. C Configure the Ops Agent with a metrics receiver.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào tình huống bảo mật và giám sát trong Google Cloud Platform (GCP):
Bạn đang chạy một ứng dụng web triển khai trên Compute Engine managed instance group (MIG). Ops Agent đã được cài đặt trên tất cả các instance. Gần đây, bạn phát hiện hoạt động đáng ngờ từ một địa chỉ IP cụ thể, và cần cấu hình Cloud Monitoring để xem số lượng requests từ IP đó với overhead hoạt động tối thiểu (minimal operational overhead).

📌 Mục tiêu chính:

  • Theo dõi metrics (số lượng requests) từ IP nghi ngờ mà không cần thay đổi code ứng dụng lớn, không script thủ công, và tận dụng các công cụ sẵn có của GCP để giảm thiểu công sức vận hành.
  • Ops Agent là agent đa năng (host metrics, logs, processes) được khuyến nghị thay thế cho các agent cũ như Stackdriver Agent (tính đến 2024-2026, Ops Agent là phiên bản mới nhất, hỗ trợ fluentbit cho logs và otelcol cho metrics).
  • Cloud Monitoring cho phép tạo logs-based metrics từ logs để chuyển thành metrics có thể query và alert.

✅ Đáp án đúng: Configure the Ops Agent with a logging receiver. Create a logs-based metric.

Lý do lựa chọn:
🛠️ Phương án này tối ưu nhất vì:

  • Ops Agent với logging receiver (dùng Fluent Bit) thu thập logs từ web server (ví dụ: nginx/apache access logs) mà không cần thay đổi ứng dụng. Logs được gửi đến Cloud Logging.
  • Từ logs, tạo logs-based metric (counter metric) trong Cloud Monitoring để đếm số requests matching IP cụ thể (sử dụng filter như jsonPayload.client_ip = "suspicious_ip").
  • Minimal overhead: Chỉ config YAML của Ops Agent (thêm receiver cho log path), restart agent, và tạo metric qua UI/CLI – không code, không script, scale tự động trên MIG.
  • Hỗ trợ real-time monitoring, alerting, và dashboard với chi phí thấp (logs-based metrics miễn phí ingest nếu từ Ops Agent).

📘 Tài liệu tham khảo:

📋 Giải thích tất cả các phương án

  • Configure the Ops Agent with a logging receiver. Create a logs-based metric.
    ✅ Đúng 🏆: Như giải thích trên, tận dụng logging receiver của Ops Agent để thu thập web logs (access logs chứa IP), sau đó logs-based metric đếm chính xác requests từ IP cụ thể. Overhead thấp, scale với MIG, phù hợp best practice GCP 2026.

  • B Create a script to scrape the web server log. Export the IP address request metrics to the Cloud Monitoring API.
    ❌ Sai 🚫: Phương án này yêu cầu viết script custom để scrape logs thủ công và push metrics qua API – overhead cao (phải maintain script, cron job, handle scale MIG, error-prone). Không tận dụng Ops Agent sẵn có, vi phạm yêu cầu "minimal operational overhead".

  • C Update the application to export the IP address request metrics to the Cloud Monitoring API.
    ❌ Sai 🚫: Cần sửa code ứng dụng (thêm OpenTelemetry hoặc custom exporter) để push metrics – overhead lớn (deploy lại app, test, maintain code). Không phù hợp cho monitoring nhanh suspicious activity, và không dùng logs sẵn có từ web server.

  • D Configure the Ops Agent with a metrics receiver.
    ❌ Sai 🚫: Metrics receiver của Ops Agent chỉ thu thập metrics hệ thống/host sẵn có (CPU, disk, network) hoặc app metrics qua OpenTelemetry – không parse logs để đếm IP requests. Không giải quyết được yêu cầu extract từ web logs, dẫn đến thiếu dữ liệu chính xác.

🧠 Kết luận: Phương án A là best practice cho DevOps Engineer trên GCP, đảm bảo zero-downtime monitoring với công cụ native! 🚀