Ngân hàng đề — Google Cloud Professional Cloud DevOps Engineer
Tìm thấy 269 câu.
- A Deploy the application as a new Cloud Run service.
- B Deploy a new Cloud Run revision with a tag and use the --no-traffic option.
- C Deploy a new Cloud Run revision without a tag and use the --no-traffic option.
- D Deploy the new application version and use the --no-traffic option. Route production traffic to the revision’s URL.
- E Deploy the new application version, and split traffic to the new version.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi
Câu hỏi tập trung vào việc triển khai một phiên bản mới của ứng dụng đang chạy trên Cloud Run (dịch vụ serverless container của Google Cloud). Mục tiêu chính là:
- Sử dụng live production traffic (lưu lượng truy cập thực tế từ sản xuất) để kiểm tra phiên bản mới, đồng thời cho phép đội QA thực hiện manual testing (kiểm tra thủ công).
- Hạn chế tác động tiềm ẩn nếu phiên bản mới có vấn đề (limit potential impact).
- Khả năng rollback (quay về phiên bản cũ) dễ dàng khi cần.
Tình huống cụ thể: Ứng dụng hiện tại đang phục vụ traffic sản xuất. Khi deploy phiên bản mới, cần cách thức cho phép test với traffic thực tế nhưng không ảnh hưởng toàn bộ hệ thống ngay lập tức. Cloud Run quản lý deploy qua revisions (phiên bản con), mỗi deploy tạo revision mới với URL riêng (https://revision-hash.run.app), và traffic có thể được điều hướng linh hoạt mà không downtime. Câu hỏi yêu cầu chọn hai phương án đúng để đạt yêu cầu trên, dựa trên best practices của Cloud Run (cập nhật đến 2026: Cloud Run hỗ trợ traffic splitting, tags cho revisions, và --no-traffic flag để deploy "im lặng").
📘 Tài liệu tham khảo:
- Cloud Run Documentation: Revisions and traffic management (Google Cloud, cập nhật 2025).
- gcloud run deploy flags: --no-traffic (CLI reference, phiên bản mới nhất 2026).
- Best practices for canary deployments on Cloud Run.
✅ Đáp án đúng (Chọn hai)
Hai phương án đúng là:
- Deploy a new Cloud Run revision with a tag and use the --no-traffic option.
- Deploy the new application version and use the --no-traffic option. Route production traffic to the revision’s URL.
Lý do lựa chọn: 🛠️ Sử dụng --no-traffic khi deploy (qua gcloud run deploy --no-traffic) tạo revision mới mà không tự động route traffic sản xuất đến nó, giữ nguyên revision cũ phục vụ production (hạn chế impact).
- Revision mới có URL riêng, cho phép route thủ công production traffic đến URL đó để test live traffic (ví dụ: proxy hoặc direct call từ production env).
- Tag (như
--tag=new-version) giúp identify và quản lý revision dễ dàng (e.g.,service:tag). - Rollback: Chỉ cần set traffic về revision cũ (
gcloud run services update-traffic), zero-downtime. Điều này lý tưởng cho QA manual testing + live traffic test với rủi ro thấp. Hoàn hảo khớp yêu cầu!
🧐 Giải thích tất cả các phương án (Đúng/Sai)
-
❌ Deploy the application as a new Cloud Run service.
Sai vì tạo service mới hoàn toàn riêng biệt (không chia sẻ domain/traffic với service cũ). Không dễ dàng dùng production traffic cho test (phải config DNS/load balancer riêng), khó limit impact và rollback (phải migrate traffic thủ công, phức tạp hơn revisions). Không phù hợp best practices cho incremental testing trên cùng service. -
✅ Deploy a new Cloud Run revision with a tag and use the --no-traffic option.
Đúng như giải thích trên: Deploy revision "im lặng" (--no-traffic), tag giúp label (e.g.,gcloud run deploy --image=new-img --tag=v2 --no-traffic). Test live traffic qua URL revision (https://service-hash-tag.run.app), QA manual dễ dàng, rollback nhanh bằng traffic shift. Hạn chế impact tối đa! -
❌ Deploy a new Cloud Run revision without a tag and use the --no-traffic option.
Sai dù dùng --no-traffic tốt, nhưng thiếu tag làm khó quản lý/identify revision (chỉ có hash dài, không human-readable). Không tiện cho QA manual testing hoặc route traffic chính xác, vi phạm best practices (Google khuyến nghị tag cho canary/blue-green deploys). -
✅ Deploy the new application version and use the --no-traffic option. Route production traffic to the revision’s URL.
Đúng hoàn hảo: --no-traffic deploy "ẩn", sau đó route production traffic thủ công đến URL revision (e.g., via API Gateway, proxy, hoặc direct calls). Limit impact (traffic cũ vẫn chạy), test live + manual QA, rollback bằngupdate-trafficvề revision cũ. Linh hoạt cao! -
❌ Deploy the new application version, and split traffic to the new version.
Sai vì split traffic (e.g.,--update-traffic=50/50) tự động gửi % traffic sản xuất đến version mới ngay lập tức, không limit impact tốt (có thể ảnh hưởng users thực nếu bug). Không ưu tiên manual testing (traffic split là automated canary), khó kiểm soát chính xác như route thủ công. Phù hợp blue-green hơn là yêu cầu ở đây.
- A Notify the team about the lack of error budget and ensure that all their tests are successful so the launch will not further risk the error budget
- B Notify the team that their error budget is used up. Negotiate with the team for a launch freeze or tolerate a slightly worse user experience.
- C Escalate the situation and request additional error budget.
- D Look through other metrics related to the product and find SLOs with remaining error budget. Reallocate the error budgets and allow the feature launch.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi này xoay quanh Site Reliability Engineering (SRE) – một thực hành kỹ thuật đáng tin cậy từ Google, được áp dụng rộng rãi trong DevOps trên các nền tảng đám mây như AWS (tích hợp qua AWS Well-Architected Framework và các công cụ như CloudWatch cho SLO/SLI).
Tình huống cụ thể:
Một dịch vụ của bạn đã vượt quá error budget (ngân sách lỗi) trong rolling window period (khoảng thời gian trượt, thường 28-30 ngày để đo lường SLO – Service Level Objective). Nhóm sản phẩm sắp ra mắt tính năng mới, có nguy cơ làm tăng lỗi thêm. Bạn cần tuân thủ SRE practices: Error budget là phần "dư địa" cho lỗi (ví dụ: nếu SLO là 99.9%, error budget là 0.1% thời gian downtime). Khi hết budget, ưu tiên reliability (độ tin cậy) hơn velocity (tốc độ phát triển).
Mục tiêu: Quyết định hành động đúng để cân bằng giữa phát triển và ổn định hệ thống, tránh rủi ro SLO bị vi phạm.
(Kiến thức cập nhật: SRE principles vẫn giữ nguyên đến 2026, AWS tích hợp qua Amazon Managed Grafana và CloudWatch SLO cho error budget tracking – theo AWS re:Post 2025 updates).
✅ Đáp án đúng
Notify the team that their error budget is used up. Negotiate with the team for a launch freeze or tolerate a slightly worse user experience.
Lý do chọn đáp án này (theo SRE best practices):
🛠️ Khi error budget cạn kiệt, SRE yêu cầu thông báo ngay cho dev/product team và đàm phán thay vì cho phép launch tự do. Các lựa chọn:
- Launch freeze (tạm dừng ra mắt) để phục hồi reliability.
- Hoặc tolerate slightly worse UX (chấp nhận trải nghiệm người dùng kém hơn nhẹ, ví dụ: enable feature flags để rollback nhanh).
Điều này thúc đẩy shared ownership giữa SRE và dev team, ưu tiên user experience lâu dài. Không có "magic fix" – phải trade-off rõ ràng.
📘 Nguồn tham khảo:
- Google SRE Book (Chapter 5: Error Budgets): sre.google/sre-book/error-budgets.
- AWS Well-Architected Reliability Pillar (2025): docs.aws.amazon.com/wellarchitected/latest/reliability-pillar – khuyến nghị error budget cho SLO.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn một, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên SRE principles (error budget policy: "No launches when budget exhausted").
-
Notify the team about the lack of error budget and ensure that all their tests are successful so the launch will not further risk the error budget
❌ Sai: Phương án này giả định tests thành công = an toàn launch, nhưng SRE không dựa vào tests (unit/integration) vì chúng không dự đoán production failures. Error budget đã hết → không launch, dù tests pass. Tests chỉ là hygiene, không thay thế SLO monitoring. Rủi ro: Launch có thể burn thêm budget thực tế. -
Notify the team that their error budget is used up. Negotiate with the team for a launch freeze or tolerate a slightly worse user experience
✅ Đúng: Như giải thích trên. Notify + negotiate là core SRE: Tạo accountability, khuyến khích team tự quyết định trade-off (freeze hoặc degrade gracefully). Phù hợp AWS practices với feature flags (via CodeDeploy/AppConfig). -
Escalate the situation and request additional error budget
❌ Sai: Không tồn tại "additional error budget" trong SRE – budget là fixed dựa trên SLO (ví dụ: 43.2 phút/tháng cho 99.9%). Escalate chỉ dùng cho outages lớn, không phải "xin thêm budget". Làm vậy vi phạm nguyên tắc: Budget phải strict để enforce reliability. -
Look through other metrics related to the product and find SLOs with remaining error budget. Reallocate the error budgets and allow the feature launch
❌ Sai: Error budgets không thể reallocate giữa các SLO/services khác nhau – mỗi service có SLO độc lập (tree structure có thể liên kết, nhưng không shuffle budget). Làm vậy tạo dependency ảo, vi phạm isolation principle. Launch vẫn rủi ro nếu service chính đã hết budget.
Kết luận 🚀: Áp dụng SRE giúp hệ thống AWS scalable & reliable. Nếu implement trên AWS, dùng CloudWatch SLO + X-Ray để track error budget real-time (cập nhật 2026: AI-powered SLO forecasting).
- A Encourage new employees to conduct postmortems to team through practice.
- B Create a designated team that is responsible for conducting all postmortems.
- C Encourage your senior leadership to acknowledge and participate in postmortems.
- D Ensure that writing effective postmortems is a rewarded and celebrated practice.
- E Provide your organization with a forum to critique previous postmortems.
Xem giải thích
🔍 Giải thích nội dung câu hỏi
🧩 Câu hỏi này thuộc chủ đề DevOps và SRE (Site Reliability Engineering) trên AWS, tập trung vào việc giới thiệu quy trình postmortem (phân tích sự cố sau khi xảy ra) vào tổ chức một cách hiệu quả để được đón nhận tốt. Postmortem là thực hành blameless (không đổ lỗi), nhằm học hỏi từ sự cố để cải thiện hệ thống, theo AWS Well-Architected Framework - Operational Excellence Pillar.
Câu hỏi yêu cầu chọn hai hành động đúng (Choose two) để đảm bảo postmortem được chấp nhận rộng rãi, nhấn mạnh vào văn hóa tổ chức, lãnh đạo tham gia và khuyến khích học hỏi. Điều này giúp xây dựng sự tin tưởng, tránh sợ hãi khi chia sẻ lỗi, phù hợp với best practices AWS cập nhật đến 2026 (Reliability Pillar khuyến nghị postmortem là phần cốt lõi của incident management).
✅ Đáp án đúng và lý do lựa chọn
Hai đáp án đúng là:
Encourage your senior leadership to acknowledge and participate in postmortems.
Ensure that writing effective postmortems is a rewarded and celebrated practice.
🛠️ Lý do chọn:
- Theo AWS, để postmortem được đón nhận, cần lãnh đạo cấp cao dẫn dắt bằng ví dụ (lead by example) để tạo văn hóa tin cậy và khuyến khích toàn tổ chức tham gia. Đồng thời, thưởng và kỷ niệm postmortem hiệu quả giúp biến nó thành thói quen tích cực, thúc đẩy học hỏi liên tục (continuous improvement). Những hành động này giải quyết rào cản văn hóa, đảm bảo postmortem không bị coi là "trừng phạt" mà là cơ hội phát triển.
📋 Phân tích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Tôi sử dụng ✅ cho đúng, ❌ cho sai, dựa trên best practices AWS mới nhất (2026).
-
Encourage new employees to conduct postmortems to team through practice.
❌ Sai: Việc giao postmortem cho nhân viên mới sẽ làm họ thiếu kinh nghiệm, dẫn đến phân tích kém chất lượng và giảm uy tín quy trình. AWS khuyến nghị postmortem do đội ngũ có kinh nghiệm dẫn dắt để đảm bảo tính khách quan và hiệu quả, tránh làm nản lòng người mới. -
Create a designated team that is responsible for conducting all postmortems.
❌ Sai: Tạo đội ngũ chuyên trách sẽ tạo silo (phân cách), làm giảm sự tham gia của toàn tổ chức và trách nhiệm cá nhân. AWS nhấn mạnh mọi đội ngũ phải sở hữu postmortem của sự cố mình gặp, thúc đẩy trách nhiệm phân tán (decentralized ownership) theo Operational Excellence. -
Encourage your senior leadership to acknowledge and participate in postmortems.
✅ Đúng: Lãnh đạo cấp cao tham gia và công nhận giúp xây dựng văn hóa blameless, truyền cảm hứng cho nhân viên. AWS Reliability Pillar (2026) coi đây là yếu tố then chốt để postmortem được đón nhận, vì "lãnh đạo làm gương" tạo động lực từ trên xuống. -
Ensure that writing effective postmortems is a rewarded and celebrated practice.
✅ Đúng: Thưởng và kỷ niệm postmortem hiệu quả củng cố hành vi tích cực, biến nó thành văn hóa học hỏi. AWS khuyến nghị tích hợp vào OKR/KPI, giúp tổ chức coi postmortem là thành tựu, không phải nghĩa vụ. -
Provide your organization with a forum to critique previous postmortems.
❌ Sai: Diễn đàn phê phán sẽ tạo sợ hãi, làm nhân viên ngại chia sẻ postmortem tương lai (fear of blame). AWS nhấn mạnh postmortem phải blameless và confidential, tập trung học hỏi chứ không critique công khai để tránh cản trở quy trình.
📘 Tài liệu tham khảo
- AWS Well-Architected Framework - Operational Excellence Pillar (cập nhật 2026): docs.aws.amazon.com/wellarchitected/latest/operational-excellence-pillar – Phần "Respond to incidents" nhấn mạnh leadership và rewards cho postmortems.
- AWS Reliability Pillar (2026): docs.aws.amazon.com/wellarchitected/latest/reliability-pillar – Hướng dẫn blameless postmortems và incident response.
- Google SRE Book (ảnh hưởng AWS practices): Chapter 6 "Tracking Outages" – Áp dụng tương tự cho AWS DevOps.
Những nguồn này xác nhận best practices không thay đổi cơ bản đến 2026. 🚀
- A Set up a GitHub action to trigger Cloud Build when there is a parameter change. In Cloud Build, run a gcloud CLI command to apply the change.
- B When there is a change in GitHub. use a web hook to send a request to Anthos Service Mesh, and apply the change.
- C Configure Anthos Config Management with the GitHub repository. When there is a change in the repository, use Anthos Config Management to apply the change.
- D Configure Config Connector with the GitHub repository. When there is a change in the repository, use Config Connector to apply the change.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc thực thi (enforce) các constraint templates trên nhiều Google Kubernetes Engine (GKE) clusters. Các constraint này bao gồm policy parameters (tham số chính sách), chẳng hạn như việc hạn chế Kubernetes API. Yêu cầu chính là:
- Lưu trữ các policy parameters trong một GitHub repository.
- Tự động áp dụng (automatically applied) các thay đổi khi có cập nhật trong repository.
📌 Mục tiêu chính: Đảm bảo tính nhất quán (consistency) và tự động hóa (automation) trong việc quản lý chính sách (policies) trên các GKE clusters, sử dụng GitOps approach (quản lý config qua Git). Đây là tình huống điển hình trong Anthos platform của Google Cloud, nơi Policy Controller (dựa trên Open Policy Agent - OPA Gatekeeper) được sử dụng để enforce constraints. Các constraint templates cần sync từ Git repo một cách liên tục (continuous sync).
🛠️ Bối cảnh kiến thức cập nhật (đến 2026): Theo tài liệu AWS không liên quan ở đây (câu hỏi thuộc Google Cloud), mà dựa trên Anthos Policy Controller và Config Management phiên bản mới nhất (Anthos 1.15+), hỗ trợ Git sync cho constraints với referential integrity checks và auto-remediation.
📘 Tài liệu tham khảo:
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Configure Anthos Config Management with the GitHub repository. When there is a change in the repository, use Anthos Config Management to apply the change.
Lý do:
- Anthos Config Management (trước đây gọi là Config Sync) là giải pháp chính thức của Google Cloud để sync cấu hình từ Git repository (hỗ trợ GitHub) vào GKE clusters một cách tự động và liên tục (hierarchical hoặc fleet-wide).
- Nó tích hợp trực tiếp với Policy Controller, cho phép enforce constraint templates và parameters (như restrict Kubernetes API) trên nhiều clusters.
- Khi có thay đổi trong repo, Config Management sẽ tự động detect và apply qua GitOps pull-based model, đảm bảo tính nhất quán mà không cần webhook hay CI/CD thủ công.
- Hoàn hảo cho scale multi-cluster trong Anthos. ✅
📋 Giải thích tất cả các phương án (đúng/sai)
-
[SAI] Set up a GitHub action to trigger Cloud Build when there is a parameter change. In Cloud Build, run a gcloud CLI command to apply the change.
❌ Sai vì: GitHub Actions + Cloud Build chỉ là CI/CD pipeline push-based, không phải GitOps chuẩn cho config management.gcloud CLIkhông hỗ trợ enforce constraints tự động trên GKE (cầnanthosclihoặc Config Management). Không đảm bảo sync liên tục, dễ lỗi thủ công, và không tích hợp Policy Controller. Không phù hợp cho multi-cluster. -
[SAI] When there is a change in GitHub. use a web hook to send a request to Anthos Service Mesh, and apply the change.
❌ Sai vì: Anthos Service Mesh (nay là GKE Enterprise với Istio) chỉ quản lý traffic management, security, observability (mTLS, circuit breaking), không liên quan đến config/policy enforcement. Webhook chỉ trigger events, không apply constraints. Sai hoàn toàn về công cụ. -
[ĐÚNG] Configure Anthos Config Management with the GitHub repository. When there is a change in the repository, use Anthos Config Management to apply the change.
✅ Đúng vì: Như đã giải thích ở trên, đây là cách chuẩn và tự động nhất, hỗ trợ Git sync cho Policy Controller constraints/parameters trên GKE/Anthos clusters. -
[SAI] Configure Config Connector with the GitHub repository. When there is a change in the repository, use Config Connector to apply the change.
❌ Sai vì: Config Connector dùng để quản lý GCP resources qua Kubernetes CRDs (như IAM, Cloud Storage từ YAML K8s), không sync Git config cho policies/constraints. Nó không hỗ trợ Policy Controller hay GitOps cho Kubernetes resources/policies. Sai mục đích sử dụng. 🛠️
-
A
1. Communicate your intent to the incident team.
2. Perform a load analysis to determine if the remaining nodes can handle the increase in traffic offloaded from the removed node, and scale appropriately.
3. When any new nodes report healthy, drain traffic from the unhealthy node, and remove the unhealthy node from service. -
B
1. Communicate your intent to the incident team.
2. Add a new node to the pool, and wait for the new node to report as healthy.
3. When traffic is being served on the new node, drain traffic from the unhealthy node, and remove the old node from service. -
C
1. Drain traffic from the unhealthy node and remove the node from service.
2. Monitor traffic to ensure that the error is resolved and that the other nodes in the pool are handling the traffic appropriately.
3. Scale the pool as necessary to handle the new load.
4. Communicate your actions to the incident team. -
D
1. Drain traffic from the unhealthy node and remove the old node from service.
2. Add a new node to the pool, wait for the new node to report as healthy, and then serve traffic to the new node.
3. Monitor traffic to ensure that the pool is healthy and is handling traffic appropriately.
4. Communicate your actions to the incident team.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả tình huống bạn là Operations Lead (Trưởng nhóm vận hành) đang xử lý một incident (sự cố) với một dịch vụ thường chạy ở mức 70% capacity (tải khoảng 70%). Bạn phát hiện một node (máy chủ) đang trả về 5xx errors (lỗi server-side) cho tất cả các requests (yêu cầu), đồng thời có tăng đột biến các support cases từ khách hàng. Nhiệm vụ là loại bỏ node lỗi khỏi load balancer pool (nhóm máy chủ cân bằng tải) để isolate và investigate (cách ly và điều tra), đồng thời giảm thiểu tác động đến người dùng. Bạn phải tuân thủ Google-recommended practices (các thực hành khuyến nghị của Google) trong quản lý incident, dựa trên nguyên tắc SRE (Site Reliability Engineering) của Google Cloud.
🛠️ Mục tiêu chính: Ưu tiên giao tiếp rõ ràng, phân tích tải trước khi hành động, tránh làm gián đoạn dịch vụ, và scale (mở rộng) an toàn. Theo tài liệu SRE mới nhất (cập nhật đến 2026 từ Google Cloud), quy trình incident management nhấn mạnh ECSC (Event, Communicate, Stabilize, Clean-up) và capacity planning trước khi remove node để tránh overload các node còn lại.
📘 Nguồn tham khảo:
- Google SRE Workbook (2023-2026 editions): Chapter on Incident Management.
- Google Cloud Documentation: "Troubleshoot incidents" và "Autoscaling best practices" (cloud.google.com/sre).
- Google Cloud Load Balancing docs: Health checks và draining traffic (cập nhật 2025).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng là phương án đầu tiên (được đánh dấu [ĐÚNG]).
Lý do: Phương án này tuân thủ nghiêm ngặt Google SRE practices cho incident management:
- Bước 1: Giao tiếp intent trước để phối hợp team, tránh hành động bất ngờ (theo nguyên tắc "Communicate early and often").
- Bước 2: Phân tích tải (load analysis) để đảm bảo các node còn lại chịu nổi traffic offload, rồi scale nếu cần – tránh overload vì service đang 70% capacity, remove 1 node có thể đẩy lên cao.
- Bước 3: Chờ node mới healthy rồi drain traffic từ node lỗi và remove – an toàn, graceful, giảm thiểu downtime. Điều này giảm impact tối đa, phù hợp với capacity-aware scaling trong Google Cloud (như Managed Instance Groups với health checks).
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng phương án, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể dựa trên Google-recommended practices (không phải AWS, dù câu hỏi có yếu tố chung).
-
✅ Phương án ĐÚNG:
- Communicate your intent to the incident team.
- Perform a load analysis to determine if the remaining nodes can handle the increase in traffic offloaded from the removed node, and scale appropriately.
- When any new nodes report healthy, drain traffic from the unhealthy node, and remove the unhealthy node from service.
Giải thích: Hoàn hảo theo SRE! 🏆 Giao tiếp đầu tiên đảm bảo team đồng bộ. Load analysis + scale trước khi remove tránh overload (critical vì 70% capacity). Drain sau khi có node mới healthy là graceful removal (sử dụng connection draining trong Google Cloud Load Balancer). Giảm MTTR (Mean Time to Recovery) và tuân thủ "stabilize first".
-
❌ Phương án SAI:
- Communicate your intent to the incident team.
- Add a new node to the pool, and wait for the new node to report as healthy.
- When traffic is being served on the new node, drain traffic from the unhealthy node, and remove the old node from service.
Giải thích: Gần đúng nhưng thiếu load analysis! 😕 Bước 2 chỉ add node mà không kiểm tra capacity của pool hiện tại (có thể overload tạm thời). Google SRE yêu cầu phân tích tải trước để scale proactive, không chỉ reactive add node. Có thể gây spike latency nếu traffic tăng đột ngột.
-
❌ Phương án SAI:
- Drain traffic from the unhealthy node and remove the node from service.
- Monitor traffic to ensure that the error is resolved and that the other nodes in the pool are handling the traffic appropriately.
- Scale the pool as necessary to handle the new load.
- Communicate your actions to the incident team.
Giải thích: Drain và remove NGAY LẬP TỨC mà không giao tiếp trước! 🚫 Vi phạm nguyên tắc SRE "Communicate intent BEFORE action" (dẫn đến confusion trong incident team). Scale ở bước 3 là muộn (reactive), không dự đoán capacity như khuyến nghị. Communicate ở cuối là sai thứ tự – phải đầu tiên!
-
❌ Phương án SAI:
- Drain traffic from the unhealthy node and remove the old node from service.
- Add a new node to the pool, wait for the new node to report as healthy, and then serve traffic to the new node.
- Monitor traffic to ensure that the pool is healthy and is handling traffic appropriately.
- Communicate your actions to the incident team.
Giải thích: Tệ nhất: Drain/remove trước giao tiếp và không load analysis! ⚠️ Gây rủi ro overload cao (service 70% → remove đột ngột). Add node sau là reactive, không scale upfront. Communicate cuối cùng vi phạm "early communication". Không graceful theo Google Cloud health checks.
🧠 Kết luận: Luôn ưu tiên Communicate → Analyze Capacity → Drain Gracefully trong Google Cloud incidents để đạt SLOs (Service Level Objectives). Nếu áp dụng thực tế, dùng Cloud Monitoring cho load analysis và MIGs (Managed Instance Groups) cho auto-scaling! 🚀
- A Create an attestation for the builds that pass the load test by requiring the lead quality assurance engineer to sign the attestation by using their personal private key.
- B Create an attestation for the builds that pass the load test by using a private key stored in Cloud Key Management Service (Cloud KMS) with a service account JSON key stored as a Kubernetes Secret.
- C Create an attestation for the builds that pass the load test by using a private key stored in Cloud Key Management Service (Cloud KMS) authenticated through Workload Identity.
- D Create an attestation for the builds that pass the load test by requiring the lead quality assurance engineer to sign the attestation by using a key stored in Cloud Key Management Service (Cloud KMS).
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc cấu hình pipeline CI/CD native trên Google Cloud để tự động hóa quy trình kiểm tra tải (load test) cho các build trong môi trường pre-production GKE trước khi promote lên production GKE. 🎯 Mục tiêu là chỉ cho phép deploy các build đã pass load test bằng cách sử dụng Binary Authorization (một tính năng của GKE giúp xác thực và kiểm soát việc deploy image container dựa trên policy và attestation).
Cụ thể:
- Binary Authorization yêu cầu các container image phải có attestation (chứng nhận ký số) hợp lệ trước khi deploy vào cluster.
- Quy trình: Load test pass → Tạo attestation → Policy Binary Authorization kiểm tra và approve deploy vào prod.
- Yêu cầu tuân thủ Google-recommended practices (thực hành tốt nhất của Google), nhấn mạnh vào bảo mật, tự động hóa và tránh các phương pháp thủ công hoặc kém an toàn như key cá nhân hoặc JSON key trong Secret.
Câu hỏi kiểm tra kiến thức về cách tạo attestation an toàn sử dụng Cloud KMS (dịch vụ quản lý khóa mã hóa) kết hợp với cơ chế xác thực hiện đại. 📘 (Kiến thức dựa trên tài liệu Google Cloud cập nhật đến 2026: Binary Authorization v1beta1, Workload Identity Federation được ưu tiên từ 2023+).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create an attestation for the builds that pass the load test by using a private key stored in Cloud Key Management Service (Cloud KMS) authenticated through Workload Identity.
Lý do:
- 🛠️ Đây là thực hành được Google khuyến nghị chính thức cho việc ký attestation trong Binary Authorization. Sử dụng Cloud KMS lưu trữ private key đảm bảo khóa được quản lý tập trung, mã hóa và kiểm soát truy cập chi tiết (IAM).
- Workload Identity (thay thế cho service account JSON key) cho phép workload (như Cloud Build hoặc Cloud Run) xác thực trực tiếp với KMS mà không cần lưu trữ key file, giảm rủi ro lộ key. Nó sử dụng OIDC federation, tích hợp seamless với GKE và CI/CD pipelines.
- Tự động hóa hoàn toàn: Load test pass → Service account với Workload Identity ký attestation → Binary Authorization policy verify và deploy tự động vào prod GKE.
- ✅ Phù hợp với nguyên tắc least privilege và zero-trust security trong Google Cloud (cập nhật 2026: Workload Identity là mandatory cho hầu hết workload mới).
Nguồn tham khảo:
- 📘 Binary Authorization Overview (Google Cloud Docs, 2026).
- 📘 Attestations with Cloud KMS and Workload Identity và Workload Identity Best Practices.
❌ Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên tính an toàn, tự động hóa và tuân thủ Google-recommended practices:
-
[SAI] Create an attestation for the builds that pass the load test by requiring the lead quality assurance engineer to sign the attestation by using their personal private key.
❌ Sai vì: Phương án này thủ công và không scalable (yêu cầu lead QA ký thủ công bằng private key cá nhân). Không khuyến nghị vì personal key dễ bị lộ, khó quản lý quyền truy cập, và vi phạm nguyên tắc tự động hóa CI/CD. Google cấm dùng personal keys trong production workflows để tránh single point of failure. -
[SAI] Create an attestation for the builds that pass the load test by using a private key stored in Cloud Key Management Service (Cloud KMS) with a service account JSON key stored as a Kubernetes Secret.
❌ Sai vì: Mặc dù dùng Cloud KMS là tốt, nhưng lưu service account JSON key trong Kubernetes Secret là rủi ro bảo mật cao (Secret có thể bị lộ qua misconfig hoặc attack). Google đã deprecated JSON keys từ 2022 và khuyến nghị chuyển hoàn toàn sang Workload Identity (cập nhật 2026: JSON keys chỉ dùng cho legacy). -
[ĐÚNG] Create an attestation for the builds that pass the load test by using a private key stored in Cloud Key Management Service (Cloud KMS) authenticated through Workload Identity.
✅ Đúng vì: Như đã giải thích ở trên – kết hợp hoàn hảo KMS + Workload Identity cho ký số tự động, an toàn và không cần quản lý key file. Đây là Google-recommended pattern cho CI/CD với Binary Authorization trong GKE. -
[SAI] Create an attestation for the builds that pass the load test by requiring the lead quality assurance engineer to sign the attestation by using a key stored in Cloud Key Management Service (Cloud KMS).
❌ Sai vì: Vẫn thủ công (lead QA phải ký), dù dùng KMS tốt hơn personal key. Không tự động hóa pipeline, gây bottleneck và không theo best practices (Google ưu tiên service accounts + Workload Identity cho automation, không phải manual intervention).
🧠 Tóm tắt key takeaway: Binary Authorization + Cloud KMS + Workload Identity là bộ ba "vàng" cho secure CI/CD trên GKE. Tránh manual signing và JSON keys để đạt zero-trust! 🚀
- A Store the password in Secret Manager and send the secret to the application by using environment variables.
- B Store the password in Secret Manager and mount the secret as a volume within the application.
- C Use Cloud Build to add your password into the application container at build time. Ensure that Artifact Registry is secured from public access.
- D Store the password directly in the code. Use Cloud Build to rebuild and deploy the application each time the password changes.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc triển khai một ứng dụng trên Cloud Run (dịch vụ serverless container của Google Cloud Platform - GCP), nơi ứng dụng cần một password để khởi động. Tổ chức yêu cầu xoay vòng (rotate) password mỗi 24 giờ, và ứng dụng phải luôn sử dụng password mới nhất mà không gây downtime (ngừng dịch vụ).
📌 Yêu cầu chính cần giải quyết:
- Lưu trữ password an toàn (sử dụng Secret Manager).
- Cập nhật password tự động mỗi 24 giờ mà không cần redeploy hoặc restart container (đảm bảo zero-downtime).
- Áp dụng kiến thức GCP cập nhật đến năm 2026: Cloud Run hỗ trợ tích hợp Secret Manager với hai cách chính: environment variables (biến môi trường) hoặc volume mount (gắn volume). Volume mount cho phép tự động cập nhật secret mà không restart instance (Cloud Run poll Secret Manager định kỳ ~1 phút và remount volume nếu secret thay đổi).
🛠️ Bối cảnh kỹ thuật: Cloud Run là fully managed, scale-to-zero, nên cần giải pháp động để xoay secret mà không gián đoạn traffic.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Store the password in Secret Manager and mount the secret as a volume within the application.
Lý do chi tiết:
- Secret Manager lưu trữ password an toàn, hỗ trợ rotation tự động (qua Cloud Functions hoặc scheduler mỗi 24h).
- Mount as volume: Cloud Run tự động poll Secret Manager (mỗi ~60 giây), detect thay đổi và remount volume mới cho tất cả instances mà KHÔNG restart container hoặc gây downtime. Ứng dụng đọc file từ volume
/secrets/passwordđể lấy password mới ngay lập tức. - Đảm bảo zero-downtime: Phù hợp yêu cầu, vì không cần redeploy.
- Theo docs GCP 2026: Đây là best practice cho secrets động trong Cloud Run (revision giữ nguyên, chỉ secret update).
📘 Tài liệu tham khảo:
- Cloud Run Secrets Configuration (GCP Docs, cập nhật 2025-2026).
- Secret Manager Rotation.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm giải thích chi tiết bằng tiếng Việt:
-
❌ Store the password in Secret Manager and send the secret to the application by using environment variables.
Phương án này sai vì: Mặc dù Secret Manager an toàn và hỗ trợ rotation, nhưng env vars chỉ inject giá trị tại thời điểm deploy. Khi password rotate (mỗi 24h), Cloud Run KHÔNG tự cập nhật env vars – cần redeploy revision mới để lấy secret mới, gây downtime ngắn (traffic shift sang revision mới). Không đáp ứng zero-downtime. -
✅ Store the password in Secret Manager and mount the secret as a volume within the application.
Phương án này đúng vì: Như đã giải thích ở phần đáp án đúng. Cloud Run poll Secret Manager định kỳ, remount volume tự động khi secret thay đổi → ứng dụng đọc file mới mà không restart/downtime. Hoàn hảo cho rotation 24h. -
❌ Use Cloud Build to add your password into the application container at build time. Ensure that Artifact Registry is secured from public access.
Phương án này sai vì: Bake password vào image lúc build (qua Cloud Build) chỉ static, KHÔNG hỗ trợ rotation động. Mỗi 24h phải rebuild + push image mới lên Artifact Registry + redeploy, gây downtime và phức tạp (vi phạm nguyên tắc immutable images). Bảo mật Artifact Registry không giải quyết vấn đề rotation. -
❌ Store the password directly in the code. Use Cloud Build to rebuild and deploy the application each time the password changes.
Phương án này sai vì: Hardcode password vào source code là anti-pattern bảo mật (rủi ro leak qua Git/repo). Mỗi rotation cần rebuild + redeploy thủ công, gây downtime lớn, không scale và vi phạm best practices (secret không nên vào code/image).
🧠 Kết luận nổi bật: Sử dụng volume mount với Secret Manager là giải pháp DevOps-native của GCP cho secrets động, đảm bảo security + reliability + zero-downtime. Nếu implement, dùng gcloud run deploy --update-secrets cho volume! 🚀
- A Install and configure Config Connector in Google Kubernetes Engine (GKE).
- B Configure Cloud Build with a Terraform builder to execute terraform plan and terraform apply commands.
- C Create a Pod resource with a Terraform docker image to execute terraform plan and terraform apply commands.
- D Create a Job resource with a Terraform docker image to execute terraform plan and terraform apply commands.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc triển khai Infrastructure as Code (IaC) trong môi trường Google Kubernetes Engine (GKE), theo phương pháp GitOps. Công ty đang chạy ứng dụng trên GKE, nơi các lập trình viên thường tạo tài nguyên đám mây để hỗ trợ ứng dụng. Mục tiêu là:
- Cho phép developer quản lý hạ tầng như code (IaC).
- Tuân thủ best practices của Google.
- Đảm bảo IaC tự động reconcile định kỳ để tránh configuration drift (sự lệch lạc cấu hình, khi trạng thái thực tế không khớp với code khai báo).
🛠️ Yêu cầu cốt lõi: Cần một giải pháp tích hợp với Kubernetes, hỗ trợ GitOps (pull-based, declarative), và tự động đồng bộ/reconcile liên tục (ví dụ: mỗi vài phút) để phát hiện và sửa drift.
📘 Kiến thức cập nhật (đến 2024-2026): Theo tài liệu Google Cloud mới nhất (Config Connector v2.x, tích hợp Anthos Service Mesh và GKE Enterprise), Config Connector là giải pháp chính thức cho IaC trên GKE, sử dụng Kubernetes Custom Resource Definitions (CRDs) để quản lý GCP resources. Nó hỗ trợ GitOps qua Config Sync hoặc Flux, với reconcile tự động (mặc định 60 giây).
Nguồn tham khảo:
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Install and configure Config Connector in Google Kubernetes Engine (GKE).
Lý do chi tiết 🏆:
- Config Connector biến GCP resources thành Kubernetes CRDs (như
IAMPolicy,CloudSQLInstance), cho phép developer khai báo IaC trực tiếp trong Git repo dưới dạng YAML manifests. - Tích hợp hoàn hảo với GitOps (qua Config Sync hoặc Flux CD), nơi controller tự động pull changes từ Git và reconcile định kỳ (mặc định 60 giây), sửa drift ngay lập tức.
- Đây là Google-recommended practice cho GKE IaC, an toàn, declarative, và native Kubernetes – tránh sử dụng imperative tools như Terraform CLI.
- Hỗ trợ multi-tenancy, RBAC Kubernetes, và audit logs tự động.
📋 Giải thích tất cả các phương án (đúng/sai)
-
Install and configure Config Connector in Google Kubernetes Engine (GKE).
✅ Đúng – Như giải thích ở trên, đây là giải pháp chính thức, hỗ trợ reconcile tự động qua Kubernetes operators, lý tưởng cho GitOps trên GKE. Tránh drift hoàn toàn bằng cách continuously sync với desired state từ Git. -
Configure Cloud Build with a Terraform builder to execute terraform plan and terraform apply commands.
❌ Sai – Cloud Build là CI/CD pipeline (trigger-based, push/pull requests), chạy Terraform imperative (plan/apply một lần). Không hỗ trợ reconcile định kỳ tự động; chỉ chạy khi trigger, dễ gây drift nếu có thay đổi thủ công ngoài pipeline. Không phải GitOps thuần túy và không native GKE. -
Create a Pod resource with a Terraform docker image to execute terraform plan and terraform apply commands.
❌ Sai – Pod là workload long-running hoặc one-off, nhưng chạy Terraform CLI không tự động reconcile. Pod crash/restart không trigger reconcile định kỳ, thiếu GitOps pull-model, dễ drift. Không scalable và không theo best practices (Terraform nên declarative qua CRDs). -
Create a Job resource with a Terraform docker image to execute terraform plan and terraform apply commands.
❌ Sai – Job chạy một lần rồi hoàn thành (one-shot), không reconcile định kỳ. Nếu schedule qua CronJob, vẫn thiếu GitOps integration (không pull từ Git tự động), dễ miss drift giữa các lần chạy. Không an toàn cho production IaC trên GKE.
🎯 Kết luận: Config Connector là lựa chọn tối ưu, giúp developer tự quản lý IaC qua kubectl apply từ Git, với reconcile tự động – đúng tinh thần GitOps và Google best practices! 🚀
-
A
• Cloud Infrastructure (Terraform) repository is shared: different directories are different environments
• GKE Infrastructure (Anthos Config Management Kustomize manifests) repository is shared: different overlay directories are different environments
• Application (app source code) repositories are separated: different branches are different features -
B
• Cloud Infrastructure (Terraform) repository is shared: different directories are different environments
• GKE Infrastructure (Anthos Config Management Kustomize manifests) repositories are separated: different branches are different environments
• Application (app source code) repositories are separated: different branches are different features -
C
• Cloud Infrastructure (Terraform) repository is shared: different branches are different environments
• GKE Infrastructure (Anthos Config Management Kustomize manifests) repository is shared: different overlay directories are different environments
• Application (app source code) repository is shared: different directories are different features -
D
• Cloud Infrastructure (Terraform) repositories are separated: different branches are different environments
• GKE Infrastructure (Anthos Config Management Kustomize manifests) repositories are separated: different overlay directories are different environments
• Application (app source code) repositories are separated: different branches are different
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi này xoay quanh việc thiết kế hệ thống với ba môi trường khác nhau: development (dev), quality assurance (QA) và production (prod) trên Google Cloud Platform (GCP). Mỗi môi trường sẽ được triển khai bằng Terraform để tạo Google Kubernetes Engine (GKE) cluster, sau đó các đội ngũ ứng dụng có thể deploy ứng dụng của họ. Anthos Config Management (nay là phần của Google Cloud's ConfigSync trong Anthos Service Mesh, cập nhật đến 2026) sẽ được sử dụng với các template để deploy tài nguyên cấp infrastructure trong từng GKE cluster. Tất cả người dùng (như operator infrastructure và owner ứng dụng) đều sử dụng GitOps (quản lý deployment qua Git repository).
Mục tiêu chính: Xác định cấu trúc source control repositories (repo Git) tối ưu cho:
- Infrastructure as Code (IaC) bằng Terraform (tạo cloud infra và GKE clusters).
- GKE Infrastructure bằng Anthos Config Management (sử dụng Kustomize manifests để config namespace, RBAC, network policies... trong GKE).
- Application code (source code của app).
Best practices GitOps trên GCP (cập nhật 2026):
- Sử dụng monorepo (repo chia sẻ) cho infra để dễ quản lý chung, tránh drift giữa envs.
- Với Anthos Config Management (ConfigSync), dùng hierarchical structure với base + overlays (directories khác nhau cho từng env) để customize config mà không duplicate code.
- App repos thường tách riêng theo team/feature để hỗ trợ CI/CD độc lập và branching strategy (như GitFlow).
📘 Tài liệu tham khảo:
- Anthos Config Management best practices (Google Cloud Docs, 2026).
- Terraform on GCP for multi-env.
- GitOps with GKE.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng là lựa chọn đầu tiên:
- Cloud Infrastructure (Terraform) repository is shared: different directories are different environments
- GKE Infrastructure (Anthos Config Management Kustomize manifests) repository is shared: different overlay directories are different environments
- Application (app source code) repositories are separated: different branches are different features
Lý do:
- 🛠️ Terraform repo shared với directories khác nhau cho envs: Phù hợp multi-env IaC, dễ quản lý state files (dùng remote backend như Cloud Storage) và promote changes qua directories (dev/qa/prod). Tránh repo riêng dẫn đến khó sync.
- 🛠️ Anthos Config Management repo shared với overlay directories: ConfigSync hỗ trợ Kustomize overlays chuẩn (base/ + overlays/dev|qa|prod/), tự động apply config khác nhau mà không cần branch phức tạp, giảm risk drift.
- 🛠️ App repos separated với branches cho features: Mỗi team/app có repo riêng, dùng branches (feature/dev/main) cho GitOps workflow (ArgoCD hoặc native ConfigSync), hỗ trợ independent deployment.
Cấu trúc này tuân thủ GitOps principles (declarative, observable) và multi-tenancy trên GKE/Anthos.
📋 Giải thích tất cả các phương án (đúng/sai)
-
✅ Phương án ĐÚNG (lựa chọn 1):
- Cloud Infrastructure (Terraform) repository is shared: different directories are different environments
- GKE Infrastructure (Anthos Config Management Kustomize manifests) repository is shared: different overlay directories are different environments
- Application (app source code) repositories are separated: different branches are different features
Giải thích: Như trên, đây là best practice cho scalability và maintainability. Directories/overlays dễ audit/promote, app separate tránh monolith repo lớn.
-
❌ Phương án SAI (lựa chọn 2):
- Cloud Infrastructure (Terraform) repository is shared: different directories are different environments
- GKE Infrastructure (Anthos Config Management Kustomize manifests) repositories are separated: different branches are different environments
- Application (app source code) repositories are separated: different branches are different features
Giải thích: GKE infra repos separated không tối ưu vì Anthos ConfigSync khuyến nghị single repo với overlays để sync hierarchical configs. Repos riêng dễ gây inconsistency giữa envs và phức tạp merge branches.
-
❌ Phương án SAI (lựa chọn 3):
- Cloud Infrastructure (Terraform) repository is shared: different branches are different environments
- GKE Infrastructure (Anthos Config Management Kustomize manifests) repository is shared: different overlay directories are different environments
- Application (app source code) repository is shared: different directories are different features
Giải thích: Terraform dùng branches cho envs kém vì branch-per-env khó maintain long-lived branches (merge hell). App repo shared không phù hợp multi-team, vi phạm separation of concerns; nên separate cho CI/CD độc lập.
-
❌ Phương án SAI (lựa chọn 4):
- Cloud Infrastructure (Terraform) repositories are separated: different branches are different environments
- GKE Infrastructure (Anthos Config Management Kustomize manifests) repositories are separated: different overlay directories are different environments
- Application (app source code) repositories are separated: different branches are different
Giải thích: Tất cả repos separated với branches/overlays gây fragmented management, khó enforce policies chung (như OPA/Gatekeeper). Terraform separated dễ state drift; không leverage monorepo benefits cho infra.
- A Export the service account key and configure the agents to use the key.
- B Update the instance to use the default Compute Engine service account.
- C Add the Logs Writer role to the service account.
- D Enable Private Google Access on the subnet that the instance is in.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi này thuộc lĩnh vực Google Cloud Platform (GCP), cụ thể là cấu hình Cloud Logging cho một ứng dụng mới chạy trên Compute Engine instance có địa chỉ IP công khai (public IP). Instance đã gắn một user-managed service account, và các agent cần thiết (như Ops Agent hoặc Cloud Logging agent) đã được xác nhận đang chạy trên instance. Tuy nhiên, không thấy bất kỳ log entry nào xuất hiện trong Cloud Logging.
Mục tiêu là giải quyết vấn đề theo các thực hành được Google khuyến nghị (Google-recommended practices). Vấn đề cốt lõi nằm ở quyền truy cập (IAM permissions): Service account cần có quyền ghi log vào Cloud Logging, vì các agent sử dụng instance metadata service để xác thực mà không cần key file. Instance có public IP nên không liên quan đến private networking. (Lưu ý: Đây là kiến thức GCP cập nhật đến năm 2026, theo tài liệu chính thức của Google Cloud về Cloud Logging và Compute Engine service accounts).
📘 Tài liệu tham khảo:
- Cloud Logging troubleshooting (Google Cloud Docs, cập nhật 2024-2026).
- Compute Engine logging và IAM roles for Logging (roles/logging.logWriter là role cần thiết).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Add the Logs Writer role to the service account.
🛠️ Lý do: Theo best practices của Google, user-managed service account gắn vào Compute Engine instance cần được cấp role roles/logging.logWriter (Logs Writer) để các agent có thể ghi log vào Cloud Logging. Các agent (như fluentbit trong Ops Agent) sử dụng service account token từ metadata server của instance để authenticate, không cần key file. Thiếu role này sẽ khiến log không được gửi lên, dù agent đang chạy. Đây là giải pháp trực tiếp, an toàn và được khuyến nghị (không khuyến khích dùng default service account vì lý do security).
🔍 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá với emoji tương ứng:
-
❌ [SAI] Export the service account key and configure the agents to use the key.
🧨 Giải thích sai: Không nên export và sử dụng service account key file vì vi phạm nguyên tắc least privilege và tăng rủi ro bảo mật (key có thể bị lộ). Google khuyến nghị sử dụng Workload Identity Federation hoặc metadata service trên Compute Engine để authenticate tự động. Sử dụng key chỉ dành cho trường hợp legacy/off-instance, không phải best practice cho instance-attached SA (theo docs IAM 2026). -
❌ [SAI] Update the instance to use the default Compute Engine service account.
🔒 Giải thích sai: Default Compute Engine service account (<project-number>-compute@developer.gserviceaccount.com) có thể có quyền Logs Writer mặc định, nhưng Google không khuyến nghị chuyển sang dùng default SA vì lý do security (default SA thường có quyền rộng hơn cần thiết). Best practice là giữ user-managed SA và cấp quyền cụ thể (principle of least privilege), tránh thay đổi instance metadata. -
✅ [ĐÚNG] Add the Logs Writer role to the service account.
🎯 Giải thích đúng: Như đã nêu ở phần trên, đây là giải pháp chính xác. Sử dụnggcloud projects add-iam-policy-binding <project> --member="serviceAccount:<sa-email>" --role="roles/logging.logWriter"để cấp quyền. Agent sẽ tự động sử dụng token từ metadata, log sẽ xuất hiện sau vài phút (xác nhận qua Logs Explorer). -
❌ [SAI] Enable Private Google Access on the subnet that the instance is in.
🌐 Giải thích sai: Private Google Access chỉ cần cho instance không có public IP để truy cập Google APIs qua private IP (10.x range). Instance ở đây có public IP, nên đã có thể reach metadata service và Cloud Logging APIs công khai. Bật PGA không giải quyết vấn đề quyền IAM và có thể không ảnh hưởng đến logging agent (theo VPC docs 2026).
💡 Lời khuyên thực hành: Sau khi cấp role, kiểm tra bằng gcloud logging logs list hoặc Logs Explorer. Nếu vẫn lỗi, verify agent config qua sudo systemctl status google-cloud-ops-agent. Áp dụng monitoring với Cloud Monitoring để tránh vấn đề tương tự!