Ngân hàng đề — Google Professional Cloud Architect
Tìm thấy 420 câu.
Your company handles large datasets for real-time analytics. You are tasked with designing a data processing pipeline that ingests raw data into Google Cloud Storage (GCS) and then processes this data using BigQuery. The data is ingested in various formats like JSON, CSV, and Avro, and needs to be cleaned and transformed before analysis. The processed data must be available in BigQuery within 10 minutes of ingestion to meet business requirements. Additionally, the system should be cost-effective and scalable, given the unpredictability of data volume. Which of the following design choices would best meet the requirements for this data processing pipeline?
-
A
Use Cloud Dataflow for data ingestion and transformation, storing the intermediate results in GCS, and loading the final cleaned data into BigQuery via a scheduled Dataflow job.
-
B
Deploy a Dataproc cluster to process the data, storing the intermediate data in GCS, and then using BigQuery’s federated queries to access the processed data.
-
C
Use a scheduled Cloud Function triggered by new objects in GCS to ingest and transform data, storing the results directly into BigQuery.
-
D
Use Pub/Sub to ingest the data, triggering Dataflow for real-time processing, and then load the cleaned data into BigQuery.
Xem giải thích
Đáp án
D — Pub/Sub nhận dữ liệu, kích hoạt Dataflow xử lý thời gian thực
Vì sao đúng
Đề nói phân tích thời gian thực, và điểm phân biệt nằm ở lớp nhận dữ liệu. Pub/Sub làm lớp đệm: nó hấp thụ được tải bùng phát, đảm bảo không mất tin, và tách rời khâu sinh dữ liệu khỏi khâu xử lý — bên xử lý tạm chậm thì thông điệp nằm chờ chứ không dội ngược về nguồn. Dataflow xử lý luồng đó với cửa sổ thời gian và khả năng xử lý dữ liệu tới muộn.
Vì sao các phương án khác sai
- A. Dataflow đọc thẳng từ Cloud Storage — kho đối tượng làm nguồn nghĩa là xử lý theo lô; mất tính thời gian thực.
- B. Cụm Dataproc — dựng cho xử lý theo lô, và phải vận hành cụm.
- C. Cloud Function kích hoạt theo đối tượng mới trong GCS — vẫn theo tệp nên có độ trễ, và chi phí gọi hàm ở khối lượng lớn là không khả thi.
Your company runs multiple internal tools for HR, finance, and engineering, each hosted on Google Cloud with APIs exposed through Apigee. Leadership wants to enhance employee workflows by using Gemini AI Agents to automate cross-application tasks such as generating onboarding checklists, summarizing expense reports, and creating project updates. You are designing an integration strategy that maintains data governance, security, and scalability across the organization. Which approach best meets these requirements?
-
A
Use Gemini AI Agents to fetch all HR, finance, and engineering data into a BigQuery dataset for central summarization before generating responses.
-
B
Use Gemini AI Agents integrated with Vertex AI Search and Conversation to orchestrate API calls through Apigee, applying IAM-based access control and Vertex AI private endpoints.
-
C
Directly expose all internal APIs to Gemini AI Agents using public endpoints secured with API keys for simplicity.
-
D
Connect Gemini AI Agents to each tool via Cloud Functions triggers that forward raw data to the agent for contextual analysis.
Xem giải thích
Đáp án
B — Gemini AI Agents tích hợp với Vertex AI Search and Conversation để điều phối các API
Vì sao đúng
Bài toán là một trợ lý nội bộ trả lời được câu hỏi trải trên nhiều hệ thống — nhân sự, tài chính, kỹ thuật. Cách đúng là để tác nhân gọi tới các API sẵn có qua Apigee thay vì sao chép dữ liệu đi nơi khác: dữ liệu luôn mới, và phân quyền vẫn do chính các hệ thống gốc quyết định — ai không được xem lương thì hỏi trợ lý cũng không ra.
Vì sao các phương án khác sai
- A. Kéo hết dữ liệu của ba hệ thống vào một kho chung — nhân bản dữ liệu nhạy cảm và làm mất mô hình phân quyền gốc; đây là rủi ro lớn nhất.
- C. Phơi toàn bộ API nội bộ ra điểm cuối công khai — mở bề mặt tấn công ra Internet.
- D. Nối qua Cloud Functions chuyển tiếp thẳng — bỏ qua lớp quản trị API mà Apigee đang cung cấp: xác thực, giới hạn tần suất, ghi vết.
As a cloud architect, you have been tasked with setting up a robust and flexible content delivery system. The requirement is to have different Compute Engine instances serve content based on the URL path. For example, requests to www.example.com/audio/* should be served by one set of instances, and requests to www.example.com/video/* should be served by another set of instances. Which of the following would be the most appropriate approach to achieve this?
-
A
Create an HTTPS load balancer and define URL maps to route traffic to the appropriate set of instances based on the path.
-
B
Create a single instance group and use instance templates to define the content to be served based on the URL path.
-
C
Create an HTTPS load balancer with a single backend service. Use instance tagging to route traffic to the correct set of instances.
-
D
Create two separate HTTPS load balancers, one for each path (
/audioand/video), each pointing to a different set of instances.
Xem giải thích
Đáp án
A — Dựng HTTPS load balancer và khai URL map để định tuyến theo đường dẫn
Vì sao đúng
URL map là tính năng cốt lõi của load balancer tầng 7: nó đọc đường dẫn của yêu cầu rồi chuyển tới backend service tương ứng — /audio sang nhóm máy phục vụ âm thanh, /video sang nhóm khác. Tất cả nằm sau một địa chỉ IP và một chứng chỉ TLS duy nhất, nên người dùng chỉ thấy một tên miền.
Vì sao các phương án khác sai
- D. Dựng hai load balancer riêng — thành hai địa chỉ IP, hai chứng chỉ, và người dùng phải biết gọi đúng nơi; đúng thứ URL map sinh ra để tránh.
- C. Một backend service duy nhất rồi phân biệt bằng tag máy — tag dùng cho luật tường lửa, không phải cơ chế định tuyến theo đường dẫn.
- B. Một instance group với instance template — template quyết định máy trông như thế nào, không quyết định yêu cầu nào đi đâu.
You are responsible for architecting a three-tier web application on Google Cloud Platform. This application includes a web server tier, an application server tier, and a database tier, each tier hosted on separate Compute Engine instances. You need to control network traffic between these tiers and from the public internet. To achieve this, you decide to use tags and firewall rules. Which of the following options would be the most effective way to apply these tags and set up firewall rules?
-
A
Add a tag only to the database tier and create a firewall rule that only allows traffic from the IP addresses of the other two tiers.
-
B
Assign each tier a unique tag and create individual firewall rules to control traffic between tags and from the public internet.
-
C
Add a single tag to all instances irrespective of the tier and create one firewall rule to allow all inbound traffic.
-
D
Assign the same tag to the web server and application server tier, and a different tag to the database tier. Then, create a firewall rule to allow all traffic within the same tag and limited traffic to the database tag.
Xem giải thích
Đáp án
B — Gán cho mỗi tầng một tag riêng và viết luật tường lửa cho từng luồng
Vì sao đúng
Network tag mô tả vai trò nên bám theo máy khi nhóm co giãn: máy mới thừa hưởng tag và luật áp dụng ngay. Mỗi tầng một tag cho phép khai chính xác từng luồng được phép — giao diện tới ứng dụng, ứng dụng tới CSDL — và vì VPC mặc định chặn mọi lưu lượng đi vào, những gì không khai đều bị chặn mà không phải liệt kê.
Vì sao các phương án khác sai
- A. Chỉ gắn tag cho tầng CSDL — bảo vệ được tầng dữ liệu nhưng để ngỏ luồng giữa giao diện và tầng ứng dụng.
- C. Một tag chung cho tất cả với một luật duy nhất — không phân biệt được tầng nào, nên không phân đoạn được gì.
- D. Gộp giao diện và tầng ứng dụng chung một tag — mất ranh giới giữa hai tầng đó; máy chủ web bị chiếm là chạm thẳng vào tầng ứng dụng.
You are a cloud architect responsible for a microservices-based application deployed in Google Kubernetes Engine (GKE). The application consists of several interacting containers. To optimize performance, you want to ensure that these containers are located as close to each other as possible. Which of the following would be the most effective way to achieve this?
-
A
Use the
PodAffinityandPodAntiAffinityrules in your Kubernetes pod specification. -
B
Increase the number of replicas for each container to ensure they are distributed across all nodes in the cluster.
-
C
Use node taints and tolerations to force certain containers to run on specific nodes.
-
D
Use GKE Regional Clusters to deploy the containers across multiple zones in the same region.
Xem giải thích
Đáp án
A — Dùng luật PodAffinity và PodAntiAffinity trong đặc tả pod
Vì sao đúng
Hai luật này cho bạn nói rõ pod nào nên hoặc không nên nằm cùng chỗ:
- PodAffinity — đặt các pod hay gọi nhau lên cùng node hoặc cùng zone, giảm độ trễ mạng giữa chúng.
- PodAntiAffinity — trải các bản sao của cùng một dịch vụ ra các node khác nhau, nên mất một node không mất hết bản sao.
Đây đúng là công cụ cho ứng dụng microservice có nhiều thành phần tương tác lẫn nhau.
Vì sao các phương án khác sai
- B. Tăng số bản sao và mong chúng được trải đều — Kubernetes không đảm bảo điều đó nếu không khai anti-affinity; hai bản sao có thể rơi vào cùng một node.
- C. Dùng taint và toleration — cơ chế để dành riêng node cho một loại workload, không phải để diễn đạt quan hệ giữa các pod.
- D. Dùng cụm theo vùng — trải node qua nhiều zone, nhưng không quyết định pod nào đặt cạnh pod nào.
You are tasked with setting up secure and scalable network communication between multiple VPCs in different projects within your Google Cloud organization. Each project is managed by a separate team. You want to enable centralized network control while allowing service teams to manage their workloads independently. What should you do?
-
A
Use Cloud VPN between each project’s VPC and a central VPC.
-
B
Use Shared VPC and host the network in a central host project.
-
C
Use VPC peering between all VPCs in a full mesh topology.
-
D
Create a default VPC in each project and rely on public IPs for communication.
Xem giải thích
Đáp án
B — Dùng Shared VPC và đặt mạng ở một dự án chủ trung tâm
Vì sao đúng
Đây là mô hình chuẩn cho tổ chức nhiều dự án: dự án chủ giữ VPC, subnet, luật tường lửa và các kết nối lai; các dự án dịch vụ gắn vào đó và dùng chung mạng. Nhờ vậy đội hạ tầng quản lý mạng ở một chỗ, còn mỗi đội vẫn có dự án riêng để quản lý tài nguyên, hạn mức và chi phí của mình. Máy ở các dự án khác nhau nói chuyện bằng IP riêng, không đi ra ngoài.
Vì sao các phương án khác sai
- C. Peering đầy đủ giữa mọi VPC — số cặp peering tăng theo bình phương số dự án, và peering không có tính bắc cầu nên phải khai từng cặp một.
- A. VPN từ mỗi dự án về một VPC trung tâm — thêm cổng VPN phải vận hành, băng thông có trần, và tốn hơn hẳn khi mọi thứ đều ở trong Google Cloud.
- D. VPC mặc định ở mỗi dự án rồi dùng IP công khai — đẩy lưu lượng nội bộ ra Internet; sai cả về bảo mật lẫn chi phí.
For this question, refer to the Cymbal Retail case study.
https://services.google.com/fh/files/misc/v6.1_pca_cymbal_retail_case_study_english.pdf
Cymbal wants to automate product catalog enrichment to improve product descriptions, attributes, and categorization across millions of SKUs. Product data currently resides in multiple databases (MySQL, SQL Server, MongoDB) and is updated through batch ETL processes and SFTP file transfers. The new solution must:
-
Minimize operational overhead
-
Scale automatically as the catalog grows
-
Enable AI-driven enrichment (for example, extracting attributes from product descriptions and images)
-
Integrate cleanly with existing Kubernetes-based applications
Which architecture best meets Cymbal’s requirements while following best practices?
-
A
Deploy a monolithic enrichment service on Google Kubernetes Engine that polls all databases and performs enrichment synchronously.
-
B
Use Dataflow to ingest product data into BigQuery, apply enrichment using Vertex AI APIs, and publish enriched catalog data through Pub/Sub to downstream services.
-
C
Export product data nightly via SFTP to Compute Engine VMs, run custom Python scripts for enrichment, and write results back to on-premises databases.
-
D
Stream product data changes into BigQuery, use scheduled SQL queries for enrichment, and push enriched data back to transactional databases.
Xem giải thích
Đáp án
B — Dùng Dataflow nạp dữ liệu sản phẩm vào BigQuery, làm giàu bằng Vertex AI
Vì sao đúng
Đề cần tự động làm giàu dữ liệu sản phẩm, và phương án này ghép đúng ba mảnh: Dataflow lo phần đường ống chạy liên tục và tự co giãn, BigQuery làm nơi chứa và truy vấn, còn Vertex AI cung cấp phần suy luận để sinh mô tả, phân loại hay thuộc tính còn thiếu. Cả ba đều là dịch vụ được quản lý nên không có cụm hay máy chủ nào phải vận hành.
Vì sao các phương án khác sai
- D. Đưa dữ liệu vào BigQuery rồi làm giàu bằng truy vấn SQL theo lịch — SQL không sinh được mô tả sản phẩm hay phân loại theo ngữ nghĩa; đó là việc của mô hình.
- A. Một dịch vụ nguyên khối trên GKE hỏi vòng liên tục — hỏi vòng vừa tốn vừa có độ trễ, và bạn phải vận hành cụm.
- C. Xuất qua SFTP hằng đêm rồi chạy script Python trên máy ảo — dữ liệu luôn cũ một ngày, và đây là kiểu quy trình hỏng âm thầm.
For this question, refer to the Cymbal Retail case study.
https://services.google.com/fh/files/misc/v6.1_pca_cymbal_retail_case_study_english.pdf
Cymbal plans to introduce conversational commerce features where customers can ask for products in different colors or styles during a chat session. When a customer requests a variant that does not already exist, the system should generate a new image dynamically and store it for future reuse. The solution must be cost-efficient, avoid unnecessary recomputation, and integrate cleanly with Cymbal’s existing data stores and microservices. What is the most appropriate design to support this requirement?
-
A
Trigger batch image generation jobs nightly using Cloud Composer and overwrite existing product images
-
B
Pre-generate all possible product image variants using Vertex AI and store them permanently in Cloud Storage
-
C
Generate images on demand using Vertex AI, cache generated images in Cloud Storage, and store metadata references in an existing product database
-
D
Generate images in real time using Cloud Run and discard them after the chat session ends
Xem giải thích
Đáp án
C — Sinh ảnh theo yêu cầu bằng Vertex AI, rồi nhớ đệm ảnh đã sinh vào Cloud Storage
Vì sao đúng
Phương án này cân bằng đúng giữa chi phí và trải nghiệm. Sinh theo yêu cầu nghĩa là chỉ trả tiền cho ảnh thật sự có người xem — quan trọng vì số tổ hợp sản phẩm rất lớn mà phần lớn không bao giờ được hỏi tới. Nhớ đệm lại khiến lần yêu cầu thứ hai cho cùng một ảnh trả về ngay và không tốn thêm lần suy luận nào.
Vì sao các phương án khác sai
- B. Sinh trước mọi biến thể có thể có — số tổ hợp bùng nổ; trả tiền sinh và lưu cho rất nhiều ảnh chẳng ai xem.
- A. Sinh theo lô hằng đêm rồi ghi đè — ảnh cho sản phẩm mới phải chờ tới đêm, và vẫn sinh cả những thứ không ai cần.
- D. Sinh thời gian thực rồi bỏ đi sau phiên trò chuyện — mỗi lần hỏi lại là một lần trả tiền suy luận cho đúng thứ vừa sinh xong.
For this question, refer to the EHR Healthcare case study.
https://services.google.com/fh/files/misc/v6.1_pca_ehr_healthcare_case_study_english.pdf
HR Healthcare is expanding their operations into a new geographic region. They plan to deploy a new patient analytics application on Google Cloud in that region. The application must have low-latency access to a new on-premises database hosted in their regional data center. The network team confirms the data center is located less than 10 km from a Google point of presence, and the existing routers in the data center support VLAN attachments. The team also wants an SLA-backed, high-throughput connection with secure, private IP communication. What should you recommend?
-
A
Use Cloud VPN with dynamic (BGP) routing.
-
B
Use Partner Interconnect with a high-availability configuration.
-
C
Use Dedicated Interconnect.
-
D
Use Cloud VPN with static routing.
Xem giải thích
Đáp án
B — Partner Interconnect ở cấu hình sẵn sàng cao
Vì sao đúng
EHR cần kết nối riêng tư và ổn định cho dữ liệu y tế, nhưng chưa tới mức phải có mặt tại cơ sở colocation và cam kết đường 10 Gbps. Partner Interconnect lấp đúng khoảng giữa: nối qua nhà cung cấp đối tác, chọn băng thông từ 50 Mbps tới 50 Gbps, mà vẫn vào thẳng VPC riêng. Cấu hình sẵn sàng cao là phần bắt buộc — một đường là một điểm hỏng duy nhất.
Vì sao các phương án khác sai
- C. Dedicated Interconnect — cho băng thông cao nhất nhưng đòi hiện diện tại colocation và bắt đầu từ 10 Gbps; quá nặng khi nhu cầu nhỏ hơn.
- A và D. Cloud VPN — chạy trên Internet công cộng nên độ trễ và băng thông biến động; riêng bản dùng định tuyến tĩnh còn tệ hơn vì tuyến không tự tính lại khi đường hỏng.
You are a cloud architect assigned to manage access to resources in a Google Cloud project for a team. The team consists of a Data Scientist who needs to run BigQuery jobs, a Cloud Engineer who deploys applications on App Engine, a Network Engineer who manages VPCs, and an Auditor who checks IAM permissions and resource usage. Which of the following sets of predefined roles should you assign to each team member to grant them the least privilege they need to do their job effectively?
-
A
BigQuery Data Editor to the Data Scientist, App Engine Service Admin to the Cloud Engineer, Compute Network Viewer to the Network Engineer, and Security Reviewer to the Auditor.
-
B
BigQuery User to the Data Scientist, App Engine Viewer to the Cloud Engineer, Compute Network Admin to the Network Engineer, and Viewer to the Auditor.
-
C
BigQuery User to the Data Scientist, App Engine Admin to the Cloud Engineer, Compute Network User to the Network Engineer, and Security Reviewer to the Auditor.
-
D
BigQuery User to the Data Scientist, App Engine Admin to the Cloud Engineer, Compute Network Admin to the Network Engineer, and Security Reviewer to the Auditor.
Xem giải thích
Đáp án
D — BigQuery User / App Engine Admin / Compute Network Admin / Security Reviewer
Vì sao đúng
Bốn người, bốn nhu cầu, và mỗi vai phải khớp đúng động từ trong đề:
- Chạy job BigQuery →
BigQuery User. Vai này cho tạo job và chạy truy vấn, khácdataViewer(đọc được dữ liệu nhưng không tạo được job). - Triển khai ứng dụng lên App Engine →
App Engine Admin. Triển khai là hành động ghi. - Quản lý VPC →
Compute Network Admin. "Quản lý" là quyền sửa. - Kiểm tra quyền IAM và mức dùng tài nguyên →
Security Reviewer, vai chỉ đọc dựng riêng cho việc rà soát cấu hình bảo mật.
Vì sao các phương án khác sai
- C.
Compute Network Usercho kỹ sư mạng — vai này chỉ cho sử dụng mạng đã có (gắn máy vào subnet), không cho tạo hay sửa VPC. - B.
App Engine Viewercho kỹ sư triển khai — chỉ xem được, không triển khai nổi. - A.
BigQuery Data Editorcho nhà khoa học dữ liệu — cho quyền sửa dữ liệu nhưng vẫn không tạo được job, nên vẫn không chạy truy vấn được.