Ngân hàng đề — Google Cloud Professional Data Engineer
Tìm thấy 429 câu.
Company Overview -
MJTelco is a startup that plans to build networks in rapidly growing, underserved markets around the world. The company has patents for innovative optical communications hardware. Based on these patents, they can create many reliable, high-speed backbone links with inexpensive hardware.
Company Background -
Founded by experienced telecom executives, MJTelco uses technologies originally developed to overcome communications challenges in space. Fundamental to their operation, they need to create a distributed data infrastructure that drives real-time analysis and incorporates machine learning to continuously optimize their topologies. Because their hardware is inexpensive, they plan to overdeploy the network allowing them to account for the impact of dynamic regional politics on location availability and cost.
Their management and operations teams are situated all around the globe creating many-to-many relationship between data consumers and provides in their system. After careful consideration, they decided public cloud is the perfect environment to support their needs.
Solution Concept -
MJTelco is running a successful proof-of-concept (PoC) project in its labs. They have two primary needs:
✑ Scale and harden their PoC to support significantly more data flows generated when they ramp to more than 50,000 installations.
Refine their machine-learning cycles to verify and improve the dynamic models they use to control topology definition.
MJTelco will also use three separate operating environments `" development/test, staging, and production `" to meet the needs of running experiments, deploying new features, and serving production customers.
Business Requirements -
✑ Scale up their production environment with minimal cost, instantiating resources when and where needed in an unpredictable, distributed telecom user community.
✑ Ensure security of their proprietary data to protect their leading-edge machine learning and analysis.
✑ Provide reliable and timely access to data for analysis from distributed research workers
✑ Maintain isolated environments that support rapid iteration of their machine-learning models without affecting their customers.
Technical Requirements -
Ensure secure and efficient transport and storage of telemetry data
Rapidly scale instances to support between 10,000 and 100,000 data providers with multiple flows each.
Allow analysis and presentation against data tables tracking up to 2 years of data storing approximately 100m records/day
Support rapid iteration of monitoring infrastructure focused on awareness of data pipeline problems both in telemetry flows and in production learning cycles.
CEO Statement -
Our business model relies on our patents, analytics and dynamic machine learning. Our inexpensive hardware is organized to be highly reliable, which gives us cost advantages. We need to quickly stabilize our large distributed data pipelines to meet our reliability and capacity commitments.
CTO Statement -
Our public cloud services must operate as advertised. We need resources that scale and keep our data secure. We also need environments in which our data scientists can carefully study and quickly adapt our models. Because we rely on automation to process our data, we also need our development and test environments to work as we iterate.
CFO Statement -
The project is too large for us to maintain the hardware and software required for the data and analysis. Also, we cannot afford to staff an operations team to monitor so many data feeds, so we will rely on automation and infrastructure. Google Cloud's machine learning will allow our quantitative researchers to work on our high-value problems instead of problems with our data pipelines.
You need to compose visualization for operations teams with the following requirements:
✑ Telemetry must include data from all 50,000 installations for the most recent 6 weeks (sampling once every minute)
✑ The report must not be more than 3 hours delayed from live data.
✑ The actionable report should only show suboptimal links.
✑ Most suboptimal links should be sorted to the top.
Suboptimal links can be grouped and filtered by regional geography.
✑ User response time to load the report must be <5 seconds.
You create a data source to store the last 6 weeks of data, and create visualizations that allow viewers to see multiple date ranges, distinct geographic regions, and unique installation types. You always show the latest data without any changes to your visualizations. You want to avoid creating and updating new visualizations each month. What should you do?
- A Look through the current data and compose a series of charts and tables, one for each possible combination of criteria.
- B Look through the current data and compose a small set of generalized charts and tables bound to criteria filters that allow value selection.
- C Export the data to a spreadsheet, compose a series of charts and tables, one for each possible combination of criteria, and spread them across multiple tabs.
- D Load the data into relational database tables, write a Google App Engine application that queries all rows, summarizes the data across each criteria, and then renders results using the Google Charts and visualization API.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi thuộc case study MJTelco – một startup viễn thông xây dựng mạng lưới ở các thị trường đang phát triển nhanh chóng, sử dụng công nghệ quang học giá rẻ để tạo backbone links tốc độ cao. Họ cần hạ tầng dữ liệu phân tán trên public cloud (Google Cloud) để xử lý dữ liệu telemetry thời gian thực từ hàng chục nghìn installations (lên đến 50.000+), tích hợp machine learning để tối ưu topology mạng.
Yêu cầu cụ thể cho visualization dành cho operations teams:
- 📊 Dữ liệu telemetry từ tất cả 50.000 installations, trong 6 tuần gần nhất (sampling mỗi phút) → Khối lượng dữ liệu khổng lồ (~100 triệu records/ngày, theo Technical Requirements).
- ⏱️ Báo cáo không chậm quá 3 giờ so với dữ liệu live.
- 🔍 Chỉ hiển thị suboptimal links (liên kết không tối ưu), sắp xếp suboptimal nhất lên top.
- 🗺️ Nhóm và filter theo regional geography (khu vực địa lý).
- ⚡ User response time <5 giây khi load báo cáo.
- 💾 Tạo data source lưu 6 tuần dữ liệu gần nhất.
- 📈 Visualization phải hỗ trợ xem multiple date ranges, distinct geographic regions, và unique installation types.
- 🔄 Luôn hiển thị dữ liệu mới nhất mà không cần thay đổi visualization mỗi tháng (tránh tạo mới/update hàng tháng).
Mục tiêu chính: Xây dựng báo cáo actionable, scalable, an toàn, với chi phí thấp, phù hợp môi trường dev/test/staging/prod riêng biệt. Dữ liệu được lưu trữ và query hiệu quả (gợi ý BigQuery cho storage/query nhanh), visualization dùng công cụ như Looker Studio (trước là Data Studio, cập nhật đến 2026 hỗ trợ filters động, parameters mạnh mẽ).
🛠️ Vấn đề cốt lõi: Cần visualization linh hoạt, generalized với filters động để user tự chọn criteria (date range, region, type), tránh hard-code charts cho từng combo (vì dữ liệu thay đổi monthly và khối lượng lớn → không scalable, tốn thời gian).
📘 Tài liệu tham khảo:
- Google Cloud Professional Data Engineer Exam Guide (2024-2026): cloud.google.com/certification/guides/data-engineer.
- Looker Studio Documentation (cập nhật 2026): lookerstudio.google.com – Hỗ trợ dynamic filters, parameters cho BigQuery data sources.
- BigQuery Best Practices for Viz: cloud.google.com/bigquery/docs/visualize-data-studio.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Look through the current data and compose a small set of generalized charts and tables bound to criteria filters that allow value selection.
Lý do 🏆:
- Phương án này sử dụng charts/tables generalized (tổng quát) kết nối với filters động (như date range picker, region dropdown, installation type selector) trong Looker Studio hoặc Connected Sheets.
- ✅ Đáp ứng đầy đủ requirements: Hiển thị latest data tự động (data source refresh scheduled <3h), filter chỉ suboptimal links (sắp xếp top bằng sort/filter), group by region, load <5s (BigQuery query optimized), không cần update monthly vì filters tự adapt dữ liệu mới.
- 🧩 Scalable & cost-effective: Chỉ cần small set charts (ví dụ: 1 dashboard với 3-5 viz chính + controls), phù hợp 50k installations, 6 weeks data (~hàng TB), tận dụng BigQuery partitioning/clustering cho query nhanh.
- Theo CTO/CFO: Tự động hóa, không cần ops team monitor, data scientists iterate nhanh.
📋 Giải thích tất cả các phương án
-
❌ [SAI] Look through the current data and compose a series of charts and tables, one for each possible combination of criteria.
Lý do sai: Tạo hàng loạt charts riêng lẻ cho từng combo (date x region x type) → Không scalable với 50k installations và dữ liệu thay đổi monthly (phải recreate hàng tháng). Load time >5s vì quá nhiều viz static, không filter động suboptimal links, vi phạm "avoid creating/updating new visualizations each month". Không hiệu quả cho real-time analysis. -
✅ [ĐÚNG] Look through the current data and compose a small set of generalized charts and tables bound to criteria filters that allow value selection.
Lý do đúng: Như đã giải thích ở trên. Sử dụng filters bound (parameters/filters trong Looker Studio) để user chọn criteria → Linh hoạt, always latest data, chỉ show suboptimal (dùng calculated fields/sort), group/filter region dễ dàng. Optimized cho BigQuery, refresh tự động <3h. -
❌ [SAI] Export the data to a series of charts and tables, one for each possible combination of criteria, and spread them across multiple tabs.
Lý do sai: Export sang spreadsheet (Google Sheets?) → Không handle được 6 weeks data từ 50k sources (quá lớn, ~100m records/day → timeout/export fail). Static tabs per combo → Không real-time (<3h delay ok nhưng không latest monthly), load chậm >5s, không filter động, phải manual update hàng tháng. Không secure cho proprietary data. -
❌ [SAI] Load the data into relational database tables, write a Google App Engine application that queries all rows, summarizes the data across each criteria, and then renders results using the Google Charts and visualization API.
Lý do sai: Relational DB (Cloud SQL?) không phù hợp petabyte-scale telemetry (BigQuery tốt hơn). App Engine app query all rows → Latency cao (>5s với 100m records/day x 6 weeks), cost cao (full scan), phức tạp maintain (code custom summarize/filter). Không native support latest data auto-refresh, khó iterate ML cycles, vi phạm "minimal cost" và automation.
🔥 Kết luận: Phương án đúng tận dụng Looker Studio + BigQuery (standard GCP solution 2026) cho dashboard động, giúp MJTelco scale production mà không ảnh hưởng customers! 🚀
Company Overview -
MJTelco is a startup that plans to build networks in rapidly growing, underserved markets around the world. The company has patents for innovative optical communications hardware. Based on these patents, they can create many reliable, high-speed backbone links with inexpensive hardware.
Company Background -
Founded by experienced telecom executives, MJTelco uses technologies originally developed to overcome communications challenges in space. Fundamental to their operation, they need to create a distributed data infrastructure that drives real-time analysis and incorporates machine learning to continuously optimize their topologies. Because their hardware is inexpensive, they plan to overdeploy the network allowing them to account for the impact of dynamic regional politics on location availability and cost.
Their management and operations teams are situated all around the globe creating many-to-many relationship between data consumers and provides in their system. After careful consideration, they decided public cloud is the perfect environment to support their needs.
Solution Concept -
MJTelco is running a successful proof-of-concept (PoC) project in its labs. They have two primary needs:
✑ Scale and harden their PoC to support significantly more data flows generated when they ramp to more than 50,000 installations.
✑ Refine their machine-learning cycles to verify and improve the dynamic models they use to control topology definition.
MJTelco will also use three separate operating environments `" development/test, staging, and production `" to meet the needs of running experiments, deploying new features, and serving production customers.
Business Requirements -
✑ Scale up their production environment with minimal cost, instantiating resources when and where needed in an unpredictable, distributed telecom user community.
✑ Ensure security of their proprietary data to protect their leading-edge machine learning and analysis.
✑ Provide reliable and timely access to data for analysis from distributed research workers
✑ Maintain isolated environments that support rapid iteration of their machine-learning models without affecting their customers.
Technical Requirements -
Ensure secure and efficient transport and storage of telemetry data
Rapidly scale instances to support between 10,000 and 100,000 data providers with multiple flows each.
Allow analysis and presentation against data tables tracking up to 2 years of data storing approximately 100m records/day
Support rapid iteration of monitoring infrastructure focused on awareness of data pipeline problems both in telemetry flows and in production learning cycles.
CEO Statement -
Our business model relies on our patents, analytics and dynamic machine learning. Our inexpensive hardware is organized to be highly reliable, which gives us cost advantages. We need to quickly stabilize our large distributed data pipelines to meet our reliability and capacity commitments.
CTO Statement -
Our public cloud services must operate as advertised. We need resources that scale and keep our data secure. We also need environments in which our data scientists can carefully study and quickly adapt our models. Because we rely on automation to process our data, we also need our development and test environments to work as we iterate.
CFO Statement -
The project is too large for us to maintain the hardware and software required for the data and analysis. Also, we cannot afford to staff an operations team to monitor so many data feeds, so we will rely on automation and infrastructure. Google Cloud's machine learning will allow our quantitative researchers to work on our high-value problems instead of problems with our data pipelines.
Given the record streams MJTelco is interested in ingesting per day, they are concerned about the cost of Google BigQuery increasing. MJTelco asks you to provide a design solution. They require a single large data table called tracking_table. Additionally, they want to minimize the cost of daily queries while performing fine-grained analysis of each day's events. They also want to use streaming ingestion. What should you do?
- A Create a table called tracking_table and include a DATE column.
- B Create a partitioned table called tracking_table and include a TIMESTAMP column.
- C Create sharded tables for each day following the pattern tracking_table_YYYYMMDD.
- D Create a table called tracking_table with a TIMESTAMP column to represent the day.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi thuộc case study MJTelco – một startup viễn thông sử dụng Google Cloud (đặc biệt BigQuery) để xử lý dữ liệu telemetry lớn từ hàng chục nghìn thiết bị (khoảng 100 triệu records/ngày). Họ cần một giải pháp thiết kế bảng dữ liệu duy nhất tên tracking_table để:
- Ingest dữ liệu streaming (dữ liệu thời gian thực từ các nguồn phân tán).
- Phân tích fine-grained từng ngày (phân tích chi tiết sự kiện theo ngày).
- Giảm thiểu chi phí query hàng ngày trên BigQuery, vì lượng dữ liệu khổng lồ có thể làm chi phí tăng vọt (BigQuery tính phí dựa trên bytes scanned khi query).
Vấn đề cốt lõi: BigQuery là data warehouse serverless, chi phí query phụ thuộc vào lượng dữ liệu quét (scanned). Không tối ưu hóa sẽ scan toàn bộ bảng lớn (lên đến 2 năm dữ liệu ~73 tỷ records), dẫn đến tốn kém. Giải pháp cần hỗ trợ partitioning hoặc sharding để partition pruning (chỉ quét partition cần thiết khi query theo thời gian), kết hợp streaming ingestion, và phù hợp với môi trường dev/test/staging/prod.
Yêu cầu kỹ thuật liên quan (cập nhật BigQuery đến 2026):
- Streaming inserts hỗ trợ partitioned tables (từ 2018, tối ưu hóa liên tục).
- Partitioning theo
_PARTITIONTIME(pseudo-column TIMESTAMP) hoặc ingestion-time partitioning là best practice cho dữ liệu thời gian (time-series như telemetry). - Không cần sharding thủ công vì BigQuery tự động cluster và partition.
📘 Tài liệu tham khảo:
- BigQuery Partitioned Tables (Google Cloud Docs, cập nhật 2025).
- Streaming Ingestion into Partitioned Tables (hỗ trợ đầy đủ TIMESTAMP partitioning).
- MJTelco Case Study chính thức từ Google Cloud Professional Data Engineer exam guide.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create a partitioned table called tracking_table and include a TIMESTAMP column.
Lý do 🛠️:
- Partitioned table sử dụng ingestion-time partitioning (dựa trên
_PARTITIONTIME– pseudo-column TIMESTAMP tự động tạo theo ngày ingest), cho phép partition pruning khi query theo ngày/thời gian → chỉ scan dữ liệu cần thiết, giảm chi phí query hàng ngày lên đến 90-99% cho phân tích fine-grained (ví dụ:WHERE _PARTITIONTIME = TIMESTAMP("2025-01-01")). - TIMESTAMP column (khuyến nghị cho dữ liệu telemetry chính xác đến giây/phút) làm partition key hoặc clustering column, hỗ trợ time-unit column partitioning (ngày/tháng/giờ), tối ưu streaming ingestion (BigQuery tự partition dữ liệu stream theo thời gian ingest).
- Phù hợp quy mô: Scale từ 10k-100k data providers, lưu 2 năm dữ liệu, isolated envs, và ML cycles nhanh.
- Cập nhật 2026: BigQuery hỗ trợ optimized partitioning với SLAs cao hơn cho streaming (low-latency <1s), không downtime khi scale.
📋 Giải thích tất cả các phương án (đúng/sai)
-
✅ [ĐÚNG] Create a partitioned table called tracking_table and include a TIMESTAMP column.
🟢 Đúng vì: Như giải thích trên, partitioning + TIMESTAMP là giải pháp chuẩn của BigQuery cho time-series data, hỗ trợ streaming, pruning hiệu quả, giảm cost tối đa mà không cần quản lý thủ công. Lý tưởng cho MJTelco với dữ liệu telemetry real-time và query daily fine-grained. -
❌ [SAI] Create a table called tracking_table and include a DATE column.
🔴 Sai vì: Bảng không partitioned chỉ scan toàn bộ dữ liệu (100m records/ngày × 2 năm = hàng trăm TB), chi phí query cao ngất. DATE column chỉ hữu ích nếu dùng làm partition key, nhưng thiếu partitioning thì vô ích. Streaming vẫn work nhưng không pruning → không giải quyết vấn đề cost. -
❌ [SAI] Create sharded tables for each day following the pattern tracking_table_YYYYMMDD.
🔴 Sai vì: Sharding thủ công (tạo bảng riêng mỗi ngày) phức tạp quản lý (hàng nghìn bảng/năm), không hỗ trợ tốt streaming ingestion (phải route stream thủ công), query cross-day khó khăn (UNION ALL), tăng ops overhead. BigQuery khuyến cáo tránh sharding vì partitioning tự động tốt hơn, rẻ hơn (không cần code automation). -
❌ [SAI] Create a table called tracking_table with a TIMESTAMP column to represent the day.
🔴 Sai vì: TIMESTAMP column chỉ "đại diện ngày" không partitioning → vẫn full table scan khi query theo ngày (dùngWHERE DATE(timestamp_col) = '2025-01-01'kém hiệu quả, scan ~toàn bộ). Không tận dụng pruning thực sự. TIMESTAMP tốt nhưng thiếu partitioning là fail cost optimization.
Kết luận 🎯: Giải pháp partitioned table là best practice Google Cloud cho MJTelco, đảm bảo scale, secure, low-cost, và rapid iteration! Nếu implement, dùng SQL DDL: CREATE TABLE tracking_table (..., event_timestamp TIMESTAMP) PARTITION BY DATE(event_timestamp).
Company Overview -
Flowlogistic is a leading logistics and supply chain provider. They help businesses throughout the world manage their resources and transport them to their final destination. The company has grown rapidly, expanding their offerings to include rail, truck, aircraft, and oceanic shipping.
Company Background -
The company started as a regional trucking company, and then expanded into other logistics market. Because they have not updated their infrastructure, managing and tracking orders and shipments has become a bottleneck. To improve operations, Flowlogistic developed proprietary technology for tracking shipments in real time at the parcel level. However, they are unable to deploy it because their technology stack, based on Apache Kafka, cannot support the processing volume. In addition, Flowlogistic wants to further analyze their orders and shipments to determine how best to deploy their resources.
Solution Concept -
Flowlogistic wants to implement two concepts using the cloud:
✑ Use their proprietary technology in a real-time inventory-tracking system that indicates the location of their loads
✑ Perform analytics on all their orders and shipment logs, which contain both structured and unstructured data, to determine how best to deploy resources, which markets to expand info. They also want to use predictive analytics to learn earlier when a shipment will be delayed.
Existing Technical Environment -
Flowlogistic architecture resides in a single data center:
✑ Databases
- 8 physical servers in 2 clusters
- SQL Server `" user data, inventory, static data
- 3 physical servers
- Cassandra `" metadata, tracking messages
10 Kafka servers `" tracking message aggregation and batch insert
✑ Application servers `" customer front end, middleware for order/customs
- 60 virtual machines across 20 physical servers
- Tomcat `" Java services
- Nginx `" static content
- Batch servers
✑ Storage appliances
- iSCSI for virtual machine (VM) hosts
- Fibre Channel storage area network (FC SAN) `" SQL server storage
Network-attached storage (NAS) image storage, logs, backups
✑ 10 Apache Hadoop /Spark servers
- Core Data Lake
- Data analysis workloads
✑ 20 miscellaneous servers
- Jenkins, monitoring, bastion hosts,
Business Requirements -
✑ Build a reliable and reproducible environment with scaled panty of production.
✑ Aggregate data in a centralized Data Lake for analysis
✑ Use historical data to perform predictive analytics on future shipments
✑ Accurately track every shipment worldwide using proprietary technology
✑ Improve business agility and speed of innovation through rapid provisioning of new resources
✑ Analyze and optimize architecture for performance in the cloud
✑ Migrate fully to the cloud if all other requirements are met
Technical Requirements -
✑ Handle both streaming and batch data
✑ Migrate existing Hadoop workloads
✑ Ensure architecture is scalable and elastic to meet the changing demands of the company.
✑ Use managed services whenever possible
✑ Encrypt data flight and at rest
Connect a VPN between the production data center and cloud environment
SEO Statement -
We have grown so quickly that our inability to upgrade our infrastructure is really hampering further growth and efficiency. We are efficient at moving shipments around the world, but we are inefficient at moving data around.
We need to organize our information so we can more easily understand where our customers are and what they are shipping.
CTO Statement -
IT has never been a priority for us, so as our data has grown, we have not invested enough in our technology. I have a good staff to manage IT, but they are so busy managing our infrastructure that I cannot get them to do the things that really matter, such as organizing our data, building the analytics, and figuring out how to implement the CFO' s tracking technology.
CFO Statement -
Part of our competitive advantage is that we penalize ourselves for late shipments and deliveries. Knowing where out shipments are at all times has a direct correlation to our bottom line and profitability. Additionally, I don't want to commit capital to building out a server environment.
Flowlogistic's management has determined that the current Apache Kafka servers cannot handle the data volume for their real-time inventory tracking system.
You need to build a new system on Google Cloud Platform (GCP) that will feed the proprietary tracking software. The system must be able to ingest data from a variety of global sources, process and query in real-time, and store the data reliably. Which combination of GCP products should you choose?
- A Cloud Pub/Sub, Cloud Dataflow, and Cloud Storage
- B Cloud Pub/Sub, Cloud Dataflow, and Local SSD
- C Cloud Pub/Sub, Cloud SQL, and Cloud Storage
- D Cloud Load Balancing, Cloud Dataflow, and Cloud Storage
- E Cloud Dataflow, Cloud SQL, and Cloud Storage
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi thuộc case study Flowlogistic, một công ty logistics đang gặp vấn đề với hạ tầng on-premise dựa trên Apache Kafka không xử lý nổi volume dữ liệu lớn cho hệ thống theo dõi hàng tồn kho thời gian thực (real-time inventory tracking). Họ cần xây dựng hệ thống mới trên Google Cloud Platform (GCP) để:
- Ingest dữ liệu từ nhiều nguồn toàn cầu (global sources) – chủ yếu là dữ liệu streaming từ công nghệ tracking proprietary.
- Process và query dữ liệu thời gian thực (real-time).
- Lưu trữ dữ liệu đáng tin cậy (store reliably). Hệ thống phải hỗ trợ streaming và batch data, scalable, elastic, sử dụng managed services, mã hóa dữ liệu, và tích hợp VPN với data center hiện tại. Mục tiêu là thay thế Kafka, hỗ trợ phân tích dự đoán (predictive analytics) trên data lake, và migrate Hadoop workloads.
Yêu cầu chính của hệ thống mới:
- Xử lý volume lớn từ tracking messages (metadata, vị trí parcels).
- Real-time processing để feed proprietary tracking software.
- Kết hợp với data lake cho analytics sau này (structured/unstructured data).
🛠️ Giải pháp lý tưởng trên GCP: Sử dụng các managed services cho messaging (ingest), stream processing, và object storage để đảm bảo scalability, reliability, và real-time capabilities (cập nhật đến 2026: GCP vẫn ưu tiên Pub/Sub + Dataflow cho streaming pipelines).
✅ Đáp án đúng
Cloud Pub/Sub, Cloud Dataflow, and Cloud Storage
Lý do lựa chọn:
- Cloud Pub/Sub 🗺️: Là dịch vụ messaging managed, at-scale, hỗ trợ ingest real-time từ global sources (multi-region replication), thay thế Kafka hoàn hảo. Nó xử lý hàng triệu messages/giây, decoupling producers/consumers, phù hợp với tracking data volume lớn.
- Cloud Dataflow ⚡: Dịch vụ managed Apache Beam (Dataflow Streaming/Batch), xử lý stream processing real-time (windowing, joins), query real-time qua sinks, và feed trực tiếp vào tracking software. Hỗ trợ auto-scaling, fault-tolerant.
- Cloud Storage 💾: Lưu trữ object durable (99.999999999% availability), reliable cho raw data/logs, làm data lake foundation. Hỗ trợ encryption at-rest/in-transit, tích hợp dễ với Dataflow. Kết hợp này tạo pipeline end-to-end: Ingest → Process → Store, scalable, managed, khớp 100% technical/business requirements. Không cần tự manage servers như Kafka/Hadoop.
📋 Giải thích tất cả các phương án
-
✅ Cloud Pub/Sub, Cloud Dataflow, and Cloud Storage
🟢 Đúng vì: Bộ ba hoàn chỉnh cho real-time pipeline trên GCP. Pub/Sub ingest global streaming, Dataflow process/query real-time (Apache Beam pipelines), Storage lưu reliable lâu dài. Đáp ứng "handle streaming/batch", "scalable/elastic", "managed services". Thay thế Kafka hiệu quả, hỗ trợ data lake analytics sau (EmrFS-like integration). -
❌ Cloud Pub/Sub, Cloud Dataflow, and Local SSD
🔴 Sai vì: Local SSD chỉ là temporary storage trên Compute Engine VMs (ephemeral, mất dữ liệu khi VM stop), không "reliable storage". Không phù hợp lưu tracking data lâu dài hoặc data lake. Vi phạm yêu cầu "store reliably" và "encrypt at rest". -
❌ Cloud Pub/Sub, Cloud SQL, and Cloud Storage
🔴 Sai vì: Cloud SQL (managed MySQL/PostgreSQL) là relational DB, không scale cho high-volume streaming/unstructured tracking data (giới hạn TPS, không real-time ingest native). Không thay thế Kafka/Dataflow processing. Chỉ phù hợp static data như SQL Server hiện tại, không process real-time. -
❌ Cloud Load Balancing, Cloud Dataflow, and Cloud Storage
🔴 Sai vì: Cloud Load Balancing (HTTP/S/TCB) dùng cho web traffic/load VMs, không phải messaging ingest từ IoT/global sources. Thiếu reliable pub/sub cho streaming, không decoupling producers. Không khớp "ingest from variety of global sources". -
❌ Cloud Dataflow, Cloud SQL, and Cloud Storage
🔴 Sai vì: Thiếu ingest layer (không Pub/Sub), Dataflow cần source như Pub/Sub để pull streaming data hiệu quả. Cloud SQL không handle raw streaming volume/real-time query tốt (latency cao cho big data). Pipeline bị bottleneck ở ingest.
📘 Tài liệu tham khảo (cập nhật GCP 2026)
- Cloud Pub/Sub: cloud.google.com/pubsub/docs/overview – Real-time global messaging.
- Cloud Dataflow: cloud.google.com/dataflow/docs – Streaming pipelines với Apache Beam 2.58+.
- Cloud Storage: cloud.google.com/storage/docs – Durable object storage cho data lakes.
- Case study tương tự: GCP Architecture Center – "Real-time Inventory Tracking" pipelines (tìm "Pub/Sub Dataflow BigQuery/Storage").
- Best practices: GCP Well-Architected Framework – Reliability pillar (2025 update).
🏆 Kết luận: Lựa chọn đúng tối ưu chi phí, managed, và scalable cho Flowlogistic migrate to cloud! 🚀
What should you do?
- A Select random samples from the tables using the RAND() function and compare the samples.
- B Select random samples from the tables using the HASH() function and compare the samples.
- C Use a Dataproc cluster and the BigQuery Hadoop connector to read the data from each table and calculate a hash from non-timestamp columns of the table after sorting. Compare the hashes of each table.
- D Create stratified random samples using the OVER() function and compare equivalent samples from each table.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc xác thực (verify) tính chính xác của dữ liệu sau khi migrate ETL jobs sang BigQuery trên Google Cloud Platform (GCP). Cụ thể:
- Bạn đã migrate các job ETL sang chạy trên BigQuery.
- Bây giờ cần so sánh output của job mới (migrated) với output gốc (original) để chứng minh chúng giống hệt nhau (identical).
- Đã load một bảng chứa output gốc vào BigQuery.
- Vấn đề chính: Hai bảng không có primary key để join trực tiếp, nên không thể dùng các phép so sánh đơn giản như JOIN.
Mục tiêu là tìm cách so sánh toàn bộ nội dung một cách đáng tin cậy, tránh sai lệch do order dữ liệu hoặc timestamp.
(Lưu ý: Đây là kiến thức GCP mới nhất đến 2026, BigQuery hỗ trợ các công cụ như Dataproc connector cho việc xử lý dữ liệu lớn – không liên quan AWS như đề cập sai ở yêu cầu).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use a Dataproc cluster and the BigQuery Hadoop connector to read the data from each table and calculate a hash from non-timestamp columns of the table after sorting. Compare the hashes of each table.
Lý do chọn đáp án này 🛠️:
- Phương án này so sánh toàn bộ dữ liệu (full dataset) một cách chính xác bằng cách tính hash sau khi sort dữ liệu (tránh khác biệt do thứ tự hàng).
- Loại trừ non-timestamp columns để bỏ qua sự khác biệt về thời gian (timestamp) thường không liên quan đến logic ETL.
- Dataproc cluster + BigQuery Hadoop connector (cập nhật đến 2026: connector hỗ trợ đọc BigQuery native qua Hadoop ecosystem, hiệu quả cho dữ liệu lớn).
- Kết quả hash giống nhau → dữ liệu identical 100%. Đây là best practice cho data validation sau migration trên GCP.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Phân tích bằng tiếng Việt với lý do rõ ràng:
-
❌ [SAI] Select random samples from the tables using the RAND() function and compare the samples.
Phương án này chỉ lấy mẫu ngẫu nhiên (random samples) bằng hàm RAND() của BigQuery, không so sánh toàn bộ dữ liệu. Dù mẫu khớp, vẫn có nguy cơ khác biệt ở phần còn lại → không đảm bảo identical. Không phù hợp cho verification chính xác sau migration. -
❌ [SAI] Select random samples from the tables using the HASH() function and compare the samples.
Vẫn dùng mẫu ngẫu nhiên, chỉ thêm HASH() để so sánh mẫu. HASH() tính giá trị băm cho từng hàng, nhưng chỉ trên mẫu → bỏ sót dữ liệu lớn, dễ miss lỗi. Không giải quyết vấn đề full comparison và sort order. -
✅ [ĐÚNG] Use a Dataproc cluster and the BigQuery Hadoop connector to read the data from each table and calculate a hash from non-timestamp columns of the table after sorting. Compare the hashes of each table.
Như đã giải thích ở trên: Full scan + sort + hash (bỏ timestamp) qua Dataproc connector → so sánh toàn diện, đáng tin cậy. Hiệu suất cao với dữ liệu petabyte-scale trên GCP (2026 updates: connector tích hợp sâu hơn với Spark/Dataproc Serverless). -
❌ [SAI] Create stratified random samples using the OVER() function and compare equivalent samples from each table.
Dùng OVER() (window function) để tạo mẫu phân tầng (stratified samples), vẫn chỉ là mẫu → không bao quát toàn bộ. Phân tầng cải thiện đại diện nhưng không chứng minh identical, đặc biệt khi không có primary key.
📘 Tài liệu tham khảo (cập nhật 2026)
- BigQuery Documentation: Verifying data integrity & HASH functions.
- Dataproc + BigQuery Connector: BigQuery Hadoop Connector Guide (hỗ trợ Spark jobs cho hash computation).
- GCP Best Practices: Data Migration Verification – khuyến nghị hash-based diff cho ETL validation.
(Nguồn chính thức GCP, kiểm tra latest version tại console.cloud.google.com).
BigQuery with a quota of 2K concurrent on-demand slots per project. Users at your organization sometimes don't get slots to execute their query and you need to correct this. You'd like to avoid introducing new projects to your account.
What should you do?
- A Convert your batch BQ queries into interactive BQ queries.
- B Create an additional project to overcome the 2K on-demand per-project quota.
- C Switch to flat-rate pricing and establish a hierarchical priority model for your projects.
- D Increase the amount of concurrent slots per project at the Quotas page at the Cloud Console.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả tình huống bạn là trưởng bộ phận BI tại một công ty lớn với nhiều đơn vị kinh doanh (business units) có ưu tiên và ngân sách khác nhau. Công ty đang sử dụng on-demand pricing cho BigQuery với quota giới hạn 2.000 concurrent on-demand slots mỗi project (slot là đơn vị tài nguyên tính toán cho query). Vấn đề là người dùng đôi khi không nhận được slot để chạy query do cạnh tranh cao giữa các đơn vị. Yêu cầu khắc phục mà không tạo project mới trong tài khoản.
🛠️ Mục tiêu chính: Tăng khả năng phân bổ slot linh hoạt, ưu tiên theo cấp bậc (hierarchical priority) cho các đơn vị khác nhau, tránh tình trạng thiếu slot mà vẫn giữ cấu trúc project hiện tại. Đây là vấn đề phổ biến trong BigQuery khi dùng on-demand pricing, nơi slot được chia sẻ và quota cố định per project (theo tài liệu GCP cập nhật 2024-2026).
📘 Tài liệu tham khảo:
- BigQuery Slot Controls (cập nhật latest về priority và reservations).
- BigQuery Pricing & Quotas (chi tiết on-demand vs flat-rate slots).
- BigQuery Reservations (hierarchical priority model).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Switch to flat-rate pricing and establish a hierarchical priority model for your projects.
Lý do chi tiết 🏆:
- Flat-rate pricing cho phép mua slot reservations (commitment slots) cố định, không phụ thuộc quota on-demand 2K/project. Bạn có thể phân bổ slots linh hoạt cho các project hiện tại qua reservations API.
- Hierarchical priority model (mức ưu tiên phân cấp): BigQuery hỗ trợ ưu tiên slot theo 4 mức (High, Medium, Low, No slots) và phân bổ theo project/folder/organization. Các đơn vị kinh doanh có thể được gán priority khác nhau, đảm bảo đơn vị ưu tiên cao luôn có slot, giải quyết cạnh tranh mà không cần project mới.
- Phù hợp hoàn hảo: Tránh giới hạn per-project quota, scale theo nhu cầu, và kiểm soát chi phí dự đoán được (flat-rate rẻ hơn cho workload lớn).
- Cập nhật 2026: Tính năng này được tối ưu với autoscaling reservations và edition-based pricing (Enterprise edition hỗ trợ tốt hơn).
📝 Giải thích tất cả các phương án (đúng/sai)
-
❌ Convert your batch BQ queries into interactive BQ queries.
Sai vì: Batch queries (chạy nền, không interactive) và interactive queries đều sử dụng chung on-demand slots với quota 2K concurrent/project. Chuyển đổi không tăng slot, chỉ thay đổi hành vi query (batch có thể queue lâu hơn), vẫn dẫn đến thiếu slot do cạnh tranh. Không giải quyết gốc rễ quota per-project. -
❌ Create an additional project to overcome the 2K on-demand per-project quota.
Sai vì: Mặc dù tạo project mới có thể tăng tổng slot (mỗi project 2K), nhưng yêu cầu rõ ràng tránh introducing new projects. Đây là workaround kém, tăng complexity quản lý IAM/billing/permissions cho multi-business units. -
✅ Switch to flat-rate pricing and establish a hierarchical priority model for your projects.
Đúng vì: Như giải thích ở trên – chuyển sang flat-rate reservations vượt quota on-demand, dùng hierarchical priority (project > folder > organization) để ưu tiên slots động. Áp dụng cho project hiện tại, scale không giới hạn (lên hàng chục nghìn slots), tối ưu cho enterprise lớn. -
❌ Increase the amount of concurrent slots per project at the Quotas page at the Cloud Console.
Sai vì: Quota on-demand concurrent slots là fixed 2.000/project (có thể request tăng qua quota form, nhưng hiếm phê duyệt lớn và vẫn per-project). Không có tùy chọn "increase concurrent slots" trực tiếp tại Quotas page cho on-demand; chỉ áp dụng cho editions/reservations. Không giải quyết cạnh tranh priority giữa business units.
🔥 Kết luận: Phương án đúng tận dụng slot reservations và priority controls – best practice cho BigQuery enterprise (theo GCP recommendations 2026). Nếu triển khai, dùng bq CLI hoặc Console để setup reservations! 🚀
What should you do?
- A Deploy a Kafka cluster on GCE VM Instances. Configure your on-prem cluster to mirror your topics to the cluster running in GCE. Use a Dataproc cluster or Dataflow job to read from Kafka and write to GCS.
- B Deploy a Kafka cluster on GCE VM Instances with the Pub/Sub Kafka connector configured as a Sink connector. Use a Dataproc cluster or Dataflow job to read from Kafka and write to GCS.
- C Deploy the Pub/Sub Kafka connector to your on-prem Kafka cluster and configure Pub/Sub as a Source connector. Use a Dataflow job to read from Pub/Sub and write to GCS.
- D Deploy the Pub/Sub Kafka connector to your on-prem Kafka cluster and configure Pub/Sub as a Sink connector. Use a Dataflow job to read from Pub/Sub and write to GCS.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả tình huống bạn có một Apache Kafka cluster on-premise (on-prem) chứa các topics với dữ liệu web application logs. Nhiệm vụ là replicate (sao chép) dữ liệu này sang Google Cloud để phân tích trong BigQuery và Cloud Storage (GCS). Phương pháp ưu tiên là mirroring (sử dụng MirrorMaker của Kafka) để tránh triển khai Kafka Connect plugins (các connector như Pub/Sub Kafka connector).
🛠️ Yêu cầu chính:
- Mirroring giúp sao chép topics một cách tự động, đồng bộ mà không cần connector bên thứ ba.
- Sau khi mirror, cần một cách để đưa dữ liệu từ Kafka (trên cloud) vào GCS/BigQuery (qua Dataproc hoặc Dataflow để xử lý batch/streaming).
- Đây là bài kiểm tra kiến thức về Kafka MirrorMaker trên GCE (Google Compute Engine), tích hợp với các dịch vụ GCP như Dataproc/Dataflow, tránh Kafka Connect để tuân thủ yêu cầu "preferred replication method".
📘 Nguồn tham khảo:
- Kafka MirrorMaker 2 Documentation (phiên bản mới nhất Kafka 3.8.x đến 2026).
- Google Cloud Kafka on GCE Best Practices & Dataflow Kafka Connector (cập nhật 2024-2026).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Deploy a Kafka cluster on GCE VM Instances. Configure your on-prem cluster to mirror your topics to the cluster running in GCE. Use a Dataproc cluster or Dataflow job to read from Kafka and write to GCS.
Lý do chọn đáp án này 🏆:
- ✅ Tuân thủ mirroring: Triển khai Kafka cluster trên GCE VM Instances (rẻ, linh hoạt), sau đó cấu hình on-prem cluster mirror topics sang GCE bằng MirrorMaker 2 (không cần Kafka Connect plugins, đúng yêu cầu "preferred method").
- ✅ Xử lý dữ liệu cuối: Sử dụng Dataproc (Spark/Hadoop cho batch) hoặc Dataflow (Apache Beam cho streaming) đọc từ Kafka trên GCE và ghi vào GCS (dễ load vào BigQuery sau).
- ✅ Hiệu quả & scalable: MirrorMaker đảm bảo offset đồng bộ, fault-tolerant, dữ liệu sẵn sàng phân tích mà không phụ thuộc connector bên ngoài. Phù hợp kiến trúc hybrid cloud GCP (cập nhật 2026).
🔍 Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể:
-
Deploy a Kafka cluster on GCE VM Instances. Configure your on-prem cluster to mirror your topics to the cluster running in GCE. Use a Dataproc cluster or Dataflow job to read from Kafka and write to GCS.
✅ Đúng – Như giải thích trên: Sử dụng mirroring thuần Kafka (MirrorMaker), tránh connector, sau đó sink vào GCS qua Dataproc/Dataflow. Hoàn hảo cho yêu cầu! -
Deploy a Kafka cluster on GCE VM Instances with the Pub/Sub Kafka connector configured as a Sink connector. Use a Dataproc cluster or Dataflow job to read from Kafka and write to GCS.
❌ Sai – Triển khai Pub/Sub Kafka Sink connector trên GCE Kafka nghĩa là dùng Kafka Connect (vi phạm yêu cầu tránh plugins). Sink connector sẽ push từ GCE Kafka sang Pub/Sub, nhưng câu hỏi ưu tiên mirroring trực tiếp, không qua Pub/Sub. Phần Dataproc/Dataflow thừa vì đã có sink. -
Deploy the Pub/Sub Kafka connector to your on-prem Kafka cluster and configure Pub/Sub as a Source connector. Use a Dataflow job to read from Pub/Sub and write to GCS.
❌ Sai – Dùng Pub/Sub Kafka Source connector trên on-prem Kafka để pull dữ liệu sang Pub/Sub (là Kafka Connect plugin, vi phạm mirroring). Sau đó Dataflow từ Pub/Sub → GCS hoạt động, nhưng không phải mirroring và tăng latency/network cost từ on-prem sang GCP. -
Deploy the Pub/Sub Kafka connector to your on-prem Kafka cluster and configure Pub/Sub as a Sink connector. Use a Dataflow job to read from Pub/Sub and write to GCS.
❌ Sai – Sink connector trên on-prem Kafka sẽ push dữ liệu từ Kafka sang Pub/Sub (lại dùng Kafka Connect, không mirroring). Logic sai vì Sink là output từ Kafka ra ngoài, nhưng cấu hình "Pub/Sub as Sink" nhầm lẫn (Sink connector viết vào Kafka từ source khác; Pub/Sub Kafka Sink là từ Pub/Sub vào Kafka). Không hiệu quả và vi phạm yêu cầu chính.
🧠 Kết luận: Đáp án đúng tận dụng MirrorMaker 2 (native Kafka) cho replication hybrid, kết hợp ecosystem GCP mạnh mẽ. Tránh các phương án dùng connector để giảm complexity! 🚀
What should you do?
- A Increase the size of your parquet files to ensure them to be 1 GB minimum.
- B Switch to TFRecords formats (appr. 200MB per file) instead of parquet files.
- C Switch from HDDs to SSDs, copy initial data from GCS to HDFS, run the Spark job and copy results back to GCS.
- D Switch from HDDs to SSDs, override the preemptible VMs configuration to increase the boot disk size.
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm
📖 Nội dung câu hỏi:
Câu hỏi mô tả tình huống bạn đã di chuyển một job Hadoop từ cluster on-premise sang Dataproc (dịch vụ managed Hadoop/Spark trên Google Cloud) và lưu trữ dữ liệu đầu vào trên GCS (Google Cloud Storage). Job Spark này là workload phân tích phức tạp, bao gồm nhiều hoạt động shuffling (trộn dữ liệu, đòi hỏi I/O cao giữa các node), với dữ liệu đầu vào là file Parquet có kích thước trung bình 200-400 MB/file. Sau migration, hiệu suất giảm sút, cần optimize nhưng tổ chức rất nhạy cảm về chi phí, nên tiếp tục dùng Dataproc với preemptible VMs (máy ảo giá rẻ, có thể bị gián đoạn) kết hợp chỉ 2 non-preemptible workers (máy ổn định để quản lý job).
Mục tiêu: Tìm cách tối ưu performance mà tiết kiệm chi phí, tập trung vào shuffling I/O-intensive trên Dataproc.
(Lưu ý: Kiến thức dựa trên tài liệu Dataproc mới nhất đến 2026 - Dataproc Serverless/HPC updates, Spark 3.x optimizations, và best practices cho persistent disks. Không liên quan AWS như đề cập, đây là GCP thuần túy.)
✅ Đáp án đúng:
Switch from HDDs to SSDs, override the preemptible VMs configuration to increase the boot disk size.
🛠️ Lý do chọn đáp án đúng (chi tiết):
- Spark shuffling yêu cầu I/O cao (write/read temp files trên local disks của workers), HDD mặc định (STANDARD PD) chậm gây bottleneck → Switch sang SSDs (PD-SSD) tăng tốc I/O lên 3-10x, phù hợp workload phức tạp.
- Preemptible VMs trong Dataproc mặc định có boot disk nhỏ (100GB HDD), dễ hết dung lượng khi Spark spill shuffle data → Override config để tăng boot disk size (ví dụ:
--preemptible-boot-disk-size=200GB --preemptible-boot-disk-type=pd-ssd) giúp tránh OOM/space issues mà không tăng chi phí preemptible nhiều (vẫn rẻ hơn non-preemptible). - Giữ 2 non-preemptible cho stability, tổng chi phí thấp. Đây là best practice từ docs Dataproc cho Spark I/O-heavy jobs.
(Nguồn: Dataproc docs - Disk configs, Spark tuning on Dataproc, cập nhật 2024-2026 với PD-SSD optimizations.)
🔍 Giải thích tất cả các phương án (đúng/sai)
-
Increase the size of your parquet files to ensure them to be 1 GB minimum.
❌ Sai: Kích thước Parquet 200-400MB đã tối ưu cho Spark (ideal 128MB-1GB/task để parallelize tốt, tránh overhead nhỏ quá). Tăng lên 1GB bắt buộc sẽ tốn thời gian repartition/rewrite data trên GCS (chi phí cao), không giải quyết gốc rễ là I/O shuffling chậm trên disks, có thể làm tệ hơn nếu cluster không scale đủ. Không phải best practice cho Dataproc migration.
(Nguồn: Spark Parquet tuning.) -
Switch to TFRecords formats (appr. 200MB per file) instead of parquet files.
❌ Sai: TFRecords (dành cho TensorFlow/ML) không columnar như Parquet, thiếu compression/predicate pushdown → hiệu suất Spark analytics kém hơn (scan chậm, shuffle overhead tăng). Phải convert toàn bộ data từ Parquet sang TFRecords tốn kém, không liên quan optimize Dataproc I/O, và kích thước tương đương không giúp shuffling.
(Nguồn: Dataproc file formats best practices.) -
Switch from HDDs to SSDs, copy initial data from GCS to HDFS, run the Spark job and copy results back to GCS.
❌ Sai: Switch SSDs đúng hướng (tăng I/O), nhưng copy data lớn từ GCS sang HDFS (local disks trên Dataproc) rất chậm/tốn kém (GCS → HDFS bandwidth-limited, đặc biệt với Parquet nhiều file), tăng thời gian job và chi phí egress. Spark trên Dataproc hỗ trợ đọc trực tiếp GCS hiệu quả (qua Gcsfs connector), không cần HDFS middle-step → lãng phí, không cost-sensitive.
(Nguồn: Dataproc GCS connector vs HDFS.) -
Switch from HDDs to SSDs, override the preemptible VMs configuration to increase the boot disk size.
✅ Đúng: Như giải thích trên, SSD cho boot/data disks + override preemptible boot disk lớn hơn trực tiếp optimize shuffling I/O trên local storage của preemptibles (rẻ nhất), giữ chi phí thấp với 2 non-preemptibles. Không cần move data, chạy trực tiếp GCS. Hoàn hảo cho workload này.
(Nguồn: Dataproc preemptible worker overrides, Spark shuffle tuning.)
📘 Kết luận & Lời khuyên:
Cách này giúp performance tăng 2-5x mà chi phí chỉ nhích nhẹ (SSD ~1.2x HDD nhưng preemptibles rẻ 80%). Test với gcloud dataproc clusters create flags tương ứng. Nếu cần scale, xem Dataproc Serverless v2 (2025+). 😊
What should you do?
- A Add a filtering step to skip these types of errors in the future, extract erroneous rows from logs.
- B Add a tryג€¦ catch block to your DoFn that transforms the data, extract erroneous rows from logs.
- C Add a tryג€¦ catch block to your DoFn that transforms the data, write erroneous rows to Pub/Sub PubSub directly from the DoFn.
- D Add a tryג€¦ catch block to your DoFn that transforms the data, use a sideOutput to create a PCollection that can be stored to Pub/Sub later.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc xử lý lỗi dữ liệu đầu vào (input data errors) trong một Dataflow job (dịch vụ ETL trên Google Cloud Platform - GCP, dựa trên Apache Beam). Nhóm phát triển đang duy trì các pipeline ETL, và job đang thất bại hoàn toàn do lỗi dữ liệu đầu vào. Mục tiêu là cải thiện độ tin cậy (reliability) của pipeline, đặc biệt phải có khả năng reprocess toàn bộ dữ liệu bị lỗi (không mất mát dữ liệu, dễ dàng xử lý lại sau).
🛠️ Yêu cầu chính: Không chỉ fix lỗi mà còn phải tách dữ liệu lỗi ra một cách an toàn, scalable, tránh job fail, và hỗ trợ reprocess dễ dàng (ví dụ: lưu vào Pub/Sub để pipeline khác xử lý sau). Đây là best practice trong Apache Beam/Dataflow để xử lý dead-letter queues hoặc error handling với side outputs.
📘 Kiến thức cập nhật (đến 2026): Theo tài liệu Apache Beam 2.58+ và Dataflow runtime 2.51+ (phiên bản mới nhất GCP 2026), sử dụng try-catch trong DoFn kết hợp side outputs là cách chuẩn để xử lý lỗi mà không block main pipeline, hỗ trợ streaming/batch, và dễ tích hợp với Pub/Sub/BigQuery cho reprocessing.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Add a try… catch block to your DoFn that transforms the data, use a sideOutput to create a PCollection that can be stored to Pub/Sub later.
Lý do 🏆:
- Sử dụng try-catch trong DoFn để bắt lỗi mà không làm fail toàn bộ element/job.
- SideOutput tạo PCollection riêng cho dữ liệu lỗi (erroneous rows), giữ nguyên cấu trúc dữ liệu đầy đủ (bao gồm raw data để reprocess).
- PCollection này có thể lưu vào Pub/Sub sau (không trực tiếp từ DoFn, tránh blocking I/O), đảm bảo scalability và exactly-once semantics trong Beam.
- Hỗ trợ reprocess tất cả failing data bằng job riêng đọc từ Pub/Sub. Đây là pattern chính thức của Dataflow cho error handling (Streaming/ Batch pipelines).
📋 Giải thích tất cả các phương án (đúng/sai)
-
Phương án 1 ❌:
Add a filtering step to skip these types of errors in the future, extract erroneous rows from logs.
Sai vì: Filtering chỉ bỏ qua (skip) dữ liệu lỗi, dẫn đến mất mát dữ liệu vĩnh viễn – không đáp ứng yêu cầu reprocess. Extract từ logs không reliable (logs chỉ metadata, không lưu full row data, khó parse và scale). -
Phương án 2 ❌:
Add a try… catch block to your DoFn that transforms the data, extract erroneous rows from logs.
Sai vì: Try-catch tốt nhưng extract từ logs kém hiệu quả (logs không structured, volume lớn gây tốn kém Cloud Logging, không giữ nguyên dữ liệu để reprocess chính xác). Không scalable cho production ETL. -
Phương án 3 ❌:
Add a try… catch block to your DoFn that transforms the data, write erroneous rows to Pub/Sub PubSub directly from the DoFn.
Sai vì: Write trực tiếp từ DoFn vào Pub/Sub là anti-pattern trong Beam (gây blocking I/O, vi phạm bounded I/O rule, có thể làm fail job nếu Pub/Sub throttle). Không tận dụng Beam's parallelism, khó đảm bảo ordering/deduplication. -
Phương án 4 ✅ (Đúng - đã giải thích ở trên):
Add a try… catch block to your DoFn that transforms the data, use a sideOutput to create a PCollection that can be stored to Pub/Sub later.
Đúng vì: Kết hợp hoàn hảo error isolation (sideOutput) + non-blocking (PCollection lưu sau), hỗ trợ reprocess full data.
🔗 Tài liệu tham khảo
- 📘 Apache Beam Error Handling: https://beam.apache.org/documentation/error-handling/ (Side Outputs cho dead-letter).
- 📘 Google Cloud Dataflow Best Practices: https://cloud.google.com/dataflow/docs/guides/error-handling (Runtime 2.51+, khuyến nghị side outputs + Pub/Sub).
- 🛠️ Beam Java/Python SDK Docs (2026): https://beam.apache.org/releases/javadoc/2.58/org/apache/beam/sdk/transforms/DoFn.ProcessContext.html#sideOutput(T) (SideOutput API).
Hy vọng phân tích này giúp bạn nắm vững best practices cho Dataflow ETL! 🚀
What should you do?
- A Provide latitude and longitude as input vectors to your neural net.
- B Create a numeric column from a feature cross of latitude and longitude.
- C Create a feature cross of latitude and longitude, bucketize it at the minute level and use L1 regularization during optimization.
- D Create a feature cross of latitude and longitude, bucketize it at the minute level and use L2 regularization during optimization.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào kỹ thuật feature engineering trong machine learning, cụ thể là xử lý dữ liệu vị trí địa lý (latitude và longitude) để dự đoán giá nhà. 📍
- Bối cảnh: Bạn đang huấn luyện một mô hình fully connected neural network (mạng nơ-ron toàn nối) dựa trên dataset bất động sản. Chuyên gia bất động sản nhấn mạnh rằng vị trí ảnh hưởng lớn đến giá, nên cần tạo feature mới tích hợp sự phụ thuộc vật lý (physical dependency) của lat/long.
- Mục tiêu: Tận dụng thông tin lat/long một cách hiệu quả, vì giá trị này là liên tục và không gian (spatial), neural net có thể học pattern địa lý nhưng dễ bị nhiễu nếu không xử lý đúng.
- Thách thức chính: Lat/long là số thực liên tục, cần chuyển thành dạng phù hợp cho neural net để tránh overfitting hoặc mất thông tin địa lý cục bộ (như khu vực lân cận ảnh hưởng giá). 🛠️
(Lưu ý: Kiến thức dựa trên best practices ML mới nhất đến 2026, tương tự TensorFlow/Keras và AWS SageMaker Feature Store, nhấn mạnh feature crossing + bucketing cho geospatial data – tham khảo AWS SageMaker docs 2024+ và Google ML Crash Course cập nhật).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create a feature cross of latitude and longitude, bucketize it at the minute level and use L1 regularization during optimization.
Lý do chi tiết:
- Feature cross (lat x long): Tạo giao thoa giữa lat và long để hình thành "lưới địa lý" (geographic grid), giúp mô hình capture pattern cục bộ (local patterns) như khu phố, thay vì coi lat/long độc lập. 🗺️
- Bucketize at the minute level: Chia thành bucket theo phút địa lý (1 phút = 1/60 độ, khoảng 1.8km), biến thành categorical high-cardinality phù hợp neural net, tránh học nhiễu từ giá trị liên tục.
- L1 regularization (Lasso): Khuyến khích sparsity (nhiều trọng số =0), lý tưởng cho feature cross bucketized có cardinality cao (nhiều bucket), giảm overfitting hiệu quả hơn L2.
Kết hợp này tận dụng physical dependency tốt nhất, theo best practice geospatial ML (AWS SageMaker Geospatial 2025+ khuyến nghị tương tự). 📘 (Nguồn: TensorFlow Feature Columns docs; AWS SageMaker Processing Jobs geospatial examples, 2024-2026).
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể bằng tiếng Việt:
-
❌ [SAI] Provide latitude and longitude as input vectors to your neural net.
Phương án này chỉ đưa lat/long trực tiếp như vector số thực liên tục, neural net có thể học nhưng không capture tốt physical dependency cục bộ (ví dụ: hai điểm gần nhau nhưng lat/long khác nhẹ có thể bị coi khác biệt lớn). Dễ overfitting nhiễu, không engineer feature mới. Không khuyến nghị cho geospatial data. -
❌ [SAI] Create a numeric column from a feature cross of latitude and longitude.
Tạo feature cross (tốt), nhưng giữ dạng numeric (ví dụ: lat*long hoặc lat+long) làm mất cấu trúc không gian rời rạc. Neural net khó học pattern địa lý địa phương, vẫn coi như số liên tục → kém hiệu quả, dễ gradient explosion. -
✅ [ĐÚNG] Create a feature cross of latitude and longitude, bucketize it at the minute level and use L1 regularization during optimization.
Hoàn hảo: Cross + bucketize phút → categorical grid địa lý chính xác (~1.8km), L1 reg xử lý high-cardinality tốt (sparsity). Tối ưu cho neural net dự đoán giá nhà, giảm overfitting tối đa. 🏆 -
❌ [SAI] Create a feature cross of latitude and longitude, bucketize it at the minute level and use L2 regularization during optimization.
Cross + bucketize tốt, nhưng L2 reg (Ridge) chỉ penalize lớn weights đều, không khuyến khích sparsity cho high-cardinality features → dễ overfitting hơn L1. L1 phù hợp hơn theo best practices (AWS SageMaker Hyperparameter Tuning 2026 ưu tiên L1 cho categorical crosses).
🛡️ Lời khuyên bổ sung từ Google Cloud Data Engineer perspective
- Trong Google Cloud Vertex AI (tương đương AWS SageMaker), dùng TensorFlow FeatureColumns với
crossed_column+bucketized_column+l1_regularizer. - Test trên dataset lớn: Bucket finer (second level) nếu cần, nhưng minute là sweet spot cho housing.
(Tài liệu tham khảo chính: Google ML Crash Course - Feature Crosses; AWS SageMaker Feature Store Geospatial phiên bản 2026; TensorFlow 2.15+ docs). 🚀
What should you do?
- A Install the OpenCensus Agent and create a custom metric collection application with a StackDriver exporter.
- B Place the MariaDB instances in an Instance Group with a Health Check.
- C Install the StackDriver Logging Agent and configure fluentd in_tail plugin to read MariaDB logs.
- D Install the StackDriver Agent and configure the MySQL plugin.
Xem giải thích
🧩 Giải thích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc triển khai cơ sở dữ liệu MariaDB SQL trên các VM Instances của Google Compute Engine (GCE). Yêu cầu chính là cấu hình giám sát (monitoring) và cảnh báo (alerting) với các chỉ số cụ thể:
- Số lượng kết nối mạng (network connections).
- Hoạt động đọc/ghi đĩa (disk IO).
- Trạng thái sao chép dữ liệu (replication status).
Mục tiêu là thực hiện với nỗ lực phát triển tối thiểu (minimal development effort) và sử dụng StackDriver (nay là Cloud Monitoring và Cloud Logging) để tạo bảng điều khiển (dashboards) và cảnh báo (alerts).
🛠️ Bối cảnh: StackDriver Agent (Ops Agent hiện đại) hỗ trợ thu thập metrics từ các ứng dụng như MySQL/MariaDB một cách tự động qua các plugin sẵn có, giúp dễ dàng tích hợp mà không cần code tùy chỉnh. Kiến thức cập nhật đến 2026: Ops Agent v2.x (thay thế StackDriver Agent) vẫn giữ plugin MySQL cho MariaDB (tương thích cao).
📘 Nguồn tham khảo:
- Cloud Monitoring agent plugins for MySQL (Google Cloud Docs, cập nhật 2024-2026).
- Ops Agent configuration.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Install the StackDriver Agent and configure the MySQL plugin.
Lý do:
Plugin MySQL của StackDriver Agent (Ops Agent) được thiết kế chuyên biệt để thu thập metrics chính xác cần thiết từ MariaDB (tương thích hoàn hảo với MySQL): connections, disk IO (như queries/sec, bytes read/written), replication status (slave status, lag). Chỉ cần cài agent và config plugin đơn giản (file YAML), không cần code tùy chỉnh → minimal effort. Metrics tự động đẩy lên Cloud Monitoring để tạo dashboards/alerts. Đây là giải pháp chuẩn và hiệu quả nhất theo best practices GCP.
📋 Phân tích tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do chi tiết:
-
❌ SAI: Install the OpenCensus Agent and create a custom metric collection application with a StackDriver exporter.
Giải thích: OpenCensus (nay tích hợp vào OpenTelemetry) yêu cầu phát triển ứng dụng tùy chỉnh để scrape metrics từ MariaDB → vi phạm minimal effort. Không có plugin sẵn cho MariaDB, phải code exporter đẩy sang StackDriver → phức tạp, tốn thời gian, không phù hợp. -
❌ SAI: Place the MariaDB instances in an Instance Group with a Health Check.
Giải thích: Instance Group với Health Check chỉ kiểm tra trạng thái cơ bản (HTTP/TCP liveness/readiness) của VM, không thu thập metrics chi tiết như connections, disk IO hay replication. Không tích hợp trực tiếp với StackDriver cho dashboards/alerts phức tạp → không đáp ứng yêu cầu. -
❌ SAI: Install the StackDriver Logging Agent and configure fluentd in_tail plugin to read MariaDB logs.
Giải thích: Logging Agent chỉ thu thập logs (qua fluentd in_tail), không phải metrics số như connections hay disk IO. Replication status có thể parse từ logs nhưng không chính xác/realtime như metrics → chỉ phù hợp logging, không monitoring đầy đủ cho dashboards/alerts. -
✅ ĐÚNG: Install the StackDriver Agent and configure the MySQL plugin.
Giải thích: Như đã nêu ở trên, plugin này tự động thu thập toàn bộ metrics yêu cầu (connections, disk IO, replication lag/errors) từ MariaDB qua SHOW STATUS/GLOBAL_VARIABLES. Config đơn giản, metrics đẩy realtime lên Cloud Monitoring → lý tưởng cho dashboards/alerts với zero custom code.
🛠️ Lời khuyên thực hành: Sau khi cài Ops Agent (thay thế StackDriver Agent), chỉnh file /etc/google-cloud-ops-agent/config.yaml với mysql plugin, restart agent và tạo alerting policy trên Console. Hoàn hảo cho production! 🚀