Ngân hàng đề — Google Cloud Professional Data Engineer
Tìm thấy 429 câu.
Furthermore, data from Cloud Storage buckets and BigQuery datasets must be shared for use in other projects in an ad hoc way. You want to simplify access control management by minimizing the number of policies. Which two steps should you take? (Choose two.)
- A Use Cloud Deployment Manager to automate access provision.
- B Introduce resource hierarchy to leverage access control policy inheritance.
- C Create distinct groups for various teams, and specify groups in Cloud IAM policies.
- D Only use service accounts when sharing data for Cloud Storage buckets and BigQuery datasets.
- E For each Cloud Storage bucket or BigQuery dataset, decide which projects need access. Find all the active members who have access to these projects, and create a Cloud IAM policy to grant access to all these users.
Xem giải thích
📋 Phân tích câu hỏi trắc nghiệm GCP: Quản lý Access Control cho Projects và Dữ liệu Chia sẻ
🧩 Giải thích nội dung câu hỏi một cách chi tiết và rõ ràng:
Câu hỏi mô tả tình huống tổ chức đang mở rộng sử dụng Google Cloud Platform (GCP), với nhiều team tự tạo projects riêng biệt. Số lượng projects tăng nhanh để hỗ trợ các giai đoạn triển khai (như dev, staging, prod) và đối tượng người dùng khác nhau. Mỗi project đòi hỏi cấu hình access control (quyền truy cập) độc lập. Đội ngũ central IT cần quyền truy cập toàn bộ projects.
Ngoài ra, dữ liệu từ Cloud Storage buckets và BigQuery datasets cần được chia sẻ ad hoc (tạm thời, linh hoạt) giữa các projects khác nhau.
Mục tiêu: Đơn giản hóa quản lý access control bằng cách giảm thiểu số lượng policies (IAM policies). Câu hỏi yêu cầu chọn hai bước phù hợp nhất để đạt được điều này, dựa trên các best practices của GCP IAM (Identity and Access Management) và Resource Manager (phiên bản cập nhật đến 2026, với hỗ trợ Resource Hierarchy và Cloud Identity Groups).
✅ Đáp án đúng và lý do lựa chọn:
Hai đáp án đúng là:
- Introduce resource hierarchy to leverage access control policy inheritance.
- Create distinct groups for various teams, and specify groups in Cloud IAM policies.
Lý do chọn (theo kiến thức GCP mới nhất 2026):
🛠️ Introduce resource hierarchy...: Sử dụng cấu trúc phân cấp tài nguyên (Organization → Folders → Projects) cho phép IAM policies kế thừa từ cấp cao hơn xuống cấp thấp hơn (inheritance). Central IT chỉ cần gán quyền tại Organization/Folder level (ví dụ: roles/viewer hoặc roles/owner), tự động áp dụng cho tất cả projects con, giảm số policies cần tạo thủ công. Điều này lý tưởng cho central IT truy cập toàn bộ và quản lý projects phân tán.
🛠️ Create distinct groups...: Tạo Google Groups hoặc Cloud Identity Groups cho từng team (ví dụ: group-dev@org.com), sau đó gán quyền IAM cho group thay vì từng user cá nhân. Một policy duy nhất cho group có thể cover hàng trăm users, giảm thiểu số lượng bindings và dễ quản lý khi teams thay đổi (thêm/xóa members chỉ tại group level). Kết hợp với hierarchy, đây là cách tối ưu để share data ad hoc giữa projects mà không cần policies riêng lẻ.
Hai bước này trực tiếp minimize policies theo nguyên tắc principle of least privilege và centralized management trong GCP IAM v2 (cập nhật 2024-2026).
📘 Phân tích tất cả các phương án (đúng/sai):
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Phần giải thích sử dụng kiến thức GCP cập nhật đến 2026:
-
❌ [SAI] Use Cloud Deployment Manager to automate access provision.
Cloud Deployment Manager dùng để tự động hóa triển khai infrastructure (như tạo VMs, networks) qua templates YAML. Nó không xử lý IAM policies hoặc access control trực tiếp, chỉ gián tiếp qua custom resources. Sử dụng nó sẽ không giảm số policies mà còn phức tạp hóa, không phù hợp cho quản lý access động/ad hoc. -
✅ [ĐÚNG] Introduce resource hierarchy to leverage access control policy inheritance.
Như giải thích trên, resource hierarchy (Organization/Folders/Projects) cho phép policy inheritance từ parent xuống child, giúp central IT gán quyền một lần cho toàn bộ. Đây là best practice chính thức, giảm policies từ hàng trăm xuống chỉ vài chục (ví dụ: gán tại Folder cho dev/staging/prod). -
✅ [ĐÚNG] Create distinct groups for various teams, and specify groups in Cloud IAM policies.
Tạo groups (qua Cloud Identity hoặc Google Workspace) và bind vào IAM policies giúp quản lý tập trung members, dễ scale cho teams lớn. Một policy cho group thay vì per-user, hỗ trợ share data giữa projects mà không cần duplicate bindings. -
❌ [SAI] Only use service accounts when sharing data for Cloud Storage buckets and BigQuery datasets.
Service accounts phù hợp cho machine-to-machine hoặc cross-project access (qua IAM hoặc domain-wide delegation), nhưng không phải cách duy nhất và không linh hoạt cho ad hoc human access. Nó yêu cầu tạo/manage keys phức tạp, không giảm policies tổng thể và bỏ qua user/group-based sharing (như bucket policies hoặc dataset IAM). -
❌ [SAI] For each Cloud Storage bucket or BigQuery dataset, decide which projects need access. Find all the active members who have access to these projects, and create a Cloud IAM policy to grant access to all these users.
Cách này tăng số policies khổng lồ (per bucket/dataset, liệt kê từng user từ projects), vi phạm mục tiêu minimize. Nó thủ công, dễ lỗi (users thay đổi liên tục), và không tận dụng inheritance hay groups – trái ngược best practices GCP.
🔗 Tài liệu tham khảo (cập nhật 2026):
- 📖 GCP Resource Hierarchy & Inheritance (Google Cloud Docs).
- 📖 IAM Best Practices & Groups (bao gồm Cloud Identity Groups v2).
- 📖 Cross-project Data Sharing & BigQuery Authorized Views/Datasets.
- 🎓 Google Cloud Professional Data Engineer Exam Guide (phần IAM & Resource Management).
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần thêm ví dụ thực tế, hãy hỏi nhé.
✑ Single global endpoint
✑ ANSI SQL support
✑ Consistent access to the most up-to-date data
What should you do?
- A Implement BigQuery with no region selected for storage or processing.
- B Implement Cloud Spanner with the leader in North America and read-only replicas in Asia and Europe.
- C Implement Cloud SQL for PostgreSQL with the master in North America and read replicas in Asia and Europe.
- D Implement Bigtable with the primary cluster in North America and secondary clusters in Asia and Europe.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một công ty có trụ sở tại Mỹ phát triển ứng dụng đánh giá và phản hồi hành động người dùng, với lượng dữ liệu chính trong bảng tăng 250.000 bản ghi/giây (rất cao, đòi hỏi hệ thống OLTP chịu tải lớn). Nhiều bên thứ ba sử dụng API của ứng dụng để tích hợp vào frontend của họ. Các yêu cầu chính cho API:
- Single global endpoint: Một điểm cuối toàn cầu duy nhất, dễ truy cập từ mọi nơi mà không cần routing phức tạp. 📍
- ANSI SQL support: Hỗ trợ chuẩn SQL ANSI đầy đủ, phù hợp cho query phức tạp và giao dịch. 🔤
- Consistent access to the most up-to-date data: Truy cập nhất quán vào dữ liệu mới nhất (strong consistency), không chấp nhận dữ liệu cũ hoặc eventual consistency. ⚡
Mục tiêu: Chọn dịch vụ lưu trữ/database phù hợp trên Google Cloud để đáp ứng throughput cao, tính toàn cầu, SQL chuẩn và consistency mạnh. Tốc độ tăng dữ liệu cực lớn loại trừ các hệ thống batch/OLAP thuần túy. 🛠️
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Implement Cloud Spanner with the leader in North America and read-only replicas in Asia and Europe.
Lý do:
- Cloud Spanner là distributed SQL database toàn cầu của Google Cloud, hỗ trợ ANSI SQL chuẩn (fully compliant), single global endpoint qua cấu hình multi-region. 📘
- Với leader ở North America (gần trụ sở công ty) và read-only replicas ở Asia/Europe, nó đảm bảo strong consistency (dữ liệu up-to-date nhất quán toàn cầu nhờ TrueTime và Paxos replication). Hỗ trợ throughput cao lên đến hàng triệu QPS và 250k writes/giây dễ dàng (theo docs GCP 2024-2026). 🌍
- Phù hợp hoàn hảo cho workload OLTP real-time cao, API global với latency thấp. Không có dịch vụ nào khác đáp ứng tất cả 3 yêu cầu cùng lúc. 🚀
Nguồn tham khảo:
- Google Cloud Spanner Documentation (cập nhật 2025: Multi-region configs với leader/replicas).
- Spanner Global Distribution (throughput & consistency).
📋 Giải thích tất cả các phương án (đúng/sai)
-
[SAI] Implement BigQuery with no region selected for storage or processing.
❌ Sai vì: BigQuery là OLAP analytics warehouse (columnar storage), không hỗ trợ ANSI SQL đầy đủ cho transactional workloads (chỉ query analytics, không ACID transactions thực thụ). Không có "single global endpoint" cho writes real-time; multi-region chỉ cho storage/query, nhưng ingestion streaming giới hạn ~1MB/s/table, không chịu nổi 250k records/giây với consistency mạnh (eventual consistency cho append-only). Phù hợp batch analytics, không OLTP API. 📊 -
[ĐÚNG] Implement Cloud Spanner with the leader in North America and read-only replicas in Asia and Europe.
✅ Đúng như phân tích ở trên: Hoàn hảo với global distribution, SQL ANSI, strong consistency và scale cao. Leader NA giảm latency writes, replicas đọc toàn cầu. 🌐 -
[SAI] Implement Cloud SQL for PostgreSQL with the master in North America and read replicas in Asia and Europe.
❌ Sai vì: Cloud SQL là managed relational DB regional (PostgreSQL/MySQL), không có single global endpoint thực sự (read replicas cross-region có latency cao, eventual consistency cho reads, không strong global). Không scale tự động đến 250k writes/giây (giới hạn ~65k IOPS/instance). Yêu cầu cross-region replicas thủ công, không native global như Spanner. 🗺️ -
[SAI] Implement Bigtable with the primary cluster in North America and secondary clusters in Asia and Europe.
❌ Sai vì: Bigtable là NoSQL wide-column store (HBase-compatible), không hỗ trợ ANSI SQL (chỉ API NoSQL, cần connector riêng như BigQuery federation nhưng không native). Replication multi-cluster là eventual consistency (không "most up-to-date" nhất quán), phù hợp analytics/IoT cao throughput nhưng không cho SQL queries phức tạp hoặc transactional API global. ⚙️
Kết luận: Cloud Spanner là lựa chọn tối ưu cho workload global OLTP SQL high-throughput trên GCP (cập nhật 2026). Nếu cần tư vấn thêm, hãy hỏi! 💡
- A Add a WHERE clause to the query, and grant the BigQuery Data Viewer role to the application service account.
- B Create an Authorized View with the provided query. Share the dataset that contains the view with the application service account.
- C Create a Dataflow pipeline using BigQueryIO to read results from the query. Grant the Dataflow Worker role to the application service account.
- D Create a Dataflow pipeline using BigQueryIO to read predictions for all users from the query. Write the results to Bigtable using BigtableIO. Grant the Bigtable Reader role to the application service account so that the application can read predictions for individual users from Bigtable.
Xem giải thích
🧩 Phân tích câu hỏi trắc nghiệm
📘 Nội dung câu hỏi được giải thích chi tiết:
Câu hỏi xoay quanh việc xây dựng một ML pipeline để phục vụ dự đoán (predictions) từ mô hình BigQuery ML mà data scientist đã tạo. Ứng dụng là một REST API cần phục vụ dự đoán cho một user ID cá nhân với độ trễ (latency) dưới 100 milliseconds. Query được cung cấp là:SELECT predicted_label, user_id FROM ML.PREDICT (MODEL 'dataset.model', table user_features').
🛠️ Vấn đề cốt lõi:
- Hàm
ML.PREDICTtrong BigQuery thường mất thời gian để xử lý (scan dữ liệu và chạy mô hình), không phù hợp cho real-time inference với latency thấp (<100ms). - Yêu cầu là tạo pipeline để tối ưu hóa, cho phép API query nhanh chóng cho từng user ID mà không phải chạy query BigQuery mỗi lần (vì BigQuery không phải là low-latency store cho single lookups).
- Giải pháp cần pre-compute predictions cho tất cả users và lưu vào store có độ trễ thấp như Bigtable (NoSQL key-value store của GCP, hỗ trợ <10ms latency cho lookups).
(Kiến thức cập nhật đến 2026: BigQuery ML hỗ trợ batch predictions tốt, Dataflow v2+ với Apache Beam 2.52+ tích hợp BigQueryIO và BigtableIO mượt mà cho streaming/batch pipelines. Bigtable hỗ trợ single-row reads <5ms với proper schema – theo GCP docs 2025).
✅ Đáp án đúng:
Create a Dataflow pipeline using BigQueryIO to read predictions for all users from the query. Write the results to Bigtable using BigtableIO. Grant the Bigtable Reader role to the application service account so that the application can read predictions for individual users from Bigtable.
🟢 Lý do chọn đáp án này (chi tiết):
- Dataflow pipeline chạy batch job định kỳ (hoặc triggered) để thực thi query
ML.PREDICTcho tất cả users một lần (pre-compute), tránh chạy query real-time. - BigQueryIO đọc kết quả query từ BigQuery một cách hiệu quả.
- BigtableIO ghi predictions vào Bigtable với key là
user_id(cho phép lookup O(1) siêu nhanh <100ms, thường <10ms). - Bigtable Reader role cho service account của app, API chỉ cần query Bigtable bằng
user_id→ latency thấp. - Đây là best practice cho online serving từ BigQuery ML (theo pattern "batch predict → low-latency store").
📚 Tài liệu tham khảo:
- BigQuery ML PREDICT (GCP Docs 2025).
- Dataflow BigQueryIO & BigtableIO (Apache Beam 2.52, 2026).
- Bigtable Low-Latency Serving (GCP Best Practices 2025).
🔍 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] Add a WHERE clause to the query, and grant the BigQuery Data Viewer role to the application service account.
ThêmWHERE user_id = ?vào query và cấp quyền Data Viewer cho app.
Lý do sai: Mỗi request API vẫn phải chạyML.PREDICTvới filter → BigQuery scan full table/mô hình mỗi lần, latency thường >1s (không dưới 100ms). Không giải quyết vấn đề real-time, chỉ là workaround kém hiệu quả. -
❌ [SAI] Create an Authorized View với the provided query. Share the dataset that contains the view with the application service account.
Tạo Authorized View chứa query và share dataset cho service account.
Lý do sai: View trong BigQuery là logical view, query vẫn chạy on-the-fly mỗi lần truy vấn → latency cao giống hệt query gốc (ML.PREDICT chậm cho single row). Authorized View chỉ kiểm soát quyền, không tối ưu performance. -
❌ [SAI] Create a Dataflow pipeline using BigQueryIO to read results from the query. Grant the Dataflow Worker role to the application service account.
Dataflow đọc kết quả query qua BigQueryIO, cấp Dataflow Worker role cho app.
Lý do sai: Pipeline chỉ đọc từ BigQuery, nhưng không lưu trữ kết quả vào store low-latency. App vẫn phải trigger Dataflow hoặc query BigQuery gián tiếp → không đạt <100ms (Dataflow Worker role không liên quan đến app serving, chỉ cho pipeline run). -
✅ [ĐÚNG] Create a Dataflow pipeline using BigQueryIO to read predictions for all users from the query. Write the results to Bigtable using BigtableIO. Grant the Bigtable Reader role to the application service account so that the application can read predictions for individual users from Bigtable.
(Đã giải thích chi tiết ở phần đáp án đúng ở trên – hoàn hảo cho low-latency serving! 🚀)
Consumers will receive the data in the following ways:
✑ Real-time event stream
✑ ANSI SQL access to real-time stream and historical data
✑ Batch historical exports
Which solution should you use?
- A Cloud Dataflow, Cloud SQL, Cloud Spanner
- B Cloud Pub/Sub, Cloud Storage, BigQuery
- C Cloud Dataproc, Cloud Dataflow, BigQuery
- D Cloud Pub/Sub, Cloud Dataproc, Cloud SQL
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một ứng dụng xây dựng để chia sẻ dữ liệu thị trường tài chính thời gian thực với người dùng cuối (consumers). Dữ liệu được thu thập từ thị trường thời gian thực (real-time). Người dùng sẽ nhận dữ liệu qua ba cách chính:
- Real-time event stream: Dữ liệu được truyền trực tiếp dưới dạng luồng sự kiện thời gian thực (như streaming data feeds).
- ANSI SQL access to real-time stream and historical data: Truy cập dữ liệu thời gian thực và lịch sử qua ngôn ngữ SQL chuẩn ANSI (hỗ trợ query linh hoạt trên cả dữ liệu mới và cũ).
- Batch historical exports: Xuất dữ liệu lịch sử theo lô (batch), thường lưu trữ để tải xuống hoặc xử lý sau.
📌 Yêu cầu giải pháp: Cần một hệ thống Google Cloud (GCP) tích hợp để xử lý ingestion thời gian thực, query SQL chuẩn, và xuất batch, với khả năng scale lớn cho dữ liệu tài chính cao tần suất (high-velocity data). Kiến thức dựa trên phiên bản GCP mới nhất đến 2026, nơi BigQuery hỗ trợ streaming inserts nhanh chóng (sub-second latency) và ANSI SQL đầy đủ.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Cloud Pub/Sub, Cloud Storage, BigQuery
🛠️ Lý do chi tiết:
- Cloud Pub/Sub ✅: Dịch vụ messaging quản lý hoàn toàn cho real-time event stream, hỗ trợ publish/subscribe với độ trễ thấp (<100ms), scale đến hàng triệu message/giây – lý tưởng cho dữ liệu thị trường tài chính real-time.
- BigQuery ✅: Data warehouse serverless hỗ trợ ANSI SQL chuẩn để query real-time stream (qua streaming inserts từ Pub/Sub) và historical data (lưu trữ columnar). Hỗ trợ federated queries và materialized views cho hiệu suất cao.
- Cloud Storage ✅: Lưu trữ object rẻ tiền cho batch historical exports từ BigQuery (export jobs nhanh chóng dưới dạng CSV/Avro/Parquet).
🔗 Tích hợp hoàn hảo: Pub/Sub → Dataflow/BigQuery streaming → Storage exports. Đây là blueprint chuẩn GCP cho streaming analytics (xem GCP Data Streaming best practices 2024-2026).
Nguồn tham khảo:
- Cloud Pub/Sub docs
- BigQuery streaming inserts
- BigQuery exports to Storage
- GCP Architecture: Financial Services Data Pipeline
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn (giữ nguyên văn bản gốc tiếng Anh). Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), với lý do cụ thể dựa trên yêu cầu câu hỏi:
-
Cloud Dataflow, Cloud SQL, Cloud Spanner ❌
❌ Sai vì: Cloud Dataflow chỉ là stream/batch processing (Apache Beam), không phải dịch vụ messaging gốc cho real-time event stream (thiếu pub/sub native). Cloud SQL (MySQL/PostgreSQL) và Cloud Spanner (distributed SQL DB) hỗ trợ ANSI SQL nhưng không scale cho real-time high-velocity data (giới hạn throughput, chi phí cao), và thiếu batch exports native. Không khớp full requirements. -
Cloud Pub/Sub, Cloud Storage, BigQuery ✅
✅ Đúng hoàn hảo (như giải thích ở trên). Pub/Sub lo stream, BigQuery lo SQL real-time/historical, Storage lo batch exports. Tích hợp seamless, chi phí tối ưu, scale vô hạn. -
Cloud Dataproc, Cloud Dataflow, BigQuery ❌
❌ Sai vì: Cloud Dataproc (managed Hadoop/Spark) dành cho batch processing lớn, không hỗ trợ real-time event stream (latency cao, không pub/sub). Dataflow bổ trợ processing nhưng thừa thãi và thiếu storage exports trực tiếp. BigQuery tốt nhưng combo không cover real-time ingestion gốc. -
Cloud Pub/Sub, Cloud Dataproc, Cloud SQL ❌
❌ Sai vì: Cloud Pub/Sub tốt cho stream ✅, nhưng Cloud Dataproc chỉ batch (không SQL real-time/historical hiệu quả), Cloud SQL là relational DB không phù hợp big data streaming (scale kém, ANSI SQL hạn chế so BigQuery, thiếu batch exports columnar). Không hỗ trợ query historical data lớn quy mô tài chính.
Kết luận tổng quát 🎯: Giải pháp đúng tận dụng serverless-native GCP services cho real-time + analytics + storage, tránh managed clusters tốn kém như Dataproc/Spanner. Đây là pattern chuẩn cho fintech workloads trên GCP!
✑ Decoupling producer from consumer
✑ Space and cost-efficient storage of the raw ingested data, which is to be stored indefinitely
✑ Near real-time SQL query
✑ Maintain at least 2 years of historical data, which will be queried with SQL
Which pipeline should you use to meet these requirements?
- A Create an application that provides an API. Write a tool to poll the API and write data to Cloud Storage as gzipped JSON files.
- B Create an application that writes to a Cloud SQL database to store the data. Set up periodic exports of the database to write to Cloud Storage and load into BigQuery.
- C Create an application that publishes events to Cloud Pub/Sub, and create Spark jobs on Cloud Dataproc to convert the JSON data to Avro format, stored on HDFS on Persistent Disk.
- D Create an application that publishes events to Cloud Pub/Sub, and create a Cloud Dataflow pipeline that transforms the JSON event payloads to Avro, writing the data to Cloud Storage and BigQuery.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả tình huống xây dựng một ứng dụng mới cần thu thập dữ liệu một cách có thể mở rộng (scalable). Dữ liệu đến liên tục suốt ngày dưới dạng JSON, dự kiến đạt khoảng 150 GB/ngày vào cuối năm. Các yêu cầu chính bao gồm:
- Tách rời producer khỏi consumer (Decoupling producer from consumer): Producer (ứng dụng tạo dữ liệu) không cần biết consumer (hệ thống xử lý) đang ở đâu.
- Lưu trữ raw data hiệu quả về không gian và chi phí, lưu trữ vô thời hạn (Space and cost-efficient storage of raw ingested data, stored indefinitely): Dữ liệu thô cần nén tốt, chi phí thấp, giữ mãi mãi.
- Truy vấn SQL gần thời gian thực (Near real-time SQL query): Có thể chạy SQL nhanh chóng trên dữ liệu mới.
- Giữ ít nhất 2 năm dữ liệu lịch sử để truy vấn SQL (Maintain at least 2 years of historical data, queried with SQL): Hỗ trợ lưu trữ dài hạn và phân tích SQL.
📊 Tổng quan pipeline cần thiết: Cần một hệ thống streaming decoupling (như Pub/Sub), chuyển đổi định dạng hiệu quả (JSON → Avro để nén), lưu raw data rẻ tiền (Cloud Storage), và kho dữ liệu SQL nhanh (BigQuery). Kiến thức cập nhật đến 2026 vẫn giữ nguyên các dịch vụ cốt lõi GCP như Dataflow (Apache Beam) hỗ trợ streaming real-time, BigQuery với streaming inserts cho near real-time queries (latency <1 phút).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create an application that publishes events to Cloud Pub/Sub, and create a Cloud Dataflow pipeline that transforms the JSON event payloads to Avro, writing the data to Cloud Storage and BigQuery.
Lý do chi tiết:
- 🛤️ Decoupling: Pub/Sub là message queue managed, producer publish events mà không cần biết consumer.
- 💾 Lưu trữ raw data hiệu quả: Dataflow chuyển JSON sang Avro (nén tốt hơn JSON ~70-80%, chi phí thấp trên Cloud Storage, lưu indefinitely với lifecycle rules).
- ⚡ Near real-time SQL: Dataflow streaming pipeline write trực tiếp vào BigQuery (hỗ trợ streaming buffers từ 2023-2026, query latency giây-phút).
- 📈 Lưu trữ lịch sử 2 năm: BigQuery scale petabyte, columnar storage tối ưu SQL queries trên historical data.
Pipeline này fully managed, auto-scale, phù hợp 150GB/day (~1.7MB/s).
📘 Nguồn tham khảo:
- Google Cloud Pub/Sub Docs (decoupling).
- Dataflow Streaming Guide (JSON to Avro, write to Storage/BigQuery).
- BigQuery Streaming Inserts (near real-time, cập nhật 2026).
❌ Giải thích tất cả các phương án
-
[SAI] Create an application that provides an API. Write a tool to poll the API and write data to Cloud Storage as gzipped JSON files.
❌ Sai vì: Không decoupling (phải poll API liên tục, tight coupling, dễ mất dữ liệu nếu tool crash). Gzipped JSON vẫn kém hiệu quả nén so Avro, không hỗ trợ near real-time SQL (chỉ raw Storage, cần ETL riêng). Không giữ historical data cho SQL dễ dàng. Không scalable cho continuous data. -
[SAI] Create an application that writes to a Cloud SQL database to store the data. Set up periodic exports of the database to write to Cloud Storage and load into BigQuery.
❌ Sai vì: Không decoupling (direct write vào Cloud SQL, single point failure). Cloud SQL (OLTP) không scale cho 150GB/day (~52TB/năm), chi phí cao, latency cao. Periodic exports không near real-time (batch delay hàng giờ). Không efficient cho raw storage indefinitely. -
[SAI] Create an application that publishes events to Cloud Pub/Sub, and create Spark jobs on Cloud Dataproc to convert the JSON data to Avro format, stored on HDFS on Persistent Disk.
❌ Sai vì: Decoupling tốt với Pub/Sub, nhưng HDFS trên Persistent Disk kém efficient (không serverless, chi phí cao, không auto-scale như Storage). Không near real-time SQL (không có BigQuery). Dataproc Spark cần manage cluster thủ công, không lý tưởng cho streaming continuous (batch-oriented hơn Dataflow). Không phù hợp lưu indefinitely dài hạn. -
[ĐÚNG] Create an application that publishes events to Cloud Pub/Sub, and create a Cloud Dataflow pipeline that transforms the JSON event payloads to Avro, writing the data to Cloud Storage and BigQuery.
✅ Đúng vì: Hoàn hảo khớp tất cả yêu cầu như phân tích trên. Dataflow (Apache Beam) là lựa chọn best practice cho streaming ETL trên GCP đến 2026, fully managed, fault-tolerant.
🛠️ Kết luận: Pipeline đúng tận dụng sức mạnh GCP streaming ecosystem, tối ưu chi phí (~$0.01/GB processed) và performance!
- A Increase the number of max workers
- B Use a larger instance type for your Dataflow workers
- C Change the zone of your Dataflow pipeline to run in us-central1
- D Create a temporary table in Bigtable that will act as a buffer for new data. Create a new step in your pipeline to write to this table first, and then create a new pipeline to write from Bigtable to BigQuery
- E Create a temporary table in Cloud Spanner that will act as a buffer for new data. Create a new step in your pipeline to write to this table first, and then create a new pipeline to write from Cloud Spanner to BigQuery
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi này thuộc chủ đề Google Cloud Dataflow (không phải AWS như mô tả ban đầu, có thể là nhầm lẫn), tập trung vào việc tối ưu hóa hiệu suất của một pipeline streaming. Cụ thể:
- Pipeline hiện tại: Nhận dữ liệu từ Pub/Sub topic, xử lý và ghi kết quả vào BigQuery dataset nằm ở khu vực EU (multi-region).
- Cấu hình: Chạy ở region europe-west4 (một region EU), giới hạn tối đa 3 workers với loại instance n1-standard-1 (1 vCPU, 3.75 GB RAM).
- Vấn đề: Trong giờ cao điểm (peak periods), pipeline chậm xử lý dữ liệu kịp thời, vì tất cả 3 workers đạt 100% CPU utilization.
- Yêu cầu: Chọn hai hành động để tăng performance (tăng tốc độ xử lý).
Mục tiêu là giải quyết tình trạng CPU bottleneck bằng cách scale pipeline Dataflow một cách hiệu quả. Dataflow hỗ trợ autoscaling tự động, nhưng cần điều chỉnh tham số để xử lý tải cao hơn. (Kiến thức cập nhật đến 2024-2026: Dataflow sử dụng Unified Worker từ 2023, hỗ trợ scale nhanh hơn với Flex Templates và Streaming Engine).
📘 Tài liệu tham khảo:
✅ Đáp án đúng (Chọn hai)
Hai phương án đúng là:
- Increase the number of max workers
- Use a larger instance type for your Dataflow workers
Lý do lựa chọn:
- Pipeline đang bị giới hạn bởi số lượng workers ít (chỉ 3) và CPU yếu (n1-standard-1), dẫn đến overload ở peak.
- Tăng max workers (horizontal scaling): Cho phép Dataflow tự động scale lên nhiều workers hơn (ví dụ:
--maxNumWorkers=10), xử lý song song nhiều records từ Pub/Sub, giảm backlog. Dataflow autoscaling sẽ kích hoạt khi CPU cao. - Dùng instance type lớn hơn (vertical scaling): Nâng cấp lên n1-standard-4 hoặc n2-highcpu-4 (thêm vCPU/RAM), tăng sức mạnh mỗi worker để xử lý nhanh hơn mà không cần nhiều workers. Kết hợp cả hai là tối ưu cho streaming workloads.
- Đây là các hành động trực tiếp, đơn giản từ docs chính thức, không ảnh hưởng latency đến BigQuery EU.
🛠️ Giải thích chi tiết tất cả các phương án
-
✅ Increase the number of max workers
Đúng: Tăng--maxNumWorkerscho phép Dataflow scale ngang, phân tán tải CPU qua nhiều workers hơn. Trong peak, autoscaler sẽ thêm workers tự động (giới hạn mặc định thấp gây bottleneck). Hiệu quả cao cho Pub/Sub streaming, chi phí theo usage. -
✅ Use a larger instance type for your Dataflow workers
Đúng: Thay--machineType=n1-standard-4(hoặc cao hơn như custom-2-4096) tăng CPU/RAM mỗi worker. Giải quyết trực tiếp CPU 100%, xử lý nhanh hơn mà không thay đổi kiến trúc. Dataflow hỗ trợ từ n1/n2/e2 series (cập nhật 2024: ưu tiên Tau workers cho perf tốt hơn). -
❌ Change the zone of your Dataflow pipeline to run in us-central1
Sai: Thay region từ europe-west4 (EU) sang us-central1 (US) sẽ tăng latency đáng kể khi ghi vào BigQuery EU (cross-region traffic chậm 100-200ms+). Dataflow ưu tiên chạy gần data sink/source để giảm network overhead. Không giải quyết CPU issue, chỉ làm chậm hơn (vi phạm best practice multi-region). -
❌ Create a temporary table in Bigtable that will act as a buffer for new data. Create a new step in your pipeline to write to this table first, and then create a new pipeline to write from Bigtable to BigQuery
Sai: Bigtable là NoSQL cho high-throughput OLTP, không phù hợp buffer streaming (thiếu ACID, schema-less gây phức tạp sync với BigQuery). Tạo hai pipelines tăng latency/chi phí (double processing), OOM risk cao. Dataflow ghi trực tiếp BigQuery streaming inserts hiệu quả hơn (hàng triệu rows/sec). -
❌ Create a temporary table in Cloud Spanner that will act as a buffer for new data. Create a new step in your pipeline to write to this table first, and then create a new pipeline to write from Cloud Spanner to BigQuery
Sai: Spanner là distributed SQL cho OLTP global, quá nặng/overkill cho buffer (chi phí cao ~10x BigQuery, throughput giới hạn). Thêm bước trung gian làm pipeline phức tạp, chậm hơn (Spanner write latency 10-50ms). Dataflow có Streaming Inserts native cho BigQuery, không cần buffer ngoài.
🚀 Khuyến nghị bổ sung
- Kết hợp:
--maxNumWorkers=10 --machineType=n1-standard-4 --enableStreamingEngineđể perf tối đa. - Monitor qua Dataflow Monitoring UI hoặc Cloud Monitoring metrics (CPU, backlog elements).
- Test với Dataflow Flex Templates cho streaming updates 2025+.
Hy vọng phân tích giúp bạn ôn thi chứng chỉ! 💪
- A Configure your Dataflow pipeline to use local execution
- B Increase the maximum number of Dataflow workers by setting maxNumWorkers in PipelineOptions
- C Increase the number of nodes in the Bigtable cluster
- D Modify your Dataflow pipeline to use the Flatten transform before writing to Bigtable
- E Modify your Dataflow pipeline to use the CoGroupByKey transform before writing to Bigtable
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một data pipeline sử dụng Dataflow job để tổng hợp (aggregate) và ghi (write) dữ liệu time series metrics vào Bigtable. Vấn đề là dữ liệu cập nhật chậm trong Bigtable, và dữ liệu này cung cấp cho một dashboard được hàng nghìn người dùng truy cập đồng thời. Yêu cầu là hỗ trợ thêm người dùng đồng thời (concurrent users) và giảm thời gian ghi dữ liệu. Đây là câu hỏi trắc nghiệm chọn hai hành động (Choose two) để giải quyết, tập trung vào việc tối ưu hóa hiệu suất của Dataflow và Bigtable trong Google Cloud Platform (GCP).
📘 Bối cảnh kỹ thuật: Dataflow là dịch vụ batch/stream processing dựa trên Apache Beam, Bigtable là NoSQL database columnar scale-out cho dữ liệu lớn. Vấn đề chậm thường do bottleneck ở worker Dataflow (parallelism) hoặc capacity Bigtable (throughput write).
✅ Đáp án đúng và lý do lựa chọn
Hai đáp án đúng là:
- Increase the maximum number of Dataflow workers by setting maxNumWorkers in PipelineOptions
- Increase the number of nodes in the Bigtable cluster
Lý do chọn:
🛠️ Tăng maxNumWorkers trong Dataflow: Dataflow scale tự động, nhưng giới hạn workers mặc định có thể gây bottleneck khi write lớn vào Bigtable. Tăng workers (qua PipelineOptions) tăng parallelism, xử lý và ghi dữ liệu nhanh hơn, hỗ trợ concurrent users tốt hơn. Đây là best practice cho high-throughput pipelines (cập nhật Dataflow 2024+ hỗ trợ autoscaling tốt hơn với Unified Worker).
🛠️ Tăng nodes Bigtable cluster: Bigtable scale horizontally bằng nodes (mỗi node ~1-2k QPS write tùy config). Write chậm do cluster undersized; thêm nodes tăng throughput tổng (lên đến hàng triệu ops/s), giảm latency cho dashboard. Phù hợp time series metrics cao tải (Bigtable replica sets mới 2025+ tối ưu hơn).
Kết hợp hai hành động này giải quyết cả phía producer (Dataflow) và consumer (Bigtable), đảm bảo end-to-end performance.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh:
-
❌ Configure your Dataflow pipeline to use local execution
Phương án này sai vì local execution chạy trên máy local (direct runner), không scale cho production pipeline lớn. Nó thiếu autoscaling workers trên GCP, dẫn đến write chậm hơn, không hỗ trợ concurrent users. Chỉ dùng cho testing/debug, không phải giải pháp cho high-load Bigtable writes. -
✅ Increase the maximum number of Dataflow workers by setting maxNumWorkers in PipelineOptions
Phương án này đúng như giải thích trên: Tăng workers nâng cao parallelism Dataflow, đẩy nhanh aggregation và write batch vào Bigtable. Config qua--maxNumWorkershoặc PipelineOptions API (Dataflow docs khuyến nghị cho >1TB data/day). -
✅ Increase the number of nodes in the Bigtable cluster
Phương án này đúng như giải thích trên: Bigtable throughput tỷ lệ thuận với nodes (mỗi node tăng ~1000 write QPS). Dùng GCP Console/CLI scale cluster production (hỗ trợ SSD/HDD nodes mới 2026). -
❌ Modify your Dataflow pipeline to use the Flatten transform before writing to Bigtable
Phương án này sai vì Flatten chỉ merge multiple PCollections cùng schema (ví dụ multi-windows), không cải thiện write performance vào Bigtable. Time series metrics đã aggregated, Flatten có thể tạo overhead không cần, làm chậm hơn thay vì nhanh. Không liên quan đến scaling. -
❌ Modify your Dataflow pipeline to use the CoGroupByKey transform before writing to Bigtable
Phương án này sai vì CoGroupByKey dùng để join/group multiple PCollections theo key, tăng shuffle overhead và memory usage ở Dataflow. Với time series đã aggregated, nó làm phức tạp pipeline, giảm throughput write thay vì tăng, gây bottleneck lớn hơn.
📚 Tài liệu tham khảo (cập nhật mới nhất GCP 2026)
- Dataflow scaling: cloud.google.com/dataflow/docs/guides/scaling (maxNumWorkers & autoscaling).
- Bigtable performance: cloud.google.com/bigtable/docs/performance (nodes & throughput).
- Best practices pipelines: cloud.google.com/dataflow/docs/guides/bigtable-io (write optimization).
Lưu ý: Kiến thức dựa trên GCP docs 2024-2026, Bigtable v2+ và Dataflow 2.60+ với Flex Templates.
- A Create a Cloud Dataproc Workflow Template
- B Create an initialization action to execute the jobs
- C Create a Directed Acyclic Graph in Cloud Composer
- D Create a Bash script that uses the Cloud SDK to create a cluster, execute jobs, and then tear down the cluster
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm
✅ Giải thích nội dung câu hỏi:
Câu hỏi mô tả tình huống bạn có nhiều công việc Spark (Spark jobs) chạy trên một cụm Cloud Dataproc (Dataproc cluster) theo lịch trình (on a schedule). Một số job chạy tuần tự (in sequence), nghĩa là phải chờ job trước hoàn thành mới chạy job sau; một số job chạy song song (concurrently), nghĩa là có thể chạy cùng lúc mà không phụ thuộc lẫn nhau. Nhiệm vụ là tự động hóa toàn bộ quy trình này, bao gồm việc lập lịch, quản lý thứ tự thực thi và tính song song. Đây là yêu cầu điển hình về orchestration (điều phối workflow) trong Google Cloud, nơi cần một công cụ hỗ trợ Directed Acyclic Graph (DAG) để mô hình hóa dependencies phức tạp và scheduling tự động. (Kiến thức cập nhật đến 2026: Cloud Dataproc và Cloud Composer vẫn là các dịch vụ cốt lõi cho big data processing trên GCP, với Composer dựa trên Apache Airflow 2.x+ hỗ trợ DAGs động và scaling tốt hơn).
🟢 Đáp án đúng:
Create a Directed Acyclic Graph in Cloud Composer
Lý do lựa chọn: Cloud Composer (dịch vụ managed Apache Airflow trên GCP) là công cụ orchestration lý tưởng để tự động hóa quy trình phức tạp như vậy. Bạn có thể định nghĩa một DAG (Directed Acyclic Graph) để mô hình hóa chính xác các job Spark: các task chạy sequence (qua dependencies >> hoặc <<) và concurrent (qua BranchPythonOperator hoặc parallel tasks). Composer hỗ trợ scheduling tự động qua cron expressions, tích hợp native với Dataproc (qua DataprocSubmitJobOperator), retry logic, monitoring và scaling. Điều này đảm bảo quy trình chạy đáng tin cậy mà không cần can thiệp thủ công. (📘 Tài liệu tham khảo: Cloud Composer DAGs Documentation và Dataproc Operators in Airflow - phiên bản mới nhất 2026).
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ Create a Cloud Dataproc Workflow Template
Phân tích sai: Workflow Templates của Dataproc cho phép định nghĩa graph jobs (hỗ trợ sequence và parallel qua steps với dependencies), nhưng chúng không hỗ trợ scheduling tự động (chỉ submit thủ công hoặc qua API). Để schedule, bạn vẫn cần công cụ ngoài như cron job hoặc Composer, làm quy trình không hoàn toàn tự động. Không phù hợp cho multi-schedule phức tạp. (📘 Tài liệu: Dataproc Workflow Templates - giới hạn không có built-in scheduler). -
❌ Create an initialization action to execute the jobs
Phân tích sai: Initialization actions chỉ chạy một lần lúc khởi tạo cluster (bootstrap script), không hỗ trợ multiple jobs với sequence/concurrent hay scheduling định kỳ. Chúng dành cho setup môi trường (cài package), không phải orchestration workflow. Cluster phải luôn chạy, tốn kém và không linh hoạt. (📘 Tài liệu: Dataproc Init Actions). -
✅ Create a Directed Acyclic Graph in Cloud Composer
Phân tích đúng: Như đã giải thích ở trên, DAG trong Composer hoàn hảo cho dependencies phức tạp (sequence/concurrent), scheduling tự động, và tích hợp seamless với Dataproc jobs. Hỗ trợ error handling, monitoring qua UI, và scale theo nhu cầu (phiên bản Composer 3+ đến 2026 có dynamic DAGs và GKE autoscaling). (🛠️ Ưu điểm: Giảm chi phí bằng cách submit jobs on-demand mà không cần cluster luôn-on). -
❌ Create a Bash script that uses the Cloud SDK to create a cluster, execute jobs, and then tear down the cluster
Phân tích sai: Bash script với gcloud SDK có thể tạo cluster → submit jobs → xóa cluster, nhưng không hỗ trợ tự động hóa đầy đủ: thiếu native sequence/concurrent logic (phải hard-code if/while), không có scheduling (cần cron bên ngoài), dễ lỗi, khó monitor/scale/retry. Không phải giải pháp enterprise cho production. (📘 Tài liệu: gcloud Dataproc Commands - chỉ là CLI cơ bản, không thay thế orchestrator).
🧠 Kết luận: Sử dụng Cloud Composer với DAG là cách best practice trên GCP cho data pipeline orchestration, đặc biệt với Spark trên Dataproc. Nếu cần tùy chỉnh sâu, kết hợp với Workflow Templates bên trong DAG tasks! 🚀
- A Create an API using App Engine to receive and send messages to the applications
- B Use a Cloud Pub/Sub topic to publish jobs, and use subscriptions to execute them
- C Create a table on Cloud SQL, and insert and delete rows with the job information
- D Create a table on Cloud Spanner, and insert and delete rows with the job information
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả tình huống xây dựng một data pipeline mới để chia sẻ dữ liệu (cụ thể là jobs) giữa hai loại ứng dụng: job generators (ứng dụng tạo ra các công việc) và job runners (ứng dụng thực thi các công việc). Yêu cầu chính bao gồm:
- Tự động scale để xử lý tăng trưởng sử dụng (increases in usage).
- Hỗ trợ thêm ứng dụng mới mà không ảnh hưởng hiệu suất của các ứng dụng hiện tại.
Đây là vấn đề kinh điển về decoupled architecture (kiến trúc tách rời), nơi cần một hệ thống trung gian messaging (trao đổi thông điệp) để tránh coupling chặt chẽ giữa producer (generators) và consumer (runners). Giải pháp phải asynchronous, scalable, và multi-tenant friendly (hỗ trợ nhiều subscriber độc lập).
📘 Nguồn tham khảo: Google Cloud Documentation - Pub/Sub Overview (cập nhật 2024, vẫn áp dụng đến 2026): cloud.google.com/pubsub/docs/overview.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use a Cloud Pub/Sub topic to publish jobs, and use subscriptions to execute them
Lý do:
🛠️ Cloud Pub/Sub là dịch vụ publish-subscribe messaging của Google Cloud, được thiết kế chính xác cho các pipeline như thế này.
- Job generators publish (gửi) jobs vào topic (chủ đề chung).
- Job runners subscribe (đăng ký) vào subscriptions riêng biệt để pull hoặc push jobs mà không ảnh hưởng lẫn nhau.
- Scale tự động: Pub/Sub xử lý hàng triệu message/giây, auto-scale theo demand, hỗ trợ at-least-once delivery và dead letter queues cho reliability.
- Thêm app mới: Tạo subscription mới chỉ mất vài giây, không impact existing ones (zero-downtime, decoupled hoàn toàn).
✅ Hoàn hảo khớp yêu cầu, theo best practices GCP cho event-driven architecture (cập nhật Pub/Sub v2 với enhanced features đến 2026).
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể:
-
Create an API using App Engine to receive and send messages to the applications
❌ Sai. App Engine phù hợp cho web apps stateless, nhưng tạo API synchronous sẽ tạo tight coupling giữa generators và runners (phải gọi API trực tiếp). Khi scale, API dễ thành single point of failure/bottleneck (quota limits ~28 instances miễn phí, cần scale manual). Thêm app mới yêu cầu update API logic, ảnh hưởng toàn bộ hệ thống. Không lý tưởng cho high-throughput messaging.
📘 Nguồn: App Engine Quotas (2024): cloud.google.com/appengine/quotas. -
Use a Cloud Pub/Sub topic to publish jobs, and use subscriptions to execute them
✅ Đúng (như đã giải thích ở trên). Đây là golden standard cho decoupled, scalable pipelines ở GCP, hỗ trợ fan-out (nhiều subscribers/topic) và regional/global replication mới nhất (2024+). -
Create a table on Cloud SQL, and insert and delete rows with the job information
❌ Sai. Cloud SQL (MySQL/PostgreSQL) là relational DB, không phải messaging system. Generators insert rows, runners phải poll (query liên tục) để lấy jobs → inefficient (high latency, wasteful CPU/network), dễ thundering herd khi scale. Delete rows gây race conditions và không atomic cho multi-app. Không scale tốt cho bursts (max ~100k IOPS), thêm app mới cần schema changes.
📘 Nguồn: Cloud SQL Best Practices (2024): cloud.google.com/sql/docs/mysql/best-practices. -
Create a table on Cloud Spanner, and insert and delete rows with the job information
❌ Sai. Cloud Spanner là distributed SQL cho global scale transactions, mạnh về consistency (strong ACID), nhưng vẫn yêu cầu polling giống Cloud SQL → không hiệu quả cho queuing/messaging (chi phí cao ~$0.9/node/giờ). Insert/delete lặp lại gây hotspots và throughput limits nếu không sharding thủ công. Thêm app cần coordination, vẫn coupling cao. Pub/Sub rẻ hơn và optimized hơn cho use case này.
📘 Nguồn: Spanner vs Pub/Sub Comparison (2024): cloud.google.com/spanner/docs/compare.
- A The current epoch time
- B A concatenation of the product name and the current epoch time
- C A random universally unique identifier number (version 4 UUID)
- D The original order identification number from the sales system, which is a monotonically increasing integer
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc thiết kế primary key cho một bảng giao dịch mới trong Cloud Spanner (dịch vụ cơ sở dữ liệu phân tán, có khả năng mở rộng toàn cầu của Google Cloud), dùng để lưu trữ dữ liệu bán hàng sản phẩm (product sales data).
📌 Bối cảnh chính: Cloud Spanner sử dụng kiến trúc phân tán với các split (phân vùng dữ liệu) tự động dựa trên primary key. Để đạt hiệu suất tối ưu (performance perspective), primary key phải đảm bảo:
- Phân bố đều các bản ghi mới trên các split (tránh hot spots – nơi một split nhận quá nhiều write operations cùng lúc).
- Hỗ trợ high throughput cho các giao dịch viết (inserts) đồng thời cao, đặc biệt với dữ liệu thời gian thực như sales data.
- Tránh các key có tính monotonic (tăng dần) hoặc predictable (dễ dự đoán), vì chúng gây ra write hotspots ở cuối bảng (tail).
Mục tiêu: Chọn chiến lược primary key giúp tối ưu hóa hiệu suất khi insert dữ liệu lớn, liên tục.
(Lưu ý: Dù người dùng đề cập "AWS", câu hỏi thực tế thuộc Google Cloud Spanner. Tôi sử dụng kiến thức cập nhật đến 2026 từ tài liệu Google Cloud Spanner v2.x, với các best practices về schema design không thay đổi cơ bản).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: A random universally unique identifier number (version 4 UUID)
🛠️ Lý do chi tiết:
- UUID v4 là random 128-bit identifier, được tạo ngẫu nhiên (sử dụng random number generator), đảm bảo phân bố đều trên toàn bộ không gian key.
- Trong Cloud Spanner, điều này giúp tránh hot spots hoàn toàn, vì các insert mới sẽ được phân bổ ngẫu nhiên vào các split khác nhau, hỗ trợ high QPS (queries per second) cho write-heavy workloads như sales transactions.
- Ưu điểm hiệu suất: Giảm latency, tăng throughput lên đến hàng triệu ops/giây. Google khuyến nghị cho bảng transaction tables (xem best practices 2024-2026).
- Dễ implement: Sử dụng hàm
UUID()trong SQL hoặc client libraries (Java, Python, Go).
📘 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn một cách chi tiết, với đánh giá đúng/sai dựa trên nguyên tắc thiết kế key trong Cloud Spanner:
-
❌ [SAI] The current epoch time
Giải thích: Epoch time (thời gian Unix timestamp) là giá trị monotonic tăng dần (ví dụ: 1728000000). Khi nhiều giao dịch insert cùng lúc (common trong sales data), tất cả sẽ có key gần giống nhau → hot spot nghiêm trọng ở split cuối cùng. Gây throttling, cao latency (có thể >100ms/write). Không phù hợp với workload cao. -
❌ [SAI] A concatenation of the product name and the current epoch time
Giải thích: Ghép product name + epoch time (ví dụ: "iPhone_1728000000") vẫn semi-monotonic theo thời gian. Các giao dịch cùng sản phẩm/same time sẽ cluster vào cùng split → per-product hot spots. Không giải quyết vấn đề phân bố đều, dẫn đến imbalance và giảm throughput toàn hệ thống. -
✅ [ĐÚNG] A random universally unique identifier number (version 4 UUID)
Giải thích: Như đã nêu ở trên, random distribution lý tưởng cho Cloud Spanner. Không có pattern dự đoán → splits tự động cân bằng, tối ưu performance cho insert-heavy tables. Hỗ trợ global consistency mà không bottleneck. -
❌ [SAI] The original order identification number from the sales system, which is a monotonically increasing integer
Giải thích: Order ID tăng dần (monotonically increasing, ví dụ: 1,2,3...) gây sequential writes → extreme hot spot ở split chứa key lớn nhất (tail). Cloud Spanner cảnh báo rõ ràng chống lại kiểu key này trong docs, vì throughput write giảm mạnh (chỉ ~100-1000 QPS/split). Phù hợp hơn cho read-only, không phải transaction sales.
📚 Tài liệu tham khảo (cập nhật 2026)
- Google Cloud Spanner Best Practices: Schema Design - Choosing a Primary Key (khuyến nghị UUID cho random keys).
- Spanner Performance Tuning: Avoiding Hot Spots (ví dụ UUID v4).
- Sample Code: Spanner client libraries (v2.20+), hàm
GENERATE_UUID_V4(). - Certification Guide: Google Cloud Professional Data Engineer study guide (2024 edition, nhấn mạnh anti-hotspot strategies).
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần ví dụ code hoặc case study, hãy hỏi thêm.