Ngân hàng đề — Google Cloud Professional Data Engineer
Tìm thấy 429 câu.
Company Overview -
MJTelco is a startup that plans to build networks in rapidly growing, underserved markets around the world. The company has patents for innovative optical communications hardware. Based on these patents, they can create many reliable, high-speed backbone links with inexpensive hardware.
Company Background -
Founded by experienced telecom executives, MJTelco uses technologies originally developed to overcome communications challenges in space. Fundamental to their operation, they need to create a distributed data infrastructure that drives real-time analysis and incorporates machine learning to continuously optimize their topologies. Because their hardware is inexpensive, they plan to overdeploy the network allowing them to account for the impact of dynamic regional politics on location availability and cost.
Their management and operations teams are situated all around the globe creating many-to-many relationship between data consumers and provides in their system. After careful consideration, they decided public cloud is the perfect environment to support their needs.
Solution Concept -
MJTelco is running a successful proof-of-concept (PoC) project in its labs. They have two primary needs:
✑ Scale and harden their PoC to support significantly more data flows generated when they ramp to more than 50,000 installations.
✑ Refine their machine-learning cycles to verify and improve the dynamic models they use to control topology definition.
MJTelco will also use three separate operating environments `" development/test, staging, and production `" to meet the needs of running experiments, deploying new features, and serving production customers.
Business Requirements -
✑ Scale up their production environment with minimal cost, instantiating resources when and where needed in an unpredictable, distributed telecom user community.
✑ Ensure security of their proprietary data to protect their leading-edge machine learning and analysis.
✑ Provide reliable and timely access to data for analysis from distributed research workers
✑ Maintain isolated environments that support rapid iteration of their machine-learning models without affecting their customers.
Technical Requirements -
Ensure secure and efficient transport and storage of telemetry data
✑ Rapidly scale instances to support between 10,000 and 100,000 data providers with multiple flows each.
✑ Allow analysis and presentation against data tables tracking up to 2 years of data storing approximately 100m records/day
✑ Support rapid iteration of monitoring infrastructure focused on awareness of data pipeline problems both in telemetry flows and in production learning cycles.
CEO Statement -
Our business model relies on our patents, analytics and dynamic machine learning. Our inexpensive hardware is organized to be highly reliable, which gives us cost advantages. We need to quickly stabilize our large distributed data pipelines to meet our reliability and capacity commitments.
CTO Statement -
Our public cloud services must operate as advertised. We need resources that scale and keep our data secure. We also need environments in which our data scientists can carefully study and quickly adapt our models. Because we rely on automation to process our data, we also need our development and test environments to work as we iterate.
CFO Statement -
The project is too large for us to maintain the hardware and software required for the data and analysis. Also, we cannot afford to staff an operations team to monitor so many data feeds, so we will rely on automation and infrastructure. Google Cloud's machine learning will allow our quantitative researchers to work on our high-value problems instead of problems with our data pipelines.
MJTelco needs you to create a schema in Google Bigtable that will allow for the historical analysis of the last 2 years of records. Each record that comes in is sent every 15 minutes, and contains a unique identifier of the device and a data record. The most common query is for all the data for a given device for a given day.
Which schema should you use?
- A Rowkey: date#device_id Column data: data_point
- B Rowkey: date Column data: device_id, data_point
- C Rowkey: device_id Column data: date, data_point
- D Rowkey: data_point Column data: device_id, date
- E Rowkey: date#data_point Column data: device_id
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi thuộc case study MJTelco – một startup viễn thông sử dụng Google Cloud để xây dựng hạ tầng dữ liệu phân tán, xử lý dữ liệu telemetry thời gian thực từ các thiết bị (data providers). Họ cần lưu trữ lịch sử 2 năm dữ liệu (~100 triệu records/ngày), với mỗi record được gửi mỗi 15 phút, chứa device_id (định danh thiết bị duy nhất) và data record (dữ liệu đo lường).
Yêu cầu chính cho schema Bigtable:
- Hỗ trợ phân tích lịch sử 2 năm.
- Truy vấn phổ biến nhất (most common query): Lấy tất cả dữ liệu của một thiết bị cụ thể trong một ngày cụ thể (given device for a given day).
- Bigtable là cơ sở dữ liệu NoSQL wide-column, row key quyết định hiệu suất truy vấn (locality, range scan). Dữ liệu lưu trong column qualifiers (dynamic), row key cần sắp xếp để các record liên quan nằm gần nhau, tránh wide rows (quá nhiều cells/row gây chậm), hotspots (ghi đọc tập trung).
Mục tiêu: Schema tối ưu scale (10k-100k thiết bị, multiple flows), secure, rapid iteration ML. Không dùng SQL (Bigtable không hỗ trợ full scan toàn bộ table hiệu quả).
✅ Đáp án đúng: Rowkey: date#device_id Column data: data_point
Lý do chọn (dựa kiến thức Bigtable cập nhật 2026):
- Row key = date#device_id (vd: "2023-01-01#device123"): Tạo một row duy nhất mỗi ngày mỗi thiết bị.
- Column data: data_point: Column family
data_pointchứa giá trị dữ liệu, qualifier = timestamp trong ngày (vd: "12:00:00" → giá trị data_record). Ngày có ~96 records (24h*4), → row chỉ ~96 cells, nhỏ gọn, tránh wide row. - Truy vấn device + day siêu hiệu quả:
GetRow("date#device_id")→ đọc 1 row duy nhất, scan 96 cells theo thứ tự lex (qualifiers sorted). Latency thấp <10ms, scale tốt 100m records/ngày. - Phù hợp best practice Bigtable cho time-series IoT/telemetry (shard theo time + entity): TTL dễ xóa dữ liệu cũ, range scan prefix
date#device_idđếndate#device_id\xff. Không hotspot vì date phân bố đều. - Với 100k thiết bị, ~100k rows/ngày (~73M rows/2 năm), Bigtable xử lý dễ dàng (hàng nghìn tỷ rows OK).
🛠️ Giải thích tất cả các phương án
-
✅ Rowkey: date#device_id Column data: data_point
Phương án tối ưu nhất. Row key time-first + device đảm bảo locality hoàn hảo cho truy vấn phổ biến (1 row/ device-day). Nhỏ gọn (96 cells/row), dễ GC cũ dữ liệu (TTL trên prefix date). Hoàn hảo cho scale và ML analysis. 📈 -
❌ Rowkey: date Column data: device_id, data_point
Row key chỉdate(1 row/ngày toàn table). Columns lưudevice_id:data_point→ overwrite nếu cùng device nhiều records, hoặc wide row khổng lồ (100m cells/ngày → impossible, hotspot nghiêm trọng). Truy vấn device+day cần full scan row → chậm, tốn chi phí. Không scale. 🚫 -
❌ Rowkey: device_id Column data: date, data_point
Row key chỉdevice_id(1 row/thiết bị). Columns lưudate#data_point(qualifier="2023-01-01_12:00:data") → wide row cực lớn (~70k qualifiers/row sau 2 năm, scan chậm >1s/query). Bigtable khuyên tránh >vài nghìn cells/row. Truy vấn OK nhưng performance kém dài hạn, hotspot ghi mới. Không phù hợp historical data lớn. 📉 -
❌ Rowkey: data_point Column data: device_id, date
Row key dựadata_point(giá trị dữ liệu?) → vô nghĩa, không locality (records rải rác theo giá trị data). Truy vấn device+day cần full table scan hoặc nhiều GetRow → không khả thi với 100m records/ngày. Hoàn toàn sai best practice. 💥 -
❌ Rowkey: date#data_point Column data: device_id
Row keydate#data_pointưu tiên time + value → records theo data value group, không liên quan device. Truy vấn device+day: prefix scan rộng (tất cả data_point ngày đó), bỏ sót hoặc chậm. Không hỗ trợ "all data for device". Lãng phí. 🔄
📘 Tài liệu tham khảo (cập nhật 2026)
- Google Cloud Bigtable Schema Design: https://cloud.google.com/bigtable/docs/schema-design#schema-time-series (best practices IoT: time#entity hoặc entity#reversed-time).
- Bigtable Limits: https://cloud.google.com/bigtable/quotas (rows: unlimited thực tế; cells/row <100M nhưng tránh wide rows).
- GCP Professional Data Engineer Exam Guide: Case MJTelco Q7 (examtopics.com hoặc Google practice exams).
- IoT Telemetry Patterns: https://cloud.google.com/solutions/iot/reference-architecture (shard device-day).
Schema này giúp MJTelco đạt business req: scale minimal cost, secure (IAM + CMEK), isolated envs. Nếu cần refine ML, dùng Dataflow + BigQuery export! 🚀
- A Rewrite the job in Pig.
- B Rewrite the job in Apache Spark.
- C Increase the size of the Hadoop cluster.
- D Decrease the size of the Hadoop cluster but also rewrite the job in Hive.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi mô tả tình huống một công ty đang phát triển nhanh chóng, dẫn đến lượng dữ liệu đầu vào (ingesting data) tăng cao hơn trước đây. Bạn đang quản lý các công việc phân tích batch hàng ngày sử dụng MapReduce trên Apache Hadoop. Tuy nhiên, do dữ liệu tăng, các batch job này bị chậm trễ (falling behind). Nhiệm vụ là khuyến nghị cách tăng tính responsive (phản hồi nhanh hơn) của các phân tích mà KHÔNG tăng chi phí (without increasing costs).
🛠️ Điểm cốt lõi: Cần tối ưu hóa hiệu suất xử lý dữ liệu lớn trên Hadoop mà không mở rộng tài nguyên (như cluster size), tập trung vào việc cải thiện công cụ xử lý để chạy nhanh hơn trên cùng hạ tầng hiện tại. Đây là vấn đề phổ biến trong big data, nơi MapReduce truyền thống (disk-based) chậm so với các công nghệ mới hơn (in-memory).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Rewrite the job in Apache Spark.
Lý do: Apache Spark là engine xử lý dữ liệu phân tán in-memory (xử lý trong RAM), nhanh hơn MapReduce (disk-based) tới 100 lần trong các workload memory-bound và 10 lần trong compute-bound (theo benchmark AWS EMR đến 2026). Việc rewrite job từ MapReduce sang Spark giúp tăng tốc độ đáng kể mà không cần tăng kích thước cluster, từ đó giữ nguyên chi phí. AWS EMR (Elastic MapReduce) hỗ trợ Spark native, dễ migrate và scale tự động. Điều này phù hợp hoàn hảo với yêu cầu "tăng responsiveness without increasing costs" – nhanh hơn trên cùng tài nguyên! 🚀
📋 Phân tích tất cả các phương án
-
❌ Rewrite the job in Pig.
Sai vì: Pig là ngôn ngữ scripting cao cấp chạy trên MapReduce, chỉ đơn giản hóa code nhưng không cải thiện tốc độ xử lý cốt lõi (vẫn disk-based, chậm tương đương MapReduce). Không giải quyết được vấn đề batch job falling behind, và không đảm bảo tăng responsiveness mà không tốn thêm chi phí. Pig phù hợp cho ETL nhanh prototype, nhưng không phải giải pháp performance. -
✅ Rewrite the job in Apache Spark.
Đúng vì: Như đã giải thích ở trên, Spark sử dụng RDD (Resilient Distributed Datasets) và DAG (Directed Acyclic Graph) scheduler để xử lý in-memory, tối ưu hóa đáng kể so với MapReduce 2-stage. Trên AWS EMR (phiên bản mới nhất 6.x+ đến 2026), Spark tích hợp seamless với Hadoop ecosystem, hỗ trợ dynamic allocation để tiết kiệm tài nguyên. Kết quả: Job chạy nhanh hơn, responsive cao, chi phí không đổi! 💥 -
❌ Increase the size of the Hadoop cluster.
Sai vì: Tăng kích thước cluster (thêm node) sẽ tăng chi phí trực tiếp (EC2 instances, storage), vi phạm yêu cầu "without increasing costs". Dù có thể tăng throughput, đây chỉ là scale vertically/horizontally thô, không hiệu quả lâu dài và không giải quyết bottleneck của MapReduce chậm. -
❌ Decrease the size of the Hadoop cluster but also rewrite the job in Hive.
Sai vì: Giảm cluster size sẽ làm chậm hơn nữa do ít tài nguyên, còn Hive (trước Hive 4.x với Spark/Tez) chủ yếu dựa trên MapReduce, chỉ cải thiện query SQL-like nhưng không nhanh bằng Spark cho batch jobs phức tạp. Kết hợp "decrease cluster" + Hive sẽ tệ hơn tình trạng hiện tại, không tăng responsiveness mà còn rủi ro failure cao.
📘 Tài liệu tham khảo
- AWS EMR Documentation (2026): Apache Spark on Amazon EMR – So sánh performance Spark vs. MapReduce.
- AWS Big Data Blog: Migrating from Hadoop MapReduce to Spark – Benchmark tăng tốc 100x.
- Apache Spark Official: Spark vs. Hadoop – Lý thuyết in-memory processing.
- Google Cloud tương đương (tham chiếu): Dataproc với Spark cho tương tự EMR.
Hy vọng phân tích này giúp bạn nắm vững! Nếu cần deep dive thêm, hỏi nhé! 🔍
- A Create a view in BigQuery that concatenates the FirstName and LastName field values to produce the FullName.
- B Add a new column called FullName to the Users table. Run an UPDATE statement that updates the FullName column for each user with the concatenation of the FirstName and LastName values.
- C Create a Google Cloud Dataflow job that queries BigQuery for the entire Users table, concatenates the FirstName value and LastName value for each user, and loads the proper values for FirstName, LastName, and FullName into a new table in BigQuery.
- D Use BigQuery to export the data for the table to a CSV file. Create a Google Cloud Dataproc job to process the CSV file and output a new CSV file containing the proper values for FirstName, LastName and FullName. Run a BigQuery load job to load the new CSV file into BigQuery.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả một chuỗi nhà hàng thức ăn nhanh lớn với hơn 400.000 nhân viên, lưu trữ thông tin nhân viên trong bảng Users của Google BigQuery. Bảng này chỉ có hai trường: FirstName và LastName. Một thành viên IT đang xây dựng ứng dụng cần truy vấn trường FullName được tạo bằng cách nối FirstName + khoảng trắng + LastName cho từng nhân viên.
Yêu cầu chính: Sửa schema và dữ liệu trong BigQuery để ứng dụng có thể truy vấn FullName, đồng thời tối ưu hóa chi phí (minimizing cost).
🛠️ Thách thức: Bảng lớn (400k+ rows), cần tránh chi phí lưu trữ/compute cao như cập nhật dữ liệu thực tế hoặc ETL phức tạp. BigQuery tính phí dựa trên dữ liệu quét (query) và lưu trữ, nên giải pháp phải không lưu trữ thêm dữ liệu và chỉ tính phí khi query.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create a view in BigQuery that concatenates the FirstName and LastName field values to produce the FullName.
Lý do:
- View trong BigQuery là logical view (chỉ định nghĩa truy vấn, không lưu trữ dữ liệu vật lý), sử dụng hàm
CONCAT(FirstName, ' ', LastName) AS FullName. - Ứng dụng query view như bảng thật, BigQuery tự động thực thi truy vấn gốc.
- Tối ưu chi phí 💰: Không tốn lưu trữ thêm, chỉ tính phí on-demand query (dựa trên bytes quét, ~$5/TB đầu tiên). Với 400k rows, chi phí rất thấp nếu query không thường xuyên. Không cần ETL hay update dữ liệu gốc.
- Phù hợp phiên bản BigQuery mới nhất (2026): Views hỗ trợ tốt cho transformation đơn giản, có thể kết hợp authorized views cho security.
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do cụ thể:
-
✅ Create a view in BigQuery that concatenates the FirstName and LastName field values to produce the FullName.
Đúng vì: Đây là giải pháp đơn giản, serverless, zero-storage. View chỉ là metadata (query definition), BigQuery compute on-the-fly khi query. Tiết kiệm nhất cho workload read-heavy như app truy vấn. Không thay đổi schema/dữ liệu gốc. -
❌ Add a new column called FullName to the Users table. Run an UPDATE statement that updates the FullName column for each user with the concatenation of the FirstName and LastName values.
Sai vì: ALTER TABLE thêm cột tốn lưu trữ vĩnh viễn (~$0.02/GB/tháng cho 400k rows + FullName string). UPDATE toàn bộ bảng quét toàn bộ dữ liệu (slot-based pricing, có thể >$100 cho table lớn). Không scalable nếu dữ liệu thay đổi thường xuyên (cần re-update). -
❌ Create a Google Cloud Dataflow job that queries BigQuery for the entire Users table, concatenates the FirstName value and LastName value for each user, and loads the proper values for FirstName, LastName, and FullName into a new table in BigQuery.
Sai vì: Dataflow (Apache Beam) là ETL mạnh nhưng overkill cho concat đơn giản. Chi phí cao: Query BigQuery ($5/TB) + Dataflow vCPU/memory (~$0.01-0.1/vCPU-hour) + load mới vào BigQuery (miễn phí load nhưng lưu trữ mới). Tổng chi phí gấp 10-100x view cho job one-time. -
❌ Use BigQuery to export the data for the table to a CSV file. Create a Google Cloud Dataproc job to process the CSV file and output a new CSV file containing the proper values for FirstName, LastName and FullName. Run a BigQuery load job to load the new CSV file into BigQuery.
Sai vì: Quy trình phức tạp nhất, chi phí "khủng": Export BigQuery ($5/TB) + Dataproc cluster (VM + Spark, ~$0.5-2/giờ) + lưu trữ GCS + load back (miễn phí load nhưng lưu trữ mới). Không hiệu quả cho transformation trivial, dễ lỗi với 400k rows.
📘 Tài liệu tham khảo (Cập nhật đến 2026)
- BigQuery Views: cloud.google.com/bigquery/docs/views – Chi tiết tạo view với CONCAT.
- Pricing: cloud.google.com/bigquery/pricing – On-demand storage/query vs. slot-based.
- Best practices: cloud.google.com/bigquery/docs/best-practices-performance – Khuyến nghị views cho derived data để minimize cost.
- Materialized Views (nâng cao, không cần ở đây): cloud.google.com/bigquery/docs/materialized-views-intro – Chỉ dùng nếu query frequent.
🛠️ Khuyến nghị: Nếu query FullName rất thường xuyên (>1TB/tháng), cân nhắc materialized view để cache, nhưng view thường vẫn rẻ nhất!
'tags' have multiple values but the property 'date released' does not. A typical query would ask for all movies with actor=<actorname> ordered by date_released or all movies with tag=Comedy ordered by date_released. How should you avoid a combinatorial explosion in the number of indexes?
-
A
Manually configure the index in your index config as follows:
Indexes: - kind: Movie Properties: - name: actors name: date_released - kind: Movie Properties: - name: tags name: date_released -
B
Manually configure the index in your index config as follows:
Indexes: - kind: Movie Properties: - name: actors - name: tags - name: date_published - C Set the following in your entity options: exclude_from_indexes = 'actors, tags'
- D Set the following in your entity options: exclude_from_indexes = 'date_published'
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi xoay quanh việc thiết kế hệ thống lưu trữ dữ liệu cho ứng dụng di động dịch vụ streaming media sử dụng Google Cloud Datastore (nay là một phần của Firestore in Datastore mode). Các entity kiểu 'Movie' có các thuộc tính (properties) đa giá trị (multi-valued) như 'actors' (danh sách diễn viên) và 'tags' (danh sách thẻ tag), trong khi 'date_released' là thuộc tính đơn giá trị (single-valued).
Các truy vấn điển hình bao gồm:
- Lấy tất cả phim có diễn viên cụ thể (
actor=<actorname>), sắp xếp theodate_released. - Lấy tất cả phim có tag cụ thể (
tag=Comedy), sắp xếp theodate_released.
Vấn đề chính là tránh "combinatorial explosion" (bùng nổ tổ hợp chỉ mục) trong số lượng indexes của Datastore. Datastore tự động tạo composite indexes cho các truy vấn, nhưng với properties đa giá trị (arrays), nếu không cấu hình thủ công, nó sẽ tạo ra số lượng indexes khổng lồ (ví dụ: mỗi tổ hợp actors × tags × sort field), dẫn đến chi phí lưu trữ và hiệu suất kém. Giải pháp cần tập trung vào việc định nghĩa indexes exploding thủ công chỉ cho các cặp truy vấn cần thiết, sử dụng file index.yaml.
📘 Kiến thức cập nhật (đến 2026): Theo tài liệu Google Cloud Datastore/Firestore mới nhất (Firestore in Datastore mode v1.0+), hỗ trợ exploding indexes cho multi-valued properties (arrays) bằng cách đặt property đa giá trị đầu tiên trong index config với sort field đơn sau đó. Điều này tạo ra indexes "nổ" cho từng giá trị riêng lẻ trong array mà không cần tất cả tổ hợp. (Nguồn: Datastore Indexes Documentation, Firestore Exploding Indexes).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Manually configure the index in your index config as follows:
Indexes:
- kind: Movie
Properties:
- name: actors
name: date_released (direction mặc định ASC cho sort)
- kind: Movie
Properties:
- name: tags
name: date_released
Lý do 🛠️:
- Cấu hình này tạo hai indexes riêng biệt (exploding indexes): Một cho
actors(multi-valued) +date_released(sort field), và một chotags(multi-valued) +date_released. - Datastore sẽ "nổ" (explode) từng giá trị riêng lẻ trong
actorshoặctagskết hợp vớidate_released, hỗ trợ chính xác các truy vấn điển hình mà không tạo tổ hợp thừa (như actors × tags). - Tránh bùng nổ chỉ mục vì chỉ index các cặp cần thiết, tối ưu chi phí và hiệu suất. Đây là best practice cho Datastore với array queries + sorting.
❌ Giải thích tất cả các phương án (đúng/sai)
-
✅ Phương án ĐÚNG (như trên):
Hoàn hảo cho exploding indexes trên từng property đa giá trị riêng lẻ với sort field chung, tránh combinatorial explosion bằng cách không kết hợp actors và tags cùng lúc. Hỗ trợ queryWHERE actors = value ORDER BY date_releasedvà tương tự cho tags. -
❌ Phương án SAI 1: Manually configure the index in your index config as follows:
Indexes: - kind: Movie Properties: - name: actors - name: tags - name: date_publishedLý do sai ❌: Tạo một composite index ba fields (actors × tags × date_published), dẫn đến bùng nổ tổ hợp cực lớn (mỗi diễn viên × mỗi tag × date), chính là vấn đề cần tránh. Ngoài ra, sai tên field (
date_publishedthay vìdate_released), làm index không khớp query. -
❌ Phương án SAI 2: Set the following in your entity options: exclude_from_indexes = 'actors, tags'
Lý do sai ❌: Loại trừactorsvàtagskhỏi tất cả indexes, khiến không thể query trên chúng (ví dụ:WHERE actors = valuesẽ fail). Query điển hình yêu cầu index trên các field này, nên phương án này phá hủy chức năng cốt lõi. -
❌ Phương án SAI 3: Set the following in your entity options: exclude_from_indexes = 'date_published'
Lý do sai ❌: Loại trừdate_published(sai tên field, nên làdate_released) khỏi indexes, khiến không thể sort ORDER BY date_released (Datastore yêu cầu index cho sort fields). Query sẽ fail với lỗi "no matching index". Hơn nữa, không giải quyết vấn đề multi-valued explosion.
💡 Lời khuyên thực tế: Sau khi deploy index.yaml qua gcloud datastore indexes deploy, kiểm tra indexes trong Google Cloud Console để xác nhận. Nếu dùng Firestore native mode (từ 2023+), syntax tương tự nhưng quản lý tự động hơn với single-field indexes. (Nguồn tham khảo: Google Cloud Datastore Best Practices, cập nhật 2025).
Dataflow job to process that log file. You need to make sure the log file in processed once per day as inexpensively as possible. What should you do?
- A Change the processing job to use Google Cloud Dataproc instead.
- B Manually start the Cloud Dataflow job each morning when you get into the office.
- C Create a cron job with Google App Engine Cron Service to run the Cloud Dataflow job.
- D Configure the Cloud Dataflow job as a streaming job so that it processes the log data immediately.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả tình huống bạn làm việc tại một nhà máy sản xuất, nơi các file log ứng dụng được batch (gom nhóm) thành một file log duy nhất mỗi ngày một lần vào lúc 2:00 AM. Bạn đã viết một Google Cloud Dataflow job để xử lý file log đó. Yêu cầu chính là đảm bảo file log được xử lý đúng một lần mỗi ngày và với chi phí thấp nhất có thể (inexpensively as possible).
🛠️ Vấn đề cốt lõi: Dataflow job hiện tại là batch job (xử lý theo lô), phù hợp cho dữ liệu hàng ngày. Cần một cơ chế tự động kích hoạt (trigger) job đúng giờ (2:00 AM hàng ngày), không chạy liên tục để tiết kiệm chi phí, và đáng tin cậy (không thủ công). Đây là bài toán về scheduling (lập lịch) cho Dataflow trên Google Cloud Platform (GCP), không liên quan đến AWS như mô tả ban đầu (có thể là nhầm lẫn).
📘 Kiến thức cập nhật (đến 2026): Theo tài liệu GCP mới nhất (Dataflow 2.x, App Engine flexible/standard environment), Dataflow hỗ trợ batch jobs hiệu quả cho dữ liệu lớn. Scheduling qua Cloud Scheduler (thay thế dần App Engine Cron từ 2023+) hoặc App Engine Cron vẫn hợp lệ, nhưng App Engine Cron được đề cập trực tiếp trong câu hỏi. Nguồn: Google Cloud Dataflow Docs và App Engine Cron Service.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create a cron job with Google App Engine Cron Service to run the Cloud Dataflow job.
Lý do (🧩 Phân tích sâu):
- Tự động hóa hoàn hảo: App Engine Cron Service cho phép lập lịch cron job (ví dụ:
0 2 * * *cho 2:00 AM hàng ngày) để gọi API Dataflow (gcloud dataflow jobs runhoặc REST API), kích hoạt job chính xác một lần/ngày mà không cần can thiệp thủ công. - Tiết kiệm chi phí nhất (💰): App Engine Cron miễn phí cho các job cơ bản (dưới 9 jobs/ngày), chỉ tính phí khi chạy (rẻ hơn streaming). Dataflow batch chỉ chạy khi trigger, không lãng phí tài nguyên idle.
- Đáng tin cậy: Tích hợp GCP, retry tự động nếu fail, phù hợp quy mô production. Không thay đổi logic job hiện tại.
- So với các lựa chọn khác, đây là giải pháp native GCP rẻ nhất cho batch daily (không cần Dataproc đắt đỏ hơn hoặc streaming tốn kém).
❌ Giải thích tất cả các phương án (đúng/sai)
-
[SAI] Change the processing job to use Google Cloud Dataproc instead.
❌ Sai vì: Dataproc là managed Hadoop/Spark cluster, đắt đỏ hơn Dataflow cho batch logs (phải quản lý cluster, chi phí idle cao). Dataflow (Apache Beam) tối ưu hơn cho batch/streaming logs, tự scale/auto-optimize. Không giải quyết scheduling, chỉ thay công cụ không cần thiết. (Nguồn: Dataproc vs Dataflow Comparison). -
[SAI] Manually start the Cloud Dataflow job each morning when you get into the office.
❌ Sai vì: Không tự động, phụ thuộc con người (quên, muộn giờ, nghỉ phép → miss batch 2AM). Không "inexpensively" vì tốn thời gian lao động, không scalable cho production. Vi phạm yêu cầu "processed once per day" reliably. -
[ĐÚNG] Create a cron job with Google App Engine Cron Service to run the Cloud Dataflow job.
✅ Đúng như phân tích trên: Scheduling tự động, rẻ, chính xác. Hoàn hảo cho batch daily! -
[SAI] Configure the Cloud Dataflow job as a streaming job so that it processes the log data immediately.
❌ Sai vì: Streaming mode chạy liên tục 24/7 (watermark, unbounded data), tốn kém gấp nhiều lần (chi phí worker luôn on, dù log chỉ có 1 file/ngày). Không khớp "once per day" và batch nature. Batch mode rẻ hơn cho dữ liệu định kỳ. (Nguồn: Dataflow Batch vs Streaming).
🛡️ Lời khuyên thực tế: Trong thực tế 2026, ưu tiên Cloud Scheduler (thay thế App Engine Cron, hỗ trợ Pub/Sub trigger Dataflow trực tiếp, miễn phí tier lớn hơn). Test job với gcloud scheduler jobs create http cho production!
What should you do?
- A Load the data every 30 minutes into a new partitioned table in BigQuery.
- B Store and update the data in a regional Google Cloud Storage bucket and create a federated data source in BigQuery
- C Store the data in Google Cloud Datastore. Use Google Cloud Dataflow to query BigQuery and combine the data programmatically with the data stored in Cloud Datastore
- D Store the data in a file in a regional Google Cloud Storage bucket. Use Cloud Dataflow to query BigQuery and combine the data programmatically with the data stored in Google Cloud Storage.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả tình huống bạn làm việc tại một công ty tư vấn kinh tế, sử dụng Google BigQuery để phân tích và tương quan (correlate) dữ liệu khách hàng với giá trung bình của 100 mặt hàng phổ biến (như bánh mì, xăng dầu, sữa, v.v.). Dữ liệu giá này được cập nhật mỗi 30 phút, và bạn cần đảm bảo dữ liệu luôn up-to-date để kết hợp (combine) với các dữ liệu khác trong BigQuery một cách rẻ nhất có thể.
🛠️ Yêu cầu chính:
- Giữ dữ liệu giá luôn mới mẻ mà không tốn kém chi phí lưu trữ hoặc xử lý trong BigQuery.
- Tránh load dữ liệu liên tục vào BigQuery vì sẽ phát sinh chi phí cao (query và storage fees).
- Giải pháp phải hỗ trợ federated queries hoặc cách kết nối trực tiếp để query mà không cần copy dữ liệu.
📘 Kiến thức cập nhật đến 2026: Theo tài liệu Google Cloud mới nhất (BigQuery v2.0+), BigQuery hỗ trợ federated queries với Cloud Storage (CSV, JSON, Parquet, ORC, Avro) qua external tables, cho phép query dữ liệu từ GCS mà không load vào BQ, tiết kiệm >90% chi phí so với load đầy đủ. Nguồn: BigQuery Federated Queries Docs & External Tables.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Store and update the data in a regional Google Cloud Storage bucket and create a federated data source in BigQuery.
🧩 Lý do chi tiết:
- Lưu dữ liệu giá vào regional GCS bucket (rẻ, scalable, cập nhật dễ dàng mỗi 30 phút bằng overwrite file).
- Tạo federated data source (external table) trong BigQuery để query trực tiếp từ GCS mà không load dữ liệu vào BQ → Tiết kiệm chi phí storage/query (chỉ scan dữ liệu cần thiết, không duplicate).
- Hỗ trợ update nhanh (overwrite file GCS), phù hợp tần suất 30 phút, và integrate mượt mà với BigQuery qua SQL chuẩn.
- Regional bucket đảm bảo low-latency, high-availability, phù hợp data trends thời gian thực.
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] Load the data every 30 minutes into a new partitioned table in BigQuery.
Phương án này tốn kém vì load liên tục (mỗi 30 phút = ~17k lần/năm) tạo table partitioned mới → phát sinh chi phí storage cao (duplicate data) và query fees khi rewrite partitions. Không hiệu quả cho data nhỏ, update thường xuyên; vi phạm yêu cầu "rẻ nhất". -
✅ [ĐÚNG] Store and update the data in a regional Google Cloud Storage bucket and create a federated data source in BigQuery
Như giải thích ở trên: Rẻ nhất, query federated không load data, update GCS dễ dàng. Hoàn hảo cho use case! -
❌ [SAI] Store the data in Google Cloud Datastore. Use Google Cloud Dataflow to query BigQuery and combine the data programmatically with the data stored in Cloud Datastore
Datastore (Firestore successor) không tối ưu cho analytical queries (thiếu SQL federation với BQ), phải dùng Dataflow (Apache Beam) để ETL phức tạp → chi phí cao (Dataflow compute + BQ query). Không rẻ, không native integrate như GCS. -
❌ [SAI] Store the data in a file in a regional Google Cloud Storage bucket. Use Cloud Dataflow to query BigQuery and combine the data programmatically with the data stored in Google Cloud Storage.
Dùng Dataflow để query BQ và combine với GCS → phức tạp và đắt (Dataflow pipeline chạy định kỳ, compute fees cao). Không tận dụng federated queries native của BQ, vi phạm "combine as cheaply as possible".
🛠️ Kết luận khuyến nghị: Chọn federated GCS để tối ưu chi phí và performance. Test bằng lệnh: CREATE EXTERNAL TABLE ... WITH CONNECTION ... trong BQ Console! 🚀
✑ The user profile: What the user likes and doesn't like to eat
✑ The user account information: Name, address, preferred meal times
✑ The order information: When orders are made, from where, to whom
The database will be used to store all the transactional data of the product. You want to optimize the data schema. Which Google Cloud Platform product should you use?
- A BigQuery
- B Cloud SQL
- C Cloud Bigtable
- D Cloud Datastore
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi yêu cầu thiết kế schema cơ sở dữ liệu cho một dịch vụ đặt món ăn dựa trên machine learning (ML), nhằm dự đoán món ăn mà người dùng muốn. Dữ liệu cần lưu trữ bao gồm:
- Hồ sơ người dùng: Những món ăn người dùng thích hoặc không thích (dữ liệu phân tích hành vi).
- Thông tin tài khoản: Tên, địa chỉ, thời gian ăn uống ưa thích (dữ liệu cá nhân hóa).
- Thông tin đơn hàng: Thời gian đặt hàng, nguồn gốc, người nhận (dữ liệu giao dịch - transactional data).
Mục tiêu là tối ưu hóa schema dữ liệu cho toàn bộ dữ liệu giao dịch của sản phẩm, sử dụng sản phẩm Google Cloud Platform (GCP). Đây là tình huống điển hình cho analytical workload với ML, cần xử lý dữ liệu lớn, query phức tạp để train model dự đoán, thay vì chỉ OLTP (Online Transaction Processing). Kiến thức cập nhật đến 2026: BigQuery hỗ trợ columnar storage, partitioning/clustering để optimize schema, tích hợp BigQuery ML cho inference trực tiếp (phiên bản mới nhất nhấn mạnh vector search và Gemini integration cho ML food recommendation).
📘 Tài liệu tham khảo:
- BigQuery Documentation (cập nhật 2025: Enhanced ML capabilities).
- GCP Data Engineer Best Practices (cho ML-based apps).
✅ Đáp án đúng: BigQuery
Lý do lựa chọn:
BigQuery là serverless data warehouse lý tưởng cho dữ liệu analytical lớn, transactional data với schema tối ưu qua columnar storage, partitioning, clustering (giảm chi phí scan, tăng tốc query ML). Phù hợp lưu user profile (cho recommendation models), account info, order data để train/predict "what users want to eat" via BigQuery ML (hỗ trợ SQL-based ML như logistic regression hoặc Vertex AI integration). Không cần quản lý schema rigid như RDBMS, scale petabyte mà không downtime. Trong 2026, BigQuery hỗ trợ multi-region tables và vector embeddings cho food similarity search – hoàn hảo cho optimize schema ML food ordering.
🛠️ Ví dụ: Sử dụng CREATE TABLE với clustering trên user_id và order_time để query nhanh "predict next order".
📋 Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Phân tích dựa trên use case analytical/ML/transactional data optimization:
-
BigQuery ✅ Đúng:
Như giải thích trên, BigQuery tối ưu schema cho analytical queries và ML trên dữ liệu lớn (user prefs + orders). Hỗ trợ schema auto-detection từ JSON/CSV, partitioning tự động, chi phí pay-per-query. Không phù hợp OLTP thuần túy nhưng lý tưởng cho "predict what users want" với BigQuery ML (cập nhật 2026: Remote functions cho custom ML). -
Cloud SQL ❌ Sai:
Cloud SQL là managed relational DB (MySQL/PostgreSQL), phù hợp OLTP nhỏ-giữa (insert/update nhanh cho orders). Không optimize schema cho analytical/ML lớn (scan toàn bảng chậm với petabyte data), thiếu columnar storage/ML native. Dùng cho app nhỏ, không scale cho "transactional data of the product" toàn diện. -
Cloud Bigtable ❌ Sai:
Cloud Bigtable là NoSQL wide-column store cho high-throughput/low-latency (e.g., time-series như orders real-time). Schema linh hoạt nhưng không phải data warehouse, thiếu SQL analytics/ML tools (query kém cho complex joins user profile + orders). Phù hợp IoT/big data ingestion, không optimize cho ML prediction schema. -
Cloud Datastore ❌ Sai:
Cloud Datastore (nay Firestore in Datastore mode) là NoSQL document DB cho app web/mobile, schema-less documents (user profile tốt). Nhưng không dành analytical/transactional lớn (query giới hạn, không partitioning/clustering như BigQuery). Scale kém cho ML queries trên orders, thiếu columnar optimization (cập nhật 2026: Firestore tốt real-time sync, không thay thế warehouse).
🧩 Kết luận: BigQuery là lựa chọn tối ưu nhất cho workload ML analytical với schema linh hoạt, scale lớn trên GCP! 🚀
- A The CSV data loaded in BigQuery is not flagged as CSV.
- B The CSV data has invalid rows that were skipped on import.
- C The CSV data loaded in BigQuery is not using BigQuery's default encoding.
- D The CSV data has not gone through an ETL phase before loading into BigQuery.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào vấn đề tải dữ liệu từ file CSV vào Google BigQuery. Cụ thể:
- Công ty đang tải file CSV (dùng dấu phẩy làm delimiter) vào BigQuery.
- Quá trình tải thành công hoàn toàn (fully imported successfully).
- Tuy nhiên, dữ liệu đã tải vào BigQuery không khớp byte-to-byte (tức là so sánh nhị phân chính xác từng byte) với file nguồn gốc.
📌 Vấn đề cốt lõi: Dữ liệu trông có vẻ giống nhau về mặt nội dung (có thể đọc được), nhưng khi so sánh ở mức byte (binary level), chúng khác biệt. Điều này thường xảy ra do sự khác biệt về encoding (mã hóa ký tự), vì encoding ảnh hưởng trực tiếp đến cách byte được diễn giải và lưu trữ. BigQuery sử dụng UTF-8 làm encoding mặc định cho CSV (theo tài liệu chính thức cập nhật đến 2024-2026), nên nếu file nguồn dùng encoding khác (như UTF-16 hoặc ISO-8859-1), dữ liệu sẽ bị decode sai, dẫn đến mismatch byte-to-byte dù nội dung logic vẫn đúng.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: The CSV data loaded in BigQuery is not using BigQuery's default encoding.
🛠️ Lý do chi tiết:
- BigQuery mặc định sử dụng UTF-8 khi tải CSV nếu không chỉ định encoding khác (qua tham số
encodingtrong load job). - Nếu file CSV nguồn dùng encoding khác (ví dụ: UTF-16LE, phổ biến ở Windows), BigQuery sẽ cố gắng decode theo UTF-8, dẫn đến dữ liệu byte không khớp chính xác với source (dù có thể hiển thị đúng ký tự).
- Đây là nguyên nhân phổ biến nhất cho mismatch byte-to-byte, vì quá trình import thành công nhưng byte bị thay đổi do reinterpretation.
- Theo docs BigQuery 2026: Hỗ trợ encoding CSV bao gồm UTF-8 (default), ISO-8859-1, UTF-16BE/LE, UTF-32BE/LE. Phải chỉ định đúng để match byte-perfect.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Tôi đánh dấu ✅ cho đúng và ❌ cho sai, kèm lý do bằng tiếng Việt rõ ràng:
-
✅ [ĐÚNG] The CSV data loaded in BigQuery is not using BigQuery's default encoding.
🧩 Như đã giải thích ở trên: Encoding không khớp (không dùng UTF-8 default) là nguyên nhân trực tiếp gây mismatch byte-to-byte. Giải quyết bằng cách chỉ địnhencoding="UTF-16"(hoặc tương ứng) trong load config. Đây là best practice từ BigQuery Load Jobs API. -
❌ [SAI] The CSV data loaded in BigQuery is not flagged as CSV.
📘 Sai vì: BigQuery tự động detect format CSV qua extension (.csv) hoặc source_format=CSV. Nếu không flagged đúng, import sẽ thất bại ngay (không "fully imported successfully"). Vấn đề flagged chỉ ảnh hưởng đến parsing, không phải byte match. -
❌ [SAI] The CSV data has invalid rows that were skipped on import.
🛠️ Sai vì: Nếu có invalid rows (ví dụ: quá nhiều cột), BigQuery skip chúng và báo lỗi (allow_invalid_rows=false default), nhưng dữ liệu valid vẫn import byte-exact nếu encoding đúng. Câu hỏi nhấn "fully imported successfully" và mismatch toàn bộ data, không chỉ partial skip. -
❌ [SAI] The CSV data has not gone through an ETL phase before loading into BigQuery.
🚫 Sai vì: ETL (Extract-Transform-Load) không bắt buộc cho byte-to-byte match. BigQuery hỗ trợ direct load CSV mà không cần ETL (qua gs:// URI hoặc bq load). Thiếu ETL chỉ ảnh hưởng đến data quality/transformation, không thay đổi byte gốc trừ khi ETL alter encoding (nhưng câu hỏi không đề cập).
📚 Tài liệu tham khảo (cập nhật mới nhất 2026)
- BigQuery Documentation - Loading CSV data: Loading CSV data into BigQuery – Chi tiết về encoding và default UTF-8.
- BigQuery Load Job Reference: Load jobs | BigQuery – Encoding options.
- Troubleshooting Data Loading: Troubleshoot data loading errors – Đề cập encoding mismatch.
- Kiến thức từ Google Cloud Professional Data Engineer cert (2024-2026 exam guide): Nhấn mạnh encoding trong BigQuery data ingestion.
Hy vọng phân tích này giúp bạn nắm vững! Nếu cần demo code bq load với encoding, hãy hỏi thêm nhé! 🚀
You are told that due to seasonality, your company expects the number of files to double for the next three months. Which two actions should you take? (Choose two.)
- A Introduce data compression for each file to increase the rate file of file transfer.
- B Contact your internet service provider (ISP) to increase your maximum bandwidth to at least 100 Mbps.
- C Redesign the data ingestion process to use gsutil tool to send the CSV files to a storage bucket in parallel.
- D Assemble 1,000 files into a tape archive (TAR) file. Transmit the TAR files instead, and disassemble the CSV files in the cloud upon receiving them.
- E Create an S3-compatible storage endpoint in your network, and use Google Cloud Storage Transfer Service to transfer on-premises data to the designated storage bucket.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một tình huống thực tế trong Google Cloud Platform (GCP):
Công ty sản xuất 20.000 file CSV (mỗi file <4KB) mỗi giờ, tổng cộng khoảng 480.000 file/ngày. Các file cần được ingest (tiếp nhận) vào GCP trước khi xử lý.
- Vấn đề hiện tại:
- Latency từ site công ty đến GCP: 200ms.
- Bandwidth Internet: 50 Mbps (thấp, nhưng utilization thấp).
- Phương pháp hiện tại: SFTP server trên Compute Engine VM, client SFTP local truyền file từng cái → Chỉ vừa đủ đáp ứng volume hiện tại, nhưng sắp double (40.000 file/giờ) trong 3 tháng tới do seasonality.
- Mục tiêu: Báo cáo dữ liệu ngày trước sẵn sàng trước 10:00 AM hôm sau (khoảng 14 giờ xử lý sau khi ingest xong).
Vấn đề cốt lõi: Hiệu suất truyền file thấp do overhead cao của SFTP (connection overhead, protocol inefficiency với file nhỏ), không tận dụng parallelization hoặc compression hiệu quả. Cần 2 actions để scale up mà không tăng chi phí lớn.
(Kiến thức cập nhật GCP 2026: gsutil hỗ trợ parallel upload mạnh mẽ với multi-threading; Cloud Storage Transfer Service ưu tiên cho large-scale transfer, nhưng không phù hợp on-prem nhỏ lẻ như thế này. Xem docs: gsutil docs, Cloud Storage best practices.)
✅ Đáp án đúng (Chọn 2)
Hai lựa chọn đúng là:
- Redesign the data ingestion process to use gsutil tool to send the CSV files to a storage bucket in parallel.
- Assemble 1,000 files into a tape archive (TAR) file. Transmit the TAR files instead, and disassemble the CSV files in the cloud upon receiving them.
Lý do chọn:
- Cả hai đều giải quyết overhead truyền file nhỏ (small file problem): SFTP kém hiệu quả với hàng nghìn file nhỏ do latency 200ms gây bottleneck (mỗi connection tốn thời gian handshake).
- gsutil parallel: Tận dụng multi-threaded upload (mặc định 8-16 threads, config lên cao hơn), giảm thời gian tổng nhờ parallel I/O, phù hợp bandwidth 50Mbps hiện tại (utilization thấp → room để optimize protocol).
- TAR bundling: Giảm số lượng file truyền từ 20k/giờ xuống 20 file/giờ (1k file/TAR), overhead giảm 50x, dễ disassemble bằng Cloud Functions/Compute Engine/Dataflow sau.
- Kết hợp: Scale double volume dễ dàng, đảm bảo ingest kịp trước 10AM (thời gian truyền giảm từ giờ xuống phút). Không cần tăng bandwidth hay thay ISP.
📋 Giải thích tất cả các phương án
-
❌ Introduce data compression for each file to increase the rate file of file transfer.
Sai vì: Compression (gzip/bzip2) trên file CSV nhỏ (<4KB) tốn CPU local nhiều hơn lợi ích (overhead compress/decompress > tiết kiệm bandwidth). Với 20k file/giờ, latency 200ms vẫn bottleneck chính (không parallel), utilization bandwidth thấp → không giải quyết root cause. Compression phù hợp file lớn hơn (>1MB). -
❌ Contact your internet service provider (ISP) to increase your maximum bandwidth to at least 100 Mbps.
Sai vì: Bandwidth hiện tại 50Mbps chưa saturate (utilization thấp), vấn đề là protocol overhead SFTP + small files + latency, không phải bandwidth thiếu. Tăng bandwidth tốn kém, không scale dài hạn, bỏ qua optimize GCP native tools. -
✅ Redesign the data ingestion process to use gsutil tool to send the CSV files to a storage bucket in parallel.
Đúng vì: gsutil cp/rsync hỗ trợ parallel composite uploads (multi-part, threading tự động), tối ưu cho small files. Ví dụ:gsutil -m cp -r local_dir gs://bucket/→ chia file thành threads, giảm thời gian tổng 10-20x so SFTP. Dễ script tự động, tích hợp Cloud Storage (durable, scalable). (Nguồn: gsutil parallel uploads). -
✅ Assemble 1,000 files into a tape archive (TAR) file. Transmit the TAR files instead, and disassemble the CSV files in the cloud upon receiving them.
Đúng vì: Bundling TAR giảm drastic số transfers (20k → 20 file/giờ), TAR không compress (nhẹ CPU), truyền nhanh qua SFTP hoặc gsutil. Cloud side: Dùng Cloud Storage + Cloud Functions hoặc Dataflow extract (tar -xf), atomic và scalable. Phù hợp volume double. (Nguồn: GCP best practices for small files). -
❌ Create an S3-compatible storage endpoint in your network, and use Google Cloud Storage Transfer Service to transfer on-premises data to the designated storage bucket.
Sai vì: Storage Transfer Service (STS) thiết kế cho large-scale, scheduled transfers từ cloud-to-cloud hoặc S3, không tối ưu on-prem nhỏ lẻ (overhead setup S3-compatible như MinIO tốn thời gian, STS polling-based kém real-time). Với 200ms latency + 50Mbps, STS không parallel tốt như gsutil, phức tạp hóa thay vì simplify. (Nguồn: STS limitations).
🛠️ Khuyến nghị triển khai
- Test ngay: Benchmark gsutil -m vs SFTP với 40k files.
- Scale thêm: Kết hợp Pub/Sub + Cloud Functions cho ingest real-time nếu cần.
- Monitoring: Dùng Cloud Monitoring theo dõi ingestion latency/throughput.
Tổng kết: Tập trung optimize protocol + bundling thay vì hardware/ISP! 🚀
TB per year, and each data entry has about 100 attributes. The data processing pipeline does not require atomicity, consistency, isolation, and durability (ACID).
However, high availability and low latency are required.
You need to analyze the data by querying against individual fields. Which three databases meet your requirements? (Choose three.)
- A Redis
- B HBase
- C MySQL
- D MongoDB
- E Cassandra
- F HDFS with Hive
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi yêu cầu chọn ba cơ sở dữ liệu NoSQL phù hợp để xử lý dữ liệu telemetry từ hàng triệu thiết bị IoT. 📊
- Quy mô dữ liệu: Tăng trưởng 100TB/năm, mỗi bản ghi có khoảng 100 thuộc tính (attributes) → Cần database scale lớn, xử lý dữ liệu phi cấu trúc hoặc semi-structured.
- Yêu cầu kỹ thuật:
- Không cần ACID (Atomicity, Consistency, Isolation, Durability) → Chấp nhận eventual consistency, ưu tiên throughput cao.
- High availability (HA): Khả năng sẵn sàng cao, chịu lỗi tốt.
- Low latency: Truy vấn nhanh, đặc biệt khi query theo individual fields (truy vấn trên từng trường cụ thể).
- Mục đích: Phân tích dữ liệu bằng cách query trực tiếp trên các trường riêng lẻ, phù hợp với NoSQL hỗ trợ indexing và querying linh hoạt.
Câu hỏi thuộc chủ đề AWS Big Data/NoSQL (phiên bản cập nhật 2024-2026: AWS EMR cho HBase/Cassandra, DocumentDB cho MongoDB, Keyspaces cho Cassandra).
Đáp án đúng (chọn 3): ✅ HBase, ✅ MongoDB, ✅ Cassandra.
Lý do chọn: Đây là các NoSQL database phân tán, scale ngang tốt cho petabyte-scale data từ IoT, hỗ trợ query fields với low latency (<10ms), HA qua replication multi-AZ, eventual consistency (không ACID). Phù hợp AWS EMR (HBase/Cassandra) và DocumentDB/Keyspaces (MongoDB/Cassandra). 🛠️
📋 Giải thích tất cả các phương án (đúng/sai)
-
Redis ❌
Sai vì: Redis (Amazon ElastiCache Redis) là in-memory key-value store, ưu tiên tốc độ cực cao nhưng không phù hợp lưu trữ persistent 100TB (dữ liệu chủ yếu RAM, persistence là optional và kém scale). Không hỗ trợ query complex trên 100 attributes hiệu quả, dễ mất dữ liệu nếu crash dù có HA. -
HBase ✅
Đúng vì: HBase (AWS EMR HBase) là wide-column NoSQL trên HDFS, scale hoàn hảo cho 100TB+ IoT data, query individual fields qua row-key/column families với low latency (millions QPS). HA qua HMaster replication, eventual consistency (không ACID). Cập nhật 2026: HBase 2.5+ hỗ trợ vector search cho telemetry. -
MySQL ❌
Sai vì: MySQL (Amazon RDS/Aurora MySQL) là RDBMS SQL, yêu cầu ACID transactions, không scale tốt cho 100TB NoSQL workloads (shard khó, latency cao khi query large tables). Không phù hợp IoT high-throughput mà không ACID. -
MongoDB ✅
Đúng vì: MongoDB (Amazon DocumentDB) là document NoSQL, lưu trữ JSON/BSON với 100 attributes dễ dàng, query fields linh hoạt (aggregation pipeline) low latency. HA multi-AZ replicas, eventual consistency option. Cập nhật 2026: DocumentDB 5.0+ hỗ trợ time-series cho IoT telemetry. -
Cassandra ✅
Đúng vì: Cassandra (AWS Keyspaces hoặc EMR) là wide-column NoSQL, thiết kế cho write-heavy IoT (100TB/năm), query CQL trên fields với low latency (tunable consistency). HA ring topology, không ACID. Cập nhật 2026: Keyspaces serverless scale auto, zero-ETL to S3. -
HDFS with Hive ❌
Sai vì: HDFS (Hadoop Distributed File System trên EMR) + Hive là batch processing system, không phải database real-time (latency cao hàng phút/giờ). Hive query SQL-like nhưng không low latency cho IoT, tập trung analytics lớn chứ không HA/low-latency queries.
📘 Tài liệu tham khảo (AWS cập nhật 2024-2026)
- AWS EMR HBase/Cassandra: docs.aws.amazon.com/emr/latest/ReleaseGuide/emr-hbase.html
- Amazon DocumentDB (MongoDB): docs.aws.amazon.com/documentdb
- Amazon Keyspaces (Cassandra): docs.aws.amazon.com/keyspaces
- IoT Data Storage Best Practices: aws.amazon.com/blogs/iot (Time-series NoSQL recommendations).
🧐 Kết luận: Các lựa chọn đúng tối ưu cho workload IoT AWS-native!