Ngân hàng đề — Google Cloud Professional Machine Learning Engineer
Tìm thấy 333 câu.
- A Use the Cloud Data Loss Prevention (DLP) API to de-identify the PII before performing data exploration and preprocessing.
- B Use customer-managed encryption keys (CMEK) to encrypt the PII data at rest, and decrypt the PII data during data exploration and preprocessing.
- C Use a VM inside a VPC Service Controls security perimeter to perform data exploration and preprocessing.
- D Use Google-managed encryption keys to encrypt the PII data at rest, and decrypt the PII data during data exploration and preprocessing.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc xử lý dữ liệu khách hàng cho một tổ chức y tế, lưu trữ trong Cloud Storage (Google Cloud Storage - GCS). Dữ liệu chứa PII (Personally Identifiable Information) - thông tin cá nhân có thể nhận dạng cá nhân, như tên, địa chỉ, số bảo hiểm y tế, rất nhạy cảm trong lĩnh vực y tế. Nhiệm vụ là thực hiện data exploration (khám phá dữ liệu) và preprocessing (tiền xử lý dữ liệu), đồng thời đảm bảo bảo mật và quyền riêng tư cho các trường dữ liệu nhạy cảm.
🔑 Yêu cầu cốt lõi: Không chỉ mã hóa dữ liệu tại chỗ (at rest), mà cần xử lý an toàn khi dữ liệu được đọc và phân tích (in use), tránh lộ PII trong quá trình exploration/preprocessing. Giải pháp phải tuân thủ các tiêu chuẩn bảo mật cao như HIPAA cho y tế.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use the Cloud Data Loss Prevention (DLP) API to de-identify the PII before performing data exploration and preprocessing.
Lý do:
- Cloud DLP API là công cụ chuyên dụng của Google Cloud để phát hiện và de-identify (ẩn danh hóa) PII tự động (như mask, redact, replace, tokenize).
- Quy trình: Scan dữ liệu trong GCS → De-identify PII trước → Sau đó mới exploration/preprocessing trên dữ liệu đã an toàn.
- Điều này đảm bảo zero-trust cho PII trong quá trình xử lý, phù hợp với quy định y tế (HIPAA, GDPR). Không cần decrypt hay expose PII gốc.
- Hiệu quả cao với ML workflows, tích hợp dễ dàng với Dataflow, BigQuery. 📈 (Cập nhật 2024-2026: DLP hỗ trợ AI-based inspection với Gemini models cho accuracy cao hơn).
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh:
-
✅ Use the Cloud Data Loss Prevention (DLP) API to de-identify the PII before performing data exploration and preprocessing.
Đúng vì: Như đã giải thích ở trên, DLP trực tiếp giải quyết vấn đề de-identify PII trước khi xử lý, tránh rủi ro lộ thông tin trong exploration. 🛡️ Đây là best practice cho sensitive data trong GCS. -
❌ Use customer-managed encryption keys (CMEK) to encrypt the PII data at rest, and decrypt the PII data during data exploration and preprocessing.
Sai vì: CMEK chỉ mã hóa dữ liệu tại chỗ (at rest) trong GCS, nhưng khi exploration/preprocessing, dữ liệu phải decrypt để đọc → PII vẫn bị expose đầy đủ. Không bảo vệ trong quá trình xử lý (in transit/in use). 🔒 Không phù hợp cho privacy PII. -
❌ Use a VM inside a VPC Service Controls security perimeter to perform data exploration and preprocessing.
Sai vì: VPC Service Controls (nay là VPC SC Perimeter) ngăn data exfiltration (rò rỉ dữ liệu ra ngoài perimeter), nhưng không de-identify hay che giấu PII. VM vẫn đọc PII gốc → rủi ro insider threat hoặc log exposure cao. 🛡️ Chỉ bảo vệ perimeter, không xử lý nội dung dữ liệu. -
❌ Use Google-managed encryption keys to encrypt the PII data at rest, and decrypt the PII data during data exploration and preprocessing.
Sai vì: Tương tự CMEK, Google-managed keys chỉ mã hóa at rest mặc định cho GCS. Decrypt khi xử lý → PII vẫn lộ. Bạn mất kiểm soát key so với CMEK, nhưng vấn đề cốt lõi là không de-identify. 🔑 Không giải quyết privacy trong exploration.
📘 Tài liệu tham khảo (cập nhật mới nhất 2024-2026)
- Cloud DLP API: Google Cloud DLP Documentation - Hướng dẫn de-identify PII cho GCS (v1.5+ với Gemini integration).
- GCS Encryption: Encryption at Rest - So sánh CMEK vs Google-managed.
- VPC Service Controls: VPC SC Overview - Focus trên perimeter, không de-identify.
- Best Practices for Healthcare: Google Cloud Healthcare API & Compliance - Khuyến nghị DLP cho PII.
🧰 Lời khuyên: Trong ML pipeline thực tế, kết hợp DLP với Confidential Computing (như Confidential VMs) để bảo vệ end-to-end!
- A Use scikit-learn to build a tree-based model, and use SHAP values to explain the model output.
- B Use scikit-learn to build a tree-based model, and use partial dependence plots (PDP) to explain the model output.
- C Use TensorFlow to create a deep learning-based model, and use Integrated Gradients to explain the model output.
- D Use TensorFlow to create a deep learning-based model, and use the sampled Shapley method to explain the model output.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc xây dựng một mô hình dự đoán bảo trì dự phòng (predictive maintenance) để phát hiện sớm các khuyết điểm trên các bộ phận của cầu (bridges). Đầu vào của mô hình là hình ảnh độ phân giải cao (high definition images), đòi hỏi xử lý dữ liệu hình ảnh phức tạp. Mục tiêu chính là giải thích đầu ra của mô hình (model output) một cách rõ ràng cho các bên liên quan (stakeholders) để họ có thể hành động kịp thời.
🛠️ Thách thức chính:
- Dữ liệu hình ảnh yêu cầu mô hình học sâu (deep learning) như CNN để trích xuất đặc trưng hiệu quả, không phù hợp với mô hình cây (tree-based) truyền thống.
- Cần phương pháp giải thích mô hình (XAI - Explainable AI) đáng tin cậy, đặc biệt trong AWS SageMaker Clarify (công cụ giải thích mô hình của AWS), hỗ trợ tốt cho dữ liệu hình ảnh.
- Kiến thức cập nhật đến 2026: AWS SageMaker Clarify (phiên bản mới nhất hỗ trợ image/text data qua Integrated Gradients, SHAP cho tree-based, PDP cho global explanations).
📘 Tài liệu tham khảo:
- AWS SageMaker Clarify Documentation: Model Explainability (cập nhật 2025).
- AWS re:Invent 2024/2025 sessions về XAI cho vision models.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use TensorFlow to create a deep learning-based model, and use Integrated Gradients to explain the model output.
Lý do:
- TensorFlow lý tưởng cho deep learning trên hình ảnh (CNN/ResNet), xử lý HD images hiệu quả trong SageMaker.
- Integrated Gradients là phương pháp gradient-based chuẩn cho DL vision models trong SageMaker Clarify, tạo heatmap trực quan chỉ ra pixel quan trọng gây ra dự đoán (ví dụ: vị trí khuyết điểm trên cầu). Phương pháp này nhanh, chính xác cao với dữ liệu hình ảnh, và được AWS khuyến nghị cho image classification/object detection từ 2023+.
🧩 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] Use scikit-learn to build a tree-based model, and use SHAP values to explain the model output.
Phương án này sai vì scikit-learn với tree-based (như Random Forest/XGBoost) không phù hợp xử lý hình ảnh HD (tabular data mới hiệu quả). SHAP hoạt động tốt cho tree (exact computation), nhưng mô hình gốc đã kém → không giải thích được chính xác đặc trưng hình ảnh. SageMaker Clarify hỗ trợ SHAP cho tree, nhưng không dùng cho vision. -
❌ [SAI] Use scikit-learn to build a tree-based model, and use partial dependence plots (PDP) to explain the model output.
Sai tương tự: Tree-based scikit-learn không xử lý images (cần flatten → mất thông tin không gian). PDP chỉ hiển thị tương tác global features (tốt cho tabular), không trực quan cho hình ảnh. SageMaker Clarify hỗ trợ PDP cho tree-based, nhưng không giải quyết vấn đề input data. -
✅ [ĐÚNG] Use TensorFlow to create a deep learning-based model, and use Integrated Gradients to explain the model output.
Đúng hoàn toàn: TensorFlow DL (Keras/TF2.x) tối ưu cho images trong SageMaker Training. Integrated Gradients tạo attribution maps chi tiết (superpixels/pixels), dễ giải thích cho stakeholders (ví dụ: "Mô hình phát hiện vết nứt tại vị trí X"). AWS Clarify tích hợp native, scalable đến 2026. -
❌ [SAI] Use TensorFlow to create a deep learning-based model, and use the sampled Shapley method to explain the model output.
Sai vì dù TensorFlow DL đúng, sampled Shapley (approximation của SHAP) chậm, noisy với high-dim images (KernelSHAP/DeepLIFT kém hiệu quả). SageMaker Clarify ưu tiên Integrated Gradients cho images thay vì SHAP approximations (chỉ dùng cho tabular/non-image DL). Không trực quan bằng heatmap.
The data includes the following variables for each day:
•Number of scheduled surgeries
•Number of beds occupied
•Date
You want to maximize the speed of model development and testing. What should you do?
- A Create a BigQuery table. Use BigQuery ML to build a regression model, with number of beds as the target variable, and number of scheduled surgeries and date features (such as day of week) as the predictors.
- B Create a BigQuery table. Use BigQuery ML to build an ARIMA model, with number of beds as the target variable, and date as the time variable.
- C Create a Vertex AI tabular dataset. Train an AutoML regression model, with number of beds as the target variable, and number of scheduled minor surgeries and date features (such as day of the week) as the predictors.
- D Create a Vertex AI tabular dataset. Train a Vertex AI AutoML Forecasting model, with number of beds as the target variable, number of scheduled surgeries as a covariate and date as the time variable.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc xây dựng mô hình dự đoán số giường bệnh cần thiết hàng ngày cho bệnh viện dựa trên số lượng phẫu thuật đã lên lịch. Dữ liệu có 365 dòng (một năm) với các biến:
- Number of scheduled surgeries (số phẫu thuật lên lịch - đây là thông tin biết trước cho tương lai).
- Number of beds occupied (số giường đang chiếm dụng - biến mục tiêu cần dự đoán).
- Date (ngày - yếu tố thời gian quan trọng).
Mục tiêu: Dự đoán trước số giường cần dựa trên lịch phẫu thuật, và tối ưu hóa tốc độ phát triển + kiểm tra mô hình (maximize speed of model development and testing). Đây là bài toán time series forecasting (dự báo chuỗi thời gian) với covariate đã biết trước (scheduled surgeries), phù hợp với các công cụ tự động hóa nhanh trên Google Cloud như Vertex AI hoặc BigQuery ML.
📘 Dẫn nguồn:
- Vertex AI Forecasting documentation (cập nhật 2024-2026) - Hỗ trợ time series với known future covariates.
- BigQuery ML ARIMA_PLUS (cập nhật 2025) - Univariate time series.
✅ Đáp án đúng và lý do lựa chọn
Create a Vertex AI tabular dataset. Train a Vertex AI AutoML Forecasting model, with number of beds as the target variable, number of scheduled surgeries as a covariate and date as the time variable.
Lý do:
- Đây là lựa chọn nhanh nhất cho phát triển và kiểm tra vì Vertex AI AutoML Forecasting được thiết kế chuyên biệt cho time series forecasting với known future covariates (scheduled surgeries biết trước).
- Target: Number of beds (dự đoán).
- Covariate: Scheduled surgeries (tăng/giảm giường theo lịch).
- Time variable: Date (xử lý seasonality, trend tự động).
- Tốc độ cao: Upload dataset → AutoML train tự động → Test nhanh, không cần code phức tạp. Phù hợp dữ liệu nhỏ (365 rows).
🛠️ Ưu điểm nổi bật: Hỗ trợ cross-validation thời gian, forecast horizon linh hoạt (cập nhật Vertex AI 2025).
📋 Giải thích tất cả các phương án (đúng/sai)
-
❌ [SAI] Create a BigQuery table. Use BigQuery ML to build a regression model, with number of beds as the target variable, and number of scheduled surgeries and date features (such as day of week) as the predictors.
Lý do sai: BigQuery ML regression là mô hình tabular regression thông thường, không xử lý time series (bỏ qua thứ tự thời gian, autocorrelation). Việc engineer features thủ công từ date (day of week) mất thời gian, không tối ưu tốc độ. Không dự báo "ahead" tốt cho future dates mà không có dữ liệu tương lai đầy đủ. -
❌ [SAI] Create a BigQuery ML to build an ARIMA model, with number of beds as the target variable, and date as the time variable.
Lý do sai: BigQuery ML ARIMA (ARIMA_PLUS) chỉ hỗ trợ univariate time series (không dùng covariates như scheduled surgeries). Bỏ qua thông tin lịch phẫu thuật quan trọng, dẫn đến dự đoán kém chính xác. Tuy nhanh nhưng không fit bài toán có covariate known future (cập nhật BQML 2025 vẫn hạn chế multivariate). -
❌ [SAI] Create a Vertex AI tabular dataset. Train an AutoML regression model, with number of beds as the target variable, and number of scheduled minor surgeries and date features (such as day of the week) as the predictors.
Lý do sai:- AutoML Tabular Regression không dành cho time series (xử lý như dữ liệu tĩnh, ignore temporal dependencies).
- Sai dữ liệu: "Scheduled minor surgeries" không tồn tại trong dataset (chỉ có "scheduled surgeries").
- Phải engineer date features thủ công → Chậm hơn AutoML Forecasting. Không dự báo future tốt.
-
✅ [ĐÚNG] Create a Vertex AI tabular dataset. Train a Vertex AI AutoML Forecasting model, with number of beds as the target variable, number of scheduled surgeries as a covariate and date as the time variable.
Lý do đúng: Như phần trên - Tối ưu tốc độ (no-code, auto-handle time series + covariates), chính xác cao cho bài toán dự đoán giường dựa lịch phẫu thuật. Vertex AI Forecasting (2024-2026) hỗ trợ context windows, exogenous variables hoàn hảo.
🧠 Kết luận: Chọn Vertex AI AutoML Forecasting để nhanh - chính xác - dễ scale! Nếu cần custom, có thể dùng Vertex AI Pipelines sau.
- A Use the Kubeflow Pipelines SDK to implement the pipeline. Use the BigQueryJobOp component to run the preprocessing script and the CustomTrainingJobOp component to launch a Vertex AI training job.
- B Use the Kubeflow Pipelines SDK to implement the pipeline. Use the DataflowPythonJobOp component to preprocess the data and the CustomTrainingJobOp component to launch a Vertex AI training job.
- C Use the TensorFlow Extended SDK to implement the pipeline Use the ExampleGen component with the BigQuery executor to ingest the data the Transform component to preprocess the data, and the Trainer component to launch a Vertex AI training job.
- D Use the TensorFlow Extended SDK to implement the pipeline Implement the preprocessing steps as part of the input_fn of the model. Use the ExampleGen component with the BigQuery executor to ingest the data and the Trainer component to launch a Vertex AI training job.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi này tập trung vào việc xây dựng một training pipeline tự động trên Google Cloud Platform (GCP), cụ thể là Vertex AI, để retrain một mô hình wide and deep được phát triển bằng TensorFlow.
-
Bối cảnh chính:
- Dữ liệu thô được preprocess bằng SQL script trên BigQuery, thực hiện các instance-level transformations (chuyển đổi ở mức từng bản ghi dữ liệu cá nhân, không phải batch-level như feature statistics).
- Pipeline cần retrain mô hình hàng tuần, nhưng mô hình được sử dụng để tạo recommendations hàng ngày.
- Mục tiêu cốt lõi: Giảm thiểu thời gian phát triển mô hình và training (minimize model development and training time), nghĩa là ưu tiên tái sử dụng script SQL hiện có, tránh viết lại logic preprocess phức tạp.
-
Yêu cầu kỹ thuật:
- Pipeline phải tích hợp preprocessing từ BigQuery và training trên Vertex AI.
- Phù hợp với quy trình production ML trên GCP, tận dụng các component sẵn có để nhanh chóng implement mà không cần custom code nhiều.
Câu hỏi kiểm tra kiến thức về Kubeflow Pipelines (KFP) và TensorFlow Extended (TFX) trên Vertex AI (cập nhật đến 2026: Vertex AI Pipelines vẫn dựa trên KFP SDK v2+, hỗ trợ các Op như BigQueryJobOp và CustomTrainingJobOp cho custom TF jobs).
📘 Tài liệu tham khảo:
- Vertex AI Pipelines Documentation (Google Cloud, cập nhật 2025).
- Kubeflow Pipelines Components: BigQueryJobOp & CustomTrainingJobOp.
- BigQuery ML Integration.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use the Kubeflow Pipelines SDK to implement the pipeline. Use the BigQueryJobOp component to run the preprocessing script and the CustomTrainingJobOp component to launch a Vertex AI training job.
Lý do 🛠️:
- BigQueryJobOp cho phép chạy SQL script trực tiếp trên BigQuery mà không cần viết thêm code Python hay Dataflow job, tái sử dụng hoàn toàn script preprocess instance-level hiện có → giảm tối đa thời gian phát triển.
- CustomTrainingJobOp tích hợp mượt mà với Vertex AI để launch custom TensorFlow training job (phù hợp wide & deep model).
- Pipeline chạy hàng tuần qua Vertex AI Pipelines (dựa KFP SDK), tự động hóa end-to-end, hỗ trợ scheduling cron job → minimize training time nhờ scale tự động trên GCP infra.
- Đây là cách best practice trên Vertex AI (2026), tránh overhead của TFX cho trường hợp preprocess đơn giản bằng SQL.
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn (giữ nguyên văn bản gốc bằng tiếng Anh), đánh dấu ✅ đúng hoặc ❌ sai, kèm lý do cụ thể bằng tiếng Việt:
-
✅ Use the Kubeflow Pipelines SDK to implement the pipeline. Use the BigQueryJobOp component to run the preprocessing script and the CustomTrainingJobOp component to launch a Vertex AI training job.
🛠️ Đúng vì: Như giải thích trên, BigQueryJobOp chạy SQL script gốc trực tiếp (output table cho training data), kết hợp CustomTrainingJobOp train TF model trên Vertex AI. Tối ưu dev time (không rewrite preprocess), hỗ trợ weekly retrain qua KFP scheduler. Phù hợp instance-level SQL transforms. -
❌ Use the Kubeflow Pipelines SDK to implement the pipeline. Use the DataflowPythonJobOp component to preprocess the data and the CustomTrainingJobOp component to launch a Vertex AI training job.
🧩 Sai vì: DataflowPythonJobOp yêu cầu viết Python job để đọc BQ → transform → write output, không tái sử dụng SQL script gốc → tăng dev time (vi phạm minimize requirement). Dataflow phù hợp batch/stream processing phức tạp, thừa cho SQL đơn giản trên BQ. -
❌ Use the TensorFlow Extended SDK to implement the pipeline Use the ExampleGen component with the BigQuery executor to ingest the data the Transform component to preprocess the data, and the Trainer component to launch a Vertex AI training job.
🚫 Sai vì: TFX (TF Extended) dùng ExampleGen (BigQuery executor) chỉ ingest raw data từ BQ thành TF Examples, Transform component yêu cầu tf.transform (Apache Beam) cho statistical transformations (như compute schema, vocabulary) → không phù hợp instance-level SQL (cần custom SQL logic riêng). Overhead cao, dev time dài hơn KFP đơn giản. -
❌ Use the TensorFlow Extended SDK to implement the pipeline Implement the preprocessing steps as part of the input_fn of the model. Use the ExampleGen component with the BigQuery executor to ingest the data and the Trainer component to launch a Vertex AI training job.
🚫 Sai vì: Đẩy preprocessing vào input_fn của model (eager/on-the-fly) làm training chậm (preprocess mỗi epoch), không tạo dataset tĩnh từ SQL → vi phạm minimize training time. ExampleGen chỉ ingest raw, không chạy SQL script → mất tái sử dụng, không scalable cho weekly retrain lớn.
Kết luận 🎯: Lựa chọn đầu tiên là optimal cho GCP ML Engineer, tận dụng native components Vertex AI để nhanh chóng productionize! Nếu cần code sample, tham khảo Vertex AI SDK samples trên GitHub.
- A Configure the machines of the first two worker pools to have GPUs, and to use a container image where your training code runs. Configure the third worker pool to have GPUs, and use the reductionserver container image.
- B Configure the machines of the first two worker pools to have GPUs and to use a container image where your training code runs. Configure the third worker pool to use the reductionserver container image without accelerators, and choose a machine type that prioritizes bandwidth.
- C Configure the machines of the first two worker pools to have TPUs and to use a container image where your training code runs. Configure the third worker pool without accelerators, and use the reductionserver container image without accelerators, and choose a machine type that prioritizes bandwidth.
- D Configure the machines of the first two pools to have TPUs, and to use a container image where your training code runs. Configure the third pool to have TPUs, and use the reductionserver container image.
Xem giải thích
🧩 Giải thích nội dung câu hỏi
Câu hỏi tập trung vào việc cấu hình worker pools cho công việc huấn luyện phân tán (distributed training job) trên Vertex AI của Google Cloud, sử dụng chiến lược Reduction Server.
- Bối cảnh: Bạn đang huấn luyện một mô hình ngôn ngữ tùy chỉnh (custom language model) với bộ dữ liệu lớn. Reduction Server là kỹ thuật tối ưu hóa giao tiếp all-reduce giữa các worker bằng cách sử dụng các server chuyên biệt để giảm độ trễ và tăng băng thông, đặc biệt hữu ích cho huấn luyện quy mô lớn trên GPU.
- Yêu cầu chính: Cấu hình ba worker pools:
- Hai pool đầu: Chạy mã huấn luyện thực tế (training code).
- Pool thứ ba: Chạy Reduction Server để xử lý giao tiếp.
- Mục tiêu: Đảm bảo hiệu suất cao, tránh lãng phí tài nguyên (như không dùng accelerator không cần thiết cho Reduction Server), và ưu tiên băng thông mạng cao cho pool thứ ba.
(Kiến thức dựa trên tài liệu Vertex AI mới nhất 2025-2026: Reduction Server yêu cầu pool chính có accelerator (GPU/TPU), pool reduction không cần accelerator mà ưu tiên machine type high-bandwidth như n1-highcpu hoặc c3-highcpu với InfiniBand/10Gbps+ network. Xem docs: Vertex AI Distributed Training). 📘
✅ Đáp án đúng
Configure the machines of the first two worker pools to have GPUs and to use a container image where your training code runs. Configure the third worker pool to use the reductionserver container image without accelerators, and choose a machine type that prioritizes bandwidth.
Lý do chọn đáp án này 🏆:
- Hai pool đầu có GPUs để chạy training code (container image tùy chỉnh chứa mã huấn luyện), phù hợp cho mô hình ngôn ngữ lớn cần compute cao.
- Pool thứ ba dùng reductionserver container image (image chính thức từ Google), không cần accelerator (tiết kiệm chi phí vì chỉ xử lý giao tiếp), và machine type ưu tiên bandwidth (như
n1-highcpu-16hoặcc3series với high network throughput) để tối ưu all-reduce operations. - Đây là cấu hình chuẩn theo best practices Vertex AI cho Reduction Server, giảm bottleneck giao tiếp lên đến 50-70% so với NCCL thuần. ✅
❌ Giải thích tất cả các phương án
-
[SAI] Configure the machines of the first two worker pools to have GPUs, and to use a container image where your training code runs. Configure the third worker pool to have GPUs, and use the reductionserver container image.
Phân tích sai: Pool thứ ba không cần GPUs vì Reduction Server chỉ xử lý giao tiếp dữ liệu (gradient aggregation), không cần compute accelerator. Việc thêm GPU gây lãng phí tài nguyên và chi phí cao, vi phạm nguyên tắc tối ưu của Vertex AI. ❌ -
[ĐÚNG] Configure the machines of the first two worker pools to have GPUs and to use a container image where your training code runs. Configure the third worker pool to use the reductionserver container image without accelerators, and choose a machine type that prioritizes bandwidth.
(Đã giải thích ở phần đáp án đúng) ✅ -
[SAI] Configure the machines of the first two worker pools to have TPUs and to use a container image where your training code runs. Configure the third worker pool without accelerators, and use the reductionserver container image without accelerators, and choose a machine type that prioritizes bandwidth.
Phân tích sai: Hai pool đầu dùng TPUs thay vì GPUs là không phù hợp. Reduction Server trên Vertex AI hỗ trợ tốt nhất với GPUs cho custom language model (như dựa trên Transformer), vì TPU yêu cầu code tối ưu XLA-specific và ít linh hoạt hơn cho mô hình tùy chỉnh. Lặp "without accelerators" ở pool thứ ba là thừa, nhưng vấn đề chính là TPU ở pool chính. ❌ (Docs: TPU chủ yếu cho TensorFlow/JAX pre-built, GPU linh hoạt hơn - Vertex AI Accelerators). -
[SAI] Configure the machines of the first two pools to have TPUs, and to use a container image where your training code runs. Configure the third pool to have TPUs, and use the reductionserver container image.
Phân tích sai: Tương tự option trước, TPUs không phải lựa chọn tối ưu cho pool chính với custom LLM (GPU phổ biến hơn). Pool thứ ba thêm TPUs là sai hoàn toàn vì Reduction Server không cần accelerator, chỉ cần CPU + high-bandwidth network. Không chỉ định bandwidth là thiếu sót lớn, dẫn đến hiệu suất kém. ❌
Tóm tắt best practices 🛠️: Luôn dùng GPU cho worker chính, CPU high-bandwidth cho reduction, và test với gcloud ai custom-jobs create để verify. Nếu scale lớn, kết hợp với Multi-node Pod hoặc Pathways! 🚀
- A Perform data validation to ensure that the input data to the pipeline is the same format as the input data to the endpoint.
- B Refactor the transformation code in the batch data pipeline so that it can be used outside of the pipeline. Use the same code in the endpoint.
- C Refactor the transformation code in the batch data pipeline so that it can be used outside of the pipeline. Share this code with the end users of the endpoint.
- D Batch the real-time requests by using a time window and then use the Dataflow pipeline to preprocess the batched requests. Send the preprocessed requests to the endpoint.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi này tập trung vào một tình huống thực tế trong quy trình Machine Learning trên Google Cloud Platform (GCP) (không phải AWS như đề cập, vì Dataflow là dịch vụ Apache Beam trên GCP dùng cho batch/streaming data processing).
- Bối cảnh: Bạn đã huấn luyện (training) một mô hình ML sử dụng dữ liệu được tiền xử lý (preprocessed) qua một pipeline batch Dataflow.
- Yêu cầu: Triển khai real-time inference (suy luận thời gian thực) tại endpoint serving.
- Mục tiêu chính: Đảm bảo logic tiền xử lý dữ liệu (preprocessing logic) được áp dụng nhất quán (consistent) giữa giai đoạn training và serving. Điều này rất quan trọng để tránh data drift hoặc mismatch dẫn đến hiệu suất mô hình kém (ví dụ: mô hình train trên dữ liệu normalized nhưng serving nhận dữ liệu raw).
- Thách thức: Dataflow batch phù hợp cho training lớn, nhưng real-time inference cần tốc độ cao, không thể dùng batch pipeline trực tiếp.
Mục đích câu hỏi: Kiểm tra kiến thức best practice về MLOps trên GCP, đặc biệt là reuse code để đảm bảo consistency giữa training và serving (theo nguyên tắc "same code, same logic").
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng:
Refactor the transformation code in the batch data pipeline so that it can be used outside of the pipeline. Use the same code in the endpoint.
Lý do chi tiết:
- 🛠️ Refactor code (tái cấu trúc mã biến đổi dữ liệu từ pipeline batch Dataflow) để nó có thể chạy ngoài pipeline (ví dụ: thành hàm Python độc lập hoặc TensorFlow Transform - TFT).
- ✅ Sử dụng cùng code đó trực tiếp trong endpoint serving (như Vertex AI Prediction, Cloud Run, hoặc custom container). Điều này đảm bảo 100% consistency về logic preprocessing (không có sai lệch do implement lại thủ công).
- 📈 Đây là best practice MLOps trên GCP (cập nhật đến 2026): Theo Vertex AI pipelines và TensorFlow Extended (TFX), reuse transformation code giúp tránh "training-serving skew". Ví dụ: Sử dụng
@tf.functionhoặc Beam transforms modular để deploy vào serving.
📝 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc:
-
Perform data validation to ensure that the input data to the pipeline is the same format as the input data to the endpoint.
❌ Sai: Chỉ validate format (kiểm tra định dạng dữ liệu đầu vào) không đảm bảo logic preprocessing giống nhau. Validation chỉ phát hiện mismatch muộn (runtime error), không giải quyết gốc rễ (preprocessing khác biệt giữa training/serving). Không hiệu quả cho real-time và không reuse code. -
Refactor the transformation code in the batch data pipeline so that it can be used outside of the pipeline. Use the same code in the endpoint.
✅ Đúng: Như đã giải thích ở trên. Đây là cách tối ưu nhất, đảm bảo consistency tuyệt đối bằng cách chia sẻ chính xác code (modular design). Phù hợp real-time (low latency) và scalable trên Vertex AI (hỗ trợ custom preprocessing trong Prediction SDK đến 2026). -
Refactor the transformation code in the batch data pipeline so that it can be used outside of the pipeline. Share this code with the end users of the endpoint.
❌ Sai: Refactor code là đúng, nhưng chia sẻ code với end users (người dùng cuối endpoint) là không an toàn và không thực tế. End users không nên tự implement preprocessing (dẫn đến inconsistency, security risk). Thay vào đó, integrate code trực tiếp vào endpoint service (nhà phát triển quản lý). -
Batch the real-time requests by using a time window and then use the Dataflow pipeline to preprocess the batched requests. Send the preprocessed requests to the endpoint.
❌ Sai: Batch real-time requests (gom theo time window) làm mất tính real-time (latency cao, không phù hợp use case yêu cầu inference nhanh). Dataflow batch không optimize cho low-latency serving; dùng streaming Dataflow (Apache Beam) phức tạp hơn và vẫn không đảm bảo consistency dễ dàng như refactor code.
📘 Tài liệu tham khảo (cập nhật mới nhất đến 2026)
- GCP Official Docs: Vertex AI Pipelines - Consistent Preprocessing & TensorFlow Transform (TFT) for Reuse – Hướng dẫn refactor Beam transforms cho serving.
- Best Practices MLOps: Google Cloud ML Blog - Training-Serving Consistency (2024 update).
- Apache Beam Docs: Modular Transforms for Batch/Stream – Refactor cho outside-pipeline use.
- Certification Reference: Google Cloud Professional ML Engineer Exam Guide (2025 edition) – Phần "MLOps & Serving".
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần ví dụ code cụ thể, hãy hỏi thêm nhé!
- A Create a BigQuery script to preprocess the data, and write the result to another BigQuery table.
- B Create a pipeline in Vertex AI Pipelines to read the data from BigQuery and preprocess it using a custom preprocessing component.
- C Create a preprocessing function that reads and transforms the data from BigQuery. Create a Vertex AI custom prediction routine that calls the preprocessing function at serving time.
- D Create an Apache Beam pipeline to read the data from BigQuery and preprocess it by using TensorFlow Transform and Dataflow.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc phát triển một mô hình TensorFlow tùy chỉnh (custom TensorFlow model) để sử dụng cho dự đoán trực tuyến (online predictions) trên Google Cloud. Dữ liệu huấn luyện được lưu trữ trong BigQuery, và yêu cầu chính là áp dụng các biến đổi dữ liệu ở mức instance-level (biến đổi từng mẫu dữ liệu riêng lẻ) cho cả giai đoạn huấn luyện (training) và phục vụ mô hình (serving).
Quan trọng nhất: Phải sử dụng cùng một quy trình tiền xử lý (preprocessing routine) cho cả hai giai đoạn để đảm bảo tính nhất quán (consistency), tránh sự khác biệt dẫn đến hiệu suất kém hoặc lỗi dự đoán.
Mục tiêu là cấu hình quy trình tiền xử lý sao cho:
- Đọc dữ liệu từ BigQuery.
- Áp dụng biến đổi instance-level (ví dụ: normalization, encoding, scaling từng hàng dữ liệu).
- Tích hợp mượt mà vào Vertex AI cho training và serving online (như qua Vertex AI Prediction endpoints).
✅ Đây là best practice trong Vertex AI (cập nhật đến 2026): Sử dụng TensorFlow Transform (TFT) kết hợp Apache Beam trên Dataflow để tạo preprocessing graph có thể export và reuse cho serving, đảm bảo tính đồng nhất.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Create an Apache Beam pipeline to read the data from BigQuery and preprocess it by using TensorFlow Transform and Dataflow.
Lý do:
- 🛠️ TensorFlow Transform (TFT) chuyên xử lý biến đổi instance-level (như tf.transform.normalize, vocabulary accumulation) và tạo preprocessing graph (SavedModel) có thể sử dụng lại cho cả training và serving.
- 📊 Apache Beam pipeline trên Dataflow đọc dữ liệu từ BigQuery, chạy TFT để phân tích thống kê toàn bộ dataset (stats analysis), sau đó áp dụng biến đổi và lưu kết quả (transformed data + graph).
- 🔄 Tính nhất quán hoàn hảo: Graph từ TFT được tích hợp vào TensorFlow model serving trên Vertex AI Endpoints, đảm bảo cùng logic preprocessing cho online predictions.
- 🚀 Scalable và production-ready: Dataflow xử lý dữ liệu lớn hiệu quả, hỗ trợ Vertex AI Training/Pipelines (phiên bản mới nhất 2026 vẫn khuyến nghị TFT + Beam cho custom TF models).
Nguồn tham khảo:
- 📘 TensorFlow Transform docs & Vertex AI Preprocessing.
- 📘 Google Cloud Vertex AI Training (cập nhật 2025-2026).
❌ Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá dựa trên khả năng đáp ứng yêu cầu cùng preprocessing routine cho training và serving với dữ liệu BigQuery và custom TF model.
-
[SAI] Create a BigQuery script to preprocess the data, and write the result to another BigQuery table.
❌ Sai vì: BigQuery script (SQL UDF hoặc query) chỉ phù hợp cho batch preprocessing đơn giản, không hỗ trợ instance-level transforms phức tạp như TFT cần stats toàn dataset. Không tạo graph tái sử dụng cho serving, dẫn đến phải viết lại logic riêng cho prediction time → không đảm bảo consistency. Không tích hợp tốt với Vertex AI serving. -
[SAI] Create a pipeline in Vertex AI Pipelines to read the data from BigQuery and preprocess it using a custom preprocessing component.
❌ Sai vì: Vertex AI Pipelines tốt cho training workflow, nhưng custom component chỉ chạy lúc training. Serving time (online predictions) không tự động gọi component này → phải implement riêng preprocessing cho endpoint, vi phạm yêu cầu "same routine". Không scalable cho instance-level transforms lớn. -
[SAI] Create a preprocessing function that reads and transforms the data from BigQuery. Create a Vertex AI custom prediction routine that calls the preprocessing function at serving time.
❌ Sai vì: Hàm preprocessing đọc trực tiếp BigQuery lúc serving sẽ chậm (latency cao) cho online predictions (real-time), không hiệu quả với dữ liệu lớn. Không tích hợp stats analysis toàn dataset như TFT → không consistent (training dùng full data stats, serving chỉ per-instance). Custom prediction routine phức tạp, không best practice cho TF models. -
[ĐÚNG] Create an Apache Beam pipeline to read the data from BigQuery and preprocess it by using TensorFlow Transform and Dataflow.
✅ Đúng vì: Như giải thích ở trên, đây là cách chuẩn và được Google khuyến nghị cho custom TF với BigQuery data. TFT + Beam/Dataflow tạo transformed dataset + serving graph, deploy dễ dàng lên Vertex AI cho online serving với cùng logic. Hoàn hảo cho instance-level transforms! 🎉
- A Use the Apache Airflow SDK to create multiple operators that use Dataflow and Vertex AI services. Deploy the workflow on Cloud Composer.
- B Use the MLFlow SDK and deploy it on a Google Kubernetes Engine cluster. Create multiple components that use Dataflow and Vertex AI services.
- C Use the Kubeflow Pipelines (KFP) SDK to create multiple components that use Dataflow and Vertex AI services. Deploy the workflow on Vertex AI Pipelines.
- D Use the TensorFlow Extended (TFX) SDK to create multiple components that use Dataflow and Vertex AI services. Deploy the workflow on Vertex AI Pipelines.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi tập trung vào việc xây dựng một luồng công việc (workflow) tự động hóa, bảo trì thấp cho mô hình generative TensorFlow text-to-image (chuyển văn bản thành hình ảnh) sử dụng dataset khổng lồ (hàng tỷ hình ảnh kèm caption). Các yêu cầu cụ thể bao gồm:
- 📥 Đọc dữ liệu từ Cloud Storage bucket (GCS).
- 📊 Thu thập thống kê (statistics) về dataset.
- 🔀 Chia dataset thành training/validation/test.
- 🔄 Thực hiện biến đổi dữ liệu (data transformations).
- 🏋️ Huấn luyện mô hình bằng training/validation datasets.
- ✅ Xác thực mô hình (validate) bằng test dataset.
Mục tiêu là tạo workflow tự động, ít bảo trì, tận dụng các dịch vụ Google Cloud như Dataflow (xử lý dữ liệu lớn quy mô) và Vertex AI (nền tảng ML end-to-end). Đây là kịch bản điển hình cho ML pipelines production-grade trên Google Cloud, nhấn mạnh tính tích hợp, scalability và automation. (Kiến thức cập nhật đến 2026: Vertex AI Pipelines v2+ hỗ trợ TFX native, Dataflow 2.0+ với Apache Beam cho batch/streaming).
Nguồn tham khảo:
- Vertex AI Pipelines Documentation
- TFX Official Guide (TFX v1.13+ tích hợp Vertex AI).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use the TensorFlow Extended (TFX) SDK to create multiple components that use Dataflow and Vertex AI services. Deploy the workflow on Vertex AI Pipelines.
Lý do:
- 🛠️ TFX (TensorFlow Extended) là framework chính thức của TensorFlow dành cho end-to-end ML pipelines production, được thiết kế tối ưu cho các bước: ExampleGen (đọc data từ GCS), StatisticsGen (thu thập stats), ExampleSplit (chia train/val/test), Transform (data transformations), Trainer (huấn luyện với Vertex AI), Evaluator (validate với test set).
- 📈 Tích hợp hoàn hảo với Dataflow (cho data processing scalable) và Vertex AI (training/custom jobs), hỗ trợ low maintenance nhờ orchestration tự động, versioning, và monitoring built-in.
- 🚀 Deploy trên Vertex AI Pipelines (hosted Kubeflow-based) cho phép chạy serverless, scalable, với retry/restart tự động – lý tưởng cho dataset billions-scale.
- So với các lựa chọn khác, TFX là native TensorFlow, giảm boilerplate code và đảm bảo consistency cho text-to-image models.
❌ Giải thích tất cả các phương án (đúng/sai)
-
[SAI] Use the Apache Airflow SDK to create multiple operators that use Dataflow and Vertex AI services. Deploy the workflow on Cloud Composer.
❌ Sai vì: Airflow (qua Cloud Composer) là orchestrator tổng quát cho workflows, không chuyên biệt cho ML pipelines. Nó thiếu components ML-native như stats collection, data splitting tự động, hoặc TFX-style transforms. Dù tích hợp Dataflow/Vertex AI, workflow sẽ phức tạp, high maintenance (custom operators nhiều), không tối ưu cho TensorFlow generative models. Cloud Composer phù hợp data ETL hơn ML end-to-end. -
[SAI] Use the MLFlow SDK and deploy it on a Google Kubernetes Engine cluster. Create multiple components that use Dataflow and Vertex AI services.
❌ Sai vì: MLFlow là tracking/experiment tool (logging metrics, models), không phải pipeline orchestrator đầy đủ. Deploy trên GKE yêu cầu self-managed cluster (high maintenance, scaling thủ công), thiếu automation cho data stats/split/transform. Không native hỗ trợ Dataflow/Vertex AI components cho billions-scale data, dễ lỗi khi integrate TensorFlow workflows. -
[SAI] Use the Kubeflow Pipelines (KFP) SDK to create multiple components that use Dataflow and Vertex AI services. Deploy the workflow on Vertex AI Pipelines.
❌ Sai vì: KFP (Kubeflow Pipelines) là general-purpose ML pipelines, linh hoạt nhưng không optimized cho TensorFlow như TFX. Bạn phải custom build components cho stats/split/transform/trainer (nhiều code hơn), dẫn đến less portable và high maintenance. Vertex AI Pipelines hỗ trợ KFP, nhưng TFX là recommended cho TensorFlow (KFP phù hợp multi-framework như PyTorch). -
[ĐÚNG] Use the TensorFlow Extended (TFX) SDK to create multiple components that use Dataflow and Vertex AI services. Deploy the workflow on Vertex AI Pipelines.
✅ Đúng vì: Như giải thích ở trên, TFX cung cấp pre-built components chuẩn cho toàn bộ pipeline (từ data ingestion đến validation), tích hợp seamless Dataflow (Beam pipelines) và Vertex AI (custom training). Vertex AI Pipelines làm fully-managed orchestrator, hỗ trợ scheduling, caching, và cost-optimization – perfect cho low-maintenance, large-scale TensorFlow workflows (xác nhận từ Google Cloud ML best practices 2026).
Tóm tắt khuyến nghị 🎯: Chọn TFX cho TensorFlow pipelines trên GCP để tối ưu hiệu suất và bảo trì. Nếu cần code sample, tham khảo TFX + Vertex AI Tutorial.
- A Use the Vertex AI REST API within a custom component based on a vertex-ai/prediction/xgboost-cpu image
- B Use the Vertex AI ModelEvaluationOp component to evaluate the model
- C Use the Vertex AI SDK for Python within a custom component based on a python:3.10 image
- D Chain the Vertex AI ModelUploadOp and ModelDeployOp components together
Xem giải thích
🧩 Phân tích chi tiết câu hỏi trắc nghiệm
✅ Giải thích nội dung câu hỏi:
Câu hỏi tập trung vào việc phát triển một pipeline ML sử dụng Vertex AI Pipelines (một dịch vụ của Google Cloud Vertex AI). Mục tiêu là pipeline phải upload một phiên bản mới của mô hình XGBoost vào Vertex AI Model Registry (nơi lưu trữ và quản lý các phiên bản mô hình) và sau đó deploy mô hình đó lên Vertex AI Endpoints để phục vụ suy luận trực tuyến (online inference). Yêu cầu sử dụng cách tiếp cận đơn giản nhất (simplest approach).
Đây là tình huống thực tế trong quy trình MLOps trên Google Cloud, nơi bạn cần tự động hóa việc đăng ký và triển khai mô hình mà không cần viết code phức tạp. Vertex AI Pipelines hỗ trợ các component sẵn có để chain (kết nối) các bước này một cách dễ dàng, giúp pipeline chạy tự động và có thể tái sử dụng. Kiến thức dựa trên phiên bản Vertex AI mới nhất (cập nhật đến 2026), với các op như ModelUploadOp và ModelDeployOp được tối ưu cho XGBoost và các mô hình khác.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Chain the Vertex AI ModelUploadOp and ModelDeployOp components together
🛠️ Lý do chi tiết:
- ModelUploadOp là component sẵn có trong Vertex AI Pipelines, dùng để upload artifact mô hình (như XGBoost) vào Model Registry dưới dạng phiên bản mới, tự động xử lý metadata và serving container phù hợp (ví dụ: XGBoost CPU/GPU).
- ModelDeployOp kết nối trực tiếp sau đó để deploy mô hình từ Registry lên một Endpoint mới hoặc cập nhật endpoint hiện có cho online inference.
- Việc chain (kết nối) hai op này là cách đơn giản nhất vì không cần custom code, chỉ cần định nghĩa pipeline bằng Kubeflow DSL hoặc Vertex AI SDK, pipeline sẽ tự động chạy end-to-end. Điều này phù hợp với best practice MLOps trên Vertex AI, giảm thời gian phát triển và lỗi.
📘 Nguồn tham khảo: - Vertex AI Pipelines Components (Google Cloud Docs, cập nhật 2025).
- ModelUploadOp & ModelDeployOp và ModelDeployOp.
📋 Giải thích tất cả các phương án (đúng và sai)
-
Use the Vertex AI REST API within a custom component based on a vertex-ai/prediction/xgboost-cpu image
❌ Sai vì: Phương án này yêu cầu tạo custom component sử dụng REST API của Vertex AI để gọi các endpoint upload/deploy thủ công. Imagevertex-ai/prediction/xgboost-cpuchỉ dành cho serving prediction, không tối ưu cho upload/deploy. Cách này phức tạp hơn (cần viết code gọi API, xử lý auth, error handling), vi phạm yêu cầu "simplest approach". Không tận dụng component sẵn có. -
Use the Vertex AI ModelEvaluationOp component to evaluate the model
❌ Sai vì: ModelEvaluationOp chỉ dùng để đánh giá mô hình (tính metrics như accuracy, ROC-AUC từ test data), không hỗ trợ upload vào Model Registry hay deploy lên Endpoint. Nó thường dùng ở bước trước upload, không thay thế được quy trình end-to-end cần thiết. Sử dụng op này sẽ không đạt mục tiêu. -
Use the Vertex AI SDK for Python within a custom component based on a python:3.10 image
❌ Sai vì: Yêu cầu custom component với SDK Python (gọiaiplatform.Models.upload()vàdeploy()), dựa trên image cơ bảnpython:3.10thiếu pre-built serving cho XGBoost. Cách này không đơn giản vì phải viết code đầy đủ (import SDK, config container, handle artifacts), dễ lỗi và tốn công hơn chain các op sẵn có. -
Chain the Vertex AI ModelUploadOp and ModelDeployOp components together
✅ Đúng vì: Như đã giải thích ở trên, đây là cách tích hợp sẵn, đơn giản nhất trong Vertex AI Pipelines. Chỉ cần pipe output của ModelUploadOp (model resource) vào input của ModelDeployOp, pipeline chạy tự động hỗ trợ XGBoost với container optimized. Hoàn hảo cho production MLOps!
🧠 Lưu ý bổ sung: Trong thực tế (Vertex AI v2025+), bạn có thể compile pipeline như sau:
upload_op = ModelUploadOp(...)
deploy_op = ModelDeployOp(model=upload_op.outputs["upload_model"])
Pipeline này có thể trigger từ Cloud Build hoặc schedule tự động!
- A Use Prophet on Vertex AI Training to build a custom model.
- B Use Vertex AI Forecast to build a NN-based model.
- C Use BigQuery ML to build a statistical ARIMA_PLUS model.
- D Use TensorFlow on Vertex AI Training to build a custom model.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả tình huống bạn làm việc cho một nhà bán lẻ trực tuyến với vài nghìn sản phẩm có lifecycle ngắn (nghĩa là sản phẩm tồn tại ngắn hạn, không kéo dài). Công ty có 5 năm dữ liệu doanh số lưu trữ trong BigQuery. Nhiệm vụ là xây dựng mô hình dự đoán doanh số hàng tháng cho từng sản phẩm. Yêu cầu chính: giải pháp triển khai nhanh chóng với nỗ lực tối thiểu (minimal effort).
📈 Đây là bài toán dự báo chuỗi thời gian (time series forecasting) trên dữ liệu lớn trong BigQuery, cần công cụ tích hợp sẵn, không yêu cầu code phức tạp hay training thủ công để tiết kiệm thời gian và công sức. Các lựa chọn tập trung vào các dịch vụ Google Cloud như Vertex AI và BigQuery ML.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use BigQuery ML to build a statistical ARIMA_PLUS model.
🛠️ Lý do: BigQuery ML cho phép tạo mô hình ARIMA_PLUS (một mô hình thống kê nâng cao cho time series) trực tiếp bằng SQL trong BigQuery, không cần di chuyển dữ liệu hay code Python phức tạp. Với dữ liệu đã sẵn trong BigQuery, bạn chỉ cần chạy lệnh CREATE MODEL là có thể train và predict nhanh chóng cho hàng nghìn sản phẩm. ARIMA_PLUS tự động xử lý seasonality, trend, và hỗ trợ nhiều series cùng lúc (multi-series forecasting), lý tưởng cho sản phẩm lifecycle ngắn. Đây là giải pháp nhanh nhất, ít effort nhất theo best practices của Google Cloud (cập nhật đến 2026, ARIMA_PLUS vẫn là lựa chọn hàng đầu cho time series đơn giản trong BigQuery ML).
📘 Nguồn tham khảo: BigQuery ML Time Series Documentation và ARIMA_PLUS Best Practices.
🔍 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Tôi sử dụng ✅ cho đúng và ❌ cho sai, kèm giải thích rõ ràng:
-
❌ Use Prophet on Vertex AI Training to build a custom model.
Phương án này yêu cầu custom training trên Vertex AI với thư viện Prophet (của Facebook cho time series). Bạn phải viết code Python, chuẩn bị dữ liệu, upload vào Vertex AI Training – tốn nhiều effort (setup pipeline, tuning hyperparameters). Không phù hợp với yêu cầu "minimal effort" vì Prophet không tích hợp sẵn trong BigQuery, phải di chuyển dữ liệu. -
❌ Use Vertex AI Forecast to build a NN-based model.
Vertex AI Forecasting (AutoML Tables cho time series) sử dụng neural network (NN) để dự báo, nhưng cần import dữ liệu vào Vertex AI dataset, cấu hình schema time series, và chờ training tự động (có thể mất hàng giờ với nghìn sản phẩm). Mặc dù ít code hơn custom model, vẫn phức tạp hơn BigQuery ML và không "nhanh nhất" vì phải rời khỏi BigQuery. Phiên bản 2026 vẫn ưu tiên BigQuery ML cho dữ liệu native. -
✅ Use BigQuery ML to build a statistical ARIMA_PLUS model.
Như đã giải thích ở trên: Tích hợp hoàn hảo với dữ liệu BigQuery, chỉ dùng SQL đơn giản (ML.FORECAST), hỗ trợ scale cho hàng nghìn series, xử lý tốt dữ liệu lịch sử 5 năm với lifecycle ngắn. Triển khai trong phút, zero data movement – đúng yêu cầu "quickly with minimal effort". -
❌ Use TensorFlow on Vertex AI Training to build a custom model.
Yêu cầu xây dựng mô hình TensorFlow từ đầu trên Vertex AI Training, bao gồm code deep learning cho time series (LSTM/Transformer), tuning, distributed training. Rất tốn effort (code phức tạp, debug, optimize cho scale), không phù hợp cho quick implementation dù mạnh mẽ cho custom NN.
🏆 Kết luận và khuyến nghị
BigQuery ML ARIMA_PLUS là lựa chọn tối ưu nhất, giúp bạn train và predict ngay trong query BigQuery mà không cần ML engineer chuyên sâu. Nếu dữ liệu phức tạp hơn (ví dụ: external features), có thể kết hợp với Vertex AI sau.
📚 Tài liệu bổ sung: Vertex AI Forecasting vs BigQuery ML Comparison (cập nhật 2026: BigQuery ML vẫn ưu tiên cho SQL-first workflows). Nếu cần code mẫu, thử query sau trong BigQuery:
CREATE OR REPLACE MODEL `project.dataset.sales_model`
OPTIONS(model_type='ARIMA_PLUS', time_series_timestamp_col='date', time_series_data_col='sales', time_series_id_col='product_id') AS
SELECT * FROM `project.dataset.sales_data`;