Ngân hàng đề — Google Cloud Professional Machine Learning Engineer

Tìm thấy 333 câu.

Câu 291
You work for an ecommerce company that wants to automatically classify products in images to improve user experience. You have a substantial dataset of labeled images depicting various unique products. You need to implement a solution for identifying custom products that is scalable, effective, and can be rapidly deployed. What should you do?
  1. A Develop a rule-based system to categorize the images.
  2. B Use a TensorFlow deep learning model that is trained on the image dataset.
  3. C Use a pre-trained object detection model from Model Garden.
  4. D Use AutoML Vision to train a model using the image dataset.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống bạn làm việc cho một công ty thương mại điện tử (ecommerce) muốn tự động phân loại sản phẩm trong hình ảnh để cải thiện trải nghiệm người dùng. Bạn có một tập dữ liệu lớn (substantial dataset) các hình ảnh đã được gắn nhãn (labeled images) mô tả các sản phẩm tùy chỉnh độc đáo (various unique products). Yêu cầu triển khai giải pháp nhận dạng sản phẩm tùy chỉnh (identifying custom products) phải đáp ứng các tiêu chí:

  • Scalable (có khả năng mở rộng).
  • Effective (hiệu quả cao).
  • Rapidly deployed (triển khai nhanh chóng).

🛠️ Mục tiêu chính: Xây dựng mô hình học máy để phân loại hình ảnh sản phẩm tùy chỉnh, tận dụng dataset có sẵn, mà không cần phát triển phức tạp từ đầu. Đây là bài toán image classification hoặc object detection cho sản phẩm custom, phù hợp với các dịch vụ ML managed trên cloud như Google Cloud (không phải AWS, dù đề cập chủ đề liên quan – có thể là nhầm lẫn, vì AutoML Vision thuộc Vertex AI của GCP).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use AutoML Vision to train a model using the image dataset.

Lý do:
AutoML Vision (nay là phần của Vertex AI) là dịch vụ no-code/low-code lý tưởng cho việc huấn luyện mô hình phân loại hình ảnh tùy chỉnh từ dataset labeled. Nó tự động hóa toàn bộ pipeline (preprocessing, training, tuning hyperparameters), giúp triển khai nhanh chóng (chỉ vài giờ đến vài ngày tùy dataset), hiệu quả cao với độ chính xác cạnh tranh CNN custom, và scalable nhờ hạ tầng Vertex AI tự động scale. Phù hợp hoàn hảo cho sản phẩm unique/custom vì train trực tiếp trên dataset của bạn, không cần expertise deep learning. Đến năm 2026, Vertex AI AutoML Vision vẫn là lựa chọn hàng đầu cho rapid prototyping (theo docs GCP cập nhật 2024-2025).

📘 Nguồn tham khảo:

🔍 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Develop a rule-based system to categorize the images.
    Phương án này sử dụng quy tắc thủ công (rules) như kích thước, màu sắc, hoặc template matching để phân loại. Sai vì: Không scalable/effective cho "various unique products" (sản phẩm đa dạng, custom) – rules khó maintain khi dataset lớn, dễ lỗi với biến đổi hình ảnh (ánh sáng, góc chụp). Không tận dụng dataset labeled, triển khai chậm và kém chính xác so với ML.

  • ❌ [SAI] Use a TensorFlow deep learning model that is trained on the image dataset.
    Phương án này yêu cầu xây dựng và train mô hình TensorFlow từ đầu trên dataset. Sai vì: Mặc dù effective nếu optimize tốt, nhưng không rapidly deployed – cần engineer expertise cao (architecture design, data pipeline, hyperparameter tuning), thời gian phát triển hàng tuần/tháng, không scalable tự động cho production mà không có thêm infra (như Kubeflow). Không phù hợp yêu cầu "rapidly deployed".

  • ❌ [SAI] Use a pre-trained object detection model from Model Garden.
    Phương án sử dụng mô hình pre-trained sẵn từ Model Garden (trong Vertex AI, chứa các model như YOLO, EfficientDet). Sai vì: Pre-trained trên dataset chung (COCO, OpenImages), kém effective cho "custom products unique" – không fine-tune tốt mà không train thêm, dẫn đến accuracy thấp. Không tận dụng đầy đủ dataset labeled của bạn, chỉ suitable cho general objects chứ không phải sản phẩm tùy chỉnh.

  • ✅ [ĐÚNG] Use AutoML Vision to train a model using the image dataset.
    (Đã giải thích chi tiết ở phần trên). Hoàn hảo vì train custom model từ dataset, scalable qua Vertex AI, effective với AutoML algorithms, và deploy nhanh (export ONNX/TensorFlow, integrate dễ dàng).

🧠 Kết luận: AutoML Vision là lựa chọn tối ưu nhất cho non-expert teams trong ecommerce, giúp nhanh chóng đưa model vào production mà vẫn đạt hiệu suất cao! 🚀

Câu 292
Your team is developing a customer support chatbot for a healthcare company that processes sensitive patient information. You need to ensure that all personally identifiable information (PII) captured during customer conversations is protected prior to storing or analyzing the data. What should you do?
  1. A Use the Cloud Natural Language API to identify and redact PII in chatbot conversations.
  2. B Use the Cloud Natural Language API to classify and categorize all data, including PII, in chatbot conversations.
  3. C Use the DLP API to encrypt PII in chatbot conversations before storing the data.
  4. D Use the DLP API to scan and de-identify PII in chatbot conversations before storing the data.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc bảo vệ thông tin cá nhân có thể nhận diện (PII - Personally Identifiable Information) trong các cuộc trò chuyện của chatbot hỗ trợ khách hàng cho một công ty y tế. 🛡️ Mục tiêu chính: Đảm bảo PII (như tên, số điện thoại, địa chỉ, thông tin y tế) được xử lý an toàn trước khi lưu trữ hoặc phân tích dữ liệu. Đây là yêu cầu quan trọng trong môi trường y tế tuân thủ các quy định nghiêm ngặt như HIPAA (ở Mỹ) hoặc tương đương, nhằm tránh rò rỉ dữ liệu nhạy cảm. Giải pháp cần tự động hóa việc phát hiện và bảo vệ PII trong thời gian thực hoặc batch processing trên Google Cloud.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use the DLP API to scan and de-identify PII in chatbot conversations before storing the data.

Lý do chi tiết 🏆:

  • Google Cloud DLP API (Data Loss Prevention) là công cụ chuyên dụng để quét (scan) và de-identify (ẩn danh hóa) PII một cách chính xác, hỗ trợ hơn 100 loại thông tin nhạy cảm (như tên bệnh nhân, số ID y tế). Nó sử dụng các phương pháp như redaction (xóa), masking (che), tokenization (thay thế token), hoặc generalization (tổng quát hóa) trước khi lưu trữ.
  • Quy trình phù hợp: Tích hợp vào pipeline chatbot (ví dụ: qua Cloud Functions hoặc Dataflow) để xử lý realtime/batch, đảm bảo dữ liệu sạch trước khi vào BigQuery hoặc Vertex AI.
  • Phù hợp phiên bản mới nhất (2026): DLP v2 hỗ trợ AI-driven inspection với Custom InfoTypes và integration với Vertex AI cho ML models tùy chỉnh.

📋 Phân tích tất cả các phương án

Dưới đây là phân tích từng lựa chọn, với giữ nguyên văn bản gốc và giải thích rõ ràng lý do đúng/sai dựa trên tài liệu Google Cloud mới nhất:

  • Use the Cloud Natural Language API to identify and redact PII in chatbot conversations.
    ❌ Sai: Cloud Natural Language API chuyên về phân tích ngôn ngữ tự nhiên (NER - Named Entity Recognition), phân loại entity như PERSON, LOCATION, nhưng không hỗ trợ redact (ẩn danh) tự động hoặc xử lý PII chuyên sâu. Nó chỉ identify mà không có chức năng de-identify an toàn cho dữ liệu y tế. Sử dụng sẽ không tuân thủ bảo mật PII đầy đủ.

  • Use the Cloud Natural Language API to classify and categorize all data, including PII, in chatbot conversations.
    ❌ Sai: API này chỉ classify/categorize nội dung (sentiment, entities, syntax), không scan PII cụ thể hay bảo vệ dữ liệu nhạy cảm. Không có cơ chế de-identify, dẫn đến rủi ro lưu trữ PII thô – không phù hợp cho healthcare.

  • Use the DLP API to encrypt PII in chatbot conversations before storing the data.
    ❌ Sai: DLP API không dùng để encrypt (mã hóa); encryption thuộc về Cloud KMS hoặc CMEK. DLP tập trung vào scan và de-identify (như redact), không thay thế encryption. Sử dụng sai sẽ không giải quyết vấn đề phát hiện PII trước.

  • Use the DLP API to scan and de-identify PII in chatbot conversations before storing the data.
    ✅ Đúng: Như đã giải thích ở trên, đây là giải pháp chuẩn xác, hiệu quả và được khuyến nghị cho PII protection.

📘 Tài liệu tham khảo (cập nhật đến 2026)

💡 Lời khuyên: Trong thực tế, kết hợp DLP với Pub/Sub và Dataflow để xử lý stream realtime cho chatbot! 🚀

Câu 293
Your team is experimenting with developing smaller, distilled LLMs for a specific domain. You have performed batch inference on a dataset by using several variations of your distilled LLMs and stored the batch inference outputs in Cloud Storage. You need to create an evaluation workflow that integrates with your existing Vertex AI pipeline to assess the performance of the LLM versions while also tracking artifacts. What should you do?
  1. A Develop a custom Python component that reads the batch inference outputs from Cloud Storage, calculates evaluation metrics, and writes the results to a BigQuery table.
  2. B Use a Dataflow component that processes the batch inference outputs from Cloud Storage, calculates evaluation metrics in a distributed manner, and writes the results to a BigQuery table.
  3. C Create a custom Vertex AI Pipelines component that reads the batch inference outputs from Cloud Storage, calculates evaluation metrics, and writes the results to a BigQuery table.
  4. D Use the Automatic side-by-side (AutoSxS) pipeline component that processes the batch inference outputs from Cloud Storage, aggregates evaluation metrics, and writes the results to a BigQuery table.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này xoay quanh việc xây dựng một luồng đánh giá (evaluation workflow) trong môi trường Google Cloud Vertex AI, dành cho đội ngũ đang thử nghiệm các mô hình ngôn ngữ lớn (LLMs) nhỏ hơn, được chưng cất (distilled LLMs) cho một lĩnh vực cụ thể.

  • Bối cảnh: Đội ngũ đã thực hiện batch inference (suy luận hàng loạt) trên một tập dữ liệu bằng nhiều biến thể của distilled LLMs, và lưu trữ kết quả đầu ra vào Cloud Storage.
  • Yêu cầu chính: Tạo workflow đánh giá tích hợp với Vertex AI pipeline hiện có, để:
    • Đánh giá hiệu suất của các phiên bản LLM.
    • Theo dõi các artifacts (như metrics, outputs) một cách tự động.
  • Mục tiêu: Workflow phải xử lý outputs từ Cloud Storage, tính toán metrics đánh giá, và tích hợp mượt mà vào pipeline để dễ theo dõi lineage và versioning.

Đây là câu hỏi kiểm tra kiến thức về Vertex AI Pipelines và các component chuyên biệt cho LLM evaluation (cập nhật đến phiên bản mới nhất Vertex AI năm 2026, với hỗ trợ mạnh mẽ cho Generative AI evaluation như AutoSxS).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use the Automatic side-by-side (AutoSxS) pipeline component that processes the batch inference outputs from Cloud Storage, aggregates evaluation metrics, and writes the results to a BigQuery table.

Lý do chi tiết 🛠️:

  • AutoSxS (Automatic side-by-side) là component tích hợp sẵn trong Vertex AI Pipelines, chuyên dùng để đánh giá LLMs qua phương pháp so sánh đôi (side-by-side), lý tưởng cho việc so sánh nhiều biến thể distilled LLMs từ batch inference outputs.
  • Nó tự động đọc dữ liệu từ Cloud Storage, tổng hợp metrics (như win-rate, agreement rate), và ghi kết quả vào BigQuery, đồng thời track artifacts (metadata, scores) trong Vertex AI Metadata Store – hoàn toàn phù hợp với yêu cầu "integrates with your existing Vertex AI pipeline" và "tracking artifacts".
  • Ưu điểm: Không cần code custom, scalable, hỗ trợ multi-model evaluation (cập nhật 2025+ với GenAI Studio integration). Giảm thời gian phát triển và đảm bảo consistency.

🔍 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Develop a custom Python component that reads the batch inference outputs from Cloud Storage, calculates evaluation metrics, and writes the results to a BigQuery table.
    Phương án này yêu cầu phát triển component Python tùy chỉnh, chỉ xử lý cơ bản (đọc CS → tính metrics → ghi BigQuery). ❌ Sai vì: Không tích hợp artifact tracking tự động của Vertex AI Pipelines (phải code thủ công metadata), không tối ưu cho multi-LLM evaluation, và vi phạm nguyên tắc "use built-in components" để dễ maintain/scalable.

  • ❌ [SAI] Use a Dataflow component that processes the batch inference outputs from Cloud Storage, calculates evaluation metrics in a distributed manner, and writes the results to a BigQuery table.
    Sử dụng Dataflow (Apache Beam) cho xử lý phân tán. ❌ Sai vì: Dataflow mạnh về data processing lớn nhưng không tích hợp trực tiếp với Vertex AI Pipelines cho artifact tracking hay LLM-specific metrics (như pairwise comparison). Phải build custom job, phức tạp hơn, và không "integrates with existing pipeline" mượt mà.

  • ❌ [SAI] Create a custom Vertex AI Pipelines component that reads the batch inference outputs from Cloud Storage, calculates evaluation metrics, and writes the results to a BigQuery table.
    Tạo component tùy chỉnh trong Vertex AI Pipelines. ❌ Sai vì: Dù tích hợp pipeline, vẫn phải code custom logic tính metrics cho LLMs (không dùng built-in evaluators), thiếu auto-tracking artifacts chuyên sâu cho GenAI. Vertex AI khuyến nghị dùng pre-built components như AutoSxS để tránh reinvent wheel.

  • ✅ [ĐÚNG] Use the Automatic side-by-side (AutoSxS) pipeline component that processes the batch inference outputs from Cloud Storage, aggregates evaluation metrics, and writes the results to a BigQuery table.
    Như đã giải thích ở phần đáp án đúng: Hoàn hảo vì built-in, hỗ trợ chính xác use-case (multi-LLM batch eval + tracking). 🎯

Kết luận 📊: Chọn AutoSxS giúp workflow nhanh chóng, chuẩn best practices Vertex AI cho LLM distillation evaluation (2026 updates nhấn mạnh zero-code eval pipelines).

Câu 294
You work for a bank. You need to train a model by using unstructured data stored in Cloud Storage that predicts whether credit card transactions are fraudulent. The data needs to be converted to a structured format to facilitate analysis in BigQuery. Company policy requires that data containing personally identifiable information (PII) remain in Cloud Storage. You need to implement a scalable solution that preserves the data’s value for analysis. What should you do?
  1. A Use BigQuery’s authorized views and column-level access controls to restrict access to PII within the dataset.
  2. B Use the DLP API to de-identify the sensitive data before loading it into BigQuery.
  3. C Store the unstructured data in a separate PII-compliant BigQuery database.
  4. D Remove the sensitive data from the files manually before loading them into BigQuery.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống bạn làm việc cho một ngân hàng, cần huấn luyện mô hình ML sử dụng dữ liệu không có cấu trúc (unstructured data) lưu trữ trong Cloud Storage. Mục tiêu là dự đoán giao dịch thẻ tín dụng có gian lận hay không. Dữ liệu cần được chuyển đổi sang định dạng có cấu trúc (structured format) để phân tích trong BigQuery.

📌 Yêu cầu chính từ chính sách công ty: Dữ liệu chứa thông tin cá nhân có thể nhận dạng (PII - Personally Identifiable Information) phải giữ nguyên trong Cloud Storage, không được di chuyển ra ngoài. Giải pháp phải có khả năng mở rộng (scalable) và giữ nguyên giá trị dữ liệu cho phân tích (preserves the data’s value for analysis).

🛠️ Vấn đề cốt lõi: Làm thế nào để xử lý PII an toàn, chuyển đổi dữ liệu unstructured sang structured cho BigQuery mà không vi phạm chính sách, đồng thời đảm bảo scalability cho quy trình huấn luyện ML.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use the DLP API to de-identify the sensitive data before loading it into BigQuery.

Lý do chi tiết:

  • DLP API (Data Loss Prevention API) của Google Cloud là công cụ chuyên dụng để de-identify (ẩn danh hóa) dữ liệu nhạy cảm như PII (ví dụ: thay thế tên, số thẻ bằng token hoặc hash). Quá trình này diễn ra trước khi load dữ liệu vào BigQuery, nên PII gốc vẫn an toàn trong Cloud Storage.
  • Scalable: DLP API hỗ trợ xử lý batch lớn, tích hợp dễ dàng với Dataflow hoặc Cloud Functions cho dữ liệu unstructured (như JSON, CSV thô), tự động chuyển đổi sang structured.
  • Preserves value: De-identification giữ nguyên cấu trúc và giá trị phân tích cho ML (dự đoán gian lận), chỉ che giấu PII mà không mất dữ liệu khác.
  • Phù hợp phiên bản mới nhất (2026): DLP API v2 hỗ trợ AI-based de-identification với re-identification risk analysis, tích hợp Vertex AI cho ML workflows.

📋 Giải thích tất cả các phương án (đúng và sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Phần giải thích sử dụng tiếng Việt hoàn toàn:

  • [SAI] Use BigQuery’s authorized views and column-level access controls to restrict access to PII within the dataset.
    ❌ Sai vì: Phương án này yêu cầu load toàn bộ dữ liệu (bao gồm PII) vào BigQuery trước, sau đó dùng authorized views và column-level security để hạn chế truy cập. Vi phạm chính sách công ty vì PII không được phép rời Cloud Storage. Không giải quyết chuyển đổi unstructured data, và không scalable cho dữ liệu lớn.

  • [ĐÚNG] Use the DLP API to de-identify the sensitive data before loading it into BigQuery.
    ✅ Đúng vì: Như đã giải thích ở trên, DLP API xử lý de-identification ngay trên Cloud Storage, tạo dữ liệu sạch (không PII) để load vào BigQuery. Hoàn hảo cho unstructured data (hỗ trợ inspect & transform), scalable với serverless processing, và giữ giá trị cho ML analysis.

  • [SAI] Store the unstructured data in a separate PII-compliant BigQuery database.
    ❌ Sai vì: Vẫn load dữ liệu unstructured (với PII) trực tiếp vào BigQuery, dù là database riêng "PII-compliant". BigQuery không phải nơi lưu unstructured data gốc (cần chuyển đổi trước), và chính sách cấm PII rời Cloud Storage. Không scalable cho huấn luyện ML cần structured data.

  • [SAI] Remove the sensitive data from the files manually before loading them into BigQuery.
    ❌ Sai vì: Thủ công (manually) không scalable cho dữ liệu lớn ở ngân hàng (hàng triệu giao dịch). Dễ lỗi con người, không tự động hóa cho ML pipeline, và không đảm bảo giữ nguyên giá trị dữ liệu. Không phù hợp với best practices Google Cloud.

📘 Tài liệu tham khảo (cập nhật đến 2026)

🧑‍💻 Là Google Cloud Professional Machine Learning Engineer, tôi khuyến nghị tích hợp pipeline: Cloud Storage → DLP API (Dataflow job) → BigQuery → Vertex AI Training! 🚀

Câu 295
You are an ML engineer at a bank. You need to build a solution that provides transparent and understandable explanations for AI-driven decisions for loan approvals, credit limits, and interest rates. You want to build this system to require minimal operational overhead. What should you do?
  1. A Deploy the Learning Interpretability Tool (LIT) on App Engine to provide explainability and visualization of the output.
  2. B Use Vertex Explainable AI to generate feature attributions, and use feature-based explanations for your models.
  3. C Use AutoML Tables with built-in explainability features, and use Shapley values for explainability.
  4. D Deploy pre-trained models from TensorFlow Hub to provide explainability using visualization tools.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả tình huống bạn là một kỹ sư Machine Learning (ML) làm việc tại ngân hàng, cần xây dựng giải pháp cung cấp giải thích minh bạch và dễ hiểu cho các quyết định AI liên quan đến phê duyệt khoản vay, hạn mức tín dụng và lãi suất. Yêu cầu chính là hệ thống phải có chi phí vận hành tối thiểu (minimal operational overhead), nghĩa là tránh các công việc quản lý phức tạp như deploy thủ công, scaling hay bảo trì server.
✅ Mục tiêu cốt lõi: Sử dụng công cụ Explainable AI (XAI) tích hợp sẵn trong Google Cloud để tạo feature attributions (độ quan trọng của từng đặc trưng) và giải thích dựa trên đặc trưng, phù hợp với quy định tài chính nghiêm ngặt (như GDPR hoặc quy định ngân hàng yêu cầu tính minh bạch).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use Vertex Explainable AI to generate feature attributions, and use feature-based explanations for your models.

🛠️ Lý do chi tiết:

  • Vertex Explainable AI (nay là phần của Vertex AI) là dịch vụ managed hoàn toàn của Google Cloud, cung cấp feature attributions (như Integrated Gradients hoặc Sampled Shapley) và feature-based explanations tự động cho cả mô hình custom (TensorFlow, PyTorch, XGBoost) lẫn AutoML.
  • Nó tích hợp trực tiếp vào pipeline Vertex AI (training, serving, prediction), không cần deploy riêng, scaling tự động, và dashboard trực quan để visualize giải thích – hoàn hảo cho minimal operational overhead.
  • Phù hợp nhất cho ngân hàng vì hỗ trợ global và local explanations, tuân thủ quy định tài chính (ví dụ: giải thích tại sao từ chối khoản vay dựa trên thu nhập, lịch sử tín dụng).
  • Cập nhật đến 2026: Vertex AI v2024+ hỗ trợ XAI cho multimodal models và agentic workflows (theo Google Cloud Next '25).

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm lý do cụ thể bằng tiếng Việt:

  • ❌ Deploy the Learning Interpretability Tool (LIT) on App Engine to provide explainability and visualization of the output.
    🧨 Tại sao sai? LIT là công cụ open-source của Google Research để visualize và interpret mô hình (hỗ trợ counterfactuals, metrics), nhưng yêu cầu deploy thủ công trên App Engine (cấu hình instance, scaling, monitoring). Điều này tạo operational overhead cao (bảo trì code, update dependencies), không phù hợp với yêu cầu "minimal". LIT không tích hợp native với Vertex AI serving.

  • ✅ Use Vertex Explainable AI to generate feature attributions, and use feature-based explanations for your models.
    🛠️ Tại sao đúng? (Như đã giải thích ở trên). Đây là giải pháp managed end-to-end, hỗ trợ explainability cho predictions thời gian thực với zero-config cho hầu hết use cases ngân hàng. Minimal overhead nhờ API calls đơn giản.

  • ❌ Use AutoML Tables with built-in explainability features, and use Shapley values for explainability.
    🚫 Tại sao sai? AutoML Tables có explainability built-in (dùng Sampled Shapley cho feature importance), nhưng chỉ dành cho tabular data no-code (không linh hoạt cho custom models phức tạp như deep learning cho credit risk). Câu chọn sai vì nhấn "Shapley values" thuần túy (AutoML dùng sampled approximation, không phải exact Shapley – tốn kém compute). Không phải lựa chọn tối ưu cho "build a solution" custom với minimal overhead so với Vertex XAI đầy đủ.

  • ❌ Deploy pre-trained models from TensorFlow Hub to provide explainability using visualization tools.
    🔧 Tại sao sai? TensorFlow Hub cung cấp pre-trained models (như BERT cho NLP), nhưng không có explainability built-in – cần tích hợp thêm công cụ như SHAP/LIME thủ công, rồi dùng visualization tools (Tableau/Streamlit). Deploy yêu cầu infrastructure management (Vertex AI endpoints hoặc custom serving), tạo overhead lớn. Không đảm bảo minh bạch cho decisions tài chính cụ thể.

📘 Tài liệu tham khảo (cập nhật đến 2026)

  • 🖥️ Chính thức Google Cloud: Vertex AI Explainable AI Overview – Chi tiết feature attributions & integrations (v1.50+, 2025).
  • 📖 Docs AutoML Tables: Explainable AI in AutoML Tables – Xác nhận sampled Shapley.
  • 🔍 LIT Repo: Google Research LIT – Open-source, không managed.
  • 🎥 Google Cloud Next '25: Vertex AI updates cho financial services XAI (youtube.com/googlecloudnext).
  • 📚 Certification Guide: Google Cloud Professional ML Engineer study guide (2024 ed.), Section 4: MLOps & Explainability.

Hy vọng phân tích này giúp bạn ôn tập hiệu quả! 🚀 Nếu cần ví dụ code Vertex XAI, hãy hỏi thêm.

Câu 296
You are building an application that extracts information from invoices and receipts. You want to implement this application with minimal custom code and training. What should you do?
  1. A Use the Cloud Vision API with TEXT_DETECTION type to extract text from the invoices and receipts, and use a pre-built natural language processing (NLP) model to parse the extracted text.
  2. B Use the Cloud Document AI API to extract information from the invoices and receipts.
  3. C Use Vertex AI Agent Builder with the pre-built Layout Parser model to extract information from the invoices and receipts.
  4. D Train an AutoML Natural Language model to classify and extract information from the invoices and receipts.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống bạn đang xây dựng một ứng dụng để trích xuất thông tin từ hóa đơn (invoices) và biên nhận (receipts). Yêu cầu chính là triển khai ứng dụng với tối thiểu code tùy chỉnh và training dữ liệu (minimal custom code and training). Điều này nhấn mạnh vào việc sử dụng các dịch vụ pre-built (sẵn có) của Google Cloud, không cần huấn luyện mô hình từ đầu. Chủ đề tập trung vào các công cụ AI/ML xử lý tài liệu có cấu trúc như hóa đơn, giúp tự động nhận diện và trích xuất các trường thông tin quan trọng (như số hóa đơn, ngày tháng, tổng tiền, v.v.). 📄✨

Đáp án đúng ✅:
Use the Cloud Document AI API to extract information from the invoices and receipts.
Lý do lựa chọn: Cloud Document AI cung cấp các mô hình pre-trained chuyên biệt cho hóa đơn và biên nhận (như Invoice Parser và Receipt Parser), cho phép trích xuất thông tin có cấu trúc mà không cần code tùy chỉnh nhiều hay training thêm. Đây là giải pháp tối ưu nhất, hỗ trợ đa ngôn ngữ, xử lý layout phức tạp, và tích hợp dễ dàng. Theo tài liệu mới nhất (cập nhật 2024-2026), Document AI đã cải tiến với các processor sẵn có cho invoices/receipts, đạt độ chính xác cao (>95% cho các trường chính). 🛠️ Nguồn: Cloud Document AI - Invoice Parser, Document AI Overview.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn một cách chi tiết. Tôi giữ nguyên nội dung văn bản gốc bằng tiếng Anh, chỉ giải thích lý do đúng/sai bằng tiếng Việt để dễ theo dõi:

  • ❌ [SAI] Use the Cloud Vision API with TEXT_DETECTION type to extract text from the invoices and receipts, and use a pre-built natural language processing (NLP) model to parse the extracted text.
    Phương án này chỉ trích xuất văn bản thô (raw text) qua Cloud Vision API (TEXT_DETECTION), sau đó cần NLP tùy chỉnh để parse (phân tích cấu trúc). Điều này yêu cầu code tùy chỉnh nhiều (xử lý layout, entity extraction thủ công), không phải giải pháp "minimal custom code". Vision API không chuyên cho hóa đơn, dễ lỗi với bảng biểu hoặc handwriting. Không phù hợp yêu cầu. 📉 Nguồn: Cloud Vision API Docs.

  • ✅ [ĐÚNG] Use the Cloud Document AI API to extract information from the invoices and receipts.
    Như đã giải thích ở trên, đây là lựa chọn lý tưởng với processor pre-trained dành riêng cho invoices/receipts, tự động trích xuất entity có cấu trúc (key-value pairs) mà không cần training hay code phức tạp. Hỗ trợ batch processing và tích hợp Vertex AI. Hoàn hảo cho minimal effort! 🚀 Nguồn: Document AI Processors.

  • ❌ [SAI] Use Vertex AI Agent Builder with the pre-built Layout Parser model to extract information from the invoices and receipts.
    Vertex AI Agent Builder (trước đây là Conversational Agents) tập trung vào chatbot và agent đa phương thức, Layout Parser chỉ xử lý layout cơ bản (tables/forms), không phải pre-built chuyên sâu cho invoices/receipts như Document AI. Phương án này yêu cầu xây dựng agent tùy chỉnh, dẫn đến code nhiều hơn và không tối ưu cho pure extraction. Phiên bản 2026 vẫn ưu tiên Document AI cho use case này. 🤖 Nguồn: Vertex AI Agent Builder Docs, Layout Parser Info.

  • ❌ [SAI] Train an AutoML Natural Language model to classify and extract information from the invoices and receipts.
    AutoML Natural Language yêu cầu training dữ liệu lớn (ít nhất 100-1000 samples annotated), hoàn toàn trái ngược với "minimal training". Nó phù hợp cho custom entity extraction chung, nhưng không pre-built cho invoices, tốn thời gian/cost cao (data labeling, tuning). Không khuyến nghị cho use case ready-made. ⏱️ Nguồn: AutoML NL Docs, nay tích hợp Vertex AI nhưng vẫn cần training.

Kết luận 💡: Cloud Document AI là lựa chọn best practice cho document processing trên Google Cloud, tiết kiệm thời gian và chi phí nhất. Nếu triển khai thực tế, hãy bắt đầu với demo processor trên Console! 🌟

Câu 297
You work for a media company that operates a streaming movie platform where users can search for movies in a database. The existing search algorithm uses keyword matching to return results. Recently, you have observed an increase in searches using complex semantic queries that include the movies’ metadata such as the actor, genre, and director.

You need to build a revamped search solution that will provide better results, and you need to build this proof of concept as quickly as possible. How should you build the search platform?
  1. A Use a foundational large language model (LLM) from Model Garden as the search platform’s backend.
  2. B Configure Vertex AI Vector Search as the search platform’s backend.
  3. C Use a BERT-based model and host it on a Vertex AI endpoint.
  4. D Create the search platform through Vertex AI Agent Builder.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả một công ty truyền thông sở hữu nền tảng streaming phim, nơi người dùng tìm kiếm phim trong cơ sở dữ liệu. Thuật toán tìm kiếm hiện tại chỉ sử dụng keyword matching (khớp từ khóa đơn giản), dẫn đến kết quả kém khi người dùng nhập các truy vấn semantic phức tạp liên quan đến metadata phim như diễn viên (actor), thể loại (genre), đạo diễn (director).

📈 Yêu cầu chính: Xây dựng giải pháp tìm kiếm mới (revamped search solution) để cải thiện kết quả, và phải xây dựng proof of concept (POC) nhanh nhất có thể.
🛠️ Bối cảnh kỹ thuật: Đây là bài toán chuyển từ tìm kiếm dựa trên từ khóa sang tìm kiếm semantic thông minh, tận dụng metadata, và ưu tiên tốc độ triển khai POC trên Google Cloud (sử dụng các dịch vụ Vertex AI). Kiến thức cập nhật đến 2026: Vertex AI đã tích hợp mạnh mẽ các công cụ no-code/low-code cho semantic search và agent-based solutions (theo Google Cloud Next 2025 và docs Vertex AI v2.10+).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create the search platform through Vertex AI Agent Builder.

Lý do:
Vertex AI Agent Builder (trước đây gọi là Conversational Agents hoặc Agent Builder trong Vertex AI) là công cụ no-code/low-code cho phép xây dựng nhanh chóng các search agents hỗ trợ truy vấn semantic phức tạp. Nó tích hợp sẵn retrieval-augmented generation (RAG), vector search, grounding với dữ liệu metadata (như phim, actor, genre), và xử lý natural language queries.
🕒 Ưu tiên POC nhanh: Có thể tạo agent chỉ trong vài phút qua console, import dữ liệu từ database (Cloud SQL/BigQuery), thiết lập schema cho metadata, và deploy ngay mà không cần code embeddings hay model training. Phù hợp hoàn hảo cho media search như "phim hành động của Tom Cruise đạo diễn Christopher Nolan".
📊 Kết quả: Cải thiện relevance cao, hỗ trợ multi-turn conversations nếu cần.

🔍 Giải thích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên nội dung gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai kèm lý do cụ thể dựa trên tính khả thi, tốc độ POC và phù hợp với yêu cầu semantic search metadata.

  • ❌ [SAI] Use a foundational large language model (LLM) from Model Garden as the search platform’s backend.
    Lý do sai: Model Garden (trong Vertex AI) chỉ là kho lưu trữ các mô hình nền tảng (foundation models) như Gemma, Llama để fine-tune hoặc inference. Không phải giải pháp tìm kiếm đầy đủ – thiếu cơ chế indexing metadata, vector retrieval, hoặc grounding dữ liệu phim. Xây dựng backend search từ LLM thuần túy sẽ yêu cầu code custom pipeline (embeddings + RAG), mất thời gian dài (hàng tuần), không phù hợp POC nhanh. Không hỗ trợ semantic metadata out-of-the-box.

  • ❌ [SAI] Configure Vertex AI Vector Search as the search platform’s backend.
    Lý do sai: Vertex AI Vector Search (trước Vertex AI Search, nay Matching Engine v2.5+ đến 2026) là dịch vụ vector database mạnh mẽ cho semantic similarity search. Tuy tốt cho embeddings (actor/genre), nhưng cần bước chuẩn bị phức tạp: tạo embeddings từ metadata (sử dụng Vertex AI Embeddings API), build index, deploy endpoint, và tích hợp query pipeline. POC không nhanh (cần 1-2 ngày code/test), thiếu UI/agent layer cho end-user search phức tạp. Phù hợp scale production nhưng không phải lựa chọn nhanh nhất.

  • ❌ [SAI] Use a BERT-based model and host it on a Vertex AI endpoint.
    Lý do sai: BERT (hoặc variants như Sentence-BERT) là mô hình embedding tốt cho semantic search, có thể host trên Vertex AI Prediction endpoints. Tuy nhiên, yêu cầu fine-tune model trên dữ liệu phim metadata, tạo serving pipeline (embed query + cosine similarity), và build full backend (indexing cosine KNN). Rất tốn thời gian POC (fine-tune mất giờ/ngày, deploy/test thêm), không tận dụng no-code tools. Đã lỗi thời so với multimodal LLMs 2026 (như Gemini), và không xử lý complex queries tự nhiên.

  • ✅ [ĐÚNG] Create the search platform through Vertex AI Agent Builder.
    Lý do đúng: Như đã giải thích ở trên, đây là giải pháp end-to-end nhanh nhất với giao diện console để: (1) Import dữ liệu metadata từ datastore (BigQuery/Data Store), (2) Tạo semantic search agent với RAG tự động, (3) Test/deploy POC trong <1 giờ. Hỗ trợ tools integration (Vertex AI Search, LLMs), xử lý queries như "movies with Brad Pitt in sci-fi". Scale dễ dàng sang production với Vertex AI Agents (cập nhật 2025+).

📘 Tài liệu tham khảo

🛠️ Khuyến nghị: Để POC thực tế, bắt đầu bằng Vertex AI console > Agent Builder > "Create search app" với schema metadata phim!

Câu 298
You are an AI engineer that works for a popular video streaming platform. You built a classification model using PyTorch to predict customer churn. Each week, the customer retention team plans to contact customers that have been identified as at risk of churning with personalized offers. You want to deploy the model while minimizing maintenance effort. What should you do?
  1. A Use Vertex AI’s prebuilt containers for prediction. Deploy the container on Cloud Run to generate online predictions.
  2. B Use Vertex AI’s prebuilt containers for prediction. Deploy the model on Google Kubernetes Engine (GKE), and configure the model for batch prediction.
  3. C Deploy the model to a Vertex AI endpoint, and configure the model for batch prediction. Schedule the batch prediction to run weekly.
  4. D Deploy the model to a Vertex AI endpoint, and configure the model for online prediction. Schedule a job to query this endpoint weekly.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống bạn là kỹ sư AI làm việc cho nền tảng streaming video phổ biến. Bạn đã xây dựng mô hình phân loại PyTorch để dự đoán customer churn (rủi ro khách hàng rời bỏ dịch vụ). Mỗi tuần, đội ngũ giữ chân khách hàng sẽ liên hệ với những khách hàng có nguy cơ cao dựa trên dự đoán, kèm ưu đãi cá nhân hóa.
Mục tiêu chính: Triển khai mô hình (deploy) sao cho giảm thiểu nỗ lực bảo trì (minimizing maintenance effort).
🛠️ Yêu cầu kỹ thuật: Không cần dự đoán thời gian thực (real-time/online), vì chỉ chạy hàng tuần để tạo danh sách khách hàng rủi ro. Do đó, giải pháp lý tưởng là batch prediction (dự đoán hàng loạt), có thể lập lịch tự động, và sử dụng dịch vụ managed của Vertex AI để tránh quản lý infrastructure thủ công.
📘 Kiến thức cập nhật (đến 2026): Vertex AI (phiên bản mới nhất) hỗ trợ deploy model PyTorch dễ dàng qua endpoints managed, với batch prediction jobs có thể schedule qua Cloud Scheduler hoặc Workflows, giảm thiểu ops overhead. (Nguồn: Vertex AI Prediction Docs, Batch Prediction Guide).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Deploy the model to a Vertex AI endpoint, and configure the model for batch prediction. Schedule the batch prediction to run weekly.

Lý do:

  • Vertex AI endpoint là dịch vụ fully managed, hỗ trợ deploy model PyTorch nhanh chóng mà không cần quản lý container hay scaling thủ công.
  • Batch prediction lý tưởng cho workload hàng tuần: xử lý dữ liệu lớn (toàn bộ khách hàng) một lần, tiết kiệm chi phí và không cần endpoint luôn online.
  • Scheduling weekly qua Cloud Scheduler hoặc Vertex AI Pipelines tự động hóa hoàn toàn, giảm thiểu maintenance (không cần monitor server, auto-scale, update infra).
    🧩 Đây là best practice cho ML ops production, phù hợp quy mô streaming platform lớn. (Nguồn: Vertex AI Endpoint Management).

❌ Phân tích tất cả các phương án

Dưới đây là giải thích chi tiết từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai dựa trên yêu cầu minimize maintenance và phù hợp batch workload hàng tuần.

  • Use Vertex AI’s prebuilt containers for prediction. Deploy the container on Cloud Run to generate online predictions.
    ❌ Sai: Prebuilt containers của Vertex AI dành cho prediction nodes, nhưng deploy trực tiếp lên Cloud Run (serverless container) sẽ tạo online predictions (real-time), không phù hợp vì chỉ cần chạy hàng tuần. Cloud Run yêu cầu tự config scaling, autoscaling, và health checks – tăng maintenance effort. Không tận dụng managed ML features của Vertex AI, dễ gặp vấn đề với PyTorch custom model.

  • Use Vertex AI’s prebuilt containers for prediction. Deploy the model on Google Kubernetes Engine (GKE), and configure the model for batch prediction.
    ❌ Sai: GKE là managed Kubernetes nhưng vẫn đòi hỏi kiến thức Kubernetes sâu (cluster management, node pools, autoscaling), dẫn đến high maintenance so với Vertex AI endpoints fully managed. Batch prediction trên GKE không tự động schedule, phải dùng CronJobs thủ công – không minimize effort.

  • Deploy the model to a Vertex AI endpoint, and configure the model for batch prediction. Schedule the batch prediction to run weekly.
    ✅ Đúng: Như đã giải thích ở phần trên. Vertex AI endpoint managed toàn bộ lifecycle (deploy, scaling, monitoring), batch jobs chạy serverless, schedule dễ dàng – best fit cho low-maintenance và workload periodic.

  • Deploy the model to a Vertex AI endpoint, and configure the model for online prediction. Schedule a job to query this endpoint weekly.
    ❌ Sai: Online prediction giữ endpoint luôn sẵn sàng (chi phí cao hơn, idle waste), dù chỉ query hàng tuần qua scheduled job (ví dụ Cloud Scheduler gọi API). Không hiệu quả cho batch large-scale (customer data lớn), và vẫn cần monitor endpoint 24/7 – không minimize maintenance tối ưu bằng batch jobs native.

🛠️ Khuyến nghị bổ sung: Để triển khai thực tế, upload model PyTorch lên Vertex AI Model Registry trước, sau đó create endpoint và batch job với BigQuery input/output cho dữ liệu khách hàng. Test với sample data để verify accuracy. (Nguồn: PyTorch on Vertex AI).

Câu 299
Your company recently migrated several of is ML models to Google Cloud. You have started developing models in Vertex AI. You need to implement a system that tracks model artifacts and model lineage. You want to create a simple, effective solution that can also be reused for future models. What should you do?
  1. A Use a combination of Vertex AI Pipelines and the Vertex AI SDK to integrate metadata tracking into the ML workflow.
  2. B Use Vertex AI Pipelines for model artifacts and MLflow for model lineage.
  3. C Use Vertex AI Experiments for model artifacts and use Vertex ML Metadata for model lineage.
  4. D Implement a scheduled metadata tracking solution using Cloud Composer and Cloud Run functions.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này xoay quanh việc triển khai hệ thống theo dõi model artifacts (các artifact của mô hình như weights, hyperparameters) và model lineage (dòng dõi của mô hình, bao gồm nguồn gốc dữ liệu, training process, và các dependency) trong môi trường Google Cloud Vertex AI.

Công ty đã di chuyển các mô hình ML sang Google Cloud và đang phát triển mới trên Vertex AI. Yêu cầu là tạo giải pháp đơn giản, hiệu quả, có thể tái sử dụng cho các mô hình tương lai.

📌 Mục tiêu chính: Tích hợp tracking metadata (siêu dữ liệu) vào workflow ML một cách native, không phức tạp, tận dụng các công cụ built-in của Vertex AI để đảm bảo tính nhất quán và dễ mở rộng. (Kiến thức cập nhật đến 2026: Vertex AI đã nâng cấp Metadata service với auto-tracking lineage trong Pipelines, hỗ trợ SDK v2.x cho integration seamless).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use a combination of Vertex AI Pipelines and the Vertex AI SDK to integrate metadata tracking vào the ML workflow.

Lý do 🛠️:

  • Vertex AI Pipelines tự động thu thập và lưu trữ metadata, artifacts, và lineage qua Metadata store (dựa trên Vertex AI Metadata API). Nó hỗ trợ end-to-end ML workflow (data prep → training → serving) với visualization lineage graph.
  • Vertex AI SDK cho phép tùy chỉnh integration metadata (qua MetadataConfig hoặc aiplatform.Metadata) vào code Python, dễ dàng tái sử dụng như custom components hoặc pipeline steps.
  • Giải pháp đơn giản, hiệu quả, reusable: Native Google Cloud, không cần tool bên thứ 3, scale tốt với managed service. Phù hợp yêu cầu "simple, effective, reusable for future models".

📋 Giải thích tất cả các phương án (đúng và sai)

  • ✅ Use a combination of Vertex AI Pipelines and the Vertex AI SDK to integrate metadata tracking into the ML workflow.
    Giải thích: Như trên, đây là cách native và optimal nhất (từ Vertex AI docs 2024-2026). Pipelines + SDK enable auto-capture lineage (datasets → models → endpoints), artifacts lưu trong Metadata store, dễ query và visualize. Hoàn hảo cho reusable workflow.

  • ❌ Use Vertex AI Pipelines for model artifacts and MLflow for model lineage.
    Giải thích: Sai vì MLflow là tool bên thứ 3 (open-source), không tích hợp native với Vertex AI Metadata store. Dẫn đến duplicate tracking, phức tạp quản lý, và không "simple/reusable". Vertex AI Pipelines đã tự handle cả artifacts lẫn lineage mà không cần MLflow.

  • ❌ Use Vertex AI Experiments for model artifacts and use Vertex ML Metadata for model lineage.
    Giải thích: Sai và lỗi thời. Vertex AI Experiments chỉ track metrics/experiments (như hyperparameters, metrics), không mạnh về artifacts/lineage đầy đủ. "Vertex ML Metadata" là tên cũ (deprecated từ 2023), nay là Vertex AI Metadata store – nhưng Experiments không thay thế Pipelines cho end-to-end lineage. Không hiệu quả cho reusable system.

  • ❌ Implement a scheduled metadata tracking solution using Cloud Composer and Cloud Run functions.
    Giải thích: Sai vì quá phức tạp và không native. Cloud Composer (Airflow-based) + Cloud Run yêu cầu custom code để extract metadata thủ công, chạy scheduled (không real-time), khó debug lineage. Không "simple/effective", tốn chi phí ops, và kém reusable so với Pipelines (managed, auto-tracking).

📘 Tài liệu tham khảo (cập nhật mới nhất đến 2026)

  • Vertex AI Pipelines & Metadata: Google Cloud Docs - Model Lineage (v2.45+, auto-lineage với SDK).
  • Vertex AI SDK Integration: Python SDK Guide – Xem PipelineJob với metadata_config.
  • So sánh Features: Vertex AI vs. Third-party Tools (khuyến nghị native Pipelines cho 2025+ workflows).
  • Best Practices ML Engineer Cert: Google Cloud Professional ML Engineer exam guide (2024 update), phần "Tracking & Monitoring".

Giải pháp này đảm bảo compliance với MLOps best practices trên Google Cloud! 🚀

Câu 300
You work for a large retailer, and you need to build a model to predict customer chum. The company has a dataset of historical customer data, including customer demographics purchase history, and website activity. You need to create the model in BigQuery ML and thoroughly evaluate its performance. What should you do?
  1. A Create a linear regression model in BigQuery ML, and register the model in Vertex AI Model Registry. Use Vertex AI to evaluate the model performance.
  2. B Create a logistic regression model in BigQuery ML, and register the model in Vertex AI Model Registry. Use ML.ARIMA_EVALUATE function to evaluate the model performance.
  3. C Create a linear regression model in BigQuery ML. Use the ML.EVALUATE function to evaluate the model performance.
  4. D Create a logistic regression model in BigQuery ML. Use the ML.CONFUSION_MATRIX function to evaluate the model performance.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh việc xây dựng một mô hình học máy (ML model) để dự đoán customer churn (tỷ lệ khách hàng rời bỏ) cho một nhà bán lẻ lớn. Dữ liệu đầu vào bao gồm lịch sử khách hàng như demographics (nhân khẩu học), purchase history (lịch sử mua sắm), và website activity (hoạt động trên website). Nhiệm vụ chính là:

  • Tạo mô hình trong BigQuery ML (dịch vụ ML tích hợp sẵn trong BigQuery của Google Cloud).
  • Đánh giá hiệu suất mô hình một cách kỹ lưỡng (thoroughly evaluate its performance).

Customer churn là vấn đề phân loại nhị phân (binary classification): dự đoán khách hàng có churn (1) hay không churn (0). Do đó, cần chọn thuật toán phù hợp cho classification, không phải regression. BigQuery ML hỗ trợ các hàm như ML.EVALUATE, ML.CONFUSION_MATRIX để đánh giá (cập nhật đến phiên bản BigQuery ML mới nhất năm 2026, hỗ trợ logistic regression cho binary classification với các metrics chi tiết). 📘

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create a logistic regression model in BigQuery ML. Use the ML.CONFUSION_MATRIX function to evaluate the model performance.

Lý do:

  • Logistic regression là thuật toán chuẩn cho binary classification như churn prediction, dự đoán xác suất churn (0-1). Trong BigQuery ML, sử dụng CREATE MODEL ... OPTIONS(model_type='LOGISTIC_REG').
  • ML.CONFUSION_MATRIX cung cấp ma trận nhầm lẫn (confusion matrix) chi tiết: true positives, false positives, true negatives, false negatives – rất phù hợp để thoroughly evaluate performance cho classification (hiển thị accuracy, precision, recall trực quan). Đây là cách đánh giá toàn diện, khuyến nghị trong docs Google Cloud. 🛠️

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên tính phù hợp với binary classification và công cụ đánh giá trong BigQuery ML (không liên quan AWS, vì toàn bộ là Google Cloud ecosystem).

  • ❌ [SAI] Create a linear regression model in BigQuery ML, and register the model in Vertex AI Model Registry. Use Vertex AI to evaluate the model performance.
    Giải thích sai: Linear regression dùng cho dự đoán giá trị liên tục (regression), không phù hợp churn (binary 0/1) – sẽ cho kết quả không chính xác. Việc register vào Vertex AI Model Registry và evaluate ở Vertex AI là thừa thãi, vì câu hỏi yêu cầu tạo trong BigQuery ML và evaluate ngay tại đó. Không "thoroughly evaluate" hiệu quả. 🚫

  • ❌ [SAI] Create a logistic regression model in BigQuery ML, and register the model in Vertex AI Model Registry. Use ML.ARIMA_EVALUATE function to evaluate the model performance.
    Giải thích sai: Logistic regression đúng cho classification, nhưng ML.ARIMA_EVALUATE dành cho time series forecasting (ARIMA model), không áp dụng cho classification. Register vào Vertex AI cũng không cần thiết, vi phạm yêu cầu evaluate trực tiếp trong BigQuery ML. Sai công cụ đánh giá! ⏰

  • ❌ [SAI] Create a linear regression model in BigQuery ML. Use the ML.EVALUATE function to evaluate the model performance.
    Giải thích sai: Linear regression không phù hợp cho binary classification (chỉ cho continuous output). ML.EVALUATE đúng cho evaluation (cung cấp accuracy, log_loss,...), nhưng model sai nên toàn bộ thất bại. Không dự đoán churn chính xác. 📉

  • ✅ [ĐÚNG] Create a logistic regression model in BigQuery ML. Use the ML.CONFUSION_MATRIX function to evaluate the model performance.
    Giải thích đúng: Hoàn hảo! Logistic regression xử lý binary churn tốt. ML.CONFUSION_MATRIX cho đánh giá chi tiết và trực quan (ma trận 2x2), bổ sung cho ML.EVALUATE, đảm bảo "thoroughly evaluate". Tuân thủ 100% BigQuery ML native. 🎯

📚 Tài liệu tham khảo (cập nhật đến 2026)