Ngân hàng đề — Google Cloud Professional Machine Learning Engineer

Tìm thấy 333 câu.

Câu 211
You are creating a model training pipeline to predict sentiment scores from text-based product reviews. You want to have control over how the model parameters are tuned, and you will deploy the model to an endpoint after it has been trained. You will use Vertex AI Pipelines to run the pipeline. You need to decide which Google Cloud pipeline components to use. What components should you choose?
  1. A TabularDatasetCreateOp, CustomTrainingJobOp, and EndpointCreateOp
  2. B TextDatasetCreateOp, AutoMLTextTrainingOp, and EndpointCreateOp
  3. C TabularDatasetCreateOp. AutoMLTextTrainingOp, and ModelDeployOp
  4. D TextDatasetCreateOp, CustomTrainingJobOp, and ModelDeployOp
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc xây dựng một pipeline huấn luyện mô hình trên Vertex AI Pipelines (Google Cloud) để dự đoán điểm sentiment từ các đánh giá sản phẩm dựa trên văn bản (text-based product reviews). Các yêu cầu chính:

  • Kiểm soát thủ công cách tuning tham số mô hình (không dùng tự động).
  • Triển khai mô hình lên endpoint sau khi huấn luyện xong.
  • Sử dụng các pipeline components phù hợp của Vertex AI để tạo dataset, huấn luyện và deploy.

Mục tiêu là chọn bộ 3 components chính xác nhất cho quy trình: tạo dataset text → huấn luyện custom → deploy model. Đây là kịch bản thực tế trong Vertex AI (cập nhật đến 2026, theo docs Vertex AI Pipelines v2+), nơi dữ liệu là text thuần nên cần dataset chuyên biệt, huấn luyện tùy chỉnh để control hyperparameters, và deploy chuẩn.

✅ Đáp án đúng: TextDatasetCreateOp, CustomTrainingJobOp, and ModelDeployOp

Lý do lựa chọn:

  • TextDatasetCreateOp: Hoàn hảo cho dữ liệu văn bản (text reviews), import từ CSV/JSON với annotations sentiment. 📊
  • CustomTrainingJobOp: Cho phép kiểm soát hoàn toàn tuning hyperparameters (ví dụ: dùng TensorFlow/PyTorch custom code), không tự động như AutoML. 🛠️
  • ModelDeployOp: Triển khai model đã train lên endpoint Vertex AI một cách chuẩn, tự động scale và serve predictions. 🚀 Bộ này khớp 100% yêu cầu, theo Vertex AI Pipelines best practices (2024-2026).

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do chi tiết:

  • ❌ [SAI] TabularDatasetCreateOp, CustomTrainingJobOp, and EndpointCreateOp
    Phương án này sai vì:

    • TabularDatasetCreateOp chỉ dành cho dữ liệu bảng (structured/tabular như CSV số), không phù hợp với text reviews thuần (unstructured text). Sẽ lỗi khi import text sentiment. ❌
    • CustomTrainingJobOp đúng (control tuning). ✅
    • EndpointCreateOp là component cũ/lỗi thời (deprecated ở Pipelines v2+), chỉ tạo endpoint trống mà không deploy model trực tiếp. Phải dùng ModelDeployOp để attach model. Không khớp deploy sau train. 🗑️
  • ❌ [SAI] TextDatasetCreateOp, AutoMLTextTrainingOp, and EndpointCreateOp
    Phương án này sai vì:

    • TextDatasetCreateOp đúng cho text data. ✅
    • AutoMLTextTrainingOp không cho control tuning (AutoML tự động hóa toàn bộ, dùng managed models), vi phạm yêu cầu "control over how the model parameters are tuned". ❌
    • EndpointCreateOp sai như trên, không deploy model hiệu quả. 🛑
  • ❌ [SAI] TabularDatasetCreateOp. AutoMLTextTrainingOp, and ModelDeployOp
    Phương án này sai vì:

    • TabularDatasetCreateOp sai cho text (như phân tích đầu). ❌
    • AutoMLTextTrainingOp sai vì không control tuning thủ công. ❌
    • ModelDeployOp đúng (deploy chuẩn). ✅
      (Lưu ý: Có dấu chấm lạ "TabularDatasetCreateOp.", nhưng không ảnh hưởng phân tích). Tổng thể không khớp text + custom.
  • ✅ [ĐÚNG] TextDatasetCreateOp, CustomTrainingJobOp, and ModelDeployOp
    Hoàn toàn chính xác như giải thích ở phần đáp án đúng: Text dataset → Custom train (control params) → Deploy endpoint. Đây là flow chuẩn Vertex AI cho NLP tasks như sentiment analysis. 🎯

📘 Tài liệu tham khảo (cập nhật mới nhất 2026)

  • Vertex AI Pipelines Components: Google Cloud Docs - Pipeline Components (TextDatasetCreateOp, CustomTrainingJobOp, ModelDeployOp chi tiết ở v2024.10+).
  • Custom Training Guide: Vertex AI Custom Training – Nhấn mạnh control hyperparameters.
  • Model Deployment: Deploy Models to Endpoint – ModelDeployOp là recommended.
  • Samples: Kubeflow Pipelines repo trên GitHub (google-cloud-pipelines-components), ví dụ sentiment text pipeline.

Nếu cần code sample hoặc pipeline YAML, hãy cho tôi biết! 🚀

Câu 212
Your team frequently creates new ML models and runs experiments. Your team pushes code to a single repository hosted on Cloud Source Repositories. You want to create a continuous integration pipeline that automatically retrains the models whenever there is any modification of the code. What should be your first step to set up the CI pipeline?
  1. A Configure a Cloud Build trigger with the event set as "Pull Request"
  2. B Configure a Cloud Build trigger with the event set as "Push to a branch"
  3. C Configure a Cloud Function that builds the repository each time there is a code change
  4. D Configure a Cloud Function that builds the repository each time a new branch is created
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc thiết lập dòng chảy tích hợp liên tục (CI pipeline) trên Google Cloud Platform (GCP), cụ thể là sử dụng Cloud Source Repositories để lưu trữ mã nguồn. Đội ngũ thường xuyên tạo mô hình ML mới và chạy thí nghiệm, đồng thời đẩy code lên một kho lưu trữ duy nhất. Mục tiêu là tự động retrain các mô hình ML mỗi khi có bất kỳ thay đổi code nào. Câu hỏi yêu cầu bước đầu tiên để thiết lập CI pipeline này.

📌 Yếu tố chính cần lưu ý:

  • Sự kiện kích hoạt (trigger) phải phản ứng với modification of the code (thay đổi code), thường là khi push code lên branch.
  • Sử dụng Cloud Build – dịch vụ CI/CD chuẩn của GCP – là lựa chọn tối ưu cho automation builds từ repo.
  • Kiến thức cập nhật đến 2026: Theo tài liệu GCP mới nhất (Cloud Build v2024+), triggers hỗ trợ các event như PUSH, PULL_REQUEST, với PUSH_TO_BRANCH là phổ biến cho CI tự động trên main/default branch.

Nguồn tham khảo chính:

✅ Đáp án đúng

Configure a Cloud Build trigger with the event set as "Push to a branch"

Lý do lựa chọn:

  • Đây là bước đầu tiên và chính xác nhất để thiết lập CI pipeline. Khi code được push lên branch (ví dụ: main hoặc default branch), trigger sẽ tự động kích hoạt build pipeline, bao gồm retrain ML models.
  • Phù hợp với kịch bản "single repository" và "any modification of the code" – PUSH event capture mọi thay đổi trực tiếp trên branch, không yêu cầu PR.
  • Cloud Build tích hợp native với Cloud Source Repositories, hỗ trợ regex cho branch matching (ví dụ: ^main$), đảm bảo scalability và serverless.

🛠️ Giải thích chi tiết từng phương án

  • Configure a Cloud Build trigger with the event set as "Pull Request" ❌
    Sai vì: Event "Pull Request" chỉ kích hoạt khi tạo hoặc cập nhật PR (pull request), không phải khi merge hoặc push trực tiếp code. Kịch bản yêu cầu trigger mỗi khi có modification code (push lên repo), không đề cập PR workflow. Sử dụng PR trigger phù hợp cho CD (continuous deployment) sau review, không phải CI thuần cho mọi thay đổi.

  • Configure a Cloud Build trigger with the event set as "Push to a branch" ✅
    Đúng vì: Như đã giải thích ở trên, đây là event lý tưởng cho CI tự động. PUSH to branch detect mọi commit/push lên branch chỉ định, kích hoạt build ngay lập tức để retrain models. Đây là best practice cho ML workflows trên GCP (Vertex AI + Cloud Build).

  • Configure a Cloud Function that builds the repository each time there is a code change ❌
    Sai vì: Cloud Functions không phải công cụ CI/CD chuẩn; chúng là serverless functions cho event-driven tasks ngắn hạn, không tối ưu cho build pipeline phức tạp như retrain ML (cần Docker, resources lớn). Phải tự code logic poll repo hoặc dùng Pub/Sub hacky, kém hiệu quả so với Cloud Build triggers native. Không phải "first step" đơn giản.

  • Configure a Cloud Function that builds the repository each time a new branch is created ❌
    Sai vì: Event "new branch created" quá hẹp, chỉ trigger khi tạo branch mới (hiếm xảy ra), không cover "any modification of the code" (push/commit). Cloud Functions vẫn không phù hợp cho CI build, và GCP không có trigger native cho "branch creation" trực tiếp từ Source Repos mà không qua Cloud Build.

💡 Lời khuyên thực hành: Sau bước trigger này, tiếp theo là định nghĩa cloudbuild.yaml với steps như gcloud ai-platform models retrain hoặc Vertex AI pipelines cho ML ops đầy đủ! 🚀

Câu 213
You have built a custom model that performs several memory-intensive preprocessing tasks before it makes a prediction. You deployed the model to a Vertex AI endpoint, and validated that results were received in a reasonable amount of time. After routing user traffic to the endpoint, you discover that the endpoint does not autoscale as expected when receiving multiple requests. What should you do?
  1. A Use a machine type with more memory
  2. B Decrease the number of workers per machine
  3. C Increase the CPU utilization target in the autoscaling configurations.
  4. D Decrease the CPU utilization target in the autoscaling configurations
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả tình huống bạn đã xây dựng một mô hình tùy chỉnh (custom model) thực hiện các tác vụ xử lý trước dữ liệu tiêu tốn nhiều bộ nhớ (memory-intensive preprocessing tasks) trước khi đưa ra dự đoán. Mô hình được triển khai lên Vertex AI endpoint trên Google Cloud, và bạn đã kiểm tra (validated) rằng kết quả trả về trong thời gian hợp lý khi test đơn lẻ. Vấn đề xảy ra: Khi chuyển hướng lưu lượng người dùng thực tế (routing user traffic), endpoint không tự động mở rộng quy mô (autoscale) như mong đợi khi nhận nhiều yêu cầu đồng thời.
📌 Mục tiêu: Tìm giải pháp khắc phục vấn đề autoscaling không hoạt động hiệu quả dưới tải cao, tập trung vào đặc thù của task memory-intensive (bộ nhớ là nút thắt chính, không phải CPU).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Decrease the CPU utilization target in the autoscaling configurations

Lý do chi tiết:
🛠️ Vertex AI endpoint sử dụng autoscaling dựa trên CPU utilization làm metric chính (mặc định 60%). Với các task memory-intensive, CPU usage thường thấp vì bottleneck ở RAM (preprocessing ngốn bộ nhớ, dẫn đến swap hoặc chậm nhưng CPU không max). Kết quả: Hệ thống không trigger scale-up kịp thời.
Giảm CPU utilization target (ví dụ: từ 60% xuống 40-50%) sẽ khiến autoscaling kích hoạt sớm hơn khi CPU chạm ngưỡng thấp, tăng replica kịp thời để xử lý tải. Đây là giải pháp trực tiếp và hiệu quả nhất theo tài liệu Vertex AI (cập nhật đến 2026).
📘 Nguồn tham khảo:

❌ Phân tích tất cả các phương án

Dưới đây là giải thích từng phương án một cách chi tiết, đánh dấu đúng/sai dựa trên ngữ cảnh memory-intensive preprocessing và autoscaling Vertex AI:

  • ❌ [SAI] Use a machine type with more memory
    🧩 Phương án này tăng bộ nhớ máy (machine type lớn hơn, ví dụ n1-standard-8 lên n1-standard-16). Nó có thể giảm latency đơn lẻ bằng cách tránh OOM (Out of Memory), nhưng KHÔNG giải quyết vấn đề autoscaling. Endpoint vẫn scale dựa trên CPU (không phải memory metric trực tiếp), nên dưới tải cao, CPU thấp vẫn không trigger replica mới. Đây chỉ là fix tạm thời, không tối ưu cho scale-out.

  • ❌ [SAI] Decrease the number of workers per machine
    🛠️ Giảm số worker/node (ví dụ từ 4 xuống 2) sẽ tăng bộ nhớ/request (vì ít worker chia sẻ RAM), giúp tránh memory pressure từng request. Tuy nhiên, KHÔNG ảnh hưởng trực tiếp đến autoscaling trigger. Vấn đề cốt lõi là CPU threshold cao khiến scale chậm; phương án này chỉ tweak resource allocation, không làm endpoint scale nhanh hơn dưới tải đột biến.

  • ❌ [SAI] Increase the CPU utilization target in the autoscaling configurations.
    📉 Tệ hơn nữa: Tăng target (ví dụ từ 60% lên 80%) sẽ khiến autoscaling kích hoạt MUỘN HƠN, vì yêu cầu CPU cao hơn mới scale. Với memory-intensive task (CPU thấp), endpoint sẽ overload lâu hơn, dẫn đến latency cao hoặc failure. Hoàn toàn ngược với nhu cầu!

  • ✅ [ĐÚNG] Decrease the CPU utilization target in the autoscaling configurations
    (Đã giải thích chi tiết ở phần trên). Đây là best practice cho trường hợp CPU không phản ánh tải thực (memory-bound workloads).

🏆 Kết luận và lưu ý

🔍 Mẹo pro: Trong Vertex AI (cập nhật 2026), bạn có thể kết hợp custom metrics (như memory utilization qua Cloud Monitoring) nếu cần scale dựa trên RAM, nhưng CPU target là cách nhanh nhất. Test với gcloud ai endpoints deploy-model và theo dõi Vertex AI Dashboard. Nếu vấn đề persist, kiểm tra minReplicas/maxReplicas hoặc concurrency per replica!
📚 Tài liệu bổ sung: Vertex AI Resource Management.

Câu 214
Your company manages an ecommerce website. You developed an ML model that recommends additional products to users in near real time based on items currently in the user’s cart. The workflow will include the following processes:

1. The website will send a Pub/Sub message with the relevant data and then receive a message with the prediction from Pub/Sub
2. Predictions will be stored in BigQuery
3. The model will be stored in a Cloud Storage bucket and will be updated frequently

You want to minimize prediction latency and the effort required to update the model. How should you reconfigure the architecture?
  1. A Write a Cloud Function that loads the model into memory for prediction. Configure the function to be triggered when messages are sent to Pub/Sub.
  2. B Create a pipeline in Vertex AI Pipelines that performs preprocessing, prediction, and postprocessing. Configure the pipeline to be triggered by a Cloud Function when messages are sent to Pub/Sub.
  3. C Expose the model as a Vertex AI endpoint. Write a custom DoFn in a Dataflow job that calls the endpoint for prediction.
  4. D Use the RunInference API with WatchFilePattern in a Dataflow job that wraps around the model and serves predictions.
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi mô tả một hệ thống thương mại điện tử (ecommerce website) sử dụng mô hình ML để gợi ý sản phẩm bổ sung gần thời gian thực (near real-time) dựa trên các mặt hàng trong giỏ hàng của người dùng. Quy trình hiện tại bao gồm:

  • Bước 1: Website gửi tin nhắn Pub/Sub chứa dữ liệu liên quan, sau đó nhận dự đoán (prediction) từ Pub/Sub.
  • Bước 2: Các dự đoán được lưu trữ vào BigQuery.
  • Bước 3: Mô hình ML được lưu trong Cloud Storage bucket và cập nhật thường xuyên.

Mục tiêu là tái cấu trúc kiến trúc (reconfigure the architecture) để giảm thiểu độ trễ dự đoán (prediction latency) và giảm nỗ lực cập nhật mô hình (effort required to update the model). Đây là bài toán điển hình trong Google Cloud Platform (GCP) về online inference với yêu cầu low-latency, streaming data qua Pub/Sub, và model serving tự động update từ GCS. (Lưu ý: Mặc dù người dùng đề cập "AWS", nội dung câu hỏi hoàn toàn thuộc GCP với các dịch vụ như Pub/Sub, BigQuery, Vertex AI, Dataflow).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use the RunInference API with WatchFilePattern in a Dataflow job that wraps around the model and serves predictions.

Lý do:

  • RunInference API trong Dataflow (tính năng mới nhất từ Vertex AI, cập nhật đến 2024-2026) cho phép chạy inference trực tiếp trong Dataflow job mà không cần endpoint riêng, giảm latency xuống mức sub-second (gần real-time) nhờ xử lý streaming trực tiếp từ Pub/Sub.
  • WatchFilePattern tự động theo dõi và tải model mới từ Cloud Storage khi có thay đổi (pattern matching file), giảm nỗ lực update (không cần restart job hay redeploy).
  • Phù hợp hoàn hảo: Input từ Pub/Sub → Dataflow job xử lý prediction → Output lưu BigQuery, đảm bảo low-latency và scalable cho ecommerce traffic cao.
  • 📘 Tài liệu tham khảo: Cloud Dataflow RunInference Docs & Vertex AI Model Serving (cập nhật 2024).

🛠️ Phân tích chi tiết tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá đúng/sai kèm lý do cụ thể dựa trên yêu cầu low-latency và easy-update:

  • ❌ [SAI] Write a Cloud Function that loads the model into memory for prediction. Configure the function to be triggered when messages are sent to Pub/Sub.

    • Lý do sai: Cloud Functions có cold start latency cao (có thể >1s), không phù hợp near real-time. Việc load model lớn từ GCS vào memory mỗi lần trigger tốn thời gian và giới hạn memory (128MB-10GB). Update model yêu cầu redeploy function, tăng effort. Không scalable cho traffic ecommerce cao.
  • ❌ [SAI] Create a pipeline in Vertex AI Pipelines that performs preprocessing, prediction, and postprocessing. Configure the pipeline to be triggered by a Cloud Function when messages are sent to Pub/Sub.

    • Lý do sai: Vertex AI Pipelines dành cho batch ML workflows (training/retraining), không phải real-time inference → latency cao (phút thay vì ms). Trigger qua Cloud Function thêm overhead. Update model phức tạp, yêu cầu rerun pipeline, không giảm effort và không tối ưu cho Pub/Sub streaming.
  • ❌ [SAI] Expose the model as a Vertex AI endpoint. Write a custom DoFn in a Dataflow job that calls the endpoint for prediction.

    • Lý do sai: Gọi Vertex AI endpoint từ Dataflow DoFn tạo network latency (HTTP calls, ~100-500ms), không đạt near real-time. Custom DoFn phức tạp, coding nhiều. Update model trên endpoint dễ dàng nhưng vẫn cần quản lý endpoint riêng, tăng effort so với auto-watch.
  • ✅ [ĐÚNG] Use the RunInference API with WatchFilePattern in a Dataflow job that wraps around the model and serves predictions.

    • Lý do đúng (như đã giải thích ở trên): Low-latency streaming inference native trong Dataflow, auto-update model qua WatchFilePattern từ GCS, dễ integrate Pub/Sub → BigQuery. Scalable, serverless, và là best practice mới nhất cho production ML serving (2024+). Giảm effort tối đa! 🚀
Câu 215
You are collaborating on a model prototype with your team. You need to create a Vertex AI Workbench environment for the members of your team and also limit access to other employees in your project. What should you do?
  1. A 1. Create a new service account and grant it the Notebook Viewer role
    2. Grant the Service Account User role to each team member on the service account
    3. Grant the Vertex AI User role to each team member
    4. Provision a Vertex AI Workbench user-managed notebook instance that uses the new service account
  2. B 1. Grant the Vertex AI User role to the default Compute Engine service account
    2. Grant the Service Account User role to each team member on the default Compute Engine service account
    3. Provision a Vertex AI Workbench user-managed notebook instance that uses the default Compute Engine service account.
  3. C 1. Create a new service account and grant it the Vertex AI User role
    2. Grant the Service Account User role to each team member on the service account
    3. Grant the Notebook Viewer role to each team member.
    4. Provision a Vertex AI Workbench user-managed notebook instance that uses the new service account
  4. D 1. Grant the Vertex AI User role to the primary team member
    2. Grant the Notebook Viewer role to the other team members
    3. Provision a Vertex AI Workbench user-managed notebook instance that uses the primary user’s account
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi yêu cầu thiết lập môi trường Vertex AI Workbench (một phần của Google Cloud Vertex AI) để các thành viên trong nhóm có thể cộng tác phát triển mô hình prototype, đồng thời hạn chế quyền truy cập cho các nhân viên khác trong cùng project.

  • Mục tiêu chính:

    • Tạo môi trường notebook chia sẻ an toàn cho team cụ thể.
    • Sử dụng user-managed notebook instance trong Vertex AI Workbench (nay là một phần của Vertex AI Studio, cập nhật đến 2026 với các tính năng IAM tinh chỉnh hơn).
    • Áp dụng nguyên tắc least privilege (quyền tối thiểu) qua IAM roles và service accounts để tránh rò rỉ quyền truy cập project-wide.
  • Bối cảnh kỹ thuật: Vertex AI Workbench hỗ trợ hai loại notebook: user-managed (tự quản lý với service account tùy chỉnh) và managed (Google quản lý). Để chia sẻ an toàn, ưu tiên user-managed với service account riêng, tránh dùng default Compute Engine SA (dễ bị lạm dụng). Team members cần quyền impersonate service account (qua Service Account User role) và quyền xem/chỉnh notebook cụ thể (qua Notebook Viewer/Editor).

📘 Tài liệu tham khảo:

✅ Đáp án đúng: Phương án 3

Lý do chọn 🛠️:

  • Phương án này tuân thủ best practices GCP: Tạo service account riêng (new SA) để isolate quyền, grant Vertex AI User role (roles/aiplatform.user) cho SA để chạy notebooks/đào tạo model.
  • Team members được grant Service Account User role (roles/iam.serviceAccountUser) trên SA → họ có thể impersonate SA mà không cần quyền project-wide.
  • Thêm Notebook Viewer role (roles/aiplatform.notebookViewer) cho team → cho phép xem notebook mà không chỉnh sửa project.
  • Cuối cùng, provision user-managed notebook dùng SA mới → môi trường chia sẻ an toàn, hạn chế access cho người ngoài team.
  • Hoàn hảo cho collaboration prototype mà không expose quyền rộng! (Cập nhật 2026: Vertex AI hỗ trợ fine-grained access cho Workbench instances).

❌ Phân tích tất cả các phương án

  • Phương án 1:

    1. Create a new service account and grant it the Notebook Viewer role
    2. Grant the Service Account User role to each team member on the service account
    3. Grant the Vertex AI User role to each team member
    4. Provision a Vertex AI Workbench user-managed notebook instance that uses the new service account
      Sai vì ❌: Bước 1 chỉ grant Notebook Viewer cho SA (chỉ xem notebook, không đủ quyền chạy Vertex AI jobs/model training). Bước 3 grant Vertex AI User trực tiếp cho từng member → vi phạm least privilege, họ có quyền full project-wide, không hạn chế access như yêu cầu.
  • Phương án 2:

    1. Grant the Vertex AI User role to the default Compute Engine service account
    2. Grant the Service Account User role to each team member on the default Compute Engine service account
    3. Provision a Vertex AI Workbench user-managed notebook instance that uses the default Compute Engine service account.
      Sai vì ❌: Sử dụng default Compute Engine SA (như project@.iam.gserviceaccount.com) không an toàn – nó có quyền Compute Engine rộng, dễ bị exploit nếu team member lạm dụng. GCP khuyến cáo KHÔNG dùng default SA cho Workbench chia sẻ (docs 2025 cảnh báo security risks).
  • Phương án 3 (Đúng ✅):

    1. Create a new service account and grant it the Vertex AI User role
    2. Grant the Service Account User role to each team member on the service account
    3. Grant the Notebook Viewer role to each team member.
    4. Provision a Vertex AI Workbench user-managed notebook instance that uses the new service account
      Đúng vì ✅: Như giải thích ở trên – SA riêng với quyền chính xác, impersonation an toàn, và Notebook Viewer bổ sung cho collaboration/viewing. Hoàn chỉnh và secure!
  • Phương án 4:

    1. Grant the Vertex AI User role to the primary team member
    2. Grant the Notebook Viewer role to the other team members
    3. Provision a Vertex AI Workbench user-managed notebook instance that uses the primary user’s account
      Sai vì ❌: Không dùng service account → primary member phải share credentials cá nhân (rủi ro cao). Các member khác chỉ Viewer (không chạy code), và instance dùng user account → không scale cho team, vi phạm isolation. Không hỗ trợ impersonation đúng cách.
Câu 216
You work at a leading healthcare firm developing state-of-the-art algorithms for various use cases. You have unstructured textual data with custom labels. You need to extract and classify various medical phrases with these labels. What should you do?
  1. A Use the Healthcare Natural Language API to extract medical entities
  2. B Use a BERT-based model to fine-tune a medical entity extraction model
  3. C Use AutoML Entity Extraction to train a medical entity extraction model
  4. D Use TensorFlow to build a custom medical entity extraction model
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống bạn làm việc tại một công ty y tế hàng đầu, phát triển các thuật toán tiên tiến cho nhiều trường hợp sử dụng. Bạn có dữ liệu văn bản không cấu trúc (unstructured textual data) kèm theo nhãn tùy chỉnh (custom labels). Nhiệm vụ là trích xuất và phân loại các cụm từ y tế (medical phrases) bằng chính các nhãn này.
📌 Yêu cầu chính: Cần một giải pháp xử lý entity extraction (trích xuất thực thể) cho dữ liệu y tế với nhãn tùy chỉnh, nghĩa là phải huấn luyện mô hình dựa trên dữ liệu đã gắn nhãn sẵn, thay vì dùng mô hình pre-trained cố định.
🛠️ Bối cảnh Google Cloud: Đây là vấn đề về Vertex AI (trước đây là AutoML Natural Language), phù hợp cho dữ liệu không cấu trúc với custom labels trong lĩnh vực y tế.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Use AutoML Entity Extraction to train a medical entity extraction model
Lý do:

  • AutoML Entity Extraction (nay thuộc Vertex AI Entity Extraction) cho phép huấn luyện mô hình tùy chỉnh chỉ với dữ liệu văn bản đã gắn nhãn (custom labels), không cần code phức tạp.
  • Nó chuyên xử lý entity extraction trên dữ liệu không cấu trúc, hỗ trợ y tế với độ chính xác cao.
  • Phù hợp nhất vì dễ dàng, nhanh chóng (no-code/low-code), tận dụng dữ liệu custom mà không cần kiến thức sâu về ML.
    📘 Nguồn tham khảo: Vertex AI Documentation - Entity Extraction (cập nhật 2024-2026, hỗ trợ custom training với unlabeled data augmentation).

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai với lý do chi tiết:

  • Use the Healthcare Natural Language API to extract medical entities
    ❌ Sai: API này là mô hình pre-trained (đã huấn luyện sẵn) chỉ trích xuất các thực thể y tế chuẩn (như bệnh, thuốc, triệu chứng) theo schema cố định của Google. Không hỗ trợ custom labels – bạn không thể huấn luyện thêm với dữ liệu riêng. Phù hợp cho dữ liệu chung, không phải custom.
    📘 Nguồn: Cloud Healthcare API - NLP (không custom training).

  • Use a BERT-based model to fine-tune a medical entity extraction model
    ❌ Sai: BERT (hoặc BioBERT/MedBERT) có thể fine-tune cho entity extraction, nhưng yêu cầu kiến thức chuyên sâu về ML, code TensorFlow/PyTorch, quản lý training/inference. Quá phức tạp và tốn tài nguyên so với nhu cầu đơn giản có custom labels. Không phải lựa chọn tối ưu cho non-expert.
    🛠️ Lý do loại: Vi phạm nguyên tắc "ít công sức nhất" trong Google Cloud best practices.

  • Use AutoML Entity Extraction to train a medical entity extraction model
    ✅ Đúng: Như đã giải thích ở trên, đây là giải pháp tự động hóa hoàn hảo cho custom entity extraction trên dữ liệu văn bản y tế. Hỗ trợ upload dataset với labels, auto-train, deploy dễ dàng. Hiệu suất cao với dữ liệu y tế (tích hợp de-identification).
    📘 Nguồn: Vertex AI - Train custom entity extraction model (hỗ trợ đến 2026 với improvements như multimodal).

  • Use TensorFlow to build a custom medical entity extraction model
    ❌ Sai: TensorFlow cho phép xây dựng từ đầu (custom NER model), nhưng đòi hỏi code thủ công toàn bộ (data prep, training, evaluation, deployment). Tốn thời gian, chi phí cao, và không tận dụng managed service. Không phù hợp khi có AutoML sẵn.
    📘 Nguồn: TensorFlow Hub - NER models (low-level, không managed).

Câu 217
You developed a custom model by using Vertex AI to predict your application's user churn rate. You are using Vertex AI Model Monitoring for skew detection. The training data stored in BigQuery contains two sets of features - demographic and behavioral. You later discover that two separate models trained on each set perform better than the original model. You need to configure a new model monitoring pipeline that splits traffic among the two models. You want to use the same prediction-sampling-rate and monitoring-frequency for each model. You also want to minimize management effort. What should you do?
  1. A Keep the training dataset as is. Deploy the models to two separate endpoints, and submit two Vertex AI Model Monitoring jobs with appropriately selected feature-thresholds parameters.
  2. B Keep the training dataset as is. Deploy both models to the same endpoint and submit a Vertex AI Model Monitoring job with a monitoring-config-from-file parameter that accounts for the model IDs and feature selections.
  3. C Separate the training dataset into two tables based on demographic and behavioral features. Deploy the models to two separate endpoints, and submit two Vertex AI Model Monitoring jobs.
  4. D Separate the training dataset into two tables based on demographic and behavioral features. Deploy both models to the same endpoint, and submit a Vertex AI Model Monitoring job with a monitoring-config-from-file parameter that accounts for the model IDs and training datasets.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc cấu hình pipeline giám sát mô hình (Vertex AI Model Monitoring) cho phát hiện skew (sự lệch dữ liệu) trong Vertex AI (Google Cloud).

  • Bối cảnh: Bạn đã huấn luyện một mô hình tùy chỉnh trên Vertex AI để dự đoán tỷ lệ churn (khách hàng rời bỏ) của ứng dụng. Dữ liệu huấn luyện lưu trong BigQuery bao gồm hai bộ features: demographic (dân số học) và behavioral (hành vi). Sau đó, phát hiện ra rằng hai mô hình riêng biệt (một dùng demographic, một dùng behavioral) hoạt động tốt hơn mô hình gốc.
  • Yêu cầu chính:
    • Cấu hình pipeline giám sát mới để chia traffic (phân bổ lưu lượng) giữa hai mô hình này.
    • Sử dụng cùng prediction-sampling-rate (tỷ lệ lấy mẫu dự đoán) và monitoring-frequency (tần suất giám sát) cho cả hai.
    • Tối thiểu hóa nỗ lực quản lý (minimize management effort) – nghĩa là tránh tạo nhiều endpoint/job riêng lẻ, tận dụng tính năng hỗ trợ multi-model của Vertex AI.
  • Mục tiêu: Giám sát skew detection trên dữ liệu dự đoán so với training data, tập trung vào features tương ứng cho từng mô hình, mà không cần thay đổi lớn dataset gốc.

📘 Kiến thức cập nhật (Vertex AI phiên bản mới nhất 2024-2026): Vertex AI hỗ trợ multi-model endpoints (deploy nhiều mô hình vào một endpoint duy nhất), và Model Monitoring job có thể monitor nhiều mô hình cùng lúc qua file config YAML/JSON với tham số monitoring-config-from-file. Điều này cho phép chỉ định modelIds và featureSelections riêng cho từng mô hình, giữ nguyên dataset gốc trong BigQuery. Không cần tách dataset vì skew detection dựa trên features subset. (Nguồn: Vertex AI Model Monitoring docs, Multi-model endpoints, cập nhật Q1 2026).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Keep the training dataset as is. Deploy both models to the same endpoint and submit a Vertex AI Model Monitoring job with a monitoring-config-from-file parameter that accounts for the model IDs and feature selections.

🛠️ Lý do chi tiết:

  • Giữ nguyên training dataset: Không cần tách BigQuery table vì skew detection so sánh features subset (demographic/behavioral) từ cùng dataset gốc.
  • Deploy both models to the same endpoint: Vertex AI hỗ trợ Online Prediction với multi-model endpoint, tự động split traffic dựa trên model ID hoặc routing, giảm effort quản lý (chỉ 1 endpoint thay vì 2).
  • One Model Monitoring job với monitoring-config-from-file: File config chỉ định modelIds (cho từng mô hình), featureSelections (features riêng cho demographic/behavioral), cùng predictionSamplingRate và monitoringFrequency. Điều này tối ưu hóa effort – chỉ 1 job duy nhất monitor cả hai, tự động áp dụng config phù hợp.
  • Lợi ích: Minimize management (1 endpoint + 1 job), đảm bảo skew detection chính xác trên features tương ứng.

📋 Phân tích tất cả các phương án (A, B, C, D)

Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá ✅ (đúng) hoặc ❌ (sai), kèm giải thích bằng tiếng Việt.

  • Phương án 1: Keep the training dataset as is. Deploy the models to two separate endpoints, and submit two Vertex AI Model Monitoring jobs with appropriately selected feature-thresholds parameters.
    ❌ Sai vì: Tạo hai endpoint riêng và hai monitoring jobs làm tăng effort quản lý (deploy/maintain nhiều tài nguyên). Vertex AI không khuyến khích cách này khi có multi-model endpoint. feature-thresholds chỉ là tham số skew, không giải quyết split traffic hiệu quả.

  • Phương án 2: Keep the training dataset as is. Deploy both models to the same endpoint and submit a Vertex AI Model Monitoring job with a monitoring-config-from-file parameter that accounts for the model IDs and feature selections.
    ✅ Đúng hoàn toàn (như giải thích ở phần đáp án trên). Đây là cách tối ưu nhất theo best practices Vertex AI 2026, tận dụng config file để customize cho từng model mà không tách dataset hay nhân đôi job/endpoint.

  • Phương án 3: Separate the training dataset into two tables based on demographic and behavioral features. Deploy the models to two separate endpoints, and submit two Vertex AI Model Monitoring jobs.
    ❌ Sai vì: Tách dataset thành hai BigQuery tables là không cần thiết và tốn kém (duplicate data, quản lý schema phức tạp). Kết hợp với hai endpoint + hai jobs càng tăng effort cao, vi phạm yêu cầu minimize management. Skew detection có thể dùng subset features từ dataset gốc.

  • Phương án 4: Separate the training dataset into two tables based on demographic and behavioral features. Deploy both models to the same endpoint, and submit a Vertex AI Model Monitoring job with a monitoring-config-from-file parameter that accounts for the model IDs and training datasets.
    ❌ Sai vì: Tách dataset vẫn thừa thãi và phức tạp hóa (phải migrate data). Config file nên dùng featureSelections (features subset), không phải trainingDatasets riêng – Vertex AI Model Monitoring ưu tiên so sánh features từ một dataset gốc để tránh inconsistency.

🔍 Tóm tắt khuyến nghị: Chọn phương án 2 để scale dễ dàng, tiết kiệm chi phí (ít tài nguyên hơn), và phù hợp production. Nếu cần code sample config YAML, tham khảo Vertex AI Monitoring Pipeline API.

Câu 218
You work for a pharmaceutical company based in Canada. Your team developed a BigQuery ML model to predict the number of flu infections for the next month in Canada. Weather data is published weekly, and flu infection statistics are published monthly. You need to configure a model retraining policy that minimizes cost. What should you do?
  1. A Download the weather and flu data each week. Configure Cloud Scheduler to execute a Vertex AI pipeline to retrain the model weekly.
  2. B Download the weather and flu data each month. Configure Cloud Scheduler to execute a Vertex AI pipeline to retrain the model monthly.
  3. C Download the weather and flu data each week. Configure Cloud Scheduler to execute a Vertex AI pipeline to retrain the model every month.
  4. D Download the weather data each week, and download the flu data each month. Deploy the model to a Vertex AI endpoint with feature drift monitoring, and retrain the model if a monitoring alert is detected.
Xem giải thích

🧩 Phân tích chi tiết câu hỏi trắc nghiệm

✅ Nội dung câu hỏi:
Câu hỏi xoay quanh việc quản lý và tối ưu hóa chính sách huấn luyện lại (retraining policy) cho một mô hình Machine Learning (ML) được xây dựng bằng BigQuery ML để dự đoán số ca nhiễm cúm (flu infections) trong tháng tới tại Canada. Bạn làm việc cho một công ty dược phẩm ở Canada. Dữ liệu thời tiết (weather data) được công bố hàng tuần, còn dữ liệu thống kê nhiễm cúm (flu infection statistics) được công bố hàng tháng. Mục tiêu chính là tối ưu hóa chi phí (minimize cost) khi cấu hình chính sách retraining. Câu hỏi yêu cầu chọn phương án phù hợp nhất để xử lý dữ liệu mới và retrain mô hình một cách hiệu quả, tránh lãng phí tài nguyên tính toán trên Google Cloud.

🛠️ Bối cảnh kỹ thuật:

  • BigQuery ML cho phép xây dựng và huấn luyện mô hình ML trực tiếp trong BigQuery mà không cần di chuyển dữ liệu.
  • Để deploy và monitor mô hình, cần sử dụng Vertex AI (nền tảng ML end-to-end của Google Cloud).
  • Cloud Scheduler dùng để lên lịch tự động, nhưng retraining định kỳ có thể tốn kém nếu không cần thiết.
  • Giải pháp lý tưởng phải cân bằng giữa cập nhật dữ liệu mới (weather hàng tuần, flu hàng tháng) và chỉ retrain khi mô hình thực sự suy giảm (drift), sử dụng feature drift monitoring trên Vertex AI endpoint để giảm chi phí.

📘 Đáp án đúng và lý do lựa chọn:
Đáp án đúng: Download the weather data each week, and download the flu data each month. Deploy the model to a Vertex AI endpoint with feature drift monitoring, and retrain the model if a monitoring alert is detected.

✅ Lý do chọn đáp án này:
Phương án này tối ưu chi phí nhất vì:

  • Cập nhật dữ liệu đúng tần suất: weather hàng tuần (để nắm bắt thay đổi nhanh), flu hàng tháng (tránh tải dữ liệu thừa).
  • Deploy mô hình lên Vertex AI endpoint với feature drift monitoring (giám sát sự thay đổi phân bố đặc trưng dữ liệu mới so với dữ liệu huấn luyện ban đầu). Chỉ retrain khi phát hiện alert (cảnh báo drift), tránh retraining định kỳ không cần thiết – điều này giảm đáng kể chi phí compute trên Vertex AI (theo giá tính theo giờ sử dụng).
  • Phù hợp với best practice của Google Cloud (cập nhật đến 2026): Vertex AI hỗ trợ Model Monitoring với drift detection tự động, tích hợp BigQuery ML models qua export/import. Không dùng Cloud Scheduler định kỳ, nên tiết kiệm hơn hẳn.

🔍 Giải thích tất cả các phương án (đúng và sai)

  • Phương án 1: Download the weather and flu data each week. Configure Cloud Scheduler to execute a Vertex AI pipeline to retrain the model weekly.
    ❌ Sai vì: Tải flu data hàng tuần là lãng phí (flu chỉ cập nhật hàng tháng, dữ liệu còn lại không thay đổi). Retrain hàng tuần qua Cloud Scheduler + Vertex AI pipeline rất tốn kém (chi phí compute lặp lại cao, dù model chỉ cần cập nhật monthly). Không minimize cost, vi phạm yêu cầu chính.

  • Phương án 2: Download the weather and flu data each month. Configure Cloud Scheduler to execute a Vertex AI pipeline to retrain the model monthly.
    ❌ Sai vì: Tải weather data chỉ hàng tháng bỏ lỡ cập nhật hàng tuần (thời tiết thay đổi nhanh, ảnh hưởng lớn đến dự đoán flu). Retrain định kỳ monthly qua Scheduler vẫn tốn kém hơn monitoring (không linh hoạt, retrain dù model chưa drift). Không tận dụng dữ liệu tối ưu.

  • Phương án 3: Download the weather and flu data each week. Configure Cloud Scheduler to execute a Vertex AI pipeline to retrain the model every month.
    ❌ Sai vì: Tải flu data hàng tuần thừa thãi (dữ liệu không thay đổi). Dù retrain chỉ monthly, vẫn dùng Cloud Scheduler định kỳ → chi phí cố định cao, không "minimize" thực sự (retrain dù không cần). Thiếu cơ chế thông minh như drift monitoring.

  • Phương án 4 (Đúng): Download the weather data each week, and download the flu data each month. Deploy the model to a Vertex AI endpoint with feature drift monitoring, and retrain the model if a monitoring alert is detected.
    ✅ Đúng vì: Như giải thích ở trên – cập nhật dữ liệu chính xác tần suất, dùng drift monitoring để retrain on-demand (chỉ khi cần), giảm chi phí lên đến 70-90% so với scheduled retraining (dựa trên case studies Google Cloud).

📚 Tài liệu tham khảo (cập nhật đến 2026)

  • Vertex AI Model Monitoring: Vertex AI Documentation - Model Monitoring (hỗ trợ feature drift từ 2021, cải tiến alerting 2024-2026 với AutoML integration).
  • BigQuery ML Deployment to Vertex AI: BigQuery ML Guide & Vertex AI Pipelines.
  • Best Practices MLOps: Google Cloud Well-Architected Framework for ML (2025 update), nhấn mạnh monitoring over scheduled retraining để minimize cost.
  • Pricing Insight: Vertex AI endpoint monitoring rẻ hơn pipeline chạy định kỳ (xem Vertex AI Pricing).

Hy vọng phân tích này giúp bạn nắm vững! 🚀

Câu 219
You are building a MLOps platform to automate your company’s ML experiments and model retraining. You need to organize the artifacts for dozens of pipelines. How should you store the pipelines’ artifacts?
  1. A Store parameters in Cloud SQL, and store the models’ source code and binaries in GitHub.
  2. B Store parameters in Cloud SQL, store the models’ source code in GitHub, and store the models’ binaries in Cloud Storage.
  3. C Store parameters in Vertex ML Metadata, store the models’ source code in GitHub, and store the models’ binaries in Cloud Storage.
  4. D Store parameters in Vertex ML Metadata and store the models’ source code and binaries in GitHub.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi tập trung vào việc xây dựng nền tảng MLOps (Machine Learning Operations) trên Google Cloud Platform (GCP) để tự động hóa các thí nghiệm ML (ML experiments) và huấn luyện lại mô hình (model retraining). Bạn cần tổ chức artifacts (các sản phẩm/phần tử) cho hàng chục pipelines (dòng chảy công việc ML). Cụ thể, artifacts bao gồm:

  • Parameters (tham số huấn luyện, cấu hình).
  • Models’ source code (mã nguồn mô hình).
  • Models’ binaries (file mô hình đã huấn luyện, như pickle, SavedModel).

Mục tiêu là chọn cách lưu trữ tối ưu, đảm bảo traceability (theo dõi nguồn gốc), versioning (quản lý phiên bản), và scalability (mở rộng) cho nhiều pipelines, phù hợp với best practices MLOps trên GCP (Vertex AI Pipelines).

📘 Tài liệu tham khảo:

✅ Đáp án đúng

Store parameters in Vertex ML Metadata, store the models’ source code in GitHub, and store the models’ binaries in Cloud Storage.

Lý do lựa chọn 🏆:
Đây là cách tối ưu nhất theo best practices MLOps trên GCP.

  • Vertex ML Metadata (nay là Vertex AI Metadata Store) chuyên lưu parameters, metrics, và lineage (nguồn gốc dữ liệu/mô hình) cho pipelines, hỗ trợ query và visualize tự động qua Vertex AI Experiments. Nó tích hợp liền mạch với Vertex AI Pipelines, giúp theo dõi hàng chục pipelines mà không cần database thủ công.
  • GitHub lý tưởng cho source code (version control với Git).
  • Cloud Storage phù hợp lưu binaries lớn (model artifacts), hỗ trợ versioning qua object naming và tích hợp với Vertex AI Model Registry. Kết hợp này đảm bảo end-to-end traceability, scalable cho production MLOps. ✅

🔍 Giải thích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc:

  • Store parameters in Cloud SQL, and store the models’ source code and binaries in GitHub.
    ❌ Sai. Cloud SQL (database quan hệ) không phù hợp lưu parameters động của ML pipelines vì thiếu tích hợp metadata/ML-specific (như lineage tracking). GitHub chỉ tốt cho source code, không phải binaries lớn (model files có thể hàng GB, gây chậm và không scalable). Thiếu Cloud Storage cho artifacts lớn. 🛠️ Không theo best practices Vertex AI.

  • Store parameters in Cloud SQL, store the models’ source code in GitHub, and store the models’ binaries in Cloud Storage.
    ❌ Sai. Phân tách source code (GitHub) và binaries (Cloud Storage) là tốt, nhưng Cloud SQL vẫn kém cho parameters ML – không hỗ trợ auto-logging từ pipelines, khó query lineage/experiments. Vertex AI Metadata mới là lựa chọn chuẩn. 📊

  • Store parameters in Vertex ML Metadata, store the models’ source code in GitHub, and store the models’ binaries in Cloud Storage.
    ✅ Đúng (như đã giải thích ở trên). Hoàn hảo cho MLOps: metadata tracking + source control + scalable storage. Vertex AI Metadata (cập nhật 2024+) hỗ trợ rich queries và integration với Kubeflow/Cloud Composer. 🚀

  • Store parameters in Vertex ML Metadata and store the models’ source code and binaries in GitHub.
    ❌ Sai. Vertex ML Metadata đúng cho parameters, GitHub tốt cho source code, nhưng binaries không nên lưu ở GitHub (repo bị phình to, giới hạn kích thước ~1GB/repo, không tối ưu versioning cho files lớn). Nên dùng Cloud Storage hoặc Artifact Registry. 💾

Câu 220
You work for a telecommunications company. You’re building a model to predict which customers may fail to pay their next phone bill. The purpose of this model is to proactively offer at-risk customers assistance such as service discounts and bill deadline extensions. The data is stored in BigQuery and the predictive features that are available for model training include:

- Customer_id
- Age
- Salary (measured in local currency)
- Sex
- Average bill value (measured in local currency)
- Number of phone calls in the last month (integer)
- Average duration of phone calls (measured in minutes)

You need to investigate and mitigate potential bias against disadvantaged groups, while preserving model accuracy.

What should you do?
  1. A Determine whether there is a meaningful correlation between the sensitive features and the other features. Train a BigQuery ML boosted trees classification model and exclude the sensitive features and any meaningfully correlated features.
  2. B Train a BigQuery ML boosted trees classification model with all features. Use the ML.GLOBAL_EXPLAIN method to calculate the global attribution values for each feature of the model. If the feature importance value for any of the sensitive features exceeds a threshold, discard the model and tram without this feature.
  3. C Train a BigQuery ML boosted trees classification model with all features. Use the ML.EXPLAIN_PREDICT method to calculate the attribution values for each feature for each customer in a test set. If for any individual customer, the importance value for any feature exceeds a predefined threshold, discard the model and train the model again without this feature.
  4. D Define a fairness metric that is represented by accuracy across the sensitive features. Train a BigQuery ML boosted trees classification model with all features. Use the trained model to make predictions on a test set. Join the data back with the sensitive features, and calculate a fairness metric to investigate whether it meets your requirements.
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi này thuộc lĩnh vực Machine Learning Fairness trên Google Cloud BigQuery ML, tập trung vào việc xây dựng mô hình dự đoán khách hàng có nguy cơ không trả hóa đơn điện thoại (churn prediction kiểu payment failure). Mục tiêu là phát hiện và giảm thiểu bias (thiên kiến) chống lại các nhóm yếu thế (disadvantaged groups), đồng thời giữ nguyên độ chính xác của mô hình.

Dữ liệu lưu trữ trong BigQuery, với các features bao gồm:

  • Customer_id (ID khách hàng)
  • Age (tuổi)
  • Salary (lương, đơn vị tiền tệ địa phương)
  • Sex (giới tính)
  • Average bill value (giá trị hóa đơn trung bình)
  • Number of phone calls (số cuộc gọi tháng qua)
  • Average duration of phone calls (thời lượng cuộc gọi trung bình)

Sensitive features tiềm năng gây bias: Sex (giới tính), Age (tuổi), Salary (có thể đại diện cho nhóm thu nhập thấp).
Vấn đề cốt lõi: Làm thế nào để kiểm tra bias (ví dụ: mô hình kém chính xác hơn với nhóm nữ/giới tính thiểu số, người trẻ/cao tuổi, hoặc thu nhập thấp) mà không làm giảm accuracy tổng thể. Đây là best practice theo Google Cloud ML Fairness guidelines (cập nhật đến 2026, tích hợp trong Vertex AI và BigQuery ML với các metric như disparate impact, equalized odds).

📘 Tài liệu tham khảo:

  • BigQuery ML Documentation: Fairness in BigQuery ML (phiên bản 2026 hỗ trợ ML.EVALUATE với fairness metrics).
  • Google ML Crash Course: Fairness Module (cập nhật 2025-2026 với BigQuery integration).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Define a fairness metric that is represented by accuracy across the sensitive features. Train a BigQuery ML boosted trees classification model with all features. Use the trained model to make predictions on a test set. Join the data back with the sensitive features, and calculate a fairness metric to investigate whether it meets your requirements.

Lý do chọn đáp án này 🛠️:

  • Đây là cách chuẩn và scalable nhất trong BigQuery ML để đo fairness constraints. Định nghĩa metric như accuracy parity (độ chính xác ngang nhau giữa các nhóm sensitive: ví dụ accuracy nam = accuracy nữ ± threshold).
  • Train model với tất cả features (bao gồm sensitive) để giữ accuracy cao, sau đó predict trên test set, join lại sensitive features (qua Customer_id), và tính metric (sử dụng SQL trong BigQuery như ML.EVALUATE hoặc custom query với GROUP BY Sex/Age).
  • Nếu metric không đạt (ví dụ accuracy nữ < 80% accuracy nam), mới mitigate (retrain với regularization hoặc preprocessing). Không loại feature ngay, tránh underfitting.
  • Phù hợp best practice 2026: BigQuery ML hỗ trợ post-training fairness evaluation qua ML.GLOBAL_EXPLAIN kết hợp custom metrics, đảm bảo preserving model accuracy như yêu cầu.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ Phương án SAI: Determine whether there is a meaningful correlation between the sensitive features and the other features. Train a BigQuery ML boosted trees classification model and exclude the sensitive features and any meaningfully correlated features.
    Giải thích sai: Loại bỏ sensitive features (Sex, Age) và các features correlated (ví dụ Salary correlate với Age) sẽ giảm accuracy mạnh vì mất thông tin dự đoán quan trọng (tuổi/lương ảnh hưởng churn). Correlation ≠ bias; bias là disparity in performance giữa groups, không phải correlation. Cách này là preprocessing thô, vi phạm nguyên tắc giữ accuracy.

  • ❌ Phương án SAI: Train a BigQuery ML boosted trees classification model with all features. Use the ML.GLOBAL_EXPLAIN method to calculate the global attribution values for each feature of the model. If the feature importance value for any of the sensitive features exceeds a threshold, discard the model and tram without this feature.
    Giải thích sai: ML.GLOBAL_EXPLAIN chỉ cho global feature importance (tầm quan trọng trung bình), không đo group-wise bias (ví dụ model dùng Sex quan trọng nhưng fair nếu accuracy cân bằng). Threshold arbitrary dẫn đến over-removal features, giảm accuracy. "Tram" là lỗi đánh máy của "train". Không phải cách mitigate bias chuẩn.

  • ❌ Phương án SAI: Train a BigQuery ML boosted trees classification model with all features. Use the ML.EXPLAIN_PREDICT method to calculate the attribution values for each feature for each customer in a test set. If for any individual customer, the importance value for any feature exceeds a predefined threshold, discard the model and train the model again without this feature.
    Giải thích sai: ML.EXPLAIN_PREDICT là local explain (per-instance), kiểm tra mỗi customer quá strict và không scalable (hàng triệu records → discard model chỉ vì 1 cá nhân). Không đo systematic bias ở group level (như disadvantaged groups). Threshold per-instance impractical, dễ overfitting hoặc discard model vô ích.

  • ✅ Phương án ĐÚNG (như đã giải thích ở trên): Define a fairness metric...
    Giải thích đúng: Cách holistic và thực tế, tập trung group-level fairness (accuracy/precision theo sensitive slices), dễ implement bằng SQL join trong BigQuery. Hỗ trợ iterate (tính metric → adjust nếu cần) mà giữ accuracy.

🧠 Kết luận: Best practice là measure first, mitigate if needed với fairness metrics cụ thể, thay vì loại features mù quáng. Áp dụng ngay trong BigQuery ML để production-ready! 🚀