Ngân hàng đề — Google Cloud Professional Machine Learning Engineer

Tìm thấy 333 câu.

Câu 181
You have developed a BigQuery ML model that predicts customer chum, and deployed the model to Vertex AI Endpoints. You want to automate the retraining of your model by using minimal additional code when model feature values change. You also want to minimize the number of times that your model is retrained to reduce training costs. What should you do?
  1. A 1 Enable request-response logging on Vertex AI Endpoints
    2. Schedule a TensorFlow Data Validation job to monitor prediction drift
    3. Execute model retraining if there is significant distance between the distributions
  2. B 1. Enable request-response logging on Vertex AI Endpoints
    2. Schedule a TensorFlow Data Validation job to monitor training/serving skew
    3. Execute model retraining if there is significant distance between the distributions
  3. C 1. Create a Vertex AI Model Monitoring job configured to monitor prediction drift
    2. Configure alert monitoring to publish a message to a Pub/Sub queue when a monitoring alert is detected
    3. Use a Cloud Function to monitor the Pub/Sub queue, and trigger retraining in BigQuery
  4. D 1. Create a Vertex AI Model Monitoring job configured to monitor training/serving skew
    2. Configure alert monitoring to publish a message to a Pub/Sub queue when a monitoring alert is detected
    3. Use a Cloud Function to monitor the Pub/Sub queue, and trigger retraining in BigQuery
Xem giải thích

🧩 Giải thích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc tự động hóa việc huấn luyện lại (retraining) mô hình BigQuery ML dự đoán churn khách hàng, đã được triển khai trên Vertex AI Endpoints. Mục tiêu chính là:

  • Sử dụng ít code bổ sung nhất (minimal additional code).
  • Giảm số lần retraining để tiết kiệm chi phí (minimize training costs).
  • Kích hoạt retraining khi giá trị đặc trưng (feature values) thay đổi (ví dụ: phân phối dữ liệu input thay đổi theo thời gian).

Vấn đề cốt lõi là giám sát sự thay đổi dữ liệu (data drift) trên endpoint production, phát hiện bất thường, và tự động trigger retraining mà không cần can thiệp thủ công nhiều. Vertex AI cung cấp các công cụ tích hợp như Model Monitoring để làm điều này một cách tự động và hiệu quả. 📈

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng là lựa chọn thứ 3:

  1. Create a Vertex AI Model Monitoring job configured to monitor prediction drift
  2. Configure alert monitoring to publish a message to a Pub/Sub queue when a monitoring alert is detected
  3. Use a Cloud Function to monitor the Pub/Sub queue, and trigger retraining in BigQuery

Lý do chọn đáp án này 🏆:

  • Vertex AI Model Monitoring (tính năng mới nhất đến 2026) hỗ trợ giám sát prediction drift – chính xác phát hiện sự thay đổi phân phối đặc trưng input (feature values) so với dữ liệu lúc train, phù hợp hoàn hảo với yêu cầu "model feature values change". ✅
  • Tích hợp alert qua Pub/Sub và Cloud Functions để tự động trigger retraining BigQuery ML với minimal code (chỉ cần config job và function đơn giản).
  • Giảm retraining bằng cách chỉ trigger khi drift significant (threshold configurable), tiết kiệm chi phí. 🛡️
  • Đây là best practice của Google Cloud, không cần code phức tạp như TFDV thủ công.

📋 Giải thích tất cả các phương án (đúng/sai)

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Tôi đánh dấu ✅ đúng hoặc ❌ sai, kèm lý do cụ thể dựa trên tài liệu Vertex AI mới nhất (2024-2026). 🧐

  • Lựa chọn 1 (❌ SAI):

    1 Enable request-response logging on Vertex AI Endpoints
    2. Schedule a TensorFlow Data Validation job to monitor prediction drift
    3. Execute model retraining if there is significant distance between the distributions

    Phân tích: Request-response logging chỉ ghi log input/output, không tự động monitor drift. TensorFlow Data Validation (TFDV) yêu cầu schedule thủ công và code nhiều để so sánh phân phối – không minimal code. Không tích hợp sẵn với Vertex AI endpoints cho automation, dễ dẫn đến retraining thừa. 🚫

  • Lựa chọn 2 (❌ SAI):

    1. Enable request-response logging on Vertex AI Endpoints
    2. Schedule a TensorFlow Data Validation job to monitor training/serving skew
    3. Execute model retraining if there is significant distance between the distributions

    Phân tích: Tương tự lựa chọn 1, logging + TFDV không phải giải pháp tự động. Training/serving skew là khác biệt giữa training data và serving data, nhưng câu hỏi nhấn "feature values change" (prediction drift ở input mới), không khớp chính xác. Vẫn cần code nhiều, không tối ưu chi phí. ❌

  • Lựa chọn 3 (✅ ĐÚNG):

    1. Create a Vertex AI Model Monitoring job configured to monitor prediction drift
    2. Configure alert monitoring to publish a message to a Pub/Sub queue when a monitoring alert is detected
    3. Use a Cloud Function to trigger retraining in BigQuery

    Phân tích: Hoàn hảo! Model Monitoring tự động thu thập dữ liệu từ endpoints, detect prediction drift (thay đổi feature distributions), alert qua Pub/Sub, và Cloud Function trigger BigQuery retraining chỉ khi cần. Minimal code (giao diện console/UI), giảm retraining thừa nhờ threshold. Best practice! 🎯

  • Lựa chọn 4 (❌ SAI):

    1. Create a Vertex AI Model Monitoring job configured to monitor training/serving skew
    2. Configure alert monitoring to publish a message to a Pub/Sub queue when a monitoring alert is detected
    3. Use a Cloud Function to monitor the Pub/Sub queue, and trigger retraining in BigQuery

    Phân tích: Gần đúng nhưng sai ở training/serving skew – skew so sánh training vs serving data, trong khi câu hỏi là "feature values change" (drift ở input mới so với baseline lúc train → prediction drift). Skew phù hợp hơn nếu nghi ngờ data pipeline issue, không chính xác 100%. Vẫn tốt về automation nhưng không khớp yêu cầu. 🔄

📘 Tài liệu tham khảo (cập nhật mới nhất 2026)

Hy vọng phân tích này giúp bạn ôn thi hiệu quả! Nếu cần ví dụ code Cloud Function, hỏi thêm nhé. 🚀

Câu 182
You have been tasked with deploying prototype code to production. The feature engineering code is in PySpark and runs on Dataproc Serverless. The model training is executed by using a Vertex AI custom training job. The two steps are not connected, and the model training must currently be run manually after the feature engineering step finishes. You need to create a scalable and maintainable production process that runs end-to-end and tracks the connections between steps. What should you do?
  1. A Create a Vertex AI Workbench notebook. Use the notebook to submit the Dataproc Serverless feature engineering job. Use the same notebook to submit the custom model training job. Run the notebook cells sequentially to tie the steps together end-to-end.
  2. B Create a Vertex AI Workbench notebook. Initiate an Apache Spark context in the notebook and run the PySpark feature engineering code. Use the same notebook to run the custom model training job in TensorFlow. Run the notebook cells sequentially to tie the steps together end-to-end.
  3. C Use the Kubeflow pipelines SDK to write code that specifies two components:
    - The first is a Dataproc Serverless component that launches the feature engineering job
    - The second is a custom component wrapped in the create_custom_training_job_from_component utility that launches the custom model training job
    Create a Vertex AI Pipelines job to link and run both components
  4. D Use the Kubeflow pipelines SDK to write code that specifies two components
    - The first component initiates an Apache Spark context that runs the PySpark feature engineering code
    - The second component runs the TensorFlow custom model training code
    Create a Vertex AI Pipelines job to link and run both components.
Xem giải thích

🧠 Phân tích câu hỏi trắc nghiệm về quy trình triển khai ML production trên Google Cloud

✅ Giải thích nội dung câu hỏi một cách chi tiết và rõ ràng:
Câu hỏi mô tả tình huống bạn cần triển khai mã prototype vào production một cách scalable (có khả năng mở rộng) và maintainable (dễ bảo trì). Cụ thể:

  • Bước 1 (Feature Engineering): Sử dụng mã PySpark chạy trên Dataproc Serverless (dịch vụ serverless cho Spark trên GCP, không cần quản lý cluster).
  • Bước 2 (Model Training): Chạy qua Vertex AI custom training job (công việc huấn luyện mô hình tùy chỉnh trên Vertex AI).
    Hiện tại, hai bước không kết nối tự động, phải chạy manual bước training sau khi feature engineering hoàn tất.
    Yêu cầu: Tạo quy trình end-to-end (từ đầu đến cuối tự động), track connections between steps (theo dõi liên kết giữa các bước) để đảm bảo production-ready.
    🛠️ Mục tiêu chính: Sử dụng orchestration tool để liên kết, chạy tự động, theo dõi lineage (dòng dõi dữ liệu/mô hình), và scale tốt (không dùng notebook vì notebook không phù hợp production).
    📘 Kiến thức cập nhật (Vertex AI Pipelines v2+, Kubeflow Pipelines SDK 2.x đến 2026): Vertex AI Pipelines là giải pháp chính thức cho ML workflows trên GCP, hỗ trợ components sẵn có như Dataproc Serverless và custom training jobs.

✅ Đáp án đúng và lý do lựa chọn:
Đáp án đúng là lựa chọn thứ 3:
Use the Kubeflow pipelines SDK to write code that specifies two components:
- The first is a Dataproc Serverless component that launches the feature engineering job
- The second is a custom component wrapped in the create_custom_training_job_from_component utility that launches the custom model training job
Create a Vertex AI Pipelines job to link and run both components

Lý do chọn (🧩 Phân tích chi tiết):

  • Sử dụng Kubeflow Pipelines SDK (tích hợp sâu với Vertex AI Pipelines) để định nghĩa hai components chính xác:
    • Component 1: Dataproc Serverless component (component sẵn có trong KFP SDK, launch job PySpark serverless – khớp yêu cầu gốc).
    • Component 2: Custom component bọc bằng create_custom_training_job_from_component (utility chính thức của Vertex AI để wrap custom training job).
  • Tạo Vertex AI Pipelines job để link và run end-to-end, tự động hóa, track metadata/lineage (qua Pipeline Runs UI), scalable (chạy parallel/scale theo nhu cầu), maintainable (code-based, version control).
  • ✅ Hoàn hảo cho production: Không manual, hỗ trợ retry, monitoring, và tích hợp Vertex AI Metadata Store (cập nhật 2024-2026).

❌ Giải thích tất cả các phương án (đúng/sai):
🧩 Lựa chọn 1 (SAI):
Create a Vertex AI Workbench notebook. Use the notebook to submit the Dataproc Serverless feature engineering job. Use the same notebook to submit the custom model training job. Run the notebook cells sequentially to tie the steps together end-to-end.
Lý do sai: Vertex AI Workbench (Jupyter notebook) chỉ phù hợp prototyping, không scalable/maintainable cho production (không auto-scale, khó schedule, không track lineage tốt, phải manual run cells). Không phải end-to-end production process thực thụ. ❌

🧩 Lựa chọn 2 (SAI):
Create a Vertex AI Workbench notebook. Initiate an Apache Spark context in the notebook and run the PySpark feature engineering code. Use the same notebook to run the custom model training job in TensorFlow. Run the notebook cells sequentially to tie the steps together end-to-end.
Lý do sai: Tương tự lựa chọn 1, notebook không production-ready. Hơn nữa, khởi tạo Spark context trực tiếp trong notebook thay vì dùng Dataproc Serverless (yêu cầu gốc), gây tốn tài nguyên và không serverless. Giả định training dùng TensorFlow (không chỉ định). ❌

🧩 Lựa chọn 3 (ĐÚNG):
Use the Kubeflow pipelines SDK to write code that specifies two components:
- The first is a Dataproc Serverless component that launches the feature engineering job
- The second is a custom component wrapped in the create_custom_training_job_from_component utility that launches the custom model training job
Create a Vertex AI Pipelines job to link and run both components

Lý do đúng: Như phân tích trên, khớp 100% yêu cầu: Dataproc Serverless component chính xác, custom training wrapper chuẩn, Vertex AI Pipelines orchestrate end-to-end với tracking. ✅ Best practice 2026!

🧩 Lựa chọn 4 (SAI):
Use the Kubeflow pipelines SDK to write code that specifies two components
- The first component initiates an Apache Spark context that runs the PySpark feature engineering code
- The second component runs the TensorFlow custom model training code
Create a Vertex AI Pipelines job to link and run both components.

Lý do sai: Dù dùng Kubeflow/Vertex AI Pipelines (tốt), nhưng component 1 khởi tạo Spark context trực tiếp (không dùng Dataproc Serverless – vi phạm yêu cầu gốc). Component 2 giả định TensorFlow (không chỉ định), không dùng utility wrap custom job chuẩn. ❌

📘 Tài liệu tham khảo (cập nhật mới nhất 2026):

Câu 183
You recently deployed a scikit-learn model to a Vertex AI endpoint. You are now testing the model on live production traffic. While monitoring the endpoint, you discover twice as many requests per hour than expected throughout the day. You want the endpoint to efficiently scale when the demand increases in the future to prevent users from experiencing high latency. What should you do?
  1. A Deploy two models to the same endpoint, and distribute requests among them evenly
  2. B Configure an appropriate minReplicaCount value based on expected baseline traffic
  3. C Set the target utilization percentage in the autoscailngMetricSpecs configuration to a higher value
  4. D Change the model’s machine type to one that utilizes GPUs
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi mô tả tình huống thực tế khi triển khai một mô hình scikit-learn lên Vertex AI endpoint (dịch vụ triển khai mô hình ML trên Google Cloud). Bạn đang kiểm tra mô hình với lưu lượng truy cập sản xuất thực tế (live production traffic), và phát hiện lưu lượng yêu cầu gấp đôi so với dự kiến (twice as many requests per hour) suốt cả ngày. Vấn đề chính là cần tối ưu hóa khả năng scale của endpoint để xử lý nhu cầu tăng đột biến trong tương lai, tránh tình trạng độ trễ cao (high latency) cho người dùng.

Mục tiêu: Đảm bảo endpoint tự động mở rộng (autoscaling) hiệu quả, dựa trên cơ chế scaling của Vertex AI như replicas (số lượng instance song song), minReplicaCount (số replica tối thiểu), và các metric như CPU utilization. Đây là vấn đề phổ biến trong production ML, nơi baseline traffic cao hơn dự đoán có thể gây cold start (khởi động chậm) nếu không cấu hình đúng.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Configure an appropriate minReplicaCount value based on expected baseline traffic

Lý do:
Trong Vertex AI, minReplicaCount quy định số lượng replica (instance) tối thiểu luôn chạy sẵn để xử lý lưu lượng cơ bản (baseline traffic). Với lưu lượng thực tế gấp đôi dự kiến, việc đặt minReplicaCount phù hợp (ví dụ: tăng từ 1 lên 2-4 dựa trên peak traffic) sẽ đảm bảo endpoint có đủ tài nguyên sẵn sàng ngay lập tức, tránh tình trạng scale từ 0 (cold start gây latency cao). Autoscaling sẽ tự động tăng replica khi traffic vượt ngưỡng, nhưng minReplicaCount giúp phản ứng nhanh chóng với demand tăng. Đây là best practice theo docs Vertex AI mới nhất (cập nhật 2025-2026), giúp tối ưu chi phí và performance.

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Deploy two models to the same endpoint, and distribute requests among them evenly
    Phương án này sai vì Vertex AI endpoint chỉ hỗ trợ triển khai một model version duy nhất tại một thời điểm (không phải multiple models cùng endpoint). Việc "deploy two models" và phân phối request sẽ yêu cầu multi-model endpoints hoặc separate endpoints, phức tạp hóa quản lý và không giải quyết scaling traffic. Thay vào đó, dùng traffic splitting giữa các endpoint versions, nhưng không phải cách chính để scale.

  • ✅ [ĐÚNG] Configure an appropriate minReplicaCount value based on expected baseline traffic
    Như đã giải thích ở trên: Đây là giải pháp trực tiếp, hiệu quả nhất. Vertex AI autoscaling dựa trên minReplicaCount, maxReplicaCount, và metrics như CPU. Đặt minReplicaCount cao hơn baseline (dựa trên monitoring từ Cloud Monitoring) đảm bảo zero cold start, scale mượt mà. Ví dụ: Nếu baseline là 100 req/hour, set min=2 để handle gấp đôi.

  • ❌ [SAI] Set the target utilization percentage in the autoscailngMetricSpecs configuration to a higher value
    Sai vì target utilization (thường 60-80% CPU) là ngưỡng kích hoạt autoscaling. Tăng giá trị này (ví dụ từ 60% lên 90%) sẽ làm endpoint scale muộn hơn, dẫn đến overload và latency cao hơn chính xác vấn đề cần tránh. Nên giữ hoặc giảm target để scale sớm hơn, kết hợp minReplicaCount.

  • ❌ [SAI] Change the model’s machine type to one that utilizes GPUs
    Sai vì mô hình scikit-learn chủ yếu dùng CPU (không tận dụng GPU hiệu quả cho inference nhẹ). Chuyển sang GPU machine type (như n1-standard với GPU) sẽ tăng chi phí vô ích (GPU đắt hơn 5-10x), không giải quyết scaling traffic mà chỉ tăng throughput/instance. Scale đúng cách là tăng replicas, không phải hardware mạnh hơn.

📘 Tài liệu tham khảo (cập nhật mới nhất đến 2026)

🛠️ Khuyến nghị thêm: Sử dụng Cloud Monitoring để phân tích traffic patterns (QPS, latency percentiles), và test với Vertex AI Batch Prediction trước production. Nếu traffic biến động cao, kết hợp maxReplicaCount và concurrency settings!

Câu 184
You work at a bank. You have a custom tabular ML model that was provided by the bank’s vendor. The training data is not available due to its sensitivity. The model is packaged as a Vertex AI Model serving container, which accepts a string as input for each prediction instance. In each string, the feature values are separated by commas. You want to deploy this model to production for online predictions and monitor the feature distribution over time with minimal effort. What should you do?
  1. A 1. Upload the model to Vertex AI Model Registry, and deploy the model to a Vertex AI endpoint
    2. Create a Vertex AI Model Monitoring job with feature drift detection as the monitoring objective, and provide an instance schema
  2. B 1. Upload the model to Vertex AI Model Registry, and deploy the model to a Vertex AI endpoint
    2. Create a Vertex AI Model Monitoring job with feature skew detection as the monitoring objective, and provide an instance schema
  3. C 1. Refactor the serving container to accept key-value pairs as input format
    2. Upload the model to Vertex AI Model Registry, and deploy the model to a Vertex AI endpoint
    3. Create a Vertex AI Model Monitoring job with feature drift detection as the monitoring objective.
  4. D 1. Refactor the serving container to accept key-value pairs as input format
    2. Upload the model to Vertex AI Model Registry, and deploy the model to a Vertex AI endpoint
    3. Create a Vertex AI Model Monitoring job with feature skew detection as the monitoring objective
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả tình huống bạn làm việc tại một ngân hàng, có một mô hình ML tùy chỉnh dạng bảng (tabular ML model) do nhà cung cấp cung cấp. Dữ liệu huấn luyện không khả dụng vì lý do bảo mật nhạy cảm. Mô hình được đóng gói sẵn dưới dạng Vertex AI Model serving container, nhận đầu vào là chuỗi string với các giá trị đặc trưng (feature) phân cách bằng dấu phẩy (comma-separated, giống định dạng CSV).
Mục tiêu: Triển khai mô hình lên production để thực hiện online predictions (dự đoán thời gian thực), đồng thời giám sát phân bố đặc trưng (feature distribution) theo thời gian với nỗ lực tối thiểu (minimal effort).
🔑 Thách thức chính: Không có dữ liệu huấn luyện → không thể so sánh với baseline huấn luyện; đầu vào là string đơn giản → cần công cụ hỗ trợ giám sát mà không cần thay đổi container; ưu tiên giải pháp dễ dàng nhất trên Vertex AI (Google Cloud).

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng là phương án đầu tiên:

  1. Upload the model to Vertex AI Model Registry, and deploy the model to a Vertex AI endpoint
  2. Create a Vertex AI Model Monitoring job with feature drift detection as the monitoring objective, and provide an instance schema

Lý do chọn đáp án này 🛠️:

  • Bước 1: Upload mô hình vào Vertex AI Model Registry và deploy lên Vertex AI endpoint là quy trình chuẩn, đơn giản nhất để đưa mô hình custom container vào production cho online predictions (không cần refactor).
  • Bước 2: Tạo Vertex AI Model Monitoring job với feature drift detection để giám sát sự thay đổi phân bố đặc trưng giữa các khoảng thời gian serving data (không cần dữ liệu huấn luyện). Cung cấp instance schema (JSON schema mô tả cấu trúc string input, ví dụ: định nghĩa các field comma-separated) giúp Vertex AI parse dữ liệu unstructured một cách tự động, hỗ trợ tabular data dạng string mà không cần thay đổi container.
  • Minimal effort: Không refactor, tận dụng tính năng mới nhất của Vertex AI (từ 2023-2026), phù hợp khi training data unavailable.
    📘 Nguồn tham khảo:
  • Vertex AI Model Monitoring Overview (cập nhật 2025: hỗ trợ drift cho serving data over time).
  • Monitoring Unstructured Data (instance schema cho string/CSV input).

📋 Giải thích chi tiết tất cả các phương án

Dưới đây là phân tích từng phương án, giữ nguyên nội dung văn bản gốc bằng tiếng Anh. Mỗi phương án được đánh giá đúng/sai kèm lý do cụ thể dựa trên kiến thức Vertex AI mới nhất (2026).

  • Phương án ĐÚNG ✅:

    1. Upload the model to Vertex AI Model Registry, and deploy the model to a Vertex AI endpoint
    2. Create a Vertex AI Model Monitoring job with feature drift detection as the monitoring objective, and provide an instance schema
      Lý do đúng: Như giải thích ở trên. Feature drift lý tưởng để monitor feature distribution over time (so sánh serving data giữa các window thời gian), instance schema giải quyết input string mà không cần training data hay refactor. Hoàn hảo cho minimal effort! 🏆
  • Phương án SAI ❌ (Phương án 2):

    1. Upload the model to Vertex AI Model Registry, and deploy the model to a Vertex AI endpoint
    2. Create a Vertex AI Model Monitoring job with feature skew detection as the monitoring objective, and provide an instance schema
      Lý do sai: Feature skew yêu cầu baseline từ training data để so sánh với serving data, nhưng training data không available (sensitive). Dù có instance schema, vẫn không thể chạy monitoring skew → thất bại mục tiêu giám sát. (Vertex AI docs: skew cần training statistics).
  • Phương án SAI ❌ (Phương án 3):

    1. Refactor the serving container to accept key-value pairs as input format
    2. Upload the model to Vertex AI Model Registry, and deploy the model to a Vertex AI endpoint
    3. Create a Vertex AI Model Monitoring job with feature drift detection as the monitoring objective.
      Lý do sai: Refactor container để nhận key-value pairs là nỗ lực lớn (vi phạm minimal effort), không cần thiết vì Vertex AI hỗ trợ string input qua instance schema. Bỏ qua schema → monitoring drift không parse được dữ liệu tabular string. Phức tạp hóa vấn đề!
  • Phương án SAI ❌ (Phương án 4):

    1. Refactor the serving container to accept key-value pairs as input format
    2. Upload the model to Vertex AI Model Registry, and deploy the model to a Vertex AI endpoint
    3. Create a Vertex AI Model Monitoring job with feature skew detection as the monitoring objective
      Lý do sai: Kết hợp 2 lỗi: Refactor không cần thiết (minimal effort bị phá vỡ) + feature skew không dùng được vì thiếu training data. Đầu vào key-value chỉ làm phức tạp thêm, không giải quyết gốc rễ.

Tóm tắt nhanh 🚀: Chọn drift + schema để "plug-and-play" giám sát serving data over time, tận dụng Vertex AI mà không đụng đến code hay data nhạy cảm!

Câu 185
You are implementing a batch inference ML pipeline in Google Cloud. The model was developed using TensorFlow and is stored in SavedModel format in Cloud Storage. You need to apply the model to a historical dataset containing 10 TB of data that is stored in a BigQuery table. How should you perform the inference?
  1. A Export the historical data to Cloud Storage in Avro format. Configure a Vertex AI batch prediction job to generate predictions for the exported data
  2. B Import the TensorFlow model by using the CREATE MODEL statement in BigQuery ML. Apply the historical data to the TensorFlow model
  3. C Export the historical data to Cloud Storage in CSV format. Configure a Vertex AI batch prediction job to generate predictions for the exported data
  4. D Configure a Vertex AI batch prediction job to apply the model to the historical data in BigQuery
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc triển khai batch inference (suy luận hàng loạt) cho một pipeline ML trên Google Cloud. Cụ thể:

  • Model được phát triển bằng TensorFlow, lưu dưới định dạng SavedModel trong Cloud Storage (GCS).
  • Dữ liệu lịch sử có dung lượng lớn 10 TB, lưu trong một bảng BigQuery.
  • Mục tiêu: Áp dụng model lên toàn bộ dữ liệu này một cách hiệu quả, tối ưu chi phí và thời gian (vì dữ liệu khổng lồ, tránh export không cần thiết).

Vấn đề cốt lõi: Với dữ liệu lớn ở BigQuery, cần chọn cách inference trực tiếp trên BigQuery để tránh chi phí export dữ liệu ra GCS (rất tốn kém với 10TB), đồng thời tận dụng tích hợp sẵn của Google Cloud. Kiến thức cập nhật đến 2026: BigQuery ML hỗ trợ import model TensorFlow SavedModel từ GCS và predict trực tiếp trên bảng BigQuery mà không cần di chuyển dữ liệu. Vertex AI batch prediction thì yêu cầu input ở GCS (CSV/JSON Lines), không hỗ trợ trực tiếp BigQuery.

📘 Tài liệu tham khảo:

✅ Đáp án đúng: Import the TensorFlow model by using the CREATE MODEL statement in BigQuery ML. Apply the historical data to the TensorFlow model

Lý do chọn đáp án này:

  • BigQuery ML cho phép import model TensorFlow SavedModel trực tiếp từ GCS bằng lệnh CREATE MODEL ... FROM 'gs://path/to/savedmodel'.
  • Sau đó, sử dụng ML.PREDICT(MODEL my_model, TABLE my_table) để inference trực tiếp trên bảng BigQuery 10TB, không cần export dữ liệu → Tiết kiệm chi phí, thời gian và tài nguyên (BigQuery xử lý song song quy mô lớn).
  • Hoàn hảo cho batch inference lớn, scalable đến petabyte mà không di chuyển data.

🛠️ Giải thích tất cả các phương án (đúng/sai)

  • ✅ [ĐÚNG] Import the TensorFlow model by using the CREATE MODEL statement in BigQuery ML. Apply the historical data to the TensorFlow model
    Như đã giải thích ở trên: Phương án tối ưu nhất cho dữ liệu lớn ở BigQuery. Hỗ trợ đầy đủ TensorFlow SavedModel, predict in-place, chi phí thấp (chỉ tính query BigQuery). ✅ Hoàn toàn phù hợp.

  • ❌ [SAI] Export the historical data to Cloud Storage in Avro format. Configure a Vertex AI batch prediction job to generate predictions for the exported data
    Sai vì: Vertex AI batch prediction không hỗ trợ input Avro (chỉ CSV/JSON Lines cho tabular data). Export 10TB Avro từ BigQuery rất tốn kém (chi phí storage + egress + thời gian dài). Không hiệu quả so với BigQuery ML. ❌ Không khả thi về format và hiệu suất.

  • ❌ [SAI] Export the historical data to Cloud Storage in CSV format. Configure a Vertex AI batch prediction job to generate predictions for the exported data
    Sai vì: Mặc dù Vertex AI hỗ trợ CSV từ GCS, nhưng vẫn phải export 10TB từ BigQuery → Chi phí cao (storage GCS ~$0.023/GB/tháng, export query lớn), thời gian lâu (hàng giờ/ngày). Không tận dụng BigQuery native. ❌ Có thể làm nhưng không phải best practice.

  • ❌ [SAI] Configure a Vertex AI batch prediction job to apply the model to the historical data in BigQuery
    Sai vì: Vertex AI batch prediction KHÔNG hỗ trợ input trực tiếp từ BigQuery (phải là file ở GCS). Tài liệu chính thức xác nhận chỉ GCS/Cloud Storage. ❌ Không được hỗ trợ, sẽ lỗi.

Câu 186
You recently deployed a model to a Vertex AI endpoint. Your data drifts frequently, so you have enabled request-response logging and created a Vertex AI Model Monitoring job. You have observed that your model is receiving higher traffic than expected. You need to reduce the model monitoring cost while continuing to quickly detect drift. What should you do?
  1. A Replace the monitoring job with a DataFlow pipeline that uses TensorFlow Data Validation (TFDV)
  2. B Replace the monitoring job with a custom SQL script to calculate statistics on the features and predictions in BigQuery
  3. C Decrease the sample_rate parameter in the RandomSampleConfig of the monitoring job
  4. D Increase the monitor_interval parameter in the ScheduleConfig of the monitoring job
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào quản lý chi phí giám sát mô hình (Model Monitoring) trên Vertex AI của Google Cloud. Cụ thể:

  • Bạn đã triển khai một mô hình lên Vertex AI endpoint.
  • Dữ liệu thường xuyên bị drift (sự thay đổi phân bố dữ liệu đầu vào hoặc dự đoán), nên bạn đã kích hoạt request-response logging (ghi log yêu cầu và phản hồi) và tạo Vertex AI Model Monitoring job để theo dõi.
  • Vấn đề: Lưu lượng truy cập (traffic) cao hơn mong đợi, dẫn đến chi phí giám sát tăng cao.
  • Mục tiêu: Giảm chi phí giám sát nhưng vẫn phát hiện drift nhanh chóng.

🛠️ Yêu cầu hành động: Cần điều chỉnh cấu hình monitoring job để tối ưu chi phí mà không làm chậm quá trình phát hiện drift (tức là giữ nguyên tần suất kiểm tra, chỉ giảm lượng dữ liệu xử lý).

📘 Tài liệu tham khảo:

✅ Đáp án đúng

Decrease the sample_rate parameter in the RandomSampleConfig of the monitoring job

Lý do lựa chọn:

  • RandomSampleConfig cho phép lấy mẫu ngẫu nhiên từ log request-response để tính toán thống kê (như phân bố features, predictions).
  • Giảm sample_rate (tỷ lệ lấy mẫu, ví dụ từ 1.0 xuống 0.1) sẽ giảm lượng dữ liệu xử lý, từ đó giảm chi phí Compute Engine/BigQuery mà không ảnh hưởng đến tần suất kiểm tra drift (vẫn chạy theo lịch ScheduleConfig hiện tại).
  • Drift vẫn được phát hiện nhanh chóng vì job vẫn chạy định kỳ, chỉ xử lý ít dữ liệu hơn → cân bằng hoàn hảo giữa chi phí và hiệu suất. ✅

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ [SAI] Replace the monitoring job with a DataFlow pipeline that uses TensorFlow Data Validation (TFDV)
    Phương án này sai vì thay thế hoàn toàn Vertex AI Model Monitoring bằng DataFlow pipeline + TFDX sẽ tăng chi phí và phức tạp hóa. DataFlow là dịch vụ batch/stream processing tốn kém (dựa trên Compute Engine), TFDV chỉ validate schema/data drift thủ công, không tích hợp tự động alerting như Vertex AI job. Không giải quyết traffic cao mà còn yêu cầu dev ops thêm → không giảm chi phí hiệu quả.

  • ❌ [SAI] Replace the monitoring job with a custom SQL script to calculate statistics on the features and predictions in BigQuery
    Phương án sai vì dùng SQL tùy chỉnh trên BigQuery đòi hỏi tự build pipeline (query log tables), thiếu tự động hóa alerting/drift detection của Vertex AI. Với traffic cao, query BigQuery sẽ tăng chi phí scan dữ liệu lớn, không scalable và không nhanh bằng sampling tích hợp → vi phạm yêu cầu "giảm chi phí" và "detect nhanh".

  • ✅ [ĐÚNG] Decrease the sample_rate parameter in the RandomSampleConfig of the monitoring job
    (Như giải thích ở trên: Giảm tỷ lệ lấy mẫu → giảm dữ liệu xử lý → tiết kiệm chi phí, giữ nguyên tốc độ detect drift). Đây là cách tối ưu nhất theo best practices Vertex AI. 🏆

  • ❌ [SAI] Increase the monitor_interval parameter in the ScheduleConfig of the monitoring job
    Phương án sai vì tăng monitor_interval (ví dụ từ 1 giờ lên 24 giờ) sẽ làm job chạy ít thường xuyên hơn, dẫn đến phát hiện drift chậm (có thể mất hàng ngày). Mặc dù giảm chi phí, nhưng vi phạm yêu cầu "continuing to quickly detect drift" → không phù hợp.

🧠 Lời khuyên bổ sung: Theo docs 2026, kết hợp với alerting channels (Pub/Sub, email) để notify drift ngay lập tức, và dùng BigQuery views để preview log trước khi scale sampling! 🚀

Câu 187
You work for a retail company. You have created a Vertex AI forecast model that produces monthly item sales predictions. You want to quickly create a report that will help to explain how the model calculates the predictions. You have one month of recent actual sales data that was not included in the training dataset. How should you generate data for your report?
  1. A Create a batch prediction job by using the actual sales data. Compare the predictions to the actuals in the report.
  2. B Create a batch prediction job by using the actual sales data, and configure the job settings to generate feature attributions. Compare the results in the report.
  3. C Generate counterfactual examples by using the actual sales data. Create a batch prediction job using the actual sales data and the counterfactual examples. Compare the results in the report.
  4. D Train another model by using the same training dataset as the original, and exclude some columns. Using the actual sales data create one batch prediction job by using the new model and another one with the original model. Compare the two sets of predictions in the report.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi thuộc lĩnh vực Vertex AI (Google Cloud Platform - GCP), không phải AWS như mô tả ban đầu (có thể là nhầm lẫn). Bạn đang làm việc cho một công ty bán lẻ, đã xây dựng mô hình dự báo (forecast model) trên Vertex AI để dự đoán doanh số bán hàng theo tháng cho từng mặt hàng. Mục tiêu là tạo nhanh một báo cáo giải thích cách mô hình tính toán các dự đoán (explanability của model). Bạn có dữ liệu doanh số thực tế (actual sales data) của 1 tháng gần nhất, chưa được dùng trong tập huấn luyện (held-out data).

Vấn đề cốt lõi: Cần phương pháp nhanh chóng để generate dữ liệu cho báo cáo, tập trung vào cách model tính predictions (không chỉ accuracy), sử dụng dữ liệu thực tế này. Vertex AI hỗ trợ feature attributions (giải thích đóng góp của từng feature vào prediction) cho batch prediction jobs, đặc biệt với forecasting models (cập nhật đến Vertex AI v2026, hỗ trợ XAI cho Time Series Forecasting).

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Create a batch prediction job by using the actual sales data, and configure the job settings to generate feature attributions. Compare the results in the report.

Lý do 🛠️:

  • Đây là cách nhanh nhất và trực tiếp nhất để giải thích cách model tính toán predictions. Vertex AI cho phép enable feature attributions ngay trong batch prediction job trên held-out data (actual sales).
  • Feature attributions (dựa trên Integrated Gradients hoặc Sampled Shapley) sẽ output đóng góp của từng feature (ví dụ: giá cả, mùa vụ, lịch sử bán hàng) vào prediction cụ thể cho từng item/tháng.
  • Sau đó, so sánh results (predictions + attributions vs actuals) trong báo cáo để minh họa rõ ràng, dễ hiểu. Không cần train model mới, tiết kiệm thời gian.
  • Phù hợp với best practice XAI trên Vertex AI cho forecasting (cập nhật 2026: hỗ trợ full cho AutoML Tabular Forecasting).

❌ Giải thích tất cả các phương án (đúng/sai)

  • [SAI] Create a batch prediction job by using the actual sales data. Compare the predictions to the actuals in the report.
    ❌ Sai vì: Chỉ tạo batch prediction và so sánh predictions vs actuals chỉ đánh giá accuracy (như MAE/RMSE), không giải thích cách model tính toán (no explainability). Không generate dữ liệu attributions để báo cáo minh họa đóng góp features. Thiếu yếu tố cốt lõi của yêu cầu "explain how the model calculates".

  • [ĐÚNG] Create a batch prediction job by using the actual sales data, and configure the job settings to generate feature attributions. Compare the results in the report.
    ✅ Đúng vì: Như giải thích ở trên. Configure job settings để enable feature attributions (qua console/API: explainParameters), output attributions matrix cho held-out data. Rất nhanh (batch job chỉ mất phút-giờ), trực tiếp dùng actuals để contextualize explanations trong báo cáo. Hoàn hảo cho forecasting models trên Vertex AI.

  • [SAI] Generate counterfactual examples by using the actual sales data. Create a batch prediction job using the actual sales data and the counterfactual examples. Compare the results in the report.
    ❌ Sai vì: Counterfactuals (what-if scenarios, ví dụ: "nếu giá thay đổi thì sao?") là advanced XAI, nhưng Vertex AI chưa hỗ trợ native cho forecasting models (chỉ experimental ở một số tabular classifiers đến 2026). Phức tạp, tốn thời gian generate counterfactuals riêng, không trực tiếp explain model calculations mà chỉ simulate scenarios. Không "quickly" như yêu cầu.

  • [SAI] Train another model by using the same training dataset as the original, and exclude some columns. Using the actual sales data create one batch prediction job by using the new model and another one with the original model. Compare the two sets of predictions in the report.
    ❌ Sai vì: Đây là permutation feature importance thủ công (train model mới thiếu columns để đo delta performance), rất tốn kém (train lại model đầy đủ), không chính xác bằng native feature attributions, và không explain per-prediction mà chỉ global importance. Không phù hợp "quickly create" – mất hàng giờ/ngày, vi phạm best practice Vertex AI (dùng built-in XAI thay vì hack).

Kết luận 🎯: Phương án đúng tận dụng built-in explainability của Vertex AI, đảm bảo nhanh, chính xác và scalable cho production forecasting! Nếu cần code sample, tôi có thể cung cấp Python SDK ví dụ.

Câu 188
Your team has a model deployed to a Vertex AI endpoint. You have created a Vertex AI pipeline that automates the model training process and is triggered by a Cloud Function. You need to prioritize keeping the model up-to-date, but also minimize retraining costs. How should you configure retraining?
  1. A Configure Pub/Sub to call the Cloud Function when a sufficient amount of new data becomes available
  2. B Configure a Cloud Scheduler job that calls the Cloud Function at a predetermined frequency that fits your team’s budget
  3. C Enable model monitoring on the Vertex AI endpoint. Configure Pub/Sub to call the Cloud Function when anomalies are detected
  4. D Enable model monitoring on the Vertex AI endpoint. Configure Pub/Sub to call the Cloud Function when feature drift is detected
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi tập trung vào việc cấu hình quy trình retraining (huấn luyện lại mô hình) cho một mô hình đã triển khai trên Vertex AI endpoint trong Google Cloud. Đội ngũ đã xây dựng một Vertex AI pipeline tự động hóa quá trình huấn luyện mô hình, và pipeline này được kích hoạt bởi Cloud Function.

Mục tiêu chính:

  • ✅ Ưu tiên giữ mô hình luôn cập nhật (up-to-date) để đảm bảo hiệu suất tốt với dữ liệu mới.
  • 💰 Giảm thiểu chi phí retraining bằng cách tránh huấn luyện không cần thiết.

Vấn đề cần giải quyết: Làm thế nào để kích hoạt retraining một cách thông minh, chỉ khi thực sự cần thiết (dựa trên sự thay đổi dữ liệu hoặc hiệu suất mô hình), thay vì lịch cố định hoặc ngẫu nhiên. Đây là tình huống thực tế trong MLOps trên Vertex AI, nơi model monitoring giúp phát hiện các vấn đề như feature drift (sự thay đổi phân phối đặc trưng giữa dữ liệu huấn luyện và dữ liệu phục vụ).

Bối cảnh kiến thức cập nhật (đến 2026): Theo tài liệu Vertex AI mới nhất (Google Cloud Vertex AI documentation, cập nhật 2024-2026), Model Monitoring hỗ trợ phát hiện feature drift, prediction drift, ground truth drift, và anomalies. Nó tích hợp với Pub/Sub để gửi thông báo, kích hoạt Cloud Function/Pipeline một cách tự động. Điều này phù hợp với best practices trong Vertex AI Pipelines và Automated ML Operations.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Enable model monitoring on the Vertex AI endpoint. Configure Pub/Sub to call the Cloud Function when feature drift is detected.

Lý do:

  • 🛠️ Feature drift là sự thay đổi trong phân phối dữ liệu đầu vào (features) giữa lúc huấn luyện và lúc phục vụ, thường dẫn đến mô hình kém hiệu suất. Phát hiện drift này qua Model Monitoring sẽ kích hoạt retraining chính xác khi cần, giúp mô hình luôn up-to-date mà tiết kiệm chi phí (chỉ retrain khi drift vượt ngưỡng, không lãng phí).
  • Đây là cách tối ưu nhất theo best practices Google Cloud: Tích hợp monitoring với Pub/Sub → Cloud Function → Pipeline, hỗ trợ tự động hóa MLOps end-to-end.
  • Phù hợp mục tiêu: Up-to-date (phát hiện thay đổi dữ liệu kịp thời) + Minimize costs (tránh retrain định kỳ vô ích).

📋 Giải thích tất cả các phương án (đúng/sai)

  • ❌ Phương án SAI: Configure Pub/Sub to call the Cloud Function when a sufficient amount of new data becomes available
    Giải thích: Phương án này dựa vào lượng dữ liệu mới tích lũy (ví dụ: đếm records trong BigQuery), nhưng không đảm bảo dữ liệu mới có ý nghĩa thay đổi (có thể chỉ là dữ liệu tương tự). Dẫn đến retrain không cần thiết, tăng chi phí mà không ưu tiên up-to-date thực sự. Vertex AI không có cơ chế "sufficient data" tự động chuẩn như model monitoring.

  • ❌ Phương án SAI: Configure a Cloud Scheduler job that calls the Cloud Function at a predetermined frequency that fits your team’s budget
    Giải thích: Sử dụng Cloud Scheduler chạy định kỳ (ví dụ: hàng tuần) có thể tiết kiệm chi phí nếu tần suất thấp, nhưng không thông minh: Có thể miss thay đổi dữ liệu đột ngột (model lỗi sớm) hoặc retrain thừa khi dữ liệu ổn định. Không ưu tiên up-to-date, chỉ dựa lịch cố định – kém hiệu quả so với monitoring-based triggers.

  • ❌ Phương án SAI: Enable model monitoring on the Vertex AI endpoint. Configure Pub/Sub to call the Cloud Function when anomalies are detected
    Giải thích: Anomalies (dữ liệu bất thường/outliers) hữu ích phát hiện lỗi runtime, nhưng không trực tiếp chỉ ra nhu cầu retrain mô hình (anomalies có thể do input xấu, không phải model cũ). Retrain dựa anomalies có thể lãng phí (model vẫn tốt), không tập trung vào drift – nguyên nhân chính làm model outdated.

  • ✅ Phương án ĐÚNG: Enable model monitoring on the Vertex AI endpoint. Configure Pub/Sub to call the Cloud Function when feature drift is detected
    Giải thích: Như đã nêu ở phần đáp án đúng. Feature drift là metric chính xác nhất để quyết định retrain trong Vertex AI (threshold configurable, ví dụ: Jensen-Shannon divergence > 0.1). Tích hợp Pub/Sub notification tự động, kích hoạt pipeline chỉ khi cần → Cân bằng hoàn hảo up-to-date & low cost. Hỗ trợ đầy đủ trong Vertex AI v1beta1+ (2024-2026).

Câu 189
Your company stores a large number of audio files of phone calls made to your customer call center in an on-premises database. Each audio file is in wav format and is approximately 5 minutes long. You need to analyze these audio files for customer sentiment. You plan to use the Speech-to-Text API You want to use the most efficient approach. What should you do?
  1. A 1. Upload the audio files to Cloud Storage
    2. Call the speech:longrunningrecognize API endpoint to generate transcriptions
    3. Call the predict method of an AutoML sentiment analysis model to analyze the transcriptions.
  2. B 1. Upload the audio files to Cloud Storage.
    2. Call the speech:longrunningrecognize API endpoint to generate transcriptions
    3. Create a Cloud Function that calls the Natural Language API by using the analyzeSentiment method
  3. C 1. Iterate over your local files in Python
    2. Use the Speech-to-Text Python library to create a speech.RecognitionAudio object, and set the content to the audio file data
    3. Call the speech:recognize API endpoint to generate transcriptions
    4. Call the predict method of an AutoML sentiment analysis model to analyze the transcriptions.
  4. D 1. Iterate over your local files in Python
    2. Use the Speech-to-Text Python Library to create a speech.RecognitionAudio object and set the content to the audio file data
    3. Call the speech:longrunningrecognize API endpoint to generate transcriptions.
    4. Call the Natural Language API by using the analyzeSentiment method
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi mô tả tình huống công ty bạn lưu trữ số lượng lớn file audio (định dạng WAV, mỗi file khoảng 5 phút) từ các cuộc gọi đến trung tâm chăm sóc khách hàng trong cơ sở dữ liệu on-premises. Nhiệm vụ là phân tích cảm xúc khách hàng (sentiment analysis) bằng cách sử dụng Speech-to-Text API để chuyển đổi audio thành văn bản (transcription), sau đó phân tích văn bản đó. Mục tiêu là chọn cách tiếp cận hiệu quả nhất (most efficient approach).

Các yếu tố chính cần xem xét (dựa trên tài liệu Google Cloud cập nhật đến 2026):

  • File audio dài 5 phút → Phải dùng API asynchronous như speech:longrunningrecognize (hỗ trợ file >60 giây, xử lý batch lớn hiệu quả).
  • Số lượng lớn file → Tránh xử lý local (Python iterate) vì không scalable, tốn tài nguyên on-premises; ưu tiên Cloud Storage để lưu trữ và xử lý cloud-native.
  • Sentiment analysis: Natural Language API (analyzeSentiment) là dịch vụ managed sẵn dùng, không cần train model, phù hợp ngay lập tức. AutoML yêu cầu huấn luyện model trước, kém efficient hơn.
  • Hiệu quả tổng thể: Sử dụng serverless như Cloud Functions để tự động hóa quy trình transcription → sentiment, tránh blocking.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng là phương án thứ 2:

1. Upload the audio files to Cloud Storage.
2. Call the speech:longrunningrecognize API endpoint to generate transcriptions
3. Create a Cloud Function that calls the Natural Language API by using the analyzeSentiment method

Lý do chi tiết 🛠️:

  • Bước 1: Upload lên Cloud Storage là cách scalable và efficient nhất cho large files từ on-premises (dùng gsutil hoặc Transfer Service). Cloud Storage hỗ trợ WAV trực tiếp, chi phí thấp, tích hợp dễ với Speech-to-Text.
  • Bước 2: speech:longrunningrecognize là API asynchronous lý tưởng cho file dài 5 phút và batch lớn – không block, xử lý song song, kết quả lưu về Cloud Storage hoặc Pub/Sub.
  • Bước 3: Cloud Functions (serverless) trigger tự động sau transcription (qua Pub/Sub hoặc Storage events), gọi Natural Language API analyzeSentiment – dịch vụ pre-trained, nhanh, không cần model custom. Toàn bộ quy trình cloud-native, auto-scale, cost-efficient.
  • Tổng thể: Efficient nhất vì batch processing, không local compute, tích hợp mượt mà (giá Speech-to-Text ~$0.006/phút đến 2026).

❌ Phân tích tất cả các phương án (đúng/sai)

  • Phương án 1 (SAI):

    1. Upload the audio files to Cloud Storage
    2. Call the speech:longrunningrecognize API endpoint to generate transcriptions
    3. Call the predict method of an AutoML sentiment analysis model to analyze the transcriptions.
    

    ❌ Sai vì: Bước 1-2 đúng (upload + longrunningrecognize efficient). Nhưng bước 3 dùng AutoML predict yêu cầu model đã train trước (Vertex AI AutoML), không ready-to-use như NL API. Phải tốn thời gian thu thập data/label để train → không efficient cho task nhanh.

  • Phương án 2 (ĐÚNG): ✅ Đã giải thích ở trên – Hoàn hảo về scalability, async, serverless.

  • Phương án 3 (SAI):

    1. Iterate over your local files in Python
    2. Use the Speech-to-Text Python library to create a speech.RecognitionAudio object, and set the content to the audio file data
    3. Call the speech:recognize API endpoint to generate transcriptions
    4. Call the predict method of an AutoML sentiment analysis model to analyze the transcriptions.
    

    ❌ Sai nghiêm trọng vì:

    • Local iterate Python trên on-premises → Không scalable cho "large number" files, tốn bandwidth/memory.
    • speech:recognize chỉ cho file ngắn <60 giây (sync, blocking) → Lỗi với 5 phút, timeout dễ xảy ra.
    • AutoML predict → Cần train model, không efficient. → Toàn bộ low-efficiency, không phù hợp production.
  • Phương án 4 (SAI):

    1. Iterate over your local files in Python
    2. Use the Speech-to-Text Python Library to create a speech.RecognitionAudio object and set the content to the audio file data
    3. Call the speech:longrunningrecognize API endpoint to generate transcriptions.
    4. Call the Natural Language API by using the analyzeSentiment method
    

    ❌ Sai vì: Bước 3-4 đúng về API (longrunningrecognize + NL analyzeSentiment). Nhưng local iterate + load content vào memory → Không efficient cho large files/batch (tốn RAM on-premises, sequential processing chậm, bandwidth cao). Nên upload Cloud Storage để parallel/async thực sự.

Kết luận 🎯: Phương án 2 là optimal theo best practices GCP 2026 – cloud-first, serverless, pre-trained ML!

Câu 190
You work for a social media company. You want to create a no-code image classification model for an iOS mobile application to identify fashion accessories. You have a labeled dataset in Cloud Storage. You need to configure a training workflow that minimizes cost and serves predictions with the lowest possible latency. What should you do?
  1. A Train the model by using AutoML, and register the model in Vertex AI Model Registry. Configure your mobile application to send batch requests during prediction.
  2. B Train the model by using AutoML Edge, and export it as a Core ML model. Configure your mobile application to use the .mlmodel file directly.
  3. C Train the model by using AutoML Edge, and export the model as a TFLite model. Configure your mobile application to use the .tflite file directly.
  4. D Train the model by using AutoML, and expose the model as a Vertex AI endpoint. Configure your mobile application to invoke the endpoint during prediction.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này thuộc lĩnh vực Machine Learning trên Google Cloud Platform (GCP), cụ thể là Vertex AI và AutoML. Bạn đang làm việc cho một công ty mạng xã hội, cần xây dựng mô hình phân loại hình ảnh no-code (không cần code) cho ứng dụng di động iOS để nhận diện phụ kiện thời trang. Dataset đã được gắn nhãn sẵn và lưu trong Cloud Storage. Yêu cầu chính là:

  • Tối ưu chi phí (minimize cost): Tránh chi phí server, endpoint deployment.
  • Độ trễ thấp nhất (lowest possible latency): Ưu tiên inference on-device để không phụ thuộc mạng. Mục tiêu là thiết lập workflow training phù hợp, tận dụng các công cụ no-code như AutoML cho image classification.

📘 Tài liệu tham khảo:

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Train the model by using AutoML Edge, and export it as a Core ML model. Configure your mobile application to use the .mlmodel file directly.

Lý do 🛠️:

  • AutoML Edge chuyên dụng cho mô hình on-device (edge devices như mobile), hỗ trợ no-code training image classification từ dataset Cloud Storage.
  • Export Core ML (.mlmodel) là định dạng native cho iOS, tích hợp trực tiếp vào app qua Core ML framework → latency thấp nhất (inference local, không mạng), cost thấp (không cần endpoint/server).
  • Phù hợp hoàn hảo: No-code, on-device cho iOS, tối ưu yêu cầu.

❌ Phân tích tất cả các phương án (đúng/sai)

  • [SAI] Train the model by using AutoML, and register the model in Vertex AI Model Registry. Configure your mobile application to send batch requests during prediction.
    ❌ Sai vì: AutoML (không phải Edge) dùng cho cloud-based models, cần deploy endpoint → chi phí cao (Vertex AI serving fees), latency cao (batch requests qua mạng, không real-time cho mobile). Batch không phù hợp iOS app (thường cần online prediction). Model Registry chỉ quản lý, không giải quyết latency/cost.

  • [ĐÚNG] Train the model by using AutoML Edge, and export it as a Core ML model. Configure your mobile application to use the .mlmodel file directly.
    ✅ Đúng vì: Như phân tích trên – AutoML Edge tối ưu on-device, Core ML native iOS → latency thấp nhất (local inference <1ms), cost gần 0 (chỉ training fee). Hoàn toàn no-code, tích hợp dễ dàng qua Xcode.

  • [SAI] Train the model by using AutoML Edge, and export the model as a TFLite model. Configure your mobile application to use the .tflite file directly.
    ❌ Sai vì: AutoML Edge đúng cho on-device, nhưng TFLite (.tflite) dành cho Android (TensorFlow Lite). iOS không hỗ trợ native TFLite (cần wrapper phức tạp như TensorFlow Lite Swift), dẫn đến latency cao hơn, tích hợp khó → không tối ưu cho iOS so với Core ML.

  • [SAI] Train the model by using AutoML, and expose the model as a Vertex AI endpoint. Configure your mobile application to invoke the endpoint during prediction.
    ❌ Sai vì: AutoML cloud-based yêu cầu Vertex AI endpoint → chi phí serving liên tục (compute hours), latency cao (network round-trip 100-500ms+), không phù hợp mobile real-time. Online prediction qua API kém hơn on-device.

🧩 Tóm tắt insight: Chọn on-device với Core ML là key để cân bằng cost/low-latency trên iOS. Nếu Android, TFLite sẽ đúng hơn! 🚀