Ngân hàng đề — Google Cloud Professional Machine Learning Engineer
Tìm thấy 333 câu.
- A 1 = Dataflow, 2 = AI Platform, 3 = BigQuery
- B 1 = DataProc, 2 = AutoML, 3 = Cloud Bigtable
- C 1 = BigQuery, 2 = AutoML, 3 = Cloud Functions
- D 1 = BigQuery, 2 = AI Platform, 3 = Cloud Storage
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả tình huống xây dựng một mô hình ML để phát hiện bất thường (anomaly detection) trên dữ liệu cảm biến thời gian thực (real-time sensor data).
- Dữ liệu đầu vào đến qua Pub/Sub (dịch vụ messaging streaming của Google Cloud, xử lý incoming requests một cách scalable).
- Mục tiêu: Xây dựng pipeline để xử lý dữ liệu, chạy inference ML, và lưu kết quả phục vụ analytics & visualization (phân tích và trực quan hóa dữ liệu).
Pipeline cần 3 bước chính:
- Bước 1: Xử lý streaming data từ Pub/Sub (cần tool mạnh về real-time ETL/processing).
- Bước 2: Chạy mô hình ML (inference cho anomaly detection).
- Bước 3: Lưu trữ kết quả để query, phân tích (hỗ trợ SQL, viz như Looker/Data Studio).
📘 Kiến thức cập nhật (2024-2026): Google Cloud ưu tiên Dataflow cho streaming, Vertex AI (tiếp nối AI Platform) cho ML serving, BigQuery cho analytics. (Nguồn: Google Cloud Dataflow Docs, Vertex AI Prediction).
✅ Đáp án đúng: 1 = Dataflow, 2 = AI Platform, 3 = BigQuery
Lý do chọn đáp án này:
- Đây là pipeline tối ưu cho real-time ML trên GCP:
🛠️ Dataflow (bước 1): Xử lý streaming từ Pub/Sub bằng Apache Beam, transform data real-time (windowing, anomaly prep).
🧠 AI Platform (bước 2): Deploy & serve custom ML model cho inference thời gian thực (bây giờ là Vertex AI Prediction endpoints).
📊 BigQuery (bước 3): Lưu structured results, hỗ trợ streaming inserts, SQL analytics & viz tích hợp (Looker Studio).
Pipeline flow: Pub/Sub → Dataflow (process) → AI Platform (predict) → BigQuery (store). Hoàn hảo, scalable, serverless.
(Nguồn: GCP Streaming ML Pipeline Best Practices).
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng phương án (giữ nguyên text gốc bằng tiếng Anh). Mỗi phương án đánh dấu ✅ (đúng) hoặc ❌ (sai), với lý do cụ thể bằng tiếng Việt.
-
1 = Dataflow, 2 = AI Platform, 3 = BigQuery
✅ Đúng hoàn toàn. Như giải thích trên: Dataflow lý tưởng cho real-time từ Pub/Sub, AI Platform cho ML inference, BigQuery cho analytics/viz (streaming inserts nhanh, columnar storage tối ưu query). Không có sai sót nào. -
1 = DataProc, 2 = AutoML, 3 = Cloud Bigtable
❌ Sai.- DataProc: Chỉ batch processing (Hadoop/Spark), không phù hợp real-time streaming từ Pub/Sub (thiếu low-latency).
- AutoML: Tool no-code training (không phải serving inference custom model).
- Cloud Bigtable: NoSQL wide-column cho high-throughput OLTP, kém analytics/SQL/viz so với BigQuery (khó query phức tạp).
-
1 = BigQuery, 2 = AutoML, 3 = Cloud Functions
❌ Sai.- BigQuery (bước 1): Là data warehouse, không phải streaming processor (chỉ ingest sau khi process, không transform real-time từ Pub/Sub).
- AutoML: Không serve model real-time (chỉ build/train).
- Cloud Functions (bước 3): Serverless FaaS ngắn hạn, không lưu trữ analytics lâu dài (không viz tốt).
-
1 = BigQuery, 2 = AI Platform, 3 = Cloud Storage
❌ Sai.- BigQuery (bước 1): Không xử lý streaming ETL như Dataflow (chỉ storage/query).
- AI Platform: Đúng cho ML, nhưng pipeline sai từ đầu.
- Cloud Storage (bước 3): Object storage rẻ cho unstructured/raw data, không hỗ trợ SQL analytics/viz nhanh (cần export thêm tool khác).
🧠 Tóm tắt insight: Chọn pipeline phải khớp real-time → ML inference → analytics. Dataflow + AI Platform/Vertex AI + BigQuery là pattern chuẩn GCP 2026! (Nguồn: GCP Well-Architected Framework - ML).
- A 1. Build a tree-based regression model that predicts how many passengers will be picked up at each shuttle station. 2. Dispatch an appropriately sized shuttle and provide the map with the required stops based on the prediction.
- B 1. Build a tree-based classification model that predicts whether the shuttle should pick up passengers at each shuttle station. 2. Dispatch an available shuttle and provide the map with the required stops based on the prediction.
- C 1. Define the optimal route as the shortest route that passes by all shuttle stations with confirmed attendance at the given time under capacity constraints. 2. Dispatch an appropriately sized shuttle and indicate the required stops on the map.
- D 1. Build a reinforcement learning model with tree-based classification models that predict the presence of passengers at shuttle stops as agents and a reward function around a distance-based metric. 2. Dispatch an appropriately sized shuttle and provide the map with the required stops based on the simulated outcome.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi này thuộc chủ đề Machine Learning Engineer trên Google Cloud Platform (GCP), cụ thể là ứng dụng optimization và routing problems trong các tình huống thực tế như quản lý dịch vụ shuttle (xe đưa đón nội bộ).
- Bối cảnh vấn đề 🚀: Tổ chức muốn tối ưu hóa tuyến đường shuttle nội bộ để hiệu quả hơn. Hiện tại, shuttle dừng tại tất cả các điểm đón khách (shuttle stations) khắp thành phố mỗi 30 phút từ 7h sáng đến 10h sáng – điều này lãng phí thời gian và tài nguyên vì không phải lúc nào cũng có khách.
- Dữ liệu sẵn có 📱: Đội ngũ phát triển đã xây dựng ứng dụng trên Google Kubernetes Engine (GKE), yêu cầu người dùng xác nhận sự hiện diện (confirm presence) và điểm đón (shuttle station) trước 1 ngày. Nghĩa là dữ liệu nhu cầu đã được xác nhận chính xác trước, không phải dự đoán.
- Mục tiêu 🎯: Tìm approach tốt nhất để dispatch shuttle với kích thước phù hợp, chỉ dừng tại các điểm cần thiết, dựa trên dữ liệu confirmed, nhằm giảm thời gian di chuyển và tối ưu hóa tuyến đường.
- Liên quan ML/Optimization trên GCP 🛠️: Đây là bài toán Vehicle Routing Problem (VRP) cổ điển với ràng buộc capacity và confirmed demand. Không cần mô hình dự đoán (prediction) vì dữ liệu đã known; thay vào đó, dùng optimization algorithms như OR-Tools (tích hợp sẵn trong Vertex AI) để tính shortest path qua các điểm confirmed.
Kiến thức cập nhật đến 2026 ⚡: Theo tài liệu GCP ML mới nhất (Vertex AI v2.15+, 2025-2026), các vấn đề routing như VRP được giải quyết bằng Cloud Operations Suite kết hợp Vertex AI Optimization hoặc OR-Tools Python API trên GKE, hỗ trợ capacity constraints và real-time dispatch. Không liên quan AWS (có lẽ là nhầm lẫn; AWS dùng Amazon SageMaker Canvas hoặc Route Optimization API).
Nguồn tham khảo 📘:
- Google Cloud Professional ML Engineer Exam Guide (2024-2026): https://cloud.google.com/learn/certification/machine-learning-engineer
- Vertex AI Documentation - Optimization: https://cloud.google.com/vertex-ai/docs/solutions/operations-research
- OR-Tools VRP Solver: https://developers.google.com/optimization/routing/vrp
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng là lựa chọn thứ 3:
- Define the optimal route as the shortest route that passes by all shuttle stations with confirmed attendance at the given time under capacity constraints. 2. Dispatch an appropriately sized shuttle and indicate the required stops on the map.
Lý do chọn ✅:
- 🧮 Phù hợp hoàn hảo với dữ liệu: Ứng dụng GKE đã cung cấp confirmed attendance (dữ liệu chắc chắn), nên không cần ML prediction mà dùng optimization trực tiếp để tính shortest route qua tất cả stations có confirmed, với ràng buộc capacity (kích thước shuttle phù hợp số khách).
- 🚀 Hiệu quả cao nhất: Giải quyết chính xác VRP với known demand, giảm lãng phí so với stopping everywhere. Trên GCP, deploy dễ dàng qua Vertex AI Pipelines hoặc GKE Autoscaling cho real-time dispatch.
- 📈 Best practice ML Engineer: Theo cert GCP, ưu tiên operations research (OR) trước ML khi data deterministic (xác định).
❌ Phân tích tất cả các phương án (đúng/sai)
Dưới đây là giải thích từng phương án một cách chi tiết, giữ nguyên văn bản gốc tiếng Anh. Mỗi cái được đánh dấu ✅ (đúng) hoặc ❌ (sai), kèm lý do dựa trên logic ML/Optimization.
-
[SAI] ❌
- Build a tree-based regression model that predicts how many passengers will be picked up at each shuttle station. 2. Dispatch an appropriately sized shuttle and provide the map with the required stops based on the prediction.
Giải thích sai ❌: Phương án này dùng regression (dự đoán số lượng hành khách) – không cần thiết vì dữ liệu confirmed đã known từ app GKE. Tree-based (như XGBoost trên Vertex AI) chỉ phù hợp historical data không chắc chắn, dẫn đến prediction error (over/under-estimate), lãng phí compute và không optimal. Không giải quyết routing constraints.
- Build a tree-based regression model that predicts how many passengers will be picked up at each shuttle station. 2. Dispatch an appropriately sized shuttle and provide the map with the required stops based on the prediction.
-
[SAI] ❌
- Build a tree-based classification model that predicts whether the shuttle should pick up passengers at each shuttle station. 2. Dispatch an available shuttle and provide the map with the required stops based on the prediction.
Giải thích sai ❌: Classification (phân loại có/không pick up) thừa thãi vì confirmed attendance đã quyết định rõ. Tree-based model (Random Forest) trên BigQuery ML sẽ tạo bias từ historical patterns, bỏ qua real-time confirm. Không handle capacity hoặc shortest route, chỉ predict binary – kém hiệu quả cho VRP.
- Build a tree-based classification model that predicts whether the shuttle should pick up passengers at each shuttle station. 2. Dispatch an available shuttle and provide the map with the required stops based on the prediction.
-
[ĐÚNG] ✅
- Define the optimal route as the shortest route that passes by all shuttle stations with confirmed attendance at the given time under capacity constraints. 2. Dispatch an appropriately sized shuttle and indicate the required stops on the map.
Giải thích đúng ✅: Như đã nêu ở phần đáp án. Optimization thuần túy (không ML predict), dùng tools như OR-Tools để solve TSP/VRP variant với confirmed stops + capacity. Deploy trên GKE: real-time, scalable, tiết kiệm nhất (giảm stops không cần).
- Define the optimal route as the shortest route that passes by all shuttle stations with confirmed attendance at the given time under capacity constraints. 2. Dispatch an appropriately sized shuttle and indicate the required stops on the map.
-
[SAI] ❌
- Build a reinforcement learning model with tree-based classification models that predict the presence of passengers at shuttle stops as agents and a reward function around a distance-based metric. 2. Dispatch an appropriately sized shuttle and provide the map with the required stops based on the simulated outcome.
Giải thích sai ❌: Reinforcement Learning (RL) phức tạp quá mức (agents = stops, reward = distance), kết hợp tree-classification – overkill và không stable cho confirmed data. RL (như trên Vertex AI RL Suite) cần simulation dài hạn, training heavy (GPU-intensive), dễ overfit noisy historical. Không deterministic, kém cho shuttle real-time; GCP recommend OR trước RL.
- Build a reinforcement learning model with tree-based classification models that predict the presence of passengers at shuttle stops as agents and a reward function around a distance-based metric. 2. Dispatch an appropriately sized shuttle and provide the map with the required stops based on the simulated outcome.
Kết luận tổng quát 🏆: Approach đúng tập trung optimization trên known data thay vì predict không cần. Nếu deploy production trên GCP, dùng Vertex AI + OR-Tools để automate dispatch! 🚀
- A Use the class distribution to generate 10% positive examples.
- B Use a convolutional neural network with max pooling and softmax activation.
- C Downsample the data with upweighting to create a sample with 10% positive examples.
- D Remove negative examples until the numbers of positive and negative examples are equal.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào vấn đề class imbalance (mất cân bằng lớp) trong huấn luyện mô hình phân loại (classification). Cụ thể:
- Bạn đang phân tích lỗi linh kiện dây chuyền sản xuất dựa trên dữ liệu cảm biến (sensor readings).
- Dataset có ít hơn 1% mẫu positive (đại diện cho sự cố lỗi - failure incidents), còn lại hầu hết là negative (bình thường).
- Các mô hình phân loại đã thử (như logistic regression, tree-based, neural nets) không hội tụ (không converge), thường do mô hình bị bias về lớp majority (negative), dẫn đến accuracy cao giả tạo nhưng recall/precision cho positive kém.
- Mục tiêu: Giải quyết imbalance để mô hình học tốt hơn lớp thiểu số (positive), giúp detect failure hiệu quả.
Vấn đề này phổ biến trong ML thực tế (như anomaly detection), và giải pháp cần cân bằng dataset mà không làm mất thông tin quan trọng. (Kiến thức cập nhật: AWS SageMaker, Vertex AI hỗ trợ imbalance qua built-in algorithms như XGBoost vớiscale_pos_weightđến 2026).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Downsample the data with upweighting to create a sample with 10% positive examples.
🛠️ Lý do chi tiết:
- Downsample: Giảm số lượng mẫu negative (majority class) để cân bằng tỷ lệ, tránh overfitting vào lớp đa số.
- Upweighting: Tăng trọng số (weight) cho lớp positive trong loss function (ví dụ:
sample_weighttrong scikit-learn/XGBoost), giúp mô hình penalize lỗi trên positive mạnh hơn mà không cần oversample (tạo dữ liệu giả). - Kết hợp tạo dataset con với 10% positive (từ <1% lên hợp lý, tránh extreme balance 50/50 gây bias ngược).
- Phương pháp này hiệu quả cho convergence, được khuyến nghị trong AWS SageMaker Autopilot/BlazingText (2026 updates hỗ trợ automatic resampling + weighting). Không mất dữ liệu gốc, dễ implement.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn một cách chi tiết, giữ nguyên văn bản gốc bằng tiếng Anh:
-
Use the class distribution to generate 10% positive examples.
❌ Sai vì: Phương án này gợi ý oversampling (generate synthetic positive dựa trên phân phối lớp), nhưng không rõ cách generate (SMOTE? random?). Với <1% positive gốc, generate dễ tạo noise/artifacts, dẫn đến overfitting và mô hình kém generalize. Không giải quyết trực tiếp convergence (models vẫn bias nếu không upweight). AWS/ML best practices khuyên tránh pure oversampling cho imbalance cực đoan. -
Use a convolutional neural network with max pooling and softmax activation.
❌ Sai vì: Đây là kiến trúc CNN (phù hợp image/tabular data với pooling + softmax cho multi-class), nhưng không liên quan đến class imbalance. CNN không tự fix imbalance; vẫn cần resampling/weighting bổ sung. Dữ liệu sensor thường tabular/time-series, không phải image → irrelevant, tốn tài nguyên mà không converge. (AWS Forecast/EC2 hỗ trợ CNN nhưng không phải giải pháp core cho imbalance). -
Downsample the data with upweighting to create a sample with 10% positive examples.
✅ Đúng như đã giải thích ở trên: Kết hợp downsample (giảm negative) + upweighting (tăng tầm quan trọng positive) là best practice cho imbalance <1%, giúp convergence nhanh, balance recall/F1-score. Đạt tỷ lệ 10% positive là realistic target. -
Remove negative examples until the numbers of positive and negative examples are equal.
❌ Sai vì: Pure downsampling cực đoan (remove negative đến 50/50) làm mất quá nhiều dữ liệu (với <1% positive, chỉ giữ ~200x positive samples → dataset quá nhỏ, underfitting). Không upweight → mô hình vẫn kém detect real-world (negative dominance ngoài production). AWS docs cảnh báo undersampling thuần túy kém cho extreme imbalance.
📘 Tài liệu tham khảo
- Google Cloud ML Engineer Exam Guide (câu hỏi tương tự): cloud.google.com/certification/guides/machine-learning-engineer – Section: Data Preparation & Imbalance.
- AWS SageMaker Documentation (2026 updates): docs.aws.amazon.com/sagemaker/latest/dg/handle-imbalanced.html – Khuyến nghị class weights + resampling.
- Scikit-learn Imbalanced-learn: imbalanced-learn.org/stable/references/index.html –
RandomUnderSampler+class_weight='balanced'. - Paper tham khảo: "Learning from Imbalanced Data" (IEEE, 2009, vẫn core đến 2026).
Hy vọng phân tích giúp bạn nắm vững! 🚀 Nếu cần code demo trên SageMaker/Vertex AI, hỏi thêm nhé!
- A Use Data Fusion's GUI to build the transformation pipelines, and then write the data into BigQuery.
- B Convert your PySpark into SparkSQL queries to transform the data, and then run your pipeline on Dataproc to write the data into BigQuery.
- C Ingest your data into Cloud SQL, convert your PySpark commands into SQL queries to transform the data, and then use federated queries from BigQuery for machine learning.
- D Ingest your data into BigQuery using BigQuery Load, convert your PySpark commands into BigQuery SQL queries to transform the data, and then write the transformations to a new table.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi tập trung vào việc xây dựng lại pipeline ML cho dữ liệu có cấu trúc (structured data) trên Google Cloud, với các yêu cầu chính:
- Hiện tại: Sử dụng PySpark để transform dữ liệu quy mô lớn, nhưng pipeline mất hơn 12 giờ để chạy → Cần tăng tốc phát triển và thời gian chạy.
- Yêu cầu mới: Sử dụng công cụ serverless (không quản lý server), cú pháp SQL (dễ dàng hơn PySpark), dữ liệu raw đã nằm trong Cloud Storage (GCS).
- Mục tiêu: Xây dựng pipeline nhanh, hiệu quả cho ML, tận dụng dữ liệu đã sẵn sàng.
Vấn đề cốt lõi là chuyển từ PySpark (Spark-based, managed nhưng tốn thời gian) sang serverless SQL-based để scale nhanh, chi phí thấp, và phát triển dễ dàng hơn. BigQuery là lựa chọn lý tưởng vì nó serverless, hỗ trợ SQL mạnh mẽ, load dữ liệu từ GCS siêu nhanh (batch load jobs), và phù hợp ML (tích hợp Vertex AI). Kiến thức cập nhật đến 2026: BigQuery vẫn là data warehouse serverless hàng đầu GCP, với BigQuery Storage Write API và ML integrations mới (như BigQuery ML v2 với Gemini models).
✅ Đáp án đúng
Ingest your data into BigQuery using BigQuery Load, convert your PySpark commands into BigQuery SQL queries to transform the data, and then write the transformations to a new table.
Lý do lựa chọn:
- ✅ Serverless hoàn toàn: BigQuery Load jobs từ GCS là serverless, tự động scale, load hàng TB dữ liệu trong phút (không cần cluster như Spark).
- ✅ SQL syntax: Chuyển PySpark logic (DataFrames) sang BigQuery SQL (dễ dàng, nhanh phát triển hơn code Python/Scala).
- ✅ Tốc độ cao: Load từ GCS → BigQuery chỉ mất giây/phút, transform bằng SQL query (parallel execution), write ra table mới ngay lập tức. Giảm từ 12h xuống <1h dễ dàng.
- ✅ Phù hợp ML pipeline: BigQuery table mới sẵn sàng cho Vertex AI, AutoML, hoặc BigQuery ML.
- 🛠️ Quy trình:
bq load --source_format=... gs://bucket/raw.csv dataset.table→CREATE TABLE new_table AS SELECT ... FROM raw_table.
📋 Giải thích tất cả các phương án
Dưới đây là phân tích từng lựa chọn (giữ nguyên văn bản gốc tiếng Anh), đánh dấu ✅ đúng hoặc ❌ sai, với lý do chi tiết bằng tiếng Việt:
-
❌ Use Data Fusion's GUI to build the transformation pipelines, and then write the data into BigQuery.
Sai vì: Data Fusion (Cloud Data Fusion) là ETL tool dựa trên Apache Beam/CDAP, sử dụng Spark dưới hood → Không serverless thuần (cần quản lý instances), GUI giúp build nhanh nhưng vẫn chậm như PySpark (12h+), không ưu tiên SQL syntax chính. Phù hợp low-code nhưng không đáp ứng "serverless tool and SQL syntax" trực tiếp. -
❌ Convert your PySpark into SparkSQL queries to transform the data, and then run your pipeline on Dataproc to write the data into BigQuery.
Sai vì: Dataproc là managed Spark cluster (không serverless), SparkSQL chỉ là SQL trên Spark → Vẫn tốn thời gian khởi tạo cluster, chạy job (giống PySpark hiện tại, không giảm 12h). Không phải serverless, phải convert thủ công và quản lý cluster. -
❌ Ingest your data into Cloud SQL, convert your PySpark commands into SQL queries to transform the data, and then use federated queries from BigQuery for machine learning.
Sai vì: Cloud SQL là managed relational DB (MySQL/PostgreSQL), không serverless cho big data scale (giới hạn size/throughput), ingest từ GCS chậm và tốn kém. Federated queries từ BigQuery → Chậm (external tables), không hiệu quả cho transform lớn (12h+), không phù hợp structured data quy mô ML. -
✅ Ingest your data into BigQuery using BigQuery Load, convert your PySpark commands into BigQuery SQL queries to transform the data, and then write the transformations to a new table.
(Đã giải thích chi tiết ở phần trên – hoàn hảo khớp yêu cầu!)
📘 Tài liệu tham khảo (cập nhật 2026)
- BigQuery Load Jobs: Google Cloud Docs - Loading Data – Hỗ trợ batch load từ GCS nhanh chóng.
- BigQuery SQL vs Spark: Migrate from Spark to BigQuery.
- Serverless ML Pipeline: Vertex AI & BigQuery Integration – Phiên bản mới với BigQuery ML Gen2.
- Benchmark: BigQuery thường nhanh gấp 10-100x Spark cho SQL workloads (theo GCP case studies 2025).
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần ví dụ code SQL convert từ PySpark, hãy hỏi thêm nhé! 😊
- A Use the AI Platform custom containers feature to receive training jobs using any framework.
- B Configure Kubeflow to run on Google Kubernetes Engine and receive training jobs through TF Job.
- C Create a library of VM images on Compute Engine, and publish these images on a centralized repository.
- D Set up Slurm workload manager to receive jobs that can be scheduled to run on your cloud infrastructure.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả tình huống một quản lý đội ngũ data scientists đang sử dụng hệ thống backend dựa trên cloud để submit các job huấn luyện mô hình ML. Hệ thống hiện tại khó quản lý (difficult to administer), và họ muốn chuyển sang dịch vụ managed service (dịch vụ được quản lý hoàn toàn bởi cloud provider). Đội ngũ sử dụng nhiều framework đa dạng: Keras, PyTorch, Theano, Scikit-learn, và cả custom libraries (thư viện tùy chỉnh).
📌 Mục tiêu chính: Tìm giải pháp managed service hỗ trợ bất kỳ framework nào, dễ quản lý, không cần tự admin infrastructure.
🛠️ Đây là câu hỏi điển hình trong Google Cloud về Vertex AI (trước đây là AI Platform), tập trung vào tính linh hoạt cho ML training jobs với custom containers. (Cập nhật đến 2026: Vertex AI Training hỗ trợ custom containers đầy đủ cho mọi framework).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Use the AI Platform custom containers feature to receive training jobs using any framework.
Lý do:
- Vertex AI (AI Platform) cung cấp custom containers cho training jobs, cho phép sử dụng bất kỳ framework hoặc custom libraries nào mà không bị giới hạn (Keras, PyTorch, Scikit-learn, v.v.).
- Đây là managed service hoàn toàn: Google tự quản lý scaling, infrastructure, monitoring – giải quyết vấn đề "difficult to administer".
- Data scientists chỉ cần submit job qua container image (Docker), rất linh hoạt và dễ integrate.
📘 Tài liệu tham khảo: Vertex AI Custom Training Containers (Google Cloud Docs, cập nhật 2025).
📋 Giải thích tất cả các phương án
Dưới đây là phân tích chi tiết từng lựa chọn, với ✅ cho đúng và ❌ cho sai. Tôi giữ nguyên văn bản gốc bằng tiếng Anh, chỉ giải thích bằng tiếng Việt:
-
✅ Use the AI Platform custom containers feature to receive training jobs using any framework.
🟢 Đúng vì: Như đã giải thích ở trên, đây là giải pháp managed, hỗ trợ mọi framework/custom libs qua Docker containers. Hoàn hảo cho đa dạng tools của data scientists, tự động scale và managed end-to-end. -
❌ Configure Kubeflow to run on Google Kubernetes Engine and receive training jobs through TF Job.
🔴 Sai vì: Kubeflow trên GKE là managed nhưng TFJob chỉ hỗ trợ TensorFlow chủ yếu, không linh hoạt cho PyTorch/Keras/Scikit-learn/custom libs. Bạn vẫn phải tự quản lý K8s cluster (không fully managed như Vertex AI), vi phạm yêu cầu "managed service dễ admin". -
❌ Create a library of VM images on Compute Engine, and publish these images on a centralized repository.
🔴 Sai vì: Compute Engine VM images yêu cầu tự quản lý toàn bộ (provisioning, scaling, patching, networking). Không phải managed ML service, data scientists phải handle infrastructure – làm tình hình khó admin hơn, không hỗ trợ job submission tự động cho nhiều frameworks. -
❌ Set up Slurm workload manager to receive jobs that can be scheduled to run on your cloud infrastructure.
🔴 Sai vì: Slurm là HPC workload manager (tự cài trên GCE/GKE), không phải managed service của Google Cloud. Bạn phải tự setup, maintain cluster – rất phức tạp cho ML jobs đa framework, không tích hợp native với ML tools như Vertex AI.
🏆 Kết luận & Lời khuyên
Giải pháp tối ưu là Vertex AI Custom Containers để chuyển đổi mượt mà sang managed ML platform. Nếu triển khai thực tế, bắt đầu với Quickstart Vertex AI Training để test với PyTorch/Keras. 🚀 Nếu cần hỗ trợ thêm về migration, hãy cung cấp chi tiết workload!
- A Keep the original test dataset unchanged even if newer products are incorporated into retraining.
- B Extend your test dataset with images of the newer products when they are introduced to retraining.
- C Replace your test dataset with images of the newer products when they are introduced to retraining.
- D Update your test dataset with images of the newer products when your evaluation metrics drop below a pre-decided threshold.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả một công ty bán lẻ trực tuyến đang xây dựng động cơ tìm kiếm hình ảnh (visual search engine) trên Google Cloud. Họ đã thiết lập đường ống ML end-to-end để phân loại hình ảnh có chứa sản phẩm của công ty hay không. Với việc sắp ra mắt sản phẩm mới, họ đã cấu hình chức năng retraining để đưa dữ liệu mới vào mô hình ML. Đồng thời, họ muốn sử dụng dịch vụ continuous evaluation của AI Platform (nay là Vertex AI) để đảm bảo mô hình duy trì độ chính xác cao trên tập test dataset.
Vấn đề cốt lõi: Làm thế nào để quản lý test dataset khi có dữ liệu sản phẩm mới, nhằm đảm bảo đánh giá liên tục (continuous evaluation) phản ánh đúng hiệu suất mô hình mà không bị lệch (data shift hoặc concept drift). Đây là thực hành MLOps chuẩn trên Google Cloud Vertex AI, nơi continuous evaluation theo dõi metrics như accuracy trên test data mới nhất. Kiến thức cập nhật đến 2026: Vertex AI (thay thế AI Platform từ 2022) hỗ trợ Model Monitoring và continuous evaluation để phát hiện drift, yêu cầu test set phải đại diện cho data thực tế đang thay đổi.
📘 Tài liệu tham khảo:
- Vertex AI Continuous Evaluation & Monitoring (Google Cloud Docs, cập nhật 2025).
- MLOps on Google Cloud: Managing Data Drift (Best practices 2024-2026).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Extend your test dataset with images of the newer products when they are introduced to retraining.
Lý do:
- Khi retraining với sản phẩm mới, test dataset cần được mở rộng (extend) để bao gồm cả hình ảnh sản phẩm cũ và mới, đảm bảo đánh giá toàn diện, tránh data shift.
- Continuous evaluation của Vertex AI yêu cầu test set đại diện cho distribution data hiện tại (training data), nên phải thêm data mới đồng bộ với retraining để metrics (như accuracy) luôn chính xác và đáng tin cậy.
- Không replace hay giữ nguyên sẽ dẫn đến đánh giá sai lệch. Đây là best practice MLOps: versioning test set song song với model retraining. ✅
🛠️ Giải thích tất cả các phương án
-
[SAI] Keep the original test dataset unchanged even if newer products are incorporated into retraining.
❌ Sai vì: Giữ nguyên test set cũ sẽ không đại diện cho data mới (sản phẩm mới), dẫn đến data distribution shift. Continuous evaluation sẽ báo cáo metrics sai (accuracy cao giả tạo trên data cũ), không phát hiện drift thực tế. Vertex AI khuyến nghị cập nhật test set để khớp training distribution. -
[ĐÚNG] Extend your test dataset with images of the newer products when they are introduced to retraining.
✅ Đúng vì: Mở rộng test set bằng cách thêm (extend) hình ảnh sản phẩm mới đồng bộ với retraining, giữ coverage data cũ + mới. Đảm bảo continuous evaluation đo lường chính xác trên toàn bộ distribution hiện tại, phù hợp với Vertex AI Model Monitoring (threshold-based alerts). -
[SAI] Replace your test dataset with images of the newer products when they are introduced to retraining.
❌ Sai vì: Thay thế hoàn toàn (replace) sẽ mất dữ liệu cũ, khiến đánh giá chỉ tập trung vào sản phẩm mới, bỏ qua hiệu suất trên sản phẩm legacy. Dẫn đến under-evaluation data cũ, vi phạm nguyên tắc test set phải comprehensive trong continuous evaluation. -
[SAI] Update your test dataset with images of the newer products when your evaluation metrics drop below a pre-decided threshold.
❌ Sai vì: Việc cập nhật chỉ khi metrics drop là phản ứng (reactive), không chủ động. Continuous evaluation cần test set luôn cập nhật đồng bộ với retraining để dự đoán drift sớm, không chờ threshold. Điều này có thể bỏ lỡ vấn đề sớm, trái với MLOps proactive trên Vertex AI. 🧩
- A Configure AutoML Tables to perform the classification task.
- B Run a BigQuery ML task to perform logistic regression for the classification.
- C Use AI Platform Notebooks to run the classification model with pandas library.
- D Use AI Platform to run the classification model job configured for hyperparameter tuning.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi yêu cầu xây dựng luồng phân loại (classification workflows) trên nhiều bộ dữ liệu có cấu trúc (structured datasets) đang lưu trữ trong BigQuery. Người dùng muốn thực hiện các bước sau mà không cần viết code:
- Phân tích dữ liệu thăm dò (exploratory data analysis - EDA),
- Chọn lọc đặc trưng (feature selection),
- Xây dựng và huấn luyện mô hình (model building & training),
- Tối ưu siêu tham số (hyperparameter tuning),
- Triển khai phục vụ mô hình (serving).
🔍 Yêu cầu cốt lõi: Toàn bộ quy trình phải no-code (không code), lặp lại nhiều lần trên dữ liệu BigQuery. Đây là tình huống điển hình cho các công cụ ML tự động hóa trên Google Cloud Platform (GCP), tập trung vào dữ liệu bảng (tabular data).
📘 Tài liệu tham khảo:
- Vertex AI (trước là AutoML Tables) - GCP Docs (cập nhật 2024-2026: AutoML Tables đã tích hợp vào Vertex AI, hỗ trợ đầy đủ no-code cho tabular classification).
- BigQuery ML Overview.
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Configure AutoML Tables to perform the classification task.
Lý do chi tiết 🛠️:
- AutoML Tables (nay là phần của Vertex AI Tabular) là giải pháp no-code hoàn chỉnh dành riêng cho dữ liệu bảng trong BigQuery. Nó tự động hóa tất cả các bước yêu cầu: EDA (tự phân tích dữ liệu), feature selection (tối ưu đặc trưng), model building/training (hỗ trợ classification như logistic regression, random forest,...), hyperparameter tuning (tự động), và serving (deploy endpoint dễ dàng).
- Hỗ trợ trực tiếp import dữ liệu từ BigQuery, lặp lại workflow nhanh chóng mà không cần code. Đây là lựa chọn tối ưu cho structured data classification theo best practices GCP đến 2026.
- ✅ Hoàn hảo khớp 100% yêu cầu no-code và toàn diện.
📋 Giải thích tất cả các phương án (đúng/sai)
Dưới đây là phân tích từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Mỗi phương án được đánh giá dựa trên khả năng đáp ứng no-code và toàn bộ các bước yêu cầu.
-
Configure AutoML Tables to perform the classification task.
✅ Đúng 🏆: Như đã giải thích, đây là công cụ no-code chuyên biệt cho tabular classification trên BigQuery, tự động hóa EDA → tuning → serving. Không cần code, phù hợp lặp lại nhiều lần. (Vertex AI update 2026: Vẫn hỗ trợ seamless migration từ AutoML Tables). -
Run a BigQuery ML task to perform logistic regression for the classification.
❌ Sai 🚫: BigQuery ML (BQML) hỗ trợ classification qua SQL (như CREATE MODEL với logistic_reg), nhưng yêu cầu viết SQL code cho training/tuning. Không có EDA/feature selection tự động đầy đủ, và hyperparameter tuning hạn chế (không no-code hoàn chỉnh). Không khớp yêu cầu "without writing code". -
Use AI Platform Notebooks to run the classification model with pandas library.
❌ Sai 📝: AI Platform Notebooks (nay là Vertex AI Workbench) là môi trường Jupyter, bắt buộc viết code Python với pandas để EDA, feature selection, training (ví dụ scikit-learn), tuning. Không phải no-code, chỉ phù hợp dev thủ công, không tự động hóa workflow. -
Use AI Platform to run the classification model job configured for hyperparameter tuning.
❌ Sai ⚙️: AI Platform (nay là Vertex AI Training) hỗ trợ custom training jobs với hyperparameter tuning, nhưng yêu cầu viết code (Python scripts, containerized jobs). Không tự động EDA/feature selection/model building, chỉ tuning phần cuối. Không đáp ứng no-code toàn diện.
🧠 Kết luận nổi bật: AutoML Tables là lựa chọn duy nhất no-code end-to-end cho tabular classification trên BigQuery. Các phương án khác đều đòi hỏi code, vi phạm yêu cầu chính! 🚀
- A Configure Kubeflow Pipelines to schedule your multi-step workflow from training to deploying your model.
- B Use a model trained and deployed on BigQuery ML, and trigger retraining with the scheduled query feature in BigQuery.
- C Write a Cloud Functions script that launches a training and deploying job on AI Platform that is triggered by Cloud Scheduler.
- D Use Cloud Composer to programmatically schedule a Dataflow job that executes the workflow from training to deploying your model.
Xem giải thích
🧩 Phân tích chi tiết nội dung câu hỏi
Câu hỏi mô tả tình huống bạn làm việc cho một công ty vận tải công cộng, cần xây dựng mô hình dự đoán thời gian trễ (delay times) cho nhiều tuyến đường khác nhau. Các dự đoán phải được phục vụ thời gian thực (real-time) trực tiếp cho người dùng qua ứng dụng di động. Dữ liệu bị ảnh hưởng bởi mùa vụ và tăng dân số, nên mô hình cần được retrain hàng tháng. Yêu cầu tuân thủ best practices được Google khuyến nghị cho kiến trúc end-to-end của mô hình dự đoán, từ huấn luyện đến triển khai.
Mục tiêu chính: Xây dựng pipeline ML tự động hóa quy trình multi-step (đa bước: data prep, training, evaluation, deployment), hỗ trợ retrain định kỳ, và phục vụ inference real-time. Đây là kịch bản điển hình trong Google Cloud MLOps, nhấn mạnh orchestration (điều phối) workflow ML toàn diện.
📘 Tài liệu tham khảo:
- Google Cloud ML Best Practices: Architecting with Google Kubernetes Engine & Vertex AI Pipelines/Kubeflow (cập nhật 2024-2026, Kubeflow vẫn là nền tảng core cho GKE-based workflows).
✅ Đáp án đúng và lý do lựa chọn
Đáp án đúng: Configure Kubeflow Pipelines to schedule your multi-step workflow from training to deploying your model.
Lý do chọn 🛠️:
- Kubeflow Pipelines là best practice được Google khuyến nghị cho end-to-end ML workflows trên Google Cloud (chạy trên GKE). Nó hỗ trợ đa bước tự động (training, validation, deployment), lập lịch retrain hàng tháng qua cron jobs, và tích hợp seamless với real-time serving (qua Vertex AI Endpoints hoặc Seldon/KFServing).
- Phù hợp hoàn hảo với yêu cầu: Xử lý dữ liệu thay đổi theo mùa, scalable cho multiple routes, và tuân thủ MLOps principles (CI/CD for ML).
- Cập nhật mới nhất (2026): Kubeflow 2.x + Vertex AI Pipelines hỗ trợ hybrid, đảm bảo reproducibility và monitoring qua TensorBoard/ML Metadata.
📋 Giải thích tất cả các phương án
-
Configure Kubeflow Pipelines to schedule your multi-step workflow from training to deploying your model.
✅ Đúng 🏆: Như giải thích trên, đây là giải pháp toàn diện nhất, được Google ưu tiên cho production ML pipelines. Hỗ trợ custom components, versioning, và deployment real-time serving. Lý tưởng cho retrain định kỳ và multi-model (nhiều tuyến đường). -
Use a model trained and deployed on BigQuery ML, and trigger retraining with the scheduled query feature in BigQuery.
❌ Sai 🚫: BigQuery ML chỉ phù hợp cho SQL-based models đơn giản (như linear regression), không hỗ trợ custom ML frameworks (TensorFlow/PyTorch) hoặc end-to-end deployment real-time phức tạp. Scheduled queries chỉ retrain cơ bản, thiếu orchestration đa bước và serving latency thấp cho app. -
Write a Cloud Functions script that launches a training and deploying job on AI Platform that is triggered by Cloud Scheduler.
❌ Sai ⚠️: Cloud Functions (serverless) + Cloud Scheduler chỉ là trigger đơn giản, không đủ cho multi-step workflow (không có dependency management, error handling, hoặc parallelism). AI Platform (nay là Vertex AI) tốt cho training/deploy riêng lẻ, nhưng thiếu pipeline orchestration toàn diện như Kubeflow. -
Use Cloud Composer to programmatically schedule a Dataflow job that executes the workflow from training to deploying your model.
❌ Sai 🔧: Cloud Composer (Airflow-based) mạnh cho DAGs data processing, nhưng không phải best practice cho ML (thiếu ML-specific components như hyperparam tuning, model serving). Dataflow chỉ batch/stream processing, không tối ưu cho training/deploy ML end-to-end so với Kubeflow.
Kết luận 🌟: Kubeflow là lựa chọn tối ưu theo Google Cloud Professional ML Engineer certification (exam guide 2024-2026), đảm bảo scalability, reliability và best practices cho real-time ML apps!
- A Use Cloud Functions to identify changes to your code in Cloud Storage and trigger a retraining job.
- B Use the gcloud command-line tool to submit training jobs on AI Platform when you update your code.
- C Use Cloud Build linked with Cloud Source Repositories to trigger retraining when new code is pushed to the repository.
- D Create an automated workflow in Cloud Composer that runs daily and looks for changes in code in Cloud Storage using a sensor.
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi mô tả tình huống bạn đang phát triển các mô hình ML cho nhiệm vụ phân đoạn hình ảnh (image segmentation) trên ảnh chụp CT scan, sử dụng AI Platform (nay là một phần của Vertex AI trên Google Cloud). Bạn thường xuyên cập nhật kiến trúc mô hình dựa trên các bài báo nghiên cứu mới nhất, và cần huấn luyện lại (retrain) trên cùng một bộ dữ liệu để so sánh hiệu suất (benchmark).
Yêu cầu chính cần đáp ứng:
- Giảm thiểu chi phí tính toán (minimize computation costs): Tránh huấn luyện không cần thiết.
- Giảm thiểu can thiệp thủ công (minimize manual intervention): Tự động hóa quy trình.
- Kiểm soát phiên bản code (version control for your code): Quản lý thay đổi code một cách có hệ thống.
📘 Bối cảnh cập nhật 2026: AI Platform đã được tích hợp sâu vào Vertex AI (phiên bản mới nhất), nhưng các dịch vụ như Cloud Build, Cloud Source Repositories vẫn hỗ trợ trigger tự động cho training jobs. Cloud Source Repositories là Git repo managed bởi Google Cloud, tích hợp tốt với Cloud Build cho CI/CD.
✅ Đáp án đúng
Use Cloud Build linked with Cloud Source Repositories to trigger retraining when new code is pushed to the repository.
Lý do chọn đáp án này:
- 🛠️ Tự động hóa hoàn hảo: Khi push code mới lên Cloud Source Repositories (repo Git), Cloud Build tự động trigger pipeline, chạy training job trên AI Platform/Vertex AI mà không cần can thiệp thủ công.
- 💰 Tiết kiệm chi phí: Chỉ huấn luyện khi code thực sự thay đổi (on push), tránh lãng phí tài nguyên.
- 🔄 Version control đầy đủ: Repo Git cho phép track thay đổi code, branch, commit – lý tưởng cho việc benchmark các kiến trúc mới từ research papers.
- 🏆 Phù hợp nhất: Đáp ứng tất cả yêu cầu, là best practice cho ML workflows trên GCP (CI/CD for ML).
📋 Giải thích chi tiết tất cả các phương án
-
❌ [SAI] Use Cloud Functions to identify changes to your code in Cloud Storage and trigger a retraining job.
Lý do sai: Cloud Functions chỉ phát hiện thay đổi file trong Cloud Storage (không phải repo Git), thiếu version control thực sự (Storage chỉ lưu file, không track history/commit). Dễ miss thay đổi nhỏ, và vẫn cần config thủ công để detect – không tối ưu chi phí/manual. -
❌ [SAI] Use the gcloud command-line tool to submit training jobs on AI Platform when you update your code.
Lý do sai: Đây là cách thủ công hoàn toàn (chạy lệnh gcloud mỗi lần), đòi hỏi can thiệp liên tục – vi phạm yêu cầu "minimize manual intervention". Không tự động, không version control tích hợp. -
✅ [ĐÚNG] Use Cloud Build linked with Cloud Source Repositories to trigger retraining when new code is pushed to the repository.
Lý do đúng: Như đã giải thích ở trên – tự động trên push code, version control qua Git repo, tiết kiệm chi phí và không manual. Đây là giải pháp CI/CD chuẩn cho ML trên Vertex AI. -
❌ [SAI] Create an automated workflow in Cloud Composer that runs daily and looks for changes in code in Cloud Storage using a sensor.
Lý do sai: Cloud Composer (Airflow managed) chạy hàng ngày → tốn kém computation (check + train không cần thiết). Dùng Storage sensor thiếu version control, và vẫn gián tiếp (không trigger real-time on change).
📚 Tài liệu tham khảo
- Cloud Build Triggers: docs.cloud.google.com/build/docs/automating-builds/build-triggers (cập nhật 2026: Hỗ trợ Vertex AI training).
- Cloud Source Repositories: cloud.google.com/source-repositories/docs – Tích hợp Git + triggers.
- Vertex AI Pipelines (thay thế AI Platform): cloud.google.com/vertex-ai/docs – Best practices for ML CI/CD.
- Google Cloud ML Certification Guide: Các case study về automating retraining với Cloud Build (Professional ML Engineer exam topics).
Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần thêm ví dụ code, hãy hỏi nhé!
- A Categorical hinge
- B Binary cross-entropy
- C Categorical cross-entropy
- D Sparse categorical cross-entropy
Xem giải thích
🧩 Phân tích nội dung câu hỏi
Câu hỏi thuộc lĩnh vực Machine Learning trên nền tảng đám mây (liên quan đến Google Cloud Vertex AI hoặc TensorFlow/Keras, dù đề cập AWS nhưng nội dung tập trung vào loss functions chuẩn của TensorFlow).
✅ Tình huống: Nhóm của bạn cần xây dựng một mô hình phân loại đa lớp (multi-class classification) để dự đoán xem hình ảnh chứa driver's license (giấy phép lái xe), passport (hộ chiếu), hay credit card (thẻ tín dụng).
📊 Dataset: Đã được pipeline tạo sẵn với 10.000 ảnh driver's license (lớp chính, chiếm đa số), 1.000 ảnh passport, và 1.000 ảnh credit card → Dataset không cân bằng (imbalanced), nhưng điều này không ảnh hưởng trực tiếp đến lựa chọn loss function chính (có thể xử lý bằng class weights).
🔑 Label map: ['drivers_license', 'passport', 'credit_card'] → Đây là 3 lớp độc quyền (mutually exclusive), mỗi ảnh chỉ thuộc một lớp duy nhất (single-label multi-class). Labels thường được mã hóa thành one-hot vectors hoặc integer indices (0,1,2) trong TensorFlow.
🎯 Yêu cầu: Chọn loss function phù hợp để train mô hình (thường dùng với output layer softmax cho xác suất các lớp).
🛠️ Ngữ cảnh cập nhật 2026: Dựa trên TensorFlow 2.16+ / Keras 3.x (tích hợp Vertex AI trên Google Cloud hoặc SageMaker trên AWS), loss functions được tối ưu cho hiệu suất GPU/TPU, khuyến nghị Sparse cho integer labels để tiết kiệm bộ nhớ.
✅ Đáp án đúng: Categorical cross-entropy
Lý do lựa chọn:
- Đây là multi-class classification với one-hot encoded labels (mỗi label là vector xác suất [1,0,0], [0,1,0], [0,0,1]).
CategoricalCrossentropytính toán loss dựa trên phân phối xác suất softmax từ model và one-hot targets, rất phù hợp cho bài toán này khi labels đã được encode đầy đủ.- Trong TensorFlow/Keras (cập nhật 2026), đây là lựa chọn chuẩn cho categorical targets, hỗ trợ tốt imbalance qua
class_weighthoặcsample_weight. - Không dùng cho binary/multi-label, và khớp với label map string-to-one-hot thông qua
tf.keras.utils.to_categorical().
📋 Giải thích chi tiết tất cả các phương án
Dưới đây là phân tích từng lựa chọn giữ nguyên văn bản gốc bằng tiếng Anh, kèm giải thích sai/đúng bằng tiếng Việt. Sử dụng kiến thức TensorFlow/Keras mới nhất (TF 2.16+, Keras 3.x).
-
Categorical hinge ❌ SAI
🧩 Đây là loss cho multi-class SVM-like (hinge loss categorical), phù hợp với linear classifiers hoặc margin-based (như trong tf.nn.softmax_cross_entropy_with_logits nhưng hinge variant). Không dùng cho softmax output chuẩn trong neural nets phân loại xác suất. Với dataset hình ảnh này, nó kém hiệu quả hơn cross-entropy vì không tối ưu hóa xác suất trực tiếp, dễ overfit imbalance. -
Binary cross-entropy ❌ SAI
🛠️ Loss dành cho binary classification hoặc multi-label (mỗi sample có nhiều labels true/false độc lập, ví dụ sigmoid output). Ở đây là single-label multi-class (chỉ 1/3 lớp đúng), nên không phù hợp – sẽ tính loss sai trên 3 outputs, dẫn đến model dự đoán chồng chéo labels. Thường dùng vớifrom_logits=Falsecho sigmoid. -
Categorical cross-entropy ✅ ĐÚNG
📘 Hoàn hảo cho multi-class single-label với one-hot targets (shape: [batch_size, num_classes=3]). Model output logits/softmax (3 dims), loss = -sum(y_true * log(y_pred)). Xử lý tốt imbalance, và là default trong Keras Sequential/Model cho Image classification (như ResNet trên Vertex AI). Khuyến nghị trong docs TF 2026. -
Sparse categorical cross-entropy ❌ SAI
🔢 Dành cho integer-encoded labels (sparse targets: shape [batch_size], giá trị 0/1/2), tự động one-hot hóa bên trong để tiết kiệm memory (tốt cho >100 classes). Tuy phù hợp nếu labels là integers từ label map, nhưng câu hỏi nhấn label map strings ngụ ý one-hot categorical chuẩn; dùng sai format sẽ báo lỗi shape mismatch. Trong exam context (Google ML Engineer), ưu tiên Categorical cho explicit one-hot.
📚 Tài liệu tham khảo
- TensorFlow Docs (cập nhật 2026): tf.keras.losses.CategoricalCrossentropy & SparseCategoricalCrossentropy.
- Google Cloud Vertex AI Guide: Custom Training with Keras (loss cho image classification).
- AWS SageMaker (tương đương): Built-in Algorithms - Image Classification sử dụng cross-entropy tương tự.
- Khuyến nghị xử lý imbalance:
class_weight={0:1.0, 1:10.0, 2:10.0}dựa trên tỷ lệ 10k:1k:1k.
Hy vọng phân tích giúp bạn ôn thi hiệu quả! 🚀 Nếu cần code sample, hỏi thêm nhé!