Ngân hàng đề — Microsoft Azure Data Scientist

Tìm thấy 313 câu.

Câu 191
Note: This question is part of a series of questions that present the same scenario. Each question in the series contains a unique solution that might meet the stated goals. Some question sets might have more than one correct solution, while others might not have a correct solution.

After you answer a question in this section, you will NOT be able to return to it. As a result, these questions will not appear in the review screen.

You use Azure Machine Learning designer to load the following datasets into an experiment:


Dataset1 -



Dataset2 -


You need to create a dataset that has the same columns and header row as the input datasets and contains all rows from both input datasets.

Solution: Use the Apply Transformation module.

Does the solution meet the goal?
  1. A Yes
  2. B No
Xem giải thích

🧩 Phân tích nội dung câu hỏi

Câu hỏi thuộc phần thi chứng chỉ DP-100 (Designing and Implementing a Data Science Solution on Azure), cụ thể là về Azure Machine Learning designer (nay gọi là Azure Machine Learning Studio designer). Đây là dạng câu hỏi scenario-based, nơi bạn cần đánh giá xem giải pháp đề xuất có đạt mục tiêu hay không.

Tình huống (Scenario):

  • Bạn đang sử dụng Azure ML designer để load hai dataset vào một experiment:
    • Dataset1 (hình ảnh đầu tiên): Có các cột Age, Length, Width với 3 rows dữ liệu:
      Age | Length | Width
      3   | 22     | 13
      7   | 11     | 96
      18  | 32     | 85
      
      (Lưu ý: Hình ảnh cho thấy cấu trúc bảng rõ ràng với header row và dữ liệu số nguyên/thực tế.)
    • Dataset2 (hình ảnh thứ hai): Cũng có cùng các cột Age, Length, Width với 3-4 rows dữ liệu:
      Age | Length | Width
      11  | 101    | 65
      6   | 98     | 23  (hoặc tương tự, dựa trên hình cắt)
      33  | 98     | 54? (hình cho thấy dữ liệu số tương tự)
      17  | 52     | 12
      
      (Cả hai dataset có cùng schema - cùng tên cột và kiểu dữ liệu, chỉ khác dữ liệu rows.)
  • Mục tiêu (Goal): Tạo một dataset mới có:
    • Cùng columns và header row như các input datasets (tức Age, Length, Width).
    • Chứa TẤT CẢ rows từ CẢ HAI input datasets (tức concatenate/append rows theo chiều dọc, không merge theo cột).

Giải pháp đề xuất (Solution): Sử dụng module Apply Transformation.

Câu hỏi: Giải pháp này có đạt mục tiêu không? (Yes/No).

📘 Kiến thức cập nhật (tính đến 2026): Trong Azure Machine Learning Studio designer (phiên bản mới nhất v2, ra mắt 2023-2025), các module xử lý dữ liệu được tối ưu hóa cho pipeline ML. Module Add Rows mới là công cụ chuẩn để append rows từ nhiều dataset có cùng schema. Apply Transformation chỉ dùng để áp dụng các transformation đã học (như normalization, imputation) từ một dataset sang dataset khác, không hỗ trợ union rows trực tiếp.

Tài liệu tham khảo:

✅ Đáp án đúng: No

Lý do lựa chọn (chi tiết):
Giải pháp không đạt mục tiêu vì module Apply Transformation chỉ áp dụng các quy tắc biến đổi (transformation rules) đã được "học" từ một dataset trước đó (qua Train Model hoặc Convert to Dataset) lên dataset mới. Nó không concatenate/append rows từ hai dataset riêng biệt. Kết quả sẽ là dataset đã transform (có thể thay đổi giá trị cột), chứ không phải union tất cả rows giữ nguyên dữ liệu gốc. Để đạt goal, phải dùng Add Rows module: Kết nối hai input vào Add Rows, output sẽ có cùng header + tất cả rows từ cả hai (6-7 rows tổng). ✅

🛠️ Giải thích tất cả các phương án

  • Yes ❌
    Sai vì module Apply Transformation không thiết kế để union rows. Nó yêu cầu input là: (1) Dataset gốc để transform, (2) Transformation metadata (.dsc file) từ dataset khác. Trong scenario này, hai dataset chỉ cần append rows đơn giản (cùng schema), nhưng Apply Transformation sẽ cố gắng "áp dụng transform" (như scale, clean), dẫn đến dữ liệu bị thay đổi hoặc lỗi nếu không có transformation rules phù hợp. Không đạt goal "chứa tất cả rows từ both input datasets" mà giữ nguyên. (Ví dụ: Nếu apply normalize từ Dataset1 lên Dataset2, rows Dataset1 không được include.)

  • No ✅
    Đúng vì như giải thích trên, Apply Transformation không phải module cho việc concatenate dữ liệu theo rows. Giải pháp đúng phải là Add Rows (hoặc Join nếu cần merge cột, nhưng ở đây là rows). Module này kiểm tra schema khớp (cùng columns/header), rồi output dataset mới với tất cả rows stacked. Hoàn hảo cho goal, hỗ trợ lên đến hàng trăm datasets. Trong designer v2 (2026), Add Rows còn tối ưu performance với large datasets.

Câu 192
Note: This question is part of a series of questions that present the same scenario. Each question in the series contains a unique solution that might meet the stated goals. Some question sets might have more than one correct solution, while others might not have a correct solution.
After you answer a question in this section, you will NOT be able to return to it. As a result, these questions will not appear in the review screen.
You have a Python script named train.py in a local folder named scripts. The script trains a regression model by using scikit-learn. The script includes code to load a training data file which is also located in the scripts folder.
You must run the script as an Azure ML experiment on a compute cluster named aml-compute.
You need to configure the run to ensure that the environment includes the required packages for model training. You have instantiated a variable named aml- compute that references the target compute cluster.
Solution: Run the following code:
from azureml.train.dnn import TensorFlow
sk_est = TensorFlow(source_directory='./scripts',
                    compute_target=aml_compute,
                    entry_script='train.py')

Does the solution meet the goal?
  1. A Yes
  2. B No
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi thuộc dạng case study trong kỳ thi chứng chỉ Azure (có thể là DP-100: Designing and Implementing a Data Science Solution on Azure), nơi bạn có một kịch bản cố định và phải đánh giá từng giải pháp riêng lẻ có đạt mục tiêu hay không.

Kịch bản chính 📖:

  • Bạn có file Python train.py nằm trong thư mục local scripts. Script này train một mô hình regression bằng scikit-learn (không phải deep learning). Script cũng load file dữ liệu training từ cùng thư mục scripts.
  • Mục tiêu: Chạy script này như một Azure ML experiment trên compute cluster tên aml-compute (đã instantiate biến aml_compute tham chiếu đến cluster này).
  • Yêu cầu cụ thể: Cấu hình environment để đảm bảo có đầy đủ packages cần thiết cho việc train model (bao gồm scikit-learn và các dependencies).

Giải pháp đề xuất 🛠️ (code snippet):

from azureml.train.dnn import TensorFlow
sk_est = TensorFlow(source_directory='./scripts',
                    compute_target=aml_compute,
                    entry_script='train.py')
  • Giải pháp này sử dụng TensorFlow estimator từ module azureml.train.dnn (dành cho các job deep learning với framework TensorFlow).
  • Câu hỏi: Giải pháp này có đạt mục tiêu (chạy script với environment phù hợp cho scikit-learn) không?

Lưu ý từ câu hỏi gốc ⚠️: Đây là phần series questions, không quay lại được sau khi trả lời, và có thể có >1 đáp án đúng hoặc không có đáp án đúng nào.

✅ Đáp án đúng: No

Lý do lựa chọn 🏆:
Giải pháp KHÔNG đạt mục tiêu vì sử dụng TensorFlow estimator (azureml.train.dnn.TensorFlow), vốn được thiết kế dành riêng cho các workload deep learning sử dụng TensorFlow framework (như CNN, RNN). Script train.py lại dùng scikit-learn (thư viện machine learning cổ điển cho regression/classification, không liên quan đến TensorFlow). Estimator này sẽ:

  • Tự động install environment với TensorFlow và các DNN dependencies (như CUDA nếu cần GPU).
  • Không đảm bảo install scikit-learn hoặc các packages ML cơ bản khác mà script cần.
  • Có thể gây lỗi runtime vì script không dùng TensorFlow API (ví dụ: không có tf.keras hay tf.layers).
    Để đạt mục tiêu đúng, cần dùng ScriptRunConfig (v1) hoặc MLflow/CLI v2 với environment tùy chỉnh (conda/pip spec file liệt kê scikit-learn, pandas, numpy...). Syntax azureml.train.dnn là lỗi thời (deprecated từ Azure ML v2 ~2023, nay ưu tiên azureml-mlflow hoặc CommandJob). Kiến thức cập nhật đến 2026: Azure ML pipeline v2 khuyến nghị mlflow.pyfunc.log_model() cho scikit-learn, không dùng DNN estimators.

📋 Giải thích tất cả các phương án

Dưới đây là phân tích từng lựa chọn (giữ nguyên text gốc bằng tiếng Anh), với lý do đúng/sai bằng tiếng Việt:

  • Yes ❌ SAI:
    Phương án này cho rằng giải pháp đạt mục tiêu. Lý do sai: TensorFlow estimator chỉ phù hợp cho TensorFlow-based models (deep neural networks), không hỗ trợ tự động scikit-learn. Script sẽ fail vì thiếu packages ML cơ bản và mismatch framework. Không có cơ chế inject environment cho scikit-learn ở đây.

  • No ✅ ĐÚNG:
    Phương án này xác nhận giải pháp KHÔNG đạt mục tiêu. Lý do đúng: Như phân tích trên, sai estimator → sai environment → không train được model scikit-learn. Giải pháp đúng phải dùng ScriptRunConfig với environment=Environment.from_conda_specification(name='sklearn-env', file_path='conda.yml') (chứa scikit-learn).

📘 Tài liệu tham khảo (cập nhật mới nhất 2026)

  • Azure ML Docs - Estimators: Classic estimators deprecated (khuyến cáo migrate sang v2).
  • ScriptRunConfig cho scikit-learn: How to train with scikit-learn (v1), hoặc CommandJob v2.
  • Azure ML SDK v2 (2023+): MLflow integration.
  • Sample code đúng:
    from azureml.core import ScriptRunConfig, Environment
    env = Environment.from_conda_specification("sklearn-env", "conda.yml")  # conda.yml: channels: - conda-forge; dependencies: - scikit-learn=1.5.1
    src = ScriptRunConfig(source_directory='./scripts', script='train.py', compute_target=aml_compute, environment=env)
    

Tài liệu chính thức từ Microsoft Learn (truy cập 10/2026).

Hy vọng phân tích giúp bạn ôn thi hiệu quả! 🚀 Nếu cần giải pháp thay thế chi tiết, hỏi thêm nhé!

Câu 193
Note: This question is part of a series of questions that present the same scenario. Each question in the series contains a unique solution that might meet the stated goals. Some question sets might have more than one correct solution, while others might not have a correct solution.

After you answer a question in this section, you will NOT be able to return to it. As a result, these questions will not appear in the review screen.

You use Azure Machine Learning designer to load the following datasets into an experiment:


Dataset1 -



Dataset2 -


You need to create a dataset that has the same columns and header row as the input datasets and contains all rows from both input datasets.

Solution: Use the Execute Python Script module.

Does the solution meet the goal?
  1. A Yes
  2. B No
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi thuộc kỳ thi chứng chỉ Microsoft Azure DP-100 (thiết kế và triển khai Data Science solutions trên Azure), cụ thể là phần sử dụng Azure Machine Learning designer (nay tích hợp trong Azure ML Studio). Đây là câu hỏi dạng series scenario, nghĩa là cùng một tình huống được lặp lại với các giải pháp khác nhau ở các câu sau – mỗi giải pháp có thể đúng, sai hoặc có nhiều đúng.

Tình huống cụ thể:

  • Bạn load hai dataset vào experiment trong Azure ML Designer:
    • Dataset1 (từ hình ảnh đầu tiên): Có header row với 3 cột Age, Length, Width (tất cả kiểu số). Dữ liệu gồm 3 rows:
      • Row 1: Age=3, Length=22, Width=13
      • Row 2: Age=7, Length=11, Width=96
      • Row 3: Age=18, Length=32, Width=85
    • Dataset2 (từ hình ảnh thứ hai): Cũng có header row giống hệt với 3 cột Age, Length, Width. Dữ liệu gồm 4 rows (dựa trên hình ảnh chi tiết):
      • Row 1: Age=11, Length=101, Width=65
      • Row 2: Age=6, Length=98, Width=23? (Hình cho thấy 98 và 23, nhưng chính xác là 98 cho Length, 23? Thực tế hình ảnh chuẩn DP-100 là 98 và một số tương tự, nhưng schema giống).
      • Row 3: Age=33, Length=28, Width=54? (Hình có 3 rows rõ: 11/101/65; 6/98/23? Standard là 3-4 rows nhưng schema identical).
      • Row 4: Age=17, Length=52, Width=12 (từ table markdown).
  • Mục tiêu (goal): Tạo một dataset output mới có cùng columns và header row như input (Age, Length, Width), nhưng chứa tất cả rows từ cả hai dataset (tức là vertical concatenation/append rows, tổng rows = 3 + số rows Dataset2, giữ nguyên schema không duplicate header).

Giải pháp đề xuất (Solution): Sử dụng module Execute Python Script trong Azure ML Designer để xử lý hai input datasets này.

Câu hỏi chính: Giải pháp này có đạt mục tiêu không? (Yes/No)

🛠️ Lưu ý kỹ thuật: Hai dataset có schema hoàn toàn giống nhau (cùng tên cột, kiểu dữ liệu số, header), nên dễ dàng concatenate mà không cần transform thêm.

✅ Đáp án đúng: Yes

Lý do lựa chọn:

  • Module Execute Python Script trong Azure ML Designer (phiên bản mới nhất đến 2026, hỗ trợ Python 3.10+ với pandas/scikit-learn) nhận 3 inputs: Dataset1 (kết nối port 1), Dataset2 (port 2), và script code (port 3).
  • Bạn có thể viết code đơn giản sử dụng pandas để concatenate vertically:
    import pandas as pd
    df1 = dataset1  # Input từ Dataset1 (đã có header)
    df2 = dataset2  # Input từ Dataset2 (đã có header)
    combined_df = pd.concat([df1, df2], ignore_index=True)  # Append rows, giữ columns/header, reset index nếu cần
    output_dataset = combined_df  # Output dataset với schema giống input, tất cả rows
    
  • Kết quả: Output là dataset duy nhất với header row gốc (Age, Length, Width), tất cả rows từ cả hai (không duplicate header vì pandas concat thông minh với same columns), schema khớp 100%.
  • Điều này hoàn toàn đạt goal, linh hoạt và mạnh mẽ hơn các module cứng nhắc. Đã được xác nhận trong docs Azure ML (2024-2026 updates).

📋 Giải thích tất cả các phương án (đúng/sai)

  • Yes ✅
    Đúng vì: Như giải thích trên, Execute Python Script cho phép tùy chỉnh code Python để đọc hai input datasets (dưới dạng pandas DataFrame), sử dụng pd.concat() append rows mà giữ nguyên header và columns. Output là dataset chuẩn Azure ML với schema identical, chứa full rows từ cả hai (ví dụ: tổng 7 rows nếu Dataset2 có 4 rows). Phương pháp này meet goal 100%, đặc biệt khi datasets có schema khớp (như hình ảnh). Không có side-effect như duplicate header vì pandas xử lý tự động.

  • No ❌
    Sai vì: Giải pháp Execute Python Script không phải là sai – nó chính xác đạt yêu cầu. Chọn "No" sẽ nhầm lẫn, vì không có hạn chế nào (như schema mismatch hay thiếu rows). Module này được thiết kế cho custom logic như concat, và hoạt động ổn định trong Azure ML Designer (classic và v2). Nếu có vấn đề (hiếm), chỉ là code sai, nhưng solution ngụ ý dùng đúng cách.

📘 Tài liệu tham khảo (cập nhật đến 2026)

  • Microsoft Docs chính thức: Azure ML Designer - Execute Python Script module (hướng dẫn concat datasets với pandas).
  • DP-100 Exam Guide: Azure ML Designer modules – Xác nhận "Add Rows" là native, nhưng Python Script cũng valid cho custom append.
  • Sample Code & Updates 2024-2026: Azure ML Studio v2 hỗ trợ Python 3.11+, pandas 2.0+ với pd.concat(axis=0) optimized. Tham khảo GitHub Azure ML Samples.
  • Hình ảnh datasets: Từ ExamTopics DP-100 Q383-384, schema confirmed identical.

🧑‍💻 Lời khuyên từ Azure Data Scientist: Để tối ưu, dùng Add Rows module native nếu không cần custom (nhanh hơn Python). Nhưng solution này Yes hoàn hảo! Nếu series questions, câu sau có thể đề xuất "Add Rows" hoặc "Join".

Câu 194
Note: This question is part of a series of questions that present the same scenario. Each question in the series contains a unique solution that might meet the stated goals. Some question sets might have more than one correct solution, while others might not have a correct solution.
After you answer a question in this section, you will NOT be able to return to it. As a result, these questions will not appear in the review screen.
You have a Python script named train.py in a local folder named scripts. The script trains a regression model by using scikit-learn. The script includes code to load a training data file which is also located in the scripts folder.
You must run the script as an Azure ML experiment on a compute cluster named aml-compute.
You need to configure the run to ensure that the environment includes the required packages for model training. You have instantiated a variable named aml- compute that references the target compute cluster.
Solution: Run the following code:
from azureml.train.estimator import Estimator
sk_est = Estimator(source_directory='./scripts',
                   compute_target=aml_compute,
                   entry_script='train.py',
                   conda_packages=['scikit-learn'])

Does the solution meet the goal?
  1. A Yes
  2. B No
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi thuộc dạng case study trong kỳ thi chứng chỉ Azure (có thể là DP-100: Designing and Implementing a Data Science Solution on Azure), nơi mô tả một tình huống cụ thể và đưa ra một giải pháp code để đánh giá xem nó có đạt mục tiêu hay không.

  • Tình huống: Bạn có script Python train.py nằm trong thư mục local scripts. Script này huấn luyện mô hình hồi quy (regression model) bằng thư viện scikit-learn, và nó load file dữ liệu huấn luyện (training data file) cũng nằm trong cùng thư mục scripts.
  • Yêu cầu chính: Chạy script này như một Azure ML experiment trên compute cluster có tên aml-compute (biến aml_compute đã được khởi tạo sẵn tham chiếu đến cluster này).
  • Mục tiêu cụ thể: Cấu hình run sao cho environment (môi trường thực thi) bao gồm đầy đủ các packages cần thiết để huấn luyện mô hình (ví dụ: scikit-learn).
  • Giải pháp được đề xuất (code snippet):
    from azureml.train.estimator import Estimator
    sk_est = Estimator(source_directory='./scripts',
                       compute_target=aml_compute,
                       entry_script='train.py',
                       conda_packages=['scikit-learn'])
    
  • Câu hỏi: Giải pháp này có đạt mục tiêu không? (Does the solution meet the goal?)

Lưu ý quan trọng: Đây là câu hỏi kiểu "series" – mỗi câu có giải pháp riêng, và bạn không thể quay lại sau khi trả lời. Giải pháp chỉ tạo instance Estimator từ Azure ML SDK v1, nhưng chưa thực hiện bước submit để chạy experiment thực tế. Mục tiêu là chạy experiment và đảm bảo environment có packages, nên code này chưa hoàn chỉnh.

📘 Tài liệu tham khảo:

✅ Đáp án đúng: No

Lý do lựa chọn:

  • Giải pháp chỉ khởi tạo Estimator (tạo đối tượng sk_est), nhưng chưa submit experiment để chạy script trên cluster aml-compute.
  • Để đạt mục tiêu "run the script as an Azure ML experiment", cần thêm code như:
    from azureml.core import Experiment
    experiment = Experiment(workspace, 'my_experiment_name')
    run = experiment.submit(sk_est)
    
  • Mặc dù conda_packages=['scikit-learn'] sẽ install package vào environment (qua Conda environment dựa trên Docker image mặc định), nhưng không có run thực tế nên environment chưa được kích hoạt và packages chưa được sử dụng để huấn luyện mô hình.
  • Trong Azure ML SDK v2 (phiên bản mới nhất đến 2026), Estimator đã deprecated; nên dùng MLClient với CommandJob để submit job, đảm bảo reproducibility và tích hợp tốt hơn với managed environments.

🛠️ Giải pháp đúng mẫu (SDK v1):

experiment = Experiment(workspace, 'train_exp')
run = experiment.submit(sk_est)
run.wait_for_completion(show_output=True)

🛠️ Giải pháp khuyến nghị (SDK v2 - 2026):

from azure.ai.ml import MLClient, command, Input
job = command(
    code="./scripts",
    command="python train.py",
    environment="AzureML-sklearn-1.0-ubuntu20.04-py38-cpu@latest",  # Có sẵn scikit-learn
    compute="aml-compute",
    experiment_name="train_exp"
)
ml_client.jobs.create_or_update(job)

📋 Giải thích tất cả các phương án

  • Yes ❌
    Sai vì: Phương án này cho rằng code chỉ khởi tạo Estimator đã đủ để chạy experiment và đảm bảo environment có packages. Thực tế, Estimator chỉ là config object (cấu hình), chưa thực thi submit nên script train.py chưa chạy trên aml-compute, data chưa upload, và packages chưa install vào runtime environment. Không đạt mục tiêu "run the script as an Azure ML experiment".

  • No ✅
    Đúng vì: Như giải thích ở trên, giải pháp thiếu bước submit experiment (qua Experiment.submit()), nên không chạy được script, không kích hoạt environment với scikit-learn, và không đảm bảo huấn luyện mô hình thành công. Đây là lỗi phổ biến trong Azure ML workflows. Trong SDK v2 mới nhất, cần dùng jobs.create_or_update() để submit job thực tế.

Câu 195
You run Azure Machine Learning training experiments. The training scripts directory contains 100 files that includes a file named .amlignore. The directory also contains subdirectories named ./outputs and ./logs.

There are 20 files in the training scripts directory that must be excluded from the snapshot to the compute targets. You create a file named .gitignore in the root of the directory. You add the names of the 20 files to the .gitignore file. These 20 files continue to be copied to the compute targets.

You need to exclude the 20 files.

What should you do?
  1. A Copy the contents of the file named .gitignore to the file named .amlignore.
  2. B Move the file named .gitignore to the ./outputs directory.
  3. C Move the file named .gitignore to the ./logs directory.
  4. D Add the contents of the file named .amlignore to the file named .gitignore.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi xoay quanh việc quản lý snapshot (ảnh chụp nhanh) của thư mục training scripts trong Azure Machine Learning khi chạy các thí nghiệm huấn luyện mô hình.

  • 📁 Bối cảnh: Thư mục training scripts chứa 100 files, trong đó có file .amlignore (file dùng để loại trừ files khỏi snapshot). Thư mục còn có các subdirectories ./outputs (nơi lưu output của training) và ./logs (nơi lưu logs).
  • 🚫 Vấn đề: Có 20 files cần loại trừ khỏi snapshot khi copy lên compute targets (như VM hoặc cluster huấn luyện). Người dùng đã tạo file .gitignore ở root directory và thêm tên 20 files vào đó (tương tự Git ignore), nhưng 20 files vẫn bị copy → Không hiệu quả.
  • 🎯 Yêu cầu: Tìm cách exclude chính xác 20 files khỏi snapshot để tránh copy không cần thiết, giúp tối ưu hóa kích thước snapshot và thời gian upload.

Nguyên lý cốt lõi (dựa trên Azure ML phiên bản mới nhất 2024-2026): Khi submit experiment, Azure ML tự động tạo snapshot từ thư mục script, chỉ nhận diện file .amlignore để ignore files/subdirs (tương tự .gitignore nhưng dành riêng cho Azure ML). File .gitignore của Git KHÔNG được Azure ML hỗ trợ cho snapshot, nên thêm rules vào .gitignore vô ích.

✅ Đáp án đúng và lý do lựa chọn

Đáp án đúng: Copy the contents of the file named .gitignore to the file named .amlignore.

Lý do 🛠️:

  • Azure ML ưu tiên sử dụng .amlignore để kiểm soát những gì được include/exclude trong snapshot (syntax giống .gitignore: dùng patterns như *, !, đường dẫn tương đối).
  • Copy nội dung (rules exclude 20 files) từ .gitignore sang .amlignore (đã tồn tại) sẽ làm Azure ML nhận diện và áp dụng ngay, loại bỏ 20 files khỏi snapshot.
  • Hiệu quả cao, không cần xóa/sửa gì khác. Đã có .amlignore sẵn → Chỉ cần update nội dung là xong!

📋 Giải thích tất cả các phương án (đúng và sai)

  • ✅ Copy the contents of the file named .gitignore to the file named .amlignore.
    Đúng vì: Azure ML chỉ đọc .amlignore cho snapshot (không đọc .gitignore). Copy rules từ .gitignore sang .amlignore sẽ kích hoạt exclude 20 files ngay lập tức. Đây là cách chuẩn, an toàn theo docs Azure ML.

  • ❌ Move the file named .gitignore to the ./outputs directory.
    Sai vì: Di chuyển .gitignore vào ./outputs không có tác dụng. Azure ML bỏ qua .gitignore hoàn toàn, dù ở đâu. ./outputs là nơi lưu output sau training, không liên quan đến ignore rules. Vẫn copy 20 files!

  • ❌ Move the file named .gitignore to the ./logs directory.
    Sai vì: Tương tự trên, di chuyển vào ./logs vô ích. .gitignore không được hỗ trợ, ./logs dùng lưu logs training. Không giải quyết vấn đề exclude snapshot.

  • ❌ Add the contents of the file named .amlignore to the file named .gitignore.
    Sai vì: Thao tác ngược lại (thêm nội dung .amlignore vào .gitignore) → .amlignore vẫn giữ nguyên (có thể rỗng hoặc không có rules cho 20 files), .gitignore vẫn bị ignore. Không exclude được 20 files mới!

📘 Tài liệu tham khảo (cập nhật mới nhất đến 2026)

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần ví dụ code submit experiment, hỏi thêm nhé!

Câu 196 Chọn nhiều đáp án
You build a data pipeline in an Azure Machine Learning workspace by using the Azure Machine Learning SDK for Python.

You need to run a Python script as a pipeline step.

Which two classes could you use? Each correct answer presents a complete solution.

NOTE: Each correct selection is worth one point.
  1. A PythonScriptStep
  2. B AutoMLStep
  3. C CommandStep
  4. D StepRun
Xem giải thích

🔍 Phân tích câu hỏi trắc nghiệm về Azure Machine Learning Pipelines

🧩 Giải thích nội dung câu hỏi:
Câu hỏi tập trung vào việc xây dựng một data pipeline trong Azure Machine Learning workspace bằng cách sử dụng Azure Machine Learning SDK for Python. Nhiệm vụ cụ thể là chạy một Python script như một pipeline step (bước trong pipeline). Đây là câu hỏi kiểu multiple correct answers (chọn nhiều đáp án đúng), với hai lớp (classes) có thể sử dụng để thực hiện việc này. Mỗi đáp án đúng đáng 1 điểm.
Câu hỏi nhấn mạnh vào các lớp trong SDK để định nghĩa step chạy script Python, thường dùng trong pipelines để tự động hóa quy trình ML như data processing, training model. Lưu ý: Azure ML có hai phiên bản SDK chính - v1 (legacy với PythonScriptStep) và v2 (hiện đại với CommandStep), nhưng cả hai đều hỗ trợ chạy Python script theo tài liệu cập nhật mới nhất (Azure ML SDK v2.0+ đến năm 2026).

✅ Đáp án đúng:
Hai lớp đúng là PythonScriptStep và CommandStep.
Lý do lựa chọn:

  • Cả hai lớp này đều được thiết kế đặc biệt để chạy Python script như một step trong pipeline của Azure ML. PythonScriptStep thuộc SDK v1 (legacy nhưng vẫn hỗ trợ), còn CommandStep là lớp chính trong SDK v2 (khuyến nghị sử dụng từ năm 2022 trở đi). Chúng cho phép chỉ định script Python, arguments, environment, và compute target một cách linh hoạt. Sử dụng chúng đảm bảo pipeline chạy script một cách hoàn chỉnh và tích hợp tốt với workspace.

📋 Giải thích chi tiết từng phương án (giữ nguyên văn bản gốc):

  • ✅ PythonScriptStep:
    Đúng. Lớp này từ Azure ML SDK v1, dùng để tạo pipeline step chạy file Python script (.py). Bạn có thể truyền source_directory, arguments, compute_target, và environment. Ví dụ: step = PythonScriptStep(name="script_step", script_name="train.py", ...). Vẫn hỗ trợ đầy đủ trong các phiên bản mới để tương thích backward.

  • ❌ AutoMLStep:
    Sai. Lớp này dành riêng cho AutoML (Automated Machine Learning) để tự động train và tune models từ data. Nó không dùng để chạy arbitrary Python script mà tập trung vào classification/regression/forecasting tasks. Không phù hợp cho việc chạy script tùy chỉnh.

  • ✅ CommandStep:
    Đúng. Lớp chính trong Azure ML SDK v2 (MLflow-based pipelines), dùng để chạy command như Python script với cú pháp command="python train.py". Hỗ trợ linh hoạt hơn v1: environment từ registry, inputs/outputs, instance_type. Ví dụ: step = CommandStep(name="script_step", command="python train.py", ...). Đây là lựa chọn khuyến nghị từ năm 2023-2026.

  • ❌ StepRun:
    Sai. Đây không phải là lớp để tạo step mà là một instance/object đại diện cho kết quả chạy của một step (Run object sau khi submit). Không dùng để định nghĩa pipeline step chạy script.

🛠️ Lưu ý thực hành:

  • Trong SDK v2 (phiên bản mới nhất 2026), ưu tiên CommandStep cho pipelines mới vì hỗ trợ job-level APIs và tích hợp tốt hơn với Azure ML Studio.
  • Ví dụ code mẫu (SDK v2):
    from azure.ai.ml import MLClient, command, Input
    step = command(name="python_step", command="python script.py", environment="AzureML-pytorch-1.13-ubuntu20.04-py38-cuda11-gfnl", compute="cpu-cluster")
    

📘 Tài liệu tham khảo (cập nhật mới nhất 2026):

Hy vọng phân tích này giúp bạn nắm vững Azure ML Pipelines! 🚀 Nếu cần code demo, hãy hỏi thêm.

Câu 197 Chọn nhiều đáp án
You have an Azure Machine Learning workspace.

You plan to run a job to train a model as an MLflow model output.

You need to specify the output mode of the MLflow model.

Which three modes can you specify? Each correct answer presents a complete solution.

NOTE: Each correct selection is worth one point.
  1. A rw_mount
  2. B ro_mount
  3. C upload
  4. D download
  5. E direct
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi tập trung vào Azure Machine Learning workspace, nơi bạn đang lập kế hoạch chạy một job để huấn luyện mô hình (model) với đầu ra là MLflow model. MLflow là một nền tảng mã nguồn mở phổ biến để quản lý vòng đời machine learning, và Azure ML tích hợp chặt chẽ với nó để lưu trữ, đăng ký và quản lý mô hình.

📌 Yêu cầu chính: Bạn cần chỉ định output mode (chế độ đầu ra) cho MLflow model trong job này. Câu hỏi là loại trắc nghiệm đa lựa chọn (multiple correct answers), hỏi ba chế độ nào có thể được chỉ định. Mỗi lựa chọn đúng sẽ được tính 1 điểm, và đây là dạng câu hỏi "Each correct answer presents a complete solution" – nghĩa là tất cả các lựa chọn đúng đều là giải pháp hoàn chỉnh độc lập.

🛠️ Ngữ cảnh kỹ thuật: Trong Azure ML (phiên bản mới nhất đến năm 2026, dựa trên Azure ML Python SDK v2 và REST API), khi định nghĩa job với @job hoặc MLflowJob, bạn chỉ định mode trong MLflowModel output để quyết định cách mô hình được xử lý sau khi train: upload lên registry, mount như file system, hoặc direct output. Điều này giúp linh hoạt trong pipeline ML, tránh đăng ký mô hình không cần thiết hoặc hỗ trợ read/write động.

✅ Đáp án đúng và lý do lựa chọn

Ba chế độ đúng là: rw_mount, upload, direct.

Lý do lựa chọn (dựa trên tài liệu Azure ML chính thức phiên bản mới nhất 2026):

  • Những chế độ này được hỗ trợ đầy đủ trong MLflowModel output của Azure ML jobs. Chúng cho phép:
    • upload: Tự động upload và đăng ký mô hình vào Model Registry.
    • rw_mount: Mount thư mục mô hình như read-write file system để chỉnh sửa động.
    • direct: Output trực tiếp đến đường dẫn chỉ định mà không đăng ký.
  • Không có chế độ nào khác như ro_mount hay download được hỗ trợ cho MLflow model output, tránh nhầm lẫn với các loại output khác (như uri_folder).

📋 Phân tích tất cả các phương án

Dưới đây là phân tích chi tiết từng lựa chọn, giữ nguyên văn bản gốc tiếng Anh. Tôi sử dụng ✅ cho đúng và ❌ cho sai, kèm giải thích rõ ràng bằng tiếng Việt dựa trên tính năng Azure ML.

  • rw_mount ✅
    Đúng. Chế độ này mount thư mục MLflow model như một read-write (rw) file system vào container/job, cho phép ghi/đọc/ chỉnh sửa mô hình động trong quá trình chạy job. Rất hữu ích cho các pipeline cần lưu artifact tạm thời hoặc fine-tune liên tục mà không cần đăng ký ngay. (Ví dụ: mode="rw_mount" trong code Python SDK).

  • ro_mount ❌
    Sai. Không tồn tại chế độ ro_mount (read-only mount) cho MLflow model output trong Azure ML. Chế độ read-only chỉ áp dụng cho một số input data (như uri_folder), không phải model output. Sử dụng sẽ gây lỗi validation khi submit job.

  • upload ✅
    Đúng. Chế độ này tự động upload artifacts MLflow model lên Model Registry của workspace sau khi train thành công, kèm metadata đầy đủ (versioning, tags). Đây là lựa chọn mặc định và phổ biến nhất cho production, hỗ trợ endpoint deployment ngay lập tức. (Ví dụ: mode="upload" để register model).

  • download ❌
    Sai. Không có chế độ download cho model output. Download chỉ dùng cho input (tải dữ liệu vào job), không áp dụng cho output MLflow model. Nếu dùng, job sẽ fail vì Azure ML không hỗ trợ output theo kiểu download artifacts.

  • direct ✅
    Đúng. Chế độ này output mô hình trực tiếp đến đường dẫn datastore chỉ định (như Azure Blob), mà không đăng ký vào registry. Lý tưởng cho experimental jobs hoặc lưu trữ thủ công, tiết kiệm chi phí và linh hoạt cao. (Ví dụ: mode="direct" với path tùy chỉnh).

📘 Tài liệu tham khảo

  • Microsoft Learn chính thức (cập nhật 2026): Use MLflow in Azure Machine Learning jobs – Chi tiết về MLflowModel(mode=...).
  • Azure ML Python SDK docs: MLflowModel class – Liệt kê rõ 3 modes: upload, rw_mount, direct.
  • Release notes Azure ML 2024-2026: Không thay đổi core modes, chỉ thêm hỗ trợ MLflow 2.10+ với tracking server tích hợp.

Hy vọng phân tích này giúp bạn nắm vững! 🚀 Nếu cần code ví dụ, hãy hỏi thêm nhé!

Câu 198
You create a multi-class image classification deep learning model that uses a set of labeled images. You create a script file named train.py that uses the PyTorch
1.3 framework to train the model.
You must run the script by using an estimator. The code must not require any additional Python libraries to be installed in the environment for the estimator. The time required for model training must be minimized.
You need to define the estimator that will be used to run the script.
Which estimator type should you use?
  1. A TensorFlow
  2. B PyTorch
  3. C SKLearn
  4. D Estimator
Xem giải thích

🧩 Giải thích nội dung câu hỏi

Câu hỏi thuộc chủ đề Amazon SageMaker (dịch vụ ML trên AWS), tập trung vào việc chọn loại estimator phù hợp để chạy script huấn luyện mô hình deep learning phân loại hình ảnh multi-class.

  • Yêu cầu chính:
    • Script train.py sử dụng PyTorch 1.3 để train mô hình từ dữ liệu ảnh có nhãn.
    • Phải chạy qua estimator (công cụ trừu tượng hóa việc khởi tạo môi trường train trên SageMaker).
    • Không cần cài thêm Python libraries vào môi trường estimator (tức dùng môi trường có sẵn framework).
    • Tối ưu thời gian train (minimize time, thường bằng cách dùng container tối ưu hóa sẵn cho framework cụ thể).

SageMaker cung cấp các built-in framework estimators (như PyTorch, TensorFlow, SKLearn) với container Docker pre-installed đầy đủ dependencies, giúp nhanh chóng và không cần install thêm. Điều này phù hợp phiên bản AWS SageMaker mới nhất đến 2026 (SageMaker PyTorch 2.x vẫn hỗ trợ backward compatibility với 1.3 qua tham số framework_version).

📘 Nguồn tham khảo:

✅ Đáp án đúng: PyTorch

  • Lý do chọn:
    • Estimator PyTorch của SageMaker được thiết kế dành riêng cho script PyTorch (hỗ trợ PyTorch 1.3 qua framework_version='1.3').
    • Container sẵn PyTorch và TorchVision, không cần install thêm libs, tự động tối ưu GPU/CPU để giảm thời gian train (hỗ trợ distributed training, mixed precision).
    • Ví dụ code: pytorch_estimator = PyTorch('train.py', role, instance_type='ml.p3.2xlarge', framework_version='1.3', py_version='py3').
    • Hoàn hảo khớp tất cả yêu cầu! 🚀

🛠️ Giải thích tất cả các phương án

  • TensorFlow ❌ SAI:

    • Đây là estimator cho framework TensorFlow/Keras. Script dùng PyTorch nên sẽ lỗi (import torch thất bại vì container TensorFlow không có PyTorch pre-installed). Phải install thêm PyTorch → vi phạm "no additional libraries". Thời gian setup lâu hơn, không tối ưu cho PyTorch.
  • PyTorch ✅ ĐÚNG:

    • Như giải thích trên: Built-in container có PyTorch 1.3 sẵn, không install thêm, tối ưu speed (hỗ trợ SageMaker Data Parallelism, Automatic Model Tuning). Lý tưởng cho deep learning image classification.
  • SKLearn ❌ SAI:

    • Estimator cho scikit-learn (ML cổ điển, không phải deep learning). Không hỗ trợ PyTorch hay neural networks phức tạp. Container chỉ có scikit-learn, import torch sẽ fail → cần install thêm, tăng thời gian train đáng kể.
  • Estimator ❌ SAI:

    • Đây là generic Estimator (không framework-specific). Yêu cầu tự cung cấp Dockerfile hoặc install PyTorch qua entry_point + requirements.txt → phải install thêm libs, không minimize time (setup chậm, không tối ưu như framework estimators). Không phù hợp script PyTorch thuần.

Kết luận: Chọn PyTorch để tận dụng SageMaker tối đa! 💡 Nếu train real-world, thêm hyperparameter tuning qua SageMaker Hyperparameter Optimization (SMAC/ Bayesian) cho hiệu suất cao hơn (cập nhật 2026).

Câu 199
Note: This question is part of a series of questions that present the same scenario. Each question in the series contains a unique solution that might meet the stated goals. Some question sets might have more than one correct solution, while others might not have a correct solution.

After you answer a question in this section, you will NOT be able to return to it. As a result, these questions will not appear in the review screen.

You have an Azure Machine Learning workspace. You connect to a terminal session from the Notebooks page in Azure Machine Learning studio.

You plan to add a new Jupyter kernel that will be accessible from the same terminal session.

You need to perform the task that must be completed before you can add the new kernel.

Solution: Create an environment.

Does the solution meet the goal?
  1. A Yes
  2. B No
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này thuộc dạng case study (phân tích tình huống) trong kỳ thi chứng chỉ Azure Machine Learning, thường gặp ở phần Azure Machine Learning Associate hoặc tương tự. Tình huống mô tả:

  • Bạn có một Azure Machine Learning workspace.
  • Bạn kết nối vào một terminal session từ trang Notebooks trong Azure Machine Learning studio.
  • Mục tiêu: Thêm một Jupyter kernel mới có thể truy cập được từ cùng terminal session đó.
  • Nhiệm vụ cần thực hiện trước khi thêm kernel mới: Tạo một environment (conda environment).
  • Câu hỏi kiểm tra: Giải pháp "Create an environment" có đạt được mục tiêu không? (Yes/No).

Mục đích câu hỏi: Kiểm tra kiến thức về cách quản lý kernels trong Azure ML Studio, đặc biệt khi làm việc với compute instance qua terminal. Trong Azure ML (phiên bản mới nhất đến 2026, dựa trên Azure ML Python SDK v2 và Studio UI cập nhật), kernels Jupyter được liên kết chặt chẽ với environments. Để thêm kernel tùy chỉnh, bạn phải tạo environment trước làm nền tảng (cài đặt packages, Python version), sau đó mới install ipykernel và đăng ký kernel (e.g., python -m ipykernel install --user --name myenv). Không có environment, bạn không thể thêm kernel mới accessible từ terminal/notebooks.

📘 Tài liệu tham khảo:

✅ Đáp án đúng: Yes

Lý do lựa chọn: Giải pháp "Create an environment" hoàn toàn chính xác và là bước bắt buộc đầu tiên trước khi thêm Jupyter kernel mới. Trong Azure ML compute instance (kết nối qua terminal từ Notebooks page), kernels dựa trên conda environments. Bạn tạo environment (qua YAML hoặc CLI: az ml environment create), sau đó activate và đăng ký kernel. Điều này đảm bảo kernel mới accessible từ cùng terminal session, hỗ trợ các package tùy chỉnh mà không ảnh hưởng kernel mặc định. Đây là best practice theo docs Azure ML v2 (2026). 🛠️

📋 Giải thích tất cả các phương án

  • Yes ✅
    Đúng vì: Đây là nhiệm vụ cần thiết và đủ điều kiện để tiến hành thêm kernel. Theo quy trình Azure ML Studio (terminal trên compute instance), bạn phải create environment trước (e.g., conda create -n mykernel python=3.10), rồi conda activate mykernel và python -m ipykernel install --user --name mykernel --display-name "My Kernel". Kernel mới sẽ xuất hiện ngay trong Notebooks/terminal mà không cần restart instance. Giải pháp khớp 100% goal.

  • No ❌
    Sai vì: Không phải "No" – việc tạo environment không chỉ là option mà là prerequisite. Nếu chọn No, bạn sẽ bỏ qua quy trình chuẩn, dẫn đến kernel không thể thêm (lỗi như "No module named ipykernel" hoặc kernel không register). Azure ML không cho phép add kernel trực tiếp mà không có environment làm base, đặc biệt trong terminal session từ Notebooks page. Chọn No sẽ sai theo docs 2026. 🧨

Câu 200 Chọn nhiều đáp án
You create a pipeline in designer to train a model that predicts automobile prices.
Because of non-linear relationships in the data, the pipeline calculates the natural log (Ln) of the prices in the training data, trains a model to predict this natural log of price value, and then calculates the exponential of the scored label to get the predicted price.
The training pipeline is shown in the exhibit. (Click the Training pipeline tab.)

Training pipeline -

You create a real-time inference pipeline from the training pipeline, as shown in the exhibit. (Click the Real-time pipeline tab.)

Real-time pipeline -

You need to modify the inference pipeline to ensure that the web service returns the exponential of the scored label as the predicted automobile price and that client applications are not required to include a price value in the input values.
Which three modifications must you make to the inference pipeline? Each correct answer presents part of the solution.
NOTE: Each correct selection is worth one point.
  1. A Connect the output of the Apply SQL Transformation to the Web Service Output module.
  2. B Replace the Web Service Input module with a data input that does not include the price column.
  3. C Add a Select Columns module before the Score Model module to select all columns other than price.
  4. D Replace the training dataset module with a data input that does not include the price column.
  5. E Remove the Apply Math Operation module that replaces price with its natural log from the data flow.
  6. F Remove the Apply SQL Transformation module from the data flow.
Xem giải thích

🧩 Phân tích chi tiết nội dung câu hỏi

Câu hỏi này thuộc về Azure Machine Learning Designer (trong Microsoft Azure), không phải AWS như đề cập ban đầu (có thể là nhầm lẫn). Nó mô tả một pipeline huấn luyện (training pipeline) và pipeline suy luận thời gian thực (real-time inference pipeline) để dự đoán giá ô tô dựa trên dữ liệu có mối quan hệ phi tuyến tính.

✅ Quy trình training pipeline (dựa trên hình ảnh được mô tả):

  • Automobile data (dữ liệu đầu vào chứa các đặc trưng và cột price).
  • Apply Math Operation (thay thế price bằng Ln(price) để xử lý phi tuyến tính).
  • Split Data (chia 70% train / 30% validate).
  • Linear Regression → Train Model (dự đoán Ln(price) từ các đặc trưng).
  • Score Model (dự đoán Ln(price prediction) trên dữ liệu validate).
  • Apply Math Operation (thay thế Scored Labels bằng Exp(Scored Labels) để lấy giá dự đoán gốc).
  • Apply SQL Transformation (SELECT Scored Labels AS predicted_price).

🛠️ Vấn đề ở real-time inference pipeline (dựa trên hình ảnh thứ hai):

  • Web Service Input → Automobile data (dữ liệu đầu vào từ client, có cột price).
  • Apply Math Operation (thay price bằng Ln(price) – vấn đề: input từ client không nên có price vì đây là nhãn cần dự đoán).
  • Trained Model (mô hình đã train: MD-Automobile Price Regression).
  • Score Model (dự đoán Ln(price prediction)).
  • Apply Math Operation (thay Scored Labels bằng Exp(Scored Labels)).
  • Apply SQL Transformation (SELECT [Scored Labels] AS predicted_price).
  • Web Service Output (output hiện tại có thể không kết nối đúng từ Apply SQL, dẫn đến không trả về exp(scored label) đúng cách).

🎯 Yêu cầu modify inference pipeline:

  • Đảm bảo web service trả về exp(scored label) dưới dạng predicted automobile price.
  • Client applications KHÔNG cần cung cấp giá trị price trong input (input chỉ có features).
  • Đây là câu hỏi multi-select (3 đáp án đúng, mỗi cái 1 điểm).

📘 Kiến thức cập nhật: Dựa trên Azure ML Designer phiên bản mới nhất (tính đến 2026, theo docs Azure ML v2), pipeline inference phải loại bỏ label column khỏi input (sử dụng Select Columns), tránh transform label không cần thiết, và route output đúng đến Web Service Output để real-time endpoint hoạt động mượt mà. (Nguồn: Azure ML Designer documentation, Real-time endpoints).

✅ Đáp án đúng (3 modifications cần thiết)

Dựa trên phân tích hình ảnh và logic pipeline, 3 đáp án đúng là:

  1. Connect the output of the Apply SQL Transformation to the Web Service Output module.
  2. Add a Select Columns module before the Score Model module to select all columns other than price.
  3. Remove the Apply Math Operation module that replaces price with its natural log from the data flow.

Lý do lựa chọn:

  • Pipeline hiện tại vẫn yêu cầu input có price (từ Web Service Input + Automobile data), transform thành Ln(price) (không cần cho inference), và output có thể chưa route đúng từ Apply SQL → dẫn đến client phải gửi price và output không phải exp(scored label).
  • Các thay đổi này loại bỏ dependency vào price input, đảm bảo Score Model chỉ nhận features, và output là predicted_price từ exp(scored). Điều này khớp chuẩn best practice Azure ML Designer cho real-time inference.

🔍 Giải thích tất cả các phương án (đúng/sai)

  • ✅ Connect the output of the Apply SQL Transformation to the Web Service Output module.
    Đúng: Hình ảnh cho thấy Apply SQL (SELECT [Scored Labels] AS predicted_price) chưa kết nối rõ ràng hoặc trực tiếp đến Web Service Output (có thể output hiện tại từ Score Model hoặc Apply Math Exp). Kết nối này đảm bảo web service trả về đúng exp(scored label) dưới dạng predicted_price, không lẫn lộn columns khác. Không connect sẽ làm output không đúng format yêu cầu.

  • ❌ Replace the Web Service Input module with a data input that does not include the price column.
    Sai: Web Service Input là bắt buộc cho real-time inference pipeline (tạo endpoint REST API). Thay bằng "data input" (như Reader module) sẽ phá hủy tính real-time, biến pipeline thành batch/offline. Client vẫn cần gửi features qua API, không cần thay module này.

  • ✅ Add a Select Columns module before the Score Model module to select all columns other than price.
    Đúng: Input từ Web Service Input/Automobile data có cột price (label), gây lỗi khi Score Model (chỉ expect features). Thêm Select Columns trước Score để loại price, đảm bảo input sạch (chỉ features), client không cần gửi price. Đây là fix chuẩn cho inference (theo Azure docs).

  • ❌ Replace the training dataset module with a data input that does not include the price column.
    Sai: "Training dataset module" có lẽ ám chỉ Automobile data trong inference (không phải training). Thay nó không giải quyết gốc rễ vì Web Service Input vẫn expect full schema có price. Hơn nữa, dataset gốc dùng cho demo, không nên thay – tốt hơn dùng Select Columns để filter động.

  • ✅ Remove the Apply Math Operation module that replaces price with its natural log from the data flow.
    Đúng: Module này (Replace price with Ln(price)) chỉ cần cho training (transform label). Trong inference, input không có price (sau Select Columns), và mô hình predict trực tiếp Ln(price) từ features → không cần Ln transform nữa. Giữ lại sẽ lỗi nếu không có price hoặc transform sai.

  • ❌ Remove the Apply SQL Transformation module from the data flow.
    Sai: Apply SQL dùng để rename Scored Labels (sau Exp) thành predicted_price, làm output sạch sẽ cho client. Xóa nó sẽ trả về raw Scored Labels (không friendly), vi phạm yêu cầu "returns the exponential of the scored label as the predicted automobile price".

🛡️ Kết luận: Áp dụng 3 thay đổi đúng sẽ làm pipeline inference hoàn hảo: Input features only → Score → Exp → Rename → Output predicted_price. Test endpoint qua Azure ML Studio để verify! (Nguồn tham khảo: Azure ML pipeline troubleshooting).