Skip to content

Repository files navigation

🤖 ML.NET.Classifier

ML.NET.Classifier is a Windows Forms application built with ML.NET for exploring numeric binary classification and SMS spam detection using real datasets. It brings data preparation, model training, and visual performance evaluation into one workspace.

The application enables users to:

  • Load and preprocess CSV datasets for numeric or text classification
  • Train and compare classification models
  • Inspect predictions and evaluate performance using key metrics
  • Visualize results through ROC curves, confusion matrices, and cumulative gains charts

Build overview · Getting started · Results · Charts and metrics · Implementation · Project structure

✨ Features

  • 🤖 Four classifiers: Logistic Regression, Averaged Perceptron, Decision Tree, and Naive Bayes.
  • 🔀 Two workflows: numeric binary classification and text classification, with included diabetes and SMS datasets.
  • 🎯 Reproducible data preparation: repeatable splits that preserve class proportions, with a separate test set.
  • 🧹 Preprocessing fitted on training data: numeric scaling, class weights, text features, and Naive Bayes balancing and probability calibration.
  • 📈 Visual evaluation: ROC/AUC, positive and negative precision/recall, confusion matrices, and cumulative gains for each class.
  • 🧩 Separate responsibilities: projects for data loading, preparation, model training, visualization, and the desktop interface.

🛠️ Build overview

The solution separates data loading, preparation, model training, and visualization into individual projects. The Windows Forms interface connects them into one workflow.

🧹 Preparing the data

  • Numeric data is split into training, validation, and test sets while preserving class proportions. Numeric scaling and class weights are calculated from training data.
  • Text data is read with CsvHelper, cleaned, checked for conflicting labels, and deduplicated before splitting into training and test sets. Text features are learned from training data only.

🧠 Training and evaluation

  • Logistic Regression and Averaged Perceptron classify numeric data, with decision thresholds selected using validation F1 scores. Logistic Regression also uses class weights.
  • For text classification, Decision Tree uses FastForest to combine multiple trees. Naive Bayes uses oversampling to balance training classes and calibration to adjust probability estimates. All models are evaluated on a separate test set.

📈 Displaying the results

  • Dataset charts summarize the loaded data. Performance metrics and confusion matrices show classification results, while precision and recall charts compare positive and negative predictions.
  • ROC/AUC evaluates ranking for numeric models, and cumulative gains shows how quickly text models find ham or spam. Both curve calculations are implemented in C# and group equal scores together.

🖥️ Application preview

Numeric binary classification

Logistic Regression interface with binary classification metrics

Text classification

Text classification interface with FastForest results, labelled Decision Tree in the application

🚀 Getting started

🛠️ Requirements

  • .NET 8 SDK, or a compatible newer SDK with the .NET 8 Windows Desktop runtime available.
  • Optional: Visual Studio 2022 with .NET 8 support and the .NET desktop development workload, or VS Code with C# Dev Kit.

▶️ Build and run

git clone https://github.com/Bgajski/ML.NET.Classifier.git
cd ML.NET.Classifier

dotnet restore ML.NET.Classifier.sln
dotnet build ML.NET.Classifier.sln --configuration Release
dotnet run --project ML.Design/ML.Design.csproj --configuration Release

🧪 Try the included datasets

  1. Load a CSV file from ML.Dataset.
  2. Review the detected data and classification type.
  3. Select a compatible algorithm and train it.
  4. Inspect the metrics and select an available chart.
  5. Train the other algorithm for the same dataset to compare results.
Dataset Task Models
diabetes_bin.csv Predict the binary Outcome label from numeric features Logistic Regression, Averaged Perceptron
spam_txt.csv Classify SMS messages as ham or spam Decision Tree, Naive Bayes

📊 Evaluation results

The following results were obtained using the application and the included datasets. Each table describes one test set evaluation, rather than an average across multiple splits. These are example results, not independently verified benchmarks. Percentages use two decimal places. AUC and log loss are shown as decimal values.

Compare models within the same task. The diabetes and SMS datasets have different labels, class distributions, and difficulty.

Numeric classification - diabetes

Here, the positive class is Outcome = 1 and the negative class is Outcome = 0.

Metric Logistic Regression Averaged Perceptron
Accuracy 72.08% 72.08%
Precision / PPV 58.21% 57.97%
Recall / TPR 72.22% 74.07%
Negative predictive value / NPV 82.76% 83.53%
Specificity / TNR 72.00% 71.00%
ROC AUC 0.80 0.81

Reading the comparison: both models achieve the same accuracy. Averaged Perceptron has slightly higher recall and AUC, while Logistic Regression has slightly higher precision and specificity. These small differences apply to this test set. They do not show that either model will consistently perform better on other datasets or splits.

This dataset is used as a classification example. These results do not establish clinical suitability.

Text classification - SMS spam detection

Metric Decision Tree Naive Bayes
Macro accuracy 93.08% 95.31%
Micro accuracy 97.87% 98.84%
Log loss 0.0803 0.0324

Both models were evaluated on 1,031 test messages: 903 ham (legitimate messages) and 128 spam. Simply predicting ham for every message would achieve 87.58% accuracy, so accuracy should be considered alongside the number of spam messages detected.

Reading the comparison: Decision Tree detects 111 spam messages and flags 5 ham messages as spam. Naive Bayes makes 12 errors, compared with 22 for Decision Tree. It detects 116 of 128 spam messages and flags no ham messages as spam. Naive Bayes also has lower reported log loss in this evaluation.

Confusion matrices

Rows show the actual labels. Columns show the predicted labels.

Decision Tree Predicted ham Predicted spam
Actual ham 898 5
Actual spam 17 111
Naive Bayes Predicted ham Predicted spam
Actual ham 903 0
Actual spam 12 116

The application also reports results separately for each class. In the ham rows, “positive” means ham. In the spam rows, “positive” means spam. TP and TN count correct predictions, FP and FN count incorrect predictions.

Model Target class TP TN FP FN
Decision Tree Class 0 - ham 898 111 17 5
Decision Tree Class 1 - spam 111 898 5 17
Naive Bayes Class 0 - ham 903 116 12 0
Naive Bayes Class 1 - spam 116 903 0 12

Detection rates for each class

Rate Decision Tree Naive Bayes
Ham true positive rate 99.45% 100.00%
Ham false positive rate 13.28% 9.38%
Spam true positive rate 86.72% 90.62%
Spam false positive rate 0.55% 0.00%

These rates describe the model's class decisions. Although displayed alongside cumulative gains in the application, they are not cumulative gains values. For example, ham FPR measures the proportion of actual spam incorrectly classified as ham.

🔎 Understanding the results

Classification metrics

For binary metrics, first identify which class counts as positive. TP means true positive, TN means true negative, FP means false positive, and FN means false negative.

Metric Calculation or meaning
Accuracy (TP + TN) / (TP + TN + FP + FN)
Precision / PPV TP / (TP + FP) - how often a positive prediction is correct
Recall / TPR TP / (TP + FN) - proportion of actual positives detected
NPV TN / (TN + FN) - how often a negative prediction is correct
Specificity / TNR TN / (TN + FP) - proportion of actual negatives correctly classified as negative
FPR FP / (FP + TN) - proportion of actual negatives incorrectly flagged
Macro accuracy Average recall across classes, giving each class equal weight
Micro accuracy Overall proportion of correctly classified examples
Log loss Measures how well predicted probabilities match actual labels, lower is better

Precision is also called PPV. Recall is also called TPR. Each pair refers to one metric.

ROC curve and AUC

The ROC curve plots false positive rate on the X-axis against true positive rate on the Y-axis as the decision threshold changes.

  • Curves closer to the upper left corner indicate stronger separation.
  • The diagonal shows the expected performance of random ranking.
  • AUC summarizes how well the model ranks positive examples above negative ones. A value of 1.0 means perfect separation on the evaluated data, 0.5 corresponds to random ranking performance.

The application calculates the ROC curve directly from prediction scores in C#. It sorts the scores, groups equal values, and calculates the area between consecutive curve points using the trapezoidal rule:

Area between points = (FPR₂ − FPR₁) × (TPR₁ + TPR₂) / 2
AUC = sum of these areas

Both numeric models use raw prediction scores for ROC calculation. Each score is paired with its actual label, and a higher score indicates the positive class. These scores do not need to fall between 0 and 1. Using only the final PredictedLabel would lose the information needed to compare different thresholds.

Cumulative gains

Cumulative gains answers: “If I inspect the highest ranked messages first, what fraction of the target class will I find?”

  • X-axis: fraction of all messages inspected.
  • Y-axis: fraction of the target class found.
  • 🟢 Green - ham: messages most likely to be ham are inspected first. This reverses the spam score order.
  • 🔴 Red - spam: messages ranked from highest to lowest spam score.
  • ⚪ Gray dashed line - random selection: inspecting 20% of messages finds about 20% of either target class on average.

For example, a spam curve passing through (20%, 90%) would mean that inspecting the highest ranked 20% of all messages finds 90% of all spam. This example explains how to read the axes, it is not a measured result from the tables above.

The red and green curves inspect messages in different orders. Red starts with the highest spam scores, green starts with the lowest. Both can therefore rise above the random baseline. Equal scores are grouped together so tied messages do not produce different curves just because their row order changes.

The ham curve naturally rises more slowly because ham makes up most of this test set. Even with perfect ranking, finding all 903 ham messages requires inspecting at least 87.58% of the 1,031 messages. Finding all 128 spam messages requires at least 12.42%.

Every complete gains curve ends at (100%, 100%): once all messages have been inspected, every target message has been found. This does not mean the model classified every message correctly.

⚙️ Implementation details

Numeric pipeline

  1. Detect the binary target column and combine the remaining numeric columns into Features.
  2. Reserve approximately 20% of rows for testing using a stratified split with seed 42.
  3. Split the remaining rows again, reserving 20% of that portion for validation: approximately 64% training / 16% validation / 20% testing, subject to rounding.
  4. Learn numeric scaling parameters from training data using mean-variance normalization. Logistic Regression also uses class weights calculated from the training set.
  5. Train LbfgsLogisticRegression or AveragedPerceptron.
  6. Select a decision threshold by maximizing F1 on validation data, then evaluate on the separate test set.

Logistic Regression applies the selected threshold to Probability. Averaged Perceptron applies it to the raw Score. Accuracy, precision, and recall are calculated from the resulting classification decisions. ROC/AUC uses the scores before applying that threshold.

Text pipeline

  1. Parse the label and message columns with CsvHelper, preserving quoted fields.
  2. Trim values, normalize label casing, skip empty entries, reject conflicting labels, and remove duplicate messages.
  3. Create a seeded, stratified 80% training / 20% testing split before balancing or model fitting.
  4. Learn the mapping from labels to class IDs and extract text features using training data only.
  5. Train the selected classifier, then evaluate it on the test set without balancing or oversampling the test messages.

Decision Tree

The text features include individual words, pairs of consecutive words, and sequences of three characters. They are weighted using TF-IDF, which takes into account how often a feature appears in a message and across the training messages.

Setting Value
Number of trees 120
Maximum leaves per tree 40
Minimum examples per leaf 3
Feature fraction 0.8
Random seed 42

Model explanation: The “Decision Tree” option uses FastForest, a random forest made up of multiple decision trees.

ML.NET’s OneVersusAll trainer builds a classifier for each class against the others. For this dataset, that means ham versus spam and spam versus ham. The cumulative gains chart ranks messages by normalized spam probability: highest first for the spam curve and lowest first for the ham curve.

Naive Bayes

ML.NET's default text feature extraction converts messages into numeric inputs. Examples from smaller classes are repeated to balance the training data, this is called oversampling.

Probability calibration uses three fold cross validation within the training set. Each fold is scored by a model trained on the other two folds. These predictions are used to fit temperature scaling, which adjusts confidence through a softmax calculation without changing the predicted class. The final classifier is then trained on the full training set.

Calibrated probabilities are used for log loss. For cumulative gains, the model uses the difference between its spam and ham log scores. This preserves ranking detail that can be lost when probabilities round to 0 or 1.

Input formats

Numeric CSV: include a header, a binary target column, and numeric feature columns. Recognized target names include Label, Outcome, Class, Target, and Y. If none of these names is found, the application looks for a column containing binary values. Use 0/1 or false/true labels.

Text CSV: include a header, with the label in the first column and message in the second. Text preparation ignores extra columns, so the complete message must be contained in the second column. The included dataset uses v1 and v2. Quote messages containing commas:

label,text
ham,"Hi, are we still meeting today?"
spam,"Congratulations! Claim your prize now."

The specialized cumulative gains chart requires exactly the ham and spam classes.

🗂️ Project structure

Project Responsibility
ML.Design Windows Forms startup project and application workflow
ML.Data Data loading and dataset characteristic checks
ML.DataPreparation CSV preparation, stratified splits, validation, and class balancing/weighting
ML.Model Trainers, threshold optimization, probability calibration, and evaluation
ML.Graph Dataset chart rendering and chart type selection
ML.Performance Performance descriptions, confusion matrices, ROC/AUC, and cumulative gains
ML.Dataset Included numeric and text CSV datasets
ML.Tests NUnit data preparation tests
ML.ReadmeExtra README screenshots

📦 Dependencies

The project uses the following framework and package versions.

Dependency Version
Target framework .NET 8 for Windows
Microsoft.ML 4.0.2
Microsoft.ML.CpuMath 4.0.2
Microsoft.ML.LightGbm 4.0.2
WinForms.DataVisualization 1.9.2
CsvHelper 32.0.3
Microsoft.NET.Test.Sdk 17.6.0
NUnit 3.13.3

🗃️ Datasets and attribution

Dataset licensing and attribution are governed by the respective dataset sources, separately from the application license.

📄 License

The application is distributed under the MIT License.

About

ML.NET.Classifier is a .NET Windows Forms application that utilizes the ML.NET library to demonstrate binary and textual data classification process using relevant metrics and visual charts.

Topics

Resources

Stars

2 stars

Watchers

1 watching

Forks

Used by

Contributors

Languages