This teaching lab shows why accuracy can be dangerously misleading when the event being predicted is extremely rare. It uses C#, .NET 10, ML.NET, and AutoML to compare a naive baseline, a conventional classifier, controlled undersampling, automated model search, threshold selection, and business-cost policies.
The central lesson is that the useful solution is not merely the model with the best technical metric. It is the combination of model, threshold, and decision policy that produces the best business outcome.
This is Exercise 4 in the AInDotNet Predictive AI lab series.
The application demonstrates how to:
- Recognize why accuracy is misleading for rare-event classification.
- Establish a meaningful majority-class baseline before evaluating ML models.
- Measure precision, recall, F1, ROC-AUC, PR-AUC, and confusion-matrix results.
- Verify that calculated metrics agree with the confusion matrix.
- Keep rebalanced training data separate from realistic validation and test populations.
- Compare a manually selected FastTree model with AutoML candidates.
- Inspect score distributions before choosing an operational threshold.
- Separate model quality from decision policy.
- Compare model-and-threshold combinations using illustrative business costs.
- Reserve an untouched final test population for the selected policy.
- Save and reload the selected model and its decision policy.
- Business question: Is this transaction fraudulent?
- Prediction: A fraud-risk score and a fraud/not-fraud decision.
- Historical label:
Class, where1represents fraud and0represents a legitimate transaction. - Teaching prediction point: The moment a transaction is evaluated for authorization or review.
- Available inputs:
Time,Amount, and anonymized transformed variablesV1throughV28. - Possible action: Approve, hold, decline, or route the transaction for manual review.
- Success criterion: Reduce estimated fraud loss while controlling investigation cost and customer friction.
Because V1 through V28 are anonymized, this dataset cannot prove that every input would genuinely be available at the prediction point. A production project must validate point-in-time feature availability and target leakage using the original business definitions.
| Run | Experiment question |
|---|---|
| A | What happens if every transaction is predicted legitimate? |
| B | How well does default FastTree perform on the original imbalanced data? |
| C | Does 10:1 training-set undersampling improve fraud recall? |
| D | Can AutoML find a stronger candidate when explicitly optimizing PR-AUC? |
| E | How does the AutoML model behave at different decision thresholds? |
| F | Which model-and-threshold policy minimizes illustrative business cost on validation data? |
| G | How does the selected policy perform on the untouched final test population? |
The application profiles the raw dataset before modeling. The profile includes:
- Class prevalence
- Majority-class baseline accuracy
- Missing and invalid values
- Selected feature distributions
- Transaction amounts by class
- Exact duplicate observations
Exact duplicates are defined using all model inputs and the historical label. Before the data is split, the application keeps the first occurrence of each exact observation and removes subsequent copies.
This matters because copies of the same observation must not appear in both training and evaluation data. Cross-split duplicates can make validation and test performance look better than performance on genuinely unseen transactions.
The console output reports the original row count, number of duplicates removed, and final modeling population.
After duplicate removal, the application uses approximately:
- 60% of the rows for training
- 20% for validation
- 20% for final testing
These populations serve different purposes.
AutoML uses internal validation while searching candidate trainers. Its reported internal PR-AUC is used by AutoML to choose its winning trial.
The separate validation population is used to:
- Compare Runs B, C, and D
- Inspect threshold behavior
- Select the operational threshold
- Compare illustrative business-cost policies
The validation population is not the final test population.
Run G evaluates the selected model-and-threshold policy against the untouched final test population. The final test data is not used to choose the model, threshold, or business policy.
Download the Credit Card Fraud Detection dataset from Kaggle:
https://www.kaggle.com/datasets/mlg-ulb/creditcardfraud
Place the downloaded file here:
Data/creditcard.csv
The CSV is intentionally excluded from the repository. See Data/README.md for the expected schema, limitations, and usage notes.
AInDotNet.MLNET.CreditCardFraud.sln
AInDotNet.MLNET.CreditCardFraud.csproj
Program.cs
Data/
DataAnalysis/
Evaluation/
Model/
ModelArtifacts/
Training/
- Visual Studio 2026 or another environment supporting .NET 10
- .NET 10 SDK
- Kaggle account or another authorized method of obtaining the dataset
Run these commands from the repository root:
dotnet restore
dotnet build --configuration Release
dotnet run --project .\AInDotNet.MLNET.CreditCardFraud.csproj --configuration ReleaseThe application expects Data/creditcard.csv relative to the project directory.
The application uses a fixed MLContext seed and fixed train, validation, and test splits. Controlled undersampling also uses a fixed random seed.
AutoML uses a time budget. The number of explored trials, winning trainer, model metrics, selected threshold, and final policy can differ across machines or runs even when the data split and ML.NET seed are fixed. Treat included output as a documented example rather than a guarantee of identical numbers.
Before drawing conclusions from a changed version of the exercise, run the complete application three to five times and verify that the central findings remain stable.
- Accuracy: Percentage of all predictions that are correct. It can be dangerously misleading when fraud is extremely rare.
- Precision: Percentage of fraud alerts that are actually fraud.
- Recall: Percentage of actual fraud transactions detected by the model.
- F1: Balance between precision and recall at a specific threshold.
- ROC-AUC: Ranking performance across thresholds, but sometimes overly optimistic for rare events.
- PR-AUC: Precision-recall performance across thresholds and generally more informative for rare positive classes.
- False positive: A legitimate transaction incorrectly flagged as fraud.
- False negative: A fraudulent transaction incorrectly allowed as legitimate.
The application recalculates accuracy, precision, and recall from the confusion matrix and stops if those values do not agree with ML.NET's metrics.
A classifier produces a score or probability-like value. A threshold converts that value into a fraud/not-fraud decision.
Changing the threshold changes the operational policy without retraining the model. A lower threshold may detect more fraud while creating more false positives. A higher threshold may reduce false alerts while missing more fraud.
A model's Probability output should not automatically be interpreted as a calibrated real-world fraud probability. Inspect its score distribution and validate calibration before using the number as an actual probability.
Run F evaluates combinations of:
- Model
- Threshold
- False-positive unit cost
- Missed-fraud loss multiplier
The simulation uses transaction Amount as an instructional proxy for potential fraud loss. It assumes that a detected fraud transaction avoids the modeled loss and that a missed fraud transaction incurs the modeled loss.
Those assumptions are deliberately simple. They are not a validated banking loss model. A production implementation would require institution-specific information about fraud recovery, chargebacks, review expense, customer friction, false-decline consequences, and analyst capacity.
After a successful run, the application writes:
ModelArtifacts/CreditCardFraudModel.zip
ModelArtifacts/CreditCardFraudDecisionPolicy.json
The model and decision policy are stored separately because a trained model does not define the operational threshold or economic assumptions by itself.
The application reloads both artifacts and verifies that:
- The model can be loaded successfully.
- The reloaded model reproduces the original prediction for a verification record.
- The decision policy can be deserialized successfully.
- The reloaded policy matches the selected in-memory policy.
The generated artifacts are excluded from Git and should normally be recreated locally.
- Change the legitimate-to-fraud undersampling ratio from 10:1 to 5:1 or 20:1.
- Compare results with and without exact duplicate removal.
- Expand or narrow the threshold range after inspecting the winning model's score distribution.
- Change the false-positive cost and missed-fraud multiplier. Determine when the preferred policy changes.
- Replace the random split with chronological validation using
Time. - Compare undersampling with class weighting or another imbalance strategy.
- Run AutoML three to five times and compare the winning trainer, validation PR-AUC, and selected policy.
- Add probability calibration and compare calibrated probabilities with ranking scores.
- Add review-capacity constraints, such as a maximum number of transactions that investigators can review each day.
This repository is a teaching implementation, not a production fraud platform.
A production implementation would require:
- Original business definitions for every feature
- Point-in-time feature validation
- Time-aware backtesting
- Entity and duplicate leakage controls
- Current operational data
- Delayed-label handling
- Model and decision-policy versioning
- Data and concept-drift monitoring
- Prediction, decision, and outcome logging
- Security and access controls
- Review-capacity constraints
- Deployment, rollback, and champion/challenger controls
- Institution-specific loss and customer-friction assumptions
- Ongoing actual-versus-predicted analysis
The source code is licensed under the MIT License.
The Kaggle dataset is not covered by this repository's MIT License. Obtain the dataset from an authorized source and follow the dataset publisher's current terms.
https://aindotnet.com/forecasting/ https://aindotnet.com/2026/09/credit-card-fraud-detection-mlnet/