| isFraud | Cases | Total_Amount | Case_Pct | Amount_Pct |
|---|---|---|---|---|
| Not Fraud | 773946 | $104,924,052 | 98% | 97% |
| Fraud | 12417 | $2,796,506 | 2% | 3% |
Fraud is a major challenge for banks
Only ~2% of transactions are fraudulent
Goal: Build a model to predict fraud in real time
Using transactional data, we built a model that can detect fraud in real-time, reducing potential losses while controlling false alerts. The model identifies up to 70% of fraudulent transactions, capturing approximately 90% of the total fraud value.
Note
We need to find the right balance:
Missing fraud vs Flagging normal transactions
financial loss vs poor customer experience

It would be 98% accurate β and still useless.
| isFraud | Cases | Total_Amount | Case_Pct | Amount_Pct |
|---|---|---|---|---|
| Not Fraud | 773946 | $104,924,052 | 98% | 97% |
| Fraud | 12417 | $2,796,506 | 2% | 3% |
| Mean.Transaction | median.Transaction | Min.Transaction | max.Transaction |
|---|---|---|---|
| 157.27 | 50 | 1 | 32850 |
The dataset is transaction-level financial data
| accountNumber | customerId | creditLimit | availableMoney | transactionDateTime | transactionAmount | merchantName | acqCountry | merchantCountryCode | posEntryMode | posConditionCode | merchantCategoryCode | currentExpDate | accountOpenDate | dateOfLastAddressChange | cardCVV | enteredCVV | cardLast4Digits | transactionType | currentBalance | cardPresent | expirationDateKeyInMatch | isFraud | transactionTime |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 737265056 | 737265056 | 5000 | 5000 | 2016-08-13T14:27:32 | 98.55 | Uber | US | US | 02 | 01 | rideshare | 06/2023 | 2015-03-14 | 2015-03-14 | 414 | 414 | 1803 | PURCHASE | 0 | FALSE | FALSE | 0 | 2016-08-13 14:27:32 |
| Cases | Total Amount ($) | Cases (%) | Amount (%) |
|---|---|---|---|
| 17757 | 2666479 | 2.26 | 2.48 |
| dup_key | purchase_count | reversal_count |
|---|---|---|
| 101376441_0_cheapfast.com_128_6683 | 0 | 1 |
Multi swipe, when a vendor accidentally charges a customerβs card multiple times within a short time span.
| Multi-Swipe | Cases | Total Amount ($) | Cases (%) | Amount (%) |
|---|---|---|---|---|
| 1 | 20156 | 2201824 | 2.56 | 2.04 |
We explored patterns to understand what fraud looks like:
Transactions
Merchant types
CVV
Transactions:
Fraudulent transactions tend to be larger than usual
Fraud Varies by Merchant Type
E.g. Online purchases or travel-related transactions show higher fraud rates
Errors in transaction details, such as incorrect CVV entries, are more common in fraud cases
Higher mismatch between entered and real CVV = higher fraud risk.
Think of fraud as behavior that looks different from a customerβs normal habits.
Logistic Regression β simple, interpretable
Random Forest β more flexible, captures complex patterns

Rebalance the data
We also addressed the imbalance problem using techniques like SMOTE, which helps the model learn rare fraud cases better.

| Model Performance Comparison | ||||
| Fraud Detection vs Customer Impact | ||||
| Model | Recall (All Fraud) | Fraud Captured ($) | False Alerts / 10K | Net Benefit |
|---|---|---|---|---|
| Random Forest | 78.6% | 90.6% | 255 | $2,050 |
| Logistic Regression | 21.4% | 37.0% | 73 | $977 |
If the primary goal is to minimize false negatives and ensure that most fraud cases are identified, then the logit model may be the better choice despite its higher false positive rate.
Conversely, if the focus is on achieving greater overall accuracy and reducing false alarms, the random forest model could be better, provided that the cost of missed fraud is acceptable.
Fraud leaves patterns β our job is to detect them early
Thank you!
The idea is:
A meta-model learns when to trust each model
βTwo fraud analysts review the transaction independently, and a third analyst makes the final decision.β
Ensemble / Stacking
| Model Performance Comparison | ||||
| Fraud Detection vs Customer Impact | ||||
| Model | Recall (All Fraud) | Fraud Captured ($) | False Alerts / 1K | Net Benefit |
|---|---|---|---|---|
| Random Forest | 78.6% | 90.6% | 255 | $2,050 |
| Logistic Regression | 21.4% | 37.0% | 73 | $977 |
| Stacked Ensemble | 78.6% | 90.6% | 219 | $2,212 |
While the ensemble combined insights from both Logistic Regression and Random Forest, the Random Forest model delivered the strongest fraud detection performance in this analysis. Nonetheless, the ensemble demonstrates how multiple models can be combined into a single risk score, potentially increasing robustness and reducing dependence on a single algorithm.