Pipeline before modelling
Executive Summary
Following the Reproducible framework, I structured the pipeline to ensure transparency and reproducibility.
I started with data auditing, explicitly handled missing values, and flagged outliers instead of removing them.
Then I engineered behavioral features like transaction timing and CVV mismatch, and placed all transformations inside a pipeline to prevent leakage.
Finally, I used a stratified train-test split and evaluated performance using ROC AUC.
This Pipline is a summary of the one used in the case-study. Fraud detection.
0. Set up Libreries
import pandas as pd import numpy as np
from sklearn.model_selection import train_test_split from sklearn.pipeline import Pipeline from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.ensemble import RandomForestClassifier from sklearn.linear_model import LogisticRegression
from sqlalchemy import create_engine
from sklearn.metrics import roc_auc_score
from imblearn.pipeline import Pipeline as ImbPipeline from imblearn.over_sampling import SMOTE
from sklearn.model_selection import GridSearchCV
Reproducibility
RANDOM_STATE = 42 np.random.seed(RANDOM_STATE)
1.Data Collection
engine = create_engine( “postgresql://username:password@host:5432/database_name” )
Query
query = ““” SELECT * FROM transactions ““”
Load into DataFrame
df = pd.read_sql(query, engine)
Basic audit
print(“Shape:”, df.shape) print(df.dtypes) print(df.head())
Missing values audit
missing_report = df.isnull().mean().sort_values(ascending=False) print(missing_report)
Duplicate check
print(“Duplicates:”, df.duplicated().sum())
2.Missing Values
Separate types
num_cols = df.select_dtypes(include=np.number).columns.tolist() cat_cols = df.select_dtypes(include=“object”).columns.tolist()
Remove target if needed
target = “isFraud” num_cols = [c for c in num_cols if c != target]
Define imputers
num_imputer = SimpleImputer(strategy=“median”) cat_imputer = SimpleImputer(strategy=“most_frequent”)
3. Outliers
def add_outlier_flag(df, cols): df_out = df.copy() for col in cols: Q1 = df[col].quantile(0.25) Q3 = df[col].quantile(0.75) IQR = Q3 - Q1 df_out[f”{col}_outlier”] = ( (df[col] < Q1 - 1.5 * IQR) | (df[col] > Q3 + 1.5 * IQR) ).astype(int) return df_out
df = add_outlier_flag(df, num_cols)
4. Feature Engineering
principle: features should reflect behavior and context
Time features
df[“transactionTime”] = pd.to_datetime(df[“transactionTime”])
df[“hour”] = df[“transactionTime”].dt.hour df[“weekday”] = df[“transactionTime”].dt.weekday
Behavior feature
df[“cvv_match”] = (df[“cardCVV”] == df[“enteredCVV”]).astype(int)
Drop raw timestamp after extracting info
df = df.drop(columns=[“transactionTime”])
5.Encoding + Scaling
transformations must be contained (no leakage)
numeric_features = df.select_dtypes(include=np.number).columns.tolist() categorical_features = df.select_dtypes(include=“object”).columns.tolist()
Remove target
numeric_features = [c for c in numeric_features if c != target]
preprocessor = ColumnTransformer( transformers=[ (“num”, Pipeline([ (“imputer”, SimpleImputer(strategy=“median”)), (“scaler”, StandardScaler()) ]), numeric_features),
("cat", Pipeline([
("imputer", SimpleImputer(strategy="most_frequent")),
("encoder", OneHotEncoder(handle_unknown="ignore"))
]), categorical_features)
]
)
6. Train-Test split (STRTIFIED)
target = “isFraud”
X = df.drop(columns=[target]) y = df[target]
X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.3, stratify=y, random_state=RANDOM_STATE )
print(“Fraud rate train:”, y_train.mean()) print(“Fraud rate test:”, y_test.mean())
7. Model Pipeline
Hyperparameter tuning is performed entirely on the training set using cross-validation.
Once the best parameters are identified, the model is refit on the full training data.
The test set is only used once at the end to evaluate the final model, ensuring an unbiased estimate of performance.
Tune (inside TRAIN only)
Logistic → RandomizedSearchCV
Random Forest → Grid/Random Search
Logit
logit_pipeline = ImbPipeline(steps=[ (“preprocessing”, preprocessor), (“smote”, SMOTE(random_state=RANDOM_STATE)), (“model”, LogisticRegression(max_iter=1000))])
param_dist_logit = { “model__C”: [0.01, 0.1, 1, 10], # regularization strength “model__penalty”: [“l2”], # (keep simple) “model__solver”: [“lbfgs”] }
logit_search = RandomizedSearchCV( logit_pipeline, param_distributions=param_dist_logit, n_iter=5, cv=3, scoring=“roc_auc”, n_jobs=-1, random_state=42 )
logit_search.fit(X_train, y_train)
best_logit = logit_search.best_estimator_
Random Forest
rf_pipeline = ImbPipeline(steps=[ (“preprocessing”, preprocessor), (“smote”, SMOTE(random_state=42)), (“model”, RandomForestClassifier(random_state=42))])
I tuned the Random Forest model using GridSearchCV within a pipeline that also includes SMOTE. This ensures that oversampling is applied only within each training fold, preventing data leakage.
GridSearch
param_grid = { “model__n_estimators”: [100, 200], “model__max_depth”: [5, 10, None], “model__min_samples_split”: [2, 5], “model__min_samples_leaf”: [1, 2] }
grid_search = GridSearchCV( rf_pipeline, param_grid, cv=3, scoring=“roc_auc”,
n_jobs=-1, verbose=1 )
grid_search.fit(X_train, y_train)
print(“Best Parameters:”, grid_search.best_params_) print(“Best AUC:”, grid_search.best_score_)
best_rf_model = grid_search.best_estimator_
Randomized Search
from sklearn.model_selection import RandomizedSearchCV
param_dist = { “model__n_estimators”: [100, 200, 300], “model__max_depth”: [5, 10, 20, None], “model__min_samples_split”: [2, 5, 10], “model__min_samples_leaf”: [1, 2, 4] }
random_search = RandomizedSearchCV( rf_pipeline, # Your full pipeline (preprocessing + SMOTE + model) param_distributions=param_dist, #Dictionary of hyperparameters to try n_iter=10, #Number of random combinations to test cv=3, #Data is split into 3 parts: Train on 2, validate on 1 (repeat 3 times) scoring=“roc_auc”, #Metric used to pick best model n_jobs=-1, #Use all CPU cores random_state=42, #Makes results reproducible verbose=1 #Prints training progress )
random_search.fit(X_train, y_train)
This does:
Train many models on TRAIN data Cross-validate internally Select best parameters
models = { “Random Forest”: best_rf_model, “Logistic”: best_logit }
for name, model in models.items(): probs = model.predict_proba(X_test)[:, 1] auc = roc_auc_score(y_test, probs) print(f”{name}: {auc:.3f}“)
8. Evaluation
Probabilities
logit_probs = best_logit.predict_proba(X_test)[:, 1] rf_probs = best_rf_model.predict_proba(X_test)[:, 1]
ROC AUC
logit_auc = roc_auc_score(y_test, logit_probs) rf_auc = roc_auc_score(y_test, rf_probs)
print(f”Logistic AUC: {logit_auc:.3f}“) print(f”Random Forest AUC: {rf_auc:.3f}“)
import matplotlib.pyplot as plt from sklearn.metrics import roc_curve
fpr_logit, tpr_logit, _ = roc_curve(y_test, logit_probs) fpr_rf, tpr_rf, _ = roc_curve(y_test, rf_probs)
plt.figure(figsize=(5, 4)) plt.plot(fpr_logit, tpr_logit, label=“Logistic”) plt.plot(fpr_rf, tpr_rf, label=“Random Forest”) plt.plot([0, 1], [0, 1], linestyle=“–”)
plt.xlabel(“False Positive Rate”) plt.ylabel(“True Positive Rate”) plt.title(“ROC Curve Comparison”) plt.legend() plt.tight_layout() plt.show()
You are solving:
Highly imbalanced problem (~2% fraud)
Business cost trade-off (missed fraud vs false alerts)
ROC: Measures ranking ability (how well model separates fraud vs non-fraud), but does not explain How many frauds you actually catch. Recall: Of all frauds, how many did we detect? Precision: Of what we flagged, how much is actually fraud?