Pipeline before modelling

Author

Cecilia Regueira

Executive Summary

Note

Following the Reproducible framework, I structured the pipeline to ensure transparency and reproducibility.

I started with data auditing, explicitly handled missing values, and flagged outliers instead of removing them.

Then I engineered behavioral features like transaction timing and CVV mismatch, and placed all transformations inside a pipeline to prevent leakage.

Finally, I used a stratified train-test split and evaluated performance using ROC AUC.

This Pipline is a summary of the one used in the case-study. Fraud detection.

0. Set up Libreries

import pandas as pd import numpy as np

from sklearn.model_selection import train_test_split from sklearn.pipeline import Pipeline from sklearn.compose import ColumnTransformer

from sklearn.impute import SimpleImputer from sklearn.preprocessing import OneHotEncoder, StandardScaler

from sklearn.ensemble import RandomForestClassifier from sklearn.linear_model import LogisticRegression

from sqlalchemy import create_engine

from sklearn.metrics import roc_auc_score

from imblearn.pipeline import Pipeline as ImbPipeline from imblearn.over_sampling import SMOTE

from sklearn.model_selection import GridSearchCV

Reproducibility

RANDOM_STATE = 42 np.random.seed(RANDOM_STATE)

1.Data Collection

engine = create_engine( “postgresql://username:password@host:5432/database_name” )

Query

query = ““” SELECT * FROM transactions ““”

Load into DataFrame

df = pd.read_sql(query, engine)

Basic audit

print(“Shape:”, df.shape) print(df.dtypes) print(df.head())

Missing values audit

missing_report = df.isnull().mean().sort_values(ascending=False) print(missing_report)

Duplicate check

print(“Duplicates:”, df.duplicated().sum())

2.Missing Values

Separate types

num_cols = df.select_dtypes(include=np.number).columns.tolist() cat_cols = df.select_dtypes(include=“object”).columns.tolist()

Remove target if needed

target = “isFraud” num_cols = [c for c in num_cols if c != target]

Define imputers

num_imputer = SimpleImputer(strategy=“median”) cat_imputer = SimpleImputer(strategy=“most_frequent”)

3. Outliers

def add_outlier_flag(df, cols): df_out = df.copy() for col in cols: Q1 = df[col].quantile(0.25) Q3 = df[col].quantile(0.75) IQR = Q3 - Q1 df_out[f”{col}_outlier”] = ( (df[col] < Q1 - 1.5 * IQR) | (df[col] > Q3 + 1.5 * IQR) ).astype(int) return df_out

df = add_outlier_flag(df, num_cols)

4. Feature Engineering

principle: features should reflect behavior and context

Time features

df[“transactionTime”] = pd.to_datetime(df[“transactionTime”])

df[“hour”] = df[“transactionTime”].dt.hour df[“weekday”] = df[“transactionTime”].dt.weekday

Behavior feature

df[“cvv_match”] = (df[“cardCVV”] == df[“enteredCVV”]).astype(int)

Drop raw timestamp after extracting info

df = df.drop(columns=[“transactionTime”])

5.Encoding + Scaling

transformations must be contained (no leakage)

numeric_features = df.select_dtypes(include=np.number).columns.tolist() categorical_features = df.select_dtypes(include=“object”).columns.tolist()

Remove target

numeric_features = [c for c in numeric_features if c != target]

preprocessor = ColumnTransformer( transformers=[ (“num”, Pipeline([ (“imputer”, SimpleImputer(strategy=“median”)), (“scaler”, StandardScaler()) ]), numeric_features),

    ("cat", Pipeline([
        ("imputer", SimpleImputer(strategy="most_frequent")),
        ("encoder", OneHotEncoder(handle_unknown="ignore"))
    ]), categorical_features)
]

)

6. Train-Test split (STRTIFIED)

target = “isFraud”

X = df.drop(columns=[target]) y = df[target]

X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.3, stratify=y, random_state=RANDOM_STATE )

print(“Fraud rate train:”, y_train.mean()) print(“Fraud rate test:”, y_test.mean())

7. Model Pipeline

Hyperparameter tuning is performed entirely on the training set using cross-validation.

Once the best parameters are identified, the model is refit on the full training data.

The test set is only used once at the end to evaluate the final model, ensuring an unbiased estimate of performance.

Tune (inside TRAIN only)

  • Logistic → RandomizedSearchCV

  • Random Forest → Grid/Random Search

Logit

logit_pipeline = ImbPipeline(steps=[ (“preprocessing”, preprocessor), (“smote”, SMOTE(random_state=RANDOM_STATE)), (“model”, LogisticRegression(max_iter=1000))])

param_dist_logit = { “model__C”: [0.01, 0.1, 1, 10], # regularization strength “model__penalty”: [“l2”], # (keep simple) “model__solver”: [“lbfgs”] }

logit_search = RandomizedSearchCV( logit_pipeline, param_distributions=param_dist_logit, n_iter=5, cv=3, scoring=“roc_auc”, n_jobs=-1, random_state=42 )

logit_search.fit(X_train, y_train)

best_logit = logit_search.best_estimator_

Random Forest

rf_pipeline = ImbPipeline(steps=[ (“preprocessing”, preprocessor), (“smote”, SMOTE(random_state=42)), (“model”, RandomForestClassifier(random_state=42))])

I tuned the Random Forest model using GridSearchCV within a pipeline that also includes SMOTE. This ensures that oversampling is applied only within each training fold, preventing data leakage.

GridSearch

param_grid = { “model__n_estimators”: [100, 200], “model__max_depth”: [5, 10, None], “model__min_samples_split”: [2, 5], “model__min_samples_leaf”: [1, 2] }

grid_search = GridSearchCV( rf_pipeline, param_grid, cv=3, scoring=“roc_auc”,
n_jobs=-1, verbose=1 )

grid_search.fit(X_train, y_train)

print(“Best Parameters:”, grid_search.best_params_) print(“Best AUC:”, grid_search.best_score_)

best_rf_model = grid_search.best_estimator_

This does:

Train many models on TRAIN data Cross-validate internally Select best parameters

models = { “Random Forest”: best_rf_model, “Logistic”: best_logit }

for name, model in models.items(): probs = model.predict_proba(X_test)[:, 1] auc = roc_auc_score(y_test, probs) print(f”{name}: {auc:.3f}“)

8. Evaluation

Probabilities

logit_probs = best_logit.predict_proba(X_test)[:, 1] rf_probs = best_rf_model.predict_proba(X_test)[:, 1]

ROC AUC

logit_auc = roc_auc_score(y_test, logit_probs) rf_auc = roc_auc_score(y_test, rf_probs)

print(f”Logistic AUC: {logit_auc:.3f}“) print(f”Random Forest AUC: {rf_auc:.3f}“)

import matplotlib.pyplot as plt from sklearn.metrics import roc_curve

fpr_logit, tpr_logit, _ = roc_curve(y_test, logit_probs) fpr_rf, tpr_rf, _ = roc_curve(y_test, rf_probs)

plt.figure(figsize=(5, 4)) plt.plot(fpr_logit, tpr_logit, label=“Logistic”) plt.plot(fpr_rf, tpr_rf, label=“Random Forest”) plt.plot([0, 1], [0, 1], linestyle=“–”)

plt.xlabel(“False Positive Rate”) plt.ylabel(“True Positive Rate”) plt.title(“ROC Curve Comparison”) plt.legend() plt.tight_layout() plt.show()

You are solving:
  • Highly imbalanced problem (~2% fraud)

  • Business cost trade-off (missed fraud vs false alerts)

ROC: Measures ranking ability (how well model separates fraud vs non-fraud), but does not explain How many frauds you actually catch. Recall: Of all frauds, how many did we detect? Precision: Of what we flagged, how much is actually fraud?