← Back to projects

Dec 2024 - Present

Behavioral Efficiency Scoring with Double Machine Learning

Built a causal evaluation framework that separated behavioral signal from environmental confounders, producing fair rankings and actionable coaching insights for operations teams.

PythonPyTorchScikit-LearnSHAPAWS S3Statistical Inference

Cohort Evaluated

Team-wide

Efficiency Gap Quantified

Confounder-adjusted

Savings Opportunity

Significant, recurring

Fairness Stabilization Trend

Normalized confounder-adjusted scoring performance over release cycles.

S1S2S3S4S5S6
Cohort evaluatedTeam-wide
Efficiency gap quantifiedConfounder-adjusted
Savings opportunitySignificant

As confounder controls improved, ranking consistency increased and recommendations became easier to operationalize.

Causal DAG

causal_dag.py — DML Behavioral Efficiency
Confounder
Treatment
Mediator
Outcome
Instrument
· click to expand ·

Problem

Performance comparisons across a large population of subjects were noisy because environmental factors like route topography, traffic, and weather heavily biased raw efficiency measurements.

Approach

I designed a Double Machine Learning pipeline to produce operation-conditioned effect estimates that separate behavioral signal from environmental covariates.

  • Gathered and cleaned high-volume telemetry, then engineered environmental and behavioral features for fair comparison.
  • Estimated nuisance components with five-fold cross-fitted XGBoost/LightGBM learners, so no observation informs its own nuisance prediction.
  • Added mixed-effects modeling so scores remained stable across cohorts and repeated observations.
  • Applied matched-pair validation, cluster-robust variance estimation, and SHAP diagnostics to confirm directionality and model consistency.

Outcome

The final scorecard surfaced a measurable, confounder-adjusted efficiency gap across the full cohort with statistically defensible confidence, enabling targeted coaching on behavioral patterns.

Why it mattered

This work moved discussions from anecdotal feedback to causal evidence and quantified a significant recurring savings opportunity, improving trust from operations stakeholders and accelerating coaching adoption.