End-to-end analytical pipeline built to identify what factors influence crime case resolution
in Los Angeles. The project combines exploratory analysis, feature engineering, and
supervised classification to support better resource allocation decisions.
Dataset
LAPD historical crime records with victim, district, and event attributes.
Prepared data by handling missing values, duplicates, and categorical normalization.
Built exploratory views to understand crime trends by location, time, and victim profile.
Created engineered features to improve signal quality for classification models.
Trained and compared Logistic Regression and Random Forest baselines.
Evaluated performance using accuracy, recall, F1-score, and AUC-ROC.
Key Findings
Demographic and contextual variables showed strong explanatory power for case outcomes.
District and crime-type patterns revealed significant variation in resolution rates.
77%Best accuracy (Random Forest)
0.82AUC-ROC
High impactVictim and district features
Visual Output
Main analytic dashboard with case-solvability context.Distribution of crime events across city districts.Feature importance used to explain prediction drivers.