Recidivism & Reentry
Predicting Criminal Recidivism Using Specialized Feature Engineering and XGBoost (NIJ Recidivism Forecasting Challenge, Georgia parolees released 2013) (NCJ 305039, 2021)
This NIJ-funded winning paper applies XGBoost machine learning to predict recidivism among approximately 26,000 individuals released from Georgia prisons on discretionary parole in 2013. The authors achieved best-in-class Brier scores for Year 2 recidivism prediction, with the 'Early Recidivism' feature (Recidivism_Arrest_PrevYear) identified as a key driver of model performance. The paper documents feature importance rankings, model parameters, and performance metrics, providing methodological insights for GPS research on Georgia parolee recidivism forecasting.
Key Findings
The most impactful data from this research collection.
0.2
Year 1 Brier Score 0.1837
StatisticEarly Recidivism Feature Boosted Performance
Finding0.1
Oracle Team Winning Brier 0.1233
StatisticAll Data Points
46 verified data points extracted from primary sources.
Dataset size: ~26,000 Georgia parolees released 2013 Statistic
The aggregated dataset provided by the NIJ contains approximately 26,000 individuals released from Georgia prisons on discretionary parole for post-incarceration supervision between January 1st, 2013 and December 31st, 2013.
26,000 individuals
NIJ train/test split proportion 70/30 Methodology note
NIJ split the dataset into a training and test set with a 70/30 proportion.
Data providers: GDCS and Georgia Bureau of Investigation Methodology note
Both the GDCS and the Georgia Bureau of Investigation provided data. The GDCS data included demographics, prison and parole case information, prior community supervision history, and supervision activities. The Georgia Bureau of Investigation provid…
Recidivism measure definition: new felony or misdemeanor arrest within 3 years Methodology note
GCIC data also provides the recidivism measure, defined as a new felony or misdemeanor arrest episode within three years of parole supervision start date. This recidivism measure includes three dichotomous variables measuring if an individual recidi…
Early Recidivism feature added for Year 2 and Year 3 models Methodology note
For Year 2 and Year 3, a feature called Early Recidivism was added. This feature delineated whether the individual already returned to prison. This could only be added to Year 2 training and Year 3 training.
Low feature importance variables retained in model Methodology note
There were multiple variables that had relatively low feature importance: Gender, Prior_Conviction_Episodes_PPViolationCharge, Prior_Conviction_Episodes_DomesticViolenceCharges, to name a few. However, because computational speed wasn't a priority f…
Software: Python 3.8 and scikit-learn Methodology note
Preprocessing and model construction were performed on Python 3.8. Preprocessing functions and ML models were imported from the Python library scikit-learn.
Missing values imputed using SimpleImputer Methodology note
Missing values were imputed using the SimpleImputer library.
XGBoost selected as best performing model Finding
In Round 1 of the competition, we analyzed different model performances using 10 fold cross validation and found that on average XGBoost performed the best.
Year 1 model Brier score 0.1837 Statistic
Year 1 model achieved a Brier score of 0.1837 and F1 score of 0.3009.
0.2 Brier score
Year 2 model Brier score 0.1172 Statistic
Year 2 model achieved a Brier score of 0.1172 and F1 score of 0.2512.
0.1 Brier score
Year 3 model Brier score 0.0720 Statistic
Year 3 model achieved a Brier score of 0.0720 and F1 score of 0.0492.
0.1 Brier score
Year 2 model performed very well on hidden dataset Finding
In terms of our performance in the challenge, we found that our Year 2 model performed very well on the hidden dataset when compared with other submissions.
Random Forest performed well but more variable than XGBoost Finding
The Random Forest performed well but had a much more variable performance than the XGBoost. Other regression and classification methods were assessed but performances varied widely, none achieving the success of the XGBoost.
Class imbalance affected all models Finding
All models suffered from heavy class imbalance due to the relative proportion of the non-recidiviated class to the recidivated class. Synthetic augmentation methods like SMOTE were used to mitigate this class imbalance, but did not provide significa…
Optimal probability threshold 0.5 for XGBoost Finding
Different probability thresholds were experimented to optimize XGBoost predictions. With the training and testing data split that was used, optimal brier scores were achieved by using a 0.5 threshold. Increasing or decreasing the threshold leads to …
Early Recidivism feature significantly improved model performance Finding
More specifically the addition of our Early Recidivism feature (Recidivism_Arrest_PrevYear) played a significant role in helping our model performance improve for the later rounds when compared with other submissions.
F1 score recommended for future imbalanced data studies Finding
One metric that could be used in future studies is the F1 score. The F1 score is the harmonic mean of the precision and recall, making it an effective score for delineating the performance of models which were trained on highly imbalanced data.
Top 5 features for Year 1 recidivism prediction Finding
Top five most important features for prediction of recidivism in Year 1: Age_at_Release, Prior_Arrest_Episodes_Felony, Gang_Affiliated, Prison_Years, Prior_Arrest_Episodes_Property.
Top 5 features for Year 2 recidivism prediction Finding
Top five most important features for prediction of recidivism in Year 2: Recidivism_Arrest_PrevYear, Percent_Days_Employed, Jobs_Per_Year, Age_at_Release, Avg_Days_per_DrugTest.
Top 5 features for Year 3 recidivism prediction Finding
Top five most important features for prediction of recidivism in Year 3: Recidivism_Arrest_PrevYear, Percent_Days_Employed, Jobs_Per_Year, Age_at_Release, Avg_Days_per_DrugTest.
Neural Network 10-fold validation Brier score mean 0.1525 Statistic
Neural Network with hyperparameter tuning achieved a mean Brier score of 0.1525 (SD 0.00336) and mean F1 score of 0.7850 (SD 0.0090) in 10-fold validation.
0.2 Brier score
Random Forest 10-fold validation Brier score mean 0.1087 Statistic
Random Forest with hyperparameter tuning achieved a mean Brier score of 0.1087 (SD 0.000126) and mean F1 score of 0.8390 (SD 0.00152) in 10-fold validation.
0.1 Brier score
XGBoost 10-fold validation Brier score mean 0.0945 Statistic
XGBoost with hyperparameter tuning achieved a mean Brier score of 0.0945 (SD 0.000335) and mean F1 score of 0.8559 (SD 0.00128) in 10-fold validation.
0.1 Brier score
Oracle team 1st place Year 2 female parolees Brier score 0.1233 Statistic
In Year 2 female parolees competition, Oracle placed 1st with a Brier score of 0.1233.
0.1 Brier score
MCHawks team 2nd place Year 2 female parolees Brier score 0.1242 Statistic
In Year 2 female parolees competition, MCHawks placed 2nd with a Brier score of 0.1242.
0.1 Brier score
VT-ISE team 3rd place Year 2 female parolees Brier score 0.1260 Statistic
In Year 2 female parolees competition, VT-ISE placed 3rd with a Brier score of 0.1260.
0.1 Brier score
DEAP team 4th place Year 2 female parolees Brier score 0.1263 Statistic
In Year 2 female parolees competition, DEAP placed 4th with a Brier score of 0.1263.
0.1 Brier score
MCHawks team 1st place Year 2 male & female parolees Brier score 0.1405 Statistic
In Year 2 male & female parolees competition, MCHawks placed 1st with a Brier score of 0.1405.
0.1 Brier score
Oracle team 2nd place Year 2 male & female parolees Brier score 0.1451 Statistic
In Year 2 male & female parolees competition, Oracle placed 2nd with a Brier score of 0.1451.
0.1 Brier score
VT-ISE team 3rd place Year 2 male & female parolees Brier score 0.1472 Statistic
In Year 2 male & female parolees competition, VT-ISE placed 3rd with a Brier score of 0.1472.
0.1 Brier score
DEAP team 4th place Year 2 male & female parolees Brier score 0.1481 Statistic
In Year 2 male & female parolees competition, DEAP placed 4th with a Brier score of 0.1481.
0.1 Brier score
NIJ Recidivism Forecasting Challenge goals Policy
With this Challenge, NIJ aims to: 1) encourage 'non-criminal justice' forecasting researchers to compete against more 'traditional' criminal justice forecasting researchers, building upon the current knowledge base while infusing innovative, new per…
Recidivism definition per US Department of Justice Legal fact
Recidivism is defined by the US Department of Justice as 'a person's relapse into criminal behavior, often after the person receives sanctions or undergoes intervention for a previous crime'. It is measured by the criminal acts that result in rearre…
COMPAS system identified as early criminal justice AI tool Finding
One of the earliest tools developed that utilized these new technologies was the Correctional Offender Management Profiling for Alternative Sanctions (COMPAS) system by Northpoint.
Team entered Small Team category of NIJ challenge Methodology note
Our team entered into the Small Team category of the challenge and aimed to utilize state of the art machine learning techniques to assist in this field.
XGBoost parameters: colsample_bytree 0.75 Methodology note
Optimized XGBoost parameter colsample_bytree set to 0.75.
XGBoost parameters: learning_rate 0.1456 Methodology note
Optimized XGBoost parameter learning_rate set to 0.1456.
XGBoost parameters: max_depth 6 Methodology note
Optimized XGBoost parameter max_depth set to 6.
XGBoost parameters: min_child_weight 2 Methodology note
Optimized XGBoost parameter min_child_weight set to 2.
XGBoost parameters: n_estimators 925 Methodology note
Optimized XGBoost parameter n_estimators set to 925.
XGBoost parameters: subsample 0.7 Methodology note
Optimized XGBoost parameter subsample set to 0.7.
Grid search hyperparameter tuning used for XGBoost Methodology note
Parameters of the XGBoost were optimized using grid search hyperparameter tuning. Parameters that were optimized for are number of estimators, learning rate, and max depth of XGBoost trees.
10-fold validation performed for each model Methodology note
For each model trained, 10-fold validation was performed to measure average performance. Metrics captured were brier score and F1 score.
Categorical variables one-hot-encoded, ordinal integer-encoded Methodology note
Variables were split into categorical, ordinal, and numerical. Categorical variables were one-hot-encoded and ordinal variables were integer-encoded. Numerical variables were scaled to be between 0 and 1.
Models can supplement existing recidivism technologies Finding
While the models built suffer from class imbalance, data points with a high probability of being positive for recidivism are likely going to exhibit recidivism in reality. Hence, the models can be used as a supplement to existing technologies in the…
Sources
6 cited sources backing this research.
Primary
Academic
Primary
Official report
Primary
Academic
Primary
Official report
Secondary
Academic
Primary
Academic
Key Entities
Organizations, people, facilities, and other named entities referenced in this research.
COMPAS
[program]
DEAP
[organization]
Georgia Bureau of Investigation
[organization]
Georgia Crime Information Center
[organization]
Georgia Department of Corrections
[organization]
MCHawks
[organization]
National Institute of Justice
[organization]
Northpoint
[organization]
Oracle
[organization]
Prathic Sundararajan
[person]
Suraj Rajendran
[person]
US Department of Justice
[organization]
VT-ISE
[organization]
Related Topics
Research topics that draw on data from this collection.
Parole & Sentencing
Georgia's parole system has contracted to a fraction of its former output: parole releases fell 42% between FY19 and FY24, the Board's overall grant rate hit a record-low 28% in FY24, and only 4.5% of the 2,046 life-sentence cases decided that year ended in release. At the same time, the average time served on a life sentence before release rose from under nine years in 1973 to 29.6 years in FY25. The result is a release regime in which most people now leave prison by serving out their maximum sentence rather than by parole, and in which people die waiting — including people whose release dates were already set.
11,774 data points
Racial Disparities
Georgia's prison population is roughly 58 to 61 percent Black in a state that is roughly 31 to 33 percent Black, and the gap is reproduced at every stage of the system — arrest, plea bargaining, probation revocation, life sentencing, solitary confinement, and exoneration. This page synthesizes 23 GPS research collections documenting the disparity's origins in the Black Codes and convict leasing, its present-day mechanics, and the significant gaps in the data used to measure it.
2,134 data points
Recidivism & Reentry
Georgia reports one of the lowest recidivism rates in the country — an official 25–27% three-year felony reconviction rate — but that figure counts only reconvictions, only within three years, and excludes people who die, who return on technical violations, or who are rearrested without conviction. National data that count arrests find 83% of released state prisoners rearrested within nine years, and GPS's own research library estimates Georgia's real return-to-incarceration rate is closer to 50%. This page tracks what the state measures, what it doesn't, what the evidence says actually reduces recidivism, and how thin Georgia's reentry infrastructure remains relative to the 12,000–16,000 people it releases each year.
12,264 data points
Reform Models & Programs
Georgia operates a thin rehabilitation infrastructure against a deep evidence base: MRT and Thinking for a Change as core cognitive programs, 12 reentry centers with 2,344 beds, and a vocational education budget of $172,000 statewide — $3.44 per person. The programs that do exist show results — Georgia's own vocational completers recidivate at 13.64% against a 26% general rate, and the state's Reasoning and Rehabilitation experiment produced a statistically significant 17% reduction in returns to prison for completers — but completion, staffing, and funding collapse before scale. National models from California, Texas, Maine, Michigan, and Vera's Restoring Promise demonstrate measurable reductions in recidivism and violence; Georgia's own STEP program and the state's audit standards sit unused at the policy floor while the DOJ documents programming 'slashed rather than expanded.'
12,440 data points