By Stephanie L. Fleming, Ph.D., MS, and Janna M. Broaddus
Corrections staff have always had to explain discretionary calls. A classification officer who places someone in a more restrictive unit than a colleague might have chosen has always been expected to justify that decision if asked. That much isn’t new.
What is new is the kind of input staff are now working with. When a colleague or supervisor made a recommendation, staff could always ask why and get an answer grounded in a person’s actual reasoning. When an AI-supported tool generates a risk score or a recommendation, that option is often more limited: many tools don’t surface a full account of their own reasoning, and even where a platform does offer audit logs or documentation features, those logs typically capture what the tool recommended, not what the staff member who acted on it actually weighed. Most agencies haven’t built a habit of writing down that second part. That gap, not the discretion itself, is the actual operational problem, and it’s specific enough to fix with a specific set of steps.
Below is a walkthrough of how a composite facility — call it Greenriver County Correctional Facility, a 400-bed county jail — could realistically build this practice over four months, without new software and without slowing staff down. Greenriver isn’t a single real site; the details below reflect patterns we’ve seen across multiple facility rollouts, compressed into one narrative so the sequence is easier to follow.
Worth noting up front: this practice doesn’t depend on which AI platform a facility uses, or how good that platform’s own logging is. Even a vendor tool with strong audit logs and documentation features is logging the tool’s side of the interaction. The agency still needs its own process for capturing what staff actually did with that output, because no vendor system was built to track a specific officer’s reasoning at a specific facility. The four-field note below is meant to sit alongside whatever the platform already provides, not replace or compete with it.
Greenriver’s classification unit processes roughly 30 intakes a month. Historically, four to six of those involve a staff member overriding the AI-supported risk score, a number the facility didn’t actually know until it started counting.
Month one: Pick one decision point and build a one-page note
Trying to document every AI-touched decision agency-wide on day one guarantees the effort collapses. Start with a single decision point where an AI-supported tool is already in use — classification is a common starting place — and build a one-page note tied specifically to that.
The note needs four fields, nothing more. What did the tool recommend? What did the staff member decide? What specific fact, observation, or concern drove any difference between the two? Who reviewed or countersigned it, if anyone did at the time? Each field takes a sentence or two, filled out at the point of decision, not reconstructed later. A staff member should be able to complete it in under two minutes.
This isn’t a new form buried in a binder. It gets attached directly to the same file the AI recommendation already lives in, so the two are read together from day one.
Here’s what a completed note looked like at Greenriver during its first week:
- Tool recommendation: Low-risk, general-population-suitable.
- Staff decision: Placed in restrictive housing pending further review.
- Reasoning: A disciplinary note from the prior facility flagged an unresolved incident three weeks before transfer, and intake staff observed elevated agitation that was inconsistent with the detainee’s stated history.
- Countersigned by: Shift supervisor, same day.
Four lines, written in under two minutes and legible enough that anyone reading the file six months later understands exactly what happened and why.
Whoever leads this first month, often a shift supervisor or a classification unit lead, should also decide up front what happens to a note once it’s written. Does it get scanned into the same electronic file as the AI output? Does a paper copy go in a physical folder pending a records system update? The answer matters less than having one, because a note with no defined home tends to disappear within a few weeks, however good the intentions were when it was designed.
Month two: Pilot with one unit, then listen
Roll the note out with a single shift or unit rather than the whole facility. This does two things. It surfaces friction early: staff may find a field confusing or find the note redundant with something else they’re already filling out, and it’s far easier to fix that with 15 people than with the whole roster. It also gives supervisors a small, manageable set of notes to actually read and react to, rather than a flood they’ll skim or ignore.
During this month, a shift supervisor should read every note completed, not to audit staff, but to catch confusion about the form itself before it hardens into bad habits. Adjust the wording of the fields based on what staff actually write. If the “what drove the difference” field keeps getting one-word answers, that’s a sign staff need a prompt or an example, not a stricter policy.
Expect some resistance in this month, and expect it to be reasonable. Staff who have never been asked to explain a judgment call in writing may read the note as a sign they’re being second-guessed, especially if a similar-sounding form has been used punitively somewhere else in their career. The way to work through this is direct: the note exists to protect a well-reasoned decision, not to flag it for discipline, and the clearest way to demonstrate that is for supervisors to say so out loud in roll call or shift briefings, not just in a memo nobody reads closely. A pilot that skips this conversation tends to produce vague, defensive notes; one that has it tends to produce notes staff are actually willing to stand behind.
At Greenriver, this played out directly during the second week of the pilot. Two officers on the same shift began writing notes so brief they were nearly useless, something closer to “gut feeling” than an actual account. A supervisor pulled both aside individually rather than addressing it at roll call, and learned that a prior “explain your decision” form at another facility years earlier had been used to build a disciplinary case against a coworker. Once the supervisor explained plainly that these notes were reviewed for patterns, not individual performance, and that a well-reasoned override was something the agency wanted on record, not something to be punished, both officers’ notes became substantially more detailed within the week. The fix wasn’t a policy change; it was one direct conversation addressing a specific, reasonable fear.
Month three: Add supervisory review, on a schedule
Once the note is working at the unit level, add a standing monthly review. A supervisor pulls every note from the past month for that unit and looks for two things. First, the supervisor examines whether overrides are clustering. Against Greenriver’s baseline of four to six a month, a sudden jump to 12 tied to one officer or one shift is a signal worth a conversation, while the same four to six spread evenly across the roster looks more like ordinary judgment at work. Second, the supervisor determines whether the reasoning captured actually holds up: Is it specific enough that someone outside the room could understand the decision, or is it vague enough that it wouldn’t survive being read back in a deposition?
This is the point where the practice starts protecting staff as much as the agency. A pattern of well-reasoned overrides, documented consistently, is evidence of sound judgment. The same pattern with no documentation looks like inconsistency by default, whether or not it actually was.
The review itself doesn’t need to be elaborate. Thirty minutes once a month, with the notes from that unit in hand, is enough for a supervisor to spot a pattern worth a conversation. What matters more than the length of the review is that it happens on a fixed schedule rather than only when something has already gone wrong, because a review that only ever gets triggered by an incident will always look, in hindsight, like it was reacting to that incident rather than preventing the next one. A facility that can point to a running log of routine monthly reviews, most of which found nothing notable, is in a far stronger position during an audit than one that can only produce a review from the week after a lawsuit was filed.
Month four: write the policy, then expand
Only after the pilot has run for a few months should the agency write a formal policy statement covering when overriding an AI-supported recommendation is expected, when it’s discouraged, and when it needs a second signature. Writing this too early means guessing at problems instead of addressing the ones that actually surfaced during the pilot.
Once the policy and the note are both proven at the unit level, expand to a second decision point — a housing determination or a program eligibility screen, for instance — using the same four-field note as a starting template, adjusted for what that specific decision requires.
Common pitfalls worth avoiding
A few mistakes show up often enough in early rollouts to flag directly. The first is over-designing the note before ever piloting it, adding fields for edge cases that haven’t come up yet, until it’s long enough that staff start skipping fields or filling them in after the fact from memory. Four fields is a starting point precisely because it’s short enough to survive contact with a busy shift.
The second is skipping the pilot and rolling out facility-wide immediately, usually out of a sense that the practice needs to look comprehensive from day one to satisfy an auditor or a board. In practice, a facility-wide rollout with no pilot tends to produce wildly inconsistent notes, because there was never a feedback loop to catch confusing wording before it spread everywhere at once.
The third is treating the monthly review as a compliance checkbox rather than an actual read-through. A supervisor who signs off on a stack of notes without reading them defeats the purpose as thoroughly as not collecting the notes at all, and it tends to surface at the worst possible moment, when an auditor asks a supervisor to explain a specific note and the supervisor has never actually seen it before.
What this produces, six months in
By month six, a facility running this practice has something it didn’t have before: a retrievable, timestamped record of staff reasoning that sits alongside every AI-supported recommendation it acted on differently. That record is what turns up in an accreditation review, a PREA audit, or a discovery request, produced in minutes rather than reconstructed from memory months or years after the fact.
It also produces something less tangible but just as valuable: a shared, agency-wide understanding of how staff are actually using these tools day to day, rather than relying on the tool’s validation study alone. A classification supervisor who reviews four months of notes learns things about how the tool performs on this facility’s specific population that a vendor study, however thorough, is unlikely to capture, because that study wasn’t built to track this facility’s staff decisions, and that knowledge only exists because someone wrote it down consistently enough to notice the pattern.
This kind of practice is part of a broader discipline some in the field are calling Justice Decision Observability, treating the human judgment layered on top of AI-supported tools as something worth documenting in its own right, the same way the tools themselves are already validated and audited. But the label matters less than the habit. An agency that builds the four-field note, the unit pilot, and the monthly review described above has the substance of it already, whatever it’s called.
References
- American Correctional Association (ACA) accreditation standards
- Commission on Accreditation for Law Enforcement Agencies (CALEA) standards
- National Commission on Correctional Health Care (NCCHC) standards
- Prison Rape Elimination Act (PREA) audit standards
About the authors
Stephanie L. Fleming, PhD, MS, is founder and principal of Justice Beacon Solutions, an independent operational governance and documentation firm serving the justice and public safety ecosystem. She is the co-creator of Justice Decision Observability (JDO), an operational governance discipline focused on documenting and reconstructing how consequential human decisions are made in AI-supported and traditional operational environments. Dr. Fleming has been featured on podcasts discussing AI governance, operational accountability, and decision documentation in corrections and public safety. She is based in Indianapolis, Indiana, and can be reached at stephanie@justicebeacon.com.
Janna M. Broaddus is co-creator of Justice Decision Observability (JDO) and director of operations at Justice Beacon Solutions. She leads operational strategy, client engagement, and implementation planning while helping organizations strengthen decision visibility and governance documentation across justice and public safety operations.
Justice Beacon Solutions partners with agencies seeking to better understand, document, and strengthen how consequential decisions are made in practice. The firm supports organizations before, during, and after significant operational events by reconstructing decision pathways, documenting governance processes, and helping agencies demonstrate how authority was exercised. Rather than auditing technology or evaluating policy compliance, Justice Beacon Solutions focuses on making human decision-making observable, explainable, and defensible when organizations are required to account for their actions.