Decision science under uncertainty

How can we make better decisions when information is highly incomplete?

I build decision systems where human judgment and algorithmic inference reinforce each other, and methods for testing whether those systems actually reason well. My work connects decision theory to practical questions about how people and AI systems should reason from incomplete information.

Portrait of Jeff Helzner

Bayesian modeling (Stan, PyMC) · LLM decision evaluation · Formal methods (Lean 4) · Python, R, SQL

About

I'm a decision scientist with over a decade of experience building data systems, probabilistic models, and decision tools in the insurance industry. Before that, after earning my PhD at Carnegie Mellon, I was a philosophy professor at Columbia teaching and researching the foundations of decision theory, formal epistemology, and logic. Both careers have been shaped by a single question: how should decision makers—whether human or algorithmic—form beliefs, choose actions, and revise their thinking in light of evidence?

The bridge between them was noise. In 2014 I left Columbia to found one of the insurance industry's first behavioral science teams, where I worked closely with Daniel Kahneman on reducing unwanted variability in expert judgment. That work turned out to be the same problem I had been studying formally, now with adjusters and underwriters in place of idealized agents.

Since then I have focused on designing systems where human expertise and machine intelligence complement each other: Bayesian models that blend actuarial benchmarks with underwriter judgment, monitoring that surfaces unexpected patterns for human review, and AI-powered tools that augment—rather than replace—the judgment of claims adjusters and underwriters.

Selected research

This work extends my industry focus on human–machine decision systems to AI systems themselves — most recently to Ellsberg-style ambiguity problems, the same territory as my 2014 work on risky and ambiguous gambles, approached now with a Bayesian measurement framework and an LLM in the decision seat. I tend to reach for tools that resonate with my background in logic and formal epistemology, drawing on both probabilistic modeling and formal proof. The aim is to take a claim like "this system reasons well under uncertainty" and turn it into something you can test, and then improve.

Working paper · 2026

Sensitivity to Subjective Expected Utility Maximization: A Methodological Study, with an Illustrative Application to LLM Decision-Making

A framework for estimating how strongly a decision maker's choices respond to differences in expected utility, together with the adequacy checks that constrain when such an estimate can be meaningfully interpreted. Illustrated on LLM decision-making under varying sampling temperature.

arXiv:2607.11920 Decision making Rationality

Selected publications

  1. The Description–Experience Gap in Risky and Ambiguous Gambles

    with Varun Dutt, Horacio Arló-Costa, and Cleotilde Gonzalez

    Journal of Behavioral Decision Making 27 (4): 316–327. 2014.

  2. Rationalizing two-tiered choice functions through conditional choice

    Synthese 190 (6): 929–951. 2013.

  3. On the Application of Multiattribute Utility Theory to Models of Choice

    Theory and Decision 66 (4): 301–315. 2009.

  4. Transfer principles in nonstandard intuitionistic arithmetic

    with Jeremy Avigad

    Archive for Mathematical Logic 41 (6): 581–602. 2002.

Full publication list on PhilPeople

Projects

SEU Sensitivity

A Bayesian framework modeling sensitivity to Subjective Expected Utility maximization, with a full workflow from prior calibration through model adequacy checks.

Source Stan Bayesian workflow

Conditional Admissibility Judgments

Technical reports and a Lean 4 formalization of conditional judgments of admissibility under uncertainty, pairing machine-checked proof with applied decision theory.

Source Lean 4 Formalization

Experience

Across four organizations I have built decision-science capability from zero: founding teams, platforms, and data infrastructure where none existed. Fully remote since 2020, including team leadership.

  1. 2024 — present

    Director of Data Strategy & Analytics

    Berkley Enterprise Risk Solutions (a W. R. Berkley company)

    A California-based startup writing large workers' compensation policies. I lead the data systems, probabilistic models, and AI tools behind underwriting and claims decisions — including five Bayesian and machine learning models in production, automated monitoring that surfaces anomalous shifts in claim frequency and severity for human review, and LLM-powered drafting tools adopted broadly by claims adjusters.

  2. 2020 — 2024

    Decision Scientist

    Joyn Insurance

    At an early-stage insurtech, helped build a data-driven underwriting platform from the ground up, embedding the integration of model output with underwriter judgment into both the technical infrastructure and the culture.

  3. 2018 — 2020

    Decision Scientist, Blackboard Insurance

    AIG

    At AIG's insurtech subsidiary, developed methods for extracting behavioral insight from operational data, and worked with R&D on NLP and OCR pipelines for structuring unstructured documents.

  4. 2014 — 2018

    Head of Behavioral Science

    AIG

    Founded and led a five-person team including one PhD and two MS-level statisticians — among the first behavioral science teams in the insurance industry. Worked closely with Daniel Kahneman on noise-reduction studies that demonstrably reduced variance in adjuster case-reserve estimates. Presented findings to the CEO and Chief Underwriting Officer, and represented the team at board meetings.

  5. 2003 — 2014

    Assistant, then Associate Professor of Philosophy

    Columbia University

    Taught undergraduate and graduate courses in logic, probability, decision theory, and the philosophy of science. Research focused on the foundations of decision theory under uncertainty.

Education