I work on AI evaluation for knowledge work, mostly in life sciences and medical devices — also energy, legal, and finance.
I run Raycaster, which builds benchmarks and environments for that work. I also support evalscience.org, a nonprofit focused on the science of agent evaluation.
Friends know me as an existential pragmatic with a daoist metaphysical temperament and a tragic-humanist sense of life.
Papers
Designing Benchmarks for Knowledge Work
Earlier
- Using POMDPs to Learn Language in a Spatial Reference Game
- Fair Machine: Human Conceptions of Algorithmic Fairness
- Reimagining the Abstract Image with GANs
- Accountable: Achieving Goals with Accountability
- Real-time Acoustic Modeling with Convolutional Neural Networks
- Toxic Comment Classification with Unintended Bias