What I work on

I conduct generative AI research for Meta on contract, evaluating frontier LLM reasoning across mathematics and other STEM domains. The work includes examining correctness, assumptions, proof validity, ambiguity, and alignment with the intended task.

I partner with researchers on the questions worth asking, how to structure evaluation workflows, and which findings should guide the next stage of the work.

Evaluation is a research problem

A useful evaluation needs a well-defined task, a sound reference, and criteria that distinguish a real error from a different valid approach. I review prompts, rubrics, and reference solutions with those distinctions in mind.

  • Does the reasoning support the conclusion?
  • Which assumptions are stated, implicit, or missing?
  • Is the reference solution correct and sufficiently clear?
  • Does the evaluation measure the ability the task was intended to test?

From findings to better data

My role also involves translating quality findings into recommendations for external data vendors and annotation workflows. I contribute to early-stage research on diagnosing and repairing annotation issues, checking proposed changes, preserving the original task intent, and recognizing when human review is needed.

A connection to model risk

My earlier work at JPMorgan involved challenging quantitative models and documenting their limitations. That experience informs how I approach AI evaluation: make the assumptions explicit, investigate failure modes, and be precise about what the evidence supports.

Professional backgroundDiscuss a research question