Language models write fluently, but they do not work the way experts do. When we compared lawyers' recorded workflows with an LLM's on the same legal task, the lawyers followed the one branch that mattered and revised their plan as the case developed; the model expanded every branch it could find.
Most models learn only from finished text. We also record the process behind it, then use those traces to train models, test them against people, and build tools that fit how experts already work.
Three connected pillars. Data and models from the first feed the tools in the second; the third asks whose perspectives all of it represents. Pick a topic to see the projects behind it.
Writing a paper takes months of planning, drafting, abandoning, and revising, yet a model trained on the final PDF sees none of that. We record that process, use it as training signal, and borrow methods from cognitive science to measure where models still differ from people.
In science and law, full automation tends to flatten the reasoning that makes expert work valuable. We build tools that help at the moment of difficulty, adapt to how a particular expert works, and leave the judgment with the person.
People disagree, and their perspectives differ by background, identity, and culture; a model trained toward an average answer can erase that. I founded the Pluralistic Alignment workshop (NeurIPS 2024; second edition at ICML 2026) to work on this with ML, HCI, and social-science researchers.
The work draws on linguistics, cognitive science, and the social sciences, with collaborators in computer science, law, psychology, education, journalism, design, and medicine. To connect these communities I co-organized the first CtrlGen workshop on controllable generative modeling (NeurIPS 2021), founded the In2Writing workshop series on intelligent and interactive writing assistants (ACL 2022, CHI 2023, CHI 2024), and founded the Pluralistic Alignment workshop (NeurIPS 2024).