Xiao Chen

Educational data science, measurement, and human-centered AI.
Seattle, WA

Portrait of Xiao Chen

For three and a half years I built AI decision-support systems inside a regulated pharmacy workflow, where a wrong answer has consequences and a pharmacist makes the final call. In parallel, I worked on quantitative education and psychology studies using PISA and other large-scale survey data. Both lines led me to the same problem: the things we most want to measure — self-regulation, comprehension, attention — are usually captured by a single self-report item, and I wanted to see what a record of actual behavior would show instead.

Research

Measurement beyond self-report

Surveys ask people to summarize behavior they may not remember accurately. I work on measuring the same constructs from what people actually do — response logs, navigation traces, device use — and on testing whether those measures mean what we say they mean.

Evaluating AI where the answer matters

When a system’s output changes what a person does, accuracy is not the whole question. I’m interested in how you define success when the target is a latent construct, how failures propagate through systems with several components, and when a system should hand a decision back to a person.

Building the instruments

Most of the measurement questions I care about need data that doesn’t exist yet. I build the collection and analysis systems that produce it, and I treat the decisions inside them — what counts as active use, what happens to an unanswered prompt — as part of the research rather than implementation detail.

Read more about my research

Selected work

PhoneMood Today screen showing recorded phone use, mood ratings, and check-in counts, using synthetic demonstration data

PhoneMood

2026 · Android · sole author

An Android app that records phone and app use and asks for a mood rating close to the moment of use. A direct follow-up to the PISA 2022 study below, built to measure what a once-a-year questionnaire item could not.

GitHub repository
Prescription verification screen listing specific mismatches found between the label, the order, and the dispensed product, using demonstration data

Pharmacy verification copilot

2022–2026 · Amazon Health

A multimodal, multi-agent system that cross-checks prescription images, extracted text, order data, and database records before a pharmacist reviews them. 3,000+ images a day; pharmacists keep final authority.

Pharmacy workflow platform home screen showing fill, verify, and pack task entry points, using demonstration data

Pharmacy verification platform

2022–2026 · Amazon Health

The review system pharmacists actually work in: structured decisions, auditable states, and a workflow shaped by iterative research with the people using it.

All projects

Publications

Rongxiu Wu, Xiao Chen

Predicting STEM Career Interest at the End of High School: A Machine Learning Approach to In-School and Out-of-School Time Experiences

Under review

Rongxiu Wu, Xiao Chen, Rong Su

Problems With Self-Efficacy Across Educational Systems: Gender Differences and School Contexts in PISA 2022

Current PsychologyUnder review
All publications

About

I studied mathematics and economics at Boston University, then applied data science at USC. As an undergraduate I worked on a family-planning field study in Malawi, which is where I first handled survey data that came from talking to people rather than from a repository. I spent the years after that writing production software, and I’ve been working my way back toward research since.