PhD Proposal: Towards Validity and Fairness in AI-Assisted Assessment Development
IRB-5111 https://umd.zoom.us/my/rudinger
Traditional assessment development is costly and time-consuming because it requires multiple rounds of review, piloting, and refinement. We use AI-assisted assessment development to refer to the use of LLMs to augment or fully automate any stage of the assessment development process. However, LLMs generate highly polished and plausible-sounding text, which does not necessarily suggest high quality. We therefore expect that these AI-assisted assessment developments be structured around an evidence-centered design (ECD) to establish psychometric validity and fairness. This thesis asks which LLM outputs can provide defensible evidence for pre-pilot mathematics item development in an effort to meet the rigorous standards of educational measurement.
The completed work identifies two problems that must be addressed before LLMs can provide defensible evidence for assessment development: cultural inconsistency and unreliable simulation of student reasoning. First, LLMs responded less reliably from Ghanaian than from United States cultural perspectives, particularly when the intended culture was underspecified. Also, simulated classroom correctness rates correlated with the proportions of real NAEP students answering mathematics items correctly, reaching ρ = 0.82, but the simulations were less reliable in reproducing students’ incorrect answer choices and reasoning. These findings show that LLMs may support preliminary screening of relative item difficulty, but matching aggregate outcomes does not establish that the models reproduce the intended reasoning processes or perform consistently across cultural contexts. This problem is relevant to the broader proposal because assessment developers need evidence about item outcomes, student reasoning, and cultural fairness before using LLM simulations to make pre-pilot decisions.
The proposed work examines when LLM-generated evidence is sufficiently valid and fair to support pre-pilot mathematics item development. The first proposed study will test whether persona-conditioned simulated respondents produce reasoning that is mathematically consistent with their selected answers and whether their errors align with documented student errors and distractor-selection patterns. The second proposed study will treat cultural adaptation as a controlled measurement transformation and use a “detect-diagnose-decompose” framework to determine whether estimated difficulty changes reflect the adapted item or model-family bias. We expect to identify which properties, including relative difficulty, reasoning consistency, error patterns, and cultural stability, can be defensibly pre-screened with LLMs and which still require evidence from real students. The contribution is an evidence-centered foundation for using and evaluating LLMs in assessment development while preserving validity and fairness across diverse educational contexts.