LSE Statistics PhD Reading Group

Logo

A super simple site to organize meetings for our reading group

View My GitHub Profile

LLM Benchmarking via Representation Multi-task Learning

Large language models (LLM) are commonly evaluated using benchmark scores aggregated across items and domains, but such summaries may obscure differences in item difficulty, domain structure, and cross-domain dependence. We develop a statistical framework for measuring both general and domain-specific LLM capabilities from item-level benchmark responses. Our approach combines item response theory with multi-task learning, allowing information to be shared across related domains while retaining domain-specific heterogeneity. We propose a regularized estimator that adapts to the degree of similarity across domains, and show that it is minimax optimality under suitable regimes. Simulations demonstrate the benefit of adaptive information sharing. An application to the MMLU benchmark further reveals a strong common capability component alongside substantial domain-specific variation across LLMs.