Case Study: GPT-5 Can Assist but Not Replace Statisticians in Research
This case study evaluates ChatGPT-5 and ScholarAI on real statistical research tasks—literature review, gap identification, and R code generation—focused on dynamic treatment regime estimation via dynamic weighted ordinary least squares. The authors find current GenAI models lack the depth and contextual understanding to complete these tasks unsupervised, though they can boost efficiency in search, summarization, and basic code debugging when guided by a knowledgeable researcher. Their conclusion: GenAI works as a research tool under expert supervision, not as a substitute for methodological expertise. On Twitter, commentary highlighted the study's practical approach of testing GenAI across the full research pipeline against actual peer-reviewed benchmarks, with the takeaway that individual tasks can be sped up but full automation without expert oversight remains premature.
Discussion: 1 tweets from 1 authors · @TJO_datasci