Large Language Models Evaluated for Meta-Analysis
Researchers tested four AI models on their capacity to autonomously perform complex meta-analyses.
Updated on Sept. 25, 2026 in Artificial Intelligence

Live Poll
Do you trust artificial intelligence to perform complex scientific data analysis without human oversight?
Researchers evaluated the performance of Claude Sonnet 5, Gemini 3.1 Pro, GPT 5.6, and GPT 5.3 in conducting autonomous meta-analyses. The models were tested on their ability to extract data from ophthalmic literature and calculate key statistical values.
Why it matters
This study examined whether large language models can reliably assist in scientific research by automating time-intensive meta-analysis tasks. The findings highlight current capabilities and limitations of AI in reproducing standardized statistical results.
GPT 5.6 correctly extracted 87.5% of cells and reproduced 85.0% of 2x2 tables, while Claude Sonnet 5 achieved 85.0% cell accuracy. In contrast, Gemini 3.1 Pro matched 40.0% of pooled odds ratios and GPT 5.3 matched 20.0% of pooled odds ratios.
The players
GPT 5.6
This is a large language model developed by OpenAI that was included in the comparative analysis.
Claude Sonnet 5
This is a large language model developed by Anthropic that was evaluated for its meta-analysis performance.
Gemini 3.1 Pro
This is a large language model developed by Google that was tested against statistical reference standards.
GPT 5.3
This is an earlier version of an OpenAI large language model included as a comparator in the study.
The details
The models were provided with primary study PDFs and prompts to extract specific data for five meta-analyses. Researchers assessed the performance of these AI tools against established statistical reference standards to determine the accuracy of pooled results and I2 values.
Timeline
The evaluation results were published on September 25, 2026.
The Tech Race
The emergence of AI-driven meta-analysis tools marks a significant shift in scientific methodology, potentially replacing manual data extraction workflows. This study follows the pattern set by the Cochrane Handbook for Systematic Reviews of Interventions standards to benchmark AI performance.
For researchers and data analysts, these findings suggest that while AI can significantly accelerate data extraction, human verification remains essential for high-stakes statistical output. Users should monitor updates as these models improve in their ability to match complex pooled ratios.
The takeaway
The study demonstrates that current large language models show varying levels of proficiency in executing autonomous statistical synthesis. Researchers should treat AI-generated meta-analyses as supportive tools that still require rigorous human oversight to ensure data integrity.
Further reading
For more information on the evolving role of machine learning in research, visit the Artificial Intelligence section.
Live Poll
Do you trust artificial intelligence to perform complex scientific data analysis without human oversight?







