Skip to Content

Cohan wins NSF CAREER Award to build more reliable AI for science

For his work on building more trustworthy and capable AI tools for science, Arman Cohan has won a Faculty Early Career Development (CAREER) Award from the National Science Foundation (NSF).

Cohan, assistant professor of computer science, will use the five-year, $600,000 grant to develop stronger foundations for using large language models (LLMs), the technology behind modern AI assistants and agents, as reliable aids in scientific research. The NSF CAREER Award is a prestigious honor for young faculty members and supports the early career activities of teachers and scholars who are considered most likely to become the academic leaders of the future.

Large language models have already shown real promise as tools for scientific work. They help researchers sift through enormous bodies of literature, summarize prior findings, answer technical questions, and generate new hypotheses. But these systems come with significant limitations in high-stakes scientific settings. They can overlook key publications, draw on weak or irrelevant evidence, express unwarranted confidence in uncertain claims, or produce answers that sound authoritative but are not supported by current scientific evidence.

“Researchers are already using AI agents to help with scientific work: searching the literature, analyzing evidence, using tools, running experiments, and reasoning through difficult research questions,” Cohan said. “That creates enormous opportunities to accelerate science. But in scientific settings, reliability and validity are critical. If these systems give misleading answers, cite evidence that does not actually support the task, overlook important prior work, or fail in ways that are hard to detect, they can send research in wrong directions. This project is about building the foundations needed to make AI systems more trustworthy, evidence-grounded, and useful for science.”

Cohan addresses these challenges along three fronts. The first focuses on evaluation – developing new frameworks to rigorously measure how well AI models perform on scientific tasks, including their ability to retrieve relevant evidence, reason through challenging problems, and produce detailed analyses. The second develops new methods for adapting LLMs to the particular demands of scientific domains, drawing on the rich structure of scientific literature and experimental data. The third focuses on reliability and interpretability: developing methods to better understand where AI systems succeed or fail on scientific tasks, make their outputs more understandable and transparent, and generally improve their trustworthiness in research settings.

The research focuses on core LLM methodologies while having broad applications across scientific disciplines, including computer science, medicine, biology, and physics – fields where scientific knowledge is expanding rapidly and the consequences of error can be high.

A key emphasis of the project is openness and accessibility. Cohan's team will release publicly available datasets, benchmarks, software, and tutorials designed to help researchers across institutions use these technologies more effectively and responsibly.

“The challenges we're addressing aren't unique to any one scientific field,” he said. “As researchers increasingly use AI systems in their work, we need to understand when these systems can be trusted, where they fail, and how to make them more reliable and interpretable. Our goal is to develop core AI methods that can support science broadly, while training students to use and build these systems responsibly.”

The project also includes an educational component with plans to mentor graduate and undergraduate students through hands-on research experiences, contributing to a scientifically literate workforce equipped to work responsibly at the intersection of AI and science.

More Details

Published Date

Jun 17, 2026

Featured Departments

Areas of Impact