When researchers screen candidate tuberculosis drugs, they often have too many options. Some look promising but later turn out to be costly dead ends.
“You might get thousands of compounds from a screen, but then you’ll have to decide which compounds to work on,” said James Sacchettini, Ph.D., Texas A&M AgriLife research scientist, professor in the Texas A&M College of Agriculture and Life Sciences, Department of Chemistry and Biophysics, College of Arts and Sciences, Department of Chemistry, and Roger J. Wolf Welch Foundation Scientific Chair.
His lab recently built an artificial intelligence tool that allows scientists to focus on research after the initial screening. The team also uses AI to organize years of collaborative data into a searchable format.
What information really helps us make decisions? It would be great if we could use AI to shorten the time from idea to actual treatment. ”
James Sacchettini, Ph.D., Professor, Department of Chemistry and Biophysics, Texas A&M College of Agriculture and Life Sciences, and Department of Chemistry, College of Arts and Sciences
persistent enemy
The lab’s work with AI is even more important because of the problems it aims to solve. According to the World Health Organization, tuberculosis is the world’s deadliest infectious disease. It has been with humanity for thousands of years. Standard treatment takes several months, and even longer if the disease involves drug-resistant strains or co-infection with HIV.
Dr. James Sacchettini (centre) and researchers Saswati Panda (left) and Siddhant Rath (right) are developing new AI tools for tuberculosis drug discovery.
Most affected people live in low-income areas, where long treatment times and limited care make the disease particularly difficult to control.
In the United States, the infectious disease outbreak that occurred in New York in the 1990s served as a wake-up call.
“People thought, ‘Oh, we fixed that years ago,'” he says. “Then it turned out that Rikers Island was completely covered in tuberculosis. People were coming out of jail, getting into elevators with six other people, and by the time they got to the sixth floor, five other people were infected.”
Finding better treatments presents many challenges. Bacteria are coated with a thick waxy coating that prevents most drugs from reaching their intracellular targets. Bacteria also grow slowly, so experiments take time.
“Experiments for tuberculosis can take months, but for staphylococcus and streptococcus it can take a week,” Sacchettini said. “This is part of why the drug discovery pipeline is relatively slow. This is a perfect area to work with AI.”
Build the foundation first
Before AI tools, the problem was much simpler. Research data in academic laboratories is often stored on network drives, slide decks, and in the memories of the people who performed the experiments.
Sacchettini’s lab built DAIKON, an open-source platform launched in 2023, to track drug targets from genes to years of chemical research in one place. The Gates Foundation-backed Tuberculosis Drug Accelerator (TBDA) uses DAIKON across its laboratory and industry partnerships. New tools are directly connected, including the latest AI systems from the Sacchettini Institute.
“We don’t want AI to give us exactly the right answer. But it tells us what not to work on. And it tells us what to work on. And that’s a really big time saver.”
Reducing false positives in early screening
During the early stages of drug development, researchers test numerous compounds against the protein of interest. Some appear to be working. Some things interfere with the test itself.
“These ‘nuisance elements’ have cost us a lot of time,” Sacchettini said. “One of the goals is to identify them so that you don’t have to spend months, years, or even hundreds of thousands of dollars working on them.”
To fix this, Sacchettini’s team developed an AI model that identifies these false signals.
The team’s model, called CAGE-Fusion, learns from publicly available screening data and categorizes compounds into four types of trouble. That is, compounds that aggregate, compounds that fool the chemical signals of the test, compounds that react instead of binding, and compounds that attach to many targets instead of one. The published research received funding from the Gates Foundation and the Welch Foundation.
Given one nuisance compound and one clean compound, the model ranks the nuisance compound as more suspicious about 94% of the time. Some categories are better than others. Compounds that react are the easiest to capture, and compounds that attach to many targets are the most difficult to capture. Inside DAIKON, a model is automatically run based on the test data it receives, flagging potential problems before the compound reaches more expensive stages.
Chemists do not have to take results at face value.
“This model can walk you through the process and show you which areas of the molecule are likely to be at fault,” said Siddhant Rath, an AgriLife Research scientist in the Sacchettini lab who led the effort.
Early identification of misleading compounds is just one piece of the puzzle. Drug discovery spans many steps across multiple institutions. The Sacchettini lab is also leveraging AI to help TBDA better leverage shared knowledge.
Turning mountains of results into shared knowledge
Rath and colleague Saswati Panda, also in Sacchettini’s lab, have developed an AI system to categorize TBDA data to make it easier to use. This project received funding from the Gates Foundation, the National Institutes of Health, and the Welch Foundation.
The consortium’s documents over the years include many molecular structures.
“You might come across a molecule in your research and think, ‘I’ve seen this before,’ but there was no easy way to find it,” Sacchettini said.
This system collects TBDA data throughout the drug discovery pipeline. Now you can visually track molecules throughout your project, even where you get stuck. Researchers can query data through a chat interface.
“It tells you in seconds who presented something, what they said about it, and it gives you the presentation so you can look at the slides,” Sacchettini said.
New tools and old diseases
In 2026, scientists will have access to processing power that was previously unavailable. This has allowed AI to play a larger role in research, Russ and Panda said.
Sacchettini said scientists around the world are implementing AI tools at various stages of drug design. Models created by his lab are already helping to ease the process.
“We don’t want AI to give us exact answers,” he says. “But it tells us what not to work on. And it tells us what we should work on. And it’s a really big time saver.”
sauce:
Texas A&M AgriLife Communications
Reference magazines:
Russ, S. and others. (2026). Deep learning for assay nuisance compound detection using gated co-attention graph embedding model (CAGE-Fusion). Chemoinformatics magazine. DOI: 10.1186/s13321-026-01207-4. https://link.springer.com/article/10.1186/s13321-026-01207-4

