Science

AI in drug discovery faces data quality and bias challenges

2 min read

AI in drug discovery faces data quality and bias challenges
Photo: Logan Gutierrez · Unsplash
0 0
XWhatsAppTelegramLinkedIn

The high cost and long timelines of drug development are driving pharmaceutical companies to adopt artificial intelligence (AI) to improve success rates and reduce risks. AI is being used to design drug candidates from scratch and predict their interactions with disease targets, replacing traditional physical screening of vast molecular libraries. This shift allows companies to identify promising compounds faster and eliminate low-quality candidates before lab testing, saving time and resources.

However, AI-generated candidates still require laboratory validation because current models cannot reliably predict kinetics or developability. This has increased pressure on lab teams, who must test and characterize a growing volume of diverse, AI-generated compounds. Traditional screening workflows, designed for binary yes-or-no results, struggle to provide the detailed data needed for these complex candidates.

A major challenge is the quality and completeness of training data. Many AI models rely on public datasets that lack structure, labeling, and diversity, leading to a "data wall" where models reach similar conclusions with diminishing returns. Publication bias exacerbates this, as most public data focuses on positive results, while negative data—failed experiments and non-binding compounds—remains unpublished. Paul Belcher, director of protein research strategy at Cytiva, notes that this bias limits models' ability to learn from failures and make reliable predictions.

Data integrity is another concern. Fabrication of images, such as Western blots, has become easier with generative AI. Research by microbiologist Elisabeth Bik found that almost 4% of biomedical papers contained duplicated or manipulated images as of 2016. Belcher warns that manipulated data used to train AI could have disastrous consequences. Some vendors are developing tools to verify data authenticity, such as Cytiva's Image Integrity Checker, which uses secure hash algorithms to detect tampering. Publishing houses are showing interest in adopting such tools as standard practice.

Looking ahead, the industry must address these data challenges to fully realize AI's potential in drug discovery. This includes creating more comprehensive datasets that include negative results, improving data labeling and structure, and implementing verification tools to ensure data integrity. Without these steps, AI models risk perpetuating biases and producing unreliable predictions, undermining their promise to accelerate drug development.

Sources

Report / request removal

Related

Comments

No comments yet. Be the first.