AI in drug discovery faces data quality and bias challenges
The high cost and long timelines of drug development are driving pharmaceutical companies to adopt artificial intelligence (AI) to improve success rates and reduce risks. AI is being used to design drug candidates from scratch and predict their interactions with disease targets, replacing traditional physical screening of vast molecular libraries. This shift allows companies to identify promising compounds faster and eliminate low-quality candidates before lab testing, saving time and resources.
However, AI-generated candidates still require laboratory validation because current models cannot reliably predict kinetics or developability. This has increased pressure on lab teams, who must test and characterize a growing volume of diverse, AI-generated compounds. Traditional screening workflows, designed for binary yes-or-no results, struggle to provide the detailed data needed for these complex candidates.
A major challenge is the quality and completeness of training data. Many AI models rely on public datasets that lack structure, labeling, and diversity, leading to a "data wall" where models reach similar conclusions with diminishing returns. Publication bias exacerbates this, as most public data focuses on positive results, while negative data—failed experiments and non-binding compounds—remains unpublished. Paul Belcher, director of protein research strategy at Cytiva, notes that this bias limits models' ability to learn from failures and make reliable predictions.
Data integrity is another concern. Fabrication of images, such as Western blots, has become easier with generative AI. Research by microbiologist Elisabeth Bik found that almost 4% of biomedical papers contained duplicated or manipulated images as of 2016. Belcher warns that manipulated data used to train AI could have disastrous consequences. Some vendors are developing tools to verify data authenticity, such as Cytiva's Image Integrity Checker, which uses secure hash algorithms to detect tampering. Publishing houses are showing interest in adopting such tools as standard practice.
Looking ahead, the industry must address these data challenges to fully realize AI's potential in drug discovery. This includes creating more comprehensive datasets that include negative results, improving data labeling and structure, and implementing verification tools to ensure data integrity. Without these steps, AI models risk perpetuating biases and producing unreliable predictions, undermining their promise to accelerate drug development.
Sources
- MIT Technology ReviewSecondary
Related
Chicxulub impact charbroiled dinosaurs with superheated dust cloud
Asteroid dust cloud roasted dinosaurs to death within hours
NASA reveals Earth's geoid varies by 191 meters in gravity model
All living things emit a faint glow, scientists explore medical uses
Amateur astronomer finds 390-million-year-old meteorite crater via Google Maps
Extreme solar storms may have no upper intensity limit, new research warns
Jodrell Bank faces funding crisis risking UK physics brain drain
Fossil footprints show 1.4-million-year-old human relative traveled in all-male groups
Trending now
- Cyera acquires Oasis Security for $1B in third deal this year
- NASA Swift rescue mission hits attitude control trouble
- American Airlines grounds all flights nationwide after IT outage
- SK Hynix Q2 profit surges 557% to record high
- 1,100 AI staffers urge US to pace tech growth
- ChatGPT hack overwhelms tech firm, emergency call held
- NASA’s Swift rescue satellite loses 2 of 3 reaction wheels
- Chicxulub impact charbroiled dinosaurs with superheated dust cloud
Comments
No comments yet. Be the first.