Insilico Medicine (HKEX: 3696) has launched Drug Discovery and Development Benchmark as a Service (DDD BaaS), a commercial platform that enables AI developers and pharmaceutical companies to evaluate foundation models against drug discovery and development tasks using standardized pharmaceutical benchmarks.
The service comprises two evaluation suites. Drug Discovery Foundations includes more than 300 benchmark tasks spanning disease biology, molecular property prediction, retrosynthesis, structure-based drug design, and clinical development, while Drug Candidate Essentials assesses a model's ability to progress an end-to-end discovery program from hit identification through preclinical candidate (PCC) nomination. According to Insilico, both suites use decontaminated, out-of-distribution datasets and reference baselines derived from the company's own validated drug discovery programs.
Models accessible through a standard chat-completions API can be submitted for evaluation. Participants receive a verified scorecard benchmarked against expert reference outputs, with the option of appearing on a public leaderboard. The platform is available immediately through dddbench.insilico.com and builds on the company's Pharma.AI platform and MMAI Gym scientific AI training environment.
The launch addresses a growing challenge in AI-driven drug discovery: publicly available benchmarks can be vulnerable to training-data contamination, allowing models to achieve artificially high scores by memorizing benchmark content rather than demonstrating real-world scientific capability. Insilico says its framework is designed around proprietary programs that have generated more than 30 PCC nominations, more than 10 investigational new drug (IND) clearances, and Rentosertib (ISM001-055), which is currently in Phase III development for idiopathic pulmonary fibrosis.
The service expands Pharma.AI beyond Insilico's internal research and platform licensing business, creating a new commercial offering aimed at independent model evaluation. Potential users include pharmaceutical companies, AI drug discovery firms, and developers of large language and molecular foundation models seeking standardized assessment of model performance on pharmaceutical research tasks.