Analytical scientist reviewing chromatographic data on screen beside pharmaceutical samples in the Molkem Labs R&D laboratory

Direct answer

Generative artificial intelligence (AI) supports drug discovery by proposing molecular structures, protein designs and research hypotheses for scientific evaluation. Combined with predictive models and laboratory testing, it can guide candidate selection and optimisation. Its outputs require experimental confirmation and do not, by themselves, establish safety, efficacy or pharmaceutical development readiness.

Generative AI at a Glance

Molecular design

Generative models propose structures or sequences for a defined scientific objective.

Experimental evidence

Predictions help prioritise work; measured results establish whether a proposal merits further investigation.

Development readiness

Candidate selection connects to formulation, analytical, stability, safety and manufacturing considerations.

Scientific accountability

Data quality, model limitations and documented human review determine how outputs can be used.

The Role of Generative AI in Drug Discovery

The role of generative AI in drug discovery is to propose scientifically relevant candidates for investigation. It extends the range of structures or sequences that researchers can consider, while relying on predictive tools and experiments to evaluate their properties.

Generation and prediction serve different purposes. A generative model may propose a molecular structure. A predictive model may estimate that structure’s activity or solubility. Molecular docking assesses possible binding arrangements within a target structure. None of these outputs, considered alone, demonstrates that a compound will treat a disease.

For small molecules, generative methods can explore alternatives to a known chemical series or suggest structures under specified constraints. For proteins, they can propose backbones or sequences for a desired function. The appropriate representation and evaluation method depend on the research question.

Early work on continuous molecular representations demonstrated how a learned representation could support generation and property-guided exploration. It illustrates a design capability, rather than evidence of therapeutic benefit.

A useful distinction is between a proposed structure, a confirmed hit, a lead and a development candidate. A hit shows activity in a defined experimental system. A lead undergoes further optimisation and characterisation. Candidate nomination is a programme decision based on a broader evidence package; it is not a designation supplied by the model.

Applications Across the Discovery Workflow

Generative AI is most useful within an iterative scientific workflow. Researchers define the objective, generate proposals, assess practical constraints, test selected candidates and use the results to guide further work. The relevant output is a better-supported decision about the next experiment or candidate.

Target selection can draw on biological data, literature and predictive analyses. Generative methods may contribute hypotheses, but identifying a plausible target and validating its disease relevance remain distinct scientific tasks. Similarly, a predicted interaction should not be described as confirmed binding.

Swipe horizontally to view the full table.

Illustrative workflow: activities, evidence and decisions
ActivityAI contributionEvidence and decision
Define the research objectiveSupport exploration of biological hypotheses and design constraints.Review target rationale, available assays and relevant data before defining the generation task.
Generate molecular proposalsPropose structures, analogues or protein designs under selected constraints.Check representation validity, diversity and relevance before choosing candidates for further assessment.
Prioritise candidatesCombine generation with property predictions and other computational assessments.Examine uncertainty, synthesis feasibility and property trade-offs to select a manageable experimental set.
Synthesise and testInform design choices and, where suitable, support synthesis planning.Confirm identity and relevant biological properties with appropriate controls; decide which findings warrant follow-up.
Refine the designUse measured outcomes to guide subsequent proposals.Assess whether improvements are reproducible and whether new liabilities have emerged.
Evaluate candidate readinessContribute to the combined assessment of candidate options.Review pharmacology, safety, developability and supply considerations before advancing the programme.

This table is a practical synthesis, not a fixed regulatory sequence. Workstreams overlap, and an unexpected finding may require changes to an earlier design assumption. Research on reinforcement learning for molecular design demonstrates how generation can be guided by selected objectives; the objectives themselves still require scientific justification.

Generative Models Used in Drug Discovery

Model families differ in how they represent molecules and generate proposals. Selection should reflect the available data, the output required and the evaluation method. A model’s technical sophistication does not, by itself, establish the quality of the candidates it produces.

  • Variational autoencoders: learn a continuous representation from which molecular proposals can be reconstructed or sampled.
  • Autoregressive and transformer models: generate sequences progressively. Chemical and protein applications require domain-specific representations and assessment.
  • Diffusion models: learn to recover structured outputs from noise and can be adapted to molecular or protein design tasks.
  • Generative adversarial networks: train a generator alongside a discriminator that evaluates its outputs against training examples.

Reinforcement learning is often used to steer generation towards specified objectives; it should not be treated as interchangeable with a particular model architecture. Graph representations, sequence representations and three-dimensional structures also describe different aspects of a modelling system.

In RFdiffusion, researchers adapted a structure-prediction network for generative protein design. This is distinct from using a structure predictor only to estimate how an existing sequence may fold.

Published Evidence and Its Interpretation

Evidence for generative AI in discovery ranges from computational benchmarks to laboratory experiments and clinical studies. These evidence levels answer different questions. A successful design benchmark cannot establish patient benefit, and a clinical result for one molecule cannot establish the performance of all generative methods.

Protein design: RFdiffusion

The RFdiffusion study reported generative protein-design tasks and experimental characterisation of selected designs, including protein binders. Its significance lies in connecting computational proposals with measured molecular behaviour. The reported results apply to the designs and conditions evaluated; therapeutic suitability requires further investigation.

Clinical investigation: rentosertib

A randomised phase 2a study published in 2025 evaluated rentosertib, a generative AI-designed small molecule, in 71 adults with idiopathic pulmonary fibrosis over 12 weeks. The primary endpoint concerned treatment-emergent adverse events. Lung-function measurements were among the secondary endpoints.

The authors reported findings that supported further investigation. The study’s size and duration limit conclusions about long-term safety and efficacy. It provides an example of an AI-origin candidate undergoing clinical evaluation, rather than proof that AI removes development risk or establishes regulatory approval.

The study design also matters when interpreting productivity claims. This trial compared treatment groups; it did not compare AI-assisted discovery with a conventional discovery programme. It therefore cannot establish a general improvement in discovery cost, speed or clinical success rate.

Benefits and Practical Limitations

Generative methods can expand design options and help researchers consider several objectives together. Potential value includes exploring alternative structures, identifying property trade-offs and prioritising experiments. Benefits should be assessed against the programme’s actual results and resources, rather than the number of structures generated.

For example, improving predicted target activity may be unhelpful if the same changes reduce solubility or make synthesis impractical. Absorption, distribution, metabolism, excretion and toxicity (ADMET) estimates can support prioritisation, but their reliability depends on the endpoint, training data and candidate chemistry.

Data quality and applicability

Training data may contain inconsistent assay conditions, missing negative results or uneven representation of chemical series. A model can perform well on familiar examples and less reliably on new chemistry. Evaluation should therefore examine relevant independent data and the limits of the intended application.

Scoring functions and experimental relevance

A generator can improve its numerical score without improving the property that matters experimentally. Published work on the limits of molecular generation with reinforcement learning illustrates why optimisation objectives and evaluation design require scrutiny.

GuacaMol benchmarks provide structured ways to assess molecular-generation performance. Such benchmarks support comparison, but do not replace testing of synthesis feasibility, biological activity or candidate performance in the intended setting.

Synthesis, novelty and reproducibility

A proposed structure must be practically accessible before its properties can be tested. Synthetic-accessibility estimates do not establish a reproducible process, acceptable impurity profile or reliable supply. Structural novelty within a dataset also does not establish patentability or freedom to operate; those require separate assessment.

Traceability should connect the model version, input data, candidate identity, selection rationale and experimental results. Without that record, teams may struggle to reproduce a result or understand why a design decision changed.

These limitations explain why the role of generative AI in drug discovery should be assessed through experimentally supported progress. Faster computation is useful when it improves the quality or efficiency of the scientific decisions that follow.

Development Requirements After Candidate Selection

An AI-origin molecule must undergo evaluation of exposure, quality, stability and reproducible supply, as required for other development candidates. Formulation, analytical and process evidence must develop alongside the programme’s safety and clinical evidence.

Preformulation and formulation development

For a small-molecule candidate, early assessment may examine solubility, solid-state properties, degradation pathways and excipient compatibility. These findings inform formulation development and process choices. The appropriate work depends on the molecule, intended route, dose and stage of development.

The International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use (ICH) develops harmonised technical guidelines. Its Q8(R2) pharmaceutical-development guideline provides principles for understanding product and process performance. Application remains subject to regional implementation and product context.

Analytical methods and stability

Analytical development establishes how relevant quality attributes will be measured. Depending on the product, this may include identity, assay, related substances and dissolution. Methods and their performance evidence should be suitable for the decision they support.

ICH Q14 addresses science- and risk-based analytical procedure development. Stability studies separately examine how quality changes over time under relevant conditions. The ICH Q1A(R2) guideline describes the role of these data in supporting storage conditions and shelf-life or retest-period decisions.

Integrated CMC, clinical and manufacturing planning

Chemistry, manufacturing and controls (CMC) work connects material attributes, formulation, process, analytical methods and specifications. Nonclinical and clinical planning must account for the material being evaluated. A change in formulation or process may require additional assessment of its effect on quality or performance.

These activities may progress in parallel. Early analytical findings can change formulation choices; stability results can alter packaging requirements; scale-up can reveal process sensitivities. The broader pharmaceutical drug development process depends on maintaining continuity across these decisions.

Responsibilities should be explicit when work involves external partners. Technology transfer needs documented process and method knowledge. Manufacturing or clinical activities performed by qualified partners should be distinguished from work undertaken directly by the development organisation.

AI Evidence and Regulatory Expectations

Regulatory relevance depends on how an AI output is used. Exploring candidate ideas internally is different from relying on an AI-derived result to support a regulatory decision about safety, effectiveness or quality. The intended use determines the evidence and oversight that need consideration.

The United States Food and Drug Administration (FDA) has published draft guidance on AI supporting drug and biological-product regulatory decisions. It proposes assessing model credibility within a defined context of use through a risk-based framework. The cited document is draft guidance, not a final universal requirement for exploratory discovery models.

The joint FDA–European Medicines Agency (EMA) principles for good AI practice address human oversight, data governance, context of use, performance assessment and lifecycle management. These principles support responsible practice; they do not create a single international approval route.

For a specific programme, teams should document which outputs influence decisions, what limitations apply and who reviews them. Current authority expectations should be checked for the product, development stage and market. The role of regulatory affairs in drug development includes connecting those expectations with the planned evidence package.

Considerations for Programmes in India

For India-based pharmaceutical teams, AI-assisted discovery should connect to an evidence and development plan appropriate to the intended markets. Model origin does not establish that the candidate, its manufacturing process or its clinical programme meets applicable requirements.

The Central Drugs Standard Control Organization (CDSCO) and the applicable New Drugs and Clinical Trials Rules, 2019, and amendments are relevant to new-drug and clinical-trial planning in India. Required permissions and documentation depend on the product and proposed activity. Teams should verify the current pathway rather than infer an exemption or accelerated route from the use of AI.

For programmes involving multiple organisations, project agreements and quality arrangements should define data access, confidentiality, deliverables and review responsibilities. Candidate records should remain connected to source data and experimental reports as work moves into formulation, analytical and regulatory development.

Where other markets are intended, the India plan should be coordinated with the relevant authorities’ requirements. Scientific principles may be shared across jurisdictions, while procedures and submission expectations differ.

Assessment Priorities for Development Teams

A development assessment should separate what is proposed, what has been measured and what remains uncertain. The following priorities help organise that review without assuming that every programme requires the same tests or acceptance criteria.

  • Defined objective: the intended target, candidate properties and decision are clear.
  • Traceable data: sources, assay conditions and relevant limitations are documented.
  • Relevant evaluation: computational performance is distinguished from independent experimental results.
  • Practical feasibility: synthesis or production considerations receive scientific review.
  • Balanced properties: activity is assessed alongside selectivity, exposure and developability.
  • Evidence continuity: candidate identity, analytical results and development records remain connected.
  • Regulatory context: intended markets and uses of AI-derived evidence are identified.
  • Clear accountability: sponsor, development-team and partner responsibilities are defined.

Conclusion

The role of generative AI in drug discovery is to expand and guide the design options available to researchers. Its value depends on the quality of the scientific objective, the reliability of the data and the experimental evidence generated from its proposals.

Progress towards a medicine requires continuity between discovery decisions and formulation, analytical, safety, clinical and manufacturing evidence. Generative capability contributes to that process when its outputs remain interpretable, traceable and subject to appropriate scientific review.

Frequently Asked Questions

What is the role of generative AI in drug discovery?

Generative AI proposes molecular structures, protein designs or research hypotheses for scientific evaluation. It can help explore candidate options and guide optimisation when combined with predictive models and laboratory testing. Its output does not establish that a candidate is safe, effective or suitable for development.

How does generative AI differ from predictive AI?

Generative AI produces candidate structures or other new outputs. Predictive AI estimates properties of an input, such as activity or solubility. Discovery workflows often combine them: a generator proposes options, predictors help rank them, and experiments test the resulting hypotheses.

Can generative AI replace laboratory experiments?

No. Computational proposals need appropriate experimental assessment. Depending on the programme, this includes confirming identity, biological activity, selectivity and relevant properties. Clinical benefit and product quality require evidence beyond a molecular design or predicted score.

What are the benefits of using AI in drug discovery?

Potential benefits include exploring more candidate options, identifying property trade-offs earlier and prioritising experiments. Their value depends on data quality, model performance and experimental results. A faster design task does not establish a shorter overall development programme or a higher probability of approval.

What are the main limitations of generative AI in drug discovery?

Important limitations include incomplete training data, inaccurate predictions outside familiar chemistry, uncertain synthesis feasibility and optimisation against imperfect scoring functions. Scientific review, prospective testing and traceable records help determine whether a proposed candidate is useful for the intended programme.

Has a generative AI-designed molecule been studied in humans?

Yes. A published randomised phase 2a study evaluated rentosertib in adults with idiopathic pulmonary fibrosis. Its primary endpoint concerned safety, with lung-function measures among the secondary endpoints. This provides a clinical research example; it does not establish the general success of AI-designed medicines.

How should India-based teams assess an AI-origin candidate?

Assess the experimental evidence and intended development route, including quality, nonclinical and clinical requirements where applicable. Check current CDSCO requirements for the product and proposed activity. Document how AI outputs influenced decisions and keep responsibilities clear across the sponsor and its partners.

Where does Molkem Labs fit within this lifecycle?

Molkem Labs supports pharmaceutical development through its confirmed formulation, analytical, stability, regulatory and technology-transfer scope. Its role should be defined for the product and project. This article does not represent Molkem as a provider of drug discovery or generative AI-model development services.

About Molkem Labs

Molkem Labs is an integrated R&D and pharmaceutical development platform designed to take products from concept to market through a seamless combination of formulation development, analytical development, regulatory expertise and advanced technology.

Spread across a 45,000 sq. ft. facility with an integrated 100 MT warehouse, the centre is equipped with advanced capabilities along with dedicated facilities for microbiology. Supported by decentralized HVAC systems, classified clean rooms, cGMP-compliant infrastructure, 21 CFR & EQFAR-compliant analytical laboratories, and NABL-accredited capabilities, Molkem Labs is engineered to handle complex products including hygroscopic, thermolabile and deliquescent molecules, while enabling development of patented molecules under POC as per QbD principles.

From early-stage development and analytical characterization to scale-up, technology transfer and regulatory support, Molkem Labs offers an integrated pathway to accelerate pharmaceutical innovation and bring quality products from development to market.

Explore Pharmaceutical Development Support

Molkem Labs supports formulation, analytical, stability, regulatory and technology-transfer activities within a defined development programme. Enquiries are assessed against the product requirements and confirmed project scope.

RELATED POSTS

Expert Tips, Latest News, and Innovations