AI in Life Sciences Glossary: 86 Artificial Intelligence Terms Explained



Key Takeaways
  • AI vocabulary in life sciences now spans four distinct worlds: computation, drug discovery, regulation, and data infrastructure
  • Regulatory terminology is the fastest-moving category, with FDA guidance on predetermined change control plans finalized in 2025 and EU AI Act obligations phasing in
  • The first fully AI-designed drugs entered Phase III trials in 2026, but none has completed one or reached approval, so treat AI outputs as accelerated hypotheses rather than validated answers
  • Several core terms, including TechBio and AI-ready data, have genuinely competing definitions and are still being negotiated across the industry
  • This glossary is a companion to the California Life Sciences Glossary and cross-links to it throughout

Artificial intelligence entered life sciences through the side door. It arrived first as a computational chemistry tool, then as a clinical trial operations tool, then as a regulated medical device, and now as a line item in nearly every investor conversation in California. Each of those arrivals brought its own vocabulary, and none of them agreed to standardize.

The result is a communication problem that founders, operators, and investors feel daily. A regulatory lead and a computational chemist can use the word “validation” in the same meeting and mean two completely different things. This glossary covers 86 terms across AI foundations, drug discovery and molecular design, regulatory and governance, data infrastructure, clinical development, omics, diagnostics, and manufacturing — defined for the business of life sciences rather than for machine learning engineers.

See California Life Sciences resources


How to Use This Glossary

This reference is organized alphabetically, not by category. Use the jump navigation below to reach any letter instantly. See also links connect related terms and help build context around unfamiliar concepts.

Foundational terms — the ones worth knowing before the rest make sense — are marked with a .

This glossary is a companion to the California Life Sciences Glossary, which covers the core vocabulary of drug development, regulatory pathways, and life sciences financing. The two are designed to be read together. AI terminology sits on top of that existing vocabulary rather than replacing it, so where an entry here depends on a core concept such as Phase III, orphan drug status, or bioinformatics, the See also line points to the main glossary. Definitions there are not repeated here.


# · A · B · C · D · E · F · G · H · I · L · M · N · O · P · Q · R · S · T · V


Nearly half of these terms sit in drug discovery and regulatory governance — the two areas where AI vocabulary carries the most practical consequence.


Why AI Vocabulary Matters Right Now

Terminology confusion is not a cosmetic problem in this field. It has direct consequences for diligence, regulatory submissions, and partnership negotiations.

Consider the word “validation.” To a computational scientist, validating a model means holding back a slice of the training data and checking predictive accuracy against it. To a regulatory affairs professional, validation means demonstrating that a system consistently performs as intended for its specific stated use in the real clinical environment. Both are correct within their discipline. A company that submits the first when a regulator expects the second has a problem that no amount of good science can fix.

The same gap opens around “AI-ready,” “explainable,” “real-world evidence,” and “drift.” These are not abstractions. They are the words in the diligence questionnaire.


Our observation: In California Life Sciences’ FAST California program, the early-stage companies that stumble on AI questions rarely stumble on the science. They stumble on translation — describing a model’s performance in terms an investor’s technical advisor finds credible, or explaining which parts of a workflow a human still reviews. The companies that answer those questions crisply move faster, and it has more to do with vocabulary than with the underlying technology.


No-Cost Startup Mentorship

FAST California is a 12-week, no-cost, equity-free mentorship program for early-stage life sciences founders. Applications are accepted on a rolling basis.

Learn More and Apply →


Where AI Actually Sits in the Drug Development Pipeline

One reason the vocabulary feels chaotic is that AI is not one thing applied at one point. It appears at nearly every stage of the pipeline, doing genuinely different jobs, with genuinely different levels of maturity and regulatory scrutiny at each.

In discovery, AI proposes: it predicts protein structures, ranks compounds, and suggests targets. In clinical development, AI selects and monitors: it stratifies patients, models enrollment, and reads trial data. In regulated products such as diagnostic software, AI decides or assists in deciding, which is where the FDA’s device framework applies. In manufacturing, AI watches: it detects anomalies and predicts equipment failure. Those are four different risk profiles, and conflating them is the single most common source of confusion in this vocabulary.



The same three letters describe seven different jobs. Knowing which one a vendor or a paper means is most of the work.


The maturity gradient matters as much as the map. Protein structure prediction is a solved-enough problem that it has changed daily practice in structural biology. AI-designed therapeutics are in human trials but have not completed the full journey. AI-enabled diagnostic software has a well-established regulatory pathway. AI-generated regulatory documentation is still finding its footing.


Our observation: Across the California Life Sciences membership, the AI deployments that have produced measurable results in the past two years are mostly unglamorous — document processing, trial site selection, assay image analysis, and quality control. The molecular design work draws the headlines and the capital. The operational work is where most companies are actually seeing returns today.



The compression is real and it is confined to discovery. Attrition in Phase I and II has not moved, which is why an AI-designed molecule still faces the same odds once it reaches patients. Source: Frontiers in Pharmacology (2026).


Explore CLS Membership

California Life Sciences membership connects you to California’s leading life sciences network — from advocacy and cost savings to exclusive events and programming.

See Benefits →


Several of These Terms Are Still Being Argued Over

Most glossaries present definitions as settled. In AI for life sciences, several of the most commonly used terms are not settled at all, and pretending otherwise makes a reference less useful rather than more.

“TechBio” is the clearest example. Depending on who is speaking, it describes a company built on a computational platform rather than a single asset, a company where engineers outnumber bench scientists, or simply a biotech company that wants a technology valuation multiple. “AI-ready data” is contested in a similar way, split between a technical reading focused on machine readability and a governance reading focused on provenance and consent. “Explainability” has at least three working definitions across computer science, clinical practice, and regulation.

Where a term is contested, the entries below say so and describe the competing readings. The practical advice is simple: in any contract, diligence questionnaire, or regulatory submission, define the term inside the document rather than assuming shared meaning.


Our observation: The most expensive vocabulary mistakes we see are not misunderstandings between a company and a regulator. They are misunderstandings between a company and a vendor, discovered after signature, over what “validated,” “compliant,” or “your data” was supposed to mean.


Life Sciences Savings

CLS members save an average of $20,783 per year through the California Life Sciences Savings Program, with larger organizations saving as much as $587,566 per year.

Save With CLS →


The Full Glossary


#

21 CFR Part 11
The section of US federal regulation that sets FDA requirements for electronic records and electronic signatures in regulated activities. It governs audit trails, system access controls, and record integrity. Part 11 applies to AI systems that generate or hold records supporting a regulated submission, which is why model outputs used in filings are subject to the same recordkeeping expectations as any other regulated record. See also: Data Provenance, Model Card.


A

Adaptive Trial Design
A clinical trial design that allows prespecified modifications — such as dropping an arm, adjusting sample size, or reallocating patients — based on interim data, without undermining statistical validity. AI is used to simulate candidate designs and identify which adaptations preserve power. All adaptation rules must be defined before the trial begins. See also: Predictive Enrollment, Protocol Optimization, Clinical Trial (Life Sciences Glossary).

ADMET Prediction
The computational prediction of a compound’s absorption, distribution, metabolism, excretion, and toxicity properties before it is synthesized. Poor ADMET properties are a leading cause of late-stage attrition, so predicting them early is one of the highest-value applications of machine learning in discovery. Predictions narrow the candidate list; they do not replace toxicology studies. See also: QSAR, Lead Optimization, Toxicology (Life Sciences Glossary).

Agentic AI
An AI system that plans and executes multi-step tasks with limited human direction, calling tools, querying databases, and acting on intermediate results. In life sciences, agentic systems are being tested for literature review, regulatory document assembly, and laboratory workflow orchestration. Because the system acts rather than just answers, the review and audit requirements are correspondingly higher. See also: Human-in-the-Loop, Self-Driving Lab.

AI-Designed Antibody
A therapeutic antibody whose sequence or binding region was generated or substantially optimized by a computational model rather than discovered through animal immunization or library screening. Several AI-designed antibodies have entered human trials. None has completed Phase III, so the category remains a demonstrated acceleration of the design stage rather than a proven end-to-end substitute. See also: De Novo Protein Design, Binding Affinity Prediction, Antibody (Life Sciences Glossary).

★ AI-Ready Data
Data that is complete, consistently structured, well-documented, and legally usable for model training or inference. The term is contested: a technical reading emphasizes machine readability and standardized formats, while a governance reading emphasizes provenance, consent, and permitted use. Both readings matter in practice, and vendors frequently use the term to mean only the first. See also: FAIR Data, Data Provenance, Structured vs. Unstructured Data.

AI Vendor Diligence
The evaluation process a life sciences organization applies before adopting a third-party AI tool, covering model performance claims, training data sources, validation evidence, data ownership, security posture, and regulatory support. Diligence here differs from ordinary software procurement because performance can degrade after deployment and because the buyer usually carries the regulatory burden. See also: Build vs. Buy, Model Drift, Model Card, Due Diligence (Life Sciences Glossary).

★ Algorithmic Bias
Systematic error in a model’s outputs that produces different performance across populations, typically because the training data underrepresented some groups. In diagnostics, this can mean lower sensitivity for patients whose demographic or clinical profile was scarce in training. Detecting it requires performance to be reported by subgroup rather than in aggregate. See also: Training Data, Sensitivity and Specificity, Model Validation vs. Verification.

★ AlphaFold
An artificial intelligence system from Google DeepMind that predicts a protein’s three-dimensional structure from its amino acid sequence. It has predicted more than 200 million structures, released freely, and the underlying work shared the 2024 Nobel Prize in Chemistry. Predictions are strong hypotheses rather than measurements. See our full explainer: What Is AlphaFold? See also: Protein Structure Prediction, In Silico.

Ambient Clinical Documentation
Software that listens to a clinical encounter and drafts the visit note automatically, using speech recognition and language models. It is among the fastest-adopted AI applications in healthcare because the workflow benefit is immediate and the clinician reviews and signs every note. The clinician remains responsible for accuracy. See also: Natural Language Processing (NLP) for EHR, Human-in-the-Loop, Hallucination.

Anomaly Detection
A machine learning approach that flags data points deviating from an established normal pattern, without needing labeled examples of every possible failure. In biomanufacturing it identifies batch deviations and sensor faults; in pharmacovigilance it surfaces unexpected safety signals. Its main practical limitation is false positives, which erode trust when alert volume is high. See also: Process Analytical Technology (PAT), Predictive Maintenance, Pharmacovigilance and Signal Detection.

★ Artificial Intelligence (AI)
Computer systems that perform tasks normally requiring human intelligence, such as recognizing patterns, making predictions, interpreting language, or generating new content. In life sciences the term covers everything from a decades-old statistical classifier to a modern generative model, which is precisely why it carries so little information on its own. Always ask which method is meant. See also: Machine Learning, Deep Learning, Generative AI.


B

Binding Affinity Prediction
The computational estimation of how tightly a candidate molecule attaches to a target protein, usually expressed as a dissociation constant. Stronger predicted affinity guides which compounds are synthesized and tested first. Predicted affinities are useful for ranking candidates against one another and are considerably less reliable as absolute values. See also: Molecular Docking, Virtual Screening, Hit-to-Lead.

Biological Foundation Model
A large model trained on broad biological data — genomic sequences, single-cell profiles, or protein sequences — that can be adapted to many downstream tasks rather than built for one. Examples include sequence models such as Evo. The category is early, and independent benchmarking of biological claims is still thinner than in language or vision. See also: Foundation Model, Protein Language Model, Genomics (Life Sciences Glossary).

Build vs. Buy
The decision of whether to develop an AI capability internally or license it from a vendor. Building offers control over data and model behavior but requires machine learning talent, infrastructure, and ongoing maintenance. Buying moves faster but concentrates risk in the vendor’s data practices and validation evidence, which the adopting company usually still has to defend. See also: AI Vendor Diligence, Model Card.


C

Clinical Decision Support (CDS)
Software that provides clinicians with patient-specific recommendations, alerts, or risk scores to inform care decisions. Whether a given CDS tool is regulated as a medical device depends on factors including how much the clinician can independently review the basis for the recommendation. Transparency of reasoning is therefore a regulatory question, not only a design preference. See also: Software as a Medical Device (SaMD), Explainability, Human-in-the-Loop.

Computer-Aided Detection (CADe/CADx)
Software that marks suspicious findings in medical images (CADe) or characterizes them as likely benign or malignant (CADx). It is the oldest regulated category of clinical AI, with cleared products in mammography and radiology dating back decades. Performance is reported as sensitivity and specificity against a reference standard. See also: Computer Vision, Sensitivity and Specificity, Diagnostics (Life Sciences Glossary).

Computer Vision
The field of AI concerned with extracting information from images and video. In life sciences it underpins digital pathology, radiology software, high-content screening, and automated colony counting in the lab. It is the most operationally mature branch of AI in the sector, with a well-established regulatory track record in imaging. See also: Digital Pathology, High-Content Screening, Radiomics.

Context of Use
A precise statement of the specific question a model is intended to answer, for which population, and within which decision. It is the anchor of AI credibility assessment: the same model can be well supported for one context of use and unsupported for another. Widening the context of use after validation restarts the evidence burden. See also: Credibility Assessment, Model Validation vs. Verification.

Credibility Assessment
A structured evaluation of whether the evidence supporting a computational model is sufficient for its stated context of use and the risk of the decision it informs. The framework scales evidence to consequence: a model influencing an exploratory analysis requires less support than one influencing a dosing decision. See also: Context of Use, Model Validation vs. Verification, In Silico.


D

Data Provenance
The documented record of where a dataset came from, how it was collected, what consent or license governs it, and every transformation applied since. Provenance determines whether data can lawfully be used to train a model and whether resulting outputs can be defended in a submission. Gaps in provenance are frequently discovered during diligence, not before. See also: AI-Ready Data, FAIR Data, 21 CFR Part 11.

De-identification
The removal or alteration of information that could identify an individual in a health dataset, so the data can be used for research or model training. De-identification reduces re-identification risk but does not eliminate it, particularly for rich longitudinal or genomic data where combinations of attributes can be distinctive. See also: HIPAA Safe Harbor, Differential Privacy, Synthetic Data.

De Novo Protein Design
The computational design of proteins that do not exist in nature, built to a specified structure or function rather than derived from a known sequence. The field’s foundational work shared the 2024 Nobel Prize in Chemistry. Designed proteins are entering therapeutic and industrial development, and each still requires the full experimental and clinical validation path. See also: Protein Structure Prediction, AI-Designed Antibody, Generative AI.

★ Deep Learning
A subset of machine learning that uses neural networks with many layers to learn patterns directly from raw data, without hand-engineered features. Deep learning is what made protein structure prediction, medical image analysis, and language models work at their current level. It typically requires substantially more data than classical statistical methods. See also: Machine Learning, Neural Network, Training Data.

Differential Privacy
A mathematical guarantee that the presence or absence of any single individual’s record cannot be detected from a model’s outputs, achieved by adding calibrated noise. It provides a formal, quantifiable privacy assurance rather than a procedural one. The trade-off is direct: stronger privacy guarantees reduce the precision of results. See also: De-identification, Federated Learning, Synthetic Data.

★ Digital Biomarker
A physiological or behavioral measure collected by a sensor or digital device — such as gait, sleep pattern, or heart rate variability — used as an indicator of disease state or treatment response. Digital biomarkers can be captured continuously outside the clinic. Each requires its own validation before it can support a regulatory endpoint. See also: Real-World Evidence (RWE), Digital Health (Life Sciences Glossary).

★ Digital Pathology
The digitization of tissue slides into high-resolution images that can be reviewed remotely and analyzed by software. It is a prerequisite for applying computer vision to pathology, enabling automated tumor grading, biomarker quantification, and consistency checks across readers. Slide scanning and storage costs are the usual barrier to adoption rather than the algorithms. See also: Computer Vision, High-Content Screening, Diagnostics (Life Sciences Glossary).

★ Digital Twin
A computational model of a specific biological system, patient, process, or facility that is updated with real data and used to simulate outcomes before acting in the physical world. In clinical development, patient-level digital twins are being explored as a way to model an individual’s likely control-arm trajectory. See also: Synthetic Control Arm, In Silico, Process Analytical Technology (PAT).


E

★ EU AI Act
European Union legislation establishing a risk-tiered framework for AI systems placed on the EU market, adopted in 2024 with obligations phasing in over several years. AI that forms part of a regulated medical device generally falls in the high-risk tier, carrying requirements for risk management, data governance, documentation, human oversight, and post-market monitoring. It applies alongside existing device rules. See also: Software as a Medical Device (SaMD), Human-in-the-Loop, Model Card.

★ Explainability (Black Box Problem)
The degree to which a model’s reasoning can be understood and inspected by a human. Deep learning models are often called black boxes because their internal logic is not directly readable. The term is contested — computer scientists, clinicians, and regulators each mean something different — so specify whether you mean feature attribution, clinical rationale, or development transparency. See also: Model Card, Clinical Decision Support (CDS), Neural Network.

External Control Arm
A comparator group drawn from patients outside the trial, whether from historical trials, registries, or routine care records. It is the broader category that includes synthetic control arms. External controls are most defensible where the disease course is well characterized and the outcome is objective. See also: Synthetic Control Arm, Real-World Evidence (RWE), Phase III Clinical Trial (Life Sciences Glossary).


F

★ FAIR Data
A set of principles stating that data should be Findable, Accessible, Interoperable, and Reusable — by machines as well as people. FAIR predates the current AI wave and is now frequently cited as the practical foundation for AI readiness. Note that FAIR concerns how data is described and structured, not whether it is open. See also: AI-Ready Data, Interoperability and FHIR, Data Provenance.

Federated Learning
A training approach in which a model is sent to each data holder, trained locally, and only the resulting model updates are shared centrally — the underlying records never leave the institution. This makes multi-site collaboration possible where patient data cannot be pooled. It adds substantial coordination and infrastructure complexity. See also: Differential Privacy, De-identification, Training Data.

★ Fine-Tuning
Adapting an already-trained general model to a specific task or domain by continuing training on a smaller, targeted dataset. Fine-tuning a general language model on regulatory documents, for example, improves performance without the cost of training from scratch. It is far cheaper than pretraining and still requires quality-controlled, well-governed data. See also: Foundation Model, Training Data, Retrieval-Augmented Generation (RAG).

★ Foundation Model
A large model trained on broad data at scale that serves as a general-purpose starting point for many downstream applications, rather than being built for a single task. Language models are the familiar case; protein and genomic foundation models apply the same idea to biological sequence. Adaptation to a specific use is a separate step. See also: Fine-Tuning, Biological Foundation Model, Large Language Model (LLM).


G

★ Generative AI
AI systems that produce new content — text, images, molecular structures, or protein sequences — rather than classifying or scoring existing inputs. In life sciences the same underlying technique writes a draft protocol summary and proposes a novel compound, which is why the term spans wildly different risk profiles. Always ask what is being generated and who reviews it. See also: Generative Chemistry, Large Language Model (LLM), Hallucination.

Generative Chemistry
The use of generative models to propose novel molecular structures with desired properties, rather than screening existing compound libraries. It expands the searchable chemical space enormously and produces candidates that must still be synthesized, assayed, and taken through the full development path. Generated molecules are proposals, and synthesizability remains a real constraint. See also: Virtual Screening, Retrosynthesis Prediction, ADMET Prediction.

★ Good Machine Learning Practice (GMLP)
A set of guiding principles for developing machine learning-enabled medical devices, published jointly by the FDA, Health Canada, and the UK’s MHRA. The principles address multidisciplinary expertise, data quality and representativeness, independence of training and test datasets, human factors, and monitoring after deployment. GMLP is a framework of principles rather than a binding checklist. See also: Software as a Medical Device (SaMD), Model Validation vs. Verification, Algorithmic Bias.


H

★ Hallucination
Output from a generative model that is fluent, confident, and factually wrong — a fabricated citation, a misstated dosage, an invented study result. Hallucination is a property of how these models generate text, not a bug that has been fixed. It is the single strongest argument for human review of any generative output that enters a regulated or clinical workflow. See also: Retrieval-Augmented Generation (RAG), Human-in-the-Loop, Large Language Model (LLM).

High-Content Screening
Automated microscopy combined with image analysis to extract many measurements per cell across large numbers of experimental conditions. Machine learning is used to classify phenotypes that are difficult to define by explicit rules. It generates very large image datasets, which makes storage and provenance practical concerns rather than afterthoughts. See also: Computer Vision, Single-Cell Analysis, Assay (Life Sciences Glossary).

HIPAA Safe Harbor
A method under US health privacy rules for treating health data as de-identified by removing 18 specified categories of identifiers, including names, dates more precise than year, and geographic detail below state level. It is a prescriptive checklist rather than a risk assessment. The alternative route relies on a qualified expert determination of re-identification risk. See also: De-identification, Data Provenance, Differential Privacy.

Hit-to-Lead
The discovery stage where confirmed active compounds, or hits, are evaluated and refined into a smaller set of lead series worth serious optimization. AI contributes by ranking hits on predicted potency, selectivity, and developability at once. It remains a stage defined by laboratory confirmation. See also: Lead Optimization, Virtual Screening, Lead Compound (Life Sciences Glossary).

★ Human-in-the-Loop
A system design in which a person reviews, approves, or can override the AI’s output before it takes effect. The term appears throughout regulatory frameworks as a control that reduces the risk of an AI-enabled system. Its effectiveness depends on whether the human genuinely has the information and time to disagree, rather than approving by default. See also: Clinical Decision Support (CDS), Agentic AI, Explainability.


I

★ In Silico
Performed by computer simulation, as distinct from in vitro (in glassware) and in vivo (in a living organism). In silico methods span molecular docking, ADMET prediction, trial simulation, and physiology modeling. Regulators evaluate in silico evidence through a credibility framework tied to the specific question being answered. See also: Credibility Assessment, Virtual Screening, Digital Twin.

Inference
The stage at which a trained model is actually used to produce an output on new data, as opposed to training, when it is learning. The distinction matters commercially and operationally: training is a large one-time cost, while inference is a recurring per-use cost that scales with adoption. See also: Training Data, Fine-Tuning, Model Drift.

Interoperability and FHIR
The ability of different health systems to exchange and use data, with FHIR (Fast Healthcare Interoperability Resources) as the dominant modern standard for structuring that exchange. Interoperability is the precondition for assembling training data across institutions. Adopting the standard does not by itself make records consistent, since local implementation practices vary. See also: FAIR Data, Structured vs. Unstructured Data, AI-Ready Data.


L

★ Large Language Model (LLM)
A model trained on very large volumes of text that generates and interprets language by predicting likely continuations. In life sciences, LLMs are applied to literature review, regulatory drafting, protocol summarization, and extracting structured facts from clinical notes. They generate plausible language rather than verified fact, which sets the review requirement. See also: Foundation Model, Hallucination, Retrieval-Augmented Generation (RAG).

Lead Optimization
The iterative refinement of a lead compound to improve potency, selectivity, safety, and pharmacokinetic properties before it becomes a development candidate. Machine learning shortens each design-make-test cycle by predicting which structural changes are most likely to improve the profile. Every cycle still ends in the laboratory. See also: ADMET Prediction, Hit-to-Lead, Lead Compound (Life Sciences Glossary).

Locked vs. Adaptive Algorithm
A locked algorithm produces the same output for the same input every time, changing only through a deliberate, documented update. An adaptive algorithm continues to learn after deployment and can change behavior in the field. The distinction is central to how AI-enabled devices are reviewed, because the second requires a plan describing permitted changes. See also: Predetermined Change Control Plan (PCCP), Model Drift, Software as a Medical Device (SaMD).


M

★ Machine Learning
A branch of AI in which systems learn patterns from data rather than following explicitly programmed rules. It ranges from classical methods such as random forests and regression to modern deep learning. Much of what is marketed as AI in life sciences is machine learning that predates the current wave, which is neither a criticism nor a selling point on its own. See also: Deep Learning, Training Data, Overfitting.

Model Card
A standardized summary document describing a model’s intended use, training data, performance across subgroups, known limitations, and evaluation methods. Model cards are becoming a default expectation in vendor diligence and in regulatory documentation because they make comparison possible. The absence of one is itself informative. See also: AI Vendor Diligence, Explainability, Algorithmic Bias.

★ Model Drift
The degradation of a model’s performance over time as real-world conditions diverge from the conditions represented in its training data. Drift can come from changed patient populations, new instruments, updated coding practices, or shifts in clinical workflow. Detecting it requires ongoing performance monitoring after deployment, not just validation at launch. See also: Locked vs. Adaptive Algorithm, Predetermined Change Control Plan (PCCP), Inference.

★ Model Validation vs. Verification
Verification asks whether a model was built correctly to its specification; validation asks whether it performs its intended job in the real setting of use. The two words are used interchangeably in casual conversation and are distinct in regulated contexts. Confusing them is a frequent source of misaligned expectations between technical and regulatory teams. See also: Context of Use, Credibility Assessment, Good Machine Learning Practice (GMLP).

Molecular Docking
A computational method that predicts how a small molecule fits into a target protein’s binding site and estimates the resulting pose and score. Docking has been used for decades; machine learning has improved both its speed and its scoring accuracy. Docking scores rank candidates and are not reliable predictions of experimental affinity. See also: Binding Affinity Prediction, Virtual Screening, In Silico.

Multi-omics Integration
The computational combination of genomic, transcriptomic, proteomic, metabolomic, and other layers of biological data into a single analytical view. Machine learning is used because the layers differ in scale, noise, and structure. Integration supports target discovery and patient stratification. See also: Single-Cell Analysis, Patient Stratification, Multi-omics (Life Sciences Glossary).


N

Natural Language Processing (NLP) for EHR
The application of language models to the unstructured narrative content of electronic health records — progress notes, discharge summaries, pathology reports — to extract structured, analyzable facts. A large share of clinically meaningful detail exists only in this free text. Extraction accuracy varies by document type and institution. See also: Structured vs. Unstructured Data, Real-World Evidence (RWE), Large Language Model (LLM).

Neural Network
A model structure loosely inspired by the brain, in which layers of simple connected units transform input into output, with the connection strengths learned from data. It is the architecture underlying deep learning. The biological analogy is a naming convention and should not be read as a claim about how the system reasons. See also: Deep Learning, Machine Learning, Explainability.


O

Overfitting
When a model learns the noise and idiosyncrasies of its training data rather than the underlying signal, producing excellent results on familiar data and poor results on new data. It is the most common reason an impressive internal benchmark fails to reproduce at a partner site. Independent test data is the standard defense. See also: Training Data, Model Validation vs. Verification, Model Drift.


P

Patient Stratification
Dividing a patient population into subgroups likely to respond differently to a treatment, using molecular, clinical, imaging, or real-world data. Better stratification can rescue a therapy that failed in an unselected population and is central to precision medicine. It also narrows the eligible population, which has commercial consequences. See also: Multi-omics Integration, Precision Medicine (Life Sciences Glossary).

Pharmacovigilance and Signal Detection
The monitoring of safety data after a product reaches the market to identify previously unrecognized adverse effects. Machine learning is applied to case volume, literature, and claims data to surface potential signals earlier. Every flagged signal requires human clinical assessment before any action. See also: Anomaly Detection, Real-World Evidence (RWE), Adverse Event (Life Sciences Glossary).

★ Predetermined Change Control Plan (PCCP)
A plan submitted to the FDA specifying in advance what modifications a manufacturer may make to an AI-enabled device software function after authorization, how they will be implemented, and how their impact will be assessed. The FDA issued final guidance on PCCPs in 2025. Changes covered by an authorized PCCP need no new marketing submission. See also: Locked vs. Adaptive Algorithm, Software as a Medical Device (SaMD), Model Drift.

Predictive Enrollment
The use of models to forecast how quickly a trial will recruit patients at a given set of sites, and to identify where recruitment is likely to stall. Enrollment shortfalls are among the most common causes of trial delay and cost overrun. Forecasts depend heavily on the quality of historical site data. See also: Site Feasibility, Protocol Optimization, Adaptive Trial Design.

Predictive Maintenance
The use of sensor data and machine learning to anticipate equipment failure before it occurs and schedule intervention accordingly. In biomanufacturing, unplanned downtime can jeopardize an entire batch, so early warning has direct financial value. It is one of the lower-risk, better-established AI applications in the sector. See also: Anomaly Detection, Process Analytical Technology (PAT).

Process Analytical Technology (PAT)
A framework for designing and controlling manufacturing by measuring critical quality attributes in real time during production rather than testing the finished batch. Machine learning models interpret the resulting sensor streams and support in-process adjustment. PAT is an established regulatory concept that predates current AI methods. See also: Predictive Maintenance, Anomaly Detection, Biomanufacturing (Life Sciences Glossary).

Protein Language Model
A model trained on large collections of protein sequences using techniques developed for natural language, learning statistical patterns that correspond to structure and function. Protein language models support structure prediction, variant interpretation, and design. Sequence is treated as a language here by analogy, and the model has no explicit representation of physics. See also: Biological Foundation Model, Protein Structure Prediction, Variant Effect Prediction.

★ Protein Structure Prediction
The computational determination of a protein’s three-dimensional shape from its amino acid sequence. Because shape determines function, this was one of biology’s longest-standing challenges before deep learning methods largely resolved it for well-behaved proteins. Predictions remain hypotheses: proteins flex, disordered regions stay difficult, and confidence scores must be read. See also: AlphaFold, De Novo Protein Design, In Silico.

Protocol Optimization
The use of data and modeling to test a draft trial protocol against historical trial performance before it is finalized, identifying eligibility criteria or assessment schedules likely to slow recruitment or burden sites. Small protocol changes have outsized effects on enrollment. Amendments after a trial starts are expensive. See also: Predictive Enrollment, Site Feasibility, Adaptive Trial Design.


Q

QSAR (Quantitative Structure-Activity Relationship)
A modeling approach that relates a molecule’s structural and physicochemical descriptors to its biological activity, allowing activity to be predicted for untested compounds. QSAR dates to the 1960s and is among the oldest machine learning applications in the sector. Its predictions hold only within the chemical space of the training set. See also: ADMET Prediction, Virtual Screening, Overfitting.


R

Radiomics
The extraction of large numbers of quantitative features from medical images — texture, shape, intensity distribution — that are not apparent to the human eye, then correlating them with clinical outcomes. It aims to turn a routine scan into a data-rich source. Feature values are sensitive to scanner and acquisition settings, which complicates generalization across sites. See also: Computer Vision, Digital Biomarker, Computer-Aided Detection (CADe/CADx).

★ Real-World Evidence (RWE)
Clinical evidence derived from real-world data — electronic health records, claims, registries, and device data — generated outside a controlled trial. RWE supports regulatory decisions in defined circumstances, informs payer negotiations, and underpins post-market safety work. Its central challenge is confounding, because patients were not randomized. See also: External Control Arm, Natural Language Processing (NLP) for EHR, Outcomes (Life Sciences Glossary).

Retrieval-Augmented Generation (RAG)
An architecture in which a language model retrieves relevant documents from a defined source before generating an answer, and grounds its response in what it retrieved. RAG reduces hallucination and allows answers to cite a traceable source. It does not eliminate error, since the model can still misread or misattribute retrieved material. See also: Hallucination, Large Language Model (LLM), Vector Database.

Retrosynthesis Prediction
The computational prediction of a viable synthetic route to a target molecule, working backward from the desired product to available starting materials. It addresses a practical bottleneck in generative chemistry, where models can propose molecules faster than chemists can determine how to make them. Predicted routes still require bench validation. See also: Generative Chemistry, Lead Optimization.


S

★ Self-Driving Lab
A laboratory in which experimental design, robotic execution, and analysis run in a closed loop, with an AI system choosing the next experiment based on the last result. The appeal is a design-make-test cycle that runs continuously. Adoption is limited by capital cost and by how many assays can be genuinely automated. See also: Agentic AI, High-Content Screening, Digital Twin.

Sensitivity and Specificity
Sensitivity is the proportion of true positives a test correctly identifies; specificity is the proportion of true negatives. Together they describe diagnostic performance far more usefully than a single accuracy figure, especially for rare conditions where a test can be highly accurate and nearly useless. Both should be reported by subgroup. See also: Algorithmic Bias, Computer-Aided Detection (CADe/CADx), Diagnostics (Life Sciences Glossary).

Single-Cell Analysis
Measurement of gene expression or other molecular features in individual cells rather than in bulk tissue, revealing population structure that averaging conceals. The datasets are large and sparse, which is why machine learning is standard for clustering, cell-type assignment, and trajectory inference. See also: Multi-omics Integration, Biological Foundation Model, Genomics (Life Sciences Glossary).

Site Feasibility
The assessment of whether a clinical trial site can realistically enroll and retain the required patients, given its patient population, staffing, and competing trials. Models trained on historical site performance increasingly supplement questionnaire-based assessment. Site selection is one of the highest-leverage decisions in trial operations. See also: Predictive Enrollment, Protocol Optimization, CRO (Life Sciences Glossary).

★ Software as a Medical Device (SaMD)
Software intended for a medical purpose that performs that purpose without being part of a hardware medical device. Most standalone clinical AI products fall here, and the pathway depends on risk classification and intended use. The FDA’s AI-Enabled Medical Device List passed 1,500 entries in early 2026, roughly three-quarters of them radiology. See also: Clinical Decision Support (CDS), Predetermined Change Control Plan (PCCP), 510(k) (Life Sciences Glossary).

Structured vs. Unstructured Data
Structured data is organized into defined fields, such as a lab value or an ICD code; unstructured data is free-form, such as a clinical note, a pathology report, or an image. Most clinically meaningful detail lives in unstructured form, which is why language and vision models unlocked so much health data. See also: Natural Language Processing (NLP) for EHR, AI-Ready Data, Interoperability and FHIR.

★ Synthetic Control Arm
A comparator group constructed from historical trial or real-world patient data rather than by randomizing living patients to a placebo. Regulators have accepted synthetic control arms in narrow settings — typically rare disease or oncology, where withholding treatment raises ethical concerns or where recruiting a control arm is impractical. See also: External Control Arm, Real-World Evidence (RWE), Rare Disease (Life Sciences Glossary).

Synthetic Data
Artificially generated data that reproduces the statistical structure of a real dataset without containing any real individual’s record. It is used to share data across institutions, augment scarce training sets, and test systems safely. Fidelity and privacy pull against each other: the closer it resembles the source, the more it can leak. See also: Differential Privacy, De-identification, Digital Twin.


T

★ Target Identification
The process of determining which biological molecule to act on to treat a disease. AI contributes by mining genomic, proteomic, literature, and clinical datasets for associations that human review would take far longer to surface. Identification is only the first half; target validation still requires experimental confirmation. See also: Multi-omics Integration, Variant Effect Prediction, Target (Life Sciences Glossary).

★ TechBio
A contested label for companies that treat computation and data as the core platform rather than as support for a single therapeutic asset. Competing definitions emphasize engineering headcount, platform-versus-asset business model, or simply a preference for technology-sector valuation framing. Because usage varies by speaker, treat it as a signal of positioning rather than a category with fixed criteria. See also: Foundation Model, Build vs. Buy, Biotechnology (Life Sciences Glossary).

Training Data
The dataset a model learns from, which determines what it can recognize and where it will fail. In life sciences the practical questions are representativeness, labeling quality, provenance, and whether the licensing permits the intended use. A model is bounded by its training data in ways that no amount of tuning corrects. See also: AI-Ready Data, Algorithmic Bias, Data Provenance.


V

Variant Effect Prediction
The computational assessment of whether a specific genetic variant is likely to disrupt protein function or cause disease. It addresses the large backlog of variants of uncertain significance produced by clinical sequencing. Predictions inform interpretation and are weighed alongside other evidence rather than used alone. See also: Protein Language Model, Genomics (Life Sciences Glossary).

Vector Database
A database that stores data as numerical embeddings and retrieves items by similarity rather than exact match. It is the retrieval layer beneath most retrieval-augmented generation systems, allowing a model to find conceptually related documents rather than keyword matches. It is infrastructure, not intelligence. See also: Retrieval-Augmented Generation (RAG), Inference.

Virtual Screening
The computational evaluation of large compound libraries to prioritize which molecules are worth physical testing. Machine learning has extended screening to libraries far larger than any laboratory could handle. Screening produces a ranked shortlist, and the hit rate is confirmed only at the bench. See also: Molecular Docking, Hit-to-Lead, In Silico.


Conclusion

Artificial intelligence is now part of the working vocabulary of California’s life sciences sector, and the terminology will keep moving. Some of these terms will settle into standard usage, several will be replaced, and the regulatory entries in particular will change as agencies publish new guidance.

Bookmark this glossary alongside the main California Life Sciences Glossary and come back as you encounter new terms. We update both as the industry and its language evolve.

This glossary is maintained by California Life Sciences. Definitions reflect common industry usage and are intended as a starting-point reference, not legal or regulatory guidance. Last updated: August 2026.



FAQ: AI in Life Sciences Glossary

An AI in life sciences glossary defines the artificial intelligence terminology used across drug discovery, clinical development, regulation, and manufacturing. This one is written for founders, operators, and investors rather than for machine learning engineers or clinicians. It covers 86 terms and cross-links to the main California Life Sciences Glossary.

Start with in silico, target identification, virtual screening, and protein structure prediction. Those four describe the sequence most AI discovery work follows: model the biology, choose what to act on, narrow the chemistry, then confirm at the bench. Generative chemistry and ADMET prediction come next.

No fully AI-designed drug has completed a Phase III trial or reached approval. The first candidates entered Phase III in 2026, and AI has genuinely compressed the discovery stage that precedes them. The clinical and regulatory path itself remains unchanged.

A predetermined change control plan, or PCCP, is a plan submitted to the FDA describing what modifications a manufacturer may make to an AI-enabled device software function after authorization. The FDA issued final guidance on PCCPs in 2025. It allows specified updates without a new marketing submission.

There is no single agreed definition. A technical reading focuses on whether data is complete, consistently formatted, and machine readable. A governance reading focuses on documented provenance, consent, and whether the intended use is permitted. Buyers should confirm which meaning a vendor is using before signing.