FAIR Data Principles, Explained for Life Sciences Startups

Key Takeaways
  • The FAIR data principles say data should be findable, accessible, interoperable, and reusable, by machines as well as people. These rules predate the current AI wave by a decade.
  • FAIR does not mean public: proprietary data behind authentication is entirely FAIR provided the conditions for reaching it are clear to those who need to know.
  • Metadata is the highest-leverage place to start to ensure your data is FAIR. Metadata has to be captured while the work is happening rather than reconstructed later.
  • The European Commission put the minimum cost of non-FAIR data at €10.2 billion a year, with 96% of that cost in redundant storage and time lost searching.
  • Buying software does not make data FAIR: in the Pistoia Alliance’s 2025 survey, 81% of labs had an electronic lab notebook and 57% still named data silos their top obstacle.
  • FAIR is becoming a funding condition: California’s stem cell agency, CIRM, expects it of the companies it funds, and the National Institutes of Health built its data-sharing policy on the same four principles, including for small-business (SBIR) awards.
  • Treat FAIR as a maturity path rather than a pass or fail: fix the most valuable active dataset first to start building the muscles that make FAIR data a habit.

Data becomes hard to use long before it becomes big. Assay files accumulate in shared folders, sample names shift between experiments, `final_final_v3.xlsx` becomes load-bearing, and a promising result cannot be traced back to the inputs that produced it.

FAIR data principles are the best solution life sciences organizations have to prevent that data drift.

FAIR principles were first outlined in a 2016 paper in Scientific Data and the principles hold true today, but they’re not written for the real life realities of a life sciences start-up.

At California Life Sciences, we work with companies at every stage, from small, early-stage teams in FAST California to members with 20 years of accumulated research data. Data hygiene is a persistent challenge. If start-ups can take one lesson from more established players it’s to get out ahead of your data before it drifts. It’s not just a best practice. Increasingly it’s a requirement for funding: The California Institute for Regenerative Medicine expects its awardees to adhere to FAIR principles, and NIH built its data-sharing policy on the same four principles we’ll outline.

The goal for a small team probably shouldn’t be to build a perfectly structured institutional data repository. But the data teams generate needs structure and consistency enough that it can be found, understood, checked, and reused when it matters.

Here is what the four FAIR principles require, what ignoring them might cost, where accepting public funding already makes them a contractual condition, and what to do first.

1. What the FAIR data principles actually say

FAIR stands for findable, accessible, interoperable, and reusable. The principles were published in 2016 by Mark Wilkinson and 52 co-authors as The FAIR Guiding Principles for scientific data management and stewardship in Scientific Data.

Four words, with 15 sub-principles that get into the specifics: Data needs a globally unique and persistent identifier, and metadata indexed somewhere searchable. Protocols to access it must be open (not in the “publicly available” sense of the word) and documented. Vocabularies must be established and universally applied, and licenses for using that data must be clear.

Principle Sub-principle What it asks for
Findable F1 Data and metadata carry a globally unique, persistent identifier
F2 Data is described with rich metadata
F3 The metadata states the identifier of the data it describes
F4 Data and metadata are registered or indexed somewhere searchable
Accessible A1 Data is retrievable by its identifier over a standard communications protocol
A1.1 That protocol is open, free, and universally implementable
A1.2 The protocol allows authentication and authorization where necessary
A2 Metadata stays accessible even when the data itself is gone
Interoperable I1 Data uses a formal, shared, broadly applicable language for representing knowledge
I2 Its vocabularies themselves follow FAIR principles
I3 It includes qualified references to other data
Reusable R1 Data is richly described with many accurate, relevant attributes
R1.1 It is released with a clear, accessible usage license
R1.2 It carries detailed provenance
R1.3 It meets domain-relevant community standards

There’s a clear emphasis on machine-actionability: a computational system that’s given access must be able to understand the data to find, combine, and reuse it with minimal human intervention. The basic idea is that the volume of data will inevitably outpace what people can handle. FAIR wasn’t designed with machine learning in mind but machine learning specifically and AI generally are a clear case for why ensuring data is FAIR is a universal best practice in life sciences and many other disciplines.

The definitions are easier to understand as operational questions:

  • Findable Can the team locate a dataset and identify it reliably? A filename is not an identifier. A record with a stable dataset ID, a title, keywords, sample type, assay type, owner, and date is.
  • Accessible Can the data or its metadata be retrieved by a known method? The question is not “can everyone download it?” It is “can an authorized person tell where it is, who controls access, and how to obtain access?”
  • Interoperable Can data be combined across tools, teams, and partners without manual translation at every handoff? If one team records sex as `F`, another as `female`, and a third as XX, later analysis needs judgment calls nobody can see or repeat.
  • Reusable means someone can judge whether the data fit a new purpose, which depends on provenance, usage terms, and method detail. A result without a protocol version, instrument context, or processing history might be available but also scientifically weak.

FAIR data practices insulate from a very real challenge that faces start-ups: critical knowledge leaving with one person, or sitting in a file structure that only one person understands.

2. FAIR does not mean open

This is the point that decides whether a company engages with FAIR at all. GO FAIR, the implementation network for the principles, says so directly: FAIR is not equivalent to open, there are legitimate reasons data may be available only under conditions and only to certain users, and as long as those access conditions are properly described, non-open data can be entirely FAIR. The Accessible principle carries this explicitly, because the retrieval protocol it asks for is allowed to include authentication and authorization. Commercial confidentiality, patient privacy, and competitive position are all recognized grounds for gating access. If the gate is intentional and documented, the data is FAIR.

The principles also separate data from metadata. They expect metadata to be the more open of the two. Think of metadata as the label on a locked archive box. The label tells an authorized colleague what is inside, who created it, which study it belongs to, and how to request access. It does not reveal its contents to everyone. That is the arrangement controlled-access repositories such as dbGaP already run on, and the principles go further still: sub-principle A2 asks that metadata stay accessible even when the data itself is no longer available.

Question A FAIR answer What FAIR does not require
Can someone discover the dataset exists? Yes, through a searchable record or catalog Public download access
Can an approved user obtain it? Yes, through a documented access process Anonymous access
Can software read the key fields? Yes, through standard formats and terms where practical A custom enterprise platform
Can another team judge whether it fits a new use? Yes, through provenance, terms, and method detail Releasing confidential data

3. The cost of unFAIRness

That scattered data has a cost is clear. Specifics are harder to pin down. The best attempt to attach a price tag was commissioned by the European Commission and carried out by PwC in 2018. Cost of not having FAIR research data put the minimum cost to the European economy at €10.2 billion per year across five measured indicators.

Two causes account for 96% of the total: €5.3 billion in redundant storage, meaning paying repeatedly to keep copies of data nobody can tell apart, and €4.5 billion in time, meaning researchers searching for data that’s playing hard to get. Licence costs, duplicated funding, and retracted research make up the remaining 4%. The study also estimated that 31.52% of the time spent finding data could be saved under FAIR. The fact is that waste concentrates when data isn’t findable. Conveniently, findability is the cheapest of the four principles to act on.

A lack of findability in data might show up as this: Six months after comparing two candidate biomarkers, leadership asks whether the gap came from biology, sample handling, a revised protocol, or a pipeline update. Without a record the answer rests on memory. With one it traces to sample identifiers, protocol versions, and processing code, which is exactly what a partner’s diligence team asks for later.

The two causes that dominate the bill are both findability problems, and the three that barely register are the ones most compliance advice starts with.

No-Cost Startup Mentorship

FAST California is a 12-week, no-cost, equity-free mentorship program for early-stage life sciences founders. Applications are accepted on a rolling basis.

Learn More and Apply →

4. Start with metadata, not a new platform

Metadata is what explains a dataset: what it contains, how it was generated, which samples and methods it covers, who is responsible, what restrictions apply, and how it relates to other records. When that context lives in email threads and one scientist’s memory, the data gets less reusable with every project. A workable minimum record can start in a spreadsheet, a LIMS, or an electronic lab notebook. The format matters less than consistency and being able to search or export it.

Field Why it matters Example
Dataset ID A stable reference that survives a move `BIO-2026-0042`
Plain-language title Makes it searchable and reviewable Plasma proteomics pilot cohort
Owner and contributors Fixes accountability and credit Study lead, analyst, lab contact
Study purpose States the intended use Exploratory biomarker discovery
Sample and subject context Carries the biological meaning Human plasma, consented research participants
Protocol and assay version Makes the result interpretable Immunoassay protocol 2.1
Instrument and software detail Supports technical review Instrument model, pipeline version
Dates and processing history Documents provenance Collection date, transformation steps
Access conditions Prevents misuse Internal only; controlled collaborator access
Reuse terms States what others may do Internal research use with approval

Add fields because an informed person will need them later. A discovery assay needs reagent lot, sample preparation, and normalization method; a cell-line image set needs cell line identity, passage range, stain, and annotation rules. And so on.

5. Use identifiers that survive a team change

A persistent identifier is a reference that does not change when a file moves, a folder is reorganized, or the person who made it leaves.

Deposited data gets this for free: Zenodo issues DOIs at no cost, the NCBI repositories issue accessions, and contributors can use ORCID. For everything in-house, which is most of what a small company holds, the identifier is whatever you decide, and the requirement collapses into one discipline: assign an ID from a single documented scheme, once, and never silently reuse it.

The identifier should point to a record, not a file path. A path breaks when storage changes; a record can be repointed while keeping the dataset’s identity and history intact. So: assign the ID at intake, carry it into the metadata record and analysis notebooks, record relationships such as “derived from” or “used in analysis,” and keep prior versions visible rather than overwriting them.

You’re FAIR enough when: a new scientist can locate last year’s results without knowing who produced them.

6. Choose standards based on the data you actually have

Interoperability is where small teams overreach. No 15-person company needs every ontology in biomedical research. The question is which standards reduce friction for your next likely collaboration, submission, or analysis.

Use published identifiers and vocabularies where ambiguity could change meaning:

  • UniProt accessions for proteins
  • InChI or a registered identifier for compounds
  • OBO Foundry ontology terms for phenotypes and assays
  • CDISC standards for clinical data
  • FHIR where health records are involved

For genomics that also means recording genome assembly, reference version, and pipeline details; for multiomics, tying every layer to the same sample identifier. Adopting an existing vocabulary beats building one because this way, the decoding key is known. It doesn’t live with a team member who might take another job.

Machine-actionability is a direction but it’s not a day-one requirement. It does not start with a knowledge graph. The easiest way to get consistent is to replace free text fields with preset lists, standard units, and required fields wherever precision matters. A field called “collection time” should not hold a mixture of `am`, `morning`, `08:00`, and blanks. Decide the format, document it, and validate it at entry. That one decision prevents a large cleanup later.

7. Restricted is not the same as invisible

Sensitive human data create a real tension. You have to honor consent terms, privacy obligations, and contracts, and making the data invisible creates duplicate work and blocks legitimate collaboration. A tiered split usually resolves it:

  • Discoverable metadata: high-level study description, data type, access contact, and eligibility.
  • Controlled metadata: deeper technical detail for qualified internal users and approved collaborators.
  • Restricted data: participant-level or commercially sensitive files, reachable only through approved systems.

Where the lines fall depends on the study, the consent language, and your agreements, and no field list is safe for every human dataset. A broad cohort description is usually fine; rare disease characteristics, recruitment locations, or small subgroup counts can create disclosure risk on their own. Our checklist on whether your data is AI-ready covers the consent and ownership questions that set your own lines, because those are decided in agreements long before anyone builds a catalog.

8. Buying software will not make your data FAIR

Software and system shopping is a common response to a data findability problem. It’s reasonable but it’s not sufficient. The Pistoia Alliance’s Lab of the Future 2025 survey of 206 experts across pharma, biotech, software, services, academia, and non-profits in Europe, the Americas, and Asia-Pacific found electronic lab notebook use had risen to 81% from 66% the year before, and cloud data platform use to 80% from 70%. Data silos were still the top obstacle to making better use of lab data, named by 57% of respondents.

Tooling is not irrelevant, and silos did fall back from 2023. But even in the “Lab of the Future,” adoption ran well ahead of the problem it was bought to solve. In other words, the constraint lies elsewhere: whether identifiers are consistent, whether metadata is captured at generation, and whether what goes where is documented and understood. A new platform that faithfully reproduces data based on decisions that haven’t been documented doesn’t help. Conventions come before a system to enforce them. Reverse that order and you’re just paying more for data storage.

Four in five labs now run an electronic lab notebook, and a majority still cannot get at their own data.

9. FAIR is becoming a condition of funding, including in California

CIRM. The California Institute for Regenerative Medicine requires a Data Sharing and Management Plan, and its guidelines state that DISC awardees are expected to share data consistent with FAIR and CARE principles, reflecting practices in their research community. CARE, covering collective benefit, authority to control, responsibility, and ethics, sits alongside FAIR rather than inside it. CIRM runs its own findability layer, the CIRM Data Explorer, and expects your plan’s information to travel with the data when it is deposited.

NIH. The 2023 Data Management and Sharing Policy was built on the four FAIR principles and applies to research generating scientific data, including SBIR and STTR awards. It changed shape this year: under NOT-OD-26-046, issued in February 2026, the two-page narrative plan became a shorter form of largely yes-or-no commitments, required for due dates on or after 25 May 2026. The underlying policy did not change but the new form asks applicants to commit rather than describe.

Small businesses have one accommodation worth knowing. Under the Small Business Act, SBIR awardees may withhold data for 20 years from the award date. That protects the asset; it does not remove the expectation that the application addresses sharing. A plan stating that data will be withheld under the SBIR provision, with the metadata still described and a defined route for requests, is a FAIR answer. Doing the work retroactively is expensive.

Funder What it asks for Who it applies to
CIRM A Data Sharing and Management Plan, with data shared consistent with FAIR and CARE principles DISC awardees
NIH A Data Management and Sharing Plan on the shorter 2026 form, built on the four FAIR principles Any award generating scientific data, including SBIR and STTR
SBIR provision Sharing still has to be addressed, even where the data itself is withheld SBIR awardees withholding data up to 20 years from the award date

The Latest Life Sciences Updates

Stay current on the issues shaping California life sciences. Follow CLS on LinkedIn to never miss an update.

Follow on LinkedIn →

10. Where to start to make things FAIR

Nobody implements 15 sub-principles from a standing start, and FAIR is not a certification. A dataset can be easy to find and hard to reuse. Pick one active, valuable dataset, the one behind a lead program decision or heading into a partner review, and run it through five questions.

What to check

  • Can an authorized colleague find the dataset within ten minutes?
  • Can they tell what it contains without asking whoever made it?
  • Can they determine which protocol, sample set, software version, and transformations produced it?
  • Can they tell whether access is open, internal, restricted, or prohibited?
  • Can they reuse it without guessing what the labels, units, or missing values mean?

If the answer to any of these questions is no, there’s some work to do. Addressing metadata records and outlining clear ownership of the project is the easiest way to get to yes.

Level What it looks like Best next move
Basic Files exist; naming and ownership are inconsistent Create a dataset register and assign IDs
Managed Core metadata, owners, and access rules are documented Standardize required fields and versioning
Exchange-ready Common formats and controlled terms support handoffs Map fields to relevant community standards
Reuse-ready Provenance, terms, quality notes, and machine-readable metadata are maintained Automate validation on high-value workflows

Where to start

THIS MONTH

  • Pick one identifier scheme for samples, runs, and datasets, write it down, and apply it to all new work.
  • Register every dataset you have, what it holds, where it lives, and which agreement governs its use. A spreadsheet is a legitimate first register.
  • Document and share where data goes when it is generated within the team.

THIS QUARTER

  • Adopt published identifiers and ontologies for core entities, and record units explicitly.
  • Write the access process for each sensitive dataset, with a named owner.
  • Capture metadata at the point of collection.
  • Review data capture at decision points rather than on a calendar: before a handoff, a diligence request, or a model training run.

NOT YET

  • Retroactive cleanup of historical data, unless a submission, partnership, or grant requires it.
  • A data platform purchase, until the conventions exist for it to enforce.

STOP NOW

  • Generating data without metadata captured at the same time.
  • Inventing internal vocabularies where a published one exists.
  • Treating private or proprietary data as falling outside FAIR principles.

Legacy datasets often have provenance that cannot be reconstructed with confidence. If the gap can’t be filled accurately, it should be documented: a record noting that a processing version is unknown can save a lot of time and trouble down the road.

Three Free FAIR Resources for Life Sciences Organizations

  • The FAIR Cookbook, maintained by ELIXIR and the FAIRplus project, is life sciences specific and organized as more than 70 practical recipes by difficulty and reading time, and described in Scientific Data.
  • The Pistoia Alliance FAIR Toolkit and its FAIR Maturity Matrix were built by pharmaceutical companies and SMEs for industry rather than academic use.
  • F-UJI scores a public dataset against the core FAIR metrics at no cost, the fastest way to see the criteria applied to something real.

Conclusion: Making Your Data FAIR

FAIR arrived as a research-stewardship framework and is turning into infrastructure: a funder requirement, a diligence question, and the substrate every AI project runs on. The two most expensive failures — paying for redundant storage and time spent looking for data — are both findability problems, and findability is mostly a matter of choosing conventions and ensuring the whole team sticks with them. If the vocabulary here is unfamiliar, our AI in life sciences glossary covers FAIR data, provenance, and interoperability in plain English.

Explore CLS Membership

California Life Sciences membership connects you to California’s leading life sciences network — from advocacy and cost savings to exclusive events and programming.

See Benefits →

FAQ: FAIR data principles in life sciences

The FAIR data principles state that data should be findable, accessible, interoperable, and reusable, by machines as well as people. They were published in 2016 in Scientific Data by Mark Wilkinson and 52 co-authors, and sit on top of 15 sub-principles covering identifiers, metadata, access protocols, vocabularies, and licensing.

No. Open data concerns whether data can be publicly accessed. FAIR concerns whether data are described and structured well enough to be discovered and reused. A restricted clinical dataset can be fully FAIR if its metadata, access process, identifiers, provenance, and reuse conditions are clear to authorized users.

At minimum: a persistent dataset ID, title, purpose, owner, dates, sample context, methods, protocol versions, instrument and software detail, provenance, access conditions, and reuse terms. Add domain-specific fields where they change scientific interpretation, such as genome assembly for sequence data, or permitted values for clinical variables.

It is the document a funder asks for describing how you will handle and share the data an award generates. NIH calls it a Data Management and Sharing Plan and CIRM calls it a Data Sharing and Management Plan; both are built on the FAIR principles. As of NOT-OD-26-046, NIH applications with due dates on or after 25 May 2026 use a shorter form of largely yes-or-no commitments rather than a two-page narrative.

Yes. The NIH Data Management and Sharing Policy covers research generating scientific data, including SBIR and STTR awards. Under the Small Business Act, SBIR awardees may withhold data for 20 years from the award date, but NIH still expects the application to address sharing, so the plan should cite the withholding provision and describe the metadata and request route rather than say nothing.

FAIR makes data easier to discover, access, interpret, and reuse. Reproducibility asks whether a result can be obtained again from the same data, methods, code, and decisions. FAIR supports reproducibility by preserving context and provenance, but it cannot correct a weak study design, an unstable assay, or missing controls.

FAIR describes the machine-actionability a model needs from the data underneath it, which is why a 2016 stewardship framework became the reference for AI readiness. It covers the technical half. The other half is legal usability, meaning consent scope and data ownership, which our AI-ready data checklist addresses alongside it.