What Genomic Testing Companies Actually Do With Your DNA Data
July 7, 2026
Tens of millions of people have submitted saliva samples to direct-to-consumer genomic testing companies—23andMe, AncestryDNA, MyHeritage, and others—receiving ancestry reports, health predisposition assessments, and trait analyses in return. The appeal is straightforward: a fascinating personalised glimpse at your genetic data, delivered in a readable interface for a modest one-time fee.
The business model behind that simple transaction is more complicated than the marketing conveys. The DNA you submit generates data that has value beyond the report you receive—value for pharmaceutical research, aggregate studies, law enforcement, and future applications that don’t exist yet. The terms of service that govern what companies do with that data are detailed, frequently updated, and understood by a small fraction of the people who accept them. Here’s what’s actually happening with genomic data after you mail in the sample.
What the Companies Collect and Store
When you submit a DNA sample, the company genotypes it—identifying specific genetic variants at hundreds of thousands of positions in your genome. This is not sequencing (reading your entire DNA code); it’s targeted genotyping using microarray chips that look at known variable positions. The result is a dataset of your specific variants at those positions, typically several hundred kilobytes to a few megabytes in size.
This genotype data is stored, along with:
- The personal information you provided when registering (name, email, date of birth)
- Your survey responses and health information (which companies encourage you to provide for more personalised results)
- Your family connections within the service (relatives who have also tested)
- Your browsing and interaction data within the platform
The data is stored on cloud infrastructure. 23andMe, before its bankruptcy and subsequent acquisition proceedings, stored data on Amazon Web Services. Data encryption at rest and in transit is standard; internal access controls and audit logging are claimed in privacy policies. The practical security of this data against breach, legal process, and internal misuse depends on the company’s actual implementation of these controls—which users have no direct way to audit.
Research Data Sales and Partnerships
The most commercially significant use of aggregated genomic data is pharmaceutical and research partnerships. 23andMe signed a $300 million deal with GlaxoSmithKline in 2018, giving GSK access to the company’s customer database for drug target research. The principle: if you have a genotype dataset for millions of people linked to health survey responses, you can identify genetic variants associated with specific diseases and use those associations to identify potential drug targets.
The terms of 23andMe’s user agreement at the time of the GSK deal included opt-in consent for research participation—users who consented (the default was opt-in) had their data included in aggregated research datasets shared with partners. The “aggregated” qualifier matters: companies typically share derived research data (statistical associations, variant frequencies across populations) rather than individual-level data. But the meaningful analytical work—the research results—comes from the individual-level data the company holds.
23andMe’s therapeutic research division, which emerged from this model, attempted to use its proprietary dataset as a drug development asset—identifying genetic associations and building a pipeline of drug candidates based on the genomic insights from its customer database. This is a legitimate and scientifically interesting approach to drug discovery; it’s also a fundamentally different value proposition than “send us your spit and get an ancestry report,” and the connection between the two is not foregrounded in consumer marketing.
Ancestry and MyHeritage have similar research agreements with different terms and partners. The specific opt-in/opt-out mechanisms and the nature of data sharing vary by company and have changed over time as policies evolve.

Law Enforcement Access
Law enforcement use of consumer genomic databases has grown from theoretical concern to documented practice. The most significant legal mechanism is the warrant or court order compelling the company to provide individual customer data. The companies’ stated policies vary: 23andMe and AncestryDNA both publish transparency reports and claim they contest overly broad requests—but when presented with a valid warrant, they comply, as any company must under US law.
More significant is investigative genetic genealogy (IGG)—a technique where law enforcement uploads a suspect’s DNA profile to a consumer database and searches for relatives whose DNA matches. Those relatives can then be used to identify the suspect through conventional genealogy methods. GEDmatch, a platform that allows users to upload their genome files from any testing company, became a well-known tool for this technique after investigators used it to identify the Golden State Killer in 2018.
The legal framework for IGG is unsettled. GEDmatch, after the Golden State Killer case, changed its default to opt-out for law enforcement searches (the original default allowed it). Several states have passed laws regulating IGG; no comprehensive federal legislation governs it. The effectiveness of IGG as an investigative tool—it has solved numerous cold cases—exists in tension with the privacy implications of submitting one’s genome creating a searchable profile that implicates one’s relatives without their consent.
The relative consent problem is specific to genomic data: by submitting your DNA, you reveal information about your biological relatives who never consented to any interaction with the company. Your variants at risk-relevant positions partially reflect your parents’, siblings’, and children’s variants. Family members who haven’t consented to any testing have reduced genetic privacy because of a consenting relative’s decision.
The Bankruptcy Risk: What Happens to Data When Companies Fail
23andMe’s 2025 bankruptcy proceedings brought the data disposition question into sharp focus. The company’s genomic database—containing data from over 14 million customers—was identified as a significant asset in bankruptcy proceedings, raising the question of who would acquire that asset and under what data governance terms.
This is not a hypothetical scenario: when a company with a large data asset undergoes bankruptcy, that asset typically goes to the highest bidder, who may have entirely different privacy practices and business intentions than the original company. 23andMe’s terms of service addressed this scenario with language indicating that data could be transferred in a merger, acquisition, or bankruptcy, with the company promising to require the acquirer to maintain the same privacy practices—a promise that is difficult to enforce after the company ceases to exist.
Several state attorneys general sent letters to 23andMe during the bankruptcy proceedings urging customers to delete their data, and the California AG specifically noted that 23andMe’s privacy policy language created uncertainty about what protections would apply to data transferred in a bankruptcy sale. The episode is a concrete illustration of a risk that exists for any company holding large-scale sensitive personal data: the company you trusted may not be the entity that ultimately controls your data.
What Deleting Your Account Actually Means
Most genomic testing companies offer account deletion and data deletion options. Understanding what these options actually do requires reading the fine print carefully.
Account deletion typically means: the consumer-facing account is closed, the personal information associated with the account is deleted, and the raw genotype file is deleted (or moved to a state where it’s scheduled for deletion). What this may not mean: deletion of data that has already been included in research datasets or shared with research partners, deletion of backup copies on storage systems with retention periods, or deletion of aggregated derivative data.
The research opt-out process is separate from account deletion on most platforms. If you consented to research participation and then delete your account without first opting out of research, your data may remain in research datasets even after your account is closed. The order of operations matters: opt out of research first, then delete your account.
Physical sample destruction is another option that most companies offer at extra cost or upon specific request. The physical saliva sample (stored in a biobank after genotyping) is separate from the digital genotype data. Requesting physical sample destruction removes the ability to retest or generate additional data from your sample; it doesn’t affect already-generated digital data.

The Regulatory Landscape
Direct-to-consumer genomic data occupies a regulatory gap that has been slow to close. In the United States:
HIPAA (the Health Insurance Portability and Accountability Act) does not apply to consumer genomic testing companies unless they’re operating as covered healthcare entities—which direct-to-consumer companies like 23andMe typically are not. The HIPAA protections that people often assume apply to health data don’t apply to your genomic testing results.
GINA (the Genetic Information Nondiscrimination Act of 2008) prohibits discrimination by health insurers and employers based on genetic information—a meaningful protection. But GINA doesn’t apply to life insurance, disability insurance, or long-term care insurance, which can legally use genetic information in underwriting decisions in most US states. This creates a scenario where your genomic data, if accessible to an insurer in those categories, could affect your insurability or rates.
California’s CCPA (California Consumer Privacy Act) provides California residents with rights over their personal data including genetic data—the right to know what data is collected, the right to delete it, and the right to opt out of sale. Several states have passed or are considering similar legislation. Federal comprehensive privacy legislation in the US remains absent.
The EU’s GDPR provides stronger protections for European users: genetic data is classified as “special category” data requiring explicit consent for processing, with stronger deletion rights and restrictions on third-party sharing. European operations of these companies must comply with these higher standards.
The Risk Profile: Who Should Care Most
The practical privacy risk from consumer genomic testing is neither negligible nor catastrophic—it depends significantly on individual circumstances:
For most users, the primary risks are: data breach exposing genetic information to unknown parties, law enforcement access for those with legal concerns, and the long-term uncertainty of who controls the data in 10 or 20 years under circumstances no current policy can anticipate.
For specific populations, the risks are more acute: individuals with undisclosed biological relationships (adoptees, people with non-paternity situations, donor-conceived people) face identity-related risks from relatives discovering matches. Individuals with health conditions associated with specific genetic markers face potential discrimination risks in insurance contexts not covered by GINA. Individuals with immigration concerns face risks from law enforcement genealogy searches.
The consumer decision about whether to submit a DNA sample to these services is a legitimate personal choice with real tradeoffs. Making that choice with accurate information about what happens to the data—not just the marketing version—requires reading beyond the ancestry percentages on the homepage and into the privacy policy, research consent terms, and law enforcement transparency reports that companies publish but don’t emphasise.