Types of Genetic Data: What Lies Behind SNPs, Haplotypes, and Polygenic Profiles

When you receive a genetic report, every result in it rests on a specific type of genetic data. For example, the predisposition to vitamin D deficiency shown in the report is not based on a single variant but on dozens of SNPs (single-nucleotide polymorphisms) at once: each one individually contributes very little, but together they form your individual picture. This data is not uniform: some indicators rest on isolated point variants in the genome, others on patterns of inheritance across entire chromosomal regions, and others still on the simultaneous analysis of hundreds or thousands of variants. The most common mistake is to assume that "genetics" is something singular and indivisible. In practice, the method by which data is collected determines what can and cannot be learned from it.
To understand where your result comes from, three concepts are enough: what a SNP is as the basic unit of genetic analysis, how SNPs group into haplotypes, and how a polygenic profile is built from them. We will go through each in turn.
How this data appears in your personal report and what the result scale means — How to Read the Apixmed Prism Report: What the Results Mean
SNP — The Smallest Unit of Genetic Difference
The human genome contains approximately 3 billion nucleotide pairs — the "letters" of the genetic code. In most of them, people are identical: differences between any two individuals account for less than 0.1% of the genome (1000 Genomes Project, Nature, 2015). Yet even this small fraction of variation spans millions of positions across the genome.
Think of the genome as a book of 3 billion letters. A SNP (Single Nucleotide Polymorphism) is a position where you have the letter A (adenine) while someone else has G (guanine). Most of the text is identical, but these differences shape your unique biology. Each such position has its own identifier — an rs number (for example, rs1801133): these are the numbers you see in the marker table of your genetic report.
Most SNPs do not "cause" disease or directly determine a trait on their own. They are genetic markers — signals associated with particular biological characteristics or predispositions. A single SNP typically explains a small fraction of the variability of a complex trait, which is why most indicators in the report are based not on one marker but on a whole set of them.
If you want to look up a specific marker in open databases such as dbSNP or ClinVar, the rs number from your report is a direct search key. There is no need to interpret each marker individually, however: the report already integrates them into a single assessment, and it is that assessment which carries clinical significance.

Haplotypes: SNPs That Are Inherited Together
SNPs are not evenly distributed across the genome. Variants located close together on the same chromosome tend to be inherited together. This phenomenon is called linkage disequilibrium, and a set of SNPs that is statistically passed from parents to children as a single block forms a haplotype.
Haplotypes are useful for a practical reason: instead of analysing each SNP separately, it is possible to identify a characteristic pattern of variants and use it to reconstruct information about the adjacent chromosomal region. This is the principle on which most commercial genotyping microarray chips are built: they read a limited set of "tag" SNPs, and use reference databases — such as the 1000 Genomes Project — to reconstruct the full picture of the region.
In medical genetics, haplotypes are particularly important for analysing genes with complex structures — for example, HLA genes (Human Leukocyte Antigen, responsible for the major histocompatibility complex), which determine predisposition to autoimmune diseases. Here it is not a single variant but the combination of alleles across the entire haplotype that produces the biological effect.
Polygenic Profile: From Individual SNPs to an Integrated Score
Most complex traits — including blood pressure, cholesterol levels, and predisposition to type 2 diabetes — are not determined by a single gene or a single SNP. They are shaped by thousands of variants across the genome, each contributing a small amount. The method that aggregates these contributions into a single numerical score is called a polygenic risk score, or PRS (Polygenic Risk Score).
The logic of PRS is straightforward: each SNP associated with a given trait in large population studies such as GWAS (Genome-Wide Association Study) is assigned a weight coefficient reflecting the strength of the association. These weights are then summed across the genome, and the resulting total is the individual polygenic score. A higher score means that, taken together, your genome contains more variants associated with an elevated risk of a given condition in the population.
A polygenic profile reflects relative predisposition compared with a reference population. This is exactly what you see on the result scale in the report: "lower than 76% of the population" or "above average." It is not a prediction of what will happen to a specific individual, because phenotype is also influenced by lifestyle, environment, and gene–gene interactions that PRS does not fully capture (Wand et al., Nat. Med., 2021).
How a polygenic profile is built, how to interpret it, and in which clinical contexts it is used — What PRS Is and What the Numbers in Your Report Mean

How These Data Types Connect in Your Report
SNPs, haplotypes, and polygenic profiles are not alternative approaches — they are levels of a single analytical process:
-
SNPs — the basic units read during genotyping or sequencing of your sample.
-
Haplotypes — patterns of adjacent SNPs that allow the full picture of a chromosomal region to be reconstructed from fewer direct measurements.
-
Polygenic profile — the result of aggregating hundreds or thousands of SNPs weighted by coefficients derived from population studies.
The indicators in the Apixmed Prism report are built primarily on polygenic profiles — that is, on the simultaneous analysis of many variants rather than on a single "key" SNP. The marker table on each indicator page displays the specific rs numbers that contributed most to the score: this ensures methodological transparency so you can see which variants underlie the conclusion.
Why the Data Collection Method Matters
Which SNPs are analysed and how depends on the technology. There are two main approaches:
-
Genotyping (microarray chip) reads a pre-defined set of hundreds of thousands or millions of selected SNPs. It is a fast, accessible method, optimal for polygenic analysis, GWAS, and most DTC tests (Direct-to-Consumer, meaning without a physician's referral). It does not, however, detect rare or novel variants that lie outside the chip.
-
Sequencing — in WGS (Whole Genome Sequencing) or WES (Whole Exome Sequencing, covering only the coding regions) format — reads DNA de novo, identifying variants regardless of whether they are known in advance. It produces a more complete picture but requires greater computational resources for interpretation.
The Apixmed Prism report supports both approaches. For most health traits covered by the report, polygenic analysis based on genotyping and on whole genome sequencing yields a comparable, clinically meaningful result, since the key variants used in the calculation are well represented by both methods. The difference between approaches becomes more apparent when searching for rare or novel variants, which falls outside the scope of standard polygenic analysis.
The technical differences between genetic data formats — VCF (Variant Call Format), BAM (Binary Alignment Map), FASTQ — and what happens to your sample after the laboratory — are covered in the article "Genetic Data Formats: What VCF, BAM Mean and Why It Matters" (Article #69).
Data Is the Foundation; Interpretation Is the Value
SNPs capture individual differences in the genome. Haplotypes show how those differences are inherited in blocks. A polygenic profile converts hundreds of individual signals into one summarised predisposition score. Each level provides its part of the answer, and together they form what you see as the result in the report.
Understanding this logic changes the way you read the report: the result is not one "key mutation" but a statistical portrait of your genome relative to a reference population. Genetic predisposition describes the biological context in which decisions about lifestyle, nutrition, and preventive check-ups are made — but it does not replace any of them.
If you have not yet chosen a test, explore the directions of genetic testing with Apixmed Prism.
Genetic test results are not a diagnosis and do not replace a consultation with a doctor. The Apixmed Prism report provides genetic context that complements clinical test results and supports informed decision-making together with your physician.
Sources
1. 1000 Genomes Project Consortium (2015). A global reference for human genetic variation. Nature, 526, 68–74.https://doi.org/10.1038/nature15393
2. Wand, H., Lambert, S. A., Tamburro, C., et al. (2021). Improving reporting standards for polygenic scores in risk prediction studies. Nature Medicine, 27, 1742–1748.https://doi.org/10.1038/s41591-021-01549-6
3. Sudlow, C., Gallacher, J., Allen, N., et al. (2015). UK Biobank: An Open Access Resource for Identifying the Causes of a Wide Range of Complex Diseases of Middle and Old Age. PLOS Medicine, 12(3), e1001779.https://doi.org/10.1371/journal.pmed.1001779












