Untitled
Apixmed Prism — an innovative next-generation genetic analytics platform
Home keyboard_arrow_right Blog keyboard_arrow_right
From Raw Data to Report Metrics: How Genotyping Data Becomes a Genetic Profile
Technology
Publication date:
5 minutes of reading

From Raw Data to Report Metrics: How Genotyping Data Becomes a Genetic Profile

A wave of scattered light-blue microspheres organizing into a structured surface with an orange glow — main cover image for the genetic data processing article.

A number in a genetic report looks simple. But behind it lie several stages of data processing: from determining genotypes and checking their quality to statistical models and comparison with a reference population.

For polygenic metrics, this process includes genotyping or whole-genome sequencing, quality control, imputation when needed, the use of GWAS data, and the calculation of PRS. The result is then compared with a reference population to determine its relative position — for example, as a percentile.

Let's look at this process step by step — from a saliva sample and obtaining genotyping data to the metric that ultimately appears in your report.

From a Saliva Sample to Genotyping Data

After collecting a saliva sample, the genetic material is most often analyzed using genotyping, which determines genotypes at a large number of specific DNA positions and generates a dataset for further analysis.

The resulting genotype dataset undergoes quality control: the technical reliability of the signal is checked for each sample and marker, along with other parameters that determine whether the data is fit for further analysis (Zhao et al., J. Genomics, 2022). Such checks usually rely on several statistical criteria at once rather than a single metric (Treccani et al., BIOCELL, 2023). The results of quality control determine which data can be used in the next processing stages. Data that fail these checks are excluded from further analysis: this may be a single marker with an unreliable signal or, less often, an entire sample if the proportion of problematic results is too high.

What Genotyping Actually Determines

Genotyping determines genotypes at a large number of specific DNA positions — including single nucleotide polymorphisms (SNPs), that is, positions where individual nucleotides can differ between people. Unlike sequencing, which determines the DNA sequence across a wide range of positions, genotyping works with a predefined set of genetic markers.

Apixmed Prism analyzes about 700,000 genetic variants — the resulting dataset covers positions linked to various traits and conditions, and becomes the basis for the further statistical interpretation of metrics.

For more on how to choose a genetic testing method, see the article WGS, WES, and Genotyping: How to Choose a Genetic Testing Method.

Image 2

How Imputation Expands Genotyping Data

Genotyping covers a large number of DNA positions, but certain statistical models may require variants that were not directly determined. Imputation makes it possible to statistically estimate such genotypes using patterns of co-inheritance of genetic variants and a reference panel — a set of previously fully characterized genomes (Treccani et al., BIOCELL, 2023).

The result of imputation can be probabilities of different genotypes or estimates of the copy number of a particular allele — numerical values that are then used in statistical models (Treccani et al., BIOCELL, 2023). Imputed values are of a different nature than directly measured genotypes: they are statistical estimates of the genotype at a position that was not measured directly.

Where the Statistical Weights for a Metric Come From

To combine information on thousands of genetic variants into a single metric, models use weights that reflect the magnitude and direction of each variant's effect on a given trait. These estimates come from genome-wide association studies (GWAS) — analyses of large samples that look for associations between genetic variants and specific traits or conditions (Uffelmann et al., Nat. Rev. Methods Primers, 2021).

For example, if a particular variant is associated with only a slight shift in the trait in the studied population, its weight in the model will be small. A variant with a more pronounced statistical association will receive a greater weight. Whether it is linked to an increase or a decrease in the trait depends on the sign of the effect (Uffelmann et al., Nat. Rev. Methods Primers, 2021).

A single variant usually explains only a small part of the genetic variability of a complex trait. A polygenic model combines the contributions of a large number of variants, using an appropriate statistical weight for each one (Uffelmann et al., Nat. Rev. Methods Primers, 2021). The quality of these estimates depends, among other things, on the size and characteristics of the GWAS sample, how well the studied trait is defined, and the statistical power of the study.

Image 3

What a Metric in the Apixmed Prism Report Means

A metric in the Apixmed Prism genetic report is the result of sequential processing and statistical interpretation of genetic data. For polygenic models, it is calculated from genotypes and statistical weights: each risk variant is weighted proportionally, and the contributions of all associated variants are summed into a single metric (Choi et al., Nat. Protoc., 2020). Its interpretation depends, among other things, on the reference population — comparison with it is what turns the raw score into a percentile, i.e., a person's relative position among the values of that population, rather than the probability of developing a particular condition (Kullo, Nat. Rev. Genet., 2025).

For details on what specific metrics and sections of the finished report mean, see the article How to Read an Apixmed Prism Report: What the Results Mean.

Genetic data can be used for further interpretation without collecting a new sample: a person's DNA profile remains stable throughout life, while scientific data and statistical models continue to evolve. This means that over time the same genotyping data can become the basis for new metrics or for revising existing results — for example, as GWAS studies refine statistical weights or as models emerge for traits that were not previously analyzed.

For practical steps after receiving your report, see the article What to Do With a Genetic Report: From Data to Decisions.

Genetic data becomes part of a broader picture of health only when it can be interpreted in context: the genotype dataset alone, without a model and a reference population, does not provide practically applicable information. Check out the available Apixmed Prism testing directions, to see which genetic data can complement your overall picture of health information.

 

The results of a genetic test are not a diagnosis and do not replace a doctor's consultation. The Apixmed Prism report provides genetic context that complements examination results and helps you make decisions together with your doctor.

 

Sources

1. Zhao, S., Jiang, L., Yu, H., Guo, Y. (2022). GTQC: Automated Genotyping Array Quality Control and Report. Journal of Genomics, 10, 39–44. https://doi.org/10.7150/jgen.69860 

2. Treccani, M., Locatelli, E., Patuzzo, C., Malerba, G. (2023). A broad overview of genotype imputation: Standard guidelines, approaches, and future investigations in genomic association studies. BIOCELL, 47(6), 1225–1241. https://doi.org/10.32604/biocell.2023.027884 

3. Uffelmann, E., Huang, Q. Q., Munung, N. S., de Vries, J., Okada, Y., Martin, A. R., Martin, H. C., Lappalainen, T., Posthuma, D. (2021). Genome-wide association studies. Nature Reviews Methods Primers, 1, 59. https://doi.org/10.1038/s43586-021-00056-9 

4. Choi, S. W., Mak, T. S. H., O'Reilly, P. F. (2020). Tutorial: a guide to performing polygenic risk score analyses. Nature Protocols, 15(9), 2759–2772. https://doi.org/10.1038/s41596-020-0353-1 

5. Kullo, I. J. (2025). Clinical use of polygenic risk scores: current status, barriers and future directions. Nature Reviews Genetics. https://doi.org/10.1038/s41576-025-00900-8 





Notes

Image 1: illustration of transformation — scattered, randomly arranged small elements (dots, particles) gradually shift into an ordered, structured form, for example toward a clear gradient or an organized pattern. The idea is movement from raw, scattered data to an ordered result.

Title: From Scattered Genetic Data to an Ordered Metric

Alt: Conceptual image of the transformation of genetic data into a report metric

 

Image 2: illustration of incompleteness and completion — a pattern or grid in which most elements are sharp and solid, while some are softer, semi-transparent, as if “filled in”. The idea is the contrast between what was directly measured and what was statistically estimated.

Title: Direct Measurement and Statistical Completion of Data

Alt: Conceptual image of the principle of genetic data imputation

 

Image 3: illustration of combination — many small elements of varying size or intensity converge into a single point or shape, symbolizing the summation of many individual contributions into one result. 

Title: From Many Statistical Contributions to One Metric

Alt: Conceptual image of the summation of genetic data into a metric

A 3D lattice of translucent molecular nodes with blue elements and a single central node highlighted in orange, illustrating genotype imputation.

Як імпутація розширює дані генотипування

Генотипування охоплює велику кількість позицій ДНК, однак для окремих статистичних моделей можуть бути потрібні варіанти, які безпосередньо не визначалися. Імпутація дає змогу статистично оцінити такі генотипи, використовуючи закономірності спільного успадкування генетичних варіантів і референтну панель — масив раніше повністю охарактеризованих геномів (Treccani et al., BIOCELL, 2023).

Результатом імпутації можуть бути ймовірності різних генотипів або оцінки кількості копій певного алеля — числові значення, які надалі використовують у статистичних моделях (Treccani et al., BIOCELL, 2023). Імпутовані значення мають іншу природу, ніж безпосередньо виміряні генотипи: це статистичні оцінки генотипу в позиції, яка не була виміряна напряму.

Звідки беруться статистичні ваги для показника

Щоб об'єднати інформацію про тисячі генетичних варіантів для одного показника, моделі використовують ваги, які відображають величину та напрямок ефекту кожного варіанта на певну ознаку. Такі оцінки дають повногеномні асоціативні дослідження (GWAS) — аналізи великих вибірок, у яких шукають зв'язок між генетичними варіантами та конкретними ознаками чи станами (Uffelmann et al., Nat. Rev. Methods Primers, 2021).

Наприклад, якщо певний варіант асоційований із незначним зсувом ознаки в дослідженій популяції, його вага в моделі буде малою. Варіант із помітнішим статистичним зв'язком отримає більшу вагу. Чи пов’язаний він зі збільшенням прояву ознаки, чи зі зменшенням, залежить від знака ефекту (Uffelmann et al., Nat. Rev. Methods Primers, 2021).

Окремий варіант зазвичай пояснює лише невелику частину генетичної варіативності складної ознаки. Полігенна модель об'єднує внески великої кількості варіантів, використовуючи для кожного відповідну статистичну вагу (Uffelmann et al., Nat. Rev. Methods Primers, 2021). Якість цих оцінок залежить, зокрема, від розміру та характеристик вибірки GWAS, якості визначення досліджуваної ознаки та статистичної потужності дослідження.

Multiple blue and transparent spheres converging into a single glowing orange focal sphere, symbolizing polygenic risk score calculation.

Що означає показник у звіті Apixmed Prism

Показник у генетичному звіті Apixmed Prism — це результат послідовної обробки та статистичної інтерпретації генетичних даних. Для полігенних моделей він формується на основі генотипів і статистичних ваг: кожен варіант ризику враховується пропорційно до своєї ваги, і внески усіх асоційновиних варіантів підсумовуються в єдиний показник (Choi et al., Nat. Protoc., 2020). Його інтерпретація залежить, зокрема, від референтної популяції — саме порівняння з нею перетворює сирий показник на перцентиль, тобто відносне положення результату людини серед значень цієї популяції, а не ймовірність розвитку певного стану (Kullo, Nat. Rev. Genet., 2025).

Про те, що означають конкретні показники та розділи готового звіту, — у статті Як читати звіт Apixmed Prism: що означають результати.

Генетичні дані можуть використовуватися для подальшої інтерпретації без повторного збору зразка: ДНК-профіль людини залишається стабільним протягом життя, тоді як наукові дані та статистичні моделі продовжують розвиватися. Це означає, що з часом ті самі генотипні дані можуть стати основою для нових показників або для перегляду вже наявних результатів — наприклад, у міру того, як GWAS-дослідження уточнюють статистичні ваги або з'являються моделі для ознак, які раніше не аналізувалися.

Про практичні кроки після отримання звіту — у статті Що робити з генетичним звітом: від даних до рішень.

Генетичні дані стають частиною ширшої картини здоров'я лише тоді, коли їх можна інтерпретувати в контексті: сам масив генотипів, без моделі й референтної популяції, не дає практично застосовної інформації. Перегляньте доступні напрями тестування Apixmed Prism, щоб побачити, які генетичні дані можуть доповнити вашу систему інформації про здоров'я.

Результати генетичного тесту — це не діагноз і не заміна консультації лікаря. Звіт Apixmed Prism дає генетичний контекст, який доповнює результати обстежень і допомагає приймати рішення разом із лікарем.

Джерела

1. Zhao, S., Jiang, L., Yu, H., Guo, Y. (2022). GTQC: Automated Genotyping Array Quality Control and Report. Journal of Genomics, 10, 39–44. https://doi.org/10.7150/jgen.69860 

2. Treccani, M., Locatelli, E., Patuzzo, C., Malerba, G. (2023). A broad overview of genotype imputation: Standard guidelines, approaches, and future investigations in genomic association studies. BIOCELL, 47(6), 1225–1241. https://doi.org/10.32604/biocell.2023.027884 

3. Uffelmann, E., Huang, Q. Q., Munung, N. S., de Vries, J., Okada, Y., Martin, A. R., Martin, H. C., Lappalainen, T., Posthuma, D. (2021). Genome-wide association studies. Nature Reviews Methods Primers, 1, 59. https://doi.org/10.1038/s43586-021-00056-9 

4. Choi, S. W., Mak, T. S. H., O'Reilly, P. F. (2020). Tutorial: a guide to performing polygenic risk score analyses. Nature Protocols, 15(9), 2759–2772. https://doi.org/10.1038/s41596-020-0353-1 

5. Kullo, I. J. (2025). Clinical use of polygenic risk scores: current status, barriers and future directions. Nature Reviews Genetics. https://doi.org/10.1038/s41576-025-00900-8 

Any questions left?

Leave your contact details — our specialists will contact you shortly. We will help you understand the offers, select a test according to your request, and answer any questions about the product and process.