[G×I] predicts 200+ candidate regulators of GLP1R expression
[G×I] expression model identifies more than 200 variants with strong predicted effects on GLP1R expression.
In Part 1, we described a genetic variant associated with greater weight loss on Ozempic and other GLP-1 medications. We also introduced a workflow that lets people search their own genomic data for related variants.
We started with rs10305420, the strongest treatment-efficacy signal reported in the Nature study. The paper proposed that the variant alters the GLP-1 receptor protein. We asked whether it might also change how much of the receptor is produced. In this second part of the series, we describe the analysis in detail and expand it from a single variant to more than 2,700 nearby DNA changes.
rs10305420 variant sits in the GLP1R control region
We located rs10305420 on the GRCh38 human genome assembly at chromosome 6, position 39,048,860.
The GLP1R transcript used in the analysis, ENST00000373256, begins at position 39,048,781. This places rs10305420 only 79 base pairs from the point where GLP1R transcription starts, inside the gene’s first exon and promoter region.
This location is important. A promoter is a DNA region that helps control how much of a gene is expressed. Although rs10305420 changes the GLP-1 receptor protein, it is also positioned where it could affect GLP1R activity.
Public regulatory data support this possibility.
Firstly, ENCODE annotations classify the region as an active promoter-like element.
rs10305420 overlaps an active promoter-like regulatory element near the GLP1R transcription start site.Secondly, the coordinate of rs10305420 overlaps:
- ATAC-seq and DNase-seq signals, which mark accessible DNA;
- transcription-factor and chromatin-regulator binding sites;
- H3K4me3 and H3K27ac histone marks associated with active promoters;
- regulatory signals observed in pancreas and endocrine-pancreas samples.
rs10305420 is accessible and carries active-promoter marks, including in pancreas-related samples.Predicting effect of C to T substitution in rs10305420 variant with [G×I] expression
The [G×I] expression model takes a DNA sequence around the start of a gene, together with a description of the biological context, and predicts gene expression as log(TPM+1). Here, TPM, or transcripts per million, is a normalized measure of RNA abundance that estimates how strongly a gene is expressed within a sample. The logarithmic transformation compresses large differences in expression and allows genes with zero measured expression to be included.
For GLP1R, we used a 9,198-base-pair sequence centered on the transcription start site. We created two versions that differed at only one position:
- the reference sequence with the C allele;
- the alternative sequence with the T allele associated with greater weight loss.
The model predicted higher GLP1R expression for the T allele in all three pancreas-related contexts we tested.
rs10305420 T allele increases predicted GLP1R expression in the pancreas-related contexts tested. The model also predicts that the effect depends on biological context, including a small shift in the opposite direction in a brain context.The changes are modest but consistent. Expressed on the more intuitive TPM scale, the predicted increase ranges from approximately 4% to 9%, depending on the pancreas context.
| Context | Ref C log(TPM+1) | Alt T log(TPM+1) | Delta | Predicted TPM change |
|---|---|---|---|---|
| endocrine pancreas tissue | 0.2520 | 0.2715 | +0.0195 | +8.9% |
| body of pancreas tissue | 0.6367 | 0.6641 | +0.0274 | +5.9% |
| pancreas tissue | 0.6914 | 0.7109 | +0.0195 | +4.0% |
This suggests a second possible mechanism for rs10305420. The same DNA change may affect both the structure of the GLP-1 receptor and the amount of receptor produced. Thus the both could contribute to the observed difference in treatment response.
Searching for other genomic variants increasing GLP1R expression
Large population studies are best at detecting common variants. A rare variant may have an important biological effect but remain invisible because too few carriers are present in the study.
Sequence models offer a complementary approach: they can estimate the effect of a variant directly from its DNA sequence, without requiring thousands of people who carry it.
We screened 2,734 single-base changes within the fixed 9,198-base-pair GLP1R model window. Each alternative allele was scored in the endocrine-pancreas context.
We used the predicted effect of rs10305420, |delta log(TPM+1)| = 0.0195, as a benchmark.
[G×I] scan of GLP1R variants found 229 candidates exceeding expression effect of rs10305420.
The model identified 229 variants with a predicted effect larger than rs10305420. Of these, 111 have already been observed in the gnomAD human variation database.
The strongest rare variants
Several of the highest-scoring variants already observed in human populations are located within a few hundred bases of the GLP1R transcription start site.
| Variant | rsID | Distance to TSS | Delta log(TPM+1) | gnomAD AF |
|---|---|---|---|---|
6-39048550-C-T | rs1183935109 | -231 bp | +0.1738 | 0.0000066 |
6-39048794-C-G | - | +13 bp | +0.1074 | 0.0000014 |
6-39048582-T-G | rs1768014910 | -199 bp | +0.1035 | 0.0000066 |
6-39048794-C-A | rs10305419 | +13 bp | +0.1015 | 0.0000068 |
6-39048579-G-A | rs182808054 | -202 bp | +0.0898 | 0.00047 |
Each variant is rare by itself. Together, however, the benchmark-exceeding rare variants correspond to an estimated carrier frequency of approximately 6.5% - roughly 1 in 15 people, or around 527 million people worldwide.
This estimate uses a simple Hardy-Weinberg and independence approximation. It ignores linkage between variants and should be interpreted as a prioritization estimate, not as the number of people with altered medication response.
An important warning from rs880067
The scan also identified a common promoter variant, rs880067, with a large predicted increase in GLP1R expression.
Because this variant is common and located near the transcription start site, it would be easy to present it as an explanation for differences in GLP-1 medication response. We do not think the evidence supports that conclusion yet.
An inheritance-correlation analysis found that rs880067 and rs10305420 have high D’ but only moderate r^2 in European ancestry data, at approximately 0.22. In simple terms, the variants are often inherited in related genetic backgrounds, but one does not reliably predict the other.
Variants that track very closely with rs880067 also had much smaller predicted expression effects. This makes the rs880067 result look more like a position-specific model prediction requiring experimental validation than a ready-made explanation of treatment response.
This is an important example of why model predictions must be interpreted carefully rather than treated as biological facts.
From in silico screening to experimental validation
New sequence to expression genomic models make it possible to screen thousands of variants in silico before committing to costly laboratory experiments. Here, the [G×I] expression model suggests that rs10305420 may act through two compatible mechanisms: (1) altering the GLP-1 receptor through the p.Pro7Leu substitution and (2) changing GLP1R expression.
The same approach extends beyond a single GWAS hit. By scoring 2,734 nearby variants directly from sequence, we identified hundreds with predicted effects comparable to or greater than that of rs10305420, including rare variants that population studies may be underpowered to detect. These predictions provide a practical way to prioritize candidates for functional testing.
The next step is wet-lab validation. Pancreatic islet and endocrine-cell studies could test whether these alleles alter GLP1R expression, while MPRA, reporter assays, base editing and CRISPR experiments could evaluate their regulatory effects directly.
We are seeking experimental partners with relevant cell models, functional-genomics platforms or expertise in GLP-1 biology to help test these hypotheses. If your lab is interested in validating selected variants or developing a larger screening study, we would welcome a collaboration. Dataset of 229 variants with a predicted effect larger than rs10305420 is avalable over this link.
You can download the workflow and apply the same analysis to your own genomic data. You can also try the [G×I] expression model at app.genomicintelligence.ai/expression, or run the workflow through the REST API or MCP server.
Disclaimer: These results are computational hypotheses, not clinical predictions or a medical test. They should not be used to select, compare, dose, start or stop any medication.