Help & methods

Use Soybean mGWAS DB to move between metabolite features, GWAS results and related genes, then follow external resources to investigate biological hypotheses.

Getting started

Start with a metabolite feature

  1. Open Metabolites. Search a P identifier, candidate name, molecular formula or compound ID. A name search uses the stored candidate annotations.
  2. Open a result and check its deionized mass, average m/z, average retention time and adduct. The P identifier is the stable feature identifier.
  3. Review Candidate annotations. Compound IDs open their source records. Search PubChem and Soybean literature use the candidate name to search further.
  4. Inspect the Manhattan plot and GWAS summary, then follow the associated gene links. Compare Overlap and Within ±5 kb as two views of the supplied gene–feature relationships.

Start with a gene

  1. Open Genes. Search a Glyma identifier, alias, description or an Arabidopsis ID in the supplied annotation.
  2. Choose an association scope and inspect the related features. On the gene page, the text filter narrows the displayed features; the summary cards describe the full selected scope.
  3. Use Research resources to compare gene records, protein annotations and publications. The expandable section also links to gene-name translation and expression resources.
  4. Open promising features to examine their individual GWAS evidence and compound candidates. Export the gene's feature list to keep a record of the selected scope.

Develop a research hypothesis

Bring together the association pattern, the plausibility of a candidate compound, gene or protein annotations, and expression in relevant tissues. Record the feature ID, gene ID, scope, genome version and external record identifiers. Prioritisation can guide targeted metabolite identification or genetic and biochemical experiments.

Study context

The metabolite features originate from untargeted LC–MS profiling of soybean leaves. Soybean mGWAS DB connects these features with genetic association results and gene annotations. For metabolomics information, follow the Soybean Metabolome Repository link on an individual feature page.

Associations were analysed with a linear mixed model in the R package gaston, accounting for population structure, sampling date and genetic relatedness. Association testing used the Wald test.

The genotyping workflow used the soybean reference genome Gmax_275_v2.0. Check the genome assembly and gene annotation version when comparing genomic positions with external resources.

Data and terminology

Metabolite feature
A P identifier represents a measured feature. Several candidates may share a compatible mass, and several measured features may relate to the same compound.
Deionized mass and average m/z
Deionized mass is the supplied value used for candidate searching. Average m/z describes the measured ion; its interpretation depends on the adduct. Keep these fields distinct.
Average RT
The supplied average retention time. Compare retention times only with compatible chromatographic methods and units.
Candidate annotation
A result from molecular-mass searching, rather than confirmation of identity. The stored candidate annotations were generated with a 20 ppm tolerance. On-demand results are ordered by absolute mass difference (Δppm); this order is not an identification probability. Search again starts from the supplied deionized mass, with 20 ppm as the default tolerance. The search form accepts whole-number tolerances from 0 to 100 ppm.
Candidate databases
The configured mass search uses MFSearcher/UC2 through the Metabolome Repository service and queries FlavonoidViewer, HMDB, KEGG, KNApSAcK and LIPID MAPS. UC2 combines records by atom connectivity after charge normalisation; grouped database IDs do not establish stereochemistry. A database hit alone does not show that the compound occurs in soybean.
Associated genes and scope
SNP records below the selected P-value threshold are mapped to GFF3 gene intervals (Wm82.a2.v1, 1-based inclusive). Overlap includes positions from the gene start through its end; Within ±5 kb also includes positions up to 5,000 bp outside either boundary, on either strand. Only gene features are used; transcripts, CDS and UTR rows are not counted separately. This view does not recalculate LD blocks, fine-map a locus or establish a causal gene.
Supporting SNP records
The number of indexed SNP records below the chosen threshold mapped to a feature–gene pair in the selected scope. A gene or metabolite counts once in the relevant summary if it has at least one such record. One SNP may map to more than one gene; repeated records at a coordinate remain separate records. This is distinct from an effect size, gene-specific P value or count of statistically independent signals.
Minimum P
The smallest contributing SNP P value for that feature–gene pair and scope. It is not a gene-level test or a corrected P value.
Gene description and Arabidopsis ortholog
Annotations supplied with the gene table. External links help examine their supporting evidence; the site does not transfer a function from an ortholog automatically.

MFSearcher/UC2 documentation ↗

Reading GWAS evidence

Manhattan plots place SNP records along chromosomes and show −log₁₀(P) on the vertical axis. Higher points have smaller P values. The default display threshold is P < 1 × 10−10. In Settings, choose P < 10−10, 10−15, …, 10−50, in steps of five in the exponent; smaller P values are more stringent. SNPs exactly equal to the threshold are excluded. On narrow screens, fewer chromosome tick labels are shown to avoid overlap; all chromosome records remain plotted.

P values are read from the supplied GWAS files. The application does not recalculate association tests or apply a multiple-testing correction.

The display-file preparation retains SNP coordinates with P ≤ 1 × 10−6 and randomly selects one twentieth of the remaining coordinates, rounded down. This sampling threshold is separate from the application's display threshold. Counts and summaries refer to the records in the indexed file.

Plots and summaries describe the indexed GWAS file for the feature. The minimum P value is a feature-level summary across that file, not a P value assigned to each related gene. Nearby associated SNPs can be correlated through linkage disequilibrium; their count is not a count of independent loci.

File coverage, Q–Q plots and zero P values

Files with “downsample” in their names are marked as downsampled. Their record counts and Q–Q plots describe the available records. Selection based on P values changes their distribution, so a Q–Q plot of such selected records cannot be used to assess calibration across all tested variants. Use the complete test results for genome-wide P-value distribution and population-structure diagnostics.

Source-reported P = 0 values are preserved in downloads, plotted at −log₁₀(P) = 300 and excluded from Q–Q calculations. The plot cap does not supply their unknown exact P values. This release does not estimate genomic inflation (λGC).

Functional annotation, expression and biochemical context provide evidence for follow-up. Distinguishing a causal gene or metabolite requires evidence beyond proximity and a small association P value.

Tables and numeric display

Tables default to 10 rows; choose 30, 50 or 100 when useful. Explorer and gene-page filters and page settings are retained in their URLs. On a metabolite page, paging the gene, candidate and SNP tables stays within the page so the GWAS plot state is retained.

Overview and explorer results are ordered by the relevant association count. Counts are not a statistical-significance ranking. Ordinary non-integer values use four decimal places; integers have no decimal suffix and very small values use scientific notation. Missing values are shown as an em dash. API and download values are not rounded to the display format.

Downloads and stable links

OutputWhat it contains
Gene → Export all at this threshold (TSV)All associated features satisfying the selected P threshold and scope, including features beyond the current page. The text filter on the gene page does not restrict this export.
GWAS → Download Summary TSVSummary values for the feature's indexed GWAS file, including its threshold and reported source coverage.
GWAS → Download SNPs below threshold (TSV)All below-threshold SNP records returned from that indexed file, across all displayed table pages.

Keep the feature URL when sharing a record. The /metabolite/Pxxxxx/ and /Pxxxxx/ paths refer to the same feature. Gene pages use /gene/Glyma.…/. Retain the site's deployment prefix when constructing links.

Display settings

Settings stores your GWAS display threshold on this browser for this site. The same choice is used when you open another metabolite from a search or gene page. Domains and test/production paths keep separate preferences; clearing site data restores P < 10−10 for URLs without an explicit p_exp. Shared links retain their explicit threshold, including legacy values from earlier releases. Older saved browser preferences outside the nine current choices return to the default.

The threshold updates the Manhattan highlight and line, below-threshold SNP count and percentage, chromosome counts, SNP table and GWAS downloads. It filters the available P values without recalculating them. Background points and the Q–Q plot still use all applicable indexed records. The least stringent available choice is P < 10−10.

Gene–metabolite associations, their supporting SNP counts, overview totals, search filters, result ordering and threshold-aware TSV exports follow the same setting. Genes are linked by genomic proximity to below-threshold SNPs; no gene-level significance test is performed. This filter is not a new statistical test or multiple-testing correction. Features without an indexed GWAS file and genes without a matching GFF interval show unavailable counts (—), not zero. Gene-side totals cover the 30345 indexed features out of 55607 catalogue features.

GWAS download filenames include the selected exponent; the summary TSV also records the threshold. The JSON API retains its fixed default threshold for compatibility and does not read browser preferences.

The threshold-aware gene export records the threshold, scope, genome annotation and minimum contributing P. The original /gene/ID/associations.tsv route still returns the legacy source-table records for compatibility; its values must not be interpreted as dynamically filtered counts. The paginated /api/metabolite/Pxxxxx/associations/?p_exp=10&scope=overlap route returns threshold-dependent gene links. API calls without p_exp use exponent 10 and do not read browser preferences.

Programmatic access

The existing JSON routes /api/metabolite/Pxxxxx/cached-annotations/ and /api/metabolite/Pxxxxx/manhattan/ expose stored candidates and indexed GWAS records. Use the actual P identifier and the same site prefix. A missing GWAS file returns a missing-resource response; it is not evidence of a nonsignificant result.

The API fields significant_points and is_significant retain the fixed P < 10−10 condition. In browser-generated GWAS downloads, labels containing “Significant” refer to the selected display threshold, which is also recorded in the summary. Existing field and column names are retained for compatibility.

Research resources

Links marked ↗ open another website in a new tab. Identifier searches can return several records or no result; inspect the species, assembly, identifiers and evidence in the destination.

Further reading: interpreting plant association results

These studies provide background for interpretation and do not describe the source dataset of this database.

Common questions

Why does one feature have several candidate compounds?

Different compounds can have compatible masses. Examine the original compound records and use additional analytical evidence, such as suitable MS/MS and reference-standard comparisons, when identifying a feature.

Why is a compound name found on one page but not in a later name search?

Explorer searches use the stored annotation list. Search again queries the external service for the current mass and tolerance; these results are not written back to the persistent explorer index by this action.

What does an empty result mean?

No candidate annotation means no available candidate in that search result. For a feature with indexed GWAS data, no associated gene means no SNP record below the selected threshold maps to a gene in that scope. Unavailable data are shown separately. Neither result proves absence of the compound or absence of genetic regulation.

Why can the gene and SNP tables show different counts?

Both use the selected P threshold. SNPs without a gene in the selected scope remain in the SNP count; several SNPs may map to one gene, and one SNP may map to several genes. A gene count is a count of distinct linked gene identifiers, not a SNP count or a gene-level test. Changing the positional scope changes gene links but does not change which SNPs pass the P threshold.

Can supporting records or minimum P tell me the direction of an effect?

No. Neither field gives an allele effect or direction. Effect interpretation requires the appropriate effect estimates, allele definitions and study methods.