Summary
On the allele data model branch (RT), a CSV export shows one ClinVar record per allele per release, but nothing guarantees there is only one. When there are two, the export keeps whichever the database happens to return last, so the same download can show a different clinical significance from one request to the next. This investigation finds out whether it happens in real data and decides which record wins. RT is explained in the glossary on #808.
What to find out
- On staging's backfilled data, count alleles with more than one live ClinVar link in the same release:
clinvar_allele_links joined to clinical_controls on clinvar_control_id, grouped by allele and db_version, with valid_to IS NULL.
- For a sample of them, whether the records disagree on clinical significance.
- Whether ClinVar itself can link one allele to two records in one release, or whether this only comes from our ingestion.
What to produce
A comment here with the counts, three examples, and a rule for choosing between records: the higher review status, the more recent record, or listing all of them. Plus the enforcement, either a tighter unique index or an ordering in the query.
Acceptance criteria
Background
_clinvar_by_allele in lib/csv/fetch.py builds {allele_id: control} with no ORDER BY, so the last row wins. The model comment on ClinvarAlleleLink says "one live link per (allele, release)", but uq_clinvar_allele_links_live is unique on (allele_id, clinvar_control_id), which allows two controls of the same release.
On main, release-2026.3.0 and 754 the same last-row-wins pattern applies per mapped variant. release-2026.3.0 adds a unique constraint on (db_name, db_identifier, db_version), which removes identical duplicates but not two different records.
Summary
On the allele data model branch (RT), a CSV export shows one ClinVar record per allele per release, but nothing guarantees there is only one. When there are two, the export keeps whichever the database happens to return last, so the same download can show a different clinical significance from one request to the next. This investigation finds out whether it happens in real data and decides which record wins. RT is explained in the glossary on #808.
What to find out
clinvar_allele_linksjoined toclinical_controlsonclinvar_control_id, grouped by allele anddb_version, withvalid_to IS NULL.What to produce
A comment here with the counts, three examples, and a rule for choosing between records: the higher review status, the more recent record, or listing all of them. Plus the enforcement, either a tighter unique index or an ordering in the query.
Acceptance criteria
Background
_clinvar_by_alleleinlib/csv/fetch.pybuilds{allele_id: control}with noORDER BY, so the last row wins. The model comment onClinvarAlleleLinksays "one live link per (allele, release)", butuq_clinvar_allele_links_liveis unique on(allele_id, clinvar_control_id), which allows two controls of the same release.On main, release-2026.3.0 and 754 the same last-row-wins pattern applies per mapped variant. release-2026.3.0 adds a unique constraint on
(db_name, db_identifier, db_version), which removes identical duplicates but not two different records.