Skip to content

feat: Add Allele data model and MappingRecord migration #739

Description

@bencap

Context

Part of the Better Reverse Translation feature. This issue establishes the complete new data model, replacing MappedVariant with MappingRecord + a flat Allele table, and formalising the annotation architecture.

New Data Model

MappingRecord (replaces MappedVariant)

Provenance record — one per AssayedVariant per mapping run.

  • FK to AssayedVariant
  • vrs_digest — VRS digest of the pre-mapped (assayed-level) variant; indexed for fast search without JSONB parsing
  • pre_mapped JSONB (raw pipeline input blob)
  • assay_level (genomic | cdna | protein) — the level at which this variant was assayed
  • hgvs_assay_level — HGVS string at the assayed level (stable, lives here not on the allele)
  • mapping_api_version, mapped_date, vrs_version, current
  • at_mismatched_locus, near_gap, target_gene_mapping_id (QC fields from dcd-mapping)
  • M:N to Allele rows via mapping_record_alleles association table

Allele (flat table — no inheritance)

One row per unique variant, deduplicated by VRS digest across all score sets.

Column Type Notes
id PK
vrs_digest unique
level enum genomic | coding | protein
transcript NOT NULL present for all levels — anchors the translation
hgvs_g nullable genomic HGVS string; populated in post-processing
hgvs_c nullable coding HGVS string; populated in post-processing
hgvs_p nullable protein HGVS string; populated in post-processing
clingen_allele_id nullable populated where available, all levels
post_mapped JSONB raw mapper output for this allele at this level

HGVS fields are nullable at insert time and filled by a post-processing step. Correctness (e.g. that a genomic allele eventually has hgvs_g) is enforced at the application layer, not via DB check constraints.

Schema rule: Fields that are stable by construction (HGVS strings, ClinGen IDs once assigned) live as columns on allele/mapping record tables. Fields that are external interpretations subject to independent revision (VEP consequence, gnomAD frequency, ClinVar significance) live in first-class annotation tables with temporal support.

FK Chain

AssayedVariant (1:N) → MappingRecord (M:N via mapping_record_alleles) → Allele

Annotation Architecture

AnnotationStatus is scoped strictly to QC and pipeline audit — it tracks job runs, success/failure/skipped status, and error detail. It is not the store for annotation data values.

Annotation data lives in first-class per-type tables, each with superseded_at for temporal queries:

VEPAnnotation (new)

  • allele_id (FK → alleles.id)
  • consequence, impact, affected_transcripts (structured columns)
  • source_version (VEP release)
  • created_at, superseded_at (nullable — null means current)

GnomADVariant (existing, updated)

  • FK migrated from mapped_variant_id → allele_id
  • Add superseded_at

ClinicalControl (existing, updated)

  • FK migrated from mapped_variant_id → allele_id
  • Add superseded_at

Temporal query pattern for any annotation table:

WHERE allele_id = :id
  AND created_at <= :point_in_time
  AND (superseded_at IS NULL OR superseded_at > :point_in_time)

VariantAnnotationStatus — scoped to QC

  • FK updated to reference allele_id (not variant_id)
  • annotation_metadata JSONB repurposed as job debug output only — not used for serving annotation data
  • Remove or deprecate fields that duplicated annotation data values

PublishedVariantsMV materialized view

The existing published_variants_materialized_view joins Variant → MappedVariant → ScoreSet and is queried by the statistics router. It must be updated to join AssayedVariant → MappingRecord → ScoreSet. The migration must explicitly drop, redefine, and refresh this view — it does not update automatically.

Acceptance Criteria

  • mapping_records table with Alembic migration (replaces mapped_variants); vrs_digest column indexed
  • Flat alleles table with vrs_digest unique constraint, level enum, transcript NOT NULL, post_mapped JSONB, and nullable HGVS columns
  • mapping_record_alleles M:N association table
  • New VEPAnnotation table with allele_id FK and superseded_at
  • GnomADVariant and ClinicalControl FK migrated to allele_id; superseded_at added to both
  • VariantAnnotationStatus FK updated to allele_id; scoped to QC only
  • PublishedVariantsMV updated to reflect new schema
  • SQLAlchemy models with appropriate relationships for all new tables
  • AnnotationStatusManager accepts any Allele row as a target
  • Basic model test coverage

Activity

  1. bencap commented on May 18, 2026

    @bencap
    CollaboratorAuthor

    Implementation note on current flag semantics:

    A DerivedAllele's current status is a function of its parent MappedVariant's current status — when a MappedVariant is superseded on remapping, its derived alleles are implicitly non-current. This should be enforced at the application layer: whenever MappedVariant.current is set to False, all linked DerivedAllele records should be updated to current=False in the same transaction. No separate lifecycle management is needed.

  2. bencap commented on May 18, 2026

    @bencap
    CollaboratorAuthor

    Implementation note on the annotation data model

    We should ask ourselves whether, given we now would store annotation information directly on two items, whether extracting all annotations to dedicated tables would be prudent.

  3. changed the title [-]feat: Add DerivedAllele data model and migration[/-] [+]feat: Add Allele data model and MappingRecord migration[/+] on May 19, 2026
  4. self-assigned this
    on May 29, 2026
  5. added theissue type on Aug 7, 2026
  6. added this to the R2026.3 milestone on Aug 12, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

app: backendTask implementation touches the backendapp: databaseTask implementation requires database changes

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions