GRCh38 is used here as Source Layer 1 — the initial reference frame from which first-pass feature typing and uncertainty mapping will be anchored. It is not treated as a complete model of human genomic diversity or as a final biological authority.
// Normalized source model — GRCh38 / GRCh38.p14 { "source_id": "GRCh38_p14", "source_name": "Genome Reference Consortium Human Build 38, patch release 14", "source_type": "reference_assembly", "priority_layer": 1, "role": "baseline_reference_frame", "species": "Homo sapiens", "provider": "Genome Reference Consortium / NCBI", "assembly_accession_genbank": "GCA_000001405.29", "assembly_accession_refseq": "GCF_000001405.40", "approximate_haploid_size_bp": 3100000000, "chromosome_count": "22 autosomes + X + Y + MT + unlocalized/unplaced scaffolds", "known_gap_count_approx": 151, "centromere_status": "modeled (not fully sequenced in GRCh38)", "telomere_status": "partial", "status": "active_baseline_reference", "limitations": [ "Not a complete human diversity model — single mosaic haplotype", "~151 sequence gaps in primary assembly", "Centromere regions are modeled, not fully sequenced", "Telomeres only partially represented", "Highly repetitive and segmentally duplicated regions are challenging", "T2T-CHM13 fills ~8% of genome not fully resolved in GRCh38", "Human Pangenome Reference required for population-level diversity modeling" ], "notes": [ "Use as first-pass coordinate and comparison anchor", "Reference sequence does not equal full biological meaning", "Expand with T2T-CHM13 (Layer 2) and Human Pangenome (Layer 3)" ] }
| Property | Value | Implication for comparison work |
|---|---|---|
| Assembly type | Linear reference, mosaic | All coordinates are GRCh38-relative; liftover required for other assemblies |
| Known gaps | ~151 | Gap-flanking features must carry reference_limitation uncertainty flag |
| Centromeres | Modeled, not fully sequenced | Centromeric features are structurally uncertain by construction |
| Telomeres | Partial | Sub-telomeric features have reduced coverage confidence |
| Diversity coverage | Single mosaic haplotype | Allele frequency and population diversity require Human Pangenome layer |
| Functional annotation | Not intrinsic to assembly | ENCODE/GTEx required for regulatory and expression feature layers |
These feature classes define what the first-pass comparison engine recognizes. All classes include an explicit status field that can hold KNOWN, PARTIAL, PLACEHOLDER, or UNKNOWN_UNRESOLVED. Placeholder classes are structural commitments — they define the type even before data is available to fill them.
The uncertainty model separates two independent failure modes that must never be merged. This is the core categorical distinction of the whole engine — from file 002_BIOLOGICAL_VS_MEASUREMENT_UNRESOLVEDNESS.md.
| Uncertainty Class | Category | Effect on comparison | May raise similarity score? |
|---|---|---|---|
| reference_limitation | A — Biological | Increases uncertainty; feature typed as partial or unknown | No |
| biological_unresolvedness | A — Biological | Increases uncertainty; preserves possibility space | No |
| functional_annotation_absent | A — Biological | Feature cannot be scored functionally; type = placeholder | No |
| structural_variant_unresolved | A — Biological | SV class held as placeholder; consequence unknown | No |
| measurement_unresolvedness | B — Measurement | Uncertainty increase + may force comparison invalidation | No |
| low_coverage | B — Measurement | Coverage confidence flag set; comparison weight reduced | No |
| provenance_unresolvedness | B — Measurement | Provenance flag set; may escalate to hard override | No |
| contamination_risk | B — Measurement | Contamination flag set; comparison must be suspended pending review | No |
| chain_of_custody_failure | B — Measurement | Hard override candidate; comparison result must be flagged as potentially invalid | No |
| ambiguous_state | A+B — Mixed | Category of ambiguity must be declared; cannot be left untagged | No |
| not_measured | A+B — Mixed | Feature is absent from data; must not be treated as reference-match | No |
This draft schema defines how a single genomic feature is represented in the first-pass system. All uncertainty-carrying fields are typed separately and tagged with their category (A = biological, B = measurement/provenance). No field may remain implicitly unset — all flags default to explicit values.
// Draft feature object schema — GRCh38 first-pass { // --- Identity --- "feature_id": "string // globally unique ID for this feature record", "source_id": "string // e.g. GRCh38_p14", "reference_frame": "string // e.g. GRCh38 | GRCh38_p14 | T2T-CHM13 (when added)", "created_at": "ISO8601 timestamp", // --- Coordinate / Region --- "chromosome": "string // e.g. chr1, chrX, chrMT, chrUn_...", "position_start_bp": "int // 1-based GRCh38 coordinate", "position_end_bp": "int // inclusive end", "strand": "enum // +, -, . (unstranded)", "region_type": "enum // primary_assembly | alt_locus | unlocalized | unplaced | gap", // --- Feature Type --- "feature_type": "enum // snp | indel | sv_placeholder | cnv_placeholder | repeat_region_placeholder | regulatory_placeholder | cds_region | noncoding_rna_placeholder | epigenetic_placeholder | gap_region | unknown_unresolved_feature | intergenic_region", "feature_subtype": "string // optional detail e.g. missense | synonymous | frameshift | intron | 3UTR", "status": "enum // known | partial | placeholder | ambiguous | unknown_unresolved", // --- Biological Uncertainty (Category A) --- "biological_unknown_class": "enum // none | unknown_function | regulatory_role_unclear | sv_consequence_undetermined | expression_context_unknown | epigenetic_effect_unmapped | gene_env_interaction_unknown | population_frequency_unknown", "functional_annotation_present": "bool // false = functional role not annotated in this source", "reference_completeness_flag": "enum // complete | gap_adjacent | centromere_region | telomere_region | repeat_dense | segdup_region", // --- Measurement / Provenance Uncertainty (Category B) --- "measurement_unknown_class": "enum // none | low_coverage | degraded_sample | mixed_sample | contamination_suspected | platform_limitation | extraction_problem | method_limitation", "provenance_class": "enum // verified | unverified | metadata_incomplete | chain_of_custody_broken | origin_uncertain | not_documented", "contamination_flag": "bool // true = contamination suspected or confirmed", "coverage_confidence": "enum // high | medium | low | not_assessed | gap (cannot be assessed)", // --- Reference Confidence --- "reference_confidence": "enum // high | medium | low | gap | modeled_only", "reference_limitation_note": "string // free text if reference_confidence is not high", // --- Scoring Gates --- "continuity_weight_allowed": "bool // false = this feature may NOT contribute to positive similarity score", "uncertainty_weight": "float // 0.0 (no uncertainty) to 1.0 (fully uncertain); contributes to uncertainty accumulation only", "comparison_valid": "enum // valid | caution | suspended | invalidated", "hard_override_triggered": "bool // true = Category B failure severe enough to invalidate comparison", // --- Notes --- "notes": "string[] // free-text annotations", "linked_source_records": "string[] // accessions or IDs from upstream databases" }
continuity_weight_allowed defaults to false for all features where feature_type ends in _placeholder or unknown_unresolved_feature, where biological_unknown_class ≠ none, or where measurement_unknown_class ≠ none. A feature may not contribute to positive similarity scoring unless all uncertainty fields are clean and reference_confidence is high.
These rules govern what the first-pass engine is permitted to do with each feature class and uncertainty class. This is not a final scoring engine — it is the boundary logic that any future scoring engine must operate within.
| Rule | Trigger condition | Effect |
|---|---|---|
| R-01 | feature_type = unknown_unresolved_feature | continuity_weight_allowed := false unconditionally |
| R-02 | contamination_flag = true | comparison_valid := suspended; hard_override_triggered := true |
| R-03 | provenance_class = chain_of_custody_broken | hard_override_triggered := true; comparison_valid := invalidated |
| R-04 | reference_confidence = gap | continuity_weight_allowed := false; reference_limitation_note required |
| R-05 | biological_unknown_class ≠ none | uncertainty_weight += contribution; continuity_weight_allowed := false |
| R-06 | measurement_unknown_class ≠ none | uncertainty_weight += contribution; comparison_valid downgrade to caution minimum |
| R-07 | feature_type ends in _placeholder | status := placeholder; continuity_weight_allowed := false |
| R-08 | coverage_confidence = not_assessed | Cannot default to high; must remain not_assessed; caution applied |
| R-09 | Any new feature from T2T or Pangenome | Must be re-evaluated independently; GRCh38 classification does not transfer automatically |
| R-10 (HIR) | Any identity, continuity, or family-relation inference attempted | Blocked. May not proceed beyond declared evidence limits. Respect invariant enforced. |
| HIR Gate | First-Pass Implementation |
|---|---|
| Honesty (H) | Every feature record must carry explicit status, uncertainty class, reference confidence, and coverage confidence. No field may be left implicitly "clean." Unknown must be labeled unknown. |
| Integrity (I) | Category A (biological) and Category B (measurement/provenance) uncertainty classes must remain in separate fields at all times. No merging, implicit promotion, or collapsing of these into a single uncertainty score. |
| Respect (R) | No comparison output may be used to assert identity, family relation, person-level continuity, or any claim beyond the declared evidence limits of verified features with high reference confidence and clean provenance. |
Each source layer is recommended in order of dependency. No layer is added until its predecessor has been fully normalized, its uncertainty classes have been documented, and the schema has been validated with a first-pass feature set. Later layers do not replace earlier ones — they are additive.
Baseline coordinate system. Chromosome-level organization. First-pass feature typing. Establishes schema and uncertainty framework. Must be fully normalized before Layer 2 is added.
Adds ~8% of human sequence not fully resolved in GRCh38, including complete centromere sequences, fully resolved telomeres, and previously gap-filled regions. Used to compare GRCh38 feature calls against a gap-free reference. New features discovered in T2T require their own uncertainty evaluation — GRCh38 gap-region flags do not automatically resolve.
Moves beyond single linear reference to graph-based pangenome representing variation across 47+ diverse individuals. Required for any work involving allele frequency, population specificity, or variation beyond what a single mosaic haplotype captures. Adds substantial complexity — schema must be extended to handle graph coordinates before this layer is introduced.
Provides regulatory element annotations: promoters, enhancers, CTCF binding sites, open chromatin regions, transcription factor binding sites. Required to fill regulatory_region_placeholder and epigenetic_placeholder classes. Cell-type specificity must be tracked — ENCODE annotations are not genome-wide constants.
Provides tissue-specific expression quantitative trait loci (eQTLs) and gene expression data. Required to resolve expression_context_unknown flags for features with unclear expression consequences. Tissue specificity means a single feature may have different uncertainty categories across tissues — schema must accommodate this.
ClinVar provides variant-level clinical significance assertions. ClinGen provides gene-disease validity evidence. These layers should be added only after the comparison schema is stable, because clinical assertions carry their own provenance, review status, and conflicting interpretation complexities that must be handled separately from reference uncertainty.
| Layer | Source | Prerequisite | Primary addition |
|---|---|---|---|
| 1 | GRCh38.p14 | None | Baseline coordinates, SNP/indel typing, gap catalog |
| 2 | T2T-CHM13v2.0 | Layer 1 normalized | Gap resolution, centromere/telomere features |
| 3 | HPRC pangenome | Layer 2 + graph schema extension | Population diversity, allele frequency, pangenome graph coordinates |
| 4 | ENCODE | Layer 1 stable | Regulatory element annotation, chromatin state |
| 5 | GTEx | Layer 4 | Tissue-specific expression, eQTL resolution |
| 6 | ClinVar / ClinGen | Layers 1–3 stable | Clinical significance, gene-disease validity |
If this GRCh38 first-pass architecture plan is accepted, the next OSF-ready packet should contain the following files. It should not contain raw genome data — it should contain structured, schema-validated metadata that can be reviewed and verified independently.
What should come next, in order of value:
| Priority | Next action | Reason |
|---|---|---|
| 1 — Now | Formalize the typed schema (Task 4) as a validated JSON Schema file with required fields, enum values, and default rules | Everything downstream depends on the schema being stable and machine-verifiable. A JSON Schema file lets any future code validate records against it automatically. |
| 2 — Now | Encode the HIR rule set (Task 5 rules R-01 through R-10) as a machine-readable rule table (JSON or YAML) | Rules expressed only in prose will drift in interpretation. A structured rule table can be executed, tested, and versioned independently of narrative documentation. |
| 3 — Now | Produce the GRCh38 gap region catalog as a structured TSV (chromosome, start, end, reference_confidence = gap, notes) | This is a bounded, publicly documentable data set. It defines which coordinate ranges must automatically carry reference_limitation flags. It is the first "real" data slice that can be ingested and tested without raw genome files. |
| 4 — Next packet | Write 10–20 sample feature records in JSONL format conforming to the v0.2 schema | Sample records serve simultaneously as schema tests, documentation, and design validation. They reveal schema gaps before a full pipeline is built. |
| 5 — Next packet | Prepare a T2T-CHM13 source model document following the same structure as this GRCh38 first pass | Layer 2 normalization should mirror Layer 1 exactly. Preparing it as a parallel document ensures the schema is portable before any comparative analysis is attempted. |
| 6 — Later | Design the liftover / coordinate-alignment spec between GRCh38 and T2T-CHM13 | Comparison across reference frames requires explicit coordinate mapping. This is non-trivial for centromeres and previously-gap regions. It should be designed before data from both layers exists. |
| 7 — Later | Extend the schema with a pangenome_graph_coordinate field before adding HPRC Layer 3 | Graph-based coordinates are structurally different from linear coordinates. Adding HPRC data to a linear-coordinate schema without this extension will produce ambiguous records. |