ARCHITECTURE PLAN — NOT VALIDATED SYSTEM. Indexed for review/source-return only. Not empirical validation, clinical/legal/scientific authority, ontology proof, production readiness, or domain certification. v0.9.2 rule: branch documents read outside the registry must be re-anchored to their claim ladder and source registry.
ARCHITECTURE / REVIEW SOURCE — NOT A VALIDATED SYSTEM. This branch is indexed for inspection only. It is not empirical validation, clinical/legal/scientific certification, production deployment, ontology proof, consciousness proof, or domain authority. SHA-256/provenance integrity does not make the claim true.
Primordial DNA Layer · Architecture Planning Document · v0.1
GRCh38 First-Pass
Genomic Uncertainty Engine
HIR-Governed · Source Layer 1 Architecture Plan · GRCh38 / GRCh38.p14
Created and Developed by Collin D. Weber April 30, 2026 Source Layer 1 of N First-Pass Planning Only
Non-Negotiable Invariant (from 002_BIOLOGICAL_VS_MEASUREMENT_UNRESOLVEDNESS.md):
Unknown biology may preserve possibility space. Untrusted data may invalidate the comparison.
Neither may be converted into hidden positive evidence.
Biological unresolvedness ≠ measurement/provenance unresolvedness. These categories must remain separate throughout all schemas, scoring, and downstream layers.
Task 1
Normalized GRCh38 Source Model

GRCh38 is used here as Source Layer 1 — the initial reference frame from which first-pass feature typing and uncertainty mapping will be anchored. It is not treated as a complete model of human genomic diversity or as a final biological authority.

// Normalized source model — GRCh38 / GRCh38.p14
{
  "source_id": "GRCh38_p14",
  "source_name": "Genome Reference Consortium Human Build 38, patch release 14",
  "source_type": "reference_assembly",
  "priority_layer": 1,
  "role": "baseline_reference_frame",
  "species": "Homo sapiens",
  "provider": "Genome Reference Consortium / NCBI",
  "assembly_accession_genbank": "GCA_000001405.29",
  "assembly_accession_refseq": "GCF_000001405.40",
  "approximate_haploid_size_bp": 3100000000,
  "chromosome_count": "22 autosomes + X + Y + MT + unlocalized/unplaced scaffolds",
  "known_gap_count_approx": 151,
  "centromere_status": "modeled (not fully sequenced in GRCh38)",
  "telomere_status": "partial",
  "status": "active_baseline_reference",
  "limitations": [
    "Not a complete human diversity model — single mosaic haplotype",
    "~151 sequence gaps in primary assembly",
    "Centromere regions are modeled, not fully sequenced",
    "Telomeres only partially represented",
    "Highly repetitive and segmentally duplicated regions are challenging",
    "T2T-CHM13 fills ~8% of genome not fully resolved in GRCh38",
    "Human Pangenome Reference required for population-level diversity modeling"
  ],
  "notes": [
    "Use as first-pass coordinate and comparison anchor",
    "Reference sequence does not equal full biological meaning",
    "Expand with T2T-CHM13 (Layer 2) and Human Pangenome (Layer 3)"
  ]
}
PropertyValueImplication for comparison work
Assembly typeLinear reference, mosaicAll coordinates are GRCh38-relative; liftover required for other assemblies
Known gaps~151Gap-flanking features must carry reference_limitation uncertainty flag
CentromeresModeled, not fully sequencedCentromeric features are structurally uncertain by construction
TelomeresPartialSub-telomeric features have reduced coverage confidence
Diversity coverageSingle mosaic haplotypeAllele frequency and population diversity require Human Pangenome layer
Functional annotationNot intrinsic to assemblyENCODE/GTEx required for regulatory and expression feature layers

Task 2
First-Pass Feature-Class Model

These feature classes define what the first-pass comparison engine recognizes. All classes include an explicit status field that can hold KNOWN, PARTIAL, PLACEHOLDER, or UNKNOWN_UNRESOLVED. Placeholder classes are structural commitments — they define the type even before data is available to fill them.

snp
Single nucleotide polymorphism. Single-base substitution relative to reference. Most common variant class. Well-supported in GRCh38 coordinate space.
GRCh38 native
indel
Small insertion or deletion (<50 bp by convention). Well-supported in GRCh38. Alignment and left-normalization conventions must be documented per-source.
GRCh38 native
structural_variant_placeholder
Large variant (≥50 bp): deletions, duplications, inversions, translocations. GRCh38 has known SV limitations. Full SV atlas requires additional sources.
placeholder
copy_number_variant_placeholder
Copy number gain or loss. Distinct from SV in scoring semantics. Requires dedicated calling pipeline and additional source layers for population-level interpretation.
placeholder
repeat_region_placeholder
Tandem repeats, STRs, SINEs, LINEs, segmental duplications. These regions are particularly challenging in GRCh38; many require T2T-CHM13 for full resolution.
placeholder
regulatory_region_placeholder
Promoters, enhancers, silencers, insulators, CTCF binding sites. Not intrinsic to reference assembly; requires ENCODE or similar annotation layer.
annotation-dependent
cds_region
Protein-coding sequence. Supported by GRCh38 gene annotation (GENCODE/RefSeq). Functional consequence depends on additional annotation and variant classification layers.
annotation-dependent
noncoding_rna_placeholder
lncRNA, miRNA, snoRNA, piRNA. Functional annotation is incomplete in GRCh38. GTEx and dedicated ncRNA databases needed for expression-level evidence.
annotation-dependent
epigenetic_placeholder
Methylation, histone modification, chromatin accessibility. Not resolvable from GRCh38 reference sequence alone. Requires ENCODE or bisulfite-seq data layers.
future layer
gap_region
Known sequence gap in GRCh38 primary assembly (~151 gaps). Any feature overlapping a gap region must carry reference_limitation flag and cannot support high-confidence comparison.
structural limit
unknown_unresolved_feature
Catch-all for features that cannot be confidently typed. Must not be promoted to any positive evidence class. Preserves position for future resolution.
unknown
intergenic_region
Regions between annotated genes. Functional role largely undefined in this source layer. May carry regulatory elements discoverable in later annotation layers.
underannotated

Task 3
First-Pass Uncertainty Model

The uncertainty model separates two independent failure modes that must never be merged. This is the core categorical distinction of the whole engine — from file 002_BIOLOGICAL_VS_MEASUREMENT_UNRESOLVEDNESS.md.

Category A — Biological Unresolvedness
Feature may be real, measurable, reproducible — but its biological meaning is not yet understood.

Does NOT automatically invalidate comparison.
DOES increase uncertainty weight.
May NOT be converted into positive similarity score.

Examples:
· known sequence, unknown function
· regulatory role unclear
· SV consequence undetermined
· tissue-specific expression unresolved
· epigenetic effect unmapped
· gene-environment interaction unknown
Category B — Measurement / Provenance Unresolvedness
Evidence itself may be unreliable, degraded, contaminated, or not auditable.

MAY invalidate the comparison directly (hard override).
DOES increase uncertainty weight.
May NOT be converted into positive similarity score.

Examples:
· low sequencing coverage
· degraded or mixed sample
· contamination
· platform limitations
· metadata gaps
· chain-of-custody failure
· uncertain sample origin
Uncertainty ClassCategoryEffect on comparisonMay raise similarity score?
reference_limitationA — BiologicalIncreases uncertainty; feature typed as partial or unknownNo
biological_unresolvednessA — BiologicalIncreases uncertainty; preserves possibility spaceNo
functional_annotation_absentA — BiologicalFeature cannot be scored functionally; type = placeholderNo
structural_variant_unresolvedA — BiologicalSV class held as placeholder; consequence unknownNo
measurement_unresolvednessB — MeasurementUncertainty increase + may force comparison invalidationNo
low_coverageB — MeasurementCoverage confidence flag set; comparison weight reducedNo
provenance_unresolvednessB — MeasurementProvenance flag set; may escalate to hard overrideNo
contamination_riskB — MeasurementContamination flag set; comparison must be suspended pending reviewNo
chain_of_custody_failureB — MeasurementHard override candidate; comparison result must be flagged as potentially invalidNo
ambiguous_stateA+B — MixedCategory of ambiguity must be declared; cannot be left untaggedNo
not_measuredA+B — MixedFeature is absent from data; must not be treated as reference-matchNo

Task 4
Proposed Typed Schema

This draft schema defines how a single genomic feature is represented in the first-pass system. All uncertainty-carrying fields are typed separately and tagged with their category (A = biological, B = measurement/provenance). No field may remain implicitly unset — all flags default to explicit values.

// Draft feature object schema — GRCh38 first-pass
{
  // --- Identity ---
  "feature_id":               "string  // globally unique ID for this feature record",
  "source_id":                "string  // e.g. GRCh38_p14",
  "reference_frame":          "string  // e.g. GRCh38 | GRCh38_p14 | T2T-CHM13 (when added)",
  "created_at":               "ISO8601 timestamp",

  // --- Coordinate / Region ---
  "chromosome":               "string  // e.g. chr1, chrX, chrMT, chrUn_...",
  "position_start_bp":        "int     // 1-based GRCh38 coordinate",
  "position_end_bp":          "int     // inclusive end",
  "strand":                   "enum    // +, -, . (unstranded)",
  "region_type":              "enum    // primary_assembly | alt_locus | unlocalized | unplaced | gap",

  // --- Feature Type ---
  "feature_type":             "enum    // snp | indel | sv_placeholder | cnv_placeholder | repeat_region_placeholder | regulatory_placeholder | cds_region | noncoding_rna_placeholder | epigenetic_placeholder | gap_region | unknown_unresolved_feature | intergenic_region",
  "feature_subtype":          "string  // optional detail e.g. missense | synonymous | frameshift | intron | 3UTR",
  "status":                   "enum    // known | partial | placeholder | ambiguous | unknown_unresolved",

  // --- Biological Uncertainty (Category A) ---
  "biological_unknown_class":  "enum    // none | unknown_function | regulatory_role_unclear | sv_consequence_undetermined | expression_context_unknown | epigenetic_effect_unmapped | gene_env_interaction_unknown | population_frequency_unknown",
  "functional_annotation_present": "bool // false = functional role not annotated in this source",
  "reference_completeness_flag": "enum   // complete | gap_adjacent | centromere_region | telomere_region | repeat_dense | segdup_region",

  // --- Measurement / Provenance Uncertainty (Category B) ---
  "measurement_unknown_class": "enum    // none | low_coverage | degraded_sample | mixed_sample | contamination_suspected | platform_limitation | extraction_problem | method_limitation",
  "provenance_class":          "enum    // verified | unverified | metadata_incomplete | chain_of_custody_broken | origin_uncertain | not_documented",
  "contamination_flag":        "bool   // true = contamination suspected or confirmed",
  "coverage_confidence":       "enum    // high | medium | low | not_assessed | gap (cannot be assessed)",

  // --- Reference Confidence ---
  "reference_confidence":      "enum    // high | medium | low | gap | modeled_only",
  "reference_limitation_note": "string  // free text if reference_confidence is not high",

  // --- Scoring Gates ---
  "continuity_weight_allowed": "bool   // false = this feature may NOT contribute to positive similarity score",
  "uncertainty_weight":        "float  // 0.0 (no uncertainty) to 1.0 (fully uncertain); contributes to uncertainty accumulation only",
  "comparison_valid":          "enum    // valid | caution | suspended | invalidated",
  "hard_override_triggered":   "bool   // true = Category B failure severe enough to invalidate comparison",

  // --- Notes ---
  "notes":                    "string[]  // free-text annotations",
  "linked_source_records":    "string[]  // accessions or IDs from upstream databases"
}
Critical schema constraint: The field continuity_weight_allowed defaults to false for all features where feature_type ends in _placeholder or unknown_unresolved_feature, where biological_unknown_class ≠ none, or where measurement_unknown_class ≠ none. A feature may not contribute to positive similarity scoring unless all uncertainty fields are clean and reference_confidence is high.

Task 5
HIR-Safe First-Pass Scoring Rule Set

These rules govern what the first-pass engine is permitted to do with each feature class and uncertainty class. This is not a final scoring engine — it is the boundary logic that any future scoring engine must operate within.

Permitted — May contribute to known comparison
· feature_type ∈ {snp, indel, cds_region} AND status = known
· biological_unknown_class = none AND measurement_unknown_class = none
· provenance_class = verified AND contamination_flag = false
· coverage_confidence ∈ {high, medium} AND reference_confidence ∈ {high}
· continuity_weight_allowed = true (set explicitly, not inferred)
· comparison_valid = valid
Caution — Contributes to uncertainty accumulation only; must not raise similarity score
· feature_type ends in _placeholder OR status = partial
· biological_unknown_class ≠ none (unknown function, regulatory role unclear, etc.)
· reference_completeness_flag ∈ {gap_adjacent, centromere_region, telomere_region, repeat_dense, segdup_region}
· reference_confidence ∈ {medium, low, modeled_only}
· coverage_confidence = low AND measurement_unknown_class = none (low coverage alone = caution)
· functional_annotation_present = false
· comparison_valid = caution
Hard stop — Must not contribute to comparison; may force comparison_valid = suspended or invalidated
· feature_type = gap_region OR feature_type = unknown_unresolved_feature
· measurement_unknown_class ∈ {contamination_suspected, mixed_sample, chain_of_custody context} → escalate
· contamination_flag = true → comparison_valid := suspended; hard_override_triggered := true
· provenance_class ∈ {chain_of_custody_broken, not_documented} → hard_override_triggered := true
· hard_override_triggered = true → comparison_valid := invalidated
· Any attempt to convert either biological_unknown_class OR measurement_unknown_class into positive similarity → rejected
RuleTrigger conditionEffect
R-01feature_type = unknown_unresolved_featurecontinuity_weight_allowed := false unconditionally
R-02contamination_flag = truecomparison_valid := suspended; hard_override_triggered := true
R-03provenance_class = chain_of_custody_brokenhard_override_triggered := true; comparison_valid := invalidated
R-04reference_confidence = gapcontinuity_weight_allowed := false; reference_limitation_note required
R-05biological_unknown_class ≠ noneuncertainty_weight += contribution; continuity_weight_allowed := false
R-06measurement_unknown_class ≠ noneuncertainty_weight += contribution; comparison_valid downgrade to caution minimum
R-07feature_type ends in _placeholderstatus := placeholder; continuity_weight_allowed := false
R-08coverage_confidence = not_assessedCannot default to high; must remain not_assessed; caution applied
R-09Any new feature from T2T or PangenomeMust be re-evaluated independently; GRCh38 classification does not transfer automatically
R-10 (HIR)Any identity, continuity, or family-relation inference attemptedBlocked. May not proceed beyond declared evidence limits. Respect invariant enforced.
HIR GateFirst-Pass Implementation
Honesty (H)Every feature record must carry explicit status, uncertainty class, reference confidence, and coverage confidence. No field may be left implicitly "clean." Unknown must be labeled unknown.
Integrity (I)Category A (biological) and Category B (measurement/provenance) uncertainty classes must remain in separate fields at all times. No merging, implicit promotion, or collapsing of these into a single uncertainty score.
Respect (R)No comparison output may be used to assert identity, family relation, person-level continuity, or any claim beyond the declared evidence limits of verified features with high reference confidence and clean provenance.

Task 6
Staged Ingest Recommendation

Each source layer is recommended in order of dependency. No layer is added until its predecessor has been fully normalized, its uncertainty classes have been documented, and the schema has been validated with a first-pass feature set. Later layers do not replace earlier ones — they are additive.

1
GRCh38 / GRCh38.p14
ACTIVE — Source Layer 1 (this document)

Baseline coordinate system. Chromosome-level organization. First-pass feature typing. Establishes schema and uncertainty framework. Must be fully normalized before Layer 2 is added.

2
T2T-CHM13v2.0 (Telomere-to-Telomere, CHM13)
NEXT — Complete gapless reference comparison

Adds ~8% of human sequence not fully resolved in GRCh38, including complete centromere sequences, fully resolved telomeres, and previously gap-filled regions. Used to compare GRCh38 feature calls against a gap-free reference. New features discovered in T2T require their own uncertainty evaluation — GRCh38 gap-region flags do not automatically resolve.

3
Human Pangenome Reference (HPRC)
NEXT AFTER T2T — Population diversity representation

Moves beyond single linear reference to graph-based pangenome representing variation across 47+ diverse individuals. Required for any work involving allele frequency, population specificity, or variation beyond what a single mosaic haplotype captures. Adds substantial complexity — schema must be extended to handle graph coordinates before this layer is introduced.

4
ENCODE (Encyclopedia of DNA Elements)
LATER — Functional annotation layer

Provides regulatory element annotations: promoters, enhancers, CTCF binding sites, open chromatin regions, transcription factor binding sites. Required to fill regulatory_region_placeholder and epigenetic_placeholder classes. Cell-type specificity must be tracked — ENCODE annotations are not genome-wide constants.

5
GTEx (Genotype-Tissue Expression)
LATER — Expression context layer

Provides tissue-specific expression quantitative trait loci (eQTLs) and gene expression data. Required to resolve expression_context_unknown flags for features with unclear expression consequences. Tissue specificity means a single feature may have different uncertainty categories across tissues — schema must accommodate this.

6
ClinVar / ClinGen
LATER — Clinical variant interpretation layer

ClinVar provides variant-level clinical significance assertions. ClinGen provides gene-disease validity evidence. These layers should be added only after the comparison schema is stable, because clinical assertions carry their own provenance, review status, and conflicting interpretation complexities that must be handled separately from reference uncertainty.

LayerSourcePrerequisitePrimary addition
1GRCh38.p14NoneBaseline coordinates, SNP/indel typing, gap catalog
2T2T-CHM13v2.0Layer 1 normalizedGap resolution, centromere/telomere features
3HPRC pangenomeLayer 2 + graph schema extensionPopulation diversity, allele frequency, pangenome graph coordinates
4ENCODELayer 1 stableRegulatory element annotation, chromatin state
5GTExLayer 4Tissue-specific expression, eQTL resolution
6ClinVar / ClinGenLayers 1–3 stableClinical significance, gene-disease validity

Task 7
Bounded Output Recommendation — Next OSF-Ready Packet

If this GRCh38 first-pass architecture plan is accepted, the next OSF-ready packet should contain the following files. It should not contain raw genome data — it should contain structured, schema-validated metadata that can be reviewed and verified independently.

Primordial_DNA_GRCh38_Source_Layer_1_v0.2_Collin_D_Weber/ │ ├── 000_READ_ME_FIRST.md ← entry point, scope statement, HIR boundary ├── 001_SOURCE_MODEL.json ← normalized GRCh38.p14 source object (from Task 1) ├── 002_FEATURE_CLASS_REGISTRY.json ← all 12 feature classes with status and rules ├── 003_UNCERTAINTY_TAXONOMY.json ← Category A and B uncertainty classes with rules ├── 004_FEATURE_SCHEMA_v0.2.json ← full typed schema (from Task 4), with field validation rules ├── 005_HIR_RULE_SET_v0.1.json ← R-01 through R-10 formalized as machine-readable rules ├── 006_SAMPLE_FEATURES.jsonl ← 10–20 example feature records conforming to schema ├── 007_GAP_REGION_CATALOG.tsv ← GRCh38 known gaps with chromosome, coordinates, reference_confidence=gap ├── 008_STAGED_INGEST_PLAN.md ← layers 1–6 with prerequisites and additions ├── 009_BIOLOGICAL_VS_MEASUREMENT_BOUNDARY.md ← v0.2 of 002_BIOLOGICAL_VS_MEASUREMENT_UNRESOLVEDNESS.md ├── 010_MANIFEST.md ← file inventory └── 011_SHA256_CHECKSUMS.txt ← integrity checksums
What 006_SAMPLE_FEATURES.jsonl should contain: 10–20 hand-crafted feature records covering each major feature_type, each major uncertainty class, at least one gap_region feature, at least one contamination_flag=true case, and at least one fully-clean comparison-valid=valid case. These serve as schema validation tests and as design documentation simultaneously.

Task 8
Concise Implementation Note

What should come next, in order of value:

PriorityNext actionReason
1 — Now Formalize the typed schema (Task 4) as a validated JSON Schema file with required fields, enum values, and default rules Everything downstream depends on the schema being stable and machine-verifiable. A JSON Schema file lets any future code validate records against it automatically.
2 — Now Encode the HIR rule set (Task 5 rules R-01 through R-10) as a machine-readable rule table (JSON or YAML) Rules expressed only in prose will drift in interpretation. A structured rule table can be executed, tested, and versioned independently of narrative documentation.
3 — Now Produce the GRCh38 gap region catalog as a structured TSV (chromosome, start, end, reference_confidence = gap, notes) This is a bounded, publicly documentable data set. It defines which coordinate ranges must automatically carry reference_limitation flags. It is the first "real" data slice that can be ingested and tested without raw genome files.
4 — Next packet Write 10–20 sample feature records in JSONL format conforming to the v0.2 schema Sample records serve simultaneously as schema tests, documentation, and design validation. They reveal schema gaps before a full pipeline is built.
5 — Next packet Prepare a T2T-CHM13 source model document following the same structure as this GRCh38 first pass Layer 2 normalization should mirror Layer 1 exactly. Preparing it as a parallel document ensures the schema is portable before any comparative analysis is attempted.
6 — Later Design the liftover / coordinate-alignment spec between GRCh38 and T2T-CHM13 Comparison across reference frames requires explicit coordinate mapping. This is non-trivial for centromeres and previously-gap regions. It should be designed before data from both layers exists.
7 — Later Extend the schema with a pangenome_graph_coordinate field before adding HPRC Layer 3 Graph-based coordinates are structurally different from linear coordinates. Adding HPRC data to a linear-coordinate schema without this extension will produce ambiguous records.
What NOT to do next
Do not begin ingesting raw .fasta, .vcf, or .bam files until the schema is validated with sample records and the HIR rule set is machine-readable.
Do not add T2T or Pangenome layers until GRCh38 Layer 1 has passed schema validation testing.
Do not attempt any similarity scoring, comparison output, or identity inference until all three of: feature_type clean, uncertainty classes clean, and provenance verified.