Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
54 commits
Select commit Hold shift + click to select a range
4aac9f7
chore: ralph prd.json for mapping-slice0; archive data-unification run
deanban Jun 30, 2026
9521b21
refactor: neutralize OMOP/OncoTree leaks in engine core (US-000)
deanban Jun 30, 2026
7ec5e77
feat: executable R29 OMOP-coupling guard (US-001)
deanban Jun 30, 2026
d1592e4
feat: slice-0 gold set + mapping eval harness (US-002)
deanban Jun 30, 2026
ac70d99
feat: vocabulary-agnostic resolver policy object (US-004)
deanban Jun 30, 2026
2f2df25
feat: SQLGlot vocab store query layer over vocabulary_omop (US-003)
deanban Jun 30, 2026
b292d28
feat: resolved value-mapping store (DuckDB, §1.5 frozen columns) (US-…
deanban Jun 30, 2026
6189af5
feat: deterministic OncoTree->SNOMED vocabulary resolver (§4 algorith…
deanban Jun 30, 2026
7e154be
feat: OMOP condition_occurrence TARGET via authored manifest (US-007)
deanban Jun 30, 2026
78d8616
feat: Slice-0 staging obligation + concrete PlanAssembler (US-008)
deanban Jun 30, 2026
0b5de4e
feat: deterministic VOCAB_LOOKUP assertion producer + graph edges (US…
deanban Jun 30, 2026
e83c76d
feat: transform compiler MappingPlan->SQLGlot->staging write (US-010)
deanban Jun 30, 2026
04549c2
feat: Gate D-lite staging QA (row count, null-rate, NO_MAP accounting…
deanban Jun 30, 2026
807781b
feat: mapping eval report with acceptance thresholds (US-012)
deanban Jun 30, 2026
9221f01
chore: mark US-012 passing in prd.json
deanban Jun 30, 2026
ef1551e
feat: sema fit DuckDB end-to-end spine smoke (US-012A)
deanban Jun 30, 2026
abe5def
feat: live Databricks staging backend for sema fit (US-013)
deanban Jun 30, 2026
481b5c7
feat: no-vocab manifest target proves the R29 boundary (US-014)
deanban Jun 30, 2026
4bca18c
fix: enforce real 3-field staging obligation in live fit + sema fit -…
deanban Jul 1, 2026
17c2cbe
docs: slice-0 next-session handoff + resolver-independent oracle sheet
deanban Jul 1, 2026
8949bb5
fix: scope resume cleanup to re-run tables only (bug-368)
deanban Jul 2, 2026
18a399e
feat: bucketed NO_MAP reasons distinguish absent/no-crosswalk/domain-…
deanban Jul 2, 2026
a4dd969
feat: --strict gates on contract conformance + label contradiction, n…
deanban Jul 2, 2026
cea5455
feat: standard_domain_governed flag makes the OMOP-not-SNOMED target …
deanban Jul 2, 2026
8ac0ee2
docs: tracked decision record — Slice-0 target is OMOP not SNOMED (bu…
deanban Jul 2, 2026
0400699
fix: scope --strict eval report to the current run, not the whole sto…
deanban Jul 2, 2026
e2bb1e1
fix: scope staging compile to the current run, not the whole store (b…
deanban Jul 2, 2026
fda15b3
docs: plan the full Slice-0 repoint (retire SNOMED legacy anchors)
deanban Jul 2, 2026
2c8007c
fix: scope assertion production + column_names to the current run/cat…
deanban Jul 2, 2026
adc7516
fix: normalize run_mappings to the store grain before staging (bug-386)
deanban Jul 2, 2026
256a019
docs: revise repoint plan per codex adversarial review
deanban Jul 2, 2026
96fe219
refactor: rename resolver policy ref to omop.oncotree_condition (reti…
deanban Jul 3, 2026
46116ea
refactor: repoint slice-0 manifest vocabulary to OMOP-Condition sentinel
deanban Jul 3, 2026
e151c85
docs: mark Slice-0 full repoint executed; log to wolf cerebrum/memory
deanban Jul 3, 2026
1194504
test: align slice-0 tests with OMOP-Condition target vocabulary
deanban Jul 4, 2026
a2d6f11
fix: detect git-lfs pointer in cbioportal fetch, fail instead of inge…
deanban Jul 22, 2026
9c58cce
docs: add slice-1 OMOP shape+identity plan and session handoff
deanban Jul 22, 2026
384a172
feat: generic identity registry with two-level get-or-create (S1-01)
deanban Jul 22, 2026
1c50cb2
feat: deterministic identity resolver with missing-key review disposi…
deanban Jul 22, 2026
145ac5d
feat: materialize omop.person from identity registry via projection p…
deanban Jul 22, 2026
19232dc
refactor: version omop condition manifest to 0.2.0 with nullable cond…
deanban Jul 22, 2026
157305b
feat: deterministic source-row surrogate PK for condition rows (S1-05)
deanban Jul 22, 2026
6d944b6
feat: FK-closed multi-table compiler for ordered person then conditio…
deanban Jul 22, 2026
ad51a15
feat: extend Gate-D-lite with FK-closure and required-field checks fo…
deanban Jul 22, 2026
4fd86a1
docs: record Stage A completion (S1-01..S1-07) and real-data gate; sc…
deanban Jul 22, 2026
e1f3cbf
feat: multi-warehouse FkBackend for the FK-closed compiler (S1-08)
deanban Jul 22, 2026
ff370ee
feat: bridge the DuckDB-canonical identity registry into Databricks (…
deanban Jul 22, 2026
d4ff727
feat: FK-closed OMOP-shape live-run harness and sema fit-omop-shape (…
deanban Jul 22, 2026
0b095c9
docs: record S1-08 live Databricks run complete (Stage A done)
deanban Jul 22, 2026
b58c231
feat: deterministic Stage B identity collapse with FK-closed rebuild
deanban Jul 22, 2026
a2578b6
refactor: move cBioPortal-to-OMOP glue into showcase package
deanban Jul 22, 2026
459986e
docs: record S1-10 collapse engine done and showcase relocation
deanban Jul 22, 2026
6b3f83d
fix: register showcase CLI commands from installed console script
deanban Jul 22, 2026
88da67e
docs: record slice-1 S1-10 live collapse complete
deanban Jul 22, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .last-branch
Original file line number Diff line number Diff line change
@@ -1 +1 @@
ralph/fix/data-unification
ralph/feat/mapping-slice0
206 changes: 206 additions & 0 deletions archive/2026-06-30-feat/mapping-slice0/prd.json

Large diffs are not rendered by default.

415 changes: 415 additions & 0 deletions archive/2026-06-30-feat/mapping-slice0/progress.txt

Large diffs are not rendered by default.

206 changes: 206 additions & 0 deletions archive/2026-06-30-fix-data-unification/prd.json

Large diffs are not rendered by default.

415 changes: 415 additions & 0 deletions archive/2026-06-30-fix-data-unification/progress.txt

Large diffs are not rendered by default.

386 changes: 386 additions & 0 deletions archive/2026-06-30-fix/data-unification/prd.json

Large diffs are not rendered by default.

3 changes: 3 additions & 0 deletions archive/2026-06-30-fix/data-unification/progress.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
# Ralph Progress Log
Started: Tue Jun 30 14:14:58 EDT 2026
---
412 changes: 296 additions & 116 deletions prd.json

Large diffs are not rendered by default.

1,313 changes: 938 additions & 375 deletions progress.txt

Large diffs are not rendered by default.

140 changes: 140 additions & 0 deletions scripts/check_engine_coupling.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,140 @@
"""R29 guard: no OMOP/OncoTree literal may leak into the engine core or the
generic mapping spine.

The 'any ontology' promise (R29) is enforced mechanically here, not in review.
A single policy constant (:data:`R29_POLICY`) owns the denylist, the scanned
core paths, and the allowlist of locations where domain literals legitimately
live (per-vocabulary policy modules, manifest adapters, authored manifests,
eval, tests). The same scan runs as a unit test (``tests/unit/
test_engine_coupling.py``) so ``uv run pytest`` fails on a leak, and as a CLI:

uv run python scripts/check_engine_coupling.py

Configured core paths that do not exist yet (e.g. ``src/sema/resolve/engine.py``,
``src/sema/compile/``) are skipped so the guard is green today and tightens
automatically as later Slice-0 stories create them.
"""
from __future__ import annotations

import sys
from dataclasses import dataclass
from pathlib import Path
from typing import Iterator, NamedTuple, Sequence


@dataclass(frozen=True)
class EngineCouplingPolicy:
"""Frozen R29 policy: what is forbidden, where it is checked, what is exempt.

``core_paths`` and ``allowlist`` are repo-relative POSIX paths. A path is a
file (scanned directly) or a directory (scanned recursively for ``*.py``).
A scanned file is exempt when its repo-relative path is, or lives under, an
allowlist entry.
"""

denylist: tuple[str, ...]
core_paths: tuple[str, ...]
allowlist: tuple[str, ...]


# Case-explicit: a casing variant must not bypass the US-000 migration.
R29_POLICY = EngineCouplingPolicy(
denylist=(
"OncoTree",
"oncotree",
"ONCOTREE",
"cBioPortal",
"cbioportal",
"cBio",
"cbio",
"Maps to",
"standard_concept",
"condition_occurrence",
"concept_id",
),
core_paths=(
"src/sema/engine",
"src/sema/resolve/engine.py",
"src/sema/resolve/assembler.py",
"src/sema/resolve/identity_registry.py",
"src/sema/resolve/identity_registry_utils.py",
"src/sema/resolve/identity_resolver.py",
"src/sema/resolve/identity_collapse.py",
"src/sema/resolve/identity_collapse_utils.py",
"src/sema/compile",
"src/sema/targets",
),
allowlist=(
"src/sema/resolve/policy.py",
"src/sema/resolve/policies",
"src/sema/targets/adapters",
"src/sema/eval",
"tests",
),
)

REPO_ROOT = Path(__file__).resolve().parents[1]


class Violation(NamedTuple):
path: str
lineno: int
literal: str


def _is_allowlisted(rel_path: str, allowlist: Sequence[str]) -> bool:
return any(
rel_path == entry or rel_path.startswith(f"{entry}/") for entry in allowlist
)


def _iter_scanned_files(
root: Path, policy: EngineCouplingPolicy
) -> Iterator[Path]:
"""Yield ``*.py`` files under the core paths, skipping allowlisted ones."""
for core in policy.core_paths:
target = root / core
if not target.exists():
continue
candidates = [target] if target.is_file() else sorted(target.rglob("*.py"))
for candidate in candidates:
if candidate.suffix != ".py":
continue
rel = candidate.relative_to(root).as_posix()
if not _is_allowlisted(rel, policy.allowlist):
yield candidate


def scan(root: Path, policy: EngineCouplingPolicy) -> list[Violation]:
"""Return every denylist hit across the policy's core paths under ``root``."""
violations: list[Violation] = []
for py_file in _iter_scanned_files(root, policy):
rel = py_file.relative_to(root).as_posix()
lines = py_file.read_text(encoding="utf-8").splitlines()
for lineno, line in enumerate(lines, start=1):
for literal in policy.denylist:
if literal in line:
violations.append(Violation(rel, lineno, literal))
return violations


def find_violations(
root: Path = REPO_ROOT, policy: EngineCouplingPolicy = R29_POLICY
) -> list[Violation]:
"""Scan the live repository tree with the default R29 policy."""
return scan(root, policy)


def main(argv: Sequence[str] | None = None) -> int:
violations = find_violations()
if not violations:
print("R29 engine-coupling guard: clean")
return 0
print("R29 engine-coupling guard: VIOLATIONS")
for v in violations:
print(f" {v.path}:{v.lineno}: {v.literal!r}")
return 1


if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
60 changes: 57 additions & 3 deletions showcase/cbioportal_to_omop/cbioportal_fetch_utils.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@
from __future__ import annotations

import json
import os
from pathlib import Path
from urllib.request import Request, urlopen

Expand All @@ -15,6 +16,12 @@
)
from sema.log import logger

# Git-LFS pointer files begin with this line. The datahub stores large tables in
# LFS; when GitHub's media CDN lacks an object it 404s and the raw host serves the
# pointer instead of the content. Accepting that pointer silently loads a 2-row
# garbage table, so we detect it and fail loudly (bug: silent LFS-pointer ingest).
_LFS_POINTER_MAGIC = b"version https://git-lfs.github.com/spec/v1"


def fetch_study_files(study_id: str, cache_dir: Path) -> Path:
study_cache = cache_dir / study_id
Expand Down Expand Up @@ -43,11 +50,58 @@ def fetch_lfs_or_raw(study_id: str, filename: str, target: Path) -> None:
media_url = MEDIA_URL_TEMPLATE.format(study_id=study_id, filename=filename)
try:
download_url_to(media_url, target)
return
if not _is_lfs_pointer(target):
return
except Exception as media_err:
logger.debug("media URL failed for {}: {}; falling back to raw", filename, media_err)
raw_url = RAW_URL_TEMPLATE.format(study_id=study_id, filename=filename)
download_url_to(raw_url, target)
raw_url = RAW_URL_TEMPLATE.format(study_id=study_id, filename=filename)
download_url_to(raw_url, target)
if not _is_lfs_pointer(target):
return
# The content we have is a git-lfs pointer, not the real file. Try an
# authenticated GitHub API fetch (LFS is resolved server-side when authorized),
# then fail loudly rather than ingest the pointer as data.
if _try_authenticated_lfs(study_id, filename, target) and not _is_lfs_pointer(target):
return
raise RuntimeError(
f"cBioPortal file '{filename}' resolved to a git-lfs POINTER, not content "
f"(GitHub LFS unavailable — media CDN 404, no/failed auth). Set GITHUB_TOKEN "
f"to authenticate, or retry when the datahub LFS object is available. "
f"Ingesting the pointer would create a 2-row garbage table."
)


def _is_lfs_pointer(target: Path) -> bool:
try:
with target.open("rb") as fh:
return fh.read(len(_LFS_POINTER_MAGIC)) == _LFS_POINTER_MAGIC
except OSError:
return False


def _try_authenticated_lfs(study_id: str, filename: str, target: Path) -> bool:
token = os.environ.get("GITHUB_TOKEN") or os.environ.get("GH_TOKEN")
if not token:
return False
api_url = (
f"https://api.github.com/repos/cBioPortal/datahub/contents/"
f"public/{study_id}/{filename}?ref=master"
)
req = Request(
api_url,
headers={
"Accept": "application/vnd.github.raw",
"Authorization": f"Bearer {token}",
},
)
try:
with urlopen(req) as resp:
data = resp.read()
except Exception as err: # noqa: BLE001 - report and fall through to fail-fast
logger.debug("authenticated LFS fetch failed for {}: {}", filename, err)
return False
target.write_bytes(data)
return True


def list_study_entries(study_id: str) -> list[dict[str, str]]:
Expand Down
Loading
Loading