Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 15 additions & 2 deletions ai/src/aiutils/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,16 @@ When writing a new script in `ai/scripts/` or a new module in
3. **JSONL iteration:** Use `iter_jsonl()` from `aiutils.jsonl_utils` to read
newline-delimited JSON, not inline generators.
4. **Atomic Parquet writes:** Use `write_parquet_atomic()` from `aiutils.parquet_utils`
to safely write DataFrames with cleanup on failure.
to safely write DataFrames with cleanup on failure. As of ADR-0926 the helper
produces **schema v2** by default: zstd-3 compression, canonical column order
(`clip_id`, `frame_idx`, sorted features, labels, metadata), and pyarrow file
metadata that carries `vmafx_schema_version` + `vmafx_pipeline_hash`. Pass
`compression="snappy"` if a downstream consumer cannot read zstd, and pass
explicit `labels=` / `metadata=` to override the column-classification
heuristics. To detect what produced an input file, use
`read_parquet_with_schema(path)` which returns `(df, schema_version)` — v1
for legacy files written by raw `df.to_parquet(...)`, v2 for files written
by this helper. **Do not** call `df.to_parquet(...)` directly in new code.
5. **Run provenance:** Use `aiutils.run_manifest.write_run_manifest()` for
script-specific sidecars that need stable entrypoint, args, input, and
output metadata plus adapter-specific counts/config. Use
Expand All @@ -37,6 +46,10 @@ When writing a new script in `ai/scripts/` or a new module in
- `file_utils.py` — `sha256(path) -> str`
- `time_utils.py` — `now_iso_8601() -> str`
- `jsonl_utils.py` — `iter_jsonl(path) -> Iterator[tuple[int, dict]]`
- `parquet_utils.py` — `write_parquet_atomic(df, output, **kwargs) -> None`
- `parquet_utils.py` — `write_parquet_atomic(df, output, **kwargs) -> None`,
`read_parquet_with_schema(path) -> (df, int)`,
`detect_schema_version(path) -> int`,
`apply_standard_column_order(df, *, labels=None, metadata=None) -> DataFrame`
(ADR-0926; schema v2 is the on-disk default)
- `run_manifest.py` — deterministic `run_provenance` sidecar helpers
- `cli_helpers.py` — shared parser/raw-argv/batch-manifest argument helpers
17 changes: 14 additions & 3 deletions ai/src/aiutils/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,24 +21,35 @@

__all__ = [
"add_batch_manifest_arguments",
"apply_standard_column_order",
"build_run_manifest_payload",
"build_run_provenance",
"collect_cli_argv",
"describe_path",
"detect_schema_version",
"iter_jsonl",
"make_argument_parser",
"now_iso_8601",
"read_parquet_with_schema",
"sha256",
"write_manifest_json",
"write_parquet_atomic",
"write_run_manifest",
]


_LAZY_PARQUET_EXPORTS = {
"apply_standard_column_order",
"detect_schema_version",
"read_parquet_with_schema",
"write_parquet_atomic",
}


def __getattr__(name: str):
"""Import optional heavy helpers only when the caller asks for them."""
if name == "write_parquet_atomic":
from aiutils.parquet_utils import write_parquet_atomic
if name in _LAZY_PARQUET_EXPORTS:
import aiutils.parquet_utils as _pq

return write_parquet_atomic
return getattr(_pq, name)
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
Loading
Loading