Skip to content

Merge/consolidate entities whose computed id was affected by a data mistake (typo, inconsistent title/name) #182

Description

@DutchJaFO

Background

Every entity in this database takes its id from its own content. A change to the content a hash is
built from therefore produces a different entity rather than a changed one, and nothing re-points the
foreign keys that referred to the old id:

Entity Id derived from What moves when it changes
Quote quote text, source title QuoteGenre, QuoteTranslation, ConversationLine
Source title, type, and date since #374 every CharacterId derived from it, plus Quote.SourceId
Character source id, name, source type Quote.CharacterId, the Character to Source links
Person name Quote.PersonId
Series name every SeasonId derived from it, plus Source.SeriesId
Season series id, number Source.SeasonId
Universe name Series.UniverseId

Two of those cascade two levels: a Series rename moves every Season beneath it, and a Source retitling
moves every Character resolved through it.

The fields a rule may change are exactly the governing ones. A ConflictResolutionRule governs
source, date, type, character, author and quoteText; a SourceAliasRule corrects a title
and a date. So an id-moving change is the normal case during seeding, not an edge one.

Nothing in the codebase re-keys a row or follows a foreign key. Sql.Quotes.UpdateOnNewestWins
rewrites QuoteText and SourceId in place while matching on Id, so a row's identity and its
content diverge silently, and no code anywhere assigns a new Id to an existing row.

This is not fixable via the existing Modify/decidability pipeline (#162, #165, #171-#176): that
machinery corrects a field's value on an already-identified single row, matched either by explicit id
or natural key. It has no concept of "these two different ids are actually the same entity, please
consolidate them." A corrective re-import does not help either, since for a Quote the id incorporates
the source title, so correcting that field produces a third, differently-computed id rather than
fixing the original two.

What the mechanism has to decide, carried here because every sub-issue depends on the answers:

  1. How a consolidation or re-key is triggered: an admin action, a curated declarative file mirroring
    Minimal per-source conflict-resolution rule file + curated field-override preload #181's per-source rule files, a rule applied during seeding, or some combination.
  2. How dependents are re-pointed, and how conflicting field values between two rows are resolved.
    Check whether FieldMergeResolver's existing FieldResolutionChoice vocabulary is directly
    reusable or needs extending.
  3. Whether it needs its own audit and reversibility story, mirroring Admin: targeted soft-reset and restore by import batch #59's reverse-an-applied-batch
    precedent. Merging is destructive to one of the two rows, unlike a field-level Modify.
  4. Whether a partial re-key, where some links moved and others did not, is repairable by the same
    mechanism or needs its own.
  5. Whether Character: migrate to global identity via new Series/Universe schema (ADR + migration) #174's Character merge mechanics generalise here, or stay separate. The triggers differ:
    Character: migrate to global identity via new Series/Universe schema (ADR + migration) #174 merges by design, same Name, Type-anchored and Series-scoped; this merges by correction,
    recognising an existing content mistake. Do not assume the answer before Character: migrate to global identity via new Series/Universe schema (ADR + migration) #174 ships.

The triggering example remains the one found while scoping #180:
NikhilNamal17_popular-movie-quotes.json carries Lord Of The Ring - The Fellowships Of The Ring
alongside the correctly titled The Lord of the Rings: The Fellowship of the Ring, both dated 2001,
both type: movie, unmistakably the same film, computing to two permanently separate Source rows.

Sub-issues

# Scope Depends on
#421 A dedupe keeps the row with the lowest rowid rather than the one whose id a curated file declares, so a quote's id depends on whether the install was upgraded or fresh n/a
#440 The mechanism: consolidate two rows that are one entity, re-key a row whose governing content changed, and move every dependent with it n/a
#220 Two Quote rows for one quote, because StableId hashed the raw pre-correction source text, so no later correction can merge them #440, the mechanism: the ids were already fixed at conversion time, so there is nothing a per-quote correction can match on

Scope boundary

Definition of done

  • Every sub-issue listed above is closed
  • Findings summarised in a closing comment
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions