Skip to content

scarecrow: make stale NSFW_CARD_IMAGE URL verdict cleanup reachable - #8

Merged
Pitchfork-and-Torch merged 1 commit into
mainfrom
cursor/nsfw-card-verdict-cleanup-114b
Sep 4, 2026
Merged

Pitchfork-and-Torch merged 1 commit into
mainfrom
cursor/nsfw-card-verdict-cleanup-114b

Conversation

@Pitchfork-and-Torch

Copy link
Copy Markdown
Owner

Problem

botmaker-rules/scarecrow/bot/NSFW_Card_Image_Media_To_URL_Verdict.bot (rule 7429) runs on media_update events for link-preview card images. When the image scores near-perfect NSFW it writes an NSFW_CARD_IMAGE verdict for the card URL with a 7-day TTL. Rule 7413 (NSFW_Card_Image_URL_to_Tweet_Verdict.bot) then labels every post linking that URL, and visibility-filtering puts those posts behind an interstitial and drops them from out-of-network recommendations (NsfwCardImageOonDropRule).

Rule 7429 also has a cleanup branch that is meant to delete the verdict when a later score for the same card comes back clean, so that a false positive does not sit on the URL for the full week. That branch could never run, for three independent reasons visible in the repo:

  1. Contradictory guards. The rule condition requires IsNearPerfectNsfw, which is precisionNewNsfw >= 0.999 (IsNearPerfectNsfwNewModelLogic.df). The cleanup branch requires !IsHighPrecisionNsfw, which is precisionNewNsfw < 0.95 (IsHighPrecisionNsfwNewModelLogic.df). Both cannot be true for the same event, so the else if was dead code.
  2. Inverted age check. :verdictCreationTime - (CurrentTimeMs()/1000) > ... computes creation minus now, which is never positive for a verdict written in the past.
  3. Wrong units on the threshold. 60 * 60 * 1000 is compared against a difference in seconds, so it means about 41 days, longer than the verdict's own 7-day TTL. Even with the sign fixed it would never fire before the verdict expired on its own.

Net effect: a card URL wrongly scored as near-perfect NSFW once stayed labelled for seven days, and every post linking it stayed interstitialed and out of non-follower recommendations, even if the image was re-scored as clean the next hour. The URLs_removed_as_NSFW_CARD_IMAGE_in_Rep_Store counter could never increment.

Change

Single file, NSFW_Card_Image_Media_To_URL_Verdict.bot:

  • condition: also admit card-image updates whose score is clean (Exists(precisionNewNsfw) && !IsHighRecallNsfw && !IsHighPrecisionNsfw). Events in the middle band (high-precision but not near-perfect, or high-recall) still do not fire, as before.
  • Cleanup branch: compute the age as now_seconds - creationTimestampInSec and compare against 60 * 60 (one hour, matching the 1-hour label TTL used by rule 7413).
  • Cleanup branch: require Exists(precisionNewNsfw), so a media update that carries no NSFW score cannot clear a verdict. Without this guard, widening the condition would have let score-less updates delete verdicts, which the original (unreachable) code would also have done.

The write path is unchanged: it fires on exactly the same inputs and writes exactly the same verdict, TTL, and source.

Before / after

Modelled from the derived-feature thresholds in derived-feature/*.df (hold = branch entered, verdict younger than 1 h, nothing deleted):

Scenario Before After
Near-perfect score, no verdict yet WRITE WRITE
Near-perfect score, verdict already set noop noop
Near-perfect score, URL already BAD noop noop
High-precision (0.97) score, any verdict state no-fire no-fire
Clean re-score (0.10), verdict 30 min old no-fire hold
Clean re-score (0.10), verdict 2 h old no-fire DELETE
Clean re-score (0.10), verdict 6 days old no-fire DELETE
Clean precision but high-recall flag set, verdict 2 h old no-fire no-fire
Clean re-score, no verdict on URL no-fire noop (one verdict lookup)
Media update with no NSFW score, verdict 2 h old no-fire no-fire

Sweeping precision 0.000–1.000, several recall values, and verdict ages from 0 to 100 weeks, the old cleanup branch produced zero DELETE outcomes.

Verification

  • Generated the parser from the repo's own grammar (botmaker/src/antlr/com/twitter/botmaker/antlr/BotMaker.g, ANTLR 3.5.2) and parsed the condition and action blocks of all 20 rules in botmaker-rules/scarecrow/bot/, before and after this change. All parse; the AST for the changed rule matches the intended shape (DISJUNCTION in the condition, DIFFERENCE(QUOTIENT(CurrentTimeMs, 1000), :verdictCreationTimeInSec) compared with PRODUCT(60, 60)).
  • CurrentTimeMs is confirmed as milliseconds in botmaker/src/java/com/twitter/botmaker/function/datetime/CurrentTimeMs.java; creationTimestampInSec is named in seconds by the rule itself.
  • The botmaker runtime and the GetUrlVerdictV2 / DeleteUrlVerdictIf functions are not in the repo, so the rule cannot be executed end to end here. The branch logic above was checked with a scratch model of the rule; nothing from that model is committed.

Trade-off

The rule now also fires for card-image updates that score clean, which adds one GetUrlVerdictV2 read per such event. That read is what the original design's cleanup branch required; it was never paid before only because the branch was unreachable.

Upstream

The file is byte-identical on xai-org/x-algorithm main (same blob), so this commit cherry-picks cleanly there.

Open in Web Open in Cursor 

Rule 7429 writes an NSFW_CARD_IMAGE verdict (7-day TTL) for a card URL
when its image scores near-perfect NSFW, and has a cleanup branch meant
to delete that verdict when a later score for the same card comes back
clean. The cleanup branch could never run:

- The rule condition required IsNearPerfectNsfw (precision >= 0.999),
  but the cleanup guard required !IsHighPrecisionNsfw (precision <
  0.95). Both cannot hold, so the branch was dead code.
- The age check computed creation - now, which is never positive.
- The threshold 60 * 60 * 1000 was compared against seconds, i.e. about
  41 days, longer than the verdict's own 7-day TTL.

Widen the condition to also admit clean card-image scores, compute the
age as now - creation in seconds, and use a one-hour threshold. Require
a present precision score on the cleanup path so a media update with no
NSFW score cannot clear a verdict. The write path is unchanged.

Co-authored-by: Jon Bailey <Pitchfork-and-Torch@users.noreply.github.com>
@Pitchfork-and-Torch
Pitchfork-and-Torch merged commit 866d4cf into main Sep 4, 2026
@Pitchfork-and-Torch
Pitchfork-and-Torch deleted the cursor/nsfw-card-verdict-cleanup-114b branch September 4, 2026 11:14
cursor Bot pushed a commit that referenced this pull request Sep 8, 2026
Rule 7429 writes an NSFW_CARD_IMAGE verdict (7-day TTL) for a card URL
when its image scores near-perfect NSFW, and has a cleanup branch meant
to delete that verdict when a later score for the same card comes back
clean. The cleanup branch could never run:

- The rule condition required IsNearPerfectNsfw (precision >= 0.999),
  but the cleanup guard required !IsHighPrecisionNsfw (precision <
  0.95). Both cannot hold, so the branch was dead code.
- The age check computed creation - now, which is never positive.
- The threshold 60 * 60 * 1000 was compared against seconds, i.e. about
  41 days, longer than the verdict's own 7-day TTL.

Widen the condition to also admit clean card-image scores, compute the
age as now - creation in seconds, and use a one-hour threshold. Require
a present precision score on the cleanup path so a media update with no
NSFW score cannot clear a verdict. The write path is unchanged.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Jon Bailey <Pitchfork-and-Torch@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants