Skip to content

Whether an all-caps word is normalized or kept as an acronym depends on what else is in the string #72

Description

@matt-edmondson

What's wrong

ToTitleCase normalizes all-caps words only when the entire string is all caps, because IsAllCaps is tested against the whole string rather than per word. Everything built on it — ToPascalCase, and ToCamelCase through it — inherits that, so the same token converts two different ways depending on its neighbours.

Measured at main (5c3fa37, v1.5.0), .NET SDK 10.0.401:

input ToTitleCase ToPascalCase ToMacroCase
MAX_SIZE Max _ Size MaxSize MAX_SIZE
set MAX_SIZE Set MAX _ SIZE SetMAXSIZE SET_MAX_SIZE
get MAX_SIZE Get MAX _ SIZE GetMAXSIZE GET_MAX_SIZE
URL Url Url URL
my URL handler My URL Handler MyURLHandler MY_URL_HANDLER
HTTP Http Http HTTP
parse HTTP header Parse HTTP Header ParseHTTPHeader PARSE_HTTP_HEADER

MAX_SIZE Pascal-cases to MaxSize on its own and to MAXSIZE the moment any lowercase word joins it. URL behaves the same way. The macro/snake path is consistent in both cases, so the divergence is specific to the title-case-derived converters.

Why it happens

ToTitleCase (CaseConverter.cs:163) reasons about the whole string:

// If the input is all caps, we want to convert it to lowercase before converting to title case,
// as TextInfo.ToTitleCase preserves words that are all caps assuming they are acronyms.
if (IsAllCaps(output))
{
    output = output.ToLowerInvariant();
}

return CultureInfo.InvariantCulture.TextInfo.ToTitleCase(output);

One lowercase character anywhere makes IsAllCaps false, so nothing is lowered and TextInfo.ToTitleCase keeps every all-caps word as an acronym.

This is a design question, not an obvious fix

Acronym preservation in a mixed string is deliberate and pinned by two existing tests:

  • ToTitleCaseShouldHandleMultipleSpaces — " the quick brown FOX " → "The Quick Brown FOX"
  • ToTitleCaseShouldConvertToTitleCase — "the quick Brown FOX" → "The Quick Brown FOX"

So the current rule is coherent: preserve all-caps words as acronyms, unless the whole string is all caps, in which case it is shouting rather than an acronym. It is just surprising at the boundary, and the boundary is where identifiers live.

Both consistent answers are defensible and they are not interchangeable:

  1. Normalize per word. set MAX_SIZE → SetMaxSize, matching MAX_SIZE → MaxSize and the macro/snake path. Better for identifiers, and closer to .NET naming guidance (HttpHeader, not HTTPHeader). Changes my URL handler → MyUrlHandler and would need the two FOX tests rewritten.
  2. Preserve per word. MAX_SIZE → MAXSIZE, keeping acronyms everywhere. Leaves the FOX tests alone but changes the single-word results, and loses the word boundary in MAX_SIZE.

A third option is to split them: leave ToTitleCase alone, since preserving FOX in prose is reasonable and tested, and give ToPascalCase/ToCamelCase per-word normalization, since an identifier converter has different needs from an English title converter. That fixes the reported case without touching the pinned behaviour, at the cost of the two no longer agreeing.

Either way it changes output for existing callers, which is why I have not picked one.

Where this was found

ktsu-dev/Coder#71 measured this while evaluating whether NameStyles.Spell could delegate to this library. It recorded set MAX_SIZE → SetMAXSIZE alongside MAX_SIZE → MaxSize and noted it looked like a defect here rather than something for that repo to work around. Filing it so it is on the record independently of what happens to that issue.

That same triage recorded a second finding — that a supplementary-plane letter was deleted (𝒳value → Value, set 𐐀abc → SetAbc). That one is already fixed as of v1.5.0 and is not part of this issue; re-measured on current main, 𝒳value → 𝒳value, set 𐐀abc → Set𐐀abc, and 𐐀abc.ToSnakeCase() → 𐐨abc with the correct Deseret lowercase. It is covered by ToPascalCaseShouldPreserveLettersOutsideTheBasicMultilingualPlane and its siblings.

Acceptance criteria

  • An all-caps word converts the same way regardless of what else is in the string, under whichever of the options above is chosen.
  • The chosen rule is stated in the XML docs on the affected methods, since it is not inferable from the method name.
  • Tests cover a single all-caps word and the same word beside a lowercase one, for each affected converter.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions