What's wrong
ToUpperInvariant('ß') returns ß, and so do other lowercase letters with no single-character uppercase form, such as fi and ʼn. As a result ToMacroCase can produce words like STRAßE that still contain a lowercase (Ll) letter. When text like that is converted again:
IsWordBoundary (CaseConverter/CaseConverter.cs:118-145) treats ß as a lowercase letter. The acronym-tail rule (upper, upper, then lower) inserts a break before A, and the lower→upper rule inserts a break before the E that follows ß.
IsAllCaps (:281) returns false for STRAßE, so ToTitleCase doesn't normalize it as an all-caps word.
Repro (net10.0, at 0c1a1d0)
"straße".ToMacroCase(); // "STRAßE"
"straße".ToMacroCase().ToMacroCase(); // "STR_Aß_E" expected "STRAßE"
"STRAßE".ToSnakeCase(); // "str_aß_e" expected "straße"
"STRAßE".ToPascalCase(); // "StrAßE" expected "Straße"
"STRAßE".ToTitleCase(); // "Str Aß E" expected "Straße"
"MAßNAHME".ToSnakeCase(); // "m_aß_nahme" expected "maßnahme"
"MAX_GRÖßE".ToCamelCase(); // "maxGrÖßE" expected "maxGröße"
"maßnahmeLimit".ToMacroCase().ToSnakeCase(); // "m_aß_nahme_limit" expected "maßnahme_limit"
"file".ToMacroCase().ToMacroCase(); // "fi_LE" expected "fiLE"
Why it matters
German identifiers and config keys (MAX_GRÖßE, STRAßE) are ordinary input. Any pipeline that converts twice, or reads back a macro-case constant it generated itself, gets names split in the wrong places. The library's own output isn't a fixed point of its own conversion.
Suggested fix / acceptance criteria
- Treat a letter that uppercasing leaves unchanged as caseless for word-boundary purposes, for example
IsCasedLower(c) => char.IsLower(c) && char.ToUpperInvariant(c) != c. Use it in place of char.IsLower in the acronym-tail rule (:125) and in the lower→upper rule (:132).
IsAllCaps skips letters that have no uppercase mapping.
- Add regression tests for the examples above, including
x.ToMacroCase().ToMacroCase() == x.ToMacroCase().
This is distinct from #84, which is about the upper→lower round trip changing lowercase letters (final sigma, micro sign).
What's wrong
ToUpperInvariant('ß')returnsß, and so do other lowercase letters with no single-character uppercase form, such asfiandʼn. As a resultToMacroCasecan produce words likeSTRAßEthat still contain a lowercase (Ll) letter. When text like that is converted again:IsWordBoundary(CaseConverter/CaseConverter.cs:118-145) treatsßas a lowercase letter. The acronym-tail rule (upper, upper, then lower) inserts a break beforeA, and the lower→upper rule inserts a break before theEthat followsß.IsAllCaps(:281) returns false forSTRAßE, soToTitleCasedoesn't normalize it as an all-caps word.Repro (net10.0, at 0c1a1d0)
Why it matters
German identifiers and config keys (
MAX_GRÖßE,STRAßE) are ordinary input. Any pipeline that converts twice, or reads back a macro-case constant it generated itself, gets names split in the wrong places. The library's own output isn't a fixed point of its own conversion.Suggested fix / acceptance criteria
IsCasedLower(c) => char.IsLower(c) && char.ToUpperInvariant(c) != c. Use it in place ofchar.IsLowerin the acronym-tail rule (:125) and in the lower→upper rule (:132).IsAllCapsskips letters that have no uppercase mapping.x.ToMacroCase().ToMacroCase() == x.ToMacroCase().This is distinct from #84, which is about the upper→lower round trip changing lowercase letters (final sigma, micro sign).