What's wrong
Several static ConcurrentDictionary caches have no size limit and never evict anything:
Frontmatter/Frontmatter.cs:26: ProcessedFrontmatterCache, keyed on (Content, Options). It keeps the full text of every input document plus the full output string.
Frontmatter/YamlSerializer.cs:22: ParsedYamlCache. It keeps the text of every frontmatter block plus a parsed dictionary for each.
PropertyMerger.cs:17 and NameStandardizer.cs:16 hold one entry per property key. These are smaller, but they are also unbounded.
Failure scenario
The library gets used in long-running processes: a static-site generator in watch/serve mode, a web service that normalizes uploaded Markdown, or a bulk run over a large docs repository. Every distinct document, and every edited revision of it, stays reachable until the process exits.
A watcher re-processing a 50 KB page on every save adds about 100 KB (input plus output) per save and never frees it. A bulk run over 10,000 documents keeps all of them in memory after the run.
A cache hit is also only useful when the exact same content is processed again, which is uncommon in those workloads. The document-level cache therefore mostly costs memory.
Because the state is global, it has already caused wrong results: #130 is a stale-cache bug. There is also no way to reset it between tests.
Suggested fix / acceptance criteria
- Drop
ProcessedFrontmatterCache, or give it a size limit (a small LRU, or clear it once it passes N entries). Do the same for ParsedYamlCache.
- Optionally add a public
ClearCaches() so tests and hosts can reset state.
- Add a test showing that processing many distinct documents does not grow the cache past the limit.
What's wrong
Several static
ConcurrentDictionarycaches have no size limit and never evict anything:Frontmatter/Frontmatter.cs:26:ProcessedFrontmatterCache, keyed on(Content, Options). It keeps the full text of every input document plus the full output string.Frontmatter/YamlSerializer.cs:22:ParsedYamlCache. It keeps the text of every frontmatter block plus a parsed dictionary for each.PropertyMerger.cs:17andNameStandardizer.cs:16hold one entry per property key. These are smaller, but they are also unbounded.Failure scenario
The library gets used in long-running processes: a static-site generator in watch/serve mode, a web service that normalizes uploaded Markdown, or a bulk run over a large docs repository. Every distinct document, and every edited revision of it, stays reachable until the process exits.
A watcher re-processing a 50 KB page on every save adds about 100 KB (input plus output) per save and never frees it. A bulk run over 10,000 documents keeps all of them in memory after the run.
A cache hit is also only useful when the exact same content is processed again, which is uncommon in those workloads. The document-level cache therefore mostly costs memory.
Because the state is global, it has already caused wrong results: #130 is a stale-cache bug. There is also no way to reset it between tests.
Suggested fix / acceptance criteria
ProcessedFrontmatterCache, or give it a size limit (a small LRU, or clear it once it passes N entries). Do the same forParsedYamlCache.ClearCaches()so tests and hosts can reset state.