Skip to content

Delegate property-name fuzzy matching to ktsu.FuzzySearch #113

Description

@matt-edmondson

What's hand-rolled

Two independent, hand-rolled approximate string-matching implementations exist in this repo for the same underlying problem — matching a non-standard frontmatter property name against a set of known/existing names:

1. NameStandardizer.FindStandardPropertyMatch (used by StandardizePropertyNames, the NameStandardizer step of CombineFrontmatter):

https://github.com/ktsu-dev/Frontmatter/blob/main/Frontmatter/NameStandardizer.cs#L91-L152

if (normalizedKey.Contains(normalizedStandard) || normalizedStandard.Contains(normalizedKey))
{
    return standardProperty;
}

(NameStandardizer.cs#L105) — a plain substring check, returning the first standard property tested, not the best.

2. PropertyMerger.FindSemanticCanonicalName/CalculateWordMatchScore (used by the PropertyMerger step, a separate merge pass):

https://github.com/ktsu-dev/Frontmatter/blob/main/Frontmatter/PropertyMerger.cs#L251-L329

private static int CalculateWordMatchScore(string[] words1, string[] words2)
{
    int score = 0;
    foreach (string word1 in words1)
    {
        foreach (string word2 in words2)
        {
            if (word1 == word2) { score += 2; }
            else if (word1.Contains(word2) || word2.Contains(word1)) { score += 1; }
        }
    }
    return score;
}

(PropertyMerger.cs#L313-L329) — its own word-tokenized scoring loop, feeding an OrderByDescending(match => match.Score) pick of the best existing key (PropertyMerger.cs#L262-L273).

Both files also carry their own separate, near-identical NormalizePropertyName helper (NameStandardizer.cs#L114-L152 and PropertyMerger.cs#L280-L310).

The class doc comment on NameStandardizer says "using fuzzy matching" (NameStandardizer.cs#L18) and PropertyNameCache is documented as a "Cache for fuzzy matched property names" (NameStandardizer.cs#L15) — but neither implementation is a scored approximate match in the usual sense; they're substring containment and word-overlap counting.

Notably, this repo already declares ktsu.FuzzySearch as a package reference:

https://github.com/ktsu-dev/Frontmatter/blob/main/Frontmatter/Frontmatter.csproj#L9

but no .cs file in the repo calls into it — it's an unused dependency. The repo's own CLAUDE.md already flags this: "ktsu.FuzzySearch - Fuzzy matching (referenced but main matching uses PropertyMappings)".

What ktsu.FuzzySearch provides

ktsu.FuzzySearch.Fuzzy.Contains(ReadOnlySpan<char> subject, ReadOnlySpan<char> pattern, out int outScore):

https://github.com/ktsu-dev/FuzzySearch/blob/main/FuzzySearch/Fuzzy.cs#L84-L88

It reports whether every character of pattern occurs in subject in sequence, and produces a score that rewards consecutive matches, matches after _/space separators, and camelCase-boundary matches, while penalizing unmatched characters. Either hand-rolled matcher above could rank candidates by this score instead of by containment or raw word-overlap count.

Why it's worth it

Both current implementations have the same correctness gap: they don't rank candidates by similarity, so the result depends on iteration order (NameStandardizer, StandardOrder.PropertyNames's order) or on a coarse, unweighted count (PropertyMerger's +2/+1 scoring, which e.g. scores "date" vs ["updated", "at"] the same as "date" vs ["created", "date"] in some cases since it sums per-word-pair hits rather than accounting for word position or adjacency). Fuzzy.Contains's scoring — consecutive-match bonus, separator-boundary bonus, camelCase bonus, unmatched-character penalty — is exactly the kind of deliberate ranking both call sites currently approximate by hand, and it already ships in this repo's dependency graph, so this isn't a new package to add — just wiring up code that's already referenced.

Compatibility

  • Subject (ktsu.Frontmatter) targets: the ktsu.Sdk-provided default multi-target set (no <TargetFrameworks> override in Frontmatter.csproj; the project conditions packages on netstandard2.0 and netstandard2.1, so both are in the set it builds).
  • ktsu.FuzzySearch targets: the same ktsu.Sdk default set (its .csproj likewise has no override and conditions packages on netstandard2.0/netstandard2.1).
  • No gap in practice: ktsu.FuzzySearch is already a working PackageReference of Frontmatter.csproj, so it already compiles against every framework Frontmatter targets.
  • Dependency direction: ktsu.FuzzySearch has zero ktsu.* dependencies (its Directory.Packages.props lists only Polyfill, System.Memory, System.Threading.Tasks.Extensions), so there's no cycle.

Sketch

// NameStandardizer.FindStandardPropertyMatch — before
foreach (string standardProperty in standardProperties)
{
    string normalizedStandard = NormalizePropertyName(standardProperty);
    if (normalizedKey.Contains(normalizedStandard) || normalizedStandard.Contains(normalizedKey))
    {
        return standardProperty;
    }
}
return null;

// after
string? best = null;
int bestScore = int.MinValue;
foreach (string standardProperty in standardProperties)
{
    string normalizedStandard = NormalizePropertyName(standardProperty);
    if (Fuzzy.Contains(normalizedKey, normalizedStandard, out int score) && score > bestScore)
    {
        best = standardProperty;
        bestScore = score;
    }
}
return best;
// PropertyMerger.FindSemanticCanonicalName — before
(string Key, int Score)? bestMatch = existingKeys
    .Select(existingKey => (Key: existingKey, Score: CalculateWordMatchScore(keyWords, ...)))
    .Where(match => match.Score > 0)
    .OrderByDescending(match => match.Score)
    .Cast<(string Key, int Score)?>()
    .FirstOrDefault();

// after
(string Key, int Score)? bestMatch = existingKeys
    .Select(existingKey => (Key: existingKey, Matched: Fuzzy.Contains(NormalizePropertyName(key), NormalizePropertyName(existingKey), out int score), Score: score))
    .Where(match => match.Matched)
    .OrderByDescending(match => match.Score)
    .Select(match => (match.Key, match.Score))
    .Cast<(string Key, int Score)?>()
    .FirstOrDefault();

Caveats

  • Fuzzy.Contains requires every pattern character to appear in sequence in the subject; NameStandardizer's current check accepts a match in either direction (key-contains-standard or standard-contains-key), and PropertyMerger's scores word-by-word rather than whole-string. Adapting either call site needs a decision on which span is pattern vs. subject (and whether to check both directions and keep the better score), since it changes which non-standard keys match successfully.
  • Both call sites are internal (NameStandardizer, PropertyMerger, and their methods are all internal), so there's no public API to break — but the change in matching behavior could change which standard/existing property some non-standard frontmatter keys get mapped to, which is an observable behavior change for consumers of CombineFrontmatter.
  • Neither Frontmatter.Test/NameStandardizerTests.cs nor Frontmatter.Test/PropertyMergerTests.cs appears to pin the exact fallback-match behavior beyond what exact/known mappings cover (based on the test file list), so this is worth double-checking before changing either scoring.
  • Adopting Fuzzy.Contains in both places would also be a chance to fold the two near-duplicate NormalizePropertyName helpers into one, though that's a separate cleanup from the library swap itself.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions