紛らわしい文字を安全なテキストに正規化する
紛らわしい文字の正規化とは、そっくり文字をそれがまねている通常の文字に置き換え、見た目が同じ二つの文字列が比較でも等しくなるようにすることです。ユーザー名の比較、ブロックリストの照合、識別子の重複排除の前に行うと効果的です。
具体例
- 入力
- Cоnfig file: аdmin2, Noёl, café
- 検出された文字体系
- ラテン文字 (Latn), キリル文字 (Cyrl)
- 検出された文字数
- 7
| 位置 | 文字 | コードポイント | 文字体系 | 置換後 | ルール |
|---|---|---|---|---|---|
| 0 | C | U+FF23 | ラテン文字 | C | NFKC 正規化 |
| 1 | о | U+043E | キリル文字 | o | 紛らわしい文字の対応表 |
| 7 | fi | U+FB01 | ラテン文字 | fi | NFKC 正規化 |
| 10 | : | U+FF1A | 共通文字 | : | 紛らわしい文字の対応表 |
| 12 | а | U+0430 | キリル文字 | a | 紛らわしい文字の対応表 |
| 17 | 2 | U+FF12 | 共通文字 | 2 | 紛らわしい文字の対応表 |
| 22 | ё | U+0451 | キリル文字 | ë | 紛らわしい文字の対応表 |
- 読み取り可能な Unicode を保持する
- Config file: admin2, Noël, café
- 厳密な ASCII フォールバック
- Config file: admin2, Noel, café
仕組み
- まず既知のそっくり文字を対応表に従って置き換えます。対応表にない文字は NFKC で正規化され、全角形、合字、その他の互換文字が標準の文字に変換されます。
- 「読み取り可能な Unicode を保持する」は、対応表で定義されている場合にアクセント付きの文字を残します(例:キリル文字の ё → ë)。「厳密な ASCII フォールバック」は代わりに通常の ASCII 文字を使います(ё → e)。
- 対応表になく NFKC でも変化しない文字はそのまま残るため、café はどちらのモードでもアクセントを保ちます。正規化したテキストは比較用のキーであり、セキュリティの判定ではありません。元のテキストも保存し、検出された文字を確認してください。
変換はベストエフォート型です。マッピングされた混同可能要素と NFKC フォールディングは決定的ですが、一部の正当な Unicode にはフラグが立てられません。
あなたのテキスト
貼り付けまたは入力 - 入力すると結果が更新されます (長い入力の場合は軽くデバウンスされます)。
30 文字をスキャン
7 件の疑わしい文字
厳密な ASCII フォールバック
オリジナル (疑わしい文字にマークが付いています)
元のビュー内の疑わしい文字には下線が引かれ、「疑わしい」というラベルが付けられます。ハイライトカラーに加えて。
suspicious character Csuspicious character оnfig suspicious character filesuspicious character : suspicious character аdminsuspicious character 2, Nosuspicious character ёl, café
クリーンアップ後の出力
文字の分析
| インデックス (0 ベース) | 元の文字 | 置換後 | コードポイント | 理由 |
|---|---|---|---|---|
| 0 | C | C | U+FF23 | NFKC normalization changed this character (compatibility or width folding). |
| 1 | о | o | U+043E | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 2 | n | n | U+006E | Not flagged as a confusable or compatibility character. |
| 3 | f | f | U+0066 | Not flagged as a confusable or compatibility character. |
| 4 | i | i | U+0069 | Not flagged as a confusable or compatibility character. |
| 5 | g | g | U+0067 | Not flagged as a confusable or compatibility character. |
| 6 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 7 | fi | fi | U+FB01 | NFKC normalization changed this character (compatibility or width folding). |
| 8 | l | l | U+006C | Not flagged as a confusable or compatibility character. |
| 9 | e | e | U+0065 | Not flagged as a confusable or compatibility character. |
| 10 | : | : | U+FF1A | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 11 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 12 | а | a | U+0430 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 13 | d | d | U+0064 | Not flagged as a confusable or compatibility character. |
| 14 | m | m | U+006D | Not flagged as a confusable or compatibility character. |
| 15 | i | i | U+0069 | Not flagged as a confusable or compatibility character. |
| 16 | n | n | U+006E | Not flagged as a confusable or compatibility character. |
| 17 | 2 | 2 | U+FF12 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 18 | , | , | U+002C | Not flagged as a confusable or compatibility character. |
| 19 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 20 | N | N | U+004E | Not flagged as a confusable or compatibility character. |
| 21 | o | o | U+006F | Not flagged as a confusable or compatibility character. |
| 22 | ё | e | U+0451 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 23 | l | l | U+006C | Not flagged as a confusable or compatibility character. |
| 24 | , | , | U+002C | Not flagged as a confusable or compatibility character. |
| 25 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 26 | c | c | U+0063 | Not flagged as a confusable or compatibility character. |
| 27 | a | a | U+0061 | Not flagged as a confusable or compatibility character. |
| 28 | f | f | U+0066 | Not flagged as a confusable or compatibility character. |
| 29 | é | é | U+00E9 | Not flagged as a confusable or compatibility character. |