MD095 - Link destinations should not contain confusable non-ASCII characters¶
Aliases: link-confusables
This rule is opt-in. Enable it with extend-enable = ["MD095"] in the
[global] section of your config.
What this rule does¶
This rule detects selected Cyrillic, Greek, Latin, Armenian, and styled
characters that can look like ASCII characters in link destinations. It checks Markdown links, reference
definitions, HTML links, scheme-qualified bare URLs, and angle-bracket autolinks. Character
references such as е and е are decoded in Markdown destinations
and HTML href values. Diagnostics highlight the full reference in the source.
Bare URLs and angle-bracket autolinks are checked literally, matching their
Markdown semantics.
Why this matters¶
This is a defensive security measure against homograph, or IDN homograph,
attacks. An attacker can register a domain or create a path containing a
character from another writing system that resembles an ASCII character.
In rendered Markdown, a destination such as https://еxample.com can look very
similar to https://example.com even though the first URL contains a Cyrillic
е and may lead to a different host.
The same technique can hide a malicious path, query parameter, or fragment inside an otherwise familiar domain. This is especially risky in documentation, changelogs, issue templates, and generated pages where readers reasonably trust links and may not inspect the raw source or Unicode code points before clicking.
MD095 makes the suspicious character visible during review and CI. It does not claim that the link is malicious, resolve the destination, or replace the character automatically. A finding should prompt the author or reviewer to compare the destination with the intended URL, verify the domain through a trusted channel, and replace the link with the exact intended spelling when appropriate.
The rule only analyzes the URL text present in the Markdown document. It does not make network calls, fetch URLs, or otherwise check whether destinations are reachable or safe.
The rule reports Cyrillic, Greek, Latin, Armenian, Mathematical Alphanumeric Symbols, fullwidth, and selected URL-punctuation characters with printable ASCII targets in the raw confusables table of Unicode Technical Standard #39, using its version 18.0.0 confusables data. Normalization-only mappings and styled letters outside the Mathematical Alphanumeric Symbols block are outside this focused list. Image destinations are not checked. U+3002 (ideographic full stop) is additionally checked as a domain-separator look-alike. Mathematical and fullwidth blocks are filtered to actual ASCII mappings; unrelated characters in those blocks are not reported. Legitimate internationalized URLs can contain non-ASCII characters, and no static list can identify every visually confusable character for every font or language. URL reputation, certificate validation, and destination scanning remain separate security controls.
Defensive review guidance¶
When this rule reports a link:
- Inspect the raw destination and identify the reported Unicode code point.
- Compare the hostname and path with the intended address character by character.
- Confirm ownership or provenance using a trusted source, such as the project's established domain or repository configuration.
- Replace the destination or remove the link if the character was unexpected.
Do not silence a finding solely because the rendered link looks correct. If the non-ASCII character is intentional, document that decision in review and keep the rule disabled for that finding through the normal inline configuration mechanism.
Configuration¶
This rule has no options.
Examples¶
Correct¶
[Project](https://example.com/docs)
<https://example.com/docs>
<a href="https://example.com/docs">Project</a>
Incorrect¶
[Project](https://еxample.com/docs)
<https://example.com/dοcs>
<a href="https://example.com/docs">Project</a>
What this rule leaves alone¶
Non-ASCII characters that are not in its confusable set. Accented Latin characters, CJK characters, and other legitimate internationalized URL content are not reported just because they are non-ASCII.
Text outside link destinations. A confusable character in visible link text, headings, or prose is outside this rule's scope. The risk addressed here is a destination that appears to lead to a trusted location.
Percent escapes and Punycode. The rule does not decode URL percent escapes or IDNA/Punycode hostnames. It checks the characters in the parsed source destination and is not a complete URL-spoofing detector.
Code and comments. Links written in code blocks, code spans, or HTML comments are not rendered destinations and are not reported.
Ordinary non-link text. A URL-shaped string is only relevant when the parser or HTML scan treats it as a destination. Confusable characters in prose, headings, or other text are outside this rule's scope.
Automatic fixes¶
None, by design. A confusable may be intentional, and the correct replacement cannot be inferred safely from the source. Reviewers should verify the intended destination and replace it manually when necessary.
Related rules¶
- MD084 - Invisible and discouraged Unicode characters, for hidden and structurally discouraged Unicode characters
- MD034 - Bare URLs, for bare URLs that should use explicit link syntax