Mojibake is a low-level Unicode library written in C11.
-
Updated
Sep 28, 2026 - C
Mojibake is a low-level Unicode library written in C11.
AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude Code skill + Python CLI. MIT.
The world's first font-by-font confusables dataset: which Unicode characters look alike, measured from the outlines of 322 fonts at the size people read them. CC BY, used in Mozilla's add-on name checks.
Research-only AI watermark & provenance robustness toolkit: local reverse proxy (OpenAI/Anthropic/Gemini) + CLI stripping C2PA/EXIF/XMP, zero-width & homoglyph Unicode, KGW text watermarks, DWT/Tree-Ring image stego, AudioSeal, PDF/DOCX/PPTX/XLSX metadata, and Trojan Source (CVE-2021-42574) code scanning.
Audit invisible AI text watermarks and test Claude watermark claims with reproducible, local-first analysis.
Reveal and remove hidden Unicode such as zero-width spaces, look-alike whitespace, and bidi controls. Runs locally.
SMUGGLR is a browser-based Unicode steganography and text obfuscation toolkit for experimenting with invisible characters, homoglyphs, whitespace ciphers, and Unicode-based payload hiding techniques — all fully client-side with zero network calls. 🔓
Detect and neutralize Unicode concealment codepoints in MCP tool metadata (covert-channel control). AGPL-3.0-or-later.
Phishing-domain detection from Certificate Transparency logs, with a measured false-positive rate. Deterministic: UTS#39 homoglyph skeletons, weighted edit distance, combosquat segmentation. No ML.
Unicode Homoglyph Detector — Offensive Unicode Security Analysis
Zero-dependency heuristic scanner for prompt-injection indicators in untrusted text, for the boundary of LLM agent pipelines.
Audit whether an AI tool approval view faithfully represents its raw payload.
Deterministic Unicode 17 text integrity tools for software and Agents.
Deterministic offline scanner for repositories and package artifacts, detecting prompt injection, refusal bait, Trojan Source, obfuscation, and supply-chain anti-analysis before AI code review.
SE·EMU is a browser-based social engineering simulation designed to help cybersecurity learners understand and practice the human side of offensive security.
'İGNORE'.lower() != 'ignore'. Turkish case-folding & Unicode confusables bypass naive prompt-injection filters 94.6% of the time — measured, with dataset + a one-line NFKC fix.
TrustNoChar is a zero-dependency browser-based lab that demonstrates how Unicode homoglyphs and typosquatting attacks exploit human visual perception. It transforms text into deceptive lookalike variants in real time to help researchers, red teams, and security learners study phishing, rendering quirks, and cognitive security risks. 🛡️👁️
To associate your repository with the unicode-security topic, visit your repo's landing page and select "manage topics."