Skip to content

Latest commit

 

History

76 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

genderfluid-tiny

Tiny offline Python name-gender classifier. Predicts gender associations from names using ML. 3 MB model, CPU only, no API needed.

pip install genderfluid-tiny


PyPI version Python 3.10+ Tests License Model size


Documentation | PyPI | GitHub


What is genderfluid-tiny?

genderfluid-tiny is a lightweight Python library that predicts whether a name is statistically associated with feminine or masculine naming conventions. It uses a character n-gram classifier trained on 138,101 real names from U.S. Social Security Administration data (1880-2023), U.S. Census 2020, and French INSEE first-name statistics (1900-2024) - 911 million recorded births in total.

Unlike API-based gender detection services, genderfluid-tiny runs entirely offline. No data leaves your machine. No API key required. The entire model is 3 MB.

Property Value
Architecture Character n-gram + logistic regression
Model size 3 MB
Training data 138,101 names (SSA + Census + INSEE France)
Inference CPU only, no GPU needed
Internet Not required
License Polyform Noncommercial
Python 3.10+

Install

pip install genderfluid-tiny

That's it. The genderfluid command and Python API are available immediately.

Quick start

genderfluid predict "Emma"
Name: Emma

Girl-associated: 97.5%
Boy-associated:  0.0%
Uncertain:       2.5%

Classification: girl-associated
Confidence:     high

Python API

Simple one-liners

from genderfluid import classify_name, is_girl_name, is_boy_name, name_probability

classify_name("Emma")       # "girl-associated"
classify_name("James")      # "boy-associated"
classify_name("Alex")       # "uncertain"

is_girl_name("Emma")        # True
is_boy_name("James")        # True

name_probability("Emma")    # 0.9731

Full result dict

from genderfluid import predict_name, predict_names

result = predict_name("Isabella")
# {"name": "Isabella",
#  "girl_associated_probability": 0.8929,
#  "boy_associated_probability": 0.0486,
#  "uncertain_probability": 0.0585,
#  "classification": "girl-associated",
#  "confidence": "medium"}

results = predict_names(["Emma", "James", "Alex"])
for r in results:
    print(f"{r['name']}: {r['classification']}")

Model instance (for repeated use)

from genderfluid import GenderfluidModel

model = GenderfluidModel()  # loads once, cached
model.predict("Olivia")
model.predict_batch(["Emma", "James", "Alex", "Max", "Taylor"])

CLI

genderfluid predict "Olivia"                       # human-readable
genderfluid predict --json "Alex"                   # JSON output
genderfluid predict --compare "Emma" "James" "Alex" # comparison table
genderfluid predict --file names.txt                # batch from file
genderfluid interactive                             # interactive mode
genderfluid stats                                   # model info
genderfluid benchmark                               # performance test

How it works

Input name
  |
Unicode normalization + lowercase
  |
Character n-gram extraction (2-5 grams)
  |
Hashing trick (4096-dim feature vector)
  |
Logistic regression (3 classes)
  |
Sigmoid calibration
  |
Output: girl-associated / boy-associated / uncertain

The classifier extracts character-level patterns from names. Names ending in -a, -ia, -ine tend to be feminine. Names ending in -o, -us, -er tend to be masculine. The model learns these patterns from real data rather than hard-coding rules.

Accuracy

Tested on held-out test data (10,797 names):

Metric Value
Accuracy 79.6%
Macro F1 0.680
Girl-associated F1 0.887
Boy-associated F1 0.816
Uncertain F1 0.335

The model is trained on U.S./English naming conventions. Accuracy varies by cultural context.

Benchmark

Measured on a 2-core x86_64 Linux machine. Results vary by hardware.

Model size:       1.50 MB
Single name:      2.4 ms
Batch (10):       7.3 ms   (1,377 names/sec)
Batch (100):     19.2 ms   (5,214 names/sec)

Run genderfluid benchmark on your own hardware.

Use cases

  • Data pipelines: Classify gender associations in CSV/spreadsheet data
  • Name validation: Check if a name follows typical gender patterns
  • Research: Analyze naming trends across datasets
  • Privacy-sensitive applications: Process names without sending data to external APIs
  • Offline applications: Works without internet connectivity
  • Embedded systems: 3 MB model runs on low-resource devices

Training data

Built from real public data:

  1. U.S. Social Security Administration baby names (1880-2020): 100,364 unique names
  2. U.S. Census Bureau 2020 Census first names: 53,616 unique names

Combined: 138,101 unique names (911 million recorded births). Names with 85%+ statistical association are labeled girl-associated or boy-associated. Below that threshold: uncertain.

Training from source

python process_real_data.py   # download and process SSA + Census data
python prepare_data.py        # validate and split data
python train.py               # train and save model
python evaluate.py            # evaluate on validation/test splits

Sweep training on GitHub Actions (recommended for retraining): the Train Model workflow runs all 12 configurations in parallel on GitHub runners (one job per config, ~4 minutes total) and a finalize job picks the best config by validation F1, evaluates it on the test split, and commits the model. Trigger it under Actions > Train Model > Run workflow, or run the two steps locally:

python train_config.py 262144 2-6 20.0 lbfgs   # one configuration
python train_finalize.py                      # pick best, evaluate, save

Dataset format

JSONL, one entry per line:

{"name": "Emma", "label": "girl-associated"}
{"name": "James", "label": "boy-associated"}
{"name": "Alex", "label": "uncertain"}

Optional fields: weight, country, language, year.

Comparison with alternatives

Feature genderfluid-tiny gender-guesser chicksexer
Model size 3 MB 600 KB+ 10 MB+
License Polyform NC GPLv3 --
Last updated 2026 2016 --
Approach ML (n-gram + LR) Lookup table ML
Uncertain category Yes Partial No
pip install Yes Yes Yes
Offline Yes Yes Yes

Limitations

  • Works with full names: first, middle, and last
  • U.S./English-centric training data
  • Name associations vary by culture, language, and generation
  • The uncertain category exists for genuinely ambiguous names
  • Not suitable for high-stakes decisions

Privacy

All inference runs locally. Names are not transmitted to any external service. Logging of names is disabled by default.

FAQ

What data is it trained on?

138,101 unique names from U.S. SSA baby names (1880-2023), U.S. Census 2020, and French INSEE first names (1900-2024) - 911 million recorded births.

How accurate is it?

79.6% accuracy on held-out test data (macro F1 0.68). Girl-associated names: 89% F1. Boy-associated names: 82% F1. Uncertain/ambiguous names: 34% F1.

Does it work offline?

Yes. After pip install genderfluid-tiny, no internet connection is needed.

What Python versions are supported?

Python 3.10, 3.11, 3.12, 3.13.

Can I retrain the model?

Yes. See the Training from source section above. The training pipeline is included.

Does it work with non-English names?

The model is trained on U.S./English naming data. It may not work well for names from other cultural contexts. The preprocessing preserves Unicode characters, so names with accents and special characters are handled.

Repository structure

genderfluid-tiny/
├── genderfluid/          # Python package
│   ├── __init__.py       # Public API
│   ├── cli.py            # Command-line interface
│   ├── inference.py      # GenderfluidModel class
│   ├── classifier.py     # Logistic regression + calibration
│   ├── features.py       # Character n-gram extraction
│   ├── preprocessing.py  # Name normalization
│   └── model_io.py       # Binary save/load
├── data/                 # Training dataset
├── models/               # Trained model
├── native/               # C++ inference (optional)
├── tests/                # 29 tests
├── pyproject.toml        # Package config
└── README.md

License

Polyform Noncommercial License 1.0.0. Free for personal, educational, and noncommercial use. Commercial use requires a license. See COMMERCIAL_LICENSE.md.

About

Offline Python classifier that predicts gender associations from names. 49KB model, no API needed.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages