Your coding LLM is bad at picking files from grep results, and it's expensive. Jev is better and costs about 1% as much.
When Claude Code or any other coding agent looks for the code behind a feature, it greps. Then the LLM reads every grep hit and decides which files matter. That's where things go wrong: the LLM misses many of the right files, and reading hundreds of hits burns tokens.
jev-filter takes that decision away from the LLM and gives it to Jev, TypeSafe's classifier. Jev reads each hit's full text and answers yes or no with a probability. The LLM then reads only the files Jev kept.
| Claude (Opus 5.5) decides | Jev decides | |
|---|---|---|
| Right files found, grep lines (benchmark) | 183 of 392 | 228 of 392 (+25%) |
| Right files found, outlines | 165 of 422 | 210 of 422 (+27%) |
| Precision at the 0.7 threshold, outlines | 43% | 45% |
| Cost to judge 4,177 candidate files | ≈ $16–21 | $0.17 |
| Saleor "How are taxes calculated?", 231 files, full text | ≈ $6.81 | $0.03 |
On the same grep candidates from six real codebases, Jev found 25–27% more of the right files than Opus 5.5 for about 1% of the cost. With the stricter 0.7 threshold it matched Opus's precision while still finding more.
> /jev-find How are taxes calculated?
grep /tax/ found 231 files. Jev kept 43 of 231 candidate files (threshold 0.7).
1.00 saleor/checkout/calculations.py
1.00 saleor/checkout/webhooks/calculate_taxes.py
1.00 saleor/tax/calculations/checkout.py
1.00 saleor/tax/calculations/order.py
0.99 saleor/plugins/avatax/plugin.py
...
Borderline (0.5 to 0.7), not kept:
0.64 saleor/checkout/models.py
...
Jev cost: $0.0305 for 231 calls (726,500 tokens)
If Claude had read every candidate in full to decide:
anthropic/claude-opus-5.5: ~$6.8077 (~1,688,069 input tokens, 223x Jev)
anthropic/claude-opus-5: ~$8.5096 (~1,688,069 input tokens, 279x Jev)
anthropic/claude-sonnet-5.5: ~$3.4039 (~1,688,069 input tokens, 112x Jev)
A real run on Saleor.
Every run shows what Jev actually cost, taken from OpenRouter's billing, next to what Claude would have charged to read the same files.
jev-filter works on top of grep. It replaces one step: the LLM deciding which grep hits matter. The LLM is weak at that step, and it's the most expensive one.
| Step | Without jev-filter | With jev-filter |
|---|---|---|
| Understand the question and choose what to grep for (pattern, folder, file glob) | Your AI agent | Your AI agent (Claude Code, or any coding copilot that can call MCP tools) |
| Run grep with those arguments | Your AI agent | The agent's own grep, or jev-filter's optional built-in grep (jev_grep_filter) |
| Decide which grep hits are relevant | The LLM reads every hit: misses right files, costs the most | Jev, one file at a time: finds 25–27% more right files, ~1% of the cost |
| Read the relevant files and answer | Your AI agent | Your AI agent, reading only what Jev kept |
jev-filter has no LLM of its own that plans searches. It never picks grep terms, rewrites the question, or searches beyond what it's given. Jev is a classifier: for each file it answers one yes/no question with a probability. If the agent's grep misses a file, jev-filter can't find it.
question ──► agent chooses the grep pattern ──► grep ──► each hit's full text ──► Jev: yes / no + probability ──► keep p ≥ 0.7 ──► agent reads only those
(Claude Code or another copilot) (agent's or built-in)
- Grep. The agent chooses the pattern. Either the agent runs grep itself and passes the hits to
jev_filter, or it passes the pattern tojev_grep_filter, which runsgit grep(orrg, or a pure-Python fallback) with exactly that pattern. Test files are skipped unless the agent asks for them. - Classify. Each file's full text, up to 24,000 tokens, goes to
typesafe/jev-1.13on OpenRouter with one question: "Is this file part of this feature, i.e. a file a developer would likely need to edit when this feature changes?" Eight requests run in parallel, with retries on rate limits. - Keep. Files at or above the threshold are kept. Files just below it are listed as borderline, so Claude can check them if the answer looks thin.
- Compare. The result puts Jev's billed cost next to an estimate of what Opus 5.5, Opus 5 and Sonnet 5.5 would have paid to read every candidate in full and decide.
Decisions are cached on disk by question, path and file contents, so asking the same question again costs nothing until a file changes.
You need Python 3.9 or newer (nothing to pip install) and an OpenRouter API key with a little credit. A typical question costs a few cents.
export OPENROUTER_API_KEY=sk-or-... # in the shell you start Claude Code from
claude plugin marketplace add ByteBell/jev-filter
claude plugin install jev-filter@bytebellAsk normally. The bundled skill tells Claude to use Jev whenever it's looking for the files behind a feature:
Where is tax calculated, and which files would I change to add a new tax provider?
Or use the command:
/jev-find How does certificate renewal work?
Claude reads the question and chooses the grep pattern (here something like renew|expir|reissu). jev-filter runs grep with that pattern and Jev judges the hits.
The two tools. Your agent calls these. In both, the agent supplies the search arguments:
| Tool | The agent passes | jev-filter does |
|---|---|---|
jev_filter |
The question and the paths from a grep the agent already ran |
Classifies those files with Jev |
jev_grep_filter |
The question and the grep pattern it chose (plus optional root, glob, include_tests) |
Runs grep with exactly that pattern, then classifies the hits with Jev |
Both also take threshold (default 0.7; use 0.5 when missing a file is costly) and compare_models (for example ["opus-5.5", "sonnet-5.5"]).
Or run it from the terminal, without an agent. You give the pattern yourself:
python3 plugin/server/jev_filter.py grep "How are taxes calculated?" "tax|vat" --root ~/src/saleor
python3 plugin/server/jev_filter.py files "How are taxes calculated?" saleor/tax/utils.py saleor/core/taxes.py
python3 plugin/server/jev_filter.py grep "How do leases work?" "lease" --glob "*.go" --threshold 0.5 --json| Threshold | Use it when | Benchmark, outlines |
|---|---|---|
| 0.5 | Missing a file is costly: refactors, impact analysis, security review | 50% recall, 41% precision |
| 0.7 (default) | You want a short, accurate list to read | 44% recall, 45% precision |
Opus 5.5 scored 39% recall and 43% precision on the same files.
We tested this on 31 questions across six public codebases: Saleor, Kubernetes, Istio, Argo CD, cert-manager and etcd, with 5,075 grep candidates in total. Answer keys come from git history, so no model's opinion counts as the truth. A file is a right answer if developers changed it in at least three commits about that feature since 2023.
| Depth | What the judge read | Opus 5.5 recall / precision | Jev ≥0.5 | Jev ≥0.7 |
|---|---|---|---|---|
| D1 (20 sets) | The lines grep matched | 47% / 39% | 58% / 29% | 53% / 32% |
| D2 (18 sets) | An outline of each file | 39% / 43% | 50% / 41% | 44% / 45% |
| D3 (16 sets) | The full file | 45% / 38% | 51% / 35% | 46% / 38% |
Both judges were scored on exactly the same files in every row. The method, all 31 questions, per-question results, every judge's picks and a script to recompute the tables are in benchmark/.
All optional, set as environment variables:
| Variable | Default | Meaning |
|---|---|---|
OPENROUTER_API_KEY |
(required) | Your OpenRouter key |
JEV_THRESHOLD |
0.7 |
Keep files at or above this probability |
JEV_MAX_TOKENS |
24000 |
Longest file text sent to Jev; longer files are cut and flagged |
JEV_MAX_FILES |
400 |
Most candidates judged in one call |
JEV_WORKERS |
8 |
Parallel requests to OpenRouter |
JEV_COMPARE_MODELS |
Opus 5.5, Opus 5, Sonnet 5.5 | Claude models to price against (OpenRouter ids, comma-separated) |
JEV_CLAUDE_PROMPT_TOKENS |
2000 |
Prompt tokens assumed for a Claude judge |
JEV_CLAUDE_PER_FILE_TOKENS |
2000 |
Tool-call overhead assumed per file a Claude judge reads |
JEV_CACHE_DIR |
~/.cache/jev-filter |
Where decisions and prices are cached |
JEV_MODEL |
typesafe/jev-1.13 |
The OpenRouter decisions model |
Model names accept short forms: opus, opus-5.5, opus-5, sonnet, sonnet-5.5, sonnet-5 and haiku.
The Claude figure is what a Claude model would pay to read every candidate in full and decide, without prompt caching:
- Input tokens: file characters ÷ 3.5, plus 2,000 prompt tokens, plus 2,000 tokens of tool-call overhead per file
- Output tokens: 12 per file, for listing the kept paths
- Prices: OpenRouter's live list prices, refreshed daily
It's an estimate, not a bill. Jev's cost is exactly what OpenRouter charged.
- Jev only judges what grep found, and the agent chooses the grep. Code that never uses the agent's grep terms, such as a subclass that never names its base class, won't be a candidate. If the result looks thin, the agent should grep again with other terms.
- Very long files are cut to their first 24,000 tokens and listed as truncated.
- Your code leaves your machine. Each candidate file's text is sent to OpenRouter and TypeSafe. Don't use this on code you can't share with them.
jev-filter/
├── .claude-plugin/marketplace.json the "bytebell" marketplace; points at ./plugin
├── plugin/ the only folder Claude Code installs
│ ├── .claude-plugin/plugin.json
│ ├── .mcp.json starts the MCP server
│ ├── server/jev_filter.py the whole plugin: MCP server + CLI, no dependencies
│ ├── skills/jev-filter/SKILL.md tells Claude when to use Jev
│ └── commands/jev-find.md /jev-find
├── benchmark/ method, data, picks and results (not installed)
└── tests/ offline tests (not installed)
Claude Code installs only the folder the marketplace points to, so benchmark/ and tests/ stay out of your plugin install.
python3 -m unittest discover -s tests # offline; fakes the Jev API, no key needed
python3 benchmark/compare.py # recompute the benchmark tables from the committed picksMIT