Test Jev on your own evals with Claude Code
The step-by-step setup and the exact prompts I used to test Jev against my own scores, put Sonnet 5.5 head to head, and find out where to trust it.
"Jev, Jev, Jev. But what is it, and can it grade my evals?"
Jev is TypeSafe's new model, and its whole job is judging. You give it text and a question, and it gives back a score plus how confident it is. It does not write anything, which makes it fast and cheap. In my test, 60 grading calls ran in about a second for under a penny.
The setup that worked: Claude Code drives (it builds and runs the test), Jev grades, and your own scores are the answer key. You do not need to write code. You do need to be comfortable typing into Claude Code.
What you need
Claude Code installed and working.
A TypeSafe account and API key. Create the key at console.typesafe.ai.
15 to 20 AI outputs you have already graded by hand on your rubric, in a spreadsheet, Notion, or a doc. Mine were customer emails scored 1 to 5 on clarity, tone, relevance, correctness, conciseness, and brand voice.
Optional, for the head-to-head: an Anthropic API key from console.anthropic.com.
Store your API key safely
Save your TypeSafe key in a password manager (I use 1Password) instead of pasting it into Claude Code or your terminal. Anything you paste can end up in chat history or your shell history. Claude Code can read the key from the password manager each time the test runs, so it never shows on screen.
Install the TypeSafe skill in Claude Code
This gives Claude Code TypeSafe's own instructions for building with Jev. Paste this into Claude Code:
Before installing any plugin, ask Claude Code what it can access and where your data goes. This one is documentation only, with no hooks or background tools.
Kick off the test
Start a new folder for the project, open Claude Code there, and give it the full picture. This is the prompt to start with (fill in the brackets):
Claude Code will build a few small files: the rubric as Jev questions, a scoring script, and a report. Ask it to explain anything you do not follow.
Run it and read the report
Say "run it live". Look at three things for each rubric part: within a point of you, harsher or kinder, and confidence on the misses. My first run was close on tone and clarity, caught an add-on we cannot legally sell in Utah, and was much harsher than me on conciseness.
One rubric part way off? Ask Claude Code to rewrite those score levels using your own grading notes, then rerun. Each rerun costs pennies.
Put a frontier model head to head
Where Jev disagrees with you, check whether a bigger model would do better. Store your Anthropic key the same way as in step 1, then:
Sonnet agreed with Jev, not me: the same score on 15 of 20 and within a point on all 20.
When the models agree, fix your rubric
Read the reasons Sonnet gave on the examples where you scored highest. Then tell Claude Code, in your own words, what you were really rewarding. Mine:
With those rules, Jev matched me within a point on 14 of 15. Sonnet stayed at 10, at about 80 times the cost per score.
Then stop tuning. I added a fourth rule and Jev got worse, so I took it out. Rules written from the same few examples make the score look better without proving anything.
Test on fresh drafts, graded blind
The real test is outputs Jev has never seen. Ask Claude Code:
Use your current instructions as the reference. Mine first graded against last year's voice guide, and brand voice matched me on only 1 of 15. With my current skill, it landed within a point on 12 to 15 of 15 for every rubric part.
Use confidence to choose what you review
Every Jev score comes with a confidence number. Ask Claude Code: "Would reviewing only Jev's lowest-confidence scores have caught its misses?" On my fresh set, the lowest-confidence quarter caught 4 of its 6 big misses. That is a small sample, so check it on your own data before relying on it.
What it cost
Jev charges for input only, and every Jev run in my test cost about half a cent. All the Jev runs together came to a few cents. The Sonnet head-to-head was about 6 to 7 cents per run, for one rubric part.
Where Jev fits, and where it doesn't
Good fit: high-volume, repeatable judgments, like scoring every draft on tone or flagging anything that breaks a policy.
Keep for yourself or a bigger model: the subtle calls. Jev missed my "be more casual with someone I work with every week" note. Its own docs list math and dates as weak spots, so keep those in code.
If you use Claude Code's plugin evals: they cannot call an outside grader directly, so run Jev on their results as a step afterward.
Want help setting up evals for your own AI work?
We can build your first rubric and answer key together in one hour.
Get Started with AI ($150)