DeSpain Consulting

Test Jev on your own evals with Claude Code

The step-by-step setup and the exact prompts I used to test Jev against my own scores, put Sonnet 5.5 head to head, and find out where to trust it.

"Jev, Jev, Jev. But what is it, and can it grade my evals?"

Jev is TypeSafe's new model, and its whole job is judging. You give it text and a question, and it gives back a score plus how confident it is. It does not write anything, which makes it fast and cheap. In my test, 60 grading calls ran in about a second for under a penny.

The setup that worked: Claude Code drives (it builds and runs the test), Jev grades, and your own scores are the answer key. You do not need to write code. You do need to be comfortable typing into Claude Code.

What you need

Claude Code installed and working.

A TypeSafe account and API key. Create the key at console.typesafe.ai.

15 to 20 AI outputs you have already graded by hand on your rubric, in a spreadsheet, Notion, or a doc. Mine were customer emails scored 1 to 5 on clarity, tone, relevance, correctness, conciseness, and brand voice.

Optional, for the head-to-head: an Anthropic API key from console.anthropic.com.

1

Store your API key safely

Save your TypeSafe key in a password manager (I use 1Password) instead of pasting it into Claude Code or your terminal. Anything you paste can end up in chat history or your shell history. Claude Code can read the key from the password manager each time the test runs, so it never shows on screen.

2

Install the TypeSafe skill in Claude Code

This gives Claude Code TypeSafe's own instructions for building with Jev. Paste this into Claude Code:

Paste into Claude CodeInstall the TypeSafe skill: run claude plugin marketplace add typesafe-ai/skills, then claude plugin install typesafe@typesafe-ai.

Before installing any plugin, ask Claude Code what it can access and where your data goes. This one is documentation only, with no hooks or background tools.

3

Kick off the test

Start a new folder for the project, open Claude Code there, and give it the full picture. This is the prompt to start with (fill in the brackets):

Paste into Claude CodeI want to see if Jev, TypeSafe's judging model, can grade my evals the way I do. My graded examples are in [where they live]. Each one has the AI output, the request it answered, and my 1 to 5 scores on [your rubric parts]. My brand voice guide and business facts are in [where]. Set up a calibration test: turn each rubric part into a Jev Score question, send Jev only the output, the request, and my guides, never my scores. Read my TypeSafe key from [your password manager] at run time and never print it. Run the requests in parallel and record time and cost. Then show me, for each rubric part, how often Jev is within a point of me, whether it runs harsher or kinder, and its confidence when it misses. Do a dry run with fake answers first, then wait for my go-ahead before the live run.

Claude Code will build a few small files: the rubric as Jev questions, a scoring script, and a report. Ask it to explain anything you do not follow.

4

Run it and read the report

Say "run it live". Look at three things for each rubric part: within a point of you, harsher or kinder, and confidence on the misses. My first run was close on tone and clarity, caught an add-on we cannot legally sell in Utah, and was much harsher than me on conciseness.

One rubric part way off? Ask Claude Code to rewrite those score levels using your own grading notes, then rerun. Each rerun costs pennies.

5

Put a frontier model head to head

Where Jev disagrees with you, check whether a bigger model would do better. Store your Anthropic key the same way as in step 1, then:

Paste into Claude CodeScore [the rubric part Jev missed] on the same examples with Sonnet 5.5, using the exact same instructions and levels Jev got. Keep my scores out of the prompt, read the Anthropic key from [your password manager], and ask for a short reason with each score. Then compare me, Jev, and Sonnet side by side, with time and cost.

Sonnet agreed with Jev, not me: the same score on 15 of 20 and within a point on all 20.

6

When the models agree, fix your rubric

Read the reasons Sonnet gave on the examples where you scored highest. Then tell Claude Code, in your own words, what you were really rewarding. Mine:

Paste into Claude CodeRepeated information counts, repeated feeling doesn't. Warmth for a frustrated customer earns its place. Missing information is not a conciseness problem. Add those rules to the question and rerun Jev and Sonnet.

With those rules, Jev matched me within a point on 14 of 15. Sonnet stayed at 10, at about 80 times the cost per score.

Then stop tuning. I added a fourth rule and Jev got worse, so I took it out. Rules written from the same few examples make the score look better without proving anything.

7

Test on fresh drafts, graded blind

The real test is outputs Jev has never seen. Ask Claude Code:

Paste into Claude CodeGenerate 15 fresh drafts from my current [skill or system prompt] for [real or sample requests], and give me a grading sheet with only the request and the draft. Once I've filled in my scores, run Jev on them using my current [skill] as the voice guide, and compare.

Use your current instructions as the reference. Mine first graded against last year's voice guide, and brand voice matched me on only 1 of 15. With my current skill, it landed within a point on 12 to 15 of 15 for every rubric part.

8

Use confidence to choose what you review

Every Jev score comes with a confidence number. Ask Claude Code: "Would reviewing only Jev's lowest-confidence scores have caught its misses?" On my fresh set, the lowest-confidence quarter caught 4 of its 6 big misses. That is a small sample, so check it on your own data before relying on it.

What it cost

Jev charges for input only, and every Jev run in my test cost about half a cent. All the Jev runs together came to a few cents. The Sonnet head-to-head was about 6 to 7 cents per run, for one rubric part.

Where Jev fits, and where it doesn't

Good fit: high-volume, repeatable judgments, like scoring every draft on tone or flagging anything that breaks a policy.

Keep for yourself or a bigger model: the subtle calls. Jev missed my "be more casual with someone I work with every week" note. Its own docs list math and dates as weak spots, so keep those in code.

If you use Claude Code's plugin evals: they cannot call an outside grader directly, so run Jev on their results as a step afterward.

Want help setting up evals for your own AI work?

We can build your first rubric and answer key together in one hour.

Get Started with AI ($150)