
TLDR
- I tested Jev (TypeSafe AI) as a PII guardrail on 2,000 texts from a public dataset. Jev is a small model that returns a yes/no decision plus a confidence score instead of writing text.
- I compared it to five other cheap, fast models.
- Jev ties Gemini 3.7 Flash and Claude Haiku 4.5 on accuracy. It does that at about 1/25th to 1/33rd of the cost and is faster.
- Its confidence scores are meaningful: when it says “not sure,” it really is less accurate. That was the claim I wanted to check.
- I still wouldn’t send real PII to an external service to find out whether it’s PII. More on that below.
The PII guardrail claim I wanted to test
Jev is pitched as a cheap, fast “traffic cop” for AI agent pipelines. It doesn’t generate anything. It answers a yes/no question and says how sure it is.
Cheap and fast are easy to check. The word I cared about was calibrated. A calibrated model that says “90% sure” is right about 90% of the time. That’s what lets you act on the number, for example “send the unsure ones to a person.”
I’m not affiliated with TypeSafe. This is an outside check of their own claim, and that’s the point of the post.
The test
- Task: does this text contain personal information (names, emails, phone numbers, addresses, ID numbers and so on)? This is a typical PII guardrail check. This is related to my recent post on PolicyPal on judgment and confidence.
- Data: 2,000 texts from
ai4privacy/pii-masking-300kon Hugging Face. 1,000 contain PII and 1,000 don’t, across 6 languages. - Models: Jev, Gemini 3.7 Flash, Claude Haiku 4.5, DeepSeek v4.1 Flash, Qwen 3.8 Flash, GPT-5.6 Luna.
- The number I care about most: false alarms. Of the 1,000 clean texts, how many did the model wrongly flag? For a guardrail, each one is a legitimate user getting blocked.
Attempt 1: send everything through one gateway
I planned to run all six models through Tailscale Aperture, a gateway that gives you one place for logging and access control.
It worked for five. Jev has its own API and isn’t an Aperture option, so I called it directly.
Dev Note
This looks like a small detail, but it isn’t. Calls to Jev skip the gateway’s logging and access control. For a prototype, that’s a footnote. For a team adopting it, it’s an architecture decision. It also means Jev’s speed wasn’t measured on the same path as the others.
Attempt 2: read the leaderboard
Here are the results. Lower is better in every column except accuracy.
| model | accuracy | false alarms (of 1,000 clean) | misses (of 1,000 with PII) | calibration error | cost per 1,000 calls | seconds per call |
|---|---|---|---|---|---|---|
| Jev | 94.5% | 83 | 26 | 0.020 | $0.021 | 0.20 |
| Gemini 3.7 Flash | 94.5% | 78 | 33 | 0.047 | $0.687 | 2.27 |
| Claude Haiku 4.5 | 93.4% | 93 | 39 | 0.016 | $0.525 | 1.18 |
| Qwen 3.8 Flash | 93.8% | 96 | 29 | 0.041 | $0.051 | 2.12 |
| DeepSeek v4.1 Flash | 93.2% | 93 | 43 | 0.050 | $0.050 | 1.63 |
| GPT-5.6 Luna | 92.8% | 122 | 22 | 0.067 | $0.080 | 0.86 |
Calibration error is the average gap between how sure the model says it is and how often it’s actually right. Zero is perfect. Jev’s cost is calculated from its list price ($0.042 per million input tokens), not from a bill.
Jev looks great as a PII guardrail: tied for the best accuracy, a low calibration error, and the lowest cost. I nearly wrote “Jev wins.”
That would have been wrong. Here’s why.
Here’s the same result as a picture. Each marker is a model. Further left means fewer false alarms, and higher means fewer misses. The cross through each marker is the range the true number could plausibly fall in, given only 2,000 texts.

Where Jev does stand apart is cost.

What worked: asking “could this be luck?”
Gemini and Jev differ by about one point of accuracy. With only 2,000 texts, differences that small can easily be luck. So I ran a significance test. In plain terms, it asks: if these two models were really equally good, how often would I see a gap this big by chance?
How the test works (McNemar’s test):
- Most of the time, two models agree. Those texts tell you nothing, so the test ignores them.
- It only looks at the texts where one model was right and the other was wrong.
- If the models are equally good, those disagreements should split roughly 50/50, like a coin.
Two examples:
- Jev vs Gemini: Jev was right and Gemini wrong 26 times. Gemini was right and Jev wrong 24 times. That’s a coin flip. They tie.
- Jev vs Luna: Jev was right and Luna wrong 53 times. Luna was right and Jev wrong only 18 times. That’s very unlikely to be luck. Jev is really better here.
What a p-value is: the chance of seeing a split this lopsided if the models were actually equal. Small is good. Jev vs Luna is p=0.0006, which is about 6 in 10,000. By convention, below 0.05 counts as “probably real.”
The catch: multiple comparisons. With 6 models there are 15 pairs. If you run 15 tests, a few will look significant by pure luck, like flipping a coin 15 times and being surprised by a streak. So I used a standard correction (Holm-Bonferroni) that raises the bar for each test.
It changed the answer. Jev vs Haiku looked significant at first (p=0.015, with Jev winning 53 to 30). After the correction it became p=0.17, which isn’t convincing. Jev vs Qwen went from 0.023 to 0.23. That’s the gap between “Jev beats Haiku” and “can’t tell them apart.”
Dev Note
This is the step I’d have skipped if I were in a hurry. The raw table says one thing and the corrected tests say another. I think most model comparisons published without a test like this are reading noise.
What survives the correction (accuracy):
- 12 of 15 pairs are indistinguishable, including Jev vs Gemini, Haiku and Qwen.
- Jev beats Luna and DeepSeek (p=0.0006 and 0.0095).
- Gemini beats Luna (p=0.002).
Calibration, with the same kind of test. I compared each text’s confidence error (the Brier score, which punishes a model more for being confidently wrong than for being hesitantly wrong):
- Jev ties Gemini and Haiku.
- Jev beats DeepSeek and Qwen (p=0.0001 and 0.0008).
- Luna is worse than Haiku, Gemini and Qwen.
So the claim I can defend is: Jev ties the best models here on accuracy and calibration as a PII guardrail, at a fraction of the cost and latency. It does not beat Gemini or Haiku.
Is the confidence number real?
This was the actual question. Here’s how often Jev is right at each confidence level:
| how sure Jev says it is | texts | how often it’s right |
|---|---|---|
| under 90% | 214 | 76% |
| 90–95% | 151 | 90% |
| 95%+ | 1,635 | 97% |
The scores aren’t random. The more confident Jev is, the more often it’s right.
Here are all six models. The dashed diagonal is perfect calibration: a model that says 80% is right 80% of the time. Dot size shows how many texts fell in that group, and hollow dots mean fewer than 10 texts, so don’t read much into them.

That makes a simple routing rule work. Send everything under 90% confidence to a person or a slower model. That’s about 11% of traffic, and it catches 48% of Jev’s mistakes.
The other models are less useful for this:
- Gemini and Luna say “99% sure” about almost everything. At the same cutoff they flag only 2% of traffic and catch 16% and 10% of their mistakes. Luna is only 94.5% right when it says 99%+.
- Haiku’s confidence is useful (it catches 77% of its mistakes), but the same cutoff sends 27% of traffic to review.
Two honest caveats:
- 82% of Jev’s predictions fall in the top group, so its good score is mostly that group. One smaller group (70–75% confidence, 29 texts) was only 59% right.
- Even when Jev is 99%+ sure, it’s wrong 2.3% of the time. Calibrated doesn’t mean perfect.
Where I tried to break it
Do you even need a model? I wrote a simple regex function (no AI) for the easy cases: emails, phone numbers, IP addresses, coordinates and long ID numbers.
- It missed 240 of the 1,000 PII texts (76% caught).
- It only had 36 false alarms per 1,000, fewer than every model.
- It costs about $0.0003 per 1,000 calls and responds in about 71 ms.
It isn’t a replacement, but it’s a useful floor. If most of what you catch is structured, regex does a lot.
Does the wording of the question matter? I asked Jev the same task five ways, including asking the same prompt twice to see how much results move on their own.
| prompt wording | false alarms (of 1,000) | misses (of 1,000) |
|---|---|---|
| A: “PII” plus a list of examples | 87 | 28 |
| A again (measures run-to-run noise) | 84 | 27 |
| B: same, phrased “for example” | 88 | 24 |
| C: “PII” definition, no list | 78 | 54 |
| D: list only, never says “PII” | 120 | 27 |
- Running the same prompt twice moved false alarms by 3. That’s the noise.
- Dropping the word “PII” (D) added about 36 false alarms.
- Dropping the list (C) doubled the misses.
The wording is a real tuning knob. Test it.
Why I wouldn’t use this PII guardrail as a security control
I wouldn’t use this for real PII detection as a security control. Two reasons:
Dev Note
The test data is synthetic and template-like, so real traffic will probably score lower. That’s the smaller problem.
- It’s an outside service. To find out whether text contains PII, you send the text to a third party. I haven’t verified TypeSafe’s data handling. For a compliance control, accuracy isn’t enough. You also need to be able to show where the data goes. That’s harder with an external service than with a function you own or a model you host. It’s also where the gateway skip hurts, since those calls aren’t logged.
- No reasoning. You get a decision and a number, not an explanation. That’s thin for an audit trail.
So the confidence number from a PII guardrail like this is better used for routing: the unsure slice goes to a person, or to a slower, more expensive model that can explain itself.
What this means if you run an engineering team
- Run the cheap decision first. A fast yes/no with a confidence score goes in front. Only the unsure cases pay for the slower path. PII is one example of a pattern I think of as LLM-lambda decisions as a service. I tested one case, so I’m not claiming the rest.
- When accuracy ties, cost and speed decide. About 25x cheaper than Haiku, 33x cheaper than Gemini, and roughly 4–11x faster.
- Check that the confidence means something before you route on it. Two of these models would have given me a router that almost never fires.
- Count false alarms for a PII guardrail. They ranged from 78 to 122 per 1,000 clean texts, and those are your blocked users.
- Test before you trust a leaderboard. Twelve of 15 comparisons were noise.
- Try the dumb baseline first. Regex beat every model on false alarms.
- Treat data handling as a separate decision. Accuracy doesn’t settle it.
Thanks to the ai4privacy team for the dataset, Tailscale for Aperture, and TypeSafe for an API that was easy to test.
If you’ve put a cheap classifier in front of an expensive pipeline, what did you route on, and how did you decide where to draw the line?

Leave a Reply