Definition
The Jev-as-a-judge test scores every row against criteria you write in plain language. In development that is the validation set. In monitoring it is the production window. A calibrated classifier returns a verdict for each row, so the test can cover the whole set instead of a sample. LLM-as-a-judge sends a sample of rows to a generative model and asks it to grade them and explain the grade. The project sample is often about 50 rows. Jev-as-a-judge judges every row unless you opt into a sample, and it returns a score rather than a written explanation.Taxonomy
- Task types: LLM.
- Availability: and .
Why it matters
- A sample can miss a failure that shows up on only a few rows. Judging every row means those failures count.
- Several criteria share one request per row, so adding a criterion costs little beyond the first.
- You choose whether the test watches the share of rows that pass, or the average verdict. On a range criterion those two numbers tell different stories.
When to use it
Use Jev-as-a-judge when you want a verdict on every row and a generative judge would be too expensive to run that way. Use LLM-as-a-judge when you need a written rationale, or when the judge has to be a specific LLM provider and model you already connect.Criteria
A test takes one to five criteria. Each one needscriteria: what the judge
should decide about the row, up to 8192 characters.
You can also set:
name— the display name, up to 50 characters.key— a stable id, up to 50 characters. Scores stay attached to this id, so you can rename the criterion without dropping verdicts already computed for it. Keys on one test must be unique. Set one when you expect to rename or reorder criteria. The measurement name (criteria0PassRate,criteria1MeanScore, and so on) follows list order, notkey.kind—noul(the default) orscore.threshold— the value a row must reach to count as passing. The default is0.5. For a binary criterion this is the probability of a1. For a range criterion it is a position on that criterion’s scale.passesAbove— which side of the threshold passes. The default istrue. Set it tofalsewhen a higher verdict is worse, such as a severity rubric that should pass below its threshold.
noul) scores each row 0 or 1. Write criteria so the judge
knows what each level means, in the same rubric style as
LLM-as-a-judge. Name both scores in
the text, for example 1 if the answer is fully supported by the retrieved context, 0 if it is not, or 1 if profanity is flagged, 0 if no profanity.
The create form placeholder uses the same shape: describe what the judge
should decide, for example 1 if the condition is met, 0 if not.
After judging, the row is 1 when the probability of a 1 is on the passing
side of threshold, and 0 otherwise. With the defaults, a probability of
0.5 or higher scores 1.
Range (score) rates the row on a numeric scale and returns that score.
Set rangeMin and rangeMax as inclusive integers. The scale must cover 2 to
10 whole numbers. The default is 0 to 1. The numerals do not mean anything
on their own, so say what the ends of the scale mean in criteria. A range
verdict can land between the ends. A binary verdict is only 0 or 1.
The create form writes a single criterion. Binary is tagged 0, 1 and
Range is tagged 0–1. Wider scales, a custom row threshold, and additional
criteria are set in tests.json.
Pass rate and mean
Each criterion publishes two measurements.N is its position in the criteria
list, starting at 0 (criteria0 … criteria4):
criteriaNPassRate— Pass rate, the fraction of judged rows whose verdict cleared that criterion’s threshold. In the create form this threshold, and the mean, accept a value from 0 to 1.criteriaNMeanScore— the average of the raw verdicts.
0 or 1, so the mean and the pass
rate are the same number.
On a range criterion they are not. Ninety-five rows at 1.0 and five rows at
0.0 have the same mean (0.95) as one hundred rows at 0.95. If the row
threshold is 0.5, the first window has five failures and the second has
none. The pass rate shows that. The mean does not. Use the pass rate when any
failed row should count, even if the rest of the window scores highly. Use the
mean when you care how far the typical verdict sits on the scale.
For a pass rate, the test threshold value is a fraction between 0 and 1. For
a mean, value is a number on that criterion’s own scale.
Rows the judge declines are left out of both measurements. They are not
counted as failures.
API key
Jev-as-a-judge readsTYPESAFE_API_KEY from
environment variables. Openlayer resolves a
key in this order: the project, then the workspace, then a platform key.
A platform key still runs the test. Those rows are judged through Openlayer’s
TypeSafe account. Set your own key when the data and the spend need to stay on
your account, including data-residency and bring-your-own-key requirements.
If no key resolves, the test is skipped rather than run. The create form shows
which case you are in. With no key, you can still create the test, and it is
skipped on every run until you add TYPESAFE_API_KEY. With only the platform
key, the form tells you the rows are judged through Openlayer’s TypeSafe
account.
Judge and sampling
classifier_evaluator chooses the judge, and how much of the data it sees.
Omit it to judge every row with the default model.
provider—typesafe. This is the only provider.model—jev-latest(the default) orjev-preview.jev-latesttracks the current stable release.jev-previewis the next release, ahead of general availability. Any other id is rejected, including a partial version such asjev-1.13.sample— omit it to judge every row. Set it only to trade coverage for cost.fraction— the share of rows to judge, from0to1. The default is1. The same rows stay in the sample if a run is retried or a monitoring window resumes.limit— a cap on rows judged for this test, counted across resumed runs rather than once per run. Omit it for no cap.
Test configuration examples
If you are writing atests.json, here are valid configurations for the
Jev-as-a-judge test. The development example judges every row. The monitoring
example judges a quarter of the window, capped at 10,000 rows.
Related
- LLM-as-a-judge — grade a sample of rows with a generative model and a written explanation.
- Environment variables — set
TYPESAFE_API_KEYon the project or the workspace.